EDBT 2026 Demo / reviewers in the wild / expert
Christoph Rhemann
dblp:01/4107
· DBLP profile ↗
42ranked-venue papers
5as first author
1since 2021 · last 2021
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 34 · 5 first-author · 1 since 2021Artificial intelligence and machine learning · 21 · 5 first-authorHuman-computer interaction and ubiquitous computing · 6Systems, architecture and hardware · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
19 papers |
3D vision · 90% Video understanding and tracking · 6% Segmentation and scene understanding · 4% | |
| Computer graphics and multimedia
22 papers |
Image and video processing · 27% Rendering · 26% Computational photography and imaging · 22% | |
| Human-computer interaction and pervasive computing
6 papers |
Interaction techniques and input · 63% Immersive interaction · 18% Collaborative and social computing · 12% |
Topics — the 30 heaviest of 80, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › 3D vision › stereo vision
stereo matching |
1.7 | 7 | 2018 | ActiveStereoNet: End-to-End Self-supervised Learning for Active Stereo Systems · ECCV (8) 2018 StereoNet: Guided Hierarchical Refinement for Real-Time Edge-Aware Depth Prediction · ECCV (15) 2018 Low Compute and Fully Parallel Computer Vision with HashMatch · ICCV 2017 |
Computer vision › 3D vision
depth estimation |
1.3 | 5 | 2018 | The need 4 speed in real-time dense visual tracking · ACM Trans. Graph. 2018 StereoNet: Guided Hierarchical Refinement for Real-Time Edge-Aware Depth Prediction · ECCV (15) 2018 UltraStereo: Efficient Learning-Based Matching for Active Stereo Systems · CVPR 2017 |
Computer vision › 3D vision
3d reconstruction |
1.0 | 5 | 2018 | Fusion4D: real-time performance capture of challenging scenes · ACM Trans. Graph. 2016 Holoportation: Virtual 3D Teleportation in Real-time · UIST 2016 Real-time non-rigid reconstruction using an RGB-D camera · ACM Trans. Graph. 2014 |
Rendering
neural rendering |
0.9 | 3 | 2020 | Deep relightable textures: volumetric performance capture with neural rendering · ACM Trans. Graph. 2020 LookinGood: enhancing performance capture with real-time neural re-rendering · ACM Trans. Graph. 2018 Light stage super-resolution: continuous high-frequency relighting · ACM Trans. Graph. 2020 |
Rendering › relighting
portrait relighting |
0.9 | 2 | 2021 | Total relighting: learning to relight portraits for background replacement · ACM Trans. Graph. 2021 Single image portrait relighting · ACM Trans. Graph. 2019 |
Computer vision › 3D vision › stereo vision
active stereo |
0.6 | 2 | 2018 | ActiveStereoNet: End-to-End Self-supervised Learning for Active Stereo Systems · ECCV (8) 2018 UltraStereo: Efficient Learning-Based Matching for Active Stereo Systems · CVPR 2017 |
Image and video processing
image matting |
0.6 | 6 | 2011 | A global sampling method for alpha matting · CVPR 2011 A spatially varying PSF-based prior for alpha matting · CVPR 2010 New appearance models for natural image matting · CVPR 2009 |
Computer vision › 3D vision › motion capture
human performance capture |
0.5 | 2 | 2017 | Motion2fusion: real-time volumetric performance capture · ACM Trans. Graph. 2017 Fusion4D: real-time performance capture of challenging scenes · ACM Trans. Graph. 2016 |
Visual content generation and editing › image editing › image compositing
background substitution |
0.5 | 1 | 2021 | Total relighting: learning to relight portraits for background replacement · ACM Trans. Graph. 2021 |
Computer vision › 3D vision › 3d reconstruction
non-rigid reconstruction |
0.4 | 2 | 2016 | Fusion4D: real-time performance capture of challenging scenes · ACM Trans. Graph. 2016 Real-time non-rigid reconstruction using an RGB-D camera · ACM Trans. Graph. 2014 |
Image and video processing
super-resolution |
0.4 | 1 | 2020 | Light stage super-resolution: continuous high-frequency relighting · ACM Trans. Graph. 2020 |
Image and video processing › image matting
alpha matting |
0.4 | 4 | 2011 | A global sampling method for alpha matting · CVPR 2011 A spatially varying PSF-based prior for alpha matting · CVPR 2010 A stereo approach that handles the matting problem via image warping · CVPR 2009 |
Computational photography and imaging
image relighting |
0.4 | 1 | 2019 | Single image portrait relighting · ACM Trans. Graph. 2019 |
Computational photography and imaging › 3d scanning
volumetric performance capture |
0.4 | 1 | 2019 | The relightables: volumetric performance capture of humans with realistic relighting · ACM Trans. Graph. 2019 |
Virtual and augmented reality
telepresence |
0.3 | 2 | 2018 | Holoportation: Virtual 3D Teleportation in Real-time · UIST 2016 LookinGood: enhancing performance capture with real-time neural re-rendering · ACM Trans. Graph. 2018 |
Computer vision › Video understanding and tracking › motion tracking
dense tracking |
0.3 | 1 | 2018 | The need 4 speed in real-time dense visual tracking · ACM Trans. Graph. 2018 |
Computer vision › 3D vision › depth estimation
monocular depth estimation |
0.3 | 1 | 2018 | Depth from motion for smartphone AR · ACM Trans. Graph. 2018 |
Computer vision › Video understanding and tracking
object tracking |
0.3 | 1 | 2018 | The need 4 speed in real-time dense visual tracking · ACM Trans. Graph. 2018 |
Computer vision › 3D vision › stereo vision › stereo matching › deep stereo matching
self-supervised stereo matching |
0.3 | 1 | 2018 | ActiveStereoNet: End-to-End Self-supervised Learning for Active Stereo Systems · ECCV (8) 2018 |
Virtual and augmented reality › tracking
6DOF tracking |
0.3 | 1 | 2018 | Depth from motion for smartphone AR · ACM Trans. Graph. 2018 |
Image and video processing
image restoration |
0.3 | 1 | 2018 | LookinGood: enhancing performance capture with real-time neural re-rendering · ACM Trans. Graph. 2018 |
Virtual and augmented reality › augmented reality
mobile augmented reality |
0.3 | 1 | 2018 | Depth from motion for smartphone AR · ACM Trans. Graph. 2018 |
Image and video processing › image restoration › multi-task image restoration
super-resolution and denoising |
0.3 | 1 | 2018 | LookinGood: enhancing performance capture with real-time neural re-rendering · ACM Trans. Graph. 2018 |
Computer vision › 3D vision › geometric estimation › registration
non-rigid registration |
0.3 | 1 | 2017 | Motion2fusion: real-time volumetric performance capture · ACM Trans. Graph. 2017 |
Computer vision › Segmentation and scene understanding › dense prediction
pixel labeling |
0.3 | 1 | 2017 | Low Compute and Fully Parallel Computer Vision with HashMatch · ICCV 2017 |
Computer vision › 3D vision › depth estimation
stereo depth estimation |
0.3 | 1 | 2017 | Low Compute and Fully Parallel Computer Vision with HashMatch · ICCV 2017 |
Computer vision › 3D vision › 3d reconstruction
volumetric reconstruction |
0.3 | 1 | 2017 | Motion2fusion: real-time volumetric performance capture · ACM Trans. Graph. 2017 |
Multimedia analysis and retrieval
image retrieval |
0.3 | 1 | 2017 | Low Compute and Fully Parallel Computer Vision with HashMatch · ICCV 2017 |
Computer vision › 3D vision
correspondence estimation |
0.2 | 1 | 2016 | The Global Patch Collider · CVPR 2016 |
Computer vision › 3D vision › feature matching
correspondence problem |
0.2 | 1 | 2016 | HyperDepth: Learning Depth from Structured Light without Matching · CVPR 2016 |
Methods — techniques the papers use, named apart from their topics
neural network · 0.8deep neural network · 0.8color gradient illumination · 0.8per-pixel lighting representation · 0.5deep learning · 0.5alpha matting · 0.5patchmatch · 0.5regularized interpolation · 0.4geometric reconstruction pipeline · 0.4machine learning reconstruction pipeline · 0.4active depth sensing · 0.4visual-inertial odometry · 0.3space-time feature matching · 0.3self-supervised learning · 0.3monocular depth computation · 0.3machine learning depth refinement · 0.3hierarchical refinement · 0.3edge-aware filtering · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | Total relighting: learning to relight portraits for background replacementabstractWe propose a novel system for portrait relighting and background replacement, which maintains high-frequency boundary details and accurately synthesizes the subject's appearance as lit by novel illumination, thereby producing realistic composite images for any desired scene. Our technique includes foreground estimation via alpha matting, relighting, and compositing. We demonstrate that each of these stages can be tackled in a sequential pipeline without the use of priors (e.g. known background or known illumination) and with no specialized acquisition techniques, using only a single RGB portrait image and a novel, target HDR lighting environment as inputs. We train our model using relit portraits of subjects captured in a light stage computational illumination system, which records multiple lighting conditions, high quality geometry, and accurate alpha mattes. To perform realistic relighting for compositing, we introduce a novel per-pixel lighting representation in a deep learning framework, which explicitly models the diffuse and the specular components of appearance, producing relit portraits with convincingly rendered non-Lambertian effects like specular highlights. Multiple experiments and comparisons show the effectiveness of the proposed approach when applied to in-the-wild images. Rohit Pandey, Sergio Orts, Chloe LeGendre, Christian Häne, Sofien Bouaziz, Christoph Rhemann, Paul E. Debevec, Sean Ryan Fanello |
ACM Trans. Graph. | 6 |
| 2020 | Deep relightable textures: volumetric performance capture with neural renderingabstractThe increasing demand for 3D content in augmented and virtual reality has motivated the development of volumetric performance capture systemsnsuch as the Light Stage. Recent advances are pushing free viewpoint relightable videos of dynamic human performances closer to photorealistic quality. However, despite significant efforts, these sophisticated systems are limited by reconstruction and rendering algorithms which do not fully model complex 3D structures and higher order light transport effects such as global illumination and sub-surface scattering. In this paper, we propose a system that combines traditional geometric pipelines with a neural rendering scheme to generate photorealistic renderings of dynamic performances under desired viewpoint and lighting. Our system leverages deep neural networks that model the classical rendering process to learn implicit features that represent the view-dependent appearance of the subject independent of the geometry layout, allowing for generalization to unseen subject poses and even novel subject identity. Detailed experiments and comparisons demonstrate the efficacy and versatility of our method to generate high-quality results, significantly outperforming the existing state-of-the-art solutions. Abhimitra Meka, Rohit Pandey, Christian Häne, Sergio Orts, Peter Barnum, Philip L. Davidson, Daniel Erickson, Yinda Zhang 0001, Jonathan Taylor 0001, Sofien Bouaziz, Chloe LeGendre, Wan-Chun Ma, Ryan S. Overbeck, Thabo Beeler, Paul E. Debevec, Shahram Izadi, Christian Theobalt, Christoph Rhemann, Sean Ryan Fanello |
ACM Trans. Graph. | 18 |
| 2020 | Light stage super-resolution: continuous high-frequency relightingabstractThe light stage has been widely used in computer graphics for the past two decades, primarily to enable the relighting of human faces. By capturing the appearance of the human subject under different light sources, one obtains the light transport matrix of that subject, which enables image-based relighting in novel environments. However, due to the finite number of lights in the stage, the light transport matrix only represents a sparse sampling on the entire sphere. As a consequence, relighting the subject with a point light or a directional source that does not coincide exactly with one of the lights in the stage requires interpolation and resampling the images corresponding to nearby lights, and this leads to ghosting shadows, aliased specularities, and other artifacts. To ameliorate these artifacts and produce better results under arbitrary high-frequency lighting, this paper proposes a learning-based solution for the "super-resolution" of scans of human faces taken from a light stage. Given an arbitrary "query" light direction, our method aggregates the captured images corresponding to neighboring lights in the stage, and uses a neural network to synthesize a rendering of the face that appears to be illuminated by a "virtual" light source at the query location. This neural network must circumvent the inherent aliasing and regularity of the light stage data that was used for training, which we accomplish through the use of regularized traditional interpolation methods within our network. Our learned model is able to produce renderings for arbitrary light directions that exhibit realistic shadows and specular highlights, and is able to generalize across a wide variety of subjects. Our super-resolution approach enables more accurate renderings of human subjects under detailed environment maps, or the construction of simpler light stages that contain fewer light sources while still yielding comparable quality renderings as light stages with more densely sampled lights. Tiancheng Sun, Zexiang Xu, Xiuming Zhang, Sean Ryan Fanello, Christoph Rhemann, Paul E. Debevec, Yun-Ta Tsai, Jonathan T. Barron, Ravi Ramamoorthi |
ACM Trans. Graph. | 5 |
| 2019 | The relightables: volumetric performance capture of humans with realistic relightingabstractWe present "The Relightables", a volumetric capture system for photorealistic and high quality relightable full-body performance capture. While significant progress has been made on volumetric capture systems, focusing on 3D geometric reconstruction with high resolution textures, much less work has been done to recover photometric properties needed for relighting. Results from such systems lack high-frequency details and the subject's shading is prebaked into the texture. In contrast, a large body of work has addressed relightable acquisition for image-based approaches, which photograph the subject under a set of basis lighting conditions and recombine the images to show the subject as they would appear in a target lighting environment. However, to date, these approaches have not been adapted for use in the context of a high-resolution volumetric capture system. Our method combines this ability to realistically relight humans for arbitrary environments, with the benefits of free-viewpoint volumetric capture and new levels of geometric accuracy for dynamic performances. Our subjects are recorded inside a custom geodesic sphere outfitted with 331 custom color LED lights, an array of high-resolution cameras, and a set of custom high-resolution depth sensors. Our system innovates in multiple areas: First, we designed a novel active depth sensor to capture 12.4 MP depth maps, which we describe in detail. Second, we show how to design a hybrid geometric and machine learning reconstruction pipeline to process the high resolution input and output a volumetric video. Third, we generate temporally consistent reflectance maps for dynamic performers by leveraging the information contained in two alternating color gradient illumination images acquired at 60Hz. Multiple experiments, comparisons, and applications show that The Relightables significantly improves upon the level of realism in placing volumetrically captured human performances into arbitrary CG scenes. Peter Lincoln, Philip Davidson, Jay Busch, Xueming Yu, Matt Whalen, Geoff Harvey, Sergio Orts, Rohit Pandey, Jason Dourgarian, Danhang Tang, Anastasia Tkach, Adarsh Kowdle, Emily Cooper, Mingsong Dou, Sean Ryan Fanello, Graham Fyffe, Christoph Rhemann, Jonathan Taylor 0001, Paul E. Debevec, Shahram Izadi |
ACM Trans. Graph. | 18 |
| 2019 | Deep reflectance fields: high-quality facial reflectance field inference from color gradient illuminationabstractWe present a novel technique to relight images of human faces by learning a model of facial reflectance from a database of 4D reflectance field data of several subjects in a variety of expressions and viewpoints. Using our learned model, a face can be relit in arbitrary illumination environments using only two original images recorded under spherical color gradient illumination. The output of our deep network indicates that the color gradient images contain the information needed to estimate the full 4D reflectance field, including specular reflections and high frequency details. While capturing spherical color gradient illumination still requires a special lighting setup, reduction to just two illumination conditions allows the technique to be applied to dynamic facial performance capture. We show side-by-side comparisons which demonstrate that the proposed system outperforms the state-of-the-art techniques in both realism and speed. Abhimitra Meka, Christian Häne, Rohit Pandey, Michael Zollhöfer, Sean Ryan Fanello, Graham Fyffe, Adarsh Kowdle, Xueming Yu, Jay Busch, Jason Dourgarian, Peter Denny, Sofien Bouaziz, Peter Lincoln, Matt Whalen, Geoff Harvey, Jonathan Taylor 0001, Shahram Izadi, Andrea Tagliasacchi, Paul E. Debevec, Christian Theobalt, Julien P. C. Valentin, Christoph Rhemann |
ACM Trans. Graph. | 22 |
| 2019 | Single image portrait relightingabstractLighting plays a central role in conveying the essence and depth of the subject in a portrait photograph. Professional photographers will carefully control the lighting in their studio to manipulate the appearance of their subject, while consumer photographers are usually constrained to the illumination of their environment. Though prior works have explored techniques for relighting an image, their utility is usually limited due to requirements of specialized hardware, multiple images of the subject under controlled or known illuminations, or accurate models of geometry and reflectance. To this end, we present a system for portrait relighting : a neural network that takes as input a single RGB image of a portrait taken with a standard cellphone camera in an unconstrained environment, and from that image produces a relit image of that subject as though it were illuminated according to any provided environment map. Our method is trained on a small database of 18 individuals captured under different directional light sources in a controlled light stage setup consisting of a densely sampled sphere of lights. Our proposed technique produces quantitatively superior results on our dataset's validation set compared to prior works, and produces convincing qualitative relighting results on a dataset of hundreds of real-world cellphone portraits. Because our technique can produce a 640 × 640 image in only 160 milliseconds, it may enable interactive user-facing photographic applications in the future. Tiancheng Sun, Jonathan T. Barron, Yun-Ta Tsai, Zexiang Xu, Xueming Yu, Graham Fyffe, Christoph Rhemann, Jay Busch, Paul E. Debevec, Ravi Ramamoorthi |
ACM Trans. Graph. | 7 |
| 2018 | StereoNet: Guided Hierarchical Refinement for Real-Time Edge-Aware Depth Prediction
Sameh Khamis, Sean Ryan Fanello, Christoph Rhemann, Adarsh Kowdle, Julien P. C. Valentin, Shahram Izadi |
ECCV (15) | 3 |
| 2018 | ActiveStereoNet: End-to-End Self-supervised Learning for Active Stereo Systems
Yinda Zhang 0001, Sameh Khamis, Christoph Rhemann, Julien P. C. Valentin, Adarsh Kowdle, Vladimir Tankovich, Michael Schoenberg, Shahram Izadi, Thomas A. Funkhouser, Sean Ryan Fanello |
ECCV (8) | 3 |
| 2018 | SOS: Stereo Matching in O(1) with Slanted Support WindowsabstractDepth cameras have accelerated research in many areas of computer vision. Most triangulation-based depth cameras, whether structured light systems like the Kinect or active (assisted) stereo systems, are based on the principle of stereo matching. Depth from stereo is an active research topic dating back 30 years. Despite recent advances, algorithms usually trade-off accuracy for speed. In particular, efficient methods rely on fronto-parallel assumptions to reduce the search space and keep computation low. We present SOS (Slanted O(1) Stereo), the first algorithm capable of leveraging slanted support windows without sacrificing speed or accuracy. We use an active stereo configuration, where an illuminator textures the scene. Under this setting, local methods - such as PatchMatch Stereo - obtain state of the art results by jointly estimating disparities and slant, but at a large computational cost. We observe that these methods typically exploit local smoothness to simplify their initialization strategies. Our key insight is that local smoothness can in fact be used to amortize the computation not only within initialization, but across the entire stereo pipeline. Building on these insights, we propose a novel hierarchical initialization that is able to efficiently perform search over disparity and slants. We then show how this structure can be leveraged to provide high quality depth maps. Extensive quantitative evaluations demonstrate that the proposed technique yields significantly more precise results than current state of the art, but at a fraction of the computational cost. Our prototype implementation runs at 4000 fps on modern GPU architectures. Vladimir Tankovich, Michael Schoenberg, Sean Ryan Fanello, Adarsh Kowdle, Christoph Rhemann, Maksym Dzitsiuk, Mirko Schmidt, Julien P. C. Valentin, Shahram Izadi |
IROS | 5 |
| 2018 | The need 4 speed in real-time dense visual trackingabstractThe advent of consumer depth cameras has incited the development of a new cohort of algorithms tackling challenging computer vision problems. The primary reason is that depth provides direct geometric information that is largely invariant to texture and illumination. As such, substantial progress has been made in human and object pose estimation, 3D reconstruction and simultaneous localization and mapping. Most of these algorithms naturally benefit from the ability to accurately track the pose of an object or scene of interest from one frame to the next. However, commercially available depth sensors (typically running at 30fps) can allow for large inter-frame motions to occur that make such tracking problematic. A high frame rate depth camera would thus greatly ameliorate these issues, and further increase the tractability of these computer vision problems. Nonetheless, the depth accuracy of recent systems for high-speed depth estimation [Fanello et al. 2017b] can degrade at high frame rates. This is because the active illumination employed produces a low SNR and thus a high exposure time is required to obtain a dense accurate depth image. Furthermore in the presence of rapid motion, longer exposure times produce artifacts due to motion blur, and necessitates a lower frame rate that introduces large inter-frame motion that often yield tracking failures. In contrast, this paper proposes a novel combination of hardware and software components that avoids the need to compromise between a dense accurate depth map and a high frame rate. We document the creation of a full 3D capture system for high speed and quality depth estimation, and demonstrate its advantages in a variety of tracking and reconstruction tasks. We extend the state of the art active stereo algorithm presented in Fanello et al. [2017b] by adding a space-time feature in the matching phase. We also propose a machine learning based depth refinement step that is an order of magnitude faster than traditional postprocessing methods. We quantitatively and qualitatively demonstrate the benefits of the proposed algorithms in the acquisition of geometry in motion. Our pipeline executes in 1.1ms leveraging modern GPUs and off-the-shelf cameras and illumination components. We show how the sensor can be employed in many different applications, from [non-]rigid reconstructions to hand/face tracking. Further, we show many advantages over existing state of the art depth camera technologies beyond framerate, including latency, motion artifacts, multi-path errors, and multi-sensor interference. Adarsh Kowdle, Christoph Rhemann, Sean Ryan Fanello, Andrea Tagliasacchi, Jonathan Taylor 0001, Philip Davidson, Mingsong Dou, Cem Keskin, Sameh Khamis, David Kim 0002, Danhang Tang, Vladimir Tankovich, Julien P. C. Valentin, Shahram Izadi |
ACM Trans. Graph. | 2 |
| 2018 | LookinGood: enhancing performance capture with real-time neural re-renderingabstractMotivated by augmented and virtual reality applications such as telepresence, there has been a recent focus in real-time performance capture of humans under motion. However, given the real-time constraint, these systems often suffer from artifacts in geometry and texture such as holes and noise in the final rendering, poor lighting, and low-resolution textures. We take the novel approach to augment such real-time performance capture systems with a deep architecture that takes a rendering from an arbitrary viewpoint, and jointly performs completion, super resolution, and denoising of the imagery in real-time. We call this approach neural (re-)rendering , and our live system "LookinGood". Our deep architecture is trained to produce high resolution and high quality images from a coarse rendering in real-time. First, we propose a self-supervised training method that does not require manual ground-truth annotation. We contribute a specialized reconstruction error that uses semantic information to focus on relevant parts of the subject, e.g. the face. We also introduce a salient reweighing scheme of the loss function that is able to discard outliers. We specifically design the system for virtual and augmented reality headsets where the consistency between the left and right eye plays a crucial role in the final user experience. Finally, we generate temporally stable results by explicitly minimizing the difference between two consecutive frames. We tested the proposed system in two different scenarios: one involving a single RGB-D sensor, and upper body reconstruction of an actor, the second consisting of full body 360° capture. Through extensive experimentation, we demonstrate how our system generalizes across unseen sequences and subjects. Ricardo Martin-Brualla, Rohit Pandey, Shuoran Yang, Pavel Pidlypenskyi, Jonathan Taylor 0001, Julien P. C. Valentin, Sameh Khamis, Philip Davidson, Anastasia Tkach, Peter Lincoln, Adarsh Kowdle, Christoph Rhemann, Dan B. Goldman, Cem Keskin, Steven M. Seitz, Shahram Izadi, Sean Ryan Fanello |
ACM Trans. Graph. | 12 |
| 2018 | Depth from motion for smartphone ARabstractAugmented reality (AR) for smartphones has matured from a technology for earlier adopters, available only on select high-end phones, to one that is truly available to the general public. One of the key breakthroughs has been in low-compute methods for six degree of freedom (6DoF) tracking on phones using only the existing hardware (camera and inertial sensors). 6DoF tracking is the cornerstone of smartphone AR allowing virtual content to be precisely locked on top of the real world. However, to really give users the impression of believable AR, one requires mobile depth. Without depth, even simple effects such as a virtual object being correctly occluded by the real-world is impossible. However, requiring a mobile depth sensor would severely restrict the access to such features. In this article, we provide a novel pipeline for mobile depth that supports a wide array of mobile phones, and uses only the existing monocular color sensor. Through several technical contributions, we provide the ability to compute low latency dense depth maps using only a single CPU core of a wide range of (medium-high) mobile phones. We demonstrate the capabilities of our approach on high-level AR applications including real-time navigation and shopping. Julien P. C. Valentin, Adarsh Kowdle, Jonathan T. Barron, Neal Wadhwa, Maksym Dzitsiuk, Michael Schoenberg, Ambrus Csaszar, Eric Turner 0001, Ivan Dryanovski, João Afonso, Jose Pascoal, Konstantine Tsotsos, Mira Leung, Mirko Schmidt, Onur G. Guleryuz, Sameh Khamis, Vladimir Tankovich, Sean Ryan Fanello, Shahram Izadi, Christoph Rhemann |
ACM Trans. Graph. | 21 |
| 2017 | UltraStereo: Efficient Learning-Based Matching for Active Stereo SystemsabstractEfficient estimation of depth from pairs of stereo images is one of the core problems in computer vision. We efficiently solve the specialized problem of stereo matching under active illumination using a new learning-based algorithm. This type of active stereo i.e. stereo matching where scene texture is augmented by an active light projector is proving compelling for designing depth cameras, largely due to improved robustness when compared to time of flight or traditional structured light techniques. Our algorithm uses an unsupervised greedy optimization scheme that learns features that are discriminative for estimating correspondences in infrared images. The proposed method optimizes a series of sparse hyperplanes that are used at test time to remap all the image patches into a compact binary representation in O(1). The proposed algorithm is cast in a PatchMatch Stereo-like framework, producing depth maps at 500Hz. In contrast to standard structured light methods, our approach generalizes to different scenes, does not require tedious per camera calibration procedures and is not adversely affected by interference from overlapping sensors. Extensive evaluations show we surpass the quality and overcome the limitations of current depth sensing technologies. Sean Ryan Fanello, Julien P. C. Valentin, Christoph Rhemann, Adarsh Kowdle, Vladimir Tankovich, Philip Davidson, Shahram Izadi |
CVPR | 3 |
| 2017 | Low Compute and Fully Parallel Computer Vision with HashMatchabstractNumerous computer vision problems such as stereo depth estimation, object-class segmentation and fore-ground/background segmentation can be formulated as per-pixel image labeling tasks. Given one or many images as input, the desired output of these methods is usually a spatially smooth assignment of labels. The large amount of such computer vision problems has lead to significant research efforts, with the state of art moving from CRF-based approaches to deep CNNs and more recently, hybrids of the two. Although these approaches have significantly advanced the state of the art, the vast majority has solely focused on improving quantitative results and are not designed for low-compute scenarios. In this paper, we present a new general framework for a variety of computer vision labeling tasks, called HashMatch. Our approach is designed to be both fully parallel, i.e. each pixel is independently processed, and low-compute, with a model complexity an order of magnitude less than existing CNN and CRF-based approaches. We evaluate HashMatch extensively on several problems such as disparity estimation, image retrieval, feature approximation and background subtraction, for which HashMatch achieves high computational efficiency while producing high quality results. Sean Ryan Fanello, Julien P. C. Valentin, Adarsh Kowdle, Christoph Rhemann, Vladimir Tankovich, Carlo Ciliberto, Philip Davidson, Shahram Izadi |
ICCV | 4 |
| 2017 | Motion2fusion: real-time volumetric performance captureabstractWe present Motion2Fusion, a state-of-the-art 360 performance capture system that enables *real-time* reconstruction of arbitrary non-rigid scenes. We provide three major contributions over prior work: 1) a new non-rigid fusion pipeline allowing for far more faithful reconstruction of high frequency geometric details, avoiding the over-smoothing and visual artifacts observed previously. 2) a high speed pipeline coupled with a machine learning technique for 3D correspondence field estimation reducing tracking errors and artifacts that are attributed to fast motions. 3) a backward and forward non-rigid alignment strategy that more robustly deals with topology changes but is still free from scene priors. Our novel performance capture system demonstrates real-time results nearing 3x speed-up from previous state-of-the-art work on the exact same GPU hardware. Extensive quantitative and qualitative comparisons show more precise geometric and texturing results with less artifacts due to fast motions or topology changes than prior art. Mingsong Dou, Philip Davidson, Sean Ryan Fanello, Sameh Khamis, Adarsh Kowdle, Christoph Rhemann, Vladimir Tankovich, Shahram Izadi |
ACM Trans. Graph. | 6 |
| 2016 | HyperDepth: Learning Depth from Structured Light without MatchingabstractStructured light sensors are popular due to their robustness to untextured scenes and multipath. These systems triangulate depth by solving a correspondence problem between each camera and projector pixel. This is often framed as a local stereo matching task, correlating patches of pixels in the observed and reference image. However, this is computationally intensive, leading to reduced depth accuracy and framerate. We contribute an algorithm for solving this correspondence problem efficiently, without compromising depth accuracy. For the first time, this problem is cast as a classification-regression task, which we solve extremely efficiently using an ensemble of cascaded random forests. Our algorithm scales in number of disparities, and each pixel can be processed independently, and in parallel. No matching or even access to the corresponding reference pattern is required at runtime, and regressed labels are directly mapped to depth. Our GPU-based algorithm runs at a 1KHz for 1.3MP input/output images, with disparity error of 0.1 subpixels. We show a prototype high framerate depth camera running at 375Hz, useful for solving tracking-related problems. We demonstrate our algorithmic performance, creating high resolution real-time depth maps that surpass the quality of current state of the art depth technologies, highlighting quantization-free results with reduced holes, edge fattening and other stereo-based depth artifacts. Sean Ryan Fanello, Christoph Rhemann, Vladimir Tankovich, Adarsh Kowdle, Sergio Orts, David Kim 0002, Shahram Izadi |
CVPR | 2 |
| 2016 | The Global Patch ColliderabstractThis paper proposes a novel extremely efficient, fully-parallelizable, task-specific algorithm for the computation of global point-wise correspondences in images and videos. Our algorithm, the Global Patch Collider, is based on detecting unique collisions between image points using a collection of learned tree structures that act as conditional hash functions. In contrast to conventional approaches that rely on pairwise distance computation, our algorithm isolates distinctive pixel pairs that hit the same leaf during traversal through multiple learned tree structures. The split functions stored at the intermediate nodes of the trees are trained to ensure that only visually similar patches or their geometric or photometric transformed versions fall into the same leaf node. The matching process involves passing all pixel positions in the images under analysis through the tree structures. We then compute matches by isolating points that uniquely collide with each other ie. fell in the same empty leaf in multiple trees. Our algorithm is linear in the number of pixels but can be made constant time on a parallel computation architecture as the tree traversal for individual image points is decoupled. We demonstrate the efficacy of our method by using it to perform optical flow matching and stereo matching on some challenging benchmarks. Experimental results show that not only is our method extremely computationally efficient, but it is also able to match or outperform state of the art methods that are much more complex. Shenlong Wang, Sean Ryan Fanello, Christoph Rhemann, Shahram Izadi, Pushmeet Kohli |
CVPR | 3 |
| 2016 | Holoportation: Virtual 3D Teleportation in Real-timeabstractWe present an end-to-end system for augmented and virtual reality telepresence, called Holoportation. Our system demonstrates high-quality, real-time 3D reconstructions of an entire space, including people, furniture and objects, using a set of new depth cameras. These 3D models can also be transmitted in real-time to remote users. This allows users wearing virtual or augmented reality displays to see, hear and interact with remote participants in 3D, almost as if they were present in the same physical space. From an audio-visual perspective, communicating and interacting with remote users edges closer to face-to-face communication. This paper describes the Holoportation technical system in full, its key interactive capabilities, the application scenarios it enables, and an initial qualitative study of using this new communication medium. Sergio Orts, Christoph Rhemann, Sean Ryan Fanello, Wayne Chang, Adarsh Kowdle, Yury Degtyarev, David Kim 0002, Philip Davidson, Sameh Khamis, Mingsong Dou, Vladimir Tankovich, Charles T. Loop, Qin Cai, Philip A. Chou, Sarah Mennicken, Julien P. C. Valentin, Vivek Pradeep, Shenlong Wang, Sing Bing Kang, Pushmeet Kohli, Yuliya Lutchyn, Cem Keskin, Shahram Izadi |
UIST | 2 |
| 2016 | Fusion4D: real-time performance capture of challenging scenesabstractWe contribute a new pipeline for live multi-view performance capture, generating temporally coherent high-quality reconstructions in real-time. Our algorithm supports both incremental reconstruction, improving the surface estimation over time, as well as parameterizing the nonrigid scene motion. Our approach is highly robust to both large frame-to-frame motion and topology changes, allowing us to reconstruct extremely challenging scenes. We demonstrate advantages over related real-time techniques that either deform an online generated template or continually fuse depth data nonrigidly into a single reference model. Finally, we show geometric reconstruction results on par with offline methods which require orders of magnitude more processing time and many more RGBD cameras. Mingsong Dou, Sameh Khamis, Yury Degtyarev, Philip Davidson, Sean Ryan Fanello, Adarsh Kowdle, Sergio Orts, Christoph Rhemann, David Kim 0002, Jonathan Taylor 0001, Pushmeet Kohli, Vladimir Tankovich, Shahram Izadi |
ACM Trans. Graph. | 8 |
| 2015 | Accurate, Robust, and Flexible Real-time Hand TrackingabstractWe present a new real-time hand tracking system based on a single depth camera. The system can accurately reconstruct complex hand poses across a variety of subjects. It also allows for robust tracking, rapidly recovering from any temporary failures. Most uniquely, our tracker is highly flexible, dramatically improving upon previous approaches which have focused on front-facing close-range scenarios. This flexibility opens up new possibilities for human-computer interaction with examples including tracking at distances from tens of centimeters through to several meters (for controlling the TV at a distance), supporting tracking using a moving depth camera (for mobile scenarios), and arbitrary camera placements (for VR headsets). These features are achieved through a new pipeline that combines a multi-layered discriminative reinitialization strategy for per-frame pose estimation, followed by a generative model-fitting stage. We provide extensive technical details and a detailed qualitative and quantitative analysis. Toby Sharp, Cem Keskin, Duncan P. Robertson, Jonathan Taylor 0001, Jamie Shotton, David Kim 0002, Christoph Rhemann, Ido Leichter, Alon Vinnikov, Daniel Freedman, Pushmeet Kohli, Eyal Krupka, Andrew W. Fitzgibbon, Shahram Izadi |
CHI | 7 |
| 2015 | A light transport model for mitigating multipath interference in Time-of-flight sensorsabstractContinuous-wave Time-of-flight (TOF) range imaging has become a commercially viable technology with many applications in computer vision and graphics. However, the depth images obtained from TOF cameras contain scene dependent errors due to multipath interference (MPI). Specifically, MPI occurs when multiple optical reflections return to a single spatial location on the imaging sensor. Many prior approaches to rectifying MPI rely on sparsity in optical reflections, which is an extreme simplification. In this paper, we correct MPI by combining the standard measurements from a TOF camera with information from direct and global light transport. We report results on both simulated experiments and physical experiments (using the Kinect sensor). Our results, evaluated against ground truth, demonstrate a quantitative improvement in depth accuracy. Nikhil Naik 0003, Achuta Kadambi, Christoph Rhemann, Shahram Izadi, Ramesh Raskar, Sing Bing Kang |
CVPR | 3 |
| 2014 | RetroDepth: 3D silhouette sensing for high-precision input on and above physical surfacesabstractWe present RetroDepth, a new vision-based system for accurately sensing the 3D silhouettes of hands, styluses, and other objects, as they interact on and above physical surfaces. Our setup is simple, cheap, and easily reproducible, comprising of two infrared cameras, diffuse infrared LEDs, and any off-the-shelf retro-reflective material. The retro-reflector aids image segmentation, creating a strong contrast between the surface and any object in proximity. A new highly efficient stereo matching algorithm precisely estimates the 3D contours of interacting objects and the retro-reflective surfaces. A novel pipeline enables 3D finger, hand and object tracking, as well as gesture recognition, purely using these 3D contours. We demonstrate high-precision sensing, allowing robust disambiguation between a finger or stylus touching, pressing or interacting above the surface. This allows many interactive scenarios that seamlessly mix together freehand 3D interactions with touch, pressure and stylus input. As shown, these rich modalities of input are enabled on and above any retro-reflective surface, including custom "physical widgets" fabricated by users. We compare our system with Kinect and Leap Motion, and conclude with limitations and future work. David Kim 0002, Shahram Izadi, Jakub Dostal, Christoph Rhemann, Cem Keskin, Christopher Zach, Jamie Shotton, Timothy A. Large, Steven Bathiche, Matthias Nießner, Alex Butler, Sean Ryan Fanello, Vivek Pradeep |
CHI | 4 |
| 2014 | FlexSense: a transparent self-sensing deformable surfaceabstractWe present FlexSense, a new thin-film, transparent sensing surface based on printed piezoelectric sensors, which can reconstruct complex deformations without the need for any external sensing, such as cameras. FlexSense provides a fully self-contained setup which improves mobility and is not affected from occlusions. Using only a sparse set of sensors, printed on the periphery of the surface substrate, we devise two new algorithms to fully reconstruct the complex deformations of the sheet, using only these sparse sensor measurements. An evaluation shows that both proposed algorithms are capable of reconstructing complex deformations accurately. We demonstrate how FlexSense can be used for a variety of 2.5D interactions, including as a transparent cover for tablets where bending can be performed alongside touch to enable magic lens style effects, layered input, and mode switching, as well as the ability to use our device as a high degree-of-freedom input controller for gaming and beyond. Christian Rendl, David Kim 0002, Sean Ryan Fanello, Patrick Parzer, Christoph Rhemann, Jonathan Taylor 0001, Martin Zirkl, Gregor Scheipl, Thomas Rothländer, Michael Haller, Shahram Izadi |
UIST | 5 |
| 2014 | 3D-board: a whole-body remote collaborative whiteboardabstractThis paper presents 3D-Board, a digital whiteboard capable of capturing life-sized virtual embodiments of geographically distributed users. When using large-scale screens for remote collaboration, awareness for the distributed users' gestures and actions is of particular importance. Our work adds to the literature on remote collaborative workspaces, it facilitates intuitive remote collaboration on large scale interactive whiteboards by preserving awareness of the full-body pose and gestures of the remote collaborator. By blending the front-facing 3D embodiment of a remote collaborator with the shared workspace, an illusion is created as if the observer was looking through the transparent whiteboard into the remote user's room. The system was tested and verified in a usability assessment, showing that 3D-Board significantly improves the effectiveness of remote collaboration on a large interactive surface. Jakob Zillner, Christoph Rhemann, Shahram Izadi, Michael Haller |
UIST | 2 |
| 2014 | Real-time non-rigid reconstruction using an RGB-D cameraabstractWe present a combined hardware and software solution for markerless reconstruction of non-rigidly deforming physical objects with arbitrary shape in real-time . Our system uses a single self-contained stereo camera unit built from off-the-shelf components and consumer graphics hardware to generate spatio-temporally coherent 3D models at 30 Hz. A new stereo matching algorithm estimates real-time RGB-D data. We start by scanning a smooth template model of the subject as they move rigidly. This geometric surface prior avoids strong scene assumptions, such as a kinematic human skeleton or a parametric shape model. Next, a novel GPU pipeline performs non-rigid registration of live RGB-D data to the smooth template using an extended non-linear as-rigid-as-possible (ARAP) framework. High-frequency details are fused onto the final mesh using a linear deformation model. The system is an order of magnitude faster than state-of-the-art methods, while matching the quality and robustness of many offline algorithms. We show precise real-time reconstructions of diverse scenes, including: large deformations of users' heads, hands, and upper bodies; fine-scale wrinkles and folds of skin and clothing; and non-rigid interactions performed by users on flexible objects such as toys. We demonstrate how acquired models can be used for many interactive scenarios, including re-texturing, online performance capture and preview, and real-time shape and motion re-targeting. Michael Zollhöfer, Matthias Nießner, Shahram Izadi, Christoph Rhemann, Christopher Zach, Matthew Fisher, Chenglei Wu, Andrew W. Fitzgibbon, Charles T. Loop, Christian Theobalt, Marc Stamminger |
ACM Trans. Graph. | 4 |
| 2013 | Depth Super Resolution by Rigid Body Self-Similarity in 3DabstractWe tackle the problem of jointly increasing the spatial resolution and apparent measurement accuracy of an input low-resolution, noisy, and perhaps heavily quantized depth map. In stark contrast to earlier work, we make no use of ancillary data like a color image at the target resolution, multiple aligned depth maps, or a database of high-resolution depth exemplars. Instead, we proceed by identifying and merging patch correspondences within the input depth map itself, exploiting patch wise scene self-similarity across depth such as repetition of geometric primitives or object symmetry. While the notion of 'single-image' super resolution has successfully been applied in the context of color and intensity images, we are to our knowledge the first to present a tailored analogue for depth images. Rather than reason in terms of patches of 2D pixels as others have before us, our key contribution is to proceed by reasoning in terms of patches of 3D points, with matched patch pairs related by a respective 6 DoF rigid body motion in 3D. In support of obtaining a dense correspondence field in reasonable time, we introduce a new 3D variant of Patch Match. A third contribution is a simple, yet effective patch up scaling and merging technique, which predicts sharp object boundaries at the target resolution. We show that our results are highly competitive with those of alternative techniques leveraging even a color image at the target resolution or a database of high-resolution depth exemplars. Michael Hornacek, Christoph Rhemann, Margrit Gelautz, Carsten Rother |
CVPR | 2 |
| 2013 | MonoFusion: Real-time 3D reconstruction of small scenes with a single web cameraabstractMonoFusion allows a user to build dense 3D reconstructions of their environment in real-time, utilizing only a single, off-the-shelf web camera as the input sensor. The camera could be one already available in a tablet, phone, or a standalone device. No additional input hardware is required. This removes the need for power intensive active sensors that do not work robustly in natural outdoor lighting. Using the input stream of the camera we first estimate the 6DoF camera pose using a sparse tracking method. These poses are then used for efficient dense stereo matching between the input frame and a key frame (extracted previously). The resulting dense depth maps are directly fused into a voxel-based implicit model (using a computationally inexpensive method) and surfaces are extracted per frame. The system is able to recover from tracking failures as well as filter out geometrically inconsistent noise from the 3D reconstruction. Our method is both simple to implement and efficient, making such systems even more accessible. This paper details the algorithmic components that make up our system and a GPU implementation of our approach. Qualitative results demonstrate high quality reconstructions even visually comparable to active depth sensor-based systems such as KinectFusion. Vivek Pradeep, Christoph Rhemann, Shahram Izadi, Christopher Zach, Michael Bleyer, Steven Bathiche |
ISMAR | 2 |
| 2013 | Fast Cost-Volume Filtering for Visual Correspondence and BeyondabstractMany computer vision tasks can be formulated as labeling problems. The desired solution is often a spatially smooth labeling where label transitions are aligned with color edges of the input image. We show that such solutions can be efficiently achieved by smoothing the label costs with a very fast edge-preserving filter. In this paper, we propose a generic and simple framework comprising three steps: 1) constructing a cost volume, 2) fast cost volume filtering, and 3) Winner-Takes-All label selection. Our main contribution is to show that with such a simple framework state-of-the-art results can be achieved for several computer vision applications. In particular, we achieve 1) disparity maps in real time whose quality exceeds those of all other fast (local) approaches on the Middlebury stereo benchmark, and 2) optical flow fields which contain very fine structures as well as large displacements. To demonstrate robustness, the few parameters of our framework are set to nearly identical values for both applications. Also, competitive results for interactive image segmentation are presented. With this work, we hope to inspire other researchers to leverage this framework to other application areas. Asmaa Hosni, Christoph Rhemann, Michael Bleyer, Carsten Rother, Margrit Gelautz |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2012 | Extracting 3D Scene-Consistent Object Proposals and Depth from Stereo Images
Michael Bleyer, Christoph Rhemann, Carsten Rother |
ECCV (5) | 2 |
| 2012 | User-Centric Learning and Evaluation of Interactive Segmentation Systems
Pushmeet Kohli, Hannes Nickisch, Carsten Rother, Christoph Rhemann |
Int. J. Comput. Vis. | 4 |
| 2011 | PatchMatch Stereo - Stereo Matching with Slanted Support WindowsabstractCommon local stereo methods match support windows at integer-valued disparities. The implicit assumption that pixels within the support region have constant disparity does not hold for slanted surfaces and leads to a bias towards reconstructing fronto-parallel surfaces. This work overcomes this bias by estimating an individual 3D plane at each pixel onto which the support region is projected. The major challenge of this ap-proach is to find a pixel’s optimal 3D plane among all possible planes whose number is infinite. We show that an ideal algorithm to solve this problem is PatchMatch [1] that we extend to find an approximate nearest neighbor according to a plane. In addition to Patch-Match’s spatial propagation scheme, we propose (1) view propagation where planes are propagated among left and right views of the stereo pair and (2) temporal propagation where planes are propagated from preceding and consecutive frames of a video when doing temporal stereo. Adaptive support weights are used in matching cost aggregation to improve results at disparity borders. We also show that our slanted support windows can be used to compute a cost volume for global stereo methods, which allows for ex-plicit treatment of occlusions and can handle large untextured regions. In the results we demonstrate that our method reconstructs highly slanted surfaces and achieves impres-sive disparity details with sub-pixel precision. In the Middlebury table, our method is currently top-performer among local methods and takes rank 2 among approximately 110 competitors if sub-pixel precision is considered. 1 Michael Bleyer, Christoph Rhemann, Carsten Rother |
BMVC | 2 |
| 2011 | A global sampling method for alpha mattingabstractAlpha matting refers to the problem of softly extracting the foreground from an image. Given a trimap (specifying known foreground/background and unknown pixels), a straightforward way to compute the alpha value is to sample some known foreground and background colors for each unknown pixel. Existing sampling-based matting methods often collect samples near the unknown pixels only. They fail if good samples cannot be found nearby. In this paper, we propose a global sampling method that uses all samples available in the image. Our global sample set avoids missing good samples. A simple but effective cost function is defined to tackle the ambiguity in the sample selection process. To handle the computational complexity introduced by the large number of samples, we pose the sampling task as a correspondence problem. The correspondence search is efficiently achieved by generalizing a randomized algorithm previously designed for patch matching[3]. A variety of experiments show that our global sampling method produces both visually and quantitatively high-quality matting results. Kaiming He, Christoph Rhemann, Carsten Rother, Xiaoou Tang, Jian Sun 0001 |
CVPR | 2 |
| 2011 | Fast cost-volume filtering for visual correspondence and beyondabstractMany computer vision tasks can be formulated as labeling problems. The desired solution is often a spatially smooth labeling where label transitions are aligned with color edges of the input image. We show that such solutions can be efficiently achieved by smoothing the label costs with a very fast edge preserving filter. In this paper we propose a generic and simple framework comprising three steps: (i) constructing a cost volume (ii) fast cost volume filtering and (iii) winner-take-all label selection. Our main contribution is to show that with such a simple framework state-of-the-art results can be achieved for several computer vision applications. In particular, we achieve (i) disparity maps in real-time, whose quality exceeds those of all other fast (local) approaches on the Middlebury stereo benchmark, and (ii) optical flow fields with very fine structures as well as large displacements. To demonstrate robustness, the few parameters of our framework are set to nearly identical values for both applications. Also, competitive results for interactive image segmentation are presented. With this work, we hope to inspire other researchers to leverage this framework to other application areas. Christoph Rhemann, Asmaa Hosni, Michael Bleyer, Carsten Rother, Margrit Gelautz |
CVPR | 1 |
| 2011 | REal-time local stereo matching using guided image filteringabstractAdaptive support weight algorithms represent the state-of the-art in local stereo matching. Their limitation is a high computational demand, which makes them unattractive for many (real-time) applications. To our knowledge, the algorithm proposed in this paper is the first local method which is both fast (real-time) and produces results comparable to global algorithms. A key insight is that the aggregation step of adaptive support weight algorithms is equivalent to smoothing the stereo cost volume with an edge-preserving filter. From this perspective, the original adaptive support weight algorithm [1] applies bilateral filtering on cost volume slices, and the reason for its poor computational behavior is that bilateral filtering is a relatively slow process. We suggest to use the recently proposed guided filter [2] to overcome this limitation. Analogously to the bilateral filter, this filter has edge preserving properties, but can be implemented in a very fast way, which makes our stereo algorithm independent of the size of the match window. The GPU implementation of our stereo algorithm can process stereo images with a resolution of 640 × 480 pixels and a disparity range of 26 pixels at 25 fps. According to the Middlebury on-line ranking, our algorithm achieves rank 14 out of over 100 submissions and is not only the best performing local stereo matching method, but also the best performing real-time method. Asmaa Hosni, Michael Bleyer, Christoph Rhemann, Margrit Gelautz, Carsten Rother |
ICME | 3 |
| 2011 | Temporally Consistent Disparity and Optical Flow via Efficient Spatio-temporal Filtering
Asmaa Hosni, Christoph Rhemann, Michael Bleyer, Margrit Gelautz |
PSIVT (1) | 2 |
| 2010 | A spatially varying PSF-based prior for alpha mattingabstractIn this paper we considerably improve on a state-of-the-art alpha matting approach by incorporating a new prior which is based on the image formation process. In particular, we model the prior probability of an alpha matte as the convolution of a high-resolution binary segmentation with the spatially varying point spread function (PSF) of the camera. Our main contribution is a new and efficient de-convolution approach that recovers the prior model, given an approximate alpha matte. By assuming that the PSF is a kernel with a single peak, we are able to recover the binary segmentation with an MRF-based approach, which exploits flux and a new way of enforcing connectivity. The spatially varying PSF is obtained via a partitioning of the image into regions of similar defocus. Incorporating our new prior model into a state-of-the-art matting technique produces results that outperform all competitors, which we confirm using a publicly available benchmark. Christoph Rhemann, Carsten Rother, Pushmeet Kohli, Margrit Gelautz |
CVPR | 1 |
| 2009 | A stereo approach that handles the matting problem via image warpingabstractWe propose an algorithm that simultaneously extracts disparities and alpha matting information given a stereo image pair. Our method divides the reference image into a set of overlapping, partially transparent color segments. Each segment pixel is assigned an alpha value and color. The disparity inside the segment is modeled via a plane. The goodness of alphas, colors and disparity planes is measured by a new energy function. Its basic idea is to use the three parameters for generating artificial views representing the left and right images. If alphas, colors and disparity planes are correct, these artificial images should be very similar to the real ones. For generating the artificial right view, we warp all pixels of the left into the geometry of the right image using the disparity planes. We introduce the assumption of constant solidity in order to correctly model how pixels' alpha values are affected by the warping operation. Experimental results on the Middlebury set show that our algorithm gives good results in comparison to the state-of-the-art in stereo matching. Michael Bleyer, Margrit Gelautz, Carsten Rother, Christoph Rhemann |
CVPR | 4 |
| 2009 | A perceptually motivated online benchmark for image mattingabstractThe availability of quantitative online benchmarks for low-level vision tasks such as stereo and optical flow has led to significant progress in the respective fields. This paper introduces such a benchmark for image matting. There are three key factors for a successful benchmarking system: (a) a challenging, high-quality ground truth test set; (b) an online evaluation repository that is dynamically updated with new results; (c) perceptually motivated error functions. Our new benchmark strives to meet all three criteria. We evaluated several matting methods with our benchmark and show that their performance varies depending on the error function. Also, our challenging test set reveals problems of existing algorithms, not reflected in previously reported results. We hope that our effort will lead to considerable progress in the field of image matting, and welcome the reader to visit our benchmark at www.aIphamatting.com. Christoph Rhemann, Carsten Rother, Jue Wang 0001, Margrit Gelautz, Pushmeet Kohli, Pamela Rott |
CVPR | 1 |
| 2009 | New appearance models for natural image mattingabstractImage matting is the task of estimating a fore- and background layer from a single image. To solve this ill posed problem, an accurate modeling of the scene's appearance is necessary. Existing methods that provide a closed form solution to this problem, assume that the colors of the foreground and background layers are locally linear. In this paper, we show that such models can be an overfit when the colors of the two layers are locally constant. We derive new closed form expressions in such cases, and show that our models are more compact than existing ones. In particular, the null space of our cost function is a subset of the null space constructed by existing approaches. We discuss the bias towards specific solutions for each formulation. Experiments on synthetic and real data confirm that our compact models estimate alpha mattes more accurately than existing techniques, without the need of additional user interaction. Dheeraj Singaraju, Carsten Rother, Christoph Rhemann |
CVPR | 3 |
| 2009 | Local stereo matching using geodesic support weightsabstractLocal stereo matching has recently experienced large progress by the introduction of new support aggregation schemes. These approaches estimate a pixel's support region via color segmentation. Our contribution lies in an improved method for accomplishing this segmentation. Inside a square support window, we compute the geodesic distance from all pixels to the window's center pixel. Pixels of low geodesic distance are given high support weights and therefore large influence in the matching process. In contrast to previous work, we enforce connectivity by using the geodesic distance transform. For obtaining a high support weight, a pixel must have a path to the center point along which the color does not change significantly. This connectivity property leads to improved segmentation results and consequently to improved disparity maps. The success of our geodesic approach is demonstrated on the Middlebury images. According to the Middlebury benchmark, the proposed algorithm is the top performer among local stereo methods at the current state-of-the-art. Asmaa Hosni, Michael Bleyer, Margrit Gelautz, Christoph Rhemann |
ICIP | 4 |
| 2008 | Improving Color Modeling for Alpha MattingabstractThis paper addresses the problem of extracting an alpha matte from a single photograph given a user-defined trimap. A crucial part of this task is the color modeling step where for each pixel the optimal alpha value, together with its confidence, is estimated individually. This forms the data term of the objective function. It comprises of three steps: (i) Collecting a candidate set of potential fore- and background colors; (ii) Selecting high confidence samples from the candidate set; (iii) Estimating a sparsity prior to remove blurry artifacts. We introduce novel ideas for each of these steps and show that our approach considerably improves over state-of-the-art techniques by evaluating it on a large database of 54 images with known high-quality ground truth. 1 Christoph Rhemann, Carsten Rother, Margrit Gelautz |
BMVC | 1 |
| 2008 | High resolution matting via interactive trimap segmentationabstractWe present a new approach to the matting problem which splits the task into two steps: interactive trimap extraction followed by trimap-based alpha matting. By doing so we gain considerably in terms of speed and quality and are able to deal with high resolution images. This paper has three contributions: (i) a new trimap segmentation method using parametric max-flow; (ii) an alpha matting technique for high resolution images with a new gradient preserving prior on alpha; (iii) a database of 27 ground truth alpha mattes of still objects, which is considerably larger than previous databases and also of higher quality. The database is used to train our system and to validate that both our trimap extraction and our matting method improve on state-of-the-art techniques. Christoph Rhemann, Carsten Rother, Alex Rav-Acha, Toby Sharp |
CVPR | 1 |