EDBT 2026 Demo / reviewers in the wild / expert
Jonathan Taylor 0001
dblp:68/6729-1
· DBLP profile ↗
29ranked-venue papers
4as first author
3since 2021 · last 2024
—ORCID · unresolved
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 24 · 4 first-author · 3 since 2021Artificial intelligence and machine learning · 13 · 2 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 3
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
15 papers |
3D vision · 72% Face, body and person analysis · 16% Video understanding and tracking · 6% | |
| Computer graphics and multimedia
10 papers |
Rendering · 22% Geometric modeling and processing · 17% Image and video coding · 15% | |
| Human-computer interaction and pervasive computing
4 papers |
Interaction techniques and input · 84% Immersive interaction · 16% |
Topics — the 30 heaviest of 51, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › 3D vision
3d reconstruction |
1.9 | 7 | 2023 | Spectral Graphormer: Spectral Graph-based Transformer for Egocentric Two-Hand Reconstruction using Multi-View Color Images · ICCV 2023 Fusion4D: real-time performance capture of challenging scenes · ACM Trans. Graph. 2016 Learning an efficient model of hand shape variation from depth images · CVPR 2015 |
Computer vision › Face, body and person analysis › human pose estimation › articulated pose estimation
hand pose estimation |
1.3 | 4 | 2023 | Spectral Graphormer: Spectral Graph-based Transformer for Egocentric Two-Hand Reconstruction using Multi-View Color Images · ICCV 2023 Opening the Black Box: Hierarchical Sampling Optimization for Hand Pose Estimation · IEEE Trans. Pattern Anal. Mach. Intell. 2019 Opening the Black Box: Hierarchical Sampling Optimization for Estimating Human Hand Pose · ICCV 2015 |
Rendering
neural rendering |
0.8 | 2 | 2020 | Deep relightable textures: volumetric performance capture with neural rendering · ACM Trans. Graph. 2020 LookinGood: enhancing performance capture with real-time neural re-rendering · ACM Trans. Graph. 2018 |
Interaction techniques and input › input sensing › tracking
hand tracking |
0.8 | 3 | 2017 | Articulated distance fields for ultra-fast tracking of hands interacting · ACM Trans. Graph. 2017 Efficient and precise interactive hand tracking through joint, continuous optimization of pose and correspondences · ACM Trans. Graph. 2016 Accurate, Robust, and Flexible Real-time Hand Tracking · CHI 2015 |
Computer vision › 3D vision
pose estimation |
0.7 | 3 | 2019 | Opening the Black Box: Hierarchical Sampling Optimization for Hand Pose Estimation · IEEE Trans. Pattern Anal. Mach. Intell. 2019 Opening the Black Box: Hierarchical Sampling Optimization for Estimating Human Hand Pose · ICCV 2015 The need 4 speed in real-time dense visual tracking · ACM Trans. Graph. 2018 |
Computer vision › 3D vision › 3d human reconstruction › 3d hand reconstruction
two-hand reconstruction |
0.7 | 1 | 2023 | Spectral Graphormer: Spectral Graph-based Transformer for Egocentric Two-Hand Reconstruction using Multi-View Color Images · ICCV 2023 |
Computer vision › 3D vision › 3d reconstruction
non-rigid reconstruction |
0.7 | 3 | 2016 | Fusion4D: real-time performance capture of challenging scenes · ACM Trans. Graph. 2016 3D scanning deformable objects with a single RGBD sensor · CVPR 2015 User-Specific Hand Modeling from Monocular Depth Sequences · CVPR 2014 |
Computer vision › 3D vision › 3d shape modeling
hand shape modeling |
0.5 | 2 | 2016 | Fits Like a Glove: Rapid and Reliable Hand Shape Personalization · CVPR 2016 Learning an efficient model of hand shape variation from depth images · CVPR 2015 |
Image and video coding
texture compression |
0.4 | 1 | 2020 | Deep Implicit Volume Compression · CVPR 2020 |
Computer vision › 3D vision › novel view synthesis
free-viewpoint rendering |
0.4 | 1 | 2019 | Volumetric Capture of Humans With a Single RGBD Camera via Semi-Parametric Learning · CVPR 2019 |
Computer vision › 3D vision
novel view synthesis |
0.4 | 1 | 2019 | Volumetric Capture of Humans With a Single RGBD Camera via Semi-Parametric Learning · CVPR 2019 |
Computational photography and imaging › 3d scanning
volumetric performance capture |
0.4 | 1 | 2019 | The relightables: volumetric performance capture of humans with realistic relighting · ACM Trans. Graph. 2019 |
Computer vision › Video understanding and tracking › motion tracking
dense tracking |
0.3 | 1 | 2018 | The need 4 speed in real-time dense visual tracking · ACM Trans. Graph. 2018 |
Computer vision › 3D vision
depth estimation |
0.3 | 1 | 2018 | The need 4 speed in real-time dense visual tracking · ACM Trans. Graph. 2018 |
Computer vision › Video understanding and tracking
object tracking |
0.3 | 1 | 2018 | The need 4 speed in real-time dense visual tracking · ACM Trans. Graph. 2018 |
Image and video processing
image restoration |
0.3 | 1 | 2018 | LookinGood: enhancing performance capture with real-time neural re-rendering · ACM Trans. Graph. 2018 |
Multimedia systems and quality of experience › video streaming
real-time streaming |
0.3 | 1 | 2018 | Real-time compression and streaming of 4D performances · ACM Trans. Graph. 2018 |
Image and video processing › image restoration › multi-task image restoration
super-resolution and denoising |
0.3 | 1 | 2018 | LookinGood: enhancing performance capture with real-time neural re-rendering · ACM Trans. Graph. 2018 |
Virtual and augmented reality › immersive interaction
hand tracking |
0.3 | 1 | 2017 | Articulated distance fields for ultra-fast tracking of hands interacting · ACM Trans. Graph. 2017 |
Computer vision › 3D vision › motion capture
human performance capture |
0.2 | 1 | 2016 | Fusion4D: real-time performance capture of challenging scenes · ACM Trans. Graph. 2016 |
Computer vision › 3D vision › 3d reconstruction
multi-view stereo |
0.2 | 1 | 2016 | Fusion4D: real-time performance capture of challenging scenes · ACM Trans. Graph. 2016 |
Computer vision › 3D vision › 3d reconstruction › volumetric reconstruction
volumetric fusion |
0.2 | 1 | 2016 | Fusion4D: real-time performance capture of challenging scenes · ACM Trans. Graph. 2016 |
Geometric modeling and processing
model fitting |
0.2 | 1 | 2016 | Efficient and precise interactive hand tracking through joint, continuous optimization of pose and correspondences · ACM Trans. Graph. 2016 |
Rendering
relighting |
0.2 | 2 | 2019 | Deep reflectance fields: high-quality facial reflectance field inference from color gradient illumination · ACM Trans. Graph. 2019 The relightables: volumetric performance capture of humans with realistic relighting · ACM Trans. Graph. 2019 |
Computer vision › 3D vision › range sensing
3d scanning |
0.2 | 1 | 2015 | 3D scanning deformable objects with a single RGBD sensor · CVPR 2015 |
Computer vision › 3D vision
3d shape modeling |
0.2 | 1 | 2015 | Learning an efficient model of hand shape variation from depth images · CVPR 2015 |
Computer vision › 3D vision
correspondence estimation |
0.2 | 1 | 2015 | Metric Regression Forests for Correspondence Estimation · Int. J. Comput. Vis. 2015 |
Machine learning › Kernel, tree and ensemble methods › ensemble learning › tree ensembles
random forest |
0.2 | 1 | 2015 | Metric Regression Forests for Correspondence Estimation · Int. J. Comput. Vis. 2015 |
Machine learning › Kernel, tree and ensemble methods › ensemble learning › tree ensembles › random forest
regression forests |
0.2 | 1 | 2015 | Metric Regression Forests for Correspondence Estimation · Int. J. Comput. Vis. 2015 |
Interaction techniques and input › input sensing › tracking › hand tracking
hand pose estimation |
0.2 | 1 | 2015 | Accurate, Robust, and Flexible Real-time Hand Tracking · CHI 2015 |
Methods — techniques the papers use, named apart from their topics
signed distance function · 0.9deep neural network · 0.8color gradient illumination · 0.8transformer · 0.7spectral graph convolution · 0.7optimization-based refinement · 0.7neural network compression · 0.4geometric reconstruction pipeline · 0.4UV parameterization · 0.4semi-parametric learning · 0.4neural network · 0.4machine learning reconstruction pipeline · 0.4kinematic hierarchy · 0.4decision forests · 0.4convolutional neural network · 0.4active depth sensing · 0.4semantic reconstruction error · 0.3self-supervised learning · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | MACS: Mass Conditioned 3D Hand and Object Motion SynthesisabstractThe physical properties of an object, such as mass, significantly affect how we manipulate it with our hands. Surprisingly, this aspect has so far been neglected in prior work on 3D motion synthesis. To improve the naturalness of the synthesized 3D hand-object motions, this work proposes MACS–the first MAss Conditioned 3D hand and object motion Synthesis approach. Our approach is based on cascaded diffusion models and generates interactions that plausibly adjust based on the object’s mass and interaction type. MACS also accepts a manually drawn 3D object trajectory as input and synthesizes the natural 3D hand motions conditioned by the object’s mass. This flexibility enables MACS to be used for various downstream applications, such as generating synthetic training data for ML tasks, fast animation of hands for graphics workflows, and generating character interactions for computer games. We show experimentally that a small-scale dataset is sufficient for MACS to reasonably generalize across interpolated and extrapolated object masses unseen during the training. Furthermore, MACS shows moderate generalization to unseen objects, thanks to the mass-conditioned contact labels generated by our surface contact synthesis model ConNet. Our comprehensive user study confirms that the synthesized 3D hand-object interactions are highly plausible and realistic. Soshi Shimada, Franziska Mueller 0001, Jan Bednarík, Bardia Doosti, Bernd Bickel, Danhang Tang, Vladislav Golyanik, Jonathan Taylor 0001, Christian Theobalt, Thabo Beeler |
3DV | 8 |
| 2023 | Spectral Graphormer: Spectral Graph-based Transformer for Egocentric Two-Hand Reconstruction using Multi-View Color ImagesabstractWe propose a novel transformer-based framework that reconstructs two high fidelity hands from multi-view RGB images. Unlike existing hand pose estimation methods, where one typically trains a deep network to regress hand model parameters from single RGB image, we consider a more challenging problem setting where we directly regress the absolute root poses of two-hands with extended forearm at high resolution from egocentric view. As existing datasets are either infeasible for egocentric viewpoints or lack background variations, we create a large-scale synthetic dataset with diverse scenarios and collect a real dataset from multi-calibrated camera setup to verify our proposed multi-view image feature fusion strategy. To make the reconstruction physically plausible, we propose two strategies: (i) a coarse-to-fine spectral graph convolution decoder to smoothen the meshes during upsampling and (ii) an optimisation-based refinement stage at inference to prevent self-penetrations. Through extensive quantitative and qualitative evaluations, we show that our framework is able to produce realistic two-hand reconstructions and demonstrate the generalisation of synthetic-trained models to real data, as well as real-time AR/VR applications. Tze Ho Elden Tse, Franziska Mueller 0001, Zhengyang Shen, Danhang Tang, Thabo Beeler, Mingsong Dou, Yinda Zhang 0001, Sasa Petrovic, Hyung Jin Chang, Jonathan Taylor 0001, Bardia Doosti |
ICCV | 10 |
| 2023 | Sandwiched Video Compression: Efficiently Extending the Reach of Standard Codecs with Neural WrappersabstractWe propose sandwiched video compression – a video compression system that wraps neural networks around a standard video codec. The sandwich framework consists of a neural pre- and post-processor with a standard video codec between them. The networks are trained jointly to optimize a rate-distortion loss function with the goal of significantly improving over the standard codec in various compression scenarios. End-to-end training in this setting requires a differentiable proxy for the standard video codec, which incorporates temporal processing with motion compensation, inter/intra mode decisions, and in-loop filtering. We propose differentiable approximations to key video codec components and demonstrate that, in addition to providing meaningful compression improvements over the standard codec, the neural codes of the sandwich lead to significantly better rate-distortion performance in two important scenarios. When transporting high-resolution video via low-resolution HEVC, the sandwich system obtains 6.5 dB improvements over standard HEVC. More importantly, using the well-known perceptual similarity metric, LPIPS, we observe 30% improvements in rate at the same quality over HEVC. Last but not least, we show that pre- and post-processors formed by very modestly-parameterized, light-weight networks can closely approximate these results. Berivan Isik, Onur G. Guleryuz, Danhang Tang, Jonathan Taylor 0001, Philip A. Chou |
ICIP | 4 |
| 2020 | Deep Implicit Volume CompressionabstractWe describe a novel approach for compressing truncated signed distance fields (TSDF) stored in 3D voxel grids, and their corresponding textures. To compress the TSDF, our method relies on a block-based neural network architecture trained end-to-end, achieving state-of-the-art rate-distortion trade-off. To prevent topological errors, we losslessly com- press the signs of the TSDF, which also upper bounds the reconstruction error by the voxel size. To compress the corresponding texture, we designed a fast block-based UV parameterization, generating coherent texture maps that can be effectively compressed using existing video compression algorithms. We demonstrate the performance of our algo- rithms on two 4D performance capture datasets, reducing bitrate by 66% for the same distortion, or alternatively re- ducing the distortion by 50% for the same bitrate, compared to the state-of-the-art. Danhang Tang, Philip A. Chou, Christian Häne, Mingsong Dou, Sean Ryan Fanello, Jonathan Taylor 0001, Philip L. Davidson, Onur G. Guleryuz, Yinda Zhang 0001, Shahram Izadi, Andrea Tagliasacchi, Sofien Bouaziz, Cem Keskin |
CVPR | 7 |
| 2020 | Deep relightable textures: volumetric performance capture with neural renderingabstractThe increasing demand for 3D content in augmented and virtual reality has motivated the development of volumetric performance capture systemsnsuch as the Light Stage. Recent advances are pushing free viewpoint relightable videos of dynamic human performances closer to photorealistic quality. However, despite significant efforts, these sophisticated systems are limited by reconstruction and rendering algorithms which do not fully model complex 3D structures and higher order light transport effects such as global illumination and sub-surface scattering. In this paper, we propose a system that combines traditional geometric pipelines with a neural rendering scheme to generate photorealistic renderings of dynamic performances under desired viewpoint and lighting. Our system leverages deep neural networks that model the classical rendering process to learn implicit features that represent the view-dependent appearance of the subject independent of the geometry layout, allowing for generalization to unseen subject poses and even novel subject identity. Detailed experiments and comparisons demonstrate the efficacy and versatility of our method to generate high-quality results, significantly outperforming the existing state-of-the-art solutions. Abhimitra Meka, Rohit Pandey, Christian Häne, Sergio Orts, Peter Barnum, Philip L. Davidson, Daniel Erickson, Yinda Zhang 0001, Jonathan Taylor 0001, Sofien Bouaziz, Chloe LeGendre, Wan-Chun Ma, Ryan S. Overbeck, Thabo Beeler, Paul E. Debevec, Shahram Izadi, Christian Theobalt, Christoph Rhemann, Sean Ryan Fanello |
ACM Trans. Graph. | 9 |
| 2019 | Volumetric Capture of Humans With a Single RGBD Camera via Semi-Parametric LearningabstractVolumetric (4D) performance capture is fundamental for AR/VR content generation. Whereas previous work in 4D performance capture has shown impressive results in studio settings, the technology is still far from being accessible to a typical consumer who, at best, might own a single RGBD sensor. Thus, in this work, we propose a method to synthesize free viewpoint renderings using a single RGBD camera. The key insight is to leverage previously seen “calibration” images of a given user to extrapolate what should be rendered in a novel viewpoint from the data available in the sensor. Given these past observations from multiple viewpoints, and the current RGBD image from a fixed view, we propose an end-to-end framework that fuses both these data sources to generate novel renderings of the performer. We demonstrate that the method can produce high fidelity images, and handle extreme changes in subject pose and camera viewpoints. We also show that the system generalizes to performers not seen in the training data. We run exhaustive experiments demonstrating the effectiveness of the proposed semi-parametric model (i.e. calibration images available to the neural network) compared to other state of the art machine learned solutions. Further, we compare the method with more traditional pipelines that employ multi-view capture. We show that our framework is able to achieve compelling results, with substantially less infrastructure than previously required. Rohit Pandey, Anastasia Tkach, Shuoran Yang, Pavel Pidlypenskyi, Jonathan Taylor 0001, Ricardo Martin-Brualla, Andrea Tagliasacchi, George Papandreou, Philip Davidson, Cem Keskin, Shahram Izadi, Sean Ryan Fanello |
CVPR | 5 |
| 2019 | Opening the Black Box: Hierarchical Sampling Optimization for Hand Pose EstimationabstractHand pose estimation, formulated as an inverse problem, is typically optimized by an energy function over pose parameters using a 'black box' image generation procedure, knowing little about either the relationships between the parameters or the form of the energy function. In this paper, we show significant improvement upon such black box optimization by exploiting high-level knowledge of the parameter structure and using a local surrogate energy function. Our new framework, hierarchical sampling optimization (HSO), consists of a sequence of discriminative predictors organized into a kinematic hierarchy. Each predictor is conditioned on its ancestors, and generates a set of samples over a subset of the pose parameters, with only one selected by the highly-efficient surrogate energy. The selected partial poses are concatenated to generate a full-pose hypothesis. Repeating the same process, several hypotheses are generated and the full energy function selects the best result. Under the same kinematic hierarchy, two methods based on decision forest and convolutional neural network are proposed to generate the samples and two optimization methods are studied when optimizing these samples. Experimental evaluations on three publicly available datasets show that our method is particularly impressive in low-compute scenarios where it significantly outperforms all other state-of-the-art methods. Danhang Tang, Qi Ye 0001, Shanxin Yuan, Jonathan Taylor 0001, Pushmeet Kohli, Cem Keskin, Tae-Kyun Kim 0001, Jamie Shotton |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2019 | The relightables: volumetric performance capture of humans with realistic relightingabstractWe present "The Relightables", a volumetric capture system for photorealistic and high quality relightable full-body performance capture. While significant progress has been made on volumetric capture systems, focusing on 3D geometric reconstruction with high resolution textures, much less work has been done to recover photometric properties needed for relighting. Results from such systems lack high-frequency details and the subject's shading is prebaked into the texture. In contrast, a large body of work has addressed relightable acquisition for image-based approaches, which photograph the subject under a set of basis lighting conditions and recombine the images to show the subject as they would appear in a target lighting environment. However, to date, these approaches have not been adapted for use in the context of a high-resolution volumetric capture system. Our method combines this ability to realistically relight humans for arbitrary environments, with the benefits of free-viewpoint volumetric capture and new levels of geometric accuracy for dynamic performances. Our subjects are recorded inside a custom geodesic sphere outfitted with 331 custom color LED lights, an array of high-resolution cameras, and a set of custom high-resolution depth sensors. Our system innovates in multiple areas: First, we designed a novel active depth sensor to capture 12.4 MP depth maps, which we describe in detail. Second, we show how to design a hybrid geometric and machine learning reconstruction pipeline to process the high resolution input and output a volumetric video. Third, we generate temporally consistent reflectance maps for dynamic performers by leveraging the information contained in two alternating color gradient illumination images acquired at 60Hz. Multiple experiments, comparisons, and applications show that The Relightables significantly improves upon the level of realism in placing volumetrically captured human performances into arbitrary CG scenes. Peter Lincoln, Philip Davidson, Jay Busch, Xueming Yu, Matt Whalen, Geoff Harvey, Sergio Orts, Rohit Pandey, Jason Dourgarian, Danhang Tang, Anastasia Tkach, Adarsh Kowdle, Emily Cooper, Mingsong Dou, Sean Ryan Fanello, Graham Fyffe, Christoph Rhemann, Jonathan Taylor 0001, Paul E. Debevec, Shahram Izadi |
ACM Trans. Graph. | 19 |
| 2019 | Deep reflectance fields: high-quality facial reflectance field inference from color gradient illuminationabstractWe present a novel technique to relight images of human faces by learning a model of facial reflectance from a database of 4D reflectance field data of several subjects in a variety of expressions and viewpoints. Using our learned model, a face can be relit in arbitrary illumination environments using only two original images recorded under spherical color gradient illumination. The output of our deep network indicates that the color gradient images contain the information needed to estimate the full 4D reflectance field, including specular reflections and high frequency details. While capturing spherical color gradient illumination still requires a special lighting setup, reduction to just two illumination conditions allows the technique to be applied to dynamic facial performance capture. We show side-by-side comparisons which demonstrate that the proposed system outperforms the state-of-the-art techniques in both realism and speed. Abhimitra Meka, Christian Häne, Rohit Pandey, Michael Zollhöfer, Sean Ryan Fanello, Graham Fyffe, Adarsh Kowdle, Xueming Yu, Jay Busch, Jason Dourgarian, Peter Denny, Sofien Bouaziz, Peter Lincoln, Matt Whalen, Geoff Harvey, Jonathan Taylor 0001, Shahram Izadi, Andrea Tagliasacchi, Paul E. Debevec, Christian Theobalt, Julien P. C. Valentin, Christoph Rhemann |
ACM Trans. Graph. | 16 |
| 2018 | TwinFusion: High Framerate Non-rigid Fusion through Fast Correspondence TrackingabstractReal time non-rigid reconstruction pipelines are extremely computationally expensive and easily saturate the highest end GPUs currently available. This requires careful strategic choices to be made about a set of highly interconnected parameters that divide up the limited compute. At the same time, offline systems, prove the value of increasing voxel resolution, more iterations, and higher frame rates. To this end, we demonstrate a set of remarkably simple but effective modifications to these algorithms that significantly reduce the average per-frame computation cost allowing these parameters to be increased. Specifically, we divide the depth stream into sub-frames and fusion-frames, disabling both model accumulation (fusion) and non-rigid alignment (model tracking) on the former. Instead, we efficiently track point correspondences across neighboring sub-frames. We then leverage these correspondences to initialize the standard non-rigid alignment to a fusion-frame where data can then be accumulated into the model. As a result, compute resources in the modified non-rigid reconstruction pipeline can be immediately re-purposed to increase voxel resolution, use more iterations or to increase the frame rate. To demonstrate the latter, we leverage recent high frame rate depth algorithms to build a novel “twin” sensor consisting of a low-res/highfps sub-frame camera and a second low-fps/high-res fusion camera. Jonathan Taylor 0001, Sean Ryan Fanello, Andrea Tagliasacchi, Mingsong Dou, Philip Davidson, Adarsh Kowdle, Shahram Izadi |
3DV | 2 |
| 2018 | The need 4 speed in real-time dense visual trackingabstractThe advent of consumer depth cameras has incited the development of a new cohort of algorithms tackling challenging computer vision problems. The primary reason is that depth provides direct geometric information that is largely invariant to texture and illumination. As such, substantial progress has been made in human and object pose estimation, 3D reconstruction and simultaneous localization and mapping. Most of these algorithms naturally benefit from the ability to accurately track the pose of an object or scene of interest from one frame to the next. However, commercially available depth sensors (typically running at 30fps) can allow for large inter-frame motions to occur that make such tracking problematic. A high frame rate depth camera would thus greatly ameliorate these issues, and further increase the tractability of these computer vision problems. Nonetheless, the depth accuracy of recent systems for high-speed depth estimation [Fanello et al. 2017b] can degrade at high frame rates. This is because the active illumination employed produces a low SNR and thus a high exposure time is required to obtain a dense accurate depth image. Furthermore in the presence of rapid motion, longer exposure times produce artifacts due to motion blur, and necessitates a lower frame rate that introduces large inter-frame motion that often yield tracking failures. In contrast, this paper proposes a novel combination of hardware and software components that avoids the need to compromise between a dense accurate depth map and a high frame rate. We document the creation of a full 3D capture system for high speed and quality depth estimation, and demonstrate its advantages in a variety of tracking and reconstruction tasks. We extend the state of the art active stereo algorithm presented in Fanello et al. [2017b] by adding a space-time feature in the matching phase. We also propose a machine learning based depth refinement step that is an order of magnitude faster than traditional postprocessing methods. We quantitatively and qualitatively demonstrate the benefits of the proposed algorithms in the acquisition of geometry in motion. Our pipeline executes in 1.1ms leveraging modern GPUs and off-the-shelf cameras and illumination components. We show how the sensor can be employed in many different applications, from [non-]rigid reconstructions to hand/face tracking. Further, we show many advantages over existing state of the art depth camera technologies beyond framerate, including latency, motion artifacts, multi-path errors, and multi-sensor interference. Adarsh Kowdle, Christoph Rhemann, Sean Ryan Fanello, Andrea Tagliasacchi, Jonathan Taylor 0001, Philip Davidson, Mingsong Dou, Cem Keskin, Sameh Khamis, David Kim 0002, Danhang Tang, Vladimir Tankovich, Julien P. C. Valentin, Shahram Izadi |
ACM Trans. Graph. | 5 |
| 2018 | LookinGood: enhancing performance capture with real-time neural re-renderingabstractMotivated by augmented and virtual reality applications such as telepresence, there has been a recent focus in real-time performance capture of humans under motion. However, given the real-time constraint, these systems often suffer from artifacts in geometry and texture such as holes and noise in the final rendering, poor lighting, and low-resolution textures. We take the novel approach to augment such real-time performance capture systems with a deep architecture that takes a rendering from an arbitrary viewpoint, and jointly performs completion, super resolution, and denoising of the imagery in real-time. We call this approach neural (re-)rendering , and our live system "LookinGood". Our deep architecture is trained to produce high resolution and high quality images from a coarse rendering in real-time. First, we propose a self-supervised training method that does not require manual ground-truth annotation. We contribute a specialized reconstruction error that uses semantic information to focus on relevant parts of the subject, e.g. the face. We also introduce a salient reweighing scheme of the loss function that is able to discard outliers. We specifically design the system for virtual and augmented reality headsets where the consistency between the left and right eye plays a crucial role in the final user experience. Finally, we generate temporally stable results by explicitly minimizing the difference between two consecutive frames. We tested the proposed system in two different scenarios: one involving a single RGB-D sensor, and upper body reconstruction of an actor, the second consisting of full body 360° capture. Through extensive experimentation, we demonstrate how our system generalizes across unseen sequences and subjects. Ricardo Martin-Brualla, Rohit Pandey, Shuoran Yang, Pavel Pidlypenskyi, Jonathan Taylor 0001, Julien P. C. Valentin, Sameh Khamis, Philip Davidson, Anastasia Tkach, Peter Lincoln, Adarsh Kowdle, Christoph Rhemann, Dan B. Goldman, Cem Keskin, Steven M. Seitz, Shahram Izadi, Sean Ryan Fanello |
ACM Trans. Graph. | 5 |
| 2018 | Real-time compression and streaming of 4D performancesabstractWe introduce a realtime compression architecture for 4D performance capture that is two orders of magnitude faster than current state-of-the-art techniques, yet achieves comparable visual quality and bitrate. We note how much of the algorithmic complexity in traditional 4D compression arises from the necessity to encode geometry using an explicit model (i.e. a triangle mesh). In contrast, we propose an encoder that leverages an implicit representation (namely a Signed Distance Function) to represent the observed geometry, as well as its changes through time. We demonstrate how SDFs, when defined over a small local region (i.e. a block), admit a low-dimensional embedding due to the innate geometric redundancies in their representation. We then propose an optimization that takes a Truncated SDF (i.e. a TSDF), such as those found in most rigid/non-rigid reconstruction pipelines, and efficiently projects each TSDF block onto the SDF latent space. This results in a collection of low entropy tuples that can be effectively quantized and symbolically encoded. On the decoder side, to avoid the typical artifacts of block-based coding, we also propose a variational optimization that compensates for quantization residuals in order to penalize unsightly discontinuities in the decompressed signal. This optimization is expressed in the SDF latent embedding, and hence can also be performed efficiently. We demonstrate our compression/decompression architecture by realizing, to the best of our knowledge, the first system for streaming a real-time captured 4D performance on consumer-level networks. Danhang Tang, Mingsong Dou, Peter Lincoln, Philip Davidson, Jonathan Taylor 0001, Sean Ryan Fanello, Cem Keskin, Adarsh Kowdle, Sofien Bouaziz, Shahram Izadi, Andrea Tagliasacchi |
ACM Trans. Graph. | 6 |
| 2017 | Articulated distance fields for ultra-fast tracking of hands interactingabstractThe state of the art in articulated hand tracking has been greatly advanced by hybrid methods that fit a generative hand model to depth data, leveraging both temporally and discriminatively predicted starting poses. In this paradigm, the generative model is used to define an energy function and a local iterative optimization is performed from these starting poses in order to find a "good local minimum" (i.e. a local minimum close to the true pose). Performing this optimization quickly is key to exploring more starting poses, performing more iterations and, crucially, exploiting high frame rates that ensure that temporally predicted starting poses are in the basin of convergence of a good local minimum. At the same time, a detailed and accurate generative model tends to deepen the good local minima and widen their basins of convergence. Recent work, however, has largely had to trade-off such a detailed hand model with one that facilitates such rapid optimization. We present a new implicit model of hand geometry that mostly avoids this compromise and leverage it to build an ultra-fast hybrid hand tracking system. Specifically, we construct an articulated signed distance function that, for any pose, yields a closed form calculation of both the distance to the detailed surface geometry and the necessary derivatives to perform gradient based optimization. There is no need to introduce or update any explicit "correspondences" yielding a simple algorithm that maps well to parallel hardware such as GPUs. As a result, our system can run at extremely high frame rates (e.g. up to 1000fps). Furthermore, we demonstrate how to detect, segment and optimize for two strongly interacting hands, recovering complex interactions at extremely high framerates. In the absence of publicly available datasets of sufficiently high frame rate, we leverage a multiview capture system to create a new 180fps dataset of one and two hands interacting together or with objects. Jonathan Taylor 0001, Vladimir Tankovich, Danhang Tang, Cem Keskin, David Kim 0002, Philip Davidson, Adarsh Kowdle, Shahram Izadi |
ACM Trans. Graph. | 1 |
| 2016 | Fits Like a Glove: Rapid and Reliable Hand Shape PersonalizationabstractWe present a fast, practical method for personalizing a hand shape basis to an individual user's detailed hand shape using only a small set of depth images. To achieve this, we minimize an energy based on a sum of render-and-compare cost functions called the golden energy. However, this energy is only piecewise continuous, due to pixels crossing occlusion boundaries, and is therefore not obviously amenable to efficient gradient-based optimization. A key insight is that the energy is the combination of a smooth low-frequency function with a high-frequency, low-amplitude, piecewisecontinuous function. A central finite difference approximation with a suitable step size can therefore jump over the discontinuities to obtain a good approximation to the energy's low-frequency behavior, allowing efficient gradient-based optimization. Experimental results quantitatively demonstrate for the first time that detailed personalized models improve the accuracy of hand tracking and achieve competitive results in both tracking and model registration. David Joseph Tan, Thomas J. Cashman 0001, Jonathan Taylor 0001, Andrew W. Fitzgibbon, Daniel Tarlow, Sameh Khamis, Shahram Izadi, Jamie Shotton |
CVPR | 3 |
| 2016 | ShadowHands: High-Fidelity Remote Hand Gesture Visualization using a Hand TrackerabstractThis paper presents ShadowHands - a novel technique for visualizing a remote user's hand gestures using a single depth sensor and hand tracking system. Previous work has shown that making distributed users better aware of each other's gestures facilitates remote collaboration. These systems presented virtual embodiments as a stream of raw 2D or 3D data -- this data is noisy, and requires high bandwidth and favorable camera positions. Instead, our work uses a hand tracker to capture gestures which we visualize with a high-fidelity hand model. Our system is practical, requiring only a single depth sensor placed below the screen, and can be used without per-user calibration. As we use a 3D model rather than raw data, we can augment the hand's appearance to improve saliency and aesthetics. We alpha-blend this visualization over a shared workspace, so the local user perceives the remote user's hand as if they were separated by a transparent display. We conducted an experiment to compare traditional hand embodiments with our new technique, showing a quantitative improvement in selection accuracy and qualitative improvements in feelings of mutual understanding and engagement. Erroll Wood, Jonathan Taylor 0001, John Fogarty, Andrew W. Fitzgibbon, Jamie Shotton |
ISS | 2 |
| 2016 | Fusion4D: real-time performance capture of challenging scenesabstractWe contribute a new pipeline for live multi-view performance capture, generating temporally coherent high-quality reconstructions in real-time. Our algorithm supports both incremental reconstruction, improving the surface estimation over time, as well as parameterizing the nonrigid scene motion. Our approach is highly robust to both large frame-to-frame motion and topology changes, allowing us to reconstruct extremely challenging scenes. We demonstrate advantages over related real-time techniques that either deform an online generated template or continually fuse depth data nonrigidly into a single reference model. Finally, we show geometric reconstruction results on par with offline methods which require orders of magnitude more processing time and many more RGBD cameras. Mingsong Dou, Sameh Khamis, Yury Degtyarev, Philip Davidson, Sean Ryan Fanello, Adarsh Kowdle, Sergio Orts, Christoph Rhemann, David Kim 0002, Jonathan Taylor 0001, Pushmeet Kohli, Vladimir Tankovich, Shahram Izadi |
ACM Trans. Graph. | 10 |
| 2016 | Efficient and precise interactive hand tracking through joint, continuous optimization of pose and correspondencesabstractFully articulated hand tracking promises to enable fundamentally new interactions with virtual and augmented worlds, but the limited accuracy and efficiency of current systems has prevented widespread adoption. Today's dominant paradigm uses machine learning for initialization and recovery followed by iterative model-fitting optimization to achieve a detailed pose fit. We follow this paradigm, but make several changes to the model-fitting, namely using: (1) a more discriminative objective function; (2) a smooth-surface model that provides gradients for non-linear optimization; and (3) joint optimization over both the model pose and the correspondences between observed data points and the model surface. While each of these changes may actually increase the cost per fitting iteration, we find a compensating decrease in the number of iterations. Further, the wide basin of convergence means that fewer starting points are needed for successful model fitting. Our system runs in real-time on CPU only, which frees up the commonly over-burdened GPU for experience designers. The hand tracker is efficient enough to run on low-power devices such as tablets. We can track up to several meters from the camera to provide a large working volume for interaction, even using the noisy data from current-generation depth cameras. Quantitative assessments on standard datasets show that the new approach exceeds the state of the art in accuracy. Qualitative results take the form of live recordings of a range of interactive experiences enabled by this new approach. Jonathan Taylor 0001, Lucas Bordeaux, Thomas J. Cashman 0001, Bob Corish, Cem Keskin, Toby Sharp, Eduardo Soto, David Sweeney, Julien P. C. Valentin, Benjamin Luff, Arran Topalian, Erroll Wood, Sameh Khamis, Pushmeet Kohli, Shahram Izadi, Richard Banks, Andrew W. Fitzgibbon, Jamie Shotton |
ACM Trans. Graph. | 1 |
| 2015 | Accurate, Robust, and Flexible Real-time Hand TrackingabstractWe present a new real-time hand tracking system based on a single depth camera. The system can accurately reconstruct complex hand poses across a variety of subjects. It also allows for robust tracking, rapidly recovering from any temporary failures. Most uniquely, our tracker is highly flexible, dramatically improving upon previous approaches which have focused on front-facing close-range scenarios. This flexibility opens up new possibilities for human-computer interaction with examples including tracking at distances from tens of centimeters through to several meters (for controlling the TV at a distance), supporting tracking using a moving depth camera (for mobile scenarios), and arbitrary camera placements (for VR headsets). These features are achieved through a new pipeline that combines a multi-layered discriminative reinitialization strategy for per-frame pose estimation, followed by a generative model-fitting stage. We provide extensive technical details and a detailed qualitative and quantitative analysis. Toby Sharp, Cem Keskin, Duncan P. Robertson, Jonathan Taylor 0001, Jamie Shotton, David Kim 0002, Christoph Rhemann, Ido Leichter, Alon Vinnikov, Daniel Freedman, Pushmeet Kohli, Eyal Krupka, Andrew W. Fitzgibbon, Shahram Izadi |
CHI | 4 |
| 2015 | 3D scanning deformable objects with a single RGBD sensorabstractWe present a 3D scanning system for deformable objects that uses only a single Kinect sensor. Our work allows considerable amount of nonrigid deformations during scanning, and achieves high quality results without heavily constraining user or camera motion. We do not rely on any prior shape knowledge, enabling general object scanning with freeform deformations. To deal with the drift problem when nonrigidly aligning the input sequence, we automatically detect loop closures, distribute the alignment error over the loop, and finally use a bundle adjustment algorithm to optimize for the latent 3D shape and nonrigid deformation parameters simultaneously. We demonstrate high quality scanning results in some challenging sequences, comparing with state of art nonrigid techniques, as well as ground truth data. Mingsong Dou, Jonathan Taylor 0001, Henry Fuchs, Andrew W. Fitzgibbon, Shahram Izadi |
CVPR | 2 |
| 2015 | Large-scale and drift-free surface reconstruction using online subvolume registrationabstractDepth cameras have helped commoditize 3D digitization of the real-world. It is now feasible to use a single Kinect-like camera to scan in an entire building or other large-scale scenes. At large scales, however, there is an inherent chal-lenge of dealing with distortions and drift due to accumu-lated pose estimation errors. Existing techniques suffer from one or more of the following: a) requiring an expensive offline global optimization step taking hours to compute; b) needing a full second pass over the input depth frames to correct for accumulated errors; c) relying on RGB data alongside depth data to optimize poses; or d) requiring the user to create explicit loop closures to allow gross alignment errors to be resolved. In this paper, we present a method that addresses all of these issues. Our method supports online model correction, without needing to reprocess or store any input depth data. Even while performing global correction of a large 3D model, our method takes only minutes rather than hours to compute. Our model does not require any explicit loop closures to be detected and, finally, relies on depth data alone, allowing operation in low-lighting conditions. We show qualitative results on many large scale scenes, high-lighting the lack of error and drift in our reconstructions. We compare to state of the art techniques and demonstrate large-scale dense surface reconstruction “in the dark”, a capability not offered by RGB-D techniques. 1. Nicola Fioraio, Jonathan Taylor 0001, Andrew W. Fitzgibbon, Luigi Di Stefano, Shahram Izadi |
CVPR | 2 |
| 2015 | Learning an efficient model of hand shape variation from depth imagesabstractWe describe how to learn a compact and efficient model of the surface deformation of human hands. The model is built from a set of noisy depth images of a diverse set of subjects performing different poses with their hands. We represent the observed surface using Loop subdivision of a control mesh that is deformed by our learned parametric shape and pose model. The model simultaneously accounts for variation in subject-specific shape and subject-agnostic pose. Specifically, hand shape is parameterized as a linear combination of a mean mesh in a neutral pose with a small number of offset vectors. This mesh is then articulated using standard linear blend skinning (LBS) to generate the control mesh of a subdivision surface. We define an energy that encourages each depth pixel to be explained by our model, and the use of a smooth subdivision surface allows us to optimize for all parameters jointly from a rough initialization. The efficacy of our method is demonstrated using both synthetic and real data, where it is shown that hand shape variation can be represented using only a small number of basis components. We compare with other approaches including PCA and show a substantial improvement in the representational power of our model, while maintaining the efficiency of a linear shape basis. Sameh Khamis, Jonathan Taylor 0001, Jamie Shotton, Cem Keskin, Shahram Izadi, Andrew W. Fitzgibbon |
CVPR | 2 |
| 2015 | Opening the Black Box: Hierarchical Sampling Optimization for Estimating Human Hand PoseabstractWe address the problem of hand pose estimation, formulated as an inverse problem. Typical approaches optimize an energy function over pose parameters using a 'black box' image generation procedure. This procedure knows little about either the relationships between the parameters or the form of the energy function. In this paper, we show that we can significantly improving upon black box optimization by exploiting high-level knowledge of the structure of the parameters and using a local surrogate energy function. Our new framework, hierarchical sampling optimization, consists of a sequence of predictors organized into a kinematic hierarchy. Each predictor is conditioned on its ancestors, and generates a set of samples over a subset of the pose parameters. The highly-efficient surrogate energy is used to select among samples. Having evaluated the full hierarchy, the partial pose samples are concatenated to generate a full-pose hypothesis. Several hypotheses are generated using the same procedure, and finally the original full energy function selects the best result. Experimental evaluation on three publically available datasets show that our method is particularly impressive in low-compute scenarios where it significantly outperforms all other state-of-the-art methods. Danhang Tang, Jonathan Taylor 0001, Pushmeet Kohli, Cem Keskin, Tae-Kyun Kim 0001, Jamie Shotton |
ICCV | 2 |
| 2015 | Metric Regression Forests for Correspondence Estimation
Gerard Pons-Moll, Jonathan Taylor 0001, Jamie Shotton, Aaron Hertzmann, Andrew W. Fitzgibbon |
Int. J. Comput. Vis. | 2 |
| 2014 | Real-Time Face Reconstruction from a Single Depth ImageabstractThis paper contributes a real time method for recovering facial shape and expression from a single depth image. The method also estimates an accurate and dense correspondence field between the input depth image and a generic face model. Both outputs are a result of minimizing the error in reconstructing the depth image, achieved by applying a set of identity and expression blend shapes to the model. Traditionally, such a generative approach has shown to be computationally expensive and non-robust because of the non-linear nature of the reconstruction error. To overcome this problem, we use a discriminatively trained prediction pipeline that employs random forests to generate an initial dense but noisy correspondence field. Our method then exploits a fast ICP-like approximation to update these correspondences, allowing us to quickly obtain a robust initial fit of our model. The model parameters are then fine tuned to minimize the true reconstruction error using a stochastic optimization technique. The correspondence field resulting from our hybrid generative-discriminative pipeline is accurate and useful for a variety of applications such as mesh deformation and retexturing. Our method works in real-time on a single depth image i.e. Without temporal tracking, is free from per-user calibration, and works in low-light conditions. Vahid Kazemi, Cem Keskin, Jonathan Taylor 0001, Pushmeet Kohli, Shahram Izadi |
3DV | 3 |
| 2014 | User-Specific Hand Modeling from Monocular Depth SequencesabstractThis paper presents a method for acquiring dense nonrigid shape and deformation from a single monocular depth sensor. We focus on modeling the human hand, and assume that a single rough template model is available. We combine and extend existing work on model-based tracking, subdivision surface fitting, and mesh deformation to acquire detailed hand models from as few as 15 frames of depth data. We propose an objective that measures the error of fit between each sampled data point and a continuous model surface defined by a rigged control mesh, and uses as-rigid-as-possible (ARAP) regularizers to cleanly separate the model and template geometries. A key contribution is our use of a smooth model based on subdivision surfaces that allows simultaneous optimization over both correspondences and model parameters. This avoids the use of iterated closest point (ICP) algorithms which often lead to slow convergence. Automatic initialization is obtained using a regression forest trained to infer approximate correspondences. Experiments show that the resulting meshes model the user's hand shape more accurately than just adapting the shape parameters of the skeleton, and that the retargeted skeleton accurately models the user's articulations. We investigate the effect of various modeling choices, and show the benefits of using subdivision surfaces and ARAP regularization. Jonathan Taylor 0001, Richard V. Stebbing, Varun Ramakrishna, Cem Keskin, Jamie Shotton, Shahram Izadi, Aaron Hertzmann, Andrew W. Fitzgibbon |
CVPR | 1 |
| 2014 | FlexSense: a transparent self-sensing deformable surfaceabstractWe present FlexSense, a new thin-film, transparent sensing surface based on printed piezoelectric sensors, which can reconstruct complex deformations without the need for any external sensing, such as cameras. FlexSense provides a fully self-contained setup which improves mobility and is not affected from occlusions. Using only a sparse set of sensors, printed on the periphery of the surface substrate, we devise two new algorithms to fully reconstruct the complex deformations of the sheet, using only these sparse sensor measurements. An evaluation shows that both proposed algorithms are capable of reconstructing complex deformations accurately. We demonstrate how FlexSense can be used for a variety of 2.5D interactions, including as a transparent cover for tablets where bending can be performed alongside touch to enable magic lens style effects, layered input, and mode switching, as well as the ability to use our device as a high degree-of-freedom input controller for gaming and beyond. Christian Rendl, David Kim 0002, Sean Ryan Fanello, Patrick Parzer, Christoph Rhemann, Jonathan Taylor 0001, Martin Zirkl, Gregor Scheipl, Thomas Rothländer, Michael Haller, Shahram Izadi |
UIST | 6 |
| 2013 | Metric Regression Forests for Human Pose EstimationabstractTraditionally, human pose estimation algorithms could be classified into generative [2] and discriminative [4] approaches. Generative approaches model the likelihood of the observations given a pose estimate, however, they are susceptible to local minima and thus require good initial pose estimates. Discriminative approaches learn a direct mapping from image features to pose space from training data, however, they struggle to generalize to unseen poses. Building on previous work [3], Taylor et al. [5] bypass some of these limitations using a hybrid-approach that discriminatively predicts, for each pixel in a depth image, a corresponding point on the surface of a humanoid mesh model. This mesh model is then robustly fit to the resulting set of correspondences using local optimization. Surprisingly though, these correspondences are actually inferred using a random forest whose structure was trained using a classification objective that arbitrarily equates target model points belonging to the same predefined body part [3]. In this paper, we address Taylor et al.’s use of this proxy classification objective by proposing Metric Space Information Gain (MSIG), a replacement objective function for training a random forest to directly minimize the uncertainty over the target model points, naturally encoding the correlation between these points as a function of the geodesic distance. To this end, we view the surface of the model U as a metric space (U,dU) defined by the geodesic distance metric dU (see first panel of Figure 1). The natural objective function to minimize the uncertainty in the resulting true distributions that result from a split function s in such a space, is the information gain I(s) [1]. This is generally approximated using an empirical distribution Q = {ui} ⊆U drawn from the true unsplit distribution pU as I(s)≈ I(s;Q) = Ĥ(Q)− ∑ i∈{L,R} |Qi| |Q| Ĥ(Qi), (1) Gerard Pons-Moll, Jonathan Taylor 0001, Jamie Shotton, Aaron Hertzmann, Andrew W. Fitzgibbon |
BMVC | 2 |
| 2012 | The Vitruvian manifold: Inferring dense correspondences for one-shot human pose estimationabstractFitting an articulated model to image data is often approached as an optimization over both model pose and model-to-image correspondence. For complex models such as humans, previous work has required a good initialization, or an alternating minimization between correspondence and pose. In this paper we investigate one-shot pose estimation: can we directly infer correspondences using a regression function trained to be invariant to body size and shape, and then optimize the model pose just once? We evaluate on several challenging single-frame data sets containing a wide variety of body poses, shapes, torso rotations, and image cropping. Our experiments demonstrate that one-shot pose estimation achieves state of the art results and runs in real-time. Jonathan Taylor 0001, Jamie Shotton, Toby Sharp, Andrew W. Fitzgibbon |
CVPR | 1 |