Julien P. C. Valentin

dblp:135/4936 · DBLP profile ↗
← Back
25ranked-venue papers
5as first author
4since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 22 · 5 first-author · 4 since 2021Artificial intelligence and machine learning · 14 · 2 first-author · 3 since 2021Systems, architecture and hardware · 1Human-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
14 papers
3D vision · 75% Robot navigation and mapping · 11% Video understanding and tracking · 8%
Computer graphics and multimedia
8 papers
Rendering · 33% Geometric modeling and processing · 24% Virtual and augmented reality · 21%
Human-computer interaction and pervasive computing
3 papers
Interaction techniques and input · 46% Ubiquitous computing and smart environments · 40% Immersive interaction · 14%

Topics — the 30 heaviest of 38, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision › stereo vision
stereo matching
1.242018
ActiveStereoNet: End-to-End Self-supervised Learning for Active Stereo Systems · ECCV (8) 2018
StereoNet: Guided Hierarchical Refinement for Real-Time Edge-Aware Depth Prediction · ECCV (15) 2018
Low Compute and Fully Parallel Computer Vision with HashMatch · ICCV 2017
Computer vision › 3D vision
depth estimation
0.932018
The need 4 speed in real-time dense visual tracking · ACM Trans. Graph. 2018
StereoNet: Guided Hierarchical Refinement for Real-Time Edge-Aware Depth Prediction · ECCV (15) 2018
UltraStereo: Efficient Learning-Based Matching for Active Stereo Systems · CVPR 2017
Rendering
neural rendering
0.822021
FastNeRF: High-Fidelity Neural Rendering at 200FPS · ICCV 2021
LookinGood: enhancing performance capture with real-time neural re-rendering · ACM Trans. Graph. 2018
Geometric modeling and processing
model fitting
0.822022
Learning to Fit Morphable Models · ECCV (6) 2022
Efficient and precise interactive hand tracking through joint, continuous optimization of pose and correspondences · ACM Trans. Graph. 2016
Computer vision › 3D vision › stereo vision
active stereo
0.622018
ActiveStereoNet: End-to-End Self-supervised Learning for Active Stereo Systems · ECCV (8) 2018
UltraStereo: Efficient Learning-Based Matching for Active Stereo Systems · CVPR 2017
Computer vision › 3D vision
3d face reconstruction
0.612022
3D Face Reconstruction with Dense Landmarks · ECCV (13) 2022
Geometric modeling and processing
3d morphable model
0.612022
Learning to Fit Morphable Models · ECCV (6) 2022
Robotics › Robot navigation and mapping
SLAM
0.522020
Real-Time RGB-D Camera Pose Estimation in Novel Scenes Using a Relocalisation Cascade · IEEE Trans. Pattern Anal. Mach. Intell. 2020
On-the-Fly Adaptation of Regression Forests for Online Camera Relocalisation · CVPR 2017
Computer vision › 3D vision › visual localization
camera relocalization
0.522017
On-the-Fly Adaptation of Regression Forests for Online Camera Relocalisation · CVPR 2017
Exploiting uncertainty in regression forests for accurate camera relocalization · CVPR 2015
Rendering
neural radiance fields
0.512021
FastNeRF: High-Fidelity Neural Rendering at 200FPS · ICCV 2021
Rendering
real-time rendering
0.512021
FastNeRF: High-Fidelity Neural Rendering at 200FPS · ICCV 2021
Computer vision › 3D vision
camera pose estimation
0.412020
Real-Time RGB-D Camera Pose Estimation in Novel Scenes Using a Relocalisation Cascade · IEEE Trans. Pattern Anal. Mach. Intell. 2020
Robotics › Robot navigation and mapping › localization › global localization
relocalization
0.412020
Real-Time RGB-D Camera Pose Estimation in Novel Scenes Using a Relocalisation Cascade · IEEE Trans. Pattern Anal. Mach. Intell. 2020
Computer vision › 3D vision
3d scene understanding
0.422015
SemanticPaint: Interactive 3D Labeling and Learning at your Fingertips · ACM Trans. Graph. 2015
Mesh Based Semantic Modelling for Indoor and Outdoor Scenes · CVPR 2013
Computer vision › 3D vision
3d reconstruction
0.322018
Holoportation: Virtual 3D Teleportation in Real-time · UIST 2016
The need 4 speed in real-time dense visual tracking · ACM Trans. Graph. 2018
Virtual and augmented reality
telepresence
0.322018
Holoportation: Virtual 3D Teleportation in Real-time · UIST 2016
LookinGood: enhancing performance capture with real-time neural re-rendering · ACM Trans. Graph. 2018
Computer vision › Video understanding and tracking › motion tracking
dense tracking
0.312018
The need 4 speed in real-time dense visual tracking · ACM Trans. Graph. 2018
Computer vision › 3D vision › depth estimation
monocular depth estimation
0.312018
Depth from motion for smartphone AR · ACM Trans. Graph. 2018
Computer vision › Video understanding and tracking
object tracking
0.312018
The need 4 speed in real-time dense visual tracking · ACM Trans. Graph. 2018
Computer vision › 3D vision › stereo vision › stereo matching › deep stereo matching
self-supervised stereo matching
0.312018
ActiveStereoNet: End-to-End Self-supervised Learning for Active Stereo Systems · ECCV (8) 2018
Virtual and augmented reality › tracking
6DOF tracking
0.312018
Depth from motion for smartphone AR · ACM Trans. Graph. 2018
Image and video processing
image restoration
0.312018
LookinGood: enhancing performance capture with real-time neural re-rendering · ACM Trans. Graph. 2018
Virtual and augmented reality › augmented reality
mobile augmented reality
0.312018
Depth from motion for smartphone AR · ACM Trans. Graph. 2018
Image and video processing › image restoration › multi-task image restoration
super-resolution and denoising
0.312018
LookinGood: enhancing performance capture with real-time neural re-rendering · ACM Trans. Graph. 2018
Computer vision › Segmentation and scene understanding › dense prediction
pixel labeling
0.312017
Low Compute and Fully Parallel Computer Vision with HashMatch · ICCV 2017
Computer vision › 3D vision › depth estimation
stereo depth estimation
0.312017
Low Compute and Fully Parallel Computer Vision with HashMatch · ICCV 2017
Multimedia analysis and retrieval
image retrieval
0.312017
Low Compute and Fully Parallel Computer Vision with HashMatch · ICCV 2017
Computer vision › 3D vision › 3d reconstruction › volumetric reconstruction
real-time volumetric reconstruction
0.212016
Holoportation: Virtual 3D Teleportation in Real-time · UIST 2016
Virtual and augmented reality › telepresence
3d telepresence
0.212016
Holoportation: Virtual 3D Teleportation in Real-time · UIST 2016
Interaction techniques and input › input sensing › tracking
hand tracking
0.212016
Efficient and precise interactive hand tracking through joint, continuous optimization of pose and correspondences · ACM Trans. Graph. 2016

Methods — techniques the papers use, named apart from their topics

differentiable rendering · 1.1deep learning · 1.1regression forest · 0.7self-supervised learning · 0.7dense landmark regression · 0.6radiance map caching · 0.5neural radiance field · 0.5relocalization cascade · 0.4RANSAC · 0.4deep neural network · 0.4color gradient illumination · 0.4temporal consistency · 0.3semantic reconstruction error · 0.3monocular depth computation · 0.3hierarchical refinement · 0.3edge-aware filtering · 0.3active stereo · 0.3smooth-surface model · 0.2
YearPublicationVenuePosition
2023 DigiFace-1M: 1 Million Digital Face Images for Face Recognition
abstract
State-of-the-art face recognition models show impressive accuracy, achieving over 99.8% on Labeled Faces in the Wild (LFW) dataset. Such models are trained on large-scale datasets that contain millions of real human face images collected from the internet. Web-crawled face images are severely biased (in terms of race, lighting, makeup, etc) and often contain label noise. More importantly, the face images are collected without explicit consent, raising ethical concerns. To avoid such problems, we introduce a large-scale synthetic dataset for face recognition, obtained by rendering digital faces using a computer graphics pipeline1. We first demonstrate that aggressive data augmentation can significantly reduce the synthetic-to-real domain gap. Having full control over the rendering pipeline, we also study how each attribute (e.g., variation in facial pose, accessories and textures) affects the accuracy. Compared to Syn-Face, a recent method trained on GAN-generated synthetic faces, we reduce the error rate on LFW by 52.5% (accuracy from 91.93% to 96.17%). By fine-tuning the network on a smaller number of real face images that could reason-ably be obtained with consent, we achieve accuracy that is comparable to the methods trained on millions of real face images.
Gwangbin Bae, Martin de La Gorce, Tadas Baltrusaitis, Charlie Hewitt, Dong Chen 0003, Julien P. C. Valentin, Roberto Cipolla, Jingjing Shen
WACV6
2022 Learning to Fit Morphable Models
Vasileios Choutas, Federica Bogo, Jingjing Shen, Julien P. C. Valentin
ECCV (6)4
2022 3D Face Reconstruction with Dense Landmarks
Erroll Wood, Tadas Baltrusaitis, Charlie Hewitt, Matthew Johnson 0003, Jingjing Shen, Nikola Milosavljevic, Daniel Wilde, Stephan J. Garbin, Toby Sharp, Ivan Stojiljkovic, Thomas J. Cashman 0001, Julien P. C. Valentin
ECCV (13)12
2021 FastNeRF: High-Fidelity Neural Rendering at 200FPS
abstract
Recent work on Neural Radiance Fields (NeRF) showed how neural networks can be used to encode complex 3D environments that can be rendered photorealistically from novel viewpoints. Rendering these images is very computationally demanding and recent improvements are still a long way from enabling interactive rates, even on high-end hardware. Motivated by scenarios on mobile and mixed reality devices, we propose FastNeRF, the first NeRF-based system capable of rendering high fidelity photorealistic images at 200Hz on a high-end consumer GPU. The core of our method is a graphics-inspired factorization that allows for (i) compactly caching a deep radiance map at each position in space, (ii) efficiently querying that map using ray directions to estimate the pixel values in the rendered image. Extensive experiments show that the proposed method is 3000 times faster than the original NeRF algorithm and at least an order of magnitude faster than existing work on accelerating NeRF, while maintaining visual quality and extensibility.
Stephan J. Garbin, Marek Kowalski, Matthew Johnson 0003, Jamie Shotton, Julien P. C. Valentin
ICCV5
2020 Real-Time RGB-D Camera Pose Estimation in Novel Scenes Using a Relocalisation Cascade
abstract
Camera pose estimation is an important problem in computer vision, with applications as diverse as simultaneous localisation and mapping, virtual/augmented reality and navigation. Common techniques match the current image against keyframes with known poses coming from a tracker, directly regress the pose, or establish correspondences between keypoints in the current image and points in the scene in order to estimate the pose. In recent years, regression forests have become a popular alternative to establish such correspondences. They achieve accurate results, but have traditionally needed to be trained offline on the target scene, preventing relocalisation in new environments. Recently, we showed how to circumvent this limitation by adapting a pre-trained forest to a new scene on the fly. The adapted forests achieved relocalisation performance that was on par with that of offline forests, and our approach was able to estimate the camera pose in close to real time, which made it desirable for systems that require online relocalisation. In this paper, we present an extension of this work that achieves significantly better relocalisation performance whilst running fully in real time. To achieve this, we make several changes to the original approach: (i) instead of simply accepting the camera pose hypothesis produced by RANSAC without question, we make it possible to score the final few hypotheses it considers using a geometric approach and select the most promising one; (ii) we chain several instantiations of our relocaliser (with different parameter settings) together in a cascade, allowing us to try faster but less accurate relocalisation first, only falling back to slower, more accurate relocalisation as necessary; and (iii) we tune the parameters of our cascade, and the individual relocalisers it contains, to achieve effective overall performance. Taken together, these changes allow us to significantly improve upon the performance our original state-of-the-art method was able to achieve on the well-known 7-Scenes and Stanford 4 Scenes benchmarks. As additional contributions, we present a novel way of visualising the internal behaviour of our forests, and use the insights gleaned from this to show how to entirely circumvent the need to pre-train a forest on a generic scene.
Tommaso Cavallari, Stuart Golodetz, Nicholas A. Lord, Julien P. C. Valentin, Victor Adrian Prisacariu, Luigi Di Stefano, Philip Torr 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2019 LSMAT Least Squares Medial Axis Transform
abstract
Abstract The medial axis transform has applications in numerous fields including visualization, computer graphics, and computer vision. Unfortunately, traditional medial axis transformations are usually brittle in the presence of outliers, perturbations and/or noise along the boundary of objects. To overcome this limitation, we introduce a new formulation of the medial axis transform which is naturally robust in the presence of these artefacts. Unlike previous work which has approached the medial axis from a computational geometry angle, we consider it from a numerical optimization perspective. In this work, we follow the definition of the medial axis transform as ‘the set of maximally inscribed spheres’. We show how this definition can be formulated as a least squares relaxation where the transform is obtained by minimizing a continuous optimization problem. The proposed approach is inherently parallelizable by performing independent optimization of each sphere using Gauss–Newton, and its least‐squares form allows it to be significantly more robust compared to traditional computational geometry approaches. Extensive experiments on 2D and 3D objects demonstrate that our method provides superior results to the state of the art on both synthetic and real‐data.
Daniel Rebain, Baptiste Angles, Julien P. C. Valentin, Nicholas Vining, Jiju Poovvancheri, Shahram Izadi, Andrea Tagliasacchi
Comput. Graph. Forum3
2019 Deep reflectance fields: high-quality facial reflectance field inference from color gradient illumination
abstract
We present a novel technique to relight images of human faces by learning a model of facial reflectance from a database of 4D reflectance field data of several subjects in a variety of expressions and viewpoints. Using our learned model, a face can be relit in arbitrary illumination environments using only two original images recorded under spherical color gradient illumination. The output of our deep network indicates that the color gradient images contain the information needed to estimate the full 4D reflectance field, including specular reflections and high frequency details. While capturing spherical color gradient illumination still requires a special lighting setup, reduction to just two illumination conditions allows the technique to be applied to dynamic facial performance capture. We show side-by-side comparisons which demonstrate that the proposed system outperforms the state-of-the-art techniques in both realism and speed.
Abhimitra Meka, Christian Häne, Rohit Pandey, Michael Zollhöfer, Sean Ryan Fanello, Graham Fyffe, Adarsh Kowdle, Xueming Yu, Jay Busch, Jason Dourgarian, Peter Denny, Sofien Bouaziz, Peter Lincoln, Matt Whalen, Geoff Harvey, Jonathan Taylor 0001, Shahram Izadi, Andrea Tagliasacchi, Paul E. Debevec, Christian Theobalt, Julien P. C. Valentin, Christoph Rhemann
ACM Trans. Graph.21
2018 StereoNet: Guided Hierarchical Refinement for Real-Time Edge-Aware Depth Prediction
Sameh Khamis, Sean Ryan Fanello, Christoph Rhemann, Adarsh Kowdle, Julien P. C. Valentin, Shahram Izadi
ECCV (15)5
2018 ActiveStereoNet: End-to-End Self-supervised Learning for Active Stereo Systems
Yinda Zhang 0001, Sameh Khamis, Christoph Rhemann, Julien P. C. Valentin, Adarsh Kowdle, Vladimir Tankovich, Michael Schoenberg, Shahram Izadi, Thomas A. Funkhouser, Sean Ryan Fanello
ECCV (8)4
2018 SOS: Stereo Matching in O(1) with Slanted Support Windows
abstract
Depth cameras have accelerated research in many areas of computer vision. Most triangulation-based depth cameras, whether structured light systems like the Kinect or active (assisted) stereo systems, are based on the principle of stereo matching. Depth from stereo is an active research topic dating back 30 years. Despite recent advances, algorithms usually trade-off accuracy for speed. In particular, efficient methods rely on fronto-parallel assumptions to reduce the search space and keep computation low. We present SOS (Slanted O(1) Stereo), the first algorithm capable of leveraging slanted support windows without sacrificing speed or accuracy. We use an active stereo configuration, where an illuminator textures the scene. Under this setting, local methods - such as PatchMatch Stereo - obtain state of the art results by jointly estimating disparities and slant, but at a large computational cost. We observe that these methods typically exploit local smoothness to simplify their initialization strategies. Our key insight is that local smoothness can in fact be used to amortize the computation not only within initialization, but across the entire stereo pipeline. Building on these insights, we propose a novel hierarchical initialization that is able to efficiently perform search over disparity and slants. We then show how this structure can be leveraged to provide high quality depth maps. Extensive quantitative evaluations demonstrate that the proposed technique yields significantly more precise results than current state of the art, but at a fraction of the computational cost. Our prototype implementation runs at 4000 fps on modern GPU architectures.
Vladimir Tankovich, Michael Schoenberg, Sean Ryan Fanello, Adarsh Kowdle, Christoph Rhemann, Maksym Dzitsiuk, Mirko Schmidt, Julien P. C. Valentin, Shahram Izadi
IROS8
2018 The need 4 speed in real-time dense visual tracking
abstract
The advent of consumer depth cameras has incited the development of a new cohort of algorithms tackling challenging computer vision problems. The primary reason is that depth provides direct geometric information that is largely invariant to texture and illumination. As such, substantial progress has been made in human and object pose estimation, 3D reconstruction and simultaneous localization and mapping. Most of these algorithms naturally benefit from the ability to accurately track the pose of an object or scene of interest from one frame to the next. However, commercially available depth sensors (typically running at 30fps) can allow for large inter-frame motions to occur that make such tracking problematic. A high frame rate depth camera would thus greatly ameliorate these issues, and further increase the tractability of these computer vision problems. Nonetheless, the depth accuracy of recent systems for high-speed depth estimation [Fanello et al. 2017b] can degrade at high frame rates. This is because the active illumination employed produces a low SNR and thus a high exposure time is required to obtain a dense accurate depth image. Furthermore in the presence of rapid motion, longer exposure times produce artifacts due to motion blur, and necessitates a lower frame rate that introduces large inter-frame motion that often yield tracking failures. In contrast, this paper proposes a novel combination of hardware and software components that avoids the need to compromise between a dense accurate depth map and a high frame rate. We document the creation of a full 3D capture system for high speed and quality depth estimation, and demonstrate its advantages in a variety of tracking and reconstruction tasks. We extend the state of the art active stereo algorithm presented in Fanello et al. [2017b] by adding a space-time feature in the matching phase. We also propose a machine learning based depth refinement step that is an order of magnitude faster than traditional postprocessing methods. We quantitatively and qualitatively demonstrate the benefits of the proposed algorithms in the acquisition of geometry in motion. Our pipeline executes in 1.1ms leveraging modern GPUs and off-the-shelf cameras and illumination components. We show how the sensor can be employed in many different applications, from [non-]rigid reconstructions to hand/face tracking. Further, we show many advantages over existing state of the art depth camera technologies beyond framerate, including latency, motion artifacts, multi-path errors, and multi-sensor interference.
Adarsh Kowdle, Christoph Rhemann, Sean Ryan Fanello, Andrea Tagliasacchi, Jonathan Taylor 0001, Philip Davidson, Mingsong Dou, Cem Keskin, Sameh Khamis, David Kim 0002, Danhang Tang, Vladimir Tankovich, Julien P. C. Valentin, Shahram Izadi
ACM Trans. Graph.14
2018 LookinGood: enhancing performance capture with real-time neural re-rendering
abstract
Motivated by augmented and virtual reality applications such as telepresence, there has been a recent focus in real-time performance capture of humans under motion. However, given the real-time constraint, these systems often suffer from artifacts in geometry and texture such as holes and noise in the final rendering, poor lighting, and low-resolution textures. We take the novel approach to augment such real-time performance capture systems with a deep architecture that takes a rendering from an arbitrary viewpoint, and jointly performs completion, super resolution, and denoising of the imagery in real-time. We call this approach neural (re-)rendering , and our live system "LookinGood". Our deep architecture is trained to produce high resolution and high quality images from a coarse rendering in real-time. First, we propose a self-supervised training method that does not require manual ground-truth annotation. We contribute a specialized reconstruction error that uses semantic information to focus on relevant parts of the subject, e.g. the face. We also introduce a salient reweighing scheme of the loss function that is able to discard outliers. We specifically design the system for virtual and augmented reality headsets where the consistency between the left and right eye plays a crucial role in the final user experience. Finally, we generate temporally stable results by explicitly minimizing the difference between two consecutive frames. We tested the proposed system in two different scenarios: one involving a single RGB-D sensor, and upper body reconstruction of an actor, the second consisting of full body 360° capture. Through extensive experimentation, we demonstrate how our system generalizes across unseen sequences and subjects.
Ricardo Martin-Brualla, Rohit Pandey, Shuoran Yang, Pavel Pidlypenskyi, Jonathan Taylor 0001, Julien P. C. Valentin, Sameh Khamis, Philip Davidson, Anastasia Tkach, Peter Lincoln, Adarsh Kowdle, Christoph Rhemann, Dan B. Goldman, Cem Keskin, Steven M. Seitz, Shahram Izadi, Sean Ryan Fanello
ACM Trans. Graph.6
2018 Depth from motion for smartphone AR
abstract
Augmented reality (AR) for smartphones has matured from a technology for earlier adopters, available only on select high-end phones, to one that is truly available to the general public. One of the key breakthroughs has been in low-compute methods for six degree of freedom (6DoF) tracking on phones using only the existing hardware (camera and inertial sensors). 6DoF tracking is the cornerstone of smartphone AR allowing virtual content to be precisely locked on top of the real world. However, to really give users the impression of believable AR, one requires mobile depth. Without depth, even simple effects such as a virtual object being correctly occluded by the real-world is impossible. However, requiring a mobile depth sensor would severely restrict the access to such features. In this article, we provide a novel pipeline for mobile depth that supports a wide array of mobile phones, and uses only the existing monocular color sensor. Through several technical contributions, we provide the ability to compute low latency dense depth maps using only a single CPU core of a wide range of (medium-high) mobile phones. We demonstrate the capabilities of our approach on high-level AR applications including real-time navigation and shopping.
Julien P. C. Valentin, Adarsh Kowdle, Jonathan T. Barron, Neal Wadhwa, Maksym Dzitsiuk, Michael Schoenberg, Ambrus Csaszar, Eric Turner 0001, Ivan Dryanovski, João Afonso, Jose Pascoal, Konstantine Tsotsos, Mira Leung, Mirko Schmidt, Onur G. Guleryuz, Sameh Khamis, Vladimir Tankovich, Sean Ryan Fanello, Shahram Izadi, Christoph Rhemann
ACM Trans. Graph.1
2017 On-the-Fly Adaptation of Regression Forests for Online Camera Relocalisation
abstract
Camera relocalisation is an important problem in computer vision, with applications in simultaneous localisation and mapping, virtual/augmented reality and navigation. Common techniques either match the current image against keyframes with known poses coming from a tracker, or establish 2D-to-3D correspondences between keypoints in the current image and points in the scene in order to estimate the camera pose. Recently, regression forests have become a popular alternative to establish such correspondences. They achieve accurate results, but must be trained offline on the target scene, preventing relocalisation in new environments. In this paper, we show how to circumvent this limitation by adapting a pre-trained forest to a new scene on the fly. Our adapted forests achieve relocalisation performance that is on par with that of offline forests, and our approach runs in under 150ms, making it desirable for real-time systems that require online relocalisation.
Tommaso Cavallari, Stuart Golodetz, Nicholas A. Lord, Julien P. C. Valentin, Luigi Di Stefano, Philip Torr 0001
CVPR4
2017 UltraStereo: Efficient Learning-Based Matching for Active Stereo Systems
abstract
Efficient estimation of depth from pairs of stereo images is one of the core problems in computer vision. We efficiently solve the specialized problem of stereo matching under active illumination using a new learning-based algorithm. This type of active stereo i.e. stereo matching where scene texture is augmented by an active light projector is proving compelling for designing depth cameras, largely due to improved robustness when compared to time of flight or traditional structured light techniques. Our algorithm uses an unsupervised greedy optimization scheme that learns features that are discriminative for estimating correspondences in infrared images. The proposed method optimizes a series of sparse hyperplanes that are used at test time to remap all the image patches into a compact binary representation in O(1). The proposed algorithm is cast in a PatchMatch Stereo-like framework, producing depth maps at 500Hz. In contrast to standard structured light methods, our approach generalizes to different scenes, does not require tedious per camera calibration procedures and is not adversely affected by interference from overlapping sensors. Extensive evaluations show we surpass the quality and overcome the limitations of current depth sensing technologies.
Sean Ryan Fanello, Julien P. C. Valentin, Christoph Rhemann, Adarsh Kowdle, Vladimir Tankovich, Philip Davidson, Shahram Izadi
CVPR2
2017 Low Compute and Fully Parallel Computer Vision with HashMatch
abstract
Numerous computer vision problems such as stereo depth estimation, object-class segmentation and fore-ground/background segmentation can be formulated as per-pixel image labeling tasks. Given one or many images as input, the desired output of these methods is usually a spatially smooth assignment of labels. The large amount of such computer vision problems has lead to significant research efforts, with the state of art moving from CRF-based approaches to deep CNNs and more recently, hybrids of the two. Although these approaches have significantly advanced the state of the art, the vast majority has solely focused on improving quantitative results and are not designed for low-compute scenarios. In this paper, we present a new general framework for a variety of computer vision labeling tasks, called HashMatch. Our approach is designed to be both fully parallel, i.e. each pixel is independently processed, and low-compute, with a model complexity an order of magnitude less than existing CNN and CRF-based approaches. We evaluate HashMatch extensively on several problems such as disparity estimation, image retrieval, feature approximation and background subtraction, for which HashMatch achieves high computational efficiency while producing high quality results.
Sean Ryan Fanello, Julien P. C. Valentin, Adarsh Kowdle, Christoph Rhemann, Vladimir Tankovich, Carlo Ciliberto, Philip Davidson, Shahram Izadi
ICCV2
2016 Learning to Navigate the Energy Landscape
abstract
In this paper, we present a novel, general, and efficient architecture for addressing computer vision problems that are approached from an 'Analysis by Synthesis' standpoint. Analysis by synthesis involves the minimization of reconstruction error, which is typically a non-convex function of the latent target variables. State-of-the-art methods adopt a hybrid scheme where discriminatively trained predictors like Random Forests or Convolutional Neural Networks are used to initialize local search algorithms. While these hybrid methods have been shown to produce promising results, they often get stuck in local optima. Our method goes beyond the conventional hybrid architecture by not only proposing multiple accurate initial solutions but by also defining a navigational structure over the solution space that can be used for extremely efficient gradient-free local search. We demonstrate the efficacy and generalizability of our approach on tasks as diverse as Hand Pose Estimation, RGB Camera Relocalization, and Image Retrieval.
Julien P. C. Valentin, Angela Dai, Matthias Nießner, Pushmeet Kohli, Philip Torr 0001, Shahram Izadi, Cem Keskin
3DV1
2016 Knowing who to listen to: Prioritizing experts from a diverse ensemble for attribute personalization
abstract
Learning attribute models for applications like Zero-Shot Learning (ZSL) and image search is challenging because they require attribute classifiers to generalize to test data that may be very different from the training data. A typical scenario is when the notion of an attribute may differ from one user to another, e.g. one user may find a shoe formal whereas another user may not. In this case, the distribution of labels at test time is different from that at training time. We argue that due to the uncertainty in what the test distribution might be, committing to one attribute model during training is not advisable. We propose a novel framework for attribute learning which involves training an ensemble of diverse models for attributes and identifying experts from them at test time given a small amount of personalized annotations from a user. Our approach for attribute personalization is not specific to any classification model and we show results using Random Forest and SVM ensembles. We experiment with 2 datasets: SUN Attributes and Shoes and show significant improvements over baselines.
Shrenik Lad, Bernardino Romera-Paredes, Julien P. C. Valentin, Philip Torr 0001, Devi Parikh
ICIP3
2016 Holoportation: Virtual 3D Teleportation in Real-time
abstract
We present an end-to-end system for augmented and virtual reality telepresence, called Holoportation. Our system demonstrates high-quality, real-time 3D reconstructions of an entire space, including people, furniture and objects, using a set of new depth cameras. These 3D models can also be transmitted in real-time to remote users. This allows users wearing virtual or augmented reality displays to see, hear and interact with remote participants in 3D, almost as if they were present in the same physical space. From an audio-visual perspective, communicating and interacting with remote users edges closer to face-to-face communication. This paper describes the Holoportation technical system in full, its key interactive capabilities, the application scenarios it enables, and an initial qualitative study of using this new communication medium.
Sergio Orts, Christoph Rhemann, Sean Ryan Fanello, Wayne Chang, Adarsh Kowdle, Yury Degtyarev, David Kim 0002, Philip Davidson, Sameh Khamis, Mingsong Dou, Vladimir Tankovich, Charles T. Loop, Qin Cai, Philip A. Chou, Sarah Mennicken, Julien P. C. Valentin, Vivek Pradeep, Shenlong Wang, Sing Bing Kang, Pushmeet Kohli, Yuliya Lutchyn, Cem Keskin, Shahram Izadi
UIST16
2016 Efficient and precise interactive hand tracking through joint, continuous optimization of pose and correspondences
abstract
Fully articulated hand tracking promises to enable fundamentally new interactions with virtual and augmented worlds, but the limited accuracy and efficiency of current systems has prevented widespread adoption. Today's dominant paradigm uses machine learning for initialization and recovery followed by iterative model-fitting optimization to achieve a detailed pose fit. We follow this paradigm, but make several changes to the model-fitting, namely using: (1) a more discriminative objective function; (2) a smooth-surface model that provides gradients for non-linear optimization; and (3) joint optimization over both the model pose and the correspondences between observed data points and the model surface. While each of these changes may actually increase the cost per fitting iteration, we find a compensating decrease in the number of iterations. Further, the wide basin of convergence means that fewer starting points are needed for successful model fitting. Our system runs in real-time on CPU only, which frees up the commonly over-burdened GPU for experience designers. The hand tracker is efficient enough to run on low-power devices such as tablets. We can track up to several meters from the camera to provide a large working volume for interaction, even using the noisy data from current-generation depth cameras. Quantitative assessments on standard datasets show that the new approach exceeds the state of the art in accuracy. Qualitative results take the form of live recordings of a range of interactive experiences enabled by this new approach.
Jonathan Taylor 0001, Lucas Bordeaux, Thomas J. Cashman 0001, Bob Corish, Cem Keskin, Toby Sharp, Eduardo Soto, David Sweeney, Julien P. C. Valentin, Benjamin Luff, Arran Topalian, Erroll Wood, Sameh Khamis, Pushmeet Kohli, Shahram Izadi, Richard Banks, Andrew W. Fitzgibbon, Jamie Shotton
ACM Trans. Graph.9
2015 Joint Object-Material Category Segmentation from Audio-Visual Cues
abstract
It is not always possible to recognise objects and infer material properties for a scene from visual cues alone, since objects can look visually similar whilst being made of very different materials. In this paper, we therefore present an approach that augments the available dense visual cues with sparse auditory cues in order to estimate dense object and material labels. Since estimates of object class and material properties are mutually informative, we optimise our multi-output labelling jointly using a random-field framework. We evaluate our system on a new dataset with paired visual and auditory data that we make publicly available. We demonstrate that this joint estimation of object and material labels significantly outperforms the estimation of either category in isolation.
Anurag Arnab, Michael Sapienza, Stuart Golodetz, Julien P. C. Valentin, Ondrej Miksik, Shahram Izadi, Philip Torr 0001
BMVC4
2015 Exploiting uncertainty in regression forests for accurate camera relocalization
abstract
Recent advances in camera relocalization use predictions from a regression forest to guide the camera pose optimization procedure. In these methods, each tree associates one pixel with a point in the scene's 3D world coordinate frame. In previous work, these predictions were point estimates and the subsequent camera pose optimization implicitly assumed an isotropic distribution of these estimates. In this paper, we train a regression forest to predict mixtures of anisotropic 3D Gaussians and show how the predicted uncertainties can be taken into account for continuous pose optimization. Experiments show that our proposed method is able to relocalize up to 40% more frames than the state of the art.
Julien P. C. Valentin, Matthias Nießner, Jamie Shotton, Andrew W. Fitzgibbon, Shahram Izadi, Philip Torr 0001
CVPR1
2015 SemanticPaint: Interactive 3D Labeling and Learning at your Fingertips
abstract
We present a new interactive and online approach to 3D scene understanding. Our system, SemanticPaint , allows users to simultaneously scan their environment whilst interactively segmenting the scene simply by reaching out and touching any desired object or surface. Our system continuously learns from these segmentations, and labels new unseen parts of the environment. Unlike offline systems where capture, labeling, and batch learning often take hours or even days to perform, our approach is fully online. This provides users with continuous live feedback of the recognition during capture, allowing to immediately correct errors in the segmentation and/or learning—a feature that has so far been unavailable to batch and offline methods. This leads to models that are tailored or personalized specifically to the user's environments and object classes of interest, opening up the potential for new applications in augmented reality, interior design, and human/robot navigation. It also provides the ability to capture substantial labeled 3D datasets for training large-scale visual recognition systems.
Julien P. C. Valentin, Vibhav Vineet, Ming-Ming Cheng, David Kim 0002, Jamie Shotton, Pushmeet Kohli, Matthias Nießner, Antonio Criminisi, Shahram Izadi, Philip Torr 0001
ACM Trans. Graph.1
2013 Mesh Based Semantic Modelling for Indoor and Outdoor Scenes
abstract
Semantic reconstruction of a scene is important for a variety of applications such as 3D modelling, object recognition and autonomous robotic navigation. However, most object labelling methods work in the image domain and fail to capture the information present in 3D space. In this work we propose a principled way to generate object labelling in 3D. Our method builds a triangulated meshed representation of the scene from multiple depth estimates. We then define a CRF over this mesh, which is able to capture the consistency of geometric properties of the objects present in the scene. In this framework, we are able to generate object hypotheses by combining information from multiple sources: geometric properties (from the 3D mesh), and appearance properties (from images). We demonstrate the robustness of our framework in both indoor and outdoor scenes. For indoor scenes we created an augmented version of the NYU indoor scene dataset (RGBD images) with object labelled meshes for training and evaluation. For outdoor scenes, we created ground truth object labellings for the KITTY odometry dataset (stereo image sequence). We observe a significant speed-up in the inference stage by performing labelling on the mesh, and additionally achieve higher accuracies.
Julien P. C. Valentin, Sunando Sengupta, Jonathan Warrell, Ali Shahrokni, Philip Torr 0001
CVPR1
2012 A Robust Stereo Prior for Human Segmentation
Glenn Sheasby, Julien P. C. Valentin, Nigel T. Crook, Philip Torr 0001
ACCV (2)2