Vladimir Tankovich

dblp:183/9318 · DBLP profile ↗
← Back
13ranked-venue papers
2as first author
2since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
10 papers
3D vision · 89% Video understanding and tracking · 8% Segmentation and scene understanding · 3%
Computer graphics and multimedia
5 papers
Virtual and augmented reality · 80% Multimedia analysis and retrieval · 16% Computer animation and physical simulation · 4%
Human-computer interaction and pervasive computing
2 papers
Interaction techniques and input · 79% Immersive interaction · 21%

Topics — the 28 heaviest of 30, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision › stereo vision
stereo matching
1.752021
HITNet: Hierarchical Iterative Tile Refinement Network for Real-time Stereo Matching · CVPR 2021
ActiveStereoNet: End-to-End Self-supervised Learning for Active Stereo Systems · ECCV (8) 2018
Low Compute and Fully Parallel Computer Vision with HashMatch · ICCV 2017
Computer vision › 3D vision
depth estimation
0.932018
The need 4 speed in real-time dense visual tracking · ACM Trans. Graph. 2018
UltraStereo: Efficient Learning-Based Matching for Active Stereo Systems · CVPR 2017
HyperDepth: Learning Depth from Structured Light without Matching · CVPR 2016
Computer vision › 3D vision › stereo vision
active stereo
0.622018
ActiveStereoNet: End-to-End Self-supervised Learning for Active Stereo Systems · ECCV (8) 2018
UltraStereo: Efficient Learning-Based Matching for Active Stereo Systems · CVPR 2017
Computer vision › 3D vision
3d reconstruction
0.632018
Fusion4D: real-time performance capture of challenging scenes · ACM Trans. Graph. 2016
Holoportation: Virtual 3D Teleportation in Real-time · UIST 2016
The need 4 speed in real-time dense visual tracking · ACM Trans. Graph. 2018
Computer vision › 3D vision › motion capture
human performance capture
0.522017
Motion2fusion: real-time volumetric performance capture · ACM Trans. Graph. 2017
Fusion4D: real-time performance capture of challenging scenes · ACM Trans. Graph. 2016
Computer vision › Video understanding and tracking › motion tracking
dense tracking
0.312018
The need 4 speed in real-time dense visual tracking · ACM Trans. Graph. 2018
Computer vision › 3D vision › depth estimation
monocular depth estimation
0.312018
Depth from motion for smartphone AR · ACM Trans. Graph. 2018
Computer vision › Video understanding and tracking
object tracking
0.312018
The need 4 speed in real-time dense visual tracking · ACM Trans. Graph. 2018
Computer vision › 3D vision › stereo vision › stereo matching › deep stereo matching
self-supervised stereo matching
0.312018
ActiveStereoNet: End-to-End Self-supervised Learning for Active Stereo Systems · ECCV (8) 2018
Virtual and augmented reality › tracking
6DOF tracking
0.312018
Depth from motion for smartphone AR · ACM Trans. Graph. 2018
Virtual and augmented reality › augmented reality
mobile augmented reality
0.312018
Depth from motion for smartphone AR · ACM Trans. Graph. 2018
Computer vision › 3D vision › geometric estimation › registration
non-rigid registration
0.312017
Motion2fusion: real-time volumetric performance capture · ACM Trans. Graph. 2017
Computer vision › Segmentation and scene understanding › dense prediction
pixel labeling
0.312017
Low Compute and Fully Parallel Computer Vision with HashMatch · ICCV 2017
Computer vision › 3D vision › depth estimation
stereo depth estimation
0.312017
Low Compute and Fully Parallel Computer Vision with HashMatch · ICCV 2017
Computer vision › 3D vision › 3d reconstruction
volumetric reconstruction
0.312017
Motion2fusion: real-time volumetric performance capture · ACM Trans. Graph. 2017
Virtual and augmented reality › immersive interaction
hand tracking
0.312017
Articulated distance fields for ultra-fast tracking of hands interacting · ACM Trans. Graph. 2017
Multimedia analysis and retrieval
image retrieval
0.312017
Low Compute and Fully Parallel Computer Vision with HashMatch · ICCV 2017
Interaction techniques and input › input sensing › tracking
hand tracking
0.312017
Articulated distance fields for ultra-fast tracking of hands interacting · ACM Trans. Graph. 2017
Computer vision › 3D vision › feature matching
correspondence problem
0.212016
HyperDepth: Learning Depth from Structured Light without Matching · CVPR 2016
Computer vision › 3D vision › 3d reconstruction
multi-view stereo
0.212016
Fusion4D: real-time performance capture of challenging scenes · ACM Trans. Graph. 2016
Computer vision › 3D vision › 3d reconstruction
non-rigid reconstruction
0.212016
Fusion4D: real-time performance capture of challenging scenes · ACM Trans. Graph. 2016
Computer vision › 3D vision › 3d reconstruction › volumetric reconstruction
real-time volumetric reconstruction
0.212016
Holoportation: Virtual 3D Teleportation in Real-time · UIST 2016
Computer vision › 3D vision › range sensing
structured light
0.212016
HyperDepth: Learning Depth from Structured Light without Matching · CVPR 2016
Computer vision › 3D vision › 3d reconstruction › volumetric reconstruction
volumetric fusion
0.212016
Fusion4D: real-time performance capture of challenging scenes · ACM Trans. Graph. 2016
Virtual and augmented reality › telepresence
3d telepresence
0.212016
Holoportation: Virtual 3D Teleportation in Real-time · UIST 2016
Virtual and augmented reality
telepresence
0.212016
Holoportation: Virtual 3D Teleportation in Real-time · UIST 2016
Computer vision › 3D vision
pose estimation
0.112018
The need 4 speed in real-time dense visual tracking · ACM Trans. Graph. 2018
Immersive interaction
augmented reality interaction
0.112016
Holoportation: Virtual 3D Teleportation in Real-time · UIST 2016

Methods — techniques the papers use, named apart from their topics

visual-inertial odometry · 0.7monocular depth computation · 0.7signed distance function · 0.6hybrid tracking · 0.6gradient-based optimization · 0.6neural network · 0.5geometric propagation · 0.5depth camera · 0.5space-time feature matching · 0.3self-supervised learning · 0.3machine learning depth refinement · 0.3active stereo · 0.3patchmatch · 0.3hashing · 0.3fully parallel per-pixel processing · 0.3binary patch representation · 0.3volumetric fusion · 0.2qualitative study · 0.2
YearPublicationVenuePosition
2023 NeuralBF: Neural Bilateral Filtering for Top-down Instance Segmentation on Point Clouds
abstract
We introduce a method for instance proposal generation for 3D point clouds. Existing techniques typically directly regress proposals in a single feed-forward step, leading to inaccurate estimation. We show that this serves as a critical bottleneck, and propose a method based on iterative bilateral filtering with learned kernels. Following the spirit of bilateral filtering, we consider both the deep feature embeddings of each point, as well as their locations in the 3D space. We show via synthetic experiments that our method brings drastic improvements when generating instance proposals for a given point of interest. We further validate our method on the challenging ScanNet benchmark, achieving the best instance segmentation performance amongst the sub-category of top-down methods.
Daniel Rebain, Renjie Liao 0001, Vladimir Tankovich, Soroosh Yazdani, Kwang Moo Yi, Andrea Tagliasacchi
WACV4
2021 HITNet: Hierarchical Iterative Tile Refinement Network for Real-time Stereo Matching
abstract
This paper presents HITNet, a novel neural network architecture for real-time stereo matching. Contrary to many recent neural network approaches that operate on a full cost volume and rely on 3D convolutions, our approach does not explicitly build a volume and instead relies on a fast multi-resolution initialization step, differentiable 2D geometric propagation and warping mechanisms to infer disparity hypotheses. To achieve a high level of accuracy, our network not only geometrically reasons about disparities but also infers slanted plane hypotheses allowing to more accurately perform geometric warping and upsampling operations. Our architecture is inherently multi-resolution allowing the propagation of information across different levels. Multiple experiments prove the effectiveness of the proposed approach at a fraction of the computation required by state-of-the-art methods. At the time of writing, HITNet ranks 1st-3rdon all the metrics published on the ETH3D website for two view stereo, ranks 1ston most of the metrics amongst all the end-to-end learning approaches on Middlebury-v3, ranks 1ston the popular KITTI 2012 and 2015 benchmarks among the published methods faster than 100 ms.
Vladimir Tankovich, Christian Häne, Yinda Zhang 0001, Adarsh Kowdle, Sean Ryan Fanello, Sofien Bouaziz
CVPR1
2018 ActiveStereoNet: End-to-End Self-supervised Learning for Active Stereo Systems
Yinda Zhang 0001, Sameh Khamis, Christoph Rhemann, Julien P. C. Valentin, Adarsh Kowdle, Vladimir Tankovich, Michael Schoenberg, Shahram Izadi, Thomas A. Funkhouser, Sean Ryan Fanello
ECCV (8)6
2018 SOS: Stereo Matching in O(1) with Slanted Support Windows
abstract
Depth cameras have accelerated research in many areas of computer vision. Most triangulation-based depth cameras, whether structured light systems like the Kinect or active (assisted) stereo systems, are based on the principle of stereo matching. Depth from stereo is an active research topic dating back 30 years. Despite recent advances, algorithms usually trade-off accuracy for speed. In particular, efficient methods rely on fronto-parallel assumptions to reduce the search space and keep computation low. We present SOS (Slanted O(1) Stereo), the first algorithm capable of leveraging slanted support windows without sacrificing speed or accuracy. We use an active stereo configuration, where an illuminator textures the scene. Under this setting, local methods - such as PatchMatch Stereo - obtain state of the art results by jointly estimating disparities and slant, but at a large computational cost. We observe that these methods typically exploit local smoothness to simplify their initialization strategies. Our key insight is that local smoothness can in fact be used to amortize the computation not only within initialization, but across the entire stereo pipeline. Building on these insights, we propose a novel hierarchical initialization that is able to efficiently perform search over disparity and slants. We then show how this structure can be leveraged to provide high quality depth maps. Extensive quantitative evaluations demonstrate that the proposed technique yields significantly more precise results than current state of the art, but at a fraction of the computational cost. Our prototype implementation runs at 4000 fps on modern GPU architectures.
Vladimir Tankovich, Michael Schoenberg, Sean Ryan Fanello, Adarsh Kowdle, Christoph Rhemann, Maksym Dzitsiuk, Mirko Schmidt, Julien P. C. Valentin, Shahram Izadi
IROS1
2018 The need 4 speed in real-time dense visual tracking
abstract
The advent of consumer depth cameras has incited the development of a new cohort of algorithms tackling challenging computer vision problems. The primary reason is that depth provides direct geometric information that is largely invariant to texture and illumination. As such, substantial progress has been made in human and object pose estimation, 3D reconstruction and simultaneous localization and mapping. Most of these algorithms naturally benefit from the ability to accurately track the pose of an object or scene of interest from one frame to the next. However, commercially available depth sensors (typically running at 30fps) can allow for large inter-frame motions to occur that make such tracking problematic. A high frame rate depth camera would thus greatly ameliorate these issues, and further increase the tractability of these computer vision problems. Nonetheless, the depth accuracy of recent systems for high-speed depth estimation [Fanello et al. 2017b] can degrade at high frame rates. This is because the active illumination employed produces a low SNR and thus a high exposure time is required to obtain a dense accurate depth image. Furthermore in the presence of rapid motion, longer exposure times produce artifacts due to motion blur, and necessitates a lower frame rate that introduces large inter-frame motion that often yield tracking failures. In contrast, this paper proposes a novel combination of hardware and software components that avoids the need to compromise between a dense accurate depth map and a high frame rate. We document the creation of a full 3D capture system for high speed and quality depth estimation, and demonstrate its advantages in a variety of tracking and reconstruction tasks. We extend the state of the art active stereo algorithm presented in Fanello et al. [2017b] by adding a space-time feature in the matching phase. We also propose a machine learning based depth refinement step that is an order of magnitude faster than traditional postprocessing methods. We quantitatively and qualitatively demonstrate the benefits of the proposed algorithms in the acquisition of geometry in motion. Our pipeline executes in 1.1ms leveraging modern GPUs and off-the-shelf cameras and illumination components. We show how the sensor can be employed in many different applications, from [non-]rigid reconstructions to hand/face tracking. Further, we show many advantages over existing state of the art depth camera technologies beyond framerate, including latency, motion artifacts, multi-path errors, and multi-sensor interference.
Adarsh Kowdle, Christoph Rhemann, Sean Ryan Fanello, Andrea Tagliasacchi, Jonathan Taylor 0001, Philip Davidson, Mingsong Dou, Cem Keskin, Sameh Khamis, David Kim 0002, Danhang Tang, Vladimir Tankovich, Julien P. C. Valentin, Shahram Izadi
ACM Trans. Graph.13
2018 Depth from motion for smartphone AR
abstract
Augmented reality (AR) for smartphones has matured from a technology for earlier adopters, available only on select high-end phones, to one that is truly available to the general public. One of the key breakthroughs has been in low-compute methods for six degree of freedom (6DoF) tracking on phones using only the existing hardware (camera and inertial sensors). 6DoF tracking is the cornerstone of smartphone AR allowing virtual content to be precisely locked on top of the real world. However, to really give users the impression of believable AR, one requires mobile depth. Without depth, even simple effects such as a virtual object being correctly occluded by the real-world is impossible. However, requiring a mobile depth sensor would severely restrict the access to such features. In this article, we provide a novel pipeline for mobile depth that supports a wide array of mobile phones, and uses only the existing monocular color sensor. Through several technical contributions, we provide the ability to compute low latency dense depth maps using only a single CPU core of a wide range of (medium-high) mobile phones. We demonstrate the capabilities of our approach on high-level AR applications including real-time navigation and shopping.
Julien P. C. Valentin, Adarsh Kowdle, Jonathan T. Barron, Neal Wadhwa, Maksym Dzitsiuk, Michael Schoenberg, Ambrus Csaszar, Eric Turner 0001, Ivan Dryanovski, João Afonso, Jose Pascoal, Konstantine Tsotsos, Mira Leung, Mirko Schmidt, Onur G. Guleryuz, Sameh Khamis, Vladimir Tankovich, Sean Ryan Fanello, Shahram Izadi, Christoph Rhemann
ACM Trans. Graph.18
2017 UltraStereo: Efficient Learning-Based Matching for Active Stereo Systems
abstract
Efficient estimation of depth from pairs of stereo images is one of the core problems in computer vision. We efficiently solve the specialized problem of stereo matching under active illumination using a new learning-based algorithm. This type of active stereo i.e. stereo matching where scene texture is augmented by an active light projector is proving compelling for designing depth cameras, largely due to improved robustness when compared to time of flight or traditional structured light techniques. Our algorithm uses an unsupervised greedy optimization scheme that learns features that are discriminative for estimating correspondences in infrared images. The proposed method optimizes a series of sparse hyperplanes that are used at test time to remap all the image patches into a compact binary representation in O(1). The proposed algorithm is cast in a PatchMatch Stereo-like framework, producing depth maps at 500Hz. In contrast to standard structured light methods, our approach generalizes to different scenes, does not require tedious per camera calibration procedures and is not adversely affected by interference from overlapping sensors. Extensive evaluations show we surpass the quality and overcome the limitations of current depth sensing technologies.
Sean Ryan Fanello, Julien P. C. Valentin, Christoph Rhemann, Adarsh Kowdle, Vladimir Tankovich, Philip Davidson, Shahram Izadi
CVPR5
2017 Low Compute and Fully Parallel Computer Vision with HashMatch
abstract
Numerous computer vision problems such as stereo depth estimation, object-class segmentation and fore-ground/background segmentation can be formulated as per-pixel image labeling tasks. Given one or many images as input, the desired output of these methods is usually a spatially smooth assignment of labels. The large amount of such computer vision problems has lead to significant research efforts, with the state of art moving from CRF-based approaches to deep CNNs and more recently, hybrids of the two. Although these approaches have significantly advanced the state of the art, the vast majority has solely focused on improving quantitative results and are not designed for low-compute scenarios. In this paper, we present a new general framework for a variety of computer vision labeling tasks, called HashMatch. Our approach is designed to be both fully parallel, i.e. each pixel is independently processed, and low-compute, with a model complexity an order of magnitude less than existing CNN and CRF-based approaches. We evaluate HashMatch extensively on several problems such as disparity estimation, image retrieval, feature approximation and background subtraction, for which HashMatch achieves high computational efficiency while producing high quality results.
Sean Ryan Fanello, Julien P. C. Valentin, Adarsh Kowdle, Christoph Rhemann, Vladimir Tankovich, Carlo Ciliberto, Philip Davidson, Shahram Izadi
ICCV5
2017 Motion2fusion: real-time volumetric performance capture
abstract
We present Motion2Fusion, a state-of-the-art 360 performance capture system that enables *real-time* reconstruction of arbitrary non-rigid scenes. We provide three major contributions over prior work: 1) a new non-rigid fusion pipeline allowing for far more faithful reconstruction of high frequency geometric details, avoiding the over-smoothing and visual artifacts observed previously. 2) a high speed pipeline coupled with a machine learning technique for 3D correspondence field estimation reducing tracking errors and artifacts that are attributed to fast motions. 3) a backward and forward non-rigid alignment strategy that more robustly deals with topology changes but is still free from scene priors. Our novel performance capture system demonstrates real-time results nearing 3x speed-up from previous state-of-the-art work on the exact same GPU hardware. Extensive quantitative and qualitative comparisons show more precise geometric and texturing results with less artifacts due to fast motions or topology changes than prior art.
Mingsong Dou, Philip Davidson, Sean Ryan Fanello, Sameh Khamis, Adarsh Kowdle, Christoph Rhemann, Vladimir Tankovich, Shahram Izadi
ACM Trans. Graph.7
2017 Articulated distance fields for ultra-fast tracking of hands interacting
abstract
The state of the art in articulated hand tracking has been greatly advanced by hybrid methods that fit a generative hand model to depth data, leveraging both temporally and discriminatively predicted starting poses. In this paradigm, the generative model is used to define an energy function and a local iterative optimization is performed from these starting poses in order to find a "good local minimum" (i.e. a local minimum close to the true pose). Performing this optimization quickly is key to exploring more starting poses, performing more iterations and, crucially, exploiting high frame rates that ensure that temporally predicted starting poses are in the basin of convergence of a good local minimum. At the same time, a detailed and accurate generative model tends to deepen the good local minima and widen their basins of convergence. Recent work, however, has largely had to trade-off such a detailed hand model with one that facilitates such rapid optimization. We present a new implicit model of hand geometry that mostly avoids this compromise and leverage it to build an ultra-fast hybrid hand tracking system. Specifically, we construct an articulated signed distance function that, for any pose, yields a closed form calculation of both the distance to the detailed surface geometry and the necessary derivatives to perform gradient based optimization. There is no need to introduce or update any explicit "correspondences" yielding a simple algorithm that maps well to parallel hardware such as GPUs. As a result, our system can run at extremely high frame rates (e.g. up to 1000fps). Furthermore, we demonstrate how to detect, segment and optimize for two strongly interacting hands, recovering complex interactions at extremely high framerates. In the absence of publicly available datasets of sufficiently high frame rate, we leverage a multiview capture system to create a new 180fps dataset of one and two hands interacting together or with objects.
Jonathan Taylor 0001, Vladimir Tankovich, Danhang Tang, Cem Keskin, David Kim 0002, Philip Davidson, Adarsh Kowdle, Shahram Izadi
ACM Trans. Graph.2
2016 HyperDepth: Learning Depth from Structured Light without Matching
abstract
Structured light sensors are popular due to their robustness to untextured scenes and multipath. These systems triangulate depth by solving a correspondence problem between each camera and projector pixel. This is often framed as a local stereo matching task, correlating patches of pixels in the observed and reference image. However, this is computationally intensive, leading to reduced depth accuracy and framerate. We contribute an algorithm for solving this correspondence problem efficiently, without compromising depth accuracy. For the first time, this problem is cast as a classification-regression task, which we solve extremely efficiently using an ensemble of cascaded random forests. Our algorithm scales in number of disparities, and each pixel can be processed independently, and in parallel. No matching or even access to the corresponding reference pattern is required at runtime, and regressed labels are directly mapped to depth. Our GPU-based algorithm runs at a 1KHz for 1.3MP input/output images, with disparity error of 0.1 subpixels. We show a prototype high framerate depth camera running at 375Hz, useful for solving tracking-related problems. We demonstrate our algorithmic performance, creating high resolution real-time depth maps that surpass the quality of current state of the art depth technologies, highlighting quantization-free results with reduced holes, edge fattening and other stereo-based depth artifacts.
Sean Ryan Fanello, Christoph Rhemann, Vladimir Tankovich, Adarsh Kowdle, Sergio Orts, David Kim 0002, Shahram Izadi
CVPR3
2016 Holoportation: Virtual 3D Teleportation in Real-time
abstract
We present an end-to-end system for augmented and virtual reality telepresence, called Holoportation. Our system demonstrates high-quality, real-time 3D reconstructions of an entire space, including people, furniture and objects, using a set of new depth cameras. These 3D models can also be transmitted in real-time to remote users. This allows users wearing virtual or augmented reality displays to see, hear and interact with remote participants in 3D, almost as if they were present in the same physical space. From an audio-visual perspective, communicating and interacting with remote users edges closer to face-to-face communication. This paper describes the Holoportation technical system in full, its key interactive capabilities, the application scenarios it enables, and an initial qualitative study of using this new communication medium.
Sergio Orts, Christoph Rhemann, Sean Ryan Fanello, Wayne Chang, Adarsh Kowdle, Yury Degtyarev, David Kim 0002, Philip Davidson, Sameh Khamis, Mingsong Dou, Vladimir Tankovich, Charles T. Loop, Qin Cai, Philip A. Chou, Sarah Mennicken, Julien P. C. Valentin, Vivek Pradeep, Shenlong Wang, Sing Bing Kang, Pushmeet Kohli, Yuliya Lutchyn, Cem Keskin, Shahram Izadi
UIST11
2016 Fusion4D: real-time performance capture of challenging scenes
abstract
We contribute a new pipeline for live multi-view performance capture, generating temporally coherent high-quality reconstructions in real-time. Our algorithm supports both incremental reconstruction, improving the surface estimation over time, as well as parameterizing the nonrigid scene motion. Our approach is highly robust to both large frame-to-frame motion and topology changes, allowing us to reconstruct extremely challenging scenes. We demonstrate advantages over related real-time techniques that either deform an online generated template or continually fuse depth data nonrigidly into a single reference model. Finally, we show geometric reconstruction results on par with offline methods which require orders of magnitude more processing time and many more RGBD cameras.
Mingsong Dou, Sameh Khamis, Yury Degtyarev, Philip Davidson, Sean Ryan Fanello, Adarsh Kowdle, Sergio Orts, Christoph Rhemann, David Kim 0002, Jonathan Taylor 0001, Pushmeet Kohli, Vladimir Tankovich, Shahram Izadi
ACM Trans. Graph.12