Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Adarsh Kowdle

dblp:25/8107 · DBLP profile ↗
← Back
36ranked-venue papers
9as first author
4since 2021 · last 2026
0000-0002-4428-889XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 27 · 8 first-author · 2 since 2021Artificial intelligence and machine learning · 17 · 5 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 4 · 2 since 2021Systems, architecture and hardware · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
19 papers
3D vision · 75% Segmentation and scene understanding · 7% Video understanding and tracking · 5%
Computer graphics and multimedia
14 papers
Virtual and augmented reality · 29% Image and video processing · 24% Computational photography and imaging · 13%
Human-computer interaction and pervasive computing
7 papers
Immersive interaction · 45% Interaction techniques and input · 23% User interface design and tools · 20%

Topics — the 30 heaviest of 67, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision › stereo vision
stereo matching
2.062021
HITNet: Hierarchical Iterative Tile Refinement Network for Real-time Stereo Matching · CVPR 2021
ActiveStereoNet: End-to-End Self-supervised Learning for Active Stereo Systems · ECCV (8) 2018
StereoNet: Guided Hierarchical Refinement for Real-Time Edge-Aware Depth Prediction · ECCV (15) 2018
Computer vision › 3D vision
depth estimation
1.462018
The need 4 speed in real-time dense visual tracking · ACM Trans. Graph. 2018
StereoNet: Guided Hierarchical Refinement for Real-Time Edge-Aware Depth Prediction · ECCV (15) 2018
UltraStereo: Efficient Learning-Based Matching for Active Stereo Systems · CVPR 2017
Computer vision › 3D vision › depth estimation
monocular depth estimation
0.922022
Learned Monocular Depth Priors in Visual-Inertial Initialization · ECCV (22) 2022
Depth from motion for smartphone AR · ACM Trans. Graph. 2018
Computer vision › 3D vision
3d reconstruction
0.742018
Fusion4D: real-time performance capture of challenging scenes · ACM Trans. Graph. 2016
Holoportation: Virtual 3D Teleportation in Real-time · UIST 2016
Active learning for piecewise planar 3D reconstruction · CVPR 2011
User interface design and tools
visual programming
0.712023
Rapsai: Accelerating Machine Learning Prototyping of Multimedia Applications through Visual Programming · CHI 2023
Computer vision › 3D vision › stereo vision
active stereo
0.622018
ActiveStereoNet: End-to-End Self-supervised Learning for Active Stereo Systems · ECCV (8) 2018
UltraStereo: Efficient Learning-Based Matching for Active Stereo Systems · CVPR 2017
Robotics › Robot navigation and mapping › visual odometry
visual-inertial odometry
0.612022
Learned Monocular Depth Priors in Visual-Inertial Initialization · ECCV (22) 2022
Machine learning › Deep learning architectures and training
weight initialization
0.612022
Learned Monocular Depth Priors in Visual-Inertial Initialization · ECCV (22) 2022
Computer vision › 3D vision › motion capture
human performance capture
0.522017
Motion2fusion: real-time volumetric performance capture · ACM Trans. Graph. 2017
Fusion4D: real-time performance capture of challenging scenes · ACM Trans. Graph. 2016
Immersive interaction
augmented reality interaction
0.522020
DepthLab: Real-time 3D Interaction with Depth Maps for Mobile Augmented Reality · UIST 2020
Holoportation: Virtual 3D Teleportation in Real-time · UIST 2016
Virtual and augmented reality › augmented reality
mobile augmented reality
0.522020
Depth from motion for smartphone AR · ACM Trans. Graph. 2018
DepthLab: Real-time 3D Interaction with Depth Maps for Mobile Augmented Reality · UIST 2020
Interaction techniques and input › sensor-based interaction
depth camera-based interaction
0.412020
DepthLab: Real-time 3D Interaction with Depth Maps for Mobile Augmented Reality · UIST 2020
Computer vision › 3D vision › 3d reconstruction
volumetric reconstruction
0.422018
Motion2fusion: real-time volumetric performance capture · ACM Trans. Graph. 2017
Real-time compression and streaming of 4D performances · ACM Trans. Graph. 2018
Computational photography and imaging › 3d scanning
volumetric performance capture
0.412019
The relightables: volumetric performance capture of humans with realistic relighting · ACM Trans. Graph. 2019
Virtual and augmented reality
telepresence
0.322018
Holoportation: Virtual 3D Teleportation in Real-time · UIST 2016
LookinGood: enhancing performance capture with real-time neural re-rendering · ACM Trans. Graph. 2018
Computer vision › Video understanding and tracking › motion tracking
dense tracking
0.312018
The need 4 speed in real-time dense visual tracking · ACM Trans. Graph. 2018
Computer vision › Video understanding and tracking
object tracking
0.312018
The need 4 speed in real-time dense visual tracking · ACM Trans. Graph. 2018
Computer vision › 3D vision › stereo vision › stereo matching › deep stereo matching
self-supervised stereo matching
0.312018
ActiveStereoNet: End-to-End Self-supervised Learning for Active Stereo Systems · ECCV (8) 2018
Virtual and augmented reality › tracking
6DOF tracking
0.312018
Depth from motion for smartphone AR · ACM Trans. Graph. 2018
Image and video processing
image restoration
0.312018
LookinGood: enhancing performance capture with real-time neural re-rendering · ACM Trans. Graph. 2018
Rendering
neural rendering
0.312018
LookinGood: enhancing performance capture with real-time neural re-rendering · ACM Trans. Graph. 2018
Multimedia systems and quality of experience › video streaming
real-time streaming
0.312018
Real-time compression and streaming of 4D performances · ACM Trans. Graph. 2018
Image and video processing › image restoration › multi-task image restoration
super-resolution and denoising
0.312018
LookinGood: enhancing performance capture with real-time neural re-rendering · ACM Trans. Graph. 2018
Multimedia analysis and retrieval
image retrieval
0.322017
Low Compute and Fully Parallel Computer Vision with HashMatch · ICCV 2017
Interactively Co-segmentating Topically Related Images with Intelligent Scribble Guidance · Int. J. Comput. Vis. 2011
Audio and music processing
source separation
0.312026
MoXaRt: Audio-Visual Object-Guided Sound Interaction for XR · CHI 2026
Computer vision › 3D vision › geometric estimation › registration
non-rigid registration
0.312017
Motion2fusion: real-time volumetric performance capture · ACM Trans. Graph. 2017
Computer vision › Segmentation and scene understanding › dense prediction
pixel labeling
0.312017
Low Compute and Fully Parallel Computer Vision with HashMatch · ICCV 2017
Computer vision › 3D vision › depth estimation
stereo depth estimation
0.312017
Low Compute and Fully Parallel Computer Vision with HashMatch · ICCV 2017
Virtual and augmented reality › immersive interaction
hand tracking
0.312017
Articulated distance fields for ultra-fast tracking of hands interacting · ACM Trans. Graph. 2017
Interaction techniques and input › input sensing › tracking
hand tracking
0.312017
Articulated distance fields for ultra-fast tracking of hands interacting · ACM Trans. Graph. 2017

Methods — techniques the papers use, named apart from their topics

cascaded neural networks · 2.0audio-visual learning · 2.0signed distance function · 0.9physics-based collision · 0.9geometry-aware rendering · 0.9depth map processing · 0.9color gradient illumination · 0.8formative study · 0.7neural network · 0.5geometric propagation · 0.5machine learning reconstruction pipeline · 0.4deep neural network · 0.4active depth sensing · 0.4space-time feature matching · 0.3self-supervised learning · 0.3quantization · 0.3machine learning depth refinement · 0.3latent space embedding · 0.3
YearPublicationVenuePosition
2026 MoXaRt: Audio-Visual Object-Guided Sound Interaction for XR
abstract
In Extended Reality (XR), complex acoustic environments often overwhelm users, compromising both scene awareness and social engagement due to entangled sound sources. We introduce MoXaRt, a real-time XR system that uses audio-visual cues to separate these sources and enable fine-grained sound interaction. MoXaRt’s core is a cascaded architecture that performs coarse, audio-only separation in parallel with visual detection of sources (e.g., faces, instruments). These visual anchors then guide refinement networks to isolate individual sources, separating complex mixes of up to 5 concurrent sources (e.g., 2 voices + 3 instruments) with ∼ 2 second processing latency. We validate MoXaRt through a technical evaluation on a new dataset of 30 one-minute recordings featuring concurrent speech and music, and a 22-participant user study. Empirical results indicate that our system significantly enhances speech intelligibility, yielding a 36.2% (p < 0.01) increase in listening comprehension within adversarial acoustic environments while substantially reducing cognitive load (p < 0.001), thereby paving the way for more perceptive and socially adept XR experiences.
Tianyu Xu 0008, Qianhui Zheng, Tejasvi Ravi, Anuva Kulkarni, Katrina Passarella-Ward, Junyi Zhu 0001, Adarsh Kowdle
CHI9
2023 Rapsai: Accelerating Machine Learning Prototyping of Multimedia Applications through Visual Programming
abstract
In recent years, there has been a proliferation of multimedia applications that leverage machine learning (ML) for interactive experiences. Prototyping ML-based applications is, however, still challenging, given complex workflows that are not ideal for design and experimentation. To better understand these challenges, we conducted a formative study with seven ML practitioners to gather insights about common ML evaluation workflows.
Ruofei Du, Na Li 0034, Michelle Carney, Scott Miles, Maria Kleiner, Xiuxiu Yuan, Yinda Zhang 0001, Anuva Kulkarni, Xingyu Liu 0002, Ahmed Sabie, Sergio Orts, Abhishek Kar, Ram Iyengar, Adarsh Kowdle, Alex Olwal
CHI16
2022 Learned Monocular Depth Priors in Visual-Inertial Initialization
Yunwen Zhou, Abhishek Kar, Eric Turner 0001, Adarsh Kowdle, Chao X. Guo, Ryan DuToit, Konstantine Tsotsos
ECCV (22)4
2021 HITNet: Hierarchical Iterative Tile Refinement Network for Real-time Stereo Matching
abstract
This paper presents HITNet, a novel neural network architecture for real-time stereo matching. Contrary to many recent neural network approaches that operate on a full cost volume and rely on 3D convolutions, our approach does not explicitly build a volume and instead relies on a fast multi-resolution initialization step, differentiable 2D geometric propagation and warping mechanisms to infer disparity hypotheses. To achieve a high level of accuracy, our network not only geometrically reasons about disparities but also infers slanted plane hypotheses allowing to more accurately perform geometric warping and upsampling operations. Our architecture is inherently multi-resolution allowing the propagation of information across different levels. Multiple experiments prove the effectiveness of the proposed approach at a fraction of the computation required by state-of-the-art methods. At the time of writing, HITNet ranks 1st-3rdon all the metrics published on the ETH3D website for two view stereo, ranks 1ston most of the metrics amongst all the end-to-end learning approaches on Middlebury-v3, ranks 1ston the popular KITTI 2012 and 2015 benchmarks among the published methods faster than 100 ms.
Vladimir Tankovich, Christian Häne, Yinda Zhang 0001, Adarsh Kowdle, Sean Ryan Fanello, Sofien Bouaziz
CVPR4
2020 Multimodal Active Speaker Detection and Virtual Cinematography for Video Conferencing
abstract
Active speaker detection (ASD) and virtual cinematography (VC) can significantly improve the experience of a video conference by automatically panning, tilting and zooming of a camera: subjectively users rate an expert video cinematographer significantly higher than the unedited video. We describe a new automated ASD and VC that performs within 0.3 MOS of an expert cinematographer based on subjective ratings with a 1-5 scale. This system uses a 4K wide-FOV camera, a depth camera, and a microphone array, extracts features from each modality and trains an ASD using an AdaBoost machine learning system that is very efficient and runs in real-time. A VC is similarly trained using machine learning. To avoid distracting the room participants the system has no moving parts - the VC works by cropping and zooming the 4K wide-FOV video stream. The system was tuned and evaluated using extensive crowdsourcing techniques and evaluated on a system with N=100 meetings, each 25 minutes in length.
Ross Cutler, Ramin Mehran, Sam Johnson, Cha Zhang, Adam Kirk, Oliver Whyte, Adarsh Kowdle
ICASSP7
2020 DepthLab: Real-time 3D Interaction with Depth Maps for Mobile Augmented Reality
abstract
Mobile devices with passive depth sensing capabilities are ubiquitous, and recently active depth sensors have become available on some tablets and AR/VR devices. Although real-time depth data is accessible, its rich value to mainstream AR applications has been sorely under-explored. Adoption of depth-based UX has been impeded by the complexity of performing even simple operations with raw depth data, such as detecting intersections or constructing meshes. In this paper, we introduce DepthLab, a software library that encapsulates a variety of depth-based UI/UX paradigms, including geometry-aware rendering (occlusion, shadows), surface interaction behaviors (physics-based collisions, avatar path planning), and visual effects (relighting, 3D-anchored focus and aperture effects). We break down the usage of depth into localized depth, surface depth, and dense depth, and describe our real-time algorithms for interaction and rendering tasks. We present the design process, system, and components of DepthLab to streamline and centralize the development of interactive depth features. We have open-sourced our software at https://github.com/googlesamples/arcore-depth-lab to external developers, conducted performance evaluation, and discussed how DepthLab can accelerate the workflow of mobile AR designers and developers. With DepthLab we aim to help mobile developers to effortlessly integrate depth into their AR experiences and amplify the expression of their creative vision.
Ruofei Du, Eric Turner 0001, Maksym Dzitsiuk, Luca Prasso, Ivo Duarte, Jason Dourgarian, João Afonso, Jose Pascoal, Josh Gladstone, Nuno Cruces, Shahram Izadi, Adarsh Kowdle, Konstantine Tsotsos, David Kim 0002
UIST12
2019 The relightables: volumetric performance capture of humans with realistic relighting
abstract
We present "The Relightables", a volumetric capture system for photorealistic and high quality relightable full-body performance capture. While significant progress has been made on volumetric capture systems, focusing on 3D geometric reconstruction with high resolution textures, much less work has been done to recover photometric properties needed for relighting. Results from such systems lack high-frequency details and the subject's shading is prebaked into the texture. In contrast, a large body of work has addressed relightable acquisition for image-based approaches, which photograph the subject under a set of basis lighting conditions and recombine the images to show the subject as they would appear in a target lighting environment. However, to date, these approaches have not been adapted for use in the context of a high-resolution volumetric capture system. Our method combines this ability to realistically relight humans for arbitrary environments, with the benefits of free-viewpoint volumetric capture and new levels of geometric accuracy for dynamic performances. Our subjects are recorded inside a custom geodesic sphere outfitted with 331 custom color LED lights, an array of high-resolution cameras, and a set of custom high-resolution depth sensors. Our system innovates in multiple areas: First, we designed a novel active depth sensor to capture 12.4 MP depth maps, which we describe in detail. Second, we show how to design a hybrid geometric and machine learning reconstruction pipeline to process the high resolution input and output a volumetric video. Third, we generate temporally consistent reflectance maps for dynamic performers by leveraging the information contained in two alternating color gradient illumination images acquired at 60Hz. Multiple experiments, comparisons, and applications show that The Relightables significantly improves upon the level of realism in placing volumetrically captured human performances into arbitrary CG scenes.
Peter Lincoln, Philip Davidson, Jay Busch, Xueming Yu, Matt Whalen, Geoff Harvey, Sergio Orts, Rohit Pandey, Jason Dourgarian, Danhang Tang, Anastasia Tkach, Adarsh Kowdle, Emily Cooper, Mingsong Dou, Sean Ryan Fanello, Graham Fyffe, Christoph Rhemann, Jonathan Taylor 0001, Paul E. Debevec, Shahram Izadi
ACM Trans. Graph.13
2019 Deep reflectance fields: high-quality facial reflectance field inference from color gradient illumination
abstract
We present a novel technique to relight images of human faces by learning a model of facial reflectance from a database of 4D reflectance field data of several subjects in a variety of expressions and viewpoints. Using our learned model, a face can be relit in arbitrary illumination environments using only two original images recorded under spherical color gradient illumination. The output of our deep network indicates that the color gradient images contain the information needed to estimate the full 4D reflectance field, including specular reflections and high frequency details. While capturing spherical color gradient illumination still requires a special lighting setup, reduction to just two illumination conditions allows the technique to be applied to dynamic facial performance capture. We show side-by-side comparisons which demonstrate that the proposed system outperforms the state-of-the-art techniques in both realism and speed.
Abhimitra Meka, Christian Häne, Rohit Pandey, Michael Zollhöfer, Sean Ryan Fanello, Graham Fyffe, Adarsh Kowdle, Xueming Yu, Jay Busch, Jason Dourgarian, Peter Denny, Sofien Bouaziz, Peter Lincoln, Matt Whalen, Geoff Harvey, Jonathan Taylor 0001, Shahram Izadi, Andrea Tagliasacchi, Paul E. Debevec, Christian Theobalt, Julien P. C. Valentin, Christoph Rhemann
ACM Trans. Graph.7
2018 TwinFusion: High Framerate Non-rigid Fusion through Fast Correspondence Tracking
abstract
Real time non-rigid reconstruction pipelines are extremely computationally expensive and easily saturate the highest end GPUs currently available. This requires careful strategic choices to be made about a set of highly interconnected parameters that divide up the limited compute. At the same time, offline systems, prove the value of increasing voxel resolution, more iterations, and higher frame rates. To this end, we demonstrate a set of remarkably simple but effective modifications to these algorithms that significantly reduce the average per-frame computation cost allowing these parameters to be increased. Specifically, we divide the depth stream into sub-frames and fusion-frames, disabling both model accumulation (fusion) and non-rigid alignment (model tracking) on the former. Instead, we efficiently track point correspondences across neighboring sub-frames. We then leverage these correspondences to initialize the standard non-rigid alignment to a fusion-frame where data can then be accumulated into the model. As a result, compute resources in the modified non-rigid reconstruction pipeline can be immediately re-purposed to increase voxel resolution, use more iterations or to increase the frame rate. To demonstrate the latter, we leverage recent high frame rate depth algorithms to build a novel “twin” sensor consisting of a low-res/highfps sub-frame camera and a second low-fps/high-res fusion camera.
Jonathan Taylor 0001, Sean Ryan Fanello, Andrea Tagliasacchi, Mingsong Dou, Philip Davidson, Adarsh Kowdle, Shahram Izadi
3DV7
2018 StereoNet: Guided Hierarchical Refinement for Real-Time Edge-Aware Depth Prediction
Sameh Khamis, Sean Ryan Fanello, Christoph Rhemann, Adarsh Kowdle, Julien P. C. Valentin, Shahram Izadi
ECCV (15)4
2018 ActiveStereoNet: End-to-End Self-supervised Learning for Active Stereo Systems
Yinda Zhang 0001, Sameh Khamis, Christoph Rhemann, Julien P. C. Valentin, Adarsh Kowdle, Vladimir Tankovich, Michael Schoenberg, Shahram Izadi, Thomas A. Funkhouser, Sean Ryan Fanello
ECCV (8)5
2018 SOS: Stereo Matching in O(1) with Slanted Support Windows
abstract
Depth cameras have accelerated research in many areas of computer vision. Most triangulation-based depth cameras, whether structured light systems like the Kinect or active (assisted) stereo systems, are based on the principle of stereo matching. Depth from stereo is an active research topic dating back 30 years. Despite recent advances, algorithms usually trade-off accuracy for speed. In particular, efficient methods rely on fronto-parallel assumptions to reduce the search space and keep computation low. We present SOS (Slanted O(1) Stereo), the first algorithm capable of leveraging slanted support windows without sacrificing speed or accuracy. We use an active stereo configuration, where an illuminator textures the scene. Under this setting, local methods - such as PatchMatch Stereo - obtain state of the art results by jointly estimating disparities and slant, but at a large computational cost. We observe that these methods typically exploit local smoothness to simplify their initialization strategies. Our key insight is that local smoothness can in fact be used to amortize the computation not only within initialization, but across the entire stereo pipeline. Building on these insights, we propose a novel hierarchical initialization that is able to efficiently perform search over disparity and slants. We then show how this structure can be leveraged to provide high quality depth maps. Extensive quantitative evaluations demonstrate that the proposed technique yields significantly more precise results than current state of the art, but at a fraction of the computational cost. Our prototype implementation runs at 4000 fps on modern GPU architectures.
Vladimir Tankovich, Michael Schoenberg, Sean Ryan Fanello, Adarsh Kowdle, Christoph Rhemann, Maksym Dzitsiuk, Mirko Schmidt, Julien P. C. Valentin, Shahram Izadi
IROS4
2018 The need 4 speed in real-time dense visual tracking
abstract
The advent of consumer depth cameras has incited the development of a new cohort of algorithms tackling challenging computer vision problems. The primary reason is that depth provides direct geometric information that is largely invariant to texture and illumination. As such, substantial progress has been made in human and object pose estimation, 3D reconstruction and simultaneous localization and mapping. Most of these algorithms naturally benefit from the ability to accurately track the pose of an object or scene of interest from one frame to the next. However, commercially available depth sensors (typically running at 30fps) can allow for large inter-frame motions to occur that make such tracking problematic. A high frame rate depth camera would thus greatly ameliorate these issues, and further increase the tractability of these computer vision problems. Nonetheless, the depth accuracy of recent systems for high-speed depth estimation [Fanello et al. 2017b] can degrade at high frame rates. This is because the active illumination employed produces a low SNR and thus a high exposure time is required to obtain a dense accurate depth image. Furthermore in the presence of rapid motion, longer exposure times produce artifacts due to motion blur, and necessitates a lower frame rate that introduces large inter-frame motion that often yield tracking failures. In contrast, this paper proposes a novel combination of hardware and software components that avoids the need to compromise between a dense accurate depth map and a high frame rate. We document the creation of a full 3D capture system for high speed and quality depth estimation, and demonstrate its advantages in a variety of tracking and reconstruction tasks. We extend the state of the art active stereo algorithm presented in Fanello et al. [2017b] by adding a space-time feature in the matching phase. We also propose a machine learning based depth refinement step that is an order of magnitude faster than traditional postprocessing methods. We quantitatively and qualitatively demonstrate the benefits of the proposed algorithms in the acquisition of geometry in motion. Our pipeline executes in 1.1ms leveraging modern GPUs and off-the-shelf cameras and illumination components. We show how the sensor can be employed in many different applications, from [non-]rigid reconstructions to hand/face tracking. Further, we show many advantages over existing state of the art depth camera technologies beyond framerate, including latency, motion artifacts, multi-path errors, and multi-sensor interference.
Adarsh Kowdle, Christoph Rhemann, Sean Ryan Fanello, Andrea Tagliasacchi, Jonathan Taylor 0001, Philip Davidson, Mingsong Dou, Cem Keskin, Sameh Khamis, David Kim 0002, Danhang Tang, Vladimir Tankovich, Julien P. C. Valentin, Shahram Izadi
ACM Trans. Graph.1
2018 LookinGood: enhancing performance capture with real-time neural re-rendering
abstract
Motivated by augmented and virtual reality applications such as telepresence, there has been a recent focus in real-time performance capture of humans under motion. However, given the real-time constraint, these systems often suffer from artifacts in geometry and texture such as holes and noise in the final rendering, poor lighting, and low-resolution textures. We take the novel approach to augment such real-time performance capture systems with a deep architecture that takes a rendering from an arbitrary viewpoint, and jointly performs completion, super resolution, and denoising of the imagery in real-time. We call this approach neural (re-)rendering , and our live system "LookinGood". Our deep architecture is trained to produce high resolution and high quality images from a coarse rendering in real-time. First, we propose a self-supervised training method that does not require manual ground-truth annotation. We contribute a specialized reconstruction error that uses semantic information to focus on relevant parts of the subject, e.g. the face. We also introduce a salient reweighing scheme of the loss function that is able to discard outliers. We specifically design the system for virtual and augmented reality headsets where the consistency between the left and right eye plays a crucial role in the final user experience. Finally, we generate temporally stable results by explicitly minimizing the difference between two consecutive frames. We tested the proposed system in two different scenarios: one involving a single RGB-D sensor, and upper body reconstruction of an actor, the second consisting of full body 360° capture. Through extensive experimentation, we demonstrate how our system generalizes across unseen sequences and subjects.
Ricardo Martin-Brualla, Rohit Pandey, Shuoran Yang, Pavel Pidlypenskyi, Jonathan Taylor 0001, Julien P. C. Valentin, Sameh Khamis, Philip Davidson, Anastasia Tkach, Peter Lincoln, Adarsh Kowdle, Christoph Rhemann, Dan B. Goldman, Cem Keskin, Steven M. Seitz, Shahram Izadi, Sean Ryan Fanello
ACM Trans. Graph.11
2018 Real-time compression and streaming of 4D performances
abstract
We introduce a realtime compression architecture for 4D performance capture that is two orders of magnitude faster than current state-of-the-art techniques, yet achieves comparable visual quality and bitrate. We note how much of the algorithmic complexity in traditional 4D compression arises from the necessity to encode geometry using an explicit model (i.e. a triangle mesh). In contrast, we propose an encoder that leverages an implicit representation (namely a Signed Distance Function) to represent the observed geometry, as well as its changes through time. We demonstrate how SDFs, when defined over a small local region (i.e. a block), admit a low-dimensional embedding due to the innate geometric redundancies in their representation. We then propose an optimization that takes a Truncated SDF (i.e. a TSDF), such as those found in most rigid/non-rigid reconstruction pipelines, and efficiently projects each TSDF block onto the SDF latent space. This results in a collection of low entropy tuples that can be effectively quantized and symbolically encoded. On the decoder side, to avoid the typical artifacts of block-based coding, we also propose a variational optimization that compensates for quantization residuals in order to penalize unsightly discontinuities in the decompressed signal. This optimization is expressed in the SDF latent embedding, and hence can also be performed efficiently. We demonstrate our compression/decompression architecture by realizing, to the best of our knowledge, the first system for streaming a real-time captured 4D performance on consumer-level networks.
Danhang Tang, Mingsong Dou, Peter Lincoln, Philip Davidson, Jonathan Taylor 0001, Sean Ryan Fanello, Cem Keskin, Adarsh Kowdle, Sofien Bouaziz, Shahram Izadi, Andrea Tagliasacchi
ACM Trans. Graph.9
2018 Depth from motion for smartphone AR
abstract
Augmented reality (AR) for smartphones has matured from a technology for earlier adopters, available only on select high-end phones, to one that is truly available to the general public. One of the key breakthroughs has been in low-compute methods for six degree of freedom (6DoF) tracking on phones using only the existing hardware (camera and inertial sensors). 6DoF tracking is the cornerstone of smartphone AR allowing virtual content to be precisely locked on top of the real world. However, to really give users the impression of believable AR, one requires mobile depth. Without depth, even simple effects such as a virtual object being correctly occluded by the real-world is impossible. However, requiring a mobile depth sensor would severely restrict the access to such features. In this article, we provide a novel pipeline for mobile depth that supports a wide array of mobile phones, and uses only the existing monocular color sensor. Through several technical contributions, we provide the ability to compute low latency dense depth maps using only a single CPU core of a wide range of (medium-high) mobile phones. We demonstrate the capabilities of our approach on high-level AR applications including real-time navigation and shopping.
Julien P. C. Valentin, Adarsh Kowdle, Jonathan T. Barron, Neal Wadhwa, Maksym Dzitsiuk, Michael Schoenberg, Ambrus Csaszar, Eric Turner 0001, Ivan Dryanovski, João Afonso, Jose Pascoal, Konstantine Tsotsos, Mira Leung, Mirko Schmidt, Onur G. Guleryuz, Sameh Khamis, Vladimir Tankovich, Sean Ryan Fanello, Shahram Izadi, Christoph Rhemann
ACM Trans. Graph.2
2017 UltraStereo: Efficient Learning-Based Matching for Active Stereo Systems
abstract
Efficient estimation of depth from pairs of stereo images is one of the core problems in computer vision. We efficiently solve the specialized problem of stereo matching under active illumination using a new learning-based algorithm. This type of active stereo i.e. stereo matching where scene texture is augmented by an active light projector is proving compelling for designing depth cameras, largely due to improved robustness when compared to time of flight or traditional structured light techniques. Our algorithm uses an unsupervised greedy optimization scheme that learns features that are discriminative for estimating correspondences in infrared images. The proposed method optimizes a series of sparse hyperplanes that are used at test time to remap all the image patches into a compact binary representation in O(1). The proposed algorithm is cast in a PatchMatch Stereo-like framework, producing depth maps at 500Hz. In contrast to standard structured light methods, our approach generalizes to different scenes, does not require tedious per camera calibration procedures and is not adversely affected by interference from overlapping sensors. Extensive evaluations show we surpass the quality and overcome the limitations of current depth sensing technologies.
Sean Ryan Fanello, Julien P. C. Valentin, Christoph Rhemann, Adarsh Kowdle, Vladimir Tankovich, Philip Davidson, Shahram Izadi
CVPR4
2017 Low Compute and Fully Parallel Computer Vision with HashMatch
abstract
Numerous computer vision problems such as stereo depth estimation, object-class segmentation and fore-ground/background segmentation can be formulated as per-pixel image labeling tasks. Given one or many images as input, the desired output of these methods is usually a spatially smooth assignment of labels. The large amount of such computer vision problems has lead to significant research efforts, with the state of art moving from CRF-based approaches to deep CNNs and more recently, hybrids of the two. Although these approaches have significantly advanced the state of the art, the vast majority has solely focused on improving quantitative results and are not designed for low-compute scenarios. In this paper, we present a new general framework for a variety of computer vision labeling tasks, called HashMatch. Our approach is designed to be both fully parallel, i.e. each pixel is independently processed, and low-compute, with a model complexity an order of magnitude less than existing CNN and CRF-based approaches. We evaluate HashMatch extensively on several problems such as disparity estimation, image retrieval, feature approximation and background subtraction, for which HashMatch achieves high computational efficiency while producing high quality results.
Sean Ryan Fanello, Julien P. C. Valentin, Adarsh Kowdle, Christoph Rhemann, Vladimir Tankovich, Carlo Ciliberto, Philip Davidson, Shahram Izadi
ICCV3
2017 Motion2fusion: real-time volumetric performance capture
abstract
We present Motion2Fusion, a state-of-the-art 360 performance capture system that enables *real-time* reconstruction of arbitrary non-rigid scenes. We provide three major contributions over prior work: 1) a new non-rigid fusion pipeline allowing for far more faithful reconstruction of high frequency geometric details, avoiding the over-smoothing and visual artifacts observed previously. 2) a high speed pipeline coupled with a machine learning technique for 3D correspondence field estimation reducing tracking errors and artifacts that are attributed to fast motions. 3) a backward and forward non-rigid alignment strategy that more robustly deals with topology changes but is still free from scene priors. Our novel performance capture system demonstrates real-time results nearing 3x speed-up from previous state-of-the-art work on the exact same GPU hardware. Extensive quantitative and qualitative comparisons show more precise geometric and texturing results with less artifacts due to fast motions or topology changes than prior art.
Mingsong Dou, Philip Davidson, Sean Ryan Fanello, Sameh Khamis, Adarsh Kowdle, Christoph Rhemann, Vladimir Tankovich, Shahram Izadi
ACM Trans. Graph.5
2017 Articulated distance fields for ultra-fast tracking of hands interacting
abstract
The state of the art in articulated hand tracking has been greatly advanced by hybrid methods that fit a generative hand model to depth data, leveraging both temporally and discriminatively predicted starting poses. In this paradigm, the generative model is used to define an energy function and a local iterative optimization is performed from these starting poses in order to find a "good local minimum" (i.e. a local minimum close to the true pose). Performing this optimization quickly is key to exploring more starting poses, performing more iterations and, crucially, exploiting high frame rates that ensure that temporally predicted starting poses are in the basin of convergence of a good local minimum. At the same time, a detailed and accurate generative model tends to deepen the good local minima and widen their basins of convergence. Recent work, however, has largely had to trade-off such a detailed hand model with one that facilitates such rapid optimization. We present a new implicit model of hand geometry that mostly avoids this compromise and leverage it to build an ultra-fast hybrid hand tracking system. Specifically, we construct an articulated signed distance function that, for any pose, yields a closed form calculation of both the distance to the detailed surface geometry and the necessary derivatives to perform gradient based optimization. There is no need to introduce or update any explicit "correspondences" yielding a simple algorithm that maps well to parallel hardware such as GPUs. As a result, our system can run at extremely high frame rates (e.g. up to 1000fps). Furthermore, we demonstrate how to detect, segment and optimize for two strongly interacting hands, recovering complex interactions at extremely high framerates. In the absence of publicly available datasets of sufficiently high frame rate, we leverage a multiview capture system to create a new 180fps dataset of one and two hands interacting together or with objects.
Jonathan Taylor 0001, Vladimir Tankovich, Danhang Tang, Cem Keskin, David Kim 0002, Philip Davidson, Adarsh Kowdle, Shahram Izadi
ACM Trans. Graph.7
2016 HyperDepth: Learning Depth from Structured Light without Matching
abstract
Structured light sensors are popular due to their robustness to untextured scenes and multipath. These systems triangulate depth by solving a correspondence problem between each camera and projector pixel. This is often framed as a local stereo matching task, correlating patches of pixels in the observed and reference image. However, this is computationally intensive, leading to reduced depth accuracy and framerate. We contribute an algorithm for solving this correspondence problem efficiently, without compromising depth accuracy. For the first time, this problem is cast as a classification-regression task, which we solve extremely efficiently using an ensemble of cascaded random forests. Our algorithm scales in number of disparities, and each pixel can be processed independently, and in parallel. No matching or even access to the corresponding reference pattern is required at runtime, and regressed labels are directly mapped to depth. Our GPU-based algorithm runs at a 1KHz for 1.3MP input/output images, with disparity error of 0.1 subpixels. We show a prototype high framerate depth camera running at 375Hz, useful for solving tracking-related problems. We demonstrate our algorithmic performance, creating high resolution real-time depth maps that surpass the quality of current state of the art depth technologies, highlighting quantization-free results with reduced holes, edge fattening and other stereo-based depth artifacts.
Sean Ryan Fanello, Christoph Rhemann, Vladimir Tankovich, Adarsh Kowdle, Sergio Orts, David Kim 0002, Shahram Izadi
CVPR4
2016 Holoportation: Virtual 3D Teleportation in Real-time
abstract
We present an end-to-end system for augmented and virtual reality telepresence, called Holoportation. Our system demonstrates high-quality, real-time 3D reconstructions of an entire space, including people, furniture and objects, using a set of new depth cameras. These 3D models can also be transmitted in real-time to remote users. This allows users wearing virtual or augmented reality displays to see, hear and interact with remote participants in 3D, almost as if they were present in the same physical space. From an audio-visual perspective, communicating and interacting with remote users edges closer to face-to-face communication. This paper describes the Holoportation technical system in full, its key interactive capabilities, the application scenarios it enables, and an initial qualitative study of using this new communication medium.
Sergio Orts, Christoph Rhemann, Sean Ryan Fanello, Wayne Chang, Adarsh Kowdle, Yury Degtyarev, David Kim 0002, Philip Davidson, Sameh Khamis, Mingsong Dou, Vladimir Tankovich, Charles T. Loop, Qin Cai, Philip A. Chou, Sarah Mennicken, Julien P. C. Valentin, Vivek Pradeep, Shenlong Wang, Sing Bing Kang, Pushmeet Kohli, Yuliya Lutchyn, Cem Keskin, Shahram Izadi
UIST5
2016 Fusion4D: real-time performance capture of challenging scenes
abstract
We contribute a new pipeline for live multi-view performance capture, generating temporally coherent high-quality reconstructions in real-time. Our algorithm supports both incremental reconstruction, improving the surface estimation over time, as well as parameterizing the nonrigid scene motion. Our approach is highly robust to both large frame-to-frame motion and topology changes, allowing us to reconstruct extremely challenging scenes. We demonstrate advantages over related real-time techniques that either deform an online generated template or continually fuse depth data nonrigidly into a single reference model. Finally, we show geometric reconstruction results on par with offline methods which require orders of magnitude more processing time and many more RGBD cameras.
Mingsong Dou, Sameh Khamis, Yury Degtyarev, Philip Davidson, Sean Ryan Fanello, Adarsh Kowdle, Sergio Orts, Christoph Rhemann, David Kim 0002, Jonathan Taylor 0001, Pushmeet Kohli, Vladimir Tankovich, Shahram Izadi
ACM Trans. Graph.6
2014 Putting the User in the Loop for Image-Based Modeling
Adarsh Kowdle, Yao-Jen Chang, Andrew C. Gallagher, Dhruv Batra, Tsuhan Chen
Int. J. Comput. Vis.1
2013 Revisiting Depth Layers from Occlusions
abstract
In this work, we consider images of a scene with a moving object captured by a static camera. As the object (human or otherwise) moves about the scene, it reveals pairwise depth-ordering or occlusion cues. The goal of this work is to use these sparse occlusion cues along with monocular depth occlusion cues to densely segment the scene into depth layers. We cast the problem of depth-layer segmentation as a discrete labeling problem on a spatio-temporal Markov Random Field (MRF) that uses the motion occlusion cues along with monocular cues and a smooth motion prior for the moving object. We quantitatively show that depth ordering produced by the proposed combination of the depth cues from object motion and monocular occlusion cues are superior to using either feature independently, and using a naive combination of the features.
Adarsh Kowdle, Andrew C. Gallagher, Tsuhan Chen
CVPR1
2012 Learning to Segment a Video to Clips Based on Scene and Camera Motion
Adarsh Kowdle, Tsuhan Chen
ECCV (3)1
2012 Multiple View Object Cosegmentation Using Appearance and Stereo Cues
Adarsh Kowdle, Sudipta N. Sinha, Richard Szeliski
ECCV (5)1
2012 Recovering depth of a dynamic scene using real world motion prior
abstract
Given a video of a dynamic scene captured using a dynamic camera, we present a method to recover a dense depth map of the scene with a focus on estimating the depth of the dynamic objects. We assume that the static portions of the scene help estimate the pose of the cameras. We recover a dense depth map of the scene via a plane sweep stereo approach. The relative motion of the dynamic object in the scene however, results in an inaccurate depth estimate. Estimating the accurate depth of the dynamic object is an ambiguous problem since both the depth and the real world speed of the object are unknown. In this work, we show that by using occlusions and putting constraints on the speed of the object we can bound the depth of the object. We can then incorporate this real world motion into the plane sweep stereo framework to obtain a more accurate depth for the dynamic object. We focus on videos with people walking in the scene and show the effectiveness of our approach through quantitative and qualitative results.
Adarsh Kowdle, Noah Snavely, Tsuhan Chen
ICIP1
2012 Toward Holistic Scene Understanding: Feedback Enabled Cascaded Classification Models
abstract
Scene understanding includes many related subtasks, such as scene categorization, depth estimation, object detection, etc. Each of these subtasks is often notoriously hard, and state-of-the-art classifiers already exist for many of them. These classifiers operate on the same raw image and provide correlated outputs. It is desirable to have an algorithm that can capture such correlation without requiring any changes to the inner workings of any classifier. We propose Feedback Enabled Cascaded Classification Models (FE-CCM), that jointly optimizes all the subtasks while requiring only a "black box" interface to the original classifier for each subtask. We use a two-layer cascade of classifiers, which are repeated instantiations of the original ones, with the output of the first layer fed into the second layer as input. Our training method involves a feedback step that allows later classifiers to provide earlier classifiers information about which error modes to focus on. We show that our method significantly improves performance in all the subtasks in the domain of scene understanding, where we consider depth estimation, scene categorization, event categorization, object detection, geometric labeling, and saliency detection. Our method also improves performance in two robotic applications: an object-grasping robot and an object-finding robot.
Adarsh Kowdle, Ashutosh Saxena, Tsuhan Chen
IEEE Trans. Pattern Anal. Mach. Intell.2
2011 Active learning for piecewise planar 3D reconstruction
abstract
This paper presents an active-learning algorithm for piecewise planar 3D reconstruction of a scene. While previous interactive algorithms require the user to provide tedious interactions to identify all the planes in the scene, we build on successful ideas from the automatic algorithms and introduce the idea of active learning, thereby improving the reconstructions while considerably reducing the effort. Our algorithm first attempts to obtain a piecewise planar reconstruction of the scene automatically through an energy minimization framework. The proposed active-learning algorithm then uses intuitive cues to quantify the uncertainty of the algorithm and suggest regions, querying the user to provide support for the uncertain regions via simple scribbles. These interactions are used to suitably update the algorithm, leading to better reconstructions. We show through machine experiments and a user study that the proposed approach can intelligently query users for interactions on informative regions, and users can achieve better reconstructions of the scene faster, especially for scenes with texture-less surfaces lacking cues like lines which automatic algorithms rely on.
Adarsh Kowdle, Yao-Jen Chang, Andrew C. Gallagher, Tsuhan Chen
CVPR1
2011 Scribble based interactive 3D reconstruction via scene co-segmentation
abstract
In this paper, we present a novel interactive 3D reconstruction algorithm which renders a planar reconstruction of the scene. We consider a scenario where the user has taken a few images of a scene from multiple poses. The goal is to obtain a dense and visually pleasing reconstruction of the scene, including non-planar objects. Using simple user interactions in the form of scribbles indicating the surfaces in the scene, we develop an idea of 3D scribbles to propagate scene geometry across multiple views and perform co-segmentation of all the images into the different surfaces and non-planar objects in the scene. We show that this allows us to render a complete and pleasing reconstruction of the scene along with a volumetric rendering of the non-planar objects. We demonstrate the effectiveness of our algorithm on both outdoor and indoor scenes including the ability to handle featureless surfaces.
Adarsh Kowdle, Yao-Jen Chang, Dhruv Batra, Tsuhan Chen
ICIP1
2011 Interactively Co-segmentating Topically Related Images with Intelligent Scribble Guidance
Dhruv Batra, Adarsh Kowdle, Devi Parikh, Jiebo Luo 0001, Tsuhan Chen
Int. J. Comput. Vis.2
2010 iCoseg: Interactive co-segmentation with intelligent scribble guidance
abstract
This paper presents an algorithm for Interactive Co-segmentation of a foreground object from a group of related images. While previous approaches focus on unsupervised co-segmentation, we use successful ideas from the interactive object-cutout literature. We develop an algorithm that allows users to decide what foreground is, and then guide the output of the co-segmentation algorithm towards it via scribbles. Interestingly, keeping a user in the loop leads to simpler and highly parallelizable energy functions, allowing us to work with significantly more images per group. However, unlike the interactive single image counterpart, a user cannot be expected to exhaustively examine all cutouts (from tens of images) returned by the system to make corrections. Hence, we propose iCoseg, an automatic recommendation system that intelligently recommends where the user should scribble next. We introduce and make publicly available the largest co-segmentation datasetyet, the CMU-Cornell iCoseg Dataset, with 38 groups, 643 images, and pixelwise hand-annotated groundtruth. Through machine experiments and real user studies with our developed interface, we show that iCoseg can intelligently recommend regions to scribble on, and users following these recommendations can achieve good quality cutouts with significantly lower time and effort than exhaustively examining all cutouts.
Dhruv Batra, Adarsh Kowdle, Devi Parikh, Jiebo Luo 0001, Tsuhan Chen
CVPR2
2010 Video categorization using object of interest detection
abstract
Object of Interest (OOI) detection has been widely used in many recent works in video analysis, especially in video similarity and video retrieval. In this paper, we describe a generic video classification algorithm using object of interest detection. We use online user-submitted videos and aim to categorize the videos into six broad categories hot star, news, anime, pets, sports and commercials. We show through our experiments that, detecting and describing the object of interest improves the video classification accuracy by about 10 percentage points.
Adarsh Kowdle, Kuo-Wei Chang, Tsuhan Chen
ICIP1
2010 Towards Holistic Scene Understanding: Feedback Enabled Cascaded Classification Models
abstract
In many machine learning domains (such as scene understanding), several related sub-tasks (such as scene categorization, depth estimation, object detection) operate on the same raw data and provide correlated outputs. Each of these tasks is often notoriously hard, and state-of-the-art classifiers already exist for many sub-tasks. It is desirable to have an algorithm that can capture such correlation without requiring to make any changes to the inner workings of any classifier. We propose Feedback Enabled Cascaded Classification Models (FE-CCM), that maximizes the joint likelihood of the sub-tasks, while requiring only a ‘black-box’ interface to the original classifier for each sub-task. We use a two-layer cascade of classifiers, which are repeated instantiations of the original ones, with the output of the first layer fed into the second layer as input. Our training method involves a feedback step that allows later classifiers to provide earlier classifiers information about what error modes to focus on. We show that our method significantly improves performance in all the sub-tasks in two different domains: (i) scene understanding, where we consider depth estimation, scene categorization, event categorization, object detection, geometric labeling and saliency detection, and (ii) robotic grasping, where we consider grasp point detection and object classification.
Adarsh Kowdle, Ashutosh Saxena, Tsuhan Chen
NIPS2
2009 Seed Image Selection in interactive cosegmentation
abstract
Interactive image segmentation is a powerful paradigm that allows users to direct the segmentation algorithm towards a desired output. However, marking scribbles on multiple images is a cumbersome process. Recent works show that statistics collected from user input in a single image can be shared among a group of related images to perform interactive cosegmentation. Most works use a naive heuristic of requesting the user input on a random image from the group. We show that in practice, selecting the right image to scribble on is critical to the resulting segmentation quality. In this paper, we address the problem of seed image selection, i.e., deciding which image among a group of related images should be presented to the user for scribbling. We formulate our approach as a classification problem and show that our approach outperforms the naive heuristic used by other works.
Dhruv Batra, Devi Parikh, Adarsh Kowdle, Tsuhan Chen, Jiebo Luo 0001
ICIP3