Kiran Varanasi

dblp:92/1624 · DBLP profile ↗
← Back
21ranked-venue papers
1as first author
0since 2021 · last 2019
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 21 · 1 first-authorArtificial intelligence and machine learning · 7 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
3D vision · 80% Deep learning architectures and training · 20%
Computer graphics and multimedia
5 papers
Geometric modeling and processing · 49% Computer animation and physical simulation · 38% Rendering · 9%

Topics — the 16 heaviest of 20, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Deep learning architectures and training
convolutional neural network
0.422018
CNN-Based Patch Matching for Optical Flow with Thresholded Hinge Embedding Loss · CVPR 2017
Learning 3D Shapes as Multi-layered Height-Maps Using 2D Convolutional Networks · ECCV (16) 2018
Computer vision › 3D vision
3d shape representation
0.312018
Learning 3D Shapes as Multi-layered Height-Maps Using 2D Convolutional Networks · ECCV (16) 2018
Computer vision › 3D vision › motion estimation
optical flow
0.312017
CNN-Based Patch Matching for Optical Flow with Thresholded Hinge Embedding Loss · CVPR 2017
Computer vision › 3D vision › feature matching › local feature matching
patch matching
0.312017
CNN-Based Patch Matching for Optical Flow with Thresholded Hinge Embedding Loss · CVPR 2017
Computer animation and physical simulation
facial animation
0.212016
Reconstruction of Personalized 3D Face Rigs from Monocular Video · ACM Trans. Graph. 2016
Geometric modeling and processing › shape deformation
deformation component analysis
0.212013
Sparse localized deformation components · ACM Trans. Graph. 2013
Computer animation and physical simulation
mesh animation
0.212013
Sparse localized deformation components · ACM Trans. Graph. 2013
Geometric modeling and processing
mesh deformation
0.212013
Sparse localized deformation components · ACM Trans. Graph. 2013
Computer animation and physical simulation
performance capture
0.112012
Full Body Performance Capture under Uncontrolled and Varying Illumination: A Shading-Based Approach · ECCV (4) 2012
Computer vision › 3D vision › 3d scene reconstruction
dynamic scene reconstruction
0.112011
Shading-based dynamic shape refinement from multi-view video under general illumination · ICCV 2011
Computer vision › 3D vision › 3d reconstruction
multi-view reconstruction
0.112011
Shading-based dynamic shape refinement from multi-view video under general illumination · ICCV 2011
Computer vision › 3D vision › shape from shading
shading-based refinement
0.112011
Shading-based dynamic shape refinement from multi-view video under general illumination · ICCV 2011
Computer vision › 3D vision › 3d shape modeling
shape refinement
0.112011
Shading-based dynamic shape refinement from multi-view video under general illumination · ICCV 2011
Geometric modeling and processing
mesh processing
0.112009
Surface feature detection and description with applications to mesh matching · CVPR 2009
Computer vision › 3D vision
surface tracking
0.112008
Temporal Surface Tracking Using Mesh Evolution · ECCV (2) 2008
Computer animation and physical simulation
character animation
0.012013
Sparse localized deformation components · ACM Trans. Graph. 2013

Methods — techniques the papers use, named apart from their topics

2d convolutional networks · 0.3siamese network · 0.3hinge embedding loss · 0.3variational fitting · 0.2sparse linear regression · 0.2parametric shape prior · 0.2scale-invariant feature detection · 0.2MeshHOG descriptor · 0.2MeshDOG detector · 0.2sparse matrix decomposition · 0.2optimization · 0.2mesh evolution · 0.2shading-based shape reconstruction · 0.1maximum a posteriori inference · 0.1illumination estimation · 0.1albedo estimation · 0.1
YearPublicationVenuePosition
2019 Face It!: A Pipeline for Real-Time Performance-Driven Facial Animation
abstract
This paper presents a new lightweight approach for real-time performance-driven facial animation from monocular videos. We transfer facial expressions from 2D images to a 3D virtual character, by estimating the rigid head pose and non-rigid face deformation from detected and tracked 2D facial landmarks. We map the input face into the facial expression space of the 3D head model using blendshape models and formulate a lightweight energy-based optimization problem, which is solved by non-linear least squares at 18 FPS on a single CPU. Our method robustly handles varying head poses and different facial expressions, including moderately asymmetric ones. Compared to related methods, our approach does not require training data, specialised camera setups or graphics cards, and is suitable for embedded systems. We support our claims with several experiments.
Jilliam María Díaz Barros, Vladislav Golyanik, Kiran Varanasi, Didier Stricker
ICIP3
2018 DeepHPS: End-to-end Estimation of 3D Hand Pose and Shape by Learning from Synthetic Depth
abstract
Articulated hand pose and shape estimation is an important problem for vision-based applications such as augmented reality and animation.In contrast to the existing methods which optimize only for joint positions, we propose a fully supervised deep network which learns to jointly estimate a full 3D hand mesh representation and pose from a single depth image.To this end, a CNN architecture is employed to estimate parametric representations i.e. hand pose, bone scales and complex shape parameters. Then, a novel hand pose and shape layer, embedded inside our deep framework, produces 3D joint positions and hand mesh. Lack of sufficient training data with varying hand shapes limits the generalized performance of learning based methods. Also, manually annotating real data is suboptimal. Therefore, we present SynHand5M: a million-scale synthetic benchmark with accurate joint annotations, segmentation masks and mesh files of depth maps. Among model based learning (hybrid) methods, we show improved results on two of the public benchmarks i.e. NYU and ICVL. Also, by employing a joint training strategy with real and synthetic data, we recover 3D hand mesh and pose from real images in 30ms.
Jameel Malik, Ahmed Elhayek, Fabrizio Nunnari, Kiran Varanasi, Kiarash Tamaddon, Alexis Héloir, Didier Stricker
3DV4
2018 Structured Low-Rank Matrix Factorization for Point-Cloud Denoising
abstract
In this work we address the problem of point-cloud denoising where we assume that a given point-cloud comprises (noisy) points that were sampled from an underlying surface that is to be denoised. We phrase the point-cloud denoising problem in terms of a dictionary learning framework. To this end, for a given point-cloud we (robustly) extract planar patches covering the entire point-cloud, where each patch contains a (noisy) description of the local structure of the underlying surface. Based on the general assumption that many of the local patches (in the noise-free point-cloud) contain redundant information (e.g. due to smoothness of the surface, or due to repetitive structures), we find a low-dimensional affine subspace that (approximately) explains the extracted (noisy) patches. Computationally, this is achieved by solving a structured low-rank matrix factorization problem, where we impose smoothness on the patch dictionary and sparsity on the coefficients. We experimentally demonstrate that our method outperforms existing denoising approaches in various noise scenarios.
Kripasindhu Sarkar, Florian Bernard 0001, Kiran Varanasi, Christian Theobalt, Didier Stricker
3DV3
2018 Learning 3D Shapes as Multi-layered Height-Maps Using 2D Convolutional Networks
Kripasindhu Sarkar, Basavaraj Hampiholi, Kiran Varanasi, Didier Stricker
ECCV (16)3
2018 Fusion of Keypoint Tracking and Facial Landmark Detection for Real-Time Head Pose Estimation
abstract
In this paper, we address the problem of extreme head pose estimation from intensity images, in a monocular setup. We introduce a novel fusion pipeline to integrate into a dedicated Kalman Filter the pose estimated from a tracking scheme in the prediction stage and the pose estimated from a detection scheme in the correction stage. To that end, the measurement covariance of the Kalman Filter is updated in every frame. The tracking scheme is performed using a set of keypoints extracted in the area of the head along with a simple 3D geometric model. The detection scheme, on the other hand, relies on the alignment of facial landmarks in each frame combined with 3D features extracted on a head mesh. The head pose in each scheme is estimated by minimizing the reprojection error from the 3D-2D correspondences. By combining both frameworks, we extend the applicability of head pose estimation from facial landmarks to cases where these features are no longer visible. We compared the proposed method to other related approaches, showing that it can achieve state-of-the-art performance. We also demonstrate that our approach is suitable for cases with extreme head rotations and (self-) occlusions, besides being suitable for real time applications.
Jilliam María Díaz Barros, Bruno Mirbach, Frederic Garcia, Kiran Varanasi, Didier Stricker
WACV4
2018 Human Shape Capture and Tracking at Home
abstract
Human body tracking typically requires specialized capture set-ups. Although pose tracking is available in consumer devices like Microsoft Kinect, it is restricted to stick figures visualizing body part detection. In this paper, we propose a method for full 3D human body shape and motion capture of arbitrary movements from the depth channel of a single Kinect, when the subject wears casual clothes. We do not use the RGB channel or an initialization procedure that requires the subject to move around in front of the camera. This makes our method applicable for arbitrary clothing textures and lighting environments, with minimal subject intervention. Our method consists of 3D surface feature detection and articulated motion tracking, which is regularized by a statistical human body model [26]. We also propose the idea of a Consensus Mesh (CMesh) which is the 3D template of a person created from a single view point. We demonstrate tracking results on challenging poses and argue that using CMesh along with statistical body models can improve tracking accuracies. Quantitative evaluation of our dense body tracking shows that our method has very little drift which is improved by the usage of CMesh.
Saurabh Saini, Kiran Varanasi, P. J. Narayanan
WACV3
2018 3D Shape Processing by Convolutional Denoising Autoencoders on Local Patches
abstract
We propose a system for surface completion and inpainting of 3D shapes using denoising autoencoders with convolutional layers, learnt on local patches. Our method uses height map based local patches parameterized using 3D mesh quadrangulation of the low resolution input shape. This provides us sufficient amount of local 3D patch dataset to learn deep generative Convolutional Neural Networks (CNNs) for the task of repairing moderate sized holes. We design generative networks specifically suited for the 3D encoding following ideas from the recent progress in 2D inpainting, and show our results to be better than the previous methods of surface inpainting that use linear dictionary. We validate our method on both synthetic shapes and real world scans.
Kripasindhu Sarkar, Kiran Varanasi, Didier Stricker
WACV2
2017 Learning Quadrangulated Patches for 3D Shape Parameterization and Completion
abstract
We propose a novel 3D shape parameterization by surface patches, that are oriented by 3D mesh quadrangulation of the shape. By encoding 3D surface detail on local patches, we learn a patch dictionary that identifies principal surface features of the shape. Unlike previous methods, we are able to encode surface patches of variable size as determined by the user. We propose novel methods for dictionary learning and patch reconstruction based on the query of a noisy input patch with holes. We evaluate the patch dictionary towards various applications in 3D shape inpainting, denoising and compression. Our method is able to predict missing vertices and inpaint moderately sized holes. We demonstrate a complete pipeline for reconstructing the 3D mesh from the patch encoding. We validate our shape parameterization and reconstruction methods on both synthetic shapes and real world scans. We show that our patch dictionary performs successful shape completion of complicated surface textures.
Kripasindhu Sarkar, Kiran Varanasi, Didier Stricker
3DV2
2017 Fast dense feature extraction with convolutional neural networks that have pooling or striding layers
Christian Bailer, Tewodros Habtegebrial, Kiran Varanasi, Didier Stricker
BMVC3
2017 CNN-Based Patch Matching for Optical Flow with Thresholded Hinge Embedding Loss
abstract
Learning based approaches have not yet achieved their full potential in optical flow estimation, where their performance still trails heuristic approaches. In this paper, we present a CNN based patch matching approach for optical flow estimation. An important contribution of our approach is a novel thresholded loss for Siamese networks. We demonstrate that our loss performs clearly better than existing losses. It also allows to speed up training by a factor of 2 in our tests. Furthermore, we present a novel way for calculating CNN based features for different image scales, which performs better than existing methods. We also discuss new ways of evaluating the robustness of trained features for the application of patch matching for optical flow. An interesting discovery in our paper is that low-pass filtering of feature maps can increase the robustness of features created by CNNs. We proved the competitive performance of our approach by submitting it to the KITTI 2012, KITTI 2015 and MPI-Sintel evaluation portals where we obtained state-of-the-art results on all three datasets.
Christian Bailer, Kiran Varanasi, Didier Stricker
CVPR2
2016 Reconstruction of Personalized 3D Face Rigs from Monocular Video
abstract
We present a novel approach for the automatic creation of a personalized high-quality 3D face rig of an actor from just monocular video data (e.g., vintage movies). Our rig is based on three distinct layers that allow us to model the actor’s facial shape as well as capture his person-specific expression characteristics at high fidelity, ranging from coarse-scale geometry to fine-scale static and transient detail on the scale of folds and wrinkles. At the heart of our approach is a parametric shape prior that encodes the plausible subspace of facial identity and expression variations. Based on this prior, a coarse-scale reconstruction is obtained by means of a novel variational fitting approach. We represent person-specific idiosyncrasies, which cannot be represented in the restricted shape and expression space, by learning a set of medium-scale corrective shapes. Fine-scale skin detail, such as wrinkles, are captured from video via shading-based refinement, and a generative detail formation model is learned. Both the medium- and fine-scale detail layers are coupled with the parametric prior by means of a novel sparse linear regression formulation. Once reconstructed, all layers of the face rig can be conveniently controlled by a low number of blendshape expression parameters, as widely used by animation artists. We show captured face rigs and their motions for several actors filmed in different monocular video formats, including legacy footage from YouTube, and demonstrate how they can be used for 3D animation and 2D video editing. Finally, we evaluate our approach qualitatively and quantitatively and compare to related state-of-the-art methods.
Pablo Garrido 0001, Michael Zollhöfer, Dan Casas, Levi Valgaerts, Kiran Varanasi, Patrick Pérez, Christian Theobalt
ACM Trans. Graph.5
2015 VDub: Modifying Face Video of Actors for Plausible Visual Alignment to a Dubbed Audio Track
abstract
Abstract In many countries, foreign movies and TV productions are dubbed, i.e., the original voice of an actor is replaced with a translation that is spoken by a dubbing actor in the country's own language. Dubbing is a complex process that requires specific translations and accurately timed recitations such that the new audio at least coarsely adheres to the mouth motion in the video. However, since the sequence of phonemes and visemes in the original and the dubbing language are different, the video‐to‐audio match is never perfect, which is a major source of visual discomfort. In this paper, we propose a system to alter the mouth motion of an actor in a video, so that it matches the new audio track. Our paper builds on high‐quality monocular capture of 3D facial performance, lighting and albedo of the dubbing and target actors, and uses audio analysis in combination with a space‐time retrieval method to synthesize a new photo‐realistically rendered and highly detailed 3D shape model of the mouth region to replace the target performance. We demonstrate plausible visual quality of our results compared to footage that has been professionally dubbed in the traditional way, both qualitatively and through a user study.
Pablo Garrido 0001, Levi Valgaerts, H. Sarmadi, Ingmar Steiner, Kiran Varanasi, Patrick Pérez, Christian Theobalt
Comput. Graph. Forum5
2014 Compressed Manifold Modes for Mesh Processing
abstract
Abstract This paper introduces compressed eigenfunctions of the Laplace‐Beltrami operator on 3D manifold surfaces. They constitute a novel functional basis, called the compressed manifold basis, where each function has local support. We derive an algorithm, based on the alternating direction method of multipliers (ADMM), to compute this basis on a given triangulated mesh. We show that compressed manifold modes identify key shape features, yielding an intuitive understanding of the basis for a human observer, where a shape can be processed as a collection of parts. We evaluate compressed manifold modes for potential applications in shape matching and mesh abstraction. Our results show that this basis has distinct advantages over existing alternatives, indicating high potential for a wide range of use‐cases in mesh processing.
Thomas Neumann 0007, Kiran Varanasi, Christian Theobalt, Marcus A. Magnor, Markus Wacker
Comput. Graph. Forum2
2014 Interactive motion mapping for real-time character control
abstract
Abstract It is now possible to capture the 3D motion of the human body on consumer hardware and to puppet in real time skeleton‐based virtual characters. However, many characters do not have humanoid skeletons. Characters such as spiders and caterpillars do not have boned skeletons at all, and these characters have very different shapes and motions. In general, character control under arbitrary shape and motion transformations is unsolved ‐ how might these motions be mapped? We control characters with a method which avoids the rigging‐skinning pipeline — source and target characters do not have skeletons or rigs. We use interactively‐defined sparse pose correspondences to learn a mapping between arbitrary 3D point source sequences and mesh target sequences. Then, we puppet the target character in real time. We demonstrate the versatility of our method through results on diverse virtual characters with different input motion controllers. Our method provides a fast, flexible, and intuitive interface for arbitrary motion mapping which provides new ways to control characters for real‐time animation.
Helge Rhodin, James Tompkin 0001, Kwang In Kim, Kiran Varanasi, Hans-Peter Seidel, Christian Theobalt
Comput. Graph. Forum4
2013 Capture and Statistical Modeling of Arm-Muscle Deformations
abstract
Abstract We present a comprehensive data‐driven statistical model for skin and muscle deformation of the human shoulder‐arm complex. Skin deformations arise from complex bio‐physical effects such as non‐linear elasticity of muscles, fat, and connective tissue; and vary with physiological constitution of the subjects and external forces applied during motion. Thus, they are hard to model by direct physical simulation. Our alternative approach is based on learning deformations from multiple subjects performing different exercises under varying external forces. We capture the training data through a novel multi‐camera approach that is able to reconstruct fine‐scale muscle detail in motion. The resulting reconstructions from several people are aligned into one common shape parametrization, and learned using a semi‐parametric non‐linear method. Our learned data‐driven model is fast, compact and controllable with a small set of intuitive parameters – pose, body shape and external forces, through which a novice artist can interactively produce complex muscle deformations. Our method is able to capture and synthesize fine‐scale muscle bulge effects to a greater level of realism than achieved previously. We provide quantitative and qualitative validation of our method.
Thomas Neumann 0007, Kiran Varanasi, Nils Hasler, Markus Wacker, Marcus A. Magnor, Christian Theobalt
Comput. Graph. Forum2
2013 Capturing Relightable Human Performances under General Uncontrolled Illumination
abstract
Abstract We present a novel approach to create relightable free‐viewpoint human performances from multi‐view video recorded under general uncontrolled and uncalibated illumination. We first capture a multi‐view sequence of an actor wearing arbitrary apparel and reconstruct a spatio‐temporal coherent coarse 3D model of the performance using a marker‐less tracking approach. Using these coarse reconstructions, we estimate the low‐frequency component of the illumination in a spherical harmonics (SH) basis as well as the diffuse reflectance, and then utilize them to estimate the dynamic geometry detail of human actors based on shading cues. Given the high‐quality time‐varying geometry, the estimated illumination is extended to the all‐frequency domain by re‐estimating it in the wavelet basis. Finally, the high‐quality all‐frequency illumination is utilized to reconstruct the spatially‐varying BRDF of the surface. The recovered time‐varying surface geometry and spatially‐varying non‐Lambertian reflectance allow us to generate high‐quality model‐based free view‐point videos of the actor under novel illumination conditions. Our method enables plausible reconstruction of relightable dynamic scene models without a complex controlled lighting apparatus, and opens up a path towards relightable performance capture in less constrained environments and using less complex acquisition setups.
Chenglei Wu, Carsten Stoll, Yebin Liu, Kiran Varanasi, Qionghai Dai, Christian Theobalt
Comput. Graph. Forum5
2013 Sparse localized deformation components
abstract
We propose a method that extracts sparse and spatially localized deformation modes from an animated mesh sequence. To this end, we propose a new way to extend the theory of sparse matrix decompositions to 3D mesh sequence processing, and further contribute with an automatic way to ensure spatial locality of the decomposition in a new optimization framework. The extracted dimensions often have an intuitive and clear interpretable meaning. Our method optionally accepts user-constraints to guide the process of discovering the underlying latent deformation space. The capabilities of our efficient, versatile, and easy-to-implement method are extensively demonstrated on a variety of data sets and application contexts. We demonstrate its power for user friendly intuitive editing of captured mesh animations, such as faces, full body motion, cloth animations, and muscle deformations. We further show its benefit for statistical geometry processing and biomechanically meaningful animation editing. It is further shown qualitatively and quantitatively that our method outperforms other unsupervised decomposition methods and other animation parameterization approaches in the above use cases.
Thomas Neumann 0007, Kiran Varanasi, Stephan Wenger, Markus Wacker, Marcus A. Magnor, Christian Theobalt
ACM Trans. Graph.2
2012 Full Body Performance Capture under Uncontrolled and Varying Illumination: A Shading-Based Approach
Chenglei Wu, Kiran Varanasi, Christian Theobalt
ECCV (4)2
2011 Shading-based dynamic shape refinement from multi-view video under general illumination
abstract
We present an approach to add true fine-scale spatio-temporal shape detail to dynamic scene geometry captured from multi-view video footage. Our approach exploits shading information to recover the millimeter-scale surface structure, but in contrast to related approaches succeeds under general unconstrained lighting conditions. Our method starts off from a set of multi-view video frames and an initial series of reconstructed coarse 3D meshes that lack any surface detail. In a spatio-temporal maximum a posteriori probability (MAP) inference framework, our approach first estimates the incident illumination and the spatially-varying albedo map on the mesh surface for every time instant. Thereafter, albedo and illumination are used to estimate the true geometric detail visible in the images and add it to the coarse reconstructions. The MAP framework uses weak temporal priors on lighting, albedo and geometry which improve reconstruction quality yet allow for temporal variations in the data.
Chenglei Wu, Kiran Varanasi, Yebin Liu, Hans-Peter Seidel, Christian Theobalt
ICCV2
2009 Surface feature detection and description with applications to mesh matching
abstract
In this paper we revisit local feature detectors/descriptors developed for 2D images and extend them to the more general framework of scalar fields defined on 2D manifolds. We provide methods and tools to detect and describe features on surfaces equiped with scalar functions, such as photometric information. This is motivated by the growing need for matching and tracking photometric surfaces over temporal sequences, due to recent advancements in multiple camera 3D reconstruction. We propose a 3D feature detector (MeshDOG) and a 3D feature descriptor (MeshHOG) for uniformly triangulated meshes, invariant to changes in rotation, translation, and scale. The descriptor is able to capture the local geometric and/or photometric properties in a succinct fashion. Moreover, the method is defined generically for any scalar function, e.g., local curvature. Results with matching rigid and non-rigid meshes demonstrate the interest of the proposed framework.
Andrei Zaharescu, Edmond Boyer, Kiran Varanasi, Radu Horaud
CVPR3
2008 Temporal Surface Tracking Using Mesh Evolution
Kiran Varanasi, Andrei Zaharescu, Edmond Boyer, Radu Horaud
ECCV (2)1