VLDB 2026 Research / reviewers in the wild / expert
Kripasindhu Sarkar
dblp:183/3617
· DBLP profile ↗
21ranked-venue papers
8as first author
15since 2021 · last 2024
0000-0002-0220-0853ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 21 · 8 first-author · 15 since 2021Artificial intelligence and machine learning · 10 · 2 first-author · 8 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Egocentric Whole-Body Motion Capture with FisheyeViT and Diffusion-Based Motion RefinementabstractIn this work, we explore egocentric whole-body motion capture using a single fisheye camera, which simultane-ously estimates human body and hand motion. This task presents significant challenges due to three factors: the lack of high-quality datasets, fisheye camera distortion, and hu-man body self-occlusion. To address these challenges, we propose a novel approach that leverages Fisheye ViT to ex-tract fisheye image features, which are subsequently con-verted into pixel-aligned 3D heatmap representations for 3D human body pose prediction. For hand tracking, we incorporate dedicated hand detection and hand pose esti-mation networks for regressing 3D hand poses. Finally, we develop a diffusion-based whole-body motion prior model to refine the estimated whole-body motion while accounting for joint uncertainties. To train these networks, we col-lect a large synthetic dataset, Ego WholeBody, comprising 840,000 high-quality egocentric images captured across a diverse range of whole-body motion sequences. Quantitative and qualitative evaluations demonstrate the effective-ness of our method in producing high-quality whole-body motion estimates from a single egocentric camera. Jian Wang 0042, Zhe Cao 0003, Diogo C. Luvizon, Lingjie Liu, Kripasindhu Sarkar, Danhang Tang, Thabo Beeler, Christian Theobalt |
CVPR | 5 |
| 2024 | Cafca: High-quality Novel View Synthesis of Expressive Faces from Casual Few-shot CapturesabstractVolumetric modeling and neural radiance field representations have revolutionized 3D face capture and photorealistic novel view synthesis. However, these methods often require hundreds of multi-view input images and are thus inapplicable to cases with less than a handful of inputs. We present a novel volumetric prior on human faces that allows for high-fidelity expressive face modeling from as few as three input views captured in the wild. Our key insight is that an implicit prior trained on synthetic data alone can generalize to extremely challenging real-world identities and expressions and render novel views with fine idiosyncratic details like wrinkles and eyelashes. We leverage a 3D Morphable Face Model to synthesize a large training set, rendering each identity with different expressions, hair, clothing, and other assets. We then train a conditional Neural Radiance Field prior on this synthetic dataset and, at inference time, fine-tune the model on a very sparse set of real images of a single subject. On average, the fine-tuning requires only three inputs to cross the synthetic-to-real domain gap. The resulting personalized 3D model reconstructs strong idiosyncratic facial expressions and outperforms the state-of-the-art in high-quality novel view synthesis of faces from sparse inputs in terms of perceptual and photo-metric quality. Marcel C. Bühler, Gengyan Li 0001, Erroll Wood, Leonhard Helminger, Xu Chen 0025, Tanmay Shah, Daoye Wang, Stephan J. Garbin, Sergio Orts, Otmar Hilliges, Dmitry Lagun, Jérémy Riviere, Paulo F. U. Gotardo, Thabo Beeler, Abhimitra Meka, Kripasindhu Sarkar |
SIGGRAPH Asia | 16 |
| 2024 | GANtlitz: Ultra High Resolution Generative Model for Multi-Modal Face TexturesabstractAbstract High‐resolution texture maps are essential to render photoreal digital humans for visual effects or to generate data for machine learning. The acquisition of high resolution assets at scale is cumbersome, it involves enrolling a large number of human subjects, using expensive multi‐view camera setups, and significant manual artistic effort to align the textures. To alleviate these problems, we introduce GANtlitz (A play on the german noun Antlitz, meaning face), a generative model that can synthesize multi‐modal ultra‐high‐resolution face appearance maps for novel identities. Our method solves three distinct challenges: 1) unavailability of a very large data corpus generally required for training generative models, 2) memory and computational limitations of training a GAN at ultra‐high resolutions, and 3) consistency of appearance features such as skin color, pores and wrinkles in high‐resolution textures across different modalities. We introduce dual‐style blocks, an extension to the style blocks of the StyleGAN2 architecture, which improve multi‐modal synthesis. Our patch‐based architecture is trained only on image patches obtained from a small set of face textures (<100) and yet allows us to generate seamless appearance maps of novel identities at 6k × 4k resolution. Extensive qualitative and quantitative evaluations and baseline comparisons show the efficacy of our proposed system. (see https://www.acm.org/publications/class-2012 ) Aurel Gruber, Edo Collins, Abhimitra Meka, Franziska Mueller 0001, Kripasindhu Sarkar, Sergio Orts, Luca Prasso, Jay Busch, Markus Gross 0001, Thabo Beeler |
Comput. Graph. Forum | 5 |
| 2024 | ShellNeRF: Learning a Controllable High-resolution Model of the Eye and Periocular RegionabstractAbstract Eye gaze and expressions are crucial non‐verbal signals in face‐to‐face communication. Visual effects and telepresence demand significant improvements in personalized tracking, animation, and synthesis of the eye region to achieve true immersion. Morphable face models, in combination with coordinate‐based neural volumetric representations, show promise in solving the difficult problem of reconstructing intricate geometry (eyelashes) and synthesizing photorealistic appearance variations (wrinkles and specularities) of eye performances. We propose a novel hybrid representation ‐ ShellNeRF ‐ that builds a discretized volume around a 3DMM face mesh using concentric surfaces to model the deformable ‘periocular’ region. We define a canonical space using the UV layout of the shells that constrains the space of dense correspondence search. Combined with an explicit eyeball mesh for modeling corneal light‐transport, our model allows for animatable photorealistic 3D synthesis of the whole eye region. Using multi‐view video input, we demonstrate significant improvements over state‐of‐the‐art in expression re‐enactment and transfer for high‐resolution close‐up views of the eye region. Gengyan Li 0001, Kripasindhu Sarkar, Abhimitra Meka, Marcel C. Bühler, Franziska Mueller 0001, Paulo F. U. Gotardo, Otmar Hilliges, Thabo Beeler |
Comput. Graph. Forum | 2 |
| 2024 | GroomCap: High-Fidelity Prior-Free Hair CaptureabstractDespite recent advances in multi-view hair reconstruction, achieving strand-level precision remains a significant challenge due to inherent limitations in existing capture pipelines. We introduce GroomCap , a novel multi-view hair capture method that reconstructs faithful and high-fidelity hair geometry without relying on external data priors. To address the limitations of conventional reconstruction algorithms, we propose a neural implicit representation for hair volume that encodes high-resolution 3D orientation and occupancy from input views. This implicit hair volume is trained with a new volumetric 3D orientation rendering algorithm, coupled with 2D orientation distribution supervision, to effectively prevent the loss of structural information caused by undesired orientation blending. We further propose a Gaussian-based hair optimization strategy to refine the traced hair strands with a novel chained Gaussian representation, utilizing direct photometric supervision from images. Our results demonstrate that GroomCap is able to capture high-quality hair geometries that are not only more precise and detailed than existing methods but also versatile enough for a range of applications. Yuxiao Zhou 0001, Menglei Chai, Daoye Wang, Sebastian Winberg, Erroll Wood, Kripasindhu Sarkar, Markus Gross 0001, Thabo Beeler |
ACM Trans. Graph. | 6 |
| 2023 | Learning Personalized High Quality Volumetric Head Avatars from Monocular RGB VideosabstractWe propose a method to learn a high-quality implicit 3D head avatar from a monocular RGB video captured in the wild. The learnt avatar is driven by a parametric face model to achieve user-controlled facial expressions and head poses. Our hybrid pipeline combines the geometry prior and dynamic tracking of a 3DMM with a neural radiance field to achieve fine-grained control and photorealism. To reduce over-smoothing and improve out-of-model expressions synthesis, we propose to predict local features anchored on the 3DMM geometry. These learnt features are driven by 3DMM deformation and interpolated in 3D space to yield the volumetric radiance at a designated query point. We further show that using a Convolutional Neural Network in the UV space is critical in incorporating spatial context and producing representative local features. Extensive experiments show that we are able to reconstruct high-quality avatars, with more accurate expression-dependent details, good generalization to out-of-training expressions, and quantitatively superior renderings compared to other state-of-the-art approaches. Ziqian Bai, Feitong Tan, Zeng Huang, Kripasindhu Sarkar, Danhang Tang, Di Qiu, Abhimitra Meka, Ruofei Du, Mingsong Dou, Sergio Orts, Rohit Pandey, Ping Tan 0002, Thabo Beeler, Sean Ryan Fanello, Yinda Zhang 0001 |
CVPR | 4 |
| 2023 | Scene-Aware Egocentric 3D Human Pose EstimationabstractEgocentric 3D human pose estimation with a single head-mounted fisheye camera has recently attracted attention due to its numerous applications in virtual and augmented reality. Existing methods still struggle in challenging poses where the human body is highly occluded or is closely interacting with the scene. To address this issue, we propose a scene-aware egocentric pose estimation method that guides the prediction of the egocentric pose with scene constraints. To this end, we propose an egocentric depth estimation network to predict the scene depth map from a wide-view egocentric fisheye camera while mitigating the occlusion of the human body with a depth-inpainting network. Next, we propose a scene-aware pose estimation network that projects the 2D image features and estimated depth map of the scene into a voxel space and regresses the 3D pose with a V2V network. The voxel-based feature representation provides the direct geometric connection between 2D image features and scene geometry, and further facilitates the V2V network to constrain the predicted pose based on the estimated scene geometry. To enable the training of the aforementioned networks, we also generated a synthetic dataset, called EgoGTA, and an in-the-wild dataset based on EgoPW, called EgoPW-Scene. The experimental results of our new evaluation sequences show that the predicted 3D egocentric poses are accurate and physically plausible in terms of human-scene interaction, demonstrating that our method outperforms the state-of-the-art methods both quantitatively and qualitatively. Jian Wang 0042, Diogo C. Luvizon, Weipeng Xu, Lingjie Liu, Kripasindhu Sarkar, Christian Theobalt |
CVPR | 5 |
| 2023 | Preface: A Data-driven Volumetric Prior for Few-shot Ultra High-resolution Face SynthesisabstractNeRFs have enabled highly realistic synthesis of human faces including complex appearance and reflectance effects of hair and skin. These methods typically require a large number of multi-view input images, making the process hardware intensive and cumbersome, limiting applicability to unconstrained settings. We propose a novel volumetric human face prior that enables the synthesis of ultra high-resolution novel views of subjects that are not part of the prior’s training distribution. This prior model consists of an identity-conditioned NeRF, trained on a dataset of low-resolution multi-view images of diverse humans with known camera calibration. A simple sparse landmark-based 3D alignment of the training dataset allows our model to learn a smooth latent space of geometry and appearance despite a limited number of training identities. A high-quality volumetric representation of a novel subject can be obtained by model fitting to 2 or 3 camera views of arbitrary resolution. Importantly, our method requires as few as two views of casually captured images as input at inference time. Marcel C. Bühler, Kripasindhu Sarkar, Tanmay Shah, Gengyan Li 0001, Daoye Wang, Leonhard Helminger, Sergio Orts, Dmitry Lagun, Otmar Hilliges, Thabo Beeler, Abhimitra Meka |
ICCV | 2 |
| 2023 | LitNeRF: Intrinsic Radiance Decomposition for High-Quality View Synthesis and Relighting of FacesabstractHigh-fidelity, photorealistic 3D capture of a human face is a long-standing problem in computer graphics – the complex material of skin, intricate geometry of hair, and fine scale textural details make it challenging. Traditional techniques rely on very large and expensive capture rigs to reconstruct explicit mesh geometry and appearance maps, and are limited by the accuracy of hand-crafted reflectance models. More recent volumetric methods (e.g., NeRFs) have enabled view-synthesis and sometimes relighting by learning an implicit representation of the density and reflectance basis, but suffer from artifacts and blurriness due to the inherent ambiguities in volumetric modeling. These problems are further exacerbated when capturing with few cameras and light sources. We present a novel technique for high-quality capture of a human face for 3D view synthesis and relighting using a sparse, compact capture rig consisting of 15 cameras and 15 lights. Our method combines a neural volumetric representation with traditional mesh reconstruction from multiview stereo. The proxy geometry allows us to anchor the 3D density field to prevent artifacts and guide the disentanglement of intrinsic radiance components of the face appearance such as diffuse and specular reflectance, and incident radiance (shadowing) fields. Our hybrid representation significantly improves the state-of-the-art quality for arbitrarily dense renders of a face from desired camera viewpoint as well as environmental, directional, and near-field lighting. Kripasindhu Sarkar, Marcel C. Bühler, Gengyan Li 0001, Daoye Wang, Delio Vicini, Jérémy Riviere, Yinda Zhang 0001, Sergio Orts, Paulo F. U. Gotardo, Thabo Beeler, Abhimitra Meka |
SIGGRAPH Asia | 1 |
| 2022 | Estimating Egocentric 3D Human Pose in the Wild with External Weak SupervisionabstractEgocentric 3D human pose estimation with a single fisheye camera has drawn a significant amount of attention recently. However, existing methods struggle with pose estimation from in-the-wild images, because they can only be trained on synthetic data due to the unavailability of large-scale in-the-wild egocentric datasets. Furthermore, these methods easily fail when the body parts are occluded by or interacting with the surrounding scene. To address the shortage of in-the-wild data, we collect a large-scale in-the-wild egocentric dataset called Egocentric Poses in the Wild (EgoPW). This dataset is captured by a head-mounted fisheye camera and an auxiliary external camera, which provides an additional observation of the human body from a third-person perspective during training. We present a new egocentric pose estimation method, which can be trained on the new dataset with weak external supervision. Specifically, we first generate pseudo labels for the EgoPW dataset with a spatio-temporal optimization method by incorporating the external-view supervision. The pseudo labels are then used to train an egocentric pose estimation network. To facilitate the network training, we propose a novel learning strategy to supervise the egocentric features with the high-quality features extracted by a pretrained external-view pose estimation model. The experiments show that our method predicts accurate 3D poses from a single in-the-wild egocentric image and outperforms the state-of-the-art methods both quantitatively and qualitatively. Jian Wang 0042, Lingjie Liu, Weipeng Xu, Kripasindhu Sarkar, Diogo C. Luvizon, Christian Theobalt |
CVPR | 4 |
| 2021 | HumanGAN: A Generative Model of Human ImagesabstractGenerative adversarial networks achieve great performance in photorealistic image synthesis in various domains, including human images. However, they usually employ latent vectors that encode the sampled outputs globally. This does not allow convenient control of semantically relevant individual parts of the image, and cannot draw samples that only differ in partial aspects, such as clothing style. We address these limitations and present a generative model for images of dressed humans offering control over pose, local body part appearance and garment style. This is the first method to solve various aspects of human image generation, such as global appearance sampling, pose transfer, parts and garment transfer, and part sampling jointly in a unified framework. As our model encodes part-based latent appearance vectors in a normalized pose-independent space and warps them to different poses, it preserves body and clothing appearance under varying posture. Experiments show that our flexible and general generative method outperforms task-specific baselines for pose-conditioned image generation, pose transfer and part sampling in terms of realism and output resolution. Kripasindhu Sarkar, Lingjie Liu, Vladislav Golyanik, Christian Theobalt |
3DV | 1 |
| 2021 | Pose-Guided Human Animation From a Single Image in the WildabstractWe present a new pose transfer method for synthesizing a human animation from a single image of a person controlled by a sequence of body poses. Existing pose transfer methods exhibit significant visual artifacts when applying to a novel scene, resulting in temporal inconsistency and failures in preserving the identity and textures of the person. To address these limitations, we design a compositional neural network that predicts the silhouette, garment labels, and textures. Each modular network is explicitly dedicated to a subtask that can be learned from the synthetic data. At the inference time, we utilize the trained network to produce a unified representation of appearance and its labels in UV coordinates, which remains constant across poses. The unified representation provides an incomplete yet strong guidance to generating the appearance in response to the pose change. We use the trained network to complete the appearance and render it with the background. With these strategies, we are able to synthesize human animations that can preserve the identity and appearance of the person in a temporally coherent way without any fine-tuning of the network on the testing scene. Experiments show that our method outperforms the state-of-the-arts in terms of synthesis quality, temporal coherence, and generalization ability. Jae Shin Yoon, Lingjie Liu, Vladislav Golyanik, Kripasindhu Sarkar, Hyun Soo Park, Christian Theobalt |
CVPR | 4 |
| 2021 | Estimating Egocentric 3D Human Pose in Global SpaceabstractEgocentric 3D human pose estimation using a single fisheye camera has become popular recently as it allows capturing a wide range of daily activities in unconstrained environments, which is difficult for traditional outside-in motion capture with external cameras. However, existing methods have several limitations. A prominent problem is that the estimated poses lie in the local coordinate system of the fisheye camera, rather than in the world coordinate system, which is restrictive for many applications. Furthermore, these methods suffer from limited accuracy and temporal instability due to ambiguities caused by the monocular setup and the severe occlusion in a strongly distorted egocentric perspective. To tackle these limitations, we present a new method for egocentric global 3D body pose estimation using a single head-mounted fish-eye camera. To achieve accurate and temporally stable global poses, a spatio-temporal optimization is performed over a sequence of frames by minimizing heatmap reprojection errors and enforcing local and global body motion priors learned from a mocap dataset. Experimental results show that our approach outperforms state-of-the-art methods both quantitatively and qualitatively. Jian Wang 0042, Lingjie Liu, Weipeng Xu, Kripasindhu Sarkar, Christian Theobalt |
ICCV | 4 |
| 2021 | EgoRenderer: Rendering Human Avatars from Egocentric Camera ImagesabstractWe present EgoRenderer, a system for rendering full-body neural avatars of a person captured by a wearable, egocentric fisheye camera that is mounted on a cap or a VR headset. Our system renders photorealistic novel views of the actor and her motion from arbitrary virtual camera locations. Rendering full-body avatars from such egocentric images come with unique challenges due to the top-down view and large distortions. We tackle these challenges by decomposing the rendering process into several steps, including texture synthesis, pose construction, and neural image translation. For texture synthesis, we propose Ego-DPNet, a neural network that infers dense correspondences between the input fisheye images and an underlying parametric body model, and to extract textures from egocentric inputs. In addition, to encode dynamic appearances, our approach also learns an implicit texture stack that captures detailed appearance variation across poses and viewpoints. For correct pose generation, we first estimate body pose from the egocentric view using a parametric model. We then synthesize an external free-viewpoint pose image by projecting the parametric model to the user-specified target viewpoint. We next combine the target pose image and the textures into a combined feature image, which is transformed into the output color image using a neural image translation network. Experimental evaluations show that EgoRenderer is capable of generating realistic free-viewpoint avatars of a person wearing an egocentric camera. Comparisons to several baselines demonstrate the advantages of our approach. Tao Hu 0006, Kripasindhu Sarkar, Lingjie Liu, Matthias Zwicker, Christian Theobalt |
ICCV | 2 |
| 2021 | Neural actor: neural free-view synthesis of human actors with pose controlabstractWe propose Neural Actor (NA), a new method for high-quality synthesis of humans from arbitrary viewpoints and under arbitrary controllable poses. Our method is developed upon recent neural scene representation and rendering works which learn representations of geometry and appearance from only 2D images. While existing works demonstrated compelling rendering of static scenes and playback of dynamic scenes, photo-realistic reconstruction and rendering of humans with neural implicit methods, in particular under user-controlled novel poses, is still difficult. To address this problem, we utilize a coarse body model as a proxy to unwarp the surrounding 3D space into a canonical pose. A neural radiance field learns pose-dependent geometric deformations and pose- and view-dependent appearance effects in the canonical space from multi-view video input. To synthesize novel views of high-fidelity dynamic geometry and appearance, NA leverages 2D texture maps defined on the body model as latent variables for predicting residual deformations and the dynamic appearance. Experiments demonstrate that our method achieves better quality than the state-of-the-arts on playback as well as novel pose synthesis, and can even generalize well to new poses that starkly differ from the training poses. Furthermore, our method also supports shape control on the free-view synthesis of human actors. Lingjie Liu, Marc Habermann, Viktor Rudnev, Kripasindhu Sarkar, Jiatao Gu, Christian Theobalt |
ACM Trans. Graph. | 4 |
| 2020 | Neural Re-rendering of Humans from a Single Image
Kripasindhu Sarkar, Dushyant Mehta, Weipeng Xu, Vladislav Golyanik, Christian Theobalt |
ECCV (11) | 1 |
| 2019 | Simple Domain Adaptation for CAD based Object RecognitionabstractWe present a simple method of domain adaptation between synthetic images and real images - by high quality rendering of the 3D models and correlation alignment. Using this method, we solve the problem of 3D object recognition in 2D images by fine-tuning existing pretrained CNN models for the object categories using the rendered images. Experimentally, we show that our rendering pipeline along with the correlation alignment improve the recognition accuracy of existing CNN based recognition trained on rendered images - by a canonical renderer - by a large margin. Using the same idea we present a general image classifier of common objects which is trained only on the 3D models from the publicly available databases, and show that a small number of training models are sufficient to capture different variations within and across the classes. Kripasindhu Sarkar, Didier Stricker |
ICPRAM | 1 |
| 2018 | Structured Low-Rank Matrix Factorization for Point-Cloud DenoisingabstractIn this work we address the problem of point-cloud denoising where we assume that a given point-cloud comprises (noisy) points that were sampled from an underlying surface that is to be denoised. We phrase the point-cloud denoising problem in terms of a dictionary learning framework. To this end, for a given point-cloud we (robustly) extract planar patches covering the entire point-cloud, where each patch contains a (noisy) description of the local structure of the underlying surface. Based on the general assumption that many of the local patches (in the noise-free point-cloud) contain redundant information (e.g. due to smoothness of the surface, or due to repetitive structures), we find a low-dimensional affine subspace that (approximately) explains the extracted (noisy) patches. Computationally, this is achieved by solving a structured low-rank matrix factorization problem, where we impose smoothness on the patch dictionary and sparsity on the coefficients. We experimentally demonstrate that our method outperforms existing denoising approaches in various noise scenarios. Kripasindhu Sarkar, Florian Bernard 0001, Kiran Varanasi, Christian Theobalt, Didier Stricker |
3DV | 1 |
| 2018 | Learning 3D Shapes as Multi-layered Height-Maps Using 2D Convolutional Networks
Kripasindhu Sarkar, Basavaraj Hampiholi, Kiran Varanasi, Didier Stricker |
ECCV (16) | 1 |
| 2018 | 3D Shape Processing by Convolutional Denoising Autoencoders on Local PatchesabstractWe propose a system for surface completion and inpainting of 3D shapes using denoising autoencoders with convolutional layers, learnt on local patches. Our method uses height map based local patches parameterized using 3D mesh quadrangulation of the low resolution input shape. This provides us sufficient amount of local 3D patch dataset to learn deep generative Convolutional Neural Networks (CNNs) for the task of repairing moderate sized holes. We design generative networks specifically suited for the 3D encoding following ideas from the recent progress in 2D inpainting, and show our results to be better than the previous methods of surface inpainting that use linear dictionary. We validate our method on both synthetic shapes and real world scans. Kripasindhu Sarkar, Kiran Varanasi, Didier Stricker |
WACV | 1 |
| 2017 | Learning Quadrangulated Patches for 3D Shape Parameterization and CompletionabstractWe propose a novel 3D shape parameterization by surface patches, that are oriented by 3D mesh quadrangulation of the shape. By encoding 3D surface detail on local patches, we learn a patch dictionary that identifies principal surface features of the shape. Unlike previous methods, we are able to encode surface patches of variable size as determined by the user. We propose novel methods for dictionary learning and patch reconstruction based on the query of a noisy input patch with holes. We evaluate the patch dictionary towards various applications in 3D shape inpainting, denoising and compression. Our method is able to predict missing vertices and inpaint moderately sized holes. We demonstrate a complete pipeline for reconstructing the 3D mesh from the patch encoding. We validate our shape parameterization and reconstruction methods on both synthetic shapes and real world scans. We show that our patch dictionary performs successful shape completion of complicated surface textures. Kripasindhu Sarkar, Kiran Varanasi, Didier Stricker |
3DV | 1 |