VLDB 2026 Research / reviewers in the wild / expert
Kartik Teotia
dblp:243/6711
· DBLP profile ↗
7ranked-venue papers
3as first author
5since 2021 · last 2026
0009-0007-6985-7159ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GRMM: Real-Time High-Fidelity Gaussian Morphable Head Model with Learned Residualsabstract3D Morphable Models (3DMMs) enable controllable editing of facial geometry and expressions for reconstruction, animation, and AR/VR, but traditional PCA-based mesh models are limited in resolution, detail, and photorealism. Neural volumetric methods improve realism but remain too slow for interactive use. Recent Gaussian Splatting (3DGS)- based facial models achieve fast, high-quality rendering but still rely solely on a meshbased 3DMM prior for expression control, thereby limiting their ability to capture fine-grained geometry, expressions, and full-head coverage. We introduce GRMM, the first full-head Gaussian 3D morphable model that augments a base 3DMM with residual geometry and appearance components, additive refinements that recover high-frequency details such as wrinkles, fine skin texture, and hairline variations. GRMM provides disentangled control through low-dimensional, interpretable parameters (e.g., identity shape, facial expressions) while separately modelling residuals that capture subject- and expression-specific detail beyond the base model's capacity. Coarse decoders produce vertex-level mesh deformations, fine decoders represent per-Gaussian appearance,and a lightweight CNN refines rasterised images for enhanced realism, all while maintaining 75 FPS real-time rendering. To learn consistent, high-fidelity residuals, we present EXPRESS-50, the first dataset with 60 aligned expressions across 50 identities, enabling robust disentanglement of identity and expression in Gaussian-based 3DMMs. Across monocular 3D face reconstruction, novel-view synthesis, and expression transfer, GRMM surpasses state-of-the-art methods in fidelity and expression accuracy while delivering interactive real-time performance. Project page: https://mohitm1994.github.io/GRMM/ Mohit Mendiratta, Mayur Deshmukh, Kartik Teotia, Vladislav Golyanik, Adam Kortylewski, Christian Theobalt |
3DV | 3 |
| 2025 | Audio Driven Universal Gaussian Head AvatarsabstractWe introduce the first method for audio-driven universal photorealistic avatar synthesis, combining a person-agnostic speech model with our novel Universal Head Avatar Prior (UHAP). UHAP is trained on cross-identity multi-view videos. In particular, our UHAP is supervised with neutral scan data, enabling it to capture the identity-specific details at high fidelity. In contrast to previous approaches, which predominantly map audio features to geometric deformations only while ignoring audio-dependent appearance variations, our universal speech model directly maps raw audio inputs into the UHAP latent expression space. This expression space inherently encodes, both, geometric and appearance variations. For efficient personalization to new subjects, we employ a monocular encoder, which enables lightweight regression of dynamic expression variations across video frames. By accounting for these expression-dependent changes, it enables the subsequent model fine-tuning stage to focus exclusively on capturing the subject’s global appearance and geometry. Decoding these audio-driven expression codes via UHAP generates highly realistic avatars with precise lip synchronization and nuanced expressive details, such as eyebrow movement, gaze shifts, and realistic mouth interior appearance as well as motion. Extensive evaluations demonstrate that our method is not only the first generalizable audio-driven avatar model that can account for detailed appearance modeling and rendering, but it also outperforms competing (geometry-only) methods across metrics measuring lip-sync accuracy, quantitative image quality, and perceptual realism. Kartik Teotia, Helge Rhodin, Mohit Mendiratta, Hyeongwoo Kim, Marc Habermann, Christian Theobalt |
SIGGRAPH Asia | 1 |
| 2024 | GaussianHeads: End-to-End Learning of Drivable Gaussian Head Avatars from Coarse-to-fine RepresentationsabstractReal-time rendering of human head avatars is a cornerstone of many computer graphics applications, such as augmented reality, video games, and films, to name a few. Recent approaches address this challenge with computationally efficient geometry primitives in a carefully calibrated multi-view setup. Albeit producing photorealistic head renderings, they often fail to represent complex motion changes, such as the mouth interior and strongly varying head poses. We propose a new method to generate highly dynamic and deformable human head avatars from multi-view imagery in real time. At the core of our method is a hierarchical representation of head models that can capture the complex dynamics of facial expressions and head movements. First, with rich facial features extracted from raw input frames, we learn to deform the coarse facial geometry of the template mesh. We then initialize 3D Gaussians on the deformed surface and refine their positions in a fine step. We train this coarse-to-fine facial avatar model along with the head pose as learnable parameters in an end-to-end framework. This enables not only controllable facial animation via video inputs but also high-fidelity novel view synthesis of challenging facial expressions, such as tongue deformations and fine-grained teeth structure under large motion changes. Moreover, it encourages the learned head avatar to generalize towards new facial expressions and head poses at inference time. We demonstrate the performance of our method with comparisons against the related methods on different datasets, spanning challenging facial expression sequences across multiple identities. We also show the potential application of our approach by demonstrating a cross-identity facial performance transfer application. We make the code available on our project page. Kartik Teotia, Hyeongwoo Kim, Pablo Garrido 0001, Marc Habermann, Mohamed A. Elgharib, Christian Theobalt |
ACM Trans. Graph. | 1 |
| 2024 | HQ3DAvatar: High-quality Implicit 3D Head AvatarabstractMulti-view volumetric rendering techniques have recently shown great potential in modeling and synthesizing high-quality head avatars. A common approach to capture full head dynamic performances is to track the underlying geometry using a mesh-based template or 3D cube-based graphics primitives. While these model-based approaches achieve promising results, they often fail to learn complex geometric details such as the mouth interior, hair, and topological changes over time. This article presents a novel approach to building highly photorealistic digital head avatars. Our method learns a canonical space via an implicit function parameterized by a neural network. It leverages multiresolution hash encoding in the learned feature space, allowing for high quality, faster training, and high-resolution rendering. At test time, our method is driven by a monocular RGB video. Here, an image encoder extracts face-specific features that also condition the learnable canonical space. This encourages deformation-dependent texture variations during training. We also propose a novel optical flow-based loss that ensures correspondences in the learned canonical space, thus encouraging artifact-free and temporally consistent renderings. We show results on challenging facial expressions and show free-viewpoint renderings at interactive real-time rates for a resolution of 480 x 270. Our method outperforms related approaches both visually and numerically. We will release our multiple-identity dataset to encourage further research. Kartik Teotia, Mallikarjun B. R. 0001, Xingang Pan, Hyeongwoo Kim, Pablo Garrido 0001, Mohamed A. Elgharib, Christian Theobalt |
ACM Trans. Graph. | 1 |
| 2023 | AvatarStudio: Text-Driven Editing of 3D Dynamic Human Head AvatarsabstractCapturing and editing full-head performances enables the creation of virtual characters with various applications such as extended reality and media production. The past few years witnessed a steep rise in the photorealism of human head avatars. Such avatars can be controlled through different input data modalities, including RGB, audio, depth, IMUs, and others. While these data modalities provide effective means of control, they mostly focus on editing the head movements such as the facial expressions, head pose, and/or camera viewpoint. In this paper, we propose AvatarStudio, a text-based method for editing the appearance of a dynamic full head avatar. Our approach builds on existing work to capture dynamic performances of human heads using Neural Radiance Field (NeRF) and edits this representation with a text-to-image diffusion model. Specifically, we introduce an optimization strategy for incorporating multiple keyframes representing different camera viewpoints and time stamps of a video performance into a single diffusion model. Using this personalized diffusion model, we edit the dynamic NeRF by introducing view-and-time-aware Score Distillation Sampling (VT-SDS) following a model-based guidance approach. Our method edits the full head in a canonical space and then propagates these edits to the remaining time steps via a pre-trained deformation network. We evaluate our method visually and numerically via a user study, and results show that our method outperforms existing approaches. Our experiments validate the design choices of our method and highlight that our edits are genuine, personalized, as well as 3D- and time-consistent. Mohit Mendiratta, Xingang Pan, Mohamed A. Elgharib, Kartik Teotia, Mallikarjun B. R. 0001, Ayush Tewari, Vladislav Golyanik, Adam Kortylewski, Christian Theobalt |
ACM Trans. Graph. | 4 |
| 2019 | Automatic Segmentation of Common Carotid Artery in Longitudinal Mode Ultrasound Images Using Active OblongsabstractWe propose a fully automated algorithm for the segmentation of common carotid artery in longitudinal mode ultrasound images using active oblongs. The problem of segmentation and subsequent delineation of lumen-intima layer is solved as an optimization of a locally defined contrast function with respect to five degrees-of-freedom that characterize the active oblong. The detection of the common carotid artery and subsequent initialization of the active oblong inside the common carotid artery region has been done using a combination of binary thresholding, Hough transform, and pixel-offset operations. The algorithm has been validated on the Brno university signal processing lab B-mode ultrasound image database, which contains 84 longitudinal mode ultrasound images of the common carotid artery. The segmentation results are validated against the ground truth provided by two practising radiologists using Jaccard and Dice similarity measures. We have achieved a detection and segmentation accuracy of 95.2% and 97.5%, respectively. J. R. Harish Kumar, Kartik Teotia, P. Kevin Raj, Jasbon Andrade, Rajagopal Kadavigere, Chandra Sekhar Seelamantula |
ICASSP | 2 |
| 2019 | Automatic Pupil Segmentation Based On Circular Active DiscsabstractBiometric iris recognition technology is one of the popular recognition technologies in biometric authentication systems. Biometric iris recognition has been observed to provide high identification efficiency relative to the other biometric characteristic. This high performance is possible only through accurately segmenting the iris region from the given ocular image. Segmentation of pupil is one of the primary steps in iris segmentation. We propose a pupil segmentation method based on a circular active disc. As a prototype, this technique uses the active discs. The initialization of the active disc is automated using normalized template matching technique. Initialization is followed by an affine transformation of the active disc which makes the active disc lock on to the pupil. The optimum outline around the pupil region is achieved by making use of the Gradient Descent algorithm which is supplemented by efficiently computing the partial derivatives by using the Green's theorem. The algorithm has been validated on two publicly available databases Casia-V3 and the IIT Delhi iris database spanning 4598 iris images in total. The corresponding Dice index scores 0.9463 and 0.9308, respectively for the databases show the robustness of the algorithm, and a high degree of agreement with the ground truth. J. R. Harish Kumar, Kartik Teotia |
TENCON | 2 |