VLDB 2026 Research / reviewers in the wild / expert
Paulo F. U. Gotardo
dblp:33/4206
· DBLP profile ↗
38ranked-venue papers
8as first author
20since 2021 · last 2026
0000-0001-8217-5848ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 34 · 6 first-author · 19 since 2021Artificial intelligence and machine learning · 14 · 6 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GIGA: Generalizable Sparse Image-Driven Gaussian HumansabstractDriving a high-quality and photorealistic full-body virtual human from a few RGB cameras is a challenging problem that has become increasingly relevant with emerging virtual reality technologies. A promising solution to democratize such technology would be a generalizable method that takes sparse multi-view images of any person and then generates photoreal free-view renderings of them. However, the state-of-the-art approaches are not scalable to very large datasets and, thus, lack diversity and photorealism. To address this problem, we propose GIGA, a novel, generalizable full-body model for rendering photoreal humans in free viewpoint, driven by a single-view or sparse multi-view video. Notably, GIGA can scale training to a few thousand subjects while maintaining high photorealism and synthesizing dynamic appearance. At the core, we introduce a MultiHeadUNet architecture, which takes an approximate RGB texture accumulated from a single or multiple sparse views and predicts 3D Gaussian primitives represented as 2D texels on top of a human body mesh. At test time, our method performs novel view synthesis of a virtual 3D Gaussian-based human from 1 to 4 input views and a tracked body template for unseen identities. Our method excels over prior works by a significant margin in terms of identity generalization capability and photorealism. Anton Zubekhin, Heming Zhu, Paulo F. U. Gotardo, Thabo Beeler, Marc Habermann, Christian Theobalt |
3DV | 3 |
| 2025 | Synthetic Prior for Few-Shot Drivable Head Avatar InversionabstractWe present SynShot, a novel method for the few-shot inversion of a drivable head avatar based on a synthetic prior. We tackle three major challenges. First, training a controllable 3D generative network requires a large number of diverse sequences, for which pairs of images and high-quality tracked meshes are not always available. Second, the use of real data is strictly regulated (e.g., under the General Data Protection Regulation, which mandates frequent deletion of models and data to accommodate a situation when participant’s consent is withdrawn). Synthetic data, free from these constraints, is an appealing alternative. Third, state-of-the-art monocular avatar models struggle to generalize to new views and expressions, lacking a strong prior and often overfitting to a specific viewpoint distribution. Inspired by machine learning models trained solely on synthetic data, we propose a method that learns a prior model from a large dataset of synthetic heads with diverse identities, expressions, and viewpoints. With few input images, SynShot fine-tunes the pretrained synthetic prior to bridge the domain gap, modeling a photorealistic head avatar that generalizes to novel expressions and viewpoints. We model the head avatar using 3D Gaussian splatting and a convolutional encoder-decoder that outputs Gaussian parameters in UV texture space. To account for the different modeling complexities over parts of the head (e.g., skin vs hair), we embed the prior with explicit control for upsampling the number of per-part primitives. Compared to SOTA monocular and GAN-based methods, SynShot significantly improves novel view and expression synthesis. Wojciech Zielonka, Stephan J. Garbin, Alexander Lattas, Georgios Kopanas, Paulo F. U. Gotardo, Thabo Beeler, Justus Thies, Timo Bolkart |
CVPR | 5 |
| 2024 | MagicMirror: Fast and High-Quality Avatar Generation with a Constrained Search Space
Armand Comas Massague, Di Qiu, Menglei Chai, Marcel C. Bühler, Amit Raj, Ruiqi Gao, Qiangeng Xu, Mark Matthews, Paulo F. U. Gotardo, Sergio Orts, Thabo Beeler |
ECCV (66) | 9 |
| 2024 | Cafca: High-quality Novel View Synthesis of Expressive Faces from Casual Few-shot CapturesabstractVolumetric modeling and neural radiance field representations have revolutionized 3D face capture and photorealistic novel view synthesis. However, these methods often require hundreds of multi-view input images and are thus inapplicable to cases with less than a handful of inputs. We present a novel volumetric prior on human faces that allows for high-fidelity expressive face modeling from as few as three input views captured in the wild. Our key insight is that an implicit prior trained on synthetic data alone can generalize to extremely challenging real-world identities and expressions and render novel views with fine idiosyncratic details like wrinkles and eyelashes. We leverage a 3D Morphable Face Model to synthesize a large training set, rendering each identity with different expressions, hair, clothing, and other assets. We then train a conditional Neural Radiance Field prior on this synthetic dataset and, at inference time, fine-tune the model on a very sparse set of real images of a single subject. On average, the fine-tuning requires only three inputs to cross the synthetic-to-real domain gap. The resulting personalized 3D model reconstructs strong idiosyncratic facial expressions and outperforms the state-of-the-art in high-quality novel view synthesis of faces from sparse inputs in terms of perceptual and photo-metric quality. Marcel C. Bühler, Gengyan Li 0001, Erroll Wood, Leonhard Helminger, Xu Chen 0025, Tanmay Shah, Daoye Wang, Stephan J. Garbin, Sergio Orts, Otmar Hilliges, Dmitry Lagun, Jérémy Riviere, Paulo F. U. Gotardo, Thabo Beeler, Abhimitra Meka, Kripasindhu Sarkar |
SIGGRAPH Asia | 13 |
| 2024 | Infinite 3D Landmarks: Improving Continuous 2D Facial Landmark DetectionabstractAbstract In this paper, we examine three important issues in the practical use of state‐of‐the‐art facial landmark detectors and show how a combination of specific architectural modifications can directly improve their accuracy and temporal stability. First, many facial landmark detectors require a face normalization step as a pre‐process, often accomplished by a separately trained neural network that crops and resizes the face in the input image. There is no guarantee that this pre‐trained network performs optimal face normalization for the task of landmark detection. Thus, we instead analyse the use of a spatial transformer network that is trained alongside the landmark detector in an unsupervised manner, jointly learning an optimal face normalization and landmark detection by a single neural network. Second, we show that modifying the output head of the landmark predictor to infer landmarks in a canonical 3D space rather than directly in 2D can further improve accuracy. To convert the predicted 3D landmarks into screen‐space, we additionally predict the camera intrinsics and head pose from the input image. As a side benefit, this allows to predict the 3D face shape from a given image only using 2D landmarks as supervision, which is useful in determining landmark visibility among other things. Third, when training a landmark detector on multiple datasets at the same time, annotation inconsistencies across datasets forces the network to produce a sub‐optimal average. We propose to add a semantic correction network to address this issue. This additional lightweight neural network is trained alongside the landmark detector, without requiring any additional supervision. While the insights of this paper can be applied to most common landmark detectors, we specifically target a recently proposed continuous 2D landmark detector to demonstrate how each of our additions leads to meaningful improvements over the state‐of‐the‐art on standard benchmarks. Prashanth Chandran, Gaspard Zoss, Paulo F. U. Gotardo, Derek Bradley |
Comput. Graph. Forum | 3 |
| 2024 | ShellNeRF: Learning a Controllable High-resolution Model of the Eye and Periocular RegionabstractAbstract Eye gaze and expressions are crucial non‐verbal signals in face‐to‐face communication. Visual effects and telepresence demand significant improvements in personalized tracking, animation, and synthesis of the eye region to achieve true immersion. Morphable face models, in combination with coordinate‐based neural volumetric representations, show promise in solving the difficult problem of reconstructing intricate geometry (eyelashes) and synthesizing photorealistic appearance variations (wrinkles and specularities) of eye performances. We propose a novel hybrid representation ‐ ShellNeRF ‐ that builds a discretized volume around a 3DMM face mesh using concentric surfaces to model the deformable ‘periocular’ region. We define a canonical space using the UV layout of the shells that constrains the space of dense correspondence search. Combined with an explicit eyeball mesh for modeling corneal light‐transport, our model allows for animatable photorealistic 3D synthesis of the whole eye region. Using multi‐view video input, we demonstrate significant improvements over state‐of‐the‐art in expression re‐enactment and transfer for high‐resolution close‐up views of the eye region. Gengyan Li 0001, Kripasindhu Sarkar, Abhimitra Meka, Marcel C. Bühler, Franziska Mueller 0001, Paulo F. U. Gotardo, Otmar Hilliges, Thabo Beeler |
Comput. Graph. Forum | 6 |
| 2023 | Continuous Landmark Detection with 3D QueriesabstractNeural networks for facial landmark detection are notoriously limited to a fixed set of landmarks in a dedicated layout, which must be specified at training time. Dedicated datasets must also be hand-annotated with the corresponding landmark configuration for training. We propose the first facial landmark detection network that can predict continuous, unlimited landmarks, allowing to specify the number and location of the desired landmarks at inference time. Our method combines a simple image feature extractor with a queried landmark predictor, and the user can specify any continuous query points relative to a 3D template face mesh as input. As it is not tied to a fixed set of landmarks, our method is able to leverage all pre-existing 2D landmark datasets for training, even if they have inconsistent landmark configurations. As a result, we present a very powerful facial landmark detector that can be trained once, and can be used readily for numerous applications like 3D face reconstruction, arbitrary face segmentation, and is even compatible with helmeted mounted cameras, and therefore could vastly simplify face tracking workflows for media and entertainment applications. Prashanth Chandran, Gaspard Zoss, Paulo F. U. Gotardo, Derek Bradley |
CVPR | 3 |
| 2023 | ReNeRF: Relightable Neural Radiance Fields with Nearfield LightingabstractRecent work on radiance fields and volumetric inverse rendering (e.g., NeRFs) has provided excellent results in building data-driven models of real scenes for novel view synthesis with high photorealism. While full control over viewpoint is achieved, scene lighting is typically "baked" into the model and cannot be changed; other methods only capture limited variation in lighting or make restrictive assumptions about the captured scene. These limitations prevent the application on arbitrary materials and novel 3D environments with complex, distinct lighting. In this paper, we target the application scenario of capturing high-fidelity assets for neural relighting in controlled studio conditions, but without requiring a dense light stage. Instead, we leverage a small number of area lights commonly used in photogrammetry. We propose ReNeRF, a relightable radiance field model based on the intuitive and powerful approach of image-based relighting, which implicitly captures global light transport (for arbitrary objects) without complex, error-prone simulations. Thus, our new method is simple and provides full control over viewpoint and lighting, without simplistic assumptions about how light interacts with the scene. In addition, ReNeRF does not rely on the usual assumption of distant lighting – during training, we explicitly account for the distance between 3D points in the volume and point samples on the light sources. Thus, at test time, we achieve better generalization to novel, continuous lighting directions, including nearfield lighting effects. Yingyan Xu, Gaspard Zoss, Prashanth Chandran, Markus Gross 0001, Derek Bradley, Paulo F. U. Gotardo |
ICCV | 6 |
| 2023 | LitNeRF: Intrinsic Radiance Decomposition for High-Quality View Synthesis and Relighting of FacesabstractHigh-fidelity, photorealistic 3D capture of a human face is a long-standing problem in computer graphics – the complex material of skin, intricate geometry of hair, and fine scale textural details make it challenging. Traditional techniques rely on very large and expensive capture rigs to reconstruct explicit mesh geometry and appearance maps, and are limited by the accuracy of hand-crafted reflectance models. More recent volumetric methods (e.g., NeRFs) have enabled view-synthesis and sometimes relighting by learning an implicit representation of the density and reflectance basis, but suffer from artifacts and blurriness due to the inherent ambiguities in volumetric modeling. These problems are further exacerbated when capturing with few cameras and light sources. We present a novel technique for high-quality capture of a human face for 3D view synthesis and relighting using a sparse, compact capture rig consisting of 15 cameras and 15 lights. Our method combines a neural volumetric representation with traditional mesh reconstruction from multiview stereo. The proxy geometry allows us to anchor the 3D density field to prevent artifacts and guide the disentanglement of intrinsic radiance components of the face appearance such as diffuse and specular reflectance, and incident radiance (shadowing) fields. Our hybrid representation significantly improves the state-of-the-art quality for arbitrarily dense renders of a face from desired camera viewpoint as well as environmental, directional, and near-field lighting. Kripasindhu Sarkar, Marcel C. Bühler, Gengyan Li 0001, Daoye Wang, Delio Vicini, Jérémy Riviere, Yinda Zhang 0001, Sergio Orts, Paulo F. U. Gotardo, Thabo Beeler, Abhimitra Meka |
SIGGRAPH Asia | 9 |
| 2023 | An Implicit Physical Face Model Driven by Expression and Styleabstract3D facial animation is often produced by manipulating facial deformation models (or rigs), that are traditionally parameterized by expression controls. A key component that is usually overlooked is expression “style", as in, how a particular expression is performed. Although it is common to define a semantic basis of expressions that characters can perform, most characters perform each expression in their own style. To date, style is usually entangled with the expression, and it is not possible to transfer the style of one character to another when considering facial animation. We present a new face model, based on a data-driven implicit neural physics model, that can be driven by both expression and style separately. At the core, we present a framework for learning implicit physics-based actuations for multiple subjects simultaneously, trained on a few arbitrary performance capture sequences from a small set of identities. Once trained, our method allows generalized physics-based facial animation for any of the trained identities, extending to unseen performances. Furthermore, it grants control over the animation style, enabling style transfer from one character to another or blending styles of different characters. Lastly, as a physics-based model, it is capable of synthesizing physical effects, such as collision handling, setting our method apart from conventional approaches. Lingchen Yang, Gaspard Zoss, Prashanth Chandran, Paulo F. U. Gotardo, Markus Gross 0001, Barbara Solenthaler, Eftychios Sifakis, Derek Bradley |
SIGGRAPH Asia | 4 |
| 2023 | A Perceptual Shape Loss for Monocular 3D Face ReconstructionabstractAbstract Monocular 3D face reconstruction is a wide‐spread topic, and existing approaches tackle the problem either through fast neural network inference or offline iterative reconstruction of face geometry. In either case carefully‐designed energy functions are minimized, commonly including loss terms like a photometric loss, a landmark reprojection loss, and others. In this work we propose a new loss function for monocular face capture, inspired by how humans would perceive the quality of a 3D face reconstruction given a particular image. It is widely known that shading provides a strong indicator for 3D shape in the human visual system. As such, our new ‘perceptual’ shape loss aims to judge the quality of a 3D face estimate using only shading cues. Our loss is implemented as a discriminator‐style neural network that takes an input face image and a shaded render of the geometry estimate, and then predicts a score that perceptually evaluates how well the shaded render matches the given image. This ‘critic’ network operates on the RGB image and geometry render alone, without requiring an estimate of the albedo or illumination in the scene. Furthermore, our loss operates entirely in image space and is thus agnostic to mesh topology. We show how our new perceptual shape loss can be combined with traditional energy terms for monocular 3D face optimization and deep neural network regression, improving upon current state‐of‐the‐art results. Christopher Otto, Prashanth Chandran, Gaspard Zoss, Markus Gross 0001, Paulo F. U. Gotardo, Derek Bradley |
Comput. Graph. Forum | 5 |
| 2023 | Graph-Based Synthesis for Skin Micro WrinklesabstractAbstract We present a novel graph‐based simulation approach for generating micro wrinkle geometry on human skin, which can easily scale up to the micro‐meter range and millions of wrinkles. The simulation first samples pores on the skin and treats them as nodes in a graph. These nodes are then connected and the resulting edges become candidate wrinkles. An iterative optimization inspired by pedestrian trail formation is then used to assign weights to those edges, i.e., to carve out the wrinkles. Finally, we convert the graph to a detailed skin displacement map using novel shape functions implemented in graphics shaders. Our simulation and displacement map creation steps expose fine controls over the appearance at real‐time framerates suitable for interactive exploration and design. We demonstrate the effectiveness of the generated wrinkles by enhancing state‐of‐art 3D reconstructions of real human subjects with simulated micro wrinkles, and furthermore propose an artist‐driven design flow for adding micro wrinkles to fictional characters. Sebastian Weiss, J. Moulin, Prashanth Chandran, Gaspard Zoss, Paulo F. U. Gotardo, Derek Bradley |
Comput. Graph. Forum | 5 |
| 2022 | Shape Transformers: Topology-Independent 3D Shape Models Using TransformersabstractAbstract Parametric 3D shape models are heavily utilized in computer graphics and vision applications to provide priors on the observed variability of an object's geometry (e.g., for faces). Original models were linear and operated on the entire shape at once. They were later enhanced to provide localized control on different shape parts separately. In deep shape models, nonlinearity was introduced via a sequence of fully‐connected layers and activation functions, and locality was introduced in recent models that use mesh convolution networks. As common limitations, these models often dictate, in one way or another, the allowed extent of spatial correlations and also require that a fixed mesh topology be specified ahead of time. To overcome these limitations, we present Shape Transformers, a new nonlinear parametric 3D shape model based on transformer architectures. A key benefit of this new model comes from using the transformer's self‐attention mechanism to automatically learn nonlinear spatial correlations for a class of 3D shapes. This is in contrast to global models that correlate everything and local models that dictate the correlation extent. Our transformer 3D shape autoencoder is a better alternative to mesh convolution models, which require specially‐crafted convolution, and down/up‐sampling operators that can be difficult to design. Our model is also topologically independent: it can be trained once and then evaluated on any mesh topology, unlike most previous methods. We demonstrate the application of our model to different datasets, including 3D faces, 3D hand shapes and full human bodies. Our experiments demonstrate the strong potential of our Shape Transformer model in several applications in computer graphics and vision. Prashanth Chandran, Gaspard Zoss, Markus Gross 0001, Paulo F. U. Gotardo, Derek Bradley |
Comput. Graph. Forum | 4 |
| 2022 | Facial Animation with Disentangled Identity and Motion using TransformersabstractAbstract We propose a 3D+time framework for modeling dynamic sequences of 3D facial shapes, representing realistic non‐rigid motion during a performance. Our work extends neural 3D morphable models by learning a motion manifold using a transformer architecture. More specifically, we derive a novel transformer‐based autoencoder that can model and synthesize 3D geometry sequences of arbitrary length. This transformer naturally determines frame‐to‐frame correlations required to represent the motion manifold, via the internal self‐attention mechanism. Furthermore, our method disentangles the constant facial identity from the time‐varying facial expressions in a performance, using two separate codes to represent neutral identity and the performance itself within separate latent subspaces. Thus, the model represents identity‐agnostic performances that can be paired with an arbitrary new identity code and fed through our new identity‐modulated performance decoder; the result is a sequence of 3D meshes for the performance with the desired identity and temporal length. We demonstrate how our disentangled motion model has natural applications in performance synthesis, performance retargeting, key‐frame interpolation and completion of missing data, performance denoising and retiming, and other potential applications that include full 3D body modeling. Prashanth Chandran, Gaspard Zoss, Markus Gross 0001, Paulo F. U. Gotardo, Derek Bradley |
Comput. Graph. Forum | 4 |
| 2022 | Learning Dynamic 3D Geometry and Texture for Video Face SwappingabstractAbstract Face swapping is the process of applying a source actor's appearance to a target actor's performance in a video. This is a challenging visual effect that has seen increasing demand in film and television production. Recent work has shown that data‐driven methods based on deep learning can produce compelling effects at production quality in a fraction of the time required for a traditional 3D pipeline. However, the dominant approach operates only on 2D imagery without reference to the underlying facial geometry or texture, resulting in poor generalization under novel viewpoints and little artistic control. Methods that do incorporate geometry rely on pre‐learned facial priors that do not adapt well to particular geometric features of the source and target faces. We approach the problem of face swapping from the perspective of learning simultaneous convolutional facial autoencoders for the source and target identities, using a shared encoder network with identity‐specific decoders. The key novelty in our approach is that each decoder first lifts the latent code into a 3D representation, comprising a dynamic face texture and a deformable 3D face shape, before projecting this 3D face back onto the input image using a differentiable renderer. The coupled autoencoders are trained only on videos of the source and target identities, without requiring 3D supervision. By leveraging the learned 3D geometry and texture, our method achieves face swapping with higher quality than when using off‐the‐shelf monocular 3D face reconstruction, and overall lower FID score than state‐of‐the‐art 2D methods. Furthermore, our 3D representation allows for efficient artistic control over the result, which can be hard to achieve with existing 2D approaches. Christopher Otto, Jacek Naruniec, Leonhard Helminger, Thomas Etterlin, Graziana Mignone, Prashanth Chandran, Gaspard Zoss, Christopher Schroers, Markus Gross 0001, Paulo F. U. Gotardo, Derek Bradley, Romann M. Weber |
Comput. Graph. Forum | 10 |
| 2022 | Facial hair tracking for high fidelity performance captureabstractFacial hair is a largely overlooked topic in facial performance capture. Most production pipelines in the entertainment industry do not have a way to automatically capture facial hair or track the skin underneath it. Thus, actors are asked to shave clean before face capture, which is very often undesirable. Capturing the geometry of individual facial hairs is very challenging, and their presence makes it harder to capture the deforming shape of the underlying skin surface. Some attempts have already been made at automating this task, but only for static faces with relatively sparse 3D hair reconstructions. In particular, current methods lack the temporal correspondence needed when capturing a sequence of video frames depicting facial performance. The problem of robustly tracking the skin underneath also remains unaddressed. In this paper, we propose the first multiview reconstruction pipeline that tracks both the dense 3D facial hair, as well as the underlying 3D skin for entire performances. Our method operates with standard setups for face photogrammetry, without requiring dense camera arrays. For a given capture subject, our algorithm first reconstructs a dense, high-quality neutral 3D facial hairstyle by registering sparser hair reconstructions over multiple frames that depict a neutral face under quasi-rigid motion. This custom-built, reference facial hairstyle is then tracked throughout a variety of changing facial expressions in a captured performance, and the result is used to constrain the tracking of the 3D skin surface underneath. We demonstrate the proposed capture pipeline on a variety of different facial hairstyles and lengths, ranging from sparse and short to dense full-beards. Sebastian Winberg, Gaspard Zoss, Prashanth Chandran, Paulo F. U. Gotardo, Derek Bradley |
ACM Trans. Graph. | 4 |
| 2022 | Production-Ready Face Re-Aging for Visual EffectsabstractPhotorealistic digital re-aging of faces in video is becoming increasingly common in entertainment and advertising. But the predominant 2D painting workflow often requires frame-by-frame manual work that can take days to accomplish, even by skilled artists. Although research on facial image re-aging has attempted to automate and solve this problem, current techniques are of little practical use as they typically suffer from facial identity loss, poor resolution, and unstable results across subsequent video frames. In this paper, we present the first practical, fully-automatic and production-ready method for re-aging faces in video images. Our first key insight is in addressing the problem of collecting longitudinal training data for learning to re-age faces over extended periods of time, a task that is nearly impossible to accomplish for a large number of real people. We show how such a longitudinal dataset can be constructed by leveraging the current state-of-the-art in facial re-aging that, although failing on real images, does provide photoreal re-aging results on synthetic faces. Our second key insight is then to leverage such synthetic data and formulate facial re-aging as a practical image-to-image translation task that can be performed by training a well-understood U-Net architecture, without the need for more complex network designs. We demonstrate how the simple U-Net, surprisingly, allows us to advance the state of the art for re-aging real faces on video, with unprecedented temporal stability and preservation of facial identity across variable expressions, viewpoints, and lighting conditions. Finally, our new face re-aging network (FRAN) incorporates simple and intuitive mechanisms that provides artists with localized control and creative freedom to direct and fine-tune the re-aging effect, a feature that is largely important in real production pipelines and often overlooked in related research work. Gaspard Zoss, Prashanth Chandran, Eftychios Sifakis, Markus Gross 0001, Paulo F. U. Gotardo, Derek Bradley |
ACM Trans. Graph. | 5 |
| 2021 | Adaptive Convolutions for Structure-Aware Style TransferabstractStyle transfer between images is an artistic application of CNNs, where the ‘style’ of one image is transferred onto another image while preserving the latter’s content. The state of the art in neural style transfer is based on Adaptive Instance Normalization (AdaIN), a technique that transfers the statistical properties of style features to a content image, and can transfer a large number of styles in real time. However, AdaIN is a global operation; thus local geometric structures in the style image are often ignored during the transfer. We propose Adaptive Convolutions (AdaConv), a generic extension of AdaIN, to allow for the simultaneous transfer of both statistical and structural styles in real time. Apart from style transfer, our method can also be readily extended to style-based image generation, and other tasks where AdaIN has already been adopted. Prashanth Chandran, Gaspard Zoss, Paulo F. U. Gotardo, Markus Gross 0001, Derek Bradley |
CVPR | 3 |
| 2021 | Single Day Outdoor Photometric StereoabstractPhotometric Stereo (PS) under outdoor illumination remains a challenging, ill-posed problem due to insufficient variability in illumination. Months-long capture sessions are typically used in this setup, with little success on shorter, single-day time intervals. In this paper, we investigate the solution of outdoor PS over a single day, under different weather conditions. First, we investigate the relationship between weather and surface reconstructability in order to understand when natural lighting allows existing PS algorithms to work. Our analysis reveals that partially cloudy days improve the conditioning of the outdoor PS problem while sunny days do not allow the unambiguous recovery of surface normals from photometric cues alone. We demonstrate that calibrated PS algorithms can thus be employed to reconstruct Lambertian surfaces accurately under partially cloudy days. Second, we solve the ambiguity arising in clear days by combining photometric cues with prior knowledge on material properties, local surface geometry and the natural variations in outdoor lighting through a CNN-based, weakly-calibrated PS technique. Given a sequence of outdoor images captured during a single sunny day, our method robustly estimates the scene surface normals with unprecedented quality for the considered scenario. Our approach does not require precise geolocation and significantly outperforms several state-of-the-art methods on images with real lighting, showing that our CNN can combine efficiently learned priors and photometric cues available during a single sunny day. Yannick Hold-Geoffroy, Paulo F. U. Gotardo, Jean-François Lalonde |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2021 | Rendering with style: combining traditional and neural approaches for high-quality face renderingabstractFor several decades, researchers have been advancing techniques for creating and rendering 3D digital faces, where a lot of the effort has gone into geometry and appearance capture, modeling and rendering techniques. This body of research work has largely focused on facial skin, with much less attention devoted to peripheral components like hair, eyes and the interior of the mouth. As a result, even with the best technology for facial capture and rendering, in most high-end productions a lot of artist time is still spent modeling the missing components and fine-tuning the rendering parameters to combine everything into photo-real digital renders. In this work we propose to combine incomplete, high-quality renderings showing only facial skin with recent methods for neural rendering of faces, in order to automatically and seamlessly create photo-realistic full-head portrait renders from captured data without the need for artist intervention. Our method begins with traditional face rendering, where the skin is rendered with the desired appearance, expression, viewpoint, and illumination. These skin renders are then projected into the latent space of a pre-trained neural network that can generate arbitrary photo-real face images (StyleGAN2). The result is a sequence of realistic face images that match the identity and appearance of the 3D character at the skin level, but is completed naturally with synthesized hair, eyes, inner mouth and surroundings. Notably, we present the first method for multi-frame consistent projection into this latent space, allowing photo-realistic rendering and preservation of the identity of the digital human over an animated performance sequence, which can depict different expressions, lighting conditions and viewpoints. Our method can be used in new face rendering pipelines and, importantly, in other deep learning applications that require large amounts of realistic training data with ground-truth 3D geometry, appearance maps, lighting, and viewpoint. Prashanth Chandran, Sebastian Winberg, Gaspard Zoss, Jérémy Riviere, Markus Gross 0001, Paulo F. U. Gotardo, Derek Bradley |
ACM Trans. Graph. | 6 |
| 2020 | Single-shot high-quality facial geometry and skin appearance captureabstractWe propose a new light-weight face capture system capable of reconstructing both high-quality geometry and detailed appearance maps from a single exposure. Unlike currently employed appearance acquisition systems, the proposed technology does not require active illumination and hence can readily be integrated with passive photogrammetry solutions. These solutions are in widespread use for 3D scanning humans as they can be assembled from off-the-shelf hardware components, but lack the capability of estimating appearance. This paper proposes a solution to overcome this limitation, by adding appearance capture to photogrammetry systems. The only additional hardware requirement to these solutions is that a subset of the cameras are cross-polarized with respect to the illumination, and the remaining cameras are parallel-polarized. The proposed algorithm leverages the images with the two different polarization states to reconstruct the geometry and to recover appearance properties. We do so by means of an inverse rendering framework, which solves per texel diffuse albedo, specular intensity, and high-resolution normals, as well as global specular roughness considering the subsurface scattering nature of skin. We show results for a variety of human subjects of different ages and skin typology, illustrating how the captured fine-detail skin surface and subsurface scattering effects lead to realistic renderings of their digital doubles, also in different illumination conditions. Jérémy Riviere, Paulo F. U. Gotardo, Derek Bradley, Abhijeet Ghosh, Thabo Beeler |
ACM Trans. Graph. | 2 |
| 2018 | From Faces to Outdoor Light ProbesabstractAbstract Image‐based lighting has allowed the creation of photo‐realistic computer‐generated content. However, it requires the accurate capture of the illumination conditions, a task neither easy nor intuitive, especially to the average digital photography enthusiast. This paper presents an approach to directly estimate an HDR light probe from a single LDR photograph, shot outdoors with a consumer camera, without specialized calibration targets or equipment. Our insight is to use a person's face as an outdoor light probe. To estimate HDR light probes from LDR faces we use an inverse rendering approach which employs data‐driven priors to guide the estimation of realistic, HDR lighting. We build compact, realistic representations of outdoor lighting both parametrically and in a data‐driven way, by training a deep convolutional autoencoder on a large dataset of HDR sky environment maps. Our approach can recover high‐frequency, extremely high dynamic range lighting environments. For quantitative evaluation of lighting estimation accuracy and relighting accuracy, we also contribute a new database of face photographs with corresponding HDR light probes. We show that relighting objects with HDR light probes estimated by our method yields realistic results in a wide variety of settings. Dan Andrei Calian, Jean-François Lalonde, Paulo F. U. Gotardo, Tomas Simon, Iain A. Matthews, Kenny Mitchell |
Comput. Graph. Forum | 3 |
| 2018 | Practical dynamic facial appearance modeling and acquisitionabstractWe present a method to acquire dynamic properties of facial skin appearance, including dynamic diffuse albedo encoding blood flow, dynamic specular intensity, and per-frame high resolution normal maps for a facial performance sequence. The method reconstructs these maps from a purely passive multi-camera setup, without the need for polarization or requiring temporally multiplexed illumination. Hence, it is very well suited for integration with existing passive systems for facial performance capture. To solve this seemingly underconstrained problem, we demonstrate that albedo dynamics during a facial performance can be modeled as a combination of: (1) a static, high-resolution base albedo map, modeling full skin pigmentation; and (2) a dynamic, one-dimensional component in the CIE L*a*b* color space, which explains changes in hemoglobin concentration due to blood flow. We leverage this albedo subspace and additional constraints on appearance and surface geometry to also estimate specular reflection parameters and resolve high-resolution normal maps with unprecedented detail in a passive capture system. These constraints are built into an inverse rendering framework that minimizes the difference of the rendered face to the captured images, incorporating constraints from multiple views for every texel on the face. The presented method is the first system capable of capturing high-quality dynamic appearance maps at full resolution and video framerates, providing a major step forward in the area of facial appearance acquisition. Paulo F. U. Gotardo, Jérémy Riviere, Derek Bradley, Abhijeet Ghosh, Thabo Beeler |
ACM Trans. Graph. | 1 |
| 2018 | Appearance capture and modeling of human teethabstractRecreating the appearance of humans in virtual environments for the purpose of movie, video game, or other types of production involves the acquisition of a geometric representation of the human body and its scattering parameters which express the interaction between the geometry and light propagated throughout the scene. Teeth appearance is defined not only by the light and surface interaction, but also by its internal geometry and the intra-oral environment, posing its own unique set of challenges. Therefore, we present a system specifically designed for capturing the optical properties of live human teeth such that they can be realistically re-rendered in computer graphics. We acquire our data in vivo in a conventional multiple camera and light source setup and use exact geometry segmented from intra-oral scans. To simulate the complex interaction of light in the oral cavity during inverse rendering we employ a novel pipeline based on derivative path tracing with respect to both optical properties and geometry of the inner dentin surface. The resulting estimates of the global derivatives are used to extract parameters in a joint numerical optimization. The final appearance faithfully recreates the acquired data and can be directly used in conventional path tracing frameworks for rendering virtual humans. Zdravko Velinov, Marios Papas, Derek Bradley, Paulo F. U. Gotardo, Parsa Mirdehghan, Steve Marschner, Jan Novák, Thabo Beeler |
ACM Trans. Graph. | 4 |
| 2015 | x-Hour Outdoor Photometric StereoabstractWhile Photometric Stereo (PS) has long been confined to the lab, there has been a recent interest in applying this technique to reconstruct outdoor objects and scenes. Un-fortunately, the most successful outdoor PS techniques typically require gathering either months of data, or waiting for a particular time of the year. In this paper, we analyze the illumination requirements for single-day outdoor PS to yield stable normal reconstructions, and determine that these requirements are often available in much less than a full day. In particular, we show that the right set of conditions for stable PS solutions may be observed in the sky within short time intervals of just above one hour. This work provides, for the first time, a detailed analysis of the factors affecting the performance of x-hour outdoor photometric stereo. Yannick Hold-Geoffroy, Paulo F. U. Gotardo, Jean-François Lalonde |
3DV | 3 |
| 2015 | What Is a Good Day for Outdoor Photometric Stereo?abstractPhotometric Stereo has been explored extensively in laboratory conditions since its inception. Recently, attempts have been made at applying this technique under natural outdoor lighting. Outdoor photometric stereo presents additional challenges as one does not have control over illumination anymore. In this paper, we explore the stability of surface normals reconstructed outdoors. We present a data-driven analysis based on a large database of outdoor HDR environment maps. Given a sequence of object images and corresponding sky maps captured in a single day, we investigate natural factors that impact the uncertainty in the estimated surface normals. Quantitative evidence reveals strong dependencies between expected accuracy and the normal orientation, cloud coverage, and sun elevation. In particular, we show that partially cloudy days yield greater accuracy than sunny days with clear skies; furthermore, high sun elevation--recommended in previous work--is in fact not necessarily optimal when taking more elaborate illumination models into account. Yannick Hold-Geoffroy, Paulo F. U. Gotardo, Jean-François Lalonde |
ICCP | 3 |
| 2015 | Photogeometric Scene Flow for High-Detail Dynamic 3D ReconstructionabstractPhotometric stereo (PS) is an established technique for high-detail reconstruction of 3D geometry and appearance. To correct for surface integration errors, PS is often combined with multiview stereo (MVS). With dynamic objects, PS reconstruction also faces the problem of computing optical flow (OF) for image alignment under rapid changes in illumination. Current PS methods typically compute optical flow and MVS as independent stages, each one with its own limitations and errors introduced by early regularization. In contrast, scene flow methods estimate geometry and motion, but lack the fine detail from PS. This paper proposes photogeometric scene flow (PGSF) for high-quality dynamic 3D reconstruction. PGSF performs PS, OF, and MVS simultaneously. It is based on two key observations: (i) while image alignment improves PS, PS allows for surfaces to be relit to improve alignment, (ii) PS provides surface gradients that render the smoothness term in MVS unnecessary, leading to truly data-driven, continuous depth estimates. This synergy is demonstrated in the quality of the resulting RGB appearance, 3D geometry, and 3D motion. Paulo F. U. Gotardo, Tomas Simon, Yaser Sheikh, Iain A. Matthews |
ICCV | 1 |
| 2014 | Salient and non-salient fiducial detection using a probabilistic graphical model
Carlos F. Benitez-Quiroz, Samuel Rivera, Paulo F. U. Gotardo, Aleix Martinez |
Pattern Recognit. | 3 |
| 2012 | Learning Spatially-Smooth Mappings in Non-Rigid Structure From Motion
Onur C. Hamsici, Paulo F. U. Gotardo, Aleix Martinez |
ECCV (4) | 2 |
| 2011 | Non-rigid structure from motion with complementary rank-3 spacesabstractNon-rigid structure from motion (NR-SFM) is a difficult, underconstrained problem in computer vision. This paper proposes a new algorithm that revises the standard matrix factorization approach in NR-SFM. We consider two alternative representations for the linear space spanned by a small number K of 3D basis shapes. As compared to the standard approach using general rank-3K matrix factors, we show that improved results are obtained by explicitly modeling K complementary spaces of rank-3. Our new method is positively compared to the state-of-the-art in NR-SFM, providing improved results on high-frequency deformations of both articulated and simpler deformable shapes. We also present an approach for NR-SFM with occlusion. Paulo F. U. Gotardo, Aleix Martinez |
CVPR | 1 |
| 2011 | Kernel non-rigid structure from motionabstractNon-rigid structure from motion (NRSFM) is a difficult, underconstrained problem in computer vision. The standard approach in NRSFM constrains 3D shape deformation using a linear combination of K basis shapes; the solution is then obtained as the low-rank factorization of an input observation matrix. An important but overlooked problem with this approach is that non-linear deformations are often observed; these deformations lead to a weakened low-rank constraint due to the need to use additional basis shapes to linearly model points that move along curves. Here, we demonstrate how the kernel trick can be applied in standard NRSFM. As a result, we model complex, deformable 3D shapes as the outputs of a non-linear mapping whose inputs are points within a low-dimensional shape space. This approach is flexible and can use different kernels to build different non-linear models. Using the kernel trick, our model complements the low-rank constraint by capturing non-linear relationships in the shape coefficients of the linear model. The net effect can be seen as using non-linear dimensionality reduction to further compress the (shape) space of possible solutions. Paulo F. U. Gotardo, Aleix Martinez |
ICCV | 1 |
| 2011 | Computing Smooth Time Trajectories for Camera and Deformable Shape in Structure from Motion with OcclusionabstractWe address the classical computer vision problems of rigid and nonrigid structure from motion (SFM) with occlusion. We assume that the columns of the input observation matrix W describe smooth 2D point trajectories over time. We then derive a family of efficient methods that estimate the column space of W using compact parameterizations in the Discrete Cosine Transform (DCT) domain. Our methods tolerate high percentages of missing data and incorporate new models for the smooth time trajectories of 2D-points, affine and weak-perspective cameras, and 3D deformable shape. We solve a rigid SFM problem by estimating the smooth time trajectory of a single camera moving around the structure of interest. By considering a weak-perspective camera model from the outset, we directly compute euclidean 3D shape reconstructions without requiring postprocessing steps such as euclidean upgrade and bundle adjustment. Our results on real SFM data sets with high percentages of missing data compared positively to those in the literature. In nonrigid SFM, we propose a novel 3D shape trajectory approach that solves for the deformable structure as the smooth time trajectory of a single point in a linear shape space. A key result shows that, compared to state-of-the-art algorithms, our nonrigid SFM method can better model complex articulated deformation with higher frequency DCT components while still maintaining the low-rank factorization constraint. Finally, we also offer an approach for nonrigid SFM when W is presented with missing data. Paulo F. U. Gotardo, Aleix Martinez |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2006 | A New Deformable Model for Boundary Tracking in Cardiac MRI and Its Application to the Detection of Intra-Ventricular DyssynchronyabstractWe present a new deformable model technique following a snake-like approach and using a complex Fourier shape descriptors parameterization to efficiently formulate the forces that constrain contour deformation. The method was successfully applied to track the left ventricle’s (LV) endocardial and epicardial boundaries in sequences of shortaxis magnetic resonance images depicting complete cardiac cycles. The extracted shapes show the method’s robustness to weak contrast, noisy edge maps and to papillary muscle anatomy. Our second contribution is a statistical pattern recognition approach for the detection of asynchronous activation of the LV walls. We applied our deformable model method to provide spatio-temporal characterizations of complete cardiac cycles and then designed a linear classifier using the popular combination of Principal Component Analysis and Linear Discriminant Analysis. From a database comprising 33 patients, our approach provided a correct classification performance of 90.9% showing its potential in providing improved dyssynchrony characterization as an adjunct to current criteria for selecting patients for therapy, which provided a classification accuracy of just 62.5% on the same database. Paulo F. U. Gotardo, Kim L. Boyer, Joel H. Saltz, Subha V. Raman |
CVPR (1) | 1 |
| 2004 | Range image segmentation into planar and quadric surfaces using an improved robust estimator and genetic algorithmabstractThis paper presents a novel range image segmentation method employing an improved robust estimator to iteratively detect and extract distinct planar and quadric surfaces. Our robust estimator extends M-estimator Sample Consensus/Random Sample Consensus (MSAC/RANSAC) to use local surface orientation information, enhancing the accuracy of inlier/outlier classification when processing noisy range data describing multiple structures. An efficient approximation to the true geometric distance between a point and a quadric surface also contributes to effectively reject weak surface hypotheses and avoid the extraction of false surface components. Additionally, a genetic algorithm was specifically designed to accelerate the optimization process of surface extraction, while avoiding premature convergence. We present thorough experimental results with quantitative evaluation against ground truth. The segmentation algorithm was applied to three real range image databases and competes favorably against eleven other segmenters using the most popular evaluation framework in the literature. Our approach lends itself naturally to parallel implementation and application in real-time tasks. The method fits well, into several of today's applications in man-made environments, such as target detection and autonomous navigation, for which obstacle detection, but not description or reconstruction, is required. It can also be extended to process point clouds resulting from range image registration. Paulo F. U. Gotardo, Olga R. P. Bellon, Kim L. Boyer, Luciano Silva |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 2003 | Range Image Segmentation by Surface Extraction Using an Improved Robust EstimatorabstractThe paper presents a novel range image segmentation algorithm based on planar surface extraction. The algorithm was applied to common range image databases and was favorably compared against seven other segmentation algorithms using a popular evaluation framework. The experimental results show that, as compared to the other methods, our algorithm presents a good performance in preserving small regions and edge locations when processing noisy images. Our main contribution is an improved robust estimator, derived from the RANSAC and MSAC estimators, whose optimization process is accelerated by a genetic algorithm with a new set of parameters and operations designed to avoid premature convergence. Paulo F. U. Gotardo, Olga R. P. Bellon, Luciano Silva |
CVPR (2) | 1 |
| 2003 | Range image registration using enhanced genetic algorithmsabstractMost range image registration techniques are based on variants of the ICP (iterative closest point) algorithm. The ICP algorithm has two main drawbacks, the possibility of convergence to a local minimum and the need to prealign the images. Genetic algorithms (GAs) are known to be robust in relation to search and optimization problems and were recently applied to range image registration, providing good convergence results without the constraints observed in the ICP approaches. To improve range image registration by GAs, we explored 3 novel approaches: a hybrid algorithm that combines a GA with hillclimbing heuristics (GH), a parallel migration GA (MGA), and a MGA using hillclimbing (MGH). We also define a new robust evaluation measure, called the surface interpenetration, to compare the obtained registration results. Up to now, interpenetration has been evaluated only qualitatively; we define the first quantitative measure for it. The experimental results show that our methods yield more accurate registration results than either ICP or standard GA approaches. Luciano Silva, Olga R. P. Bellon, Paulo F. U. Gotardo, Kim L. Boyer |
ICIP (2) | 3 |
| 2002 | A global-to-local approach for robust range image segmentationabstractWe present a range image segmentation algorithm based on a robust estimation technique, the M-estimator sample consensus (MSAC). The algorithm is a parallelizable "global-to-local" approach for the extraction of planar surfaces directly from range images. Solutions to some problems faced when extracting planar surfaces globally are also proposed. Experimental results show the algorithm is robust to image noise in the sense that it is able to preserve object shapes so that neither presmoothing, nor postprocessing steps are required. It also does not rely on MSAC and can be easily adapted to use other robust estimators. Thus, it may be used as a framework to compare robust estimators. Luciano Silva, Olga R. P. Bellon, Paulo F. U. Gotardo |
ICIP (1) | 3 |
| 2001 | Edge-based image segmentation using curvature sign maps from reflectance and range imagesabstractA new approach to image segmentation by edge detection is proposed for preserving objects topology and shape while retrieving precisely located, one-pixel-wide edges. The method is based on mean (H) and Gaussian (K) surface curvatures sign maps (HK-sign maps) computed from both registered reflectance and range images, provided by a single sensor. HK-sign maps have been used to identify objects regions on range and intensity images, but not edges, as presented in this work. The combination of the computed range and reflectance edge maps has led to more accurate segmentation results than just by using either of them alone. The proposed algorithm has been tested on real images and compared to four traditional range image segmentation algorithms. Experimental results demonstrate the viability and usefulness of our approach. Luciano Silva, Olga R. P. Bellon, Paulo F. U. Gotardo |
ICIP (1) | 3 |