VLDB 2026 Research / reviewers in the wild / expert
Gabriel Schwartz
dblp:68/11348
· DBLP profile ↗
12ranked-venue papers
3as first author
7since 2021 · last 2024
0000-0002-8781-5573ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Relightable Gaussian Codec AvatarsabstractThe fidelity of relighting is bounded by both geometry and appearance representations. For geometry, both mesh and volumetric approaches have difficulty modeling intri-cate structures like 3D hair geometry. For appearance, existing relighting models are limited in fidelity and often too slow to render in real-time with high-resolution contin-uous environments. In this work, we present Relightable Gaussian Codec Avatars, a method to build high-fidelity relightable head avatars that can be animated to generate novel expressions. Our geometry model based on 3D Gaus-sians can capture 3D-consistent sub-millimeter details such as hair strands and pores on dynamic face sequences. To support diverse materials of human heads such as the eyes, skin, and hair in a unified manner, we present a novel re-lightable appearance model based on learnable radiance transfer. Together with global illumination-aware spheri-cal harmonics for the diffuse components, we achieve real-time relighting with all-frequency reflections using spheri-cal Gaussians. This appearance model can be efficiently relit under both point light and continuous illumination. We further improve the fidelity of eye reflections and enable ex-plicit gaze control by introducing relightable explicit eye models. Our method outperforms existing approaches with-out compromising real-time performance. We also demon-strate real-time relighting of avatars on a tethered con-sumer VR headset, showcasing the efficiency and fidelity of our avatars. Shunsuke Saito, Gabriel Schwartz, Tomas Simon, Giljoo Nam |
CVPR | 2 |
| 2024 | Rasterized Edge Gradients: Handling Discontinuities Differentiably
Stanislav Pidhorskyi, Tomas Simon, Gabriel Schwartz, He Wen 0001, Yaser Sheikh, Jason M. Saragih |
ECCV (83) | 3 |
| 2024 | Codec Avatar Studio: Paired Human Captures for Complete, Driveable, and Generalizable AvatarsabstractTo build photorealistic avatars that users can embody, human modelling must be complete (cover the full body), driveable (able to reproduce the current motion and appearance from the user), and generalizable (i.e., easily adaptable to novel identities).Towards these goals, paired captures, that is, captures of the same subject obtained from systems of diverse quality and availability, are crucial.However, paired captures are rarely available to researchers outside of dedicated industrial labs: Codec Avatar Studio is our proposal to close this gap.Towards generalization and driveability, we introduce a dataset of 256 subjects captured in two modalities: high resolution multi-view scans of their heads, and video from the internal cameras of a headset.Towards completeness, we introduce a dataset of 4 subjects captured in eight modalities: high quality relightable multi-view captures of heads and hands, full body multi-view captures with minimal and regular clothes, and corresponding head, hands and body phone captures.Together with our data, we also provide code and pre-trained models for different state-of-the-art human generation models.Our datasets and code are available at https://github.com/facebookresearch/ava-256 and https://github.com/facebookresearch/goliath. Julieta Martinez 0001, Emily Kim, Javier Romero 0002, Timur M. Bagautdinov, Shunsuke Saito, Shoou-I Yu, Michael Zollhöfer, Te-Li Wang, Shaojie Bai, Chenghui Li, Shih-En Wei, Rohan Joshi, Wyatt Borsos, Tomas Simon, Jason M. Saragih, Paul Theodosis, Alexander Greene, Anjani Josyula, Silvio Maeta, Andrew Jewett, Simion Venshtain, Christopher Heilman, Yueh-Tung Chen, Sidi Fu, Mohamed Elshaer, Tingfang Du, Longhua Wu, Shen-Chi Chen, Youssef Emad, Steven Longay, Ashley Brewer, Hitesh Shah, Taylor Koska, Kayla Haidle, Matthew Andromalos, Joanna Hsu, Thomas Dauer, Peter Selednik, Timothy Godisart, Scott Ardisson, Matthew Cipperly, Ben Humberston, Lon Farr, Bob Hansen, Peihong Guo, Dave Braun, Steven Krenn, He Wen 0001, Lucas Evans, Natalia Fadeeva, Matthew Stewart, Gabriel Schwartz, Divam Gupta, Gyeongsik Moon, Takaaki Shiratori, Fabian Prada, Bernardo Pires, Julia Buffalini, Autumn Trimble, Kevyn McPhail, Melissa Schoeller, Yaser Sheikh |
NeurIPS | 56 |
| 2024 | URAvatar: Universal Relightable Gaussian Codec Avatars
Chen Cao 0001, Gabriel Schwartz, Rawal Khirodkar, Christian Richardt, Tomas Simon, Yaser Sheikh, Shunsuke Saito |
SIGGRAPH Asia | 3 |
| 2024 | Universal Facial Encoding of Codec Avatars from VR HeadsetsabstractFaithful real-time facial animation is essential for avatar-mediated telepresence in Virtual Reality (VR). To emulate authentic communication, avatar animation needs to be efficient and accurate: able to capture both extreme and subtle expressions within a few milliseconds to sustain the rhythm of natural conversations. The oblique and incomplete views of the face, variability in the donning of headsets, and illumination variation due to the environment are some of the unique challenges in generalization to unseen faces. In this paper, we present a method that can animate a photorealistic avatar in realtime from head-mounted cameras (HMCs) on a consumer VR headset. We present a self-supervised learning approach, based on a cross-view reconstruction objective, that enables generalization to unseen users. We present a lightweight expression calibration mechanism that increases accuracy with minimal additional cost to run-time efficiency. We present an improved parameterization for precise ground-truth generation that provides robustness to environmental variation. The resulting system produces accurate facial animation for unseen users wearing VR headsets in realtime. We compare our approach to prior face-encoding methods demonstrating significant improvements in both quantitative metrics and qualitative results. Shaojie Bai, Te-Li Wang, Chenghui Li, Akshay Venkatesh, Tomas Simon, Chen Cao 0001, Gabriel Schwartz, Jason M. Saragih, Yaser Sheikh, Shih-En Wei |
ACM Trans. Graph. | 7 |
| 2022 | Depth of Field Aware Differentiable RenderingabstractCameras with a finite aperture diameter exhibit defocus for scene elements that are not at the focus distance, and have only a limited depth of field within which objects appear acceptably sharp. In this work we address the problem of applying inverse rendering techniques to input data that exhibits such defocus blurring. We present differentiable depth-of-field rendering techniques that are applicable to both rasterization-based methods using mesh representations, as well as ray-marching-based methods using either explicit [Yu et al. 2021] or implicit volumetric radiance fields [Mildenhall et al. 2020]. Our approach learns significantly sharper scene reconstructions on data containing blur due to depth of field, and recovers aperture and focus distance parameters that result in plausible forward-rendered images. We show applications to macro photography, where typical lens configurations result in a very narrow depth of field, and to multi-camera video capture, where maintaining sharp focus across a large capture volume for a moving subject is difficult. Stanislav Pidhorskyi, Timur M. Bagautdinov, Shugao Ma, Jason M. Saragih, Gabriel Schwartz, Yaser Sheikh, Tomas Simon |
ACM Trans. Graph. | 5 |
| 2021 | Mixture of volumetric primitives for efficient neural renderingabstractReal-time rendering and animation of humans is a core function in games, movies, and telepresence applications. Existing methods have a number of drawbacks we aim to address with our work. Triangle meshes have difficulty modeling thin structures like hair, volumetric representations like Neural Volumes are too low-resolution given a reasonable memory budget, and high-resolution implicit representations like Neural Radiance Fields are too slow for use in real-time applications. We present Mixture of Volumetric Primitives (MVP), a representation for rendering dynamic 3D content that combines the completeness of volumetric representations with the efficiency of primitive-based rendering, e.g., point-based or mesh-based methods. Our approach achieves this by leveraging spatially shared computation with a convolutional architecture and by minimizing computation in empty regions of space with volumetric primitives that can move to cover only occupied regions. Our parameterization supports the integration of correspondence and tracking constraints, while being robust to areas where classical tracking fails, such as around thin or translucent structures and areas with large topological variability. MVP is a hybrid that generalizes both volumetric and primitive-based representations. Through a series of extensive experiments we demonstrate that it inherits the strengths of each, while avoiding many of their limitations. We also compare our approach to several state-of-the-art methods and demonstrate that MVP produces superior results in terms of quality and runtime performance. Stephen Lombardi, Tomas Simon, Gabriel Schwartz, Michael Zollhöfer, Yaser Sheikh, Jason M. Saragih |
ACM Trans. Graph. | 3 |
| 2020 | Recognizing Material Properties from ImagesabstractHumans implicitly rely on the properties of materials to guide our interactions. Grasping smooth materials, for example, requires more care than rough ones. We may even visually infer non-visual properties (e.g., softness is a physical material property). We refer to visually-recognizable material properties as visual material attributes. Recognizing these attributes in images can provide valuable information for scene understanding and material recognition. Unlike typical object and scene attributes, however, visual material attributes are local (i.e., "fuzziness" does not have a shape). Given full supervision, we may accurately recognize such attributes from purely local information (small image patches). Obtaining consistent full supervision at scale, however, is challenging. To solve this problem, we probe the human visual perception of materials. By asking simple yes/no questions comparing pairs of image patches, we obtain the weak supervision required to build a set of classifiers for attributes that, while unnamed, function similarly to the attributes with which we describe materials. Furthermore, we integrate this method in the end-to-end learning of a CNN that simultaneously recognizes materials and their visual attributes. Experiments show that visual material attributes serve as both a useful representation for known material categories and as a basis for transfer learning. Gabriel Schwartz, Ko Nishino |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2020 | The eyes have it: an integrated eye and face model for photorealistic facial animationabstractInteracting with people across large distances is important for remote work, interpersonal relationships, and entertainment. While such face-to-face interactions can be achieved using 2D video conferencing or, more recently, virtual reality (VR), telepresence systems currently distort the communication of eye contact and social gaze signals. Although methods have been proposed to redirect gaze in 2D teleconferencing situations to enable eye contact, 2D video conferencing lacks the 3D immersion of real life. To address these problems, we develop a system for face-to-face interaction in VR that focuses on reproducing photorealistic gaze and eye contact. To do this, we create a 3D virtual avatar model that can be animated by cameras mounted on a VR headset to accurately track and reproduce human gaze in VR. Our primary contributions in this work are a jointly-learnable 3D face and eyeball model that better represents gaze direction and upper facial expressions, a method for disentangling the gaze of the left and right eyes from each other and the rest of the face allowing the model to represent entirely unseen combinations of gaze and expression, and a gaze-aware model for precise animation from headset-mounted cameras. Our quantitative experiments show that our method results in higher reconstruction quality, and qualitative results show our method gives a greatly improved sense of presence for VR avatars. Gabriel Schwartz, Shih-En Wei, Te-Li Wang, Stephen Lombardi, Tomas Simon, Jason M. Saragih, Yaser Sheikh |
ACM Trans. Graph. | 1 |
| 2019 | Neural volumes: learning dynamic renderable volumes from imagesabstractModeling and rendering of dynamic scenes is challenging, as natural scenes often contain complex phenomena such as thin structures, evolving topology, translucency, scattering, occlusion, and biological motion. Mesh-based reconstruction and tracking often fail in these cases, and other approaches (e.g., light field video) typically rely on constrained viewing conditions, which limit interactivity. We circumvent these difficulties by presenting a learning-based approach to representing dynamic objects inspired by the integral projection model used in tomographic imaging. The approach is supervised directly from 2D images in a multi-view capture setting and does not require explicit reconstruction or tracking of the object. Our method has two primary components: an encoder-decoder network that transforms input images into a 3D volume representation, and a differentiable ray-marching operation that enables end-to-end training. By virtue of its 3D representation, our construction extrapolates better to novel viewpoints compared to screen-space rendering techniques. The encoder-decoder architecture learns a latent representation of a dynamic scene that enables us to produce novel content sequences not seen during training. To overcome memory limitations of voxel-based representations, we learn a dynamic irregular grid structure implemented with a warp field during ray-marching. This structure greatly improves the apparent resolution and reduces grid-like artifacts and jagged motion. Finally, we demonstrate how to incorporate surface-based representations into our volumetric-learning framework for applications where the highest resolution is required, using facial performance capture as a case in point. Stephen Lombardi, Tomas Simon, Jason M. Saragih, Gabriel Schwartz, Andreas M. Lehrmann, Yaser Sheikh |
ACM Trans. Graph. | 4 |
| 2015 | Automatically discovering local visual material attributesabstractShape cues play an important role in computer vision, but shape is not the only information available in images. Materials, such as fabric and plastic, are discernible in images even when shapes, such as those of an object, are not. We argue that it would be ideal to recognize materials without relying on object cues such as shape. This would allow us to use materials as a context for other vision tasks, such as object recognition. Humans are intuitively able to find visual cues that describe materials. Previous frameworks attempt to recognize these cues (as visual material traits) using fully-supervised learning. This requirement is not feasible when multiple annotators and large quantities of images are involved. In this paper, we derive a framework that allows us to discover locally-recognizable material attributes from crowdsourced perceptual material distances. We show that the attributes we discover do in fact separate material categories. Our learned attributes exhibit the same desirable properties as material traits, despite the fact that they are discovered using only partial supervision. Gabriel Schwartz, Ko Nishino |
CVPR | 1 |
| 2012 | 3D Geometric Scale Variability in Range Images: Features and Descriptors
Prabin Bariya, John Novatnack, Gabriel Schwartz, Ko Nishino |
Int. J. Comput. Vis. | 3 |