VLDB 2026 Research / reviewers in the wild / expert
Bernhard Egger 0001
dblp:40/2471-1
· DBLP profile ↗
38ranked-venue papers
5as first author
27since 2021 · last 2026
0000-0002-4736-2397ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 27 · 4 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 21 · 3 first-author · 15 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 6 since 2021Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SurfFill: Completion of LiDAR point clouds via Gaussian surfel splattingabstractLiDAR-captured point clouds are often considered the gold standard in active 3D reconstruction. While their accuracy is exceptional in flat regions, the capturing is susceptible to missing small geometric structures, thin edges, and structures exhibiting challenging surface properties. Alternatively, capturing multiple photos of the scene and applying 3D photogrammetry can infer these details as they often represent feature-rich regions. However, the accuracy of LiDAR for featureless regions is rarely reached. Therefore, we suggest combining the strengths of LiDAR and camera-based capture by introducing SurfFill: a Gaussian surfel-based LiDAR completion scheme. We analyze LiDAR capturings and attribute LiDAR beam divergence as a main factor for artifacts, manifesting mostly at thin structures and edges. We use this insight to introduce an ambiguity heuristic for completed scans by evaluating the change in density in the point cloud. This allows us to identify points close to missed areas, which we can then use to grow additional points from to complete the scan. For this point growing, we employ Gaussian surfels and focus optimization and densification on these ambiguous areas. Finally, Gaussian primitives of the reconstruction in ambiguous areas are extracted and sampled for points to complete the point cloud. To address the challenges of large-scale reconstruction, we extend this pipeline with a divide-and-conquer scheme for building-sized point cloud completion. We evaluate on the task of LiDAR point cloud completion of synthetic and real-world scenes and find that our method outperforms previous reconstruction methods. Svenja Strobel, Matthias Innmann, Bernhard Egger 0001, Marc Stamminger, Linus Franke |
Comput. Graph. | 3 |
| 2026 | Multi-Spectral Gaussian Splatting with Neural Color RepresentationabstractAbstract 3D Gaussian Splatting (3DGS) [KKLD23] has transformed novel‐view synthesis from RGB images, yet remains restricted to the visible spectrum. Many applications, including agricultural monitoring, rely on multi‐spectral imaging, where spectral camera alignment and scalability pose major challenges. We present MS‐Splatting—a multi‐spectral 3DGS framework enabling unified multi‐view consistent reconstruction and rendering across both visible and invisible spectra. Our key component is a neural color representation that encodes per‐primitive features shared across spectral bands, decoded through a shallow multi‐layer perceptron into spectrum‐specific radiance. By leveraging inter‐band correlations, this formulation enhances detail while reducing memory consumption compared to independent band modeling via per‐channel modeling with spherical harmonics. Our method enables accurate parallax‐free novel‐view vegetation index rendering for plant monitoring and enhances RGB novel view synthesis quality by exploiting details revealed through multi‐spectral bands. Our evaluation demonstrates that MS‐Splatting exceeds the current leading methods in both categories. In addition, we introduce a multi‐spectral dataset from aerial captures covering outdoor environments, specifically designed for evaluating these applications. We will release our code and dataset to facilitate further research. The project page is located at: https://meyerls.github.io/ms_splatting Lukas Meyer, Josef Grün, Maximilian Weiherer, Bernhard Egger 0001, Marc Stamminger, Linus Franke |
Comput. Graph. Forum | 4 |
| 2026 | RENI++: A Rotation-Equivariant, Scale-Invariant, Natural Illumination PriorabstractInverse rendering is an ill-posed problem. Previous work has sought to resolve this by focussing on priors for object or scene shape or appearance. In this work, we instead focus on a prior for natural illuminations. Current methods rely on spherical harmonic lighting or other generic representations and, at best, a simplistic prior on the parameters. This results in limitations for the inverse setting in terms of the expressivity of the illumination conditions, especially when taking specular reflections into account. We propose a conditional neural field representation based on a variational auto-decoder and a transformer decoder. We extend Vector Neurons to build equivariance directly into our architecture, and leveraging insights from depth estimation through a scale-invariant loss function, we enable the accurate representation of High Dynamic Range (HDR) images. The result is a compact, rotation-equivariant HDR neural illumination model capable of capturing complex, high-frequency features in natural environment maps. Training our model on a curated dataset of 1.6 K HDR environment maps of natural scenes, we compare it against traditional representations, demonstrate its applicability for an inverse rendering task and show environment map completion from partial observations. James A. D. Gardner, Bernhard Egger 0001, William A. P. Smith |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | Gen3DSR: Generalizable 3D Scene Reconstruction Via Divide and Conquer From a Single ViewabstractSingle-view 3D reconstruction is currently approached from two dominant perspectives: reconstruction of scenes with limited diversity using 3D data supervision or reconstruction of diverse singular objects using large image priors. However, real-world scenarios are far more complex and exceed the capabilities of these methods. We therefore propose a hybrid method following a divide-and-conquer strategy. We first process the scene holistically, extracting depth and semantic information, and then leverage an object-level method for the detailed reconstruction of individual components. By splitting the problem into simpler tasks, our system is able to generalize to various types of scenes without retraining or fine-tuning. We purposely design our pipeline to be highly modular with independent, self-contained modules, to avoid the need for end-to-end training of the whole system. This enables the pipeline to naturally improve as future methods can replace the individual modules. We demonstrate the reconstruction performance of our approach on both synthetic and real-world scenes, comparing favorable against prior works. Project page: https://andreeadogaru.github.io/Gen3DSR Andreea Ardelean, Mert Özer, Bernhard Egger 0001 |
3DV | 3 |
| 2025 | Landmark-Constrained Multi-Object Model Fitting. An Application for 3D Reconstruction of X-Ray Images of the Human FootabstractIn developing countries, two-dimensional (2D) Xrays are widely used in radiology due to their cost-effectiveness, accessibility, and lower radiation exposure compared to three-dimensional (3D) computed tomography (CT) scans. Despite the challenges in interpreting 3D anatomy from 2D images, these methods remain crucial in resource-limited settings. To improve diagnostic capabilities, techniques have been developed to reconstruct 3D images from 2D X-rays, utilising data-efficient and explainable model-to-modality approaches. This study presents simulation of landmark-constrained model fitting steps for 3D reconstruction of complex anatomy using a multi-object statistical shape, pose, and intensity model (SSPIM) and bi-planar digitally reconstructed radiographs (DRRs) forming the model-to-modality approach for clinical adoption. The resulting mean error from the 3D coordinate point reconstruction was 0.28 millimetres (mm) for calibration frame control points, 2.82 and 4.18 mm for the reconstruction of anatomical landmarks of the SSPIM sample and foot CT image, respectively. The average distances obtained from fitting the multi-object SSPIM to the SSPIM sample and the foot CT image anatomical landmarks were 0.32 and 1.94 mm, and 1.94 and 2.54 mm for the 12 ground-truth anatomical landmarks and the 12 reprojected landmarks, respectively. Although the results of this study indicate that there is still a need to improve the 3D-2D algorithm, it also confirms the capability and feasibility of reconstructing complex anatomy from X-rays using 2D landmarks and statistical models. The use of this pipeline potentially offers a cost-effective and clinically useful solution to improve diagnostic accuracy with available low-cost imaging, especially in underserved regions. Catherine Namayega, Tinashe E. M. Mutsvangwa, Bernhard Egger 0001, Bhushan Borotikar, Lindie du Plessis |
BIBE | 3 |
| 2025 | Matérn Kernels for Tunable Implicit Surface ReconstructionabstractWe propose to use the family of Matérn kernels for implicit surface reconstruction, building upon the recent success of kernel methods for 3D reconstruction of oriented point clouds. As we show from a theoretical and practical perspective, Matérn kernels have some appealing properties which make them particularly well suited for surface reconstruction---outperforming state-of-the-art methods based on the arc-cosine kernel while being significantly easier to implement, faster to compute, and scalable. Being stationary, we demonstrate that Matérn kernels allow for tunable surface reconstruction in the same way as Fourier feature mappings help coordinate-based MLPs overcome spectral bias. Moreover, we theoretically analyze Matérn kernels' connection to SIREN networks as well as their relation to previously employed arc-cosine kernels. Finally, based on recently introduced Neural Kernel Fields, we present data-dependent Matérn kernels and conclude that especially the Laplace kernel (being part of the Matérn family) is extremely competitive, performing almost on par with state-of-the-art methods in the noise-free case while having a more than five times shorter training time. Maximilian Weiherer, Bernhard Egger 0001 |
ICLR | 2 |
| 2024 | Improving Throughput-oriented LLM Inference with CPU ComputationsabstractLarge language models (LLMs) have recently captured the attention of a broad audience. To a large part, their exceptional performance in text generation was made possible by an exponential growth of the model parameters. This growth, however, comes at the expense of significantly higher operational costs and a decreased processing speed. Recent research has focused on running LLMs on commodity hardware, for example, by employing the memory hierarchy to augment throughput by increasing the number of batches. These studies, however, tend to overlook or inefficiently utilize the additional computational resources provided by the CPU. In this paper, we introduce a technique capable of efficiently harnessing all available computational resources through a finely tuned and dynamic workload allocation approach. This technique applies to decoder-based models on standard general-purpose hardware, effectively minimizing idle periods for both the CPU and the GPU. We conducted experiments involving various large language models, each representing distinct decoder-based architectures. Compared to the state-of-the-art, the results demonstrate a potential for an increase of up to 105% in throughput with the OPT-30B model. Daon Park, Bernhard Egger 0001 |
PACT | 2 |
| 2024 | Processing of scene intrinsics in the ventral visual stream for object recognition
Shreya Kapoor, Bernhard Egger 0001 |
CogSci | 2 |
| 2024 | TI2V-Zero: Zero-Shot Image Conditioning for Text-to-Video Diffusion ModelsabstractText-conditioned image-to-video generation (TI2V) aims to synthesize a realistic video starting from a given image (e.g., a woman's photo) and a text description (e.g., “a woman is drinking water.”). Existing TI2V frameworks often require costly training on video-text datasets and spe-cific model designs for text and image conditioning. In this paper, we propose TI2V-Zero, a zero-shot, tuning-free method that empowers a pretrained text-to-video (T2V) diffusion model to be conditioned on a provided image, enabling TI2V generation without any optimization, fine-tuning, or introducing external modules. Our approach leverages a pretrained T2V diffusion foundation model as the generative prior. To guide video generation with the additional image input, we propose a “repeat-and-slide” strategy that modulates the reverse denoising process, al-lowing the frozen diffusion model to synthesize a video frame-by-frame starting from the provided image. To ensure temporal continuity, we employ a DDPM inversion strategy to initialize Gaussian noise for each newly synthesized frame and a resampling technique to help preserve visual details. We conduct comprehensive experiments on both domain-specific and open-domain datasets, where TI2V-Zero consistently outperforms a recent open-domain TI2V model. Furthermore, we show that TI2V-Zero can seam-lessly extend to other tasks such as video infilling and pre-diction when provided with more images. Its autoregressive design also supports long video generation. Haomiao Ni, Bernhard Egger 0001, Suhas Lohit, Anoop Cherian, Ye Wang 0001, Toshiaki Koike-Akino, Sharon X. Huang, Tim K. Marks |
CVPR | 2 |
| 2024 | RANRAC: Robust Neural Scene Representations via Random Ray Consensus
Benno Buschmann, Andreea Dogaru, Elmar Eisemann, Michael Weinmann, Bernhard Egger 0001 |
ECCV (76) | 5 |
| 2024 | The Sky's the Limit: Relightable Outdoor Scenes via a Sky-Pixel Constrained Illumination Prior and Outside-In Visibility
James A. D. Gardner, Evgenii Kashin, Bernhard Egger 0001, William A. P. Smith |
ECCV (54) | 3 |
| 2024 | GBOT: Graph-Based 3D Object Tracking for Augmented Reality-Assisted Assembly GuidanceabstractGuidance for assemblable parts is a promising field for augmented reality. Augmented reality assembly guidance requires 6D object poses of target objects in real time. Especially in time-critical medical or industrial settings, continuous and markerless tracking of individual parts is essential to visualize instructions superimposed on or next to the target object parts. In this regard, occlusions by the user’s hand or other objects and the complexity of different assembly states complicate robust and real-time markerless multi-object tracking. To address this problem, we present Graph-based Object Tracking (GBOT), a novel graph-based single-view RGB-D tracking approach. The real-time markerless multi-object tracking is initialized via 6D pose estimation and updates the graph-based assembly poses. The tracking through various assembly states is achieved by our novel multi-state assembly graph. We update the multi-state assembly graph by utilizing the relative poses of the individual assembly parts. Linking the individual objects in this graph enables more robust object tracking during the assembly process. For evaluation, we introduce a synthetic dataset of publicly available and 3D printable assembly assets as a benchmark for future work. Quantitative experiments in synthetic data and further qualitative study in real test data show that GBOT can outperform existing work towards enabling context-aware augmented reality assembly guidance. Dataset and code will be made publically available.****https://github.com/roth-hex-lab/gbot Shiyu Li 0003, Hannah Schieber, Niklas Corell, Bernhard Egger 0001, Julian Kreimeier, Daniel Roth 0001 |
VR | 4 |
| 2024 | Approximating Intersections and Differences Between Linear Statistical Shape Models Using Markov Chain Monte CarloabstractTo date, the comparison of Statistical Shape Models (SSMs) is often solely performance-based, carried out by means of simplistic metrics such as compactness, generalization, or specificity. Any similarities or differences between the actual shape spaces can neither be visualized nor quantified. In this paper, we present a new method to qualitatively compare two linear SSMs in dense correspondence by computing approximate intersection spaces and set-theoretic differences between the (hyper-ellipsoidal) allowable shape domains spanned by the models. To this end, we approximate the distribution of shapes lying in the intersection space using Markov chain Monte Carlo and subsequently apply Principal Component Analysis (PCA) to the posterior samples, eventually yielding a new SSM of the intersection space. We estimate differences between linear SSMs in a similar manner; here, however, the resulting spaces are no longer convex and we do not apply PCA but instead use the posterior samples for visualization. We showcase the proposed algorithm qualitatively by computing and analyzing intersection spaces and differences between publicly available face models, focusing on gender-specific male and female as well as identity and expression models. Our quantitative evaluation based on SSMs built from synthetic and real-world data sets provides detailed evidence that the introduced method is able to recover ground-truth intersection spaces and differences accurately. Maximilian Weiherer, Finn Klein, Bernhard Egger 0001 |
WACV | 3 |
| 2024 | NeRFtrinsic Four: An end-to-end trainable NeRF jointly optimizing diverse intrinsic and extrinsic camera parametersabstractNovel view synthesis using neural radiance fields (NeRF) is the state-of-the-art technique for generating high-quality images from novel viewpoints. Existing methods require a priori knowledge about extrinsic and intrinsic camera parameters. This limits their applicability to synthetic scenes, or real-world scenarios with the necessity of a preprocessing step. Current research on the joint optimization of camera parameters and NeRF focuses on refining noisy extrinsic camera parameters and often relies on the preprocessing of intrinsic camera parameters. Further approaches are limited to cover only one single camera intrinsic. To address these limitations, we propose a novel end-to-end trainable approach called NeRFtrinsic Four. We utilize Gaussian Fourier features to estimate extrinsic camera parameters and dynamically predict varying intrinsic camera parameters through the supervision of the projection error. Our approach outperforms existing joint optimization methods on LLFF and BLEFF. In addition to these existing datasets, we introduce a new dataset called iFF with varying intrinsic camera parameters. NeRFtrinsic Four is a step forward in joint optimization NeRF-based view synthesis and enables more realistic and flexible rendering in real-world scenarios with varying camera parameters. • A dynamic joint end-to-end trainable optimization framework, capable of handling diverse cameras. • A pose-multilayer perceptron (MLP), using Gaussian Fourier features for the handling of challenging poses. • Our novel iFF dataset focusing on the challenge of diverse cameras, on which we demonstrate the advantages of NeRFtrinsic Four. Hannah Schieber, Fabian Deuser, Bernhard Egger 0001, Norbert Oswald, Daniel Roth 0001 |
Comput. Vis. Image Underst. | 3 |
| 2024 | Building 3D Generative Models from Minimal DataabstractWe propose a method for constructing generative models of 3D objects from a single 3D mesh and improving them through unsupervised low-shot learning from 2D images. Our method produces a 3D morphable model that represents shape and albedo in terms of Gaussian processes. Whereas previous approaches have typically built 3D morphable models from multiple high-quality 3D scans through principal component analysis, we build 3D morphable models from a single scan or template. As we demonstrate in the face domain, these models can be used to infer 3D reconstructions from 2D data (inverse graphics) or 3D data (registration). Specifically, we show that our approach can be used to perform face recognition using only a single 3D template (one scan total, not one per person). We extend our model to a preliminary unsupervised learning framework that enables the learning of the distribution of 3D faces using one 3D template and a small number of 2D images. Our approach is motivated as a potential model for the origins of face perception in human infants, who appear to start with an innate face template and subsequently develop a flexible system for perceiving the 3D structure of any novel face from experience with only 2D images of a relatively small number of familiar faces. Skylar Sutherland, Bernhard Egger 0001, Josh Tenenbaum |
Int. J. Comput. Vis. | 2 |
| 2024 | Survey and systematization of 3D object detection models and methodsabstractAbstract Strong demand for autonomous vehicles and the wide availability of 3D sensors are continuously fueling the proposal of novel methods for 3D object detection. In this paper, we provide a comprehensive survey of recent developments from 2012–2021 in 3D object detection covering the full pipeline from input data, over data representation and feature extraction to the actual detection modules. We introduce fundamental concepts, focus on a broad range of different approaches that have emerged over the past decade, and propose a systematization that provides a practical framework for comparing these approaches with the goal of guiding future development, evaluation, and application activities. Specifically, our survey and systematization of 3D object detection models and methods can help researchers and practitioners to get a quick overview of the field by decomposing 3DOD solutions into more manageable pieces. Moritz Drobnitzky, Jonas Friederich, Bernhard Egger 0001, Patrick Zschech |
Vis. Comput. | 3 |
| 2023 | Perception of Mooney Faces: Extreme Generalization through Inverse Rendering?
Shreya Kapoor, Maximilian Weiherer, Max H. Siegel, Amir Arsalan Soltani, Ilker Yildirim, Josh Tenenbaum, Bernhard Egger 0001 |
CogSci | 7 |
| 2023 | Robust Model-based Face Reconstruction through Weakly-Supervised Outlier SegmentationabstractIn this work, we aim to enhance model-based face reconstruction by avoiding fitting the model to outliers, i.e. regions that cannot be well-expressed by the model such as occluders or makeup. The core challenge for localizing outliers is that they are highly variable and difficult to annotate. To overcome this challenging problem, we introduce a joint Face-autoencoder and outlier segmentation approach (FOCUS). In particular, we exploit the fact that the outliers cannot be fitted well by the face model and hence can be localized well given a high-quality model fitting. The main challenge is that the model fitting and the outlier segmentation are mutually dependent on each other, and need to be inferred jointly. We resolve this chicken-and-egg problem with an EM-type training strategy, where a face autoencoder is trained jointly with an outlier segmentation network. This leads to a synergistic effect, in which the segmentation network prevents the face encoder from fitting to the outliers, enhancing the reconstruction quality. The improved 3D face reconstruction, in turn, enables the segmentation network to better predict the outliers. To resolve the ambiguity between outliers and regions that are difficult to fit, such as eyebrows, we build a statistical prior from synthetic data that measures the systematic bias in model fitting. Experiments on the NoW testset demonstrate that FOCUS achieves SOTA 3D face reconstruction performance among all baselines trained without 3D annotation. Moreover, our results on CelebA-HQ and AR database show that the segmentation network can localize occluders accurately despite being trained without any segmentation annotation. Chunlu Li, Andreas Morel-Forster, Thomas Vetter, Bernhard Egger 0001, Adam Kortylewski |
CVPR | 4 |
| 2023 | PLIKS: A Pseudo-Linear Inverse Kinematic Solver for 3D Human Body EstimationabstractWe introduce PLIKS (Pseudo-Linear Inverse Kinematic Solver) for reconstruction of a 3D mesh of the human body from a single 2D image. Current techniques directly regress the shape, pose, and translation of a parametric model from an input image through a non-linear mapping with minimal flexibility to any external influences. We approach the task as a model-in-the-loop optimization problem. PLIKS is built on a linearized formulation of the parametric SMPL model. Using PLIKS, we can analytically reconstruct the human model via 2D pixel-aligned vertices. This enables us with the flexibility to use accurate camera calibration information when available. PLIKS offers an easy way to introduce additional constraints such as shape and translation. We present quantitative evaluations which confirm that PLIKS achieves more accurate reconstruction with greater than 10% improvement compared to other state-of-the-art methods with respect to the standard 3D human pose and shape benchmarks while also obtaining a reconstruction error improvement of 12.9 mm on the newer AGORA dataset. Karthik Shetty, Annette Birkhold, Srikrishna Jaganathan, Norbert Strobel, Markus Kowarschik, Andreas K. Maier, Bernhard Egger 0001 |
CVPR | 7 |
| 2023 | State of the Art in Dense Monocular Non-Rigid 3D ReconstructionabstractAbstract 3D reconstruction of deformable (ornon‐rigid) scenes from a set of monocular 2D image observations is a long‐standing and actively researched area of computer vision and graphics. It is an ill‐posed inverse problem, since—without additional prior assumptions—it permits infinitely many solutions leading to accurate projection to the input 2D images. Non‐rigid reconstruction is a foundational building block for downstream applications like robotics, AR/VR, or visual content creation. The key advantage of using monocular cameras is their omnipresence and availability to the end users as well as their ease of use compared to more sophisticated camera set‐ups such as stereo or multi‐view systems. This survey focuses on state‐of‐the‐art methods for dense non‐rigid 3D reconstruction of various deformable objects and composite scenes from monocular videos or sets of monocular views. It reviews the fundamentals of 3D reconstruction and deformation modeling from 2D image observations. We then start from general methods—that handle arbitrary scenes and make only a few prior assumptions—and proceed towards techniques making stronger assumptions about the observed objects and types of deformations (e.g. human faces, bodies, hands, and animals). A significant part of this STAR is also devoted to classification and a high‐level comparison of the methods, as well as an overview of the datasets for training and evaluation of the discussed techniques. We conclude by discussing open challenges in the field and the social aspects associated with the usage of the reviewed methods. Edith Tretschk, Navami Kairanda, Mallikarjun B. R. 0001, Rishabh Dabral, Adam Kortylewski, Bernhard Egger 0001, Marc Habermann, Pascal Fua, Christian Theobalt, Vladislav Golyanik |
Comput. Graph. Forum | 6 |
| 2023 | Learning the shape of female breasts: an open-access 3D statistical shape model of the female breast built from 110 breast scansabstractAbstract We present theRegensburg Breast Shape Model(RBSM)—a 3D statistical shape model of the female breast built from 110 breast scans acquired in a standing position, and the first publicly available. Together with the model, a fully automated, pairwise surface registration pipeline used to establish dense correspondence among 3D breast scans is introduced. Our method is computationally efficient and requires only four landmarks to guide the registration process. A major challenge when modeling female breasts from surface-only 3D breast scans is the non-separability of breast and thorax. In order to weaken the strong coupling between breast and surrounding areas, we propose to minimize thevarianceoutside the breast region as much as possible. To achieve this goal, a novel concept calledbreast probability masks(BPMs) is introduced. A BPM assigns probabilities to each point of a 3D breast scan, telling howlikelyit is that a particular point belongs to the breast area. During registration, we use BPMs to align the template to the target as accurately as possibleinsidethe breast region and only roughly outside. This simple yet effective strategy significantly reduces the unwanted variance outside the breast region, leading to better statistical shape models in which breast shapes are quite well decoupled from the thorax. The RBSM is thus able to produce a variety of different breast shapes as independently as possible from the shape of the thorax. Our systematic experimental evaluation reveals a generalization ability of 0.17 mm and a specificity of 2.8 mm. To underline the expressiveness of the proposed model, we finally demonstrate in two showcase applications how the RBSM can be used for surgical outcome simulation and the prediction of a missing breast from the remaining one. Our model is available at https://www.rbsm.re-mic.de/ . Maximilian Weiherer, Andreas Eigenberger, Bernhard Egger 0001, Vanessa Brébant, Lukas Prantl, Christoph Palm |
Vis. Comput. | 3 |
| 2022 | MOST-GAN: 3D Morphable StyleGAN for Disentangled Face Image ManipulationabstractRecent advances in generative adversarial networks (GANs) have led to remarkable achievements in face image synthesis. While methods that use style-based GANs can generate strikingly photorealistic face images, it is often difficult to control the characteristics of the generated faces in a meaningful and disentangled way. Prior approaches aim to achieve such semantic control and disentanglement within the latent space of a previously trained GAN. In contrast, we propose a framework that a priori models physical attributes of the face such as 3D shape, albedo, pose, and lighting explicitly, thus providing disentanglement by design. Our method, MOST-GAN, integrates the expressive power and photorealism of style-based GANs with the physical disentanglement and flexibility of nonlinear 3D morphable models, which we couple with a state-of-the-art 2D hair manipulation network. MOST-GAN achieves photorealistic manipulation of portrait images with fully disentangled 3D control over their physical attributes, enabling extreme manipulation of lighting, facial expression, and pose variations up to full profile view. Safa C. Medin, Bernhard Egger 0001, Anoop Cherian, Ye Wang 0001, Josh Tenenbaum, Xiaoming Liu 0002, Tim K. Marks |
AAAI | 2 |
| 2022 | Benchmarking mid-level vision with texture-defined 3D objects
Yoni Friedman, Thomas P. O'Connell, Max H. Siegel, Daniel Bear, Tuan Anh Le 0001, Bernhard Egger 0001, Josh Tenenbaum |
CogSci | 6 |
| 2022 | Rotation-Equivariant Conditional Spherical Neural Fields for Learning a Natural Illumination PriorabstractInverse rendering is an ill-posed problem. Previous work has sought to resolve this by focussing on priors for object or scene shape or appearance. In this work, we instead focus on a prior for natural illuminations. Current methods rely on spherical harmonic lighting or other generic representations and, at best, a simplistic prior on the parameters. We propose a conditional neural field representation based on a variational auto-decoder with a SIREN network and, extending Vector Neurons, build equivariance directly into the network. Using this, we develop a rotation-equivariant, high dynamic range (HDR) neural illumination model that is compact and able to express complex, high-frequency features of natural environment maps. Training our model on a curated dataset of 1.6K HDR environment maps of natural scenes, we compare it against traditional representations, demonstrate its applicability for an inverse rendering task and show environment map completion from partial observations. James A. D. Gardner, Bernhard Egger 0001, William A. P. Smith |
NeurIPS | 2 |
| 2021 | Explaining the Gestalt principle of common fate as amortized inference
Yoni Friedman, Tuan Anh Le 0001, Bernhard Egger 0001, Max H. Siegel, Josh Tenenbaum |
CogSci | 3 |
| 2021 | Seeing in the dark: Testing deep neural network and analysis-by-synthesis accounts of 3D shape perception with highly degraded images
Hakan Yilmaz, Gargi Singh, Bernhard Egger 0001, Josh Tenenbaum, Ilker Yildirim |
CogSci | 3 |
| 2021 | Identity-Expression Ambiguity in 3D Morphable Face Modelsabstract3D Morphable Models are a class of generative models commonly used to model faces. They are typically applied to ill-posed problems such as 3D reconstruction from 2D data. Several ambiguities in this problem's image formation process have been studied explicitly. We demonstrate that nonorthogonality of the variation in identity and expression can cause identity-expression ambiguity in 3D Morphable Models, and that in practice expression and identity are far from orthogonal and can explain each other surprisingly well. Whilst previously reported ambiguities only arise in an inverse rendering setting, identity-expression ambiguity emerges in the 3D shape generation process itself. We demonstrate this effect with 3D shapes directly as well as through an inverse rendering task, and use two popular models built from high quality 3D scans as well as a model built from a large collection of 2D images and videos. We explore this issue's implications for inverse rendering and observe that it cannot be resolved by a purely statistical prior on identity and expression deformations. Bernhard Egger 0001, Skylar Sutherland, Safa C. Medin, Josh Tenenbaum |
FG | 1 |
| 2020 | Inverse Rendering Best Explains Face Perception Under Extreme Illuminations
Bernhard Egger 0001, Max H. Siegel, Riya Arora, Amir Arsalan Soltani, Ilker Yildirim, Josh Tenenbaum |
CogSci | 1 |
| 2020 | Learning a Generative Model of Human Faces Through Inverse Rendering
Skylar Sutherland, Bernhard Egger 0001, Josh Tenenbaum |
CogSci | 2 |
| 2020 | A Morphable Face Albedo ModelabstractIn this paper, we bring together two divergent strands of research: photometric face capture and statistical 3D face appearance modelling. We propose a novel lightstage capture and processing pipeline for acquiring ear-to-ear, truly intrinsic diffuse and specular albedo maps that fully factor out the effects of illumination, camera and geometry. Using this pipeline, we capture a dataset of 50 scans and combine them with the only existing publicly available albedo dataset (3DRFE) of 23 scans. This allows us to build the first morphable face albedo model. We believe this is the first statistical analysis of the variability of facial specular albedo maps. This model can be used as a plug in replacement for the texture model of the Basel Face Model and we make our new albedo model publicly available. We ensure careful spectral calibration such that our model is built in a linear sRGB space, suitable for inverse rendering of images taken by typical cameras. We demonstrate our model in a state of the art analysis-by-synthesis 3DMM fitting pipeline, are the first to integrate specular map estimation and outperform the Basel Face Model in albedo reconstruction. William A. P. Smith, Alassane Seck, Hannah M. Dee, Bernard Tiddeman, Josh Tenenbaum, Bernhard Egger 0001 |
CVPR | 6 |
| 2020 | 3D Morphable Face Models - Past, Present, and FutureabstractIn this article, we provide a detailed survey of 3D Morphable Face Models over the 20 years since they were first proposed. The challenges in building and applying these models, namely, capture, modeling, image formation, and image analysis, are still active research topics, and we review the state-of-the-art in each of these areas. We also look ahead, identifying unsolved challenges, proposing directions for future research, and highlighting the broad range of current and future applications. Bernhard Egger 0001, William A. P. Smith, Ayush Tewari, Stefanie Wuhrer, Michael Zollhöfer, Thabo Beeler, Florian Bernard 0001, Timo Bolkart, Adam Kortylewski, Sami Romdhani, Christian Theobalt, Volker Blanz, Thomas Vetter |
ACM Trans. Graph. | 1 |
| 2018 | Morphable Face Models - An Open FrameworkabstractIn this paper, we present a novel open-source pipeline for face registration based on Gaussian processes as well as an application to face image analysis. Non-rigid registration of faces is significant for many applications in computer vision, such as the construction of 3D Morphable face models (3DMMs). Gaussian Process Morphable Models (GPMMs) unify a variety of non-rigid deformation models with B-splines and PCA models as examples. GPMM separate problem specific requirements from the registration algorithm by incorporating domain-specific adaptions as a prior model. The novelties of this paper are the following: (i) We present a strategy and modeling technique for face registration that considers symmetry, multi-scale and spatially-varying details. The registration is applied to neutral faces and facial expressions. (ii) We release an open-source software framework for registration model-building demonstrated on the publicly available BU3D-FE database. The released pipeline also contains an implementation of an Analysis-by-Synthesis model adaption of 2D face images, tested on the Multi-PIE and LFW database. This enables the community to reproduce, evaluate and compare the individual steps of registration to model-building and 3D/2D model fitting. (iii) Along with the framework release, we publish a new version of the Basel Face Model (BFM-2017) with an improved age distribution and an additional facial expression model. Thomas Gerig, Andreas Morel-Forster, Clemens Blumer, Bernhard Egger 0001, Marcel Lüthi, Sandro Schönborn, Thomas Vetter |
FG | 4 |
| 2018 | A Parametric Freckle Model for FacesabstractWe propose a novel stochastic generative parametric freckle model for the analysis and synthesis of human faces. Morphable Models are the state-of-the-art generative parametric face models. However, they are unable to synthesize freckles which are part of atural face variation. The deficiency lies in requiring point-to-point correspondence on the texture pixels. We propose to assume a correspondence between freckle density and not the freckles themselves. We propose a model that is stochastic, generative, and parametric and generates freckles with a point process according to a density and size distribution. The resulting model can synthesize photo-realistic freckles according to observations as well as add freckles to existing faces. We create more realistic faces than with Morphable Models alone and allow for detailed face pigment analysis. Bernhard Egger 0001, Thomas Vetter |
FG | 2 |
| 2018 | Occlusion-Aware 3D Morphable Models and an Illumination Prior for Face Image Analysis
Bernhard Egger 0001, Sandro Schönborn, Adam Kortylewski, Andreas Morel-Forster, Clemens Blumer, Thomas Vetter |
Int. J. Comput. Vis. | 1 |
| 2017 | Efficient Global Illumination for Morphable ModelsabstractWe propose an efficient self-shadowing illumination model for Morphable Models. Simulating self-shadowing with ray casting is computationally expensive which makes them impractical in Analysis-by-Synthesis methods for object reconstruction from single images. Therefore, we propose to learn self-shadowing for Morphable Model parameters directly with a linear model. Radiance transfer functions are a powerful way to represent self-shadowing used within the precomputed radiance transfer framework (PRT). We build on PRT to render deforming objects with self-shadowing at interactive frame rates. It can be illuminated efficiently by environment maps represented with spherical harmonics. The result is an efficient global illumination method for Morphable Models, exploiting an approximated radiance transfer. We apply the method to fitting Morphable Model parameters to a single image of a face and demonstrate that considering self-shadowing improves shape reconstruction. Sandro Schönborn, Bernhard Egger 0001, Lavrenti Frobeen, Thomas Vetter |
ICCV | 3 |
| 2017 | Markov Chain Monte Carlo for Automated Face Image Analysis
Sandro Schönborn, Bernhard Egger 0001, Andreas Morel-Forster, Thomas Vetter |
Int. J. Comput. Vis. | 2 |
| 2016 | Occlusion-aware 3D Morphable Face Models
Bernhard Egger 0001, Clemens Blumer, Andreas Morel-Forster, Sandro Schönborn, Thomas Vetter |
BMVC | 1 |
| 2015 | Background modeling for generative image models
Sandro Schönborn, Bernhard Egger 0001, Andreas Morel-Forster, Thomas Vetter |
Comput. Vis. Image Underst. | 2 |