VLDB 2026 Research / reviewers in the wild / expert
Zhixin Shu
dblp:129/3987
· DBLP profile ↗
37ranked-venue papers
6as first author
27since 2021 · last 2025
0009-0004-3878-4805ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 34 · 6 first-author · 25 since 2021Artificial intelligence and machine learning · 29 · 3 first-author · 22 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Rig3DGS: Creating Controllable Portraits From Casual Monocular VideosabstractWe present Rig3DGS, a novel technique for creating reanimatable 3D portraits from short monocular smartphone videos. Rig3DGS learns to reconstruct a set of controllable 3D Gaussians from a monocular video of a dynamic subject captured with varying head poses and facial expressions in an in-the-wild scene. In contrast to synchronized multi-view studio captures, this in-the-wild, single camera setup brings fresh challenges to learning high quality 3D Gaussians. We address these challenges by learning to deform 3D Gaussians from a fixed canonical space to the deformed space that is consistent with the target facial expression and headpose. Our key contribution is a carefully designed deformation model that is guided by a 3D face morphable model. This deformation not only enables control over facial expression and head-poses but also allows our method to generates high-quality photorealistic renders of the whole scene. Once trained, Rig3DGS is able to generate photorealistic renders of a subject and their scene for novel facial expression, head-poses, and viewing directions. Through extensive experiments we demonstrate that Rig3DGS significantly outperforms prior art while being orders of magnitude faster. Alfredo Rivero, Shahrukh Athar, Zhixin Shu, Dimitris Samaras |
3DV | 3 |
| 2025 | SynthLight: Portrait Relighting with Diffusion Model by Learning to Re-render Synthetic FacesabstractWe introduce SynthLight, a diffusion model for portrait relighting. We frame image relighting as a re-rendering problem, where pixels are transformed in response to changes in environmental lighting. Using a physically-based rendering engine, we create a dataset to simulate this lighting-conditioned transformation with 3D head assets under varying lighting. We propose training and inference strategies to bridge the gap between the synthetic and real image domains: (1) multi-task training leveraging real human portraits without lighting labels; (2) an inference time diffusion sampling scheme based on classifier-free guidance leveraging the input portrait to better preserve details. Our method generalizes to diverse real portraits and produces realistic illumination effects, including specular highlights and cast shadows, while preserving identity. Our quantitative experiments on light stage data demonstrate results comparable to state-of-the-art relighting methods. Our qualitative results on in-the-wild images showcase rich and unprecedented illumination effects. Sumit Chaturvedi, Mengwei Ren, Yannick Hold-Geoffroy, Julie Dorsey, Zhixin Shu |
CVPR | 6 |
| 2025 | Free-viewpoint Human Animation with Pose-correlated Reference SelectionabstractDiffusion-based human animation aims to animate a human character based on a source human image as well as driving signals such as a sequence of poses. Leveraging the generative capacity of diffusion model, existing approaches are able to generate high-fidelity poses, but struggle with significant viewpoint changes, especially in zoom-in/zoom-out scenarios where camera-character distance varies. This limits the applications such as cinematic shot type plan or camera control. We propose a pose-correlated reference selection diffusion network, supporting substantial viewpoint variations in human animation. Our key idea is to enable the network to utilize multiple reference images as input, since significant viewpoint changes often lead to missing appearance details on the human body. To eliminate the computational cost, we first introduce a novel pose correlation module to compute similarities between non-aligned target and source poses, and then propose an adaptive reference selection strategy, utilizing the attention map to identify key regions for animation generation. To train our model, we curated a large dataset from public TED talks featuring varied shots of the same character, helping the model learn synthesis for different perspectives. Our experimental results show that with the same number of reference images, our model performs favorably compared to the current SOTA methods under large viewpoint changes. We further show that the adaptive reference selection is able to choose the most relevant reference regions to generate humans under free viewpoints. Fa-Ting Hong, Qinjie Lin, Luchuan Song, Zhixin Shu, Duygu Ceylan, Dan Xu 0002 |
CVPR | 6 |
| 2025 | MegaSynth: Scaling Up 3D Scene Reconstruction with Synthesized DataabstractWe propose scaling up 3D scene reconstruction by training with synthesized data. At the core of our work is MegaSynth, a procedurally generated 3D dataset comprising 700K scenes—over 50 times larger than the prior real dataset DL3DV—dramatically scaling the training data. To enable scalable data generation, our key idea is eliminating semantic information, removing the need to model complex semantic priors such as object affordances and scene composition. Instead, we model scenes with basic spatial structures and geometry primitives, offering scalability. Besides, we control data complexity to facilitate training while loosely aligning it with real-world data distribution to benefit real-world generalization. We explore training LRMs with both MegaSynth and available real data. Experiment results show that joint training or pre-training with MegaSynth improves reconstruction quality by 1.2 to 1.8 dB PSNR across diverse image domains. Moreover, models trained solely on MegaSynth perform comparably to those trained on real data, underscoring the low-level nature of 3D reconstruction. Additionally, we provide an in-depth analysis of MegaSynth’s properties for enhancing model capability, training stability, and generalization, as well as application to other tasks. Hanwen Jiang, Zexiang Xu, Desai Xie, Haian Jin, Fujun Luan, Zhixin Shu, Kai Zhang 0045, Sai Bi, Xin Sun 0014, Jiuxiang Gu, Qixing Huang, Georgios Pavlakos, Hao Tan 0002 |
CVPR | 7 |
| 2025 | Comprehensive Relighting: Generalizable and Consistent Monocular Human Relighting and HarmonizationabstractThis paper introduces Comprehensive Relighting, the first all-in-one approach that can both control and harmonize the lighting from an image or video of humans with arbitrary body parts from any scene. Building such a generalizable model is extremely challenging due to the lack of dataset, restricting existing image-based relighting models to a specific scenario (e.g., face or static human). To address this challenge, we repurpose a pre-trained diffusion model as a general image prior and jointly model the human relighting and background harmonization in the coarse-to-fine framework. To further enhance the temporal coherence of the relighting, we introduce an unsupervised temporal lighting model that learns the lighting cycle consistency from many real-world videos without any ground truth. In inference time, our temporal lighting module is combined with the diffusion models through the spatio-temporal feature blending algorithms without extra training; and we apply a new guided refinement as a post-processing to pre-serve the high-frequency details from the input image. In the experiments, Comprehensive Relighting shows a strong generalizability and lighting temporal coherence, outperforming existing image-based human relighting and harmonization methods. Xin Sun 0014, Krishna Kumar Singh, Zhixin Shu, He Zhang 0004, Jimei Yang, Nanxuan Zhao, Tuanfeng Y. Wang, Simon S. Chen, Ulrich Neumann, Jae Shin Yoon |
CVPR | 5 |
| 2025 | Augmented Mass-Spring Model for Real-Time Dense Hair Simulation
Jorge Alejandro Amador Herrera, Yi Zhou 0023, Xin Sun 0014, Zhixin Shu, Chengan He, Sören Pirk, Dominik L. Michels |
ICCV | 4 |
| 2025 | FaceLift: Learning Generalizable Single Image 3D Face Reconstruction From Synthetic Heads
Weijie Lyu, Yi Zhou 0023, Ming-Hsuan Yang 0001, Zhixin Shu |
ICCV | 4 |
| 2025 | Perm: A Parametric Representation for Multi-Style 3D Hair ModelingabstractWe present Perm, a learned parametric representation of human 3D hair designed to facilitate various hair-related applications. Unlike previous work that jointly models the global hair structure and local curl patterns, we propose to disentangle them using a PCA-based strand representation in the frequency domain, thereby allowing more precise editing and output control. Specifically, we leverage our strand representation to fit and decompose hair geometry textures into low- to high-frequency hair structures, termed guide textures and residual textures, respectively. These decomposed textures are later parameterized with different generative models, emulating common stages in the hair grooming process. We conduct extensive experiments to validate the architecture design of Perm, and finally deploy the trained model as a generic prior to solve task-agnostic problems, further showcasing its flexibility and superiority in tasks such as single-view hair reconstruction, hairstyle editing, and hair-conditioned image generation. More details can be found on our project page: https://cs.yale.edu/homes/che/projects/perm/. Chengan He, Xin Sun 0014, Zhixin Shu, Fujun Luan, Sören Pirk, Jorge Alejandro Amador Herrera, Dominik L. Michels, Tuanfeng Y. Wang, Meng Zhang 0043, Holly E. Rushmeier, Yi Zhou 0023 |
ICLR | 3 |
| 2024 | Controllable Dynamic Appearance for Neural 3D PortraitsabstractRecent advances in Neural Radiance Fields (NeRFs) have made it possible to reconstruct and reanimate dynamic portrait scenes with control over head-pose, facial expressions and viewing direction. However, training such models assumes photometric consistency over the deformed region e.g. the face must be evenly lit as it deforms with changing head-pose and facial expression. Such photometric consistency across frames of a video is hard to maintain, even in studio environments, thus making the created reanimatable neural portraits prone to artefacts during reanimation. In this work, we propose CoDyNeRF, a system that enables the creation of fully controllable 3D portraits in real-world capture conditions. CoDyNeRF learns to approximate illumination dependent effects via a dynamic appearance model in the canonical space that is conditioned on predicted surface normals and the facial expressions and head-pose deformations. The surface normals prediction is guided using 3DMM normals that act as a coarse prior for the normals of the human head, where direct prediction of normals is hard due to rigid and non-rigid deformations induced by head-pose and facial expression changes. Using only a smartphone-captured short video of a subject for training, we demonstrate the effectiveness of our method on free view synthesis of a portrait scene with explicit head pose and expression controls, and realistic lighting effects. Shahrukh Athar, Zhixin Shu, Zexiang Xu, Fujun Luan, Sai Bi, Kalyan Sunkavalli, Dimitris Samaras |
3DV | 2 |
| 2024 | Holo-Relighting: Controllable Volumetric Portrait Relighting from a Single ImageabstractAt the core of portrait photography is the search for ideal lighting and viewpoint. The process often requires advanced knowledge in photography and an elaborate studio setup. In this work, we propose Holo-Relighting, a volumetric relighting method that is capable of synthesizing novel viewpoints, and novel lighting from a single image. Holo-Relighting leverages the pretrained 3D GAN (EG3D) to reconstruct geometry and appearance from an input portrait as a set of 3D-aware features. We design a relighting module conditioned on a given lighting to process these features, and predict a relit 3D representation in the form of a tri-plane, which can render to an arbitrary viewpoint through volume rendering. Besides viewpoint and lighting control, Holo-Relighting also takes the head pose as a condition to enable head-pose-dependent lighting effects. With these novel designs, Holo-Relighting can generate complex non-Lambertian lighting effects (e.g., specular highlights and cast shadows) without using any explicit physical lighting priors. We train Holo-Relighting with data captured with a light stage, and propose two data-rendering techniques to improve the data quality for training the volumetric relighting system. Through quantitative and qualitative experiments, we demonstrate Holo-Relighting can achieve state-of-the-arts relighting quality with better photorealism, 3D consistency and controllability. Yiqun Mei, Yu Zeng 0001, He Zhang 0004, Zhixin Shu, Xuaner Cecilia Zhang, Sai Bi, Jianming Zhang 0001, Hyunjoon Jung, Vishal M. Patel |
CVPR | 4 |
| 2024 | Relightful Harmonization: Lighting-Aware Portrait Background ReplacementabstractPortrait harmonization aims to composite a subject into a new background, adjusting its lighting and color to ensure harmony with the background scene. Existing harmo-nization techniques often only focus on adjusting the global color and brightness of the foreground and ignore crucial illumination cues from the background such as apparent lighting direction, leading to unrealistic compositions. We introduce Relightful Harmonization, a lighting-aware diffusion model designed to seamlessly harmonize sophisticated lighting effect for the foreground portrait using any back-ground image. Our approach unfolds in three stages. First, we introduce a lighting representation module that allows our diffusion model to encode lighting information from target image background. Second, we introduce an alignment network that aligns lighting features learned from image background with lighting features learned from panorama environment maps, which is a complete representation for scene illumination. Last, to further boost the photorealism of the proposed method, we introduce a novel data simulation pipeline that generates synthetic training pairs from a diverse range of natural images, which are used to refine the model. Our method outperforms existing benchmarks in visual fidelity and lighting coherence, showing superior generalization in real-world testing scenarios, highlighting its versatility and practicality. Mengwei Ren, Wei Xiong 0008, Jae Shin Yoon, Zhixin Shu, Jianming Zhang 0001, Hyunjoon Jung, Guido Gerig, He Zhang 0004 |
CVPR | 4 |
| 2024 | Carve3D: Improving Multi-view Reconstruction Consistency for Diffusion Models with RL FinetuningabstractMulti-view diffusion models, obtained by applying Su-pervised Finetuning (SFT) to text-to-image diffusion mod-els, have driven recent breakthroughs in text-to-3D re-search. However, due to the limited size and quality of ex-isting 3D datasets, they still suffer from multi-view incon-sistencies and Neural Radiance Field (NeRF) reconstruction artifacts. We argue that multi-view diffusion models can benefit from further Reinforcement Learning Finetuning (RLFT), which allows models to learn from the data generated by themselves and improve beyond their dataset limitations during SFT. To this end, we introduce Carve3D, an improved RLFT algorithm coupled with a novel Multi-view Reconstruction Consistency (MRC) metric, to enhance the consistency of multi-view diffusion models. To mea-sure the MRC metric on a set of multi-view images, we compare them with their corresponding NeRF renderings at the same camera viewpoints. The resulting model, which we denote as Carve3DM, demonstrates superior multi-view consistency and NeRF reconstruction quality than existing models. Our results suggest that pairing SFT with Carve3D's RLFT is essential for developing multi-view-consistent diffusion models, mirroring the standard Large Language Model (LLM) alignment pipeline. Our code, training and testing data, and video results are available at: https://desaixie.github.io/carve-3d. Desai Xie, Hao Tan 0002, Xin Sun 0014, Zhixin Shu, Yi Zhou 0023, Sai Bi, Sören Pirk, Arie E. Kaufman |
CVPR | 5 |
| 2024 | In-N-Out: Faithful 3D GAN Inversion with Volumetric Decomposition for Face Editingabstract3D-aware GANs offer new capabilities for view synthe-sis while preserving the editing functionalities of their 2D counterparts. GAN inversion is a crucial step that seeks the latent code to reconstruct input images or videos, subsequently enabling diverse editing tasks through manipulation of this latent code. However, a model pretrained on a particular dataset (e.g., FFHQ) often has difficulty re-constructing images with out-of-distribution (OOD) objects such as faces with heavy make-up or occluding objects. We address this issue by explicitly modeling OOD objects from the input in 3D-aware GANs. Our core idea is to represent the image using two individual neural radiance fields: one for the in-distribution content and the other for the out-of-distribution object. The final reconstruction is achieved by optimizing the composition of these two radiance fields with carefully designed regularization. We demonstrate that our explicit decomposition alleviates the inherent tradeoff between reconstruction fidelity and editability. We evaluate reconstruction accuracy and editability of our method on challenging real face images and videos and showcase favorable results against other baselines. More results can found at https://in-n-out-3d.github.io/. Zhixin Shu, Cameron Smith, Seoung Wug Oh, Jia-Bin Huang 0001 |
CVPR | 2 |
| 2024 | COMPOSE: Comprehensive Portrait Shadow Editing
Andrew Hou, Zhixin Shu, Xuaner Cecilia Zhang, He Zhang 0004, Yannick Hold-Geoffroy, Jae Shin Yoon, Xiaoming Liu 0002 |
ECCV (61) | 2 |
| 2024 | LRM-Zero: Training Large Reconstruction Models with Synthesized DataabstractWe present LRM-Zero, a Large Reconstruction Model (LRM) trained entirely on synthesized 3D data, achieving high-quality sparse-view 3D reconstruction. The core of LRM-Zero is our procedural 3D dataset, Zeroverse, which is automatically synthesized from simple primitive shapes with random texturing and augmentations (e.g., height fields, boolean differences, and wireframes). Unlike previous 3D datasets (e.g., Objaverse) which are often captured or crafted by humans to approximate real 3D data, Zeroverse completely ignores realistic global semantics but is rich in complex geometric and texture details that are locally similar to or even more intricate than real objects. We demonstrate that our LRM-Zero, trained with our fully synthesized Zeroverse, can achieve high visual quality in the reconstruction of real-world objects, competitive with models trained on Objaverse. We also analyze several critical design choices of Zeroverse that contribute to LRM-Zero's capability and training stability. Our work demonstrates that 3D reconstruction, one of the core tasks in 3D vision, can potentially be addressed without the semantics of real-world objects. The Zeroverse's procedural synthesis code and interactive visualization are available at: https://desaixie.github.io/lrm-zero/. Desai Xie, Sai Bi, Zhixin Shu, Kai Zhang 0045, Zexiang Xu, Yi Zhou 0023, Sören Pirk, Arie E. Kaufman, Xin Sun 0014, Hao Tan 0002 |
NeurIPS | 3 |
| 2024 | Consistent Multimodal Generation via A Unified GAN FrameworkabstractWe investigate how to generate multimodal image outputs, such as RGB, depth, and surface normals, with a single generative model. The challenge is to produce outputs that are realistic, and also consistent with each other. Our solution builds on the StyleGAN3 architecture, with a shared backbone and modality-specific branches in the last layers of the synthesis network, and we propose per-modality fidelity discriminators and a cross-modality consistency discriminator. In experiments on the Stanford2D3D dataset, we demonstrate realistic and consistent generation of RGB, depth, and normal images. We also show a training recipe to easily extend our pretrained model on a new domain, even with a few pairwise data. We further evaluate the use of synthetically generated RGB and depth pairs for training or fine-tuning depth estimators. Code will be available at here. Zhen Zhu 0006, Yijun Li 0001, Weijie Lyu, Krishna Kumar Singh, Zhixin Shu, Sören Pirk, Derek Hoiem |
WACV | 5 |
| 2024 | Generative Portrait Shadow RemovalabstractWe introduce a high-fidelity portrait shadow removal model that can effectively enhance the image of a portrait by predicting its appearance under disturbing shadows and highlights. Portrait shadow removal is a highly ill-posed problem where multiple plausible solutions can be found based on a single image. For example, disentangling complex environmental lighting from original skin color is a non-trivial problem. While existing works have solved this problem by predicting the appearance residuals that can propagate local shadow distribution, such methods are often incomplete and lead to unnatural predictions, especially for portraits with hard shadows. We overcome the limitations of existing local propagation methods by formulating the removal problem as a generation task where a diffusion model learns to globally rebuild the human appearance from scratch as a condition of an input portrait image. For robust and natural shadow removal, we propose to train the diffusion model with a compositional repurposing framework: a pre-trained text-guided image generation model is first fine-tuned to harmonize the lighting and color of the foreground with a background scene by using a background harmonization dataset; and then the model is further fine-tuned to generate a shadow-free portrait image via a shadow-paired dataset. To overcome the limitation of losing fine details in the latent diffusion model, we propose a guided-upsampling network to restore the original high-frequency details (e.g. , wrinkles and dots) from the input image. To enable our compositional training framework, we construct a high-fidelity and large-scale dataset using a lightstage capturing system and synthetic graphics simulation. Our generative framework effectively removes shadows caused by both self and external occlusions while maintaining original lighting distribution and high-frequency details. Our method also demonstrates robustness to diverse subjects captured in real environments. Jae Shin Yoon, Zhixin Shu, Mengwei Ren, Cecilia Zhang, Yannick Hold-Geoffroy, Krishna Kumar Singh, He Zhang 0004 |
ACM Trans. Graph. | 2 |
| 2023 | PaletteNeRF: Palette-based Appearance Editing of Neural Radiance FieldsabstractRecent advances in neural radiance fields have enabled the high-fidelity 3D reconstruction of complex scenes for novel view synthesis. However, it remains underexplored how the appearance of such representations can be efficiently edited while maintaining photorealism. In this work, we present PaletteNeRF, a novel method for photorealistic appearance editing of neural radiance fields (NeRF) based on 3D color decomposition. Our method decom-poses the appearance of each 3D point into a linear combination of palette-based bases (i.e., 3D segmentations defined by a group of NeRF-type functions) that are shared across the scene. While our palette-based bases are view-independent, we also predict a view-dependent function to capture the color residual (e.g., specular shading). During training, we jointly optimize the basis functions and the color palettes, and we also introduce novel regulariz-ers to encourage the spatial coherence of the decomposition. Our method allows users to efficiently edit the appearance of the 3D scene by modifying the color palettes. We also extend our framework with compressed semantic features for semantic-aware appearance editing. We demonstrate that our technique is superior to baseline methods both quantitatively and qualitatively for appearance editing of complex real-world scenes. Our project page is https://palettenerf.github.io. Zhengfei Kuang, Fujun Luan, Sai Bi, Zhixin Shu, Gordon Wetzstein, Kalyan Sunkavalli |
CVPR | 4 |
| 2023 | LightPainter: Interactive Portrait Relighting with Freehand ScribbleabstractRecent portrait relighting methods have achieved realistic results of portrait lighting effects given a desired lighting representation such as an environment map. However, these methods are not intuitive for user interaction and lack precise lighting control. We introduce LightPainter, a scribble-based relighting system that allows users to interactively manipulate portrait lighting effect with ease. This is achieved by two conditional neural networks, a delighting module that recovers geometry and albedo optionally conditioned on skin tone, and a scribble-based module for re-lighting. To train the relighting module, we propose a novel scribble simulation procedure to mimic real user scribbles, which allows our pipeline to be trained without any human annotations. We demonstrate high-quality and flexible portrait lighting editing capability with both quantitative and qualitative experiments. User study comparisons with commercial lighting editing tools also demonstrate consistent user preference for our method. Yiqun Mei, He Zhang 0004, Xuaner Cecilia Zhang, Jianming Zhang 0001, Zhixin Shu, Yilin Wang 0002, Zijun Wei, Hyunjoon Jung, Vishal M. Patel |
CVPR | 5 |
| 2023 | FLAME-in-NeRF: Neural control of Radiance Fields for Free View Face AnimationabstractThis paper presents a neural rendering method for controllable portrait video synthesis. Recent advances in volumetric neural rendering, such as neural radiance fields (NeRF), have enabled the photorealistic novel view synthesis of static scenes with impressive results. However, modeling dynamic and controllable objects as part of a scene with such scene representations is still challenging. In this work, we design a system that enables 1) novel view synthesis for portrait video, of both the human subject and the scene they are in and 2) explicit control of the facial expressions through a low-dimensional expression representation. We represent the distribution of human facial expressions using the expression parameters of a 3D Morphable Model (3DMM) and condition the NeRF volumetric function on them. In order to guide the network to learn disentangled control for static scene appearance and dynamic facial actions, we impose a spatial prior via 3DMM fitting. We show the effectiveness of our method on free view synthesis of portrait videos with expression controls. To train a scene, our method only requires a short video of a subject captured by a mobile device. Shahrukh Athar, Zhixin Shu, Dimitris Samaras |
FG | 2 |
| 2022 | RigNeRF: Fully Controllable Neural 3D PortraitsabstractVolumetric neural rendering methods, such as neural radiance fields (NeRFs), have enabled photo-realistic novel view synthesis. However, in their standard form, NeRFs do not support the editing of objects, such as a human head, within a scene. In this work, we propose RigNeRF, a system that goes beyond just novel view synthesis and enables full control of head pose and facial expressions learned from a single portrait video. We model changes in head pose and facial expressions using a deformation field that is guided by a 3D morphable face model (3DMM). The 3DMM effectively acts as a prior for RigNeRF that learns to predict only residuals to the 3DMM deformations and allows us to render novel (rigid) poses and (non-rigid) expressions that were not present in the input sequence. Using only a smartphone-captured short video of a subject for training, we demonstrate the effectiveness of our method on free view synthesis of a portrait scene with explicit head pose and expression controls. Shahrukh Athar, Zexiang Xu, Kalyan Sunkavalli, Eli Shechtman, Zhixin Shu |
CVPR | 5 |
| 2022 | Point-NeRF: Point-based Neural Radiance FieldsabstractVolumetric neural rendering methods like NeRF [34] generate high-quality view synthesis results but are optimized per-scene leading to prohibitive reconstruction time. On the other hand, deep multi-view stereo methods can quickly reconstruct scene geometry via direct network inference. Point-NeRF combines the advantages of these two approaches by using neural 3D point clouds, with associated neural features, to model a radiance field. Point-NeRF can be rendered efficiently by aggregating neural point features near scene surfaces, in a ray marching-based rendering pipeline. Moreover, Point-NeRF can be initialized via direct inference of a pre-trained deep network to produce a neural point cloud; this point cloud can be finetuned to surpass the visual quality of NeRF with 30× faster training time. Point-NeRF can be combined with other 3D re-construction methods and handles the errors and outliers in such methods via a novel pruning and growing mechanism. Qiangeng Xu, Zexiang Xu, Julien Philip, Sai Bi, Zhixin Shu, Kalyan Sunkavalli, Ulrich Neumann |
CVPR | 5 |
| 2022 | Learning Motion-Dependent Appearance for High-Fidelity Rendering of Dynamic Humans from a Single CameraabstractAppearance of dressed humans undergoes a complex geometric transformation induced not only by the static pose but also by its dynamics, i.e., there exists a number of cloth geometric configurations given a pose depending on the way it has moved. Such appearance modeling conditioned on motion has been largely neglected in existing human rendering methods, resulting in rendering of physically implausible motion. A key challenge of learning the dynamics of the appearance lies in the requirement of a prohibitively large amount of observations. In this paper, we present a compact motion representation by enforcing equivariance—a representation is expected to be transformed in the way that the pose is transformed. We model an equivariant encoder that can generate the generalizable representation from the spatial and temporal derivatives of the 3D body surface. This learned representation is decoded by a compositional multi-task decoder that renders high fidelity time-varying appearance. Our experiments show that our method can generate a temporally coherent video of dynamic humans for unseen body poses and novel views given a single view video. Jae Shin Yoon, Duygu Ceylan, Tuanfeng Y. Wang, Jingwan Lu, Jimei Yang, Zhixin Shu, Hyun Soo Park |
CVPR | 6 |
| 2022 | Learning an Isometric Surface Parameterization for Texture Unwrapping
Sagnik Das, Ke Ma 0001, Zhixin Shu, Dimitris Samaras |
ECCV (37) | 3 |
| 2022 | 3D-FM GAN: Towards 3D-Controllable Face Manipulation
Yuchen Liu 0002, Zhixin Shu, Yijun Li 0001, Zhe Lin 0001, Richard Zhang 0001, Sun-Yuan Kung |
ECCV (15) | 2 |
| 2021 | Content-Aware GAN CompressionabstractGenerative adversarial networks (GANs), e.g., StyleGAN2, play a vital role in various image generation and synthesis tasks, yet their notoriously high computational cost hinders their efficient deployment on edge devices. Directly applying generic compression approaches yields poor results on GANs, which motivates a number of recent GAN compression works. While prior works mainly accelerate conditional GANs, e.g., pix2pix and Cycle-GAN, compressing state-of-the-art unconditional GANs has rarely been explored and is more challenging. In this paper, we propose novel approaches for unconditional GAN compression. We first introduce effective channel pruning and knowledge distillation schemes specialized for unconditional GANs. We then propose a novel content-aware method to guide the processes of both pruning and distillation. With content-awareness, we can effectively prune channels that are unimportant to the contents of interest, e.g., human faces, and focus our distillation on these regions, which significantly enhances the distillation quality. On StyleGAN2 and SN-GAN, we achieve a substantial improvement over the state-of-the-art compression method. Notably, we reduce the FLOPs of StyleGAN2 by 11× with visually negligible image quality loss compared to the full-size model. More interestingly, when applied to various image manipulation tasks, our compressed model forms a smoother and better disentangled latent manifold, making it more effective for image editing. Yuchen Liu 0002, Zhixin Shu, Yijun Li 0001, Zhe Lin 0001, Federico Perazzi, Sun-Yuan Kung |
CVPR | 2 |
| 2021 | Pose with style: detail-preserving pose-guided image synthesis with conditional StyleGANabstractWe present an algorithm for re-rendering a person from a single image under arbitrary poses. Existing methods often have difficulties in hallucinating occluded contents photo-realistically while preserving the identity and fine details in the source image. We first learn to inpaint the correspondence field between the body surface texture and the source image with a human body symmetry prior. The inpainted correspondence field allows us to transfer/warp local features extracted from the source to the target view even under large pose changes. Directly mapping the warped local features to an RGB image using a simple CNN decoder often leads to visible artifacts. Thus, we extend the StyleGAN generator so that it takes pose as input (for controlling poses) and introduces a spatially varying modulation for the latent space using the warped local features (for controlling appearances). We show that our method compares favorably against the state-of-the-art algorithms in both quantitative evaluation and visual comparison. Badour AlBahar, Jingwan Lu, Jimei Yang, Zhixin Shu, Eli Shechtman, Jia-Bin Huang 0001 |
ACM Trans. Graph. | 4 |
| 2020 | Self-supervised Deformation Modeling for Facial Expression EditingabstractDeep generative models have recently demonstrated impressive results in photo-realistic facial image synthesis and editing. Existing neural network-based approaches usually only rely on texture generation to edit expressions and largely neglect the motion information. However, facial expressions are inherently the result of muscle movement. In this work, we propose a novel end-to-end network that disentangles the task of facial editing into two steps: a “motionediting” step and a “texture-editing” step. In the “motion-editing” step, we explicitly model facial movement through an image deformation, warping the image into the desired expression. In the “texture-editing” step, we generate the necessary textures, such as teeth and shading effects, for a photorealistic result. Our physically-based task-disentanglement system design allows each step to learn a focused task, and thus need not generate texture to hallucinate motion. Our system is trained in a self-supervised manner, requiring no ground truth deformation annotation. Using Action Units [8] as the representation for facial expression, our method improves the state-of-the-art facial expression editing performance in both qualitative and quantitative evaluations. Shahrukh Athar, Zhixin Shu, Dimitris Samaras |
FG | 2 |
| 2020 | Learning Monocular Face Reconstruction using Multi-View SupervisionabstractWe present a method to reconstruct faces from a single portrait image. While traditional face reconstruction methods fit low-dimensional 3D morphable models to images, we train a deep network to regress depth from a single image directly. We do so by combining supervised losses on synthetic data with indirect supervision on real data using a novel multi-view photo-consistency loss. Furthermore, we regularize the depth estimation using a 3D morphable model (3DMM). We demonstrate that this leads to results that preserve facial features, capture facial geometry that goes beyond 3DMMs, and is also robust to viewpoint conditions. We evaluate our method on various datasets and via ablation studies, and demonstrate that it outperforms previous work significantly. Zhixin Shu, Duygu Ceylan, Kalyan Sunkavalli, Eli Shechtman, Sunil Hadap, Dimitris Samaras |
FG | 1 |
| 2019 | DewarpNet: Single-Image Document Unwarping With Stacked 3D and 2D Regression NetworksabstractCapturing document images with hand-held devices in unstructured environments is a common practice nowadays. However, "casual" photos of documents are usually unsuitable for automatic information extraction, mainly due to physical distortion of the document paper, as well as various camera positions and illumination conditions. In this work, we propose DewarpNet, a deep-learning approach for document image unwarping from a single image. Our insight is that the 3D geometry of the document not only determines the warping of its texture but also causes the illumination effects. Therefore, our novelty resides on the explicit modeling of 3D shape for document paper in an end-to-end pipeline. Also, we contribute the largest and most comprehensive dataset for document image unwarping to date - Doc3D. This dataset features multiple ground-truth annotations, including 3D shape, surface normals, UV map, albedo image, etc. Training with Doc3D, we demonstrate state-of-the-art performance for DewarpNet with extensive qualitative and quantitative evaluations. Our network also significantly improves OCR performance on captured document images, decreasing character error rate by 42% on average. Both the code and the dataset are released. Sagnik Das, Ke Ma 0005, Zhixin Shu, Dimitris Samaras, Roy Shilkrot |
ICCV | 3 |
| 2019 | An Adversarial Neuro-Tensorial Approach for Learning Disentangled RepresentationsabstractSeveral factors contribute to the appearance of an object in a visual scene, including pose, illumination, and deformation, among others. Each factor accounts for a source of variability in the data, while the multiplicative interactions of these factors emulate the entangled variability, giving rise to the rich structure of visual object appearance. Disentangling such unobserved factors from visual data is a challenging task, especially when the data have been captured in uncontrolled recording conditions (also referred to as “in-the-wild”) and label information is not available. In this paper, we propose a pseudo-supervised deep learning method for disentangling multiple latent factors of variation in face images captured in-the-wild. To this end, we propose a deep latent variable model, where the multiplicative interactions of multiple latent factors of variation are explicitly modelled by means of multilinear (tensor) structure. We demonstrate that the proposed approach indeed learns disentangled representations of facial expressions and pose, which can be used in various applications, including face editing, as well as 3D face reconstruction and classification of facial expression, identity and pose. Mengjiao Wang 0002, Zhixin Shu, Shiyang Cheng 0001, Yannis Panagakis, Dimitris Samaras, Stefanos Zafeiriou |
Int. J. Comput. Vis. | 2 |
| 2018 | DocUNet: Document Image Unwarping via a Stacked U-NetabstractCapturing document images is a common way for digitizing and recording physical documents due to the ubiquitousness of mobile cameras. To make text recognition easier, it is often desirable to digitally flatten a document image when the physical document sheet is folded or curved. In this paper, we develop the first learning-based method to achieve this goal. We propose a stacked U-Net [25] with intermediate supervision to directly predict the forward mapping from a distorted image to its rectified version. Because large-scale real-world data with ground truth deformation is difficult to obtain, we create a synthetic dataset with approximately 100 thousand images by warping non-distorted document images. The network is trained on this dataset with various data augmentations to improve its generalization ability. We further create a comprehensive benchmark1 that covers various real-world conditions. We evaluate the proposed model quantitatively and qualitatively on the proposed benchmark, and compare it with previous non-learning-based methods. Ke Ma 0001, Zhixin Shu, Jue Wang 0001, Dimitris Samaras |
CVPR | 2 |
| 2018 | Deforming Autoencoders: Unsupervised Disentangling of Shape and Appearance
Zhixin Shu, Mihir Sahasrabudhe, Riza Alp Güler, Dimitris Samaras, Nikos Paragios, Iasonas Kokkinos |
ECCV (10) | 1 |
| 2018 | Portrait Lighting Transfer Using a Mass Transport ApproachabstractLighting is a critical element of portrait photography. However, good lighting design typically requires complex equipment and significant time and expertise. Our work simplifies this task using a relighting technique that transfers the desired illumination of one portrait onto another. The novelty in our approach to this challenging problem is our formulation of relighting as a mass transport problem. We start from standard color histogram matching that only captures the overall tone of the illumination, and we show how to use the mass-transport formulation to make it dependent on facial geometry. We fit a three-dimensional (3D) morphable face model to the portrait, and for each pixel, we combine the color value with the corresponding 3D position and normal. We then solve a mass-transport problem in this augmented space to generate a color remapping that achieves localized, geometry-aware relighting. Our technique is robust to variations in facial appearance and small errors in face reconstruction. As we demonstrate, this allows our technique to handle a variety of portraits and illumination conditions, including scenarios that are challenging for previous methods. Zhixin Shu, Sunil Hadap, Eli Shechtman, Kalyan Sunkavalli, Sylvain Paris, Dimitris Samaras |
ACM Trans. Graph. | 1 |
| 2017 | Neural Face Editing with Intrinsic Image DisentanglingabstractTraditional face editing methods often require a number of sophisticated and task specific algorithms to be applied one after the other - a process that is tedious, fragile, and computationally intensive. In this paper, we propose an end-to-end generative adversarial network that infers a face-specific disentangled representation of intrinsic face properties, including shape (i.e. normals), albedo, and lighting, and an alpha matte. We show that this network can be trained on “in-the-wild” images by incorporating an in-network physically-based image formation module and appropriate loss functions. Our disentangling latent representation allows for semantically relevant edits, where one aspect offacial appearance can be manipulated while keeping orthogonal properties fixed, and we demonstrate its use for a number offacial editing applications. Zhixin Shu, Ersin Yumer, Sunil Hadap, Kalyan Sunkavalli, Eli Shechtman, Dimitris Samaras |
CVPR | 1 |
| 2017 | Portrait lighting transfer using a mass transport approachabstractLighting is a critical element of portrait photography. However, good lighting design typically requires complex equipment and significant time and expertise. Our work simplifies this task using a relighting technique that transfers the desired illumination of one portrait onto another. The novelty in our approach to this challenging problem is our formulation of relighting as a mass transport problem. We start from standard color histogram matching that only captures the overall tone of the illumination, and we show how to use the mass-transport formulation to make it dependent on facial geometry. We fit a three-dimensional (3D) morphable face model to the portrait, and for each pixel, we combine the color value with the corresponding 3D position and normal. We then solve a mass-transport problem in this augmented space to generate a color remapping that achieves localized, geometry-aware relighting. Our technique is robust to variations in facial appearance and small errors in face reconstruction. As we demonstrate, this allows our technique to handle a variety of portraits and illumination conditions, including scenarios that are challenging for previous methods. Zhixin Shu, Sunil Hadap, Eli Shechtman, Kalyan Sunkavalli, Sylvain Paris, Dimitris Samaras |
ACM Trans. Graph. | 1 |
| 2017 | EyeOpener: Editing Eyes in the WildabstractClosed eyes and look-aways can ruin precious moments captured in photographs. In this article, we present a new framework for automatically editing eyes in photographs. We leverage a user’s personal photo collection to find a “good” set of reference eyes and transfer them onto a target image. Our example-based editing approach is robust and effective for realistic image editing. A fully automatic pipeline for realistic eye editing is challenging due to the unconstrained conditions under which the face appears in a typical photo collection. We use crowd-sourced human evaluations to understand the aspects of the target-reference image pair that will produce the most realistic results. We subsequently train a model that automatically selects the top-ranked reference candidate(s) by narrowing the gap in terms of pose, local contrast, lighting conditions, and even expressions. Finally, we develop a comprehensive pipeline of three-dimensional face estimation, image warping, relighting, image harmonization, automatic segmentation, and image compositing in order to achieve highly believable results. We evaluate the performance of our method via quantitative and crowd-sourced experiments. Zhixin Shu, Eli Shechtman, Dimitris Samaras, Sunil Hadap |
ACM Trans. Graph. | 1 |