EDBT 2026 Demo / reviewers in the wild / expert
Xuan Wang 0009
dblp:34/4799-9
· DBLP profile ↗
42ranked-venue papers
3as first author
34since 2021 · last 2026
0000-0001-5813-3875ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 35 · 1 first-author · 30 since 2021Artificial intelligence and machine learning · 31 · 2 first-author · 28 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MaTe3D: Mask-Guided Text-based 3D-Aware Portrait Editing
Kangneng Zhou, Daiheng Gao, Xuan Wang 0009, Jie Zhang 0090, Peng Zhang 0080, Xusen Sun, Longhao Zhang, Shiqi Yang 0002, Bang Zhang, Liefeng Bo, Yaxing Wang, Ming-Ming Cheng |
Int. J. Comput. Vis. | 3 |
| 2025 | HERA: Hybrid Explicit Representation for Ultra-Realistic Head AvatarsabstractWe introduce a novel approach to creating ultra-realistic head avatars and rendering them in real time (≥ 30 fps at 2048 × 1334 resolution). First, we propose a hybrid explicit representation that combines the advantages of two primitive based efficient rendering techniques. UV-mapped 3D mesh is utilized to capture sharp and rich textures on smooth surfaces, while 3D Gaussian Splatting is employed to represent complex geometric structures. In the pipeline of modeling an avatar, after tracking parametric models based on captured multi-view RGB videos, our goal is to simultaneously optimize the texture and opacity map of mesh, as well as a set of 3D Gaussian splats localized and rigged onto the mesh facets. Specifically, we perform α-blending on the color and opacity values based on the merged and reordered z-buffer from the rasterization results of mesh and 3DGS. This process involves the mesh and 3DGS adaptively fitting the captured visual information to outline a high-fidelity digital avatar. To avoid artifacts caused by Gaussian splats crossing the mesh facets, we design a stable hybrid depth sorting strategy. Experiments illustrate that our modeled results exceed those of state-of-the-art approaches. Hongrui Cai, Xuan Wang 0009, Jiafei Li, Yanbo Fan, Shenghua Gao, Juyong Zhang |
CVPR | 3 |
| 2025 | AvatarArtist: Open-Domain 4D AvatarizationabstractThis work focuses on open-domain 4D avatarization, with the purpose of creating a 4D avatar from a portrait image in an arbitrary style. We select parametric triplanes as the intermediate 4D representation, and propose a practical training paradigm that takes advantage of both generative adversarial networks (GANs) and diffusion models. Our design stems from the observation that 4D GANs excel at bridging images and triplanes without supervision yet usually face challenges in handling diverse data distributions. A robust 2D diffusion prior emerges as the solution, assisting the GAN in transferring its expertise across various domains. The synergy between these experts permits the construction of a multi-domain image-triplane dataset, which drives the development of a general 4D avatar creator. Extensive experiments suggest that our model, termed AvatarArtist, is capable of producing high-quality 4D avatars with strong robustness to various source image domains. The code, the data, and the models will be made publicly available to facilitate future studies. Xuan Wang 0009, Ziyu Wan, Yue Ma 0016, Jingye Chen, Yanbo Fan, Yujun Shen, Yibing Song, Qifeng Chen 0001 |
CVPR | 2 |
| 2025 | DualTalk: Dual-Speaker Interaction for 3D Talking Head ConversationsabstractIn face-to-face conversations, individuals need to switch between speaking and listening roles seamlessly. Existing 3D talking head generation models focus solely on speaking or listening, neglecting the natural dynamics of interactive conversation, which leads to unnatural interactions and awkward transitions. To address this issue, we propose a new task—multi-round dual-speaker interaction for 3D talking head generation—which requires models to handle and generate both speaking and listening behaviors in continuous conversation. To solve this task, we introduce DualTalk, a novel unified framework that integrates the dynamic behaviors of speakers and listeners to simulate realistic and coherent dialogue interactions. This framework not only synthesizes lifelike talking heads when speaking but also generates continuous and vivid non-verbal feedback when listening, effectively capturing the interplay between the roles. We also create a new dataset featuring 50 hours of multi-round conversations with over 1,000 characters, where participants continuously switch between speaking and listening roles. Extensive experiments demonstrate that our method significantly enhances the naturalness and expressiveness of 3D talking heads in dual-speaker conversations. We recommend watching the supplementary video: https://ziqiaopeng.github.io/dualtalk Ziqiao Peng, Yanbo Fan, Xuan Wang 0009, Hongyan Liu 0002, Jun He 0008, Zhaoxin Fan |
CVPR | 4 |
| 2025 | Diffusion-based Realistic Listening Head Generation via Hybrid Motion ModelingabstractListening head generation aims to synthesize non-verbal responsive listening head videos that naturally react to a certain speaker, for which, both realistic head movements, expressive facial expressions, and high visual qualities are expected. Previous approaches typically follow a two-stage pipeline that first generates intermediate 3D motion signals such as 3DMM coefficients, and then synthesizes the videos by deterministic rendering, suffering from limited motion expressiveness and low visual quality (e.g. 256×256). In this work, we propose a novel listening head generation method that harnesses the generative capabilities of the diffusion model for both motion generation and high-quality rendering. Crucially, we propose an effective hybrid motion modeling module that addresses training difficulties caused by the scarcity of listening head data while preserving the intricate details that may be lost in explicit motion representations. We further develop a tailored control guidance for head pose and facial expression, by integrating their intrinsic motion characteristics. Our method enables high-fidelity video generation with 512 × 512 resolution and delivers vivid listener motion feedback. We conduct comprehensive experiments and obtain superior performance in terms of both visual quality and motion expressiveness compared with existing methods. Yanbo Fan, Xuan Wang 0009, Yu Guo 0006, Fei Wang 0008 |
CVPR | 3 |
| 2025 | 3D Gaussian Head Avatars with Expressive Dynamic Appearances by Compact Tensorial RepresentationsabstractRecent studies have combined 3D Gaussian and 3D Morphable Models (3DMM) to construct high-quality 3D head avatars. In this line of research, existing methods either fail to capture the dynamic textures or incur significant overhead in terms of runtime speed or storage space. To this end, we propose a novel method that addresses all the aforementioned demands. In specific, we introduce an expressive and compact representation that encodes texture-related attributes of the 3D Gaussians in the tensorial format. We store appearance of neutral expression in static tri-planes, and represents dynamic texture details for different expressions using lightweight 1D feature lines, which are then decoded into opacity offset relative to the neutral face. We further propose adaptive truncated opacity penalty and class-balanced sampling to improve generalization across different expressions. Experiments show this design enables accurate face dynamic details capturing while maintains real-time rendering and significantly reduces storage costs, thus broadening the applicability to more scenarios. Xuan Wang 0009, Ran Yi 0002, Yanbo Fan, Jichen Hu, Jingcheng Zhu, Lizhuang Ma |
CVPR | 2 |
| 2025 | DGTalker: Disentangled Generative Latent Space Learning for Audio-Driven Gaussian Talking Heads
Xiaoxi Liang, Yanbo Fan, Qiya Yang, Xuan Wang 0009, Wei Gao 0003, Ge Li 0002 |
ICCV | 4 |
| 2025 | Fine-Grained 3D Gaussian Head Avatars Modeling from Static Captures Via Joint Reconstruction and Registration
Yuan Sun 0003, Xuan Wang 0009, WeiLi Zhang, Yanbo Fan, Yu Guo 0006, Fei Wang 0008 |
ICCV | 2 |
| 2025 | Self-Ensembling Gaussian Splatting for Few-Shot Novel View Synthesisabstract3D Gaussian Splatting (3DGS) has demonstrated remarkable effectiveness in novel view synthesis (NVS). However, 3DGS tends to overfit when trained with sparse views, limiting its generalization to novel viewpoints. In this paper, we address this overfitting issue by introducing Self-Ensembling Gaussian Splatting (SE-GS). We achieve self-ensembling by incorporating an uncertainty-aware perturbation strategy during training. A $\mathbfΔ$-model and a $\mathbfΣ$-model are jointly trained on the available images. The $\mathbfΔ$-model is dynamically perturbed based on rendering uncertainty across training steps, generating diverse perturbed models with negligible computational overhead. Discrepancies between the $\mathbfΣ$-model and these perturbed models are minimized throughout training, forming a robust ensemble of 3DGS models. This ensemble, represented by the $\mathbfΣ$-model, is then used to generate novel-view images during inference. Experimental results on the LLFF, Mip-NeRF360, DTU, and MVImgNet datasets demonstrate that our approach enhances NVS quality under few-shot training conditions, outperforming existing state-of-the-art methods. The code is released at: https://sailor-z.github.io/projects/SEGS.html. Chen Zhao 0025, Xuan Wang 0009, Tong Zhang 0023, Saqib Javed, Mathieu Salzmann |
ICCV | 2 |
| 2025 | Echo: Enhancing Conversational Behavior Generation via Hierarchical Semantic Comprehension with Large Language ModelsabstractConversational behavior generation, being a crucial capability of embodied agents, is a significant factor influencing human-computer interaction. Generating high-quality conversational motions requires not only appropriate audio-motion mapping but also interactive responses to interlocutor behaviors and comprehensive understanding of conversational semantics. Existing methods primarily rely on audio signals and interlocutor motions for main agent motion generation, lacking high-level semantic understanding of the conversational content, leading to moderate quality motions that are not appropriate for the dialogue. To address these limitations, we leverage the powerful semantic understanding capabilities of large language models, to comprehend complex conversational contexts. Inspired by human conversation processes that conversational motions are highly related to both global and local semantic factors, including the conversational context, and the intentions, emotions, and passive or active states of the participants, we propose an agentic system named Echo that analyzes such information. To achieve comprehensive conversational understanding, Echo leverages multiple prompts and test-time recipes to guide large language models in decomposing conversational structures and extracting fine-grained semantic information. Furthermore, we design a hierarchical feature fusion network that systematically integrates from frame-level audio-motion features to sentence-level semantic understanding and finally to conversation-level contextual comprehension, organically combining fine-grained semantic features from large language models with audio and motion characteristics. Experimental results demonstrate that our framework can be effectively integrated with several state-of-the-art motion generation models to enhance their performance in generating high-quality conversational behaviors. Haiwei Xue, Yanbo Fan, Xuan Wang 0009, Zhiyong Wu 0001 |
SIGGRAPH Asia | 3 |
| 2025 | SparseSCIGaussian: Sparse snapshot compressed images 3D Gaussian splatting
Xuan Wang 0009, Nanning Zheng 0001, Caigui Jiang |
Neurocomputing | 2 |
| 2025 | Foodfusion: A Novel Approach for Food Image Composition via Diffusion ModelsabstractFood image composition requires the use of existing dish images and background images to synthesize a natural new image, while diffusion models have made significant advancements in image generation, enabling the construction of end-to-end architectures that yield promising results. However, existing diffusion models face challenges in processing and fusing information from multiple images and lack access to high-quality publicly available datasets, which prevents the application of diffusion models in food image composition. In this paper, we introduce a large-scale, high-quality food image composite dataset,FC22 k, which comprises 22,000 foreground, background, and ground truth ternary image pairs. Additionally, we propose a novel food image composition method,Foodfusion, which leverages the capabilities of the pre-trained diffusion models and incorporates a Fusion Module for processing and integrating foreground and background information. This fused information aligns the foreground features with the background structure by merging the global structural information at the cross-attention layer of the denoising UNet. To further enhance the content and structure of the background, we also integrate a Content-Structure Control Module. Extensive experiments demonstrate the effectiveness and scalability of our proposed method. Chaohua Shi, Xuan Wang 0009, Xule Wang, Mingrui Zhu, Nannan Wang 0001, Xinbo Gao 0001 |
IEEE Trans. Multim. | 2 |
| 2024 | A Pre-convolved Representation for Plug-and-Play Neural Illumination FieldsabstractRecent advances in implicit neural representation have demonstrated the ability to recover detailed geometry and material from multi-view images. However, the use of simplified lighting models such as environment maps to represent non-distant illumination, or using a network to fit indirect light modeling without a solid basis, can lead to an undesirable decomposition between lighting and material. To address this, we propose a fully differentiable framework named Neural Illumination Fields (NeIF) that uses radiance fields as a lighting model to handle complex lighting in a physically based way. Together with integral lobe encoding for roughness-adaptive specular lobe and leveraging the pre-convolved background for accurate decomposition, the proposed method represents a significant step towards integrating physically based rendering into the NeRF representation. The experiments demonstrate the superior performance of novel-view rendering compared to previous works, and the capability to re-render objects under arbitrary NeRF-style environments opens up exciting possibilities for bridging the gap between virtual and real-world scenes. Yiyu Zhuang, Qi Zhang 0029, Xuan Wang 0009, Hao Zhu 0004, Xiaoyu Li 0002, Ying Shan, Xun Cao |
AAAI | 3 |
| 2024 | Real-Time 3D-Aware Portrait Editing from a Single Image
Qingyan Bai, Zifan Shi, Yinghao Xu 0001, Hao Ouyang, Qiuyu Wang, Ceyuan Yang, Xuan Wang 0009, Gordon Wetzstein, Yujun Shen, Qifeng Chen 0001 |
ECCV (51) | 7 |
| 2023 | High-fidelity Facial Avatar Reconstruction from Monocular Video with Generative PriorsabstractHigh-fidelity facial avatar reconstruction from a monocular video is a significant research problem in computer graphics and computer vision. Recently, Neural Radiance Field (NeRF) has shown impressive novel view rendering results and has been considered for facial avatar reconstruction. However, the complex facial dynamics and missing 3D information in monocular videos raise significant challenges for faithful facial reconstruction. In this work, we propose a new method for NeRF-based facial avatar reconstruction that utilizes 3D-aware generative prior. Different from existing works that depend on a conditional deformation field for dynamic modeling, we propose to learn a personalized generative prior, which is formulated as a local and low dimensional subspace in the latent space of 3D-GAN. We propose an efficient method to construct the personalized generative prior based on a small set of facial images of a given individual. After learning, it allows for photo-realistic rendering with novel views, and the face reenactment can be realized by performing navigation in the latent space. Our proposed method is applicable for different driven signals, including RGB images, 3DMM coefficients, and audio. Compared with existing works, we obtain superior novel view synthesis results and faithfully face reenactment performance. The code is available here https://github.com/bbaaii/HFA-GP. Yunpeng Bai, Yanbo Fan, Xuan Wang 0009, Yong Zhang 0034, Jingxiang Sun, Chun Yuan 0003, Ying Shan |
CVPR | 3 |
| 2023 | Local-to-Global Registration for Bundle-Adjusting Neural Radiance FieldsabstractNeural Radiance Fields (NeRF) have achieved photorealistic novel views synthesis; however, the requirement of accurate camera poses limits its application. Despite analysis-by-synthesis extensions for jointly learning neural3D representations and registering camera frames exist, they are susceptible to suboptimal solutions if poorly initialized. We propose L2G-NeRF, a Local-to-Global registration method for bundle-adjusting Neural Radiance Fields: first, a pixel-wise flexible alignment, followed by a framewise constrained parametric alignment. Pixel-wise local alignment is learned in an unsupervised way via a deep network which optimizes photometric reconstruction errors. framewise global alignment is performed using differentiable parameter estimation solvers on the pixel-wise correspondences to find a global transformation. Experiments on synthetic and real-world data show that our method outperforms the current state-of-the-art in terms of high-fidelity reconstruction and resolving large camera pose misalignment. Our module is an easy-to-use plugin that can be applied to NeRF variants and other neural field applications. The Code and supplementary materials are available at https://rover-xingyu.github.io/L2G-NeRF/. Xuan Wang 0009, Qi Zhang 0029, Yu Guo 0006, Ying Shan, Fei Wang 0008 |
CVPR | 3 |
| 2023 | UV Volumes for Real-time Rendering of Editable Free-view Human PerformanceabstractNeural volume rendering enables photo-realistic renderings of a human performer in free-view, a critical task in immersive VR/AR applications. But the practice is severely limited by high computational costs in the rendering process. To solve this problem, we propose the UV Volumes, a new approach that can render an editable free-view video of a human performer in real-time. It separates the high-frequency (i.e., non-smooth) human appearance from the 3D volume, and encodes them into 2D neural texture stacks (NTS). The smooth UV volumes allow much smaller and shallower neural networks to obtain densities and texture coordinates in 3D while capturing detailed appearance in 2D NTS. For editability, the mapping between the parameterized human model and the smooth texture coordinates allows us a better generalization on novel poses and shapes. Furthermore, the use of NTS enables interesting applications, e.g., retexturing. Extensive experiments on CMU Panoptic, ZJU Mocap, and H36M datasets show that our model can render$960\times 540$images in 30FPS on average with comparable photo-realism to state-of-the-art methods. The project and supplementary materials are available at https://fanegg.github.io/UV-Volumes. Xuan Wang 0009, Qi Zhang 0029, Xiaoyu Li 0002, Yu Guo 0006, Jue Wang 0001, Fei Wang 0008 |
CVPR | 2 |
| 2023 | Local Implicit Ray Function for Generalizable Radiance Field RepresentationabstractWe propose LIRF (Local Implicit Ray Function), a generalizable neural rendering approach for novel view rendering. Current generalizable neural radiance fields (NeRF) methods sample a scene with a single ray per pixel and may therefore render blurred or aliased views when the input views and rendered views capture scene content with different resolutions. To solve this problem, we propose LIRF to aggregate the information from conical frustums to construct a ray. Given 3D positions within conical frustums, LIRF takes 3D coordinates and the features of conical frustums as inputs and predicts a local volumetric radiance field. Since the coordinates are continuous, LIRF renders high-quality novel views at a continuously-valued scale via volume rendering. Besides, we predict the visible weights for each input view via transformer-based feature matching to improve the performance in occluded areas. Experimental results on real-world scenes validate that our method outperforms state-of-the-art methods on novel view rendering of unseen scenes at arbitrary scales. Xin Huang 0021, Qi Zhang 0029, Xiaoyu Li 0002, Xuan Wang 0009, Qing Wang 0006 |
CVPR | 5 |
| 2023 | High-Fidelity Clothed Avatar Reconstruction from a Single ImageabstractThis paper presents a framework for efficient 3D clothed avatar reconstruction. By combining the advantages of the high accuracy of optimization-based methods and the efficiency of learning-based methods, we propose a coarse-to-fine way to realize a high-fidelity clothed avatar reconstruction (CAR) from a single image. At the first stage, we use an implicit model to learn the general shape in the canonical space of a person in a learning-based way, and at the second stage, we refine the surface detail by estimating the non-rigid deformation in the posed space in an optimization way. A hyper-network is utilized to generate a good initialization so that the convergence of the optimization process is greatly accelerated. Extensive experiments on various datasets show that the proposed CAR successfully produces high-fidelity avatars for arbitrarily clothed humans in real scenes. The codes will be released in https://github.com/TingtingLiao/CAR. Tingting Liao, Yuliang Xiu, Hongwei Yi, Xudong Liu 0006, Guo-Jun Qi, Yong Zhang 0034, Xuan Wang 0009, Xiangyu Zhu 0001, Zhen Lei 0001 |
CVPR | 8 |
| 2023 | Next3D: Generative Neural Texture Rasterization for 3D-Aware Head Avatarsabstract3D-aware generative adversarial networks (GANs) synthesize high-fidelity and multi-view-consistent facial images using only collections of single-view 2D imagery. Towards fine-grained control over facial attributes, recent efforts incorporate 3D Morphable Face Model (3DMM) to describe deformation in generative radiance fields either explicitly or implicitly. Explicit methods provide fine-grained expression control but cannot handle topological changes caused by hair and accessories, while implicit ones can model varied topologies but have limited generalization caused by the unconstrained deformation fields. We propose a novel 3D GAN framework for unsupervised learning of generative, high-quality and 3D-consistent facial avatars from unstructured 2D images. To achieve both deformation accuracy and topological flexibility, we propose a 3D representation called Generative Texture-Rasterized Tri-planes. The proposed representation learns Generative Neural Textures on top of parametric mesh templates and then projects them into three orthogonal-viewed feature planes through rasterization, forming a tri-plane feature representation for volume rendering. In this way, we combine both fine-grained expression control of mesh-guided explicit deformation and the flexibility of implicit volumetric representation. We further propose specific modules for modeling mouth interior which is not taken into account by 3DMM. Our method demonstrates state-of-the-art 3D-aware synthesis quality and animation ability through extensive experiments. Furthermore, serving as 3D prior, our animatable 3D representation boosts multiple applications including one-shot facial avatars and 3D-aware stylization. Project page: https://mrtornado24.github.io/Next3D/. Code: https://github.com/MrTornado24/Next3D. Jingxiang Sun, Xuan Wang 0009, Lizhen Wang 0002, Xiaoyu Li 0002, Yong Zhang 0034, Hongwen Zhang 0001, Yebin Liu |
CVPR | 2 |
| 2023 | 3D GAN Inversion with Facial Symmetry PriorabstractRecently, a surge of high-quality 3D-aware GANs have been proposed, which leverage the generative power of neural rendering. It is natural to associate 3D GANs with GAN inversion methods to project a real image into the generator's latent space, allowing free-view consistent synthesis and editing, referred as 3D GAN inversion. Although with the facial prior preserved in pre-trained 3D GANs, reconstructing a 3D portrait with only one monocular image is still an ill-pose problem. The straightforward application of 2D GAN inversion methods focuses on texture similarity only while ignoring the correctness of 3D geometry shapes. It may raise geometry collapse effects, especially when reconstructing a side face under an extreme pose. Besides, the synthetic results in novel views are prone to be blurry. In this work, we propose a novel method to promote 3D GAN inversion by introducing facial symmetry prior. We design a pipeline and constraints to make full use of the pseudo auxiliary view obtained via image flipping, which helps obtain a view-consistent and well-structured geometry shape during the inversion process. To enhance texture fidelity in unobserved viewpoints, pseudo labels from depth-guided 3D warping can provide extra supervision. We design constraints to filter out conflict areas for optimization in asymmetric situations. Comprehensive quantitative and qualitative evaluations on image reconstruction and editing demonstrate the superiority of our method. Yong Zhang 0034, Xuan Wang 0009, Tengfei Wang 0002, Xiaoyu Li 0002, Yuan Gong 0002, Yanbo Fan, Xiaodong Cun, Ying Shan, A. Cengiz Öztireli, Yujiu Yang 0001 |
CVPR | 3 |
| 2023 | SadTalker: Learning Realistic 3D Motion Coefficients for Stylized Audio-Driven Single Image Talking Face AnimationabstractGenerating talking head videos through a face image and a piece of speech audio still contains many challenges. i.e., unnatural head movement, distorted expression, and identity modification. We argue that these issues are mainly caused by learning from the coupled 2D motion fields. On the other hand, explicitly using 3D information also suffers problems of stiff expression and incoherent video. We present SadTalker, which generates 3D motion coefficients (head pose, expression) of the 3DMM from audio and implicitly modulates a novel 3D-aware face render for talking head generation. To learn the realistic motion coefficients, we explicitly model the connections between audio and different types of motion coefficients individually. Precisely, we present ExpNet to learn the accurate facial expression from audio by distilling both coefficients and 3D-rendered faces. As for the head pose, we design PoseVAE via a conditional VAE to synthesize head motion in different styles. Finally, the generated 3D motion coefficients are mapped to the unsupervised 3D keypoints space of the proposed face render to synthesize the final video. We conducted extensive experiments to demonstrate the superiority of our method in terms of motion and video quality.11The code and demo videos are available at https://sadtalker.github.io. Xiaodong Cun, Xuan Wang 0009, Yong Zhang 0034, Xi Shen 0001, Yu Guo 0006, Ying Shan, Fei Wang 0008 |
CVPR | 3 |
| 2023 | ToonTalker: Cross-Domain Face ReenactmentabstractWe target cross-domain face reenactment in this paper, i.e., driving a cartoon image with the video of a real person and vice versa. Recently, many works have focused on one-shot talking face generation to drive a portrait with a real video, i.e., within-domain reenactment. Straightforwardly applying those methods to cross-domain animation will cause inaccurate expression transfer, blur effects, and even apparent artifacts due to the domain shift between cartoon and real faces. Only a few works attempt to settle cross-domain face reenactment. The most related work AnimeCeleb [13] requires constructing a dataset with pose vector and cartoon image pairs by animating 3D characters, which makes it inapplicable anymore if no paired data is available. In this paper, we propose a novel method for cross-domain reenactment without paired data. Specifically, we propose a transformer-based framework to align the motions from different domains into a common latent space where motion transfer is conducted via latent code addition. Two domain-specific motion encoders and two learnable motion base memories are used to capture domain properties. A source query transformer and a driving one are exploited to project domain-specific motion to the canonical space. The edited motion is projected back to the domain of the source with a transformer. Moreover, since no paired data is provided, we propose a novel cross-domain training scheme using data from two domains with the designed analogy constraint. Besides, we contribute a cartoon dataset in Disney style. Extensive evaluations demonstrate the superiority of our method over competing methods. Yuan Gong 0002, Yong Zhang 0034, Xiaodong Cun, Yanbo Fan, Xuan Wang 0009, Baoyuan Wu, Yujiu Yang 0001 |
ICCV | 6 |
| 2023 | Robust Pose Transfer With Dynamic Details Using Neural Video RenderingabstractPose transfer of human videos aims to generate a high-fidelity video of a target person imitating actions of a source person. A few studies have made great progress either through image translation with deep latent features or neural rendering with explicit 3D features. However, both of them rely on large amounts of training data to generate realistic results, and the performance degrades on more accessible Internet videos due to insufficient training frames. In this paper, we demonstrate that the dynamic details can be preserved even when trained from short monocular videos. Overall, we propose a neural video rendering framework coupled with an image-translation-based dynamic details generation network (D$^{2}$G-Net), which fully utilizes both the stability of explicit 3D features and the capacity of learning components. To be specific, a novel hybrid texture representation is presented to encode both the static and pose-varying appearance characteristics, which is then mapped to the image space and rendered as a detail-rich frame in the neural rendering stage. Through extensive comparisons, we demonstrate that our neural human video renderer is capable of achieving both clearer dynamic details and more robust performance even on accessible short videos with only 2 k$\sim$4 k frames, as illustrated in Fig. 1. Yang-Tian Sun, Hao-Zhi Huang 0001, Xuan Wang 0009, Yukun Lai, Wei Liu 0005, Lin Gao 0004 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Hallucinated Neural Radiance Fields in the WildabstractNeural Radiance Fields (NeRF) has recently gained popularity for its impressive novel view synthesis ability. This paper studies the problem of hallucinated NeRF: i.e., recovering a realistic NeRF at a different time of day from a group of tourism images. Existing solutions adopt NeRF with a controllable appearance embedding to render novel views under various conditions, but they cannot render view-consistent images with an unseen appearance. To solve this problem, we present an end-to-end framework for constructing a hallucinated NeRF, dubbed as Ha-NeRF. Specifically, we propose an appearance hallucination module to handle time-varying appearances and transfer them to novel views. Considering the complex occlusions of tourism images, we introduce an anti-occlusion module to decompose the static subjects for visibility accurately. Experimental results on synthetic data and real tourism photo collections demonstrate that our method can hallucinate the desired appearances and render occlusion-free images from different views. The project and supplementary materials are available at https://rover-xingyu.github.io/Ha-NeRF/. Qi Zhang 0029, Xiaoyu Li 0002, Xuan Wang 0009, Jue Wang 0001 |
CVPR | 6 |
| 2022 | HDR-NeRF: High Dynamic Range Neural Radiance FieldsabstractWe present High Dynamic Range Neural Radiance Fields (HDR-NeRF) to recover an HDR radiance field from a set of low dynamic range (LDR) views with different exposures. Using the HDR-NeRF, we are able to generate both novel HDR views and novel LDR views under different exposures. The key to our method is to model the simplified physical imaging process, which dictates that the radiance of a scene point transforms to a pixel value in the LDR image with two implicit functions: a radiance field and a tone mapper. The radiance field encodes the scene radiance (values vary from 0 to$+\infty$), which outputs the density and radiance of a ray by giving corresponding ray origin and ray direction. The tone mapper models the mapping process that a ray hitting on the camera sensor becomes a pixel value. The color of the ray is predicted by feeding the radiance and the corresponding exposure time into the tone mapper. We use the classic volume rendering technique to project the output radiance, colors and densities into HDR and LDR images, while only the input LDR images are used as the supervision. We collect a new forward-facing HDR dataset to evaluate the proposed method. Experimental results on synthetic and real-world scenes validate that our method can not only accurately control the exposures of synthesized views but also render views with a high dynamic range. Xin Huang 0021, Qi Zhang 0029, Hongdong Li, Xuan Wang 0009, Qing Wang 0006 |
CVPR | 5 |
| 2022 | Deblur-NeRF: Neural Radiance Fields from Blurry ImagesabstractNeural Radiance Field (NeRF) has gained considerable attention recently for 3D scene reconstruction and novel view synthesis due to its remarkable synthesis quality. However, image blurriness caused by defocus or motion, which often occurs when capturing scenes in the wild, significantly degrades its reconstruction quality. To address this problem, We propose Deblur-NeRF, the first method that can recover a sharp NeRF from blurry input. We adopt an analysis-by-synthesis approach that reconstructs blurry views by simulating the blurring process, thus making NeRF robust to blurry inputs. The core of this simulation is a novel Deformable Sparse Kernel (DSK) module that models spatially-varying blur kernels by deforming a canonical sparse kernel at each spatial location. The ray origin of each kernel point is Jointly optimized, inspired by the physical blurring process. This module is parameterized as an MLP that has the ability to be generalized to various blur types. Jointly optimizing the NeRF and the DSK module allows us to restore a sharp NeRF. We demonstrate that our method can be used on both camera motion blur and defocus blur: the two most common types of blur in real scenes. Evaluation results on both synthetic and real-world data show that our method outperforms several baselines. The synthetic and real datasets along with the source code is publicly available at https://limacv.github.io/deblurNeRF/. Xiaoyu Li 0002, Jing Liao 0001, Qi Zhang 0029, Xuan Wang 0009, Jue Wang 0001, Pedro V. Sander |
CVPR | 5 |
| 2022 | FENeRF: Face Editing in Neural Radiance FieldsabstractPrevious portrait image generation methods roughly fall into two categories: 2D GANs and 3D-aware GANs. 2D GANs can generate high fidelity portraits but with low view consistency. 3D-aware GAN methods can maintain view consistency but their generated images are not locally editable. To overcome these limitations, we propose FENeRF, a 3D-aware generator that can produce view-consistent and locally-editable portrait images. Our method uses two decoupled latent codes to generate corresponding facial semantics and texture in a spatial-aligned 3D volume with shared geometry. Benefiting from such underlying 3D representation, FENeRF can Jointly render the boundary-aligned image and semantic mask and use the semantic mask to edit the 3D volume via GAN inversion. We further show such 3D representation can be learned from widely available monocular image and semantic mask pairs. Moreover, we reveal that Joint learning semantics and texture helps to generate finer geometry. Our experiments demonstrate that FENeRF outperforms state-of-the-art methods in various face editing tasks. Code is available at https://github.com/MrTornado24/FENeRF. Jingxiang Sun, Xuan Wang 0009, Yong Zhang 0034, Xiaoyu Li 0002, Qi Zhang 0029, Yebin Liu, Jue Wang 0001 |
CVPR | 2 |
| 2022 | StyleHEAT: One-Shot High-Resolution Editable Talking Face Generation via Pre-trained StyleGAN
Yong Zhang 0034, Xiaodong Cun, Mingdeng Cao, Yanbo Fan, Xuan Wang 0009, Qingyan Bai, Baoyuan Wu, Jue Wang 0001, Yujiu Yang 0001 |
ECCV (17) | 6 |
| 2022 | VideoReTalking: Audio-based Lip Synchronization for Talking Head Video Editing In the WildabstractWe present VideoReTalking, a new system to edit the faces of a real-world talking head video according to input audio, producing a high-quality and lip-syncing output video even with a different emotion. Our system disentangles this objective into three sequential tasks: (1) face video generation with a canonical expression; (2) audio-driven lip-sync; and (3) face enhancement for improving photo-realism. Given a talking-head video, we first modify the expression of each frame according to the same expression template using the expression editing network, resulting in a video with the canonical expression. This video, together with the given audio, is then fed into the lip-sync network to generate a lip-syncing video. Finally, we improve the photo-realism of the synthesized faces through an identity-aware face enhancement network and post-processing. We use learning-based approaches for all three steps and all our modules can be tackled in a sequential pipeline without any user intervention. Furthermore, our system is a generic approach that does not need to be retrained to a specific person. Evaluations on two widely-used datasets and in-the-wild examples demonstrate the superiority of our framework over other state-of-the-art methods in terms of lip-sync accuracy and visual quality. Xiaodong Cun, Yong Zhang 0034, Menghan Xia, Mingrui Zhu, Xuan Wang 0009, Jue Wang 0001, Nannan Wang 0001 |
SIGGRAPH Asia | 7 |
| 2022 | Neural Parameterization for Dynamic Human Head EditingabstractImplicit radiance functions emerged as a powerful scene representation for reconstructing and rendering photo-realistic views of a 3D scene. These representations, however, suffer from poor editability. On the other hand, explicit representations such as polygonal meshes allow easy editing but are not as suitable for reconstructing accurate details in dynamic human heads, such as fine facial features, hair, teeth, and eyes. In this work, we present Neural Parameterization (NeP), a hybrid representation that provides the advantages of both implicit and explicit methods. NeP is capable of photo-realistic rendering while allowing fine-grained editing of the scene geometry and appearance. We first disentangle the geometry and appearance by parameterizing the 3D geometry into 2D texture space. We enable geometric editability by introducing an explicit linear deformation blending layer. The deformation is controlled by a set of sparse key points, which can be explicitly and intuitively displaced to edit the geometry. For appearance, we develop a hybrid 2D texture consisting of an explicit texture map for easy editing and implicit view and time-dependent residuals to model temporal and view variations. We compare our method to several reconstruction and editing baselines. The results show that the NeP achieves almost the same level of rendering accuracy while maintaining high editability. Xiaoyu Li 0002, Jing Liao 0001, Xuan Wang 0009, Qi Zhang 0029, Jue Wang 0001, Pedro V. Sander |
ACM Trans. Graph. | 4 |
| 2022 | IDE-3D: Interactive Disentangled Editing for High-Resolution 3D-Aware Portrait SynthesisabstractExisting 3D-aware facial generation methods face a dilemma in quality versus editability: they either generate editable results in low resolution, or high-quality ones with no editing flexibility. In this work, we propose a new approach that brings the best of both worlds together. Our system consists of three major components: (1) a 3D-semantics-aware generative model that produces view-consistent, disentangled face images and semantic masks; (2) a hybrid GAN inversion approach that initializes the latent codes from the semantic and texture encoder, and further optimizes them for faithful reconstruction; and (3) a canonical editor that enables efficient manipulation of semantic masks in canonical view and produces high-quality editing results. Our approach is competent for many applications, e.g. free-view face drawing, editing and style control. Both quantitative and qualitative results show that our method reaches the state-of-the-art in terms of photorealism, faithfulness and efficiency. Jingxiang Sun, Xuan Wang 0009, Yichun Shi, Lizhen Wang 0002, Jue Wang 0001, Yebin Liu |
ACM Trans. Graph. | 2 |
| 2021 | Monocular 3D multi-person pose estimation via predicting factorized correction factors
Yu Guo 0006, Lichen Ma, Zhi Li 0055, Xuan Wang 0009, Fei Wang 0008 |
Comput. Vis. Image Underst. | 4 |
| 2021 | UniFaceGAN: A Unified Framework for Temporally Consistent Facial Video EditingabstractRecent research has witnessed advances in facial image editing tasks including face swapping and face reenactment. However, these methods are confined to dealing with one specific task at a time. In addition, for video facial editing, previous methods either simply apply transformations frame by frame or utilize multiple frames in a concatenated or iterative fashion, which leads to noticeable visual flickers. In this paper, we propose a unified temporally consistent facial video editing framework termed UniFaceGAN. Based on a 3D reconstruction model and a simple yet efficient dynamic training sample selection mechanism, our framework is designed to handle face swapping and face reenactment simultaneously. To enforce the temporal consistency, a novel 3D temporal loss constraint is introduced based on the barycentric coordinate interpolation. Besides, we propose a region-aware conditional normalization layer to replace the traditional AdaIN or SPADE to synthesize more context-harmonious results. Compared with the state-of-the-art facial image editing methods, our framework generates video portraits that are more photo-realistic and temporally smooth. Meng Cao 0002, Hao-Zhi Huang 0001, Hao Wang 0050, Xuan Wang 0009, Li Shen 0008, Linchao Bao, Zhifeng Li 0001, Jiebo Luo 0001 |
IEEE Trans. Image Process. | 4 |
| 2019 | On Boosting Single-Frame 3D Human Pose Estimation via Monocular VideosabstractThe premise of training an accurate 3D human pose estimation network is the possession of huge amount of richly annotated training data. Nonetheless, manually obtaining rich and accurate annotations is, even not impossible, tedious and slow. In this paper, we propose to exploit monocular videos to complement the training dataset for the single-image 3D human pose estimation tasks. At the beginning, a baseline model is trained with a small set of annotations. By fixing some reliable estimations produced by the resulting model, our method automatically collects the annotations across the entire video as solving the 3D trajectory completion problem. Then, the baseline model is further trained with the collected annotations to learn the new poses. We evaluate our method on the broadly-adopted Human3.6M and MPI-INF-3DHP datasets. As illustrated in experiments, given only a small set of annotations, our method successfully makes the model to learn new poses from unlabelled monocular videos, promoting the accuracies of the baseline model by about 10%. By contrast with previous approaches, our method does not rely on either multi-view imagery or any explicit 2D keypoint annotations. Zhi Li 0055, Xuan Wang 0009, Fei Wang 0008, Peilin Jiang |
ICCV | 2 |
| 2019 | Stacked Mixed-Scale Networks for Human Pose Estimation
Xuan Wang 0009, Zhi Li 0055, Peilin Jiang, Fei Wang 0008 |
PRICAI (1) | 1 |
| 2018 | Robust real-time visual object tracking via multi-scale fully convolutional Siamese networks
Longchao Yang, Peilin Jiang, Fei Wang 0008, Xuan Wang 0009 |
Multim. Tools Appl. | 4 |
| 2017 | Recovering complex non-rigid 3D structures from monocular images by union of nonlinear subspacesabstractNon-rigid structure from motion (NRSfM) is a well-known challenging task due to its inherent ambiguities. Most existing approaches rely on kinds of low-rank linear subspaces assumption to make the problem well-constrained. In this paper, we make two contributions. First, we empirically present that relying on the assumption, 3D shapes lie on a union of non-linear subspaces, can better model the complex non-rigid motion than its linear counterparts. Second, we introduce the nonlinear low-rank representation as a regularizer to the objective function for NRSfM and show that it can be solved by alternating direction multiplier method (ADMM). Our experiments demonstrate that our method yields more accurate reconstruction and more reasonable clustering results than state-of-the-art methods, on CMU MoCap and UMPM datasets. Fei Wang 0008, Xuan Wang 0009 |
ICIP | 3 |
| 2017 | Region-based fully convolutional siamese networks for robust real-time visual trackingabstractPartial occlusions and deformations in visual object tracking are still very challenging. Existing Convolutional Neural Networks (CNNs) trackers either fail to handle these issues or can just run in low speed. In this paper, we present a real-time tracker which is robust to occlusions and deformations based on a Region-based, Fully Convolutional Siamese Network (R-FCSN). In the proposed R-FCSN, the information of regions is extracted separately by the proposition of position-sensitive score maps. Combining these score maps via adaptive weights leads to accurate location of the target on a new frame. The experiments illustrate that our method outperforms state-of-the-art approaches, and can handle the cases of object deformation and occlusion at about 51 FPS. Longchao Yang, Peilin Jiang, Fei Wang 0008, Xuan Wang 0009 |
ICIP | 4 |
| 2016 | Template-Free 3D Reconstruction of Poorly-Textured Nonrigid Surfaces
Xuan Wang 0009, Mathieu Salzmann, Fei Wang 0008, Jizhong Zhao |
ECCV (7) | 1 |
| 2014 | Monocular 3D Shape Recovery of Inextensibility Deformable Surface by Using DE-Based Niching Algorithm with Partial Reinitialization
Xuan Wang 0009, Fei Wang 0008 |
ICIC (2) | 1 |
| 2014 | Clustering-Based Latent Variable Models for Monocular Non-rigid 3D Shape Recovery
Fei Wang 0008, Xuan Wang 0009 |
ICIC (2) | 4 |