VLDB 2026 Research / reviewers in the wild / expert
Xuan Gao 0003
dblp:34/9793-3
· DBLP profile ↗
6ranked-venue papers
3as first author
6since 2021 · last 2026
0000-0001-9396-0761ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GSwap: Realistic Head Swapping With Dynamic Neural Gaussian FieldabstractWe present GSwap, a novel consistent and realistic video head-swapping system empowered by dynamic neural Gaussian portrait priors, which significantly advances the state of the art in face and head replacement. Unlike previous methods that rely primarily on 2D generative models or 3D Morphable Face Models (3DMM), our approach overcomes their inherent limitations, including poor 3D consistency, unnatural facial expressions, and restricted synthesis quality. Moreover, existing techniques struggle with full head-swapping tasks due to insufficient holistic head modeling and ineffective background blending, often resulting in visible artifacts and misalignments. To address these challenges, GSwap introduces an intrinsic 3D Gaussian feature field embedded within a full-body SMPL-X surface, effectively elevating 2D portrait videos into a dynamic neural Gaussian field. This innovation ensures high-fidelity, 3D-consistent portrait rendering while preserving natural head-torso relationships and seamless motion dynamics. To facilitate training, we adapt a pretrained 2D portrait generative model to the source head domain using only a few reference images, enabling efficient domain adaptation. Furthermore, we propose a neural re-rendering strategy that harmoniously integrates the synthesized foreground with the original background, eliminating blending artifacts and enhancing realism. Extensive experiments demonstrate that GSwap surpasses existing methods in multiple aspects, including visual quality, temporal coherence, identity preservation, and 3D consistency. Xuan Gao 0003, Dongyu Liu, Junhui Hou, Juyong Zhang |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2025 | Constructing Diffusion Avatar with Learnable EmbeddingsabstractRecent advances in diffusion models have made significant progress in digital human generation. However, most existing models still struggle to maintain 3D consistency, temporal coherence, and motion accuracy. These limitations primarily stem from two key factors: the limited representation ability of commonly used control signals (e.g., landmarks, depth maps), and the lack of diversity in identity and pose variations within publicly available datasets. In this paper, we construct a powerful head model from both aspects by constructing learnable control signals and enabling the model to adaptively leverage synthetic data. Firstly, we introduce a novel control signal representation that is learnable, dense, expressive, and 3D consistent. Our method embeds learnable Gaussians onto a parametric head surface, which significantly enhances the consistency and expressiveness of diffusion-based head models. Secondly, in terms of data, we synthesize a large-scale dataset covering diverse poses and identities. To reduce the negative impact of artifacts in synthetic data, we introduce real/synthetic embeddings that allow the model to distinguish between real and synthetic samples and learn to utilize them adaptively. Extensive experiments show that our model outperforms existing methods in terms of realism, expressiveness, and 3D consistency. Our code, synthetic datasets, and pre-trained models will be released at https://ustc3dv.github.io/Learn2Control. Xuan Gao 0003, Dongyu Liu, Yuqi Zhou 0004, Juyong Zhang |
SIGGRAPH Asia | 1 |
| 2024 | FlashAvatar: High-Fidelity Head Avatar with Efficient Gaussian EmbeddingabstractWe propose FlashAvatar, a novel and lightweight 3D animatable avatar representation that could reconstruct a digital avatar from a short monocular video sequence in minutes and render high-fidelity photo-realistic images at 300FPS on a consumer-grade GPU. To achieve this, we maintain a uniform 3D Gaussian field embedded in the surface of a parametric face model and learn extra spatial offset to model non-surface regions and subtle facial details. While full use of geometric priors can capture high-frequency facial details and preserve exaggerated expressions, proper initialization can help reduce the number of Gaussians, thus enabling super-fast rendering speed. Extensive experimental results demonstrate that FlashAvatar outperforms existing works regarding visual quality and personalized details and is almost an order of magnitude faster in rendering speed. Project page: https://ustc3dv.github.io/FlashAvatar/ Xuan Gao 0003, Juyong Zhang |
CVPR | 2 |
| 2024 | Portrait Video Editing Empowered by Multimodal Generative Priors
Xuan Gao 0003, Haiyao Xiao, Chenglai Zhong, Shimin Hu 0004, Juyong Zhang |
SIGGRAPH Asia | 1 |
| 2024 | IntrinsicNGP: Intrinsic Coordinate Based Hash Encoding for Human NeRFabstractRecently, many works have been proposed to use the neural radiance field for novel view synthesis of human performers. However, most of these methods require hours of training, making them difficult for practical use. To address this challenging problem, we propose IntrinsicNGP, which can be trained from scratch and achieve high-fidelity results in a few minutes with videos of a human performer. To achieve this goal, we introduce a continuous and optimizable intrinsic coordinate instead of the original explicit euclidean coordinate in the hash encoding module of InstantNGP. With this novel intrinsic coordinate, IntrinsicNGP can aggregate interframe information for dynamic objects using proxy geometry shapes. Moreover, the results trained with the given rough geometry shapes can be further refined with an optimizable offset field based on the intrinsic coordinate. Extensive experimental results on several datasets demonstrate the effectiveness and efficiency of IntrinsicNGP. We also illustrate the ability of our approach to edit the shape of reconstructed objects. Bo Peng 0020, Xuan Gao 0003, Juyong Zhang |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2022 | Reconstructing Personalized Semantic Facial NeRF Models from Monocular VideoabstractWe present a novel semantic model for human head defined with neural radiance field. The 3D-consistent head model consist of a set of disentangled and interpretable bases, and can be driven by low-dimensional expression coefficients. Thanks to the powerful representation ability of neural radiance field, the constructed model can represent complex facial attributes including hair, wearings, which can not be represented by traditional mesh blendshape. To construct the personalized semantic facial model, we propose to define the bases as several multi-level voxel fields. With a short monocular RGB video as input, our method can construct the subject's semantic facial NeRF model with only ten to twenty minutes, and can render a photorealistic human head image in tens of miliseconds with a given expression coefficient and view direction. With this novel representation, we apply it to many tasks like facial retargeting and expression editing. Experimental results demonstrate its strong representation ability and training/inference speed. Demo videos and released code are provided in our project page: https://ustc3dv.github.io/NeRFBlendShape/ Xuan Gao 0003, Chenglai Zhong, Yang Hong 0003, Juyong Zhang |
ACM Trans. Graph. | 1 |