Hao Xu 0049

dblp:43/6008-49 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
11since 2021 · last 2026
0000-0001-5690-367XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 3 first-author · 11 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Lightmap Compression with Color-Coherent UV Clustering and Cascade Texture Optimization
abstract
Abstract To address the storage overhead of lightmaps and the limitations of existing compression techniques, we propose a novel UV‐space compression framework based on per‐triangle processing. By mapping triangles to a standardized domain, we cluster and repack color‐coherent regions into a compact atlas, generating a cascade texture refined via differentiable rendering. Experimental results show an average storage reduction of 83% with approximately 10 dB higher PSNR than existing methods. Our approach is the first dedicated lightmap compression framework compatible with standard block‐based formats, offering an effective solution for memory‐efficient 3D asset delivery.
Dehan Chen, Hongyu Huang 0001, Yuzhe Luo, Hao Xu 0049, Yuqing Zhang 0005, Sipeng Yang, Xifeng Gao, Heng Cai, Xiaogang Jin 0001
Comput. Graph. Forum4
2025 LegoACE: Autoregressive Construction Engine for Expressive LEGO® Assemblies
abstract
Automated LEGO® design is challenging due to the extensive variety of LEGO® brick types and the necessity of constructing semantically meaningful models from individually meaningless components. Current automatic LEGO® generation methods face two key challenges: i) They typically rely on explicit modeling of brick connectivity to ensure structural validity. However, this requires extensive manual annotation, which is labor-intensive as the variety of LEGO® primitives increases. This limits training data diversity, restricting the variety of LEGO® bricks that can be effectively utilized. ii) To facilitate learning within neural networks, current methods often employ either volume or text-based descriptions to represent LEGO® models. However, volumetric representations are computationally expensive and hamper large-scale generative training, while text-based approaches rely on large language models and dedicated text-to-brick mapping rules, introducing a semantic gap between language tokens and 3D brick structures.
Hao Xu 0049, Yuqing Zhang 0005, Xinyang Zheng, Xiangjun Tang, Yunhan Yang, Ding Liang, Yingtian Liu, Yan-Pei Cao 0001, Xiaogang Jin 0001
SIGGRAPH Asia1
2025 OmniPart: Part-Aware 3D Generation with Semantic Decoupling and Structural Cohesion
abstract
The creation of 3D assets with explicit, editable part structures is crucial for advancing interactive applications, yet most generative methods produce only monolithic shapes, limiting their utility. We introduce OmniPart, a novel framework for part-aware 3D object generation designed to achieve high semantic decoupling among components while maintaining robust structural cohesion. OmniPart uniquely decouples this complex task into two synergistic stages: (1) an autoregressive structure planning module generates a controllable, variable-length sequence of 3D part bounding boxes, critically guided by flexible 2D part masks that allow for intuitive control over part decomposition without requiring direct correspondences or semantic labels; and (2) a spatially-conditioned rectified flow model, efficiently adapted from a pre-trained holistic 3D generator, synthesizes all 3D parts simultaneously and consistently within the planned layout. Our approach supports user-defined part granularity, precise localization, and enables diverse downstream applications. Extensive experiments demonstrate that OmniPart achieves state-of-the-art performance, paving the way for more interpretable, editable, and versatile 3D content.
Yunhan Yang, Yufan Zhou 0004, Zixin Zou, Ying-Tian Liu, Hao Xu 0049, Ding Liang, Yan-Pei Cao 0001, Xihui Liu
SIGGRAPH Asia7
2025 GSFaceMorpher: High-Fidelity 3D Face Morphing via Gaussian Splatting
abstract
ABSTRACT High‐fidelity 3D face morphing aims to achieve seamless transitions between realistic 3D facial representations of different identities. Although 3D Gaussian Splatting (3DGS) excels in high‐quality rendering, its application to morphing is hindered by the lack of Gaussian primitive correspondence and variations in primitive quantities. To address this, we propose GSFaceMorpher, which is a novel framework for high‐fidelity 3D face morphing based on 3DGS. Our method constructs an auxiliary model that bridges the source and target face models by aligning the geometry through Radial Basis Function (RBF) warping and optimizing the appearance in the image space. This auxiliary model enables smooth parameter interpolation, whereas a diffusion‐based refinement step enhances critical facial details through attention replacement from the reference faces. Experiments demonstrate that our method produces visually coherent and high‐fidelity morphing sequences, significantly outperforming NeRF‐based baselines in terms of both quantitative metrics and user preferences. Our work establishes a new benchmark for high‐fidelity 3D face morphing with applications in visual effects, animation, and immersive experiences.
Xiwen Shi, Hao Xu 0049, Ziyi Yang 0008, Xiaogang Jin 0001
Comput. Animat. Virtual Worlds4
2025 3DPortraitGAN: Learning One-Quarter Headshot 3D GANs From a Single-View Portrait Dataset With Diverse Body Poses
abstract
3D-aware face generators are typically trained on 2D real-life face image datasets that primarily consist of near-frontal face data. Due to data limitations, these generators cannot generateone-quarter headshot3D portraits with head, neck, and shoulder geometry, which is crucial for applications like talking heads. Two reasons account for this issue: First, existing facial recognition methods struggle with extracting facial data captured from large camera angles or back views. Second, it is challenging to learn a distribution of 3D portraits covering the one-quarter headshot region from single-view data due to significant geometric deformation caused by diverse body poses. To this end, we first create the dataset360°-Portrait-HQ(360°PHQfor short) which consists of high-quality single-view real portraits annotated with a variety of camera parameters (the yaw angles span the entire 360° range) and body poses. We then propose3DPortraitGAN, the first 3D-aware one-quarter headshot portrait generator that learns a canonical 3D avatar distribution from the360°PHQ dataset with body pose self-learning. Our model can generate view-consistent portrait images from all camera angles with a canonical one-quarter headshot 3D representation. Our experiments show that the proposed framework can accurately predict portrait body poses and generate view-consistent, realistic portrait images with complete geometry from all camera angles.
Hao Xu 0049, Xiangjun Tang, Yue Shangguan, Hongbo Fu 0001, Xiaogang Jin 0001
IEEE Trans. Circuits Syst. Video Technol.2
2025 AlignTex: Pixel-Precise Texture Generation from Multi-view Artwork
abstract
Current 3D asset creation pipelines typically consist of three stages: creating multi-view concept art, producing 3D meshes based on the artwork, and painting textures for the meshes—an often labor-intensive process. Automated texture generation offers significant acceleration, but prior methods, which fine-tune 2D diffusion models with multi-view input images, often fail to preserve pixel-level details. These methods primarily emphasize semantic and subject consistency, which do not meet the requirements of artwork-guided texture workflows. To address this, we present AlignTex , a novel framework for generating high-quality textures from 3D meshes and multi-view artwork, ensuring both appearance detail and geometric consistency. AlignTex operates in two stages: aligned image generation and texture refinement. The core of our approach, AlignNet , resolves complex misalignments by extracting information from both the artwork and the mesh, generating images compatible with orthographic projection while maintaining geometric and visual fidelity. After projecting aligned images into the texture space, further refinement addresses seams and self-occlusion using an inpainting model and a geometry-aware texture dilation method. Experimental results demonstrate that AlignTex outperforms baseline methods in generation quality and efficiency, offering a practical solution to enhance 3D asset creation in gaming and film production.
Yuqing Zhang 0005, Hao Xu 0049, Sirui Lin, Xiang Li 0130, Xifeng Gao, Xiaogang Jin 0001
ACM Trans. Graph.2
2025 Geometry guidance diffusion image morphing with large shape difference
Hao Xu 0049, Xiwen Shi, Xiaogang Jin 0001
Vis. Comput.3
2024 Portrait3D: Text-Guided High-Quality 3D Portrait Generation Using Pyramid Representation and GANs Prior
abstract
Existing neural rendering-based text-to-3D-portrait generation methods typically make use of human geometry prior and diffusion models to obtain guidance. However, relying solely on geometry information introduces issues such as the Janus problem, over-saturation, and over-smoothing. We present Portrait3D , a novel neural rendering-based framework with a novel joint geometry-appearance prior to achieve text-to-3D-portrait generation that overcomes the aforementioned issues. To accomplish this, we train a 3D portrait generator, 3DPortraitGAN, as a robust prior. This generator is capable of producing 360° canonical 3D portraits, serving as a starting point for the subsequent diffusion-based generation process. To mitigate the "grid-like" artifact caused by the high-frequency information in the feature-map-based 3D representation commonly used by most 3D-aware GANs, we integrate a novel pyramid tri-grid 3D representation into 3DPortraitGAN. To generate 3D portraits from text, we first project a randomly generated image aligned with the given prompt into the pre-trained 3DPortraitGAN's latent space. The resulting latent code is then used to synthesize a pyramid tri-grid. Beginning with the obtained pyramid tri-grid , we use score distillation sampling to distill the diffusion model's knowledge into the pyramid tri-grid. Following that, we utilize the diffusion model to refine the rendered images of the 3D portrait and then use these refined images as training data to further optimize the pyramid tri-grid , effectively eliminating issues with unrealistic color and unnatural artifacts. Our experimental results show that Portrait3D can produce realistic, high-quality, and canonical 3D portraits that align with the prompt.
Hao Xu 0049, Xiangjun Tang, Xien Chen, Siyu Tang 0001, Zhebin Zhang, Chen Li 0062, Xiaogang Jin 0001
ACM Trans. Graph.2
2024 FusionDeformer: text-guided mesh deformation using diffusion models
Hao Xu 0049, Xiangjun Tang, Jing Zhang 0038, Zhebin Zhang, Chen Li 0062, Xiaogang Jin 0001
Vis. Comput.1
2024 Publisher Correction: FusionDeformer: text-guided mesh deformation using diffusion models
Hao Xu 0049, Xiangjun Tang, Jing Zhang 0038, Zhebin Zhang, Chen Li 0062, Xiaogang Jin 0001
Vis. Comput.1
2022 Effective Eyebrow Matting with Domain Adaptation
abstract
Abstract We present the first synthetic eyebrow matting datasets and a domain adaptation eyebrow matting network for learning domain‐robust feature representation using synthetic eyebrow matting data and unlabeled in‐the‐wild images with adversarial learning. Different from existing matting methods that may suffer from the lack of ground‐truth matting datasets, which are typically labor‐intensive to annotate or even worse, unable to obtain, we train the matting network in a semi‐supervised manner using synthetic matting datasets instead of ground‐truth matting data while achieving high‐quality results. Specifically, we first generate a large‐scale synthetic eyebrow matting dataset by rendering avatars and collect a real‐world eyebrow image dataset while maximizing the data diversity as much as possible. Then, we use the synthetic eyebrow dataset to train a multi‐task network, which consists of a regression task to estimate the eyebrow alpha mattes and an adversarial task to adapt the learned features from synthetic data to real data. As a result, our method can successfully train an eyebrow matting network using synthetic data without the need to label any real data. Our method can accurately extract eyebrow alpha mattes from in‐the‐wild images without any additional prior and achieves state‐of‐the‐art eyebrow matting performance. Extensive experiments demonstrate the superior performance of our method with both qualitative and quantitative results.
Luyuan Wang, Qinjie Xiao, Hao Xu 0049, Chunhua Shen, Xiaogang Jin 0001
Comput. Graph. Forum4