Zhuo Chen 0060

dblp:29/6497-60 · DBLP profile ↗
← Back
12ranked-venue papers
4as first author
12since 2021 · last 2026
0009-0004-1068-6525ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 10 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021
YearPublicationVenuePosition
2026 ExpDiff: Generating High-Fidelity 3D Facial Expression Meshes and BRDF Textures via Diffusion Model
abstract
3D face generation is a critical task for immersive multimedia applications, where a key challenge is the joint synthesis of expressive geometry and BRDF textures. Existing methods often struggle with geometric-textural coherence and corresponding reflectance modeling. To overcome these limitations, we present ExpDiff, a framework that generates expression meshes and corresponding BRDF textures from a single neutral-expression face. Our method employs an attention-based diffusion model to learn the semantic transition across expressions. To ensure correspondence between geometry and texture, we introduce a unified representation that explicitly models geometric-textural interaction, which is encoded into a shared latent space by models pre-trained on a vast dataset for strong generalization. To achieve semantically coherent and physically consistent generation, we propose to guide the denoising direction with specially designed textual prompts. We further construct two novel facial expression datasets, J-Reflectance, for ultra-high-quality assets, and FFHQ-BRDFExp for diverse identities, both of which are publicly released to advance the community. Extensive experiments demonstrate our method's superior performance in photo-realistic facial expression synthesis. Project page:https://cyh-sj.github.io/expdiff/.
Yuhao Cheng, Xuanchen Li, Xingyu Ren, Zhuo Chen 0060, Chenghui Ke, Xiaokang Yang 0001, Yichao Yan
IEEE Trans. Multim.4
2026 Relightable and Animatable Gaussian Head Avatar From Monocular Videos
abstract
In the realm of virtual avatar creation, accurate relighting capabilities are key to enhancing realism and immersion. We propose a novel pipeline for building personalized and relightable avatars from a monocular video captured under unknown lighting. This minimal input poses challenges in material entanglement and novel-view inconsistency. To tackle these, we introduce a disentangled dynamic 3D Gaussian representation that models diverse material properties and supports photorealistic rendering and animation via a parametric face model. To resolve material ambiguity under uncontrolled lighting, we train a 2D diffusion-based model to predict canonical-lighting images and physically-based material maps from casually lit portraits. These predictions serve as supervisory signals to guide the 3D disentanglement process. Additionally, we incorporate a 3D prior to enhance novel-view consistency, improving geometry and appearance in unseen views. Experiments demonstrate that our approach significantly boosts reconstruction quality and relighting fidelity, offering a practical and cost-effective solution for creating high-quality personalized avatars.
Zhuo Chen 0060, Yichao Yan, Jingnan Gao, Zhuo Su 0008, Zhaohu Li, Yuhao Cheng, Xueying Lee, Yutong Leng, Yikun Zeng, Guidong Wang, Xiaokang Yang 0001
IEEE Trans. Vis. Comput. Graph.1
2025 Multimodal Latent Diffusion Model for Complex Sewing Pattern Generation
abstract
Generating sewing patterns in garment design is receiving increasing attention due to its CG-friendly and flexible-editing nature. Previous sewing pattern generation methods have been able to produce exquisite clothing, but struggle to design complex garments with detailed control. To address these issues, we propose SewingLDM, a multi-modal generative model that generates sewing patterns controlled by text prompts, body shapes, and garment sketches. Initially, we extend the original vector of sewing patterns into a more comprehensive representation to cover more intricate details and then compress them into a compact latent space. To learn the sewing pattern distribution in the latent space, we design a two-step training strategy to inject the multi-modal conditions, \ie, body shapes, text prompts, and garment sketches, into a diffusion model, ensuring the generated garments are body-suited and detail-controlled. Comprehensive qualitative and quantitative experiments show the effectiveness of our proposed method, significantly surpassing previous approaches in terms of complex garment design and various body adaptability. Our project page: https://shengqiliu1.github.io/SewingLDM.
Shengqi Liu, Yuhao Cheng, Zhuo Chen 0060, Xingyu Ren, Wenhan Zhu, Lincheng Li, Mengxiao Bi, Xiaokang Yang 0001, Yichao Yan
ICCV3
2025 AniSDF: Fused-Granularity Neural Surfaces with Anisotropic Encoding for High-Fidelity 3D Reconstruction
abstract
Neural radiance fields have recently revolutionized novel-view synthesis and achieved high-fidelity renderings. However, these methods sacrifice the geometry for the rendering quality, limiting their further applications including relighting and deformation. How to synthesize photo-realistic rendering while reconstructing accurate geometry remains an unsolved problem. In this work, we present AniSDF, a novel approach that learns fused-granularity neural surfaces with physics-based encoding for high-fidelity 3D reconstruction. Different from previous neural surfaces, our fused-granularity geometry structure balances the overall structures and fine geometric details, producing accurate geometry reconstruction. To disambiguate geometry from reflective appearance, we introduce blended radiance fields to model diffuse and specularity following the anisotropic spherical Gaussian encoding, a physics-based rendering pipeline. With these designs, AniSDF can reconstruct objects with complex structures and produce high-quality renderings. Furthermore, our method is a unified model that does not require complex hyperparameter tuning for specific objects. Extensive experiments demonstrate that our method boosts the quality of SDF-based methods by a great scale in both geometry reconstruction and novel-view synthesis.
Jingnan Gao, Zhuo Chen 0060, Xiaokang Yang 0001, Yichao Yan
ICLR2
2025 Revealing Directions for Text-Guided 3D Face Editing
abstract
3D face editing is a significant task in multimedia, aimed at the manipulation of 3D face models across various control signals. The success of 3D-aware GAN provides expressive 3D models learned from 2D single-view images only, encouraging researchers to discover semantic editing directions in its latent space. However, previous methods face challenges in balancing quality, efficiency, and generalization. To solve the problem, we explore the possibility of introducing the strength of diffusion model into 3D-aware GANs. In this paper, we presentFace Clan, a fast and text-general approach for generating and manipulating 3D faces based on arbitrary attribute descriptions. To achieve disentangled editing, we propose to diffuse on the latent space under a pair of opposite prompts to estimate the mask indicating the region of interest on latent codes. Based on the mask, we then apply denoising to the masked latent codes to reveal the editing direction. Our method offers a precisely controllable manipulation method, allowing users to intuitively customize regions of interest with the text description. Experiments demonstrate the effectiveness and generalization of our Face Clan for various pre-trained GANs. It offers an intuitive and wide application for text-guided face editing that contributes to the landscape of multimedia content creation. Our project page:https://windlikestone.github.io/Face_clan_website/.
Zhuo Chen 0060, Yichao Yan, Shengqi Liu, Yuhao Cheng, Weiming Zhao, Lincheng Li, Mengxiao Bi, Xiaokang Yang 0001
IEEE Trans. Multim.1
2025 EvaSurf: Efficient View-Aware Implicit Textured Surface Reconstruction
abstract
Reconstructing real-world 3D objects has numerous applications in computer vision, such as virtual reality, video games, and animations. Ideally, 3D reconstruction methods should generate high-fidelity results with 3D consistency in real-time. Traditional methods match pixels between images using photo-consistency constraints or learned features, while differentiable rendering methods like Neural Radiance Fields (NeRF) use differentiable volume rendering or surface-based representation to generate high-fidelity scenes. However, these methods require excessive runtime for rendering, making them impractical for daily applications. To address these challenges, we present EvaSurf, an Efficient View-Aware implicit textured Surface reconstruction method on mobile devices. In our method, we first employ an efficient surface-based model with a multi-view supervision module to ensure accurate mesh reconstruction. To enable high-fidelity rendering, we learn an implicit texture embedded with view-aware encoding to capture view-dependent information. Furthermore, with the explicit geometry and the implicit texture, we can employ a lightweight neural shader to reduce the expense of computation and further support real-time rendering on common mobile devices. Extensive experiments demonstrate that our method can reconstruct high-quality appearance and accurate mesh on both synthetic and real-world datasets. Moreover, our method can be trained in just 1-2 hours using a single GPU and run on mobile devices at over 40 FPS (Frames Per Second), with a final package required for rendering taking up only 40-50 MB.
Jingnan Gao, Zhuo Chen 0060, Yichao Yan, Bowen Pan, Jiangjing Lyu, Xiaokang Yang 0001
IEEE Trans. Vis. Comput. Graph.2
2024 3D-Aware Face Editing via Warping-Guided Latent Direction Learning
abstract
3D facial editing, a longstanding task in computer vision with broad applications, is expected to fast and intuitively manipulate any face from arbitrary viewpoints following the user's will. Existing works have limitations in terms of intuitiveness, generalization, and efficiency. To overcome these challenges, we propose FaceEdit3D, which allows users to directly manipulate 3D points to edit a 3D face, achieving natural and rapid face editing. After one or several points are manipulated by users, we propose the tri-plane warping to directly deform the view-independent 3D representation. To address the problem of distortion caused by tri-plane warping, we train a warp-aware encoder to project the warped face onto a standardized latent space. In this space, we further propose directional latent editing to mitigate the identity bias caused by the encoder and realize the disentangled editing of various attributes. Extensive experiments show that our method achieves superior results with rich facial details and nice identity preservation. Our approach also supports general applications like multi-attribute continuous editing and cat/car editing. The project website is https://cyh-sj.github.io/FaceEdit3DI.
Yuhao Cheng, Zhuo Chen 0060, Xingyu Ren, Wenhan Zhu, Zhengqin Xu, Di Xu 0012, Changpeng Yang, Yichao Yan
CVPR2
2024 Infusion: Preventing Customized Text-to-Image Diffusion from Overfitting
abstract
Text-to-image (T2I) customization aims to create images that embody specific visual concepts delineated in textual descriptions. However, existing works still face a main challenge, concept overfitting. To tackle this challenge, we first analyze overfitting, categorizing it into concept-agnostic overfitting, which undermines non-customized concept knowledge, and concept-specific overfitting, which is confined to customize on limited diversities, i.e, backgrounds, layouts, styles. To evaluate the overfitting degree, we further introduce two metrics, i.e, Latent Fisher divergence and Wasserstein metric to measure the distribution changes of non-customized and customized concept respectively. Drawing from the analysis, we propose Infusion, a T2I customization method that enables the learning of target concepts to avoid being constrained by limited training diversities, while preserving non-customized knowledge. Remarkably, Infusion achieves this feat with remarkable efficiency, requiring a mere 11KB of trained parameters. Extensive experiments also demonstrate that our approach outperforms state-of-the-art methods in both single and multi-concept customized generation. Project page: https://zwl666666.github.io/infusion/.
Yichao Yan, Zhuo Chen 0060, Pengzhi Chu, Weiming Zhao, Xiaokang Yang 0001
ACM Multimedia4
2024 Multi-times Monte Carlo Rendering for Inter-reflection Reconstruction
abstract
Inverse rendering methods have achieved remarkable performance in reconstructing high-fidelity 3D objects with disentangled geometries, materials, and environmental light. However, they still face huge challenges in reflective surface reconstruction. Although recent methods model the light trace to learn specularity, the ignorance of indirect illumination makes it hard to handle inter-reflections among multiple smooth objects. In this work, we propose Ref-MC2 that introduces the multi-time Monte Carlo sampling which comprehensively computes the environmental illumination and meanwhile considers the reflective light from object surfaces. To address the computation challenge as the times of Monte Carlo sampling grow, we propose a specularity-adaptive sampling strategy, significantly reducing the computational complexity. Besides the computational resource, higher geometry accuracy is also required because geometric errors accumulate multiple times. Therefore, we further introduce a reflection-aware surface model to initialize the geometry and refine it during inverse rendering. We construct a challenging dataset containing scenes with multiple objects and inter-reflections. Experiments show that our method outperforms other inverse rendering methods on various object groups. We also show downstream applications, e.g., relighting and material editing, to illustrate the disentanglement ability of our method.
Tengjie Zhu, Zhuo Chen 0060, Jingnan Gao, Yichao Yan, Xiaokang Yang 0001
NeurIPS2
2024 Directional Texture Editing for 3D Models
abstract
Abstract Texture editing is a crucial task in 3D modelling that allows users to automatically manipulate the surface materials of 3D models. However, the inherent complexity of 3D models and the ambiguous text description lead to the challenge of this task. To tackle this challenge, we propose ITEM3D, a Texture Editing Model designed for automatic 3D object editing according to the text Instructions. Leveraging the diffusion models and the differentiable rendering, ITEM3D takes the rendered images as the bridge between text and 3D representation and further optimizes the disentangled texture and environment map. Previous methods adopted the absolute editing direction, namely score distillation sampling (SDS) as the optimization objective, which unfortunately results in noisy appearances and text inconsistencies. To solve the problem caused by the ambiguous text, we introduce a relative editing direction, an optimization objective defined by the noise difference between the source and target texts, to release the semantic ambiguity between the texts and images. Additionally, we gradually adjust the direction during optimization to further address the unexpected deviation in the texture domain. Qualitative and quantitative experiments show that our ITEM3D outperforms the state‐of‐the‐art methods on various 3D objects. We also perform text‐guided relighting to show explicit control over lighting. Our project page: https://shengqiliu1.github.io/ITEM3D/ .
Shengqi Liu, Zhuo Chen 0060, Jingnan Gao, Yichao Yan, Wenhan Zhu, Jiangjing Lyu, Xiaokang Yang 0001
Comput. Graph. Forum2
2024 HyperStyle3D: Text-Guided 3D Portrait Stylization via Hypernetworks
abstract
Portrait stylization is a long-standing task enabling extensive applications. Although 2D-based methods have made great progress in recent years, real-world applications such as metaverse and games often demand 3D content. On the other hand, the requirement of 3D data, which is costly to acquire, significantly impedes the development of 3D portrait stylization methods. In this paper, inspired by the success of 3D-aware GANs that bridge 2D and 3D domains with 3D fields as the intermediate representation for rendering 2D images, we propose a novel method, dubbed HyperStyle3D, based on 3D-aware GANs for 3D portrait stylization. At the core of our method is a hyper-network learned to manipulate the parameters of the generator in a single forward pass. It not only offers a strong capacity to handle multiple styles with a single model, but also enables flexible fine-grained stylization that affects only texture, shape, or local part of the portrait. While the use of 3D-aware GANs bypasses the requirement of 3D data, we further alleviate the necessity of style images with the CLIP model being the style guidance. We conduct an extensive set of experiments across the style, attribute, and shape, and meanwhile, measure the 3D consistency. These experiments demonstrate the superior capability of our HyperStyle3D model in rendering 3D-consistent images in diverse styles, deforming the face shape, and editing various attributes.
Zhuo Chen 0060, Xudong Xu, Yichao Yan, Wenhan Zhu, Wayne Wu, Bo Dai 0002, Xiaokang Yang 0001
IEEE Trans. Circuits Syst. Video Technol.1
2023 HiStyle: Reinventing Historic Portrait for 3D Colorization and Zero-Shot Stylization
abstract
Restoring and reinventing historical portraits has long been a challenging task in computer vision. In order to faithfully explore the portrait images, it is necessary to restore the color and reconstruct the 3D geometry. Furthermore, the ability to stylize historic portraits is crucial for extending their use to different forms of media. Existing methods for each specific task make huge progress. However, they struggle with conflicts of multi-tasks, which hinders their ability to meet the requirement in a unified model. To achieve this goal, we propose HiStyle, a generative model for generally reinventing historic portraits, which simultaneously realizes 2D-to-3D reconstruction, gray-to-RGB restoration, and photo-to-style image translation. We introduce a GAN inversion technique to transfer a gray historical portrait into the latent space of a 3D generator, restoring the lost color information and meanwhile lifting the 2D image to 3D representation. Besides, we incorporate the power of CLIP model with 3D-aware GANs to achieve zero-shot text-driven style transfer. The results demonstrate the superior of HiStyle in the quality and diversity of synthesized images and also highlight the potential of 3D-aware GANs for preserving cultural heritage.
Zhuo Chen 0060, Zhu Li 0001
MMSP1