Weijian Cao

dblp:192/2786 · DBLP profile ↗
← Back
12ranked-venue papers
0as first author
12since 2021 · last 2026
0009-0006-2554-1294ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 12 since 2021Artificial intelligence and machine learning · 6 · 6 since 2021
YearPublicationVenuePosition
2026 SwiftVideo: A Unified Framework for Few-Step Video Generation Through Trajectory-Distribution Alignment
abstract
Diffusion-based or flow-based models have achieved significant progress in video synthesis but require multiple iterative sampling steps, which incurs substantial computational overhead. While many distillation methods that are solely based on trajectory-preserving or distribution-matching have been developed to accelerate video generation models, these approaches often suffer from performance breakdown or increased artifacts in few-step settings. To address these limitations, we propose SwiftVideo, a unified and stable distillation framework that combines the advantages of trajectory-preserving and distribution-matching strategies. Our approach introduces continuous-time consistency distillation to ensure precise preservation of ODE trajectories. Subsequently, We propose a dual-perspective alignment encompassing distribution alignment between synthetic and real data along with trajectory alignment across different inference steps. Our method maintains high-quality video generation while substantially reducing the number of inference steps. Quantitative evaluations on the OpenVid-1M benchmark demonstrate that our method significantly outperforms existing approaches in few-step video generation.
Yanxiao Sun, Jiafu Wu, Yun Cao 0002, Chengming Xu 0001, Yabiao Wang, Weijian Cao, Donghao Luo 0001, Chengjie Wang 0001, Yanwei Fu 0001
AAAI6
2026 Semantic Frame Interpolation
abstract
Generating intermediate video content of varying lengths based on given first and last frames, along with text prompt information, offers significant research and application potential. However, traditional frame interpolation tasks primarily focus on scenarios with a small number of frames, no text control, and minimal differences between the first and last frames. Recent community developers have utilized large video models represented by Wan to endow frame-to-frame capabilities. However, these models can only generate a fixed number of frames and often fail to produce satisfactory results for certain frame lengths, while this setting lacks a clear official definition and a well-established benchmark. In this paper, we first propose a new practical Semantic Frame Interpolation (SFI) task from the perspective of academic definition, which covers the above two settings and supports inference at multiple frame rates. To achieve this goal, we propose a novel SemFi model building upon Wan2.1, which incorporates a Mixture-of-LoRA module to ensure the generation of high-consistency content that aligns with control conditions across various frame length limitations. Furthermore, we propose SFI-300K, the first general-purpose dataset and benchmark specifically designed for SFI. To support this, we collect and process data from the perspective of SFI, carefully designing evaluation metrics and methods to evaluate the performance of the model in multiple dimensions, including images and videos, and various aspects, including consistency and diversity. Through extensive experiments on SFI-300K, we demonstrate that our method is particularly well-suited to meet the requirements of the SFI task.
Yijia Hong, Jiangning Zhang, Ran Yi 0002, Weijian Cao, Xiaobin Hu, Lizhuang Ma, Shuicheng Yan
IEEE Trans. Image Process.4
2025 ID-Sculpt: ID-aware 3D Head Generation from Single In-the-wild Portrait Image
abstract
While recent works have achieved great success on one-shot 3D common object generation, high quality and fidelity 3D head generation from a single image remains a great challenge. Previous text-based methods for generating 3D heads were limited by text descriptions and image-based methods struggled to produce high-quality head geometry. To handle this challenging problem, we propose a novel framework, ID-Sculpt, to generate high-quality 3D heads while preserving their identities. Our work incorporates the identity information of the portrait image into three parts: 1) geometry initialization, 2) geometry sculpting, and 3) texture generation stages. Given a reference portrait image, we first align the identity features with text features to realize ID-aware guidance enhancement, which contains the control signals representing the face information. We then use the canny map, ID features of the portrait image, and a pre-trained text-to-normal/depth diffusion model to generate ID-aware geometry supervision and 3D-GAN inversion is employed to generate ID-aware geometry initialization. Furthermore, with the ability to inject identity information into 3D head generation, we use ID-aware guidance to calculate ID-aware Score Distillation (ISD) for geometry sculpting. For texture generation, we adopt the ID Consistent Texture Inpainting and Refinement which progressively expands the view for texture inpainting to obtain an initialization UV texture map. We then use the id-aware guidance to provide image-level supervision for noisy multi-view images to obtain a refined texture map. Extensive experiments demonstrate that we can generate high-quality 3D heads with accurate geometry and texture from a single in-the-wild portrait image.
Jinkun Hao, Junshu Tang, Jiangning Zhang, Ran Yi 0002, Yijia Hong, Moran Li, Weijian Cao, Chengjie Wang 0001, Lizhuang Ma
AAAI7
2025 VividPose: Vividly 3D-driven Stable Pose Diffusion of High Facial Fidelity
abstract
Human image animation aims to generate a video from a static image by following a specified pose sequence. Existing methods typically adopt a multi-stage pipeline that separately learns appearance and motion, leading to appearance degradation and temporal inconsistencies. To address these issues, we propose VividPose, an innovative end-to-end pipeline based on Stable Video Diffusion (SVD) that ensures superior temporal stability. To enhance the retention of human identity, we propose an identity-aware appearance controller that integrates additional facial information without compromising other appearance details such as clothing texture and background. To accommodate diverse human body shapes and hand movements, we introduce a geometry-aware pose controller that utilizes both rendering maps and skeleton maps. Extensive qualitative and quantitative experiments on the UBCFashion and TikTok benchmarks demonstrate that our method achieves state-of-the-art performance. Furthermore, VividPose exhibits superior generalization capabilities on our proposed in-the-wild dataset. Our project page is available at https://kelu007.github.io/vivid-pose.
Zhengkai Jiang 0001, Chengming Xu 0001, Jiangning Zhang, Yabiao Wang, Xinyi Zhang 0005, Yun Cao 0002, Weijian Cao, Chengjie Wang 0001, Zhanxiong Wang, Yanwei Fu 0001
ICME8
2025 Identity-Preserving Text-to-Video Generation Guided by Simple yet Effective Spatial-Temporal Decoupled Representations
abstract
Identity-preserving text-to-video (IPT2V) generation, which aims to create high-fidelity videos with consistent human identity, has become crucial for downstream applications. However, current end-to-end frameworks suffer a critical spatial-temporal trade-off: optimizing for spatially coherent layouts of key elements ( e.g., character identity preservation) often compromises instruction-compliant temporal smoothness, while prioritizing dynamic realism risks disrupting the spatial coherence of visual structures. To tackle this issue, we propose a simple yet effective spatial-temporal decoupled framework that decomposes representations into spatial features for layouts and temporal features for motion dynamics. Specifically, our paper proposes a semantic prompt optimization mechanism and stage-wise decoupled generation paradigm. The former module decouples the prompt into spatial and temporal components. Aligned with the subsequent stage-wise decoupled approach, the spatial prompts guide the text-to-image (T2I) stage to generate coherent spatial features, while the temporal prompts direct the sequential image-to-video (I2V) stage to ensure motion consistency. Experimental results validate that our approach achieves excellent spatiotemporal consistency, demonstrating outstanding performance in identity preservation, text relevance, and video quality. By leveraging this simple yet robust mechanism, our algorithm secures the runner-up position in 2025 ACM Multimedia Challenge. Our code is available at https://github.com/rain152/IPVG.
Yuji Wang, Moran Li, Xiaobin Hu, Ran Yi 0002, Jiangning Zhang, Weijian Cao, Yabiao Wang, Chengjie Wang 0001, Lizhuang Ma
ACM Multimedia7
2025 StrandDesigner: Towards Practical Strand Generation with Sketch Guidance
Moran Li, Chengming Xu 0001, Xiaobin Hu, Jiangning Zhang, Weijian Cao, Chengjie Wang 0001, Yanwei Fu 0001
ACM Multimedia7
2025 MagicFace: Slot-Driven High-Fidelity One-Shot Facial Appearance Editing
abstract
Facial Appearance Editing (FAE) focuses on modifying physical attributes like pose, expression, and lighting in facial images while preserving identity and background, which is crucial in photography. Despite significant progress, current research faces three main challenges: low generation fidelity, poor attribute preservation, and inefficient inference. To address these issues, this paper introduces MagicFace, a slot-driven, high-fidelity, one-shot FAE framework. Particularly, MagicFace employs Space-sensitive Physical Customization (SPC) for accurate query attribute transfer, utilizing rendering textures from the 3D Morphable Model (3DMM) to ensure fidelity and generalization. To preserve source attributes, we propose the Region-responsive Semantic Composition (RSC) module, guided by slot attention to learn and preserve decoupled source features. This approach can successfully maintain identity and reduces artifacts related to non-facial attributes such as hair, clothes, and background. Additionally, we introduce a consistency regularization technique to enhance editing controllability by leveraging prior knowledge from attention matrices in the diffusion model. Extensive experiments demonstrate that MagicFace achieves state-of-the-art performance in FAE task, outperforming existing methods.
Jiangning Zhang, Chengming Xu 0001, Weijian Cao, Yanwei Fu 0001
IEEE Signal Process. Lett.4
2024 FreeMotion: A Unified Framework for Number-Free Text-to-Motion Synthesis
Junshu Tang, Weijian Cao, Ran Yi 0002, Moran Li, Jingyu Gong, Jiangning Zhang, Yabiao Wang, Chengjie Wang 0001, Lizhuang Ma
ECCV (8)3
2024 TexDreamer: Towards Zero-Shot High-Fidelity 3D Human Texture Generation
Junshu Tang, Jiangning Zhang, Weijian Cao, Chengjie Wang 0001, Yunsheng Wu, Dongjin Huang
ECCV (47)6
2024 MambaGesture: Enhancing Co-Speech Gesture Generation with Mamba and Disentangled Multi-Modality Fusion
abstract
Co-speech gesture generation is crucial for producing synchronized and realistic human gestures that accompany speech, enhancing the animation of lifelike avatars in virtual environments. While diffusion models have shown impressive capabilities, current approaches often overlook a wide range of modalities and their interactions, resulting in less dynamic and contextually varied gestures. To address these challenges, we present MambaGesture, a novel framework integrating a Mamba-based attention block, MambaAttn, with a multi-modality feature fusion module, SEAD. The MambaAttn block combines the sequential data processing strengths of the Mamba model with the contextual richness of attention mechanisms, enhancing the temporal coherence of generated gestures. SEAD adeptly fuses audio, text, style, and emotion modalities, employing disentanglement to deepen the fusion process and yield gestures with greater realism and diversity. Our approach, rigorously evaluated on the multi-modal BEAT dataset, demonstrates significant improvements in Fréchet Gesture Distance (FGD), diversity scores, and beat alignment, achieving state-of-the-art performance in co-speech gesture generation.
Chencan Fu, Yabiao Wang, Jiangning Zhang, Zhengkai Jiang 0001, Xiaofeng Mao, Jiafu Wu, Weijian Cao, Chengjie Wang 0001, Yanhao Ge, Yong Liu 0007
ACM Multimedia7
2023 Learning Neural Proto-Face Field for Disentangled 3D Face Modeling in the Wild
abstract
Generative models show good potential for recovering 3D faces beyond limited shape assumptions. While plausible details and resolutions are achieved, these models easily fail under extreme conditions of pose, shadow or appearance, due to the entangled fitting or lack of multi-view priors. To address this problem, this paper presents a novel Neural Proto-face Field (NPF) for unsupervised robust 3D face modeling. Instead of using constrained images as Neural Radiance Field (NeRF), NPF disentangles the common/specific facial cues, i.e., ID, expression and scene-specific details from in-the-wild photo collections. Specifically, NPF learns a face prototype to aggregate 3D-consistent identity via uncertainty modeling, extracting multi-image priors from a photo collection. NPF then learns to deform the prototype with the appropriate facial expressions, constrained by a loss of expression consistency and personal idiosyncrasies. Finally, NPF is optimized to fit a target image in the collection, recovering specific details of appearance and geometry. In this way, the generative model benefits from multi-image priors and meaningful facial structures. Extensive experiments on benchmarks show that NPF recovers superior or competitive facial shapes and textures, compared to state-of-the-art methods.
Zhenyu Zhang 0005, Renwang Chen, Weijian Cao, Ying Tai, Chengjie Wang 0001
CVPR3
2022 Physically-guided Disentangled Implicit Rendering for 3D Face Modeling
abstract
This paper presents a novel Physically-guided Disentangled Implicit Rendering (PhyDIR) framework for highfidelity 3D face modeling. The motivation comes from two observations: Widely-used graphics renderers yield excessive approximations against photo-realistic imaging, while neural rendering methods produce superior appearances but are highly entangled to perceive 3D-aware operations. Hence, we learn to disentangle the implicit rendering via explicit physical guidance, while guaranteeing the properties of: (1) 3D-aware comprehension and (2) high-reality image formation. For the former one, PhyDIR explicitly adopts 3D shading and rasterizing modules to control the renderer, which disentangles the light, facial shape, and viewpoint from neural reasoning. Specifically, PhyDIR proposes a novel multi-image shading strategy to compensate for the monocular limitation, so that the lighting variations are accessible to the neural renderer. For the latter, PhyDIR learns the face-collection implicit texture to avoid ill-posed intrinsic factorization, then leverages a series of consistency losses to constrain the rendering robustness. With the disentangled method, we make 3D face modeling benefit from both kinds of rendering strategies. Extensive experiments on benchmarks show that PhyDIR obtains superior performance than state-of-the-art explicit/implicit methods on geometry/texture modeling.
Zhenyu Zhang 0005, Yanhao Ge, Ying Tai, Weijian Cao, Renwang Chen, Kunlin Liu, Hao Tang 0005, Chengjie Wang 0001, Dongjin Huang
CVPR4