Quanwei Yang

dblp:329/0636 · DBLP profile ↗
← Back
14ranked-venue papers
4as first author
14since 2021 · last 2026
0000-0003-3997-0031ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 2 first-author · 10 since 2021Artificial intelligence and machine learning · 8 · 3 first-author · 8 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Cosh-DiT: Co-Speech Gesture Video Synthesis via Hybrid Audio-Visual Diffusion Transformers
Yasheng Sun, Hang Zhou 0009, Jiazhi Guan, Quanwei Yang, Kaisiyuan Wang, Borong Liang, Haocheng Feng, Jingdong Wang 0001, Ziwei Liu 0002, Hideki Koike
Int. J. Comput. Vis.5
2026 SwapController: Toward Improving Identity and Attribute Control for Diffusion-Based Face Swapping
abstract
Face swapping efforts strive to achieve high-fidelity and well-controlled generation effects. Owing to the remarkable generative capabilities, diffusion models deliver promising high-fidelity solutions. However, their intrinsic stochastic properties complicate the accurate modeling of facial representations, introducing new challenges for identity and attribute consistency of the generated faces. In this paper, we introduce a novel diffusion-based face-swapping framework, named SwapController, which achieves high-fidelity generation via careful facial identity and attribute modeling. Specifically, our facial modeling mainly involves facial structure and facial texture. For structure modeling, 3D facial priors are leveraged to provide explicit structure supervision, enabling accurate head structure control. On this basis, two novel components are proposed to deeply mine facial textural representations from identity and attribute aspects. To enhance identity control, multi-grained source identity embeddings are obtained from various functional encoders to convey critical global identity and fine-grained identity details. To improve attribute modeling, identity-shifted attribute embeddings are derived by applying identity modulation to the most salient textural attribute features of the target face. Moreover, in line with the diffusion denoising characteristics, a timestep-aware identity optimization objective is introduced to optimize identity consistency guidelines and overall fidelity. Extensive experiments demonstrate the effectiveness of our SwapController in generating identity-consistent portrait images while faithfully preserving target attributes, which obtains a 98.32 ID Retrieval, exceeding the SOTA DiffSFSR by 7.32 $\uparrow$↑.
Lingyun Yu 0002, Quanwei Yang, Runxin Liu, Yongdong Zhang 0001, Hongtao Xie 0001
IEEE Trans. Vis. Comput. Graph.3
2025 Re-HOLD: Video Hand Object Interaction Reenactment via adaptive Layout-instructed Diffusion Model
abstract
Current digital human studies focusing on lip-syncing and body movement are no longer sufficient to meet the growing industrial demand, while human video generation techniques that support interacting with real-world environments (e.g., objects) have not been well investigated. Despite human hand synthesis already being an intricate problem, generating objects in contact with hands and their interactions presents an even more challenging task, especially when the objects exhibit obvious variations in size and shape. To tackle these issues, we present a novel video reenactment framework focusing on Human-Object Interaction (HOI) via an adaptive Layout-instructed Diffusion model (Re-HOLD). Our key insight is to employ specialized layout representation for hands and objects, respectively. Such representations enable effective disentanglement of hand modeling and object adaptation to diverse motion sequences. To further improve the quality of the HOI generation, we design an interactive textural enhancement module for both hands and objects by introducing two independent memory banks. We also propose a layout adjustment strategy for the cross-object reenactment scenario to adaptively adjust unreasonable layouts caused by diverse object sizes during inference. Comprehensive qualitative and quantitative evaluations demonstrate that our proposed framework significantly outperforms existing methods. Project page: https://fyycs.github.io/Re-HOLD.
Quanwei Yang, Kaisiyuan Wang, Hang Zhou 0009, Haocheng Feng, Errui Ding, Yu Wu 0011, Jingdong Wang 0001
CVPR2
2025 AudCast: Audio-Driven Human Video Generation by Cascaded Diffusion Transformers
abstract
Despite the recent progress of audio-driven video generation, existing methods mostly focus on driving facial movements, leading to non-coherent head and body dynamics. Moving forward, it is desirable yet challenging to generate holistic human videos with both accurate lip-sync and delicate co-speech gestures w.r.t. given audio. In this work, we propose AudCast, a generalized audio-driven human video generation framework adopting a cascade Diffusion-Transformers (DiTs) paradigm, which synthesizes holistic human videos based on a reference image and a given audio. 1) Firstly, an audio-conditioned Holistic Human DiT architecture is proposed to directly drive the movements of any human body with vivid gesture dynamics. 2) Then to enhance hand and face details that are well-knownly difficult to handle, a Regional Refinement DiT leverages regional 3D fitting as the bridge to reform the signals, producing the final results. Extensive experiments demonstrate that our framework generates high-fidelity audio-driven holistic human videos with temporal coherence and fine facial and hand details. Resources can be found at https://guanjz20.github.io/projects/AudCast.
Jiazhi Guan, Kaisiyuan Wang, Quanwei Yang, Yasheng Sun, Shengyi He, Borong Liang, Haocheng Feng, Errui Ding, Jingdong Wang 0001, Youjian Zhao, Hang Zhou 0009, Ziwei Liu 0002
CVPR4
2025 Forensic-MoE: Exploring Comprehensive Synthetic Image Detection Traces With Mixture of Experts
Mingqi Fang, Ziguang Li, Lingyun Yu 0002, Quanwei Yang, Hongtao Xie 0001, Yongdong Zhang 0001
ICCV4
2025 GestureHYDRA: Semantic Co-Speech Gesture Synthesis via Hybrid Modality Diffusion Transformer and Cascaded-Synchronized Retrieval-Augmented Generation
Quanwei Yang, Luying Huang, Kaisiyuan Wang, Jiazhi Guan, Shengyi He, Fengguo Li, Hang Zhou 0009, Lingyun Yu 0002, Haocheng Feng, Hongtao Xie 0001
ICCV1
2025 THGS: Lifelike Talking Human Avatar Synthesis From Monocular Video Via 3D Gaussian Splatting
abstract
Abstract Despite the remarkable progress in 3D talking head generation, directly generating 3D talking human avatars still suffers from rigid facial expressions, distorted hand textures and out‐of‐sync lip movements. In this paper, we extend speaker‐specific talking head generation task to talking human avatar synthesis and propose a novel pipeline, THGS, that animates lifelike Talking Human avatars using 3D Gaussian Splatting (3DGS). Given speech audio, expression and body poses as input, THGS effectively overcomes the limitations of 3DGS human re‐construction methods in capturing expressive dynamics, such as mouth movements, facial expressions and hand gestures, from a short monocular video. Firstly, we introduce a simple yet effective Learnable Expression Blendshapes (LEB) for facial dynamics re‐construction, where subtle facial dynamics can be generated by linearly combining the static head model and expression blendshapes. Secondly, a Spatial Audio Attention Module (SAAM) is proposed for lip‐synced mouth movement animation, building connections between speech audio and mouth Gaussian movements. Thirdly, we employ a body pose, expression and skinning weights joint optimization strategy to optimize these parameters on the fly, which aligns hand movements and expressions better with video input. Experimental results demonstrate that THGS can achieve high‐fidelity 3D talking human avatar animation at 150+ fps on a web‐based rendering system, improving the requirements of real‐time applications. Our project page is at https://sora158.github.io/THGS.github.io/ .
Lingyun Yu 0002, Quanwei Yang, Aihua Zheng, Hongtao Xie 0001
Comput. Graph. Forum3
2025 TalkingAvatar: Learning 3D talking human avatar via NeRF
Lingyun Yu 0002, Chuanbin Liu 0001, Wu Liu 0005, Quanwei Yang, Meng Shao
Neurocomputing5
2025 High Fidelity Face Swapping via Facial Texture and Structure Consistency Mining
Lingyun Yu 0002, Quanwei Yang, Meng Shao, Hongtao Xie 0001
IEEE Trans. Multim.3
2024 ShowMaker: Creating High-Fidelity 2D Human Video via Fine-Grained Diffusion Modeling
abstract
Although significant progress has been made in human video generation, most previous studies focus on either human facial animation or full-body animation, which cannot be directly applied to produce realistic conversational human videos with frequent hand gestures and various facial movements simultaneously. To address these limitations, we propose a 2D human video generation framework, named ShowMaker, capable of generating high-fidelity half-body conversational videos via fine-grained diffusion modeling. We leverage dual-stream diffusion models as the backbone of our framework and carefully design two novel components for crucial local regions (i.e., hands and face) that can be easily integrated into our backbone. Specifically, to handle the challenging hand generation caused by sparse motion guidance, we propose a novel Key Point-based Fine-grained Hand Modeling module by amplifying positional information from raw hand key points and constructing a corresponding key point-based codebook. Moreover, to restore richer facial details in generated results, we introduce a Face Recapture module, which extracts facial texture features and global identity features from the aligned human face and integrates them into the diffusion process for face enhancement. Extensive quantitative and qualitative experiments demonstrate the superior visual quality and temporal consistency of our method.
Quanwei Yang, Jiazhi Guan, Kaisiyuan Wang, Lingyun Yu 0002, Wenqing Chu, Hang Zhou 0009, ZhiQiang Feng, Haocheng Feng, Errui Ding, Jingdong Wang 0001, Hongtao Xie 0001
NeurIPS1
2024 TALK-Act: Enhance Textural-Awareness for 2D Speaking Avatar Reenactment with Diffusion Model
Jiazhi Guan, Quanwei Yang, Kaisiyuan Wang, Hang Zhou 0009, Shengyi He, Haocheng Feng, Errui Ding, Jingdong Wang 0001, Hongtao Xie 0001, Youjian Zhao, Ziwei Liu 0002
SIGGRAPH Asia2
2024 Symmetrical Siamese Network for pose-guided person synthesis
Quanwei Yang, Lingyun Yu 0002, Yun Song, Meng Shao, Guoqing Jin, Hongtao Xie 0001
Comput. Vis. Image Underst.1
2023 High Fidelity Face Swapping via Semantics Disentanglement and Structure Enhancement
abstract
In this paper, we propose a novel Semantics and Structure-aware face Swapping framework (S2Swap) that exploits semantics disentanglement and structure enhancement for high fidelity face generation. Different from previous methods that either 1) suffer from degraded generation fidelity due to insufficient identity-attributes disentanglement or 2) neglect the importance of structure information for identity consistency, our approach can achieve local facial semantics disentanglement beyond global identity while boosting identity consistency through structure enhancement. Specifically, to achieve identity-attributes disentanglement, our S2Swap is designed from global-local perspectives. Firstly, an Oriented Identity Transfer module is proposed to globally disentangle target identity and attributes under global identity semantics prior. Such global disentanglement enables source identity transfer to the individual target identity. Secondly, a Local Semantics Disentanglement module is devised to disentangle local identity and identity-irrelevant facial semantics, providing local semantic compensation for the global counterpart. Moreover, to boost identity consistency, a Structure-Aware Head Modeling module is introduced to provide the desired face structure enhancement through an intuitive face sketch. Finally, considering the identity-attributes trade-off, we adaptively integrate semantics and structure information in a self-learning manner. Extensive experiments qualitatively and quantitatively show that our method outperforms SOTA face swapping methods in terms of both identity transfer and attribute preservation.
Lingyun Yu 0002, Hongtao Xie 0001, Chuanbin Liu 0001, Zhiguo Ding 0006, Quanwei Yang, Yongdong Zhang 0001
ACM Multimedia6
2022 REMOT: A Region-to-Whole Framework for Realistic Human Motion Transfer
abstract
Human Video Motion Transfer (HVMT) aims to, given an image of a source person, generate his/her video that imitates the motion of the driving person. Existing methods for HVMT mainly exploit Generative Adversarial Networks (GANs) to perform the warping operation based on the flow estimated from the source person image and each driving video frame. However, these methods always generate obvious artifacts due to the dramatic differences in poses, scales, and shifts between the source person and the driving person. To overcome these challenges, this paper presents a novel REgion-to-whole human MOtion Transfer (REMOT) framework based on GANs. To generate realistic motions, the REMOT adopts a progressive generation paradigm: it first generates each body part in the driving pose without flow-based warping, then composites all parts into a complete person of the driving motion. Moreover, to preserve the natural global appearance, we design a Global Alignment Module to align the scale and position of the source person with those of the driving person based on their layouts. Furthermore, we propose a Texture Alignment Module to keep each part of the person aligned according to the similarity of the texture. Finally, through extensive quantitative and qualitative experiments, our REMOT achieves state-of-the-art results on two public benchmarks.
Quanwei Yang, Xinchen Liu, Wu Liu 0005, Hongtao Xie 0001, Xiaoyan Gu 0001, Lingyun Yu 0002, Yongdong Zhang 0001
ACM Multimedia1