Lingyun Yu 0002

dblp:47/3963-2 · DBLP profile ↗
← Back
40ranked-venue papers
9as first author
29since 2021 · last 2026
0000-0001-6403-761XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 28 · 7 first-author · 18 since 2021Artificial intelligence and machine learning · 12 · 2 first-author · 10 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 4 since 2021Security and privacy · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 SwapController: Toward Improving Identity and Attribute Control for Diffusion-Based Face Swapping
abstract
Face swapping efforts strive to achieve high-fidelity and well-controlled generation effects. Owing to the remarkable generative capabilities, diffusion models deliver promising high-fidelity solutions. However, their intrinsic stochastic properties complicate the accurate modeling of facial representations, introducing new challenges for identity and attribute consistency of the generated faces. In this paper, we introduce a novel diffusion-based face-swapping framework, named SwapController, which achieves high-fidelity generation via careful facial identity and attribute modeling. Specifically, our facial modeling mainly involves facial structure and facial texture. For structure modeling, 3D facial priors are leveraged to provide explicit structure supervision, enabling accurate head structure control. On this basis, two novel components are proposed to deeply mine facial textural representations from identity and attribute aspects. To enhance identity control, multi-grained source identity embeddings are obtained from various functional encoders to convey critical global identity and fine-grained identity details. To improve attribute modeling, identity-shifted attribute embeddings are derived by applying identity modulation to the most salient textural attribute features of the target face. Moreover, in line with the diffusion denoising characteristics, a timestep-aware identity optimization objective is introduced to optimize identity consistency guidelines and overall fidelity. Extensive experiments demonstrate the effectiveness of our SwapController in generating identity-consistent portrait images while faithfully preserving target attributes, which obtains a 98.32 ID Retrieval, exceeding the SOTA DiffSFSR by 7.32 $\uparrow$↑.
Lingyun Yu 0002, Quanwei Yang, Runxin Liu, Yongdong Zhang 0001, Hongtao Xie 0001
IEEE Trans. Vis. Comput. Graph.2
2025 IDseq: Decoupled and Sequentially Detecting and Grounding Multi-Modal Media Manipulation
abstract
Detecting and grounding multi-modal media manipulation aims to categorize the type and localize the region of manipulation for image-text pairs in both two modalities. Existing methods have not sufficiently explored the intrinsic properties of the manipulated images, which contain both forgery and content features, leading to inefficient utilization. To address this problem, we propose an Image-Driven Decoupled Sequential Framework (IDseq), designed to decouple image features and rationally integrate them to accomplish different sub-tasks effectively. Specifically, IDseq employs two specially designed disentangled losses to guide the disentangled learning of forgery and content features. To efficiently leverage these features, we propose a Decoupled Image Manipulation Decoder (DIMD) that processes image tasks within a decoupled schema. We mitigate their exclusive competition by separating the image tasks into forgery-relevant and content-relevant components and training them without gradient interaction. Additionally, we utilize content features enhanced by the proposed Manipulation Indicator Generator (MIG) for the text tasks, which provide the maximal visual information as a reference while eliminating interference from unverified image data. Extensive experiments show the superiority of our IDseq, where it notably outperforms SOTA methods on the fine-grained classification by 3.8% in mAP and the forgery face grounding by 8.7% in IoUmean, even 1.3% in F1 on the most challenging manipulated text grounding.
Runxin Liu, Jiaming Li 0017, Lingyun Yu 0002, Hongtao Xie 0001
AAAI4
2025 Forensic-MoE: Exploring Comprehensive Synthetic Image Detection Traces With Mixture of Experts
Mingqi Fang, Ziguang Li, Lingyun Yu 0002, Quanwei Yang, Hongtao Xie 0001, Yongdong Zhang 0001
ICCV3
2025 GestureHYDRA: Semantic Co-Speech Gesture Synthesis via Hybrid Modality Diffusion Transformer and Cascaded-Synchronized Retrieval-Augmented Generation
Quanwei Yang, Luying Huang, Kaisiyuan Wang, Jiazhi Guan, Shengyi He, Fengguo Li, Hang Zhou 0009, Lingyun Yu 0002, Haocheng Feng, Hongtao Xie 0001
ICCV8
2025 Proactive Deepfake Detection via Self-Verifiable Semantic Watermarking
abstract
Malicious Deepfakes pose serious security risks by producing highly realistic forged faces. While numerous countermeasures have been developed to train binary Deepfake classifiers, their limited generalization capacity restricts practical deployment. To proactively defend against Deepfakes, we propose SVS-WM, a Self-Verifiable Semantic Watermarking strategy. The core idea behind SVS-WM is to embed pairs of correlated watermarks within facial semantics, leveraging the inherent fragility of these features, i.e., any semantic modification will disrupt the watermark correlation, thereby enabling robust Deepfake detection. SVS-WM employs a facial semantic disentanglement and reconstruction network, allowing semi-fragile watermarks to be embedded concurrently across multiple semantic levels, including identity and multi-levels of attributes. Specifically, pairs of pseudo-random noise watermarks are adaptively injected into facial attribute and identity features. During propagation stage, the protected image may encounter identity or facial attributes manipulations, we then detect Deepfakes by verifying the correlation result between the decoded attribute watermark and the extracted identity vector. This unique cross-verification mechanism enables authentication without requiring original reference watermark, thereby realizing blind Deepfake detection. Extensive experiments validate the effectiveness of our approach, achieving an average detection accuracy of 98.19% across diverse Deepfake manipulations.
Peiqi Jiang, Bohan Lei, Lingyun Yu 0002, Zhineng Chen, Hongtao Xie 0001, Yongdong Zhang 0001
ACM Multimedia4
2025 THGS: Lifelike Talking Human Avatar Synthesis From Monocular Video Via 3D Gaussian Splatting
abstract
Abstract Despite the remarkable progress in 3D talking head generation, directly generating 3D talking human avatars still suffers from rigid facial expressions, distorted hand textures and out‐of‐sync lip movements. In this paper, we extend speaker‐specific talking head generation task to talking human avatar synthesis and propose a novel pipeline, THGS, that animates lifelike Talking Human avatars using 3D Gaussian Splatting (3DGS). Given speech audio, expression and body poses as input, THGS effectively overcomes the limitations of 3DGS human re‐construction methods in capturing expressive dynamics, such as mouth movements, facial expressions and hand gestures, from a short monocular video. Firstly, we introduce a simple yet effective Learnable Expression Blendshapes (LEB) for facial dynamics re‐construction, where subtle facial dynamics can be generated by linearly combining the static head model and expression blendshapes. Secondly, a Spatial Audio Attention Module (SAAM) is proposed for lip‐synced mouth movement animation, building connections between speech audio and mouth Gaussian movements. Thirdly, we employ a body pose, expression and skinning weights joint optimization strategy to optimize these parameters on the fly, which aligns hand movements and expressions better with video input. Experimental results demonstrate that THGS can achieve high‐fidelity 3D talking human avatar animation at 150+ fps on a web‐based rendering system, improving the requirements of real‐time applications. Our project page is at https://sora158.github.io/THGS.github.io/ .
Lingyun Yu 0002, Quanwei Yang, Aihua Zheng, Hongtao Xie 0001
Comput. Graph. Forum2
2025 TalkingAvatar: Learning 3D talking human avatar via NeRF
Lingyun Yu 0002, Chuanbin Liu 0001, Wu Liu 0005, Quanwei Yang, Meng Shao
Neurocomputing1
2025 A Detail-Aware Transformer to Generalizable Face Forgery Detection
abstract
Generalisable face forgery detectors strive to detect forgeries generated by unseen manipulations. Recently advanced detection methods have managed to capture subtle blending traces, but their neglect of the diversity of blending traces in different regions leads to limited generalization. Towards this, transformer with global receptive fields and dynamic weight mechanism is a promising solution, but vanilla transformer is weak at capturing subtle blending traces. In this paper, we propose a novel Detail-Aware Transformer (DAT) able to focus on both diverse and subtle blending traces caused by inconsistencies in the low-level image details. The intrinsic multi-head self-attention mechanism of the transformer allows our DAT to adaptively capture diverse blending traces in different regions. Furthermore, we improve the transformer’s capability of capturing subtle blending traces by two inference overhead-free measures,$i.e$., self-supervised pre-training based on patch augmentation and region-level contrastive learning. Specifically, the self-supervised pre-training encourages the model to focus on the inconsistencies in low-level image details through a patch number prediction task. The region-level contrastive learning employs a contrastive loss on representations of regions with different low-level details to further improve the transformer’s ability to handle subtle blending traces. Extensive experiments show that our method substantially improves the generalization performance and outperforms the state-of-the-art methods on CDF, DFDC, DFDCP, FFIW, and WildDeepfake datasets.
Jiaming Li 0017, Lingyun Yu 0002, Runxin Liu, Hongtao Xie 0001
IEEE Trans. Circuits Syst. Video Technol.2
2025 Distilling Multi-Level Semantic Cues Across Multi-Modalities for Face Forgery Detection
abstract
Existing face forgery detection methods attempt to identify low-level forgery artifacts (e.g., blending boundary, flickering) in spatial-temporal domains or high-level semantic inconsistencies (e.g., abnormal lip movements) between visual-auditory modalities for generalized face forgery detection. However, they still suffer from significant performance degradation when dealing with out-of-domain artifacts, as they only consider single semantic mode inconsistencies, but ignore the complementarity of forgery traces at different levels and different modalities. In this paper, we propose a novel Multi-modal Multi-level Semantic Cues Distillation Detection framework that adopts the teacher-student protocol to focus on both spatial-temporal artifacts and visual-auditory incoherence to capture multi-level semantic cues. Specifically, our framework primarily comprises the Spatial-Temporal Pattern Learning module and the Visual-Auditory Consistency Modeling module. The Spatial-Temporal Pattern Learning module employs a mask-reconstruction strategy, in which the student network learns diverse spatial-temporal patterns from a pixel-wise teacher network to capture low-level forgery artifacts. The Visual-Auditory Consistency Modeling module is designed to enhance the student network’s ability to identify high-level semantic irregularities, with a visual-auditory consistency modeling expert serving as a guide. Furthermore, a novel Real-Similarity loss is proposed to enhance the proximity of real faces in feature space without explicitly penalizing the distance from manipulated faces, which prevents the overfitting in particular manipulation methods and improves the generalization capability. Extensive experiments show that our method substantially improves the generalization and robustness performance. Particularly, our approach outperforms the SOTA detector by 1.4% in generalization performance on DFDC with large domain gaps, and by 2.0% in the robustness evaluation on the FF++ dataset under various extreme settings. Our code is available athttps://github.com/TianXie834/M2SD.
Lingyun Yu 0002, Chuanbin Liu 0001, Guoqing Jin, Zhiguo Ding 0006, Hongtao Xie 0001
IEEE Trans. Circuits Syst. Video Technol.1
2025 High Fidelity Face Swapping via Facial Texture and Structure Consistency Mining
Lingyun Yu 0002, Quanwei Yang, Meng Shao, Hongtao Xie 0001
IEEE Trans. Multim.2
2024 DiffAM: Diffusion-Based Adversarial Makeup Transfer for Facial Privacy Protection
abstract
With the rapid development of face recognition (FR) systems, the privacy of face images on social media is facing severe challenges due to the abuse of unauthorized FR systems. Some studies utilize adversarial attack techniques to defend against malicious FR systems by generating adversarial examples. However, the generated adversarial examples, i.e., the protected face images, tend to suffer from sub-par visual quality and low transferability. In this paper, we propose a novel face protection approach, dubbed DiffAM, which leverages the powerful generative ability of diffusion models to generate high-quality protected face images with adversarial makeup transferred from reference images. To be specific, we first introduce a makeup removal module to generate non-makeup images utilizing a fine-tuned diffusion model with guidance of textual prompts in CLIP space. As the inverse process of makeup transfer, makeup removal can make it easier to establish the deterministic relationship between makeup domain and non-makeup domain regardless of elaborate text prompts. Then, with this relationship, a CLIP-based makeup loss along with an ensemble attack strategy is introduced to jointly guide the direction of adversarial makeup domain, achieving the generation of protected face images with natural-looking makeup and high black-box transferability. Extensive experiments demonstrate that DiffAM achieves higher visual quality and attack success rates with a gain of 12.98% under black-box setting compared with the state of the arts. The code will be available at https://github.com/HansSunYIDiffAM.
Lingyun Yu 0002, Hongtao Xie 0001, Jiaming Li 0017, Yongdong Zhang 0001
CVPR2
2024 Control-Talker: A Rapid-Customization Talking Head Generation Method for Multi-Condition Control and High-Texture Enhancement
abstract
In recent years, the field of talking head generation has made significant strides. However, the need for substantial computational resources for model training, coupled with a scarcity of high-quality video data, poses challenges for the rapid customization of model to specific individual. Additionally, existing models usually only support single-modal control, lacking the ability to generate vivid facial expressions and controllable head poses based on multiple conditions such as audio, video, etc. These limitations restricts the models' widespread application. In this paper, we introduce a two-stage method called Control-Talker to achieve rapid customization of identity in talking head model and high-quality generation based on multimodal conditions. Specifically, we divide the training process into two stages: prior learning stage and identity rapid-customization stage. 1) In the prior learning stage, we leverage a diffusion-based model pre-trained on the high-quality image dataset to acquire a robust controllable facial prior. Meanwhile, we innovatively propose a high-frequency ControlNet structure to enhance the fidelity of the synthesized results. This structure adeptly extracts a high-frequency feature map from the source image, serving as a facial texture prior, thereby excellently preserving facial texture of the source image. 2) In the identity rapid-customization stage, the identity is fixed by fine-tuning the U-Net part of the diffusion model on merely several images of a specific individual. The entire fine-tuning process for identity customization can be completed within approximately ten minutes, thereby significantly reducing training costs. Further, we propose a unified driving method for both audio and video, enabling the model to precisely control expressions, poses, and lighting under multi conditions. Extensive experiments and visual results demonstrate that our method outperforms other state-of-the-art models. Additionally, our model demonstrates reduced training costs and lower data requirements.
Yiding Li, Lingyun Yu 0002, Li Wang 0154, Hongtao Xie 0001
ACM Multimedia2
2024 ShowMaker: Creating High-Fidelity 2D Human Video via Fine-Grained Diffusion Modeling
abstract
Although significant progress has been made in human video generation, most previous studies focus on either human facial animation or full-body animation, which cannot be directly applied to produce realistic conversational human videos with frequent hand gestures and various facial movements simultaneously. To address these limitations, we propose a 2D human video generation framework, named ShowMaker, capable of generating high-fidelity half-body conversational videos via fine-grained diffusion modeling. We leverage dual-stream diffusion models as the backbone of our framework and carefully design two novel components for crucial local regions (i.e., hands and face) that can be easily integrated into our backbone. Specifically, to handle the challenging hand generation caused by sparse motion guidance, we propose a novel Key Point-based Fine-grained Hand Modeling module by amplifying positional information from raw hand key points and constructing a corresponding key point-based codebook. Moreover, to restore richer facial details in generated results, we introduce a Face Recapture module, which extracts facial texture features and global identity features from the aligned human face and integrates them into the diffusion process for face enhancement. Extensive quantitative and qualitative experiments demonstrate the superior visual quality and temporal consistency of our method.
Quanwei Yang, Jiazhi Guan, Kaisiyuan Wang, Lingyun Yu 0002, Wenqing Chu, Hang Zhou 0009, ZhiQiang Feng, Haocheng Feng, Errui Ding, Jingdong Wang 0001, Hongtao Xie 0001
NeurIPS4
2024 Symmetrical Siamese Network for pose-guided person synthesis
Quanwei Yang, Lingyun Yu 0002, Yun Song, Meng Shao, Guoqing Jin, Hongtao Xie 0001
Comput. Vis. Image Underst.2
2024 Generalizable Speech Spoofing Detection Against Silence Trimming With Data Augmentation and Multi-Task Meta-Learning
abstract
A major difficulty in speech spoofing detection lies in improving the generalization ability to detect unknown forgery methods. However, most previous methods do not consider the interference of silence information on the generalization performance of speech spoofing detection. Notably, we experimentally observe that the generalization performance of existing methods drops sharply when silence segments are trimmed. This indicates that previous works have two problems: a) they do not remove the interference of silence and over-rely on silence information, and b) they lack the ability to uncover general forgery traces in utterance segments. To solve the above two problems, we propose a novel Silence-Agnostic Speech Spoofing Detection (SASSD) framework. To be specific, unlike previous methods trained on speech samples with silence information, we completely remove the leading and trailing silence segments from all speech samples to eliminate the interference of silence and focus on utterance information. Meanwhile, to uncover general forgery traces in utterance segments and improve the generalization ability, we view speech spoofing detection as a domain generalization problem and employ meta-learning to simulate the actual domain shift scenarios, which can reduce overfitting to specific forgery methods. In addition, to improve the domain generalization of meta-learning, a novel data augmentation method named ShuffleMix is proposed. Unlike previous methods that only consider inter-speech patterns, our method additionally introduces an intra-speech augmentation technique, which performs enhancements within a single speech and across multiple speech to generate more diverse forged samples. Extensive experiments show that our method achieves SOTA on the ASVspoof 2019LA dataset. In particular, our method achieves 0.231% EER and 2.529% EER on the original dataset with silence information and the silence-trimmed dataset, respectively.
Li Wang 0154, Lingyun Yu 0002, Yongdong Zhang 0001, Hongtao Xie 0001
IEEE ACM Trans. Audio Speech Lang. Process.2
2024 STIDNet: Identity-Aware Face Forgery Detection With Spatiotemporal Knowledge Distillation
abstract
The impressive development of facial manipulation techniques has raised severe public concerns. Identity-aware methods, especially suitable for protecting celebrities, are seen as one of promising face forgery detection approaches with additional reference video. However, without in-depth observation of fake video’s characteristics, most existing identity-aware algorithms are just naive imitation of face verification model and fail to exploit discriminative information. In this article, we argue that it is necessary to take both spatial and temporal perspectives into consideration for adequate inconsistency clues and propose a novel forgery detector named SpatioTemporal IDentity network (STIDNet). To effectively capture heterogeneous spatiotemporal information in a unified formulation, our STIDNet is following a knowledge distillation architecture that the student identity extractor receives supervision from a spatial information encoder (SIE) and a temporal information encoder (TIE) through multiteacher training. Specifically, a regional sensitive identity modelling paradigm is proposed in SIE by introducing facial blending augmentation but with uniform identity label, thus encourage model to focus on spatial discriminative region like outer face. Meanwhile, considering the strong temporal correlation between audio and talking face video, our TIE is devised in a cross-modal pattern that the audio information is introduced to supervise model exploiting temporal personalized movements. Benefit from knowledge transfer from SIE and TIE, STIDNet is able to capture individual’s essential spatiotemporal identity attributes and sensitive to even subtle identity deviation caused by manipulation. Extensive experiments indicate the superiority of our STIDNet compared with previous works. Moreover, we also demonstrate STIDNet is more suitable for real-world implementation in terms of model complexity and reference set size.
Mingqi Fang, Lingyun Yu 0002, Hongtao Xie 0001, Qingfeng Tan, Zhiyuan Tan 0001, Amir Hussain 0001, Zezheng Wang 0002, Zhihong Tian 0001
IEEE Trans. Comput. Soc. Syst.2
2024 Exploring Bi-Level Inconsistency via Blended Images for Generalizable Face Forgery Detection
abstract
The challenge of generalization in face forgery detection has become increasingly prominent as manipulation techniques continue to evolve. Although recent image blending-based methods have demonstrated remarkable potential, they often encounter a significant performance drop when applied to datasets exhibiting significant domain gaps. This limitation stems from the exclusive reliance of prior methods on blending unaltered faces with various augmentations to produce common artifacts, which ignores the inherent characteristics of the forged regions. To fully exploit the potential of image blending-based methods for generalizable Deepfake detection, we propose a novel image synthesis framework called Bi-Level Inconsistency Generator (Bi-LIG) to introduce bi-level inconsistency in the synthesized images. Specifically, Bi-LIG generates synthetic images by blending source and target images from both pristine and forged image sets, introducing a) Extrinsic-Inconsistency between real and pseudo-forged regions, and b) Inherent-Inconsistency between real and manipulated areas. In this way, Bi-LIG creates a diverse synthesized image set and establishes a generalizable training domain. Furthermore, we propose a novel face forgery detection network named Token Consistency Constrained Vision Transformer, in which two modules are developed based on patch consistency learning. Firstly, a Patch Token Contrast module is employed to learn the bi-level patch inconsistencies. Secondly, a Progressive Patch Token Assemble module is adopted to aggregate local patch relations and enhance the inconsistency representations. Experimental results demonstrate the effectiveness and superiority of our method on both in-dataset and cross-dataset evaluations. Notably, our approach outperforms state-of-the-art methods by 5.09% and 10.15% on cross-dataset evaluations in DFDCp and DFDC, respectively.
Peiqi Jiang, Hongtao Xie 0001, Lingyun Yu 0002, Guoqing Jin, Yongdong Zhang 0001
IEEE Trans. Inf. Forensics Secur.3
2024 IEIRNet: Inconsistency Exploiting Based Identity Rectification for Face Forgery Detection
abstract
Face forgery detection has attracted much attention due to the ever-increasing social concerns caused by facial manipulation techniques. Recently, identity-based detection methods have made considerable progress, which is especially suitable in the celebrity protection scenario. However, they still suffer from two main limitations: (a) generic identity extractor is not specifically designed for forgery detection, leading to nonnegligibleIdentity Representation Biasto forged images. (b) existing methods only analyze the identity representation of each image individually, but ignores the query-reference interaction for inconsistency exploiting. To address these issues, a novelInconsistency Exploiting based Identity Rectification Network(IEIRNet) is proposed in this paper. Firstly, for the identity bias rectification, the IEIRNet follows an effective two-branches structure. Besides theGeneric Identity Extractor(GIE) branch, an essentialBias Diminishing Module(BDM) branch is proposed to eliminate the identity bias through a novelAttention-based Bias Rectification(ABR) component, accordingly acquiring the ultimate discriminative identity representation. Secondly, for query-reference inconsistency exploiting, anInconsistency Exploiting Module(IEM) is applied in IEIRNet to comprehensively exploit the inconsistency clues from both spatial and channel perspectives. In the spatial aspect, an innovative region-aware kernel is derived to activate the local region inconsistency with deep spatial interaction. Afterward in the channel aspect, a coattention mechanism is utilized to model the channel interaction meticulously, and accordingly highlight the channel-wise inconsistency with adaptive weight assignment and channel-wise dropout. Our IEIRNet has shown effectiveness and superiority in various generalization and robustness experiments.
Mingqi Fang, Lingyun Yu 0002, Yun Song, Yongdong Zhang 0001, Hongtao Xie 0001
IEEE Trans. Multim.2
2023 RAIRNet: Region-Aware Identity Rectification for Face Forgery Detection
abstract
The malicious usage of facial manipulation techniques boosts the desire of face forgery detection research. Recently, identity-based approaches have attracted much attention due to the effective observation of identity inconsistency. However, there are still several nonnegligible problems: (1) generic identity extractor is totally trained on real images, leading to enormous identity representation bias during processing forged content; (2) the identity information of forged image is hybrid and presents regional distribution, while the single global identity feature is hard to reflect this local identity inconsistency. To solve the above problems, in this paper a novel Region-Aware Identity Rectification Network (RAIRNet) is proposed to effectively rectify the identity bias and adaptively exploit the inconsistency local region. Firstly, for the identity bias problem, our RAIRNet is devised in a two-branch architecture, which consists of a Generic Identity Extractor (GIE) branch and a Bias Diminishing Module (BDM) branch. The BDM branch is designed to rectify the bias introduced by GIE branch through a prototype-based training schema. This two-branch architecture effectively promotes model to adapt to forged content while maintaining the focus on identity space. Secondly, for local identity inconsistency exploiting, a novel Meta Identity Filter Generator (MIFG) is devised in a meta-learning way to generate the region-aware filter based on identity prior. This region-aware filter can adaptively exploit the local inconsistency clues and activate the discriminative local region. Moreover, to balance the local-global information and highlight the forensic clues, an Adaptive Weight Assignment Mechanism (AWAM) is proposed to assign adaptive importance weight to two branches. Extensive experiments on various datasets show the superiority of our RAIRNet. In particular, on the challenging DFDCp dataset, our approach outperforms previous binary-based and identity-based methods by 10.3% and 5.5% respectively.
Mingqi Fang, Lingyun Yu 0002, Hongtao Xie 0001, Junqiang Wu, Zezheng Wang 0002, Yongdong Zhang 0001
ACM Multimedia2
2023 High Fidelity Face Swapping via Semantics Disentanglement and Structure Enhancement
abstract
In this paper, we propose a novel Semantics and Structure-aware face Swapping framework (S2Swap) that exploits semantics disentanglement and structure enhancement for high fidelity face generation. Different from previous methods that either 1) suffer from degraded generation fidelity due to insufficient identity-attributes disentanglement or 2) neglect the importance of structure information for identity consistency, our approach can achieve local facial semantics disentanglement beyond global identity while boosting identity consistency through structure enhancement. Specifically, to achieve identity-attributes disentanglement, our S2Swap is designed from global-local perspectives. Firstly, an Oriented Identity Transfer module is proposed to globally disentangle target identity and attributes under global identity semantics prior. Such global disentanglement enables source identity transfer to the individual target identity. Secondly, a Local Semantics Disentanglement module is devised to disentangle local identity and identity-irrelevant facial semantics, providing local semantic compensation for the global counterpart. Moreover, to boost identity consistency, a Structure-Aware Head Modeling module is introduced to provide the desired face structure enhancement through an intuitive face sketch. Finally, considering the identity-attributes trade-off, we adaptively integrate semantics and structure information in a self-learning manner. Extensive experiments qualitatively and quantitatively show that our method outperforms SOTA face swapping methods in terms of both identity transfer and attribute preservation.
Lingyun Yu 0002, Hongtao Xie 0001, Chuanbin Liu 0001, Zhiguo Ding 0006, Quanwei Yang, Yongdong Zhang 0001
ACM Multimedia2
2023 Discriminative Feature Mining Based on Frequency Information and Metric Learning for Face Forgery Detection
abstract
Face forgery detection has received considerable attention due to security concerns about abnormal faces generated by face forgery technology. While recent researches have made prominent progress, they still suffer from two limitations: a) the learned features supervised by softmax loss are insufficiently discriminative, since the softmax loss fails to explicitly boost inter-class separability and intra-class compactness; b) hand-crafted features are unable to effectively mine forgery patterns from frequency domain. To address the two problems, this paper proposes a novel frequency-aware discriminative feature learning framework. Specifically, we design an innovative single-center loss which compresses mere intra-class variations of natural faces while encouraging inter-class differences between natural and manipulated faces in the embedding space. Supervised by such a loss, more discriminative features can be learned with less optimization difficulty. As for frequency-related features, a frequency feature adaptively generated module is developed to capture frequency clues in a data-driven manner. Besides, to better fuse the features of both RGB domain and frequency domain, this paper devises a fusion module based on positional correlation of features. The effectiveness and superiority of our framework have been proved by extensive experiments and our approach achieves state-of-the-art performance in both in-dataset and cross-dataset evaluation.
Jiaming Li 0017, Hongtao Xie 0001, Lingyun Yu 0002, Xingyu Gao 0001, Yongdong Zhang 0001
IEEE Trans. Knowl. Data Eng.3
2023 Constructing Spatio-Temporal Graphs for Face Forgery Detection
abstract
Recently, advanced development of facial manipulation techniques threatens web information security, thus, face forgery detection attracts a lot of attention. It is clear that both spatial and temporal information of facial videos contains the crucial manipulation traces, which are inevitably created during the generation process. However, most existing face forgery detectors only focus on the spatial artifacts or the temporal incoherence, and they are struggling to learn a significant and general kind of representations for manipulated facial videos. In this work, we propose to construct spatial-temporal graphs for fake videos to capture the spatial inconsistency and the temporal incoherence at the same time. To model the spatial-temporal relationship among the graph nodes, a novel forgery detector named Spatio-Temporal Graph Network (STGN) is proposed, which contains two kinds of graph-convolution-based units, the Spatial Relation Graph Unit (SRGU) and the Temporal Attention Graph Unit (TAGU). To exploit spatial information, the SRGU models the inconsistency between each pair of patches in the same frame, instead of focusing on the low-level local spatial artifacts which are vulnerable to samples created by unseen manipulation methods. And, the TAGU is proposed to model the long-distance temporal relation among the patches at the same spatial position in different frames with a graph attention mechanism based on the inter-node similarity. With the SRGU and the TAGU, our STGN can combine the discriminative power of spatial inconsistency and the generalization capacity of temporal incoherence for face forgery detection. Our STGN achieves state-of-the-art performances on several popular forgery detection datasets. Extensive experiments demonstrate both the superiority of our STGN on intra manipulation evaluation and the effectiveness for new sorts of face forgery videos on cross manipulation evaluation.
Zhihua Shang, Hongtao Xie 0001, Lingyun Yu 0002, Zhengjun Zha, Yongdong Zhang 0001
ACM Trans. Web3
2022 Wavelet-enhanced Weakly Supervised Local Feature Learning for Face Forgery Detection
abstract
Face forgery detection is getting increasing attention due to the security threats caused by forged faces. Recently, local patch-based approaches have achieved sound achievements due to effective attention to local details. However, there are still unignorable problems: a) local feature learning requires patch-level labels to circumvent label noise, which is not practical in real-world scenarios; b) the commonly used DCT (FFT) transform loses all spatial information, which brings difficulty in handling local details. To compensate for such limitations, a novel wavelet-enhanced weakly supervised local feature learning framework is proposed in this paper. Specifically, to supervise the learning of local features with only image-level labels, two modules are devised based on the idea of multi-instance learning: local relation constraint module (LRCM) and category knowledge-guided local feature aggregation module (CKLFA). LRCM constrains the maximum distance between local features of forged face images greater than that of real face images. CKLFA adaptively aggregates local features based on their correlation to global embedding containing global category information. Combining these two modules, the network is encouraged to learn discriminative local features supervised only by image-level labels. Besides, a multi-level wavelet-powered feature enhancement module is developed to promote the network mining local forgery artifacts from spatio-frequency domain, which is beneficial to learning discriminative local features. Extensive experiments show that our approach outperforms previous state-of-the-art methods when only image-level labels are available and achieves comparable or even better performance than counterparts using patch-level labels.
Jiaming Li 0017, Hongtao Xie 0001, Lingyun Yu 0002, Yongdong Zhang 0001
ACM Multimedia3
2022 REMOT: A Region-to-Whole Framework for Realistic Human Motion Transfer
abstract
Human Video Motion Transfer (HVMT) aims to, given an image of a source person, generate his/her video that imitates the motion of the driving person. Existing methods for HVMT mainly exploit Generative Adversarial Networks (GANs) to perform the warping operation based on the flow estimated from the source person image and each driving video frame. However, these methods always generate obvious artifacts due to the dramatic differences in poses, scales, and shifts between the source person and the driving person. To overcome these challenges, this paper presents a novel REgion-to-whole human MOtion Transfer (REMOT) framework based on GANs. To generate realistic motions, the REMOT adopts a progressive generation paradigm: it first generates each body part in the driving pose without flow-based warping, then composites all parts into a complete person of the driving motion. Moreover, to preserve the natural global appearance, we design a Global Alignment Module to align the scale and position of the source person with those of the driving person based on their layouts. Furthermore, we propose a Texture Alignment Module to keep each part of the person aligned according to the similarity of the texture. Finally, through extensive quantitative and qualitative experiments, our REMOT achieves state-of-the-art results on two public benchmarks.
Quanwei Yang, Xinchen Liu, Wu Liu 0005, Hongtao Xie 0001, Xiaoyan Gu 0001, Lingyun Yu 0002, Yongdong Zhang 0001
ACM Multimedia6
2022 Attention-guided transformation-invariant attack for black-box adversarial examples
abstract
With the development of media convergence, information acquisition is no longer limited to traditional media, such as newspapers and televisions, but more from digital media on the Internet, where media contents should be under supervision by platforms. At present, the media content analysis technology of Internet platforms relies on deep neural networks (DNNs). However, DNNs show vulnerability to adversarial examples, which results in security risks. Therefore, it is necessary to adequately study the internal mechanism of adversarial examples to build more effective supervision models. When coming to practical applications, supervision models are mostly faced with black-box attacks, where cross-model transferability of adversarial examples has attracted increasing attention. In this paper, to improve the transferability of adversarial examples, we propose an attention-guided transformation-invariant adversarial attack method, which incorporates an attention mechanism to disrupt the most distinctive features and simultaneously ensures adversarial attack invariance under different transformations. Specifically, we dynamically weight the latent features according to an attention mechanism and disrupt them accordingly. Meanwhile, considering the lack of semantics in low-level features, high-level semantics are introduced as spatial guidance to make low-level feature perturbations concentrate on the most discriminative regions. Moreover, since the attention heatmaps may vary significantly across different models, a transformation-invariant aggregated attack strategy is proposed to alleviate overfitting to the proxy model attention. Comprehensive experimental results show that the proposed method can significantly improve the transferability of adversarial examples.
Lingyun Yu 0002, Hongtao Xie 0001, Bo Wu 0018, Yongdong Zhang 0001
Int. J. Intell. Syst.3
2022 Dynamic-Aware Federated Learning for Face Forgery Video Detection
abstract
The spread of face forgery videos is a serious threat to information credibility, calling for effective detection algorithms to identify them. Most existing methods have assumed a shared or centralized training set. However, in practice, data may be distributed on devices of different enterprises that cannot be centralized to share due to security and privacy restrictions. In this article, we propose a Federated Learning face forgery detection framework to train a global model collaboratively while keeping data on local devices. In order to make the detection model more robust, we propose a novel Inconsistency-Capture module (ICM) to capture the dynamic inconsistencies between adjacent frames of face forgery videos. The ICM contains two parallel branches. The first branch takes the whole face of adjacent frames as input to calculate a global inconsistency representation. The second branch focuses only on the inter-frame variation of critical regions to capture the local inconsistency. To the best of our knowledge, this is the first work to apply federated learning to face forgery video detection, which is trained with decentralized data. Extensive experiments show that the proposed framework achieves competitive performance compared with existing methods that are trained with centralized data, with higher-level security and privacy guarantee.
Ziheng Hu, Hongtao Xie 0001, Lingyun Yu 0002, Xingyu Gao 0001, Zhihua Shang, Yongdong Zhang 0001
ACM Trans. Intell. Syst. Technol.3
2022 Multimodal Learning for Temporally Coherent Talking Face Generation With Articulator Synergy
abstract
Talking face generation is a demanding task to synthesize a high quality video with accurate lip synchronization and rhythmic head motion. Any subtle artifacts could be sensitively captured by humans and lead to poor visual quality. Existing methods tend to employ a conditional generation solution, which introduces facial landmarks to bridge the input information and output videos. However, these methods always suffer from unrealistic facial animations, because 1) they only take single-mode input, but ignore the complementarity of multimodal inputs for lip-sync improvement; 2) they only explore lip movements, but ignore the articulator synergy between lips and jaw; 3) they generate each video frame in a temporal-independent way, but ignore the temporal continuity among the entire video. To address these limitations, in this paper, we present a novel method to generate realistic and temporally coherent talking heads by considering multimodal inputs, articulator synergy, inter-frame consistency and intra-frame consistency. Firstly, for landmark prediction, a novel Multiple Synergy Network (MSN) is proposed to improve the accuracy of landmark prediction by incorporating multimodal inputs (i.e., audio and text inputs). Besides, instead of merely considering lip landmarks, we also explore the jaw movements to ensure articulator synergy among lips and jaw. Secondly, for realistic video generation, a Video Consistency Network (VCN) is proposed conditioned on the predicted landmarks. In VCN, the optical flow is adopted to model the temporal continuity between frames to ensure inter-frame consistency. Meanwhile, a mouth generation branch is proposed to enhance mouth texture and the corresponding mouth mask is employed to ensure intra-frame consistency between the mouth area and the others. Extensive experiments demonstrate that our approach exhibits excellent superiority on lip-sync and can generate photo-realistic facial animations. Project is available at http://imcc.ustc.edu.cn/project/tfgen/.
Lingyun Yu 0002, Hongtao Xie 0001, Yongdong Zhang 0001
IEEE Trans. Multim.1
2021 PRRNet: Pixel-Region relation network for face forgery detection
Zhihua Shang, Hongtao Xie 0001, Zhengjun Zha, Lingyun Yu 0002, Yan Li 0068, Yongdong Zhang 0001
Pattern Recognit.4
2021 Multimodal Inputs Driven Talking Face Generation With Spatial-Temporal Dependency
abstract
Given an arbitrary speech clip or text information as input, the proposed work aims to generate a talking face video with accurate lip synchronization. Existing works mainly have three limitations. (1) A single-modal learning is adopted with either audio or text as input, hence it lacks the complementarity ofmultimodal inputs. (2) Each frame is generated independently, hence it ignores thetemporal dependencybetween consecutive frames. (3) Each face image is generated by the traditional convolution neural network (CNN) with a local receptive field, hence it cannot effectively capture thespatial dependencywithin internal representations of face images. To overcome these problems above, we decompose the talking face generation task into two steps: mouth landmarks prediction and video synthesis. First, a multimodal learning method is proposed to generate accurate mouth landmarks with multimedia inputs (both text and audio). Second, a network named Face2Vid is proposed to generate video frames conditioned on the predicted mouth landmarks. In Face2Vid, the optical flow is employed to model the temporal dependency between frames, meanwhile, a self-attention mechanism is introduced to model the spatial dependency across image regions. Extensive experiments demonstrate that our approach can generate photo-realistic video frames with the background, and exhibit the superiorities on accurate synchronization of lip movements and smooth transition of facial movements.
Lingyun Yu 0002, Jun Yu 0001, Qiang Ling 0001
IEEE Trans. Circuits Syst. Video Technol.1
2020 Filtration and Distillation: Enhancing Region Attention for Fine-Grained Visual Categorization
abstract
Delicate attention of the discriminative regions plays a critical role in Fine-Grained Visual Categorization (FGVC). Unfortunately, most of the existing attention models perform poorly in FGVC, due to the pivotal limitations in discriminative regions proposing and region-based feature learning. 1) The discriminative regions are predominantly located based on the filter responses over the images, which can not be directly optimized with a performance metric. 2) Existing methods train the region-based feature extractor as a one-hot classification task individually, while neglecting the knowledge from the entire object. To address the above issues, in this paper, we propose a novel “Filtration and Distillation Learning” (FDL) model to enhance the region attention of discriminate parts for FGVC. Firstly, a Filtration Learning (FL) method is put forward for discriminative part regions proposing based on the matchability between proposing and predicting. Specifically, we utilize the proposing-predicting matchability as the performance metric of Region Proposal Network (RPN), thus enable a direct optimization of RPN to filtrate most discriminative regions. Go in detail, the object-based feature learning and region-based feature learning are formulated as “teacher” and “student”, which can furnish better supervision for region-based feature learning. Accordingly, our FDL can enhance the region attention effectively, and the overall framework can be trained end-to-end without neither object nor parts annotations. Extensive experiments verify that FDL yields state-of-the-art performance under the same backbone with the most competitive approaches on several FGVC tasks.
Chuanbin Liu 0001, Hongtao Xie 0001, Zhengjun Zha, Lingfeng Ma, Lingyun Yu 0002, Yongdong Zhang 0001
AAAI5
2020 Bidirectional Attention-Recognition Model for Fine-Grained Object Classification
abstract
Fine-grained object classification (FGOC) is a challenging research topic in multimedia computing with machine learning, which faces two pivotal conundrums: focusing attention on the discriminate part regions, and then processing recognition with the part-based features. Existing approaches generally adopt a unidirectional two-step structure, that first locate the discriminate parts and then recognize the part-based features. However, they neglect the truth that part localization and feature recognition can be reinforced in a bidirectional process. In this paper, we propose a novel bidirectional attention-recognition model (BARM) to actualize the bidirectional reinforcement for FGOC. The proposed BARM consists of one attention agent for discriminate part regions proposing and one recognition agent for feature extraction and recognition. Meanwhile, a feedback flow is creatively established to optimize the attention agent directly by recognition agent. Therefore, in BARM the attention agent and the recognition agent can reinforce each other in a bidirectional way and the overall framework can be trained end-to-end without neither object nor parts annotations. Moreover, a novel Multiple Random Erasing data augmentation is proposed, and it exhibits impressive pertinency and superiority for FGOC. Conducted on several extensive FGOC benchmarks, BARM outperforms the present state-of-the-art methods in classification accuracy. Furthermore, BARM exhibits a clear interpretability and keeps consistent with the human perception in visualization experiments.
Chuanbin Liu 0001, Hongtao Xie 0001, Zhengjun Zha, Lingyun Yu 0002, Zhineng Chen, Yongdong Zhang 0001
IEEE Trans. Multim.4
2019 Mining Audio, Text and Visual Information for Talking Face Generation
abstract
Providing methods to support audio-visual interaction with growing volumes of video data is an increasingly important challenge for data mining. To this end, there has been some success in speech-driven lip motion generation or talking face generation. Among them, talking face generation aims to generate realistic talking heads synchronized with the audio or text input. This task requires mining the relationship between audio signal/text and lip-sync video frames and ensures the temporal continuity between frames. Due to the issues such as polysemy, ambiguity, and fuzziness of sentences, creating visual images with lip synchronization is still challenging. To overcome the problems above, we present a data-mining framework to learn the synchronous pattern between different channels from large recorded audio/text dataset and visual dataset, and apply it to generate realistic talking face animations. Specifically, we decompose this task into two steps: mouth landmarks prediction and video synthesis. First, a multimodal learning method is proposed to generate accurate mouth landmarks with multimedia inputs (both text and audio). Second, a network named Face2Vid is proposed to generate video frames conditioned on the predicted mouth landmarks. In Face2Vid, optical flow is employed to model the temporal dependency between frames, meanwhile, a self-attention mechanism is introduced to model the spatial dependency across image regions. Extensive experiments demonstrate that our method can generate realistic videos with background, and exhibit the superiorities on accurate synchronization of lip movements and smooth transition of facial movements.
Lingyun Yu 0002, Jun Yu 0001, Qiang Ling 0001
ICDM1
2019 Beauty Product Retrieval Based on Regional Maximum Activation of Convolutions with Generalized Attention
abstract
Beauty and Personal care product retrieval has attracted more and more research attention for its value in real life. However, suffering from data variants and complex background, this task has been very challenging. In this paper, we propose a novel Generalized-attention Regional Maximal Activation of Convolutions (GRMAC) descriptor which helps to generate image features for retrieval. This method introduces attention mechanism to reduce the influence of clustered background and highlight the target, and thus contributes to enhancing the effectiveness of features and boosting the retrieval performance. Different from other attention-based methods, our method supports adjusting mask with a hyperparameter p, which is more flexible and accurate in real application. To demonstrate its effectiveness, we conduct experiments on the dataset containing more than half million personal care products (Perfect-500K) and obtain remarkable results. Furthermore, we try to fuse multiple features from different models for more improvements. And finally, our team (USTC_NELSLIP) ranked 1st in the Grand Challenge of AI Meets Beauty in ACM Multimedia 2019 with a MAP score of 0.408614. Our code is available at: https://github.com/gniknoil/Perfect500K-Beauty-and-Personal-Care-Products-Retrieval-Challenge
Jun Yu 0001, Guochen Xie, Haonian Xie, Lingyun Yu 0002
ACM Multimedia5
2019 Deep Neural Network Based 3D Articulatory Movement Prediction Using Both Text and Audio Inputs
Lingyun Yu 0002, Jun Yu 0001, Qiang Ling 0001
MMM (1)1
2019 BLTRCNN-Based 3-D Articulatory Movement Prediction: Learning Articulatory Synchronicity From Both Text and Audio Inputs
abstract
Predicting articulatory movements from audio or text has diverse applications, such as speech visualization. Various approaches have been proposed to solve the acoustic-articulatory mapping problem. However, their precision is not high enough with only acoustic features available. Recently, deep neural network (DNN) has brought tremendous success in various fields, like speech recognition and image processing. To increase the accuracy, we propose a new network architecture for articulatory movement prediction with both text and audio inputs, called a bottleneck long-term recurrent convolutional neural network (BLTRCNN). To the best of our knowledge, it is the first time to predict articulatory movements based on DNN by fusing text and audio inputs. Our BLTRCNN consists of two networks. The first is the bottleneck network, generating a compact bottleneck features of text information for each frame independently. The second, including convolutional neural network, long short-term memory and skip connection, is called the long-term recurrent convolutional neural network (LTRCNN). LTRCNN is used for articulatory movement prediction when bottleneck features, acoustic features, and text features are integrated as inputs together. Experiments show that the proposed BLTRCNN achieves the state-of-the-art root-mean-square error (RMSE) 0.528 mm and the correlation coefficient 0.961. Moreover, we also demonstrate how text information complements acoustic features in this prediction task.
Lingyun Yu 0002, Jun Yu 0001, Qiang Ling 0001
IEEE Trans. Multim.1
2018 Synthesizing Photo-Realistic 3D Talking Head: Learning Lip Synchronicity and Emotion from Audio and Video
abstract
We propose a speech driven emotional 3D talking head. First, based on face prior knowledge, an individual head model is reconstructed by integrating the internal articulators from Magnetic Resonance Imaging slices with the appearance from visible images. Second, the blendshapes of each phoneme and the facial expression are synthesized by statistical learning. Third, visual coarticulation is modeled by training a deep neural network on parallel audio-visual data. Finally, the articulatory animations of continuous phonemes are fused by visual co articulation model and facial expressions to produce expressive speech synchronized animations. Experiments show that the system can not only distinguish the visual differences among phonemes in real-time, but also significantly increase the ability of human perception.
Jun Yu 0001, Lingyun Yu 0002
ICIP2
2018 Synthesizing 3D Acoustic-Articulatory Mapping Trajectories: Predicting Articulatory Movements by Long-Term Recurrent Convolutional Neural Network
abstract
Robust and accurate predicting of articulatory movements has various important applications, such as 3D articulatory animations and visual communication. Various approaches have been proposed to solve the acoustic-articulatory mapping problem. However, their precision is not high enough. Recently, deep neural network (DNN), especially convolutional neural network (CNN) and recurrent neural network (RNN), has brought tremendous success in speech recognition and synthesis. To increase the accuracy, we propose a new network architecture for acoustic-articulatory mapping, called long-term recurrent convolutional neural network (LTRCNN). The network consists of CNN, RNN and a skip connection. CNN can model the spectral correlation among acoustic features efficiently. RNN, like long short-term memory (LSTM), can learn the temporal context information from sequential data powerfully. Besides, skip connections can increase the input representation from different levels to preserve the feature information. Experiments show that LTRCNN achieves the state-of-the-art root-mean-squared error (RMSE) with 0.690 mm and the correlation coefficient with 0.949 in this prediction task.
Lingyun Yu 0002, Jun Yu 0001, Qiang Ling 0001
VCIP1
2017 A realistic 3D articulatory animation system for emotional visual pronunciation
Lingyun Yu 0002, Jun Yu 0001, Zengfu Wang
Multim. Tools Appl.1
2016 Facial video coding/decoding at ultra-low bit-rate: a 2D/3D model-based approach
Jun Yu 0001, Changwei Luo, Lingyun Yu 0002, Ling-yan Li, Zengfu Wang
Multim. Tools Appl.3
2015 Video Based Face Tracking and Animation
Changwei Luo, Jun Yu 0001, Zhigang Zheng, Lingyun Yu 0002, Zengfu Wang
ICIG (3)5