EDBT 2026 Demo / reviewers in the wild / expert
Yang Yu 0039
dblp:46/2181-39
· DBLP profile ↗
15ranked-venue papers
7as first author
14since 2021 · last 2025
0000-0002-5204-2929ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 6 first-author · 12 since 2021Computer networks · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Mining Generalized Multi-timescale Inconsistency for Detecting Deepfake Videos
Yang Yu 0039, Siyuan Yang 0001, Yu Ni, Yao Zhao 0001, Alex Chichung Kot |
Int. J. Comput. Vis. | 1 |
| 2025 | NPVForensics: Learning VA correlations in non-critical phoneme-viseme regions for deepfake detection
Yu Chen 0049, Yang Yu 0039, Haoliang Li, Wei Wang 0108, Yao Zhao 0001 |
Image Vis. Comput. | 2 |
| 2025 | FaceDefend: Copyright Protection to Prevent Face EmbezzleabstractWith the rapid evolution of deep learning and the advent of AI, the metaverse has emerged as a significant technology. Within the metaverse, diverse elements such as rich applications and realistic digital avatars provide users with immersive experiences, but it poses a series of security problems. Current research predominantly focuses on the data storage and transmission processes from the perspective of blockchain and the Internet of Things to achieve the protection of the metaverse. However, there exists a gap in security research on the digital avatar generation process. Given that digital avatars are the primary entities engaging in social activities within the metaverse and are crafted based on real face images, the virtual character can be generated easily by stealing the user’s face image and controlled to interact with others. In order to deal with the above problems, we propose a novel method to prevent the misuse of faces, which maintains the security of the metaverse by protecting facial data and thus preventing its misuse. We explore the common architecture of generative models and propose a defense method based on copyright protection to prevent face embezzling. Firstly, we utilize the copyright protection module to obtain copyright protection information. Secondly, we utilized the defense control module to ensure the representation of the protected images occurs errors in the latent space of the generation model. Therefore, the subsequent generation task output fails, which effectively protects the face data and prevents the generation of digital avatars. Furthermore, the results on public datasets and across multiple generative models present unnatural outputs, indicating the excellence of our defense and transfer capabilities. Yang Yu 0039, Yao Zhao 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | MCL: Multimodal Contrastive Learning for Deepfake DetectionabstractAdvancements in computer vision and deep learning have led to difficulty in distinguishing Deepfake and real videos. In particular, forgery audios are also generated to accompany fake videos and make them more realistic, which makes Deepfake detection more difficult. Existing Deepfake detection methods that use multimodal information ignore the representation gap between different modalities, resulting in limited performance. To address this problem, in this paper, a novel Deepfake detection method utilizing multimodal contrastive learning (MCL) is proposed to better explore intra-modal and cross-modal forgery clues. To reduce the cross-modal gap and explore multimodal forgery artifacts, a cross-modal contrastive learning strategy is designed to learn a compositional embedding from multimodal information, which facilitates pulling together representations across uni-modalities and multi-modalities. Moreover, to supplement the intra-frame forgery clues mining ability of the video network, the frame knowledge is distilled to the video network without adding additional computation. Specifically, to mine intra-modal clues, three modality features are first extracted from audio, frame and video, respectively. Secondly, the audio and frame features are separately composed with the video feature to derive two cross-modal representations. Subsequently, these cross-modal features are contrastive with the intra-modal features to reduce cross-modal gap. By jointly pulling together the unimodal and multimodal features through MCL, a more effective representation that contains intra-modal and cross-modal forgery artifacts can be learned. Finally, a noise-based feature augmentation (NFA) module is proposed to adaptively perturb the audio-visual feature and further improve generalization performance. Extensive experiments demonstrate that the proposed framework outperforms SOTA methods. Yang Yu 0039, Xiaolong Li 0001, Yao Zhao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | PVASS-MDD: Predictive Visual-Audio Alignment Self-Supervision for Multimodal Deepfake DetectionabstractDeepfake techniques can forge the visual or audio signals in the video, which leads to inconsistencies between visual and audio (VA) signals. Therefore, multimodal detection methods expose deepfake videos by extracting VA inconsistencies. Recently, deepfake technology has started VA collaborative forgery to obtain more realistic deepfake videos, which poses new challenges for extracting VA inconsistencies. Recent multimodal detection methods propose to first extract natural VA correspondences in real videos in a self-supervised manner, and then use the learned real correspondences as targets to guide the extraction of VA inconsistencies in the subsequent deepfake detection stage. However, the inherent VA relations are difficult to extract due to the modality gap, which leads to the limited auxiliary performance of the aforementioned self-supervised methods. In this paper, we propose Predictive Visual-audio Alignment Self-supervision for Multimodal Deepfake Detection (PVASS-MDD), which consists of PVASS auxiliary and MDD stages. In the PVASS auxiliary stage in real videos, we first devise a three-stream network to associate two augmented visual views with corresponding audio clues, leading to explore common VA correspondences based on cross-view learning. Secondly, we introduce a novel cross-modal predictive align module for eliminating VA gaps to provide inherent VA correspondences. In the MDD stage, we propose to the auxiliary loss to utilize the frozen PVASS network to align VA features of real videos, to better assist multimodal deepfake detector for capturing subtle VA inconsistencies. We conduct extensive experiments on existing widely used and latest multimodal deepfake datasets. Our method obtains a significant performance improvement compared to state-of-the-art methods. Yang Yu 0039, Siyuan Yang 0001, Yao Zhao 0001, Alex Chichung Kot |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Narrowing Domain Gaps With Bridging Samples for Generalized Face Forgery DetectionabstractFace forgery technology has developed rapidly, causing severe security issues in society. Recently, with the continuous emergence of forgery techniques and types, most forensics methods suffer from the generalization problem. In particular, it is difficult for existing generalized methods to detect fake faces with unseen fake types. The reason is that the distribution gaps among cross-forgery types are too large. In this article, we propose a novel generalized framework to narrow large gaps based on bridging cross-domain alignment to solve this problem. Specifically, our framework consists of three key steps: preventing, bridging and aligning distribution gaps. Firstly, in the feature mining stage, taking advantage of the ability of Instance Normalization (IN) to better tolerate domain gaps, we design Adaptive Batch and Instance Normalization (ABIN) to replace the commonly used BN to adaptively extract features to preliminarily prevent domain gaps. Secondly, we propose to generate bridging samples distributed among the inter-domains to fill large gaps based on progressive linear interpolation operation. Finally, with the help of bridging samples, the cross-domain alignment is performed to better narrow distribution gaps to refine data distribution, which helps to learn a more generalized framework. Extensive experiments show that our proposed framework achieves the state-of-the-art generalized performance. Yang Yu 0039, Siyuan Yang 0001, Yao Zhao 0001, Alex Chichung Kot |
IEEE Trans. Multim. | 1 |
| 2023 | Magnifying multimodal forgery clues for Deepfake detection
Yang Yu 0039, Xiaolong Li 0001, Yao Zhao 0001 |
Signal Process. Image Commun. | 2 |
| 2023 | Defending Fake via Warning: Universal Proactive Defense Against Face ManipulationabstractThe emergence of deep learning has led to the rise of malicious face manipulation applications, which pose a significant threat to face security. In order to prevent the generation of forgery fundamentally, researchers have proposed proactive methods to disrupt the process of manipulation models. However, these methods output distorted images, which exhibit unacceptable black shadows or distorted facial features causing facial stigmatization. To address this issue, we propose a Universal Proactive Warning Defense (UPWD) method, which leads fake images to present a warning pattern against multiple manipulation models. Specifically, we proposed an Invisible Protection Module that generates protection messages and a feature-level measure strategy that enhances the salience of warning patterns. Furthermore, we improve the universality of the method based on Hard Model Meta-learning. Extensive experimental results on CelebA and LFWA datasets demonstrate that our proposed UPWD method effectively defends against multiple manipulation models and outperforms existing methods. Yu Chen 0049, Yang Yu 0039, Yao Zhao 0001 |
IEEE Signal Process. Lett. | 4 |
| 2023 | MSVT: Multiple Spatiotemporal Views Transformer for DeepFake Video DetectionabstractRecently, DeepFake videos have developed rapidly, causing new security issues in society. Due to the rough spatiotemporal view, existing video-based detection methods struggle to capture fine-grained spatiotemporal information, resulting in limited generalization ability. In addition, although the transformer has achieved great success in the past few years, the application of transformer on deepfake video detection still needs to be studied. To solve this problem, in this paper, we propose a novel Multiple Spatiotemporal Views Transformer (MSVT) with Local Spatiotemporal View (LSV) and Global Spatiotemporal View (GSV), to mine more detailed spatiotemporal information. Firstly, for establishing the LSV, different from existing works that sparsely sample a single frame to build the input sequence, we employ the local-consecutive temporal view to capture vital dynamic inconsistency. Furthermore, the extracted frame features within each group are fed to the temporal transformer followed by the feature fusion module, to generate group-level spatiotemporal features. Then, we further establish Global Spatiotemporal View (GSV) by feeding all the frame features within the whole video to the temporal transformer followed by the feature fusion module. Finally, we propose a novel global-local transformer (GLT) to effectively integrate these multi-level features for mining more subtle and comprehensive features. Extensive experiments on six large datasets demonstrate that our MSVT outperforms state-of-the-art detection methods. Yang Yu 0039, Yao Zhao 0001, Siyuan Yang 0001, Fen Xia |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | Augmented Multi-Scale Spatiotemporal Inconsistency Magnifier for Generalized DeepFake DetectionabstractRecently, realistic DeepFake videos have raised severe security concerns in society. Existing video-based detection methods observe local spatial regions with the coarse temporal view, thus it is difficult to obtain subtle spatiotemporal information, resulting in limited generalization ability. In this paper, we propose a novel Augmented Multi-scale Spatiotemporal Inconsistency Magnifier (AMSIM) with a Global Inconsistency View (GIV) and a more meticulous Multi-timescale Local Inconsistency View (MLIV), focusing on mining comprehensive and more subtle spatiotemporal cues. Firstly, the GIV that includs the global spatial and long-term temporal views is established to ensure comprehensive spatiotemporal clues are captured. Then, the MLIV with the critical local spatial and multi-timescale local temporal views is designed for magnifying the indetectable spatiotemporal abnormality. Subsequently, GIV is utilized to guide MLIV to dynamically find local spatiotemporal anomalies that are highly relevant to the overall video. Finally, to further obtain a generalized framework, the adversarial data augmentation is specially designed to expand source domains and simulate unseen forgery domains. Extensive experiments on six large-scale datasets show that our AMSIM outperforms state-of-the-art detection methods and remains effective when applied to unseen forgery techniques and datasets. Yang Yu 0039, Siyuan Yang 0001, Yao Zhao 0001, Alex Chichung Kot |
IEEE Trans. Multim. | 1 |
| 2023 | TCSD: Triple Complementary Streams Detector for Comprehensive Deepfake DetectionabstractAdvancements in computer vision and deep learning have made it difficult to distinguish deepfake visual media. While existing detection frameworks have achieved significant performance on challenging deepfake datasets, these approaches consider only a single perspective. More importantly, in urban scenes, neither complex scenarios can be covered by a single view nor can the correlation between multiple datasets of information be well utilized. In this article, to mine the new view for deepfake detection and utilize the correlation of multi-view information contained in images, we propose a novel triple complementary streams detector (TCSD). First, a novel depth estimator is designed to extract depth information (DI), which has not been used in previous methods. Then, to supplement depth information for obtaining comprehensive forgery clues, we consider the incoherence between image foreground and background information (FBI) and the inconsistency between local and global information (LGI). In addition, we designed an attention-based multi-scale feature extraction (MsFE) module to extract more complementary features from DI, FBI, and LGI. Finally, two attention-based feature fusion modules are proposed to adaptively fuse information. Extensive experiment results show that the proposed approach achieves state-of-the-art performance on detecting deepfakes. Yang Yu 0039, Xiaolong Li 0001, Yao Zhao 0001, Guodong Guo |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2022 | Exploring Complementarity of Global and Local Spatiotemporal Information for Fake Face Video DetectionabstractThe spread of fake face videos leads to severe social concerns, which promotes the development of detection methods for these videos. Existing patch-based methods focus on local regions to find forgery common clues, while ignoring the important role of the global information. In this paper, a novel spatiotemporal network is proposed which can better utilize the implicit complementary advantages of global and local information. Specifically, the spatial module consists of the global information stream and local information stream extracted from patches selected by attention layers. Then, the fusion features of these two streams are fed into the temporal module to further capture temporal clues. Besides, a regularization loss is designed to guide the selection of local information and the extraction of fusion temporal information with the reference to the global information. Extensive experiments on different datasets demonstrate the superiority of our framework. Yang Yu 0039, Yao Zhao 0001 |
ICASSP | 2 |
| 2022 | Information Adversarial Disentanglement for Face Swapping
Yang Yu 0039, Yao Zhao 0001 |
PRCV (4) | 2 |
| 2022 | Detection of AI-Manipulated Fake Faces via Mining Generalized FeaturesabstractRecently, AI-manipulated face techniques have developed rapidly and constantly, which has raised new security issues in society. Although existing detection methods consider different categories of fake faces, the performance on detecting the fake faces with “unseen” manipulation techniques is still poor due to the distribution bias among cross-manipulation techniques. To solve this problem, we propose a novel framework that focuses on mining intrinsic features and further eliminating the distribution bias to improve the generalization ability. First, we focus on mining the intrinsic clues in the channel difference image (CDI) and spectrum image (SI) view of two different aspects, including the camera imaging process and the indispensable step in AI manipulation process. Then, we introduce the Octave Convolution and an attention-based fusion module to effectively and adaptively mine intrinsic features from CDI and SI view of these two different but intrinsic aspects. Finally, we design an alignment module to eliminate the bias of manipulation techniques to obtain a more generalized detection framework. We evaluate the proposed framework on four categories of fake faces datasets with the most popular and state-of-the-art manipulation techniques and achieve very competitive performances. We further conduct experiments on cross-manipulation techniques, and the results of our method show the superior advantages on improving generalization ability. Yang Yu 0039, Yao Zhao 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2020 | Detection of fake high definition for HEVC videos based on prediction mode feature
Yang Yu 0039, Haichao Yao, Yao Zhao 0001 |
Signal Process. | 1 |