VLDB 2026 Research / reviewers in the wild / expert
Qilin Yin
dblp:236/3223
· DBLP profile ↗
13ranked-venue papers
5as first author
12since 2021 · last 2026
0000-0001-7571-046XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 6 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Context-Aware TFL: A Universal Context-Aware Contrastive Learning Framework for Temporal Forgery LocalizationabstractMost research efforts in the multimedia forensics domain have focused on detecting forgery audio-visual content and reached sound achievements. However, these works only consider deepfake detection as a classification task and ignore the case where partial segments of the video are tampered with. Temporal forgery localization (TFL) of small fake audio-visual clips embedded in real videos is still challenging and more in line with realistic application scenarios. To resolve this issue, we propose a universal context-aware contrastive learning framework (UniCa-CLF) for TFL. Our approach leverages supervised contrastive learning to discover and identify forged instants by means of anomaly detection, allowing for the precise localization of temporal forged segments. To this end, we propose a specialized context-aware perception layer that utilizes a heterogeneous activation operation and an adaptive context updater to construct a context-aware contrastive objective, which enhances the discriminability of forged instant features by contrasting them with genuine instant features in terms of their distances to the global context. An efficient context-aware contrastive coding is introduced to further push the limit of instant feature distinguishability between genuine and forged instants in a supervised sample-by-sample manner, suppressing the cross-sample influence to improve temporal forgery localization performance. Extensive experimental results over five public datasets demonstrate that our proposed UniCaCLF significantly outperforms the state-of-the-art competing algorithms. The source code and pre-trained models of our proposed UniCaCLF are made publicly available at GitHub repository.TimeTimeTimeTime Qilin Yin, Wei Lu 0001, Xiangyang Luo 0001, Rui Yang 0006, Xiaochun Cao |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2025 | GLCF: A Global-Local Multimodal Coherence Analysis Framework for Talking Face Generation DetectionabstractTalking face generation (TFG) allows for producing lifelike talking videos of any character using only facial images and accompanying text. Abuse of this technology could pose significant risks to society, creating the urgent need for research into corresponding detection methods. However, research in this field has been hindered by the lack of public datasets. In this paper, we construct the first large-scale multi-scenario talking face dataset (MSTF), which contains 22 audio and video forgery techniques, filling the gap of datasets in this field. The dataset covers 11 generation scenarios and more than 20 semantic scenarios, closer to the practical application scenario of TFG. Besides, we also propose a TFG detection framework, which leverages the analysis of both global and local coherence in the multimodal content of TFG videos. Therefore, a region-focused smoothness detection module (RSFDM) and a discrepancy capture-time frame aggregation module (DCTAM) are introduced to evaluate the global temporal coherence of TFG videos, aggregating multi-grained spatial information. Additionally, a visual-audio fusion module (V-AFM) is designed to evaluate audiovisual coherence within a localized temporal perspective. Comprehensive experiments demonstrate the reasonableness and challenges of our datasets, while also indicating the superiority of our proposed method compared to the state-of-the-art deepfake detection approaches. Xiaocan Chen, Qilin Yin, Jiarui Liu 0002, Wei Lu 0001, Xiangyang Luo 0001, Jiantao Zhou 0001 |
AAAI | 2 |
| 2025 | Towards Open-world Generalized Deepfake Detection: General Feature Extraction via Unsupervised Domain AdaptationabstractWith the development of generative artificial intelligence, new forgery methods are rapidly emerging. Social platforms are flooded with vast amounts of unlabeled synthetic data and authentic data, making it increasingly challenging to distinguish real from fake. Due to the lack of labels, existing supervised detection methods struggle to effectively address the detection of unknown deepfake methods. Moreover, in open world scenarios, the amount of unlabeled data greatly exceeds that of labeled data. Therefore, we define a new deepfake detection generalization task which focuses on how to achieve efficient detection of large amounts of unlabeled data based on limited labeled data to simulate a open world scenario. To solve the above mentioned task, we propose a novel Open-World Deepfake Detection Generalization Enhancement Training Strategy (OWG-DS) to improve the generalization ability of existing methods. Our approach aims to transfer deepfake detection knowledge from a small amount of labeled source domain data to large-scale unlabeled target domain data. Specifically, we introduce the Domain Distance Optimization (DDO) module to align different domain features by optimizing both inter-domain and intra-domain distances. Additionally, the Similarity-based Class Boundary Separation (SCBS) module is used to enhance the aggregation of similar samples to ensure clearer class boundaries, while an adversarial training mechanism is adopted to learn the domain-invariant features. Extensive experiments show that the proposed deepfake detection generalization enhancement training strategy excels in cross-method and cross-dataset scenarios, improving the model's generalization. Midou Guo, Qilin Yin, Wei Lu 0001, Xiangyang Luo 0001 |
ACM Multimedia | 2 |
| 2025 | Deepfake Detection and Localization Using Multi-View Inconsistency MeasurementabstractAs deepfake technology advances, forgery detection techniques have evolved beyond simple classification to include fine-grained localization. However, existing deepfake localization methods struggle with with real-world deepfake videos, which are often multi-face scenarios with only some parts manipulated. To address the above-mentioned problems, we propose a Multi-View Inconsistency Measurement (MVIM) network that simultaneously measures inconsistencies from noise and temporal view to detect and locate tampered regions. Specifically, considering the noise inconsistencies in multi-face scenarios where fake faces have inconsistent noise patterns compared to real faces and backgrounds, we design a Noise Inconsistency Measurement (Noise-IM) module that measures noise similarity among faces and between faces and backgrounds using a masked attention mechanism to identify suspected tampered regions in noise domain. Since facial jitter of tampered regions in deepfake videos is observed to be more intense than that of real regions, we design a Temporal Inconsistency Measurement (Temporal-IM) module which adopts self-attention mechanism and fine-grained bi-direction convolutions to capture tampering traces between frames in temporal domain. Inconsistency features obtained by the two modules are fused for detecting and locating tampered regions. The superiority of our MVIM network is verified by extensive experiments with many state-of-the-art methods in different benchmark datasets. Qilin Yin, Wei Lu 0001, Xiangyang Luo 0001 |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2024 | Fine-Grained Multimodal DeepFake Classification via Heterogeneous Graphs
Qilin Yin, Wei Lu 0001, Xiaochun Cao, Xiangyang Luo 0001, Yicong Zhou, Jiwu Huang |
Int. J. Comput. Vis. | 1 |
| 2024 | Deepfake detection via inter-frame inconsistency recomposition and enhancement
Chuntao Zhu, Qilin Yin, Chengxi Yin, Wei Lu 0001 |
Pattern Recognit. | 3 |
| 2024 | Audio Multi-View Spoofing Detection Framework Based on Audio-Text-Emotion CorrelationsabstractIn recent years, audio spoofing detection has received widespread attention for protecting personal privacy and social security. Despite the significant progress achieved in audio single-view spoofing detection, challenges remain with regard to addressing unknown spoofing attacks in realistic scenarios. To solve these challenging problems, in this paper, we introduce a novel audio multi-view spoofing detection framework (AMSDF), whose goal is to capture both intra-view and inter-view cues by measuring correlations within audio multi-view features (i.e., audio-emotion-text) for audio spoofing detection. In general, different view features are inherently interconnected in the real patterns, while they may present unnatural correlations in the spoofing patterns. Therefore, more discriminative cues can be mined by utilizing their complex interactions, which is beneficial to the audio spoofing detection task. To this end, an intra-view graph attention mechanism (IGAM) is first utilized to aggregate each intra-view node within the same view. Subsequently, a heterogeneous graph fusion module (HGFM) is applied to measure correlations within inter-view nodes, which are enhanced with a master node for comprehensive analysis purposes. Finally, a group-based readout scheme (GRS) is designed to capture and preserve the most distinctive cues by leveraging the strengths of different feature sets, thereby effectively distinguishing subtle differences between real and spoofing audio. The experimental results show that our proposed framework can achieve better performance than that of the state-of-the-art methods, especially in realistic scenarios. The code and pre-trained models are available athttps://github.com/ItzJuny/AMSDF. Junyan Wu, Qilin Yin, Ziqi Sheng, Wei Lu 0001, Jiwu Huang, Bin Li 0011 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2023 | Dynamic Difference Learning With Spatio-Temporal Correlation for Deepfake Video DetectionabstractWith the rapid development of face forgery techniques, the existing frame-based deepfake video detection methods have fell into a dilemma that frame-based methods may fail when encountering extremely realistic images. To overcome the above problem, many approaches attempted to model the spatio-temporal inconsistency of videos to distinguish real and fake videos. However, current works model spatio-temporal inconsistency by combining intra-frame and inter-frame information, but ignore the disturbance caused by facial motions that would limit further improvement in detection performance. To address this issue, we investigate into long and short range inter-frame motions and propose a novel dynamic difference learning method to distinguish between the inter-frame differences caused by face manipulation and the inter-frame differences caused by facial motions in order to model precise spatio-temporal inconsistency for deepfake video detection. Moreover, we elaborately design a dynamic fine-grained difference capture module (DFDC-module) and a multi-scale spatio-temporal aggregation module (MSA-module) to collaboratively model spatio-temporal inconsistency. Specifically, the DFDC-module applies self-attention mechanism and fine-grained denoising operation to eliminate the differences caused by facial motions and generates long range difference attention maps. The MSA-module is devised to aggregate multi-direction and multi-scale temporal information to model spatio-temporal inconsistency. The existing 2D CNNs can be extended into dynamic spatio-temporal inconsistency capture networks by integrating the proposed two modules. Extensive experimental results demonstrate that our proposed algorithm steadily outperforms state-of-the-art methods by a clear margin in different benchmark datasets. Qilin Yin, Wei Lu 0001, Bin Li 0011, Jiwu Huang |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2022 | GAN-generated fake face detection via two-stream CNN with PRNU in the wild
Kehui Zeng, Bin Ma 0003, Xiangyang Luo 0001, Qilin Yin, Guangjie Liu 0001, Sunil Kr. Jha |
Multim. Tools Appl. | 5 |
| 2022 | Contrastive Learning based Multi-task Network for Image Manipulation Detection
Qilin Yin, Wei Lu 0001, Xiangyang Luo 0001 |
Signal Process. | 1 |
| 2022 | SmsNet: A New Deep Convolutional Neural Network Model for Adversarial Example DetectionabstractThe emergence of adversarial examples has had a significant impact on the development and application of deep learning. In this paper, a novel convolutional neural network model, the stochastic multifilter statistical network (SmsNet), is proposed for the detection of adversarial examples. A feature statistical layer is constructed to collect statistical data of feature map output from each convolutional layer in SmsNet by combining manual features with a neural network. The entire model is an end-to-end detection model, so the feature statistical layer is not independent of the network, and its output is directly transmitted to the fully connected layer by a short-cut connection called the SmsConnection. Additionally, a dynamic pruning strategy is introduced to simplify the model structure for better performance. The experiments demonstrate the effectiveness of the network structure and pruning strategy, and the proposed model achieves high detection rates against state-of-the-art adversarial attacks. Qilin Yin, Xiangyang Luo 0001, Yuhui Zheng, Yun Q. Shi 0001, Sunil Kr. Jha |
IEEE Trans. Multim. | 3 |
| 2021 | Modify the Quantization Table in the JPEG Header File for Forensics and Anti-forensics
Hao Wang 0060, Xiangyang Luo 0001, Qilin Yin, Bin Ma 0003, Jinsheng Sun |
IWDW | 4 |
| 2020 | A Hybrid Loss Network for Localization of Image Manipulation
Qilin Yin, Xiangyang Luo 0001 |
IWDW | 1 |