VLDB 2026 Research / reviewers in the wild / expert
Yuzhen Lin
dblp:219/4153
· DBLP profile ↗
18ranked-venue papers
6as first author
16since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 5 first-author · 13 since 2021Artificial intelligence and machine learning · 8 · 1 first-author · 8 since 2021Computer networks · 1Security and privacy · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MOSAIC: Orchestrating Collaborative Knowledge Tracing with Hierarchical Semantic Alignment
Xinjin Li, Mengyue Wang, Yuzhen Lin, Pengbin Feng, Ziqi Sha, Yeyang Zhou |
ICPR (15) | 3 |
| 2026 | CEMG: Collaborative-Enhanced Multimodal Generative Recommendation
Yuzhen Lin, Xuanjing Chen, Ivonne Xu, Dongming Jiang |
MMM (1) | 1 |
| 2025 | Standing on the Shoulders of Giants: Reprogramming Visual-Language Model for General Deepfake DetectionabstractThe proliferation of deepfake faces poses huge potential negative impacts on our daily lives. Despite substantial advancements in deepfake detection over these years, the generalizability of existing methods against forgeries from unseen datasets or created by emerging generative models remains constrained. In this paper, inspired by the zero-shot advantages of Vision-Language Models (VLMs), we propose a novel approach that repurposes a well-trained VLM for general deepfake detection. Motivated by the model reprogramming paradigm that manipulates the model prediction via input perturbations, our method can reprogram a pre-trained VLM model (e.g., CLIP) solely based on manipulating its input without tuning the inner parameters. First, learnable visual perturbations are used to refine feature extraction for deepfake detection. Then, we exploit information of face embedding to create sample-level adaptative text prompts, improving the performance. Extensive experiments on several popular benchmark datasets demonstrate that (1) the cross dataset and cross-manipulation performances of deepfake detection can be significantly and consistently improved (e.g., over 88% AUC in cross-dataset setting from FF++ to Wild-Deepfake); (2) the superior performances are achieved with fewer trainable parameters, making it a promising approach for real-world applications. Kaiqing Lin, Yuzhen Lin, Weixiang Li, Taiping Yao, Bin Li 0011 |
AAAI | 2 |
| 2025 | Reinforced Multi-teacher Knowledge Distillation for Efficient General Image Forgery Detection and LocalizationabstractImage forgery detection and localization (IFDL) is of vital importance as forged images can spread misinformation that poses potential threats to our daily life. However, previous methods still struggled to effectively handle forged images processed with diverse forgery operations in real-world scenarios. In this paper, we propose a novel Reinforced Multi-teacher Knowledge Distillation (Re-MTKD) framework for the IFDL task, structured around an encoder-decoder ConvNeXt-UperNet along with Edge-Aware Module, named Cue-Net. First, three Cue-Net models are separately trained for the three main types of image forgeries, i.e., copy-move, splicing and inpainting, which then serve as the multi-teacher models to train the target student model with Cue-Net through self-knowledge distillation. A Reinforced Dynamic Teacher Selection (Re-DTS) strategy is developed to dynamically assign weights to the involved teacher models, which facilitates specific knowledge transfer and enables the student model to effectively learn both the common and specific natures of diverse tampering traces. Extensive experiments demonstrate that, compared with other state-of-the-art methods, the proposed method achieves superior performance on several recently emerged datasets comprised of various kinds of image forgeries. Zeqin Yu, Jiangqun Ni, Jian Zhang 0086, Haoyi Deng, Yuzhen Lin |
AAAI | 5 |
| 2025 | Optimization of Deep Learning Models for Dynamic Market Behavior Prediction
Shenghan Zhao, Yuzhen Lin, Ximeng Yang, Qiaochu Lu, Haozhong Xue, Gaozhe Jiang |
IEEE Big Data | 2 |
| 2025 | Federated Learning for Heterogeneous Data Integration and Privacy ProtectionabstractFederated learning (FL) represents a promising approach that enables the collaborative training of machine learning models without compromising data privacy. This approach is particularly advantageous when handling heterogeneous data dispersed across numerous institutions or devices, as centralized data aggregation is often constrained by privacy concerns and data regulations. In order to address the challenges posed by heterogeneous data, we have devised an adaptive data integration mechanism. This mechanism maps the features of disparate data sources to a unified feature space through the use of feature alignment technology, thereby facilitating the effective fusion of data. This fusion is achieved through the application of statistical alignment and multi-perspective learning technology. Furthermore, in order to safeguard the confidentiality of data, we integrate differential privacy and homomorphic encryption techniques, thereby preventing the disclosure of information during model updates and data transfers. Furthermore, a multi-level privacy protection strategy is proposed, which employs de-identification, secure multi-party computation, and federated averaging technologies at the three stages of data preprocessing, model training, and result aggregation, respectively. This approach ensures data security and facilitates effective model updates. The experimental results demonstrate that the proposed framework exhibits enhanced model performance and robustness in comparison to traditional federated learning methods on a multitude of real-world heterogeneous datasets. Chenwei Gong, Yuzhen Lin, Pei-Chiang Su |
CSCWD | 3 |
| 2025 | Guard Me If You Know Me: Protecting Specific Face-Identity from DeepfakesabstractSecuring personal identity against deepfake attacks is increasingly critical in the digital age, especially for celebrities and political figures whose faces are easily accessible and frequently targeted.
Most existing deepfake detection methods focus on general-purpose scenarios and often ignore the valuable prior knowledge of known facial identities, e.g., "VIP individuals" whose authentic facial data are already available.
In this paper, we propose **VIPGuard**, a unified multimodal framework designed to capture fine-grained and comprehensive facial representations of a given identity, compare them against potentially fake or similar-looking faces, and reason over these comparisons to make accurate and explainable predictions.
Specifically, our framework consists of three main stages. First, we fine-tune a multimodal large language model (MLLM) to learn detailed and structural facial attributes.
Second, we perform identity-level discriminative learning to enable the model to distinguish subtle differences between highly similar faces, including real and fake variations. Finally, we introduce user-specific customization, where we model the unique characteristics of the target face identity and perform semantic reasoning via MLLM to enable personalized and explainable deepfake detection.
Our framework shows clear advantages over previous detection works, where traditional detectors mainly rely on low-level visual cues and provide no human-understandable explanations, while other MLLM-based models often lack a detailed understanding of specific face identities.
To facilitate the evaluation of our method, we build a comprehensive identity-aware benchmark called **VIPBench** for personalized deepfake detection, involving the latest 7 face-swapping and 7 entire face synthesis techniques for generation.
Extensive experiments show that our model outperforms existing methods in both detection and explanation.
The code is available at https://github.com/KQL11/VIPGuard . Kaiqing Lin, Zhiyuan Yan 0002, Ke-Yue Zhang, Yuzhen Lin, Weixiang Li, Taiping Yao, Shouhong Ding, Bin Li 0011 |
NeurIPS | 6 |
| 2024 | DiffForensics: Leveraging Diffusion Prior to Image Forgery Detection and LocalizationabstractAs manipulating images may lead to misinterpretation of the visual content, addressing the image forgery detection and localization (IFDL) problem has drawn serious public concerns. In this work, we propose a simple assumption that the effective forensic method should focus on the mesoscopic properties of images. Base on the assumption, a novel two-stage self-supervised framework leveraging the diffusion model for IFDL task, i.e., DiffForensics, is proposed in this paper. The DiffForensics begins with self-supervised denoising diffusion paradigm equipped with the module of encoder-decoder structure, by freezing the pre-trained encoder (e.g., in ADE-20K) to inherit macroscopic features for general image characteristics, while encour-aging the decoder to learn microscopic feature represen-tation of images, enforcing the whole model to focus the mesoscopic representations. The pre-trained model as a prior, is then further fine-tuned for IFDL task with the customized Edge Cue Enhancement Module (ECEM), which progressively highlights the boundary features within the manipulated regions, thereby refining tampered area local-ization with better precision. Extensive experiments on several public challenging datasets demonstrate the effectiveness of the proposed method compared with other state-of-the-art methods. The proposed DiffForensics could significantly improve the model's capabilities for both accurate tamper detection and precise tamper localization while con-currently elevating its generalization and robustness. Zeqin Yu, Jiangqun Ni, Yuzhen Lin, Haoyi Deng, Bin Li 0011 |
CVPR | 3 |
| 2024 | Fake It till You Make It: Curricular Dynamic Forgery Augmentations Towards General Deepfake Detection
Yuzhen Lin, Wentang Song, Bin Li 0011, Yuezun Li, Jiangqun Ni, Qiushi Li 0001 |
ECCV (86) | 1 |
| 2024 | Towards Generic Deepfake Detection with Dynamic CurriculumabstractMost previous deepfake detection methods bent their efforts to discriminate artifacts by end-to-end training. However, the learned networks often fail to mine the generic face forgery information efficiently due to ignoring data diversity. In this work, we propose to introduce sample hardness into the training of deepfake detectors via a curriculum learning paradigm. Specifically, we present a novel simple yet effective strategy, named Dynamic Facial Forensic Curriculum (DFFC), which makes the model gradually focus on hard samples during the training. To this end, we propose Dynamic Forensic Hardness (DFH) which integrates the facial quality score and instantaneous instance loss to dynamically measure sample hardness during training. Besides, we present a pacing function to construct data subsets from easy to hard throughout the training process based on DFH. Comprehensive experiments show that DFFC can improve both within- and cross-dataset performance of various kinds of end-to-end deepfake detectors in a plug-and-play manner. It indicates that DFFC can help deepfake detectors learn generic forgery discriminative features more efficiently by exploiting the information from hard samples. Wentang Song, Yuzhen Lin, Bin Li 0011 |
ICASSP | 2 |
| 2023 | Learning to Locate the Text Forgery in Smartphone ScreenshotsabstractIn this paper, we present the Screenshot Text Forgery Dataset (STFD), which is the first public dataset for the smartphone screenshot text forgery localization task. To address such a task, we propose a novel Screenshot Text Forgery Localization Network (STFL-Net). Specifically, we introduce the OCR (Optical Character Recognition) stream as the complementary of the RGB stream, and propose a novel dual-stream Y-net architecture to collaboratively learn the representations focused on the traces on text regions of the image. Considering the text forgery is often subtle and local, we introduce a multi-teacher knowledge distillation learning strategy for training the STFL-Net, which makes the model less prone to over-fit one specific forgery trace. Comprehensive experimental results on STFD show that our method outperforms several previous methods designed for image forgery localization. We believe that, with our STFD dataset and STFL-Net, more advanced countermeasures against screenshot text forgeries can be developed in the future. Zeqin Yu, Bin Li 0011, Yuzhen Lin, Jinhua Zeng, Jishen Zeng |
ICASSP | 3 |
| 2023 | Learning Features of Intra-Consistency and Inter-Diversity: Keys Toward Generalizable Deepfake DetectionabstractPublic concerns about deepfake face forgery are continually rising in recent years. Most deepfake detection approaches attempt to learn discriminative features between real and fake faces through end-to-end trained deep neural networks. However, the majorities of them suffer from the problem of poor generalization across different data sources, forgery methods, and/or post-processing operations. In this paper, following the simple but effective principle in discriminative representation learning, i.e., towards learning features of intra-consistency within classes and inter-diversity between classes, we leverage a novel transformer-based self-supervised learning method and an effective data augmentation strategy towards generalizable deepfake detection. Considering the differences between the real and fake images are often subtle and local, the proposed method firstly utilizes Self Prediction Learning (SPL) to learn rich hidden representations by predicting masked patches at a pre-training stage. Intra-class consistency clues in images can be mined without deepfake labels. After pre-training, the discrimination model is then fine-tuned via multi-task learning, including a deepfake classification task and a forgery mask estimation task. It is facilitated by our new data augmentation method called Adjustable Forgery Synthesizer (AFS), which can conveniently simulate the process of synthesizing deepfake images with various levels of visual reality in an explicit manner. AFS greatly prevents overfitting due to insufficient diversity in training data. Comprehensive experiments demonstrate that our method outperforms the state-of-the-art competitors on several popular benchmark datasets in terms of generalization to unseen forgery methods and untrained datasets. Yuzhen Lin, Bin Li 0011, Shunquan Tan |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Exposing Face Forgery Clues via Retinex-Based Image Enhancement
Yuzhen Lin, Bin Li 0011 |
ACCV (4) | 2 |
| 2022 | Towards Generalizable DEEPFAKE Face Forgery Detection with Semi-Supervised Learning and Knowledge DistillationabstractExisting methods for deepfake face forgery detection have already achieved tremendous progress in well-controlled laboratory conditions. However, under wild scenarios where the training and testing forgeries are synthesized by different algorithms and when labeled data are insufficient, the performance always drops greatly. In this work, we present a Semi-supervised Contrastive Learning and Knowledge Distillation-based framework (SCL-KD) for deepfake detection to reduce the aforementioned performance gap. Our proposed framework contains three stages: self-supervised pre-training, supervised training, and knowledge distillation. Specifically, a feature encoder is firstly trained in a self-supervised manner with a large number of unlabeled samples through a momentum contrastive mechanism. Secondly, a fully-connected classifier on top of the feature encoder is trained in a supervised manner with a small amount of labeled samples to build a teacher model. Finally, a compact student model is trained with the help of the teacher model using knowledge distillation, in order to avoid overfitting to labeled data and have better generalizability on mismatched datasets. Evaluations on several benchmark datasets corroborate the good performance of our approach in cross-dataset situations and few labeled data scenarios. It reveals the potential of our proposed method for real-world deepfake detection. Yuzhen Lin, Bin Li 0011, Junqiang Wu |
ICIP | 1 |
| 2022 | Source-ID-Tracker: Source Face Identity Protection in Face SwappingabstractSwapping faces with deep learning technology to generate realistic fake videos/images (a.k.a, deepfakes) has drawn great public con-cerns recently. Numerous approaches have been proposed to iden-tify fake contents; however, less work has been dedicated to pro-tecting the source faces in an active way. In this paper, we stand for a legitimate faceswap service provider and present an approach called Source-ID- Tracker (SIDT), which aims to protect the identity of source faces in deepfakes from malicious uses. As a plug-in, the encoder of SIDT implicitly embeds a source face image into a deep-fake image while ensuring the resultant encoded image is visually indistinguishable from the deepfake image. After sharing through social media, the embedded source face and its identity can still be recovered with a decoder. Experimental results show that the pro-posed model achieves a promising performance, in terms of reconstruction quality and attribution inference accuracy, in revealing the hidden source face. Yuzhen Lin, Emanuele Maiorana, Patrizio Campisi, Bin Li 0011 |
ICME | 1 |
| 2021 | Tackling the Cover Source Mismatch Problem in Audio Steganalysis With Unsupervised Domain AdaptationabstractNowadays, the convolutional neural network (CNN) based steganalysis has achieved remarkable performance in the well-controlled lab environment. However, the cover source mismatch (CSM) problem, which can be attributed to the discrepancy between the training, and evaluation datasets, is still one of the pivotal obstacles for adapting the steganalysis into real-world applications. In this letter, we propose to merge the domain adaptation strategy into CNN-based audio steganalysis for handling the CSM problem. Specifically, the proposed framework contains three components: feature extractor, steganalytic classifier, and domain discriminator. The cascade of feature extractor, and steganalytic classifier compose the typical supervised steganalysis model. The unsupervised domain adaptation is implemented by the domain adversarial training between the feature extractor, and domain discriminator. Ultimately, the feature extractor is trained to extract the steganalytic, and domain-invariant features. It aims to reduce the domain gap between the training data, and testing data. The experimental results show that our approach could effectively mitigate the CSM impact caused by the diversity of audio recording devices. Yuzhen Lin, Rangding Wang, Li Dong 0006, Diqun Yan, Jie Wang 0028 |
IEEE Signal Process. Lett. | 1 |
| 2020 | Towards Designing an Effective Complexity Indicator for Audio SteganographyabstractIn the field of steganography, to effectively hide the secret message, it is of great importance to determine which part of the steganographic cover is suitable for embedding. Currently, most of the existing works focus on the image cover, while few works touch the audio cover case. In this work, we attempt to characterize the complexity of audio for selecting the steganographic cover. Specifically, the original cover is first convoluted with a specially designed adaptive convolution kernel. Based on the residual between the original and the convoluted audio, we derive a quantity for measuring the complexity of each frame for a given audio clip. Experimental results verify the usability of the proposed complexity indicator, suggesting high-complexity audio cover is favorable for data embedding. It is also found that the proposed complexity indicator could further boost the steganographic performance of the state-of the-art audio steganography methods. The source code is publicly available at https://github.com/capzxy/audio-complexity. Xueyuan Zhang, Rangding Wang, Li Dong 0006, Diqun Yan, Yuzhen Lin, Jie Wang 0028 |
ICC | 5 |
| 2019 | Audio Steganalysis with Improved Convolutional Neural NetworkabstractDeep learning, especially the convolutional neural network (CNN), has enjoyed significant success in many fields, e.g., image recognition. Recently, CNN has successfully applied to multimedia steganalysis. However, the detection performance is still unsatisfactory. In this work, we propose an improved CNN-based method for audio steganalysis. Specifically, a special convolutional layer is first carefully designed, which could capture the minor steganographic noise. Then, a truncated linear unit is adapted to activate the output of shallow convolutional layer. In addition, we employ the average pooling to minimize the over-fitting risk. Finally, a parameter transfer strategy is adopted, aiming to boost the detection performance for the low embedding-rate cases. The experimental results evaluated on 30,000 audio clips verify the effectiveness of our method for a variety of embedding rates. Compared with the existing CNN-based steganalysis methods, our proposed method could achieve superior performance. To facilitate the reproducible research, the source code will be released at GitHub. Yuzhen Lin, Rangding Wang, Diqun Yan, Li Dong 0006, Xueyuan Zhang |
IH&MMSec | 1 |