EDBT 2026 Demo / reviewers in the wild / expert
Peipeng Yu
dblp:258/5964
· DBLP profile ↗
20ranked-venue papers
5as first author
20since 2021 · last 2026
0000-0003-0056-4300ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 11 since 2021Artificial intelligence and machine learning · 9 · 1 first-author · 9 since 2021Security and privacy · 3 · 2 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CLIP-FTI: Fine-Grained Face Template Inversion via CLIP-Driven Attribute ConditioningabstractFace recognition systems store face templates for efficient matching. Once leaked, these templates pose a threat: inverting them can yield photorealistic surrogates that compromise privacy and enable impersonation. Although existing research has achieved relatively realistic face template inversion, the reconstructed facial images exhibit over-smoothed facial-part attributes (eyes, nose, mouth) and limited transferability. To address this problem, we present CLIP-FTI, a CLIP-driven fine-grained attribute conditioning framework for face template inversion. Our core idea is to use the CLIP model to obtain the semantic embeddings of facial features, in order to realize the reconstruction of specific facial feature attributes. Specifically, facial feature attribute embeddings extracted from CLIP are fused with the leaked template via a cross-modal feature interaction network and projected into the intermediate latent space of a pretrained Style- GAN. The StyleGAN generator then synthesizes face images with the same identity as the templates but with more finegrained facial feature attributes. Experiments across multiple face recognition backbones and datasets show that our reconstructions (i) achieve higher identification accuracy and attribute similarity, (ii) recover sharper component-level attribute semantics, and (iii) improve cross-model attack transferability compared to prior reconstruction attacks. To the best of our knowledge, ours is the first method to use additional information besides the face template attack to realize face template inversion and obtains SOTA results. Longchen Dai, Zixuan Shen, Peipeng Yu, Zhihua Xia |
AAAI | 4 |
| 2026 | One for All: Synthesis-Free Fingerprint Learning for Attribution of In-the-Wild Synthetic ImagesabstractAttributing synthetic images to their source generative models is critical for digital forensics and security. While most existing attribution methods can distinguish images produced by known models and reject those from unknown ones, they are unable to verify whether a given image was produced by a specific, previously unseen model. To address this limitation, we formulate an open-set verification problem: determining whether a given image was generated by a specific model. Our key insight is that synthetic images from different models show consistent, content-independent fingerprints in their amplitude spectrum. Based on this insight, we design a dynamic fingerprint simulator capable of simulating over 1.6 trillion generative model architectures. We further train an extractor to capture model-specific fingerprint representations with supervised contrastive learning, enabling accurate attribution of synthetic images, even from previously unseen models. Our method does not rely on any synthetic images, instead, it is trained solely on real images. On DMDetection and AIGCBenchmark, which comprises dozens of state-of-the-art and in-the-wild generative models, our method improves the attribution performance (AUC) of the prior method from random level to 94.05% and 83.05%, respectively. On GenImage and OSMA datasets, we obtain 85.08%, and 88.48% OSCR, outperforming the SOTA methods by 4.30% and 9.37% under the same settings. Jianwei Fei, Yunshu Dai, Peipeng Yu, Zhihua Xia, Dasara Shullani, Daniele Baracchi, Alessandro Piva |
AAAI | 3 |
| 2026 | Fine-Grained DINO Tuning with Dual Supervision for Face Forgery DetectionabstractThe proliferation of sophisticated deepfakes poses significant threats to information integrity. While DINOv2 shows promise for detection, existing fine-tuning approaches treat it as generic binary classification, overlooking distinct artifacts inherent to different deepfake methods. To address this, we propose a DeepFake Fine-Grained Adapter (DFF-Adapter) for DINOv2. Our method incorporates lightweight multi-head LoRA modules into every transformer block, enabling efficient backbone adaptation. DFF-Adapter simultaneously addresses authenticity detection and fine-grained manipulation type classification, where classifying forgery methods enhances artifact sensitivity. We introduce a shared branch propagating fine-grained manipulation cues to the authenticity head. This enables multi-task cooperative optimization, explicitly enhancing authenticity discrimination with manipulation-specific knowledge. Utilizing only 3.5M trainable parameters, our parameter-efficient approach achieves detection accuracy comparable to or even surpassing that of current complex state-of-the-art methods. Peipeng Yu, Zhihua Xia, Longchen Dai |
AAAI | 2 |
| 2026 | BAM: Backdoor defense based on adversarial mitigation
Run Lu, Peipeng Yu, Zhihua Xia |
Pattern Recognit. | 2 |
| 2026 | SWFTI: Facial template inversion via StyleSwin mapping
Zixuan Shen, Zhihua Xia, Kaikai Gan, Peipeng Yu |
Pattern Recognit. | 4 |
| 2026 | RAFS: Reversible Identity-Anonymization Face Swapping for Provenance TrackingabstractFace-swapping technologies have rapidly emerged as a mainstream AI service across entertainment, social media, and virtual platforms. While current face-swapping methods offer highly realistic results, they also introduce significant privacy risks, as most current approaches require clear target faces, exposing users’ identities. Moreover, the absence of built-in authorization and forensic mechanisms renders these systems incapable of tracing or verifying manipulated content, raising critical issues over accountability and potential misuse. To address these challenges, we propose a privacy-preserving and forensics-enabled face-swapping framework that simultaneously safeguards user identity and enables robust post-hoc face recovery. Instead of relying on visible target faces, our method operates on non-facial target images, fundamentally preventing identity exposure at the source. To ensure provenance traceability, we embed the target face’s features into non-facial regions of the generated image via an imperceptible and reversible encoding scheme. To further enhance robustness, we introduce a mask-distortion simulation layer that bridges pre-/post-swap mask discrepancies and stabilizes face recovery under perturbations. Extensive experiments demonstrate that our method produces realistic face-swapped images without revealing the original identity, while enabling high-fidelity recovery under various adversarial conditions—validating its effectiveness in both privacy protection and forensic traceability. Jiancheng Li, Peipeng Yu, Chip-Hong Chang, Zhangjie Fu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | DFREC: DeepFake Identity Recovery Based on Identity-Aware Masked AutoencoderabstractRecent advances in deepfake forensics have primarily focused on improving the classification accuracy and generalization performance. Despite enormous progress in detection accuracy across a wide variety of forgery algorithms, existing algorithms lack intuitive interpretability and identity traceability to help with forensic investigation. In this paper, we introduce a novel DeepFake Identity Recovery scheme (DFREC) to fill this gap. DFREC aims to recover the pair of source and target faces from a deepfake image to facilitate deepfake identity tracing and reduce the risk of deepfake attacks. It comprises three key components: an Identity Segmentation Module (ISM), a Source Identity Reconstruction Module (SIRM), and a Target Identity Reconstruction Module (TIRM). The ISM segments the input face into distinct source and target face information, and the SIRM reconstructs the source face and extracts latent target identity features with the segmented source information. The background context and latent target identity features are synergetically fused by a Masked Autoencoder in the TIRM to reconstruct the target face. We evaluate DFREC on different high-fidelity face-swapping attacks on FaceForensics++, CelebaMegaFS, FFHQ-E4S, and Celeb-DFv2 datasets, which demonstrate its superior recovery performance over state-of-the-art deepfake recovery algorithms. In addition, DFREC is the only scheme that can recover both pristine source and target faces directly from the forgery image with high fidelity. Peipeng Yu, Jianwei Fei, Zhihua Xia, Chip-Hong Chang |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2025 | IdentityLock: An Identity-aware Backdoor strategy for Face Swapping DefenseabstractDeepFakes have become capable of producing highly realistic fabricated faces, posing significant threats to personal privacy and social security. The uncontrolled spread of such forged content, especially when influential figures are targeted, could lead to catastrophic consequences for society. Existing defense algorithms primarily rely on passive detection and perturbation-based strategies but fall short in preventing the generation of DeepFake content. Inspired by the concept of backdoor attacks, this paper proposes IdentityLock, a Face Swapping defense algorithm based on backdoor strategy. It leverages the identities of protected individuals as triggers, embedding defense backdoors during the training process of face swapping models. For images of ordinary individuals, the backdoor model produces typical face-swapped outputs. However, when processing images of protected individuals, the model outputs the original target images, thereby safeguarding protected identities. Extensive experiments on the VGGFace2 dataset demonstrate the superior protection performance and robustness of IdentityLock, which effectively shields protected individuals without additional information. Peipeng Yu, Zhihua Xia, Run Lu |
ICASSP | 2 |
| 2025 | Scalable Dual Fingerprinting for Hierarchical Attribution of Text-to-Image Models
Jianwei Fei, Yunshu Dai, Peipeng Yu, Zhe Kong, Zhihua Xia |
ICCV | 3 |
| 2025 | Unlocking the Capabilities of Large Vision-Language Models for Generalizable and Explainable Deepfake DetectionabstractCurrent Large Vision-Language Models (LVLMs) have demonstrated remarkable capabilities in understanding multimodal data, but their potential remains underexplored for deepfake detection due to the misalignment of their knowledge and forensics patterns. To this end, we present a novel framework that unlocks LVLMs’ potential capabilities for deepfake detection. Our framework includes a Knowledge-guided Forgery Detector (KFD), a Forgery Prompt Learner (FPL), and a Large Language Model (LLM). The KFD is used to calculate correlations between image features and pristine/deepfake image description embeddings, enabling forgery classification and localization. The outputs of the KFD are subsequently processed by the Forgery Prompt Learner to construct fine-grained forgery prompt embeddings. These embeddings, along with visual and question prompt embeddings, are fed into the LLM to generate textual detection responses. Extensive experiments on multiple benchmarks, including FF++, CDF2, DFD, DFDCP, DFDC, and DF40, demonstrate that our scheme surpasses state-of-the-art methods in generalization performance, while also supporting multi-turn dialogue capabilities. Peipeng Yu, Jianwei Fei, Xuan Feng 0002, Zhihua Xia, Chip-Hong Chang |
ICML | 1 |
| 2025 | MADPHash: Manipulation-Aware Deep Perceptual Hashing using Feature ConsistencyabstractPerceptual hashing has garnered significant attention for its wide-ranging applications in image retrieval and authentication domains. However, existing algorithms often struggle to detect subtle manipulations confined to small regions of an image. In this paper, we introduce a novel framework, Manipulation-Aware Deep Perceptual Hashing (MADPHash), which leverages feature consistency to enhance sensitivity to such subtle manipulations. MADPHash explicitly treats tampered images as a distinct category, incorporates a tampering detection objective into the perceptual hash generation process, and employs a Consistency Constraint Module to amplify discrepancies between tampered and untampered regions. Comprehensive experiments conducted on five benchmark datasets demonstrate that MADPHash significantly improves the detection of subtle manipulations while maintaining robustness against content-preserving transformations, outperforming several state-of-the-art perceptual hashing methods. Lizhi Xiong, Peipeng Yu |
ACM Multimedia | 2 |
| 2025 | CHEAT: A Large-Scale Dataset for Detecting CHatGPT-writtEn AbsTractsabstractThe powerful ability of ChatGPT has caused widespread concern in the academic community. Malicious users could synthesize dummy academic content through ChatGPT, which is extremely harmful to academic rigor and originality. The need to develop ChatGPT-written content detection algorithms calls for large-scale datasets. In this paper, we initially investigate the possible negative impact of ChatGPT on academia, and present a large-scale CHatGPT-writtEn AbsTract dataset (CHEAT) to support the development of detection algorithms. In particular, the ChatGPT-written abstract dataset contains 35,304 synthetic abstracts, with$Generation$,$Polish$, and$Fusion$as prominent representatives. Based on these data, we perform a thorough analysis of the existing text synthesis detection algorithms. We show that ChatGPT-written abstracts are detectable with well-trained detectors, while the detection difficulty increases with more human guidance involved. Peipeng Yu, Xuan Feng 0002, Zhihua Xia |
IEEE Trans. Big Data | 1 |
| 2024 | Learning spatial-frequency interaction for generalizable deepfake detectionabstractAbstract In recent years, face forgery detection has gained significant attention, resulting in considerable advancements. However, most existing methods rely on CNNs to extract artefacts from the spatial domain, overlooking the pervasive frequency‐domain artefacts present in deepfake content, which poses challenges in achieving robust and generalized detection. To address these issues, we propose the dual‐stream frequency—spatial fusion network is proposed for deepfake detection. The dual‐stream frequency‐spatial fusion network consists of three components: the spatial forgery feature extraction module, the frequency forgery feature extraction module, and the spatial–frequency feature fusion module. The spatial forgery feature extraction module employs spatial‐channel attention to extract spatial domain features, targeting artefacts in the spatial domain. The frequency forgery feature extraction module leverages the focused linear attention to detect frequency domain anomalies in internal regions, enabling the identification of generated content. The spatial–frequency feature fusion module then fuses forgery features extracted from both the spatial and frequency domains, facilitating accurate detection of splicing artefacts and internally generated forgeries. This approach enhances the model's ability to more accurately capture forgery characteristics. Extensive experiments on several widely‐used benchmarks demonstrate that our carefully designed network exhibits superior generalization and robustness, significantly improving deepfake detection performance. Tianbo Zhai, Kaiyin Lu, Peipeng Yu, Zhihua Xia |
IET Image Process. | 6 |
| 2023 | A black-box reversible adversarial example for authorizable recognition to shared images
Lizhi Xiong, Peipeng Yu, Yuhui Zheng |
Pattern Recognit. | 3 |
| 2023 | A Privacy-Preserving JPEG Image Retrieval Scheme Using the Local Markov Feature and Bag-of-Words Model in Cloud ComputingabstractThe development of cloud computing attracts a great deal of image owners to upload their images to the cloud server to save the local storage. But privacy becomes a great concern to the owner. A forthright way is to encrypt the images before uploading, which, however, would obstruct the efficient usage of image, such as the Content-Based Image Retrieval (CBIR). In this paper, we propose a privacy-preserving JPEG image retrieval scheme. The image content is protected by a specially-designed image encryption method, which is compatible to JPEG compression and makes no expansion to the final JPEG files. Then, the encrypted JPEG files are uploaded to the cloud, and the cloud can directly extract the features from the encrypted JPEG files for searching similar images. Specifically, big-blocks are first assembled with adjacent 8×8 discrete cosine transform (DCT) coefficient blocks. Then, the big-blocks are permuted and the binary code of DCT coefficients are substituted, so as to disturb the content of image. After receiving the encrypted images, local Markov features are extracted from the encrypted big-blocks, and then the Bag-Of-Words (BOW) model is applied to construct a feature vector with these local features to represent the image, so as to provide the CBIR service to image owner. Experimental results and security analysis demonstrate the retrieval performance and security of our scheme. Peipeng Yu, Jian Tang 0009, Zhihua Xia, Zhetao Li, Jian Weng 0001 |
IEEE Trans. Cloud Comput. | 1 |
| 2023 | Secure Outsourced SIFT: Accurate and Efficient Privacy-Preserving Image SIFT Feature ExtractionabstractCloud computing has become an important IT infrastructure in the big data era; more and more users are motivated to outsource the storage and computation tasks to the cloud server for convenient services. However, privacy has become the biggest concern, and tasks are expected to be processed in a privacy-preserving manner. This paper proposes a secure SIFT feature extraction scheme with better integrity, accuracy and efficiency than the existing methods. SIFT includes lots of complex steps, including the construction of DoG scale space, extremum detection, extremum location adjustment, rejecting of extremum point with low contrast, eliminating of the edge response, orientation assignment, and descriptor generation. These complex steps need to be disassembled into elementary operations such as addition, multiplication, comparison for secure implementation. We adopt a serial of secret-sharing protocols for better accuracy and efficiency. In addition, we design a secure absolute value comparison protocol to support absolute value comparison operations in the secure SIFT feature extraction. The SIFT feature extraction steps are completely implemented in the ciphertext domain. And the communications between the clouds are appropriately packed to reduce the communication rounds. We carefully analyzed the accuracy and efficiency of our scheme. The experimental results show that our scheme outperforms the existing state-of-the-art. Xiang Liu 0020, Xueli Zhao, Zhihua Xia, Peipeng Yu, Jian Weng 0001 |
IEEE Trans. Image Process. | 5 |
| 2022 | Learning Second Order Local Anomaly for General Face Forgery DetectionabstractIn this work, we propose a novel method to improve the generalization ability of CNN-based face forgery detectors. Our method considers the feature anomalies of forged faces caused by the prevalent blending operations in face forgery algorithms. Specifically, we propose a weakly supervised Second Order Local Anomaly (SOLA) learning module to mine anomalies in local regions using deep feature maps. SOLA first decomposes the neighborhood of local features by different directions and distances and then calculates the first and second order local anomaly maps which provide more general forgery traces for the classifier. We also propose a Local Enhancement Module (LEM) to improve the discrimination between local features of real and forged regions, so as to ensure accuracy in calculating anomalies. Besides, an improved Adaptive Spatial Rich Model (ASRM) is introduced to help mine subtle noise features via learnable high pass filters. With neither pixel level annotations nor external synthetic data, our method using a simple ResNet18 backbone achieves competitive performances compared with state-of-the-art works when evaluated on unseen forgeries. Jianwei Fei, Yunshu Dai, Peipeng Yu, Tianrun Shen, Zhihua Xia, Jian Weng 0001 |
CVPR | 3 |
| 2022 | Improving Generalization by Commonality Learning in Face Forgery DetectionabstractThis paper proposes a commonality learning strategy for face video forgery detection to improve the generalization. Considering various face forgery methods could leave certain similar forgery traces in videos, we attempt to learn the common forgery features from different forgery databases, so as to achieve better generalization in the detection of unknown forgery methods. Firstly, the Specific Forgery Feature Extractors (SFFExtractors) are trained separately for each of given forgery methods. We utilize the U-net structure and consider the triplet loss, location loss, classification loss, and automatic weighted loss to ensure the detection ability of SFFExtractors on the corresponding forgery methods. Next, the Common Forgery Feature Extractor (CFFExtractor) is trained under the supervision of SFFExtractors to explore the commonality of the forgery traces caused by different forgery methods. The extracted common forgery feature is expected to have a good generalization. The experimental results on FaceForensic++ show that the SFFExtractors outperform many state-of-the-arts in face forgery detection. The generalization performance of the CFFExtractor is verified on FaceForensic++, DFDC, and CelebDF. It is proved that commonality learning can be an effective strategy to improve generalization. Peipeng Yu, Jianwei Fei, Zhihua Xia, Zhili Zhou 0001, Jian Weng 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2021 | Exposing AI-generated videos with motion magnification
Jianwei Fei, Zhihua Xia, Peipeng Yu, Fengjun Xiao |
Multim. Tools Appl. | 3 |
| 2021 | PLDP: Personalized Local Differential Privacy for Multidimensional Data AggregationabstractThe collection of multidimensional crowdsourced data has caused a public concern because of the privacy issues. To address it, local differential privacy (LDP) is proposed to protect the crowdsourced data without much loss of usage, which is popularly used in practice. However, the existing LDP protocols ignore users’ personal privacy requirements in spite of offering good utility for multidimensional crowdsourced data. In this paper, we consider the personality of data owners in protection and utilization of their multidimensional data by introducing the notion of personalized LDP (PLDP). Specifically, we design personalized multiple optimized unary encoding (PMOUE) to perturb data owners’ data, which satisfies ϵ total -PLDP. Then, the aggregation algorithm for frequency estimation on multidimensional data under PLDP is developed, which is described in two situations. Experiments are conducted on four real datasets, and the results show that the proposed aggregation algorithm yields high utility. Moreover, case studies with four real datasets demonstrate the efficiency and superiority of the proposed scheme. Zixuan Shen, Zhihua Xia, Peipeng Yu |
Secur. Commun. Networks | 3 |