Decheng Liu

dblp:191/6552 · DBLP profile ↗
← Back
72ranked-venue papers
25as first author
64since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 33 · 10 first-author · 27 since 2021Artificial intelligence and machine learning · 24 · 11 first-author · 21 since 2021Security and privacy · 15 · 5 first-author · 15 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 MGFFD-VLM: Multi-Granularity Prompt Learning for Face Forgery Detection with VLM
abstract
The proliferation of face forgery content on the Web poses a severe threat to online trust, social media security, and the credibility of digital information. Existing detection approaches often fail to generalize across diverse forgery types and unseen scenarios commonly encountered in web-scale applications. Recent studies have utilized visual large language models (VLMs) to answer not only ''Is this face a forgery?'' but also ''Why is the face a forgery?'' These studies introduced forgery-related attributes, such as forgery location and type, to construct deepfake VQA datasets and train VLMs, achieving high accuracy while providing human-understandable explanatory text descriptions. However, these methods still have limitations. For example, they do not fully leverage face quality-related attributes, which are often abnormal in forged faces, and they lack effective training strategies for forgery-aware VLMs. In this paper, we extend the VQA dataset to create DD-VQA+, which features a richer set of attributes and a more diverse range of samples. Furthermore, we introduce a novel forgery detection framework, MGFFD-VLM, which integrates an Attribute-Driven Hybrid LoRA Strategy to enhance the capabilities of Visual Large Language Models (VLMs). Additionally, our framework incorporates Multi-Granularity Prompt Learning and a Forgery-Aware Training Strategy. By transforming classification and forgery segmentation results into prompts, our method not only improves forgery classification but also enhances interpretability. To further boost detection performance, we design multiple forgery-related auxiliary losses. Experimental results demonstrate that our approach surpasses existing methods in both text-based forgery judgment and analysis, achieving superior accuracy.
Decheng Liu, Chunlei Peng
WWW3
2026 Symmetrical bidirectional knowledge alignment for zero-shot sketch-based image retrieval
abstract
This paper studies the problem of zero-shot sketch-based image retrieval (ZS-SBIR), which aims to use sketches from unseen categories as queries to match the images of the same category. Due to the large cross-modality discrepancy, ZS-SBIR is still a challenging task and mimics realistic zero-shot scenarios. The key is to leverage transferable knowledge from the pre-trained model to improve generalizability. Existing researchers often utilize the simple fine-tuning training strategy or knowledge distillation from a teacher model with fixed parameters, lacking efficient bidirectional knowledge alignment between student and teacher models simultaneously for better generalization. In this paper, we propose a novel Symmetrical Bidirectional Knowledge Alignment for zero-shot sketch-based image retrieval (SBKA). The symmetrical bidirectional knowledge alignment learning framework is designed to effectively learn mutual rich discriminative information between teacher and student models to achieve the goal of knowledge alignment. Instead of the former one-to-one cross-modality matching in the testing stage, a one-to-many cluster cross-modality matching method is proposed to leverage the inherent relationship of intra-class images to reduce the adverse effects of the existing modality gap. Experiments on several representative ZS-SBIR datasets (Sketchy Ext dataset, TU-Berlin Ext dataset and QuickDraw Ext dataset) prove the proposed algorithm can achieve superior performance compared with state-of-the-art methods. The source code is publicly available at https://github.com/zermatt-luo/SBKA.
Decheng Liu, Chunlei Peng, Nannan Wang 0001, Ruimin Hu, Xinbo Gao 0001
Neural Networks1
2026 TransFA: Transformer-based representation for face attribute evaluation
Decheng Liu, Chunlei Peng, Nannan Wang 0001, Jie Li 0001, Xinbo Gao 0001
Pattern Recognit.1
2026 FST: Improving adversarial robustness via feature similarity-based targeted adversarial training
Dawei Zhou 0004, Decheng Liu, Nannan Wang 0001, Xinbo Gao 0001
Pattern Recognit.3
2026 Adaptive Consensus Multi-Teacher Distillation for Generalizable Face Forgery Detection
abstract
Recent advancements in deep learning have significantly lowered the cost of generating and processing facial images, but this progress has also presented a growing challenge for face forgery detection. Existing detection methods, however, often struggle to generalize effectively to forged samples generated by previously unseen forgery techniques. Additionally, transferring knowledge through knowledge distillation to enhance generalization remains a persistent challenge. To address these issues, this paper proposes Adaptive Consensus Multi-teacher Knowledge Distillation (ACMD), a novel framework aimed at improving the generalization capabilities of face forgery detection models. ACMD leverages multiple teacher models with varying network architectures and introduces a consensus mechanism to resolve knowledge conflicts arising from the diversity of teacher models. It further adapts the assignment of sample-specific distillation weights, based on each teacher model’s prediction confidence. By combining this with feature-based knowledge distillation, the student model gains a deeper understanding of the teacher’s knowledge. Extensive experiments on the DeepfakeBench benchmark demonstrate that ACMD not only outperforms a single teacher model but also achieves state-of-the-art performance in dataset evaluations. Specifically, in a cross-dataset evaluation from FaceForensics++ to CDFv2 and DFDC, ACMD achieves frame-level AUCs of 84.22% and 74.43%, respectively, surpassing all baseline models. Ablation studies further validate the effectiveness of each component and highlight their complementary contributions to the overall detection performance.
Jiuyao Jing, Chunlei Peng, Decheng Liu, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Circuits Syst. Video Technol.4
2026 DeepFidelity: Perceptual Forgery Fidelity Assessment for Deepfake Detection
abstract
Deepfake detection refers to detecting artificially generated or edited faces in images or videos, which plays an essential role in visual information security. Despite promising progress in recent years, Deepfake detection remains a challenging problem due to the complexity and variability of face forgery techniques. Existing Deepfake detection methods are often devoted to extracting features by designing sophisticated networks but ignore the influence of perceptual quality of faces. Considering the complexity of the quality distribution of real and fake faces, we propose a deepfake detection framework called DeepFidelity, which mines the perceptual forgery fidelity of face images and introduces a quality-aware scoring mechanism to distinguish real and fake faces of different image qualities. Specifically, we improve the model’s ability to identify complex samples by mapping real and fake face data of different qualities to different scores to distinguish them in a more detailed way. In addition, we propose a network structure called Symmetric Spatial Attention Augmentation based vision Transformer (SSAAFormer), which uses the symmetry of face images to promote the network to model the geographic long-distance relationship at the shallow level and augment local features. Extensive experiments on multiple benchmark datasets demonstrate the superiority of the proposed method over state-of-the-art methods. The code is available athttps://github.com/shimmer-ghq/DeepFidelity.
Chunlei Peng, Huiqing Guo, Decheng Liu, Nannan Wang 0001, Ruimin Hu, Xinbo Gao 0001
IEEE Trans. Circuits Syst. Video Technol.3
2026 RFA-Tex: Range-Flexible Adaptive Physical Adversarial Texture Against Real-World Person Detectors
Mengyao Zhu 0004, Xinghua Li 0001, Decheng Liu, Shunjie Yuan, Yigang Li, Yinbin Miao, Robert H. Deng
IEEE Trans. Inf. Forensics Secur.3
2025 Thinking Racial Bias in Fair Forgery Detection: Models, Datasets and Evaluations
abstract
Due to the successful development of deep image generation technology, forgery detection plays a more important role in social and economic security. Racial bias has not been explored thoroughly in the deep forgery detection field. In the paper, we first contribute a dedicated dataset called the Fair Forgery Detection (FairFD) dataset, where we prove the racial bias of public state-of-the-art (SOTA) methods. Different from existing forgery detection datasets, the self-constructed FairFD dataset contains a balanced racial ratio and diverse forgery generation images with the largest-scale subjects. Additionally, we identify the problems with naive fairness metrics when benchmarking forgery detection models. To comprehensively evaluate fairness, we design novel metrics including Approach Averaged Metric and Utility Regularized Metric, which can avoid deceptive results. We also present an effective and robust post-processing technique, Bias Pruning with Fair Activations (BPFA), which improves fairness without requiring retraining or weight updates. Extensive experiments conducted with 12 representative forgery detection models demonstrate the value of the proposed dataset and the reasonability of the designed fairness metrics. By applying the BPFA to the existing fairest detector, we achieve a new SOTA. Furthermore, we conduct more in-depth analyses to offer more insights to inspire researchers in the community.
Decheng Liu, Zongqi Wang, Chunlei Peng, Nannan Wang 0001, Ruimin Hu, Xinbo Gao 0001
AAAI1
2025 Mitigating Feature Gap for Adversarial Robustness by Feature Disentanglement
abstract
Adversarial fine-tuning methods enhance adversarial robustness via fine-tuning the pre-trained model in an adversarial training manner. However, we identify that some specific latent features of adversarial samples are confused by adversarial perturbation and lead to an unexpectedly increasing gap between features in the last hidden layer of natural and adversarial samples. To address this issue, we propose a disentanglement-based approach to explicitly model and further remove the specific latent features. We introduce a feature disentangler to separate out the specific latent features from the features of the adversarial samples, thereby boosting robustness by eliminating the specific latent features. Besides, we align clean features in the pre-trained model with features of adversarial samples in the fine-tuned model, to benefit from the intrinsic features of natural samples. Empirical evaluations on three benchmark datasets demonstrate that our approach surpasses existing adversarial fine-tuning methods and adversarial training baselines.
Nuoyan Zhou, Dawei Zhou 0004, Decheng Liu, Nannan Wang 0001, Xinbo Gao 0001
AAAI3
2025 Phase and Amplitude-aware Prompting for Enhancing Adversarial Robustness
abstract
Deep neural networks are found to be vulnerable to adversarial perturbations. The prompt-based defense has been increasingly studied due to its high efficiency. However, existing prompt-based defenses mainly exploited mixed prompt patterns, where critical patterns closely related to object semantics lack sufficient focus. The phase and amplitude spectra have been proven to be highly related to specific semantic patterns and crucial for robustness. To this end, in this paper, we propose a Phase and Amplitude-aware Prompting (PAP) defense. Specifically, we construct phase-level and amplitude-level prompts for each class, and adjust weights for prompting according to the model’s robust performance under these prompts during training. During testing, we select prompts for each image using its predicted label to obtain the prompted image, which is inputted to the model to get the final prediction. Experimental results demonstrate the effectiveness of our method.
Dawei Zhou 0004, Decheng Liu, Nannan Wang 0001
ICML3
2025 Cross-Domain Matrix Compression Adaptation for Zero-Shot Sketch-Based Image Retrieval
abstract
This paper explores an efficient parameter fine-tuning strategy for zero-shot sketch-based image retrieval (ZS-SBIR). We highlight a key finding: through efficient parameter fine-tuning, competitive retrieval accuracy can be achieved using only 12% of the full parameters. This superior performance is attributed to low-rank matrix compression in the redundant parameter space, which extracts more discriminative effective feature dimensions. Specifically, to fully leverage the potential of pretrained models, we propose a cross-domain matrix compression adaptation method. For lower-level features, we apply a general low-rank decomposition to extract shared basic shapes or contours information across modalities. To mitigate overfitting to local similarities, we propose a domain-specific matrix compression module that guides the model in learning high-level abstractions essential for sketch retrieval. Our method is simple and effective, balancing both general semantic information and feature variations across domains within the same category. Experimental results on the ZS-SBIR benchmark dataset show that our method not only outperforms existing state-of-the-art methods, but also requires significantly fewer training parameters.
Decheng Liu, Yu Zheng 0006, Chunlei Peng
IJCNN1
2025 Convergence and Optimization of Wireless Federated Low-Rank Adaptation with Imperfect CSI
abstract
The substantial number of parameters in large artificial intelligence (AI) models imposes significant transmission burden during the fine-tuning process over wireless networks. In this paper, we propose a wireless federated low-rank adaptation (Fed-LoRA) framework for the fine-tuning of large AI models under imperfect channel state information (CSI). By analyzing the fine-tuning error caused by transmission outages and rank integrity in gradient aggregation, we capture the convergence behavior of the proposed Fed-LoRA framework. Based on the analysis results, we formulate an optimization problem to enhance the fine-tuning performance of Fed-LoRA by minimizing an optimality gap. Then, we design a low-complexity resource allocation scheme to solve this problem. Simulation results on two large AI models demonstrate that Fed-LoRA outperforms other baselines on both convergence performance and fine-tuning loss by effectively balancing transmission outage and the rank integrity in gradient aggregation.
Zhitong Zhou, Haofeng Sun, Decheng Liu
PIMRC4
2025 Imperceptible Face Forgery Attack via Adversarial Semantic Mask
Qixuan Su, Decheng Liu, Chunlei Peng
PRCV (7)3
2025 Toward trustworthy identity tracing via multi-attribute synergistic identification
Wenbin Feng, Decheng Liu, Ruimin Hu
Inf. Sci.2
2025 Revisiting face forgery detection towards generalization
Chunlei Peng, Decheng Liu, Huiqing Guo, Nannan Wang 0001, Xinbo Gao 0001
Neural Networks3
2025 FairForensics: mitigating attribute bias in deepfake detection by integrating texture and attribute features
Chunlei Peng, Yinyin Chen, Decheng Liu, Nannan Wang 0001, Xinbo Gao 0001
Neural Networks3
2025 Frequency-Based Comprehensive Prompt Learning for Vision-Language Models
abstract
This paper targets to learn multiple comprehensive text prompts that can describe the visual concepts from coarse to fine, thereby endowing pre-trained VLMs with better transfer ability to various downstream tasks. We focus on exploring this idea on transformer-based VLMs since this kind of architecture achieves more compelling performances than CNN-based ones. Unfortunately, unlike CNNs, the transformer-based visual encoder of pre-trained VLMs cannot naturally provide discriminative and representative local visual information. To solve this problem, we propose Frequency-based Comprehensive Prompt Learning (FCPrompt) to excavate representative local visual information from the redundant output features of the visual encoder. FCPrompt transforms these features into frequency domain via Discrete Cosine Transform (DCT). Taking the advantages of energy concentration and information orthogonality of DCT, we can obtain compact, informative and disentangled local visual information by leveraging specific frequency components of the transformed frequency features. To better fit with transformer architectures, FCPrompt further adopts and optimizes different text prompts to respectively align with the global and frequency-based local visual information via a dual-branch framework. Finally, the learned text prompts can thus describe the entire visual concepts from coarse to fine comprehensively. Extensive experiments indicate that FCPrompt achieves the state-of-the-art performances on various benchmarks.
Liangchen Liu 0001, Nannan Wang 0001, Chen Chen 0128, Decheng Liu, Xi Yang 0011, Xinbo Gao 0001, Tongliang Liu
IEEE Trans. Pattern Anal. Mach. Intell.4
2025 Devil in Shadow: Attacking NIR-VIS Heterogeneous Face Recognition via Adversarial Shadow
abstract
Near infrared-visible (NIR-VIS) heterogeneous face recognition aims to match face identities in cross-modality settings, which has achieved significant development recently. The work on adversarial attack and security issues of the heterogeneous face recognition task is still lacking. Existing adversarial face generation methods can’t deploy directly because of the inevitable large modality discrepancy. Besides, the ideal adversarial attacking generated images should maintain both high capabilities and low detectability. Considering the properties of near-infrared face images, our basic idea is to construct adversarial shadows for good stealthiness and high attack capability. In this paper, we propose a novel face adversarial shadow generation framework for NIR-VIS heterogeneous face recognition, which can synthesize fine-crafted lighting conditions containing strong identity attacking ability. Specifically, we design the variance consistency-based symmetric face attacking loss to improve the attacking generalization and the synthesized image quality. Extensive qualitative and quantitative experiments on the public large-scale NIR-VIS heterogeneous face dataset prove the proposed method achieves superior performance compared with the state-of-the-art methods. The source code is publicly available athttps://github.com/GEaMU/Devil-in-Shadow.
Decheng Liu, Rong Sheng, Chunlei Peng, Nannan Wang 0001, Ruimin Hu, Xinbo Gao 0001
IEEE Trans. Circuits Syst. Video Technol.1
2025 A Knowledge-Guided Adversarial Defense for Resisting Malicious Visual Manipulation
abstract
Malicious applications of visual manipulation have raised serious threats to the security and reputation of users in many fields. To alleviate these issues, adversarial noise-based defenses have been enthusiastically studied in recent years. However, “data-only” methods tend to distort fake samples in the low-level feature space rather than the high-level semantic space, leading to limitations in resisting malicious manipulation. Frontier research has shown that integrating knowledge in deep learning can produce reliable and generalizable solutions. Inspired by these, we propose aknowledge-guided adversarial defense (KGAD) to actively force malicious manipulation models to output semantically confusing samples. Specifically, in the process of generating protective adversarial noise, we focus on constructing significant semantic confusions at the domain-specific knowledge level, and exploit a metric closely related to visual perception to replace the general pixel-wise metrics. The generated adversarial noise can actively interfere with the malicious manipulation model by triggering knowledge-guided and perception-related disruptions in the fake samples. To validate the effectiveness of the proposed method, we conduct qualitative and quantitative experiments on human perception and visual quality assessment. The results on two different tasks both show our defense achieves competitive performances and generalizability, indicating that it can effectively resist malicious visual manipulation.
Dawei Zhou 0004, Zhigang Su, Decheng Liu, Tongliang Liu, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Dependable Secur. Comput.3
2025 Masked Text Adversarial Training for Cloth-Changing Person Re-Identification
Chengrui Hao, Chunlei Peng, Decheng Liu, Nannan Wang 0001, Ruimin Hu, Xinbo Gao 0001
IEEE Trans. Inf. Forensics Secur.4
2025 Attention Consistency Refined Masked Frequency Forgery Representation for Generalizing Face Forgery Detection
abstract
Due to the successful development of deep image generation technology, visual data forgery detection would play a more important role in social and economic security. Existing forgery detection methods suffer from unsatisfactory generalization ability to determine the authenticity in the unseen domain. In this paper, we propose a novel Attention Consistency Refined masked frequency forgery representation model toward a generalizing face forgery detection algorithm (ACMF). Most forgery technologies always bring in high-frequency aware cues, which make it easy to distinguish source authenticity but difficult to generalize to unseen artifact types. The masked frequency forgery representation module is designed to explore robust forgery cues by randomly discarding high-frequency information. In addition, we find that the forgery saliency map inconsistency through the detection network could affect the generalizability. Thus, the forgery attention consistency is introduced to force detectors to focus on similar attention regions for better generalization ability. Experiment results on several public face forgery datasets (FaceForensic++, DFD, Celeb-DF, WDF and DFDC datasets) demonstrate the superior performance of the proposed method compared with the state-of-the-art methods. The source code and models are publicly available athttps://github.com/chenboluo/ACMF.
Decheng Liu, Chunlei Peng, Nannan Wang 0001, Ruimin Hu, Xinbo Gao 0001
IEEE Trans. Inf. Forensics Secur.1
2025 Improving Adversarial Robustness via Decoupled Visual Representation Masking
abstract
Deep neural networks are proven to be vulnerable to finely designed adversarial examples, and adversarial defense algorithms draw more and more attention nowadays. Pre-processing based defense is a major strategy, as well as learning robust feature representation, has been proven an effective way to boost generalization. However, existing defense works lack considering different depth-level visual features in the training process. In this paper, we first highlight two novel properties of robust features from the feature distribution perspective: 1) Diversity (robust features within the same class should maintain appropriate variety). 2) Discriminability (robust features from different classes should be sufficiently separated). We find that state-of-the-art defense methods aim to address both of these mentioned issues well. It motivates us to increase intra-class variance and decrease inter-class discrepancy simultaneously in adversarial training. Specifically, we propose a simple but effective defense based on decoupled visual representation masking. The designed Decoupled Visual Feature Masking (DFM) block can adaptively disentangle visual discriminative features and non-visual features with diverse mask strategies, while the suitable discarding information can disrupt adversarial noise to improve robustness. Our work provides a generic and easy-to-plugin block unit for any former adversarial training algorithm to achieve better protection integrally. Extensive experimental results prove that the proposed method can achieve superior performance compared with state-of-the-art defense approaches. The code is publicly available at https://github.com/chenboluo/Adversarial-defense.
Decheng Liu, Chunlei Peng, Nannan Wang 0001, Ruimin Hu, Xinbo Gao 0001
IEEE Trans. Inf. Forensics Secur.1
2025 Toward Fair Adversarial Defense via Class Encourage-Suppress Robust Learning
abstract
Deep Neural Networks with natural training are quite vulnerable to adversarial attacks, so it’s necessary to defend these attacks with effective defense methods like adversarial training. However, while defending against adversarial attacks, adversarially trained models’ robustness between classes shows severe unfairness. To mitigate the disparity, a lot of methods have been proposed, while many of them sacrifice the overall accuracy to leverage the worst-class accuracy. Inspired by previous works, we propose a new fairness methodology, and we name it Class Encourage-suppress Robust Learning (CRL). Based on the overall accuracy and the class-wise accuracies of the dataset, we introduce a new measurement named Class Diversity Ratio to adjust the weights of different classes in the loss function. Additionally, we propose a new learning strategy called the Competitor Encourage-suppress Strategy to mitigate the disparity between diverse classes, which is simple but effective. Experimental results on representative datasets show that our method outperforms state-of-the-art (SOTA) methods.
Decheng Liu, Chunlei Peng, Nannan Wang 0001, Ruimin Hu, Xinbo Gao 0001
IEEE Trans. Inf. Forensics Secur.1
2025 Semantic Token Transformer for Face Forgery Detection
abstract
In the era of digital media, the proliferation of forged images and videos poses a significant threat to societal stability. With the rapid advancement of deep learning, the generation of realistic fake images has become increasingly simple, presenting unprecedented challenges in discerning the authenticity of images. While some existing methods have shown promising results in forgery detection, they often underutilize facial semantic information. To address this issue, this paper introduces the Semantic Token Transformer for Face Forgery Detection. By incorporating facial semantic information with a transformer network, the input tokens of the transformer are transformed into tokens of varying shapes and sizes based on their importance, thereby enhancing the accuracy of the detector. To achieve this objective, we first employ an image processing stage to manipulate the image based on facial semantic information. Subsequently, we introduce a scoring network, guided by prior knowledge, which adaptively categorizes tokens into different clusters based on their importance and relevance to the results of the preprocessing stage. Finally, we merge the tokens within the clusters using an attention mechanism and input them into the detector for forgery detection. Through experiments conducted on multiple datasets and cross-dataset evaluations, we demonstrate that our approach outperforms state-of-the-art detection methods.
Chunlei Peng, Xiaoyi Luo, Decheng Liu, Nannan Wang 0001, Ruimin Hu, Xinbo Gao 0001
IEEE Trans. Inf. Forensics Secur.3
2025 Within 3DMM Space: Exploring Inherent 3D Artifact for Video Forgery Detection
abstract
Recently, the breathtaking development and potential misuse of deepfake technology has raised numerous privacy and security concerns, triggering widespread apprehension. Existing deepfake detection methods focus on the analysis of local regions for faces, such as mouth movement, eye blinking frequency, etc., which, however, are limited in their ability to capture the global inconsistencies present in forged faces. Some researchers attempt to seize 3D artifacts related to facial global information, but typically treat the 3D information as mere input, lacking the in-depth analysis. To address these shortcomings and mine the inherent and delicate 3D artifacts in the forged faces, this paper innovatively proposes the 3D Artifact Detector (3DAD) method, which leverages the spatio-temporal inconsistency on the 3D semantic space in the forgery videos to uncover the deepfake clues. Specifically, we employ 3D Analysis Unit (3DAU) to pre-train the face reconstruction task within 3D Morphable Model (3DMM) space, thereby obtaining the high-level inherent 3d representation. Concurrently, for the multi-levels of information in the face, we utilize the Texture Perception Unit (TPU) to extract the texture information in the low-level semantic space of the images. Ultimately we feed the two distinct modalities into the spatiotemporal fusion model for final detection. Through extensive intra- and cross-dataset experiments on publicly available datasets, we demonstrate the effectiveness and generalizability of the proposed method. The source code is available at https://github.com/Cookie-XT/3DAD.
Chunlei Peng, Decheng Liu, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Inf. Forensics Secur.3
2025 PrivacyHFR: Visual Privacy Preserving for Heterogeneous Face Recognition
abstract
Face recognition has achieved remarkable progress and is widely deployed in real-world scenarios. Recently more and more attention has been given to individual privacy protection, due to unauthorized sensitive image leakage by malicious attackers. Multi-modality face images captured by diverse sensors, also called heterogeneous faces, bring in more challenges in face privacy protection while lacking related research. In this paper, we propose a novel visual Privacy preserving method for Heterogeneous Face Recognition (Privacy-HFR) to protect perceptual visual information and maintain essential identity information in multi-modality face analysis scenarios. Frequency domain analysis is a vital strategy to bridge the inevitable modality gap for heterogeneous face images. Meanwhile, recent theoretical insights also inspire us to design a suitable frequency component adjustment to balance human visual sensitivity and identity discriminative information. In addition, the ability to defend against recovery attacks has emerged as an essential criterion for privacy preserving face recognition. Noting that there seems to exist a dilemma that reducing accessible information by the attack model will affect the extracted identity information for recognition. It is because these two kinds of information are mutually blended in the frequency domain, which makes it a challenge to simultaneously maintain visual privacy and identity distinguishability. Thus, we provide a novel perspective to leverage the randomly optimal solutions and design the specific adversarial perturbations against the recovery attack. Experiments on several large-scale heterogeneous face datasets (CASIA NIR-VIS 2.0, LAMP-HQ, Tufts Face and CUFSF datasets) prove that the proposed method outperforms existing privacy-preserving face recognition methods in terms of recognition accuracy and privacy protection capability. The code is available in https://github.com/xiyin11/Privacy-HFR.
Decheng Liu, Weizhao Yang, Chunlei Peng, Nannan Wang 0001, Ruimin Hu, Xinbo Gao 0001
IEEE Trans. Image Process.1
2025 SketchAging: Face Photo-Sketch Synthesis and Aging With Multi-Scale Feature Extraction
abstract
With the rapid development of Generative Adversarial Networks (GANs), facial sketch generation and age transformation have advanced considerably. These technologies show great potential in digital media, entertainment, and forensic applications, particularly in helping law enforcement reconstruct the appearance of long-term fugitives. However, current methodologies exhibit notable limitations: existing approaches typically specialize in either facial sketch generation or age progression independently, lacking an effective integration for cross-domain synthesis. Moreover, preserving identity information while ensuring high-quality image generation remains a challenge. This paper proposes Multi-Scale Feature Extraction Networks (MSFE), an image-to-image translation framework that enables continuous age transformation while maintaining the stylistic characteristics of sketch domains. The core of the MSFS network uses a Dual Conditional Normalization Attention (DCNA) architecture to extract sketch features and encode facial images into the latent space of a pre-trained StyleGAN based on the desired age change. Experimental results on public datasets demonstrate that our approach outperforms existing methods, achieving superior facial photo-sketch synthesis with enhanced realism, identity preservation, and age accuracy.
Chunlei Peng, Zhuang Tang, Decheng Liu, Nannan Wang 0001, Ruimin Hu, Xinbo Gao 0001
IEEE Trans. Image Process.3
2025 Face Forgery Detection With CLIP-Enhanced Multi-Encoder Distillation
abstract
With the development of face forgery technology, fake faces are rampant, threatening the security and authenticity of many fields. Therefore, it is of great significance to study face forgery detection. At present, existing detection methods have deficiencies in the comprehensiveness of feature extraction and model adaptability, and it is difficult to accurately deal with complex and changeable forgery scenarios. However, the rise of multimodal models provides new insights for current forgery detection methods. At present, most methods use relatively simple text prompts to describe the difference between real and fake faces. However, these researchers ignore that the CLIP model itself does not have the relevant knowledge of forgery detection. Therefore, our paper proposes a face forgery detection method based on multi-encoder fusion and cross-modal knowledge distillation. On the one hand, the prior knowledge of the CLIP model and the forgery model is fused. On the other hand, through the alignment distillation, the student model can learn the visual abnormal patterns and semantic features of the forged samples captured by the teacher model. Specifically, our paper extracts the features of face photos by fusing the CLIP text encoder and the CLIP image encoder, and uses the dataset in the field of forgery detection to pretrain and fine-tune the Deepfake-V2-Model to enhance the detection ability, which are regarded as the teacher model. At the same time, the visual and language patterns of the teacher model are aligned with the visual patterns of the pretrained student model, and the aligned representations are refined to the student model. This not only combines the rich representation of the CLIP image encoder and the excellent generalization ability of text embedding, but also enables the original model to effectively acquire relevant knowledge for forgery detection. Experiments show that our method effectively improves the performance on face forgery detection.
Chunlei Peng, Tianzhe Yan, Decheng Liu, Nannan Wang 0001, Ruimin Hu, Xinbo Gao 0001
IEEE Trans. Image Process.3
2025 Generalizable Prompt Learning via Gradient Constrained Sharpness-Aware Minimization
abstract
This paper targets a novel trade-off problem in generalizable prompt learning for vision-language models (VLM), i.e., improving the performance on unseen classes while maintaining the performance on seen classes. Comparing with existing generalizable methods that neglect the seen classes degradation, the setting of this problem is stricter and fits more closely with practical applications. To solve this problem, we start from the optimization perspective, and leverage the relationship between loss landscape geometry and model generalization ability. By analyzing the loss landscapes of the state-of-the-art method and vanilla Sharpness-aware Minimization (SAM) based method, we conclude that the trade-off performance correlates to bothloss valueandloss sharpness, while each of them is indispensable. However, we find the optimizing gradient of existing methods cannot maintain high relevance to both loss value and loss sharpness during optimization, which severely affects their trade-off performance. To this end, we propose a novel SAM-based method for prompt learning, denoted as Gradient Constrained Sharpness-aware Context Optimization (GCSCoOp), to dynamically constrain the optimizing gradient, thus achieving above two-fold optimization objective simultaneously. Extensive experiments verify the effectiveness of GCSCoOp in the trade-off problem.
Liangchen Liu 0001, Nannan Wang 0001, Dawei Zhou 0004, Decheng Liu, Xi Yang 0011, Xinbo Gao 0001, Tongliang Liu
IEEE Trans. Multim.4
2025 Masked Attribute Description Embedding for Cloth-Changing Person Re-Identification
abstract
Cloth-changing person re-identification (CC-ReID) aims to match persons who change clothes over long periods. The key challenge in CC-ReID is to extract cloth-irrelated features, such as face, hairstyle, body shape, and gait. Current research mainly focuses on modeling body shape using multi-modal biological features (such as silhouettes and sketches). However, it does not fully leverage the personal description information hidden in the original RGB image. Considering that there are certain attribute descriptions that remain unchanged after the changing of cloth, we propose a Masked Attribute Description Embedding (MADE) method that unifies personal visual appearance and attribute description for CC-ReID. Specifically, handling variable cloth-sensitive information, such as color and type, is challenging for effective modeling. To address this, we mask the clothes type and color information (upper body type, upper body color, lower body type, and lower body color) in the personal attribute description extracted through an attribute detection model. The masked attribute description is then connected and embedded into Transformer blocks at various levels, fusing it with the low-level to high-level features of the image. This approach compels the model to discard cloth information. Experiments are conducted on several CC-ReID benchmarks, including PRCC, LTCC, Celeb-reID-light, and LaST. Results demonstrate that MADE effectively utilizes attribute description, enhancing cloth-changing person re-identification performance, and compares favorably with state-of-the-art methods.
Chunlei Peng, Decheng Liu, Nannan Wang 0001, Ruimin Hu, Xinbo Gao 0001
IEEE Trans. Multim.3
2024 Adv-Diffusion: Imperceptible Adversarial Face Identity Attack via Latent Diffusion Model
abstract
Adversarial attacks involve adding perturbations to the source image to cause misclassification by the target model, which demonstrates the potential of attacking face recognition models. Existing adversarial face image generation methods still can’t achieve satisfactory performance because of low transferability and high detectability. In this paper, we propose a unified framework Adv-Diffusion that can generate imperceptible adversarial identity perturbations in the latent space but not the raw pixel space, which utilizes strong inpainting capabilities of the latent diffusion model to generate realistic adversarial images. Specifically, we propose the identity-sensitive conditioned diffusion generative model to generate semantic perturbations in the surroundings. The designed adaptive strength-based adversarial perturbation algorithm can ensure both attack transferability and stealthiness. Extensive qualitative and quantitative experiments on the public FFHQ and CelebA-HQ datasets prove the proposed method achieves superior performance compared with the state-of-the-art methods without an extra generative model training process. The source code is available at https://github.com/kopper-xdu/Adv-Diffusion.
Decheng Liu, Xijun Wang 0005, Chunlei Peng, Nannan Wang 0001, Ruimin Hu, Xinbo Gao 0001
AAAI1
2024 Advancing Generalized Deepfake Detector with Forgery Perception Guidance
abstract
One of the serious impacts brought by artificial intelligence is the abuse of deepfake techniques. Despite the proliferation of deepfake detection methods aimed at safeguarding the authenticity of media across the Internet, they mainly consider the improvement of detector architecture or the synthesis of forgery samples. The forgery perceptions, including the feature responses and prediction scores for forgery samples, have not been well considered. As a result, the generalization across multiple deepfake techniques always comes with complicated detector structures and expensive training costs. In this paper, we shift the focus to real-time perception analysis in the training process and generalize deepfake detectors through an efficient method dubbed Forgery Perception Guidance (FPG). In particular, after investigating the deficiencies of forgery perceptions, FPG adopts a sample refinement strategy to pertinently train the detector, thereby elevating the generalization efficiently. Moreover, FPG introduces more sample information as explicit optimizations, which makes the detector further adapt the sample diversities. Experiments demonstrate that FPG improves the generality of deepfake detectors with small training costs, minor detector modifications, and the acquirement of real data only. In particular, our approach not only outperforms the state-of-the-art on both the cross-dataset and cross-manipulation evaluation but also surpasses the baseline that needs more than 3× training time.
Ruiyang Xia, Dawei Zhou 0004, Decheng Liu, Lin Yuan 0002, Shuodi Wang, Jie Li 0001, Nannan Wang 0001, Xinbo Gao 0001
ACM Multimedia3
2024 Spatial-Frequency Dual-Stream Reconstruction for Deepfake Detection
Chunlei Peng, Decheng Liu, Yu Zheng 0006, Nannan Wang 0001
PRCV (11)3
2024 Audio-Driven Face Photo-Sketch Video Generation
Siyue Zhou, Qun Guan, Chunlei Peng, Decheng Liu, Yu Zheng 0006
PRICAI (3)4
2024 Pyramid-resolution person restoration for cross-resolution person re-identification
Chunlei Peng, Decheng Liu, Nannan Wang 0001, Xinbo Gao 0001
Sci. China Inf. Sci.3
2024 GazeForensics: DeepFake detection via gaze-guided spatial inconsistency learning
Qinlin He, Chunlei Peng, Decheng Liu, Nannan Wang 0001, Xinbo Gao 0001
Neural Networks3
2024 Local artifacts amplification for deepfakes augmentation
Chunlei Peng, Decheng Liu, Nannan Wang 0001, Xinbo Gao 0001
Neural Networks3
2024 MRLReID: Unconstrained Cross-Resolution Person Re-Identification With Multi-Task Resolution Learning
abstract
Cross-resolution person re-identification (ReID) is a challenging task that addresses the issue of matching individuals across different resolution conditions. Traditional person ReID methods often assume that images have sufficiently high resolution and overlook the practical scenarios involving low-resolution or blurry images. Existing cross-resolution ReID approaches either utilize image super-resolution techniques to improve the quality of low-resolution images or extract and learn resolution invariant features for person representation. Although multi-task learning has been applied in ReID to integrate auxiliary tasks including attribute recognition, image super-resolution, and so on, how to incorporate the vital resolution learning task into cross-resolution ReID has rarely explored before. Therefore, we propose a novel multi-task resolution learning based ReID network named MRLReID. Our approach treats ross-resolution person ReID as the primary task and the resolution estimation as an auxiliary task. Our network simultaneously learns the resolution information and person identity information of images, aiming to improve cross-resolution person ReID performance. Considering that existing similuated cross-resolution datasets are too simple to mimic unconstrained scenario, we further employ image degradation technique to simulate more realistic cross-resolution ReID datasets. We evaluate our method on two real-world cross-resolution datasets and two newly simulated cross-resolution datasets, and both intra-dataset and cross-dataset evaluations demonstrate the effectiveness and superiority of our method in cross-resolution person ReID. The codes and datasets are available at https://github.com/amateurbo/MRLReID.
Chunlei Peng, Decheng Liu, Nannan Wang 0001, Ruimin Hu, Xinbo Gao 0001
IEEE Trans. Circuits Syst. Video Technol.3
2024 Universal Heterogeneous Face Analysis via Multi-Domain Feature Disentanglement
abstract
Heterogeneous face analysis is an important and challenge problem in face recognition community, because of the large modality discrepancy between heterogeneous face images. Existing methods either focus on transforming heterogeneous faces into the same style via face synthesis process, or intend to directly recognize heterogeneous face via modality invariant descriptors. However, the tasks of cross modality face synthesis and face recognition share a common purpose, which is to disentangle an inherent explainable representation. To this end, we propose a novel universal heterogenous face analysis method via multi-domain feature disentanglement, which does not need any face domain label. The proposed method explores to disentangle factors of variations of cross modality faces in an unsupervised manner. Then we could translate cross modality faces through modifying semantic factors, and the extracted inherent explainable representation still maintains being discriminative for heterogeneous face recognition. Experimental results on multiple cross modality face databases demonstrate the effectiveness of the proposed method. These experimental results also inspire us that the unsupervised disentangled module could help to analyze the interpretability of heterogenous face representation.
Decheng Liu, Xinbo Gao 0001, Chunlei Peng, Nannan Wang 0001, Jie Li 0001
IEEE Trans. Inf. Forensics Secur.1
2024 Where Deepfakes Gaze at? Spatial-Temporal Gaze Inconsistency Analysis for Video Face Forgery Detection
abstract
With the continuous development of generative models on face generation, how to distinguish the real and fake face has become an important problem for security. Because of the continuous improvement on the detection accuracy by facial physiological signals, video face forgery detection based on facial physiological signal analysis has received more and more attention, which has become an important research branch in the field of face forgery detection. Currently, most of the research on forgery detection based on physiological signal analysis use biometric features such as blinking patterns, head swings, heart rate signals, and lip movements. However, there hasn’t been much exploration on the usage of gaze features in face forgery detection. Through the analysis of gaze directions in face videos, we have observed differences in the distribution of gaze direction pattern between the real and forged videos. Specifically, real videos tend to have more concentrated gaze distribution within a short period of time, while forged videos have more dispersed gaze distributions. In this paper, we present a novel Deepfake gaze analysis method named DFGaze, to explore spatial-temporal gaze inconsistency for video face forgery detection. Our method uses the gaze analysis model (GAM) to analyze the gaze features of face video frames, and then applies a spatial-temporal feature aggregator to realize authenticity classification based on gaze features. In order to better mine the authenticity clues in the videos, we further use the texture analysis model (TAM) and attribute analysis model (AAM) to improve the representation ability of spatial-temporal feature differences between real and forged faces. Extensive experiments show that our method can achieve state-of-the-art performance with the help of gaze analysis. The source code is available at https://github.com/ziminMIAO/DFGaze.
Chunlei Peng, Zimin Miao, Decheng Liu, Nannan Wang 0001, Ruimin Hu, Xinbo Gao 0001
IEEE Trans. Inf. Forensics Secur.3
2024 MMNet: Multi-Collaboration and Multi-Supervision Network for Sequential Deepfake Detection
abstract
Advanced manipulation techniques have provided criminals with opportunities to make social panic or gain illicit profits through the generation of deceptive media, such as forgery face images. In response, various deepfake detection methods have been proposed to assess image authenticity. Sequential deepfake detection, which is an extension of deepfake detection, aims to identify forged facial regions with the correct sequence for recovery. Nonetheless, due to the different combinations of spatial and sequential manipulations, forgery face images exhibit substantial discrepancies that severely impact detection performance. Additionally, the recovery of forged images requires knowledge of the manipulation model to implement inverse transformations, which is difficult to ascertain as relevant techniques are often concealed by attackers. To address these issues, we propose Multi-Collaboration and Multi-Supervision Network (MMNet) that handles various spatial scales and sequential permutations in forgery face images and achieve recovery without requiring knowledge of the corresponding manipulation method. Furthermore, existing evaluation metrics only consider detection accuracy at a single inferring step, without accounting for the matching degree with ground-truth under continuous multiple steps. To overcome this limitation, we propose a novel evaluation metric called Complete Sequence Matching (CSM), which considers the detection accuracy at multiple inferring steps, reflecting the ability to detect integrally forged sequences. Extensive experiments on several typical datasets demonstrate that MMNet achieves state-of-the-art detection performance and independent recovery performance. Code will be available at https://github.com/xarryon/MMNet.
Ruiyang Xia, Decheng Liu, Jie Li 0001, Lin Yuan 0002, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Inf. Forensics Secur.2
2024 Inspector for Face Forgery Detection: Defending Against Adversarial Attacks From Coarse to Fine
abstract
The emergence of face forgery has raised global concerns on social security, thereby facilitating the research on automatic forgery detection. Although current forgery detectors have demonstrated promising performance in determining authenticity, their susceptibility to adversarial perturbations remains insufficiently addressed. Given the nuanced discrepancies between real and fake instances are essential in forgery detection, previous defensive paradigms based on input processing and adversarial training tend to disrupt these discrepancies. For the detectors, the learning difficulty is thus increased, and the natural accuracy is dramatically decreased. To achieve adversarial defense without changing the instances as well as the detectors, a novel defensive paradigm called Inspector is designed specifically for face forgery detectors. Specifically, Inspector defends against adversarial attacks in a coarse-to-fine manner. In the coarse defense stage, adversarial instances with evident perturbations are directly identified and filtered out. Subsequently, in the fine defense stage, the threats from adversarial instances with imperceptible perturbations are further detected and eliminated. Experimental results across different types of face forgery datasets and detectors demonstrate that our method achieves state-of-the-art performances against various types of adversarial perturbations while better preserving natural accuracy. Code is available on https://github.com/xarryon/Inspector.
Ruiyang Xia, Dawei Zhou 0004, Decheng Liu, Jie Li 0001, Lin Yuan 0002, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Image Process.3
2024 Towards Specific Domain Prompt Learning via Improved Text Label Optimization
abstract
Prompt learning has emerged as a thriving parameter-efficient fine-tuning technique for adapting pre-trained vision-language models (VLMs) to various downstream tasks. However, existing prompt learning approaches still exhibit limited capability for adapting foundational VLMs to specific domains that require specialized and expert-level knowledge. Since this kind of specific knowledge is primarily embedded in the pre-defined text labels, we infer that foundational VLMs cannot directly interpret semantic meaningful information from these specific text labels, which causes the above limitation. From this perspective, this paper additionally models text labels with learnable tokens and casts this operation into traditional prompt learning framework. By optimizing label tokens, semantic meaningful text labels are automatically learned for each class. Nevertheless, directly optimizing text label still remains two critical problems, i.e., insufficient optimization and biased optimization. We further address these problems by proposing Modality Interaction Text Label Optimization (MITLOp) and Color-based Consistency Augmentation (CCAug) respectively, thereby effectively improving the quality of the optimized text labels. Extensive experiments indicate that our proposed method achieves significant improvements in VLM adaptation on specific domains.
Liangchen Liu 0001, Nannan Wang 0001, Decheng Liu, Xi Yang 0011, Xinbo Gao 0001, Tongliang Liu
IEEE Trans. Multim.3
2024 Hierarchical Forgery Classifier on Multi-Modality Face Forgery Clues
abstract
Face forgery detection plays an important role in personal privacy and social security. With the development of adversarial generative models, high-quality forgery images become more and more indistinguishable from real to humans. Existing methods always regard as forgery detection task as the common binary or multi-label classification, and ignore exploring diverse multi-modality forgery image types, e.g. visible light spectrum and near-infrared scenarios. In this article, we propose a novelHierarchicalForgeryClassifier forMulti-modalityFaceForgeryDetection(HFC-MFFD), which could effectively learn robust patches-based hybrid domain representation to enhance forgery authentication in multiple modality scenarios. The local hybrid domain representation is designed to explore strong discriminative forgery clues both in the image and frequency domain with the intra-attention mechanism. Furthermore, the specific hierarchical face forgery classifier is designed through the authenticity feedback strategy to integrate diverse discriminative clues. Experimental results on representative multi-modality face forgery datasets demonstrate the superior performance of the proposed HFC-MFFD compared with state-of-the-art algorithms.
Decheng Liu, Zeyang Zheng, Chunlei Peng, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Multim.1
2024 Disguised Heterogeneous Face Generation With Iterative-Adversarial Style Unification
abstract
Heterogeneous face recognition (HFR), which refers to matching face images with different modalities, is essential to public safety. Although HFR has made promising progress in recent years, disguised faces in HFR scenarios still remain a major challenge for the following reasons. First, most existing HFR methods focus on traditional scenarios without disguised accessories, and the performance degrades when dealing directly with disguised faces. Second, there is a need for disguised heterogeneous face datasets, which is essential for developing the related research community. Third, colorful accessories are distinct from heterogeneous face images in terms of their modalities, and their direct combination results in style inconsistency and poor quality. Therefore, we propose a disguised heterogeneous face generation method based on an iterative-adversarial style unification framework. Our approach aims to gradually learn frame textures to detail textures in multiple confrontation iterations, resulting in style unification for disguised accessories and heterogeneous faces. We also construct a disguised heterogeneous face dataset, which contains a disguised NIR-VIS subset and a disguised sketch-photo subset. Moreover, we provide benchmark evaluations conducted on our proposed dataset with face recognition and image quality assessment, demonstrating the superiority of our method over direct addition and two representative disguised face generation techniques.
Chunlei Peng, Zimo Kong, Decheng Liu, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Multim.3
2023 Hiding Visual Information via Obfuscating Adversarial Perturbations
abstract
Growing leakage and misuse of visual information raise security and privacy concerns, which promotes the development of information protection. Existing adversarial perturbations-based methods mainly focus on the de-identification against deep learning models. However, the inherent visual information of the data has not been well protected. In this work, inspired by the Type-I adversarial attack, we propose an Adversarial Visual Information Hiding (AVIH) method to protect the visual privacy of data. Specifically, the method generates obfuscating adversarial perturbations to obscure the visual information of the data. Meanwhile, it maintains the hidden objectives to be correctly predicted by models. In addition, our method does not modify the parameters of the applied model, which makes it flexible for different scenarios. Experimental results on the recognition and classification tasks demonstrate that the proposed method can effectively hide visual information and hardly affect the performances of models. The code is available at https://github.com/suzhigangssz/AVIH.
Zhigang Su, Dawei Zhou 0004, Nannan Wang 0001, Decheng Liu, Zhen Wang 0037, Xinbo Gao 0001
ICCV4
2023 Eliminating Adversarial Noise via Information Discard and Robust Representation Restoration
abstract
Deep neural networks (DNNs) are vulnerable to adversarial noise. Denoising model-based defense is a major protection strategy. However, denoising models may fail and induce negative effects in fully white-box scenarios. In this work, we start from the latent inherent properties of adversarial samples to break the limitations. Unlike solely learning a mapping from adversarial samples to natural samples, we aim to achieve denoising by destroying the spatial characteristics of adversarial noise and preserving the robust features of natural information. Motivated by this, we propose a defense based on information discard and robust representation restoration. Our method utilize complementary masks to disrupt adversarial noise and guided denoising models to restore robust-predictive representations from masked samples. Experimental results show that our method has competitive performance against white-box attacks and effectively reverses the negative effect of denoising models.
Dawei Zhou 0004, Nannan Wang 0001, Decheng Liu, Xinbo Gao 0001, Tongliang Liu
ICML4
2023 Modality-agnostic Augmented Multi-Collaboration Representation for Semi-supervised Heterogenous Face Recognition
abstract
Heterogeneous face recognition (HFR) aims to match input face identity across different image modalities. Due to the existing large modality gap and the limited number of training data, HFR is still a challenging problem in biometrics and draws more and more attention. Existing researchers always extract modality invariant features or generate homogeneous images to decrease the modality gap, lacking abundant labeled data to avoid the overfitting problem. In this paper, we proposed a novel Modality-Agnostic Augmented Multi-Collaboration representation for Heterogeneous Face Recognition (MAMCO-HFR) in a semi-supervised manner. The modality-agnostic augmentation strategy is proposed to generate adversarial perturbations to map unlabeled faces into the modality-agnostic domain. The multi-collaboration feature constraint is designed to mine the inherent relationships between diverse layers for discriminative representation. Experiments on several large-scale heterogeneous face datasets (CASIA NIR-VIS 2.0, LAMP-HQ and Tufts Face dataset) prove the proposed algorithm can achieve superior performance compared with state-of-the-art methods. The source code is available at https://github.com/xiyin11/Semi-HFR.
Decheng Liu, Weizhao Yang, Chunlei Peng, Nannan Wang 0001, Ruimin Hu, Xinbo Gao 0001
ACM Multimedia1
2023 Face photo-sketch synthesis via intra-domain enhancement
Chunlei Peng, Congyu Zhang, Decheng Liu, Nannan Wang 0001, Xinbo Gao 0001
Knowl. Based Syst.3
2023 Spatial-Temporal Frequency Forgery Clue for Video Forgery Detection in VIS and NIR Scenario
abstract
In recent years, with the rapid development of face editing and generation, more and more fake videos are circulating on social media, which has caused extreme public concerns. Existing face forgery detection methods based on frequency domain find that the GAN forged images have obvious grid-like visual artifacts in the frequency spectrum. But for synthesized videos, these methods only confine to a single frame and pay little attention to the most discriminative part and temporal frequency clue among different frames. To take full advantage of the rich information in video sequences, this paper performs video forgery detection on both spatial and temporal frequency domains and proposes a Discrete Cosine Transform-based Forgery Clue Augmentation Network (FCAN-DCT) to achieve a more comprehensive spectrum spatial-temporal feature representation. FCAN-DCT totally consists of a backbone network and two branches: Compact Feature Extraction (CFE) module and Frequency Temporal Attention (FTA) module. We conduct thorough experimental assessments on three visible light (VIS) based datasets (i.e.,, FaceForensics++, Celeb-DF (v2), WildDeepfake), and our self-built video forgery dataset DeepfakeNIR, which is the first video forgery dataset on near-infrared (NIR) modality. The experimental results demonstrate the effectiveness and robustness of our method for detecting forgery videos in both VIS and NIR scenarios.DeepfakeNIR and code are available athttps://github.com/AEP-WYK/DeepfakeNIR.
Chunlei Peng, Decheng Liu, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Circuits Syst. Video Technol.3
2023 FedForgery: Generalized Face Forgery Detection With Residual Federated Learning
abstract
With the continuous development of deep learning in the field of image generation models, a large number of vivid forged faces have been generated and spread on the Internet. These high-authenticity artifacts could grow into a threat to society security. Existing face forgery detection methods directly utilize the obtained public shared or centralized data for training but ignore the personal privacy and security issues when personal data couldn’t be centralizedly shared in real-world scenarios. Additionally, different distributions caused by diverse artifact types would further bring adverse influences on the forgery detection task. To solve the mentioned problems, the paper proposes a novel generalized residual Federated learning for face Forgery detection (FedForgery). The designed variational autoencoder aims to learn robust discriminative residual feature maps to detect forgery faces (with diverse or even unknown artifact types). Furthermore, the general federated learning strategy is introduced to construct distributed detection model trained collaboratively with multiple local decentralized devices, which could further boost the representation generalization. Experiments conducted on publicly available face forgery detection datasets prove the superior performance of the proposed FedForgery. The designed novel generalized face forgery detection protocols and source code would be publicly available at https://github.com/GANG370/FedForgery.
Decheng Liu, Zhan Dang, Chunlei Peng, Yu Zheng 0006, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Inf. Forensics Secur.1
2023 HiFiSketch: High Fidelity Face Photo-Sketch Synthesis and Manipulation
abstract
With the rapid development of generative adversarial networks, face photo-sketch synthesis has achieved promising performance and playing an increasingly important role in law enforcement as well as entertainment. However, most of the existing methods only work under the condition of no interference, and lack of generalization ability in wild scenes. The fidelity of the images generated by the existing methods are insufficient, and the manipulation ability according to text description is unavailable. Directly applying existing text-based image manipulation methods on face photo-sketch scenario may lead to severe distortions due to the cross-domain challenges. Therefore, we propose a novel cross-domain face photo-sketch synthesis framework named HiFiSketch, a network that learns to adjust the weights of generators for high-fidelity synthesis and manipulation. It can realize the translation of images between the photo domain and the sketch domain, and modify results according to the text input in the meanwhile. We further propose a cross-domain loss function, which can effectively preserve facial details during face photo-sketch synthesis. Extensive experiments on four public face sketch datasets show the superiority of our method compared to existing methods. We further present text-based face photo-sketch manipulation and sequential face photo-sketch manipulation for the first time to demonstrate the effectiveness of our method on high fidelity face photo-sketch synthesis and manipulation.
Chunlei Peng, Congyu Zhang, Decheng Liu, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Image Process.3
2022 SketchCLIP: Text-based Attribute Manipulation for Face Sketch Synthesis
abstract
This paper proposes a method of modifying the face sketch with text descriptions. Face sketch is widely used in the criminal field and digital entertainment field. Forensic painters usually draw face sketches based on descriptions provided by witnesses or clients. However, drawing a face sketch often takes lots of time and effort. Existing face sketch synthesis studies have not considered text-based sketch manipulation, and we find that applying text-driven editing methods on natural images directly to face sketches causes severe distortion of generated results. Therefore, this paper proposes a novel text-based attribute manipulation method for face sketch synthesis, named SketchCLIP. Our approach adopts text-driven attribute manipulation by using the powerful Contrastive Language-Image Pre-Training (CLIP) model, which not only conforms to the current drawing process of face sketches but also does not require tedious manual operations and allows for more diverse modifications. Besides, we design an intra-modality fine-tuning module to eliminate distortion and improve the quality of the modified face sketch. Through extensive comparison experiments on public face sketch datasets, our method is demonstrated to be very excellent in the effectiveness of the face sketch processing and the quality of modified results.
Mengdi Dong, Chunlei Peng, Decheng Liu, Yu Zheng 0006, Nannan Wang 0001, Xinbo Gao 0001
IJCB3
2022 A robust adaptive blind color image watermarking for resisting geometric attacks
Qingtang Su, Decheng Liu, Yehan Sun
Inf. Sci.2
2022 ForgeryNIR: Deep Face Forgery and Detection in Near-Infrared Scenario
abstract
Deep face forgery and detection is an emerging topic due to the development of GANs. Face forgery detection relies greatly on existing databases for evaluation and adequate training examples for data-hungry machine learning algorithms. However, considering the wide application of face recognition in near-infrared scenarios, there is no publicly available face forgery database that includes near-infrared modality currently. In this paper, we present an attempt at constructing a large-scale dataset for face forgery detection in the near-infrared modality and propose a new forgery detection method based on knowledge distillation named cross-modality knowledge distillation aiming to use a teacher model which is pre-trained on the visible light-based (VIS) big data to guide the student model with a small amount of near-infrared (NIR) data. The proposed near-infrared face forgery dataset, named ForgeryNIR, contains a total of over 50,000 real and fake identities. A number of perturbations are applied to help simulate real-world scenarios. All source images in ForgeryNIR are collected from CASIA NIR-VIS 2.0, and fake images are generated via multiple GAN techniques. The proposed dataset fills the gap of face forgery detection research in the near-infrared modality. A comprehensive study on six representative detection baselines is conducted to evaluate the performance of face forgery detection algorithms in the NIR domain. We further construct a hard testing set, named ForgeryNIR+, which contains forged images that have bypassed existing face forgery detection methods. The proposed datasets will be publicly available and aim to help boost further research on face forgery detection, as well as NIR face detection and recognition.
Chunlei Peng, Decheng Liu, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Inf. Forensics Secur.3
2022 Edge Aware Domain Transformation for Face Sketch Synthesis
abstract
With the development of generative adversarial networks (GAN), the field of face sketch synthesis has received extensive attention. Face sketch synthesis (FSS) has promising prospects in the fields of entertainment and law enforcement, where it plays an increasingly important role. We propose a novel generative adversarial network for synthesizing sketches with similar shapes and rich details to photos. This problem is challenging because it involves the transition between the sketch domain and the photo domain. Many methods have been used for face sketch synthesis in recent years, but existing methods cannot fully exploit the semantic information between different domains. To this end, we use a cross-domain face sketch synthesis framework based on edge-preserving filters to make the boundaries of different semantics in semantic layouts have a smooth transition. We further propose a new spatially adaptive denormalization module named edge-aware enhancement Spatially Adaptive DEnormalization (eaeSPADE), which can make full use of the semantic information in the semantic layout of faces and improve the details of the synthesized face images. Extensive experiments demonstrate that our method outperforms existing face sketch synthesis methods.
Congyu Zhang, Decheng Liu, Chunlei Peng, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Inf. Forensics Secur.2
2022 Feature Ensemble Net: A Deep Framework for Detecting Incipient Faults in Dynamical Processes
abstract
How to detect incipient faults has been an important problem in the field of fault detection. Although many types of machine and deep learning methods have been proposed, their performance is not as good as expected. In this article, a novel feature ensemble net (FENet) was developed, particularly for faults 3, 9, and 15 in the Tennessee Eastman process (TEP), which are notoriously difficult to detect. For the input feature layer, features extracted by the basic detectors are integrated to expand the detection ability of FENet. For the hidden feature transformer layers, with sliding-window patches and principal component analysis (PCA), the previous feature matrix is transformed. The sliding-window patches can be used to generate singular values, whereas the patches in the well-known convolution technique can only be vectorized, primarily for performing PCA in PCA-based networks. This enhances the sensitivity of the FENet to incipient faults. For the output feature layer, all feature matrices in the last hidden layer are completely stacked into a large feature matrix. The sliding technique is performed at the decision layer, and a detection index is designed with normalized singular values. The superiority of FENet can be completely verified by a continuous stirred tank heater and TEP. As compared with deep PCA, PCA-based monitoring network, and typical ensemble strategies, such as averaging, voting, stacking, and Bayesian inference, FENet can effectively detect Faults 3, 9, and 15 in TEP.
Decheng Liu, Min Wang 0041, Mao-Yin Chen
IEEE Trans. Ind. Informatics1
2022 Heterogeneous Face Interpretable Disentangled Representation for Joint Face Recognition and Synthesis
abstract
Heterogeneous faces are acquired with different sensors, which are closer to real-world scenarios and play an important role in the biometric security field. However, heterogeneous face analysis is still a challenging problem due to the large discrepancy between different modalities. Recent works either focus on designing a novel loss function or network architecture to directly extract modality-invariant features or synthesizing the same modality faces initially to decrease the modality gap. Yet, the former always lacks explicit interpretability, and the latter strategy inherently brings in synthesis bias. In this article, we explore to learn the plain interpretable representation for complex heterogeneous faces and simultaneously perform face recognition and synthesis tasks. We propose the heterogeneous face interpretable disentangled representation (HFIDR) that could explicitly interpret dimensions of face representation rather than simple mapping. Benefited from the interpretable structure, we further could extract latent identity information for cross-modality recognition and convert the modality factor to synthesize cross-modality faces. Moreover, we propose a multimodality heterogeneous face interpretable disentangled representation (M-HFIDR) to extend the basic approach suitable for the multimodality face recognition and synthesis. To evaluate the ability of generalization, we construct a novel large-scale face sketch data set. Experimental results on multiple heterogeneous face databases demonstrate the effectiveness of the proposed method.
Decheng Liu, Xinbo Gao 0001, Chunlei Peng, Nannan Wang 0001, Jie Li 0001
IEEE Trans. Neural Networks Learn. Syst.1
2021 A fusion-domain color image watermarking based on Haar transform and image correction
Decheng Liu, Qingtang Su, Zihan Yuan
Expert Syst. Appl.1
2021 Iterative local re-ranking with attribute guided synthesis for face sketch recognition
Decheng Liu, Xinbo Gao 0001, Nannan Wang 0001, Chunlei Peng, Jie Li 0001
Pattern Recognit.1
2021 A blind color digital image watermarking method based on image correction and eigenvalue decomposition
Decheng Liu, Qingtang Su, Zihan Yuan
Signal Process. Image Commun.1
2021 Principal Component Analysis-Based Ensemble Detector for Incipient Faults in Dynamic Processes
abstract
The significant advancement in data-driven fault detection has been made, but incipient faults such as faults 3, 9, and 15 in Tennessee Eastern process (TEP) still remain difficult for the current approaches. In this article, a powerful principal component analysis (PCA)-based ensemble detector (PCAED) is developed for detecting incipient faults. To begin with, multiple PCA-based detectors are designed based on bootstrap sampling in the training dataset. It can generate two matrices according to principal component and residual subspaces. Then, two sensitive detection indices are developed using maximal singular values of one-step sliding windows along the rows of the above two matrices. With this kind of detection index, PCAED can effectively detect incipient faults, specially faults 3, 9, and 15 in TEP, which cannot be detected by an individual PCA detector. Simulations of TEP and a practical coal pulverizing system fully verify the effectiveness of PCAED. Faults can be successfully detected at the incipient stage, which is very helpful to avoid possible economic or human loss.
Decheng Liu, Jun Shang, Mao-Yin Chen
IEEE Trans. Ind. Informatics1
2021 A color watermarking scheme in frequency domain based on quaternary coding
Decheng Liu, Qingtang Su, Zihan Yuan
Vis. Comput.1
2021 A blind image watermarking scheme combining spatial domain and frequency domain
Zihan Yuan, Qingtang Su, Decheng Liu
Vis. Comput.3
2020 Fast and robust image watermarking method in the spatial domain
abstract
To solve the copyright protection problem of a colour image, a new blind colour image watermarking method combining a discrete cosine transform (DCT) in the spatial domain is presented in this study. The advantages of the spatial‐domain watermarking algorithm and frequency‐domain one are made full use in this scheme. Based on the different quantisation steps in red, green, and blue three‐layer images, the processes of watermark embedding and blind extraction are completed in the spatial domain without a real DCT domain. The scheme is realised by using the unique features of the direct current (DC) coefficient and the relativity of DC coefficients between adjacent pixel blocks. This scheme can effectively solve the problems of the large‐capacity colour image watermarking algorithm, such as long‐running time and weak robustness. Comparing with other advanced watermarking algorithms, the presented scheme has better invisibility, stronger robustness, and higher real‐time performance.
Zihan Yuan, Qingtang Su, Decheng Liu
IET Image Process.3
2020 A blind color image watermarking scheme with variable steps based on Schur decomposition
Decheng Liu, Zihan Yuan, Qingtang Su
Multim. Tools Appl.1
2020 A combined domain watermarking algorithm of color image
Qingtang Su, Huanying Wang, Decheng Liu, Zihan Yuan
Multim. Tools Appl.3
2020 DCT-based color digital image blind watermarking method with variable steps
Zihan Yuan, Decheng Liu, Huanying Wang, Qingtang Su
Multim. Tools Appl.2
2020 Coupled Attribute Learning for Heterogeneous Face Recognition
abstract
Heterogeneous face recognition (HFR) is a challenging problem in face recognition and subject to large textural and spatial structure differences of face images. Different from conventional face recognition in homogeneous environments, there exist many face images taken from different sources (including different sensors or different mechanisms) in reality. In addition, limited training samples of cross-modality pairs make HFR more challenging due to the complex generation procedure of these images. Despite the great progress that has been achieved in recent years, existing works mainly focus on HFR from only cross-modality image matching. However, it is more practical to obtain both facial images and semantic descriptions about facial attributes in real-world situations, in which the semantic description clues are nearly always obtained during the process of image generation. Motivated by human cognitive mechanisms, we naturally utilize the explicit invariant semantic description, i.e., face attributes, to help address the gap among face images of different modalities. Existing facial attributes-related face recognition methods primarily regard attributes as the high-level features used to enhance recognition performance, ignoring the inherent relationship between face attributes and identities. In this article, we propose novel coupled attribute learning for the HFR (CAL-HFR) method without labeling the attributes manually. Deep convolutional networks are employed to directly map face images in heterogeneous scenarios to a compact common space where distances are taken as dissimilarities of pairs. Coupled attribute guided triplet loss (CAGTL) is designed to train an end-to-end HFR network that can effectively eliminate defects of incorrectly estimated attributes. Extensive experiments on multiple heterogeneous scenarios demonstrate that the proposed method achieves superior performance compared with that of state-of-the-art methods. Furthermore, we make publicly available our generated pairwise annotated heterogeneous facial attribute database for evaluation and promoting related research.
Decheng Liu, Xinbo Gao 0001, Nannan Wang 0001, Jie Li 0001, Chunlei Peng
IEEE Trans. Neural Networks Learn. Syst.1
2019 A new watermarking scheme for colour image using QR decomposition and ternary coding
Qingtang Su, Decheng Liu, Zihan Yuan, Hongye Ning
Multim. Tools Appl.3
2018 Deep Attribute Guided Representation for Heterogeneous Face Recognition
abstract
Heterogeneous face recognition (HFR) is a challenging problem in face recognition, subject to large texture and spatial structure differences of face images. Different from conventional face recognition in homogeneous environments, there exist many face images taken from different sources (including different sensors or different mechanisms) in reality. Motivated by human cognitive mechanism, we naturally utilize the explicit invariant semantic information (face attributes) to help address the gap of different modalities. Existing related face recognition methods mostly regard attributes as the high level feature integrated with other engineering features enhancing recognition performance, ignoring the inherent relationship between face attributes and identities. In this paper, we propose a novel deep attribute guided representation based heterogeneous face recognition method (DAG-HFR) without labeling attributes manually. Deep convolutional networks are employed to directly map face images in heterogeneous scenarios to a compact common space where distances mean similarities of pairs. An attribute guided triplet loss (AGTL) is designed to train an end-to-end HFR network which could effectively eliminate defects of incorrectly detected attributes. Extensive experiments on multiple heterogeneous scenarios (composite sketches, resident ID cards) demonstrate that the proposed method achieves superior performances compared with state-of-the-art methods.
Decheng Liu, Nannan Wang 0001, Chunlei Peng, Jie Li 0001, Xinbo Gao 0001
IJCAI1
2018 Composite components-based face sketch recognition
Decheng Liu, Jie Li 0001, Nannan Wang 0001, Chunlei Peng, Xinbo Gao 0001
Neurocomputing1