VLDB 2026 Research / reviewers in the wild / expert
Chunlei Peng
dblp:148/8269
· DBLP profile ↗
76ranked-venue papers
26as first author
63since 2021 · last 2026
0000-0003-3448-2514ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 30 · 8 first-author · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 29 · 12 first-author · 24 since 2021Security and privacy · 19 · 5 first-author · 18 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Computer networks · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ResProto-FD: Visual-Language Residual Prototype Sets for Generalized Face Forgery DetectionabstractWith the rapid development of generative models, such as generative adversarial networks and diffusion models, the task of face forgery detection has emerged, aiming to identify forged faces in real-world scenarios. A key challenge for current face forgery detection models is improving generalization to unknown forgeries. To address this, we propose ResProto-FD, a framework that constructs residual prototype sets to capture diverse forgery cues and discriminative differences from real faces. Our novel perspective collects prototypes from the most informative residual features generated during training, enabling better representation of various forgery traces and real-vs-fake distinctions. First, we introduce a Visual-Language Residual Learning (VLRL) module based on the CLIP model. This module constructs residual features between image and text embeddings to capture inconsistencies between visual features and associated textual semantics. In doing so, it guides the model to attend to subtle visual forgery clues and enhances the discriminative power of image representations. Furthermore, we design a Gradient-aware Residual Prototypes (GRP) mechanism— a dynamic collection strategy that selectively stores uncertain residual features based on gradient signals to build the prototype sets. This enhances the model’s ability to generalize to unknown forgery types. Extensive experiments across various datasets and forgery methods demonstrate that ResProto-FD significantly improves generalization performance and consistently outperforms state-of-the-art methods. Jiuyao Jing, Yu Zheng 0006, Chunlei Peng |
AAAI | 3 |
| 2026 | MGFFD-VLM: Multi-Granularity Prompt Learning for Face Forgery Detection with VLMabstractThe proliferation of face forgery content on the Web poses a severe threat to online trust, social media security, and the credibility of digital information. Existing detection approaches often fail to generalize across diverse forgery types and unseen scenarios commonly encountered in web-scale applications. Recent studies have utilized visual large language models (VLMs) to answer not only ''Is this face a forgery?'' but also ''Why is the face a forgery?'' These studies introduced forgery-related attributes, such as forgery location and type, to construct deepfake VQA datasets and train VLMs, achieving high accuracy while providing human-understandable explanatory text descriptions. However, these methods still have limitations. For example, they do not fully leverage face quality-related attributes, which are often abnormal in forged faces, and they lack effective training strategies for forgery-aware VLMs. In this paper, we extend the VQA dataset to create DD-VQA+, which features a richer set of attributes and a more diverse range of samples. Furthermore, we introduce a novel forgery detection framework, MGFFD-VLM, which integrates an Attribute-Driven Hybrid LoRA Strategy to enhance the capabilities of Visual Large Language Models (VLMs). Additionally, our framework incorporates Multi-Granularity Prompt Learning and a Forgery-Aware Training Strategy. By transforming classification and forgery segmentation results into prompts, our method not only improves forgery classification but also enhances interpretability. To further boost detection performance, we design multiple forgery-related auxiliary losses. Experimental results demonstrate that our approach surpasses existing methods in both text-based forgery judgment and analysis, achieving superior accuracy. Decheng Liu, Chunlei Peng |
WWW | 4 |
| 2026 | Symmetrical bidirectional knowledge alignment for zero-shot sketch-based image retrievalabstractThis paper studies the problem of zero-shot sketch-based image retrieval (ZS-SBIR), which aims to use sketches from unseen categories as queries to match the images of the same category. Due to the large cross-modality discrepancy, ZS-SBIR is still a challenging task and mimics realistic zero-shot scenarios. The key is to leverage transferable knowledge from the pre-trained model to improve generalizability. Existing researchers often utilize the simple fine-tuning training strategy or knowledge distillation from a teacher model with fixed parameters, lacking efficient bidirectional knowledge alignment between student and teacher models simultaneously for better generalization. In this paper, we propose a novel Symmetrical Bidirectional Knowledge Alignment for zero-shot sketch-based image retrieval (SBKA). The symmetrical bidirectional knowledge alignment learning framework is designed to effectively learn mutual rich discriminative information between teacher and student models to achieve the goal of knowledge alignment. Instead of the former one-to-one cross-modality matching in the testing stage, a one-to-many cluster cross-modality matching method is proposed to leverage the inherent relationship of intra-class images to reduce the adverse effects of the existing modality gap. Experiments on several representative ZS-SBIR datasets (Sketchy Ext dataset, TU-Berlin Ext dataset and QuickDraw Ext dataset) prove the proposed algorithm can achieve superior performance compared with state-of-the-art methods. The source code is publicly available at https://github.com/zermatt-luo/SBKA. Decheng Liu, Chunlei Peng, Nannan Wang 0001, Ruimin Hu, Xinbo Gao 0001 |
Neural Networks | 3 |
| 2026 | TransFA: Transformer-based representation for face attribute evaluation
Decheng Liu, Chunlei Peng, Nannan Wang 0001, Jie Li 0001, Xinbo Gao 0001 |
Pattern Recognit. | 3 |
| 2026 | IDCFace: Identity Consistent Face anonymization for secure recognition
Ruiying Lu, Shuang Wan, Zimin Miao, Nannan Wang 0001, Chunlei Peng |
Pattern Recognit. | 5 |
| 2026 | Cross-color space feature fusion for anomaly detection-based face anti-spoofing
Yu Zheng 0006, Chunlei Peng |
Pattern Recognit. Lett. | 4 |
| 2026 | Adaptive Consensus Multi-Teacher Distillation for Generalizable Face Forgery DetectionabstractRecent advancements in deep learning have significantly lowered the cost of generating and processing facial images, but this progress has also presented a growing challenge for face forgery detection. Existing detection methods, however, often struggle to generalize effectively to forged samples generated by previously unseen forgery techniques. Additionally, transferring knowledge through knowledge distillation to enhance generalization remains a persistent challenge. To address these issues, this paper proposes Adaptive Consensus Multi-teacher Knowledge Distillation (ACMD), a novel framework aimed at improving the generalization capabilities of face forgery detection models. ACMD leverages multiple teacher models with varying network architectures and introduces a consensus mechanism to resolve knowledge conflicts arising from the diversity of teacher models. It further adapts the assignment of sample-specific distillation weights, based on each teacher model’s prediction confidence. By combining this with feature-based knowledge distillation, the student model gains a deeper understanding of the teacher’s knowledge. Extensive experiments on the DeepfakeBench benchmark demonstrate that ACMD not only outperforms a single teacher model but also achieves state-of-the-art performance in dataset evaluations. Specifically, in a cross-dataset evaluation from FaceForensics++ to CDFv2 and DFDC, ACMD achieves frame-level AUCs of 84.22% and 74.43%, respectively, surpassing all baseline models. Ablation studies further validate the effectiveness of each component and highlight their complementary contributions to the overall detection performance. Jiuyao Jing, Chunlei Peng, Decheng Liu, Nannan Wang 0001, Xinbo Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | DeepFidelity: Perceptual Forgery Fidelity Assessment for Deepfake DetectionabstractDeepfake detection refers to detecting artificially generated or edited faces in images or videos, which plays an essential role in visual information security. Despite promising progress in recent years, Deepfake detection remains a challenging problem due to the complexity and variability of face forgery techniques. Existing Deepfake detection methods are often devoted to extracting features by designing sophisticated networks but ignore the influence of perceptual quality of faces. Considering the complexity of the quality distribution of real and fake faces, we propose a deepfake detection framework called DeepFidelity, which mines the perceptual forgery fidelity of face images and introduces a quality-aware scoring mechanism to distinguish real and fake faces of different image qualities. Specifically, we improve the model’s ability to identify complex samples by mapping real and fake face data of different qualities to different scores to distinguish them in a more detailed way. In addition, we propose a network structure called Symmetric Spatial Attention Augmentation based vision Transformer (SSAAFormer), which uses the symmetry of face images to promote the network to model the geographic long-distance relationship at the shallow level and augment local features. Extensive experiments on multiple benchmark datasets demonstrate the superiority of the proposed method over state-of-the-art methods. The code is available athttps://github.com/shimmer-ghq/DeepFidelity. Chunlei Peng, Huiqing Guo, Decheng Liu, Nannan Wang 0001, Ruimin Hu, Xinbo Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2026 | Component-Specific Prompt Tuning for Deepfake DetectionabstractWith the development of deep learning technology, the facial images generated by deepfake technology have reached a level of authenticity that is difficult to distinguish, posing a serious threat to personal privacy and data security. Therefore, it is of great significance to develop efficient and reliable deepfake detection technology. In recent years, Visual Language Models (VLM) have been applied to deepfake detection tasks due to their powerful multimodal understanding capabilities. However, the existing VLM have not been specifically optimized for deepfake detection tasks. When directly applied to this task, there are problems such as insufficient model accuracy and insufficient feature extraction, especially when dealing with complex forgery scenes. In response to these challenges, this paper proposes an innovative deepfake face detection method based on VLM and component-specific prompt tuning. We transform the deepfake detection task into a Visual Question Answering (VQA) task, making full use of the multimodal understanding capabilities of VLM and the flexibility of prompt tuning technology. This method uses a local prompt strategy to customize specific prompt questions for key facial components such as eyes, nose, and mouth, guiding the model to focus on the local features of these areas, thereby accurately capturing forgery traces. In addition, we introduced a feature extraction module Q-Former based on instructions, which can flexibly adjust the focus area of visual features according to prompts, significantly improving the model’s perception of locally forged features. By fusing these local features extracted by Q-Former and combining them with the language model to judge the authenticity of the overall face image, we can finally generate accurate prediction results. A large number of experimental results show that our method is significantly better than existing technologies in terms of detection accuracy and robustness. Yinyin Chen, Huiqing Guo, Chunlei Peng, Nannan Wang 0001, Xinbo Gao 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2026 | Contextual Masking Distillation for Network Traffic Anomaly DetectionabstractNetwork traffic anomaly detection is critical for cybersecurity but faces challenges in accurately identifying malicious activities. Recent zero-positive approaches, which use only normal training data under the reconstruction paradigm, have shown progress. However, encrypted network traffic obscures normal–anomalous distinctions, causing confused modeling. In addition, the “identical shortcut” problem, where models reconstruct any input with similar fidelity, produces suboptimal representations and indistinguishable detection. To address these limitations, this paper introduces ConMD, a novel Contextual Masking Knowledge Distillation framework. ConMD features distillation paradigm for discriminative representations and then pursues two objectives: effective contextual information modeling and a comprehensive anomaly metric. Specifically, we introduce context-aware local-global attention mechanisms for the student network's backbone, which capture both intra-packet and inter-packet dependencies. Additionally, a context-enhanced masking training strategy is designed to facilitate contextual interactions in normal flows. Given the structural characteristics of network traffic, we also present a new anomaly scoring with multi-view awareness, which perceive comprehensive traffic patterns. ConMD combines insights from both packet- and flow-level views to highlight deviations in anomalous network flows, thereby improving detection accuracy. Extensive experiments on three real-world datasets validate the effectiveness of ConMD, yielding consistent improvements over state-of-the-art baselines, achieving up to 2.8% and 5.1% AUC gains on the DataCon2020 and CIC-IDS2017 datasets, respectively. Our model code will be released at https://github.com/ikun0124/ConMD. Xinglin Lian, Yu Zheng 0006, Fan Zhou 0002, Chunlei Peng, Xinbo Gao 0001 |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2025 | Thinking Racial Bias in Fair Forgery Detection: Models, Datasets and EvaluationsabstractDue to the successful development of deep image generation technology, forgery detection plays a more important role in social and economic security. Racial bias has not been explored thoroughly in the deep forgery detection field. In the paper, we first contribute a dedicated dataset called the Fair Forgery Detection (FairFD) dataset, where we prove the racial bias of public state-of-the-art (SOTA) methods. Different from existing forgery detection datasets, the self-constructed FairFD dataset contains a balanced racial ratio and diverse forgery generation images with the largest-scale subjects. Additionally, we identify the problems with naive fairness metrics when benchmarking forgery detection models. To comprehensively evaluate fairness, we design novel metrics including Approach Averaged Metric and Utility Regularized Metric, which can avoid deceptive results. We also present an effective and robust post-processing technique, Bias Pruning with Fair Activations (BPFA), which improves fairness without requiring retraining or weight updates. Extensive experiments conducted with 12 representative forgery detection models demonstrate the value of the proposed dataset and the reasonability of the designed fairness metrics. By applying the BPFA to the existing fairest detector, we achieve a new SOTA. Furthermore, we conduct more in-depth analyses to offer more insights to inspire researchers in the community. Decheng Liu, Zongqi Wang, Chunlei Peng, Nannan Wang 0001, Ruimin Hu, Xinbo Gao 0001 |
AAAI | 3 |
| 2025 | Cross-Domain Matrix Compression Adaptation for Zero-Shot Sketch-Based Image RetrievalabstractThis paper explores an efficient parameter fine-tuning strategy for zero-shot sketch-based image retrieval (ZS-SBIR). We highlight a key finding: through efficient parameter fine-tuning, competitive retrieval accuracy can be achieved using only 12% of the full parameters. This superior performance is attributed to low-rank matrix compression in the redundant parameter space, which extracts more discriminative effective feature dimensions. Specifically, to fully leverage the potential of pretrained models, we propose a cross-domain matrix compression adaptation method. For lower-level features, we apply a general low-rank decomposition to extract shared basic shapes or contours information across modalities. To mitigate overfitting to local similarities, we propose a domain-specific matrix compression module that guides the model in learning high-level abstractions essential for sketch retrieval. Our method is simple and effective, balancing both general semantic information and feature variations across domains within the same category. Experimental results on the ZS-SBIR benchmark dataset show that our method not only outperforms existing state-of-the-art methods, but also requires significantly fewer training parameters. Decheng Liu, Yu Zheng 0006, Chunlei Peng |
IJCNN | 4 |
| 2025 | Face Anti-spoofing based on Contour-constrained Anomaly DetectionabstractFace recognition systems are increasingly susceptible to spoofing attacks, which pose serious security risks in biometric authentication. Traditional face anti-spoofing is typically framed as a binary classification problem, but the diversity and evolving nature of spoofing techniques hinder generalization. To address this, we reformulate face anti-spoofing as an anomaly detection task, training only on normal data. Existing reconstruction-based anomaly detection methods often rely on reconstruction errors to identify anomalies, but they tend to produce blurry images that lack high-frequency details such as edges and textures. To overcome this limitation, we propose a contour-guided face anti-spoofing approach that enhances reconstruction by preserving fine-grained details. Extensive experiments demonstrate that our method outperforms state-of-the-art techniques in face anti-spoofing. Chunlei Peng, Yu Zheng 0006 |
ICMR | 3 |
| 2025 | Metal Surface Defect Detection based on Variable Mask Ratio Multi-scale ReconstructionabstractMetal surface defect detection plays a critical role in ensuring product quality in industrial manufacturing. Traditional detection methods often suffer from incomplete feature extraction and limited defective samples, hindering the performance of supervised learning. This paper explores the potential of Vision Transformer (ViT) for metal surface defect detection by addressing domain shift issues and improving training speed. We propose a semi-supervised anomaly detection method based on a Transformer network, using a variable masking ratio to integrate generative modeling with representation learning. By fusing features from a convolutional-Transformer encoder, we utilize a feature pyramid network and a frozen pre-trained hierarchical encoder to generate multi-scale features. By employing these technologies, the proposed method captures fine-grained anomalies effectively, providing a robust solution for detecting diverse defect types. The experimental results on benchmark datasets demonstrate that the proposed method outperforms state-of-the-art defect detection techniques. Yu Zheng 0006, Chunlei Peng |
ICMR | 4 |
| 2025 | Imperceptible Face Forgery Attack via Adversarial Semantic Mask
Qixuan Su, Decheng Liu, Chunlei Peng |
PRCV (7) | 4 |
| 2025 | Fooling human detectors via robust and visually natural adversarial patches
Dawei Zhou 0004, Hongbin Qu, Nannan Wang 0001, Chunlei Peng, Zhuoqi Ma, Xi Yang 0011, Xinbo Gao 0001 |
Neurocomputing | 4 |
| 2025 | Revisiting face forgery detection towards generalization
Chunlei Peng, Decheng Liu, Huiqing Guo, Nannan Wang 0001, Xinbo Gao 0001 |
Neural Networks | 1 |
| 2025 | FairForensics: mitigating attribute bias in deepfake detection by integrating texture and attribute features
Chunlei Peng, Yinyin Chen, Decheng Liu, Nannan Wang 0001, Xinbo Gao 0001 |
Neural Networks | 1 |
| 2025 | Semi-supervised anomaly traffic detection via multi-frequency reconstruction
Xinglin Lian, Yu Zheng 0006, Zhangxuan Dang, Chunlei Peng, Xinbo Gao 0001 |
Pattern Recognit. | 4 |
| 2025 | Devil in Shadow: Attacking NIR-VIS Heterogeneous Face Recognition via Adversarial ShadowabstractNear infrared-visible (NIR-VIS) heterogeneous face recognition aims to match face identities in cross-modality settings, which has achieved significant development recently. The work on adversarial attack and security issues of the heterogeneous face recognition task is still lacking. Existing adversarial face generation methods can’t deploy directly because of the inevitable large modality discrepancy. Besides, the ideal adversarial attacking generated images should maintain both high capabilities and low detectability. Considering the properties of near-infrared face images, our basic idea is to construct adversarial shadows for good stealthiness and high attack capability. In this paper, we propose a novel face adversarial shadow generation framework for NIR-VIS heterogeneous face recognition, which can synthesize fine-crafted lighting conditions containing strong identity attacking ability. Specifically, we design the variance consistency-based symmetric face attacking loss to improve the attacking generalization and the synthesized image quality. Extensive qualitative and quantitative experiments on the public large-scale NIR-VIS heterogeneous face dataset prove the proposed method achieves superior performance compared with the state-of-the-art methods. The source code is publicly available athttps://github.com/GEaMU/Devil-in-Shadow. Decheng Liu, Rong Sheng, Chunlei Peng, Nannan Wang 0001, Ruimin Hu, Xinbo Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Masked Text Adversarial Training for Cloth-Changing Person Re-Identification
Chengrui Hao, Chunlei Peng, Decheng Liu, Nannan Wang 0001, Ruimin Hu, Xinbo Gao 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2025 | Attention Consistency Refined Masked Frequency Forgery Representation for Generalizing Face Forgery DetectionabstractDue to the successful development of deep image generation technology, visual data forgery detection would play a more important role in social and economic security. Existing forgery detection methods suffer from unsatisfactory generalization ability to determine the authenticity in the unseen domain. In this paper, we propose a novel Attention Consistency Refined masked frequency forgery representation model toward a generalizing face forgery detection algorithm (ACMF). Most forgery technologies always bring in high-frequency aware cues, which make it easy to distinguish source authenticity but difficult to generalize to unseen artifact types. The masked frequency forgery representation module is designed to explore robust forgery cues by randomly discarding high-frequency information. In addition, we find that the forgery saliency map inconsistency through the detection network could affect the generalizability. Thus, the forgery attention consistency is introduced to force detectors to focus on similar attention regions for better generalization ability. Experiment results on several public face forgery datasets (FaceForensic++, DFD, Celeb-DF, WDF and DFDC datasets) demonstrate the superior performance of the proposed method compared with the state-of-the-art methods. The source code and models are publicly available athttps://github.com/chenboluo/ACMF. Decheng Liu, Chunlei Peng, Nannan Wang 0001, Ruimin Hu, Xinbo Gao 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2025 | Improving Adversarial Robustness via Decoupled Visual Representation MaskingabstractDeep neural networks are proven to be vulnerable to finely designed adversarial examples, and adversarial defense algorithms draw more and more attention nowadays. Pre-processing based defense is a major strategy, as well as learning robust feature representation, has been proven an effective way to boost generalization. However, existing defense works lack considering different depth-level visual features in the training process. In this paper, we first highlight two novel properties of robust features from the feature distribution perspective: 1) Diversity (robust features within the same class should maintain appropriate variety). 2) Discriminability (robust features from different classes should be sufficiently separated). We find that state-of-the-art defense methods aim to address both of these mentioned issues well. It motivates us to increase intra-class variance and decrease inter-class discrepancy simultaneously in adversarial training. Specifically, we propose a simple but effective defense based on decoupled visual representation masking. The designed Decoupled Visual Feature Masking (DFM) block can adaptively disentangle visual discriminative features and non-visual features with diverse mask strategies, while the suitable discarding information can disrupt adversarial noise to improve robustness. Our work provides a generic and easy-to-plugin block unit for any former adversarial training algorithm to achieve better protection integrally. Extensive experimental results prove that the proposed method can achieve superior performance compared with state-of-the-art defense approaches. The code is publicly available at https://github.com/chenboluo/Adversarial-defense. Decheng Liu, Chunlei Peng, Nannan Wang 0001, Ruimin Hu, Xinbo Gao 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2025 | Toward Fair Adversarial Defense via Class Encourage-Suppress Robust LearningabstractDeep Neural Networks with natural training are quite vulnerable to adversarial attacks, so it’s necessary to defend these attacks with effective defense methods like adversarial training. However, while defending against adversarial attacks, adversarially trained models’ robustness between classes shows severe unfairness. To mitigate the disparity, a lot of methods have been proposed, while many of them sacrifice the overall accuracy to leverage the worst-class accuracy. Inspired by previous works, we propose a new fairness methodology, and we name it Class Encourage-suppress Robust Learning (CRL). Based on the overall accuracy and the class-wise accuracies of the dataset, we introduce a new measurement named Class Diversity Ratio to adjust the weights of different classes in the loss function. Additionally, we propose a new learning strategy called the Competitor Encourage-suppress Strategy to mitigate the disparity between diverse classes, which is simple but effective. Experimental results on representative datasets show that our method outperforms state-of-the-art (SOTA) methods. Decheng Liu, Chunlei Peng, Nannan Wang 0001, Ruimin Hu, Xinbo Gao 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2025 | Posture-Aware Robust Person Re-Identification via Optimal Transport Calibration
Ruiying Lu, Yalin Sun, Chunlei Peng, Yu Zheng 0006 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2025 | Semantic Token Transformer for Face Forgery DetectionabstractIn the era of digital media, the proliferation of forged images and videos poses a significant threat to societal stability. With the rapid advancement of deep learning, the generation of realistic fake images has become increasingly simple, presenting unprecedented challenges in discerning the authenticity of images. While some existing methods have shown promising results in forgery detection, they often underutilize facial semantic information. To address this issue, this paper introduces the Semantic Token Transformer for Face Forgery Detection. By incorporating facial semantic information with a transformer network, the input tokens of the transformer are transformed into tokens of varying shapes and sizes based on their importance, thereby enhancing the accuracy of the detector. To achieve this objective, we first employ an image processing stage to manipulate the image based on facial semantic information. Subsequently, we introduce a scoring network, guided by prior knowledge, which adaptively categorizes tokens into different clusters based on their importance and relevance to the results of the preprocessing stage. Finally, we merge the tokens within the clusters using an attention mechanism and input them into the detector for forgery detection. Through experiments conducted on multiple datasets and cross-dataset evaluations, we demonstrate that our approach outperforms state-of-the-art detection methods. Chunlei Peng, Xiaoyi Luo, Decheng Liu, Nannan Wang 0001, Ruimin Hu, Xinbo Gao 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2025 | Within 3DMM Space: Exploring Inherent 3D Artifact for Video Forgery DetectionabstractRecently, the breathtaking development and potential misuse of deepfake technology has raised numerous privacy and security concerns, triggering widespread apprehension. Existing deepfake detection methods focus on the analysis of local regions for faces, such as mouth movement, eye blinking frequency, etc., which, however, are limited in their ability to capture the global inconsistencies present in forged faces. Some researchers attempt to seize 3D artifacts related to facial global information, but typically treat the 3D information as mere input, lacking the in-depth analysis. To address these shortcomings and mine the inherent and delicate 3D artifacts in the forged faces, this paper innovatively proposes the 3D Artifact Detector (3DAD) method, which leverages the spatio-temporal inconsistency on the 3D semantic space in the forgery videos to uncover the deepfake clues. Specifically, we employ 3D Analysis Unit (3DAU) to pre-train the face reconstruction task within 3D Morphable Model (3DMM) space, thereby obtaining the high-level inherent 3d representation. Concurrently, for the multi-levels of information in the face, we utilize the Texture Perception Unit (TPU) to extract the texture information in the low-level semantic space of the images. Ultimately we feed the two distinct modalities into the spatiotemporal fusion model for final detection. Through extensive intra- and cross-dataset experiments on publicly available datasets, we demonstrate the effectiveness and generalizability of the proposed method. The source code is available at https://github.com/Cookie-XT/3DAD. Chunlei Peng, Decheng Liu, Nannan Wang 0001, Xinbo Gao 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2025 | PrivacyHFR: Visual Privacy Preserving for Heterogeneous Face RecognitionabstractFace recognition has achieved remarkable progress and is widely deployed in real-world scenarios. Recently more and more attention has been given to individual privacy protection, due to unauthorized sensitive image leakage by malicious attackers. Multi-modality face images captured by diverse sensors, also called heterogeneous faces, bring in more challenges in face privacy protection while lacking related research. In this paper, we propose a novel visual Privacy preserving method for Heterogeneous Face Recognition (Privacy-HFR) to protect perceptual visual information and maintain essential identity information in multi-modality face analysis scenarios. Frequency domain analysis is a vital strategy to bridge the inevitable modality gap for heterogeneous face images. Meanwhile, recent theoretical insights also inspire us to design a suitable frequency component adjustment to balance human visual sensitivity and identity discriminative information. In addition, the ability to defend against recovery attacks has emerged as an essential criterion for privacy preserving face recognition. Noting that there seems to exist a dilemma that reducing accessible information by the attack model will affect the extracted identity information for recognition. It is because these two kinds of information are mutually blended in the frequency domain, which makes it a challenge to simultaneously maintain visual privacy and identity distinguishability. Thus, we provide a novel perspective to leverage the randomly optimal solutions and design the specific adversarial perturbations against the recovery attack. Experiments on several large-scale heterogeneous face datasets (CASIA NIR-VIS 2.0, LAMP-HQ, Tufts Face and CUFSF datasets) prove that the proposed method outperforms existing privacy-preserving face recognition methods in terms of recognition accuracy and privacy protection capability. The code is available in https://github.com/xiyin11/Privacy-HFR. Decheng Liu, Weizhao Yang, Chunlei Peng, Nannan Wang 0001, Ruimin Hu, Xinbo Gao 0001 |
IEEE Trans. Image Process. | 3 |
| 2025 | SketchAging: Face Photo-Sketch Synthesis and Aging With Multi-Scale Feature ExtractionabstractWith the rapid development of Generative Adversarial Networks (GANs), facial sketch generation and age transformation have advanced considerably. These technologies show great potential in digital media, entertainment, and forensic applications, particularly in helping law enforcement reconstruct the appearance of long-term fugitives. However, current methodologies exhibit notable limitations: existing approaches typically specialize in either facial sketch generation or age progression independently, lacking an effective integration for cross-domain synthesis. Moreover, preserving identity information while ensuring high-quality image generation remains a challenge. This paper proposes Multi-Scale Feature Extraction Networks (MSFE), an image-to-image translation framework that enables continuous age transformation while maintaining the stylistic characteristics of sketch domains. The core of the MSFS network uses a Dual Conditional Normalization Attention (DCNA) architecture to extract sketch features and encode facial images into the latent space of a pre-trained StyleGAN based on the desired age change. Experimental results on public datasets demonstrate that our approach outperforms existing methods, achieving superior facial photo-sketch synthesis with enhanced realism, identity preservation, and age accuracy. Chunlei Peng, Zhuang Tang, Decheng Liu, Nannan Wang 0001, Ruimin Hu, Xinbo Gao 0001 |
IEEE Trans. Image Process. | 1 |
| 2025 | Face Forgery Detection With CLIP-Enhanced Multi-Encoder DistillationabstractWith the development of face forgery technology, fake faces are rampant, threatening the security and authenticity of many fields. Therefore, it is of great significance to study face forgery detection. At present, existing detection methods have deficiencies in the comprehensiveness of feature extraction and model adaptability, and it is difficult to accurately deal with complex and changeable forgery scenarios. However, the rise of multimodal models provides new insights for current forgery detection methods. At present, most methods use relatively simple text prompts to describe the difference between real and fake faces. However, these researchers ignore that the CLIP model itself does not have the relevant knowledge of forgery detection. Therefore, our paper proposes a face forgery detection method based on multi-encoder fusion and cross-modal knowledge distillation. On the one hand, the prior knowledge of the CLIP model and the forgery model is fused. On the other hand, through the alignment distillation, the student model can learn the visual abnormal patterns and semantic features of the forged samples captured by the teacher model. Specifically, our paper extracts the features of face photos by fusing the CLIP text encoder and the CLIP image encoder, and uses the dataset in the field of forgery detection to pretrain and fine-tune the Deepfake-V2-Model to enhance the detection ability, which are regarded as the teacher model. At the same time, the visual and language patterns of the teacher model are aligned with the visual patterns of the pretrained student model, and the aligned representations are refined to the student model. This not only combines the rich representation of the CLIP image encoder and the excellent generalization ability of text embedding, but also enables the original model to effectively acquire relevant knowledge for forgery detection. Experiments show that our method effectively improves the performance on face forgery detection. Chunlei Peng, Tianzhe Yan, Decheng Liu, Nannan Wang 0001, Ruimin Hu, Xinbo Gao 0001 |
IEEE Trans. Image Process. | 1 |
| 2025 | Masked Attribute Description Embedding for Cloth-Changing Person Re-IdentificationabstractCloth-changing person re-identification (CC-ReID) aims to match persons who change clothes over long periods. The key challenge in CC-ReID is to extract cloth-irrelated features, such as face, hairstyle, body shape, and gait. Current research mainly focuses on modeling body shape using multi-modal biological features (such as silhouettes and sketches). However, it does not fully leverage the personal description information hidden in the original RGB image. Considering that there are certain attribute descriptions that remain unchanged after the changing of cloth, we propose a Masked Attribute Description Embedding (MADE) method that unifies personal visual appearance and attribute description for CC-ReID. Specifically, handling variable cloth-sensitive information, such as color and type, is challenging for effective modeling. To address this, we mask the clothes type and color information (upper body type, upper body color, lower body type, and lower body color) in the personal attribute description extracted through an attribute detection model. The masked attribute description is then connected and embedded into Transformer blocks at various levels, fusing it with the low-level to high-level features of the image. This approach compels the model to discard cloth information. Experiments are conducted on several CC-ReID benchmarks, including PRCC, LTCC, Celeb-reID-light, and LaST. Results demonstrate that MADE effectively utilizes attribute description, enhancing cloth-changing person re-identification performance, and compares favorably with state-of-the-art methods. Chunlei Peng, Decheng Liu, Nannan Wang 0001, Ruimin Hu, Xinbo Gao 0001 |
IEEE Trans. Multim. | 1 |
| 2025 | Semi-Supervised Learning for Anomaly Traffic Detection via Bidirectional Normalizing FlowsabstractWith the rapid development of the Internet, various types of anomaly traffic are threatening network security. However, the difficulty of collecting and labelling anomalous traffic is a significant challenge, so this paper proposes a semi-supervised anomaly detection framework. Considering normal and abnormal traffic have different data distributions, our framework can generate pseudo anomaly samples without prior knowledge of anomalies to achieve the detection of anomaly data. The framework comprises three principal components. Firstly, a pre-trained feature extractor is employed to extract a feature representation of the network traffic. Secondly, a bidirectional normalizing flow module establishes a reversible transformation between the latent data distribution and a Gaussian space. Through this bidirectional mapping, samples first undergo transformation manipulation within the Gaussian distribution space, and are then transported through the generative direction of normalizing flows, translating mathematical transformations into semantic feature evolutions in the latent data space. Finally, a simple classifier explicitly learns the potential differences between anomaly and normal samples to facilitate better anomaly detection. During inference, our framework requires only two modules to detect anomalous samples, leading to a considerable reduction in model size. According to the experiments, our method achieves the state-of-the-art results on the common benchmarking datasets of anomaly network traffic detection. Furthermore, it exhibits good generalisation performance across datasets. Zhangxuan Dang, Yu Zheng 0006, Xinglin Lian, Chunlei Peng, Qiuyu Chen, Xinbo Gao 0001 |
IEEE Trans. Netw. Serv. Manag. | 4 |
| 2024 | Adv-Diffusion: Imperceptible Adversarial Face Identity Attack via Latent Diffusion ModelabstractAdversarial attacks involve adding perturbations to the source image to cause misclassification by the target model, which demonstrates the potential of attacking face recognition models. Existing adversarial face image generation methods still can’t achieve satisfactory performance because of low transferability and high detectability. In this paper, we propose a unified framework Adv-Diffusion that can generate imperceptible adversarial identity perturbations in the latent space but not the raw pixel space, which utilizes strong inpainting capabilities of the latent diffusion model to generate realistic adversarial images. Specifically, we propose the identity-sensitive conditioned diffusion generative model to generate semantic perturbations in the surroundings. The designed adaptive strength-based adversarial perturbation algorithm can ensure both attack transferability and stealthiness. Extensive qualitative and quantitative experiments on the public FFHQ and CelebA-HQ datasets prove the proposed method achieves superior performance compared with the state-of-the-art methods without an extra generative model training process. The source code is available at https://github.com/kopper-xdu/Adv-Diffusion. Decheng Liu, Xijun Wang 0005, Chunlei Peng, Nannan Wang 0001, Ruimin Hu, Xinbo Gao 0001 |
AAAI | 3 |
| 2024 | Spatial-Frequency Dual-Stream Reconstruction for Deepfake Detection
Chunlei Peng, Decheng Liu, Yu Zheng 0006, Nannan Wang 0001 |
PRCV (11) | 1 |
| 2024 | Face Anti-spoofing Based on Multi-view Anomaly Detection
Jiuyao Jing, Chunlei Peng |
PRCV (15) | 4 |
| 2024 | Audio-Driven Face Photo-Sketch Video Generation
Siyue Zhou, Qun Guan, Chunlei Peng, Decheng Liu, Yu Zheng 0006 |
PRICAI (3) | 3 |
| 2024 | Face Anti-spoofing based on Multi-modal Dual-stream Anomaly DetectionabstractContemporary research often addresses the face anti-spoofing challenge through a classification paradigm. However, due to the rapidly changing and diverse characteristics of spoofing faces, it is unreasonable to regard all spoofing faces as a single category. Moreover, the swift evolution of spoofing techniques can render trained detectors ineffective. Anomaly detection offers a solution to these challenges by training exclusively on normal samples, distinguishing living samples as normal and non-living samples as anomalies. This paper presents a novel face anti-spoofing approach grounded in anomaly detection. It devises an RGB-D dual-stream network architecture that integrates multi-scale features from RGB and depth modalities via intermediate fusion. In addition, it also incorporates adversarial learning to facilitate network training. Furthermore, it proposes a novel multi-stage anomaly score generation technique for face anti-spoofing. Our experiments on three public datasets demonstrate the superiority of our method over comparative approaches. Jiuyao Jing, Yu Zheng 0006, Chunlei Peng |
TrustCom | 4 |
| 2024 | Pyramid-resolution person restoration for cross-resolution person re-identification
Chunlei Peng, Decheng Liu, Nannan Wang 0001, Xinbo Gao 0001 |
Sci. China Inf. Sci. | 1 |
| 2024 | Multi-view multi-label network traffic classification based on MLP-Mixer neural network
Yu Zheng 0006, Zhangxuan Dang, Xinglin Lian, Chunlei Peng, Xinbo Gao 0001 |
Comput. Networks | 4 |
| 2024 | GazeForensics: DeepFake detection via gaze-guided spatial inconsistency learning
Qinlin He, Chunlei Peng, Decheng Liu, Nannan Wang 0001, Xinbo Gao 0001 |
Neural Networks | 2 |
| 2024 | Local artifacts amplification for deepfakes augmentation
Chunlei Peng, Decheng Liu, Nannan Wang 0001, Xinbo Gao 0001 |
Neural Networks | 1 |
| 2024 | Hiding in Plain Sight: Adversarial Attack via Style Transfer on Image BordersabstractDeep Convolution Neural Networks (CNNs) have become the cornerstone of image classification, but the emergence of adversarial image attacks brings serious security risks to CNN-based applications. As a local perturbation attack, the border attack can achieve high success rates by only modifying the pixels around the border of an image, which is a novel attack perspective. However, existing border attacks have shortcomings in stealthiness and are easily detected. In this article, we propose a novel stealthy border attack method based on deep feature alignment. Specifically, we propose a deep feature alignment algorithm based on style transfer to guarantee the stealthiness of adversarial borders. The algorithm takes the deep feature difference between the adversarial and the original borders as the stealthiness loss and thus ensures good stealthiness of the generated adversarial images. To ensure high attack success rates simultaneously, we apply cross entropy to design the targeted attack loss and use margin loss as well as Leaky ReLU to design the untargeted attack loss. Experiments show that the structural similarity between the generated adversarial images and the original images is 8.8% higher than the state-of-art border attack method, indicating that our proposed adversarial images have better stealthiness. At the same time, the success rate of our attack in the face of defense methods is much higher, which is about four times that of the state-of-art border attack under the adversarial training defense. Xinghua Li 0001, Chunlei Peng, Yunwei Wang, Ning Zhang 0017, Yinbin Miao, Ximeng Liu, Kim-Kwang Raymond Choo |
IEEE Trans. Computers | 4 |
| 2024 | MRLReID: Unconstrained Cross-Resolution Person Re-Identification With Multi-Task Resolution LearningabstractCross-resolution person re-identification (ReID) is a challenging task that addresses the issue of matching individuals across different resolution conditions. Traditional person ReID methods often assume that images have sufficiently high resolution and overlook the practical scenarios involving low-resolution or blurry images. Existing cross-resolution ReID approaches either utilize image super-resolution techniques to improve the quality of low-resolution images or extract and learn resolution invariant features for person representation. Although multi-task learning has been applied in ReID to integrate auxiliary tasks including attribute recognition, image super-resolution, and so on, how to incorporate the vital resolution learning task into cross-resolution ReID has rarely explored before. Therefore, we propose a novel multi-task resolution learning based ReID network named MRLReID. Our approach treats ross-resolution person ReID as the primary task and the resolution estimation as an auxiliary task. Our network simultaneously learns the resolution information and person identity information of images, aiming to improve cross-resolution person ReID performance. Considering that existing similuated cross-resolution datasets are too simple to mimic unconstrained scenario, we further employ image degradation technique to simulate more realistic cross-resolution ReID datasets. We evaluate our method on two real-world cross-resolution datasets and two newly simulated cross-resolution datasets, and both intra-dataset and cross-dataset evaluations demonstrate the effectiveness and superiority of our method in cross-resolution person ReID. The codes and datasets are available at https://github.com/amateurbo/MRLReID. Chunlei Peng, Decheng Liu, Nannan Wang 0001, Ruimin Hu, Xinbo Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Universal Heterogeneous Face Analysis via Multi-Domain Feature DisentanglementabstractHeterogeneous face analysis is an important and challenge problem in face recognition community, because of the large modality discrepancy between heterogeneous face images. Existing methods either focus on transforming heterogeneous faces into the same style via face synthesis process, or intend to directly recognize heterogeneous face via modality invariant descriptors. However, the tasks of cross modality face synthesis and face recognition share a common purpose, which is to disentangle an inherent explainable representation. To this end, we propose a novel universal heterogenous face analysis method via multi-domain feature disentanglement, which does not need any face domain label. The proposed method explores to disentangle factors of variations of cross modality faces in an unsupervised manner. Then we could translate cross modality faces through modifying semantic factors, and the extracted inherent explainable representation still maintains being discriminative for heterogeneous face recognition. Experimental results on multiple cross modality face databases demonstrate the effectiveness of the proposed method. These experimental results also inspire us that the unsupervised disentangled module could help to analyze the interpretability of heterogenous face representation. Decheng Liu, Xinbo Gao 0001, Chunlei Peng, Nannan Wang 0001, Jie Li 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2024 | Detecting and Quantifying Crowd-Level Abnormal Behaviors in Crowd EventsabstractDetecting and quantifying abnormal crowd motion emerging from complex interactions of individuals is paramount to ensure the safety of crowds. Crowd-level abnormal behaviors (CABs), e.g., counter flow and crowd turbulence, are proven to be the crucial causes of many crowd disasters. Unlike individual-level anomaly, CABs usually do not exhibit salient difference from the normal behaviors when observed locally and the scale of CABs could vary from one scenario to another. It is also challenging to quantify the risk level of these CABs from video surveillance. In this paper, we present an improved version of our crowd motion learning framework for CABs detection, multi-scale motion consistency network (MSMC-Net) with a dual-attention fusion process to accommodate both the spatio-temporal and scale variations of different CABs. In addition, we propose an assessment method to quantify the risk level of detected CABs based on the anomaly score generated from our MSMC-Net. The risk quantification is performed in an online and accumulated manner and it can reflect the risk level of CABs consistent with other offline assessment metrics (e.g., crowd pressure), but without the extraction of detailed crowd data (e.g., pedestrian trajectories). For empirical study, we evaluate our method on large-scale crowd event datasets, including UMN, Hajj and Love Parade. Experimental results show that MSMC-Net could improve the AUC performance by 7.9%, 12.2% and 29.5% on three datasets respectively, compared to the best results of the state-of-the-art methods. Linbo Luo 0001, Shangwei Xie, Haiyan Yin, Chunlei Peng, Yew-Soon Ong |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2024 | Where Deepfakes Gaze at? Spatial-Temporal Gaze Inconsistency Analysis for Video Face Forgery DetectionabstractWith the continuous development of generative models on face generation, how to distinguish the real and fake face has become an important problem for security. Because of the continuous improvement on the detection accuracy by facial physiological signals, video face forgery detection based on facial physiological signal analysis has received more and more attention, which has become an important research branch in the field of face forgery detection. Currently, most of the research on forgery detection based on physiological signal analysis use biometric features such as blinking patterns, head swings, heart rate signals, and lip movements. However, there hasn’t been much exploration on the usage of gaze features in face forgery detection. Through the analysis of gaze directions in face videos, we have observed differences in the distribution of gaze direction pattern between the real and forged videos. Specifically, real videos tend to have more concentrated gaze distribution within a short period of time, while forged videos have more dispersed gaze distributions. In this paper, we present a novel Deepfake gaze analysis method named DFGaze, to explore spatial-temporal gaze inconsistency for video face forgery detection. Our method uses the gaze analysis model (GAM) to analyze the gaze features of face video frames, and then applies a spatial-temporal feature aggregator to realize authenticity classification based on gaze features. In order to better mine the authenticity clues in the videos, we further use the texture analysis model (TAM) and attribute analysis model (AAM) to improve the representation ability of spatial-temporal feature differences between real and forged faces. Extensive experiments show that our method can achieve state-of-the-art performance with the help of gaze analysis. The source code is available at https://github.com/ziminMIAO/DFGaze. Chunlei Peng, Zimin Miao, Decheng Liu, Nannan Wang 0001, Ruimin Hu, Xinbo Gao 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2024 | Hierarchical Forgery Classifier on Multi-Modality Face Forgery CluesabstractFace forgery detection plays an important role in personal privacy and social security. With the development of adversarial generative models, high-quality forgery images become more and more indistinguishable from real to humans. Existing methods always regard as forgery detection task as the common binary or multi-label classification, and ignore exploring diverse multi-modality forgery image types, e.g. visible light spectrum and near-infrared scenarios. In this article, we propose a novelHierarchicalForgeryClassifier forMulti-modalityFaceForgeryDetection(HFC-MFFD), which could effectively learn robust patches-based hybrid domain representation to enhance forgery authentication in multiple modality scenarios. The local hybrid domain representation is designed to explore strong discriminative forgery clues both in the image and frequency domain with the intra-attention mechanism. Furthermore, the specific hierarchical face forgery classifier is designed through the authenticity feedback strategy to integrate diverse discriminative clues. Experimental results on representative multi-modality face forgery datasets demonstrate the superior performance of the proposed HFC-MFFD compared with state-of-the-art algorithms. Decheng Liu, Zeyang Zheng, Chunlei Peng, Nannan Wang 0001, Xinbo Gao 0001 |
IEEE Trans. Multim. | 3 |
| 2024 | Disguised Heterogeneous Face Generation With Iterative-Adversarial Style UnificationabstractHeterogeneous face recognition (HFR), which refers to matching face images with different modalities, is essential to public safety. Although HFR has made promising progress in recent years, disguised faces in HFR scenarios still remain a major challenge for the following reasons. First, most existing HFR methods focus on traditional scenarios without disguised accessories, and the performance degrades when dealing directly with disguised faces. Second, there is a need for disguised heterogeneous face datasets, which is essential for developing the related research community. Third, colorful accessories are distinct from heterogeneous face images in terms of their modalities, and their direct combination results in style inconsistency and poor quality. Therefore, we propose a disguised heterogeneous face generation method based on an iterative-adversarial style unification framework. Our approach aims to gradually learn frame textures to detail textures in multiple confrontation iterations, resulting in style unification for disguised accessories and heterogeneous faces. We also construct a disguised heterogeneous face dataset, which contains a disguised NIR-VIS subset and a disguised sketch-photo subset. Moreover, we provide benchmark evaluations conducted on our proposed dataset with face recognition and image quality assessment, demonstrating the superiority of our method over direct addition and two representative disguised face generation techniques. Chunlei Peng, Zimo Kong, Decheng Liu, Nannan Wang 0001, Xinbo Gao 0001 |
IEEE Trans. Multim. | 1 |
| 2023 | A Semi-Supervised Anomaly Network Traffic Detection Framework via Multimodal Traffic Information FusionabstractAnomaly traffic detection is a crucial issue in the cyber-security field. Previously, many researchers regarded anomaly traffic detection as a supervised classification problem. However, in real scenarios, anomaly network traffic is unpredictable, dynamically changing and difficult to collect. To address these limitations, we employ anomaly detection setting to propose a novel semi-supervised anomaly network traffic detection framework. It only learns features of normal samples during the training phase. Our framework utilizes low-pass filtering to extract multi-scale low-frequency information from 2-D traffic image. Furthermore, we design a two-stage fusion scheme to incorporate information from original and multi-scale low-frequency traffic image modalities. We conduct experiments on two public datasets: ISCX Tor-nonTor and USTC-TFC2016. The experimental results show that our method outperforms current state-of-the-art anomaly detection methods. Yu Zheng 0006, Xinglin Lian, Zhangxuan Dang, Chunlei Peng, Chao Yang 0016, Jianfeng Ma 0001 |
CIKM | 4 |
| 2023 | Modality-agnostic Augmented Multi-Collaboration Representation for Semi-supervised Heterogenous Face RecognitionabstractHeterogeneous face recognition (HFR) aims to match input face identity across different image modalities. Due to the existing large modality gap and the limited number of training data, HFR is still a challenging problem in biometrics and draws more and more attention. Existing researchers always extract modality invariant features or generate homogeneous images to decrease the modality gap, lacking abundant labeled data to avoid the overfitting problem. In this paper, we proposed a novel Modality-Agnostic Augmented Multi-Collaboration representation for Heterogeneous Face Recognition (MAMCO-HFR) in a semi-supervised manner. The modality-agnostic augmentation strategy is proposed to generate adversarial perturbations to map unlabeled faces into the modality-agnostic domain. The multi-collaboration feature constraint is designed to mine the inherent relationships between diverse layers for discriminative representation. Experiments on several large-scale heterogeneous face datasets (CASIA NIR-VIS 2.0, LAMP-HQ and Tufts Face dataset) prove the proposed algorithm can achieve superior performance compared with state-of-the-art methods. The source code is available at https://github.com/xiyin11/Semi-HFR. Decheng Liu, Weizhao Yang, Chunlei Peng, Nannan Wang 0001, Ruimin Hu, Xinbo Gao 0001 |
ACM Multimedia | 3 |
| 2023 | Face photo-sketch synthesis via intra-domain enhancement
Chunlei Peng, Congyu Zhang, Decheng Liu, Nannan Wang 0001, Xinbo Gao 0001 |
Knowl. Based Syst. | 1 |
| 2023 | Spatial-Temporal Frequency Forgery Clue for Video Forgery Detection in VIS and NIR ScenarioabstractIn recent years, with the rapid development of face editing and generation, more and more fake videos are circulating on social media, which has caused extreme public concerns. Existing face forgery detection methods based on frequency domain find that the GAN forged images have obvious grid-like visual artifacts in the frequency spectrum. But for synthesized videos, these methods only confine to a single frame and pay little attention to the most discriminative part and temporal frequency clue among different frames. To take full advantage of the rich information in video sequences, this paper performs video forgery detection on both spatial and temporal frequency domains and proposes a Discrete Cosine Transform-based Forgery Clue Augmentation Network (FCAN-DCT) to achieve a more comprehensive spectrum spatial-temporal feature representation. FCAN-DCT totally consists of a backbone network and two branches: Compact Feature Extraction (CFE) module and Frequency Temporal Attention (FTA) module. We conduct thorough experimental assessments on three visible light (VIS) based datasets (i.e.,, FaceForensics++, Celeb-DF (v2), WildDeepfake), and our self-built video forgery dataset DeepfakeNIR, which is the first video forgery dataset on near-infrared (NIR) modality. The experimental results demonstrate the effectiveness and robustness of our method for detecting forgery videos in both VIS and NIR scenarios.DeepfakeNIR and code are available athttps://github.com/AEP-WYK/DeepfakeNIR. Chunlei Peng, Decheng Liu, Nannan Wang 0001, Xinbo Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | FedForgery: Generalized Face Forgery Detection With Residual Federated LearningabstractWith the continuous development of deep learning in the field of image generation models, a large number of vivid forged faces have been generated and spread on the Internet. These high-authenticity artifacts could grow into a threat to society security. Existing face forgery detection methods directly utilize the obtained public shared or centralized data for training but ignore the personal privacy and security issues when personal data couldn’t be centralizedly shared in real-world scenarios. Additionally, different distributions caused by diverse artifact types would further bring adverse influences on the forgery detection task. To solve the mentioned problems, the paper proposes a novel generalized residual Federated learning for face Forgery detection (FedForgery). The designed variational autoencoder aims to learn robust discriminative residual feature maps to detect forgery faces (with diverse or even unknown artifact types). Furthermore, the general federated learning strategy is introduced to construct distributed detection model trained collaboratively with multiple local decentralized devices, which could further boost the representation generalization. Experiments conducted on publicly available face forgery detection datasets prove the superior performance of the proposed FedForgery. The designed novel generalized face forgery detection protocols and source code would be publicly available at https://github.com/GANG370/FedForgery. Decheng Liu, Zhan Dang, Chunlei Peng, Yu Zheng 0006, Nannan Wang 0001, Xinbo Gao 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2023 | HiFiSketch: High Fidelity Face Photo-Sketch Synthesis and ManipulationabstractWith the rapid development of generative adversarial networks, face photo-sketch synthesis has achieved promising performance and playing an increasingly important role in law enforcement as well as entertainment. However, most of the existing methods only work under the condition of no interference, and lack of generalization ability in wild scenes. The fidelity of the images generated by the existing methods are insufficient, and the manipulation ability according to text description is unavailable. Directly applying existing text-based image manipulation methods on face photo-sketch scenario may lead to severe distortions due to the cross-domain challenges. Therefore, we propose a novel cross-domain face photo-sketch synthesis framework named HiFiSketch, a network that learns to adjust the weights of generators for high-fidelity synthesis and manipulation. It can realize the translation of images between the photo domain and the sketch domain, and modify results according to the text input in the meanwhile. We further propose a cross-domain loss function, which can effectively preserve facial details during face photo-sketch synthesis. Extensive experiments on four public face sketch datasets show the superiority of our method compared to existing methods. We further present text-based face photo-sketch manipulation and sequential face photo-sketch manipulation for the first time to demonstrate the effectiveness of our method on high fidelity face photo-sketch synthesis and manipulation. Chunlei Peng, Congyu Zhang, Decheng Liu, Nannan Wang 0001, Xinbo Gao 0001 |
IEEE Trans. Image Process. | 1 |
| 2022 | SketchCLIP: Text-based Attribute Manipulation for Face Sketch SynthesisabstractThis paper proposes a method of modifying the face sketch with text descriptions. Face sketch is widely used in the criminal field and digital entertainment field. Forensic painters usually draw face sketches based on descriptions provided by witnesses or clients. However, drawing a face sketch often takes lots of time and effort. Existing face sketch synthesis studies have not considered text-based sketch manipulation, and we find that applying text-driven editing methods on natural images directly to face sketches causes severe distortion of generated results. Therefore, this paper proposes a novel text-based attribute manipulation method for face sketch synthesis, named SketchCLIP. Our approach adopts text-driven attribute manipulation by using the powerful Contrastive Language-Image Pre-Training (CLIP) model, which not only conforms to the current drawing process of face sketches but also does not require tedious manual operations and allows for more diverse modifications. Besides, we design an intra-modality fine-tuning module to eliminate distortion and improve the quality of the modified face sketch. Through extensive comparison experiments on public face sketch datasets, our method is demonstrated to be very excellent in the effectiveness of the face sketch processing and the quality of modified results. Mengdi Dong, Chunlei Peng, Decheng Liu, Yu Zheng 0006, Nannan Wang 0001, Xinbo Gao 0001 |
IJCB | 2 |
| 2022 | ForgeryNIR: Deep Face Forgery and Detection in Near-Infrared ScenarioabstractDeep face forgery and detection is an emerging topic due to the development of GANs. Face forgery detection relies greatly on existing databases for evaluation and adequate training examples for data-hungry machine learning algorithms. However, considering the wide application of face recognition in near-infrared scenarios, there is no publicly available face forgery database that includes near-infrared modality currently. In this paper, we present an attempt at constructing a large-scale dataset for face forgery detection in the near-infrared modality and propose a new forgery detection method based on knowledge distillation named cross-modality knowledge distillation aiming to use a teacher model which is pre-trained on the visible light-based (VIS) big data to guide the student model with a small amount of near-infrared (NIR) data. The proposed near-infrared face forgery dataset, named ForgeryNIR, contains a total of over 50,000 real and fake identities. A number of perturbations are applied to help simulate real-world scenarios. All source images in ForgeryNIR are collected from CASIA NIR-VIS 2.0, and fake images are generated via multiple GAN techniques. The proposed dataset fills the gap of face forgery detection research in the near-infrared modality. A comprehensive study on six representative detection baselines is conducted to evaluate the performance of face forgery detection algorithms in the NIR domain. We further construct a hard testing set, named ForgeryNIR+, which contains forged images that have bypassed existing face forgery detection methods. The proposed datasets will be publicly available and aim to help boost further research on face forgery detection, as well as NIR face detection and recognition. Chunlei Peng, Decheng Liu, Nannan Wang 0001, Xinbo Gao 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2022 | Edge Aware Domain Transformation for Face Sketch SynthesisabstractWith the development of generative adversarial networks (GAN), the field of face sketch synthesis has received extensive attention. Face sketch synthesis (FSS) has promising prospects in the fields of entertainment and law enforcement, where it plays an increasingly important role. We propose a novel generative adversarial network for synthesizing sketches with similar shapes and rich details to photos. This problem is challenging because it involves the transition between the sketch domain and the photo domain. Many methods have been used for face sketch synthesis in recent years, but existing methods cannot fully exploit the semantic information between different domains. To this end, we use a cross-domain face sketch synthesis framework based on edge-preserving filters to make the boundaries of different semantics in semantic layouts have a smooth transition. We further propose a new spatially adaptive denormalization module named edge-aware enhancement Spatially Adaptive DEnormalization (eaeSPADE), which can make full use of the semantic information in the semantic layout of faces and improve the details of the synthesized face images. Extensive experiments demonstrate that our method outperforms existing face sketch synthesis methods. Congyu Zhang, Decheng Liu, Chunlei Peng, Nannan Wang 0001, Xinbo Gao 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2022 | Towards Multi-Domain Face Synthesis Via Domain-Invariant Representations and Multi-Level Feature PartsabstractCross-domain face synthesis plays a positive role in the real world. It is challenging to synthesize high-quality faces across multiple domains based on limited paired data because the multiple mappings between different domains may interfere with each other. Cognitive science investigates that the brain can recognize the same person with multiple different expressions by extracting invariant information on the face and we humans perceive instances by decomposing them into parts. Motivated by these cognition, we propose a unified semi-supervised framework for multi-domain face synthesis by extracting a domain-invariant representation and exploiting parts of multi-level features. Specifically, realized by adversarial training with additional ability to utilize domain-specific information, a encoder is trained to remove domain-specific information and extract the domain-invariant representation from multiple inputs. Then, we utilize the multi-level feature parts extracted from inputs and reconstructed faces via a pre-trained recognition model to ensure that the domain-invariant representation contains enough useful semantic information. we also utilize the feature parts extracted from inputs and limited paired data to compose pseudo features in target domain for supervising the synthesis, which makes our framework suitable for large amounts of unpaired training data. By exploiting this framework, we can achieve face synthesis between multiple domains using some paired data together with a large training database without ground truth target faces. Experimental results demonstrate our framework achieves great performances on qualitative and quantitative evaluations under both artificial and uncontrolled environments, and our framework has competitive performances in single translation compared with specialized methods for translation between two specific domains. Dawei Zhou 0004, Nannan Wang 0001, Chunlei Peng, Yi Yu 0001, Xi Yang 0011, Xinbo Gao 0001 |
IEEE Trans. Multim. | 3 |
| 2022 | Heterogeneous Face Interpretable Disentangled Representation for Joint Face Recognition and SynthesisabstractHeterogeneous faces are acquired with different sensors, which are closer to real-world scenarios and play an important role in the biometric security field. However, heterogeneous face analysis is still a challenging problem due to the large discrepancy between different modalities. Recent works either focus on designing a novel loss function or network architecture to directly extract modality-invariant features or synthesizing the same modality faces initially to decrease the modality gap. Yet, the former always lacks explicit interpretability, and the latter strategy inherently brings in synthesis bias. In this article, we explore to learn the plain interpretable representation for complex heterogeneous faces and simultaneously perform face recognition and synthesis tasks. We propose the heterogeneous face interpretable disentangled representation (HFIDR) that could explicitly interpret dimensions of face representation rather than simple mapping. Benefited from the interpretable structure, we further could extract latent identity information for cross-modality recognition and convert the modality factor to synthesize cross-modality faces. Moreover, we propose a multimodality heterogeneous face interpretable disentangled representation (M-HFIDR) to extend the basic approach suitable for the multimodality face recognition and synthesis. To evaluate the ability of generalization, we construct a novel large-scale face sketch data set. Experimental results on multiple heterogeneous face databases demonstrate the effectiveness of the proposed method. Decheng Liu, Xinbo Gao 0001, Chunlei Peng, Nannan Wang 0001, Jie Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2021 | Removing Adversarial Noise in Class Activation Feature SpaceabstractDeep neural networks (DNNs) are vulnerable to adversarial noise. Pre-processing based defenses could largely remove adversarial noise by processing inputs. However, they are typically affected by the error amplification effect, especially in the front of continuously evolving attacks. To solve this problem, in this paper, we propose to remove adversarial noise by implementing a self-supervised adversarial training mechanism in a class activation feature space. To be specific, we first maximize the disruptions to class activation features of natural examples to craft adversarial examples. Then, we train a denoising model to minimize the distances between the adversarial examples and the natural examples in the class activation feature space. Empirical evaluations demonstrate that our method could significantly enhance adversarial robustness in comparison to previous state-of-the-art approaches, especially against unseen adversarial attacks and adaptive attacks. Dawei Zhou 0004, Nannan Wang 0001, Chunlei Peng, Xinbo Gao 0001, Xiaoyu Wang 0002, Jun Yu 0001, Tongliang Liu |
ICCV | 3 |
| 2021 | Towards Defending against Adversarial Examples via Attack-Invariant FeaturesabstractDeep neural networks (DNNs) are vulnerable to adversarial noise. Their adversarial robustness can be improved by exploiting adversarial examples. However, given the continuously evolving attacks, models trained on seen types of adversarial examples generally cannot generalize well to unseen types of adversarial examples. To solve this problem, in this paper, we propose to remove adversarial noise by learning generalizable invariant features across attacks which maintain semantic classification information. Specifically, we introduce an adversarial feature learning mechanism to disentangle invariant features from adversarial noise. A normalization term has been proposed in the encoded space of the attack-invariant features to address the bias issue between the seen and unseen types of attacks. Empirical evaluations demonstrate that our method could provide better protection in comparison to previous state-of-the-art approaches, especially against unseen types of attacks and adaptive attacks. Dawei Zhou 0004, Tongliang Liu, Bo Han 0003, Nannan Wang 0001, Chunlei Peng, Xinbo Gao 0001 |
ICML | 5 |
| 2021 | Iterative local re-ranking with attribute guided synthesis for face sketch recognition
Decheng Liu, Xinbo Gao 0001, Nannan Wang 0001, Chunlei Peng, Jie Li 0001 |
Pattern Recognit. | 4 |
| 2021 | Soft Semantic Representation for Cross-Domain Face RecognitionabstractThe problem of cross-domain face recognition aims to identify facial images obtained across different domains, which attracts increasing attentions because of its wide applications on law-enforcement identification and camera surveillance. The problem is challenging due to the huge domain discrepancy. Despite great progress achieved in recent years, existing algorithms usually fail to fully exploit the semantic information for identifying cross-domain faces, which could be a strong clue for recognition. In this article, we propose an effective algorithm for cross-domain face recognition by exploiting semantic information integrated with deep convolutional neural networks (CNN). We first introduce a soft face parsing algorithm where the boundaries of facial components are measured as probabilistic values. By taking the original face image as the guidance to improve face parsing result, each pixel may belong partially to the facial component to avoid inaccurate segmentation around component boundaries. We then propose a hierarchical soft semantic representation framework for cross-domain face recognition. Both the soft semantic level and contour level deep features obtained via CNN are computed and combined together, which could fully exploit the identical semantic clue among cross-domain faces. We provide extensive experiments to demonstrate that the proposed soft semantic representation algorithm performs superior against state-of-the-art methods. Chunlei Peng, Nannan Wang 0001, Jie Li 0001, Xinbo Gao 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2020 | Face Sketch Synthesis in the Wild via Deep Patch Representation-Based Probabilistic Graphical ModelabstractThis paper considers the problem of face sketch synthesis in the wild, which transforms a face photo into a face sketch. Face sketch synthesis is widely applied in law enforcement as well as digital entertainment fields. However, the existing methods either focus on hand-crafted techniques where prior human experience is relied on or adopt deep learning techniques as an end-to-end framework, where facial details cannot be well represented. In this paper, we propose a novel approach for face sketch synthesis in the wild via a deep patch representation-based probabilistic graphical model (DeepPGM). A Siamese network is constructed to extract deep patch representation from a raw facial patch, where the representative detail information for robust face sketch synthesis can be exploited. The generated deep patch representation and facial image patches are then optimally combined through a probabilistic graphical model. The proposed DeepPGM approach not only outperforms the state-of-the-art on public face sketch datasets but also can cope with forensic photos in the wild conditions, including varying lightings, poses, occlusions, skin colors, and ethnic origins. The superiority of the proposed method is demonstrated by extensive experiments on two public face sketch datasets and real-world forensic photos in the wild. Chunlei Peng, Nannan Wang 0001, Jie Li 0001, Xinbo Gao 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2020 | Universal Face Photo-Sketch Style Transfer via Multiview Domain TranslationabstractFace photo-sketch style transfer aims to convert a representation of a face from the photo (or sketch) domain to the sketch (respectively, photo) domain while preserving the character of the subject. It has wide-ranging applications in law enforcement, forensic investigation and digital entertainment. However, conventional face photo-sketch synthesis methods usually require training images from both the source domain and the target domain, and are limited in that they cannot be applied to universal conditions where collecting training images in the source domain that match the style of the test image is unpractical. This problem entails two major challenges: 1) designing an effective and robust domain translation model for the universal situation in which images of the source domain needed for training are unavailable, and 2) preserving the facial character while performing a transfer to the style of an entire image collection in the target domain. To this end, we present a novel universal face photo-sketch style transfer method that does not need any image from the source domain for training. The regression relationship between an input test image and the entire training image collection in the target domain is inferred via a deep domain translation framework, in which a domain-wise adaption term and a local consistency adaption term are developed. To improve the robustness of the style transfer process, we propose a multiview domain translation method that flexibly leverages a convolutional neural network representation with hand-crafted features in an optimal way. Qualitative and quantitative comparisons are provided for universal unconstrained conditions of unavailable training images from the source domain, demonstrating the effectiveness and superiority of our method for universal face photo-sketch style transfer. Chunlei Peng, Nannan Wang 0001, Jie Li 0001, Xinbo Gao 0001 |
IEEE Trans. Image Process. | 1 |
| 2020 | Coupled Attribute Learning for Heterogeneous Face RecognitionabstractHeterogeneous face recognition (HFR) is a challenging problem in face recognition and subject to large textural and spatial structure differences of face images. Different from conventional face recognition in homogeneous environments, there exist many face images taken from different sources (including different sensors or different mechanisms) in reality. In addition, limited training samples of cross-modality pairs make HFR more challenging due to the complex generation procedure of these images. Despite the great progress that has been achieved in recent years, existing works mainly focus on HFR from only cross-modality image matching. However, it is more practical to obtain both facial images and semantic descriptions about facial attributes in real-world situations, in which the semantic description clues are nearly always obtained during the process of image generation. Motivated by human cognitive mechanisms, we naturally utilize the explicit invariant semantic description, i.e., face attributes, to help address the gap among face images of different modalities. Existing facial attributes-related face recognition methods primarily regard attributes as the high-level features used to enhance recognition performance, ignoring the inherent relationship between face attributes and identities. In this article, we propose novel coupled attribute learning for the HFR (CAL-HFR) method without labeling the attributes manually. Deep convolutional networks are employed to directly map face images in heterogeneous scenarios to a compact common space where distances are taken as dissimilarities of pairs. Coupled attribute guided triplet loss (CAGTL) is designed to train an end-to-end HFR network that can effectively eliminate defects of incorrectly estimated attributes. Extensive experiments on multiple heterogeneous scenarios demonstrate that the proposed method achieves superior performance compared with that of state-of-the-art methods. Furthermore, we make publicly available our generated pairwise annotated heterogeneous facial attribute database for evaluation and promoting related research. Decheng Liu, Xinbo Gao 0001, Nannan Wang 0001, Jie Li 0001, Chunlei Peng |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2019 | DLFace: Deep local descriptor for cross-modality face recognition
Chunlei Peng, Nannan Wang 0001, Jie Li 0001, Xinbo Gao 0001 |
Pattern Recognit. | 1 |
| 2019 | Sparse graphical representation based discriminant analysis for heterogeneous face recognition
Chunlei Peng, Xinbo Gao 0001, Nannan Wang 0001, Jie Li 0001 |
Signal Process. | 1 |
| 2019 | Re-Ranking High-Dimensional Deep Local Representation for NIR-VIS Face RecognitionabstractHeterogeneous face recognition refers to matching facial images captured from different sensors or sources, which has wide applications in public security and law enforcement. Because of the great differences in sensing and creating procedure, there are huge feature gap between heterogeneous facial images. Existing methods merely focus on comparing the probe image with the gallery in feature space, while the true target may not appear at the first rank due to the appearance variations caused by different sensing patterns. In order to exploit valuable information from initial ranking result, this paper proposes to re-rank high-dimensional deep local representation for matching near-infrared (NIR) and visual (VIS) facial images, i.e. NIR-VIS face recognition. A high-dimensional deep local representation is firstly constructed by extracting and concatenating deep features on local facial patches via a convolutional neural network (CNN). The initial NIR-VIS recognition ranking results can be obtained by comparing the compressed deep features. We then propose a novel and efficient locally linear re-ranking (LLRe-Rank) technique to refine the initial ranking results, which can explore valuable information from initial ranking result. The proposed re-ranking method does not require any human interaction or data annotation, and can be served as an unsupervised post processing technique. Experimental results on the most challenging Oulu-CASIA NIR-VIS database and CASIA NIR-VIS 2.0 database demonstrate the effectiveness of our method. Chunlei Peng, Nannan Wang 0001, Jie Li 0001, Xinbo Gao 0001 |
IEEE Trans. Image Process. | 1 |
| 2018 | Deep Attribute Guided Representation for Heterogeneous Face RecognitionabstractHeterogeneous face recognition (HFR) is a challenging problem in face recognition, subject to large texture and spatial structure differences of face images. Different from conventional face recognition in homogeneous environments, there exist many face images taken from different sources (including different sensors or different mechanisms) in reality. Motivated by human cognitive mechanism, we naturally utilize the explicit invariant semantic information (face attributes) to help address the gap of different modalities. Existing related face recognition methods mostly regard attributes as the high level feature integrated with other engineering features enhancing recognition performance, ignoring the inherent relationship between face attributes and identities. In this paper, we propose a novel deep attribute guided representation based heterogeneous face recognition method (DAG-HFR) without labeling attributes manually. Deep convolutional networks are employed to directly map face images in heterogeneous scenarios to a compact common space where distances mean similarities of pairs. An attribute guided triplet loss (AGTL) is designed to train an end-to-end HFR network which could effectively eliminate defects of incorrectly detected attributes. Extensive experiments on multiple heterogeneous scenarios (composite sketches, resident ID cards) demonstrate that the proposed method achieves superior performances compared with state-of-the-art methods. Decheng Liu, Nannan Wang 0001, Chunlei Peng, Jie Li 0001, Xinbo Gao 0001 |
IJCAI | 3 |
| 2018 | Composite components-based face sketch recognition
Decheng Liu, Jie Li 0001, Nannan Wang 0001, Chunlei Peng, Xinbo Gao 0001 |
Neurocomputing | 4 |
| 2018 | Face recognition from multiple stylistic sketches: Scenarios, datasets, and evaluation
Chunlei Peng, Xinbo Gao 0001, Nannan Wang 0001, Jie Li 0001 |
Pattern Recognit. | 1 |
| 2017 | Adaptive representation-based face sketch-photo synthesis
Jie Li 0001, Xinye Yu, Chunlei Peng, Nannan Wang 0001 |
Neurocomputing | 3 |
| 2017 | Graphical Representation for Heterogeneous Face RecognitionabstractHeterogeneous face recognition (HFR) refers to matching face images acquired from different sources (i.e., different sensors or different wavelengths) for identification. HFR plays an important role in both biometrics research and industry. In spite of promising progresses achieved in recent years, HFR is still a challenging problem due to the difficulty to represent two heterogeneous images in a homogeneous manner. Existing HFR methods either represent an image ignoring the spatial information, or rely on a transformation procedure which complicates the recognition task. Considering these problems, we propose a novel graphical representation based HFR method (G-HFR) in this paper. Markov networks are employed to represent heterogeneous image patches separately, which takes the spatial compatibility between neighboring image patches into consideration. A coupled representation similarity metric (CRSM) is designed to measure the similarity between obtained graphical representations. Extensive experiments conducted on multiple HFR scenarios (viewed sketch, forensic sketch, near infrared image, and thermal infrared image) show that the proposed method outperforms state-of-the-art methods. Chunlei Peng, Xinbo Gao 0001, Nannan Wang 0001, Jie Li 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2017 | Superpixel-Based Face Sketch-Photo SynthesisabstractFace sketch-photo synthesis technique has attracted growing attention in many computer vision applications, such as law enforcement and digital entertainment. Existing methods either simply perform the face sketch-photo synthesis on the holistic image or divide the face image into regular rectangular patches ignoring the inherent structure of the face image. In view of such situations, this paper presents a novel superpixel-based face sketch-photo synthesis method by estimating the face structures through image segmentation. In our proposed method, face images are first segmented into superpixels, which are then dilated to enhance the compatibility of neighboring superpixels. Each input face image induces a specific graphical structure modeled by Markov networks. We employ a two-stage synthesis process to learn the face structures through Markov networks constructed from two scales of dilation, respectively. Experiments on several public databases demonstrate that our proposed face sketch-photo synthesis method achieves superior performance compared with the state-of-the-art methods. Chunlei Peng, Xinbo Gao 0001, Nannan Wang 0001, Jie Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2016 | Multiple Representations-Based Face Sketch-Photo SynthesisabstractFace sketch-photo synthesis plays an important role in law enforcement and digital entertainment. Most of the existing methods only use pixel intensities as the feature. Since face images can be described using features from multiple aspects, this paper presents a novel multiple representations-based face sketch-photo-synthesis method that adaptively combines multiple representations to represent an image patch. In particular, it combines multiple features from face images processed using multiple filters and deploys Markov networks to exploit the interacting relationships between the neighboring image patches. The proposed framework could be solved using an alternating optimization strategy and it normally converges in only five outer iterations in the experiments. Our experimental results on the Chinese University of Hong Kong (CUHK) face sketch database, celebrity photos, CUHK Face Sketch FERET Database, IIIT-D Viewed Sketch Database, and forensic sketches demonstrate the effectiveness of our method for face sketch-photo synthesis. In addition, cross-database and database-dependent style-synthesis evaluations demonstrate the generalizability of this novel method and suggest promising solutions for face identification in forensic science. Chunlei Peng, Xinbo Gao 0001, Nannan Wang 0001, Dacheng Tao, Xuelong Li 0001, Jie Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |