Ajian Liu 0001

dblp:232/2295-1 · DBLP profile ↗
← Back
43ranked-venue papers
11as first author
41since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 24 · 5 first-author · 22 since 2021Artificial intelligence and machine learning · 21 · 6 first-author · 19 since 2021Security and privacy · 14 · 3 first-author · 13 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 PointMC: Multi-view Consistent Encoding and Center-Global Feature Fusion for Point Clouds Understanding
abstract
Point cloud tasks have recently benefited from Mamba-based architecture, which leverage state space modeling to achieve strong performance. Previous studies have primarily focused on network design while overlooking the importance of position encoding and relying on coarse-grained geometric feature aggregation. The former leads to semantic ambiguity due to inconsistent spatial relationships, while the latter results in geometric feature dispersion by overlooking fine-grained local geometric details. To tackle the above problem, we propose a novel framework, PointMC, including Multi-view Consistent Learnable Position Encoding (MCLPE) and Center-Global Feature Fusion (CGFF), to provide semantically coherent positional guidance for inter-patch and enable fine-grained geometric structure aggregation within intra-patch regions. Specifically, the proposed MCLPE module is inspired by a spatial structure modeling mechanism guided by physical constraints, leverages multi-view virtual reconstruction and a learnable strategy to dynamically constrain spatial relationships along patch boundaries, thereby enhancing the semantic consistency and representational clarity across inter-patch regions. Furthermore, considering the lack of local structural information within each patch, the CGFF module employs a dual-guidance mechanism based on center and global structures to effectively promote the aggregation of local geometric features. Extensive experiments on multiple benchmark datasets validate the effectiveness of PointMC, consistently outperforming existing state-of-the-art methods, and demonstrating superior capability in capturing both inter-patch semantic consistency and intra-patch geometric details.
Xinxing Yu, Ajian Liu 0001, Sunyuan Qiang, Hui Ma 0018, Yanyan Liang 0001
AAAI2
2026 UniAttack: Unified Physical-Digital Face Attack Detection
Shunxin Chen, Ajian Liu 0001, Haocheng Yuan, Junze Zheng, Dingheng Zeng, Jiankang Deng, Sergio Escalera, Xiaoming Liu 0002, Jun Wan 0001, Zhen Lei 0001
Int. J. Comput. Vis.2
2026 ICPE-FAS: Instance and Category Prompts Engineering for Generalizable Face Anti-Spoofing
Ajian Liu 0001, Xun Lin, Hui Ma 0018, Xinxing Yu, Jiabao Guo, Zitong Yu, Jun Wan 0001, Zhanchuan Cai, Zhen Lei 0001, Yanyan Liang 0001
Int. J. Comput. Vis.1
2026 Context-aware knowledge distillation for anomaly detection
Ning Li 0035, Xuxin Lin, Ajian Liu 0001, Chaohao Jiang, Zhenwei Zhu, Yanyan Liang 0001
Knowl. Based Syst.3
2026 DGPDL: Domain-Guided Prompt Distribution Learning for Generalizable Face Anti-Spoofing
abstract
The overfitting of domain signals results in poor domain generalization of face anti-spoofing. The current methods usually improve the diversity of source domains to alleviate this overfitting. However, this benefit is minimal, as even the most diverse domain signals will also be absent in the target domain. In this work, we propose a Domain-Guided Prompt Distribution Learning (DGPDL) built on Vision-Language Models like CLIP, which explores a unified representation of domain signals as a prompt across the source and target domain to alleviate the understanding bias caused by domain gaps. Specifically, we first define a learnable Domain-Specific Distribution (DSD) that covers as many domain elements as possible, such as image quality, color tone, camera settings, etc., which establish connections between different domains and linearly combinable prompt in any domain; Then, based on the style statistics of the given sample, we construct its optimal Domain-Specific Prompts (DSPs) from the defined DSD through the designed Prompt Assemble Attention (PAA) with the similarity matching; Finally, the assembled DSPs will act as carrier or agent to perform on both the vision and language branches, synergistically improving the model's recognition of domain signals. By using the prompt to represent domain signals uniformly, if the model can be robust to DSPs in the source domain, it should be applicable to target domain, as they share the same DSD. By representing domain signals as prompts rather than instantiation features, DGPDL effectively reduces the reliance on specific domain appearances. This design enables the model to dynamically adapt to unseen target domains without the need for retraining. Extensive experiments show that the DGPDL is effective and outperforms the state-of-the-art methods on several cross-domain benchmarks.
Ajian Liu 0001, Xun Lin, Ruicong Zhi, Yanyan Liang 0001, Xinshan Zhu, Zhanchuan Cai, Jun Wan 0001, Sergio Escalera, Zhen Lei 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2026 Flexible Modal Mixture-of-Experts With Inter-Modal Knowledge Distillation for Face Anti-Spoofing
Hui Ma 0018, Ajian Liu 0001, Ning Li 0035, Boyun Wang, Hang Zou 0002, Yuan Zhang 0023, Jing Huang 0017, Zhiqiang Pu, Jun Wan 0001, Zhanchuan Cai, Zhen Lei 0001, Yanyan Liang 0001
IEEE Trans. Inf. Forensics Secur.2
2026 HySpeFAS: A Hyperspectral Face Anti-Spoofing Dataset Based on Snapshot Compressive Imaging
abstract
Face anti-spoofing, which aims to prevent the attacks of widely-used face recognition systems, is highly related to personal privacy and property security. However, existing benchmarks on face anti-spoofing mainly focus on RGB images, further challenged by consistently developed 3D high-fidelity (HiFi) masks. To facilitate the research on multimodal face anti-spoofing, we construct the HyperSpectral Face Anti-Spoofing (HySpeFAS) dataset. We introduce the newly-developed snapshot spectral imaging (SSI) technology to capture real and spoof faces, as well as identify unknown HiFi masks. Specifically, hyperspectral images (HSIs) acquired by SSI sensor contain rich information about the chemical composition of the targets, which can be used to effectively distinguish live human skin and various spoof materials. The HySpeFAS dataset contains 22,368 multimodal images (i.e., RGB, SSI, HSI) of 17 live subjects and 60 spoof subjects. Moreover, extensive experiments with baseline deep learning models validate the special features of the SSI images and the potential of SSI in FAS. By publishing the dataset as well as the baseline models, we encourage the community to foster the algorithm study associated with hyperspectral images and the development of SSI-equipped intelligent systems.
Shijie Rao, Yidong Huang, Xueqian Zhang, Ajian Liu 0001, Jun Wan 0001, Kaiyu Cui, Yali Li 0001
IEEE Trans. Inf. Forensics Secur.5
2026 CoGA: A Collaborative Gray-Box Adversarial Attack for Multimodal Language Models
abstract
Multimodal language models (LMs) have shown significant potential for applications across various domains but remain vulnerable to adversarial attacks. Current research in white-box or black-box settings generally struggles with unrealistic attack assumptions and limited efficacy of targeted attacks. This paper introduces CoGA, a novel gray-box collaborative adversarial attack method for multimodal LMs. Under our gray-box settings, attackers have access only to the victim model’s input encoders. With the guidance of different modalities, we perturb the embedding representations from encoders to disrupt the semantic alignment across modalities, ultimately causing inaccurate outputs on various downstream tasks. Specifically, we integrate text embeddings into the loss calculations of the image attack and utilize image embeddings to guide the ranking of vulnerable words and the selection of final samples. Extensive experiments demonstrate that our method achieves superior attack performance across diverse models and tasks, suggesting the shared vulnerability of multimodal LMs in confronting adversarial challenges. Our work provides new insights into the security of multimodal LMs, facilitating the deployment of more robust and secure models in practical applications.
Feng Lin 0004, Gaojian Wang, Tiantian Liu 0002, Zhibo Wang 0001, Weizhi Meng 0001, Ajian Liu 0001, Kui Ren 0001
IEEE Trans. Inf. Forensics Secur.7
2026 CPG-PAD: Concept-Informed Prompts Guided Presentation Attack Detection
Xiangyu Zhu 0001, Ajian Liu 0001, Siran Peng, Zhen Lei 0001
IEEE Trans. Inf. Forensics Secur.4
2026 Domain Generalization for Face Anti-Spoofing via Content-Aware Composite Prompt Engineering
abstract
The challenge of Domain Generalization (DG) in Face Anti-Spoofing (FAS) is the significant interference of domain-specific signals on subtle spoofing clues. Recently, some CLIP-based algorithms have been developed to alleviate this interference by adjusting the weights of visual classifiers. How-ever, our analysis of this class-wise prompt engineering suffers from two shortcomings for DG FAS: (1) The categories of facial categories, such as real or spoof, have no semantics for the CLIP model, making it difficult to learn accurate category descriptions. (2) A single form of prompt cannot portray the various types of spoofing. In this work, instead of class-wise prompts, we propose a novel Content-aware Composite Prompt Engineering (CCPE) that generates instance-wise composite prompts, including both fixed template and learnable prompts. Specifically, our CCPE constructs content-aware prompts from two branches: (1) Inherent content prompt explicitly benefits from abundant transferred knowledge from the instruction-based Large Language Model (LLM). (2) Learnable content prompts implicitly extract the most informative visual content via Q-Former. Moreover, we design a Cross-Modal Guidance Module (CGM) that dynamically adjusts unimodal features for fusion to achieve better generalized FAS. Finally, our CCPE has been validated for its effectiveness in multiple cross-domain experiments and achieves state-of-the-art (SOTA) results.
Jiabao Guo, Ajian Liu 0001, Yunfeng Diao, Hui Ma 0018, Bo Zhao 0023, Richang Hong, Meng Wang 0001
IEEE Trans. Multim.2
2025 Mixture-of-Attack-Experts with Class Regularization for Unified Physical-Digital Face Attack Detection
abstract
Unified detection of digital and physical attacks in facial recognition systems has become a focal point of research in recent years. However, current multi-modal methods typically ignore the intra-class and inter-class variability across different types of attacks, leading to degraded performance. To address this limitation, we propose MoAE-CR, a framework that effectively leverages class-aware information for improved attack detection. Our improvements manifest at two levels, i.e., the feature and loss level. At the feature level, we propose Mixture-of-Attack-Experts (MoAEs) to capture more subtle differences among various types of fake faces. At the loss level, we introduce Class Regularization (CR) through the Disentanglement Module (DM) and the Cluster Distillation Module (CDM). The DM enhances class separability by increasing the distance between the centers of live and fake face classes. However, center-to-center constraints alone are insufficient to ensure distinctive representations for individual features. Thus, we propose the CDM to further cluster features around their class centers while maintaining separation from other classes. Moreover, specific attacks that significantly deviate from common attack patterns are often overlooked. To address this issue, our distance calculation prioritizes more distant features. Extensive experiments on two unified physical-digital attack datasets demonstrate the state-of-the-art performance of the proposed method.
Shunxin Chen, Ajian Liu 0001, Junze Zheng, Jun Wan 0001, Kailai Peng, Sergio Escalera, Zhen Lei 0001
AAAI2
2025 Interpretable Face Anti-Spoofing: Enhancing Generalization with Multimodal Large Language Models
abstract
Face Anti-Spoofing (FAS) is essential for ensuring the security and reliability of facial recognition systems. Most existing FAS methods are formulated as binary classification tasks, providing confidence scores without interpretation. They exhibit limited generalization in out-of-domain scenarios, such as new environments or unseen spoofing types. In this work, we introduce a multimodal large language model (MLLM) framework for FAS, termed Interpretable Face Anti-Spoofing (I-FAS), which transforms the FAS task into an interpretable visual question answering (VQA) paradigm. Specifically, we propose a Spoof-aware Captioning and Filtering (SCF) strategy to generate high-quality captions for FAS images, enriching the model's supervision with natural language interpretations. To mitigate the impact of noisy captions during training, we develop a Lopsided Language Model (L-LM) loss function that separates loss calculations for judgment and interpretation, prioritizing the optimization of the former. Furthermore, to enhance the model's perception of global visual features, we design a Globally Aware Connector (GAC) to align multi-level visual representations with the language model. Extensive experiments on standard and newly devised One to Eleven cross-domain benchmarks, comprising 12 public datasets, demonstrate that our method significantly outperforms state-of-the-art methods.
Keyao Wang, Haixiao Yue, Ajian Liu 0001, Errui Ding, Jingdong Wang 0001
AAAI4
2025 Recover and Match: Open-Vocabulary Multi-Label Recognition through Knowledge-Constrained Optimal Transport
abstract
Identifying multiple novel classes in an image, known as open-vocabulary multi-label recognition, is a challenging task in computer vision. Recent studies explore the transfer of powerful vision-language models such as CLIP. However, these approaches face two critical challenges: (1) The local semantics of CLIP are disrupted due to its global pre-training objectives, resulting in unreliable regional predictions. (2) The matching property between image regions and candidate labels has been neglected, relying instead on naive feature aggregation such as average pooling, which leads to spurious predictions from irrelevant regions. In this paper, we present RAM (Recover And Match), a novel framework that effectively addresses the above issues. To tackle the first problem, we propose Ladder Local Adapter (LLA) to enforce refocusing on local regions, recovering local semantics in a memory-friendly way. For the second issue, we propose Knowledge-Constrained Optimal Transport (KCOT) to suppress meaningless matching to non-GT labels by formulating the task as an optimal transport problem. As a result, RAM achieves state-of-the-art performance on various datasets from three distinct domains, and shows great potential to boost the existing methods. Code: https://github.com/EricTan7/RAM.
Zichang Tan, Jun Li 0033, Ajian Liu 0001, Jun Wan 0001, Zhen Lei 0001
CVPR4
2025 Not All Frame Features are Equal: Video-to-4D Generation via Decoupling Dynamic-Static Features
abstract
Recently, the generation of dynamic 3D objects from a video has shown impressive results. Existing methods directly optimize Gaussians using whole information in frames. However, when dynamic regions are interwoven with static regions within frames, particularly if the static regions account for a large proportion, existing methods often overlook information in dynamic regions and are prone to overfitting on static regions. This leads to producing results with blurry textures. We consider that decoupling dynamic-static features to enhance dynamic representations can alleviate this issue. Thus, we propose a dynamic-static feature decoupling module (DSFD). Along temporal axes, it regards the regions of current frame features that possess significant differences relative to reference frame features as dynamic features. Conversely, the remaining parts are the static features. Then, we acquire decoupled features driven by dynamic features and current frame features. Moreover, to further enhance the dynamic representation of decoupled features from different viewpoints and ensure accurate motion prediction, we design a temporal-spatial similarity fusion module (TSSF). Along spatial axes, it adaptively selects similar information of dynamic regions. Hinging on the above, we construct a novel approach, DS4D. Experimental results verify our method achieves state-of-the-art (SOTA) results in video-to-4D. In addition, the experiments on a real-world scenario dataset demonstrate its effectiveness on the 4D scene. Our code will be publicly available.
Zhenwei Zhu, Ajian Liu 0001, Hui Ma 0018, Jian Nong, Yanyan Liang 0001
ICCV4
2025 TASAR: Transfer-based Attack on Skeletal Action Recognition
abstract
Skeletal sequence data, as a widely employed representation of human actions, are crucial in Human Activity Recognition (HAR). Recently, adversarial attacks have been proposed in this area, which exposes potential security concerns, and more importantly provides a good tool for model robustness test. Within this research, transfer-based attack is an important tool as it mimics the real-world scenario where an attacker has no knowledge of the target model, but is under-explored in Skeleton-based HAR (S-HAR). Consequently, existing S-HAR attacks exhibit weak adversarial transferability and the reason remains largely unknown. In this paper, we investigate this phenomenon via the characterization of the loss function. We find that one prominent indicator of poor transferability is the low smoothness of the loss function. Led by this observation, we improve the transferability by properly smoothening the loss when computing the adversarial examples. This leads to the first Transfer-based Attack on Skeletal Action Recognition, TASAR. TASAR explores the smoothened model posterior of pre-trained surrogates, which is achieved by a new post-train Dual Bayesian optimization strategy. Furthermore, unlike existing transfer-based methods which overlook the temporal coherence within sequences, TASAR incorporates motion dynamics into the Bayesian attack, effectively disrupting the spatial-temporal coherence of S-HARs. For exhaustive evaluation, we build the first large-scale robust S-HAR benchmark, comprising 7 S-HAR models, 10 attack methods, 3 S-HAR datasets and 2 defense models. Extensive results demonstrate the superiority of TASAR. Our benchmark enables easy comparisons for future studies, with the code available in the https://github.com/yunfengdiao/Skeleton-Robustness-Benchmark.
Yunfeng Diao, Baiqi Wu, Ajian Liu 0001, Xiaoshuai Hao, Meng Wang 0001, He Wang 0002
ICLR4
2025 SUEDE: Shared Unified Experts for Physical- Digital Face Attack Detection Enhancement
abstract
Face recognition systems are vulnerable to physical attacks (e.g., printed photos) and digital threats (e.g., DeepFake), which are currently being studied as independent visual tasks, such as Face Anti-Spoofing and Forgery Detection. The inherent differences among various attack types present significant challenges in identifying a common feature space, making it difficult to develop a unified framework for detecting data from both attack modalities simultaneously. Inspired by the efficacy of Mixture-of-Experts (MoE) in learning across diverse domains, we explore utilizing multiple experts to learn the distinct features of various attack types. However, the feature distributions of physical and digital attacks overlap and differ. This suggests that relying solely on distinct experts to learn the unique features of each attack type may overlook shared knowledge between them. To address these issues, we propose SUEDE, the Shared Unified Experts for Physical-Digital Face Attack Detection Enhancement. SUEDE combines a shared expert (always activated) to capture common features for both attack types and multiple routed experts (selectively activated) for specific attack types. Further, we integrate CLIP as the base network to ensure the shared expert benefits from prior visual knowledge and align visual-text representations in a unified space. Extensive results demonstrate SUEDE achieves superior performance compared to state-of-the-art unified detection methods.
Zuying Xie, Changtao Miao, Ajian Liu 0001, Jiabao Guo, Feng Li 0037, Dan Guo 0001, Yunfeng Diao
ICME3
2025 PESTalk: Speech-Driven 3D Facial Animation with Personalized Emotional Styles
Tianshun Han, Benjia Zhou, Ajian Liu 0001, Yanyan Liang 0001, Zhen Lei 0001, Jun Wan 0001
ACM Multimedia3
2025 Dynamic Analysis and Adaptive Discriminator for Fake News Detection
abstract
In current web environment, fake news spreads rapidly across online social networks, posing serious threats to society. Existing multimodal fake news detection methods can generally be classified into knowledge-based and semantic-based approaches. However, these methods are heavily rely on human expertise and feedback, lacking flexibility. To address this challenge, we propose a Dynamic Analysis and Adaptive Discriminator (DAAD) approach for fake news detection. For knowledge-based methods, we introduce the Monte Carlo Tree Search algorithm to leverage the self-reflective capabilities of large language models (LLMs) for prompt optimization, providing richer, domain-specific details and guidance to the LLMs, while enabling more flexible integration of LLM comment on news content. For semantic-based methods, we define four typical deceit patterns: emotional exaggeration, logical inconsistency, image manipulation, and semantic inconsistency, to reveal the mechanisms behind fake news creation. To detect these patterns, we carefully design four discriminators and expand them in depth and breadth, using the soft-routing mechanism to explore optimal detection models. Experimental results on three real-world datasets demonstrate the superiority of our approach.
Xinqi Su, Zitong Yu, Yawen Cui, Ajian Liu 0001, Xun Lin, Haochen Liang, Wenhui Li 0001, Li Shen 0008, Xiaochun Cao
ACM Multimedia4
2025 Reliable and Balanced Transfer Learning for Generalized Multimodal Face Anti-Spoofing
abstract
Face Anti-Spoofing (FAS) is essential for securing face recognition systems against presentation attacks. Recent advances in sensor technology and multimodal learning have enabled the development of multimodal FAS systems. However, existing methods often struggle to generalize to unseen attacks and diverse environments due to two key challenges: (1) Modality unreliability, where sensors such as depth and infrared suffer from severe domain shifts, impairing the reliability of cross-modal fusion; and (2) Modality imbalance, where over-reliance on a dominant modality weakens the model's robustness against attacks that affect other modalities. To overcome these issues, we propose MMDG++, a multimodal domain-generalized FAS framework built upon the vision-language model CLIP. In MMDG++, we design the Uncertainty-Guided Cross-Adapter++ (U-Adapter++) to filter out unreliable regions within each modality, enabling more reliable multimodal interactions. Additionally, we introduce Rebalanced Modality Gradient Modulation (ReGrad) for adaptive gradient modulation to balance modality convergence. To further enhance generalization, propose Asymmetric Domain Prompts (ADPs) that leverage CLIP's language priors to learn generalized decision boundaries across modalities. We also develop a novel multimodal FAS benchmark to evaluate generalizability under various deployment conditions. Extensive experiments across this benchmark show our method outperforms state-of-the-art FAS methods, demonstrating superior generalization capability.
Xun Lin, Ajian Liu 0001, Zitong Yu, Rizhao Cai, Shuai Wang 0049, Yi Yu 0011, Jun Wan 0001, Zhen Lei 0001, Xiaochun Cao, Alex Chichung Kot
IEEE Trans. Pattern Anal. Mach. Intell.2
2025 Visual Prompt Flexible-Modal Face Anti-Spoofing
abstract
Recently, vision transformer based multimodal learning methods have been proposed to improve the robustness of face anti-spoofing (FAS) systems. However, multimodal face data collected from the real world is often imperfect due to missing modalities from various imaging sensors. Recently, flexible-modal FAS (Yu et al. 2023) has attracted more attention, which aims to develop a unified multimodal FAS model using complete multimodal face data but is insensitive to test-time missing modalities. In this paper, we tackle one main challenge in flexible-modal FAS, i.e., when missing modality occurs either during training or testing in real-world situations. Inspired by the recent success of the prompt learning in language models, we proposeVisualPrompt flexible-modalFAS(VP-FAS), which learns the modal-relevant prompts to adapt the frozen pre-trained foundation model to downstream flexible-modal FAS task. Specifically, both vanilla visual prompts and residual contextual prompts are plugged into multimodal transformers to handle general missing-modality cases, while only requiring less than 4% learnable parameters compared to training the entire model. Furthermore, missing-modality regularization is proposed to force models to learn consistent multimodal feature embeddings when missing partial modalities. Extensive experiments conducted on two multimodal FAS benchmark datasets demonstrate the effectiveness of our VP-FAS framework that improves the performance under various missing-modality cases while alleviating the requirement of heavy model re-training.
Zitong Yu, Rizhao Cai, Yawen Cui, Ajian Liu 0001, Changsheng Chen 0001
IEEE Trans. Dependable Secur. Comput.4
2025 FA3-CLIP: Frequency-Aware Cues Fusion and Attack-Agnostic Prompt Learning for Unified Face Attack Detection
abstract
Facial recognition systems are vulnerable to physical (e.g., printed photos) and digital (e.g., DeepFake) face attacks. Existing methods struggle to simultaneously detect physical and digital attacks due to: 1) significant intra-class variations between these attack types, and 2) the inadequacy of spatial information alone to comprehensively capture live and fake cues. To address these issues, we propose a unified attack detection model termed Frequency-Aware and Attack-Agnostic CLIP (FA3-CLIP), which introduces attack-agnostic prompt learning to express generic live and fake cues derived from the fusion of spatial and frequency features, enabling unified detection of live faces and all categories of attacks. Specifically, the attack-agnostic prompt module generates generic live and fake prompts within the language branch to extract corresponding generic representations from both live and fake faces, guiding the model to learn a unified feature space for unified attack detection. Meanwhile, the module adaptively generates the live/fake conditional bias from the original spatial and frequency information to optimize the generic prompts accordingly, reducing the impact of intra-class variations. We further propose a dual-stream cues fusion framework in the vision branch, which leverages frequency information to complement subtle cues that are difficult to capture in the spatial domain. In addition, a frequency compression block is utilized in the frequency stream, which reduces redundancy in frequency features while preserving the diversity of crucial cues. We also establish new challenging protocols to facilitate unified face attack detection effectiveness. Experimental results on multiple benchmarks demonstrate that FA3-CLIP significantly improves performance, reducing ACER by over 1.2% on UniAttackData, and increasing AUC by more than 3% as well as reducing EER by over 4% on the JFSFDB dataset.
Yongze Li, Ning Li 0035, Ajian Liu 0001, Hui Ma 0018, Xihong Chen, Zhiyao Liang, Yanyan Liang 0001, Jun Wan 0001, Zhen Lei 0001
IEEE Trans. Inf. Forensics Secur.3
2025 Toward Generalized Iris Presentation Attack Detection: A Mask-and-Distill Mixture of Experts Approach
abstract
Iris Presentation Attack Detection (PAD) is critical for securing recognition systems, yet its practical deployment is severely hindered by the poor generalization of models across different acquisition devices and diverse datasets. To address this persistent cross-domain challenge, we first introduce a comprehensive evaluation framework, the Iris Presentation Attack Detection Cross-Domain-Testing (IPAD-CDT) Protocol, designed to evaluate the model robustness in these scenarios. Our core contribution is a novel Masked Mixture-of-Experts (MMoE) method, which enhances the generalization of Transformer-based architectures. MMoE introduces a structured information asymmetry, where "student" Experts learn robust features from masked inputs by distilling knowledge from an unmasked "teacher" Expert via a cosine distance loss. This mask-and-distill mechanism effectively mitigates overfitting and guides the model to learn domain-invariant cues. By integrating MMoE into a CLIP-based model, we conduct extensive experiments on our IPAD-CDT protocol. The results demonstrate that our method sets a new state-of-the-art, significantly outperforming existing models, especially in the challenging cross-dataset and cross-device settings.
Hang Zou 0002, Chenxi Du, Ajian Liu 0001, Yuan Zhang 0023, Jing Liu 0062, Jun Wan 0001, Hui Zhang 0061, Zhenan Sun
IEEE Trans. Inf. Forensics Secur.3
2025 Knowledge Distillation-Based Anomaly Detection via Adaptive Discrepancy Optimization
abstract
Knowledge distillation has emerged as a primary solution for anomaly detection, leveraging feature discrepancies between teacher–student (T–S) networks to locate anomalies. However, previous approaches suffer from ambiguous feature discrepancies, which hinder effective anomaly detection due to two main challenges: 1) overgeneralization, where the student network excessively mimics teacher features in anomalous regions, and 2) semantic bias between T–S networks in normal regions. To address these issues, we propose an Adaptive Discrepancy Optimization (Ado) block. The Ado block adaptively calibrates feature discrepancies by reducing overgeneralization in anomalous regions and selectively aligning semantic features in normal regions via learnable feature offsets. This versatile block can be seamlessly integrated into various distillation-based methods. Experimental results demonstrate that the Ado block significantly enhances performance across 11 different knowledge distillation frameworks on two widely used datasets. Notably, when integrated with the Ado block, RD4AD achieves a 22% relative improvement in pixel-level PRO on the VisA dataset. In addition, a real-world keyboard inspection application further validates the effectiveness of the Ado block.
Ning Li 0035, Ajian Liu 0001, Zhenwei Zhu, Xuxin Lin, Hui Ma 0018, Hongning Dai, Yanyan Liang 0001
IEEE Trans. Ind. Informatics2
2025 Vision Transformer With Relation Exploration for Pedestrian Attribute Recognition
abstract
Pedestrian attribute recognition has achieved high accuracy by exploring the relations between image regions and attributes. However, existing methods typically adopt features directly extracted from the backbone or utilize a single structure (e.g., transformer) to explore the relations, leading to inefficient and incomplete relation mining. To overcome these limitations, this paper proposes a comprehensive relationship framework called Vision Transformer with Relation Exploration (ViT-RE) for pedestrian attribute recognition, which includes two novel modules, namely Attribute and Contextual Feature Projection (ACFP) and Relation Exploration Module (REM). In ACFP, attribute-specific features and contextual-aware features are learned individually to capture discriminative information tailored for attributes and image regions, respectively. Then, REM employs Graph Convolutional Network (GCN) Blocks and Transformer Blocks to concurrently explore attribute, contextual, and attribute-contextual relations. To enable fine-grained relation mining, a Dynamic Adjacency Module (DAM) is further proposed to construct instance-wise adjacency matrix for the GCN Block. Equipped with comprehensive relation information, ViT-RE achieves promising performance on three popular benchmarks, including PETA, RAP, and PA-100 K datasets. Moreover, ViT-RE achieves the first place in theWACV 2023 UPAR Challenge.
Zichang Tan, Dunfang Weng, Ajian Liu 0001, Jun Wan 0001, Zhen Lei 0001, Stan Z. Li
IEEE Trans. Multim.4
2024 Multi-Domain Incremental Learning for Face Presentation Attack Detection
abstract
Previous face Presentation Attack Detection (PAD) methods aim to improve the effectiveness of cross-domain tasks. However, in real-world scenarios, the original training data of the pre-trained model is not available due to data privacy or other reasons. Under these constraints, general methods for fine-tuning single-target domain data may lose previously learned knowledge, leading to a catastrophic forgetting problem. To address these issues, we propose a multi-domain incremental learning (MDIL) method for PAD, which not only learns knowledge well from the new domain but also maintains the performance of previous domains stably. Specifically, we propose an adaptive domain-specific experts (ADE) framework based on the vision transformer to preserve the discriminability of previous domains. Furthermore, an asymmetric classifier is designed to keep the output distribution of different classifiers consistent, thereby improving the generalization ability. Extensive experiments show that our proposed method achieves state-of-the-art performance compared to prior methods of incremental learning. Excitingly, under more stringent setting conditions, our method approximates or even outperforms the DA/DG-based methods.
Keyao Wang, Haixiao Yue, Ajian Liu 0001, Haocheng Feng, Junyu Han, Errui Ding, Jingdong Wang 0001
AAAI4
2024 CFPL-FAS: Class Free Prompt Learning for Generalizable Face Anti-Spoofing
abstract
Domain generalization (DG) based Face Anti-Spoofing (FAS) aims to improve the model's performance on unseen domains. Existing methods either rely on domain labels to align domain-invariant feature spaces, or disentangle generalizable features from the whole sample, which inevitably lead to the distortion of semantic feature structures and achieve limited generalization. In this work, we make use of large-scale VLMs like CLIP and leverage the textual feature to dynamically adjust the classifier's weights for exploring generalizable visual features. Specifically, we propose a novel Class Free Prompt Learning (CFPL) paradigm for DG FAS, which utilizes two lightweight transformers, namely Content Q-Former (CQF) and Style Q-Former (SQF), to learn the different semantic prompts conditioned on content and style features by using a set of learnable query vectors, respectively. Thus, the generalizable prompt can be learned by two improvements: (1) A Prompt-Text Matched (PTM) supervision is introduced to ensure CQF learns visual representation that is most informative of the content description. (2) A Diversified Style Prompt (DSP) technology is proposed to diversify the learning of style prompts by mixing feature statistics between instance-specific styles. Finally, the learned text features modulate visual features to generalization through the designed Prompt Modulation (PM). Extensive experiments show that the CFPL is effective and outperforms the state-of-the-art methods on several cross-domain datasets.
Ajian Liu 0001, Jianwen Gan, Jun Wan 0001, Yanyan Liang 0001, Jiankang Deng, Sergio Escalera, Zhen Lei 0001
CVPR1
2024 VL-FAS: Domain Generalization via Vision-Language Model For Face Anti-Spoofing
abstract
Recent approaches have demonstrated the effectiveness of Vision Transformer (ViT) with attention mechanisms for domain generalization of Face Anti-Spoofing (FAS). However, current attention algorithms highlight all the salient objects (e.g., background objects, hair, glasses), which results in the feature learned by the model containing face-irrelevant noisy information. Inspired by existing Vision-language works, we propose the VL-FAS to extract more generalized and cleaner discriminative features. Specifically, we leverage fine-grained natural language descriptions of the face region to act as a task-oriented teacher, directing the model’s attention towards the face region through top-down attention regulation. Furthermore, to enhance the domain generalization ability of the model, we propose a Sample-Level Vision-Text optimization module (SLVT). SLVT uses sample-level image-text pairs for contrastive learning, allowing the visual coder to comprehend the intrinsic semantics of each image sample, thereby reducing the dependence on domain information. Extensive experiments show that our approach significantly outperforms the state-of-the-art and improves the performance of the ViT by about twice.
Ajian Liu 0001, Jun Wan 0001
ICASSP2
2024 CPL-CLIP: Compound Prompt Learning for Flexible-Modal Face Anti-Spoofing
abstract
Face anti-spoofing (FAS) is pivotal in safeguarding the integrity of face recognition systems. Flexible-modal FAS utilizes multi-modal data and trains a unified model adaptable to any single-modal testing scenario. This innovation addresses the shortcomings of conventional multi-modal FAS approaches, which typically demand separate model training and deployment for each modality. However, existing flexible-modal FAS approaches activate specific network branches based on the modality of the tested sample. This not only increases the model’s parameters but also necessitates the provision of the image’s modality for testing, thereby constraining deployment flexibility. To address the issue, we present Compound Prompt Learning CLIP (CPL-CLIP), a novel method for flexible-modal FAS. This approach capitalizes on a learned textual prompt that is nearly independent of modality, thus bolstering class-based classification across arbitrary modalities. Specifically, our CPL-CLIP introduces a Dual-Branch Prompt (DBP), consisting of class and modal prompts that describe and guide classification, where each prompt is composed of learnable vectors and fixed templates. To further render the class prompt as modality-agnostic as possible, a Cosine Similarity Loss (CSL) is proposed to facilitate the maximal separation of the class prompt from the modality prompt. With only the class prompt utilized during testing, CPL-CLIP enables deployment in diverse modal testing scenarios without the necessity of the test image’s modality to be known. Extensive experiments demonstrate CPL-CLIP’s superiority over existing methods on several flexible-modal FAS benchmarks.
Xiangyu Zhu 0001, Ajian Liu 0001, Xun Lin, Jun Wan 0001, Zhen Lei 0001
IJCB3
2024 La-SoftMoE CLIP for Unified Physical-Digital Face Attack Detection
abstract
Facial recognition systems are susceptible to both physical and digital attacks, posing significant security risks. Traditional approaches often treat these two attack types separately due to their distinct characteristics. Thus, when being combined attacked, almost all methods could not deal. Some studies attempt to combine the sparse data from both types of attacks into a single dataset and try to find a common feature space, which is often impractical due to the space is difficult to be found or even non-existent. To overcome these challenges, we propose a novel approach that uses the sparse model to handle sparse data, utilizing different parameter groups to process distinct regions of the sparse feature space. Specifically, we employ the Mixture of Experts (MoE) framework in our model, expert parameters are matched to tokens with varying weights during training and adaptively activated during testing. However, the traditional MoE struggles with the complex and irregular classification boundaries of this problem. Thus, we introduce a flexible self-adapting weighting mechanism, enabling the model to better fit and adapt. In this paper, we proposed La-SoftMoE CLIP, which allows for more flexible adaptation to the Unified Attack Detection (UAD) task, significantly enhancing the model’s capability to handle diversity attacks. Experiment results demonstrate that our proposed method has SOTA performance.
Hang Zou 0002, Chenxi Du, Hui Zhang 0061, Yuan Zhang 0023, Ajian Liu 0001, Jun Wan 0001, Zhen Lei 0001
IJCB5
2024 Unified Physical-Digital Face Attack Detection
Ajian Liu 0001, Haocheng Yuan, Junze Zheng, Dingheng Zeng, Jiankang Deng, Sergio Escalera, Xiaoming Liu 0002, Jun Wan 0001, Zhen Lei 0001
IJCAI2
2024 HideMIA: Hidden Wavelet Mining for Privacy-Enhancing Medical Image Analysis
Xun Lin, Yi Yu 0011, Zitong Yu, Ruohan Meng, Jiale Zhou 0001, Ajian Liu 0001, Yizhong Liu, Shuai Wang 0049, Wenzhong Tang, Zhen Lei 0001, Alex Chichung Kot
ACM Multimedia6
2024 FM-CLIP: Flexible Modal CLIP for Face Anti-Spoofing
abstract
In this work, borrowing a solution from the large-scale vision-language models (VLMs) instead of directly removing modality-specific signals from visual features, we propose a novel Flexible Modal CLIP (FM-CLIP) for flexible modal FAS, that can utilize text features to dynamically adjust visual features to be modality independent. In the visual branch, considering the huge visual differences of the same attack in different modalities, which makes it difficult for classifiers to flexibly identify subtle spoofing clues in different test modalities, we propose Cross-Modal Spoofing Enhancer (CMS-Enhancer). It includes a Frequency Extractor (FE) and Cross-Modal Interactor (CMI), aiming to map different modal attacks in a shared frequency space to reduce interference from modality-specific signals and enhance spoofing clues by leveraging cross-modal learning from the shared frequency space. In the text branch, we introduce a Language-Guided Patch Alignment (LGPA) based on prompt learning, which further guides the image encoder to focus on patch-level spoofing representations through dynamic weighting by text features. Thus, our FM-CLIP can flexibly test different modal samples by identifying and enhancing modality-agnostic spoofing cues. Finally, extensive experiments show that FM-CLIP is effective and outperforms state-of-the-art methods on multiple multi-modal datasets.
Ajian Liu 0001, Hui Ma 0018, Junze Zheng, Haocheng Yuan, Xiaoyuan Yu, Yanyan Liang 0001, Sergio Escalera, Jun Wan 0001, Zhen Lei 0001
ACM Multimedia1
2024 CA-MoEiT: Generalizable Face Anti-spoofing via Dual Cross-Attention and Semi-fixed Mixture-of-Expert
Ajian Liu 0001
Int. J. Comput. Vis.1
2024 Surveillance Face Anti-Spoofing
abstract
Face Anti-spoofing (FAS) is essential to secure face recognition systems from various physical attacks. However, recent research generally focuses on short-distance applications (i.e., phone unlocking) while lacking consideration of long-distance scenes (i.e., surveillance security checks). In order to promote relevant research and fill this gap in the community, we collect a large-scale Su rveillance Hi gh-Fi delity Mask (SuHiFiMask) dataset captured under 40 surveillance scenes, which has 101 subjects from different age groups with$232~3\text{D}$attacks (high-fidelity masks),$200~2\text{D}$attacks (posters, portraits, and screens), and 2 adversarial attacks. In this scene, low image resolution and noise interference are new challenges faced in surveillance FAS. Together with the SuHiFiMask dataset, we propose a Contrastive Quality-Invariance Learning (CQIL) network to alleviate the performance degradation caused by image quality from three aspects: 1) An Image Quality Variable module (IQV) is introduced to recover image information associated with discrimination by combining the super-resolution network. 2) Using generated sample pairs to simulate quality variance distributions to help contrastive learning strategies obtain robust feature representation under quality variation. 3) A Separate Quality Network (SQN) is designed to learn discriminative features independent of image quality. Finally, a large number of experiments verify the quality of the SuHiFiMask dataset and the superiority of the proposed CQIL.
Ajian Liu 0001, Jun Wan 0001, Sergio Escalera, Stan Z. Li, Zhen Lei 0001
IEEE Trans. Inf. Forensics Secur.2
2023 FM-ViT: Flexible Modal Vision Transformers for Face Anti-Spoofing
abstract
The availability of handy multi-modal (i.e., RGB-D) sensors has brought about a surge of face anti-spoofing research. However, the current multi-modal face presentation attack detection (PAD) has two defects: (1) The framework based on multi-modal fusion requires providing modalities consistent with the training input, which seriously limits the deployment scenario. (2) The performance of ConvNet-based model on high fidelity datasets is increasingly limited. In this work, we present a pure transformer-based framework, dubbed the Flexible Modal Vision Transformer (FM-ViT), for face anti-spoofing to flexibly target any single-modal (i.e., RGB) attack scenarios with the help of available multi-modal data. Specifically, FM-ViT retains a specific branch for each modality to capture different modal information and introduces the Cross-Modal Transformer Block (CMTB), which consists of two cascaded attentions named Multi-headed Mutual-Attention (MMA) and Fusion-Attention (MFA) to guide each modal branch to mine potential features from informative patch tokens, and to learn modality-agnostic liveness features by enriching the modal information of own CLS token, respectively. Experiments demonstrate that the single model trained based on FM-ViT can not only flexibly evaluate different modal samples, but also outperforms existing single-modal frameworks by a large margin, and approaches the multi-modal frameworks introduced with smaller FLOPs and model parameters.
Ajian Liu 0001, Zichang Tan, Zitong Yu, Jun Wan 0001, Yanyan Liang 0001, Zhen Lei 0001, Stan Z. Li, Guodong Guo
IEEE Trans. Inf. Forensics Secur.1
2022 Disentangling Facial Pose and Appearance Information for Face Anti-spoofing
abstract
Face Anti-spoofing aims to determine whether the captured face from a face recognition system is real or fake. However, the facial pose and local significant spoofing traces (i.e., the boundary and reflection spot in presentation attack instruments) seriously affects the performance and stability of the current algorithms. Due to they regard the face image as an indivisible unit, and process it holistically, rarely consider excluding these liveness-irrelated factors. Unlike it, we design a Pose-Independent Face Anti-Spoofing (PIFAS) framework to disentangle face into an appearance information and a pose code to capture liveness and liveness-irrelated features, respectively. Specifically, the PIFAS consists of an Unsupervised Pose Switching (UPS) module and a Mutual Information Averaged Defense (MIAD) module, which are used to control the facial pose and suppress the local significant attack traces by averaging the local and global knowledge. Extensive experimental evaluations on multiple face anti-spoofing datasets verify that the proposed method can improve the generalization and stabilize the performance of each testing video through alleviating the interference from liveness-irrelated factors.
Ajian Liu 0001, Jun Wan 0001, Yanyan Liang 0001
ICPR1
2022 MA-ViT: Modality-Agnostic Vision Transformers for Face Anti-Spoofing
abstract
The existing multi-modal face anti-spoofing (FAS) frameworks are designed based on two strategies: halfway and late fusion. However, the former requires test modalities consistent with the training input, which seriously limits its deployment scenarios. And the latter is built on multiple branches to process different modalities independently, which limits their use in applications with low memory or fast execution requirements. In this work, we present a single branch based Transformer framework, namely Modality-Agnostic Vision Transformer (MA-ViT), which aims to improve the performance of arbitrary modal attacks with the help of multi-modal data. Specifically, MA-ViT adopts the early fusion to aggregate all the available training modalities’ data and enables flexible testing of any given modal samples. Further, we develop the Modality-Agnostic Transformer Block (MATB) in MA-ViT, which consists of two stacked attentions named Modal-Disentangle Attention (MDA) and Cross-Modal Attention (CMA), to eliminate modality-related information for each modal sequences and supplement modality-agnostic liveness features from another modal sequences, respectively. Experiments demonstrate that the single model trained based on MA-ViT can not only flexibly evaluate different modal samples, but also outperforms existing single-modal frameworks by a large margin, and approaches the multi-modal frameworks introduced with smaller FLOPs and model parameters.
Ajian Liu 0001, Yanyan Liang 0001
IJCAI1
2022 Contrastive Context-Aware Learning for 3D High-Fidelity Mask Face Presentation Attack Detection
abstract
Face presentation attack detection (PAD) is essential to secure face recognition systems primarily from high-fidelity mask attacks. Most existing 3D mask PAD benchmarks suffer from several drawbacks: 1) a limited number of mask identities, types of sensors, and a total number of videos; 2) low-fidelity quality of facial masks. Basic deep models and remote photoplethysmography (rPPG) methods achieved acceptable performance on these benchmarks but still far from the needs of practical scenarios. To bridge the gap to real-world applications, we introduce a large-scale High-Fidelity Mask dataset, namely HiFiMask. Specifically, a total amount of 54,600 videos are recorded from 75 subjects with 225 realistic masks by 7 new kinds of sensors. Along with the dataset, we propose a novel Contrastive Context-aware Learning (CCL) framework. CCL is a new training methodology for supervised PAD tasks, which is able to learn by leveraging rich contexts accurately (e.g., subjects, mask material and lighting) among pairs of live faces and high-fidelity mask attacks. Extensive experimental evaluations on HiFiMask and three additional 3D mask datasets demonstrate the effectiveness of our method. The codes and dataset will be released soon.
Ajian Liu 0001, Zitong Yu, Jun Wan 0001, Anyang Su, Zichang Tan, Sergio Escalera, Junliang Xing, Yanyan Liang 0001, Guodong Guo, Zhen Lei 0001, Stan Z. Li
IEEE Trans. Inf. Forensics Secur.1
2022 Cross-Batch Hard Example Mining With Pseudo Large Batch for ID vs. Spot Face Recognition
abstract
In our daily life, a large number of activities require identity verification, e.g., ePassport gates. Most of those verification systems recognize who you are by matching the ID document photo (ID face) to your live face image (spot face). The ID vs. Spot (IvS) face recognition is different from general face recognition where each dataset usually contains a small number of subjects and sufficient images for each subject. In IvS face recognition, the datasets usually contain massive class numbers (million or more) while each class only has two image samples (one ID face and one spot face), which makes it very challenging to train an effective model (e.g., excessive demand on GPU memory if conducting the classification on such massive classes, hardly capture the effective features for bisample data of each identity, etc.). To avoid the excessive demand on GPU memory, a two-stage training method is developed, where we first train the model on the dataset in general face recognition (e.g., MS-Celeb-1M) and then employ the metric learning losses (e.g., triplet and quadruplet losses) to learn the features on IvS data with million classes. To extract more effective features for IvS face recognition, we propose two novel algorithms to enhance the network by selecting harder samples for training. Firstly, a Cross-Batch Hard Example Mining (CB-HEM) is proposed to select the hard triplets from not only the current mini-batch but also past dozens of mini-batches (for convenience, we use batch to denote a mini-batch in the following), which can significantly expand the space of sample selection. Secondly, a Pseudo Large Batch (PLB) is proposed to virtually increase the batch size with a fixed GPU memory. The proposed PLB and CB-HEM can be employed simultaneously to train the network, which dramatically expands the selecting space by hundreds of times, where the very hard sample pairs especially the hard negative pairs can be selected for training to enhance the discriminative capability. Extensive comparative evaluations conducted on multiple IvS benchmarks demonstrate the effectiveness of the proposed method.
Zichang Tan, Ajian Liu 0001, Jun Wan 0001, Hao Li 0030, Zhen Lei 0001, Guodong Guo, Stan Z. Li
IEEE Trans. Image Process.2
2021 CASIA-SURF CeFA: A Benchmark for Multi-modal Cross-ethnicity Face Anti-spoofing
abstract
The issue of ethnic bias has proven to affect the performance of face recognition in previous works, while it still remains to be vacant in face anti-spoofing. Therefore, in order to study the ethnic bias for face anti-spoofing, we introduce the largest CASIA-SURF Cross-ethnicity Face Anti-spoofing (CeFA) dataset, covering 3 ethnicities, 3 modalities, 1,607 subjects, and 2D plus 3D attack types. Five protocols are introduced to measure the affect under varied evaluation conditions, such as cross-ethnicity, unknown spoofs or both of them. As our knowledge, CASIA-SURF CeFA is the first dataset including explicit ethnic labels in current released datasets. Then, we propose a novel multi-modal fusion method as a strong baseline to alleviate the ethnic bias, which employs a partially shared fusion strategy to learn complementary information from multiple modalities. Extensive experiments have been conducted on the proposed dataset to verify its significance and generalization capability for other existing datasets, i.e., CASIA-SURF, OULU-NPU and SiW datasets. The dataset is available at https://sites.google.com/qq.com/face-anti-spoofing/welcome/challengecvpr2020?authuser=0.
Ajian Liu 0001, Zichang Tan, Jun Wan 0001, Sergio Escalera, Guodong Guo, Stan Z. Li
WACV1
2021 Face Anti-Spoofing via Adversarial Cross-Modality Translation
abstract
Face Presentation Attack Detection (PAD) approaches based on multi-modal data have been attracted increasingly by the research community. However, they require multi-modal face data consistently involved in both the training and testing phases. It would severely limit the applicability due to the most Face Anti-spoofing (FAS) systems are only equipped with Visible (VIS) imaging devices, i.e., RGB cameras. Therefore, how to use other modality (i.e., Near-Infrared (NIR)) to assist the performance improvement of VIS-based PAD is significant for FAS. In this work, we first discuss the big gap of performances among different modalities even though the same backbone network is applied. Then, we propose a novel Cross-modal Auxiliary (CMA) framework for the VIS-based FAS task. The main trait of CMA is that the performance can be greatly improved with the help of other modality while no other modality is required in the testing stage. The proposed CMA consists of a Modality Translation Network (MT-Net) and a Modality Assistance Network (MA-Net). The former aims to close the visible gap between different modalities via a generative model that maps inputs from one modality (i.e., RGB) to another (i.e., NIR). The latter focuses on how to use the translated modality (i.e., target modality) and RGB modality (i.e., source modality) together to train a discriminative PAD model. Extensive experiments are conducted to demonstrate that the proposed framework can push the state-of-the-art (SOTA) performances on both multi-modal datasets (i.e., CASIA-SURF, CeFA, and WMCA) and RGB-based datasets (i.e., OULU-NPU, and SiW).
Ajian Liu 0001, Zichang Tan, Jun Wan 0001, Yanyan Liang 0001, Zhen Lei 0001, Guodong Guo, Stan Z. Li
IEEE Trans. Inf. Forensics Secur.1
2020 3DPC-Net: 3D Point Cloud Network for Face Anti-spoofing
abstract
Face anti-spoofing plays a vital role in face recognition systems. Most deep learning-based methods directly use 2D images assisted with temporal information (i.e., motion, rPPG) or pseudo-3D information (i.e., Depth). The main drawback of the mentioned methods is that another extra network is needed to generate the depth/rPPG information to assist the backbone network for face anti-spoofing. Different from these methods, we propose a novel method named 3D Point Cloud Network (3DPC-Net). It is an encoder-decoder network that can predict the 3DPC maps to discriminate live faces from spoofing ones. The main traits of the proposed method are that: 1) It is the first time that 3DPC is used for face anti-spoofing; 2) 3DPC-Net is simple and effective and it only relies on 3DPC supervision. Extensive experiments on four databases (i.e., Oulu-NPU, SiW, CASIA-FASD, Replay Attack) have demonstrated that the 3DPC-Net is comparative to the state-of-the-art methods.
Jun Wan 0001, Yi Jin 0001, Ajian Liu 0001, Guodong Guo, Stan Z. Li
IJCB4
2019 A Dataset and Benchmark for Large-Scale Multi-Modal Face Anti-Spoofing
abstract
Face anti-spoofing is essential to prevent face recognition systems from a security breach. Much of the progresses have been made by the availability of face anti-spoofing benchmark datasets in recent years. However, existing face anti-spoofing benchmarks have limited number of subjects (≤170) and modalities (≤2), which hinder the further development of the academic community. To facilitate face anti-spoofing research, we introduce a large-scale multi-modal dataset, namely CASIA-SURF, which is the largest publicly available dataset for face anti-spoofing in terms of both subjects and visual modalities. Specifically, it consists of 1,000 subjects with 21,000 videos and each sample has 3 modalities (i.e., RGB, Depth and IR). We also provide a measurement set, evaluation protocol and training/validation/testing subsets, developing a new benchmark for face anti-spoofing. Moreover, we present a new multi-modal fusion method as baseline, which performs feature re-weighting to select the more informative channel features while suppressing the less useful ones for each modal. Extensive experiments have been conducted on the proposed dataset to verify its significance and generalization capability. The dataset is available at https://sites.google.com/qq.com/chalearnfacespoofingattackdete/.
Xiaobo Wang 0001, Ajian Liu 0001, Jun Wan 0001, Sergio Escalera, Hailin Shi, Stan Z. Li
CVPR3