Rizhao Cai

dblp:250/6021 · DBLP profile ↗
← Back
22ranked-venue papers
7as first author
21since 2021 · last 2025
0000-0002-7114-8462ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 9 · 3 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author · 9 since 2021Artificial intelligence and machine learning · 7 · 4 first-author · 7 since 2021
YearPublicationVenuePosition
2025 Forensicability Assessment: Not All Samples Qualify for Recapture Detection
abstract
Recapture detection is critical in forensic tasks, especially for authentication of face and document images in electronic Know Your Customer (e-KYC) processes. While deep learning has advanced face anti-spoofing (FAS) and document presentation attack detection (DPAD), weak forensic cues still hinder reliability. We propose the Forensicability Assessment Network (FANet) to assess sample forensicability and reject low-forensicability samples before recapture detection. This enhances the overall performance and reliability of e-KYC systems. FANet operates independently, without relying on real training data or actual forensic features, ensuring strong generalization across various scenarios. It combines image quality and forensic task cues, defining three forensicability classes based on domain knowledge. FANet is trained with cross-entropy loss, updating centers using a momentum-based approach. Experimental results demonstrate significant improvements in reliability by filtering out low-forensicability samples. This work introduces the first comprehensive approach to assessing forensicability, ensuring more reliable recapture detection in face and document images. The source codes are available at https://github.com/chenlewis/FANet.
Lin Zhao 0017, Rizhao Cai, Zitong Yu, Changsheng Chen 0001, Bin Li 0011
ICME3
2025 Towards Data-Centric Face Anti-spoofing: Improving Cross-Domain Generalization via Physics-Based Data Synthesis
Rizhao Cai, Cecelia Soh, Zitong Yu, Haoliang Li, Wenhan Yang, Alex Chichung Kot
Int. J. Comput. Vis.1
2025 Rehearsal-Free and Efficient Continual Learning for Cross-Domain Face Anti-Spoofing
abstract
Face Anti-Spoofing (FAS) is constantly challenged by new attack types and mediums, and thus it is crucial for a FAS model to not only mitigate Catastrophic Forgetting (CF) of previously learned spoofing knowledge on the training data during continual learning but also enhance the model's generalization ability to potential spoofing attacks. In this paper, we first highlight that current strategies for catastrophic forgetting are not well-suited to the imperceptible nature of spoofing information in FAS and lack the focus on improving generalization capability. Then, the instance-wise dynamic central difference convolutional adapter module with the weighted ensemble strategy for Vision Transformer (ViT) is proposed for efficiently fine-tuning with low-shot data by extracting generalized spoofing texture information. Furthermore, we find that catastrophic forgetting in FAS can be reflected through the inconsistent attention matrices of ViT between different continual sessions, as the attention matrices embody relationships of spoofing clues between different patch tokens. Hence, we introduce attention consistency regularization by learning and reusing attention matrices to alleviate catastrophic forgetting. Finally, we devise new protocols and conduct extensive experiments to validate the superior performance of alleviating catastrophic forgetting and generalization on unseen domains.
Rizhao Cai, Yawen Cui, Zitong Yu, Xun Lin, Changsheng Chen 0001, Alex Chichung Kot
IEEE Trans. Pattern Anal. Mach. Intell.1
2025 Reliable and Balanced Transfer Learning for Generalized Multimodal Face Anti-Spoofing
abstract
Face Anti-Spoofing (FAS) is essential for securing face recognition systems against presentation attacks. Recent advances in sensor technology and multimodal learning have enabled the development of multimodal FAS systems. However, existing methods often struggle to generalize to unseen attacks and diverse environments due to two key challenges: (1) Modality unreliability, where sensors such as depth and infrared suffer from severe domain shifts, impairing the reliability of cross-modal fusion; and (2) Modality imbalance, where over-reliance on a dominant modality weakens the model's robustness against attacks that affect other modalities. To overcome these issues, we propose MMDG++, a multimodal domain-generalized FAS framework built upon the vision-language model CLIP. In MMDG++, we design the Uncertainty-Guided Cross-Adapter++ (U-Adapter++) to filter out unreliable regions within each modality, enabling more reliable multimodal interactions. Additionally, we introduce Rebalanced Modality Gradient Modulation (ReGrad) for adaptive gradient modulation to balance modality convergence. To further enhance generalization, propose Asymmetric Domain Prompts (ADPs) that leverage CLIP's language priors to learn generalized decision boundaries across modalities. We also develop a novel multimodal FAS benchmark to evaluate generalizability under various deployment conditions. Extensive experiments across this benchmark show our method outperforms state-of-the-art FAS methods, demonstrating superior generalization capability.
Xun Lin, Ajian Liu 0001, Zitong Yu, Rizhao Cai, Shuai Wang 0049, Yi Yu 0011, Jun Wan 0001, Zhen Lei 0001, Xiaochun Cao, Alex Chichung Kot
IEEE Trans. Pattern Anal. Mach. Intell.4
2025 Forgery-Aware Adaptive Learning With Vision Transformer for Generalized Face Forgery Detection
abstract
With the rapid progress of generative models, the current challenge in face forgery detection is how to effectively detect realistic manipulated faces from different unseen domains. Though previous studies show that pre-trained Vision Transformer (ViT) based models can achieve some promising results after fully fine-tuning on the Deepfake dataset, their generalization performances are still unsatisfactory. To this end, we present a Forgery-aware Adaptive Vision Transformer (FA-ViT) under the adaptive learning paradigm for generalized face forgery detection, where the parameters in the pre-trained ViT are kept fixed while the designed adaptive modules are optimized to capture forgery features. Specifically, a global adaptive module is designed to model long-range interactions among input tokens, which takes advantage of self-attention mechanism to mine global forgery clues. To further explore essential local forgery clues, a local adaptive module is proposed to expose local inconsistencies by enhancing the local contextual association. In addition, we introduce a fine-grained adaptive learning module that emphasizes the common compact representation of genuine faces through relationship learning in fine-grained pairs, driving these proposed adaptive modules to be aware of fine-grained forgery-aware information. Extensive experiments demonstrate that our FA-ViT achieves state-of-the-arts results in the cross-dataset evaluation, and enhances the robustness against unseen perturbations. Particularly, FA-ViT achieves 93.83% and 78.32% AUC scores on Celeb-DF and DFDC datasets in the cross-dataset evaluation. The code and trained model have been released at:https://github.com/LoveSiameseCat/FAViT.
Anwei Luo, Rizhao Cai, Chenqi Kong, Yakun Ju, Xiangui Kang, Jiwu Huang, Alex Chichung Kot
IEEE Trans. Circuits Syst. Video Technol.2
2025 Visual Prompt Flexible-Modal Face Anti-Spoofing
abstract
Recently, vision transformer based multimodal learning methods have been proposed to improve the robustness of face anti-spoofing (FAS) systems. However, multimodal face data collected from the real world is often imperfect due to missing modalities from various imaging sensors. Recently, flexible-modal FAS (Yu et al. 2023) has attracted more attention, which aims to develop a unified multimodal FAS model using complete multimodal face data but is insensitive to test-time missing modalities. In this paper, we tackle one main challenge in flexible-modal FAS, i.e., when missing modality occurs either during training or testing in real-world situations. Inspired by the recent success of the prompt learning in language models, we proposeVisualPrompt flexible-modalFAS(VP-FAS), which learns the modal-relevant prompts to adapt the frozen pre-trained foundation model to downstream flexible-modal FAS task. Specifically, both vanilla visual prompts and residual contextual prompts are plugged into multimodal transformers to handle general missing-modality cases, while only requiring less than 4% learnable parameters compared to training the entire model. Furthermore, missing-modality regularization is proposed to force models to learn consistent multimodal feature embeddings when missing partial modalities. Extensive experiments conducted on two multimodal FAS benchmark datasets demonstrate the effectiveness of our VP-FAS framework that improves the performance under various missing-modality cases while alleviating the requirement of heavy model re-training.
Zitong Yu, Rizhao Cai, Yawen Cui, Ajian Liu 0001, Changsheng Chen 0001
IEEE Trans. Dependable Secur. Comput.2
2025 CMoA: Contrastive Mixture of Adapters for Generalized Few-Shot Continual Learning
abstract
The goal of Few-Shot Continual Learning (FSCL) is to incrementally learn novel tasks with limited labeled samples and preserve previous capabilities simultaneously. However, current FSCL works lack research on domain increment and domain generalization ability, which cannot cope with changes in the visual perception environment. In this paper, we set up a Generalized FSCL (GFSCL) protocol involving both class- and domain-incremental scenarios together with domain generalization assessment. Firstly, two benchmark datasets and protocols are newly arranged, and detailed baselines are provided for this unexplored configuration. Furthermore, we find that common continual learning methods have poor generalization ability on unseen domains and cannot better tackle catastrophic forgetting issue in cross-incremental tasks. Hence, we propose a rehearsal-free framework based on Vision Transformer (ViT) named Contrastive Mixture of Adapters (CMoA). It contains two non-conflicting parts: (1) By applying the fast-adaptation characteristic of adapter-embedded ViT, the mixture of Adapters (MoA) module is incorporated into ViT. For stability purpose, cosine similarity regularization and dynamic weighting are designed to make each adapter learn specific knowledge and concentrate on particular classes. (2) To further enhance domain generalization ability, we alleviate the intra-class variation by prototype-calibrated contrastive learning to improve domain-invariant representation learning. Finally, six evaluation indicators showing the overall performance and forgetting are compared by comprehensive experiments on two benchmark datasets to validate the efficacy of CMoA, and the results illustrate that CMoA can achieve comparative performance with rehearsal-based continual learning methods.
Yawen Cui, Jian Zhao 0006, Zitong Yu, Rizhao Cai, Lei Jin 0003, Alex Chichung Kot, Li Liu 0002, Xuelong Li 0001
IEEE Trans. Multim.4
2025 ContextualCoder: Adaptive In-Context Prompting for Programmatic Visual Question Answering
abstract
Visual Question Answering (VQA) presents a challenging task at the intersection of computer vision and natural language processing, aiming to bridge the semantic gap between visual perception and linguistic comprehension. Traditional VQA approaches do not distinguish between data processing and reasoning, limiting their interpretability and generalizability in complex and diverse scenarios. Conversely, Programmatic Visual Question Answering (PVQA) models leverage large language models (LLMs) to generate executable codes, providing answers with detailed and interpretable reasoning processes. However, existing PVQA models typically rely on simplistic input-output prompting, which struggles to elicit domain-specific knowledge from LLMs and often produces unclear or extraneous outputs. Furthermore, PVQA models typically rely on a basic in-context example (ICE) selection methodology that is heavily influenced by individual word similarity rather than the overall sentence context. This leads to suboptimal ICE selection and a reliance on dataset-specific ICE candidates. In this paper, we propose ContextualCoder, a novel prompting framework tailored for PVQA models. ContextualCoder leverages frozen LLMs for code generation and pre-trained visual models for code execution, eliminating the need for extensive training and enhancing model flexibility. By incorporating an innovative prompting methodology and a novel ICE selection strategy, ContextualCoder facilitates the use of diverse in-context information for code generation, thereby improving the performance of PVQA models. Our approach surpasses state-of-the-art models, as evidenced by comprehensive experiments across diverse VQA datasets, including multilingual scenarios.
Ruoyue Shen, Nakamasa Inoue, Dayan Guan, Rizhao Cai, Alex Chichung Kot, Koichi Shinoda
IEEE Trans. Multim.4
2024 Suppress and Rebalance: Towards Generalized Multi-Modal Face Anti-Spoofing
abstract
Face Anti-Spoofing (FAS) is crucial for securing face recognition systems against presentation attacks. With ad-vancements in sensor manufacture and multi-modal learning techniques, many multi-modal FAS approaches have emerged. However, they face challenges in generalizing to unseen attacks and deployment conditions. These chal-lenges arise from (1) modality unreliability, where some modality sensors like depth and infrared undergo signifi-cant domain shifts in varying environments, leading to the spread of unreliable information during cross-modal feature fusion, and (2) modality imbalance, where training overly relies on a dominant modality hinders the conver-gence of others, reducing effectiveness against attack types that are indistinguishable by sorely using the dominant modality. To address modality unreliability, we propose the Uncertainty-Guided Cross-Adapter (U-Adapter) to recognize unreliably detected regions within each modality and suppress the impact of unreliable regions on other modal-ities. For modality imbalance, we propose a Rebalanced Modality Gradient Modulation (ReGrad) strategy to rebal-ance the convergence speed of all modalities by adaptively adjusting their gradients. Besides, we provide the first large-scale benchmark for evaluating multi-modal FAS per-formance under domain generalization scenarios. Exten-sive experiments demonstrate that our method outperforms state-of-the-art methods. Source codes and protocols are released on https://github.com/OMGGGGG/mmdg.
Xun Lin, Shuai Wang 0049, Rizhao Cai, Yizhong Liu, Ying Fu 0001, Wenzhong Tang, Zitong Yu, Alex Chichung Kot
CVPR3
2024 BenchLMM: Benchmarking Cross-Style Visual Capability of Large Multimodal Models
Rizhao Cai, Zirui Song, Dayan Guan, Zhenhao Chen, Yaohang Li, Chenyu Yi, Alex Chichung Kot
ECCV (50)1
2024 Controllable and Gradual Facial Blemishes Retouching Via Physics-Based Modelling
abstract
Face retouching aims to remove facial blemishes, such as pigmentation and acne, and still retain fine-grain texture details. Nevertheless, existing methods just remove the blemishes but focus little on realism of the intermediate process, limiting their use more to beautifying facial images on social media rather than being effective tools for simulating changes in facial pigmentation and ance. Motivated by this limitation, we propose our Controllable and Gradual Face Retouching (CGFR). Our CGFR is based on physical modelling, adopting Sum-of-Gaussians to approximate skin subsurface scattering in a decomposed melanin and haemoglobin color space. Our CGFR offers a user-friendly control over the facial blemishes, achieving realistic and gradual blemishes retouching. Experimental results based on actual clinical data shows that CGFR can realistically simulate the blemishes’ gradual recovering process.
Chenhao Shuai, Rizhao Cai, Bandara Dissanayake, Amanda Newman, Dayan Guan, Dennis Sng, Ling Li 0012, Alex Chichung Kot
ICME2
2024 Rethinking Vision Transformer and Masked Autoencoder in Multimodal Face Anti-Spoofing
abstract
Abstract Recently, vision transformer (ViT) based multimodal learning methods have been proposed to improve the robustness of face anti-spoofing (FAS) systems. However, there are still no works to explore the fundamental natures (e.g., modality-aware inputs, suitable multimodal pre-training, and efficient finetuning) in vanilla ViT for multimodal FAS. In this paper, we investigate three key factors (i.e., inputs, pre-training, and finetuning) in ViT for multimodal FAS with RGB, Infrared (IR), and Depth. First, in terms of the ViT inputs, we find that leveraging local feature descriptors (such as histograms of oriented gradients) benefits the ViT on IR modality but not RGB or Depth modalities. Second, in consideration of the task (FAS vs. generic object classification) and modality (multimodal vs. unimodal) gaps, ImageNet pre-trained models might be sub-optimal for the multimodal FAS task. Finally, in observation of the inefficiency on direct finetuning the whole or partial ViT, we design an adaptive multimodal adapter (AMA), which can efficiently aggregate local multimodal features while freezing majority of ViT parameters. To bridge these gaps, we propose the modality-asymmetric masked autoencoder (M $$^{2}$$ 2 A $$^{2}$$ 2 E) for multimodal FAS self-supervised pre-training without costly annotated labels. Compared with the previous modality-symmetric autoencoder, the proposed M $$^{2}$$ 2 A $$^{2}$$ 2 E is able to learn more intrinsic task-aware representation and compatible with modality-agnostic (e.g., unimodal, bimodal, and trimodal) downstream settings. Extensive experiments with both unimodal (RGB, Depth, IR) and multimodal (RGB+Depth, RGB+IR, Depth+IR, RGB+Depth+IR) settings conducted on multimodal FAS benchmarks demonstrate the superior performance of the proposed methods. One highlight is that the proposed method is robust under various missing-modality cases where previous multimodal FAS models suffer serious performance drops. We hope these findings and solutions can facilitate the future research for ViT-based multimodal FAS.
Zitong Yu, Rizhao Cai, Yawen Cui, Xin Liu 0012, Yongjian Hu, Alex Chichung Kot
Int. J. Comput. Vis.2
2024 Benchmarking Joint Face Spoofing and Forgery Detection With Visual and Physiological Cues
abstract
Face anti-spoofing (FAS) and face forgery detection play vital roles in securing face biometric systems from presentation attacks (PAs) and vicious digital manipulation (e.g., deepfakes). Despite satisfactory performance upon large-scale data and powerful deep models, recent advances in face spoofing and forgery detection approaches usually focus on 1) unimodal visual appearance or physiological (i.e., remote photoplethysmography (rPPG)) cues; and 2) separated feature representation for FAS or face forgery detection. On one side, unimodal appearance and rPPG features are respectively vulnerable to high-fidelity face 3D mask and video replay attacks, inspiring us to design reliable multi-modal fusion mechanisms for generalized FAS. On the other side, there are rich common features across FAS and face forgery detection tasks (e.g., periodic rPPG rhythms and vanilla appearance for bonafides), providing solid evidence to design a joint FAS and face forgery detection system in a multi-task learning fashion. In this paper, we establish the first joint face spoofing and forgery detection benchmark using both visual appearance and physiological rPPG cues. To enhance the rPPG periodicity discrimination, we design a two-branch physiological network using both facial spatio-temporal rPPG signal map and its continuous wavelet transformed counterpart as inputs. To mitigate the modality bias and improve the fusion efficacy, we conduct a weighted batch and layer normalization for both appearance and rPPG features before multi-modal fusion. We also investigate prevalent deep models, feature fusion strategies and multi-task learning configurations for joint face spoofing and forgery detection. We find that the generalization capacities of both unimodal (appearance or rPPG) and multi-modal (appearance+rPPG) models can be obviously improved via joint training on these two tasks. We hope this new benchmark will facilitate the future research of both FAS and deepfake detection communities. The codes will be released athttps://github.com/ZitongYu/Benchmarking.
Zitong Yu, Rizhao Cai, Zhi Li 0054, Wenhan Yang, Jingang Shi, Alex Chichung Kot
IEEE Trans. Dependable Secur. Comput.2
2024 S-Adapter: Generalizing Vision Transformer for Face Anti-Spoofing With Statistical Tokens
abstract
Face Anti-Spoofing (FAS) aims to detect malicious attempts to invade a face recognition system by presenting spoofed faces. State-of-the-art FAS techniques predominantly rely on deep learning models but their cross-domain generalization capabilities are often hindered by the domain shift problem, which arises due to different distributions between training and testing data. In this study, we develop a generalized FAS method under the Efficient Parameter Transfer Learning (EPTL) paradigm, where we adapt the pre-trained Vision Transformer models for the FAS task. During training, the adapter modules are inserted into the pre-trained ViT model, and the adapters are updated while other pre-trained parameters remain fixed. We find the limitations of previous vanilla adapters in that they are based on linear layers, which lack a spoofing-aware inductive bias and thus restrict the cross-domain generalization. To address this limitation and achieve cross-domain generalized FAS, we propose a novel Statistical Adapter (S-Adapter) that gathers local discriminative and statistical information from localized token histograms. To further improve the generalization of the statistical tokens, we propose a novel Token Style Regularization (TSR), which aims to reduce domain style variance by regularizing Gram matrices extracted from tokens across different domains. Our experimental results demonstrate that our proposed S-Adapter and TSR provide significant benefits in both zero-shot and few-shot cross-domain testing, outperforming state-of-the-art methods on several benchmark tests. We will release the source code upon acceptance.
Rizhao Cai, Zitong Yu, Chenqi Kong, Haoliang Li, Changsheng Chen 0001, Yongjian Hu, Alex Chichung Kot
IEEE Trans. Inf. Forensics Secur.1
2024 Distortion Model-Based Spectral Augmentation for Generalized Recaptured Document Detection
abstract
Document recapturing is a presentation attack that covers the forensic traces in the digital domain. Document presentation attack detection (DPAD) is an important step in the document authentication pipeline. Existing DPAD methods suffer from low generalization performance under the cross-domain scenario with different types of documents. Data augmentation is a de facto technique to reduce the risk of overfitting the training data and improve the generalizability of a trained model. In this work, we improve the generalization performance of DPAD approaches by addressing two important limitations of the existing frequency domain augmentation (FDA) methods. First, contrary to the existing FDA methods that treat different spectral bands equally, we establish a band-of-interest localization (BOIL) method that locates the spectral band-of-interest (BOI) related to the recapturing operation by domain knowledge from the theoretical distortion models. Second, we propose a frequency-domain halftoning augmentation (FHAG) strategy that enhances the halftoning features in the BOI with considerations of different halftoning distortions. To evaluate the generalization performance of our FHAG with BOIL method on different types of document images, we have constructed a diverse recaptured document image dataset with 162 types of documents (RDID162), consisting of 5346 samples. The proposed method has been evaluated on the generic deep learning models and a state-of-the-art DPAD approach under both cross-device and cross-domain protocols for the DPAD task. Compared to the existing FDA methods, our method has improved the models with ResNet50 backbone by reducing more than 25% or 5 percentage points in EERs. The source code and data in this work is available athttps://github.com/chenlewis/FHAG-with-BOIL.
Changsheng Chen 0001, Bokang Li, Rizhao Cai, Jishen Zeng, Jiwu Huang
IEEE Trans. Inf. Forensics Secur.3
2024 Category-Conditional Gradient Alignment for Domain Adaptive Face Anti-Spoofing
abstract
In view of inconsistent face acquisition procedure in face anti-spoofing, the detection performance on the target domain generally suffers severe degradation under source-specific gradient optimization. Existing domain adaptation face anti-spoofing methods focus on improving model generalization capability through feature matching, which do not consider the gradient discrepancy between the source and target domains. To this end, this work develops a category-conditional gradient alignment guided face anti-spoofing algorithm (CCGA-FAS) from a novel perspective of gradient discrepancy elimination. Technically, the category-conditional gradient alignment mechanism maximizes the cosine similarity of the gradient vectors generated by source and target samples within the live and spoof categories separately, which promotes the source and target domains to follow similar gradient descent directions during optimization. Considering that the gradient vector generation and alignment is computationally dependent on reliable category information, a temporal knowledge and flexible threshold based dynamic category measurer is devised to provide pseudo category information for unlabelled target samples in an easy-to-hard manner. The optimization for CCGA-FAS is implemented under the teacher-student structure, where the student model serves as the gradient optimization backbone, and the category prediction simultaneously benefits from the teacher and student models to consolidate the alignment stability. Experimental results and analysis demonstrate that the proposed method outperforms the state-of-the-art methods in both unsupervised and K-shot semi-supervised domain adaptive face anti-spoofing scenarios.
Fei Peng 0001, Rizhao Cai, Zitong Yu, Min Long 0003, Kwok-Yan Lam
IEEE Trans. Inf. Forensics Secur.3
2023 Rehearsal-Free Domain Continual Face Anti-Spoofing: Generalize More and Forget Less
abstract
Face Anti-Spoofing (FAS) is recently studied under the continual learning setting, where the FAS models are expected to evolve after encountering data from new domains. However, existing methods need extra replay buffers to store previous data for rehearsal, which becomes infeasible when previous data is unavailable because of privacy issues. In this paper, we propose the first rehearsal-free method for Domain Continual Learning (DCL) of FAS, which deals with catastrophic forgetting and unseen domain generalization problems simultaneously. For better generalization to unseen domains, we design the Dynamic Central Difference Convolutional Adapter (DCDCA) to adapt Vision Transformer (ViT) models during the continual learning sessions. To alleviate the forgetting of previous domains without using previous data, we propose the Proxy Prototype Contrastive Regularization (PPCR) to constrain the continual learning with previous domain knowledge from the proxy prototypes. Simulating practical DCL scenarios, we devise two new protocols which evaluate both generalization and anti-forgetting performance. Extensive experimental results show that our proposed method can improve the generalization performance in unseen domains and alleviate the catastrophic forgetting of previous knowledge. The code and protocol files are released on https://github.com/RizhaoCai/DCL-FAS-ICCV2023.
Rizhao Cai, Yawen Cui, Zitong Yu, Haoliang Li, Yongjian Hu, Alex Chichung Kot
ICCV1
2022 Image Inpainting Detection via Enriched Attentive Pattern with Near Original Image Augmentation
abstract
As deep learning-based inpainting methods have achieved increasingly better results, its malicious use, e.g. removing objects to report fake news or to provide fake evidence, is becoming threatening. Previous works have provided rich discussions on network architectures, e.g. even performing Neural Architecture Search to obtain the optimal model architecture. However, there are rooms in other aspects. In our work, we provide comprehensive efforts from data and feature aspects. From the data aspect, as harder samples in the training data usually lead to stronger detection models, we propose near original image augmentation that pushes the inpainted images closer to the original ones (without distortion and inpainting) as the input images, which is proved to improve the detection accuracy. From the feature aspect, we propose to extract the attentive pattern. With the designed attentive pattern, the knowledge of different inpainting methods can be better exploited during the training phase. Finally, extensive experiments are conducted. In our evaluation, we consider the scenarios where the inpainting masks, which are used to generate the testing set, have a distribution gap from those masks used to produce the training set. Thus, the comparisons are conducted on a newly proposed dataset, where testing masks are inconsistent with the training ones. The experimental results show the superiority of the proposed method and the effectiveness of each component. All our codes and data will be online available.
Wenhan Yang, Rizhao Cai, Alex Chichung Kot
ACM Multimedia2
2022 Learning Meta Pattern for Face Anti-Spoofing
abstract
Face Anti-Spoofing (FAS) is essential to secure face recognition systems and has been extensively studied in recent years. Although deep neural networks (DNNs) for the FAS task have achieved promising results in intra-dataset experiments with similar distributions of training and testing data, the DNNs’ generalization ability is limited under the cross-domain scenarios with different distributions of training and testing data. To improve the generalization ability, recent hybrid methods have been explored to extract task-aware handcrafted features (e.g., Local Binary Pattern) as discriminative information for the input of DNNs. However, the handcrafted feature extraction relies on experts’ domain knowledge, and how to choose appropriate handcrafted features is underexplored. To this end, we propose a learnable network to extract Meta Pattern (MP) in our learning-to-learn framework. By replacing handcrafted features with the MP, the discriminative information from MP is capable of learning a more generalized model. Moreover, we devise a two-stream network to hierarchically fuse the input RGB image and the extracted MP by using our proposed Hierarchical Fusion Module (HFM). We conduct comprehensive experiments and show that our MP outperforms the compared handcrafted features. Also, our proposed method with HFM and the MP can achieve state-of-the-art performance on two different domain generalization evaluation benchmarks.
Rizhao Cai, Zhi Li 0054, Renjie Wan, Haoliang Li, Yongjian Hu, Alex Chichung Kot
IEEE Trans. Inf. Forensics Secur.1
2022 One-Class Knowledge Distillation for Face Presentation Attack Detection
abstract
Face presentation attack detection (PAD) has been extensively studied by research communities to enhance the security of face recognition systems. Although existing methods have achieved good performance on testing data with similar distribution as the training data, their performance degrades severely in application scenarios with data of unseen distributions. In situations where the training and testing data are drawn from different domains, a typical approach is to apply domain adaptation techniques to improve face PAD performance with the help of target domain data. However, it has always been a non-trivial challenge to collect sufficient data samples in the target domain, especially for attack samples. This paper introduces a teacher-student framework to improve the cross-domain performance of face PAD with one-class domain adaptation. In addition to the source domain data, the framework utilizes only a few genuine face samples of the target domain. Under this framework, a teacher network is trained with source domain samples to provide discriminative feature representations for face PAD. Student networks are trained to mimic the teacher network and learn similar representations for genuine face samples of the target domain. In the test phase, the similarity score between the representations of the teacher and student networks is used to distinguish attacks from genuine ones. To evaluate the proposed framework under one-class domain adaptation settings, we devised two new protocols and conducted extensive experiments. The experimental results show that our method outperforms baselines under one-class domain adaptation settings and even state-of-the-art methods with unsupervised domain adaptation.
Zhi Li 0054, Rizhao Cai, Haoliang Li, Kwok-Yan Lam, Yongjian Hu, Alex Chichung Kot
IEEE Trans. Inf. Forensics Secur.2
2021 DRL-FAS: A Novel Framework Based on Deep Reinforcement Learning for Face Anti-Spoofing
abstract
Inspired by the philosophy employed by human beings to determine whether a presented face example is genuine or not, i.e., to glance at the example globally first and then carefully observe the local regions to gain more discriminative information, for the face anti-spoofing problem, we propose a novel framework based on the Convolutional Neural Network (CNN) and the Recurrent Neural Network (RNN). In particular, we model the behavior of exploring face-spoofing-related information from image sub-patches by leveraging deep reinforcement learning. We further introduce a recurrent mechanism to learn representations of local information sequentially from the explored sub-patches with an RNN. Finally, for the classification purpose, we fuse the local information with the global one, which can be learned from the original input image through a CNN. Moreover, we conduct extensive experiments, including ablation study and visualization analysis, to evaluate our proposed framework on various public databases. The experiment results show that our method can generally achieve state-of-the-art performance among all scenarios, demonstrating its effectiveness.
Rizhao Cai, Haoliang Li, Shiqi Wang 0001, Changsheng Chen 0001, Alex Chichung Kot
IEEE Trans. Inf. Forensics Secur.1
2020 A Copy-Proof Scheme Based on the Spectral and Spatial Barcoding Channel Models
abstract
The traditional two-dimensional (2D) barcode has been employed in anti-counterfeiting systems as a storage media for serial numbers. However, an attack can be initiated by simply copying the 2D barcode and attaching it to a counterfeit product. In this paper, we aim at proposing an authentication scheme with a mobile imaging device for a 2D barcode. This work presents a competitive solution among the 2D barcode authentication schemes that have been verified under mobile imaging conditions. The proposed copy-proof scheme is composed of two sets of features which are extracted by exploiting the characteristics of barcoding channel models. The proposed features identify the intrinsic differences between genuine and counterfeit barcode images in the frequency and spatial domains. An efficient two-stage barcode authentication framework is then proposed by combining the two sets of features in a cascading manner. To evaluate the practicality of the proposed authentication scheme, four databases with different devices (printers, scanners, mobile cameras), barcode sizes, and barcode designs are considered in the experiments. By comparing with the existing texture descriptors and some deep learning-based approaches, it is shown that the proposed scheme has a higher authentication accuracy under various conditions, such as cross-database, cross-size and cross-pattern experiments which study the generalities of a pre-trained model towards challenging conditions commonly found in real-world scenarios. Last but not least, the proposed scheme has been evaluated under some state-of-the-art attack scenarios where the attacker employs several realizations of genuine patterns or the deep learning-based technique to produce a counterfeit copy. The source code and data for producing the results in our experiments are available at https://bit.ly/2FOlJH7.
Changsheng Chen 0001, Mulin Li, Anselmo Ferreira, Jiwu Huang, Rizhao Cai
IEEE Trans. Inf. Forensics Secur.5