VLDB 2026 Research / reviewers in the wild / expert
Changsheng Chen 0001
dblp:12/8842-1
· DBLP profile ↗
36ranked-venue papers
14as first author
28since 2021 · last 2026
0000-0003-4857-6810ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 18 · 6 first-author · 14 since 2021Security and privacy · 12 · 5 first-author · 10 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 4 since 2021Computer networks · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Generalized Document Tampering Localization via Color and Semantic DisentanglementabstractDocument images are vulnerable to tampering attacks from image editing tools and deep models. Therefore, the Document Tampering Localization (DTL) task has received increasing attention in recent years. However, given the wide variety of document types (e.g., contracts, certificates, ID cards), our analysis shows that existing DTL methods struggle with document images containing diverse background colors and varying semantic contents. Further analysis and experiments verify that the varying background color and semantic contents interfere with the forensic feature extraction process in the existing DTL methods. To address this issue, we propose two disentanglement modules to mitigate such interference and improve the ability of forgery trace detection. First, we design a Color Disentanglement (CD) module that applies disentangled learning representation to forensic features. The CD module, grounded in real-world prior knowledge, effectively decouples color information from forensic features, thereby improving robustness against varying background colors. Second, we propose the Semantic Disentanglement (SD) module, which performs image-level clustering on the tampering probability map during the inference process. The SD module focuses on tampering probabilities for each pixel, while discarding local semantic information (e.g., font, location, and shape). It leads to strong robustness against variations in document content. The evaluations demonstrate that our CD-SD method outperforms existing methods by 45.12% or 0.162 on the F1 metric in cross-dataset tests. Ablation studies show that the CD and SD modules improve the F1 score by 7.98% and 13.38%, respectively, across different backbones. Our method delivers consistent and stable improvements across various experimental protocols. Moreover, it is compatible with many DTL methods in a plug-and-play fashion. Shiqiang Zheng 0002, Changsheng Chen 0001, Shen Chen 0004, Taiping Yao, Shouhong Ding, Bin Li 0011, Jiwu Huang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | FALCON-Net: Feature Aggregation of Local Patterns for AI-Generated Image DetectionabstractWith the rapid development of generative models, the visual quality of generated images has become almost indistinguishable from real images, which poses a huge challenge to content authenticity verification. A key limitation of existing detectors is their reliance on model-specific cues, resulting in poor generalization to unseen models. Based on the observation of local differences in the generated images, we found that the generated images lack device-specific sensor noise and unnatural pixel intensity variations caused by the oversimplified generation process. These discrepancies provide important forensic cues for distinguishing between real and generated images. We propose the Feature Aggregation for Localized Context and Noise Network (FALCON-Net), which leverages these discrepancies to enhance detection capabilities. FALCON-Net integrates two complementary modules to enhance detection capabilities: the Intrinsic Noise Pattern Isolation (INP) module isolates device-specific noise patterns by analyzing high-frequency features in the frequency domain, while the Local Variation Pattern (LVP) module models the complex relationships between local pixels to capture directional intensity variations and reveal unnatural regularities in generated images. By combining these sensor-level and local structural cues, FALCON-Net identifies fundamental generative inconsistencies, ensuring robustness to post-processing and strong generalization to unseen models. Extensive experimental results show that FALCON-Net achieves the state-of-the-art performance in detecting generated images and shows good generalization ability to unseen generative models. The code is available at https://github.com/humiaomiaohaha/FALCON-Net. Dengyong Zhang, Changsheng Chen 0001, Jin Wang 0001, Yun Song, Gaobo Yang, Xin Liao 0001, Xiangling Ding |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2026 | DiffEraser: Generalized Text Erasure Based on Latent Diffusion PriorabstractText removal is an important task in processing both scene and document images. However, existing scene text removal (STR) methods are primarily focus on scene text images. The STR models (trained by scene text images) perform poorly on document images with dense, complex textured backgrounds. We discover that the limitations of existing methods can be attributed to the difficuties in background features estimation in the regions to be erased, which is based on the knowledge from neighboring regions in the input images and priors learned from the training data. The background features estimation performance degrades under the cross-domain scenarios, and compromises the quality of STR results. To address these issues, we introduce DiffEraser, a novel text removal framework that leverages prior knowledge from the Latent Diffusion Model (LDM) for removing text in both scene and document images. Our DiffEraser incorporates two key innovations to fully exploit the prior knowledge of LDM. First, we replace the conventional Variational Auto-Encoders (VAE) encoder with a Diffusion-Prior (DP) encoder, aiming to integrate the heterogeneous information from the LDM prior knowledge in latent space with the multi-level encoded features of the input image. Second, we introduce a Latent-Fusion (LF) decoder that integrates the heterogeneous features from both the LDM and DP encoders to generate high-quality text-erased results. To evaluate the generalization performance of our DiffEraser, we focus on the cross-domain protocols and construct a document image dataset, NPID295, which contains 295 types of passports and identity cards. Notably, when trained on a scene text dataset, DiffEraser significantly outperforms existing STR methods in the challenging NPID295 dataset. The resources of this work will be available online upon acceptance. Zhihao Chen 0011, Changsheng Chen 0001, Shunquan Tan, Jiwu Huang |
IEEE Trans. Image Process. | 3 |
| 2025 | Unmask Tampering: Efficient Document Tampering Localization under Recapturing Attacks with Real Distortion Knowledge
Changsheng Chen 0001, Yinyin Lin, Bin Li 0011, Jiwu Huang |
CCS | 1 |
| 2025 | Forensicability Assessment: Not All Samples Qualify for Recapture DetectionabstractRecapture detection is critical in forensic tasks, especially for authentication of face and document images in electronic Know Your Customer (e-KYC) processes. While deep learning has advanced face anti-spoofing (FAS) and document presentation attack detection (DPAD), weak forensic cues still hinder reliability. We propose the Forensicability Assessment Network (FANet) to assess sample forensicability and reject low-forensicability samples before recapture detection. This enhances the overall performance and reliability of e-KYC systems. FANet operates independently, without relying on real training data or actual forensic features, ensuring strong generalization across various scenarios. It combines image quality and forensic task cues, defining three forensicability classes based on domain knowledge. FANet is trained with cross-entropy loss, updating centers using a momentum-based approach. Experimental results demonstrate significant improvements in reliability by filtering out low-forensicability samples. This work introduces the first comprehensive approach to assessing forensicability, ensuring more reliable recapture detection in face and document images. The source codes are available at https://github.com/chenlewis/FANet. Lin Zhao 0017, Rizhao Cai, Zitong Yu, Changsheng Chen 0001, Bin Li 0011 |
ICME | 5 |
| 2025 | DITL2: Dual-Stage Invariance Transfer Learning for Generalizable Document Image Tampering LocalizationabstractDocument Image Tampering Localization (DITL) advances considerably, yet achieving robust cross-dataset generalization remains a formidable challenge for practical applications. Expanding existing document datasets for training is labor-intensive, making it appealing to incorporate data from non-document domains such as natural scene images. However, domain-specific variations, including differences in color distribution and texture, compromise the performance of joint training. To address this issue, we propose DITL2, a Dual-stage Invariance Transfer Learning framework for Document Image Tampering Localization that consists of Cross-Domain Invariance Pre-training (CDIP) and Frequency Decoupling Parameter Adaptation (FDPA). In the pre-training stage, CDIP employs style transfer and texture consistency learning to suppress domain-specific influences from tampered natural scene images, and tampering trace commonality learning to acquire domain-invariant features. In the fine-tuning stage, FDPA adapts the parameters of the pre-trained model, leveraging the general knowledge from the pre-trained model to address DITL tasks while reducing the risk of overfitting. Experiments show that this approach effectively leverages external data resources to boost model performance, achieving state-of-the-art results across a variety of cross-dataset settings. Shen Chen 0004, Bin Li 0011, Kaiqing Lin, Changsheng Chen 0001, Haodong Li 0001, Taiping Yao, Shouhong Ding |
ACM Multimedia | 6 |
| 2025 | Rehearsal-Free and Efficient Continual Learning for Cross-Domain Face Anti-SpoofingabstractFace Anti-Spoofing (FAS) is constantly challenged by new attack types and mediums, and thus it is crucial for a FAS model to not only mitigate Catastrophic Forgetting (CF) of previously learned spoofing knowledge on the training data during continual learning but also enhance the model's generalization ability to potential spoofing attacks. In this paper, we first highlight that current strategies for catastrophic forgetting are not well-suited to the imperceptible nature of spoofing information in FAS and lack the focus on improving generalization capability. Then, the instance-wise dynamic central difference convolutional adapter module with the weighted ensemble strategy for Vision Transformer (ViT) is proposed for efficiently fine-tuning with low-shot data by extracting generalized spoofing texture information. Furthermore, we find that catastrophic forgetting in FAS can be reflected through the inconsistent attention matrices of ViT between different continual sessions, as the attention matrices embody relationships of spoofing clues between different patch tokens. Hence, we introduce attention consistency regularization by learning and reusing attention matrices to alleviate catastrophic forgetting. Finally, we devise new protocols and conduct extensive experiments to validate the superior performance of alleviating catastrophic forgetting and generalization on unseen domains. Rizhao Cai, Yawen Cui, Zitong Yu, Xun Lin, Changsheng Chen 0001, Alex Chichung Kot |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | Moiré Spectral Augmentation and Masked Frequency Modeling for Document Presentation Attack DetectionabstractDocument Presentation Attack is an anti-forensic operation that conceals the forgery traces of image manipulation in the digital domain. Existing document presentation attack detection (DPAD) methods show unsatisfactory performance under samples with different contents and qualities. In this work, we focus on the DPAD task on screen-recapturing channel and exploit the prior knowledge of distortion (i.e., moire pattern) in the spectral domain to address these limitations. We propose a frequency-domain moir ´ e´ augmentation (FMAG) strategy that enhances the spectral components contributed to the moire distortion, improving the generalization ´ performance under different document contents. We devise the mask moire frequency modeling (M ´ 2FM) scheme to reconstruct the moire-related spectral components in low-quality samples under the guidance of the spectral distortion model and a pre-trained DPAD ´ classifier. To evaluate the generalization performance, we collect the diverse Screen Recaptured Document Image Dataset with 162 different document contents (SRDID162) consisting of 162 genuine document images, as well as 2592 low and high-quality recaptured document images, respectively. Our experimental protocol involves training with high-quality ID images and testing with SRDID162 dataset of diverse contents and image qualities. Compared to a SOTA data augmentation approach for recaptured natural images, our FMAG & M2FM approach achieves a significant improvement of 49.15% or 22.50 percentage points in average EER on the generic deep learning backbones. The data and code of this work will be available at Github Changsheng Chen 0001, Youjie Li, Bokang Li, Weifan Yu, Baoying Chen, Bin Li 0011, Jiwu Huang |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2025 | Visual Prompt Flexible-Modal Face Anti-SpoofingabstractRecently, vision transformer based multimodal learning methods have been proposed to improve the robustness of face anti-spoofing (FAS) systems. However, multimodal face data collected from the real world is often imperfect due to missing modalities from various imaging sensors. Recently, flexible-modal FAS (Yu et al. 2023) has attracted more attention, which aims to develop a unified multimodal FAS model using complete multimodal face data but is insensitive to test-time missing modalities. In this paper, we tackle one main challenge in flexible-modal FAS, i.e., when missing modality occurs either during training or testing in real-world situations. Inspired by the recent success of the prompt learning in language models, we proposeVisualPrompt flexible-modalFAS(VP-FAS), which learns the modal-relevant prompts to adapt the frozen pre-trained foundation model to downstream flexible-modal FAS task. Specifically, both vanilla visual prompts and residual contextual prompts are plugged into multimodal transformers to handle general missing-modality cases, while only requiring less than 4% learnable parameters compared to training the entire model. Furthermore, missing-modality regularization is proposed to force models to learn consistent multimodal feature embeddings when missing partial modalities. Extensive experiments conducted on two multimodal FAS benchmark datasets demonstrate the effectiveness of our VP-FAS framework that improves the performance under various missing-modality cases while alleviating the requirement of heavy model re-training. Zitong Yu, Rizhao Cai, Yawen Cui, Ajian Liu 0001, Changsheng Chen 0001 |
IEEE Trans. Dependable Secur. Comput. | 5 |
| 2025 | Image Copy-Move Forgery Detection via Deep PatchMatch and Pairwise Ranking LearningabstractRecent advances in deep learning algorithms have shown impressive progress in image copy-move forgery detection (CMFD). However, these algorithms lack generalizability in practical scenarios where the copied regions are not present in the training images, or the cloned regions are part of the background. Additionally, these algorithms utilize convolution operations to distinguish source and target regions, leading to unsatisfactory results when the target regions blend well with the background. To address these limitations, this study proposes a novel end-to-end CMFD framework that integrates the strengths of conventional and deep learning methods. Specifically, the study develops a deep cross-scale PatchMatch (PM) method that is customized for CMFD to locate copy-move regions. Unlike existing deep models, our approach utilizes features extracted from high-resolution scales to seek explicit and reliable point-to-point matching between source and target regions. Furthermore, we propose a novel pairwise rank learning framework to separate source and target regions. By leveraging the strong prior of point-to-point matches, the framework can identify subtle differences and effectively discriminate between source and target regions, even when the target regions blend well with the background. Our framework is fully differentiable and can be trained end-to-end. Comprehensive experimental results highlight the remarkable generalizability of our scheme across various copy-move scenarios, significantly outperforming existing methods. Yuanman Li, Yingjie He 0003, Changsheng Chen 0001, Li Dong 0006, Bin Li 0011, Jiantao Zhou 0001, Xia Li 0006 |
IEEE Trans. Image Process. | 3 |
| 2024 | CMA: A Chromaticity Map Adapter for Robust Detection of Screen-Recapture Document ImagesabstractThe rebroadcasting of screen-recaptured document images introduces a significant risk to the confidential docu-ments processed in government departments and commer-cial companies. However, detecting recaptured document images subjected to distortions from online social networks (OSNs) is challenging since the common forensics cues, such as moiré pattern, are weakened during transmission. In this work, we first devise a pixel-level distortion model of the screen-recaptured document image to identify the robust features of color artifacts. Then, we extract a chromaticity map from the recaptured image to highlight the presence of color artifacts even under low-quality samples. Based on the prior understanding, we design a chromaticity map adapter (CMA) to efficiently extract the chromaticity map, and feed it into the transformer backbone as multi-modal prompt tokens. To evaluate the performance of the pro-posed method, we collect a recaptured office document im-age dataset with over 10K diverse samples. Experimental results demonstrate that the proposed CMA method outper-forms a SOTA approach (with RGB modality only), reducing the average EER from 26.82% to 16.78%. Robustness eval-uation shows that our method achieves 0.8688 and 0.7554 AUCs under samples with JPEG compression$(QF=70)$and resolution as low as$534\times 503$pixels. Changsheng Chen 0001, Liangwei Lin, Bin Li 0011, Jishen Zeng, Jiwu Huang |
CVPR | 1 |
| 2024 | Robust Document Presentation Attack Detection via Diffusion Models and Knowledge Distillation
Bokang Li, Changsheng Chen 0001 |
PRCV (11) | 2 |
| 2024 | A distortion model guided adversarial surrogate for recaptured document detection
Changsheng Chen 0001, Xijin Li, Baoying Chen, Haodong Li 0001 |
Pattern Recognit. | 1 |
| 2024 | S-Adapter: Generalizing Vision Transformer for Face Anti-Spoofing With Statistical TokensabstractFace Anti-Spoofing (FAS) aims to detect malicious attempts to invade a face recognition system by presenting spoofed faces. State-of-the-art FAS techniques predominantly rely on deep learning models but their cross-domain generalization capabilities are often hindered by the domain shift problem, which arises due to different distributions between training and testing data. In this study, we develop a generalized FAS method under the Efficient Parameter Transfer Learning (EPTL) paradigm, where we adapt the pre-trained Vision Transformer models for the FAS task. During training, the adapter modules are inserted into the pre-trained ViT model, and the adapters are updated while other pre-trained parameters remain fixed. We find the limitations of previous vanilla adapters in that they are based on linear layers, which lack a spoofing-aware inductive bias and thus restrict the cross-domain generalization. To address this limitation and achieve cross-domain generalized FAS, we propose a novel Statistical Adapter (S-Adapter) that gathers local discriminative and statistical information from localized token histograms. To further improve the generalization of the statistical tokens, we propose a novel Token Style Regularization (TSR), which aims to reduce domain style variance by regularizing Gram matrices extracted from tokens across different domains. Our experimental results demonstrate that our proposed S-Adapter and TSR provide significant benefits in both zero-shot and few-shot cross-domain testing, outperforming state-of-the-art methods on several benchmark tests. We will release the source code upon acceptance. Rizhao Cai, Zitong Yu, Chenqi Kong, Haoliang Li, Changsheng Chen 0001, Yongjian Hu, Alex Chichung Kot |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2024 | Distortion Model-Based Spectral Augmentation for Generalized Recaptured Document DetectionabstractDocument recapturing is a presentation attack that covers the forensic traces in the digital domain. Document presentation attack detection (DPAD) is an important step in the document authentication pipeline. Existing DPAD methods suffer from low generalization performance under the cross-domain scenario with different types of documents. Data augmentation is a de facto technique to reduce the risk of overfitting the training data and improve the generalizability of a trained model. In this work, we improve the generalization performance of DPAD approaches by addressing two important limitations of the existing frequency domain augmentation (FDA) methods. First, contrary to the existing FDA methods that treat different spectral bands equally, we establish a band-of-interest localization (BOIL) method that locates the spectral band-of-interest (BOI) related to the recapturing operation by domain knowledge from the theoretical distortion models. Second, we propose a frequency-domain halftoning augmentation (FHAG) strategy that enhances the halftoning features in the BOI with considerations of different halftoning distortions. To evaluate the generalization performance of our FHAG with BOIL method on different types of document images, we have constructed a diverse recaptured document image dataset with 162 types of documents (RDID162), consisting of 5346 samples. The proposed method has been evaluated on the generic deep learning models and a state-of-the-art DPAD approach under both cross-device and cross-domain protocols for the DPAD task. Compared to the existing FDA methods, our method has improved the models with ResNet50 backbone by reducing more than 25% or 5 percentage points in EERs. The source code and data in this work is available athttps://github.com/chenlewis/FHAG-with-BOIL. Changsheng Chen 0001, Bokang Li, Rizhao Cai, Jishen Zeng, Jiwu Huang |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2023 | Image Copy-Move Forgery Detection via Deep Cross-Scale PatchMatchabstractThe recently developed deep algorithms achieve promising progress in the field of image copy-move forgery detection (CMFD). However, they have limited generalizability in some practical scenarios, where the copy-move objects may not appear in the training images or cloned regions are from the background. To address the above issues, in this work, we propose a novel end-to-end CMFD framework by integrating merits from both conventional and deep methods. Specifically, we design a deep cross-scale patchmatch method tailored for CMFD to localize copy-move regions. In contrast to existing deep models, our scheme aims to seek explicit and reliable point-to-point matching between source and target regions using features extracted from high-resolution scales. Further, we develop a manipulation region location branch for source/target separation. The proposed CMFD framework is completely differentiable and can be trained in an end-to-end manner. Extensive experimental results demonstrate the high generalizability of our method to different copy-move contents, and the proposed scheme achieves significantly better performance than existing approaches. Yingjie He 0003, Yuanman Li, Changsheng Chen 0001, Xia Li 0006 |
ICME | 3 |
| 2023 | Reading Multilevel 2-D Barcodes Using a Machine Learning ApproachabstractThis article addresses the reading problem of multilevel 2-D barcodes over a print-and-capture (PC) channel. The prior reading schemes have different limitations to hinder their applications, e.g., suffering from quantization error, being sensitive to the predetermined decision boundaries, and being sensitive to the selection of initial parameters. In this article, we introduce a machine learning approach to address the above limitations using a new ensemble clustering (EC) algorithm. Based on the new EC algorithm, we propose two reading schemes of a multilevel 2-D barcode. Specifically, the first proposed scheme is named the EC reading scheme. In the EC reading scheme, we introduce a weighted ensemble mechanism to assign different weights to different base clustering results. Then, we propose the second scheme, named the enhanced EC (EEC) reading scheme, to further improve the reading performance with the help of the reference symbols. We implement our approach and conduct extensive performance comparisons through an actual excremental platform under various multilevel 2-D barcodes and various capturing devices. From experimental results, we observe that both proposed reading schemes have better performance than the prior reading schemes. Moreover, the EEC reading scheme has better performance than the EC reading scheme, and their performance gap becomes more apparent as the distortion of a PC channel increases. Jiaheng Zhang, Le Ou-Yang, Changsheng Chen 0001, Ning Xie 0007 |
IEEE Internet Things J. | 4 |
| 2022 | Learning General Gaussian Mixture Model with Integral Cosine SimilarityabstractGaussian mixture model (GMM) is a powerful statistical tool in data modeling, especially for unsupervised learning tasks. Traditional learning methods for GMM such as expectation maximization (EM) require the covariance of the Gaussian components to be non-singular, a condition that is often not satisfied in real-world applications. This paper presents a new learning method called G$^2$M$^2$ (General Gaussian Mixture Model) by fitting an unnormalized Gaussian mixture function (UGMF) to a data distribution. At the core of G$^2$M$^2$ is the introduction of an integral cosine similarity (ICS) function for comparing the UGMF and the unknown data density distribution without having to explicitly estimate it. By maximizing the ICS through Monte Carlo sampling, the UGMF can be made to overlap with the unknown data density distribution such that the two only differ by a constant scalar, and the UGMF can be normalized to obtain the data density distribution. A Siamese convolutional neural network is also designed for optimizing the ICS function. Experimental results show that our method is more competitive in modeling data having correlations that may lead to singular covariance matrices in GMM, and it outperforms state-of-the-art methods in unsupervised anomaly detection. Guanglin Li 0001, Bin Li 0011, Changsheng Chen 0001, Shunquan Tan, Guoping Qiu |
IJCAI | 3 |
| 2022 | Regional People Counting Approach Based on Multi-Motion States Models Using MIMO RadarabstractThe existing radar-based regional people counting methods usually perform a single prediction model to classify the number of people. However, signals of different motion states of the target vary greatly, which results in some features aliasing among different numbers of people. It is difficult to find an optimal single prediction model to adapt the different motion states of the target, especially in the case of limited training data. To address this problem, we propose a regional people counting approach based on multi-motion states models using multi-input multioutput (MIMO) radar. We first use the state classification model with the 2-D spatial attribute features to classify the motion state of the target and then cascade the corresponding motion state model to classify the number of people. In particular, we incorporate the mixed state into the proposed approach to deal with the situations of mixed motion states of targets. The experimental results show that the proposed approach can be implemented in an embedded operating system, and its average classification accuracy of zero person, one person, and multiple persons is above 97.45% and much better than that of the existing regional people counting methods in the metal ramp with strong static clutter and strong multipath, which is a small room and too small for many people. Xiaoze Huang, Zhaocheng Yang, Changsheng Chen 0001 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2022 | A distortion model-based pre-screening method for document image tampering localization under recapturing attack
Changsheng Chen 0001, Lin Zhao 0017, Jiabin Yan, Haodong Li 0001 |
Signal Process. | 1 |
| 2022 | Detection of Information Hiding at Anti-Copying 2D BarcodesabstractThis paper addresses the issue of detecting the use of information hiding at anti-copying 2D barcodes. Prior hidden information detection schemes have their roots in either heuristic-based or Machine Learning (ML). However, prior heuristics-based schemes lack a rigorous theoretical analysis. Prior ML-based information schemes lack robustness because a printed 2D barcode is very much environmentally dependent. Thus, an information hiding detection scheme trained in one environment often does not work well in another environment. In this paper, we propose two hidden information detection schemes for existing anti-copying 2D barcodes. The first scheme directly uses the pixel distance to detect the use of an information hiding scheme in a 2D barcode, referred to as the Pixel Distance Based Detection (PDBD) scheme. The second scheme first calculates the variance of raw signal and the covariance between the recovered and raw signals, and then based on the variance results, detects the use of information hiding scheme in a 2D barcode, referred to as the Pixel Variance Based Detection (PVBD) scheme. Moreover, we design advanced Illegitimately-Copying (IC) attacks to evaluate the security of two existing anti-copying 2D barcodes. We conduct extensive performance comparisons among the proposed schemes and prior schemes under different capturing devices,e.g., different scanners or camera phones. Our experimental results show that the PVBD scheme can correctly detect the existence of hidden information at both the 2LQR code and the LCAC 2D barcode. Moreover, the successful attacking probability of the proposed IC attacks achieves 0.6538 for the 2LQR code and 1 for the LCAC 2D barcode. Ning Xie 0007, Yicong Chen, Changsheng Chen 0001, Lei Huang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2022 | Domain-Agnostic Document Authentication Against Practical Recapturing AttacksabstractRecapturing attack can be employed as a simple but effective anti-forensic tool for digital document images. Inspired by the document inspection process that compares a questioned document against some known samples, we proposed a document recapture detection scheme by employing a Siamese network to compare and extract distinct features in a recaptured document image. The proposed algorithm takes advantage of both metric learning and image forensic techniques, and forms triplets by considering some important factors in document authentication, e.g., document types, resolutions, and content in each image patch. After training with our triplet selection strategy, the resulting feature embedding clusters the genuine samples near the reference while pushing the recaptured samples apart. In the experiment, we consider practical settings under domain differences, such as the variations in printing/imaging devices, substrates, recapturing channels, and document types. To evaluate the robustness of different approaches, we benchmark some popular off-the-shelf machine learning-based approaches, a state-of-the-art document image detection scheme, and the proposed schemes with different network backbones under various experimental protocols. Experimental results show that the proposed scheme consistently outperforms the state-of-the-art approaches under different experimental settings. Specifically, under the most challenging scenario in our experiment, i.e., evaluation across different types of documents (produced by different manufacturers, devices, and substrates), we have achieved 6.92% APCER (Attack Presentation Classification Error Rate) and 8.51% BPCER (Bona Fide Presentation Classification Error Rate) by the proposed network with ResNeXt101 backbone at 5.00% BPCER decision threshold. Changsheng Chen 0001, Shuzheng Zhang, Fengbo Lan, Jiwu Huang |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2022 | Document Recapture Detection Based on a Unified Distortion Model of Halftone CellsabstractIn recent years, digital copies of paper documents are used widely with the prevalence of various online services. As a result, it is critical to validate the authenticity of the uploaded document images to protect against attacks from malicious users. Out of various types of attacks, the recapture attack (by reprinting and recapturing) is effective in concealing the trace of document forgeries. However, detecting the recaptured document images is challenging. To address this problem, we first study the halftone cell distortion introduced in both the genuine and recaptured document images. Based on our study, a unified model that characterizes the halftone cell distortion (e.g., errors in size and displacement) is then proposed for accurate estimation of the distortion parameters. The statistics of the estimated parameters are then exploited in a hypothesis testing framework to detect the recaptured document images. The questioned document image can be authenticated by testing against the null hypothesis, i.e., the image is a genuine sample. To evaluate the performance of the proposed approach under different application scenarios, extensive experiments are conducted with different prior knowledge of printers (known printer model, known printing technique, and In-The-Wild (unknown printing device and document contents)). The experiment results show that the proposed approach outperforms the data-driven benchmark approaches by a significant margin. Specifically, under the In-The-Wild experiment protocol, the Area Under the Receiver Operating Characteristic (ROC) Curve (AUC) of the proposed approach is above 0.87 while the AUC of the benchmark approaches (even some utilize both genuine and recaptured samples) degrades to less than 0.77. Zhaoxu Hu, Changsheng Chen 0001, Wai Ho Mow, Jiwu Huang |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2021 | Open-Set Product Authentication Based on Deep Texture Verification
Sudao Cai, Lin Zhao 0017, Changsheng Chen 0001 |
ICIG (1) | 3 |
| 2021 | Security Model of Authentication at the Physical Layer and Performance Analysis over Fading ChannelsabstractCompared with conventional authentication schemes, which utilize cryptographic mechanisms and operate at an upper layer, authentication at the physical (PHY) layer possesses many advantages, such as, higher information-theoretic security by introducing uncertainty to an adversary, better efficiency, and better compatibility by avoiding any operations at the upper layer. An effective PHY-layer authentication scheme should consider covertness, security, and robustness together. However, in published literature, these three properties have been separately analyzed. This paper proposes a generic security model for PHY-layer authentication, and we design a new systematic metric, which is referred to as the probability of security authentication (PSA). According to the proposed generic security model, we can not only systematically analyze the effect of the parameter of a certain PHY-layer authentication scheme on the final performance, but also fairly compare the performance of different PHY-layer authentication schemes under the same channel conditions at the legitimate and adversary receivers. For performance analysis, the probability of detection (PD) and the probability of false alarm (PFA) of two PHY-layer authentication schemes over fading channels are defined, and their closed-form expressions are well derived. On the basis of the proposed generic security model, a systematic performance comparison between different PHY-layer schemes is provided. Ning Xie 0007, Changsheng Chen 0001, Zhong Ming 0001 |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2021 | DRL-FAS: A Novel Framework Based on Deep Reinforcement Learning for Face Anti-SpoofingabstractInspired by the philosophy employed by human beings to determine whether a presented face example is genuine or not, i.e., to glance at the example globally first and then carefully observe the local regions to gain more discriminative information, for the face anti-spoofing problem, we propose a novel framework based on the Convolutional Neural Network (CNN) and the Recurrent Neural Network (RNN). In particular, we model the behavior of exploring face-spoofing-related information from image sub-patches by leveraging deep reinforcement learning. We further introduce a recurrent mechanism to learn representations of local information sequentially from the explored sub-patches with an RNN. Finally, for the classification purpose, we fuse the local information with the global one, which can be learned from the original input image through a CNN. Moreover, we conduct extensive experiments, including ablation study and visualization analysis, to evaluate our proposed framework on various public databases. The experiment results show that our method can generally achieve state-of-the-art performance among all scenarios, demonstrating its effectiveness. Rizhao Cai, Haoliang Li, Shiqi Wang 0001, Changsheng Chen 0001, Alex Chichung Kot |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2021 | Deep Learning-Based Forgery Attack on Document ImagesabstractWith the ongoing popularization of online services, the digital document images have been used in various applications. Meanwhile, there have emerged some deep learning-based text editing algorithms which alter the textual information of an image in an end-to-end fashion. In this work, we present a low-cost document forgery algorithm by the existing deep learning-based technologies to edit practical document images. To achieve this goal, the limitations of existing text editing algorithms towards complicated characters and complex background are addressed by a set of network design strategies. First, the unnecessary confusion in the supervision data is avoided by disentangling the textual and background information in the source images. Second, to capture the structure of some complicated components, the text skeleton is provided as auxiliary information and the continuity in texture is considered explicitly in the loss function. Third, the forgery traces induced by the text editing operation are mitigated by some post-processing operations which consider the distortions from the print-and-scan channel. Quantitative comparisons of the proposed method and the exiting approach have shown the advantages of our design by reducing the about 2/3 reconstruction error measured in MSE, improving reconstruction quality measured in PSNR and in SSIM by 4 dB and 0.21, respectively. Qualitative experiments have confirmed that the reconstruction results of the proposed method are visually better than the existing approach in both complicated characters and complex texture. More importantly, we have demonstrated the performance of the proposed document forgery algorithm under a practical scenario where an attacker is able to alter the textual information in an identity document using only one sample in the target domain. The forged-and-recaptured samples created by the proposed text editing attack and recapturing operation have successfully fooled some existing document authentication systems. Lin Zhao 0017, Changsheng Chen 0001, Jiwu Huang |
IEEE Trans. Image Process. | 2 |
| 2021 | Low-Cost Anti-Copying 2D Barcode by Exploiting Channel Noise CharacteristicsabstractIn this paper, to overcome the drawbacks of prior approaches for defending against Illegally-Copying (IC) attacks, such as low generality, high cost, and high overhead, we propose a Low-Cost Anti-Copying (LCAC) 2D barcode by exploiting the difference between the noise characteristics of legal and illegal channels. An embedding strategy is proposed, and for a variant of this strategy, we also perform a corresponding analysis. To accurately evaluate the performance of our approach, a theoretical model of the noise in an illegal channel is established by using a generalized Gaussian distribution. By comparing the experimental results based on various printers, scanners, and mobile phones, it can be found that the sample histogram and fitting curve of the theoretical model match well, so it can be concluded that the theoretical model works well. To evaluate the security of the proposed LCAC code, in addition to the Direct-Copying (DC) attack, an improved version, called the Synthesized-Copying (SC) attack, is also considered in this paper. Based on the theoretical model, we build a prediction function to optimize the parameters of our approach. Parameters optimization seeks a tradeoff between the production cost and the cost of IC attacks. The experimental results show that the proposed LCAC code with two printers and two scanners can detect DC attacks effectively and resist SC attacks up to the access of 14 legal copies. Ning Xie 0007, Yicong Chen, Changsheng Chen 0001 |
IEEE Trans. Multim. | 6 |
| 2020 | A Copy-Proof Scheme Based on the Spectral and Spatial Barcoding Channel ModelsabstractThe traditional two-dimensional (2D) barcode has been employed in anti-counterfeiting systems as a storage media for serial numbers. However, an attack can be initiated by simply copying the 2D barcode and attaching it to a counterfeit product. In this paper, we aim at proposing an authentication scheme with a mobile imaging device for a 2D barcode. This work presents a competitive solution among the 2D barcode authentication schemes that have been verified under mobile imaging conditions. The proposed copy-proof scheme is composed of two sets of features which are extracted by exploiting the characteristics of barcoding channel models. The proposed features identify the intrinsic differences between genuine and counterfeit barcode images in the frequency and spatial domains. An efficient two-stage barcode authentication framework is then proposed by combining the two sets of features in a cascading manner. To evaluate the practicality of the proposed authentication scheme, four databases with different devices (printers, scanners, mobile cameras), barcode sizes, and barcode designs are considered in the experiments. By comparing with the existing texture descriptors and some deep learning-based approaches, it is shown that the proposed scheme has a higher authentication accuracy under various conditions, such as cross-database, cross-size and cross-pattern experiments which study the generalities of a pre-trained model towards challenging conditions commonly found in real-world scenarios. Last but not least, the proposed scheme has been evaluated under some state-of-the-art attack scenarios where the attacker employs several realizations of genuine patterns or the deep learning-based technique to produce a counterfeit copy. The source code and data for producing the results in our experiments are available at https://bit.ly/2FOlJH7. Changsheng Chen 0001, Mulin Li, Anselmo Ferreira, Jiwu Huang, Rizhao Cai |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2020 | Face Spoofing Detection Based on Local Ternary Label Supervision in Fully Convolutional NetworksabstractFace verification systems are prone to spoofing attacks on photos, videos, and 3D masks. Face spoofing detection, i.e., face anti-spoofing, face liveness detection, or face presentation attack detection, is an important task for securing face verification systems in practice and presents many challenges. In this paper, a state-of-the-art face spoofing detection method based on a depth-based Fully Convolutional Network (FCN) is revisited. Different supervision schemes, including global and local label supervisions, are comprehensively investigated. A generic theoretical analysis and associated simulation are provided to demonstrate that local label supervision is more suitable than global label supervision for local tasks with insufficient training samples, such as the face spoofing detection task. Based on the analysis, the Spatial Aggregation of Pixel-level Local Classifiers (SAPLC), which is composed of an FCN part and an aggregation part, is proposed. The FCN part predicts the pixel-level ternary labels, which include the genuine foreground, the spoofed foreground, and the undetermined background. Then, these labels are aggregated together to yield an accurate image-level decision. Furthermore, to quantitatively evaluate the proposed SAPLC, experiments are carried out on the CASIA-FASD, Replay-Attack, OULU-NPU, and SiW datasets. The experiments show that the proposed SAPLC outperforms the representative deep networks, including two globally supervised CNNs, one depth-based FCN, two FCNs with binary labels, and two FCNs with ternary labels, and achieves competitive performances close to some state-of-the-art method performances under various common protocols. Overall, the results empirically verify the advantage of the proposed pixel-level local label supervision scheme. Wenyun Sun, Changsheng Chen 0001, Jiwu Huang, Alex Chichung Kot |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2018 | RA Code: A Robust and Aesthetic Code for Resolution-Constrained ApplicationsabstractRecently, some picture-embedding schemes have been proposed to improve the aesthetic appearance of 2D barcodes. However, these aesthetic 2D barcodes are not robust to the distortions incurred by the print/display-and-capture channel under limited rendering space and resolution. In this paper, a picture-embedding 2D barcode named the Robust and Aesthetic (RA) Code is proposed to counter the channel impairments, such as downsampling error and inter-symbol interference. Experimental results demonstrate that the proposed RA Code is more robust than the existing picture-embedding 2D barcodes, especially under limited printing/displaying space. The RA Code of a size as small as 1.5 × 1.5 cm2can be successfully decoded with a demodulated bit error probability of less than 0.1, which corresponds to more than 50% improvement over those of the state-of-the-art picture-embedding 2D barcodes. In addition, the correct decoding probability is improved significantly from as low as 0 to more than 0.8650. The decoding algorithm has been implemented on the Android platform, and the practicality of the RA Code has been successfully demonstrated. Changsheng Chen 0001, Baojian Zhou, Wai Ho Mow |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2017 | QR Code Authentication with Embedded Message Authentication Code
Changsheng Chen 0001 |
Mob. Networks Appl. | 1 |
| 2016 | PiCode: A New Picture-Embedding 2D BarcodeabstractNowadays, 2D barcodes have been widely used as an interface to connect potential customers and advertisement contents. However, the appearance of a conventional 2D barcode pattern is often too obtrusive for integrating into an aesthetically designed advertisement. Besides, no human readable information is provided before the barcode is successfully decoded. This paper proposes a new picture-embedding 2D barcode, called PiCode, which mitigates these two limitations by equipping a scannable 2D barcode with a picturesque appearance. PiCode is designed with careful considerations on both the perceptual quality of the embedded image and the decoding robustness of the encoded message. Comparisons with the existing beautified 2D barcodes show that PiCode achieves one of the best perceptual qualities for the embedded image, and maintains a better tradeoff between image quality and decoding robustness in various application conditions. PiCode has been implemented in the MATLAB on a PC and some key building blocks have also been ported to Android and iOS platforms. Its practicality for real-world applications has been successfully demonstrated. Changsheng Chen 0001, Wenjian Huang 0003, Baojian Zhou, Wai Ho Mow |
IEEE Trans. Image Process. | 1 |
| 2013 | A two-stage quality measure for mobile phone captured 2D barcode images
Changsheng Chen 0001, Alex Chichung Kot, Huijuan Yang |
Pattern Recognit. | 1 |
| 2012 | Step-edge reconstruction using 2D finite rate of innovation principleabstractParametric signals that have a finite number of degrees of freedom per unit of time are defined as signals with Finite Rate of Innovation (FRI). Sampling and reconstruction schemes have been developed based on the 1D FRI principle and applied to reconstructing step edge images on a row by row basis. In this paper, we derive the 2D FRI principle by exploiting the separability of the B-spline sampling kernel. The proposed 2D FRI principle regards the sampling and reconstruction as block by block operations. The step-edge parameters can be retrieved in high accuracy with no post-processing. The performance on synthetic images shows that our proposed technique is more precise than the row by row approaches on Signal-to-Noise Ratio (SNR) levels larger than 4 dB. Experimental results on real images demonstrate that the proposed method can reconstruct the step-edge precisely under noisy and practical sampling conditions. Changsheng Chen 0001, Pina Marziliano, Alex Chichung Kot |
ICASSP | 1 |
| 2010 | A quality measure of mobile phone captured 2D barcode imagesabstract2D barcode based mobile tagging technique has not fully matured yet. One of the difficulties is that mobile phone cameras induce inevitable distortions in captured barcode images. For a variety of reasons to be discussed herein, evaluation of the decodability of a barcode image is needed before the decoding process. We propose a novel blurriness measure to classify the samples according to their decodabilities. The proposed quality measure of 2D barcode image is computed by utilizing distinctive histogram features. Experimental results show that the classification accuracy is as high as 94% in our database which consists of samples captured with different phones and conditions. Comparisons with state-of-art quality measures are also conducted. The proposed method is not sensitive to noise, rotation and scaling. It is also applicable to many popular 2D barcode patterns. Changsheng Chen 0001, Alex Chichung Kot, Huijuan Yang |
ICIP | 1 |