Gang Cao 0001

dblp:49/982-1 · DBLP profile ↗
← Back
31ranked-venue papers
8as first author
20since 2021 · last 2027
0000-0002-4549-0125ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 22 · 6 first-author · 15 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Security and privacy · 3 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 1Theory of computation · 1
YearPublicationVenuePosition
2027 Progressive feedback-enhanced transformer for image forgery localization
Haochen Zhu, Weiheng Zhu, Gang Cao 0001, Xianglin Huang
Expert Syst. Appl.3
2026 Transferable Dual-Domain Feature Importance Attack Against AI-Generated Image Detector
abstract
Recent AI-generated image (AIGI) detectors achieve impressive accuracy under clean condition. In view of anti-forensics, it is significant to develop advanced adversarial attacks for evaluating the security of such detectors, which remains unexplored sufficiently. This letter proposes a Dual-domain Feature Importance Attack (DuFIA) scheme to invalidate AIGI detectors to some extent. Forensically important features are captured by the spatially interpolated gradient and frequency-aware perturbation. The adversarial transferability is enhanced by jointly modeling spatial and frequency-domain feature importances, which are fused to guide the optimization-based adversarial example generation. Extensive experiments across various AIGI detectors verify the cross-model transferability, transparency and robustness of DuFIA.
Weiheng Zhu, Gang Cao 0001, Lifang Yu, ShaoWei Weng
IEEE Signal Process. Lett.2
2025 A Lightweight and Effective Image Tampering Localization Network With Vision Mamba
abstract
Current image tampering localization methods primarily rely on Convolutional Neural Networks (CNNs) and Transformers. While CNNs suffer from limited local receptive fields, Transformers offer global context modeling at the expense of quadratic computational complexity. Recently, the state space model Mamba has emerged as a competitive alternative, enabling linear-complexity global dependency modeling. Inspired by it, we propose a lightweight and effective FORensic network based on vision MAmba (ForMa) for blind image tampering localization. Firstly, ForMa captures multi-scale global features that achieves efficient global dependency modeling through linear complexity. Then the pixel-wise localization map is generated by a lightweight decoder, which employs a parameter-free pixel shuffle layer for upsampling. Additionally, a noise-assisted decoding strategy is proposed to integrate complementary manipulation traces from tampered images, boosting decoder sensitivity to forgery cues. Experimental results on 10 standard datasets demonstrate that ForMa achieves state-of-the-art generalization ability and robustness, while maintaining the lowest computational complexity. Code is available at https://github.com/multimediaFor/ForMa.
Kun Guo 0010, Gang Cao 0001, Zijie Lou, Xianglin Huang, Jiaoyun Liu
IEEE Signal Process. Lett.2
2025 Image Forgery Localization With State Space Models
abstract
Pixel dependency modeling from tampered images is pivotal for image forgery localization. Current approaches predominantly rely on Convolutional Neural Networks (CNNs) or Transformer-based models, which often either lack sufficient receptive fields or entail significant computational overheads. Recently, State Space Models (SSMs), exemplified by Mamba, have emerged as a promising approach. They not only excel in modeling long-range interactions but also maintain a linear computational complexity. In this paper, we propose LoMa, a novel image forgery localization method that leverages the selective SSMs. Specifically, LoMa initially employs atrous selective scan to traverse the spatial domain and convert the tampered image into ordered patch sequences, and subsequently applies multi-directional state space modeling. In addition, an auxiliary convolutional branch is introduced to enhance local feature extraction. Extensive experimental results validate the superiority of LoMa over CNN-based and Transformer-based state-of-the-arts. To our best knowledge, this is the first image forgery localization model constructed based on the SSM-based model. We aim to establish a baseline and provide valuable insights for the future development of more efficient and effective SSM-based forgery localization models.
Zijie Lou, Gang Cao 0001, Kun Guo 0010, ShaoWei Weng, Lifang Yu
IEEE Signal Process. Lett.2
2025 Video Inpainting Localization With Contrastive Learning
abstract
Video inpainting techniques typically serve to restore destroyed or missing regions in digital videos. However, such techniques may also be illegally used to remove important objects for creating forged videos. This letter proposes a simple yet effective forensic scheme for Video Inpainting LOcalization with ContrAstive Learning (ViLocal). A 3D Uniformer encoder is applied to the video noise residual for learning effective spatiotemporal features. To enhance discriminative power, supervised contrastive learning is adopted to capture the local regional inconsistency through separating the pristine and inpainted pixels. The pixel-wise inpainting localization map is yielded by a lightweight convolution decoder with two-stage training. To prepare enough training samples, we build a video object segmentation dataset (VOS2k5) of 2500 videos with pixel-level annotations per frame. Extensive experimental results validate the superiority of ViLocal over the state-of-the-arts.
Zijie Lou, Gang Cao 0001, Man Lin
IEEE Signal Process. Lett.2
2025 A Labeled Intermediate Domain Aided Two-Domain Correlation Fusion for Mismatched Steganalysis
abstract
When the target images to be detected and the source images used to train the steganalyzer come from different distributions, the cover-source mismatch (CSM) occurs, which often leads to a sharp decrease in detection accuracy. To alleviate the problem, this letter proposes a four-stage steganalyzer, named ICSNet. The first three stages concentrate on generating a labeled intermediate domain to build a bridge between source/target domains. To be specific, the labeled intermediate domain is constructed by first adding the noise to target samples using a noise adding module to generate intermediate samples following the source domain distribution, and subsequently performing data embedding to these intermediate samples to generate the stego intermediate samples. The last stage focuses on strengthening the fusion of channel-wise and spatial correlations between source/target domains by presenting a coarse-to-fine two-domain channel-wise correlation fusion (CCF) module and a source-guided two-domain spatial correlation fusion (SCF) module. In CCF, the labeled intermediate domain works as a supplementary to the target domain, so that source/target domains can guide each other to reinforce the fusion of channel-wise correlations. In SCF, the labeled source domain guides the target domain to consolidate the fusion of spatial correlations. The four stages work together to reduce the distribution shift between source/target domains, thereby bringing performance improvement. Experimental results demonstrate that ICSNet significantly outperforms existing methods across various CSM scenarios.
ShaoWei Weng, Yang Li 0205, Lifang Yu, Gang Cao 0001
IEEE Signal Process. Lett.4
2025 WL-WEM Combining Low-Cost Watermark Enhancement Modules for In-Generation Watermarking
abstract
In general, modifying the latent diffusion model (LDM) decoder to achieve in-generation watermarking cannot introduce tremendous computational burden, which easily leads to non-convergence. This necessarily increases the difficulty of embedding the watermark into the LDM decoder due to the need to strike a balance among imperceptibility, robustness and computational cost. We realize the difficulty and design two lightweight watermarking modules, namely a low-cost watermark redundancy enhancement module (WREM) and a latent-guided watermark enhancement module (LWEM), aiming at reducing the modifications to the LDM decoder as much as possible while maintaining the generation quality and enhancing the robustness. Specifically, WREM, specially designed for shallow layers, utilizes a small number of repetition operations to strengthen the robustness of the watermark, and adopts a low-cost sub-pixel convolution layer to achieve dimension consistency between the watermark residual and the input latent, greatly reducing the computational cost while enhancing the integration of watermark features and the latent feature. LWEM, tailored for deep layers, innovatively exploits a simple bilinear interpolation to strengthen the robustness of the watermark, and fuses watermark features and the latent feature using a cheap convolution layer so as to generate the watermark residual with relatively low impact on the input latent. Combining WREM and LWEM, we construct a lightweight encoder-noiselayer-decoder in-generation watermarking method dubbed WL-WEM pursuing a satisfactory balance among three metrics including computational cost, generation quality and robustness. Experimental results also demonstrate that the proposed WL-WEM outperforms several related works in balancing three metrics.
Lifang Yu, Xinchen Geng, ShaoWei Weng, Yang Li 0205, Gang Cao 0001
IEEE Signal Process. Lett.5
2025 Trusted Video Inpainting Localization via Deep Attentive Noise Learning
abstract
Digital video inpainting technique has been substantially improved with deep learning in recent years. It may be used as malicious manipulation to remove important objects for creating forged videos. As such it is significant to blindly identify the inpainted regions in videos. In this paper, we present a Trusted Video Inpainting Localization network (TruVIL) with excellent robustness and generalization ability. Observing that high-frequency noise can effectively unveil the inpainted regions, we design deep attentive noise learning in multiple stages to capture the inpainting traces. Firstly, a multiscale noise extraction module based on 3D High Pass (HP3D) layers is used to create the noise modality from input RGB frames. Then the correlation between such two complementary modalities are explored by a cross-modality attentive fusion module to facilitate mutual feature learning. Lastly, spatial details are selectively enhanced by an attentive noise decoding module to boost the localization performance of the network. To prepare enough training samples, we also build a frame-level video object segmentation dataset (VOS2k5) with 2500 videos and pixel-level annotation for all frames. Both quantitative and qualitative evaluations on various inpainted videos verify the robustness against video compression and generalization ability of TruVIL.
Zijie Lou, Gang Cao 0001, Man Lin, Lifang Yu, ShaoWei Weng
IEEE Trans. Dependable Secur. Comput.2
2025 Exploring Multi-View Pixel Contrast for General and Robust Image Forgery Localization
abstract
Image forgery localization, which aims to segment tampered regions in an image, is a fundamental yet challenging digital forensic task. While some deep learning-based forensic methods have achieved impressive results, they directly learn pixel-to-label mappings without fully exploiting the relationship between pixels in the feature space. To address such deficiency, we propose a Multi-view Pixel-wise Contrastive algorithm (MPC) for image forgery localization. Specifically, we first pre-train the feature extraction backbone network with a supervised contrastive loss to model pixel relationships in view of within-image, cross-scale and cross-modality. That is aimed at increasing intra-class compactness and inter-class separability. Then the localization head is fine-tuned using cross-entropy loss, resulting in a better forged pixel localizer. The MPC is trained on three different scale training datasets to make a comprehensive and fair comparison with existing image forgery localization algorithms. Extensive test results on over ten public datasets show that the proposed MPC achieves higher generalization performance and robustness than the state-of-the-arts. It is particularly noteworthy that our approach maintains a high level of localization accuracy under various post-processing combinations that approximate real-world scenarios, as well as when confronted with novel intelligent editing techniques. Finally, comprehensive and detailed ablation experiments demonstrate the reasonableness of MPC.
Zijie Lou, Gang Cao 0001, Kun Guo 0010, Lifang Yu, ShaoWei Weng
IEEE Trans. Inf. Forensics Secur.2
2024 Effective Image Tampering Localization Via Enhanced Transformer and Co-Attention Fusion
abstract
Powerful manipulation techniques have made digital image forgeries be easily created and widespread without leaving visual anomalies. The blind localization of tampered regions becomes quite significant for image forensics. In this paper, we propose an effective image tampering localization network (EITLNet) based on a two-branch enhanced transformer encoder with attention-based feature fusion. Specifically, a feature enhancement module is deployed to enhance the feature representation ability of the transformer encoder. The features extracted from RGB and noise streams are fused effectively by the coordinate attention-based fusion module at multiple scales. Extensive experimental results verify that the proposed scheme achieves the state-of-the-art generalization ability and robustness in various benchmark datasets. Code is public at https://github.com/multimediaFor/EITLNet.
Kun Guo 0010, Haochen Zhu, Gang Cao 0001
ICASSP3
2024 AI-Generated Video Detection via Spatial-Temporal Anomaly Learning
Jianfa Bai, Man Lin, Gang Cao 0001, Zijie Lou
PRCV (10)3
2024 A deep steganalysis network combining source-supervised and target-unsupervised information for cover-source mismatch
Lifang Yu, Zhuwei Zhang, ShaoWei Weng, Gang Cao 0001
Expert Syst. Appl.5
2024 Transferable adversarial attack on image tampering localization
Gang Cao 0001, Haochen Zhu, Zijie Lou, Lifang Yu
J. Vis. Commun. Image Represent.1
2024 Effective image tampering localization with multi-scale ConvNeXt feature fusion
Haochen Zhu, Gang Cao 0001, Mo Zhao, Huawei Tian, Weiguo Lin
J. Vis. Commun. Image Represent.2
2024 Attention-enhanced joint learning network for micro-video venue classification
Bing Wang 0013, Xianglin Huang, Gang Cao 0001, Lifang Yang, Zhulin Tao
Multim. Tools Appl.3
2024 Discriminability-Aware Intermediate Domains for Mismatched Steganalysis
abstract
This letter proposes GDNet equipped with the generation of discriminative mixing regions (GDMR) and discriminability-aware local image mixing (DLIM), a steganalysis network aiming at alleviating significant accuracy degradation caused by cover-source mismatch (CSM), which pertains to the situation where source and target domains come from different distributions. GDNet guides a steganalyzer trained on the source domain to the target domain by mixing the source and target images at the region-level and pixel-level to construct a discriminative intermediate domain. On the one hand, GDMR designs an epoch-related region-level mixing ratio to control the size of the mixed region, and based on this ratio, selects the regions within the target image strongly related to the stego signal to participate in the generation of the intermediate domain, while suppressing other regions weakly related to the stego signal. On the other hand, DLIM utilizes the pixel-level mixing ratio to reduce the impact of the regions weakly related to the stego signal on the discriminability of the intermediate domain as the region-level mixing ratio increases, thereby increasing the diversity of the intermediate domain. Experimental results demonstrate that GDNet significantly outperforms existing methods across various CSM scenarios.
Yang Li 0205, Lifang Yu, ShaoWei Weng, Huawei Tian, Gang Cao 0001
IEEE Signal Process. Lett.5
2024 High-Precision Reversible Data Hiding Predictor: UCANet
abstract
Existing convolutional neural network-based reversible data hiding (RDH) predictors typically stack the standard convolution blocks with stride 1 for feature extraction, and keep the sizes of input and output feature maps unchanged through padding. This suggests that only a limited range of contextual spatial information is obtained. To remedy this problem above, a U-Net-like RDH predictor named UCANet is proposed in this paper to capture rich multi-scale contextual information by gradually downsampling feature maps. To fuse two feature maps at different levels along the channel dimension, we put forward the channel adaptive attention (CAA). By merely combining cheap pointwise convolution operations, CAA achieves the integration of non-linear and linear features as well as implicitly enhances channel dimensionality with low computational burden, thereby effectively enriching the expression of the channel information. The design of UCANet considers the characteristics of RDH from two aspects. On the one hand, instead of maxpooling or average pooling commonly used for downsampling, a stride-2 convolution block that can adaptively adjust the weights of convolution kernels and select useful information is utilized to downsample feature maps. On the other hand, UCANet removes the batch normalization layers to avoid their influence on the distribution of feature maps, which helps to strengthen the network's prediction capability. Extensive experiments also demonstrate that the proposed UCANet achieves better prediction performance, compared to several state-of-the-art methods.
Haiyang Rao, ShaoWei Weng, Lifang Yu, Li Li 0014, Gang Cao 0001
IEEE Signal Process. Lett.5
2024 Universal Mismatched Steganalysis Equipped With Progressive Intermediate Domains
abstract
In general, cover source mismatch (CSM) inevitably leads to a significant decrease in detection accuracy in image steganalysis because the source and target domains have different distributions. To remedy this problem, a universal mismatched steganalyzer ISNet equipped with generated local mixing positions, a local feature-level mixup-related patchup (LFMP), and domain factors is proposed for both spatial and JPEG domains in this paper. Unlike existing deep steganalysis networks, which simply minimize the domain discrepancy between source and target to address the problem of CSM, and thus, cannot handle large domain discrepancy, IDGM based on LFMP generates diverse intermediate domains to bridge the two extreme domains so as to alleviate the decrease in detection accuracy caused by CSM. Moreover, ISNet enables the intermediate domain distribution to progressively transit from source to target by adjusting the domain factor sampled from Beta(α, 1), so that the classifier gradually adapted to target provides an improvement in discriminability on the target domain. The experimental results show that ISNet achieves the best performance in various CSM cases, compared with the most advanced deep learning-based steganalysis network.
ShaoWei Weng, Zhuwei Zhang, Lifang Yu, Gang Cao 0001
IEEE Signal Process. Lett.5
2023 Black-box attack against GAN-generated image detector with contrastive perturbation
Zijie Lou, Gang Cao 0001, Man Lin
Eng. Appl. Artif. Intell.2
2022 Hybrid Transformer-CNN for Real Image Denoising
abstract
Transformer typically enjoys larger model capacity but higher computational loads than convolutional neural network (CNN) in vision tasks. In this letter, the advantages of such two networks are fused for achieving effective and efficient real image denoising. We propose a hybrid denoising model based on Transformer Encoder and Convolutional Decoder Network (TECDNet). The Transformer based on novel radial basis function (RBF) attention is used as encoder to improve the representation capability of overall model. In decoder, the residual CNN instead of Transformer is adopted to greatly reduce computational complexity of the whole denoising network. Extensive experimental results on real images show that TECDNet achieves the state-of-the-art denosing performance with relatively low computational cost.
Mo Zhao, Gang Cao 0001, Xianglin Huang, Lifang Yang
IEEE Signal Process. Lett.2
2020 Multi-modal sequence model with gated fully convolutional blocks for micro-video venue classification
Wei Liu 0084, Xianglin Huang, Gang Cao 0001, Jianglong Zhang, Gege Song, Lifang Yang
Multim. Tools Appl.3
2019 Better Word Representations with Word Weight
abstract
As a fundamental task of natural language processing, text classification has been widely used in various applications such as sentiment analysis and spam detection. In recent years, the continuous-valued word embedding learned by neural network attaches extensive attentions. Although word embedding achieves impressive results in capturing similarities and regularities between words, it fails to highlight important words for identifying text category. Such deficiency could be attenuated by word weight, which conveys word contribution in text categorization. Toward this end, we propose an effective text classification scheme by incorporating word weight into word embedding in this paper. Specifically, in order to enrich word representation, the bidirectional gated recurrent units (Bi-GRU) is first employed to grasp context information of words. Then the word weights yielded by term frequency (TF) are used to modulate the word representation of Bi-GRU for constructing text representation. Extensive experimental results on several large text datasets verify that the accuracy of our proposed text classification scheme outperforms the state-of-the-art ones.
Gege Song, Xianglin Huang, Gang Cao 0001, Zhulin Tao, Wei Liu 0084, Lifang Yang
MMSP3
2018 Acceleration of histogram-based contrast enhancement via selective downsampling
abstract
The authors propose a general framework to accelerate the universal histogram‐based image contrast enhancement (CE) algorithms. Both spatial and grey‐level selective downsampling of digital images are adopted to decrease computational cost, while the visual quality of enhanced images is still preserved and without apparent degradation. Mapping function calibration is proposed to reconstruct the pixel mapping on the grey levels missed by downsampling. As two case studies, the accelerations of histogram equalisation (HE) and the state‐of‐the‐art global CE algorithm, i.e. spatial mutual information and PageRank (SMIRANK), are presented in detail. Both quantitative and qualitative assessment results have verified the effectiveness of their proposed CE acceleration framework. In typical tests, the computational efficiencies of HE and SMIRANK have been increased by about 3.9 and 13.5 times, respectively.
Gang Cao 0001, Huawei Tian, Lifang Yu, Xianglin Huang, Yongbin Wang
IET Image Process.1
2018 Translation and scale invariants of Krawtchouk moments
Ruicong Zhi, Lianyu Cao, Gang Cao 0001
Inf. Process. Lett.3
2014 Attacking contrast enhancement forensics in digital images
Gang Cao 0001, Yao Zhao 0001, Huawei Tian, Lifang Yu
Sci. China Inf. Sci.1
2014 A channel selection rule for YASS
Lifang Yu, Yao Zhao 0001, Gang Cao 0001
Sci. China Inf. Sci.4
2014 Contrast Enhancement-Based Forensics in Digital Images
abstract
As a retouching manipulation, contrast enhancement is typically used to adjust the global brightness and contrast of digital images. Malicious users may also perform contrast enhancement locally for creating a realistic composite image. As such it is significant to detect contrast enhancement blindly for verifying the originality and authenticity of the digital images. In this paper, we propose two novel algorithms to detect the contrast enhancement involved manipulations in digital images. First, we focus on the detection of global contrast enhancement applied to the previously JPEG-compressed images, which are widespread in real applications. The histogram peak/gap artifacts incurred by the JPEG compression and pixel value mappings are analyzed theoretically, and distinguished by identifying the zero-height gap fingerprints. Second, we propose to identify the composite image created by enforcing contrast adjustment on either one or both source regions. The positions of detected blockwise peak/gap bins are clustered for recognizing the contrast enhancement mappings applied to different source regions. The consistency between regional artifacts is checked for discovering the image forgeries and locating the composition boundary. Extensive experiments have verified the effectiveness and efficacy of the proposed techniques.
Gang Cao 0001, Yao Zhao 0001, Xuelong Li 0001
IEEE Trans. Inf. Forensics Secur.1
2011 Unsharp Masking Sharpening Detection via Overshoot Artifacts Analysis
abstract
In this letter, we propose a new method in detecting unsharp masking (USM) sharpening operation in digital images. Overshoot artifacts are found to occur around side-planar edges in the sharpened images. Such artifacts, measured by a sharpening detector, can serve as a rather unique feature for identifying the previous performance of sharpening operation. Test results on photograph images with regard to various sharpening operators show the effectiveness of our proposed method.
Gang Cao 0001, Yao Zhao 0001, Alex Chichung Kot
IEEE Signal Process. Lett.1
2010 Forensic estimation of gamma correction in digital images
abstract
In the digital era, digital photographs become pervasive and are frequently used to record event facts. Authenticity and integrity of such photos can be ascertained by discovering more information about the previously applied operations. In this paper, we propose a forensic scheme for identifying and reconstructing gamma correction operations in digital images. Statistical abnormity on image grayscale histograms, which is caused by the contrast enhancement, is analyzed theoretically and measured effectively. Graylevel mapping functions involved in gamma correction can be estimated blindly. Experiments both on globally and locally applied corrected images show the validity of our proposed gamma estimation algorithm.
Gang Cao 0001, Yao Zhao 0001
ICIP1
2010 Forensic detection of median filtering in digital images
abstract
In digital image forensics, prior works are prone to the detection of malicious tampering. However, there is also a need for developing techniques to identify general content-preserved manipulations, which are employed to conceal tampering trails frequently. In this paper, we propose a blind forensic algorithm to detect median filtering (MF), which is applied extensively for signal denoising and digital image enhancement. The probability of zero values on the first order difference map in texture regions can serve as MF statistical fingerprint, which distinguishes MF from other operations. Since anti-forensic techniques enjoy utilizing MF to attack the linearity assumption of existing forensics algorithms, blind detection of the non-linear MF becomes especially significant. Both theoretically reasoning and experimental results verify the effectiveness of our proposed MF forensics scheme.
Gang Cao 0001, Yao Zhao 0001, Lifang Yu, Huawei Tian
ICME1
2009 Detection of image sharpening based on histogram aberration and ringing artifacts
abstract
With the wide use of sophisticated photo manipulation, capability of digital images for recording authentic scene information has been addressed with public suspicion. So it is urgent to develop forensics techniques for verifying photographs' originality and authenticity. In this paper, a blind forensic algorithm is proposed to detect sharpening manipulation in digital images. Gradient aberration of the gray histogram generated from unsaturated luminance regions of an image, is measured and employed to capture trails of sharpening operation. Ringing artifacts around step edges are exploited to provide another complementary clue, especially useful when the histogram-based features are not available. Tests on plenty of photo images show the effecttiveness of our proposed sharpening detection scheme.
Gang Cao 0001, Yao Zhao 0001
ICME1