Lifang Yu

dblp:53/374 · DBLP profile ↗
← Back
36ranked-venue papers
9as first author
29since 2021 · last 2026
0000-0002-0508-7526ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 27 · 5 first-author · 24 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 3 since 2021Security and privacy · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author
YearPublicationVenuePosition
2026 Block and frequency-band guided AC coefficient expandability estimation for JPEG reversible data hiding
Lifang Yu, ShaoWei Weng, Yao Zhao 0001
J. Vis. Commun. Image Represent.1
2026 An image steganalyzer with an end-to-end trainable preprocessing and an efficient feature extraction module
Hongrui Lin, ShaoWei Weng, Lifang Yu
Signal Process.3
2026 DTBF: Combining Local Statistical Artifacts and Concept Alignment for Synthetic Image Detection
ShaoWei Weng, Lifang Yu, Gaobo Yang, Pei-Wei Tsai
IEEE Signal Process. Lett.3
2026 A Dual-Reward Guided 2D Mapping Generation Network for JPEG Reversible Data Hiding
abstract
Recently, researchers have shifted focus to reversible data hiding (RDH) schemes for JPEG images. The reinforcement learning (RL) is a solution for RDH to automatically acquire the optimal two-dimensional (2D) mapping for 2D histograms of non-zero quantized alternating current coefficients. However, merely utilizing the payload-distortion reward mechanism (PDRM) in RL cannot inject the payload guidance to the 2D mapping generation process. To tackle this issue, we propose a payload supplementary reward mechanism (PSRM) and incorporate PDRM and PSRM into RL to construct DR-2DNet, a dual-reward guided 2D mapping generation network with considering additional payload guidance. DR-2DNet generates two candidate 2D mappings, one with low distortion generated by merely utilizing PDRM and the other with low distortion and high payload obtained by jointly using PDRM and PSRM. Finally, according to the required payload, the one with the lower distortion selected from two acquired 2D mappings is used for achieving data embedding. To priorly select the frequency bands with low costs for data embedding, a frequency selection strategy combining the smoothness and embedding performance of the frequency band is designed to evaluate the cost of each frequency band, reducing image distortion and preserving the file size. Extensive experiments are conducted on the Kodak dataset and 100 images randomly chosen from the BOSSBase dataset, and the results demonstrate that the proposed method is superior to several related state-of-the-art RDH schemes for JPEG images.
Yao Zhao 0001, ShaoWei Weng, Lifang Yu
IEEE Signal Process. Lett.4
2026 Transferable Dual-Domain Feature Importance Attack Against AI-Generated Image Detector
abstract
Recent AI-generated image (AIGI) detectors achieve impressive accuracy under clean condition. In view of anti-forensics, it is significant to develop advanced adversarial attacks for evaluating the security of such detectors, which remains unexplored sufficiently. This letter proposes a Dual-domain Feature Importance Attack (DuFIA) scheme to invalidate AIGI detectors to some extent. Forensically important features are captured by the spatially interpolated gradient and frequency-aware perturbation. The adversarial transferability is enhanced by jointly modeling spatial and frequency-domain feature importances, which are fused to guide the optimization-based adversarial example generation. Extensive experiments across various AIGI detectors verify the cross-model transferability, transparency and robustness of DuFIA.
Weiheng Zhu, Gang Cao 0001, Lifang Yu, ShaoWei Weng
IEEE Signal Process. Lett.4
2026 DCNet: Learning Similarity and Spatial Complementary Features for Generalized AI-Generated Image Detection
ShaoWei Weng, Lifang Yu
IEEE Trans. Circuits Syst. Video Technol.3
2026 DV-Net: Detecting and Distinguishing Copy-Move Regions From Dual Views
abstract
Copy-move forgery detection (CMFD) is a technique tailored to detect the existence of copy-move regions in a query image. In this paper, a dual-view CMFD network named DV-Net is proposed, which integrates the combination of the similarity information and tampered features from shallow features conducive to copy-move region localization by using dual-view self-correlation calculation (DV-SCC) and the shallow similarity attention module (SSAM), and strengthens the ability of distinguishing source/target regions by making deep features pass through three serial multiple serial adaptive receptive field selection modules (ARFSMs). The SCC plays an irreplaceable role in identifying copy-move regions. However, single-view SCC, such as the cosine similarity or the Euclidean distance, can solely capture the similarly information from a single perspective. DV-SCC, a combination of Euclidean distance and cosine similarity, provides more comprehensive similarity information from numerical and directional perspectives. In addition, different from previous CMFD networks that only utilize the similarity information to locate similar regions while neglecting tampered features contained in the shallow features, which are of vital importance to CMFD, we innovatively convert the similarity information into the SSAM and apply SSAM on the shallow features to emphasize the similarity information while preserving tampered features, significantly enhancing the localization accuracy of source/target regions. Multiple serial ARFSMs, each containing two parallel branches controlled by a soft attention, can adaptively select appropriate receptive fields according to the scales of tampered regions, improving the classification accuracy of source/target regions. The experimental results show that DV-Net outperforms several advanced algorithms in source/target region localization and discrimination on three publicly available datasets.
ShaoWei Weng, Lifang Yu, Tangguo Zhu
IEEE Trans. Multim.3
2025 Image Forgery Localization With State Space Models
abstract
Pixel dependency modeling from tampered images is pivotal for image forgery localization. Current approaches predominantly rely on Convolutional Neural Networks (CNNs) or Transformer-based models, which often either lack sufficient receptive fields or entail significant computational overheads. Recently, State Space Models (SSMs), exemplified by Mamba, have emerged as a promising approach. They not only excel in modeling long-range interactions but also maintain a linear computational complexity. In this paper, we propose LoMa, a novel image forgery localization method that leverages the selective SSMs. Specifically, LoMa initially employs atrous selective scan to traverse the spatial domain and convert the tampered image into ordered patch sequences, and subsequently applies multi-directional state space modeling. In addition, an auxiliary convolutional branch is introduced to enhance local feature extraction. Extensive experimental results validate the superiority of LoMa over CNN-based and Transformer-based state-of-the-arts. To our best knowledge, this is the first image forgery localization model constructed based on the SSM-based model. We aim to establish a baseline and provide valuable insights for the future development of more efficient and effective SSM-based forgery localization models.
Zijie Lou, Gang Cao 0001, Kun Guo 0010, ShaoWei Weng, Lifang Yu
IEEE Signal Process. Lett.5
2025 A Labeled Intermediate Domain Aided Two-Domain Correlation Fusion for Mismatched Steganalysis
abstract
When the target images to be detected and the source images used to train the steganalyzer come from different distributions, the cover-source mismatch (CSM) occurs, which often leads to a sharp decrease in detection accuracy. To alleviate the problem, this letter proposes a four-stage steganalyzer, named ICSNet. The first three stages concentrate on generating a labeled intermediate domain to build a bridge between source/target domains. To be specific, the labeled intermediate domain is constructed by first adding the noise to target samples using a noise adding module to generate intermediate samples following the source domain distribution, and subsequently performing data embedding to these intermediate samples to generate the stego intermediate samples. The last stage focuses on strengthening the fusion of channel-wise and spatial correlations between source/target domains by presenting a coarse-to-fine two-domain channel-wise correlation fusion (CCF) module and a source-guided two-domain spatial correlation fusion (SCF) module. In CCF, the labeled intermediate domain works as a supplementary to the target domain, so that source/target domains can guide each other to reinforce the fusion of channel-wise correlations. In SCF, the labeled source domain guides the target domain to consolidate the fusion of spatial correlations. The four stages work together to reduce the distribution shift between source/target domains, thereby bringing performance improvement. Experimental results demonstrate that ICSNet significantly outperforms existing methods across various CSM scenarios.
ShaoWei Weng, Yang Li 0205, Lifang Yu, Gang Cao 0001
IEEE Signal Process. Lett.3
2025 GURNet: A Gated U-Shaped Encoder-Decoder Predictor for Reversible Data Hiding
abstract
The existing deep learning based reversible data hiding (RDH) predictors typically adopt standard convolutions for extracting features, which inherently fails to capture contextual information across different scales, making the model have difficulty to fully understand the image content. To this end, a gated multi-scale module (GMM) is proposed to enrich and strengthen feature representations by collecting multi-scale features with less computational cost using a set of parallel depthwise convolutions, and customizing the gated convolution (GConv) for RDH to weight the importance of features in channel and spatial dimensions. Considering that directly utilizing the addition or concatenation operations cannot better fuse two types of features with different receptive fields, a gated feature fusion and refinement module (GFFRM) is tailored to employ the standard convolutions of different sizes to shorten the receptive field differences between deep and shallow features. GFFRM also constructs depthwise separable convolution followed by GConv to enrich and refine the expression of features at low computational cost and enhance the information exchange across channel and spatial dimensions, thereby improving the fusion effect of features at different levels. A two-path multi-dimensional feature interaction module (MFIM) is designed, where one branch utilizes a pointwise convolution to obtain low-dimensional representations of features, whereas the other branch fuse two linearly transformed features through element- wise multiplication constructs to generate implicit high-dimensional features. GFFRM and MFIM are complementary for each other to enhance the prediction performance. Three modules, namely GMM, GFFRM and MFIM, are embedded in U -shaped encoder-decoder architecture to establish a novel RDH predictor GURNet. Extensive experiments implemented on four publicly available datasets demonstrate the superiority of GURNet, compared with state-of-the-art RDH predictors.
ShaoWei Weng, Haiyang Rao, Lifang Yu
IEEE Signal Process. Lett.3
2025 WL-WEM Combining Low-Cost Watermark Enhancement Modules for In-Generation Watermarking
abstract
In general, modifying the latent diffusion model (LDM) decoder to achieve in-generation watermarking cannot introduce tremendous computational burden, which easily leads to non-convergence. This necessarily increases the difficulty of embedding the watermark into the LDM decoder due to the need to strike a balance among imperceptibility, robustness and computational cost. We realize the difficulty and design two lightweight watermarking modules, namely a low-cost watermark redundancy enhancement module (WREM) and a latent-guided watermark enhancement module (LWEM), aiming at reducing the modifications to the LDM decoder as much as possible while maintaining the generation quality and enhancing the robustness. Specifically, WREM, specially designed for shallow layers, utilizes a small number of repetition operations to strengthen the robustness of the watermark, and adopts a low-cost sub-pixel convolution layer to achieve dimension consistency between the watermark residual and the input latent, greatly reducing the computational cost while enhancing the integration of watermark features and the latent feature. LWEM, tailored for deep layers, innovatively exploits a simple bilinear interpolation to strengthen the robustness of the watermark, and fuses watermark features and the latent feature using a cheap convolution layer so as to generate the watermark residual with relatively low impact on the input latent. Combining WREM and LWEM, we construct a lightweight encoder-noiselayer-decoder in-generation watermarking method dubbed WL-WEM pursuing a satisfactory balance among three metrics including computational cost, generation quality and robustness. Experimental results also demonstrate that the proposed WL-WEM outperforms several related works in balancing three metrics.
Lifang Yu, Xinchen Geng, ShaoWei Weng, Yang Li 0205, Gang Cao 0001
IEEE Signal Process. Lett.1
2025 Steganalysis Network With Two-Branch Preprocessing for Spatial and JPEG Domains
abstract
Considering that the nature of the stego signal caused by spatial domain steganography and joint photographic experts group (JPEG) domain steganography is different, existing deep-learning steganalysis networks typically cannot work well in both spatial and JPEG domains. We propose a unified steganalysis network named ESNet to effectively preserve and identify the stego signal from spatial and JPEG domains. Specifically, dual-branch preprocessing extracts noise residuals by using fixed SRM kernels (branch 1) and randomly initialized kernels (branch 2), fuses the features from two branches and exchanges the fused complementary information through two carefully designed bidirectional fusion blocks, thereby effectively enhancing the signal-to-noise ratio. During feature extraction, considering that low-level features, such as texture and edge, are indispensable for steganalysis, we gather multi-level feature maps at different layers of the network to provide richer feature representations and merge them by using a multi-level feature fusion module, which learns the weight of different features in single-level feature map to enhance the expression of steganographic features. During classification, the multi-scale attention pooling module is employed to extract multi-scale features by designing convolution kernels of different sizes. After concatenating features of different scales, gated channel transformation is exploited to weight the importance of each channel to further strengthen the representations of steganographic features. Finally, stylepooling in combination with global standard deviation pooling and global average pooling, is used to compress channels and preserve the representation ability of channels as much as possible for classification. The experimental results show that the proposed ESNet exhibits state-of-the-art detection performance in both spatial and JPEG domains, and achieves satisfactory robustness against the cover source mismatch.
ShaoWei Weng, Lifang Yu, Dewang Chen
IEEE Trans. Circuits Syst. Video Technol.3
2025 Trusted Video Inpainting Localization via Deep Attentive Noise Learning
abstract
Digital video inpainting technique has been substantially improved with deep learning in recent years. It may be used as malicious manipulation to remove important objects for creating forged videos. As such it is significant to blindly identify the inpainted regions in videos. In this paper, we present a Trusted Video Inpainting Localization network (TruVIL) with excellent robustness and generalization ability. Observing that high-frequency noise can effectively unveil the inpainted regions, we design deep attentive noise learning in multiple stages to capture the inpainting traces. Firstly, a multiscale noise extraction module based on 3D High Pass (HP3D) layers is used to create the noise modality from input RGB frames. Then the correlation between such two complementary modalities are explored by a cross-modality attentive fusion module to facilitate mutual feature learning. Lastly, spatial details are selectively enhanced by an attentive noise decoding module to boost the localization performance of the network. To prepare enough training samples, we also build a frame-level video object segmentation dataset (VOS2k5) with 2500 videos and pixel-level annotation for all frames. Both quantitative and qualitative evaluations on various inpainted videos verify the robustness against video compression and generalization ability of TruVIL.
Zijie Lou, Gang Cao 0001, Man Lin, Lifang Yu, ShaoWei Weng
IEEE Trans. Dependable Secur. Comput.4
2025 Exploring Multi-View Pixel Contrast for General and Robust Image Forgery Localization
abstract
Image forgery localization, which aims to segment tampered regions in an image, is a fundamental yet challenging digital forensic task. While some deep learning-based forensic methods have achieved impressive results, they directly learn pixel-to-label mappings without fully exploiting the relationship between pixels in the feature space. To address such deficiency, we propose a Multi-view Pixel-wise Contrastive algorithm (MPC) for image forgery localization. Specifically, we first pre-train the feature extraction backbone network with a supervised contrastive loss to model pixel relationships in view of within-image, cross-scale and cross-modality. That is aimed at increasing intra-class compactness and inter-class separability. Then the localization head is fine-tuned using cross-entropy loss, resulting in a better forged pixel localizer. The MPC is trained on three different scale training datasets to make a comprehensive and fair comparison with existing image forgery localization algorithms. Extensive test results on over ten public datasets show that the proposed MPC achieves higher generalization performance and robustness than the state-of-the-arts. It is particularly noteworthy that our approach maintains a high level of localization accuracy under various post-processing combinations that approximate real-world scenarios, as well as when confronted with novel intelligent editing techniques. Finally, comprehensive and detailed ablation experiments demonstrate the reasonableness of MPC.
Zijie Lou, Gang Cao 0001, Kun Guo 0010, Lifang Yu, ShaoWei Weng
IEEE Trans. Inf. Forensics Secur.4
2025 A Copy-Move Forgery Detection Network Based on Selective Sampling Attention and Low-Cost Two-Step Self-Correlation Calculation
abstract
The commonly used standard convolutional layers cannot adaptively adjust the number and locations of sampling points according to the scales and shapes of tampered regions, which increases the difficulty of detecting images containing tampered regions of different sizes. Therefore, the selective sampling attention (SSA) is proposed to automatically learn the number and locations of sampling points as well as the weight of each sampling point within a certain context range of the input feature map through backpropagation, which can help the network better adapt to tampered regions of different scales and shapes. In addition, the self-correlation calculation (SCC), aiming at calculating the similarity between every two feature points in a feature map, necessarily incurs an expensive computational burden when used for high-resolution feature maps. To remedy the problem, the two-step SCC (TS-SCC) with low computation burden is proposed to pick out highly similar regions by means of the feature similarity obtained from low-resolution version of the input feature map, so that the high-resolution version merely needs to calculate the similarity between every two feature points within its high-similarity regions. Finally, to predict the edges and interiors of copy-move tampered regions more precisely, adaptive dual-branch feature fusion module is proposed to employ a lightweight multi-scale atrous convolutional module to adaptively fuse multi-level features before TS-SCC and the correlation features after TS-SCC, thereby improving the detection performance. Combining these three structures, a lightweight, fast, low-cost and high-precision CMFD network, ST-Net, is designed in this paper. Experimental results on four publicly available datasets verify that ST-Net outperforms several related CMFD networks in terms of detection accuracy, number of parameters, computational cost and inference time.
ShaoWei Weng, Lifang Yu, Li Li 0014
IEEE Trans. Multim.3
2025 DCM-Net: A Diffusion Model-Based Detection Network Integrating the Characteristics of Copy-Move Forgery
abstract
Essentially, directly introducing any object detection network to perform copy-move forgery detection (CMFD) inevitably leads to low detection accuracy. Therefore, DCM-Net, an object detection network dominated by diffusion model that incorporates the characteristics of copy-move forgery, is proposed in this paper for obviously enhancing CMFD performance. DCM-Net, as the first diffusion model-based CMFD network, has the following three improvements. Firstly, the high-similarity box padding strategy pads high-similarity boxes, rather than random boxes used in diffusion model, to ground truth boxes, better guiding subsequent dual-attention detection heads (DDHs) to focus more on high-similarity regions. Secondly, different from previous deep learning based CMFD networks that utilize self-correlation calculation to indiscriminately transform all classification features extracted from feature extraction into high-similarly features, an adaptive feature combination strategy is proposed to obtain the optimal feature transformation capable of achieving the best detection performance, enabling DDHs to more effectively distinguish source and target regions. Finally, to make detection heads have more accurate source/target localization and distinguishment, DDHs equipped with efficient multi-scale attention and contextual transformer, are proposed to generate tampered features fusing the entire precise spatial position information and rich contextual global information. The experimental results carried out on three publicly available datasets including USC-ISI, CoMoFoD, and COVERAGE, demonstrate that DCM-Net outperforms several advanced algorithms in terms of similarity detection ability and source/target differentiation ability.
ShaoWei Weng, Tanguo Zhu, Lifang Yu
IEEE Trans. Multim.4
2025 Adaptive PUPM-Based HEVC Video Steganography Balancing Embedding Performance and Security
abstract
For the prediction unit partition modes (PUPM)-based steganography, a mainstream branch of high efficiency video coding (HEVC) video steganography, striking a balance between embedding performance and security is very challenging. Including the$2\mathcal {N} \times 2\mathcal {N}$PUPMs having the maximum number of PUPMs into data embedding is indeed an effective way of enlarging the embedding capacity, but it necessarily causes a significant decline in security. Therefore, a multi-factor-involved cost function (MFICF) is proposed in this paper to evaluate the embedding cost for modifying each PUPM by comprehensively considering four different aspects affecting the embedding performance and security. With the assistance of MFICF, the 7-ary notational system is combined to use all the 7 types of PUPMs containing$2\mathcal {N} \times 2\mathcal {N}$for data embedding, thus enlarging the embedding capacity as well as enhancing the embedding efficiency. The syndrome-trellis code driven by MFICF, named CFSTC, is designed to preferentially select PUPMs with low embedding costs for data embedding, so that the embedding efficiency is largely enhanced. The security is effectively guaranteed by allocating a large embedding cost for modifying$2\mathcal {N} \times 2\mathcal {N}$to another type of PUPM. Finally, a lightweight convolutional neural network in combination with gated channel transformation, called GSCNet, is proposed to replace the in-loop filter in HEVC, further optimizing the visual distortion and bitrate increase caused by data embedding. Combining these components above, we design a PUPM-based steganography algorithm, GSAPM. Experimental results show that GSAPM effectively enhances the embedding performance while maintaining high security.
Lifang Yu, ShaoWei Weng, Dewang Chen
IEEE Trans. Multim.1
2024 A two-stream-network based steganalysis network: TSNet
Shiyao Sun, ShaoWei Weng, Lifang Yu
Expert Syst. Appl.4
2024 RCDD: Contrastive domain discrepancy with reliable steganalysis labeling for cover source mismatch
Lifang Yu, ShaoWei Weng, Mengfei Chen, Yunchao Wei
Expert Syst. Appl.1
2024 A deep steganalysis network combining source-supervised and target-unsupervised information for cover-source mismatch
Lifang Yu, Zhuwei Zhang, ShaoWei Weng, Gang Cao 0001
Expert Syst. Appl.1
2024 Transferable adversarial attack on image tampering localization
Gang Cao 0001, Haochen Zhu, Zijie Lou, Lifang Yu
J. Vis. Commun. Image Represent.5
2024 Discriminability-Aware Intermediate Domains for Mismatched Steganalysis
abstract
This letter proposes GDNet equipped with the generation of discriminative mixing regions (GDMR) and discriminability-aware local image mixing (DLIM), a steganalysis network aiming at alleviating significant accuracy degradation caused by cover-source mismatch (CSM), which pertains to the situation where source and target domains come from different distributions. GDNet guides a steganalyzer trained on the source domain to the target domain by mixing the source and target images at the region-level and pixel-level to construct a discriminative intermediate domain. On the one hand, GDMR designs an epoch-related region-level mixing ratio to control the size of the mixed region, and based on this ratio, selects the regions within the target image strongly related to the stego signal to participate in the generation of the intermediate domain, while suppressing other regions weakly related to the stego signal. On the other hand, DLIM utilizes the pixel-level mixing ratio to reduce the impact of the regions weakly related to the stego signal on the discriminability of the intermediate domain as the region-level mixing ratio increases, thereby increasing the diversity of the intermediate domain. Experimental results demonstrate that GDNet significantly outperforms existing methods across various CSM scenarios.
Yang Li 0205, Lifang Yu, ShaoWei Weng, Huawei Tian, Gang Cao 0001
IEEE Signal Process. Lett.2
2024 High-Precision Reversible Data Hiding Predictor: UCANet
abstract
Existing convolutional neural network-based reversible data hiding (RDH) predictors typically stack the standard convolution blocks with stride 1 for feature extraction, and keep the sizes of input and output feature maps unchanged through padding. This suggests that only a limited range of contextual spatial information is obtained. To remedy this problem above, a U-Net-like RDH predictor named UCANet is proposed in this paper to capture rich multi-scale contextual information by gradually downsampling feature maps. To fuse two feature maps at different levels along the channel dimension, we put forward the channel adaptive attention (CAA). By merely combining cheap pointwise convolution operations, CAA achieves the integration of non-linear and linear features as well as implicitly enhances channel dimensionality with low computational burden, thereby effectively enriching the expression of the channel information. The design of UCANet considers the characteristics of RDH from two aspects. On the one hand, instead of maxpooling or average pooling commonly used for downsampling, a stride-2 convolution block that can adaptively adjust the weights of convolution kernels and select useful information is utilized to downsample feature maps. On the other hand, UCANet removes the batch normalization layers to avoid their influence on the distribution of feature maps, which helps to strengthen the network's prediction capability. Extensive experiments also demonstrate that the proposed UCANet achieves better prediction performance, compared to several state-of-the-art methods.
Haiyang Rao, ShaoWei Weng, Lifang Yu, Li Li 0014, Gang Cao 0001
IEEE Signal Process. Lett.3
2024 Lightweight and High-Precision Network for Image Copy-Move Forgery Detection
abstract
The existing deep learning based copy-move forgery detection (DL-CMFD) networks focus on providing impressive detection accuracy for tampered regions of different sizes, but usually result in high computation cost and a large number of parameters. The focus of this letter is to propose LHCM-Net, a lightweight, high-precision DL-CMFD network, by integrating a low-cost self-correlation calculation (SCC) module (LCSCC), a gated feature fusion module (GFFM) and residual U-blocks (RSU) equipped with FasterNet blocks (FRSU). Considering that SCC, which calculates the similarity between every two pixels, inevitably leads to high computation cost, existing DL-CMFD networks have to carry out SCC only on low-resolution feature maps to reduce the computation cost. To make high-resolution features available for SCC without obviously introducing high computation cost, this letter proposes LCSCC to calculate the similarity between pixels with a certain distance. GFFM is presented to fuse feature maps of different spatial resolutions by adaptively adjusting their weights based on their respective characteristics, thereby fully integrating high-resolution and low-resolution features for subsequent LCSCC and obviously enhancing the detection accuracy. The FRSU allows LHCM-Net to keep the number of parameters (NP) and computation cost low by combining lightweight FasterNet blocks. The experimental results also demonstrate that LHCM-Net outperforms several existing DL-CMFD networks on three publicly available datasets in terms of detection accuracy, NP and computation cost.
ShaoWei Weng, Lifang Yu, Li Li 0014
IEEE Signal Process. Lett.3
2024 Universal Mismatched Steganalysis Equipped With Progressive Intermediate Domains
abstract
In general, cover source mismatch (CSM) inevitably leads to a significant decrease in detection accuracy in image steganalysis because the source and target domains have different distributions. To remedy this problem, a universal mismatched steganalyzer ISNet equipped with generated local mixing positions, a local feature-level mixup-related patchup (LFMP), and domain factors is proposed for both spatial and JPEG domains in this paper. Unlike existing deep steganalysis networks, which simply minimize the domain discrepancy between source and target to address the problem of CSM, and thus, cannot handle large domain discrepancy, IDGM based on LFMP generates diverse intermediate domains to bridge the two extreme domains so as to alleviate the decrease in detection accuracy caused by CSM. Moreover, ISNet enables the intermediate domain distribution to progressively transit from source to target by adjusting the domain factor sampled from Beta(α, 1), so that the classifier gradually adapted to target provides an improvement in discriminability on the target domain. The experimental results show that ISNet achieves the best performance in various CSM cases, compared with the most advanced deep learning-based steganalysis network.
ShaoWei Weng, Zhuwei Zhang, Lifang Yu, Gang Cao 0001
IEEE Signal Process. Lett.3
2023 Adaptive multi-teacher softened relational knowledge distillation framework for payload mismatch in image steganalysis
Lifang Yu, ShaoWei Weng, Huawei Tian
J. Vis. Commun. Image Represent.1
2023 An Image Steganoganalyzer With Comprehensive Detection Performance
abstract
Effectively enhancing the weak stego signal while striking the balance among three evaluation metrics, i.e., detection accuracy, time cost, as well as the number of parameters (NP) is indeed a huge challenge for existing deep learning-based steganalysis detectors. In this letter, a novel steganalysis detector called GFS-Net is proposed, aiming at enhancing the stego signal as much as possible while balancing the three metrics to obtain comprehensive detection performance. In preprocessing, combining highly lightweight gated channel transformation with a pointwise convolution layer used for enlarging the number of channels enriches the expression of the stego signal by promoting cooperation among enlarged channels, thereby significantly improving the signal-to-noise ratio while avoiding occupying a large NP. Moreover, two FasterNet blocks equipped with partial convolution having a small NP, rather than residual blocks, are applied to the last two layers of feature extraction to efficiently extract the stego signal by reducing the calculation of similar features, so that the computational cost and NP are saved. Finally, compared to using the global average pooling (GAP) alone, the stylepooling jointly utilizing the global standard deviation pooling and GAP helps the subsequent fully connected better identify weak stego signal, and thus improves detection accuracy. By means of the above three perspectives, GFS-Net with only 0.13 M parameters is obtained. Experimental results also demonstrate that GFS-Net achieves higher detection accuracy and lower computational cost than state-of-the-art steganalysis detectors.
ShaoWei Weng, Lifang Yu, Wei Chen 0079
IEEE Signal Process. Lett.3
2023 Fast SwT-Based Deep Steganalysis Network for Arbitrary-Sized Images
abstract
In this letter, a novel deep steganalysis network called SwT-SN is proposed by integrating directional difference adaptive combination (DDAC) followed by three residual blocks, convolutional spatial pyramid pooling equipped with size-independent detector (CSPP-SID) as well as a two-part of Swin transformer (SwT) structure suitable for steganalysis, aiming at enhancing the detection accuracy for arbitrary-sized images while significantly reducing training and test cost. In comparison to simply exploiting DDAC as preprocessing, DDAC + the residual structure can better improve the signal-to-noise ratio of the residual maps by suppressing the image content. In addition, CSPP-SID is innovatively proposed to convert feature maps of any size into feature vectors with fixed dimension, helping SwT-SN achieving high detection accuracy for arbitrary-sized images. Finally, a two-part structure of SwT with a fixed number of patches is firstly designed for image steganalysis to greatly reduce training cost by calculating the multi-head self-attention mechanism in the shifted window. Extensive experiments on two benckmark databases verify that SwT-SN has higher detection accuracy and shorter training cost compared to two prior state-of-the-art networks.
ShaoWei Weng, Shiyao Sun, Lifang Yu
IEEE Signal Process. Lett.3
2022 Lightweight and Effective Deep Image Steganalysis Network
abstract
In this letter, a lightweight and effective deep steganalysis network (DSN) with less than 400,000 parameters, called LWENet, is proposed, which focuses on increasing the performance as well as significantly reducing the number of parameters (NP) from three perspectives. Firstly, in the preprocessing part, several lightweight bottleneck residual blocks are combined into the spatial rich model filters to improve the signal-to-noise ratio of stego signals while slightly increasing NP, thereby improving the subsequent performance. Secondly, a depthwise separable convolution layer is exploited at the end of the feature extraction part to largely reduce NP and increase the performance by capturing salient correlations while ignoring trivial ones among feature maps. Finally, to keep LWENet lightweight, we have to select only one fully connected (FC) layer. Simultaneously, multi-view global pooling is employed prior to the FC layer to yield multi-view features and further improve the detection performance. Extensive experiments demonstrate that our network achieves better performance than several state-of-the-art DSNs.
ShaoWei Weng, Mengfei Chen, Lifang Yu, Shiyao Sun
IEEE Signal Process. Lett.3
2018 Acceleration of histogram-based contrast enhancement via selective downsampling
abstract
The authors propose a general framework to accelerate the universal histogram‐based image contrast enhancement (CE) algorithms. Both spatial and grey‐level selective downsampling of digital images are adopted to decrease computational cost, while the visual quality of enhanced images is still preserved and without apparent degradation. Mapping function calibration is proposed to reconstruct the pixel mapping on the grey levels missed by downsampling. As two case studies, the accelerations of histogram equalisation (HE) and the state‐of‐the‐art global CE algorithm, i.e. spatial mutual information and PageRank (SMIRANK), are presented in detail. Both quantitative and qualitative assessment results have verified the effectiveness of their proposed CE acceleration framework. In typical tests, the computational efficiencies of HE and SMIRANK have been increased by about 3.9 and 13.5 times, respectively.
Gang Cao 0001, Huawei Tian, Lifang Yu, Xianglin Huang, Yongbin Wang
IET Image Process.3
2014 Attacking contrast enhancement forensics in digital images
Gang Cao 0001, Yao Zhao 0001, Huawei Tian, Lifang Yu
Sci. China Inf. Sci.5
2014 A channel selection rule for YASS
Lifang Yu, Yao Zhao 0001, Gang Cao 0001
Sci. China Inf. Sci.1
2010 Forensic detection of median filtering in digital images
abstract
In digital image forensics, prior works are prone to the detection of malicious tampering. However, there is also a need for developing techniques to identify general content-preserved manipulations, which are employed to conceal tampering trails frequently. In this paper, we propose a blind forensic algorithm to detect median filtering (MF), which is applied extensively for signal denoising and digital image enhancement. The probability of zero values on the first order difference map in texture regions can serve as MF statistical fingerprint, which distinguishes MF from other operations. Since anti-forensic techniques enjoy utilizing MF to attack the linearity assumption of existing forensics algorithms, blind detection of the non-linear MF becomes especially significant. Both theoretically reasoning and experimental results verify the effectiveness of our proposed MF forensics scheme.
Gang Cao 0001, Yao Zhao 0001, Lifang Yu, Huawei Tian
ICME4
2010 A high-performance YASS-like scheme using randomized big-blocks
abstract
Randomly selecting 8 × 8 host blocks in big-blocks for data embedding, YASS, a recently developed advanced stegano-graphic scheme makes these blocks not coincident with the 8×8 grids used in JPEG compression. As a result, it effectively invalidates the self-calibration technique used in modern steganaly-sis. However, the randomization is not sufficient enough, i.e., some positions in an image are possible to hold host blocks and some are definitely not. Based on this observation, the newly developed specific steganalyzer can effectively defeat YASS. In this paper, a new steganographic scheme is presented. Through randomizing the size and position of each big-block, our improved steganographic method makes almost every position possible to hold a host block, which has been verified by our statistical analysis. Consequently, the proposed scheme can survive the attack made by the specific steganalyzer. Experimental results have demonstrated that the detection rate achieved by the specific steganalyzer on our proposed method is less than 58%, while that on YASS is about 95% and above.
Lifang Yu, Yao Zhao 0001, Yun Q. Shi 0001
ICME1
2009 PM1 steganography in JPEG images using genetic algorithm
Lifang Yu, Yao Zhao 0001, Zhenfeng Zhu
Soft Comput.1
2008 A High Capacity Steganographic Algorithm in Color Images
Yao Zhao 0001, Lifang Yu
IWDW4