Guopu Zhu

dblp:77/5845 · DBLP profile ↗
← Back
74ranked-venue papers
7as first author
45since 2021 · last 2026
0000-0001-7956-5343ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 37 · 2 first-author · 22 since 2021Security and privacy · 14 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 13 · 3 first-author · 8 since 2021Computer networks · 5 · 4 since 2021Databases, data management, data science and information retrieval · 5 · 4 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 TraceMark-LDM: Authenticatable watermarking for latent diffusion models via binary-guided rearrangement
Zhangyi Shen, Ye Yao 0003, Feng Ding 0007, Guopu Zhu, Weizhi Meng 0001
Expert Syst. Appl.5
2026 CL-DRW: Curriculum Learning-Based Deep Robust Watermarking for Social Networks
abstract
The development of the internet has greatly facilitated the transmission of images over social networks, while also triggering serious copyright issues. Deep robust watermarking serves as a crucial technique for image copyright protection. However, the image distortions caused by Social Network Transmission Operations (SNTOs) make existing deep robust watermarking methods fragile in real-world social network scenarios. To address this, we propose a Curriculum Learning-based Deep Robust Watermarking method, called CL-DRW, to generate watermarks that can be resilient to SNTOs. Specifically, we develop a watermarking model constructed with an invertible neural network and present a multi-stage training framework based on curriculum learning to train it effectively. We incrementally introduce noise attacks based on their disruptive impact on the watermark, from weak to strong, thereby enabling our model to build robustness against SNTOs gradually. Additionally, we design an SNTOs simulation noise layer, which is built upon a transformer-based deep network and incorporates differentiable JPEG, to simulate the black-box distortions caused by SNTOs. Extensive experiments indicate that our proposed CL-DRW outperforms state-of-the-art deep watermarking methods in terms of robustness against real-world social network transmission operations. Source code is available at https://github.com/yingshuai-zhao/CL-DRW.
Yingshuai Zhao, Guopu Zhu, Jiantao Zhou 0001, Xiaolong Li 0001, Hongli Zhang 0001, Xinpeng Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.2
2026 Test-Time Adaptation for Detecting Image Inpainting Forgeries
abstract
The rapid development of deep learning-based image inpainting poses serious challenges to image authenticity. As inpainting methods continue to evolve, the inpainted images exhibit extremely high visual fidelity, presenting recognition difficulties to the forgery detection model due to differences in operational mode and forgery traces among methods. In particular, the detection performance tends to drop significantly in the testing phase when the test samples differ from the training data. To address this issue, we propose a test-time adaptive detection framework for image inpainting forgeries. First, we propose an image gradient-based metric that quantifies model uncertainty and orchestrates the entire adaptation process. Integrating this metric with sample-specific batch normalization (BN) statistics enhances the ability of pretrained models in the inference stage. Second, we introduce a cross-attention module as a side-tuning module, enabling the model to adapt dynamically to reliable test samples without altering the backbone network. To validate the effectiveness of the proposed method, we construct a dataset comprising synthetic images of multiple inpainting methods and design experiments under two scenarios of distributional bias. The results demonstrate that our proposed framework outperforms the existing baseline method, enhancing the adaptability and detection performance of the forgery detection model in dynamic environments.
Guopu Zhu, Hongli Zhang 0001, Xinpeng Zhang 0001, Yicong Zhou, Ligang Wu 0001
IEEE Trans. Cybern.2
2026 SHL-Net: Semantics-Enhanced Network for Localizing Harmonized Image Splicing
Xiwen Fu, Guopu Zhu, Hongli Zhang 0001, Jiwu Huang, Tao Xiang 0001, Yicong Zhou, Ligang Wu 0001
IEEE Trans. Inf. Forensics Secur.2
2026 MU-MIA: Machine Unlearning for Membership Inference Attacks
abstract
The widespread deployment of deep learning models across various applications has raised significant concerns regarding data privacy. Membership inference attacks (MIAs), a major privacy threat, aim to determine whether a specific sample is used during model training, thereby posing significant risks to sensitive information. Most existing MIA methods rely on the model’s final state output, overlooking the process by which the model memorizes training samples. To better exploit model memorization for MIAs, we propose a novel attack method called machine unlearning-based membership inference attack (MU-MIA). The proposed method introduces machine unlearning to incrementally reduce the model’s memorization of specific samples, generating a forgetting trajectory for each sample. The forgetting trajectory is composed of temporal variations in different metrics of the sample during machine unlearning. To distinguish member from non-member samples, we design a BiLSTM-based binary classifier with attention, which captures discriminative temporal patterns within each forgetting trajectory. Moreover, the machine unlearning phase of our attack is conducted under a zero-shot setting, which eliminates the need for any real data during the unlearning process, thereby improving the practicality and generalizability of the attack. We evaluate the proposed MIA method across different datasets and model architectures, and the comparative experimental results show that our method outperforms existing baseline attack methods.
Hongming Yang, Guopu Zhu, Xinpeng Zhang 0001, Tao Xiang 0001
IEEE Trans. Inf. Forensics Secur.2
2026 Information Disclosure Risk of Thumbnail-Preserving Encryption
abstract
With the rapid growth of cloud services, the storage of images in cloud environments requires secure and effective data encryption methods. Many thumbnail-preserving encryption (TPE) methods have thus been proposed to balance privacy and usability of image data. However, the exposure of thumbnail information in TPE methods may introduce privacy leakage risk, and a systematic evaluation of their security has not yet been conducted. In this paper, we propose a new Mamba-Transformer cooperation Network (MTNet) to recover the original images from the limited exposed thumbnail information, highlighting the information disclosure problem in TPE. Specifically, the core model component integrates a Mamba block and a Transformer block, which employ the powerful capabilities of the Mamba for wide field dependency modeling and the Transformer for effective channel interaction. Besides, the cascade architecture incorporates an intermediate output that provides supplementary information and achieves multilevel supervision, thereby improving the quality of the final output. Finally, to better utilize the subtle details in different levels, we propose a multi-scale fusion module that adaptively integrates features from various stages of the encoding process. The experimental results achieved by our proposed MTNet reveal that the privacy risk associated with TPE is significantly underestimated and more robust defense mechanisms are required. Source code is available athttps://github.com/HITLiXincodes/MTNet.
Xin Li 0154, Guopu Zhu, Hongli Zhang 0001, Tao Xiang 0001, Xiangyang Luo 0001, Sam Kwong
IEEE Trans. Multim.2
2026 DCHVF-GAN: Synthesizing Adversarial DeepFakes with High Visual Fidelity by Multimodality Fusion
abstract
DeepFake, an AI-driven face-swapping technique, has been weaponized to spread disinformation. In response, researchers have developed forensic detectors to identify such manipulations. To circumvent these defenses, a growing body of work now focuses on generating adversarial samples—carefully perturbed forgeries designed to deceive detection tools. However, most existing adversarial generation methods sacrifice image quality to achieve undetectability, introducing perceptible artifacts that ironically make them more detectable under human scrutiny. To address this limitation, we propose a novel spectral fusion approach to multimodally synthesize forgery traces from authentic facial images. Unlike traditional noise injection methods, our technique integrates diffusion-based noise during image preprocessing, embedding perturbations in the forward process of a diffusion model. This approach not only deceives forensic detectors more effectively but also preserves high visual fidelity. Through extensive experiments, our method achieves state-of-the-art DeepFake anti-forensic performance while preserving high visual fidelity, ensuring that the adversarial samples remain indistinguishable from real images.
Feng Ding 0007, Xinan He, Rensheng Kuang, Mengyao Xiao, Xiaogang Zhu 0003, Guopu Zhu, Pradeep K. Atrey
ACM Trans. Multim. Comput. Commun. Appl.6
2025 VLForgery Face Triad: Detection, Localization and Attribution via Multimodal Large Language Models
abstract
Faces synthesized by diffusion models (DMs) with high-quality and controllable attributes pose a significant challenge for Deepfake detection. Most state-of-the-art detectors only yield a binary decision, incapable of forgery localization, attribution of forgery methods, and providing analysis on the cause of forgeries. In this work, we integrate Multimodal Large Language Models (MLLMs) within DM-based face forensics, and propose a fine-grained analysis triad framework called VLForgery, that can 1) predict falsified facial images; 2) locate the falsified face regions subjected to partial synthesis; and 3) attribute the synthesis with specific generators. To achieve the above goals, we introduce VLF (Visual Language Forensics), a novel and diverse synthesis face dataset designed to facilitate rich interactions between `Visual' and `Language' modalities in MLLMs. Additionally, we propose an extrinsic knowledge-guided description method, termed EkCot, which leverages knowledge from the image generation pipeline to enable MLLMs to quickly capture image content. Furthermore, we introduce a low-level vision comparison pipeline designed to identify differential features between real and fake that MLLMs can inherently understand. These features are then incorporated into EkCot, enhancing its ability to analyze forgeries in a structured manner, following the sequence of detection, localization, and attribution. Extensive experiments demonstrate that VLForgery outperforms other state-of-the-art forensic approaches in detection accuracy, with additional potential for falsified region localization and attribution analysis.
Xinan He, Yue Zhou 0008, Bing Fan, Guopu Zhu, Feng Ding 0007
NeurIPS5
2025 Leading Attackers Astray: Mitigating Link Flooding Attacks through Stub Node Relocation and Insertion
abstract
Link Flooding Attacks (LFA) exploit network topology knowledge to disrupt connectivity by targeting critical links and nodes. Existing defenses often presuppose an attacker with complete topological awareness and overlook the concentration of attack traffic on specific routers. Furthermore, many countermeasures rely on SDN, which can suffer from performance degradation due to the limited packet processing capabilities of switches. To address these issues, we introduce the GateLFA attacker model, which assumes that attackers lack complete topology knowledge and guide their attacks based on traffic density analysis. We propose the EqualFlow algorithm, which utilizes stub node relocation and insertion to minimize adversarial impact, balance attack traffic, and reduce defense costs. Additionally, we present the Network Topology Obfuscation System, leveraging XDP for high-speed packet processing at the network boundary to overcome the performance challenges of SDN-based solutions. Our experimental results demonstrate that EqualFlow computes high-quality virtual topologies, outperforming existing algorithms across small, medium, and large-scale networks. Moreover, the Network Topology Obfuscation System effectively disrupts prominent topology probing tools through explicit information interference at a 10 Gbps line rate. For implicit interference, the system increases the packet rate of typical traceroute probes by approximately 17% compared to traffic control methods. This research provides an efficient and practical solution for defending against LFA.
Bin Wang 0098, Kehong Liu, Yu Zhang 0036, Guopu Zhu, Binxing Fang
SMC5
2025 Multi-Level Feature Fusion Network for Shadow Removal Detection
abstract
By now, many works have been done on shadow removal for image manipulation. As a result, detecting shadow removal has become a critical part to reveal the traces of image manipulation. However, there are only a few works conducted on shadow removal detection, and these works cannot accurately localize the image regions where the shadows have been removed. In this paper, we present a novel model called Multi-level Feature Fusion Network (MFF-Net) for shadow removal detection. MFF-Net consists of two parts: a dual-branch feature extraction encoder and a dense prediction decoder. The encoder anchors the approximate position of the manipulated regions, while the decoder progressively fills in the details of the estimated shadow masks by integrating multi-level information. In the encoder part, a global modeling branch is constructed to capture long-range dependencies, while a local feature extraction branch is designed to extract local structural information. The features extracted by these two branches are integrated using a feature fusion module. In the decoder part, a multi-scale feature upsampling module is proposed to upsample the input features and integrate them with the low-level features obtained from the encoder part. Meanwhile, the cross attention mechanism is introduced to guide the multi-level feature fusion process. Finally, the features of different resolutions are employed to estimate the shadow masks in a coarse-to-fine manner. Extensive experiments on shadow removal detection demonstrate the superiority of MFF-Net over the state-of-the-art methods. The source code of MFF-Net is publicly available athttps://github.com/HITFuxiwen/MFF-Net.
Xiwen Fu, Guopu Zhu, Hongli Zhang 0001, Xinpeng Zhang 0001, Anthony Tung Shuen Ho, Sam Kwong
IEEE Trans. Circuits Syst. Video Technol.2
2025 FGMIA: Feature-Guided Model Inversion Attacks Against Face Recognition Models
abstract
Model Inversion Attacks (MIAs) against face recognition systems aim to reconstruct facial images of specific individuals from the recognition models. Existing MIA approaches commonly optimize the latent variables of Generative Adversarial Networks (GANs) iteratively, which can result in non-smooth optimizations due to the complexity and entanglement of latent space. Furthermore, the optimization guided by the target model’s gradients may generate high-confidence images with poor perceptual similarity to the target class. This paper introduces a novel perspective by reformulating the inversion attack as a conditional data distribution learning task. Based on this, we propose a Feature-Guided Model Inversion Attack (FGMIA), which learns the facial data distribution and integrates feature guidance as a conditional signal. Specifically, we treat the deconstructed target model as a feature encoder, which provides guidance during the training of a specialized feature-guided diffusion model. During the attack, feature encodings implicit in the target model are extracted and utilized to guide the reconstruction of private data. Extensive experiments demonstrate that FGMIA accurately reconstructs private data from face recognition models and significantly improves evaluation accuracy and perceptual similarity compared to state-of-the-art methods while maintaining comparable target confidence scores. Our code is available at https://github.com/MMCTTT/FGMIA_codes.
Shen Wang 0004, Guopu Zhu, Zhaoyang Zhang 0002, Jiwu Huang
IEEE Trans. Inf. Forensics Secur.3
2025 Generating Higher-Quality Anti-Forensics DeepFakes with Adversarial Sharpening Mask
abstract
DeepFake, an AI technology that can automatically synthesize facial forgeries, has recently attracted worldwide attention. While DeepFakes can be entertaining, they can also be used to spread falsified information or be weaponized as cognition warfare. Forensic researchers have been dedicated to designing defensive algorithms to combat such disinformation. However, attacking technologies have been developed to make DeepFake products more aggressive. For example, by launching anti-forensics and adversarial attacks, DeepFakes can be disguised as authentic media to evade forensic detectors. However, such manipulations often sacrifice image quality for satisfactory undetectability. To address this issue, we propose a method to generate a novel adversarial sharpening mask for launching black-box anti-forensics attacks. Unlike many existing methods, our approach injects perturbations that allow DeepFakes to achieve high anti-forensics performance while maintaining pleasant sharpening visual effects. Experimental evaluations demonstrate that our method successfully disrupts state-of-the-art DeepFake detectors. Moreover, compared to images processed by existing DeepFake anti-forensics methods, our method’s quality of anti-forensics DeepFakes rendered is significantly improved. Our code is available at https://github.com/fb-reps/HQ-AF_GAN .
Bing Fan, Feng Ding 0007, Guopu Zhu, Jiwu Huang, Sam Kwong, Pradeep K. Atrey, Siwei Lyu
ACM Trans. Multim. Comput. Commun. Appl.3
2024 CoSTA: Co-training spatial-temporal attention for blind video quality assessment
Fengchuang Xing, Yuan-Gen Wang, Weixuan Tang 0004, Guopu Zhu, Sam Kwong
Expert Syst. Appl.4
2024 Exploring trajectory embedding via spatial-temporal propagation for dynamic region representations
Hongli Zhang 0001, Guopu Zhu, Haotian Guan, Sam Kwong
Inf. Sci.3
2024 An efficient cheating-detectable secret image sharing scheme with smaller share sizes
Zuquan Liu, Guopu Zhu, Yu Zhang 0036, Hongli Zhang 0001, Sam Kwong
J. Inf. Secur. Appl.2
2024 Object-attentional untargeted adversarial attack
Yuan-Gen Wang, Guopu Zhu
J. Inf. Secur. Appl.3
2024 Disrupting Anti-Spoofing Systems by Images of Consistent Identity
abstract
Face anti-spoofing aims to distinguish between live and spoof images to ensure the authenticity and reliability of face recognition. Methods based on convolutional neural networks surpass prior approaches in accuracy but remain vulnerable to adversarial attacks. Traditional vulnerability disclosure relies on image quality metrics, which lack crucial identity details for face recognition in deceptive images. Thus, we introduce a novel framework that creates identity-consistent adversarial samples for face anti-spoofing. We also redefine image quality assessment for anti-spoofing by using detection rates from facial recognizers rather than conventional metrics. Inspired by style transfer, our generator incorporates an LKA module to enhance performance. An identity recognition module ensures consistency within synthesized images, capturing the essence of identity across live and spoof visuals. Our approach outperforms most adversarial tactics on anti-spoofing detectors while retaining a high identity recognition rate.
Feng Ding 0007, Yue Zhou 0008, Guopu Zhu
IEEE Signal Process. Lett.5
2024 Deep Reverse Attack on SIFT Features With a Coarse-to-Fine GAN Model
abstract
Recently, it has been shown that adversaries can reconstruct images from SIFT features through reverse attacks. However, the images reconstructed by existing reverse attack methods suffer from information loss and are unable to sufficiently reveal the private contents of the original images. In this paper, a two-stage deep reverse attack model called Coarse-to-Fine Generative Adversarial Network (CFGAN) is proposed to more deeply explore the information in SIFT features and further demonstrate the risk of privacy leakage associated with SIFT features. Specifically, the proposed model consists of two sub-networks, namely coarse net and fine net. The coarse net is developed to restore coarse images using SIFT features, while the fine net is responsible for refining the coarse images to obtain better reconstruction results. To effectively leverage the information contained in SIFT features, an efficient fusion strategy based on the AdaIN operation is designed in the fine net. Additionally, we introduce a new loss function called sift loss that enhances the color fidelity of reconstructed images. Extensive experiments conducted on various datasets verify that the proposed CFGAN performs favorably against state-of-the-art methods. The reconstructed images exhibit better visual quality, less texture distortion, and higher color fidelity. Source code is available at https://github.com/HITLiXincodes/CFGAN.
Xin Li 0154, Guopu Zhu, Shen Wang 0004, Yicong Zhou, Xinpeng Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.2
2024 Impulse Noise Image Restoration Using Nonconvex Variational Model and Difference of Convex Functions Algorithm
abstract
In this article, the problem of impulse noise image restoration is investigated. A typical way to eliminate impulse noise is to use an$L_{1}$norm data fitting term and a total variation (TV) regularization. However, a convex optimization method designed in this way always yields staircase artifacts. In addition, the$L_{1}$norm fitting term tends to penalize corrupted and noise-free data equally, and is not robust to impulse noise. In order to seek a solution of high recovery quality, we propose a new variational model that integrates the nonconvex data fitting term and the nonconvex TV regularization. The usage of the nonconvex TV regularizer helps to eliminate the staircase artifacts. Moreover, the nonconvex fidelity term can detect impulse noise effectively in the way that it is enforced when the observed data is slightly corrupted, while is less enforced for the severely corrupted pixels. A novel difference of convex functions algorithm is also developed to solve the variational model. Using the variational method, we prove that the sequence generated by the proposed algorithm converges to a stationary point of the nonconvex objective function. Experimental results show that our proposed algorithm is efficient and compares favorably with state-of-the-art methods.
Benxin Zhang, Guopu Zhu, Hongli Zhang 0001, Yicong Zhou, Sam Kwong
IEEE Trans. Cybern.2
2024 Adversarial Perturbation Prediction for Real-Time Protection of Speech Privacy
abstract
The widespread collection and analysis of private speech signals have become increasingly prevalent, raising significant privacy concerns. To protect speech signals from unauthorized analysis, adversarial attack methods for deceiving speaker recognition models have been proposed. While a few of these methods are specifically designed for real-time protection of speech signals, they introduce significant delays that can severely impact speech communication when applied to streaming speech data. In this paper, we present a novel approach that aims to offer real-time protection for speech signals without delays. By utilizing observed data only, we generate initial adversarial seed perturbations and refine them to obtain the necessary adversarial perturbations predicted for adjacent unobserved signals. This refinement process is conducted via a proposed model called PAPG. On the basis of perturbation prediction, we develop a streaming audio processing framework that generates perturbations in synchronization with the playback of the original signal, effectively eliminating delays. The experimental results demonstrate that under the proposed attack, the average Top-1 accuracy of various advanced speaker recognition methods is reduced by 89%, and the average equal error rate (EER) increases to 36%. Remarkably, these results are achieved without delays while maintaining superior perceptual quality.
Zhaoyang Zhang 0002, Shen Wang 0004, Guopu Zhu, Dechen Zhan, Jiwu Huang
IEEE Trans. Inf. Forensics Secur.3
2023 Secure and asynchronous filtering for piecewise homogeneous Markov jump systems with quantization and round-Robin communication
Guopu Zhu, Peng Shi 0001
Inf. Sci.2
2023 CNN-Transformer Based Generative Adversarial Network for Copy-Move Source/ Target Distinguishment
abstract
Copy-move forgery can be used for hiding certain objects or duplicating meaningful objects in images. Although copy-move forgery detection has been studied extensively in recent years, it is still a challenging task to distinguish between the source and the target regions in copy-move forgery images. In this paper, a convolutional neural network-transformer based generative adversarial network (CNN-T GAN) is proposed to distinguish the source and target regions in a copy-move forged image. A generator is first utilized to generate a mask that is similar to the groundtruth mask. Then, a discriminator is trained to discriminate the true image pairs from the false ones. When the discriminator cannot discriminate the true/false image pairs accurately, the generator can be used to obtain the final localization maps of copy-move forgery. In the generator, convolutional neural network (CNN) and transformer are exploited to extract the local features and global representations in copy-move forgery images, respectively. In addition, feature coupling layers are designed to integrate the features in CNN branch and transformer branch in an interactive way. Finally, a new Pearson correlation layer is introduced to match the similarity features in source and target regions, which can improve the performance of copy-move forgery localization, especially the localization performance on source regions. To the best of our knowledge, this is the first work to utilize transformer for feature extraction in copy-move forgery localization. The proposed method can not only detect the copy-move regions, but also distinguish the source and target regions. Extensive experimental results on several commonly used copy-move datasets have shown that the proposed method outperforms the state-of-the-art methods for copy-move detection.
Yulan Zhang, Guopu Zhu, Xiangyang Luo 0001, Yicong Zhou, Hongli Zhang 0001, Ligang Wu 0001
IEEE Trans. Circuits Syst. Video Technol.2
2023 ExS-GAN: Synthesizing Anti-Forensics Images via Extra Supervised GAN
abstract
So far, researchers have proposed many forensics tools to protect the authenticity and integrity of digital information. However, with the explosive development of machine learning, existing forensics tools may compromise against new attacks anytime. Hence, it is always necessary to investigate anti-forensics to expose the vulnerabilities of forensics tools. It is beneficial for forensics researchers to develop new tools as countermeasures. To date, one of the potential threats is the generative adversarial networks (GANs), which could be employed for fabricating or forging falsified data to attack forensics detectors. In this article, we investigate the anti-forensics performance of GANs by proposing a novel model, the ExS-GAN, which features an extra supervision system. After training, the proposed model could launch anti-forensics attacks on various manipulated images. Evaluated by experiments, the proposed method could achieve high anti-forensics performance while preserving satisfying image quality. We also justify the proposed extra supervision via an ablation study.
Feng Ding 0007, Zhangyi Shen, Guopu Zhu, Sam Kwong, Yicong Zhou, Siwei Lyu
IEEE Trans. Cybern.3
2023 Query-Efficient Adversarial Attack With Low Perturbation Against End-to-End Speech Recognition Systems
abstract
With the widespread use of automated speech recognition (ASR) systems in modern consumer devices, attack against ASR systems have become an attractive topic in recent years. Although related white-box attack methods have achieved remarkable success in fooling neural networks, they rely heavily on obtaining full access to the details of the target models. Due to the lack of prior knowledge of the victim model and the inefficiency in utilizing query results, most of the existing black-box attack methods for ASR systems are query-intensive. In this paper, we propose a new black-box attack called the Monte Carlo gradient sign attack (MGSA) to generate adversarial audio samples with substantially fewer queries. It updates an original sample based on the elements obtained by a Monte Carlo tree search. We attribute its high query efficiency to the effective utilization of the dominant gradient phenomenon, which refers to the fact that only a few elements of each origin sample have significant effect on the output of ASR systems. Extensive experiments are performed to evaluate the efficiency of MGSA and the stealthiness of the generated adversarial examples on the DeepSpeech system. The experimental results show that MGSA achieves 98% and 99% attack success rates on the LibriSpeech and Mozilla Common Voice datasets, respectively. Compared with the state-of-the-art methods, the average number of queries is reduced by 27% and the signal-to-noise ratio is increased by 31%.
Shen Wang 0004, Zhaoyang Zhang 0002, Guopu Zhu, Xinpeng Zhang 0001, Yicong Zhou, Jiwu Huang
IEEE Trans. Inf. Forensics Secur.3
2023 Estimating the Secret Key of Spread Spectrum Watermarking Based on Equivalent Keys
abstract
The security of spread spectrum (SS) watermarking largely depends on the difficulty of estimating its secret key. Some estimators have been proposed to estimate the secret key in the known-message attack (KMA) scenario. However, the estimation accuracies of existing estimators are not satisfactory when the number of observations is not large enough. Currently, it is still a challenging and open problem to design more effective estimators. In this paper, we propose an equivalent keys (EK)-based estimator to estimate the secret key for both the traditional and more secure SS watermarking methods. Equivalent keys form an equivalent region, which is the intersection of a unit hypersphere and a hypercone. According to the Monte Carlo simulation, we find that the secret key can be estimated by adding up the equivalent keys uniformly sampled from the equivalent region. Thus, the proposed estimator selects equivalent keys from randomly-generated vectors by exploiting the pairs of watermarked signals and their embedded messages. A theoretical analysis is performed for the proposed estimator to evaluate the estimation accuracy. Experimental results verify the theoretical analysis and show the superiority of the proposed estimator over existing estimation methods. Furthermore, this paper also shows the insecurity of the more secure SS watermarking methods in the KMA scenario from a practical perspective for the first time.
Jinkun You, Yuan-Gen Wang, Guopu Zhu, Ligang Wu 0001, Hongli Zhang 0001, Sam Kwong
IEEE Trans. Multim.3
2022 Starvqa: Space-Time Attention for Video Quality Assessment
abstract
Transformer based on self-attention mechanism is blooming in computer vision nowadays. However, its application to video quality assessment (VQA) has not been reported. Evaluating the quality of in-the-wild videos is challenging due to the unknown of pristine reference and shooting distortion. This paper presents a novel space-time attention network for the VQA problem, named StarVQA. StarVQA builds a Transformer by alternately concatenating the divided space-time attention. To adapt the Transformer architecture for training, StarVQA designs a vectorized regression loss by encoding the mean opinion score (MOS) to the probability vector and embedding a special vectorized label token as the learnable variable. To capture the long-range spatiotemporal dependencies of a video sequence, StarVQA encodes the space-time position information of each patch to the input of the Transformer. Various experiments are conducted on the de-facto in-the-wild video datasets, including LIVE-VQC, KoNViD-1k, LSVQ, and LSVQ-1080p. Experimental results demonstrate the superiority of StarVQA over the state-of-the-art. The source code is available at https://github.com/GZHU-DVL/StarVQA.
Fengchuang Xing, Yuan-Gen Wang, Hanpin Wang, Leida Li, Guopu Zhu
ICIP5
2022 Reversible Data Hiding for Color Images Based on Adaptive 3D Prediction-Error Expansion and Double Deep Q-Network
abstract
Reversible data hiding (RDH) for color images has attracted increasing attention in recent years. Due to its effective utilization of the correlation between prediction errors, high-dimensional prediction-error expansion (PEE) can achieve much better performance for color image RDH than low-dimensional PEE. However, existing studies only focus on high-dimensional PEE with nonadaptive embedding. To further improve the embedding performance for color images, we propose a novel three-dimensional PEE method that is adaptive to image content. Double deep Q-network (DDQN), introduced to RDH for the first time, is adopted to find the optimal mapping paths for PEE. In addition, an action selection scheme is presented for DDQN to efficiently find the reversible mapping paths. Extensive experiments show that the proposed method outperforms existing color image RDH methods in image quality.
Guopu Zhu, Hongli Zhang 0001, Yicong Zhou, Xiangyang Luo 0001, Ligang Wu 0001
IEEE Trans. Circuits Syst. Video Technol.2
2022 Forensic Analysis of JPEG-Domain Enhanced Images via Coefficient Likelihood Modeling
abstract
JPEG-domain enhancement improves the visual quality of JPEG images by directly manipulating the decoded DCT (discrete cosine transform) coefficients, which inevitably leads to mixed compression and enhancement artifacts. Existing forensic methods that merely consider JPEG artifacts are unsuitable to address such mixed artifacts and hence suffer a considerable performance decline in compression parameter estimation and lack the ability to estimate the enhancement parameter. This work attempts to explore the characterization of the mixed artifacts, and to further estimate both the enhancement and compression parameters of JPEG-domain enhanced images. First, a statistical likelihood function is proposed to characterize the periodicity of DCT coefficients, which can measure how well an enhanced image is de-enhanced back to its JPEG compressed version given the compression and enhancement parameters. The proposed likelihood function reaches its maximum if the parameters match their true values. Then, a forensic method of enhancement detection and parameter estimation is developed based on the proposed likelihood function for two kinds of classical JPEG-domain enhancement. Specifically, JPEG-domain enhanced images are detected by thresholding a scalar feature computed upon the likelihoods, and the enhancement and compression parameters are estimated by locating the maximal likelihood. In addition, mathematical proof of the de-enhancement feasibility is provided. Experimental results demonstrate that the proposed method outperforms the compared methods in both enhancement detection and parameter estimation.
Jianquan Yang, Guopu Zhu, Yao Luo, Sam Kwong, Xinpeng Zhang 0001, Yicong Zhou
IEEE Trans. Circuits Syst. Video Technol.2
2022 Truncated Robust Natural Watermarking With Hungarian Optimization
abstract
Over the past two decades, spread spectrum (SS) embedding has been widely used in digital watermarking due to its competitive performance in robustness and security. However, the robustness of existing secure SS embedding methods, such as natural watermarking (NW) and robust-NW (RNW), is still weak. In this article, we propose a new secure SS embedding method named truncated-RNW (TRNW), which improves the robustness of RNW while maintaining the same security level. The main idea of TRNW is to move RNW-watermarked correlations within an origin-centered sphere onto the spherical surface along the radial direction. Moreover, Hungarian algorithm is used to reduce embedding distortion, and the optimized method is called Hungarian-TRNW (HTRNW). A theoretical analysis and extensive experiments are conducted to validate the effectiveness of the proposed method. The results show that HTRNW achieves the same security level as RNW and a significant improvement over existing representative secure SS embeddings in terms of robustness.
Jinkun You, Yuan-Gen Wang, Guopu Zhu, Sam Kwong
IEEE Trans. Circuits Syst. Video Technol.3
2022 Multi-Task SE-Network for Image Splicing Localization
abstract
Image splicing can be easily used for illegal activities such as falsifying propaganda for political purposes and reporting false news, which may result in negative impacts on society. Hence, it is highly required to detect spliced images and localize the spliced regions. In this work, we propose a multi-task squeeze and excitation network (SE-Network) for splicing localization. The proposed network consists of two streams, namely label mask stream and edge-guided stream, both of which adopt convolutional encoder-decoder architecture. The information from the edge-guided stream is transmitted to the label mask stream for enhancing the discrimination of features between the spliced and host regions. This work has three main contributions. First, image edges, along with label masks and mask edges, are exploited to supply more comprehensive supervision for the localization of spliced regions. Second, the low-level feature maps extracted from shallow layers are fused with the high-level feature maps from deep layers to provide more reliable feature for splicing localization. Finally, several squeeze and excitation attention modules are incorporated into the network to recalibrate the fused features to enhance the feature expression. Extensive experiments show that the proposed multi-task SE-Network outperforms existing splicing localization methods evidently on two synthetic splicing datasets and four benchmark splicing datasets.
Yulan Zhang, Guopu Zhu, Ligang Wu 0001, Sam Kwong, Hongli Zhang 0001, Yicong Zhou
IEEE Trans. Circuits Syst. Video Technol.2
2022 Asynchronous Distributed Finite-Time H∞ Filtering in Sensor Networks With Hidden Markovian Switching and Two-Channel Stochastic Attack
abstract
This article investigates the asynchronous distributed finite-time$H_{\infty }$filtering problem for nonlinear Markov jump systems over sensor networks under stochastic attacks. The stochastic attacks, called two-channel deception attacks, exist not only between the Markov jump plant and the sensors but also among the sensors. It is assumed that the mode of the filter relies on, but is asynchronous with, that of the Markov jump plant. First, we establish a filtering error system that combines the Markov jump plant with the asynchronous filtering system. Then, we present an asynchronous distributed filter, which ensures the filtering error system mean-square finite-time bounded and satisfies a prescribed$H_{\infty }$performance level under the two-channel attacks. Finally, an example is given to illustrate the effectiveness of the presented filter.
Guopu Zhu, Peng Shi 0001, Ramesh K. Agarwal
IEEE Trans. Cybern.2
2022 Anti-Forensics for Face Swapping Videos via Adversarial Training
abstract
Generating falsified faces by artificial intelligence, widely known as DeepFake, has attracted attention worldwide since 2017. Given the potential threat brought by this novel technique, forensics researchers dedicated themselves to detect the video forgery. Except for exposing falsified faces, there could be extended research directions for DeepFake such as anti-forensics. It can disclose the vulnerability of current DeepFake forensics methods. Besides, it could also enable DeepFake videos as tactical weapons if the falsified faces are more subtle to be detected. In this paper, we propose a GAN model to behave as an anti-forensics tool. It features a novel architecture with additional supervising modules for enhancing image visual quality. Besides, a loss function is designed to improve the efficiency of the proposed model. After experimental evaluations, we show that the DeepFake forensics detectors are susceptible to attacks launched by the proposed method. Besides, the proposed method can efficiently produce anti-forensics videos in satisfying visual quality without noticeable artifacts. Compared with the other anti-forensics approaches, this is tremendous progress achieved for DeepFake anti-forensics. The attack launched by our proposed method can be truly regarded as DeepFake anti-forensics as it can fool detecting algorithms and human eyes simultaneously.
Feng Ding 0007, Guopu Zhu, Yingcan Li, Xinpeng Zhang 0001, Pradeep K. Atrey, Siwei Lyu
IEEE Trans. Multim.2
2022 A Novel Rank Learning Based No-Reference Image Quality Assessment Method
abstract
Recently, applying deep learning to no-reference image quality assessment (NR-IQA) has received significant attention. Especially in the last five years, an increasing interest has been drawn to the studies of rank learning since it can help mitigate the problem of small IQA datasets. However, on one hand, existing rank learning is not suitable for the authentically distorted images due to the lack of generated rank samples. On the other hand, the output of existing rank loss functions is uncontrollable, resulting in reduced performance. Motivated by these two limitations, we propose a novel rank learning based NR-IQA method, termed controllable list-wise ranking IQA (CLRIQA) in this paper. To be specific, we first present an imaging-heuristic approach, in which the over- and under-exposure is formulated as an inverse of the Weber-Fechner law, and fusion strategy and compression are adopted, to simulate the authentic distortion and generate the rank image samples. These samples are label-free yet associated with quality ranking information. Then we design a controllable list-wise ranking (CLR) loss function by setting an upper and lower bound of rank range and introducing an adaptive margin to tune rank interval. Finally, both the generated rank samples and proposed CLR are used to pre-train a convolutional neural network. Moreover, to obtain a more accurate prediction model, we take advantage of the IQA datasets to fine-tune the pre-trained network further. Various experiments are conducted on the IQA benchmark datasets, and experimental results demonstrate the effectiveness of the proposed CLRIQA method. The source code and network model can be downloaded at the following web address:https://github.com$/$GZHU-DVL$/$CLRIQA.
Fu-Zhao Ou, Yuan-Gen Wang, Jin Li 0002, Guopu Zhu, Sam Kwong
IEEE Trans. Multim.4
2022 Contrast-Enhanced Color Visual Cryptography for (k, n) Threshold Schemes
abstract
In traditional visual cryptography schemes (VCSs), pixel expansion remains to be an unsolved challenge. To alleviate the impact of pixel expansion, several colored-black-and-white VCSs, called CBW-VCSs, were proposed in recent years. Although these methods could ease the effect of pixel expansion, the reconstructed image obtained by these methods may also suffer from low contrasts. To address this issue, we propose a contrast-enhanced (k, n) CBW-VCS based on random grids, named (k,n) RG-CBW-VCS, in this article. By applying color random grids, a binary secret image is encrypted into n color shares that have no pixel expansion. When any k 1 (k 1 > k ) color shares are collected together, the stacked results of them can be identified as the secret image; whereas the superposition of any k 2 ( k 2 < k ) color shares shows nothing. Through theoretical analysis and experimental results, we justify the effectiveness of the proposed (k, n) RG-CBW-VCS. Compared with related methods in feature, contrast, and pixel expansion, the results indicate that the proposed method generally achieves better performance.
Zuquan Liu, Guopu Zhu, Feng Ding 0007, Xiangyang Luo 0001, Sam Kwong, Peng Li 0050
ACM Trans. Multim. Comput. Commun. Appl.2
2022 Stability Analysis for Hybrid Time-Delay Systems With Double Degrees
abstract
In this article, the problem of stability is investigated for a class of hybrid time-delay systems with double degrees. The systems consist of both discrete and continuous subsystems that have homogeneity and nonzero degrees. Conditions are derived to guarantee the preasymptotic stability of the systems. The relationships of solutions are established between the systems with different sizes using homogeneous properties. Then, equivalent conditions are presented for the analysis of preasymptotic stability. Furthermore, necessary and sufficient delay-independent conditions are developed to analyze the stability of the systems. Finally, two examples are provided to illustrate the effectiveness of the proposed new stability analysis methods.
Guopu Zhu, Peng Shi 0001
IEEE Trans. Syst. Man Cybern. Syst.2
2021 Reinforcement learning-based QoE-oriented dynamic adaptive streaming framework
Xuekai Wei, Mingliang Zhou 0001, Sam Kwong, Hui Yuan 0001, Shiqi Wang 0001, Guopu Zhu, Jingchao Cao
Inf. Sci.6
2021 Feature pyramid network for diffusion-based image inpainting detection
Yulan Zhang, Feng Ding 0007, Sam Kwong, Guopu Zhu
Inf. Sci.4
2021 Hybrid prediction-based pixel-value-ordering method for reversible data hiding
Feng Ding 0007, Xiaolong Li 0001, Guopu Zhu
J. Vis. Commun. Image Represent.4
2021 Weighted visual secret sharing for general access structures based on random grids
Zuquan Liu, Guopu Zhu, Feng Ding 0007, Sam Kwong
Signal Process. Image Commun.2
2021 Distributed Fault Detection and Control for Markov Jump Systems Over Sensor Networks With Round-Robin Protocol
abstract
This paper addresses the problem of simultaneous fault detection and control (SFDC) for Markov jump systems (MJSs) over sensor networks. In order to save the bandwidth of the network, round-Robin protocol is adopted for the communication between the sensors and fault detection filter. First, to fully utilize the data of sensor networks, a new distributed system is developed especially for SFDC, and a residual system, which integrates the MJSs with the distributed SFDC system, is also established. Then, by constructing a mode-dependent Lyapunov functional, the conditions for distributed SFDC are derived, which can make the residual system asymptotically mean-square stable and meanwhile satisfy the prescribed H∞performance requirements. Finally, comparison analyses are made by examples to illustrate the effectiveness and superiority of the proposed new design approach.
Guopu Zhu, Peng Shi 0001, Ramesh K. Agarwal
IEEE Trans. Circuits Syst. I Regul. Pap.2
2021 Defeating Lattice-Based Data Hiding Code Via Decoding Security Hole
abstract
Lattice code has been widely used for data hiding. It can provide security for data hiding by randomly translating its codebook with a secret dither. However, besides the secret dither, there are an infinite amount of points that are near the secret dither and can also be used for perfect decoding in the noiseless scenario. This means that lattice-based data hiding has a serious security hole, named decoding security hole (DSH) in this paper. After a theoretical analysis of DSH, we find that these points form a convex polytope and the centroid of this convex polytope is the secret dither. Based on this finding, a simple yet effective attack method is presented to estimate the secret dither of lattice-based data hiding. Extensive experimental results show that the proposed method significantly outperforms state-of-the-art attack methods, especially when the number of observations is small or the document-to-watermark ratio changes over a wide range.
Yuan-Gen Wang, Guopu Zhu, Jin Li 0002, Mauro Conti, Jiwu Huang
IEEE Trans. Circuits Syst. Video Technol.2
2021 A Clustering-Based Framework for Improving the Performance of JPEG Quantization Step Estimation
abstract
Quantization plays a pivotal role in JPEG compression with respect to the tradeoff between image fidelity and storage size, and the blind estimation of quantization parameters has attracted considerable interest in the fields of image steganalysis and forensics. Existing estimation methods have made great progress, but they usually suffer a sharp decline in accuracy when addressing small-size JPEG decompressed bitmaps due to the insufficiency of coefficients. Aiming to alleviate this issue, this paper proposes a generic clustering-based framework to improve the performance of the existing methods. The core idea is to gather as many coefficients as possible by clustering subbands before feeding them into a step estimator. The proposed framework is implemented using hierarchical clustering with two kinds of histogram-like features. Extensive experiments are conducted to validate the effectiveness of the proposed framework on a variety of images of different sizes and quality factors, and the results show that notable improvements can be achieved. In addition to quantization step estimation, we believe the idea behind the proposed framework might provide inspiration for other forensic tasks to alleviate their performance issues induced by sample insufficiency.
Jianquan Yang, Yulan Zhang, Guopu Zhu, Sam Kwong
IEEE Trans. Circuits Syst. Video Technol.3
2021 ASIF-Net: Attention Steered Interweave Fusion Network for RGB-D Salient Object Detection
abstract
Salient object detection from RGB-D images is an important yet challenging vision task, which aims at detecting the most distinctive objects in a scene by combining color information and depth constraints. Unlike prior fusion manners, we propose an attention steered interweave fusion network (ASIF-Net) to detect salient objects, which progressively integrates cross-modal and cross-level complementarity from the RGB image and corresponding depth map via steering of an attention mechanism. Specifically, the complementary features from RGB-D images are jointly extracted and hierarchically fused in a dense and interweaved manner. Such a manner breaks down the barriers of inconsistency existing in the cross-modal data and also sufficiently captures the complementarity. Meanwhile, an attention mechanism is introduced to locate the potential salient regions in an attention-weighted fashion, which advances in highlighting the salient objects and suppressing the cluttered background regions. Instead of focusing only on pixelwise saliency, we also ensure that the detected salient objects have the objectness characteristics (e.g., complete structure and sharp boundary) by incorporating the adversarial learning that provides a global semantic constraint for RGB-D salient object detection. Quantitative and qualitative experiments demonstrate that the proposed method performs favorably against 17 state-of-the-art saliency detectors on four publicly available RGB-D salient object detection datasets. The code and results of our method are available at https://github.com/Li-Chongyi/ASIF-Net.
Chongyi Li, Runmin Cong, Sam Kwong, Junhui Hou, Huazhu Fu, Guopu Zhu, Dingwen Zhang, Qingming Huang
IEEE Trans. Cybern.6
2021 A Novel (t, s, k, n)-Threshold Visual Secret Sharing Scheme Based on Access Structure Partition
abstract
Visual secret sharing (VSS) is a new technique for sharing a binary image into multiple shadows. For VSS, the original image can be reconstructed from the shadows in any qualified set, but cannot be reconstructed from those in any forbidden set. In most traditional VSS schemes, the shadows held by participants have the same importance. However, in practice, a certain number of shadows are given a higher importance due to the privileges of their owners. In this article, a novel ( t , s , k , n )-threshold VSS scheme is proposed based on access structure partition. First, we construct the basis matrix of the proposed ( t , s , k , n )-threshold VSS scheme by utilizing a new access structure partition method and sub-access structure merging method. Then, the secret image is shared by the basis matrix as n shadows, which are divided into s essential shadows and n - s non-essential shadows. To reconstruct the secret image, k or more shadows should be collected, which include at least t essential shadows; otherwise, no information about the secret image can be obtained. Compared with related schemes, our scheme achieves a smaller shadow size and a higher visual quality of the reconstructed image. Theoretical analysis and experiments indicate the effectiveness of the proposed scheme.
Zuquan Liu, Guopu Zhu, Yuan-Gen Wang, Jianquan Yang, Sam Kwong
ACM Trans. Multim. Comput. Commun. Appl.2
2021 Adaptive Event-Triggered and Double-Quantized Consensus of Leader-Follower Multiagent Systems With Semi-Markovian Jump Parameters
abstract
In practice, many dynamical systems can be modeled by multiagent systems (MASs) with different operating modes. The consensus problem of leader-follower MASs with semi-Markovian jump parameters is investigated in this article. First, to save the bandwidth resources of the MASs, adaptive event-triggered communication and double quantization are introduced to design the consensus protocol, in which the event-triggered threshold is dynamically adjusted instead of being a constant, and not only the data from the agents to the consensus controller but also that from the controller to the agents, are both quantized. Next, a sufficient condition is established to ensure the stability of the consensus error system under the adaptive event-triggered condition, and then the consensus controller is developed by matrix inequality approach. Finally, a numeral example is given to demonstrate the effectiveness of the developed consensus controller.
Guopu Zhu, Peng Shi 0001
IEEE Trans. Syst. Man Cybern. Syst.2
2020 (τ, ϵ)-Greedy Reinforcement Learning For Anti-Jamming Wireless Communications
abstract
In this article, we propose a(τ, ε)-greedy reinforcement learning algorithm for anti-jamming wireless communications, which chooses previous action with probability τ and applies ε-greedy with probability 1-τ. The key idea of our algorithm is that the more valuable the previous action is, the higher probability of directly performing it at the current time slot without learning. For this purpose, the average utility of several previous actions is first calculated as a threshold for the valuable action judgment. Then, probability τ is formulated as a Gaussian-like function with respect to the difference between the threshold and the utility of the previous action, which makes the wireless devices find the optimal action at a faster speed in the early stage, and eventually ensures the convergence. As a concrete example, the proposed algorithm is implemented in a wireless communication system against multiple jammers. Simulation results show that compared with ε-greedy, the (τ, ε)-greedy obtains faster convergence rate and slightly higher signalto-interference-plus-noise ratio when being applied to Qlearning, deep Q-networks (DQN), double DQN (DDQN), and prioritized experience reply based DDQN (PDDQN). The source code is available at https://github.com/GZHUDVL/tau-epsilon-greedy-RL.
Yuan-Gen Wang, Jin Li 0002, Liang Xiao 0003, Guopu Zhu
GLOBECOM5
2020 Intra Frame Rate Control for Versatile Video Coding with Quadratic Rate-Distortion Modelling
abstract
With numerous coding tools adopted in the forthcoming Versatile Video Coding (VVC) standard, much less work has been dedicated to study the corresponding Rate-Distortion (R-D) characteristics. This paper proposes a new quadratic R-D model for Versatile Video Coding. In particular, based on the proposed model, a new R-λ relationship is derived and used for frame level rate control. The rate control algorithm is implemented on VTM 2.0 platform for intra coding scenarios. Compared to the default rate control algorithm in VTM 2.0, experimental results show that proposed rate control algorithm can achieve 0.77% BD-BR reduction with similar control accuracy.
Yi Chen 0028, Sam Kwong, Mingliang Zhou 0001, Shiqi Wang 0001, Guopu Zhu
ICASSP5
2020 Just Noticeable Distortion Based Perceptually Lossless Intra Coding
abstract
Perceptual video coding plays a very important role in video codec optimization aiming at removing the perceptual redundancies in video content. In this paper, a just noticeable distortion (JND) guided perceptually lossless coding framework is proposed for Versatile Video Coding (VVC) intra coding. Within this framework, a pattern-based pixel wise JND model is employed to guide the distortion distribution, and subsequently the most appropriate quantization parameter is chosen for each Coding Tree Unit (CTU). The content adaptive Laplacian distribution based D-Q model based on a two pass coding framework is established to derive the most proper QP that satisfies the perceptually lossless coding criteria. The whole framework is integrated into the H.266/VVC intra coding framework. Experimental results demonstrate that the proposed scheme can achieve high accuracy prediction and efficient perceptually lossless intra coding, leading to around 10% bitrate savings comparing with the frame level QP derivation scheme.
Xuelin Shen, Xinfeng Zhang 0001, Shiqi Wang 0001, Sam Kwong, Guopu Zhu
ICASSP5
2020 A TV-log nonconvex approach for image deblurring with impulsive noise
Benxin Zhang, Guopu Zhu
Signal Process.2
2020 METEOR: Measurable Energy Map Toward the Estimation of Resampling Rate via a Convolutional Neural Network
abstract
In recent years, with the improvements in machine learning, image forensics has made considerable progress in detecting editing manipulations. This progress also raises more questions in image forensics research, such as can the parameters applied in a manipulation be estimated. Many parameter estimation works have already been performed. However, most of these works are based on mathematical analyses. In this paper, we attempt to solve a particular parameter estimation problem from a different aspect. Specifically, a new convolutional neural network (CNN) model is proposed to estimate the resampling rate for resampled images regardless of whether the image is upscaled or downscaled. This model features an original layer to generate a measurable energy map toward the estimation of resampling rate (METEOR). The METEOR layer is demonstrated to be an outstanding method that can assist in enhancing the estimation performance of the CNN. Furthermore, the METEOR layer can also increase the robustness of the CNN against JPEG compression, which makes it extremely important in realistic application scenarios. Our work has verified that machine learning, particularly CNNs, with proper optimization can also be refined to adapt to parameter estimation in digital forensics with excellent performance and robustness.
Feng Ding 0007, Hanzhou Wu, Guopu Zhu, Yun Q. Shi 0001
IEEE Trans. Circuits Syst. Video Technol.3
2019 A Novel Blind Image Quality Assessment Method Based on Refined Natural Scene Statistics
abstract
Natural scene statistics (NSS) model has received considerable attention in the image quality assessment (IQA) community due to its high sensitivity to image distortion. However, most existing NSS-based IQA methods extract features either from spatial domain or from transform domain. There is little work to simultaneously consider the features from these two domains. In this paper, a novel blind IQA method (NBIQA) based on refined NSS is proposed. The proposed NBIQA first investigates the performance of a large number of candidate features from both the spatial and transform domains. Based on the investigation, we construct a refined NSS model by selecting competitive features from existing NSS models and adding three new features. Then the refined NSS is fed into SVM tool to learn a simple regression model. Finally, the trained regression model is used to predict the scalar quality score of the image. Experimental results tested on both LIVE IQA and LIVE-C databases show that the proposed NBIQA performs better in terms of synthetic and authentic image distortion than current mainstream IQA methods. The source code is available at https://github.com/GZU-Image-Video-Lab/NBIQA.
Fu-Zhao Ou, Yuan-Gen Wang, Guopu Zhu
ICIP3
2019 Smoothing identification for digital image forensics
Feng Ding 0007, Yuxi Shi, Guopu Zhu, Yun Q. Shi 0001
Multim. Tools Appl.3
2018 An efficient weak sharpening detection method for image forensics
Feng Ding 0007, Guopu Zhu, Weiqiang Dong, Yun Q. Shi 0001
J. Vis. Commun. Image Represent.2
2018 Detecting median filtering via two-dimensional AR models of multiple filtered residuals
Jianquan Yang, Honglei Ren, Guopu Zhu, Jiwu Huang, Yun Q. Shi 0001
Multim. Tools Appl.3
2018 ℒ2-ℒ∞ filtering for stochastic time-varying delay systems based on the Bessel-Legendre stochastic inequality
Guopu Zhu, Peng Shi 0001
Signal Process.2
2018 A Study on the Security Levels of Spread-Spectrum Embedding Schemes in the WOA Framework
abstract
Security analysis is a very important issue for digital watermarking. Several years ago, according to Kerckhoffs' principle, the famous four security levels, namely insecurity, key security, subspace security, and stego-security, were defined for spread-spectrum (SS) embedding schemes in the framework of watermarked-only attack. However, up to now there has been little application of the definition of these security levels to the theoretical analysis of the security of SS embedding schemes, due to the difficulty of the theoretical analysis. In this paper, based on the security definition, we present a theoretical analysis to evaluate the security levels of five typical SS embedding schemes, which are the classical SS, the improved SS (ISS), the circular extension of ISS, the nonrobust and robust natural watermarking, respectively. The theoretical analysis of these typical SS schemes are successfully performed by taking advantage of the convolution of probability distributions to derive the probabilistic models of watermarked signals. Moreover, simulations are conducted to illustrate and validate our theoretical analysis. We believe that the theoretical and practical analysis presented in this paper can bridge the gap between the definition of the four security levels and its application to the theoretical analysis of SS embedding schemes.
Yuan-Gen Wang, Guopu Zhu, Sam Kwong, Yun Q. Shi 0001
IEEE Trans. Cybern.2
2018 Transportation Spherical Watermarking
abstract
During the past twenty years, there has been a great interest in the study of spread spectrum (SS) watermarking. However, it is still a challenging task to design a secure and robust SS watermarking method. In this paper, we first define a family of secure SS watermarking methods, named as spherical watermarking (SW). The watermarked correlation of SW is defined to be uniformly distributed on a spherical surface, and this makes SW be key-secure against the watermarked-only attack. Then, we propose an implementation of SW, called transportation SW (TSW), which is designed to decrease embedding distortion in a recursive manner using the transportation theory, meanwhile keeping the security of SW. Moreover, we present a theoretical analysis of the embedding distortion and robustness of the proposed method. Finally, extensive experiments are conducted on simulated signals and real images. The experimental results show that TSW is more robust than existing secure SS watermarking methods.
Yuan-Gen Wang, Guopu Zhu, Yun Q. Shi 0001
IEEE Trans. Image Process.2
2016 Analyzing the Effect of JPEG Compression on Local Variance of Image Intensity
abstract
The local variance of image intensity is a typical measure of image smoothness. It has been extensively used, for example, to measure the visual saliency or to adjust the filtering strength in image processing and analysis. However, to the best of our knowledge, no analytical work has been reported about the effect of JPEG compression on image local variance. In this paper, a theoretical analysis on the variation of local variance caused by JPEG compression is presented. First, the expectation of intensity variance of 8×8 non-overlapping blocks in a JPEG image is derived. The expectation is determined by the Laplacian parameters of the discrete cosine transform coefficient distributions of the original image and the quantization step sizes used in the JPEG compression. Second, some interesting properties that describe the behavior of the local variance under different degrees of JPEG compression are discussed. Finally, both the simulation and the experiments are performed to verify our derivation and discussion. The theoretical analysis presented in this paper provides some new insights into the behavior of local variance under JPEG compression. Moreover, it has the potential to be used in some areas of image processing and analysis, such as image enhancement, image quality assessment, and image filtering.
Jianquan Yang, Guopu Zhu, Yun Q. Shi 0001
IEEE Trans. Image Process.2
2015 Recursive optimization of spherical watermarking using transportation theory
abstract
In this paper, we first define a class of secure spread spectrum (SS) watermarking, called spherical watermarking (SW), of which the distribution of the watermarked correlations has a uniform distribution over a spherical surface. Then, we propose an implementation of SW, namely transportation SW (TSW), which recursively minimizes embedding distortion using transportation theory. A theoretical analysis is also presented on the performance of the proposed TSW scheme in terms of the watermark to content power ratio (WCR), and is validated by our simulations. Simulation results show that TSW performs better in terms of WCR and robustness than existing secure SS watermarking schemes.
Yuan-Gen Wang, Guopu Zhu
ICIP3
2015 An Advanced Texture Analysis Method for Image Sharpening Detection
Feng Ding 0007, Weiqiang Dong, Guopu Zhu, Yun Q. Shi 0001
IWDW3
2015 An improved AQIM watermarking method with minimum-distortion angle quantization and amplitude projection strategy
Yuan-Gen Wang, Guopu Zhu
Inf. Sci.2
2015 Edge Perpendicular Binary Coding for USM Sharpening Detection
abstract
Unsharp masking (USM) sharpening is a basic technique for image manipulation and editing. In recent years, the detection of USM sharpening has attracted attention from image forensics point of view. After USM sharpening, overshoot artifacts, which shape image texture, are generated along image edges. By utilizing the special characteristic of the texture modification caused by the USM sharpening, a novel method called edge perpendicular binary coding is proposed in this letter to detect USM sharpening. Extensive experiments have been conducted to show the superiority of the proposed method over the existing methods.
Feng Ding 0007, Guopu Zhu, Jianquan Yang, Yun Q. Shi 0001
IEEE Signal Process. Lett.2
2014 An Effective Method for Detecting Double JPEG Compression With the Same Quantization Matrix
abstract
Detection of double JPEG compression plays an important role in digital image forensics. Some successful approaches have been proposed to detect double JPEG compression when the primary and secondary compressions have different quantization matrices. However, detecting double JPEG compression with the same quantization matrix is still a challenging problem. In this paper, an effective error-based statistical feature extraction scheme is presented to solve this problem. First, a given JPEG file is decompressed to form a reconstructed image. An error image is obtained by computing the differences between the inverse discrete cosine transform coefficients and pixel values in the reconstructed image. Two classes of blocks in the error image, namely, rounding error block and truncation error block, are analyzed. Then, a set of features is proposed to characterize the statistical differences of the error blocks between single and double JPEG compressions. Finally, the support vector machine classifier is employed to identify whether a given JPEG image is doubly compressed or not. Experimental results on three image databases with various quality factors have demonstrated that the proposed method can significantly outperform the state-of-the-art method.
Jianquan Yang, Guopu Zhu, Sam Kwong, Yun Q. Shi 0001
IEEE Trans. Inf. Forensics Secur.3
2013 A Novel Method for Detecting Image Sharpening Based on Local Binary Pattern
Feng Ding 0007, Guopu Zhu, Yun Q. Shi 0001
IWDW2
2013 Detecting Non-aligned Double JPEG Compression Based on Refined Intensity Difference and Calibration
Jianquan Yang, Guopu Zhu, Yun Q. Shi 0001
IWDW2
2012 An Improved Algorithm for Reversible Data Hiding in Encrypted Image
Guopu Zhu, Xiaolong Li 0001, Jianquan Yang
IWDW2
2011 Random Gray code and its performance analysis for image hashing
Guopu Zhu, Sam Kwong, Jiwu Huang, Jianquan Yang
Signal Process.1
2010 Gradient vector flow active contours with prior directional information
Guopu Zhu, Shuqun Zhang, Qingshuang Zeng, Changhong Wang 0003
Pattern Recognit. Lett.1
2010 Fragility analysis of adaptive quantization-based image hashing
abstract
Fragility is one of the most important properties of authentication-oriented image hashing. However, to date, there has been little theoretical analysis on the fragility of image hashing. In this paper, we propose a measure called expected discriminability for the fragility of image hashing and study this fragility theoretically based on the proposed measure. According to our analysis, when Gray code is applied into the discrete-binary conversion stage of image hashing, the value of the expected discriminability, which is dominated by the quantization stage of image hashing, is no more than 1/2. We further evaluate the expected discriminability of the image-hashing scheme that uses adaptive quantization, which is the most popular quantization scheme in the field of image hashing. Our evaluation reveals that if deterministic adaptive quantization is applied, then the expected discriminability of the image-hashing scheme can reach the maximum value (i.e., 1/2). Finally, some experiments are conducted to validate our theoretical analysis and to compare the performance of several quantization schemes for image hashing.
Guopu Zhu, Jiwu Huang, Sam Kwong, Jianquan Yang
IEEE Trans. Inf. Forensics Secur.1
2009 Estimation of distribution algorithms making use of both high quality and low quality individuals
abstract
Most estimation of distribution algorithms only make use of some high quality individuals and neglect other low quality individuals. However like high quality individuals, these neglected low quality individuals also contain some important information that may be useful for guiding the search of estimation of distribution algorithms. This paper proposes a novel kind of estimation of distribution algorithms, where both high quality and low quality individuals in the old population are employed for reproducing new candidate individuals at the next generation. In particular, both the density PH(X) of high quality individuals and the density PL(X) of low quality individuals are estimated; then the new population G is obtained with the following steps employed: 1) a new candidate individual x is reproduced through sampling from the density PH(X); 2) to let PH(X = x) and PL(X = x) compare and the individual x will be stored into the new population G if and only if PH(X = x) ges PL(X = x); 3) the above steps repeat until M new individuals have been successfully generated where M is the population size. To demonstrate the usefulness of low quality individuals for estimation of distribution algorithms, estimation of distribution algorithms using both high quality and low quality individuals are tested on several benchmark problems and their results are compared with those obtained by estimation of distribution algorithms where only high quality individuals are used. The usefulness of low quality individuals for speeding up the search of estimation of distribution algorithms is confirmed by the experimental results.
Yi Hong 0002, Guopu Zhu, Sam Kwong, Qingsheng Ren
FUZZ-IEEE2
2009 A study on the randomness measure of image hashing
abstract
How to measure the security of image hashing is still an open issue in the field of image authentication. Some works have been conducted on the security measure of image hashing. One of the most important works is the randomness measure proposed by Swaminathan, which uses differential entropy as a metric to evaluate the security of randomized image features and has been applied mainly in the security analysis of the feature extraction stage of image hashing. It is meaningful to measure the randomness of the image features over the secret-key set for the security of image hashing because the image features extracted by image hashing should be generated randomly and difficult to guess. However, as is well known, differential entropy is not invariant to scaling; thus it might not be enough to evaluate the security of randomized image features. In this paper, we show the fact that if the image features of an image hash function are scaled by a constant that is large than one, then the tradeoff between the robustness and the fragility of the image hash function will not change at all, but the security indicated by the randomness measure will increase. The above-mentioned fact seems to contradict the following. First, the security of image hashing, which conflicts with robustness and fragility, cannot increase freely. Secondly, a deterministic operation, such as deterministic scaling, does not change the security of image hashing in terms of the difficulty of guessing the secret key or randomized image features. Therefore, the randomness measure should be modified to be invariant to scaling at least.
Guopu Zhu, Jiwu Huang, Sam Kwong, Jianquan Yang
IEEE Trans. Inf. Forensics Secur.1
2008 Anisotropic virtual electric field for active contours
Guopu Zhu, Shuqun Zhang, Qingshuang Zeng, Changhong Wang 0004
Pattern Recognit. Lett.1
2007 Efficient Illumination Insensitive Object Tracking by Normalized Gradient Matching
abstract
We propose a novel tracking algorithm by minimizing the sum-of-squared differences (SSD) between the normalized image gradients of the template image and the input image from the test image sequence. The proposed tracking algorithm is efficient to implement since it is based on the framework of the inverse compositional algorithm, a computationally efficient tracking technique. The experiments show that the proposed tracking algorithm is superior to the intensity-based inverse compositional algorithm in tracking objects under varying illumination conditions.
Guopu Zhu, Shuqun Zhang, Xijun Chen, Changhong Wang 0004
IEEE Signal Process. Lett.1
2006 Efficient edge-based object tracking
Guopu Zhu, Qingshuang Zeng, Changhong Wang 0004
Pattern Recognit.1