Zhihua Xia

dblp:94/8286 · DBLP profile ↗
← Back
99ranked-venue papers
13as first author
83since 2021 · last 2026
0000-0001-6860-647XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 40 · 4 first-author · 36 since 2021Security and privacy · 26 · 2 first-author · 23 since 2021Artificial intelligence and machine learning · 22 · 22 since 2021Computer networks · 7 · 1 first-author · 6 since 2021Systems, architecture and hardware · 6 · 3 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 CLIP-FTI: Fine-Grained Face Template Inversion via CLIP-Driven Attribute Conditioning
abstract
Face recognition systems store face templates for efficient matching. Once leaked, these templates pose a threat: inverting them can yield photorealistic surrogates that compromise privacy and enable impersonation. Although existing research has achieved relatively realistic face template inversion, the reconstructed facial images exhibit over-smoothed facial-part attributes (eyes, nose, mouth) and limited transferability. To address this problem, we present CLIP-FTI, a CLIP-driven fine-grained attribute conditioning framework for face template inversion. Our core idea is to use the CLIP model to obtain the semantic embeddings of facial features, in order to realize the reconstruction of specific facial feature attributes. Specifically, facial feature attribute embeddings extracted from CLIP are fused with the leaked template via a cross-modal feature interaction network and projected into the intermediate latent space of a pretrained Style- GAN. The StyleGAN generator then synthesizes face images with the same identity as the templates but with more finegrained facial feature attributes. Experiments across multiple face recognition backbones and datasets show that our reconstructions (i) achieve higher identification accuracy and attribute similarity, (ii) recover sharper component-level attribute semantics, and (iii) improve cross-model attack transferability compared to prior reconstruction attacks. To the best of our knowledge, ours is the first method to use additional information besides the face template attack to realize face template inversion and obtains SOTA results.
Longchen Dai, Zixuan Shen, Peipeng Yu, Zhihua Xia
AAAI5
2026 One for All: Synthesis-Free Fingerprint Learning for Attribution of In-the-Wild Synthetic Images
abstract
Attributing synthetic images to their source generative models is critical for digital forensics and security. While most existing attribution methods can distinguish images produced by known models and reject those from unknown ones, they are unable to verify whether a given image was produced by a specific, previously unseen model. To address this limitation, we formulate an open-set verification problem: determining whether a given image was generated by a specific model. Our key insight is that synthetic images from different models show consistent, content-independent fingerprints in their amplitude spectrum. Based on this insight, we design a dynamic fingerprint simulator capable of simulating over 1.6 trillion generative model architectures. We further train an extractor to capture model-specific fingerprint representations with supervised contrastive learning, enabling accurate attribution of synthetic images, even from previously unseen models. Our method does not rely on any synthetic images, instead, it is trained solely on real images. On DMDetection and AIGCBenchmark, which comprises dozens of state-of-the-art and in-the-wild generative models, our method improves the attribution performance (AUC) of the prior method from random level to 94.05% and 83.05%, respectively. On GenImage and OSMA datasets, we obtain 85.08%, and 88.48% OSCR, outperforming the SOTA methods by 4.30% and 9.37% under the same settings.
Jianwei Fei, Yunshu Dai, Peipeng Yu, Zhihua Xia, Dasara Shullani, Daniele Baracchi, Alessandro Piva
AAAI4
2026 Towards Provably Secure and Highly Robust Generative Image Steganography Leveraging Latent Diffusion Model
abstract
Generative image steganography has attracted significant attention for its exceptional resistance to steganalysis. However, current generative steganography methods still face limitations in terms of the lack of provable security guarantees under statistical analysis and vulnerability to real-world, unforeseen channel attacks. To address these issues, this paper proposes a novel generative image steganography framework that leverages the Latent Diffusion Model (LDM). Notably, we have uncover a consistent trend: regardless of whether an image has undergone attacks such as compression or noise addition, the sign pattern of values in its latent vector encoded by the LDM remains largely invariant. Capitalizing on this trend, we have devised an adaptive distribution-preserving mapping (ADPM) mechanism, capable of converting a secret message into a latent vector that follows standard normal distribution in an adjustable way. Since both the secret latent vector and the latent vector randomly generated during regular image generation follow the same distribution, satisfying the optimal input conditions for the diffusion model, the proposed method can achieve provable security. Experimental results demonstrate the outstanding performance of our approach in terms of robustness, security, and extraction accuracy.
Chengsheng Yuan 0001, Zhaonan Ji, Zhili Zhou 0001, Xinting Li, Zhihua Xia
AAAI6
2026 Fine-Grained DINO Tuning with Dual Supervision for Face Forgery Detection
abstract
The proliferation of sophisticated deepfakes poses significant threats to information integrity. While DINOv2 shows promise for detection, existing fine-tuning approaches treat it as generic binary classification, overlooking distinct artifacts inherent to different deepfake methods. To address this, we propose a DeepFake Fine-Grained Adapter (DFF-Adapter) for DINOv2. Our method incorporates lightweight multi-head LoRA modules into every transformer block, enabling efficient backbone adaptation. DFF-Adapter simultaneously addresses authenticity detection and fine-grained manipulation type classification, where classifying forgery methods enhances artifact sensitivity. We introduce a shared branch propagating fine-grained manipulation cues to the authenticity head. This enables multi-task cooperative optimization, explicitly enhancing authenticity discrimination with manipulation-specific knowledge. Utilizing only 3.5M trainable parameters, our parameter-efficient approach achieves detection accuracy comparable to or even surpassing that of current complex state-of-the-art methods.
Peipeng Yu, Zhihua Xia, Longchen Dai
AAAI3
2026 CREF: Concept Response Fingerprints for Large Language Models
abstract
Protecting the intellectual property of Large Language Models (LLMs) is critical because training them requires massive computational resources and data. A key challenge is determining whether a suspicious model is derived from a specific base model after fine-tuning or structural modification. Existing fingerprinting methods rely on model weights or high-dimensional representations, leading to substantial storage overhead. We propose a non-intrusive fingerprinting framework Concept REsponse Fingerprints (CREF). Inspired by activation engineering, CREF constructs a set of concept activation vectors as semantic probes. It then measures the response strength of hidden representations along these concept activation vectors using shared inputs. The resulting concept response matrix serves as a compact fingerprint, and similarity between models is measured using centered kernel alignment. Experiments on multiple LLM families show that CREF reliably distinguishes derived models from independently trained models and remains robust to fine-tuning, pruning, parameter permutation, and scaling. Moreover, the fingerprint requires only kilobyte-level storage, making it practical for large-scale deployment and ownership verification.
Haiyong Tang, Hanzhou Wu, Gejian Zhao, Li Li 0103, Zhihua Xia, Xinpeng Zhang 0001
IH&MMSec5
2026 A robust dual-pronged proactive defense framework against deepfakes via adversarial semi-fragile watermarking
Chengsheng Yuan 0001, Youqiang Cao, Zhili Zhou 0001, Zhangjie Fu 0001, Zhihua Xia, Q. M. Jonathan Wu
Expert Syst. Appl.5
2026 Secure Distribution: Anti-collusion Watermarking via Spectral Weight Modulation in Latent Diffusion Models
Yunshu Dai, Jianwei Fei, Wenhong Huang, Fangjun Huang, Zhihua Xia
Pattern Recognit.5
2026 BAM: Backdoor defense based on adversarial mitigation
Run Lu, Peipeng Yu, Zhihua Xia
Pattern Recognit.4
2026 SWFTI: Facial template inversion via StyleSwin mapping
Zixuan Shen, Zhihua Xia, Kaikai Gan, Peipeng Yu
Pattern Recognit.2
2026 A Robust Reversible Watermarking scheme using DC prediction and histogram shifting
Jiancheng Xiao, Shuaichao Wu, Bingwen Feng, Jilian Zhang, Bing Chen 0004, Zhihua Xia, Wei Lu 0001
Signal Process.6
2026 Secure Difference Contraction Watermarking for Static Deep Neural Networks
abstract
Static deep neural network (DNN) watermarking techniques typically employ irreversible methods to embed watermarks into the DNN model weights. However, this approach causes permanent damage to the watermarked model and fails to meet the requirements for integrity authentication. Reversible data hiding (RDH) methods offer a potential solution, but existing approaches suffer from limitations in usability, capacity, and fidelity, hindering their practical adoption. In this paper, we propose a secure static DNN watermarking scheme called Secure Difference Contraction (SDC). Our scheme utilizes a one-dimensional quantizer for watermark embedding and employs dithering to ensure key-dependent security, i.e., the watermark cannot be correctly extracted without the secret key used during embedding. Additionally, we design two schemes to address the challenges of integrity protection and legitimate authentication for DNNs. Simulation results on training loss and classification accuracy demonstrate the feasibility and effectiveness of our proposed methods, highlighting their advantages in capacity and fidelity over existing techniques.
Shanxiang Lyu, Junren Qin, Fan Yang 0149, Rongke Liu, Zhihua Xia, Xiaochun Cao
IEEE Trans. Dependable Secur. Comput.5
2026 Rotation, Scale, and Translation Resilient Black-Box Fingerprinting for Intellectual Property Protection of EaaS Models
abstract
Feature embedding has become a cornerstone technology for processing high-dimensional and complex data, which results in that Embedding as a Service (EaaS) models have been widely deployed in the cloud. To protect the intellectual property of EaaS models, existing methods apply digital watermarking to inject specific backdoor triggers into EaaS models by modifying training samples or network parameters. However, these methods inevitably produce detectable patterns through semantic analysis and exhibit susceptibility to geometric transformations including rotation, scaling, and translation (RST). To address this problem, we propose a novel fingerprinting framework called POSTER for EaaS models, rather than merely refining existing watermarking techniques. Different from watermarking, the proposed POSTER establishes EaaS model ownership through geometric analysis of embedding space's topological structure, rather than relying on the modified training samples or triggers. The key innovation lies in modeling the victim and suspicious embeddings as point clouds, allowing us to perform robust spatial alignment and similarity measurement, which inherently resists RST attacks. Experiments evaluated on visual and textual embedding tasks verify the superiority and applicability. This work reveals inherent characteristics of EaaS models and provides a promising solution for ownership verification of EaaS models under black-box scenarios.
Hongjie Zhang 0001, Zhiqi Zhao, Hanzhou Wu, Zhihua Xia, Athanasios V. Vasilakos
IEEE Trans. Dependable Secur. Comput.4
2026 Optimal Access Structure Partition Methods for Image Secret Sharing
abstract
Visual cryptography scheme (VCS) and polynomial-based secret image sharing (PSIS) are two primary types of secret sharing for protecting images. VCS and PSIS have their respective pros and cons. For VCS, the benefits of perfect security and easy decoding are provided. But it suffers from the limitations of lossy secret recovery and binary image-oriented. PSIS can deal with grayscale/color images and offers lossless secret reconstruction. Whereas, the secret decoding is computationally intensive (i.e.,O(klog2k) for (k,n) threshold) and the residual-image problem in PSIS compromises the security. In this paper, we are motivated to investigate a sharing technique that can preserve the advantages of both VCS and PSIS. Differing from existing VCS and PSIS, the proposed sharing method is accomplished based on the access structure partition (ASP) result. Essentially, an ASP guided image secret sharing approach is developed and three optimal ASP algorithms are designed. When compared with existing partition method, significant improvement is offered by our partition techniques especially for the (k,n) threshold with a largern. Take the (2; 15), (2; 18), and (4; 12) thresholds for example, the numbers of involved sub-access structures by our method are 4, 5, and 19, while the quantities by existing approach are 8, 10, and 45. The percentages of improvement are 100%, 100%, and 137%. Further, based on the partition result from ASP algorithms, we can employ (k,k) probabilistic VCS (PVCS) to constitute a (k,n) sharing method for encoding gray-level/color images. Experiments are demonstrated to confirm the effectiveness of the sharing method and ASP algorithms. Meanwhile, comparisons are included to show that the merits of perfect security, low decoding complexity (i.e.,O(d)), lossless secret recovery (i.e., PSNR= ∞, SSIM= 1), and grayscale/color image-oriented are provided by our sharing method.
Zhihua Xia, Ching-Nung Yang, Wei Qi Yan 0001
IEEE Trans. Inf. Forensics Secur.3
2026 EA-APO: A Universal Proactive Defense Against Facial Manipulation
abstract
The advent of deep learning has accelerated the development of facial manipulation techniques, particularly face-swapping and face attribute editing, raising serious concerns about privacy and identity-related misuse. Existing proactive defense methods predominantly target attribute editing and often generalize poorly to face-swapping models, making it difficult to provide effective protection across both tasks within a unified framework. To bridge this gap, we propose a generalized defense framework, Epoch-Adaptive Adversarial Perturbation Optimization (EA-APO). Specifically, EA-APO introduces a proactive defense mechanism that establishes optimal adversarial paths by optimizing perturbations on a white-box surrogate model to enhance adversarial transferability, and applies the resulting perturbations to source face images to disrupt both face swapping and face attribute editing, even against previously unseen target models in black-box settings. This approach mitigates identity feature tampering while adapting to changes in visual attributes and preserving high-quality adversarial examples. Experimental results show the generalization of our method across multiple face-swapping and attribute-editing models, including commercial ones, while also maintaining strong defense under various common post-processing operations and real-world social media transmission conditions, underscoring its potential for real-world deployment.
Lizhi Xiong, Ziqiang Li 0001, Weiwei Jiang 0001, Zhangjie Fu 0001, Zhihua Xia
IEEE Trans. Inf. Forensics Secur.6
2026 DFREC: DeepFake Identity Recovery Based on Identity-Aware Masked Autoencoder
abstract
Recent advances in deepfake forensics have primarily focused on improving the classification accuracy and generalization performance. Despite enormous progress in detection accuracy across a wide variety of forgery algorithms, existing algorithms lack intuitive interpretability and identity traceability to help with forensic investigation. In this paper, we introduce a novel DeepFake Identity Recovery scheme (DFREC) to fill this gap. DFREC aims to recover the pair of source and target faces from a deepfake image to facilitate deepfake identity tracing and reduce the risk of deepfake attacks. It comprises three key components: an Identity Segmentation Module (ISM), a Source Identity Reconstruction Module (SIRM), and a Target Identity Reconstruction Module (TIRM). The ISM segments the input face into distinct source and target face information, and the SIRM reconstructs the source face and extracts latent target identity features with the segmented source information. The background context and latent target identity features are synergetically fused by a Masked Autoencoder in the TIRM to reconstruct the target face. We evaluate DFREC on different high-fidelity face-swapping attacks on FaceForensics++, CelebaMegaFS, FFHQ-E4S, and Celeb-DFv2 datasets, which demonstrate its superior recovery performance over state-of-the-art deepfake recovery algorithms. In addition, DFREC is the only scheme that can recover both pristine source and target faces directly from the forgery image with high fidelity.
Peipeng Yu, Jianwei Fei, Zhihua Xia, Chip-Hong Chang
IEEE Trans. Inf. Forensics Secur.5
2026 Zero-Knowledge Proof-Based IP Protection of Visual Large Models of Autonomous Driving
Chengsheng Yuan 0001, Lvyang Cao, Xinting Li, Zhili Zhou 0001, Zhihua Xia, Zhangjie Fu 0001
IEEE Trans. Inf. Forensics Secur.5
2025 OmniMark: Efficient and Scalable Latent Diffusion Model Fingerprinting
abstract
We introduce OmniMark, a novel and efficient fingerprinting method for Latent Diffusion Models (LDM). OmniMark can encode user-specific fingerprints across diverse dimensions of the weights of the LDM, including kernels, filters, channels, and spatial domains. The LDM is fine-tuned to encode the invisible fingerprint into generated images, which can be decoded by a decoder. By altering fingerprints and re-encoding the weights, OmniMark supports efficient and scalable ad-hoc generation (
Jianwei Fei, Yunshu Dai, Zhihua Xia, Fangjun Huang
AAAI3
2025 IdentityLock: An Identity-aware Backdoor strategy for Face Swapping Defense
abstract
DeepFakes have become capable of producing highly realistic fabricated faces, posing significant threats to personal privacy and social security. The uncontrolled spread of such forged content, especially when influential figures are targeted, could lead to catastrophic consequences for society. Existing defense algorithms primarily rely on passive detection and perturbation-based strategies but fall short in preventing the generation of DeepFake content. Inspired by the concept of backdoor attacks, this paper proposes IdentityLock, a Face Swapping defense algorithm based on backdoor strategy. It leverages the identities of protected individuals as triggers, embedding defense backdoors during the training process of face swapping models. For images of ordinary individuals, the backdoor model produces typical face-swapped outputs. However, when processing images of protected individuals, the model outputs the original target images, thereby safeguarding protected identities. Extensive experiments on the VGGFace2 dataset demonstrate the superior protection performance and robustness of IdentityLock, which effectively shields protected individuals without additional information.
Peipeng Yu, Zhihua Xia, Run Lu
ICASSP3
2025 Scalable Dual Fingerprinting for Hierarchical Attribution of Text-to-Image Models
Jianwei Fei, Yunshu Dai, Peipeng Yu, Zhe Kong, Zhihua Xia
ICCV6
2025 Transparent Vision: A Theory of Hierarchical Invariant Representations
Yushu Zhang 0001, Chao Wang 0028, Zhihua Xia, Xiaochun Cao, Fenglei Fan
ICCV4
2025 Addressing Representation Collapse in Vector Quantized Models with One Linear Layer
abstract
Vector Quantization (VQ) is essential for discretizing continuous representations in unsupervised learning but suffers from representation collapse, causing low codebook utilization and limiting scalability. Existing solutions often rely on complex optimizations or reduce latent dimensionality, which compromises model capacity and fails to fully solve the problem. We identify the root cause as disjoint codebook optimization, where only a few code vectors are updated via gradient descent. To fix this, we propose \textbf{Sim}ple\textbf{VQ}, which reparameterizes code vectors through a learnable linear transformation layer over a latent basis, optimizing the \textit{entire linear space} rather than nearest \textit{individual code vectors}. Although the multiplication of two linear matrices is equivalent to applying a single linear layer, this simple approach effectively prevents collapse. Extensive experiments on image and audio tasks demonstrate that SimVQ improves codebook usage, is easy to implement, and generalizes well across modalities and architectures. The code is available at https://github.com/youngsheen/SimVQ.
Yongxin Zhu 0003, Bocheng Li, Yifei Xin, Zhihua Xia, Linli Xu 0002
ICCV4
2025 Unlocking the Capabilities of Large Vision-Language Models for Generalizable and Explainable Deepfake Detection
abstract
Current Large Vision-Language Models (LVLMs) have demonstrated remarkable capabilities in understanding multimodal data, but their potential remains underexplored for deepfake detection due to the misalignment of their knowledge and forensics patterns. To this end, we present a novel framework that unlocks LVLMs’ potential capabilities for deepfake detection. Our framework includes a Knowledge-guided Forgery Detector (KFD), a Forgery Prompt Learner (FPL), and a Large Language Model (LLM). The KFD is used to calculate correlations between image features and pristine/deepfake image description embeddings, enabling forgery classification and localization. The outputs of the KFD are subsequently processed by the Forgery Prompt Learner to construct fine-grained forgery prompt embeddings. These embeddings, along with visual and question prompt embeddings, are fed into the LLM to generate textual detection responses. Extensive experiments on multiple benchmarks, including FF++, CDF2, DFD, DFDCP, DFDC, and DF40, demonstrate that our scheme surpasses state-of-the-art methods in generalization performance, while also supporting multi-turn dialogue capabilities.
Peipeng Yu, Jianwei Fei, Xuan Feng 0002, Zhihua Xia, Chip-Hong Chang
ICML5
2025 Proof-of-GoS: An Efficient GoS-Based Consensus Algorithm for IoT
abstract
With the advancement of 5G networks, the deploy-ment of Internet of Things (IoT) technology has seen significant growth. Blockchain technology, recognized for its strong security features, is increasingly utilized within the IoT domain. However, the current IoT landscape is characterized by challenges such as substantial resource consumption, limited throughput capacity, and insufficient security protocols, which hinder its optimal per-formance. Towards addressing such problems, we propose a con-sensus algorithm called Proof-of-GoS (PoG) based on the grade of service (GoS), in which a node must have a service score over a set score threshold to be allowed to join the consensus process. The correct behavior of a node results in a reward, while any malicious actions result in penalties. Finally, we simulate a network to evalu-ate the performance and security of PoG and compare it with sev-eral existing consensus algorithms. The experimental findings in-dicate that the proposed PoG consensus algorithm retains the fun-damental security properties of blockchain and outperforms the state-of-the-art consensus mechanisms.
Guangyong Gao, Chongtao Guo, Xinyu Wan, Zhihua Xia, Yun Q. Shi 0001
IEEE Internet Things J.4
2025 Distributor-centric model watermarking for image generative models
Jianwei Fei, Yunshu Dai, Zhihua Xia
Knowl. Based Syst.4
2025 AT-diff: An adversarial diffusion model for unrestricted adversarial examples generation
Chengsheng Yuan 0001, Jingfa Pang, Jianwei Fei, Xinting Li, Zhihua Xia
Knowl. Based Syst.5
2025 One-class network leveraging spectro-temporal features for generalized synthetic speech detection
Jiahong Ye, Diqun Yan, Songyin Fu, Bin Ma 0003, Zhihua Xia
Speech Commun.5
2025 Reversible Data Hiding-Based Contrast Enhancement With Adaptive Stretching Interval for ROI of Medical Image
abstract
Contrast enhancement methods based on reversible data hiding (RDHCE) can be used for contrast enhancement of medical images, which is a hot research topic in recent years. However, the region of interest (ROI) of medical images cannot be accurately segmented using the current RDHCE algorithms and histogram pixels clustering in medical images results in incorrect localization of longer intervals, which affects the contrast enhancement effect of images. In this paper, the Unet3+ network model is used, which makes the segmented ROI region and ROI histogram clearer and more accurate than those obtained by the traditional segmentation methods and the algorithm integrates a larger embedding capacity and a better visual quality of the image. It adaptively determines and stretches the interval of the ROI greyscale histogram and at the same time enlarges the embedding capacity of the ROI to enhance the contrast of the image. The proposed algorithm improves the visual quality of medical images by 20% and enhances ROI embedding capacity by 25% compared to existing methods.
Guangyong Gao, Xiangyang Hu, Sitian Yang, Zhihua Xia
IEEE Signal Process. Lett.4
2025 Compressed Domain Invariant Adversarial Representation Learning for Robust Audio Deepfake Detection
abstract
The primary aim of audio deepfake detection (ADD) is to thwart deception arising from forged audio generated through text-to-speech or voice conversion technologies. However, encoding speech signals using diverse compression algorithms introduces discrepancies that significantly impair the performance of existing countermeasure systems. To tackle these challenges, this letter proposes a robust audio deepfake detection method based on Compressed Domain Invariant Adversarial Representation Learning with Adaptive Token Pooling (DANet-ATP). This framework incorporates a Compression Codecs Discriminator (CCD) that, through adversarial learning in tandem with the backbone network, enhances the model's ability to extract more robust features across diverse compression codecs. Moreover, to efficiently prune redundant frame-level features while retaining vital spoofing cues, the letter designs a plug-and-play, parameter-free Adaptive Token Pooling module, significantly improving detection performance. Experimental results on the ASVspoof2021 DF dataset showcase the exceptional performance of the proposed model. Furthermore, a series of ablation experiments validate the validity and effectiveness of the proposed method.
Chengsheng Yuan 0001, Yifei Chen 0013, Zhili Zhou 0001, Zhihua Xia, Yongfeng Huang 0001
IEEE Signal Process. Lett.4
2025 Paradoxical Role of Adversarial Attacks: Enabling Crosslinguistic Attacks and Information Hiding in Multilingual Speech Recognition
abstract
With the rise of automatic speech recognition (ASR) research and practical applications, enabling adversarial attacks on ASR systems via subtle perturbations has become a priority. Most prior research has focused on single-language, single-model ASR systems. However, multilingual ASR systems hold opportunities for crosslinguistic attacks and covert message transmission. This letter introduces a new approach for crosslinguistic adversarial attacks in multilingual ASR, focusing on information hiding. For example, in military settings, adversarial examples applied to eavesdropping devices can encode messages detectable only by friendly devices, leaving adversaries, even with identical methods, unable to access them. This letter examines multilingual ASR system properties and introduces a crosslinguistic adversarial example with minimal perturbation, allowing friendly classifiers to extract hidden information while being undetectable by hostile classifiers. The experimental results on 5 models and 5 datasets show that the proposed method achieves a success rate of over 90% and an SNR close to 40 dB.
Zhihua Xia, Bin Ma 0003, Diqun Yan
IEEE Signal Process. Lett.2
2025 CHEAT: A Large-Scale Dataset for Detecting CHatGPT-writtEn AbsTracts
abstract
The powerful ability of ChatGPT has caused widespread concern in the academic community. Malicious users could synthesize dummy academic content through ChatGPT, which is extremely harmful to academic rigor and originality. The need to develop ChatGPT-written content detection algorithms calls for large-scale datasets. In this paper, we initially investigate the possible negative impact of ChatGPT on academia, and present a large-scale CHatGPT-writtEn AbsTract dataset (CHEAT) to support the development of detection algorithms. In particular, the ChatGPT-written abstract dataset contains 35,304 synthetic abstracts, with$Generation$,$Polish$, and$Fusion$as prominent representatives. Based on these data, we perform a thorough analysis of the existing text synthesis detection algorithms. We show that ChatGPT-written abstracts are detectable with well-trained detectors, while the detection difficulty increases with more human guidance involved.
Peipeng Yu, Xuan Feng 0002, Zhihua Xia
IEEE Trans. Big Data4
2025 Reversible Data Hiding-Based Local Contrast Enhancement With Nonuniform Superpixel Blocks for Medical Images
abstract
Reversible data hiding-based contrast enhancement can be applied to medical images, which not only allows the storage of patient information through reversible embedding, but also achieves image contrast enhancement, thereby assisting doctors in accurately diagnosing patient diseases. In response to the existing problems of mainstream methods, a novel reversible data hiding-based local contrast enhancement method for medical images is proposed. This method utilizes superpixel segmentation to segment medical images into multiple pixel blocks, and performs reversible data embedding and contrast enhancement for the pixel blocks within the region of interest (ROI). Additionally, a new embedding strategy is proposed. According to the contrast and texture features of each pixel block, histogram expansion of different degrees is carried out to effectively enhance the pixel blocks with low contrast, while avoiding excessive enhancement of the pixel blocks with high contrast. Experimental results demonstrate that, compared with the state-of-the-art mainstream methods, the proposed method not only improves the contrast in the ROI but also ensures high visual quality of the medical images.
Guangyong Gao, Sitian Yang, Xiangyang Hu, Zhihua Xia, Yun Q. Shi 0001
IEEE Trans. Circuits Syst. Video Technol.4
2025 Reversible Data Hiding in Encrypted Images With Adaptive Multi-Directional MED and Huffman Code Based on Interval-Wise Dynamic Prediction Axes
abstract
With the popularization of digital information, reversible data hiding in ciphertext has become a critical research focus in privacy protection in cloud storage. A reversible data hiding method for encrypted images is proposed: Reversible Data Hiding in Encrypted Images with Adaptive Multi-directional MED and Huffman Code based on Interval-Wise Dynamic Prediction Axes (RDHEI-AHIDA). Firstly, the original image is predicted by the gradient Adaptive Multi-Directional Median Edge Detector (AM-MED) to obtain the critical gradient and the position of the Interval-wise Dynamic Prediction Axes (IDP-Axes). Then, information bits are allocated at intervals on the IDP-Axes. Combining the determined position of the IDP-Axes and the critical gradient, the prediction error values of the original image are calculated and recorded. After the image is encrypted, according to the distribution of prediction error values, an adaptive Huffman code rule is established, and pixel marking, classification and auxiliary information embedding are carried out. Finally, the secret data is embedded by the bit replacement method. Compared with the state-of-the-art RDHEI methods, experimental results show that RDHEI-AHIDA not only provides a higher pure payload while ensuring security but also exhibits certain robustness.
Guangyong Gao, Yimin Yu, Zhihua Xia
IEEE Trans. Circuits Syst. Video Technol.4
2025 Preview Helps Selection: Previewable Image Watermarking With Client-Side Embedding
abstract
The increasing sharing of images on social networks is prompting the involvement of digital watermarking to protect copyright and combat illegal redistribution. Owner-side embedding and client-side embedding are two modes of digital watermarking, among which the latter enables better system scalability than the former due to its higher owner-side efficiency. However, the existing client-side watermarking schemes do not take into account the preview needs of users, in which users are prevented from acquiring any visual information about the original image before decryption because it is encrypted to be fully blurred. As a result, users cannot select the desired one by previewing when a batch of encrypted images is given. To solve this problem, we overcome the incompatibility between techniques and innovatively combine client-side watermarking with thumbnail-preserving encryption to render the degraded visual perception of the original image onto the encrypted one. Specifically, the image is first fully encrypted as usual client-side watermarking, and then pixel adjustments are performed to approximate the sum of the original pixels in each block for rendering the degraded visual perception. In this way, two schemes with different performance emphasis are proposed, which implement watermark embedding based on spread spectrum and quantization index modulation separately. In terms of performance evaluation, the security of both schemes is thoroughly demonstrated, and experiments are conducted to assess their feasibility, robustness, and efficiency.
Xiangli Xiao, Yushu Zhang 0001, Zhongyun Hua, Zhihua Xia, Jian Weng 0001
IEEE Trans. Dependable Secur. Comput.4
2025 A Universal Framework for Reversible Data Hiding in Encrypted Images With Multiple Hiders
abstract
Recently, Reversible Data Hiding in Encrypted Images with Multiple Hiders (RDHEI-MH) has attracted the attention of researchers, as it can satisfy the requirements of multiparty embedding. In this paper, a universal framework for RDHEI-MH is proposed, in which most of existing RDH algorithms in the plaintext domain (including Data Expansion (DE)-based and Histogram Shifting (HS)-based algorithms) can be applied in encrypted images without performance loss. That cannot be achieved by current existing schemes. The framework contains three entities: image owner, multiple hiders, and image receiver. The cover image is shared by Additive Secret Sharing (ASS) at the image owner’s side, and then each image share is sent to the corresponding hider. Based on the Secure Multiparty Computation (MPC) protocols specially designed for RDH operations, the hiders can securely perform computation on shares to embed additional data. After receiving all the marked image shares, the image receiver can reconstruct the marked image, and further extract the data. Compared with the previous RDHEI-MH schemes, the proposed scheme achieves full-separability, i.e., the extraction and decryption operations are commutative. The experimental results and theoretical analyses demonstrate that the proposed scheme has high visual quality of the reconstructed image and high embedding capacity, while its security is proved.
Lizhi Xiong, Zhihua Xia, Jian Weng 0001
IEEE Trans. Dependable Secur. Comput.3
2025 DGADM-GIS: Deterministic Guided Additive Diffusion Model for Generative Image Steganography
abstract
In recent years, generative steganography has witnessed remarkable progress in the field of covert communication. It leverages techniques such as generative adversarial networks (GANs) or flow-based generative models (GLOW) to generate stego images. However, these approaches often grapple with the dilemma of achieving optimal steganographic capacity while ensuring the accurate extraction of hidden information. Additionally, the models occasionally still generate low-quality images that are highly vulnerable to detection by steganalysis tools. To tackle the aforementioned challenges and enhance the overall performance of generative image steganography, this paper proposes the deterministic guided additive diffusion model for generative image steganography (DGADM-GIS). Initially, we devise a reversible mapping function that is used for deterministic guided by a provided secret message, and then construct a secret latent Gaussian vector. Moreover, the proposed DGADM-GIS framework designs an additive sampling method based on the superposition principle of normal distribution to obtain a Gaussian vector that satisfies independent, random and obeys the standard normal distribution, which is transformed to a stego image in a way of maintaining the distribution by the diffusion model. Furthermore, we conduct error analysis experiments on our proposed scheme and derive methods to enhance the accuracy of secret information extraction. The experimental results show that our proposed steganographic method exhibits robust resistance to steganalysis. When embedding 3 bits of secret information per pixel, it achieves nearly 100% extraction accuracy.
Chengsheng Yuan 0001, Zhaonan Ji, Xinting Li, Zhili Zhou 0001, Zhihua Xia, Q. M. Jonathan Wu
IEEE Trans. Dependable Secur. Comput.5
2025 Robust Generative Steganography for Image Hiding Using Concatenated Mappings
abstract
Generative steganography stands as a promising technique for information hiding, primarily due to its remarkable resistance to steganalysis detection. Despite its potential, hiding a secret image using existing generative steganographic models remains a challenge, especially in lossy or noisy communication channels. This paper proposes a robust generative steganography model for hiding full-size image. It lies on three reversible concatenated mappings proposed. The first mapping uses VQGAN with an order-preserving codebook to compress an image into a more concise representation. The second mapping incorporates error correction to further convert the representation into a robust binary representation. The third mapping devises a distribution-preserving sampling mapping that transforms the binary representation into the latent representation. This latent representation is then used as input for a text-to-image Diffusion model, which generates the final stego image. Experimental results show that our proposed scheme can freely customize the stego image content. Moreover, it simultaneously attains high stego and recovery image quality, high robustness, and provable security.
Bingwen Feng, Zhihua Xia, Wei Lu 0001, Jian Weng 0001
IEEE Trans. Inf. Forensics Secur.3
2025 Screen-Shooting Robust Watermark Based on Style Transfer and Structural Re-Parameterization
abstract
In real-world applications, screen capturing represents a significant scenario where this process can induce substantial distortion to the original image. Previous methods for simulating screen-shooting distortion often involved combining different formulas. We found that these simulation methods still have a significant gap compared to real distortions, making it urgently necessary to develop a realistic and credible comprehensive noise layer to achieve robustness against screen-shooting distortion. This paper presents a watermarking scheme capable of withstanding severe screen-shooting distortion. First, a dataset is constructed to train a screen-shooting distortion simulation network based on style transfer. Subsequently, a comprehensive noise layer is built upon this network to achieve robustness against severe screen-shooting distortion. Additionally, this paper incorporates structural re-parameterization techniques into the traditional U-shaped encoder to improve the quality of encoded images. Extensive experiments demonstrate the proposed scheme’s superior performance in terms of robustness and generalization, especially under severe screen-shooting distortion conditions.
Guangyong Gao, Xiaoan Chen, Li Li 0123, Zhihua Xia, Jianwei Fei, Yun Q. Shi 0001
IEEE Trans. Inf. Forensics Secur.4
2025 JPEG Compression-Resistant Generative Image Hiding Utilizing Cascaded Invertible Networks
abstract
Generative steganography is renowned for its exceptional undetectability. However, prevalent generative methods often have insufficient capacity for concealing secret images. Furthermore, the sensitivity of commonly utilized generative models exacerbates the challenge of ensuring robustness against channel distortions such as JPEG compression. In this paper, we introduce a generative image hiding network that employs two invertible generators to transform secret images into stego images within a disparate image domain. Additionally, we seamlessly integrate an up-and-down sampling module (UDM) within these generators to facilitate efficient decoupling of the intermediate representations obtained by each generator. The UDM serves multiple purposes: preserving coherence between the intermediate representations, enhancing resilience against JPEG compression, and safeguarding the confidentiality of the concealed images. To address the complexity of mapping both uncompressed and compressed stego images to a unified intermediary representation, we implement two distinct flows for the forward and backward processes of the generator associated with the stego images. The experimental results show that our scheme offers concurrent advantages in terms of full-size image hiding ability, undetectability, confidentiality, and robustness.
Tiewei Qin, Bingwen Feng, Bingbing Zhou, Jilian Zhang, Zhihua Xia, Jian Weng 0001, Wei Lu 0001
IEEE Trans. Inf. Forensics Secur.5
2025 Efficient and Privacy-Preserving Ride Matching Over Road Networks Against Malicious ORH Server
abstract
Online ride-hailing (ORH) services have become indispensable for our travel needs, offering the convenience of easily locating the nearest driver for riders through ride matching algorithms. However, existing ORH systems, such as Lyft and Didi, require users (both riders and drivers) to disclose their real-time location information during the matching process, thus giving rise to serious privacy concerns. Despite the proposal of various privacy-preserving ride-matching schemes, they remain insufficient in addressing potential malicious behaviors from the ORH server, such as colluding with designated drivers and deviation from computation protocols to interfere with the matching process. These behaviors lead to non-optimal matching results for riders. To address these issues, we present EMPRide, an efficient and privacy-preserving ride-matching scheme resistant to malicious ORH server. In EMPRide, we design an efficient and accurate computation of distances between users protocol, which integrates road network embedding and secure two-party computation. Additionally, we design a verification protocol that allows riders to verify the correctness of computed distances and matching results. Crucially, the communication overhead for riders in EMPRide remains constant, irrelevant to the number of available drivers. Our evaluation using real-world datasets demonstrates that EMPRide significantly outperforms existing solutions. Specifically, under identical conditions, in EMPRide, the computation speed on the ORH server is$19.22\times $faster and the communication cost is$8.08\times $less than state-of-the-art approaches. Moreover, riders experience a speed improvement of 4.84 orders of magnitude with$1.30\times $less communication, while drivers benefit from a 4.79 orders of magnitude speed increase with$1.45\times $less communication.
Mingtian Zhang, Anjia Yang, Jian Weng 0001, Min-Rong Chen, Huang Zeng, Yi Liu 0053, Zhihua Xia
IEEE Trans. Inf. Forensics Secur.8
2025 A Blockchain and Improved Perception Hash Based Copyright Protection Scheme for Purely Chromatic Background Images
abstract
Purely chromatic background images are widely used in computer wallpapers and advertisements, leading to issues such as copyright infringement and the loss of interest of holders. Image hashing is a technique used for comparing the similarity between images, and is often used for image verification, search, and copy detection due to its insensitivity to subtle changes in the original image. In a purely chromatic background image, the central detail of the image is the primary part and the key for copyright authentication. As the perception hash (pHash) algorithm only retains the low-frequency portion of the discrete cosine transform (DCT) matrix, it is unsuitable for purely chromatic background images. To deal with this issue, we propose an improved perception hash (ipHash) algorithm to enhance the universality of the algorithm by extracting purely chromatic background image features. Meanwhile, the development of image hashing is restricted due to the requirement of a trusted third party. To solve this issue, a secure blockchain-based image copyright protection scheme is designed. It realizes the copyright authentication and traceability, and overcomes the issue of a lack of trusted third parties. Experimental results show that the proposed method outperforms the state-of-theart image copyright protection schemes.
Guangyong Gao, Tongchao Feng, Chongtao Guo, Zhihua Xia, Yun Q. Shi 0001
IEEE Trans. Multim.4
2025 SEDN: A Spatiotemporal Encoder-Decoder Network for End-to-End Object Removal Forgery Detection in High-Resolution Videos
abstract
With the growing popularity of high-resolution (HR) video and the continuous growth of network bandwidth, the challenge of object removal detection in HR videos has attracted significant attention. Expert forgers leverage the rich detail in HR videos for meticulous pixel manipulation and apply sophisticated postprocessing techniques to hide high-frequency artifacts, thereby making forgery detection and localization more difficult when existing schemes are used. Additionally, the end-to-end framework simplifies the detection and localization process, which has not been considered in previous work. To solve the above issues, a spatiotemporal encoder−decoder network (SEDN) is proposed for end-to-end object removal forgery detection in HR videos. In the SEDN, a new model composed of a 3D asymmetric dual-stream network (3D-ADSN) and Transformer is proposed. The 3D-ADSN is utilized as the encoder, which fully integrates the high-frequency and low-frequency spatiotemporal information of videos. Transformer is utilized as the decoder to capture the global structure spatiotemporal information of the long-range feature sequence obtained by the encoder. This network combination successfully achieves simultaneous detection in the temporal and spatial domains without any additional postprocessing calculations. The experimental results demonstrate the better performance of the SEDN at different resolutions.
Lizhi Xiong, Linsen Ding, Mengqi Cao, Zhihua Xia, Yun Q. Shi 0001
IEEE Trans. Multim.4
2025 Dual-Decoupling With Frequency-Spatial Domains for Image Manipulation Localization
abstract
Leveraging trace-rich features within embedded spaces has been established as effective in image manipulation localization (IML). Nevertheless, the feature of manipulated traces frequently comprises substantial redundant information only loosely related to IML tasks. This complexity has hindered existing methods in fully comprehending the essence of trace features. In light of this challenge, we introduce a novel decoupling representation learning network (DRN) tailored for IML. The DRN excels at decoupling intricate multidomain information and transforming it into representations directly pertinent to IML objectives. This is achieved through a meticulously designed frequency decoupling representation learning module (FDM) and spatial decoupling representation learning module (SDM). Specifically, the FDM operates by acquiring distinct low and high-frequency components to effectively decouple redundant information. The decoupled high-frequency components are then harnessed as intricate trace complements, enhancing the overall aggregation process. In addition, the redundant information is expertly separated into authentic and manipulated representations through the use of channel activation maps in SDM. Through extensive experimentation on three public benchmarks including CASIA, NIST, and Coverage, our method consistently demonstrates superior performance and enhanced robustness compared with existing state-of-the-art methods.
Wenyan Pan, Wentao Ma 0003, Tongqing Zhou, Shan Zhao 0002, Lichuan Gu, Guolong Shi, Zhihua Xia
IEEE Trans. Neural Networks Learn. Syst.7
2024 Make Privacy Renewable! Generating Privacy-Preserving Faces Supporting Cancelable Biometric Recognition
Tao Wang 0084, Yushu Zhang 0001, Xiangli Xiao, Lin Yuan 0002, Zhihua Xia, Jian Weng 0001
ACM Multimedia5
2024 Once-for-all: Efficient Visual Face Privacy Protection via Person-specific Veils
abstract
As billions of face images stored on cloud platforms contain sensitive information to human vision, the public confronts substantial threats to visual face privacy. In response, the community has proposed some perturbation-based schemes to mitigate visual privacy leakage. However, these schemes need to generate a new protective perturbation for each image, failing to satisfy the real-time requirement of cloud platforms. To address this issue, we present an efficient visual face privacy protection scheme by utilizing person-specific veils, which can be conveniently applied to all images of the same user without regeneration. The protected images exhibit significant visual differences from the originals but remain identifiable to face recognition models. Furthermore, the protected images can be recovered to originals under certain circumstances. In the process of generating the veils, we propose a feature alignment loss to promote consistency between the recognition outputs of protected and original images with approximate construction of feature subspace. Meanwhile, the block variance loss is designed to enhance the concealment of visual identity information. Extensive experimental results demonstrate that our scheme can significantly eliminate the visual appearance of original images and almost has no impact on face recognition models.
Zixuan Yang 0004, Yushu Zhang 0001, Tao Wang 0084, Zhongyun Hua, Zhihua Xia, Jian Weng 0001
ACM Multimedia5
2024 Learning spatial-frequency interaction for generalizable deepfake detection
abstract
Abstract In recent years, face forgery detection has gained significant attention, resulting in considerable advancements. However, most existing methods rely on CNNs to extract artefacts from the spatial domain, overlooking the pervasive frequency‐domain artefacts present in deepfake content, which poses challenges in achieving robust and generalized detection. To address these issues, we propose the dual‐stream frequency—spatial fusion network is proposed for deepfake detection. The dual‐stream frequency‐spatial fusion network consists of three components: the spatial forgery feature extraction module, the frequency forgery feature extraction module, and the spatial–frequency feature fusion module. The spatial forgery feature extraction module employs spatial‐channel attention to extract spatial domain features, targeting artefacts in the spatial domain. The frequency forgery feature extraction module leverages the focused linear attention to detect frequency domain anomalies in internal regions, enabling the identification of generated content. The spatial–frequency feature fusion module then fuses forgery features extracted from both the spatial and frequency domains, facilitating accurate detection of splicing artefacts and internally generated forgeries. This approach enhances the model's ability to more accurately capture forgery characteristics. Extensive experiments on several widely‐used benchmarks demonstrate that our carefully designed network exhibits superior generalization and robustness, significantly improving deepfake detection performance.
Tianbo Zhai, Kaiyin Lu, Peipeng Yu, Zhihua Xia
IET Image Process.7
2024 Conditional image hiding network based on style transfer
Fenghua Zhang, Bingwen Feng, Zhihua Xia, Jian Weng 0001, Wei Lu 0001, Bing Chen 0004
Inf. Sci.3
2024 Moiré pattern generation-based image steganography
Tiewei Qin, Bingwen Feng, Bing Chen 0004, Zecheng Peng, Zhihua Xia, Wei Lu 0001
J. Inf. Secur. Appl.5
2024 Auto-focus tracing: Image manipulation detection with artifact graph contrastive
Wenyan Pan, Zhihua Xia, Wentao Ma 0003, Yuwei Wang 0002, Lichuan Gu, Guolong Shi, Shan Zhao 0002
Knowl. Based Syst.2
2024 Spatial-frequency gradient fusion based model augmentation for high transferability adversarial attack
Jingfa Pang, Chengsheng Yuan 0001, Zhihua Xia, Xinting Li, Zhangjie Fu 0001
Knowl. Based Syst.3
2024 Face Omron Ring: Proactive defense against face forgery with identity awareness
Yunshu Dai, Jianwei Fei, Fangjun Huang, Zhihua Xia
Neural Networks4
2024 Robust image hiding network with Frequency and Spatial Attentions
Xiaobin Zeng, Bingwen Feng, Zhihua Xia, Zecheng Peng, Tiewei Qin, Wei Lu 0001
Pattern Recognit.3
2024 STR: Secure Computation on Additive Shares Using the Share-Transform-Reveal Strategy
abstract
The rapid development of cloud computing probably benefits many of us while the privacy risks brought by semi-honest cloud servers have aroused the attention of more and more people and legislatures. In the last two decades, plenty of works seek to outsource various specific tasks to servers while ensuring the security of private data. The tasks to be outsourced are countless; however, the computations involved are similar. In this article, we construct a series of novel protocols that support the secure computation of various functions on numbers (e.g., the basic elementary functions) and matrices (e.g., the calculation of eigenvectors and eigenvalues) on arbitrary$n\geq 2$servers. All protocols only require constant rounds of interactions and achieve low computation complexity. Moreover, the proposed$n$-party protocols ensure the security of private data even though$n-1$servers collude. The convolutional neural network models are utilized as the case studies to verify the protocols. The theoretical analysis and experimental results demonstrate the correctness, efficiency, and security of the proposed protocols.
Zhihua Xia, Lizhi Xiong, Jian Weng 0001, Naixue Xiong
IEEE Trans. Computers1
2024 Removing Hidden Information by Geometrical Perturbation in Frequency Domain
abstract
The risk of malicious exploitation of advanced image steganography necessitates the removal of hidden information from images. However, it is crucial to preserve the visual quality of the images undergoing processed. This paper suggests a geometrical attack in frequency domain (GAF) to address this challenge. GAF employs a thin plate spline (TPS) to slightly geometrically perturb the frequency components of the stego image. It incorporates a channel weight estimator and a frequency jammer. The channel weight estimator assigns perturbation strengths to each DCT channel, while the frequency jammer performs the TPS transform on the DCT channels using the assigned perturbation strengths. Experimental results demonstrate that the proposed approach effectively hinders secret image recovery with a little distortion to the stego images. Furthermore, it well preserves the visual quality of clear images that do not contain secret information.
Bingwen Feng, Zecheng Peng, Bing Chen 0004, Zhihua Xia, Wei Lu 0001
IEEE Trans. Circuits Syst. Video Technol.5
2024 Image Manipulation Detection With Cascade Hierarchical Graph Representation
abstract
Recent image manipulation detection approaches primarily rely on sophisticated Convolutional Neural Network (CNN)-based models for region localization, while they tend to ignore: 1) the feature correlations that exist between manipulated and non-manipulated regions. 2) the significance of multi-scale representations in detecting manipulated regions of varying sizes, consequently hampering the overall performance of image manipulation detection. To address these limitations, we propose a novel approach, called Cascade Hierarchical Graph Convolutional Network (Cas-HGCN), which comprehensively learns the feature correlations between manipulated and non-manipulated regions at different scales using the Feature Correlations Modeling (FCM) module. Specifically, the FCM module treats the grids in the hierarchical image/feature maps as nodes, constructs a fully-connected graph by connecting each node, and leverages it to learn and refine feature correlations across different scales in a cascading manner. This process results in high discriminability for distinguishing manipulated and non-manipulated regions. Extensive experiments conducted on three public datasets, namely CASIA, NIST, and Coverage, demonstrate the promising detection accuracy achieved by Cas-HGCN without the need for pre-training on large datasets, surpassing the performance of existing state-of-the-art competitors.
Wenyan Pan, Wentao Ma 0003, Shan Zhao 0002, Lichuan Gu, Guolong Shi, Zhihua Xia, Meng Wang 0001
IEEE Trans. Circuits Syst. Video Technol.6
2024 Wide Flat Minimum Watermarking for Robust Ownership Verification of GANs
abstract
We propose a novel multi-bit box-free watermarking method for the protection of Intellectual Property Rights (IPR) of GANs with improved robustness against white-box model-level attacks like fine-tuning, pruning, quantization, and surrogate model attacks. The watermark is embedded by adding an extra watermarking loss term during GAN training, ensuring that the images generated by the GAN contain an invisible watermark that can be retrieved by a pre-trained watermark decoder. In order to improve the robustness against white-box model-level attacks, we make sure that the model converges to a wide flat minimum of the watermarking loss term, in such a way that any modification of the model parameters does not erase the watermark. To do so, we add random noise vectors to the parameters of the generator and require that the watermarking loss term is as invariant as possible with respect to the presence of noise. This procedure forces the generator to converge to a wide flat minimum of the watermarking loss. The proposed method is architecture- and dataset-agnostic, thus being applicable to many different generation tasks and models, as well as to CNN-based image processing architectures. We present the results of extensive experiments showing that the presence of the watermark has a negligible impact on the quality of the generated images, and proving the superior robustness of the watermark against model modification and surrogate model attacks.
Jianwei Fei, Zhihua Xia, Benedetta Tondi, Mauro Barni
IEEE Trans. Inf. Forensics Secur.2
2024 Client-Side Embedding of Screen-Shooting Resilient Image Watermarking
abstract
The proliferation of portable camera devices, represented by smartphones, is increasing the risk of sensitive internal data being leaked by screen shooting. To trace the leak source, a lot of research has been done on screen-shooting resilient watermarking technique, which is capable of extracting the previously embedded watermark from the screen-shot image. However, all existing screen-shooting resilient watermarking schemes follow the owner-side embedding mode. In this mode, the management center will suffer heavy computational and communication burden in the case of numerous screens, which hinders the system scalability. As another embedding mode of digital watermarking, client-side embedding can solve the above scalability problem by migrating the watermark embedding operation to the same time when the screen decrypts the image. By designing a pair of image encryption and personalized decryption algorithms based on matrix operation, this paper is the first to realize the client-side embedding of screen-shooting resilient watermarking. In this implementation, challenges are overcome and the following key achievements are attained. First, our scheme embeds watermark using the algorithm of Fanget al. without modification, and thus fully inherits its robustness against screen shooting. Second, the original image is securely encrypted and the watermarked image can be directly retrieved through decryption. Third, the secrecy of the screen watermark is ensured by concealing the embedding pattern. Finally, our scheme is validated by experiments, which shows that the efficiency advantage of client-side embedding is realized while maintaining robustness.
Xiangli Xiao, Yushu Zhang 0001, Zhongyun Hua, Zhihua Xia, Jian Weng 0001
IEEE Trans. Inf. Forensics Secur.4
2024 Understanding Visual Privacy Protection: A Generalized Framework With an Instance on Facial Privacy
abstract
With the widespread application of computer vision, the scenarios in terms of visual privacy have become increasingly diverse and meanwhile numerous studies have been conducted to address privacy concerns in these scenarios. However, these studies are individually tailored for specific scenarios, making their layouts challenging to be drawn upon easily. When encountering a new scenario, it takes significant additional efforts to redesign a scheme due to the low referability of previous works. To tackle this issue, we explore commonalities among existing works and propose a generalized framework to meet the demand for visual privacy protection in various scenarios. Our framework is elaborately organized into several crucial steps, including privacy definition, scenario abstraction, algorithm design, and effect evaluation. It serves as a guide for researchers to efficiently design visual privacy protection schemes. In our framework, we establish a unified standard for quantifying privacy and introduce a novel constrained optimization theory to balance privacy and usability, which contributes to a broader understanding of visual privacy protection. Furthermore, we present an instance under the guidance of the framework that can support identity protection and attribute control scenarios through a diffusion-based model. Extensive experimental results demonstrate the effectiveness of our framework.
Yushu Zhang 0001, Junhao Ji, Wenying Wen, Youwen Zhu, Zhihua Xia, Jian Weng 0001
IEEE Trans. Inf. Forensics Secur.5
2024 Watermarking in Secure Federated Learning: A Verification Framework Based on Client-Side Backdooring
abstract
Federated learning (FL) allows multiple participants to collaboratively build deep learning (DL) models without directly sharing data. Consequently, the issue of copyright protection in FL becomes important since unreliable participants may gain access to the jointly trained model. Application of homomorphic encryption (HE) in a secure FL framework prevents the central server from accessing plaintext models. Thus, it is no longer feasible to embed the watermark at the central server using existing watermarking schemes. In this article, we propose a novel client-side FL watermarking scheme to tackle the copyright protection issue in secure FL with HE. To the best of our knowledge, it is the first scheme to embed the watermark to models under a secure FL environment. We design a black-box watermarking scheme based on client-side backdooring to embed a pre-designed trigger set into an FL model by a gradient-enhanced embedding method. Additionally, we propose a trigger set construction mechanism to ensure that the watermark cannot be forged. Experimental results demonstrate that our proposed scheme delivers outstanding protection performance and robustness against various watermark removal attacks and ambiguity attack.
Shuo Shao 0002, Yue Yang 0007, Xiyao Liu 0001, Ximeng Liu, Zhihua Xia, Gerald Schaefer, Hui Fang 0003
ACM Trans. Intell. Syst. Technol.6
2024 Reversible Data Hiding-Based Contrast Enhancement With Multi-Group Stretching for ROI of Medical Image
abstract
Reversible data hiding-based contrast enhancement (RDHCE) can be used in contrast enhancement for medical images, and it has been a popular research topic in recent years. However, the existing RDHCE methods suffer from the problem of inaccurate segmentation of the region of interest (ROI) in medical images, which can impact the contrast enhancement effect of the images. Moreover, some methods face limitations in their universality for ROI histograms with few empty bins on both sides, which results in unsatisfactory embedding capacity and contrast enhancement effect. To solve these problems, this study proposes an improved RDHCE method for medical images. The proposed method uses the UNet3+ network model, which makes the segmented ROI histograms more consistent with the subjective judgment of doctors compared to those obtained by traditional segmentation approaches. In addition, a multi-group stretching method is proposed to address the limitation of histogram expansion caused by the empty bins on both histogram sides, enabling adaptation to different ROI histograms with varying gray distributions. Compared to state-of-the-art RDHCE methods, the proposed method offers better generalizability, superior contrast enhancement performance and a larger ROI embedding capacity. It can greatly improve the visual quality of medical images in the field of medical imaging and aid doctors in making more accurate diagnoses.
Guangyong Gao, Hui Zhang 0137, Zhihua Xia, Xiangyang Luo 0001, Yun Q. Shi 0001
IEEE Trans. Multim.3
2024 A Bitcoin-based Secure Outsourcing Scheme for Optimization Problem in Multimedia Internet of Things
abstract
With the development of the Internet of Things (IoT) and cloud computing, various multimedia data such as audio, video, and images have experienced explosive growth, ushering in the era of big data. Large-scale computing tasks in the Multimedia Internet of Things (M-IoT), such as mathematical optimization problems, have begun to be outsourced from IoT devices with limited computing power to cloud servers for execution. However, outsourcing computation brings security concerns, because the behaviors of clouds are invisible to users. The leakage of privacy data in outsourced optimization problems leads to immeasurable losses. The mutual distrust between clouds and users causes that the correctness of the optimal decisions and the fairness of the payment activities are not guaranteed. Blockchain technology has the characteristic of immutability and has become a new security paradigm for eliminating multi-party trust concerns. In this article, we propose a Bitcoin-based secure outsourcing scheme to address the aforementioned security concerns. To prevent confidential data leakage, the proposed scheme designs a computable privacy-preserving method for the outsourced optimization problems. To judge the correctness of the optimal decision and reduce verification costs, the proposed scheme designs a low-cost two-layer verification mechanism based on dual theory and blockchain technology. Blockchain nodes reach a consensus on the problem solutions and trigger an automatic fair payment protocol-based Bitcoin. Security analysis and experimental results demonstrate that our scheme guarantees privacy, fairness, and computational efficiency.
Shaocong Wu, Jianwei Fei, Xianwang Zeng, Yuemin Ding, Zhihua Xia
ACM Trans. Multim. Comput. Commun. Appl.6
2023 General GAN-generated Image Detection by Data Augmentation in Fingerprint Domain
abstract
In this work, we investigate improving the generalizability of GAN-generated image detectors by performing data augmentation in the fingerprint domain. Specifically, we first separate the fingerprints and contents of the GAN-generated images using an autoencoder based GAN fingerprint extractor, followed by random perturbations of the fingerprints. Then the original fingerprints are substituted with the perturbed fingerprints and added to the original contents, to produce images that are visually invariant but with distinct fingerprints. The perturbed images can successfully imitate images generated by different GANs to improve the generalization of the detectors, which is demonstrated by the spectra visualization. To our knowledge, we are the first to conduct data augmentation in the fingerprint domain. Our work explores a novel prospect that is distinct from previous works on spatial and frequency domains augmentation. Extensive cross-GAN experiments demonstrate the effectiveness of our method compared to the state-of-the-art methods in detecting fake images generated by unknown GANs.
Huaming Wang, Jianwei Fei, Yunshu Dai, Lingyun Leng, Zhihua Xia
ICME5
2023 Secure outsourced NB: Accurate and efficient privacy-preserving Naive Bayes classification
Xueli Zhao, Zhihua Xia
Comput. Secur.2
2023 A screen-shooting resilient document image watermarking scheme using deep neural network
abstract
Abstract With the advent of the screen‐reading era, the confidential documents displayed on the screen can be easily captured by a camera without leaving any traces. Thus, this paper proposes a novel screen‐shooting resilient watermarking scheme for document image using deep neural network. By applying this scheme, when the watermarked image is displayed on the screen and captured by a camera, the watermark can be still extracted from the captured photographs. Specifically, the scheme is an end‐to‐end neural network with an encoder to embed watermark and a decoder to extract watermark. During the training process, a distortion layer between encoder and decoder is added to simulate the distortions introduced by screen‐shooting process in real scenes, such as camera distortion, shooting distortion, and light source distortion. Furthermore, a background sensitive loss and a lpips loss are used to improve visual quality of the watermarked document images in the training process. Besides, a strength factor adjustment strategy is also designed to improve the visual quality with little loss of bit extraction accuracy. The experimental results show that the proposed scheme has higher visual quality and robustness than the other two recent state‐of‐the‐art methods.
Sulong Ge, Jianwei Fei, Zhihua Xia, Jian Weng 0001, Jia-Nan Liu
IET Image Process.3
2023 A Data Integrity Authentication Scheme in WSNs Based on Double Watermark
abstract
Aiming at data security in wireless sensor networks (WSNs), a data authentication scheme based on a double watermark is proposed in this article. Double watermark includes reversible watermark and irreversible watermark. The former generated by group head data is embedded into the effective precision bit of data, and the latter generated by the effective precision bit of the data themselves and a defined array is embedded into the flag bit in two cases. Two adjacent groups form a working group. For the first case, a flag-check matrix is constructed by the special data set which includes the head of the current group and the tail of the previous group, and the data around them. Through the matrix, the watermark is embedded into the special data set to ensure the robustness of grouping. For the second case, watermark is embedded into the group data except special data set (the rest data) to ensure efficient detection of attacks. When the attacked position and type are detected, the data that are not tampered with in this group can still recover the effective precision bit. The security analysis and experimental results demonstrate that the comprehensive performance of the proposed scheme outperforms that of those state-of-the-art schemes.
Guangyong Gao, Zhihua Xia
IEEE Internet Things J.5
2023 A robust document image watermarking scheme using deep neural network
Sulong Ge, Zhihua Xia, Jianwei Fei, Jian Weng 0001, Ming Li 0049
Multim. Tools Appl.2
2023 Investigating prosodic entrainment from global conversations to local turns and tones in Mandarin conversations
Zhihua Xia, Julia Hirschberg, Rivka Levitan
Speech Commun.1
2023 Deepfake Fingerprint Detection Model Intellectual Property Protection via Ridge Texture Enhancement
abstract
In addition to relying on super computing power and professional domain knowledge, training a high-precision deepfake fingerprint detection model (DFDM) to authenticate the fingerprints also requires the support of massive private fingerprint data. To sum up, the DFDM should be deemed as intellectual property (IP) of the trainers, so it is crucial to protect IP. Currently, most watermarking-based IP protection schemes are implemented by introducing additional tasks, such as constructing trigger sets, fine-tuning model weights, etc., which severely impair the performance of the original task and increase the training cost. Inspired by the feature knowledge learned by the model, this letter proposes an IP protection (IPP) scheme for DFDM by verifying whether the suspect model contains the fingerprint ridge features learned by the victim model from another perspective based on the verifier. Firstly, the local binary pattern (LBP) is used to enhance the ridge texture on the fingerprint samples, so that DFDM can better learn the ridge features. Then, a DFDM lacking texture augmentation is employed as the adversarial model for training the meta-verifier without any alteration to the model parameters. Finally, the trained meta-verifier is used to determine whether the suspected model contains the ridge features in the victim model. The public fingerprint dataset (LivDet2017) was leveraged in the DFDM training process to validate our approach. Experimental results show that the proposed scheme can verify the IP of DFDM and is robust to some common attacks.
Chengsheng Yuan 0001, Zhili Zhou 0001, Zhangjie Fu 0001, Zhihua Xia
IEEE Signal Process. Lett.5
2023 A Universal Reversible Data Hiding Method in Encrypted Image Based on MSB Prediction and Error Embedding
abstract
Image encryption is used for privacy protection in cloud computing. Nowadays, reversible data hiding in encrypted image (RDHEI) has achieved great success with the demand of embedding additional information into encrypted image. The existing algorithms cannot implement large embedding capacity and good reconstructed image quality simultaneously. Besides, the universality of some methods is limited when they are used in images with different textural characteristics. In this work, an RDHEI method based on most significant bit (MSB) prediction and error embedding is proposed. On one hand, all types of prediction errors are considered in the proposed method, therefore all the pixels with prediction errors can be recovered correctly. On the other hand, error blocks are utilized to mark the locations of prediction errors and message blocks are utilized to embed data. Moreover, flag blocks are utilized to distinguish error and message blocks. To solve the problem of misjudgement for flag blocks in the decoding phase, special operations are conducted on error blocks and message blocks. Experimental results demonstrate that, compared with the state-of-the-art RDHEI methods, the proposed method has good universality on well-known databases.
Guangyong Gao, Shikun Tong, Zhihua Xia, Yun Q. Shi 0001
IEEE Trans. Cloud Comput.3
2023 A Privacy-Preserving JPEG Image Retrieval Scheme Using the Local Markov Feature and Bag-of-Words Model in Cloud Computing
abstract
The development of cloud computing attracts a great deal of image owners to upload their images to the cloud server to save the local storage. But privacy becomes a great concern to the owner. A forthright way is to encrypt the images before uploading, which, however, would obstruct the efficient usage of image, such as the Content-Based Image Retrieval (CBIR). In this paper, we propose a privacy-preserving JPEG image retrieval scheme. The image content is protected by a specially-designed image encryption method, which is compatible to JPEG compression and makes no expansion to the final JPEG files. Then, the encrypted JPEG files are uploaded to the cloud, and the cloud can directly extract the features from the encrypted JPEG files for searching similar images. Specifically, big-blocks are first assembled with adjacent 8×8 discrete cosine transform (DCT) coefficient blocks. Then, the big-blocks are permuted and the binary code of DCT coefficients are substituted, so as to disturb the content of image. After receiving the encrypted images, local Markov features are extracted from the encrypted big-blocks, and then the Bag-Of-Words (BOW) model is applied to construct a feature vector with these local features to represent the image, so as to provide the CBIR service to image owner. Experimental results and security analysis demonstrate the retrieval performance and security of our scheme.
Peipeng Yu, Jian Tang 0009, Zhihua Xia, Zhetao Li, Jian Weng 0001
IEEE Trans. Cloud Comput.3
2023 AVoiD-DF: Audio-Visual Joint Learning for Detecting Deepfake
abstract
Recently, deepfakes have raised severe concerns about the authenticity of online media. Prior works for deepfake detection have made many efforts to capture the intra-modal artifacts. However, deepfake videos in real-world scenarios often consist of a combination of audio and visual. In this paper, we propose an Audio-Visual Joint Learning for Detecting Deepfake (AVoiD-DF), which exploits audio-visual inconsistency for multi-modal forgery detection. Specifically, AVoiD-DF begins by embedding temporal-spatial information in Temporal-Spatial Encoder. A Multi-Modal Joint-Decoder is then designed to fuse multi-modal features and jointly learn inherent relationships. Afterward, a Cross-Modal Classifier is devised to detect manipulation with inter-modal and intra-modal disharmony. Since existing datasets for deepfake detection mainly focus on one modality and only cover a few forgery methods, we build a novel benchmark DefakeAVMiT for multi-modal deepfake detection. DefakeAVMiT contains sufficient visuals with corresponding audios, where any one of the modalities may be maliciously modified by multiple deepfake methods. The experimental results on DefakeAVMiT, FakeAVCeleb, and DFDC demonstrate that the AVoiD-DF outperforms many state-of-the-arts in deepfake detection. Our proposed method also yields superior generalization on various forgery techniques.
Bofei Guo, Zhongjie Ba, Zhihua Xia, Xiaochun Cao, Kui Ren 0001
IEEE Trans. Inf. Forensics Secur.6
2023 Secure Outsourced SIFT: Accurate and Efficient Privacy-Preserving Image SIFT Feature Extraction
abstract
Cloud computing has become an important IT infrastructure in the big data era; more and more users are motivated to outsource the storage and computation tasks to the cloud server for convenient services. However, privacy has become the biggest concern, and tasks are expected to be processed in a privacy-preserving manner. This paper proposes a secure SIFT feature extraction scheme with better integrity, accuracy and efficiency than the existing methods. SIFT includes lots of complex steps, including the construction of DoG scale space, extremum detection, extremum location adjustment, rejecting of extremum point with low contrast, eliminating of the edge response, orientation assignment, and descriptor generation. These complex steps need to be disassembled into elementary operations such as addition, multiplication, comparison for secure implementation. We adopt a serial of secret-sharing protocols for better accuracy and efficiency. In addition, we design a secure absolute value comparison protocol to support absolute value comparison operations in the secure SIFT feature extraction. The SIFT feature extraction steps are completely implemented in the ciphertext domain. And the communications between the clouds are appropriately packed to reduce the communication rounds. We carefully analyzed the accuracy and efficiency of our scheme. The experimental results show that our scheme outperforms the existing state-of-the-art.
Xiang Liu 0020, Xueli Zhao, Zhihua Xia, Peipeng Yu, Jian Weng 0001
IEEE Trans. Image Process.3
2023 RDH-DES: Reversible Data Hiding over Distributed Encrypted-Image Servers Based on Secret Sharing
abstract
Reversible Data Hiding in Encrypted Image (RDHEI) schemes may redistribute the data hiding procedure to other parties and can preserve privacy of the cover image. Recently, cloud computing technology has led to the rapid growth of networked media, and many multimedia rights are owned by multiple parties, such as a film's producer and multiple distributors. Thus, the data hiding task could be distributed to multiple distributed servers. Multi-party data hiding has become an important demand for networked media. In addition, it is essential to preserve multi-server and multi-message privacy and data integrity. However, most of the RDHEI schemes involve only one data hider. That inspired us to design the secure multi-party embedding over distributed encrypted-image servers as a solution for multi-party RDHEI applications. In this article, we propose a novel Reversible Data Hiding over Distributed Encrypted-Image Servers (RDH-DES) based on secret sharing. The Chinese remainder theorem, secret sharing, and block-level scrambling are developed as a lightweight cryptography to generate the encrypted image shares. These shares are distributed to different image servers and are used to embed secret data in the proposed framework. The marked encrypted image can be constructed through the marked encrypted shares from different parties, and the decryption and extraction can be completed by the receiver. The experimental results and theoretical analysis have demonstrated that the proposed scheme is secure and effective.
Lizhi Xiong, Ching-Nung Yang, Zhihua Xia
ACM Trans. Multim. Comput. Commun. Appl.4
2022 Learning Second Order Local Anomaly for General Face Forgery Detection
abstract
In this work, we propose a novel method to improve the generalization ability of CNN-based face forgery detectors. Our method considers the feature anomalies of forged faces caused by the prevalent blending operations in face forgery algorithms. Specifically, we propose a weakly supervised Second Order Local Anomaly (SOLA) learning module to mine anomalies in local regions using deep feature maps. SOLA first decomposes the neighborhood of local features by different directions and distances and then calculates the first and second order local anomaly maps which provide more general forgery traces for the classifier. We also propose a Local Enhancement Module (LEM) to improve the discrimination between local features of real and forged regions, so as to ensure accuracy in calculating anomalies. Besides, an improved Adaptive Spatial Rich Model (ASRM) is introduced to help mine subtle noise features via learnable high pass filters. With neither pixel level annotations nor external synthetic data, our method using a simple ResNet18 backbone achieves competitive performances compared with state-of-the-art works when evaluated on unseen forgeries.
Jianwei Fei, Yunshu Dai, Peipeng Yu, Tianrun Shen, Zhihua Xia, Jian Weng 0001
CVPR5
2022 Attentional Local Contrastive Learning for Face Forgery Detection
Yunshu Dai, Jianwei Fei, Huaming Wang, Zhihua Xia
ICANN (1)4
2022 MSPPIR: Multi-Source Privacy-Preserving Image Retrieval in cloud computing
Zhihua Xia, Xingming Sun
Future Gener. Comput. Syst.2
2022 Coverless Information Hiding Based on Probability Graph Learning for Secure Communication in IoT Environment
abstract
To securely transmit secret data between Internet of Things (IoT) nodes, it is required to the implement information hiding technique for secure communication in the IoT environment. The traditional information hiding approaches generally select a multimedia file, such as texts, images, and video clips as the cover, and then embed secret information into the cover by slight modification. However, it is not feasible to directly apply these approaches in the IoT environment for the following reasons. First, it is hard for some IoT nodes to effectively and efficiently process and transmit the complex multimedia data. Second, the modification trace left in the cover will cause the presence of hidden secret information to be easily exposed by steganalysis tools. To address the above issues, we propose a coverless information hiding scheme based on probability graph learning for secure communication in the IoT environment. Instead of modifying an existing multimedia cover, we conceal secret information in a generated sequence of IoT data to realize secure communication between different nodes in the IoT environment. According to the node-data interaction relationships, we first learn the transition probability graph (TPG) to describe the transition probabilities between IoT data elements. Then, guided by a given secret message that needs to be hidden, we sequentially select a set of highly correlated data elements from the TPG to generate the sequence. The experimental results and theoretical analysis demonstrate that the proposed information hiding scheme can achieve high hiding capacity with desirable imperceptibility and security performances in the IoT environment.
Zhili Zhou 0001, Yuecheng Su, Yulan Zhang, Zhihua Xia, Shan Du 0001, Brij B. Gupta, Lianyong Qi
IEEE Internet Things J.4
2022 Improving Generalization by Commonality Learning in Face Forgery Detection
abstract
This paper proposes a commonality learning strategy for face video forgery detection to improve the generalization. Considering various face forgery methods could leave certain similar forgery traces in videos, we attempt to learn the common forgery features from different forgery databases, so as to achieve better generalization in the detection of unknown forgery methods. Firstly, the Specific Forgery Feature Extractors (SFFExtractors) are trained separately for each of given forgery methods. We utilize the U-net structure and consider the triplet loss, location loss, classification loss, and automatic weighted loss to ensure the detection ability of SFFExtractors on the corresponding forgery methods. Next, the Common Forgery Feature Extractor (CFFExtractor) is trained under the supervision of SFFExtractors to explore the commonality of the forgery traces caused by different forgery methods. The extracted common forgery feature is expected to have a good generalization. The experimental results on FaceForensic++ show that the SFFExtractors outperform many state-of-the-arts in face forgery detection. The generalization performance of the CFFExtractor is verified on FaceForensic++, DFDC, and CelebDF. It is proved that commonality learning can be an effective strategy to improve generalization.
Peipeng Yu, Jianwei Fei, Zhihua Xia, Zhili Zhou 0001, Jian Weng 0001
IEEE Trans. Inf. Forensics Secur.3
2022 A Format-compatible Searchable Encryption Scheme for JPEG Images Using Bag-of-words
abstract
The development of cloud computing attracts enterprises and individuals to outsource their data, such as images, to the cloud server. However, direct outsourcing causes the extensive concern of privacy leakage, as images often contain rich sensitive information. A straightforward way to protect privacy is to encrypt the images using the standard cryptographic tools before outsourcing. However, in such a way the possible usage of the outsourced images would be strongly limited together with the services provided to users, like the Content-Based Image Retrieval (CBIR). In this article, we propose a secure outsourced CBIR scheme, in which an encryption scheme is designed for the widely used JPEG-format images, and the secure features can be directly extracted from such encrypted images. Specifically, the JPEG images are encrypted by the block permutation, intra-block permutation, polyalphabetic cipher, and stream cipher. Then secure local histograms are extracted from the encrypted DCT blocks and the Bag-Of-Words (BOW) model is further used to organize the encrypted local features to represent the image. The proposed image encryption gets all of the image data protected and the experimental results show that the proposed scheme achieves improved accuracy with a small file size expansion.
Zhihua Xia, Qiuju Ji, Chengsheng Yuan 0001, Fengjun Xiao
ACM Trans. Multim. Comput. Commun. Appl.1
2022 BOEW: A Content-Based Image Retrieval Scheme Using Bag-of-Encrypted-Words in Cloud Computing
abstract
Content-based Image Retrieval (CBIR) techniques have been extensively studied with the rapid growth of digital images. Generally, CBIR service is quite expensive in computational and storage resources. Thus, it is a good choice to outsource CBIR service to the cloud server that is equipped with enormous resources. However, the privacy protection becomes a big problem, as the cloud server cannot be fully trusted. In this paper, we propose an outsourced CBIR scheme based on a novel bag-of-encrypted-words (BOEW) model. The image is encrypted by color value substitution, block permutation, and intra-block pixel permutation. Then, the local histograms are calculated from the encrypted image blocks by the cloud server. All the local histograms are clustered together, and the cluster centers are used as the encrypted visual words. In this way, the bag-of-encrypted-words (BOEW) model is built to represent each image by a feature vector, i.e., a normalized histogram of the encrypted visual words. The similarity between images can be directly measured by the Manhattan distance between feature vectors on the cloud server side. Experimental results and security analysis on the proposed scheme demonstrate its search accuracy and security.
Zhihua Xia, Leqi Jiang, Byeungwoo Jeon
IEEE Trans. Serv. Comput.1
2021 Exposing AI-generated videos with motion magnification
Jianwei Fei, Zhihua Xia, Peipeng Yu, Fengjun Xiao
Multim. Tools Appl.2
2021 Channel-Wise Spatiotemporal Aggregation Technology for Face Video Forensics
abstract
Recent progress in deep learning, in particular the generative models, makes it easier to synthesize sophisticated forged faces in videos, leading to severe threats on social media about personal privacy and reputation. It is therefore highly necessary to develop forensics approaches to distinguish those forged videos from the authentic. Existing works are absorbed in exploring frame-level cues but insufficient in leveraging affluent temporal information. Although some approaches identify forgeries from the perspective of motion inconsistency, there is so far not a promising spatiotemporal feature fusion strategy. Towards this end, we propose the Channel-Wise Spatiotemporal Aggregation (CWSA) module to fuse deep features of continuous video frames without any recurrent units. Our approach starts by cropping the face region with some background remained, which transforms the learning objective from manipulations to the difference between pristine and manipulated pixels. A deep convolutional neural network (CNN) with skip connections that are conducive to the preservation of detection-helpful low-level features is then utilized to extract frame-level features. The CWSA module finally makes the real or fake decision by aggregating deep features of the frame sequence. Evaluation against a list of large facial video manipulation benchmarks has illustrated its effectiveness. On all three datasets, FaceForensics++, Celeb-DF, and DeepFake Detection Challenge Preview, the proposed approach outperforms the state-of-the-art methods with significant advantages.
Yujiang Lu, Yaju Liu, Jianwei Fei, Zhihua Xia
Secur. Commun. Networks4
2021 PLDP: Personalized Local Differential Privacy for Multidimensional Data Aggregation
abstract
The collection of multidimensional crowdsourced data has caused a public concern because of the privacy issues. To address it, local differential privacy (LDP) is proposed to protect the crowdsourced data without much loss of usage, which is popularly used in practice. However, the existing LDP protocols ignore users’ personal privacy requirements in spite of offering good utility for multidimensional crowdsourced data. In this paper, we consider the personality of data owners in protection and utilization of their multidimensional data by introducing the notion of personalized LDP (PLDP). Specifically, we design personalized multiple optimized unary encoding (PMOUE) to perturb data owners’ data, which satisfies ϵ total -PLDP. Then, the aggregation algorithm for frequency estimation on multidimensional data under PLDP is developed, which is described in two situations. Experiments are conducted on four real datasets, and the results show that the proposed aggregation algorithm yields high utility. Moreover, case studies with four real datasets demonstrate the efficiency and superiority of the proposed scheme.
Zixuan Shen, Zhihua Xia, Peipeng Yu
Secur. Commun. Networks2
2021 Reversible data hiding with automatic contrast enhancement for medical images
Guangyong Gao, Shikun Tong, Zhihua Xia, Bin Wu 0021, Liya Xu
Signal Process.3
2020 A Privacy-Preserving Outsourcing Scheme for Image Local Binary Pattern in Secure Industrial Internet of Things
abstract
In the era of Industrial Internet of Things (IIoT), huge amounts of data are generated, and companies are highly motivated to store the data on cloud servers for cost saving and efficient application. However, the IIoT data are always of great value. The direct outsourcing of such data can leak the important information of the companies and cause great business losses. A straightforward solution is to encrypt the data by using standard encryption methods before outsourcing. Nevertheless, this will make data utilization quite inconvenient. This paper focuses on the secure process of image data on cloud servers. Images are stored on cloud servers in encrypted form, and the local binary pattern feature can be directly extracted from the encrypted images for applications. The security analysis and experimental results demonstrate the security and effectiveness of our scheme.
Zhihua Xia, Leqi Jiang, Xiaohe Ma, Puzhao Ji, Naixue Xiong
IEEE Trans. Ind. Informatics1
2020 A Novel Weber Local Binary Descriptor for Fingerprint Liveness Detection
abstract
In recent years, fingerprint authentication systems have been extensively deployed in various applications, including attendance systems, authentications on smartphones, mobile payment authorizations, as well as various safety certifications. However, similar to the other biometric identification technologies, fingerprint recognition is vulnerable to artificial replicas made from cheap materials, such as silicon, gelatin, etc. Thus, it is especially necessary to distinguish whether a given fingerprint is a live or a spoof one prior to such authentication. In order to solve the problems above, a novel local descriptor named Weber local binary descriptor for fingerprint liveness detection (FLD) has been proposed in this paper. The method consists of two components: the local binary differential excitation component that extracts intensity-variance features and the local binary gradient orientation component that extracts orientation features. The co-occurrence probability of the two components is calculated to construct a discriminative feature vector, which is fed into support vector machine (SVM) classifiers. The effectiveness of the proposed method is intuitively analyzed on the image samples and numerically demonstrated by Mahalanobis distance. Experiments are performed on two public databases from FLD competitions from 2011 and 2013. The results have proved that the proposed method obtains the best detection accuracy among the existing image local descriptors in FLD.
Zhihua Xia, Chengsheng Yuan 0001, Rui Lv, Xingming Sun, Naixue Xiong, Yun Q. Shi 0001
IEEE Trans. Syst. Man Cybern. Syst.1
2019 Secure multimedia distribution in cloud computing using re-encryption and fingerprinting
Lizhi Xiong, Zhihua Xia, Xianyi Chen, Hiuk Jae Shim
Multim. Tools Appl.2
2018 Rotation-invariant Weber pattern and Gabor feature for fingerprint liveness detection
Zhihua Xia, Rui Lv, Xingming Sun
Multim. Tools Appl.1
2018 A copy-move forgery detection method based on CMFD-SIFT
Bin Yang 0025, Xingming Sun, Zhihua Xia, Xianyi Chen
Multim. Tools Appl.4
2018 Improved Encrypted-Signals-Based Reversible Data Hiding Using Code Division Multiplexing and Value Expansion
abstract
Compared to the encrypted-image-based reversible data hiding (EIRDH) method, the encrypted-signals-based reversible data hiding (ESRDH) technique is a novel way to achieve a greater embedding rate and better quality of the decrypted signals. Motivated by ESRDH using signal energy transfer, we propose an improved ESRDH method using code division multiplexing and value expansion. At the beginning, each pixel of the original image is divided into several parts containing a little signal and multiple equal signals. Next, all signals are encrypted by Paillier encryption. And then a large number of secret bits are embedded into the encrypted signals using code division multiplexing and value expansion. Since the sum of elements in any spreading sequence is equal to 0, lossless quality of directly decrypted signals can be achieved using code division multiplexing on the encrypted equal signals. Although the visual quality is reduced, high-capacity data hiding can be accomplished by conducting value expansion on the encrypted little signal. The experimental results show that our method is better than other methods in terms of the embedding rate and average PSNR.
Xianyi Chen, Haidong Zhong, Lizhi Xiong, Zhihua Xia
Secur. Commun. Networks4
2018 Towards Privacy-Preserving Content-Based Image Retrieval in Cloud Computing
abstract
Content-based image retrieval (CBIR) applications have been rapidly developed along with the increase in the quantity, availability and importance of images in our daily life. However, the wide deployment of CBIR scheme has been limited by its the severe computation and storage requirement. In this paper, we propose a privacy-preserving content-based image retrieval scheme, which allows the data owner to outsource the image database and CBIR service to the cloud, without revealing the actual content of the database to the cloud server. Local features are utilized to represent the images, and earth mover's distance (EMD) is employed to evaluate the similarity of images. The EMD computation is essentially a linear programming (LP) problem. The proposed scheme transforms the EMD problem in such a way that the cloud server can solve it without learning the sensitive information. In addition, local sensitive hash (LSH) is utilized to improve the search efficiency. The security analysis and experiments show the security and efficiency of the proposed scheme.
Zhihua Xia, Yi Zhu 0012, Xingming Sun, Zhan Qin, Kui Ren 0001
IEEE Trans. Cloud Comput.1
2017 EPCBIR: An efficient and privacy-preserving content-based image retrieval scheme in cloud computing
Zhihua Xia, Naixue Xiong, Athanasios V. Vasilakos, Xingming Sun
Inf. Sci.1
2017 Enhancing Security of FPGA-Based Embedded Systems with Combinational Logic Binding
Jiliang Zhang 0002, Xingwei Wang 0001, Zhihua Xia
J. Comput. Sci. Technol.4
2017 An Approach of Web Service Organization Using Bayesian Network Learning
Jianxiao Liu, Zhihua Xia
J. Web Eng.2
2016 Techniques for Design and Implementation of an FPGA-Specific Physical Unclonable Function
Jiliang Zhang 0002, Qiang Wu 0015, Yipeng Ding, Yongqiang Lyu 0001, Qiang Zhou 0001, Zhihua Xia, Xingming Sun, Xingwei Wang 0001
J. Comput. Sci. Technol.6
2016 Steganalysis of LSB matching using differences between nonadjacent pixels
Zhihua Xia, Xingming Sun, Quansheng Liu, Naixue Xiong
Multim. Tools Appl.1
2016 A Privacy-Preserving and Copy-Deterrence Content-Based Image Retrieval Scheme in Cloud Computing
abstract
With the increasing importance of images in people's daily life, content-based image retrieval (CBIR) has been widely studied. Compared with text documents, images consume much more storage space. Hence, its maintenance is considered to be a typical example for cloud storage outsourcing. For privacy-preserving purposes, sensitive images, such as medical and personal images, need to be encrypted before outsourcing, which makes the CBIR technologies in plaintext domain to be unusable. In this paper, we propose a scheme that supports CBIR over encrypted images without leaking the sensitive information to the cloud server. First, feature vectors are extracted to represent the corresponding images. After that, the pre-filter tables are constructed by locality-sensitive hashing to increase search efficiency. Moreover, the feature vectors are protected by the secure kNN algorithm, and image pixels are encrypted by a standard stream cipher. In addition, considering the case that the authorized query users may illegally copy and distribute the retrieved images to someone unauthorized, we propose a watermark-based protocol to deter such illegal distributions. In our watermark-based protocol, a unique watermark is directly embedded into the encrypted images by the cloud server before images are sent to the query user. Hence, when image copy is found, the unlawful query user who distributed the image can be traced by the watermark extraction. The security analysis and the experiments show the security and efficiency of the proposed scheme.
Zhihua Xia, Liangao Zhang, Zhan Qin, Xingming Sun, Kui Ren 0001
IEEE Trans. Inf. Forensics Secur.1
2016 A Secure and Dynamic Multi-Keyword Ranked Search Scheme over Encrypted Cloud Data
abstract
Due to the increasing popularity of cloud computing, more and more data owners are motivated to outsource their data to cloud servers for great convenience and reduced cost in data management. However, sensitive data should be encrypted before outsourcing for privacy requirements, which obsoletes data utilization like keyword-based document retrieval. In this paper, we present a secure multi-keyword ranked search scheme over encrypted cloud data, which simultaneously supports dynamic update operations like deletion and insertion of documents. Specifically, the vector space model and the widely-used TF x IDF model are combined in the index construction and query generation. We construct a special tree-based index structure and propose a “Greedy Depth-first Search” algorithm to provide efficient multi-keyword ranked search. The secure kNN algorithm is utilized to encrypt the index and query vectors, and meanwhile ensure accurate relevance score calculation between encrypted index and query vectors. In order to resist statistical attacks, phantom terms are added to the index vector for blinding search results. Due to the use of our special tree-based index structure, the proposed scheme can achieve sub-linear search time and deal with the deletion and insertion of documents flexibly. Extensive experiments are conducted to demonstrate the efficiency of the proposed scheme.
Zhihua Xia, Xingming Sun, Qian Wang 0002
IEEE Trans. Parallel Distributed Syst.1
2014 Steganalysis of least significant bit matching using multi-order differences
abstract
ABSTRACT This paper presents a learning‐based steganalysis/detection method to attack spatial domain least significant bit (LSB) matching steganography in grayscale images, which is the antetype of many sophisticated steganographic methods. We model the message embedded by LSB matching as the independent noise to the image, and theoretically prove that LSB matching smoothes the histogram of multi‐order differences. Because of the dependency among neighboring pixels, histogram of low order differences can be approximated by Laplace distribution. The smoothness caused by LSB matching is especially apparent at the peak of the histogram. Consequently, the low order differences of image pixels are calculated. The co‐occurrence matrix is utilized to model the differences with the small absolute value in order to extract features. Finally, support vector machine classifiers are trained with the features so as to identify a test image either an original or a stego image. The proposed method is evaluated by LSB matching and its improved version “Hugo”. In addition, the proposed method is compared with state‐of‐the‐art steganalytic methods. The experimental results demonstrate the reliability of the new detector. Copyright © 2013 John Wiley & Sons, Ltd.
Zhihua Xia, Xingming Sun, Baowei Wang
Secur. Commun. Networks1
2013 Multi-keyword ranked search supporting synonym query over encrypted data in cloud computing
abstract
Cloud computing becomes increasingly popular. To protect data privacy, sensitive data should be encrypted by the data owner before outsourcing, which makes the traditional and efficient plaintext keyword search technique useless. The existing searchable encryption schemes support only exact or fuzzy keyword search, not support semantics-based multi-keyword ranked search. In the real search scenario, it is quite common that cloud customers' searching input might be the synonyms of the predefined keywords, not the exact or fuzzy matching keywords due to the possible synonym substitution (reproduction of information content) and/or her lack of exact knowledge about the data. Therefore, synonym-based multi-keyword ranked search over encrypted cloud data remains a very challenging problem. In this paper, for the first time, we propose an effective approach to solve the problem of synonym-based multi-keyword ranked search over encrypted cloud data. We make contributions mainly in two aspects: synonym-based search for supporting synonym query and multi-keyword ranked search for achieving more accurate search result. Two secure schemes are proposed to meet privacy requirements in two threat models of known ciphertext model and known background model. In enhanced scheme, the sensitive frequency information can be well protected by introducing some dummy keywords, which is not adopted in basic scheme. We give security analysis to justify the correctness and privacy-preserving guarantee of the proposed schemes. Extensive experiments on real-world dataset validate our analysis and show that our proposed solution is very efficient and effective in supporting synonym-based searching.
Zhangjie Fu 0001, Xingming Sun, Zhihua Xia, Jiangang Shu
IPCCC3