Jiwu Huang

dblp:90/6446 · also Ji-Wu Huang · DBLP profile ↗
← Back
307ranked-venue papers
4as first author
112since 2021 · last 2026
0000-0002-7625-5689ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 154 · 4 first-author · 46 since 2021Security and privacy · 118 · 50 since 2021Artificial intelligence and machine learning · 25 · 16 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 3 since 2021Computer networks · 9 · 2 since 2021Systems, architecture and hardware · 2 · 1 since 2021
YearPublicationVenuePosition
2026 Active Adversarial Noise Suppression for Image Forgery Localization
abstract
Recent advances in deep learning have significantly propelled the development of image forgery localization. However, existing models remain highly vulnerable to adversarial attacks: imperceptible noise added to forged images can severely mislead these models. In this paper, we address this challenge with an Adversarial Noise Suppression Module (ANSM) that generates a defensive perturbation to suppress the attack effect of adversarial noise. We observe that forgery-relevant features extracted from adversarial and original forged images exhibit distinct distributions. To bridge this gap, we introduce Forgery-relevant Features Alignment (FFA) as a first-stage training strategy, which reduces distributional discrepancies by minimizing the channel-wise Kullback-Leibler divergence between these features. To further refine the defensive perturbation, we design a second-stage training strategy, termed Mask-guided Refinement (MgR), which incorporates a dual-mask constraint. MgR ensures that the defensive perturbation remains effective for both adversarial and original forged images, recovering forgery localization accuracy to their original level. Extensive experiments across various attack algorithms demonstrate that our method significantly restores the forgery localization model's performance on adversarial images. Notably, when ANSM is applied to original forged images, the performance remains nearly unaffected. To our best knowledge, this is the first report of adversarial defense in image forgery localization tasks.
Rongxuan Peng, Shunquan Tan, Xianbo Mo, Alex Chichung Kot, Jiwu Huang
IEEE Trans. Pattern Anal. Mach. Intell.5
2026 Enhancing JPEG Steganography With GANs via Adaptive Modification Loss and Random Masking
abstract
Learning adaptive embedding costs with deep learning has become an important direction for improving steganographic security. However, most existing approaches focus on spatial-domain images, while effective cost learning for JPEG images remains challenging due to the complexity introduced by block-wise discrete cosine transform and lossy quantization. In this paper, we present an enhanced GAN-based framework that learns asymmetric embedding costs for JPEG images from scratch, improving both content adaptivity and training stability. Specifically, we introduce an adaptive modification loss that directly links the generator’s predicted embedding probabilities to the actual modification behavior, enabling fully data-driven and content-aware optimization without relying on handcrafted filters or quantization-dependent heuristics. In addition, we propose a random masking strategy applied during later training stages to prevent discriminator dominance and maintain informative adversarial feedback. Extensive experiments on multiple benchmark datasets and under various steganalytic detectors demonstrate that the proposed method consistently improves steganographic security over existing JPEG-domain approaches. Ablation studies further validate the effectiveness of the proposed loss design and training strategy.
Tianrui Gu, Bohong Li, Weiqi Luo 0001, Peijia Zheng, Shunquan Tan, Jiwu Huang
IEEE Trans. Circuits Syst. Video Technol.6
2026 Generalized Document Tampering Localization via Color and Semantic Disentanglement
abstract
Document images are vulnerable to tampering attacks from image editing tools and deep models. Therefore, the Document Tampering Localization (DTL) task has received increasing attention in recent years. However, given the wide variety of document types (e.g., contracts, certificates, ID cards), our analysis shows that existing DTL methods struggle with document images containing diverse background colors and varying semantic contents. Further analysis and experiments verify that the varying background color and semantic contents interfere with the forensic feature extraction process in the existing DTL methods. To address this issue, we propose two disentanglement modules to mitigate such interference and improve the ability of forgery trace detection. First, we design a Color Disentanglement (CD) module that applies disentangled learning representation to forensic features. The CD module, grounded in real-world prior knowledge, effectively decouples color information from forensic features, thereby improving robustness against varying background colors. Second, we propose the Semantic Disentanglement (SD) module, which performs image-level clustering on the tampering probability map during the inference process. The SD module focuses on tampering probabilities for each pixel, while discarding local semantic information (e.g., font, location, and shape). It leads to strong robustness against variations in document content. The evaluations demonstrate that our CD-SD method outperforms existing methods by 45.12% or 0.162 on the F1 metric in cross-dataset tests. Ablation studies show that the CD and SD modules improve the F1 score by 7.98% and 13.38%, respectively, across different backbones. Our method delivers consistent and stable improvements across various experimental protocols. Moreover, it is compatible with many DTL methods in a plug-and-play fashion.
Shiqiang Zheng 0002, Changsheng Chen 0001, Shen Chen 0004, Taiping Yao, Shouhong Ding, Bin Li 0011, Jiwu Huang
IEEE Trans. Circuits Syst. Video Technol.7
2026 Secure Moving Object Detection in Compressed Video Using Attentions
abstract
Moving Object Detection (MOD) can be outsourced to the cloud for computational convenience, in which case the video must be encrypted to protect privacy. Secure video MOD methods designed to perform MOD on encrypted video are still in their infancy. In this paper, we present an attention-based framework for privacy-preserving MOD in compressed videos. On the user side, we adopt selective video encryption for the compressed video, while in the cloud, we extract the Compressed Video entropy-coded Syntax Elements (CVSE) from the encrypted video. Since the extracted CVSE data lacks sufficient motion information and contains noise, we introduce a two-stage training process. In the first phase, we propose a new deep learning-based approach for interpolating CVSE-based motion feature maps, addressing a significant drawback of traditional methods that rely exclusively on empirical interpolation algorithms. In the second stage, we propose a specialized backbone tailored for feature extraction from sparse CVSE data. We then design an attention-based neck that focuses on areas with denser motion and varying sizes of moving objects. Experimental results on two public datasets, VIRAT and DUKE-MTMC, show that our framework achieves state-of-the-art detection performance. Compared to previous secure solutions, the proposed method exhibits more robustness in challenging scenes.
Peijia Zheng, Yuru Song, Xianhao Tian, Wei Lu 0001, Xiaochun Cao, Jiwu Huang
IEEE Trans. Dependable Secur. Comput.6
2026 SHL-Net: Semantics-Enhanced Network for Localizing Harmonized Image Splicing
Xiwen Fu, Guopu Zhu, Hongli Zhang 0001, Jiwu Huang, Tao Xiang 0001, Yicong Zhou, Ligang Wu 0001
IEEE Trans. Inf. Forensics Secur.4
2026 VoIP Call Identification via a Dual-Level 1D-CNN With Frame and Utterance Features
abstract
The increasing use of Voice over Internet Protocol (VoIP) technology in telecom fraud has become a serious global concern. Its ability to spoof caller IDs and IP addresses, and the use of overseas or anonymized servers make VoIP-based scams difficult to trace and regulate. As a result, distinguishing VoIP calls from conventional mobile phone calls based on voice signal characteristics is crucial for enhancing anti-fraud measures. However, existing forensic techniques often struggle to accurately identify speech transmitted via VoIP. To address this challenge, we propose a dual-level 1D-CNN that leverages both frame and utterance features for effective VoIP detection. After evaluating a range of acoustic features, we primarily focus on short-frame Mel-Frequency Cepstral Coefficients (MFCCs) due to their effectiveness in capturing VoIP characteristics. Given the frame-based processing and transmission nature of VoIP, we employ a 1D-CNN, rather than the more commonly used 2D-CNN that treats spectrograms as image, to extract frame-level codec features. Finally, we propose a dual-level classification strategy: the frame-level classifier captures encoding discrepancies within individual frames, while the utterance-level classifier aggregates these frame-level features to learn global encoding patterns through global covariance pooling. Experimental results on the VoIP Phone Call Identification Database (VPCID) demonstrate that the proposed method consistently outperforms existing approaches, delivering superior accuracy and robustness across a wide range of challenging scenarios. Moreover, comprehensive ablation studies validate the effectiveness and rationale behind the design of the proposed model architecture.
Guoyuan Lin, Weiqi Luo 0001, Peijia Zheng, Jiwu Huang
IEEE Trans. Inf. Forensics Secur.4
2026 Query-Efficient Hard-Label Attacks Against Black-Box Image Forgery Localization Model via Reinforcement Learning
abstract
Deep learning-based image forgery localization models are increasingly deployed in real-world forensic services, yet their robustness against black-box adversarial manipulation remains insufficiently understood, calling for practical anti-forensics techniques to expose potential security weaknesses. Prior adversarial anti-forensics studies for forgery localization mainly assume white-box access, which limits their applicability to deployed systems where only hard, mask-like outputs are available and queries are tightly constrained. To bridge this gap, we propose AdvFor, a query-efficient black-box attack frame-work tailored for forgery localization with hard-label, mask-only, spatially dense binary feedback. AdvFor formulates the attacker–model interaction as a finite-horizon Markov Decision Process and learns a transferable attack policy from hard-mask feedback. Once trained, AdvFor can be deployed via fixed-length policy execution with onlyT=7 queries per image, avoiding per-image boundary refinement or query-dense direction/gradient estimation. The learned policy optimizes a structured objective—progressively suppressing forgery responses in the predicted localization mask so that the masks of forgery images approach an authentic-like (near-zero) output—while maintaining visual fidelity. Extensive experiments on six benchmark datasets and multiple modern forgery localization models demonstrate that AdvFor consistently achieves stronger attack performance than representative baselines under the same perturbation constraints, while operating in an ultra-low-query regime.We further validate AdvFor under common deployment-style defenses, showing its notable effectiveness in realistic settings.
Xianbo Mo, Shunquan Tan, Rongxuan Peng, Bin Li 0011, Jiwu Huang
IEEE Trans. Inf. Forensics Secur.5
2026 DiffEraser: Generalized Text Erasure Based on Latent Diffusion Prior
abstract
Text removal is an important task in processing both scene and document images. However, existing scene text removal (STR) methods are primarily focus on scene text images. The STR models (trained by scene text images) perform poorly on document images with dense, complex textured backgrounds. We discover that the limitations of existing methods can be attributed to the difficuties in background features estimation in the regions to be erased, which is based on the knowledge from neighboring regions in the input images and priors learned from the training data. The background features estimation performance degrades under the cross-domain scenarios, and compromises the quality of STR results. To address these issues, we introduce DiffEraser, a novel text removal framework that leverages prior knowledge from the Latent Diffusion Model (LDM) for removing text in both scene and document images. Our DiffEraser incorporates two key innovations to fully exploit the prior knowledge of LDM. First, we replace the conventional Variational Auto-Encoders (VAE) encoder with a Diffusion-Prior (DP) encoder, aiming to integrate the heterogeneous information from the LDM prior knowledge in latent space with the multi-level encoded features of the input image. Second, we introduce a Latent-Fusion (LF) decoder that integrates the heterogeneous features from both the LDM and DP encoders to generate high-quality text-erased results. To evaluate the generalization performance of our DiffEraser, we focus on the cross-domain protocols and construct a document image dataset, NPID295, which contains 295 types of passports and identity cards. Notably, when trained on a scene text dataset, DiffEraser significantly outperforms existing STR methods in the challenging NPID295 dataset. The resources of this work will be available online upon acceptance.
Zhihao Chen 0011, Changsheng Chen 0001, Shunquan Tan, Jiwu Huang
IEEE Trans. Image Process.5
2025 Query-efficient Attack for Black-box Image Inpainting Forensics via Reinforcement Learning
abstract
Recently, image inpainting has become a common tool for manipulating nature images in a malicious manner, which has led to the rapid advancement of inpainting forensics. Although current forensics methods have shown precise location of inpainting regions and reliable robustness against image post-processing operations, it remains unclear whether they can effectively resist the possible attacks in real-world scenarios. To identify potential flaws, we propose a novel black-box anti-forensics framework to attack inpainting forensics methods, which employs reinforcement learning to generate a query-efficient countermeasure, named RLGC. To this end, we define reinforcement learning paradigm to model the Markov Decision Process of query-based black-box anti-forensics scenario. Specifically, pixel-wise agents are used to modulate anti-forensics images based on action selection and query forensics methods to obtain corresponding outputs. Later, reward function evaluates attack effect and image distortion with these outputs. To maximize the cumulative reward, policy and value networks are integrated and trained by Asynchronous Advantage Actor-Critic algorithm. Experimental results demonstrate that, without visually detectable distortion on anti-forensics images, RLGC achieves remarkable attack effects in a highly query-effcient way against various black-box inpainting forensics methods, even outperforming the most representative white-box attack method.
Xianbo Mo, Shunquan Tan, Bin Li 0011, Jiwu Huang
AAAI4
2025 Unmask Tampering: Efficient Document Tampering Localization under Recapturing Attacks with Real Distortion Knowledge
Changsheng Chen 0001, Yinyin Lin, Bin Li 0011, Jiwu Huang
CCS5
2025 Exploiting Robust Model Watermarking Against the Model Fine-Tuning Attack via Flat Minima Aware Optimizers
abstract
With the rapid advancement of deep neural networks (DNNs), model watermarking has emerged as a widely adopted technique for safeguarding model copyrights. A prevalent method involves utilizing a watermark decoder to retrieve watermark bits from generated outputs, but such methods are often vulnerable to model fine-tuning attacks. Traditionally, this challenge is mitigated through adversarial training or data augmentation, both of which significantly increase the computational burden. In this paper, we present a solution employing Flat Minima Aware (FMA) optimizers to bolster the robustness of model watermarking without requiring additional training data. By optimizing the watermark loss with flat minima awareness, our approaches significantly enhance the robustness of watermarks against the model fine-tuning attack. Comprehensive experiments have demonstrated our method’s superior ability to preserve watermark integrity. These findings suggest that this innovative optimization strategy offers a robust and efficient pathway for protecting models, thereby contributing to more secure and reliable model copyright protection mechanisms.
Dongdong Lin, Yue Li 0041, Bin Li 0011, Jiwu Huang
ICASSP4
2025 A GAN Framework for Asymmetric Embedding Costs Learning in JPEG Steganography
abstract
A key challenge in current steganography research is automatically learning image embedding costs without relying on existing costs. To date, there has been limited work on JPEG steganography, and the reported methods primarily rely on learning symmetric embedding, which fails to fully exploit the relationships between different modification directions within an embedding unit. This limitation restricts their security, leaving room for improvements in JPEG steganography techniques. To address this issue, we propose a GAN-based framework for JPEG steganography that learns asymmetric embedding costs from scratch. Our approach extends a modern framework of spatial steganography by incorporating a key IDCT module, which facilitates the conversion between JPEG and spatial domains. This enables the integration of effective spatial steganographic and steganalytic techniques into our framework for JPEG steganography. Additionally, we use a dual-branch UNet to generate separate embedding probabilities for +1 and -1 DCT coefficients and introduce a specialized loss function to guide DCT modifications. This loss function is designed by converting the modified DCT coefficients back to the spatial domain and analyzing various spatial residuals. Extensive experiments demonstrate that our method significantly outperforms existing JPEG steganography techniques, achieving state-of-the-art security performance. Furthermore, many ablation experiments validate the rationale of our model.
Bohong Li, Weiqi Luo 0001, Peijia Zheng, Shunquan Tan, Jiwu Huang
ICME5
2025 Toward Real-world Text Image Forgery Localization: Structured and Interpretable Data Synthesis
abstract
Existing Text Image Forgery Localization (T-IFL) methods often suffer from poor generalization due to the limited scale of real-world datasets and the distribution gap caused by synthetic data that fails to capture the complexity of real-world tampering. To tackle this issue, we propose Fourier Series-based Tampering Synthesis (FSTS), a structured and interpretable framework for synthesizing tampered text images. FSTS first collects 16,750 real-world tampering instances from five representative tampering types, using a structured pipeline that records human-performed editing traces via multi-format logs (e.g., video, PSD, and editing logs). By analyzing these collected parameters and identifying recurring behavioral patterns at both individual and population levels, we formulate a hierarchical modeling framework. Specifically, each individual tampering parameter is represented as a compact combination of basis operation–parameter configurations, while the population-level distribution is constructed by aggregating these behaviors. Since this formulation draws inspiration from the Fourier series, it enables an interpretable approximation using basis functions and their learned weights. By sampling from this modeled distribution, FSTS synthesizes diverse and realistic training data that better reflect real-world forgery traces. Extensive experiments across four evaluation protocols demonstrate that models trained with FSTS data achieve significantly improved generalization on real-world datasets. Dataset is available at \href{https://github.com/ZeqinYu/FSTS}{Project Page}.
Zeqin Yu, Haotao Xie, Jiangqun Ni, Wenkang Su 0001, Jiwu Huang
NeurIPS6
2025 Identification of Generative Forged Western Blot Images
abstract
With the rapid advancement of AI generative image technologies, the quality of fabricated images has significantly improved, posing a serious challenge to research integrity. Western blot (WB) images, frequently used in scientific publications, have become a primary target for forgery, leading to serious academic misconduct. However, detecting AI-generated WB images remains particularly challenging due to the low texture complexity and structural simplicity of WB images. To address this issue, we propose a novel detection framework for generative WB forgeries by leveraging large-scale pretrained models with targeted fine-tuning. This approach retains the generalization capabilities of the backbone model while adapting it to the unique characteristics of WB images. Moreover, we introduce an adversarial feature alignment module with a direction-aware domain classifier designed according to the banded and structural nature of WB images, which enhances robustness under limited data and across unseen generative styles. Extensive experiments demonstrate that our method consistently achieves superior performance compared with existing approaches across multiple generative models, low-resource target domains, and cross-domain settings, showing higher detection accuracy, stronger generalization, and robustness to common post-processing operations. These results highlight the practical applicability of the proposed framework in real-world scientific integrity monitoring.
Yunqiao Zhang, Shunquan Tan, Jiwu Huang
TrustCom3
2025 Towards generalizable and robust image tampering localization with multi-task learning and contrastive learning
Haodong Li 0001, Peiyu Zhuang, Yang Su 0005, Jiwu Huang
Expert Syst. Appl.4
2025 Inter-frame residual frequency-based reconstruction learning for deep video frame interpolation detection
Yibin Xu, Huaquan Yang, Shan Bian, Chuntao Wang, Bin Li 0011, Jiwu Huang
Expert Syst. Appl.6
2025 CTNet: A Convolutional Transformer Network for Color Image Steganalysis
Kangkang Wei, Weiqi Luo 0001, Shunquan Tan, Jiwu Huang
J. Comput. Sci. Technol.4
2025 Universal forged image detection and localization via self-supervised data generation and large-scale model adaptation
Yang Su 0005, Shunquan Tan, Yunqiao Zhang, Jiwu Huang
Multim. Syst.4
2025 Towards JPEG-Resistant Image Forgery Detection and Localization Via Self-Supervised Domain Adaptation
abstract
With wide applications of image editing tools, forged images (splicing, copy-move, removal and etc.) have been becoming great public concerns. Although existing image forgery localization methods could achieve fairly good results on several public datasets, most of them perform poorly when the forged images are JPEG compressed as they are usually done in social networks. To tackle this issue, in this paper, a self-supervised domain adaptation network, which is composed of a backbone network with Siamese architecture and a compression approximation network (ComNet), is proposed for JPEG-resistant image forgery detection and localization. To improve the performance against JPEG compression, ComNet is customized to approximate the JPEG compression operation through self-supervised learning, generating JPEG-agent images with general JPEG compression characteristics. The backbone network is then trained with domain adaptation strategy to localize the tampering boundary and region, and alleviate the domain shift between uncompressed and JPEG-agent images. Extensive experimental results on several public datasets show that the proposed method outperforms or rivals to other state-of-the-art methods in image forgery detection and localization, especially for JPEG compression with unknown QFs.
Yuan Rao 0002, Jiangqun Ni, Weizhe Zhang, Jiwu Huang
IEEE Trans. Pattern Anal. Mach. Intell.4
2025 Few-shot based learning recaptured image detection with multi-scale feature fusion and attention
Israr Hussain, Shunquan Tan, Jiwu Huang
Pattern Recognit.3
2025 An audio watermarking method against re-recording distortions
Guoyuan Lin, Weiqi Luo 0001, Peijia Zheng, Jiwu Huang
Pattern Recognit.4
2025 Elastic Supernet with Dynamic Training for JPEG steganalysis
Qiushi Li 0001, Shunquan Tan, Bin Li 0011, Jiwu Huang
Signal Process.4
2025 Forgery-Aware Adaptive Learning With Vision Transformer for Generalized Face Forgery Detection
abstract
With the rapid progress of generative models, the current challenge in face forgery detection is how to effectively detect realistic manipulated faces from different unseen domains. Though previous studies show that pre-trained Vision Transformer (ViT) based models can achieve some promising results after fully fine-tuning on the Deepfake dataset, their generalization performances are still unsatisfactory. To this end, we present a Forgery-aware Adaptive Vision Transformer (FA-ViT) under the adaptive learning paradigm for generalized face forgery detection, where the parameters in the pre-trained ViT are kept fixed while the designed adaptive modules are optimized to capture forgery features. Specifically, a global adaptive module is designed to model long-range interactions among input tokens, which takes advantage of self-attention mechanism to mine global forgery clues. To further explore essential local forgery clues, a local adaptive module is proposed to expose local inconsistencies by enhancing the local contextual association. In addition, we introduce a fine-grained adaptive learning module that emphasizes the common compact representation of genuine faces through relationship learning in fine-grained pairs, driving these proposed adaptive modules to be aware of fine-grained forgery-aware information. Extensive experiments demonstrate that our FA-ViT achieves state-of-the-arts results in the cross-dataset evaluation, and enhances the robustness against unseen perturbations. Particularly, FA-ViT achieves 93.83% and 78.32% AUC scores on Celeb-DF and DFDC datasets in the cross-dataset evaluation. The code and trained model have been released at:https://github.com/LoveSiameseCat/FAViT.
Anwei Luo, Rizhao Cai, Chenqi Kong, Yakun Ju, Xiangui Kang, Jiwu Huang, Alex Chichung Kot
IEEE Trans. Circuits Syst. Video Technol.6
2025 Moiré Spectral Augmentation and Masked Frequency Modeling for Document Presentation Attack Detection
abstract
Document Presentation Attack is an anti-forensic operation that conceals the forgery traces of image manipulation in the digital domain. Existing document presentation attack detection (DPAD) methods show unsatisfactory performance under samples with different contents and qualities. In this work, we focus on the DPAD task on screen-recapturing channel and exploit the prior knowledge of distortion (i.e., moire pattern) in the spectral domain to address these limitations. We propose a frequency-domain moir ´ e´ augmentation (FMAG) strategy that enhances the spectral components contributed to the moire distortion, improving the generalization ´ performance under different document contents. We devise the mask moire frequency modeling (M ´ 2FM) scheme to reconstruct the moire-related spectral components in low-quality samples under the guidance of the spectral distortion model and a pre-trained DPAD ´ classifier. To evaluate the generalization performance, we collect the diverse Screen Recaptured Document Image Dataset with 162 different document contents (SRDID162) consisting of 162 genuine document images, as well as 2592 low and high-quality recaptured document images, respectively. Our experimental protocol involves training with high-quality ID images and testing with SRDID162 dataset of diverse contents and image qualities. Compared to a SOTA data augmentation approach for recaptured natural images, our FMAG & M2FM approach achieves a significant improvement of 49.15% or 22.50 percentage points in average EER on the generic deep learning backbones. The data and code of this work will be available at Github
Changsheng Chen 0001, Youjie Li, Bokang Li, Weifan Yu, Baoying Chen, Bin Li 0011, Jiwu Huang
IEEE Trans. Dependable Secur. Comput.7
2025 DiRLoc: Disentanglement Representation Learning for Robust Image Forgery Localization
abstract
Deep Learning image forgery localization methods have achieved remarkable results but cannot maintain comparable performance when the forgery images are JPEG compressed, a format that is widely used in daily information transmission. The robustness against JPEG compression has become a bottleneck to the practical application of image forgery localization. To address this issue, a robust image forgery localization framework is proposed against the performance degradation caused by JPEG compression. Specifically, a cutting-edge progressive disentanglement strategy is proposed that incorporates coarse-grained image disentanglement to mitigate the detrimental effects of general JPEG compression, while harnessing the ability of fine-grained element disentanglement to separate multi-scale artifacts, thereby minimizing interference from content information. Moreover, the decision strategy is carefully designed to reinforce subtle signals from tampered areas, including artifacts fusion block reasoning multi-scale artifacts and dual attention block that learn more about forgery-related features. Extensive visualizations and experiments demonstrate that our method can achieve competitive performance in general JPEG-resistant image forgery localization, especially in the performance of generalization experiments.
Ziqi Sheng, Zuomin Qu, Wei Lu 0001, Xiaochun Cao, Jiwu Huang
IEEE Trans. Dependable Secur. Comput.5
2025 Prompt Engineering-Assisted Malware Dynamic Analysis Using GPT-4
abstract
Malware detection remains a critical challenge due to the increasing use of code obfuscation, packing, and wrapping techniques, which hinder traditional static analysis methods. Dynamic analysis, particularly through the examination of Application Programming Interface (API) call sequences, has emerged as an effective approach for identifying malicious behaviors. However, existing deep learning models often struggle to generate high-quality representations of API calls and are unable to handle previously unseen APIs, thereby limiting detection performance and model generalization. To address these challenges, we propose a novel malware dynamic analysis framework that leveragesGPT-4prompt engineering to generate descriptive text for each API call within a sequence. These descriptions are then encoded using a pre-trained BERT model to produce rich, knowledge-enhanced representations of API sequences. Our method not only incorporates external knowledge for improved semantic understanding but also enables the representation of unknown API calls, thus enhancing generalization. We further design a CNN-based classifier to extract features from the enriched representations for malware detection and classification. Extensive experiments on five benchmark datasets demonstrate that our approach outperforms state-of-the-art methods, achieving superior detection accuracy and generalization across different datasets. Especially, the detection accuracy on the Catak dataset increased by 12.38%, which highlights the significant improvement of our method in challenging scenarios. The code is available.
Pei Yan, Shunquan Tan, Miaohui Wang, Jiwu Huang
IEEE Trans. Dependable Secur. Comput.4
2025 Accurate and Efficient Privacy-Preserving Feature Extraction on Encrypted Images
abstract
In cloud computing, it is necessary to outsource image processing algorithms securely without exposing private image content. The scale-invariant feature transform (SIFT) is a famous local descriptor widely used in computer vision. There are already some privacy-preserving schemes for computing SIFT on encrypted images. However, the state-of-the-art works have to convert fixed-point numbers into their binary representations, which reduces efficiency and accuracy. In this paper, we propose a novel privacy-preserving SIFT scheme built from secure protocols designed explicitly for fixed-point numbers to solve this problem. Specifically, using RLWE-based homomorphic encryption, we propose word-wise protocols to perform secure division, square root operation, comparison, derivation, and matrix inversion in a single-instruction multiple-data manner. These protocols allow direct processing of fixed-point numbers without converting them to binary numbers, thus achieving high computational efficiency. We have also realized critical SIFT steps missing from previous works, including Euclidean gradient amplitude computation, histogram peak interpolation, and precise interval localization, leading to improved accuracy of SIFT features in the encrypted domain. We conduct security analysis and perform extensive experiments to evaluate the execution efficiency and accuracy. The experimental results show that the proposed scheme outperforms the state-of-the-art works in terms of computational efficiency and accuracy.
Peijia Zheng, Xiongjie Fang, Rui Yang 0006, Wei Lu 0001, Xiaochun Cao, Jiwu Huang
IEEE Trans. Dependable Secur. Comput.7
2025 Evading Detection Actively: Toward Anti-Forensics Against Forgery Localization
abstract
Anti-forensics seeks to eliminate or conceal traces of tampering artifacts. Typically, anti-forensic methods are designed to deceive binary detectors and persuade them to misjudge the authenticity of an image. However, to the best of our knowledge, no attempts have been made to deceive forgery detectors at the pixel level and mis-locate forged regions. Traditional adversarial attack methods cannot be directly used against forgery localization due to the following defects: 1) they tend to just naively induce the target forensic models to flip their pixel-level pristine or forged decisions; 2) their anti-forensics performance tends to be severely degraded when faced with the unseen forensic models; 3) they lose validity once the target forensic models are retrained with the anti-forensics images generated by them. To tackle the three defects, we propose SEAR (Self-supErvised Anti-foRensics), a novel self-supervised and adversarial training algorithm that effectively trains deep-learning anti-forensic models against forgery localization. SEAR sets a pretext task to reconstruct perturbation for self-supervised learning. In adversarial training, SEAR employs a forgery localization model as a supervisor to explore tampering features and constructs a deep-learning concealer to erase corresponding traces. We have conducted large-scale experiments across diverse datasets. The experimental results demonstrate that, through the combination of self-supervised learning and adversarial learning, SEAR successfully deceives the state-of-the-art forgery localization methods, as well as tackle the three defects regarding traditional adversarial attack methods mentioned above.
Long Zhuo, Shenghai Luo, Shunquan Tan, Bin Li 0011, Jiwu Huang
IEEE Trans. Dependable Secur. Comput.6
2025 A Forensic Framework With Diverse Data Generation for Generalizable Forgery Localization
abstract
Deep learning-based forensic techniques have emerged as the leading approach for image forgery localization. However, many existing methods struggle with overfitting to the training data, which limits their generalization performance and real-world applicability. To overcome this challenge, we propose a novel forensic framework that incorporates an advanced data augmentation technique. The framework consists of two key components: a generator and a detector. The generator challenges the detector’s learned distribution under constraints of diversity and consistency, ensuring that the generated data diverges from the source domain while maintaining statistical differences related to tampering. The detector, in turn, captures tampering traces from three critical aspects of the tampered image: long-range dependency information, RGB-noise fusion information, and boundary artifacts, resulting in a more comprehensive detection process. By alternating the optimization of the generator and detector, the framework fosters mutual reinforcement, promoting diverse data generation and expanding the distributional coverage, ultimately improving performance. Extensive experiments demonstrate that the proposed method significantly surpasses state-of-the-art approaches in both generalization and robustness, with numerous ablation studies further validating the soundness of the model design.
Yuanhang Huang, Weiqi Luo 0001, Xiaochun Cao, Jiwu Huang
IEEE Trans. Inf. Forensics Secur.4
2025 FGMIA: Feature-Guided Model Inversion Attacks Against Face Recognition Models
abstract
Model Inversion Attacks (MIAs) against face recognition systems aim to reconstruct facial images of specific individuals from the recognition models. Existing MIA approaches commonly optimize the latent variables of Generative Adversarial Networks (GANs) iteratively, which can result in non-smooth optimizations due to the complexity and entanglement of latent space. Furthermore, the optimization guided by the target model’s gradients may generate high-confidence images with poor perceptual similarity to the target class. This paper introduces a novel perspective by reformulating the inversion attack as a conditional data distribution learning task. Based on this, we propose a Feature-Guided Model Inversion Attack (FGMIA), which learns the facial data distribution and integrates feature guidance as a conditional signal. Specifically, we treat the deconstructed target model as a feature encoder, which provides guidance during the training of a specialized feature-guided diffusion model. During the attack, feature encodings implicit in the target model are extracted and utilized to guide the reconstruction of private data. Extensive experiments demonstrate that FGMIA accurately reconstructs private data from face recognition models and significantly improves evaluation accuracy and perceptual similarity compared to state-of-the-art methods while maintaining comparable target confidence scores. Our code is available at https://github.com/MMCTTT/FGMIA_codes.
Shen Wang 0004, Guopu Zhu, Zhaoyang Zhang 0002, Jiwu Huang
IEEE Trans. Inf. Forensics Secur.5
2025 Adv-Inversion: Stealthy Adversarial Attacks via GAN-Inversion for Facial Privacy Protection
Weiqi Luo 0001, Xiaohua Xie, Peijia Zheng, Wenmin Huang, Jiwu Huang
IEEE Trans. Inf. Forensics Secur.6
2025 StealthPhase: Toward a Stealthy Backdoor Attack Against Speaker Recognition
abstract
Speaker recognition systems (SRS) play a vital role in identity authentication. At the same time, researchers have found that these systems are highly vulnerable to backdoor attacks, where the poisoned model will misclassify poisoned inputs. Most backdoor attack methods primarily focus on improving attack success rates (ASR), achieving ASR as high as 99%. However, these methods reveal a significant concern in terms of stealthiness. Poisoned audio often exhibits detectable differences from the clean audio, which can be detected by human listeners or through visualization. To overcome this issue, we prioritize stealthiness in our attack design and propose StealthPhase. Motivated by preliminary experiments on frequency-domain random noise backdoor attacks, our method implants a predefined trigger into the phase spectrum through frequency decomposition to ensure inherent stealth. The predefined trigger uses the natural phase pattern derived from real speech. Therefore, it is both learnable, as it addresses the challenge of designing effective phase-based triggers, and stealthy, as it remains imperceptible in both spectrogram visualizations and auditory perception. A key advantage of our method is that it avoids complex algorithms to optimize triggers and does not require an extra loss function to balance stealthiness and effectiveness. Extensive experimental results demonstrate that StealthPhase achieves 99% ASR with minimal impact on the model’s benign accuracy (BA). Meanwhile, its stealthiness is validated from three perspectives. First, visualizations show that the backdoor audio samples are nearly indistinguishable from clean samples. Second, an audio quality assessment confirms that the trigger introduces minimal perceptual distortion, preserving the overall audio quality. Finally, speech recognition performance evaluation shows that the word error rate (WER) remains largely unaffected. Furthermore, we validate the effectiveness of StealthPhase in real-world scenarios, where it achieves an ASR of 80%, and demonstrate its ability to bypass defense mechanisms.
Zhe Ye 0001, Qiben Yan 0001, Xiangui Kang, Jiwu Huang
IEEE Trans. Inf. Forensics Secur.5
2025 Privacy-Preserving CNN Inference for Image Super-Resolution Cross Multiple Ciphertexts
abstract
Online image super-resolution (SR) services have been widely used in applications such as Remini and DeepAI. However, the exposure of plaintext images raises serious privacy concerns. While secure CNN inference techniques are employed to protect images in image classification, they are not applicable to the unique challenges posed by image SR: the output resolution is significantly higher than that of the input image. In this paper, we present a secure CNN inference scheme for image SR by employing a multiple ciphertext encapsulation method. We begin by designing fundamental homomorphic operations, including addition, multiplication, and rotation across ciphertexts. Recognizing that image SR typically involves an upsampling layer-unlike image classification-we propose a fast algorithm for secure upsampling. This technique leverages pre-weight block masking and cross-ciphertext rotation, resulting in a significant speedup compared to direct homomorphic upsampling. We then present an efficient batched homomorphic two-dimensional convolution method across ciphertexts, incorporating kernel rearrangement and merging strategies. We also design a polynomial activation function specifically optimized for image SR, further enhancing performance. Extensive experiments demonstrate that our HE-friendly SR network outperforms existing secure solutions, while the proposed multiple ciphertext encapsulation technique achieves at least a 2x improvement in both computational efficiency and memory usage.
Peijia Zheng, Donger Mo, Xiaochun Cao, Jiwu Huang
IEEE Trans. Image Process.6
2025 Generating Higher-Quality Anti-Forensics DeepFakes with Adversarial Sharpening Mask
abstract
DeepFake, an AI technology that can automatically synthesize facial forgeries, has recently attracted worldwide attention. While DeepFakes can be entertaining, they can also be used to spread falsified information or be weaponized as cognition warfare. Forensic researchers have been dedicated to designing defensive algorithms to combat such disinformation. However, attacking technologies have been developed to make DeepFake products more aggressive. For example, by launching anti-forensics and adversarial attacks, DeepFakes can be disguised as authentic media to evade forensic detectors. However, such manipulations often sacrifice image quality for satisfactory undetectability. To address this issue, we propose a method to generate a novel adversarial sharpening mask for launching black-box anti-forensics attacks. Unlike many existing methods, our approach injects perturbations that allow DeepFakes to achieve high anti-forensics performance while maintaining pleasant sharpening visual effects. Experimental evaluations demonstrate that our method successfully disrupts state-of-the-art DeepFake detectors. Moreover, compared to images processed by existing DeepFake anti-forensics methods, our method’s quality of anti-forensics DeepFakes rendered is significantly improved. Our code is available at https://github.com/fb-reps/HQ-AF_GAN .
Bing Fan, Feng Ding 0007, Guopu Zhu, Jiwu Huang, Sam Kwong, Pradeep K. Atrey, Siwei Lyu
ACM Trans. Multim. Comput. Commun. Appl.4
2025 Domain-invariant and Patch-discriminative Feature Learning for General Deepfake Detection
abstract
Hyper-realistic avatars in the metaverse have already raised security concerns about deepfake techniques; deepfakes involving generated video “recording” may be mistaken for a real recording of the people it depicts. As a result, deepfake detection has drawn considerable attention in the multimedia forensic community. Though existing methods for deepfake detection achieve fairly good performance under the intra-dataset scenario, many of them gain unsatisfying results in the case of cross-dataset testing with more practical value, where the forged faces in training and testing datasets are from different domains. To tackle this issue, in this article, we propose a novel Domain-Invariant and Patch-Discriminative feature learning framework—DI&PD. For image-level feature learning, a single-side adversarial domain generalization is introduced to eliminate domain variances and learn domain-invariant features in training samples from different manipulation methods, along with the global and local random crop augmentation strategy to generate more data views of forged images at various scales. A graph structure is then built by splitting the learned image-level feature maps, with each spatial location corresponding to a local patch, which facilitates patch representation learning by message-passing among similar nodes. Two types of center losses are utilized to learn more discriminative features in both image-level and patch-level embedding spaces. Extensive experimental results on several datasets demonstrate the effectiveness and generalization of the proposed method compared with other state-of-the-art methods.
Jian Zhang 0086, Jiangqun Ni, Fan Nie, Jiwu Huang
ACM Trans. Multim. Comput. Commun. Appl.4
2024 SDGAN: Disentangling Semantic Manipulation for Facial Attribute Editing
abstract
Facial attribute editing has garnered significant attention, yet prevailing methods struggle with achieving precise attribute manipulation while preserving irrelevant details and controlling attribute styles. This challenge primarily arises from the strong correlations between different attributes and the interplay between attributes and identity. In this paper, we propose Semantic Disentangled GAN (SDGAN), a novel method addressing this challenge. SDGAN introduces two key concepts: a semantic disentanglement generator that assigns facial representations to distinct attribute-specific editing modules, enabling the decoupling of the facial attribute editing process, and a semantic mask alignment strategy that confines attribute editing to appropriate regions, thereby avoiding undesired modifications. Leveraging these concepts, SDGAN demonstrates accurate attribute editing and achieves high-quality attribute style manipulation through both latent-guided and reference-guided manners. We extensively evaluate our method on the CelebA-HQ database, providing both qualitative and quantitative analyses. Our results establish that SDGAN significantly outperforms state-of-the-art techniques, showcasing the effectiveness of our approach. To foster reproducibility and further research, we will provide the code for our method.
Wenmin Huang, Weiqi Luo 0001, Jiwu Huang, Xiaochun Cao
AAAI3
2024 Two-Tier Data Packing in RLWE-based Homomorphic Encryption for Secure Federated Learning
abstract
Homomorphic Encryption (HE) facilitates the preservation of privacy in federated learning (FL) aggregation. However, HE imposes significant computational and communication overhead. To address this problem, data encoding methods have been introduced that enable batch processing to improving the efficiency of ciphertext usage. The existing methods simply concatenate integer or coefficients assignment in polynomials, which do not fully make use of HE based on ring learning with errors (RLWE). We present a novel two-tier data encoding approach tailored for RLWE-based HE, effectively utilizing RLWE's polynomial structure. Our method involves a dual-level data packing strategy for batch processing at both integer and polynomial levels. At the first tier (integer level), we amalgamate those quantized model data into larger integers. Beyond existing concatenation-based encoding, we introduce a new encoding method derived from the Chinese Remainder Theorem (CRT). This CRT-based method effectively mitigates overflow and error propagation concerns. At the second tier (polynomial level), we transmute the large integers into a polynomial form. Additionally, we propose a new subring decomposition method, i.e., employing ring isomorphism mappings to project multiple large integers into varied sub-polynomial rings. Our dual-tier encoding strategy offers a more flexible and effective batch HE solution. We rigorously analyze the correctness, efficiency, and security of our approach. Our extensive experimental evaluations reveal that secure FL, empowered by our dual-tier encoding technique, markedly enhances computational and communication efficiencies over prevailing batch HE methods.
Peijia Zheng, Xiaochun Cao, Jiwu Huang
CCS4
2024 CMA: A Chromaticity Map Adapter for Robust Detection of Screen-Recapture Document Images
abstract
The rebroadcasting of screen-recaptured document images introduces a significant risk to the confidential docu-ments processed in government departments and commer-cial companies. However, detecting recaptured document images subjected to distortions from online social networks (OSNs) is challenging since the common forensics cues, such as moiré pattern, are weakened during transmission. In this work, we first devise a pixel-level distortion model of the screen-recaptured document image to identify the robust features of color artifacts. Then, we extract a chromaticity map from the recaptured image to highlight the presence of color artifacts even under low-quality samples. Based on the prior understanding, we design a chromaticity map adapter (CMA) to efficiently extract the chromaticity map, and feed it into the transformer backbone as multi-modal prompt tokens. To evaluate the performance of the pro-posed method, we collect a recaptured office document im-age dataset with over 10K diverse samples. Experimental results demonstrate that the proposed CMA method outper-forms a SOTA approach (with RGB modality only), reducing the average EER from 26.82% to 16.78%. Robustness eval-uation shows that our method achieves 0.8688 and 0.7554 AUCs under samples with JPEG compression$(QF=70)$and resolution as low as$534\times 503$pixels.
Changsheng Chen 0001, Liangwei Lin, Bin Li 0011, Jishen Zeng, Jiwu Huang
CVPR6
2024 A Keyless Extraction Framework Targeting at Deep Learning Based Image-Within-Image Models
abstract
Image-within-image technique aims to establish covert communication by concealing a secret image within a cover image. Compared with traditional steganography algorithms, the security of image-within-image technique has not been rigorously evaluated by steganalysis. Existing attack methods just brutally destroy the container image, resulting in the secret image cannot be revealed by the original decryption model (key). This paper introduces a novel keyless extraction framework, carrying out steganalysis on the container image without destroying it. Our approach utilizes collected pairs of container and revealed images to construct a master key, enabling us to extract secret image from container image without relying on the original key. Remarkably, the master key remains effective for multiple image-within-image techniques simultaneously, even when their encryption and decryption models are re-trained. In addition, we propose a patch-based data augmentation technique to adapt to scenarios with limited training samples, and we design a weighted loss function with three components to further enhance the visual quality of the extracted secret image. All the experiments are conducted on datasets derived from ImageNet, COCO and DIV2k. The results demonstrate that our approach can extract secret images with comparable visual quality to the original ones.
Rongxuan Peng, Xianbo Mo, Shunquan Tan, Bin Li 0011, Jiwu Huang
ICASSP5
2024 Velocity Field-Based Surveillance Video Frame Deletion Detection Using Siamese Network
Yang Su 0005, Shunquan Tan, Jiwu Huang
ICPR (22)3
2024 A Novel Universal Image Forensics Localization Model Based on Image Noise and Segment Anything Model
abstract
Maliciously manipulated images have inflicted severe negative impacts on people's lives, making the development of forgery localization techniques imperative. The goal of forgery localization is to segment regions containing tampering traces. In this paper, we treat tampered areas as distinct targets within the noise feature map, utilizing noise features to filter objects in the source image and achieve forgery localization. Our approach involves designing a transformer-based forgery feature extractor, which seamlessly integrates two key features during the extraction process: target features obtained from the image segmentation model Segment Anything Model, and noise features extracted by the steganalysis rich model filters. This extractor can effectively extract forgery features in the image and a mask was applied to localize the tampered regions. Extensive experimentation across multiple datasets underscores the robust competitiveness of our model and its strong transfer learning capabilities.
Yang Su 0005, Shunquan Tan, Jiwu Huang
IH&MMSec3
2024 GAN-based Symmetric Embedding Costs Adjustment for Enhancing Image Steganographic Security
abstract
Designing embedding costs is pivotal in modern image steganography. Many studies have shown adjusting symmetric embedding costs to asymmetric ones can enhance steganographic security. However, most existing methods heavily depend on manually defined parameters or rules, limiting security performance improvements. To overcome this limitation, we introduce an advanced GAN-based framework that transitions symmetric costs to asymmetric ones without the need for the manual intervention seen in existing approaches, such as the detailed specification of cost modulation directions and magnitudes. In our framework, we firstly achieve symmetric costs for a cover image, which is randomly split into two sub-images, with part of the secret information embedded into one. Subsequently, we design a GAN model to adjust the embedding costs of the second sub-image to asymmetric, facilitating the secure embedding of the remaining secret information. To support our phased embedding approach, our GAN's discriminator incorporates two steganalyers with different tasks: distinguishing the generator's final output, i.e., the stego image, from both the input cover image and the partially embedded stego image, providing diverse guidance to the generator. In addition, we introduce a simple yet effective update strategy to ensure a stable training process. Comprehensive experiments demonstrate that our method significantly enhances security over existing symmetric steganography techniques, achieving state-of-the-art levels compared to other methods focused on embedding costs adjustments. Additionally, detailed ablation studies validate our approach's effectiveness.
Miaoxin Ye, Saixing Zhou, Weiqi Luo 0001, Shunquan Tan, Jiwu Huang
ACM Multimedia5
2024 A semi-supervised deep learning approach for cropped image detection
Israr Hussain, Shunquan Tan, Jiwu Huang
Expert Syst. Appl.3
2024 Fine-Grained Multimodal DeepFake Classification via Heterogeneous Graphs
Qilin Yin, Wei Lu 0001, Xiaochun Cao, Xiangyang Luo 0001, Yicong Zhou, Jiwu Huang
Int. J. Comput. Vis.6
2024 An efficient distortion cost function design for image steganography in spatial domain using quaternion representation
Qingliang Liu 0001, Wenkang Su 0001, Jiangqun Ni, Xianglei Hu, Jiwu Huang
Signal Process.5
2024 Efficient JPEG image steganography using pairwise conditional random field model
Yuanfeng Pan, Jiangqun Ni, Qingliang Liu 0001, Wenkang Su 0001, Jiwu Huang
Signal Process.5
2024 A knowledge distillation based deep learning framework for cropped images detection in spatial domain
Israr Hussain, Shunquan Tan, Jiwu Huang
Signal Process. Image Commun.3
2024 Enhanced Dynamic Analysis for Malware Detection With Gradient Attack
abstract
Malware detection is an effective way to prevent the intrusion of malware into computer systems, and the API-based dynamic analysis method can effectively detect obfuscated and packaged malware. However, existing methods still suffer from limited detection accuracy and weak generalization. To address this issue, this paper presents a gradient attack-based malware dynamic analysis method. Through exerting adversarial noise into the embedding layer, the malware detection model can learn more robust representations of API sequences during training, achieving broader coverage of sample representations. The strategy of normalizing attack noise and recovering attacked representation is designed, which controls the strength of the gradient attack within a reasonable range and prevents a negative impact on the model's detection performance. The proposed method can be applied to existing API-based malware detection models to enhance their detection performance, indicating the strong generality of the proposed method. Experimental results on two benchmark datasets (i.e.,AliyunandCatak) demonstrate the effectiveness of the proposed gradient attack method, which further improves the detection performance of the mainstream API-based models, with an average accuracy increase of 2.80% and 3.66% on these two datasets, respectively.
Pei Yan, Shunquan Tan, Miaohui Wang, Jiwu Huang
IEEE Signal Process. Lett.4
2024 FairCMS: Cloud Media Sharing With Fair Copyright Protection
abstract
The onerous media sharing task prompts resource-constrained media owners to seek help from a cloud platform, i.e., storing media contents in the cloud and letting the cloud do the sharing. There are three key security/privacy problems that need to be solved in the cloud media sharing scenario, including data privacy leakage and access control in the cloud, infringement on the owner’s copyright, and infringement on the user’s rights. In view of the fact that no single technique can solve the above three problems simultaneously, two cloud media sharing schemes are proposed in this article, named FairCMS-I and FairCMS-II. By cleverly utilizing the proxy re-encryption technique and the asymmetric fingerprinting (AFP) technique, FairCMS-I and FairCMS-II solve the above three problems with different privacy/efficiency tradeoffs. Among them, FairCMS-I focuses more on cloud-side efficiency while FairCMS-II focuses more on the security of the media content, which provides owners with flexibility of choice. In addition, FairCMS-I and FairCMS-II also have advantages over existing cloud media sharing efforts in terms of optional indistinguishability under chosen-plaintext attack (IND-CPA) security and high cloud-side efficiency, as well as exemption from needing a trusted third party. Furthermore, FairCMS-I and FairCMS-II allow owners to reap significant local resource savings and thus can be seen as the privacy-preserving outsourcing of AFP. Finally, the feasibility and efficiency of FairCMS-I and FairCMS-II are demonstrated by experiments.
Xiangli Xiao, Yushu Zhang 0001, Leo Yu Zhang, Zhongyun Hua, Zhe Liu 0001, Jiwu Huang
IEEE Trans. Comput. Soc. Syst.6
2024 Interactive Generative Adversarial Networks With High-Frequency Compensation for Facial Attribute Editing
abstract
Recently, facial attribute editing has drawn increasing attention and has achieved significant progress due to Generative Adversarial Network (GAN). Since paired images before and after editing are not available, existing methods typically perform the editing and reconstruction tasks simultaneously, and transfer facial details learned from the reconstruction to the editing via sharing the latent representation space and weights. In this way, they can not preserve those non-targeted regions well during editing. In addition, they usually introduce skip connections between the encoder and decoder to improve image quality at the cost of attribute editing ability. In this paper, we propose a novel method called InterGAN with high-frequency compensation to alleviate above problems. Specifically, we first propose the cross-task interaction (CTI) to fully explore the relationships between editing and reconstruction tasks. The CTI includes two translations: style translation adjusts the mean and variance of feature maps according to style features, and conditional translation utilizes attribute vector as condition to guide feature map transformation. They provide effective information interaction to preserve the irrelevant regions unchanged. Without using skip connections between the encoder and decoder, furthermore, we propose the high-frequency compensation module (HFCM) to improve image quality. The HFCM tries to collect potentially loss information from input images and each down-sampling layers of the encoder, and then re-inject them into subsequent layers to alleviate the information loss. Ablation analysis show the effectiveness of proposed CTI and HFCM. Extensive qualitative and quantitative experiments on CelebA-HQ demonstrate that the proposed method outperforms state-of-the-art methods both in attribute editing accuracy and image quality.
Wenmin Huang, Weiqi Luo 0001, Xiaochun Cao, Jiwu Huang
IEEE Trans. Circuits Syst. Video Technol.4
2024 Secure Deep Learning Framework for Moving Object Detection in Compressed Video
abstract
In the cloud, there is an urgent need to implement intelligent video surveillance in a privacy-preserving way. Moving object detection is an important task in the intelligent surveillance system. In this paper, we propose a privacy-preserving deep learning framework to detect moving objects on compressed videos. We encrypt video bitstreams using selective video encryption to protect the private video content. We propose encrypted domain motion information (EDMI) without decryption and decompression to design three motion feature maps. Due to the sparsity of the EDMI distribution, existing convolutional backbones designed for RGB images have difficulty providing satisfactory performance. We design a novel convolutional backbone using a ”subtraction” strategy to reduce model complexity. Our backbone employs residual blocks and skipping connections to reuse the EDMI at deeper layers. We evaluate our model on two large high-definition surveillance video datasets, i.e., VIRAT and Duke-MTMC. The experimental results show that the proposed framework achieves state-of-the-art detection performance compared with the most recent works. Our approach achieves an excellent privacy-utility tradeoff. Compared to previous solutions, it performs more robustly in crowded scenarios with challenges like occlusion. To our best knowledge, this is the first reported deep learning framework for moving object detection in encrypted-compressed video.
Xianhao Tian, Peijia Zheng, Jiwu Huang
IEEE Trans. Dependable Secur. Comput.3
2024 Distortion Model-Based Spectral Augmentation for Generalized Recaptured Document Detection
abstract
Document recapturing is a presentation attack that covers the forensic traces in the digital domain. Document presentation attack detection (DPAD) is an important step in the document authentication pipeline. Existing DPAD methods suffer from low generalization performance under the cross-domain scenario with different types of documents. Data augmentation is a de facto technique to reduce the risk of overfitting the training data and improve the generalizability of a trained model. In this work, we improve the generalization performance of DPAD approaches by addressing two important limitations of the existing frequency domain augmentation (FDA) methods. First, contrary to the existing FDA methods that treat different spectral bands equally, we establish a band-of-interest localization (BOIL) method that locates the spectral band-of-interest (BOI) related to the recapturing operation by domain knowledge from the theoretical distortion models. Second, we propose a frequency-domain halftoning augmentation (FHAG) strategy that enhances the halftoning features in the BOI with considerations of different halftoning distortions. To evaluate the generalization performance of our FHAG with BOIL method on different types of document images, we have constructed a diverse recaptured document image dataset with 162 types of documents (RDID162), consisting of 5346 samples. The proposed method has been evaluated on the generic deep learning models and a state-of-the-art DPAD approach under both cross-device and cross-domain protocols for the DPAD task. Compared to the existing FDA methods, our method has improved the models with ResNet50 backbone by reducing more than 25% or 5 percentage points in EERs. The source code and data in this work is available athttps://github.com/chenlewis/FHAG-with-BOIL.
Changsheng Chen 0001, Bokang Li, Rizhao Cai, Jishen Zeng, Jiwu Huang
IEEE Trans. Inf. Forensics Secur.5
2024 Steganography Embedding Cost Learning With Generative Multi-Adversarial Network
abstract
Since the generative adversarial network (GAN) was proposed by Ian Goodfellow et al. in 2014, it has been widely used in various fields. However, there are only a few works related to image steganography so far. Existing GAN-based steganographic methods mainly focus on the design of generator, and just assign a relatively poorer steganalyzer in discriminator, which inevitably limits the performances of their models. In this paper, we propose a novel Steganographic method based on Generative Multi-Adversarial Network (Steg-GMAN) to enhance steganography security. Specifically, we first employ multiple steganalyzers rather than a single steganalyzer like existing methods to enhance the performance of discriminator. Furthermore, in order to balance the capabilities of the generator and the discriminator during training stage, we propose an adaptive way to update the parameters of the proposed GAN model according to the discriminant ability of different steganalyzers. In each iteration, we just update the poorest one among all steganalyzers in discriminator, while update the generator with the gradients derived from the strongest one. In this way, the performance of generator and discriminator can be gradually improved, so as to avoid training failure caused by gradient vanishing. Extensive comparative results show that the proposed method can achieve state-of-the-art results compared with the traditional steganography and the modern GAN-based steganographic methods. In addition, a large number of ablation experiments verify the rationality of the proposed model.
Dongxia Huang, Weiqi Luo 0001, Minglin Liu, Weixuan Tang 0004, Jiwu Huang
IEEE Trans. Inf. Forensics Secur.5
2024 One-Class Neural Network With Directed Statistics Pooling for Spoofing Speech Detection
abstract
Existing deep learning models for spoofing speech detection often struggle to effectively generalize to unseen spoofing attacks that were not present during the training stage. Moreover, the presence of class imbalance further compounds this issue by biasing the learning process towards seen attack samples. To address these challenges, we present an innovative end-to-end model called One-Class Neural Network with Directed Statistics Pooling (OCNet-DSP). Our model incorporates a feature cropping operation to attenuate high-frequency components, mitigating the risk of overfitting. Additionally, leveraging the time-frequency characteristics of speech signals, we introduce a directed statistics pooling layer that extracts more effective features for distinguishing between bonafide and spoofing classes. We also propose the Threshold One-class Softmax loss, which mitigates class imbalance by reducing the optimization weight of spoofing samples during training. Extensive comparative results demonstrate that the proposed model outperforms all existing single models, achieving an equal error rate of 0.44% and a minimum detection cost function of 0.0145 for the ASVspoof 2019 logical access database. Moreover, the proposed ensemble version, which accommodates speech inputs of varying lengths in each submodel, maintains state-of-the-art performance among reproducible ensemble models. Additionally, numerous ablation experiments, along with a cross-dataset experiment, are conducted to validate the rationality and effectiveness of the proposed model.
Guoyuan Lin, Weiqi Luo 0001, Jiwu Huang
IEEE Trans. Inf. Forensics Secur.4
2024 Beyond the Prior Forgery Knowledge: Mining Critical Clues for General Face Forgery Detection
abstract
Face forgery detection is essential in combating malicious digital face attacks. Previous methods mainly rely on prior expert knowledge to capture specific forgery clues, such as noise patterns, blending boundaries, and frequency artifacts. However, these methods tend to get trapped in local optima, resulting in limited robustness and generalization capability. To address these issues, we propose a novel Critical Forgery Mining (CFM) framework, which can be flexibly assembled with various backbones to boost their generalization and robustness performance. Specifically, we first build a fine-grained triplet and suppress specific forgery traces through prior knowledge-agnostic data augmentation. Subsequently, we propose a fine-grained relation learning prototype to mine critical information in forgeries through instance and local similarity-aware losses. Moreover, we design a novel progressive learning controller to guide the model to focus on principal feature components, enabling it to learn critical forgery features in a coarse-to-fine manner. The proposed method achieves state-of-the-art forgery detection performance under various challenging evaluation settings. The source code is available at:https://github.com/LoveSiameseCat/CFM.
Anwei Luo, Chenqi Kong, Jiwu Huang, Yongjian Hu, Xiangui Kang, Alex Chichung Kot
IEEE Trans. Inf. Forensics Secur.3
2024 Employing Reinforcement Learning to Construct a Decision-Making Environment for Image Forgery Localization
abstract
The widespread misuse of advanced image editing tools and deep generative techniques has led to a proliferation of images with altered content in real-life scenarios, often without any discernible traces of tampering. This has created a potential threat to security and credibility of images. Image forgery localization is an urgent technique. In this paper, we propose a novel reinforcement learning-based framework CoDE (Construct Decision-making Environment) that can provide reliable localization result of tampered area in forged images. We model the forgery localization task as a Markov Decision Process (MDP), where each pixel is equipped with an agent that performs Gaussian distribution-based continuous action to iteratively update the respective forgery probability, so as to achieve pixel-level image forgery localization. In order to construct the state transitions within MDP, we propose a twin-flow state encoder to handle the updated state, which consists of the forged image and its corresponding forgery probability map. What’s more, considering that the tampered area is often sparse in practical image tampering scenarios, we design a reward function specifically for these sparse tampered area. This reward function can guide the agent to more effectively learn the optimal strategy for maximizing the cumulative reward. Extensive experiments conducted on a variety of benchmark datasets demonstrate CoDE’s superior localization accuracy and robustness against image degradation caused by transmission through Online Social Networks (OSNs) and various post-processing attacks.
Rongxuan Peng, Shunquan Tan, Xianbo Mo, Bin Li 0011, Jiwu Huang
IEEE Trans. Inf. Forensics Secur.5
2024 Joint Cost Learning and Payload Allocation With Image-Wise Attention for Batch Steganography
abstract
In recent years, although cost learning methods have made great progress in single-image steganography, its development in batch steganography is relatively slower, which is a more practical communication scenario in the real world. The difficulties are capturing the full view of the image batch and building connections between cost learning and payload allocation by neural networks. To address the issues, this paper proposes a cost learning framework for batch steganography called JoCoP (Joint Cost Learning and Payload Allocation), wherein the policy network is designed to learn the optimal embedding policies for a batch of images via the collaboration between a cost learning module and a payload allocation module. In specific layers of the policy network, in the cost learning module, the intermediate feature maps of embedding costs are extracted for different images independently, which are sent to the payload allocation module. In the payload allocation module, to implement implicit payload allocation, the feature maps corresponding to different images within the same batch are adjusted by an image-wise attention mechanism. Afterwards, these adjusted feature maps are returned to the cost learning module for subsequent feature extraction in the next layer. Owing to the collaboration between the two modules and the batch-level receptive field in the image-wise attention mechanism, the embedding costs and the payload allocation can be jointly optimized in an end-to-end manner. Experimental results show that the proposed JoCoP outperforms existing methods against both single-image steganalyzers and pooled steganalyzers based on feature extraction and convolutional neural networks.
Weixuan Tang 0004, Zhili Zhou 0001, Bin Li 0011, Kim-Kwang Raymond Choo, Jiwu Huang
IEEE Trans. Inf. Forensics Secur.5
2024 Color Image Steganalysis Based on Pixel Difference Convolution and Enhanced Transformer With Selective Pooling
abstract
Current deep learning-based steganalyzers often depend on specific image dimensions, leading to inevitable adjustments in network structure when dealing with varied image sizes. This impedes their effectiveness in managing the wide range of image sizes commonly found on social media. To address this issue, our paper presents a novel steganalytic network that is optimized for fixed-size (notably,$256\times 256$) color images, but is capable of efficiently detecting stego images of arbitrary size without needing retraining or modifications to the network. Our proposed network is comprised of four modules. In the initial stem module, we calculate truncated residuals for each color channel of the input image. Diverging from existing steganalytic networks that rely on vanilla convolution, we have developed a pixel difference convolution module designed to better capture the artifacts introduced by steganography. Following this, we introduce an enhanced Transformer module with selective pooling, aimed at more effectively extracting global steganalytic features. To guarantee our network’s adaptability to different image sizes, we have developed a selective pooling strategy. This involves using global covariance pooling for fixed-size color images and spatial pyramid pooling for color images of various other sizes. This approach effectively standardizes the feature maps into uniform feature vectors. The final module is focused on classification. Extensive testing results on the ALASKA II color image dataset have demonstrated that our approach significantly improves detection performance for both fixed-size and arbitrary-size images, achieving state-of-the-art results. Additionally, we provide numerous ablation studies to confirm the effectiveness and soundness of our proposed network architecture.
Kangkang Wei, Weiqi Luo 0001, Jiwu Huang
IEEE Trans. Inf. Forensics Secur.3
2024 Audio Multi-View Spoofing Detection Framework Based on Audio-Text-Emotion Correlations
abstract
In recent years, audio spoofing detection has received widespread attention for protecting personal privacy and social security. Despite the significant progress achieved in audio single-view spoofing detection, challenges remain with regard to addressing unknown spoofing attacks in realistic scenarios. To solve these challenging problems, in this paper, we introduce a novel audio multi-view spoofing detection framework (AMSDF), whose goal is to capture both intra-view and inter-view cues by measuring correlations within audio multi-view features (i.e., audio-emotion-text) for audio spoofing detection. In general, different view features are inherently interconnected in the real patterns, while they may present unnatural correlations in the spoofing patterns. Therefore, more discriminative cues can be mined by utilizing their complex interactions, which is beneficial to the audio spoofing detection task. To this end, an intra-view graph attention mechanism (IGAM) is first utilized to aggregate each intra-view node within the same view. Subsequently, a heterogeneous graph fusion module (HGFM) is applied to measure correlations within inter-view nodes, which are enhanced with a master node for comprehensive analysis purposes. Finally, a group-based readout scheme (GRS) is designed to capture and preserve the most distinctive cues by leveraging the strengths of different feature sets, thereby effectively distinguishing subtle differences between real and spoofing audio. The experimental results show that our proposed framework can achieve better performance than that of the state-of-the-art methods, especially in realistic scenarios. The code and pre-trained models are available athttps://github.com/ItzJuny/AMSDF.
Junyan Wu, Qilin Yin, Ziqi Sheng, Wei Lu 0001, Jiwu Huang, Bin Li 0011
IEEE Trans. Inf. Forensics Secur.5
2024 Adversarial Perturbation Prediction for Real-Time Protection of Speech Privacy
abstract
The widespread collection and analysis of private speech signals have become increasingly prevalent, raising significant privacy concerns. To protect speech signals from unauthorized analysis, adversarial attack methods for deceiving speaker recognition models have been proposed. While a few of these methods are specifically designed for real-time protection of speech signals, they introduce significant delays that can severely impact speech communication when applied to streaming speech data. In this paper, we present a novel approach that aims to offer real-time protection for speech signals without delays. By utilizing observed data only, we generate initial adversarial seed perturbations and refine them to obtain the necessary adversarial perturbations predicted for adjacent unobserved signals. This refinement process is conducted via a proposed model called PAPG. On the basis of perturbation prediction, we develop a streaming audio processing framework that generates perturbations in synchronization with the playback of the original signal, effectively eliminating delays. The experimental results demonstrate that under the proposed attack, the average Top-1 accuracy of various advanced speaker recognition methods is reduced by 89%, and the average equal error rate (EER) increases to 36%. Remarkably, these results are achieved without delays while maintaining superior perceptual quality.
Zhaoyang Zhang 0002, Shen Wang 0004, Guopu Zhu, Dechen Zhan, Jiwu Huang
IEEE Trans. Inf. Forensics Secur.5
2024 Detection of Deepfake Videos Using Long-Distance Attention
abstract
With the rapid progress of deepfake techniques in recent years, facial video forgery can generate highly deceptive video content and bring severe security threats. And detection of such forgery videos is much more urgent and challenging. Most existing detection methods treat the problem as a vanilla binary classification problem. In this article, the problem is treated as a special fine-grained classification problem since the differences between fake and real faces are very subtle. It is observed that most existing face forgery methods left some common artifacts in the spatial domain and time domain, including generative defects in the spatial domain and interframe inconsistencies in the time domain. And a spatial-temporal model is proposed which has two components for capturing spatial and temporal forgery traces from a global perspective, respectively. The two components are designed using a novel long-distance attention mechanism. One component of the spatial domain is used to capture artifacts in a single frame, and the other component of the time domain is used to capture artifacts in consecutive frames. They generate attention maps in the form of patches. The attention method has a broader vision which contributes to better assembling global information and extracting local statistic information. Finally, the attention maps are used to guide the network to focus on pivotal parts of the face, just like other fine-grained classification methods. The experimental results on different public datasets demonstrate that the proposed method achieves state-of-the-art performance, and the proposed long-distance attention method can effectively capture pivotal parts for face forgery.
Wei Lu 0001, Lingyi Liu, Xianfeng Zhao, Yicong Zhou, Jiwu Huang
IEEE Trans. Neural Networks Learn. Syst.7
2023 Poster: Query-efficient Black-box Attack for Image Forgery Localization via Reinforcement Learning
abstract
Recently, deep learning has been widely used in forensics tools to detect and localize forgery images. However, its susceptibility to adversarial attacks highlights the need for the exploration of anti-forensics research. To achieve this, we introduce an innovative and query-efficient black-box anti-forensics framework tailored for the generation of adversarial forgery images. This framework is designed to simulate the query dynamics of online forensic services, utilizing a Markov Decision Process formulation within the paradigm of reinforcement learning. We further introduce a novel reward function, which evaluates the efficacy of attacks based on the disjunction between query results and attack targets. To improve the query efficiency of these attacks, an actor-critic algorithm is employed to maximize cumulative rewards. Empirical findings substantiate the efficacy of our proposed methodology. Specifically, it demonstrates pronounced adversarial effects on a range of prevailing image forgery detectors, while ensuring negligible visually perceptible distortions in the resultant anti-forensics images.
Xianbo Mo, Shunquan Tan, Bin Li 0011, Jiwu Huang
CCS4
2023 Multi-Scale Enhanced Dual-Stream Network for Facial Attribute Editing Localization
Jinkun Huang, Weiqi Luo 0001, Wenmin Huang, Ziyi Xi, Kangkang Wei, Jiwu Huang
IWDW6
2023 Automatic Asymmetric Embedding Cost Learning via Generative Adversarial Networks
abstract
In comparison to symmetric embedding, asymmetric methods generally provide better steganography security. However, the performance of existing asymmetric methods is limited by their reliance on symmetric embedding costs. In this paper, we present a novel Generative Adversarial Network (GAN)-based steganography approach that independently learns asymmetric embedding costs from scratch. Our proposed framework features a generator with a dual-branch architecture and a discriminator that integrates multiple steganalytic networks. To address the issues of model instability and non-convergence that often arise in GAN model training, we implement an adaptive strategy that updates the GAN model parameters according to the performance of multiple steganalytic networks in each iteration. Furthermore, we introduce a new adversarial loss function that effectively learns asymmetric embedding costs by utilizing features like image residuals, gradients, asymmetric embedding probability maps, and the sign of the modification map to train the dual-branch network within the generator. Our comprehensive experiments show that our method achieves state-of-the-art steganography security results, significantly outperforming existing top-performing symmetric and asymmetric methods. Additionally, numerous ablation experiments confirm the rationality of our GAN-based model design.
Dongxia Huang, Weiqi Luo 0001, Peijia Zheng, Jiwu Huang
ACM Multimedia4
2023 Reinforcement learning of non-additive joint steganographic embedding costs with attention mechanism
Weixuan Tang 0004, Bin Li 0011, Weixiang Li, Yuangen Wang, Jiwu Huang
Sci. China Inf. Sci.5
2023 Discriminative Frequency Information Learning for End-to-End Speech Anti-Spoofing
abstract
End-to-end technology is an active research topic in speech anti-spoofing. Although end-to-end methods have achieved remarkable success in the speech anti-spoofing, channel effects brought by telephony transmission and certain challenging forms of spoofing attacks still plague them. We observe that differences in the high-frequency components between bonafide and spoofed speech help detect some most troublesome attack forms and the differences also remain after the signals are affected by transmission and codecs. Based on this observation, we aim to utilize the high-frequency information of speech signals to develop better generalization ability to unknown attacks and stronger robustness against transmission and codecs. We propose a raw waveform processing module based on sinc convolution and multiple pre-emphasis to obtain discriminative shallow feature representations. Additionally, we propose an improved backbone to learn discriminative feature embeddings, and a feature classification loss to optimize intra-class and inter-class distances simultaneously. The above modules constitute the proposed Discriminative Frequency-information SincNet, namely DFSincNet. Our proposed algorithm demonstrates competitive performance on both ASVspoof 2019 and 2021 logical access (LA) scenarios.
Bingyuan Huang, Sanshuai Cui, Jiwu Huang, Xiangui Kang
IEEE Signal Process. Lett.3
2023 Non-Interactive Privacy-Preserving Frequent Itemset Mining Over Encrypted Cloud Data
abstract
Frequent itemset mining is a data mining technique widely used on massive datasets. In cloud computing, the dataset may be encrypted for privacy protection. Therefore, frequent itemset mining over encrypted data is a crucial application in secure cloud computing. In this paper, we propose an effective privacy-preserving framework where the cloud server can directly perform data mining on the encrypted database without interacting with other cloud servers. We first design three security primitives to implement subset determination, accumulation, and comparison in the encrypted domain for frequent itemset mining. Based on the proposed framework, we then propose two secure protocols that allow the cloud server to perform frequent itemset mining on encrypted cloud data with these security primitives. The first protocol leaks no information to the cloud and the second protocol has the advantage of more efficient mining performance. We then present two strategies with parallel algorithms and GPU computing to accelerate the running time. We also analyze the security of our protocols and the computational complexities. Experimental results show that our serial-based protocols achieve shorter running times and higher levels of privacy than previous solutions. Our multi-CPU (or GPU) based parallel protocol can further reduce the practical running time.
Peijia Zheng, Ziyan Cheng, Xianhao Tian, Hongmei Liu 0001, Weiqi Luo 0001, Jiwu Huang
IEEE Trans. Cloud Comput.6
2023 Anti-Rounding Image Steganography With Separable Fine-Tuned Network
abstract
Image steganographic methods based on encoder-decoder model with end-to-end network architecture recently have been proposed. However, in steganographic applications, the feature map (called stego matrix) generated by the encoder needs to be rounded as a real stego image for the receiver. The loss of precision by rounding stego matrix leads to the decline in the accuracy of extracted secret messages. The challenge of using end-to-end network to preserve robustness against rounding operation is that it is non-differentiable. In this paper, we propose an anti-rounding image steganography method with separable fine-tuning network architecture which includes the joint training stage (JT-stage) and the separable fine-tuning stage (SF-stage). Firstly, in JT-stage, an embedded generator and a stego matrix extractor are jointly learned without rounding operation. Utilizing concatenation in embedded generator can realistically fuse cover image and secret messages. And the multi-scale fusion block and residual dense block in stego matrix extractor can make secret messages more correctly decoded. Moreover, the discriminator is constructed by generative adversarial nets (GAN) in JT-stage to effectively improve the authenticity and steganalysis security. Then, in SF-stage, the embedded generator is frozen, and the stego matrix is obtained and rounded as a stego image. A stego image extractor is constructed by fine-tuning the layers of the stego matrix extractor to improve the accuracy of message extraction. As the loss will not backpropagate in the embedded generator, the non-differentiability of rounding operation can be offset. Experiments show that the proposed separation fine-tuning network is robust to rounding operation, and effectively reduces the degradation of the image quality and steganalysis performance.
Xiaolin Yin, Shaowu Wu, Wei Lu 0001, Yicong Zhou, Jiwu Huang
IEEE Trans. Circuits Syst. Video Technol.6
2023 Adversarial Steganography Embedding via Stego Generation and Selection
abstract
The recent literature has shown that adversarial embedding has promise for enhancing the security of steganography. However, existing methods achieve the final stego mainly based on a pre-trained Convolutional Neural Network (CNN)-based steganalyzer without considering any other steganalytic features. When the steganalyzer is re-trained, its performance usually drops significantly. We propose a novel adversarial embedding method via stego generation and selection. To improve the diversity of the stego images, this method first randomly generates many candidate stegos according to the amplitudes of the gradients and embedding costs of a given cover. Since the image residuals are the commonly used low-level features in many steganalyzers, the proposed method carefully designs different adaptive high-pass filters to calculate the image residuals, and then selects a final stego from among those candidate stegos which can successfully fool the pre-trained steganalyzer, according to the residual distance between stego and the cover. Extensive experimental evaluations on re-trained CNN-based and traditional steganalyzers demonstrate that the proposed method can significantly enhance the security of the modern steganographic methods in both spatial and JPEG domains, and achieve much better performance than related adversarial embedding methods.
Minglin Liu, Weiqi Luo 0001, Peijia Zheng, Jiwu Huang
IEEE Trans. Dependable Secur. Comput.5
2023 STD-NET: Search of Image Steganalytic Deep-Learning Architecture via Hierarchical Tensor Decomposition
abstract
Steganalysis aims to reveal covert communication established via steganography. In the arm race with steganography, steganalysis has evolved from the old-style hand-crafted features set to deep-learning architectures. However, recent studies show that the majority of existing deep steganalysis models have a large amount of redundancy, which leads to a huge waste of storage and computing resources. The existing model compression method cannot flexibly compress the convolutional layer in residual shortcut block so that a satisfactory shrinking rate cannot be obtained. In this paper, we propose STD-NET, an unsupervised deep-learning architecture search approach via hierarchical tensor decomposition for image steganalysis. Our proposed strategy will not be restricted by various residual connections, since this strategy does not change the number of input and output channels of the convolution block. We propose a normalized distortion threshold to evaluate the sensitivity of each involved convolutional layer of the base model to guide STD-NET to compress target network in an efficient and unsupervised approach, and obtain two network structures of different shapes with low computation cost and similar performance compared with the original one. Extensive experiments have confirmed that, on one hand, our model can achieve comparable or even better detection performance in various steganalytic scenarios due to the great adaptivity of the obtained network architecture. On the other hand, the experimental results also demonstrate that our proposed strategy is more efficient and can remove more redundancy compared with previous steganalytic network compression methods.
Shunquan Tan, Qiushi Li 0001, Laiyuan Li, Bin Li 0011, Jiwu Huang
IEEE Trans. Dependable Secur. Comput.5
2023 An Adaptive IPM-Based HEVC Video Steganography via Minimizing Non-Additive Distortion
abstract
Recently, almost all the proposed adaptive video steganographic schemes are based on minimizing an additive embedding distortion. However, they ignore the hard fact that the additive embedding distortion is not quite suitable for video steganography because of the interplay of cover elements in video steganography. In this article, an adaptive intra prediction mode based (IPM-based) video steganography is proposed by minimizing the non-additive distortion in HEVC. To reduce the complexity of minimizing the non-additive distortion, a multi-layered embedding structure combined with a proposed embedding distortion updating strategy is adopted to approximate the non-additive distortion in an additive form. First, all IPMs are decomposed into multiple layers based on the distortion drift graph to offer multi-layered embedding. Each IPM in the same layer is considered to be independent, and syndrome-trellis code (STC) can be applied to embed the message segment into each layer with an additive distortion function sequentially. Then, a distortion function composed of self-distortion and drift-distortion is proposed to initialize the distortion of modifying each IPM. Finally, after embedding the first message segment into the IPMs in the first layer with the initialized distortions, an embedding distortion updating strategy is applied to update the distortions of the IPMs in the remaining layers dynamically. Experimental results demonstrate that the proposed adaptive IPM-based video steganography can achieve much better perceptual quality and security performance than the state-of-the-art.
Jie Wang 0031, Xuemei Yin, Yifang Chen 0002, Jiwu Huang, Xiangui Kang
IEEE Trans. Dependable Secur. Comput.4
2023 Comprehensive Android Malware Detection Based on Federated Learning Architecture
abstract
Android malware and its variants are a major challenge for mobile platforms. However, there are two main problems in the existing detection methods:a) The detection method lacks the evolution ability for Android malware, which leads to the low detection rate of the detection model for malware and its variants.b) Traditional detection methods require centralized data for model training, however, the aggregation of training samples is limited due to the infectivity of malware and growing data privacy concerns, centralized detection methods are difficult to be applied in actual detection scenarios. In this paper, we propose FEDriod, a comprehensive Android malware detection method based on federated learning architecture that protects against growing Android malware or emerging Android malware variants. Specifically, we employ genetic evolution strategy to simulate the evolution of Android malware and develop potential malware variants from typical Android malware. Then, we customize the Android malware detection model based on residual neural network to achieve high detection accuracy. Finally, to achieve the protection sensitive data, we develope a federated learning framework to allows multiple Android malware detection agencies to jointly build a comprehensive Android malware detection model. We comprehensively evaluate the performance of FEDriod on the CIC, Drebin, and Contagio authoritative datasets. Experimental results show that our local model outperforms all baseline classifiers. In the federal scenario, our proposed method is superior to the state-of-the-art detection methods, especially in the cross-dataset evaluation, the F1 of FEDriod is 98.53%. More important, we performed genetic evolution experiments on the Drebin dataset, and the results showed that our proposed method has the ability to detect Android malware variants.
Wenbo Fang, Junjiang He, Wenshan Li 0001, Xiaolong Lan, Tao Li 0016, Jiwu Huang, Linlin Zhang 0005
IEEE Trans. Inf. Forensics Secur.7
2023 ReLOAD: Using Reinforcement Learning to Optimize Asymmetric Distortion for Additive Steganography
abstract
Recently, the success of non-additive steganography has demonstrated that asymmetric distortion can remarkably improve security performance compared with symmetric cost functions. However, most of current existing additive steganographic methods are still based on symmetric distortion. In this paper, for the first time we optimize asymmetric distortion for additive steganography and propose an A3C (Asynchronous Advantage Actor-Critic) based steganographic framework, called ReLOAD. ReLOAD is composed of an actor and a critic, where the former guides action selection for pixel-wise distortion modulation, and the latter evaluates the performance of modulated distortion. Meanwhile, a reward function that considers embedding effects is proposed to unify the goal of steganography and reinforcement learning, so that the minimization of embedding effects can be achieved by learning secure policy to maximize total rewards. Statistical analysis shows that compared with non-additive steganography, ReLOAD achieves lower change rates and makes embedding traces more consistent with cover image textures. Comprehensive experiments conducted on both hand-crafted feature-based and deep learning-based steganalyzers show that ReLOAD significantly promotes the state-of-the-art security performance of current additive methods and even outperforms non-additive steganography when the modification distribution gets sparser.
Xianbo Mo, Shunquan Tan, Weixuan Tang 0004, Bin Li 0011, Jiwu Huang
IEEE Trans. Inf. Forensics Secur.5
2023 Query-Efficient Adversarial Attack With Low Perturbation Against End-to-End Speech Recognition Systems
abstract
With the widespread use of automated speech recognition (ASR) systems in modern consumer devices, attack against ASR systems have become an attractive topic in recent years. Although related white-box attack methods have achieved remarkable success in fooling neural networks, they rely heavily on obtaining full access to the details of the target models. Due to the lack of prior knowledge of the victim model and the inefficiency in utilizing query results, most of the existing black-box attack methods for ASR systems are query-intensive. In this paper, we propose a new black-box attack called the Monte Carlo gradient sign attack (MGSA) to generate adversarial audio samples with substantially fewer queries. It updates an original sample based on the elements obtained by a Monte Carlo tree search. We attribute its high query efficiency to the effective utilization of the dominant gradient phenomenon, which refers to the fact that only a few elements of each origin sample have significant effect on the output of ASR systems. Extensive experiments are performed to evaluate the efficiency of MGSA and the stealthiness of the generated adversarial examples on the DeepSpeech system. The experimental results show that MGSA achieves 98% and 99% attack success rates on the LibriSpeech and Mozilla Common Voice datasets, respectively. Compared with the state-of-the-art methods, the average number of queries is reduced by 27% and the signal-to-noise ratio is increased by 31%.
Shen Wang 0004, Zhaoyang Zhang 0002, Guopu Zhu, Xinpeng Zhang 0001, Yicong Zhou, Jiwu Huang
IEEE Trans. Inf. Forensics Secur.6
2023 Dynamic Difference Learning With Spatio-Temporal Correlation for Deepfake Video Detection
abstract
With the rapid development of face forgery techniques, the existing frame-based deepfake video detection methods have fell into a dilemma that frame-based methods may fail when encountering extremely realistic images. To overcome the above problem, many approaches attempted to model the spatio-temporal inconsistency of videos to distinguish real and fake videos. However, current works model spatio-temporal inconsistency by combining intra-frame and inter-frame information, but ignore the disturbance caused by facial motions that would limit further improvement in detection performance. To address this issue, we investigate into long and short range inter-frame motions and propose a novel dynamic difference learning method to distinguish between the inter-frame differences caused by face manipulation and the inter-frame differences caused by facial motions in order to model precise spatio-temporal inconsistency for deepfake video detection. Moreover, we elaborately design a dynamic fine-grained difference capture module (DFDC-module) and a multi-scale spatio-temporal aggregation module (MSA-module) to collaboratively model spatio-temporal inconsistency. Specifically, the DFDC-module applies self-attention mechanism and fine-grained denoising operation to eliminate the differences caused by facial motions and generates long range difference attention maps. The MSA-module is devised to aggregate multi-direction and multi-scale temporal information to model spatio-temporal inconsistency. The existing 2D CNNs can be extended into dynamic spatio-temporal inconsistency capture networks by integrating the proposed two modules. Extensive experimental results demonstrate that our proposed algorithm steadily outperforms state-of-the-art methods by a clear margin in different benchmark datasets.
Qilin Yin, Wei Lu 0001, Bin Li 0011, Jiwu Huang
IEEE Trans. Inf. Forensics Secur.4
2023 ReLoc: A Restoration-Assisted Framework for Robust Image Tampering Localization
abstract
With the spread of tampered images, locating the tampered regions in digital images has drawn increasing attention. The existing tampering localization methods, however, suffer from severe performance degradation when the images are subjected to some post-processing, as the tampering traces would be distorted by the post-processing operations. The poor robustness against post-processing has become a bottleneck for the practical applications of image tampering localization techniques. In order to address this issue, this paper proposes a novelrestoration-assisted framework for image tamperinglocalization (ReLoc). The ReLoc framework mainly consists of an image restoration module and a tampering localization module. The key idea of ReLoc is to use the restoration module to recover a high-quality counterpart from the distorted tampered image, such that the distorted tampering traces can be re-enhanced, facilitating the tampering localization module to identify the tampered regions. To achieve this, the restoration module is optimized not only with the conventional constraints on image visual quality, but also with a forensics-oriented objective function. Furthermore, the restoration module and the localization module are trained alternately, which can stabilize the training process and is beneficial for improving the performance. The robustness of ReLoc has been evaluated by using several common post-processing operations, including lossy compressions, online social network transmission, and image resizing. Extensive experimental results show that ReLoc can significantly improve the localization performance compared to using a restoration-free model. In addition, we have shown that the restoration module in a well-trained ReLoc model is transferable for different localization modules and across different datasets.
Peiyu Zhuang, Haodong Li 0001, Rui Yang 0006, Jiwu Huang
IEEE Trans. Inf. Forensics Secur.4
2022 Deep Video Inpainting Localization Using Spatial and Temporal Traces
abstract
Advanced deep-learning-based video inpainting can fill a specified video region with visually plausible contents, usually leaving imperceptible traces. As inpainting can be used for malicious video manipulations, it has led to potential privacy and security issues. Therefore, it is necessary to detect and locate the video regions subjected to deep inpainting. This paper addresses this problem by exploiting the spatial and temporal traces left by inpainting. Firstly, the inpainting traces are enhanced by intra-frame and inter-frame residuals. In particular, we guide the extraction of inter-frame residual with optical-flow based frame alignment, which can better reveal the inpainting traces. Then, a dual-stream network, acting as the encoder, is designed to learn discriminative features from frame residuals. Finally, bidirectional convolutional LSTMs are embedded in the decoder network to produce pixel-wise predictions of inpainted regions for each frame. The proposed method is evaluated with tampered videos created by two state-of-the-art deep video inpainting algorithms. Extensive experimental results show that the proposed method can effectively localize the inpainted regions, outperforming existing methods.
Shujin Wei, Haodong Li 0001, Jiwu Huang
ICASSP3
2022 Keyword Spotting in the Homomorphic Encrypted Domain Using Deep Complex-Valued CNN
abstract
In this paper, we propose a non-interactive scheme to achieve end-to-end keyword spotting in the homomorphic encrypted domain using deep learning techniques. We carefully designed a complex-valued convolutional neural network (CNN) structure for the encrypted domain keyword spotting to take full advantage of the limited multiplicative depth. At the same depth, the proposed complex-valued CNN can learn more speech representations than the real-valued CNN, thus achieving higher accuracy in keyword spotting. The complex activation function of the complex-valued CNN is non-arithmetic and cannot be supported by homomorphic encryption. To implement the complex activation function in the encrypted domain without interaction, we design methods to approximate complex activation functions with low-degree polynomials while preserving the keyword spotting performance. Our scheme supports single-instruction multiple-data (SIMD), which reduces the total size of ciphertexts and improves computational efficiency. We conducted extensive experiments to investigate our performance with various metrics, such as accuracy, robustness, and F1-score. The experimental results show that our approach significantly outperforms the state-of-the-art solutions on every metric.
Peijia Zheng, Zhiwei Cai, Huicong Zeng, Jiwu Huang
ACM Multimedia4
2022 ISP-GAN: inception sub-pixel deconvolution-based lightweight GANs for colorization
Long Zhuo, Shunquan Tan, Bin Li 0011, Jiwu Huang
Multim. Tools Appl.4
2022 Efficient image tampering localization using semi-fragile watermarking and error control codes
Pascal Lefèvre, Philippe Carré, Caroline Fontaine, Philippe Gaborit, Jiwu Huang
Signal Process.5
2022 New design paradigm of distortion cost function for efficient JPEG steganography
Wenkang Su 0001, Jiangqun Ni, Xianglei Hu, Jiwu Huang
Signal Process.4
2022 Evading generated-image detectors: A deep dithering approach
Hao Xie 0002, Jiangqun Ni, Jian Zhang 0086, Weizhe Zhang, Jiwu Huang
Signal Process.5
2022 Corrigendum to 'Evading generated-image detectors: A deep dithering approach' [Signal Processing 197(2022) 108558]
Hao Xie 0002, Jiangqun Ni, Jian Zhang 0086, Weizhe Zhang, Jiwu Huang
Signal Process.5
2022 Hybrid deep-learning framework for object-based forgery detection in video
Shunquan Tan, Baoying Chen, Jishen Zeng, Bin Li 0011, Jiwu Huang
Signal Process. Image Commun.5
2022 Synthetic Speech Detection Based on Local Autoregression and Variance Statistics
Sanshuai Cui, Bingyuan Huang, Jiwu Huang, Xiangui Kang
IEEE Signal Process. Lett.3
2022 One-Class Double Compression Detection of Advanced Videos Based on Simple Gaussian Distribution Model
abstract
Passive video forensics has become an active topic in recent years. Generally, a pristine video obtained from surveillance cameras or other video recording devices is single compressed. At the same time, double lossy compression will certainly be introduced in tampered videos since it needs to go through recompression to perform tampering on a commonly used compressed video. In the video’s double compression, some traces are left due to the intrinsic effects of recompression. Traditional supervised learning is inefficient because it requires two classes (the pristine and the manipulated) that occur in the video to be exhaustively assigned labels. Actually, compared with pristine videos, manipulated videos with labels are more difficult to obtain. To address this problem, in this paper, one-class classification, which is often used for anomaly detection and only needs the target class, is introduced. We first treat all decompressed video frames as still images and extract subtractive pixel adjacency matrix (SPAM) steganalysis features to detect traces left in the double compression process. Then we adopt a Gaussian density-based one-class classifier since SPAM features extracted from pristine video frames approximately subject to the Gaussian distribution. Furthermore, we improve the robustness of the classifier by using ensemble strategy. Experimental results indicate that our proposed method exceeds other more complex one-class classification methods, and outperforms fully-supervised learning methods only by feeding features from single compressed video frames to our one-class classifier.
Qiushi Li 0001, Shengda Chen, Shunquan Tan, Bin Li 0011, Jiwu Huang
IEEE Trans. Circuits Syst. Video Technol.5
2022 A Novel Video Steganographic Scheme Incorporating the Consistency Degree of Motion Vectors
abstract
In this letter, a novel steganographic scheme in motion vector domain (MV) for H.264 video is presented, which can significantly improve the security performance against the newly emerged powerful multi-domain feature set MVC (motion vector consistency). By taking into account both the consistency degree of motion vectors for sub-blocks within a macroblock (MB) or sub-macroblock (sub-MB), and the MV statistics, the corresponding distortion function called dMVC is proposed. The proposed dMVC is also shown to be capable of integrating with existing methods to resist the joint steganalytic attacks of both MVC feature and local optimality features, e.g., NPELO, in the framework of minimal distortion embedding. Compared with other state-of-the-art MV-based steganographic schemes, experimental results on YUV sequences at various embedding rates and QPs show that the proposed method gains significant performance improvement while maintaining good coding efficiency.
Ying Liu 0062, Jiangqun Ni, Weizhe Zhang, Jiwu Huang
IEEE Trans. Circuits Syst. Video Technol.4
2022 Gradually Enhanced Adversarial Perturbations on Color Pixel Vectors for Image Steganography
abstract
Compared to element-wise embedding, vector-wise embedding based on CPV (color pixel vector) shows its superiority in color image steganography. However, when working with an adversarial embedding scheme for introducing adversarial perturbations, its success rate of deceiving a target CNN (convolutional neural network) steganalyzer dramatically drops. In this paper, inspired by the I-FGSM (iterative fast gradient sign method), we present an effective steganography for color images. Specifically, after decomposing an image into several non-overlapped sub-images, we iteratively and gradually increase the possibilities of generating adversarial perturbations for the CPVs in each sub-image by changing their adversarial costs. The costs are incrementally adjusted with a small step so that their maximum relative variation is minimized. Leveraging a new designed cost adjustment criterion, more modification patterns of CPV can participate in producing effective adversarial perturbations. Extensive experiments demonstrate that the proposed method achieves a high success rate in deceiving the target CNN steganalyzer and stably defending against the detection of other non-target steganalytic schemes for color images.
Xinghong Qin, Bin Li 0011, Shunquan Tan, Weixuan Tang 0004, Jiwu Huang
IEEE Trans. Circuits Syst. Video Technol.5
2022 Improving Cost Learning for JPEG Steganography by Exploiting JPEG Domain Knowledge
abstract
Although significant progress has been achieved recently in automatic learning of steganographic cost, the existing methods designed for spatial images cannot be directly applied to JPEG images which are more common media in daily life. The difficulties of migration are mainly caused by the characteristics of the$8\times 8$DCT mode structure. To address the issue, in this paper we extend an existing automatic cost learning scheme to JPEG, where the proposed scheme called JEC-RL (JPEG Embedding Cost with Reinforcement Learning) is explicitly designed to tailor the JPEG DCT structure. It works with the embedding action sampling mechanism under reinforcement learning, where a policy network learns the optimal embedding policies via maximizing the rewards provided by an environment network. Following a domain-transition design paradigm, the policy network is composed of three modules, i.e., pixel-level texture complexity evaluation module, DCT feature extraction module, and mode-wise rearrangement module. These modules operate in serial, gradually extracting useful features from a decompressed JPEG image and converting them into embedding policies for DCT elements, while considering JPEG characteristics including inter-block and intra-block correlations simultaneously. The environment network is designed in a gradient-oriented way to provide stable reward values by using a wide architecture equipped with a fixed preprocessing layer with$8\times 8$DCT basis filters. Extensive experiments and ablation studies demonstrate that the proposed method can achieve good security performance for JPEG images against both advanced feature-based and modern CNN-based steganalyzers.
Weixuan Tang 0004, Bin Li 0011, Mauro Barni, Jin Li 0002, Jiwu Huang
IEEE Trans. Circuits Syst. Video Technol.5
2022 Landmarking for Navigational Streaming of Stored High-Dimensional Media
abstract
Modern media data such as 360° videos and light field (LF) images are typically captured in much higher dimensions than the observers’ visual displays. To efficiently browse high-dimensional media, a navigational streaming model is considered: a client navigates the media space by dictating a navigation path to a server, who in response transmits the corresponding pre-encoded media data units (MDU) to the client one-by-one in sequence. Assuming that the MDU quality is pre-chosen and fixed, the problem resides in selecting and storing redundant representations of MDUs at the server in order to best trade off storage and transmission costs, while enabling adequate user’s random access. We address this problem with a landmark-based MDU optimization framework. The media space is divided into neighborhoods, each containing one landmark (a chosen MDU). MDUs in a neighborhood use the associated landmark as a predictor for inter-coding. Thus, for any MDU transition within the same neighborhood, only one inter-coded MDU transmission is required when the landmark resides in the decoder buffer. It results in lower transmission cost and enables navigational random access. To optimize an MDU structure, we employ tree-structured vector quantizer (TSVQ) to first optimize landmark locations, then iteratively add P-MDUs as refinements using a fast branch-and-bound technique. Taking interactive LF images and viewport adaptive 360° images as illustrative applications, and I-, P- and previously proposed merge frames to intra- and inter-code MDUs, we show experimentally that landmarked MDU structures can noticeably reduce the expected transmission cost compared with MDU structures without landmarks.
Yuan Yuan 0007, Gene Cheung, Pascal Frossard, H. Vicky Zhao, Jiwu Huang
IEEE Trans. Circuits Syst. Video Technol.5
2022 Secure Halftone Image Steganography Based on Feature Space and Layer Embedding
abstract
Syndrome-trellis codes (STCs) are commonly used in image steganographic schemes, which aim at minimizing the embedding distortion, but most distortion models cannot capture the mutual interaction of embedding modifications (MIEMs). In this article, a secure halftone image steganographic scheme based on a feature space and layer embedding is proposed. First, a feature space is constructed by a characterization method that is designed based on the statistics of 4 ×4 pixel blocks in halftone images. Upon the feature space, a generalized steganalyzer with good classification ability is proposed, which is used to measure the embedding distortion. As a result, a distortion model based on a hybrid feature space is constructed, which outperforms some state-of-the-art models. Then, as the distortion model is established on the statistics of local regions, a layer embedding strategy is proposed to reduce MIEM. It divides the host image into multiple layers according to their relative positions in 4 ×4 blocks, and the embedding procedure is executed layer by layer. In each layer, any two pixels are located at different 4 ×4 blocks in the original image, and the distortion model makes sure that the calculation of pixel distortions is independent. Between layers, the pixel distortions of the current layer are updated according to the previous embedding modifications, thus reducing the total embedding distortion. Comparisons with prior schemes demonstrate that the proposed steganographic scheme achieves high statistical security when resisting the state-of-the-art steganalysis.
Wei Lu 0001, Junjia Chen, Junhong Zhang, Jiwu Huang, Jian Weng 0001, Yicong Zhou
IEEE Trans. Cybern.4
2022 Robust Estimation of Upscaling Factor on Double JPEG Compressed Images
abstract
As one of the most important topics in image forensics, resampling detection has developed rapidly in recent years. However, the robustness to JPEG compression is still challenging for most classical spectrum-based methods, since JPEG compression severely degrades the image contents and introduces block artifacts in the boundary of the compression grid. In this article, we propose a method to estimate the upscaling factors on double JPEG compressed images in the presence of image upscaling between the two compressions. We first analyze the spectrum of scaled images and give an overall formulation of how the scaling factors along with the parameters of JPEG compression and image contents influence the appearance of tampering artifacts. The expected positions of five kinds of characteristic peaks are analytically derived. Then, we analyze the features of double JPEG compressed images in the block discrete cosine transform (BDCT) domain and present an inverse scaling strategy for the upscaling factor estimation with a detailed proof. Finally, a fusion method is proposed that through frequency-domain analysis, a candidate set of upscaling factors is given, and through analysis in the BDCT domain, the optimal estimation from all candidates is determined. The experimental results demonstrate that the proposed method outperforms other state-of-the-art methods.
Wei Lu 0001, Shangjun Luo, Yicong Zhou, Jiwu Huang, Yun Q. Shi 0001
IEEE Trans. Cybern.5
2022 Secret Sharing Based Reversible Data Hiding in Encrypted Images With Multiple Data-Hiders
abstract
The existing models of reversible data hiding in encrypted images (RDH-EI) are based on single data-hider, where the original image cannot be reconstructed when the data-hider is damaged. To address this issue, this article proposes a novel model with multiple data-hiders for RDH-EI based on secret sharing. It divides the original image into multiple different encrypted images with the same size of the original image and distributes them to multiple different data-hiders for data hiding. Each data-hider can independently embed data into the encrypted image to obtain the corresponding marked encrypted image. The original image can be losslessly recovered by collecting sufficient marked encrypted images from undamaged data-hiders when individual data-hiders are subjected to potential damage. This further protects the security of the original image. We provide four cases of the proposed model, namely, two joint cases and two separable cases. From the proposed model, we derive a separable RDH-EI method with high-capacity. Experimental results are presented to illustrate the effectiveness of the proposed method.
Bing Chen 0004, Wei Lu 0001, Jiwu Huang, Jian Weng 0001, Yicong Zhou
IEEE Trans. Dependable Secur. Comput.3
2022 Domain-Agnostic Document Authentication Against Practical Recapturing Attacks
abstract
Recapturing attack can be employed as a simple but effective anti-forensic tool for digital document images. Inspired by the document inspection process that compares a questioned document against some known samples, we proposed a document recapture detection scheme by employing a Siamese network to compare and extract distinct features in a recaptured document image. The proposed algorithm takes advantage of both metric learning and image forensic techniques, and forms triplets by considering some important factors in document authentication, e.g., document types, resolutions, and content in each image patch. After training with our triplet selection strategy, the resulting feature embedding clusters the genuine samples near the reference while pushing the recaptured samples apart. In the experiment, we consider practical settings under domain differences, such as the variations in printing/imaging devices, substrates, recapturing channels, and document types. To evaluate the robustness of different approaches, we benchmark some popular off-the-shelf machine learning-based approaches, a state-of-the-art document image detection scheme, and the proposed schemes with different network backbones under various experimental protocols. Experimental results show that the proposed scheme consistently outperforms the state-of-the-art approaches under different experimental settings. Specifically, under the most challenging scenario in our experiment, i.e., evaluation across different types of documents (produced by different manufacturers, devices, and substrates), we have achieved 6.92% APCER (Attack Presentation Classification Error Rate) and 8.51% BPCER (Bona Fide Presentation Classification Error Rate) by the proposed network with ResNeXt101 backbone at 5.00% BPCER decision threshold.
Changsheng Chen 0001, Shuzheng Zhang, Fengbo Lan, Jiwu Huang
IEEE Trans. Inf. Forensics Secur.4
2022 Document Recapture Detection Based on a Unified Distortion Model of Halftone Cells
abstract
In recent years, digital copies of paper documents are used widely with the prevalence of various online services. As a result, it is critical to validate the authenticity of the uploaded document images to protect against attacks from malicious users. Out of various types of attacks, the recapture attack (by reprinting and recapturing) is effective in concealing the trace of document forgeries. However, detecting the recaptured document images is challenging. To address this problem, we first study the halftone cell distortion introduced in both the genuine and recaptured document images. Based on our study, a unified model that characterizes the halftone cell distortion (e.g., errors in size and displacement) is then proposed for accurate estimation of the distortion parameters. The statistics of the estimated parameters are then exploited in a hypothesis testing framework to detect the recaptured document images. The questioned document image can be authenticated by testing against the null hypothesis, i.e., the image is a genuine sample. To evaluate the performance of the proposed approach under different application scenarios, extensive experiments are conducted with different prior knowledge of printers (known printer model, known printing technique, and In-The-Wild (unknown printing device and document contents)). The experiment results show that the proposed approach outperforms the data-driven benchmark approaches by a significant margin. Specifically, under the In-The-Wild experiment protocol, the Area Under the Receiver Operating Characteristic (ROC) Curve (AUC) of the proposed approach is above 0.87 while the AUC of the benchmark approaches (even some utilize both genuine and recaptured samples) degrades to less than 0.77.
Zhaoxu Hu, Changsheng Chen 0001, Wai Ho Mow, Jiwu Huang
IEEE Trans. Inf. Forensics Secur.4
2022 Universal Deep Network for Steganalysis of Color Image Based on Channel Representation
abstract
Up to now, most existing steganalytic methods were designed for grayscale images, and are not suitable for the color images that are widely used in social networks. In this paper, we design a universal color image steganalysis network (called UCNet) for the spatial and JPEG domains. The proposed method includes preprocessing, convolutional, and classification modules. To preserve the steganalytic features in each color channel, the preprocessing module first separates the input image into three channels based on the corresponding embedding spaces (i.e., RGB in the spatial domain, and YCbCr in the JPEG domain), and then extracts the image residuals with 62 fixed high-pass filters. Finally, all truncated residuals are concatenated for subsequent analysis, rather than adding them together in the first layer as in existing CNN-based steganalyzers. To accelerate network convergence and effectively reduce the number of parameters, the convolutional module contains three carefully designed types of layers with different shortcut connections and group convolution structures, to further learn the high-level steganalytic features. In the classification module, we employ global average pooling and a fully connected layer for classification. We conduct extensive experiments on ALASKA II to demonstrate that the proposed method can achieve state-of-the-art results that are comparable with other modern CNN-based steganalyzers (e.g., SRNet and LC-Net) in both the spatial and JPEG domains, with relatively few memory requirements and short training times. Furthermore, we also provide some necessary descriptions and carry out numerous ablation experiments to verify the rationality of the network design.
Kangkang Wei, Weiqi Luo 0001, Shunquan Tan, Jiwu Huang
IEEE Trans. Inf. Forensics Secur.4
2022 Self-Adversarial Training Incorporating Forgery Attention for Image Forgery Localization
abstract
Image editing techniques enable people to modify the content of an image without leaving visual traces and thus may cause serious security risks. Hence the detection and localization of these forgeries become quite necessary and challenging. Furthermore, unlike other tasks with extensive data, there is usually a lack of annotated forged images for training due to annotation difficulties. In this paper, we propose a self-adversarial training strategy and a reliable coarse-to-fine network that utilizes a self-attention mechanism to localize forged regions in forgery images. The self-attention module is based on a Channel-Wise High Pass Filter block (CW-HPF). CW-HPF leverages inter-channel relationships of features and extracts noise features by high pass filters. Based on the CW-HPF, a self-attention mechanism, calledforgery attention, is proposed to capture rich contextual dependencies of intrinsic inconsistency extracted from tampered regions. Specifically, we append two types of attention modules on top of CW-HPF respectively to model internal interdependencies in spatial dimension and external dependencies among channels. We exploit a coarse-to-fine network to enhance the noise inconsistency between original and tampered regions. More importantly, to address the issue of insufficient training data, we design a self-adversarial training strategy that expands training data dynamically to achieve more robust performance. Specifically, in each training iteration, we perform adversarial attacks against our network to generate adversarial examples and train our model on them. The proposed method is based on the assumption of content-changed manipulations. Extensive experimental results demonstrate that our proposed algorithm steadily outperforms state-of-the-art methods by a clear margin in different benchmark datasets.
Long Zhuo, Shunquan Tan, Bin Li 0011, Jiwu Huang
IEEE Trans. Inf. Forensics Secur.4
2021 Image Steganography Based on Iterative Adversarial Perturbations Onto a Synchronized-Directions Sub-Image
abstract
Nowadays a steganography has to face challenges to both feature-based staganalysis and convolutional neural network (CNN) based steganalysis. In this paper, we present a novel steganographic scheme to incorporate synchronizing modification directions and iterative adversarial perturbations to enhance steganographic performance. Firstly an existing steganographic function is employed to compute initial costs. Then the secret message bits are embedded following clustering modification directions profile. If the target CNN classifier discriminates the resulting stego image as the correct class, we change costs in adversarial manners, and then choose a sub-image to re-embed message with changed costs. Adversarial intensity will be iteratively increased until the adversarial stego image can deceive the target CNN classifier, which guarantees that applied adversarial perturbations are minimal and it is unnecessary to search the optimal adversarial intensity. Experiments demonstrate that the proposed method effectively enhances security to counter both feature-based classifiers and CNN classifiers, no matter they are targeted or non-targeted.
Xinghong Qin, Shunquan Tan, Weixuan Tang 0004, Bin Li 0011, Jiwu Huang
ICASSP5
2021 A novel deep learning framework for double JPEG compression detection of small size blocks
Israr Hussain, Shunquan Tan, Bin Li 0011, Xinghong Qin, Dostdar Hussain, Jiwu Huang
J. Vis. Commun. Image Represent.6
2021 Detecting facial manipulated videos based on set convolutional neural networks
Zhaopeng Xu, Jiarui Liu 0002, Wei Lu 0001, Bozhi Xu, Xianfeng Zhao, Bin Li 0011, Jiwu Huang
J. Vis. Commun. Image Represent.7
2021 Secure Robust JPEG Steganography Based on AutoEncoder With Adaptive BCH Encoding
abstract
Social networks are everywhere and currently transmitting very large messages. As a result, transmitting secret messages in such an environment is worth researching. However, the images used in transmitting messages are usually compressed with a JPEG compression channel, which is lossy and damages the transmitted data. Therefore, to prevent secret messages from being damaged, a robust JPEG steganography is urgently needed. In this paper, a secure robust JPEG steganographic scheme based on an autoencoder with an adaptive BCH encoding (Bose-Chaudhuri-Hocquenghem encoding) is proposed. In particular, the autoencoder is first pretrained to fit the transformation relationship between the JPEG image before and after compression by the compression channel. In addition, the BCH encoding is adaptively utilized according to the content of cover image to decrease the error rate of secret message extraction. The DCT (Discrete Cosine Transformation) coefficient adjustment based on practical JPEG channel characteristics further improves the robustness and statistical security. Comparisons with prior state-of-the-art schemes demonstrate that the proposed robust JPEG steganographic algorithm can provide a more robust performance and statistical security.
Wei Lu 0001, Junhong Zhang, Xianfeng Zhao, Weiming Zhang 0001, Jiwu Huang
IEEE Trans. Circuits Syst. Video Technol.5
2021 Defeating Lattice-Based Data Hiding Code Via Decoding Security Hole
abstract
Lattice code has been widely used for data hiding. It can provide security for data hiding by randomly translating its codebook with a secret dither. However, besides the secret dither, there are an infinite amount of points that are near the secret dither and can also be used for perfect decoding in the noiseless scenario. This means that lattice-based data hiding has a serious security hole, named decoding security hole (DSH) in this paper. After a theoretical analysis of DSH, we find that these points form a convex polytope and the centroid of this convex polytope is the secret dither. Based on this finding, a simple yet effective attack method is presented to estimate the secret dither of lattice-based data hiding. Extensive experimental results show that the proposed method significantly outperforms state-of-the-art attack methods, especially when the number of observations is small or the document-to-watermark ratio changes over a wide range.
Yuan-Gen Wang, Guopu Zhu, Jin Li 0002, Mauro Conti, Jiwu Huang
IEEE Trans. Circuits Syst. Video Technol.5
2021 Reversible Data Hiding in Halftone Images Based on Dynamic Embedding States Group
abstract
In many reversible data hiding (RDH) methods for halftone images, the traditional embedding process embeds a 1-bit secret message into each embeddable pixel or pattern. To improve the embedding efficiency and payload, we propose an RDH method used in halftone images based on the dynamic embedding states group (DESG), which can embed at least 1 bit of secret messages per embeddable pixel or pattern. First, by exploiting the statistical features of$4 \times 4$patterns and the state sequences in each image, the DESG is constructed dynamically, including$n$embedding states with their state patterns and state sequences. Then, secret messages are encoded by matching the longest common subsequence according to the DESG, which are split into several state sequences. The state sequences are embedded by Markov transitions between these$n$changing state patterns. Finally, reversibility is achieved by recording the DESG as the overhead information in RDH. Experiments show that the construction of DESG can improve the embedding efficiency under the same number of embeddable pixels or patterns, and the visual distortion is also significantly reduced by flipping fewer pixels.
Xiaolin Yin, Wei Lu 0001, Wanteng Liu, Jing-Ming Guo, Jiwu Huang, Yun Q. Shi 0001
IEEE Trans. Circuits Syst. Video Technol.5
2021 Secure Halftone Image Steganography Based on Pixel Density Transition
abstract
Most state-of-the-art halftone image steganographic techniques only consider the flipping distortion according to the human visual system, which are not always secure when they are attacked by steganalyzers. In this paper, we propose a halftone image steganographic scheme that aims to generate stego images with good visual quality and strong statistical security of anti-steganalysis. First, the concept of pixel density is proposed and a novel construction called pixel density histogram (PDH) is proposed to design a “embedding” scheme for halftone images. Then, we optimize density pair selection to select density blocks that can improve visual quality. Finally, the messages are embedded through pixel density transition, where a novel pixel flipping strategy is proposed, which can maintain the structural dependence by optimizing the pixel mesh Markov transition matrix (PMMTM). The experimental results demonstrate that the proposed steganography scheme can achieve strong statistical security of anti-steganalysis with good visual quality without degrading the embedding capacity.
Wei Lu 0001, Yingjie Xue, Yuileong Yeung, Hongmei Liu 0001, Jiwu Huang, Yun Q. Shi 0001
IEEE Trans. Dependable Secur. Comput.5
2021 Efficient JPEG Batch Steganography Using Intrinsic Energy of Image Contents
abstract
Batch steganography aims at properly allocating a large payload to multiple covers, so as to keep the whole covert communication at a satisfactory level of security. JPEG is currently one of the most widely used formats for image storage and transmission. This paper presents an efficient JPEG batch steganographic scheme, which allocates the payload in a linear manner w.r.t. a new heuristic measure - the intrinsic energy of JPEG image contents, in which more concerns are with the high frequency components, and the proposed measure could also be easily generalized to cover selection in batch steganographic applications. And a calibration strategy is elaborately designed to balance the security level when JPEG covers of various QFs are involved in JPEG batch steganography. In this way, the proposed scheme can effectively resolve the problem that the statistical undetectability fluctuates dramatically w.r.t. the size and quality factor when the batch set is involved with various image parameters, and consequently maintains the overall security of the practical JPEG batch steganographic system. Experimental results show that the proposed method exhibits security performance superior or comparable to the state-of-the-art batch schemes while maintaining a low computational cost.
Xianglei Hu, Jiangqun Ni, Weizhe Zhang, Jiwu Huang
IEEE Trans. Inf. Forensics Secur.4
2021 A New Adversarial Embedding Method for Enhancing Image Steganography
abstract
Image steganography aims to embed secret messages into cover images in an imperceptible manner. While steganalysis tries to identify stegos from covers, which is a special binary classification problem. Recently, some literatures show that the adversarial embedding can mislead the advanced steganalyzers based on convolutional neural network (CNN), and thus enhance the steganography security. Since adding perturbations to stegos may lead to messages extraction failure due to properties of syndrome-trellis codes (STC), the existing adversarial examples are derived from covers or their enhanced versions, while those stegos are not fully utilized. In this paper, we propose a new adversarial embedding scheme for image steganography. Unlike those related works, we first combine multiple gradients of cover and generated stegos to determine the directions of cost modifications. Next, instead of adjusting all or a random part of embedding costs in existing works, we carefully select the candidate costs according to the amplitudes of cover gradients and their costs. Extensive experimental results demonstrate that by adjusting a tiny part of embedding costs (less than 5% in most cases), the proposed method can significantly improve the security of five modern steganographic methods evaluated on both re-trained CNN-based and traditional steganalyzers, and achieve much better security performances compared with related methods. In addition, the security performances evaluated on different image database show that the generalization of the proposed method is good.
Minglin Liu, Weiqi Luo 0001, Peijia Zheng, Jiwu Huang
IEEE Trans. Inf. Forensics Secur.4
2021 MCTSteg: A Monte Carlo Tree Search-Based Reinforcement Learning Framework for Universal Non-Additive Steganography
abstract
Recent research has shown that non-additive image steganographic frameworks effectively improve security performance through adjusting distortion distribution. However, as far as we know, all of the existing non-additive proposals are based on handcrafted policies, and can only be applied to a specific image domain, which heavily prevent non-additive steganography from releasing its full potentiality. In this paper, we propose an automatic non-additive steganographic distortion learning framework called MCTSteg to remove the above restrictions. Guided by the reinforcement learning paradigm, we combine Monte Carlo Tree Search (MCTS) and steganalyzer-based environmental model to build MCTSteg. MCTS makes sequential decisions to adjust distortion distribution without human intervention. Our proposed environmental model is used to obtain feedbacks from each decision. Due to its self-learning characteristic and domain-independent reward function, MCTSteg has become the first reported universal non-additive steganographic framework which can work in both spatial and JPEG domains. Extensive experimental results show that MCTSteg can effectively withstand the detection of both hand-crafted feature-based and deep-learning-based steganalyzers. In both spatial and JPEG domains, the security performance of MCTSteg steadily outperforms the state of the art by a clear margin under different scenarios.
Xianbo Mo, Shunquan Tan, Bin Li 0011, Jiwu Huang
IEEE Trans. Inf. Forensics Secur.4
2021 CALPA-NET: Channel-Pruning-Assisted Deep Residual Network for Steganalysis of Digital Images
abstract
Over the past few years, detection performance improvements of deep-learning based steganalyzers have been usually achieved through structure expansion. However, excessive expanded structure results in huge computational cost, storage overheads, and consequently difficulty in training and deployment. In this paper we propose CALPA-NET, a ChAnneL-Pruning-Assisted deep residual network architecture search approach to shrink the network structure of existing vast, over-parameterized deep-learning based steganalyzers. We observe that the broad inverted-pyramid structure of existing deep-learning based steganalyzers might contradict the well-established model diversity oriented philosophy, and therefore is not suitable for steganalysis. Then a hybrid criterion combined with two network pruning schemes is introduced to adaptively shrink every involved convolutional layer in a data-driven manner. The resulting network architecture presents a slender bottleneck-like structure. We have conducted extensive experiments on BOSSBase + BOWS2 dataset, more diverse ALASKA dataset and even a large-scale subset extracted from ImageNet CLS-LOC dataset. The experimental results show that the model structure generated by our proposed CALPA-NET can achieve comparative performance with less than two percent of parameters and about one third FLOPs compared to the original steganalytic model. The new model possesses even better adaptivity, transferability, and scalability.
Shunquan Tan, Weilong Wu, Zilong Shao, Qiushi Li 0001, Bin Li 0011, Jiwu Huang
IEEE Trans. Inf. Forensics Secur.6
2021 An Automatic Cost Learning Framework for Image Steganography Using Deep Reinforcement Learning
abstract
Automatic cost learning for steganography based on deep neural networks is receiving increasing attention. Steganographic methods under such a framework have been shown to achieve better security performance than methods adopting hand-crafted costs. However, they still exhibit some limitations that prevent a full exploitation of their potentiality, including using a function-approximated neural-network-based embedding simulator and a coarse-grained optimization objective without explicitly using pixel-wise information. In this article, we propose a new embedding cost learning framework called SPAR-RL (Steganographic Pixel-wise Actions and Rewards with Reinforcement Learning) that overcomes the above limitations. In SPAR-RL, an agent utilizes a policy network which decomposes the embedding process into pixel-wise actions and aims at maximizing the total rewards from a simulated steganalytic environment, while the environment employs an environment network for pixel-wise reward assignment. A sampling process is utilized to emulate the message embedding of an optimal embedding simulator. Through the iterative interactions between the agent and the environment, the policy network learns a secure embedding policy which can be converted into pixel-wise embedding costs for practical message embedding. Experimental results demonstrate that the proposed framework achieves state-of-the-art security performance against various modern steganalyzers, and outperforms existing cost learning frameworks with regard to learning stability and efficiency.
Weixuan Tang 0004, Bin Li 0011, Mauro Barni, Jin Li 0002, Jiwu Huang
IEEE Trans. Inf. Forensics Secur.5
2021 Robust Privacy-Preserving Motion Detection and Object Tracking in Encrypted Streaming Video
abstract
Video privacy leakage is becoming an increasingly severe public problem, especially in cloud-based video surveillance systems. It leads to the new need for secure cloud-based video applications, where the video is encrypted for privacy protection. Despite some methods that have been proposed for encrypted video moving object detection and tracking, none has robust performance against complex and dynamic scenes. In this paper, we propose an efficient and robust privacy-preserving motion detection and multiple object tracking scheme for encrypted surveillance video bitstreams. By analyzing the properties of the video codec and format-compliant encryption schemes, we propose a new compressed-domain feature to capture motion information in complex surveillance scenarios. Based on this feature, we design an adaptive clustering algorithm for moving object segmentation with an accuracy of 4×4 pixels. We then propose a multiple object tracking scheme that uses Kalman filter estimation and adaptive measurement refinement. The proposed scheme does not require video decryption or full decompression and has a very low computation load. The experimental results demonstrate that our scheme achieves the best detection and tracking performance compared with existing works in the encrypted and compressed domain. Our scheme can be effectively used in complex surveillance scenarios with different challenges, such as camera movement/jitter, dynamic background, and shadows.
Xianhao Tian, Peijia Zheng, Jiwu Huang
IEEE Trans. Inf. Forensics Secur.3
2021 Image Tampering Localization Using a Dense Fully Convolutional Network
abstract
The emergence of powerful image editing software has substantially facilitated digital image tampering, leading to many security issues. Hence, it is urgent to identify tampered images and localize tampered regions. Although much attention has been devoted to image tampering localization in recent years, it is still challenging to perform tampering localization in practical forensic applications. The reasons include the difficulty of learning discriminative representations of tampering traces and the lack of realistic tampered images for training. Since Photoshop is widely used for image tampering in practice, this paper attempts to address the issue of tampering localization by focusing on the detection of commonly used editing tools and operations in Photoshop. In order to well capture tampering traces, a fully convolutional encoder-decoder architecture is designed, where dense connections and dilated convolutions are adopted for achieving better localization performance. In order to effectively train a model in the case of insufficient tampered images, we design a training data generation strategy by resorting to Photoshop scripting, which can imitate human manipulations and generate large-scale training samples. Extensive experimental results show that the proposed approach outperforms state-of-the-art competitors when the model is trained with only generated images or fine-tuned with a small amount of realistic tampered images. The proposed method also has good robustness against some common post-processing operations.
Peiyu Zhuang, Haodong Li 0001, Shunquan Tan, Bin Li 0011, Jiwu Huang
IEEE Trans. Inf. Forensics Secur.5
2021 Deep Learning-Based Forgery Attack on Document Images
abstract
With the ongoing popularization of online services, the digital document images have been used in various applications. Meanwhile, there have emerged some deep learning-based text editing algorithms which alter the textual information of an image in an end-to-end fashion. In this work, we present a low-cost document forgery algorithm by the existing deep learning-based technologies to edit practical document images. To achieve this goal, the limitations of existing text editing algorithms towards complicated characters and complex background are addressed by a set of network design strategies. First, the unnecessary confusion in the supervision data is avoided by disentangling the textual and background information in the source images. Second, to capture the structure of some complicated components, the text skeleton is provided as auxiliary information and the continuity in texture is considered explicitly in the loss function. Third, the forgery traces induced by the text editing operation are mitigated by some post-processing operations which consider the distortions from the print-and-scan channel. Quantitative comparisons of the proposed method and the exiting approach have shown the advantages of our design by reducing the about 2/3 reconstruction error measured in MSE, improving reconstruction quality measured in PSNR and in SSIM by 4 dB and 0.21, respectively. Qualitative experiments have confirmed that the reconstruction results of the proposed method are visually better than the existing approach in both complicated characters and complex texture. More importantly, we have demonstrated the performance of the proposed document forgery algorithm under a practical scenario where an attacker is able to alter the textual information in an identity document using only one sample in the target domain. The forged-and-recaptured samples created by the proposed text editing attack and recapturing operation have successfully fooled some existing document authentication systems.
Lin Zhao 0017, Changsheng Chen 0001, Jiwu Huang
IEEE Trans. Image Process.3
2020 Image processing operations identification via convolutional neural network
Haodong Li 0001, Weiqi Luo 0001, Jiwu Huang
Sci. China Inf. Sci.4
2020 Universal stego post-processing for enhancing image steganography
Weiqi Luo 0001, Peijia Zheng, Jiwu Huang
J. Inf. Secur. Appl.4
2020 Identification of deep network generated images using disparities in color components
Haodong Li 0001, Bin Li 0011, Shunquan Tan, Jiwu Huang
Signal Process.4
2020 A novel selective encryption scheme for H.264/AVC video with improved visual security
Yuzhang Xu, Weiqi Luo 0001, Shaohua Tang, Jiwu Huang
Signal Process. Image Commun.5
2020 Efficient Privacy-Preserving Anomaly Detection and Localization in Bitstream Video
abstract
In cloud computing, videos may be in an encrypted format to protect privacy. Therefore, encrypted video processing is an important application in secure cloud computing. In this paper, we focus on parameter estimation and anomaly detection in an encrypted video bitstream. By analyzing the common properties of video encoding frameworks and the format-compliant encryption schemes, we propose an anomaly detection scheme for encrypted video bitstream with format-compliant encryption. From the encrypted bitstream, we extract three types of complementary features, i.e., the macroblock sizes, the macroblock partitions, and the motion vector difference magnitude, and then propose a method to combine these three features. The proposed detection and localization scheme does not involve video decryption, full decompression, or an interactive protocol, which makes it efficient. Our scheme is also compatible with different video encryption methods. To accelerate the running time, we develop a parallel implementation for our scheme. The experimental results show that our method achieves good running time and detection rate performance.
Jianting Guo, Peijia Zheng, Jiwu Huang
IEEE Trans. Circuits Syst. Video Technol.3
2020 Downscaling Factor Estimation on Pre-JPEG Compressed Images
abstract
Resampling detection is one of the most important topics in image forensics, and the most widely used method in resampling detection is spectral analysis. Since JPEG is the most widely used image format, it is reasonable that the resampling operation is processed on JPEG images. JPEG block artifacts bring severe interference to spectrum-based methods and degrade the detection performance. In addition, the spectral characteristics of the downscaling scenarios are very weak. The detection of downscaling still presents a considerable challenge to forensic applications. In this paper, we propose a method to estimate the downscaling factors of pre-JPEG compressed images in the presence of image downscaling after JPEG compressions. We first analyze the spectrum of scaled images and give an exact formulation of how the scaling factors influence the appearance of periodic artifacts. The expected positions of the characteristic resampling peaks are analytically derived. For the downscaling scenario, the shifted JPEG block artifacts produce periodic peaks, which cause misdetection in the characteristic peak. We find that the interval between the adjacent extrema of difference images obeys the geometric distribution and the distribution has periodic peaks for JPEG images. Hence, we adopt the difference image extremum interval histogram and combine the spectral method to obtain the final estimation. The experimental results demonstrate that the proposed detection method outperforms some state-of-the-art methods.
Xianjin Liu, Wei Lu 0001, Jiwu Huang, Yun Q. Shi 0001
IEEE Trans. Circuits Syst. Video Technol.4
2020 Binary Image Steganalysis Based on Histogram of Structuring Elements
abstract
Utilizing statistical models of binary images is a common and effective means to steganalyze binary images, and the design of the statistical model is essential to the performance of steganalysis. In this paper, we propose a new model based on a histogram of pixel structuring elements (SEs), which is a suitable representation of a binary image for the task of steganalysis. The texture property and the dependency among pixels are considered inside the SEs. The SEs with different patterns will be evaluated comprehensively according to a statistical criterion, and some of them will be selected to construct the feature set for training the steganalyzer. The distributions of these selected SEs, which contain many highly flippable pixels, will be emphasized by the criterion, and they can reflect the difference between cover images and stego-images. Finally, a series of experiments are conducted on two datasets, and the results show that the proposed scheme significantly outperforms state-of-the-art schemes.
Wei Lu 0001, Lingwen Zeng, Junjia Chen, Jiwu Huang, Yun Q. Shi 0001
IEEE Trans. Circuits Syst. Video Technol.5
2020 Secure Binary Image Steganography With Distortion Measurement Based on Prediction
abstract
In this paper, a binary image steganographic scheme is presented, which aims at minimizing the embedding distortions measured by prediction. A prediction model of the center pixel's value is established in a 3 × 3 local region. A concept of “uncertainty” is introduced to represent the prediction result and the uncertainty is defined as the proximity of probabilities about whether the center pixel is black or white. A pixel with high uncertainty means that it is hard to distinguish whether it has been flipped or not, and thus the distortion introduced by flipping this pixel is small. The uncertainty is an appended statistical explanation of human visual perception and the distortion measurement based on it can evaluate the embedding changes on both vision and statistics. Benefiting from the statistics, uncertainty can evaluate the distortion influence in an extended local region. To play the advantage of distortion measurement, the syndrome-trellis code (STC) is employed to minimize the embedding distortions. Comparisons with prior schemes demonstrate that the proposed steganographic scheme achieves high vision imperceptibility and statistical security.
Yuileong Yeung, Wei Lu 0001, Yingjie Xue, Jiwu Huang, Yun Q. Shi 0001
IEEE Trans. Circuits Syst. Video Technol.4
2020 A Copy-Proof Scheme Based on the Spectral and Spatial Barcoding Channel Models
abstract
The traditional two-dimensional (2D) barcode has been employed in anti-counterfeiting systems as a storage media for serial numbers. However, an attack can be initiated by simply copying the 2D barcode and attaching it to a counterfeit product. In this paper, we aim at proposing an authentication scheme with a mobile imaging device for a 2D barcode. This work presents a competitive solution among the 2D barcode authentication schemes that have been verified under mobile imaging conditions. The proposed copy-proof scheme is composed of two sets of features which are extracted by exploiting the characteristics of barcoding channel models. The proposed features identify the intrinsic differences between genuine and counterfeit barcode images in the frequency and spatial domains. An efficient two-stage barcode authentication framework is then proposed by combining the two sets of features in a cascading manner. To evaluate the practicality of the proposed authentication scheme, four databases with different devices (printers, scanners, mobile cameras), barcode sizes, and barcode designs are considered in the experiments. By comparing with the existing texture descriptors and some deep learning-based approaches, it is shown that the proposed scheme has a higher authentication accuracy under various conditions, such as cross-database, cross-size and cross-pattern experiments which study the generalities of a pre-trained model towards challenging conditions commonly found in real-world scenarios. Last but not least, the proposed scheme has been evaluated under some state-of-the-art attack scenarios where the attacker employs several realizations of genuine patterns or the deep learning-based technique to produce a counterfeit copy. The source code and data for producing the results in our experiments are available at https://bit.ly/2FOlJH7.
Changsheng Chen 0001, Mulin Li, Anselmo Ferreira, Jiwu Huang, Rizhao Cai
IEEE Trans. Inf. Forensics Secur.4
2020 Identification of VoIP Speech With Multiple Domain Deep Features
abstract
Identifying whether a phone call comes from VoIP (Voice over Internet Protocol) is a challenging but less-investigated audio forensic issue. As shown in a previous study, existing feature based methods do not work well. In this paper, we propose a robust data-driven approach, called CNN-MLS (convolutional neural network based multi-domain learning scheme), to distinguish VoIP calls from mobile phone calls. To better explore the differences between VoIP and mobile phone calls, we first process data with high-pass filtering, and then extract deep features from both temporal domain and spectral domain. Two CNN architectures are designed for accepting data from respective domains, and some tricks such as auxiliary classifiers and individual subnet training are used for accelerating network convergence. The deep features are finally fused in a classification module for identifying the phone call type. The proposed method is evaluated on VPCID (VoIP Phone Call Identification Database) dataset, under various testing conditions. We pay particular attention to tests on data belonging to a source mismatched with the training sources. Experimental results show that, compared with existing methods, our method can achieve satisfactory and better accuracy on two-second-long inputs, implying that an alert may be activated shortly after a VoIP call is made.
Yuankun Huang, Bin Li 0011, Mauro Barni, Jiwu Huang
IEEE Trans. Inf. Forensics Secur.4
2020 Face Spoofing Detection Based on Local Ternary Label Supervision in Fully Convolutional Networks
abstract
Face verification systems are prone to spoofing attacks on photos, videos, and 3D masks. Face spoofing detection, i.e., face anti-spoofing, face liveness detection, or face presentation attack detection, is an important task for securing face verification systems in practice and presents many challenges. In this paper, a state-of-the-art face spoofing detection method based on a depth-based Fully Convolutional Network (FCN) is revisited. Different supervision schemes, including global and local label supervisions, are comprehensively investigated. A generic theoretical analysis and associated simulation are provided to demonstrate that local label supervision is more suitable than global label supervision for local tasks with insufficient training samples, such as the face spoofing detection task. Based on the analysis, the Spatial Aggregation of Pixel-level Local Classifiers (SAPLC), which is composed of an FCN part and an aggregation part, is proposed. The FCN part predicts the pixel-level ternary labels, which include the genuine foreground, the spoofed foreground, and the undetermined background. Then, these labels are aggregated together to yield an accurate image-level decision. Furthermore, to quantitatively evaluate the proposed SAPLC, experiments are carried out on the CASIA-FASD, Replay-Attack, OULU-NPU, and SiW datasets. The experiments show that the proposed SAPLC outperforms the representative deep networks, including two globally supervised CNNs, one depth-based FCN, two FCNs with binary labels, and two FCNs with ternary labels, and achieves competitive performances close to some state-of-the-art method performances under various common protocols. Overall, the results empirically verify the advantage of the proposed pixel-level local label supervision scheme.
Wenyun Sun, Changsheng Chen 0001, Jiwu Huang, Alex Chichung Kot
IEEE Trans. Inf. Forensics Secur.4
2020 An Embedding Cost Learning Framework Using GAN
abstract
Successful adaptive steganography has mainly focused on embedding the payload while minimizing an appropriately defined distortion function. The application of deep learning to steganalysis has greatly challenged present adaptive steganographic methods, but has also shown the potential for the improvement of steganography. This paper proposes a distortion function generating a framework for steganography. It has three modules: a generator with a U-Net architecture to translate a cover image into an embedding change probability map, a no-pre-training-required double-tanh function to approximate the optimal embedding simulator while preserving gradient norm during backpropagation in the adversarial training, and an enhanced steganalyzer based on a convolution neural network together with multiple high pass filters as the discriminator. Extensive experimental results on different datasets have shown that the proposed framework outperforms the current state-of-the-art steganographic schemes. Moreover, the adversarial training time is reduced dramatically compared with the GAN-based automatic steganographic distortion learning framework (ASDL-GAN).
Danyang Ruan, Jiwu Huang, Xiangui Kang, Yun Q. Shi 0001
IEEE Trans. Inf. Forensics Secur.3
2019 A New Spatial Steganographic Scheme by Modeling Image Residuals with Multivariate Gaussian Model
abstract
Embedding costs used in content-adaptive image steganographic schemes can be defined in a heuristic way or with a statistical model. Inspired by previous steganographic methods, i.e., MG (multivariate Gaussian model) and MiPOD (minimizing the power of optimal detector), we propose a model-driven scheme in this paper. Firstly, we model image residuals obtained by high-pass filtering with quantized multivariate Gaussian distribution. Then, we derive the approximated Fisher Information (FI). We show that FI is related to both Gaussian variance and filter coefficients. Lastly, by selecting the maximum FI value derived with various filters as the final FI, we obtain embedding costs. Experimental results show that the proposed scheme is comparable to existing steganographic methods in resisting steganalysis equipped with rich models and selection-channel-aware rich models. It is also computational efficient when compared to MiPOD, which is the state-of-the-art model-driven method.
Xinghong Qin, Bin Li 0011, Jiwu Huang
ICASSP3
2019 Localization of Deep Inpainting Using High-Pass Fully Convolutional Network
abstract
Image inpainting has been substantially improved with deep learning in the past years. Deep inpainting can fill image regions with plausible contents, which are not visually apparent. Although inpainting is originally designed to repair images, it can even be used for malicious manipulations, e.g., removal of specific objects. Therefore, it is necessary to identify the presence of inpainting in an image. This paper presents a method to locate the regions manipulated by deep inpainting. The proposed method employs a fully convolutional network that is based on high-pass filtered image residuals. Firstly, we analyze and observe that the inpainted regions are more distinguishable from the untouched ones in the residual domain. Hence, a high-pass pre-filtering module is designed to get image residuals for enhancing inpainting traces. Then, a feature extraction module, which learns discriminative features from image residuals, is built with four concatenated ResNet blocks. The learned feature maps are finally enlarged by an up-sampling module, so that a pixel-wise inpainting localization map is obtained. The whole network is trained end-to-end with a loss addressing the class imbalance. Extensive experimental results evaluated on both synthetic and realistic images subjected to deep inpainting have shown the effectiveness of the proposed method.
Haodong Li 0001, Jiwu Huang
ICCV2
2019 A Novel High-Capacity Reversible Data Hiding Scheme for Encrypted JPEG Bitstreams
abstract
As cloud storage becomes more common, concerns about the invasion of privacy are increasing. When images are stored in an encrypted form in the public cloud, reversible data hiding in the encrypted domain can be applied to embed additional data within the encrypted images for ease of management. Most existing works focus on uncompressed images and are not applicable to JPEG images, which are widely used throughout the Internet. Therefore, in this paper, a novel reversible data hiding scheme for encrypted JPEG bitstreams is proposed. First, an effective method of bitstream-based JPEG image encryption is employed to encrypt plaintext JPEG images. Then, we present a reversible data hiding technique for encrypted JPEG images based on invariant zero-run length in the zero-run value pairs. In the cloud, additional data, such as labels, timestamps, origins, and authentication messages, are directly embedded into the encrypted JPEG images with our proposed data hiding technique. From the marked encrypted JPEG images, the extraction of the hidden data and the recovery of the original images can be done independently. Extensive experiments performed on several typical images and three well-known image databases show that the proposed scheme can achieve much higher embedding capacity than that of most recent schemes, and the file sizes of the marked encrypted JPEG images are well preserved compared to those of the related methods. In addition, we provide further analysis to show that the proposed scheme has good format compatibility and low computational complexity.
Junxi Chen, Weiqi Luo 0001, Shaohua Tang, Jiwu Huang
IEEE Trans. Circuits Syst. Video Technol.5
2019 Patchwork-Based Audio Watermarking Robust Against De-Synchronization and Recapturing Attacks
abstract
Watermarking is a solution for copyright protection and forensics tracking, but recapturing and de-synchronization attacks may be used to effectively remove audio watermarks. Although much effort has been made in recent years, the robustness of audio watermarking against recapturing and de-synchronization attacks is still a challenging issue. Specifically, we first construct the frequency-domain coefficients logarithmic mean (FDLM) feature of digital audio. By theoretical analysis, we conclude that the residual of the two groups' FDLM feature is robust against recapturing attack. We then propose a robust audio watermarking method based on this feature using the patchwork framework. Compared with the method having the best robustness performance against recapturing attack, the BER value of our method is decreased by 7%. Besides that, the proposed method outperforms the state-of-the-art patchwork-based watermarking methods notably, under recapturing and post-processed with signal processing operations and de-synchronization attacks.
Zhenghui Liu, Yuankun Huang, Jiwu Huang
IEEE Trans. Inf. Forensics Secur.3
2019 CNN-Based Adversarial Embedding for Image Steganography
abstract
Steganographic schemes are commonly designed in a way to preserve image statistics or steganalytic features. Since most of the state-of-the-art steganalytic methods employ a machine learning (ML)-based classifier, it is reasonable to consider countering steganalysis by trying to fool the ML classifiers. However, simply applying perturbations on stego images as adversarial examples may lead to the failure of data extraction and introduce unexpected artifacts detectable by other classifiers. In this paper, we present a steganographic scheme with a novel operation called adversarial embedding (ADV-EMB), which achieves the goal of hiding a stego message while at the same time fooling a convolutional neural network (CNN)-based steganalyzer. The proposed method works under the conventional framework of distortion minimization. In particular, ADV-EMB adjusts the costs of image elements modifications according to the gradients back propagated from the target CNN steganalyzer. Therefore, modification direction has a higher probability to be the same as the inverse sign of the gradient. In this way, the so-called adversarial stego images are generated. Experiments demonstrate that the proposed steganographic scheme achieves better security performance against the target adversary-unaware steganalyzer by increasing its missed detection rate. In addition, it deteriorates the performance of other adversary-aware steganalyzers, opening the way to a new class of modern steganographic schemes capable of overcoming powerful CNN-based steganalysis.
Weixuan Tang 0004, Bin Li 0011, Shunquan Tan, Mauro Barni, Jiwu Huang
IEEE Trans. Inf. Forensics Secur.5
2019 Robust Copy-Move Detection of Speech Recording Using Similarities of Pitch and Formant
abstract
Copy-move forgery on very short speech segments, followed by post-processing operations to eliminate traces of the forgery, presents a great challenge to forensic detection. In this paper, we propose a robust method for detecting and locating a speech copy-move forgery. We found that pitch and formant can be used as the features representing a voiced speech segment, and these two features are very robust against commonly used post-processing operations. In the proposed algorithm, we first divide the speech recording into voiced speech segments and unvoiced speech segments. We then extract the pitch sequence and the first two formant sequences as the feature set of each voiced speech segment. Dynamic time warping is applied to compute the similarities of each feature set. By comparing the similarities with a threshold, we can detect and locate copy-move forgeries in speech recording. The extensive experiments show that the proposed method is very effective in detecting and locating copy-move forgeries, even on a forged speech segment as short as one voiced speech segment. The proposed method is also robust against several kinds of commonly used post-processing operations and background noise, which highlights the promising potential of the proposed method as a speech copy-move forgery localization tool in practical forensics applications.
Qi Yan 0004, Rui Yang 0006, Jiwu Huang
IEEE Trans. Inf. Forensics Secur.3
2019 Detection of Speech Smoothing on Very Short Clips
abstract
Audio editing software can easily be used to manipulate digital speech for forgery. Smoothing on the tampered boundary is usually performed to eliminate the obvious traces of forgery after tampering. This presents a considerable challenge for the forensic detection of tampered speech because the smoothing model is unknown and the smoothing operation often modifies only several tens of samples with the editing software. In this paper, we propose to apply six filtering models to approximate the smoothing in audio editing software for training the classifier. We analyze the impact of filtering operations on speech signals, especially on differential signals. On the basis of the local variance of the differential signal, we design a simple and yet efficient feature set. Theoretical analysis and extensive experiments show that the proposed features are very effective in detecting several common filtering operations on very short speech clips. The experimental results also show that the proposed method can detect unknown smoothing performed by commonly used audio editing software, such as Cooledit and Adobe Audition. This highlights the promising potential of the proposed method for use as a forgery localization tool of digital speech signals in practical forensic applications. The proposed method is capable of detecting smoothing on very short speech clips containing only several tens of samples and practical forgery used in audio editing software.
Qi Yan 0004, Rui Yang 0006, Jiwu Huang
IEEE Trans. Inf. Forensics Secur.3
2019 WISERNet: Wider Separate-Then-Reunion Network for Steganalysis of Color Images
abstract
Until recently, deep steganalyzers in the spatial domain have been all designed for gray-scale images. In this paper, we propose the wider separate-then-reunion network (WISERNet) for steganalysis of color images. We provide theoretical rationale to claim that the summation in normal convolution is one sort of linear collusion attack which reserves strong correlated patterns while impairs uncorrelated noises. Therefore, in the bottom convolutional layer which aims at suppressing correlated image contents, we adopt separate channel-wise convolution without summation instead. Conversely, in the upper convolutional layers, we believe that the summation in normal convolution is beneficial. Therefore, we adopt united normal convolution in those layers and make them remarkably wider to reinforce the effect of linear collusion attack. As a result, our proposed wide-and-shallow, separate-then-reunion network structure is specifically suitable for color image steganalysis. We have conducted extensive experiments on color image datasets generated from BOSSBase raw images and another large-scale dataset that contains 100, 000 raw images, with different demosaicking algorithms and down-sampling algorithms. The experimental results show that our proposed network outperforms other state-of-the-art color image steganalytic models either hand crafted or learned using deep networks in the literature by a clear margin. Specifically, it is noted that the detection performance gain is achieved with less than half the complexity compared to the most advanced deep-learning steganalyzer as far as we know, which is scarce in the literature.
Jishen Zeng, Shunquan Tan, Guangqing Liu, Bin Li 0011, Jiwu Huang
IEEE Trans. Inf. Forensics Secur.5
2018 Edge Detection and Image Segmentation on Encrypted Image with Homomorphic Encryption and Garbled Circuit
abstract
Edge detection is one of the most important topics of image processing. In the scenario of cloud computing, performing edge detection may also consider privacy protection. In this paper, we propose an edge detection and image segmentation scheme on an encrypted image with Sobel edge detector. We implement Gaussian filtering and Sobel operator on the image in the encrypted domain with homomorphic property. By implementing an adaptive threshold decision algorithm in the encrypted domain, we obtain a threshold determined by the image distribution. With the technique of garbled circuit, we perform comparison in the encrypted domain and obtain the edge of the image without decrypting the image in advanced. We then propose an image segmentation scheme on the encrypted image based on the detected edges. Our experiments demonstrate the viability and effectiveness of the proposed encrypted image edge detection and segmentation.
Delin Chen, Peijia Zheng, Jiwu Huang
ICME5
2018 VPCID - A VoIP Phone Call Identification Database
Yuankun Huang, Shunquan Tan, Bin Li 0011, Jiwu Huang
IWDW4
2018 Verifiable keyword search for secure big data-based mobile healthcare networks with fine-grained authorization control
Zehong Chen, Fangguo Zhang, Peng Zhang 0029, Joseph K. Liu, Jiwu Huang, Hanbang Zhao, Jian Shen 0001
Future Gener. Comput. Syst.5
2018 Riemannian competitive learning for symmetric positive definite matrices clustering
Ligang Zheng, Guoping Qiu, Jiwu Huang
Neurocomputing3
2018 Data-driven multimedia forensics and security
Anderson Rocha 0001, Shujun Li 0001, C.-C. Jay Kuo, Alessandro Piva, Jiwu Huang
J. Vis. Commun. Image Represent.5
2018 Detecting median filtering via two-dimensional AR models of multiple filtered residuals
Jianquan Yang, Honglei Ren, Guopu Zhu, Jiwu Huang, Yun Q. Shi 0001
Multim. Tools Appl.4
2018 A novel reversible data hiding method with image contrast enhancement
Shaohua Tang, Jiwu Huang, Yun Q. Shi 0001
Signal Process. Image Commun.3
2018 Identification of Various Image Operations Using Residual-Based Features
abstract
Image forensics has attracted wide attention during the past decade. However, most existing works aim at detecting a certain operation, which means that their proposed features usually depend on the investigated image operation and they consider only binary classification. This usually leads to misleading results if irrelevant features and/or classifiers are used. For instance, a JPEG decompressed image would be classified as an original or median filtered image if it was fed into a median filtering detector. Hence, it is important to develop forensic methods and universal features that can simultaneously identify multiple image operations. Based on extensive experiments and analysis, we find that any image operation, including existing anti-forensics operations, will inevitably modify a large number of pixel values in the original images. Thus, some common inherent statistics such as the correlations among adjacent pixels cannot be preserved well. To detect such modifications, we try to analyze the properties of local pixels within the image in the residual domain rather than the spatial domain considering the complexity of the image contents. Inspired by image steganalytic methods, we propose a very compact universal feature set and then design a multiclass classification scheme for identifying many common image operations. In our experiments, we tested the proposed features as well as several existing features on 11 typical image processing operations and four kinds of anti-forensic methods. The experimental results show that the proposed strategy significantly outperforms the existing forensic methods in terms of both effectiveness and universality.
Haodong Li 0001, Weiqi Luo 0001, Xiaoqing Qiu, Jiwu Huang
IEEE Trans. Circuits Syst. Video Technol.4
2018 Band Energy Difference for Source Attribution in Audio Forensics
abstract
Digital audio recordings are one of the key types of evidence used in law enforcement proceedings. As a result, the development of reliable techniques for forensic analysis of such recordings is of principal importance. One of the main problems in forensic analysis is source attribution, i.e., verifying whether a certain recording was acquired with a given device. While this problem has been widely studied for other types of multimedia signals, there are a very few techniques for audio recordings. Moreover, reported evaluation results were obtained from extremely small data sets on the order of a dozen devices. The goal of this paper is to propose a new feature set, the band energy difference (BED) descriptor, for source attribution of digital speech recordings. We demonstrate that a frequency response curve extracted from sample recordings can serve as a robust fingerprint that carries significant discriminative power and can characterize the recording device. We study two sub-problems of source attribution: 1) identification of a recording device among a list of possible candidates (device identification) and 2) confirming that a suspected device has indeed been used to acquire the recording in question (device verification). For our evaluation, we prepared two novel data sets: a controlled-conditions data set with 31 devices and an uncontrolled-conditions data set with 141 devices. Our experimental evaluation demonstrates that the proposed BED descriptor is effective for both device identification and verification. In the former task, we reached an accuracy of over 96%. In the latter, we obtained a high true positive rate of 89% while maintaining a fixed low false positive rate of 1%.
Pawel Korus, Jiwu Huang
IEEE Trans. Inf. Forensics Secur.3
2018 Large-Scale JPEG Image Steganalysis Using Hybrid Deep-Learning Framework
abstract
Adoption of deep learning in image steganalysis is still in its initial stage. In this paper, we propose a generic hybrid deep-learning framework for JPEG steganalysis incorporating the domain knowledge behind rich steganalytic models. Our proposed framework involves two main stages. The first stage is hand-crafted, corresponding to the convolution phase and the quantization and truncation phase of the rich models. The second stage is a compound deep-neural network containing multiple deep subnets, in which the model parameters are learned in the training procedure. We provided experimental evidence and theoretical reflections to argue that the introduction of threshold quantizers, though disabling the gradient-descent-based learning of the bottom convolution phase, is indeed cost-effective. We have conducted extensive experiments on a large-scale data set extracted from ImageNet. The primary data set used in our experiments contains 500 000 cover images, while our largest data set contains five million cover images. Our experiments show that the integration of quantization and truncation into deep-learning steganalyzers do boost the detection performance by a clear margin. Furthermore, we demonstrate that our framework is insensitive to JPEG blocking artifact alterations, and the learned model can be easily transferred to a different attacking target and even a different data set. These properties are of critical importance in practical applications.
Jishen Zeng, Shunquan Tan, Bin Li 0011, Jiwu Huang
IEEE Trans. Inf. Forensics Secur.4
2018 Efficient Encrypted Images Filtering and Transform Coding With Walsh-Hadamard Transform and Parallelization
abstract
Since homomorphic encryption operations have high computational complexity, image applications based on homomorphic encryption are often time consuming, which makes them impractical. In this paper, we study efficient encrypted image applications with the encrypted domain Walsh-Hadamard transform (WHT) and parallel algorithms. We first present methods to implement real and complex WHTs in the encrypted domain. We then propose a parallel algorithm to improve the computational efficiency of the encrypted domain WHT. To compare the WHT with the discrete cosine transform (DCT), integer DCT, and Haar transform in the encrypted domain, we conduct theoretical analysis and experimental verification, which reveal that the encrypted domain WHT has the advantages of lower computational complexity and a shorter running time. Our analysis shows that the encrypted WHT can accommodate plaintext data of larger values. We propose two encrypted image applications using the encrypted domain WHT. To accelerate the practical execution, we present two parallelization strategies for the proposed applications. The experimental results show that the speedup of the homomorphic encrypted image application exceeds 12.
Peijia Zheng, Jiwu Huang
IEEE Trans. Image Process.2
2018 JPEG Image Encryption With Improved Format Compatibility and File Size Preservation
abstract
Image encryption techniques can be used to ensure the security and privacy of valuable images. The related works in this field have focused more on raster images than on compressed images. Many existing JPEG image encryption schemes are not quite well compatible with the JPEG standard, or the file size of an encrypted JPEG image is apparently increased. In this paper, a novel bitstream-based JPEG image encryption method is presented. First, the groups of successive DC codes that encode the quantized DC coefficient differences with the same sign are permuted within each group. Second, the left half and the right half of a group, whose size will increase with the number of iterations, of consecutive DC codes may be swapped with each other, depending on whether an overflow of quantized DC coefficients occurs during decoding. Third, all AC codes are classified into 63 categories according to their zero-run lengths, then the AC codes within each category are, respectively, scrambled. Finally, all MCUs, except for DC codes, are randomly shuffled as a whole. Moreover, an image-content-related encryption key is employed to provide further security. The experimental results show that the file size of an encrypted JPEG image is almost the same as that of the corresponding plaintext image except for slight variations because of byte alignment. In addition, the quantized DC coefficients decoded from an encrypted JPEG image will not fall outside the valid range. Improved format compatibility is provided compared with other related methods. Moreover, it is unnecessary to perform entropy encoding again because all of the encryption operations are performed directly on the JPEG bitstream. The proposed method proves to be secure against brute-force attacks, differential cryptanalysis, known plaintext attacks, and outline attacks. Our proposed method can also be applied to color JPEG images.
Shuhao Huang, Shaohua Tang, Jiwu Huang
IEEE Trans. Multim.4
2018 Improved Audio Steganalytic Feature and Its Applications in Audio Forensics
abstract
Digital multimedia steganalysis has attracted wide attention over the past decade. Currently, there are many algorithms for detecting image steganography. However, little research has been devoted to audio steganalysis. Since the statistical properties of image and audio files are quite different, features that are effective in image steganalysis may not be effective for audio. In this article, we design an improved audio steganalytic feature set derived from both the time and Mel-frequency domains for detecting some typical steganography in the time domain, including LSB matching, Hide4PGP, and Steghide. The experiment results, evaluated on different audio sources, including various music and speech clips of different complexity, have shown that the proposed features significantly outperform the existing ones. Moreover, we use the proposed features to detect and further identify some typical audio operations that would probably be used in audio tampering. The extensive experiment results have shown that the proposed features also outperform the related forensic methods, especially when the length of the audio clip is small, such as audio clips with 800 samples. This is very important in real forensic situations.
Weiqi Luo 0001, Haodong Li 0001, Qi Yan 0004, Rui Yang 0006, Jiwu Huang
ACM Trans. Multim. Comput. Commun. Appl.5
2017 VideoSet: A large-scale compressed video quality dataset based on JND measurement
abstract
• A large-scale JND-based coded video quality dataset is presented. • The VideoSet contains 220 5-s sequences in four resolutions coded by H.264/AVC. • The subjective test procedure, JND data cleaning and properties are described. • The significance and implications of the VideoSet are discussed. • This work points out a clear path to data-driven perceptual coding. A new methodology to measure coded image/video quality using the just-noticeable-difference (JND) idea was proposed in Lin et al. (2015). Several small JND-based image/video quality datasets were released by the Media Communications Lab at the University of Southern California in Jin et al. (2016) and Wang et al. (2016) [3]. In this work, we present an effort to build a large-scale JND-based coded video quality dataset. The dataset consists of 220 5-s sequences in four resolutions (i.e., 1920 × 1080 , 1280 × 720 , 960 × 540 and 640 × 360 ). For each of the 880 video clips, we encode it using the H.264/AVC codec with QP = 1 , … , 51 and measure the first three JND points with 30 + subjects. The dataset is called the “VideoSet”, which is an acronym for “Video Subject Evaluation Test (SET)”. This work describes the subjective test procedure, detection and removal of outlying measured data, and the properties of collected JND data. Finally, the significance and implications of the VideoSet to future video coding research and standardization efforts are pointed out. All source/coded video clips as well as measured JND data included in the VideoSet are available to the public in the IEEE DataPort (Wang et al., 2016 [4]).
Haiqiang Wang, Ioannis Katsavounidis, Jiantong Zhou, Jeong-Hoon Park, Shawmin Lei, Xin Zhou 0001, Man-On Pun, Xin Jin 0002, Ronggang Wang, Xu Wang 0006, Yun Zhang 0002, Jiwu Huang, Sam Kwong, C.-C. Jay Kuo
J. Vis. Commun. Image Represent.12
2017 A security watermark scheme used for digital speech forensics
Zhenghui Liu, Jiwu Huang, Xingming Sun, Chuanda Qi
Multim. Tools Appl.2
2017 Tamper recovery algorithm for digital speech signal based on DWT and DCT
Zhenghui Liu, Jiwu Huang, Chuanda Qi
Multim. Tools Appl.3
2017 Automatic Steganographic Distortion Learning Using a Generative Adversarial Network
abstract
Generative adversarial network has shown to effectively generate artificial samples indiscernible from their real counterparts with a united framework of two subnetworks competing against each other. In this letter, we first propose an automatic steganographic distortion learning framework using a generative adversarial network, which is composed of a steganographic generative subnetwork and a steganalytic discriminative subnetwork. Via alternately training these two oppositional subnetworks, our proposed framework can automatically learn embedding change probabilities for every pixel in a given spatial cover image. The learnt embedding change probabilities can then be converted to embedding distortions, which can be adopted in the existing framework of minimal-distortion embedding. Under this framework, the distortion function is directly related to the undetectability against the oppositional evolving steganalyzer. Experimental results show that with adversarial learning, our proposed framework can effectively evolve from nearly naive random ±1 embedding at the beginning to much more advanced content-adaptive embedding which tries to embed secret bits in textural regions. The security performance is also steadily improved with increasing training iterations.
Weixuan Tang 0004, Shunquan Tan, Bin Li 0011, Jiwu Huang
IEEE Signal Process. Lett.4
2017 Data-Driven Feature Characterization Techniques for Laser Printer Attribution
abstract
Laser printer attribution is an increasing problem with several applications, such as pointing out the ownership of crime proofs and authentication of printed documents. However, as commonly proposed methods for this task are based on custom-tailored features, they are limited by modeling assumptions about printing artifacts. In this paper, we explore solutions able to learn discriminant-printing patterns directly from the available data during an investigation, without any further feature engineering, proposing the first approach based on deep learning to laser printer attribution. This allows us to avoid any prior assumption about printing artifacts that characterize each printer, thus highlighting almost invisible and difficult printer footprints generated during the printing process. The proposed approach merges, in a synergistic fashion, convolutional neural networks (CNNs) applied on multiple representations of multiple data. Multiple representations, generated through different pre-processing operations, enable the use of the small and lightweight CNNs whilst the use of multiple data enable the use of aggregation procedures to better determine the provenance of a document. Experimental results show that the proposed method is robust to noisy data and outperforms existing counterparts in the literature for this problem.
Anselmo Ferreira, Luca Bondi, Luca Baroffio, Paolo Bestagini, Jiwu Huang, Jefersson A. dos Santos, Stefano Tubaro, Anderson Rocha 0001
IEEE Trans. Inf. Forensics Secur.5
2017 Multi-Scale Analysis Strategies in PRNU-Based Tampering Localization
abstract
Accurate unsupervised tampering localization is one of the most challenging problems in digital image forensics. In this paper, we consider a photo response non-uniformity analysis and focus on the detection of small forgeries. For this purpose, we adopt a recently proposed paradigm of multi-scale analysis and discuss various strategies for its implementation. First, we consider a multi-scale fusion approach, which involves combination of multiple candidate tampering probability maps into a single, more reliable decision map. The candidate maps are obtained with sliding windows of various sizes and thus allow to exploit the benefits of both the small- and large-scale analyses. We extend this approach by introducing modulated threshold drift and content-dependent neighborhood interactions, leading to improved localization performance with superior shape representation and easier detection of small forgeries. We also discuss two novel alternative strategies: a segmentation-guided approach, which contracts the decision statistic to a central segment within each analysis window and an adaptive-window approach, which dynamically chooses analysis window size for each location in the image. We perform extensive experimental evaluation on both synthetic and realistic forgeries and discuss in detail practical aspects of parameter selection. Our evaluation shows that the multi-scale analysis leads to significant performance improvement compared with the commonly used single-scale approach. The proposed multi-scale fusion strategy delivers stable results with consistent improvement in various test scenarios.
Pawel Korus, Jiwu Huang
IEEE Trans. Inf. Forensics Secur.2
2017 Localization of Diffusion-Based Inpainting in Digital Images
abstract
Image inpainting, an image processing technique for restoring missing or damaged image regions, can be utilized by forgers for removing objects in digital images. Since no obviously perceptible artifacts are left after inpainting, it is necessary to develop methods for detecting the presence of inpainting. In general, there are two main categories of image inpainting techniques: exemplar-based and diffusion-based techniques. Although several methods have been proposed for detecting exemplar-based inpainting, there is still no effective method for detecting diffusion-based inpainting. Usually, the tampered regions manipulated by diffusion-based inpainting techniques are much smaller than those manipulated by exemplar-based ones, presenting more challenges in detecting these regions. As a pioneering attempt, this paper proposes a method for the localization of diffusion-based inpainted regions in digital images. We first analyze the diffusion process in inpainting, and observe that the changes in the image Laplacian along the direction perpendicular to the gradient are different in the inpainted and untouched regions. Following this observation, we construct a feature set based on the intra-channel and inter-channel local variances of the changes to identify the inpainted regions. Finally, two effective post-processing operations are designed for further refining of the localization result. The extensive experimental results evaluated on both synthetic and realistic inpainted images show the effectiveness of the proposed method.
Haodong Li 0001, Weiqi Luo 0001, Jiwu Huang
IEEE Trans. Inf. Forensics Secur.3
2017 Image Forgery Localization via Integrating Tampering Possibility Maps
abstract
Over the past decade, many efforts have been made in passive image forensics. Although it is able to detect tampered images at high accuracies based on some carefully designed mechanisms, localization of the tampered regions in a fake image still presents many challenges, especially when the type of tampering operation is unknown. Some researchers have realized that it is necessary to integrate different forensic approaches in order to obtain better localization performance. However, several important issues have not been comprehensively studied, for example, how to select and improve/readjust proper forensic approaches, and how to fuse the detection results of different forensic approaches to obtain good localization results. In this paper, we propose a framework to improve the performance of forgery localization via integrating tampering possibility maps. In the proposed framework, we first select and improve two existing forensic approaches, i.e., statistical feature-based detector and copy-move forgery detector, and then adjust their results to obtain tampering possibility maps. After investigating the properties of possibility maps and comparing various fusion schemes, we finally propose a simple yet very effective strategy to integrate the tampering possibility maps to obtain the final localization results. The extensive experiments show that the two improved approaches used in our framework significantly outperform the state-of-the-art techniques, and the proposed fusion results achieve the best F1-score in the IEEE IFS-TC Image Forensics Challenge.
Haodong Li 0001, Weiqi Luo 0001, Xiaoqing Qiu, Jiwu Huang
IEEE Trans. Inf. Forensics Secur.4
2017 Detection of Double Compressed AMR Audio Using Stacked Autoencoder
abstract
The adaptive multi-rate (AMR) audio codec adopted by many portable recording devices is widely used in speech compression. The use of AMR speech recordings as evidence in court is growing. Nowadays, it is easy to tamper with digital speech recordings, which makes audio forensics increasingly important. The detection of double compressed audio is one of the key issues in audio forensics. In this paper, we propose a framework for detecting double compressed AMR audio based on the stacked autoencoder (SAE) network and the universal background model-Gaussian mixture model (UBM-GMM). Instead of hand-crafted features, we used the SAE to learn the optimal features automatically from the audio waveforms. Audio frames are used as network input and the last hidden layer's output constitutes the features of a single frame. For an audio clip with many frames, the features of all the frames are aggregated and classified by UBM-GMM. Experimental results show that our method is effective in distinguishing single/double compressed AMR audio and outperforms the existing methods by achieving a detection accuracy of 98% on the TIMIT database. Exhaustive experiments demonstrate the effectiveness and robustness of the proposed method.
Rui Yang 0006, Bin Li 0011, Jiwu Huang
IEEE Trans. Inf. Forensics Secur.4
2017 Pixel-Decimation-Assisted Steganalysis of Synchronize-Embedding-Changes Steganography
abstract
This paper deals with the state-of-the-art synchronize-embedding-changes (SECs) steganography. We propose the pixel-decimation-assisted steganalytic feature set, a novel feature set construction protocol that extends upon the recent selection-channel-aware spatial rich model maxSRMd2. Our method is based on pixel decimation, a specific type of image downsampling. Based on theoretical analysis and empirical evaluation, we clearly demonstrate that our method impairs the synchronization of embedding changes in SEC steganography, and improves the accuracy of embedding change probability estimation. Our method significantly improves stego image detection performance when extended from a selection-channel-aware rich-model feature set (maxSRMd2) and is robust to different image downsampling methods. Furthermore, increasing the number of sweeps in SEC steganography has no effect to the performance of our proposed method even though it further strengthens synchronization of embedding changes. It is worth noting that with ensemble classifier, the above-mentioned performance improvements are achieved at a little extra cost.
Shunquan Tan, Bin Li 0011, Jiwu Huang
IEEE Trans. Inf. Forensics Secur.4
2017 An Efficient Motion Detection and Tracking Scheme for Encrypted Surveillance Videos
abstract
Performing detection on surveillance videos contributes significantly to the goals of safety and security. However, performing detection on unprotected surveillance video may reveal the privacy of innocent people in the video. Therefore, striking a proper balance between maintaining personal privacy while enhancing the feasibility of detection is an important issue. One promising solution to this problem is to encrypt the surveillance videos and perform detection on the encrypted videos. Most existing encrypted signal processing methods focus on still images or small data volumes; however, because videos are typically much larger, investigating how to process encrypted videos is a significant challenge. In this article, we propose an efficient motion detection and tracking scheme for encrypted H.264/AVC video bitstreams, which does not require the previous decryption on the encrypted video. The main idea is to first estimate motion information from the bitstream structure and codeword length and, then, propose a region update (RU) algorithm to deal with the loss and error drifting of motion caused by the video encryption. The RU algorithm is designed based on the prior knowledge that the object motion in the video is continuous in space and time. Compared to the existing scheme, which is based on video encryption that occurs at the pixel level, the proposed scheme has the advantages of requiring only a small storage of the encrypted video and has a low computational cost for both encryption and detection. Experimental results show that our scheme performs better regarding detection accuracy and execution speed. Moreover, the proposed scheme can work with more than one format-compliant video encryption method, provided that the positions of the macroblocks can be extracted from the encrypted video bitstream. Due to the coupling of video stream encryption and detection algorithms, our scheme can be directly connected to the video stream output (e.g., surveillance cameras) without requiring any camera modifications.
Jianting Guo, Peijia Zheng, Jiwu Huang
ACM Trans. Multim. Comput. Commun. Appl.3
2016 Clustering Symmetric Positive Definite Matrices on the Riemannian Manifolds
Ligang Zheng, Guoping Qiu, Jiwu Huang
ACCV (1)3
2016 Reversible data hiding in Paillier cryptosystem
Haotian Wu 0009, Yiu-Ming Cheung, Jiwu Huang
J. Vis. Commun. Image Represent.3
2016 Twenty years of digital audio watermarking - a comprehensive review
abstract
Digital audio watermarking is an important technique to secure and authenticate audio media. This paper provides a comprehensive review of the twenty years’ research and development works for digital audio watermarking, based on an exhaustive literature survey and careful selections of representative solutions. We generally classify the existing designs into time domain and transform domain methods , and relate all the reviewed works using two generic watermark embedding equations in the two domains. The most important designing criteria, i.e., imperceptibility and robustness, are thoroughly reviewed. For imperceptibility , the existing measurement and control approaches are classified into heuristic and analytical types, followed by intensive analysis and discussions. Then, we investigate the robustness of the existing solutions against a wide range of critical attacks categorized into basic, desynchronization, and replacement attacks, respectively. This reveals current challenges in developing a global solution robust against all the attacks considered in this paper. Some remaining problems as well as research potentials for better system designs are also discussed. In addition, audio watermarking applications in terms of US patents and commercialized solutions are reviewed. This paper serves as a comprehensive tutorial for interested readers to gain a historical, technical, and also commercial view of digital audio watermarking.
Guang Hua 0001, Jiwu Huang, Yun Q. Shi 0001, Jonathan Goh, Vrizlynn L. L. Thing
Signal Process.2
2016 Authentication and recovery algorithm for speech signal based on digital watermarking
Zhenghui Liu, Hongxia Wang 0001, Jiwu Huang
Signal Process.5
2016 Robust image watermarking based on Tucker decomposition and Adaptive-Lattice Quantization Index Modulation
Bingwen Feng, Wei Lu 0001, Wei Sun 0007, Jiwu Huang, Yun Q. Shi 0001
Signal Process. Image Commun.4
2016 Improved Tampering Localization in Digital Image Forensics Based on Maximal Entropy Random Walk
abstract
In this paper we propose to use maximal entropy random walk on a graph for tampering localization in digital image forensics. Our approach serves as an additional post-processing step after conventional sliding-window analysis with a forensic detector. Strong localization property of this random walk will highlight important regions and attenuate the background - even for noisy response maps. Our evaluation shows that the proposed method can significantly outperform both the commonly used threshold-based decision, and the recently proposed optimization-based approach with a Markovian prior.
Pawel Korus, Jiwu Huang
IEEE Signal Process. Lett.2
2016 Audio Postprocessing Detection Based on Amplitude Cooccurrence Vector Feature
abstract
Authentication of audio signals is an important problem in multimedia forensics. Tampering is typically followed by postprocessing that aims to remove the traces of the forgery. The variety of possible postprocessing operations makes tampering detection even more challenging. In this letter, we propose the amplitude cooccurrence vector features, which exploit cooccurrence patterns in audio signals. Experimental results show that our proposed features are able to distinguish between the original audio and the postprocessed audio with an average accuracy of above 95%. Furthermore, it can also effectively discriminate different kinds of postprocessing operations.
Jiwu Huang
IEEE Signal Process. Lett.3
2016 Clustering Steganographic Modification Directions for Color Components
abstract
It is conventionally assumed that steganographic schemes for gray-scale images can be directly applied to color images by embedding messages independently in different color channels. However, the correlation among color channels may be disturbed and it is unclear how to preserve the channel correlation so as to increase empirical security. In this paper, we propose a strategy called CMD-C (clustering modification directions for color components). The basic idea of the strategy is to change different color components from the same pixel location towards a positive or negative direction consistently. To implement the strategy, we decompose an image into several sub-images in which segmented hidden message bits are successively embedded. The embedding costs of a sub-image are computed by considering the correlation both within and among color channels. Experimental results show that the proposed CMD-C strategy has made great improvement over conventional methods in resisting state-of-the-art steganalytic methods.
Weixuan Tang 0004, Bin Li 0011, Weiqi Luo 0001, Jiwu Huang
IEEE Signal Process. Lett.4
2016 Automatic Detection of Object-Based Forgery in Advanced Video
abstract
Passive multimedia forensics has become an active topic in recent years. However, less attention has been paid to video forensics. Research on video forensics, and especially on automatic detection of object-based video forgery, is still in its infancy. In this paper, we develop an approach for automatic identification and forged segment localization of object-based forged video encoded with advanced frameworks. The proposed approach starts with a frame manipulation detector. An automatic algorithm is proposed to identify object-based video forgery based on the frame manipulation detector. Then, a two-stage automatic algorithm is provided to accurately locate the forged video segments in the suspicious video. To construct the proposed frame manipulation detector, motion residuals are generated from the target video frame sequence. We regard the object-based forgery in video frames as image tampering in the motion residuals and employ the feature extractors that are originally built for still image steganalysis to extract forensic features from the motion residuals. The experiments show that the proposed approach achieves excellent results in both forged video identification and automatic forged temporal segment localization.
Shengda Chen, Shunquan Tan, Bin Li 0011, Jiwu Huang
IEEE Trans. Circuits Syst. Video Technol.4
2016 Reversible Data Hiding in JPEG Images
abstract
Among various digital image formats used in daily life, the Joint Photographic Experts Group (JPEG) is the most popular. Therefore, reversible data hiding (RDH) in JPEG images is important and useful for many applications such as archive management and image authentication. However, RDH in JPEG images is considerably more difficult than that in uncompressed images because there is less information redundancy in JPEG images than that in uncompressed images, and any modification in the compressed domain may introduce more distortion in the host image. Furthermore, along with the embedding capacity and fidelity (visual quality), which have to be considered for uncompressed images, the storage size of the marked JPEG file should be considered. In this paper, based on the philosophy behind the JPEG encoder and the statistical properties of discrete cosine transform (DCT) coefficients, we present some basic insights into how to select quantized DCT coefficients for RDH. Then, a new histogram shifting-based RDH scheme for JPEG images is proposed, in which the zero coefficients remain unchanged and only coefficients with values 1 and -1 are expanded to carry message bits. Moreover, a block selection strategy based on the number of zero coefficients in each 8 × 8 block is proposed, which can be utilized to adaptively choose DCT coefficients for data hiding. Experimental results demonstrate that by using the proposed method we can easily realize high embedding capacity and good visual quality. The storage size of the host JPEG file can also be well preserved.
Fangjun Huang, Xiaochao Qu, Hyoung Joong Kim, Jiwu Huang
IEEE Trans. Circuits Syst. Video Technol.4
2016 New Framework for Reversible Data Hiding in Encrypted Domain
abstract
In the past more than one decade, hundreds of reversible data hiding (RDH) algorithms have been reported. Via exploring the correlation between the neighboring pixels (or coefficients), extra information can be embedded into the host image reversibly. However, these RDH algorithms cannot be accomplished in encrypted domain directly, since the correlation between the neighboring pixels will disappear after encryption. In order to accomplish RDH in encrypted domain, specific RDH schemes have been designed according to the encryption algorithm utilized. In this paper, we propose a new simple yet effective framework for RDH in encrypted domain. In the proposed framework, the pixels in a plain image are first divided into sub-blocks with the size of $m\times n$ . Then, with an encryption key, a key stream (a stream of random or pseudorandom bits/bytes that are combined with a plaintext message to produce the encrypted message) is generated, and the pixels in the same sub-block are encrypted with the same key stream byte. After the stream encryption, the encrypted $m\times n$ sub-blocks are randomly permutated with a permutation key. Since the correlation between the neighboring pixels in each sub-block can be well preserved in the encrypted domain, most of those previously proposed RDH schemes can be applied to the encrypted image directly. One of the main merits of the proposed framework is that the RDH scheme is independent of the image encryption algorithm. That is, the server manager (or channel administrator) does not need to design a new RDH scheme according to the encryption algorithm that has been conducted by the content owner; instead, he/she can accomplish the data hiding by applying the numerous RDH algorithms previously proposed to the encrypted domain directly.
Fangjun Huang, Jiwu Huang, Yun Q. Shi 0001
IEEE Trans. Inf. Forensics Secur.2
2016 Adaptive Steganalysis Based on Embedding Probabilities of Pixels
abstract
In modern steganography, embedding modifications are highly concentrated on the textural regions within an image, as such regions are difficult to model for steganalysis. Previous studies have shown that compared with non-adaptive strategies, this content adaptive strategy achieves stronger security against existing steganalysis. Based on the experiments and analyses, however, we found that this embedding property would inevitably lead to a large limitation in existing adaptive steganography. That is, it is possible for steganalyzers to estimate the regions that have probably been modified after data hiding. In this paper, we propose an adaptive steganalytic scheme based on embedding probabilities of pixels. The main idea of our scheme is that we assign different weights to different pixels in feature extraction. For those pixels with high embedding probabilities, their corresponding weights are larger, since they should contribute more to steganalysis and vice versa. By doing so, we can concentrate our attention on the regions that have probably been modified and significantly reduce the impact of other unchanged smooth regions. It is expected that our proposed method is an improvement on the existing steganalytic methods, which usually assume every pixel has the same contribution to steganalysis. The extensive experiments evaluated on four typical adaptive steganographic methods have shown the effectiveness of the proposed scheme, especially for low embedding rates, for example, lower than 0.20 bpp.
Weixuan Tang 0004, Haodong Li 0001, Weiqi Luo 0001, Jiwu Huang
IEEE Trans. Inf. Forensics Secur.4
2016 Multi-Scale Fusion for Improved Localization of Malicious Tampering in Digital Images
abstract
A sliding window-based analysis is a prevailing mechanism for tampering localization in passive image authentication. It uses existing forensic detectors, originally designed for a full-frame analysis, to obtain the detection scores for individual image regions. One of the main problems with a window-based analysis is its impractically low localization resolution stemming from the need to use relatively large analysis windows. While decreasing the window size can improve the localization resolution, the classification results tend to become unreliable due to insufficient statistics about the relevant forensic features. In this paper, we investigate a multi-scale analysis approach that fuses multiple candidate tampering maps, resulting from the analysis with different windows, to obtain a single, more reliable tampering map with better localization resolution. We propose three different techniques for multi-scale fusion, and verify their feasibility against various reference strategies. We consider a popular tampering scenario with mode-based first digit features to distinguish between singly and doubly compressed regions. Our results clearly indicate that the proposed fusion strategies can successfully combine the benefits of small-scale and large-scale analyses and improve the tampering localization performance.
Pawel Korus, Jiwu Huang
IEEE Trans. Image Process.2
2016 Identification of Reconstructed Speech
abstract
Both voice conversion and hidden Markov model-- (HMM) based speech synthesis can be used to produce artificial voices of a target speaker. They have shown great negative impacts on speaker verification (SV) systems. In order to enhance the security of SV systems, the techniques to detect converted/synthesized speech should be taken into consideration. During voice conversion and HMM-based synthesis, speech reconstruction is applied to transform a set of acoustic parameters to reconstructed speech. Hence, the identification of reconstructed speech can be used to distinguish converted/synthesized speech from human speech. Several related works on such identification have been reported. The equal error rates (EERs) lower than 5% of detecting reconstructed speech have been achieved. However, through the cross-database evaluations on different speech databases, we find that the EERs of several testing cases are higher than 10%. The robustness of detection algorithms to different speech databases needs to be improved. In this article, we propose an algorithm to identify the reconstructed speech. Three different speech databases and two different reconstruction methods are considered in our work, which has not been addressed in the reported works. The high-dimensional data visualization approach is used to analyze the effect of speech reconstruction on Mel-frequency cepstral coefficients (MFCC) of speech signals. The Gaussian mixture model supervectors of MFCC are used as acoustic features. Furthermore, a set of commonly used classification algorithms are applied to identify reconstructed speech. According to the comparison among different classification methods, linear discriminant analysis-ensemble classifiers are chosen in our algorithm. Extensive experimental results show that the EERs lower than 1% can be achieved by the proposed algorithm in most cases, outperforming the reported state-of-the-art identification techniques.
Haojun Wu, Jiwu Huang
ACM Trans. Multim. Comput. Commun. Appl.3
2015 Copy-move detection of audio recording with pitch similarity
abstract
The widespread availability of audio editing software has made it very easy to create forgeries without perceptual trace. Copy-move is one of popular audio forgeries. It is very important to identify audio recording with duplicated segments. However, copy-move detection in digital audio with sample by sample comparison is invalid due to post-processing after forgeries. In this paper we present a method based on pitch similarity to detect copy-move forgeries. We use a robust pitch tracking method to extract the pitch of every syllable and calculate the similarities of these pitch sequences. Then we can use the similarities to detect copy-move forgeries of digital audio recording. Experimental result shows that our method is feasible and efficient.
Qi Yan 0004, Rui Yang 0006, Jiwu Huang
ICASSP3
2015 Feature Selection for High Dimensional Steganalysis
Yanping Tan, Fangjun Huang, Jiwu Huang
IWDW3
2015 Local pixel patterns
abstract
In this paper, a new class of image texture operators is proposed. We firstly determine that the number of gray levels in each B × B subblock is a fundamental property of the local image texture. Thus, an occurrence histogram for each B × B sub-block can be utilized to describe the texture of the image. Moreover, using a new multi-bit plane strategy, i.e., representing the image texture with the occurrence histogram of the first one or more significant bit-planes of the input image, more powerful operators for describing the image texture can be obtained. The proposed approach is invariant to gray scale variations since the operators are, by definition, invariant under any monotonic transformation of the gray scale, and robust to rotation. They can also be used as supplementary operators to local binary patterns (LBP) to improve their capability to resist illuminance variation, surface transformations, etc.
Fangjun Huang, Xiaochao Qu, Hyoung Joong Kim, Jiwu Huang
Comput. Vis. Media4
2015 Secure watermarking scheme against watermark attacks in the encrypted domain
Jianting Guo, Peijia Zheng, Jiwu Huang
J. Vis. Commun. Image Represent.3
2015 A reversible data hiding method with contrast enhancement for medical images
Jiwu Huang, Yun Q. Shi 0001
J. Vis. Commun. Image Represent.2
2015 Anti-forensics of double JPEG compression with the same quantization matrix
Haodong Li 0001, Weiqi Luo 0001, Jiwu Huang
Multim. Tools Appl.3
2015 Revealing the Trace of High-Quality JPEG Compression Through Quantization Noise Analysis
abstract
To identify whether an image has been JPEG compressed is an important issue in forensic practice. The state-of-the-art methods fail to identify high-quality compressed images, which are common on the Internet. In this paper, we provide a novel quantization noise-based solution to reveal the traces of JPEG compression. Based on the analysis of noises in multiple-cycle JPEG compression, we define a quantity called forward quantization noise. We analytically derive that a decompressed JPEG image has a lower variance of forward quantization noise than its uncompressed counterpart. With the conclusion, we develop a simple yet very effective detection algorithm to identify decompressed JPEG images. We show that our method outperforms the state-of-the-art methods by a large margin especially for high-quality compressed images through extensive experiments on various sources of images. We also demonstrate that the proposed method is robust to small image size and chroma subsampling. The proposed algorithm can be applied in some practical applications, such as Internet image classification and forgery detection.
Bin Li 0011, Tian-Tsong Ng, Xiaolong Li 0001, Shunquan Tan, Jiwu Huang
IEEE Trans. Inf. Forensics Secur.5
2015 A Strategy of Clustering Modification Directions in Spatial Image Steganography
abstract
Most of the recently proposed steganographic schemes are based on minimizing an additive distortion function defined as the sum of embedding costs for individual pixels. In such an approach, mutual embedding impacts are often ignored. In this paper, we present an approach that can exploit the interactions among embedding changes in order to reduce the risk of detection by steganalysis. It employs a novel strategy, called clustering modification directions (CMDs), based on the assumption that when embedding modifications in heavily textured regions are locally heading toward the same direction, the steganographic security might be improved. To implement the strategy, a cover image is decomposed into several subimages, in which message segments are embedded with well-known schemes using additive distortion functions. The costs of pixels are updated dynamically to take mutual embedding impacts into account. Specifically, when neighboring pixels are changed toward a positive/negative direction, the cost of the considered pixel is biased toward the same direction. Experimental results show that our proposed CMD strategy, incorporated into existing steganographic schemes, can effectively overcome the challenges posed by the modern steganalyzers with high-dimensional features.
Bin Li 0011, Xiaolong Li 0001, Shunquan Tan, Jiwu Huang
IEEE Trans. Inf. Forensics Secur.5
2015 Statistical Model of JPEG Noises and Its Application in Quantization Step Estimation
abstract
In this paper, we present a statistical analysis of JPEG noises, including the quantization noise and the rounding noise during a JPEG compression cycle. The JPEG noises in the first compression cycle have been well studied; however, so far less attention has been paid on the statistical model of JPEG noises in higher compression cycles. Our analysis reveals that the noise distributions in higher compression cycles are different from those in the first compression cycle, and they are dependent on the quantization parameters used between two successive cycles. To demonstrate the benefits from the analysis, we apply the statistical model in JPEG quantization step estimation. We construct a sufficient statistic by exploiting the derived noise distributions, and justify that the statistic has several special properties to reveal the ground-truth quantization step. Experimental results demonstrate that the proposed estimator can uncover JPEG compression history with a satisfactory performance.
Bin Li 0011, Tian-Tsong Ng, Xiaolong Li 0001, Shunquan Tan, Jiwu Huang
IEEE Trans. Image Process.5
2014 Detecting double compressed AMR audio using deep learning
abstract
The Adaptive Multi-Rate (AMR) audio codec is a widely used audio data compression scheme optimized for speech and adopted by many devices. With the audio editing software, it is easy to perform tampering on digital speech recording, which makes the audio forensics become an important and urgent issue. Usually, the tampered AMR audio is double compressed AMR audio. In this paper, we proposed a method to detect the double compressed AMR audio. Such technique may be served as a tool for authenticating the originality of audio recordings and detecting the forgery positions. Our proposed method is based on deep learning algorithm and a majority voting strategy is designed for decision. The experimental results show that our method is effective to detect the double compressed AMR audio. Besides, the potential application of this technique is also discussed.
Rui Yang 0006, Jiwu Huang
ICASSP3
2014 Fast and accurate Nearest Neighbor search in the manifolds of symmetric positive definite matrices
abstract
In this paper, we present a fast and accurate Nearest Neighbor (NN) search method in the Riemannian manifolds formed by a kind of structured data - symmetric positive definite (SPD) matrices. We use an ensemble of vocabulary trees based on hierarchical k-means clustering and query these trees to find the NN candidates in sub-linear time. As generating these vocabulary trees with widely used affine-invariant Riemannian metric (AIRM) will be very time-demanding, we propose to use the second-order approximation to AIRM (SOA-AIRM). We evaluate the proposed NN search algorithm in the application scenario of near-duplicate image detection in a large database. Experimental results demonstrate that the proposed method significantly outperforms state of the art techniques in terms of both accuracy and speed.
Ligang Zheng, Guoping Qiu, Jiwu Huang, Jiang Duan
ICASSP3
2014 A new cost function for spatial image steganography
abstract
A well defined cost function is crucial to steganography under the scenario of minimizing embedding distortion. In this paper, we present a new cost function for spatial image steganography. The proposed cost function is designed by using a high-pass filter to locate the less predictable parts in an image, and then using two low-pass filters to make the low cost values more clustered. Experiments show that the steganographic method with the proposed cost function makes the embedding changes more concentrated in texture regions, and thus achieves a better performance on resisting the state-of-the-art steganalysis over prior works, including HUGO, WOW, and S-UNIWARD.
Bin Li 0011, Jiwu Huang, Xiaolong Li 0001
ICIP3
2014 Improved steganalysis algorithm against motion vector based video steganography
abstract
This paper proposes an improved steganalysis algorithm to detect the secret message hidden in the compressed video. As majority of video steganographic algorithms modify motion vectors (MV) in inter-frame encoding to hide data, aliasing effect may be caused in the distribution of the difference between MVs in two adjacent macroblocks. This phenomenon has been observed in detecting the secret data that were added to the MVs in cover video. To exploit the correlations between the neighboring MVs so as to detect the hidden data more efficiently, we consider the joint distribution of the MV differences between one macroblock and the other two macroblocks neighboring to it. The calculated joint probability mass functions are used to distinguish the stego videos from the non-stego ones. The experimental results show that significant improvement in detection accuracy can be made by using the joint distribution of MV differences instead of the statistics calculated from two neighboring MVs as features.
Yuan Liu 0021, Jiwu Huang
ICIP3
2014 Anti-forensics of JPEG Detectors via Adaptive Quantization Table Replacement
abstract
Due to the popularity of JPEG compression standard, JPEG images have been widely used in various applications. Nowadays, detection of JPEG forgeries becomes an important issue in digital image forensics, and lots of related works have been reported. However, most existing works mainly rely on a pre-trained classifier according to the quantization table shown in the file header of the suspicious JPEG image, and they assume that such a table is authentic. This assumption leaves a potential flaw for those wise forgers to confuse or even invalidate the current JPEG forensic detectors. Based on our analysis and experiments, we found that the generalization ability of most current JPEG forensic detectors is not very good. If the quantization table changes, their performances would decrease significantly. Based on this observation, we propose a universal anti-forensic scheme via replacing the quantization table adaptively. The extensive experimental results evaluated on 10,000 natural images have shown the effectiveness of the proposed scheme for confusing four typical JPEG forensic works.
Haodong Li 0001, Weiqi Luo 0001, Rui Yang 0006, Jiwu Huang
ICPR5
2014 A universal image forensic strategy based on steganalytic model
abstract
Image forensics have made great progress during the past decade. However, almost all existing forensic methods can be regarded as the specific way, since they mainly focus on detecting one type of image processing operations. When the type of operations changes, the performances of the forensic methods usually degrade significantly. In this paper, we propose a universal forensics strategy based on steganalytic model. By analyzing the similarity between steganography and image processing operation, we find that almost all image operations have to modify many image pixels without considering some inherent properties within the original image, which is similar to what in steganography. Therefore, it is reasonable to model various image processing operations as steganography and it is promising to detect them with the help of some effective universal steganalytic features. In our experiments, we evaluate several advanced steganalytic features on six kinds of typical image processing operations. The experimental results show that all evaluated steganalyzers perform well while some steganalytic methods such as the spatial rich model (SRM) [4] and LBP [19] based methods even outperform the specific forensic methods significantly. What is more, they can further identify the type of various image processing operations, which is impossible to achieve using the existing forensic methods.
Xiaoqing Qiu, Haodong Li 0001, Weiqi Luo 0001, Jiwu Huang
IH&MMSec4
2014 Adaptive steganalysis against WOW embedding algorithm
abstract
WOW (Wavelet Obtained Weights) [5] is one of the advanced steganographic methods in spatial domain, which can adaptively embed secret message into cover image according to textural complexity. Usually, the more complex of an image region, the more pixel values within it would be modified. In such a way, it can achieve good visual quality of the resulting stegos and high security against typical steganalytic detectors. Based on our analysis, however, we point out one of the limitations in the WOW embedding algorithm, namely, it is easy to narrow down those possible modified regions for a given stego image based on the embedding costs used in WOW. If we just extract features from such regions and perform analysis on them, it is expected that the detection performance would be improved compared with that of extracting steganalytic features from the whole image. In this paper, we first proposed an adaptive steganalytic scheme for the WOW method, and use the spatial rich model (SRM) based features [4] to model those possible modified regions in our experiments. The experimental results evaluated on 10,000 images have shown the effectiveness of our scheme. It is also noted that our steganalytic strategy can be combined with other steganalytic features to detect the WOW and/or other adaptive steganographic methods both in the spatial and JPEG domains.
Weixuan Tang 0004, Haodong Li 0001, Weiqi Luo 0001, Jiwu Huang
IH&MMSec4
2014 Structure Aware Visual Cryptography
abstract
Abstract Visual cryptography is an encryption technique that hides a secret image by distributing it between some shared images made up of seemingly random black‐and‐white pixels. Extended visual cryptography (EVC) goes further in that the shared images instead represent meaningful binary pictures. The original approach to EVC suffered from low contrast, so later papers considered how to improve the visual quality of the results by enhancing contrast of the shared images. This work further improves the appearance of the shared images by preserving edge structures within them using a framework of dithering followed by a detail recovery operation. We are also careful to suppress noise in smooth areas.
Ralph R. Martin, Jiwu Huang, Shi-Min Hu 0001
Comput. Graph. Forum3
2014 A framework for identifying shifted double JPEG compression artifacts with application to non-intrusive digital image forensics
Zhenhua Qu, Weiqi Luo 0001, Jiwu Huang
Sci. China Inf. Sci.3
2014 Detecting video frame-rate up-conversion based on periodic properties of inter-frame similarity
Shan Bian, Weiqi Luo 0001, Jiwu Huang
Multim. Tools Appl.3
2014 Geometric invariant features in the Radon transform domain for near-duplicate image detection
Yanqiang Lei, Ligang Zheng, Jiwu Huang
Pattern Recognit.3
2014 Exposing Fake Bit Rate Videos and Estimating Original Bit Rates
abstract
Bit rate is one of the important criterions for digital video quality. With some video tools, however, video bit rate can be easily increased without improving the video quality at all. In such a case, a claimed high bit rate video would actually have poor visual quality if it is up-converted from an original lower bit rate version. Therefore, exposing fake bit rate videos becomes an important issue for digital video forensics. To the best of our knowledge, although some methods have been proposed for exposing fake bit rate MPEG-2 videos, no relative work has been reported to further estimate their original bit rates. In this paper, we first analyze the statistical artifacts of these fake bit rate videos, including the requantization artifacts based on the first-digit law in the DCT frequency domain (12-D) and the changes of the structural similarity indexes between the query video and its sequential bit rate down-converted versions in the spatial domain (4-D), and then we propose a compact yet very effective 16-D feature vector for exposing fake bit rate videos and further estimating their original bit rates. The extensive experiments evaluated on hundreds of video sequences with four different resolutions and two typical compression schemes (i.e., MPEG-2 and H.264/AVC) have shown the effectiveness of the proposed method compared with the existing relative ones.
Shan Bian, Weiqi Luo 0001, Jiwu Huang
IEEE Trans. Circuits Syst. Video Technol.3
2014 Investigation on Cost Assignment in Spatial Image Steganography
abstract
Relating the embedding cost in a distortion function to statistical detectability is an open vital problem in modern steganography. In this paper, we take one step forward by formulating the process of cost assignment into two phases: 1) determining a priority profile and 2) specifying a cost-value distribution. We analytically show that the cost-value distribution determines the change rate of cover elements. Furthermore, when the cost-values are specified to follow a uniform distribution, the change rate has a linear relation with the payload, which is a rare property for content-adaptive steganography. In addition, we propose some rules for ranking the priority profile for spatial images. Following such rules, we propose a five-step cost assignment scheme. Previous steganographic schemes, such as HUGO, WOW, S-UNIWARD, and MG, can be integrated into our scheme. Experimental results demonstrate that the proposed scheme is capable of better resisting steganalysis equipped with high-dimensional rich model features.
Bin Li 0011, Shunquan Tan, Jiwu Huang
IEEE Trans. Inf. Forensics Secur.4
2014 Identification of Electronic Disguised Voices
abstract
Since voice disguise has been an increasing tendency in illegal applications, and it has a great negative impact on establishing the authenticity of audio evidence for audio forensics, it is important to be able to identify whether a suspected voice has been disguised or not. However, few studies on such identification have been reported. In this paper, we propose an algorithm to identify electronic disguised voices. Since voice disguise, in essence, the modification of the frequency spectrum of speech signals, and mel-frequency cepstrum coefficients (MFCCs) can be used to well describe frequency spectral properties, MFCC-based features are supposed to be effective for the identification of disguised voices. In this paper, MFCC statistical moments including mean values and correlation coefficients are extracted as acoustic features. Then, an algorithm based on the extracted features and support vector machine classifiers is proposed to separate disguised voices from original voices. Extensive experiments show that the detection rates higher than 90% of the voices from various speech databases and disguised by various methods can be achieved, indicating that the identification performance of this algorithm is remarkable.
Haojun Wu, Jiwu Huang
IEEE Trans. Inf. Forensics Secur.3
2014 Fast Near-Duplicate Image Detection Using Uniform Randomized Trees
abstract
Indexing structure plays an important role in the application of fast near-duplicate image detection, since it can narrow down the search space. In this article, we develop a cluster of uniform randomized trees (URTs) as an efficient indexing structure to perform fast near-duplicate image detection. The main contribution in this article is that we introduce “uniformity” and “randomness” into the indexing construction. The uniformity requires classifying the object images into the same scale subsets. Such a decision makes good use of the two facts in near-duplicate image detection, namely: (1) the number of categories is huge; (2) a single category usually contains only a small number of images. Therefore, the uniform distribution is very beneficial to narrow down the search space and does not significantly degrade the detection accuracy. The randomness is embedded into the generation of feature subspace and projection direction, improveing the flexibility of indexing construction. The experimental results show that the proposed method is more efficient than the popular locality-sensitive hashing and more stable and flexible than the traditional KD-tree.
Yanqiang Lei, Guoping Qiu, Ligang Zheng, Jiwu Huang
ACM Trans. Multim. Comput. Commun. Appl.4
2014 Identifying Compression History of Wave Audio and Its Applications
abstract
Audio signal is sometimes stored and/or processed in WAV (waveform) format without any knowledge of its previous compression operations. To perform some subsequent processing, such as digital audio forensics, audio enhancement and blind audio quality assessment, it is necessary to identify its compression history. In this article, we will investigate how to identify a decompressed wave audio that went through one of three popular compression schemes, including MP3, WMA (windows media audio) and AAC (advanced audio coding). By analyzing the corresponding frequency coefficients, including modified discrete cosine transform (MDCT) and Mel-frequency cepstral coefficients (MFCCs), of those original audio clips and their decompressed versions with different compression schemes and bit rates, we propose several statistics to identify the compression scheme as well as the corresponding bit rate previously used for a given WAV signal. The experimental results evaluated on 8,800 audio clips with various contents have shown the effectiveness of the proposed method. In addition, some potential applications of the proposed method are discussed.
Weiqi Luo 0001, Rui Yang 0006, Jiwu Huang
ACM Trans. Multim. Comput. Commun. Appl.4
2013 Forensic sensor pattern noise extraction from large image data set
abstract
The sensor pattern noise (SPN) can be regarded as the unique identity of a digital camera which is highly useful in digital image forensics [1, 2]. Existing methods [1, 2] which works by denoising each individual natural image often took an investigator a long time and great efforts to collect sufficient photos of diversified enough natural scenes. These processes are hard to repeat or standardized for officially using by an authority. In this work, we create noise image data set by taking photos of random noises displayed on a high definition monitor and propose a homomorphic based SPN extraction method. It offers the forensic researcher a fast way to create a large image data set in a few minutes. And the extraction method only needs to denoise once, which is highly efficient to deal with large numbers of photos. We compared the source camera identification performance of the proposed SPN extraction method to a prior state-of-art with identical experimental settings. The experimental results confirm the effectiveness of the proposed method.
Zhenhua Qu, Xiangui Kang, Jiwu Huang, Yinxiang Li
ICASSP3
2013 Blind detection of electronic disguised voice
abstract
Since voice disguise has great negative impact on establishing authenticity of audio evidence in forensics, and has shown an increasing tendency in illegal applications, it is important to identify whether a suspected voice has been disguised or not. However, research on such detection has not been reported. In this paper, we focus on blind detection of electronic disguised voice. Statistical moments of Mel-frequency cepstrum coefficients (MFCC) are extracted as acoustic features of speech signals. Then an approach for detection of disguised voice based on the extracted features and Support Vector Machine (SVM) classifiers is proposed. The extensive experiments demonstrate that detection rates higher than 95% can be achieved, indicating that detection performance of the proposed approach is good.
Haojun Wu, Jiwu Huang
ICASSP3
2013 Exposing fake bitrate video and its original bitrate
abstract
Video bitrate, as one of the important factors that reflect the video quality, can be easily manipulated via some video editing softwares. In some forensic scenarios, for example, video uploaders of video-sharing websites may increase video bitrate for seeking more commercial profits. In this paper, we try to detect those fake high bitrate videos, and then further to estimate their original bitrates. The proposed method is mainly based on the fact that if the video bitrate has been increased with the help of video editing software, its essential video quality will not increase at all. By analyzing the quality of the questionable video and a series of its re-encoded versions with different lower bitrates, we can obtain a feature curve to measure the change of the video quality, and then we propose a compact feature vector (3-D) to expose fake bitrate videos and their original bitrates. The experimental results evaluated on both CIF and QCIF raw sequences have shown the effectiveness of the proposed method.
Shan Bian, Weiqi Luo 0001, Jiwu Huang
ICIP3
2013 Mixed-strategy Nash equilibrium in the camera source identification game
abstract
Although sensor pattern noise (SPN) is recognized as a reliable device fingerprint for camera source identification (CSI), this fingerprint could have been forged by anti-forensics. In order to evaluate the performance in the case of both forensic investigator and forger exist, we model this interplay as a camera source identification game. The mixed-strategy Nash equilibrium is introduced to solve this game. The Nash equilibrium receiver operating characteristic (ROC) curves are obtained experimentally. Through our analysis, we are able to determine under which case the CSI result is reliable.
Hui Zeng 0002, Xiangui Kang, Jiwu Huang
ICIP3
2013 Distortion function designing for JPEG steganography with uncompressed side-image
abstract
In this paper, we present a new framework for designing distortion functions of joint photographic experts group (JPEG) steganography with uncompressed side-image. In our framework, the discrete cosine transform (DCT) coefficients, including all direct current (DC) coefficients and alternating current (AC) coefficients, are divided into two groups: first-priority group (FPG) and second-priority group (SPG). Different strategies are established to associate the distortion values to the coefficients in FPG and SPG, respectively. In this paper, three scenarios for dividing the coefficients into FPG and SPG are exemplified, which can be utilized to form a series of new distortion functions. Experimental results demonstrate that while applying these generated distortion functions to JPEG steganography, the intrinsic statistical characteristics of the carrier image will be preserved better than the prior-art, and consequently the security performance of the corresponding JPEG steganography can be improved significantly.
Fangjun Huang, Weiqi Luo 0001, Jiwu Huang, Yun Q. Shi 0001
IH&MMSec3
2013 Improved Algorithm of Edge Adaptive Image Steganography Based on LSB Matching Revisited Algorithm
Fangjun Huang, Yane Zhong, Jiwu Huang
IWDW3
2013 An efficient image homomorphic encryption scheme with small ciphertext expansion
abstract
The field of image processing in the encrypted domain has been given increasing attention for the extensive potential applications, for example, providing efficient and secure solutions for privacy-preserving applications in untrusted environment. One obstacle to the widespread use of these techniques is the ciphertext expansion of high orders of magnitude caused by the existing homomorphic encryptions. In this paper, we provide a way to tackle this issue for image processing in the encrypted domain. By using characteristics of image format, we develop an image encryption scheme to limit ciphertext expansion while preserving the homomorphic property. The proposed encryption scheme first encrypts image pixels with an existing probabilistic homomorphic cryptosystem, and then compresses the whole encrypted image in order to save storage space. Our scheme has a much smaller ciphertext expansion factor compared with the element-wise encryption scheme, while preserving the homomorphic property. It is not necessary to require additional interactive protocols when applying secure signal processing tools to the compressed encrypted image. We present a fast algorithm for the encryption and the compression of the proposed image encryption scheme, which speeds up the computation and makes our scheme much more efficient. The analysis on the security, ciphertext expansion ratio, and computational complexity are also conducted. Our experiments demonstrate the validity of the proposed algorithms. The proposed scheme is suitable to be employed as an image encryption method for the applications in secure image processing.
Peijia Zheng, Jiwu Huang
ACM Multimedia2
2013 Region duplication detection based on Harris corner points and step sector statistics
Likai Chen, Wei Lu 0001, Jiangqun Ni, Wei Sun 0007, Jiwu Huang
J. Vis. Commun. Image Represent.5
2013 Blind Detection of Median Filtering in Digital Images: A Difference Domain Based Approach
abstract
Recently, the median filtering (MF) detector as a forensic tool for the recovery of images' processing history has attracted wide interest. This paper presents a novel method for the blind detection of MF in digital images. Following some strongly indicative analyses in the difference domain of images, we introduce two new feature sets that allow us to distinguish a median-filtered image from an untouched image or average-filtered one. The effectiveness of the proposed features is verified with evidence from exhaustive experiments on a large composite image database. Compared with prior arts, the proposed method achieves significant performance improvement in the case of low resolution and strong JPEG post-compression. In addition, it is demonstrated that our method is more robust against additive noise than other existing MF detectors. With analyses and extensive experimental researches presented in this paper, we hope that the proposed method will add a new tool to the arsenal of forensic analysts.
Chenglong Chen, Jiangqun Ni, Jiwu Huang
IEEE Trans. Image Process.3
2013 Discrete Wavelet Transform and Data Expansion Reduction in Homomorphic Encrypted Domain
abstract
Signal processing in the encrypted domain is a new technology with the goal of protecting valuable signals from insecure signal processing. In this paper, we propose a method for implementing discrete wavelet transform (DWT) and multiresolution analysis (MRA) in homomorphic encrypted domain. We first suggest a framework for performing DWT and inverse DWT (IDWT) in the encrypted domain, then conduct an analysis of data expansion and quantization errors under the framework. To solve the problem of data expansion, which may be very important in practical applications, we present a method for reducing data expansion in the case that both DWT and IDWT are performed. With the proposed method, multilevel DWT/IDWT can be performed with less data expansion in homomorphic encrypted domain. We propose a new signal processing procedure, where the multiplicative inverse method is employed as the last step to limit the data expansion. Taking a 2-D Haar wavelet transform as an example, we conduct a few experiments to demonstrate the advantages of our method in secure image processing. We also provide computational complexity analyses and comparisons. To the best of our knowledge, there has been no report on the implementation of DWT and MRA in the encrypted domain.
Peijia Zheng, Jiwu Huang
IEEE Trans. Image Process.2
2012 Identifying Shifted Double JPEG Compression Artifacts for Non-intrusive Digital Image Forensics
Zhenhua Qu, Weiqi Luo 0001, Jiwu Huang
CVM3
2012 Compression history identification for digital audio signal
abstract
Compression history identification plays a very important role in digital multimedia forensics. However, most existing literatures mainly focus on digital image forensics, and just a few works consider digital audio. In this paper, we investigate two popular compression schemes in digital audio, that is, MP3 and WMA, and try to reveal the compression history for a questionable audio signal in the original uncompressed WAV format via analyzing some statistical characteristics of the modified discrete cosine transform coefficients of the audio. The extensive experimental results have shown that the proposed method can effectively identify whether the given audio has been previously compressed with MP3 and/or WMA, and can further estimate the hidden compression rates, even the compression rate is as high as 128 K bps (bits per second).
Weiqi Luo 0001, Rui Yang 0006, Jiwu Huang
ICASSP4
2012 Efficient coarse-to-fine near-duplicate image detection in riemannian manifold
abstract
This paper presents an efficient coarse-to-fine strategy for near duplicate image detection in a Riemannian space. At the coarse level, we use the faster but less accurate log-Euclidean Riemannian metric to search the entire database to retrieve a subset of the images that are likely to contain the near duplicates of the querying image; and at the fine level, we use the more accurate but computationally more demanding affine-invariant Riemannian metric to search the coarse level results to accurately identify near-duplicates. We present experimental results to show that the new coarse to fine strategy can be over 20 times faster than existing techniques using affine-invariant Riemannian metric without sacrificing accuracy.
Ligang Zheng, Guoping Qiu, Jiwu Huang
ICASSP3
2012 Countering anti-JPEG compression forensics
abstract
The quantization artifacts and blocking artifacts are the two significant properties in the JPEG compressed images. Most relative forensic techniques usually use such inherent properties to provide some evidences on how image data is acquired and/or processed. A wise attacker, however, may perform some post-operations to confuse the two artifacts to fool current forensic techniques. Recently, Stamm et al. in [1] propose a novel anti-JPEG compression method via adding anti-forensic dither to the DCT coefficients and further reducing the blocking artifacts. In this paper, we found that the dithering operation will inevitably destroy the statistical correlations among the 8 × 8 intrablock and interblock within an image. In the view of JPEG steganalysis, we employ the transition probability matrix of the DCT coefficients to measure such modifications for identifying the forged images from those original JPEG decompressed images and uncompressed ones. On average, we can obtain a detection accuracy as high as 99% on the image database of UCID [2].
Haodong Li 0001, Weiqi Luo 0001, Jiwu Huang
ICIP3
2012 Blind Detection of Electronic Voice Transformation with Natural Disguise
Yanhong Deng, Haojun Wu, Jiwu Huang
IWDW4
2012 Perceptual video hashing robust against geometric distortions
Shijun Xiang, Jianquan Yang, Jiwu Huang
Sci. China Inf. Sci.3
2012 Digital image splicing detection based on Markov features in DCT and DWT domain
Zhongwei He, Wei Lu 0001, Wei Sun 0007, Jiwu Huang
Pattern Recognit.4
2012 Reversible image watermarking on prediction errors by efficient histogram modification
Jiwu Huang
Signal Process.2
2012 Video Sequence Matching Based on the Invariance of Color Correlation
abstract
Video sequence matching aims to locate a query video clip in a video database. It plays an important role in reducing storage redundancy and detecting video copies for copyright protection. In this paper, we propose an effective method for video sequence matching based on the invariance of color correlation. The proposed method first splits each key-frame into nonoverlapping blocks. For each block, we sort the red, green, and blue color components according to their average intensities, and use the percentage of the color correlation to generate a frame feature with a small size. Finally, the resulting video feature is made up of the consecutive frame features, which is demonstrated to be robust against most typical video content-preserving operations, including geometric distortion, blurring, noise contamination, contrast enhancement, and strong re-encoding. The experimental results show that the proposed method outperforms the existing methods in the literature, as well as the method based on the traditional color histogram. Furthermore, the time and space complexity of our algorithm are both satisfactory, which are very important for many real-time applications.
Yanqiang Lei, Weiqi Luo 0001, Yuan-Gen Wang, Jiwu Huang
IEEE Trans. Circuits Syst. Video Technol.4
2012 Controllable Secure Watermarking Technique for Tradeoff Between Robustness and Security
abstract
The circular watermarking (CW) technique has attracted increasing attention because it can resist the estimation of secret carriers in the watermarked only attack (WOA) framework. However, the existing CW schemes are not applicable whenever a malicious watermark removal attack can take place. This is because they either have low security because the attacker can disclose the embedding subspace or have low robustness. Based on an existing CW scheme called transportation natural watermarking (TNW), this correspondence presents a new CW technique for the tradeoff between robustness and security, which we refer to as controllable secure watermarking (CSW). The idea behind the CSW is that by altering the host signal in the orthogonal complement of the embedding subspace, we can make the watermarked signal have an orthogonally invariant distribution in a higher dimensional subspace including the embedding subspace. Orthogonally invariant distribution essentially requires that the distribution does not change if multiplied by any freely chosen orthogonal matrix, and the higher dimensional subspace is referred to as invariant subspace. We prove that the attacker can only reduce the uncertainty of secret carriers up to the invariant subspace. The dimension of the invariant subspace can be used for the tradeoff between robustness and security. Further, the experiment results show that the robustness-security tradeoff provided by the CSW is efficient. In particular, with the increase of the dimension of the invariant subspace, the security of the CSW will increase quickly while its robustness will only decrease slowly.
Jiwu Huang
IEEE Trans. Inf. Forensics Secur.2
2012 New Channel Selection Rule for JPEG Steganography
abstract
In this paper, we present a new channel selection rule for joint photographic experts group (JPEG) steganography, which can be utilized to find the discrete cosine transform (DCT) coefficients that may introduce minimal detectable distortion for data hiding. Three factors are considered in our proposed channel selection rule, i.e., the perturbation error (PE), the quantization step (QS), and the magnitude of quantized DCT coefficient to be modified (MQ). Experimental results demonstrate that higher security performance can be obtained in JPEG steganography via our new channel selection rule.
Fangjun Huang, Jiwu Huang, Yun Q. Shi 0001
IEEE Trans. Inf. Forensics Secur.2
2012 Enhancing Source Camera Identification Performance With a Camera Reference Phase Sensor Pattern Noise
abstract
Sensor pattern noise (SPN) extracted from digital images has been proved to be a unique fingerprint of digital cameras. However, SPN can be contaminated largely in the frequency domain by image content and nonunique artefacts of JPEG compression, on-sensor signal transfer, sensor design, color interpolation. The source camera identification (CI) performance based on SPN needs to be improved for small sizes of images and especially in resisting JPEG compression. Because the SPN is modelled as an additive white Gaussian noise (AWGN) in its extraction process from an image, it is reasonable to assume the camera reference SPN to be a white noise signal in order to remove the interference mentioned above. The noise residues (SPN) extracted from the original images are whitened first, then they are averaged to generate the camera reference SPN. Motivated by Goljan 's test statistic peak to correlation energy (PCE), we propose to use correlation to circular correlation norm (CCN) as the test statistic, which can lower the false positive rate to be a half of that with PCE. Theoretical analysis shows that the proposed CI method can remove the interference and raise the CCN value of a positive sample and thus achieve greater CI performance, CCN values of the negative sample class with the proposed method follow the normal distributionN(0,1) and the false positive rate can be calculated. Compared with the existing state of the art on seven cameras, 1400 photos totally (200 for each camera), the experimental results show that the proposed CI method achieves the best receiver operating characteristic (ROC) performance among all CI methods in all cases and especially achieves much better resistance to JPEG compression than all of the existing state-of-the-art CI methods.
Xiangui Kang, Yinxiang Li, Zhenhua Qu, Jiwu Huang
IEEE Trans. Inf. Forensics Secur.4
2012 An Informed Watermarking Scheme Using Hidden Markov Model in the Wavelet Domain
abstract
Achieving robustness, imperceptibility and high capacity simultaneously is of great importance in digital watermarking. This paper presents a new informed image watermarking scheme with high robustness and simplified complexity at an information rate of 1/64 bit/pixel. Firstly, a Taylor series approximated locally optimum test (TLOT) detector based on the hidden Markov model (HMM) in the wavelet domain is developed to tackle the problem of unavailability of exact embedding strength in the receiver due to informed embedding. Then based on the TLOT detector and the concept of dirty-paper code design, new HMM-based spherical codes are constructed to provide an effective tradeoff between robustness and distortion. The process of informed embedding is formulated as an optimization problem under the robustness and distortion constraints and the genetic algorithm (GA) is then employed to solve this problem. Moreover, the perceptual distance in the wavelet domain is also developed and incorporated into the GA-based optimization. Simulation results demonstrate that the proposed informed watermarking algorithm has high robustness against common attacks in signal processing and shows a comparable performance to the state-of-the-art scheme with a greatly reduced arithmetic complexity.
Chuntao Wang, Jiangqun Ni, Jiwu Huang
IEEE Trans. Inf. Forensics Secur.3
2012 Near-Duplicate Image Detection in a Visually Salient Riemannian Space
abstract
This paper presents a framework for near-duplicate image detection in a visually salient Riemannian space. A visual saliency model is first used to identify salient regions of the image and then the salient region covariance matrix (SCOV) of various image features is computed. SCOV, which lies in a Riemannian manifold, is used as a robust and compact image content descriptor. An efficient coarse-to-fine Riemannian (CTOFR) image search strategy has been developed to improve efficiency while maintaining accuracy. CTOFR first uses a computationally fast but less accurate log-Euclidean Riemannian metric to do a coarse level search of the entire database and retrieve a subset of likely targets and then uses a computationally expensive but more accurate affine-invariant Riemannian metric to search the returns from the coarse search. We present experimental results to demonstrate that SCOV is a very compact, robust, and discriminative descriptor which is competitive to other state-of-the-art descriptors for near-duplicate image and video detection. We show that CTOFR can yield significant speedups over traditional full search methods without sacrificing accuracy, and that the larger the database the higher the speedup factor.
Ligang Zheng, Yanqiang Lei, Guoping Qiu, Jiwu Huang
IEEE Trans. Inf. Forensics Secur.4
2012 Reference index-based H.264 video watermarking scheme
abstract
Video watermarking has received much attention over the past years as a promising solution to copy protection. Watermark robustness is still a key issue of research, especially when a watermark is embedded in the compressed video domain. In this article, a robust watermarking scheme for H.264 video is proposed. During video encoding, the watermark is embedded in the index of the reference frame, referred to as reference index, a bitstream syntax element newly proposed in the H.264 standard. Furthermore, the video content (current coded blocks) is modified based on an optimization model, aiming at improving watermark robustness without unacceptably degrading the video's visual quality or increasing the video's bit rate. Compared with the existing schemes, our method has the following three advantages: (1) The bit rate of the watermarked video is adjustable; (2) the robustness against common video operations can be achieved; (3) the watermark embedding and extraction are simple. Extensive experiments have verified the good performance of the proposed watermarking scheme.
Jian Li 0034, Hongmei Liu 0001, Jiwu Huang, Yun Q. Shi 0001
ACM Trans. Multim. Comput. Commun. Appl.3
2012 Exposing MP3 audio forgeries using frame offsets
abstract
Audio recordings should be authenticated before they are used as evidence. Although audio watermarking and signature are widely applied for authentication, these two techniques require accessing the original audio before it is published. Passive authentication is necessary for digital audio, especially for the most popular audio format: MP3. In this article, we propose a passive approach to detect forgeries of MP3 audio. During the process of MP3 encoding the audio samples are divided into frames, and thus each frame has its own frame offset after encoding. Forgeries lead to the breaking of framing grids. So the frame offset is a good indication for locating forgeries, and it can be retrieved by the identification of the quantization characteristic. In this way, the doctored positions can be automatically located. Experimental results demonstrate that the proposed approach is effective in detecting some common forgeries, such as deletion, insertion, substitution, and splicing. Even when the bit rate is as low as 32 kbps, the detection rate is above 99%.
Rui Yang 0006, Zhenhua Qu, Jiwu Huang
ACM Trans. Multim. Comput. Commun. Appl.3
2011 Secure JPEG steganography by LSB+ matching and multi-band embedding
abstract
A new steganographic algorithm is proposed for JPEG images by modifying the block DCT coefficients. Firstly, an embedding algorithm called LSB+matching is generated to approximately preserve the marginal distribution of DCT coefficients. We further divide the DCT coefficients into four frequency bands, including the direct current (DC), low-frequency, middle-frequency, and high-frequency. Via matrix encoding, low data hiding rate and high embedding efficiency are achieved in high-frequency band, while the hiding rate is increased in the middle-frequency and DC bands, and highest in the low-frequency band. In addition, a coefficient selection strategy is employed to make the hidden message less detectable. The proposed algorithm is implemented on a set of 10000 images and tested with four steganalytic algorithms. The experimental results show that it outperforms the F5, nsF5 and LSB+algorithms in terms of detection accuracy for the JPEG images at various levels of quality.
Jiwu Huang
ICIP2
2011 Salient covariance for near-duplicate image and video detection
abstract
This paper introduces the covariance matrix of visually salient image features as a compact and robust descriptor for near duplicate image and video copy detection. We make two novel contributions. We first present a fast method for computing information theoretic based visual saliency maps using a data independent fast transform to replace the conventional data dependent computationally demanding transforms. We then introduce salient covariance (SCOV) - the covariance matrix of various image features within the visually salient regions and use SCOV for near duplicate image and video copy detection. We present experimental results to show that our new fast visual saliency computation technique improves efficiency without compromising performances. We demonstrate that SCOV is a very compact and robust feature for near duplicate image and video copy detection. Compared to popular features such as GIST, SCOV is not only more robust against various manipulations but also can be over 20 times more compact whilst achieving the same or better performances.
Ligang Zheng, Guoping Qiu, Jiwu Huang, Hao Fu 0001
ICIP3
2011 Three Novel Algorithms for Hiding Data in PDF Files Based on Incremental Updates
Hongmei Liu 0001, Jian Li 0034, Jiwu Huang
IWDW4
2011 Implementation of the discrete wavelet transform and multiresolution analysis in the encrypted domain
abstract
Signal processing in the encrypted domain is a new technology for protecting valuable signals from insecure signal processing. Although there has been some research in the area, this field of research is still in its infancy.
Peijia Zheng, Jiwu Huang
ACM Multimedia2
2011 Steganalysis of JPEG steganography with complementary embedding strategy
abstract
Recently, a new high-performance JPEG steganography with a complementary embedding strategy (JPEG-CES) was presented. It can disable many specific steganalysers such as the Chi-square family and S family detectors, which have been used to attack J-Steg, JPHide, F5 and OutGuess successfully. In this work, a study on the security performance of JPEG-CES is reported. Our theoretical analysis demonstrates that in this algorithm, the number of the different quantised discrete cosine transform (qDCT) coefficients and the symmetry of the qDCT coefficient histogram both will be disturbed when the secret message is embedded. Moreover, the intrinsic sign and magnitude dependencies that existed in intra-block and inter-block qDCT coefficients will be disturbed too. Thus it may be detected by some modern universal steganalysers which can catch these disturbances. In this work, the authors have proposed two new steganalytic approaches. Through exploring the distortions that have been introduced into the qDCT coefficient histogram and the dependencies existed in the intra-block and inter-block sense, respectively, these two alternative steganalysers can detect JPEG-CES effectively. In addition, via merging the features of these two steganalysers, a more reliable classifier can be obtained.
Fangjun Huang, Weiqi Luo 0001, Jiwu Huang
IET Inf. Secur.3
2011 Minority codes with improved embedding efficiency for large payloads
Hongmei Liu 0001, Xinzhi Yao, Jiwu Huang
Multim. Tools Appl.3
2011 A more secure steganography based on adaptive pixel-value differencing scheme
Weiqi Luo 0001, Fangjun Huang, Jiwu Huang
Multim. Tools Appl.3
2011 Random Gray code and its performance analysis for image hashing
Guopu Zhu, Sam Kwong, Jiwu Huang, Jianquan Yang
Signal Process.3
2011 Robust image hash in Radon transform domain for authentication
Yanqiang Lei, Yuan-Gen Wang, Jiwu Huang
Signal Process. Image Commun.3
2011 Security Analysis on Spatial ± 1 Steganography for JPEG Decompressed Images
abstract
Although many existing steganalysis works have shown that the spatial ±1 steganography on JPEG pre-compressed images is relatively easier to be detected compared with that on the never-compressed images, most experimental results seem not very convincing since these methods usually assume that the quantization table of the JPEG stegos previously used is known before detection and/or the length of embedded message is fixed. Furthermore, there are just few effective quantitative algorithms for further estimating the spatial modifications. In this letter, we firstly propose an effective method to detect the quantization table from the contaminated digital images which are originally stored as JPEG format based on our recently developed work about JPEG compression error analysis , and then we present a quantitative method to reliably estimate the length of spatial modifications in those gray-scale JPEG stegos by using data fitting technology. The extensive experimental results show that our estimators are very effective, and the order of magnitude of prediction error can remain around measured by the mean absolute difference.
Weiqi Luo 0001, Yuan-Gen Wang, Jiwu Huang
IEEE Signal Process. Lett.3
2011 Geometric Invariant Audio Watermarking Based on an LCM Feature
abstract
The development of a geometric invariant audio watermarking scheme without degrading acoustical quality is challenging work. This paper proposes a multi-bit spread-spectrum audio watermarking scheme based on a geometric invariant log coordinate mapping (LCM) feature. The LCM feature is very robust to audio geometric distortions. The watermark is embedded in the LCM feature, but it is actually embedded in the Fourier coefficients which are mapped to the feature via LCM, so the embedding is actually performed in the DFT domain without interpolation, thus eliminating completely the severe distortion resulted from the non-uniform interpolation mapping. The watermarked audio achieves high auditory quality in both objective and subjective quality assessments. A mixed correlation between the LCM feature and a key-generated PN tracking sequence is proposed to align the log-coordinate mapping, thus synchronizing the watermark efficiently with only one FFT and one IFFT. Both the theoretical analysis and experimental results show that the proposed audio watermarking scheme is not only resilient against common signal processing operations, including low-pass filtering, MP3 recompression, echo addition, volume change, normalization, test functions in the Stirmark benchmark, and DA/AD conversion, but also has conquered the challenging audio geometric distortion and achieves the best robustness against simultaneous geometric distortions, such as pitch invariant time-scale modification (TSM) by ±20%, tempo invariant pitch shifting by 20%, resample TSM with scaling factors between 75% and 140%, and random cropping by 95%. This is mainly contributed by the proposed geometric invariant LCM feature. To our best knowledge, audio watermarking based on LCM has not been reported before.
Xiangui Kang, Rui Yang 0006, Jiwu Huang
IEEE Trans. Multim.3
2010 An image copy detection scheme based on radon transform
abstract
Copy detection aims to identify various content-preserving copies from the same origin and distinguish the different sources. In this paper, we propose a novel feature for image copy detection based on the radon transform (RT). We first theoretically analyze that the proposed feature is invariant against rotation, scaling and translation (RST) operations and robust to lossy JPEG compression and additive noises. The feature can meet the robustness requirements of copy detection applications. The proposed feature is statistically dependent of image content and is capable of capturing unique information of the image for fragility of the copy detection system. The experimental results show that the proposed scheme outperforms the existing methods, when detecting image copies subjected to various transformations, in terms of receiver operating characteristic (ROC) curves.
Yuan-Gen Wang, Yanqiang Lei, Jiwu Huang
ICIP3
2010 A geometrically resilient robust image watermarking scheme using deformable multi-scale transform
abstract
The robust performance against geometrical manipulations is still one of major concerns in robust watermarking although significant improvement has been achieved in past decades. In this paper, we tackle the global geometrical attacks by designing a deformable multi-scale transform (DMST) that has joint shiftability in position, orientation, and scale. Via DMST, we both derive theoretically the principles for geometrical synchronization and develop a template-based scheme to efficiently estimate geometrical parameters. Also, the hidden Markov model in the standard wavelet domain is extended to the steerable wavelet domain and further used to improve the performance of watermark extraction. Experimental simulation demonstrates that the proposed watermarking scheme is quite robust to the common signal processing, geometrical attacks, and their joint attacks.
Chuntao Wang, Jiangqun Ni, Huashuo Zhuo, Jiwu Huang
ICIP4
2010 New JPEG Steganographic Scheme with High Security Performance
Fangjun Huang, Yun Q. Shi 0001, Jiwu Huang
IWDW3
2010 Discriminating Computer Graphics Images and Natural Images Using Hidden Markov Tree Model
Jiwu Huang
IWDW2
2010 Robust AVS audio watermarking
Jiwu Huang
Sci. China Inf. Sci.2
2010 An experimental study on the security performance of YASS
abstract
This paper presents an experimental study on the security performance of Yet Another Steganographic Scheme (YASS). It reports: 1) YASS's security performance with different input images, i.e., uncompressed images and JPEG compressed images; 2) YASS's security performance compared with two other JPEG steganographic schemes MB1 and F5; and 3) some experimental results about extended YASS.
Fangjun Huang, Jiwu Huang, Yun Q. Shi 0001
IEEE Trans. Inf. Forensics Secur.2
2010 Detecting Double JPEG Compression With the Same Quantization Matrix
abstract
Detection of double joint photographic experts group (JPEG) compression is of great significance in the field of digital forensics. Some successful approaches have been presented for detecting double JPEG compression when the primary compression and the secondary compression have different quantization matrixes. However, when the primary compression and the secondary compression have the same quantization matrix, no detection method has been reported yet. In this paper, we present a method which can detect double JPEG compression with the same quantization matrix. Our algorithm is based on the observation that in the process of recompressing a JPEG image with the same quantization matrix over and over again, the number of different JPEG coefficients, i.e., the quantized discrete cosine transform coefficients between the sequential two versions will monotonically decrease in general. For example, the number of different JPEG coefficients between the singly and doubly compressed images is generally larger than the number of different JPEG coefficients between the corresponding doubly and triply compressed images. Via a novel random perturbation strategy implemented on the JPEG coefficients of the recompressed test image, we can find a “proper” randomly perturbed ratio. For different images, this universal “proper” ratio will generate a dynamically changed threshold, which can be utilized to discriminate the singly compressed image and doubly compressed image. Furthermore, our method has the potential to detect triple JPEG compression, four times JPEG compression, etc.
Fangjun Huang, Jiwu Huang, Yun Q. Shi 0001
IEEE Trans. Inf. Forensics Secur.2
2010 Efficient general print-scanning resilient data hiding based on uniform log-polar mapping
abstract
This paper proposes an efficient, blind, and robust data hiding scheme which is resilient to both geometric distortion and the general print-scan process, based on a near uniform log-polar mapping (ULPM). In contrast to performing inverse log-polar mapping (a mapping from the log-polar system to the Cartesian system) to the watermark signal or its index as done in the prior works, we apply ULPM to the frequency index (u,v) in the Cartesian system to obtain the discrete log-polar coordinate (l1,l2), then embed one watermark bitw(l1,l2) in the corresponding discrete Fourier transform coefficientc(u,v). This mapping of index from the Cartesian system to the log-polar system but embedding the corresponding watermark directly in the Cartesian domain not only completely removes the interpolation distortion and the interference distortion introduced to the watermark signal as observed in some prior works, but also largely expands the cardinality of watermark in the log-polar mapping domain. Both theoretical analysis and experimental results show that the proposed watermarking scheme achieves excellent robustness to geometric distortion, normal signal processing, and the general print-scan process. Compared to existing watermarking schemes, our algorithm offers significant improvement in terms of robustness against general print-scan, receiver operating characteristic (ROC) performance, and efficiency of blind resynchronization.
Xiangui Kang, Jiwu Huang, Wenjun Zeng 0001
IEEE Trans. Inf. Forensics Secur.2
2010 Edge adaptive image steganography based on LSB matching revisited
abstract
The least-significant-bit (LSB)-based approach is a popular type of steganographic algorithms in the spatial domain. However, we find that in most existing approaches, the choice of embedding positions within a cover image mainly depends on a pseudorandom number generator without considering the relationship between the image content itself and the size of the secret message. Thus the smooth/flat regions in the cover images will inevitably be contaminated after data hiding even at a low embedding rate, and this will lead to poor visual quality and low security based on our analysis and extensive experiments, especially for those images with many smooth regions. In this paper, we expand the LSB matching revisited image steganography and propose an edge adaptive scheme which can select the embedding regions according to the size of secret message and the difference between two consecutive pixels in the cover image. For lower embedding rates, only sharper edge regions are used while keeping the other smoother regions as they are. When the embedding rate increases, more edge regions can be released adaptively for data hiding by adjusting just a few parameters. The experimental results evaluated on 6000 natural images with three specific and four universal steganalytic algorithms show that the new scheme can enhance the security significantly compared with typical LSB-based approaches as well as their edge adaptive ones, such as pixel-value-differencing-based approaches, while preserving higher visual quality of stego images at the same time.
Weiqi Luo 0001, Fangjun Huang, Jiwu Huang
IEEE Trans. Inf. Forensics Secur.3
2010 JPEG error analysis and its applications to digital image forensics
abstract
JPEG is one of the most extensively used image formats. Understanding the inherent characteristics of JPEG may play a useful role in digital image forensics. In this paper, we introduce JPEG error analysis to the study of image forensics. The main errors of JPEG include quantization, rounding, and truncation errors. Through theoretically analyzing the effects of these errors on single and double JPEG compression, we have developed three novel schemes for image forensics including identifying whether a bitmap image has previously been JPEG compressed, estimating the quantization steps of a JPEG image, and detecting the quantization table of a JPEG image. Extensive experimental results show that our new methods significantly outperform existing techniques especially for the images of small sizes. We also show that the new method can reliably detect JPEG image blocks which are as small as 8 × 8 pixels and compressed with quality factors as high as 98. This performance is important for analyzing and locating small tampered regions within a composite image.
Weiqi Luo 0001, Jiwu Huang, Guoping Qiu
IEEE Trans. Inf. Forensics Secur.2
2010 Detection of Quantization Artifacts and Its Applications to Transform Encoder Identification
abstract
Quantization is one of the commonly used techniques in most lossy image source encoders. It is observed that the quantization operation usually introduces some obvious artifacts into the histogram of the corresponding transform coefficients under various compression schemes. By investigating such inherent artifacts over all candidate transform coefficients, it is possible to identify the transform, as well as some parameters previously employed in the transform-based encoder from a decompressed image. In this paper, we first analyze the properties of the quantized coefficients and present a simple yet effective way to detect the quantization artifacts, and then we propose an approach to identify the transform-based encoder based on the quantization artifacts detection. The simulation results evaluated on thousands of natural images with some popular compression schemes demonstrate the effectiveness of our method.
Weiqi Luo 0001, Yuan-Gen Wang, Jiwu Huang
IEEE Trans. Inf. Forensics Secur.3
2010 Fragility analysis of adaptive quantization-based image hashing
abstract
Fragility is one of the most important properties of authentication-oriented image hashing. However, to date, there has been little theoretical analysis on the fragility of image hashing. In this paper, we propose a measure called expected discriminability for the fragility of image hashing and study this fragility theoretically based on the proposed measure. According to our analysis, when Gray code is applied into the discrete-binary conversion stage of image hashing, the value of the expected discriminability, which is dominated by the quantization stage of image hashing, is no more than 1/2. We further evaluate the expected discriminability of the image-hashing scheme that uses adaptive quantization, which is the most popular quantization scheme in the field of image hashing. Our evaluation reveals that if deterministic adaptive quantization is applied, then the expected discriminability of the image-hashing scheme can reach the maximum value (i.e., 1/2). Finally, some experiments are conducted to validate our theoretical analysis and to compare the performance of several quantization schemes for image hashing.
Guopu Zhu, Jiwu Huang, Sam Kwong, Jianquan Yang
IEEE Trans. Inf. Forensics Secur.2
2009 Content-based authentication algorithm for binary images
abstract
This paper proposes a content-based authentication scheme for tampering detection and localization of binary images. The watermark is generated by the feature vector of the original binary image, and embedded back into the image by a structural method. The feature vector we used is Zernike moments. It can tell the degree of the tamper in the binary image. A distance between two feature vectors is defined. We authenticate image by the distance between the extracted watermark and the feature vector of the test image. Once the tampering is detected, we can locate the tampered areas by comparing different components of the distance. The experimental results show that the algorithm can locate the tamper effectively.
Xinzhi Yao, Hongmei Liu 0001, Wei Rui, Jiwu Huang
ICIP4
2009 Temporal Statistic Based Video Watermarking Scheme Robust against Geometric Attacks and Frame Dropping
Jiangqun Ni, Jiwu Huang
IWDW3
2009 Robust AVS Audio Watermarking
Jiwu Huang
IWDW2
2009 Calibration based universal JPEG steganalysis
Fangjun Huang, Jiwu Huang
Sci. China Ser. F Inf. Sci.2
2009 Non-ambiguity of blind watermarking: a revisit with analytical resolution
Xiangui Kang, Jiwu Huang, Wenjun Zeng 0001, Yun Q. Shi 0001
Sci. China Ser. F Inf. Sci.2
2009 Discriminating between photorealistic computer graphics and natural images using fractal geometry
JiongBin Chen, Jiwu Huang
Sci. China Ser. F Inf. Sci.3
2009 Steganalysis of YASS
abstract
A promising steganographic method-yet another steganography scheme (YASS)-was designed to resist blind steganalysis via embedding data in randomized locations. In addition to a concrete realization which is named the YASS algorithm in this paper, a few strategies were proposed to work with the YASS algorithm in order to enhance the data embedding rate and security. In this work, the YASS algorithm and these strategies, together referred to as YASS, have been analyzed from a warden's perspective. It is observed that the embedding locations chosen by YASS are not randomized enough and the YASS embedding scheme causes detectable artifacts. We present a steganalytic method to attack the YASS algorithm, which is facilitated by a specifically selected steganalytic observation domain (SO-domain), a term to define the domain from which steganalytic features are extracted. The proposed SO-domain is not exactly, but partially accesses, the domain where the YASS algorithm embeds data. Statistical features generated from the SO-domain have demonstrated high effectiveness in detecting the YASS algorithm and identifying some embedding parameters. In addition, we discuss how to defeat the above-mentioned strategies of YASS and demonstrate a countermeasure to a new case in which the randomness of the embedding locations is enhanced. The success of detecting YASS by the proposed method indicates a properly selected SO-domain is beneficial for steganalysis and confirms that the embedding locations are of great importance in designing a secure steganographic scheme.
Bin Li 0011, Jiwu Huang, Yun Q. Shi 0001
IEEE Trans. Inf. Forensics Secur.2
2009 A study on the randomness measure of image hashing
abstract
How to measure the security of image hashing is still an open issue in the field of image authentication. Some works have been conducted on the security measure of image hashing. One of the most important works is the randomness measure proposed by Swaminathan, which uses differential entropy as a metric to evaluate the security of randomized image features and has been applied mainly in the security analysis of the feature extraction stage of image hashing. It is meaningful to measure the randomness of the image features over the secret-key set for the security of image hashing because the image features extracted by image hashing should be generated randomly and difficult to guess. However, as is well known, differential entropy is not invariant to scaling; thus it might not be enough to evaluate the security of randomized image features. In this paper, we show the fact that if the image features of an image hash function are scaled by a constant that is large than one, then the tradeoff between the robustness and the fragility of the image hash function will not change at all, but the security indicated by the randomness measure will increase. The above-mentioned fact seems to contradict the following. First, the security of image hashing, which conflicts with robustness and fragility, cannot increase freely. Secondly, a deterministic operation, such as deterministic scaling, does not change the security of image hashing in terms of the difficulty of guessing the secret key or randomized image features. Therefore, the randomness measure should be modified to be invariant to scaling at least.
Guopu Zhu, Jiwu Huang, Sam Kwong, Jianquan Yang
IEEE Trans. Inf. Forensics Secur.2
2008 A convolutive mixing model for shifted double JPEG compression with application to passive image authentication
abstract
The artifacts by JPEG recompression have been demonstrated to be useful in passive image authentication. In this paper, we focus on the shifted double JPEG problem, aiming at identifying if a given JPEG image has ever been compressed twice with inconsistent block segmentation. We formulated the shifted double JPEG compression (SD-JPEG) as a noisy convolutive mixing model mostly studied in blind source separation (BSS). In noise free condition, the model can be solved by directly applying the independent component analysis (ICA) method with minor constraint to the contents of natural images. In order to achieve robust identification in noisy condition, the asymmetry of the independent value map (IVM) is exploited to obtain a normalized criteria of the independency. We generate a total of 13 features to fully represent the asymmetric characteristic of the independent value map and then feed to a support vector machine (SVM) classifier. Experiment results on a set of 1000 images, with various parameter settings, demonstrated the effectiveness of our method.
Zhenhua Qu, Weiqi Luo 0001, Jiwu Huang
ICASSP3
2008 Universal JPEG steganalysis based on microscopic and macroscopic calibration
abstract
In this paper, we present a new universal steganalysis scheme to effectively attack some recently proposed JPEG steganography. Different from the other steganalyzers, not only the magnitude but also the sign dependencies existed in intra-block and inter-block quantized DCT (discrete cosine transform) coefficients are exploited by the Markov empirical transition matrices. Moreover, a new microscopic and macroscopic calibration method is proposed to calibrate the local and global distribution of the quantized DCT coefficients of the test image, thus improve the detecting performance. Experimental results demonstrate that our proposed scheme outperforms some existing steganalyzers in attacking the advanced JPEG steganography such as F5, MB1 and Outguess.
Fangjun Huang, Bin Li 0011, Jiwu Huang
ICIP3
2008 A study on security performance of YASS
abstract
YASS (Yet another steganographic scheme) is a newly developed JPEG steganographic method. Through embedding data in the randomized 8×8 blocks which do not coincide with the 8×8 grid used in JPEG compression, it effectively disables the self-calibration process popularly used in today’s JPEG steganalyzers. However, with YASS’ complicated embedding procedure, the intra- and inter-block dependency among the quantized DCT coefficients belonging to the original image is disturbed after the secret message embedding. Furthermore, because of the randomly selection of an 8×8 block within a large block and the necessary utilization of error correction code, the amount of information that YASS can embed is largely reduced. In this paper a study on security performance of YASS is reported. Our experimental results have demonstrated that 1) the steganalyzers which utilizes intra- and/or inter-block correlation of JPEG coefficients can break YASS, 2) with embedding the same amount of information bits, the security of YASS is not stronger than that of MB1 when some today’s blind JPEG steganalyzers are used.
Fangjun Huang, Yun Q. Shi 0001, Jiwu Huang
ICIP3
2008 An efficient print-scanning resilient data hiding scheme based on a novel LPM
abstract
Print-scan resilient data hiding has not been extensively researched. This paper presents an efficient multi-bit blind watermarking scheme based on a novel Fourier log-polar mapping (LPM). The watermark resynchronization after print-scanning is efficiently solved by an embedded tracking pattern which cannot be removed by template removing attacks and is not detectable for a malicious part. Experimental results show that the proposed watermarking scheme has excellent robustness to print-scanning, cropping, geometric distortion and JPEG compression etc. The obtained success ratios of extraction 60 bits message without error from the combination attack of JPEG compressed with quality factor of 50–100 and then print-scanning were at least 95%.
Xiangui Kang, Xiong Zhong, Jiwu Huang, Wenjun Zeng 0001
ICIP3
2008 A Robust Watermarking Scheme for H.264
Jian Li 0034, Hongmei Liu 0001, Jiwu Huang
IWDW3
2008 A Novel Method for Block Size Forensics Based on Morphological Operations
Weiqi Luo 0001, Jiwu Huang, Guoping Qiu
IWDW2
2008 Robust Audio Watermarking Based on Log-Polar Frequency Index
Rui Yang 0006, Xiangui Kang, Jiwu Huang
IWDW3
2008 GSM Based Security Analysis for Add-SS Watermarking
Dong Zhang 0002, Jiangqun Ni, Dah-Jye Lee, Jiwu Huang
IWDW4
2008 Detecting doubly compressed JPEG images by using Mode Based First Digit Features
abstract
In this paper, we utilize the probabilities of the first digits of quantized DCT (Discrete Cosine Transform) coefficients from individual AC (Alternate Current) modes to detect doubly compressed JPEG images. Our proposed features, named by Mode Based First Digit Features (MBFDF), have been shown to outperform all previous methods on discriminating doubly compressed JPEG images from singly compressed JPEG images. Furthermore, combining the MBFDF with a multi-class classification strategy can be exploited to identify the quality factor in the primary JPEG compression, thus successfully revealing the double JPEG compression history of a given JPEG image.
Bin Li 0011, Yun Q. Shi 0001, Jiwu Huang
MMSP3
2008 Efficient Tate pairing computation using double-base chains
Changan Zhao, Fangguo Zhang, Jiwu Huang
Sci. China Ser. F Inf. Sci.3
2008 Audio watermarking robust against time-scale modification and MP3 compression
Shijun Xiang, Hyoung Joong Kim, Jiwu Huang
Signal Process.3
2008 Steganalysis of Multiple-Base Notational System Steganography
abstract
This letter presents a method for attacking multiple-base notational system (MBNS) steganography . In the MBNS steganography, secret data are converted into symbols in a notational system with multiple bases. The pixels of a host image are then modified such that when the pixel values are divided by the bases, their remainders are equal to the symbols. Through analysis, we prove that the amount of small remainders increases due to the modification. Based on this observation, we propose a steganalytic approach which is effective in not only detecting MBNS steganography but also estimating its embedding rate.
Bin Li 0011, Yanmei Fang, Jiwu Huang
IEEE Signal Process. Lett.3
2008 Invariant Image Watermarking Based on Statistical Features in the Low-Frequency Domain
abstract
Watermark resistance to geometric attacks is an important issue in the image watermarking community. Most countermeasures proposed in the literature usually focus on the problem of global affine transforms such as rotation, scaling and translation (RST), but few are resistant to challenging cropping and random bending attacks (RBAs). The main reason is that in the existing watermarking algorithms, those exploited robust features are more or less related to the pixel position. In this paper, we present an image watermarking scheme by the use of two statistical features (the histogram shape and the mean) in the Gaussian filtered low-frequency component of images. The two features are: 1) mathematically invariant to scaling the size of images; 2) independent of the pixel position in the image plane; 3)statistically resistant to cropping; and 4) robust to interpolation errors during geometric transformations, and common image processing operations. As a result, the watermarking system provides a satisfactory performance for those content-preserving geometric deformations and image processing operations, including JPEG compression, lowpass filtering, cropping and RBAs.
Shijun Xiang, Hyoung Joong Kim, Jiwu Huang
IEEE Trans. Circuits Syst. Video Technol.3
2008 Improving Robustness of Quantization-Based Image Watermarking via Adaptive Receiver
abstract
In this paper, the watermarking channel is modeled as a generalized channel with fading andnonzeromeanadditive noise. In order to improve the watermark robustness against the generalized channel, we present an optimized watermark extraction scheme by using an adaptive receiver for quantization-based watermarking. In the proposed extraction scheme, we adaptively estimate the decision zone of the binary data bits and the quantization step size. A training sequence is embedded into the original image together with the informative watermark. The estimation of the decision zone takes advantage of the response function of the training sequence. Compared to those watermarking schemes without receiver adaptation, the main improvement is the enhanced robustness against median filtering, image intensity Direct Current (DC) change, histogram equalization, color reduction, image intensity linear scaling, image intensity nonlinear scaling such as Gamma correction etc.
Xiangui Kang, Jiwu Huang, Wenjun Zeng 0001
IEEE Trans. Multim.2
2007 A Novel Method for Detecting Cropped and Recompressed Image Block
abstract
One of the most common practices in image tampering involves cropping a patch from a source and pasting it onto a target. In this paper, we present a novel method for the detection of such tampering operations in JPEG images. The lossy JPEG compression introduces inherent blocking artifacts into the image and our method exploits such artifacts to serve as a 'watermark' for the detection of image tampering. We develop the blocking artifact characteristics matrix (BACM) and show that, for the original JPEG images, the BACM exhibits regular symmetrical shape; for images that are cropped from another JPEG image and re-saved as JPEG images, the regular symmetrical property of the BACM is destroyed. We fully exploit this property of the BACM and derive representation features from the BACM to train a support vector machine (SVM) classifier for recognizing whether an image is an original JPEG image or it has been cropped from another JPEG image and re-saved as a JPEG image. We present experiment results to show the efficacy of our method.
Weiqi Luo 0001, Zhenhua Qu, Jiwu Huang, Guoping Qiu
ICASSP (2)3
2007 A GA-Based Joint Coding and Embedding Optimization for Robust and High Capacity Image Watermarking
abstract
A new informed image watermarking algorithm is presented in this paper, which can achieve the information rate of 1/64 bits/pixel with high robustness. Firstly, a LOT (local optimal test) detector based on HMM in wavelet domain is developed to tackle the issue that the exact strength for informed embedding is unknown to the receiver. Then based on the LOT detector, the dirty-paper code for informed coding is constructed and the metric for the robustness is defined accordingly. Unlike the previous approaches of informed watermarking which take the informed coding and embedding process separately, the proposed algorithm implements a joint coding and embedding optimization for high capacity and robust watermarking. The genetic algorithm (GA) is employed to optimize the robustness and distortion constraints simultaneously. Experimental results show that the proposed algorithm achieves significant improvements in performance against JPEG, gain attack, low-pass filtering and etc.
Jiangqun Ni, Chuntao Wang, Jiwu Huang, Rongyue Zhang, Meiying Huang
ICASSP (2)3
2007 Attack LSB Matching Steganography by Counting Alteration Rate of the Number of Neighbourhood Gray Levels
abstract
In this paper, we propose a new method for attacking the LSB (least significant bit) matching based steganography. Different from the LSB substitution, the least two or more significant bit-planes of the cover image would be changed during the embedding in LSB matching steganography and thus the pairs of values do not exist in stego image. In our proposed method, we get an image by combining the least two significant bit-planes and divide it into 3x3 overlapped subimages. The subimages are grouped into four types, i.e.T1,T2,T3andT4according to the count of gray levels. Via embedding a random sequence by LSB matching and then computing the alteration rate of the number of elements inT1, we find that normally the alteration rate is higher in cover image than in the corresponding stego image. This new finding is used as the discrimination rule in our method. Experimental results demonstrate that the proposed algorithm is efficient to detect the LSB matching stegonagraphy on uncompressed gray scale images.
Fangjun Huang, Bin Li 0011, Jiwu Huang
ICIP (1)3
2007 Steganalysis of LSB Greedy Embedding Algorithm for JPEG Images using Coefficient Symmetry
abstract
A recently developed LSB greedy embedding algorithm for JPEG images is capable of resisting the chi-square attack. By carefully studying the quantized DCT (discrete cosine transform) coefficients of the cover and stego images, we find that the embedding algorithm does not preserve the histogram of the DCT coefficients well. In this paper, we define a new chi-square statistic which is used to measure whether the image under scrutiny is like the cover or the stego. Our proposed steganalytic method is based on the symmetry property of the DCT coefficients in JPEG images. It can also be used in the scenario where the cover images are double JPEG compressed. The reliability of this specific steganalytic scheme depends on the embedding rate and it is influenced by the JPEG quality factor. Experimental results show that when the embedding rate exceeds half of the maximal embedding capacity, the steganographic algorithm is detectable with a very low false negative rate, whatever the quality factor is.
Bin Li 0011, Fangjun Huang, Jiwu Huang
ICIP (1)3
2007 Binary Image Authentication using Zernike Moments
abstract
In this paper, we propose a content-based binary image authentication scheme. At first, we use Zernike moments magnitudes (ZMM) to generate the feature vector and demonstrate that this feature vector can represent the binary image and decide its authenticity effectively. Then the watermark is generated by quantizing ZMMs and embedded into the image. The authentication doesn't need the original watermark. The decision depends on the distance between the extracted watermark and the feature vector of the test image and a metric measure. To decrease the influence of watermarking on the feature vector, we split the binary image into two parts by a random mask, one for generating feature vector and the other for embedding watermark. Zernike moments are usually computationally expensive, so we propose a fast algorithm. Extensive experiments show that our scheme can detect malicious attacks effectively.
Hongmei Liu 0001, Wei Rui, Jiwu Huang
ICIP (1)3
2007 Effect of Different Coding Patterns on Compressed Frequency Domain Based Universal JPEG Steganalysis
Bin Li 0011, Fangjun Huang, Shunquan Tan, Jiwu Huang, Yun Q. Shi 0001
IWDW4
2007 Steganalysis of Enhanced BPCS Steganography Using the Hilbert-Huang Transform Based Sequential Analysis
Shunquan Tan, Jiwu Huang, Yun Q. Shi 0001
IWDW2
2007 Survey of information security
Changxiang Shen, Huanguo Zhang, Dengguo Feng, Zhenfu Cao, Jiwu Huang
Sci. China Ser. F Inf. Sci.5
2007 A survey of passive technology for digital image forensics
Weiqi Luo 0001, Zhenhua Qu, Jiwu Huang
Frontiers Comput. Sci. China4
2007 Robust Image Watermarking Based on Multiband Wavelets and Empirical Mode Decomposition
abstract
In this paper, we propose a blind image watermarking algorithm based on the multiband wavelet transformation and the empirical mode decomposition. Unlike the watermark algorithms based on the traditional two-band wavelet transform, where the watermark bits are embedded directly on the wavelet coefficients, in the proposed scheme, we embed the watermark bits in the mean trend of some middle-frequency subimages in the wavelet domain. We further select appropriate dilation factor and filters in the multiband wavelet transform to achieve better performance in terms of perceptually invisibility and the robustness of the watermark. The experimental results show that the proposed blind watermarking scheme is robust against JPEG compression, Gaussian noise, salt and pepper noise, median filtering, and ConvFilter attacks. The comparison analysis demonstrate that our scheme has better performance than the watermarking schemes reported recently.
Ning Bi, Qiyu Sun, Daren Huang, Zhihua Yang, Jiwu Huang
IEEE Trans. Image Process.5
2007 Histogram-Based Audio Watermarking Against Time-Scale Modification and Cropping Attacks
abstract
In audio watermarking area, the robustness against desynchronization attacks, such as TSM (Time-Scale Modification) and random cropping operations, is still one of the most challenging issues. In this paper, we present a multibit robust audio watermarking solution for such a problem by using the insensitivity of the audio histogram shape and the modified mean to TSM and cropping operations. We address the insensitivity property in both mathematical analysis and experimental testing by representing the histogram shape as the relative relations in the number of samples among groups of three neighboring bins. By reassigning the number of samples in groups of three neighboring bins, the watermark sequence is successfully embedded. In the embedding process, the histogram is extracted from a selected amplitude range by referring to the mean in such a way that the watermark will be able to be resistant to amplitude scaling and avoid exhaustive search in the extraction process. The watermarked audio signal is perceptibly similar to the original one. Experimental results demonstrate that the hidden message is very robust to TSM and random cropping attacks, and also has a satisfactory robustness for those common audio signal processing operations.
Shijun Xiang, Jiwu Huang
IEEE Trans. Multim.2
2006 A Hybrid Watermarking Scheme for Video Authentication
abstract
In this paper, we present a hybrid watermarking scheme for video authentication based on wavelet domain. It embeds one robust watermark for temporal authentication distinguishing different inter-attacks, such as frame loss, inserting and reordering. Other two watermarks are used for intra authentication, one for discriminating the malicious attacks from acceptable manipulations, and the other for locating malicious attacks. The latter two watermarks are content-based. Zernike moments magnitudes (ZMMs) of the lowpass wavelet band of the host video frames are chosen as features. Experimental results show that ZMMs are robust to MPEG-2 compression and slight noise, while fragile to content change. This semi-fragile property is used to tell malicious attacks from non-malicious attacks. We also use structure of the embedded ZMMs to locate the tampered area. Experimental results show that this scheme can authenticate the video effectively.
Hongmei Liu 0001, Jiwu Huang
ICIP3
2006 Performance Enhancement for DWT-HMM Image Watermarking with Content-Adaptive Approach
abstract
A DWT-HMM (hidden Markov model in wavelet domain) image watermarking algorithm with content-adaptive approach is proposed in this paper to optimized the trade-off between robustness and visual quality, which is characterized as follows: the entropy mask proposed by Watson is constructed in wavelet domain; the entropy mask and the new developed integrated HVS are used as the measures to adaptively select image components for watermarking; repeat-accumulation (RA) code with erasure and error correction is employed to synchronize the watermarked image; and a posterior HMM is utilized in watermark detection. Considerable improvement in robustness performance with the proposed adaptive algorithm is obtained over the previous DWT-HMM watermarking algorithm with stochastic embedding.
Jiangqun Ni, Chuntao Wang, Jiwu Huang, Rongyue Zhang
ICIP3
2006 Steganalysis of JPEG2000 Lazy-Mode Steganography using the Hilbert-Huang Transform Based Sequential Analysis
abstract
In this paper, we present a steganalytic method to attack JPEG2000 lazy-mode steganography proposed by Su et al. The key element of the method is the Hilbert-Huang transform based analysis of the code-block noise variance sequences of stego images and non-stego noisy images. The Hilbert transform based characteristic vectors are constructed via empirical mode decomposition of the sequences and the support vector machine classifier is used in classification. Experimental results have demonstrated effectiveness of the proposed steganalytic method. According to our best knowledge, this method is the first successful attack of JPEG2000 lazy-mode steganography. And furthermore, the proposed method takes first step towards the application of Hilbert-Huang transform in steganalysis and proves its great advantage.
Shunquan Tan, Jiwu Huang, Zhihua Yang, Yun Q. Shi 0001
ICIP2
2006 A Rotation-Invariant Secure Image Watermarking Algorithm Incorporating Steerable Pyramid Transform
Jiangqun Ni, Rongyue Zhang, Jiwu Huang, Chuntao Wang, Quanbo Li
IWDW3
2006 Robust Audio Watermarking Based on Low-Order Zernike Moments
Shijun Xiang, Jiwu Huang, Rui Yang 0006, Chuntao Wang, Hongmei Liu 0001
IWDW2
2006 Steganalysis of stochastic modulation steganography
Jiwu Huang
Sci. China Ser. F Inf. Sci.2
2006 An algorithm for removable visible watermarking
abstract
A visible watermark may convey ownership information that identifies the originator of image and video. A potential application scenario for visible watermarks was proposed by IBM where an image is originally embedded with a visible watermark before posting on the web for free observation and download. The watermarked image which serves as a "teaser." The watermark can be removed to recreate the unmarked image by request of interested buyers. Before we can design an algorithm for satisfying this application, three basic problems should be solved. First, we need to find a strategy suitable for producing large amount of visually same but numerically different watermarked versions of the image for different users. Second, the algorithm should let the embedding parameters reachable for any legal user to make the embedding process invertible. Third, an unauthorized user should be prevented from removing the embedded watermark pattern. In this letter, we propose a user-key-dependent removable visible watermarking system (RVWS). The user key structure decides both the embedded subset of watermark and the host information adopted for adaptive embedding. The neighbor-dependent embedder adjusts the marking strength to host features and makes unauthorized removal very difficult. With correct user keys, watermark removal can be accomplished in "informed detection" and the high quality unmarked image can be restored. In contrast, unauthorized operation either overly or insufficiently removes the watermark due to wrong estimation of embedding parameters, and thus, the resulting image has apparent defect.
Yongjian Hu, Sam Kwong, Jiwu Huang
IEEE Trans. Circuits Syst. Video Technol.3
2005 Revaluation of Error Correcting Coding in Watermarking Channel
Limin Gu, Yanmei Fang, Jiwu Huang
CANS3
2005 The M-band wavelets in image watermarking
abstract
Multi-band (M-band) wavelet domain presents a novelty to host the watermark. In this paper, a new family of M-band wavelets, which is symmetric and parameterized with a variable /spl lambda/, is proposed and applied to image watermarking. The parameter /spl lambda/ also can be used as a key in watermark detection to improve the security of watermark. The multi-resolution analysis (MRA) of M-band wavelet transform, integrating with the CDMA (code division multiple access) encoding techniques is studied and employed to watermarking. The security, imperceptibility, and the robustness against JPEG compression and Gaussian noise, are analyzed for the proposed watermarking scheme. The experiments of watermarking based on M-band wavelet transform provide more encouraging results than those based on 2-band wavelets.
Yanmei Fang, Ning Bi, Daren Huang, Jiwu Huang
ICIP (1)4
2005 Performance Analysis of CDMA-Based Watermarking with Quantization Scheme
Yanmei Fang, Limin Gu, Jiwu Huang
ISPEC3
2005 A New Approach to Estimating Hidden Message Length in Stochastic Modulation Steganography
Jiwu Huang, Guoping Qiu
IWDW2
2005 Multi-band Wavelet Based Digital Watermarking Using Principal Component Analysis
Xiangui Kang, Yun Q. Shi 0001, Jiwu Huang, Wenjun Zeng 0001
IWDW3
2005 A Robust Multi-bit Image Watermarking Algorithm Based on HMM in Wavelet Domain
Jiangqun Ni, Rongyue Zhang, Jiwu Huang, Chuntao Wang
IWDW3
2005 A RST-Invariant Robust DWT-HMM Watermarking Algorithm Incorporating Zernike Moments and Template
Jiangqun Ni, Chuntao Wang, Jiwu Huang, Rongyue Zhang
KES (1)3
2005 Semi-fragile Watermarking Based on Zernike Moments and Integer Wavelet Transform
Xiaoyun Wu, Hongmei Liu 0001, Jiwu Huang
KES (2)3
2005 Analysis of Quantization-Based Audio Watermarking in DA/AD Conversions
Shijun Xiang, Jiwu Huang, Xiaoyun Feng
KES (2)2
2005 Watermarking Parameters Feasible Region Model
abstract
In this paper, the effectiveness and the parameters space of the public watermark algorithm based on modulus of congruence are considered, and the model of the watermark effectiveness feasible region is proposed. The main contributions of the paper are as follows: (1) by distinction parameter, the relations of embedding, extracting watermark and distinction parameter are established, and the necessary and sufficient conditions of lossless watermark detecting or extracting and its parameters feasible region are presented, (2) by parameters feasible region, the control conditions of the watermark robustness and fragileness are given, and the optimality of distinction parameter of the algorithm is proved, (3) under Gaussian noises attack, the mathematical model of the relations of the lossy watermark BER (bit error rate), attack strength and embedding strength is proposed
Changzhen Xiong, Dongxu Qi, Jiwu Huang
MMSP4
2004 Provably Secure and ID-Based Group Signature Scheme
abstract
There are two important directions in group signature scheme: ID-based group signature scheme and provably secure group signature scheme. We first analyze the advantages and flaws of these two directions, and propose a provably secure ID-based group signature scheme basing on the ACJT group signature scheme. Compared with prior ID-Based group signature scheme, our scheme is provably secure in random oracle model. Compared with the ACJT scheme, our scheme is ID-Based and strong nonrepudiation.
Zewen Chen, Jiwu Huang, Daren Huang, Jianhong Zhang 0001, Yumin Wang
AINA (2)2
2004 Improve security of fragile watermarking via parameterized wavelet
abstract
The security is an important issue in watermarking. It has not, however, received enough attention yet. In this paper, we propose a secure fragile watermarking algorithm based on parameterized integer wavelet transform, and the rational range of the parameter is derived theoretically. Without the parameter of the wavelet base used for watermarking, it is hard for attacker to recover or attack the hidden watermark. Multiresolution tamper detection is developed for the accurate detection. Both security and lower computational complexity of the generated fragile watermark are achieved.
Jiwu Huang, Junquan Hu, Daren Huang, Yun Q. Shi 0001
ICIP1
2004 Improve robustness of image watermarking via adaptive receiving
abstract
Almost all the existing popular watermarking schemes model the watermarking channel noises as additive noise with zero-mean. However, our experiments show that this is not reasonable in the case of channel noise introduced by image filtering. This paper presents a new quantization-based watermarking scheme with enhanced robustness via adaptive receiving and turbo coding in addition to other measures. The present algorithm can successfully resist almost all the StirMark testing functions including both common signal processing and geometric distortions in StirMark 4.0 except for random distortion.
Xiangui Kang, Jiwu Huang, Yun Q. Shi 0001
ICIP2
2003 Image Fusion Based Visible Watermarking Using Dual-Tree Complex Wavelet Transform
Yongjian Hu, Jiwu Huang, Sam Kwong, Yiu-Keung Chan
IWDW2
2003 Robust Watermarking with Adaptive Receiving
Xiangui Kang, Jiwu Huang, Yun Q. Shi 0001, Jianxiang Zhu
IWDW2
2003 A DWT-DFT composite watermarking scheme robust to both affine transform and JPEG compression
abstract
Robustness is a crucially important issue in watermarking. Robustness against geometric distortion and JPEG compression at the same time with blind extraction remains especially challenging. A blind discrete wavelet transform-discrete Fourier transform (DWT-DFT) composite image watermarking algorithm that is robust against both affine transformation and JPEG compression is proposed. The algorithm improves robustness by using a new embedding strategy, watermark structure, 2D interleaving, and synchronization technique. A spread-spectrum-based informative watermark with a training sequence is embedded in the coefficients of the LL subband in the DWT domain while a template is embedded in the middle frequency components in the DFT domain. In watermark extraction, we first detect the template in a possibly corrupted watermarked image to obtain the parameters of an affine transform and convert the image back to its original shape. Then, we perform translation registration using the training sequence embedded in the DWT domain, and, finally, extract the informative watermark. Experimental work demonstrates that the proposed algorithm generates a more robust watermark than other reported watermarking algorithms. Specifically it is robust simultaneously against almost all affine transform related testing functions in StirMark 3.1 and JPEG compression with quality factor as low as 10. While the approach is presented for gray-level images, it can also be applied to color images and video sequences.
Xiangui Kang, Jiwu Huang, Yun Q. Shi 0001
IEEE Trans. Circuits Syst. Video Technol.2
2002 A DWT-Based Fragile Watermarking Tolerant of JPEG Compression
Junquan Hu, Jiwu Huang, Daren Huang, Yun Q. Shi 0001
IWDW2
2002 An Image Watermarking Algorithm Robust to Geometric Distortion
Xiangui Kang, Jiwu Huang, Yun Q. Shi 0001
IWDW2
2002 Reliable information bit hiding
abstract
One of challenges encountered in information bit hiding is the reliability of information bit detection. This paper addresses the issue and presents an algorithm in the discrete cosing transform (DCT) domain with a communication theory approach. It embeds information bits (first) in the DC and (then in the) low-frequency AC coefficients. To extract the hidden information bits from a possibly corrupted marked image with a low error probability, we model information hiding as a digital communication problem and apply Bose-Chaudhuri-Hocquenghen channel coding with soft-decision decoding based on matched filtering. The robustness of the hidden bits has been tested with StirMark. The experimental results demonstrate that the embedded information bits are perceptually transparent and can successfully resist common signal processing procedures, jitter attack, aspect ratio variation, scaling change, small angle rotation, small amount cropping, and JPEG compression with quality factor as low as 10. Compared with some information hiding algorithms reported in the literature, it appears that the hidden information bits with the proposed approach are relatively more robust. While the approach is presented for gray level images, it can also be applied to color images and video sequences.
Jiwu Huang, Yun Q. Shi 0001
IEEE Trans. Circuits Syst. Video Technol.1
2001 A Dwt-Based Image Watermarking Algorithm
abstract
In this paper, a new embedding strategy for DWT-based watermarking is proposed. Different from the existing watermarking schemes in which low frequency coefficients are explicitly excluded from watermark embedding, we claim that watermarks should be embedded in the low frequency subband firstly, and the remains should be embedded in high frequency subbands according to the significance of subbands. We also claim that different embedding formula should be applied on the low frequency subband and high frequency subbands respectively. Applying this strategy, an adaptive algorithm incorporating the feature of visual masking of human vision system into watermarking is proposed. In the algorithm, a novel method to classify wavelet blocks is presented. The experimental results demonstrate that the watermarks generated with the proposed algorithm are invisible and robust against noise and commonly used image processing techniques.
Daren Huang, Jiufen Liu, Jiwu Huang, Hongmei Liu 0001
ICME3
2001 An Adaptive Video Watermarking Algorithm
abstract
Abstract: Robustness is one of the major issues of digital watermarking algorithm. An effective method to improve the robustness of the watermark is to embed the watermark adaptively based on the perceptual property and signal characteristics. In this paper, an adaptive video watermarking algorithm based on wavelet domain is proposed. According to the properties of the 2-D wavelet coefficients, the watermark is inserted into the low frequency subband coefficients to achieve better robustness. In order to improve the strength of the components of the watermark, we propose to classify the coefficients of the low frequency subband based on the motion of the object and the texture complexity of the content in the video sequence. According to the result of classification, the strength of the watermark component is adjusted adaptively. The experimental results show that the watermark just generated is robust to video degradation and distortions, e.g., those that result from additive Gaussian noise, MPEG-2 coding at high compress-ratio, temporal downsampling and spatial downsampling, while the transparency of watermark is guaranteed.
Hongmei Liu 0001, Jiwu Huang, Zi-mei Xiao
ICME2
2000 Embedding image watermarks in dc components
abstract
Both watermark structure and embedding strategy affect robustness of image watermarks. Where should watermarks be embedded in the discrete cosine transform (DCT) domain in order for the invisible image watermarks to be robust? Though many papers in the literature agree that watermarks should be embedded in perceptually significant components, dc components are explicitly excluded from watermark embedding. In this letter, a new embedding strategy for watermarking is proposed based on a quantitative analysis on the magnitudes of DCT components of host images. We argue that more robustness can be achieved if watermarks are embedded in dc components since dc components have much larger perceptual capacity than any ac components. Based on this idea, an adaptive watermarking algorithm is presented. We incorporate the feature of texture masking and luminance masking of the human visual system into watermarking. Experimental results demonstrate that the invisible watermarks embedded with the proposed watermark algorithm are very robust.
Jiwu Huang, Yun Q. Shi 0001
IEEE Trans. Circuits Syst. Video Technol.1
1998 Power constrained multiple signaling in digital image watermarking
abstract
A watermark signal, which is a unique sequence of random variables, by itself does not give a decisive indication of ownership. This work addresses the need to include meaningful information such that a string of English characters, numbers, and punctuation is embedded within the watermark signal, while the signal remains a sequence of random variables. We attempt to provide answers for (1) the maximum number of symbols the watermark signal can carry, and (2) the most practical and reliable way to implement this. We approached this problem as power constrained multiple signaling over an AWGN channel.
Jiwu Huang, George F. Elmasry, Yun Q. Shi 0001
MMSP1