EDBT 2026 Demo / reviewers in the wild / expert
Xiangui Kang
dblp:75/2824
· DBLP profile ↗
88ranked-venue papers
15as first author
36since 2021 · last 2026
0000-0002-3134-0353ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 50 · 8 first-author · 24 since 2021Security and privacy · 27 · 6 first-author · 7 since 2021Artificial intelligence and machine learning · 10 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Sentences Based Adversarial Attack on AI-Generated Text DetectorsabstractThe widespread use of AI-generated text has introduced significant security concerns, driving the need for reliable detection systems. However, recent studies reveal that neural network-based detectors are vulnerable to adversarial examples. To improve the robustness of such classifiers, a number of adversarial attack strategies have been developed, particularly in the context of text sentiment classification. Most existing adversarial attack methods focus on the semantics of individual words or sentences, often neglecting the broader contextual semantics of the entire text-particularly in the case of long AI-generated text. This limitation frequently results in adversarial examples that lack fluency and coherence. In this paper, we propose a novel method calledSentence-based Adversarial attack on AI-Generated Text detectors (SAGT), which generates linguistically fluent adversarial examples by inserting model-generated sentences into the original text. To ensure contextual semantic consistency, we extract important keywords from the original text-selected based on changes in the detector's confidence score-and incorporate them into the generated sentences. Extensive experimental results demonstrate that adversarial examples crafted bySAGTcan effectively evade AI-generated text detectors. Rongxin Tu, Xiangui Kang, Chee-Wei Tan 0001, Chihung Chi, Kwok-Yan Lam |
IEEE Trans. Big Data | 2 |
| 2026 | High-Capacity Generative Image Steganography Approach for Hiding Multiple Secret ImagesabstractCurrent high-capacity image steganography methods face challenges in balancing hidden capacity, imperceptibility, and recovery quality. Existing embedding-based image-in-image steganography approaches tend to produce detectable artifacts when hiding multiple images, whereas existing generative methods struggle to conceal full-sized secret images and often generate unrealistic stego images. To address these issues, this paper proposes a novel generative steganography approach that hides multiple secret images in a single realistic generated image. Our main contributions include a meticulously designed autoencoder that compresses and injects secret images into the shallow layer of the generator to increase hidden capacity, a three-stage optimization strategy for stable training to enhance the recovery quality of secret images, and an automatic image selection procedure which explores the advantage of generation diversity to enhance the imperceptibility of stego images. Experimental results demonstrate that our method outperforms embedding-based approaches by achieving higher recovered image quality with a PSNR value of 30.45 dB when concealing four images while maintaining stronger resistance against steganalysis tools, with an accuracy of 50%. Against generative approaches, our method achieves a higher hidden capacity while preserving a superior visual quality of stego images, with a FID of 6.97, surpassing the suboptimal method's FID of 22.72. Xiyao Liu 0001, Lian Zhong, Xiangui Kang, Gerald Schaefer, Da Huang 0002, Ziping Ma 0002 |
IEEE Trans. Multim. | 3 |
| 2025 | Speed Master: Quick or Slow Play to Attack Speaker RecognitionabstractBackdoor attacks pose a significant threat during the model's training phase. Attackers craft pre-defined triggers to break deep neural networks, ensuring the model accurately classifies clean samples during inference yet erroneously classifies samples added with these triggers. Recent studies have shown that speaker recognition systems trained on large-scale data are susceptible to backdoor attacks. Existing attackers employ unnoticed ambient sounds as triggers. However, these sounds are not inherently part of the training samples themselves. In essence, triggers can be designed to maintain an intrinsic connection with the original speech to enhance stealthiness. Our paper presents a novel attack methodology named Speed Master, which undermines deep neural networks by manipulating the speed of speech samples. Specifically, we execute poison-only backdoor attacks using speed or tempo adjustment. Changes in speech rate have become a common occurrence, as seen on platforms that allow users to adjust playback speed. In real-world scenarios, people naturally adjust their speaking rate depending on the context. As a result, changes in a speaker’s speech rate are typically perceived as normal and are unlikely to raise suspicion. Furthermore, detecting such subtle adjustments becomes challenging for users without reference speech. Our comprehensive experiments demonstrate that Speed Master can achieve an ASR over 99% in the digital domain, with only a 0.6% poisoning rate. Additionally, we validate the feasibility of Speed Master in the real world and its resistance to typical defensive measures. Zhe Ye 0001, Ying Ren, Xiangui Kang, Diqun Yan, Bin Ma 0003, Shiqi Wang 0001 |
AAAI | 4 |
| 2025 | PGD-Imp: Rethinking and Unleashing Potential of Classic PGD with Dual Strategies for Imperceptible Adversarial AttacksabstractImperceptible adversarial attacks have recently attracted increasing research interests. Existing methods typically incorporate external modules or loss terms other than a simple lp-norm into the attack process to achieve imperceptibility, while we argue that such additional designs may not be necessary. In this paper, we rethink the essence of imperceptible attacks and propose two simple yet effective strategies to unleash the potential of PGD, the common and classical attack, for imperceptibility from an optimization perspective. Specifically, the Dynamic Step Size is introduced to find the optimal solution with minimal attack cost towards the decision boundary of the attacked model, and the Adaptive Early Stop strategy is adopted to reduce the redundant strength of adversarial perturbations to the minimum level. The proposed PGD-Imperceptible (PGD-Imp) attack achieves state-of-the-art results in imperceptible adversarial attacks for both untargeted and targeted scenarios. When performing untargeted attacks against ResNet-50, PGD-Imp attains 100% (+0.3%) ASR, 0.89 (-1.76) l2distance, and 52.93 (+9.2) PSNR with 57s (-371s) running time, significantly outperforming existing methods. Zitong Yu, Ziqiang He, Z. Jane Wang 0001, Xiangui Kang |
ICASSP | 5 |
| 2025 | CA-UAP: Content-Agnostic Universal Adversarial Perturbation for Enhanced GeneralizationabstractDeep Neural Networks (DNNs) have been shown vulnerable to universal adversarial perturbation (UAP), which are imperceptible and capable of fooling the target model for most samples. Existing universal attack methods mainly focus on aggregating the gradient obtained from global image features to directly optimize (noise-based) or indirectly generate (generator-based) UAP. However, such methods do not yet consider improving the generalization of UAP from the perspective of making the perturbation irrelevant to image content. We note that minimizing self-similarity is helpful to make the UAP irrelevant to the image content. Therefore, we propose a novel Content-Agnostic UAP (CA-UAP), which combines global image features and local patch features to optimize UAP. Specifically, we introduce a self-similarity loss that encourages minimizing the similarity between adversarial perturbed global images and their randomly cropped local regions, making the UAP agnostic to image content and consequently enhancing UAP generalization. Extensive experiments on the ILSVRC 2012 dataset demonstrate that our proposed method outperforms existing methods in both untargeted and targeted attacks, e.g., improving the average fooling rate from 79.12% (achieved by the state-of-the-art method) to 82.25% in targeted attacks. Ziqiang He, Jingyang Wen, Xiangui Kang, Z. Jane Wang 0001 |
ICASSP | 4 |
| 2025 | INN-based Secure Steganography Using Lost Information as Adversarial PerturbationsabstractRecently image steganography methods based on invertible neural networks (INNs) demonstrated the capability to automatically embed and extract secret messages while maintaining high visual quality in stego images. However, there remain concerns about security and invertibility of such methods. In this paper, for the first time, we introduce adversarial hiding into INN-based image steganography method to simultaneously perform steganographic embedding and adversarial perturbation generation, resulting in improved security. Our method enhances the invertibility of the INN structure: It utilizes the lost information of the INN to generate perturbations, which are then combined with the gradient of the cover image to produce an adversarial stego image. Also, a learnable noise layer is proposed to mitigate information loss caused by rounding and truncation during image storage. Therefore, the proposed method significantly improves security while enhancing extraction performance of INN-based steganography approach, as supported by our experimental results. For example, the steganalysis detection accuracy of SRNet decreases from 96.86% to 51.77% at a payload of 0.2 bits per pixel (bpp). Fei Shang, Weixiang Zhao, Xiangui Kang, Z. Jane Wang 0001 |
ICASSP | 3 |
| 2025 | Prototype-Based Communication Topology Optimization for Decentralized Federated LearningabstractDecentralized Federated Learning (DFL) is a privacy-preserving framework that eliminates the need for a central server, enabling direct communication between clients to save communication resources. However, existing approaches do not adequately address the high data heterogeneity among communication neighbors and often employ inefficient neighbor selection strategies. We propose ProToDFL, a novel framework that leverages prototype learning for dynamic topology optimization in DFL. ProToDFL constructs communication topologies through selecting 1-hop and 2-hop neighbors based on the Kullback-Leibler divergence between prototype vectors, ensuring alignment with local data distributions. To prevent topology splitting during optimization, it develops a cut-edge detection mechanism. Additionally, ProToDFL incorporates a Stop-and-Wait strategy inspired by exponential back-off to minimize prototype vector transmissions. We provide theoretical convergence analysis of ProToDFL within a general non-convex framework for decentralized training. Experimental results on multiple real-world datasets show that ProToDFL significantly outperforms existing DFL methods, particularly in heterogeneous and sparse environments. Xinlin Leng, Kangyu Hu, Hanlin Gu, Xiangui Kang |
ICME | 4 |
| 2025 | GM-DF: Generalized Multi-Scenario Deepfake DetectionabstractRecent advances in face forgery detection have shown strong in-domain performance but often fail to generalize to out-of-distribution data, especially when confronted with unseen manipulation techniques or domain shifts (e.g., lighting conditions, camera noise). We propose a novel Mixture-of-Experts framework, termed GM-DF, that decouples domain-specific and domain-invariant features to tackle cross-domain face forgery detection. Our method builds upon a foundation model (CLIP) and incorporates three key modules: (1) Dataset-Embedding Generator that leverages lightweight expert layers and database-aware feature normalization to adaptively modulate features at a per-domain level, capturing idiosyncratic cues without overfitting; (2) Multi-Dataset Representation mechanism that fuses these expert embeddings using scaled dot-product attention and integrates a mask image modeling (MIM) task to amplify local forgery artifacts; (3) Meta-Domain-Embedding Optimizer, inspired by MAML, which alternates between domain-specific (inner-loop) and domain-invariant (outer-loop) updates to facilitate rapid adaptation on new domains. Additionally, inspired by [13] (Yossi Gandelsman, Alexei A Efros, and Jacob Steinhardt. 2024. Interpreting the second-order effects of neurons in clip. arXiv preprint arXiv:2406.04341 (2024)) we introduce second-order feature propagation in the intermediate layers of CLIP to enhance fine-grained artifact cues and propose domain-class disentangled prompts to flexibly encode multi-domain text representations. Together, these strategies enable GM-DF to learn robust, shared forgery cues while preserving essential domain nuances. Our extensive experiments on multiple cross-domain benchmarks demonstrate that GM-DF significantly outperforms state-of-the-art approaches in both detection accuracy and domain transferability, reducing reliance on superficial artifacts and improving generalization to unseen forgeries. Importantly, our design requires minimal overhead beyond standard CLIP, making GM-DF both effective and computationally efficient for real-world face forgery detection. Yingxin Lai, Hongyang Wang 0001, Xiangui Kang, Bin Li 0011, LinLin Shen, Zitong Yu |
ACM Multimedia | 4 |
| 2025 | Secure INN-based Steganography via Model Smoothing and Adversarial AttacksabstractIn recent years, image steganography methods based on invertible neural networks (INNs) have received significant attention due to their invertible structure, which offers advantages in embedding and extracting secret messages. However, current INN-based image steganography methods face challenges, particularly their limited tolerance against noise interference (e.g., added Gaussian noise, adversarial perturbations, and JPEG compression) and vulnerability to detection by advanced deep steganalyzers. To address these concerns, we present a novel steganography framework that combines Median Smoothing Training (MST) with dynamic Projected Gradient Descent (d-PGD). Specifically, our method begins with employing an MST strategy during the training phase to improve the INN’s tolerance to noise, ensuring that accurate message extraction even under noise interference. Subsequently, to improve the security of INN-based steganography, we propose a d-PGD algorithm that can generate minimal adversarial perturbations capable of deceiving deep steganalyzers, thereby improving security without compromising extraction accuracy. Experimental results demonstrate that our method achieves state-of-the-art secret message extraction accuracy while significantly improving resistance against deep steganalyzers. Weixiang Zhao, Fei Shang, Jingyang Wen, Xiangui Kang, Z. Jane Wang 0001 |
MMSP | 5 |
| 2025 | JPEG Image Steganography With Automatic Embedding Cost LearningabstractA great challenge to steganography has arisen with the wide application of steganalysis methods based on convolutional neural networks (CNNs). To this end, embedding cost learning frameworks based on generative adversarial networks (GANs) has been proposed and achieved success for spatial image steganography. However, the application of GAN to JPEG steganography is still in the prototype stage; its antidetectability and training efficiency should be improved. In conventional steganography, research has shown that the side information calculated from the precover can be used to enhance security. However, it is hard to calculate the side information without the spatial domain image. In this work, an embedding cost learning framework for JPEG image steganography via a GAN (JS–GAN) has been proposed, the learned embedding cost can be further adjusted asymmetrically according to the estimated side information (ESI). Experimental results have demonstrated that the proposed method can automatically learn a content‐adaptive embedding cost function, and using the ESI properly can effectively improve the security performance. For example, under the attack of a classic steganalyzer GFR with a quality factor of 75 and 0.4 bpnzAC, the proposed JS–GAN can increase the detection error by 2.58% over J‐UNIWARD, and the ESI–aided version JS–GAN (ESI) can further increase the security performance by 11.25% over JS–GAN. Fei Shang, Xiangui Kang, Yifang Chen 0002, Yun Q. Shi 0001 |
Int. J. Intell. Syst. | 4 |
| 2025 | CAT+: Investigating and Enhancing Audio-Visual Understanding in Large Language ModelsabstractMultimodal Large Language Models (MLLMs) have gained significant attention due to their rich internal implicit knowledge for cross-modal learning. Although advances in bringing audio-visuals into LLMs have resulted in boosts for a variety of Audio-Visual Question Answering (AVQA) tasks, they still face two crucial challenges: 1) audio-visual ambiguity, and 2) audio-visual hallucination. Existing MLLMs can respond to audio-visual content, yet sometimes fail to describe specific objects due to the ambiguity or hallucination of responses. To overcome the two aforementioned issues, we introduce the CAT+, which enhances MLLM to ensure more robust multimodal understanding. We first propose the Sequential Question-guided Module (SQM), which combines tiny transformer layers and cascades Q-Formers to realize a solid audio-visual grounding. After feature alignment and high-quality instruction tuning, we introduce Ambiguity Scoring Direct Preference Optimization (AS-DPO) to correct the problem of CAT+ bias toward ambiguous descriptions. To explore the hallucinatory deficits of MLLMs in dynamic audio-visual scenes, we build a new Audio-visual Hallucination Benchmark, named AVHbench. This benchmark detects the extent of MLLM's hallucinations across three different protocols in the perceptual object, counting, and holistic description tasks. Extensive experiments across video-based understanding, open-ended, and close-ended AVQA demonstrate the superior performance of our method. The AVHbench is released at https://github.com/rikeilong/Bay-CAT. Qilang Ye, Zitong Yu, Rui Shao 0001, Yawen Cui, Xiangui Kang, Xin Liu 0012, Philip Torr 0001, Xiaochun Cao |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | Lp-norm distortion-efficient adversarial attack
Yuan-Gen Wang, Zijia Wang 0001, Xiangui Kang |
Signal Process. Image Commun. | 4 |
| 2025 | Forgery-Aware Adaptive Learning With Vision Transformer for Generalized Face Forgery DetectionabstractWith the rapid progress of generative models, the current challenge in face forgery detection is how to effectively detect realistic manipulated faces from different unseen domains. Though previous studies show that pre-trained Vision Transformer (ViT) based models can achieve some promising results after fully fine-tuning on the Deepfake dataset, their generalization performances are still unsatisfactory. To this end, we present a Forgery-aware Adaptive Vision Transformer (FA-ViT) under the adaptive learning paradigm for generalized face forgery detection, where the parameters in the pre-trained ViT are kept fixed while the designed adaptive modules are optimized to capture forgery features. Specifically, a global adaptive module is designed to model long-range interactions among input tokens, which takes advantage of self-attention mechanism to mine global forgery clues. To further explore essential local forgery clues, a local adaptive module is proposed to expose local inconsistencies by enhancing the local contextual association. In addition, we introduce a fine-grained adaptive learning module that emphasizes the common compact representation of genuine faces through relationship learning in fine-grained pairs, driving these proposed adaptive modules to be aware of fine-grained forgery-aware information. Extensive experiments demonstrate that our FA-ViT achieves state-of-the-arts results in the cross-dataset evaluation, and enhances the robustness against unseen perturbations. Particularly, FA-ViT achieves 93.83% and 78.32% AUC scores on Celeb-DF and DFDC datasets in the cross-dataset evaluation. The code and trained model have been released at:https://github.com/LoveSiameseCat/FAViT. Anwei Luo, Rizhao Cai, Chenqi Kong, Yakun Ju, Xiangui Kang, Jiwu Huang, Alex Chichung Kot |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | All Points Guided Adversarial Generator for Targeted Attack Against Deep Hashing RetrievalabstractDeep hashing has been widely used in image retrieval tasks, while deep hashing networks are vulnerable to adversarial example attacks. To improve the deep hashing networks’ robustness, it is essential to investigate adversarial attacks on the networks, especially targeted attacks. Among the existing targeted attacks for hashing, the generation-based targeted attack methods have attracted increasing attention due to their efficiency in generating adversarial examples. However, these methods supervise the generation of adversarial examples solely with the hash codes of positive samples, without employing the hash codes of all points in the training set to directly participate in supervisory training, thereby making the attack less effective. Since the hash codes of the training set samples are generated by a well-trained hashing model, these hash codes retain rich semantic information of their corresponding samples, highlighting the necessity of sufficiently utilizing them. Therefore, in this paper, we propose a targeted attack method that utilizes all points’ hash codes in the training set to guide the generation of adversarial attack examples directly. Specifically, we first decode the target label to obtain the corresponding feature map. Then, we concatenate the feature map with the query image and feed them into an encoder-decoder network that employs a skip-connection strategy to obtain a perturbed example. Furthermore, to guide adversarial example generation, we introduce a loss function that exploits the similarities between the perturbed example’s hash code and all points’ hash codes in the training set, thereby making sufficient utilization of the rich semantic information in these hash codes. Experimental results illustrate that our method outperforms the state-of-the-art targeted attack methods in targeted attack effectiveness and transferability. The code is available athttps://github.com/rongxintu3/APGA. Rongxin Tu, Xiangui Kang, Chee-Wei Tan 0001, Chihung Chi, Kwok-Yan Lam |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2025 | StealthPhase: Toward a Stealthy Backdoor Attack Against Speaker RecognitionabstractSpeaker recognition systems (SRS) play a vital role in identity authentication. At the same time, researchers have found that these systems are highly vulnerable to backdoor attacks, where the poisoned model will misclassify poisoned inputs. Most backdoor attack methods primarily focus on improving attack success rates (ASR), achieving ASR as high as 99%. However, these methods reveal a significant concern in terms of stealthiness. Poisoned audio often exhibits detectable differences from the clean audio, which can be detected by human listeners or through visualization. To overcome this issue, we prioritize stealthiness in our attack design and propose StealthPhase. Motivated by preliminary experiments on frequency-domain random noise backdoor attacks, our method implants a predefined trigger into the phase spectrum through frequency decomposition to ensure inherent stealth. The predefined trigger uses the natural phase pattern derived from real speech. Therefore, it is both learnable, as it addresses the challenge of designing effective phase-based triggers, and stealthy, as it remains imperceptible in both spectrogram visualizations and auditory perception. A key advantage of our method is that it avoids complex algorithms to optimize triggers and does not require an extra loss function to balance stealthiness and effectiveness. Extensive experimental results demonstrate that StealthPhase achieves 99% ASR with minimal impact on the model’s benign accuracy (BA). Meanwhile, its stealthiness is validated from three perspectives. First, visualizations show that the backdoor audio samples are nearly indistinguishable from clean samples. Second, an audio quality assessment confirms that the trigger introduces minimal perceptual distortion, preserving the overall audio quality. Finally, speech recognition performance evaluation shows that the word error rate (WER) remains largely unaffected. Furthermore, we validate the effectiveness of StealthPhase in real-world scenarios, where it achieves an ASR of 80%, and demonstrate its ability to bypass defense mechanisms. Zhe Ye 0001, Qiben Yan 0001, Xiangui Kang, Jiwu Huang |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2024 | Transferable and high-quality adversarial example generation leveraging diffusion modelabstractIn recent years, adversarial example methods in deep learning have proliferated. Meanwhile diffusion models have gained wide applications across various tasks due to their superior distribution reconstruction capability. By leveraging the design experience from prior adversarial examples, combining it with the modeling proficiency of diffusion models, and employing a cost function to evaluate image smoothness for regulating the regional distribution of adversarial noise, we propose a novel adversarial example design method. Experiments demonstrate that our generated adversarial examples exhibit both high attack success rate and superior image quality. The controlled distribution region of adversarial noises significantly enhances the subjective visual quality of our generated images. Kangze Xu, Ziqiang He, Xiangui Kang, Z. Jane Wang 0001 |
ICME | 3 |
| 2024 | AdvAD: Exploring Non-Parametric Diffusion for Imperceptible Adversarial AttacksabstractImperceptible adversarial attacks aim to fool DNNs by adding imperceptible perturbation to the input data. Previous methods typically improve the imperceptibility of attacks by integrating common attack paradigms with specifically designed perception-based losses or the capabilities of generative models. In this paper, we propose Adversarial Attacks in Diffusion (AdvAD), a novel modeling framework distinct from existing attack paradigms. AdvAD innovatively conceptualizes attacking as a non-parametric diffusion process by theoretically exploring basic modeling approach rather than using the denoising or generation abilities of regular diffusion models requiring neural networks. At each step, much subtler yet effective adversarial guidance is crafted using only the attacked model without any additional network, which gradually leads the end of diffusion process from the original image to a desired imperceptible adversarial example. Grounded in a solid theoretical foundation of the proposed non-parametric diffusion process, AdvAD achieves high attack efficacy and imperceptibility with intrinsically lower overall perturbation strength. Additionally, an enhanced version AdvAD-X is proposed to evaluate the extreme of our novel framework under an ideal scenario. Extensive experiments demonstrate the effectiveness of the proposed AdvAD and AdvAD-X. Compared with state-of-the-art imperceptible attacks, AdvAD achieves an average of 99.9% (+17.3%) ASR with 1.34 (-0.97) $l_2$ distance, 49.74 (+4.76) PSNR and 0.9971 (+0.0043) SSIM against four prevalent DNNs with three different architectures on the ImageNet-compatible dataset. Code is available at https://github.com/XianguiKang/AdvAD. Ziqiang He, Anwei Luo, Jianfang Hu, Z. Jane Wang 0001, Xiangui Kang |
NeurIPS | 6 |
| 2024 | Boosting Deepfake Feature Extractors Using Unsupervised Domain AdaptationabstractTo make deepfake detectors generalizable to different target domains, one effective way is to let the source domain for training be similar to the target domain for detection. This letter tackles the problem from the perspective of domain adaptation and achieves both image-level and feature-level domains alignments. The proposed unsupervised domain adapter accomplishes the image-level domain alignment relying on the combination of cross-domain style feature mixing and diffusion model, and the feature-level domain alignment relying on prototypical consistency guided supervision and adversarial learning. The style is transferred from source to target for the generation of target data proxy in the form of stylized images. The features of the stylized images are further aligned with the target prototypical features. We apply the domain adapter as a feature booster to four current deepfake detectors. Experimental results show that all the detectors get a significant increase in AUC values on cross-dataset testings. We further propose a deepfake detector based on the Xception backbone with our booster. Compared with five state-of-the-art detectors, the proposed detector performs best in all experiments. Yongjian Hu, Zhaolong Gong, Xiangui Kang |
IEEE Signal Process. Lett. | 5 |
| 2024 | ResNeXt+: Attention Mechanisms Based on ResNeXt for Malware Detection and ClassificationabstractMalware detection and classification are crucial for protecting digital devices and information systems. Accurate identification of malware enables researchers and incident responders to take prompt measures against malware and mitigate its damage. With the development of attention mechanisms in the field of computer vision, attention mechanism-based malware detection techniques are also rapidly evolving. The essence of the attention mechanism is to focus on the information of interest and suppress the useless information. In this paper, we develop different plug-and-play attention mechanisms based on the ResNeXt tagging model, where the designed model is trained to focus on the malware features by capturing the malware image channel perception field of view and is also able to provide more helpful and flexible information than other methods. We have named this designed neural network ResNeXt+, and its core modules are built with different plug-and-play attention mechanisms. Extensive experimental results show that ResNeXt+ is effective and efficient in malware detection and classification with high classification accuracy. The proposed methods outperform the state-of-the-art techniques with seven benchmark datasets. Cross-dataset experiments conducted on the Windows and Android datasets, with an accuracy of 90.64% on cross-dataset detection of the android. Ablation experiments are also conducted on seven datasets, which demonstrate that attention mechanisms can improve malware detection and classification accuracy. Yuewang He, Xiangui Kang, Qiben Yan 0001, Enping Li |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2024 | Beyond the Prior Forgery Knowledge: Mining Critical Clues for General Face Forgery DetectionabstractFace forgery detection is essential in combating malicious digital face attacks. Previous methods mainly rely on prior expert knowledge to capture specific forgery clues, such as noise patterns, blending boundaries, and frequency artifacts. However, these methods tend to get trapped in local optima, resulting in limited robustness and generalization capability. To address these issues, we propose a novel Critical Forgery Mining (CFM) framework, which can be flexibly assembled with various backbones to boost their generalization and robustness performance. Specifically, we first build a fine-grained triplet and suppress specific forgery traces through prior knowledge-agnostic data augmentation. Subsequently, we propose a fine-grained relation learning prototype to mine critical information in forgeries through instance and local similarity-aware losses. Moreover, we design a novel progressive learning controller to guide the model to focus on principal feature components, enabling it to learn critical forgery features in a coarse-to-fine manner. The proposed method achieves state-of-the-art forgery detection performance under various challenging evaluation settings. The source code is available at:https://github.com/LoveSiameseCat/CFM. Anwei Luo, Chenqi Kong, Jiwu Huang, Yongjian Hu, Xiangui Kang, Alex Chichung Kot |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2023 | Robust Image Steganography: Hiding Messages in Frequency CoefficientsabstractSteganography is a technique that hides secret messages into a public multimedia object without raising suspicion from third parties. However, most existing works cannot provide good robustness against lossy JPEG compression while maintaining a relatively large embedding capacity. This paper presents an end-to-end robust steganography system based on the invertible neural network (INN). Instead of hiding in the spatial domain, our method directly hides secret messages into the discrete cosine transform (DCT) coefficients of the cover image, which significantly improves the robustness and anti-steganalysis security. A mutual information loss is first proposed to constrain the flow of information in INN. Besides, a two-way fusion module (TWFM) is implemented, utilizing spatial and DCT domain features as auxiliary information to facilitate message extraction. These two designs aid in recovering secret messages from the DCT coefficients losslessly. Experimental results demonstrate that our method yields significantly lower error rates than other existing hiding methods. For example, our method achieves reliable extraction with 0 error rate for 1 bit per pixel (bpp) embedding payload; and under the JPEG compression with quality factor QF=10, the error rate of our method is about 22% lower than the state-of-the-art robust image hiding methods, which demonstrates remarkable robustness against JPEG compression. Yuhang Lan, Fei Shang, Xiangui Kang, Enping Li |
AAAI | 4 |
| 2023 | Double Compression Detection Based on the De-Blocking Filtering of HEVC VideosabstractInstead of detecting whether the whole video sequence is double compressed, a frame-level detection result can provide more precise information for video forensic tasks, such as locate tamper point and restore compression history, et al. But the research on frame-level double compression detection is still in its infancy. Therefore we aim to provide a frame-level detection method for HEVC videos in this paper. The relocated I(RI) frame belongs to different GOP groups from its reference frame at the first compression and may cause more severe blocking effects than other types of P frames. Hence, this paper proposes an algorithm based on the de-blocking filtering feature mode to detect RI frames in the double compressed HEVC videos with shifted GOP structure. Firstly, the abnormal traces of the de-blocking filtering parameters, such as boundary strength, filtering switch and filtering mode, in the RI frame are analyzed. Then, the de-blocking filtering feature is constructed by mapping the different combinations of the three parameters into a single numerical value. Finally, the de-blocking filtering feature of the video clips is adopted as the input of the proposed mini_MobileViT network, which is the combination of Convolutional Neural Network (CN-N) and Transformer, to learn spatial and temporal representations to identify the RI frames. Experimental results demonstrate the advantages of the proposed algorithm in detecting RI frames in the double compressed HEVC videos. Compared with the state-of-art work He’s method, the proposed method has a 1.72% improvement in the accuracy of detecting RI frames. Compared with other traditional methods, there is a more than 10% improvement. Xiangui Kang, Pengcheng Su, Zisheng Huang, Yifang Chen 0002, Jie Wang 0031 |
ICASSP | 1 |
| 2023 | Boosting Transferability of Adversarial Example via an Enhanced Euler's MethodabstractAdversarial examples are intentionally designed images to force convolution neural networks to give error classification outputs. Existing attacks have constructed transferable adversarial examples from the base attack algorithm, data augmentation, ensemble model, etc. Nevertheless, under the black-box case especially facing defense models, the transferability of adversarial examples still needs to be improved. In this paper, we try to develop a better base attack to boost the transferability of adversarial examples. Through analyzing the baseline gradient-based attacks, we found their iterative procedures of updating gradients are similar to numerical Euler’s methods. From the perspective of numerical analysis, we employ an enhanced Euler’s method, with less approximate errors and thus more accurate, to search a better approximate optimal solution to construct a more transferable gradient-based attack. To this end, we apply two-step gradient calculations of the enhanced Euler’s method to correct gradient descent directions. As a base attack, our attacks can be easily integrated with data augmentations and ensemble model augmentations. Experimental results show the proposed augmented attack significantly improves the transferability of adversarial examples and achieves an average attack success rate at least 3% higher than state-of-the-arts under black-box settings with defense mechanisms. Anjie Peng, Hui Zeng 0002, Wenxin Yu 0001, Xiangui Kang |
ICASSP | 5 |
| 2023 | Transferable Waveform-level Adversarial Attack against Speech Anti-spoofing ModelsabstractSpeech anti-spoofing models protect media from malicious fake speech but are vulnerable to adversarial attacks. Studies of adversarial attacks are conducive to developing robust speech anti-spoofing systems. Existing transfer-based attack methods mainly craft adversarial speech examples at the handcrafted-feature level, which have limited attack ability against the real-world anti-spoofing systems, as these systems only have raw waveform input interfaces. In this work, we propose a waveform-level input data transformation, called the temporal smoothing method, to generate more transferable adversarial speech examples. In the optimization iterations of the adversarial perturbation, we randomly smooth input waveforms to prevent the adversarial examples from overfitting white-box surrogate models. The proposed transformation can be combined with any iterative gradient-based attack method. Extensive experiments demonstrate that our method significantly enhances the transferability of waveform-level adversarial speech examples. Bingyuan Huang, Sanshuai Cui, Xiangui Kang, Enping Li |
ICME | 3 |
| 2023 | Adversarial Attacks on Generated Text DetectorsabstractGenerated text detectors can effectively detect the machine-generated texts which aim to produce false information to destroy the credibility of the media platform. However, generated text detectors are vulnerable to adversarial example attacks which focus on char-level perturbations to produce many word errors. In this paper, we design a sentence granularity based black-box attack model Sentence-Keyword-Attack (SK-Attack), which can effectively generate semantics-preserved, fluent, and grammatical adversarial examples. SK-Attack adaptively truncates the input examples based on sentence granularity and searches for the essential sentences to apply a sequence of contextualized perturbations with strict constraints. SK-Attack also applies keyword protection to preserve the keywords from being perturbed. Experiments show that SK-Attack outperforms the baselines when attacking the RoBERTa detector with various challenging generated text datasets and also has strong transferability to other attack models. Pengcheng Su, Rongxin Tu, Hongmei Liu 0001, Yue Qing, Xiangui Kang |
ICME | 5 |
| 2023 | Robust data hiding for JPEG images with invertible neural network
Fei Shang, Yuhang Lan, Enping Li, Xiangui Kang |
Neural Networks | 5 |
| 2023 | Discriminative Frequency Information Learning for End-to-End Speech Anti-SpoofingabstractEnd-to-end technology is an active research topic in speech anti-spoofing. Although end-to-end methods have achieved remarkable success in the speech anti-spoofing, channel effects brought by telephony transmission and certain challenging forms of spoofing attacks still plague them. We observe that differences in the high-frequency components between bonafide and spoofed speech help detect some most troublesome attack forms and the differences also remain after the signals are affected by transmission and codecs. Based on this observation, we aim to utilize the high-frequency information of speech signals to develop better generalization ability to unknown attacks and stronger robustness against transmission and codecs. We propose a raw waveform processing module based on sinc convolution and multiple pre-emphasis to obtain discriminative shallow feature representations. Additionally, we propose an improved backbone to learn discriminative feature embeddings, and a feature classification loss to optimize intra-class and inter-class distances simultaneously. The above modules constitute the proposed Discriminative Frequency-information SincNet, namely DFSincNet. Our proposed algorithm demonstrates competitive performance on both ASVspoof 2019 and 2021 logical access (LA) scenarios. Bingyuan Huang, Sanshuai Cui, Jiwu Huang, Xiangui Kang |
IEEE Signal Process. Lett. | 4 |
| 2023 | An Adaptive IPM-Based HEVC Video Steganography via Minimizing Non-Additive DistortionabstractRecently, almost all the proposed adaptive video steganographic schemes are based on minimizing an additive embedding distortion. However, they ignore the hard fact that the additive embedding distortion is not quite suitable for video steganography because of the interplay of cover elements in video steganography. In this article, an adaptive intra prediction mode based (IPM-based) video steganography is proposed by minimizing the non-additive distortion in HEVC. To reduce the complexity of minimizing the non-additive distortion, a multi-layered embedding structure combined with a proposed embedding distortion updating strategy is adopted to approximate the non-additive distortion in an additive form. First, all IPMs are decomposed into multiple layers based on the distortion drift graph to offer multi-layered embedding. Each IPM in the same layer is considered to be independent, and syndrome-trellis code (STC) can be applied to embed the message segment into each layer with an additive distortion function sequentially. Then, a distortion function composed of self-distortion and drift-distortion is proposed to initialize the distortion of modifying each IPM. Finally, after embedding the first message segment into the IPMs in the first layer with the initialized distortions, an embedding distortion updating strategy is applied to update the distortions of the IPMs in the remaining layers dynamically. Experimental results demonstrate that the proposed adaptive IPM-based video steganography can achieve much better perceptual quality and security performance than the state-of-the-art. Jie Wang 0031, Xuemei Yin, Yifang Chen 0002, Jiwu Huang, Xiangui Kang |
IEEE Trans. Dependable Secur. Comput. | 5 |
| 2022 | An adversarial learning framework with cross-domain loss for median filtered image restoration and anti-forensics
Jianyuan Wu, Tianyao Tong, Yifang Chen 0002, Xiangui Kang, Wei Sun 0007 |
Comput. Secur. | 4 |
| 2022 | Synthetic Speech Detection Based on Local Autoregression and Variance Statistics
Sanshuai Cui, Bingyuan Huang, Jiwu Huang, Xiangui Kang |
IEEE Signal Process. Lett. | 4 |
| 2022 | Boosting Query Efficiency of Meta Attack With Dynamic Fine-TuningabstractIn black-box attack, excessive queries to target model may cause suspicion and expose attacker's identity. Equipped with advanced meta learning technique, Meta Attack simulates the target model with a surrogate model, significantly reducing the queries. However, it queries for ZOO-gradients to correct the estimated meta-gradients with a fixed frequency, thereby still leading to massive unnecessary queries. To overcome this limitation, this letter takes the dynamic changes of the accuracy of the estimated gradients as a starting point, and develops a Dynamic Meta Attack (DMA). At the beginning of each fine-tuning round, DMA computes the distance between the above two types of gradients. Such distance metric can reflect the accuracy of the meta-gradients, and guide the dynamic adjustment of query frequency for the ZOO-gradients. Moreover, the working flow of the dynamic fine-tuning process can be controlled by a set of parameters, which are of physical significance and easy to be tuned. By this means, DMA merely launches queries at critical moments, greatly saving query resource. Experiments conducted on MNIST and CIFAR10 show that the proposed DMA requires far fewer queries than existing methods while maintaining a satisfying attack success rate and distortion. Yuan-Gen Wang, Weixuan Tang 0004, Xiangui Kang |
IEEE Signal Process. Lett. | 4 |
| 2021 | A Layered Embedding-Based Scheme to Cope with Intra-Frame Distortion Drift In IPM-Based HEVC SteganographyabstractThe spatial correlation of the intra-frame prediction units brings great challenges when minimizing embedding distortions using syndrome-trellis coding (STC) in High Efficiency Video Coding (HEVC) steganography. To solve this problem, we propose a layered embedding scheme which embeds information into the intra-prediction modes (IPMs) of 4×4 intra-frame prediction units (PUs) in HEVC. Firstly we divide the PUs of the intra-frame into different layers using Hasse diagram and make modification decisions for PUs in each layer respectively to decorrelate the correlated PUs. Secondly we make a statistics on more than 100,000 sampling PU pairs to quantitatively analyze the impacts between the distortions of PUs and then design a distortion function which takes mutual impacts of PUs into account. Experimental results show that our method can significantly reduce the embedding distortion and improve the security compared with the existing STC-based steganography methods embedding in IPMs. Xiaoqing Jia, Jie Wang 0031, Yongliang Liu, Xiangui Kang, Yun Q. Shi 0001 |
ICASSP | 4 |
| 2021 | A Capsule Network Based Approach for Detection of Audio Spoofing AttacksabstractAudio spoofing attacks not only increasingly pose a threat to automatic speaker verification systems but also have the potential to destabilize national security (e.g., by creating fake audio of influential politicians). The main purpose of anti-spoofing is to detect fake audios synthesized by advanced methods, while current algorithms using convolutional neural networks as classifiers exposed poor generalization to the unknown attacks. In this paper, as the first attempt, we introduce a capsule network to enhance the generalization of the detection system. To make the capsule network suitable for anti-spoofing tasks, we modified the original dynamic routing algorithm to force the model to pay more attention to artifacts and thus yield better detection performance for text-to-speech/voice conversion attacks. Furthermore, replay attack detection is also investigated, and the results indicate that our proposed approach is also highly capable of detecting replay attacks. Anwei Luo, Enlei Li, Yongliang Liu, Xiangui Kang, Z. Jane Wang 0001 |
ICASSP | 4 |
| 2021 | Face Forgery Detection Based On Segmentation NetworkabstractRecent progress in facial manipulation technologies have made it hard to distinguish the sophisticated face swapped images/videos. Due to the diversity of generation software and data sources, it is extremely challenging to devise an efficient generality framework. Instead of regarding the detection process as a vanilla binary classification task, we proposed a detection framework based on pixel-level classification. Considering that the acquisition of real pixel-level ground-truth is somehow expensive or even impractical, we proposed a pseudo ground-truth generation pipeline with prior knowledge of facial manipulation. Besides, we added a new module into the neural network to capture frequency clues, while the ablation experiment verified the effectiveness of this module. The experimental results on several public datasets demonstrated that our proposed framework is effective and superior to other existing similar detection networks. Yingbin Zhou, Anwei Luo, Xiangui Kang, Siwei Lyu |
ICIP | 3 |
| 2021 | A framework of generative adversarial networks with novel loss for JPEG restoration and anti-forensics
Jianyuan Wu, Xiangui Kang, Wei Sun 0007 |
Multim. Syst. | 2 |
| 2021 | SpecView: Malware Spectrum Visualization Framework With Singular Spectrum TransformationabstractWith the rapid development of automation tools including polymorphic and metamorphic engines, generic packers, and genetic programming, many variants of malware have emerged, which pose a significant threat to the Internet security. To effectively detect malware variants, researchers have developed visualization-based approaches that can visualize malware adaptations for in-depth malware analysis. However, most existing visualization approaches rely on the binary image of a malware sample, which fail to provide an effective texture feature representation and thus often result in low efficiency in coping with challenging malware samples. In this paper, we proposeSpecView, a malware spectrum visualization framework with singular spectrum transformation. SpecView converts malware binary code into one-dimensional time series spectrum data, and leverages the singular spectrum transformation method to obtain the structural changes preserved in the time series spectrum data. Then, we utilize the particle swarm optimization algorithm to optimize the singular spectrum transformation performance in SpecView. We apply SpecView in the task of malware classification. Extensive experimental results show that SpecView is effective and efficient in malware classification on the Malimg, Malheur, Drebin, and PRAGuard Malgenome Class Encryption datasets, with classification accuracy exceeding 99%, and it can effectively identify malware variants that use evasive techniques such as packer and encryption obfuscation. The proposed method outperforms the state-of-the-art methods on all datasets and the classification accuracy reaches 100% for 5 malware families packed by the UPX packer on the Malimg dataset, as well as 9 malware families that use Class Encryption obfuscation techniques on the PRAGuard Malgenome Class Encryption datasets. Jian Yu 0006, Yuewang He, Qiben Yan 0001, Xiangui Kang |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2020 | Approaching Optimal Embedding In Audio Steganography With GANabstractAudio steganography is a technology that embeds messages into audio without raising any suspicion from hearing it. Current steganography methods are based on heuristic cost designs. In this work, we proposed a framework based on Generative Adversarial Network (GAN) to approach optimal embedding for audio steganography in the temporal domain. This is the first attempt to approach optimal embedding with GAN and automatically learn the embedding probability/cost for audio steganography. The embedding framework consists of three parts: a U-Net based generator, an embedding simulator, and a discriminator. For practical applications, Syndrome-Trellis Coding (STC) is used to generate stego audio with the learned embedding probability. Experimental results on the UME-ERJ and WSJ speech datasets have shown that the proposed framework can automatically learn the adaptive embedding probabilities for audio steganogra- phy and has a considerable advantage in terms of resisting steganalyzers in comparison with the existing conventional method. Huilin Zheng, Xiangui Kang, Yun Q. Shi 0001 |
ICASSP | 3 |
| 2020 | Autoregressive Model Based Smoothing Forensics Of Very Short Speech ClipsabstractSmoothing is a post-processing widely used in speech tampering. Thus, we may determine whether a speech signal is original by smoothing forensics. However, in existing smoothing forensics methods, the detection performance of very short speech clips is much worse than that of long speech clips, and MP3 compression may lead to performance degradation, especially when the length of the smoothing window becomes small. Based on the observation that a very short speech clips can be considered as a stationary autoregressive (AR) process model, we proposed a robust smoothing forensics method of very short speech clips using the AR model coefficients. Experimental results on the TIMIT speech dataset demonstrate that the proposed method significantly outperforms the state-of-the-art method in terms of accuracy and robustness against various MP3 compression. Sanshuai Cui, Enlei Li, Xiangui Kang |
ICME | 3 |
| 2020 | Auto-Generating Neural Networks with Reinforcement Learning for Multi-Purpose Image ForensicsabstractDesigning a forensic convolutional neural network (CNN) is usually based on some ad-hoc intuition and domain knowledge. Many methods to automate neural network design have been proposed for computer vision tasks, but they may not be directly applied to image forensic problems, which tend to detect weak traces signals left by image operations rather than strong image content signals. In this paper, we propose an approach to learn an optimal forensic CNN structure with reinforcement learning for detecting multiple image tampering operations. A learning agent is introduced to select CNN layers sequentially in a limited state-action space using Q-learning with an $\epsilon$-greedy strategy and experience replay. The experiments demonstrate that the auto-generated network performs better than other classic image forensic methods and shows more robustness against JPEG compression. To our knowledge, this is the first attempt to design forensic deep neural networks automatically with reinforcement learning. Yujun Wei, Yifang Chen 0002, Xiangui Kang, Z. Jane Wang 0001, Liang Xiao 0003 |
ICME | 3 |
| 2020 | Electric Network Frequency Based Audio Forensics Using Convolutional Neural Networks
Maoyu Mao, Zhongcheng Xiao, Xiangui Kang, Liang Xiao 0003 |
IFIP Int. Conf. Digital Forensics | 3 |
| 2020 | Reinforcement Learning Aided Network Architecture Generation for JPEG Image SteganalysisabstractThe architectures of convolutional neural networks used in steganalysis have been designed heuristically. In this paper, an automatic Network Architecture Generation algorithm based on reinforcement learning for JPEG image Steganalysis (JS-NAG) has been proposed. Different from the automatic neural network generation methods in computer vision which are based on the strong content signals, steganalysis is based on the weak embedded signals, thus needs specific design. In the proposed method, the agent is trained to sequentially select some high-performing blocks using Q-learning to generate networks. An early stop strategy and a well-designed performance prediction function have been utilized to reduce the search time. To generate the optimal networks, hundreds of networks have been searched and trained on 3 GPUs for 15 days. To further improve the detection accuracy, we make an ensemble classifier out of the generated convolutional neural networks. The experimental results have shown that the proposed method significantly outperforms the current state-of-the-art CNN based methods. Beiling Lu, Liang Xiao 0003, Xiangui Kang, Yun Q. Shi 0001 |
IH&MMSec | 4 |
| 2020 | General and Improved Five-Step Discrete-Time Zeroing Neural Dynamics Solving Linear Time-Varying Matrix Equation with Unknown Transpose
Chaowei Hu, Yunong Zhang, Xiangui Kang |
Neural Process. Lett. | 3 |
| 2020 | An Embedding Cost Learning Framework Using GANabstractSuccessful adaptive steganography has mainly focused on embedding the payload while minimizing an appropriately defined distortion function. The application of deep learning to steganalysis has greatly challenged present adaptive steganographic methods, but has also shown the potential for the improvement of steganography. This paper proposes a distortion function generating a framework for steganography. It has three modules: a generator with a U-Net architecture to translate a cover image into an embedding change probability map, a no-pre-training-required double-tanh function to approximate the optimal embedding simulator while preserving gradient norm during backpropagation in the adversarial training, and an enhanced steganalyzer based on a convolution neural network together with multiple high pass filters as the discriminator. Extensive experimental results on different datasets have shown that the proposed framework outperforms the current state-of-the-art steganographic schemes. Moreover, the adversarial training time is reduced dramatically compared with the GAN-based automatic steganographic distortion learning framework (ASDL-GAN). Danyang Ruan, Jiwu Huang, Xiangui Kang, Yun Q. Shi 0001 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2019 | Towards Automatic Embedding Cost Learning for JPEG SteganographyabstractCurrent mainstream methods for digital image steganography are content adaptive. That is, the secret messages are embedded in the complicated region in the cover image while minimizing the embedding distortion so as to suppress statistical detectability. Since there is already a practical encoding scheme for data embedding near the payload-distortion bound, the design of the embedding cost function becomes a deterministic part in steganography. Unlike the traditional heuristic hand-crafted method, this paper proposes a novel generative adversarial network based framework to automatically learn the embedding cost function for JPEG steganography. The proposed framework consists of a generator, a gradient-descent friendly inverse discrete cosine transformation module, an embedding simulator and a discriminator for steganalysis. Through training the generator and discriminator in alternation, the embedding cost function can finally be obtained by the trained generator. Experimental results demonstrate that our method can automatically learn a reasonable embedding cost function and achieve a satisfying performance. Danyang Ruan, Xiangui Kang, Yun Q. Shi 0001 |
IH&MMSec | 3 |
| 2019 | IStego100K: Large-Scale Image Steganalysis Dataset
Zhongliang Yang, Ke Wang 0033, Yongfeng Huang 0001, Xiangui Kang, Xianfeng Zhao |
IWDW | 5 |
| 2019 | Depthwise Separable Convolutional Neural Network for Image ForensicsabstractGeneral-purpose forensics on small image patches appears to be feasible and important, but in fact poses a challenge due to insufficient statistics. Furthermore, there is a need to develop a forensic approach that can automatically learn effective and robust features related to image forensics with high parameter efficiency. In this paper, we propose a depthwise separable convolutional neural network (CNN) for the simultaneous detection of eleven types of image manipulations in image patches. Different from the previous CNNs based on standard convolution, depthwise separable convolution is introduced in the proposed CNN to adaptively extract forensics-related features from image patches with better parameter efficiency. When compared with four state-of-the-art methods, experiments demonstrate that the proposed CNN architecture can achieve better performance, e.g., the improvement in terms of accuracy in the detection of 32 × 32 images is up to 7.33%. It also achieves significantly better overall performance for different databases and better robustness against JPEG compression. Yifang Chen 0002, Xiangui Kang, Z. Jane Wang 0001 |
VCIP | 3 |
| 2019 | JPEG steganalysis with combined dense connected CNNs and SCA-GFR
Xiangui Kang, Edward K. Wong, Yun Q. Shi 0001 |
Multim. Tools Appl. | 2 |
| 2018 | A Rotation-Invariant Convolutional Neural Network for Image Enhancement ForensicsabstractMany proposed complex convolutional neural network (CNN) models in image forensics are with a large number of parameters, requiring a huge number of training data and having the risk of being overfitting. Considering the desired rotation invariance in the detection of some specific image manipulations, i.e., image enhancement, we propose employing convolutional filters with an isotropic architecture in the CNN model which can significantly reduce the required number of CNN parameters. With the same weights in symmetric positions, the proposed filter can extract rotation-invariant features for image enhancement forensics. Experimental results show that the proposed rotation-invariant CNN models with much less parameters can achieve much better performance, e.g., yielding more than 13% improvement in terms of detection accuracy in Gamma correction forensics. It also achieves significantly better generalization performances on different databases and better robustness against JPEG compression when compared with the popular BayarNet in [16]. Yifang Chen 0002, Zi Xian Lyu, Xiangui Kang, Z. Jane Wang 0001 |
ICASSP | 3 |
| 2018 | Densely Connected Convolutional Neural Network for Multi-purpose Image Forensics under Anti-forensic AttacksabstractMultiple-purpose forensics has been attracting increasing attention worldwide. However, most of the existing methods based on hand-crafted features often require domain knowledge and expensive human labour and their performances can be affected by factors such as image size and JPEG compression. Furthermore, many anti-forensic techniques have been applied in practice, making image authentication more difficult. Therefore, it is of great importance to develop methods that can automatically learn general and robust features for image operation detectors with the capability of countering anti-forensics. In this paper, we propose a new convolutional neural network (CNN) approach for multi-purpose detection of image manipulations under anti-forensic attacks. The dense connectivity pattern, which has better parameter efficiency than the traditional pattern, is explored to strengthen the propagation of general features related to image manipulation detection. When compared with three state-of-the-art methods, experiments demonstrate that the proposed CNN architecture can achieve a better performance (i.e., with a 11% improvement in terms of detection accuracy under anti-forensic attacks). The proposed method can also achieve better robustness against JPEG compression with maximum improvement of 13% on accuracy under low-quality JPEG compression. Yifang Chen 0002, Xiangui Kang, Z. Jane Wang 0001 |
IH&MMSec | 2 |
| 2018 | Comparison of DCT and Gabor Filters in Residual Extraction of CNN Based JPEG Steganalysis
Huilin Zheng, Danyang Ruan, Xiangui Kang, Yun Q. Shi 0001 |
IWDW | 4 |
| 2018 | Countering JPEG anti-forensics based on noise level estimation
Hui Zeng 0002, Xiangui Kang, Siwei Lyu |
Sci. China Inf. Sci. | 3 |
| 2018 | Three-step general discrete-time Zhang neural network design and application to time-variant matrix inversion
Chaowei Hu, Xiangui Kang, Yunong Zhang |
Neurocomputing | 2 |
| 2018 | Robust Electric Network Frequency Estimation with Rank Reduction and Linear PredictionabstractThis article deals with the problem of Electric Network Frequency (ENF) estimation where Signal to Noise Ratio (SNR) is an essential challenge. By exploiting the low-rank structure of the ENF signal from the audio spectrogram, we propose an approach based on robust principle component analysis to get rid of the interference from speech contents and some of the background noise, which in our case can be regarded as sparse in nature. Weighted linear prediction is enforced on the low-rank signal subspace to gain accurate ENF estimation. The performance of the proposed scheme is analyzed and evaluated as a function of SNR, and the Cramér-Rao Lower Bound (CRLB) is approached at an SNR level above -10 dB. Experiments on real datasets have demonstrated the advantages of the proposed method over state-of-the-art work in terms of estimation accuracy. Specifically, the proposed scheme can effectively capture the ENF fluctuations along the time axis using small numbers of signal observations while preserving sufficient frequency precision. Xiaodan Lin, Xiangui Kang |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2017 | Supervised audio tampering detection using an autoregressive modelabstractSplicing, cutting and insertion are the most common operations imposed on audio files when the adversary intends to modify or fabricate the content. The detection of such kinds of tampering is still challenging in real-world applications. In this paper, a generic approach for the detection of audio tampering is proposed via the analysis of electric network frequency (ENF). Based on the fact that tampering with an audio leads to anomalous variations of the underlying ENF signal, a wavelet-filtered ENF signal is generated to highlight the abnormal ENF variations. An autoregressive (AR) model is then fitted to the detail part of the ENF signal and the resulting AR coefficients are employed to train the classifier under a supervised-learning framework. Experimental results show that our proposed method significantly outperforms the state-of-art methods in the context where moderate or high levels of noise are present. Moreover, robustness against MP3 compression can be achieved. Xiaodan Lin, Xiangui Kang |
ICASSP | 2 |
| 2017 | Image Forensics Based on Transfer Learning and Convolutional Neural NetworkabstractThere have been a growing number of interests in using the convolutional neural network(CNN) in image forensics, where some excellent methods have been proposed. Training the randomly initialized model from scratch needs a big amount of training data and computational time. To solve this issue, we present a new method of training an image forensic model using prior knowledge transferred from the existing steganalysis model. We also find out that CNN models tend to show poor performance when tested on a different database. With knowledge transfer, we are able to easily train an excellent model for a new database with a small amount of training data from the new database. Performance of our models are evaluated on Bossbase and BOW by detecting five forensic types, including median filtering, resampling, JPEG compression, contrast enhancement and additive Gaussian noise. Through a series of experiments, we demonstrate that our proposed method is very effective in two scenario mentioned above, and our method based on transfer learning can greatly accelerate the convergence of CNN model. The results of these experiments show that our proposed method can detect five different manipulations with an average accuracy of 97.36%. Yifeng Zhan, Yifang Chen 0002, Xiangui Kang |
IH&MMSec | 4 |
| 2017 | Steganalysis Based on Awareness of Selection-Channel and Deep Learning
Xiangui Kang, Edward K. Wong, Yun Q. Shi 0001 |
IWDW | 3 |
| 2017 | Image splicing localization using PCA-based noise level estimation
Hui Zeng 0002, Yifeng Zhan, Xiangui Kang, Xiaodan Lin |
Multim. Tools Appl. | 3 |
| 2017 | A Framework of Camera Source Identification Bayesian GameabstractImage forensics with the presence of an adversary, such as the interplay between the sensor-based camera source identification (CSI) and the fingerprint-copy attack, has attracted increasing attention recently. In this paper, we propose a framework of CSI game with both complete information and incomplete information. A noise level-based counter anti-forensic method is presented to detect the potential fingerprint-copy attack, and unlike the state-of-the-art countermeasure of the triangle test, it does not need to collect the candidate image set. With the existence of countermeasure, a rational forger needs to balance the tradeoff between synthesizing source information and leaving new detectable evidence of raising the noise level of a forged image. The mixed-strategy other than the sequential-move assumption is adopted to solve the games. The Bayesian game is introduced to address the information asymmetry in practice. The Nash equilibrium of both the complete information game and Bayesian game are theoretically analyzed, and the expected Nash equilibrium payoff of a Bayesian game is obtained. Nash equilibrium receiver operating characteristic curves are adopted to evaluate the detection performance. Simulation results show that the information asymmetry can remarkably affect the final detection performance. To our knowledge, this paper is the first attempt in analyzing a Bayesian forensic game with practical information asymmetry. Hui Zeng 0002, Jingxian Liu, Xiangui Kang, Yun Q. Shi 0001, Z. Jane Wang 0001 |
IEEE Trans. Cybern. | 4 |
| 2016 | Pedestrian detection via a leg-driven physiology frameworkabstractIn this paper, we propose a leg-driven physiology framework for pedestrian detection. The framework is introduced to reduce the search space of candidate regions of pedestrians. Given a set of vertical line segments, we can generate a space of rectangular candidate regions, based on a model of body proportions. The proposed framework can be either integrated with or without learning-based pedestrian detection methods to validate the candidate regions. A symmetry constraint is then applied to validate each candidate region to decrease the false positive rate. The experiment demonstrates the promising results of the proposed method by comparing it with Dalal & Triggs method. For example, rectangular regions detected by the proposed method has much similar area to the ground truth than regions detected by Dalal & Triggs method. Gongbo Liang, Qi Li 0001, Xiangui Kang |
ICIP | 3 |
| 2016 | A Multi-purpose Image Counter-anti-forensic Method Using Convolutional Neural Networks
Yifeng Zhan, Xiangui Kang |
IWDW | 4 |
| 2016 | A Multi-purpose countermeasure against image anti-forensics using autoregressive model
Hui Zeng 0002, Xiangui Kang, Anjie Peng |
Neurocomputing | 2 |
| 2016 | Forensics and counter anti-forensics of video inter-frame forgery
Xiangui Kang, Jingxian Liu, Hongmei Liu 0001, Z. Jane Wang 0001 |
Multim. Tools Appl. | 1 |
| 2016 | Audio Recapture Detection With Convolutional Neural NetworksabstractIn this paper, we investigate how features can be effectively learned by deep neural networks for audio forensic problems. By providing a preliminary feature preprocessing based on electric network frequency (ENF) analysis, we propose a convolutional neural network (CNN) for training and classification of genuine and recaptured audio recordings. Hierarchical representations which contain levels of details of the ENF components are learned from the deep neural networks and can be used for further classification. The proposed method works for small audio clips of 2 second duration, whereas the state of the art may fail with such small audio clips. Experimental results demonstrate that the proposed network yields high detection accuracy with each ENF harmonic component represented as a single-channel input. The performance can be further improved by a combined input representation which incorporates both the fundamental ENF and its harmonics. The convergence property of the network and the effect of using an analysis window with various sizes are also studied. Performance comparison against the support tensor machine demonstrates the advantage of using CNN for the task of audio recapture detection. Moreover, visualization of the intermediate feature maps provides some insight into what the deep neural networks actually learn and how they make decisions. Xiaodan Lin, Jingxian Liu, Xiangui Kang |
IEEE Trans. Multim. | 3 |
| 2015 | Countering anti-forensics of image resamplingabstractImage resampling leaves behind periodical artifacts which are used as fingerprints for the forensics. A knowledgeable anti-forensic method erases such artifacts by irregular sampling. We observe that the irregular sampling followed by interpolation causes changes in local linear correlations, and propose a novel method to detect the anti-forensic method of resampling via partial autocorrelation coefficients. Experimental results on a large set of images show that the proposed method could effectively detect the anti-forensics of resampling with a low dimensional feature set. Anjie Peng, Hui Zeng 0002, Xiaodan Lin, Xiangui Kang |
ICIP | 4 |
| 2015 | Removing camera fingerprint to disguise photograph sourceabstractSensor-based camera source identification (CSI) is believed to be an effective tool for linking a photograph to its source camera. In this paper, we propose a photo response non-uniformity (PRNU) removing attack on CSI. The PRNU fingerprint is a kind of multiplicative noise, and is believed to be difficult to be removed from an image. In this work, both the pattern and magnitude of the PRNU fingerprint of an unaltered image are estimated first, then it is subtracted from the target image. Theoretical analysis and experimental results show that the PRNU fingerprint can be removed using only common processing techniques without introducing visual artefact, therefore CSI can be easily defeated. Hui Zeng 0002, Xiangui Kang, Wenjun Zeng 0001 |
ICIP | 3 |
| 2015 | Median Filtering Forensics Based on Convolutional Neural NetworksabstractMedian filtering detection has recently drawn much attention in image editing and image anti-forensic techniques. Current image median filtering forensics algorithms mainly extract features manually. To deal with the challenge of detecting median filtering from small-size and compressed image blocks, by taking into account of the properties of median filtering, we propose a median filtering detection method based on convolutional neural networks (CNNs), which can automatically learn and obtain features directly from the image. To our best knowledge, this is the first work of applying CNNs in median filtering image forensics. Unlike conventional CNN models, the first layer of our CNN framework is a filter layer that accepts an image as the input and outputs its median filtering residual (MFR). Then, via alternating convolutional layers and pooling layers to learn hierarchical representations, we obtain multiple features for further classification. We test the proposed method on several experiments. The results show that the proposed method achieves significant performance improvements, especially in the cut-and-paste forgery detection. Xiangui Kang, Z. Jane Wang 0001 |
IEEE Signal Process. Lett. | 2 |
| 2014 | Countering anti-forensics of median filteringabstractThe statistical fingerprints left by median filtering can be a valuable clue for image forensics. However, these fingerprints may be maliciously erased by a forger. Recently, a tricky anti-forensic method has been proposed to remove median filtering traces by restoring images' pixel difference distribution. In this paper, we analyze the traces of this anti-forensic technique and propose a novel counter method. The experimental results show that our method could reveal this anti-forensics effectively at low computation load. According to our best knowledge, it's the first work on countering anti-forensics of median filtering. Hui Zeng 0002, Tengfei Qin, Xiangui Kang, Li Liu 0009 |
ICASSP | 3 |
| 2013 | Forensic sensor pattern noise extraction from large image data setabstractThe sensor pattern noise (SPN) can be regarded as the unique identity of a digital camera which is highly useful in digital image forensics [1, 2]. Existing methods [1, 2] which works by denoising each individual natural image often took an investigator a long time and great efforts to collect sufficient photos of diversified enough natural scenes. These processes are hard to repeat or standardized for officially using by an authority. In this work, we create noise image data set by taking photos of random noises displayed on a high definition monitor and propose a homomorphic based SPN extraction method. It offers the forensic researcher a fast way to create a large image data set in a few minutes. And the extraction method only needs to denoise once, which is highly efficient to deal with large numbers of photos. We compared the source camera identification performance of the proposed SPN extraction method to a prior state-of-art with identical experimental settings. The experimental results confirm the effectiveness of the proposed method. Zhenhua Qu, Xiangui Kang, Jiwu Huang, Yinxiang Li |
ICASSP | 2 |
| 2013 | Mixed-strategy Nash equilibrium in the camera source identification gameabstractAlthough sensor pattern noise (SPN) is recognized as a reliable device fingerprint for camera source identification (CSI), this fingerprint could have been forged by anti-forensics. In order to evaluate the performance in the case of both forensic investigator and forger exist, we model this interplay as a camera source identification game. The mixed-strategy Nash equilibrium is introduced to solve this game. The Nash equilibrium receiver operating characteristic (ROC) curves are obtained experimentally. Through our analysis, we are able to determine under which case the CSI result is reliable. Hui Zeng 0002, Xiangui Kang, Jiwu Huang |
ICIP | 2 |
| 2013 | Camera Source Identification Game with Incomplete Information
Hui Zeng 0002, Xiangui Kang |
IWDW | 2 |
| 2013 | Robust Median Filtering Forensics Using an Autoregressive ModelabstractIn order to verify the authenticity of digital images, researchers have begun developing digital forensic techniques to identify image editing. One editing operation that has recently received increased attention is median filtering. While several median filtering detection techniques have recently been developed, their performance is degraded by JPEG compression. These techniques suffer similar degradations in performance when a small window of the image is analyzed, as is done in localized filtering or cut-and-paste detection, rather than the image as a whole. In this paper, we propose a new, robust median filtering forensic technique. It operates by analyzing the statistical properties of the median filter residual (MFR), which we define as the difference between an image in question and a median filtered version of itself. To capture the statistical properties of the MFR, we fit it to an autoregressive (AR) model. We then use the AR coefficients as features for median filter detection. We test the effectiveness of our proposed median filter detection techniques through a series of experiments. These results show that our proposed forensic technique can achieve important performance gains over existing methods, particularly at low false-positive rates, with a very small dimension of features. Xiangui Kang, Matthew C. Stamm, Anjie Peng, K. J. Ray Liu |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2012 | Scalable Lossy Compression for Pixel-Value Encrypted ImagesabstractCompression of encrypted data draws much attention in recent years due to the security concerns in a service oriented environment such as cloud computing. We propose a scalable lossy compression scheme for images having their pixel value encrypted with a standard stream cipher. The encrypted data are simply compressed by transmitting a uniformly sub sampled portion of the encrypted data and some bit-planes of another uniformly sub sampled portion of the encrypted data. With a proposed content adaptive interpolation prediction method with side information, at the receiver side, a decoder performs content adaptive interpolation based on the decrypted partial information, where the received bit-plane information serves as the side information that reflects the image edge information, making the image reconstruction more precise. When more bit-planes are transmitted, higher quality of the decompressed image can be achieved. The experimental results show that our proposed scheme achieves much better performance than the existing lossy compression scheme for pixel value encrypted images, and also similar performance as the state-of-the-art lossy compression for pixel permutation based encrypted images. In addition, our proposed scheme has the following advantages: at the decoder side, no computationally intensive iteration and no additional public orthogonal matrix is needed. It works well for both smooth and texture-rich images. Xiangui Kang, Xianyu Xu, Anjie Peng, Wenjun Zeng 0001 |
DCC | 1 |
| 2012 | A context adaptive predictor of sensor pattern noise for camera source identificationabstractSensor pattern noise (SPN) is a noise-like spread-spectrum signal inherently cast onto every digital image by each imaging device and has been recognised as a reliable device fingerprint for camera source identification (CSI) and image origin verification. It can be estimated as the noise residual between the image content and its denoised version. However, the SPN extracted from a single image can be contaminated largely by image scene because image edge noise is usually much stronger than the SPN. So the identification performance is heavily dependent upon the purity of the estimated SPN, especially for small size images because they have less and weaker SPN. Although there are some existing works dedicated to improving the performance of source camera identification, an effective method to eliminate the contamination of image scene and extract an accurate SPN is currently lacking. In this paper, we will propose an edge adaptive SPN predictor based on context adaptive interpolation (PCAI) to exclude the contamination of image scene. Different from most of the existing methods extracting SPN from wavelet high frequency coefficients, we extract SPN directly from the spatial domain with a pixel-wise adaptive Wiener filter, based on the assumption that the SPN is a white signal. Extensive experiments show that our proposed PCAI method achieves the best receiver operating characteristic (ROC) performance among all of the state-of-the-art CSI schemes on different sizes of images, and has the best performance in resisting JPEG compression (e.g. with a quality factor of 90%) simultaneously. Guangdong Wu, Xiangui Kang, K. J. Ray Liu |
ICIP | 2 |
| 2012 | Robust Median Filtering Detection Based on Filtered Residual
Anjie Peng, Xiangui Kang |
IWDW | 2 |
| 2012 | Enhancing Source Camera Identification Performance With a Camera Reference Phase Sensor Pattern NoiseabstractSensor pattern noise (SPN) extracted from digital images has been proved to be a unique fingerprint of digital cameras. However, SPN can be contaminated largely in the frequency domain by image content and nonunique artefacts of JPEG compression, on-sensor signal transfer, sensor design, color interpolation. The source camera identification (CI) performance based on SPN needs to be improved for small sizes of images and especially in resisting JPEG compression. Because the SPN is modelled as an additive white Gaussian noise (AWGN) in its extraction process from an image, it is reasonable to assume the camera reference SPN to be a white noise signal in order to remove the interference mentioned above. The noise residues (SPN) extracted from the original images are whitened first, then they are averaged to generate the camera reference SPN. Motivated by Goljan 's test statistic peak to correlation energy (PCE), we propose to use correlation to circular correlation norm (CCN) as the test statistic, which can lower the false positive rate to be a half of that with PCE. Theoretical analysis shows that the proposed CI method can remove the interference and raise the CCN value of a positive sample and thus achieve greater CI performance, CCN values of the negative sample class with the proposed method follow the normal distributionN(0,1) and the false positive rate can be calculated. Compared with the existing state of the art on seven cameras, 1400 photos totally (200 for each camera), the experimental results show that the proposed CI method achieves the best receiver operating characteristic (ROC) performance among all CI methods in all cases and especially achieves much better resistance to JPEG compression than all of the existing state-of-the-art CI methods. Xiangui Kang, Yinxiang Li, Zhenhua Qu, Jiwu Huang |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2011 | Geometric Invariant Audio Watermarking Based on an LCM FeatureabstractThe development of a geometric invariant audio watermarking scheme without degrading acoustical quality is challenging work. This paper proposes a multi-bit spread-spectrum audio watermarking scheme based on a geometric invariant log coordinate mapping (LCM) feature. The LCM feature is very robust to audio geometric distortions. The watermark is embedded in the LCM feature, but it is actually embedded in the Fourier coefficients which are mapped to the feature via LCM, so the embedding is actually performed in the DFT domain without interpolation, thus eliminating completely the severe distortion resulted from the non-uniform interpolation mapping. The watermarked audio achieves high auditory quality in both objective and subjective quality assessments. A mixed correlation between the LCM feature and a key-generated PN tracking sequence is proposed to align the log-coordinate mapping, thus synchronizing the watermark efficiently with only one FFT and one IFFT. Both the theoretical analysis and experimental results show that the proposed audio watermarking scheme is not only resilient against common signal processing operations, including low-pass filtering, MP3 recompression, echo addition, volume change, normalization, test functions in the Stirmark benchmark, and DA/AD conversion, but also has conquered the challenging audio geometric distortion and achieves the best robustness against simultaneous geometric distortions, such as pitch invariant time-scale modification (TSM) by ±20%, tempo invariant pitch shifting by 20%, resample TSM with scaling factors between 75% and 140%, and random cropping by 95%. This is mainly contributed by the proposed geometric invariant LCM feature. To our best knowledge, audio watermarking based on LCM has not been reported before. Xiangui Kang, Rui Yang 0006, Jiwu Huang |
IEEE Trans. Multim. | 1 |
| 2010 | Efficient general print-scanning resilient data hiding based on uniform log-polar mappingabstractThis paper proposes an efficient, blind, and robust data hiding scheme which is resilient to both geometric distortion and the general print-scan process, based on a near uniform log-polar mapping (ULPM). In contrast to performing inverse log-polar mapping (a mapping from the log-polar system to the Cartesian system) to the watermark signal or its index as done in the prior works, we apply ULPM to the frequency index (u,v) in the Cartesian system to obtain the discrete log-polar coordinate (l1,l2), then embed one watermark bitw(l1,l2) in the corresponding discrete Fourier transform coefficientc(u,v). This mapping of index from the Cartesian system to the log-polar system but embedding the corresponding watermark directly in the Cartesian domain not only completely removes the interpolation distortion and the interference distortion introduced to the watermark signal as observed in some prior works, but also largely expands the cardinality of watermark in the log-polar mapping domain. Both theoretical analysis and experimental results show that the proposed watermarking scheme achieves excellent robustness to geometric distortion, normal signal processing, and the general print-scan process. Compared to existing watermarking schemes, our algorithm offers significant improvement in terms of robustness against general print-scan, receiver operating characteristic (ROC) performance, and efficiency of blind resynchronization. Xiangui Kang, Jiwu Huang, Wenjun Zeng 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2009 | Non-ambiguity of blind watermarking: a revisit with analytical resolution
Xiangui Kang, Jiwu Huang, Wenjun Zeng 0001, Yun Q. Shi 0001 |
Sci. China Ser. F Inf. Sci. | 1 |
| 2008 | An efficient print-scanning resilient data hiding scheme based on a novel LPMabstractPrint-scan resilient data hiding has not been extensively researched. This paper presents an efficient multi-bit blind watermarking scheme based on a novel Fourier log-polar mapping (LPM). The watermark resynchronization after print-scanning is efficiently solved by an embedded tracking pattern which cannot be removed by template removing attacks and is not detectable for a malicious part. Experimental results show that the proposed watermarking scheme has excellent robustness to print-scanning, cropping, geometric distortion and JPEG compression etc. The obtained success ratios of extraction 60 bits message without error from the combination attack of JPEG compressed with quality factor of 50–100 and then print-scanning were at least 95%. Xiangui Kang, Xiong Zhong, Jiwu Huang, Wenjun Zeng 0001 |
ICIP | 1 |
| 2008 | Robust Audio Watermarking Based on Log-Polar Frequency Index
Rui Yang 0006, Xiangui Kang, Jiwu Huang |
IWDW | 2 |
| 2008 | A Practical Print-and-Scan Resilient Watermarking for High Resolution Images
Xiangui Kang, Philipp Zhang |
IWDW | 2 |
| 2008 | Improving Robustness of Quantization-Based Image Watermarking via Adaptive ReceiverabstractIn this paper, the watermarking channel is modeled as a generalized channel with fading andnonzeromeanadditive noise. In order to improve the watermark robustness against the generalized channel, we present an optimized watermark extraction scheme by using an adaptive receiver for quantization-based watermarking. In the proposed extraction scheme, we adaptively estimate the decision zone of the binary data bits and the quantization step size. A training sequence is embedded into the original image together with the informative watermark. The estimation of the decision zone takes advantage of the response function of the training sequence. Compared to those watermarking schemes without receiver adaptation, the main improvement is the enhanced robustness against median filtering, image intensity Direct Current (DC) change, histogram equalization, color reduction, image intensity linear scaling, image intensity nonlinear scaling such as Gamma correction etc. Xiangui Kang, Jiwu Huang, Wenjun Zeng 0001 |
IEEE Trans. Multim. | 1 |
| 2005 | Multi-band Wavelet Based Digital Watermarking Using Principal Component Analysis
Xiangui Kang, Yun Q. Shi 0001, Jiwu Huang, Wenjun Zeng 0001 |
IWDW | 1 |
| 2005 | SNR-Based Frame-Level Video Bit Rate AllocationabstractQuality fluctuation has a major negative effect on perceptive video quality. In [1], we derived accurate approximations in close-form for the highly nonlinear rate-distortion (R-D) and distortion-quantization (D-Q) relationships, all at the frame-level. Based on the two close forms, we can allocate the bit rate at the frame-level rather easily as far as a target distortion for each frame could be established. In [1], a target distortion was set up for each frame based on a hypothesis that maintaining constant distortion over frames would boast video quality smoothing and extensive experiments showed the constant-distortion bit allocation (CDBA) scheme significantly outperformed the popular constant bit allocation (CBA) scheme in terms of delivered video quality. Maintaining constant distortion is no different from maintaining constant Peak-signal-to-Noise-Ratio (PSNR). In scene changes, however, the picture energy often dramatically changes, producing significantly different Signal-to-Noise-Ratio (SNR) if constant distortion or constant PSNR is maintained. Although computationally more complex, SNR represents a more objective measure than PSNR in assessing picture/video quality. In the paper, an SNR-based bit allocation scheme is developed for video quality smoothing. The algorithm uses a single pass and attempts to maintain constant SNR at the frame level throughout the video sequence. Experimental results on all testing video sequences show that the proposed CSNRBA scheme provides smooth video quality in terms of natural color and sharp objects and silhouette significantly better than both the CBA and CDBA schemes. Xinhua Zhuang, Xiangui Kang, Li Liu 0009, Junqiang Lan, Guang Zhou, Guangdong Wu |
MMSP | 2 |
| 2004 | Improve robustness of image watermarking via adaptive receivingabstractAlmost all the existing popular watermarking schemes model the watermarking channel noises as additive noise with zero-mean. However, our experiments show that this is not reasonable in the case of channel noise introduced by image filtering. This paper presents a new quantization-based watermarking scheme with enhanced robustness via adaptive receiving and turbo coding in addition to other measures. The present algorithm can successfully resist almost all the StirMark testing functions including both common signal processing and geometric distortions in StirMark 4.0 except for random distortion. Xiangui Kang, Jiwu Huang, Yun Q. Shi 0001 |
ICIP | 1 |
| 2003 | Robust Watermarking with Adaptive Receiving
Xiangui Kang, Jiwu Huang, Yun Q. Shi 0001, Jianxiang Zhu |
IWDW | 1 |
| 2003 | A DWT-DFT composite watermarking scheme robust to both affine transform and JPEG compressionabstractRobustness is a crucially important issue in watermarking. Robustness against geometric distortion and JPEG compression at the same time with blind extraction remains especially challenging. A blind discrete wavelet transform-discrete Fourier transform (DWT-DFT) composite image watermarking algorithm that is robust against both affine transformation and JPEG compression is proposed. The algorithm improves robustness by using a new embedding strategy, watermark structure, 2D interleaving, and synchronization technique. A spread-spectrum-based informative watermark with a training sequence is embedded in the coefficients of the LL subband in the DWT domain while a template is embedded in the middle frequency components in the DFT domain. In watermark extraction, we first detect the template in a possibly corrupted watermarked image to obtain the parameters of an affine transform and convert the image back to its original shape. Then, we perform translation registration using the training sequence embedded in the DWT domain, and, finally, extract the informative watermark. Experimental work demonstrates that the proposed algorithm generates a more robust watermark than other reported watermarking algorithms. Specifically it is robust simultaneously against almost all affine transform related testing functions in StirMark 3.1 and JPEG compression with quality factor as low as 10. While the approach is presented for gray-level images, it can also be applied to color images and video sequences. Xiangui Kang, Jiwu Huang, Yun Q. Shi 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2002 | An Image Watermarking Algorithm Robust to Geometric Distortion
Xiangui Kang, Jiwu Huang, Yun Q. Shi 0001 |
IWDW | 1 |