Li Dong 0006

dblp:85/5090-6 · DBLP profile ↗
← Back
57ranked-venue papers
8as first author
45since 2021 · last 2026
0000-0003-2002-8249ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 30 · 8 first-author · 22 since 2021Artificial intelligence and machine learning · 13 · 13 since 2021Security and privacy · 8 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 6 since 2021Computer networks · 4 · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Deferred Poisoning: Making the Model More Vulnerable via Hessian Singularization
abstract
Recent studies have shown that deep learning models are very vulnerable to poisoning attacks. Many defense methods have been proposed to address this issue. However, traditional poisoning attacks are not as threatening as commonly believed. This is because they often cause differences in how the model performs on the training set compared to the validation set. Such inconsistency can alert defenders that their data has been poisoned, allowing them to take the necessary defensive actions. In this paper, we introduce a more threatening type of poisoning attack called the Deferred Poisoning Attack. This new attack allows the model to function normally during the training and validation phases but makes it very sensitive to evasion attacks or even natural noise. We achieve this by ensuring the poisoned model's loss function has a similar value as a normally trained model at each input sample but with a large local curvature. A similar model loss ensures that there is no obvious inconsistency between the training and validation accuracy, demonstrating high stealthiness. On the other hand, the large curvature implies that a small perturbation may cause a significant increase in model loss, leading to substantial performance degradation, which reflects a worse robustness. We fulfill this purpose by making the model have singular Hessian information at the optimal point via our proposed Singularization Regularization term. We have conducted both theoretical and empirical analyses of the proposed method and validated its effectiveness through experiments on image classification tasks. Furthermore, we have confirmed the hazards of this form of poisoning attack under more general scenarios using natural noise, offering a new perspective for research in the field of security.
Yuhao He 0001, Jinyu Tian 0001, Xianwei Zheng, Li Dong 0006, Yuanman Li, Jiantao Zhou 0001
AAAI4
2026 Detecting diffusion-based text tampering in scene images based on multi-scale feature fusion
Qingwen Zhu, Li Dong 0006, Yuanman Li, Haiwei Wu, Yushu Zhang 0001
Expert Syst. Appl.2
2026 Deep Robust Reversible Watermarking
abstract
Robust Reversible Watermarking (RRW) enables perfect recovery of cover images and watermarks in lossless channels while ensuring robust watermark extraction under lossy channels. However, existing RRW methods, mostly non-deep learning-based, suffer from complex designs, high computational costs, and poor robustness limiting their practical applications. To address these issues, this paper proposes Deep Robust Reversible Watermarking (DRRW), a deep learning-based RRW scheme. DRRW introduces an Integer Invertible Watermark Network (iIWN) to achieve an invertible mapping between integer data distributions, fundamentally addressing the limitations of conventional RRW approaches. Unlike traditional RRW methods requiring task-specific designs for different distortions, DRRW adopts an encoder-noise layer-decoder framework, enabling adaptive robustness against various distortions through end-to-end training. During inference, the cover image and watermark are mapped into an overflowed stego image and latent variables. Arithmetic coding efficiently compresses these into a compact bitstream, which is embedded via reversible data hiding to ensure lossless recovery of both the image and watermark. To reduce pixel overflow, we introduce an overflow penalty loss, significantly shortening the auxiliary bitstream while improving both robustness and stego image quality. Additionally, we propose an adaptive weight adjustment strategy that eliminates the need to manually preset the watermark loss weight, ensuring improved training stability and performance. Experiments on multiple datasets demonstrate that the proposed DRRW addresses key challenges in current RRW methods and significantly advances the practical deployment of RRW.
Wei Wang 0077, Chongyang Shi 0001, Li Dong 0006, Yuanman Li, Xiping Hu
IEEE Trans. Pattern Anal. Mach. Intell.4
2026 Universal and Quality-Preserving Watermark Removal Based on Unpaired Learning
abstract
Invisible image watermarking plays a critical role in safeguarding AI-generated images, yet current removal methods face practical limitations in real-world settings. They compromise image quality, are tailored to specific watermarking schemes, or depend on original-watermarked image pairs. These limitations hinder reliable evaluations of watermark robustness. In this work, we propose a universal and quality-preserving watermark removal method based on unpaired learning. Specifically, we implement a three-stage training framework in which we first pre-train the remover to denoise corrupted images. Then the discriminator is trained to distinguish watermarked images from original ones. Finally, the remover and the discriminator are jointly trained in an adversarial manner to further strengthen the watermark elimination capability of the remover while enhancing image quality. Evaluated across various watermarking schemes, our method achieves watermark extraction error rates close to random guessing while maintaining high visual quality. This work reveals that most existing watermarking methods lack sufficient robustness.
Fangjun Yan, Xiaojian Ji, Li Dong 0006, Weiwei Sun 0009, Yuanman Li
IEEE Trans. Dependable Secur. Comput.3
2026 DINVMark: A Deep Invertible Network for Video Watermarking
abstract
With the wide spread of video, video watermarking has become increasingly crucial for copyright protection and content authentication. However, video watermarking still faces numerous challenges. For example, existing methods typically have shortcomings in terms of watermarking capacity and robustness, and there is a lack of specialized noise layer for High Efficiency Video Coding(HEVC) compression. To address these issues, this paper introduces a Deep Invertible Network for Video watermarking (DINVMark) and designs a noise layer to simulate HEVC compression. This approach not only increases watermarking capacity but also enhances robustness. DINVMark employs an Invertible Neural Network (INN), where the encoder and decoder share the same network structure for both watermark embedding and extraction. This shared architecture ensures close coupling between the encoder and decoder, thereby improving the accuracy of the watermark extraction process. Experimental results demonstrate that the proposed scheme significantly enhances watermark robustness, preserves video quality, and substantially increases watermark embedding capacity.
Jianbin Ji, Dawen Xu 0001, Li Dong 0006, Lin Yang 0024, Songhan He
IEEE Trans. Multim.3
2025 Learning Robust Image Watermarking with Lossless Cover Recovery
Wei Wang 0077, Chongyang Shi 0001, Li Dong 0006, Xiping Hu
ICCV4
2025 Diversity-Preserving Robust Watermarking for Diffusion Model Generated Images
abstract
This paper introduces a robust watermarking technique for diffusion model-generated images, which effectively balances watermark robustness, image fidelity, and diversity preservation. Unlike traditional post-hoc approaches, the proposed method embeds watermark information directly into the latent noise of the diffusion model, ensuring seamless integration into the image generation process. This approach minimizes perceptual impact while maintaining high visual quality and diversity of the generated images. Experimental results demonstrate the method’s resilience to various image distortions, including noise, compression, blurring etc., significantly outperforming existing watermarking techniques. The proposed method supports both watermark detection and bit-level extraction, providing a practical solution for secure content protection and traceability in generative models without compromising the integrity of the image generation process.
Linghong Wan, Li Dong 0006, Diqun Yan, Rangding Wang
ISCAS2
2025 An efficient and scalable semi-supervised framework for semantic segmentation
Huazheng Hao, Hui Xiao 0005, Li Dong 0006, Diqun Yan, Dongtai Liang, Jiayan Zhuang, Chengbin Peng 0001
Neural Comput. Appl.4
2025 Multiscale Feature-Guided Adversarial Examples Quality Assessment via Hierarchical Perception of Human Visual System
abstract
Deep neural networks (DNNs) reveal significant robustness deficiencies due to their susceptibility to being misled by small and imperceptible adversarial examples, thus it is crucial to improve the robustness of DNNs against such harmful perturbations. The current$L_{p}$specification ignores differences in human visual perception when measuring similarity, and most existing image quality assessment (IQA) methods and adversarial example datasets lack subjective scores for evaluation. In this paper, we construct a new database of adversarial examples, called the AED, which contains 35 original images, 1050 adversarial examples, and the corresponding subjective scores of adversarial examples. Then, a novel full-reference IQA model for the quality evaluation of the adversarial examples is proposed by taking into full consideration the hierarchical perception of human visual system (HVS) and the outstanding capabilities of the multi-scale feature extraction network in feature extraction. Specifically, a feature encoding network that uses continuous convolution layers to pre-extract features and expand the receptive field of the image is employed. To simulate the HVS hierarchical perception, the features of different scales are further obtained by designing a multi-scale feature extraction network. The structural similarity scores of the feature maps at different scales are calculated for jointly arriving at the final IQA score of the adversarial examples. Experimental results have demonstrated that our proposed model is closer to the perception of HVS in small imperceptible distortions evaluation of adversarial examples compared with other classical and state-of-the-art models.
Wenying Wen, Minghui Huang, Li Dong 0006, Yushu Zhang 0001, Yuming Fang 0001
IEEE Trans. Big Data3
2025 WaveRecovery: Screen-Shooting Watermarking Based on Wavelet and Recovery
abstract
The demand for resilient watermarking technology in the context of the screen-shooting scenario is steadily on the rise. The principal objective of this technique is to embed messages into the cover image, with the ability to effectively recover the message from the screen-captured image at the extraction end. However, current watermarking methods result in low visual quality watermarked images and are insufficiently robust in screen-shooting scenarios. This is mainly because they only utilize spatial domain information during embedding, and they do not consider the impact of noise that introduced during screen capturing. This paper introduces an innovative network framework, including the wavelet domain concatenation and recovery mechanism, to overcome the dual challenges encountered in robust watermarking, namely visual fidelity and robustness. For fidelity, we present a cascade network operating in the wavelet domain. This network excel at detecting watermark information in the wavelet domain. This capability makes it more sensitive to high and low-frequency details. Discrete wavelet transform can make CNN focus on different frequency characteristics, and the use of discrete inverse wavelet transform in upsampling can make the information high fidelity. As a result, it can more accurately identify and preserve critical visual details in this frequency domain, leading to an overall enhancement in visual quality. For robustness, a recovery network is specifically designed to mitigate the influence of noise introduced during screen-shooting on watermark information extraction. Experimental validation of our proposed method substantiates its effectiveness in significantly enhancing the visual quality and the accuracy of the watermarked images.
Linbo Fu, Xin Liao 0001, Jinlin Guo, Li Dong 0006, Zheng Qin 0001
IEEE Trans. Circuits Syst. Video Technol.4
2025 Mixed-Bit Sampling Marking: Toward Unifying Document Authentication in Copy-Sensitive Graphical Codes
abstract
Combating counterfeit products is crucial for maintaining a healthy market. Recently, Copy Sensitive Graphical Codes (CSGC) have garnered significant attention due to their high sensitivity to illegal physical copying. Copy Detection Patterns (CDP) and Two-Level QR Codes (2LQR code) are two representative methods. CDP offers high efficiency and low cost, enabling use in document authentication and product anti-counterfeiting, and has achieved broad commercial adoption. In contrast, 2LQR code, as a consumer-grade document authentication solution, provides additional private message sharing functionalities. We observe that both the CDP and 2LQR code can be synthesized using textured patterns. To this end, we propose a flexible framework that integrates the stochastic anti-counterfeiting properties of CDP with the private message sharing of 2LQR code. Specifically, we model CDP as a random noise image composed of multiple textured patterns similar to those in 2LQR code, where each pattern represents an informative digit. Thus, both codes can be generated through textured pattern design. We formulate this as a constrained optimization framework called Mixed-Bit Sampling Marking (MSM). The objective incorporates white pixel ratio and spatial randomness, with constraints defined by a flexible modulation function (e.g., DCT or Pearson similarity), customizable to user needs. A two-step sampling algorithm solves the optimization. We demonstrate CDP and 2LQR codes generated via MSM and validate their ability to inherit advantages from both approaches. Experiments show that MSM-generated texture patterns effectively synthesize both CDPs and 2LQR codes, preserving their advantages while offering a novel, flexible solution for document authentication.
Li Dong 0006, Wei Wang 0077, Rangding Wang, Weiwei Sun 0009, Yushu Zhang 0001, Jiantao Zhou 0001
IEEE Trans. Inf. Forensics Secur.2
2025 Image Copy-Move Forgery Detection via Deep PatchMatch and Pairwise Ranking Learning
abstract
Recent advances in deep learning algorithms have shown impressive progress in image copy-move forgery detection (CMFD). However, these algorithms lack generalizability in practical scenarios where the copied regions are not present in the training images, or the cloned regions are part of the background. Additionally, these algorithms utilize convolution operations to distinguish source and target regions, leading to unsatisfactory results when the target regions blend well with the background. To address these limitations, this study proposes a novel end-to-end CMFD framework that integrates the strengths of conventional and deep learning methods. Specifically, the study develops a deep cross-scale PatchMatch (PM) method that is customized for CMFD to locate copy-move regions. Unlike existing deep models, our approach utilizes features extracted from high-resolution scales to seek explicit and reliable point-to-point matching between source and target regions. Furthermore, we propose a novel pairwise rank learning framework to separate source and target regions. By leveraging the strong prior of point-to-point matches, the framework can identify subtle differences and effectively discriminate between source and target regions, even when the target regions blend well with the background. Our framework is fully differentiable and can be trained end-to-end. Comprehensive experimental results highlight the remarkable generalizability of our scheme across various copy-move scenarios, significantly outperforming existing methods.
Yuanman Li, Yingjie He 0003, Changsheng Chen 0001, Li Dong 0006, Bin Li 0011, Jiantao Zhou 0001, Xia Li 0006
IEEE Trans. Image Process.4
2024 Breaking Speaker Recognition with Paddingback
abstract
Machine Learning as a Service (MLaaS) has gained popularity due to advancements in Deep Neural Networks (DNNs). However, untrusted third-party platforms have raised concerns about AI security, particularly in backdoor attacks. Recent research has shown that speech backdoors can utilize transformations as triggers, similar to image backdoors. However, human ears can easily be aware of these transformations, leading to suspicion. In this paper, we propose PaddingBack, an inaudible backdoor attack that utilizes malicious operations to generate poisoned samples, rendering them indistinguishable from clean ones. Instead of using external perturbations as triggers, we exploit the widely-used speech signal operation, padding, to break speaker recognition systems. Experimental results demonstrate the effectiveness of our method, achieving a significant attack success rate while retaining benign accuracy. Furthermore, Padding-Back demonstrates the ability to resist defense methods and maintain its stealthiness against human perception.
Zhe Ye 0001, Diqun Yan, Li Dong 0006, Kailai Shen
ICASSP3
2024 A multi-view consistency framework with semi-supervised domain adaptation
Yuting Hong, Li Dong 0006, Xiaojie Qiu, Hui Xiao 0005, Baochen Yao, Siming Zheng, Chengbin Peng 0001
Eng. Appl. Artif. Intell.2
2024 C²F²: Cross-Task Cross-Domain Feature Fusion for Semi-Supervised Change Detection
abstract
Semi-supervised learning for change detection (CD), which significantly reduces the labor costs associated with data annotation, has recently garnered substantial attention. In this study, we propose to enhance traditional semi-supervised learning frameworks by leveraging cross-task cross-domain (CTCD) models, which generate complementary features that differ from standard hidden features. The procedure is as follows. First, the standard features obtained from a traditional encoding–decoding structure are fused with attention-augmented complementary features. Second, a secondary decoder maps the fused heterogeneous features into the label space to obtain high-quality pseudo-labels, offering more precise guidance for semi-supervised learning on traditional structures. This approach improves pseudo-labels by leveraging the strength of CTCD models, including large pretrained models, to enhance the semi-supervised learning process of domain-specific and task-specific models. Experimental results on benchmark datasets demonstrate that our proposed approach surpasses state-of-the-art methods.
Dongjie Zhang 0001, Yuting Hong, Xiaojie Qiu, Li Dong 0006, Diqun Yan, Chengbin Peng 0001
IEEE Geosci. Remote. Sens. Lett.4
2024 Mixed-Bit Sampling Graphic: When Watermarking Meets Copy Detection Pattern
abstract
Copy Detection Pattern (CDP) is a high-density random noise-alike image that exhibits a different noise pattern after physical copying, and is thus treated as a promising anti-counterfeiting solution. However, CDP cannot convey any message, and it is often used in combination with additional carriers, such as QR codes. In this letter, we take the first step towards extending CDP with watermarking functionality. Specifically, we devise a scheme called Mixed-bit Sampling Graphic (MSG), which could realize invisible watermarking and anti-counterfeiting simultaneously. Compared with conventional CDP, the noise pattern generation of MSG is controlled by the portions of sampling over two bit templates. We formulate this mixed-bit sampling process as an optimization problem and solve it using a block coordinate descent sampling algorithm. Experimental results validate that the proposed MSG can effectively communicate watermark bits while retaining the anti-counterfeiting capability of CDP.
Li Dong 0006, Rangding Wang, Diqun Yan, Chengbin Peng 0001
IEEE Signal Process. Lett.2
2024 Transformer-Based Image Inpainting Detection via Label Decoupling and Constrained Adversarial Training
abstract
Image inpainting based on generative adversarial networks (GANs) has achieved great success in producing visually plausible images and plays an important role in many real tasks. However, the techniques of image inpainting might also be maliciously used, e.g., altering or removing interesting objects to report fake news. Despite the promising performance of recently developed inpainting detection algorithms, they are built on convolutional neural networks (CNNs) with limited receptive fields. Consequently, they fail to fully capture the disparity between the inpainted regions and untouched regions and thus are ineffective in obtaining fine-grained detection results. In this work, we develop a new image inpainting detection approach. First, we propose a locally enhanced transformer architecture tailored for image inpainting detection. Unlike previous CNN-based methods, our approach leverages both the short-range and long-range dependencies of pixels, enabling the learning of diverse statistical behaviors of inpainted and untouched regions. Second, to mitigate the distraction caused by near-edge pixels with a mixed nature during training, we propose decoupling the label into a body map and a soft-edge map, and then a cross-modality attention module is designed to propagate their information interactively. It demonstrates that our decoupling strategy outperforms the conventional edge supervision in enhancing detection accuracy. Finally, we devise a constrained adversarial training methodology in consideration of the confrontational generation procedure of deep image inpainting methods. It shows that our constrained adversarial training further enhances the detection performance by adaptively introducing interference noise in the inpainted regions. Extensive experiments validate the superiority of our scheme compared to existing CNN-based methods, showcasing its desirable detection generalizability for both deep inpainting and traditional inpainting algorithms.
Yuanman Li, Liangpei Hu, Li Dong 0006, Haiwei Wu, Jinyu Tian 0001, Jiantao Zhou 0001, Xia Li 0006
IEEE Trans. Circuits Syst. Video Technol.3
2024 Uncertainty-Guided Contrastive Learning for Weakly Supervised Point Cloud Segmentation
abstract
Three-dimensional point cloud data are widely used in many fields, as they can be easily obtained and contain rich semantic information. Recently, weakly supervised segmentation has attracted lots of attention, because it only requires very few labels, thus reducing time-consuming and expensive data annotation efforts for huge amounts of point cloud data. The existing approaches typically adopt softmax scores from the last layer as the confidence for selecting high-confident point predictions. However, such approaches can ignore the potential value of a large number of low-confidence point predictions under traditional metrics. In this work, we propose an uncertainty-guided contrastive learning (UCL) framework for weakly supervised point cloud segmentation. A novel uncertainty metric based on prototype entropy (PE) is presented to estimate the reliability of model predictions. With this metric, we propose a negative contrastive learning module exploiting negative pseudo-labels of predictions with low reliability and an active contrastive learning module enhancing feature learning of segmentation models by predictions with high reliability. We also propose a generic multiscale feature perturbation method to expand a wider perturbation space. Extensive experimental results on indoor and outdoor point cloud datasets demonstrate that the proposed method achieves competitive performance.
Baochen Yao, Li Dong 0006, Xiaojie Qiu, Kangkang Song, Diqun Yan, Chengbin Peng 0001
IEEE Trans. Geosci. Remote. Sens.2
2024 Multi-Level Label Correction by Distilling Proximate Patterns for Semi-Supervised Semantic Segmentation
abstract
Semi-supervised semantic segmentation relieves the reliance on large-scale labeled data by leveraging unlabeled data. Recent semi-supervised semantic segmentation approaches mainly resort to pseudo-labeling methods to exploit unlabeled data. However, unreliable pseudo-labeling can undermine the semi-supervision processes. In this paper, we propose an algorithm called Multi-Level Label Correction (MLLC), which aims to use graph neural networks to capture structural relationships in Semantic-Level Graphs (SLGs) and Class-Level Graphs (CLGs) to rectify erroneous pseudo-labels. Specifically, SLGs represent semantic affinities between pairs of pixel features, and CLGs describe classification consistencies between pairs of pixel labels. With the support of proximate pattern information from graphs, MLLC can rectify incorrectly predicted pseudo-labels and can facilitate discriminative feature representations. We design an end-to-end network to train and perform this effective label corrections mechanism. Experiments demonstrate that MLLC can significantly improve supervised baselines and outperforms state-of-the-art approaches in different scenarios on Cityscapes and PASCAL VOC 2012 datasets. Specifically, MLLC improves the supervised baseline by at least 5% and 2% with DeepLabV2 and DeepLabV3+ respectively under different partition protocols.
Hui Xiao 0005, Yuting Hong, Li Dong 0006, Diqun Yan, Jiayan Zhuang, Dongtai Liang, Chengbin Peng 0001
IEEE Trans. Multim.3
2024 Centralized Error Distribution-Preserving Adaptive Steganography for HEVC
abstract
Distortion compensation method is a common way to cope with the distortion drift problem in coefficient domain HEVC steganography. However, it will leave obvious steganographic traces called centralized error (CER). The current coefficient domain HEVC steganography is fragile to CER-based steganalysis. In this article, a novel adaptive HEVC steganography that can resist CER-based steganalysis is proposed. First, the difference of CER between H.264/AVC and HEVC is introduced, and the CER feature in HEVC is re-modeled. Then, from two aspects of overall average distribution and single-frame distribution, we conclude that there is a strong correlation among four components of the CER feature. Last, an adaptive cost function is proposed by maintaining one component distribution to resist steganalysis. Experimental results show that the proposed cost function can effectively improve the security compared with other coefficient-based HEVC steganography. In addition, the proposed steganography outperforms other HEVC steganography in visual quality and bit rate increase.
Lin Yang 0024, Rangding Wang, Dawen Xu 0001, Li Dong 0006, Songhan He
IEEE Trans. Multim.4
2024 Amount-Based Covert Communication Over Blockchain
abstract
Recent years have witnessed the booming growth of 5G and 6G technology, which has brought unprecedented massive data transmission, causing severe privacy issues. However, traditional information encryption and multimedia covert communication fail to protect the identities of communication parties and the originality of messages. The emergence of blockchain provides a promising solution to solve these problems. Its anonymity manages to hide the identities of communication parties, and immutability ensures the message is undestroyable. However, the existing blockchain-based covert communication schemes suffer the issues of low embedding capacity and high time cost. In this paper, an amount-based covert communication scheme over the blockchain is proposed, in which a unique coding method is devised for hiding messages into transaction amounts to improve the embedding capacity. Compared with existing address-based methods, the proposed scheme can apply any address and reduce the time of obtaining special addresses. Besides, we innovate the way to prove the concealment by calculating the relative entropy of the transaction amount between Bitcoin and the proposed scheme. The security of our method is demonstrated by comparing the probability of attackers acquiring secret messages under different adversary capabilities. The experimental results verify that the proposed approach outperforms the existing schemes regarding embedding capacity, time costs, number of transactions, concealment, and security.
Yang Tian 0004, Xin Liao 0001, Li Dong 0006, Yang Xu 0013, Hongbo Jiang 0001
IEEE Trans. Netw. Serv. Manag.3
2023 SQAT-LD: SPeech Quality Assessment Transformer Utilizing Listener Dependent Modeling for Zero-Shot Out-of-Domain MOS Prediction
abstract
In this paper, we propose the speech quality assessment transformer utilizing listener dependent modeling (SQAT-LD) mean opinion score (MOS) prediction system, which was submitted to the 2023 VoiceMOS Challenge. The system is based on a combination of self-supervised learning (SSL) models and listener-dependent modeling. Due to this challenge’s emphasis on real-world and challenging zero-shot out-of-domain MOS prediction in three different voice evaluation scenarios, we specifically designed a two-branch module to predict scores and weights for each frame, aiming to achieve better generalization. In the challenge, our system achieved fourth place in Track 1a, second place in Track 1b and first place in Track 2. Additionally, we conducted an ablation study to investigate the effectiveness of our proposed method.
Kailai Shen, Diqun Yan, Li Dong 0006, Ying Ren, Xiaoxun Wu
ASRU3
2023 Synthetic Feature Assessment for Zero-Shot Object Detection
abstract
Zero-shot object detection aims to simultaneously identify and localize classes that were not presented during training. Many generative model-based methods have shown promising performance by synthesizing the visual features of unseen classes from semantic embeddings. However, these synthetic features are inevitably of varied quality, which may be far from the ground truth. It degrades the performance of trained unseen classifier. Instead of tweaking the generative model, a new idea of feature quality assessment is proposed to utilize both the good and bad features to optimize the classifier in the right direction. Moreover, contrastive learning is also introduced to enhance the feature uniqueness between unseen and seen classes, which helps the feature assessment implicitly. To demonstrate the effectiveness of the proposed algorithm, comprehensive experiments are conducted on the MS COCO dataset and PASCAL VOC dataset, the state-of-the-art performance is achieved. Our code is available at: https://github.com/Dai1029/SFA-ZSD.
Xinmiao Dai, Chong Wang 0001, Haohe Li, Sunqi Lin, Li Dong 0006, Jiafei Wu, Jun Wang 0071
ICME5
2023 A Pseudo-Dual Self-Rectification Framework for Semantic Segmentation
abstract
Semantic segmentation has achieved remarkable success in various applications. However, the training process for such techniques necessitates a significant amount of labeled data. Although semi-supervised frameworks can alleviate this issue, traditional approaches typically require multiple baseline models to form a dual model. To allow a semi-supervised semantic segmentation framework to be used in robotic systems with precious computation and memory resources, we propose a framework utilizing a single baseline model only. The overall framework is composed of three parts: an encoder, a shallow decoder, and a deep decoder. It distills knowledge from the ensemble of two decoders to improve the encoder, which can implicitly form a pseudo-dual model. It also calculates class-wise likelihoods according to the similarity between features and class prototypes learned from different decoders and rectifies low-confidence pseudo-labels. Our framework outperforms state-of-the-art frameworks on benchmark datasets with a significant amount of decrease in using computing resources.
Huazheng Hao, Hui Xiao 0005, Li Dong 0006, Diqun Yan, Dongtai Liang, Jiayan Zhuang, Chengbin Peng 0001
ICME3
2023 Gradient Sign Inversion: Making an Adversarial Attack a Good Defense
abstract
Deep neural networks have been proven vulnerable to deliberately crafted adversarial example, which cause serious safety and security concerns. Many defense approaches were proposed to resist such threats. However, existing defenses such as pre-compression or adversarial training would degrade the model performance on clean images or incur heavy computational costs. In this work, we propose a plug-and-play defensive module Gradient Sign Inversion (GSI) to defend gradient-based attack. Essentially, GSI attempts to inverse the direction of the backpropagated gradient for the victim model, disturbing the adversarial example generation of the attacking while retaining the performance of the vanilla network on genuine inputs. Specifically, an additive model based on periodic trigonometric function is established by investigating the necessary conditions that a suitable defensive module should have. By enforcing constraints on the defensive module, the parameters of GSI are determined, accompanied by a theoretical justification. Interestingly, we observe that the proposed GSI not only prevents the gradient-based adversarial attack, but can even improve the confidence of the ground-truth label when initiating an attack, making the attack betray as a defense. Source code is publicly available at https://github.com/JidaDiao/GSI.
Xiaojian Ji, Li Dong 0006, Rangding Wang, Diqun Yan, Yang Yin, Jinyu Tian 0001
IJCNN2
2023 Fake the Real: Backdoor Attack on Deep Speech Classification via Voice Conversion
abstract
Deep speech classification has achieved tremendous success and greatly promoted the emergence of many real-world applications. However, backdoor attacks present a new security threat to it, particularly with untrustworthy third-party platforms, as pre-defined triggers set by the attacker can activate the backdoor. Most of the triggers in existing speech backdoor attacks are sample-agnostic, and even if the triggers are designed to be unnoticeable, they can still be audible. This work explores a backdoor attack that utilizes sample-specific triggers based on voice conversion. Specifically, we adopt a pre-trained voice conversion model to generate the trigger, ensuring that the poisoned samples does not introduce any additional audible noise. Extensive experiments on two speech classification tasks demonstrate the effectiveness of our attack. Furthermore, we analyzed the specific scenarios that activated the proposed backdoor and verified its resistance against fine-tuning.
Zhe Ye 0001, Terui Mao, Li Dong 0006, Diqun Yan
INTERSPEECH3
2023 Universal Defensive Underpainting Patch: Making Your Text Invisible to Optical Character Recognition
abstract
Optical Character Recognition (OCR) enables automatic text extraction from scanned or digitized text images, but it also makes it easy to pirate valuable or sensitive text from these images. Previous methods to prevent OCR piracy by distorting characters in text images are impractical in real-world scenarios, as pirates can capture arbitrary portions of the text images, rendering the defenses ineffective. In this work, we propose a novel and effective defense mechanism termed the Universal Defensive Underpainting Patch (UDUP) that modifies the underpainting of text images instead of the characters. UDUP is created through an iterative optimization process to craft a small, fixed-size defensive patch that can generate non-overlapping underpainting for text images of any size. Experimental results show that UDUP effectively defends against unauthorized OCR under the setting of any screenshot range or complex image background. It is agnostic to the content, size, colors, and languages of characters, and is robust to typical image operations such as scaling and compressing. In addition, the transferability of UDUP is demonstrated by evading several off-the-shelf OCRs. The code is available at https://github.com/QRICKDD/UDUP.
Jiacheng Deng 0001, Li Dong 0006, Diqun Yan, Rangding Wang, Dengpan Ye, Lingchen Zhao, Jinyu Tian 0001
ACM Multimedia2
2023 Adaptive-SpEx: Local and Global Perceptual Modeling with Speaker Adaptation for Target Speaker Extraction
abstract
Target speaker extraction aims to extract a target speaker's speech from a multi-talker environment with the help of the target speaker's reference speech. However, the simple fusion of different features and local perceptual modeling lead to limited extraction performance. In this work, we propose a new speaker extraction model called Adaptive-SpEx. The correlation between mixed speech features and speaker embedding is fully exploited, and a dual-path structure is used for local and global perceptual modeling. We evaluate the model on the WSJ0-2mix-extr dataset in terms of its ability to reconstruct signal quality. Experimental results show that the proposed model outperforms other baseline systems on WSJ0-2mix-extr and achieves better generalizability on the Libri-2talker dataset. Furthermore, the proposed model can significantly reduce the word error rate of mixed speech on speech recognition from 79.49% to 32.73%.
Xianbo Xu, Diqun Yan, Li Dong 0006
SMC3
2023 Corrigendum to "Semi-supervised semantic segmentation with cross teacher training" [Neurocomputing 508 (2022) 36-46]
Hui Xiao 0005, Li Dong 0006, Shuibo Fu, Diqun Yan, Kangkang Song, Chengbin Peng 0001
Neurocomputing2
2023 Semi-supervised learning with pseudo-negative labels for image classification
Hui Xiao 0005, Huazheng Hao, Li Dong 0006, Xiaojie Qiu, Chengbin Peng 0001
Knowl. Based Syst.4
2023 Imperceptible adversarial audio steganography based on psychoacoustic model
Lang Chen, Rangding Wang, Li Dong 0006, Diqun Yan
Multim. Tools Appl.3
2023 Stealthy Backdoor Attack Against Speaker Recognition Using Phase-Injection Hidden Trigger
abstract
Deep learning has achieved significant breakthroughs in speaker recognition, driven by continual advancements in foundation models. However, malicious third-party platforms have introduced a severe security concern through backdoor attacks, in which attackers can manipulate a model to output a specific label by implanting a trigger. Existing speech backdoor attack methods typically utilize fixed and unnoticeable perturbations as triggers, but these may still be audible and thus detected during training and inference stages. To overcome this limitation, we propose a novel backdoor attack paradigm (PhaseBack) injecting triggers in the phase spectrum. PhaseBack exhibits sufficient stealth by leveraging the fact that the human ear is insensitive to phase information. Besides, injecting partial perturbations in the frequency domain results in global perturbations throughout the time domain, making the attack more effective. Extensive experiments on the Voxceleb1 dataset demonstrate the effectiveness and stealthiness of PhaseBack. Moreover, it has strong resistance to bypass several defense methods.
Zhe Ye 0001, Diqun Yan, Li Dong 0006, Jiacheng Deng 0001, Shui Yu 0001
IEEE Signal Process. Lett.3
2022 Synchronous Bi-directional Pedestrian Trajectory Prediction with Error Compensation
Ce Xie, Yuanman Li, Rongqin Liang, Li Dong 0006, Xia Li 0006
ACCV (6)4
2022 Watermark-Preserving Keypoint Enhancement for Screen-Shooting Resilient Watermarking
abstract
Screen-shooting resilient (SSR) watermark is a special kind of robust watermarking. One can extract the watermark message even the embedded image communicates via a physical screen to the camera channel. The keypoint-based SSR watermarking is one promising solution to realize such screen-to-camera communication. The enhanced keypoints were used to locate the embedding region and then perform watermark embedding. However, the keypoint-based SSR watermarking treats the critical two steps, keypoint enhancement and watermark embedding, independently, neglecting their inter-play. This work proposes a watermark-preserving keypoint enhancement algorithm for SSR watermarking. Specifically, we resort to a convex constrained optimization framework to unify keypoint enhancement and watermark embedding. Multiple constraints are imposed to simultaneously ensure the watermark validity and blind synchronization of embedding regions. Our method enables jointly optimizing the watermarking distortion and keypoint enhancement. The proposed method achieves superior watermark extraction accuracy while retaining better watermarked image quality when compared with previous works.
Li Dong 0006, Chengbin Peng 0001, Yuanman Li, Weiwei Sun 0009
ICME1
2022 Physical Anti-copying Semi-robust Random Watermarking for QR Code
Li Dong 0006, Rangding Wang, Diqun Yan, Weiwei Sun 0009, Hang-Yu Fan
IWDW2
2022 High-Capacity Adaptive Steganography Based on Transform Coefficient for HEVC
Lin Yang 0024, Rangding Wang, Dawen Xu 0001, Li Dong 0006, Songhan He, Fang Liu 0002
IWDW4
2022 On Attacking Deep Image Quality Evaluator Via Spatial Transform
abstract
Adversarial examples fool the neural networks by adding slightly-perturbed noise to the original image, which barriers the usability of deep models. Most of the works focused on the adversarial attack on the classification task. We, in this work, attempt to develop an adversarial example generation method for attacking neural-based image quality assessment (IQA). Specifically, instead of employing conventional additive adversarial noise generation methods, we propose an image content deformation approach, avoiding the loss of adversarial noise after compression. The deformation component is designed as neural layers. The given image is firstly deformed and then undergoes compression; an existing IQA evaluates the compressed image. The deformation layers are trained by back-propagating the differences between the targeted IQA score and the originally-evaluated one. Experimental results demonstrate that the proposed method can produce compression-resistant adversarial images for image quality evaluators. The generated adversarial examples could effectively attack neural-based image quality evaluators with less distortion. The data and code of this work are available at https://github.com/luning409/Attack_IQA.
Li Dong 0006, Diqun Yan, Xianliang Jiang
SMC2
2022 Robust Document Image Forgery Localization Against Image Blending
abstract
Digital documents, as a twin of hard copy, are increasingly being used as credible evidence. Unfortunately, digital document images easily suffer forgery or malicious manipulation, with the availability of sophisticated image editing tools. To verify and detect the possible forgeries for a given document, a number of forensic schemes have been developed. However, in the real-world scenario, the doctored image could be further processed or transmitted over a channel with unknown distortion, which dramatically degrade the forgery detection performance. In this work, we make the first step towards designing a robust document image forgery localization against image blending. Specifically, we propose an encoder-decoder neural network architecture consisting of three modules. The first module is responsible for capturing the multi-scale features from the high-level feature maps, and the remaining two attention-based modules aim to extract low-level local features and high-level global features. For training the model, we construct a dedicated forgery document database processed by several recent image blending procedures. Extensive experiments demonstrate the effectiveness and superiority of the proposed method in detecting the forgery that undergoes image blending. The source code, models and the constructed image dataset are publicly available at https://github.com/lwp0201/Image-Forgery-Localization-Against-Image-Blending.
Weipeng Liang, Li Dong 0006, Rangding Wang, Diqun Yan, Yuanman Li
TrustCom2
2022 Semi-supervised semantic segmentation with cross teacher training
Hui Xiao 0005, Li Dong 0006, Shuibo Fu, Diqun Yan, Kangkang Song, Chengbin Peng 0001
Neurocomputing2
2022 Decision-Based Attack to Speaker Recognition System via Local Low-Frequency Perturbation
abstract
Despite neural network-based speaker recognition systems (SRS) have enjoyed significant success, they are proved to be quite vulnerable to adversarial examples. In practice, the SRS model parameters are not always available. Attackers have to probe the model only via querying, and such decision-based attacking merely relies on the output label is quite challenging. This letter proposes a two-step query-efficient decision-based attack based on local low-frequency perturbation. Specifically, instead of imposing perturbation on the entire audio sample, a local attacking region is firstly sought, confining the perturbed distortion to a local region. Second, considering that the majority of energy concentrates on the low-frequency bands, the proposed method suggests performing perturbation generation in the low-frequency domain. Experimental results demonstrate that, compared with the recent methods, our method could implement target attacking to SRS with a higher attacking success rate, at the cost of much lower queries and adversarial perturbation.
Jiacheng Deng 0001, Li Dong 0006, Rangding Wang, Rui Yang 0006, Diqun Yan
IEEE Signal Process. Lett.2
2022 Robust Matrix Factorization via Minimum Weighted Error Entropy Criterion
abstract
Learning the intrinsic low-dimensional subspace from high-dimensional data is a key step for many social systems of artificial intelligence. In practical scenarios, the observed data are usually corrupted by many types of noise, which brings a great challenge for social systems to analyze data. As a commonly utilized subspace learning technique, robust low-rank matrix factorization (LRMF) focuses on recovering the underlying subspaces in a noisy environment. However, most of the existing approaches simply assume that the noise contaminating the data is independent identically distributed (i.i.d.), such as Gaussian and Laplacian noises. This assumption, though greatly simplifies the underlying learning problem, may not hold for more complex non-i.i.d. noise widely existed in social systems. In this work, we suggest a robust LRMF approach to deal with various types of noise in a unified manner. Different from traditional algorithms, noise in our framework is modeled using an independent and piecewise identically distributed (i.p.i.d.) source, which employs a collection of distributions, instead of a single one to characterize the statistical behavior of the underlying noise. Assisted by the generic noise model, we then design a robust LRMF algorithm under the information-theoretic learning (ITL) framework through a new minimization criterion. By adopting the half-quadratic optimization paradigm, we further deliver an optimization strategy for our proposed method. Experimental results on both synthetic and real data are provided to demonstrate the superiority of our proposed scheme.
Yuanman Li, Jiantao Zhou 0001, Junyang Chen 0001, Jinyu Tian 0001, Li Dong 0006, Xia Li 0006
IEEE Trans. Comput. Soc. Syst.5
2021 Fast speech adversarial example generation for keyword spotting system with conditional GAN
Donghua Wang 0001, Li Dong 0006, Rangding Wang, Diqun Yan
Comput. Commun.2
2021 Multi-windowed vertex-frequency analysis for signals on undirected graphs
Xianwei Zheng, Cuiming Zou, Li Dong 0006, Jiantao Zhou 0001
Comput. Commun.3
2021 Tackling the Cover Source Mismatch Problem in Audio Steganalysis With Unsupervised Domain Adaptation
abstract
Nowadays, the convolutional neural network (CNN) based steganalysis has achieved remarkable performance in the well-controlled lab environment. However, the cover source mismatch (CSM) problem, which can be attributed to the discrepancy between the training, and evaluation datasets, is still one of the pivotal obstacles for adapting the steganalysis into real-world applications. In this letter, we propose to merge the domain adaptation strategy into CNN-based audio steganalysis for handling the CSM problem. Specifically, the proposed framework contains three components: feature extractor, steganalytic classifier, and domain discriminator. The cascade of feature extractor, and steganalytic classifier compose the typical supervised steganalysis model. The unsupervised domain adaptation is implemented by the domain adversarial training between the feature extractor, and domain discriminator. Ultimately, the feature extractor is trained to extract the steganalytic, and domain-invariant features. It aims to reduce the domain gap between the training data, and testing data. The experimental results show that our approach could effectively mitigate the CSM impact caused by the diversity of audio recording devices.
Yuzhen Lin, Rangding Wang, Li Dong 0006, Diqun Yan, Jie Wang 0028
IEEE Signal Process. Lett.3
2021 Optimal Pre-Filtering for Improving Facebook Shared Images
abstract
Online Social Networks (OSNs) have attracted a huge number of users, who store and share various images on a daily basis. As a well-known fact, most OSN platforms apply a series of lossy operations on the uploaded images, which could severely degrade the quality of the shared images, negatively affecting the user experiences. In this work, we consider the problem of significantly improving OSN-shared images through applying an optimal pre-filtering prior to image sharing, without any cooperation from the OSN platform itself. Facebook, as one of the most popular and representative OSNs, is chosen as the platform to present our designed pre-filtering strategy. We first treat Facebook as a black box, and thoroughly recover its mechanism of processing color images. Based on the precise knowledge on the image processing pipeline on Facebook, we design the pre-filter under an optimization framework, minimizing the end-to-end distortion between the shared image and the original one. Compared with the directly shared images, our proposed pre-filtering-then-sharing strategy brings significant improvements in terms of both quantitative and qualitative metrics. Extensive experimental results are provided to show the superiority of our proposed method. Finally, we discuss the strategy on how to extend our proposed technique to other OSN platforms.
Weiwei Sun 0009, Jiantao Zhou 0001, Li Dong 0006, Jinyu Tian 0001, Jun Liu 0071
IEEE Trans. Image Process.3
2020 Towards Designing an Effective Complexity Indicator for Audio Steganography
abstract
In the field of steganography, to effectively hide the secret message, it is of great importance to determine which part of the steganographic cover is suitable for embedding. Currently, most of the existing works focus on the image cover, while few works touch the audio cover case. In this work, we attempt to characterize the complexity of audio for selecting the steganographic cover. Specifically, the original cover is first convoluted with a specially designed adaptive convolution kernel. Based on the residual between the original and the convoluted audio, we derive a quantity for measuring the complexity of each frame for a given audio clip. Experimental results verify the usability of the proposed complexity indicator, suggesting high-complexity audio cover is favorable for data embedding. It is also found that the proposed complexity indicator could further boost the steganographic performance of the state-of the-art audio steganography methods. The source code is publicly available at https://github.com/capzxy/audio-complexity.
Xueyuan Zhang, Rangding Wang, Li Dong 0006, Diqun Yan, Yuzhen Lin, Jie Wang 0028
ICC3
2020 Efficient Generation of Speech Adversarial Examples with Generative Model
Donghua Wang 0001, Rangding Wang, Li Dong 0006, Diqun Yan
IWDW3
2020 An Antiforensic Method against AMR Compression Detection
abstract
Adaptive multirate (AMR) compression audio has been exploited as an effective forensic evidence to justify audio authenticity. Little consideration has been given, however, to antiforensic techniques capable of fooling AMR compression forensic algorithms. In this paper, we present an antiforensic method based on generative adversarial network (GAN) to attack AMR compression detectors. The GAN framework is utilized to modify double AMR compressed audio to have the underlying statistics of single compressed one. Three state-of-the-art detectors of AMR compression are selected as the targets to be attacked. The experimental results demonstrate that the proposed method is capable of removing the forensically detectable artifacts of AMR compression under various ratios with an average successful attack rate about 94.75%, which means the modified audios generated by our well-trained generator can treat the forensic detector effectively. Moreover, we show that the perceptual quality of the generated AMR audio is well preserved.
Diqun Yan, Li Dong 0006, Rangding Wang
Secur. Commun. Networks3
2019 Audio Steganalysis with Improved Convolutional Neural Network
abstract
Deep learning, especially the convolutional neural network (CNN), has enjoyed significant success in many fields, e.g., image recognition. Recently, CNN has successfully applied to multimedia steganalysis. However, the detection performance is still unsatisfactory. In this work, we propose an improved CNN-based method for audio steganalysis. Specifically, a special convolutional layer is first carefully designed, which could capture the minor steganographic noise. Then, a truncated linear unit is adapted to activate the output of shallow convolutional layer. In addition, we employ the average pooling to minimize the over-fitting risk. Finally, a parameter transfer strategy is adopted, aiming to boost the detection performance for the low embedding-rate cases. The experimental results evaluated on 30,000 audio clips verify the effectiveness of our method for a variety of embedding rates. Compared with the existing CNN-based steganalysis methods, our proposed method could achieve superior performance. To facilitate the reproducible research, the source code will be released at GitHub.
Yuzhen Lin, Rangding Wang, Diqun Yan, Li Dong 0006, Xueyuan Zhang
IH&MMSec4
2019 Content-Adaptive Noise Estimation for Color Images With Cross-Channel Noise Modeling
abstract
Noise estimation is crucial in many image processing tasks such as denoising. Most of the existing noise estimation methods are specially developed for grayscale images. For color images, these methods simply handle each color channel independently, without considering the correlation across channels. Moreover, these methods often assume a globally fixed noise model throughout the entire image, neglecting the adaptation to the local structures. In this work, we propose a contentadaptive multivariate Gaussian approach to model the noise in color images, in which we explicitly consider both the contentdependence and the inter-dependence among color channels. We design an effective method for estimating the noise covariance matrices within the proposed model. Specifically, a patch selection scheme is first introduced to select weakly textured patches via thresholding the texture strength indicators. Noticing that the patch selection actually depends on the unknown noise covariance, we present an iterative noise covariance estimation algorithm, where the patch selection and the covariance estimation are conducted alternately. For the remaining textured regions, we estimate a distinct covariance matrix associated with each pixel using a linear shrinkage estimator, which adaptively fuses the estimate coming from the weakly textured region and the sample covariance estimated from the local region. Experimental results show that our method can effectively estimate the noise covariance. The usefulness of our method is demonstrated with several image processing applications such as color image denoising and noise-robust superpixel.
Li Dong 0006, Jiantao Zhou 0001, Yuan Yan Tang
IEEE Trans. Image Process.1
2018 Color Image Noise Covariance Estimation with Cross-Channel Image Noise Modeling
abstract
Noise estimation is crucial in many image processing tasks such as denoising. Most of the existing noise estimation methods are specially developed for grayscale images. For color images, these methods simply handle each color channel independently, without considering the correlation across channels. In this work, we propose a multivariate Gaussian approach to model the noise in color images, in which we explicitly consider the inter-dependence among color channels. We design a practical method for estimating the noise covariance matrix within the proposed model. Specifically, a patch selection scheme is first introduced to select weakly textured patches through thresholding the texture strength indicators. Noticing that the patch selection actually depends on the unknown noise covariance, we present an iterative noise covariance estimation algorithm, where the patch selection and the covariance estimation are conducted alternately. Experimental results show that our method can effectively estimate the noise covariance. The practical usage is demonstrated with color image denoising.
Li Dong 0006, Jiantao Zhou 0001, Tao Dai 0001
ICME1
2018 Effective and Fast Estimation for Image Sensor Noise Via Constrained Weighted Least Squares
abstract
Noise estimation is crucial in many image processing algorithms such as image denoising. Conventionally, the noise is assumed as a signal-independent additive white Gaussian process. However, for the real raw data of image sensor, the present noise should be practically modeled as signal dependent. In this paper, we propose an effective and fast image sensor noise estimation method for a single raw image. The noise model parameters are estimated via constrained weighted least squares (WLS) fitting on a number of data samples, each of which is generated from a group of weakly textured patches. Specifically, we first design a fast scheme for selecting weakly textured patches, with the guidance of image histogram. To robustly fit the data samples, we then explicitly account for the credibility of each sample by measuring the texture strength of the grouped patches. The image sensor noise estimation is finally formulated as a constrained WLS optimization problem, which can be solved efficiently. Experimental results demonstrate that our method could run much faster than the existing schemes, while retaining the state-of-the-art estimation performance.
Li Dong 0006, Jiantao Zhou 0001, Yuan Yan Tang
IEEE Trans. Image Process.1
2017 Efficient image sensor noise estimation via iterative re-weighted least squares
abstract
Noise estimation is crucial in many image processing algorithms such as image denoising. Conventionally, the noise is assumed as signal-independent additive white Gaussian process. However, for the real raw-data of imaging sensors, the present noise is better modeled as signal-dependent noise. In this work, we propose an efficient image sensor noise estimation method based on iterative re-weighted least squares optimization. Specifically, the image patches are first clustered into different groups, each of which will generate a data sample. To fit those observations robustly, we introduce a weighting matrix to reflect the credibility of each sample. Unfortunately, this setting of weighting matrix in turn depends on the unknown noise parameters. We then develop an iterative re-weighted least squares optimization procedure, in which the weighting matrix and parameter estimates can be updated alternately. Experimental results show that our method outperforms the state-of-the-art works, in terms of both estimation accuracy and computational efficiency.
Li Dong 0006, Jiantao Zhou 0001, Guangtao Zhai
ICME1
2017 Noise Level Estimation for Natural Images Based on Scale-Invariant Kurtosis and Piecewise Stationarity
abstract
Noise level estimation is crucial in many image processing applications, such as blind image denoising. In this paper, we propose a novel noise level estimation approach for natural images by jointly exploiting the piecewise stationarity and a regular property of the kurtosis in bandpass domains. We design a K-means-based algorithm to adaptively partition an image into a series of non-overlapping regions, each of whose clean versions is assumed to be associated with a constant, but unknown kurtosis throughout scales. The noise level estimation is then cast into a problem to optimally fit this new kurtosis model. In addition, we develop a rectification scheme to further reduce the estimation bias through noise injection mechanism. Extensive experimental results show that our method can reliably estimate the noise level for a variety of noise types, and outperforms some state-of-the-art techniques, especially for non-Gaussian noises.
Li Dong 0006, Jiantao Zhou 0001, Yuan Yan Tang
IEEE Trans. Image Process.1
2016 Estimating noise level for natural images based on scale-invariant kurtosis and piecewise stationarity
abstract
Noise level estimation is crucial in many image processing applications such as blind image denoising. In this work, we propose a novel noise level estimation approach for natural images by jointly exploiting the piecewise stationarity and a regular property of the kurtosis in band-pass domains. We design a K-means based algorithm to adaptively partition an image into a series of non-overlapping regions, each of whose clean versions is assumed to be associated with a constant kurtosis throughout scales. The noise level estimation is then formulated as a problem to optimally fit this new kurtosis model. Experimental results show that our method can reliably estimate the noise level for a variety of noise types, and outperforms some state-of-the-art techniques, especially for non-Gaussian noises.
Li Dong 0006, Jiantao Zhou 0001
ICIP1
2016 Secure Reversible Image Data Hiding Over Encrypted Domain via Key Modulation
abstract
This paper proposes a novel reversible image data hiding scheme over encrypted domain. Data embedding is achieved through a public key modulation mechanism, in which access to the secret encryption key is not needed. At the decoder side, a powerful two-class SVM classifier is designed to distinguish encrypted and nonencrypted image patches, allowing us to jointly decode the embedded message and the original image signal. Compared with the state-of-the-art methods, the proposed approach provides higher embedding capacity and is able to perfectly reconstruct the original image as well as the embedded message. Extensive experimental results are provided to validate the superior performance of our scheme.
Jiantao Zhou 0001, Weiwei Sun 0009, Li Dong 0006, Xianming Liu 0005, Oscar C. Au, Yuan Yan Tang
IEEE Trans. Circuits Syst. Video Technol.3
2014 Estimation of capacity parameters for dynamic histogram shifting (DHS)-based reversible image watermarking
abstract
Dynamic histogram shifting (DHS) is a generation of the conventional histogram shifting (HS) technique for reversible image watermarking. Its superior embedding performance is achieved at the cost of significantly increased computational burden incurred by estimating the capacity parameters via multi-rounds of embedding iterations. In this work, we propose an analytical framework on estimating the optimal capacity parameters for DHS-based reversible image watermarking. We demonstrate that such parameter estimation can be cast as a convex optimization problem, which can be numerically solved in an efficient manner. The estimated values can then be utilized to facilitate a local search algorithm to obtain the truly optimal ones with much lowered complexity. Experimental results are provided to verify the validity of our findings.
Li Dong 0006, Jiantao Zhou 0001, Yuan Yan Tang, Xianming Liu 0005
ICME1