Haodong Li 0001

dblp:126/4508-1 · DBLP profile ↗
← Back
27ranked-venue papers
9as first author
12since 2021 · last 2026
0000-0003-0532-9481ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 6 first-author · 7 since 2021Security and privacy · 9 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 2 since 2021Computer networks · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Document image forgery detection and localization in desensitization scenarios
Weixiang Li, Bin Li 0011, Kengtao Zheng, Haodong Li 0001
Signal Process.5
2025 DITL2: Dual-Stage Invariance Transfer Learning for Generalizable Document Image Tampering Localization
abstract
Document Image Tampering Localization (DITL) advances considerably, yet achieving robust cross-dataset generalization remains a formidable challenge for practical applications. Expanding existing document datasets for training is labor-intensive, making it appealing to incorporate data from non-document domains such as natural scene images. However, domain-specific variations, including differences in color distribution and texture, compromise the performance of joint training. To address this issue, we propose DITL2, a Dual-stage Invariance Transfer Learning framework for Document Image Tampering Localization that consists of Cross-Domain Invariance Pre-training (CDIP) and Frequency Decoupling Parameter Adaptation (FDPA). In the pre-training stage, CDIP employs style transfer and texture consistency learning to suppress domain-specific influences from tampered natural scene images, and tampering trace commonality learning to acquire domain-invariant features. In the fine-tuning stage, FDPA adapts the parameters of the pre-trained model, leveraging the general knowledge from the pre-trained model to address DITL tasks while reducing the risk of overfitting. Experiments show that this approach effectively leverages external data resources to boost model performance, achieving state-of-the-art results across a variety of cross-dataset settings.
Shen Chen 0004, Bin Li 0011, Kaiqing Lin, Changsheng Chen 0001, Haodong Li 0001, Taiping Yao, Shouhong Ding
ACM Multimedia7
2025 Towards generalizable and robust image tampering localization with multi-task learning and contrastive learning
Haodong Li 0001, Peiyu Zhuang, Yang Su 0005, Jiwu Huang
Expert Syst. Appl.1
2024 A distortion model guided adversarial surrogate for recaptured document detection
Changsheng Chen 0001, Xijin Li, Baoying Chen, Haodong Li 0001
Pattern Recognit.4
2024 WebP-JPEG Transcoding Detection by Spotting Re-Compression Artifacts With CNN-ViT for Processing Dual-Domain Features
abstract
The trace of double compression can serve as a crucial evidence of image manipulation for forensic investigation. With the ever-increasing popularity of WebP format, a new type of double compression case, WebP-JPEG transcoding, has emerged. However, distinguishing it from two common compression cases, single JPEG (SJPEG) and double JPEG (DJPEG) has not yet been studied. In this paper, we propose a specialized method for the new task. Firstly, a detailed analysis is conducted to reveal the differences in compression artifacts between WebP-JPEG and SJPEG/DJPEG, which manifests in the distributions of$4\times 4$/$8\times 8$DCT coefficients and the high-frequency portions of image spectrum. Then, multi-modality DCT histograms (MMDH) and high-pass-filtered image residuals (HPFIR) are proposed as front-end dual-domain forensic features to expose the above differences. An indispensable part of these features are extracted through a novel frequency-isolation module (FIM), offering additional information based on the derived relationship between$4\times 4$and$8\times 8$DCT coefficients. Finally, a CNN-ViT (Convolutional Neural Network-Vision Transformer) dual-stream network is designed to learn back-end deep features for a reliable detection, where a CNN stream is used to process statistical features in MMDH while a ViT stream to learn spatial correlations in HPFIR. Extensive experimental results demonstrate that the proposed method significantly outperforms state-of-the-art double compression detection methods in distinguishing WebP-JPEG from SJPEG/DJPEG and is more effective in tampering localization. In specific, the proposed method achieves an average detection accuracy of 0.942 for small images of size$128\times 128$.
Bin Li 0011, Weixiang Li, Haodong Li 0001
IEEE Trans. Circuits Syst. Video Technol.4
2023 ReLoc: A Restoration-Assisted Framework for Robust Image Tampering Localization
abstract
With the spread of tampered images, locating the tampered regions in digital images has drawn increasing attention. The existing tampering localization methods, however, suffer from severe performance degradation when the images are subjected to some post-processing, as the tampering traces would be distorted by the post-processing operations. The poor robustness against post-processing has become a bottleneck for the practical applications of image tampering localization techniques. In order to address this issue, this paper proposes a novelrestoration-assisted framework for image tamperinglocalization (ReLoc). The ReLoc framework mainly consists of an image restoration module and a tampering localization module. The key idea of ReLoc is to use the restoration module to recover a high-quality counterpart from the distorted tampered image, such that the distorted tampering traces can be re-enhanced, facilitating the tampering localization module to identify the tampered regions. To achieve this, the restoration module is optimized not only with the conventional constraints on image visual quality, but also with a forensics-oriented objective function. Furthermore, the restoration module and the localization module are trained alternately, which can stabilize the training process and is beneficial for improving the performance. The robustness of ReLoc has been evaluated by using several common post-processing operations, including lossy compressions, online social network transmission, and image resizing. Extensive experimental results show that ReLoc can significantly improve the localization performance compared to using a restoration-free model. In addition, we have shown that the restoration module in a well-trained ReLoc model is transferable for different localization modules and across different datasets.
Peiyu Zhuang, Haodong Li 0001, Rui Yang 0006, Jiwu Huang
IEEE Trans. Inf. Forensics Secur.2
2022 Deep Video Inpainting Localization Using Spatial and Temporal Traces
abstract
Advanced deep-learning-based video inpainting can fill a specified video region with visually plausible contents, usually leaving imperceptible traces. As inpainting can be used for malicious video manipulations, it has led to potential privacy and security issues. Therefore, it is necessary to detect and locate the video regions subjected to deep inpainting. This paper addresses this problem by exploiting the spatial and temporal traces left by inpainting. Firstly, the inpainting traces are enhanced by intra-frame and inter-frame residuals. In particular, we guide the extraction of inter-frame residual with optical-flow based frame alignment, which can better reveal the inpainting traces. Then, a dual-stream network, acting as the encoder, is designed to learn discriminative features from frame residuals. Finally, bidirectional convolutional LSTMs are embedded in the decoder network to produce pixel-wise predictions of inpainted regions for each frame. The proposed method is evaluated with tampered videos created by two state-of-the-art deep video inpainting algorithms. Extensive experimental results show that the proposed method can effectively localize the inpainted regions, outperforming existing methods.
Shujin Wei, Haodong Li 0001, Jiwu Huang
ICASSP2
2022 DS-UNet: A dual streams UNet for refined image forgery localization
Yuanhang Huang, Shan Bian, Haodong Li 0001, Chuntao Wang, Kangshun Li
Inf. Sci.3
2022 A distortion model-based pre-screening method for document image tampering localization under recapturing attack
Changsheng Chen 0001, Lin Zhao 0017, Jiabin Yan, Haodong Li 0001
Signal Process.4
2021 Image Tampering Localization Using Unified Two-Stream Features Enhanced with Channel and Spatial Attention
Haodong Li 0001, Peiyu Zhuang, Bin Li 0011
PRCV (2)1
2021 Detail-enhanced image inpainting based on discrete wavelet transforms
Bin Li 0011, Bowei Zheng, Haodong Li 0001, Yanran Li
Signal Process.3
2021 Image Tampering Localization Using a Dense Fully Convolutional Network
abstract
The emergence of powerful image editing software has substantially facilitated digital image tampering, leading to many security issues. Hence, it is urgent to identify tampered images and localize tampered regions. Although much attention has been devoted to image tampering localization in recent years, it is still challenging to perform tampering localization in practical forensic applications. The reasons include the difficulty of learning discriminative representations of tampering traces and the lack of realistic tampered images for training. Since Photoshop is widely used for image tampering in practice, this paper attempts to address the issue of tampering localization by focusing on the detection of commonly used editing tools and operations in Photoshop. In order to well capture tampering traces, a fully convolutional encoder-decoder architecture is designed, where dense connections and dilated convolutions are adopted for achieving better localization performance. In order to effectively train a model in the case of insufficient tampered images, we design a training data generation strategy by resorting to Photoshop scripting, which can imitate human manipulations and generate large-scale training samples. Extensive experimental results show that the proposed approach outperforms state-of-the-art competitors when the model is trained with only generated images or fine-tuned with a small amount of realistic tampered images. The proposed method also has good robustness against some common post-processing operations.
Peiyu Zhuang, Haodong Li 0001, Shunquan Tan, Bin Li 0011, Jiwu Huang
IEEE Trans. Inf. Forensics Secur.2
2020 Image processing operations identification via convolutional neural network
Haodong Li 0001, Weiqi Luo 0001, Jiwu Huang
Sci. China Inf. Sci.2
2020 Identification of deep network generated images using disparities in color components
Haodong Li 0001, Bin Li 0011, Shunquan Tan, Jiwu Huang
Signal Process.1
2019 Localization of Deep Inpainting Using High-Pass Fully Convolutional Network
abstract
Image inpainting has been substantially improved with deep learning in the past years. Deep inpainting can fill image regions with plausible contents, which are not visually apparent. Although inpainting is originally designed to repair images, it can even be used for malicious manipulations, e.g., removal of specific objects. Therefore, it is necessary to identify the presence of inpainting in an image. This paper presents a method to locate the regions manipulated by deep inpainting. The proposed method employs a fully convolutional network that is based on high-pass filtered image residuals. Firstly, we analyze and observe that the inpainted regions are more distinguishable from the untouched ones in the residual domain. Hence, a high-pass pre-filtering module is designed to get image residuals for enhancing inpainting traces. Then, a feature extraction module, which learns discriminative features from image residuals, is built with four concatenated ResNet blocks. The learned feature maps are finally enlarged by an up-sampling module, so that a pixel-wise inpainting localization map is obtained. The whole network is trained end-to-end with a loss addressing the class imbalance. Extensive experimental results evaluated on both synthetic and realistic images subjected to deep inpainting have shown the effectiveness of the proposed method.
Haodong Li 0001, Jiwu Huang
ICCV1
2018 Identification of Various Image Operations Using Residual-Based Features
abstract
Image forensics has attracted wide attention during the past decade. However, most existing works aim at detecting a certain operation, which means that their proposed features usually depend on the investigated image operation and they consider only binary classification. This usually leads to misleading results if irrelevant features and/or classifiers are used. For instance, a JPEG decompressed image would be classified as an original or median filtered image if it was fed into a median filtering detector. Hence, it is important to develop forensic methods and universal features that can simultaneously identify multiple image operations. Based on extensive experiments and analysis, we find that any image operation, including existing anti-forensics operations, will inevitably modify a large number of pixel values in the original images. Thus, some common inherent statistics such as the correlations among adjacent pixels cannot be preserved well. To detect such modifications, we try to analyze the properties of local pixels within the image in the residual domain rather than the spatial domain considering the complexity of the image contents. Inspired by image steganalytic methods, we propose a very compact universal feature set and then design a multiclass classification scheme for identifying many common image operations. In our experiments, we tested the proposed features as well as several existing features on 11 typical image processing operations and four kinds of anti-forensic methods. The experimental results show that the proposed strategy significantly outperforms the existing forensic methods in terms of both effectiveness and universality.
Haodong Li 0001, Weiqi Luo 0001, Xiaoqing Qiu, Jiwu Huang
IEEE Trans. Circuits Syst. Video Technol.1
2018 Improved Audio Steganalytic Feature and Its Applications in Audio Forensics
abstract
Digital multimedia steganalysis has attracted wide attention over the past decade. Currently, there are many algorithms for detecting image steganography. However, little research has been devoted to audio steganalysis. Since the statistical properties of image and audio files are quite different, features that are effective in image steganalysis may not be effective for audio. In this article, we design an improved audio steganalytic feature set derived from both the time and Mel-frequency domains for detecting some typical steganography in the time domain, including LSB matching, Hide4PGP, and Steghide. The experiment results, evaluated on different audio sources, including various music and speech clips of different complexity, have shown that the proposed features significantly outperform the existing ones. Moreover, we use the proposed features to detect and further identify some typical audio operations that would probably be used in audio tampering. The extensive experiment results have shown that the proposed features also outperform the related forensic methods, especially when the length of the audio clip is small, such as audio clips with 800 samples. This is very important in real forensic situations.
Weiqi Luo 0001, Haodong Li 0001, Qi Yan 0004, Rui Yang 0006, Jiwu Huang
ACM Trans. Multim. Comput. Commun. Appl.2
2017 Audio Steganalysis with Convolutional Neural Network
abstract
In recent years, deep learning has achieved breakthrough results in various areas, such as computer vision, audio recognition, and natural language processing. However, just several related works have been investigated for digital multimedia forensics and steganalysis. In this paper, we design a novel CNN (convolutional neural networks) to detect audio steganography in the time domain. Unlike most existing CNN based methods which try to capture media contents, we carefully design the network layers to suppress audio content and adaptively capture the minor modifications introduced by ±1 LSB based steganography. Besides, we use a mix of convolutional layer and max pooling to perform subsampling to achieve good abstraction and prevent over-fitting. In our experiments, we compared our network with six similar network architectures and two traditional methods using handcrafted features. Extensive experimental results evaluated on 40,000 speech audio clips have shown the effectiveness of the proposed convolutional network.
Weiqi Luo 0001, Haodong Li 0001
IH&MMSec3
2017 Adaptive Audio Steganography Based on Advanced Audio Coding and Syndrome-Trellis Coding
Weiqi Luo 0001, Haodong Li 0001
IWDW3
2017 Localization of Diffusion-Based Inpainting in Digital Images
abstract
Image inpainting, an image processing technique for restoring missing or damaged image regions, can be utilized by forgers for removing objects in digital images. Since no obviously perceptible artifacts are left after inpainting, it is necessary to develop methods for detecting the presence of inpainting. In general, there are two main categories of image inpainting techniques: exemplar-based and diffusion-based techniques. Although several methods have been proposed for detecting exemplar-based inpainting, there is still no effective method for detecting diffusion-based inpainting. Usually, the tampered regions manipulated by diffusion-based inpainting techniques are much smaller than those manipulated by exemplar-based ones, presenting more challenges in detecting these regions. As a pioneering attempt, this paper proposes a method for the localization of diffusion-based inpainted regions in digital images. We first analyze the diffusion process in inpainting, and observe that the changes in the image Laplacian along the direction perpendicular to the gradient are different in the inpainted and untouched regions. Following this observation, we construct a feature set based on the intra-channel and inter-channel local variances of the changes to identify the inpainted regions. Finally, two effective post-processing operations are designed for further refining of the localization result. The extensive experimental results evaluated on both synthetic and realistic inpainted images show the effectiveness of the proposed method.
Haodong Li 0001, Weiqi Luo 0001, Jiwu Huang
IEEE Trans. Inf. Forensics Secur.1
2017 Image Forgery Localization via Integrating Tampering Possibility Maps
abstract
Over the past decade, many efforts have been made in passive image forensics. Although it is able to detect tampered images at high accuracies based on some carefully designed mechanisms, localization of the tampered regions in a fake image still presents many challenges, especially when the type of tampering operation is unknown. Some researchers have realized that it is necessary to integrate different forensic approaches in order to obtain better localization performance. However, several important issues have not been comprehensively studied, for example, how to select and improve/readjust proper forensic approaches, and how to fuse the detection results of different forensic approaches to obtain good localization results. In this paper, we propose a framework to improve the performance of forgery localization via integrating tampering possibility maps. In the proposed framework, we first select and improve two existing forensic approaches, i.e., statistical feature-based detector and copy-move forgery detector, and then adjust their results to obtain tampering possibility maps. After investigating the properties of possibility maps and comparing various fusion schemes, we finally propose a simple yet very effective strategy to integrate the tampering possibility maps to obtain the final localization results. The extensive experiments show that the two improved approaches used in our framework significantly outperform the state-of-the-art techniques, and the proposed fusion results achieve the best F1-score in the IEEE IFS-TC Image Forensics Challenge.
Haodong Li 0001, Weiqi Luo 0001, Xiaoqing Qiu, Jiwu Huang
IEEE Trans. Inf. Forensics Secur.1
2016 Adaptive Steganalysis Based on Embedding Probabilities of Pixels
abstract
In modern steganography, embedding modifications are highly concentrated on the textural regions within an image, as such regions are difficult to model for steganalysis. Previous studies have shown that compared with non-adaptive strategies, this content adaptive strategy achieves stronger security against existing steganalysis. Based on the experiments and analyses, however, we found that this embedding property would inevitably lead to a large limitation in existing adaptive steganography. That is, it is possible for steganalyzers to estimate the regions that have probably been modified after data hiding. In this paper, we propose an adaptive steganalytic scheme based on embedding probabilities of pixels. The main idea of our scheme is that we assign different weights to different pixels in feature extraction. For those pixels with high embedding probabilities, their corresponding weights are larger, since they should contribute more to steganalysis and vice versa. By doing so, we can concentrate our attention on the regions that have probably been modified and significantly reduce the impact of other unchanged smooth regions. It is expected that our proposed method is an improvement on the existing steganalytic methods, which usually assume every pixel has the same contribution to steganalysis. The extensive experiments evaluated on four typical adaptive steganographic methods have shown the effectiveness of the proposed scheme, especially for low embedding rates, for example, lower than 0.20 bpp.
Weixuan Tang 0004, Haodong Li 0001, Weiqi Luo 0001, Jiwu Huang
IEEE Trans. Inf. Forensics Secur.2
2015 Anti-forensics of double JPEG compression with the same quantization matrix
Haodong Li 0001, Weiqi Luo 0001, Jiwu Huang
Multim. Tools Appl.1
2014 Anti-forensics of JPEG Detectors via Adaptive Quantization Table Replacement
abstract
Due to the popularity of JPEG compression standard, JPEG images have been widely used in various applications. Nowadays, detection of JPEG forgeries becomes an important issue in digital image forensics, and lots of related works have been reported. However, most existing works mainly rely on a pre-trained classifier according to the quantization table shown in the file header of the suspicious JPEG image, and they assume that such a table is authentic. This assumption leaves a potential flaw for those wise forgers to confuse or even invalidate the current JPEG forensic detectors. Based on our analysis and experiments, we found that the generalization ability of most current JPEG forensic detectors is not very good. If the quantization table changes, their performances would decrease significantly. Based on this observation, we propose a universal anti-forensic scheme via replacing the quantization table adaptively. The extensive experimental results evaluated on 10,000 natural images have shown the effectiveness of the proposed scheme for confusing four typical JPEG forensic works.
Haodong Li 0001, Weiqi Luo 0001, Rui Yang 0006, Jiwu Huang
ICPR2
2014 A universal image forensic strategy based on steganalytic model
abstract
Image forensics have made great progress during the past decade. However, almost all existing forensic methods can be regarded as the specific way, since they mainly focus on detecting one type of image processing operations. When the type of operations changes, the performances of the forensic methods usually degrade significantly. In this paper, we propose a universal forensics strategy based on steganalytic model. By analyzing the similarity between steganography and image processing operation, we find that almost all image operations have to modify many image pixels without considering some inherent properties within the original image, which is similar to what in steganography. Therefore, it is reasonable to model various image processing operations as steganography and it is promising to detect them with the help of some effective universal steganalytic features. In our experiments, we evaluate several advanced steganalytic features on six kinds of typical image processing operations. The experimental results show that all evaluated steganalyzers perform well while some steganalytic methods such as the spatial rich model (SRM) [4] and LBP [19] based methods even outperform the specific forensic methods significantly. What is more, they can further identify the type of various image processing operations, which is impossible to achieve using the existing forensic methods.
Xiaoqing Qiu, Haodong Li 0001, Weiqi Luo 0001, Jiwu Huang
IH&MMSec2
2014 Adaptive steganalysis against WOW embedding algorithm
abstract
WOW (Wavelet Obtained Weights) [5] is one of the advanced steganographic methods in spatial domain, which can adaptively embed secret message into cover image according to textural complexity. Usually, the more complex of an image region, the more pixel values within it would be modified. In such a way, it can achieve good visual quality of the resulting stegos and high security against typical steganalytic detectors. Based on our analysis, however, we point out one of the limitations in the WOW embedding algorithm, namely, it is easy to narrow down those possible modified regions for a given stego image based on the embedding costs used in WOW. If we just extract features from such regions and perform analysis on them, it is expected that the detection performance would be improved compared with that of extracting steganalytic features from the whole image. In this paper, we first proposed an adaptive steganalytic scheme for the WOW method, and use the spatial rich model (SRM) based features [4] to model those possible modified regions in our experiments. The experimental results evaluated on 10,000 images have shown the effectiveness of our scheme. It is also noted that our steganalytic strategy can be combined with other steganalytic features to detect the WOW and/or other adaptive steganographic methods both in the spatial and JPEG domains.
Weixuan Tang 0004, Haodong Li 0001, Weiqi Luo 0001, Jiwu Huang
IH&MMSec2
2012 Countering anti-JPEG compression forensics
abstract
The quantization artifacts and blocking artifacts are the two significant properties in the JPEG compressed images. Most relative forensic techniques usually use such inherent properties to provide some evidences on how image data is acquired and/or processed. A wise attacker, however, may perform some post-operations to confuse the two artifacts to fool current forensic techniques. Recently, Stamm et al. in [1] propose a novel anti-JPEG compression method via adding anti-forensic dither to the DCT coefficients and further reducing the blocking artifacts. In this paper, we found that the dithering operation will inevitably destroy the statistical correlations among the 8 × 8 intrablock and interblock within an image. In the view of JPEG steganalysis, we employ the transition probability matrix of the DCT coefficients to measure such modifications for identifying the forged images from those original JPEG decompressed images and uncompressed ones. On average, we can obtain a detection accuracy as high as 99% on the image database of UCID [2].
Haodong Li 0001, Weiqi Luo 0001, Jiwu Huang
ICIP1