Hao Wang 0060

dblp:181/2812-60 · DBLP profile ↗
← Back
16ranked-venue papers
6as first author
15since 2021 · last 2026
0000-0003-4139-8193ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 8 since 2021Security and privacy · 5 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 A survey on JPEG image forensics: Exploring key advances and persistent challenges in compression and quantization analysis
Hao Wang 0060, Xin Cheng 0018, Jiawei Zhang 0011, Hao Wu 0078, Xue Xie, Xiangyang Luo 0001, Bin Ma 0003
Comput. Secur.1
2026 Anti-forensic for quantization steps estimation based on direct and preemptive adversarial attacks
Jiawei Zhang 0011, Hao Wu 0078, Xin Cheng 0018, Xiangyang Luo 0001, Bin Ma 0003, Hao Wang 0060
Eng. Appl. Artif. Intell.7
2026 Re-Cropping Framework: A Grid Recovery Method for Quantization Step Estimation in Non-Aligned Recompressed Images
abstract
The manipulation history of Joint Photographic Experts Group (JPEG) compression plays an important role in JPEG image forensics and information hiding. For non-aligned recompressed images, different cropping methods produce non-aligned outputs with varying feature distributions. One such important factor is the shifts of the discrete cosine transform (DCT) grid (i.e., the misalignment parameters) between two compression processes. Although many methods have been proposed to estimate the misalignment parameters, the limited amount of useful information available in small-sized images leads to low accuracy of these methods. To enhance the accuracy of misalignment parameter estimation for small-sized non-aligned images, we propose a novel two-branch network structure that accounts for the unique horizontal and vertical characteristics of non-aligned images. This structure employs convolution to simulate second-order difference (SOD) and incorporates it throughout the training process to optimize the difference parameters dynamically. Based on the insight that cropping operations leave traces in all color channels, we derive the Cg channel through a color space transformation. This approach expands the input dimensionality to four channels (Y, Cb, Cr, and Cg), thereby compensating for the information scarcity in small-sized images. The experimental results show that our method outperforms existing methods on different image sizes, regardless of the known or unknown quality factor (QF) of the first compression. Finally, we propose a re-cropping framework based on the estimated misalignment parameters. The influence of the first cropping is counteracted by a re-cropping operation, which improves the accuracy of existing methods in estimating the first quantization step for non-aligned recompressed images.
Xin Cheng 0018, Hao Wang 0060, Xiangyang Luo 0001, Qingxiao Guan, Bin Ma 0003
IEEE Trans. Circuits Syst. Video Technol.2
2026 Rethinking Cross-Table Quantization Step Estimation: From Global and Local Perspectives
abstract
The quantization step is a crucial parameter in the JPEG compression process, and provides prior knowledge for JPEG image steganography and forensics. Existing neural network-based methods typically estimate the quantization steps for all discrete cosine transform (DCT) subbands jointly, by treating the entire quantization table as a unified input and leveraging the inter-subband relationships. However, subband relationships vary across different quantization tables, leading to poor generalization for methods that rely heavily on such relationships. To address the above issues, we depart from the strategy that relies on inter-subband relationships and instead train the model on a specific single subband. To compensate for the possible decrease in accuracy due to the lack of relationships between subbands, we extract the ranking features and histogram features from the DCT coefficient histograms of the subbands. Ranking features capture local patterns in DCT histograms by modeling the relative relationships between neighboring coefficients, thereby compensating for the absence of local detail. On the other hand, histogram features represent the overall distribution pattern of the DCT coefficient histograms and capture the global trends and statistical properties in the subbands. We subsequently employ convolutional groups and multilayer perceptron (MLP) structures to extract compression artifacts from these two features. Finally, we introduce a comprehensive evaluation metric, called GenAQt, to quantify the algorithm’s generalization ability across quantization tables. The experimental results demonstrate that our method maintains high accuracy across quantization tables, with RelGenAQt (relative accuracy decrease) exceeding 81% and AbsGenAQt (absolute accuracy decrease) being less than 0.38.
Xin Cheng 0018, Hao Wang 0060, Xiangyang Luo 0001, Bin Ma 0003, Baowei Wang, Bin Li 0011
IEEE Trans. Inf. Forensics Secur.2
2026 HENet: A Heterogeneous Encoding Network for General and Robust Adversarial Example Generation
abstract
Generator-based adversarial attack methods aim to fool deep neural networks (DNNs) by training a generator for crafting adversarial examples (AEs). However, as DNNs evolve from Convolutional Neural Networks (CNNs) to Transformers, the existing generator-based methods can hardly achieve satisfactory attack performance against different target model architectures in semi-whitebox attack scenarios. In addition, the generated AEs are susceptible to various distortions (especially for JPEG compression with low quality factors), which deteriorate the attack ability and increase the unreliability. To address these issues, we propose a dual-branch guided generative model called Heterogeneous Encoding Network (HENet) to form a robust generator-based adversarial attack framework. Specifically, our HENet introduces an Adaptive Feature Fusion Module (AFFM) to solve the dimensions and representativeness contradictions between CNNs and Transformers, which steers the perturbation generation based on a richer latent space and achieves better general attack ability. To further improve the robustness against JPEG compression, we design and integrate a Dynamic Differentiable JPEG Simulator (DDJS), which introduces an adaptive quantization mask to determine the flow of the gradient backpropagation in each frequency position. Extensive experiments prove the proposed method achieves a better attack success rate, lower perturbation magnitude, and higher robustness for various target network architectures under compressed, distorted, and lossless scenarios. Our codes will be made publicly available.
Jiawei Zhang 0011, Hao Wang 0060, Hao Wu 0078, Bin Li 0011, Xiangyang Luo 0001, Bin Ma 0003
IEEE Trans. Inf. Forensics Secur.2
2025 DVW: Diffusion Visible Watermark
abstract
With the rapid development of the diffusion models, numerous exquisitely generated images have significantly increased the risk of image misuse and abuse. Despite various AI parties and companies having devoted themselves to embedding watermarks into the generated images to curb the potential detriments, the isolated embedding from the generation process makes the watermarks vulnerable to watermark removal networks. To address this issue, we propose a novel generative image watermark scheme, dubbed Diffusion Visible Watermark (DVW), which can generate watermarked images in one step without additional training or fine-tuning of the diffusion models. Specifically, DVW introduces a masked distribution alignment strategy to fuse the watermark distribution with a Gaussian noise distribution. By iterative denoising the fused aligned distribution with the pretraining diffusion models, the watermarked images with coordinated and unified distribution can be generated with natural robustness against removal. In addition, we design and integrate a dynamic transparency module to adaptively control the watermark coverage degree for better visual quality. Comprehensive experiments and analysis are conducted on two representative kinds of diffusion models, GLIDE and StableDiffusion, to prove the superior and generic robustness of our DVW against watermark removal without sacrificing the generation ability of the diffusion models.
Jiawei Zhang 0011, Xiaoli Jiang, Hao Wang 0060, Lin Yuan 0002, Xiangyang Luo 0001, Bin Ma 0003
ACM Multimedia3
2025 LDSGAN: Unsupervised Image-to-Image Translation With Long-Domain Search GAN for Generating High-Quality Anime Images
abstract
Image‐to‐image ( I2I ) translation has emerged as a valuable tool for privacy protection in the digital age, offering effective ways to safeguard portrait rights in cyberspace. In addition, I2I translation is applied in real‐world tasks such as image synthesis, super‐resolution, virtual fitting, and virtual live streaming. Traditional I2I translation models demonstrate strong performance when handling similar datasets. However, when the domain distance between two datasets is large, translation quality may degrade significantly due to notable differences in image shape and edges. To address this issue, we propose Long‐Domain Search GAN ( LDSGAN ), an unsupervised I2I translation network that employs a GAN structure as its backbone, incorporating a novel Real‐Time Routing Search ( RTRS ) module and Sketch Loss. Specifically, RTRS aids in expanding the search space within the target domain, aligning feature projection with images closest to the optimization target. Additionally, Sketch Loss retains human visual similarity during long‐domain distance translation. Experimental results indicate that LDSGAN surpasses existing I2I translation models in both image quality and semantic similarity between input and generated images, as reflected by its mean FID and LPIPS scores of 31.509 and 0.581, respectively.
Hao Wang 0060, Chenbin Wang, Xin Cheng 0018, Hao Wu 0078, Jiawei Zhang 0011, Xiangyang Luo 0001, Bin Ma 0003
Int. J. Intell. Syst.1
2025 A GAN-based anti-forensics method by modifying the quantization table in JPEG header file
Hao Wang 0060, Xin Cheng 0018, Hao Wu 0078, Xiangyang Luo 0001, Bin Ma 0003, Hui Zong, Jiawei Zhang 0011
J. Vis. Commun. Image Represent.1
2024 Advancing Quantization Steps Estimation: A Two-Stream Network Approach for Enhancing Robustness
Xin Cheng 0018, Hao Wang 0060, Xiangyang Luo 0001, Bin Ma 0003
ACM Multimedia2
2024 Trustworthy adaptive adversarial perturbations in social networks
Jiawei Zhang 0011, Hao Wang 0060, Xiangyang Luo 0001, Bin Ma 0003
J. Inf. Secur. Appl.3
2024 Quantization Step Estimation of Color Images Based on Res2Net-C With Frequency Clustering Prior Knowledge
abstract
The quantization step is a crucial parameter in JPEG compression, that can reveal the compression history of a JPEG image. Estimating the quantization steps for single compressed and recompressed images is attracting considerable interest in the field of image forensics and steganalysis. Several effective methods have been proposed, but the performance of these methods still needs to be improved on small-sized and low-quality images. To solve the above problems, feature enrichment is performed on images in the frequency domain, resulting in clustering discrete cosine transform (DCT) coefficients of the same frequency. Then, we construct a hierarchical connection within the residual blocks of the network to represent multi-scale features, enabling the network to learn deep features of the image. At the same time, we use multiple small-sized convolution kernels instead of one large-sized convolution kernel to minimize the impact of block artifacts. Based on the above two ideas, we construct a network model, Res2Net-C, to discover information about the quantization steps in the frequency domain. The integration of multi-channel information of color images is achieved by multi-channel convolution, and the quantization steps of the chrominance and luminance channels of the color images are estimated. The experimental results show that the accuracy of the proposed method for estimating the quantization steps is 29.97% better than that of the existing algorithm with a single compressed dataset and 4.87% better than that of the existing algorithm with a recompressed image dataset. In addition, the method has good performance with mixed datasets that contain both single compressed and recompressed images.
Xin Cheng 0018, Hao Wang 0060, Xiangyang Luo 0001, Bin Ma 0003
IEEE Trans. Circuits Syst. Video Technol.3
2024 General Forensics for Aligned Double JPEG Compression Based on the Quantization Interference
abstract
Detection of aligned double Joint Photographic Experts Group (JPEG) compressed images is a crucial area of research within the field of digital image forensics. The detection tasks for aligned double JPEG compression can be categorized into two sub-tasks, namely detecting double JPEG images with the same quantization matrix (DJSQM) or double JPEG images with different quantization matrices (DJDQM). Existing methods for one of these sub-tasks may not be effective for the other. To address this issue, a novel approach is proposed by recompressing both DJDQM and DJSQM using modified quantization coefficients. The perturbation in the recompression process results in a perturbed error image, which is valid for both DJDQM and DJSQM. Subsequently, the relative change rate is used to combine the perturbed error image, the original error image, and the quantization error to derive the interference error and the interference quantization error. The interference error and interference quantization error further expand the difference between single and double compressed images by preserving the general validity of the original image information. Furthermore, the recompression process of DJDQM and DJSQM results in the conversion of truncation and rounding errors at the pixel level, which can be represented by the pixel state map. The pixel state map characterizes the differing transformation relationships between single and double compressed images and provides additional valid features, thereby enhancing the performance of the proposed method. The empirical results demonstrate that the proposed method outperforms existing methods on detecting aligned double JPEG compressed images.
Hao Wang 0060, Jiawei Zhang 0011, Xiangyang Luo 0001, Bin Ma 0003, Bin Li 0011, Jinsheng Sun
IEEE Trans. Circuits Syst. Video Technol.1
2023 Self-Recoverable Adversarial Examples: A New Effective Protection Mechanism in Social Networks
abstract
Nowadays, users upload numerous photos to social network platforms to share their daily lives. These photos contain numerous personal information, which can be easily captured by intelligent algorithms. To improve privacy security, we aim to form a protection mechanism by exploiting adversarial examples, which can mislead and disrupt intelligent algorithms. However, the existing adversarial attack lacks the study on recoverability and reversibility, which makes them unable to serve as an effective protection mechanism. To address this issue, we propose a recoverable generative adversarial network to generate self-recoverable adversarial examples. By modeling the adversarial attack and recovery as a united task, our method can minimize the error of the recovered examples while maximizing the attack ability, resulting in better recoverability of adversarial examples. To further boost the recoverability of these examples, we exploit a dimension reducer to optimize the distribution of adversarial perturbation. The experimental results prove that the adversarial examples generated by the proposed method present superior recoverability, attack ability, and robustness on different datasets and network architectures, which ensure its effectiveness as a protection mechanism in social networks.
Jiawei Zhang 0011, Hao Wang 0060, Xiangyang Luo 0001
IEEE Trans. Circuits Syst. Video Technol.3
2022 Detecting Aligned Double JPEG Compressed Color Image With Same Quantization Matrix Based on the Stability of Image
abstract
Joint photographic experts group (JPEG) compression is widely used in image processing and computer vision. Detecting double compressed JPEG images is a common problem in forensics and detecting compressed images with the same quantization matrix remains a challenging task. However, most existing methods were designed for detection in grayscale images and cannot fully use the unique characteristics of color images (such as the relationship between channels and color information). In addition, the performance of existing methods is unsatisfactory for low JPEG quality factors and in cross detection experiments. To solve these problems, we analyze the stability of a color image to obtain the convergence error and transposition error. According to the convergence characteristics of color JPEG images, the continuous compression by the same quantization matrix can make the JPEG image tend to be stable. The final stable state and the convergence process are determined by the number of compressions of the original image. Thus, continuously compressed JPEG images can be regarded as a continuous frame to obtain the convergence error. As the color image converges, its ability to resist interference decreases. To reflect the changes in anti-interference ability, the transposition operation is used to disturb the color JPEG image to obtain the transposition error. In addition, quaternion mapping is used to retain the relationship between continuously compressed JPEG images and enlarge the influence caused by transposition operation. In our experiments on several image databases, the proposed method outperforms existing methods in different settings.
Hao Wang 0060, Xiangyang Luo 0001, Yuhui Zheng, Bin Ma 0003, Jinsheng Sun, Sunil Kr. Jha
IEEE Trans. Circuits Syst. Video Technol.1
2021 Modify the Quantization Table in the JPEG Header File for Forensics and Anti-forensics
Hao Wang 0060, Xiangyang Luo 0001, Qilin Yin, Bin Ma 0003, Jinsheng Sun
IWDW1
2020 Detecting Double JPEG Compressed Color Images With the Same Quantization Matrix in Spherical Coordinates
abstract
Detection of double Joint Photographic Experts Group (JPEG) compression is an important part of image forensics. Although methods in the past studies have been presented for detecting the double JPEG compression with a different quantization matrix, the detection of double JPEG compression with the same quantization matrix is still a challenging problem. In this paper, an effective method to detect the recompression in the color images by using the conversion error, rounding error, and truncation error on the pixel in the spherical coordinate system is proposed. The randomness of truncation errors, rounding errors, and quantization errors result in random conversion errors. The pixel number of the conversion error is used to extract six-dimensional features. Truncation error and rounding error on the pixel in its three channels are mapped to the spherical coordinate system based on the relation of a color image to the pixel values in the three channels. The former is converted into amplitude and angles to extract 30-dimensional features and 8-dimensional auxiliary features are extracted from the number of special points and special blocks. As a result, a total of 44-dimensional features have been used in the classification by using the support vector machine (SVM) method. Thereafter, the support vector machine recursive feature elimination (SVMRFE) method is used to improve the classification accuracy. The experimental results show that the performance of the proposed method is better than the existing methods.
Hao Wang 0060, Jian Li 0034, Xiangyang Luo 0001, Yun Q. Shi 0001, Sunil Kr. Jha
IEEE Trans. Circuits Syst. Video Technol.2