Wenkang Su 0001

dblp:169/0709-1 · DBLP profile ↗
← Back
20ranked-venue papers
5as first author
17since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 13 · 4 first-author · 12 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Security and privacy · 3 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 SemiDDM-weather: A semi-supervised learning framework for all-in-one adverse weather removal
Fang Long, Wenkang Su 0001, Mingjie Li 0004, Yuan-Gen Wang, Xiaochun Cao
Neural Networks2
2026 Enhancing Cross-Domain Correspondence for Unsupervised Image-to-Image Translation
abstract
UNsupervised Image-to-image Translation (UNIT) aims to translate images across visual domains without paired training data, which has been widely used in style transfer, image processing, game design, etc. However, ensuring the correspondence (e.g., target category, pose, or head orientation) between generated images and inputs remains a formidable challenge. To this end, we present a new scheme, named EC-UNIT, which comprises three innovative designs aiming to Enhance cross domain Correspondence for UNIT. Specifically, 1) we propose Multi-level Style Embedding to extract multi-level style features for fusion while imposing our newly designed Hierarchical Consistency Constraints on both the content and style features (MSE&HCC), aiming to retain more style representations and facilitate feature disentanglement; 2) we develop Semantic Perceptual Matching (SPM) to minimize the semantic distribution discrepancy between the generated image and the input image by leveraging the multimodal model CLIP, dedicated to enhancing semantic consistency; 3) considering that previous works have struggled to control the image translation using pixel-level visual consistency constraints, we design Visual Perceptual Guidance (VPG) to reduce the perceptual distance between the generated image and the style input in VGG feature space, devoted to enhancing visual perceptual correspondence, thereby preventing the generation of unrealistic image details. Extensive experiments demonstrate that our EC-UNIT is more stable and outperforms current SOTA competitors in terms of image quality and diversity as well as both content and style consistency.
Binxin Lai, Wenkang Su 0001, Yuying Liang, Yuan-Gen Wang, Mingjie Li 0004, Jiantao Zhou 0001
IEEE Trans. Multim.2
2026 Uncertainty-driven Progressive Single Image De-raining
abstract
Over the past years, progressive methods have demonstrated promising performance in single image de-raining task. Nonetheless, current methods still struggle to precisely remove rain and preserve more image details during the progressive de-raining process, resulting in undesirable local artifacts or image detail loss. To tackle these limitations, a novel progressive approach, called Uncertainty-driven Progressive Single Image De-raining (UPSID), is proposed. Firstly, a powerful internal-and-external dense sub-network is devised, which effectively integrates three proper and flexible components, including dense connection, long short-term memory, and channel attention. Subsequently, the sub-network is further unfolded into multiple recurrent stages to form a progressive de-raining network. Finally, the overall progressive de-raining network is trained with an adaptive weighted loss to focus more on challenging pixels that characterize rain or texture/edge regions. Extensive quantitative and qualitative experiments confirm that the proposed UPSID outperforms multiple state-of-the-art algorithms, including single-stage, progressive, and uncertainty-driven single image de-raining methods. Additionally, this article also demonstrates the superiority of UPSID for other similar image restoration tasks such as single image de-snowing. The code will be publicly available at https://github.com/Lcai-QZ/UPSID .
Jianqing Zhu, Huanqiang Zeng, Tao Zhu 0002, Jing Chen 0001, Wenkang Su 0001
ACM Trans. Multim. Comput. Commun. Appl.7
2025 Toward Real-world Text Image Forgery Localization: Structured and Interpretable Data Synthesis
abstract
Existing Text Image Forgery Localization (T-IFL) methods often suffer from poor generalization due to the limited scale of real-world datasets and the distribution gap caused by synthetic data that fails to capture the complexity of real-world tampering. To tackle this issue, we propose Fourier Series-based Tampering Synthesis (FSTS), a structured and interpretable framework for synthesizing tampered text images. FSTS first collects 16,750 real-world tampering instances from five representative tampering types, using a structured pipeline that records human-performed editing traces via multi-format logs (e.g., video, PSD, and editing logs). By analyzing these collected parameters and identifying recurring behavioral patterns at both individual and population levels, we formulate a hierarchical modeling framework. Specifically, each individual tampering parameter is represented as a compact combination of basis operation–parameter configurations, while the population-level distribution is constructed by aggregating these behaviors. Since this formulation draws inspiration from the Fourier series, it enables an interpretable approximation using basis functions and their learned weights. By sampling from this modeled distribution, FSTS synthesizes diverse and realistic training data that better reflect real-world forgery traces. Extensive experiments across four evaluation protocols demonstrate that models trained with FSTS data achieve significantly improved generalization on real-world datasets. Dataset is available at \href{https://github.com/ZeqinYu/FSTS}{Project Page}.
Zeqin Yu, Haotao Xie, Jiangqun Ni, Wenkang Su 0001, Jiwu Huang
NeurIPS5
2025 DRR: A new method for multiple adverse weather removal
Fang Long, Wenkang Su 0001, Yuan-Gen Wang, Qingxiao Guan
Expert Syst. Appl.3
2025 HiFiMSFA: Robust and High-Fidelity Image Watermarking Using Attention Augmented Deep Network
abstract
In recent years, the popularity of digital media sharing, especially high-quality images through online social networks (OSNs) has spurred an increasing demand for digital rights management (DRM) with watermarking. Although the most recent watermarking schemes with deep networks have exhibited considerable performance improvement, they still fall short in resisting multiple attacks with high-fidelity watermarking. To tackle this issue, a customized framework with encoder/decoder structure is proposed in this letter, aiming to consistently improve the robustness performance against multiple attacks. In specific, theMulti-scaleSalientFeatureAttentionBlock(MSFABlock) is exploited to effectively extract the robust image features with the encoder and decoder by taking advantage of the salient features, e.g., the image features obtained with difference of Gaussian (DoG) and other gradient operators. In addition, an adaptive squared Hinge function is developed as message loss to encourage adaptive watermark embedding. Experimental results demonstrate excellent performance in terms of robustness and perceptual fidelity as well as high efficiency of the proposed scheme in comparison to other SOTA methods.
Jiangqun Ni, Wenkang Su 0001
IEEE Signal Process. Lett.3
2024 StegaStyleGAN: Towards Generic and Practical Generative Image Steganography
abstract
The recent advances in generative image steganography have drawn increasing attention due to their potential for provable security and bulk embedding capacity. However, existing generative steganographic schemes are usually tailored for specific tasks and are hardly applied to applications with practical constraints. To address this issue, this paper proposes a generic generative image steganography scheme called Steganography StyleGAN (StegaStyleGAN) that meets the practical objectives of security, capacity, and robustness within the same framework. In StegaStyleGAN, a novel Distribution-Preserving Secret Data Modulator (DP-SDM) is used to achieve provably secure generative image steganography by preserving the data distribution of the model inputs. Additionally, a generic and efficient Secret Data Extractor (SDE) is invented for accurate secret data extraction. By choosing whether to incorporate the Image Attack Simulator (IAS) during the training process, one can obtain two models with different parameters but the same structure (both generator and extractor) for lossless and lossy channel covert communication, namely StegaStyleGAN-Ls and StegaStyleGAN-Ly. Furthermore, by mating with GAN inversion, conditional generative steganography can be achieved as well. Experimental results demonstrate that, whether for lossless or lossy communication channels, the proposed StegaStyleGAN can significantly outperform the corresponding state-of-the-art schemes.
Wenkang Su 0001, Jiangqun Ni, Yiyan Sun
AAAI1
2024 Model-Based Non-Independent Distortion Cost Design for Effective JPEG Steganography
abstract
Recent achievements have shown that model-based steganographic schemes hold promise for better security than heuristic-based ones, as they can provide theoretical guarantees on secure steganography under a given statistical model. However, it remains a challenge to exploit the correlations between DCT coefficients for secure steganography in practical scenarios where only a single compressed JPEG image is available. To cope with this, we propose a novel model-based steganographic scheme using the Conditional Random Field (CRF) model with four-element cross-neighborhood to capture the dependencies among DCT coefficients for JPEG steganography with symmetric embedding. Specifically, the proposed CRF model is characterized by the delicately designed energy function, which is defined as the weighted sum of a series of unary and pairwise potentials, where the potentials associated with the statistical detectability of steganography are formulated as the KL divergence between the statistical distributions of cover and stego. By optimizing the constructed energy function with the given payload constraint, the non-independent distortion cost corresponding to the least detectability can be accordingly obtained. Extensive experimental results validate the effectiveness of our proposed scheme, especially outperforming the previous independent art J-MiPOD.
Yuanfeng Pan, Wenkang Su 0001, Jiangqun Ni, Qingliang Liu 0001, Donghua Jiang 0001
ACM Multimedia2
2024 Unifying Homophily and Heterophily for Spectral Graph Neural Networks via Triple Filter Ensembles
abstract
Polynomial-based learnable spectral graph neural networks (GNNs) utilize polynomial to approximate graph convolutions and have achieved impressive performance on graphs. Nevertheless, there are three progressive problems to be solved. Some models use polynomials with better approximation for approximating filters, yet perform worse on real-world graphs. Carefully crafted graph learning methods, sophisticated polynomial approximations, and refined coefficient constraints leaded to overfitting, which diminishes the generalization of the models. How to design a model that retains the ability of polynomial-based spectral GNNs to approximate filters while it possesses higher generalization and performance? In this paper, we propose a spectral GNN with triple filter ensemble (TFE-GNN), which extracts homophily and heterophily from graphs with different levels of homophily adaptively while utilizing the initial features. Specifically, the first and second ensembles are combinations of a set of base low-pass and high-pass filters, respectively, after which the third ensemble combines them with two learnable coefficients and yield a graph convolution (TFE-Conv). Theoretical analysis shows that the approximation ability of TFE-GNN is consistent with that of ChebNet under certain conditions, namely it can learn arbitrary filters. TFE-GNN can be viewed as a reasonable combination of two unfolded and integrated excellent spectral GNNs, which motivates it to perform well. Experiments show that TFE-GNN achieves high generalization and new state-of-the-art performance on various real-world datasets.
Rui Duan 0003, Mingjian Guang, Junli Wang 0001, ChunGang Yan, Hongda Qi, Wenkang Su 0001, Can Tian, Haoran Yang 0003
NeurIPS6
2024 An efficient distortion cost function design for image steganography in spatial domain using quaternion representation
Qingliang Liu 0001, Wenkang Su 0001, Jiangqun Ni, Xianglei Hu, Jiwu Huang
Signal Process.2
2024 Efficient JPEG image steganography using pairwise conditional random field model
Yuanfeng Pan, Jiangqun Ni, Qingliang Liu 0001, Wenkang Su 0001, Jiwu Huang
Signal Process.4
2024 DWW: Robust Deep Wavelet-Domain Watermarking With Enhanced Frequency Mask
abstract
This letter concentrates on the challenges of deep learning-based robust image watermarking against print-scanning, print-camera, and screen-shooting attacks for “physical channel transmission”. Given the excellent performance demonstrated by wavelet domain watermarking, in this paper, we incorporate the wavelet integrated convolutional neural networks (CNNs) and propose a Deep Wavelet-domain Watermarking (DWW) model, which is dedicated to embedding watermarks in the wavelet domain rather than the spatial domain of the previous arts. In addition, a frequency-domain enhanced mask loss is developed to increase the loss weight in the high-frequency regions of the image during back-propagation, thereby encouraging the model to embed the message in low-frequency components with priority so as to improve the robustness performance. Experiment results show that the proposed DWW consistently outperforms other state-of-the-art (SOTA) schemes by a clear margin in terms of embedding capacity, imperceptibility, and robustness.
Shiyuan Tang, Jiangqun Ni, Wenkang Su 0001
IEEE Signal Process. Lett.3
2024 Efficient Audio Steganography Using Generalized Audio Intrinsic Energy With Micro-Amplitude Modification Suppression
abstract
Recent advances in content-adaptive Audio Steganography in Temporal Domain (ASTD) suggest that modification of micro-amplitude samples may compromise its security. To prevent the micro-amplitude samples from being modified, a targeted Large Amplitude First (LAF) rule was adopted in some audio steganographic schemes, e.g., DFR. However, it is observed that the results with LAF rule are often unstable across different datasets, we thus propose a new Micro-Amplitude Suppression (MAS) rule in this paper following the design philosophy of wet paper coding. Unlike DFR where the audio steganographic performance heavily depends on the adopted heuristic filters, we propose to evaluate the embedding cost of cover audio with the Generalized Audio Intrinsic Energy (GAIE), which is obtained by calculating the weighted sum of squared DCT coefficients for each segmented audio clip with carefully designed weights. Extensive experimental results demonstrate that the proposed MAS rule tends to be more general and consistent than the LAF rule, and the proposed GAIE also shows better empirical security performance and audio quality compared to the advanced AAC and DFR_res (a variant of DFR). In addition, by preventing the micro-amplitude samples from being modified, the proposed GAIE_MAS can not only outperform other hand-crafted audio steganographic schemes but also the recently emerged deep learning-based schemes, e.g., IAA.
Wenkang Su 0001, Jiangqun Ni, Xianglei Hu, Bin Li 0011
IEEE Trans. Inf. Forensics Secur.1
2023 A Novel Deep Video Watermarking Framework with Enhanced Robustness to H.264/AVC Compression
abstract
The recent success of deep image watermarking has demonstrated the potential of deep learning for watermarking, which has drawn increasing attention to deep video watermarking with the objective to improve its robustness and perceptual quality. Compared to images, video watermarking is much more challenging due to the rich structures of video data and the diversity of attacks in video transmission pipeline. The existing deep video watermarking schemes are far from satisfactory in dealing with temporal attacks, e.g., frame averaging, frame dropping and transcoding. To this end, a novel deep framework for Robustness Enhanced Video watermarking (REVMark) is proposed in this paper, aiming at improving the overall robustness, especially in dealing with H.264/AVC compression, while maintaining good visual quality. REVMark has an encoder/decoder structure with a pre-processing block (TAsBlock) to effectively extract the temporal-associated features on aligned frames. To ensure the end-to-end robust training, a distortion layer is integrated into the REVMark to resemble various attacks in real-world scenarios, among which, a new differentiable simulator of video compression, namely DiffH264, is developed to approximately simulate the process of H.264/AVC compression. In addition, the mask loss is incorporated to guide the encoder to embed the watermark in the human-imperceptible regions, thus improving the perceptual quality of the watermarked video. Experimental results demonstrate that the proposed scheme can outperform other SOTA methods while achieving 10X faster inference.
Jiangqun Ni, Wenkang Su 0001, Xin Liao 0001
ACM Multimedia3
2022 Wavelet-Based CNN for Robust and High-Capacity Image Watermarking
abstract
“Physical” watermarking, i.e., print-scanning, print-camera, and screen-shooting resilient watermarking, has drawn great attention these years. Recent studies show that watermarking with deep networks, e.g., StegaStamp, is particularly suitable for these tasks with strong robustness against the attacks from “physical transmissions” in the wild, although it usually has a relatively small amount of embedding capacity. Recognizing the conventional CNNs are vulnerable to input noise inter-ruptions, the wavelet-based CNNs are adopted in our work by replacing their down-sampling (pooling) and up-sampling layers with Discrete Wavelet Transform (DWT) and Inverse Discrete Wavelet Transform (IDWT), respectively, to learn the stable feature representation from noise-corrupted sam-ples. A new residual regularization loss function incorporating the texture complexity is also proposed to significantly improve the visual quality of watermarked images. Exper-imental results show that the proposed wavelet-based CNN model significantly outperforms the state-of-the-art StegaS-tamp in terms of embedding capacity, imperceptibility, and robustness.
Junxiong Lu, Jiangqun Ni, Wenkang Su 0001, Hao Xie 0002
ICME3
2022 New design paradigm of distortion cost function for efficient JPEG steganography
Wenkang Su 0001, Jiangqun Ni, Xianglei Hu, Jiwu Huang
Signal Process.1
2021 Image Steganography With Symmetric Embedding Using Gaussian Markov Random Field Model
abstract
Recent advances on adaptive steganography show that the performance of image steganographic communication can be improved by incorporating the non-additive models that capture the dependencies among adjacent pixels. In this paper, a Gaussian Markov Random Field model (GMRF) with four-element cross neighborhood is proposed to characterize the interactions among local elements of cover images, and the problem of secure image steganography is formulated as the one of minimization of KL-divergence in terms of a series of low-dimensional clique structures associated with GMRF by taking advantages of the conditional independence of GMRF. The adoption of the proposed GMRF tessellates the cover image into two disjoint subimages, and an alternating iterative optimization scheme is developed to effectively embed the given payload while minimizing the total KL-divergence between cover and stego, i.e., the statistical detectability. Experimental results demonstrate that the proposed GMRF outperforms the prior arts of model based schemes, e.g., MiPOD, and rivals the state-of-the-art HiLL for practical steganography, where the selection channel knowledges are unavailable to steganalyzers.
Wenkang Su 0001, Jiangqun Ni, Xianglei Hu, Jessica J. Fridrich
IEEE Trans. Circuits Syst. Video Technol.1
2019 Image Steganography Using an Eight-Element Neighborhood Gaussian Markov Random Field Model
Yichen Tong, Jiangqun Ni, Wenkang Su 0001
IWDW3
2018 A New Distortion Function Design for JPEG Steganography Using the Generalized Uniform Embedding Strategy
abstract
Nowadays, the most prevailing approach to steganography is the minimal embedding distortion framework, which includes an optimizable distortion function for each cover element and an encoding method to minimize the distortion. With the emergence of Syndrome-Trellis Code, the distortion function plays an increasingly important role in modern adaptive image steganography. In this letter, a new distortion function called generalized uniform embedding distortion (GUED) is proposed for JPEG steganography. The proposed GUED consists of the new distortion measures for both Alternating Current (AC) mode and Discrete Cosine Transform (DCT) block, which are represented in a more general exponential model, aiming to flexibly allocate the embedding data so as to minimize the global changes of the statistics of quantized DCT coefficients after embedding. In addition, an empirical rule is developed to determine the parameters of the exponential function according to the payload and quality factor. By exploring the statistics of both DCT and spatial domains, the proposed GUED is shown to be more consistent with the objective of generalized uniform embedding strategy, i.e., maintaining the relative changes of DCT coefficients to be proportional to their coefficients of variations. Extensive experiments demonstrate that the proposed GUED gains significant performance improvements when compared with its original UERD, and outperforms the state-of-the-art J-UNIWARD with markedly reduced computation time.
Wenkang Su 0001, Jiangqun Ni, Yun Q. Shi 0001
IEEE Trans. Circuits Syst. Video Technol.1
2015 Using Statistical Image Model for JPEG Steganography: Uniform Embedding Revisited
abstract
Uniform embedding was first introduced in 2012 for non-side-informed JPEG steganography, and then extended to the side-informed JPEG steganography in 2014. The idea behind uniform embedding is that, by uniformly spreading the embedding modifications to the quantized discrete cosine transform (DCT) coefficients of all possible magnitudes, the average changes of the first-order and the second-order statistics can be possibly minimized, which leads to less statistical detectability. The purpose of this paper is to refine the uniform embedding by considering the relative changes of statistical model for digital images, aiming to make the embedding modifications to be proportional to the coefficient of variation. Such a new strategy can be regarded as generalized uniform embedding in substantial sense. Compared with the original uniform embedding distortion (UED), the proposed method uses all the DCT coefficients (including the DC, zero, and non-zero AC coefficients) as the cover elements. We call the corresponding distortion function uniform embedding revisited distortion (UERD), which incorporates the complexities of both the DCT block and the DCT mode of each DCT coefficient (i.e., selection channel), and can be directly derived from the DCT domain. The effectiveness of the proposed scheme is verified with the evidence obtained from the exhaustive experiments using a popular steganalyzer with rich models on the BOSSbase database. The proposed UERD gains a significant performance improvement in terms of secure embedding capacity when compared with the original UED, and rivals the current state-of-the-art with much reduced computational complexity.
Linjie Guo, Jiangqun Ni, Wenkang Su 0001, Chengpei Tang, Yun Q. Shi 0001
IEEE Trans. Inf. Forensics Secur.3