Wenbo Wan

dblp:142/0297 · DBLP profile ↗
← Back
60ranked-venue papers
6as first author
37since 2021 · last 2026
0000-0003-1447-0524ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 32 · 3 first-author · 14 since 2021Artificial intelligence and machine learning · 17 · 3 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 7 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Security and privacy · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1
YearPublicationVenuePosition
2026 SCIA-GAN: Robust image watermarking via spatial-channel interaction attention and feature preservation
Lingchen Gu, Jun Wang 0061, Wenbo Wan, Jiande Sun 0001, Sen-Ching S. Cheung
Expert Syst. Appl.5
2026 Task-driven infrared and visible image fusion via detail and semantic dual injection
Kai Zhang 0010, Ludan Sun, Feng Zhang 0028, Wenbo Wan, Jiande Sun 0001
Neural Networks5
2026 A universal pansharpening network via spatial-spectral contrastive learning
Kai Zhang 0010, Yunlong Liu 0005, Feng Zhang 0028, Wenbo Wan, Lingchen Gu, Jiande Sun 0001
Pattern Recognit.5
2026 Quality-Guided Forgery Adapter for Generalizable AIGC Image Detection
abstract
The rapid advancement of AI-generated content (AIGC) presents significant challenges for digital forensics, necessitating robust and generalizable detection frameworks. Existing detection methods primarily rely on visual feature extraction, while vision-language model-based approaches are limited to class-label prompts, failing to capture quality-related artifacts introduced by different generative models. To address this limitation, we introduce QAFD, a novel Quality-Assisted Forgery Detection framework that incorporates image quality information into the detection process. Specifically, we design a quality queried attention block to effectively fuse class-based content prompts with quality-aware text prompts. This integration enhances the model’s ability to capture semantic artifacts related to degradation patterns commonly associated with AI-generated images. Furthermore, we introduce the Quality-Guided Forgery Adapter (QGFA) to incorporate quality-aware textual cues into the visual domain, improving feature extraction for both spatial and frequency-based forgery artifacts. This synergy allows frequency cues to enhance low-level artifact perception, while quality-aware guidance strengthens high-level discriminative representation. Extensive experiments demonstrate that QAFD achieves superior generalization to unseen generative models over three datasets and significantly maintains its robustness against common image post-processing operations.The codes will be released at github.
Jun Wang 0061, Zitong Yu, Chaomeng Chen, Lingchen Gu, Wenbo Wan, Jiantao Zhou 0001, Weiming Zhang 0001
IEEE Trans. Inf. Forensics Secur.5
2025 WPM-GAN: Watermark-Preserving Module for GAN-Based Robust Industrial Image Watermarking
abstract
In this paper, a robust watermarking framework is proposed for ensuring copyright and authenticity of industrial imagery. Leveraging Squeeze-and-Excitation (SE) blockbased encoder-decoder network, our method embeds and extracts watermarks in a more efficient and imperceptible manner, and our method introduces a novel Watermark-Preserving Module (WPM) at the receiving end, which maximizes the preservation of watermark features in noise-corrupted images to assist the decoder in watermark extraction. And we adopt a dualstage training strategy to capture and learn specific watermark features, along with a decoder-based watermark-preserving loss to further enhance robustness. Experimental results demonstrate that WPM-GAN achieves superior visual quality while effectively resisting various attacks.
Yingchao Yang, Lingchen Gu, Wenbo Wan
ICPADS5
2025 Orientation-Aware Reversible Data Hiding With Brainstorming Optimization for UAV Aerial Images
abstract
In recent years, with the rapid development of unmanned aerial vehicle (UAV), aerial images have extended across various industries such as intelligent building, agriculture, transportation, and Industry 4.0. Notably, the security of UAV‐assisted data acquisition during transmission has become a critical concern. The reversible data hiding (RDH) method can hide data in aerial images for transmission and ensure secure communication. In general, an aerial image may exhibit substantially different orientation regularity from a natural scene image. This casts major challenges to the RDH method, for which existing approaches lack effective mechanisms to capture such content type variations, and thus are difficult to generalize from one type to another. In this paper, the orientation‐aware selectivity mechanism is introduced to achieve an accurate orientation‐aware prediction along different directions in local regions with different structure regularity. Furthermore, we propose a progressive brainstorming optimization algorithm (BSO)‐guided optimal PSNR value strategy, which can obtain a superior perceptual performance and the corresponding thresholds by further exploring the pixel correlations within the UAV aerial images. Experimental results on the USC‐SIPI Miscellaneous dataset and two challenging aerial datasets, including the USC‐SIPI High Altitude Aerial Imagery dataset and the Kaggle dataset, demonstrate that the proposed framework enhances the imperceptibility powerfully in marked UAV aerial images and ensures sufficient embedding capacity effectively. The average PSNR of the marked image obtained by the proposed method is 63.85 dB when embedded with 30,000 bits of data, which is an improvement of 0.59 dB compared to the current state‐of‐the‐art RDH methods.
Xiaodan Tai, Yannan Ren, Jing Li 0046, Jiande Sun 0001, Kai Zhang 0010, Wenbo Wan
Int. J. Intell. Syst.6
2025 Multiscale Integration Network With Quaternion Convolution for Pansharpening
abstract
In this letter, we proposed a multiscale integration network with quaternion convolution (MQ-Net) for the fusion of low spatial resolution multispectral (LRMS) and panchromatic (PAN) images. In this network, LRMS and PAN images are resampled at different scales and fed into feature fusion modules (FFMs) to merge the spatial and spectral information among them. Then, multiscale feature enhancement modules (MFEMs) are designed to sufficiently learn the spatial and spectral information at different scales. Meanwhile, we employ a quaternion convolution module (QCM) to better capture the dependencies within spectral bands of LRMS images. Then, the quaternion features are introduced into MFEMs for efficient feature enhancement. Finally, all information from different scales is integrated for the reconstruction of high LRMS images. Reduced- and full-resolution experiments are performed on GeoEye-1 and WorldView-2 satellite datasets. Compared to some state-of-the-art pansharpening methods, the proposed MQ-Net obtains better results in terms of qualitative and quantitative evaluations. The code is available athttps://github.com/RSMagneto/MQ-Net.
Yingjie Kong, Xuquan Wang, Kai Zhang 0010, Hong Li 0005, Wenbo Wan, Jiande Sun 0001
IEEE Geosci. Remote. Sens. Lett.5
2025 Spatial-spectral unfolding network with mutual guidance for multispectral and hyperspectral image fusion
Kai Zhang 0010, Qinzhu Sun, Chiru Ge, Wenbo Wan, Jiande Sun 0001, Huaxiang Zhang 0001
Pattern Recognit.5
2025 Dual Prototypes-Based Personalized Federated Adversarial Cross-Modal Hashing
abstract
With the rapid advances in wireless communication and IoT platforms, it is increasingly difficult to analyze relevant multi-modal data distributed across geographically diverse and heterogeneous platforms. One promising approach is to rely on federated learning to build compact cross-modal hash codes. However, existing federated learning methods easily exhibit degenerative performance in the global model due to the distributed data being derived from diverse domains. In addition, directly forcing each client to adopt the same global parameters as local parameters, without effective local training, significantly reduces the performance of each client. To overcome these challenges, we propose a novel federated adversarial cross-modal hashing, called Dual Prototypes-based personalized Federated Adversarial (DP-FeAd), which provides iterated training of shared dual prototypes. Specifically, aiming to expand local hashing models beyond their knowledge realms, DP-FeAd enables participating clients to engage in cooperative learning through two constructions: cluster prototypes and unbiased prototypes, instead of the traditional global prototypes, ensuring both generalization and stability. Specifically, the cluster prototypes are derived from local class-level prototypes and adversarially trained with local approximate hash codes to align their distributions. The unbiased prototypes are averaged from cluster prototypes and integrated into the training of local hashing models to maintain consistency across different local class-level prototypes further. The experiments conducted on two benchmark datasets demonstrate that our proposed method significantly enhances the performance of deep cross-modal hashing models in both IID (Independent and Identically Distributed) and non-IID scenarios.
Lingchen Gu, Xiaojuan Shen, Jiande Sun 0001, Jing Li 0046, Zhihui Li 0001, Sen-Ching S. Cheung, Wenbo Wan
IEEE Trans. Circuits Syst. Video Technol.8
2025 RP-ASAF: Anonymous Submission of Application Framework Using RDHSI and Polynomial Interpolation
abstract
Reversible data hiding (RDH) is one special type of data hiding, and is widely used for many intended applications. Moreover, to protect the cover image, RDH in encrypted image (RDHEI) schemes are accordingly proposed. In RDHEI, if the marked image is lost, the cover image and the secret data cannot be restored. To address the issue of losing sub marked images, RDH in shared image (RDHSI) by secret sharing (SS) is proposed to achieve fault-tolerance. Recently, an anonymous application scenario using RDHSI is introduced, on which multiple data hiders in RDHSI serve as reviewers in a committee. When the number of agreement votes is above the threshold of SS, the application is approved. However, there are weaknesses in this application framework: (i) reviewers cannot provide equal right to vote (namely, they cannot cast “Yes” and “No” votes, respectively), and (ii) anonymous submission is not really achieved. In the paper, we propose a new RDHSI to allow data hiders can hide approved sub data or disapproved sub data into sub images. Therefore, how to perform recovery and extraction from n marked images (some embedded with “Yes” votes and some with “No” votes) should be carefully designed. Based on the consistency property of polynomial interpolation, we conduct verification, recovery and extraction algorithms from n marked images. In addition, we use Hamming codewords to represent pixel difference instead of directly hiding pixel difference for reversibility, and this improvement also improves the embedding capacity. Compare with the current anonymous application scheme using RDHSI, the embedding rate in this paper can reach 3.5 bits per pixel (bpp), which is an improvement of 2 bpp. Thus, it is more suitable for the anonymous submission of application framework.
Ching-Nung Yang, Lizhi Xiong, Shu-Yu Liu, Chih-Yueh Tseng, Xiaodan Tai, Wenbo Wan
IEEE Trans. Circuits Syst. Video Technol.6
2025 Texture-Content Dual Guided Network for Visible and Infrared Image Fusion
abstract
The preservation and enhancement of texture information is crucial for the fusion of visible and infrared images. However, most current deep neural network (DNN)-based methods ignore the differences between texture and content, leading to unsatisfactory fusion results. To further enhance the quality of fused images, we propose a texture-content dual guided (TCDG-Net) network, which produces the fused image by the guidance inferred from source images. Specifically, a texture map is first estimated jointly by combining the gradient information of visible and infrared images. Then, the features learned by the shallow feature extraction (SFE) module are enhanced with the guidance of the texture map. To effectively model the texture information in the long-range dependencies, we design the texture-guided enhancement (TGE) module, in which the texture-guided attention mechanism is utilized to capture the global similarity of the texture regions in source images. Meanwhile, we employ the content-guided enhancement (CGE) module to refine the content regions in the fused result by utilizing the complement of the texture map. Finally, the fused image is generated by adaptively integrating the enhanced texture and content information. Extensive experiments on three benchmark datasets demonstrate the effectiveness of the proposed TCDG-Net in terms of qualitative and quantitative evaluations. Besides, the fused images generated by our proposed TCDG-Net also show better performance in downstream tasks, such as objection detection and semantic segmentation.
Kai Zhang 0010, Ludan Sun, Wenbo Wan, Jiande Sun 0001, Shuyuan Yang 0001, Huaxiang Zhang 0001
IEEE Trans. Multim.4
2024 Efficient Image Harmonization via RGB Transformation
abstract
Image harmonization aims to adjust the appearance of the foreground to make it harmonious with the background, thereby maintaining visual consistency in composite images. Previous deep learning-based methods have mainly focused on reconstructing harmonized images with the same size as the input composite images, often leading to complex network structures and a large number of parameters. In this paper, we propose a simple yet effective lightweight image harmonization network architecture. First, we generate a low-resolution 3-channel feature map to represent the rough variations in the RGB channels of a composite image, which is then upsampled and added to this composite image to obtain a preliminary harmonization result. Then, a refinement module is applied to refine the preliminary result and output the final harmonized image. Additionally, we design a dynamic data generation and training strategy to pre-train our model on another dataset. Experimental results on the iHarmony4 dataset show that our method indicates a significant reduction in the number of parameters compared to other methods, yet it still achieved competitive performance.
Jiande Sun 0001, Wenbo Wan, Kai Zhang 0010, Jian Wang 0004
MMSP4
2024 Triple disentangled network with dual attention for remote sensing image fusion
Feng Zhang 0028, Guishuo Yang, Jiande Sun 0001, Wenbo Wan, Kai Zhang 0010
Expert Syst. Appl.4
2024 Deep image watermarking with loss-driven modification
Wenqing Yang, Jing Li 0046, Jiande Sun 0001, Wenbo Wan
Multim. Tools Appl.7
2024 Building Change Detection in Earthquake: A Multiscale Interaction Network With Offset Calibration and a Dataset
abstract
As one of the most destructive natural disasters, earthquakes have struck many countries around the world in recent years, causing serious economic losses. Change detection (CD) can be applied to postearthquake building CD as it can infer interested change regions from multitemporal remote sensing (RS) images. Furthermore, the CD with short imaging intervals will better satisfy the needs of the emergency rescues after earthquakes. However, the capability of current methods built on deep neural networks (DNNs) is limited because the dataset with short imaging intervals is absent. To meet postdisaster immediate relief, we create a CD dataset, the Turkey earthquake CD dataset (TUE-CD), for the detection of building collapse in the short term after an earthquake. Due to the high requirement for timeliness of postevent images, the orbit of the satellite during postevent imaging deviates from that during preevent imaging, which leads to a side-looking problem between bitemporal images. To deal with these challenges, we present a multiscale feature interaction network (MSI-Net) for efficient interaction between bitemporal features, as well as mitigating the effect of side-looking problems. Specifically, the proposed MSI-Net consists of joint cross-attention (JCA) modules, multiscale offset calibration (MOC) modules, and feature integration (FeI) modules. The JCA module unifies channel cross-attention (CCA) and spatial joint attention (SJA) for sufficient feature interaction. The MOC module further estimates the offsets to align the bitemporal image with the multiscale features. Finally, calibrated features and multiscale features are fused by FeI modules for the prediction of changed areas. The best mF1 and mIoU scores are achieved on two public datasets and the constructed TUE-CD dataset: WHU-CD (95.58%, 91.81%), CLCD (82.96%, 73.53%), and TUE-CD (78.02%, 68.48%). Experimental results demonstrate that the proposed MSI-Net provides competitive performance compared to the state-of-the-art CD methods. The TUE-CD dataset and the code of MSI-Net will be available athttps://github.com/RSMagneto/MSI-Net.
Yunlong Liu 0005, Kai Zhang 0010, Chunan Guan, Shanxin Zhang, Hong Li 0005, Wenbo Wan, Jiande Sun 0001
IEEE Trans. Geosci. Remote. Sens.6
2024 Spectral-Spatial Dual Graph Unfolding Network for Multispectral and Hyperspectral Image Fusion
abstract
Recently, deep neural network (DNN)-based methods have achieved good results in terms of the fusion of low spatial resolution hyperspectral (LR HS) and high spatial resolution multispectral (HR MS) images. However, the spectral band correlation (SBC) and the spatial nonlocal similarity (SNS) in hyperspectral (HS) images are not sufficiently exploited by them. To model the two priors efficiently, we propose a spectral-spatial dual graph unfolding network (SDGU-Net), which is derived from the optimization of graph regularized restoration models. Specifically, we introduce spectral and spatial graphs to regularize the reconstruction of the desired high spatial resolution hyperspectral (HR HS) image. To explore the SBC and SNS priors of HS images in feature space and utilize the powerful learning ability of DNNs simultaneously, the iterative optimization of the spectral and spatial graph regularized models is unfolded as a network, which is composed of spectral and spatial graph unfolding modules. The two kinds of modules are designed according to the solutions of the spectral and spatial graph regularized models. In these modules, we employ graph convolution networks (GCNs) to capture the SBC and SNS in the fused image. Then, the learned features are integrated by the corresponding feature fusion modules and fed into the feature condense module to generate the HR HS image. We conduct extensive experiments on three benchmark datasets and the results demonstrate the effectiveness of our proposed SDGU-Net.
Kai Zhang 0010, Feng Zhang 0028, Chiru Ge, Wenbo Wan, Jiande Sun 0001, Huaxiang Zhang 0001
IEEE Trans. Geosci. Remote. Sens.5
2024 Entropy-Optimized Deep Weighted Product Quantization for Image Retrieval
abstract
Hashing and quantization have greatly succeeded by benefiting from deep learning for large-scale image retrieval. Recently, deep product quantization methods have attracted wide attention. However, representation capability of codewords needs to be further improved. Moreover, since the number of codewords in the codebook depends on experience, representation capability of codewords is usually imbalanced, which leads to redundancy or insufficiency of codewords and reduces retrieval performance. Therefore, in this paper, we propose a novel deep product quantization method, named Entropy Optimized deep Weighted Product Quantization (EOWPQ), which not only encodes samples into the weighted codewords in a new flexible manner but also balances the codeword assignment, improving while balancing representation capability of codewords. Specifically, we encode samples using the linear weighted sum of codewords instead of a single codeword as traditionally. Meanwhile, we establish the linear relationship between the weighted codewords and semantic labels, which effectively maintains semantic information of codewords. Moreover, in order to balance the codeword assignment, that is, avoiding some codewords representing most samples or some codewords representing very few samples, we maximize the entropy of the coding probability distribution and obtain the optimal coding probability distribution of samples by utilizing optimal transport theory, which achieves the optimal assignment of codewords and balances representation capability of codewords. The experimental results on three benchmark datasets show that EOWPQ can achieve better retrieval performance and also show the improvement of representation capability of codewords and the balance of codeword assignment.
Lingchen Gu, Wenbo Wan, Jiande Sun 0001
IEEE Trans. Image Process.4
2024 Deep Rank-N Decomposition Network for Image Fusion
abstract
Existing deep neural network (DNN)-based image fusion methods seldom consider low-rank priors for the decomposition of source images, which cannot efficiently model base and detail components in images. To exploit the low-rank priors better, we propose a deep rank-Ndecomposition network (DRDec-Net) according to the rank-Ndecomposition of source images. Specifically, a rank-Ndecomposition model is first established by imposing low-rank priors on the base component of source images. Then, based on the decomposition model, we construct DRDec-Net, which is composed of low-rank decomposition (LRD) modules, a detail fusion (DetailF) module, and a low-rank fusion (LRF) module. In DRDec-Net, it is assumed that source images share the same base component, which is expressed as the sum of rank-1 components. We employNcascaded LRD modules to extract these rank-1 components from source images. Meanwhile, detail components are obtained by subtracting the base component from source images. Next, the extracted rank-1 components and detail components are integrated by LRF and DetailF modules to produce the base component and detail component of the fused image. Finally, the sum of the two obtained components is regarded as the fused image. Compared to some state-of-the-art methods, experimental results demonstrate that the proposed DRDec-Net can produce a better performance on three image fusion tasks, including infrared and visible images, multi-exposure images, and multi-focus images.
Ludan Sun, Kai Zhang 0010, Feng Zhang 0028, Wenbo Wan, Jiande Sun 0001
IEEE Trans. Multim.4
2023 S-Feature Pyramid Network and Attention Model for Drone Detection
abstract
The issue of aviation safety has always received a great of attention and focus, and birds are also an important issue in aviation safety. Nowadays, drones have emerged and share the same airspace with birds at low altitudes. The problems associated with drones should also be taken into account. For example, small drones can be misused for illegal activities and the threat from them is on the rise. Driven by this situation, we used data provided by the ICASSP Drone-vs-Bird detection Grand Challenge for drone detection and used the method of adding shallow feature pyramid network and attention model on SSD [1] (SFA-SSD) to solve the problem of drone detection in competition. Out of 30 test videos, our method was able to detect drones in 11 videos, with 8 videos scoring above 0.1 and only 3 videos scoring above 0.7.
Pengcheng Dong, Chuntao Wang, Zhenyong Lu, Kai Zhang 0010, Wenbo Wan, Jiande Sun 0001
ICASSP5
2023 AFcIHNet: Attention feature-constrained network for single image information hiding
Xingwang Jia, Hua-Mei Xin 0001, Lingchen Gu, Jiande Sun 0001, Wenbo Wan
Eng. Appl. Artif. Intell.6
2023 Multi-dimensional constraints-based PPVO for high fidelity reversible data hiding
Wenxiu Liu, Lili Meng, Jiande Sun 0001, Wenbo Wan
Expert Syst. Appl.5
2023 Towards perceptual image watermarking with robust texture measurement
Yunming Zhang, Yuxin Gong, Jun Wang 0061, Jiande Sun 0001, Wenbo Wan
Expert Syst. Appl.5
2023 A cover selection-based reversible data hiding method by learning cross-modal hashing
Liming Zou, Jiande Sun 0001, Wenbo Wan, Jing Li 0046, Q. M. Jonathan Wu
Multim. Tools Appl.3
2023 CanBiPT: Cancelable biometrics with physical template
Youjun Gao, Jiande Sun 0001, Huaxiang Zhang 0001, Wenbo Wan
Pattern Recognit. Lett.7
2023 Multispectral and hyperspectral image fusion based on low-rank unfolding network
Kai Zhang 0010, Feng Zhang 0028, Chiru Ge, Wenbo Wan, Jiande Sun 0001
Signal Process.5
2023 3D geometrical total variation regularized low-rank matrix factorization for hyperspectral image denoising
Feng Zhang 0028, Kai Zhang 0010, Wenbo Wan, Jiande Sun 0001
Signal Process.3
2023 Spatial-Spectral Dual Back-Projection Network for Pansharpening
abstract
Deep unfolding networks have obtained satisfactory performance in the pansharpening task owing to their sufficient interpretability. Inspired by the back-projection (BP) mechanism, we propose a BP-driven model, spatial-spectral dual back-project network (S2DBPN), to fuse the low spatial resolution multispectral (LR MS) and the high spatial resolution panchromatic (PAN) images by exploiting the BP in spatial and spectral domains. Specifically, the proposed S2DBPN is made up of a spatial BP network, a spectral BP network, and a reconstruction network. In the spatial BP network, spatial down- and up-projection modules are derived from BP, which is responsible for the projection of the LR MS image into the spatial domain. By analogy with the spatial BP, we reformulate the degradation between high spatial resolution multispectral (HR MS) and PAN images as spectral down- and up-projections. Then, the spectral BP network is constructed for the projection of the PAN image along the channel dimension. Finally, the features from spatial and spectral BP networks are integrated to produce the desired HR MS image through the reconstruction network. Compared to the state-of-the-art methods, extensive experiments on QuickBird, GeoEye-1, and WorldView-2 datasets demonstrate that our S2DBPN produces better HR MS images in terms of qualitative and quantitative evaluation metrics. The code of S2DBPN is released at: https://github.com/RSMagneto/S2DBPN.
Kai Zhang 0010, Anfei Wang, Feng Zhang 0028, Wenbo Wan, Jiande Sun 0001, Lorenzo Bruzzone
IEEE Trans. Geosci. Remote. Sens.4
2023 Learning Deep Multiscale Local Dissimilarity Prior for Pansharpening
abstract
Various deep neural networks (DNNs) have been constructed to inject the spatial information of the panchromatic (PAN) image into the low spatial resolution multispectral (LR MS) image. However, most of them ignore the local dissimilarity (LD) prior between MS and PAN images, which has a negative influence on the fused image. Considering the above-mentioned issues, we propose a deep multiscale local dissimilarity network (DMLD-Net) to learn the LD prior at different scales and enhance the spatial and spectral information in the fused image better. Specifically, we first synthesize a downsampled PAN image from the original PAN image to match the scale of the LR MS image. Then, a LD metric is designed to calculate the dissimilarity map between the two images in feature space. According to the learned dissimilarity map, we utilize a LD-guided attention block (LDGAB) to suppress the impact of LD, which filters out the dissimilar information in the features of the PAN image. To learn the LD prior between MS and PAN images sufficiently, the multiscale architecture is considered and we infer the dissimilar maps hierarchically and inject filtered features into the LR MS image progressively. Finally, the fused image is generated by a reconstruction block. Through the LD learning at different scales, reasonable spatial information is extracted from the PAN image, by which the distortions in the fused image caused by LD can be reduced efficiently. Extensive experiments are conducted on GeoEye-1 and WorldView-2 datasets and the results demonstrate the effectiveness of the proposed DMLD-Net in terms of spatial and spectral preservation. The code is available at https://github.com/RSMagneto/DMLD-Net.
Kai Zhang 0010, Guishuo Yang, Feng Zhang 0028, Wenbo Wan, Man Zhou 0003, Jiande Sun 0001, Huaxiang Zhang 0001
IEEE Trans. Geosci. Remote. Sens.4
2023 Robust Coverless Image Steganography Based on Neglected Coverless Image Dataset Construction
abstract
Most of the existing image selection-based coverless image steganography methods mainly focus on improving the capacity and robustness under the assumption that the corresponding dataset is available. But they ignore how to successfully construct the coverless image dataset, which is the foundation of such methods and has a critical impact on the capacity. In this paper, a coverless image steganography is proposed that considers how to efficiently construct the coverless image dataset. In the proposed method, the CNN-based deep hash is extracted from the image and a specific mapping rule is designed to map the high-dimensional deep hash to the low-dimensional secret message. In addition, an unsupervised clustering algorithm is adopted to construct the coverless image dataset, which makes the construction of the coverless image dataset efficient and improves the robustness of the proposed steganography method. To our best knowledge, this is the first attempt to improve the construction efficiency of the coverless image dataset in the field of coverless image steganography. Experimental results show that the construction of a large coverless image dataset is feasible and reliable, and the proposed method has better robustness and higher dataset utilization rate compared with the state-of-the-art methods.
Liming Zou, Jing Li 0046, Wenbo Wan, Q. M. Jonathan Wu, Jiande Sun 0001
IEEE Trans. Multim.3
2022 Robust watermarking based on blur-guided JND model for macrophotography images
abstract
Macrophotography Images (MPIs) have recently emerged as an active topic due to the development of mobile phone camera technology. A large number of MPIs have been rapidly increasing in many rich visual services, such as smartphones or high-definition monitors. MPIs are often composed of sharp macroimage and blur background, which exhibit different perceptual properties that often lead to different just noticeable difference (JND) estimation. Inspired by this, we formulate the blur concealment (BC) as another factor to determine the total masking effect: the interaction is relatively straightforward with a limited masking effect in the sharp regions, and is complicated with a strong masking effect in the blur parts. Furthermore, texture and orientation adaption and color information weighting are separately incorporated into the contrast masking and color masking. Finally, considering both BC and masking effects, a novel robust watermarking framework based on the proposed blur-guided JND model for MPIs, targeting at further improving the MPIs copyright protection performance. Extensive experiments on MP2020 and Blur Detection data sets show that the applicability of the proposed JND model in the scenario of perceptually MPIs watermarking, and our proposed scheme can outperform the state-of-the-art watermarking schemes by providing better robustness performance at the uniform visual quality.
Wenbo Wan, Wenqian Shan, Wenxiu Liu, Zihan Diao, Jiande Sun 0001
Int. J. Intell. Syst.1
2022 A comprehensive survey on robust image watermarking
Wenbo Wan, Jun Wang 0061, Yunming Zhang, Jing Li 0046, Hui Yu 0001, Jiande Sun 0001
Neurocomputing1
2022 HLF-Net: Pansharpening Based on High- and Low-Frequency Fusion Networks
abstract
Many deep neural networks have been constructed for the pansharpening task. However, the differences between the high and low frequencies in images are not considered in some DNN-based pansharpening methods. As high and low frequencies have different information of images, it is difficult for the same network to learn and reconcile the two kinds of frequencies. Considering the aforementioned differences, we propose a new pansharpening network to fuse the high and low frequencies in low spatial resolution multispectral and panchromatic images separately. Specifically, a high and low frequency fusion network is constructed, which is composed of a high-frequency fusion network and a low-frequency fusion network. In the high-frequency fusion network, skip attention is introduced into U-Net to better retain the high frequencies in feature maps. The low-frequency fusion network uses the involution to capture the dependency among the channels of feature maps. Experiments on the GeoEye-1 dataset reveal that the proposed network outperforms some state-of-the-art methods. The code can be accessed at https://github.com/RSMagneto/HLF-Net.
Wenxiu Diao, Feng Zhang 0028, Haitao Wang 0023, Wenbo Wan, Jiande Sun 0001, Kai Zhang 0010
IEEE Geosci. Remote. Sens. Lett.4
2022 Pan-Sharpening Based on Transformer With Redundancy Reduction
abstract
Pan-sharpening methods based on deep neural network (DNN) have produced the state-of-the-art results. However, the common information in the panchromatic (PAN) image and the low spatial resolution multispectral (LRMS) image is not sufficiently explored. As PAN and LRMS images are collected from the same scene, there exists some common information among them, in addition to their respective unique information. The direct concatenation of extracted features leads to some redundancy in the feature space. To reduce the redundancy among features and exploit the global information in source images, we proposed a novel pan-sharpening method by combining the convolution neural network and transformer. Specifically, PAN and LRMS images are encoded as unique features and common features by the subnetworks consisting of convolution blocks and transformer blocks. Then, the common features are averaged and combined with unique features from source images for the reconstruction of the fused image. To extract accurate common features, the equality constraint is imposed on them. Experimental results show that the proposed method outperforms the state-of-the-art methods on both reduced-scale and full-scale datasets. The source code is available athttps://github.com/RSMagneto/TRRNet.
Kai Zhang 0010, Feng Zhang 0028, Wenbo Wan, Jiande Sun 0001
IEEE Geosci. Remote. Sens. Lett.4
2022 Multiple description coding network based on semantic segmentation
Xue Li 0001, Lili Meng, Yanyan Tan, Jia Zhang 0028, Wenbo Wan, Huaxiang Zhang 0001
Multim. Tools Appl.5
2021 JND-aware robust image watermarking with tri-directional inter-block correlation
abstract
A novel block-level perceptual image watermarking framework is proposed in this study, including tri-directional correlation and a block-level just noticeable difference (JND) model. Specifically, the difference in the discrete cosine transform (DCT) coefficients of two blocks is calculated based on three directions in the neighborhood, called the tri-directional correlation (TriDC). Additionally, the representative alternating current (AC) coefficients along horizontal, vertical, and diagonal directions, which can describe structural patterns, are projected and merged for TriDC differences. Then, the difference of the DCT coefficient is modulated to a predefined zone depending on the JND-based offset. Finally, the extent of the watermarked AC coefficients is determined with perceptual JND adjustment. The experimental results demonstrate that the proposed scheme can protect most common image processing attacks; and has better robustness compared with recent zone modulation watermarking schemes and traditional watermarking methods.
Yunming Zhang, Zhenhua Wang 0004, Yantong Zhan, Lili Meng, Jiande Sun 0001, Wenbo Wan
Int. J. Intell. Syst.6
2021 Deep semantic segmentation-based multiple description coding
Xue Li 0001, Lili Meng, Yanyan Tan, Jia Zhang 0028, Wenbo Wan, Huaxiang Zhang 0001
Multim. Tools Appl.5
2021 Visual Security Assessment via Saliency-Weighted Structure and Orientation Similarity for Selective Encrypted Images
abstract
Selective encryption has been widely used in image privacy protection. Visual security assessment is necessary for the effectiveness and practicability of image encryption methods, and there have been a series of research studies on this aspect. However, these methods do not take into account perceptual factors. In this paper, we propose a new visual security assessment (VSA) by saliency-weighted structure and orientation similarity. Considering that the human visual perception is sensitive to the characteristics of selective encrypted images, we extract the structure and orientation feature maps, and then similarity measurements are conducted on these feature maps to generate the structure and orientation similarity maps. Next, we compute the saliency map of the original image. Then, a simple saliency-based pooling strategy is subsequently used to combine these measurements and generate the final visual security score. Extensive experiments are conducted on two public encryption databases, and the results demonstrate the superiority and robustness of our proposed VSA compared with the existing most advanced work.
Zhengguo Wu, Kai Zhang 0010, Yannan Ren, Jing Li 0046, Jiande Sun 0001, Wenbo Wan
Secur. Commun. Networks6
2020 Color image watermarking based on orientation diversity and color complexity
Jun Wang 0061, Wenbo Wan, Xiao Xiao Li, Jiande Sun 0001, Huaxiang Zhang 0001
Expert Syst. Appl.2
2020 Two-class 3D-CNN classifiers combination for video copy detection
Jing Li 0046, Huaxiang Zhang 0001, Wenbo Wan, Jiande Sun 0001
Multim. Tools Appl.3
2020 Hash length: a neglected element
Haifeng Qi, Jing Li 0046, Qiang Wu 0009, Wenbo Wan, Jiande Sun 0001
Multim. Tools Appl.4
2020 Saccadic trajectory-based identity authentication
Huiru Shao, Jing Li 0046, Wenbo Wan, Huaxiang Zhang 0001, Jiande Sun 0001
Multim. Tools Appl.3
2020 Hybrid JND model-guided watermarking method for screen content images
Wenbo Wan, Jun Wang 0061, Jing Li 0046, Jiande Sun 0001, Huaxiang Zhang 0001
Multim. Tools Appl.1
2020 A novel attention-guided JND Model for improving robust image watermarking
Jun Wang 0061, Wenbo Wan
Multim. Tools Appl.2
2020 Pattern complexity-based JND estimation for quantization watermarking
Wenbo Wan, Jun Wang 0061, Jing Li 0046, Lili Meng, Jiande Sun 0001, Huaxiang Zhang 0001
Pattern Recognit. Lett.1
2020 Multi-class joint subspace learning for cross-modal retrieval
En Yu, Jing Li 0046, Li Wang 0148, Jia Zhang 0028, Wenbo Wan, Jiande Sun 0001
Pattern Recognit. Lett.5
2020 Blind Photograph Watermarking with Robust Defocus-Based JND Model
abstract
Just noticeable distortion (JND) is widely employed to describe the perception redundancy in the quantization-based watermarking framework. However, the existing JND models are generally constructed to treat every region of the photograph with an equal focus level, whereas the defocus effect has never been considered. In this paper, the defocus feature, which can portray the aesthetic emphasis in the photograph, is provided to improve the perceptual JND model. Firstly, two indicators which consider the block energy in the defocus measurement (DM) are proposed. Then, the defocus feature map (DFM) is obtained by integrating the influence of the circumambient blocks, and it is applied to the proposed JND contrast masking (CM) processing. In this way, a new blind photograph watermarking method, with emphasis on defocus-JND estimation combined with the proposed CM, is presented. Simulations show that the proposed JND is more suitable for watermarking framework than some exiting JND models, and the proposed watermarking scheme with the improved defocus-based JND model has superior robustness compared with some watermarking schemes.
Chun-Xing Wang, Meiling Xu, Jun Wang 0061, Wenbo Wan
Wirel. Commun. Mob. Comput.5
2019 Weighted locality collaborative representation based on sparse subspace
Huaxiang Zhang 0001, Lei Zhu 0002, Wenbo Wan, Zhenhua Wang 0004, Qiang Wang 0015, Peilian Guo, Jiande Sun 0001
J. Vis. Commun. Image Represent.4
2019 Coupled feature selection based semi-supervised modality-dependent cross-modal retrieval
En Yu, Jiande Sun 0001, Li Wang 0148, Wenbo Wan, Huaxiang Zhang 0001
Multim. Tools Appl.4
2019 A novel coverless information hiding method based on the average pixel value of the sub-images
Liming Zou, Jiande Sun 0001, Min Gao 0001, Wenbo Wan, Brij B. Gupta
Multim. Tools Appl.4
2018 Semi-supervised modality-dependent cross-media retrieval
Jiande Sun 0001, Peiyong Duan, Lili Meng, Yanyan Tan, Wenbo Wan, Hongchen Wu, Bin Zhang 0050, Huaxiang Zhang 0001
Multim. Tools Appl.6
2018 View-invariant gait recognition based on kinect skeleton feature
Jiande Sun 0001, Jing Li 0046, Wenbo Wan, De Cheng, Huaxiang Zhang 0001
Multim. Tools Appl.4
2018 Joint graph regularization based modality-dependent cross-media retrieval
Jihong Yan, Huaxiang Zhang 0001, Jiande Sun 0001, Qiang Wang 0015, Peilian Guo, Lili Meng, Wenbo Wan
Multim. Tools Appl.7
2018 Discriminative correlation hashing for supervised cross-modal retrieval
Xu Lu 0004, Huaxiang Zhang 0001, Jiande Sun 0001, Zhenhua Wang 0004, Peilian Guo, Wenbo Wan
Signal Process. Image Commun.6
2017 A two-stage learning approach to face recognition
Huaxiang Zhang 0001, Jiande Sun 0001, Wenbo Wan
J. Vis. Commun. Image Represent.4
2016 A frame rate up-conversion method with quadruple motion vector post-processing
abstract
We propose a frame rate up-conversion method with quadruple motion vector post-processing scheme (QMVPS). By considering occlusion and hole problems, we adopt the hybrid motion estimation (ME) method which is a combination of bilateral ME and unilateral ME in both directions. To produce better quality-interpolated frames, the motion vector (MV) outliers of the both unilateral MV fields are detected and corrected by prior-information-based MV refining method and the bilateral MV field is smoothed by using vector extrapolation and weighted summation. Moreover, the selection and determination criterion, which is composed of matching ratio of blocks and ladder-style analysis, is proposed to select and decide the denser bilateral MV field. Experimental results verify the superiority of our work in both objective and subjective performances compared with other conventional ME and MV post-processing methods.
Aixi Qu, Wenbo Wan, Yifan Xiao
ICASSP3
2016 Video hashing based on appearance and attention features fusion via DBN
Jiande Sun 0001, Xiaocui Liu, Wenbo Wan, Jing Li 0046, Dong Zhao 0017, Huaxiang Zhang 0001
Neurocomputing3
2016 Improved logarithmic spread transform dither modulation using a robust perceptual model
Wenbo Wan, Jiande Sun 0001
Multim. Tools Appl.1
2015 Improved Spread Transform Dither Modulation Using Luminance-Based JND Model
Wenhua Tang, Wenbo Wan, Jiande Sun 0001
ICIG (2)2
2015 Frame Rate Up-Conversion Using Motion Vector Angular for Occlusion Detection
Guoxia Sun, Wenbo Wan
ICIG (2)5
2013 Logarithmic Spread-Transform Dither Modulation watermarking Based on Perceptual Model
abstract
Logarithmic Quantization Index Modulation (LQIM) is an important extension of the original quantization-based watermarking method. However, it is well known that it is sensitive to valumetric scaling attack and easy to result in sign error after quantization and attacks. For that, in this paper, we propose a new method, namely Logarithmic Spread-Transform Dither Modulation Based on Perceptual Model (LSTDM-WM). In this regard the host signal is first projected onto a random vector and transformed using a novel Logarithmic Quantization function. Then the transformed signal is quantized regarding the watermark data and the watermarked signal is obtained by applying inverse transform to the quantized signal. The perceptual model is further exploited to adjust the quantization step adaptively for watermark embedding. Experimental results indicate that our proposed scheme overcomes two challenges cited above and has superior performance in comparison with conventional LQIM and former proposed schemes of STDM.
Wenbo Wan, Jiande Sun 0001, Xiushan Nie
ICIP1