VLDB 2026 Research / reviewers in the wild / expert
Zhenjun Tang
dblp:09/7638
· DBLP profile ↗
126ranked-venue papers
25as first author
91since 2021 · last 2027
0000-0003-3664-1363ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 70 · 11 first-author · 51 since 2021Artificial intelligence and machine learning · 23 · 3 first-author · 20 since 2021Databases, data management, data science and information retrieval · 14 · 3 first-author · 10 since 2021Computer networks · 12 · 2 first-author · 11 since 2021Security and privacy · 12 · 4 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 3 first-author · 4 since 2021Theory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | HRFusion: A dual-branch infrared and visible image fusion network with hybrid dilated residual coordinate attention and restormer-Res2net blocks
Xiangdong Yin, Xianquan Zhang, Chunqiang Yu, Zhenjun Tang |
Expert Syst. Appl. | 4 |
| 2026 | Self-Supervised Video Hashing with Consistent Short-Term and Long-Range Temporal ModelingabstractWith the explosive growth of video data, learning compact discriminative binary representations for efficient video retrieval has become essential. Most existing self-supervised video hashing methods struggle to achieve consistent modeling across short-term and long-range temporal dependencies, resulting in hash codes that are less discriminative and more redundant. To address this, we propose a novel self-supervised video hashing method called CSL-Hash, based on a bidirectional Mamba encoder-decoder network. The encoder integrates residual convolution, channel attention, and bidirectional Mamba to achieve more comprehensive feature modeling, while the decoder employs bidirectional Mamba to preserve temporal semantics. Additionally, we design a new loss function consisting of reconstruction loss, contrastive loss, and orthogonality constraint loss to enhance the independence and discriminability of hash codes. Experimental results validate the effectiveness of CSL-Hash on multiple video retrieval datasets. Likai Yang, Xiaoping Liang, Zhenjun Tang |
ICMR | 4 |
| 2026 | Video Hashing with Robust Secondary Frames and Local Tangent Space Alignment for Copy DetectionabstractVideo hashing is an effective method to solve copy detection problems. This paper proposes a video hashing with robust secondary frames (RSFs) and local tangent space alignment (LTSA) for copy detection. In the proposed scheme, the input video is uniformly grouped, and RSFs are calculated based on the video frame groups. Then, a tensor is obtained by stacking all RSFs, and the tensor is divided into multiple smaller tensors. The mean calculation is used to construct a matrix from these small tensors, and the low-dimensional features are extracted through LTSA. Moreover, the MobileNetV2 is used to calculate two feature maps from each RSF. The two feature maps are stacked to form a dual-channel feature map, which is then divided into multiple feature map blocks. The mean calculation is used to extract deep features from these blocks. Finally, the low-dimensional and deep features are binarized and concatenated to obtain a hash. Comparative experimental results show that our hashing scheme has the best copy detection and classification performance. Zixuan Yu, Xiaoping Liang, Zhenjun Tang |
ICMR | 3 |
| 2026 | Video Hashing via a Mamba-Transformer Network for Retrieval
Likai Yang, Nianqiao Li, Xiaoping Liang, Lv Chen, Zhenjun Tang |
MMM (2) | 5 |
| 2026 | No-Reference Image Quality Assessment via Attention-Based Feature Enhancement and Feature Interaction
Qiqun Yu, Yihua Chen 0001, Jiliang Ma, Zhenjun Tang |
MMM (1) | 4 |
| 2026 | Bias mitigation label and aware context for debiased visual question answering
Runlin Cao, Zhixin Li 0001, Zhenjun Tang, Huifang Ma |
Expert Syst. Appl. | 3 |
| 2026 | UAPFinger: One-to-many Deep Neural Network Fingerprinting via Universal Adversarial Perturbations
Deyang Wu, Xianquan Zhang, Zhenjun Tang |
Expert Syst. Appl. | 5 |
| 2026 | Multi-scale and global feature fusion network with multiple attentions for no-reference image quality assessment
Qiqun Yu, Yihua Chen 0001, Jiliang Ma, Xiaoping Liang, Zhenjun Tang |
Expert Syst. Appl. | 5 |
| 2026 | Robust video hashing with DWT and tensor SVD for copy detection
Zixuan Yu, Xiaoping Liang, Lv Chen, Xianquan Zhang, Zhenjun Tang |
Expert Syst. Appl. | 5 |
| 2026 | Robust image hashing based on adaptive weighted feature space for copy detection
Hanyun Zhang, Xiaoping Liang, Lv Chen, Xianquan Zhang, Zhenjun Tang |
Expert Syst. Appl. | 5 |
| 2026 | Reversible Data Hiding in Encrypted Images Based on Bit-Plane Classification and Adaptive Group CodingabstractABSTRACT Reversible data hiding in encrypted images (RDH‐EI) is a key technique for secure communication, secure cloud storage and privacy protection. Most existing RDH‐EI algorithms rely on a single and static bit‐plane encoding strategy, failing to fully exploit the structural differences of bit‐planes from the block level to the sequence level, thereby limiting further improvements on embedding capacity. To address this issue, this paper proposes a novel RDH‐EI algorithm based on Bit‐plane Classification and Adaptive Group Coding (hereafter BCAGC algorithm). The core contributions are as follows: (1) A bit‐plane classification mechanism is proposed. It categorises bit‐planes into simple or complex types according to the number of non‐all‐zero 4‐bit sequence (NAZ‐4BS) patterns. (2) An adaptive group coding scheme is proposed. It dynamically selects between fixed‐length coding and Huffman coding based on bit‐plane types and the frequency distribution of NAZ‐4BS patterns, thereby achieving compact coding tables and efficient bit‐plane compression. Experimental results demonstrate that the proposed BCAGC algorithm achieves average embedding rates of 4.1472, 4.0565 and 3.4465 bpp on the BOSSbase, BOWS‐2 and UCID datasets, respectively, outperforming several state‐of‐the‐art RDH‐EI methods. Guoyan Zhou, Nianqiao Li, Chunqiang Yu, Xianquan Zhang, Zhenjun Tang |
IET Image Process. | 5 |
| 2026 | No-Reference Screen Content Image Quality Assessment via Edge and Visual Salient Feature Fusion Network
Xiaoping Liang, Hongting Pan, Yihua Chen 0001, Zhenjun Tang |
IEEE Internet Things J. | 4 |
| 2026 | Semantic-Guided Channel Cross-Attention Integration Network for No-Reference Image Quality AssessmentabstractNo-Reference Image Quality Assessment (NR-IQA) serves as a fundamental task in computer vision that aims to predict image quality consistent with human perception. Currently, numerous NR-IQA methods often use simplistic fusion strategies to integrate features from different backbones. However, these methods typically neglect the semantic differences and intricate inter-channel interactions among different features, thereby limiting their abilities to represent features effectively. To address this issue, we propose a Semantic-guided Channel Cross-attention Integration Network for NR-IQA (SCCIN-IQA), which enables more effective integration of complementary information from different backbones. The core module of our method is the fusion-semantic channel cross-attention. It first generates a semantic feature by integrating features from different backbones, then utilizes this semantic feature as a query to integrate the original backbone features via a channel cross-attention mechanism, thereby adaptively highlighting quality-relevant channel activations. Additionally, a space-channel enhancement module is introduced to further enhance the learned features in both space and channel dimensions, enabling comprehensive modeling of multi-dimensional contextual dependencies. Extensive experiments conducted on multiple public datasets demonstrate that the proposed SCCIN-IQA achieves state-of-the-art performance, consistently surpassing several mainstream methods while exhibiting strong generalization. Jiliang Ma, Yihua Chen 0001, Xiaoping Liang, Xianquan Zhang, Zhenjun Tang |
IEEE Internet Things J. | 5 |
| 2026 | Bimodal-frequency-aware hashing network for screen content images
Ziqing Huang, Lanxiang Guo, Shuo Zhang 0014, Zhenjun Tang |
J. Vis. Commun. Image Represent. | 5 |
| 2026 | GCL-MIH: A Generative-Based Coverless Multi-Image Hiding MethodabstractSecure and high-capacity secret information transmission is an important task of the image hiding research. The existing image hiding methods face some critical issues: cover-based methods offer high capacity but introduce image distortion and security risks, whereas secure coverless methods have low capacity. To address these issues, this paper proposes a novel generative-based coverless multi-image hiding method called GCL-MIH, which can achieve high capacity and high security. The GCL-MIH first utilizes a feature reverse module to compress multiple secret images into multiple feature vectors and then normalizes them to generate a vector that conforms to a standard normal distribution, and finally inputs this vector into an invertible generative network (Flow-GAN) to generate a face image, enabling coverless multiple-image hiding without a predefined cover image. Experimental results demonstrate that the GCL-MIH successfully hides up to four images within a single generated face image, achieving a maximum embedding rate of 32 bpp. This capacity far exceeds those of the existing coverless methods. On the COCO test set, the generated stego images of the GCL-MIH are highly realistic (FID score: 11.98), and the recovered secret images exhibit satisfactory fidelity (the average PSNR and SSIM of four recovered secret images are 33.18 dB and 0.9412). Xianquan Zhang, Chunqiang Yu, Xinpeng Zhang 0001, Ching-Nung Yang, Zhenjun Tang |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2026 | HAR-HFNet: Hybrid-Attention Refinement and Hierarchical Fusion Network for No-Reference Image Quality AssessmentabstractNo-Reference Image Quality Assessment (NR-IQA) is an important task in the field of computer vision. Most methods utilize the pre-trained features with information irrelevant to image quality. In addition, some methods directly regress the pre-trained features without interaction or simply concatenate all features for score regression. They ignore the differences between local and global features. These issues lead to the limited IQA performance. To address this, we propose a Hybrid-Attention Refinement and Hierarchical Fusion Network (HAR-HFNet) for NR-IQA, which consists of a feature extraction module, a hybrid attention refinement module, a hierarchical dilated-attention fusion module and a quality prediction module. Firstly, the hybrid attention refinement module filters and refines the multi-stage features extracted by the pre-trained Swin Transformer, which enhances distortion-related information. Secondly, the hierarchical dilated-attention fusion module fuses the deep global feature with local features. It enables effective hierarchical integration of global semantics and local details. Finally, the quality prediction module predicts the score through weighted feature aggregation. Experiments on six public IQA datasets demonstrate that the HAR-HFNet outperforms some baseline NR-IQA methods in prediction accuracy and generalization ability. Chunyu Wu, Yihua Chen 0001, Kejing Wu, Xiaoping Liang, Zhenjun Tang |
IEEE Signal Process. Lett. | 5 |
| 2026 | Reversible Data Hiding in Shared Images Using Overlapped Coefficients in PolynomialsabstractReversible data hiding (RDH) in shared images is an effective technique for securely storing and managing confidential images. However, most existing methods suffer from a noticeable data expansion and cannot achieve a good trade-off between data expansion and embedding rate. To address this issue, we propose a novel RDH in shared images (RDHSI) using overlapped coefficients in polynomials. In the proposed method, an original image is compressed losslessly to reduce its size before sharing, and then the compressed image is divided into a series of groups, where any two adjacent groups have an overlapped part. Next, each group is shared by our proposed (k,n)-threshold based sharing technique, which is performed by the polynomial over Galois field GF(28) with overlapped coefficients. Finally, data embedding is performed on each shared image by bit replacement according to a constructed 0-1 matrix. Experimental results demonstrate that the proposed method can effectively reduce the sizes of the shared images and achieve a high embedding capacity. Chunqiang Yu, Xianquan Zhang, Ching-Nung Yang, Xinpeng Zhang 0001, Zhenjun Tang |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | Watermarking for Model Ownership Verification:Invisible at Deployment, Activated by UpdatesabstractDeep neural networks for image classification require protection against unauthorized use and redistribution. Existing watermarking methods suffer from a critical vulnerability: watermarks are always active and detectable, allowing adversaries to identify and remove them before deployment. We propose DormMark, a novel framework for image classification models that introduces delayed-activation watermarks which remain dormant and hidden under deployment-time black-box query auditing during initial deployment, but automatically activate upon fine-tuning. Our approach employs a three-stage training paradigm: (1) embedding watermarks using triggered samples, (2) masking to suppress watermark functionality while preserving its latent presence, and (3) activation through standard fine-tuning without owner intervention. This mechanism exploits neural networks’ forgetting-remembering behaviors during continued training, creating a fragile equilibrium that behaves similarly to a clean model under deployment-time black-box auditing but reliably manifests ownership indicators after modification. We consider a private-key black-box verification setting in which the owner keeps the concrete trigger instances secret. Experiments across multiple architectures (VGG19, ResNet-18/56, DenseNet-121, WideResNet-34) and datasets (CIFAR-10, CIFAR-100, GTSRB) demonstrate 100% watermark success rates, high imperceptibility (PSNR > 38 dB, SSIM = 0.99), negligible accuracy loss (< 0.04%), and robustness against 80% parameter pruning. DormMark represents a paradigm shift from static to conditionally-activated ownership verification, providing a more robust framework for intellectual property protection. Hewang Nie, Jue Xiao, Renfei Shen, Chunqiang Yu, Zhenjun Tang |
ACM Trans. Priv. Secur. | 6 |
| 2025 | Structure-Preserving Video Hashing via Self-Supervised Transformer for RetrievalabstractSelf-supervised video hashing aims at generating hash codes and performing fast video content retrieval by leveraging the visual content information inherent in the videos themselves. Most existing methods often overlook the structure-preserving information within the visual content of the videos and thus cannot learn an effective discriminative video representation. In this paper, a Structure-Preserving Video Hashing (SPVH) via a self-supervised Transformer for retrieval is proposed by exploring the global relationships, local relationships, and inter-video relationships in the visual content of videos. In the proposed SPVH, a Transformer-based autoencoder model is used to extract the deep features of the videos. Moreover, a new structure-preserving loss function with the clustering loss, constraint loss, contrastive loss, and reconstruction loss is designed to capture the structural information of the videos. Extensive experiments are conducted on two large-scale video datasets. The results demonstrate the superior performance of our SPVH compared to some state-of-the-art methods. Lixia Du, Xiaoping Liang, Likai Yang, Zhenjun Tang |
ICASSP | 4 |
| 2025 | Attention-Enhanced Feature Fusion Network for No-Reference Image Quality AssessmentabstractNo-Reference Image Quality Assessment (NR-IQA) is a fundamental computer vision task. In this paper, we propose a new Attention-Enhanced Feature Fusion Network for NR-IQA (AEFF-IQA) that integrates multi-scale local and non-local features. Firstly, multi-scale local and non-local features containing distortion information at semantic level are extracted by a feature extraction module consisting of two pre-trained deep neural networks. Secondly, a self-attention enhanced fusion module with four self-attention enhanced fusion components fuses the local and non-local features at the same scale to obtain multi-scale fusion features. Then, a cross-attention enhanced fusion module containing three cross-attention enhanced fusion components integrates these fusion features across multiple scales. Finally, the quality score is obtained by using a prediction module. Extensive experiments are conducted on five public datasets. The results show that our AEFF-IQA outperforms some state-of-the-art models and exhibits good generalization performance. Jiliang Ma, Yihua Chen 0001, Pengsheng Huang, Zhenjun Tang |
ICASSP | 4 |
| 2025 | Artistic Image Aesthetics Assessment Assisted by Photographic Visual AttributesabstractMost data-driven deep learning-based Artistic Image Aesthetics Assessment (AIAA) methods cannot effectively extract visual attributes from art images since the existing artistic image datasets don’t provide any information about visual attributes. The lack of visual attributes reduces the interpretability of AIAA methods and limits their performance. To address these problems, a novel artistic image aesthetics assessment assisted by photographic visual attributes is proposed. The proposed method consists of a feature extraction module and a joint prediction module. The feature extraction module pre-trained on a photographic dataset and an artistic image dataset can learn the information of photographic attributes and the generic artistic aesthetic information. The joint prediction module uses a non-local self-attention block to fuse the photographic visual attribute features with general artistic aesthetic features. The fused features are fed into an FC layer for calculating the artistic image aesthetic score. Experimental results indicate that our proposed method outperforms some state-of-the-art AIAA methods. Haiyong Tang, Yihua Chen 0001, Xiaoping Liang, Lv Chen, Pengsheng Huang, Zhenjun Tang |
ICASSP | 6 |
| 2025 | HGNet: Hash Generation Network Guided by High Frequency Information for Fine-Grained Image RetrievalabstractFine-grained image retrieval (FGIR) is an important topic of image retrieval, and its challenge lies in the accurate identification of image objects with minor inter-class differences and considerable intraclass differences. Most existing methods exploit Convolutional Neural Networks (CNNs) to capture fine-grained and coarse-grained information while overlooking the scale variations. To address these issues, a novel method named Hash Generation Network (HGNet) guided by high frequency information is developed to learn crucial details across different scales. The HGNet consists of a High-Frequency Guidance Module (HFGM) and a Hash Generation Module (HGM). The key contribution is the proposed HFGM which integrates the high-frequency information and multi-scale features extracted from the Swin Transformer. As the Swin Transformer can effectively capture global contextual information, its multi-scale features, guided by high-frequency information that contains fine-grained texture details, can represent both fine-grained and coarse-grained details, thereby guiding the HGM in generating discriminative hash codes. Experimental results show that the HGNet outperforms several SOTA FGIR methods in retrieval performance. Hanyun Zhang, Yihua Chen 0001, Xiaoping Liang, Lv Chen, Zhenjun Tang |
ICASSP | 5 |
| 2025 | A Plug-and-Play and Invisible Multi-Bit Watermarking Scheme for Deep Neural Networks
Jingyu Ye, Zhenjun Tang, Hanzhou Wu |
ICIC (18) | 3 |
| 2025 | Secret image restoration with high-bit correction and symbiotic organisms search
Jianzhong Yang, Xianquan Zhang, Chunqiang Yu, Guoxiang Li, Zhenjun Tang |
Expert Syst. Appl. | 5 |
| 2025 | Reversible Data Hiding via Bit-Plane Block Rearrangement and Intra-Block Compression Coding for Encrypted ImagesabstractABSTRACT Reversible data hiding in encrypted images (RDHEI) enables secret data embedding within encrypted images while allowing for the lossless recovery of the original image after data extraction. This technique holds significant applications in various domains such as cloud storage and data security. However, many existing RDHEI methods suffer from limited embedding capacity. To address this limitation, we present a novel and high capacity RDHEI algorithm via bit‐plane block rearrangement and intra‐block compression coding (hereafter BRBCC algorithm). First, the prediction error (PE) image is generated by using a median edge detection predictor, and the high‐order zero‐valued bit‐planes are compressed. The non‐zero‐valued bit‐planes are then separated into non‐overlapping blocks that can be classified as all‐zero blocks, embeddable blocks, or non‐embeddable blocks. These blocks are then sorted and grouped in terms of block type for block coding. Finally, a new intra‐block compression coding technique with small coded data for locating block elements is proposed to conduct effective compression and thereby reserve more space for embedding secret data. Experimental results indicate that the embedding rates of the BRBCC algorithm reach 3.9381 and 3.8436 bpp on the public datasets of BOSSbase and BOWS‐2, respectively, outperforming some state‐of‐the‐art RDHEI algorithms and exhibiting good application potential. Shuyi Deng, Nianqiao Li, Chunqiang Yu, Xianquan Zhang, Zhenjun Tang |
IET Image Process. | 5 |
| 2025 | Secret image restoration with interpolation and social network search
Jianzhong Yang, Xianquan Zhang, Chunqiang Yu, Xuemao Zhang, Guoxiang Li, Zhenjun Tang |
Neurocomputing | 6 |
| 2025 | Unifying Statistical and Refined Semantic Features for Lightweight No-Reference Image Quality AssessmentabstractNo-Reference Image Quality Assessment (NR-IQA) is an important task of computer vision. Most deep neural networks based NR-IQA methods have the ability of accurate quality predictions, but they have large-scale parameters and high computational complexity. To alleviate these problems, we propose a lightweight NR-IQA method by unifying statistical and refined semantic features. Our proposed method consists of a lightweight feature extractor, a Statistical Semantic Feature Extraction (SSFE) module, and a Refined Semantic Feature Extraction (RSFE) module. The lightweight feature extractor is used to extract semantic features with perceptual distortion information. The SSFE module is designed to obtain statistical information of the semantic features for capturing the local and global changes of distorted image. The RSFE module is designed to refine the semantic features for measuring complex distortions. Extensive experiments on many IQA datasets are done and the results indicate that our proposed method outperforms some baseline NR-IQA methods in IQA performance, generalization ability, and model complexity. Yihua Chen 0001, Lv Chen, Xiaoping Liang, Haiyong Tang, Zhenjun Tang |
IEEE Internet Things J. | 5 |
| 2025 | Enhancing robust VQA via contrastive and self-supervised learningabstractVisual Question Answering (VQA) aims to evaluate the reasoning abilities of an intelligent agent using visual and textual information. However, recent research indicates that many VQA models rely primarily on learning the correlation between questions and answers in the training dataset rather than demonstrating actual reasoning ability. To address this limitation, we propose a novel training approach called Enhancing Robust VQA via Contrastive and Self-supervised Learning (CSL-VQA) to construct a more robust VQA model. Our approach involves generating two types of negative samples to balance the biased data, using self-supervised auxiliary tasks to help the base VQA model overcome language priors, and filtering out biased training samples . In addition, we construct positive samples by removing spurious correlations in biased samples and perform auxiliary training through contrastive learning . Our approach does not require additional annotations and is compatible with different VQA backbones. Experimental results demonstrate that CSL-VQA significantly outperforms current state-of-the-art approaches, achieving an accuracy of 62.30% on the VQA-CP v2 dataset, while maintaining robust performance on the in-distribution VQA v2 dataset. Moreover, our method shows superior generalization capabilities on challenging datasets such as GQA-OOD and VQA-CE, proving its effectiveness in reducing language bias and enhancing the overall robustness of VQA models. Runlin Cao, Zhixin Li 0001, Zhenjun Tang, Canlong Zhang, Huifang Ma |
Pattern Recognit. | 3 |
| 2025 | A quantum reversible color-to-grayscale conversion scheme via image encryption based on true random numbers and two-dimensional quantum walks
Nianqiao Li, Zhenjun Tang |
Signal Process. | 2 |
| 2025 | Reversible data hiding in encrypted images using prediction error modification and basic block compression
Xuemao Zhang, Xianquan Zhang, Chunqiang Yu, Guoxiang Li, Zhenjun Tang |
Signal Process. | 5 |
| 2025 | Perceptual Screen Content Image Hashing Using Adaptive Texture and Shape FeaturesabstractWith the flourishing development of multi-client interactive systems, a new type of digital image known as Screen Content Image (SCI) has emerged. Unlike traditional natural scene images, SCI encompasses various visual contents, including natural images, graphics, and text. Because the multi-region distribution characteristics of screen content images result in the presence of blank regions, malicious modifications are easier to operate and harder to perceive, making a serious threat to visual content security. To this end, this paper proposes a color screen content image hashing algorithm using adaptive text regions features and global shape features. Specifically, the text regions are adaptively collected by calculating the local standard deviation of sub-blocks. Then, quaternion Fourier significant maps are computed for the text regions, and texture statistical features are further extracted to reflect the essential visual content robustness. Moreover, the global shape features are represented from the entire color SCI to ensure the discrimination. Finally, the hash sequence with a length of 142 bits is derived from the above features. Importantly, a specialized tampering dataset for SCIs has been established, and the proposed hashing shows highly sensitive to malicious modifications with a satisfactory detection accuracy. Meanwhile, the ROC curve analysis indicates that the proposed method outperforms existing hashing algorithms. Xue Yang 0019, Ziqing Huang, Shuo Zhang 0014, Zhenjun Tang |
IEEE Signal Process. Lett. | 5 |
| 2025 | MB-FAENet: Multi-Branch Feature and Attention Enhancement Network for No-Reference Image Quality Assessment
Qiqun Yu, Pengsheng Huang, Yihua Chen 0001, Xiaoping Liang, Zhenjun Tang |
IEEE Signal Process. Lett. | 5 |
| 2025 | Reversible Data Hiding in Encrypted Images With Secret Sharing and Multivariate Linear Equation
Chunqiang Yu, Xianquan Zhang, Guoxiang Li, Peng Liu 0044, Xinpeng Zhang 0001, Zhenjun Tang |
IEEE Trans. Dependable Secur. Comput. | 6 |
| 2025 | Robust Image Hashing With Weighted Saliency Map and Laplacian EigenmapsabstractCopy detection is crucial for protecting image copyright. This paper proposes a robust image hashing approach via Weighted Saliency Map (WSM) and Laplacian Eigenmaps (LE) (hereafter WSM-LE approach). An important contribution is the WSM construction via the edge map and the saliency map. As the WSM can indicate the interest regions of image, hash calculation based on WSM can provide robustness of our WSM-LE approach. Another contribution is the low-dimensional feature learning by the LE technique. As the LE technique can effectively learn the internal geometric relationships of image, the extracted low-dimensional features can improve discrimination of our WSM-LE approach. In addition, the low-dimensional features are treated as vectors and the vector distances are used to create a compact and encrypted hash. Numerous experiments and comparisons are conducted to confirm the effectiveness and superiority of our WSM-LE approach. The results indicate that our WSM-LE approach has excellent classification and copy detection performances than some baseline approaches. Xiaoping Liang, Zhenjun Tang, Xianquan Zhang, Xinpeng Zhang 0001, Ching-Nung Yang |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2025 | A novel image hashing with low-rank sparse matrix decomposition and feature distance
Zixuan Yu, Zhenjun Tang, Xiaoping Liang, Hanyun Zhang, Ronghai Sun, Xianquan Zhang |
Vis. Comput. | 2 |
| 2024 | Adaptive Prompt Construction Method for Relation ExtractionabstractPrompt learning was proposed to solve the problem of inconsistency between the upstream and downstream tasks and has achieved State-Of-The-Art (SOTA) results in various Natural Language Processing (NLP) tasks. However, Relation Extraction (RE) is more complex than other text classification tasks, which makes it more difficult to design a suitable prompt template for each dataset manually. To solve this issue, we propose a Adaptive Prompt Construction method (APC) for relation extraction. Our method entails obtaining context-aware prompt tokens by extracting and generating trigger words associated with the entities. Furthermore, to alleviate the issue of instability in the prompt-tuning framework during training, we introduce a novel joint contrastive loss to optimize our model. Our method not only effectively reduces the human effort used for prompt template construction, but also achieves better performance in RE. We conduct the experiment on four public RE datasets, which demonstrate the proposed method outperforms the existing SOTA results in both datasets and experimental settings. Zhenbin Chen, Zhixin Li 0001, Zhenjun Tang |
ICASSP | 4 |
| 2024 | Fragile Model Watermark for integrity protection: leveraging boundary volatility and sensitive sample-pairingabstractNeural networks have increasingly influenced people’s lives. Ensuring the faithful deployment of neural networks as designed by their model owners is crucial, as they may be susceptible to various malicious or unintentional modifications, such as backdooring and poisoning attacks. Fragile model watermarks aim to prevent unexpected tampering that could lead DNN models to make incorrect decisions. They ensure the detection of any tampering with the model as sensitively as possible. However, prior watermarking methods suffered from inefficient sample generation and insufficient sensitivity, limiting their practical applicability. Our approach employs a sample-pairing technique, placing the model boundaries between pairs of samples, while simultaneously maximizing logits. This ensures that the model’s decision results of sensitive samples change as much as possible and the Top-1 labels easily alter regardless of the direction it moves. Experimental evaluations conducted across multiple models and datasets demonstrate the superior sensitivity and generation efficiency of our method compared to the current approaches. Zhenzhe Gao, Zhenjun Tang, Zhao-Xia Yin, Baoyuan Wu, Yue Lu 0001 |
ICME | 2 |
| 2024 | Integrating Subjective and Objective Features for Image Aesthetics AssessmentabstractImage aesthetics assessment is commonly treated as a classification or regression task and its performance bottleneck mainly depends on the effective utilization of aesthetic features. Traditional methods tend to use only subjective or objective features. Only one type of feature overlooks the inherent unity of these features in aesthetic characteristics. To comprehensively utilize image aesthetic characteristics, we propose an innovative image aesthetics assessment method that integrates subjective and objective features. Specifically, the proposed method comprises two modules: feature extraction module and feature fusion module. In the feature extraction module, a dual-branch neural network is proposed to extract subjective and objective aesthetic features of the image. In the feature fusion module, a feature fusion module based on a deep feedforward neural network is proposed to combine these two types of features and generate the final image aesthetic score. The experimental results indicate that our method gets a high performance in aesthetics assessment. Haiyong Tang, Yihua Chen 0001, Hanyun Zhang, Lv Chen, Ronghai Sun, Zhenjun Tang |
IJCNN | 6 |
| 2024 | Unifying Pictorial and Textual Features for Screen Content Image Quality EvaluationabstractDividing a Screen Content Image (SCI) with complex components into pictorial and textual regions for predicting scores is one of the common Screen Content Image Quality Assessment (SCIQA) methods. However, how to efficiently leverage pictorial and textual features to predict quality scores for no-reference SCIQA still needs to be explored. In addition, statistical analysis reveals that labels of SCIs present a distribution. Therefore, both the distribution of quality scores of SCIQA and the distribution of labels need to be considered in the SCIQA. This paper proposes a no-reference SCIQA method unifying pictorial and textual features. One contribution is the proposed dual-branch extraction module with the parameter-free attention convolution block and the joint prediction module. The proposed method employs the dual-branch extraction module to generate efficient pictorial and textual features and then uses the joint prediction module to predict quality scores. Another contribution is the joint distribution loss. It makes the distribution of the quality scores as close as possible to the distribution of labels. Experiments on the SCIQA datasets show that the proposed method achieves excellent SCIQA performance and generalization ability. Yihua Chen 0001, Xiaoping Liang, Mengzhu Yu, Zhenjun Tang |
ICMR | 4 |
| 2024 | Robust Video Hashing with Non-negative Tensor Factorization for Copy DetectionabstractCopy detection is a key task of video copyright protection. This paper presents a robust video hashing with non-negative tensor factorization (NTF) for copy detection. In the presented video hashing scheme, secondary frames are computed from the preprocessed video by assigning weights to all frames within a video group based on color entropy. Next, the secondary frames are fed into the pre-trained MobileNetV2 and then NTF is exploited to compress the three-order tensor constructed by stacking the output feature maps for hash construction. Experiments conducted on publicly available video datasets indicate that the presented hashing scheme outperforms the evaluated hashing schemes in the performances of classification and copy detection. Mengzhu Yu, Zhenjun Tang, Huijiang Zhuang, Xiaoping Liang, Zhixin Li 0001, Xianquan Zhang |
ICMR | 2 |
| 2024 | Video Hashing with Tensor Robust PCA and Histogram of Optical Flow for Copy DetectionabstractAbstract This paper proposes a novel video hashing with tensor robust Principal Component Analysis (PCA) and Histogram of Optical Flow (HOF) for copy detection. In the proposed hashing, a video is divided into some video groups. For each video group, a low-rank secondary frame is constructed from the low-rank component decomposed by applying tensor robust PCA to the video group. Since the low-rank component can well indicate spatial-temporal intrinsic structure of the video group and it is slightly disturbed by digital operations, feature extraction from the low-rank secondary frames is discriminative and stable. Next, spatial features and temporal features are extracted from low-rank secondary frames by Charlier moments and HOF, respectively. Since the Charlier moments are robust to geometric transform and they can efficiently distinguish video frames with different contents, the use of Charlier moments can make robust and discriminative spatial features. As the HOF can measure the distribution of motion information between frames, the temporal features formed by HOFs can provide good discrimination. Hash is ultimately determined by quantizing the spatial and temporal features and concatenating the quantized results. Numerous experiments on open video datasets indicate that the proposed hashing is superior to some hashing baseline schemes in terms of classification and copy detection. Mengzhu Yu, Zhenjun Tang, Hanyun Zhang, Xiaoping Liang, Xianquan Zhang |
Comput. J. | 2 |
| 2024 | Dual-attention pyramid transformer network for No-Reference Image Quality Assessment
Jiliang Ma, Yihua Chen 0001, Lv Chen, Zhenjun Tang |
Expert Syst. Appl. | 4 |
| 2024 | Effective Image Hashing With Deep and Moment Features for Content AuthenticationabstractHashing is an efficient technology for various image tasks. This article proposes an effective image hashing with deep and moment features for content authentication. The deep features are calculated by Wavelet scattering network (ScatNet) and local tangent space alignment (LTSA). The ScatNet is used to construct a third-order tensor from the image brightness component in the polar coordinates transformation (PCT) domain and the LTSA is used to learn the compact features from the third-order tensor. The moment features are contributed by the tchebichef moments (TMs) and quaternion bessel fourier moments (QBFMs), where the TMs can measure shape features and the QBFMs can reflect color features. Extensive experiments on four public databases are done to verify performances of the proposed algorithm. The results demonstrate that the proposed algorithm is superior to some baseline algorithms in content authentication. Zixuan Yu, Lv Chen, Xiaoping Liang, Xianquan Zhang, Zhenjun Tang |
IEEE Internet Things J. | 5 |
| 2024 | Multidirectional Gradient Predictor for Region-Based Reversible Data Hiding in Encrypted ImagesabstractReversible data hiding in encrypted images (RDHEIs) has attracted considerable attention, as it can facilitate the management of massive encrypted images and can be employed for covert communication. Recent research has demonstrated that the RDHEI methods with pixel prediction can achieve a more significant embedding capacity than those that do not utilize pixel prediction. Moreover, the accuracy of predictors greatly impacts the embedding capacity. Nevertheless, current predictors have several limitations, including a lack of accuracy and insufficient flexibility. To address these issues, we propose a high-precision multidirectional gradient predictor (MDGP). Based on this predictor, a novel region-based RDHEI method is proposed. Pixel prediction, image compression, data embedding, data extraction, and image recovery are conducted independently within image regions. Extensive experiments have demonstrated that the proposed MDGP predictor outperforms the current predictors in several metrics, including the average absolute errors, information entropy, and embedding capacity. The proposed RDHEI method demonstrates superior embedding capacity on the test images, and the data sets BOSSBase and BOWS2 outperforming the several state-of-the-art methods. Furthermore, it exhibits robust resilience to a variety of attacks, including perceptual attacks, statistical analysis, and patch removal attacks. Xuemao Zhang, Xianquan Zhang, Chunqiang Yu, Jianzhong Yang, Zhenjun Tang |
IEEE Internet Things J. | 5 |
| 2024 | Lightweight transformer and multi-head prediction network for no-reference image quality assessment
Zhenjun Tang, Yihua Chen 0001, Xiaoping Liang, Xianquan Zhang |
Neural Comput. Appl. | 1 |
| 2024 | Exploring refined dual visual features cross-combination for image captioning
Junbo Hu, Zhixin Li 0001, Zhenjun Tang, Huifang Ma |
Neural Networks | 4 |
| 2024 | Image Hiding Based on Compressive Autoencoders and Normalizing FlowabstractImage hiding aims to hide the secret data in the cover image for secure transmission. Recently, with the development of deep learning, some deep learning-based image hiding methods were proposed. However, most of them do not achieve outstanding hiding performance yet. To address this issue, we propose a new image hiding framework called CAE-NF, which consists of compressive autoencoders (CAE) and normalizing flow (NF). Specifically, CAE's encoder respectively maps the secret image and cover image into the corresponding feature vectors. Image hiding and recovery can be modelled as the forward and backward processes of NF since NF is an invertible neural network. NF maps two feature vectors to a stego-image by its forward process. On the recovery side, the stego-images are mapped to two feature vectors by NF's backward process. Finally, the secret image is recovered by CAE's decoder. The proposed framework can achieve a good trade-off between the stego-image quality and recovered secret image quality, and meanwhile, improve the hiding and recovery performances. The experimental results demonstrate that the proposed framework significantly outperforms some state-of-the-art methods in terms of invisibility, security, and recovery accuracy on various datasets. Xianquan Zhang, Chunqiang Yu, Zhenjun Tang |
IEEE Signal Process. Lett. | 4 |
| 2024 | Robust Hashing With Local Tangent Space Alignment for Image Copy DetectionabstractRobust hashing is a useful technique for the image applications of watermarking, authentication, quality assessment and copy detection. This paper proposes a new robust hashing for image copy detection by using local tangent space alignment (LTSA). A key contribution is the weighted visual map computation based on the difference of Gaussian (DOG) and visual attention model. The weighted visual map can provide the proposed method with good robustness. Another contribution is the feature learning via LTSA from the feature matrix of the weighted visual map in discrete cosine transform domain. As it can maintain the local geometric relationships within image, the learned features can make the proposed method discriminative. Extensive experiments on public databases are conducted to validate the proposed robust hashing method. Compared with some famous robust hashing methods, the proposed robust hashing method demonstrates preferable classification performance in terms of discrimination and robustness. Copy detection performance is tested and the result verifies effectiveness of the proposed robust hashing method. Xiaoping Liang, Zhenjun Tang, Xianquan Zhang, Mengzhu Yu, Xinpeng Zhang 0001 |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2024 | Flexible Tensor Learning for Multi-View Clustering With Markov ChainabstractMulti-view clustering has gained great progress recently, which employs the representations from different views for improving the final performance. In this paper, we focus on the problem of multi-view clustering based on the Markov chain by considering low-rank constraints. Since most existing methods fail to simultaneously characterize the relations among different entries in a tensor from the global perspective and describe local structures of similarity matrices of a tensor, we propose a novel Flexible Tensor Learning for Multi-view Clustering with the Markov chain (FTLMCM) to solve this problem. We also construct transition probability matrices based on the Markov chain to fully utilize the connection between the Markov chain and spectral clustering. Specifically, the low-rank constraints of the tensor, the frontal slices and the lateral slices of the tensor are imposed on the objective function of the proposed method to achieve these goals. Besides, these three constraints can be optimized jointly to achieve mutual refinement. FTLMCM also uses the tensor rotation to better explore the relationships among different views. We formulate FTLMCM as a problem of low-rank tensor recovery and solve it with the augmented Lagrangian multiplier. Experiments on six different benchmark data sets under six metrics demonstrate that the proposed method is able to achieve better clustering performance. Yalan Qin, Zhenjun Tang, Hanzhou Wu, Guorui Feng |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2024 | Robust Tracking via Bidirectional Transduction With Mask InformationabstractIn the tracking literature, foreground and background information have been extensively investigated to discriminate a target from its surrounding background. However, both foreground and background possess their own spatial-temporal correlation relationship that provide significant information to separate the target from its surrounding background, which has been usually ignored by existing work. To address this issue, we propose a bidirectional transductive network based tracker, which incorporates long-range spatial-temporal and bidirectional constraints. Specifically, our tracker consists of two modules, namely the mask generation module (MGM) and the transduction attention module (TAM). MGM aggregates long-range interdependencies of a target along the history frames for generating accurate target masks. TAM retrieves back to the history frames to find patches similar to the current frame, which are then forwarded along with the target masks generated by MGM. In this manner, each position in the current frame can determine its own identity, whether belonging to either the background or the foreground, hence accurately distinguishing the target from its distractors. We conduct systematically experiments and achieve state-of-the-art performance on several benchmarks, obtaining 69.2% AO on GOT-10k and 82.1% on TrackingNet. TianYu Ning, Bineng Zhong 0001, Qihua Liang, Zhenjun Tang, Xianxian Li |
IEEE Trans. Multim. | 4 |
| 2024 | Reversible Data Hiding in Encrypted Images With Asymmetric Coding and Bit-Plane Block CompressionabstractReversible data hiding in encrypted images (RDHEI) is an effective technology of protecting private data. In this paper, a high-capacity RDHEI method with asymmetric coding and bit-plane block compression is proposed. Our major contributions are twofold. (1) We propose an asymmetric coding technique for processing prediction error (PE) blocks before encryption. The proposed asymmetric coding technique does not generate the sign bit-plane and facilitates massive 0s converging on the high bit-planes. This is beneficial to reserve the embedding room. (2) We present a bit-plane block compression technique for improving the embedding capacity. This technique divides the PE codes in a block into two parts which are both compressed and thus contribute a large embedding room. Experimental results demonstrate that the average embedding rates of the proposed method are 4.156, 4.063 and 3.450 bpp on the BOSSBase, BOWS-2 and UCID datasets, respectively. Comparisons show that our average embedding rates on the three datasets are all bigger than those of some state-of-the-art methods. Xianquan Zhang, Feiyi He, Chunqiang Yu, Xinpeng Zhang 0001, Ching-Nung Yang, Zhenjun Tang |
IEEE Trans. Multim. | 6 |
| 2024 | Robust Image Hashing via CP Decomposition and DCT for Copy DetectionabstractCopy detection is a key task of image copyright protection. This article proposes a robust image hashing algorithm by CP decomposition and discrete cosine transform (DCT) for copy detection. The first contribution is the third-order tensor construction with low-frequency coefficients in the DCT domain. Since the low-frequency DCT coefficients contain most of the image energy, they can reflect the basic visual content of the image and are less disturbed by noise. Hence, the third-order tensor construction with the low-frequency DCT coefficients can ensure robustness of our algorithm. Another contribution is the application of the CP decomposition to the third-order tensor for learning a short binary hash. As the factor matrices learned from the CP decomposition can preserve the topology of the original tensor, the binary hash derived from the factor matrices can reach good discrimination. Lots of experiments and comparisons are done to validate effectiveness and advantage of our algorithm. The results demonstrate that our algorithm has superior classification and copy detection performances than several baseline algorithms. In addition, our algorithm is also better than some baseline algorithms with regard to hash length and computational time. Xiaoping Liang, Wanting Liu, Xianquan Zhang, Zhenjun Tang |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2024 | Robust Hashing via Global and Local Invariant Features for Image Copy DetectionabstractRobust hashing is a powerful technique for processing large-scale images. Currently, many reported image hashing schemes do not perform well in balancing the performances of discrimination and robustness, and thus they cannot efficiently detect image copies, especially the image copies with multiple distortions. To address this, we exploit global and local invariant features to develop a novel robust hashing for image copy detection. A critical contribution is the global feature calculation by gray level co-occurrence moment learned from the saliency map determined by the phase spectrum of quaternion Fourier transform, which can significantly enhance discrimination without reducing robustness. Another essential contribution is the local invariant feature computation via Kernel Principal Component Analysis (KPCA) and vector distances. As KPCA can maintain the geometric relationships within image, the local invariant features learned with KPCA and vector distances can guarantee discrimination and compactness. Moreover, the global and local invariant features are encrypted to ensure security. Finally, the hash is produced via the ordinal measures of the encrypted features for making a short length of hash. Numerous experiments are conducted to show efficiency of our scheme. Compared with some well-known hashing schemes, our scheme demonstrates a preferable classification performance of discrimination and robustness. The experiments of detecting image copies with multiple distortions are tested and the results illustrate the effectiveness of our scheme. Xiaoping Liang, Zhenjun Tang, Zhixin Li 0001, Mengzhu Yu, Hanyun Zhang, Xianquan Zhang |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2024 | Reversible Data Hiding in Shared JPEG ImagesabstractReversible data hiding (RDH) in encrypted images has emerged as an effective technique for securely storing and managing confidential images in the cloud. However, most RDH methods in shared images (RDHSI) are designed for uncompressed images and cannot be applied for JPEG images. To address this issue, we propose a novel RDH in shared JPEG images. Our method consists of JPEG image sharing and data hiding in JPEG shares, which are both conducted on JPEG bit-stream. Specifically, the DC appended bits (DCA) and AC appended bits (ACA) derived from the original JPEG bit-stream are shared by ( \(k\) , \(n\) ) threshold Chinese remainder theorem-based secret sharing (CRTSS) with two different constraints, one for DC sharing and another for AC sharing. The constraint of DC sharing ensures that the DC coefficient shares do not overflow. The constraint of AC sharing ensures the sizes of ACA shares are less than the sizes of the original ACA so that the embedding room can be vacated from each shared JPEG bit-stream. Each data-hider can embed the secret data into the personal JPEG share. The original JPEG image can be recovered losslessly from any \(k\) JPEG shares. The proposed sharing and data hiding are both well compatible with the JPEG standard. Experimental results demonstrate that the proposed method not only well preserves the file size whether the JPEG shares or marked JPEG shares but also achieves outstanding security performance and a high embedding capacity. Chunqiang Yu, Shichao Cheng, Xianquan Zhang, Xinpeng Zhang 0001, Zhenjun Tang |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2024 | Robust Hashing with Deep Features and Meixner Moments for Image Copy DetectionabstractCopy detection is a key task of image copyright protection. Most robust hashing schemes do not make satisfied performance of image copy detection yet. To address this, a robust hashing scheme with deep features and Meixner moments is proposed for image copy detection. In the proposed hashing, global deep features are extracted by applying tensor Singular Value Decomposition (t-SVD) to the three-order tensor constructed in the DWT domain of the feature maps calculated by the pre-trained VGG16. Since the feature maps in the DWT domain are slightly disturbed by digital operations, the constructed three-order tensor is stable and thus the desirable robustness is guaranteed. Moreover, since t-SVD can decompose a three-order tensor into multiple low-dimensional matrices reflecting intrinsic structure, the global deep feature calculation from the low-dimensional matrices can provide good discrimination. Local features are calculated by the block-based Meixner moments. As the Meixner moments are resistant to geometric transformation and can efficiently discriminate various images, the use of the block-based Meixner moments can make discriminative and robust local features. Hash is ultimately determined by quantifying and combining global deep features and local features. The results of extensive experiments on open image datasets demonstrate that the proposed robust hashing outperforms some state-of-the-art robust hashing schemes in terms of classification and copy detection performances. Mengzhu Yu, Zhenjun Tang, Xiaoping Liang, Xianquan Zhang, Zhixin Li 0001, Xinpeng Zhang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2023 | Dual-Feature Aggregation Network for No-Reference Image Quality Assessment
Yihua Chen 0001, Mengzhu Yu, Zhenjun Tang |
MMM (1) | 4 |
| 2023 | Lightweight Image Hashing Based on Knowledge Distillation and Optimal Transport for Face Retrieval
Ping Feng, Hanyun Zhang, Zhenjun Tang |
MMM (2) | 4 |
| 2023 | Robust Image Hashing With Saliency Map And Sparse ModelabstractAbstract Image hashing is an effective technology for extensive image applications, such as retrieval, authentication and copy detection. This paper designs a new image hashing scheme based on saliency map and sparse model. The major contributions are twofold. The first contribution is the construction of a weighted image representation by combining a visual attention model called Itti model and the matrix of color vector angle (CVA). Since the Itti model can efficiently detect saliency map and CVA fully captures color information of image, they contribute to a visually robust and discriminative image representation. The second contribution is the hash extraction from the weighted image representation via sparse model. A classical sparse model called robust principal component analysis is exploited to decompose the weighted image representation into a low-rank component and a sparse component. As the low-rank component can describe intrinsic structure of image, hash calculation with low-rank component can achieve good discrimination. The efficiencies of the proposed scheme are validated by extensive experiments with open databases. The results demonstrate that the proposed scheme is superior to some state-of-the-art schemes in terms of classification performance between robustness and discrimination. Mengzhu Yu, Zhenjun Tang, Zhixin Li 0001, Xiaoping Liang, Xianquan Zhang |
Comput. J. | 2 |
| 2023 | Heterogeneous graph convolution based on In-domain Self-supervision for Multimodal Sentiment Analysis
Yufei Zeng, Zhixin Li 0001, Zhenjun Tang, Zhenbin Chen, Huifang Ma |
Expert Syst. Appl. | 3 |
| 2023 | Unifying knowledge iterative dissemination and relational reconstruction network for image-text matching
Xiumin Xie, Zhixin Li 0001, Zhenjun Tang, Huifang Ma |
Inf. Process. Manag. | 3 |
| 2023 | A survey of visual neural networks: current trends, challenges and opportunities
Ping Feng, Zhenjun Tang |
Multim. Syst. | 2 |
| 2023 | SiamBAN: Target-Aware Tracking With Siamese Box Adaptive NetworkabstractVariation of scales or aspect ratios has been one of the main challenges for tracking. To overcome this challenge, most existing methods adopt either multi-scale search or anchor-based schemes, which use a predefined search space in a handcrafted way and therefore limit their performance in complicated scenes. To address this problem, recent anchor-free based trackers have been proposed without using prior scale or anchor information. However, an inconsistency problem between classification and regression degrades the tracking performance. To address the above issues, we propose a simple yet effective tracker (named Siamese Box Adaptive Network, SiamBAN) to learn a target-aware scale handling schema in a data-driven manner. Our basic idea is to predict the target boxes in a per-pixel fashion through a fully convolutional network, which is anchor-free. Specifically, SiamBAN divides the tracking problem into classification and regression tasks, which directly predict objectiveness and regress bounding boxes, respectively. A no-prior box design is proposed to avoid tuning hyper-parameters related to candidate boxes, which makes SiamBAN more flexible. SiamBAN further uses a target-aware branch to address the inconsistency problem. Experiments on benchmarks including VOT2018, VOT2019, OTB100, UAV123, LaSOT and TrackingNet show that SiamBAN achieves promising performance and runs at 35 FPS. Zedu Chen, Bineng Zhong 0001, Guorong Li, Shengping Zhang, Rongrong Ji, Zhenjun Tang, Xianxian Li |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2023 | Secret Image Restoration With Convex Hull and Elite Opposition-Based Learning StrategyabstractDigital images are easily corrupted during transmission. Most image denoising methods cannot perform well on restoring the secret image extracted from a corrupted stego image. To deal with this issue, we propose a new secret image restoration method with convex hull and elite opposition-based learning strategy. Specifically, the pixel distortion values of the corrupted secret image are calculated and used to classify the pixels into trustable pixels or untrusted pixels. For an untrusted pixel, a convex hull is generated by its context trustable pixels due to the irregular distribution of trustable pixels. The untrusted pixel in the convex hull is restored by the trustable pixels within the convex hull. The other untrusted pixels are restored using elite opposition-based learning strategy. The experimental results show that the proposed method outperforms some state-of-the-art methods regarding recovered secret image quality. Xianquan Zhang, Chunqiang Yu, Zhenjun Tang |
IEEE Signal Process. Lett. | 4 |
| 2023 | Robust Tracking via Uncertainty-Aware Semantic ConsistencyabstractRobust tracking has a variety of practical applications. Despite many years of progress, it is still a difficult problem due to enormous uncertainties in real-world scenes. To address this issue, we propose a robust anchor-free based tracking model with uncertainty estimation. Within the model, a new data-driven uncertainty estimation strategy is proposed to generate uncertainty-aware features with promising discriminative and descriptive power. Then, a simple yet effective pyramid-wise cross correlation operation is constructed to extract multi-scale semantic features that provide rich correlation information for uncertainty-aware estimation and thus enhances the tracking robustness. Finally, a semantic consistency checking branch is designed to further estimate uncertainty of output results from the classification and regression branches by adaptively generating semantically consistent labels. Experiments on six benchmarks (i.e., OTB100, VOT2018, VOT2020, TrackingNet, GOT-10K and LaSOT) show the competing performance of our tracker with 130 FPS. Jie Ma 0006, Xiangyuan Lan, Bineng Zhong 0001, Guorong Li, Zhenjun Tang, Xianxian Li, Rongrong Ji |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | Reversible Data Hiding in Encrypted Images With Secret Sharing and Hybrid CodingabstractReversible data hiding in encrypted images (RDHEI) is an essential data security technique. Most RDHEI methods cannot perform well in embedding capacity and security. To address this issue, we propose a new RDHEI method using Chinese remainder theorem-based secret sharing (CRTSS) and hybrid coding. Specifically, a hybrid coding is first proposed for RDH to achieve high embedding capacity. At the content owner side, a novel iterative encryption is designed to conduct block based encryption for perfectly preserving the spatial correlation of original blocks in their encrypted blocks. Then, the CRTSS with the constraints is exploited to generate multiple encrypted image shares, in which spatial correlations of the encrypted blocks are also preserved. Meanwhile, the CRTSS provides good security properties for the proposed method. Since there are strong spatial correlations in the blocks of each share, the data-hider can exploit the proposed hybrid coding to perform data embedding for improving capacity. On the receiver side, even if some shares are corrupted/missing, the original image can be losslessly recovered as long as enough uncorrupted marked shares are obtained. Experiment results show that the proposed RDHEI method outperforms some state-of-the-art methods, including some secret sharing (SS) based methods in terms of embedding capacity. Chunqiang Yu, Xianquan Zhang, Chuan Qin 0001, Zhenjun Tang |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | Leveraging Local and Global Cues for Visual Tracking via Parallel Interaction NetworkabstractDespite that both local and context information are crucial for robust tracking, existing CNN-based and transformer-based methods mainly focus on one of these aspects. Consequently, the former fails to exploit rich global context information due to the limited receptive field, while the latter suffers from the deficiencies in constructing the local relationship among neighboring regions. To address this issue, we propose the SiamPIN tracker, based on our Parallel Interaction Network. It consists of two effective modules, namely Global Aggregation Block (GAB) and Local Process Block (LPB). GAB perceives the global context to capture the long-range spatial dependency through a transformer-based architecture. Meanwhile, LPB performs local information extraction using a CNN model to retain the detailed appearance information of the target. These two modules are connected consecutively to compose a Trans-Conv unit block, which transmits the global context information to the local feature extraction procedure, hence enables the interaction of global-local information flow. Several such blocks are cascaded so that our model can learn to aggregate local and context information interactively. The proposed tracker achieves state-of-the-art performance on six benchmark datasets, while maintaining a real time running speed. Yaozong Zheng, Bineng Zhong 0001, Qihua Liang, Zhenjun Tang, Rongrong Ji, Xianxian Li |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | Perceptual Image Hashing With Locality Preserving Projection for Copy DetectionabstractPerceptual image hashing is an effective and efficient way to identify images in large-scale databases, where two major performances are robustness and discrimination. A better tradeoff between robustness and discrimination is still a severe challenge for the current hashing research. Aiming at this issue, we design a novel perceptual image Hashing with Locality Preserving Projection (LPP) (hereafter HLPP). Specifically, to improve the robustness against content-preserving operations, Gabor filtering is leveraged to adaptively extract the orientation and structure features, which are consistent with the response of human visual system. The LPP is adopted to learn intrinsic local structure from the maximum Gabor filtering response. The use of LPP can discover meaningful low-dimensional information hidden in the maximum Gabor filtering response and thus improves discrimination of HLPP. During hash similarity calculation, the Hamming distance is selected as the metric. The tradeoff performance between robustness and discrimination is validated on benchmark databases, and the results indicate that the proposed HLPP is superior to some state-of-the-art algorithms. In addition, extensive experiments of copy detection also demonstrate that the proposed HLPP can provide higher accuracy than the compared algorithms. Ziqing Huang, Zhenjun Tang, Xianquan Zhang, Linlin Ruan, Xinpeng Zhang 0001 |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2023 | Efficient Hashing Method Using 2D-2D PCA for Image Copy DetectionabstractImage copy detection is an important technology of copyright protection. This paper proposes an efficient hashing method for image copy detection using 2D-2D (two-directional two-dimensional) PCA (Principal Component Analysis). The key is the discovery of the translation invariance of 2D-2D PCA. With the property of translation invariance, a novel model of extracting rotation-invariant low-dimensional features is designed by combining PCT (Polar Coordinate Transformation) and 2D-2D PCA. The PCT can convert an input rotated image to a translation matrix. Since the 2D-2D PCA is invariant to translation, the low-dimensional features learned from the translation matrix are rotation-invariant. Moreover, vector distances of low-dimensional features are stable to common digital operations and thus hash construction with the vector distances is of robustness and compactness. Three open image datasets are exploited to conduct various experiments for validating efficiencies of the proposed method. The results demonstrate that the proposed method is much better than some representative hashing methods in the performances of classification and copy detection. Xiaoping Liang, Zhenjun Tang, Ziqing Huang, Xianquan Zhang, Shichao Zhang 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2023 | Robust Image Hashing With Isomap and Saliency Map for Copy DetectionabstractCompression technology for representing image is on demand for efficiently processing images in the Big Data era. Image hashing is an effective compression technology for computing a short representation based on visual content of input image. Currently, most reported image hashing algorithms have weakness in making a desirable classification between discrimination and robustness and thus can not reach good performance in copy detection. To address these issues, this paper proposes a new robust image hashing with Isometric Mapping (Isomap) and saliency map for copy detection. A key contribution is hash generation with saliency map determined by the Frequency Tuned (FT) method, which can guarantee robustness of the proposed image hashing. Another contribution is the use of Isomap in deriving hash from the FT-based saliency map. Since Isomap can discover the internal geometry features of image, the use of Isomap can learn discriminative image features and thus discrimination of the proposed image hashing is ensured. Experiments on open image databases are carried out. Comparison results illustrate that the proposed image hashing is better than some state-of-the-art algorithms in the performances of classification and copy detection. Xiaoping Liang, Zhenjun Tang, Jingli Wu, Zhixin Li 0001, Xinpeng Zhang 0001 |
IEEE Trans. Multim. | 2 |
| 2023 | Robust Long-Term Tracking via Localizing OccludersabstractOcclusion is known as one of the most challenging factors in long-term tracking because of its unpredictable shape. Existing works devoted into the design of loss functions, training strategies or model architectures, which are considered to have not directly touched the key point. Alternatively, we came up with a direct and natural idea that is discarding things that covers the target. We propose a novel occluder-aware representation learning framework to develop this idea. First, we design a local occluders detection module (LODM) to localize the occluders, which works on the principle that discriminates the non-noumenal part from a target based on the general knowledge of this category. An extra dataset and a clustering strategy is proposed to support this general knowledge. Second, we devise a feature reconstruction module to guide the occluder-aware representation learning. With the help of above methods, our localizing occluders tracker, called LOTracker, can learn an occluder-free representation and promote the performance that tracks with occlusion scenarios. Extensive experimental results show that our LOTracker achieves a state-of-the-art performance in multiple benchmarks such as LaSOT, VOTLT2018, VOTLT2019, and OxUvALT. Binfei Chu, Bineng Zhong 0001, Zhenjun Tang, Xianxian Li, Jing Wang 0049 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2023 | Unifying Dual-Attention and Siamese Transformer Network for Full-Reference Image Quality AssessmentabstractImage Quality Assessment (IQA) is a critical task of computer vision. Most Full-Reference (FR) IQA methods have limitation in the accurate prediction of perceptual qualities of the traditional distorted images and the Generative Adversarial Networks (GANs) based distorted images. To address this issue, we propose a novel method by Unifying Dual-Attention and Siamese Transformer Network (UniDASTN) for FR-IQA. An important contribution is the spatial attention module composed of a Siamese Transformer Network and a feature fusion block. It can focus on significant regions and effectively maps the perceptual differences between the reference and distorted images to a latent distance for distortion evaluation. Another contribution is the dual-attention strategy that exploits channel attention and spatial attention to aggregate features for enhancing distortion sensitivity. In addition, a novel loss function is designed by jointly exploiting Mean Square Error (MSE), bidirectional Kullback–Leibler divergence, and rank order of quality scores. The designed loss function can offer stable training and thus enables the proposed UniDASTN to effectively learn visual perceptual image quality. Extensive experiments on standard IQA databases are conducted to validate the effectiveness of the proposed UniDASTN. The IQA results demonstrate that the proposed UniDASTN outperforms some state-of-the-art FR-IQA methods on the LIVE, CSIQ, TID2013, and PIPAL databases. Zhenjun Tang, Zhixin Li 0001, Bineng Zhong 0001, Xianquan Zhang, Xinpeng Zhang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2022 | Multi-Level Feature Aggregation Network for Full-Reference Image Quality AssessmentabstractImage quality assessment (IQA) is an important task of computer vision. Most full-reference (FR) IQA methods do not reach desirable prediction performance. To address this issue, we propose a novel multi-level feature aggregation network (MLFAN) for FR-IQA. An important contribution is an effective multi-level feature aggregation network. This network utilizes a siamese network with vision transformer for multi-level feature extraction. It compares images at the multi-level perceptual feature differences by considering the relationship among color, texture and shape information, focuses more on the salient regions by an attention aggregator and scores images by a two-branch prediction head. Another important contribution is a novel loss function. This loss function jointly utilizes Mean Square Error, KL divergence and rank order of quality scores to provide stable training. It makes the proposed MLFAN-IQA method effectively learn perceptual quality of images. Experiments are done to test IQA performance of the proposed MLFAN-IQA method. Comparisons show that the proposed MLFAN-IQA method outperforms some state-of-the-art FR-IQA methods on the datatsets of conventional distorted images. Moreover, the proposed MLFAN-IQA method also reaches comparable performance on the dataset of GAN-based synthetic distorted images. Yihua Chen 0001, Xiaoping Liang, Zhenjun Tang |
ICTAI | 4 |
| 2022 | Reversible Data Hiding via Arranging Blocks of Bit-Planes in Encrypted Images
Guoxiong Xie, Guijin Fan, Chunqiang Yu, Zhenjun Tang |
IWDW | 5 |
| 2022 | Multi-source Data Hiding in Neural NetworksabstractThis paper proposes a multi-source data hiding scheme for neural networks, in which multiple senders can simultaneously transmit different secret data to a receiver using the same neural network. In our scheme, multiple senders execute data embedding in the overlapping position of a network, so that the existence of other senders can be concealed. Each sender uses a unique embedding key to scramble the parameters for embedding, preventing an attacker or other senders from pretending him. In addition, data embedding is achieved during the training process of the neural network instead of modifying the neural network after training. As a result, the operation of data embedding has a tiny impact on the original neural network. On the receiver side, the corresponding embedding key is used to extract the secret data, while additional decoding networks are unnecessary. Experiments verified the effectiveness and security of our scheme, including embedding capacity and undetectability. Ziyun Yang, Zichi Wang, Xinpeng Zhang 0001, Zhenjun Tang |
MMSP | 4 |
| 2022 | Efficient video hashing based on low-rank framesabstractAbstract Video hashing is a useful technology for diverse video applications, such as digital watermarking, copy detection and content authentication. This paper proposes a novel efficient video hashing based on low‐rank frames. A key contribution is the low‐rank frame calculation using the low‐rank approximation of singular value decomposition (SVD). As the large singular values of SVD are stable to digital operations, video hash extraction using low‐rank frames can provide good robustness. Since most energy is contained within the large singular values, low‐rank frames also contribute to discrimination. Moreover, two‐dimensional discrete wavelet transform (DWT) is applied to every low‐rank frame and the mean of low‐frequency DWT coefficients is selected as a hash element. Since these coefficients can represent input data approximately, hash discrimination is thus ensured. Experiments with 16,850 videos are carried out to test performances of the proposed algorithm. The results show that the proposed algorithm outperforms some well‐known video hashing algorithms in computational time and classification about discrimination and robustness. Zhenhai Chen, Zhenjun Tang, Xinpeng Zhang 0004, Ronghai Sun, Xianquan Zhang |
IET Image Process. | 2 |
| 2022 | A novel hashing scheme via image feature map and 2D PCAabstractAbstract Hashing scheme is a high‐efficiency technique for processing massive images. Two critical metrics of the hashing scheme are discrimination and robustness, but most schemes do not get satisfied classification performance between them. This paper proposes a novel hashing scheme via image feature map and 2D PCA. First, the proposed scheme extracts local phase quantization (LPQ) features in the frequency domain and local ternary pattern (LTP) features in the spatial domain, and combines them to construct an image feature map. Second, the proposed scheme conducts dimension reduction via 2D PCA for learning features from the image feature map. Last, the learned features are compressed to generate the hash sequence. Performances are tested on open image datasets. The results demonstrate that the proposed scheme can make a good balance between discrimination and robustness. In addition, the classification and copy detection of the proposed scheme are both superior to those of some famous hashing schemes. Xiaoping Liang, Zhenjun Tang, Sheng Li 0006, Chunqiang Yu, Xianquan Zhang |
IET Image Process. | 2 |
| 2022 | Reversible data hiding with adaptive difference recovery for encrypted images
Chunqiang Yu, Xianquan Zhang, Guoxiang Li, Shanhua Zhan, Zhenjun Tang |
Inf. Sci. | 5 |
| 2022 | Multiple instance relation graph reasoning for cross-modal hash retrieval
Chuanwen Hou, Zhixin Li 0001, Zhenjun Tang, Xiumin Xie, Huifang Ma |
Knowl. Based Syst. | 3 |
| 2022 | Reversible data hiding with pairwise PEE and 2D-PEH decomposition
Chunqiang Yu, Xianquan Zhang, Dewang Wang, Zhenjun Tang |
Signal Process. | 4 |
| 2022 | Adaptive Path Selection for Dynamic Image CaptioningabstractImage captioning is a challenging task, i.e., given an image machine automatically generates natural language that matches its semantic content and has attracted much attention in recent years. However, most existing models are designed manually, and their performance depends heavily on the expert experience of the designer. In addition, the computational flow of the model is predefined, and hard and easy samples will share the same coding path and easily interfere with each other, thus confusing the learning of the model. In this paper, we propose a Dynamic Transformer to change the encoding procedure from sequential to adaptive, i.e., data-dependent computing paths. Specifically, we design three different types of visual feature extraction blocks and deploy them in parallel at each layer to construct a multi-layer routing space in a fully connected manner. Each block contains a calculation unit that performs the corresponding operations and a routing gate that learns to adaptively select the direction to pass the signal based on the input image. Thus, our model can achieve a robust visual representation by exploring potential visual feature extraction paths. We evaluate our method quantitatively and qualitatively using a benchmark MSCOCO image caption dataset and perform extensive ablation studies to investigate the reasons behind its effectiveness. The experimental results show that our method is significantly superior to previous state-of-the-art methods. Tiantao Xian, Zhixin Li 0001, Zhenjun Tang, Huifang Ma |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Perceptual Hashing With Complementary Color Wavelet Transform and Compressed Sensing for Reduced-Reference Image Quality AssessmentabstractImage quality assessment (IQA) is an important task of image processing and has diverse applications, such as image super-resolution reconstruction, image transmission and monitoring systems. This paper proposes a perceptual hashing algorithm with complementary color wavelet transform (CCWT) and compressed sensing (CS) for reduced-reference (RR) IQA. The CCWT is exploited to decompose input color image into different sub-bands. Since the calculation of CCWT uses all color channels without discarding any information, the distortions introduced by digital operations on color channels are preserved in the CCWT sub-bands. The block-based CS is used to extract features from the CCWT sub-bands. As the Euclidean distance between the block-based CS features is slightly influenced by content-preserving operations, perceptual features constructed by Euclidean distances are robust, discriminative and compact. Hash sequence is finally determined by quantifying the perceptual features. Effectiveness of the proposed hashing is verified by various experiments on four open image databases. Experimental results demonstrate that the proposed hashing is superior to some state-of-the-art algorithms in terms of classification and RR IQA application. Mengzhu Yu, Zhenjun Tang, Xianquan Zhang, Bineng Zhong 0001, Xinpeng Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Reversible Data Hiding With Hierarchical Embedding for Encrypted ImagesabstractReversible data hiding in encrypted images (RDHEI) is an effective technique of data security. Most state-of-the-art RDHEI methods do not achieve desirable payload yet. To address this problem, we propose a new RDHEI method with hierarchical embedding. Our contributions are twofold. (1) A novel technique of hierarchical label map generation is proposed for the bit-planes of plaintext image. The hierarchical label map is calculated by using prediction technique, and it is compressed and embedded into the encrypted image. (2) Hierarchical embedding is designed to achieve a high embedding payload. This embedding technique hierarchically divides prediction errors into three kinds: small-magnitude, medium-magnitude, and large-magnitude, which are marked by different labels. Different from the conventional techniques, pixels with small-magnitude/large-magnitude prediction errors are both used to accommodate secret bits in the hierarchical embedding technique, and therefore contribute a high embedding payload. Experiments on two standard datasets are discussed to validate the proposed RDHEI method. The results demonstrate that the proposed RDHEI method outperforms some state-of-the-art RDHEI methods in payload. The average payloads of the proposed RDHEI method are 3.4568 bpp and 3.6823 bpp for BOWS-2 dataset and BOSSbase dataset, respectively. Chunqiang Yu, Xianquan Zhang, Xinpeng Zhang 0001, Guoxiang Li, Zhenjun Tang |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | Noise Removal in Embedded Image With Bit ApproximationabstractStego-images are often contaminated by interchannel noise or active noise attack when communicating on the Web. And it is challenging to restore embedded image from corrupted stego-image. This paper studies akNN-bit approximation algorithm to remove noises in embedded image. The proposed algorithm distinguishes reliable bits from extracted bits, and estimates pixel values by keeping reliable bits unchanged and correcting unreliable bits. Specifically, the 8th (highest) unreliable bit of a pixel can be approximated with its nearest neighbor pixels. And then, if an unreliable bit locates at any one of the$5^{th}\sim 7^{th}$bits of a pixel, it is adjusted with two nearest neighbors of the pixel, where the pixel is in-between these two nearest neighbors. Finally, for other unreliable bits, each one is approximated by the maximum and minimum possible values of nearest neighbors of its pixel. We conduct experiments for illustrating the efficiency, and demonstrate that the proposed algorithm can recover the embedded images with good visual quality from corrupted stego-images. Xianquan Zhang, Xuelong Li 0001, Zhenjun Tang, Shichao Zhang 0001, Shaomin Xie |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2021 | Learning To Filter: Siamese Relation Network for Robust TrackingabstractDespite the great success of Siamese-based trackers, their performance under complicated scenarios is still not satisfying, especially when there are distractors. To this end, we propose a novel Siamese relation network, which introduces two efficient modules, i.e. Relation Detector (RD) and Refinement Module (RM). RD performs in a meta-learning way to obtain a learning ability to filter the distractors from the background while RM aims to effectively integrate the proposed RD into the Siamese framework to generate accurate tracking result. Moreover, to further improve the discriminability and robustness of the tracker, we introduce a contrastive training strategy that attempts not only to learn matching the same target but also to learn how to distinguish the different objects. Therefore, our tracker can achieve accurate tracking results when facing background clutters, fast motion, and occlusion. Experimental results on five popular benchmarks, including VOT2018, VOT2019, OTB100, LaSOT, and UAV123, show that the proposed method is effective and can achieve state-of-the-art results. The code will be available at https://github.com/hqucv/siamrn Siyuan Cheng 0003, Bineng Zhong 0001, Guorong Li, Xin Liu 0011, Zhenjun Tang, Xianxian Li, Jing Wang 0049 |
CVPR | 5 |
| 2021 | Distractor-Aware Fast Tracking via Dynamic Convolutions and MOT PhilosophyabstractA practical long-term tracker typically contains three key properties, i.e. an efficient model design, an effective global re-detection strategy and a robust distractor awareness mechanism. However, most state-of-the-art long-term trackers (e.g., Pseudo and re-detecting based ones) do not take all three key properties into account and therefore may either be time-consuming or drift to distractors. To address the issues, we propose a two-task tracking framework (named DMTrack), which utilizes two core components (i.e., one-shot detection and re-identification (re-id) association) to achieve distractor-aware fast tracking via Dynamic convolutions (d-convs) and Multiple object tracking (MOT) philosophy. To achieve precise and fast global detection, we construct a lightweight one-shot detector using a novel dynamic convolutions generation method, which provides a unified and more flexible way for fusing target information into the search field. To distinguish the target from distractors, we resort to the philosophy of MOT to reason distractors explicitly by maintaining all potential similarities’ tracklets. Benefited from the strength of high recall detection and explicit object association, our tracker achieves state-of-the-art performance on the LaSOT, Ox-UvA, TLP, VOT2018LT and VOT2019LT benchmarks and runs in real-time (3x faster than comparisons)1. Zikai Zhang 0003, Bineng Zhong 0001, Shengping Zhang, Zhenjun Tang, Xin Liu 0011, Zhaoxiang Zhang 0001 |
CVPR | 4 |
| 2021 | Robust Image Hashing With Singular Values Of Quaternion SVDabstractAbstract Image hashing is an efficient technique of many multimedia systems, such as image retrieval, image authentication and image copy detection. Classification between robustness and discrimination is one of the most important performances of image hashing. In this paper, we propose a robust image hashing with singular values of quaternion singular value decomposition (QSVD). The key contribution is the innovative use of QSVD, which can extract stable and discriminative image features from CIE L*a*b* color space. In addition, image features of a block are viewed as a point in the Cartesian coordinates and compressed by calculating the Euclidean distance between its point and a reference point. As the Euclidean distance requires smaller storage than the original block features, this technique helps to make a discriminative and compact hash. Experiments with three open image databases are conducted to validate efficiency of our image hashing. The results demonstrate that our image hashing can resist many digital operations and reaches a good discrimination. Receiver operating characteristic curve comparisons illustrate that our image hashing outperforms some state-of-the-art algorithms in classification performance. Zhenjun Tang, Mengzhu Yu, Heng Yao 0001, Hanyun Zhang, Chunqiang Yu, Xianquan Zhang |
Comput. J. | 1 |
| 2021 | Reversible data hiding for encrypted image based on adaptive prediction error codingabstractAbstract Reversible data hiding (RDH) is a useful technique of data security. Embedding capacity is one of the most important performance of RDH for encrypted image. Many existing RDH algorithms for encrypted image do not reach desirable embedding capacity yet. To address this problem, a new RDH algorithm is proposed for encrypted image based on adaptive prediction error coding. The proposed RDH algorithm uses a block‐based encryption scheme to preserve spatial correlation of original image in the encrypted domain and exploits a novel technique called adaptive prediction error coding to vacate room for data embedding. A key contribution of the proposed RDH algorithm is the adaptive prediction error coding. It can efficiently vacate room from encrypted image block by adaptively coding prediction errors according to block content and thus contributes to a large embedding capacity. Many experiments on benchmark image databases are done to validate performance of the proposed RDH algorithm. The results show that the average embedding rates on the open databases of UCID, BOSSBase and BOWS‐2 are 1.7081, 2.4437 and 2.3083 bpp, respectively. Comparison results illustrate that the proposed RDH algorithm outperforms some state‐of‐the‐art RDH algorithms in embedding capacity. Zhenjun Tang, Mingyuan Pang, Chunqiang Yu, Guijin Fan, Xianquan Zhang |
IET Image Process. | 1 |
| 2021 | Dual-JPEG-image reversible data hiding
Heng Yao 0001, Fanyu Mao, Chuan Qin 0001, Zhenjun Tang |
Inf. Sci. | 4 |
| 2021 | Video hashing with secondary frames and invariant moments
Zhenjun Tang, Shaopeng Zhang, Xianquan Zhang, Zhixin Li 0001, Zhenhai Chen, Chunqiang Yu |
J. Vis. Commun. Image Represent. | 1 |
| 2021 | Robust and fast image hashing with two-dimensional PCA
Xiaoping Liang, Zhenjun Tang, Jingli Wu, Xianquan Zhang |
Multim. Syst. | 2 |
| 2021 | Robust Video Hashing Based on Multidimensional Scaling and Ordinal MeasuresabstractMultimedia hashing is a useful technology of multimedia management, e.g., multimedia search and multimedia security. This paper proposes a robust multimedia hashing for processing videos. The proposed video hashing constructs a high-dimensional matrix via gradient features in the discrete wavelet transform (DWT) domain of preprocessed video, learns low-dimensional features from high-dimensional matrix via multidimensional scaling, and calculates video hash by ordinal measures of the learned low-dimensional features. Extensive experiments on 8300 videos are performed to examine the proposed video hashing. Performance comparisons reveal that the proposed scheme is better than several state-of-the-art schemes in balancing the performances of robustness and discrimination. Zhenjun Tang, Shaopeng Zhang, Zhenhai Chen, Xianquan Zhang |
Secur. Commun. Networks | 1 |
| 2020 | Video Hashing with DCT and NMFabstractAbstract Video hashing is a novel technique of multimedia processing and finds applications in video retrieval, video copy detection, anti-piracy search and video authentication. In this paper, we propose a robust video hashing based on discrete cosine transform (DCT) and non-negative matrix decomposition (NMF). The proposed video hashing extracts secure features from a normalized video via random partition and dominant DCT coefficients, and exploits NMF to learn a compact representation from the secure features. Experiments with 2050 videos are carried out to validate efficiency of the proposed video hashing. The results show that the proposed video hashing is robust to many digital operations and reaches good discrimination. Receiver operating characteristic (ROC) curve comparisons illustrate that the proposed video hashing outperforms some state-of-the-art algorithms in classification between robustness and discrimination. Zhenjun Tang, Lv Chen, Heng Yao 0001, Xianquan Zhang, Chunqiang Yu, Fionn Murtagh |
Comput. J. | 1 |
| 2020 | Robust image hashing with visual attention model and invariant momentsabstractImage hashing is an efficient technique of multimedia processing for many applications, such as image copy detection, image authentication, and social event detection. In this study, the authors propose a novel image hashing with visual attention model and invariant moments. An important contribution is the weighted DWT (discrete wavelet transform) representation by incorporating a visual attention model called Itti saliency model into LL sub‐band. Since the Itti saliency model can efficiently extract saliency map reflecting regions of attention focus, perceptual robustness of the proposed hashing is achieved. In addition, as invariant moments are robust and discriminative features, hash construction with invariant moments extracted from the weighted DWT representation ensures good classification performance between robustness and discrimination. Extensive experiments with open image datasets are done to validate the performances of the proposed hashing. The results demonstrate that the proposed hashing is robust and discriminative. Performance comparisons with some hashing algorithms are also conducted, and the receiver operating characteristic results illustrate that the proposed hashing outperforms the compared hashing algorithms in classification performance between robustness and discrimination. Zhenjun Tang, Hanyun Zhang, Chi-Man Pun, Mengzhu Yu, Chunqiang Yu, Xianquan Zhang |
IET Image Process. | 1 |
| 2020 | High-fidelity dual-image reversible data hiding via prediction-error shift
Heng Yao 0001, Fanyu Mao, Zhenjun Tang, Chuan Qin 0001 |
Signal Process. | 3 |
| 2020 | Relation R-CNN: A Graph Based Relation-Aware Network for Object DetectionabstractDue to the deteriorated quality of feature in the propagation process of the neural network, it may be hard for traditional detector to identify a small object by just utilizing information within one region proposal. To overcome the limitation of the traditional object detector, we proposed a graph based relation-aware network, to capture the relation information from labels, and images. The semantic relation network is proposed to mine the global semantic relation in labels, and the spatial relation network is proposed to capture the local spatial relation in images. The feature representation is further improved by aggregating the outputs of the two networks. Instead of directly disseminating visual features in the network, the relation-aware network explores more advanced feature information. Experiments on the PASCAL VOC, and MS COCO datasets demonstrate that key relation information significantly improve the performance of object detection with better ability to detect small objects, and reasonable bounding box. The results on COCO dataset demonstrate our method can detect objects robustly, increasing the detection performance of small objects from average precision, and average recall by 31.8%, and 32.3% respectively in performance relative to Faster R-CNN. Shengjia Chen, Zhixin Li 0001, Zhenjun Tang |
IEEE Signal Process. Lett. | 3 |
| 2020 | Robust Image Hashing with Low-Rank Representation and Ring PartitionabstractImage hashing has attracted much attention of the community of multimedia security in the past years. It has been successfully used in social event detection, image authentication, copy detection, image quality assessment, and so on. This paper presents a novel image hashing with low-rank representation (LRR) and ring partition. The proposed hashing finds the saliency map by the spectral residual model and exploits it to construct the visual representation of the preprocessed image. Next, the proposed hashing calculates the low-rank recovery of the visual representation by LRR and extracts the rotation-invariant hash from the low-rank recovery by ring partition. Hash similarity is finally determined by L2 norm. Extensive experiments are done to validate effectiveness of the proposed hashing. The results demonstrate that the proposed hashing can reach a good balance between robustness and discrimination and is superior to some state-of-the-art hashing algorithms in terms of the area under the receiver operating characteristic curve. Zhenjun Tang, Zixuan Yu, Zhixin Li 0001, Chunqiang Yu, Xianquan Zhang |
Wirel. Commun. Mob. Comput. | 1 |
| 2019 | Effective reversible data hiding in encrypted image with adaptive encoding strategy
Yujie Fu, Ping Kong, Heng Yao 0001, Zhenjun Tang, Chuan Qin 0001 |
Inf. Sci. | 4 |
| 2019 | Local complexity based adaptive embedding mechanism for reversible data hiding in digital images
Bowen An, Heng Yao 0001, Zhenjun Tang |
Multim. Tools Appl. | 4 |
| 2019 | Reversible data hiding with differential compression in encrypted image
Zhenjun Tang, Heng Yao 0001, Chuan Qin 0001, Xianquan Zhang |
Multim. Tools Appl. | 1 |
| 2019 | Adaptive image camouflage using human visual system model
Heng Yao 0001, Zhenjun Tang, Chuan Qin 0001 |
Multim. Tools Appl. | 3 |
| 2019 | Correction to: Adaptive image camouflage using human visual system model
Heng Yao 0001, Zhenjun Tang, Chuan Qin 0001 |
Multim. Tools Appl. | 3 |
| 2019 | Image Encryption with Double Spiral Scans and Chaotic MapsabstractImage encryption is a useful technique of image content protection. In this paper, we propose a novel image encryption algorithm by jointly exploiting random overlapping block partition, double spiral scans, Henon chaotic map, and Lü chaotic map. Specifically, the input image is first divided into overlapping blocks and pixels of every block are scrambled via double spiral scans. During spiral scans, the start-point is randomly selected under the control of Henon chaotic map. Next, image content based secret keys are generated and used to control the Lü chaotic map for calculating a secret matrix with the same size of input image. Finally, the encrypted image is obtained by calculating XOR operation between the corresponding elements of the scrambled image and the secret matrix. Experimental result shows that the proposed algorithm has good encrypted results and outperforms some popular encryption algorithms. Zhenjun Tang, Chunqiang Yu, Xianquan Zhang |
Secur. Commun. Networks | 1 |
| 2019 | Reversible Data Hiding by Using Adaptive Pixel Value Prediction and Adaptive Embedding Bin SelectionabstractIn this letter, a reversible data hiding (RDH) scheme by using adaptive pixel value prediction and adaptive embedding bin selection based on pixel-based pixel value ordering (PPVO) is proposed. Different from the previous PPVO based methods, each to-be-embedded pixel is predicted by its neighbor pixels, which are selected adaptively according to the complexity of its neighbor pixel values. Moreover, the value range of embedding bin is not constrained in our method. Especially, when the value of embedding bin is less than 0, its value is updated adaptively during the process of embedding and extracting secret bits. Our maximum embedding capacity (EC) is improved significantly due to unconstrained embedding bin. In addition, the optimal embedding bins can be selected to achieve highest visual quality under a given EC. Experimental results show that the proposed method outperforms some state-of-the-art PPVO based RDH methods. Dewang Wang, Xianquan Zhang, Chunqiang Yu, Zhenjun Tang |
IEEE Signal Process. Lett. | 4 |
| 2019 | Robust Image Hashing with Tensor DecompositionabstractThis paper presents a new image hashing that is designed with tensor decomposition (TD), referred to as TD hashing, where image hash generation is viewed as deriving a compact representation from a tensor. Specifically, a stable three-order tensor is first constructed from the normalized image, so as to enhance the robustness of our TD hashing. A popular TD algorithm, called Tucker decomposition, is then exploited to decompose the three-order tensor into a core tensor and three orthogonal factor matrices. As the factor matrices can reflect intrinsic structure of original tensor, hash construction with the factor matrices makes a desirable discrimination of the TD hashing. To examine these claims, there are 14,551 images selected for our experiments. A receiver operating characteristics (ROC) graph is used to conduct theoretical analysis and the ROC comparisons illustrate that the TD hashing outperforms some state-of-the-art algorithms in classification performance between the robustness and discrimination. Zhenjun Tang, Lv Chen, Xianquan Zhang, Shichao Zhang 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2018 | Perceptual Image Hashing with Weighted DWT Features for Reduced-Reference Image Quality AssessmentabstractWe propose a novel perceptual image hashing based on weighted discrete wavelet transform (DWT) statistical features. This hashing converts input image into a normalized image by bi-linear interpolation and color space conversion, extracts edge image of the normalized image via Canny operator, and divides the edge image into non-overlapping blocks. For each block, a three-level 2D DWT is applied to obtain different sub-bands and the weighted sum of the DWT statistics of these sub-bands is calculated. Finally, image hash is generated by concatenating and quantizing these weighted DWT features. Similarity of image hashes is measured by Euclidean distance. The Copydays dataset and the Uncompressed Color Image Database (UCID) are both used to evaluate classification between robustness and discrimination. Receiver operating characteristics curve comparisons illustrate that our hashing is superior to some state-of-the-art algorithms in classification performance with respect to robustness and discrimination. The LIVE Image Quality Assessment Database is used to validate our application in reduced-reference image quality assessment. Experimental results show that our hashing has better performance in image quality assessment than two popular measures, i.e. peak signal-to-noise ratio and structural similarity. Zhenjun Tang, Ziqing Huang, Heng Yao 0001, Xianquan Zhang, Lv Chen, Chunqiang Yu |
Comput. J. | 1 |
| 2018 | Image hashing with color vector angle
Zhenjun Tang, Xuelong Li 0001, Xianquan Zhang, Shichao Zhang 0001, Yumin Dai |
Neurocomputing | 1 |
| 2018 | Expose noise level inconsistency incorporating the inhomogeneity scoring strategy
Heng Yao 0001, Zhenjun Tang |
Multim. Tools Appl. | 3 |
| 2018 | Reversible Data Hiding with Pixel Prediction and Additive Homomorphism for Encrypted ImageabstractData hiding in encrypted image is a recent popular topic of data security. In this paper, we propose a reversible data hiding algorithm with pixel prediction and additive homomorphism for encrypted image. Specifically, the proposed algorithm applies pixel prediction to the input image for generating a cover image for data embedding, referred to as the preprocessed image. The preprocessed image is then encrypted by additive homomorphism. Secret data is finally embedded into the encrypted image via modular 256 addition. During secret data extraction and image recovery, addition homomorphism and pixel prediction are jointly used. Experimental results demonstrate that the proposed algorithm can accurately recover original image and reach high embedding capacity and good visual quality. Comparisons show that the proposed algorithm outperforms some recent algorithms in embedding capacity and visual quality. Chunqiang Yu, Xianquan Zhang, Zhenjun Tang, Jingyu Huang |
Secur. Commun. Networks | 3 |
| 2018 | Improving Alzheimer's Disease Classification by Combining Multiple MeasuresabstractSeveral anatomical magnetic resonance imaging (MRI) markers for Alzheimer's disease (AD) have been identified. Cortical gray matter volume, cortical thickness, and subcortical volume have been used successfully to assist the diagnosis of Alzheimer's disease including its early warning and developing stages, e.g., mild cognitive impairment (MCI) including MCI converted to AD (MCIc) and MCI not converted to AD (MCInc). Currently, these anatomical MRI measures have mainly been used separately. Thus, the full potential of anatomical MRI scans for AD diagnosis might not yet have been used optimally. Meanwhile, most studies currently only focused on morphological features of regions of interest (ROIs) or interregional features without considering the combination of them. To further improve the diagnosis of AD, we propose a novel approach of extracting ROI features and interregional features based on multiple measures from MRI images to distinguish AD, MCI (including MCIc and MCInc), and health control (HC). First, we construct six individual networks based on six different anatomical measures (i.e., CGMV, CT, CSA, CC, CFI, and SV) and Automated Anatomical Labeling (AAL) atlas for each subject. Then, for each individual network, we extract all node (ROI) features and edge (interregional) features, and denoted as node feature set and edge feature set, respectively. Therefore, we can obtain six node feature sets and six edge feature sets from six different anatomical measures. Next, each feature within a feature set is ranked by -score in descending order, and the top ranked features of each feature set are applied to MKBoost algorithm to obtain the best classification accuracy. After obtaining the best classification accuracy, we can get the optimal feature subset and the corresponding classifier for each node or edge feature set. Afterwards, to investigate the classification performance with only node features, we proposed a weighted multiple kernel learning (wMKL) framework to combine these six optimal node feature subsets, and obtain a combined classifier to perform AD classification. Similarly, we can obtain the classification performance with only edge features. Finally, we combine both six optimal node feature subsets and six optimal edge feature subsets to further improve the classification performance. Experimental results show that the proposed method outperforms some state-of-the-art methods in AD classification, and demonstrate that different measures contain complementary information. Jin Liu 0012, Jianxin Wang 0001, Zhenjun Tang, Bin Hu 0001, Fang-Xiang Wu, Yi Pan 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2017 | Guided filtering based color image reversible data hiding
Heng Yao 0001, Chuan Qin 0001, Zhenjun Tang |
J. Vis. Commun. Image Represent. | 3 |
| 2017 | Image encryption based on random projection partition and chaotic system
Zhenjun Tang, Xianquan Zhang |
Multim. Tools Appl. | 1 |
| 2017 | High capacity data hiding based on interpolated image
Xianquan Zhang, Zerui Sun, Zhenjun Tang, Chunqiang Yu |
Multim. Tools Appl. | 3 |
| 2017 | Robust image hashing with multidimensional scaling
Zhenjun Tang, Ziqing Huang, Xianquan Zhang, Huan Lao |
Signal Process. | 1 |
| 2017 | Improved dual-image reversible data hiding method using the selection strategy of shiftable pixels' coordinates with minimum distortion
Heng Yao 0001, Chuan Qin 0001, Zhenjun Tang |
Signal Process. | 3 |
| 2016 | Robust image hashing via DCT and LLE
Zhenjun Tang, Huan Lao, Xianquan Zhang |
Comput. Secur. | 1 |
| 2016 | Robust Image Hashing With Ring Partition and Invariant Vector DistanceabstractRobustness and discrimination are two of the most important objectives in image hashing. We incorporate ring partition and invariant vector distance to image hashing algorithm for enhancing rotation robustness and discriminative capability. As ring partition is unrelated to image rotation, the statistical features that are extracted from image rings in perceptually uniform color space, i.e., CIE L*a*b* color space, are rotation invariant and stable. In particular, the Euclidean distance between vectors of these perceptual features is invariant to commonly used digital operations to images (e.g., JPEG compression, gamma correction, and brightness/contrast adjustment), which helps in making image hash compact and discriminative. We conduct experiments to evaluate the efficiency with 250 color images, and demonstrate that the proposed hashing algorithm is robust at commonly used digital operations to images. In addition, with the receiver operating characteristics curve, we illustrate that our hashing is much better than the existing popular hashing algorithms at robustness and discrimination. Zhenjun Tang, Xianquan Zhang, Xianxian Li, Shichao Zhang 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2015 | Efficient image encryption with block shuffling and chaotic map
Zhenjun Tang, Xianquan Zhang, Weiwei Lan |
Multim. Tools Appl. | 1 |
| 2014 | Discovery of Tampered Image with Robust Hashing
Zhenjun Tang, Xianquan Zhang, Shichao Zhang 0001 |
ADMA | 1 |
| 2014 | Robust image hashing via colour vector angles and discrete wavelet transformabstractColour vector angle has been widely used in edge detection and image retrieval, but its investigation in image hashing is still limited. In this study, the authors investigate the use of colour vector angle in image hashing and propose a robust hashing algorithm combining colour vector angles with discrete wavelet transform (DWT). Specifically, the input image is firstly resized to a normalised size by bi‐cubic interpolation and blurred by a Gaussian low‐pass filter. Colour vector angles are then calculated and divided into non‐overlapping blocks. Next, block means of colour vector angles are extracted to form a feature matrix, which is further compressed by DWT. Image hash is finally formed by those DWT coefficients in the LL sub‐band. Experiments show that the proposed hashing is robust against normal digital operations, such as JPEG compression, watermarking embedding and rotation within 5°. Receiver operating characteristics curve comparisons are conducted and the results show that the proposed hashing is better than some well‐known algorithms. Zhenjun Tang, Yumin Dai, Xianquan Zhang, Liyan Huang |
IET Image Process. | 1 |
| 2014 | Robust Perceptual Image Hashing Based on Ring Partition and NMFabstractThis paper designs an efficient image hashing with a ring partition and a nonnegative matrix factorization (NMF), which has both the rotation robustness and good discriminative capability. The key contribution is a novel construction of rotation-invariant secondary image, which is used for the first time in image hashing and helps to make image hash resistant to rotation. In addition, NMF coefficients are approximately linearly changed by content-preserving manipulations, so as to measure hash similarity with correlation coefficient. We conduct experiments for illustrating the efficiency with 346 images. Our experiments show that the proposed hashing is robust against content-preserving operations, such as image rotation, JPEG compression, watermark embedding, Gaussian low-pass filtering, gamma correction, brightness adjustment, contrast adjustment, and image scaling. Receiver operating characteristics (ROC) curve comparisons are also conducted with the state-of-the-art algorithms, and demonstrate that the proposed hashing is much better than all these algorithms in classification performances with respect to robustness and discrimination. Zhenjun Tang, Xianquan Zhang, Shichao Zhang 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2013 | Learning semantic concepts from image database with hybrid generative/discriminative approach
Zhixin Li 0001, Zhongzhi Shi, Weizhong Zhao, Zhenjun Tang |
Eng. Appl. Artif. Intell. | 5 |
| 2013 | Robust image hashing using ring-based entropies
Zhenjun Tang, Xianquan Zhang, Liyan Huang, Yumin Dai |
Signal Process. | 1 |
| 2012 | Restoration of embedded image from corrupted stego image
Xianquan Zhang, Zhenjun Tang, Xuan Dai |
Signal Process. | 3 |
| 2011 | Structural Feature-Based Image Hashing and Similarity Metric for Tampering DetectionabstractStructural image features are exploited to construct perceptual image hashes in this work. The image is first preprocessed and divided into overlapped blocks. Correlation between each image block and a reference pattern is calculated. The intermediate hash is obtained from the correlation coefficients. These coefficients are finally mapped to the interval [0, 100], and scrambled to generate the hash sequence. A key component of the hashing method is a specially defined similarity metric to measure the “distance” between hashes. This similarity metric is sensitive to visually unacceptable alterations in small regions of the image, enabling the detection of small area tampering in the image. The hash is robust against content-preserving processing such as JPEG compression, moderate noise contamination, watermark embedding, re-scaling, brightness and contrast adjustment, and low-pass filtering. It has very low collision probability. Experiments are conducted to show performance of the proposed method. Zhenjun Tang, Shuozhong Wang, Xinpeng Zhang 0001, Weimin Wei |
Fundam. Informaticae | 1 |
| 2011 | Lexicographical framework for image hashing with implementation based on DCT and NMF
Zhenjun Tang, Shuozhong Wang, Xinpeng Zhang 0001, Weimin Wei |
Multim. Tools Appl. | 1 |
| 2010 | Estimation of image rotation angle using interpolation-related spectral signatures with application to blind detection of image forgeryabstractMotivated by the image rescaling estimation method proposed by Gallagher (2nd Canadian Conf. Computer & Robot Vision, 2005: 65-72), we develop an image rotation angle estimator based on the relations between the rotation angle and the frequencies at which peaks due to interpolation occur in the spectrum of the image's edge map. We then use rescaling/rotation detection and parameter estimation to detect fake objects inserted into images. When a forged image contains areas from different sources, or from another part of the same image, rescaling and/or rotation are often involved. In these geometric operations, interpolation is a necessary step. By dividing the image into blocks, detecting traces of rescaling and rotation in each block, and estimating the parameters, we can effectively reveal the forged areas in an image that have been rescaled and/or rotated. If multiple geometrical operations are involved, different processing sequences, i.e., repeated zooming, repeated rotation, rotation-zooming, or zooming-rotation, may be determined from different behaviors of the peaks due to rescaling and rotation. This may also provide a useful clue to image authentication. Weimin Wei, Shuozhong Wang, Xinpeng Zhang 0001, Zhenjun Tang |
IEEE Trans. Inf. Forensics Secur. | 4 |