Weijia Cao

dblp:123/2648 · DBLP profile ↗
← Back
28ranked-venue papers
5as first author
22since 2021 · last 2026
0000-0002-6038-8014ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 14 · 3 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 first-authorComputer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 Improved LSB substitution based semi-blind fragile watermarking for high-accuracy tamper localization
Weijia Cao
J. Inf. Secur. Appl.2
2026 RWKVSR: Receptance Weighted Key-Value Network for Hyperspectral Image Super-Resolution
abstract
Deep learning has achieved significant success in hyperspectral image super-resolution (HSISR) by leveraging advanced feature extraction techniques to reconstruct high-resolution images from low-resolution counterparts. However, existing methods predominantly utilize 2D/3D convolutions or Transformer architectures, which are often hindered by limited receptive fields, quadratic computational complexity, and inadequate fusion of spatial-spectral dependencies. To address these challenges, this paper proposes RWKVSR, a novel lightweight network that integrates a Receptance Weighted Key-Value (RWKV) architecture for efficient HSISR. The proposed RWKVSR comprises of three key components: (1) A linear-complexity RWKV module replacing quadratic self-attention, enabling efficient global spectral-spatial modeling; (2) A Spectral-Spatial Residual Module (SSRM) employing anisotropic, direction-separable 3D convolutions to hierarchically extract multi-scale features while enhancing local-global interactions; and (3) A Hyperspectral Frequency Loss (HFL) optimizing spectral consistency by prioritizing high-frequency structural alignment between reconstructed and ground-truth images in the frequency domain. Extensive experiments conducted on the CAVE and Harvard datasets demonstrate that RWKVSR outperforms the existing state-of-the-art methods, effectively balancing accuracy and efficiency, and providing a practical solution for high-quality HSI reconstruction. Our paper code is publicly available at https://github.com/backy-1/RWKVSR.git.
Xiaofei Yang 0002, Sihuan Li, Weijia Cao, Yifang Ban, Yicong Zhou
IEEE Trans. Circuits Syst. Video Technol.3
2026 Reversible Data Hiding over Encrypted Images via Intrinsic Correlation in Block-Based Secret Sharing
abstract
Reversible data hiding over encrypted images (RDH-EI) is an important technique for secure cloud image management but existing schemes often exhibit high computational complexity, low embedding rates, and excessive data expansion. This article addresses these issues by analyzing block-based secret sharing, revealing significant intra-block data redundancy. Based on this observation, we propose two space-preserving methods: the direct space-vacating method and the image-shrinking-based space-vacating method. Using these techniques, we design two novel RDH-EI schemes: a high-capacity RDH-EI scheme and a size-reduced RDH-EI scheme. The high-capacity RDH-EI scheme directly creates embedding space in encrypted images, eliminating the need for complex space-vacating operations and achieving higher and more stable embedding rates. In contrast, the size-reduced RDH-EI scheme minimizes data expansion by discarding unnecessary shares, resulting in smaller encrypted images. Experimental results show that the high-capacity RDH-EI scheme outperforms existing methods in terms of embedding capacity, while the size-reduced RDH-EI scheme achieves strong performance in minimizing data expansion. Both schemes offer effective solutions for RDH-EI challenges.
Jianhui Zou, Weijia Cao, Nankun Mu, Yifeng Zheng 0001, Zhaoquan Gu, Zhongyun Hua
ACM Trans. Multim. Comput. Commun. Appl.2
2025 Multispectral Image Recompression in Ciphertext Domain With Texture Block Decision and 3D-MDCT
abstract
The rapid advancement of remote sensing (RS) technology has posed increasing demands for secure and efficient processing of multispectral data. However, conventional joint image encryption and compression schemes, originally developed for natural images, are not well suited to the specific requirements of multispectral RS scenarios, such as managing interband redundancy, capturing spatial texture variation, and preserving spectral consistency for downstream applications. To address these challenges, we propose a joint encryption and compression algorithm for multispectral images (JECA-MS), the first joint encryption and compression framework specifically designed for multispectral RS images with support for ciphertext domain recompression. The JECA-MS incorporates four key innovations: 1) an adaptive two-size texture block decision (TBD) strategy that classifies image regions into strong and weak texture blocks (WTBs), reducing data volume in weak-texture areas by up to fourfold; 2) a modified 3-D discrete cosine transform (3D-MDCT) that enhances spatial–spectral decorrelation, particularly in homogeneous regions such as clouds and water; 3) a ciphertext domain recompression mechanism that enables flexible adjustment of compression ratios (CRs) without decryption; and 4) a dedicated JECA-MS coding format (JECA-MS-CF) for efficient data encapsulation and compatibility with RS data structures. Extensive experiments show that the JECA-MS achieves 55% and 36% improvements in CRs for water and cloud images, while reducing encoding and decoding time by 39% and 68%, compared to state-of-the-art methods. Security evaluation shows that the JECA-MS can resist statistical attacks, achieve a tradeoff between lightweight encryption and compression performance. This work offers a flexible solution for secure and efficient RS data management.
Xiaoran Leng, Weijia Cao, Tao Yu 0001, Xingfa Gu
IEEE Trans. Geosci. Remote. Sens.2
2025 ACTN: Adaptive Coupling Transformer Network for Hyperspectral Image Classification
abstract
Convolutional neural networks (CNNs) and Transformer networks have shown impressive performance in hyperspectral image (HSI) classification. However, these models usually concentrate on examining either local or global representations of HSI data, frequently falling short of capturing multidimensional representations. Furthermore, these methods fail to fully leverage the strengths of CNNs and Transformers. This article presents the adaptive coupling Transformer network (ACTN), a parallel-hybrid network aiming to improve representation learning for HSI classification. ACTN can capture different types of representation and facilitate mutual learning. Specifically, we introduce a parallel-hybrid module called the adaptive coupling module (ACM), which is designed to capture multifaceted representations from the HSI cube. The ACM consists of two branches: a CNN branch that extracts local contextual representations and a Transformer branch that captures global dependency representations. Our proposal is an adaptive response fusion module (ARFM) that interacts with the hybrid module to merge local and global representations at different resolutions in an adaptive way. In addition, we utilize a cosine similarity function to restrict the loss function in mutual learning, guaranteeing the preservation of both local and global representations to the maximum extent. Extensive experiments conducted on three public HSI datasets demonstrate that ACTN outperforms state-of-the-art methods based on Transformers and CNNs.
Xiaofei Yang 0002, Weijia Cao, Yicong Zhou, Yao Lu 0008
IEEE Trans. Geosci. Remote. Sens.2
2024 Few-shot image classification via hybrid representation
Baodi Liu, Shuai Shao 0006, Lei Xing 0005, Weifeng Liu 0001, Weijia Cao, Yicong Zhou
Pattern Recognit.6
2023 Thematic relations outperform taxonomic relations in a cued recall task
Weijia Cao, Omri Raccah, Yi-Ping Phoebe Chen, David Poeppel
CogSci1
2023 Nonlocal Correntropy Matrix Representation for Hyperspectral Image Classification
abstract
Hyperspectral image (HSI) classification is a hot topic in the remote sensing community. However, it is challenging to fully use spatial–spectral information for HSI classification due to the high dimensionality of the data, high intraclass variability, and the limited availability of training samples. To deal with these issues, we propose a novel feature extraction method called nonlocal correntropy matrix (NLCM) representation in this letter. NLCM can characterize the spectral correlation and effectively extract discriminative features for HSI classification. We verify the effectiveness of the proposed method on two widely used datasets. The results show that NLCM performs better than the state-of-the-art methods, especially when the training set size is small. Furthermore, the experimental results also demonstrate that the proposed method outperforms compared methods significantly when the land covers are complex and with irregular distributions.
Guochao Zhang, Xueting Hu, Yantao Wei, Weijia Cao, Huang Yao, Xueyang Zhang, Keyi Song
IEEE Geosci. Remote. Sens. Lett.4
2023 Shared Dictionary Learning Via Coupled Adaptations for Cross-Domain Classification
Yuying Cai, Baodi Liu, Weijia Cao, Honglong Chen, Weifeng Liu 0001
Neural Process. Lett.4
2023 Dynamic Feature Attention Network for Remote Sensing Image Dehazing
Wenzong Jiang, Weifeng Liu 0001, Weijia Cao, Baodi Liu
Neural Process. Lett.4
2023 Cross-Domain Few-Shot classification via class-shared and class-specific dictionaries
Lei Xing 0005, Baodi Liu, Dapeng Tao, Weijia Cao, Weifeng Liu 0001
Pattern Recognit.5
2023 Self-Paced Hard Task-Example Mining for Few-Shot Classification
abstract
In recent years, researchers have commonly employed assistant tasks to enhance the training phase of the few-shot classification models. Several methods have been proposed to exploit and optimize the training tasks, such as Curriculum Learning (CL) and Hard Example Mining (HEM). However, most of the existing strategies can not elaborately leverage the training tasks and share some common drawbacks, including 1) the ignorance of the target tasks’ properties, and 2) the neglect of sample relationships. In this work, we propose a Self-Paced Hard tAsk-Example Mining (SP-HAEM) method to solve these problems. Specifically, the SP-HAEM automatically chooses hard examples via the similarity between training and target tasks to optimize the support set. To represent the property of target tasks, SP-HAEM obtains a representation of the dataset, called “meta-task”. No need to apply an additional model to measure difficulty and choose hard examples like other HEM methods, SP-HAEM selects the tasks with large optimal transport distance to the meta-task as hard tasks. Thus, training with such hard tasks can not only enhances the generalization ability of the model but also eliminate the negative effect of redundancy tasks. To evaluate the effectiveness of SP-HAEM, we conduct extensive experiments on a variety of datasets, including MiniImageNet, TieredImageNet, and FC100. The results of the experiments show that SP-HAEM can achieve higher accuracy compared with the typical few-shot classification models, e.g., Prototypical Network, MAML, FEAT, and MTL.
Xinghao Yang, Xingxing Yao, Dapeng Tao, Weijia Cao, Weifeng Liu 0001
IEEE Trans. Circuits Syst. Video Technol.5
2023 QTN: Quaternion Transformer Network for Hyperspectral Image Classification
abstract
Numerous state-of-the-art transformer-based techniques with self-attention mechanisms have recently been demonstrated to be quite effective in the classification of hyperspectral images (HSIs). However, traditional transformer-based methods severely suffer from the following problems when processing HSIs with three dimensions: (1) processing the HSIs using 1D sequences misses the 3D structure information; (2) too expensive numerous parameters for hyperspectral image classification tasks; (3) only capturing spatial information while lacking the spectral information. To solve these problems, we propose a novel Quaternion Transformer Network (QTN) for recovering self-adaptive and long-range correlations in HSIs. Specially, we first develop a band adaptive selection module (BASM) for producing Quaternion data from HSIs. And then, we propose a new and novel quaternion self-attention (QSA) mechanism to capture the local and global representations. Finally, we propose a new and novel transformer method, i.e., QTN by stacking a series of QSA for hyperspectral classification. The proposed QTN could exploit computation using Quaternion algebra in hypercomplex spaces. Extensive experiments on three public datasets demonstrate that the QTN outperforms the state-of-the-art vision transformers and convolution neural networks.
Xiaofei Yang 0002, Weijia Cao, Yao Lu 0008, Yicong Zhou
IEEE Trans. Circuits Syst. Video Technol.2
2023 RAN: Region-Aware Network for Remote Sensing Image Super-Resolution
abstract
The remote sensing (RS) image super-resolution (SR) algorithm aims to reconstruct a high-resolution (HR) image with rich texture details from a given low-resolution (LR) image, improving the spatial resolution. It has been widely concerned in remote sensing image processing and application. Most current deep learning-based methods rely on paired training datasets. However, most datasets are often based on bicubic degradation. This single construction way limits the performance of the pre-trained network. Moreover, SR is an ill-posed problem in that multiple SR images are constructed from a single LR input. This paper proposes a Region-Aware Network (RAN) for remote sensing image super-resolution to alleviate the above issues. First, we introduce the contrastive learning strategy to mine the latent degraded representation of the image and serve as the prior knowledge of the network. Considering the RS images are acquired in specific scenes that have apparent self-similarity. Then, we propose a Region-Aware Module (RAM) based on attention mechanisms and the graph neural network to explore region information and cross-patch self-similarity. Extensive experiments have demonstrated that the proposed RAN adapts to RS image super-resolution tasks with various degradations and performs better in constructing texture information.
Baodi Liu, Lifei Zhao, Shuai Shao 0006, Weifeng Liu 0001, Dapeng Tao, Weijia Cao, Yicong Zhou
IEEE Trans. Geosci. Remote. Sens.6
2022 Multi-view learning for hyperspectral image classification: An overview
Baodi Liu, Kai Zhang 0029, Honglong Chen, Weijia Cao, Weifeng Liu 0001, Dapeng Tao
Neurocomputing5
2022 Multiorder Interaction Information Embedding-Based Multiview Fusion-Aided Hyperspectral Image Classification
abstract
Hyperspectral images (HSI) are obtained from hyperspectral imaging sensors, which capture information in hundreds of spectral bands of objects. However, how to take full advantage of spatial and spectral information from many spectral bands to improve the performance of HSI classification remains an open question. Many HSI classification works have recently been reported by employing multi-view learning (MVL) algorithms that can fully use complementary information between different view features and thus have received widespread attention. This paper proposes a multi-view fusion network based on multi-order interaction information embedding for HSI classification. Firstly, the correlation matrix between spectral bands is used to divide the original data into multiple subsets as local views. The subset after the Segmented-PCA process is used as the global view. Secondly, the features of different views are extracted separately using a feature extraction network and mapped to the same dimension. Pre-fusion is achieved by multi-order interaction of various view features. Finally, loss-weighted fusion is applied to each view according to its contribution to the classification task. To evaluate the effectiveness of the proposed method, complete experiments were conducted on three commonly used HSI datasets, namely Pavia University, Houston 2013, and Houston 2018. The experimental results demonstrate that the proposed method improves the classification performance of existing feature extraction networks and is more competitive with other methods in the field.
Weijia Cao, Kai Zhang 0029, Baodi Liu, Dapeng Tao, Weifeng Liu 0001
IEEE Geosci. Remote. Sens. Lett.2
2022 Rethinking Few-Shot Remote Sensing Scene Classification: A Good Embedding Is All You Need?
abstract
In recent years, few-shot remote sensing scene classification (FSRSSC) has attracted more and more attention. For FSRSSC, most methods currently focus on designing a meta-learning algorithm, which obtains meta-knowledge from limited samples and then applies it to novel tasks. In this work, on the one hand, we optimize the training pipeline of the feature extractor; on the other hand, we apply a novel model fusion method further to optimize the feature extractor capability of the feature extractor. We show a novel few-shot remote sensing scene classification baseline: learning two feature representations through using two self-supervised methods on the meta-training set and then fusing the two representations into one. Then, training a linear classifier on this representation achieves state-of-the-art performance. It shows that training a good feature extractor can be more efficient than complex meta-learning algorithms for FSRSSC. We believe that our results can inspire a rethinking of few-shot remote sensing scene classification benchmarks.
Lei Xing 0005, Yuteng Ma, Weijia Cao, Shuai Shao 0006, Weifeng Liu 0001, Baodi Liu
IEEE Geosci. Remote. Sens. Lett.3
2022 Class Shared Dictionary Learning for Few-Shot Remote Sensing Scene Classification
abstract
In the field of remote sensing, it is infeasible to collect a large number of labeled samples due to imaging equipment and the imaging environment. Few-Shot Learning (FSL) is the dominant method to alleviate this problem, which pursues quickly adapting to novel categories from a limited number of labeled samples. The few-shot Remote Sensing Scene Classification (RSSC) generally includes the pre-training and meta-test phases. However, a “negative transfer” problem exists that data categories in both phases are different. It causes the pre-trained feature extractor to be unable well-adapted to the novel data category. This paper proposes Class Shared Dictionary Learning for Few-Shot Remote Sensing Scene Classification (CSDL) to address this issue. Specifically, this paper designs the Mirror-based Feature Extractor (MFE) in the pre-training phase, constructing a self-supervised classification task to improve the feature extractor robustness. Furthermore, this paper proposes a Class Shared Dictionary classifier (CSD) based on dictionary learning. The CSD projects the novel data feature in meta-test into subspace to reconstruct more discriminative features and complete the classification task. Extensive experiments on remote sensing datasets have demonstrated that the proposed CSDL achieves the advanced classification performance.
Lei Xing 0005, Lifei Zhao, Weijia Cao, Xinmin Ge, Weifeng Liu 0001, Baodi Liu
IEEE Geosci. Remote. Sens. Lett.3
2022 Hyperspectral Image Transformer Classification Networks
abstract
Hyperspectral image (HSI) classification is an important task in earth observation missions. Convolution neural networks (CNNs) with the powerful ability of feature extraction have shown prominence in HSI classification tasks. However, existing CNN-based approaches cannot sufficiently mine the sequence attributes of spectral features, hindering the further performance promotion of HSI classification. This article presents a hyperspectral image transformer (HiT) classification network by embedding convolution operations into the transformer structure to capture the subtle spectral discrepancies and convey the local spatial context information. HiT consists of two key modules, i.e., spectral-adaptive 3-D convolution projection module and convolution permutator (ConV-Permutator) to retrieve the subtle spatial–spectral discrepancies. The spectral-adaptive 3-D convolution projection module produces the local spatial–spectral information from HSIs using two spectral-adaptive 3-D convolution layers instead of the linear projection layer. In addition, the Conv-Permutator module utilizes the depthwise convolution operations to separately encode the spatial–spectral representations along the height, width, and spectral dimensions, respectively. Extensive experiments on four benchmark HSI datasets, including Indian Pines, Pavia University, Houston2013, and Xiongan (XA) datasets, show the superiority of the proposed HiT over existing transformers and the state-of-the-art CNN-based methods. Our codes of this work are available athttps://github.com/xiachangxue/DeepHyperXfor the sake of reproducibility.
Xiaofei Yang 0002, Weijia Cao, Yao Lu 0008, Yicong Zhou
IEEE Trans. Geosci. Remote. Sens.2
2022 Self-Supervised Learning With Prediction of Image Scale and Spectral Order for Hyperspectral Image Classification
abstract
In recent years, Convolutional Neural Networks (CNNs) have achieved great success in hyperspectral image classification attributed to their unparalleled capacity to extract the local information. However, to successfully learn the high-level semantic image features, they always require massive amounts of manually labeled data during the training process, which is expensive, scarce, and impractical, and severely hinders the improvement of supervised deep learning methods. To alleviate these burdens, we present Self-Supervised Learning methods for hyperspectral image classification by a pre-training model using extensive unlabeled data and fine-tuning the hyperspectral image target classification. In this paper, we propose a new method for learning image characteristics by training a CNN to recognize the image scale that is applied to the hyperspectral images (HSIs). In addition, we propose a multi-pretext task method to learn stable and good feature representations combing two different pretext task methods and contrastive loss function. We evaluate the proposed methods in Self-Supervised Learning benchmarks on four benchmark HSIs datasets. The experiment results demonstrate that the proposed methods outperform the traditional supervised deep learning methods when large amounts of unlabeled HSIs data are used. Moreover, it demonstrates that the Self-Supervised Learning method is promising to alleviate dependence on manually labeled data of hyperspectral image classification. Finally, our research contributes to the creation and refinement of Self-Supervised Learning methods for pretextual tasks within the HSIs community.
Xiaofei Yang 0002, Weijia Cao, Yao Lu 0008, Yicong Zhou
IEEE Trans. Geosci. Remote. Sens.2
2022 Local Correntropy Matrix Representation for Hyperspectral Image Classification
abstract
The hyperspectral images (HSIs) classification technique has received widespread attention in the field of remote sensing. However, how to achieve satisfactory classification performance in the presence of a large amount of noise is still a problem worthy of consideration. In this article, a local correntropy matrix (LCEM)-based spatial–spectral feature representation method is proposed for HSI classification. Motivated by the successful application of information-theoretic learning (ITL), we propose to adopt correntropy matrix to represent the spatial–spectral features of HSI. Specifically, the dimension reduction is first performed on the original hyperspectral data. Then, for each pixel, we select its local neighbors within a sliding window using cosine distance for the construction of the LCEM. In this way, each pixel can be characterized as an LCEM. Finally, all the correntropy matrices are fed into a support vector machine (SVM) for final classification. In addition, we also propose a novel way to determine the size of the local window based on standard deviation. Because the LCEM as the feature descriptor can characterize discriminative spatial–spectral features, the proposed method has shown great interclass separability and intraclass compactness. Compared with other advanced approaches, the proposed LCEM method has achieved competitive performance in both evaluation indexes and visual effects, especially when the training size is very small.
Xinyu Zhang 0025, Yantao Wei, Weijia Cao, Huang Yao, Jiangtao Peng, Yicong Zhou
IEEE Trans. Geosci. Remote. Sens.3
2021 SAR Image Classification Using Greedy Hierarchical Learning With Unsupervised Stacked CAEs
abstract
Synthetic aperture radar (SAR) can provide stable data source for earth observation due to its advantages of all day and night, all-weather, and strong penetration. SAR image classification as a fundamental procedure has been proved its great value in plenty of remote sensing applications. Conventional classification algorithms mainly rely on hand-designed features, which are susceptible to widespread coherent speckle noise and geometric distortion in high-resolution SAR images. Inspired by the recent impressive success in data mining and deep learning, a greedy hierarchical convolutional neural network (GHCNN) is developed. It aims at obtaining optimized feature representation, relieving the effect of speckle noise, and promoting the local pattern recognition of geometric distortion in single-polarized SAR image classification. First, a series of convolutional autoencoders (CAEs) is trained in the greedy layer-wise unsupervised strategy. This step provides an unbiased regularizer anda prioridistribution derived from large volumes of unlabeled SAR patches. Then, to optimize multiple parameter subspaces globally, several CAEs are coupled together to form a deeper hierarchical structure in a stacked and unsupervised fashion. Afterward, a convolutional network with identical topology inherits the pretrained weights. After supervised finetuning, it realizes class prediction. Synchronously, t-distributed stochastic neighbor embedding (t-SNE) algorithm is applied to monitor the efficiency of feature representation during the training period. Experimental results demonstrate that the proposed method has competitive advantages over involved contrast methods.
Zhensheng Sun, Peng Liu 0024, Weijia Cao, Tao Yu 0001, Xingfa Gu
IEEE Trans. Geosci. Remote. Sens.4
2020 Designing a 2D infinite collapse map for image encryption
Weijia Cao, Yujun Mao, Yicong Zhou
Signal Process.1
2017 4 × 4 parametric integer discrete cosine transforms
abstract
One of the famous transformations is discrete cosine transform (DCT), which is always used in digital image coding standards like JPEG and MPEG. DCT has different types and all of DCTs have excellent energy compaction properties. Meanwhile, matrices of DCTs II and IV are examined by a lot of researchers. However, other types of DCTs are rarely developed. Therefore, this paper presents 4 × 4 parametric integer DCTs, which cannot only represent DCTs II and IV but other types of DCTs (i.e. DCT I, V, VIII). It shows an excellent performance in terms of mean square errors and transform coding gains while it is comparing with state of the art.
Weijia Cao, Yicong Zhou
SMC1
2017 Medical image encryption using edge maps
Weijia Cao, Yicong Zhou, C. L. Philip Chen, Liming Xia
Signal Process.1
2015 Fast Fourier transform using matrix decomposition
Yicong Zhou, Weijia Cao, Licheng Liu, Sos S. Agaian, C. L. Philip Chen
Inf. Sci.2
2014 Image encryption using binary bitplane
Yicong Zhou, Weijia Cao, C. L. Philip Chen
Signal Process.2
2012 A new image encryption algorithm using Truncated P-Fibonacci Bit-planes
abstract
Image encryption is an effective approach to protect privacy and security of images. This paper introduces a novel image encryption algorithm using the Truncated P-Fibonacci Bit-planes as security key images to encrypt images. Simulation results and security analysis are provided to show the encryption performance of the proposed algorithm.
Weijia Cao, Yicong Zhou, C. L. Philip Chen
SMC1