Wen-Shuai Hu

dblp:241/5297 · DBLP profile ↗
← Back
12ranked-venue papers
5as first author
10since 2021 · last 2027
0000-0002-4757-2765ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 9 · 4 first-author · 7 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2027 Learning tensor correlation filter with fused low-rank and smoothness priors for hyperspectral video object tracking
Wen-Shuai Hu, Jian-Li Wang, Ran Tao 0003, Qian Du 0001
Expert Syst. Appl.3
2025 Unsupervised Domain Adaptation With Hierarchical Masked Dual-Adversarial Network for End-to-End Classification of Multisource Remote Sensing Data
abstract
Although unsupervised domain adaptation (UDA) has been successfully applied for cross-scene classification of multisource remote sensing (MSRS) data, there are still some tough issues: 1) The vast majority of them are patch-based, requiring pixel by pixel processing at high complexity and ignoring the roles of unlabeled data between different domains. 2) Traditional masked autoencoder (MAE)-based methods lack effective multiscale analysis and require pre-training, ignoring the roles of low-level representations. As such, a hierarchical masked dual-adversarial DA network (HMDA-DANet) is proposed for cross-domain end-to-end classification of MSRS data. Firstly, a hierarchical asymmetric MAE (HAMAE) without pre-training is designed, containing a frequency dynamic large-scale convolutional (FDLConv) block to enhance important structural information in the frequency domain, and an intramodality enhancement and intermodality interaction (IAEIEI) block to embed some additional information beyond the domain distribution by expanding the cross-modal reconstruction space. Representative multimodal multiscale features can be extracted, while to some extent improving their generalization to the target domain. Then, a multimodal multiscale feature fusion (MMFF) block is built to model the spatial and scale dependencies for feature fusion and reduce the layer by layer transmission of redundancy or interference information. Finally, a dual-discriminator-based DA (DDA) block is designed for class-specific semantic feature and global structural alignments in both spatial and prediction spaces. It will enable HAMAE to model the cross-modal, cross-scale, and cross-domain associations, yielding more representative domain-invariant multimodal fusion features. Extensive experiments on five cross-domain MSRS datasets verify the superiority of the proposed HMDA-DANet over other state-of-the-art methods.
Wen-Shuai Hu, Wei Li 0032, Heng-Chao Li 0001, Xudong Zhao 0003, Mengmeng Zhang 0005, Ran Tao 0003
IEEE Trans. Geosci. Remote. Sens.1
2025 Chirplet Fourier Analysis Network for Cross-Scene Classification of Multisource Remote Sensing Data
abstract
The joint application of multisource remote sensing (MSRS) data, such as hyperspectral image (HSI) and light detection and ranging (LiDAR), offers significant potential for accurate land cover classification. However, the existing applications often struggle with domain shifts across scenes caused by sensor, illumination, and phase variations. Focusing on this domain adaptation problem, a Chirplet Fourier analysis network (ChirpFAN) is proposed for cross-scene classification of MSRS data in this paper. Firstly, a fractional spatial-frequency-phase feature extraction module including the fractional Fourier transform and a learnable phase-aware weighting block is proposed to capture multi-domain features. Secondly, a Chirplet swin transformer (ChirpST) block integrates a Chirplet Fourier analysis (ChirpFA) layer within a Swin transformer is designed to analyze multi-scale textural and oscillatory patterns. Finally, a modality-shared network including ChirpST blocks is designed for inter-modal fusion and alignment. Extensive experiments demonstrate that the ChirpFAN framework achieves state-of-the-art performance with 3% average improvements on three challenging cross-scene MSRS datasets. Code will be released on GitHub.
Xudong Zhao 0003, Qi Ming, Yixiao Yang, Wen-Shuai Hu, Wei Li 0032, Ran Tao 0003
IEEE Trans. Geosci. Remote. Sens.4
2025 Global Clue-Guided Cross-Memory Quaternion Transformer Network for Multisource Remote Sensing Data Classification
abstract
Multisource remote sensing data classification is a challenging research topic, and how to address the inherent heterogeneity between multimodal data while exploring their complementarity is crucial. Existing deep learning models usually directly adopt feature-level fusion designs, most of which, however, fail to overcome the impact of heterogeneity, limiting their performance. As such, a multimodal joint classification framework, called global clue-guided cross-memory quaternion transformer network (GCCQTNet), is proposed for multisource data [i.e., hyperspectral image (HSI) and synthetic aperture radar (SAR)/light detection and ranging (LiDAR)] classification. First, a three-branch structure is built to extract the local and global features, where an independent squeeze-expansion-like fusion (ISEF) structure is designed to update the local and global representations by considering the global information as an agent, suppressing the negative impact of multimodal heterogeneity layer by layer. A cross-memory quaternion transformer (CMQT) structure is further constructed to model the complex inner relationships between the intramodality and intermodality features to capture more discriminative fusion features that fully characterize multimodal complementarity. Finally, a cross-modality comparative learning (CMCL) structure is developed to impose the consistency constraint on global information learning, which, in conjunction with a classification head, is used to guide the end-to-end training of GCCQTNet. Extensive experiments on three public multisource remote sensing datasets illustrate the superiority of our GCCQTNet with regards to other state-of-the-art methods.
Wen-Shuai Hu, Wei Li 0032, Heng-Chao Li 0001, Ran Tao 0003
IEEE Trans. Neural Networks Learn. Syst.1
2022 SAR Image Data Augmentation via Residual and Attention-Based Generative Adversarial Network for Ship Detection
abstract
In recent years, generative adversarial networks (GANs) have been successfully applied to generate the SAR images. However, due to the fact that it is more difficult to generate the images than to distinguish the real or fake, GANs usually suffer from the problems of unstable training and mode collapse. As such, a residual and attention-based generative adversarial network (RAGAN) is proposed for SAR data augmentation. Firstly, the directional bounding box is used as a constraint in the RAGAN to limit the position of ship in the generated SAR image, which can be further set as the annotation of the SAR image for ship detection directly. After that, inspired by the residual and attention learning, a residual and attention block (RABlock) and a transposed RABlock (TRABlock) are designed to improve the generator of the RAGAN, thus preventing the whole model from gradient vanishing and suppressing the effects of speckle noise and background to enhance the quality of the generated SAR images. Experimental results on the HRSID data set demonstrate the effectiveness of our RAGAN model in SAR data augmentation for ship detection.
Yu-Shi Guo, Heng-Chao Li 0001, Wen-Shuai Hu, Wei-Ye Wang
IGARSS3
2022 Spatially Variant Gamma-WMM with Extended Variational Inference for Unsupervised PolSAR Classification
abstract
The Wishart mixture model (WMM) has been widely used for classification of polarimetric synthetic aperture radar (PolSAR) images; however, the WMM-based models usually fail to provide reliable classification results and explore the spatial information effectively in the heterogeneous areas. As such, an unsupervised spatially variant Gamma-WMM with extended variational inference algorithm (SVGaWMM-EVI) is proposed for classification of PolSAR images. Firstly, the Gamma prior distribution is imposed on the texture variable of the proposed model, which associates a set of unique texture variables with each data point to utilize the spatial information in the heterogeneous areas. Then, since the existing expectation maximization-based WMM algorithms usually fall into local optimal and update slowly, an extended variational inference algorithm is developed to improve the parameter estimation of our model, where a help function is designed to solve the intractable term. Experimental results on the real-world PolSAR data set demonstrate that our model can obtain better performance than some widely used unsupervised methods.
Heng-Chao Li 0001, Wen-Shuai Hu, Lei Pan 0003
IGARSS3
2022 Recurrent Feedback Convolutional Neural Network for Hyperspectral Image Classification
abstract
Deep neural networks have achieved promising performance for hyperspectral image (HSI) classification. However, due to the limitation of the available labeled samples, the traditional deeper and wider neural networks usually cause the overfitting problem and lose the detailed information. To solve this problem, a brain-like structure, namely spatial attention-driven recurrent feedback convolutional neural network (SARFNN), is proposed by utilizing the recurrent feedback and attention mechanism structures, from which two deep models are further developed for HSI classification. First, a 2-D SARFNN (SARF2DNN) model is developed to learn the spatial features from HSI data. After that, to better exploit the 3-D characteristic, the 3-D version is extended from SARF2DNN, thus constructing an SARF3DNN model to extract joint spatial-spectral features. Moreover, with the help of the idea of brain-likeness, the recurrent feedback module is designed to recover information loss caused by deeper structure and the dimension reduction operation. The experimental results conducted on two HSI data sets show that our SARFNN architecture can achieve more competitive performance than other state-of-the-art algorithms.
Heng-Chao Li 0001, Shuang-Shuang Li, Wen-Shuai Hu, Jun-Huan Feng, Weiwei Sun 0005, Qian Du 0001
IEEE Geosci. Remote. Sens. Lett.3
2022 Adaptive Cross-Attention-Driven Spatial-Spectral Graph Convolutional Network for Hyperspectral Image Classification
abstract
Recently, graph convolutional networks (GCNs) have been developed to explore the spatial relationship between pixels, achieving better classification performance of hyperspectral images (HSIs). However, these methods fail to sufficiently leverage the relationship between spectral bands in HSI data. As such, we propose an adaptive cross-attention-driven spatial–spectral graph convolutional network (ACSS-GCN), which is composed of a spatial GCN (Sa-GCN) subnetwork, a spectral GCN (Se-GCN) subnetwork, and a graph cross-attention fusion module (GCAFM). Specifically, Sa-GCN and Se-GCN are proposed to extract the spatial and spectral features by modeling the correlations between spatial pixels and between spectral bands, respectively. Then, by integrating attention mechanism into information aggregation of the graph, the GCAFM, including three parts, i.e., the spatial graph attention block, the spectral graph attention block, and the fusion block, is designed to fuse the spatial and spectral features, and suppress noise interference in Sa-GCN and Se-GCN. Moreover, the idea of the adaptive graph is introduced to explore an optimal graph through backpropagation during the training process. Experiments on two HSI datasets show that the proposed method achieves better performance than other classification methods.
Jin-Yu Yang, Heng-Chao Li 0001, Wen-Shuai Hu, Lei Pan 0003, Qian Du 0001
IEEE Geosci. Remote. Sens. Lett.3
2022 Pseudo Complex-Valued Deformable ConvLSTM Neural Network With Mutual Attention Learning for Hyperspectral Image Classification
abstract
Convolutional long short-term memory (ConvLSTM) has received much attention for hyperspectral image (HSI) classification due to its ability of modeling long-range correlations, which, however, is vulnerable to too many parameters and insufficient training, limiting its classification accuracy, especially for small samples. Different from it, traditional hand-crafted methods extract the features with basic attributes of HSIs, which can provide the lack of details and interpretability of deep semantic features. However, existing methods fail to incorporate their complementarity for HSI classification. As such, a Pseudo complex-valued (CV) Deformable ConvLSTM Neural Network with mutual Attention learning (APDCLNN) is proposed, providing a new way to realize the collaborative learning of hand-crafted and deep features for HSI classification. First, a 2-D pseudo CV deformable ConvLSTM (PDConvLSTM2D) cell is designed using deformable convolution and complex operations, with which a spatial–spectral PDConvLSTM2D neural network (SSPDCL2DNN) is built to extract scale- and spectral-enhanced deep spatial–spectral features. Then, 3-D Gabor filter is used to extract hand-crafted features, and a mutual attention-based multimodality feature learning and fusion (MAMLF) module is designed to integrate them into deep features for training and optimization of SSPDCL2DNN. Finally, an attention loss subnetwork is designed to refine the classification results. As we know, this is the first attempt to apply the idea of mutual attention learning to fuse hand-crafted and deep features for HSI classification. Extensive experiments on three widely used HSI datasets show the advantages of our model over other deep methods in terms of both quantitative and visual quality.
Wen-Shuai Hu, Heng-Chao Li 0001, Rui Wang 0090, Feng Gao 0005, Qian Du 0001, Antonio Plaza
IEEE Trans. Geosci. Remote. Sens.1
2022 A3 CLNN: Spatial, Spectral and Multiscale Attention ConvLSTM Neural Network for Multisource Remote Sensing Data Classification
abstract
The problem of effectively exploiting the information multiple data sources has become a relevant but challenging research topic in remote sensing. In this article, we propose a new approach to exploit the complementarity of two data sources: hyperspectral images (HSIs) and light detection and ranging (LiDAR) data. Specifically, we develop a new dual-channel spatial, spectral and multiscale attention convolutional long short-term memory neural network (called dual-channel$A^{3}$CLNN) for feature extraction and classification of multisource remote sensing data. Spatial, spectral, and multiscale attention mechanisms are first designed for HSI and LiDAR data in order to learn spectral- and spatial-enhanced feature representations and to represent multiscale information for different classes. In the designed fusion network, a novel composite attention learning mechanism (combined with a three-level fusion strategy) is used to fully integrate the features in these two data sources. Finally, inspired by the idea of transfer learning, a novel stepwise training strategy is designed to yield a final classification result. Our experimental results, conducted on several multisource remote sensing data sets, demonstrate that the newly proposed dual-channel$A^{\,3}$CLNN exhibits better feature representation ability (leading to more competitive classification performance) than other state-of-the-art methods.
Heng-Chao Li 0001, Wen-Shuai Hu, Wei Li 0032, Jun Li 0009, Qian Du 0001, Antonio Plaza
IEEE Trans. Neural Networks Learn. Syst.2
2020 Hyperspectral Image Classification Based on Tensor-Train Convolutional Long Short-Term Memory
abstract
In recent years, deep learning models have shown great advantages for hyperspectral images (HSIs) classification, in which long short-term memory (LSTM) has attracted plenty of attentions for its characteristic of modeling long-range dependencies. However, for the 2-D extended architecture of it (namely 2-D convolutional LSTM, ConvLSTM2D), it is the special gate structures of ConvLSTM2D that leads to a large number of training parameters and high requirements for device storage. To address this shortcoming, in this paper, a lightweight ConvLSTM2D cell is developed by using tensor-train decomposition (TTD) for the compression of training parameters, which is named TT-ConvLSTM2D and further applied to two state-of-the-art ConvLSTM2D-based HSI classification models for verifying its superiority. Experiments on a widely-used Indian Pines HSI data set are conducted, whose results demonstrate that the proposed TT-ConvLSTM2D cell can effectively reduce the number of the parameters and memory requirements of the whole models within a small range of accuracy degradation.
Wen-Shuai Hu, Heng-Chao Li 0001, Tian-Yu Ma, Qian Du 0001, Antonio Plaza, William J. Emery
IGARSS1
2020 Spatial-Spectral Feature Extraction via Deep ConvLSTM Neural Networks for Hyperspectral Image Classification
abstract
In recent years, deep learning has presented a great advance in the hyperspectral image (HSI) classification. Particularly, long short-term memory (LSTM), as a special deep learning structure, has shown great ability in modeling long-term dependencies in the time dimension of video or the spectral dimension of HSIs. However, the loss of spatial information makes it quite difficult to obtain better performance. In order to address this problem, two novel deep models are proposed to extract more discriminative spatial-spectral features by exploiting the convolutional LSTM (ConvLSTM). By taking the data patch in a local sliding window as the input of each memory cell band by band, the 2-D extended architecture of LSTM is considered for building the spatial-spectral ConvLSTM 2-D neural network (SSCL2DNN) to model long-range dependencies in the spectral domain. To better preserve the intrinsic structure information of the hyperspectral data, the spatial-spectral ConvLSTM 3-D neural network (SSCL3DNN) is proposed by extending LSTM to the 3-D version for further improving the classification performance. The experiments, conducted on three commonly used HSI data sets, demonstrate that the proposed deep models have certain competitive advantages and can provide better classification performance than the other state-of-the-art approaches.
Wen-Shuai Hu, Heng-Chao Li 0001, Lei Pan 0003, Wei Li 0032, Ran Tao 0003, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.1