Qiang Zhang 0011

dblp:72/3527-11 · DBLP profile ↗
← Back
37ranked-venue papers
13as first author
31since 2021 · last 2026
0000-0002-7116-9327ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 26 · 8 first-author · 20 since 2021Artificial intelligence and machine learning · 7 · 4 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021
YearPublicationVenuePosition
2026 SGD-SST 2.0: Seamless global daily sea surface temperature products cross-sensors generating from 2003 to 2025
Qi Wang 0183, Qiang Zhang 0011, Tongde Yang, Weizhen Sun, Qiangqiang Yuan
Expert Syst. Appl.2
2026 Deep low-rank tensor embedded network for hyperspectral image super-resolution
Qiang Zhang 0011, Xianpeng Zhang, Yi Xiao 0003, Hongjie Xie
Expert Syst. Appl.1
2026 ES-DETR: Edge-Guided State-Space DETR for Foggy Remote Sensing Object Detection
Qiang Zhang 0011, Zheng Liang 0001, Wenyi Zhao, Weidong Zhang 0007
IEEE Signal Process. Lett.2
2025 Spiking Meets Attention: Efficient Remote Sensing Image Super-Resolution with Attention Spiking Neural Networks
abstract
Spiking neural networks (SNNs) are emerging as a promising alternative to traditional artificial neural networks (ANNs), offering biological plausibility and energy efficiency. Despite these merits, SNNs are frequently hampered by limited capacity and insufficient representation power, yet remain underexplored in remote sensing image (RSI) super-resolution (SR) tasks. In this paper, we first observe that spiking signals exhibit drastic intensity variations across diverse textures, highlighting an active learning state of the neurons. This observation motivates us to apply SNNs for efficient SR of RSIs. Inspired by the success of attention mechanisms in representing salient information, we devise the spiking attention block (SAB), a concise yet effective component that optimizes membrane potentials through inferred attention weights, which, in turn, regulates spiking activity for superior feature representation. Our key contributions include: 1) we bridge the independent modulation between temporal and channel dimensions, facilitating joint feature correlation learning, and 2) we access the global self-similar patterns in large-scale remote sensing imagery to infer spatial attention weights, incorporating effective priors for realistic and faithful reconstruction. Building upon SAB, we proposed SpikeSR, which achieves state-of-the-art performance across various remote sensing benchmarks such as AID, DOTA, and DIOR, while maintaining high computational efficiency. Code of SpikeSR will be available at https://github.com/XY-boy/SpikeSR.
Yi Xiao 0003, Qiangqiang Yuan, Kui Jiang, Wenke Huang 0003, Qiang Zhang 0011, Chia-Wen Lin, Liangpei Zhang 0001
NeurIPS5
2025 Hyperspectral image mixed noised removal via jointly spatial and spectral difference constraint with low-rank tensor factorization
Qiang Zhang 0011, Yaming Zheng, Yushuai Dong, Chunyan Yu, Qiangqiang Yuan
Eng. Appl. Artif. Intell.1
2025 SGD-SST: Seamless global daily sea surface temperature products reconstruction and validation via deep spatio-temporal fusion model
Qi Wang 0143, Qiang Zhang 0011, Hongjie Xie, Zifeng Liu, Yushuai Dong
Expert Syst. Appl.2
2025 10-minute forest early wildfire detection: Fusing multi-type and multi-source information via recursive transformer
Qiang Zhang 0011, Yushuai Dong, Enyu Zhao, Meiping Song, Qiangqiang Yuan
Neurocomputing1
2025 Concern With Center-Pixel Labeling: Center-Specific Perception Transformer Network for Hyperspectral Image Classification
abstract
Self-attention-based approaches that leverage global context information for hyperspectral image (HSI) classification have gained increasing prominence. Nevertheless, due to the assignment of equivalent attention weight to all the tokens (pixels or patches), the existing self-attention mechanism inadvertently prioritizes the non-label-specified information over the instinct label-specified information, which generates attention shifts and redundancy in HSI classification. To alleviate the mentioned barrier, we propose the center-specific perception transformer network (CP-Transformer), which is the first attempt to perform class-guided attention and filter interference factors for HSI classification feature representation. Specifically, the central-pixel focus attention module (CFA) is presented to compute the label-related attention between the center and other pixels. In this manner, CFA reduces computational complexity and closely aligns with the center-pixel labeling strategy. Besides, the spectral saliency focus attention module (SSFA) is developed to capture the spectral correlation by focusing salient bands to provide a beneficial supplement for spatial features. Moreover, the hierarchical integration network (HIN) constructs the inference network to integrate and rectify spatial-spectral features for HSI classification. The experiment results on four popular HSI datasets demonstrate that the proposed method achieves robust performance compared to other state-of-the-art methods. Our code will be released at https://github.com/Chirsycy/CP-Transformer.
Chunyan Yu, Yuanchen Zhu, Yulei Wang 0002, Enyu Zhao, Qiang Zhang 0011, Xiaoqiang Lu
IEEE Trans. Geosci. Remote. Sens.5
2025 Probability-Guided Edge Enhancement Network for Remote Sensing Image Semantic Segmentation
abstract
Semantic segmentation in remote sensing images (RSI) assigns unique semantic labels to each pixel and plays a crucial role in real-world applications such as environmental change monitoring, precision agriculture, and economic assessment. Although convolutional neural networks (CNN) and Transformer-based models for semantic segmentation of RSI have achieved remarkable success, existing approaches still struggle to accurately detect weak edges and occluded objects due to the complexity and fuzziness of edges in RSI. To overcome this obstacle, we propose a novel probability-guided edge enhancement network (PEEN) for semantic segmentation of RSI, which is the first attempt to leverage the probability function to guide the segmentation model in performing edge prediction for RSI. Specifically, in the feature extraction stage of PEEN, we present a convolutional self-attention mechanism to enhance the global feature representation of the encoder-decoder network. In the edge enhancement stage of PEEN, we innovatively build an iterative probability-guided edge prediction module to refine edge prediction mathematically and iteratively. With the cooperation of the mentioned two stages, the proposed model yields precise segmentation of the objects and edge portions in RSI. Experiment results and analysis demonstrate that the PEEN model outperforms the existing popular CNN-based and Transformer-based models in semantic segmentation with 85.54% and 88.35% of Mean Intersection Over Union (mIOU) on the Vaihingen and Potsdam test datasets. Our code is available at https://github.com/Zyk517/PEEN.
Chunyan Yu, Yakun Zuo, Qiang Zhang 0011, Yulei Wang 0002
IEEE Trans. Geosci. Remote. Sens.3
2025 HSTNet: Hybrid Supervision-Driven Two-Stream Collaborative Network for Hyperspectral Wheat Variety Classification
abstract
Hyperspectral remote sensing plays an important role in agricultural monitoring, and fine-grained wheat variety classification is essential for advancing smart agriculture. However, progress is limited by the scarcity of high-quality spectral samples and the challenges associated with collecting large-scale hyperspectral data. To cope with these issues, we design a hybrid supervised-driven two-stream collaborative network (HSTNet), which consists of a semi-supervised conditional generative adversarial network for data augmentation (SCGAN) and a supervised two-stream discriminative network (STDNet) for wheat classification. In SCGAN, the generator constructs a mapping relationship between input noise and real wheat hyperspectral samples to generate fake wheat hyperspectral samples that are highly matched with the distribution of the real sample, and the discriminator with multilayer perceptions utilizes discriminative learning to identify real and fake samples. In STDNet, it collaboratively extracts the spectral, spatial and texture features of wheat hyperspectral images employing the dual-stream branch structure of 3DCNN and 2DCNN. Subsequently, it utilizes the Fast Fourier Transform and the cross-attention mechanism to refine and fuse these features to improve their capability of feature expression. Noteworthy, the individual design effectively improves the classification results of wheat varieties via collaborative optimization among modules. Besides, we built a mixed wheat hyperspectral dataset (MWHD) with 4800 samples of 20 wheat varieties. Extensive experiments on our constructed MWHD dataset demonstrate that the proposed HSTNet outperforms state-of-the-art methods in wheat variety classification. The code is publicly available at: https://github.com/bakam412/HSTNet.
Ling Zhou 0003, Shuoguo Cui, Qiang Zhang 0011, Wenyi Zhao, Zheng Liang 0001, Weidong Zhang 0007
IEEE Trans. Geosci. Remote. Sens.3
2025 Frequency-Assisted Mamba for Remote Sensing Image Super-Resolution
abstract
Recent progress in remote sensing image (RSI) super-resolution (SR) has exhibited remarkable performance using deep neural networks, e.g., Convolutional Neural Networks and Transformers. However, existing SR methods often suffer from either a limited receptive field or quadratic computational overhead, resulting in sub-optimal global representation and unacceptable computational costs in large-scale RSI. To alleviate these issues, we develop the first attempt to integrate the Vision State Space Model (Mamba) for RSI-SR, which specializes in processing large-scale RSI by capturing long-range dependency with linear complexity. To achieve better SR reconstruction, building upon Mamba, we devise a Frequency-assisted Mamba framework, dubbed FMSR, to explore the spatial and frequent correlations. In particular, our FMSR features a multi-level fusion architecture equipped with the Frequency Selection Module (FSM), Vision State Space Module (VSSM), and Hybrid Gate Module (HGM) to grasp their merits for effective spatial-frequency fusion. Considering that global and local dependencies are complementary and both beneficial for SR, we further recalibrate these multi-level features for accurate feature fusion via learnable scaling adaptors. Extensive experiments on AID, DOTA, and DIOR benchmarks demonstrate that our FMSR outperforms state-of-the-art Transformer-based methods HAT-L in terms of PSNR by 0.11 dB on average, while consuming only 28.05% and 19.08% of its memory consumption and complexity, respectively.
Yi Xiao 0003, Qiangqiang Yuan, Kui Jiang, Yuzeng Chen, Qiang Zhang 0011, Chia-Wen Lin
IEEE Trans. Multim.5
2024 Center Category Focusing Transformer Network for Hyperspectral Image Classification
abstract
Recently, the methods based on self-attention mechanisms have gained increasing prominence in hyperspectral image classification (HSIC). However, the existing self-attention mechanism suffers the challenge of attention shift and redundancy. To address the problem, we propose the center category focusing transformer network (CCSF-Transformer) for HSIC, which is designed to resolve attention shifts and redundancy by balancing the multiple category features. Specifically, the central-category-focused attention mechanism (CFA) is presented in the proposed framework to compute the category-matched attention between the center pixel and neighbor pixels, closely matching the center-pixel style labeling strategy, and reducing the computation complexity by excluding the computation between interference pixels. Besides, the spectral-salient-focused attention module (SFA) is developed to capture the spectral correlation, which concentrates on the salient bands and suppresses the expression of redundant bands. Moreover, the hierarchical integration network (HIN) is built to rectify spatial and spectral features The experiment results on two popular HSI datasets demonstrate that the proposed method achieves robust performance compared to other state-of-the-art methods.
Yuanchen Zhu, Chunyan Yu, Meiping Song, Yulei Wang 0002, Enyu Zhao, Haoyang Yu 0001, Qiang Zhang 0011
IGARSS7
2024 Forest Early Wildfire Detection via Multi-Source and Multi-Type Information Fusion
abstract
In this work, we use the near real-time data of Himawari-8/9 geostationary satellite for the forest early fire detection at the minute level. By using the recursive Transformer model, this work comprehensively considers the temporal, spatial and spectral characteristics of fire point pixels, and takes into account MODIS land cover product information. The method reduces the interference factors such as cloud and terrain, and realizes the 10-mintue detection of forest early wildfire.
Qiang Zhang 0011, Yaming Zheng
IGARSS2
2024 Frequency-Temporal Attention Network for Remote Sensing Imagery Change Detection
abstract
Change detection (CD) in remote sensing imagery is identified as a pivotal task in the field of Earth observation, while it usually confronts the dilemma of intricate data and minor alterations. To address the stated challenge, this letter presents an innovative frequency-temporal attention network for CD (FTAN), which incorporates two advanced modules including the multidimensional convolutional frequency attention module (MCFA) and the interactive attention module (IAM). Specifically, the MCFA module is essential for enhancing sensitivity in CD by merging multiscale spatial and frequency domain features. As a supplement to MCFA, the IAM aggregates category-related tokens and processes cross-attention information from different time phases. The seamless integration of MCFA and IAM empowers the FTAN network with enhanced capabilities to detect minor regions and edges accurately. Experiments on datasets like LEVIR-CD and DSIFN-CD demonstrate superior performance by outperforming existing models in F1 scores and IoU metrics. Our code and pretrained models will be released athttps://github.com/chirsycy/FTAN.
Chunyan Yu, Yabin Hu, Qiang Zhang 0011, Meiping Song, Yulei Wang 0002
IEEE Geosci. Remote. Sens. Lett.4
2024 Dual-Intervention-Constrained Mask-Adversary Framework for Unsupervised Domain Adaptation of Hyperspectral Image Classification
abstract
To mitigate the domain shift and enhance the alignment of the spatial-spectral features, this letter proposes a novel dual-intervention-constrained mask-adversary (DICMA) framework for unsupervised domain adaptation (UDA) of hyperspectral image classification (HSIC). Innovatively, DICMA integrates a generator, masker, and bi-classifier within an adversarial framework constrained by a dual intervention mechanism. Specifically, the correlation intervention module (CIM) ensures the preservation and independence of causal spatial-spectral variables, while the knowledge distillation intervention module completes the spatial-spectral generalization with constrained distillation information. Besides, with the collaborative adversarial training strategy, the proposed approach transfers effective knowledge for spatial-spectral feature alignment. Experimental results and analyses demonstrate the effectiveness of the proposed DICMA model, which yields an accuracy of 91.15% in the Pavia University (PaviaU)$\to $Pavia Center (PaviaC). Our code will be released athttps://github.com/Chirsycy/DICMA.
Chunyan Yu, Mingyang Xu, Qiang Zhang 0011, Xiaoqiang Lu
IEEE Geosci. Remote. Sens. Lett.3
2024 Local Extremum Constrained Total Variation Model for Natural and Hyperspectral Image Non-Blind Deblurring
abstract
Blurring and noise degrade the performance of image processing. To mitigate this effect, various regularization-based deblurring methods have been proposed. Total variation regularization is widely used owing to its excellent ability in preserving the salient edges, but it also tends to smooth the image details. In this paper, we propose a local extremum-constrained total variation (LECTV) framework for image deblurring. In the developed deblurring framework, we integrate prior knowledge of the dark channel with the structural features of the image into a single regularization term. Furthermore, unlike most existing methods that focus on the overall sparsity of the dark channel, the defined regularization term allows for a pixel-wise adaptive description of the image to restore its inherent spatial texture structure. Finally, a majorization-minimization-based method is designed to solve the developed LECTV framework. Experimental results on natural and hyperspectral images show that the designed framework exhibits excellent performance in removing multiple types and degrees of blurring. Extensive evaluations also further show its superiority compared to other advanced methods.
Lan Li 0005, Meiping Song, Qiang Zhang 0011, Yushuai Dong, Yulei Wang 0002, Qiangqiang Yuan
IEEE Trans. Circuits Syst. Video Technol.3
2024 Feedback Band Group and Variation Low-Rank Sparse Model for Hyperspectral Image Anomaly Detection
abstract
For scenes with complex backgrounds and weak anomalies, how to effectively distinguish anomaly targets from the background is the key to perform hyperspectral image anomaly detection (AD). Data decomposition-based methods have been widely studied due to their potential in separating background and anomaly components. However, due to its unclean background extraction and sensitivity to noise, it has an adverse effect on the detection of anomaly targets. Additionally, a large amount of spectral data can lead to an increase in computation during data decomposition. To address this issue, we propose an AD method based on a feedback band group and variation low-rank sparse model (FBGVLRS-AD). Firstly, we employ a uniform band selection strategy to partition spectral bands and perform data decomposition on the selected band group, to separate low-rank and sparse components. This decomposition on the band group can reduce computational time and mitigate the interference from spectral variability. Secondly, to preserve the integrity of abnormal target spectra during the background extraction process, theL2,1norm is employed for joint correlated total variation to extract the desired anomalous targets. Then, utilizing the detection information from the existing band groups, a feedback-driven iterative framework has been designed to consider the consistency and complementarity in AD across band groups. This framework facilitates the extraction of sparse components in subsequent band groups and reinforces the anomalous elements. Iteratively addressing these sub-problems on band groups helps prevent the loss of useful spectral information, maintaining sufficient anomaly information while reducing interference from redundant information and spectral variations. Finally, the proposed FBGVLR-AD is optimally solved by the augmented Lagrange multiplier (ALM) method. Comparison with state-of-the-art anomaly detectors on multiple data validates the competitiveness of the proposed method for AD tasks.
Lan Li 0005, Qiang Zhang 0011, Meiping Song, Chein-I Chang
IEEE Trans. Geosci. Remote. Sens.2
2024 Unseen Feature Extraction: Spatial Mapping Expansion With Spectral Compression Network for Hyperspectral Image Classification
abstract
Hyperspectral image classification (HSIC) models have made remarkable progress in the last decade. Nevertheless, the downsized mapping in the convolutional neural network (CNN) and down-sampled mechanism in the transformer-based approach amplify the loss of hidden knowledge in the subpixel that encompasses crucial yet unseen features within a single pixel. Considering this aspect, the mentioned popular solutions for HSIC contradict the inherent characteristic of hyperspectral data. To address this issue, we rethink the size factor in CNN and propose a novel spatial mapping expansion with spectral compression (SMESC) network for HSIC. Specifically, the SMESC builds a mapping expansion network to mine unseen information in subpixels with enlarged feature maps. A channel modulation residual block (CMRB) is developed to compress spectral redundancy and promote salient channels with modulation information. Moreover, we design a multiple-size training strategy to substitute the traditional multiple feature extraction (FE) branches and improve the model adaptation to the different sizes of the testing samples. The extensive experimental results and analysis of four hyperspectral image (HSI) datasets demonstrate the superiority of the proposed architecture compared to other advanced HSIC methods. Our code will be released athttps://github.com/Chirsycy/SMESC.
Chunyan Yu, Yuanchen Zhu, Meiping Song, Yulei Wang 0002, Qiang Zhang 0011
IEEE Trans. Geosci. Remote. Sens.5
2024 Three-Dimension Spatial-Spectral Attention Transformer for Hyperspectral Image Denoising
abstract
Hyperspectral image (HSI) denoising is a crucial step for its subsequent applications. In this article, we propose TDSAT, a 3-D spatial-spectral attention Transformer model designed to effectively remove noise in HSI processing while preserving essential spectral and spatial information. The primary objective of this model is to utilize the 3-D Transformer to explore the global spectral-spatial features in HSI, learn the relationships among different bands, and preserve high-quality spectral and spatial information for denoising. The proposed method consists of three main components: the multihead spectral attention (MHSA) module, the gated-dconv feedforward network (GDFN) module, and the spectral enhancement (SpeE) module. The MHSA module learns the relationships among different bands and emphasizes the local spatial information. The GDFN module explores more expressive and discriminative spectral features. The SpeE module enhances the perception of subtle differences between different spectrums. Moreover, unlike the previous Transformer denoising method that can only handle fixed bands, the proposed method combines 3-D convolution and spectral-spatial attention Transformer blocks, enabling the denoising of HSI with an arbitrary number of bands. Experimental results demonstrate that TDSAT outperforms compared methods. The code is available athttps://github.com/Featherrain/TDSAT.
Qiang Zhang 0011, Yushuai Dong, Yaming Zheng, Haoyang Yu 0001, Meiping Song, Lifu Zhang 0002, Qiangqiang Yuan
IEEE Trans. Geosci. Remote. Sens.1
2024 GACNet: Generate Adversarial-Driven Cross-Aware Network for Hyperspectral Wheat Variety Identification
abstract
Wheat variety identification from hyperspectral images holds significant importance in both fine breeding and intelligent agriculture. However, the discriminatory accuracy of some techniques is limited due to insufficient datasets, data redundancy, and noise interference. To address these issues, we propose a wheat variety identification framework called generate adversarial-driven cross-aware network (GACNet), comprising a semi-supervised generative adversarial network (GAN) for data augmentation and a cross-aware attention network (CAANet) for variety identification. First, the semi-supervised GAN (SSGAN) alleviates data scarcity by generating fake hyperspectral images as realistically as possible through learning the distribution hypothesis of real hyperspectral images, while the discriminator distinguishes between real and fake hyperspectral images. Subsequently, the CAANet is employed for wheat variety identification, which leverages a cascading cross-learning of 3-D and 2-D convolutions to fully exploit spectral, spatial, and texture features and refines the features through an embedded attention mechanism in the cross-convolutional module. Additionally, we constructed a hyperspectral wheat variety dataset (HWVD) comprising 4560 samples of 19 categories. Extensive experiments on our dataset demonstrate that our GACNet outperforms state-of-the-art methods for wheat variety identification. The HWVD will be made available.
Weidong Zhang 0007, Guohou Li, Peixian Zhuang, Guojia Hou, Qiang Zhang 0011, Chongyi Li
IEEE Trans. Geosci. Remote. Sens.6
2024 Thermal Infrared Hyperspectral Band Selection via Graph Neural Network for Land Surface Temperature Retrieval
abstract
Thermal infrared hyperspectral imagery presents a superior capability for capturing intricate spectral details of atmospheres and ground objects compared to multispectral images, thus offering a more nuanced dataset for land surface temperature (LST) retrieval. However, extensive inter-band correlations pose computational challenges and undesirable “dimension disaster” problem. To address this issue, this paper proposes a purpose-built framework of thermal infrared hyperspectral band selection using graph neural network for LST retrieval. Specifically, the thermal infrared hyperspectral data is firstly mapped onto a graph topology, followed by feeding it into a graph attention module with brightness temperature constraints to extract band features. Following this, the extracted band features undergo a comprehensive analysis through a multi-scale convolution module consisting of convolution kernels with multiple sizes, which has more variety and larger receptive fields for calculating the correlation between different bands features, assigning different weights to each band. Finally, a weight selection module is designed to filter the bands based on their assigned weights, creating a subset of bands with greater significance for LST retrieval. Training the designed model, 65100 observations are simulated utilizing MODTRAN, 80% allocated for training and 20% for testing. The experimental results validate the effectiveness of the proposed model, with a Root Mean Square Error (RMSE) of 1.85 K in practical applications on IASI imagery. This accomplishment substantiates the model’s capacity to reliably employ a judiciously selected subset of thermal infrared hyperspectral bands for LST retrieval applications, thus offering a promising contribution to the advancement of thermal infrared hyperspectral image processing methodologies.
Enyu Zhao, Nianxin Qu, Yulei Wang 0002, Caixia Gao, Sibo Duan, Qiang Zhang 0011
IEEE Trans. Geosci. Remote. Sens.7
2024 Hyperspectral Image Denoising: From Model-Driven, Data-Driven, to Model-Data-Driven
abstract
Mixed noise pollution in HSI severely disturbs subsequent interpretations and applications. In this technical review, we first give the noise analysis in different noisy HSIs and conclude crucial points for programming HSI denoising algorithms. Then, a general HSI restoration model is formulated for optimization. Later, we comprehensively review existing HSI denoising methods, from model-driven strategy (nonlocal mean, total variation, sparse representation, low-rank matrix approximation, and low-rank tensor factorization), data-driven strategy [2-D convolutional neural network (CNN), 3-D CNN, hybrid, and unsupervised networks], to model-data-driven strategy. The advantages and disadvantages of each strategy for HSI denoising are summarized and contrasted. Behind this, we present an evaluation of the HSI denoising methods for various noisy HSIs in simulated and real experiments. The classification results of denoised HSIs and execution efficiency are depicted through these HSI denoising methods. Finally, prospects of future HSI denoising methods are listed in this technical review to guide the ongoing road for HSI denoising. The HSI denoising dataset could be found at https://qzhang95.github.io.
Qiang Zhang 0011, Yaming Zheng, Qiangqiang Yuan, Meiping Song, Haoyang Yu 0001, Yi Xiao 0003
IEEE Trans. Neural Networks Learn. Syst.1
2023 A Swin Transformer-Based Fusion Approach for Hyperspectral Image Super-Resolution
abstract
Hyperspectral image (HSI) has attracted much attention because of its rich spectral information. However, due to the limitation of imaging hardware conditions, it is often difficult to directly obtain a high spatial resolution hyperspectral image (HR-HSI). To improve the resolution, it is an economical and effective method to fuse the hyperspectral image with the high spatial resolution multispectral image (HR-MSI) collected from the same scene. In recent years, with the development of deep learning, the convolutional neural network (CNN) based models have been applied to solve the super-resolution reconstruction of hyperspectral images. However, limited by the convolution kernel size, the receptive field of CNN is relatively small with more attention to the local information of the image. In order to solve this problem, this paper proposes a Swin Transformer based super-resolution reconstruction (STSR) network for hyperspectral images. Specifically, Swin Transformer structure is innovatively used in STSR as the skeleton of the network, where the Swin Transformer residuals are used to extract the global spatial feature information in the image. In addition, in order to retain the spectral details in the process of super-resolution reconstruction, a spectral attention module is introduced to preserve the original spectral information. The experimental results show that the high-resolution hyperspectral images fused by the proposed STSR method are superior to the comparison method in terms of vision and quality, which proves the superiority of this method.
Yulei Wang 0002, Enyu Zhao, Meiping Song, Qiang Zhang 0011
IGARSS5
2023 Hyperspectral Image Classification Based on Interactive Transformer and CNN With Multilevel Feature Fusion Network
abstract
Due to the powerful feature information mining ability of deep learning, models such as Convolutional Neural Network (CNN) and Transformer have gained a certain progress in hyperspectral image classification (HSIC). Characteristically, the CNN is good at extracting local information, but it has the limitation of insufficient receptive field. While the Transformer has the advantage of global representation, it ignores local details to some extent. Therefore, this letter proposes an interactive Transformer and CNN with multilevel feature fusion network (ITCNet) for HSIC. Specifically, in the image-based framework, features with different perceptual fields and depths are extracted interactively by a multi-layer Transformer and CNN, then fused through a multilevel feature fusion module for class prediction. Experimental results on two real datasets verifies its efficiency, with improvements over other related methods.
Haoyang Yu 0001, Jiaochan Hu, Tingting Tao, Qiang Zhang 0011
IEEE Geosci. Remote. Sens. Lett.6
2023 Combined Deep Priors With Low-Rank Tensor Factorization for Hyperspectral Image Restoration
abstract
Mixed noise pollution severely disturbs hyperspectral image (HSI) processing and applications. Plenty of algorithms have been developed to address this issue via two strategies: model-driven or data-driven strategy. However, model-driven methods exist in the highly time-consuming weakness of iterative optimization and unstable sensitivity of setting parameters. Data-driven methods usually perform poor due to the overfitting effects. To solve these issues, we combine both the deep denoising priors with low-rank tensor factorization (DP-LRTF) for HSI restoration. The proposed method uses Tucker tensor factorization to depict the global spectral low-rank constraint. Then the spectral orthogonal basis and spatial reduced factor are optimized by two deep denoising priors, respectively. Through this integrated strategy, we can simultaneously exploit the intrinsic low-rank property of HSI, and utilize the powerful feature extraction ability by deep learning for HSI restoration. Compared with model-driven and data-driven methods, DP-LRTF outperforms on HSI mixed noise removal and execution efficiency for various simulated/real experiments.
Qiang Zhang 0011, Yushuai Dong, Qiangqiang Yuan, Meiping Song, Haoyang Yu 0001
IEEE Geosci. Remote. Sens. Lett.1
2023 Hyperspectral Denoising via Global Variation and Local Structure Low-Rank Model
abstract
Hyperspectral images (HSIs) are often disturbed by various kinds of noises. This paper proposes a global variation and local structure low-rank model (GLLR) for HSI denoising by integrating spatial segmentation smoothing and spectral low-rank properties. Compared with existing denoising methods, the proposed method considers not only the global low-rank property but also the local structure low-rank property of HSIs. Specifically, the GLLR describes the global correlation and segmental smoothing structure of the HSI by the correlated total variation. In addition, we construct a new structural low-rank prior, called the local minimum difference (LMD) low-rank. With LMD low-rank property of HSI, GLLR can remove noise while retaining useful structural information in the HSI. Then, an ALM-based optimization algorithm is devised to solve the objective functions for the presented model. Finally, comparison experiments with existing methods are conducted on synthetic and real datasets to demonstrate the effectiveness and superiority of the proposed method.
Lan Li 0005, Meiping Song, Qiang Zhang 0011, Yushuai Dong
IEEE Trans. Geosci. Remote. Sens.3
2023 Deep Blind Super-Resolution for Satellite Video
abstract
Recent efforts have witnessed remarkable progress in Satellite Video Super-Resolution (SVSR). However, most SVSR methods usually assume the degradation is fixed and known,e.g., bicubicdownsampling, which makes them vulnerable in real-world scenes with multiple and unknown degradations. To alleviate this issue, blind SR has thus become a research hotspot. Nevertheless, existing approaches are mainly engaged in blur kernel estimation while losing sight of another critical aspect for VSR tasks: temporal compensation, especially compensating for blurry and smooth pixels with vital sharpness from severely degraded satellite videos. Therefore, this paper proposes a practical Blind SVSR algorithm (BSVSR) to explore more sharp cues by considering the pixel-wise blur levels in a coarse-to-fine manner. Specifically, we employed multi-scale deformable convolution to coarsely aggregate the temporal redundancy into adjacent frames by window-slid progressive fusion. Then the adjacent features are finely merged into mid-feature using deformable attention, which measures the blur levels of pixels and assigns more weights to the informative pixels, thus inspiring the representation of sharpness. Moreover, we devise a pyramid spatial transformation module to adjust the solution space of sharp mid-feature, resulting in flexible feature adaptation in multi-level domains. Quantitative and qualitative evaluations on both simulated and real-world satellite videos demonstrate that our BSVSR performs favorably against state-of-the-art non-blind and blind SR models. Code will be available at https://github.com/XY-boy/Blind-Satellite-VSR.
Yi Xiao 0003, Qiangqiang Yuan, Qiang Zhang 0011, Liangpei Zhang 0001
IEEE Trans. Geosci. Remote. Sens.3
2022 SSTNet: Spatial, Spectral, and Texture Aware Attention Network Using Hyperspectral Image for Corn Variety Identification
abstract
Currently, most existing methods using hyperspectral image to assist seed identification only consider the spectral information but ignore the spatial information resulting in unsatisfactory classification results. To cope with this issue, we propose a spatial, spectral, and texture-aware attention network to identify corn varieties, called SSTNet. Specifically, we first employ 3D convolution to extract the spatial and inter-spectral features. Subsequently, we utilize 2D convolution to extract the spatial and texture features. Meanwhile, we embed an attention mechanism into the 2D convolution module to further refine the spatial and texture features. The advantageous complementary properties of 3D and 2D convolutions allow the spatial and textural features of hyperspectral images to be fully exploited. Besides, we construct a hyperspectral image dataset including 1200 samples of 10 corn varieties. Experiments on our proposed dataset demonstrate that our SSTNet outperforms the state-of-the-art methods for identifying corn varieties.
Weidong Zhang 0007, Hai-Han Sun, Qiang Zhang 0011, Peixian Zhuang, Chongyi Li
IEEE Geosci. Remote. Sens. Lett.4
2022 Robust Thick Cloud Removal for Multitemporal Remote Sensing Images Using Coupled Tensor Factorization
abstract
The existing nonblind cloud and cloud shadow (cloud/shadow) removal methods for remote sensing (RS) images are based on the assumption that cloud/shadow masks are accurately given. Since the masks are usually manually labeled or detected by cloud detection methods, whose accuracy cannot be well guaranteed, the cloud/shadow removal effect may be affected. In this article, we suggest a robust thick cloud/shadow removal (RTCR) method that meets the problem with an inaccurate mask. To faithfully reconstruct the multitemporal information, a coupled tensor factorization is used to explore the relationship between the abundances of the multitemporal images in the same scene. Moreover, an efficient algorithm is developed to solve the proposed model based on the augmented Lagrange multiplier method. The experimental results under accurate masks and inaccurate masks demonstrate its robustness and superiority for thick cloud/shadow removal.
Jie Lin 0011, Ting-Zhu Huang, Xi-Le Zhao, Yong Chen 0013, Qiang Zhang 0011, Qiangqiang Yuan
IEEE Trans. Geosci. Remote. Sens.5
2022 Cooperated Spectral Low-Rankness Prior and Deep Spatial Prior for HSI Unsupervised Denoising
abstract
Model-driven methods and data-driven methods have been widely developed for hyperspectral image (HSI) denoising. However, there are pros and cons in both model-driven and data-driven methods. To address this issue, we develop a self-supervised HSI denoising method via integrating model-driven with data-driven strategy. The proposed framework simultaneously cooperates the spectral low-rankness prior and deep spatial prior (SLRP-DSP) for HSI self-supervised denoising. SLRP-DSP introduces the Tucker factorization via orthogonal basis and reduced factor, to capture the global spectral low-rankness prior in HSI. Besides, SLRP-DSP adopts a self-supervised way to learn the deep spatial prior. The proposed method doesn't need a large number of clean HSIs as the label samples. Through the self-supervised learning, SLRP-DSP can adaptively adjust the deep spatial prior from self-spatial information for reduced spatial factor denoising. An alternating iterative optimization framework is developed to exploit the internal low-rankness prior of third-order tensors and the spatial feature extraction capacity of convolutional neural network. Compared with both existing model-driven methods and data-driven methods, experimental results manifest that the proposed SLRP-DSP outperforms on mixed noise removal in different noisy HSIs.
Qiang Zhang 0011, Qiangqiang Yuan, Meiping Song, Haoyang Yu 0001, Liangpei Zhang 0001
IEEE Trans. Image Process.1
2021 Thick Cloud Removal for Sentinel-2 Time-Series Images via Combining Deep Prior and Low-Rank Tensor Completion
abstract
In this study, we combine both the deep prior with low-rank tensor completion (DP-LRTC) for thick cloud removal in Sentinel-2 time-series images. On the one hand, DP-LRTC utilizes the low-rank property of multitemporal images via 3-order tensor completion. On the other hand, DP-LRTC employs the 3D spatiotemporal feature expression ability by deep learning. Through integrating both model-driven with data-driven strategy, the proposed method can effectively removal thick cloud in Sentinel-2 time-series images.
Qiang Zhang 0011, Fujun Sun, Qiangqiang Yuan, Liangpei Zhang 0001
IGARSS1
2020 Combined the Data-Driven with Model-Driven Stragegy: A Novel Framework for Mixed Noise Removal in Hyperspectral Image
abstract
In this paper, we present a novel hyperspectral image (HSI) denoising method especially for mixed noise removal. The proposed method combines both data-driven with model-driven strategy via a deep spatio-spectral variational structure. The mixed noise estimation and removal are collaboratively derived through fusing the Bayesian spatio-spectral posterior and deep learning model. The framework can both utilize the logicality of traditional model-driven methods, and the high efficiency of data-driven methods for parameters optimizing. Simulated and actual experiments demonstrate that the presented method outperforms other existing methods for HSI mixed noise removal, on both reconstructing effects and time-consuming.
Qiang Zhang 0011, Fujun Sun, Qiangqiang Yuan, Jie Li 0022, Huanfeng Shen, Liangpei Zhang 0001
IGARSS1
2019 Cloud and Shadow Removal for Sentinel-2 by Progressively Spatiotemporal Patch Group Learning
abstract
In this work, a progressively spatio-temporal patch group learning framework for cloud and shadow removal in Sentinel-2 data is proposed. Through sorting the spatial and corresponding multi-temporal patches with masks as the patch group fashion, a spatiotemporal patch group recovering model is developed using a global-local deep CNN. Finally, all the ergodic patches are weighted aggregated with integrity measure, then updated spatial data and its mask are regenerated through progressive iteration. Two experiments have been performed to demonstrate the effectiveness of the proposed method on Sentienl-2 MSI data, with single/multiple temporal imageries in small and largescale scenarios.
Qiang Zhang 0011, Qiangqiang Yuan, Jie Li 0022, Huanfeng Shen, Liangpei Zhang 0001
IGARSS1
2019 Hyperspectral Image Denoising Employing a Spatial-Spectral Deep Residual Convolutional Neural Network
abstract
Hyperspectral image (HSI) denoising is a crucial preprocessing procedure to improve the performance of the subsequent HSI interpretation and applications. In this paper, a novel deep learning-based method for this task is proposed, by learning a nonlinear end-to-end mapping between the noisy and clean HSIs with a combined spatial-spectral deep convolutional neural network (HSID-CNN). Both the spatial and spectral information are simultaneously assigned to the proposed network. In addition, multiscale feature extraction and multilevel feature representation are, respectively, employed to capture both the multiscale spatial-spectral feature and fuse different feature representations for the final restoration. The simulated and real-data experiments demonstrate that the proposed HSID-CNN outperforms many of the mainstream methods in both the quantitative evaluation indexes, visual effects, and HSI classification accuracy.
Qiangqiang Yuan, Qiang Zhang 0011, Jie Li 0022, Huanfeng Shen, Liangpei Zhang 0001
IEEE Trans. Geosci. Remote. Sens.2
2019 Hybrid Noise Removal in Hyperspectral Imagery With a Spatial-Spectral Gradient Network
abstract
The existence of hybrid noise in hyperspectral images (HSIs) severely degrades the data quality, reduces the interpretation accuracy of HSIs, and restricts the subsequent HSI applications. In this paper, the spatial-spectral gradient network (SSGN) is presented for mixed noise removal in HSIs. The proposed method employs a spatial-spectral gradient learning strategy, in consideration of the unique spatial structure directionality of sparse noise and spectral differences with additional complementary information for effectively extracting intrinsic and deep features of HSIs. Based on a fully cascaded multiscale convolutional network, SSGN can simultaneously deal with different types of noise in different HSIs or spectra by the use of the same model. The simulated and real-data experiments undertaken in this study confirmed that the proposed SSGN outperforms at mixed noise removal compared with the other state-of-the-art HSI denoising algorithms, in evaluation indices, visual assessments, and time consumption.
Qiang Zhang 0011, Qiangqiang Yuan, Jie Li 0022, Xinxin Liu 0002, Huanfeng Shen, Liangpei Zhang 0001
IEEE Trans. Geosci. Remote. Sens.1
2018 A Unified Spatial-Temporal-Spectral Learning Framework for Reconstructing Missing Data in Remote Sensing Images
abstract
In this paper, a unified spatial-temporal-spectral framework of missing information reconstruction in remote sensing images is proposed. Based on an end-to-end non-linear mapping structure, the proposed method employs a unified deep convolutional neural network combined with joint spatial-temporal-spectral supplementary information. It should be noted that the proposed model can use multi-source data (spatial, spectral, and temporal) as the input of the unified framework. The results of real-data experiments demonstrate that the proposed model exhibits high effectiveness in missing information reconstruction tasks like dead lines in Aqua MODIS band 6, Landsat ETM+ SLC-off and thick cloud removal.
Qiang Zhang 0011, Qiangqiang Yuan, Huanfeng Shen, Liangpei Zhang 0001
IGARSS1
2018 Missing Data Reconstruction in Remote Sensing Image With a Unified Spatial-Temporal-Spectral Deep Convolutional Neural Network
abstract
Because of the internal malfunction of satellite sensors and poor atmospheric conditions such as thick cloud, the acquired remote sensing data often suffer from missing information, i.e., the data usability is greatly reduced. In this paper, a novel method of missing information reconstruction in remote sensing images is proposed. The unified spatial-temporal-spectral framework based on a deep convolutional neural network (CNN) employs a unified deep CNN combined with spatial-temporal-spectral supplementary information. In addition, to address the fact that most methods can only deal with a single missing information reconstruction task, the proposed approach can solve three typical missing information reconstruction tasks: (1) dead lines in Aqua Moderate Resolution Imaging Spectroradiometer band 6; (2) the Landsat Enhanced Thematic Mapper Plus scan line corrector-off problem; and (3) thick cloud removal. It should be noted that the proposed model can use multisource data (spatial, spectral, and temporal) as the input of the unified framework. The results of both simulated and real-data experiments demonstrate that the proposed model exhibits high effectiveness in the three missing information reconstruction tasks listed above.
Qiang Zhang 0011, Qiangqiang Yuan, Chao Zeng 0001, Xinghua Li 0002, Yancong Wei
IEEE Trans. Geosci. Remote. Sens.1