Shou Feng

dblp:212/2512 · DBLP profile ↗
← Back
58ranked-venue papers
13as first author
55since 2021 · last 2026
0000-0002-7308-9590ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 48 · 11 first-author · 46 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021
YearPublicationVenuePosition
2026 HyperPRET: Few-shot class incremental learning with precognition and retrospection for hyperspectral imagery
Bobo Xi, Tie Zheng, Jiaojiao Li 0001, Shou Feng, Yunsong Li 0001
Pattern Recognit.5
2025 Cross-Domain Few-Shot Learning Method Based on Fractional Domain Information for Hyperspectral Image Multi-Class Change Detection
abstract
Hyperspectral image multi-class change detection (HSI-MCD) based on deep learning (DL) rely significantly on the number of labeled data. Due to the high cost of manually labeling for hyperspectral images (HSIs), obtaining a large amount of labeled samples is difficult. Moreover, for multi-class change detection (MCD) tasks, there is the phenomenon of semantic cross-coupling of changes due to complex change scenarios. To solve the above problems, a cross-domain few-shot learning method based on fractional domain information for HSI-MCD (FrCFSL) is proposed. Firstly, a spectral-spatial-fractional information extraction module is proposed, which can extract spectral-spatial-fractional domain joint feature. Thus, the module can obtain more comprehensive and discriminative representations of land cover categories, alleviating the phenomenon of semantic cross-coupling between classes. Afterward, a cross-domain fewshot learning strategy is introduced, where it learns task-relevant category discrimination meta-knowledge from a pair of richly labeled very high-resolution optical images (VHRIs) dataset and transfers it to the bitemporal HSIs dataset. Thus, the model can achieve better MCD performance with a small number of labeled samples. Finally, to mitigate the domain distribution differences between VHRIs data and HSIs data, a topological structure alignment module is proposed to align the intrinsic topological relationships between land cover categories, thus narrowing the gap between the two domain distributions. Through experiments conducted on three HSI-MCD datasets and comparative analysis with six state-of-the-art methods, the validity and stability of the proposed method are indicated.
Shou Feng, Jinghe Zhang, Yuanze Fan, Xinyao Liu, Chunhui Zhao 0003, Wei Li 0032, Ran Tao 0003
IEEE Trans. Circuits Syst. Video Technol.1
2025 A Prototype-Aware Learning and Dual-View Regularization Network for Weakly Supervised Change Detection in VHR Remote Sensing Images
abstract
Change detection (CD) is a critical task for monitoring the spatiotemporal evolution of the Earth’s surface. Recently, due to the advantages of reduced annotation cost and improved labeling efficiency, weakly supervised change detection (WSCD) has attracted increasing attention. However, existing WSCD methods encounter several critical challenges, including incomplete activation of class activation maps (CAMs), interference from noisy pseudo-labels during training, and instability in change recognition caused by illumination and environmental variations. To address these issues, we propose a prototype-aware learning and dual-view regularization network (PDRNet) for image-level WSCD. Specifically, to address the issue of incomplete activation caused by the tendency of CAM to focus excessively on locally discriminative regions, PDRNet devises a prototype-aware module (PAM), which captures stable category prototypes and refines CAM quality by reactivating hierarchical features. Furthermore, to mitigate the network’s sensitivity to noisy pseudo-labels, a dual-view regularization strategy (DRS) is designed to partition pseudo-labels into clean and noisy regions. Region-specific regularization is subsequently employed to improve the robustness of the model against noisy supervision. Finally, to enhance the capability of identifying changed regions, PDRNet constructs a wavelet-based change enhancement module (WCEM) to decompose bi-temporal features into multiple frequency bands. This facilitates the comprehensive utilization of low-frequency structural semantics and high-frequency texture details. Extensive experiments and analyses conducted on three publicly available CD datasets yield the superiority of PDRNet.
Shou Feng, Chunhui Zhao 0003, Yingjie Tang, Wei Li 0032, Ran Tao 0003
IEEE Trans. Geosci. Remote. Sens.2
2025 Fractional Fourier-Enhanced Fusion Network Based on Pareto Optimization for Hyperspectral and LiDAR Data Classification
abstract
In recent years, the utilization of hyperspectral image (HSI) and light detection and ranging (LiDAR) for collaborative classification has emerged as a significant research direction in earth observation tasks, with diverse joint classification algorithms showing promising performance using varying network architectures. However, these methodologies infrequently address the challenge of fusion arising from the substantially larger volume of HSI feature information compared to LiDAR features. Moreover, the effective learning of HSI and LiDAR features while mitigating modality conflicts remains an area that necessitates further investigation. As such, a Fractional Fourier Enhanced Fusion Network based on Pareto Optimization (FrFENet) is proposed for HSI and LiDAR Data classification. To address the disparity in information volume between modalities, a weighted fractional Fourier enhanced fusion module (WFrFEF) is introduced, which applies a weighted fractional Fourier transform to HSI features, enhancing their representations and facilitating balanced fusion with LiDAR features. Furthermore, a Pareto-based soft optimization strategy, HLPareto, is designed to balance learning rates across HSI and LiDAR features in a dual-branch network, effectively avoiding optimization conflicts. Additionally, a spatial-spectral integration module (SSIM) and an elevation information enhancement module (EIEM) are developed to improve feature extraction. The SSIM enables effective spatial-spectral fusion by facilitating token-level interactions, while the EIEM enhances elevation feature representation, preserving spatial geometric information in LiDAR data. Extensive experiments and comparative analyses conducted on three widely utilized HSI and LiDAR datasets have shown that the proposed FrFENet exhibits superior classification performance.
Shou Feng, Hongtao Deng, Yabin Hu, Chunhui Zhao 0003, Wei Li 0032, Ran Tao 0003
IEEE Trans. Geosci. Remote. Sens.1
2025 Transformer-Based Cross-Domain Few-Shot Learning for Hyperspectral Target Detection
abstract
Deep learning-based methods have made significant progress in hyperspectral target detection (HTD). Unfortunately, limited target prior information and imbalance class resulting from the low occurrence probability of target leaves deep learning-based methods to confront bottlenecks. To ameliorate the abovementioned issues, a Transformer-based cross-domain few-shot learning (TCFSL) method is proposed for HTD. First, the TCFSL leverages cross-domain few-shot learning (FSL) to establish FSL tasks in both the source domain (SD) and the target domain (TD). This allows the TCFSL to learn transferable knowledge of the SD and distinguishable feature embedding model for the TD, to address the problems of target priori lacking and imbalance class. Second, feature-level and distribution-level domain adaptation (DA) is used to tackle the problem of domain shift in cross-domain FSL. The feature-level DA extracts intradomain information of the SD and TD to learn their common features to alleviate domain shift. The distribution-level DA based on cross-Transformer present interdomain distribution-level information aggregation and captures domain similarities of two data domains. By pursuing similarities between two data domains, the distribution-level DA block prompts specific FSL tasks in each domain, facilitating the target detection task. Finally, cross-domain FSL and DA blocks are trained in a unitary manner, which facilitates real-time information interaction and parameter adjustment between different blocks to achieve the optimal model. Experiments conducted on six HSI datasets indicate that the TCFSL outperforms 12 compared methods.
Shou Feng, Fengchao Xiong, Chunhui Zhao 0003, Wei Li 0032, Ran Tao 0003
IEEE Trans. Geosci. Remote. Sens.1
2025 Fractional-Domain Information-Enhanced Hyperspherical Prototype Learning Method for Hyperspectral Image Open-Set Classification
abstract
In recent years, research in the field of hyperspectral image classification (HSIC) has increasingly focused on the open-set problem. Open-set classification demands not only accurately classifying the known categories but also identifying the unknown samples that are not labeled or included within the training data during testing stage. Existing open-set methods often suffer from misclassification between the known and unknown categories due to their inadequate utilization of metric space. Moreover, relying on a single threshold strategy performs poorly for identifying unknown categories in complex open environments. In this paper, a fractional domain information enhanced hyperspherical proto-type learning method (FrHSPL) is proposed for hyperspectral image open-set classification. FrHSPL develops a hyperspherical prototype learning (HSPL) strategy that ensures the features of known categories are uniformly distributed on the hypersphere. Therefore, HSPL can effectively enhance inter-class separability and optimize the exploitation of metric space. Subsequently, to enhance the discrimination capability of spectral features, a frequency-spatial-spectral information aggregation module is devised to deeply integrate fractional domain information with spatial and spectral information. Finally, an open-set recognition module is designed to identify unknown categories by using the prototypes of each known category along with the corresponding prototype radii. Extensive experiments on four common HSI datasets indicate that the proposed FrHSPL exhibits superior performance in comparison with both closed-set and open-set methods.
Shou Feng, Cong'an Xu, Chunhui Zhao 0003, Wei Li 0032, Ran Tao 0003
IEEE Trans. Geosci. Remote. Sens.1
2025 A Nonlinear Weighted Graph Convolution Network Based on Manifold Geometric Regularization for Hyperspectral Image Classification
abstract
Extracting spatial-spectral joint features has become a critical approach for improving model classification performance in the field of hyperspectral image classification (HSIC). However, existing methods fail to fully exploit nonlinear spatial-spectral information. Unlike traditional convolutional neural networks (CNNs), graph convolutional neural networks (GCNs) can extract nonlinear spatial information. Nevertheless, both methods lack an accurate measurement of local neighborhood information, leading to blurred classification boundaries for ground objects. Additionally, the high-dimensional nature of hyperspectral data results in poor generalization and redundant information of trained models. To address these three issues, a nonlinear weighted graph convolution network based on manifold geometric regularization (MGR-NWGCN) method is devised for HSIC. Specifically, a nonlinear weighted graph convolution (NWGCN) module is designed, which utilizes a Graph-in-Graph structure based on cosine similarity-based normalized weighted graph convolution to extract nonlinear spatial-spectral information. Then, the manifold curvature regularization (C-MGR) module is implemented to improve the accuracy of similarity measurement and to enhance the generalization ability of the model, which constrains the model to form flatter feature manifold surfaces. Finally, the manifold intrinsic dimensionality regularization (ID-MGR) module is developed with the aim of eliminating redundant information, which embeds noise onto the surface of a low-dimensional manifold. The superior classification performance and robustness of the proposed MGR-NWGCN method are validated through extensive experiments on four datasets, with comparisons conducted against nine methods.
Shou Feng, Cong'an Xu, Chunhui Zhao 0003, Wei Li 0032, Ran Tao 0003
IEEE Trans. Geosci. Remote. Sens.1
2025 DSNet: Dynamic Stitchable Neural Network for Hyperspectral Image Classification
abstract
Hyperspectral image classification (HSIC) aims to identify land cover categories by leveraging the spectral and spatial information contained in hyperspectral images (HSI). Currently, many deep learning approaches utilize dual-branch networks to process spectral and spatial data separately, followed by the application of specialized modules to facilitate feature interaction or fusion. However, the design of these modules demands considerable time and effort from researchers and may not adequately capture the inherent relationships between independent spatial and spectral features in a dynamic manner. To address these issues, we propose the dynamic stitchable neural network (DSNet) for HSIC. While the DSNet maintains a dual-branch structure, it operates without traditional feature fusion or interaction. Instead, it employs a stitching network approach to integrate the two branches. Specifically, a spatial-spectral stitching module is presented to incorporates multiple stitching layers at various positions between the two network branches, creating new stitched networks that retain the strengths of both original networks. Additionally, a reinforcement learning-based strategy is designed for dynamically selecting stitching positions tailored to specific datasets, enabling the model to adaptively optimize the integration of spatial and spectral features. Recognizing the effectiveness of vision transformer (ViT) in learning spatial information and the capability of 1D convolutional neural network (1DCNN) in capturing spectral details, the DSNet directly stitches these two networks together. This fusion maximizes the utilization of both foundational networks, yielding a new hybrid network that delivers exceptional performance while also alleviating the burden on researchers to develop new architectures from scratch. Extensive experiments and analyses conducted on three public HSI datasets demonstrate the superiority of the proposed method, validating the effectiveness of our innovative modules. The codes of this work will be available from the website: https://github.com/ZZC/IEEE-TGRS-DSNet.
Shou Feng, Zicheng Zhao, Bobo Xi, Chunhui Zhao 0003, Wei Li 0032, Ran Tao 0003, Yunsong Li 0001
IEEE Trans. Geosci. Remote. Sens.1
2025 Collaborative Classification of Hyperspectral and LiDAR Date Based on Dynamic Multiple Fractional Fourier Domains Fusion
abstract
Collaboratively utilizing the complementary information provided by hyperspectral imagery and light detection and ranging (LiDAR) data will extend the applications associated with land cover recognition and mapping. Existing joint classification algorithms mainly focus on learning complementary patterns in the pure spatial domain, while paying little attention to complementary cues in the spatial-frequency domain. The model’s expressive capability of these methods may be limited by an upper bound subject to the spatial domain. To fill this gap, a Dynamic Multiple Fractional Fourier Domains Fusion (DMFraF) is proposed for joint classification of hyperspectral and LiDAR data. Firstly, to comprehensively learn the complementary patterns between HSI and LiDAR data, we transform the features of two modalities into multiple fractional domains containing different spatial-frequency components for multimodal fusion. Secondly, to obtain the optimal representation from the multimodal features of multiple fractional domains, we propose a dynamic fusion scheme guided by the optimal transport (OT) technique, which can dynamically adjust the contributions from different fractional domains. Finally, to extract purer modality-specific features, we propose a channel aggregation Transformer encoder with central cross-attention (C2AT encoder), to aggregate channel-wise features of central pixels into the spatial branch and compress interference from noisy surroundings. Extensive experiments and analysis on three hyperspectral and LiDAR datasets suggest the superiority of the proposed method.
Boao Qin, Shou Feng, Chunhui Zhao 0003, Wei Li 0032, Ran Tao 0003
IEEE Trans. Geosci. Remote. Sens.2
2025 Language-Enhanced Dual-Level Contrastive Learning Network for Open-Set Hyperspectral Image Classification
abstract
In recent years, language-supervised vision models have demonstrated impressive potential in learning open-world concepts. Some research has introduced this learning paradigm to the hyperspectral image (HSI) processing domain; however, there has been limited work integrating textual information into the hyperspectral open-set recognition task. To fill this gap, we leverage textual supervision information in open-set HSI classification (HSIC) and propose a language-enhanced dual-level contrastive learning network (LDCLNet). Specifically, we introduce a linguistic mode with prior knowledge as a supervised signal to enhance the metric distances between closed-set samples and provide supplementary semantic information for open-set samples. Second, a dual-level visual-language (V-L) contrastive learning (CL) approach, which can align visual and language embeddings separately at the instance level and manifold level, is proposed to establish a more accurate link between visual and language representations. Finally, a distance-refined open-set recognition method is proposed, which aims to effectively discover unknown class samples during testing by refining predictions of known and unknown classes. Extensive experiments and analysis on three public HSI datasets validate the effectiveness of LDCLNet.
Boao Qin, Shou Feng, Chunhui Zhao 0003, Wei Li 0032, Ran Tao 0003, Jun Zhou 0001
IEEE Trans. Geosci. Remote. Sens.2
2025 FCMMA: Fourier Conditional Mask-Based Mixed Attention Method for Hyperspectral Anomaly Detection
abstract
In recent years, reconstruction-based methods have achieved excellent detection results in the field of hyperspectral anomaly detection (HAD). These methods predominantly operate on two aspects regarding their working principles: 1) reconstructing background pixels and 2) suppressing anomalous pixels. However, most methods only tackle the HAD task from the spatial and spectral domains, making it challenging to effectively suppress anomalies. To eliminate these issues, this article proposes a Fourier conditional mask-based mixed attention (FCMMA) method. First, we propose the FCMMA method for HAD. FCMMA generates a conditional mask (CMASK) that suppresses anomalous high-frequency information and preserves background low-frequency information in the frequency domain, optimizing the anomaly detection process. In addition, to achieve fine-grained HAD, we propose the Fourier anomaly suppression filter (FASF). FASF uses Fourier techniques to manage background and anomalies, improving detection via precise frequency decoupling. Finally, a CMASK network is designed to effectively suppress anomalies. The CMASK network integrated the FASF module and the spatial-spectral multilayer perceptual (SSMLP) machine module together to enhance the transformation and representation capabilities of the generated masks, which can also help suppress anomalies. The results on five different datasets show that the proposed method is more effective and superior when compared to nine state-of-the-art methods.
Shou Feng, Nan Su 0001, Chunhui Zhao 0003, Wei Li 0032, Ran Tao 0003
IEEE Trans. Geosci. Remote. Sens.3
2025 HyLiOSR: Staged Progressive Learning for Joint Open-Set Recognition of Hyperspectral and LiDAR Data
abstract
The joint classification of hyperspectral images (HSIs) and light detection and ranging (LiDAR) data have seen significant advancements in recent research. However, it would be more practical if we could simultaneously detect the unknown classes in a more realistic open-set scenario. In this article, we introduce a novel open-set recognition (OSR) method for HSI and LiDAR data, termed HyLiOSR, which devises a staged progressive learning strategy to effectively bridge the gap between closed-set and open-set feature distributions within an autoencoder framework. Specifically, for the first stage, the reconstruction-based network is dedicated to accurately modeling each known category by learning multiple Gaussian prototypes, which facilitates OSR by disentangling the distribution of known classes. In the second stage, we actively synthesize samples of unknown classes during the feature extraction phase and create a virtual unknown classifier, enabling the network to effectively differentiate between known and unknown class samples. This approach establishes a distinct separation between known and unknown classes in the latent feature space, thereby enhancing the capability of the frameworks to distinguish between them. Comprehensive experiments conducted on three benchmark datasets demonstrate that the proposed HyLiOSR outperforms existing state-of-the-art methods. The source code will be accessible athttps://github.com/B-Xi/TGRS_2025_HyLiOSR.
Bobo Xi, Mingshuo Cai, Jiaojiao Li 0001, Zhengjue Wang, Shou Feng, Yunsong Li 0001, Jocelyn Chanussot
IEEE Trans. Geosci. Remote. Sens.5
2025 Unsupervised Model-Embedded Two-Stage Diffusion Method for Multispectral and Hyperspectral Image Fusion
abstract
The multispectral and hyperspectral image fusion tasks aim to obtain the hyperspectral image(HSI) with a high spatial resolution. However, existing fusion methods usually utilize degradation simulation for training due to the inaccessible ground truth, which inevitably leads to spatial distortion and parameter bias during application. Besides, neglect of the prior information inherent in inputs and the deep-learning network’s low-frequency preference eventually leads to lower interpretability and limited performance. To this end, we proposed a model-embedded two-stage diffusion method(MTDiff) for unsupervised reconstruction of high spatial resolution HSI. In the first stage, the degradation model is estimated by the prior information inherent input pairs. Meanwhile, in the second stage, embedded with the degradation model, the dual-resolution diffusion model reconstructs high-resolution HSI. Specifically, treating the degradation process as a fixed diffusion step, an unsupervised paradigm is established through a mapping from upsampled low-resolution HSI to high-resolution HSI. Besides, with the estimated degradation model, the well-designed dual-resolution diffusion model step-by-step perturbs and then denoises the image in both native and degraded resolutions for a fidelity reconstruction with great interpretability. Furthermore, to establish a high-frequency shortcut for network learning, a discrete cosine injection module is designed to flatten the frequency information to a clear 2D domain with a huge high-frequency area, achieving sharp textures and clear structures in fusion results. Extensive systematic experiments across three datasets indicate the superior performance of MTDiff in multispectral and hyperspectral fusion tasks.
Jialin Zhou, Shou Feng, Kuo Yuan, Xinlan Xu, Jiaqing Qiao
IEEE Trans. Geosci. Remote. Sens.2
2025 An Adaptive Weighted Metric Learning Network Based on Fractional Domain Decoupling for Hyperspectral Change Detection
abstract
Hyperspectral image change detection (HSI-CD) possesses strong capabilities in exploring subtle changes in land cover. Due to sensor noise and imaging conditions, different semantic land covers in the same spatial location may exhibit similar spectral characteristics, leading to pseudoinvariant phenomena (identification of changed areas as unchanged areas) and causing a higher rate of false negatives in the model. Existing methods primarily focus on obtaining auxiliary discriminative information from spatial correlations or temporal dependencies. However, the frequency domain, which possesses rich global gradient distribution information, is often overlooked. The fractional Fourier transform (FrFT) is an extension of the Fourier transform (FT), representing a temporal-frequency local transformation suitable for processing nonstationary signals. Furthermore, multiorder fractional Fourier domains provide more observable domains for change discrimination. In this work, the application of FrFT is extended to the field of HSI-CD, and an adaptive weighted metric learning network based on fractional domain decoupling (FrFTML) is proposed. Specifically, the fractional domain decoupling (FrDD) module transforms the original HSI into multiorder FrFT domains and extracts their rich spatial-frequency mixed information, effectively suppressing noise while enhancing the representation of subtle differences. In addition, an adaptive weighted metric learning (AWML) framework is designed to merge multiorder fractional Fourier domain information in an adaptively weighted fusion manner. It introduces deep metric learning to explore the distances between samples of different categories that have relatively high similarity, so as to guide the direction of adaptive weighted fusion. Finally, the differential mask attention (DMA) module is designed to explore global contextual differences between bitemporal HSIs, obtaining change features with well-represented differences. Some experiments conducted on three public datasets indicate that FrFTML outperforms other state-of-the-art methods. Furthermore, the proposed method exhibits superiority in dealing with land cover that may lead to pseudoinvariant phenomena (identification of changed areas as unchanged areas).
Shou Feng, Tianyu Lan, Yuanze Fan, Mengmeng Zhang 0005, Chunhui Zhao 0003, Wei Li 0032, Ran Tao 0003
IEEE Trans. Neural Networks Learn. Syst.1
2025 FDGNet: Frequency Disentanglement and Data Geometry for Domain Generalization in Cross-Scene Hyperspectral Image Classification
abstract
Cross-scene hyperspectral image classification (HSIC) poses a significant challenge in recognizing hyperspectral images (HSIs) from different domains. The current mainstream approaches based on domain adaptation (DA) methods need to access target data when aligning distributions between domains, limiting the applicability of the model. In contrast, recent domain generalization (DG) methods aim to directly generalize to unseen domains, eliminating the requirements for target data during training. Nonetheless, most DG-based methods overly focus on randomizing sample styles, leading to semantically compromised samples. In addition, broadening the source distribution without ensuring reasonable support may result in undesired extended distributions. To address these issues, we propose a novel DG network with frequency disentanglement and data geometry (FDGNet) for cross-scene HSIC. Specifically, we first develop a spectral-spatial encoder based on frequency disentanglement (FDSS encoder), which facilitates synthesized domains to preserve their semantic consistency while simulating interdomain gaps with the source domain. Second, to avoid the generation of unrealistic samples, we incorporate data geometry into adversarial training. This helps diversify new domains while keeping the data geometry of extended domains in an explainable support. To improve the learning of domain-invariant representation, we propose an intermediate domain sampling strategy based on the class-wise perceptual manifold. This strategy synthesizes reliable intermediate domains by sampling from class-wise manifold flows estimated over the source and extended domains. Extensive experiments and analysis on three public HSI datasets yield the superiority of our proposed FDGNet. The codes will be available from the website: https://github.com/Qba-heu/FDGNet.
Boao Qin, Shou Feng, Chunhui Zhao 0003, Bobo Xi, Wei Li 0032, Ran Tao 0003
IEEE Trans. Neural Networks Learn. Syst.2
2025 A Semantic Change Detection Network Based on Boundary Detection and Task Interaction for High-Resolution Remote Sensing Images
abstract
Semantic change detection (CD) not only helps pinpoint the locations where changes occur, but also identifies the specific types of changes in land cover and land use. Currently, the mainstream approach for semantic CD (SCD) decomposes the task into semantic segmentation (SS) and CD tasks. Although these methods have achieved good results, they do not consider the incentive effect of task correlation on the entire model. Given this issue, this article further elucidates the SCD task through the lens of multitask learning theory and proposes a semantic change detection network based on boundary detection and task interaction (BT-SCD). In BT-SCD, the boundary detection (BD) task is introduced to enhance the correlation between the SS task and the CD task in SCD, thereby promoting positive reinforcement between SS and CD tasks. Furthermore, to enhance the communication of information between the SS and CD tasks, the pixel-level interaction strategy and the logit-level interaction strategy are proposed. Finally, to fully capture the temporal change information of the bitemporal features and eliminate their temporal dependency, a bidirectional change feature extraction module is proposed. Extensive experimental results on three commonly used datasets and a nonagriculturalization dataset (NAFZ) show that our BT-SCD achieves state-of-the-art performance. The code is available at https://github.com/TangYJ1229/BT-SCD.
Yingjie Tang, Shou Feng, Chunhui Zhao 0003, Zhiyong Lv, Weiwei Sun 0005
IEEE Trans. Neural Networks Learn. Syst.2
2024 A Lightweight Change Detection Method Based on Feature Interaction and Transformer for High Resolution Remote Sensing Images
abstract
Change detection has consistently been a prominent direction in the field of remote sensing. As for high resolution remote sensing images (HRRSI), despite the notable achievements of change detection models, the majority of their impressive performance stems from their large scale architecture or computational requirements. To strike a balance between efficiency and efficacy, a lightweight change detection method based on transformer and feature interaction (LiFTNet) has been proposed. LiFTNet utilizes an efficient backbone, EfficientNet-B4, which is a lightweight network architecture. To fully utilize the information in features with limited model parameters, a multi scale feature interaction module (MSFI) is proposed to aggregate the shallow features and the deep features. As the network has a shallow depth, the semantic information contained in the features is incomplete. To enhance the extraction of semantic information with minimal increases in computational overhead, a lightweight semantic transformer is adopted in the model. A series of experiments indicate the superior performance of LiFTNet over other state-of-the-art (SOTA) methods, showing both efficiency and effectiveness.
Yingjie Tang, Shou Feng, Chunhui Zhao 0003, Yuanze Fan, Maosheng Wei
ICASSP2
2024 A Multi-Modality Feature Enhancement Method Based On Feature Disentanglement For Sar Image Target Detection
abstract
Synthetic Aperture Radar (SAR) ship detection algorithms have achieved extensive development in recent years. In spite of this, the insufficient data and the non-intuitive feature of SAR images still brought certain challenges. This paper proposes a multi-modality feature enhancement (MMFE) method based on feature disentanglement for SAR image target detection. By precisely exploring modality-shared features of optical and SAR images, MMFE can optimize the SAR feature representation capability. First, we propose a feature disentanglement (FD) module to acquire transferable modality-shared knowledge, thereby effectively alleviating the modality shift phenomenon in the subsequent modality alignment. Second, we introduce a multi-granularity modality alignment (MGMA) module that further eliminates inter-modality differences, ultimately achieving effective compensation for the SAR modality. Extensive experimental results convincingly demonstrate the compelling ability of MMFE.
Jiayue He, Nan Su 0001, Yanping Liao, Shou Feng, Chunhui Zhao 0003
ICIP5
2024 Multi-Modal Target Detection Method Based on Adaptive Feature Search
abstract
The optical remote sensing image has a high resolution, while the infrared image provides temperature information about the detected object. These two types of information are complementary. However, optical images often suffer from spatial misalignment issues, which make feature fusion operations challenging. To address these problems, we propose a Transformer feature fusion module that captures high-quality fusion feature information. Building upon this, we design a novel two-branch backbone network that utilizes infrared image features to adaptively screen optical image features, thereby enhancing the detection performance. Experimental results demonstrate the superiority of our approach over the baseline on multi-modal data with non-alignment problems.
Nan Su 0001, Minghui Sha, Chunhui Zhao 0003, Shou Feng, Yingshen Zhu
IGARSS6
2024 An Attention Feature Interaction Change Detection Method Based on Detail Enhancement for Dual-Temporal Hyperspectral Images
abstract
The application of hyperspectral image change detection (HSI-CD) in remote sensing is becoming increasingly widespread. However, due to the low spatial resolution of HSIs, conducting CD directly on the original HSIs does not effectively capture subtle changes. Therefore, this letter proposes an attention feature interaction CD method based on detail enhancement for dual-temporal HSIs (AIDECD). First, a detail enhancement module is designed to enhance the detail information of original HSIs. Second, considering the relationship between dual-temporal images, an attention interaction module is designed to achieve the interaction of temporal features between the dual-temporal images. Then, a multiscale feature extraction module is designed to capture features of different scales. The kappa coefficients obtained on three HSI datasets are 86.11%, 96.08%, and 97.54%, respectively. Compared with six other CD methods, this method has higher detection performance.
Shou Feng, Jinghe Zhang, Ruihui Peng, Chunhui Zhao 0003
IEEE Geosci. Remote. Sens. Lett.1
2024 Mind the Gap: Multilevel Unsupervised Domain Adaptation for Cross-Scene Hyperspectral Image Classification
abstract
Recently, cross-scene hyperspectral image classification (HSIC) has attracted increasing attention, alleviating the dilemma of no labeled samples in the target domain. Although collaborative source and target training has dominated this field, training effective feature extractors and overcoming intractable domain gaps remains challenging. To cope with this issue, we propose a multi-level unsupervised domain adaptation (MLUDA) framework, which comprises image-, feature-, and logic-level alignment between domains to fully investigate the comprehensive spectral-spatial information. Specifically, at the image level, we propose an innovative domain adaptation method named GuidedPGC based on classic image matching techniques and the guided filter. The adaptation results are physically explainable with intuitive visual observations. Regarding the feature level, we design a multi-branch cross attention structure (MBCA) specifically for HSIC, which enhances the interaction between the features from the source and target domains through dot-product attention. Finally, at the logic level, we adopt a supervised contrastive learning (SCL) approach that incorporates a pseudo-label strategy and local maximum mean discrepancy loss, increasing inter-class distance across diverse domains and further improving the classification performance. Experimental results on three benchmark cross-scene datasets demonstrate that our proposed method consistently outperforms the compared approaches. The source code is available at https://github.com/cfcys/MLUDA.
Mingshuo Cai, Bobo Xi, Jiaojiao Li 0001, Shou Feng, Yunsong Li 0001, Zan Li 0001, Jocelyn Chanussot
IEEE Trans. Geosci. Remote. Sens.4
2024 Fractional Fourier-Based Frequency-Spatial-Spectral Prototype Network for Agricultural Hyperspectral Image Open-Set Classification
abstract
At present, hyperspectral image classification (HSIC) technology has been warmly concerned in all walks of life, especially in agriculture. However, existing classification methods operate under the closed-set assumption, which deviates from the real world with open properties. At the same time, there are more serious phenomena of different crops with similar spectrum and same crops with different spectrum in agricultural hyperspectral data, which is also a great challenge to existing methods. In this work, a fractional Fourier based frequency-spatial-spectral prototype network is proposed to address the challenges of open-set hyperspectral image classification in agricultural scenarios. Firstly, fractional Fourier transform is introduced into the network to combine the information in the frequency domain with the spatial-spectral information, so as to expand the difference between different classes on the premise of ensuring the similarity between classes. Then, the prototype learning strategy is introduced into the network to improve the feature recognition capability of the network through prototype loss. Finally, in order to break the stubbornly closed-set property of closed-set classification method, the open-set recognition module is proposed. The difference between the prototype vector and the feature vector is used to judge the unknown class. Experiments on three agricultural hyperspectral datasets show that this method can effectively identify unknown class without sacrificing the classification accuracy of closed-set, and has satisfactory classification performance.
Maoyang Chen, Shou Feng, Chunhui Zhao 0003, Bo Qu, Nan Su 0001, Wei Li 0032, Ran Tao 0003
IEEE Trans. Geosci. Remote. Sens.2
2024 High-Resolution Remote Sensing Image Change Detection Based on Fourier Feature Interaction and Multiscale Perception
abstract
As a significant means of Earth observation, change detection in high-resolution remote sensing images has received extensive attention. Nevertheless, the variability in imaging conditions introduces style discrepancies and a range of pseudochange regions between bitemporal image pairs. Furthermore, changing objects possess diverse morphological representations, which makes accurately identifying change areas and delineating their boundaries within complex object distributions increasingly difficult. In response to the aforementioned challenges, we propose the Fourier feature interaction and multiscale perception (FIMP) model for effective change detection. To mitigate the impact of style discrepancies, FIMP employs the Fourier transform to adaptively filter bitemporal features in the frequency domain while mining the optimized bitemporal features relevant to the change detection task. To enhance the ability to recognize multiscale changing objects, FIMP aggregates and emphasizes the change areas with the introduced temporal change enhancement module (TCEM). By utilizing the U-fusion change perception module (UCPM) to perform multilevel bidirectional fusion of change features at different scales, FIMP can further enhance the ability to delineate complex semantic change boundaries. Experiments on three public datasets show that our approach outperforms seven state-of-the-art methods.
Shou Feng, Chunhui Zhao 0003, Nan Su 0001, Wei Li 0032, Ran Tao 0003, Jinchang Ren
IEEE Trans. Geosci. Remote. Sens.2
2024 Cross-Domain Few-Shot Learning Based on Decoupled Knowledge Distillation for Hyperspectral Image Classification
abstract
Existing cross-domain few-shot learning (FSL) methods for hyperspectral image (HSI) classification have garnered widespread attention due to their excellent performance in recognizing novel classes. To mitigate domain shift, researchers focus on designing sophisticated domain adaptation (DA) modules to directly apply biased metaknowledge in the target domain (TD). However, this paradigm proves somewhat inadequate in the face of significant differences in distribution. To cope with this dilemma, we adopted a new mindset of treating metaknowledge extraction and debiasing from the source domain (SD) as a synergistic process and proposed a cross-domain FSL framework based on decoupled knowledge distillation for HSI classification (HSIC). In general, to efficiently acquire and utilize unbiased metaknowledge, this framework centralizes on a knowledge distillation (KD) strategy. Through the effective information transfer process, the extraction and debiasing of metaknowledge were integrated into a comprehensive and productive process. Simultaneously, to release the constraints imposed by the coupled logits in the KD process on the knowledge interaction, the decoupled logit interaction (DLI) module is employed in the framework. This module decouples the traditional KD into two controllable components, making a more balanced and comprehensive interaction of task-related knowledge and data-intrinsic knowledge between models. Moreover, to facilitate the extraction of critical discriminative metaknowledge from the abundant redundant information in HSI, the discriminative information refinement (DIR) module is designed to develop distinctive features for similar bands. Extensive experiments on three public HSI datasets exhibited the superior performance of the proposed cross-domain few-shot learning method based on decoupled knowledge distillation for HSIC (DKD-FSL) method in comparison with seven state-of-the-art approaches.
Shou Feng, Hongzhe Zhang, Bobo Xi, Chunhui Zhao 0003, Yunsong Li 0001, Jocelyn Chanussot
IEEE Trans. Geosci. Remote. Sens.1
2024 Cross-Domain Few-Shot Learning Based on Feature Disentanglement for Hyperspectral Image Classification
abstract
Existing hyperspectral cross-domain few-shot learning (FSL) methods focus mainly on elaborating on training strategies or domain alignment algorithms, while paying less attention to the biased meta-knowledge introduced by a large amount of source data and the implicit encouragement of learning target domain-specific attributes. In this paper, from the perspective of disentangled representation learning, a novel cross-domain FSL method based on feature disentanglement (FDFSL) is proposed for hyperspectral image classification (HSIC). Specifically, to suppress the representation biased towards the source data and enable the model to implicitly focus on the inherent knowledge of the target domain, an orthogonal low-rank feature disentanglement method is employed to acquire desired features of source and target pipelines. Furthermore, to preserve more shared and discriminative information from the heterogeneous data space (i.e., the spectral dimensions of the source and target scenes are typically different), a multi-order spectral interaction block based on central position encoding (MICD) is proposed to fully integrate the respective features into the spectral domain, which allows the model to emphasize informative spectral dimensions in a data-driven manner. Finally, to diversify the feature representation space while preventing the model overfitting domain alignment task, a self-distillation scheme is developed to facilitate the acquisition of task-relevant feature components. Extensive experiments and analysis on three public HSI datasets suggest the superiority of the proposed method. The code will be available on the website at https://github.com/Qba-heu/FDFSL.
Boao Qin, Shou Feng, Chunhui Zhao 0003, Wei Li 0032, Ran Tao 0003, Wei Xiang 0001
IEEE Trans. Geosci. Remote. Sens.2
2024 Hyperspherical Structural-Aware Distillation Enhanced Spatial-Spectral Bidirectional Interaction Network for Hyperspectral Image Classification
abstract
The existing methods for hyperspectral image classification (HSIC) mainly focus on the extraction of spectral and spatial features while paying less attention to the interaction of each other. Besides, most of them directly use a parameterized classifier as the final layer of the network. While this design is convenient for end-to-end optimization with the backbone, it overlooks the utilization of the metric space. In this article, a novel hyperspherical structural-aware distillation enhanced spatial–spectral bidirectional interaction network (HSDBIN) is proposed for HSIC. HSDBIN uses a dual-branch design combining the 1-D CNN and transformer to separately learn the detailed spectral correlations and global spatial relationships in parallel. Then, by interacting and aggregating the independent information between two parallel branches, a bidirectional interaction block across branches is designed to explore complementary clues between spectral and spatial pipelines. Finally, to enhance the utilization of metric space and keep compact intraclass relationship, we propose a hyperspherical structural-aware distillation (HSD) to transfer the geometric relationship of hyperspherical space into the metric space of output logits. Extensive experiments and analysis on three public HSI datasets suggest the superiority of the proposed method and verify the effectiveness of the proposed modules.
Boao Qin, Shou Feng, Chunhui Zhao 0003, Bobo Xi, Wei Li 0032, Ran Tao 0003, Yunsong Li 0001
IEEE Trans. Geosci. Remote. Sens.2
2024 A Foreground-Driven Fusion Network for Gully Erosion Extraction Utilizing UAV Orthoimages and Digital Surface Models
abstract
Unmanned aerial vehicle (UAV) orthoimages and digital surface models (DSMs) can provide valuable insights for semantic segmentation methods in comprehending gully erosion (GE) from diverse perspectives. While the integration of these two modalities has the potential to improve the GE extraction performance, the extent of enhancement primarily depends on the quality of modality-specific features and the synergistic fusion manner employed for integrating features from both modalities. Toward this end, we propose a novel multimodal segmentation method, which is called foreground-driven fusion network (FFNet). Guided by the prototypes of foreground objects (i.e., gullies), the network effectively tackles the challenges from the modality itself and between different modalities, ultimately achieving high-quality GE extraction results. Specifically, a foreground prototype sampling (FPS) module is first devised for precisely sampling foreground prototypes related to gullies from two modalities. Then, a local-global hybrid purification (LHP) module is proposed to effectively mitigate the erroneous activation within each modality at multiple dimensions by leveraging foreground prototypes. Finally, a multimodal foreground synergy (MFS) module is introduced to further activate foreground features and facilitate full complementarity between multimodal foreground features. To validate our network, a comprehensive multimodal dataset for GE extraction is constructed based on UAV orthoimages and DSMs from northeastern China. Furthermore, a public road extraction dataset is employed to evaluate the generalizability of this network. In the experiments conducted on these two datasets, the proposed FFNet exhibits obvious superiority, outperforming the second-best method with an average improvement of 2.55% in terms of intersection over union (IoU) and 2.77% in terms of$F1$-score. These experimental results not only demonstrate the practicality of FFNet in GE extraction tasks, but also highlight its significant advantage in similar road extraction tasks.
Yi Shen 0013, Nan Su 0001, Chunhui Zhao 0003, Shou Feng, Wei Xiang 0001
IEEE Trans. Geosci. Remote. Sens.5
2024 An Object Fine-Grained Change Detection Method Based on Frequency Decoupling Interaction for High-Resolution Remote Sensing Images
abstract
Change detection is a prominent research direction in the field of remote sensing image processing. However, most current change detection methods focus solely on detecting changes without being able to differentiate the types of changes, such as “appear” or “disappear” of objects. Accurate detection of change types is of great significance in guiding decision-making processes. To address this issue, this article introduces the object fine-grained change detection (OFCD) task and proposes a method based on frequency decoupling interaction (FDINet). Specifically, in order to enhance the model’s ability to detect change types and improve its robustness to temporal information, a temporal exchange framework is designed. Additionally, to better capture spatial–temporal correlation in bi-temporal features, a wavelet interaction module (WIM) is proposed. This module utilizes wavelet transform for frequency decoupling, separating features into different components based on their frequency magnitudes. Then the module applies different interaction methods according to the characteristics of these frequency components. Finally, to aggregate complementary information from different-scale feature maps and enhance the representational capabilities of the extracted features, a feature aggregation and upsampling module (FAUM) is adopted. A series of experiments show the superiority of FDINet over most state-of-the-art methods, achieving good results on three different datasets.
Yingjie Tang, Shou Feng, Chunhui Zhao 0003, Yuanze Fan, Qian Shi 0001, Wei Li 0032, Ran Tao 0003
IEEE Trans. Geosci. Remote. Sens.2
2024 Full-Range Feature Extraction Network Based on Quality-Quantity-Balance Sample Enhancement for Hyperspectral Image Classification
abstract
Hyperspectral remote sensing images exhibit fine spectral curves, but they are also susceptible to spectral variations caused by factors like cloud and haze. It is evident that these issues become more pronounced when there is a limited number of labeled samples available. Thus, a full range feature extraction network (FRFENet) based on quality-quantity-balance sample enhancement is proposed for hyperspectral image classification. First, the full-range feature extraction method combines local-range, short-range, and long-range spatial-spectral features to address spectral variability and ensure accurate feature extraction, particularly in scenarios with limited labeled samples. Furthermore, the approach of balancing quality and quantity for pseudo-labeled samples allows for an increased number of pseudo-labels while maintaining their quality, effectively leveraging unlabeled samples. Additionally, the utilization of superpixel region homogeneity directly contributes to an expanded training sample set, resulting in improved classification performance of the algorithm. Experiments on three HSI datasets indicate that the FRFENet can obtain better classification performance when compared with the other ten state-of-the-art methods.
Chunhui Zhao 0003, Maoyang Chen, Shou Feng, Wenxiang Zhu, Boao Qin
IEEE Trans. Geosci. Remote. Sens.3
2023 A Hyperspectral Change Detection Method Based on Active Learning Strategy
abstract
In recent years, deep learning has demonstrated its transformative potential in the field of hyperspectral image (HSI) processing but is notoriously data-hungry. However, wanting to obtain a large number of labels is labor-intensive and time-consuming. To reduce the dependence of the model on the label samples while maintaining high detection accuracy, a hyperspectral image change detection algorithm based on active learning strategy (ALCD) is proposed. First, the active learning strategy is employed to select high-value labeled samples from the test set as additional training data, gradually enhancing the model’s detection performance. Second, the self-attention module MOAT is introduced to enable effective interaction of local information during the feature extraction process and enhance the network’s feature expression capability. Then, the feature interaction and the mixing block are used to blend the features of the bitemporal images, so that the feature distribution of the bitemporal images is more similar, which is conducive to subsequent feature extraction and classification. Experiments on two HIS datasets show that the proposed method can obtain better change detection results than the four comparison algorithms.
Mingrong Zhu, Chunhui Zhao 0003, Shou Feng, Yuanze Fan, Yingjie Tang
IGARSS4
2023 An End to End Change Detection Method Based on Deep Supervised and Feature Interaction for Erosion Gully
abstract
Erosion gullies are a prominent manifestation of soil erosion. And timely and accurate acquisition of relevant data about erosion gullies plays a crucial role in their management and control. Currently, there is a deficiency in automation within the majority of erosion gully detection methods. The post-classification comparison method using semantic segmentation techniques and the direct change detection method often struggle to ensure high accuracy. Therefore, a end to end change detection method based on deep supervised and feature interaction (DSFNet) is proposed for erosion gullies in this paper. To achieve accurate localization of erosion gully semantic information, DSFNet employs a deep supervision strategy to constrain the semantics of erosion gullies. Furthermore, in order to extract representative features related to erosion gullies and improve the detection accuracy of the model, a feature interaction and upsampling module (IUModule) is employed. Experimental results show that DSFNet exhibits better performance on erosion gully dataset.
Yingjie Tang, Mingrong Zhu, Shou Feng, Chunhui Zhao 0003, Yuanze Fan
IGARSS3
2023 Hyperspectral Image Classification Based on Masked Self-Supervisied Pretraining Network
abstract
Hyperspectral image (HSI) classification is a fundamental research in the field of HSI processing, which has made great development, especially after deep learning-based method is widely used in this field. These methods are commonly starved for labeled samples. However, it is more challenging to obtain labeled samples than HSI in practice. Fortunately, self-supervised leaning (SSL) can take advantage of unlabeled data. In this paper, an HSI classification method based on masked self-supervised network (MSSL) is proposed. To obtain a representative and label-independent representation of HSI, a novel masking and reconstruction based proxy task is designed to accomplish SSL. Furthermore, a spatial masking strategy for HSIs is employed due to the high similarity of adjacent objects in remote sensing images. Experimental results on two benchmark HSI datasets indicate that the proposed MSSL can achieve better classification results with small number of labeled samples.
Hongzhe Zhang, Shou Feng, JianFei Liu, Boao Qin, HaiYang Zhong
IGARSS2
2023 An Attention-Based Multiscale Spectral-Spatial Network for Hyperspectral Target Detection
abstract
Deep learning-based methods have made great progress in hyperspectral target detection. Unfortunately, the insufficient utilization of spatial information in most methods leaves deep learning-based methods to confront ineffectiveness. To ameliorate this issue, an attention-based multiscale spectral-spatial detector (AMSSD) for hyperspectral target detection is proposed. Firstly, the AMSSD leverages the Siamese structure to establish a similarity discrimination network, which can enlarge intraclass similarity and interclass dissimilarity to facilitate better discrimination between the target and the background. Secondly, 1D CNN and vision Transformer are used combinedly to extract spectral-spatial features more feasibly and adaptively. The joint use of spectral-spatial information can obtain more comprehensive features, which promotes subsequent similarity measurement. Finally, a multiscale spectral-spatial difference feature fusion module is devised to integrate spectral-spatial difference features of different scales to obtain more distinguishable representation and boost detection competence. Experiments conducted on two HSI datasets indicate that the AMSSD outperforms seven compared methods.
Shou Feng, Chunhui Zhao 0003, Fengchao Xiong, Lifu Zhang 0002
IEEE Geosci. Remote. Sens. Lett.1
2023 A Coarse-to-Fine Semisupervised Learning Method Based on Superpixel Graph and Breaking-Tie Sampling for Hyperspectral Image Classification
abstract
At present, hyperspectral image classification (HSIC) technology based on deep learning has been widely explored. However, the time and labor cost of obtaining enough labeled samples are expensive. To obtain higher classification performance with a few number of labeled samples, a coarse-to-fine semi-supervised classification learning (CFSSL) method is proposed in this letter. First of all, the CFSSL performs coarse-grained classification with a few number of labeled samples, and the breaking-ties (BT) criterion is introduced to sample the coarse-grained classification results to ensure that the samples with high confidence are selected to generate pseudo-labels. Then, the pseudo-labels and their corresponding unlabeled samples are sent to the feature extraction network for fine-grained classification, so as to obtain more advanced classification results. Finally, in the fine-grained classification stage, a multi-scale convolution kernel attention aggregation network (A2-MCKN) is designed to simultaneously extract the spatial-spectral features of the image and ensure clear texture boundaries of ground objects. Experimental results on two public datasets show that the CFSSL can obtain better accuracy than other methods with a few number of labeled samples.
Chunhui Zhao 0003, Maoyang Chen, Shou Feng, Boao Qin, Lifu Zhang 0002
IEEE Geosci. Remote. Sens. Lett.3
2023 DSTNet: Dynamic-Static Transformer Style Network for Cross-Resolution Vehicle Reidentification
abstract
Vehicle ReIdentification (ReID) can be applied to multi-temporal remote sensing target-matching tasks in different locations. However, due to the uncertainty of UAV height and maneuvering target motion, a huge resolution mismatch can be expected. In the traditional cross-resolution ReID method, the Super-Resolution (SR) method is generally used. However, there is still a large data difference between the Super-Resolution Recovered (SR-Recovered) image and the High-Resolution (HR) image, which leads to a decrease in matching efficiency. Therefore, a dynamic-static TransFormer style network is proposed, which is named DSTNet. DSTNet is designed to reduce the difference between the SR-Recovered image and the HR image and to obtain the identity invariant representation of the SR-Recovered image and the HR image. Firstly, CNN and TransFormer are used to extract context information statically and dynamically, respectively, to enhance the representation of the target identity. Secondly, in order to obtain the invariant information between the SR-Recovered image and the HR image, different normalization strategies are designed in different depths of the DSTNet. Finally, to obtain a consistent representation of the SR-Recovered image and the HR image, the High-Resolution Constraint (HRC) input method is applied to the network. To the experimental results, the performance of rank-5 and mAP is improved by 3% and 3.6% respectively on datasets with large resolution differences by our method.
Chunhui Zhao 0003, Nan Su 0001, Shou Feng
IEEE Geosci. Remote. Sens. Lett.5
2023 A Coarse-to-Fine Hyperspectral Target Detection Method Based on Low-Rank Tensor Decomposition
abstract
To solve the problem of low target detection accuracy caused by the related quantities such as background, target and noise contained in hyperspectral images (HSIs), considering the use of the spatial spectrum and spectral characteristics while increasing the degree of discrimination between target and background, a coarse-to-fine hyperspectral image target detection algorithm based on low-rank tensor decomposition (HTDLTD) is proposed. The HTD based on low rank sparse decomposition mainly decomposes hyperspectral images in spectral dimension, which does not make full use of the spatial information of HSIs, resulting in low detection accuracy. In order to solve this problem, in view of the fact that the hyperspectral third-order tensor can describe the spatial information and spectral information of HSIs equally, the HTD method based on low-rank tensor decomposition (LRTD) is proposed to extract pure background information. Then, in order to solve the problem of low detection accuracy in the case of low target and background discrimination, the rough target detection method based on max over (SMF-MAX) target detection method is proposed to perform rough detection on the original HSI to obtain rough detection results. Finally, in order to further improve the performance of target detection, the fine target detection method based on spectral distance is proposed. By calculating the spectral distance between the original HSI and the synthesized HSI, the final reconstructed target detection result is obtained. Experimental results on three data sets show that the proposed HTDLTD exceeds eight state-of-the-art target detection methods used for comparison.
Shou Feng, Chunhui Zhao 0003, Wei Li 0032, Ran Tao 0003
IEEE Trans. Geosci. Remote. Sens.1
2023 A Cross-Modality Feature Transfer Method for Target Detection in SAR Images
abstract
Synthetic aperture radar (SAR) ship detection methods have achieved remarkable progress in recent years. However, unlike RGB images, the characteristics of SAR imaging will result in non-intuitive feature representations. Furthermore, due to the insufficient data of SAR images, existing methods relying on plenty of labeled SAR images may be hard to achieve promising performance. To address the aforementioned issues, a cross-modality feature transfer (CMFT) method is proposed in this article, which enhances feature representations in the SAR modality by transferring rich knowledge in the RGB modality. First, we propose a multilevel modality alignment network (MMAN), which encourages the model to effectively learn modality-invariant features and alleviate the large cross-modality discrepancies by aligning features from multilevels (scene level, local level, global level, and instance level). Second, to address the underperformance of samples with non-intuitive features in the modality alignment, we introduce a hard-sample supervision module (HSM) in the stage of feature extraction, which can thoroughly exploit the feature of hard-to-align samples by giving more optimization energy for them. Third, to enhance the discriminability of instance-level features, a feature complementary module (FCM) is customized to fully explore the potential complementary clues between instance-level features and context information for the instance-level feature alignment. Extensive experimental results demonstrate that the CMFT outperforms the state-of-the-art detectors. Compared to the baseline model, CMFT improves the accuracy by 3.1% mean average precision (mAP) on the SSDD dataset and 3.4% mAP on the HRSID dataset, demonstrating its superior SAR ship detection performance.
Jiayue He, Nan Su 0001, Cong'an Xu, Yanping Liao, Chunhui Zhao 0003, Shou Feng
IEEE Trans. Geosci. Remote. Sens.8
2023 High-Resolution Remote Sensing Bitemporal Image Change Detection Based on Feature Interaction and Multitask Learning
abstract
With the development of remote sensing technology, high-resolution (HR) remote sensing optical images have gradually become the main source of change detection data. Albeit, the change detection for HR remote sensing images still faces challenges: 1) in complex scenes, a region contains a large amount of semantic information, which makes it difficult to accurately locate the boundaries between different semantics in the feature maps and 2) due to the inability to maintain consistent conditions such as light, weather, and other factors when acquiring bitemporal images, confounding factors such as the style of bitemporal data that are not related to change detection can cause detection difficulties. Therefore, a change detection method based on feature interaction and multitask learning (FMCD) is proposed in this article. To improve the ability to detect changes in complex scenes, FMCD models the context information of features through a multilevel feature interaction module, so as to obtain representative features, and to improve the sensitivity of the model to changes, the interaction between two temporal features is realized through the mix attention block (MAB). In addition, to eliminate the influence of weather and other factors, FMCD adopts a multitask learning strategy, takes domain adaptation as an auxiliary task, and maps the features of bitemporal images to the same space through the feature relationship adaptation module (FRAM) and feature distribution adaptation module (FDAM). Experiments on three datasets show that the proposed method is superior to other state-of-the-art methods.
Chunhui Zhao 0003, Yingjie Tang, Shou Feng, Yuanze Fan, Wei Li 0032, Ran Tao 0003, Lifu Zhang 0002
IEEE Trans. Geosci. Remote. Sens.3
2023 Hyperspectral Image Classification With Multi-Attention Transformer and Adaptive Superpixel Segmentation-Based Active Learning
abstract
Deep learning (DL) based methods represented by convolutional neural networks (CNNs) are widely used in hyperspectral image classification (HSIC). Some of these methods have strong ability to extract local information, but the extraction of long-range features is slightly inefficient, while others are just the opposite. For example, limited by the receptive fields, CNN is difficult to capture the contextual spectral-spatial features from a long-range spectral-spatial relationship. Besides, the success of DL-based methods is greatly attributed to numerous labeled samples, whose acquisition are time-consuming and cost-consuming. To resolve these problems, a hyperspectral classification framework based on multi-attention Transformer (MAT) and adaptive superpixel segmentation-based active learning (MAT-ASSAL) is proposed, which successfully achieves excellent classification performance, especially under the condition of small-size samples. Firstly, a multi-attention Transformer network is built for HSIC. Specifically, the self-attention module of Transformer is applied to model long-range contextual dependency between spectral-spatial embedding. Moreover, in order to capture local features, an outlook-attention module which can efficiently encode fine-level features and contexts into tokens is utilized to improve the correlation between the center spectral-spatial embedding and its surroundings. Secondly, aiming to train a excellent MAT model through limited labeled samples, a novel active learning (AL) based on superpixel segmentation is proposed to select important samples for MAT. Finally, to better integrate local spatial similarity into active learning, an adaptive superpixel (SP) segmentation algorithm, which can save SPs in uninformative regions and preserve edge details in complex regions, is employed to generate better local spatial constraints for AL. Quantitative and qualitative results indicate that the MAT-ASSAL outperforms seven state-of-the-art methods on three HSI datasets.
Chunhui Zhao 0003, Boao Qin, Shou Feng, Wenxiang Zhu, Weiwei Sun 0005, Wei Li 0032, Xiuping Jia
IEEE Trans. Image Process.3
2022 Hyperspectral Image Change Detection Based on Multi-Scale 3D Convolution Autoencoder
abstract
1Change detection has always been a hot research area in the field of hyperspectral image (HSI) processing. However, in the current change detection methods, most of them need to train a large number of labeled data to extract representative features. In this paper, a hyperspectral change detection method based on multi-scale three-dimensional (3D) convolution autoencoder network (M3CAN) is proposed. Firstly, the multi-scale 3D convolution block is adopted in the autoencoder which can extract effective spectral-spatial joint features of HSIs. Then, the autoencoder is pre-trained to obtain the trained encoder as the feature extractor. Finally, the feature maps of the bi-temporal data are obtained by the encoder and then sent to the Softmax classifier to obtain the final change detection result. In this paper, unsupervised training of autoencoder is combined with supervised training of classifier. Therefore, only a small amount of data is needed to complete the training, which avoids the difficulty of requiring many labeled training data. Experiments show that the proposed method has good results on two datasets.
Yingjie Tang, Yuanze Fan, Shou Feng, Chunhui Zhao 0003, Tianfang Luo
IGARSS3
2022 Short and Long Range Graph Convolution Network for Hyperspectral Image Classification
abstract
Nowadays, graph convolution networks are getting more and more attention in the field of hyperspectral image classification. The graph convolution can be divided into long-range and short-range graph convolution (GConv). However, the two graph convolutions cannot acquire global and local features at the same time, making the node features may not be accurate enough. Therefore, we propose a novel graph convolution approach, called short and long range graph convolution (SLGConv), which combines the advantages of long-range and short-range GConv. SLGConv can extract long-range (global) and short-range (local) spatial-spectral features, eliminating the disadvantages of each of long-range and short-range graph convolution. Furthermore, SLGConv can ensure that the features of nodes are not smoothed in the convolution process. Then, three layers of SLGConv are used to form the short and long range graph convolution network (SLGCN) for hyperspectral image classification. Experiments on three HSI datasets indicate that the SLGCN can obtain better classification performance when compared with seven state-of-the-art methods.
Wenxiang Zhu, Chunhui Zhao 0003, Boao Qin, Shou Feng
IGARSS4
2022 Hyperspectral Anomaly Detection With Total Variation Regularized Low Rank Tensor Decomposition and Collaborative Representation
abstract
Nowadays, many anomaly detection (AD) methods still have shortcomings in using the spatial information of hyperspectral images (HSIs), which leads to the inability to separate the background and anomalies well. In this letter, a hyperspectral AD (HAD) approach with total variation regularized low-rank tensor decomposition and collaborative representation (LRTDCRD) is proposed. First, the total variation regularized low-rank tensor decomposition (LRTD) model is adopted to separate an HSI into the background data part and the mixed information part. By virtue of exploiting the global and the piecewise smooth structure of an HSI, the low-rank background data obtained by the LRTD model can be very pure. Then,${l_{2},_{1}}$norm followed by the domain transform recursive filter (DTRF) is built to detect anomalies from the mixed information part. Finally, the collaborative representation-based detector (CRD) is used to extract anomalous information embedding in the low-rank data part. As the low-rank component still contains the information of some anomalies after LRTD, this procedure can be used as a support and supplement for the final detection. Using CRD to detect anomalies in low-rank data can not only ensure the stability of the whole algorithm, but also detect anomalies in low-rank data. The final detection map can be obtained by fusing the initial results of the low-rank data and the mixed information parts. Experimental results on three datasets express that the proposed LRTDCRD exceeds eight state-of-the-art anomaly methods used for comparison.
Shou Feng, Chunhui Zhao 0003
IEEE Geosci. Remote. Sens. Lett.1
2022 A Spectral-Spatial Change Detection Method Based on Simplified 3-D Convolutional Autoencoder for Multitemporal Hyperspectral Images
abstract
Change detection for multitemporal hyperspectral images (HSIs) has always been a research hotspot of remote sensing. However, most current detection methods only use spectral information or spatial information separately, and there are many false detection areas in the detection results. Besides, the feature extraction method based on neural networks needs a huge amount of training samples, but collecting labeled training samples for change detection tasks is difficult. Therefore, this letter proposes a hyperspectral change detection method based on a simplified 3-D convolutional autoencoder (S3DCAECD). First, the framework is based on deep unsupervised autoencoder (AE), which can extract deep spectral–spatial features from bitemporal images without the need for prior information. Second, by adding a 3-D convolution kernel and eliminating the pooling layer, the structure of 3-D convolutional AE is simplified, which can reduce spectral redundancy and improve data processing speed. Finally, a softmax classifier with a 2-D convolutional layer added is used to obtain the detection result, and only a few label samples are needed to train the classifier. Three HSIs’ experimental results indicate that the accuracy of the S3DCAECD is more than 95% on three experimental datasets and it has better detection results than several commonly used methods.
Chunhui Zhao 0003, Shou Feng
IEEE Geosci. Remote. Sens. Lett.3
2022 Spectral-Spatial Anomaly Detection via Collaborative Representation Constraint Stacked Autoencoders for Hyperspectral Images
abstract
Nowadays, due to the ability of extracting deep features, the deep learning-based anomaly detection (AD) methods for hyperspectral images (HSIs) have been widely studied. However, all these AD methods treat the tasks of feature extraction and AD separately. Besides, most of them also do not make use of abundant spatial information of HSIs. Thus, a spectral–spatial hyperspectral AD method via collaborative representation constraint stacked autoencoders (SSCRSAE) is proposed. First, the collaborative representation constraint is imposed on the stacked autoencoders to extract deep nonlinear features that are more suitable for the collaborative representation-based detector (CRD). Then, CRD is used to for obtaining the preliminary detection result, which is more convenient for real HSIs because of no need for assuming the distribution of the background. Finally, aiming at further improving the SSCRSAE detector’s performance, a novel spectral–spatial AD procedure is designed for calculating the final detection result by considering the spatial information of an HSI. Experimental results express that the proposed SSCRSAE exceeds eight state-of-the-art anomaly detectors used for comparison.
Chunhui Zhao 0003, Chuang Li 0005, Shou Feng, Wei Li 0032
IEEE Geosci. Remote. Sens. Lett.3
2022 Hyperspectral Image Classification Based on Kernel-Guided Deformable Convolution and Double-Window Joint Bilateral Filter
abstract
Convolutional neural networks (CNNs) have been widely used in hyperspectral image (HSI) classification. However, a shape-fixed convolution kernel cannot extract appropriate spatial-spectral features. Thus, we propose a novel two-stage classification method based on kernel-guided deformable convolution networks and double-window joint bilateral filter (KDCDWBF) for HSIs. First, according to the calculated similarity map, the shape of the kernel-guided deformable convolution (KDC) is more consistent with the real shape of land covers, so the KDC can extract more pure neighborhood spatial-spectral information. Then, using the piecewise smoothness property of the HSI, a double-window joint bilateral filter (DWJBF) is designed to complete the coarse-to-fine classification stage, which can solve the misclassification problem of single pixels and small regions. Experiments on two HSI datasets demonstrate that the proposed network can achieve better classification performance when compared with other state-of-the-art methods.
Chunhui Zhao 0003, Wenxiang Zhu, Shou Feng
IEEE Geosci. Remote. Sens. Lett.3
2022 Multilevel Feature Alignment Based on Spatial Attention Deformable Convolution for Cross-Scene Hyperspectral Image Classification
abstract
Nowadays, domain adaptation (DA) is getting more attention in cross-scene hyperspectral image (HSI) classification, and various DA algorithms have been proposed. However, regular convolution indiscriminately extracting features around the center pixel will result in the inaccurate extraction of spatial-spectral features, which significantly affect the subsequent feature alignment. Meanwhile, the method of aligning the category features of source and target domains from a single-level may not cope well with complex HSIs. Therefore, we propose a multilevel feature alignment algorithm based on spatial attention deformable convolution (MFA-SADC), which achieves multilevel feature alignment from feature to feature, feature to cluster-center, and cluster-center to cluster-center. In addition, spatial attention deformable convolution is proposed to compose the feature extraction network of MFA-SADC, which guarantees the purity of spatial-spectral features. Experiments on three HSI datasets indicate MFA-SADC can obtain better classification performance when compared with the seven state-of-the-art methods.
Wenxiang Zhu, Chunhui Zhao 0003, Shou Feng, Boao Qin
IEEE Geosci. Remote. Sens. Lett.3
2022 A Hyperspectral Anomaly Detection Method Based on Low-Rank and Sparse Decomposition With Density Peak Guided Collaborative Representation
abstract
The low-rank and sparse decomposition model (LSDM) has been widely studied by researchers and has successfully solved the problem of hyperspectral image (HSI) anomaly detection (AD). The traditional LSDM usually ignores the information of the low-rank matrix, which only detects the anomalous targets by using the sparse component. To utilize both the sparse component and the low-rank component comprehensively, an anomaly detector for HSIs based on LSDM with density peak guided collaborative representation (LSDDPCRD) is proposed in this article. First, the LSDM technique with the mixture of Gaussian model is used to decompose the original HSI, which can also alleviate the background noise contamination problem. Then, the low-rank matrix is detected by the density peak guided collaborative representation detection algorithm, while the sparse matrix is calculated according to the Manhattan distance. In addition, an entropy-based adaptive fusing method is designed to combine the results obtained from the low-rank matrix and the sparse component. It could choose the fusing weights adaptively according to the characteristics of an HSI. The experimental results indicate that the LSDDPCRD performs better than eight classical and state-of-the-art AD algorithms (GRX, LRX, SRX-Segmented, CRD, RPCA-RX, LSMAD, LRASR, and LSDM-MoG) on four real HSIs.
Shou Feng, Shulu Tang, Chunhui Zhao 0003
IEEE Trans. Geosci. Remote. Sens.1
2022 Enhanced Total Variation Regularized Representation Model With Endmember Background Dictionary for Hyperspectral Anomaly Detection
abstract
In recent years, several representation models based on total variation (TV) have been proposed for hyperspectral imagery (HSI) anomaly detection. However, the TV terms of these works are directly imposed on the representation coefficient matrix, which can destroy the spatial structure of an HSI to some extent. Besides, as the spatial resolution of an HSI is relatively low, mixed pixels existing in an HSI can lead to anomaly component contamination, which can make the difference between background and anomalies not significant enough. To address these issues, a novel enhanced TV (ETV) with an endmember background dictionary (EBD) for hyperspectral anomaly detection is proposed. The ETV is designed to be used on the row vectors of the representation coefficient matrix to enhance the spatial structure of an HSI in the presentation process. Furthermore, the proposed ETV regularized representation model with EBD (ETVEBD) method elaborates on a background dictionary constructed by endmembers of background pixels, which are pure spectral signatures of background pixels. The proposed EBD can decrease the influence of anomaly components in mixed pixels, and the coefficient matrix of the EBD has more physical meanings. The proposed method is evaluated on four hyperspectral datasets, and the experiment results show that its performance is the best compared with the other seven state-of-the-art methods.
Chunhui Zhao 0003, Chuang Li 0005, Shou Feng, Xiuping Jia
IEEE Trans. Geosci. Remote. Sens.3
2022 An Unsupervised Domain Adaptation Method Towards Multi-Level Features and Decision Boundaries for Cross-Scene Hyperspectral Image Classification
abstract
Despite success in the same-scene hyperspectral image classification (HSIC), for the cross-scene classification, samples between source and target scenes are not drawn from the independent and identical distribution, resulting in significant performance degradation. To tackle this issue, a novel unsupervised domain adaptation (UDA) framework toward multilevel features and decision boundaries (ToMF-B) is proposed for the cross-scene HSIC, which can align task-related features and learn task-specific decision boundaries in parallel. Based on the maximum classifier discrepancy, a two-stage alignment scheme is proposed to bridge the interdomain gap and generate discriminative decision boundaries. In addition, to fully learn task-related and domain-confusing features, a convolutional neural network (CNN) and Transformer-based multilevel features extractor (generator) is developed to enrich the feature representation of two domains. Furthermore, to alleviate the harm even the negative transfer to UDA caused by task-irrelevant features, a task-oriented feature decomposition method is leveraged to enhance the task-related features while suppressing task-irrelevant features, and enabling the aligned domain-invariant features can be contributed to the classification task explicitly. Extensive experiments on three cross-scene HSI benchmarks have validated the effectiveness of the proposed framework.
Chunhui Zhao 0003, Boao Qin, Shou Feng, Wenxiang Zhu, Lifu Zhang 0002, Jinchang Ren
IEEE Trans. Geosci. Remote. Sens.3
2022 Hyperspectral Target Detection Method Based on Nonlocal Self-Similarity and Rank-1 Tensor
abstract
In recent years, many target detection methods based on tensor representation theory have been proposed and achieved good results for hyperspectral images (HSIs). However, these methods still have some deficiencies. For example, 3-D hyperspectral data are first transformed into 1-D vectors in these methods, which may destroy the spatial structure of HSI data and reduce the detection performance. Besides, when the number of training samples is small, the results of the target detection method usually become worse. To solve these problems, a hyperspectral target detection method based on nonlocal self-similarity and rank-1 tensor is proposed in this article. First, different from these traditional tensor representation-based methods, the third-order tensor data are directly used as the input of the proposed method to preserve the spatial information and structure of an HSI. Second, the tensor blocks related to the class are constructed by using the nonlocal self-similarity of HSI data. Finally, by taking advantage of rank-1 canonical decomposition attribute, the process of tensor operation can be simplified, and the number of training samples can be reduced. The proposed method is compared with six state-of-the-art hyperspectral target detection methods on four HSI data sets. The experimental results show that the proposed method can have better target detection results than other compared methods, especially in the case of fewer training samples.
Chunhui Zhao 0003, Shou Feng, Nan Su 0001
IEEE Trans. Geosci. Remote. Sens.3
2022 Multiscale Short and Long Range Graph Convolutional Network for Hyperspectral Image Classification
abstract
Nowadays, graph convolution networks (GCNs) are getting more attention in hyperspectral image classification, and various algorithms based on GCNs have been proposed. However, because of hyperspectral images’ complex spatial texture information, the long-range graph convolution (GConv) and short-range GConv may cause inaccurate or over-smoothed feature extraction of some nodes. Thus, a multiscale short and long range graph convolution network (MSLGCN) is proposed for hyperspectral image classification. First, MSLGCN not only extracts spatial information of ground objects at different scales but also simultaneously captures global and local spectral features, which preserves objects’ fine boundaries. Then, the rich multiscale information is complementary, enabling the MSLGCN to take full advantage of texture structures of varying sizes. In addition, a method to determine the superpixel scale by the intrinsic properties of hyperspectral images is proposed to ensure that the segmentation boundary depicts the texture structure of the object accurately. Finally, the short-long graph convolution (SLGConv) is designed to fuse the advantages of global and local features, enabling the MSLGCN to extract accurate spatial-spectral features of nodes at any location. Experiments on three HSI datasets indicate that the MSLGCN can obtain better classification performance when compared with the other eleven state-of-the-art methods.
Wenxiang Zhu, Chunhui Zhao 0003, Shou Feng, Boao Qin
IEEE Trans. Geosci. Remote. Sens.3
2022 Superpixel Guided Deformable Convolution Network for Hyperspectral Image Classification
abstract
Convolutional neural networks are widely used in the field of hyperspectral image classification because of their excellent nonlinear feature extraction ability. However, as the sampling position of the regular convolution kernel is unchangeable, the regular convolution cannot distinctively extract the spatial and spectral information around the central pixel, which makes the classification results at the boundaries of ground objects over-smoothed and the classification performance degraded. Thus, we propose a novel superpixel guided deformable convolution network (SGDCN) for hyperspectral image classification. Firstly, the superpixel region fusion filter (SRF-Filter) is designed to fuse the initial superpixel region segmented by the simple linear iterative clustering (SLIC), making the fused superpixel region have a high homogeneity and also contain spatial features of diverse scales. Then, the superpixel guided deformable convolution (SGD-Conv) is proposed to make the shape of deformable convolution consistent with the real shape of land covers, and the SGD-Conv can extract pure neighborhood spatial-spectral features. Finally, a superpixel joint bilateral filter (SPJBF) is designed to solve the pixel-level and region-level misclassification problem, which can effectively utilize the superpixel region's homogeneity and improve the classification accuracy. Experiments on three HSI datasets indicate that the SGDCN can obtain better classification performance when compared with other twelve state-of-the-art methods.
Chunhui Zhao 0003, Wenxiang Zhu, Shou Feng
IEEE Trans. Image Process.3
2021 Hyperspectral Anomaly Detection Using Bilateral-Filtered Generative Adversarial Networks
abstract
Without any prior information of anomalies or background, hyperspectral anomaly detection has received a wide attention. However, such unsupervised style brings difficulties in training and learning effective features of hyperspectral image to perform detection. This paper proposes a novel hyperspectral anomaly detection algorithm using bilateral-filtered generative adversarial networks (BFGAN). Bilateral filter can smooth images and remove anomalous points while preserving edges. With closeness weights and similarity weights, the bilateral-filtered hyperspectral image can be considered as background data, so that hyperspectral background labels are obtained. Only with one class of labels, the structure of generative adversarial networks has an ability to solve two-class problem. By using the filtered background data and their labels, generative adversarial networks are trained to improve discriminator's discriminative capability for background data in a competing style. Finally, the model discriminator can finally output big probabilities for background samples and small probabilities for anomalous samples. Experiments on two real hyperspectral images demonstrate that the proposed method outperforms other state-of-the-art competitors.
Chunhui Zhao 0003, Chuang Li 0005, Shou Feng, Nan Su 0001
IGARSS3
2021 Hyperspectral Image Classification Based on Dense Convolution and Conditional Random Field
abstract
In the research of hyperspectral image (HSI) classification based on deep learning, the small sample problem and the lack of classification accuracy caused by not considering global information have not been well solved. In this paper, an HSI classification method based on dense convolution and conditional random field (DCRF) is proposed. First, the 1D-2D convolution kernel is used to extract the spectral-spatial features and the layers are densely connected to obtain a dense convolutional network to reduce parameters. Second, the Max Pooling layer is used as the output layer of the dense convolutional network to improve the accuracy of feature extraction, and the Softmax layer is used to calculate the probability of the category of the sample and preliminary classification. Finally, the conditional random field is used to fully integrate spatial global information to achieve HSI final classification. Extensive experimental results on two HSI data sets have demonstrated the effectiveness of the proposed DCRF when compared with other state-of-the-art methods.
Chunhui Zhao 0003, Boao Qin, Shou Feng
IGARSS4
2021 A Spectral-Spatial Method Based on Fractional Fourier Transform and Collaborative Representation for Hyperspectral Anomaly Detection
abstract
Anomaly detection (AD) is one of the most important tasks in hyperspectral image (HSI) processing. Most of the traditional AD methods fail to take the advantage of rich spatial information of HSIs and suffer the problem of noise contamination. To solve these problems, we propose a fractional Fourier transform and collaborative representation-based spectral-spatial hyperspectral anomaly detector (SSFrFTCRD). Different from the previous work, fractional Fourier transform (FrFT) is associated with collaborative representation detector (CRD) in the proposed method. FrFT can transfer HSI pixels into a FrFT domain, which can suppress noise and improve the discrimination between background and anomalies. By taking advantage of the CRD, the SSFrFTCRD can adaptively estimate the background through a sliding dual window without assuming its distribution. Furthermore, both spectral and spatial information are utilized to enhance the performance of the proposed detector. Experiments show that the proposed anomaly detector SSFrFTCRD can achieve superior results compared with the other state-of-the-art methods.
Chunhui Zhao 0003, Chuang Li 0005, Shou Feng
IEEE Geosci. Remote. Sens. Lett.3
2020 Spectral-Spatial Stacked Autoencoders Based on the Bilateral Filter for Hyperspectral Anomaly Detection
abstract
Taking advantaging of the ability to extract high-level features, the algorithms based on deep learning for hyperspectral imagery (HSI) anomaly detection have drawn great attention in recent years. In this paper, we propose a method named spectral-spatial stacked autoencoders based on the bilateral filter (SSSAE-BF). First, the bilateral filter is employed to obtain the derived anomaly components and background components. Second, stacked autoencoders (SAE) are respectively utilized on the derived anomaly component and background component for deep features. Finally, the Reed and Xiaoli detector (RXD) is used on the spectral-spatial features to calculate the detection result. Experiments on two real hyperspectral images demonstrate that the proposed method outperforms the other competitors.
Chunhui Zhao 0003, Chuang Li 0005, Shou Feng, Nan Su 0001
IGARSS3
2020 Dictionary Learning Hyperspectral Target Detection Algorithm Based on Tucker Tensor Decomposition
abstract
1As a research hotspot, hyperspectral image target detection is more and more widely used in military and civilian fields. In order to make use of the spatial and spectrum information of hyperspectral image data at the same time, a new dictionary learning hyperspectral image target detection algorithm based on Tucker tensor decomposition is proposed in this paper. The algorithm uses Tucker tensor decomposition to extract effective local image block spatial spectrum features. A detection model based on sparse representation and collaborative representation is established, and experiments are carried out on two representative hyperspectral images data. From the visual detection results, the algorithm effectively extracts the spatial spectrum features in the complex background and strong noise environment, has a good ability to suppress the background, and the detection target is significant.
Chunhui Zhao 0003, Nan Su 0001, Shou Feng
IGARSS4
2020 Video Salient Object Detection via Robust Seeds Extraction and Multi-Graphs Manifold Propagation
abstract
Video salient object detection aims at distinguishing the salient objects from the complex background and highlighting them uniformly in the spatiotemporal domain, which still suffers from the interference of the complicated dynamic background in unconstrained videos. To address this problem, we propose a novel coarse-to-fine spatiotemporal salient object detection method. Specifically, we first model a novel motion energy to exclude the motion noise by exploiting the motion magnitude and motion orientation. Then, a supervoxel-level inter-frame graph model is constructed for each pair of adjacent frames independently, and a robust graph clustering-based saliency seed generation method is proposed to produce a coarse saliency map. Furthermore, the supervoxel-level inter-frame graph model is reconstructed by considering the regional spatiotemporal consistency constraint based on the coarse saliency map. The prior information obtained from pixel clustering is also taken into account to optimize the weight of the inter-frame graph model. Finally, a multi-graphs saliency propagation method is exploited under the manifold regularization framework by fusing the motion energy and appearance feature to refine the coarse saliency map. The extensive experiments on two widely used datasets validate the effectiveness and superiority of the proposed method against 13 state-of-the-art methods in terms of PR-curves, scores of S-measure,$F_{\beta }$, and MAE.
Bing Liu 0022, Junbao Li, Yu Hen Hu, Shou Feng
IEEE Trans. Circuits Syst. Video Technol.6