Qian Du 0001

dblp:62/2443-1 · DBLP profile ↗
← Back
434ranked-venue papers
40as first author
226since 2021 · last 2027
0000-0001-8354-7500ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 356 · 30 first-author · 174 since 2021Artificial intelligence and machine learning · 54 · 8 first-author · 38 since 2021Graphics, computer vision, multimedia, augmented reality and games · 22 · 2 first-author · 14 since 2021Systems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2027 Learning tensor correlation filter with fused low-rank and smoothness priors for hyperspectral video object tracking
Wen-Shuai Hu, Jian-Li Wang, Ran Tao 0003, Qian Du 0001
Expert Syst. Appl.6
2026 Semisupervised graph U-Net with G-ConvLSTM for hyperspectral image classification
Jin-Yu Yang, Heng-Chao Li 0001, Xin-Ru Feng, Feng Gao 0005, Qian Du 0001, Antonio Plaza
Expert Syst. Appl.5
2026 HIMO: Cross-Arbitrary-Modality Image Invariant Feature Transform With Hierarchical Intrinsic Major Orientation
abstract
Invariant feature extraction is a critical challenge in intelligent image processing, particularly with the rapid advancement of multi-source/modal imaging. Cross-modal matching has attracted considerable attention, yet current studies primarily focus on targeted modalities rather than realizing a general approach. In this paper, cross-arbitrary-modal image invariant feature extraction and matching is studied. Inspired by human vision, a purely handcrafted invariant feature transform is proposed for universal cross-modal image matching, named Hierarchical Intrinsic Major Orientation (HIMO). Based on orientation information, a full-chain non-data-driven algorithm is designed that hinges on an Intrinsic Major Orientation (IMO) extraction. The HIMO incorporates a novel keypoint detector utilizing Difference-of-Feature Suppression (DoFS), a Polar-Pyramid descriptor (PolarP), and a Cascaded Dynamic Multi-scale Strategy (CDMS) to effectively address common challenges such as intensity distortion, rotation, scale differences, geometric deformation, and image noise. To validate the proposed method, two massive cross-modal datasets-General Cross-modal Zone (GCZ) and Wide-area Diverse Sources (WDS)-are introduced, alongside two practical evaluation metrics. Comprehensive experiments compared with 10 traditional and 15 deep-learning state-of-the-art algorithms on 5 datasets fully demonstrate that the proposed HIMO achieves superior performance in terms of robustness, stability, and generalization across diverse imaging conditions.
Chenzhong Gao, Wei Li 0032, Desheng Weng, Ran Tao 0003, Xiang-Gen Xia 0001, Qian Du 0001
IEEE Trans. Pattern Anal. Mach. Intell.6
2026 Causal HyperPrompter: A Framework for Unbiased Hyperspectral Camouflaged Object Tracking
abstract
Hyperspectral camouflaged object tracking remains a significant challenge due to the high similarity between objects and replicas in texture and color. Despite recent progress, the bias present in the tracker and the embedding token hinders the model training. Specifically, most methods rely on false-color three-channel images to fine-tune RGB-based trackers. However, it introduces a confounding effect within the RGB domain, potentially leading to harmful biases that misguide the model toward spurious correlations while neglecting the critical spectral discrimination inherent in hyperspectral images. Furthermore, current token-type embedding methods overlook the key correlations between templates and searches, ultimately confusing correlation and impairing tracking performance. To address these challenges, this paper proposes a new unbiased tracking framework named Causal HyperPrompter. It first introduces a structural causal model to disentangle and control exclusive causal factors during tracking, and incorporates a counterfactual intervention strategy to eliminate confounding variables and mitigate the bias inherited from RGB-based models. In addition, we present a novel token-type embedding module that integrates local spectral angle modeling to enhance the semantic link between template and search tokens, thereby improving the model's sensitivity to object localization. Lastly, to overcome the difficulty of manually initializing the bounding box and addressing data scarcity, we introduce a large-scale hyperspectral camouflaged object detection and tracking dataset, BihoT-130 k, consisting of 1,30,750 annotated frames across various camouflage scenes. Extensive experiments on multiple large-scale datasets illustrate the effectiveness of our proposed methods.
Hanzheng Wang, Wei Li 0032, Xiang-Gen Xia 0001, Qian Du 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2026 Alliance: All-in-One Spectral-Spatial-Frequency Awareness Foundation Model
abstract
Frequency domain analysis reveals fundamental image patterns difficult to observe in raw pixel values, while avoiding redundant information in original image processing. Although recent remote sensing foundation models (FMs) have made progress in leveraging spatial and spectral information, they have limitations in fully utilizing frequency characteristics that capture hidden features. Existing FMs that incorporate frequency properties often struggle to maintain connections with the original image content, creating a semantic gap that affects downstream performance. To address these challenges, we propose the All-in-One Spectral-Spatial-Frequency Awareness Foundation Model (Alliance), a framework that effectively integrates information across all three domains. Alliance introduces several key innovations: (1) a progressive frequency decoding mechanism inspired by human visual cognition that minimizes multi-domain information gaps while preserving connections between general image information and frequency characteristics, progressively reconstructing from low to mid to high frequencies to extract patterns difficult to observe in raw pixel values; (2) a triple-domain fusion attention module that separately processes amplitude, phase, and spectral-spatial relationships for comprehensive feature integration; and (3) frequency embedding with frequency-aware Cls token initialization and frequency-specific mask token initialization that achieves fine-grained modeling of different frequency band information. Additionally, to evaluate FMs generalizability, we construct the Yellow River dataset, a large-scale multi-temporal collection that introduces challenging cross-domain tasks and establishes more rigorous standards for FMs assessment. Extensive experiments across six downstream tasks demonstrate Alliance's superior performance.
Wei Li 0032, Yuxiang Zhang 0005, Ran Tao 0003, Qian Du 0001
IEEE Trans. Pattern Anal. Mach. Intell.8
2026 Unsupervised Feature Dimensionality Reduction via Latent Low-Rank Embedding Projection for Classification of Hyperspectral Images
abstract
To deal with the curse of dimensionality in hyperspectral images, numerous feature dimensionality reduction (FDR) methods have been proposed to map high-dimensional data into a low-dimensional subspace. However, most of existing FDR methods lack robustness against noise corruption. To this end, the representation-based subspace learning has been developed to find a robust projection matrix for FDR. Nevertheless, most of them only consider a single direction of the matrix, which ignore the information from other directions. Moreover, the majority of existing methods fail to account for both global structure and feature correlations effectively. To address the above problems, we propose a novel robust projection learning method called latent low-rank embedding (LatLRE), which integrates the latent low-rank representation (LatLRR) with projection learning. In particular, the proposed model can maintain the strong robustness of LatLRR and simultaneously learn a projection for FDR. Moreover, the nuclear norm and logarithmic norm are employed to approximate the two underlying rank functions and provide a more accurate measure of correlation. In addition, LatLRE is optimized using the alternating direction method of multipliers (ADMM) algorithm with the theoretical convergence guarantee. To verify the FDR performance of LatLRE, extensive experiments are conducted on three benchmark hyperspectral datasets. The experimental results demonstrate that LatLRE outperforms other FDR methods considered in this paper.
Heng-Chao Li 0001, Jun-Qiu Wang, Si-Jia Xiang, Qian Du 0001
IEEE Trans. Circuits Syst. Video Technol.5
2026 Adaptive Coarse-to-Fine Parameter Optimization for Incremental Hyperspectral Target Detection
abstract
Hyperspectral target detection effectively identifies fixed targets using specific spectral signatures but suffers from catastrophic forgetting when detecting multiple targets of interest within the same scene. Traditional data replay strategies may further exacerbate training instability due to mislabeled samples. To address these limitations, we propose an Adaptive Coarse-to-Fine Parameter Optimization framework (ACFPO) for incremental hyperspectral target detection, which enables stable continual learning via structural adaptation and parameter sensitivity–aware refinement. ACFPO formulates the task as a dual-stage process: coarse-grained matching and fine-grained detection. Specifically, an Adaptive Spectral Prior-Guided Coarse Matching (AS-PCM) module is designed to hierarchically organize detection tasks into semantic domains and construct intra- and inter-class spectral pairs for coarse-level alignment to adaptively select optimal submodels. Subsequently, a Distance-Aware Localized Fine-Grained Parameter Optimization (DA-LFPO) module is proposed to identify layer-wise sensitive parameters of the selected submodels by measuring spectral–spatial discrepancy, enabling selective retraining to preserve model stability on previously learned classes. By dynamically freezing non-sensitive parameters and optimizing critical modules, our approach mitigates inherent model drift and gradient conflicts in replay-based methods. Extensive experiments on three benchmark datasets demonstrate the superior performance of ACFPO, achieving a balanced trade-off between stability of the existing target and the adaptability of incremental targets. The code is available at https://github.com/Jiahuiqu/ACFPO.
Jiahui Qu, Wenqian Dong, Yunsong Li 0001, Qian Du 0001
IEEE Trans. Circuits Syst. Video Technol.5
2026 Domain-Aware Adversarial Domain Augmentation Network for Hyperspectral Image Classification
abstract
Classifying hyperspectral remote sensing images across different scenes has recently emerged as a significant challenge. When only historical labeled images (source domain, SD) are available, it is crucial to leverage these images effectively to train a model with strong generalization ability that can be directly applied to classify unseen samples (target domain, TD). To address these challenges, this paper proposes a novel single-domain generalization (SDG) network, termed the domain-aware adversarial domain augmentation network (DADAnet) for cross-scene hyperspectral image classification (HSIC). DADAnet involves two stages: adversarial domain augmentation (ADA) and task-specific training. ADA employs a progressive adversarial generation strategy to construct an augmented domain (AD). To enhance variability in both spatial and spectral dimensions, a domain-aware spatial-spectral mask (DSSM) encoder is constructed to increase the diversity of the generated adversarial samples. Furthermore, a two-level contrastive loss (TCC) is designed and incorporated into the ADA to ensure both the diversity and effectiveness of AD samples. Finally, DADAnet performs supervised learning jointly on the SD and AD during the task-specific training stage. Experimental results on two public hyperspectral image datasets and a new Hangzhouwan (HZW) dataset demonstrate that the proposed DADAnet outperforms existing domain adaptation (DA) and domain generalization (DG) methods, achieving overall accuracies of 80.69%, 63.75%, and 87.61% on three datasets, respectively.
Yi Huang 0021, Jiangtao Peng, Weiwei Sun 0005, Na Chen 0008, Zhijing Ye 0001, Qian Du 0001
IEEE Trans. Image Process.6
2026 Toward Memory-Efficient Hyperspectral Image Reconstruction via Consistency Learning
abstract
Spectral reconstruction (SR) aims to recover high-quality hyperspectral images (HSIs) from more readily available RGB or multispectral images (MSIs). While supervised SR has shown promising results, it is hindered by the difficulty of collecting abundant, well-registered RGB-HSI or MSI-HSI pairs. Semi-supervised SR (Semi-SR) offers a more practical solution by exploiting plentiful RGBs/MSIs together with limited HSIs. However, existing Semi-SR approaches still suffer from cross-domain discrepancies, cross-modality inconsistency, and unreliable pseudo-labels. To tackle these challenges, we propose a Manifold-aware Teacher-Student Semi-SR (MTSSR) framework, which seamlessly integrates labeled and unlabeled domains through a teacher-student paradigm and memory-efficient consistency learning. At its core, a Flexible Cross-attention Spectral Reconstruction (FCSR) network extracts scene-related spatial cues via customized self-attention and models scene-agnostic priors through dynamic quantization, thereby enhancing spectral fidelity. Furthermore, a manifold-aware dimensionality analysis derives a latent space that jointly captures spatial and spectral structures across modalities. This enables a manifold-aware alignment loss to enforce cross-modality consistency and a manifold-aware contrastive loss to progressively refine pseudo-label reliability. In addition, we develop a Threshold-adjusted Memory Bank Update (TMBU) strategy, which generates reliable negative samples by storing network-driven representations instead of memory-consuming HSIs, significantly reducing memory consumption. Extensive experiments on three visual and two remote sensing benchmarks demonstrate that MTSSR consistently outperforms state-of-the-art SR methods, achieving robust and memory-efficient spectral reconstruction.
Yihong Leng, Jiaojiao Li 0001, Rui Song 0003, Yunsong Li 0001, Qian Du 0001
IEEE Trans. Image Process.6
2026 Two-Timer-KAN: Dual-Exclusive Fourier KANs With Gaussian Fusion for Few-Shot Multimodal Remote Sensing Imagery Classification
abstract
Multimodal remote sensing imagery classification (MRSIC) aims to synergistically leverage complementary information from heterogeneous data sources, enabling precise land-cover classification. Existing MRSIC approaches predominantly rely on abundant annotated samples, facing critical performance degradation under data-scarce scenarios that are particularly exacerbated by the inherent complexity of heterogeneous multimodal data. Furthermore, effectively extracting spatial-spectral information of multimodal data and fusing the cross-modal heterogeneous features persists as a significant challenge. To address these obstacles, we propose a pioneering few-shot MRSIC network, Two-timer-KAN, which integrates modality-specific feature extraction for spectral- and spatial-dominant data. Specifically, leveraging the nonlinear power of Kolmogorov-Arnold Networks (KANs), we develop the Dual-Exclusive Fourier KAN (DEF-KAN) encoder, which captures modality-specific global features in the frequency domain, bridging spectral and spatial gaps across various datasets. Following this, a Multivariate-Gaussian-based Cross-KAN (MG-Cross-KAN) is dedicated to enhancing the robustness of cross-modality fusion by capturing modality-shared features in a distribution-based manner. Finally, to further tackle classification ambiguity under limited annotated samples, we present a visual-textual bidirectional alignment strategy, which leverages textual descriptions as supplementary semantical knowledge to clarify class feature centers. Extensive experiments demonstrate that the proposed two-timer-KAN achieves superior performance, outperforming the state-of-the-art methods in both accuracy and robustness.
Jiaojiao Li 0001, Hailong Wu, Rui Song 0003, Yunsong Li 0001, Qian Du 0001
IEEE Trans. Image Process.6
2026 Downstream Task-Inspired Underwater Image Enhancement: A Perception-Aware Study From Dataset Construction to Network Design
abstract
In real underwater environments, downstream image recognition tasks such as semantic segmentation and object detection often face challenges posed by problems like blurring and color inconsistencies. Underwater image enhancement (UIE) has emerged as a promising preprocessing approach, aiming to improve the recognizability of targets in underwater images. However, most existing UIE methods mainly focus on enhancing images for human visual perception, frequently failing to reconstruct high-frequency details that are critical for task-specific recognition. To address this issue, we propose a Downstream Task-Inspired Underwater Image Enhancement (DTI-UIE) framework, which leverages human visual perception model to enhance images effectively for underwater vision tasks. Specifically, we design an efficient two-branch network with task-aware attention module for feature mixing. The network benefits from a multi-stage training framework and a task-driven perceptual loss. Additionally, inspired by human perception, we automatically construct a Task-Inspired UIE Dataset (TI-UIED) using various task-specific networks. Experimental results demonstrate that DTI-UIE significantly improves task performance by generating preprocessed images that are beneficial for downstream tasks such as semantic segmentation, object detection, and instance segmentation. The code will be made publicly available at https://github.com/oucailab/DTIUIE.
Bosen Lin, Feng Gao 0005, Yanwei Yu, Junyu Dong, Qian Du 0001
IEEE Trans. Image Process.5
2026 Physics-Guided Time-Interactive-Frequency Network for Cross-Domain Few-Shot Hyperspectral Image Classification
abstract
Recently, domain alignment and metric-based few-shot learning (FSL) have been introduced into hyperspectral image classification (HSIC) to solve the issues of uneven data distribution and scarcity of annotated data faced in practical applications. However, existing cross-domain few-shot methods ignore pivotal frequency priors of the complex field, which contribute to better category discrimination and knowledge transfer. To address this issue, we propose a novel physics-guided time-interactive-frequency network (PTFNet) for cross-domain few-shot HSIC, enabling the extraction of both frequency priors and spatial features (termed "time domain" following Fourier convention) simultaneously through a lightweight time-interactive-frequency module (TiF-Module) as a pioneering effort. Meanwhile, a spectral Fourier-based augmentation module (SFA-Module) is designed to decouple the frequency priors and enhance the diversity of distribution of physical attributes to imitate the domain shift. Then, the physics consistency loss is introduced to regularize the diverse embeddings to approximate the center of each category's physical attributes, guiding the network to excavate more transferable knowledge of source domain (SD). Furthermore, to fully exploit the discriminant time-frequency information and further improve the accuracy of boundary pixels, a set of multiorientation homogeneous prototypes is adopted to represent each class comprehensively, and an intuitive and flexible uncertainty-rectified bidirectional random walk strategy is applied to replace the Euclidean metric for more reliable classification. The experimental results on four public datasets demonstrate the prominent performance of the proposed PTFNet.
Jiaojiao Li 0001, Hailong Wu, Rui Song 0003, Yunsong Li 0001, Qian Du 0001
IEEE Trans. Neural Networks Learn. Syst.6
2026 Multimodal Quaternion Representation Network for Multisource Remote Sensing Data Classification
abstract
The effective integration and classification of hyperspectral images (HSIs) and light detection and ranging (LiDAR) data is of great significance in Earth observation missions, which are confronted with challenges such as insufficient information utilization and feature heterogeneity. This article proposes a multimodal quaternion representation network (MMQRN) for multisource remote sensing (RS) data classification. Specifically, we first propose the multimodal quaternion representation (MMQR), which employs the orthogonal imaginary components of quaternions to model the complex nonlinear interactions among complementary features, thereby enabling their comprehensive fusion and utilization. Subsequently, we design a multimodal feature cross-fusion (MFCF) framework to integrate multisource, multimodal, and multilevel features adequately. Finally, we leverage the ability to capture long-term dependencies of transformers to design a quaternion convolutional transformer network (QCTN) for modeling global and local spatial-spectral information, respectively. Experiments conducted on three multisource RS datasets demonstrate the superior performance of the proposed MMQRN relative to other state-of-the-art classification methods.
Yu-Le Wei, Heng-Chao Li 0001, Jian-Li Wang, Yu-Bang Zheng, Qian Du 0001
IEEE Trans. Neural Networks Learn. Syst.6
2025 Equivariant Attention Graph Capsule Network for Remote Sensing Scene Recognition
abstract
In order to effectively exploit foreground object structures in remote sensing scene recognition, it is crucial to hierarchically parse foreground objects and learn invariant feature representation by adding an equivariant regularization (ER) term to the graph capsule network. Traditionally such equivariance is constructed using group convolutions, which become intractable when composing complex transformations, leading to increased inference time. In addition, global average pooling (GAP) can result in the loss of useful information in the captured features. To deal with this issue, we propose an equivariant attention graph capsule network (EA-GraCaps) in this letter. EA-GraCaps can progressively learn important cues of foreground objects and model potential spatial relations among parts in a transformation equivariance fashion. Specifically, the intragroup capsule layer is first fed to the graph pooling module for preliminary voting, then the intergroup capsules are input into the dual mixing attention (MA) module to refine the votes for coincidence filtering. With this formulation, our approach can characterize spatial hierarchies between object parts and improve the discriminative ability of class capsules. Experimental results demonstrate that the proposed EA-GraCaps can yield superior classification performance on three widely used benchmarks.
Xiaoyong Bian, Guorong Yu, Qian Du 0001
IEEE Geosci. Remote. Sens. Lett.4
2025 Spatial-Spectral Hypergraph Dynamic Gating MLP Network for Hyperspectral Image Classification
abstract
The advancement of spaceborne hyperspectral remote sensing technology has led to the widespread use of hyperspectral imaging, due to its ability to detect subtle spectral differences. Most of the traditional machine learning (ML) methods and popular deep learning (DL) architectures for hyperspectral image (HSI) classification either fail to capture global features or demand high computational resources. While multilayer perceptron (MLP)-based models offer a computationally efficient alternative, they struggle to capture manifold structures and are susceptible to overfitting. To address these challenges, we propose a novel spatial-spectral hypergraph dynamic gating MLP (S2H-DGMLP) framework tailored for HSI classification. The spatial–spectral hypergraph enhances discriminative power by modeling high-order spatial and spectral correlations, jointly optimizing local spatial features and global spectral features to produce more separable feature representations in the embedding space. Within this framework, the channel and spatial projections are statically parameterized using MLP, while the dynamic gating MLP (DGMLP) block captures global contextual information. The dynamic gating mechanism within the DGMLP block automatically adjusts the segmentation ratio to balance spatial and spectral contributions, while incorporating complex nonlinear combinations to improve feature representation. Experimental results on the Pavia University and Houston datasets demonstrate that S2H-DGMLP significantly improves classification performance, confirming its effectiveness in HSI classification tasks.
Yangjun Deng, Yanglan Li, Longfei Ren, Siqiao Tan, Qian Du 0001
IEEE Geosci. Remote. Sens. Lett.5
2025 Knowledge Graph-Guided Deep Network for Hyperspectral Remote Sensing Image Classification
abstract
For the classification of hyperspectral images (HSIs), most deep learning networks are data-driven and lack the usage of prior knowledge. In this letter, we propose a knowledge graph-guided classification network (KGNet), attempting to utilize the prior knowledge of land cover categories to enhance the classification performance. We first construct a knowledge graph on several hyperspectral scenes, which can characterize not only the attributes of land cover categories but also the rich connections between categories. Semantic features are then derived to represent the knowledge in the graph. Knowledge-guided learning is achieved by performing feature alignment between semantic and visual features. Finally, classification is performed on visual features that have contained the knowledge from semantic features. Experiments on three datasets demonstrate the effectiveness of applying the knowledge graph for the classification of hyperspectral remote sensing images.
Li Ma 0005, Yansheng Li 0001, Qian Du 0001
IEEE Geosci. Remote. Sens. Lett.4
2025 Global and Local Attention-Based Transformer for Hyperspectral Image Change Detection
abstract
Recently, transformer-based hyperspectral image (HSI) change detection methods have shown remarkable performance. Nevertheless, existing attention mechanisms in transformers have limitations in local feature representation. To address this issue, we propose global and local attention-based transformer (GLAFormer), which incorporates a global and local attention module (GLAM) to combine high-frequency and low-frequency signals. Furthermore, we introduce a cross-gating mechanism, called cross-gated feedforward network (CGFN), to emphasize salient features and suppress noise interference. Specifically, the GLAM splits attention heads into global and local attention components to capture comprehensive spatial–spectral features. The global attention component uses global attention on downsampled feature maps to capture low-frequency information, while the local attention component focuses on high-frequency details using nonoverlapping window-based local attention. The CGFN enhances the feature representation via convolutions and cross-gating mechanism in parallel paths. The proposed GLAFormer is evaluated on three HSI datasets. The results demonstrate its superiority over state-of-the-art HSI change detection methods. The source code of GLAFormer is available athttps://github.com/summitgao/GLAFormer.
Feng Gao 0005, Junyu Dong, Qian Du 0001
IEEE Geosci. Remote. Sens. Lett.4
2025 Dynamic Frequency Feature Fusion Network for Multisource Remote Sensing Data Classification
abstract
Multi-source data classification is a critical yet challenging task for remote sensing image interpretation. Existing methods lack adaptability to diverse land cover types when modeling frequency domain features. To this end, we propose a Dynamic Frequency Feature Fusion Network (DFFNet) for hyperspectral image (HSI) and Synthetic Aperture Radar (SAR) / Light Detection and Ranging (LiDAR) data joint classification. Specifically, we design a dynamic filter block to dynamically learn the filter kernels in the frequency domain by aggregating the input features. The frequency contextual knowledge is injected into frequency filter kernels. Additionally, we propose spectral-spatial adaptive fusion block for cross-modal feature fusion. It enhances the spectral and spatial attention weight interactions via channel shuffle operation, thereby providing comprehensive cross-modal feature fusion. Experiments on two benchmark datasets show that our DFFNet outperforms state-of-the-art methods in multi-source data classification. The codes will be made publicly available at https://github.com/oucailab/DFFNet.
Yikang Zhao, Feng Gao 0005, Xuepeng Jin, Junyu Dong, Qian Du 0001
IEEE Geosci. Remote. Sens. Lett.5
2025 A Generalized Tensor Formulation for Hyperspectral Image Super-Resolution Under General Spatial Blurring
abstract
Hyperspectral super-resolution is commonly accomplished by the fusing of a hyperspectral imaging of low spatial resolution with a multispectral image of high spatial resolution, and many tensor-based approaches to this task have been recently proposed. Yet, it is assumed in such tensor-based methods that the spatial-blurring operation that creates the observed hyperspectral image from the desired super-resolved image is separable into independent horizontal and vertical blurring. Recent work has argued that such separable spatial degradation is ill-equipped to model the operation of real sensors which may exhibit, for example, anisotropic blurring. To accommodate this fact, a generalized tensor formulation based on a Kronecker decomposition is proposed to handle any general spatial-degradation matrix, including those that are not separable as previously assumed. Analysis of the generalized formulation reveals conditions under which exact recovery of the desired super-resolved image is guaranteed, and a practical algorithm for such recovery, driven by a blockwise-group-sparsity regularization, is proposed. Extensive experimental results demonstrate that the proposed generalized tensor approach outperforms not only traditional matrix-based techniques but also state-of-the-art tensor-based methods; the gains with respect to the latter are especially significant in cases of anisotropic spatial blurring.
Yinjian Wang, Wei Li 0032, Yuanyuan Gui, Qian Du 0001, James E. Fowler
IEEE Trans. Pattern Anal. Mach. Intell.4
2025 Multi-Feature Interaction and Degradation Estimation Transformer for Spectral Compressive Imaging
abstract
Coded Aperture Snapshot Spectral Imaging (CASSI) systems provide an efficient approach to acquiring Hyperspectral Images (HSI), yet the reconstruction process still presents challenges. Traditional Deep Unfolding Networks (DUN) applied to CASSI often face constraints due to inadequate feature utilization and poor handling of multi-scale frequency-domain information, leading to the loss of image detail and global information. Furthermore, most DUN methodologies oversimplify degrading factors and fail to account for issues such as distortions found in actual imaging, thus affecting accuracy and robustness. This paper presents MIDET, a novel DUN tailored for CASSI systems, which integrates the fusion of band information, spatial information, and multi-scale information to meaningfully improve feature utilization and information interaction efficiency. Additionally, MIDET introduces a degradation-guided learning strategy and a frequency feature extraction module, enhancing the capability to handle real imaging distortions and preserve more details in HSI reconstruction. Experimental results demonstrate that MIDET significantly outperforms existing technologies on both simulated and real datasets, effectively enhancing the quality of HSI reconstruction.
Jiaojiao Li 0001, Ding Zhu, Rui Song 0003, Yunsong Li 0001, Qian Du 0001
IEEE Trans. Circuits Syst. Video Technol.6
2025 Contrastive MLP Network Based on Adjacent Coordinates for Cross-Domain Zero-Shot Hyperspectral Image Classification
abstract
With the breakthrough of transfer learning and meta-learning, cross-domain few-shot hyperspectral image classification (CDFSL HSIC) technology has recently achieved satisfactory performance under limited annotations. Nevertheless, the most practical applications are zero-shot scenarios, which are intractable for CDFSL technology, such as the extraterrestrial detection scene, where unexplored objects are recognized by scientists to be more valuable for research. To conquer the zero-shot problem under domain shift, a two-stage contrastive MLP network (MAC-CDZS) is proposed, which constitutes a pioneering effort in the cross-domain zero-shot (CDZS) HSIC task. Firstly, given the remarkable performance of MLPs within a diminutive model size and their enhanced capacity for extracting spatial-spectral features of HSIs, the MLP framework has been strategically chosen as the foundational backbone of the first stage in the MAC-CDZS for facilitating efficient feature extraction. Secondly, to alleviate the potential category collapse, the second-stage fine-tuning framework is introduced, which extends the first-stage backbone by incorporating the elaborate adjacent coordinate module and contrastive learning paradigm for more harmonious classification performance. Specifically, the adjacent coordinate module is creatively designed to adequately mine the adjacent coordinates among samples for ameliorating category collapse from the perspective of grasping more reliable priors. Furthermore, a contrastive learning paradigm is innovatively constructed, comprising a Spatial Augmentation (SA) module tailored for hyperspectral patches and a construction strategy of sample pair under zero-shot conditions, which aims to boost the representation capability and alleviate the class collapse. The superior performance of the MAC-CDZS is demonstrated by experimental results on four benchmark datasets.
Jiaojiao Li 0001, Rui Song 0003, Yunsong Li 0001, Qian Du 0001
IEEE Trans. Circuits Syst. Video Technol.6
2025 Open-Set Domain Adaptation for Hyperspectral Image Classification Based on Weighted Generative Adversarial Networks and Dynamic Thresholding
abstract
Recent studies have shown that the deep domain adaptation (DA) technique has achieved remarkable results in cross-domain hyperspectral image (HSI) classification task. However, these DA methods assume that the source and target domains share the same classes, which may not hold true in real-world applications. Under open-set conditions, since the target domain may contain classes unseen in the source domain, direct domain alignment can lead to negative transfer phenomena. Moreover, the presence of multiple unknown classes in the target domain makes it difficult to learn more discriminative classification boundaries between known and unknown classes. To address these issues, we propose an open-set DA (OSDA) method for HSI classification based on weighted generative adversarial networks and dynamic thresholding (WGDT). First, we introduce a class anchor (CA) strategy to learn the metric space of known classes in the source domain. By calculating the similarity between the target-domain samples and the CA, we compute the reliability weights of the samples belonging to known classes. Then, based on these weights, we design an instance-level weighted-domain adversarial learning strategy to better align samples that are more likely to belong to known classes, avoiding negative transfer phenomena. Finally, we propose a dynamic thresholding method to learn the classification boundaries between known and unknown classes in the feature space and reject unknown class samples, thereby separating known class samples in the target domain. The experimental results on four cross-scene HSI classification tasks demonstrate that our proposed method outperforms some existing methods. The code is available athttps://github.com/Li-ZK/WGDT.
Ke Bi, Zhaokui Li, Yushi Chen 0002, Qian Du 0001, Li Ma 0005, Yan Wang 0087, Zhuoqun Fang, Mingtai Qi
IEEE Trans. Geosci. Remote. Sens.4
2025 Wavelet-Assisted Mamba for Satellite-Derived Sea Surface Temperature Super-Resolution
abstract
Sea surface temperature (SST) is an essential indicator of global climate change and one of the most intuitive factors reflecting ocean conditions. Obtaining high-resolution SST data remains challenging due to limitations in physical imaging, and super-resolution via deep neural networks is a promising solution. Recently, Mamba-based approaches leveraging State Space Models (SSM) have demonstrated significant potential for long-range dependency modeling with linear complexity. However, their application to SST data super-resolution remains largely unexplored. To this end, we propose the Wavelet-assisted Mamba Super-Resolution (WMSR) framework for satellite-derived SST data. The WMSR includes two key components: the Low-Frequency State Space Module (LFSSM) and High-Frequency Enhancement Module (HFEM). The LFSSM uses 2D-SSM to capture global information of the input data, and the robust global modeling capabilities of SSM are exploited to preserve the critical temperature information in the low-frequency component. The HFEM employs the pixel difference convolution to match and correct the high-frequency feature, achieving accurate and clear textures. Through comprehensive experiments on three SST datasets, our WMSR demonstrated superior performance over state-of-the-art methods. Our codes and datasets will be made publicly available at https://github.com/oucailab/WMSR.
Wankun Chen, Feng Gao 0005, Yanhai Gan, Jingchao Cao, Junyu Dong, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.6
2025 Tensor Decomposition-Based Relaxed Linear Regression for Hyperspectral Image Classification
abstract
Linear regression and its variants have achieved considerable success in image classification. However, those methods still encounter two challenges when dealing with hyperspectral image (HSI) classification. On the one hand, the existing ones focus on mining the relationship between the label space and original data space during the classifier training, which is generally sensitive to noise corruptions. On the other hand, transforming the training samples into a strict binary label matrix makes the generalization ability of the classifier limited. To address these challenges, this paper constructs a novel integrative model called tensor decomposition-based relaxed linear regression (TDRLR) for HSI classification. Firstly, the model adopts tensor canonical polyadic (CP) decomposition to learn two dictionaries from spatial and spectral directions respectively, which can help to generate a double dictionary representation for HSI data. Then, the linear regression classifier is integrated to learn a transformation that reveals the mapping relation between the double dictionary representation and label space rather than the original data for enhancing robustness. Meanwhile, a more flexible way, label relaxation, is employed to enlarge the margins between different classes. More importantly, the learned double dictionary representation and classifier can be fine-tuned in tandem to enhance performance through the designed alternate iterative jointly learning algorithm. Experiments conducted on four real-world HSI datasets demonstrate that the proposed method achieves significant improvements in classification performance with a small size training set, when compared with state-of-the-art HSI classification methods.
Yangjun Deng, Lv-Wei Zhang, Longfei Ren, Heng-Chao Li 0001, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.6
2025 Adaptive Frequency Enhancement Network for Remote Sensing Image Semantic Segmentation
abstract
Semantic segmentation of high-resolution remote sensing images plays a crucial role in land-use monitoring and urban planning. Recent remarkable progress in deep learning-based methods makes it possible to generate satisfactory segmentation results. However, existing methods still face challenges in adapting network parameters to various land cover distributions and enhancing the interaction between spatial and frequency domain features. To address these challenges, we propose the Adaptive Frequency Enhancement Network (AFENet), which integrates two key components: the Adaptive Frequency and Spatial feature Interaction Module (AFSIM) and the Selective feature Fusion Module (SFM). AFSIM dynamically separates and modulates high- and low-frequency features according to the content of the input image. It adaptively generates two masks to separate high- and low-frequency components, therefore providing optimal details and contextual supplementary information for ground object feature representation. SFM selectively fuses global context and local detailed features to enhance the network’s representation capability. Hence, the interactions between frequency and spatial features are further enhanced. Extensive experiments on three publicly available datasets demonstrate that the proposed AFENet outperforms state-of-the-art methods. In addition, we also validate the effectiveness of AFSIM and SFM in managing diverse land cover types and complex scenarios. Our codes are available at https://github.com/oucailab/AFENet.
Feng Gao 0005, Miao Fu, Jingchao Cao, Junyu Dong, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.5
2025 MSFMamba: Multiscale Feature Fusion State Space Model for Multisource Remote Sensing Image Classification
abstract
In the field of multisource remote sensing image classification, remarkable progress has been made by using the convolutional neural network (CNN) and Transformer. While CNNs are constrained by their local receptive fields, Transformers mitigate this issue with their global attention mechanism. However, Transformers come with the tradeoff of higher computational complexity. Recently, Mamba-based methods built upon the state space model (SSM) have shown great potential for long-range dependence modeling with linear complexity, but they have rarely been explored for multisource remote sensing image classification tasks. To address this issue, we propose the Multi-Scale Feature Fusion Mamba (MSFMamba) network, a novel framework designed for the joint classification of hyperspectral image (HSI) and light detection and ranging (LiDAR)/synthetic aperture radar (SAR) data. The MSFMamba network is composed of three key components: the Multi-Scale Spatial Mamba (MSpa-Mamba) block, the Spectral Mamba (Spe-Mamba) block, and the fusion Mamba (Fus-Mamba) block. The MSpa-Mamba block employs a multiscale strategy to reduce computational cost and alleviate feature redundancy in multiple scanning routes, ensuring efficient spatial feature modeling. The Spe-Mamba block focuses on spectral feature extraction, addressing the unique challenges of HSI data representation. Finally, the Fus-Mamba block bridges the heterogeneous gap between HSI and LiDAR/SAR data by extending the original Mamba architecture to accommodate dual inputs, enhancing cross-modal feature interactions and enabling seamless data fusion. Together, these components enable MSFMamba to effectively tackle the challenges of multisource data classification, delivering improved performance with optimized computational efficiency. Comprehensive experiments on four real-world multisource remote sensing datasets (Berlin, Augsburg, Houston2018, and Houston2013) demonstrate the superiority of MSFMamba outperforms several state-of-the-art methods and achieves overall accuracies of 76.92%, 91.38%, 92.38%, and 92.86%, respectively. The source codes of MSFMamba will be publicly available athttps://github.com/oucailab/MSFMamba.
Feng Gao 0005, Xuepeng Jin, Xiaowei Zhou 0003, Junyu Dong, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.5
2025 Prototype-Based Information Compensation Network for Multisource Remote Sensing Data Classification
abstract
Multi-source remote sensing data joint classification aims to provide accuracy and reliability of land cover classification by leveraging the complementary information from multiple data sources. Existing methods confront two challenges: inter-frequency multi-source feature coupling and inconsistency of complementary information exploration. To solve these issues, we present a Prototype-based Information Compensation Network (PICNet) for land cover classification based on HSI and SAR/LiDAR data. Specifically, we first design a frequency interaction module to enhance the inter-frequency coupling in multi-source feature extraction. The multi-source features are first decoupled into high- and low-frequency components. Then, these features are recoupled to achieve efficient inter-frequency communication. Afterward, we design a prototype-based information compensation module to model the global multi-source complementary information. Two sets of learnable modality prototypes are introduced to represent the global modality information of multi-source data. Subsequently, cross-modal feature integration and alignment are achieved through cross-attention computation between the modality-specific prototype vectors and the raw feature representations. Extensive experiments on three public datasets demonstrate the significant superiority of our PICNet over state-of-the-art methods. The codes are available at https://github.com/oucailab/PICNet.
Feng Gao 0005, Chuanzheng Gong, Xiaowei Zhou 0003, Junyu Dong, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.7
2025 HZSCM: Hyperspectral Image Zero-Shot Classification via Vision-Language Models
abstract
Most hyperspectral image (HSI) classification methods assume that all classes in the test set are present during training. However, in real-world applications, acquiring labeled training samples is challenging. As a result, it is difficult for the training dataset to cover all possible land cover types, leading to the generalized zero-shot learning (GZSL) problem. Recently, vision-language models (VLMs) have provided rich semantic priors for land cover classes, offering promising potential for GZSL. However, two fundamental gaps hinder their application to HSI classification: the task paradigm gap, arising from the difference between image-level VLMs and the pixel-level HSI classification task; and the knowledge gap, due to the inconsistency between VLM features and HSI spectral–spatial representations. To bridge both gaps, a novel framework leveraging VLM semantic priors for GZSL in HSI classification is proposed, primarily using pseudo-labeling technique to provide knowledge for unseen classes. Specifically, a pseudo-label generation and enhancement module enables a paradigm transition from image-level understanding to pixel-level classification by incorporating HSI’s spatial information. A pseudo-label correction module then refines noisy labels using spectral cues to address the knowledge gap. Finally, a global learning strategy integrates pseudo-label distillation, supervised learning, and feature regularization to classify seen classes while enabling generalization to unseen ones. Experiments on benchmark HSI datasets demonstrate the proposed method’s superiority in generalized zero-shot classification. This work highlights the potential of VLMs in advancing HSI classification in practical applications.
Lingbo Huang, Yushi Chen 0002, Zhaokui Li, Pedram Ghamisi, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.5
2025 Lifelong Learning With Adaptive Knowledge Fusion and Class Margin Dynamic Adjustment for Hyperspectral Image Classification
abstract
With the rapid growth in satellite imagery acquisition and decreasing revisit intervals, efficient on-orbit processing of hyperspectral data has become critical due to limited onboard computing resources. In this context, lifelong learning (LLL) offers a promising solution to enable continuous learning from new data without storing all previous data or retraining from scratch. However, the plasticity-stability dilemma remains a significant challenge, particularly in hyperspectral image (HSI) classification under class-incremental scenarios. To address this, we propose a novel network architecture that integrates contrastive learning and an angular penalty loss. The contrastive learning module facilitates adaptive knowledge fusion, enabling the model to effectively incorporate new information while preserving prior knowledge. The angular penalty loss allows the classifier to dynamically expand for new classes while maintaining discrimination between old and new categories. Together, these components ensure robust knowledge retention, transfer, and adaptability. Experimental results on three benchmark hyperspectral datasets demonstrate that our method significantly outperforms existing approaches, highlighting its efficacy in addressing LLL challenges in HSI classification. The code is available athttps://github.com/Li-ZK/LLL-AFCA.
Zihui Jiang, Zhaokui Li, Yan Wang 0087, Wei Li 0032, Jing Tian 0003, Chuanyun Wang, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.8
2025 NukesFormers: Unpaired Hyperspectral Image Generation With Nonuniform Domain Alignment
abstract
The persistent challenge of acquiring precisely coregistered RGB-hyperspectral image (HSI) pairs has significantly impeded the practical deployment of current data-driven Hyperspectral Image Generation (HIG) networks in engineering applications. Gleichzeitig, the ill-posed nature of the aligning constraints, compounded with the complexities of mining cross-domain features, also hinders the advancement of unpaired HIG (UnHIG) tasks. In this paper, we conquer these challenges by reformulating the UnHIG through Range-Null Space Decomposition (RND), modeling range-space feature interaction and null-space compensations. Specifically, the introduced contrastive learning effectively aligns the geometric and spectral distributions of unpaired data by building the interaction of range space, considering the consistent feature in degradation process. Our Dual-Dimensional Contrastive Prior Module (DCPM) captures mutual information within RGB and HSI domains of a single scene while modeling relationships between internal representations across different scenes, thereby constructing comprehensive cross-domain constraints. Furthermore, the Gabor kernel-based multi-head self-attention (G-MSA) adaptively separates high-frequency components, guiding subsequent modules to concentrate on relevant frequency intervals for target objects. Then, we propose a novel Non-uniform Kolmogorov-Arnold Networks (Nukes) to exhaustively excavate null-space components(degraded/high-frequency representations) through dual-domain frequency mapping. The proposed method was evaluated on three established datasets: NTIRE 2020 ’Clean’ track, NTIRE 2022, CAVE and Grss_dfc_2018. To assess real-world applicability, ratio experiments were conducted by changing the proportion of RGB and HSI. These experiments demonstrate that our approach achieves state-of-the-art performance in UnHIG.
Jiaojiao Li 0001, Shiyao Duan, Rui Song 0003, Yunsong Li 0001, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.6
2025 UAT: Exploring Latent Uncertainty for Semi-Supervised Object Detection in Remote-Sensing Imagery
abstract
Object detectors for remote sensing imagery (RSI) are plagued by data-driven training approaches, underperforming when confronted with inadequate object information. Thus, semi-supervised object detection (SSOD) is introduced to alleviate high data dependence. However, it’s unfeasible to directly apply existing SSOD methods, considering the classification and regression uncertainties resulting from unique characteristics of RSI: 1) inter-class uncertainty: objects in remote sensing scenes are usually densely arranged, leading to severe feature fusion between different categories, 2) intra-class uncertainty: for same-category objects, the complicated contexts induced by various scenarios increase intra-class feature diversity, posing a challenge to clustering and 3) regression uncertainty: the inherent scale variation in RSI gives rise to inconsistent convergence rate, hindering the regression progress from effective converging. Therefore, we propose a Teacher-Student Model(TSM)-based SSOD method aiming at remote sensing scenes, termed Uncertainty-Aware Teacher (UAT), to quantify and promote the detector’s confidence, which is composed of Entropy-based Uncertain-pair Softening (EUS) policy, Multi-Gaussian Sample Fitting (MSF) module and Scale-adaptive Adjunct Loss (SAL). Specifically, EUS identifies the uncertain category pairs with class-wise information entropy, MSF remodels the intra-class distribution and sets the optimal threshold dynamically for pseudo-labels, and SAL adaptively re-scales bboxes based on the original size to accelerate the regression convergence. We have proven the efficiency of our method through extensive experiments on two public remote sensing datasets, DOTA and DIOR.
Jiaojiao Li 0001, Yuqing Ji, Kerui Cheng, Rui Song 0003, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.6
2025 Background Suppression Network With Attention Collapse Inhibited Transformer for Optical Remote Sensing Object Detection
abstract
Object detection for remote sensing imagery (RSI) has been extensively exploited in practical applications. However, similar and multiscale objects in RSI, especially small objects, pose challenges to RSI object detection methods. Particularly, existing approaches ignore irrelevant background in RSI leading to hardship in discriminative feature extraction, resulting in instances of false positive (FP) and false negative (FN) of similar objects. In this article, we propose an irrelevant background suppression network (IBS-Net), which employs the structure of a convolutional neural network (CNN) in series with a Transformer to efficiently capture local and global information in images to respond the challenge of multiscale object detection. Primarily, a background detach module (BDM) is designed behind the backbone to suppress the irrelevant background and enhance the foreground to minimize the interference of irrelevant background for object detection. Furthermore, a composite-sampler (C-S) is devised to sample the vectors describing the foreground and the context, which expands the limited receptive field of the detector to better distinguish similar objects. Especially, considering that transformer-based object detection methods suffer from an attention collapse issue that leads to a degradation of the network representation. An attention collapse inhibited transformer (ACI-former) is presented by designing a partial residual connection, which induces the network to perceive more target information and reduces the loss of small target features thus improving the detection accuracy of small objects. Ultimately, we have conducted related experiments on two benchmarks, which demonstrate that our method has achieved prominent results compared with other mainstream detection methods.
Jiaojiao Li 0001, Haile Li, Rui Song 0003, Yunsong Li 0001, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.6
2025 Adaptive Temperature-Driven Ternary Contrastive Autoencoder Framework for Hyperspectral Target Detection
abstract
Hyperspectral Target Detection (HTD) is a critical task in remote sensing, where numerous deep learning (DL)-based methods have emerged for their powerful ability to extract hierarchical and discriminative features. However, challenges such as insufficient labeled samples and spectral variability lead to formidable issues for DL-based methods like model underfitting and poor robustness. In particular, existing contrastive learning-based detectors rely on native positive-negative pair construction while overlooking hidden positive pairs (i.e., positive but mistakenly constructed as negative), which undermines the model’s ability to maintain consistent feature representation. To address the above issues, we propose an Adaptive Temperature-driven Ternary Contrastive Autoencoder (ATTCA) framework, which performs HTD in a self-supervised manner. Initially, we introduce a novel augmentation technique named Strong-Weak Frequency domain Interference (S-WFI) to expand data while establishing the ternary framework, which can balance the robustness and representation consistency of the model. Additionally, a Dual-Stream Quad-scan Mamba (DSQM) network based on a compositing selective state space model is tailored to effectively extract multi-scale spatial-spectral features and mitigate the effects of spectral variation. Ultimately, we formulate a Reconstruction Weights-driven Adaptive Temperature (RWAT) strategy to dynamically adjust parameters and suppress the separation of hidden positive pairs, which can facilitate the alignment of target features effectively. Experimental results on four real-world benchmark datasets demonstrate that our approach outperforms state-of-the-art methods in terms of detection performance and efficiency.
Jiaojiao Li 0001, Hangyun Liu, Yunsong Li 0001, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.6
2025 Hyperspectral Target Detection Using Diffusion Model and Convolutional Gated Linear Unit
abstract
Deep learning can effectively extract latent information from data to enhance target-background separation in hyperspectral target detection (HTD). However, these models typically require extensive labeled samples, while available target spectra in hyperspectral images (HSI) are scarce. Additionally, existing deep models struggle with target detection in complex backgrounds due to subtle spectral differences. To address these issues, we propose a novel HTD method based on diffusion model and convolutional gated linear unit (HTD-DMCG). First, the diffusion model is integrated with MixUp for data augmentation to generate a diverse and sufficiently large sample set. Next, a Transformer architecture utilizing a convolutional gated linear unit is designed to effectively capture global dependencies and local feature correlations, leading to more discriminative feature representations. Additionally, a new target aggregation and background separation loss is introduced, which emphasizes target sample aggregation while increasing the distance between targets and background samples to enhance separability. The HTD-DMCG method is compared against classical and state-of-the-art HTD methods on four real HSI datasets. Extensive experiments show that it can effectively outperform existing methods in target detection performance. The code is available at https://github.com/Li-ZK/HTD-DMCG.
Zhaokui Li, Xiaobin Zhao, Cuiwei Liu, Xuewei Gong, Wei Li 0032, Qian Du 0001, Bo Yuan 0013
IEEE Trans. Geosci. Remote. Sens.7
2025 SwinMatcher: Universal Cross-Modal Remote Sensing Image Matching With Interactive Swin Transformer
abstract
Cross-modal remote sensing image matching serves as a key technique for collaborative utilization of multi-source information. However, modal differences and geometric distortions between multi-source images pose challenges to existing methods in terms of robustness and generalization. To achieve feature interaction in cross-modal scenes, this paper proposes SwinMatcher, an end-to-end matching model based on the Transformer architecture. Innovatively proposing the window/shifted-window cross-attention based on the window/shifted-window self-attention mechanisms of Swin Transformer, SwinMatcher enables efficient cross-modal feature interaction and multi-scale contextual modeling. It also incorporates a learnable matching module to directly generate semi-dense correspondences. Moreover, a cross-modal remote sensing image matching dataset is generated, which encompasses four modalities: visible light, synthetic aperture radar (SAR), light detection and ranging (LiDAR), and map, distributed across four representative scenes. The dataset includes 400 samples produced via random homography transformations, designed to enhance modal diversity and scene complexity. Experiments demonstrate that SwinMatcher outperforms state-of-the-art methods on this new dataset as well as public benchmarks, exhibiting superior robustness under complex scenes involving coupled modal and geometric distortions. The proposed method and dataset provide novel solutions and evaluation benchmarks for cross-modal remote sensing image matching. The code and testing dataset will be made publicly available at https://github.com/LotrL/SwinMatcher.
Wei Li 0032, Desheng Weng, Chenzhong Gao, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.4
2025 Dynamic Cross-Modal Feature Interaction Network for Hyperspectral and LiDAR Data Classification
abstract
Hyperspectral image (HSI) and light detection and ranging (LiDAR) data joint classification is a challenging task. Existing multisource remote sensing data classification methods often rely on human-designed frameworks for feature extraction, which heavily depend on expert knowledge. To address these limitations, we propose a novel dynamic cross-modal feature interaction network (DCMNet), the first framework leveraging a dynamic routing mechanism for HSI and LiDAR classification. Specifically, our approach introduces three feature interaction blocks: bilinear spatial attention block (BSAB), bilinear channel attention block (BCAB), and integration convolutional block (ICB). These blocks are designed to effectively enhance spatial, spectral, and discriminative feature interactions. A multilayer routing space with routing gates is designed to determine optimal computational paths, enabling data-dependent feature fusion. Additionally, bilinear attention mechanisms are employed to enhance feature interactions in spatial and channel representations. Extensive experiments on three public HSI and LiDAR datasets demonstrate the superiority of DCMNet over the state-of-the-art methods. Our codes are available athttps://github.com/oucailab/DCMNet.
Junyan Lin, Feng Gao 0005, Lin Qi 0004, Junyu Dong, Qian Du 0001, Xinbo Gao 0001
IEEE Trans. Geosci. Remote. Sens.5
2025 Aerial Multiview Stereo via Adaptive Depth Range Inference and Normal Cues
abstract
Three-dimensional digital urban reconstruction from multi-view aerial images is a critical application where deep multi-view stereo (MVS) methods outperform traditional techniques. However, existing methods commonly overlook the key differences between aerial and close-range settings, such as varying depth ranges along epipolar lines and insensitive feature-matching associated with low-detailed aerial images. To address these issues, we propose an Adaptive Depth Range MVS (ADR-MVS), which integrates monocular geometric cues to improve multi-view depth estimation accuracy. The key component of ADR-MVS is the depth range predictor, which generates adaptive range maps from depth and normal estimates using cross-attention discrepancy learning. In the first stage, the range map derived from monocular cues breaks through predefined depth boundaries, improving feature-matching discriminability and mitigating convergence to local optima. In later stages, the inferred range maps are progressively narrowed, ultimately aligning with the cascaded MVS framework for precise depth regression. Moreover, a normal-guided cost aggregation operation is specially devised for aerial stereo images to improve geometric awareness within the cost volume. Finally, we introduce a normal-guided depth refinement module that surpasses existing RGB-guided techniques. Experimental results demonstrate that ADR-MVS achieves state-of-the-art performance on the WHU, LuoJia-MVS, and München datasets, while exhibits superior computational complexity.
Yimei Liu, Yakun Ju, Yuan Rao 0001, Hao Fan 0004, Junyu Dong, Feng Gao 0005, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.7
2025 Cyclic Consistency Constrained Multiview Graph Matching Network for Unsupervised Heterogeneous Change Detection
abstract
Change detection of heterogeneous remote sensing images is a crucial topic for Earth observation, which has various applications in many fields. Most of the existing heterogeneous change detection methods obtain modal-consistent feature representation without fully considering the characteristic of specific data modality, such as hyperspectral image (HSI). Moreover, the acquirement of labeled samples requires high costs of manual operation and extensive domain knowledge. To solve these problems, we propose a cyclic consistency constrained multi-view graph matching network (C3MGM-Net) for unsupervised change detection, which fully considers the spatial-spectral similarity of heterogeneous multi-temporal images from multiple views while preventing the information loss of HSI and PAN/RGB image. The C3MGM-Net transforms the heterogeneous images into three common domains for modal alignment, which not only enhances the spatial-spectral information, but also well preserves the original high-resolution spatial and spectral information in the multi-temporal images. The modal-consistent spatial and spectral information is interacted between multiple domains, so as to make the difference features more distinguishable in terms of both structural and node similarity. With the guidance of change detection results in all domains, the most informative samples are intelligently selected to enlarge the training set, and then fed back to further constrain the consistency of unchanged areas of the multi-temporal images in each domain. The experimental results on heterogeneous datasets demonstrate the effectiveness of the proposed method compared with the state-of-the-art methods. Code is available at https://github.com/Jiahuiqu/C3MGM-for-Heterogeneous-Change-Detection.
Jiahui Qu, Wenqian Dong, Qian Du 0001, Yunshuang Xu, Yunsong Li 0001
IEEE Trans. Geosci. Remote. Sens.3
2025 Weighted Spatiotemporal Fusion via Tensor Collaborative Representation
abstract
Spatiotemporal fusion of remote sensing data is one of the critical techniques for Earth’s surface dynamic monitoring and analysis, which solves the limitation of spatial resolution and temporal coverage in individual sensor. In order to establish a more accurate and physically meaningful spatiotemporal fusion model, a weighted spatiotemporal fusion method via tensor collaborative representation (W-STFTCR) is proposed. Specifically, the collaborative representation (CR) constraint is incorporated into the tensor decomposition framework to prevent overfitting and enhance model robustness. Meanwhile, the superpixel segmentation strategy is adopted to partition the input difference image into superpixel blocks, facilitating block dictionary construction and clustering effectively. In addition, the normalized difference vegetation index (NDVI) and joint information entropy are introduced for weighting bands in predicting the final image, which leads to more accurate and physically meaningful outcomes. To verify the performance of the proposed method, the spatiotemporal fusion experiments on two publicly available datasets were conducted. The experiment results show that the proposed method outperforms the previous state-of-the-art (SOTA) spatiotemporal fusion algorithms, with excellent parameter robustness.
Hongjun Su, Zhaoyue Wu, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.5
2025 Learning Cross-Task Features With Mamba for Remote Sensing Image Multitask Prediction
abstract
Multitask learning (MTL) for remote sensing (RS) image is a rapidly evolving field that requires simultaneous predictions across several related tasks. However, many existing MTL methods often overlook the exploring of cross-task features, while the strong interdependencies among tasks are critical for MTL. In this article, we propose RSMTMamba, an innovative MTL framework that integrates Mamba for multitask prediction in RS images. Our network simultaneously performs semantic segmentation, height estimation, and boundary detection within a unified architecture. The proposed architecture prioritizes the decoder, with a shared encoder for feature extraction. Specifically, a Mamba-based cross-task feature learning (MCFL) module is introduced to capture the interrelations among different tasks. Unlike transformer-based architecture, which requires significant computational resources, the MCFL module can model both local and global cross-task relationships for RS image with linear complexity. Additionally, Mamba-integrated refine decoders are utilized to aggregate features from the encoder, preliminary decoders, and the MCFL module, which enhances multitask prediction performance. The experimental results on three RS datasets demonstrate that our proposed CFLMamba achieves the state-of-the-art prediction performance, outperforming several deep neural networks in RS image analysis. The code is available athttps://github.com/sycs-2024/RSMultitaskMamba.
Liang Xiao 0001, Jianyu Chen 0003, Qian Du 0001, Qiaolin Ye
IEEE Trans. Geosci. Remote. Sens.4
2025 ULADiff: Unmixing-Guided Learnable Abundance-Latent Diffusion for Hyperspectral Image Denoising
abstract
Hyperspectral image (HSI) denoising is a critical preprocessing step in remote sensing. Recently, the denoising diffusion probabilistic models (DDPMs) have emerged as the powerful generative models. However, applying DDPM to HSI denoising task remains challenging owing to the scarcity and acquisition difficulty of HSIs. Thus, how to effectively incorporate physical priors into the DDPM to improve denoising performance remains an underexplored issue. To this end, we propose an Unmixing-Guided Learnable Abundance-Latent Diffusion for HSI Denoising (ULADiff), which is a from-scratch, task-specific diffusion framework that incorporates physically interpretable priors and conditional information into the DDPM. ULADiff comprises three key components, including a Spectral Unmixing Transformer (SUT) network, an abundance-based diffusion model, and a reconstruction module. Specifically, we employ a learnable block-based SUT module in a self-supervised manner to decompose noisy HSIs into the abundance maps and endmembers. The SUT module enables the diffusion model to operate in a lower-dimensional abundance domain that better captures the underlying structure of HSIs. Then, we incorporate the first eigenimage, the reconstructed image via Singular Value Decomposition, as a physically meaningful condition to facilitate controllable generation. Furthermore, we propose a reconstruction module that enforces a spatial-spectral consistency prior by simultaneously imposing total variation regularization on the endmembers and a sparsity constraint on the abundance maps. This design preserves the intrinsic structures of the HSI and improves reconstruction quality. Comprehensive evaluations on synthetic and real-world datasets demonstrate that ULADiff outperforms state-of-the-art methods in both quantitative performance and visual fidelity.
Zhemin Wei, Heng-Chao Li 0001, Yu-Bang Zheng, Jian-Li Wang, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.5
2025 Tensorized High-Order Hypergraph Convolutional Network for Hyperspectral Image Classification
abstract
In recent years, graph convolutional networks (GCNs) have gained increasing attention in hyperspectral image (HSI) classification due to their good ability to model the pairwise relationships between two pixels. However, it is difficult to effectively model more complex relationships among multiple pixels with simple graphs. To solve this problem, we propose a novel tensorized high-order hypergraph convolutional network (TH2GCN) for HSI classification. Specifically, the hypergraph structure is employed to effectively model complex spatial relationships between pixels in HSIs, and we propose a new tensor-based algebraic representation of hypergraphs as a powerful strategy for describing the high-order interaction structures of the hypergraph. Besides, by extending the adjacency matrix-based GCN to the tensor domain and exploiting the tensor decomposition, the TH2GCN method is designed to efficiently extract high-order discriminative information from the hypergraph at low complexity for improving HSI classification performance. Furthermore, the construction of the adjacency tensor on all the data requires a huge amount of memory, especially for large-scale remote sensing images. To this end, the TH2GCN is trained and tested for HSI data in a minibatch fashion. Experimental results on three HSI datasets prove that the performance of the proposed method outperforms the comparison methods.
Jin-Yu Yang, Heng-Chao Li 0001, Shaohui Mei, Qian Du 0001, Antonio Plaza
IEEE Trans. Geosci. Remote. Sens.5
2025 Spatial-Spectral Feature-Enhanced Mamba and SAM-Guided Hyperspectral Multiclass Change Detection
abstract
Multi-class change detection from hyperspectral image (HSI) leverages the rich spectral information of HSIs to detect and classify subtle changes of interest in an imaged scene. However, challenges arise due to limited samples in small categories, which hinder the accurate differentiation of changes. This study proposes a spatial-spectral feature-enhanced Mamba and SAM-guided hyperspectral multi-class change detection (SFMS) method. To address the challenges, a tri-plane gated Mamba is designed to complement spatial information by utilizing the abundant spectral information in HSIs. Additionally, frequency domain features are combined with state space models, enabling the detection of more accurate semantic and texture changes using integrated information from frequency domains. This approach effectively mitigates the problem of inaccurate detection in small-sample categories. Furthermore, the segment anything model (SAM) is adapted, with the features of change areas being enhanced through prior knowledge obtained from segmentation, thereby improving the multi-class change detection accuracy. The experimental results demonstrate that the proposed SFMS method outperforms state-of-the-art techniques, achieving superior multi-class change detection while overcoming the challenges associated with detecting small-sample categories.
Tianming Zhan, Jiaqiang Qi, Xiaobin Yu, Qian Du 0001, Zebin Wu 0001
IEEE Trans. Geosci. Remote. Sens.5
2025 Frequency-Compensated Network for Daily Arctic Sea Ice Concentration Prediction
abstract
Accurately forecasting sea ice concentration (SIC) in the Arctic is critical to global ecosystem health and navigation safety. However, current methods still is confronted with two challenges: 1) these methods rarely explore the long-term feature dependencies in the frequency domain. 2) they can hardly preserve the high-frequency details, and the changes in the marginal area of the sea ice cannot be accurately captured. To this end, we present a Frequency-Compensated Network (FCNet) for Arctic SIC prediction on a daily basis. In particular, we design a dual-branch network, including branches for frequency feature extraction and convolutional feature extraction. For frequency feature extraction, we design an adaptive frequency filter block, which integrates trainable layers with Fourier-based filters. By adding frequency features, the FCNet can achieve refined prediction of edges and details. For convolutional feature extraction, we propose a high-frequency enhancement block to separate high and low-frequency information. Moreover, high-frequency features are enhanced via channel-wise attention, and temporal attention unit is employed for low-frequency feature extraction to capture long-range sea ice changes. Extensive experiments are conducted on a satellite-derived daily SIC dataset, and the results verify the effectiveness of the proposed FCNet. Our codes and data will be made public available at: https://github.com/oucailab/FCNet.
Feng Gao 0005, Yanhai Gan, Junyu Dong, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.5
2025 Spectral Variability-Aware Cascaded Autoencoder for Hyperspectral Unmixing
abstract
Spectral variability inevitably presents in hyperspectral images (HSIs), resulting in significant unmixing errors when using the conventional linear mixture model (LMM). Though several variants of LMM have been proposed to encounter such spectral variability, they cannot well model the complex characteristics of spectral variability, and the performance of these variants strongly depends on the prior knowledge of the scene. In this article, spectral variability within an image is classified into class-dependent variability and class-independent one, which can be tackled by a novel fully linear mixture model (FLMM) introducing a class-dependent multiplicative scaling term, a class-dependent additive perturbation term, and a class-independent variability term into the conventional LMM. Moreover, a spectral variability-aware cascaded autoencoder (SVACA) is designed to realize the automatic learning and representation of unmixing targets and spectral variability in different hyperspectral scenarios, which consists of a class-independent variability autoencoder and a cascaded class-dependent variability autoencoder. Such a network is able to handle different spectral variability autonomously without any scene prior by parallel inference structure. Experimental results over synthetic and real hyperspectral datasets demonstrate that the proposed SVACA network not only outperforms several state-of-the-art unmixing networks but also presents a stronger capability to handle spectral variability within HSIs.
Ge Zhang 0006, Shaohui Mei, Huiyang Han, Yan Feng 0005, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.6
2025 Stepwise Deep Feature Transfer Model for Martian Landform Mapping With Small Number of Labeled Samples
abstract
The Martian surface landforms are highly related to the safe landing and traversability of Mars rovers. Furthermore, landforms associated with the presence of water/ice, minerals and biosignatures can provide valuable insights for Mars exploration missions, particularly in relation to the selection of landing or sample collection sites. The small number of Martian landform datasets and the scarcity of labelable landform samples over Mars make the precise mapping of Martian landforms a challenging task. In this article, we propose a stepwise deep feature transfer (SDFT) model for the mapping of Martian landforms with a small number of labeled samples. The SDFT model comprises two transfer steps. In the first transfer step, a deep learning model trained on a large public source dataset from Earth is transferred to a medium sized public dataset from Mars. This transfer is conducted through a standard pre-training and fine-tuning procedure utilizing a linear classifier. In the second transfer step, the model is further transferred to a small number of target datasets on Mars through a pre-training and fine-tuning procedure with a cosine distance classifier. The stepwise training technique mitigates the challenges associated with varying datasets and small training samples. The proposed SDFT model has been validated on two self-built sample sets using images from the Mars Reconnaissance Orbiter’s Context Camera (CTX). It has also been employed for landform mapping in two local regions with small samples to evaluate its effectiveness in comparison with existing state-of-the-art methods.
Sicong Liu 0001, Xiaohua Tong, Qian Du 0001, Lorenzo Bruzzone, Huan Xie 0001, Yongjiu Feng, Kecheng Du, Jie Zhang 0117, Yonggang Xiong
IEEE Trans. Geosci. Remote. Sens.4
2025 SSF-Net: Spatial-Spectral Fusion Network With Spectral Angle Awareness for Hyperspectral Object Tracking
abstract
Hyperspectral video (HSV) offers valuable spatial, spectral, and temporal information simultaneously, making it highly suitable for handling challenges such as background clutter and visual similarity in object tracking. However, existing methods primarily focus on band regrouping and rely on RGB trackers for feature extraction, resulting in limited exploration of spectral information and difficulties in achieving complementary representations of object features. In this paper, a spatial-spectral fusion network with spectral angle awareness (SSF-Net) is proposed for hyperspectral (HS) object tracking. Firstly, to address the issue of insufficient spectral feature extraction in existing networks, a spatial-spectral feature backbone ( $S^{2}$ FB) is designed. With the spatial and spectral extraction branch, a joint representation of texture and spectrum is obtained. Secondly, a spectral attention fusion module (SAFM) is presented to capture the intra- and inter-modality correlation to obtain the fused features from the HS and RGB modalities. It can incorporate the visual information into the HS context to form a robust representation. Thirdly, to ensure a more accurate response to the object position, a spectral angle awareness module (SAAM) is designed to investigate the region-level spectral similarity between the template and search images during the prediction stage. Furthermore, a novel spectral angle awareness loss (SAAL) is developed to offer guidance for the SAAM based on similar regions. Finally, to obtain the robust tracking results, a weighted prediction method is considered to combine the HS and RGB predicted motions of objects to leverage the strengths of each modality. Extensive experiments on the HOTC-2020, HOTC-2024, and BihoT datasets demonstrate the effectiveness of the proposed SSF-Net compared with state-of-the-art trackers. The source code will be available at https://github.com/hzwyhc/hsvt.
Hanzheng Wang, Wei Li 0032, Xiang-Gen Xia 0001, Qian Du 0001, Jing Tian 0003
IEEE Trans. Image Process.4
2025 Uncertainty-Guided Discriminative Priors Mining for Flexible Unsupervised Spectral Reconstruction
abstract
Existing supervised spectral reconstruction (SR) methods adopt paired RGB images and hyperspectral images (HSIs) to drive the overall paradigms. Nonetheless, in practice, "paired" requires higher device requirements such as specific well-calibrated dual cameras or more complex and exact registration processes among images with different time phases, widths, and spatial resolution. To tackle the above challenges, we propose a flexible uncertainty-aware unsupervised SR paradigm, which dynamically establishes the forceful and potent constraints with RGBs for driving unsupervised learning. As a specific plug-and-play tail in our paradigm, the uncertainty-aware saliency alignment module (USAM) calculates pixel- and spectralwise information entropy for uncertainty estimation, which attempts to represent the corresponding reflectivity or radiance to the light among different objects in various scenes, forcing the paradigm to adaptively explore the scene-agnostic prominent features. Furthermore, a progressively parallel network under our unsupervised paradigm is conducted to excavate discriminate structural and semantic priors of RGBs to assist in recovering dependable HSIs: 1) a learnable rank-guided structural representation (LRSR) flow is leveraged to characterize the latent structural priors via excavating nonzero elements in the full-rank matrix and further preserve evident boundaries in HSIs; and 2) a coarse-to-fine bandwise semantic perception (CBSP) flow is conducted to propagate perceptual bandwise affinity for aggregating and strengthening intrinsic interband dependencies, and further extract delicate semantic priors, which can recover plentiful contiguous spectral information in HSIs. Comprehensive quantitative and qualitative experimental results on three visual and two remote sensing benchmarks have shown the superiority and robustness of our method. We also conducted nine existing SR methods in our unsupervised paradigm to recover HSIs without any manual intervention, which proves the generality of our paradigm to some extent.
Yihong Leng, Jiaojiao Li 0001, Rui Song 0003, Yunsong Li 0001, Qian Du 0001
IEEE Trans. Neural Networks Learn. Syst.5
2025 Progressive Spatial Information-Guided Deep Aggregation Convolutional Network for Hyperspectral Spectral Super-Resolution
abstract
Fusion-based spectral super-resolution aims to yield a high-resolution hyperspectral image (HR-HSI) by integrating the available high-resolution multispectral image (HR-MSI) with the corresponding low-resolution hyperspectral image (LR-HSI). With the prosperity of deep convolutional neural networks, plentiful fusion methods have made breakthroughs in reconstruction performance promotions. Nevertheless, due to inadequate and improper utilization of cross-modality information, the most current state-of-the-art (SOTA) fusion-based methods cannot produce very satisfactory recovery quality and only yield desired results with a small upsampling scale, thus affecting the practical applications. In this article, we propose a novel progressive spatial information-guided deep aggregation convolutional neural network (SIGnet) for enhancing the performance of hyperspectral image (HSI) spectral super-resolution (SSR), which is decorated through several dense residual channel affinity learning (DRCA) blocks cooperating with a spatial-guided propagation (SGP) module as the backbone. Specifically, the DRCA block consists of an encoding part and a decoding part connected by a channel affinity propagation (CAP) module and several cross-layer skip connections. In detail, the CAP module is customized by exploiting the channel affinity matrix to model correlations among channels of the feature maps for aggregating the channel-wise interdependencies of the middle layers, thereby further boosting the reconstruction accuracy. Additionally, to efficiently utilize the two cross-modality information, we developed an innovative SGP module equipped with a simulation of the degradation part and a deformable adaptive fusion part, which is capable of refining the coarse HSI feature maps at pixel-level progressively. Extensive experimental results demonstrate the superiority of our proposed SIGnet over several SOTA fusion-based algorithms.
Jiaojiao Li 0001, Songcheng Du, Rui Song 0003, Yunsong Li 0001, Qian Du 0001
IEEE Trans. Neural Networks Learn. Syst.5
2025 Fully Tensorized Lightweight ConvLSTM Neural Networks for Hyperspectral Image Classification
abstract
Convolutional long short-term memory (ConvLSTM) possesses a remarkable capability of encoding spatial information and capturing long-range dependencies in sequential data. As a result, ConvLSTM has garnered success in hyperspectral image (HSI) classification. Nonetheless, the design of the special gate structures and convolution operations contributes to a high model complexity, making it challenging to deploy in resource-constrained environments. In this article, we propose a fully tensorized ConvLSTM model for HSI spatial-spectral classification under the premise of low complexity. First, we devise a novel and efficient tensor-sequenced convolution in the tensor train (TT) format, called ETTConv. ETTConv can reduce the number of parameters and computations in the standard convolutional layer by tensorizing the convolution kernels and mapping them to a series of smaller ones. Building upon this innovation, we present a novel ETTConvLSTM unit, formed by jointly compressing all weight tensors within the recurrent units. Using it as the fundamental unit, we construct the lightweight a efficient tensor train ConvLSTM 2-D neural network (ETTCL2DNN) model, characterized by reduced complexity without compromised classification performance. Furthermore, to better preserve the joint spatial-spectral structure of HSI data, we extend the ETTConv layer and the ETTConvLSTM unit to their 3-D versions, resulting in a new lightweight a efficient tensor train ConvLSTM 3-D neural network (ETTCL3DNN) model. Extensive quantitative experimental results on three widely used HSI datasets demonstrate the superiority of the proposed methods, exhibiting enhanced classification performance with reduced model complexity.
Tian-Yu Ma, Heng-Chao Li 0001, Yu-Bang Zheng, Qian Du 0001, Antonio Plaza
IEEE Trans. Neural Networks Learn. Syst.4
2025 A Principle Design of Registration-Fusion Consistency: Toward Interpretable Deep Unregistered Hyperspectral Image Fusion
abstract
For hyperspectral image (HSI) and multispectral image (MSI) fusion, it is often overlooked that multisource images acquired under different imaging conditions are difficult to be perfectly registered. Although some works attempt to fuse unregistered images, two thorny challenges remain. One is that registration and fusion are usually modeled as two independent tasks, and there is no yet a unified physical model to tightly couple them. Another is that deep learning (DL)-based methods may lack sufficient interpretability and generalization. In response to the above challenges, we propose an unregistered HSI fusion framework energized by a unified model of registration and fusion. First, a novel registration-fusion consistency physical perception model (RFCM) is designed, which uniformly models the image registration and fusion problem to greatly reduce the sensitivity of fusion performance to registration accuracy. Then, an HSI fusion framework (MoE-PNP) is proposed to learn the knowledge reasoning process for solving RFCM. Each basic module of MoE-PNP one-to-one corresponds to the operation in the optimization algorithm of RFCM, which can ensure clear interpretability of the network. Moreover, MoE-PNP captures the general fusion principle for different unregistered images and therefore has good generalization. Extensive experiments demonstrate that MoE-PNP achieves state-of-the-art performance for unregistered HSI and MSI fusion. The code is available at https://github.com/Jiahuiqu/MoE-PNP.
Jiahui Qu, Jizhou Cui, Wenqian Dong, Qian Du 0001, Song Xiao 0001, Yunsong Li 0001
IEEE Trans. Neural Networks Learn. Syst.4
2025 Cycle-Refined Multidecision Joint Alignment Network for Unsupervised Domain Adaptive Hyperspectral Change Detection
abstract
Hyperspectral change detection, which provides abundant information on land cover changes in the Earth's surface, has become one of the most crucial tasks in remote sensing. Recently, deep-learning-based change detection methods have shown remarkable performance, but the acquirement of labeled data is extremely expensive and time-consuming. It is intuitive to learn changes from the scene with sufficient labeled data and adapting them into an unlabeled new scene. However, the nonnegligible domain shift between different scenes leads to inevitable performance degradation. In this article, a cycle-refined multidecision joint alignment network (CMJAN) is proposed for unsupervised domain adaptive hyperspectral change detection, which realizes progressive alignment of the data distributions between the source and target domains with cycle-refined high-confidence labeled samples. There are two key characteristics: 1) progressively mitigate the distribution discrepancy to learn domain-invariant difference feature representation and 2) update the high-confidence training samples of the target domain in a cycle manner. The benefit is that the domain shift between the source and target domains is progressively alleviated to promote change detection performance on the target domain in an unsupervised manner. Experimental results on different datasets demonstrate that the proposed method can achieve better performance than the state-of-the-art change detection methods.
Jiahui Qu, Wenqian Dong, Tongzhen Zhang, Yunsong Li 0001, Qian Du 0001
IEEE Trans. Neural Networks Learn. Syst.6
2025 BihoT: A Large-Scale Dataset and Benchmark for Hyperspectral Camouflaged Object Tracking
abstract
Hyperspectral object tracking (HOT) has many important applications, particularly in scenes where objects are camouflaged. The existing trackers can effectively retrieve objects via band regrouping because of the bias in the existing HOT datasets, where most objects tend to have distinguishing visual appearances rather than spectral characteristics. This bias allows a tracker to directly use the visual features obtained from the false-color images generated by hyperspectral images (HSIs) without extracting spectral features. To tackle this bias, the tracker should focus on the spectral information when object appearance is unreliable. Thus, we provide a new task called hyperspectral camouflaged object tracking (HCOT) and meticulously construct a large-scale HCOT dataset, BihoT, consisting of 41912 HSIs covering 49 video sequences. The dataset covers various artificial camouflage scenes, where objects have similar appearances, diverse spectrums, and frequent occlusion (OCC), making it a challenging dataset for HCOT. Besides, a simple but effective baseline model, named spectral prompt-based distractor-aware network (SPDAN), is proposed, comprising a spectral embedding network (SEN), a spectral prompt-based backbone network (SPBN), and a distractor-aware module (DAM). Specifically, the SEN extracts spectral-spatial features via 3-D and 2-D convolutions to form a refined prompt representation. Then, the SPBN fine-tunes powerful RGB trackers with spectral prompts and alleviates the insufficiency of training samples. Moreover, the DAM utilizes a novel statistic to capture the distractor caused by occlusion from objects and background and corrects the deterioration of the tracking performance via a novel motion predictor. Extensive experiments demonstrate that our proposed SPDAN achieves the state-of-the-art performance on the proposed BihoT and other HOT datasets.
Hanzheng Wang, Wei Li 0032, Xiang-Gen Xia 0001, Qian Du 0001
IEEE Trans. Neural Networks Learn. Syst.4
2025 Distributed Deep Learning With Gradient Compression for Big Remote Sensing Image Interpretation
abstract
Fast and reliable interpretation of high-dimensional hyperspectral images (HSIs) can provide great support to remote sensing-based Earth observations. Targets of interest in HSI can be detected using deep neural networks (DNNs) for background learning on an acquired image where the occurrence probability of background samples is much greater than that of targets, accounting for more than 95% of the whole scene. However, there is an increasing gap between theory and feasible application, because of the contradiction between massive hyperspectral data and resource-limited Internet of Things (IoT)/edge device hardware like satellite. To facilitate the deployment of hyperspectral target detection (HTD) in an edge computing environment, we introduce distributed background learning-a decentralized deep learning approach to meet the computing requirements of exploding high-dimensional data and larger DNNs. To address the communication bottleneck caused by gradient exchange during distributed learning, the proposed gradient compression solution, named gradient compression via centroid (GCC), uniquely compresses the most replaceable gradients with redundant information, thereby reducing communication overhead while maintaining accuracy. To illustrate the feasibility of the proposed method, we test it over two very large hyperspectral datasets with a total size of about 3.2 gigabytes (GBs) on a distributed system based on Ring All-reduce. We show that HTD based on distributed background learning outperforms those developed on a single node in terms of speed. Besides, the GCC compresses 50% gradients with only 0.01% loss of target detection accuracy to greatly reduce the communication overhead, surpassing existing gradient compression methods. It is expected that this framework will accelerate the introduction of distributed training on IoT/edge devices.
Weiying Xie, Jitao Ma, Tianen Lu, Yunsong Li 0001, Jie Lei 0001, Leyuan Fang, Qian Du 0001
IEEE Trans. Neural Networks Learn. Syst.7
2025 PRF-Net: A Progressive Remote Sensing Image Registration and Fusion Network
abstract
Most of the existing fusion algorithms are not robust to unregistered input images. Even after image registration, nonlinear nonregistration may persist in the local areas of the images, leading to poor quality in the fused image. So, as to tackle these challenges, a progressive remote sensing image registration and fusion network is proposed in this article, and named PRF-Net, which is particularly useful when two images are from different platforms. First, a registration network is designed to register the input image patches, which includes a global spatial transform network (GSTN) and a local spatial warp network (LSWN). The GSTN is primarily used for coarse registration, applying rigid transformation to globally align the input images. After coarse registration, the preliminarily registered moving image is input into the LSWN for local fine-tuning to maximize correlation between the input image patches. Subsequently, the fine registered images are degraded and input into the fusion network to generate the fused image. To maintain sufficient spectral and spatial information of the fused image, a multiscale feature extraction (MSFE) block with a highly interpretable spatial details attention (SDA) block is designed, which can enhance the ability of fusion network to extract and preserve spatial details and spectral information. Three groups of experiments conducted on four types of remote sensing images give evidence of that the proposed PRF-Net exhibits excellent performance in both reduced and full resolutions, showcasing its outstanding registration and fusion quality.
Zhangxi Xiong, Wei Li 0032, Xiaobin Zhao, Baochang Zhang 0001, Ran Tao 0003, Qian Du 0001
IEEE Trans. Neural Networks Learn. Syst.6
2024 Matrix Factorization Informed Interpretable Deep Network for Unregistered Hyperspectral and Multispectral Images Fusion
abstract
Considering the existing issues in unregistered hyper-spectral images (HSI) and multispectral images (MSI) fusion methods: i) the designed registration modules introduce a significant computational burden, and registration errors accumulate in fusion errors; ii) the methods lack model guidance, resulting in poor interpretability of the network. In this paper, we propose a matrix factorization informed interpretable deep network to address the challenges of unregistered HSI and MSI fusion (IUFNet). In particular, we derive an extended matrix factorization model for unregistered fusion (EUMF), which substitutes the abundance matrix of HSI containing low-resolution and distorted spatial information by the high-resolution abundance matrix of MSI. This substitution ingeniously eliminates the dependence of fusion performance on registration accuracy. Subsequently, IUFNet is designed to unfold the iterative results obtained by proximal gradient descent into the deep learning network, where each operation has a clear physical meaning. Overall, this network achieves the fusion of unregistered HSI and MSI and exhibits inter-pretability. Experimental results on the widely used Paiva Center dataset demonstrate the effectiveness and superiority of the proposed method.
Tongzhen Zhang, Jiahui Qu, Yunsong Li 0001, Qian Du 0001, Wenqian Dong
IGARSS4
2024 Degradation Aware Unfolding Network for Spectral Super-Resolution
abstract
Currently, leading methods for spectral super-resolution (SSR) depend heavily on constructing diverse network architectures in a heuristic manner, in order to learn a full mapping from the RGB image to its corresponding hyperspectral image (HSI). Despite promising results in reconstruction performance, significant challenges remain with respect to model interpretation and the capture of long-range dependencies. In response to these issues, based on a comprehensive exploration of the physical imaging mechanism between spectral response curve (SRC) and HSI, we have developed a novel model-driven degradation-aware unfolding network (DAUNet) in an iterative way. Besides, the learning process is explicitly integrated with the intrinsic generation mechanism of the SSR task. To be specific, we unfold each step into a degradation-aware gradient decent (DAGD) module and a proximal mapping module (PMM), using the framework of maximum a posteriori (MAP) theory. Additionally, to introduce more discriminative learning capabilities to our network, we have further enhanced the PMM architecture by incorporating a fine-grained multihead spectral-wise transformer (FMST) block, which improves global feature representation compared to the channel-wise transformer block. Extensive experiments over several spectral datasets finely demonstrate the superior performance of our method beyond the current representative state-of-the-art (SOTA) SSR methods.
Songcheng Du, Yihong Leng, Xinyi Liang, Jiaojiao Li 0001, Wei Liu 0004, Qian Du 0001
IEEE Geosci. Remote. Sens. Lett.6
2024 Hybrid Convolutional and Attention Network for Hyperspectral Image Denoising
abstract
Hyperspectral image (HSI) denoising is critical for the effective analysis and interpretation of hyperspectral data. However, simultaneously modeling global and local features is rarely explored to enhance HSI denoising. In this letter, we propose a hybrid convolution and attention network (HCANet), which leverages both the strengths of convolution neural networks (CNNs) and Transformers. To enhance the modeling of both global and local features, we have devised a convolution and attention fusion module aimed at capturing long-range dependencies and neighborhood spectral correlations. Furthermore, to improve multi-scale information aggregation, we design a multi-scale feed-forward network to enhance denoising performance by extracting features at different scales. Experimental results on mainstream HSI datasets demonstrate the rationality and effectiveness of the proposed HCANet. The proposed model is effective in removing various types of complex noise. Our codes are available at https://github.com/summitgao/HCANet.
Shuai Hu, Feng Gao 0005, Xiaowei Zhou 0003, Junyu Dong, Qian Du 0001
IEEE Geosci. Remote. Sens. Lett.5
2024 PCViT: A Pyramid Convolutional Vision Transformer Detector for Object Detection in Remote-Sensing Imagery
abstract
Remote sensing object detection (RSOD) is a fundamental and valuable task in Earth monitoring. However, remote sensing images (RSIs) are typically acquired from a bird’s eye perspective, resulting in intrinsic properties such as the complex backgrounds, random and dense distribution of objects, and multiscale objects. These properties hinder the direct application of well-performed detection methods in the natural images (NIs) domain to the RSIs domain, thereby limiting the attainment of desired performance. To address this, we propose a pyramid convolutional vision transformer (PCViT) that gets rid of the limitations of existing transformer methods. Firstly, we employ a pyramid architecture to effectively capture the multiscale information present in RSIs. To enhance the feature extraction capabilities of the transformer, we introduce a parallel convolution module (PCM) that complements the local information that may be missed by the transformer. Furthermore, we propose a self-supervised pretraining strategy called multi-perspective pretraining (MPP) to pretrain the model and subsequently finetune it on the downstream detection task. During the finetuning stage, we introduce a Local/globalk-NN attention (LGKA) to improve the token relationship establishment. In the neck part, we propose a feature-reflowing pyramid network (FRPN) to facilitate contextual information interaction and further enhance our PCViT’s ability to process multiscale information. Experimental results on two representative datasets, namely NWPU VHR-10 and DIOR, demonstrate the effectiveness of our PCViT, as it achieves outstanding performance. These results highlight the suitability of PCViT for RSOD tasks.
Jiaojiao Li 0001, Penghao Tian, Rui Song 0003, Yunsong Li 0001, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.6
2024 Feature Dimensionality Reduction With L2,p-Norm-Based Robust Embedding Regression for Classification of Hyperspectral Images
abstract
The curse of dimensionality and noise corruption are two tough problems that need to be solved in hyperspectral image (HSI) classification. However, the current feature dimensionality reduction methods, including both feature extraction and feature selection ones, cannot simultaneously solve the above two problems well. To address this issue, this paper proposes a novel method calledL2,p-norm-based robust embedding regression (L2,p-RER) for robust feature dimensionality reduction of HSI, which can effectively suppress the impact of noises and reduce the feature dimensions. Specifically,L2,p-RER first integrates projection learning with robust principle component analysis (RPCA) to remove noise in a low-dimensional space. Secondly, an embedding regression regularization is proposed to improve the discriminability of the extracted low-dimensional features. Thirdly, aL2,1-norm constraint is imposed to improve the interpretability of the learned projection matrix, which can jointly extract the key features from all bands with their physical meanings certainly preserved. Last but most important, theL2,p-norm that can adaptively balance the sparsity and the convexity is employed to model the noise and regression residual in the embedded low-dimensional space, which can further enhance the robustness and generalization of the proposed method. In addition, extensive experiments conducted on three benchmark HSI datasets validated the effectiveness of the proposed method.
Yangjun Deng, Menglong Yang, Heng-Chao Li 0001, Chen-Feng Long, Kui Fang, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.6
2024 Masked Self-Distillation Domain Adaptation for Hyperspectral Image Classification
abstract
Deep learning-based unsupervised domain adaptation (UDA) has shown potential in cross-scene hyperspectral image (HSI) classification. However, existing methods often experience reduced feature discriminability during domain alignment due to the difficulty of extracting semantic information from unlabeled target domain data. This challenge is exacerbated by ambiguous categories with similar material compositions and the underutilization of target domain samples. To address these issues, we propose a novel masked self-distillation domain adaptation (MSDA) framework, which enhances feature discriminability by integrating masked self-distillation (MSD) into domain adaptation. A class-separable adversarial training (CSAT) module is introduced to prevent misclassification between ambiguous categories by decreasing class correlation. Simultaneously, CSAT reduces the discrepancy between source and target domains through biclassifier adversarial training. Furthermore, the MSD module performs a pretext task on target domain samples to extract class-relevant knowledge. Specifically, MSD enforces consistency between outputs generated from masked target images, where spatial-spectral portions of an HSI patch are randomly obscured, and predictions are produced based on the complete patches by an exponential moving average (EMA) teacher. By minimizing consistency loss, the network learns to associate categorical semantics with unmasked regions. Notably, MSD is tailored for HSI data by preserving the samples’ central pixel and the object to be classified, thus maintaining class information. Consequently, MSDA extracts highly discriminative features by improving class separability and learning class-relevant knowledge, ultimately enhancing UDA performance. Experimental results on four datasets demonstrate that MSDA surpasses the existing state-of-the-art UDA methods for HSI classification. The code is available athttps://github.com/Li-ZK/MSDA-2024.
Zhuoqun Fang, Wenqiang He, Zhaokui Li, Qian Du 0001, Qiusheng Chen
IEEE Trans. Geosci. Remote. Sens.4
2024 Foundation Model-Based Multimodal Remote Sensing Data Classification
abstract
With the increasing availability and openness of remote sensing (RS) data collected from diverse sensors, there has been a growing interest in multimodal RS data classification. Nowadays, in the area of deep learning, there is a paradigm shift with the rise of foundation models, which are trained on large-scale datasets and are adaptable to a wide range of downstream tasks. In this study, the potential and effectiveness of foundation models for multimodal RS data classification is investigated. The training datasets of foundation models and multimodal RS datasets are quite different, and therefore, it is difficult to use a pretrained foundation model for multimodal RS data classification directly. To mitigate this difficulty, this article proposes a foundation model adaptation (FMA) framework for multimodal RS data classification without fine-tuning the parameters. Specifically, two learnable modules, i.e., cross-spatial interaction module and cross-channel interaction module, are proposed to add to the foundation model for extracting multimodal-specific representations. The cross-spatial and cross-channel interaction modules extract the characteristics of unimodal features along the spatial dimension and channel dimension, respectively. To effectively tackle the disparities among various RS modalities, an alignment approach (FMA2) is further explored based on the FMA. The FMA2 describes dependencies between different modalities by establishing a coupling score function, which can further enhance classification performance. To demonstrate the effectiveness and superiority of the FMA framework, comprehensive experiments are conducted on three multimodal RS datasets, showing improvement over the advanced multimodal RS data classification image methods.
Xin He 0004, Yushi Chen 0002, Lingbo Huang, Danfeng Hong, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.5
2024 Adversarial Domain Adaptation Network With Calibrated Prototype and Dynamic Instance Convolution for Hyperspectral Image Classification
abstract
Recently, the adversarial domain adaptation (ADA) methods have been widely investigated and applied in cross-domain hyperspectral image (HSI) classification. However, most ADA algorithms aim to align the cross-domain distribution without focusing on the class separability of the aligned target features and the information of samples within the domain. To address these issues, a new ADA framework based on calibrated prototype and dynamic instance convolution (CPDIC) is proposed in this paper for cross domain HSI classification. The CPDIC is composed of a generator, a calibrated discriminator and a classifier. The generator includes a static 3D convolutional network (SCN) and a dynamic instance convolutional network (DICN), where the SCN is used to extract coarse-grained features of HSI and the DICN can extract sample-specific fine-grained features using instance convolutions generated from dynamic instance convolution kernel generation (DCKG) module. As for the generator, the static and dynamic interactive feature extraction network extracts robust domain-invariant features with discriminability. The calibrated discriminator aligns the marginal distribution between domains and calibrate the predicted pseudo labels of target domain. For classification, a calibrated prototype loss (CPL) is introduced to align the class distribution across domains. The results of three cross-domain HSI classification tasks show that the proposed CPDIC outperforms existing unsupervised domain adaptation (UDA) algorithms.
Yi Huang 0021, Jiangtao Peng, Genwei Zhang, Weiwei Sun 0005, Na Chen 0008, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.6
2024 Residual Mask in Cascaded Convolutional Transformer for Spectral Reconstruction
abstract
A significant challenge of spectral reconstruction (SR) task is the lower performance reconstructed in foreground regions compared to background regions, which can be attributed to the marked difference in diversity of objects and disparity of adjacent scene characteristics. Moreover, the reconstruction of edge regions is often fraught with substantial errors due to the transitional nature of these regions, an issue conventional single convolutional neural networks (CNNs) and transformers struggle to handle. To address these challenges, we introduce the residual mask in cascaded convolutional transformer (RC2T) to iteratively improve the reconstruction of hyperspectral images (HSIs). Specifically, we propose a residual-predict mask generator (RMG) to generate a residual mask that retains band properties to separate feature with different complexities. Meanwhile, to achieve band expansion of mask features within the autoencoder, we approximate it to a Markov process and exploit the multistage spectral-aware Markov transfer (MMT) for its lightweight implementation. Next, we introduce the parallel convolutional multihead self-attention module (PSM), in which CNN runs parallel to the transformer to handle simple and complex features separately. Additionally, the residual mask loss function uses the established relationship between complexity of feature and reconstruction accuracy to generate residual mask in a self-supervised manner for providing complex high-frequency prior. We have validated our approach using three published datasets (NTIRE 2020 “Clean” track, NTIRE 2022, and CAVE). Additionally, we also conducted experiments with the proposed method on remote sensing dataset grss_dfc_2018 and a satellite-borne remote sensing dataset, achieving optimal performance. The experimental results demonstrate that our RC2T method is state-of-the-art (SOTA) in the field of SR.
Jiaojiao Li 0001, Shiyao Duan, Yihong Leng, Rui Song 0003, Yunsong Li 0001, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.6
2024 MMIF: Interpretable Hyperspectral and Multispectral Image Fusion via Maximum Mutual Information
abstract
Fusion-based hyperspectral image (HSI) super-resolution (SR) aims to recover high-resolution HSI (hr-HSI) from its two degraded modalities, that is, low-resolution HSI (lr-HSI) and high-resolution multispectral image (hr-MSI). The resulting HSI-SR image can be explained from two viewpoints, an inverse problem solution or a spatial–spectral information fusion product. Recent methods focusing on the former point are limited by demands of accurate degradation parameters and precise prior assumptions. On the contrary, the other research line avoids these drawbacks and offers a more flexible design. However, recent methods implicitly handle information fusion. The interaction between lr-HSI and hr-MSI only lies in the extracted feature domain, not directly on the data. Moreover, the proposed modules in recent methods promote information fusion from the perspective of deep learning and neglect the specialty of the HSI domain, leading to weak interpretability and poor reliability in practice. Considering the essence of the HSI-SR problem and the inherent property of HSIs, in this article, we propose a maximum mutual information (MMI) strategy. At first, we model the HSI-SR problem in a new variance inference (VI) architecture. This new VI model simulates the physical process of HSI pairs imaging and provides convenience for the MMI strategy. Then, we insert the MMI strategy in the VI model. The MMI strategy promotes information fusion with quantitative constraints on the information interaction between lr-HSI and hr-MSI. Finally, we implement the VI architecture with a neural network. Benefitting from the MMI strategy, a simple network structure can achieve efficient fusion performance, which indicates that MMI frees DL-based HSI-SR methods from the complicated structure design. The experimental results on synthetic and real datasets demonstrate the superiority of our method to state-of-the-art methods in terms of effectiveness, generalization, and interpretability.
Yunsong Li 0001, Wen-jin Guo, Weiying Xie, Tao Jiang 0031, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.5
2024 IDA-SiamNet: Interactive- and Dynamic-Aware Siamese Network for Building Change Detection
abstract
Building change detection (BCD) is a critical task in remote sensing which aims to identify the building changes within the same geographical area over time. The complexity of BCD is heightened when utilizing very high-resolution (VHR) remote sensing images, leading to two primary challenges: distinguishing between building and nonbuilding changes and accommodating the diverse range of building shapes and sizes. The existing mainstream methods neglect interactions between encoders, thereby compromising the ability to recognize building and nonbuilding changes. Additionally, most BCD methods overlook feature alignment and fusion which hinders the precise extraction of buildings with varying shapes and sizes. To address these limitations, we propose an interactive- and dynamic-aware Siamese network (IDA-SiamNet) for BCD. Our method comprises the spatial exchange feature interaction (SEFI) module, the channel exchange feature interaction (CEFI) module, and the dynamic-deformable dual-alignment fusion (D3AF) module. The SEFI and CEFI modules play a pivotal role in facilitating mutual information exchange between Siamese encoders, enhancing discrimination between building and nonbuilding changes. Furthermore, the D3AF module dynamically aggregates multiple parallel convolutional kernels to improve feature alignment and fusion for accurate building outline extraction. D3AF adapts its receptive field (RF) based on object size and covers diverse building shapes without introducing excessive background information. Experimental evaluations on three widely used BCD datasets, learning, vision, and remote sensing change detection (LEVIR-CD), satellite side-looking (S2Looking), and WHU BCD (WHU-CD), demonstrate the superior performance of our proposed method over state-of-the-art alternatives. Code and weights are made available athttps://github.com/SUPERMAN123000/IDA-SiamNet.
Yun-Cheng Li, Sen Lei, Nanqing Liu, Heng-Chao Li 0001, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.5
2024 HyperMLP: Superpixel Prior and Feature Aggregated Perceptron Networks for Hyperspectral and LiDAR Hybrid Classification
abstract
Hyperspectral images have excellent spectral combining capabilities and LiDAR images have fine stereooscopic elevation information. Therefore, the multi-modal fusion classification of hyperspectral and LiDAR images is inevitably improves the interpretation ability of remote sensing images. In recent years, the MLP-Mixer, an image processing network based on MLP, has flourished in the field of image processing. In this work, we propose an innovative HyperMLP network based on the deep learning framework MLP-Mixer architecture to address the lack of spatial feature construction capability and the locality of multi-modal feature fusion in naive networks. Specifically,(1) The adoption of unsupervised superpixel embedding provides additional shallow morphological spatial feature information for the network, reduces the pressure of the feature extraction network, and enhances feature discrimination capabilities. (2) The feature scrambling strategy improves the diversity of features and strengthens generalization of the network by enhancing interactions between different spatial features. (3) By implementing the bilateral modulation strategy, feature fusion is applied at every stage of the deep network, reducing semantic drift between features. On three fiducial remote sensing datasets, classification tests are performed on the proposed HyperMLP network to verify its performance, and the results are definitely impressive.
Jiaojiao Li 0001, Rui Song 0003, Wei Li 0032, Yunsong Li 0001, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.6
2024 Cross-Domain Few-Shot Hyperspectral Image Classification With Cross-Modal Alignment and Supervised Contrastive Learning
abstract
Recently, metric-based few-shot learning (FSL) methods have achieved good performance in hyperspectral image (HSI) classification. However, existing methods suffer from two problems: over-reliance on image modality information leads to inaccurate prototype representation, where a prototype refers to the centroid of each class in the dataset, and the impact of redundant and noisy pixels on model discriminability is rarely considered. These problems result in insufficient discriminability of the model for the target domain. To address the above issues, we propose a cross-domain few-shot HSI classification framework with cross-modal alignment and supervised contrastive learning (CDFS-CASCL). It is well known that human visual learning greatly benefits from the input of various modal information such as vision, language and video. Inspired by the way humans abstract image class concepts in language form and understand the essence of classes, we perform cross-modal alignment (CA) between similar image and text prototypes, and use abstract text semantics to guide the model to learn semantic related features with good generalization ability in images, so as to improve the accuracy of image prototypes representation of the prototypes. In addition, through supervised contrastive learning (SCL) based on neighborhood pixel mask in the target domain, the enhanced sample features belonging to the same class are closer, while the enhanced sample features belonging to different classes are pulled further, enabling the model to learn mask-robust discriminative feature representations, suppressing the negative impact of redundant and noisy pixels, and improving the model’s discriminability. The experimental results demonstrate the superiority of the proposed CDFS-CASCL. The code is available at https://github.com/Li-ZK/CDFS-CASCL-2024.
Zhaokui Li, Yan Wang 0087, Wei Li 0032, Qian Du 0001, Zhuoqun Fang, Yushi Chen 0002
IEEE Trans. Geosci. Remote. Sens.5
2024 Semi-Supervised Dynamic Ensemble Learning With Balancing Diversity and Consistency for Hyperspectral Image Classification
abstract
Hyperspectral coastal wetland classification requires an extensive quantity of labeled samples, which are hard to acquire. Therefore, a novel semi-supervised dynamic ensemble learning (SSDEL) framework is proposed to overcome the limitations of labeled samples in wetland hyperspectral classification. Firstly, a collaborative relationship is established between labeled and unlabeled samples in the sample augmentation stage. Based on this relationship, unlabeled samples were assigned to the region to which the most similar samples belonged. Then, multiple classifiers are trained using labeled samples and predict unlabeled samples in the same region to obtain higher confidence pseudo-label results. Secondly, based on the assumption that different classifiers should produce similar classification results for a specific target sample, an objective function is designed to unify the classification behavior of multiple classifiers. The representation coefficients of multiple classifiers in the same region are constrained by optimizing the objective function through thel2norm. Finally, a complete SSDEL framework is constructed by applying consistency learning again to the augmented samples. The proposed method is evaluated using three wetland hyperspectral images of China, and the experiments results demonstrate its effectiveness.
Hongjun Su, Hengyi Zheng, Zhaohui Xue, Weiwei Sun 0005, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.8
2024 Self-Training-Based Unsupervised Domain Adaptation for Object Detection in Remote Sensing Imagery
abstract
We propose a novel two-stage cross-domain self-training (CDST) framework for unsupervised domain adaptive object detection in remote sensing. The first stage introduces the generative adversarial network (GAN)-based domain transfer strategy to preliminarily mitigate the domain shift for higher quality initial pseudo-labeled images, which utilizes the CycleGAN to transfer source-domain images to match the target domain. Moreover, the key issue in tailoring the self-training (ST) to unsupervised domain adaptive detection lies in the quality of pseudo-labeled images. To select high-quality pseudo-labeled images under the domain-shift circumstance, we propose hard example selection-based self-training (HES-ST) with the three key steps: 1) detector-based example division (DED), which divides the detected examples into easy examples and hard ones according to their confidence level; 2) confidence and relation joint score (CRJS)-based hard example selection, which combines two reliability levels calculated, respectively, by the detector and relation network (RN) module to mine reliable examples; and 3) union example (UE)-based training image selection, which combines both easy and reliable hard examples to choose target-domain images that may contain fewer detection errors. The experimental results on several remote sensing datasets demonstrate the effectiveness of our proposed framework. Compared with the baseline detector trained on the source dataset, our approach consistently improves the detection performance on the target dataset by 15.7%–16.8% mean average precision (mAP) and achieves the state-of-the-art (SOTA) results under various domain adaptation scenarios.
Sihao Luo, Li Ma 0005, Xiaoquan Yang, Dapeng Luo, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.5
2024 Physical Knowledge Analytic Framework for Sea Surface Temperature Prediction
abstract
Recently, the methods that combine the merits of the numerical model and the deep learning to improve the prediction accuracy of the sea surface temperature (SST) have received considerable attention. Existing methods usually apply the output of the numerical model as the physical knowledge to guide the training of the deep learning models. However, the physical knowledge in the observed data has not been fully exploited. With the development of observational instruments and techniques, an increasing amount of observational data has been collected. These data can be utilized for the exploration of physical knowledge. Toward this end, we propose novel scheme for SST prediction, which applies generative adversarial networks (GANs) to analyze the physical knowledge in the historical data. In particular, two GAN models are trained with numerical model data and observed data separately. Afterward, the physical knowledge is extracted from the observed data which is not contained in the data generated by the numerical model by comparing the learned physical feature from the two pre-trained GAN models. Finally, to validate the relevance of the physical knowledge which we have discovered, the extracted features are added into the numerical model data which are called newly corrected data. Besides, we train two spatial-temporal models over the newly corrected dataset and the original numerical model data for SST prediction, respectively. The experimental results show that the newly corrected dataset performs better than using the original numerical model for SST prediction.
Yuxin Meng 0003, Feng Gao 0005, Eric Rigall, Junyu Dong, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.5
2024 Domain Invariant and Compact Prototype Contrast Adaptation for Hyperspectral Image Classification
abstract
Contrastive learning achieves good performance on hyperspectral image classification (HSIC), but its application on cross-scene classification is still challenging due to domain shift. The emergence of domain adaptation (DA) techniques can reduce domain discrepancy and transfer a model between two domains. Recently, instance-level contrast adaptation methods can connect two related domains, and domain-invariant features are extracted. However, it is sensitive to noisy samples and only learns low-level discriminative features. To solve these problems, a novel domain invariant and compact prototype contrast adaptation (DIC-proCA) framework is proposed for HSIC. About the proposed DIC-proCA, the prototype is introduced into the contrastive learning framework, which serves as a representative embedding of semantically similar samples, has class representativeness and can alleviate the negative impact of outliers. Taking into account the class representativeness of the prototype and the discriminability of the sample itself, a bidirectional inter-domain instance-to-prototype contrastive loss is proposed. It explicitly expresses feature relationships between categories in different domains, and then extracts domain-invariant features. Meanwhile, the mining of compact discriminative features within the target domain is facilitated by instance-level contrastive learning after data augmentation. In addition, the strategy of label smoothing promotes the clusters in the domain to be more compact and evenly separated, making the model more generalizable. Three cross-scene HSIC tasks demonstrate that the proposed DIC-proCA exhibits superior performance compared to some advanced DA algorithms.
Yujie Ning, Jiangtao Peng, Quanyong Liu, Weiwei Sun 0005, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.5
2024 Feature Mutual Representation-Based Graph Domain Adaptive Network for Unsupervised Hyperspectral Change Detection
abstract
Recently, deep neural networks (DNNs) have been widely used in hyperspectral image change detection (HSI-CD). Generally, training such a DNN-based HSI-CD network often requires a large number of labeled training samples. However, it is time-consuming, labor-intensive, or even infeasible to label training samples in practice. In this article, we propose a feature mutual representation-based graph domain adaptive network (FGDANet) for unsupervised HSI-CD. This method constructs a pseudosiamese backbone consisting of two customized unsupervised learning domains, which can make full use of the information from different domains through the graph domain adaptation strategy to improve the feature expression capability and generalization. There are three key characteristics: first, in each customized unsupervised learning domain, a graph convolutional network (GCN)-based difference feature extraction architecture is designed to model the local and global dependence among the features of multitemporal HSIs; second, a progressive graph-to-pixel joint constraint strategy (PJCS) is proposed to provide the high-confidence training sample labels for the unsupervised learning of the network in each domain; and third, the homogeneous mutual representation joint graph feature alignment (HJGFA) module of the graph domain adaptation strategy can make full use of the difference features from the two domains through the information interaction to facilitate the model to capture the changed and unchanged essential characteristics. The experimental results on four HSI datasets demonstrate the superiority of the proposed FGDANet. Code is available athttps://github.com/Jiahuiqu/FGDANet.
Jiahui Qu, Jingyu Zhao 0011, Wenqian Dong, Song Xiao 0001, Yunsong Li 0001, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.6
2024 Iterative Semi-Supervised Learning With Few-Shot Samples for Coastal Wetland Land Cover Classification
abstract
A novel approach is proposed in this study that combines superpixel (SP) segmentation and multiclassifier ensemble learning (EL) to address the limited availability of labeled samples in coastal wetland land cover classification. First, the SP segmentation technique is employed to partition unknown samples into multiple homogeneous regions, thereby facilitating the effective capture of spatial information pertaining to land cover. Subsequently, a multiclassifier EL strategy is employed within these regions to process the samples, effectively leading to a reduction in classification errors and an improvement in accuracy. To enhance the performance of semi-supervised learning (SSL), a sample iteration selection metric is introduced to optimize the training samples based on the consistency of sample types within homogeneous regions and the results obtained from the multiclassifier ensemble, thus enhancing the reliability of pseudo-labels. Additionally, multiscale SP segmentation is utilized to augment the ensemble strategy for samples in order to reduce the necessity for hyperparameter adjustments and increase the automation and reliability of the model. Overall, the accuracy of coastal wetland classification is improved by this approach while simultaneously mitigating the complexity of SSL in terms of hyperparameter tuning. The effectiveness of the proposed approach has been assessed through experiments conducted on three GF-5 hyperspectral images of coastal wetlands in China. In particular, the proposed methods provide superior performance compared with the state-of-the-art classification methods.
Hongjun Su, Hengyi Zheng, Zhaohui Xue, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.6
2024 Transformer-Based Band Regrouping With Feature Refinement for Hyperspectral Object Tracking
abstract
Hyperspectral videos (HSVs) offer not only spatial information but also diagnostic spectral features. Due to the fact that spectral features are only related to the material of the object, this advantage can address the issue of RGB video tracking failure when the object and background are visually similar. However, the effectiveness of deep learning models is limited due to insufficient HSV training data. Existing methods tend to divide a hyperspectral image (HSI) into several three-channel false-color images to leverage the existing RGB trackers for transfer learning. Nonetheless, these methods lack adequate exploration of band interrelations and overlook correlation among objects prior to similarity calculation. In this article, a transformer-based band regrouping and feature refinement network (TBR-Net) is introduced, which is specifically tailored for hyperspectral object tracking. To maximize the potential of the RGB tracker and enhance the use of available training data, we propose a transformer-based band regrouping (TBR) method. By modeling long-range spectral dependencies, the inherent context information among bands is captured, which is subsequently utilized to reorganize bands into several false-color images. Furthermore, to combine the relationship of the template and the search (T & S) frames into a correlation calculation, a feature refinement module (FRM) is designed. The cross-attention mechanism enables mutual relation modeling, allowing similar regions to be perceived and form discriminative feature representation. As a result, a hyperspectral tracker can be efficiently trained via transfer learning to address the data insufficiency challenge, while the mutual perception between objects further enhances the tracking performance. Its effectiveness is validated by extensive benchmark experiments, which demonstrate that the TBR-Net surpasses state-of-the-art methods.
Hanzheng Wang, Wei Li 0032, Xiang-Gen Xia 0001, Qian Du 0001, Jing Tian 0003, Qing Shen 0002
IEEE Trans. Geosci. Remote. Sens.4
2024 Bridging CNN and Transformer With Cross-Attention Fusion Network for Hyperspectral Image Classification
abstract
Feature representation is crucial for hyperspectral image (HSI) classification. However, existing convolutional neural network (CNN)-based methods are limited by the convolution kernel and only focus on local features, which causes it to ignore the global properties of HSIs. Transformer-based networks can make up for the limitations of CNNs because they emphasize the global features of HSIs. How to combine the advantages of these two networks in feature extraction is of great importance in improving classification accuracy. Therefore, a cross-attention fusion network bridging CNN and Transformer (CAF-Former) is proposed, which can fully utilize the advantages of CNN in local features and Transformer’s long time-dependent feature learning for hyperspectral classification. In order to fully explore the local and global information within an HSI, a Dynamic-CNN branch is proposed to effectively encode local features of pixels, while a Gaussian Transformer branch is constructed to accurately model the global features and long-range dependencies. Moreover, in order to fully interact with local and global features, a cross-attention fusion (CAF) module is proposed as a bridge to fuse the features extracted by the two branches. Experiments over several benchmark datasets demonstrate that the proposed CAF-Former significantly outperforms both CNN-based and Transformer-based state-of-the-art networks for HSI classification.
Fulin Xu, Shaohui Mei, Ge Zhang 0006, Nan Wang 0026, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.5
2024 Multifrequency Graph Convolutional Network With Cross-Modality Mutual Enhancement for Multisource Remote Sensing Data Classification
abstract
The mining of meaningful features and effective fusion of multisource remote sensing (RS) data have always been the challenging research problems in the joint classification of hyperspectral image (HSI) and light detection and ranging (LiDAR) data. In this paper, we propose a Multi-Frequency Graph Convolutional Network with Cross-modality Mutual Enhancement (MFGCN-CME) for multisource RS data classification. Specifically, we design an adaptive multi-frequency graph feature learning module to capture the low- and high-frequency multiscale features of HSI and LiDAR in parallel and further adaptively aggregate them. Then, we propose a bipartite graph enhancement learning module to obtain the spatial-enhanced HSI features and spectral-enhanced LiDAR features by propagating inter-modality information. To the best of our knowledge, the bipartite graph is first used to multisource RS data classification task. Furthermore, compared with traditional fusion methods, a gated fusion module is used to fully explore the complementarity of two data sources. Finally, a joint loss function combing a classification loss and a semi-supervised contrastive loss is developed to improve the model robustness. Comprehensive experiments on different HSI and LiDAR datasets demonstrate that our proposed method can yield better performance compared with several state-of-the-art multisource RS data classification methods.
Jin-Yu Yang, Heng-Chao Li 0001, Lei Pan 0003, Qian Du 0001, Antonio Plaza
IEEE Trans. Geosci. Remote. Sens.5
2024 Self-Paced Probabilistic Collaborative Representation for Anomaly Detection of Hyperspectral Images
abstract
In recent years, hyperspectral anomaly detection methods based on representation models has attracted much attention. However, when the dictionary is polluted by anomalous pixels, their performance is greatly affected. To adjust the contributions of different dictionary atoms, traditional methods usually predefine a distance weighting matrix and impose it on the dictionary matrix or coefficient vector, which may not be accurate enough. To solve this problem, a self-paced probabilistic collaborative representation detector (SP-ProCRD) is proposed in this article. It assigns weights for each atom loss term according to the probability that the pixel under test (PUT) belongs to the same class as each dictionary atom. Unlike the predefined weight matrix approach, a self-paced learning (SPL) strategy is used for iterative optimization, so that dictionary atoms participate in the representation from "good" to "bad" ones when solving the model. The representation residuals are utilized to accelerate the convergence. The proposed model can optimally represent each PUT using similar dictionary atoms and minimize the negative impact caused by anomalous atoms contained in the dictionary. In terms of weighting for SPL, an adaptive weighting scheme based on the polynomial self-paced (SP) regularizer is proposed to address the generalization issues of most previous weighting schemes. This scheme improves the generalization and automation of the model. Experimental results reveal that the proposed method produces more accurate result than existing methods and runs efficiently.
Chendi Zhang, Hongjun Su, Zhaoyue Wu, Zhaohui Xue, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.7
2024 MarsMapNet: A Novel Superpixel-Guided Multiview Feature Fusion Network for Efficient Martian Landform Mapping
abstract
Landform classification and mapping of the Martian surface using Mars orbiter images can provide an important reference for landing site selection and rovers’ traversability evaluation in Mars exploration. Moreover, specific Martian landforms are closely associated with the evidences of water-related activities and Martian life, thus have crucial research importance. This article proposes a novel superpixel-guided multiview feature fusion network (MarsMapNet) for efficient mapping of the Martian landforms. In particular, the proposed MarsMapNet first generates the superpixel-level segments from Mars orbiter images by considering local morphological homogeneity of landforms. Then, a multiview feature extraction and fusion (MVF) network is developed, where abstract convolutional features are extracted based on scene-level patches, and multitextures are extracted based on local landform from shallow-to-deep feature learning. After the network being trained on scene-level samples and guided by the superpixel segmentation, Martian landforms can be correctly classified in an efficient way, whose mapping time cost sharply decreased when compared to the reference methods. The proposed MarsMapNet has been validated on three real landing sites from several Mars missions (i.e., the Jezero Crater, the Southern Utopia Planitia, and the Oxia Planum) by using the Mars Reconnaissance Orbiter’s Context Camera (CTX) images. Qualitative and quantitative analyses on the obtained experimental results confirm the effectiveness and efficiency of the proposed MarsMapNet when compared with the state-of-the-art (SOTA) methods, demonstrating its potential for supporting a Martian global landform mapping in the future.
Sicong Liu 0001, Xiaohua Tong, Qian Du 0001, Lorenzo Bruzzone, Kecheng Du, Jie Zhang 0117, Xuanning Lu
IEEE Trans. Geosci. Remote. Sens.4
2024 Graph Convolutional Network With Relaxed Collaborative Representation for Hyperspectral Image Classification
abstract
Graph convolutional networks (GCNs) have been skillfully employed in hyperspectral image (HSI) classification, exhibiting remarkable performance owing to their unique superiority in handling non-Euclidean graph-structured data. However, the inherent absence of predefined connections between pixels in HSI results in the underutilization of the structural and attribute information of the graph edges. Furthermore, the construction of adjacency matrices for large-scale HSI data imposes a huge computational burden on traditional GCNs. Therefore, in this article, a novel method combining relaxed collaborative representation (RCR) and GCN (RCR-GCN) for hyperspectral classification is proposed. Specifically, RCR is adopted to compute the representation coefficients of each feature, reflecting the similarity and diversity among different sample features. Meanwhile, the representation coefficients are applied as edge attributes in the graph, denoting the weights of the connections between neighboring nodes. After that, GCN is employed to classify the graph nodes. Moreover, an efficient version of the RCR-GCN method is developed to boost the computation, which constructs the graph based on superpixel nodes instead of the pixel nodes by using simple linear iterative clustering (SLIC). Extensive experiments on three HSI image datasets demonstrate that the proposed method outperforms other state-of-the-art methods and achieves more efficiency and feasibility in HSI image classification.
Hengyi Zheng, Hongjun Su, Zhaoyue Wu, Mercedes Eugenia Paoletti, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.5
2024 Oriented Object Detection for Remote Sensing Images via Object-Wise Rotation-Invariant Semantic Representation
abstract
Oriented object detection (OOD) in remote sensing images (RSIs) remains a challenging work due to an arbitrary orientation of instance. Learning rotation-invariant features is critical in modeling a fixed descriptor for instances with its rotated variants. However, most existing methods construct the descriptor from the perspectives of data or feature augmentation, but ignore the exploration of potentially useful supervision information inside the detection algorithm. In this paper, we propose an object-wise rotation-invariant semantic representation (ORSR) framework, which synergizes the exploration of latent supervision, rotation-invariant learning, and guided attention mechanism into a unified network to boost the performance of OOD in RSIs. First, supervised by our constructed pseudo ground truth of segmentation masks, a semantic segmentation branch is built along with the detection algorithm to refine the representation of backbone features. Moreover, a consistency loss function is proposed to encourage the segmentation branch to make the fixed predictions for backbone features with its rotated variants. Considering that segmentation predictions remain the same affine transformations before and after rotating, we further construct a Kullback-Leibler (KL) Divergence based similarity loss function for encouraging the network to model the rotation-invariant features. Finally, we separate the ”object” descriptor from the segmentation predictions to extend the implicit constraint in our proposed semantic segmentation branch. The separated ”object” descriptor not only involves the spatial regularizer to emphasize the high-responsive regions in image, but also can be guided by the constructed consistency loss function. We evaluate our proposed ORSR on the challenging DOTA, DIOR-R, and HRSC2016 datasets. Extensive experiments demonstrate that the proposed ORSR achieves competitive performance compared to other single-scale and multi-scale detection methods.
Shangdong Zheng, Zebin Wu 0001, Qian Du 0001, Yang Xu 0006, Zhihui Wei
IEEE Trans. Geosci. Remote. Sens.3
2024 SWFormer: Stochastic Windows Convolutional Transformer for Hybrid Modality Hyperspectral Classification
abstract
Joint classification of hyperspectral images with hybrid modality can significantly enhance interpretation potentials, particularly when elevation information from the LiDAR sensor is integrated for outstanding performance. Recently, the transformer architecture was introduced to the HSI and LiDAR classification task, which has been verified as highly efficient. However, the existing naive transformer architectures suffer from two main drawbacks: 1) Inadequacy extraction for local spatial information and multi-scale information from HSI simultaneously. 2) The matrix calculation in the transformer consumes vast amounts of computing power. In this paper, we propose a novel Stochastic Window Transformer (SWFormer) framework to resolve these issues. First, the effective spatial and spectral feature projection networks are built independently based on hybrid-modal heterogeneous data composition using parallel feature extraction, which is conducive to excavating the perceptual features more representative along different dimensions. Furthermore, to construct local-global nonlinear feature maps more flexibly, we implement multi-scale strip convolution coupled with a transformer strategy. Moreover, in an innovative random window transformer structure, features are randomly masked to achieve sparse window pruning, alleviating the problem of information density redundancy, and reducing the parameters required for intensive attention. Finally, we designed a plug-and-play feature aggregation module that adapts domain offset between modal features adaptively to minimize semantic gaps between them and enhance the representational ability of the fusion feature. Three fiducial datasets demonstrate the effectiveness of the SWFormer in determining classification results.
Jiaojiao Li 0001, Rui Song 0003, Yunsong Li 0001, Qian Du 0001
IEEE Trans. Image Process.6
2024 SCFormer: Spectral Coordinate Transformer for Cross-Domain Few-Shot Hyperspectral Image Classification
abstract
Cross-domain (CD) hyperspectral image classification (HSIC) has been significantly boosted by methods employing Few-Shot Learning (FSL) based on CNNs or GCNs. Nevertheless, the majority of current approaches disregard the prior information of spectral coordinates with limited interpretability, leading to inadequate robustness and knowledge transfer. In this paper, we propose an asymmetric encoder-decoder architecture, Spectral Coordinate Transformer (SCFormer), for the CDFSL HSIC task. Several dense Spectral Coordinate blocks (SC blocks) are embedded in the backbone of the encoder to establish feature representation with better generalization, which integrates spectral coordinates via Rotary Position Embedding (RoPE) to minimize spectral position disturbance caused by the convolution operation. Due to a large amount of hyperspectral image data and the high demand for model generalization ability in cross-domain scenarios, we design two mask patterns (Random Mask and Sequential Mask) built on unexploited spectral coordinates within the SC blocks, which are unified with the asymmetric structure to learn high-capacity models efficiently and effectively with satisfactory generalization. Besides, from the perspective of the loss function, we devise an intra-domain loss function founded on the Orthogonal Complement Space Projection (OCSP) theory to facilitate the aggregation of samples in the metric space, which promotes intra-domain consistency and increases interpretability. Finally, the strengthened class expression capacity of the intra-domain loss function contributes to the inter-domain loss function constructed by Wasserstein Distance (WD) for realizing domain alignment. Experimental results on four benchmark data sets demonstrate the superiority of the SCFormer.
Jiaojiao Li 0001, Rui Song 0003, Yunsong Li 0001, Qian Du 0001
IEEE Trans. Image Process.5
2024 Graph Embedding Interclass Relation-Aware Adaptive Network for Cross-Scene Classification of Multisource Remote Sensing Data
abstract
The unsupervised domain adaptation (UDA) based cross-scene remote sensing image classification has recently become an appealing research topic, since it is a valid solution to unsupervised scene classification by exploiting well-labeled data from another scene. Despite its good performance in reducing domain shifts, UDA in multisource data scenarios is hindered by several critical challenges. The first one is the heterogeneity inherent in multisource data complicates domain alignment. The second challenge is the incomplete representation of feature distribution caused by the neglect of the contribution from global information. The third challenge is the inaccuracies in alignment due to errors in establishing target domain conditional distributions. Since UDA does not guarantee the complete consistency of the distribution of the two domains, networks using simple classifiers are still affected by domain shifts, resulting in poor performance. In this paper, we propose a graph embedding interclass relation-aware adaptive network (GeIraA-Net) for unsupervised classification of multi-source remote sensing data, which facilitates knowledge transfer at the class level for two domains by leveraging aligned features to perceive inter-class relation. More specifically, a graph-based progressive hierarchical feature extraction network is constructed, capable of capturing both local and global features of multisource data, thereby consolidating comprehensive domain information within a unified feature space. To deal with the imprecise alignment of data distribution, a joint de-scrambling alignment strategy is designed to utilize the features obtained by a three-step pseudo-label generation module for more delicate domain calibration. Moreover, an adaptive inter-class topology based classifier is constructed to further improve the classification accuracy by making the classifier domain adaptive at the category level. The experimental results show that GeIraA-Net has significant advantages over the current state-of-the-art cross-scene classification methods.
Song Xiao 0001, Jiahui Qu, Wenqian Dong, Qian Du 0001, Yunsong Li 0001
IEEE Trans. Image Process.5
2024 HTD-TS3: Weakly Supervised Hyperspectral Target Detection Based on Transformer via Spectral-Spatial Similarity
abstract
As an advanced technique in remote sensing, hyperspectral target detection (HTD) is widely concerned in civilian and military applications. However, the limitation of prior and heterogeneous backgrounds makes HTD models sensitive to data corruption under various interference from the environment. In this article, a novel united HTD framework based on the concept of transformer is proposed to extract [HTD based on transformer via spectral-spatial similarity (HTD-TS3)] under weak supervision, which opens up more flexible ways to study HTD. For the first time, the transformer mechanism is introduced into the HTD task to extract spectral and spatial features in a unified optimization procedure. By modeling long-range dependence among spectra, it realizes spectral-spatial joint inference based on long-range context, which addresses the issues of insufficient utilization of spatial information. To provide samples for weakly supervised learning (WSL), the coarse sample selection and spectral sequence construction in an efficient way are proposed, which makes full use of limited prior information. Finally, an exponential constrained nonlinear function is adopted to acquire pixel-level prediction via combining discriminative spectral-spatial features and coarse spatial information. Experiments on real hyperspectral images (HSIs) captured by different sensors at various scenes verify the effectiveness and efficiency of HTD-TS3.
Weiying Xie, Yunsong Li 0001, Qian Du 0001
IEEE Trans. Neural Networks Learn. Syst.4
2024 A Spatio-Spectral Fusion Method for Hyperspectral Images Using Residual Hyper-Dense Network
abstract
Spatio-spectral fusion of panchromatic (PAN) and hyperspectral (HS) images is of great importance in improving spatial resolution of images acquired by many commercial HS sensors. DenseNets have recently achieved great success for image super-resolution because they facilitate gradient flow by concatenating all the feature outputs in a feedforward manner. In this article, we propose a residual hyper-dense network (RHDN) that extends the DenseNet to solve the spatio-spectral fusion problem. The overall structure of the proposed RHDN method is a two-branch network, which allows the network to capture the features of HS images within and outside the visible range separately. At each branch of the network, a two-stream strategy of feature extraction is designed to process PAN and HS images individually. A convolutional neural network (CNN) with cascade residual hyper-dense blocks (RHDBs), which allows direct connections between the pairs of layers within the same stream and those across different streams, is proposed to learn more complex combinations between the HS and PAN images. The residual learning is adopted to make the network efficient. Extensive benchmark evaluations well demonstrate that the proposed RHDN fusion method yields significant improvements over many widely accepted state-of-the-art approaches.
Jiahui Qu, Zhangchun Xu, Wenqian Dong, Song Xiao 0001, Yunsong Li 0001, Qian Du 0001
IEEE Trans. Neural Networks Learn. Syst.6
2024 Representation-Enhanced Status Replay Network for Multisource Remote-Sensing Image Classification
abstract
Deep-learning-based methods are widely used in multisource remote-sensing image classification, and the improvement in their performance confirms the effectiveness of deep learning for classification tasks. However, the inherent underlying problems of deep-learning models still hinder the further improvement of classification accuracy. For example, after multiple rounds of optimization learning, representation bias and classifier bias are accumulated, which prevents the further optimization of network performance. In addition, the imbalance of fusion information among multisource images also leads to insufficient information interaction throughout the fusion process, thus making it difficult to fully utilize the complementary information of multisource data. To address these issues, a Representation-enhanced Status Replay Network (RSRNet) is proposed. First, a dual augmentation including modal augmentation and semantic augmentation is proposed to enhance the transferability and discreteness of feature representation, to reduce the impact of representation bias in the feature extractor. Then, to alleviate the classifier bias and maintain the stability of the decision boundary, a status replay strategy (SRS) is built to regulate the learning and optimization of the classifier. Finally, aiming to improve the interactivity of modal fusion, a novel cross-modal interactive fusion (CMIF) method is employed to jointly optimize the parameters of different branches by combining multisource information. Quantitative and qualitative results on three datasets demonstrate the superiority of RSRNet in multisource remote-sensing image classification, and its outperformance compared with other state-of-the-art methods.
Wei Li 0032, Yinjian Wang, Ran Tao 0003, Qian Du 0001
IEEE Trans. Neural Networks Learn. Syst.5
2024 HPRN: Holistic Prior-Embedded Relation Network for Spectral Super-Resolution
abstract
Spectral super-resolution (SSR) refers to the hyperspectral image (HSI) recovery from an RGB counterpart. Due to the one-to-many nature of the SSR problem, a single RGB image can be reprojected to many HSIs. The key to tackle this ill-posed problem is to plug into multisource prior information such as the natural spatial context prior of RGB images, deep feature prior, or inherent statistical prior of HSIs so as to effectively alleviate the degree of ill-posedness. However, most current approaches only consider the general and limited priors in their customized convolutional neural networks (CNNs), which leads to the inability to guarantee the confidence and fidelity of reconstructed spectra. In this article, we propose a novel holistic prior-embedded relation network (HPRN) to integrate comprehensive priors to regularize and optimize the solution space of SSR. Basically, the core framework is delicately assembled by several multiresidual relation blocks (MRBs) that fully facilitate the transmission and utilization of the low-frequency content prior of RGBs. Innovatively, the semantic prior of RGB inputs is introduced to mark category attributes, and a semantic-driven spatial relation module (SSRM) is invented to perform the feature aggregation of clustered similar ranges for refining recovered characteristics. In addition, we develop a transformer-based channel relation module (TCRM), which breaks the habit of employing scalars as the descriptors of channelwise relations in the previous deep feature prior and replaces them with certain vectors to make the mapping function more robust and smoother. In order to maintain the mathematical correlation and spectral consistency between hyperspectral bands, the second-order prior constraints (SOPCs) are incorporated into the loss function to guide the HSI reconstruction. Finally, extensive experimental results on four benchmarks demonstrate that our HPRN can reach the state-of-the-art performance for SSR quantitatively and qualitatively. Furthermore, the effectiveness and usefulness of the reconstructed spectra are verified by the classification results on the remote sensing dataset. Codes are available at https://github.com/Deep-imagelab/HPRN.
Chaoxiong Wu, Jiaojiao Li 0001, Rui Song 0003, Yunsong Li 0001, Qian Du 0001
IEEE Trans. Neural Networks Learn. Syst.5
2024 Graph Information Aggregation Cross-Domain Few-Shot Learning for Hyperspectral Image Classification
abstract
Most domain adaptation (DA) methods in cross-scene hyperspectral image classification focus on cases where source data (SD) and target data (TD) with the same classes are obtained by the same sensor. However, the classification performance is significantly reduced when there are new classes in TD. In addition, domain alignment, as one of the main approaches in DA, is carried out based on local spatial information, rarely taking into account nonlocal spatial information (nonlocal relationships) with strong correspondence. A graph information aggregation cross-domain few-shot learning (Gia-CFSL) framework is proposed, intending to make up for the above-mentioned shortcomings by combining FSL with domain alignment based on graph information aggregation. SD with all label samples and TD with a few label samples are implemented for FSL episodic training. Meanwhile, intradomain distribution extraction block (IDE-block) and cross-domain similarity aware block (CSA-block) are designed. The IDE-block is used to characterize and aggregate the intradomain nonlocal relationships and the interdomain feature and distribution similarities are captured in the CSA-block. Furthermore, feature-level and distribution-level cross-domain graph alignments are used to mitigate the impact of domain shift on FSL. Experimental results on three public HSI datasets demonstrate the superiority of the proposed method. The codes will be available from the website: https://github.com/YuxiangZhang-BIT/IEEE_TNNLS_Gia-CFSL.
Yuxiang Zhang 0005, Wei Li 0032, Mengmeng Zhang 0005, Ran Tao 0003, Qian Du 0001
IEEE Trans. Neural Networks Learn. Syst.6
2024 Block-Wise Partner Learning for Model Compression
abstract
Despite the great potential of convolutional neural networks (CNNs) in various tasks, the resource-hungry nature greatly hinders their wide deployment in cost-sensitive and low-powered scenarios, especially applications in remote sensing. Existing model pruning approaches, implemented by a "subtraction" operation, impose a performance ceiling on the slimmed model. Self-knowledge distillation (Self-KD) resorts to auxiliary networks that are only active in the training phase for performance improvement. However, the knowledge is holistic and crude, and the learning-based knowledge transfer is mediate and lossy. Here, we propose a novel model-compression method, termed block-wise partner learning (BPL), which comprises "extension" and "fusion" operations and liberates the compressed model from the bondage of baseline. Different from the Self-KD, the proposed BPL creates a partner for each block for performance enhancement in training. For the model to absorb more diverse information, a diversity loss (DL) is designed to evaluate the difference between the original block and the partner. Besides, the partner is fused equivalently instead of being discarded directly. After training, we can simply adopt the fused compressed model that contains the enhancement information of partners but with fewer parameters and less inference cost. As validated using the UC Merced land-use, NWPU-RESISC45, and RSD46-WHU datasets, the BPL demonstrates superiority over other compared model-compression approaches. For example, it attains a substantial floating-point operations (FLOPs) reduction of 73.97% with only 0.24 accuracy (ACC.) loss for ResNet-50 on the UC Merced land-use dataset. The code is available at https://github.com/zhangxin-xd/BPL.
Xin Zhang 0092, Weiying Xie, Yunsong Li 0001, Jie Lei 0001, Kai Jiang 0001, Leyuan Fang, Qian Du 0001
IEEE Trans. Neural Networks Learn. Syst.7
2024 Cross-Scene Joint Classification of Multisource Data With Multilevel Domain Adaption Network
abstract
Domain adaption (DA) is a challenging task that integrates knowledge from source domain (SD) to perform data analysis for target domain. Most of the existing DA approaches only focus on single-source-single-target setting. In contrast, multisource (MS) data collaborative utilization has been extensively used in various applications, while how to integrate DA with MS collaboration still faces great challenges. In this article, we propose a multilevel DA network (MDA-NET) for promoting information collaboration and cross-scene (CS) classification based on hyperspectral image (HSI) and light detection and ranging (LiDAR) data. In this framework, modality-related adapters are built, and then a mutual-aid classifier is used to aggregate all the discriminative information captured from different modalities for boosting CS classification performance. Experimental results on two cross-domain datasets show that the proposed method consistently provides better performance than other state-of-the-art DA approaches.
Mengmeng Zhang 0005, Xudong Zhao 0003, Wei Li 0032, Yuxiang Zhang 0005, Ran Tao 0003, Qian Du 0001
IEEE Trans. Neural Networks Learn. Syst.6
2023 A survey on hyperspectral image restoration: from the view of low-rank tensor approximation
Na Liu 0014, Wei Li 0032, Yinjian Wang, Ran Tao 0003, Qian Du 0001, Jocelyn Chanussot
Sci. China Inf. Sci.5
2023 Semantic Segmentation Network for Classification of Hyperspectral Images With Small Size Samples
abstract
A sparse label oriented semantic segmentation network (SL-SSNet) is proposed for classification of hyperspectral images in this paper. Since semantic segmentation network performs pixel-level classification and can extract long-range contextual information, we apply it to the task of hyperspectral image classification. To mitigate the small size sample problem, we not only design a lightweight fully convolutional network, but also explore the usefulness of unlabeled data by introducing two constraints. Firstly, an adversarial learning based multi-classifier consistency strategy is employed to improve the classification of unlabeled data. It constrains two different classifiers to have consistent prediction results on unlabeled data. As a result, the extracted features of unlabeled data can be more discriminative and the predictions are more reliable. Secondly, a manifold regularizer is applied to constrain the classification results of unlabeled data to be smooth with respect to the data manifold, which can further exploit the unlabeled data and alleviate the small size sample problem. The experimental results using multiple hyperspectral data demonstrate the efficiency of the proposed method.
Li Ma 0005, Shuyue Li, Zhiyong Zhou 0002, Yafeng Yao, Qian Du 0001
IEEE Geosci. Remote. Sens. Lett.5
2023 Convolution and Attention Mixer for Synthetic Aperture Radar Image Change Detection
abstract
Synthetic aperture radar (SAR) image change detection is a critical task and has received increasing attentions in the remote sensing community. However, existing SAR change detection methods are mainly based on convolutional neural networks (CNNs), with limited consideration of global attention mechanism. In this letter, we explore Transformer-like architecture for SAR change detection to incorporate global attention. To this end, we propose a convolution and attention mixer (CAMixer). First, to compensate the inductive bias for Transformer, we combine self-attention with shift convolution in a parallel way. The parallel design effectively captures the global semantic information via the self-attention and performs local feature extraction through shift convolution simultaneously. Second, we adopt a gating mechanism in the feed-forward network to enhance the non-linear feature transformation. The gating mechanism is formulated as the element-wise multiplication of two parallel linear layers. Important features can be highlighted, leading to high-quality representations against speckle noise. Extensive experiments conducted on three SAR datasets verify the superior performance of the proposed CAMixer. The source codes will be publicly available at https://github.com/summitgao/CAMixer.
Haopeng Zhang 0016, Zijing Lin, Feng Gao 0005, Junyu Dong, Qian Du 0001, Heng-Chao Li 0001
IEEE Geosci. Remote. Sens. Lett.5
2023 Weakly supervised adversarial learning via latent space for hyperspectral target detection
Weiying Xie, Yunsong Li 0001, Kai Jiang 0001, Jie Lei 0001, Qian Du 0001
Pattern Recognit.6
2023 An Interpretable Unsupervised Unrolling Network for Hyperspectral Pansharpening
abstract
Existing deep convolutional neural networks (CNNs) have recently achieved great success in pansharpening. However, most deep CNN-based pansharpening models are based on "black-box" architecture and require supervision, making these methods rely heavily on the ground-truth data and lose their interpretability for specific problems during network training. This study proposes a novel interpretable unsupervised end-to-end pansharpening network, called as IU2PNet, which explicitly encodes the well-studied pansharpening observation model into an unsupervised unrolling iterative adversarial network. Specifically, we first design a pansharpening model, whose iterative process can be computed by the half-quadratic splitting algorithm. Then, the iterative steps are unfolded into a deep interpretable iterative generative dual adversarial network (iGDANet). Generator in iGDANet is interwoven by multiple deep feature pyramid denoising modules and deep interpretable convolutional reconstruction modules. In each iteration, the generator establishes an adversarial game with the spatial and spectral discriminators to update both spectral and spatial information without ground-truth images. Extensive experiments show that, compared with the state-of-the-art methods, our proposed IU2PNet exhibits very competitive performance in terms of quantitative evaluation metrics and qualitative visual effects.
Jiahui Qu, Wenqian Dong, Yunsong Li 0001, Shaoxiong Hou, Qian Du 0001
IEEE Trans. Cybern.5
2023 Hyperspectral and LiDAR Data Classification Based on Structural Optimization Transmission
abstract
With the development of the sensor technology, complementary data of different sources can be easily obtained for various applications. Despite the availability of adequate multisource observation data, for example, hyperspectral image (HSI) and light detection and ranging (LiDAR) data, existing methods may lack effective processing on structural information transmission and physical properties alignment, weakening the complementary ability of multiple sources in the collaborative classification task. The complementary information collaboration manner and the redundancy exclusion operator need to be redesigned for strengthening the semantic relatedness of multisources. As a remedy, we propose a structural optimization transmission framework, namely, structural optimization transmission network (SOT-Net), for collaborative land-cover classification of HSI and LiDAR data. Specifically, the SOT-Net is developed with three key modules: 1) cross-attention module; 2) dual-modes propagation module; and 3) dynamic structure optimization module. Based on above designs, SOT-Net can take full advantage of the reflectance-specific information of HSI and the detailed edge (structure) representations of multisource data. The inferred transmission plan, which integrates a self-alignment regularizer into the classification task, enhances the robustness of the feature extraction and classification process. Experiments show consistent outperformance of SOT-Net over baselines across three benchmark remote sensing datasets, and the results also demonstrate that the proposed framework can yield satisfying classification result even with small-size training samples.
Mengmeng Zhang 0005, Wei Li 0032, Yuxiang Zhang 0005, Ran Tao 0003, Qian Du 0001
IEEE Trans. Cybern.5
2023 Filter Pruning via Learned Representation Median in the Frequency Domain
abstract
In this article, we propose a novel filter pruning method for deep learning networks by calculating the learned representation median (RM) in frequency domain (LRMF). In contrast to the existing filter pruning methods that remove relatively unimportant filters in the spatial domain, our newly proposed approach emphasizes the removal of absolutely unimportant filters in the frequency domain. Through extensive experiments, we observed that the criterion for "relative unimportance" cannot be generalized well and that the discrete cosine transform (DCT) domain can eliminate redundancy and emphasize low-frequency representation, which is consistent with the human visual system. Based on these important observations, our LRMF calculates the learned RM in the frequency domain and removes its corresponding filter, since it is absolutely unimportant at each layer. Thanks to this, the time-consuming fine-tuning process is not required in LRMF. The results show that LRMF outperforms state-of-the-art pruning methods. For example, with ResNet110 on CIFAR-10, it achieves a 52.3% FLOPs reduction with an improvement of 0.04% in Top-1 accuracy. With VGG16 on CIFAR-100, it reduces FLOPs by 35.9% while increasing accuracy by 0.5%. On ImageNet, ResNet18 and ResNet50 are accelerated by 53.3% and 52.7% with only 1.76% and 0.8% accuracy loss, respectively. The code is based on PyTorch and is available at https://github.com/zhangxin-xd/LRMF.
Xin Zhang 0092, Weiying Xie, Yunsong Li 0001, Jie Lei 0001, Qian Du 0001
IEEE Trans. Cybern.5
2023 t-Linear Tensor Subspace Learning for Robust Feature Extraction of Hyperspectral Images
abstract
Subspace learning has been widely applied for feature extraction of hyperspectral images (HSIs) and achieved great success. However, the current methods still leave two problems that need to be further investigated. First, those methods mainly focus on finding one or multiple projection matrices for mapping the high-dimensional data into a low-dimensional subspace, which can only capture the information from each direction of high-order hyperspectral data separately. Second, the performance of feature extraction is barely satisfactory when the hyperspectral data is severely corrupted by noise. To address these issues, this article presents a t-linear tensor subspace learning (tLTSL) model for robust feature extraction of HSIs based on t-product projection. In the model, t-product projection is a new defined tensor transformation way similar to linear transformation in vector space, which can maximally capture the intrinsic structure of tensor data. The integrated tensor low-rank and sparse decomposition can effectively remove the noise corruption and the learned t-product projection can directly transform the high-order hyperspectral data into a subspace with information from all modes comprehensively considered. Moreover, a proposition related to tensor rank is proofed for interpreting the meaning of the tLTSL model. Extensive experiments are conducted on two different kinds of noise (i.e., simulated and real noise) corrupted HSI data, which validate the effectiveness of tLTSL.
Yangjun Deng, Heng-Chao Li 0001, Siqiao Tan, Junhui Hou, Qian Du 0001, Antonio Plaza
IEEE Trans. Geosci. Remote. Sens.5
2023 CBFF-Net: A New Framework for Efficient and Accurate Hyperspectral Object Tracking
abstract
Visual object tracking is a fundamental task in computer vision, and thrived in recent decades. With the development of snapshot hyperspectral sensors, efforts have been made to exploit tracking the object with hyperspectral (HS) videos to overcome the inherent limitation of RGB images. Existing HS tracking algorithms extract the deep features from image data separately, which break the interaction information between bands. Therefore, the discrimination ability of HS trackers is limited and the efficiency of the existing HS algorithms is low. In this paper, a novel algorithm (CBFF-Net) is proposed for HS object tracking to improve the discrimination ability and reduce the computational complexity. Specifically, the backbone and head network are implemented with modules of a transferred RGB object tracking network to carry out the HS target tracking task while maintaining the discrimination ability learned from RGB data. Moreover, a bi-directional multiple deep feature fusion (BMDFF) module is proposed to fuse the features extracted from different bands of the HS images, and a cross-band group attention (CBGA) module is introduced to learn interaction information across bands of the HS images. Experiments results indicate the superiority in performance of CBFF-Net, and it runs at 24 frames per second.
Pan Liu 0009, Weiying Xie, Jie Lei 0001, Yunsong Li 0001, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.7
2023 Hyperspectral Anomaly Detection Based on Chessboard Topology
abstract
Without any prior information, hyperspectral anomaly detection is devoted to locating targets of interest within a specific scene by exploiting differences in spectral characteristics between various land covers. Traditional methods originated from the signal processing perspective, and most of them rely heavily on specific model assumptions. Because of the model-driven attributes, such methods cannot mine the deep-level features of data to adapt to the variability of scenes and cannot fully extract the information of land covers contained in images to accurately separate anomalies from the background. By independently designing a chessboard-shaped topological framework that avoids making any distribution assumptions but directly mines high-dimensional data features to break through the limitations of traditional detectors, this article proposes a novel chessboard topology-based anomaly detection (CTAD) method to dissect images and extract detailed information of land covers adaptively, thereby enabling highly accurate detection. Extensive experimental results on hyperspectral images (HSIs) in real scenes demonstrate that the proposed CTAD can be adapted to the variability of scenes by autonomously learning data features and exhibiting strong generalization and detection capabilities, facilitating practical applications.
Lianru Gao, Xu Sun 0005, Lina Zhuang, Qian Du 0001, Bing Zhang 0001
IEEE Trans. Geosci. Remote. Sens.5
2023 A Semantic Transferred Priori for Hyperspectral Target Detection With Spatial-Spectral Association
abstract
Hyperspectral target detection is a crucial application that encompasses military, environmental, and civil needs. Target detection algorithms that have prior knowledge often assume a fixed laboratory target spectrum, which can differ significantly from the test image in the scene. This discrepancy can be attributed to various factors such as atmospheric conditions and sensor internal effects, resulting in decreased detection accuracy. To address this challenge, this article introduces a novel method for detecting hyperspectral image (HSI) targets with certain spatial information, referred to as the semantic transferred priori for hyperspectral target detection with spatial–spectral association (SSAD). Considering that the spatial textures of the HSI remain relatively constant compared to the spectral features, we propose to extract a unique and precise target spectrum from each image data via target detection in its spatial domain. Specifically, employing transfer learning, we designed a semantic segmentation network adapted for HSIs to discriminate the spatial areas of targets and then aggregated a customized target spectrum with those spectral pixels localized. With the extracted target spectrum, spectral dimensional target detection is performed subsequently by the constrained energy minimization (CEM) detector. The final detection results are obtained by combining an attention generator module to aggregate target features and deep stacked feature fusion (DSFF) module to hierarchically reduce the false alarm rate. Experiments demonstrate that our proposed method achieves higher detection accuracy and superior visual performance compared to the other benchmark methods.
Jie Lei 0001, Simin Xu, Weiying Xie, Yunsong Li 0001, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.6
2023 Class-Specific Autoaugment Architecture Based on Schmidt Mathematical Theory for Imbalanced Hyperspectral Classification
abstract
Hyperspectral image classification (HSIC) often suffers from severe imbalanced category distribution in real applications, which causes bias toward the dominated categories. As an effective method, the deep generative model (DGM) can be used to augment the features of imbalanced data through a learnable method to achieve superior classification performance. However, the features extracted by DGM are preset as a standard Gaussian distribution which results in low interclass difference. Besides, the generated features are too consistent with the original ones, which cannot play a positive role in the discriminability of minority categories (MCs). To conquer these drawbacks, we propose a class-specific autoaugment architecture based on Schmidt mathematical theory (CACS) for the challenging of imbalanced data which consists of two stages: one is training a superior features extractor, and the other one is augmenting features. The class-specific features of the whole HSI are extracted in stage one that supports the following feature augmented module. Specifically, we weighted the classifier in the first phase according to cost-sensitive learning, to prevent the classifier from overfitting. To expand the dispersion between categories, we construct feature prototypes obeying different Gaussian distributions for each class, respectively, and generate class-specific features. Then, the features are augmented in the second phase based on Schmidt’s mathematical theory, which enhances the discriminability of minority class features, thus further improving the classification accuracy with interpretability. Extensive experimental results on three benchmarking datasets demonstrate that CACS is outstanding in comparison algorithms, especially in MCs.
Jiaojiao Li 0001, Yan Diao, Rui Song 0003, Bobo Xi, Yunsong Li 0001, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.6
2023 Few-Shot Hyperspectral Image Classification With Self-Supervised Learning
abstract
Recently, few-shot learning (FSL) has been introduced for hyperspectral image (HSI) classification with few labeled samples. However, existing FSL-based HSI classification methods mainly focus on the meta-knowledge transfer between HSIs. Compared with HSIs, natural images have sufficient annotated data. To utilize natural images (base class data) to achieve accurate classification of HSIs (novel class data), we propose a novel few-shot classification framework with SSL (FSCF-SSL) for HSIs in this article. The orientation of objects in natural images is relatively unitary, whereas the objects of image patches for each pixel in HSIs have diverse orientations in the spatial domain. To make better use of base classes, we design an SSL with geometric transformations (SSLGTs), which sets rotation labels as supervision to extract low-level features that can better represent diverse orientations, and then conduct SSLGT and FSL on base classes to learn transferable spatial meta-knowledge. Next, a spectral-spatial feature extraction network is carefully designed to better utilize the spatial and spectral information of HSIs, where the weights of the first seven layers of the spatial part are initialized by the weights of the corresponding layers trained on base classes. Finally, to fully explore the few annotated data from novel classes, we design an SSL with contrastive learning (SSLCL) that can mine the category-invariant features contained in the novel class data itself, and then perform SSLCL and FSL on novel classes to learn more discriminative individual knowledge. Experimental results on four HSI datasets show that FSCF-SSL offers a significant improvement over state-of-the-art methods. The code is available athttps://github.com/Li-ZK/FSCF-SSL-2023.
Zhaokui Li, Yushi Chen 0002, Cuiwei Liu, Qian Du 0001, Zhuoqun Fang, Yan Wang 0087
IEEE Trans. Geosci. Remote. Sens.5
2023 A Model-Driven Deep Mixture Network for Robust Hyperspectral Anomaly Detection
abstract
Hyperspectral anomaly detection (HAD) aims to identify samples with unknown atypical spectra from the background. Deep learning (DL)-based methods, particularly autoencoders (AEs), have proven effective in uncovering the underlying profiles for HAD. However, in real-world applications of hyperspectral images (HSIs), complex background land-covers and anomaly corruptions are common, leading to two issues: 1) A low-dimensional manifold characterized by DL-based HAD methods can only reveal a few underlying variation factors of the background distribution and cannot capture the complex structures behind land-covers of all categories. 2) DL-based HAD methods trained on anomaly-contaminated HSIs tend to overfit specific anomalies, resulting in poor background characterization. To tackle these issues, this study presents a novel and robust framework for HAD called Model-Driven Deep Mixture Network (MDMN) that combines the strengths of model-driven and data-driven approaches while emphasizing interpretability. By assuming that the background, consisting of various land-covers, arises from a mixture of low-dimensional manifolds, the MDMN incorporates a novel deep mixture module to comprehensively characterize the background. This module utilizes a low-dimensional manifold learned by an AE to represent a specific category of background land-covers. To mitigate the impact of anomaly corruptions, the MDMN incorporates a convex relaxation of a sparse constraint, which helps prevent overfitting anomalies. Extensive experimental results demonstrate that the proposed MDMN offers more satisfactory and robust detection performance.
Yunsong Li 0001, Kai Jiang 0001, Weiying Xie, Jie Lei 0001, Xin Zhang 0092, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.6
2023 Sal²RN: A Spatial-Spectral Salient Reinforcement Network for Hyperspectral and LiDAR Data Fusion Classification
abstract
Hyperspectral image (HSI) and light detection and ranging (LiDAR) data fusion have been widely employed in HSI classification to promote interpreting performance. In the existing deep learning methods based on spatial–spectral features, the features extracted from different layers are treated fairly in the learning process. In reality, features extracted from the continuous layers contribute differentially to the final classification, such as large tracts of woodland and agriculture typically count on shallow contour features, whereas deep semantic spectral features have meaningful constraints for small entities like vehicles. Furthermore, the majority of existing classification algorithms employ a patch input scheme, which has a high probability to introduce pixels of different categories at the boundary. To acquire more accurate classification results, we propose a spatial–spectral saliency reinforcement network (Sal2RN) in this article. In spatial dimension, a novel cross-layer interaction module (CIM) is presented to adaptively alter the significance of features between various layers and integrate these diversified features. Moreover, a customized center spectrum correction module (CSCM) integrates neighborhood information and adaptively modifies the center spectrum to reduce intraclass variance and further improve the classification accuracy of the network. Finally, a statistically based feature weighted combination module is constructed to effectively fuse spatial, spectral, and LiDAR features. Compared with traditional and advanced classification methods, the Sal2RN achieves the state-of-the-art classification performance on three open benchmark datasets.
Jiaojiao Li 0001, Rui Song 0003, Yunsong Li 0001, Kailiang Han, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.6
2023 MFormer: Taming Masked Transformer for Unsupervised Spectral Reconstruction
abstract
Spectral reconstruction (SR) aims to recover the hyperspectral images (HSIs) from the corresponding RGB images directly. Most SR studies based on supervised learning require massive data annotations to achieve superior reconstruction performance, which are limited by complicated imaging techniques and laborious annotation calibration in practice. Thus, unsupervised strategies attract attention of the community, however, existing unsupervised SR works still face a fatal bottleneck from low accuracy. Besides, traditional CNN-based models are good at capturing local features but experience difficulty in global features. To ameliorate these drawbacks, we propose an unsupervised SR architecture with strong constraints, especially constructing a novel Masked Transformer (MFormer) to excavate latent hyperspectral characteristics to restore realistic HSIs further. Concretely, a Dual Spectral-wise Multi-head Self-attention (DSSA) mechanism embedded in transformer is proposed to firmly associate multi-head and channel dimensions and then capture the spectral representation in the implicit solution spaces. Furthermore, a plug-and-play Mask-guided Band Augment (MBA) module is presented to extract and further enhance the band-wise correlation and continuity to boost the robustness of the model. Innovatively, a customized loss based on the intrinsic mapping from HSIs to RGB images and the inherent spectral structural similarity is designed to restrain spectral distortion. Extensive experimental results on three benchmarks verify that our MFormer achieves superior performance over other state-of-the-art supervised and unsupervised methods under a no-label training process equally.
Jiaojiao Li 0001, Yihong Leng, Rui Song 0003, Wei Li 0032, Yunsong Li 0001, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.6
2023 Integrated Spatio-Spectral-Temporal Fusion via Anisotropic Sparsity Constrained Low-Rank Tensor Approximation
abstract
Although spatio-spectral and spatio-temporal fusion has been well explored, few efforts are made on integrating spatio-spectral-temporal features. As an intrinsic prior, low tensor-rank has been successfully taken into effect by current fusion models, most of which, however, resort to establishing an overall low-rank norm without performing factorization techniques thus have trouble capturing the latent high-order structure of hyperspectral data cube. To address that, a novel Anisotropicly Sparse (AS) tensor norm is developed to make the rank minimization a learnable process under Tucker decomposition. The AS norm enables the model to minimize the multi-linear tensor ranks if imposed on the core tensor after factorization, hence it significantly improves the model’s fusion performance. In the temporal domain, a Hadamard-product based variability descriptor is incorporated into the fusion model to map the former information to current time. Additionally, piece-wise smooth prior of the Tucker factors is employed by extra regularizers as supplement to the loss spatial information. With the Proximal Differential Matrix being developed for optimization, the proposed method reaches state-of-the-art results on both spatio-spectral and spatio-spectral-temporal fusion at low computational cost.
Wei Li 0032, Yinjian Wang, Na Liu 0014, Chenchao Xiao, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.6
2023 Supervised Contrastive Learning-Based Unsupervised Domain Adaptation for Hyperspectral Image Classification
abstract
Deep domain adaptation has achieved promising results in cross-domain hyperspectral image (HSI) classification. However, existing methods often focus on aligning data distributions without sufficient consideration of separability of source and target domain data themselves. In addition, current adversarial domain adaptation methods aim to achieve similar distributions between domains by confusing the discriminator, rather than obtaining a more compact distribution. In particular, existing methods are not discriminative enough for the target domain due to the difficulty of obtaining high-confidence labeled samples of the target domain. To address the above challenges, we propose a supervised contrastive learning-based unsupervised domain adaptation for HSI classification. A supervised contrastive learning strategy is then performed in both the source and target domains, which allows samples from the same category to be pulled closer together and samples from different categories to be pushed further apart, thus enhancing the separability of the data within the domain. The domain adaptation task is treated as a one-class classification (OCC) task, and a novel domain similarity loss based on OCC is introduced to reduce the discrepancy between domains. Finally, a confidence learning-based sample selection strategy is designed to select high-confidence labeled samples from the target domain to fine-tune the domain adaptation model, which can enhance the discrimination of the model to the target domain. Experimental results on three cross-domain datasets demonstrate that our proposed method outperforms existing domain adaptation methods. Our source code is available at https://github.com/Li-ZK/SCLUDA-2023.
Zhaokui Li, Li Ma 0005, Zhuoqun Fang, Yan Wang 0087, Wenqiang He, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.7
2023 SS-MAE: Spatial-Spectral Masked Autoencoder for Multisource Remote Sensing Image Classification
abstract
Masked image modeling (MIM) is a highly popular and effective self-supervised learning method for image understanding. Existing MIM-based methods mostly focus on spatial feature modeling, neglecting spectral feature modeling. Meanwhile, existing MIM-based methods use Transformer for feature extraction, some local or high-frequency information may get lost. To this end, we propose a spatial-spectral masked auto-encoder (SS-MAE) for HSI and LiDAR/SAR data joint classification. Specifically, SS-MAE consists of a spatial-wise branch and a spectral-wise branch. The spatial-wise branch masks random patches and reconstructs missing pixels, while the spectral-wise branch masks random spectral channels and reconstructs missing channels. Our SS-MAE fully exploits the spatial and spectral representations of the input data. Furthermore, to complement local features in the training stage, we add two lightweight CNNs for feature extraction. Both global and local features are taken into account for feature modeling. To demonstrate the effectiveness of the proposed SS-MAE, we conduct extensive experiments on three publicly available datasets. Extensive experiments on three multi-source datasets verify the superiority of our SS-MAE compared with several state-of-the-art baselines. The source codes are available at https://github.com/summitgao/SS-MAE.
Junyan Lin, Feng Gao 0005, Xiaochen Shi, Junyu Dong, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.5
2023 Spectral Variability Bayesian Unmixing for Hyperspectral Sequence in Wavelet Domain
abstract
For unmixing of sequences of hyperspectral images (SHS), spectral variability is an important factor to be considered. However, most existing unmixing methods tend to model the endmember and its variability in spatial domain rather than transform domain. In fact, the intrinsic and invariant features of the spectral curve can be effectively represented by wavelet transform. Therefore, this paper proposes to perform SHS unmixing in the wavelet domain by combing the Bayesian method. Firstly, the assumption of abundance being invariability in both the spatial and wavelet domains is made, then the formulation of unmixing in the wavelet domain using Perturbed Linear Mixing Model (PLMM) is presented. Secondly, based on the Bayesian framework, the likelihood and prior are both given, in which the parameter priors are divided into two parts: low and high frequency wavelet coefficients. Moreover, by considering the sparsity of the high-frequency wavelet coefficients of endmembers, a non-informative prior with zero-mean is designed. Meanwhile, for the coefficients of endmember variability, Gaussian distributions are utilized to represent the steady fluctuation along the temporal dimension. Finally, using the maximum a posterior (MAP) rule, a hierarchical spectral variability unmixing model in wavelet domain is built and solved by the Markov chain Monte Carlo (MCMC) sampling algorithm. Numerical experiments show that the proposed method generates more accurate estimates for endmembers and their variation.
Hongyi Liu 0001, Youkang Lu, Zebin Wu 0001, Qian Du 0001, Jocelyn Chanussot, Zhihui Wei
IEEE Trans. Geosci. Remote. Sens.4
2023 Category-Specific Prototype Self-Refinement Contrastive Learning for Few-Shot Hyperspectral Image Classification
abstract
Deep learning has been extensively used for hyperspectral image (HSI) classification with significant success, but the classification of high-dimensional HSI datasets with a limited amount of labeled samples is still a great challenge. Few-shot learning (FSL) has shown excellent performance in solving small-sample classification problems. However, most of the existing FSL methods usually suffer from the prototype instability and domain shift. In order to address these problems, this paper proposes a category-specific prototype self-refinement contrastive learning (CPSRCL) method for cross-domain FSL of HSIs. Our method uses a supervised contrastive learning (SCL) strategy to promote intra-class compactness and inter-class dispersion of features in the metric space. To stabilize and refine the prototypes of the support set, a category-specific prototype self-refinement (CSPSR) module is designed to adaptively learn different updating rules for different category prototypes using rich labeled information in the query set. Furthermore, a local discriminative domain adaptation (LDDA) method is constructed to align the global distribution between source and target domains while preserving domain-specific discriminative information. Experimental results on four public HSI datasets demonstrate that CPSRCL outperforms existing FSL and deep learning methods for HSI classification.
Quanyong Liu, Jiangtao Peng, Na Chen 0008, Weiwei Sun 0005, Yujie Ning, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.6
2023 Refined Prototypical Contrastive Learning for Few-Shot Hyperspectral Image Classification
abstract
Recently, prototypical network based few-shot learning (FSL) has been introduced for small-sample hyperspectral image (HSI) classification and shown good performance. However, existing prototypical-based FSL methods have two problems: prototype instability and domain shift between training and testing datasets. To solve these problems, we propose a refined prototypical contrastive learning network for few-shot learning (RPCL-FSL) in this paper, which incorporates supervised contrastive learning and FSL into an end-to-end network to perform small-sample HSI classification. To stabilize and refine the prototypes, RPCL-FSL imposes triple constraints on prototypes of the support set, i.e., contrastive learning (CL), self-calibration (SC) and cross-calibration (CC) based constraints. The CL module imposes internal constraint on the prototypes aiming to directly improve the prototypes using support set samples in the CL framework, and the SC and CC modules impose external constraints on the prototypes by using the prediction loss of support set samples and the query set prototypes, respectively. To alleviate domain shift in the FSL, a fusion training strategy is designed to reduce the feature differences between training and testing datasets. Experimental results on three HSI datasets demonstrate that the proposed RPCL-FSL outperforms existing state-of-the-art deep learning and FSL methods.
Quanyong Liu, Jiangtao Peng, Yujie Ning, Na Chen 0008, Weiwei Sun 0005, Qian Du 0001, Yicong Zhou
IEEE Trans. Geosci. Remote. Sens.6
2023 A Probabilistic Sample Boosting Approach With Adaptive Representation Coefficient Consistency for China Coastal Wetland Land Cover Classification Using GF-5 Hyperspectral Imagery
abstract
Wetland contains numerous features, and label acquisition is time-consuming, laborious, and inaccurate. Coastal wetland land cover classification with limited labeled training samples has become a significant challenge. In this study, a novel probabilistic ensemble sample selection framework (ProESS) is proposed for coastal wetland land cover classification. First, a sample probabilistic confidence index (SPCI) is proposed, which is defined by probabilistic output of each base classifier. Then the prediction confidences of unknown samples are obtained by joint probability of ensemble base classifiers, which can select high-quality samples to improve classification performance. However, the low accuracy of base classifiers will affect the confidence of samples selected by SPCI, thus reducing the classification accuracy. Based on this observation, an adaptive representation coefficient consistency learning (AdaRCCL) is proposed to help define SPCI. Finally, a ProESS is constructed through SPCI and AdaRCCL which can obtain training samples with high confidence from unknown samples. To evaluate the effectiveness of proposed methods, the three wetland hyperspectral datasets of China, i.e., Yangtze River Delta, Jiangsu Dafeng Natural Reserve, and Yellow River Delta, are used for classification experiments in the paper. Experimental results show that the proposed algorithms achieve higher performance and are robust to parameters in comparison to the baseline and the state-of-the-art ensemble algorithms. The extensibility and transferability of proposed methods are also discussed in the paper. Better results on multiple machine learning models with new samples show the extensibility of SPCI and ProESS. The great performance on Botswana dataset also demonstrates the transferability of proposed methods.
Hongjun Su, Hengyi Zheng, Weiwei Sun 0005, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.7
2023 Spectral Correlation-Based Diverse Band Selection for Hyperspectral Image Classification
abstract
Band selection which can reduce the spectral dimensionality effectively, has become one of the most popular topics in hyperspectral image (HSI) analysis. Recently, sparse representation based band selection (BS) has emerged as a popular tool. The existing sparse models mainly focus on minimizing reconstruction error and sparsity, while do not fully exploit the unique correlations among hundreds of continuous bands, which may cause representative bands missed and highly-correlated bands selected. Therefore, this paper proposes the spectral correlation based diverse band selection (SCDBS) for HSIs to improve representativeness and diversity of the selected bands. Specifically, a correlation derived weight is used to perform weighted sparse reconstruction to select the bands that are more correlated to the whole HSI, and a correlation minimization term is designed to remove the highly-correlated bands simultaneously. In addition, the proposed method imposes an adjustable sparse constraint by using an ℓ2,0
Mingyang Ma 0004, Shaohui Mei, Fan Li 0003, Yaoyang Ge, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.5
2023 Physical Knowledge-Enhanced Deep Neural Network for Sea Surface Temperature Prediction
abstract
Traditionally, numerical models have been deployed in oceanography studies to simulate ocean dynamics by representing physical equations. However, many factors pertaining to ocean dynamics seem to be ill-defined. We argue that transferring physical knowledge from observed data could further improve the accuracy of numerical models when predicting Sea Surface Temperature (SST). Recently, the advances in earth observation technologies have yielded a monumental growth of data. Consequently, it is imperative to explore ways in which to improve and supplement numerical models utilizing the ever-increasing amounts of historical observational data. To this end, we introduce a method for SST prediction that transfers physical knowledge from historical observations to numerical models. Specifically, we use a combination of an encoder and a generative adversarial network (GAN) to capture physical knowledge from the observed data. The numerical model data is then fed into the pre-trained model to generate physics-enhanced data, which can then be used for SST prediction. Experimental results demonstrate that the proposed method considerably enhances SST prediction performance when compared to several state-of-the-art baselines.
Yuxin Meng 0003, Feng Gao 0005, Eric Rigall, Ran Dong, Junyu Dong, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.6
2023 Contrastive Learning Based on Category Matching for Domain Adaptation in Hyperspectral Image Classification
abstract
Cross-scene hyperspectral image classification (HSIC) is a challenging topic in remote sensing, especially when there are no labels in target domain. Domain adaptation (DA) techniques for cross-scene HSIC aim to label a target domain by associating it with a labeled source domain. Most existing DA methods learn domain-invariant features by reducing feature distance across domains. Recently, contrastive learning has shown excellent performance in computer vision tasks, but there is little or no research on the performance of cross-scene HSIC. Considering that its idea is similar to reducing feature distance, this paper attempts to explore whether contrastive learning can achieve cross-scene HSIC. In this work, an instance-to-instance contrastive learning framework based on category matching (CLCM) is designed. The main idea is to take the category information as the premise in the feature space, regard the source sample as an anchor, and find its positive and negative matching samples across domains. The instance-level discriminative feature embeddings are learned through positive matching pairs attracting each other and negative matching pairs repelling each other. Among them, the target label is a pseudo-label. To further improve the quality of contrastive learning, it is considered to focus on extracting the spectral-spatial features of HSI to more accurately represent semantic information. Simultaneously, high-confidence target samples are screened to update the network. Three DA tasks confirm the effectiveness and feature discriminativeness of CLCM, while also providing new ideas for cross-scene image classification.
Yujie Ning, Jiangtao Peng, Quanyong Liu, Yi Huang 0021, Weiwei Sun 0005, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.6
2023 Multiview Spatial-Spectral Two-Stream Network for Hyperspectral Image Unmixing
abstract
Linear spectral unmixing is an important technique in the analysis of mixed pixels in hyperspectral images. In recent years, deep learning-based methods have been garnering increasing attention in hyperspectral unmixing; especially, unsupervised autoencoder (AE) networks that have achieved excellent unmixing performance are a recent trend. While most approaches use spatial information, it is well known that hyperspectral data are characterized by a large number of narrow spectral bands. In order to take full advantage of the hyperspectral bands in unmixing and the spatial information, in this article, we explore multiview spectral and spatial information in an AE-based unmixing framework. We introduce multiview spectral information through spectral partitioning and propose a multiview spatial–spectral two-stream network, called MSSS-Net, which simultaneously learns a spatial stream network and a multiview spectral stream network in an end-to-end fashion for more efficient unmixing. The MSSS-Net is a two-stream deep unmixing network sharing a decoder, where its two AE networks employ recurrent neural networks (RNNs) to collaboratively utilize multiview spectral and spatial information. The spatial stream network branch extracts the spatial features of pixels and its neighbors, while the multiview spectral stream network branch exploits the multiview spectral bands of a pixel. Meanwhile, we design a cascaded bidirectional and unidirectional RNNs’ encoder structure for multiview spatial–spectral information to learn more discriminative deep patch-pixel features. Extensive ablation studies and experiments on both synthetic and real datasets demonstrate the superiority of the MSSS-Net over state-of-the-art unmixing methods.
Lin Qi 0004, Feng Gao 0005, Junyu Dong, Xinbo Gao 0001, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.6
2023 Spectral-Spatial Morphological Attention Transformer for Hyperspectral Image Classification
abstract
In recent years, convolutional neural networks (CNNs) have drawn significant attention for the classification of hyperspectral images (HSIs). Due to their self-attention mechanism, the vision transformer (ViT) provides promising classification performance compared to CNNs. Many researchers have incorporated ViT for HSI classification purposes. However, its performance can be further improved because the current version does not use spatial–spectral features. In this article, we present a new morphological transformer (morphFormer) that implements a learnable spectral and spatial morphological network, where spectral and spatial morphological convolution operations are used (in conjunction with the attention mechanism) to improve the interaction between the structural and shape information of the HSI token and theCLStoken. Experiments conducted on widely used HSIs demonstrate the superiority of the proposed morphFormer over the classical CNN models and state-of-the-art transformer models. The source will be made available publicly athttps://github.com/mhaut/morphFormer.
Swalpa Kumar Roy, Ankur Deria, Chiranjibi Shah, Juan Mario Haut, Qian Du 0001, Antonio Plaza
IEEE Trans. Geosci. Remote. Sens.5
2023 Probabilistic Collaborative Representation Based Ensemble Learning for Classification of Wetland Hyperspectral Imagery
abstract
Protection of wetlands is important for ecosystem in recent years, and the classification of wetland ground cover is the foundation of investigation and protection work. Probabilistic collaborative representation classifier (ProCRC) is one of the best performing classifiers which has been applied in hyperspectral image (HSI) classification. However, its performance is greatly limited for wetland data where spectrums are highly similar. Moreover, the complex distribution of ground objects in wetlands have not been wisely utilized in the classification. In this article the intrinsic mechanism of ProCRC is found and its kernel version is proposed to solve the problems of wetlands classification. Then, a new ensemble learning strategy that considers neighborhood information are proposed, which largely alleviates the problem of sample collection in wetlands. Under the guidance of this strategy, two specific ensemble learning algorithms, i.e., LNE and LNSAE, are proposed. The superiority of proposed methods is validated using three typical HSI data sets of China coastal wetland with few samples.
Hongjun Su, Fu Shao, Weiwei Sun 0005, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.6
2023 Nearest Neighbor-Based Contrastive Learning for Hyperspectral and LiDAR Data Classification
abstract
The joint hyperspectral image (HSI) and light detection and ranging (LiDAR) data classification aims to interpret ground objects at more detailed and precise level. Although deep learning methods have shown remarkable success in the multisource data classification task, self-supervised learning has rarely been explored. It is commonly nontrivial to build a robust self-supervised learning model for multisource data classification, due to the fact that the semantic similarities of neighborhood regions are not exploited in the existing contrastive learning framework. Furthermore, the heterogeneous gap induced by the inconsistent distribution of multisource data impedes the classification performance. To overcome these disadvantages, we propose a nearest neighbor-based contrastive learning network (NNCNet), which takes full advantage of large amounts of unlabeled data to learn discriminative feature representations. Specifically, we propose a nearest neighbor-based data augmentation scheme to use enhanced semantic relationships among nearby regions. The intermodal semantic alignments can be captured more accurately. In addition, we design a bilinear attention module to exploit the second-order and even high-order feature interactions between the HSI and LiDAR data. Extensive experiments on four public datasets demonstrate the superiority of our NNCNet over state-of-the-art methods. The source codes are available athttps://github.com/summitgao/NNCNet.
Feng Gao 0005, Junyu Dong, Heng-Chao Li 0001, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.5
2023 RepCPSI: Coordinate-Preserving Proximity Spectral Interaction Network With Reparameterization for Lightweight Spectral Super-Resolution
abstract
Existing remarkable models for spectral super-resolution (SSR) achieve higher precision at the expense of computations with larger parameters. These algorithms require the heavy memory footprint and sufficient computing power, limiting their practical deployments and applications on portable devices. In this paper, we propose an efficient re-parameterizing coordinate-preserving proximity spectral interaction (RepCPSI) network for lightweight SSR. Specifically, the basic architecture is constituted of several polymorphic residual context restructuring (PRCR) modules to fully explore spatial and spectral contextual information with a multi-branch topology during the training stage. Using a structural re-parameterization scheme, the training-completed network is converted equivalently to a high-efficiency inference-time model, when it runs in the testing phase. To significantly improve the accuracy of SSR with an extra negligible computational overhead, a lightweight coordinate-preserving proximity spectral-aware attention (CPSA) block is developed. Such CPSA block can adaptively emphasize informative signatures and suppress useless ones among intermediate spatial-spectral features, which effectively enables the model to quickly locate features that are beneficial to the network learning and representation. Furthermore, considering the continuity of spectral variation for capturing real-world HSIs, a spectral physical consistency loss (SPCL) is added to the end-to-end network to constrain the changing trend of the spectral curve to be consistent with the ground-truth objects. Finally, our RepCPSI can accomplish a favorable balance between the reconstructed quality and model complexity. Extensive experimental results on six benchmarks demonstrate that our method obtains excellent performance with fewer parameters in terms of quantitative and qualitative measurements over the current advanced SSR approaches.
Chaoxiong Wu, Jiaojiao Li 0001, Rui Song 0003, Yunsong Li 0001, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.5
2023 Two-Branch Deeper Graph Convolutional Network for Hyperspectral Image Classification
abstract
Graph convolutional network (GCN) has recently attracted great attention in hyperspectral image (HSI) classification due to its strong ability to aggregate information of neighborhood nodes. However, a GCN model usually suffers from the over-smoothing problem (i.e., all nodes’ representations converge to a stationary point) when the number of GCN layers is increased. In addition, GCNs always work on superpixel-level nodes to reduce computational cost, so pixel-level features cannot be well captured. To deal with these problems, a novel two-branch deeper GCN (TBDGCN) is proposed to combine the advantages of superpixel-based GCN and pixel-based CNN, which can simultaneously extract superpixel-level and pixel-level features of HSIs. In the GCN branch, a GCN module with the DropEdge technique and residual connection is designed to alleviate over-smoothing and over-fitting problem, which results in a deeper network structure with more than ten layers. In the CNN branch, to capture spatial positional information and channel information, a mixed attention mechanism is constructed to extract attention-based spectral-spatial features. The features of the GCN and CNN branches are then fused for classification. Experimental results on three benchmark HSI data sets show that the classification performance of our TBDGCN is better than existing GCN models especially in the case of small sample size.
Linzhou Yu, Jiangtao Peng, Na Chen 0008, Weiwei Sun 0005, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.5
2023 SuperYOLO: Super Resolution Assisted Object Detection in Multimodal Remote Sensing Imagery
abstract
Accurately and timely detecting multiscale small objects that contain tens of pixels from remote sensing images (RSI) remains challenging. Most of the existing solutions primarily design complex deep neural networks to learn strong feature representations for objects separated from the background, which often results in a heavy computation burden. In this article, we propose an accurate yet fast object detection method for RSI, named SuperYOLO, which fuses multimodal data and performs high-resolution (HR) object detection on multiscale objects by utilizing the assisted super resolution (SR) learning and considering both the detection accuracy and computation cost. First, we utilize a symmetric compact multimodal fusion (MF) to extract supplementary information from various data for improving small object detection in RSI. Furthermore, we design a simple and flexible SR branch to learn HR feature representations that can discriminate small objects from vast backgrounds with low-resolution (LR) input, thus further improving the detection accuracy. Moreover, to avoid introducing additional computation, the SR branch is discarded in the inference stage, and the computation of the network model is reduced due to the LR input. Experimental results show that, on the widely used VEDAI RS dataset, SuperYOLO achieves an accuracy of 75.09% (in terms of$\text {mA}{{\text {P}}_{{50}}}$), which is more than 10% higher than the SOTA large models, such as YOLOv5l, YOLOv5x, and RS designed YOLOrs. Meanwhile, the parameter size and GFLOPs of SuperYOLO are about$18\times $and$3.8\times $less than YOLOv5x. Our proposed model shows a favorable accuracy–speed tradeoff compared to the state-of-the-art models. The code will be open-sourced athttps://github.com/icey-zhang/SuperYOLO.
Jie Lei 0001, Weiying Xie, Zhenman Fang, Yunsong Li 0001, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.6
2023 Morphological Transformation and Spatial-Logical Aggregation for Tree Species Classification Using Hyperspectral Imagery
abstract
Hyperspectral image (HSI) consists of abundant spectral and spatial characteristics, which contribute to a more accurate identification of materials and land covers. However, most existing methods of hyperspectral image analysis primarily focus on spectral knowledge or coarse-grained spatial information while neglecting the fine-grained morphological structures. In the classification task of complex objects, spatial morphological differences can help to search for the boundary of fine-grained classes, e.g., forestry tree species. Focusing on subtle traits extraction, a spatial-logical aggregation network (SLA-NET) is proposed with morphological transformation for tree species classification. The morphological operators are effectively embedded with the trainable structuring elements, which contributes to distinctive morphological representations. We evaluate the classification performance of the proposed method on two tree species datasets, and the results demonstrate that the proposed SLA-NET significantly outperforms the other state-of-the-art classifiers.
Mengmeng Zhang 0005, Wei Li 0032, Xudong Zhao 0003, Huan Liu 0015, Ran Tao 0003, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.6
2023 Adaptive Hypergraph Regularized Multilayer Sparse Tensor Factorization for Hyperspectral Unmixing
abstract
Hyperspectral unmixing with tensor models has received great attention in recent years. A tensor-based decomposition method can effectively represent the structural feature of hyperspectral images; however, the obtained results may be physically uninterpretable. To overcome this limitation, a novel adaptive hypergraph regularized multilayer sparse tensor factorization (AHGMLSTF) algorithm is proposed. First, a modified hypergraph is incorporated into tensor factorization, and the modified hypergraph uses spectral angle distance (SAD) instead of Euclidean distance to construct hyperedges to better represent the joint spatial and spectral information. Then, the hypergraph is constructed adaptively by hyperedges of$k$neighborhoods. Second, the concept of multilayer decomposition is introduced to explore the hierarchical features of hyperspectral images, and a sparse constraint is imposed on each layer to make the unmixing results more consistent with the physical mechanism of mixed spectral pixels. With these constraints, the proposed method established a spectral–spatial joint tensor decomposition model that represents not only the local neighborhood similarity but also the heterogeneity of adjacent edges. Experiments on simulated data and real hyperspectral data demonstrate the effectiveness of the proposed method.
Hongjun Su, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.4
2023 Single-Source Domain Expansion Network for Cross-Scene Hyperspectral Image Classification
abstract
Currently, cross-scene hyperspectral image (HSI) classification has drawn increasing attention. It is necessary to train a model only on source domain (SD) and directly transferring the model to target domain (TD), when TD needs to be processed in real time and cannot be reused for training. Based on the idea of domain generalization, a Single-source Domain Expansion Network (SDEnet) is developed to ensure the reliability and effectiveness of domain extension. The method uses generative adversarial learning to train in SD and test in TD. A generator including semantic encoder and morph encoder is designed to generate the extended domain (ED) based on encoder-randomization-decoder architecture, where spatial randomization and spectral randomization are specifically used to generate variable spatial and spectral information, and the morphological knowledge is implicitly applied as domain invariant information during domain expansion. Furthermore, the supervised contrastive learning is employed in the discriminator to learn class-wise domain invariant representation, which drives intra-class samples of SD and ED. Meanwhile, adversarial training is designed to optimize the generator to drive intra-class samples of SD and ED to be separated. Extensive experiments on two public HSI datasets and one additional multispectral image (MSI) dataset demonstrate the superiority of the proposed method when compared with state-of-the-art techniques. The codes will be available from the website:https://github.com/YuxiangZhang-BIT/IEEE_TIP_SDEnet.
Yuxiang Zhang 0005, Wei Li 0032, Ran Tao 0003, Qian Du 0001
IEEE Trans. Image Process.5
2023 Asymmetric Feature Fusion Network for Hyperspectral and SAR Image Classification
abstract
Joint classification using multisource remote sensing data for Earth observation is promising but challenging. Due to the gap of imaging mechanism and imbalanced information between multisource data, integrating the complementary merits for interpretation is still full of difficulties. In this article, a classification method based on asymmetric feature fusion, named asymmetric feature fusion network (AsyFFNet), is proposed. First, the weight-share residual blocks are utilized for feature extraction while keeping separate batch normalization (BN) layers. In the training phase, redundancy of the current channel is self-determined by the scaling factors in BN, which is replaced by another channel when the scaling factor is less than a threshold. To eliminate unnecessary channels and improve the generalization, a sparse constraint is imposed on partial scaling factors. Besides, a feature calibration module is designed to exploit the spatial dependence of multisource features, so that the discrimination capability is enhanced. Experimental results on the three datasets demonstrate that the proposed AsyFFNet significantly outperforms other competitive approaches.
Wei Li 0032, Yunhao Gao, Mengmeng Zhang 0005, Ran Tao 0003, Qian Du 0001
IEEE Trans. Neural Networks Learn. Syst.5
2023 Hyperspectral and SAR Image Classification via Multiscale Interactive Fusion Network
abstract
Due to the limitations of single-source data, joint classification using multisource remote sensing data has received increasing attention. However, existing methods still have certain shortcomings when faced with feature extraction from single-source data and feature fusion between multisource data. In this article, a method based on multiscale interactive information extraction (MIFNet) for hyperspectral and synthetic aperture radar (SAR) image classification is proposed. First, a multiscale interactive information extraction (MIIE) block is designed to extract meaningful multiscale information. Compared with traditional multiscale models, it can not only obtain richer scale information but also reduce the model parameters and lower the network complexity. Furthermore, a global dependence fusion module (GDFM) is developed to fuse features from multisource data, which implements cross attention between multisource data from a global perspective and captures long-range dependence. Extensive experiments on the three datasets demonstrate the superiority of the proposed method and the necessity of each module for accuracy improvement.
Wei Li 0032, Yunhao Gao, Mengmeng Zhang 0005, Ran Tao 0003, Qian Du 0001
IEEE Trans. Neural Networks Learn. Syst.6
2023 Semisupervised Cross-Scale Graph Prototypical Network for Hyperspectral Image Classification
abstract
In practice, the acquirement of labeled samples for hyperspectral image (HSI) is time-consuming and labor-intensive. It frequently induces the trouble of model overfitting and performance degradation for the supervised methodologies in HSI classification (HSIC). Fortunately, semisupervised learning can alleviate this deficiency, and graph convolutional network (GCN) is one of the most effective semisupervised approaches, which propagates the node information from each other in a transductive manner. In this study, we propose a cross-scale graph prototypical network (X-GPN) to achieve semisupervised high-quality HSIC. Specifically, considering the multiscale appearance of the land covers in the same remotely captured scene, we involve the neighborhoods of different scales to construct the adjacency matrices and simultaneously design a multibranch framework to investigate the abundant spectral-spatial features through graph convolutions. Furthermore, to exploit the complementary information between different scales, we simply employ the standard 1-D convolution to excavate the dependence of the intranode and concatenate the output with the features generated from other scales. Intuitively, different branches for various samples should have different importance to predict their categories. Thus, we develop a self-branch attentional addition (SBAA) module to adaptively highlight the most critical features produced by multiple branches. In addition, different from previous GCN for HSIC, we devise an innovative prototypical layer comprising a distance-based cross-entropy (DCE) loss function and a novel temporal entropy-based regularizer (TER), which can enhance the discrimination and representativeness of the node features and prototypes actively. Extensive experiments demonstrate that the proposed X-GPN is superior to the classic and state-of-the-art (SOTA) methods in terms of the classification performance.
Bobo Xi, Jiaojiao Li 0001, Yunsong Li 0001, Rui Song 0003, Yuchao Xiao, Qian Du 0001, Jocelyn Chanussot
IEEE Trans. Neural Networks Learn. Syst.6
2022 Semantic Segmentation of High-Resolution Remote Sensing Images Using an Improved Transformer
abstract
Semantic segmentation has been widely researched for high level analysis of High Spatial Resolution (HSR) remote sensing images, where Convolutional Neural Network (CNN) is the mainstream method. However, the transformer with attention mechanism has its unique capacity of extracting global information which is generally ignored by CNN models. In this paper, a Swin Transformer with UPer head (STUP) is proposed to tackle with semantic segmentation problem on a challenging remote sensing land-cover dataset called LoveDA, which owns complex background samples and inconsistent classes distributions. The proposed STUP combines the Swin Transformer with Uper Head in the form of an encoder-decoder structure, to extract features of HSR images for segmentation. Furthermore, Focal Loss is adopted to handle the unbalanced distribution problem in the training step. Experimental results demonstrate that the proposed STUP clearly outperforms several state-of-the-art models.
Shaohui Mei, Ye Wang 0020, Mingyi He, Qian Du 0001
IGARSS6
2022 HTD-VIT: Spectral-Spatial Joint Hyperspectral Target Detection with Vision Transformer
abstract
In hyperspectral images (HSIs), spatial context provides complementary information to abundant spectral features. In this paper, a united spectral-spatial framework named HTD-ViT based on vision transformer (ViT) is proposed for HTD tasks. The HTD-ViT leverages the ViT to learn discriminative spectral-spatial features of each pixel and its neighboring pixels. Meanwhile, the spectral-spatial sequence construction operation uses spectrums in the cross region centered on the selected pixel to produce the corresponding spectral-spatial sequence for ViT processing. Furthermore, the spectral-spatial sample selection procedure based on coarse detection addresses the issue of lacking well-labeled training instances in the HTD tasks. Finally, the spectral-spatial pixel-level detection combines the discriminative feature from the spectral and the spatial domains to suppress the background. In contrast to traditional spatial-spectral feature extraction methods that stack the original spectral feature with spatial neighborhood information directly, joint spectral-spatial inference in HTD-ViT can effectively discover the underlying contextual and structure information in HSIs. Experiments on real HSIs verify the effectiveness of HTD-ViT, which takes full advantage of both the variable spectral and spatial features.
Weiying Xie, Yunsong Li 0001, Qian Du 0001
IGARSS4
2022 Collaborative-Competitive Representation with Spatial Regularization for Hyperspectral Anomaly Detection
abstract
Recently, Collaborative representation (CR) has drawn much attention towards anomaly detection for hyperspectral imagery. Pixels in the background can be represented with spatial neighbors. In CR, an l2norm is implemented for estimating the weight vector in a closed-form solution. In this paper, locality constraints are imposed on the representation framework called collaborative-competitive representation for hyperspectral anomaly detection (CCRD). By incorporating competition among neighboring atoms lying in various subspaces, local information can be incorporated into the global framework of CR. In addition, distance weighted Tikhonov regularization is used to enhance the performance of CCRD, named CCRDT. Moreover, spatial information is used in the objective function of CCRDT, resulting in SCCRDT, to further enhance the performance of anomaly detection. Experimental results on several hyperspectral datasets demonstrate the superiority of proposed methods in comparison to traditional detectors, such as Reed-Xiaoli (RX) algorithm, kernel RX (KRX) algorithm, and existing collaborative representation-based anomaly detectors.
Chiranjibi Shah, Qian Du 0001
IGARSS2
2022 Laplacian Regularized Spatial-Aware Collaborative Competitive Representation for Hyperspectral Dimensionality Reduction
abstract
Recently, graph-based methods have drawn increased attention for representing a high-dimensional features into a low- dimensional data. To obtain an optimal transform for the purpose of classification, different collaborative representation-based methods are for dimensionality reduction (DR). In previous work, a spatial-aware collaborative competitive representation (SaCCPGT) based unsupervised method was investigated for DR of hyperspectral imagery (HSI). It incorporates spatial information into the representation framework. However, it can be further enhanced by considering the data manifold structure. In this paper, Laplacian regularized SaCCPGT (LapSaCCPGT) is presented for DR of HSI to better utilize data structure information into the representation framework. The experimental results observed on different hyperspectral datasets demonstrate the superiority of the proposed LapSaCCPGT than the state-of-the-art DR methods.
Chiranjibi Shah, Qian Du 0001
IGARSS2
2022 Hyperspectral Image Classification Using Hierarchical Spatial-Spectral Transformer
abstract
In recent years, convolutional neural networks (CNNs) have been successfully applied in hyperspectral image (HSI) classification tasks. However, the spatial-spectral features within an HSI have not been well explored using convolutions in CNNs. In the paper, a novel end-to-end hierarchical spatial-spectral transformer (HSST) is proposed for HSI classification, in which effective spatial-spectral features are emphasized using multi-head self-attention mechanism (MHSA). MHSA module captures better internal correlation of HSI data than the traditional convolution operation and can compute weighting scores for spatial and spectral context of pixels. Furthermore, a hierarchical architecture is designed to reduce a large number of parameters in the original transformer-style networks while still achieving satisfying classification results. Experimental results over two benchmark HSI datasets demonstrated the proposed HSST obviously outperforms several state-of-the-art deep learning-based HSI classification algorithms.
Shaohui Mei, Mingyang Ma 0004, Fulin Xu, Yifan Zhang 0006, Qian Du 0001
IGARSS6
2022 Gaussian Information Entropy based band Reduction for Unsupervised Hyperspectral Video Tracking
abstract
Hyperspectral videos, which provide extra spectral characteristics besides spatial and temporal information, can improve the performance of object tracking using spectral signatures. However, there is a lack of labeled hyperspectral videos to support deep learning based model design. On the contrary, object tracking in the color space has been well developed in the past decade with many benchmark tracking models, e.g., SiamBAN. Therefore, how to transfer models designed in the color space to the hyperspectral space is of great importance. In this paper, hyperspectral videos are reduced into 3 bands using a band reduction algorithm, by which the existing well-trained trackers can be directly used. Specifically, Gaussian Information Entropy (GIE) is used to transform a hyperspectral video into a 3-band pseudo-color video, by which hyperspectral object tracking is conducted in an unsupervised mode. Experimental results demonstrate that object trackers designed in the color space can be transferred to hyperspectral videos using band reduction algorithms and the GIE based reduction is more effective than several well-known band reduction algorithms when using SiamBAN.
Yuru Su, Shaohui Mei, Ge Zhang 0006, Ye Wang 0020, Mingyi He, Qian Du 0001
IGARSS6
2022 Learning hyperspectral images from RGB images via a coarse-to-fine CNN
Shaohui Mei, Yunhao Geng, Junhui Hou, Qian Du 0001
Sci. China Inf. Sci.4
2022 Oriented Object Detection by Searching Corner Points in Remote Sensing Imagery
abstract
Oriented object detection in remote sensing images has drawn great attention since it can provide more accurate bounding boxes. We propose a one-stage anchor-free network based on searching four corner points of an object, which can yield an arbitrary quadrilateral to fit objects with different shapes and orientations. We detect the corners by combining two strategies, where one regresses to the relative corner positions with respect to their corresponding center and the other directly detects the absolute corner positions from the corner heatmaps. By defining a candidate corner region based on the regressed results, we check whether corner points from the corner heatmaps are included in the region. If so, the closest one relative to the regressed corner is selected as the final position; otherwise, the regressed corner position is utilized. Experiments were conducted on two aerial remote sensing datasets, and the results demonstrated that the proposed method achieves superior performance to both the anchor-based and anchor-free methods.
Xueqing Chen, Li Ma 0005, Qian Du 0001
IEEE Geosci. Remote. Sens. Lett.3
2022 Hyperspectral Pansharpening via Local Intensity Component and Local Injection Gain Estimation
abstract
Hyperspectral (HS) pansharpening is an attractive topic in the field of remote sensing, which has attracted the attention of many researchers. Component substitution (CS)-based HS pansharpening algorithms are of great interest due to their simplicity and high spatial quality, and they mainly consist of two phases: detail extraction and detail injection. Detail extraction is performed by estimating the intensity component, whereas detail injection depends on the definition of injection gain. In the classic CS-based pansharpening methods, the intensity component is estimated through a global synthesis scheme, and injection gains can be obtained by a context-adaptive or a global approach. In this letter, we propose an improved CS-based HS pansharpening method in which the intensity component and the injection gain are estimated locally achieved by the binary partition tree (BPT) image segmentation algorithm. The proposed method is applied to two credible CS-based HS pansharpening algorithms, including the Gram–Schmidt adaptive (GSA) and the Brovey transform (Brovey). The experimental results show that the proposed method improves the performance of GSA and Brovey and creates promising results perceptually and quantitatively.
Wenqian Dong, Jiahui Qu, Song Xiao 0001, Qian Du 0001
IEEE Geosci. Remote. Sens. Lett.5
2022 Recurrent Feedback Convolutional Neural Network for Hyperspectral Image Classification
abstract
Deep neural networks have achieved promising performance for hyperspectral image (HSI) classification. However, due to the limitation of the available labeled samples, the traditional deeper and wider neural networks usually cause the overfitting problem and lose the detailed information. To solve this problem, a brain-like structure, namely spatial attention-driven recurrent feedback convolutional neural network (SARFNN), is proposed by utilizing the recurrent feedback and attention mechanism structures, from which two deep models are further developed for HSI classification. First, a 2-D SARFNN (SARF2DNN) model is developed to learn the spatial features from HSI data. After that, to better exploit the 3-D characteristic, the 3-D version is extended from SARF2DNN, thus constructing an SARF3DNN model to extract joint spatial-spectral features. Moreover, with the help of the idea of brain-likeness, the recurrent feedback module is designed to recover information loss caused by deeper structure and the dimension reduction operation. The experimental results conducted on two HSI data sets show that our SARFNN architecture can achieve more competitive performance than other state-of-the-art algorithms.
Heng-Chao Li 0001, Shuang-Shuang Li, Wen-Shuai Hu, Jun-Huan Feng, Weiwei Sun 0005, Qian Du 0001
IEEE Geosci. Remote. Sens. Lett.6
2022 Synthetic Aperture Radar Image Change Detection via Layer Attention-Based Noise-Tolerant Network
abstract
Recently, change detection methods for synthetic aperture radar (SAR) images based on convolutional neural networks (CNN) have gained increasing research attention. However, existing CNN-based methods neglect the interactions among multilayer convolutions, and errors involved in the preclassification restrict the network optimization. To this end, we proposed a layer attention-based noise-tolerant network, termed LANTNet. In particular, we design a layer attention module that adaptively weights the feature of different convolution layers. In addition, we design a noise-tolerant loss function that effectively suppresses the impact of noisy labels. Therefore, the model is insensitive to noisy labels in the preclassification results. The experimental results on three SAR datasets show that the proposed LANTNet performs better compared to several state-of-the-art methods. The source codes are available at https://github.com/summitgao/LANTNet.
Desen Meng, Feng Gao 0005, Junyu Dong, Qian Du 0001, Heng-Chao Li 0001
IEEE Geosci. Remote. Sens. Lett.4
2022 A General Loss-Based Nonnegative Matrix Factorization for Hyperspectral Unmixing
abstract
Nonnegative matrix factorization (NMF) is a widely used hyperspectral unmixing model which decomposes a known hyperspectral data matrix into two unknown matrices, i.e., endmember matrix and abundance matrix. Due to the use of least-squares loss, the NMF model is usually sensitive to noise or outliers. To improve its robustness, we introduce a general robust loss function to replace the traditional least-squares loss and propose a general loss-based NMF (GLNMF) model for hyperspectral unmixing in this letter. The general loss function is a superset of many common robust loss functions and is suitable for handling different types of noise. Experimental results on simulated and real hyperspectral data sets demonstrate that our GLNMF model is more accurate and robust than existing NMF methods.
Jiangtao Peng, Weiwei Sun 0005, Hong Chen 0004, Yicong Zhou, Qian Du 0001
IEEE Geosci. Remote. Sens. Lett.6
2022 Change Detection in Synthetic Aperture Radar Images Using a Dual-Domain Network
abstract
Change detection from synthetic aperture radar (SAR) imagery is a critical yet challenging task. Existing methods mainly focus on feature extraction in the spatial domain, and little attention has been paid to the frequency domain. Furthermore, in patch-wise feature analysis, some noisy features in the marginal region may be introduced. To tackle the above two challenges, we propose a dual-domain network (DDNet). Specifically, we take features from the discrete cosine transform (DCT) domain into consideration and the reshaped DCT coefficients are integrated into the proposed model as the frequency domain branch. Feature representations from both frequency and spatial domain are exploited to alleviate the speckle noise. In addition, we further propose a multi-region convolution (MRC) module, which emphasizes the central region of each patch. The contextual information and central region features are modeled adaptively. The experimental results on three SAR data sets demonstrate the effectiveness of the proposed model. Our codes are available athttps://github.com/summitgao/SAR_CD_DDNet.
Xiaofan Qu, Feng Gao 0005, Junyu Dong, Qian Du 0001, Heng-Chao Li 0001
IEEE Geosci. Remote. Sens. Lett.4
2022 Spatial-Aware Collaboration-Competition Preserving Graph Embedding for Hyperspectral Image Classification
abstract
Recently, graph-based discriminant analysis has drawn much attention in representing a high-dimensional hyperspectral data set using a low-dimensional subspace by defining high-dimensional data structure on a graph. Obtaining optimal representation coefficients for classification purposes are the key in such methods. A closed form solution can be found to solve the problem related to collaborative representation using labeled samples, which offers computational efficiency. There exists an unsupervised approach of collaboration preserving graph embedding (CPGE) for dimensionality reduction (DR), and its performance is further enhanced by imposing locality-preserving constraint in the method called collaboration–competition preserving graph embedding (CCPGE). In this letter, we introduce spatial-aware collaboration–competitive preserving graph embedding with Tikhonov (SaCCPGT) by imposing a spatial regularization term in the objective function of CCPGE with Tikhonov regularization. In this way, spectral and spatial information can be utilized in a closed form solution in the proposed method. Experimental results on different hyperspectral data sets demonstrate the superior performance of the proposed SaCCPGT in comparison to state-of-the-art graph-based discriminant analysis approaches for DR.
Chiranjibi Shah, Qian Du 0001
IEEE Geosci. Remote. Sens. Lett.2
2022 Unsupervised Robust Projection Learning by Low-Rank and Sparse Decomposition for Hyperspectral Feature Extraction
abstract
Owing to the strong correlation between the spectral bands of hyperspectral images (HSIs), many feature extraction (FE) methods have been proposed to reduce the redundancy of hyperspectral data. However, Euclidean distance-based FE methods are sensitive to noise. To address this issue, this letter proposed a new unsupervised FE method called robust projection learning (RPL) by integrating the low-rank and sparse decomposition with projection learning. Specifically, in order to enhance the discrimination of traditional robust principal component analysis (RPCA), discriminative RPCA (DRPCA) is first proposed by decomposing the raw data into a low-rank part, a discriminative sparse part, and a structured noise. Moreover, for the purpose of redundancy reduction, projection learning is integrated into DRPCA to obtain a projection matrix with robustness and discrimination. To verify the validity of RPL, two real hyperspectral data sets are used for basic comparison and robust analysis. The corresponding experimental results demonstrate that RPL outperforms the comparative FE methods.
Heng-Chao Li 0001, Lei Pan 0003, Yangjun Deng, Qian Du 0001
IEEE Geosci. Remote. Sens. Lett.7
2022 Multiscale Low-Rank Spatial Features for Hyperspectral Image Classification
abstract
This letter presents a multiscale low-rank decomposition (MSLRD) method to extract multiscale spatial structures from hyperspectral images. The MSLRD assumes that ground objects have divergent characteristics in changing spatial scales. It decomposes each band image into a series of block-wise matrices, where these low-rank blocks take detailed spatial structures at multiple scales. It formulates the low-rank matrix decomposition problem into minimizing the ranks of all block matrices and adopts the alternative direction of the multiplier method to optimize it. Experiments on Indian Pines and Pavia University data sets show that the MSLRD can greatly improve the classification performance of regular classification on spectral features (i.e., all bands) and perform better than five state-of-the-art spatial feature extraction methods.
Weiwei Sun 0005, Wenjing Shao, Jiangtao Peng, Gang Yang 0006, Xiangchao Meng, Qian Du 0001
IEEE Geosci. Remote. Sens. Lett.6
2022 Heterogeneous Few-Shot Learning for Hyperspectral Image Classification
abstract
Deep learning has achieved great success in hyperspectral image (HSI) classification. However, its success relies on the availability of sufficient training samples. Unfortunately, the collection of training samples is expensive, time-consuming, and even impossible in some cases. Natural image datasets that are different from HSI, such as Image Net and mini-ImageNet, have abundant texture and structure information. Effective knowledge transfer between two heterogeneous datasets can significantly improve the accuracy of HSI classification. In this letter, heterogeneous few-shot learning (HFSL) for HSI classification is proposed with only a few labeled samples per class. First, few-shot learning is performed on the mini-ImageNet datasets to learn the transferable knowledge. Then, to make full use of the spatial and spectral information, a spectral–spatial fusion network is devised. Spectral information is obtained by the residual network with pure 1-D operators. Spatial information is extracted by a convolution network with pure 2-D operators, and the weights of the spatial network are initialized by the weights of the model trained on the mini-ImageNet datasets. Finally, few-shot learning is fine-tuned on HSI to extract discriminative spectral–spatial features and individual knowledge, which can improve the classification performance of the new classification task. Experiments conducted on two public HSI datasets demonstrate that the HFSL outperforms the existing few-shot learning methods and supervised learning methods for HSI classification with only a few labeled samples. Our source code is available athttps://github.com/Li-ZK/HFSL.
Yan Wang 0087, Zhaokui Li, Qian Du 0001, Yushi Chen 0002, Fei Li 0018, Haibo Yang 0003
IEEE Geosci. Remote. Sens. Lett.5
2022 Extended Collaborative Representation-Based Hyperspectral Imagery Classification
abstract
Collaborative representation (CR) has been demonstrated to be very effective for hyperspectral image classification. However, insufficient diversity of training samples often results in limited classification accuracy under small-training-sample conditions, especially when diverse spectral variation is presented in testing samples. In order to alleviate such a problem, a spectral variation augmented-based linear mixed model (SV-LMM) is proposed, in which the spectral variation is extracted by conducting singular value decomposition (SVD) over training samples. Such spectral variation is further utilized to extend the CR for hyperspectral classification. Experiments over two benchmark datasets, i.e., the Pavia Center dataset and the University of Houston dataset, demonstrate that the proposed extended CR-based classifier (ECRC) clearly improves the performance of conventional CRC for hyperspectral classification and outperforms several state-of-the-art algorithms.
Bobo Xie, Shaohui Mei, Ge Zhang 0006, Yifan Zhang 0006, Yan Feng 0005, Qian Du 0001
IEEE Geosci. Remote. Sens. Lett.6
2022 Adaptive Cross-Attention-Driven Spatial-Spectral Graph Convolutional Network for Hyperspectral Image Classification
abstract
Recently, graph convolutional networks (GCNs) have been developed to explore the spatial relationship between pixels, achieving better classification performance of hyperspectral images (HSIs). However, these methods fail to sufficiently leverage the relationship between spectral bands in HSI data. As such, we propose an adaptive cross-attention-driven spatial–spectral graph convolutional network (ACSS-GCN), which is composed of a spatial GCN (Sa-GCN) subnetwork, a spectral GCN (Se-GCN) subnetwork, and a graph cross-attention fusion module (GCAFM). Specifically, Sa-GCN and Se-GCN are proposed to extract the spatial and spectral features by modeling the correlations between spatial pixels and between spectral bands, respectively. Then, by integrating attention mechanism into information aggregation of the graph, the GCAFM, including three parts, i.e., the spatial graph attention block, the spectral graph attention block, and the fusion block, is designed to fuse the spatial and spectral features, and suppress noise interference in Sa-GCN and Se-GCN. Moreover, the idea of the adaptive graph is introduced to explore an optimal graph through backpropagation during the training process. Experiments on two HSI datasets show that the proposed method achieves better performance than other classification methods.
Jin-Yu Yang, Heng-Chao Li 0001, Wen-Shuai Hu, Lei Pan 0003, Qian Du 0001
IEEE Geosci. Remote. Sens. Lett.5
2022 Remote Sensing Image Translation via Style-Based Recalibration Module and Improved Style Discriminator
abstract
Existing remote sensing change detection methods are heavily affected by seasonal variation. Since vegetation colors are different between winter and summer, such variations are inclined to be falsely detected as changes. In this letter, we proposed an image translation method to solve the problem. A style-based recalibration module is introduced to capture seasonal features effectively. Then, a new style discriminator is designed to improve the translation performance. The discriminator can not only produce a decision for the fake or real sample but also return a style vector according to the channel-wise correlations. Extensive experiments are conducted on the season-varying data set. The experimental results show that the proposed method can effectively perform image translation, thereby consistently improving the season-varying image change detection performance. Our codes and data are available athttps://github.com/summitgao/RSIT_SRM_ISD.
Tiange Zhang, Feng Gao 0005, Junyu Dong, Qian Du 0001
IEEE Geosci. Remote. Sens. Lett.4
2022 Spectral Variability Augmented Two-Stream Network for Hyperspectral Sparse Unmixing
abstract
Deep learning-based methods have drawn great attention in hyperspectral unmixing and obtained promising performance due to their powerful learning capability. However, few existing networks explicitly deal with the spectral variability inevitably present in hyperspectral images, limiting their fitting performance. In this letter, a spectral variability augmented two-stream network (SVATN) is designed to explicitly address the problem of spectral variability in a deep convolutional network for sparse unmixing. Specifically, the proposed SVATN maps a random input to coefficients of spectral variability in addition to abundances of endmembers, in which spectral variability is accommodated by the linear mixture model as an augmented item. Moreover, a spatial-spectral correlation-based variability extraction method (SSCVE) is proposed to construct a spectral variability library, which serves as priors in the loss function to optimize the proposed SVATN. Experiments over synthetic and real data sets demonstrate the superiority of the proposed SVATN over several state-of-the-art methods. The code of our proposed method is released at: https://github.com/MeiShaohui/SVATN.
Ge Zhang 0006, Shaohui Mei, Bobo Xie, Yan Feng 0005, Qian Du 0001
IEEE Geosci. Remote. Sens. Lett.5
2022 E2E-LIADE: End-to-End Local Invariant Autoencoding Density Estimation Model for Anomaly Target Detection in Hyperspectral Image
abstract
Hyperspectral anomaly target detection (also known as hyperspectral anomaly detection (HAD)] is a technique aiming to identify samples with atypical spectra. Although some density estimation-based methods have been developed, they may suffer from two issues: 1) separated two-stage optimization with inconsistent objective functions makes the representation learning model fail to dig out characterization customized for HAD and 2) incapability of learning a low-dimensional representation that preserves the inherent information from the original high-dimensional spectral space. To address these problems, we propose a novel end-to-end local invariant autoencoding density estimation (E2E-LIADE) model. To satisfy the assumption on the manifold, the E2E-LIADE introduces a local invariant autoencoder (LIA) to capture the intrinsic low-dimensional manifold embedded in the original space. Augmented low-dimensional representation (ALDR) can be generated by concatenating the local invariant constrained by a graph regularizer and the reconstruction error. In particular, an end-to-end (E2E) multidistance measure, including mean-squared error (MSE) and orthogonal projection divergence (OPD), is imposed on the LIA with respect to hyperspectral data. More important, E2E-LIADE simultaneously optimizes the ALDR of the LIA and a density estimation network in an E2E manner to avoid the model being trapped in a local optimum, resulting in an energy map in which each pixel represents a negative log likelihood for the spectrum. Finally, a postprocessing procedure is conducted on the energy map to suppress the background. The experimental results demonstrate that compared to the state of the art, the proposed E2E-LIADE offers more satisfactory performance.
Kai Jiang 0001, Weiying Xie, Jie Lei 0001, Zan Li 0001, Yunsong Li 0001, Tao Jiang 0031, Qian Du 0001
IEEE Trans. Cybern.7
2022 Context-Aware Guided Attention Based Cross-Feedback Dense Network for Hyperspectral Image Super-Resolution
abstract
Convolutional neural networks (CNNs) have shown impressive performance in computer vision due to their non-linearity. Particularly, DenseNet that facilitates feature re-use in a feedforward manner has achieved state-of-the-art reconstruction accuracy for super-resolution (SR). However, most DenseNet based SR models transfer the features generated from each layer to all the subsequent layers, inevitably introducing redundancy, especially for high-dimensional hyperspectral (HS) images. To tackle this problem, we propose a two-branch cross-feedback dense network with context-aware guided attention (CFDcagaNet) for HS super-resolution (HSSR), which allows the network to learn the attention maps of high-level features and refine the low-level features in a feedback manner across two branches. Context-aware guided attention uses high-level posterior information to provide more faithful spatial-spectral guidance for low-level features, which enables CFDcagaNet to learn more effective spatial-spectral features at low levels and yield more effective spatial-spectral transfer in the network. Extensive experiments on widely-used datasets demonstrate that the proposed method outperforms state-of-the-art methods in terms of both quantitative values and visual qualities.
Wenqian Dong, Jiahui Qu, Tongzhen Zhang, Yunsong Li 0001, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.5
2022 Confident Learning-Based Domain Adaptation for Hyperspectral Image Classification
abstract
Cross-domain hyperspectral image classification is one of the major challenges in remote sensing, especially for target domain data without labels. Recently, deep learning approaches have demonstrated effectiveness in domain adaptation. However, most of them leverage unlabeled target data only from a statistical perspective but neglect the analysis at the instance level. For better statistical alignment, existing approaches employ the entire unevaluated target data in an unsupervised manner, which may introduce noise and limit the discriminability of the neural networks. In this article, we propose confident learning-based domain adaptation (CLDA) to address the problem from a new perspective of data manipulation. To this end, a novel framework is presented to combine domain adaptation with confident learning (CL), where the former reduces the interdomain discrepancy and generates pseudo-labels for the target instances, from which the latter selects high-confidence target samples. Specifically, the confident learning part evaluates the confidence of each pseudo-labeled target sample based on the assigned labels and the predicted probabilities. Then, high-confidence target samples are selected as training data to increase the discriminative capacity of the neural networks. In addition, the domain adaptation part and the confident learning part are trained alternately to progressively increase the proportion of high-confidence labels in the target domain, thus further improving the accuracy of classification. Experimental results on four datasets demonstrate that the proposed CLDA method outperforms the state-of-the-art domain adaptation approaches. Our source code is available athttps://github.com/Li-ZK/CLDA-2022.
Zhuoqun Fang, Zhaokui Li, Wei Li 0032, Yushi Chen 0002, Li Ma 0005, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.7
2022 MS-HLMO: Multiscale Histogram of Local Main Orientation for Remote Sensing Image Registration
abstract
Multi-source image registration is challenging due to intensity, rotation, and scale differences among the images. Considering the characteristics and differences of multi-source remote sensing images, a feature-based registration algorithm named Multi-scale Histogram of Local Main Orientation (MS-HLMO) is proposed. Harris corner detection is first adopted to generate feature points. The HLMO feature of each Harris feature point is extracted on a Partial Main Orientation Map (PMOM) with a Generalized Gradient Location and Orientation Histogram-like (GGLOH) feature descriptor, which provides high intensity, rotation, and scale invariance. The feature points are matched through a multi-scale matching strategy. Comprehensive experiments on 17 multi-source remote sensing scenes demonstrate that the proposed MS-HLMO and its simplified version MS-HLMO+outperform other competitive registration algorithms in terms of effectiveness and generalization.
Chenzhong Gao, Wei Li 0032, Ran Tao 0003, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.4
2022 Hyperspectral and Multispectral Classification for Coastal Wetland Using Depthwise Feature Interaction Network
abstract
The monitoring of coastal wetlands is of great importance to the protection of marine and terrestrial ecosystems. However, due to the complex environment, severe vegetation mixture, and difficulty of access, it is impossible to accurately classify coastal wetlands and identify their species with traditional classifiers. Despite the integration of multisource remote sensing data for performance enhancement, there are still challenges with acquiring and exploiting the complementary merits from multisource data. In this article, the depthwise feature interaction network (DFINet) is proposed for wetland classification. A depthwise cross attention module is designed to extract self-correlation and cross correlation from multisource feature pairs. In this way, meaningful complementary information is emphasized for classification. DFINet is optimized by coordinating consistency loss, discrimination loss, and classification loss. Accordingly, DFINet reaches the standard solution-space under the regularity of loss functions, while the spatial consistency and feature discrimination are preserved. Comprehensive experimental results on two hyperspectral and multispectral wetland datasets demonstrate that the proposed DFINet outperforms other competitive methods in terms of overall accuracy.
Yunhao Gao, Wei Li 0032, Mengmeng Zhang 0005, Jianbu Wang, Weiwei Sun 0005, Ran Tao 0003, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.7
2022 Self-Balancing Dictionary Learning for Relaxed Collaborative Representation of Hyperspectral Image Classification
abstract
Supervised dictionary learning and representation learning framework has demonstrated its superiority for hyperspectral image classification. Relaxed collaborative representation (RCR) has also been acknowledged as an effective method in balancing the similarity and difference between features. In this paper, a new dictionary learning method is introduced to balance discrimination and reconstruction of training samples. In the dictionary learning stage, two new indicators are designed to measure the discriminability of items and can be improved by optimizing coding coefficients. The imposedl2-norm between independent item and the mean of class-specific samples constrains the similarity, and the calculated weights measure the difference. In label determination stage, considering that the residuals of RCR are adversely affected by the significant difference of features, a new classification approach is introduced. Class labels are assigned without calculating the reconstruction errors but calculating the levels of comprehensive contribution from all training samples instead. Since the effectiveness will be degraded when dealing with more complex circumstances, a region-based version is further introduced. It can further improve the discrimination of dictionary items due to the reduced categories in each sub-image and reduce the computation cost. The experimental results on several hyperspectral datasets demonstrate that our methods can effectively improve the classification performance.
Hongjun Su, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.4
2022 A Dual Global-Local Attention Network for Hyperspectral Band Selection
abstract
This article proposes a dual global–local attention network (DGLAnet), which is an end-to-end unsupervised band selection (UBS) method that fully utilizes spatial and spectral information in both global and local aspects. The DGLAnet assumes that BS can be realized using the hyperspectral image (HSI) reconstruction process. First, the DGLAnet implements a dual attention module to obtain spatial–spectral and global–local features to reweight the HSI data. It adopts bi-directional relations to grasp spatial and spectral features from a global perspective. Meanwhile, the DGLAnet extracts local features through max-pooling and mean-pooling and then merges them via the convolution operation. Global–local features are utilized to learn attention to recalibrate the original data, and the reconstruction module is adopted to restore the original image from the reweighted HSI data. Finally, a proper band subset is selected by the constructed band evaluation index. Experiments on three hyperspectral data show that the DGLAnet outperforms other state-of-the-art methods and uses all bands with a lower computational cost.
Weiwei Sun 0005, Gang Yang 0006, Xiangchao Meng, Kai Ren 0003, Jiangtao Peng, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.7
2022 Hyperspectral Change Detection Based on Multiple Morphological Profiles
abstract
With the increasing availability of multitemporal hyperspectral imagery, hyperspectral change detection under heterogeneous backgrounds is a challenging task. Due to the complexity of background features, traditional change detection algorithms in the spectral domain cannot effectively detect changed features. A novel method using multiple morphological profiles (MMPs) is proposed for hyperspectral change detection to make full use of spatial information. In the designed framework, first, the max-tree/min-tree strategy is applied to extract different attributes of multitemporal hyperspectral images (HSIs), i.e., area attribute and height attribute. Second, a spectral angle weighted-based local absolute distance (SALA) method is designed to reconstruct the discriminative spectral domain. Then, the absolute distance (AD) is adopted to extract changes in constructed feature domain. Finally, a change map is obtained by guided filtering. Experiments conducted on four real hyperspectral datasets demonstrate that the proposed detector achieves better detection performance.
Zengfu Hou, Wei Li 0032, Lu Li 0005, Ran Tao 0003, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.5
2022 Pseudo Complex-Valued Deformable ConvLSTM Neural Network With Mutual Attention Learning for Hyperspectral Image Classification
abstract
Convolutional long short-term memory (ConvLSTM) has received much attention for hyperspectral image (HSI) classification due to its ability of modeling long-range correlations, which, however, is vulnerable to too many parameters and insufficient training, limiting its classification accuracy, especially for small samples. Different from it, traditional hand-crafted methods extract the features with basic attributes of HSIs, which can provide the lack of details and interpretability of deep semantic features. However, existing methods fail to incorporate their complementarity for HSI classification. As such, a Pseudo complex-valued (CV) Deformable ConvLSTM Neural Network with mutual Attention learning (APDCLNN) is proposed, providing a new way to realize the collaborative learning of hand-crafted and deep features for HSI classification. First, a 2-D pseudo CV deformable ConvLSTM (PDConvLSTM2D) cell is designed using deformable convolution and complex operations, with which a spatial–spectral PDConvLSTM2D neural network (SSPDCL2DNN) is built to extract scale- and spectral-enhanced deep spatial–spectral features. Then, 3-D Gabor filter is used to extract hand-crafted features, and a mutual attention-based multimodality feature learning and fusion (MAMLF) module is designed to integrate them into deep features for training and optimization of SSPDCL2DNN. Finally, an attention loss subnetwork is designed to refine the classification results. As we know, this is the first attempt to apply the idea of mutual attention learning to fuse hand-crafted and deep features for HSI classification. Extensive experiments on three widely used HSI datasets show the advantages of our model over other deep methods in terms of both quantitative and visual quality.
Wen-Shuai Hu, Heng-Chao Li 0001, Rui Wang 0090, Feng Gao 0005, Qian Du 0001, Antonio Plaza
IEEE Trans. Geosci. Remote. Sens.5
2022 Two-Branch Attention Adversarial Domain Adaptation Network for Hyperspectral Image Classification
abstract
Recent studies have shown that deep domain adaptation (DA) techniques have good performance on cross-domain hyperspectral image (HSI) classification problems. However, most existing deep HSI DA approaches directly use deep networks to extract features from the data, which ignores the detailed information of HSI in spectral and spatial dimensions. To effectively exploit the spectral–spatial joint information for DA of HSIs, we propose a two-branch attention adversarial DA (TAADA) network in this article. In the TAADA network, a two-branch feature extraction (TBFE) subnetwork is first designed as a generator to extract the attention-based spectral–spatial features. Then, a discriminator based on two classifiers with the multilayer FC-BN-ReLU-Dropout structure is constructed. Based on adversarial learning between the generator and the discriminator, the ability of discriminative feature extraction and cross-domain classification is improved simultaneously. Finally, the TAADA network can adjust the distribution between the source and target domains and extract domain-invariant features. Experimental results on three cross-scene HSI classification tasks show that our proposed TAADA outperforms some existing DA methods.
Yi Huang 0021, Jiangtao Peng, Weiwei Sun 0005, Na Chen 0008, Qian Du 0001, Yujie Ning
IEEE Trans. Geosci. Remote. Sens.5
2022 Boundary Extraction Constrained Siamese Network for Remote Sensing Image Change Detection
abstract
Change detection (CD) is crucial to the understanding of relationships and interactions among multitemporal high-resolution remote sensing (RS) images. However, various inherent attributes of images have different impacts on CD judgment. How to effectively use helpful information to improve the performance of CD is still a challenge. In this article, we present a boundary extraction constrained Siamese network (BESNet) to dig out the efficacy of boundary information. BESNet is a joint learning network in which a novel multiscale boundary extraction (MSBE) module is embedded. In this way, traditional and deep learning techniques are leveraged to learn together to maximize their respective strengths through cooperation. In particular, a new boundary extraction constrained (BEC) loss function combined with a contractive loss function is used to optimize the BESNet. Considering the interaction between various extracted features, a channel-shuffle fusion strategy is developed to exploit their complementary advantages between features. Our experiments show that the proposed BESNet can significantly improve the CD performance and generate more complete and clearer object boundaries. Experiments conducted on two real datasets over different scenes demonstrate its state-of-the-art performance.
Jie Lei 0001, Yijie Gu, Weiying Xie, Yunsong Li 0001, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.5
2022 HASIC-Net: Hybrid Attentional Convolutional Neural Network With Structure Information Consistency for Spectral Super-Resolution of RGB Images
abstract
Spectral super-resolution (SSR), referring to the recovery of a reasonable hyperspectral image (HSI) from a single RGB image, has achieved satisfactory performance as part of the continued development of a convolutional neural network (CNN) in remote sensing image processing. However, the majority of existing algorithms focus on the pursuit of networks with deeper or broader architecture. Such algorithms have a poor channel or band feature extraction and fusing performance, and fail to fully leverage the input RGB images. To overcome these issues, we present a novel hybrid attentional CNN with structure information consistency (HASIC-net) that uses a two-pathway architecture. Specifically, both sides are stacked with several 2-D residual groups (2-DRGs) and residual groups (1-DRGs) equipped with channel or band attention (BA) modules, which mainly focuses on extracting channel statistics and bandwise features, respectively, by a parallel pooling architecture. We introduce several transversal connections from 2-DRG to 1-DRG to realize the interaction of information flow between both sides. In addition, we take the structure information of both RGB images and HSI into consideration and devise a structure information consistency (SIC) module to merge the structure tensor prior to the RGB images with the input of each 2-DRG. We then combine spectral gradient constraint loss with mean relative absolute error as a novel loss function to further restrain the spectral distortion and smooth the reconstructed spectral response curves. Experimental results on four benchmark datasets (i.e., NTIRE 2020, NTIRE 2018, CAVE, and Harvard) demonstrate that our proposed HASIC-net achieves state-of-the-art performance.
Jiaojiao Li 0001, Songcheng Du, Rui Song 0003, Chaoxiong Wu, Yunsong Li 0001, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.6
2022 Self-Supervised Robust Deep Matrix Factorization for Hyperspectral Unmixing
abstract
Hyperspectral unmixing is a critical step to process hyperspectral images (HSIs). Nonnegative matrix factorization (NMF) has drawn extensive attention in remotely sensed hyperspectral unmixing since it does not require prior knowledge about the pure spectral constituents (endmembers) in the scene. However, this approach is normally implemented as a single-layer procedure, which does not allow for a refinement of the obtained endmember abundances. In addition, HSIs suffer from the interference of sparse noise (besides Gaussian noise), which brings challenges when pursuing efficient hyperspectral unmixing. To address these issues, we propose a new self-supervised robust deep matrix factorization (SSRDMF) model for hyperspectral unmixing, which consists of two parts:encoderanddecoder. In theencoder, a multilayer nonlinear structure is designed to directly map the observed HSI data to the corresponding abundances. The abundances are then decoded by thedecoder, in which the connected weights are treated as the extracted endmembers. By modeling the sparse noise explicitly, the proposed method can reduce the effect caused by both Gaussian and sparse noise. Furthermore, a self-supervised constraint is included for exploring the spectral information, which is beneficial to further improve unmixing performance. To validate our method, we have conducted extensive experiments on both synthetic and real datasets. Our experiments reveal that our newly developed SSRDMF achieves superior unmixing performance compared to other state-of-the-art methods.
Heng-Chao Li 0001, Xin-Ru Feng, Donghai Zhai, Qian Du 0001, Antonio Plaza
IEEE Trans. Geosci. Remote. Sens.4
2022 Sparse Coding-Inspired GAN for Hyperspectral Anomaly Detection in Weakly Supervised Learning
abstract
Anomaly detection (AD) from hyperspectral images (HSIs) is of great importance in both space exploration and Earth observations. However, the challenges caused by insufficient datasets, no labels, and noise corruption substantially downgrade the accuracy of detection. To solve these problems, this article proposes a sparse coding (SC)-inspired generative adversarial network (GAN) for weakly supervised hyperspectral AD (HAD), named sparseHAD. It can learn a discriminative latent reconstruction with small errors for background pixels and large errors for anomalous ones. First, a background-category searching step is built to alleviate the difficulty of data annotation. Then, an SC-inspired regularized network is integrated into an end-to-end GAN to form a weakly supervised spectral mapping model consisting of two encoders, a decoder, and a discriminator. This model not only makes the network more robust and interpretable experimentally and theoretically but also develops a new SC-inspired path for HAD. Subsequently, the proposed sparseHAD detects anomalies in a latent space rather than the original space, which also contributes to its noise robustness. Quantitative assessments and experiments over real HSIs demonstrate the unique promise of the proposed sparseHAD. The code, data, and trained models are available athttps://github.com/JiangThea/HAD.
Yunsong Li 0001, Tao Jiang 0031, Weiying Xie, Jie Lei 0001, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.5
2022 Deep Cross-Domain Few-Shot Learning for Hyperspectral Image Classification
abstract
One of the challenges in hyperspectral image (HSI) classification is that there are limited labeled samples to train a classifier for very high-dimensional data. In practical applications, we often encounter an HSI domain (called target domain) with very few labeled data, while another HSI domain (called source domain) may have enough labeled data. Classes between the two domains may not be the same. This article attempts to use source class data to help classify the target classes, including the same and new unseen classes. To address this classification paradigm, a meta-learning paradigm for few-shot learning (FSL) is usually adopted. However, existing FSL methods do not account for domain shift between source and target domain. To solve the FSL problem under domain shift, a novel deep cross-domain few-shot learning (DCFSL) method is proposed. For the first time, DCFSL tackles FSL and domain adaptation issues in a unified framework. Specifically, a conditional adversarial domain adaptation strategy is utilized to overcome domain shift, which can achieve domain distribution alignment. In addition, FSL is executed in source and target classes at the same time, which can not only discover transferable knowledge in the source classes but also learn a discriminative embedding model to the target classes. Experiments conducted on four public HSI data sets demonstrate that DCFSL outperforms the existing FSL methods and deep learning methods for HSI classification. Our source code is available athttps://github.com/Li-ZK/DCFSL-2021.
Zhaokui Li, Yushi Chen 0002, Yimin Xu, Wei Li 0032, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.6
2022 A Triplet Semisupervised Deep Network for Fusion Classification of Hyperspectral and LiDAR Data
abstract
Data fusion of hyperspectral and light detection and ranging (LiDAR) is conducive to obtain more comprehensive surface information and thereby achieve better classification result in Earth Monitoring Systems. However, lack of labeled samples usually limits the performance of supervised classifiers, and the heterogeneity of multi-source data also brings great challenges to data fusion. Aiming to address these issues, we propose a triplet semi-supervised deep convolutional neural network (TSDN) for fusion classification of hyperspectral and LiDAR. Specifically, we utilize three basic pathways to extract deep learning features: 1D-CNN for spectral features in hyperspectral, 2D-CNN for spatial features in hyperspectral and Cascade Net for elevation features in LiDAR data. Furthermore, a novel label calibration module (LCM) is proposed to generate effective pseudo labels with high confidence based on the superpixel segmentation by comparing the multi-view classification results for assisting semi-supervised model training. In addition, we design a novel 3D-Cross Attention Block to enhance the complementary spatial features of multi-source data. Experiments on three public HSI-LiDAR benchmarks: Houston, Trento, and MUUFL Gulfport have demonstrated the effectiveness and superiority of our proposed method.
Jiaojiao Li 0001, Yinle Ma, Rui Song 0003, Bobo Xi, Danfeng Hong, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.6
2022 A Stepwise Domain Adaptive Segmentation Network With Covariate Shift Alleviation for Remote Sensing Imagery
abstract
Semantic segmentation for remote sensing images (RSI) is critical for the Earth monitoring system. However, the covariate shift between RSI datasets under different capture conditions cannot be alleviated by directly using the unsupervised domain adaptation (UDA) method, which negatively affects the segmentation accuracy in RSI. We propose a stepwise domain adaptive segmentation network with covariate shift alleviation (Cov-DA) for RSI parsing to solve this issue. Specifically, to alleviate domain shift generated by different sensors, both the source and target domains are projected into a colorspace with normalized distribution through an elaborate colorspace mapping unified module (CMUM). The color distributions of these two domains tend to be more uniform. Furthermore, in the target domain, the multistatistics joint evaluation module (MJEM) is proposed to capture different statistical characteristics of subscenarios for selecting plain scenarios regarded as high-confidence segmentation results to assist the further improvement of segmentation performance. In addition, a pyramid perceptual attention module (PPAM) containing omnidirectional features without computational burdens is added to our network for effectively enhancing the multiscale feature capture ability. In the cross-city DA experiments based on the International Society for Photogrammetry and Remote Sensing (ISPRS) and aerial benchmarks, the superiority of our algorithm is significantly demonstrated. Furthermore, we release a large-scale Martian terrain dataset noted as “Mars-Seg” containing 5 K images with pixel-level accurate annotations regarding issues, such as the lack of semantic segmentation datasets for unknown scenes.
Jiaojiao Li 0001, Shunyao Zi, Rui Song 0003, Yunsong Li 0001, Yinlin Hu, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.6
2022 Structure-Guided Feature Transform Hybrid Residual Network for Remote Sensing Object Detection
abstract
Object detection in remote sensing imagery (RSI) is a fundamental task for Earth monitoring. Objects captured from the bird’s eye view perspective in RSI can appear as multiscale in arbitrary orientations, most of which are small and dense. In specific, vehicles or ships only occupy a dozen pixels in the image, but are surrounded by roads and seas, which occupy thousands of pixels and comprise overwhelmingly dominant of all pixels. Although a large number of common object detection methods have been proposed, most of them cannot detect small and dense objects accurately because none of them has paid enough attention to the unique characteristic of RSI. In this work, we propose a novel structure-guided feature transform hybrid residual (SGFTHR) network, which can conquer the low performance of detection of objects at different scales, especially for small and dense objects, in an anchor-free manner. The structure-guided feature transform (SGFT) module is promoted to extract discriminative structural information and guide this information into high-level contextual feature maps, preventing the important low-level spatial and structural information from being lost when the network goes deeper. Furthermore, the hybrid residual (HR) module is embedded in the backbone to acquire multiscale features in a novel hybrid hierarchical residual-like manner. Extensive experiments are performed on the HRRSD and NWPU VHR-10 datasets to evaluate the performance of the SGFTHR network, which demonstrates that our SGFTHR network achieves state-of-the-art detection accuracy with high efficiency and robustness. Specifically, 4.12% improvements in mean average precision (mAP) on the HRRSD dataset compared with baseline powerfully demonstrate the effectiveness and superiority of the SGFTHR network.
Jiaojiao Li 0001, Huanqing Zhang, Rui Song 0003, Weiying Xie, Yunsong Li 0001, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.6
2022 Multigraph-Based Low-Rank Tensor Approximation for Hyperspectral Image Restoration
abstract
Low-rank-tensor-approximation (LRTA)-based hyperspectral image (HSI) restoration has drawn increasing attention. However, most of the methods construct a hidden low-rank tensor by utilizing the non-local self-similarity (NLSS) and global spectral correlation (GSC) inherited by HSIs. Although achieving state-of-the-art (SOTA) restoration performance, NLSS and GSC have limitations. NLSS is introduced from natural image denoising to remove spatially independent identically distributed (i.i.d.) Gaussian and impulse noise. While GSC, which is naturally possessed by HSIs, is adopted to maintain the spectral integrity and remove spectrally, i.i.d., degradations. Therefore, NLSS and GSC may not be successfully used for complex HSI restoration tasks, such as destriping, cloud removal and recovery of atmospheric absorption bands. To solve the issue, borrowing the idea from manifold learning, the geometry information characterized by proximity relationship, is integrated with the LRTA to solve the above issue, named as multi-graph-based LRTA (MGLRTA). Different with most of the existing methods, the proposed MGLRTA directly models an HSI as a low-rank tensor and efficiently explores the extra proximity information on the defined graphs that are not only inherited by the low-rank constraints but also naturally possessed in HSIs. A well-posed iterative algorithm is designed to solve the restoration problem. Experimental results on different datasets that cover several severe degradation scenarios demonstrate that the proposed MGLRTA outperforms the SOTA HSI restoration methods.
Na Liu 0014, Wei Li 0032, Ran Tao 0003, Qian Du 0001, Jocelyn Chanussot
IEEE Trans. Geosci. Remote. Sens.4
2022 Bayesian Unmixing of Hyperspectral Image Sequence With Composite Priors for Abundance and Endmember Variability
abstract
A hyperspectral image sequence can be obtained at different time in the same region from a hyperspectral sensor. The environmental change usually leads to variation in endmember reflectance, which has an important influence on unmixing process. In this article, a Bayesian unmixing model considering spectral variability for hyperspectral sequence is proposed, in which composite prior distributions of abundance and endmember variability are developed. The abundance priors consider the continuity of abundance in the temporal and spatial domains, simultaneously. Specifically, in the spatial domain, a data-adaptive variance of the abundance prior distribution is put forward based on local spatial difference. Moreover, the priors of endmember variability in temporal continuity and spectral smoothness are also exploited. Finally, a joint posterior distribution is obtained by the likelihood function and the parameter prior distributions, which can be calculated by the Markov chain Monte Carlo (MCMC) algorithm. Experiments on synthetic and real data sets demonstrate the effectiveness of the proposed approach in terms of abundance, endmember, and its variability estimation accuracy.
Hongyi Liu 0001, Youkang Lu, Zebin Wu 0001, Qian Du 0001, Jocelyn Chanussot, Zhihui Wei
IEEE Trans. Geosci. Remote. Sens.4
2022 ABNet: Adaptive Balanced Network for Multiscale Object Detection in Remote Sensing Imagery
abstract
Benefiting from the development of convolutional neural networks (CNNs), many excellent algorithms for object detection have been presented. Remote sensing object detection (RSOD) is a challenging task mainly due to: 1) complicated background of remote sensing images (RSIs) and 2) extremely imbalanced scale and sparsity distribution of remote sensing objects. Existing methods cannot effectively solve these problems with excellent detection accuracy and rapid speed. To address these issues, we propose an adaptive balanced network (ABNet) in this article. First, we design an enhanced effective channel attention (EECA) mechanism to improve the feature representation ability of the backbone, which can alleviate the obstacles of complex background on foreground objects. Then, to combine multiscale features adaptively in different channels and spatial positions, an adaptive feature pyramid network (AFPN) is designed to capture more discriminative features. Furthermore, considering that the original FPN ignores rich deep-level features, a context enhancement module (CEM) is proposed to exploit abundant semantic information for multiscale object detection. Experimental results on three public datasets demonstrate that our approach exhibits superior performance over baseline by only introducing less than 1.5M extra parameters.
Yanfeng Liu, Qiang Li 0042, Yuan Yuan 0001, Qian Du 0001, Qi Wang 0009
IEEE Trans. Geosci. Remote. Sens.4
2022 Multiscale Alternately Updated Clique Network for Hyperspectral Image Classification
abstract
Recently, deep learning has drawn significant attention in hyperspectral image (HSI) classification. With the growth of network depth and feature integration, deep learning demands abundant labeled samples to optimize many parameters. Unfortunately, most hyperspectral data are unlabeled and the available labeled samples are extremely limited. How to obtain richer features under limited training samples is a challenge for HSI classification. To tackle this issue, a new supervised multiscale alternately updated clique network (MSCN) is proposed for HSI classification to fully employ HSI features in different scales. Based on the Clique Block, we design the multiscale alternately updated clique block (MSCB) that applies convolution kernels of various sizes to adaptively exploit the multiscale HSI information and merge them within the block. Meanwhile, the recurrent feedback architecture is introduced to reuse high-level visual information and network parameters. The proposed MSCN includes two MSCBs to capture the multiscale spectral and spatial information in turn. The MSCN improves the information flow and the efficiency of parameter tuning through the feedback mechanism and the cross-utilization of multiscale feature. It not only obtains more abstract HSI information, but also reduces the network depth and the number of parameters, thereby improving the classification accuracy under limited samples. To certify the validity of the proposed MSCN, experiments are conducted on three real HSI datasets and compared with multiple state-of-the-art deep learning-based approaches. The experimental results demonstrate that the presented multiscale network achieves superior performance, especially in the case of a small number of training samples.
Qian Liu 0008, Zebin Wu 0001, Qian Du 0001, Yang Xu 0006, Zhihui Wei
IEEE Trans. Geosci. Remote. Sens.3
2022 Dual-Frequency Autoencoder for Anomaly Detection in Transformed Hyperspectral Imagery
abstract
Hyperspectral anomaly detection (HAD) is a challenging task since samples are unavailable for training. Although unsupervised learning methods have been developed, they often train the model using an original hyperspectral image (HSI) and require retraining on different HSIs, which may limit the feasibility of HAD methods in practical applications. To tackle this problem, we propose a dual-frequency autoencoder (DFAE) detection model in which the original HSI is transformed into high-frequency components (HFCs) and low-frequency components (LFCs) before detection. A novel spectral rectification is first proposed to alleviate the spectral variation problem and generate the LFCs of HSI. Meanwhile, the HFCs are extracted by the Laplacian operator. Subsequently, the proposed DFAE model is learned to detect anomalies from the LFCs and HFCs in parallel. Finally, the learned model is well-generalized for anomaly detection from other hyperspectral datasets. While breaking the dilemma of limited generalization in the sample-free HAD task, the proposed DFAE can enhance the background–anomaly separability, providing a better performance gain. Experiments on real datasets demonstrate that the DFAE method exhibits competitive performance compared with other advanced HAD methods.
Yidan Liu, Weiying Xie, Yunsong Li 0001, Zan Li 0001, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.5
2022 Novel Cross-Resolution Feature-Level Fusion for Joint Classification of Multispectral and Panchromatic Remote Sensing Images
abstract
With the increasing availability and resolution of satellite sensor data, multispectral (MS) and panchromatic (PAN) images are the most popular data that are used in remote sensing among applications. This article proposes a novel cross-resolution hidden layer feature fusion (CRHFF) approach for joint classification of multiresolution MS and PAN images. In particular, shallow spectral and spatial features at a global scale are first extracted from an MS image. Then, deep cross-resolution hidden layer features extracted from MS and PAN are fused from patches at a local scale according to an autoencoder (AE)-like deep network. Finally, the selected multiresolution hidden layer features are classified in a supervised manner. By taking advantage of integrated shallow-to-deep and global-to-local features from the high-resolution MS and PAN images, the cross-resolution latent information can be extracted and fused in order to better model imaged objects from the multimodal representation and finally increase the classification accuracy. Experimental results obtained on three real multiresolution datasets covering complex urban scenarios confirm the effectiveness of the proposed approach in terms of higher accuracy and robustness with respect to literature methods.
Sicong Liu 0001, Qian Du 0001, Lorenzo Bruzzone, Alim Samat, Xiaohua Tong
IEEE Trans. Geosci. Remote. Sens.3
2022 A Shallow-to-Deep Feature Fusion Network for VHR Remote Sensing Image Classification
abstract
With more detailed spatial information being represented in very-high-resolution (VHR) remote sensing images, stringent requirements are imposed on accurate image classification. Due to the diverse land-objects with intraclass variation and interclass similarity, efficient and fine classification of VHR images especially in complex scenes is challenging. Even for some popular deep learning (DL) frameworks, geometric details of land-object may be lost in deep feature levels, so it is difficult to maintain the highly-detailed spatial information (e.g., edges, small objects) only relying on the last high-level layer. Moreover, many of the newly developed DL methods require massive well-labeled samples, which inevitably deteriorates the model generalization ability under the few-shot learning. Therefore, in this paper, a lightweight shallow-to-deep feature fusion network (SDF2N) is proposed for VHR image classification, where the traditional machine learning (ML) and DL schemes are integrated to learn rich and representative information to improve the classification accuracy. In particular, the shallow spectral-spatial features are first extracted, and then a novel triple-stage fusion (TSF) module is designed to learn the saliency and discriminative information at different levels for classification. The TSF module includes three feature fusion stages, i.e., low-level spectral-spatial feature fusion, middle-level multi-scale feature fusion, and high-level multi-layer feature fusion. The proposed SDF2N takes advantages of the shallow-to-deep features, which can extract representative and complementary information of crossing layers. It is important to note that even with limited training samples, the SDF2N still can achieve satisfying classification performance. Experimental results obtained on three real VHR remote sensing data sets including two multispectral and one airborne hyperspectral images covering complex urban scenarios confirm the effectiveness of the proposed approach compared with the state-of-the-art methods.
Sicong Liu 0001, Qian Du 0001, Lorenzo Bruzzone, Alim Samat, Xiaohua Tong, Yanmin Jin, Chao Wang 0092
IEEE Trans. Geosci. Remote. Sens.3
2022 Lightweight Tensorized Neural Networks for Hyperspectral Image Classification
abstract
Deep learning methods have demonstrated excellent performance in hyperspectral image (HSI) classification. However, these methods mainly focus on improving the classification accuracy while ignoring their high complexity. By considering that the data formats of both HSIs and network weights can be represented in the form of tensors, we develop a new lightweight tensorized neural network for HSI classification that takes advantage of low-rank tensor decomposition techniques to reduce complexity. Firstly, inspired by tensor train (TT)-based tensorized convolutional layers, a new tensorized 2D convolutional layer based on chain calculation (with better expression ability) is introduced. Based on this innovation, a new lightweight 2D tensorized neural network (2D-TNN) is designed for HSI classification. Furthermore, to better preserve the intrinsic structure of HSI data, a new lightweight 3D tensorized neural network (3D-TNN) is proposed by extending the tensorized 2D convolutional layers to their 3D versions. Quantitative and comparative experiments on three widely used data sets show that the proposed models are able to achieve state-of-the-art performance (with a low number of model parameters) for different training sample sizes, especially for very small training sets.
Tian-Yu Ma, Heng-Chao Li 0001, Rui Wang 0090, Qian Du 0001, Xiuping Jia, Antonio Plaza
IEEE Trans. Geosci. Remote. Sens.4
2022 Hyperspectral Image Classification Using Attention-Based Bidirectional Long Short-Term Memory Network
abstract
Deep neural networks have been widely applied to hyperspectral image (HSI) classification areas, in which recurrent neural network (RNN) is one of the most typical networks. Most of the existing RNN-based classifiers treat the spectral signature of pixels as an ordered sequence, in which only unidirectional correlation along the wavelength direction of adjacent bands is considered. However, each band image is related to not only its preceding band images but also its successive band images. In order to fully explore such bidirectional spectral correlation within an HSI, in this article, a bidirectional long short-term memory (Bi-LSTM)-based network is designed for HSI classification. Moreover, a spatial–spectral attention mechanism is designed and implemented in the proposed Bi-LSTM network to emphasize the effective information and reduce the redundant information among spatial–spectral context of pixels, by which the performance of classification can be greatly improved. Experimental results over three benchmark HSIs, i.e., Salinas Valley, Pavia Centre, and Pavia University, demonstrate that our proposed Bi-LSTM obviously outperforms several state-of-the-art unidirectional RNN-based classification algorithms. Moreover, the proposed spatial–spectral attention mechanism can further improve the classification accuracy of our proposed Bi-LSTM algorithm by effectively weighting spatial and spectral context of pixels. The source code of the proposed Bi-LSTM algorithm is available athttps://github.com/MeiShaohui/Attention-based-Bidirectional-LSTM-Network.
Shaohui Mei, Huimin Cai, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.5
2022 SSCU-Net: Spatial-Spectral Collaborative Unmixing Network for Hyperspectral Images
abstract
Linear spectral unmixing is an essential technique in hyperspectral image (HSI) processing and interpretation. In recent years, deep learning-based approaches have shown great promise in hyperspectral unmixing (HU), in particular, unsupervised unmixing methods based on autoencoder (AE) networks are a recent trend. The AE model, which automatically learns low-dimensional representations (abundances) and reconstructs data with their corresponding bases (endmembers), has achieved superior performance in HU. In this article, we explore the effective utilization of spatial and spectral information in AE-based unmixing networks. Important findings on the use of spatial and spectral information in the AE framework are discussed. Inspired by these findings, we propose a spatial–spectral collaborative unmixing network, called SSCU-Net, which learns a spatial AE network and a spectral AE network in an end-to-end manner to more effectively improve the unmixing performance. SSCU-Net is a two-stream deep network and shares an alternating architecture, where the two AE networks are efficiently trained in a collaborative way for estimation of endmembers and abundances. Meanwhile, we propose a new spatial AE network by introducing a superpixel segmentation method based on abundance information, which greatly facilitates the employment of spatial information and improves the accuracy of unmixing network. Moreover, extensive ablation studies are carried out to investigate the performance gain of SSCU-Net. Experimental results on both synthetic and real hyperspectral datasets illustrate the effectiveness and competitiveness of the proposed SSCU-Net compared with several state-of-the-art HU methods.
Lin Qi 0004, Feng Gao 0005, Junyu Dong, Xinbo Gao 0001, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.5
2022 A Dual-Branch Detail Extraction Network for Hyperspectral Pansharpening
abstract
Hyperspectral (HS) pansharpening aims at creating a high-resolution hyperspectral (HR-HS) image by integrating a high spatial resolution panchromatic (HR-PAN) image with a low-resolution hyperspectral (LR-HS) image. It is an important preprocessing procedure in many remote sensing tasks. Most of the existing pansharpening methods train a specific convolutional neural network (CNN) model for each type of dataset with the same number of spectral bands. The main contribution of this study is to propose a new dual-branch detail extraction pansharpening network (called DBDENet) that can sharpen HS images with any number of spectral bands using a single pre-trained model by fine-tuning the parameters of a small module in the network. Specifically, DBDENet extracts spatial details from LR-HS and HR-PAN images by two bidirectional branches of the dual-branch detail extraction network level by level. For each level, the spatial details captured from the HR-PAN and those of the LR-HS images are fused by a spatial cross attention fusion module (SCAFM). The spatial details fused by the last SCAFM module are injected into the upsampled HS image to obtain an HR-HS image. Experimental results prove to show the proposed DBDENet is superior to other widely accepted state-of-the-art methods in terms of objective indicators and visual appearance.
Jiahui Qu, Shaoxiong Hou, Wenqian Dong, Song Xiao 0001, Qian Du 0001, Yunsong Li 0001
IEEE Trans. Geosci. Remote. Sens.5
2022 MSSL: Hyperspectral and Panchromatic Images Fusion via Multiresolution Spatial-Spectral Feature Learning Networks
abstract
The fusion of hyperspectral (HS) and panchromatic (PAN) images aims to generate a fused HS image that combines spectral information of the HS image with spatial information of the PAN image. In this article, we propose a multiresolution spatial–spectral feature learning (MSSL) framework for fusing HS and PAN images. The proposed MSSL transforms the existing deep and complex network into several simple and shallow subnetworks to simplify the feature learning process. MSSL upsamples the HS image while downsamples the PAN image and designs multiresolution 3-D convolutional autoencoder (CAEs) networks with a spectral constraint to learn complete spatial–spectral features of the HS image. MSSL designs multiresolution 2-D CAEs with spatial constraint to extract spatial features of the PAN image, with a low computational cost. In order to effectively generate the pansharpened HS image with high spatial and spectral fidelity, a multiresolution residual network is presented to reconstruct the HS image from the extracted spatial–spectral features. Extensive experiments are conducted on three widely used remote sensing data sets in comparison with state-of-the-art HS image fusion methods, demonstrating the superiority of the proposed MSSL method. Code is available athttps://github.com/Jiahuiqu/MSSL.
Jiahui Qu, Yanzi Shi, Weiying Xie, Yunsong Li 0001, Xianyun Wu, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.6
2022 Dual-Branch Difference Amplification Graph Convolutional Network for Hyperspectral Image Change Detection
abstract
Hyperspectral image (HSI) change detection aims to identify the differences in multitemporal HSIs. Recently, a graph convolutional network (GCN) has attracted increasing attention in the field of remote sensing due to its advantages in processing irregular data. In comparison with a convolutional neural network (CNN) that can only perform convolution operations on data with the assumption of the Euclidean structure, GCN adopts a graph structure to flexibly capture the characteristics and structure information of non-Euclidean data. In this article, we propose a novel dual-branch difference amplification GCN (D2AGCN) for HSI change detection with limited samples, which allows the network to fully extract and effectively amplify the difference features of multitemporal HSIs for change detection. The dual-branch structure can effectively extract sufficient different features to facilitate the detection of the changed areas. As far as we know, this is the first time that GCN has been introduced into HSI change detection. A difference magnification module is designed to suppress similar regions and highlight the feature differences between the multitemporal HSIs in the dual-branch structure, which increases the distinction between change and nonchange classes. The visual and quantitative experimental results on three real hyperspectral datasets (i.e., China, Bay Area, and Santa Barbara) show that the proposed D2AGCN outperforms most of the state-of-the-art methods in HSI change detection with limited training samples.
Jiahui Qu, Yunshuang Xu, Wenqian Dong, Yunsong Li 0001, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.5
2022 BDANet: Multiscale Convolutional Neural Network With Cross-Directional Attention for Building Damage Assessment From Satellite Images
abstract
Fast and effective responses are required when a natural disaster (e.g., earthquake and hurricane) strikes. Building damage assessment from satellite imagery is critical before relief effort is deployed. With a pair of predisaster and postdisaster satellite images, building damage assessment aims at predicting the extent of damage to buildings. With the powerful ability of feature representation, deep neural networks have been successfully applied to building damage assessment. Most existing works simply concatenate predisaster and postdisaster images as input of a deep neural network without considering their correlations. In this article, we propose a novel two-stage convolutional neural network for building damage assessment, called BDANet. In the first stage, a U-Net is used to extract the locations of buildings. Then, the network weights from the first stage are shared in the second stage for building damage assessment. In the second stage, a two-branch multiscale U-Net is employed as the backbone, where predisaster and postdisaster images are fed into the network separately. A cross-directional attention module is proposed to explore the correlations between predisaster and postdisaster images. Moreover, CutMix data augmentation is exploited to tackle the challenge of difficult classes. The proposed method achieves state-of-the-art performance on a large-scale dataset—xBD. The code is available athttps://github.com/ShaneShen/BDANet-Building-Damage-Assessment.
Sijie Zhu, Taojiannan Yang, Chen Chen 0001, Delu Pan, Jianyu Chen 0003, Liang Xiao 0001, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.8
2022 Superpixel-Based Relaxed Collaborative Representation With Band Weighting for Hyperspectral Image Classification
abstract
Representation learning methods, such as sparse representation (SR) and collaborative representation (CR), have been widely used in hyperspectral image classification. However, they merely considered the similarities between features. Due to the plentiful spatial and spectral information in hyperspectral images, the differences between features also need to be considered. Relaxed CR (RCR) is used in face recognition to accommodate the difference and similarity of features simultaneously. In this article, a novel method of RCR with band weighting based on superpixel segmentation is proposed for hyperspectral image classification. The$\boldsymbol {l}_{ \boldsymbol {2}}$norm on band coefficients and global average coefficients is exploited to ensure the similarity, and the variance determines the specific coefficient-related weight of each band. The training set is selected from each superpixel, which is considered as a subgraph rather than independent pixels. It is favorable for concentrating on the difference between similar bands since the samples in each superpixel are of high similarity. Furthermore, extended multiattribute profile (EMAP) features, Gabor features, and local binary pattern (LBP) features are employed to increase the diversity of features; thus, a method of multifeatures’ RCR based on superpixels is proposed. Three typical data are used to validate the related algorithms. The experiments demonstrate that the proposed algorithms can effectively improve classification accuracy compared to state-of-the-art classifiers.
Hongjun Su, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.3
2022 A Multiscale Spectral Features Graph Fusion Method for Hyperspectral Band Selection
abstract
This article proposes a multiscale spectral features graph fusion (MSFGF) method for selecting proper hyperspectral bands. The MSFGF regards that the selected bands should reflect diagnostic spectral information of ground objects at different scales, and it explores band selection from the aspect of multiple spatial scales. First, it adopts the multiscale low-rank decomposition (MSLRD) model to find multiscale spectral features of different ground objects. The model considers divergent spatial structures or spatial correlations of ground objects at different scales, and factorizes the hyperspectral data cube into a series of low-rank block-wise data cubes, where the blocks take spatial structures of different ground objects at increasing scales. Second, the MSFGF presents the multiscale sparse spectral clustering (MSSC) model to fuse the separate connected graphs of multiscale spectral features into a consensus graph. The consensus graph combines the complementary information of multiscale spectral features and helps to reveal the intrinsic clustering structure of all spectral bands. Finally, the MSFGF utilizes spectral clustering to find clusters from the consensus graph and selects representative bands. Experimental results on three widely used hyperspectral data prove the superiority of MSFGF in selecting bands, where it outperforms other seven state-of-the-art methods in classification with an acceptable computational cost.
Weiwei Sun 0005, Gang Yang 0006, Jiangtao Peng, Xiangchao Meng, Wei Li 0032, Heng-Chao Li 0001, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.8
2022 Corrections to "Multiscale Context-Aware Ensemble Deep KELM for Efficient Hyperspectral Image Classification"
abstract
In the above article[1],Fig. 19was incorrectly placed. The correct image and caption are provided here:
Bobo Xi, Jiaojiao Li 0001, Yunsong Li 0001, Rui Song 0003, Weiwei Sun 0005, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.6
2022 Multi-Direction Networks With Attentional Spectral Prior for Hyperspectral Image Classification
abstract
Convolutional neural networks (CNNs) have achieved prominent progress in recent years and demonstrated remarkable properties in spectral–spatial hyperspectral image (HSI) classification. However, conventional spatial-context-based CNNs commonly adopt the single patchwise scheme to represent the to-be-classified samples, which often fails to completely investigate the wealthy spectral–spatial information in complicated situations. For instance, it has great probability to cause misclassifications on the irregular or inhomogeneous areas, especially for the borders across different classes. To counteract this deficiency, we propose a unified multi-direction network (MDN) for HSI Classification (HSIC), which can exhaustively explore the abundant spectral and detailed spatial-context information through multi-direction samples. Additionally, considering the image-spectrum merged structure of the HSI, 3-D Squeeze-and-Excitation residual (3DSERes) blocks are devised in each stream of the framework to consecutively learn the spectral and spatial from low-level to high-level features. Specifically, 3DSERes can not only facilitate fluent gradient in backpropagation through skip connections, but also emphasize the significant spectral–spatial features and constrain the futile ones. This characteristic is beneficial to enhance the model’s generalization capability even with limited training samples. Furthermore, for properly aggregating the multi-direction deep features, we exploit the simple, yet effective attentional spectral prior (ASP) creatively through leveraging the original spectral correlations. Extensive experimental results on three benchmark data sets indicate that the proposed MDN-ASP can achieve promising classification performance compared to the state-of-the-art methods.
Bobo Xi, Jiaojiao Li 0001, Yunsong Li 0001, Rui Song 0003, Yuchao Xiao, Yanzi Shi, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.7
2022 Spectral Distribution-Aware Estimation Network for Hyperspectral Anomaly Detection
abstract
Recently developed deep learning-based hyperspectral anomaly detection (HAD) methods typically include two steps where the deep feature extraction is not designed specifically for the HAD task. In this article, we propose a spectral distribution-aware estimation network (SDEN) that does not conduct feature extraction and anomaly detection in two separate steps but instead learns both jointly to estimate anomalies directly in an end-to-end manner without postprocessing. The unified framework can ensure that the extracted features serve better for anomaly detection. To preserve the distribution of hyperspectral images (HSIs) during dimensionality reduction, the SDEN introduces a spectral distribution (SD)-aware module imposed with a local-invariant constraint. More specifically, we adopt Markov chain Monte Carlo (MCMC) that enables the SD module to better estimate the distribution of the complex HSIs. Considering the powerful representation capability of Gaussian mixture model (GMM), the SDEN leverages it to establish an estimation module in the deep latent space where the anomaly resides in low density while the background not. We demonstrate that the SDEN yields competitive and highly promising results in comparison with the anomaly detection benchmarks.
Weiying Xie, Shuran Fan, Jiahui Qu, Xianyun Wu, Yanli Lu, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.6
2022 Spectral Variation Augmented Representation for Hyperspectral Imagery Classification With Few Labeled Samples
abstract
Due to variation of imaging conditions, spectra of the same type of ground objects usually exhibit certain discrepancy, leading to intra-class spectral distance increase and inter-class distance decrease. As a result, classification accuracy is greatly affected, especially in cases with few labeled samples. For representation based classifiers, the spectral variability within limited training samples is far from sufficient to represent diverse variations within testing ones. To handle this problem, a spectral variation augmented representation for hyperspectral imagery classification (SVARC) with few labeled samples is proposed in this article. Firstly, a novel class-independent and class-dependent components based linear representation model (CICD-LRM) is proposed to emphasize the representation of spectral variation. Secondly, depending on spatial and spectral correlation, the CICD-LRM guided global and local spectral variation extraction schemes are designed, and a fused spectral variation dictionary is constructed by concatenation. Finally, a classifier for hyperspectral images based on the CICD-LRM and spectral variation dictionary is proposed, and specifically three different spectral variation reconstruction strategies are designed. Similar to most of the representation based classifiers, residual-driven decision is also employed in the proposed classifier. Comparative experiments are conducted with eight classical and state-of-the-art methods using two benchmark datasets. The experimental results demonstrate that the proposed SVARC method significantly outperforms the compared ones in cases with few labeled samples.
Bobo Xie, Yifan Zhang 0006, Shaohui Mei, Ge Zhang 0006, Yan Feng 0005, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.6
2022 Dual-Channel Residual Network for Hyperspectral Image Classification With Noisy Labels
abstract
Hyperspectral image (HSI) classification has drawn increasing attention recently. However, it suffers from noisy labels that may occur during field surveys due to a lack of prior information or human mistakes. To address this issue, this article proposes a novel dual-channel residual network (DCRN) to resolve HSI classification with noisy labels. Currently, the influence of noisy labels is reduced by simply detecting and removing those anomalous samples. Different from such a specifically designed noise cleansing method, DCRN is easy to implement but highly effective. It enhances its model robustness to noisy labels to a great extent by employing a novel dual-channel structure and a noise-robust loss function. In this way, DCRN can mitigate influence from noisy labels while fully utilizing useful information from mislabeled samples for augmented training. Experiments are conducted on several hyperspectral data sets with manually generated noisy labels to demonstrate its excellent performance. The code is available athttps://github.com/Li-ZK/DCRN-2021.
Yimin Xu, Zhaokui Li, Wei Li 0032, Qian Du 0001, Cuiwei Liu, Zhuoqun Fang, Lin Zhai
IEEE Trans. Geosci. Remote. Sens.4
2022 A Deep Multiscale Pyramid Network Enhanced With Spatial-Spectral Residual Attention for Hyperspectral Image Change Detection
abstract
Change detection plays an important role in Earth surface observation and has been extensively investigated over recent decades. A hyperspectral image (HSI) with high spectral resolution provides abundant ground object information, which is expected by finer change detection. The existing convolutional neural network (CNN)-based methods extract image features with a fixed kernel, which is incompetent to cope with complicated object details at diverse scales in HSI. In this article, we propose a deep multiscale pyramid network enhanced with spatial–spectral residual attention (DMP$\text {s}^{2} $raN) for HSI change detection, which has strong capability to mine multilevel and multiscale spatial–spectral features, improving the performance in complex changed regions. There are two key characteristics: 1) the multiscale spatial–spectral features are extracted by the multiscale pyramid convolution and enhanced by spatial–spectral residual attention module ($\text {S}^{2} $RAM) of each scale and 2) the multilevel features are obtained by aggregating the multiscale features level by level. As a result of this design, the proposed DMP$\text {s}^{2} $raN learns more discriminative features with both strong semantic information and rich spatial–spectral information. Experiments carried out on three datasets demonstrate the competitive performance of the proposed method in both qualitative and quantitative analyses.
Jiahui Qu, Song Xiao 0001, Wenqian Dong, Yunsong Li 0001, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.6
2022 A Novel Cross-Scale Octave Network for Hyperspectral and Multispectral Image Fusion
abstract
Recently, deep convolutional neural network-based low-resolution hyperspectral image (LR-HSI) and high-resolution multispectral image (HR-MSI) fusion methods have achieved significant performance improvement. However, the rich spatial and spectral information in HSIs is not fully explored. In this article, we propose a novel cross-scale octave network (CSONet) for hyperspectral and multispectral image fusion. Specifically, we adopt a progressive image fusion structure to effectively extract the spatial and spectral information of HR-MSI at multiple resolutions, thereby efficiently complementing LR-HSI’s information. In addition, the proposed cross-scale octave convolution module can extract rich multiscale spatial feature information and concentrate on more important spatial–spectral features at different scales with the multiscale spatial–spectral attention mechanism. Finally, a multisupervised loss function is used to improve the gradient propagation and enhance the representation ability of the network. Ablation analysis on the benchmark datasets shows the effectiveness of each component in the proposed method. Extensive experimental results on different hyperspectral images demonstrate that the proposed CSONet can achieve superior results and strong generalization ability in comparison with some state-of-the-art LR-HSI and HR-MSI fusion methods.
Tianming Zhan, Zuolin Bi, Huapeng Wu, Qian Du 0001, Yang Xu 0006, Zebin Wu 0001
IEEE Trans. Geosci. Remote. Sens.5
2022 Information Fusion for Classification of Hyperspectral and LiDAR Data Using IP-CNN
abstract
Joint use of multisensor information has attracted considerable attention in the remote sensing community. While applications in land-cover observation benefit from information diversity, multisensor integration technique is confronted with many challenges, including inconsistent size of data, different data structures, uncorrelated physical properties, and scarcity of training data. In this article, an information fusion network, named interleaving perception convolutional neural network (IP-CNN), is proposed for integrating heterogeneous information and improving joint classification performance of hyperspectral image (HSI) and light detection and ranging (LiDAR) data. Specifically, a bidirectional autoencoder is designed to reconstruct hyperspectral and LiDAR data together, and the reconstruction process is trained with no dependence upon annotated information. Both HSI-perception constraint and LiDAR-perception constraint are imposed on multisource structural information integration. Accordingly, fused data are fed into a two-branch CNN for final classification. To validate the effectiveness of the model, the experiments were conducted using three datasets (i.e., Muufl Gulfport data, Trento data, and Houston data). The final results demonstrate that the proposed framework can significantly outperform state-of-the-art methods even with small-size training samples.
Mengmeng Zhang 0005, Wei Li 0032, Ran Tao 0003, Heng-Chao Li 0001, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.5
2022 Spectral Variability Augmented Sparse Unmixing of Hyperspectral Images
abstract
Spectral unmixing expresses the mixed pixels existing in hyperspectral images as the product of endmembers and their corresponding fractional abundances, which has been widely used in hyperspectral imagery analysis. However, the endmember spectra even for pixels from the same material of an image may include variability due to the influence of lighting conditions and inherent properties of materials within different pixels. Though thein situspectral library has been used to accommodate such variability by using multiplein situspectra to represent each kind of material, the performance improvement may be restricted due to the limited number of endmembers for each material. Therefore, in this article, spectral variability is directly extracted from anin situendmember library and considered to be transferable among different endmembers for the first time. Furthermore, such a spectral variability is further used to augment sparse unmixing by synchronously performing endmember-based reconstruction and spectral variability-augmented reconstruction in the sparse unmixing model. By, respectively, imposing sparse and smoothness regularization over abundances and variability coefficients, a convex optimization-based spectral variability augmented sparse unmixing (SVASU) is finally proposed, and its convergence performance is also analyzed. Experiments conducted over synthetic and real-world datasets demonstrate that the proposed SVASU method not only significantly improves the unmixing performance of conventional spectral library-based unmixing but also outperforms several state-of-the-art sparse unmixing algorithms.
Ge Zhang 0006, Shaohui Mei, Bobo Xie, Mingyang Ma 0004, Yifan Zhang 0006, Yan Feng 0005, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.7
2022 Rank-Aware Generative Adversarial Network for Hyperspectral Band Selection
abstract
Traditional clustering-based band selection (BS) methods treat each band as individuals, and selection is conducted by enlarging the difference between clusters, which leads to the loss of band interaction and information saliency evaluation. In this article, we propose a BS method named rank-aware generative adversarial network (R-GAN) to address these problems. First, centralized reference feature extraction (FE) with GAN aids R-GAN to combine interpretability and interband relevance. Then, the reference feature is refined with the saliency estimation provided by the rank-aware strategy. According to data characteristics, there are two versions of rank computation including tensor and matrix. Finally, the structural similarity index measurement (SSIM) maps the saliency to the original data space to obtain the final BS result. Extensive comparison experiments with popular existing BS approaches on five hyperspectral images (HSIs) datasets show that the proposed R-GAN can address spectral saliency effectively and select more informative band subsets, which outperforms other competitors for both detection and classification tasks. For example, on the SD-1 dataset, the ten bands selected by R-GAN achieve 0.982 ± 0.003 with an improvement of 13.7% in the area under the curve (AUC) value of anomaly detection performance. The peaked accuracy surpasses the baseline by 0.46% for the classification on the PaviaU dataset.
Xin Zhang 0092, Weiying Xie, Yunsong Li 0001, Jie Lei 0001, Qian Du 0001, Geng Yang 0001
IEEE Trans. Geosci. Remote. Sens.5
2022 GCFnet: Global Collaborative Fusion Network for Multispectral and Panchromatic Image Classification
abstract
Among various multimodal remote sensing data, the pairing of multispectral (MS) and panchromatic (PAN) images is widely used in remote sensing applications. This article proposes a novel global collaborative fusion network (GCFnet) for joint classification of MS and PAN images. In particular, a global patch-free classification scheme based on an encoder-decoder deep learning (DL) network is developed to exploit context dependencies in the image. The proposed GCFnet is designed based on a novel collaborative fusion architecture, which mainly contains three parts: 1) two shallow-to-deep feature fusion branches related to individual MS and PAN images; 2) a multiscale cross-modal feature fusion branch of the two images, where an adaptive loss weighted fusion strategy is designed to calculate the total loss of two individual and the cross-modal branches; 3) a probability weighted decision fusion strategy for the fusion of the classification results of three branches to further improve the classification performance. Experimental results obtained on three real datasets covering complex urban scenarios confirm the effectiveness of the proposed GCFnet in terms of higher accuracy and robustness compared to existing methods. By utilizing both sampled and non-sampled position data in the feature extraction process, the proposed GCFnet can achieve excellent performance even in a small sample-size case. The codes will be available from the website: https://github.com/SicongLiuRS/GCFnet.
Sicong Liu 0001, Qian Du 0001, Lorenzo Bruzzone, Kecheng Du, Xiaohua Tong, Huan Xie 0001
IEEE Trans. Geosci. Remote. Sens.3
2022 Generative Dual-Adversarial Network With Spectral Fidelity and Spatial Enhancement for Hyperspectral Pansharpening
abstract
Hyperspectral (HS) pansharpening is of great importance in improving the spatial resolution of HS images for remote sensing tasks. HS image comprises abundant spectral contents, whereas panchromatic (PAN) image provides spatial information. HS pansharpening constitutes the possibility for providing the pansharpened image with both high spatial and spectral resolution. This article develops a specific pansharpening framework based on a generative dual-adversarial network (called PS-GDANet). Specifically, the pansharpening problem is formulated as a dual task that can be solved by a generative adversarial network (GAN) with two discriminators. The spatial discriminator forces the intensity component of the pansharpened image to be as consistent as possible with the PAN image, and the spectral discriminator helps to preserve spectral information of the original HS image. Instead of designing a deep network, PS-GDANet extends GANs to two discriminators and provides a high-resolution pansharpened image in a fraction of iterations. The experimental results demonstrate that PS-GDANet outperforms several widely accepted state-of-the-art pansharpening methods in terms of qualitative and quantitative assessment.
Wenqian Dong, Shaoxiong Hou, Song Xiao 0001, Jiahui Qu, Qian Du 0001, Yunsong Li 0001
IEEE Trans. Neural Networks Learn. Syst.5
2022 Weakly Supervised Discriminative Learning With Spectral Constrained Generative Adversarial Network for Hyperspectral Anomaly Detection
abstract
Anomaly detection (AD) using hyperspectral images (HSIs) is of great interest for deep space exploration and Earth observations. This article proposes a weakly supervised discriminative learning with a spectral constrained generative adversarial network (GAN) for hyperspectral anomaly detection (HAD), called weaklyAD. It can enhance the discrimination between anomaly and background with background homogenization and anomaly saliency in cases where anomalous samples are limited and sensitive to the background. A novel probability-based category thresholding is first proposed to label coarse samples in preparation for weakly supervised learning. Subsequently, a discriminative reconstruction model is learned by the proposed network in a weakly supervised fashion. The proposed network has an end-to-end architecture, which not only includes an encoder, a decoder, a latent layer discriminator, and a spectral discriminator competitively but also contains a novel Kullback-Leibler (KL) divergence-based orthogonal projection divergence (OPD) spectral constraint. Finally, the well-learned network is used to reconstruct HSIs captured by the same sensor. Our work paves a new weakly supervised way for HAD, which intends to match the performance of supervised methods without the prerequisite of manually labeled data. Assessments and generalization experiments over real HSIs demonstrate the unique promise of such a proposed approach.
Tao Jiang 0031, Weiying Xie, Yunsong Li 0001, Jie Lei 0001, Qian Du 0001
IEEE Trans. Neural Networks Learn. Syst.5
2022 A3 CLNN: Spatial, Spectral and Multiscale Attention ConvLSTM Neural Network for Multisource Remote Sensing Data Classification
abstract
The problem of effectively exploiting the information multiple data sources has become a relevant but challenging research topic in remote sensing. In this article, we propose a new approach to exploit the complementarity of two data sources: hyperspectral images (HSIs) and light detection and ranging (LiDAR) data. Specifically, we develop a new dual-channel spatial, spectral and multiscale attention convolutional long short-term memory neural network (called dual-channel$A^{3}$CLNN) for feature extraction and classification of multisource remote sensing data. Spatial, spectral, and multiscale attention mechanisms are first designed for HSI and LiDAR data in order to learn spectral- and spatial-enhanced feature representations and to represent multiscale information for different classes. In the designed fusion network, a novel composite attention learning mechanism (combined with a three-level fusion strategy) is used to fully integrate the features in these two data sources. Finally, inspired by the idea of transfer learning, a novel stepwise training strategy is designed to yield a final classification result. Our experimental results, conducted on several multisource remote sensing data sets, demonstrate that the newly proposed dual-channel$A^{\,3}$CLNN exhibits better feature representation ability (leading to more competitive classification performance) than other state-of-the-art methods.
Heng-Chao Li 0001, Wen-Shuai Hu, Wei Li 0032, Jun Li 0009, Qian Du 0001, Antonio Plaza
IEEE Trans. Neural Networks Learn. Syst.5
2022 Prior-Based Tensor Approximation for Anomaly Detection in Hyperspectral Imagery
abstract
The key to hyperspectral anomaly detection is to effectively distinguish anomalies from the background, especially in the case that background is complex and anomalies are weak. Hyperspectral imagery (HSI) as an image–spectrum merging cube data can be intrinsically represented as a third-order tensor that integrates spectral information and spatial information. In this article, a prior-based tensor approximation (PTA) is proposed for hyperspectral anomaly detection, in which HSI is decomposed into a background tensor and an anomaly tensor. In the background tensor, a low-rank prior is incorporated into spectral dimension by truncated nuclear norm regularization, and a piecewise-smooth prior on spatial dimension can be embedded by a linear total variation-norm regularization. For anomaly tensor, it is unfolded along spectral dimension coupled with spatial group sparse prior that can be represented by the${l}_{2,1}$-norm regularization. In the designed method, all the priors are integrated into a unified convex framework, and the anomalies can be finally determined by the anomaly tensor. Experimental results validated on several real hyperspectral data sets demonstrate that the proposed algorithm outperforms some state-of-the-art anomaly detection methods.
Lu Li 0005, Wei Li 0032, Ying Qu 0001, Chunhui Zhao 0003, Ran Tao 0003, Qian Du 0001
IEEE Trans. Neural Networks Learn. Syst.6
2021 A Patch Tensor-Based Change Detection Method for Hyperspectral Images
abstract
With the increasing of hyperspectral datasets, multi-temporal hyperspectral change detection has gradually attracted re-searcher's attention. Most of traditional change detection methods only consider spectral information, but ignore importance of spatial structure information, which leads to low detection accuracy. In this work, a novel patch tensor-based change detection method (PTCD) is proposed for hyperspectral imagery to make full use of spatial structure information. Firstly, the tensor decomposition and reconstruction strategies are used to eliminate influence of various factors in multi-temporal dataset. Meanwhile, patch-based strategy is adopted to incorporate the non-overlapping local similar property into the proposed method to exploit spatial structural information. Finally, a specially designed detector is adopted to further improve the detection accuracy. Experiments conducted on two real hyperspectral datasets demonstrate that the proposed detector achieves better detection performance.
Zengfu Hou, Wei Li 0032, Qian Du 0001
IGARSS3
2021 Hyperspectral Image Classification by Fractional Discrete Cosine Transform Based Feature Extraction
abstract
Feature extraction (FE) can greatly improve classification performance, when dealing with hyperspectral imagery (HSI) which is notable by its high dimensionality. Unsupervised FE methods have proven to perform competitively in this arena. Preprocessing methods have been used to further improve their performance. One such an approach to improving FE for classification of HSI is investigated in this paper. This approach is based on exploiting the intermediate nature of spectral signatures transformed via fractional fourier transform (FrFT) as the result from the transform, which carry information from both the original reflectance spectral domain and its fourier transfer domain. The fractional discrete cosine transform (FrDCT) is also investigated for the same purposes. Pairing these methods involved with dimensionality reduction, with a well-known FE algorithm, i.e., principal component analysis (PCA), is explored. Results obtained with a real hyperspectral dataset are promising with preprocessing based on selecting an appropriate fractional transform order, low-pass (LP) filtering in the fractional domain to an appropriate degree, followed by the PCA.
Helgi Hrafn Omarsson, Qian Du 0001
IGARSS2
2021 PTGAN: A Proposal-Weighted Two-Stage GAN with Attention for Hyperspectral Target Detection
abstract
In this paper, a proposal-weighted two-stage generative adversarial network (GAN) with attention mechanism is proposed for hyperspectral target detection (HTD). PTGAN leverages GAN to estimate spectral background distribution and realize mapping from the latent space to the spectral space. Meanwhile, PTGAN conducts the reversed mapping through latent-spectral-latent and spectral-latent-spectral learning. On this basis, PTGAN implements accurate reconstruction of background spectrum via latent space. Therefore, targets of interest can be detected through larger pixel-level reconstruction error. In particular, the variance attention module is designed to make full use of global information among spectral bands to selectively emphasize channel-wise spectral features. Furthermore, a proposal-weighted strategy in a two-stage manner reduces the false alarm of detection by refining the previous detection proposal. Finally, exponential nonlinear fusion combines the discriminative feature from two stages to suppress the background. Extensive experiments on two real hyperspectral images (HSIs) verify the effectiveness of PTGAN.
Weiying Xie, Yunsong Li 0001, Kai Jiang 0001, Jie Lei 0001, Qian Du 0001
IGARSS6
2021 Modified Structure-Aware Collaborative Representation for Hyperspectral Image Classification
abstract
Collaborative representation (CR) is an efficient method for hyperspectral image classification. There exists structure-aware CR with Tikhonov regularization (SaCRT) that utilizes the class label information of training samples into estimation of representation coefficients to provide better performance. It can be further enhanced by considering spatial features because neighboring pixels around the central pixel tend to belong to the same class with high probability. In this paper, a modified SaCRT is proposed for hyperspectral image classification. Its performance is analyzed on different types of spatial features (i.e., spatial averaging features), global feature (i.e., Gabor feature), shape features (i.e., derivative of extended morphological profile (DMP) features), and edge preserving feature. In addition, a majority voting-based ensemble technique is used to enhance the performance by combining different features. The experimental results illustrate that the proposed approach can yield better performance in comparison to state-of-the-art classifiers.
Chiranjibi Shah, Qian Du 0001
IGARSS2
2021 Collaborative and Low-Rank Graph for Discriminant Analysis of Hyperspectral Imagery
abstract
Sparse graph-based discriminant analysis has drawn much attention to represent high dimensional data into low-dimensional subspace by using$l_{1}$-norm optimization. There exists sparse and low-rank graph-based discriminant analysis (SLGDA) for incorporating local and global data structure together by combining both sparsity and low-rankness. Deviating from the concept of sparse representation, collaborative and low-rank representation-based discriminant analysis is proposed (CLGDA) in this paper with the assumption that collaboration among atoms is more important to estimate appropriate representation coefficients by accommodating within-class variation. The experimental results obtained on several hyperspectral datasets illustrate the superior classification performance of the proposed CLGDA in comparison to SLGDA and state-of-the-art approaches.
Chiranjibi Shah, Qian Du 0001
IGARSS2
2021 Hyperspectral Imagery Super-Resolution Based on Self-Calibrated Attention Residual Network
abstract
Hyperspectral remote sensing images are well-known for their abundant spectral characteristics to discriminate different object materials. However, due to the constraints of sensor limitations and exceedingly high acquisition costs, it is difficult to obtain high spatial resolution hyperspectral imagery. Though many methods have been focusing on the restoration of the spatial structure information, spectral information may be over-smoothed during such spatial super-resolution. In this paper, a novel self-calibrated attention residual network (SCARN) is proposed to increase spatial resolution of hyperspectral images while retain spectral consistency. In particular, a self-calibrated attention residual block (SCARB) is elaborately designed to fully exploit the spatial information and the correlation between the spectra of the hyperspectral data. Concretely, self-calibrated convolution, instead of standard convolution, is adopted to adaptively construct long-range spatial and spectral dependencies around each spatial location of hyperspectral imagery, and attention module is inserted to improve the representation ability of spectral information. Finally, global and local residual connections are designed to ease the network training difficulty and maintain a higher restoration accuracy. Experimental results over two benchmark hyperspectral datasets demonstrate the effectiveness and superiority of the proposed SCARN method against the state-of-the-art methods.
Baorui Wang, Shaohui Mei, Yan Feng 0005, Qian Du 0001
IGARSS4
2021 Siammraan: Siamese Multi-Level Residual Attention Adaptive Network for Hyperspectral Videos Tracking
abstract
The deep learning based techniques have been widely applied to object tracking in color videos. When these techniques are applied to hyperspectral videos, how to fully explore unique spectral signatures of tracking objects is of crucial importance as well as simultaneously utilizing spatial and temporal information. Different with color videos, hyperspectral videos record continuous spectral reflectance of targets in light wavelength indexed band images and it is more difficult to explore unique spectral feature of tracking objects. Aiming to take advantage of existing object tracking techniques in color videos, a Siamese Multi-level Residual Attention Adaptive Network (SiamMRAAN) is designed to handle 3-band images by using the well-trained ResNet50 as backbone. By grouping hyperspectral videos into several 3-band-image subsets, the proposed SiamMRAAN can be used to explore high-dimensional spectral information. We design a loss function to fuse the tracking results over these subsets to improve the tracking performance. Finally, experiments over 75 hyperspectral videos confirmed that using spectral information is critical to improve the performance of object tracking in color videos, and also demonstrated that the proposed SiamMRAAN based strategy outperforms several compared networks for hyperspectral videos.
Ye Wang 0020, Shaohui Mei, Qian Du 0001
IGARSS4
2021 Semi-Supervised Graph Prototypical Networks for Hyperspectral Image Classification
abstract
Graph convolutional network (GCN) is one of the most favorable semi-supervised approaches, which demonstrates encouraging performance for hyperspectral image classification (HSIC), especially under the condition of small sample sizes. In this paper, we propose a novel semi-supervised graph prototypical network (SSGPN) for high-precise HSIC. Different from prevenient GCN, we devise a prototypical layer comprising a distance-based cross-entropy (DCE) loss function and a novel temporal entropy-based regularizer (TER) in the frameworks of SSGPN. This effective layer can facilitate to generate more discriminative embedding features along with the representative prototypes to each class, so as to achieve accurate identification of various land-cover categories. Additionally, to promote computational efficiency, we present a graph normalization (G-Norm) to accelerate the convergence speed and boost the training procedure. Experimental results demonstrate that our proposed SSGPN can obtain promising performance compared with the state-of-the-art methods.
Bobo Xi, Jiaojiao Li 0001, Yunsong Li 0001, Qian Du 0001
IGARSS4
2021 Joint feature extraction for multi-source data using similar double-concentrated network
Yixuan Zhu, Wei Li 0032, Mengmeng Zhang 0005, Ran Tao 0003, Qian Du 0001
Neurocomputing6
2021 Generative Adversarial Capsule Network With ConvLSTM for Hyperspectral Image Classification
abstract
Recently, deep learning has been widely applied in hyperspectral image (HSI) classification since it can extract high-level spatial-spectral features. However, deep learning methods are restricted due to the lack of sufficient annotated samples. To address this problem, this letter proposes a novel generative adversarial network (GAN) for HSI classification that can generate artificial samples for data augmentation to improve the HSI classification performance with few training samples. In the proposed network, a new discriminator is designed by exploiting capsule network (CapsNet) and convolutional long short-term memory (ConvLSTM), which extracts the low-level features and combines them together with local space sequence information to form the high-level contextual features. In addition, a structured sparse L2,1constraint is imposed on sample generation to control the modes of data being generated and achieve more stable training. The experimental results on two real HSI data sets show that the proposed method can obtain better classification performance than the several state-of-the-art deep classification methods.
Wei-Ye Wang, Heng-Chao Li 0001, Yangjun Deng, Li-Yang Shao, Xiaoqiang Lu, Qian Du 0001
IEEE Geosci. Remote. Sens. Lett.6
2021 Parallel and Distributed Computing for Anomaly Detection From Hyperspectral Remote Sensing Imagery
abstract
Anomaly detection from remote sensing images is to detect pixels whose spectral signatures are different from their background. Anomalies are often man-made targets. With such target signatures being unknown, anomaly detection has many important applications, such as water quality monitoring, crop stress surveying, and law enforcement-related uses, where prior information of targets is often unavailable. The key to success is accurate background modeling. Anomaly detection from remote sensing images is challenging because spatial coverage is very large and the background is highly heterogeneous. For pixel-based anomaly detection, computing cost in background modeling and a spatial-convolution-type detection process is very expensive. Thus, parallel and distributed computing is critical in reducing execution time, which can fit the need for real-time or near real-time detection from airborne and spaceborne platforms in support of immediate decision-making. This article reviews the recent advances in anomaly detection from hyperspectral remote sensing images and their implementation using parallel and distributed systems. The classical methods, i.e., the Reed-Xiaoli (RX) algorithm and its variants, including its real-time processing version, are illustrated in commodity graphic processing units (GPUs), cloud, and field-programmable gate array (FPGA) implementations. Practical issues and future development trends are also discussed.
Qian Du 0001, Bo Tang 0011, Weiying Xie, Wei Li 0032
Proc. IEEE1
2021 Low-Rank and Sparse Decomposition With Mixture of Gaussian for Hyperspectral Anomaly Detection
abstract
Recently, the low-rank and sparse decomposition model (LSDM) has been used for anomaly detection in hyperspectral imagery. The traditional LSDM assumes that the sparse component where anomalies and noise reside can be modeled by a single distribution which often potentially confuses weak anomalies and noise. Actually, a single distribution cannot accurately describe different noise characteristics. In this article, a combination of a mixture noise model with low-rank background may more accurately characterize complex distribution. A modified LSDM, by modeling the sparse component as a mixture of Gaussian (MoG), is employed for hyperspectral anomaly detection. In the proposed framework, the variational Bayes (VB) algorithm is applied to infer a posterior MoG model. Once the noise model is determined, anomalies can be easily separated from the noise components. Furthermore, a simple but effective detector based on the Manhattan distance is incorporated for anomaly detection under complex distribution. The experimental results demonstrate that the proposed algorithm outperforms the classic Reed-Xiaoli (RX), and the state-of-the-art detectors, such as robust principal component analysis (RPCA) with RX.
Lu Li 0005, Wei Li 0032, Qian Du 0001, Ran Tao 0003
IEEE Trans. Cybern.3
2021 Weakly Supervised Low-Rank Representation for Hyperspectral Anomaly Detection
abstract
In this article, we propose a weakly supervised low-rank representation (WSLRR) method for hyperspectral anomaly detection (HAD), which formulates deep learning-based HAD into a low-lank optimization problem not only characterizing the complex and diverse background in real HSIs but also obtaining relatively strong supervision information. Different from the existing unsupervised and supervised methods, we first model the background in a weakly supervised manner, which achieves better performance without prior information and is not restrained by richly correct annotation. Considering reconstruction biases introduced by the weakly supervised estimation, LRR is an effective method for further exploring the intricate background structures. Instead of directly applying the conventional LRR approaches, a dictionary-based LRR, including both observed training data and hidden learned data drawn by the background estimation model, is proposed. Finally, the derived low-rank part and sparse part and the result of the initial detection work together to achieve anomaly detection. Comparative analyses validate that the proposed WSLRR method presents superior detection performance compared with the state-of-the-art methods.
Weiying Xie, Xin Zhang 0092, Yunsong Li 0001, Jie Lei 0001, Jiaojiao Li 0001, Qian Du 0001
IEEE Trans. Cybern.6
2021 More Diverse Means Better: Multimodal Deep Learning Meets Remote-Sensing Imagery Classification
abstract
Classification and identification of the materials lying over or beneath the earth's surface have long been a fundamental but challenging research topic in geoscience and remote sensing (RS), and have garnered a growing concern owing to the recent advancements of deep learning techniques. Although deep networks have been successfully applied in single-modality-dominated classification tasks, yet their performance inevitably meets the bottleneck in complex scenes that need to be finely classified, due to the limitation of information diversity. In this work, we provide a baseline solution to the aforementioned difficulty by developing a general multimodal deep learning (MDL) framework. In particular, we also investigate a special case of multi-modality learning (MML)-cross-modality learning (CML) that exists widely in RS image classification applications. By focusing on “what,” “where,” and “how” to fuse, we show different fusion strategies as well as how to train deep networks and build the network architecture. Specifically, five fusion architectures are introduced and developed, further being unified in our MDL framework. More significantly, our framework is not only limited to pixel-wise classification tasks but also applicable to spatial information modeling with convolutional neural networks (CNNs). To validate the effectiveness and superiority of the MDL framework, extensive experiments related to the settings of MML and CML are conducted on two different multimodal RS data sets. Furthermore, the codes and data sets will be available at https://github.com/danfenghong/IEEE_TGRS_MDL-RS, contributing to the RS community.
Danfeng Hong, Lianru Gao, Naoto Yokoya, Jing Yao 0002, Jocelyn Chanussot, Qian Du 0001, Bing Zhang 0001
IEEE Trans. Geosci. Remote. Sens.6
2021 Class-Wise Distribution Adaptation for Unsupervised Classification of Hyperspectral Remote Sensing Images
abstract
Class-wise adversarial adaptation networks are investigated for the classification of hyperspectral remote sensing images in this article. By adversarial learning between the feature extractor and the multiple domain discriminators, domain-invariant features are generated. Moreover, a probability-prediction-based maximum mean discrepancy (MMD) method is introduced to the adversarial adaptation network to achieve a superior feature-alignment performance. The class-wise adversarial adaptation in conjunction with the class-wise probability MMD is denoted as the class-wise distribution adaptation (CDA) network. The proposed CDA does not require labeled information in the target domain and can achieve an unsupervised classification of the target image. The experimental results using the Hyperion and Airborne Visible/Infrared Imaging Spectrometer (AVIRIS) hyperspectral data demonstrated its efficiency.
Zixu Liu, Li Ma 0005, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.3
2021 Self-Paced Nonnegative Matrix Factorization for Hyperspectral Unmixing
abstract
The presence of mixed pixels in the hyperspectral data makes unmixing to be a key step for many applications. Unsupervised unmixing needs to estimate the number of endmembers, their spectral signatures, and their abundances at each pixel. Since both endmember and abundance matrices are unknown, unsupervised unmixing can be considered as a blind source separation problem and can be solved by nonnegative matrix factorization (NMF). However, most of the existing NMF unmixing methods use a least-squares objective function that is sensitive to the noise and outliers. To deal with different types of noises in hyperspectral data, such as the noise in different bands (band noise), the noise in different pixels (pixel noise), and the noise in different elements of hyperspectral data matrix (element noise), we propose three self-paced learning based NMF (SpNMF) unmixing models in this article. The SpNMF models replace the least-squares loss in the standard NMF model with weighted least-squares losses and adopt a self-paced learning (SPL) strategy to learn the weights adaptively. In each iteration of SPL, atoms (bands or pixels or elements) with weight zero are considered as complex atoms and are excluded, while atoms with nonzero weights are considered as easy atoms and are included in the current unmixing model. By gradually enlarging the size of the current model set, SpNMF can select atoms from easy to complex. Usually, noisy or outlying atoms are complex atoms that are excluded from the unmixing model. Thus, SpNMF models are robust to noise and outliers. Experimental results on the simulated and two real hyperspectral data sets demonstrate that our proposed SpNMF methods are more accurate and robust than the existing NMF methods, especially in the case of heavy noise.
Jiangtao Peng, Yicong Zhou, Weiwei Sun 0005, Qian Du 0001, Lekang Xia
IEEE Trans. Geosci. Remote. Sens.4
2021 Anomaly Detection in Hyperspectral Imagery Based on Gaussian Mixture Model
abstract
Hyperspectral images (HSIs) with rich spectral information have been widely used in many fields. Anomaly detection is one of the most interesting and important applications. In this article, a novel Gaussian mixture model (GMM)-based anomaly detection (GMMD) method for HSI is proposed. The main contributions of this article are a new GMM-based extraction approach for extracting the anomaly pixels and an effective GMM-based weighting approach for fusing the extracted anomaly results. Specifically, based on the fact that the spectral values of anomaly pixels in some bands are different from those of background pixels, we propose a GMM-based anomaly extraction approach in which the HSI is characterized by the GMM and the anomaly pixels are extracted by a range prescribed by the GMM parameters. In order to fuse the extracted anomaly results, the GMM-based weighting method is introduced to adaptively construct the detection map. The detection map is rectified by using a guided filter to obtain the final anomaly detection map. Experimental results conducted on four hyperspectral data sets demonstrate the superior performance of the proposed GMMD method.
Jiahui Qu, Qian Du 0001, Yunsong Li 0001, Haoming Xia
IEEE Trans. Geosci. Remote. Sens.2
2021 Efficient Deep Learning of Nonlocal Features for Hyperspectral Image Classification
abstract
Deep-learning-based methods, such as convolution neural network (CNN), have demonstrated their efficiency in hyperspectral image (HSI) classification. These methods can automatically learn spectral-spatial discriminative features within local patches. However, for each pixel in an HSI, it is not only related to its nearby pixels but also has connections to pixels far away from itself. Therefore, to incorporate the long-range contextual information, a deep fully convolutional network (FCN) with an efficient nonlocal module, named ENL-FCN, is proposed for HSI classification. In the proposed framework, a deep FCN considers an entire HSI as input and extracts spectral-spatial information in a local receptive field. The efficient nonlocal module is embedded in the network as a learning unit to capture the long-range contextual information. Different from the traditional nonlocal neural networks, the long-range contextual information is extracted in a specially designed criss-cross path for computation efficiency. Furthermore, using a recurrent operation, each pixel's response is aggregated from all pixels of HSI. The benefits of our proposed ENL-FCN are threefold: 1) the long-range contextual information is incorporated effectively; 2) the efficient module can be freely embedded in a deep neural network in a plug-and-play fashion; and 3) it has much fewer learning parameters and requires less computational resources. The experiments conducted on three popular HSI data sets demonstrate that the proposed method achieves state-of-the-art classification performance with lower computational cost in comparison with several leading deep neural networks for HSI.
Sijie Zhu, Chen Chen 0001, Qian Du 0001, Liang Xiao 0001, Jianyu Chen 0003, Delu Pan
IEEE Trans. Geosci. Remote. Sens.4
2021 Sensor-Independent Hyperspectral Target Detection With Semisupervised Domain Adaptive Few-Shot Learning
abstract
Deep learning-based hyperspectral target detection (HTD) is potentially hindered by the limited training samples and sensor-dependent transferability. To address this issue, we propose a novel semisupervised domain adaptive few-shot learning (SDAFL) model to adaptively transfer similarity/dissimilarity measurement from source domain with sufficient labeled samples to target domain in an adversarial manner, where source data and target data can be collected by different sensors, i.e., sensor-independent. In order to alleviate negative transfer, residual channel attention (RCA) and weighted domain adaptation (WDA) are used to automatically select representative features and assign easy-transferred samples with higher priority. In addition, we adopt modulated deformable convolution (MDConv) to make the receptive field fit image spatial structure and also introduce a discriminatively boosted loss (DBL) function based on the prior known target signature to further enhance feature distinction, where intraclass similarity is improved, while interclass similarity is suppressed. After extracting discriminative features through the SDAFL model, guided filter and t-distribution kernel are jointly used for spatial–spectral target detection (S2TD). It should be noted that only the spectral signature of the desired object is needed in the target domain. Experimental results and analysis on three real hyperspectral images (HSIs) verify the efficiency and superiority of our proposed sensor-independent hyperspectral target detection (SIHTD) method compared with other algorithms.
Yanzi Shi, Jiaojiao Li 0001, Yunsong Li 0001, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.4
2021 Random Subspace-Based k-Nearest Class Collaborative Representation for Hyperspectral Image Classification
abstract
Recently, collaborative representation classification (CRC) has attracted extensive interest for hyperspectral images (HSIs) classification. However, for collaborative representation with Tikhonov (CRT), a testing sample is collaboratively represented by training samples from all the classes, which may result in high computational cost. In this article, we select the first$k$class training samples that are nearest to the testing sample for representation, namely,$k$-nearest class CRT (KNCCRT) algorithm. In order to improve the performance of KNCCRT for HSI classification, the idea of random subspace-based KNCCRT ensemble framework is proposed. KNCCRT is adopted as base classifier and random subspace (RS) contributes to diversity by selecting feature randomly. Moreover, to further increase the classification accuracy, shape-adaptive (SA) neighborhood constraint is utilized in RS ensemble framework to incorporate spatial information. Experimental results on three real hyperspectral data sets demonstrate the effectiveness of the proposed methods for HSI classification. The combination of KNCCRT and RS framework provides a reliable accuracy for HSI classification.
Hongjun Su, Zhaoyue Wu, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.4
2021 Adaptive DropBlock-Enhanced Generative Adversarial Networks for Hyperspectral Image Classification
abstract
In recent years, the hyperspectral image (HSI) classification based on generative adversarial networks (GANs) has achieved great progress. GAN-based classification methods can mitigate the limited training sample dilemma to some extent. However, several studies have pointed out that existing GAN-based HSI classification methods are heavily affected by the imbalanced training data problem. The discriminator in GAN always contradicts itself and tries to associate fake labels to the minority-class samples and, thus, impair the classification performance. Another critical issue is the mode collapse in GAN-based methods. The generator is only capable of producing samples within a narrow scope of the data space, which severely hinders the advancement of GAN-based HSI classification methods. In this article, we proposed an Adaptive DropBlock-enhanced Generative Adversarial Networks (ADGANs) for HSI classification. First, to solve the imbalanced training data problem, we adjust the discriminator to be a single classifier, and it will not contradict itself. Second, an adaptive DropBlock (AdapDrop) is proposed as a regularization method employed in the generator and discriminator to alleviate the mode collapse issue. The AdapDrop generated drop masks with adaptive shapes instead of a fixed size region, and it alleviates the limitations of DropBlock in dealing with ground objects with various shapes. Experimental results on three HSI data sets demonstrated that the proposed ADGAN achieved superior performance over state-of-the-art GAN-based methods. Our codes are available at https://github.com/summitgao/HC_ADGAN.
Feng Gao 0005, Junyu Dong, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.4
2021 Multiscale Context-Aware Ensemble Deep KELM for Efficient Hyperspectral Image Classification
abstract
Recently, multiscale spatial features have been widely utilized to improve the hyperspectral image (HSI) classification performance. However, fixed-size neighborhood involving the contextual information probably leads to misclassifications, especially for the boundary pixels. Additionally, it has been demonstrated that deep neural network (DNN) is practical to extract representative features for the classification tasks. Nevertheless, under the condition of high dimensionality versus small sample sizes, DNN tends to be over-fitting and it is generally time-consuming due to the deep-level feature learning process. To alleviate the aforementioned issues, we propose a multiscale context-aware ensemble deep kernel extreme learning machine (MSC-EDKELM) for efficient HSI classification. First, the scene of the HSI data set is over-segmented in multiscale via using the adaptive superpixel segmentation technique. Second, superpixel pattern (SP) and attentional neighboring superpixel pattern (ANSP) are generated by leveraging the superpixel maps, which can automatically comprise local and global contextual information, respectively. Afterward, an ensemble deep kernel extreme learning machine (EDKELM) is presented to investigate the deep-level characteristics in the SP and ANSP. Finally, the category of each pixel is accurately determined by the decision fusion and weighted output layer fusion strategy. Experimental results on four real-world HSI data sets demonstrate that the proposed frameworks outperform some classic and state-of-the-art methods with high computational efficiency, which can be employed to serve real-time applications.
Bobo Xi, Jiaojiao Li 0001, Yunsong Li 0001, Rui Song 0003, Weiwei Sun 0005, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.6
2021 HPGAN: Hyperspectral Pansharpening Using 3-D Generative Adversarial Networks
abstract
Hyperspectral (HS) pansharpening, as a special case of the superresolution (SR) problem, is to obtain a high-resolution (HR) image from the fusion of an HR panchromatic (PAN) image and a low-resolution (LR) HS image. Though HS pansharpening based on deep learning has gained rapid development in recent years, it is still a challenging task because of the following requirements: 1) a unique model with the goal of fusing two images with different dimensions should enhance spatial resolution while preserving spectral information; 2) all the parameters should be adaptively trained without manual adjustment; and 3) a model with good generalization should overcome the sensitivity to different sensor data in reasonable computational complexity. To meet such requirements, we propose a unique HS pansharpening framework based on a 3-D generative adversarial network (HPGAN) in this article. The HPGAN induces the 3-D spectral-spatial generator network to reconstruct the HR HS image from the newly constructed 3-D PAN cube and the LR HS image. It searches for an optimal HR HS image by successive adversarial learning to fool the introduced PAN discriminator network. The loss function is specifically designed to comprehensively consider global constraint, spectral constraint, and spatial constraint. Besides, the proposed 3-D training in the high-frequency domain reduces the sensitivity to different sensor data and extends the generalization of HPGAN. Experimental results on data sets captured by different sensors illustrate that the proposed method can successfully enhance spatial resolution and preserve spectral information.
Weiying Xie, Yuhang Cui, Yunsong Li 0001, Jie Lei 0001, Qian Du 0001, Jiaojiao Li 0001
IEEE Trans. Geosci. Remote. Sens.5
2021 Cross-Scene Hyperspectral Image Classification With Discriminative Cooperative Alignment
abstract
Cross-scene classification is one of the major challenges for hyperspectral image (HSI) classification, especially for target scenes without label samples. Most traditional domain adaptive methods learn a domain invariant subspace to reduce statistical shift while ignoring the fact that there may not exist a shared subspace when marginal distributions of source and target domains are very different. In addition, it is important for HSI classification to preserve discriminant information in the original space. To solve this issue, discriminative cooperative alignment (DCA) of subspace and distribution is proposed to cooperatively reduce the geometric and statistical shift. In the proposed framework, both geometrical and statistical alignments are considered to learn subspaces of the two domains with preserving discrimination information. Furthermore, a reconstruction constraint is imposed to enhance the robustness of subspace projection. Experimental results on three cross-scene HSI data sets demonstrate that the proposed DCA is significantly better than some state-of-the-art domain-adaptive approaches.
Yuxiang Zhang 0005, Wei Li 0032, Ran Tao 0003, Jiangtao Peng, Qian Du 0001, Zhaoquan Cai 0001
IEEE Trans. Geosci. Remote. Sens.5
2021 Characterization of Background-Anomaly Separability With Generative Adversarial Network for Hyperspectral Anomaly Detection
abstract
Hyperspectral images (HSIs) have unique advantages in distinguishing subtle spectral differences of different materials. However, due to complex and diverse backgrounds, unknown prior knowledge, and imbalanced samples, it is challenging to separate background and anomaly. In this article, we present a novel characterization of background-anomaly separability with a generative adversarial network (BASGAN) for hyperspectral anomaly detection. The key contribution is the proposal to explicitly constrain the background and anomaly separability by characterizing background spectral samples while avoiding anomaly reconstruction. First, we use a class saliency map extraction algorithm to obtain pseudobackground and anomaly samples for adversarial training. To further mitigate the suffering of anomaly contamination in background distribution estimation, we introduce background-anomaly separability constrained loss function to enhance the reconstruction of the background while weakening the anomaly reconstruction in a semisupervised way. Additionally, a discriminator is induced into the latent space to make the encoded representation resemble Gaussian distribution during adversarial training. The other is adversarial training in the reconstruction space so that the background estimation can be improved. Experiments conducted on real data sets illustrate the superior background-anomaly separability of the proposed method.
Jiaping Zhong, Weiying Xie, Yunsong Li 0001, Jie Lei 0001, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.5
2021 Superpixel-Guided Discriminative Low-Rank Representation of Hyperspectral Images for Classification
Shujun Yang, Junhui Hou, Yuheng Jia, Shaohui Mei, Qian Du 0001
IEEE Trans. Image Process.5
2020 Probability Fusion for Hyperspectral and LiDAR Data
abstract
In this paper, a new probability fusion strategy is proposed for hyperspectral and LiDAR data classification, which is inspired by the representation residual fusion strategy in our previous work. Unlike the residual fusion strategy utilizes a collaborative representation classifier, the probability fusion strategy deploys a deep residual network (DRN). This paper compares the two fusion strategies. The experiment results show that the probability fusion strategy with DRN is better than the residual fusion strategy in classification performance.
Chiru Ge, Qian Du 0001
IGARSS2
2020 Hyperspectral Image Classification Based on Tensor-Train Convolutional Long Short-Term Memory
abstract
In recent years, deep learning models have shown great advantages for hyperspectral images (HSIs) classification, in which long short-term memory (LSTM) has attracted plenty of attentions for its characteristic of modeling long-range dependencies. However, for the 2-D extended architecture of it (namely 2-D convolutional LSTM, ConvLSTM2D), it is the special gate structures of ConvLSTM2D that leads to a large number of training parameters and high requirements for device storage. To address this shortcoming, in this paper, a lightweight ConvLSTM2D cell is developed by using tensor-train decomposition (TTD) for the compression of training parameters, which is named TT-ConvLSTM2D and further applied to two state-of-the-art ConvLSTM2D-based HSI classification models for verifying its superiority. Experiments on a widely-used Indian Pines HSI data set are conducted, whose results demonstrate that the proposed TT-ConvLSTM2D cell can effectively reduce the number of the parameters and memory requirements of the whole models within a small range of accuracy degradation.
Wen-Shuai Hu, Heng-Chao Li 0001, Tian-Yu Ma, Qian Du 0001, Antonio Plaza, William J. Emery
IGARSS4
2020 Discriminative Semi-Supervised Generative Adversarial Network for Hyperspectral Anomaly Detection
abstract
Hyperspectral anomaly detection has been facing great challenges in the field of deep learning due to high dimensions and limited samples. To address these challenges, a novel discriminative semi-supervised generative adversarial network (GAN) method with dual RX (Reed-Xiaoli), called semiDRX, is proposed in this paper. The main contribution of the proposed method is to learn a reconstruction of background homogenization and anomaly saliency through a semi-supervised GAN. To achieve this goal, firstly, the coarse RX detection is performed to obtain a background sample set with potential anomalous pixels being removed. Secondly, the obtained coarse background set learns more comprehensive background characteristics through the network. The original hyperspectral image (HSI) is fed into the learned network to obtain reconstructions with homogeneous backgrounds and salient anomalies. The refined detection results are generated by a second RX detector. Experiments on three HSIs over different scenes demonstrate its advancement and effectiveness.
Tao Jiang 0031, Weiying Xie, Yunsong Li 0001, Qian Du 0001
IGARSS4
2020 L0-Motivated Low Rank Sparse Subspace Clustering for Hyperspectral Imagery
abstract
Hyperspectral image (HSI) Clustering is an unsupervised task, which segments pixels into different groups without using labeled samples. Low-rank sparse subspace clustering (LRSSC) is often applied to achieve the clustering of high-dimensional data such as HSI. The LRSSC combines low-rank recovery and sparse representation to capture both global and local structures of the data. Nuclear and Li-norm are often used to measure rank and sparsity in LRSSC since minimization of these two norms results in a convex optimization problem. However, the use of Nuclear and L1-norm can only approximate the original problem, and may lead to over-penalization. Thus, the direct solution of a Schatten-0 (So) and Lo quasi-norm regularized objective function has been proposed in the LRSSC for more accurate representation. This paper proposes to use the So/Lo-regulared LRSSC (So/Lo-LRSSC) for hyperspectral image clustering. To accommodate the large data size, an original HSI is pre-partitioned, and the Sn/Ln-LRSSC is implemented in a distributed way. Our experiments show that the performance of the So/Lo-LRSSC in hyperspectral image clustering is better than the original LRSSC and its variants based on Nuclear and L1-norm minimization.
Qian Du 0001, Ivica Kopriva
IGARSS2
2020 Correntropy-Based Sparse Spectral Clustering for Hyperspectral Band Selection
abstract
This letter presents a correntropy-based sparse spectral clustering (CSSC) method to select proper bands of a hyperspectral image. The CSSC first constructs an affinity matrix with the correntropy measure which considers the nonlinear characteristics of hyperspectral bands and can suppress effects from noise or outliers in measuring band similarity. The CSSC imposes the sparsity and block diagonal constraint on spectral clustering, which can further improve band clustering performance. Bands are finally selected from each cluster on the connected graph. Experimental results on two widely used hyperspectral images show that the CSSC behaves better than spectral clustering and other several state-of-the-art methods in band selection.
Weiwei Sun 0005, Jiangtao Peng, Gang Yang 0006, Qian Du 0001
IEEE Geosci. Remote. Sens. Lett.4
2020 Lateral-Slice Sparse Tensor Robust Principal Component Analysis for Hyperspectral Image Classification
abstract
This letter proposes a lateral-slice sparse tensor robust principal component analysis (LSSTRPCA) method to remove gross errors or outliers from hyperspectral images so as to promote the performance of subsequent classification. The LSSTRPCA assumes that a three-order hyperspectral tensor has a low-rank structure, and gross errors or outliers are sparsely scattered in a 2-D space (i.e., lateral-slice) of the tensor. It formulates a low-rank and sparse tensor decomposition problem into a convex problem and then implements the inexact augmented Lagrange multiplier method to solve it. The experiments on two hyperspectral data sets show that the LSSTRPCA can successfully remove outliers or gross errors and achieve higher accuracies than both the original robust principal component analysis (RPCA) and tensor robust principal component analysis (TRPCA).
Weiwei Sun 0005, Gang Yang 0006, Jiangtao Peng, Qian Du 0001
IEEE Geosci. Remote. Sens. Lett.4
2020 Hyperspectral Image Classification via Sparse Representation With Incremental Dictionaries
abstract
In this letter, we propose a new sparse representation (SR)-based method for hyperspectral image (HSI) classification, namely SR with incremental dictionaries (SRID). Our SRID boosts existing SR-based HSI classification methods significantly, especially when used for the task with extremely limited training samples. Specifically, by exploiting unlabeled pixels with spatial information and multiple-feature-based SR classifiers, we select and add some of them to dictionaries in an iterative manner, such that the representation abilities of the dictionaries are progressively augmented, and likewise more discriminative representations. In addition, to deal with large-scale data sets, we use a certainty sampling strategy to control the sizes of the dictionaries, such that the computational complexity is well balanced. Experiments over two benchmark data sets show that our proposed method achieves higher classification accuracy than the state-of-the-art methods, i.e., the overall classification accuracy can improve more than 4%.
Shujun Yang, Junhui Hou, Yuheng Jia, Shaohui Mei, Qian Du 0001
IEEE Geosci. Remote. Sens. Lett.5
2020 Remote sensing images super-resolution with deep convolution networks
Qiong Ran, Shizhi Zhao, Wei Li 0032, Qian Du 0001
Multim. Tools Appl.5
2020 Feature Extraction for Classification of Hyperspectral and LiDAR Data Using Patch-to-Patch CNN
abstract
Multisensor fusion is of great importance in Earth observation related applications. For instance, hyperspectral images (HSIs) provide wealthy spectral information while light detection and ranging (LiDAR) data provide elevation information, and using HSI and LiDAR data together can achieve better classification performance. In this paper, an unsupervised feature extraction framework, named as patch-to-patch convolutional neural network (PToP CNN), is proposed for collaborative classification of hyperspectral and LiDAR data. More specific, a three-tower PToP mapping is first developed to seek an accurate representation from HSI to LiDAR data, aiming at merging multiscale features between two different sources. Then, by integrating hidden layers of the designed PToP CNN, extracted features are expected to possess deeply fused characteristics. Accordingly, features from different hidden layers are concatenated into a stacked vector and fed into three fully connected layers. To verify the effectiveness of the proposed classification framework, experiments are executed on two benchmark remote sensing data sets. The experimental results demonstrate that the proposed method provides superior performance when compared with some state-of-the-art classifiers, such as two-branch CNN and context CNN.
Mengmeng Zhang 0005, Wei Li 0032, Qian Du 0001, Lianru Gao, Bing Zhang 0001
IEEE Trans. Cybern.3
2020 Patch Tensor-Based Multigraph Embedding Framework for Dimensionality Reduction of Hyperspectral Images
abstract
Graph-based dimensionality reduction (DR) techniques are of great interest in the field of image processing and especially on the analysis of hyperspectral images (HSIs). Considering the characteristics of hyperspectral data, many different types of graphs were designed to describe the structure of HSIs. Generally, the algorithms based on these graphs achieved promising performance. However, most of them only focus on how to improve the measurement of similarity between the data points by a single graph. Specifically, vector-based graph methods fail to capture the spatial information, while tensor-based graph methods assume that the pixels in each patch tensor belong to the same class, which is not exactly correct in practice. To overcome these shortcomings, this article proposes a patch tensor-based multigraph embedding (PTMGE) framework for the DR of HSIs, in which three different types of subgraphs are constructed to comprehensively describe the intrinsic geometrical structures of HSIs. First, a tensor subgraph is constructed to capture the spatial information and local geometrical structure. Second, for each two neighboring patch tensors in the tensor graph, a bipartite graph is designed to characterize the pixel-based relationships between the patch tensors. Then, considering that the diversity of pixels may be existed in each patch tensor, a pixel-based subgraph is built to describe the inner geometrical structures of every patch tensor. Finally, a novel graph fusion strategy is designed to calculate a final similarity matrix for projection learning. Experiments on three real hyperspectral data sets are conducted, and comparison with some state-of-the-art algorithms validated the effectiveness of our proposed PTMGE method.
Yangjun Deng, Heng-Chao Li 0001, Yong-Jian Sun, Xiangrong Zhang, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.6
2020 Spatial-Spectral Feature Extraction via Deep ConvLSTM Neural Networks for Hyperspectral Image Classification
abstract
In recent years, deep learning has presented a great advance in the hyperspectral image (HSI) classification. Particularly, long short-term memory (LSTM), as a special deep learning structure, has shown great ability in modeling long-term dependencies in the time dimension of video or the spectral dimension of HSIs. However, the loss of spatial information makes it quite difficult to obtain better performance. In order to address this problem, two novel deep models are proposed to extract more discriminative spatial-spectral features by exploiting the convolutional LSTM (ConvLSTM). By taking the data patch in a local sliding window as the input of each memory cell band by band, the 2-D extended architecture of LSTM is considered for building the spatial-spectral ConvLSTM 2-D neural network (SSCL2DNN) to model long-range dependencies in the spectral domain. To better preserve the intrinsic structure information of the hyperspectral data, the spatial-spectral ConvLSTM 3-D neural network (SSCL3DNN) is proposed by extending LSTM to the 3-D version for further improving the classification performance. The experiments, conducted on three commonly used HSI data sets, demonstrate that the proposed deep models have certain competitive advantages and can provide better classification performance than the other state-of-the-art approaches.
Wen-Shuai Hu, Heng-Chao Li 0001, Lei Pan 0003, Wei Li 0032, Ran Tao 0003, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.6
2020 Discriminative Reconstruction Constrained Generative Adversarial Network for Hyperspectral Anomaly Detection
abstract
The rich and distinguishable spectral information in hyperspectral images (HSIs) makes it possible to capture anomalous samples [i.e., anomaly detection (AD)] that deviate from background samples. However, hyperspectral anomaly detection (HAD) faces various challenges due to high dimensionality, redundant information, and unlabeled and limited samples. To address these problems, this article proposes an unsupervised discriminative reconstruction constrained generative adversarial network for HAD (HADGAN). Our solution is mainly based on the assumption that the number of normal samples is much larger than the number of abnormal ones. The key contribution of this article is to learn a discriminative background reconstruction with anomaly targets being suppressed, which produces the initial detection image (i.e., the residual image between the original image and reconstructed image) with anomaly targets being highlighted and background samples being suppressed. To accomplish this goal, first, by using an autoencoder (AE) network and an adversarial latent discriminator, the latent feature layer learns normal background distribution and AE learns a background reconstruction as much as possible. Second, consistency enhanced representation and shrink constraints are added to the latent feature layer to ensure that anomaly samples are projected to similar positions as normal samples in the latent feature layer. Third, using an adversarial image feature corrector in the input space can guarantee the reliability of the generated samples. Finally, an energy-based spatial and distance-based spectral joint anomaly detector is applied in the residual map to generate the final detection map. Experiments conducted on several data sets over different scenes demonstrate its state-of-the-art performance.
Tao Jiang 0031, Yunsong Li 0001, Weiying Xie, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.4
2020 Semisupervised Spectral Learning With Generative Adversarial Network for Hyperspectral Anomaly Detection
abstract
Limited by the anomalous spectral vectors in unlabeled hyperspectral images (HSIs), anomaly detection methods based on background distribution estimation often suffer from the contamination of anomalies, which decreases the estimation accuracy and, thus, weakens the detection performance. To address this problem, we proposed a novel semisupervised spectral learning (SSL) for the hyperspectral anomaly detection framework based on the generative adversarial network (GAN). GAN is applied and developed to estimate the background distribution in a semisupervised manner and obtain an initial spectral feature because of its strong representational capability and adversarial training advantage. In the proposed framework, an initial spatial feature is generated via morphological attribute filtering. Finally, an exponential constrained nonlinear suppression fusion technique is adopted to suppress the background and combine the complementary information in different features to obtain a fused detection map. The performance of the proposed anomaly detection technique is evaluated on a series of HSIs. Experimental results demonstrate that our method can outperform state-of-the-art anomaly detection methods.
Kai Jiang 0001, Weiying Xie, Yunsong Li 0001, Jie Lei 0001, Gang He 0002, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.6
2020 Hyperspectral Image Super-Resolution by Band Attention Through Adversarial Learning
abstract
Hyperspectral image (HSI) super-resolution (SR) is a challenging task due to the problems of texture blur and spectral distortion when the upscaling factor is large. To meet these two challenges, band attention through the adversarial learning method is proposed in this article. First, we put the SR process in a generative adversarial network (GAN) framework, so that the resulted high-resolution HSI can keep more texture details. Second, different from the other band-by-band SR method, the input of our method is of full bands. In order to explore the correlation of spectral bands and avoid the spectral distortion, a band attention mechanism is proposed in our generative network. A series of spatial-spectral constraints or loss functions is imposed to guide the training of our generative network so as to further alleviate spectral distortion and texture blur. The experiments on the Pavia and Cave data sets demonstrate that the proposed GAN-based SR method can yield very high-quality results, even under large upscaling factor (e.g., $8\times $ ). More importantly, it can outperform the other state-of-the-art methods by a margin which demonstrates its superiority and effectiveness.
Jiaojiao Li 0001, Ruxing Cui, Bo Li 0090, Rui Song 0003, Yunsong Li 0001, Yuchao Dai, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.7
2020 Spatial and Spectral Joint Super-Resolution Using Convolutional Neural Network
abstract
Many applications have benefited from the images with both high spatial and spectral resolution, such as mineralogy and surveillance. However, it is difficult to acquire such images due to the limitation of sensor technologies. Recently, super-resolution (SR) techniques have been proposed to improve the spatial or spectral resolution of images, e.g., improving the spatial resolution of hyperspectral images (HSIs) or improving spectral resolution of color images (reconstructing HSIs from RGB inputs). However, none of the researches attempted to improve both spatial and spectral resolution together. In this article, these two types of resolution are jointly improved using convolutional neural network (CNN). Specifically, two kinds of CNN-based SR are conducted, including a simultaneous spatial-spectral joint SR (SimSSJSR) that conducts SR in spectral and spatial domain simultaneously and a separated spatial-spectral joint SR (SepSSJSR) that considers spectral and spatial SR sequentially. In the proposed SimSSJSR, a full 3-D CNN is constructed to learn an end-to-end mapping between a low spatial-resolution mulitspectral image (LR-MSI) and the corresponding high spatial-resolution HSI (HR-HSI). In the proposed SepSSJSR, a spatial SR network and a spectral SR network are designed separately, and thus two different frameworks are proposed for SepSSJSR, namely SepSSJSR1 and SepSSJSR2, according to the order that spatial SR and spectral SR are applied. Furthermore, the least absolute deviation, instead of mean square error (MSE) in traditional SR networks, is chosen as the loss function for the proposed networks. Experimental results over simulated images from different sensors demonstrated that the proposed SepSSJSR1 is most effective to improve spatial and spectral resolution of MSIs sequentially by conducting spatial SR prior to spectral SR. In addition, validation on real Landsat images also indicates that the proposed SSJSR techniques can make full use of available MSIs for high-resolution-based analysis or applications.
Shaohui Mei, Ruituo Jiang, Xu Li 0010, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.4
2020 Ensemble Learning for Hyperspectral Image Classification Using Tangent Collaborative Representation
abstract
Recently, collaborative representation classification (CRC) has attracted much attention for hyperspectral image analysis. In particular, tangent space CRC (TCRC) has achieved excellent performance for hyperspectral image classification in a simplified tangent space. In this article, novel Bagging-based TCRC (TCRC-bagging) and Boosting-based TCRC (TCRC-boosting) methods are proposed. The main idea of TCRC-bagging is to generate diverse TCRC classification results using the bootstrap sample method, which can enhance the accuracy and diversity of a single classifier simultaneously. For TCRC-boosting, it can provide the most informative training samples by changing their distributions dynamically for each base TCRC learner. The effectiveness of the proposed methods is validated using three real hyperspectral data sets. The experimental results show that both TCRC-bagging and TCRC-boosting outperform their single classifier counterpart. In particular, the TCRC-boosting provides superior performance compared with the TCRC-bagging.
Hongjun Su, Qian Du 0001, Peijun Du
IEEE Trans. Geosci. Remote. Sens.3
2020 Fast and Latent Low-Rank Subspace Clustering for Hyperspectral Band Selection
abstract
This article presents a fast and latent low-rank subspace clustering (FLLRSC) method to select hyperspectral bands. The FLLRSC assumes that all the bands are sampled from a union of latent low-rank independent subspaces and formulates the self-representation property of all bands into a latent low-rank representation (LLRR) model. The assumption ensures sufficient sampling bands in representing low-rank subspaces of all bands and improves robustness to noise. The FLLRSC first implements the Hadamard random projections to reduce spatial dimensionality and lower the computational cost. It then adopts the inexact augmented Lagrange multiplier algorithm to optimize the LLRR program and estimates sparse coefficients of all the projected bands. After that, it employs a correntropy metric to measure the similarity between pairwise bands and constructs an affinity matrix based on sparse representation. The correntropy metric could better describe the nonlinear characteristics of hyperspectral bands and enhance the block-diagonal structure of the similarity matrix for correctly clustering all subspaces. The FLLRSC conducts spectral clustering on the connected graph denoted by the affinity matrix. The bands that are closest to their separate cluster centroids form the final band subset. Experimental results on three widely used hyperspectral data sets show that the FLLRSC performs better than the classical low-rank representation methods with higher classification accuracy at a low computational cost.
Weiwei Sun 0005, Jiangtao Peng, Gang Yang 0006, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.4
2020 CVA2E: A Conditional Variational Autoencoder With an Adversarial Training Process for Hyperspectral Imagery Classification
abstract
Deep generative models such as the generative adversarial network (GAN) and the variational autoencoder (VAE) have obtained increasing attention in a wide variety of applications. Nevertheless, the existing methods cannot fully consider the inherent features of the spectral information, which leads to the applications being of low practical performance. In this article, in order to better handle this problem, a novel generative model named the conditional variational autoencoder with an adversarial training process (CVA2E) is proposed for hyperspectral imagery classification by combining variational inference and an adversarial training process in the spectral sample generation. Moreover, two penalty terms are added to promote the diversity and optimize the spectral shape features of the generated samples. The performance on three different real hyperspectral data sets confirms the superiority of the proposed method.
Xue Wang 0008, Kun Tan 0001, Qian Du 0001, Yu Chen 0014, Peijun Du
IEEE Trans. Geosci. Remote. Sens.3
2020 Autoencoder and Adversarial-Learning-Based Semisupervised Background Estimation for Hyperspectral Anomaly Detection
abstract
Reliable detection of anomalies without any prior information is a critical yet challenging task in many applications, not least military and civilian fields. An intelligent anomaly detection system would use the material-specific spectral information in hyperspectral images (HSIs), thereby avoiding the loss of visually confusing objects. However, conventional hyperspectral anomaly detection methods are mainly achieved in an unsupervised way leading to limited performance due to lack of prior knowledge. In this article, we propose a novel autoencoder and adversarial-learning based semisupervised background estimation model (SBEM) that is trained only on the background spectral samples in order to accurately learn the background distribution. In particular, an unsupervised background searching method is firstly conducted on the original HSIs to search the background spectral samples. Our proposed SBEM consists of an encoder, a decoder, and a discriminator to thoroughly capture background distribution. Furthermore, jointly minimizing the reconstruction loss, spectral loss, and adversarial loss during training aids the model to learn the background distribution as required. Experiments on four real HSIs demonstrate that compared to the current state-of-the-art, the proposed framework yields higher detection capability and lower false alarm rate, which shows that it has a significant benefit in the tradeoff between detection accuracy and false alarm rate.
Weiying Xie, Baozhu Liu, Yunsong Li 0001, Jie Lei 0001, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.5
2020 Deep Latent Spectral Representation Learning-Based Hyperspectral Band Selection for Target Detection
abstract
Hyperspectral images (HSIs) can provide discriminative spectral signatures regarding the physical nature of different materials. It is this unique nature that makes HSIs to be of great interest in many fields. However, HSI application faces various challenges due to high dimensionality, redundant information, noisy bands, and insufficient samples. To address these problems, we propose an unsupervised band selection method based on deep latent spectral representation learning, called DLSRL, in this article. It imposes spectral consistency on deep latent space that resolves the issue of insufficient samples and spectral information lost in HSI interpretation. It pursues the low-dimensional optimal representation of the high-dimensional HSIs. In particular, an adaptive mapping relationship is constructed between the deep latent representation and the optimal subset to preserve physical significance optimally. Furthermore, a hierarchical optimization approach is introduced to achieve target detection with the selected subset. To verify the superiority of the proposed method, experiments have been conducted on four data sets captured by different sensors over different scenes. Comparative analyses validate that the proposed method presents superior performance in terms of high detection accuracy and low false alarm rate.
Weiying Xie, Jie Lei 0001, Yunsong Li 0001, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.5
2020 SRUN: Spectral Regularized Unsupervised Networks for Hyperspectral Target Detection
abstract
The high dimensionality of a hyperspectral image (HSI) provides the possibility of deeply capturing the underlying and intrinsic characteristics in spectra, such that targets embedded in the background can be detected. However, redundant information, deteriorated bands, and other interferences from background challenge the target detection problem. In this article, an effective feature extraction method based on unsupervised networks is proposed to mine intrinsic properties underlying HSIs. Our approach, called spectral regularized unsupervised networks (SRUN), imposes spectral regularization on autoencoder (AE) and variational AE (VAE) to emphasize spectral consistency, which is more suitable for characterizing spectral information of HSIs by hidden nodes than the original AE and VAE models. Then, we conduct a simple feature selection algorithm on the hidden nodes in the deepest code to select specific nodes that contain distinguishability between target and background, which is based on the spectral angular difference between a known target spectrum and spectra of other pixels in input. The selected nodes are further weighted adaptively to obtain a discriminative map depending on the observation that each selected node provides different contribution rates to target detection. Experimental results on several data sets illustrate that the proposed SRUN-based target detection algorithm is suitable for targets at the subpixel level and those with structural information.
Weiying Xie, Jie Lei 0001, Yunsong Li 0001, Qian Du 0001, Gang He 0002
IEEE Trans. Geosci. Remote. Sens.5
2020 Discriminative Marginalized Least-Squares Regression for Hyperspectral Image Classification
abstract
Least-squares regression (LSR)-based classifiers are effective in multiclassification tasks. However, most existing methods use limited projections, resulting in loss of much discriminant information; furthermore, they focus only on exactly fitting samples to target matrix while ignoring overfitting issue. To solve these drawbacks, discriminative marginalized LSR (DMLSR) is proposed to learn a more discriminative projection matrix with consideration of class separability and data-reconstruction ability simultaneously. In the proposed framework, an intraclass compactness graph is employed to avoid the overfitting problem and enhance class separability, and a data-reconstruction constraint is imposed to preserve discriminant information on limited projections. Experimental results on several hyperspectral data sets demonstrate that the proposed method significantly outperforms some state-of-the-art classifiers.
Yuxiang Zhang 0005, Wei Li 0032, Heng-Chao Li 0001, Ran Tao 0003, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.5
2020 Joint Classification of Hyperspectral and LiDAR Data Using Hierarchical Random Walk and Deep CNN Architecture
abstract
Earth observation using multisensor data is drawing increasing attention. Fusing remotely sensed hyperspectral imagery and light detection and ranging (LiDAR) data helps to increase application performance. In this article, joint classification of hyperspectral imagery and LiDAR data is investigated using an effective hierarchical random walk network (HRWN). In the proposed HRWN, a dual-tunnel convolutional neural network (CNN) architecture is first developed to capture spectral and spatial features. A pixelwise affinity branch is proposed to capture the relationships between classes with different elevation information from LiDAR data and confirm the spatial contrast of classification. Then in the designed hierarchical random walk layer, the predicted distribution of dual-tunnel CNN serves as global prior while pixelwise affinity reflects the local similarity of pixel pairs, which enforce spatial consistency in the deeper layers of networks. Finally, a classification map is obtained by calculating the probability distribution. Experimental results validated with three real multisensor remote sensing data demonstrate that the proposed HRWN significantly outperforms other state-of-the-art methods. For example, the two branches CNN classifier achieves an accuracy of 88.91% on the University of Houston campus data set, while the proposed HRWN classifier obtains an accuracy of 93.61%, resulting in an improvement of approximately 5%.
Xudong Zhao 0003, Ran Tao 0003, Wei Li 0032, Heng-Chao Li 0001, Qian Du 0001, Wenzi Liao, Wilfried Philips
IEEE Trans. Geosci. Remote. Sens.5
2020 Hyperspectral Pansharpening With Deep Priors
abstract
Hyperspectral (HS) image can describe subtle differences in the spectral signatures of materials, but it has low spatial resolution limited by the existing technical and budget constraints. In this paper, we propose a promising HS pansharpening method with deep priors (HPDP) to fuse a low-resolution (LR) HS image with a high-resolution (HR) panchromatic (PAN) image. Different from the existing methods, we redefine the spectral response function (SRF) based on the larger eigenvalue of structure tensor (ST) matrix for the first time that is more in line with the characteristics of HS imaging. Then, we introduce HFNet to capture deep residual mapping of high frequency across the upsampled HS image and the PAN image in a band-by-band manner. Specifically, the learned residual mapping of high frequency is injected into the structural transformed HS images, which are the extracted deep priors served as additional constraint in a Sylvester equation to estimate the final HR HS image. Comparative analyses validate that the proposed HPDP method presents the superior pansharpening performance by ensuring higher quality both in spatial and spectral domains for all types of data sets. In addition, the HFNet is trained in the high-frequency domain based on multispectral (MS) images, which overcomes the sensitivity of deep neural network (DNN) to data sets acquired by different sensors and the difficulty of insufficient training samples for HS pansharpening.
Weiying Xie, Jie Lei 0001, Yuhang Cui, Yunsong Li 0001, Qian Du 0001
IEEE Trans. Neural Networks Learn. Syst.5
2019 Collaborative Classification of Hyperspectral and Lidar Data With Information Fusion and Deep Nets
abstract
Convolutional neural network (CNN) receives extensive attention in hyperspectral image classification. While hyper-spectral images contain abundant spectral information but lack spatial information, which usually contributes to poor classification results. In this paper, a novel classification framework called information fusion based CNN (IF-CNN) is proposed to compensate for the shortcomings of hyper-spectral images. The proposed method merges hyperspectral images with abundant spectral information and LiDAR images with rich spatial information as the input of classification framework. Furthermore, the framework consists of two convolutional neural networks: one-dimensional CNN for extracting spectral features, and two-dimensional CNN for extracting spatial correlation features. Experimental results demonstrate that the proposed method achieves excellent performance compared with some existing methods.
Chen Chen 0001, Xudong Zhao 0003, Wei Li 0032, Ran Tao 0003, Qian Du 0001
IGARSS5
2019 Fast Kernel Collaborative Representation for Hyperspectral Image Classification
abstract
This paper proposes efficient kernel collaborative representation-based classifiers for hyperspectral image classification. The original collaborative classifier is very fast; however, its classification performance is limited for hyperspectral image classification. Although the traditional kernel methods have been widely applied to improve the classification performance, they usually suffer from high computational cost. We propose to adopt explicit kernel mapping to reduce the time complexity for kernel collaborative representation-based classifiers. The experimental results show our approach can remain high classification accuracy with low computational cost.
Yan Xu 0003, Qian Du 0001, Nicolas H. Younan
IGARSS2
2019 Multiple-Feature Ideal Regularized Kernel for Hyperspectral Image Classification
abstract
This paper proposes multiple-feature ideal regularized kernel for hyperspectral image classification, which offers the advantage of combining the complementary discriminative information among multiple features with the ideal regularized kernel. Four different types of features including spectral feature, shape feature (i.e., extended multiattribute profiles), local feature (i.e., local binary pattern), and global feature (i.e., Gabor feature) are investigated in this paper. Furthermore, a majority votingbased ensemble method combining different features is adopted to further increase classification performance. Experimental results demonstrate that the proposed method can provide superior performance than the state-of-the-art classifiers.
Yan Xu 0003, Jiangtao Peng, Qian Du 0001, Nicolas H. Younan
IGARSS3
2019 Hierarchical Deep Feature Representation for High-Resolution Scene Classification
abstract
High-resolution scene classification is a fundamental yet challenging problem due to rich image variations in viewpoint, object pose and spatial resolution, etc, which results in large within-class diversity and high between-class similarity. In the paper we focus on tackling the problem of how to learn appropriate feature representation for high-resolution scene classification. To achieve better scene representation, we proposed a combined CNN feature learning framework in multi-scale multi-layer based Gaussian coding (mSmL-Gcoding) manner. In addition, a novel feature coding with Gaussian descriptor is introduced to enhance the discriminative ability of CNN features. Experimental results on two publicly available challenging scene datasets validated that the effectiveness of our method and found it compared favorably with state-of-the-arts.
Xiaoyong Bian, Chunfang Chen, Chunhua Deng, Ruiyao Liu, Qian Du 0001
IGARSS5
2019 Spatial Constrained Hyperspectral Reconstruction from RGB Inputs Using Dictionary Representation
abstract
Reconstructing hyperspectral images from RGB inputs has gained great attention recently. In dictionary representation-based hyperspectral image reconstruction, dictionary representation is first carried out in RGB space and then dictionary reconstruction is conducted in hyperspectral space for per-pixel reconstruction. However, such work mainly focuses on spectral mapping from RGB space to hyperspectral space, ignoring physical distribution of objects in the image. In this paper, spatial context of pixels is used to improve the reconstruction performance. Specially, neighboring pixels are used to constrain the dictionary representation problem in RGB space, and the Simultaneous Orthogonal Matching Pursuit (SOMP) is used to improve the performance of hyperspectral reconstruction. Experimental results on two benchmark data sets demonstrate the superiority of the proposed technique.
Yunhao Geng, Shaohui Mei, Yifan Zhang 0006, Qian Du 0001
IGARSS5
2019 Dual 1D-2D Spatial-Spectral CNN for Hyperspectral Image Super-Resolution
abstract
Hyperspectral image (HSI) spatial super-resolution(SR) is a challenging task. Compared with a RGB images, the mapping between the low-high HSI pairs is more difficult since much more spectral bands are involved. In this paper, a novel dual 1D-2D spatial-spectral convolutional neural network (CNN) architecture is proposed for spatial SR of HSIs. Specifically, by differential treatment over redundancy in spectral and spatial domains of an HSI, the spectral and spatial context are first separately explored by 1D and 2D convolution. These two kinds of feature information are then fused using a novel hierarchical side connection, which impose the spectral information to the spatial path gradually. Experimental results over benchmark Pavia data set demonstrate that the proposed architecture clearly outperform state-of-the-art 3D CNN based works in terms of both visual quality and quantitative assessment.
Jiaojiao Li 0001, Ruxing Cui, Bo Li 0090, Yunsong Li 0001, Shaohui Mei, Qian Du 0001
IGARSS6
2019 Decs-Net: Convolutional Self-Encoding Network for Hyperspectral Image Denoising
abstract
Noises in hyperspectral image (HSI) degrades both spatial and spectral features of ground objects, and greately defects the following processing, such as classification, target detection and recognition. In this paper, a convolutional self-encoding network (DeCS-Net) is designed for HSI denoising, which integrates the superiority of convolutional neural network (CNN) and auto-encoder (AE) to learn multi-scale features. The noise in the observed HSI is estimated by residual learning strategy, and is removed from the observed HSI to obtain an estimation of the ideal HSI without noise. Experimental results on benchmark HSI data set illustrate that the proposed DeCS-Net is effective for HSI denoising and outperforms the state-of-the-art CNN based HSI denoising methods.
Shaohui Mei, Zhi Zhang 0023, Yifan Zhang 0006, Jingyu Ji, Qian Du 0001
IGARSS6
2019 Feature-Level Fusion of Landsat-8 OLI-SWIR and TIR Images for Fine Burned Area Change Detection
abstract
This paper proposes a novel feature-level fusion approach for burned area change detection at a fine level. The proposed approach relies on two features. The first feature is a modified normalized burn ratio (MNBR) fire index based on Landsat-8 OLI SWIR data, and the second feature is the Bright temperature (BT) based on Landsat-8 TIR data. Then two features are combined by using the gradient transfer fusion algorithm and a change detection technique to generate a fine burned area change map. A real Landsat-8 data set covering a complex fire disaster scenario is utilized to test the performance of the proposed approach. Experimental results demonstrate the effectiveness of the proposed feature-level fusion approach comparing with the reference methods in term of higher separability value and detection accuracy.
Sicong Liu 0001, Michele Dalponte, Xiaohua Tong, Qian Du 0001
IGARSS5
2019 Hyperspectral and Panchromatic Image Fusion Based on Weighted Tensor Matrix
abstract
In this paper, a new hyperspectral image (HSI) and panchromatic image (PANI) fusion approach via weighted tensor matrix is proposed. In the proposed method, homomorphic filtering is use for obtaining spatial component of HSI, and a weighted root mean squared error (RMSE)-based algorithm is proposed to extract the total intensity details of HSI. In addition, an optimized weighted tensor matrix-based method is proposed to acquire the integrated intensity details from both HSI and PANI. Comparative analyses show the proposed approach performs better than other excellent approaches in visual inspection and objective assessment.
Jiahui Qu, Qian Du 0001, Yunsong Li 0001, Wenqian Dong
IGARSS2
2019 Attention-based Domain Adaptation for Hyperspectral Image Classification
abstract
Machine learning algorithms have been extensively used to generate complex features for classification task in the hyperspectral images. However, for challenging cases like domain adaptation (DA), these algorithms tend to perform less efficiently. Recently, with the advent of deep learning algorithms, more complex but useful features can be generated for hyperspectral image classification task. However, attention-based feature generation is not explored till now, which has been found to be effective for distinguishing different classes of images than without transferring the parameters. In this paper, we have opted to use attention-based DA based on transferring different levels of attention from a supervisor network to the student network to provide useful but more complex features for improving the overall classification of the DA problem. It has been shown that the proposed attention-based transfer method outperforms the state-of-the-art domain adaptation methods.
Robiul Hossain Md. Rafi, Bo Tang 0011, Qian Du 0001, Nicolas H. Younan
IGARSS3
2019 Orthogonal Graph-regularized Non-negative Matrix Factorization for Hyperspectral Image Clustering
abstract
As an unsupervised task, hyperspectral image (HSI) clustering separates pixels into different groups. In this paper, an orthogonal graph-regularized non-negative matrix factorization (OGNMF) algorithm is proposed for HSI clustering. Because of large size of HSI, the HSI clustering problem is computational consuming. On the other hand, non-negative matrix factorization (NMF), which is a popular multivariate analysis method, has related to light computation budget. Furthermore, the orthogonal and graph regularization, which can improve the clustering performance and capture the local structure features, are applied to NMF for better performance in HSI. Moreover, both spatial and spectral information are extracted and utilized. Our experiments show that the performance of the proposed OGNMF is better than the NMF and ONMF algorithms in HSI clustering.
Qian Du 0001, Ivica Kopriva, Nicolas H. Younan
IGARSS2
2019 Discriminative CNN Via Metric Learning for Hyperspectral Classification
abstract
Convolutional neural networks (CNNs) have been demonstrated to be capable of learning effective spatial-spectral features for hyperspectral classification. However, traditional CNNs are mainly trained using classification errors in decision domain. In this paper, a metric learning based training strategy is proposed to further enhance feature separability by training CNNs in feature domain as well as decision domain. Specifically, a metric learning loss function is designed to train CNNs in the second last fully connected feature layer, instead of the last fully connected decision layer. As a result, both within-class feature similarity and between-class feature separability can be enhanced even with a small amount of training samples. Experimental results over two benchmark hyperspectral data sets demonstrate that the proposed metric learning strategy is very effective to explore more discriminative features and its performance obviously outperforms several state-of-art CNNs for classification of hyperspectral images.
Zhongqi Tian, Zhi Zhang 0023, Shaohui Mei, Ruoqiao Jiang, Shuai Wan, Qian Du 0001
IGARSS6
2019 Low-Rank and Collaborative Representation for Hyperspectral Anomaly Detection
abstract
Recently, low-rank representation and collaborative representation for hyperspectral anomaly detection are widely studied. In this paper, a novel anomaly detector which combines low-rank and collaborative representations for hyperspectral anomaly detection (LRCRD) is proposed. Different from existing anomaly detection methods using low-rank and collaborative representation, the proposed method divides an image into two parts: background and anomaly targets. A background dictionary is used to represent the background whose coefficient matrix is constrained by low-rank and l2norm minimization. The sparsely distributed anomalies are determined by the residual matrix which is constrained by l2,1norm minimization. Considering different similarities between a testing pixel and a dictionary atom, a distance-weighted matrix is adopted. Moreover, construction of the background dictionary avoids the pollution of abnormal pixels and makes the detection result more stable. Experimental results show that the LRCRD performs better than state-of-the-art anomaly detection methods.
Zhaoyue Wu, Hongjun Su, Qian Du 0001
IGARSS3
2019 Local Sparse Representation Based Spatial Preprocessing For Endmember Extraction
abstract
Hyperspectral unmixing has been widely used to decompose a mixed pixel into a collection of endmembers weighted by their corresponding fractional abundances, in which endmember extraction step is of crucial importance. Many classical endmember extraction algorithms mainly identify spectrally pure endmembers according to spectra of pixels, e.g., NFINDR and vertex component analysis (VCA), ignoring spatial distribution or structure information that has been demonstrated to be complemental for spectral information in hyperspectral image processing. In order to improve the performance of these classical endmember extraction algorithms, a novel spatial preprocessing method is proposed to explore spatial information prior to endmember extraction step. Specifically, pixels in hyperspectral images are modified using their sparse linear approximation by neighboring pixels, such that spectral variation within a local spatial neighbor-hood can be alleviated. Experimental results on both simulated and real data sets demonstrate that the proposed local sparse representation based spatial preprocessing algorithm is capable of producing better unmixing result compared to several state-of-the-art spatial preprocessing methods.
Ge Zhang 0006, Shaohui Mei, Yan Feng 0005, Qian Du 0001
IGARSS5
2019 Data Augmentation for Hyperspectral Image Classification With Deep CNN
abstract
Convolutional neural network (CNN) has been widely used in hyperspectral imagery (HSI) classification. Data augmentation is proven to be quite effective when training data size is relatively small. In this letter, extensive comparison experiments are conducted with common data augmentation methods, which draw an observation that common methods can produce a limited and up-bounded performance. To address this problem, a new data augmentation method, named as pixel-block pair (PBP), is proposed to greatly increase the number of training samples. The proposed method takes advantage of deep CNN to extract PBP features, and decision fusion is utilized for final label assignment. Experimental results demonstrate that the proposed method can outperform the existing ones.
Wei Li 0032, Chen Chen 0001, Mengmeng Zhang 0005, Heng-Chao Li 0001, Qian Du 0001
IEEE Geosci. Remote. Sens. Lett.5
2019 Unsupervised Manifold Alignment for Cross-Domain Classification of Remote Sensing Images
abstract
The original manifold alignment (MA) approach is for semisupervised domain adaptation. Since the target prior information is difficult to obtain, we conduct it in an unsupervised manner, resulting in an unsupervised MA (UMA) method. This approach utilizes the probabilistic prediction results of target data to construct the cross-domain similarity matrix, which characterizes the relationships between domains and is used for alignment. Due to the spectral drift, the prediction results may not be accurate, and thus affect the alignment. We employed spatial filtering and overall centroid alignment method as two preprocessing strategies to improve the prediction results. Furthermore, per-class maximum mean discrepancy (MMD) constraint is introduced to the UMA to further improve the alignment performance. The proposed UMA_MMD algorithm is applied for the classification of remote sensing images, and the experimental results using hyperion multitemporal remote sensing images demonstrated the effectiveness of the proposed approach.
Li Ma 0005, Chuang Luo, Jiangtao Peng, Qian Du 0001
IEEE Geosci. Remote. Sens. Lett.4
2019 Discriminative Transfer Joint Matching for Domain Adaptation in Hyperspectral Image Classification
abstract
Domain adaptation, which aims at learning an accurate classifier for a new domain (target domain) using labeled information from an old domain (source domain), has shown promising value in remote sensing fields yet still been a challenging problem. In this letter, we focus on knowledge transfer between hyperspectral remotely sensed images in the context of land-cover classification under unsupervised setting where labeled samples are available only for the source image. Specifically, a discriminative transfer joint matching (DTJM) method is proposed, which matches source and target features in the kernel principal component analysis space by minimizing the empirical maximum mean discrepancy, performs instance reweighting by imposing an ℓ2,1-norm on the embedding matrix, and preserves the local manifold structure of data from different domains and meanwhile maximizes the dependence between the embedding and labels. The proposed approach is compared with some state-of-the-art feature extraction techniques with and without using label information of source data. Experimental results on two benchmark hypersepctral data sets show the effectiveness of the proposed DTJM.
Jiangtao Peng, Weiwei Sun 0005, Li Ma 0005, Qian Du 0001
IEEE Geosci. Remote. Sens. Lett.4
2019 Efficient Probabilistic Collaborative Representation-Based Classifier for Hyperspectral Image Classification
abstract
This letter presents an efficient probabilistic collaborative representation-based classifier (PROCRC) for hyperspectral image classification. Its performance is evaluated on different types of spatial features of hyperspectral imagery (HSI) including shape feature (i.e., extended multiattribute feature), global feature (i.e., Gabor feature), and local feature [i.e., local binary pattern (LBP)]. Compared with the original collaborative representation classifier (CRC), the proposed PROCRC offers superior classification performance. The Tikhonov regularized versions of CRC have excellent classification performance but their computational cost is high. The experimental results show that the PROCRC can yield comparable classification accuracy but with much lower computational cost.
Yan Xu 0003, Qian Du 0001, Wei Li 0032, Nicolas H. Younan
IEEE Geosci. Remote. Sens. Lett.2
2019 GETNET: A General End-to-End 2-D CNN Framework for Hyperspectral Image Change Detection
abstract
Change detection (CD) is an important application of remote sensing, which provides timely change information about large-scale Earth surface. With the emergence of hyperspectral imagery, CD technology has been greatly promoted, as hyperspectral data with high spectral resolution are capable of detecting finer changes than using the traditional multispectral imagery. Nevertheless, the high dimension of the hyperspectral data makes it difficult to implement traditional CD algorithms. Besides, endmember abundance information at subpixel level is often not fully utilized. In order to better handle high-dimension problem and explore abundance information, this paper presents a general end-to-end 2-D convolutional neural network (CNN) framework for hyperspectral image CD (HSI-CD). The main contributions of this paper are threefold: 1) mixed-affinity matrix that integrates subpixel representation is introduced to mine more cross-channel gradient features and fuse multisource information; 2) 2-D CNN is designed to learn the discriminative features effectively from the multisource data at a higher level and enhance the generalization ability of the proposed CD algorithm; and 3) the new HSI-CD data set is designed for objective comparison of different methods. Experimental results on real hyperspectral data sets demonstrate that the proposed method outperforms most of the state of the arts.
Qi Wang 0009, Zhenghang Yuan, Qian Du 0001, Xuelong Li 0001
IEEE Trans. Geosci. Remote. Sens.3
2019 Local Spectral Similarity Preserving Regularized Robust Sparse Hyperspectral Unmixing
abstract
Spatial context has been demonstrated to be effective to constrain sparse unmixing (SU) of hyperspectral images. However, the existing algorithms employed simple spatial information without keeping spectral fidelity. By considering the fact that adjacent pixels own not only the endmembers with same variations but also approximated fractional abundances, in this paper, local spectral similarity preserving (LSSP) constraint is proposed to preserve spectral similarity in a local area during robust sparse unmixing (RSU). Specially, four LSSP constraints are constructed using different-norm-constrained pixel-level difference over abundance-level difference in a local area. Moreover, a convex optimization algorithm is proposed to solve the proposed LSSP-constrained RSU (LSSP-RSU). Experimental results on both synthetic and real hyperspectral data demonstrate that the developed algorithms yield better values of the signal-toreconstruction error (SRE). Especially, when using l2norm of pixel-level difference to weight the l1norm of abundance-level difference, the proposed LSSP-RSU algorithm can achieve superior unmixing performance.
Jiaojiao Li 0001, Yunsong Li 0001, Rui Song 0003, Shaohui Mei, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.5
2019 DDLPS: Detail-Based Deep Laplacian Pansharpening for Hyperspectral Imagery
abstract
In this paper, we propose a new pansharpening method called detail-based deep Laplacian pansharpening (DDLPS) to improve the spatial resolution of hyperspectral imagery. This method includes three main components: upsampling, detail injection, and optimization. In particular, a deep Laplacian pyramid super-resolution network (LapSRN) improves the resolution of each band. Then, a guided image filter and a gain matrix are used to combine the spatial and spectral details with an optimization problem, which is formed to adaptively select an injection coefficient. The DDLPS method is compared with 11 state-of-the-art or traditional pansharpening approaches. The experimental results demonstrate the superiority of the DDLPS method in terms of both quantitative indices and visual appearance. In addition, the training of LapSRN is based on the data sets of traditional RGB images, which overcomes the practical difficulty of insufficient training samples for pansharpening.
Kaiyan Li 0002, Weiying Xie, Qian Du 0001, Yunsong Li 0001
IEEE Trans. Geosci. Remote. Sens.3
2019 Structure-Aware Collaborative Representation for Hyperspectral Image Classification
abstract
Recently, collaborative representation (CR) has drawn increasing attention in hyperspectral image classification due to its simplicity and effectiveness. However, existing representation-based classifiers do not explicitly utilize class label information of training samples in estimating representation coefficients. To solve this issue, a structure-aware CR with Tikhonov regularization (SaCRT) method is proposed to consider both class label information of training samples and spectral signatures of testing pixels to estimate more discriminative representation coefficients. In the proposed framework, marginal regression is employed; furthermore, an interclass row-sparsity structure is designed to preserve the compact relationship among intraclass pixels and more separable interclass pixels, thereby enhancing class separability. The experimental results evaluated using three hyperspectral data sets demonstrate that the proposed method significantly outperforms some state-of-the-art classifiers.
Wei Li 0032, Yuxiang Zhang 0005, Na Liu 0014, Qian Du 0001, Ran Tao 0003
IEEE Trans. Geosci. Remote. Sens.4
2019 Hyperspectral Image Restoration Based on Low-Rank Recovery With a Local Neighborhood Weighted Spectral-Spatial Total Variation Model
abstract
Hyperspectral image (HSI) is often contaminated by mixed noise, which severely affects the visual quality and subsequent applications of the data. In this paper, HSI restoration based on low-rank recovery with a local neighborhood weighted spectral-spatial total variation (TV) model is proposed, which focuses on the preservation of spatial structure and spectral fidelity. The low-rank matrix model is adopted to exploit the spectral and spatial correlation information, and the l1-norm is used as a prior to remove the sparse noise. Furthermore, a local spatial neighborhood weighted spectral-spatial TV is utilized to jointly model the spectral-spatial prior information; specifically, the spectral and spatial differences are both considered in the TV term, and the weight is computed by considering the local neighborhood information in the spatial domain. Alternating direction method of multipliers optimization procedure is extended to solve the presented model. Experimental results demonstrate that the proposed method can remove the mixed noise, enhance the structural information simultaneously, and offer the best performance compared with several state-of-the-art HSI restoration methods.
Hongyi Liu 0001, Peipei Sun, Qian Du 0001, Zebin Wu 0001, Zhihui Wei
IEEE Trans. Geosci. Remote. Sens.3
2019 Unsupervised Spatial-Spectral Feature Learning by 3D Convolutional Autoencoder for Hyperspectral Classification
abstract
Feature learning technologies using convolutional neural networks (CNNs) have shown superior performance over traditional hand-crafted feature extraction algorithms. However, a large number of labeled samples are generally required for CNN to learn effective features under classification task, which are hard to be obtained for hyperspectral remote sensing images. Therefore, in this paper, an unsupervised spatial-spectral feature learning strategy is proposed for hyperspectral images using 3-Dimensional (3D) convolutional autoencoder (3D-CAE). The proposed 3D-CAE consists of 3D or elementwise operations only, such as 3D convolution, 3D pooling, and 3D batch normalization, to maximally explore spatial-spectral structure information for feature extraction. A companion 3D convolutional decoder network is also designed to reconstruct the input patterns to the proposed 3D-CAE, by which all the parameters involved in the network can be trained without labeled training samples. As a result, effective features are learned in an unsupervised mode that label information of pixels is not required. Experimental results on several benchmark hyperspectral data sets have demonstrated that our proposed 3D-CAE is very effective in extracting spatial-spectral features and outperforms not only traditional unsupervised feature extraction algorithms but also many supervised feature extraction algorithms in classification application.
Shaohui Mei, Jingyu Ji, Yunhao Geng, Zhi Zhang 0023, Xu Li 0010, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.6
2019 Self-Paced Joint Sparse Representation for the Classification of Hyperspectral Images
abstract
In this paper, a self-paced joint sparse representation (SPJSR) model is proposed for the classification of hyperspectral images (HSIs). It replaces the least-squares (LS) loss in the standard joint sparse representation (JSR) model with a weighted LS loss and adopts a self-paced learning (SPL) strategy to learn the weights for neighboring pixels. Rather than predefining a weight vector in the existing weighted JSR methods, both the weight and sparse representation (SR) coefficient associated with neighboring pixels are optimized by an alternating iterative strategy. According to the nature of SPL, in each iteration, neighboring pixels with nonzero weights (i.e., easy pixels) are included for the joint SR of a testing pixel. With the increase of iterations, the model size (i.e., the number of selected neighboring pixels) is enlarged and more neighboring pixels from easy to complex are gradually added into the JSR learning process. After several iterations, the algorithm can be terminated to produce a desirable model that includes easy homogeneous pixels and excludes complex inhomogeneous pixels. Experimental results on two benchmark hyperspectral data sets demonstrate that our proposed SPJSR is more accurate and robust than existing JSR methods, especially in the case of heavy noise.
Jiangtao Peng, Weiwei Sun 0005, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.3
2019 Kernel Collaborative Representation With Local Correlation Features for Hyperspectral Image Classification
abstract
Spatial information has widely been used in hyperspectral image (HSI) classification to improve classification accuracy. However, the structural information may not be fully explored when using spatial information, this paper proposes the joint collaborative representation classification with correlation matrix (CRC-CM) for HSI by using spatial correlation features in patches, which could keep the local intrinsic structure in band images. Considering spatial heterogeneity in a patch, local correlation matrices of a target neighborhood patch and training neighborhood patch are improved by a binary weight matrix and shape-adaptive neighborhood. To explore nonlinear nature of spatial features, corresponding kernel CRC-CM is also proposed. To evaluate the effectiveness of the proposed methods, three real HSIs with different degree of heterogeneity are used. The experimental results show that the proposed spatial correlation features outperform the original spectral feature and other spatial features which widely used in HSI classifiers.
Hongjun Su, Qian Du 0001, Peijun Du
IEEE Trans. Geosci. Remote. Sens.3
2019 Caps-TripleGAN: GAN-Assisted CapsNet for Hyperspectral Image Classification
abstract
The increase in the spectral and spatial information of hyperspectral imagery poses challenges in classification due to the fact that spectral bands are highly correlated, training samples may be limited, and high resolution may increase intraclass difference and interclass similarity. In this paper, in order to better handle these problems, a Caps-TripleGAN framework is proposed by exploring the 1-D structure triple generative adversarial network (TripleGAN) for sample generation and integrating CapsNet for hyperspectral image classification. Moreover, spatial information is utilized to verify the learning capacity and discriminative ability of the Caps-TripleGAN framework. The experimental results obtained with three real hyperspectral data sets confirm that the proposed method outperforms most of the state-of-the-art methods.
Xue Wang 0008, Kun Tan 0001, Qian Du 0001, Yu Chen 0014, Peijun Du
IEEE Trans. Geosci. Remote. Sens.3
2018 CascadeNet: Modified ResNet with Cascade Blocks
abstract
Different enhanced convolutional neural network (CNN) architectures have been proposed to surpass very deep layer bottleneck by using shortcut connections. In this paper, we present an effective deep CNN architecture modified on the typical Residual Network (ResNet), named as Cascade Network (CascadeNet), by repeating cascade building blocks. Each cascade block contains independent convolution paths to pass information in the previous layer and the middle one. This strategy exposes a concept of “cross-passing” which differs from the ResNet that stacks simple building blocks with residual connections. Traditional residual building block do not fully utilizes the middle layer information, but the designed cascade block catches cross-passing information for more complete features. There are several characteristics with CascadeNet: enhance feature propagation and reuse feature after each layer instead of each block. In order to verify the performance in CascadeNet, the proposed architecture is evaluated in different ways on two data sets (i.e., CIFAR-10 and HistoPhenotypes dataset), showing better results than its ResNet counterpart.
Xiang Li 0006, Wei Li 0032, Qian Du 0001
ICPR4
2018 A Novel Multiple Kernel Learning Framework for Remote Sensing Scene Classification
abstract
In the paper we propose a novel multiple kernel learning framework for representation-based classification (MKL-RC) of remote sensing image scenes. Unlike the existing methods that often greedily learn an optimal combined kernel from predefined base kernels by optimization method, resulting in high computation time but relatively better performance. The proposed approach is different from traditional kernel methods and characterized by multiple feature and multiple kernel learning in a representation-based classification manner. Experimental results on two real remote sensing scene datasets demonstrate that the proposed methods can achieve superior performance than the state-of-the-art classification methods.
Xiaoyong Bian, Yuxia Sheng, Yan Xu 0003, Qian Du 0001
IGARSS4
2018 FPGA Based Implementation of Convolutional Neural Network for Hyperspectral Classification
abstract
convolutional neural network (CNN) has been widely used for hyperspectral classification. Current researches of CN-N based hyperspectral image classification is mainly implemented on graphics processing unit (GPU) platform. However, GPU is not suitable for onboard processing due to the problem of space radiation and power supply on image acquiring platform. Therefore, in this paper, FPGA is selected to implement CNN based hyperspectral classification for further onboard processing. Specially, a hardware model is designed for the forward classification step of CNN using hardware description language, including computation structure for CNN, implementation of different layers, weight loading scheme, and data interfere. Simulation results over Pavia data set validate the proposed FPGA based implementation is coincide with that on GPU platform.
Jingyu Ji, Shaohui Mei, Yifan Zhang 0006, Manli Han, Qian Du 0001
IGARSS6
2018 Joint Feature Extraction for Multispectral and Panchromatic Images Based on Convolutional Neural Network
abstract
Along with very high-resolution satellites were launched frequently, such as the satellite WorldView-3, panchromatic and multispectral remote-sensing images can be acquired easily. However, it is still an interesting and challenging task to fuse and classify these images. In general, panchromatic image has a high spatial resolution, but with only one spectral band. Multispectral image usually has four or eight bands, but the spatial resolution is four times smaller than panchromatic image. In this paper, an unsupervised feature extraction framework is proposed, which combines multispectral (MS) image and panchromatic (PAN) image into convolution neural network (CNN). There is an image-to-image mapping, learning from the input source (i.e., MS) to the output source (i.e., PAN). Then, by integrating the hidden layer of deep CNN, the extracted features represent MS and PAN data. The experimental results of two practical remote sensing data sets show the validity of the framework.
Mengmeng Zhang 0005, Wei Li 0032, Qian Du 0001
IGARSS4
2018 Unsupervised Multi-Class Change Detection in Bitemporal Multispectral Images Using Band Expansion
abstract
This paper focuses on solving the multi-class change detection problem in bitemporal multispectral remote sensing images. In that case, information that represented in a small number (e.g., two) of the original bands may be insufficient for the accurate identification of a few of multi-class changes. In particular, this problem becomes more difficult in unsupervised change detection cases when ground reference data is not available. In this paper, a solution is proposed by using the potential information represented in expanded features that constructed from the original spectral bands. Experimental results obtained on a real bitemporal remote sensing data set confirm the effectiveness of the proposed approach.
Sicong Liu 0001, Qian Du 0001, Lorenzo Bruzzone, Alim Samat, Xiaohua Tong
IGARSS2
2018 Low-Complexity Hyperspectral Image Compression Using Folded PCA and JPEG2000
abstract
Hyperspectral image compression by PCA and JPEG2000 can provide excellent rate distortion performance while preserving essential information for a successive application, e.g., classification tasks. However, for onboard applications, PCA suffers from high computational complexity and large memory requirements due to the eigen-analysis of high-dimensional covariance matrix. Therefore, a computationally more efficient analysis, namely Folded Principal Component Analysis (FPCA) is adopted to perform dimension reduction and combined with JPEG2000 for compression. In FPCA, the spectral vector of hyperspectral pixels is folded into a matrix to compute covariance matrix, by which the dimension of covariance matrix is highly reduced. As a result, both computational complexity and memory requirement in subsequent eigen-analysis is reduced. Experimental results demonstrate that the proposed FPCA+JPEG2000 based compression scheme outperforms existing PCA+JPEG2000 in terms of rate distortion and classification after de-compression.
Shaohui Mei, Bakht Muhammad Khan, Yifan Zhang 0006, Qian Du 0001
IGARSS4
2018 Improved Random Projection with $K$-Means Clustering for Hyperspectral Image Classification
abstract
Random projection based dimensionality reduction methods are particularly attractive options for hyperspectral data analysis, due to their data independent representation, reduction in computation time and storage costs, while preserving data separability and important information at lower dimensions. In this work, we combine the benefits of dimensionality reduction using random projections with feature selection using k-means clustering in low dimensions to achieve a two-fold dimensionality reduction. Supervised classification using support vector machine (SVM) was done to study the classification performance. It is experimentally demonstrated that our proposed random projection based k-means feature selection methods offers superior classification performance at far fewer dimensions than original data without dimensionality reduction.
Vineetha Menon, Qian Du 0001, Sundar A. Christopher
IGARSS2
2018 Spatial-spectral Based Multi-view Low-rank Sparse Sbuspace Clustering for Hyperspectral Imagery
abstract
Hyperspectral image (HSI) Clustering is an unsupervised task, which segments pixels into different groups without using labeled samples. In this paper, spatial-spectral based multi-view low-rank sparse subspace clustering (SSMLC) algorithm is proposed. Due to significant number of spectra bands HSI contains much more information than a regular image. These spectral information can be considered as multiview. In this paper, the spectral partitioning is applied to generate spectral views which contain correlated bands. Morphological features of the original HSI are taken as another view which contains spatial features. Principal components construct another view, which eliminates the noise in the original dataset. After the multi-view dataset is formed, multi-view low-rank sparse subspace clustering is applied to segment HSI. Our experiments show that the performance of the proposed SSMLC is better than other that of clustering algorithms such as sparse subspace clustering and low-rank sparse subspace clustering.
Qian Du 0001, Ivica Kopriva, Nicolas H. Younan
IGARSS2
2018 Hyperspectral Classification Via Spatial Context Exploration with Multi-Scale CNN
abstract
Spatial context has shown to be very useful in hyperspectral image processing. Existing convolutional neural network (CNN)-based methods for hyperspectral classification explore spatial context by single-scale convolution kernels in 2D or 3D shapes. However, such single-scale convolution may not be capable to explore the complex spatial context in a hyperspectral image. In this paper, we propose a multi-scale CNN, MS-CNN to explore the spatial context in different extents, in which adaptive spatial neighborhood convolution kernels are used to simultaneously extract multiple spectral-spatial features from spatial context of pixels. These features obtained by different spatial kernels are then concatenated and fused for further feature extraction and classification. Experimental results show that the proposed adaptive spatial neighborhood convolution are more effective to explore spatial context than traditional single-scale spatial convolution and the performance of the proposed MS-CNN outperforms several state-of-art CNNs for classification of hyperspectral images.
Zhongqi Tian, Jingyu Ji, Shaohui Mei, Junhui Hou, Shuai Wan, Qian Du 0001
IGARSS6
2018 Hyperspectral Image Classification Based on Capsule Network
abstract
In this paper, we propose two novel classification frameworks for hyperspectral image (HSI) based on capsule network (CapsNet), which could address the drawbacks of convolutional neural network (CNN) and problem of limited training samples by introducing affine transformation matrix. Specifically, the proposed framework first performs the classification of HSI based on spectral information. Second, considering the importance of spatial information for HSI processing, we integrate the spatial and spectral information into the proposed framework to further improve the classification performance. Experimental results on real HSI data demonstrate the effectiveness of the proposed framework.
Wei-Ye Wang, Heng-Chao Li 0001, Lei Pan 0003, Gang Yang 0006, Qian Du 0001
IGARSS5
2018 Gabor-Filtering-Based Probabilistic Collaborative Representation for Hyperspectral Image Classification
abstract
This paper presents Gabor-filtering-based probabilistic collaborative representation for hyperspectral image classification. Compared with the original collaborative representation classifier (CRC) and the CRC using Gabor features, the proposed classifier offers superior classification performance. The regularized versions of CRC using Gabor features have excellent classification performance; however, those classifiers have high computational cost. Experimental results show that the proposed approach can generate high classification accuracy with lower computational cost.
Yan Xu 0003, Qian Du 0001, Wei Li 0032, Nicolas H. Younan
IGARSS2
2018 Hyperspectral Classification Based on Siamese Neural Network Using Spectral-Spatial Feature
abstract
Recently, the deep convolutional neural network (CNN) is of great interest in hyperspectral image classification. However, limited available training samples still prevent CNN from exploring the performance of classification. In this work, we employ a novel pixel-pair method based on Siamese neural network (SNN) to significantly enlarge the training set and better represent the spectral-spatial features. In training, two pixels are respectively fed into two branch CNNs to extract deep features, where the same weights and biases are shared. Then, the absolute difference between the two deep features is learned by linear full connection layers with a given label. In testing, pixel-pairs, constructed by combining the center pixel and each of the surrounding pixels, are classified by the trained SNN. The final prediction is then determined by a voting strategy. The proposed SNN framework is extended to learn deep patch-pixel features. Experimental performance demonstrates that the proposed strategy outperforms the traditional classifiers, such as support vector machine (SVM) and extreme learning machine (ELM).
Shizhi Zhao, Wei Li 0032, Qian Du 0001, Qiong Ran
IGARSS3
2018 Nuclear norm-based matrix regression preserving embedding for face recognition
Yangjun Deng, Heng-Chao Li 0001, Qi Wang 0009, Qian Du 0001
Neurocomputing4
2018 Modified Tensor Locality Preserving Projection for Dimensionality Reduction of Hyperspectral Images
abstract
By considering the cubic nature of hyperspectral image (HSI) to address the issue of the curse of dimensionality, we have introduced a tensor locality preserving projection (TLPP) algorithm for HSI dimensionality reduction and classification. The TLPP algorithm reveals the local structure of the original data through constructing an adjacency graph. However, the hyperspectral data are often susceptible to noise, which may lead to inaccurate graph construction. To resolve this issue, we propose a modified TLPP (MTLPP) via building an adjacency graph on a dual feature space rather than the original space. To this end, the region covariance descriptor is exploited to characterize a region of interest around each hyperspectral pixel. The resulting covariances are the symmetric positive definite matrices lying on a Riemannian manifold such that the Log-Euclidean metric is utilized as the similarity measure for the search of the nearest neighbors. Since the defined covariance feature is more robust against noise, the constructed graph can preserve the intrinsic geometric structure of data and enhance the discriminative ability of features in the low-dimensional space. The experimental results on two real HSI data sets validate the effectiveness of our proposed MTLPP method.
Yangjun Deng, Heng-Chao Li 0001, Lei Pan 0003, Li-Yang Shao, Qian Du 0001, William J. Emery
IEEE Geosci. Remote. Sens. Lett.5
2018 Object Tracking in Satellite Videos by Fusing the Kernel Correlation Filter and the Three-Frame-Difference Algorithm
abstract
Object tracking is a popular topic in the field of computer vision. The detailed spatial information provided by a very high resolution remote sensing sensor makes it possible to track targets of interest in satellite videos. In recent years, correlation filters have yielded promising results. However, in terms of dealing with object tracking in satellite videos, the kernel correlation filter (KCF) tracker achieves poor results due to the fact that the size of each target is too small compared with the entire image, and the target and the background are very similar. Therefore, in this letter, we propose a new object tracking method for satellite videos by fusing the KCF tracker and a three-frame-difference algorithm. A specific strategy is proposed herein for taking advantage of the KCF tracker and the three-frame-difference algorithm to build a strong tracker. We evaluate the proposed method in three satellite videos and show its superiority to other state-of-the-art tracking methods.
Bo Du 0001, Shihan Cai, Chen Wu 0003, Qian Du 0001
IEEE Geosci. Remote. Sens. Lett.5
2018 Classification of Hyperspectral Imagery Using a New Fully Convolutional Neural Network
abstract
With success of convolutional neural networks (CNNs) in computer vision, the CNN has attracted great attention in hyperspectral classification. Many deep learning-based algorithms have been focused on deep feature extraction for classification improvement. In this letter, a novel deep learning framework for hyperspectral classification based on a fully CNN is proposed. Through convolution, deconvolution, and pooling layers, the deep features of hyperspectral data are enhanced. After feature enhancement, the optimized extreme learning machine (ELM) is utilized for classification. The proposed framework outperforms the existing CNN and other traditional classification algorithms by including deconvolution layers and an optimized ELM. Experimental results demonstrate that it can achieve outstanding hyperspectral classification performance.
Jiaojiao Li 0001, Yunsong Li 0001, Qian Du 0001, Bobo Xi, Jing Hu 0005
IEEE Geosci. Remote. Sens. Lett.4
2018 Hyperspectral Image Reconstruction by Latent Low-Rank Representation for Classification
abstract
To effectively reduce the spectral variation that degrades classification performance, a novel low-rank subspace recovery method based on latent low-rank representation (LatLRR) is proposed for hyperspectral images in this letter. Different from the robust principal component analysis, LatLRR focuses on exploring the low-rank property from the perspective of row space and column space simultaneously through the low-rank regularization on their corresponding coefficient matrix. Following that, the self-expressiveness-based reconstruction is adopted to recover the intrinsic data from row and column spaces. More accurate subspace structure can be successfully preserved both in spectral domain and spatial domain; meanwhile, the robustness to noise is improved. Experimental results on two hyperspectral data sets demonstrate the effectiveness of the proposed method.
Lei Pan 0003, Heng-Chao Li 0001, Yong-Jian Sun, Qian Du 0001
IEEE Geosci. Remote. Sens. Lett.4
2018 Unsupervised Hyperspectral Remote Sensing Image Clustering Based on Adaptive Density
abstract
Hyperspectral remote sensing image (HSI) clustering can be defined as the process of segmenting pixels into different sets that satisfy the requirement that the differences between sets are much greater than the differences within sets. According to the fast density peak-based clustering algorithm, we propose an unsupervised HSI clustering method based on the density of pixels in the spectral space and the distance between pixels. For the metric of the density, we present an adaptive-bandwidth probability density function using pixel numbers as the input and the calculated pixel local density as the output, which determines the bandwidth on the basis of the Gaussian assumption. For the metric of the distance, in order to obtain a pixel-level spectral distance, we calculate the Euclidean distance between pixel vectors from the multiple bands. In the proposed approach: 1) use the least-squares method for the curve fitting of the two results; 2) eliminate outliers based on the Pauta criterion; 3) adopt regression calculation; and 4) obtain the cluster centers according to the classification criteria of the local density and the distance between pixel vectors. The other noncluster center points are clustered based on their similarities with the cluster centers by iteration. Finally, we compare the results with those of other unsupervised clustering methods and the reference data sets.
Huan Xie 0001, Ang Zhao, Sicong Liu 0001, Xiong Xu 0001, Xin Luo 0003, Haiyan Pan, Qian Du 0001, Xiaohua Tong
IEEE Geosci. Remote. Sens. Lett.9
2018 Tensor Low-Rank Discriminant Embedding for Hyperspectral Image Dimensionality Reduction
abstract
Recently, low-rank embedding (LRE) has yielded satisfactory results in dimensionality reduction (DR), for which low-rank representation and projection learning are integrated into one model to generate robust low-dimensional features. However, LRE requires to convert samples into vectors even if the data naturally appear in high-order form. Furthermore, LRE fails to take the label information into consideration. To address these problems, this paper proposes a novel supervised DR method based on multilinear algebra, i.e., the algebra of tensors. By the motivation of extending LRE into tensor space and simultaneously combining the tensor discriminant analysis, we establish tensor low-rank discriminant embedding (TLRDE) model for hyperspectral image (HSI) DR. The model of TLRDE is solved by an alternative iteration algorithm, whose convergence is also mathematically proven. The proposed TLRDE method employs the tensor representation to preserve the intrinsic geometrical structure, uses low-rank reconstruction to uncover the potential relationship among the data points, and combines label information to enhance the discriminability of features. Moreover, the proposed TLRDE does not suffer from the small sample size problem. The experimental results on three real HSI data sets validate the effectiveness of our proposed TLRDE method.
Yangjun Deng, Heng-Chao Li 0001, Kun Fu 0001, Qian Du 0001, William J. Emery
IEEE Trans. Geosci. Remote. Sens.4
2018 Hyperspectral Unmixing Using Sparsity-Constrained Deep Nonnegative Matrix Factorization With Total Variation
abstract
Hyperspectral unmixing is an important processing step for many hyperspectral applications, mainly including: 1) estimation of pure spectral signatures (endmembers) and 2) estimation of the abundance of each endmember in each pixel of the image. In recent years, nonnegative matrix factorization (NMF) has been highly attractive for this purpose due to the nonnegativity constraint that is often imposed in the abundance estimation step. However, most of the existing NMF-based methods only consider the information in a single layer while neglecting the hierarchical features with hidden information. To alleviate such limitation, in this paper, we propose a new sparsity-constrained deep NMF with total variation (SDNMF-TV) technique for hyperspectral unmixing. First, by adopting the concept of deep learning, the NMF algorithm is extended to deep NMF model. The proposed model consists ofpretraining stageandfine-tuning stage, where the former pretrains all factors layer by layer and the latter is used to reduce the total reconstruction error. Second, in order to exploit adequately the spectral and spatial information included in the original hyperspectral image, we enforce two constraints on the abundance matrix. Specifically, the$L_{1/2}$constraint is adopted, since the distribution of each endmember is sparse in the 2-D space. The TV regularizer is further introduced to promote piecewise smoothness in abundance maps. For the optimization of the proposed model, multiplicative update rules are derived using the gradient descent method. The effectiveness and superiority of the SDNMF-TV algorithm are demonstrated by comparing with other unmixing methods on both synthetic and real data sets.
Xin-Ru Feng, Heng-Chao Li 0001, Jun Li 0009, Qian Du 0001, Antonio Plaza, William J. Emery
IEEE Trans. Geosci. Remote. Sens.4
2018 Hyperspectral Image Classification With Imbalanced Data Based on Orthogonal Complement Subspace Projection
abstract
Conventional classification algorithms have shown great success for balanced classes. In remote sensing applications, it is often the case that classes are imbalanced. This paper proposes a novel solution to solve the problem of imbalanced training samples in hyperspectral image classification. It consists of two parts: one is for large-size sample sets and the other is for small-size sets. Specifically, an algorithm based on the orthogonal complement subspace projection (OCSP) is proposed to select samples from large-size classes, and an algorithm also based on OCSP is proposed to create artificial samples for small-size ones. The impact on representation-based classifiers, i.e., sparse and collaborative representation classifiers and traditional classifiers (e.g., support vector machine), is investigated. Experimental results demonstrate that the proposed solution can outperform other existing solutions in the literature.
Jiaojiao Li 0001, Qian Du 0001, Yunsong Li 0001, Wei Li 0032
IEEE Trans. Geosci. Remote. Sens.2
2018 Simultaneous Spatial and Spectral Low-Rank Representation of Hyperspectral Images for Classification
abstract
Arising from various environmental and atmos- pheric conditions and sensor interference, spectral variations are inevitable during hyperspectral remote sensing, which degrade the subsequent hyperspectral image analysis significantly. In this paper, we propose simultaneous spatial and spectral low-rank representation (S3LRR) that can effectively suppress the within-class spectral variations for classification purposes. The S3LRR recovers an intrinsic component with the same dimension as the original image, in which both spatial and spectral low-rank priors are adopted to regularize the intrinsic component simultaneously and compensate to each other, together with robust modeling of spectral variations. Compared with existing methods that explore only the spectral low-rank prior, the novel spatial low-rank prior (i.e., low-rank prior in band-wise) can take the spatial structure information of hyperspectral images into account, which has demonstrated to be very useful. Technically, we formulate S3LRR as a constrained convex optimization problem, and solve it using the efficient inexact augmented Lagrangian multiplier method. The resulting intrinsic component is less interfered by within-class spectral variations, and more discriminatory to offer higher classification accuracy. Comprehensive experiments on benchmark data sets demonstrate that the proposed S3LRR improves classification accuracy significantly, which outperforms state-of-the-art methods.
Shaohui Mei, Junhui Hou, Jie Chen 0026, Lap-Pui Chau, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.5
2018 Multifeature Dictionary Learning for Collaborative Representation Classification of Hyperspectral Imagery
abstract
Recently, multifeature learning in collaborative representation classification (CRC) for hyperspectral images has generated promising performance. In this paper, two novel multifeature learning algorithms that update dictionary directly and indirectly are proposed. In order to offer the complementarity of multifeature, four different types of features-global feature (i.e., Gabor feature), local feature (i.e., local binary pattern), shape feature (i.e., extended multiattribute profiles), and spectral feature-are adopted in this paper. Under the hypothesis that most of the features should share the same coding pattern in CRC, this paper proposes to learn proper dictionaries for each feature until obtaining stable codes in a linear classifier. Furthermore, to avoid the explicit mapping of infinite-dimensional dictionaries in a nonlinear kernelized classifier, an indirect approach to construct the transformation matrix from original dictionaries to learn new dictionaries is developed. Three real hyperspectral images acquired from different sensors are adopted for performance evaluation. The experimental results demonstrate that the proposed methods can provide superior performance compared with those of the state-of-the-art classifiers.
Hongjun Su, Qian Du 0001, Peijun Du, Zhaohui Xue
IEEE Trans. Geosci. Remote. Sens.3
2018 Graph-Regularized Fast and Robust Principal Component Analysis for Hyperspectral Band Selection
abstract
A fast and robust principal component analysis on Laplacian graph (FRPCALG) method is proposed to select bands of hyperspectral imagery (HSI). The FRPCALG assumes that a clean band matrix lies in a unified manifold subspace with low-rank and clustering properties, whereas sparse noise does not lie in the same subspace. It estimates the clean lowrank approximation of the original HSI band matrix while uncovering the clustering structure of all bands. Specifically, a structured random projection is adopted to reduce the high spatial dimensionality of the original data for computational cost saving, and then a Laplacian graph (LG) term is regularized into the regular robust principal component analysis (RPCA) to formulate the FRPCALG model for the submatrix of bands to be selected. The RPCA term ensures the clean and low-rank approximation of original data, and the LG term guarantees the clustering quality of a low-rank matrix in the low-dimensional manifold subspace. The alternating direction method of multipliers' algorithm is utilized to optimize the convex program of the FRPCALG. The K-means algorithm is to group all columns of submatrix into clusters, and corresponding bands closest to their cluster centroids finally constitute the desired band subset. Experimental results show that FRPCALG outperforms state-ofthe-art methods with lower computational cost. A moderate regularization parameter λ and a small μ could guarantee satisfying the classification accuracy of FRPCALG, and a small projected dimension greatly reduces the computational cost and does not affect the classification performance. Therefore, the FRPCALG can be an alternative method for hyperspectral band selection.
Weiwei Sun 0005, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.2
2018 Multisource Remote Sensing Data Classification Based on Convolutional Neural Network
abstract
As a list of remotely sensed data sources is available, how to efficiently exploit useful information from multisource data for better Earth observation becomes an interesting but challenging problem. In this paper, the classification fusion of hyperspectral imagery (HSI) and data from other multiple sensors, such as light detection and ranging (LiDAR) data, is investigated with the state-of-the-art deep learning, named the two-branch convolution neural network (CNN). More specific, a two-tunnel CNN framework is first developed to extract spectral-spatial features from HSI; besides, the CNN with cascade block is designed for feature extraction from LiDAR or high-resolution visual image. In the feature fusion stage, the spatial and spectral features of HSI are first integrated in a dual-tunnel branch, and then combined with other data features extracted from a cascade network. Experimental results based on several multisource data demonstrate the proposed two-branch CNN that can achieve more excellent classification performance than some existing methods.
Wei Li 0032, Qiong Ran, Qian Du 0001, Lianru Gao, Bing Zhang 0001
IEEE Trans. Geosci. Remote. Sens.4
2017 Hyperspectral image super-resolution via convolutional neural network
abstract
Due to the tradeoff between spatial and spectral resolution in remote sensing imaging, hyperspectral images are often acquired with a relative low spatial resolution, which limits their applications in many areas. Inspired by recent achievements in convolutional neural network (CNN) based super resolution (SR), a novel CNN based framework is constructed for SR of hyperspectral images by considering both spatial context and spectral correlation. As a result, the spectral distortion incurred by directly applying traditional SR algorithms to hyperspectral images is alleviated. Experimental results on several benchmark hyperspectral datasets have demonstrated that higher quality of reconstruction and spectral fidelity can be achieved, compared to band-wise manner based algorithms.
Shaohui Mei, Xin Yuan 0002, Jingyu Ji, Shuai Wan, Junhui Hou, Qian Du 0001
ICIP6
2017 Fusing two convolutional neural networks for high-resolution scene classification
abstract
This paper presents a novel deep convolutional feature fusion (ConvFF) approach for high-resolution scene classification, characterizing the well-known deep convolutional neural network (ConvNet) approach. The proposed ConvFF approach starts by generating an initial feature representation of the original scenes under exploration from two deep ConvNets pre-trained on two different large amount of labeled data. After the pre-training phase, we fine tune the two deep ConvNets consisting of mainly objects and scenes respectively in a supervised manner using the target training images. Then we propose to fuse the extracted two types of convolutional features provided by the last fully-connected (FC) layer, respectively. Finally, the fused convolutional features are fed as input to a SVM classifier for classification. The proposed method is evaluated by using two challenging high-resolution scene datasets. Experimental results show that the proposed method can effectively extract complementary features of the scenes and capture local spatial patterns, consistently outperforming several state-of-the-art methods.
Xiaoyong Bian, Chen Chen 0001, Yuxia Sheng, Yan Xu 0003, Qian Du 0001
IGARSS5
2017 Sparse graph embedding dimension reduction for hyperspectral image with a new spectral similarity metric
abstract
Graph embedding, as a dimensionality reduction framework, has already drawn great attention in hyperspectral image analysis. Taking locality preserving projection (LPP) as example, LPP utilizes typical Euclidean distance in heat kernel to create an affinity matrix and projects the high-dimensional data into a lower-dimensional space. However, the Euclidean distance is not sufficiently correlated with intrinsic spectral variation of a material, which may result in inappropriate graph representation. In this work, a graph-based discriminant analysis with novel spectral similarity measurement is proposed, which fully considers curves changing description among spectral bands. Experimental results based on real hyperspectral images demonstrate the proposed method is superior to traditional methods, such as supervised LPP, and the state-of-the-art sparse graph-based discriminant analysis (SGDA).
Fubiao Feng, Wei Li 0032, Qian Du 0001, Qiong Ran
IGARSS3
2017 Learning sensor-specific features for hyperspectral images via 3-dimensional convolutional autoencoder
abstract
Deep learning techniques have brought in revolutionary achievements for feature learning of images. In this paper, a novel structure of 3-Dimensional Convolutional AutoEncoder (3D-CAE) is proposed for hyperspectral spatial-spectral feature learning, in which the spatial context is considered by constructing a 3-Dimensional input using pixels in a spatial neighborhood. All the parameters involved in the 3D-CAE are trained without the need of labeled training samples such that feature learning is conducted in an unsupervised fashion. Such unsupervised spatial-spectral feature extraction is also extended to different images from the same sensor to learn sensor-specific features. As a result, spatial-spectral features of hyperspectral images are extracted for a specific sensor under an unsupervised manner. Experimental results on several benchmark hyperspectral datasets have demonstrated that our proposed 3D-CAE are very effective in extracting sensor-specific spatial-spectral features and outperform several state-of-the-art deep learning neural networks in classification application.
Jingyu Ji, Shaohui Mei, Junhui Hou, Xu Li 0010, Qian Du 0001
IGARSS5
2017 Transferred deep learning for hyperspectral target detection
abstract
An interesting target detection framework with transferred deep convolutional neural network (CNN) is proposed. For CNN, many labeled samples are needed to train the multi-layer network. However, for target detection tasks, only few target spectral signatures are available, or they are unknown in anomaly detection. In this work, we employ a reference data and further generate pixel-pairs to enlarge the sample size. A multi-layer CNN is trained by using difference between pixel-pairs generated from the reference image scene. During testing, there are two cases: (1) for anomaly detection, difference between pixel-pairs, constructed by combing the center pixel and its surrounding pixels, is classified by the trained CNN with result of similarity measurement; and (2) for supervised target detection, difference between pixel-pairs, constructed by combing the testing pixel and the known spectral signatures, is classified. The detection output is simply generated by averaging these similarity scores. Experimental performance demonstrates that the proposed strategy outperforms the classic detectors.
Wei Li 0032, Guodong Wu, Qian Du 0001
IGARSS3
2017 A spectral-spatial multiscale approach for unsupervised multiple change detection
abstract
A novel spectral-spatial joint multiscale approach is developed to address the multi-class change detection problem in bitemporal multispectral remote sensing images. The proposed approach is based on a multiscale morphological compressed change vector analysis (M2C2VA), which extend the state-of-the-art spectrum-based compressed change vector analysis (C2VA) while preserving more geometrical details of change targets. In particular, spectral change features are reconstructed according to the morphological analysis which exploiting the interaction of a pixel with its adjacent regions. Two multiscale ensemble strategies are proposed to integrate the change information represented at multiple scales in order to enhance the CD performance. The proposed approach is designed in an unsupervised fashion thus can be implemented without using ground reference data. A pair of real bitemporal remote sensing images is used to test the proposed approach and the obtained experimental results confirm its effectiveness.
Sicong Liu 0001, Qian Du 0001, Xiaohua Tong, Alim Samat, Lorenzo Bruzzone, Francesca Bovolo
IGARSS2
2017 Fusing different levels of deep features by deep stacked neural network for hyperspectral images
abstract
Deep learning techniques have been demonstrated to be a powerful tool to learn features of images automatically. In this paper, a novel deep learning structure, i.e., deep stacked neural network (DSNN), is constructed to extract different levels of deep features of hyperspectral images. Specifically, convolutional neural network (CNN) is used as basic units in the proposed DSNN for feature extraction of hyperspectral images. Then, different levels of deep features are concatenated to form a novel fused feature for classification with a typical classifier, e.g., SVM. Experimental results on two benchmark hyperspectral datasets show that the fusion of features extracted in DSNN can produce higher classification accuracy than state-of-the-art deep learning based methods, indicating its effectiveness in feature learning.
Shaohui Mei, Yanfu Chen, Jingyu Ji, Junhui Hou, Qian Du 0001
IGARSS5
2017 Tensor-based offset-sparsity decomposition for hyperspectral image classification
abstract
In this paper, the tensor-based offset-sparsity decomposition (TOSD) method, or low-rank and sparse decomposition, is applied to hyperspectral imagery, where the low-rank tensor is considered to be enhanced or pruned data and used for classification. In the tensor form of dataset, all the information of the original 3D data cube, includes spatial and spectral information, can be better reserved. To make the low-rank assumption more possibly true, spatial and spectral segmentations are conducted in a preprocessing step for the TOSD. The experimental results demonstrate the TOSD offers better performance than the matrix-based one, and the spatial-spectral segmentation can further improve the performance.
Qian Du 0001, Nicolas H. Younan, Ivica Kopriva
IGARSS2
2017 Nonlinear classification of multispectral imagery using representation-based classifiers
abstract
The paper investigates representation-based classification for multispectral imagery. Due to the limited spectral dimension, the performance may be limited, and, in general, it is difficult to discriminate different classes using multispectral imagery. Nonlinear band generation method is proposed to use which can provide additional spectral information for multispectral classification. Two classifiers, sparse representation-based classification (SRC) and Nearest Regularized Subspace (NRS) are evaluated on the generated datasets. The results show our approach can outperform other nonlinear method such as the traditional kernel method in terms of classification accuracy and computational cost.
Yan Xu 0003, Qian Du 0001, Wei Li 0032, Chen Chen 0001, Nicolas H. Younan
IGARSS2
2017 Transferred Deep Learning for Anomaly Detection in Hyperspectral Imagery
abstract
In this letter, a novel anomaly detection framework with transferred deep convolutional neural network (CNN) is proposed. The framework is designed by considering the following facts: 1) a reference data with labeled samples are utilized, because no prior information is available about the image scene for anomaly detection and 2) pixel pairs are generated to enlarge the sample size, since the advantage of CNN can be realized only if the number of training samples is sufficient. A multilayer CNN is trained by using difference between pixel pairs generated from the reference image scene. Then, for each pixel in the image for anomaly detection, difference between pixel pairs, constructed by combining the center pixel and its surrounding pixels, is classified by the trained CNN with the result of similarity measurement. The detection output is simply generated by averaging these similarity scores. Experimental performance demonstrates that the proposed algorithm outperforms the classic Reed-Xiaoli and the state-of-the-art representation-based detectors, such as sparse representation-based detector (SRD) and collaborative representation-based detector.
Wei Li 0032, Guodong Wu, Qian Du 0001
IEEE Geosci. Remote. Sens. Lett.3
2017 Random Hadamard Projections for Hyperspectral Unmixing
abstract
Dimensionality reduction based on random projections is investigated in the context of spectral unmixing of hyperspectral imagery with aims toward unmixing accuracy and computational efficiency. To this end, both Hadamard-based random projections-which significantly reduce computational costs with respect to more traditional Gaussian-driven projections-as well as a fast singular value decomposition deployed within a random-projection space are considered. Experimental results reveal that the methods based on Hadamard random projections offer abundance-estimation performance superior to other methods in conjunction with significantly reduced computational complexity.
Vineetha Menon, Qian Du 0001, James E. Fowler
IEEE Geosci. Remote. Sens. Lett.2
2017 Hyperspectral Image Classification via Low-Rank and Sparse Representation With Spectral Consistency Constraint
abstract
In this letter, a low-rank and sparse representation classifier with a spectral consistency constraint (LRSRC-SCC) is proposed. Different from the SRC that represents samples individually, LRSRC-SCC reconstructs samples jointly and is able to capture the local and global structures simultaneously. In this proposed classifier, an adaptive spectral constraint is imposed on both the low-rank and sparse terms so as to better reveal the data structure and enhance its discriminative power. In addition, the alternating direction method is introduced to solve the underlying minimization problem, in which, more importantly, the subobjective function associated with the low-rank term is optimized based on the rank equivalence between a matrix and its Gram matrix, resulting in a closed-form solution. Finally, LRSRC-SCC is extended to LRSRC-SCCE for fully exploiting the spatial information. Experimental results on two hyperspectral data sets demonstrate that the proposed LRSRC-SCC and LRSRC-SCCE methods outperform some state-of-the-art methods.
Lei Pan 0003, Heng-Chao Li 0001, Hua Meng 0001, Wei Li 0032, Qian Du 0001, William J. Emery
IEEE Geosci. Remote. Sens. Lett.5
2017 Subpixel Change Detection of Multitemporal Remote Sensed Images Using Variability of Endmembers
abstract
Due to the existence of mixed pixels in a remote sensed image, traditional change detection (CD) methods at “full-pixel level” are often unable to provide detailed changed information effectively. A subpixel change detection (SCD) technique can deal with this issue with two steps: soft classification is applied to derive proportional differences from coarse multitemporal images, and then a sharpened thematic map with fine spatial resolution is generated based on subpixel mapping. However, changes in endmember combination within pixels are ignored, which can result in flawed differences and degraded accuracy of SCD. The aim of this letter is to present a new SCD algorithm using variability of endmembers (SCD_VE), where a simple but effective model is proposed to take into consideration the real change of endmember combination. In order to evaluate the performance of the new algorithm, experiment is conducted on simulated images. Experimental results demonstrated that the proposed SCD_VE offers better performance than traditional SCD methods in providing more detailed CD map.
Ke Wu 0004, Qian Du 0001
IEEE Geosci. Remote. Sens. Lett.2
2017 Particle Swarm Optimization-Based Band Selection for Hyperspectral Target Detection
abstract
This letter proposes particle swarm optimization (PSO)-based band selection (BS) approach for hyperspectral target detection. Due to lack of training samples in a detection problem, it is more difficult than classification-purposed BS. The objective function, called maximum-submaximum-ratio (MSR) gauging target-background separation, is proposed for target detection during PSO searching. Typical target detectors such as target-constrained interference-minimized filter and adaptive coherence estimator are studied. Experimental results demonstrate that the proposed MSR-based objective function in conjunction with PSO-based searching can select a small band set while yielding similar or even better detection performance than using all the original bands, sequential forward search-based BS, or BS relying on detection map similarity assessment.
Yan Xu 0003, Qian Du 0001, Nicolas H. Younan
IEEE Geosci. Remote. Sens. Lett.2
2017 Locality Sensitive Discriminant Analysis for Group Sparse Representation-Based Hyperspectral Imagery Classification
abstract
This letter proposes to integrate the locality sensitive discriminant analysis (LSDA) with the group sparse representation (GSR) for a hyperspectral imagery classification. The LSDA is to project the data set to a lower-dimensional subspace to preserve local manifold structure and discriminant information, while the GSR is to encode the projected testing set as a sparse linear combination of group-structured training samples for classification. The proposed approach, denoted as LSDA-GSR classifier (GSRC), is evaluated using two real hyperspectral data sets. Experimental results demonstrate that it can provide considerable improvement to the original counterparts, i.e., SRC and GSRC, with a relatively low computational cost.
Haoyang Yu 0001, Lianru Gao, Wei Li 0032, Qian Du 0001, Bing Zhang 0001
IEEE Geosci. Remote. Sens. Lett.4
2017 Hyperspectral Image Classification Using Deep Pixel-Pair Features
abstract
The deep convolutional neural network (CNN) is of great interest recently. It can provide excellent performance in hyperspectral image classification when the number of training samples is sufficiently large. In this paper, a novel pixel-pair method is proposed to significantly increase such a number, ensuring that the advantage of CNN can be actually offered. For a testing pixel, pixel-pairs, constructed by combining the center pixel and each of the surrounding pixels, are classified by the trained CNN, and the final label is then determined by a voting strategy. The proposed method utilizing deep CNN to learn pixel-pair features is expected to have more discriminative power. Experimental results based on several hyperspectral image data sets demonstrate that the proposed method can achieve better classification performance than the conventional deep learning-based method.
Wei Li 0032, Guodong Wu, Fan Zhang 0007, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.4
2017 Learning Sensor-Specific Spatial-Spectral Features of Hyperspectral Images via Convolutional Neural Networks
abstract
Convolutional neural network (CNN) is well known for its capability of feature learning and has made revolutionary achievements in many applications, such as scene recognition and target detection. In this paper, its capability of feature learning in hyperspectral images is explored by constructing a five-layer CNN for classification (C-CNN). The proposed C-CNN is constructed by including recent advances in deep learning area, such as batch normalization, dropout, and parametric rectified linear unit (PReLU) activation function. In addition, both spatial context and spectral information are elegantly integrated into the C-CNN such that spatial-spectral features are learned for hyperspectral images. A companion feature-learning CNN (FL-CNN) is constructed by extracting fully connected feature layers in this C-CNN. Both supervised and unsupervised modes are designed for the proposed FL-CNN to learn sensor-specific spatial-spectral features. Extensive experimental results on four benchmark data sets from two well-known hyperspectral sensors, namely airborne visible/infrared imaging spectrometer (AVIRIS) and reflective optics system imaging spectrometer (ROSIS) sensors, demonstrate that our proposed C-CNN outperforms the state-of-the-art CNN-based classification methods, and its corresponding FL-CNN is very effective to extract sensor-specific spatial-spectral features for hyperspectral applications under both supervised and unsupervised modes.
Shaohui Mei, Jingyu Ji, Junhui Hou, Xu Li 0010, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.5
2017 Discriminant Analysis of Hyperspectral Imagery Using Fast Kernel Sparse and Low-Rank Graph
abstract
Due to the high-dimensional characteristic of hyperspectral images, dimensionality reduction (DR) is an important preprocessing step for classification. Recently, sparse and low-rank graph-based discriminant analysis (SLGDA) has been developed for DR of hyperspectral images, for which the properties of sparsity and low-rankness are simultaneously exploited to capture both local and global structures. However, SLGDA may not achieve satisfactory results when handling complex data with nonlinear nature. To address this problem, this paper presents two kernel extensions of SLGDA. In the first proposed classical kernel SLGDA (cKSLGDA), the kernel trick is exploited to implicitly map the original data into a high-dimensional space. With a totally different perspective, we further propose a Nyström-based kernel SLGDA (nKSLGDA) by constructing a virtual kernel space by the Nyström method, in which virtual samples can be explicitly obtained from the original data. Both cKSLGDA and nKSLGDA can achieve more informative graphs than SLGDA, and offer superiority over other state-of-the-art DR methods. More importantly, the nKSLGDA can outperform cKSLGDA with much lower computational cost.
Lei Pan 0003, Heng-Chao Li 0001, Wei Li 0032, Guangning Wu, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.6
2017 Robust Joint Sparse Representation Based on Maximum Correntropy Criterion for Hyperspectral Image Classification
abstract
Joint sparse representation (JSR) has been a popular technique for hyperspectral image classification, where a testing pixel and its spatial neighbors are simultaneously approximated by a sparse linear combination of all training samples, and the testing pixel is classified based on the joint reconstruction residual of each class. Due to the least-squares representation of the approximation error, the JSR model is usually sensitive to outliers, such as background, noisy pixels, and outlying bands. In order to eliminate such effects, we propose three correntropy-based robust JSR (RJSR) models, i.e., RJSR for handling pixel noise, RJSR for handling band noise, and RJSR for handling both pixel and band noise. The proposed RJSR models replace the traditional square of the Euclidean distance with the correntropy-based metric in measuring the joint approximation error. To solve the correntropy-based joint sparsity model, a half-quadratic optimization technique is developed to convert the original nonconvex and nonlinear optimization problem into an iteratively reweighted JSR problem. As a result, the optimization of our models can handle the noise in neighboring pixels and the noise in spectral bands. It can adaptively assign small weights to noisy pixels or bands and put more emphasis on noise-free pixels or bands. The experimental results using real and simulated data demonstrate the effectiveness of our models in comparison with the related state-of-the-art JSR models.
Jiangtao Peng, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.2
2016 Extended multi-structure local binary pattern for high-resolution image scene classification
abstract
This paper presents a novel extended multi-structure local binary pattern (EMSLBP) approach for high-resolution image classification, generalizing the well-known local binary pattern (LBP) approach. In the proposed EMSLBP approach, three-coupled descriptors with multi-structure sampling are proposed to extract complementary features (pixel value and radial difference) from local image patches. The anisotropic features derived from elliptical sampling are also rotation invariant by averaging the histograms over rotational angles and combined with the isotropic features extracted from circular sampling. Experimental results show that the proposed method can effectively capture local spatial pattern and local contrast, consistently outperforming several state-of-the-art classification algorithms.
Xiaoyong Bian, Chen Chen 0001, Qian Du 0001, Yuxia Sheng
IGARSS3
2016 Using CNN-based high-level features for remote sensing scene classification
abstract
In this paper, convolutional neural networks (CNNs) is employed for remote-sensing scene classification, which fully utilizes the semantic features extracted from the images while ignoring some traditional features. Consider the limited labeled samples, CaffeNet model as the pre-trained architecture is adopted. By fine-tuning the pre-trained models, the proposed method is expected to be robust and efficient. Its performance is evaluated with two remote-sensing scene datasets. From the experimental results, the proposed CNN-based scene classification method does provide more excellent performance and be superior to several state-of-the-art methods.
Zhengzheng Fang, Wei Li 0032, Jinyi Zou, Qian Du 0001
IGARSS4
2016 Representation-based hyperspectral image classification with imbalanced data
abstract
This paper proposes a novel solution to solve the problem of imbalanced training samples in hyperspectral image classification. It consists of two parts: one is for large-size sample sets and the other is for small-size sets. We exploit an orthogonal projection based algorithm to select samples from large-size ones; meanwhile, we propose an algorithm based on the orthogonal complementary subspace projection to create artificial samples for small-size ones. The impact on representation based classifiers, i.e., sparse representation based classifier and collaborative representation based classifier, are investigated. Experimental results demonstrate that it can outperform other traditional solutions.
Jiaojiao Li 0001, Qian Du 0001, Wei Li 0032, Yunsong Li 0001
IGARSS2
2016 How to fully explore the low-rank property for data recovery of hyperspectral images
abstract
The performance of hyperspectral classification is affected by within-class spectral variation since different materials may present similar spectral signatures. In this paper, we investigate how to fully use the low-rank property of hyperspectral images to alleviate spectra variation. Particulary, two effective strategies that explore the low-rank property in local spectral and spatial space are proposed. According to experimental results, we conclude that exploring the low-rank property in local spectral-spatial space can help to alleviate spectral variation and improve the performance of classification obviously for all tested data, while exploring the low-rank property in spatial space is more effective for images presenting large homogeneous areas.
Shaohui Mei, Qianqian Bi, Jingyu Ji, Junhui Hou, Qian Du 0001
IGARSS5
2016 Integrating spectral and spatial information into deep convolutional Neural Networks for hyperspectral classification
abstract
Deep convolutional neural networks (CNNs) have brought in achievements in image classification and target detection. In this paper, we propose a novel five-layer CNN for hyperspectral classification by encountering recent achievement in deep learning area, such as batch normalization, dropout, Parametric Rectified Linear Unit (PReLu) activation function. By taking advantage of the specific characteristics of hyperspectral images, spatial context and spectral information are elegantly integrated into the framework. Experimental results demonstrate that our proposed CNN out- performs the state-of-the-art methods.
Shaohui Mei, Jingyu Ji, Qianqian Bi, Junhui Hou, Qian Du 0001, Wei Li 0032
IGARSS5
2016 Hadamard-Walsh random projection for hyperspectral image classification
abstract
The rich spectral information in hyperspectral imagery gives rise to huge storage and transmission costs. Dimensionality reduction aims to reduce the space complexity in hyperspectral imagery by projecting data into a low-dimensional subspace. There has been an increasing interest in dimensionality reduction driven by random projections due to its data-independent representation as well as desirable qualities such as the preservation of important information and low computational costs. The performance of a random projection derived from a Hadamard-Walsh matrix is investigated, with experimental results demonstrating classification performance superior to other random dimensionality-reduction methods when deployed in conjunction with a composite-kernel support vector machine that exploits both spatial and spectral information for the classification of hyperspectral imagery.
Vineetha Menon, Qian Du 0001, James E. Fowler
IGARSS2
2016 Multispectral image enhancement with extended offset-sparsity decomposition
abstract
In this paper, the extended offset-sparsity decomposition (OSD) method is applied to multispectral image enhancement. Both principle component analysis (PCA) and hue-saturation-value (HSV) transform are considered before the single-band-based OSD is deployed. The objective is to enhance image details while maintaining the original spectral information. Spectral angle is used to evaluate spectral fidelity, while the extended sharpness and contrast quality measurements are for spatial quality assessment. The experimental results demonstrate that extended OSD with HSV offers better performance of enhancement with acceptable spectral distortion.
Qian Du 0001, Nicolas H. Younan, Ivica Kopriva
IGARSS2
2016 Parallel collaborative representation for hyperspectral image classification on GPUs
abstract
Collaborative representation-based classification with distance-weighted Tikhonov regularization (CRT) has offered high accuracy and efficiency. Due to its per-pixel classification nature without a training step, this paper develops a parallel implementation by using compute unified device architecture (CUDA) on graphics processing units (GPUs). To further improve classification accuracy, local binary pattern (LBP) is used for spatial feature extraction, and an unsupervised band selections approach is applied for dimensionality reduction and an optimized collaborative model combining spatial-spectral features is employed. The proposed parallel implementation is able to increase computational efficiency while not degrading classification accuracy when compared with the serial implementations on central processing units (CPUs).
Lucheng Wu, Xiaoming Xie, Wei Li 0032, Qian Du 0001
IGARSS4
2016 Particle swarm optimization-based band selection for hyperspectral target detection
abstract
This paper proposes particle swarm optimization (PSO)-based band selection approach for target detection from hyperspectral imagery. Specifically, typical target detectors such as constrained energy minimization (CEM) and adaptive coherence estimator (ACE) are studied. Due to the lack of training samples in the detection problem, it is more difficult than classification-purposed band selection. Several objective functions are proposed for target detection during PSO searching. In our experiments, we used the PSO with certain criteria to find the best solution for band selection, and show that it can outperform other searching method such as sequential forward search (SFS) in terms of target detection performance.
Yan Xu 0003, Qian Du 0001, Nicolas H. Younan
IGARSS2
2016 Tri_training for remote sensing classification based on multi-scale homogeneity
abstract
In the process of hyperspectral image classification, the number of training samples is the key problem in improvement of classification performance. However, finding training samples are generally difficult and time-consuming. In this paper, we propose a novel semi-supervised approach and attempt to utilize unlabeled samples to improve classification accuracy. Specifically, active learning (AL) and multi-scale homogeneity (MSH) are integrated in a tri_training framework, where unlabeled samples are selected using AL and the labels of unlabeled samples are predicted from rough classification results with consideration of spatial neighborhood information. The MSH method is utilized to process the classification results to generate the final classification results. Moreover, we propose a novel diversity measure to select optimal classifier combination from different classifiers including support vector machine (SVM), multinomial logistic regression (MLR), extreme learning machine (ELM), k-nearest neighbor (KNN), and random forest (RF) etc. Experiments on two real hyperspectral data indicate that the new diversity measure can select an optimal classifier combination, and the proposed approach can effectively improve classification performance.
Jishuai Zhu, Kun Tan 0001, Qian Du 0001
IGARSS3
2016 Scene classification using local and global features with collaborative representation fusion
Jinyi Zou, Wei Li 0032, Chen Chen 0001, Qian Du 0001
Inf. Sci.4
2016 Spectral Variation Alleviation by Low-Rank Matrix Approximation for Hyperspectral Image Analysis
abstract
Spectral variation is profound in remotely sensed images due to variable imaging conditions. The wide presence of such spectral variation degrades the performance of hyperspectral analysis, such as classification and spectral unmixing. In this letter, 11-based low-rank matrix approximation is proposed to alleviate spectral variation for hyperspectral image analysis. Specifically, hyperspectral image data are decomposed into a low-rank matrix and a sparse matrix, and it is assumed that intrinsic spectral features are represented by the low-rank matrix and spectral variation is accommodated by the sparse matrix. As a result, the performance of image data analysis can be improved by working on the low-rank matrix. Experiments on benchmark hyperspectral data sets demonstrate the performance of classification, and spectral unmixing can be clearly improved by the proposed approach.
Shaohui Mei, Qianqian Bi, Jingyu Ji, Junhui Hou, Qian Du 0001
IEEE Geosci. Remote. Sens. Lett.5
2016 Fast SVD With Random Hadamard Projection for Hyperspectral Dimensionality Reduction
abstract
While data-dependent dimensionality reduction has dominated in many applications of hyperspectral imagery, there is increasing interest in data-independent strategies - such as random projections - due to their promise for reduced computational complexity as well as their demonstrated ability to preserve application-important information. Such random-projection-based dimensionality reduction is investigated in the specific context of supervised hyperspectral classification. Both Hadamard- and Gaussian-based random projections are considered, applied alone as well as incorporated into a fast approximate singular value decomposition (SVD). Experimental results reveal that the proposed Hadamard-based random projection with the fast SVD (FSVD) offers a computationally attractive alternative to not only traditional SVD but also Gaussian-based FSVD for dimensionality reduction in hyperspectral classification.
Vineetha Menon, Qian Du 0001, James E. Fowler
IEEE Geosci. Remote. Sens. Lett.2
2016 Hyperspectral Band Selection Using Improved Firefly Algorithm
abstract
An improved firefly algorithm (FA)-based band selection method is proposed for hyperspectral dimensionality reduction (DR). In this letter, DR is formulated as an optimization problem that searches a small number of bands from a hyperspectral data set, and a feature subset search algorithm using the FA is developed. To avoid employing an actual classifier within the band searching process to greatly reduce computational cost, criterion functions that can gauge class separability are preferred; specifically, the minimum estimated abundance covariance and Jeffreys-Matusita distances are employed. The proposed band selection technique is compared with an FA-based method that actually employs a classifier, the well-known sequential forward selection, and particle swarm optimization algorithms. Experimental results show that the proposed algorithm outperforms others, providing an effective option for DR.
Hongjun Su, Bin Yong, Qian Du 0001
IEEE Geosci. Remote. Sens. Lett.3
2016 Tangent Distance-Based Collaborative Representation for Hyperspectral Image Classification
abstract
Recently, collaborative representation for hyperspectral image analysis has received great interest. Due to the effectiveness of local manifold in a tangent space, this letter extends the collaborative representation classification (CRC) mechanism into the tangent space. Specifically, this letter uses simplified tangent distance and a new regularization term and designs a modified classifier innovatively. Moreover, two variants with weighted diagonal matrices to adaptively adjust the regularization terms are developed to further improve the classification performance. In the experiments, two real hyperspectral images were adopted for performance evaluation, and the experimental results demonstrate that the proposed algorithms can significantly improve classification results compared with the original CRC algorithm and other related classifiers.
Hongjun Su, Qian Du 0001, Yehua Sheng
IEEE Geosci. Remote. Sens. Lett.3
2016 Special issue on advances in pattern recognition in remote sensing
Qian Du 0001, Eckart Michaelsen, Bing Zhang 0001, Jocelyn Chanussot
Pattern Recognit. Lett.1
2016 A survey on representation-based classification and detection in hyperspectral remote sensing imagery
Wei Li 0032, Qian Du 0001
Pattern Recognit. Lett.2
2016 An efficient radial basis function neural network for hyperspectral remote sensing image classification
Jiaojiao Li 0001, Qian Du 0001, Yunsong Li 0001
Soft Comput.2
2016 Laplacian Regularized Collaborative Graph for Discriminant Analysis of Hyperspectral Imagery
abstract
Collaborative graph-based discriminant analysis (CGDA) has been recently proposed for dimensionality reduction and classification of hyperspectral imagery, offering superior performance. In CGDA, a graph is constructed by ℓ2- norm minimization-based representation using available labeled samples. Different from sparse graph-based discriminant analysis (SGDA) where a graph is built by ℓ1- norm minimization, CGDA benefits from within-class sample collaboration and computational efficiency. However, CGDA does not consider data manifold structure reflecting geometric information. To improve CGDA in this regard, a Laplacian regularized CGDA (LapCGDA) framework is proposed, where a Laplacian graph of data manifold is incorporated into the CGDA. By taking advantage of the graph regularizer, the proposed method not only can offer collaborative representation but also can exploit the intrinsic geometric information. Moreover, both CGDA and LapCGDA are extended into kernel versions to further improve the performance. Experimental results on several different multiple-class hyperspectral classification tasks demonstrate the effectiveness of the proposed LapCGDA.
Wei Li 0032, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.2
2016 Sparse and Low-Rank Graph for Discriminant Analysis of Hyperspectral Imagery
abstract
Recently, sparse graph-based discriminant analysis (SGDA) has been developed for the dimensionality reduction and classification of hyperspectral imagery. In SGDA, a graph is constructed by ℓ1-norm optimization based on available labeled samples. Different from traditional methods (e.g., k-nearest neighbor with Euclidean distance), weights in an ℓ1-graph derived via a sparse representation can automatically select more discriminative neighbors in the feature space. However, the sparsity-based graph represents each sample individually, lacking a global constraint on each specific solution. As a consequence, SGDA may be ineffective in capturing the global structures of data. To overcome this drawback, a sparse and low-rank graph-based discriminant analysis (SLGDA) is proposed. Low-rank representation has been proved to be capable of preserving global data structures, although it may result in a dense graph. In SLGDA, a more informative graph is constructed by combining both sparsity and low rankness to maintain global and local structures simultaneously. Experimental results on several different multiple-class hyperspectral-classification tasks demonstrate that the proposed SLGDA significantly outperforms the state-of-the-art SGDA.
Wei Li 0032, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.3
2015 Adaptive sparse representation for hyperspectral image classification
abstract
In hyerspectral remote sensing community, sparse representation based classification (SRC) is a novel concept - a testing pixel is linearly represented by labeled data, and weight coefficients are often solved by an ℓ1-norm minimization. In this work, an extension of SRC is proposed by imposing an adaptive similarity measurement between the testing pixel and labeled data on the ℓ1-norm penalty, named as adaptive SRC (ASRC). ASRC generates more discriminative sparse codes which can represent the testing pixel more robustly. Experimental results demonstrate that the proposed ASRC outperforms the traditional SRC-based classification.
Wei Li 0032, Qian Du 0001
IGARSS2
2015 Spatial preprocessing for spectral endmember extraction by local linear embedding
abstract
Endmember extraction (EE) has been widely utilized to identify spectrally unique signatures of pure ground materials in hyperspectral images. Most of existing EE algorithms focus on spectral signature only, denoted as spectral EE (sEE) algorithms in this paper. In order to improve the performance of these sEE algorithms by considering spatial information, a novel spatial preprocessing (SPP) strategy based on Locally Linear Embedding (LLE) is proposed to alleviate the influence of spectral variation. Specifically, the LLE is adopted to revise pixels by smoothing spectral variation in their spatial neighborhood. Furthermore, anomalous pixels, which may be smoothed excessively by many current SPP algorithms, can be well retained by tuning off the spatial preprocessing if their signatures are revised unexpectively. As a result, the anomalous endmembers can be correctly identified by the proposed LLE based SPP algorithm. Experimental results on simulated benchmark dataset have demonstrated that the proposed LLE based SPP algorithm outperforms many state-of-the-art SPP algorithms.
Shaohui Mei, Qian Du 0001, Mingyi He, Yihang Wang 0001
IGARSS2
2015 Hyperspectral image classification with low-rank subspace and sparse representation
abstract
Hyperspectral image classification based on low-rank representation is considered. It is often assumed that major signals occupy a low-rank subspace, and the remaining component is sparse. Due to the mixed nature of hyperspectral data, the underlying data structure may include multiple subspaces instead of a single subspace. Therefore, in this paper, we propose to use low-rank subspace representation for classification. It can improve the performance of various classifiers, including the traditional linear discriminant analysis followed by maximum likelihood classifier. The performance of using low-rank subspace representation is much better than that of low-rank representation.
Alex Sumarsono, Qian Du 0001
IGARSS2
2015 Kernel Collaborative Representation With Tikhonov Regularization for Hyperspectral Image Classification
abstract
In this letter, kernel collaborative representation with Tikhonov regularization (KCRT) is proposed for hyperspectral image classification. The original data is projected into a high-dimensional kernel space by using a nonlinear mapping function to improve the class separability. Moreover, spatial information at neighboring locations is incorporated in the kernel space. Experimental results on two hyperspectral data prove that our proposed technique outperforms the traditional support vector machines with composite kernels and other state-of-the-art classifiers, such as kernel sparse representation classifier and kernel collaborative representation classifier.
Wei Li 0032, Qian Du 0001, Mingming Xiong
IEEE Geosci. Remote. Sens. Lett.2
2015 Collaborative-Representation-Based Nearest Neighbor Classifier for Hyperspectral Imagery
abstract
Novel collaborative representation (CR)-based nearest neighbor (NN) algorithms are proposed for hyperspectral image classification. The proposed methods are based on a CR computed by an ℓ2-norm minimization with a Tikhonov regularization matrix. More specific, a testing sample is represented as a linear combination of all the training samples, and the weights for representation are estimated by an ℓ2-norm minimization-derived closed-form solution. In the first strategy, the label of a testing sample is determined by majority voting of those with k largest representation weights. In the second strategy, local within-class CR is considered as an alternative, and the testing sample is assigned to the class producing the minimum representation residual. The experimental results show that the proposed algorithms achieve better performance than several previous algorithms, such as the original k-NN classifier and the local mean-based NN classifier.
Wei Li 0032, Qian Du 0001, Fan Zhang 0007, Wei Hu 0004
IEEE Geosci. Remote. Sens. Lett.2
2015 Semisupervised Discriminant Analysis for Hyperspectral Imagery With Block-Sparse Graph
abstract
In this letter, a semisupervised block-sparse graph is proposed for discriminant analysis of hyperspectral imagery. To overcome the difficulty of not having enough training samples in the previously developed block-sparse graph approach, unlabeled samples are selected to participate in graph construction. Both sparse and collaborative representations are used for unlabeled sample selection. The experimental results demonstrate that the proposed semisupervised block-sparse graph can significantly outperform the supervised version with limited training samples. The sparse and collaborative representation-based selection methods perform comparably with the collaborative version requiring much lower computational cost.
Kun Tan 0001, Songyang Zhou, Qian Du 0001
IEEE Geosci. Remote. Sens. Lett.3
2015 Hyperspectral Image Classification Using Weighted Joint Collaborative Representation
abstract
Recently, representation-based classifiers have gained increasing interest in hyperspectral image (HSI) classification. In this letter, based on our previously developed joint collaborative representation (JCR) classifier, an improved version, which is called weighted JCR (WJCR) classifier, is proposed. JCR adopts the same weights when extracting spatial and spectral features from surrounding pixels. Differing from JCR, WJCR attempts to utilize more appropriate weights by considering the similarity between the center pixel and its surroundings. Experimental results using two real HSIs demon strate that the proposed WJCR outperforms the original JCR and some other traditional classifiers, such as the support vector machine (SVM), the SVM with a composite kernel, and simultaneous orthogonal matching pursuit.
Mingming Xiong, Qiong Ran, Wei Li 0032, Jinyi Zou, Qian Du 0001
IEEE Geosci. Remote. Sens. Lett.5
2015 Sparse Representation-Based Nearest Neighbor Classifiers for Hyperspectral Imagery
abstract
In this letter, a sparse representation-based nearest neighbor (SRNN) classifier is proposed. Unlike the traditional k-nearest neighbor (NN) classifier that employs the Euclidean distance as similarity metric, the proposed SRNN considers sparse coefficients to determine the label of testing samples, since sparse coefficients can reflect the similarity between data and provide more discriminative information. A local SRNN (LSRNN) classifier is also proposed to utilize class-specific sparse coefficients to improve the performance. Furthermore, due to the fact that neighboring pixels tend to belong to the same class with high probability, a spatially joint version of LSRNN, called JSRNN, is developed to further improve LSRNN. The proposed SRNN, LSRNN, and JSRNN have been validated on several hyperspectral remote sensing image data sets. Experimental results demonstrate that the proposed classifiers increase the classification accuracy compared with the traditional k-NN, local mean-based NN (LMNN) classifiers, and original sparse representation classifiers using representation residuals.
Jinyi Zou, Wei Li 0032, Qian Du 0001
IEEE Geosci. Remote. Sens. Lett.3
2015 Combined sparse and collaborative representation for hyperspectral target detection
Wei Li 0032, Qian Du 0001, Bing Zhang 0001
Pattern Recognit.2
2015 Local Binary Patterns and Extreme Learning Machine for Hyperspectral Imagery Classification
abstract
It is of great interest in exploiting texture information for classification of hyperspectral imagery (HSI) at high spatial resolution. In this paper, a classification paradigm to exploit rich texture information of HSI is proposed. The proposed framework employs local binary patterns (LBPs) to extract local image features, such as edges, corners, and spots. Two levels of fusion (i.e., feature-level fusion and decision-level fusion) are applied to the extracted LBP features along with global Gabor features and original spectral features, where feature-level fusion involves concatenation of multiple features before the pattern classification process while decision-level fusion performs on probability outputs of each individual classification pipeline and soft-decision fusion rule is adopted to merge results from the classifier ensemble. Moreover, the efficient extreme learning machine with a very simple structure is employed as the classifier. Experimental results on several HSI data sets demonstrate that the proposed framework is superior to some traditional alternatives.
Wei Li 0032, Chen Chen 0001, Hongjun Su, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.4
2015 Collaborative Representation for Hyperspectral Anomaly Detection
abstract
In this paper, collaborative representation is proposed for anomaly detection in hyperspectral imagery. The algorithm is directly based on the concept that each pixel in background can be approximately represented by its spatial neighborhoods, while anomalies cannot. The representation is assumed to be the linear combination of neighboring pixels, and the collaboration of representation is reinforced by l2-norm minimization of the representation weight vector. To adjust the contribution of each neighboring pixel, a distance-weighted regularization matrix is included in the optimization problem, which has a simple and closed-form solution. By imposing the sum-to-one constraint to the weight vector, the stability of the solution can be enhanced. The major advantage of the proposed algorithm is the capability of adaptively modeling the background even when anomalous pixels are involved. A kernel extension of the proposed approach is also studied. Experimental results indicate that our proposed detector may outperform the traditional detection methods such as the classic Reed-Xiaoli (RX) algorithm, the kernel RX algorithm, and the state-of-the-art robust principal component analysis based and sparse-representation-based anomaly detectors, with low computational cost.
Wei Li 0032, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.2
2015 Low-Rank Subspace Representation for Estimating the Number of Signal Subspaces in Hyperspectral Imagery
abstract
In this paper, we consider signal subspace estimation based on low-rank representation for hyperspectral imagery. It is often assumed that major signal sources occupy a low-rank subspace. Due to the mixed nature of hyperspectral remote sensing data, the underlying data structure may include multiple subspaces instead of a single subspace. Therefore, in this paper, we propose the use of low-rank subspace representation to estimate the number of subspaces in hyperspectral imagery. In particular, we develop simple estimation approaches without user-defined parameters because these parameters can be fixed as constants. Both real data experiments and computer simulations demonstrate excellent performance of the proposed approaches over those currently in the literature.
Alex Sumarsono, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.2
2014 PSO-EM: A Hyperspectral Unmixing Algorithm Based On Normal Compositional Model
abstract
A new hyperspectral unmixing algorithm is proposed based on the normal compositional model (NCM) to estimate the endmembers and abundance parameters jointly in this paper. The NCM considers the hyperspectral imaging as a stochastic process and interprets each pixel value as a random vector, which is linearly mixed by the endmembers. More precisely, these endmembers are also treated as random variables as opposed to deterministic values in order to capture spectral variability that is not well described by the linear mixing model (LMM). However, the higher complexity of such an unmixing model leads to more difficulty in parameter estimation. A particle swarm optimization-expectation maximization (PSO-EM) algorithm, a “winner-take-all” version of the EM, is proposed to solve the parameter estimation problem, which employs a partial E step. The main contribution of the proposed PSO-EM is making optimum use of particle swarm optimization method (PSO) in the partial E step, which solves the difficulty of the integrals in the NCM model. The performance of the proposed methodology is evaluated through synthetic and real data experiments. Our obtained results demonstrate the superior performance of PSO-EM compared to other NCM-based as well as LMM-based methods.
Bing Zhang 0001, Lina Zhuang, Lianru Gao, Wenfei Luo, Qiong Ran, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.6
2014 Sparse Graph-Based Discriminant Analysis for Hyperspectral Imagery
abstract
Sparsity-preserving graph construction is investigated for the dimensionality reduction of hyperspectral imagery. In particular, a sparse graph-based discriminant analysis is proposed when labeled samples are available. By forcing the projection to be along the direction where a sample is clustered with within-class samples that best represented it, the discriminative power can be enhanced. The proposed method has no requirement on the number of labeled samples as in traditional linear discriminant analysis, and it can be solved by a simple generalized eigenproblem. The quality of the dimensionality reduction is evaluated by a support vector machine with a composite spatial-spectral kernel. Experimental results demonstrate that the proposed sparse graph-based discriminant analysis can yield superior classification performance with much lower dimensionality as compared to performance on the original data or on data transformed with other dimensionality-reduction approaches.
Nam Hoai Ly, Qian Du 0001, James E. Fowler
IEEE Trans. Geosci. Remote. Sens.2
2013 Combine labeled and unlabeled information for hyperspectral image classification
abstract
In hyperspectral image classification, semisupervised learning can be applied when labeled samples are limited. By utilizing unlabeled information, classification accuracy generally can be improved. Graph-based regularization is a widely used semisupervised learning technique, where graph construction with both labeled and unlabeled samples is very computationally expensive. In reality, samples are highly correlated; so it may be unnecessary to use all the unlabeled samples. Appropriate selection of unlabeled samples can not only help improve classification but also significantly reduce the computational cost. In this paper, we propose an unlabeled sample selection algorithm. The preliminary result from a semisupervised graph-regularized kernel classifier demonstrates its effectiveness.
Qian Du 0001, Deok Han, Nicolas H. Younan
IGARSS1
2013 Unsupervised nearest regularized subspace for anomaly detection in hyperspectral imagery
abstract
A method of unsupervised nearest regularized subspace is proposed for anomaly detection in hyperspectral imagery. Based on a dual window, an approximation of each testing pixel is a representation of surrounding data via a linear combination, for which the weight vector is calculated by distance-weighted Tikhonov regularization. Proposed detector returns the similarity measurement between the testing pixel and its approximation. Experimental results for real hyperspectral data of proposed approach are demonstrated and compared to other traditional detection techniques.
Wei Li 0032, Qian Du 0001
IGARSS2
2013 Multiscale spectral-spatial classification for hyperspectral imagery
abstract
In this paper, we explore hyperspectral classification using multiscale features. To reduce data dimensionality, principal component analysis (PCA) is applied to the original image. Then a multiscale transform technique (e.g., wavelet transform, contourlet transform, etc.) is applied to each of principal components (PCs). The resulting transform coefficients can be used as spatial features. In particular, local spatial neighbors are considered to generate smoother coefficients. Combining such spatial features with spectral features (e.g., PCs), improved performance can be achieved for hyperspectral classification. In this paper, several multiscale spatial features are also evaluated.
Zhiling Long, Qian Du 0001, Nicolas H. Younan
IGARSS2
2013 Hyperspectral target detection with sparseness constraint
abstract
A sparseness constrained approach is proposed for linear unmixing, and the results are used for hybrid detection of hyperspectral imagery. The sparseness constraint is imposed on the abundance fractions, resulting in better performance than the popular non-negative and fully constrained methods, particularly in the situations when background endmember spectra are not accurately acquired or estimated, which is very common in practical applications. To increase the dictionary incoherence required for sparse regression, the use of band selection is proposed to improve the performance of sparseness constrained linear unmixing, thereby enhancing the following detection performance.
Qian Du 0001
IGARSS2
2013 A novel endmember extraction method using modified maximum spectral screening
abstract
Endmember extraction is an important task for hyperspectral analysis; the accurate identification of endmembers enables efficient spectral unmixing and classification. In the paper, a new endmember extraction algorithm based on a modified MSS approach with LP error as initial spectrum selection algorithm, and OPD measure as similarity is proposed. The endmembers extracted by modified MSS are more similar than that of MSS algorithm; from the experiments results, it has proved that our proposed method outperforms the existed MSS and N-FINDR algorithms.
Hongjun Su, Peijun Du, Qian Du 0001
IGARSS3
2013 Using High-Resolution Airborne and Satellite Imagery to Assess Crop Growth and Yield Variability for Precision Agriculture
abstract
With increased use of precision agriculture techniques, information concerning within-field crop yield variability is becoming increasingly important for effective crop management. Despite the commercial availability of yield monitors, many crop harvesters are not equipped with them. Moreover, yield monitor data can only be collected at harvest and used for after-season management. On the other hand, remote sensing imagery obtained during the growing season can be used to generate yield maps for both within-season and after-season management. This paper gives an overview on the use of airborne multispectral and hyperspectral imagery and high-resolution satellite imagery for assessing crop growth and yield variability. The methodologies for image acquisition and processing and for the integration and analysis of image and yield data are discussed. Five application examples are provided to illustrate how airborne multispectral and hyperspectral imagery and high-resolution satellite imagery have been used for mapping crop yield variability. Image processing techniques including vegetation indices, unsupervised classification, correlation and regression analysis, principal component analysis, and supervised and unsupervised linear spectral unmixing are used in these examples. Some of the advantages and limitations on the use of different types of remote sensing imagery and analysis techniques for yield mapping are also discussed.
Chenghai Yang, James H. Everitt, Qian Du 0001, Bin Luo 0005, Jocelyn Chanussot
Proc. IEEE3
2013 A Partially Supervised Approach for Detection and Classification of Buried Radioactive Metal Targets Using Electromagnetic Induction Data
abstract
The analysis of the data obtained from electromagnetic induction (EMI) sensors is one of the most viable tools for the detection of metallic objects buried under soil. The existing detection methods usually consist of sophisticated EM modeling of the source/target geometry to build suitable discriminators. The major technical challenge in this field is the reduction of false alarms with an increase of the detection probability. In this paper, we propose a partially supervised approach to detect buried radioactive targets, i.e., depleted uranium, without sophisticated EM modeling. Using the EMI data obtained by a GEM-3 sensor for a field survey, our proposed algorithm can successfully detect and discriminate the targets from nontarget metals, compared to other unsupervised and supervised approaches.
Anish C. Turlapaty, Qian Du 0001, Nicolas H. Younan
IEEE Trans. Geosci. Remote. Sens.2
2012 Hyperspectral band selection using a collaborative sparse model
abstract
In our previous research, we have proposed band-similarity-based unsupervised band selection approaches, which are proven to be very efficient. In this paper, we propose to use a collaborative sparse model for further improvement. Specifically, the pre-selected bands using the fast method, called NFINDR+LP, are further refined using a collaborative sparse model. It not only requires that the linear regression coefficients are sparse, but also requires that the same set of active bands is shared by all the bands to be removed. With the collaborative sparseness constraint being relaxed, the final selected bands can be further improved, that is, the band subset with the same number of bands can provide better classification accuracy. Based on the preliminary result, the proposed sparse model is also capable of finding the minimum number of bands to be selected.
Qian Du 0001, José M. Bioucas-Dias, Antonio Plaza
IGARSS1
2012 An operational approach for hyperspectral image compression
abstract
In lossy compression such as PCA+JPEG2000 for hyperspectral imagery, the bitrate is usually not fixed, resulting in various rate-distortion performance. In this paper, we propose an operational approach to determine the approximately optimal bitrate to be used to preserve both the majority of the information in the dataset as well as the anomalous pixels. The classification results using the reconstructed data after compression with this bitrate are comparable to those using the original data without compression; meanwhile, detection accuracy can be 100% if all the anomalies are pre-removed before compression.
Qian Du 0001, Nam Hoai Ly, James E. Fowler
IGARSS1
2012 A New Sequential Algorithm for Hyperspectral Endmember Extraction
abstract
Endmember extraction is an important step in spectral mixture analysis when endmembers are unknown. Endmembers are usually assumed to be pure pixels present in an image scene. Under this circumstance, endmember extraction is to find the most distinctive pixels. To make the searching process more efficient, the sequential forward search (SFS) method is generally used, where the next endmember is determined with a certain criterion based on the currently extracted endmember set. This letter proposes a new criterion which is related to the estimated endmember abundances. Compared to other sequential endmember extraction algorithms, the proposed method can find all the different endmembers faster. This letter also proposes to use the sequential forward floating search method as the substitute of SFS, which can improve the performance of all the sequential endmember extraction algorithms.
Qian Du 0001
IEEE Geosci. Remote. Sens. Lett.1
2012 Anomaly Detection and Reconstruction From Random Projections
abstract
Compressed-sensing methodology typically employs random projections simultaneously with signal acquisition to accomplish dimensionality reduction within a sensor device. The effect of such random projections on the preservation of anomalous data is investigated. The popular RX anomaly detector is derived for the case in which global anomalies are to be identified directly in the random-projection domain, and it is determined via both random simulation, as well as empirical observation that strongly anomalous vectors are likely to be identifiable by the projection-domain RX detector even in low-dimensional projections. Finally, a reconstruction procedure for hyperspectral imagery is developed wherein projection-domain anomaly detection is employed to partition the data set, permitting anomaly and normal pixel classes to be separately reconstructed in order to improve the representation of the anomaly pixels.
James E. Fowler, Qian Du 0001
IEEE Trans. Image Process.2
2011 Fast Band Selection for Hyperspectral Imagery
abstract
Band selection is a common technique for dimensionality reduction of hyperspectral imagery. When the desired object information is unknown, an unsupervised band selection approach is employed to select the most distinctive and informative bands. However, it may be time-consuming for unsupervised band selection methods that need to take all pixels into consideration. Here, we propose an approach to select several pixels for unsupervised band selection and the number of pixels required can be equal to the number of bands to be selected minus 1. With whitened pixel signatures (not the original pixels), band selection performance can be comparable to or even better than that from using all the pixels. For this approach, graphics processing unit (GPU)-based parallel computing is implemented for pixel selection only to further expedite the process, since computational complexity in band selection has been greatly reduced.
Qian Du 0001
ICPADS2
2011 Random-projection-based dimensionality reduction and decision fusion for hyperspectral target detection
abstract
Random projection for dimensionality reduction of hyperspectral imagery with a goal of target detection is investigated. Random projection is attractive in this task because it is data independent and computationally more efficient than other widely-used dimensionality-reduction methods, such as principal component analysis or the maximum-noise-fraction transform. Experimental results reveal that dimensionality reduction based on random projections yields improved target detection after decision fusion across multiple instances of the projections. Parallel implementation using a graphics processing unit is also investigated.
Qian Du 0001, James E. Fowler
IGARSS1
2011 2011 GRSS Data Fusion Contest: Exploiting WorldView-2 multi-angular acquisitions
abstract
The multi-angle capabilities of WorldView-2 are discussed in the framework of the GRSS Data Fusion Contest. Multi-angular surface reflectance measurements of grass, trees, and water are derived from the atmospherically corrected image sequence and compared to top of the atmosphere reflectance values. If the atmospheric effects are ignored, the top of the atmosphere values show significant spectral distortions for higher off-nadir acquisitions not resulting only from the bidirectional reflectance effects. The contest will be open for approximately another month. So far, more than 700 participants from 95 different countries have downloaded the imagery.
Fabio Pacifici, Jocelyn Chanussot, Qian Du 0001
IGARSS3
2011 Detection and classification of buried radioactive-metal objects using wideband EMI data
abstract
Gamma-ray spectroscopy is frequently used for the detection of radioactive materials. As an alternative, we explore the use of electromagnetic induction (EMI) data for detection and classification of radioactive-metal objects, i.e., depleted uranium (DU), in this study. To reduce false alarms, a pattern recognition approach based on a decision tree structure is proposed. In an initial experiment, the DU rounds were placed in rows at three different depths in a rectangular field and EMI measurements are taken. The DU objects placed up to depth 30 cm below surface were successfully detected and identified along with the depth information. The algorithm also outperformed traditional threshold detection based method in terms of discriminating objects at 30 cm depth.
Anish C. Turlapaty, Qian Du 0001, Nicolas H. Younan
IGARSS2
2011 Particle swarm optimization-based dimensionality reduction for hyperspectral image classification
abstract
We propose a particle swarm optimization (PSO)-based dimensionality reduction approach to improve support vector machine (SVM)-based classification for high-resolution hyperspectral imagery. After a searching criterion function is well designed, PSO can find a global optimal solution much more efficiently, compared to other frequently used searching strategies. In our experiments, SVM classification accuracy using PSO-selected bands is greatly higher than using all the original bands or dimensionality-reduced data from principal component analysis (PCA) or linear discriminant analysis (LDA). In addition, misclassification incurred from trivial within-class spectral variation can be further corrected by decision fusion with an unsupervised clustering, where the improvement on SVM accuracy can bring out even more significant improvement in the final fusion output.
Qian Du 0001
IGARSS2
2011 Applying spectral unmixing and support vector machine to airborne hyperspectral imagery for detecting giant reed
abstract
This study evaluated linear spectral unmixing (LSU), mixture tuned matched filtering (MTMF) and support vector machine (SVM) techniques for detecting and mapping giant reed (Arundo donax L.), an invasive weed that presents a severe threat to agroecosystems and riparian areas throughout the southern United States and northern Mexico. Airborne hyperspectral imagery with 102 usable bands covering a spectral range of 475-845 nm was collected from a giant reed-infested site along the US-Mexican portion of the Rio Grande in 2009 and 2010. The imagery was transformed with minimum noise fraction (MFN) to reduce the spectral dimensionality and noise. The three classification techniques (LSU, MTMF and SVM) were applied to the transformed MNF imagery based 11 endmember spectra extracted from the images for each of the two years. Accuracy assessment and kappa analysis were performed to compare the differences in classification accuracies among the three classification methods. Results showed that SVM and MTMF performed better than LSU, with SVM being the best classifier in both years. The results from this study indicate that hyperspectral imagery in conjunction with image classification techniques is useful for distinguishing giant reed from associated plant species and for monitoring the progression of this invasive weed.
Chenghai Yang, John A. Goolsby, James H. Everitt, Qian Du 0001
IGARSS4
2011 Semisupervised Band Clustering for Dimensionality Reduction of Hyperspectral Imagery
abstract
Band clustering is applied to dimensionality reduction of hyperspectral imagery. Different from unsupervised clustering using all the pixels or supervised clustering requiring labeled pixels, the proposed semisupervised band clustering needs class spectral signatures only. After clustering, a cluster selection step is applied to select clusters to be used in the following data analysis. Initial conditions and distance metrics are also investigated to improve the clustering performance. The experimental results show that the proposed algorithm can outperform other existing methods with lower computational cost.
Hongjun Su, Qian Du 0001, Yehua Sheng
IEEE Geosci. Remote. Sens. Lett.3
2011 An Efficient Method for Supervised Hyperspectral Band Selection
abstract
Band selection is often applied to reduce the dimensionality of hyperspectral imagery. When the desired object information is known, it can be achieved by finding the bands that contain the most object information. It is expected that these bands can provide an overall satisfactory detection and classification performance. In this letter, we propose a new supervised band-selection algorithm that uses the known class signatures only without examining the original bands or the need of class training samples. Thus, it can complete the task much faster than traditional methods that test bands or band combinations. The experimental result shows that our approach can generally yield better results than other popular supervised band-selection methods in the literature.
Qian Du 0001, Hongjun Su, Yehua Sheng
IEEE Geosci. Remote. Sens. Lett.2
2011 Multitemporal Hyperspectral Image Compression
abstract
The compression of multitemporal hyperspectral imagery is considered, wherein the encoder uses a reference image to effectuate temporal decorrelation for the coding of the current image. Both linear prediction and a spectral concatenation of images are explored to this end. Experimental results demonstrate that, when there are few changes between two images, the gain in rate-distortion performance is achieved over the independent coding of the current image. In addition, a strategy that explicitly removes salient temporal changes and stores them losslessly in the bitstream is proposed, and it is observed that this change-removal process results in a slight decrease in the rate-distortion performance with the benefit of perfect representation of the changed pixels.
Wei Zhu 0005, Qian Du 0001, James E. Fowler
IEEE Geosci. Remote. Sens. Lett.2
2011 Foreword to the Special Issue on Spectral Unmixing of Remotely Sensed Data
abstract
The 19 papers in this special issue focus on the state-of-the-art and most recent developments in the area of spectral unmixing of remotely sensed data.
Antonio Plaza, Qian Du 0001, José M. Bioucas-Dias, Xiuping Jia, Fred A. Kruse
IEEE Trans. Geosci. Remote. Sens.2
2010 A joint optical flow and principal component analysis approach for motion detection
abstract
Optical flow and its extensions have been widely used in motion detection and computer vision. In this paper, we apply principle component analysis (PCA) to analyze optical flows for better motion detection performance. The joint optical flow and PCA approach can efficiently detect moving objects and suppress small turbulence. It is effective in both static and dynamic background. It is particularly useful for motion detection from outdoor videos with low quality and small moving objects. Preliminary results demonstrate that this approach outperforms other existing methods by extracting the moving objects more completely with lower false alarms.
Kui Liu 0003, Qian Du 0001
ICASSP4
2010 On the performance of random-projection-based dimensionality reduction for endmember extraction
abstract
In this paper, we investigate the use of random-projection-based dimensionality reduction for hyperspectral endmember extraction. It is data-independent and computationally more efficient than other widely used dimensionality reduction methods, such as principal component analysis and maximum noise fraction transform. Based on the preliminary result, random-projection-based dimensionality reduction is capable of providing better endmembers after effective decision fusion.
Qian Du 0001, James E. Fowler
IGARSS1
2010 Weighted decision fusion for supervised and unsupervised hyperspectral image classification
abstract
A decision fusion approach is proposed to combine the results from supervised and unsupervised classifiers. The final output takes advantage of the power of supervised classification in class separation and the capability of unsupervised classification in reducing spectral variation impact in homogeneous regions. This approach simply adopts the majority voting rule, but can achieve the same objective of object-based classification. In this paper, we propose a weighted majority voting rule for decision fusion, where pixels in the same segment contribute differently according to their distance to the spectral centroid. The weighted majority voting rule can further improve the performance of the majority voting rule.
Qian Du 0001
IGARSS2
2010 Nonlinear Spectral Mixture Analysis for Hyperspectral Imagery in an Unknown Environment
abstract
Nonlinear spectral mixture analysis for hyperspectral imagery is investigated without prior information about the image scene. A simple but effective nonlinear mixture model is adopted, where the multiplication of each pair of endmembers results in a virtual endmember representing multiple scattering effect during pixel construction process. The analysis is followed by linear unmixing for abundance estimation. Due to a large number of nonlinear terms being added in an unknown environment, the following abundance estimation may contain some errors if most of the endmembers do not really participate in the mixture of a pixel. We take advantage of the developed endmember variable linear mixture model (EVLMM) to search the actual endmember set for each pixel, which yields more accurate abundance estimation in terms of smaller pixel reconstruction error, smaller residual counts, and more pixel abundances satisfying sum-to-one and nonnegativity constraints.
Nareenart Raksuntorn, Qian Du 0001
IEEE Geosci. Remote. Sens. Lett.2
2010 Decision Fusion on Supervised and Unsupervised Classifiers for Hyperspectral Imagery
abstract
A decision fusion approach is developed to combine the results from supervised and unsupervised classifiers. The final output takes advantage of the power of a support-vector-machine-based supervised classification in class separation and the capability of an unsupervised classifier, such as K -means clustering, in reducing trivial spectral variation impact in homogeneous regions. This approach can simply adopt the majority voting (MV) rule to achieve the same objective of object-based classification. In this letter, we propose a weighted MV (WMV) rule for decision fusion, where pixels in the same segment contribute differently according to their distance to the spectral centroid. The WMV rule can further improve the performance of the original MV rule. A series of unsupervised classifiers is investigated in the use of decision fusion, and recommendations are provided on the best unsupervised classifiers to be selected.
Qian Du 0001
IEEE Geosci. Remote. Sens. Lett.2
2010 Feature-Driven Multilayer Visualization for Remotely Sensed Hyperspectral Imagery
abstract
Displaying the abundant information contained in a remotely sensed hyperspectral image is a challenging problem. Currently, no approach can satisfactorily render the desired information at arbitrary levels of detail. In this paper, we present a feature-driven multilayer visualization technique that automatically chooses data visualization techniques based on the spatial distribution and importance of the endmembers. It can simultaneously visualize the overall material distribution, subpixel level details, and target pixels and materials. By incorporating interactive tools, different levels of detail can be presented per users' request. This scheme employs five layers from the bottom to the top: the background layer, data-driven spot layer, pie-chart layer, oriented sliver layer, and anomaly layer. The background layer provides the basic tone of the display; the data-driven spot layer manifests the overall material distribution in an image scene; the pie-chart layer presents the precise abundances of endmember materials in each pixel; the oriented sliver layer emphasizes the distribution of important anomalous materials; and the anomaly layer highlights anomaly pixels (i.e., potential targets). Displays of the airborne AVIRIS data and spaceborne Hyperion data demonstrate that the proposed multilayer visualization scheme can efficiently display more information globally and locally.
Shangshu Cai, Qian Du 0001, Robert J. Moorhead II
IEEE Trans. Geosci. Remote. Sens.2
2009 Classification Performance of Random-projection-based Dimensionality Reduction of Hyperspectral Imagery
abstract
High-dimensional data such as hyperspectral imagery is traditionally acquired in full dimensionality before being reduced in dimension prior to processing. Conventional dimensionality reduction on-board remote devices is often prohibitive due to limited computational resources; on the other hand, integrating random projections directly into signal acquisition offers alternative dimensionality reduction without sender-side computational cost. Effective receiver-side reconstruction from such random projections has been demonstrated previously using compressive-projection principal component analysis (CPPCA). While this prior work has focused on squared-error quality measures, the present work reports experimental results illustrating preservation of statistical class separation and anomaly-detection performance for CPPCA reconstruction following random-projection-based dimensionality reduction.
James E. Fowler, Qian Du 0001, Wei Zhu 0005, Nicolas H. Younan
IGARSS (5)2
2009 High Performance Computing for Hyperspectral Image Analysis: Perspective and State-of-the-art
abstract
The main purpose of this paper is to describe available (HPC)-based implementations of remotely sensed hyperspectral image processing algorithms on multi-computer clusters, heterogeneous networks of computers, and specialized hardware architectures such as field programmable gate arrays (FPGAs) and graphic processing units (GPUs). Combined, the revision of existing techniques conducted in this paper, along with the description of performance results for a parallel hyperspectral processing chain on different architectures, delivers an excellent snapshot of the state-of-the-art in the area of HPC-based hyperspectral image processing and a thoughtful perspective of the potential and emerging challenges of applying HPC paradigms to hyperspectral imaging problems.
Antonio Plaza, Qian Du 0001, Yang-Lang Chang
IGARSS (5)2
2009 Nonlinear Mixture Analysis for Hyperspectral Imagery
abstract
Nonlinear mixture analysis for hyperspectral imagery is investigated in this paper. A simple but effective nonlinear mixture model is adopted, where the multiplication of each pair of endmembers results in another ¿endmember¿, representing nonlinear scattering effect during pixel construction process. The analysis is followed by original linear demixing process. Due to the larger number of nonlinear terms being added, the resulting abundance estimation may contain some error if most of endmembers do not really participate in the mixture of a pixel. We take advantage of the developed endmember variable linear mixture model (EVLMM) to search the actual endmember set for each pixel, which yields more accurate abundance estimation.
Nareenart Raksuntorn, Qian Du 0001
IGARSS (3)2
2009 Unsupervised Hyperspectral Band Selection using Parallel Processing
abstract
Band selection is a common technique to reducing the data dimensionality of hyperspectral imagery. When the desired object information is unknown, the objective of an unsupervised band selection approach is to select the most distinctive and informative bands. Although band selection can significantly alleviate the computational burden in the following data processing and analysis, the process itself may induce additional computation complexity. In this paper, we propose parallel processing techniques for an unsupervised band selection method without changing band selection result.
Qian Du 0001
IGARSS (5)2
2009 Decision Fusion for Supervised and Unsupervised Hyperspectral Image Classification
abstract
A decision fusion approach is proposed to combine the results from supervised and unsupervised classifiers. The final output takes advantage of the power of a support vector machine based supervised classification in class separation and the capability of the unsupervised K-means classifier in reducing spectral variation impact in homogeneous regions. This approach simply adopts the majority voting rule, but can achieve the same objective of object-based classification.
Qian Du 0001
IGARSS (4)3
2009 Dependent component analysis for blind restoration of images degraded by turbulent atmosphere
Qian Du 0001, Ivica Kopriva
Neurocomputing1
2009 Segmented Principal Component Analysis for Parallel Compression of Hyperspectral Imagery
abstract
Principal component analysis (PCA) is widely used for spectral decorrelation in the JPEG2000 compression of hyperspectral imagery. However, due to the data-dependent nature of principal components, the principal component transform matrix is stored in the JPEG2000 bitstream, constituting an overhead that is often negligible if the spatial size of the image is large. However, in parallel compression in which the data set is partitioned to multiple independent processing nodes, the overhead may no longer remain negligible. It is shown that a segmented approach to PCA can greatly mitigate the detrimental effects of transform-matrix overhead and can outperform wavelet-based decorrelation which entails no such overhead.
Qian Du 0001, Wei Zhu 0005, James E. Fowler
IEEE Geosci. Remote. Sens. Lett.1
2009 An improved box-counting method for image fractal dimension estimation
Qian Du 0001, Caixin Sun
Pattern Recognit.2
2009 On the Impact of Atmospheric Correction on Lossy Compression of Multispectral and Hyperspectral Imagery
abstract
Reflectance data are often preferred to radiance data in applications of multispectral and hyperspectral imagery in which subtle spectral features are analyzed. In such applications, atmospheric correction, the process which provides radiance-to-reflectance conversion, plays a prominent role in the data-distribution and archiving pipeline. Lossy compression, often in the form of the JPEG2000 standard, will also likely factor into the distribution and archiving data flow. The relative position of data compression with respect to atmospheric correction is considered and evaluated with experimental results on both multispectral and hyperspectral imagery, and recommendations on an appropriate order for compression in the data-flow chain are made.
Qian Du 0001, James E. Fowler, Wei Zhu 0005
IEEE Trans. Geosci. Remote. Sens.1
2008 Anomaly-Based Hyperspectral Image Compression
abstract
We propose a new lossy compression algorithm for hyperspectral images, which is based on spectral principal component analysis (PCA), followed by JPEG2000 (JP2K). The approach employs an anomaly-removal model in the compression process to preserve anomalous pixels. Results on two different hyperspectral image scenes show that the new algorithm not only provides good post-compression anomaly-detection performance but also improves rate-distortion performance.
Qian Du 0001, Wei Zhu 0005, James E. Fowler
IGARSS (2)1
2008 A New Linear Mixture Model for Hyperspectral Image Analysis
abstract
In the original linear mixture model, the same set of endmembers is used for mixture analysis of an entire image. Since not all of these endmembers participate in the mixing process of each pixel, it is more reasonable to find a subset of endmembers that is actually involved in the construction of each pixel. The resulting mixture model, referred to as multiple endmember spectral mixture analysis (MESMA), has been proposed. In this paper, we develop two algorithms to determine the optimal set of endmembers for each pixel, where the sum-to-one and non-negativity constraints can be automatically relaxed. We believe these algorithms can help to improve the accuracy of linear mixture analysis of hyperspectral imagery; it is also useful to multispectral imagery to overcome the limitation due to low data dimensionality.
Nareenart Raksuntorn, Qian Du 0001
IGARSS (3)2
2008 Parallel Data Compression for Hyperspectral Imagery
abstract
The high dimensionality of hyperspectral imagery challenges image processing and analysis. It has been shown that hyperspectral compression can be achieved by principal component analysis (PCA) for spectral decorrelation followed by the JPEG2000-based coding. This approach, referred to as PCA+JPEG2000, provides superior rate-distortion performance and can preserve useful data information. However, its main disadvantage is high computational complexity in the PCA process which entails the calculation of the data covariance matrix and its eigenvectors. Parallel processing is an appropriate approach to relieve the computation burden of such a PCA-based compression. In this paper, several parallel PCA implementations are proposed and their processing speed and resulting compression performance are investigated.
Qian Du 0001, Wei Zhu 0005, Ioana Banicescu, James E. Fowler
IGARSS (2)2
2008 Improvements to 3D-Tarp Coding for the Compression of Hyperspectral Imagery
abstract
In this paper, we propose several improvements to the 3D-tarp coder for the lossy compression of hyperspectral imagery. Specific ameliorations include use of principal component analysis instead of a wavelet transform for spectral decorrelation, use of the quincunx wavelet transform instead of the traditional dyadic decomposition in the spatial direction, and spectral partitioning with skipping of insignificant zeros. Experimental results reveal that the enhanced coder achieves improved rate-distortion performance.
James E. Fowler, Qian Du 0001, Guizhong Liu
IGARSS (2)3
2008 Dimensionality Reduction and Linear Discriminant Analysis for Hyperspectral Image Classification
Qian Du 0001, Nicolas H. Younan
KES (3)1
2008 Anomaly-Based JPEG2000 Compression of Hyperspectral Imagery
abstract
Lossy compression of hyperspectral imagery is considered, with special emphasis on the preservation of anomalous pixels. In the proposed scheme, anomalous pixels are extracted before compression and replaced with interpolation from surrounding nonanomalous pixels. The image is then coded using principal component analysis for spectral decorrelation followed by JPEG2000. The anomalous pixels do not participate in this lossy compression and are rather transmitted separately in a lossless fashion. Upon decoding, the anomalous pixels are inserted back into the image. Experimental results demonstrate that the proposed scheme improves not only anomaly detection performed subsequent to decoding but also the rate-distortion performance of the lossy-compression process.
Qian Du 0001, Wei Zhu 0005, James E. Fowler
IEEE Geosci. Remote. Sens. Lett.1
2008 Automated Target Detection and Discrimination Using Constrained Kurtosis Maximization
abstract
Exploiting hyperspectral imagery without prior information is a challenge. Under this circumstance, unsupervised target detection becomes an anomaly detection problem. We propose an effective algorithm for target detection and discrimination based on the normalized fourth central moment named kurtosis, which can measure the flatness of a distribution. Small targets in hyperspectral imagery contribute to the tail of a distribution, thus making it heavier. The Gaussian distribution is completely determined by the first two order statistics and has zero kurtosis. Consequently, kurtosis measures the deviation of a distribution from the background and is suitable for anomaly/target detection. When imposing appropriate inequality constraints on the kurtosis to be maximized, the resulting constrained kurtosis maximization (CKM) algorithm will be able to quickly detect small targets with several projections. Compared to the widely used unconstrained kurtosis maximization algorithm, i.e., fast independent component analysis, the CKM algorithm may detect small targets with fewer projections and yield a slightly higher detection rate.
Qian Du 0001, Ivica Kopriva
IEEE Geosci. Remote. Sens. Lett.1
2008 Similarity-Based Unsupervised Band Selection for Hyperspectral Image Analysis
abstract
Band selection is a common approach to reduce the data dimensionality of hyperspectral imagery. It extracts several bands of importance in some sense by taking advantage of high spectral correlation. Driven by detection or classification accuracy, one would expect that, using a subset of original bands, the accuracy is unchanged or tolerably degraded, whereas computational burden is significantly relaxed. When the desired object information is known, this task can be achieved by finding the bands that contain the most information about these objects. When the desired object information is unknown, i.e., unsupervised band selection, the objective is to select the most distinctive and informative bands. It is expected that these bands can provide an overall satisfactory detection and classification performance. In this letter, we propose unsupervised band selection algorithms based on band similarity measurement. The experimental result shows that our approach can yield a better result in terms of information conservation and class separability than other widely used techniques.
Qian Du 0001
IEEE Geosci. Remote. Sens. Lett.1
2008 Color Display for Hyperspectral Imagery
abstract
This paper investigates RGB color composition schemes for hyperspectral imagery display. A three-channel composite inevitably loses a significant amount of information contained in the original high-dimensional data. The objective here is to display the useful information as distinctively as possible for high-class separability. To achieve this objective, it is important to find an effective data processing step prior to color display. A series of supervised and unsupervised data transformation and classification algorithms are reviewed, implemented, and compared for this purpose. The resulting color displays are evaluated in terms of class separability using a statistical detector and perceptual color distance. We demonstrate that the use of the data processing step can significantly improve the quality of color display, whereas data classification generally outperforms data transformation, although the implementation is more complicated. Several instructive suggestions for practitioners are provided.
Qian Du 0001, Nareenart Raksuntorn, Shangshu Cai, Robert J. Moorhead II
IEEE Trans. Geosci. Remote. Sens.1