Xu Sun 0005

dblp:37/1971-5 · DBLP profile ↗
← Back
48ranked-venue papers
1as first author
43since 2021 · last 2026
0000-0001-5389-7251ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 40 · 1 first-author · 37 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 Spectral-Spatial Enhanced Local Contrast Strategy for Hyperspectral Small Air Target Detection
abstract
Detecting small air target is an important task in civil aviation. However, the weak characteristics of these targets make detection challenging. Hyperspectral image (HSI), provides a new approach for the small air target detection task due to its strong ability of capturing both spatial and spectral information simultaneously. In this article, we propose a spectral-spatial enhanced local contrast strategy for hyperspectral small air target detection. An unsupervised band selection step based on the local contrast strategy has been designed based on local contrast (LC-UBSM) to choose bands with better distinguish ability between the target and background in HSI. Then, we have developed an improved RX detection algorithm with combined spatial and spectral variance (CSSV-RX) to detect the target while suppressing both background and noise. Experimental results on both real GAOFEN-5 dataset and simulated dataset based on EO-1 (Earth Observing-1) satellite have validated the effectiveness and robustness of the proposed method.
He Sun 0009, Lianru Gao, Haoyang Yu 0001, Lulu Qian, Xu Sun 0005
IEEE Trans. Image Process.6
2026 SMN: Signal Modulation Network for Tiny Object Detection in Remote Sensing Imagery
abstract
Tiny object detection (TOD) in remote sensing imagery remains challenging because foreground signals are extremely weak in deep feature hierarchies and are easily overwhelmed by high-response background interference. To mitigate this observed foreground-background signal modulation imbalance (FBSMI) difficulty, we propose a signal modulation network (SMN) for remote-sensing TOD. SMN comprises two complementary components. First, an adaptive Wiener filter modulator (AWFM) is inserted after backbone stages to suppress background-dominated noise while preserving weak target-related responses at multiple resolutions. Second, we introduce the novel denoising diffusion transformer (DDT), a feature-space conditional diffusion module that operates on detector feature tensors rather than image pixels. DDT generates multiple diffusion-guided semantic feature variants from high-level fused features and expands the local representation space around weak tiny object evidence. Extensive experiments on AI-TOD, SODA-A, DOTAv2.0, and DIOR-R demonstrate that SMN not only effectively mitigates the FBSMI problem, but also improves detection accuracy, particularly for very tiny and tiny objects, compared with state-of-the-art methods.
Tianwei Zhang 0005, Longfei Ren, Lianru Gao, Xu Sun 0005, Bing Zhang 0001
IEEE Trans. Image Process.4
2026 Resolution Preserving and Utilization Network for Tiny Object Detection in Large-Size Remote Sensing Imagery
abstract
Efficient tiny object detection (TOD) in large-size remote sensing imagery (LSRSI) is particularly challenging in real-world remote sensing applications. We observe that as the input size of the remote sensing scene increases, TOD faces more severe foreground signal identification issues. To address this, we are the first to design a backbone network from the perspective of low-level spatial feature preservation and utilization, specifically for tiny object feature extraction in large-size remote sensing scene patches. The proposed architecture, referred to as the resolution preserving and utilization network (RPUN), demonstrates excellent foreground tiny object feature response identification ability when increasing the input size of remote sensing scenes, effectively maintaining detection performance comparable to that of smaller input slices. Additionally, we introduce GF2UBSv2, a large-scale panchromatic satellite imagery dataset focused on tiny urban bridge detection. Extensive experiments conducted on GF2UBSv2, DIOR, SODA-A, and DOTAv2.0 demonstrate the superior performance of RPUN compared with state-of-the-art methods. The code and dataset are available at: https://github.com//Nankle.
Tianwei Zhang 0005, Longfei Ren, Xu Sun 0005, Lianru Gao, Bing Zhang 0001
IEEE Trans. Image Process.3
2025 DMSN: A Deep Multistream Network for Hyperspectral Image Super-Resolution
abstract
Hyperspectral images (HSIs) typically have finer spectral resolution but coarser spatial resolution than multispectral images (MSIs). To obtain HSIs with enhanced spatial resolution, considerable emphasis has been placed on achieving hyperspectral super-resolution (SR) by fusing HSIs with MSIs in the same scene. However, most existing HSI-MSI fusion methods either rely on prior knowledge of degradation models or require sufficient training data, hindering their practicality and interpretability. This letter proposes a deep multistream network (DMSN) for HSI SR. Specifically, we introduce the Spa-DNet and the Spe-UNet modules to encode spatial and spectral transformations across resolutions. Furthermore, we design the Int-Net to achieve spatial and spectral information interaction, enhancing the model’s performance. Finally, the proposed approach enables high spatial and spectral resolution HSIs. Using the newly designed three-stage training strategy, the network parameters can exhibit the clear physical significance of the degradation process, thereby helping to ensure faithful reconstruction of the desired HSIs. Experimental results with real datasets demonstrate that the proposed DMSN performs better than other methods. The codes will be available athttps://github.com/yuanchaosu/dmsn-GRSL.
Yuanchao Su, Xu Sun 0005, Jiaxin Li 0002, Jianjian Gao, Mengying Jiang
IEEE Geosci. Remote. Sens. Lett.3
2025 FusGAT: Graph Attention-Based Fusion Network for Unsupervised Hyperspectral Image Super-Resolution
abstract
Unsupervised hyperspectral image super-resolution (HSI-SR) has recently emerged as a popular and active research topic in remote sensing data fusion. However, most methods neglect the non-local features of the data in representation learning, which limits their fusion performances. To overcome the issue, we propose a Graph Attention-based Fusion Network (FusGAT) in this letter. This approach first extracts local features from the input data using multi-scale convolutions, and then the graph attention mechanism is employed to model relationships between nodes in the spectral stream for deriving non-local features of the image and transferring them to the spatial stream. FusGAT will iteratively update the node connections and refine node embedding, facilitating the extraction of non-local features and enabling effective information flow between the streams. We conducted several experiments on two datasets to prove the effectiveness of the proposed method. The source code will be available at: https://github.com/yuanchaosu/FusGAT-GRSL.
Yuanchao Su, Xu Sun 0005, Jiaxin Li 0002, Jianjian Gao, Ronghua Liu
IEEE Geosci. Remote. Sens. Lett.3
2025 HF-MCD: A Heterogeneous Fusion Framework for Multimodal Change Detection
abstract
Multimodal change detection (MCD) aims to detect changed areas between the bi-temporal multimodal images such as the RGB, panchromatic (PAN), multispectral (MS), and synthetic aperture radar (SAR) images, which has attracted attention in recent years. However, existing deep learning-based methods for MCD tasks still face several heterogeneity factors, the first one is the spatial resolution differences in multimodal data, which leads to the semantic gap between multimodal features. To solve this problem, we propose the heterogeneous collaborative fusion (HCF) module to integrate the multimodal features with spatial gaps. The other one is the consistency and dissimilarity between multimodal data, which lead to unequal detection contributions. To address this dilemma, we propose the heterogeneous adaptive fusion (HAF) module to fuse multimodal decision-making jointly. In this study, we proposed a heterogeneous fusion network for MCD (HF-MCD) with the HCF and the HAF module. We validate the proposed method on four public available MCD datasets. Extensive experimental results have demonstrated the superior performance of HF-MCD over the state-of-the-art methods.
Luyang Cai, He Sun 0009, Xu Sun 0005, Huanqian Yan, Lianru Gao
IEEE Trans. Geosci. Remote. Sens.3
2025 Dilated Transformation-Guided Unsupervised Multimodal Learning for Hyperspectral and Multispectral Image Fusion
abstract
Multimodal fusion widely uses convolutional layers to capture local correlations and adjust feature dimensions. However, the progressive expansion of the receptive field in convolutional layers often compromises spatial context retention, leading to the loss of fine details. Furthermore, the fixed-size kernels typically used in standard convolution restrict the network’s ability to capture multiscale contextual details. To address this limitation, this paper develops a dilated transformation-guided unsupervised multimodal learning (DTUML) method to fuse a high-resolution multispectral image (HR-MSI) and a low-resolution hyperspectral image (LR-HSI), thereby generating a high-resolution hyperspectral image (HR-HSI). Our DTUML adopts a dual-stream encoder architecture to conduct multimodal data, where one stream focuses on preserving spectral information from LR-HSIs, while the other emphasizes the acquisition of spatial details from HR-MSIs. These complementary features are subsequently integrated to ensure spectral fidelity and retain spatial detail. Then, a convolutional layer restores dimensional consistency and outputs an HR-HSI. Extensive experiments demonstrate the effectiveness of DTUML, showing superior performance and strong competitiveness compared to state-of-the-art methods. Code:https://github.com/yuanchaosu/TGRS-DTUML.
Yuanchao Su, Yicong Zhou, Lianru Gao, Mengying Jiang, Xu Sun 0005, Enke Hou
IEEE Trans. Geosci. Remote. Sens.6
2025 Similar Category Enhancement Network for Discrimination on Small Object Detection
abstract
Object Detection is a fundamental procedure in the interpretation of remote sensing images. In large-scale remote sensing images, it is common to observe that the interesting objects only occupy a small area. Such objects provide limited information gain and exhibit unclear edges, often named as small objects. The inherent characteristics of small objects significantly hinder the precise localization and accurate classification of deep object detection networks. In this paper, we introduce a significant challenge: the presence of similar objects among these small objects, which leads to dramatic misclassification and overall accuracy decrease. To assess this phenomenon, we propose a novel metric, Similar Category Angle (SCA), for classification discrimination, which serves to intuitively describe the network’s effectiveness in discriminating similar category objects in its final predictions. We also propose a one-stage object detection network named Similar Category Enhancement Network (SCENet), designed to tackle the challenges associated with discriminating similar objects in small object detection tasks. Specifically, we design SCA Loss guided by the SCA metric, which integrates SCA into the network training process, thereby enhances the network’s capability to discriminate between similar category objects. Meanwhile, we propose Laplacian Sobel Enhancement FPN, LSE-FPN, a module that incorporates dynamic edge extraction operators into the FPN to enhance the network’s ability to detect small objects by sharpening the explicit edges of objects in the feature map. Extensive experiments conducted on SODA-A, VisDrone2019 and FAIR1M-AIR datasets demonstrate the superiority of SCENet in the small object detection task, with significant improvements in detection results for both the mAP50 and SCA metrics. The code is available at https://github.com/weiziji01/SCENet.
Ziji Wei, Tianwei Zhang 0005, Xu Sun 0005, Lina Zhuang, Andrea Marinoni, Lianru Gao
IEEE Trans. Geosci. Remote. Sens.3
2025 A Hyperspectral Change Detection Method for Small Vehicles
abstract
Small vehicles (SV) detection is crucial for urban security and traffic management. However, detecting such targets from a single image presents significant challenges due to the difficulty in discerning their dynamic movements. In this paper, we propose a deep joint image-level and feature-level processing network, IFNet, designed for detecting changes in SV using bi-temporal hyperspectral images. At the image-level, a new Gumbel Softmax trick (GS)-based band selection strategy is introduced to address the problem of inconsistent spectral resolutions of bi-temporal images. At the feature-level, to tackle the challenge of capturing edge and shape details of SV, we propose a feature-based edge enhancement module, it can extract the target edge using high-level difference features, and the refined change map will be generated with the guidance of the edge map. Moreover, current deep learning-based hyperspectral change detection (HCD) methods are limited by HCD datasets. Therefore, we propose a benchmark dataset, the Hyperspectral Vehicle Change Detection (HVCD) dataset, which consists of 201 pairs of aerial hyperspectral images, each with a size of $256\times 256$ , and exhibits inconsistent spectral resolutions across the bi-temporal data. Extensive experiments conducted on the HVCD dataset demonstrate that our IFNet obtains state-of-the-art performance with an acceptable computational cost.
Shuyi Xu, He Sun 0009, Xu Sun 0005, Lianru Gao
IEEE Trans. Image Process.3
2025 SRViT: Self-Supervised Relation-Aware Vision Transformer for Hyperspectral Unmixing
abstract
Vision transformer (ViT) has recently been a popular topic in the foundation model field, taking advantage of its strong scalability and outstanding representation capabilities. As a deep model, ViT introduces a new architecture for achieving hyperspectral image (HSI) unmixing. However, traditional ViTs overlook pixel-level spatial continuity by partitioning the input image into nonoverlapping fixed-size patches. This approach disrupts local structural relationships and hinders the model's ability to capture fine-grained spatial dependencies, resulting in suboptimal feature representation for dense prediction tasks in unmixing. To address these challenges, this article proposes the development of a self-supervised relation-aware ViT (SRViT). SRViT incorporates a self-embedded module comprising encoders, a pixel-level position encoder (PLPE), a self-supervised contrastive mechanism (SCM), and a decoder. The self-embedded module and PLPE preserve local correlations in HSI across different views, facilitating cross-view learning through SCM to ensure generalization. In addition, the decoder incorporates Kronecker-factored approximate curvature (K-FAC) to capture the local geometric structure of spectral information. Ultimately, SRViT learns endmembers and fractional abundance as the unmixing result. The effectiveness and competitiveness of SRViT have been systematically validated through comparative experiments, demonstrating its superior performance. The source code is available at the following link: https://github.com/yuanchaosu/TNNLS-SRViT.
Yuanchao Su, Lianru Gao, Antonio Plaza, Xu Sun 0005, Mengying Jiang, Guang Yang 0006
IEEE Trans. Neural Networks Learn. Syst.4
2024 Primary Modality Guided Multimodal Change Detection
abstract
Multimodal images can provide richer information for a wide range of applications. However, the physical heterogeneity resulted by the difference of spatial resolution pose great challenges for multimodal change detection. To this end, we propose a change detection method called primary modality guided deep neural network (PMGN), integrating multi-resolution and multimodal data. First, we propose the principal modality and rely more on its information. Second, PMGN compensates for the limitations of low spatial resolution modalities through the primary modality guided feature exchange module. Finally, the adaptive decision fusion module enables the multimodal decision-level features to fuse efficiently. Experiments demonstrate the effectiveness and advantages of the proposed approach.
Luyang Cai, Shuyi Xu, He Sun 0009, Xu Sun 0005, Lianru Gao
IGARSS4
2024 DAMS: Dilated Attention with Multi-Stream Learning for Super-Resolution of Hyperspectral Remote Sensing Images
abstract
Hyperspectral super-resolution (SR) can effectively enhance the spatial resolution of hyperspectral images, holding significant application value. Nevertheless, existing methods tend to overlook global information and some detailed aspects of hyperspectral images, resulting in limitations in feature extraction. In addressing this issue, we propose a Dilated Attention With Multi-Stream Learning (DAMS) network to facilitate the fusion of hyperspectral and multispectral images. The network comprises three autoencoders, incorporating dilated residual multipath feature extraction for high-resolution multispectral images and a dense convolutional neural network for low-resolution hyperspectral images. Notably, no prior knowledge of point spread function (PSF) and spectral response function (SRF) is required. Experimental results with DAMS showcase its advantages over other super-resolution fusion methods, demonstrating robust performance across diverse datasets with varying PSF and SRF.
Ruoqing Xu, Yuanchao Su, Lianru Gao, Xu Sun 0005, Longfei Ren, Zhiqing Zhu
IGARSS5
2024 Information Entropy Estimation Based on Point-Set Topology for Hyperspectral Anomaly Detection
abstract
Anomaly detection is one of the most popular research topics in hyperspectral remote sensing. A variety of traditional model-driven methods fail to reveal features of data with diversity due to monotonous, fixed analytical modes. This paper analyzes mathematical-statistical properties of hyperspectral images (HSIs) and proposes an interesting approach of information entropy estimation based on point-set topology (IEEPST) to resolve anomaly detection from a brand new perspective, thus eliminating the limitations caused by the data-model discrepancy. Specifically, the original HSI data are mapped into topological spaces to enable ordered arrangements, in preparation for revealing data features. Particularly, information entropy estimation is introduced for the first time in the adoption of point-set topology to adequately unravel the data arrangements in topological spaces, whereby the land cover information is efficiently extracted for detection. Experimental results demonstrate that IEEPST accommodates both detection accuracy and computational efficiency, and is highly competitive with other sophisticated and state-of-the-art methods.
Lina Zhuang, Lianru Gao, Hongmin Gao 0001, Xu Sun 0005, Yao Liu 0012, Bing Zhang 0001
IGARSS5
2024 Adaptive Endmembers Learning-Based Deep Unmixing Network for Hyperspectral Change Detection
abstract
Hyperspectral image (HSI) change detection can detect subtle land surface change information, which is of great significance for promoting the sustainable development of human beings. Different from traditional methods, deep learningbased methods can effectively extract more discriminative features, but the problem of mixed pixels is still a challenge due to the low spatial resolution HSI. In this study, an Adaptive Endmembers Learning (AEL)-based deep unmixing network has been proposed for the change detection task, which can perform an unsupervised unmixing through adaptive endmembers learning and then obtain both the binary and multi-class change detection results. Experiments on the China dataset and the USA dataset have shown that AEL performs better than current state-of-the-art methods.
Shuyi Xu, Luyang Cai, He Sun 0009, Xu Sun 0005, Lianru Gao
IGARSS4
2024 MTSANet: Multi-Head Two-Stream Attention Networks for Unsupervised Hyperspectral Image Super-Resolution
abstract
In recent years, deep learning has been proposed for hyperspectral images(HSIs) super-resolution, and many fusion models for HSI and multispectral images(MSIs) have been developed. However, these networks are constrained to the structure of convolutional neural networks(CNNs), and more attention needs to be paid to the disadvantage of the restricted receptive field of CNNs, such that some of the distal information needs to be included in the process of acquiring features. This approach involves capturing large-scale spatial features through multi-head spatial attention and spectral features of MSI and HSI through multi-head spectral attention. Subsequently, the features are processed by convolution kernels of different scales compactly. The effectiveness and competitiveness of MTSANet are evaluated by comparing it with some state-of-the-art (SOTA) methods.
Yuanchao Su, Lianru Gao, Xu Sun 0005, Longfei Ren, Zhiqing Zhu, Mengying Jiang
IGARSS4
2024 Cross-Modal Feature Fusion and Interaction Strategy for CNN-Transformer-Based Object Detection in Visual and Infrared Remote Sensing Imagery
abstract
Due to the complementarity of visible and infrared images, it has become more favorable to fuse these two modalities to improve the object detection accuracy in the remote sensing area. However, there are still some problems to be solved. Most of the existing algorithms focus too much on the local information and ignore long-range information when performing feature extraction on different modalities. Besides, coarse weighted fusion strategies do not fully utilize the information from different modalities, and the fusion structure ignores the importance of intermodal information exchange. To tackle these problems, a cross-modal feature fusion and interaction strategy for the convolutional neural network (CNN)-transformer-based object detection in visual and infrared remote sensing imagery is proposed. We adopt a parallel structure to extract the features of different modalities, separately. In visual and infrared modality, the convolutional layers and transformer encoders are cascaded to fully extract both local and long-range information. The cross-modal feature fusion and interaction module (CFFIM) adopts the attention mechanisms to jointly fuse different modal features at the same scale to improve the diversity of fused features, and the feature interaction enables the sharing of visible and infrared information. Experiments on the VEDAI dataset have demonstrated the effectiveness of the proposed scheme compared to other state-of-the-art algorithms.
Jinyan Nie, He Sun 0009, Xu Sun 0005, Lianru Gao
IEEE Geosci. Remote. Sens. Lett.3
2024 Global Feature-Injected Blind-Spot Network for Hyperspectral Anomaly Detection
abstract
Hyperspectral anomaly detection (HAD) poses the challenge of distinguishing anomalous targets from the majority of background objects without prior knowledge. Most existing deep learning (DL) models struggle to account for both local and global spatial-spectral features in the image, limiting their performance. In this letter, we introduce PUNNet, which integrates the patch-shuffle downsampling technique and nonlinear activation-free network (NAFNet) block with dilated convolution into an advanced blind-spot network for HAD. Specifically, PUNNet utilizes the patch-shuffle downsampling operation to extend its receptive field and exploits channel attention in the NAFNet block with dilated convolution to capture global contextual information in the image. Meanwhile, PUNNet satisfies the blind-spot requirement, meaning its receptive field excludes the center pixel’s information. This allows for reliable and precise background reconstruction in a self-supervised learning paradigm, further weakening anomalous feature expression and increasing the reconstruction error of anomalies. Experimental results demonstrate that PUNNet achieves a leading position in HAD performance. The code is available athttps://github.com/DegangWang97/IEEE_GRSL_PUNNet.
Lina Zhuang, Lianru Gao, Xu Sun 0005, Xiaobin Zhao
IEEE Geosci. Remote. Sens. Lett.4
2024 A Spatial-Spectrum Fully Attention Network for Band Selection of Hyperspectral Images
abstract
Deep learning (DL)-based unsupervised band selection (UBS) methods have received more attention, but the majority of current approaches face challenges associated with striking a balance between computational burden and the UBS performance, and the spatial-spectral information has not been fully investigated. With the aim of addressing these issues, we have proposed a novel method called spatial-spectrum fully-attention network (SSFAN), which includes a spatial-spectral samples generator (SSSG) and a nearest neighbor scoring (NNS) module. Aiming to improve the UBS performance without a huge computational burden, the SSSG can directly generate numerous nonoverlapped samples for the input of DL model, where the global spatial-spectral information is utilized in a more efficient way. For the purpose of further improving the robustness of SSFAN, the NNS can assign different weights to each band by jointly exploiting the prior knowledge in both spatial and spectral domains. Note that the NNS considered the time consumption when investigating the spatial-spectral prior information, so this does not conflict with the problem of UBS balance. We have conducted experiments on three commonly used remote sensing hyperspectral image datasets, where our proposed methods have shown a more effective and robust performance than current state-of-the-art approaches. The source code will be made publicly available at https://github.com/duang33/SSFAN.
Hongmin Gao 0001, He Sun 0009, Xu Sun 0005, Bing Zhang 0001
IEEE Geosci. Remote. Sens. Lett.4
2024 Shape-Sensitive Feature Extraction for Large-Aspect-Ratio Object Detection
abstract
The detection of objects with larger aspect ratios (OLAR) is a challenging problem in a special application scenario, such as remote sensing object recognition and scene text detection. However, current object detectors perform poorly in OLAR feature extraction because they are incapable of adaptively responding to object shapes, which leads to severe misalignment between impure feature representations and region proposals. In this letter, we aim at solving this problem by proposing our shape-sensitive convolution network (SSC-Net). SSC-Net is carefully embedded with a feature enhancement module (SSC module) specifically suitable for OLAR. This module can use fewer sampling points to achieve more intelligent feature sampling area transformation, thus achieving the goal of enhancing OLAR feature representation. Extensive experiments on benchmark datasets that are rich in OLARs have proved the superiority of our method. Besides, we further verified the plug-and-play performance of the SSC module, and the experimental results show that it can significantly improve the detection performance of the detector for OLAR.
Tianwei Zhang 0005, Xu Sun 0005, Lina Zhuang, Lianru Gao, Bing Zhang 0001
IEEE Geosci. Remote. Sens. Lett.2
2024 GraphGST: Graph Generative Structure-Aware Transformer for Hyperspectral Image Classification
abstract
Transformer holds significance in deep learning (DL) research. Node embedding (NE) and positional encoding (PE) are usually two indispensable components in a Transformer. The former can excavate hidden correlations from the data, while the latter can store locational relationships between nodes. Recently, the Transformer has been applied for hyperspectral image (HSI) classification because the model can capture long-range dependencies to aggregate global features for representation learning. In an HSI, adjacent pixels tend to be homogeneous, while the NE does not identify the positional information of pixels. Therefore, PE is crucial for Transformers to understand locational relationships between pixels. However, in this area, most Transformer-based methods randomly generate PEs without considering their physical meaning, which leads to weak representations. This article proposes a new graph generative structure-aware Transformer (GraphGST) to solve the above-mentioned PE problem when implementing HSI classification. In our GraphGST, a new absolute PE (APE) is established to acquire pixels’ absolute positional sequences (APSs) and is integrated into the Transformer architecture. Moreover, a generative mechanism with self-supervised learning is developed to achieve cross-view contrastive learning (CL), aiming to enhance the representation learning of the Transformer. The proposed GraphGST model can capture local-to-global correlations, and the extracted APSs can complement the spectral features of pixels to assist in NE. Several experiments with real HSIs are conducted to evaluate the effectiveness of our GraphGST. The proposed method demonstrates very competitive performance compared with other state-of-the-art (SOTA) approaches. Our source codes will be provided in the following linkhttps://github.com/yuanchaosu/TGRS-graphGST.
Mengying Jiang, Yuanchao Su, Lianru Gao, Antonio Plaza, Xi-Le Zhao, Xu Sun 0005, Guizhong Liu
IEEE Trans. Geosci. Remote. Sens.6
2024 HADGSM: A Unified Nonconvex Framework for Hyperspectral Anomaly Detection
abstract
Hyperspectral anomaly detection aims at distinguishing targets of interest from the background without prior knowledge. Although low-rank representation (LRR)-based methods have been broadly applied in anomaly detection tasks, how to approximate the penalties in LRR-based methods more precisely is still a problem that needs to be further investigated. To this end, this article designs a unified nonconvex framework called hyperspectral anomaly detection via generalized shrinkage mappings (HADGSMs) to better approximate the LRR-based methods. The core of the proposed framework is to design new nonconvex penalties to approximate the group sparsity,$l_{0}$gradient, and low-rankness penalties in the LRR-based anomaly detection models, which can be efficiently minimized by means of generalized shrinkage mappings (GSMs). Then, an efficient alternating direction method of multipliers (ADMM) is developed to handle the proposed model. Experiments conducted on several real hyperspectral datasets demonstrate the superiority and effectiveness of the proposed framework in enhancing detection performance with respect to state-of-the-art methods.
Longfei Ren, Lianru Gao, Xu Sun 0005, Jocelyn Chanussot
IEEE Trans. Geosci. Remote. Sens.4
2024 DAAN: A Deep Autoencoder-Based Augmented Network for Blind Multilinear Hyperspectral Unmixing
abstract
In recent years, deep learning (DL) has accelerated the development of hyperspectral image (HSI) processing, expanding the range of applications further. As a typical model of unsupervised DL, the autoencoder framework has been extensively applied for spectral unmixing due to its strong representation ability and scalability. Nowadays, most DL-based unmixing approaches adopt the linear mixture model (LMM) to estimate pure spectral signatures (endmembers) and their corresponding abundance fractions. However, since sunlight scattering is an inevitable physical phenomenon, the spectral mixture problem is inherently nonlinear. Moreover, most existing nonlinear unmixing approaches focus exclusively on spectral information, neglecting the spatial distribution of materials and the intrinsic correlation between pixels, making it challenging to explore latent features. To address these issues, this article develops a new deep autoencoder-based augmented network (DAAN). The proposed DAAN employs the multilinear mixture model (MLMM) to handle the nonlinear influence caused by multiple scattering. Meanwhile, the proposed DAAN constraints homogenous smoothing in the autoencoder architecture, enabling the aggregation of intrinsic correlations by means of spatial relationships to enhance the performance of abundance estimation. We achieve unsupervised nonlinear hyperspectral unmixing by combining spectral and spatial information. The effectiveness and advantages of DAAN are confirmed by several experiments with synthetic and real HSI datasets. The results indicate that the proposed method outperforms other DL-based unmixing approaches. The source codes of the proposed DAAN will be provided in the following linkhttps://github.com/yuanchaosu/TGRS-daan.
Yuanchao Su, Zhiqing Zhu, Lianru Gao, Antonio Plaza, Pengfei Li 0010, Xu Sun 0005, Xiang Xu 0002
IEEE Trans. Geosci. Remote. Sens.6
2024 A Point-Set Topology-Based Information Entropy Estimation Method for Hyperspectral Target Detection
abstract
With hyperspectral remote sensors (imaging spectrometers) imaging a scene, the specificity of the target of interest is manifested in the significant differences between it and the surrounding background in terms of quantity, spatial distribution, and spectral characteristics, which provides conditions for the implementation of pixel-level diagnostics for target detection. Traditional model-driven methods utilize specific model assumptions to parse hyperspectral image (HSI) data in scenes with variability and are prone to encounter limitations due to model-data discrepancy. Most data-driven methods are limited in practical applications due to the great demand for training samples, the large number of parameters to be determined, and the costly computational complexity. To address the limitations of the existing methods, this article adopts point-set topology theories to analyze the properties of hyperspectral data at the mathematical-statistical level and seek a solution for the information retrieval task of target detection, whereby a target detection method through information entropy estimation based on point-set topology is proposed. First, parallel topological spaces are constructed to order the original HSI data to ensure that the differences in data features between various classes of land covers are reflected in intuitive properties in the topological spaces. Second, in conjunction with the priori information about the target, information entropy estimation is introduced to select optimal separable spaces for the target and the background by measuring the degree of ordering of data to achieve an accurate separation. Finally, a proper way to quantify and highlight the differences in data features between various land covers in the optimal separable spaces is explored for the algorithmic output to perform the information retrieval task. The proposed target detection through information entropy estimation based on point-set topology (TD-IEEPST) exploits an innovative combination of point set topology theories and information entropy estimation to achieve efficient extraction of land cover information for detection, ensuring both theoretical interpretability and computational efficiency. Extensive experimental results on real hyperspectral datasets verify that the proposed method is ahead of other widely used and state-of-the-art methods in terms of computational cost, detection effects, and robustness, and promising to provide technical support for detection response requirements in practical applications.
Lina Zhuang, Lianru Gao, Hongmin Gao 0001, Xu Sun 0005, Bing Zhang 0001
IEEE Trans. Geosci. Remote. Sens.5
2024 Information Entropy Estimation Based on Point-Set Topology for Hyperspectral Anomaly Detection
abstract
As one of the most active research hotspots in hyperspectral remote sensing, anomaly detection is widely used because it takes effect without any priori information about the target or the background. Most of the traditional model-driven methods fail to reveal features of data with diversity due to fixed analytical modes. A variety of data-driven methods encounter difficulties in practical applications due to their costly computational complexity. In this article, an innovative combination of point-set topology and information entropy theories is utilized to analyze the mathematical–statistical properties of hyperspectral images (HSIs), thus eliminating the limitations caused by the data-model discrepancy. Specifically, the original HSI data are mapped into topological spaces in a specific form to enable ordered arrangements, in preparation for revealing data features. In particular, information entropy estimation is introduced for the first time in the adoption of point-set topology to adequately unravel the data arrangements in topological spaces, whereby the land cover information is efficiently extracted for detection. Accordingly, an interesting approach of information entropy estimation based on point-set topology (IEEPST) is proposed to resolve anomaly detection from a brand new perspective, pursuing prominent detection accuracy while ensuring computational efficiency. The experimental results on benchmark HSI datasets demonstrate that IEEPST achieves detection performance with high probabilities of detection (PD) and low false alarm rates (FARs) at an inexpensive computational cost. The proposed IEEPST is highly competitive with other sophisticated and state-of-the-art methods.
Lina Zhuang, Lianru Gao, Hongmin Gao 0001, Xu Sun 0005, Bing Zhang 0001
IEEE Trans. Geosci. Remote. Sens.5
2024 Sliding Dual-Window-Inspired Reconstruction Network for Hyperspectral Anomaly Detection
abstract
Hyperspectral anomaly detection (HAD) aims to identify anomalous objects that deviate from surrounding backgrounds in an unlabeled hyperspectral image (HSI). Most available neural networks that make use of the reconstruction error to perform HAD tend to fit both backgrounds and anomalies, resulting in small reconstruction errors for both and not being effective in separating targets from background. To address this issue, we develop DirectNet, a new background reconstruction network for HAD that seamlessly integrates a sliding dual-window model into a blind-block architecture. Concretely, DirectNet establishes an inner window within the network’s receptive field by erasing the center block information, so that the content of the inner window remains invisible during the reconstruction of the central pixel. Additionally, the depth of our reconstruction network is adaptive to the size of the input image patch, ensuring that the network’s receptive field aligns with the dimensions of the input patch. The receptive field outside the inner window is considered an outer window. This weakens the impact of anomalies on the reconstruction process, causing the reconstructed pixels to converge towards the background distribution in the outer window region. Consequently, the reconstructed HSI can be regarded as a pure background HSI, leading to further amplification of reconstruction errors for anomalous targets. This enhancement improves the discriminatory ability of DirectNet. Specifically, DirectNet solely utilizes the outer window information to predict/reconstruct the central pixel. As a result, when reconstructing pixels inside anomalous targets of different sizes, the targets primarily fall within the inner window. Comprehensive experiments (conducted on four datasets) demonstrate that DirectNet achieves competitive performance compared to other state-of-the-art detectors.
Lina Zhuang, Lianru Gao, Xu Sun 0005, Xiaobin Zhao, Antonio Plaza
IEEE Trans. Geosci. Remote. Sens.4
2024 Spectral-Spatial Out-of-Distribution-Based Unsupervised Band Selection Method for Hyperspectral Anomaly Detection
abstract
Hyperspectral anomaly detection (HAD) aims to highlight the pixels that are different from the surrounding pixels without any prior information. However, as a hyperspectral image (HSI) tends to possess a huge data volume in the spectral domain, the dimension curse is inevitable in HAD. The unsupervised band selection (UBS) method is an effective tool to avoid the dimensionality curse in the HAD task. To obtain a more robust band subset without the help of any HAD detectors, we propose a spectral–spatial out-of-distribution (OOD)-based UBS method for HAD (HADUBS), which can acquire the optimal band subset in a more straightforward way. Our key observation is that the OOD term of pixels can reveal the differences and similarities of anomaly representation ability of different bands. Hence, we developed an OOD-based feature subspace representation module to obtain latent feature spaces with a better indication of the anomaly detection ability. Moreover, we introduced a UBS strategy called mutual information (MI)-based local outlier factor (MILOF) to significantly improve the discriminative ability of the selected band subset by investigating the locally sparse prior of anomalies. Extensive experimental results on five common HAD datasets demonstrate the superior performance of HADUBS. The source code will be made publicly available athttps://github.com/duang33/HADUBS.
He Sun 0009, Xu Sun 0005, Hongmin Gao 0001, Lianru Gao, Bing Zhang 0001
IEEE Trans. Geosci. Remote. Sens.3
2023 NSCKL: Normalized Spectral Clustering With Kernel-Based Learning for Semisupervised Hyperspectral Image Classification
abstract
Spatial-spectral classification (SSC) has become a trend for hyperspectral image (HSI) classification. However, most SSC methods mainly consider local information, so that some correlations may not be effectively discovered when they appear in regions that are not contiguous. Although many SSC methods can acquire spatial-contextual characteristics via spatial filtering, they lack the ability to consider correlations in non-Euclidean spaces. To address the aforementioned issues, we develop a new semisupervised HSI classification approach based on normalized spectral clustering with kernel-based learning (NSCKL), which can aggregate local-to-global correlations to achieve a distinguishable embedding to improve HSI classification performance. In this work, we propose a normalized spectral clustering (NSC) scheme that can learn new features under a manifold assumption. Specifically, we first design a kernel-based iterative filter (KIF) to establish vertices of the undirected graph, aiming to assign initial connections to the nodes associated with pixels. The NSC first gathers local correlations in the Euclidean space and then captures global correlations in the manifold. Even though homogeneous pixels are distributed in noncontiguous regions, our NSC can still aggregate correlations to generate new (clustered) features. Finally, the clustered features and a kernel-based extreme learning machine (KELM) are employed to achieve the semisupervised classification. The effectiveness of our NSCKL is evaluated by using several HSIs. When compared with other state-of-the-art (SOTA) classification approaches, our newly proposed NSCKL demonstrates very competitive performance. The codes will be available at https://github.com/yuanchaosu/TCYB-nsckl.
Yuanchao Su, Lianru Gao, Mengying Jiang, Antonio Plaza, Xu Sun 0005, Bing Zhang 0001
IEEE Trans. Cybern.5
2023 Hyperspectral Anomaly Detection Based on Chessboard Topology
abstract
Without any prior information, hyperspectral anomaly detection is devoted to locating targets of interest within a specific scene by exploiting differences in spectral characteristics between various land covers. Traditional methods originated from the signal processing perspective, and most of them rely heavily on specific model assumptions. Because of the model-driven attributes, such methods cannot mine the deep-level features of data to adapt to the variability of scenes and cannot fully extract the information of land covers contained in images to accurately separate anomalies from the background. By independently designing a chessboard-shaped topological framework that avoids making any distribution assumptions but directly mines high-dimensional data features to break through the limitations of traditional detectors, this article proposes a novel chessboard topology-based anomaly detection (CTAD) method to dissect images and extract detailed information of land covers adaptively, thereby enabling highly accurate detection. Extensive experimental results on hyperspectral images (HSIs) in real scenes demonstrate that the proposed CTAD can be adapted to the variability of scenes by autonomously learning data features and exhibiting strong generalization and detection capabilities, facilitating practical applications.
Lianru Gao, Xu Sun 0005, Lina Zhuang, Qian Du 0001, Bing Zhang 0001
IEEE Trans. Geosci. Remote. Sens.3
2023 BS3LNet: A New Blind-Spot Self-Supervised Learning Network for Hyperspectral Anomaly Detection
abstract
Recent years have witnessed the flourishing of deep learning-based methods in hyperspectral anomaly detection (HAD). However, the lack of available supervision information persists throughout. In addition, existing unsupervised learning/semisupervised learning methods to detect anomalies utilizing reconstruction errors not only generate backgrounds but also reconstruct anomalies to some extent, complicating the identification of anomalies in the original hyperspectral image (HSI). In order to train a network able to reconstruct only background pixels (instead of anomalous pixels), in this article, we propose a new blind-spot self-supervised learning network (called BS3LNet) that generates training patch pairs with blind spots from a single HSI and trains the network in self-supervised fashion. The BS3LNet tends to generate high reconstruction errors for anomalous pixels and low reconstruction errors for background pixels due to the fact that it adopts a blind-spot architecture, i.e., the receptive field of each pixel excludes the pixel itself and the network reconstructs each pixel using its neighbors. The above characterization suits the HAD task well, considering the fact that spectral signatures of anomalous targets are significantly different from those of neighboring pixels. Our network can be considered a superb background generator, which effectively enhances the semantic feature representation of the background distribution and weakens the feature expression for anomalies. Meanwhile, the differences between the original HSI and the background reconstructed by our network are used to measure the degree of the anomaly of each pixel so that anomalous pixels can be effectively separated from the background. Extensive experiments on two synthetic and three real datasets reveal that our BS3LNet is competitive with regard to other state-of-the-art approaches.
Lianru Gao, Lina Zhuang, Xu Sun 0005, Antonio Plaza
IEEE Trans. Geosci. Remote. Sens.4
2023 Hyperspectral Sparse Unmixing via Nonconvex Shrinkage Penalties
abstract
International audience
Longfei Ren, Danfeng Hong, Lianru Gao, Xu Sun 0005, Jocelyn Chanussot
IEEE Trans. Geosci. Remote. Sens.4
2023 Orthogonal Subspace Unmixing to Address Spectral Variability for Hyperspectral Image
abstract
Hyperspectral unmixing aims at estimating pure spectral signatures and their proportions in each pixel. In practice, the atmospheric effects, intrinsic variation of the spectral signatures of the materials, illumination, and topographic changes cause what is known as spectral variability resulting in significant estimation errors being propagated throughout the unmixing task. To this end, we developed a new method, called the orthogonal subspace unmixing (OSU), to address spectral variability by utilizing the orthogonal subspace projection. The proposed OSU method jointly performs orthogonal subspace learning and the unmixing process to find a more suitable subspace for unmixing. The orthogonal subspace projection encourages the representation held in the subspace to be more distinct from each other to remove the complex spectral variability in the subspace. Furthermore, an alternating minimization (AM) was designed to solve the resulting optimization problem. An efficient and convergent symmetric Gauss–Seidel alternating direction method of multipliers (sGS-ADMM), essentially a special case of the semiproximal alternating direction method of multipliers (SPADMM), was developed to solve the subproblem. Experiments conducted on one synthetic data and two real data demonstrate the effectiveness and superiority of the proposed framework in mitigating the effects of spectral variability with respect to classical linear unmixing methods or variability accounting approaches.
Longfei Ren, Danfeng Hong, Lianru Gao, Xu Sun 0005, Jocelyn Chanussot
IEEE Trans. Geosci. Remote. Sens.4
2023 ACGT-Net: Adaptive Cuckoo Refinement-Based Graph Transfer Network for Hyperspectral Image Classification
abstract
Deep learning (DL) has brought many new trends for hyperspectral image classification (HIC). Graph neural networks (GNNs) are models that fuse DL and structured data. Although GNN-based methods have focused on modeling relations, most of them are susceptible to noise, being adverse to capturing hidden correlations from data. Moreover, existing related approaches typically adopt changeless graph structures, which might lead to poor generalization. To solve the problems mentioned above, this paper develops an adaptive cuckoo refinement-based graph transfer network (ACGT-Net) that introduces a meta-heuristic optimization strategy to refine the graph structure. Specifically, we first pre-train a graph convolutional network (GCN) to learn transferable weight parameters. In the undirected graph, nodes are associated with pixels, and edges correspond to similarities between nodes. Afterward, we integrate a cuckoo search strategy (CSS) into the trained GCN to adaptively refine the graph structure. The graph structure refinement (GSR) with the CSS can pay more attention to significant channels by global optimization to improve the generalization of the GNN. Several experiments with real datasets verify the effectiveness and competitiveness of our ACGT-Net compared with other state-of-the-art (SOTA) methods.
Yuanchao Su, Jiangyi Chen, Lianru Gao, Antonio Plaza, Mengying Jiang, Xiang Xu 0002, Xu Sun 0005, Pengfei Li 0010
IEEE Trans. Geosci. Remote. Sens.7
2023 Information Retrieval With Chessboard-Shaped Topology for Hyperspectral Target Detection
abstract
Given a priori knowledge, hyperspectral target detection aims to locate objects of interest within specific scenes by utilizing differences in spectral characteristics among various land covers. However, for those traditional model-driven detectors with monotonic analytical mode, they perform mediocrely in the disassembly of hyperspectral image (HSI) data, failing to cope with real scenes with complexity. The discrepancy between fixed model assumptions and HSI data severely reduces detection effects, leading to the inability of such methods to mine deep-level features and adapt to the variability of imaging scenes. To overcome the limitations of traditional methods, we propose a chessboard-shaped topological framework for high-dimensional data structures to disassemble an HSI from both spatial and spectral dimensions adaptively. With hyperspectral target detection is refined into an information retrieval task in a topological space, a target detection method based on chessboard-shaped topology (CTTD) is proposed. In the topological space, latent and hidden data features of original images are presented in an intuitive way. Therefore, the differences in both spatial and spectral dimensions between the two classes of objects, namely target and background, are specifically amplified and exploited to perform the information retrieval task with superior performance. Extensive experimental results on benchmark HSI data sets demonstrate that CTTD can efficiently adapt to the variability of real scenes while extracting abundant and detailed information for accurate target localization. Moreover, both detection effects and computational efficiency exhibited by the proposed method provide a strong support for its popularization in practical applications.
Lina Zhuang, Lianru Gao, Hongmin Gao 0001, Xu Sun 0005, Bing Zhang 0001
IEEE Trans. Geosci. Remote. Sens.5
2023 PDBSNet: Pixel-Shuffle Downsampling Blind-Spot Reconstruction Network for Hyperspectral Anomaly Detection
Lina Zhuang, Lianru Gao, Xu Sun 0005, Antonio Plaza
IEEE Trans. Geosci. Remote. Sens.4
2023 BockNet: Blind-Block Reconstruction Network With a Guard Window for Hyperspectral Anomaly Detection
abstract
Hyperspectral anomaly detection (HAD) aims to identify anomalous targets that deviate from the surrounding background in unlabeled hyperspectral images (HSIs). Most existing deep networks that exploit reconstruction errors to detect anomalies are prone to fit anomalous pixels, thus yielding small reconstruction errors for anomalies, which is not favorable for separating targets from HSIs. In order to achieve a superior background reconstruction network for HAD purposes, this paper proposes a self-supervised blind-block network (termed BockNet) with a guard window. BockNet creates a blind-block (guard window) in the center of the network’s receptive field, rendering it unable to see the information inside the guard window when reconstructing the central pixel. This process seamlessly embeds a sliding dual-window model into our BockNet, in which the inner window is the guard window and the outer window is the receptive field outside the guard window. Naturally, BockNet utilizes only the outer window information to predict/reconstruct the central pixel of the perceptive field. During the reconstruction of pixels inside anomalous targets of varying sizes, the targets typically fall into the guard window, weakening the contribution of anomalies to the reconstruction results so that those reconstructed pixels converge to the background distribution of the outer window area. Accordingly, the reconstructed HSI can be deemed as a pure background HSI, and the reconstruction error of anomalous pixels will be further enlarged, thus improving the discrimination ability of the BockNet model for anomalies. Extensive experiments on four datasets illustrate the competitive and satisfactory performance of our BockNet compared to other state-of-the-art detectors.
Lina Zhuang, Lianru Gao, Xu Sun 0005, Antonio Plaza
IEEE Trans. Geosci. Remote. Sens.4
2023 FFN: Fountain Fusion Net for Arbitrary-Oriented Object Detection
abstract
Arbitrary-oriented object detection (AOOD) is widely used in aerial images because of its efficient object representation. However, current detectors employ the over-standardized feature extraction structure, resulting in detectors has no ability to adaptively readjust feature representations of detection units. Meanwhile, we observe that many detection units could not focus on the objects of interest in their receptive field and are easily affected by the background information and interference targets, leading to the weaking of feature expression ability. We call them sub-optimal detection units. To address this issue, we propose a novel feature enhancement module called fountain feature enhancement module (FFEM). FFEM ingeniously uses the fountain-like structure to reconstruct the features of sub-optimal detection units, generating fountain features that can automatically condense spatial regional features, which effectively enhances detectors’ overall representation ability. Then, a high-performance AOOD detector called fountain fusion net (FFN) is proposed with FFEM embedded, and many novel AOOD components are tested for their progressiveness. We validated our FFN and FFEM using three remote sensing datasets ‒ DOTA, HRSC2016, and UCAS-AOD as well as one scene text dataset‒ICDAR 2015. Extensive experiments demonstrate the effectiveness of our proposed method on improving current detectors to achieve state-of-the-art performance based on this novel idea.
Tianwei Zhang 0005, Xu Sun 0005, Lina Zhuang, Lianru Gao, Bing Zhang 0001
IEEE Trans. Geosci. Remote. Sens.2
2022 Graph-Cut-Based Node Embedding for Dimensionality Reduction and Classification of Hyperspectral Remote Sensing Images
abstract
Dimensionality reduction (DR) is a common preprocessing technology for hyperspectral images (HSIs). Recently, many neural networks can implement DR to remove the re-dundant information by node embedding. However, numer-ous hidden-layer parameters limit the generalization ability of the node embedding. In this paper, we develop a graph-cut-based node embedding (GCNE) that can be used for DR of HSIs. The embedding can refine correlations by a graph-cut strategy, and it can avoid numerous parameters when using graph models. Moreover, we combine the graph-cut strategy and extreme learning machine (ELM) to achieve HSI classi-fication. The effectiveness of the proposed method is verified by using real HSIs. Compared with other state-of-the-art DR and classification methods, the proposed approach demon-strates very competitive performance.
Yuanchao Su, Mengying Jiang, Lianru Gao, Xueer You, Xu Sun 0005, Pengfei Li 0010
IGARSS5
2022 Graph-Cut-Based Collaborative Node Embeddings for Hyperspectral Images Classification
abstract
Node embedding (NE) is conducive to aggregating correlations and relieving the influence of the Hughes phenomenon when processing high-dimensional data. Although some graph neural networks can capture correlations during achieving NE, the application of NE still faces two rigorous challenges: numerous model parameters and poor generalization. In this letter, we propose a new approach for hyperspectral image (HSI) classification, called the graph-cut-based collaborative NEs (GCCNE). Specifically, we develop a graph-cut-based NE (GCNE) to achieve low-dimensional feature representation, which avoids numerous model parameters when using a graph structure. Considering that the graph-cut in a low-dimensional space does not need to set anchors to decrease the calculation amount, we adopt an ensemble framework based on random subspaces (RSs) to implement the GCNE to obtain the collaborative feature sets, enhancing the generalization of feature representation. Afterward, the collaborative feature sets are input in several kernel-based extreme learning machines (KELMs), respectively, classifying pixels. The number of RSs is the same as the number of KELMs. Finally, we acquire an ensemble result associated with each class. The effectiveness and competitiveness of the proposed method are evaluated by using real HSI datasets.
Yuanchao Su, Mengying Jiang, Lianru Gao, Xu Sun 0005, Xueer You, Pengfei Li 0010
IEEE Geosci. Remote. Sens. Lett.4
2022 Siamese Transformer Network for Hyperspectral Image Target Detection
abstract
Hyperspectral target detection can be described as locating targets of interest within a hyperspectral image based on prior information of targets. The complexity of actual scenes limits the performance of traditional statistical methods that rely on model assumptions, and traditional machine learning methods rely on mapping functions with limited complexity. To address these problems, we propose a Siamese transformer network for hyperspectral image target detection (STTD). The contribution of this article is threefold. First, we propose a novel method of constructing training samples using only the image itself and the limited prior information, which is suitable for target detection based on the Siamese network framework. Second, the Siamese network framework is utilized to solve the problem of similarity metric learning, i.e., make homogeneous features as close as possible and heterogeneous features as far as possible. Third, the most state-of-the-art network, transformer, is applied as the backbone of our proposed Siamese network to extract global features from spectra with long-range dependencies to achieve target detection. Furthermore, we make adaptive improvements to transformer for hyperspectral images. The proposed method shows its unique advantages in suppressing the background to a low level and highlighting the target with high probability. Experiments on five different datasets demonstrate the superiority of the proposed STTD as compared to the state-of-the-art.
Weiqiang Rao, Lianru Gao, Ying Qu 0001, Xu Sun 0005, Bing Zhang 0001, Jocelyn Chanussot
IEEE Trans. Geosci. Remote. Sens.4
2022 Ensemble-Based Information Retrieval With Mass Estimation for Hyperspectral Target Detection
abstract
Given the prior information of the target, hyperspectral target detection focuses on exploiting spectral differences to separate objects of interest from the background, which can be treated as information retrieval (IR) task in machine learning (ML). Most traditional detection methods work in the original feature space and rely heavily on specific assumptions, which cannot guarantee effective extraction of features for the target and background in hyperspectral images (HSIs). Mass estimation (ME) is a base modeling mechanism that has been proven to effectively solve problems in IR and is not restricted by specific assumptions. In this article, we propose a novel target detection method through ensemble-based IR with ME (EIRME). By directly deriving the ordering from a sample set to rank data points, ME provides a simple and straightforward ranking measure to ensure that points similar to the given target are far away from dissimilar points. For the estimation of mass distribution, the proposed method utilizes a tree-structured mapping to generate a feature space, in which the separability of the target and background is further improved. In particular, to break through the technical difficulty that the direct migration of IR methods with mass measure cannot specifically meet the high-precision requirements of target detection in HSIs, we develop a specialized measurement, topological mass, which innovatively combines the mass measure with tree topology to quantify the spectral difference for detection output. Moreover, the IR with ME based on parallel measurements through ensemble trees provides a robust solution with better generalization capacity and higher precision for hyperspectral target detection, facilitating practical applications. Experimental results on benchmark HSI datasets prove that the specialized measurement that we developed successfully overcomes the drawbacks of the direct migration of IR methods with ME and exhibits unique advantages. In addition, comparisons with the most classic and advanced detection algorithms demonstrate the superiority of the proposed method.
Ying Qu 0001, Lianru Gao, Xu Sun 0005, Hairong Qi 0001, Bing Zhang 0001, Ting Shen
IEEE Trans. Geosci. Remote. Sens.4
2021 Remote Sensing Image Super-Resolution Using Novel Dense-Sampling Networks
abstract
Super-resolution (SR) techniques play a crucial role in increasing the spatial resolution of remote sensing data and overcoming the physical limitations of the spaceborne imaging systems. Though the convolutional neural network (CNN)-based methods have obtained good performance, they show limited capacity when coping with large-scale super-resolving tasks. The more complicated spatial distribution of remote sensing data further increases the difficulty in reconstruction. This article develops a dense-sampling super-resolution network (DSSR) to explore the large-scale SR reconstruction of the remote sensing imageries. Specifically, a dense-sampling mechanism, which reuses an upscaler to upsample multiple low-dimension features, is presented to make the network jointly consider multilevel priors when performing reconstruction. A wide feature attention block (WAB), which incorporates the wide activation and attention mechanism, is introduced to enhance the representation ability of the network. In addition, a chain training strategy is proposed to optimize further the performance of the large-scale models by borrowing knowledge from the pretrained small-scale models. Extensive experiments demonstrate the effectiveness of the proposed methods and show that the DSSR outperforms the state-of-the-art models in both quantitative evaluation and visual quality.
Xu Sun 0005, Xiuping Jia, Zhihong Xi, Lianru Gao, Bing Zhang 0001
IEEE Trans. Geosci. Remote. Sens.2
2021 Remote Sensing Image Super-Resolution Using Second-Order Multi-Scale Networks
abstract
Remotely sensed images, especially in urban areas, have highly complex spatial distribution, since the ground objects have diverse ranges of sizes and shapes. This largely increases the difficulty of super-resolution (SR) tasks. Current deep convolutional neural network (CNN)-based SR methods often show limited performance when coping with complicated images. This article develops a second-order multi-scale super-resolution network (SMSR) to explore reconstruction tasks for difficult cases. Specifically, we propose a single-path feature reuse which cleverly captures multi-scale feature information through aggregating the features learned at different depths of a single path. Further, we present a second-order learning mechanism, which double reuses small-difference and large-difference features at local and global levels, makes use of the learned multi-scale information at maximum. The proposed methods achieve multi-scale learning using small-size convolution only, resulting in a lightweight and high-performance SR network. Experimental results show the superiority of our SMSR over state-of-the-art methods in super-resolving complicated image patterns. The effectiveness of SMSR is also demonstrated through its support to object recognition task.
Longguang Wang, Xu Sun 0005, Xiuping Jia, Lianru Gao, Bing Zhang 0001
IEEE Trans. Geosci. Remote. Sens.3
2021 Target Detection Through Tree-Structured Encoding for Hyperspectral Images
abstract
Target detection aims to locate targets of interest within a specific scene. The traditional model-driven detectors based on signal processing have proved to be very effective. However, the detection performance of such traditional methods relies heavily on the model assumption, which is limited by the discrepancy with real hyperspectral images (HSIs) data. In this article, a target detection method through tree-structured encoding (TD-TSE) for HSIs is proposed. Instead of modeling the target and the background to extract valid features, we construct a binary tree based on the features of the data itself and segment the HSI to improve the separability of the target and the background. For the purpose of highlighting the target and suppressing the background, a novel measurement of separation, distance on tree, is calculated via binary encoding based on the constructed tree structure, and the detection output can be obtained according to such distance. To further reduce the generalization error resulting from random subsampling, the statistical average of the distances on multiple independent trees is estimated to improve the robustness of TD-TSE. The proposed method is not constrained by any model assumptions, which is fundamentally different from the most widely used hyperspectral target detectors in the field of signal processing. Moreover, the construction of binary trees without any labeled samples and the linear complexity of the proposed method make it highly practical for the hyperspectral data in real scenes. Extensive experiments on three benchmark HSI data sets demonstrate the effectiveness of the proposed TD-TSE for hyperspectral target detection.
Ying Qu 0001, Lianru Gao, Xu Sun 0005, Hairong Qi 0001, Bing Zhang 0001, Ting Shen
IEEE Trans. Geosci. Remote. Sens.4
2020 Remote Sensing Image Super-Resolution via Enhanced Back-Projection Networks
abstract
Convolutional neural network (CNN)-based image super-resolution (SR) is one of the most active field of research in the remote sensing community. As a state-of-the-art super-resolving method, however, the dense deep back-projection network (DDBPN) ignores the mutual differences among the channel-wise features and discards the initial feature when performing reconstruction. In this paper, we develop an enhanced back-projection network (EBPN) with performance exceeding the DDBPN and other state-of-the-art methods. The performance improvement gains from introducing attention mechanism to capture the feature differences among channels and reconstructing images by using the element-wise sum of the upscaled initial feature and deep features learned at different depths. A retraining strategy is also employed to further boost the SR ability of EBPN for remote sensing images. Experimental results on a remote sensing dataset and four benchmark datasets demonstrate the superiority of EBPN.
Zhihong Xi, Xu Sun 0005
IGARSS3
2018 Ship Detection Without Sea-Land Segmentation for Large-Scale High-Resolution Optical Satellite Images
abstract
Ship detection is an important and challenging topic in remote sensing applications. In current literatures, sea-land segmentation is generally requested before ship detection. This makes the implementation of the methods highly complicated. Therefore, based on Faster R-CNN, this paper proposes a ship detection method for large-scale images, which does not need sea-land segmentation as preprocessing step and can detect ships directly from complicated background including sea and land. We use large-scale images consisting of GF-1 and GF-2 satellite images to test our network. Experimental results prove that the proposed method plays a role in removing the interference of objects on land.
Yiqun He, Xu Sun 0005, Lianru Gao, Bing Zhang 0001
IGARSS2
2016 A quantitative and comparative analysis of different preprocessing implementations of DPSO: a robust endmember extraction algorithm
Lianru Gao, Lina Zhuang, Yuanfeng Wu, Xu Sun 0005, Bing Zhang 0001
Soft Comput.4
2015 An improved artificial bee colony algorithm for optimal land-use allocation
abstract
Land-use allocation is of great importance for rapid urban planning and natural resource management. This article presents an improved artificial bee colony (ABC) algorithm to solve the spatial optimization problem. The new approach consists of a heuristic information-based pseudorandom initialization (HIPI) method for initial solutions and pseudorandom search strategy based on a long-chain (LC) mechanism for neighborhood searches; together, these methods substantially improve the search efficiency and quality when handling spatial data in large areas. We evaluated the approach via a series of land-use allocation experiments and compared it with particle swarm optimization (PSO) and genetic algorithm (GA) methods. The experimental results show that the new approach outperforms the current methods in both computing efficiency and optimization quality.
Xu Sun 0005, Tianhe Chi
Int. J. Geogr. Inf. Sci.2
2009 A Study on Spectral Characteristics Extraction using Fourier Approximation Theory
abstract
In this article, based on the theory of function series approaching, we change the spectral dimension of the hyperspectral data by using the Discrete Fourier transformation, and get a new feature space which could show the shape point of the spectrum curve. The coefficient, which hyperspectral data's component in the new feature space has against the Fourier series, could tell us the effect of different spectral function to the shape of spectrum curve. The paper especially analyzes the possible effect of this feature space in image shadow recognition and precision improvement of unsupervised classification based on the Euclid distance, and verify via experiments.
Xu Sun 0005, Bing Zhang 0001, Lianru Gao
IGARSS (3)1