Hongmin Gao 0001

dblp:00/11350-1 · DBLP profile ↗
← Back
48ranked-venue papers
13as first author
43since 2021 · last 2026
0000-0002-8404-2464ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 27 · 7 first-author · 27 since 2021Artificial intelligence and machine learning · 11 · 4 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 8 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Complementary information-guided interactive fusion network for HSI and LiDAR data joint classification
Shufang Xu, Qiyuan Xue, Zhonghao Chen, Shuyu Fei, Hongmin Gao 0001
Expert Syst. Appl.5
2026 Seed-to-Semantics: Few-Shot Prototype-Guided Progressive Learning for Hyperspectral and LiDAR Classification
abstract
Deep learning-based fusion of hyperspectral images (HSI) and LiDAR has achieved strong performance in multimodal remote sensing classification, but its success is heavily constrained by the high cost of pixel-wise annotation. In extremely label-scarce regimes, such as 2-5 labeled samples per class, conventional deep models are prone to severe overfitting, while standard semi-supervised learning (SSL) methods often suffer from confirmation bias because pseudo-labels are generated from unstable early-stage representations. To address these challenges, we propose Prototype-Guided Progressive Learning (PGPL), a unified framework for few-shot HSI-LiDAR classification. Instead of relying solely on model confidence in latent space, PGPL first constructs a reliable initialization pool directly in the original data domain using spectral-angle and elevation-consistency cues, and then progressively expands the training set through class-balanced pseudo-label admission and temporal confidence stabilization. In this way, the framework improves pseudo-label reliability during both initialization and subsequent self-training. Extensive experiments on three benchmark datasets demonstrate that PGPL consistently outperforms state-of-the-art supervised and semi-supervised baselines under the corresponding 2-5-shot settings, achieving overall accuracy gains of 4.64% points on Houston, 1.16% on Trento, and 3.92% on MUUFL over the strongest competing methods, while also yielding higher pseudo-label purity. The source code will be publicly available at https://github.com/zhangyiyan001/PGPL.
Hongmin Gao 0001, Weiping Ding 0001, Pedram Ghamisi, Zhonghao Chen, Bing Zhang 0001
IEEE Trans. Image Process.2
2025 DSSNet: An Anchor-Free Rotated Object Detection Network With Dynamic Sample Selection for Remote Sensing Images
abstract
ABSTRACT Object detection in remote sensing imagery requires precise localisation and identification of targets under challenging conditions. Facing the challenges of arbitrary target orientations, wide‐scale variations, dense distributions, and small objects in remote sensing object detection, anchor‐based methods suffer from inadequate rotated target representation using rectangular boxes. This necessitates excessive angle‐specific anchors, leading to heavy computational overhead, severe sample imbalance, and slow speeds unsuitable for mobile deployment. To address these accuracy‐efficiency trade‐offs, we propose DSSNet: an anchor‐free rotated object detection network with dynamic sample selection for remote sensing images. DSSNet replaces traditional backbones with the parameter‐efficient ConvNeXt‐T and utilises an FPN for accelerated multi‐scale feature extraction. During prediction, it employs a shape‐adaptive selection strategy combined with a contour point quality assessment strategy to dynamically refine target contour points, enabling real‐time rotated object detection. The efficacy of DSSNet has been thoroughly validated through benchmark comparisons on diverse datasets. On the DOTA dataset, DSSNet clearly outperforms baseline methods in detection performance, achieving a mean Average Precision (mAP) of 76.97% and the fastest detection speed of 26.2 frames per second (FPS).
Longbao Wang, Yongheng Yu, Xiaoliang Luo, Lvchun Wang, Yican Shen, Zhijun Zhou, Hongmin Gao 0001
IET Image Process.8
2025 MKGFA: Multimodal Knowledge Graph Construction and Fact-Assisted Reasoning for VQA
abstract
Knowledge-based visual question answering relies on open-ended external knowledge and a fine-grained comprehension of both the visual content of images and semantic information. Existing methods for utilizing knowledge have the following limitations: (1) Language pre-training methods output answers in the form of plain text, which only understand shallow visual content; (2) The knowledge retrieved by image objects as labels is represented as first-order logic, making it difficult to infer complex questions. To address the above problems, this paper integrates visual-textual multimodal information, accumulates domain-specific and external multi-modal knowledge, introduces and supplements external objective facts, and proposes a multimodal knowledge graph construction and fact-assisted reasoning network (MKGFA). The network consists of three parts: the multimodal knowledge graph construction module (MKGC), the objective fact-assisted reasoning module (FAR), and the answer inference module. The MKGC engages in the coarse-to-fine-grained learning of triplet representations for multimodal knowledge units. The FAR establishes deep cross-modal relations between visual objects and factual words for correlating real answers. The answer inference module makes the final decision based on the results of both. Among them, the former two modules employ a pre-training and fine-tuning strategy, systematically accumulating foundational and domain-specific knowledge. Compared with the state-of-the-arts, MKGFA achieves 1.09% and 0.7% higher accuracy on the two challenging OKVQA and KRVQA datasets, respectively. The experimental results demonstrate the complementary advantages of the integration of the two modules.
Longbao Wang, Libing Zhang, Shufang Xu, Hongmin Gao 0001
Int. J. Comput. Intell. Appl.7
2025 MDA-HTD: Mask-driven dual autoencoders meet hyperspectral target detection
Zhonghao Chen, Hongmin Gao 0001, Zhengtao Lu, Yao Ding 0010, Xin Li 0090, Bing Zhang 0001
Inf. Process. Manag.2
2025 A Euclidean Affinity-Augmented Hyperbolic Neural Network for Semantic Segmentation of Remote Sensing Images
abstract
Semantic segmentation of remote sensing images (RSIs) plays a pivotal role in advancing geospatial analyses and applications across diverse fields, such as urban planning and environmental monitoring. Traditional learning paradigms predominantly utilize Euclidean spaces for feature extraction. This approach can introduce spatial distortions when representing objects, as Euclidean architectures typically focus on locality and are optimized for grid data, not always yielding optimal geometrical representations for data structured in non-Euclidean spaces. To address these problems, we propose EAAHNet, the first fully hyperbolic neural network designed for semantic segmentation of RSIs. EAAHNet employs the Lorentz model to reformalize conventional Euclidean-based neural network operations, ensuring the preservation of hyperbolic properties. Furthermore, to account for the inherently Euclidean nature of ground objects, we propose a Euclidean affinity-augmented hyperbolic attention module (EAAHAM) that enriches contextual dependencies through an attention fusion manner. This enhancement significantly improves the network’s capacity to discern pixel-wise semantics. Extensive experiments conducted on the ISPRS Vaihingen, ISPRS Potsdam, and LoveDA datasets demonstrate EAAHNet’s superior performance over several state-of-the-art methods. Additionally, the ablation study verifies the impacts of EAAHAM.
Xin Li 0090, Feng Xu 0008, Fan Liu 0003, Xin Lyu 0001, Hongmin Gao 0001, Jun Zhou 0011, André Kaup
IEEE Trans. Geosci. Remote. Sens.5
2025 A Frequency Decoupling Network for Semantic Segmentation of Remote Sensing Images
abstract
Semantic segmentation of remote sensing images (RSIs) is vital for numerous geospatial applications, including land-use mapping, urban planning, and environmental monitoring. Traditional neural networks for semantic segmentation primarily focus on learning in the spatial domain, which often results in suboptimal performance due to the complexity of RSIs that exhibit diverse and intricate structures. To address this problem, we propose a novel frequency decoupling network (FDNet) that enhances feature representation by independently refining high-frequency and low-frequency components in the frequency domain. FDNet introduces three core components: a sparse-aware spectral enhancement module (SSEM) that optimizes spectral feature learning by compressing redundant information while highlighting informative spectral bands, a frequency decoupling attention module (FDAM) that precisely distinguishes and enhances high-frequency and low-frequency features and an attentive frequency context module (AFCM) that integrates SSEM and FDAM into a cohesive framework for enriched spectral context modeling. Extensive experiments conducted on four benchmark datasets demonstrate that FDNet outperforms several state-of-the-art methods, achieving superior segmentation accuracy and robustness across various terrains and imaging conditions. Ablation experiments further confirm the impacts of SSEM, FDAM, and AFCM.
Xin Li 0090, Feng Xu 0008, Anzhu Yu, Xin Lyu 0001, Hongmin Gao 0001, Jun Zhou 0001
IEEE Trans. Geosci. Remote. Sens.5
2025 Dual-Feature Attention Hybrid GCN Mamba Network for Joint Hyperspectral and LiDAR Classification
abstract
Hyperspectral images (HSIs) and light detection and ranging (LiDAR) data provide complementary spectral-spatial and elevation information, respectively, whose fusion can significantly improve classification accuracy. However, their inherent heterogeneity challenges effective spectral-geospatial integration. Although convolutional neural networks (CNNs), recurrent neural networks (RNNs), and Transformer models have advanced multimodal remote sensing classification, each shows distinct limitations. CNNs excel in spatial feature aggregation but lack global context, whereas RNNs and Transformers, despite capturing long-range spectral features, face issues such as computational inefficiency. To address these limitations, we propose a dual-feature attention hybrid graph convolutional network (GCN) Mamba network (DAHGMN) for joint HSI and LiDAR classification. Specifically, multimodal image cubes are first extracted by a CNN to obtain initial features. Subsequently, a dual-feature attention (DA) module is introduced to adaptively recalibrate spectral and spatial feature weights, enhancing discriminability. Furthermore, we propose a hybrid GCN Mamba (HGM) module with both low parameter complexity and time complexity, which combines the local geometric modeling capability of GCNs with the global long-range dependency modeling of Mamba’s state-space model (SSM). A probability-based decision fusion strategy is employed to integrate multi-level classification results, achieving an efficient combination of the spatial-spectral contextual features. Extensive experiments on three benchmark HSI-LiDAR datasets demonstrate that DAHGMN achieves superior classification accuracy while significantly reducing parameter complexity compared to state-of-the-art methods. The implementation code is publicly available at https://github.com/RogsXie/DAHGMN.
Zhenyang Xie, Hongmin Gao 0001, Shufang Xu, Haihua Xie
IEEE Trans. Geosci. Remote. Sens.3
2025 Multiscale Segmentation-Guided Fusion Network for Hyperspectral Image Classification
abstract
Convolution Neural Networks (CNNs) have demonstrated strong feature extraction capabilities in Euclidean spaces, achieving remarkable success in hyperspectral image (HSI) classification tasks. Meanwhile, Graph convolution networks (GCNs) effectively capture spatial-contextual characteristics by leveraging correlations in non-Euclidean spaces, uncovering hidden relationships to enhance the performance of HSI classification (HSIC). Methods combining GCNs with CNNs have achieved excellent results. However, existing GCN methods primarily rely on single-scale graph structures, limiting their ability to extract features across different spatial ranges. To address this issue, this paper proposes a multiscale segmentation-guided fusion network (MS2FN) for HSIC. This method constructs pixel-level graph structures based on multiscale segmentation data, enabling the GCN to extract features across various spatial ranges. Moreover, effectively utilizing features extracted from different spatial scales is crucial for improving classification performance. This paper adopts distinct processing strategies for different feature types to enhance feature representation. Comparative experiments demonstrate that the proposed method outperforms several state-of-the-art (SOTA) approaches in accuracy. The source code will be released at https://github.com/shengrunhua/MS2FN.
Hongmin Gao 0001, Runhua Sheng, Yuanchao Su, Zhonghao Chen, Shufang Xu, Lianru Gao
IEEE Trans. Image Process.1
2025 DDFformer: a dual-domain fused transformer for polyp segmentation
Xi Yong, Jingchen Liang, Yun Hu 0004, Xin Li 0090, Hongmin Gao 0001, Zuojian Zhou, Kongfa Hu
J. Supercomput.7
2024 Information Entropy Estimation Based on Point-Set Topology for Hyperspectral Anomaly Detection
abstract
Anomaly detection is one of the most popular research topics in hyperspectral remote sensing. A variety of traditional model-driven methods fail to reveal features of data with diversity due to monotonous, fixed analytical modes. This paper analyzes mathematical-statistical properties of hyperspectral images (HSIs) and proposes an interesting approach of information entropy estimation based on point-set topology (IEEPST) to resolve anomaly detection from a brand new perspective, thus eliminating the limitations caused by the data-model discrepancy. Specifically, the original HSI data are mapped into topological spaces to enable ordered arrangements, in preparation for revealing data features. Particularly, information entropy estimation is introduced for the first time in the adoption of point-set topology to adequately unravel the data arrangements in topological spaces, whereby the land cover information is efficiently extracted for detection. Experimental results demonstrate that IEEPST accommodates both detection accuracy and computational efficiency, and is highly competitive with other sophisticated and state-of-the-art methods.
Lina Zhuang, Lianru Gao, Hongmin Gao 0001, Xu Sun 0005, Yao Liu 0012, Bing Zhang 0001
IGARSS4
2024 A cross-modal feature aggregation and enhancement network for hyperspectral and LiDAR joint classification
Hongmin Gao 0001, Jun Zhou 0001, Pedram Ghamisi, Shufang Xu, Bing Zhang 0001
Expert Syst. Appl.2
2024 A dual-branch siamese spatial-spectral transformer attention network for Hyperspectral Image Change Detection
Shufang Xu, Hongmin Gao 0001
Expert Syst. Appl.5
2024 CSFFNet: Lightweight cross-scale feature fusion network for salient object detection in remote sensing images
abstract
Abstract Salient object detection (SOD), one of the most important applications in the field of computer vision, aims to extract the most visually appealing regions of scenes. However, the improvement of the accuracy of existing salient object detection in optical remote sensing images (ORSI‐SOD) is usually accompanied by an increase of network complexity, which affects the application of these models. Motivated by this, a novel lightweight edge‐supervised neural network for ORSI‐SOD is proposed, named CSFFNet. Specifically, the backbone (ResNet34) is first lightened by feature encoding module (FEM), building a lightweight subnet for feature extraction. Then, in the transformer‐based feature pyramid enhancement module (FPEM), the convolutional features obtained in the FEM are enhanced by long‐distance dependence to obtain multi‐scale features containing rich saliency cues. Based on this, the feature fusion module (FFM) is designed to capture cross‐scale long‐range dependencies and effectively fuse high‐level semantic information with low‐level detail information. Thus, the increase in network complexity due to multi‐level decoding is avoided. Finally, the segmentation results are optimized by using salient edges as auxiliary information, which effectively improves the contrast and completeness of the results. Experimental results on two public datasets demonstrate that the lightweight CSFFNet achieves competitive or even better performance compared with state‐of‐the‐art methods.
Longbao Wang, Chong Long, Xin Li 0090, Xiaodan Tang, Zhipeng Bai, Hongmin Gao 0001
IET Image Process.6
2024 Rotated points for object detection in remote sensing images
abstract
Abstract Object detection in remote sensing images poses great challenges due to the dense distribution, arbitrary orientation, and aspect ratio variations of objects. Most of the existing methods rely on aligned convolutional features, which fail to capture the geometric information of objects effectively and result in the inconsistency between the classification score and localization accuracy. Moreover, densely packed objects suffer from spatial feature aliasing caused by the intersection of reception fields between objects. To address this issue, a deformable convolution‐based method named rotated points is proposed, which consists of two modules: a point set loss module and a high‐quality sample assignment module. The point set loss module can extract geometric features of objects in arbitrary directions with fine‐grained point sets for feature representation and introduce outlier penalties to penalize outlier points. The high‐quality sample assignment module measures the classification and localization ability, orientation quality, and point‐wise correlation of point sets comprehensively to enhance the consistency of classification and regression significantly. Experiments on the DOTA and FAIR1M datasets demonstrate that the proposed method achieves significant improvements over the benchmark model.
Longbao Wang, Yican Shen, Hongmin Gao 0001
IET Image Process.5
2024 A CBAM-GAN-based method for super-resolution reconstruction of remote sensing image
abstract
Abstract As satellite imagery technology advances, remote sensing plays an increasingly prominent role in modern society. Nevertheless, the limitations of existing imaging sensors and complex atmospheric conditions constrain the quality of raw remote sensing data, posing challenges for interpretation and noise reduction. Super‐resolution technology focuses on enhancing low‐quality, low‐resolution remote sensing images. In this study, we introduce a method that utilizes a high‐order degradation model to generate low‐resolution remote sensing images. We employ a Generative Adversarial Network with a Convolutional Block Attention Module (CBAM‐GAN) to enhance these images, reducing noise interference and improving texture and feature display. Our approach outperforms other methods on the UCMerced‐LandUse, WHU‐RS19, and AID datasets. Specifically, it raises SSIM index scores to 0.9443, 0.8928, and 0.8633, respectively, exceeding baselines by 1.31%, 0.19%, and 1.30%. The MOS index also improves to 3.98, 3.96, and 3.83, respectively, representing a 2.31%, 8.20%, and 2.96% gain over the baseline. Our reconstruction produces superior results, demonstrating the effectiveness of our proposed method.
Longbao Wang, Xin Li 0090, Hongmin Gao 0001
IET Image Process.6
2024 A Cross-Domain Coupling Network for Semantic Segmentation of Remote Sensing Images
abstract
Semantic segmentation of remote sensing images (RSIs) is critical for various applications, including urban planning, agriculture, and disaster management. Existing methods often fail to capture fine-grained textures and periodic patterns in RSIs, leading to suboptimal results in complex terrains. To address these challenges, we propose a cross-domain coupling network (CDCNet) that leverages both domain-specific extraction and cross-domain coupling (CDC) to enrich contextual cues for semantic inference. Our CDCNet integrates a CDC layer within the encoder-decoder architecture to simultaneously refine representations in the frequency and spatial domains. This approach effectively models fine-grained textures and periodic patterns in the frequency domain, as well as edges, shapes, and broad structural elements in the spatial domain. Extensive experiments on the ISPRS Potsdam and LoveDA datasets demonstrate the superiority of CDCNet over several state-of-the-art methods. Ablation studies confirm the significant impact of the CDC layer, validating the effectiveness of our approach in handling RSIs.
Xin Li 0090, Feng Xu 0008, Feifei Tao, Hongmin Gao 0001, Fan Liu 0003, Xin Lyu 0001
IEEE Geosci. Remote. Sens. Lett.5
2024 A Spatial-Spectrum Fully Attention Network for Band Selection of Hyperspectral Images
abstract
Deep learning (DL)-based unsupervised band selection (UBS) methods have received more attention, but the majority of current approaches face challenges associated with striking a balance between computational burden and the UBS performance, and the spatial-spectral information has not been fully investigated. With the aim of addressing these issues, we have proposed a novel method called spatial-spectrum fully-attention network (SSFAN), which includes a spatial-spectral samples generator (SSSG) and a nearest neighbor scoring (NNS) module. Aiming to improve the UBS performance without a huge computational burden, the SSSG can directly generate numerous nonoverlapped samples for the input of DL model, where the global spatial-spectral information is utilized in a more efficient way. For the purpose of further improving the robustness of SSFAN, the NNS can assign different weights to each band by jointly exploiting the prior knowledge in both spatial and spectral domains. Note that the NNS considered the time consumption when investigating the spatial-spectral prior information, so this does not conflict with the problem of UBS balance. We have conducted experiments on three commonly used remote sensing hyperspectral image datasets, where our proposed methods have shown a more effective and robust performance than current state-of-the-art approaches. The source code will be made publicly available at https://github.com/duang33/SSFAN.
Hongmin Gao 0001, He Sun 0009, Xu Sun 0005, Bing Zhang 0001
IEEE Geosci. Remote. Sens. Lett.2
2024 A Frequency Domain Feature-Guided Network for Semantic Segmentation of Remote Sensing Images
abstract
Semantic segmentation of Remote Sensing Images (RSIs) entails assigning semantic labels to each pixel accurately. RSIs are rich in spatial and spectral data, revealing diverse material and object characteristics. Yet, current RSI-focused computer vision models struggle with significant intra-class variation and inter-class resemblance due to limited spectral data usage. We propose the Frequency Domain Feature-Guided Network (FFGNet) for RSI semantic segmentation, influenced by digital signal processing theories. FFGNet initially generates frequency domain features via patch partitioning and 2D discrete cosine transformation. Our Frequency Enhancement Attention module (FEA) then distinguishes and intensifies frequency components to retain detailed information. These enhanced features are integrated with the Spatial-Spectral Attention (SSA) for enriched spectral signals. In the inference phase, these features are upsampled and combined with decoded features, emphasizing spectral details. Additionally, our novel loss function combines frequency and cross-entropy losses. Experiments on LoveDA and ISPRS Potsdam datasets demonstrate FFGNet's effectiveness, surpassing other mainstream models. An ablation study further validates our dual-guidance design.
Xin Li 0090, Feng Xu 0008, Hongmin Gao 0001, Fan Liu 0003, Xin Lyu 0001
IEEE Signal Process. Lett.3
2024 TL2GH²T: Triple-Path Local-to-Global Network With Hybrid Head Transformer for Hyperspectral Change Detection
abstract
With the aid of transformers, significant progress has been achieved in hyperspectral image change detection (HSI-CD) in recent times. Nonetheless, most contemporary detection methods fail to incorporate diverse diagnostic features extracted from hyperspectral (HS) images. In addition, relying solely on algebraic-based techniques to extract information of difference is insufficient for achieving satisfactory detection performance. In this regard, we propose an innovative triple-path local-to-global network (TL2GN), complemented by a hybrid head transformer (HybridHT), called TL2GH2T, tailored for HSI-CD tasks. To be specific, TL2GH2T first investigates spatial, spectral, and spatial–spectral features from a local-to-global perspective. Then, a novel spatial and spectral token fusion (SSTF) module is developed to integrate the above three tokenized features, producing discriminative features from two HS images separately. Moreover, drawing inspiration from chromosomal crossover mechanisms, we propose a HybridHT. Its goal is to simultaneously learn cross correlation and self-correlation information of bitemporal features from a global perspective, producing highly discriminative distinctions. Our approach, validated through extensive experimentation on four varied HS benchmarks, exhibits exceptional performance in HSI-CD, outperforming contemporary methods in both visual and quantitative evaluations.
Zhonghao Chen, Swalpa Kumar Roy, Hongmin Gao 0001, Yao Ding 0010, Xiongwu Xiao, Bing Zhang 0001
IEEE Trans. Geosci. Remote. Sens.4
2024 Interactive Enhanced Network Based on Multihead Self-Attention and Graph Convolution for Classification of Hyperspectral and LiDAR Data
abstract
The fusion of multimodal data plays a crucial role in classification tasks. However, existing research typically mines and analyzes the individual features of each data source separately before considering how to fuse them. In contrast, our approach first constructs interactive enhanced fusion features (IEFFs) for initial fusion while considering the extraction of individual features and, finally, integrates them effectively to utilize the information from each data source more comprehensively. To this end, we propose a novel interactive enhanced network based on multihead self-attention (MSA) and graph convolution. Specifically, we extract individual features from hyperspectral image (HSI) and light detection and ranging (LiDAR) data and then construct IEFFs based on the row and column features of the central pixel. Individual features focus on the local characteristics of a single data source, while IEFFs strengthen the feature expression of the central pixel through matrix operations, integrating the complementary information of multimodal data. Subsequently, we use graph convolutional networks (GCNs) to construct graph structures for four types of features (interactive enhanced HSI features, interactive enhanced LiDAR features, HSI individual features, and LiDAR individual features), modeling the pixels as nodes and capturing spatial relationships. On this basis, we apply an MSA mechanism to mine spectral dependencies, further extracting global spectral features. Finally, we design a multimodal gated fusion module (MGFM) that effectively integrates these features through its weighting mechanism. The weight allocation is adjusted dynamically according to the characteristics of the feature, achieving optimal fusion of multimodal data. Extensive experiments on three popular HSI and LiDAR datasets verify the superior performance of our method. Our code will be available athttps://github.com/haofeng0003/MSA-GCN.
Hongmin Gao 0001, Shuyu Fei, Runhua Sheng, Shufang Xu, Bing Zhang 0001
IEEE Trans. Geosci. Remote. Sens.1
2024 Multiscale Random-Shape Convolution and Adaptive Graph Convolution Fusion Network for Hyperspectral Image Classification
abstract
Convolution neural networks (CNNs) are extensively utilized in hyperspectral image (HSI) classification due to their remarkable capability to extract features from patterns with fixed shapes. These networks have been shown to effectively capture features at the pixel level. However, the fixed shape of convolution kernels poses a challenge for CNNs to adapt to the diverse shapes found in HSIs. Graph neural networks (GNNs), particularly graph convolution networks (GCNs), possess robust feature extraction capabilities on graph structures and are extensively applied in HSI classification. However, one significant challenge in using GNNs is the selection of appropriate neighboring nodes for information aggregation. To address the existing challenges of GCN and CNN and leverage their respective advantages, this paper introduces a novel patch-based CNN-GCN fusion classification network, named multi-scale random-shape convolution and adaptive graph convolution fusion network (MRCAGCFN). It consists of a spectral transformation module and three main modules we proposed: a multi-scale random-shape convolution module for extracting convolution features, where the shape of the convolution kernel is randomized and a multi-scale approach is applied to enhance adaptability to data with diverse shapes; an adaptive feature-fusion graph convolution module for extracting graph convolution features, where the weights for neighborhood aggregation are learned adaptively to reduce feature fusion from dissimilar nodes and strengthen feature fusion from similar nodes; and an adaptive local feature processing module for processing features, where two different methods are employed to convert patch-level features to pixel-level features, thereby improving feature representation. MRCAGCFN combines the strengths of CNN and GCN while introducing enhancements to better accommodate diverse feature shapes. Experimental results on three HSI classification datasets demonstrate that our proposed MRCAGCFN outperforms some existing methods. The codes of our MRCAGCFN will be available at https://github.com/shengrunhua/MRCAGCFN.
Hongmin Gao 0001, Runhua Sheng, Zhonghao Chen, Haiyun Liu, Shufang Xu, Bing Zhang 0001
IEEE Trans. Geosci. Remote. Sens.1
2024 A Point-Set Topology-Based Information Entropy Estimation Method for Hyperspectral Target Detection
abstract
With hyperspectral remote sensors (imaging spectrometers) imaging a scene, the specificity of the target of interest is manifested in the significant differences between it and the surrounding background in terms of quantity, spatial distribution, and spectral characteristics, which provides conditions for the implementation of pixel-level diagnostics for target detection. Traditional model-driven methods utilize specific model assumptions to parse hyperspectral image (HSI) data in scenes with variability and are prone to encounter limitations due to model-data discrepancy. Most data-driven methods are limited in practical applications due to the great demand for training samples, the large number of parameters to be determined, and the costly computational complexity. To address the limitations of the existing methods, this article adopts point-set topology theories to analyze the properties of hyperspectral data at the mathematical-statistical level and seek a solution for the information retrieval task of target detection, whereby a target detection method through information entropy estimation based on point-set topology is proposed. First, parallel topological spaces are constructed to order the original HSI data to ensure that the differences in data features between various classes of land covers are reflected in intuitive properties in the topological spaces. Second, in conjunction with the priori information about the target, information entropy estimation is introduced to select optimal separable spaces for the target and the background by measuring the degree of ordering of data to achieve an accurate separation. Finally, a proper way to quantify and highlight the differences in data features between various land covers in the optimal separable spaces is explored for the algorithmic output to perform the information retrieval task. The proposed target detection through information entropy estimation based on point-set topology (TD-IEEPST) exploits an innovative combination of point set topology theories and information entropy estimation to achieve efficient extraction of land cover information for detection, ensuring both theoretical interpretability and computational efficiency. Extensive experimental results on real hyperspectral datasets verify that the proposed method is ahead of other widely used and state-of-the-art methods in terms of computational cost, detection effects, and robustness, and promising to provide technical support for detection response requirements in practical applications.
Lina Zhuang, Lianru Gao, Hongmin Gao 0001, Xu Sun 0005, Bing Zhang 0001
IEEE Trans. Geosci. Remote. Sens.4
2024 Information Entropy Estimation Based on Point-Set Topology for Hyperspectral Anomaly Detection
abstract
As one of the most active research hotspots in hyperspectral remote sensing, anomaly detection is widely used because it takes effect without any priori information about the target or the background. Most of the traditional model-driven methods fail to reveal features of data with diversity due to fixed analytical modes. A variety of data-driven methods encounter difficulties in practical applications due to their costly computational complexity. In this article, an innovative combination of point-set topology and information entropy theories is utilized to analyze the mathematical–statistical properties of hyperspectral images (HSIs), thus eliminating the limitations caused by the data-model discrepancy. Specifically, the original HSI data are mapped into topological spaces in a specific form to enable ordered arrangements, in preparation for revealing data features. In particular, information entropy estimation is introduced for the first time in the adoption of point-set topology to adequately unravel the data arrangements in topological spaces, whereby the land cover information is efficiently extracted for detection. Accordingly, an interesting approach of information entropy estimation based on point-set topology (IEEPST) is proposed to resolve anomaly detection from a brand new perspective, pursuing prominent detection accuracy while ensuring computational efficiency. The experimental results on benchmark HSI datasets demonstrate that IEEPST achieves detection performance with high probabilities of detection (PD) and low false alarm rates (FARs) at an inexpensive computational cost. The proposed IEEPST is highly competitive with other sophisticated and state-of-the-art methods.
Lina Zhuang, Lianru Gao, Hongmin Gao 0001, Xu Sun 0005, Bing Zhang 0001
IEEE Trans. Geosci. Remote. Sens.4
2024 Airborne Small Target Detection Method Based on Multimodal and Adaptive Feature Fusion
abstract
The detection of airborne small targets amidst cluttered environments poses significant challenges. Factors such as the susceptibility of a single RGB image to interference from the environment in target detection and the difficulty of retaining small target information in detection necessitate the development of a new method to improve the accuracy and robustness of airborne small target detection. This article proposes a novel approach to achieve this goal by fusing RGB and infrared (IR) images, which is based on the existing fusion strategy with the addition of an attention mechanism. The proposed method employs the YOLO-SA network, which integrates a YOLO model optimized for the downsampling step with an enhanced image set. The fusion strategy employs an early fusion method to retain as much target information as possible for small target detection. To refine the feature extraction process, we introduce the self-adaptive characteristic aggregation fusion (SACAF) module, leveraging spatial and channel attention mechanisms synergistically to focus on crucial feature information. Adaptive weighting ensures effective enhancement of valid features while suppressing irrelevant ones. Experimental results indicate 1.8% and 3.5% improvements in mean average precision (mAP) over the LRAF-Net model and Infusion-Net detection network, respectively. Additionally, ablation studies validate the efficacy of the proposed algorithm’s network structure.
Shufang Xu, Tianci Liu 0007, Zhonghao Chen, Hongmin Gao 0001
IEEE Trans. Geosci. Remote. Sens.6
2024 Dual-Feature Attention-Based Contrastive Prototypical Clustering for Multimodal Remote Sensing Data
abstract
The integrated use of multisource remote sensing (RS) data in Earth observation missions has garnered considerable attention. Hyperspectral images (HSIs) offer extensive spatial and spectral detail, whereas light detection and ranging (LiDAR) data provide elevation information. Therefore, the fusion of HSI and LiDAR data can enhance the accuracy (ACC) of image classification. However, contemporary supervised multimodal deep learning techniques depend heavily on extensive human-annotated training datasets. To address this challenge, we propose a contrastive prototypical clustering network enhanced with a dual-feature attention module. Specifically, two sets of enhanced modal views are constructed from the multimodal RS images for the subsequent contrastive learning. The proposed dual-feature attention module emphasizes channel and spatial attention separately for each modality, integrating both to adjust the feature representation across different channels and positions. By learning the importance weights of each channel and position, this module highlights the hierarchical structure and enhances the discriminative quality of the features. The learned features are utilized through an online clustering mechanism and a self-supervised training strategy that combines contrastive loss and cluster loss to achieve efficient and effective land cover classification. Extensive experiments on three widely used HSI and LiDAR datasets demonstrate that the proposed method outperforms current state-of-the-art approaches. The code for this method is openly available at:https://github.com/RogsDing/DFCPC.
Shufang Xu, Xinchen Ding, Zhen Zhang 0019, Hongmin Gao 0001, Bing Zhang 0001
IEEE Trans. Geosci. Remote. Sens.5
2024 Cognitive Fusion of Graph Neural Network and Convolutional Neural Network for Enhanced Hyperspectral Target Detection
abstract
In recent years, deep learning has emerged as a prominent technique in hyperspectral target detection (HTD). Extensive research has highlighted the potential of Graph Neural Network (GNN) as a promising framework for exploring non-Euclidean dependencies within hyperspectral imagery. However, GNN has not been introduced to HTD. Additionally, achieving a balanced training set while effectively suppressing background remains a challenge. Therefore, we propose the cognitive fusion of GNN and Convolutional Neural Network (CNN) for enhanced HTD (named as CFGC), which marks the first integration of GNN and CNN in HTD. Initially, using sparse subspace clustering and a similarity measurement strategy, we select the most representative background samples for HTD. Subsequently, linear interpolation combines the prior target with the Laplacian-weighted prior target, yielding abundant targets with meaningful transformations. Finally, a fused network of CNN and GNN is utilized for training both the prior target and the constructed training set. Significantly, the incorporation of attention mechanism in both the CNN and GNN branches stands out as a noteworthy advantage, augmenting the models’ ability to selectively prioritize crucial information. Four benchmark hyperspectral images have been used in extensive experiments, and the results demonstrate that CFGC exhibits superior performance in HTD.
Shufang Xu, Sijie Geng, Zhonghao Chen, Hongmin Gao 0001
IEEE Trans. Geosci. Remote. Sens.5
2024 Strengthened Residual Graph and Multiscale Gated Guided Convolutional Fusion Network for Hyperspectral Change Detection
abstract
Hyperspectral image (HSI) change detection (CD) focuses on identifying changes in the internal components of land cover and land use. Convolutional neural networks (CNNs) have made significant progress in HSI-CD. Concurrently, graph convolutional networks (GCNs) have gained considerable attention for their ability to utilize unlabeled data and explicitly exploit correlations between adjacent parcels. However, CNNs are constrained by fixed, small-size convolutional kernels, which severely limit their receptive field. On the other hand, GCNs use superpixels to reduce the number of nodes, which will lead to losing pixel-level features, resulting in partial feature representations from both networks. To leverage the strengths of both CNNs and GCNs, a model was proposed that incorporates two subnetworks: decomposed multiscale gated guided CNNs and strengthened residual graph convolution. The decomposed multiscale gated guided CNNs are designed to capture pixel-level features at various scales using different kernel sizes. A gated change information fusion (GCF) unit integrates these multiscale pixel-level features. Meanwhile, the strengthened residual graph convolution was used to aggregate change information, which can prevent node information from becoming homogeneous. Additionally, a feature fusion module (FFM) is employed to combine features from the two subnetworks. The proposed model effectively utilizes both multiscale convolution and graph features, facilitating the learning of multilevel contextual semantic features. The experimental results on three HSI datasets demonstrate that this model outperforms several state-of-the-art methods. The code is available athttps://github.com/zhangyiyan001/srgmgn.
Shufang Xu, Xiangfei Xia, Runhua Sheng, Hongmin Gao 0001, Bing Zhang 0001
IEEE Trans. Geosci. Remote. Sens.6
2024 Spectral-Spatial Out-of-Distribution-Based Unsupervised Band Selection Method for Hyperspectral Anomaly Detection
abstract
Hyperspectral anomaly detection (HAD) aims to highlight the pixels that are different from the surrounding pixels without any prior information. However, as a hyperspectral image (HSI) tends to possess a huge data volume in the spectral domain, the dimension curse is inevitable in HAD. The unsupervised band selection (UBS) method is an effective tool to avoid the dimensionality curse in the HAD task. To obtain a more robust band subset without the help of any HAD detectors, we propose a spectral–spatial out-of-distribution (OOD)-based UBS method for HAD (HADUBS), which can acquire the optimal band subset in a more straightforward way. Our key observation is that the OOD term of pixels can reveal the differences and similarities of anomaly representation ability of different bands. Hence, we developed an OOD-based feature subspace representation module to obtain latent feature spaces with a better indication of the anomaly detection ability. Moreover, we introduced a UBS strategy called mutual information (MI)-based local outlier factor (MILOF) to significantly improve the discriminative ability of the selected band subset by investigating the locally sparse prior of anomalies. Extensive experimental results on five common HAD datasets demonstrate the superior performance of HADUBS. The source code will be made publicly available athttps://github.com/duang33/HADUBS.
He Sun 0009, Xu Sun 0005, Hongmin Gao 0001, Lianru Gao, Bing Zhang 0001
IEEE Trans. Geosci. Remote. Sens.4
2023 Local aggregation and global attention network for hyperspectral image classification with spectral-induced aligned superpixel segmentation
Zhonghao Chen, Guoyong Wu, Hongmin Gao 0001, Yao Ding 0010, Danfeng Hong, Bing Zhang 0001
Expert Syst. Appl.3
2023 Grid Network: Feature Extraction in Anisotropic Perspective for Hyperspectral Image Classification
abstract
Abundant spectral signatures and spatial characteristics embedded in hyperspectral (HS) images enable the fine identification of land covers, attracting plenty of studies on feature extraction and feature utilization. Nevertheless, the high representative spectral and spatial features in the HS cube are unevenly distributed, which is failed to consider by many current methods. To conquer this shortcoming, we rethink the feature extraction of HS images from an anisotropic perspective and propose a novel model called grid network (GNet) for HS image classification. Beyond representing spectral-spatial features in three classic paradigms (simultaneously, hierarchically, and separately), GNet is capable of learning them in two new processes: multi-stage and multi-path. In this way, spectral and spatial features can be fully and balanced explored. More significantly, to make full use of low- and high-level features and avoid the existing semantic gap, we devise a spectral-spatial cross-level feature fusion module to model the relation between them. Extensive experiments, implemented on three HS datasets, demonstrate that the proposed GNet enables to acquire promising classification performance compared to state-of-the-art methods. The codes of this work will be available at https://github.com/zhonghaochen/GNet_Master for the sake of reproducibility.
Zhonghao Chen, Danfeng Hong, Hongmin Gao 0001
IEEE Geosci. Remote. Sens. Lett.3
2023 Adaptively Dictionary Construction for Hyperspectral Target Detection
abstract
The task of hyperspectral images (HSIs) target detection is to identify whether the target spectral sequences present in the HSI. Recently, the topic of representation models has received much interest in hyperspectral target detection. The performance of representation models depends on whether the corresponding dictionary and sparse matrix can correctly recover the original spectrum. Therefore, the background dictionary of these models should contain the spectra of all classes except the target spectrum; i.e., the dictionary should be overcomplete. However, most representation models cannot satisfy this condition. Moreover, due to the potentially large spectral similarity between the target and the background, representation models perform poorly in background suppression. Aiming to solve these issues, a novel adaptively dictionary construction (ADC) strategy with background suppression sparse representation (BSSR) module is proposed in this letter, called adaptively dictionary construction for target detection (ADCTD). Specifically, the proposed ADC is adopted to segment the HSI into superpixels consisting of pixels with similar spectra. This process can be considered as an unsupervised coarse classification process, which can construct an overcomplete background dictionary. In addition, the BSSR is adopted to improve the separation of the target and background by a linear function. Experiments on three datasets demonstrate the superiority of the proposed ADCTD.
Weibo Zhang, Zhonghao Chen, Hongmin Gao 0001
IEEE Geosci. Remote. Sens. Lett.5
2023 Depthwise Separable Convolutional Autoencoders for Hyperspectral Image Change Detection
abstract
Hyperspectral image change detection (HSI-CD) has recently become a research hotspot. Current methods rely heavily on a huge amount of training samples to perform the change detection tasks. While acquiring data from the same region of bi-temporal HSIs is extraordinarily time-consuming and laborious. Therefore, this letter proposes an unsupervised method based on three dimensional (3D) depthwise separable convolutional autoencoders (DSConvAE). First, the dual-branch symmetrical 3D DSConvAE is pre-trained with limited samples to obtain the optimal weights, which facilitates extracting discriminative spatial and spectral features subsequently. Second, we adopt the temporal-specific feature concatenation strategy to acquire comprehensive characteristics from bi-temporal HSIs. Third, the general autoencoders are employed at the end of the model to further explore the high-level and abstract feature vectors. Finally, we compare the mean square loss calculated from the spatial-spectral branches and apply threshold judgement to generate the ultimate detection maps. Experimental results on three public HSI datasets demonstrate that the proposed framework outperforms other comparative methods by significant improvements.
Yongfeng Zhou, Shufang Xu, Danfeng Hong, Hongmin Gao 0001, Qiqiang Zhong, Bing Zhang 0001
IEEE Geosci. Remote. Sens. Lett.5
2023 A high-level feature channel attention UNet network for cholangiocarcinoma segmentation from microscopy hyperspectral images
Hongmin Gao 0001, Xueying Cao, Peipei Xu
Mach. Vis. Appl.1
2023 A Multidepth and Multibranch Network for Hyperspectral Target Detection Based on Band Selection
abstract
Deep learning (DL) has recently risen to prominence in hyperspectral target detection (HTD). Nevertheless, how to tackle the extreme training sample imbalance together with achieving target highlighting and background suppression is challenging. Additionally, due to the spectral redundancy of hyperspectral imagery (HSI), it is a new course for HTD through band selection (BS) to retain crucial bands thereupon improving the subsequent detection performance. Accordingly, we propose a DL-based BS-HTD (DLBSTD) algorithm, incorporating DL-based BS with DL-based HTD for the first time. Most significantly, a multi-depth and multi-branch network (MDBN) for HTD based on a novel BS method is proposed. First of all, the BS method including an alternating local-global reconstruction network (ALGRN) and a correlation measurement strategy provides representative bands containing key target information for MDBN. For the training sample imbalance of MDBN, we develop a BS-based method to select multifarious representative background training samples and propose a target band random substitution (TBRS) strategy to augment an ample target training set. Lastly, the MDBN composed of a multi-depth feature extraction (MDFE) module, three fusion strategies, and the parallel local convolution and gated recurrent unit (Conv-GRU) fully taps the spectral feature relationships to highlight targets and suppress backgrounds. Compared with nine competitive HTD algorithms, we carry out plentiful experiments on four classical datasets exhibiting that the proposed DLBSTD has strong generalization and salient detection performance of target highlighting and background suppression.
Hongmin Gao 0001, Zhonghao Chen, Shufang Xu, Danfeng Hong, Bing Zhang 0001
IEEE Trans. Geosci. Remote. Sens.1
2023 AMSSE-Net: Adaptive Multiscale Spatial-Spectral Enhancement Network for Classification of Hyperspectral and LiDAR Data
abstract
With the abundant emergence of remote sensing data sources, multimodal remote sensing observation has become an active field. Extracting valuable information from multi-modal data has the potential to make a significant contribution to applications such as urban planning and monitoring. However, existing studies are deficient in extracting spectral and spatial features from hyperspectral remote sensing data. Meanwhile, the method of fusing multimodal features has limitations and poses a challenge to the convergence of the model loss function, which increases the complexity of the network model optimisation process. Therefore, this paper proposes an Adaptive Multi-scale Spatial–Spectral Enhancement Network for Classification of Hyperspectral and LiDAR Data called AMSSE-Net. First, we perform deep mining of spectral features in hyperspectral images by the involution operator. The main idea is to take full advantage of the involution operator in characterising spectral features by using the property that the convolution kernel shares the feature channels within the group. Furthermore, the multi-branching approach is used to extract the multi-scale information, and then the spectral-spatial features are formed with the strategy of hierarchical fusion. Meanwhile, we employ three-layer convolution for extracting shallow features from LiDAR data, offering supplementary information. Finally, we propose the ”Adaptive Feature Fusion Module,” an innovative and comprehensive mechanism designed for the fusion of features from diverse sources in multi-source data fusion. These dynamically assigned weights guide the selection of the optimal model, which is determined by the joint loss across the three methods, ultimately leading to the generation of an accurate prediction map. This approach not only helps to deeply explore the spectral spatial information in the hyperspectral data, but also effectively fuses the hyperspectral information with the elevation information from the LiDAR data. The expression ability of model features is rapidly improved by adaptive weighting, which in turn enhances the performance and generalisation ability of the model. Compared with some existing methods, extensive experiments on three popular HSI and LiDAR datasets show that our proposed AMSSE-Net can achieve better classification performance. The codes will be available at https://github.com/haofeng0003/AMSSE-Net, contributing to the RS community.
Hongmin Gao 0001, Shufang Xu, Bing Zhang 0001
IEEE Trans. Geosci. Remote. Sens.1
2023 Hyperspectral Target Detection via Spectral Aggregation and Separation Network With Target Band Random Mask
abstract
Hyperspectral target detection (HTD) is a pixel-wise detection method based on limited prior targets and spectral differences, which has been widely studied and applied in many fields. Recently, deep learning (DL) plays an important role in hyperspectral imagery (HSI) processing. However, for HTD, the severe lack of class-balanced training sets is an enormous challenge. Meanwhile, it is difficult to suppress backgrounds while highlighting targets through the deep network. To address these issues, we propose a spectral aggregation and separation network (SASN) with a target band random mask (TBRM) for HTD in this paper. For the training sets of SASN, a multifarious representative background selection strategy (MRBS) is first proposed to obtain a multifarious and representative background training set. Next, aiming at the notorious class imbalance, a data augmentation (DA) method, TBRM, is proposed to generate adequate target training set by repeating randomly zero-masking the spectral bands of a prior target. Subsequently, in the training of SASN, residual connection and squeeze-and-excitation (SE) channel attention mechanism are applied to fully extract high discriminative features and nonlinear ones in the spectra. Besides, to better separate the targets and backgrounds, a triplet-soft loss function is presented, which makes the training in the direction of spectral separation of background samples from both the prior target and target samples. During testing, the trained SASN distinguishes the spectral similarities and differences simultaneously for highlighting targets and suppressing backgrounds. Moreover, extensive experimental results validate that the proposed method has superior detection performances, background suppression capacity, and separability compared with ten cutting-edge HTD algorithms on six benchmark HSI datasets.
Hongmin Gao 0001, Zhonghao Chen, Feng Xu 0008, Danfeng Hong, Bing Zhang 0001
IEEE Trans. Geosci. Remote. Sens.1
2023 Information Retrieval With Chessboard-Shaped Topology for Hyperspectral Target Detection
abstract
Given a priori knowledge, hyperspectral target detection aims to locate objects of interest within specific scenes by utilizing differences in spectral characteristics among various land covers. However, for those traditional model-driven detectors with monotonic analytical mode, they perform mediocrely in the disassembly of hyperspectral image (HSI) data, failing to cope with real scenes with complexity. The discrepancy between fixed model assumptions and HSI data severely reduces detection effects, leading to the inability of such methods to mine deep-level features and adapt to the variability of imaging scenes. To overcome the limitations of traditional methods, we propose a chessboard-shaped topological framework for high-dimensional data structures to disassemble an HSI from both spatial and spectral dimensions adaptively. With hyperspectral target detection is refined into an information retrieval task in a topological space, a target detection method based on chessboard-shaped topology (CTTD) is proposed. In the topological space, latent and hidden data features of original images are presented in an intuitive way. Therefore, the differences in both spatial and spectral dimensions between the two classes of objects, namely target and background, are specifically amplified and exploited to perform the information retrieval task with superior performance. Extensive experimental results on benchmark HSI data sets demonstrate that CTTD can efficiently adapt to the variability of real scenes while extracting abundant and detailed information for accurate target localization. Moreover, both detection effects and computational efficiency exhibited by the proposed method provide a strong support for its popularization in practical applications.
Lina Zhuang, Lianru Gao, Hongmin Gao 0001, Xu Sun 0005, Bing Zhang 0001
IEEE Trans. Geosci. Remote. Sens.4
2023 Multimodal Transformer Network for Hyperspectral and LiDAR Classification
abstract
The land cover classification of single-modal remote sensing (RS) data has recently reached a bottleneck. The joint use of multi-modal RS data to improve classification performances has received much attention. Convolutional Neural Networks are powerful tools in feature extraction and contextual modeling. While they have attendant drawbacks to capture the sequence attributes of spectral signatures and struggle to acquire discriminative spectral-spatial features from a global perspective due to limitations inherent in their network backbones. The transformer backbone is a promising approach for addressing these challenges and generating novel insights in multi-modal RS image classification. In this article, we present a new model called Multi-modal Transformer Network (MTNet) that leverages transformer advantages to capture both the specific and shared characteristics of hyperspectral (HS) and light detection and ranging (LiDAR) data. HS images contain a wide range of bands with rich spectral information and LiDAR data provide accurate elevation information without affecting by environmental factors. The well-designed module Hyperspectral Spectral Transformer can learn spectrally local sequence information from neighbouring bands of HS images, yielding group-wise spectral embeddings comprising rich diagnostic information about land covers. Furthermore, the HS and LiDAR spatial transformers aim to mine the pixel-wise feature embedding relationships in a global manner, capturing spatial and elevation information of HS and LiDAR, respectively. Finally, the feature embedding tokens of two modalities are integrated jointly and a new transformer encoder is redesigned to explore the shared spatial characteristics between the two modalities. We evaluate the classification performances of the proposed MTNet on three public HS-LiDAR datasets by conducting extensive experiments, exhibiting superiority over conventional classifiers and state-of-the-art networks.
Shufang Xu, Danfeng Hong, Hongmin Gao 0001, Meiqiao Bi
IEEE Trans. Geosci. Remote. Sens.4
2022 Multiscale spectral-spatial cross-extraction network for hyperspectral image classification
abstract
Abstract Convolutional neural networks (CNN) are becoming increasingly popular in modern remote sensing image classification tasks and have exhibited excellent results. For the existing CNN‐based hyperspectral image (HSI) classification methods, most of which extract spatial or spectral features separately by convolution. But nearly all of these methods ignore the fact that the weighted summation of convolution may lead to appear new features in another dimension. To address this issue, a novel multiscale spectral‐spatial cross‐extraction network (MSSCEN) is proposed for HSI classification. Specifically, the proposed MSSCEN introduces spectral‐spatial features cross extraction module (SSCEM), which fed extracted features from previous layer into spatial and spectral extraction branches separately again, so that the changes that occurred in the other domain after each convolution can be fully utilized. In addition, a new independent data augmentation module based on U‐Net is designed to mitigate the problem of limited labelled samples. The paper conducts experiments on three classic hyperspectral datasets and the results demonstrate that the proposed method achieves the best classification accuracy than other state‐of‐the‐art methods.
Hongmin Gao 0001, Hongyi Wu, Zhonghao Chen
IET Image Process.1
2022 Shallow Network Based on Depthwise Overparameterized Convolution for Hyperspectral Image Classification
abstract
Recently, convolutional neural network (CNN) techniques have gained popularity as a tool for hyperspectral image classification (HSIC). To improve the feature extraction efficiency of HSIC under the condition of limited samples, the current methods generally use deep models with plenty of layers. However, deep network models are prone to overfitting and gradient vanishing problems when samples are limited. In addition, the spatial resolution decreases severely with deeper depth, which is very detrimental to spatial edge feature extraction. Therefore, this letter proposes a shallow model for HSIC, which is called a depthwise overparameterized convolutional neural network (DOCNN). To ensure the effective extraction of the shallow model, the depthwise overparameterized convolution (DO-Conv) kernel is introduced to extract the discriminative features. The DO-Conv kernel is composed of a standard convolution kernel and a depthwise convolution kernel, which can extract the spatial feature of the different channels individually and fuse the spatial features of the whole channels simultaneously. Moreover, to further reduce the loss of spatial edge features due to the convolution operation, a dense residual connection (DRC) structure is proposed to apply to the feature extraction part of the whole network. Experimental results obtained from three benchmark datasets show that the proposed method outperforms other state-of-the-art methods in terms of classification accuracy and computational efficiency.
Hongmin Gao 0001, Zhonghao Chen
IEEE Geosci. Remote. Sens. Lett.1
2022 Global to Local: A Hierarchical Detection Algorithm for Hyperspectral Image Target Detection
abstract
Hyperspectral image (HSI) has received considerable attention in the field of target detection due to its powerful ability to capture the spectral information of land covers, and plenty of detection algorithms have been explored. However, these methods generally leverage the difference between the spectrum of the target to be detected and the background spectrum to accomplish target detection, and so are susceptible to the problem of spectral variability. In this article, we propose a global-to-local hierarchical detection algorithm for HSI (G2LHTD). Firstly, extended morphological attribute profile (EMAP) is first used to model global spatial texture information from HSI. Subsequently, a diverse-direction constrained energy minimization (D2CEM) detector is developed to consider the spatial information within eight neighborhoods around each pixel in HSI, yielding comprehensive local spatial information. More substantially, to effectively discriminate the neighborhood information in diverse directions, we devise an adaptive neighborhood feature aggregation (ANFA) strategy, which will comprehensively evaluate the significance of neighborhood information in diverse directions. As a result, the spatial features of HSI can be comprehensively considered for hyperspectral target detection (HTD). Extensive experiments, conducted on four standard datasets, demonstrate the effectiveness of the proposed method. The codes of this work will be available at https://github.com/zhonghaocheng/G2LHTD_Master for the sake of reproducibility.
Zhonghao Chen, Zhengtao Lu, Hongmin Gao 0001, Jia Zhao 0001, Danfeng Hong, Bing Zhang 0001
IEEE Trans. Geosci. Remote. Sens.3
2021 Multiscale Residual Network With Mixed Depthwise Convolution for Hyperspectral Image Classification
abstract
Convolutional neural networks (CNNs) are becoming increasingly popular in modern remote sensing image processing tasks and exhibit outstanding capability for hyperspectral image (HSI) classification. However, for the existing CNN-based HSI-classification methods, most of them only consider single-scale feature extraction, which may neglect some important fine information and cannot guarantee to capture optimal spatial features. Moreover, many state-of-the-art methods have a huge number of network parameters needed to be tuned, which will cause high computational cost. To address the aforementioned two issues, a novel multiscale residual network (MSRN) is proposed for HSI classification. Specifically, the proposed MSRN introduces depthwise separable convolution (DSC) and replaces the ordinary depthwise convolution in DSC with mixed depthwise convolution (MDConv), which mixes up multiple kernel sizes in a single depthwise convolution operation. The DSC with mixed depthwise convolution (MDSConv) can not only explore features at different scales from each feature map but also greatly reduce learnable parameters in the network. In addition, a multiscale residual block (MRB) is designed by replacing the convolutional layer in an ordinary residual block with the MDSConv layer. The MRB is used as the major unit of the proposed MSRN. Furthermore, to enhance further the feature representation ability, the proposed network adds a high-level shortcut connection (HSC) on the cascaded two MRBs to aggregate lower level features and higher level features. Experimental results on three benchmark HSIs demonstrate the superiority of the proposed MSRN method over several state-of-the-art methods.
Hongmin Gao 0001, Lianru Gao, Bing Zhang 0001
IEEE Trans. Geosci. Remote. Sens.1
2020 Underwater salient object detection by combining 2D and 3D visual features
Zhe Chen 0004, Hongmin Gao 0001, Zhen Zhang 0019, Helen Zhou, Xun Wang 0007
Neurocomputing2
2020 Spatial-temporal multi-task learning for salient region detection
Zhe Chen 0004, Ruili Wang 0001, Ming Yu 0001, Hongmin Gao 0001
Pattern Recognit. Lett.4
2019 Multi-branch fusion network for hyperspectral image classification
Hongmin Gao 0001, Sheng Lei, Hui Zhou 0003, Xiaoyu Qu
Knowl. Based Syst.1
2019 Convolutional neural network for spectral-spatial classification of hyperspectral images
Hongmin Gao 0001, Jia Zhao 0001
Neural Comput. Appl.1
2019 Application of Hyperspectral Image Classification Based on Overlap Pooling
Hongmin Gao 0001, Shuo Lin
Neural Process. Lett.1