Mengmeng Zhang 0005

dblp:47/1678-5 · DBLP profile ↗
← Back
62ranked-venue papers
8as first author
55since 2021 · last 2026
0000-0002-5724-9785ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 38 · 5 first-author · 34 since 2021Artificial intelligence and machine learning · 16 · 3 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 6 since 2021
YearPublicationVenuePosition
2026 Block-Wise Contrastive Few-Shot Learning With Multiview Mixing for Hyperspectral Image Cross-Scene Classification
abstract
The cross-scene classification of hyperspectral images (HSI) faces challenges posed by different sensors and various land cover categories, and methods based on the cross-domain Few-shot Learning (FSL) framework are commonly used to address the challenge. However, current few-shot learning (FSL) methods struggle with the issue of intra-domain inductive bias, where models tend to learn simplistic and potentially erroneous correlations within samples (classes with similar backgrounds are always confused due to simple correlations with the background), leading to poor generalization ability to target domains. Furthermore, most approaches employ a patch-wise representation learning mechanism, which provides the model with all information within a sample, but inherently limits its representation capacity when dealing with the limited prior information in FSL tasks. To address these issues, we propose a Block-wise Contrastive Few-shot Learning (BCFSL) framework. First, a multi-view cross-domain mixing strategy enhances spatial diversity and constructs harder samples to mitigate inductive bias. Second, a novel block-wise representation mechanism performs fine-grained feature extraction from local regions, improving generalization. Finally, a multi-view spectral reconstruction module preserves essential high-dimensional spectral information during training and assists the block-wise representation in fully leveraging prior knowledge. Extensive experiments on three benchmark HSI datasets validate the effectiveness and superiority of the proposed method.
Yingze Xie, Zhengyi Lv, Yuxiang Zhang 0005, Mengmeng Zhang 0005
IEEE Geosci. Remote. Sens. Lett.7
2026 Single-Source Domain Defect-Aware Adaptation and Style-Modulated Generalization Network for Multispectral Image Segmentation
abstract
Multispectral remote sensing image (MSI) semantic segmentation faces challenges of limited labeled data and significant scene variability. Although domain adaptation (DA) and domain generalization (DG) methods alleviate these issues to some extent, they still have limitations. DA requires target domain (TD) data, and DG has limited task adaptability. The recently emerged segment anything model (SAM) demonstrates exceptional zero-shot generalization capabilities, yet its visible-light training data and interactive prompt requirements prevent direct application to MSI segmentation tasks. To address these challenges, this article proposes a single-source domain defect-aware adaptation and style-modulated generalization network (SDSnet), which integrates two key innovations: defect-aware prompt learning that automatically focuses on high-difficulty regions through entropy-based defect detection, and style generalization learning that enhances cross-domain adaptability via codebook-based style modulation. Through knowledge distillation, SDSnet enables efficient inference using only the base network, without additional computational overhead. Extensive experiments on three TDs demonstrate SDSnet's superiority over state-of-the-art DA, DG, and SAM-based methods. Code will be available at https://github.com/zhaoboyu34526/SDSnet.
Wei Li 0032, Mengmeng Zhang 0005, Yunhao Gao
IEEE Trans. Cybern.3
2026 LiteMFT: Lightweight Multi-Modal Fine-Tuning for Semantic Segmentation
abstract
Multi-modal image segmentation has recently attracted considerable attention due to its ability to integrate complementary information from diverse sensors, thereby enabling more accurate semantic predictions in complex or specialized scenarios. However, as data volume and model capacity continue to grow, many existing methods suffer substantial increases in parameters and computational costs, particularly with the widespread adoption of Vision Foundation Models (VFMs). To address these challenges, we introduce a Lightweight Multi-modal Fine-Tuning framework (LiteMFT) designed for efficient and generalizable adaptation of RGB-pretrained VFMs to multi-modal semantic segmentation. By incorporating only a small number of trainable parameters, LiteMFT enables effective extension of existing models to handle multi-modal image fusion tasks. The framework centers around two key components: the Modality Local Competition (MLC) module, which dynamically and efficiently fuses complementary features across modalities, and the Gated Low-Rank Adapter (GLR), which improves the backbone's adaptability to multi-modal data through content-aware low-rank transformation. Extensive experiments on both bi-modal and tri-modal segmentation tasks demonstrate that LiteMFT not only achieves competitive or superior performance but also exhibits strong scalability for additional modalities, underscoring its practicality and broad applicability in multi-modal semantic segmentation.
Chengwang Guo, Yuxiang Zhang 0005, Mengmeng Zhang 0005, Huan Liu 0015, Wei Li 0032
IEEE Trans. Image Process.3
2026 Cross-Scene Hyperspectral Image Classification via Bidirectional Mamba and Domain Mixing Network
abstract
To overcome the challenges posed by domain shift in hyperspectral image (HSI) classification, methods based on domain adaptation (DA) have been widely used. Currently, most HSI DA methods focus on designing complex strategies to align the distributions of the source domain (SD) and the target domain (TD) in the feature space after feature extraction, yielding promising results. However, when there exists a large domain shift between SD and TD, it becomes challenging to map them into the same feature space. In this article, we propose the bidirectional mamba and domain mixing network (BMDMnet). Since pure CNN architectures are constrained in local feature extraction, while transformer-based models improve global feature capturing capability at the cost of high computational complexity, we propose the bidirectional mamba module (BMM) as an efficient solution for capturing long-range dependencies. In addition, a self-distillation strategy is employed during training. By utilizing a more stable teacher model, reliable predictions can be obtained in the TD. Subsequently, a domain mixing supervised learning (DMSL) module is designed, which creates a mixed domain by selecting low-entropy sample-pseudo-label pairs from the TD and randomly combining them with sample-label pairs from the SD. DMSL aims to introduce mixed domain to mitigate the inter-domain gap in the data space, thereby enabling the model to learn TD representations more effectively. Experiments demonstrate that BMDMnet outperforms state-of-the-art algorithms across three cross-scene datasets.
Junzhe Dang, Chengwang Guo, Mengmeng Zhang 0005, Yuxiang Zhang 0005, Wen Jia, Wei Li 0032
IEEE Trans. Neural Networks Learn. Syst.3
2025 Cross-Domain Hyperspectral Image Classification Based on Bi-Directional Domain Adaptation
abstract
Utilizing hyperspectral remote sensing technology enables the extraction of fine-grained land cover classes. Typically, satellite or airborne images used for training and testing are acquired from different regions or times, where the same class has significant spectral shifts in different scenes. In this paper, we propose a Bi-directional Domain Adaptation (BiDA) framework for cross-domain hyperspectral image (HSI) classification, which focuses on extracting both domain-invariant features and domain-specific information in the independent adaptive space, thereby enhancing the adaptability and separability to the target scene. In the proposed BiDA, a triple-branch transformer architecture (the source branch, target branch, and coupled branch) with semantic tokenizer is designed as the backbone. Specifically, the source branch and target branch independently learn the adaptive space of source and target domains, a Coupled Multi-head Cross-attention (CMCA) mechanism is developed in coupled branch for feature interaction and inter-domain correlation mining. Furthermore, a bi-directional distillation loss is designed to guide adaptive space learning using inter-domain correlation. Finally, we propose an Adaptive Reinforcement Strategy (ARS) to encourage the model to focus on specific generalized feature extraction within both source and target scenes in noise condition. Experimental results on cross-temporal/scene airborne and satellite datasets demonstrate that the proposed BiDA performs significantly better than some state-of-the-art domain adaptation approaches. In the cross-temporal tree species classification task, the proposed BiDA is more than 3%∼5% higher than the most advanced method. The codes will be available from the website: https://github.com/YuxiangZhang-BIT/IEEE TCSVT BiDA.
Yuxiang Zhang 0005, Wei Li 0032, Wen Jia, Mengmeng Zhang 0005, Ran Tao 0003, Shunlin Liang
IEEE Trans. Circuits Syst. Video Technol.4
2025 MGCD: Change Detection Focuses on Multiscale Style Domain Generalization
abstract
Change detection (CD) plays a crucial role in remote sensing (RS) applications. Although deep learning (DL)-based CD methods have achieved impressive performance, they typically rely on large amounts of well-annotated data, which is often scarce in real-world scenarios. This data scarcity leads to overfitting and limited generalization ability in conventional CD models. To address these challenges, this article proposes a novel few-sample CD framework named multiscale style DG method for change detection (MGCD). The core idea is to enhance the model’s ability to learn domain-invariant features by increasing data diversity across multiple style scales. Specifically, MGCD introduces global style diversity by incorporating out-of-domain natural images, and enriches local structural styles through unsupervised clustering and randomization. Additionally, a semantic consistency supervision (SCS) strategy is designed to guide multitemporal feature learning, enabling the network to better capture changes across diverse styles and scales. Extensive experiments conducted on three benchmark datasets demonstrate the effectiveness and robustness of the proposed MGCD framework in few-sample CD tasks.
Chengwang Guo, Mengmeng Zhang 0005, Yuxiang Zhang 0005, Huan Liu 0015, Wei Li 0032
IEEE Trans. Geosci. Remote. Sens.2
2025 Unsupervised Domain Adaptation With Hierarchical Masked Dual-Adversarial Network for End-to-End Classification of Multisource Remote Sensing Data
abstract
Although unsupervised domain adaptation (UDA) has been successfully applied for cross-scene classification of multisource remote sensing (MSRS) data, there are still some tough issues: 1) The vast majority of them are patch-based, requiring pixel by pixel processing at high complexity and ignoring the roles of unlabeled data between different domains. 2) Traditional masked autoencoder (MAE)-based methods lack effective multiscale analysis and require pre-training, ignoring the roles of low-level representations. As such, a hierarchical masked dual-adversarial DA network (HMDA-DANet) is proposed for cross-domain end-to-end classification of MSRS data. Firstly, a hierarchical asymmetric MAE (HAMAE) without pre-training is designed, containing a frequency dynamic large-scale convolutional (FDLConv) block to enhance important structural information in the frequency domain, and an intramodality enhancement and intermodality interaction (IAEIEI) block to embed some additional information beyond the domain distribution by expanding the cross-modal reconstruction space. Representative multimodal multiscale features can be extracted, while to some extent improving their generalization to the target domain. Then, a multimodal multiscale feature fusion (MMFF) block is built to model the spatial and scale dependencies for feature fusion and reduce the layer by layer transmission of redundancy or interference information. Finally, a dual-discriminator-based DA (DDA) block is designed for class-specific semantic feature and global structural alignments in both spatial and prediction spaces. It will enable HAMAE to model the cross-modal, cross-scale, and cross-domain associations, yielding more representative domain-invariant multimodal fusion features. Extensive experiments on five cross-domain MSRS datasets verify the superiority of the proposed HMDA-DANet over other state-of-the-art methods.
Wen-Shuai Hu, Wei Li 0032, Heng-Chao Li 0001, Xudong Zhao 0003, Mengmeng Zhang 0005, Ran Tao 0003
IEEE Trans. Geosci. Remote. Sens.5
2025 A Lightweight Spatial-Spectral Deformable CNN for UAV Hyperspectral Image Classification
abstract
In recent years, unmanned aerial vehicle (UAV) technology has shown great potential for application in hyperspectral image (HSI) classification tasks due to its advantages of flexible scheduling and fast response. However, existing deep learning-based classification algorithms have not thoroughly studied the problems caused by the higher spatial resolution of UAV HSIs. Such as more severe intra-class variation and higher computational resource consumption. These issues limit classification performance. To address these challenges, this paper proposes a lightweight spatial-spectral deformable convolutional neural network (LS2DCNet) for UAV HSI classification. This network reduces computational resource usage and data processing time, meeting the fast-response application requirements while ensuring classification performance. First, spatial-spectral deformable convolution (S2DConv) is designed to construct a lightweight feature extraction network. This not only enhances the adaptive extraction of fine features but also reduces computational resource consumption, and improves response speed. Second, a dynamic labeling-based joint loss (DL-JLoss) is designed to dynamically learn the distribution relationship between data classes. This improves the network’s generalization performance. A more realistic experimental validation was conducted on three UAV HSI datasets using regionally divided training samples, and inference comparisons were performed on embedded devices. The results show that the proposed LS2DCNet exhibits a better overall performance in terms of classification accuracy and inference speed.The relevant code can be found at https://github.com/niuroushu/A-Lightweight-Spatial-Spectral-Deformable-CNN-for-UAV-Hyperspectral-Image-Classification.
Xiaohu Ma, Mengmeng Zhang 0005, Zheng Kan, Yunhao Gao, Wei Li 0032
IEEE Trans. Geosci. Remote. Sens.2
2025 LHAS: A Lightweight Network Based on Hierarchical Attention for Hyperspectral Image Segmentation
abstract
Deep learning has garnered extensive attention in hyperspectral image (HSI) processing. However, its application in HSI semantic segmentation tasks has been relatively limited. Although segmentation methods can often interpret images up to two orders of magnitude faster than classification methods when interpreting images of the same scene, the segmentation task requires the training data to be fully labeled, i.e., each pixel has a corresponding label. Such data are scarce in HSI data. To address this problem, this article proposes a lightweight segmentation network based on a hierarchical attention segmentation network (LHAS), in which a generalized data augmentation (GDA) method is utilized to acquire relatively sufficient data for semantic segmentation. Specifically, the hierarchical attention module is designed to extract global and local information on HSI patches from different layers. A prototype auxiliary module (PAM) of cluster contrast has also been developed to enhance feature discrimination. Across two different datasets in various scenarios, the proposed LHAS demonstrates superior segmentation performance compared to existing methods, affirming its effectiveness. In addition, experiments conducted on embedded devices validate the efficacy of LHAS.
Lujie Song, Yunhao Gao, Yuanyuan Gui, Daguang Jiang, Mengmeng Zhang 0005, Huan Liu 0015, Wei Li 0032
IEEE Trans. Geosci. Remote. Sens.5
2025 GeoFlowNet-SAR: Earthquake Displacement Estimation From Synthetic Aperture Radar Images
abstract
Displacement estimation using remote sensing images is an effective approach for assessing surface displacement caused by natural disasters like earthquakes and landslides. By employing pixel correlation algorithms, high-precision displacement maps can be generated from images taken before and after surface movement. However, traditional methods often rely on spatial regularization or frequency masking to reduce high-frequency noise, which can smooth spatial details and result in biased displacement estimates, especially near sharp discontinuities typical of earthquake surface ruptures. Moreover, sub-pixel displacement estimation using Synthetic Aperture Radar (SAR) images remains a challenge compared to optical images, due to the strong impact of speckle noise. This paper presents GeoFlowNet-SAR, an innovative sub-pixel displacement estimation method leveraging SAR images. SAR offers advantages thanks to an all-weather observation and high penetration, making it suitable for conditions typically challenging for optical systems in the visible light spectrum. This study uses Sentinel-1 SAR Single Look Complex (SLC) images with dual-polarization (VV and VH modes) and Interferometric Wide (IW) swath mode to balance coverage and resolution. By training on simulated displacement datasets with realistic sharp discontinuities, GeoFlowNet-SAR directly predicts surface displacement fields, providing highly efficient, robust, and precise results, while overcoming some limitations of traditional methods. The effectiveness of the proposed methodological contribution is first quantitatively demonstrated using synthetic simulated earthquake datasets, including comparisons with state-of-the-art correlation methods. The method is further validated using two real remote sensing images from the 2019 Ridgecrest earthquake and from the 2023 Turkey-Syria earthquake. The observed results from these real datasets confirm the effectiveness of GeoFlowNet-SAR in practical applications. The codes are available at: https://gricad-gitlab.univ-grenoble-alpes.fr/giffards/geoflownet-sar.
James Hollingsworth, Erwan Pathier, Tristan Montagnon, Wei Li 0032, Mengmeng Zhang 0005, Ran Tao 0003, Jocelyn Chanussot, Sophie Giffard-Roisin
IEEE Trans. Geosci. Remote. Sens.6
2025 An Adaptive Weighted Metric Learning Network Based on Fractional Domain Decoupling for Hyperspectral Change Detection
abstract
Hyperspectral image change detection (HSI-CD) possesses strong capabilities in exploring subtle changes in land cover. Due to sensor noise and imaging conditions, different semantic land covers in the same spatial location may exhibit similar spectral characteristics, leading to pseudoinvariant phenomena (identification of changed areas as unchanged areas) and causing a higher rate of false negatives in the model. Existing methods primarily focus on obtaining auxiliary discriminative information from spatial correlations or temporal dependencies. However, the frequency domain, which possesses rich global gradient distribution information, is often overlooked. The fractional Fourier transform (FrFT) is an extension of the Fourier transform (FT), representing a temporal-frequency local transformation suitable for processing nonstationary signals. Furthermore, multiorder fractional Fourier domains provide more observable domains for change discrimination. In this work, the application of FrFT is extended to the field of HSI-CD, and an adaptive weighted metric learning network based on fractional domain decoupling (FrFTML) is proposed. Specifically, the fractional domain decoupling (FrDD) module transforms the original HSI into multiorder FrFT domains and extracts their rich spatial-frequency mixed information, effectively suppressing noise while enhancing the representation of subtle differences. In addition, an adaptive weighted metric learning (AWML) framework is designed to merge multiorder fractional Fourier domain information in an adaptively weighted fusion manner. It introduces deep metric learning to explore the distances between samples of different categories that have relatively high similarity, so as to guide the direction of adaptive weighted fusion. Finally, the differential mask attention (DMA) module is designed to explore global contextual differences between bitemporal HSIs, obtaining change features with well-represented differences. Some experiments conducted on three public datasets indicate that FrFTML outperforms other state-of-the-art methods. Furthermore, the proposed method exhibits superiority in dealing with land cover that may lead to pseudoinvariant phenomena (identification of changed areas as unchanged areas).
Shou Feng, Tianyu Lan, Yuanze Fan, Mengmeng Zhang 0005, Chunhui Zhao 0003, Wei Li 0032, Ran Tao 0003
IEEE Trans. Neural Networks Learn. Syst.4
2025 Distribution-Independent Domain Generalization for Multisource Remote Sensing Classification
abstract
The availability of multisource remote sensing data provides the possibility for comprehensive observation. Convolutional neural networks (CNNs) naturally integrate multisource feature extractors and classifiers into an end-to-end multilayer design. However, CNN assumes data are independent and identically distributed. In practice, it is not always possible to access the labels or even data of the testing scenes. Therefore, the CNN-based methods have exposed its limitation on generalization ability. To solve the issue, a feature-distribution-independent network (FDINet) is designed for multisource remote sensing cross-domain classification without feature alignment and decoupling operations. On one hand, an elegantly designed baseline is used for extracting multisource cross-domain features. The baseline extracts the common line and texture features through shallow weight-sharing networks. More importantly, the modality prediction probability is used to measure the similarity between the source domains and the target domains, thereby improving cross-domain collaboration capabilities. On the other hand, the sharpness-aware feature discriminating (SAFD) strategy is developed for model optimization. Specifically, the generalization ability is improved by minimizing the sharpness of local optima. To avoid the decrease in feature discrimination caused by the gradient conflict between sharpness and overall loss, the discrimination constraints are designed to balance feature discrimination and generalization ability. Comprehensive experiments are conducted on two datasets, which demonstrate that the proposed FDINet outperforms other competitors in terms of quantitative and qualitative analyses.
Yunhao Gao, Mengmeng Zhang 0005, Wei Li 0032, Ran Tao 0003
IEEE Trans. Neural Networks Learn. Syst.2
2025 Domain Information Mining and State-Guided Adaptation Network for Multispectral Image Segmentation
abstract
Segment anything model (SAM), as a prompt-based image segmentation foundation model, demonstrates strong task versatility and domain generalization (DG) capabilities, providing a new direction for solving cross-scene segmentation tasks. However, SAM still has limitations in multispectral cross-domain segmentation tasks, mainly reflected in: 1) insufficient information utilization, which is reflected in the neglect of nonvisible spectral information and the shift information contained in source domain (SD) samples and target domain (TD) samples; and 2) lack of cross-domain strategies, which leads to insufficient cross-domain adaptation (DA) ability in downstream tasks. To address these challenges, we combine the respective advantages of masked autoencoder (MAE) and cross-domain strategies, propose an improved SAM DA network structure called domain information mining and state-guided adaptation network (DSAnet), aiming to enhance SAM's performance in multispectral cross-domain segmentation tasks from both data and task levels. At the data level, DSAnet incorporates a style masking learning component, which randomly masks image features and replaces them with domain-specific learnable tokens, integrated with the image reconstruction task, to mine the style information and domain invariance of the image itself. At the task level, DSAnet introduces domain state learning and style-guided segmentation: domain state learning, through a state sequence modeling approach, designs specific state representations for SD and TD to capture interdomain differences, thereby reducing task shift. Meanwhile, the learned domain state information can be directly applied to the inference stage. Style prompt segmentation guides the segmentation training process of SD images with TD style prompts, improving SAM's adaptability in cross-domain multispectral segmentation downstream tasks. Extensive experiments on three multitemporal multispectral image (MSI) datasets demonstrate the superiority of the proposed method compared to state-of-the-art cross-domain strategies and SAM variant methods.
Mengmeng Zhang 0005, Wei Li 0032, Yunhao Gao
IEEE Trans. Neural Networks Learn. Syst.2
2024 Clusterformer for Pine Tree Disease Identification Based on UAV Remote Sensing Image Segmentation
abstract
Pine wilt disease (PWD) is one of the most prevalent pine trees diseases, resulting in both ecological and economic havoc. UAV remote sensing segmentation plays a crucial role in early identifying and preventing PWD. However, deep learning segmentation models customized for PWD identification in scenarios with complex backgrounds have not received extensive exploration. In this paper, we propose a novel UAV remote sensing segmentation model called Clusterformer with a conventional encoder-decoder structure. The encoder is comprised of the specially designed Cluster Transformer, which includes a cluster token mixer and a spatial-channel feed-forward network (SC-FFN). The cluster token mixer utilizes constructed clusters from the feature maps to represent pixels, thereby reducing redundant and interfering information. The SC-FFN extracts multi-scale spatial information through depth-wise convolutions and channel information through a multilayer perceptron in sequence. The decoder primarily consists of the specially designed D-Cluster Transformer. The token mixer of the D-Cluster Transformer employs constructed clusters from high-level decoded tokens to represent low-level encoded tokens without relying on traditional upsampling methods such as interpolation, transpose convolution, or patch expansion. Consequently, more robust and less redundant features from high-level decoded feature maps are transferred to low-level encoded feature maps. Experimental results on two PWD datasets demonstrate that Clusterformer outperforms existing state-of-the-art segmentation models. This confirms the effectiveness and efficiency of Clusterformer in PWD identification. Code is available at https://github.com/huanliu233/Clusterformer.
Huan Liu 0015, Wei Li 0032, Wen Jia, Mengmeng Zhang 0005, Lujie Song, Yuanyuan Gui
IEEE Trans. Geosci. Remote. Sens.5
2024 LIRnet: Lightweight Hyperspectral Image Classification Based on Information Redistribution
abstract
Deep learning has received much attention in hyperspectral image (HSI) classification. However, most deep learning methods design relatively complex feature extraction and processing network modules for the characteristics of HSIs, which may not be necessary for relatively simple patch-based HSI classification tasks. The complex network structure and high feature channel dimension lead to large computational complexities, which limit the practical applicability of HSI. In this article, an elegant lightweight HSI classification-based information redistribution network (LIRnet) is proposed to separate and reaggregate the feature information to achieve feature information homogenization and extract discriminative feature information, respectively. The classification performance of LIRnet is better than that of existing methods on three different datasets in different scenarios, which proves its effectiveness. In addition, experiments on embedded devices verify the computational efficacy of LIRnet.
Lujie Song, Yunhao Gao, Xiangyang Jiang, Xiaofei Yin, Daguang Jiang, Mengmeng Zhang 0005, Wei Li 0032
IEEE Trans. Geosci. Remote. Sens.7
2024 FPNFormer: Rethink the Method of Processing the Rotation-Invariance and Rotation-Equivariance on Arbitrary-Oriented Object Detection
abstract
Feature pyramid network transformer decoder (FPNFormer) module, which can effectively deal with the strong rotation arbitrary of remote sensing images while improving the expressiveness and robustness of the model. It is a plug-and-play module that can be well transferred to various detection models and significantly improves performance. Specifically, we use the computational method of transformer decoder to deal with the problem that the image has any orientation, and its output weakly depends on the order of the input data. We apply it to the feature fusion stage and design two ways top-down and down-top to fuse features of different scales, which enables the model to have a more vital ability to perceive objects at different scales and angles. Experiments on commonly used benchmarks (DOTA1.0, DOTA1.5, SSDD, and RSDD) demonstrate that the proposed FPNFormer module significantly improves the performance of multiple arbitrary-oriented object detectors, such as 1.99% map improvement of rotated retinanet on DOTA’s cross-validation set. On RSDD datasets, the baseline model using FPNFormer improves the map of large objects by 5.1%. Combined with more competitive models, the proposed method can achieve a 79.39% map on the DOTA1.0 dataset. The code is available athttps://github.com/bityangtian/FPNFormer.
Mengmeng Zhang 0005, Yangfan Li 0002, Wei Li 0032
IEEE Trans. Geosci. Remote. Sens.2
2024 Unbalanced Class Learning Network With Scale-Adaptive Perception for Complicated Scene in Remote Sensing Images Segmentation
abstract
The semantic segmentation of wide-field remote sensing images plays a significant role in many fields. However, due to the complexity of the content of remote sensing images, the dataset often has an uneven distribution of land type between different classes and large gaps in the scales of different objects. This often creates great problems for fine segmentation. To solve the issues, an unbalanced class learning network with Scale-adaptive perception (UCSANet) is proposed, which can adaptively cope with Multi-scale objects and unbalanced classes. The design can be inserted in any convolution network easily and can enrich features without increasing too many parameters. The network groups feature and uses atrous convolutions with different dilated rates on different groups to extract Multi-scale features while separable convolutions reduce the amount of network parameters. Then, the fusion of features between different scales is achieved through the self-attention mechanism. Furthermore, a weight map is designed to adaptively combine the predictions of two segmentation heads with Cross-Entropy loss and Lovasz-Softmax loss respectively, which enable the network to focus on learning low-frequency classes without affecting high-frequency classes. Experimental results on GF-6 MSI datasets demonstrate that the proposed UCSANet performs significantly better than others and achieves multi-class segmentation more accurately.
Mengmeng Zhang 0005, Wei Li 0032, Yunhao Gao, Yuanyuan Gui, Yuxiang Zhang 0005
IEEE Trans. Geosci. Remote. Sens.2
2024 GCCD: A Generative Cross-Domain Change Detection Network
abstract
Change detection (CD) in hyperspectral image (HSI) is of great importance in the remote sensing area. The HSI-CD method based on deep learning (DL) has shown significant progress in achieving precise detection performance. However, many existing methods overlook cross-domain challenges in the CD task. In addition, the scarcity of annotated samples makes the DL models prone to overfitting. To address these issues, a generative cross-domain CD (GCCD) network based on a domain generalization (DG) technique is proposed. GCCD consists of a generator and a discriminator. The generator, with a Morph encoder (ME) and a Semantic encoder (SE), preserves fundamental structural information while introducing randomization to style and content. The discriminator extracts change information for discrimination through dual-temporal images and their difference map. Supervised adversarial learning between the generator and discriminator enhances the model’s ability to extract domain-invariant information. Extensive experiments on various datasets demonstrate the superior performance of the proposed method.
Mengmeng Zhang 0005, Chengwang Guo, Yuxiang Zhang 0005, Huan Liu 0015, Wei Li 0032
IEEE Trans. Geosci. Remote. Sens.1
2024 Remote Sensing Collaborative Classification Using Multimodal Adaptive Modulation Network
abstract
With the development of remote sensing technology, more and more data sources are available for landcover classification tasks, such as hyperspectral images (HSIs), light detection and ranging (LiDAR) data, and synthetic aperture radar (SAR) data. Due to the unique information carried by different sources, the collaborative use of multiple remote sensing data has become a key research direction in landcover classification tasks. Most existing methods are only capable of dealing with two types of remote-sensing images, limiting the potential for the collaboration of more sources. Also, the traditional feature fusion method uses addition or concatenation manner to integrate information, which makes it difficult to make full use of complement characteristics between modalities. As a remedy, we propose a novel multimodal adaptive modulation network (MAMNet), for landcover classification tasks using multimodal remote sensing data. First, the cross-modal interacting module (CIM) is utilized for information absorption between modalities. The feature representation is enhanced, and the modality-specific information is preserved. Second, the modal attention layer (MAL) is designed for multimodal feature fusion. Softmax attention is utilized to eliminate redundant information among the three modal features. Finally, an adaptive multimodal margin loss (AMM loss) is proposed to balance the consistency and diversity of multimodal features. It encourages adjustable decision margins between sources, which enables the model to better utilize complementary information between modalities and partially avoids model-overfitting by defining a more difficult learning target. Experimental results on two benchmark remote sensing datasets show the effectiveness of the proposed method compared with several state-of-the-art approaches.
Mengmeng Zhang 0005, Rongjie Chen, Yunhao Gao, Wei Li 0032
IEEE Trans. Geosci. Remote. Sens.1
2024 Intermediate Domain Prototype Contrastive Adaptation for Spartina alterniflora Segmentation Using Multitemporal Remote Sensing Images
abstract
As an invasive plant in wetlands, Spartina alterniflora (S. alterniflora) causes immeasurable damage to wetland ecosystems. Observing S.alterniflora using multitemporal remote sensing data helps us better understand its further development and facilitates effective containment of its invasion trend. However, inconsistent representation across remote sensing data from different time periods poses a challenge. Fortunately, the utilization of unsupervised domain adaptation (UDA) techniques helps in addressing such issues and enables the exploration of rich temporal dimension information in multitemporal remote sensing data, revealing the spatio-temporal distribution characteristics of S.alterniflora. However, existing UDA methods mostly focus on directly aligning the global or intraclass distribution representations across domains, which overlooks the issue of significant differences between extreme domains and lacks exploration of interclass relationships. To address these limitations, an intermediate domain prototype class-level learning network (IDPNet) is proposed. IDPNet utilizes dynamically generated intermediate domain (ID) features to construct class prototypes while incorporating interclass information into the prototype construction, achieving the class-centered distribution alignment for adaptation. Moreover, intermediate domain feature generation module (IFM) is employed in IDPNet to blend the latent representations from various domains and generate ID features in real time. Additionally, the hierarchical feature fusion module (HFM) is designed to enable IDPNet to learn more discriminative and robust spatio-temporal distribution features, thereby reducing the loss of information from patches. Experimental results on two cross-year multispectral datasets demonstrate that the proposed IDPNet outperforms several state-of-the-art UDA methods.
Mengmeng Zhang 0005, Wei Li 0032, Xiukai Song, Yunhao Gao, Yuxiang Zhang 0005
IEEE Trans. Geosci. Remote. Sens.2
2024 Relationship Learning From Multisource Images via Spatial-Spectral Perception Network
abstract
Advances in multisource remote sensing have allowed for the development of more comprehensive observation. The adoption of deep convolutional neural networks (CNN) naturally includes spatial-spectral information, which has achieved promising performance in multisource data classification. However, challenges are still found with the extraction of spatial distribution and spectrum relationships, which eventually limit the classification performance. To solve the issue, a spatial-spectral perception network (S2PNet) is proposed to extract the advantages of different data sources and the cross information between data sources in a targeted manner. Specifically, the spatial perception network is developed to build the spatial distribution relationship from high-resolution images, while the spectral perception network extracts the spectrum relationship from spectral images. For perceiving cross information, a memory unit is utilized to store the features from different data sources in succession. In addition, the distance loss and reconstruction loss are introduced to keep the feature integrity, and the cross-entropy loss ensures that features can distinguish different classes. The comprehensive experiments are conducted on several datasets to validate the superiority of the proposed algorithm. The proposed S2PNet outperforms the considered classifiers with an average improvement of +0.77%, +5.62%, +1.58%, and +1.79% for overall accuracy values.
Yunhao Gao, Wei Li 0032, Mengmeng Zhang 0005, Ran Tao 0003
IEEE Trans. Image Process.4
2024 SegHSI: Semantic Segmentation of Hyperspectral Images With Limited Labeled Pixels
abstract
Hyperspectral images (HSIs), with hundreds of narrow spectral bands, are increasingly used for ground object classification in remote sensing. However, many HSI classification models operate pixel-by-pixel, limiting the utilization of spatial information and resulting in increased inference time for the whole image. This paper proposes SegHSI, an effective and efficient end-to-end HSI segmentation model, alongside a novel training strategy. SegHSI adopts a head-free structure with cluster attention modules and spatial-aware feedforward networks (SA-FFN) for multiscale spatial encoding. Cluster attention encodes pixels through constructed clusters within the HSI, while SA-FFN integrates depth-wise convolution to enhance spatial context. Our training strategy utilizes a student-teacher model framework that combines labeled pixel class information with consistency learning on unlabeled pixels. Experiments on three public HSI datasets demonstrate that SegHSI not only surpasses other state-of-the-art models in segmentation accuracy but also achieves inference time at the scale of seconds, even reaching sub-second speeds for full-image classification. Code is available at https://github.com/huanliu233/SegHSI.
Huan Liu 0015, Wei Li 0032, Xiang-Gen Xia 0001, Mengmeng Zhang 0005, Zhengqi Guo, Lujie Song
IEEE Trans. Image Process.4
2024 A Multistage Information Complementary Fusion Network Based on Flexible-Mixup for HSI-X Image Classification
abstract
Mixup-based data augmentation has been proven to be beneficial to the regularization of models during training, especially in the remote-sensing field where the training data is scarce. However, in the process of data augmentation, the Mixup-based methods ignore the target proportion in different inputs and keep the linear insertion ratio consistent, which leads to the response of label space even if no effective objects are introduced in the mixed image due to the randomness of the augmentation process. Moreover, although some previous works have attempted to utilize different multimodal interaction strategies, they could not be well extended to various remote-sensing data combinations. To this end, a multistage information complementary fusion network based on flexible-mixup (Flex-MCFNet) is proposed for hyperspectral-X image classification. First, to bridge the gap between the mixed image and the label, a flexible-mixup (FlexMix) data augmentation strategy is designed, where the weight of the label increases with the ratio of the input image to prevent the negative impact on the label space because of the introduction of invalid information. More importantly, to summarize diverse remote-sensing data inputs including various modal supplements and uncertainties, a multistage information complementary fusion network (MCFNet) is developed. After extracting the features of hyperspectral and complementary modalities [X-modal, including multispectral, synthetic aperture radar (SAR), and light detection and ranging (LiDAR)] separately, the information between complementary modalities is fully interacted and enhanced through multiple stages of information complement and fusion, which is used for the final image classification. Extensive experimental results have demonstrated that Flex-MCFNet can not only effectively expand the training data, but also adequately regularize different data combinations to achieve state-of-the-art performance.
Mengmeng Zhang 0005, Wei Li 0032, Ran Tao 0003
IEEE Trans. Neural Networks Learn. Syst.2
2024 Graph Information Aggregation Cross-Domain Few-Shot Learning for Hyperspectral Image Classification
abstract
Most domain adaptation (DA) methods in cross-scene hyperspectral image classification focus on cases where source data (SD) and target data (TD) with the same classes are obtained by the same sensor. However, the classification performance is significantly reduced when there are new classes in TD. In addition, domain alignment, as one of the main approaches in DA, is carried out based on local spatial information, rarely taking into account nonlocal spatial information (nonlocal relationships) with strong correspondence. A graph information aggregation cross-domain few-shot learning (Gia-CFSL) framework is proposed, intending to make up for the above-mentioned shortcomings by combining FSL with domain alignment based on graph information aggregation. SD with all label samples and TD with a few label samples are implemented for FSL episodic training. Meanwhile, intradomain distribution extraction block (IDE-block) and cross-domain similarity aware block (CSA-block) are designed. The IDE-block is used to characterize and aggregate the intradomain nonlocal relationships and the interdomain feature and distribution similarities are captured in the CSA-block. Furthermore, feature-level and distribution-level cross-domain graph alignments are used to mitigate the impact of domain shift on FSL. Experimental results on three public HSI datasets demonstrate the superiority of the proposed method. The codes will be available from the website: https://github.com/YuxiangZhang-BIT/IEEE_TNNLS_Gia-CFSL.
Yuxiang Zhang 0005, Wei Li 0032, Mengmeng Zhang 0005, Ran Tao 0003, Qian Du 0001
IEEE Trans. Neural Networks Learn. Syst.3
2024 Cross-Scene Joint Classification of Multisource Data With Multilevel Domain Adaption Network
abstract
Domain adaption (DA) is a challenging task that integrates knowledge from source domain (SD) to perform data analysis for target domain. Most of the existing DA approaches only focus on single-source-single-target setting. In contrast, multisource (MS) data collaborative utilization has been extensively used in various applications, while how to integrate DA with MS collaboration still faces great challenges. In this article, we propose a multilevel DA network (MDA-NET) for promoting information collaboration and cross-scene (CS) classification based on hyperspectral image (HSI) and light detection and ranging (LiDAR) data. In this framework, modality-related adapters are built, and then a mutual-aid classifier is used to aggregate all the discriminative information captured from different modalities for boosting CS classification performance. Experimental results on two cross-domain datasets show that the proposed method consistently provides better performance than other state-of-the-art DA approaches.
Mengmeng Zhang 0005, Xudong Zhao 0003, Wei Li 0032, Yuxiang Zhang 0005, Ran Tao 0003, Qian Du 0001
IEEE Trans. Neural Networks Learn. Syst.1
2024 Fractional Fourier Image Transformer for Multimodal Remote Sensing Data Classification
abstract
With the recent development of the joint classification of hyperspectral image (HSI) and light detection and ranging (LiDAR) data, deep learning methods have achieved promising performance owing to their locally sematic feature extracting ability. Nonetheless, the limited receptive field restricted the convolutional neural networks (CNNs) to represent global contextual and sequential attributes, while visual image transformers (VITs) lose local semantic information. Focusing on these issues, we propose a fractional Fourier image transformer (FrIT) as a backbone network to extract both global and local contexts effectively. In the proposed FrIT framework, HSI and LiDAR data are first fused at the pixel level, and both multisource feature and HSI feature extractors are utilized to capture local contexts. Then, a plug-and-play image transformer FrIT is explored for global contextual and sequential feature extraction. Unlike the attention-based representations in classic VIT, FrIT is capable of speeding up the transformer architectures massively and learning valuable contextual information effectively and efficiently. More significantly, to reduce redundancy and loss of information from shallow to deep layers, FrIT is devised to connect contextual features in multiple fractional domains. Five HSI and LiDAR scenes including one newly labeled benchmark are utilized for extensive experiments, showing improvement over both CNNs and VITs.
Xudong Zhao 0003, Mengmeng Zhang 0005, Ran Tao 0003, Wei Li 0032, Wenzi Liao, Lianfang Tian, Wilfried Philips
IEEE Trans. Neural Networks Learn. Syst.2
2023 Multi-Modal Domain Generalization for Cross-Scene Hyperspectral Image Classification
abstract
The large-scale pre-training image-text foundation models have excelled in a number of downstream applications. The majority of domain generalization techniques, however, have never focused on mining linguistic modal knowledge to enhance model generalization performance. Additionally, text information has been ignored in hyperspectral image classification (HSI) tasks. To address the aforementioned shortcomings, a Multi-modal Domain Generalization Network (MDG) is proposed to learn cross-domain invariant representation from cross-domain shared semantic space. Only the source domain (SD) is used for training in the proposed method, after which the model is directly transferred to the target domain (TD). Visual and linguistic features are extracted using the dual-stream architecture, which consists of an image encoder and a text encoder. A generator is designed to obtain extended domain (ED) samples that are different from SD. Furthermore, linguistic features are used to construct a cross-domain shared semantic space, where visual-linguistic alignment is accomplished by supervised contrastive learning. Extensive experiments on two datasets show that the proposed method outperforms state-of-the-art approaches.
Yuxiang Zhang 0005, Mengmeng Zhang 0005, Wei Li 0032, Ran Tao 0003
ICASSP2
2023 MMS:Multi-Source Mutual Supervision Semantic Segmentation
abstract
How to use multi-source data for semantic segmentation is a hot topic. In this article, a new multi-source image data semantic segmentation based on multi-source mutual supervision (MMS) has been proposed. First, multi-source data from the same region are trained separately using identical segmentation networks for initialization; then, MMS performs mutual supervision training on these initialized networks, thus adaptively cooperating differences in information distribution between multiple sources of image data, and combines real label constraints output consistency across different networks. Experimental results demonstrate that the proposed MMS strategy can effectively coordinate information between multi-source data, improve segmentation accuracy, and is effective in different semantic segmentation networks and different types of data sets.
Shibo Guo, Yuanyuan Gui, Mengmeng Zhang 0005, Wei Li 0032
IGARSS3
2023 Hyperspectral Image Classification of Tree Species with Low-Depth Features
abstract
Classification of tree species is of great significance to forest surveys. Recently, considering the low differences of spectral information among tree species, enhancing the dependence between long-distance bands has become a research hotspot. A tree species classification method based on a convolutional (2-dimension) long short-term memory (Conv2DLSTM) network and transformer is proposed. First, the main features of HSI are retained by principal component analysis (PCA). Then, the Conv2DLSTM network obtains the global correlation information between long-distance band pixels, and the 3-dimensional convolutional neural network (3DCNN) updates the local spatial-spectral information. Finally, low-level features are converted into semantic tags to guide the modeling of high-level semantic features. The experimental results on the forest dataset demonstrate that the proposed method is superior to other competitive work.
Zhengqi Guo, Mengmeng Zhang 0005, Wen Jia, Wei Li 0032
IGARSS2
2023 Hyperspectral and LiDAR Data Classification Based on Structural Optimization Transmission
abstract
With the development of the sensor technology, complementary data of different sources can be easily obtained for various applications. Despite the availability of adequate multisource observation data, for example, hyperspectral image (HSI) and light detection and ranging (LiDAR) data, existing methods may lack effective processing on structural information transmission and physical properties alignment, weakening the complementary ability of multiple sources in the collaborative classification task. The complementary information collaboration manner and the redundancy exclusion operator need to be redesigned for strengthening the semantic relatedness of multisources. As a remedy, we propose a structural optimization transmission framework, namely, structural optimization transmission network (SOT-Net), for collaborative land-cover classification of HSI and LiDAR data. Specifically, the SOT-Net is developed with three key modules: 1) cross-attention module; 2) dual-modes propagation module; and 3) dynamic structure optimization module. Based on above designs, SOT-Net can take full advantage of the reflectance-specific information of HSI and the detailed edge (structure) representations of multisource data. The inferred transmission plan, which integrates a self-alignment regularizer into the classification task, enhances the robustness of the feature extraction and classification process. Experiments show consistent outperformance of SOT-Net over baselines across three benchmark remote sensing datasets, and the results also demonstrate that the proposed framework can yield satisfying classification result even with small-size training samples.
Mengmeng Zhang 0005, Wei Li 0032, Yuxiang Zhang 0005, Ran Tao 0003, Qian Du 0001
IEEE Trans. Cybern.1
2023 Adversarial Complementary Learning for Multisource Remote Sensing Classification
abstract
Convolutional neural networks (CNN) have attracted increasing attention in the field of multimodal cooperation. Recently, the adoption of CNN-based methods has achieved remarkable performance in multisource remote sensing data classification. However, it is still confronted with challenges in the aspect of complementarity extraction. In this paper, the adversarial complementary learning strategy is embedded into the CNN model called ACL-CNN, which is employed to extract the complementary information of the multisource data. The proposed ACL-CNN is able to filter out the common patterns and specific patterns from multisource data by conducting the adversarial max-min game. Especially, the modality-independent common patterns constitute the basic representation of the land-covers, while the specific patterns that are linearly independent of the common patterns that provide the supplementary representation. Therefore, the complementary information is mapped to a compact and discriminative representation. To eliminate the singularity noise, a learnable pattern sampling module (PSM) is designed to extract the mutual-exclusion relationship between specific patterns. Extensive experiments over three datasets demonstrate the superiority of the proposed ACL-CNN compared with several classification technologies.
Yunhao Gao, Mengmeng Zhang 0005, Wei Li 0032, Xiukai Song, Xiangyang Jiang, Yuanqing Ma
IEEE Trans. Geosci. Remote. Sens.2
2023 Cross-Scale Mixing Attention for Multisource Remote Sensing Data Fusion and Classification
abstract
Hyperspectral and multispectral images (HS/MS) fusion and classification as an important branch of data quality improvement and interpretation, has attracted increasing attention in recent years. However, the unavailable sensor prior still limits the performance of many traditional fusion methods, consequently deteriorating the classification results. Despite the unsupervised methods based on convolutional neural network (CNN) making a lot of attempts to mitigate the limitations, challenges with extracting the long-range dependencies hamper the performance. To address these impediments, a transformer-based baseline constructed by the cross-scale mixing attention (CSMFormer) is designed for HS/MS fusion and classification. Especially, the spatial-spectral mixer (SSMixer) is utilized to extract the long-range dependencies at large scale. Simultaneously, cross-scale feature calibration is achieved by combining information from the original scale. After that, nonlinear enhancement module (NLEM) is designed to encourage feature discrimination. Note that the spatial and spectral mixers can be replaced by any spatial-spectral feature extractors. Therefore, the proposed CSMFormer is flexible in data fusion, land-covers classification, segmentation, etc. Experiments about data fusion and land-covers classification on two HS/MS wetland remote sensing scenes demonstrate the superiority of the proposed CSMFormer baseline, improving the data quality and classification precision.
Yunhao Gao, Mengmeng Zhang 0005, Wei Li 0032
IEEE Trans. Geosci. Remote. Sens.2
2023 Multiarea Target Attention for Hyperspectral Image Classification
abstract
In hyperspectral image (HSI) classification, objects corresponding to pixels of different classes exhibit varying size characteristics, which causes a challenge for effective pixelwise feature extraction and classification. In this article, we propose a novel multiscale model, called multiarea target attention (MATA). The proposed MATA uses an architecture that includes a shared feature extractor (FE) and classifier to capture multiscale spectral–spatial information effectively and efficiently. The FE uses a multiscale target attention module (MSTAM) to extract spectral–spatial information from target pixels and their similar pixels across multiscale areas, while$L_{2}$-normalization is used to address discrepancies between features of different scales. The classifier adopts a classwise decision weighting strategy to account for the varying sizes of different classes and the different contributions of semantic features at each scale to each class. Experimental results on five public HSI datasets demonstrate that the proposed MATA outperforms existing state-of-the-art single- and multiscale models, confirming its effectiveness and efficiency in HSI classification. Code is available athttps://github.com/huanliu233/MATA.
Huan Liu 0015, Wei Li 0032, Xiang-Gen Xia 0001, Mengmeng Zhang 0005, Ran Tao 0003
IEEE Trans. Geosci. Remote. Sens.4
2023 Mask-Reconstruction-Based Decoupled Convolution Network for Hyperspectral Imagery Classification
abstract
Deep learning has attracted much attention in hyperspectral image(HSI) classification. However, most deep learning methods ignore the information loss during spatial-spectral feature extraction, which potentially affects the classification performance. In this article, Mask Reconstruction-Based Decoupled Convolution Network(MrDCN) is proposed, which including the decoupled feature extraction module (DFEM) to extract spectral information and spatial information of target HSI patch respectively. The reconstruction modules are designed to maintain the feature extraction ability of DFEM and ensure that discriminative information in high-dimensional features and low-dimensional features is preserved. MrDCN outperforms state-of-the-art methods in classification on three datasets of various scenarios, which indicates its effectiveness, and experiments on embedded devices are executed to affirm the efficiency of MrDCN.
Lujie Song, Mengmeng Zhang 0005, Wei Li 0032, Daguang Jiang, Huan Liu 0015, Yuxiang Zhang 0005
IEEE Trans. Geosci. Remote. Sens.2
2023 Large Kernel Sparse ConvNet Weighted by Multi-Frequency Attention for Remote Sensing Scene Understanding
abstract
Remote sensing scene understanding is a highly challenging task, and has gradually emerged as a research hotspot in the field of intelligent interpretation of remote sensing data. Recently, the use of convolutional neural networks (CNNs) has been proven to be a fruitful advancement. However, with the emergence of visual transformers (ViTs), the limitations of traditional small convolutional kernels in directly capturing a large receptive field have posed significant challenges to their dominant role. Additionally, the fixed neuron connections between different convolutional layers have weakened the practicality and adaptability of the models. Furthermore, the global average pooling also leads to the loss of effective information in the acquired features. In this work, a Large kernel Sparse ConvNet weighted by Multi-frequency Attention (LSCNet) is proposed. Firstly, unlike traditional convolutional neural networks, it utilizes two parallel rectangular convolutional kernels to approximate a large kernel, achieving comparable or even better results than ViTs-based methods. Secondly, an adaptive sparse optimization strategy is employed to dynamically optimize the fixed neuron connections between different convolutional layers, achieving a favorable connectivity pattern for capturing abstract features more accurately. Lastly, a novel multi-frequency attention (MFA) module is used to replace global average pooling (GAP), so as to preserve more useful information while weighting the recognition features, thereby enhancing the discriminative and learning abilities of the model. In the conducted experiments, LSCNet achieves the best recognition results on three well-known remote sensing aerial datasets when compared to the state-of-the-art methods (including ViTs-based methods).
Wei Li 0032, Mengmeng Zhang 0005, Jocelyn Chanussot
IEEE Trans. Geosci. Remote. Sens.3
2023 Remote-Sensing Scene Classification via Multistage Self-Guided Separation Network
abstract
In recent years, remote sensing scene classification is one of research hotspots and has played an important role in the field of intelligent interpretation of remote sensing data. However, various complex objects and backgrounds form a variety of remote sensing scenes through spatial combination and correlation, which brings great challenges to accurately classify different scenes. Among them, the insufficient feature difference brought about the unbalanced change of background and target between inter-class sample and the feature representation inconsistency caused by the difference of representation among the intra-class samples have become obstacles to effectively distinguish different scene images. To address these issues, a Multi-stage Self-Guided Separation Network (MGSNet) is proposed for remote sensing scene classification. First of all, different from the previous work, it attempts to utilize the background information outside the effective target in the image as a decision aid through a target-background separation strategy to improve the distinguish ability between target similarity-background difference samples. In addition, the diversity of feature concerns among different network branches is expanded through contrastive regularization to improve the separation of target-background information. Additionally, a self-guided network is proposed to find common features between intra-class samples and improve the consistency of feature representation. It combines the texture and morphological features of images to guide feature learning, effectively reducing the impact of intra-class differences. Extensive experimental results on three benchmark demonstrate that MGSNet can achieve better classification performance compared to the state-of-the-art approaches.
Wei Li 0032, Mengmeng Zhang 0005, Ran Tao 0003, Jocelyn Chanussot
IEEE Trans. Geosci. Remote. Sens.3
2023 Morphological Transformation and Spatial-Logical Aggregation for Tree Species Classification Using Hyperspectral Imagery
abstract
Hyperspectral image (HSI) consists of abundant spectral and spatial characteristics, which contribute to a more accurate identification of materials and land covers. However, most existing methods of hyperspectral image analysis primarily focus on spectral knowledge or coarse-grained spatial information while neglecting the fine-grained morphological structures. In the classification task of complex objects, spatial morphological differences can help to search for the boundary of fine-grained classes, e.g., forestry tree species. Focusing on subtle traits extraction, a spatial-logical aggregation network (SLA-NET) is proposed with morphological transformation for tree species classification. The morphological operators are effectively embedded with the trainable structuring elements, which contributes to distinctive morphological representations. We evaluate the classification performance of the proposed method on two tree species datasets, and the results demonstrate that the proposed SLA-NET significantly outperforms the other state-of-the-art classifiers.
Mengmeng Zhang 0005, Wei Li 0032, Xudong Zhao 0003, Huan Liu 0015, Ran Tao 0003, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.1
2023 Language-Aware Domain Generalization Network for Cross-Scene Hyperspectral Image Classification
abstract
Text information including extensive prior knowledge about land cover classes has been ignored in hyperspectral image (HSI) classification tasks. It is necessary to explore the effectiveness of linguistic mode in assisting HSI classification. In addition, the large-scale pretraining image–text foundation models have demonstrated great performance in a variety of downstream applications, including zero-shot transfer. However, most domain generalization methods have never addressed mining linguistic modal knowledge to improve the generalization performance of model. To compensate for the inadequacies listed above, a language-aware domain generalization network (LDGnet) is proposed to learn cross-domain-invariant representation from cross-domain shared prior knowledge. The proposed method only trains on the source domain (SD) and then transfers the model to the target domain (TD). The dual-stream architecture including the image encoder and text encoder is used to extract visual and linguistic features, in which coarse-grained and fine-grained text representations are designed to extract two levels of linguistic features. Furthermore, linguistic features are used as cross-domain shared semantic space, and visual–linguistic alignment is completed by supervised contrastive learning in semantic space. Extensive experiments on three datasets demonstrate the superiority of the proposed method when compared with the state-of-the-art techniques. The codes will be available from the website:https://github.com/YuxiangZhang-BIT/IEEE_TGRS_LDGnet.
Yuxiang Zhang 0005, Mengmeng Zhang 0005, Wei Li 0032, Ran Tao 0003
IEEE Trans. Geosci. Remote. Sens.2
2023 Multiple Attention Network for Spartina alterniflora Segmentation Using Multitemporal Remote Sensing Images
abstract
The semantic segmentation of multi-temporal remote sensing images to construct wetland land surface coverage is the basis for the perception and dynamic modeling of geographic scenes. However, the segmentation of Spartina alterniflora (S.alterniflora) in remote sensing images on wetlands faces the problems such as low level for cooperative interpretation in multi-temporal images and high fragmentation in the distribution of S.alterniflora. To solve the issues, a multiple attention network (MARNet) based on transfer learning is proposed. The method is designed with a plug-and-play attention module to enhance the learning of vegetation features and improve the network’s ability to focus on small areas of S.alterniflora. At the same time, MARNet designs the transfer learning architecture from both inter-domain alignment and intra-domain adaptation perspectives,aligning the statistical distribution by using the maximum mean difference (MMD) between the source and target domains, and entropy minimization within the domain of the target domain to enhance the high confidence prediction of this domain. In addition, since the samples have a serious imbalance problem, redundant cutting and splicing steps are employed for the prediction results to prevent the poor edge prediction of some image blocks. Experimental results on three cross-year RSIs datasets demonstrate that the proposed MARNet performs significantly better than other networks and is able to extract S.alterniflora in wetlands more accurately.
Mengmeng Zhang 0005, Jianbu Wang, Xiukai Song, Yuanyuan Gui, Yuxiang Zhang 0005, Wei Li 0032
IEEE Trans. Geosci. Remote. Sens.2
2023 Microscopic Hyperspectral Image Classification Based on Fusion Transformer With Parallel CNN
abstract
Microscopic hyperspectral image (MHSI) has received considerable attention in the medical field. The wealthy spectral information provides potentially powerful identification ability when combining with advanced convolutional neural network (CNN). However, for high-dimensional MHSI, the local connection of CNN makes it difficult to extract the long-range dependencies of spectral bands. Transformer overcomes this problem well because of its self-attention mechanism. Nevertheless, transformer is inferior to CNN in extracting spatial detailed features. Therefore, a classification framework integrating transformer and CNN in parallel, named as Fusion Transformer (FUST), is proposed for MHSI classification tasks. Specifically, the transformer branch is employed to extract the overall semantics and capture the long-range dependencies of spectral bands to highlight the key spectral information. The parallel CNN branch is designed to extract significant multiscale spatial features. Furthermore, the feature fusion module is developed to effectively fuse and process the features extracted by the two branches. Experimental results on three MHSI datasets demonstrate that the proposed FUST achieves superior performance when compared with state-of-the-art methods.
Weijia Zeng, Wei Li 0032, Mengmeng Zhang 0005, Hao Wang 0122, Yue Yang 0041, Ran Tao 0003
IEEE J. Biomed. Health Informatics3
2023 Asymmetric Feature Fusion Network for Hyperspectral and SAR Image Classification
abstract
Joint classification using multisource remote sensing data for Earth observation is promising but challenging. Due to the gap of imaging mechanism and imbalanced information between multisource data, integrating the complementary merits for interpretation is still full of difficulties. In this article, a classification method based on asymmetric feature fusion, named asymmetric feature fusion network (AsyFFNet), is proposed. First, the weight-share residual blocks are utilized for feature extraction while keeping separate batch normalization (BN) layers. In the training phase, redundancy of the current channel is self-determined by the scaling factors in BN, which is replaced by another channel when the scaling factor is less than a threshold. To eliminate unnecessary channels and improve the generalization, a sparse constraint is imposed on partial scaling factors. Besides, a feature calibration module is designed to exploit the spatial dependence of multisource features, so that the discrimination capability is enhanced. Experimental results on the three datasets demonstrate that the proposed AsyFFNet significantly outperforms other competitive approaches.
Wei Li 0032, Yunhao Gao, Mengmeng Zhang 0005, Ran Tao 0003, Qian Du 0001
IEEE Trans. Neural Networks Learn. Syst.3
2023 Central Attention Network for Hyperspectral Imagery Classification
abstract
In this article, the intrinsic properties of hyperspectral imagery (HSI) are analyzed, and two principles for spectral-spatial feature extraction of HSI are built, including the foundation of pixel-level HSI classification and the definition of spatial information. Based on the two principles, scaled dot-product central attention (SDPCA) tailored for HSI is designed to extract spectral-spatial information from a central pixel (i.e., a query pixel to be classified) and pixels that are similar to the central pixel on an HSI patch. Then, employed with the HSI-tailored SDPCA module, a central attention network (CAN) is proposed by combining HSI-tailored dense connections of the features of the hidden layers and the spectral information of the query pixel. MiniCAN as a simplified version of CAN is also investigated. Superior classification performance of CAN and miniCAN on three datasets of different scenarios demonstrates their effectiveness and benefits compared with state-of-the-art methods.
Huan Liu 0015, Wei Li 0032, Xiang-Gen Xia 0001, Mengmeng Zhang 0005, Chenzhong Gao, Ran Tao 0003
IEEE Trans. Neural Networks Learn. Syst.4
2023 Hyperspectral and SAR Image Classification via Multiscale Interactive Fusion Network
abstract
Due to the limitations of single-source data, joint classification using multisource remote sensing data has received increasing attention. However, existing methods still have certain shortcomings when faced with feature extraction from single-source data and feature fusion between multisource data. In this article, a method based on multiscale interactive information extraction (MIFNet) for hyperspectral and synthetic aperture radar (SAR) image classification is proposed. First, a multiscale interactive information extraction (MIIE) block is designed to extract meaningful multiscale information. Compared with traditional multiscale models, it can not only obtain richer scale information but also reduce the model parameters and lower the network complexity. Furthermore, a global dependence fusion module (GDFM) is developed to fuse features from multisource data, which implements cross attention between multisource data from a global perspective and captures long-range dependence. Extensive experiments on the three datasets demonstrate the superiority of the proposed method and the necessity of each module for accuracy improvement.
Wei Li 0032, Yunhao Gao, Mengmeng Zhang 0005, Ran Tao 0003, Qian Du 0001
IEEE Trans. Neural Networks Learn. Syst.4
2023 Topological Structure and Semantic Information Transfer Network for Cross-Scene Hyperspectral Image Classification
abstract
Domain adaptation techniques have been widely applied to the problem of cross-scene hyperspectral image (HSI) classification. Most existing methods use convolutional neural networks (CNNs) to extract statistical features from data and often neglect the potential topological structure information between different land cover classes. CNN-based approaches generally only model the local spatial relationships of the samples, which largely limits their ability to capture the nonlocal topological relationship that would better represent the underlying data structure of HSI. In order to make up for the above shortcomings, a Topological structure and Semantic information Transfer network (TSTnet) is developed. The method employs the graph structure to characterize topological relationships and the graph convolutional network (GCN) that is good at processing for cross-scene HSI classification. In the proposed TSTnet, graph optimal transmission (GOT) is used to align topological relationships to assist distribution alignment between the source domain and the target domain based on the maximum mean difference (MMD). Furthermore, subgraphs from the source domain and the target domain are dynamically constructed based on CNN features to take advantage of the discriminative capacity of CNN models that, in turn, improve the robustness of classification. In addition, to better characterize the correlation between distribution alignment and topological relationship alignment, a consistency constraint is enforced to integrate the output of CNN and GCN. Experimental results on three cross-scene HSI datasets demonstrate that the proposed TSTnet performs significantly better than some state-of-the-art domain-adaptive approaches. The codes will be available from the website: https://github.com/YuxiangZhang-BIT/IEEE_TNNLS_TSTnet.
Yuxiang Zhang 0005, Wei Li 0032, Mengmeng Zhang 0005, Ying Qu 0001, Ran Tao 0003, Hairong Qi 0001
IEEE Trans. Neural Networks Learn. Syst.3
2022 Dual Graph Cross-Domain Few-Shot Learning for Hyperspectral Image Classification
abstract
Most domain adaptation (DA) methods focus on the case where the source data (SD) and target data (TD) with the same classes are obtained by the same sensor in cross-scene hyperspectral image (HSI) classification tasks. However, the classification performance is significantly reduced when there are new classes in TD. In addition, domain alignment is carried out based on local spatial information in most methods, rarely taking into account the non-local spatial information (non-local relationships) with strong correspondence. A Dual Graph Cross-domain Few-shot Learning (DG-CFSL) framework is proposed, trying to make up for the above shortcomings by combining Few-shot Learning (FSL) with domain alignment. Both SD with all label samples and TD with a few label samples are implemented for FSL episodic training. Meanwhile, Intra-domain Distribution Extraction block (IDE-block) is designed to characterize and aggregate the intra-domain non-local relationships. Furthermore, feature- and distribution-level cross-domain graph alignments are used to mitigate the impact of domain shift on FSL. Experimental results on two public HSI data sets demonstrate the effectiveness of the proposed method.
Yuxiang Zhang 0005, Wei Li 0032, Mengmeng Zhang 0005, Ran Tao 0003
ICASSP3
2022 Multisource Cross-Scene Classification Using Fractional Fusion and Spatial-Spectral Domain Adaptation
abstract
To solve the limitation of labeled samples in hyperspectral image (HSI) classification, cross-scene learning methods are developed recently. However, the disparity caused by environmental variation between HSI scenes is still a challenge. As a supplement, light detection and ranging (LiDAR) data provides elevation and spatial information regardless the variations. In this paper, we propose a multisource cross-scene classification method using fractional fusion and spatial-spectral domain adaptation to reduce disparity between scenes. The spatial information of HSI is preserved by fractional differential masks (FrDM) firstly. Then the LiDAR data is utilized for spectral alignment of HSI. The utilization of LiDAR data reduces the pixel-level disparity between scenes. At last, a spatial-spectral domain adaptation network is proposed for feature extraction and classification. Experimental results on HSI and LiDAR scenes show 5% improvements in overall accuracy compared with state-of-the-art methods.
Xudong Zhao 0003, Mengmeng Zhang 0005, Ran Tao 0003, Wei Li 0032, Wenzi Liao, Wilfried Philips
IGARSS2
2022 Multisource Remote Sensing Data Classification Using Fractional Fourier Transformer
abstract
Focusing on joint classification of Hyperspectral image (HSI) and Light detection and ranging (LiDAR) data, a fractional Fourier image transformer (FrIT) is proposed as a backbone network in this paper. In the proposed FrIT, HSI and LiDAR data are firstly fused at pixel-level. Both multi-source and HSI feature extractors are utilized to capture local contexts. Then, a plug-and-play image transformer FrIT is explored for global contexts and sequential feature extraction. Unlike the attention-based representations in classic visual image transformer (VIT), FrIT is capable of speeding up the transformer architectures massively. To reduce the information loss from shallow to deep layers, FrIT is devised to connect contextual features in multiple fractional domains. At last, to evaluate the performance of FrIT, a new HSI and LiDAR benchmark is provided for extensive experiments, on which the proposed FrIT gains an improvement of 3% over state-of-the-art methods.
Xudong Zhao 0003, Mengmeng Zhang 0005, Ran Tao 0003, Wei Li 0032, Wenzi Liao, Wilfried Philips
IGARSS2
2022 Multi-Source Remote Sensing Data Cross Scene Classification Based on Multi-Graph Matching
abstract
Multi-source joint classification has been extensively investigated in single scenario setting; however, for cross scene (CS) classification, few studies have been conducted for evaluating the collaborative performance of multi-sources. In this paper, using hyperspectral image (HSI) and light detection and ranging (LiDAR) data, we propose a multi-source CS classification method, and build source-related alignment to reduce statistical shift. Both geometrical and statistical alignments are considered to learn common-subspaces of each source with preserving discrimination information. Finally, the aligned features from both sources are integrated for final classification. Experimental results demonstrate the superior of the proposed method over other state-of-the-art CS approaches.
Mengmeng Zhang 0005, Xudong Zhao 0003, Wei Li 0032, Yuxiang Zhang 0005
IGARSS1
2022 Visible-Assisted Infrared Image Super-Resolution Based on Spatial Attention Residual Network
abstract
Infrared images have a wide range of applications in military and civilian fields, including night vision, surveillance, and robotics. However, the most commonly used infrared images are low-resolution (LR), which lack texture details, and existing infrared image super-resolution (SR) algorithms are limited by the lack of spatial information utilization. To solve the above problems, a spatial attention residual network (SAResNet) is proposed. Specifically, the network consists of spatial attention residual block (SARB) with several short skip connections (SSCs). The SARB contains 20 spatial attention blocks (SAB), which adaptively adjusts weights of different spatial regions by considering interdependence between spatial features. Meanwhile, the visible images are considered as complementary sources; thus, a visible-assisted training strategy is designed for the infrared SR process, promoting details preservation. Furthermore, the spatial attention (SA) mechanism is utilized, which focuses more on spatial characteristics of the image and refines the main objects and target boundaries. Experimentally, the proposed method, SAResNet, is compared with existing SR methods, and the effectiveness of the proposed method is demonstrated based on both quantity and quality analyses.
Mengmeng Zhang 0005, Wei Li 0032, Ran Tao 0003
IEEE Geosci. Remote. Sens. Lett.2
2022 Hyperspectral and Multispectral Classification for Coastal Wetland Using Depthwise Feature Interaction Network
abstract
The monitoring of coastal wetlands is of great importance to the protection of marine and terrestrial ecosystems. However, due to the complex environment, severe vegetation mixture, and difficulty of access, it is impossible to accurately classify coastal wetlands and identify their species with traditional classifiers. Despite the integration of multisource remote sensing data for performance enhancement, there are still challenges with acquiring and exploiting the complementary merits from multisource data. In this article, the depthwise feature interaction network (DFINet) is proposed for wetland classification. A depthwise cross attention module is designed to extract self-correlation and cross correlation from multisource feature pairs. In this way, meaningful complementary information is emphasized for classification. DFINet is optimized by coordinating consistency loss, discrimination loss, and classification loss. Accordingly, DFINet reaches the standard solution-space under the regularity of loss functions, while the spatial consistency and feature discrimination are preserved. Comprehensive experimental results on two hyperspectral and multispectral wetland datasets demonstrate that the proposed DFINet outperforms other competitive methods in terms of overall accuracy.
Yunhao Gao, Wei Li 0032, Mengmeng Zhang 0005, Jianbu Wang, Weiwei Sun 0005, Ran Tao 0003, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.3
2022 Graph-Feature-Enhanced Selective Assignment Network for Hyperspectral and Multispectral Data Classification
abstract
Due to rich spectral and spatial information, the combination of hyperspectral and multispectral images (MSIs) has been widely used for Earth observation, such as wetland classification. However, mining of meaningful features and effective fusion of multisource remote sensing data are still urgent problems to be solved. In this article, graph-feature-enhanced selective assignment network (GSANet) is proposed. On the one hand, a graph feature extraction module (GFEM) is designed to extract topological structure information and combine with the rich spectral–spatial information. In particular, the features obtained by convolution are first mapped to the graph feature space, and the graph convolution operation is used to achieve propagation between nodes for preserving topological structure information. Moreover, to reduce the difference of graph features resulting from the mapping function and better explore the complementary properties of multisource data, a novel graph fusion strategy-graph dependence fusion is designed. A transition graph is generated to enhance the association and interaction between different graph features, so as to avoid the information loss caused by simple fusion operation. On the other hand, a selective feature assignment module (SFAM) is developed to adaptively assign weights to different discriminative features. SFAM assigns weights to different features to selectively emphasize informative features and suppress less useful ones. Extensive experiments are conducted on two multisource remote sensing datasets, and the improvement of at least 1.27% and 0.98% compared to other state-of-the-art work demonstrates the superiority of the proposed GSANet.
Wei Li 0032, Yunhao Gao, Mengmeng Zhang 0005, Ran Tao 0003, Bing Zhang 0001
IEEE Trans. Geosci. Remote. Sens.4
2022 Information Fusion for Classification of Hyperspectral and LiDAR Data Using IP-CNN
abstract
Joint use of multisensor information has attracted considerable attention in the remote sensing community. While applications in land-cover observation benefit from information diversity, multisensor integration technique is confronted with many challenges, including inconsistent size of data, different data structures, uncorrelated physical properties, and scarcity of training data. In this article, an information fusion network, named interleaving perception convolutional neural network (IP-CNN), is proposed for integrating heterogeneous information and improving joint classification performance of hyperspectral image (HSI) and light detection and ranging (LiDAR) data. Specifically, a bidirectional autoencoder is designed to reconstruct hyperspectral and LiDAR data together, and the reconstruction process is trained with no dependence upon annotated information. Both HSI-perception constraint and LiDAR-perception constraint are imposed on multisource structural information integration. Accordingly, fused data are fed into a two-branch CNN for final classification. To validate the effectiveness of the model, the experiments were conducted using three datasets (i.e., Muufl Gulfport data, Trento data, and Houston data). The final results demonstrate that the proposed framework can significantly outperform state-of-the-art methods even with small-size training samples.
Mengmeng Zhang 0005, Wei Li 0032, Ran Tao 0003, Heng-Chao Li 0001, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.1
2021 Feature Exchange for Multisource Data Classification in Wetland Scene
abstract
Wetland classification is of great significance for monitoring. Recently, collaborative analysis of multisource data has received special attention considering the limitations of single source data. In this paper, a wetland classification method based on feature exchange is proposed. Firstly, the weighting shared residual blocks are utilized for feature extraction. Then, the scaling factors in batch normalization (BN) self-determine the redundancy of current channel, which is replaced by another channel when the scaling factor is less than the threshold. To eliminate unnecessary channels and improve the generalization, sparsity constraint is employed on partial scaling factors. Experimental results on multisource wetland dataset demonstrate that the proposed method outperforms other competitive works.
Yunhao Gao, Wei Li 0032, Mengmeng Zhang 0005, Ran Tao 0003
IGARSS3
2021 Woodland Segmentation of Gaofen-6 Remote Sensing Images Based on Deep Learning
abstract
Gaofen-6 (GF-6) is a geostationary, earth-observation satellite, rely on it's multi-spectral images, GF-6 has the ability to support the monitoring of woodland resources. In this paper, the multi-spectral images sent by GF-6 are studied as dataset, and a model called Infrared Attention Network (InfAttNet) which based on semantic segmentation method is proposed to distinguish woodland from other land types to achieve the purpose of woodland extraction. To make full use of the spectral information, InfAttNet has an additional encoder to extract the features of infrared bands independently. Besides, infrared attention blocks help InfAttNet to enhance the characteristics of woodland. The experimental results proved that InfAttNet improves the accuracy of woodland extraction, and the segmentation effect is strengthened compared with classical networks.
Yuanyuan Gui, Wei Li 0032, Mengmeng Zhang 0005, Anzhi Yue
IGARSS3
2021 Joint feature extraction for multi-source data using similar double-concentrated network
Yixuan Zhu, Wei Li 0032, Mengmeng Zhang 0005, Ran Tao 0003, Qian Du 0001
Neurocomputing3
2020 Convolutional Neural Network for Coastal Wetland Classification in Hyperspectral Image
abstract
Classifying different land cover types with hyperspectral image (HSI) is significant for restoring and protecting natural resources and maintaining ecological services in coastal wetlands. This paper proposes a multi-domain features fusion convolutional neural network (MDF-CNN) based classification method for hyperspectral images of coastal wetlands. This method adopts inter-class sparsity based discriminative least square regression (ICS_DLSR) to learn a more compact and discriminative transformation, as well as fuse the high-level features of the original domain and the regression domain to obtain higher classification accuracy. Experimental results demonstrate the effectiveness of the proposed method when compared with some recent classifiers. The MDF-CNN achieved state-of-the-art performance on two latest GF-5 HSI datasets of Coastal Wetland.
Mengmeng Zhang 0005, Wei Li 0032, Weiwei Sun 0005, Ran Tao 0003
IGARSS2
2020 Collaborative Classification for Woodland Data Using Similar Multi-concentrated Network
Yixuan Zhu, Mengmeng Zhang 0005, Wei Li 0032, Ran Tao 0003, Qiong Ran
PRCV (2)2
2020 Feature Extraction for Classification of Hyperspectral and LiDAR Data Using Patch-to-Patch CNN
abstract
Multisensor fusion is of great importance in Earth observation related applications. For instance, hyperspectral images (HSIs) provide wealthy spectral information while light detection and ranging (LiDAR) data provide elevation information, and using HSI and LiDAR data together can achieve better classification performance. In this paper, an unsupervised feature extraction framework, named as patch-to-patch convolutional neural network (PToP CNN), is proposed for collaborative classification of hyperspectral and LiDAR data. More specific, a three-tower PToP mapping is first developed to seek an accurate representation from HSI to LiDAR data, aiming at merging multiscale features between two different sources. Then, by integrating hidden layers of the designed PToP CNN, extracted features are expected to possess deeply fused characteristics. Accordingly, features from different hidden layers are concatenated into a stacked vector and fed into three fully connected layers. To verify the effectiveness of the proposed classification framework, experiments are executed on two benchmark remote sensing data sets. The experimental results demonstrate that the proposed method provides superior performance when compared with some state-of-the-art classifiers, such as two-branch CNN and context CNN.
Mengmeng Zhang 0005, Wei Li 0032, Qian Du 0001, Lianru Gao, Bing Zhang 0001
IEEE Trans. Cybern.1
2020 HSI-BERT: Hyperspectral Image Classification Using the Bidirectional Encoder Representation From Transformers
abstract
Deep learning methods have been widely used in hyperspectral image classification and have achieved state-of-the-art performance. Nonetheless, the existing deep learning methods are restricted by a limited receptive field, inflexibility, and difficult generalization problems in hyperspectral image classification. To solve these problems, we propose HSI-BERT, where BERT stands for bidirectional encoder representations from transformers and HSI stands for hyperspectral imagery. The proposed HSI-BERT has a global receptive field that captures the global dependence among pixels regardless of their spatial distance. HSI-BERT is very flexible and enables the flexible and dynamic input regions. Furthermore, HSI-BERT has good generalization ability because the jointly trained HSI-BERT can be generalized from regions with different shapes without retraining. HSI-BERT is primarily built on a multihead self-attention (MHSA) mechanism in an MHSA layer. Moreover, several attentions are learned by different heads, and each head of the MHSA layer encodes the semantic context-aware representation to obtain discriminative features. Because all head-encoded features are merged, the resulting features exhibit spatial-spectral information that is essential for accurate pixel-level classification. Quantitative and qualitative results demonstrate that HSI-BERT outperforms any other CNN-based model in terms of both classification accuracy and computational time and achieves state-of-the-art performance on three widely used hyperspectral image data sets.
Ji He 0003, Lina Zhao 0002, Mengmeng Zhang 0005, Wei Li 0032
IEEE Trans. Geosci. Remote. Sens.4
2019 Data Augmentation for Hyperspectral Image Classification With Deep CNN
abstract
Convolutional neural network (CNN) has been widely used in hyperspectral imagery (HSI) classification. Data augmentation is proven to be quite effective when training data size is relatively small. In this letter, extensive comparison experiments are conducted with common data augmentation methods, which draw an observation that common methods can produce a limited and up-bounded performance. To address this problem, a new data augmentation method, named as pixel-block pair (PBP), is proposed to greatly increase the number of training samples. The proposed method takes advantage of deep CNN to extract PBP features, and decision fusion is utilized for final label assignment. Experimental results demonstrate that the proposed method can outperform the existing ones.
Wei Li 0032, Chen Chen 0001, Mengmeng Zhang 0005, Heng-Chao Li 0001, Qian Du 0001
IEEE Geosci. Remote. Sens. Lett.3
2018 Joint Feature Extraction for Multispectral and Panchromatic Images Based on Convolutional Neural Network
abstract
Along with very high-resolution satellites were launched frequently, such as the satellite WorldView-3, panchromatic and multispectral remote-sensing images can be acquired easily. However, it is still an interesting and challenging task to fuse and classify these images. In general, panchromatic image has a high spatial resolution, but with only one spectral band. Multispectral image usually has four or eight bands, but the spatial resolution is four times smaller than panchromatic image. In this paper, an unsupervised feature extraction framework is proposed, which combines multispectral (MS) image and panchromatic (PAN) image into convolution neural network (CNN). There is an image-to-image mapping, learning from the input source (i.e., MS) to the output source (i.e., PAN). Then, by integrating the hidden layer of deep CNN, the extracted features represent MS and PAN data. The experimental results of two practical remote sensing data sets show the validity of the framework.
Mengmeng Zhang 0005, Wei Li 0032, Qian Du 0001
IGARSS2
2018 Nuclei Classification Using Dual View CNNs with Multi-crop Module in Histology Images
Xiang Li 0006, Wei Li 0032, Mengmeng Zhang 0005
PRCV (2)3