Jie Geng 0005

dblp:60/10617-5 · DBLP profile ↗
← Back
60ranked-venue papers
21as first author
44since 2021 · last 2026
0000-0003-4858-823XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 43 · 18 first-author · 27 since 2021Artificial intelligence and machine learning · 10 · 1 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Remote sensing semantic change detection via semantic editing-driven change synthesis
Zhengyi Xu, Wen Jiang 0002, Jie Geng 0005
Pattern Recognit.4
2026 Multimodal Cross-City Semantic Segmentation Based on Similarity-Inspired Fusion and Invertible Transformation Learning Network
abstract
Multimodal cross-city semantic segmentation aims to adapt a network trained on multiple labeled source domains (MSDs) from one city to multiple unlabeled target domains (MTDs) in another city, where the multiple domains refer to different sensor modalities. However, remote sensing data from different sensors increases the extent of domain shift in the fused domain space, making feature alignment more challenging. Meanwhile, traditional fusion methods only consider complementarity within MSDs (or MTDs), which wastes cross-domain relevant information and neglects control over domain shift. To address the above issues, we propose a similarity-inspired fusion and invertible transformation learning network (SFITNet) for multimodal cross-city semantic segmentation. To alleviate the increasing alignment difficulty in multimodal fused domains, an invertible transformation learning strategy (ITLS) is proposed, which adopts a topological perspective on unsupervised domain adaptation. This strategy aims to simulate the potential distribution transformation function between the MSD and the MTD based on invertible neural networks (INNs) after feature fusion, thereby performing distribution alignment independently within the two feature spaces. A cross-domain similarity-inspired information interaction module (CDSiM) is also designed, which considers the correspondence between the MSD and the MTD in the fusion stage, effectively utilizes multimodal complementary information and promotes the subsequent alignment of fused domain shifts. The semantic segmentation tests are completed on the public C2Seg-AB dataset and a new multimodal cross-city Su-Wu dataset. Compared with some state-of-the-art techniques, the experimental results demonstrated the superiority of the proposed SFITNet.
Lijia Dong, Wen Jiang 0002, Zhengyi Xu, Jie Geng 0005
IEEE Trans. Neural Networks Learn. Syst.4
2026 Unlocking Pseudolabel Potential and Alignment for Unpaired Cross-Modality Adaptation in Remote Sensing Image Segmentation
abstract
With the growth of multisource sensor technology, multimodal learning has become pivotal in remote sensing (RS) image segmentation. Despite its potential, current methods face challenges in acquiring large-scale paired samples. When annotated optical images are available, but synthetic aperture radar (SAR) images lack annotations, learning discriminative features for SAR images from optical images becomes difficult. Unsupervised domain adaptation (UDA) offers a potential solution to this challenge, which we refer to as unpaired cross-modality UDA. In this article, we propose unlocking pseudolabel potential and alignment (ULPA) for unpaired cross-modality adaptation in RS image segmentation, a novel one-stage adaptation framework designed to enhance cross-modality knowledge transfer. Our approach employs a prototypical multidomain alignment (PMDA) strategy, which reduces the modality gap through contrastive learning between features and prototypes of identical classes across different modalities. In addition, we introduce the unreliable-sample-guided feature contrast (UFC) loss to address the underutilization of unreliable pixels during training. This strategy separates reliable and unreliable pixels based on prediction confidence, assigning unreliable pixels to a category-wise queue of negative samples, thus ensuring all candidate pixels contribute to the training process. Extensive experiments show that the integration of PMDA and UFC loss can lead to more effective cross-modality domain alignment and substantially boost the model's generalization capability.
Zhe Xu 0016, Jie Geng 0005, Wen Jiang 0002, Shuai Song
IEEE Trans. Neural Networks Learn. Syst.2
2025 Masked auto-encoding and scatter-decoupling transformer for polarimetric SAR image classification
Jie Geng 0005, Lijia Dong, Yuhang Zhang 0033, Wen Jiang 0002
Pattern Recognit.1
2025 Hypergraph Matching Network for Semisupervised Few-Shot Scene Classification of Remote Sensing Images
abstract
Semi-supervised few-shot learning aims to alleviate the issue of insufficient labeled data with additional unlabeled samples. As for remote sensing images, complex contextual information leads to pseudo labeling with low confidence, which weakens the effect of semi-supervised few-shot classification. To solve these issues, a hypergraph matching network is proposed for semi-supervised few-shot scene classification of remote sensing images. Specifically, a hypergraph propagation module is designed to construct a hypergraph network, which can take advantage of adjacent samples with similar semantics and improve the representation ability of class prototypes. Then, a cross-layer prototype matching module is proposed to dynamically match features of different scales and angles, which aims to predict pseudo labels with high confidences. Experimental results on three public remote sensing datasets demonstrate that the proposed method can make effective utilization of additional unlabeled samples to enhance the classification performance of few-shot learning.
Jie Geng 0005, Bohan Xue, Wen Jiang 0002
IEEE Trans. Geosci. Remote. Sens.1
2025 Unsupervised Remote Sensing Image Semantic Segmentation Based on Multiscale Contrastive Domain Adaptation
abstract
Unsupervised domain adaptation for remote sensing image semantic segmentation aims to train a deep model on the labeled source domain and apply it to the unlabeled target domain. However, resolution and scene inconsistencies of cross-domain remote sensing images lead to great distribution differences, which weakens the semantic segmentation effect. To solve the above issues, an unsupervised remote sensing image semantic segmentation method is proposed based on multi-scale contrastive domain adaptation. Firstly, the mean teacher model is introduced into the unsupervised domain adaptation paradigm to generate pseudo-labels for target domain data, thereby achieving the cross-domain segmentation capability. A dynamic class balance sampling method is proposed to mitigate the class imbalance problem in cross-domain data by increasing the sampling frequency of the categories with fewer samples. Then, a data augmentation method called cross-domain mixup is developed to reduce the gap between the source and target domains. Finally, a multi-scale cross-domain contrastive loss is developed, which introduces the contrastive learning to learn domain-consistent features across the source and target domains, resulting in a more coherent and discriminative feature representation. Experimental results show that the proposed method can yield superior performance for unsupervised remote sensing image semantic segmentation.
Jie Geng 0005, Shuai Song, Zhe Xu 0016, Wen Jiang 0002
IEEE Trans. Geosci. Remote. Sens.1
2025 Semantic Segmentation of Remote Sensing Images With Inconsistent Resolutions via a Spectral-Geometric Iterative Fusion Network
abstract
The fusion of optical, hyperspectral, and Synthetic Aperture Radar (SAR) images is essential for semantic segmentation in remote sensing, enabling more comprehensive land cover classification through multimodal data integration. However, disparities in spatial resolution and imaging characteristics across modalities impede effective feature alignment and fusion, degrading segmentation performance. To address this problem, we propose a novel Spectral-Geometric Iterative Fusion Network (SGIFN), specifically designed to handle multimodal semantic segmentation with inconsistent resolutions. The core innovation of SGIFN lies in its unified architecture that progressively aligns, integrates, and enhances multimodal features through three newly designed modules. The Spectral-Spatial Iterative Decoupling (SSID) module introduces a novel iterative mechanism to adaptively align and decouple optical and hyperspectral features. The Spectral-Geometric Synergistic Conditional Random Field (SGS-CRF) module captures both local and long-range spatial dependencies by synergizing geometric (SAR) and spectral information. The Class-Guided Multiscale Contrastive Aggregation (CG-MCA) module further strengthens feature representation across scales via multi-class, contrastive learning. We constructed a new multimodal remote sensing dataset comprising optical, hyperspectral, and SAR images with varying resolutions collected from the Wuhan and Suzhou regions. Experimental results show that SGIFN achieves an mIoU of 69.61% on the Suzhou dataset and 63.96% on the Wuhan dataset. These results demonstrate the effectiveness of SGIFN in handling multimodal data with inconsistent resolutions.
Wenqi Han, Wen Jiang 0002, Jie Geng 0005, Yanchen Bao
IEEE Trans. Geosci. Remote. Sens.3
2025 Difference-Complementary Learning and Label Reassignment for Multimodal Semi-Supervised Semantic Segmentation of Remote Sensing Images
abstract
The feature fusion of optical and Synthetic Aperture Radar (SAR) images is widely used for semantic segmentation of multimodal remote sensing images. It leverages information from two different sensors to enhance the analytical capabilities of land cover. However, the imaging characteristics of optical and SAR data are vastly different, and noise interference makes the fusion of multimodal data information challenging. Furthermore, in practical remote sensing applications, there are typically only a limited number of labeled samples available, with most pixels needing to be labeled. Semi-supervised learning has the potential to improve model performance in scenarios with limited labeled data. However, in remote sensing applications, the quality of pseudo-labels is frequently compromised, particularly in challenging regions such as blurred edges and areas with class confusion. This degradation in label quality can have a detrimental effect on the model's overall performance. In this paper, we introduce the Difference-complementary Learning and Label Reassignment (DLLR) network for multimodal semi-supervised semantic segmentation of remote sensing images. Our proposed DLLR framework leverages asymmetric masking to create information discrepancies between the optical and SAR modalities, and employs a difference-guided complementary learning strategy to enable mutual learning. Subsequently, we introduce a multi-level label reassignment strategy, treating the label assignment problem as an optimal transport optimization task to allocate pixels to classes with higher precision for unlabeled pixels, thereby enhancing the quality of pseudo-label annotations. Finally, we introduce a multimodal consistency cross pseudo-supervision strategy to improve pseudo-label utilization. We evaluate our method on two multimodal remote sensing datasets, namely, the WHU-OPT-SAR and EErDS-OPT-SAR datasets. Experimental results demonstrate that our proposed DLLR model outperforms other relevant deep networks in terms of accuracy in multimodal semantic segmentation.
Wenqi Han, Wen Jiang 0002, Jie Geng 0005, Wang Miao
IEEE Trans. Image Process.3
2024 KN-RUE: Key Nodes based Resampling Uncertainty Estimation
abstract
With the continuous development and advancement of neural networks, in the application of neural networks, users not only require neural networks to be able to complete a given task but also want to know when they can trust the network’s prediction results and when they need to be cautious about the prediction results. In response to the need for uncertainty estimation of neural networks, many researchers have invested in the study of uncertainty estimation. Existing uncertainty evaluation methods are difficult to apply to deep neural networks with large parameter scales, complex internal structures, and mappings between inputs and outputs that are hard to express. This paper proposes a key nodes based resampling uncertainty estimation method ((KN-RUE), which achieves uncertainty estimation of prediction results for arbitrarily given large-scale neural networks. In this method, the first step involves analyzing the differences in feature space between adversarial and clean samples, identifying the main nodes affected by adversarial samples, and determining the critical nodes within the network. Next, by resampling the parameters of key nodes, the model is extended while ensuring model performance as much as possible, thus completing the measurement of uncertainty in prediction results. Through experiments, the effectiveness of the extended model and the superiority of uncertainty estimation performance in KN-RUE have been verified.
Xiang Li 0018, Wen Jiang 0002, Xinyang Deng, Jie Geng 0005
FUSION4
2024 Enhancing Multimodal Fusion with Only Unimodal Data
abstract
With recent advances in remote sensing technology, a wealth of multimodal data is available for applications. However, considering the domain differences between multimodal data and the alignment challenges in practical applications, it becomes important and challenging to integrate these data effectively. In this paper, we propose a multimodal prototype representation fusion network (MPRFN) for SAR and optical image fusion segmentation. Specifically, a more robust multimodal feature representation is provided by constructing multimodal category prototype representations that better capture the characteristics and distribution of each data. Meanwhile, a prototype-consistent semi-supervised learning method is proposed to improve the effectiveness of multimodal fusion semantic segmentation using a large number of unlabelled unimodal SAR images. Experiments on SAR and optical multimodal datasets show that the proposed method achieves state-of-the-art performance.
Wenqi Han, Jie Geng 0005, Xinyang Deng, Wen Jiang 0002
IGARSS2
2024 Remote Sensing Image Scene Classification With Multi-View Collaborative Representation Network
abstract
The utilization of deep learning methods in remote sensing image scene classification (RSISC) has gained significant attention, showcasing remarkable performance. However, these methods rely solely on the network for automatic weight assignment learning, which may introduce biases in attention calculations for remote sensing images. To address this issue, we propose a multi-view collaborative representation network (MCRNet) for RSISC. Specifically, we introduce a multiview collaborative representation framework (MCRF) to evaluate the impact of local features on key information within global features by different data augmentation. Furthermore, the introduction of a semantic summarization dictionary (SSD) aims to enhance the reconstruction of global semantic features through the optimization of a low-redundancy dictionary. Experiment results on two publicly available datasets confirm that the proposed model effectively improves the classification performance.
Wang Miao, Wen Jiang 0002, Jie Geng 0005
IGARSS3
2024 Pseudo-label meta-learner in semi-supervised few-shot learning for remote sensing image scene classification
Wang Miao, Zhe Xu 0016, Jie Geng 0005, Wen Jiang 0002
Appl. Intell.5
2024 Semi-supervised few-shot class-incremental learning based on dynamic topology evolution
Wenqi Han, Jie Geng 0005, Wen Jiang 0002
Eng. Appl. Artif. Intell.3
2024 A target intention recognition method based on information classification processing and information fusion
Zhuo Zhang 0005, Wen Jiang 0002, Jie Geng 0005
Eng. Appl. Artif. Intell.4
2024 Causal Intervention and Parameter-Free Reasoning for Few-Shot SAR Target Recognition
abstract
The target recognition of synthetic aperture radar (SAR) data generally faces the issue of limited observational samples in practical applications. Recent few-shot SAR target recognition techniques based on meta-learning, which mainly focus on intricate meta-learning models without considering SAR imaging characteristics during model training, show promise. To address this issue, a novel few-shot transfer learning paradigm named causal intervention and parameter-free reasoning (CIPR) is proposed for SAR target recognition. In the proposed framework, causal intervention pretraining (CIP), which emphasizes causal features of SAR images, is developed to diminish spurious correlations caused by confounders. Moreover, variational inference approximates intricate alterations in SAR imaging angles and background clutter in a generative manner. To make predictions of the unlabelled query set without additional learnable parameters, a parameter-free label reasoning model based on optimal transport, which integrates label knowledge and effectively leverages the distribution characteristics of causal features, is introduced. Experiments on the moving and stationary target acquisition and recognition (MSTAR) dataset demonstrate that the proposed method achieves superior performance and has preferable robustness to large depression angle discrepancies.
Jie Geng 0005, Weichen Ma, Wen Jiang 0002
IEEE Trans. Circuits Syst. Video Technol.1
2024 Dual-Path Feature Aware Network for Remote Sensing Image Semantic Segmentation
abstract
Semantic segmentation is a significant task for remote sensing interpretation, which takes advantage of contextual semantic information to classify each pixel into a specific category. Most current methods apply convolutional neural networks (CNN) to learn feature representation from remote sensing images, which may ignore the global dependencies due to the limitation of convolutional kernels. Inspired by the global feature learning ability of Transformer, we propose a novel deep model called dual-path feature aware network (DPFANet), which combines the structure of CNN and Transformer for semantic segmentation of remote sensing images. DPFANet aims to learn effective modeling ability from local to global features of images. Simultaneously, an adaptive feature fusion network is developed to fuse features from dual-path networks. Moreover, an edge optimization block is applied to constrain the edge features, whose purpose is to obtain more representative features for segmentation. Experimental results on three public remote sensing datasets verify that our proposed network yields better segmentation performance compared to other related methods.
Jie Geng 0005, Shuai Song, Wen Jiang 0002
IEEE Trans. Circuits Syst. Video Technol.1
2024 Spectral-Spatial Enhancement and Causal Constraint for Hyperspectral Image Cross-Scene Classification
abstract
Hyperspectral cross-scene classification refers to using only labeled data from the source domain in training and testing directly on the target domain dataset. However, there are differences between the reflection spectra of objects with the same category, which makes the cross-scene classification performance drop significantly. The task of single-domain generalization (SDG) has received extensive attention to solve the above problem. To address the discrepancy between source and target domains, a spatial-spectral enhancement and causal constraint network (S2ECNet) in terms of both data enhancement and causal alignment is proposed in this paper. To make up for the lack of data diversity in the source domain, a generator is created inS2ECNet to simulate the spectral deviation and spatial deviation from the target domain. A causal contribution discriminator is also built inS2ECNet to solve the data bias problem caused by direct feature alignment, which constructs causal contribution vectors from a causal perspective and uses contrastive learning to constrain category labels, extracting “potential” causal invariances from spectral and spatial domains. The cross-scene classification test is completed on the Pavia dataset, the HyRANK dataset, and the Houston dataset, and compared with some advanced multimodal methods. The experimental results demonstrate the effectiveness of the proposed network.
Lijia Dong, Jie Geng 0005, Wen Jiang 0002
IEEE Trans. Geosci. Remote. Sens.2
2024 Polarimetric SAR Image Classification Based on Hierarchical Scattering-Spatial Interaction Transformer
abstract
How to fully utilize the rich but complex scattering characteristics in PolSAR data is still a challenge. In this paper, a hierarchical scattering-spatial interaction Transformer (HSSIT) for polarimetric SAR image classification is proposed to effectively combine scattering and spatial characteristics of PolSAR data. The proposed HSSIT adopts a multi-stage hierarchical structure to extract discriminative features. Specifically, spatial feature extraction branch (SFEB) is designed to improve the global information perception ability for spatial features, which combines the advantages of CNN and Transformer to extract local features and capture context dependencies between pixels. A scatter-aware branch (SAB) based on Transformer is proposed to model correlation between polarimetric scattering features. Furthermore, we further propose a cross attention based information exchange module, which aggregates the tokens from two branches to enhance the discrimination of features for land cover classification. Sufficient experiments are carried out on three widely used PolSAR datasets to certify the effectiveness and superiority of our proposed method.
Jie Geng 0005, Yuhang Zhang 0033, Wen Jiang 0002
IEEE Trans. Geosci. Remote. Sens.1
2024 CMSE: Cross-Modal Semantic Enhancement Network for Classification of Hyperspectral and LiDAR Data
abstract
The fusion of hyperspectral imagery (HSI) and light detection and ranging (LiDAR) data is widely used for land cover classification. However, due to different imaging mechanisms, HSI and LiDAR data always present significant image differences, and the dimensions and feature distributions of HSI and LiDAR are highly dissimilar. This makes it challenging to represent and correlate semantic information from multimodal data. Current methods for classifying pixel-by-pixel features, which rely on cascaded or attention-based fusion, cannot effectively use multimodal features. To achieve accurate classification results, extracting and fusing similar high-order semantic information and complementary discriminative information contained in multimodal data is vital. In this paper, we propose a cross-modal semantic enhancement network (CMSE) for multimodal semantic information mining and fusion. Our proposed CMSE framework extracts features from the image on multiple scales, capturing more representative local sparse features with different sizes of convolution kernels. To represent high-level semantic features related to land cover, we establish a Gaussian-weighted matrix and semantically transform the spatial and spectral features of distinct branches. Finally, we build a multi-level residual fusion module to incrementally fuse spectral features from HSI and elevation features from LiDAR. Additionally, we introduce a cross-modal semantically constrained loss to guide multimodal semantic feature alignment. We evaluate our approach on three multimodal remote sensing datasets, namely the Houston2013, Trento, and MUUFL datasets. The experimental results demonstrate that our proposed CMSE model achieves superior performance in terms of accuracy and robustness compared to other related deep networks.
Wenqi Han, Wang Miao, Jie Geng 0005, Wen Jiang 0002
IEEE Trans. Geosci. Remote. Sens.3
2024 Hierarchical Feature Progressive Alignment Network for Remote Sensing Image Scene Classification in Multitarget Domain Adaptation
abstract
Multitarget domain adaptation (MTDA) presents a formidable challenge in remote sensing image scene classification (RSICS), where the objective is to transfer knowledge from a labeled source domain to several unlabeled target domains. Compared to single-source-single-target domain adaptation (S3TDA), MTDA is inherently more complex due to domain shifts among multiple target domains. Directly merging the unique features of multitarget domains can result in corrupted information and poor classification performance. To address these challenges, we propose a hierarchical feature progressive alignment network (HFPAN) for RSICS in MTDA. First, our method introduces a fine-grained and contextual information extraction (FCIE) network to extract the global-local correlation in remote sensing (RS) images. Second, we construct a hierarchical feature embedding (HFE) framework that maintains hierarchical inter, intra constraints for the extracted features. Finally, we perform an alignment process for the constructed hierarchical features to minimize the differences in MTDA, progressing from coarse to fine granularity. To evaluate the efficacy of our proposed method, we conducted several cross-domain scene classification experiments on five public datasets. These experiments demonstrate the novelty of our approach and its ability to achieve improved classification performance.
Wang Miao, Wenqi Han, Jie Geng 0005, Wen Jiang 0002
IEEE Trans. Geosci. Remote. Sens.3
2024 Multidomain Constrained Translation Network for Change Detection in Heterogeneous Remote Sensing Images
abstract
In heterogeneous image change detection (HICD), preventing neural networks from distorting critical information is the main challenge of such methods based on deep translation. And most of these methods rely on a priori information to suppress the effects of changed pixels in the translation process, but the accuracy of the prior information will influence the results of translation. In this paper, we propose an end-to-end multi-domain constrained translation network (MDCTNet) for unsupervised HICD. The proposed MDCTNet utilizes an improved generative adversarial network (GAN) to generate target domain images realistically. Furthermore, to retain the content information of the source domain images, MDCTNet leverages contrastive learning to ensure the consistency of adjacent pixel relationships. Meanwhile, it employs high-frequency information consistency which preserves pivotal characteristics. We compare the proposed MDCTNet with state-of-the-art algorithms to verify the efficacy of the proposed technique. The experimental results on five real data sets demonstrate the effectiveness of the proposed method.
Jie Geng 0005, Wen Jiang 0002
IEEE Trans. Geosci. Remote. Sens.2
2024 Texture-Aware Causal Feature Extraction Network for Multimodal Remote Sensing Data Classification
abstract
The pixel-level classification of multimodal remote sensing images plays a crucial role in the intelligent interpretation of remote sensing data. However, existing methods that mainly focus on feature interaction and fusion fail to address the challenges posed by confounders – brought by sensor imaging bias, limiting their performance. In this paper, we introduce causal inference into intelligent interpretation of remote sensing and propose a new Texture-Aware Causal Feature Extraction Network (TeACFNet) for pixel-level fusion classification. Specifically, we propose a two-stage causal feature extraction framework that helps networks learn more explicit class representations by capturing the causal relationships between multimodal heterogeneous data. In addition, to solve the problem of low-resolution land cover feature representation in remote sensing images, we propose the Refined Statistical Texture Extraction (ReSTE) module. This module integrates the semantics of statistical textures in shallow feature maps through feature refinement, quantization, and encoding. Extensive experiments on two publicly available datasets with different modalities, including Houston 2013 and Berlin datasets, demonstrate the remarkable effectiveness of our proposed method, which reaches a new state-of-the-art.
Zhengyi Xu, Wen Jiang 0002, Jie Geng 0005
IEEE Trans. Geosci. Remote. Sens.3
2024 A New Data Augmentation Method Based on Mixup and Dempster-Shafer Theory
abstract
To improve the performance of deep neural networks, the Mixup method has been proposed to alleviate their memorization issues and sensitivity to adversarial samples. This provides networks with better generalization abilities. The learning principle of Mixup is essentially to train deep neural networks for regularization tasks with a convex combination of the original feature vectors and their labels. However, soft labels are generated directly using the mixing ratio without dealing with the uncertain information generated during the mixing process. Therefore, this paper proposes a new data augmentation method based on Mixup and Dempster-Shafer theory called DS-Mixup, which is a regularizer that can express and deal with the uncertainty caused by ambiguity. This method uses interval numbers to generate mass functions of mixed samples to model the distribution of set-valued random variables; then, ambiguous decision spaces are constructed, and soft labels with single-element subsets and multielement subsets are generated to further improve the delineation of decision boundaries during the training process. In addition, an evidence neural network with DS-Mixup is designed in this paper to accomplish recognition or classification tasks. Experimental results obtained on multimedia datasets, including attribute, image, text and signal data, show that the proposed method achieves more effective data augmentation effects and further improves the performance of deep neural networks.
Zhuo Zhang 0005, Jie Geng 0005, Xinyang Deng, Wen Jiang 0002
IEEE Trans. Multim.3
2023 Hyperspectral and LiDAR Data Classification Using Spatial Context and De-Redundant Fusion Network
abstract
The utilization of multi-modal data from multi-sensors (e.g., hyperspectral and light detection ranging (LiDAR) data) to classify ground objects has been an important topic in remote sensing interpretation. However, complex background leads to the difficulty of extracting context relations, at the same time, redundancy and noise among multi-modal data bring great challenges to accurate classification. In this paper, we propose a novel spatial context and de-redundant fusion network (SCDNet) to fuse hyperspectral and LiDAR data for land cover classification. Specifically, a multi-scale attention fusion module is developed in the feature extraction stage, which adaptively fuses global and local information of different scales to obtain a more accurate spatial context. In the feature fusion stage, a fusion module based on gated mechanism is proposed, which can remove the redundant information of multi-mode data and obtain discriminative fusion features. We design a series of comparisons and ablation experiments on the Houston2013 dataset and Trento dataset, and the results demonstrate the effectiveness of the proposed method.
Lijia Dong, Wen Jiang 0002, Jie Geng 0005
IEEE Geosci. Remote. Sens. Lett.3
2023 DDFN: Deblurring Dictionary Encoding Fusion Network for Infrared and Visible Image Object Detection
abstract
Both infrared and visible images have advantages for object detection, since infrared images can capture thermal characteristics of objects and visible images can provide high spatial resolution and clear texture details of objects. Combining infrared and visible images for object detection has many advantages, but how to fully utilize the inherent characteristics of these two data is still a challenging issue. To address this issue, a deblurring dictionary encoding fusion network (DDFN) is proposed for infrared and visible image object detection. Firstly, a dual-stream feature extraction backbone is structured, which aims to learn features based on the characteristics of different modalities. Then, pooling operations are applied to filter out key information and reduce the complexity of the network. Afterwards, a fuzzy compensation module is proposed, which aims to minimize the information loss of pooling process. Finally, a dictionary encoding fusion module is proposed to robustly excavate potential interactions between infrared and visible images, which can obtain fusion features with aggregating the local information of infrared features and the long-term dependent information of visible features. The proposed DDFN exhibits excellent performance on two benchmark bimodal datasets and shows superior capabilities in object detection of infrared-visible images.
Jiawei Lai, Jie Geng 0005, Xinyang Deng, Wen Jiang 0002
IEEE Geosci. Remote. Sens. Lett.2
2023 SFE-FN: A Shuffle Feature Enhancement-Based Fusion Network for Hyperspectral and LiDAR Classification
abstract
Under the background of the rapid development of remote sensing (RS) technology, multimodal RS image classification has attracted great attention. Considerable research has been devoted to designing more adequate multimodal feature-level fusion networks. However, few have noted that in the process of feature fusion, if the multimodal heterogeneous features are quite different, direct fusion may introduce noise. This greatly affects the classification performance of the network. This letter proposes a shuffle feature enhancement-based fusion network (SFE-FN) for hyperspectral and light detection and ranging (LiDAR) classification, which effectively alleviates the aforementioned problems. Specifically, first, an SFE module is proposed to achieve self-enhancement and mutual enhancement of each modal feature to preliminary reduce the feature difference. Then, a cross-layer and cross-interaction module (CLCI) is designed to further enhance the consistency of features by updating parameters across layers. Finally, the proposed shuffle feature concatenation (SFC) module and the shuffle feature fusion (SFF) module are utilized to adequately merge fewer differentiated features. Experiments on Houston2013 and Trento datasets show that the proposed method is effective.
Xinxin Shen, Xinyang Deng, Jie Geng 0005, Wen Jiang 0002
IEEE Geosci. Remote. Sens. Lett.3
2023 Foreground-Background Contrastive Learning for Few-Shot Remote Sensing Image Scene Classification
abstract
Few-shot learning (FSL) aims to train a model with limited samples for identifying novel category samples. As for remote sensing images, complex backgrounds may lead to large intra-class differences, and the number of labeled samples is quite smaller than that of large datasets, which both influence the classification performance. To solve these issues, a foreground-background contrastive learning (FBCL) is proposed for few-shot remote sensing image scene classification. Specifically, a foreground-background separation module is proposed to separate features between objects and background with supervised contrastive learning, which aims to improve the ability to distinguish foreground and background regions of remote sensing images. Moreover, a channel weight allocator is proposed to balance features of different dimensions, which can take full advantage of remote sensing image information. Experiments on three remote sensing datasets prove that the proposed few-shot method is able to produce superior classification results than other related approaches.
Jie Geng 0005, Bohan Xue, Wen Jiang 0002
IEEE Trans. Geosci. Remote. Sens.1
2023 Multigranularity Decoupling Network With Pseudolabel Selection for Remote Sensing Image Scene Classification
abstract
The existing deep networks have shown excellent performance in remote sensing scene classification (RSSC), which generally requires a large amount of class-balanced training samples. However, deep networks will result in underfitting with imbalanced training samples since they can easily bias toward the majority classes. To address these problems, a multigranularity decoupling network (MGDNet) is proposed for remote sensing image scene classification. To begin with, we design a multigranularity complementary feature representation (MGCFR) method to extract fine-grained features from remote sensing images, which utilizes region-level supervision to guide the attention of the decoupling network. Second, a class-imbalanced pseudolabel selection (CIPS) approach is proposed to evaluate the credibility of unlabeled samples. Finally, the diversity component feature (DCF) loss function is developed to force the local features to be more discriminative. Our model performs satisfactorily on three public datasets: UC Merced (UCM), NWPU-RESISC45, and Aerial Image Dataset (AID). Experimental results show that the proposed model yields superior performance compared with other state-of-the-art methods.
Wang Miao, Jie Geng 0005, Wen Jiang 0002
IEEE Trans. Geosci. Remote. Sens.2
2023 ECAE: Edge-Aware Class Activation Enhancement for Semisupervised Remote Sensing Image Semantic Segmentation
abstract
Remote sensing image semantic segmentation (RSISS) remains challenging due to the scarcity of labeled data. Semi-supervised learning can leverage pseudo-labels to enhance the model’s ability to learn from unlabeled data. However, accurately generating pseudo-labels for RSISS remains a significant challenge that severely affects the model’s performance, especially for the edges of different classes. In order to overcome these issues, we propose a semi-supervised semantic segmentation framework for remote sensing images based on edge-aware class activation enhancement (ECAE). Firstly, the baseline network is constructed based on the average teacher model, which separates the training of labeled and unlabeled data using student and teacher networks. Secondly, considering local continuity and global discreteness of object distribution in remote sensing images, the class activation mapping enhancement (CAME) network is designed to predict local areas more remarkably. Finally, the edge-aware network (EAN) is proposed to improve the performance of edge segmentation in remote sensing images. The combination of the CAME with the EAN further heightens the generation of high-confidence pseudo-labels. Experiments were performed on two publicly available remote sensing semantic segmentation datasets, Potsdam and ISPRS Vaihingen, which verify the superiorities of the proposed ECAE model.
Wang Miao, Zhe Xu 0016, Jie Geng 0005, Wen Jiang 0002
IEEE Trans. Geosci. Remote. Sens.3
2023 MMT: Mixed-Mask Transformer for Remote Sensing Image Semantic Segmentation
abstract
Remote sensing image semantic segmentation is a crucial step in the intelligent interpretation of remote sensing. Most of the current approaches are based on the attention mechanism to enhance long-range representations. However, these works ignore the key problem of foreground-background imbalance, and their performances encounter a bottleneck. In this paper, we introduce mask classification into remote sensing image interpretation for the first time, and propose a novel mixed-mask Transformer (MMT) for remote sensing image semantic segmentation. Specifically, we propose a mixed-mask attention mechanism, a simple but effective module, which assists the network to learn more explicit intraclass and interclass correlations by capturing long-range interdependent representations. In addition, a progressive multi-scale learning strategy is proposed to solve the problem of large scale-varied targets in remote sensing images, which integrates semantic and visual representations of different scale targets by efficiently utilizing large scale feature maps in Transformer. Experimental results show that the proposed MMT exceeds the existing alternative approaches and achieves state-of-the-art performance on three semantic segmentation datasets.
Zhe Xu 0016, Jie Geng 0005, Wen Jiang 0002
IEEE Trans. Geosci. Remote. Sens.2
2023 Dual-Branch Dynamic Modulation Network for Hyperspectral and LiDAR Data Classification
abstract
Deep learning algorithms that can effectively extract features from different modalities have achieved significant performance in multimodal remote sensing (RS) data classification. However, we actually found that the feature representation of one modality is likely to affect other modalities through parameter back-propagation. Even if multimodal models are superior to their uni-modal counterparts, they are likely to be underutilized. To solve the above issue, a dual-branch dynamic modulation network is proposed for hyperspectral (HS) and light detection and ranging (LiDAR) data classification. Firstly, a novel dynamic multimodal gradient optimization (DMGO) strategy is proposed to control the gradient modulation of each feature extraction branch adaptively. Then, a multimodal bi-directional enhancement (MBE) module is developed to integrate features of different modalities, which aims to enhance the complementarity of HS and LiDAR data. Furthermore, a feature distribution consistency (FDC) loss function is designed to quantify similarities between integrated features and dominant features, which can improve the consistency of features across modalities. Experimental evaluations on Houston2013 and Trento datasets demonstrate that our proposed network exceeds several state-of-the-art multimodal classification methods in terms of fusion classification performance.
Zhengyi Xu, Wen Jiang 0002, Jie Geng 0005
IEEE Trans. Geosci. Remote. Sens.3
2022 An information fusion method based on deep learning and fuzzy discount-weighting for target intention recognition
Zhuo Zhang 0005, Jie Geng 0005, Wen Jiang 0002, Xinyang Deng, Wang Miao
Eng. Appl. Artif. Intell.3
2022 Rotated Object Detection of Remote Sensing Image Based on Binary Smooth Encoding and Ellipse-Like Focus Loss
abstract
Remote sensing image object detection has been widely developed in many applications. Objects in remote sensing data have the characteristic of arbitrary directions, which leads to poor detection performance based on horizontal box detectors. To address this issue, a novel rotated object detection model based on binary smooth encoding and ellipse-like focus loss is proposed in this paper. Firstly, a multi-layer feature fusion network with attention mechanism is developed to extract features of multi-scale objects. Then, an anchor free detection module with binary smooth encoding is proposed, which aims to predict the rotated angles of objects. Moreover, an ellipse-like focus loss is proposed to obtain high-quality bounding boxes drawing near the object center. Experimental results on two public remote sensing datasets verify that the proposed method can yield superior detection performance than other related rotated object detection models.
Jie Geng 0005, Zhe Xu 0016, Zihao Zhao 0009, Wen Jiang 0002
IEEE Geosci. Remote. Sens. Lett.1
2022 IDLN: Iterative Distribution Learning Network for Few-Shot Remote Sensing Image Scene Classification
abstract
Few-shot remote sensing image scene classification has gained more attention due to its ability to recognize novel categories with several annotated samples. However, it is a great challenge to extract category characteristics with insufficient labeled samples. To address this issue, an iterative distribution learning network (IDLN) is proposed for the scene classification of few-shot remote sensing images. Specifically, the proposed model is a cyclic iterative architecture, which is composed of three modules to enhance the classification performance. In each iteration, similarity distribution learning module is proposed to calculate feature relations among instances first, then label mapping module is developed as the few-shot classifier based on instructive knowledge, and, finally, the attention-based feature calibration module is proposed to modify features based on label relations, and the calibrated features are imported to the next iteration. Experimental results on two public remote sensing datasets demonstrate that the proposed network is able to achieve superior performance for few-shot remote sensing image scene classification.
Qingjie Zeng, Jie Geng 0005, Wen Jiang 0002
IEEE Geosci. Remote. Sens. Lett.2
2022 Polarimetric SAR Image Classification Based on Feature Enhanced Superpixel Hypergraph Neural Network
abstract
Synthetic aperture radar (SAR) images can capture abundant spatial and polarimetric information of land cover objects, and thus polarimetric SAR (PolSAR) image classification has been developed for various applications. Combining the advantages of spatial and polarimetric information simultaneously is of great importance for PolSAR image classification. In this paper, a feature enhanced superpixel hypergraph neural network (FESHNN) is proposed for PolSAR image classification, which aims to take full advantage of spatial features and polarimetric features from PolSAR images. In the proposed model, superpixel hypergraph neural network is constructed for feature representation of superpixels, which aims to obtain spatial correlation and polarimetric correlation in a hypergraph. Then, a feature enhancement module is employed to refine the local features of pixels and the spatial features of superpixels, which aims to enhance the discrimination of feature representation. Experimental results on three PolSAR datasets demonstrate that the proposed method yields superior classification performance compared with other related approaches.
Jie Geng 0005, Wen Jiang 0002
IEEE Trans. Geosci. Remote. Sens.1
2022 Semi-Supervised Remote-Sensing Image Scene Classification Using Representation Consistency Siamese Network
abstract
Deep learning has achieved excellent performance in remote-sensing image scene classification, since a large number of datasets with annotations can be applied for training. However, in actual applications, there is just a few annotated samples and a large number of unannotated samples in remote-sensing images, which leads to overfitting of the deep model and affects the performance of scene classification. In order to address these problems, a semi-supervised representation consistency Siamese network (SS-RCSN) is proposed for remote-sensing image scene classification. First, considering intraclass diversity and interclass similarity of remote-sensing images, Involution-generative adversarial network (GAN) is utilized to extract the discriminative features from remote-sensing images via unsupervised learning. Then, Siamese network with a representation consistency loss is proposed for semi-supervised classification, which aims to reduce the differences of labeled and unlabeled data. Experimental results on UC Merced dataset, RESICS-45 dataset, aerial image dataset (AID), and RS dataset demonstrate that our method yields superior classification performance compared with other semi-supervised learning (SSL) methods.
Wang Miao, Jie Geng 0005, Wen Jiang 0002
IEEE Trans. Geosci. Remote. Sens.2
2021 Relation-Aware Neighborhood Aggregation for Cross-lingual Entity Alignment
Yuanna Liu, Jie Geng 0005, Xinyang Deng, Wen Jiang 0002
FUSION2
2021 Pseudo-loss Confidence Metric for Semi-supervised Few-shot Learning
abstract
Semi-supervised few-shot learning is developed to train a classifier that can adapt to new tasks with limited labeled data and a fixed quantity of unlabeled data. Most semi-supervised few-shot learning methods select pseudo-labeled data of unlabeled set by task-specific confidence estimation. This work presents a task-unified confidence estimation approach for semi-supervised few-shot learning, named pseudo-loss confidence metric (PLCM). It measures the data credibility by the loss distribution of pseudo-labels, which is synthetical considered multi-tasks. Specifically, pseudo-labeled data of different tasks are mapped to a unified metric space by mean of the pseudo-loss model, making it possible to learn the prior pseudo-loss distribution. Then, confidence of pseudo-labeled data is estimated according to the distribution component confidence of its pseudo-loss. Thus highly reliable pseudo-labeled data are selected to strengthen the classifier. Moreover, to overcome the pseudo-loss distribution shift and improve the effectiveness of classifier, we advance the multi-step training strategy coordinated with the class balance measures of class-apart selection and class weight. Experimental results on four popular benchmark datasets demonstrate that the proposed approach can effectively select pseudo-labeled data and achieve the state-of-the-art performance.
Jie Geng 0005, Wen Jiang 0002, Xinyang Deng, Zhe Xu 0016
ICCV2
2021 Self-Attention and Mutual-Attention for Few-Shot Hyperspectral Image Classification
abstract
Few-shot classification of hyperspectral image (HSI) has been increasingly abstracted attention due to its superiority of adopting to new HSI classification with only a few labeled data available. However, insufficient feature expression still bothers the improvement of performance. To address this issue, a deep self-attention and mutual-attention few-shot learning (SMA-FSL) method is proposed for HSI few-shot classification. Specifically, a deep 3D convolutional feature embedding network is utilized to extract the spectral-spatial feature at first. Then, self-attention and mutual-attention are used to ally the feature of different samples with same class and expand the class prototypes for more stable feature expression. Finally, the predicted results are obtained by calculating the distance between query set and aligned class prototypes. The experimental results on two well-know HSI datasets demonstrate that the proposed method achieves better performance compared with other related methods.
Xinyang Deng, Jie Geng 0005, Wen Jiang 0002
IGARSS3
2021 A Semi-Supervised Siamese Network with Label Fusion for Remote Sensing Image Scene Classification
abstract
Remote sensing image scene classification, which requires large amounts of labeled data, plays a critical role in a range of fields.However, in the actual complex environment, the obtained remote sensing images are sometimes unlabeled due to data perturbation and the cost of manual labeling, which limits the training effect and generalization ability. To solve this issue, a semi-supervised siamese network with label fusion is proposed for remote sensing image scene classification. The siamese network is developed to extract features from remote sensing image, where loss function based on the low-entropy principle is constructed to select the unlabeled data as pseudo-label samples. The labeled and pseudo-label samples are mixed to further train the siamese network. The results on UC Merced dataset and WHU-RS19 show that our method is capable to achieve excellent performance compared with other semi-supervised learning methods.
Wang Miao, Jie Geng 0005, Xinyang Deng, Wen Jiang 0002
IGARSS2
2021 Triplet Attention Feature Fusion Network for SAR and Optical Image Land Cover Classification
abstract
With recent advances in remote sensing, abundant multimodal data are available for applications. However, considering the redundancy and the huge domain differences among multimodal data, how to effectively integrate these data is becoming important and challenging. In this paper, we proposed a triplet attention feature fusion network (TAFFN) for SAR and optical image fusion classification. Specifically, spatial attention module and spectral attention module based on self-attention mechanism are developed to extract spatial and spectral long-range information from the SAR image and optical image respectively, at the same time, cross-attention mechanism is proposed to capture the long-range interactive representation. Triplet attentions are concatenated to further integrate the complementary information of SAR and optical images. Experiments on a SAR and optical multimodal dataset demonstrate that the proposed method can achieve the state-of-the-arts performance.
Zhe Xu 0016, Jinbiao Zhu, Jie Geng 0005, Xinyang Deng, Wen Jiang 0002
IGARSS3
2021 A new belief divergence measure for Dempster-Shafer theory based on belief and plausibility function and its application in multi-source data fusion
Xinyang Deng, Wen Jiang 0002, Jie Geng 0005
Eng. Appl. Artif. Intell.4
2021 Multi-Scale Metric Learning for Few-Shot Learning
abstract
Few-shot learning in image classification is developed to learn a model that aims to identify unseen classes with only few training samples for each class. Fewer training samples and new tasks of classification make many traditional classification models no longer applicable. In this paper, a novel few-shot learning method named multi-scale metric learning (MSML) is proposed to extract multi-scale features and learn the multi-scale relations between samples for the classification of few-shot learning. In the proposed method, a feature pyramid structure is introduced for multi-scale feature embedding, which aims to combine high-level strong semantic features with low-level but abundant visual features. Then a multi-scale relation generation network (MRGN) is developed for hierarchical metric learning, in which high-level features are corresponding to deeper metric learning while low-level features are corresponding to lighter metric learning. Moreover, a novel loss function named intra-class and inter-class relation loss (IIRL) is proposed to optimize the proposed deep network, which aims to strengthen the correlation between homogeneous groups of samples and weaken the correlation between heterogeneous groups of samples. Experimental results on mini ImageNet and tiered ImageNet demonstrate that the proposed method achieves superior performance in few-shot learning problem.
Wen Jiang 0002, Jie Geng 0005, Xinyang Deng
IEEE Trans. Circuits Syst. Video Technol.3
2021 Cross-Dataset Hyperspectral Image Classification Based on Adversarial Domain Adaptation
abstract
The cross-data set knowledge is vital for hyperspectral image classification, which can reduce the dependence on the sample quantity by transferring knowledge from other data sets and improve the training efficiency by sharing knowledge between different data sets. However, due to the capturing environment change and imaging equipment difference, domain shift troubles the exploitation of the cross-data set knowledge. To address the aforementioned issue, this article proposes an unsupervised cross-data set hyperspectral image classification method based on adversarial domain adaptation. The proposed method, which employs multiple classifiers to build a discriminator and uses variational autoencoders to constitute a generator, works in an adversarial manner to drive the target samples under the support of the source domain. In particular, the classification error and the classification disagreement are considered in the objective function, which helps to align different domains while keeping the boundaries of different classes. Experimental results of the multidomain data set demonstrate that the proposed method can transfer and share cross-data set knowledge and achieve state-of-the-art performance without using the labeled information of the target data set.
Xiaorui Ma, Xuerong Mou, Jie Wang 0003, Jie Geng 0005, Hongyu Wang 0001
IEEE Trans. Geosci. Remote. Sens.5
2020 Transfer Learning for SAR Image Classification Via Deep Joint Distribution Adaptation Networks
abstract
The problem of different characters of heterogeneous synthetic aperture radar (SAR) images leads to poor performances for transfer learning of SAR image classification. To address this issue, a semisupervised model named as deep joint distribution adaptation networks (DJDANs) is proposed for transfer learning from a source SAR image to a different but similar target SAR image, which aims to match the joint probability distributions between the source domain and target domain. In the proposed DJDAN, a marginal distribution adaptation network is developed to map features across the domains into an augmented common feature subspace, which aims to match the marginal probability distributions and unify the dimensions. Then, a conditional distribution adaptation network is proposed to transfer knowledge across the domains, which aims to reduce the discrepancies of the conditional probability distributions and enhance the effectiveness of feature representation. Moreover, one-versus-rest classification is utilized in the proposed framework, which aims to improve the discrimination between the inside and outside class. Experimental results demonstrate the effectiveness of the proposed deep networks.
Jie Geng 0005, Xinyang Deng, Xiaorui Ma, Wen Jiang 0002
IEEE Trans. Geosci. Remote. Sens.1
2019 Cross-Scene Hyperspectral Image Classification Based on Deep Conditional Distribution Adaptation Networks
abstract
Cross-scene classification of hyperspectral image (HSI) has been increasingly researched due to its crucial utilization in practical applications. However, cross-scene data generally perform distribution discrepancy, which hampers the transfer learning performance. To address this issue, deep conditional distribution adaptation networks (DCDAN) are proposed for HSI cross-scene classification, which aim to reduce the distribution shift between a source domain and a target domain. The proposed deep network adopts a conditional constraint to match the class conditional distributions across domains, where a great number of training samples from the source domain and a small number of training samples from the target domain are utilized to train the deep model. Cross-scene classification results on two HSIs demonstrate that the proposed network is able to yield superior performance compared with some related methods.
Jie Geng 0005, Xiaorui Ma, Wen Jiang 0002, Dawei Wang 0001, Hongyu Wang 0001
IGARSS1
2019 Hyperspectral Image Classification by Parameters Prediction Networks
abstract
Hyperspectral image, which contains high-resolution spectral information as well as large-scale spatial information, has been widely used in various classification applications of remote sensing area. However, due to the insufficient of the labeled samples in the training set and the unbalance of sample quantity between different classes, traditional supervised classification methods are difficult to achieve satisfying performance. In order to address above issues, this paper studies on how to predict classification parameters more effectively, and finds out that the parameters of the fully-connected layer in the classifier are closely related to the output of the feature mapping layer. Based on above fact, this paper proposes a hyperspectral image classification method base on parameter prediction network, which adapts a pre-trained neural network to novel categories by directly predicting the parameters of classifier from the feature data of the hyperspectral image. Experimental results and analysis demonstrate the competitive performance of the proposed method over other state-of-the-art classification methods based on neural network when the number of labeled samples is very small.
Sheng Ji, Xiaorui Ma, Jie Geng 0005, Hongyu Wang 0001
IGARSS5
2019 Saliency-Guided Deep Neural Networks for SAR Image Change Detection
abstract
Change detection is an important task to identify land-cover changes between the acquisitions at different times. For synthetic aperture radar (SAR) images, inherent speckle noise of the images can lead to false changed points, which affects the change detection performance. Besides, the supervised classifier in change detection framework requires numerous training samples, which are generally obtained by manual labeling. In this paper, a novel unsupervised method named saliency-guided deep neural networks (SGDNNs) is proposed for SAR image change detection. In the proposed method, to weaken the influence of speckle noise, a salient region that probably belongs to the changed object is extracted from the difference image. To obtain pseudotraining samples automatically, hierarchical fuzzy C-means (HFCM) clustering is developed to select samples with higher probabilities to be changed and unchanged. Moreover, to enhance the discrimination of sample features, DNNs based on the nonnegative- and Fisher-constrained autoencoder are applied for final detection. Experimental results on five real SAR data sets demonstrate the effectiveness of the proposed approach.
Jie Geng 0005, Xiaorui Ma, Hongyu Wang 0001
IEEE Trans. Geosci. Remote. Sens.1
2019 Hyperspectral Image Classification Based on Two-Phase Relation Learning Network
abstract
Deep learning-based classification methods are competent to achieve an excellent performance under one necessary condition, i.e., there are sufficient labeled samples in each class, which is extremely impractical in most of the remote sensing tasks. To improve the performance with small training sets, we resort to other hyperspectral images and design a two-phase relation learning network that can be transferred between different images for general information sharing and fine-trained on a specific hyperspectral image for individual information learning. Specifically, we use a relation learning method to compare samples and deal with the task inconsistency between different data sets, and we adopt an episode-based training strategy to mimic the testing setup and learn the transferable comparison ability. Benefited from these two strategies, the proposed network takes the advantage of extra knowledge for information supplement and learns to compare rather than to classify for information exploration, which guarantees a reasonable performance even with small training sets. Extensive experiments and analysis on three benchmarks demonstrate that the proposed method can provide an effective solution for hyperspectral image classification with small training sets, which makes it possible to work on large-scale applications of earth observation with less effort on field investigation.
Xiaorui Ma, Sheng Ji, Jie Wang 0003, Jie Geng 0005, Hongyu Wang 0001
IEEE Trans. Geosci. Remote. Sens.4
2018 Semisupervised Classification of Polarimetric SAR Image via Superpixel Restrained Deep Neural Network
abstract
The classification of polarimetric synthetic aperture radar (PolSAR) image is of crucial significance for SAR applications. In this letter, a superpixel restrained deep neural network with multiple decisions (SRDNN-MDs) is proposed for PolSAR image classification, which not only extracts effective superpixel spatial features and degrades the influence of speckle noises but also deals with the limited training samples. First, the polarimetric features of coherency matrix and Yamaguchi decomposition are extracted as initial features, and superpixel segmentation is conducted on the Pauli color-coded image to acquire the superpixel averaged features. Then, an SRDNN based on sparse autoencoders is proposed to capture superpixel correlative features and reduce speckle noises. After that, MDs, including nonlocal decision and local decision, are developed to select credible testing samples. Finally, our deep network is updated by the extended training set to yield the final classification map. Experimental results demonstrate that the proposed SRDNN-MD yields higher accuracies compared with other related approaches, which indicate that the proposed method is able to capture superpixel correlative information and adds the information of unlabeled samples to improve the classification performance.
Jie Geng 0005, Xiaorui Ma, Jianchao Fan, Hongyu Wang 0001
IEEE Geosci. Remote. Sens. Lett.1
2018 SAR Image Classification via Deep Recurrent Encoding Neural Networks
abstract
Synthetic aperture radar (SAR) image classification is a fundamental process for SAR image understanding and interpretation. With the advancement of imaging techniques, it permits to produce higher resolution SAR data and extend data amount. Therefore, intelligent algorithms for high-resolution SAR image classification are demanded. Inspired by deep learning technology, an end-to-end classification model from the original SAR image to final classification map is developed to automatically extract features and conduct classification, which is named deep recurrent encoding neural networks (DRENNs). In our proposed framework, a spatial feature learning network based on long-short-term memory (LSTM) is developed to extract contextual dependencies of SAR images, where 2-D image patches are transformed into 1-D sequences and imported into LSTM to learn the latent spatial correlations. After LSTM, nonnegative and Fisher constrained autoencoders (NFCAEs) are proposed to improve the discrimination of features and conduct final classification, where nonnegative constraint and Fisher constraint are developed in each autoencoder to restrict the training of the network. The whole DRENN not only combines the spatial feature learning power of LSTM but also utilizes the discriminative representation ability of our NFCAE to improve the classification performance. The experimental results tested on three SAR images demonstrate that the proposed DRENN is able to learn effective feature representations from SAR images and produce competitive classification accuracies to other related approaches.
Jie Geng 0005, Hongyu Wang 0001, Jianchao Fan, Xiaorui Ma
IEEE Trans. Geosci. Remote. Sens.1
2017 Change detection of marine reclamation using multispectral images via patch-based recurrent neural network
abstract
Marine reclamation plays an increasingly important role in expanding living space, which should be monitored to ensure legitimate development. In this paper, a patch-based recurrent neural network is developed for change detection of marine reclamation. To capture spatial difference of image patches in two images, a patch-based recurrent neural network is proposed to extract features, where patches from two multispectral images are stacked as a sequence for inputting. After training the deep network, Softmax classifier is applied to detect the changed region. It is illustrated that our network can obtain the difference of two images to improve detection accuracies. Experiments on the study area of the Jinzhou Bay demonstrate that the proposed method outperforms other approaches.
Jie Geng 0005, Jianchao Fan, Hongyu Wang 0001, Xiaorui Ma
IGARSS1
2017 Classification of fusing SAR and multispectral image via deep bimodal autoencoders
abstract
Classification of multisensor data provides potential advantages over a single sensor in accuracy. In this paper, deep bimodal autoencoders are proposed for classification of fusing synthetic aperture radar (SAR) and multispectral images. The proposed deep network based on autoencoders is trained to discover both independencies of each modality and correlations across the modalities. Specifically, the sparse encoding layers in the front are applied to learn features of each modality, then shared representation layers in the middle are developed to learn fused features of two modalities, finally softmax classifier in the top is adopted for classification. Experimental results demonstrate that the proposed network is able to yield superior classification performance compared with some related networks.
Jie Geng 0005, Hongyu Wang 0001, Jianchao Fan, Xiaorui Ma
IGARSS1
2017 Weighted Fusion-Based Representation Classifiers for Marine Floating Raft Detection of SAR Images
abstract
Detection of a marine floating raft is significant for ocean utilization, which provides a basis for marine ecosystem protection. In this case study, supervised classifiers of weighted fusion-based representation are proposed to detect marine floating raft using synthetic aperture radar images. To remove the speckle noise and obtain more discriminative features, a weighted low-rank matrix factorization (WLRMF) model is developed to optimize features before detection, where the matrix of patch features is decomposed to acquire the denoised features. Weighted fusion-based representation classifiers (WFRCs) with weighted multiplication are proposed to combine the sparse representation classifier (SRC) and the collaborative representation classifier (CRC) for floating raft detection, which can capture the competition between the floating raft and water surface as well as the collaboration within-class samples. Experiments on the study area of the Bohai Sea confirm that the proposed approach produces better results than some related methods. It is demonstrated that the WLRMF model extracts effective features and overcomes the influence of speckle noise at the same time, and the WFRC model is able to take advantages of the SRC in competition and CRC in collaboration for improving detection accuracies.
Jie Geng 0005, Jianchao Fan, Hongyu Wang 0001
IEEE Geosci. Remote. Sens. Lett.1
2017 Deep Supervised and Contractive Neural Network for SAR Image Classification
abstract
The classification of a synthetic aperture radar (SAR) image is a significant yet challenging task, due to the presence of speckle noises and the absence of effective feature representation. Inspired by deep learning technology, a novel deep supervised and contractive neural network (DSCNN) for SAR image classification is proposed to overcome these problems. In order to extract spatial features, a multiscale patch-based feature extraction model that consists of gray level-gradient co-occurrence matrix, Gabor, and histogram of oriented gradient descriptors is developed to obtain primitive features from the SAR image. Then, to get discriminative representation of initial features, the DSCNN network that comprises four layers of supervised and contractive autoencoders is proposed to optimize features for classification. The supervised penalty of the DSCNN can capture the relevant information between features and labels, and the contractive restriction aims to enhance the locally invariant and robustness of the encoding representation. Consequently, the DSCNN is able to produce effective representation of sample features and provide superb predictions of the class labels. Moreover, to restrain the influence of speckle noises, a graph-cut-based spatial regularization is adopted after classification to suppress misclassified pixels and smooth the results. Experiments on three SAR data sets demonstrate that the proposed method is able to yield superior classification performance compared with some related approaches.
Jie Geng 0005, Hongyu Wang 0001, Jianchao Fan, Xiaorui Ma
IEEE Trans. Geosci. Remote. Sens.1
2016 An iterative low-rank representation for SAR image despeckling
abstract
Speckle noises are inherent issues in synthetic aperture radar (SAR) images, which hampers the analysis and interpretation of SAR images. In this paper, we propose an iterative low-rank representation algorithm for SAR image despeckling. The original SAR image is first transformed to the logarithmic image, which is then filtered iteratively by the proposed low-rank representation model. Specifically, in each iteration, similar patches measured by the Mahalanobis distance are collected into a group, and then filtered by the nuclear regularized low-rank representation. Finally, all of the filtered patches are aggregated to form the denoised image. Experimental results demonstrate that the proposed algorithm is able to yield state-of-the-art SAR image despeckling performance.
Jie Geng 0005, Jianchao Fan, Xiaorui Ma, Hongyu Wang 0001
IGARSS1
2016 Joint collaborative representation for polarimetric SAR image classification
abstract
Polarimetric synthetic aperture radar (PolSAR) images are widely applied in terrain and ground cover classification. Feature extraction and classifier design are both important in Pol- SAR image classification. In this paper, various target decompositions are applied to obtain different polarimetric features. Since that neighboring pixels usually belong to the same species, they can be simultaneously represented through linear combinations of training samples. Therefore, a collaborative representation-based classifier with spatially joint regularization is adopted for classification. Experimental results demonstrate that the joint collaborative representation model performs better than other state-of-the-art methods, such as support vector machine and simultaneous sparse representation.
Jie Geng 0005, Jianchao Fan, Hongyu Wang 0001, Anyan Fu
IGARSS1
2016 Hyperspectral image classification with small training set by deep network and relative distance prior
abstract
This paper presents a hyperspectral image classification method based on deep network, which has shown great potential in various machine learning tasks. Since the quantity of training samples is the primary restriction of the performance of classification methods, we impose a new prior on the deep network to deal with the instability of parameter estimation under this circumstances. On the one hand, the proposed method adjusts parameters of the whole network to minimize the classification error as all supervised deep learning algorithm, on the other hand, unlike others, it also minimize the discrepancy within each class and maximize the difference between different classes. The experimental results showed that the proposed method is able to achieve great performance under small training set.
Xiaorui Ma, Hongyu Wang 0001, Jie Geng 0005, Jie Wang 0003
IGARSS3
2015 Floating raft aquaculture information automatic extraction based on high resolution SAR images
abstract
Floating raft aquaculture is an important part of the coastal marine environment monitoring. In order to achieve the accurate monitoring on the range and area of floating raft, combined with on-site underway survey, adopt the high resolution SAR satellite remote sensing data to conduct floating raft aquaculture information extraction. Choosing Beidaihe and its adjacent fields as a key demonstration of floating raft aquaculture study, verify that the proposed joint sparse representation classification approach can quickly and accurately obtain the floating raft aquaculture range and area.
Jianchao Fan, Jialan Chu, Jie Geng 0005, Fengshou Zhang
IGARSS3
2015 High-Resolution SAR Image Classification via Deep Convolutional Autoencoders
abstract
Synthetic aperture radar (SAR) image classification is a hot topic in the interpretation of SAR images. However, the absence of effective feature representation and the presence of speckle noise in SAR images make classification difficult to handle. In order to overcome these problems, a deep convolutional autoencoder (DCAE) is proposed to extract features and conduct classification automatically. The deep network is composed of eight layers: a convolutional layer to extract texture features, a scale transformation layer to aggregate neighbor information, four layers based on sparse autoencoders to optimize features and classify, and last two layers for postprocessing. Compared with hand-crafted features, the DCAE network provides an automatic method to learn discriminative features from the image. A series of filters is designed as convolutional units to comprise the gray-level cooccurrence matrix and Gabor features together. Scale transformation is conducted to reduce the influence of the noise, which integrates the correlated neighbor pixels. Sparse autoencoders seek better representation of features to match the classifier, since training labels are added to fine-tune the parameters of the networks. Morphological smoothing removes the isolated points of the classification map. The whole network is designed ingeniously, and each part has a contribution to the classification accuracy. The experiments of TerraSAR-X image demonstrate that the DCAE network can extract efficient features and perform better classification result compared with some related algorithms.
Jie Geng 0005, Jianchao Fan, Hongyu Wang 0001, Xiaorui Ma, Baoming Li, Fuliang Chen
IEEE Geosci. Remote. Sens. Lett.1