Fang Liu 0034

dblp:67/5807-34 · DBLP profile ↗
← Back
51ranked-venue papers
12as first author
46since 2021 · last 2026
0000-0001-9752-9530ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 40 · 10 first-author · 38 since 2021Artificial intelligence and machine learning · 9 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021
YearPublicationVenuePosition
2026 Enhancing few-shot segmentation via mask combination learning
Shuo Li 0010, Fang Liu 0034, Licheng Jiao, Xuejian Gou, Lingling Li 0002, Xu Liu 0006, Puhua Chen
Neurocomputing2
2026 Composite fractal scanning enhanced mamba for small object detection in remote sensing imagery
Wenjing Zhan, Yongke Li, Fang Liu 0034, Liang Xiao 0001
Neurocomputing3
2026 Concept-Aware Learning for Weakly Supervised Video Anomaly Detection
Shuo Li 0010, Fang Liu 0034, Licheng Jiao, Jiahao Wang 0002, Xu Liu 0006, Lingling Li 0002, Puhua Chen
Pattern Recognit.2
2026 Deep semi-supervised relation preserving learning model
Chenxi Tian, Lingling Li 0002, Xu Liu 0006, Licheng Jiao, Fang Liu 0034, Shuyuan Yang 0001
Pattern Recognit.5
2026 D&D-Net: A diffusion and deep priors regularized network for hyperspectral reconstruction
Jingxiang Yang, Tian Lin 0001, Wenxiu Diao, Fang Liu 0034, Jia Liu 0020, Hongyi Liu 0001, Liang Xiao 0001
Signal Process.4
2025 Semantic-Assisted Feature Integration Network for Multilabel Remote Sensing Scene Classification
abstract
With remote sensing (RS) images’ resolution increasing, a single scene label cannot adequately represent RS scenes’ contents. Therefore, multilabel RS scene classification (MLRSSC) is gradually attracting the researchers’ attention. Many methods have been proposed recently, and most use deep features or semantic connections to complete MLRSSC. However, they ignore the combination of these two aspects. In addition, the high interclass similarity and low intraclass similarity of RS images limit the robustness of these methods. In this article, we propose a semantic-assisted feature integration network (SFIN) to overcome the above limitations. It contains a dual-scale feature extractor module (DFEM), a local semantic enhance module (LSEM), a cross-scale interactive attention module (CIAM), and a classifier module (CM). DFEM utilizes the convolutional neural networks (CNNs) to extract multiscale features from RS images. LSEM extracts semantic information and establishes their relationships at different scales. CIAM enhances the feature representation by interacting with the clues across different scales. CM completes the prediction of classification (CLA) results. Integrating them into an end-to-end framework, SFIN can discover the diverse and complex land covers hidden in RS images. Furthermore, to ensure the accuracy of explored semantics and enhance the SFIN’s feature extraction ability, we design a semantic supervision (SS) loss and a semantic-based contrastive learning (SB-CL) loss. They are in charge of the correctness and discrimination of the mined semantics. Along with the typical CLA loss, SFIN can be adequately trained. Extensive experiments have been conducted on four MLRSSC datasets, and the positive results demonstrate that SFIN outperforms many existing methods in MLRSSC tasks. Our source codes are available at:https://github.com/TangXu-Group/multilabelRSSC/tree/main/SFIN.
Ruiqi Du, Xu Tang 0004, Jingjing Ma 0001, Xiangrong Zhang, Fang Liu 0034, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.5
2025 Multimodal Feature Interactive Learning for Few-Shot Hyperspectral Image Classification
abstract
Recently auxiliary cross-scene information has been widely utilized to improve the hypersperctral image classification performance by knowledge transfer. However, recognition of different objects with the same semantic category is difficult when the object types in similar scenes are different or only limited similarity knowledge is provided. In this paper, a multi-modal feature interactive learning (MMFI) method is proposed based on both hyperspectral image modality and textual modality to distinguish similar objects, which enhances the transfer capability by utilizing the semantic prior from the textual modality. First, the adversarial domain mapping (ADM) module is designed to realize cross-domain knowledge transfer across different scenes in an adversarial learning manner. In particular, the noise is simulated as data distribution in different domains through domain mapping and aggregated with source and target domain data, which is then reconstructed and optimized to learn discriminative and conducive information for transfer. Then, the adaptive interactive learning (AIL) module acts on the latent features of the encoder to mine latent associations among the aggregated features and facilitate the expression of consistent features. In addition, few-shot learning with textual embedding enables more powerful semantic priors for few-shot prototypes, making up for insufficient recognition capability in the presence of hyperspectral image modality only. Experimental results on three datasets demonstrate the superiority of our method.
Fang Liu 0034, Wenfei Gao, Jia Liu 0020, Xu Tang 0004, Liang Xiao 0001
IEEE Trans. Geosci. Remote. Sens.1
2025 Edge-Object Co-Driven Learning for Remote Sensing Change Detection
abstract
Remote sensing change detection (CD) aims to accurately reveal surface changes by comparing two temporally separated images of the same area. However, in complex environments, insufficient edge detail recognition and limited feature extraction often affect the accuracy of CD. For this purpose, we propose a novel method named the edge-object co-driven learning network (EOCLNet), which employs a combination of the Pyramid Vision Transformer (PVT) and the Fast Segment Anything Model (FastSAM) as parallel feature extractors to capture rich multilevel features. Specifically, it includes three key components which are the edge extraction module (EEM), the object revelation module (ORM), and the edge-object learning (EOL). EEM explicitly captures edge details by combining low-level spatial features with high-level semantic features, providing essential edge knowledge. ORM reveals changed objects by aggregating the highest two levels of semantic features, providing initial change guidance. EOL is designed to implicitly mine edge clues by establishing relationships between edges and changed objects across multiple levels, receiving outputs from both EEM and ORM. Furthermore, during the training process, the uncertainty from the previous level’s change map is utilized to guide the learning at the next level, thereby achieving a transition from uncertainty to certainty. The effectiveness of EOCLNet is validated on three public datasets, where it outperforms several state-of-the-art CD methods.
Yangguang Liu, Fang Liu 0034, Jia Liu 0020, Liang Xiao 0001
IEEE Trans. Geosci. Remote. Sens.2
2025 Box2Change: A Novel Weakly Supervised Way for Change Detection via Consistency Instance Segmentation
abstract
Change detection in remote sensing images aims at revealing interesting changes about the earth surface and has been one of the most important issues in earth observation. In recent years, lots of fully-supervised change detection methods have achieved good performance with the help of deep learning architectures, which rely on large amounts of pixel-level labels. However, obtaining high-quality pixel-level labels is laborious and expensive. To alleviate this problem, we propose a novel weakly-supervised change detection way via consistency instance segmentation called Box2Change, which requires only box-level labels and achieves competitive results to fully-supervised change detection method. Compared with pixel-level label, it is much more efficient to get box-level label, which locates the potential changed area by a rectangle box. There are two key components in the proposed method, the Changed Instance Segmentation (CIS) and the Self-Supervised Consistency Learning (SSCL) in affine space. The former generates multi-scale changed instances, which learns positional information from box-level labels and segments the instance boundaries within a given bounded region. The latter introduces affine transform and employs consistency constraints in a self-supervised manner to increases the robustness to pseudo-change situations caused by light or noise. In experiments, three popular public change detection datasets are tested and both visual and numerical assessment are discussed, where the proposed method exhibits competitive performance to fully-supervised methods and achieves the state-of-the-art results compared with the other weakly-supervised change detection methods.
Fang Liu 0034, Kanghua Yin, Jia Liu 0020, Jingxiang Yang, Xu Tang 0004, Liang Xiao 0001
IEEE Trans. Geosci. Remote. Sens.1
2025 Background-Driven and Foreground-Refined Network for Weakly Supervised Change Detection
abstract
Change detection (CD) in remote sensing aims to reveal meaningful surface changes and has been flourishing in recent years. Compared with fully-supervised methods based on pixel-level labels, image-level labels are easy to acquire, which reduces manual labor to a large extent. However, image-level labels lack spatial-and-shape information while containing the least semantic information, which poses a great challenge to the weakly-supervised CD task. Motivated by the prior that bi-temporal images have background semantic consistency, we propose Background-Driven and Foreground-Refined (BDFR-Net) to ameliorate the above problem. Specifically, there are two key components in the proposed method: the Background-Driven Reconstruction (BDR) with image-level supervision and the Foreground-Refined Learning (FRL) with affinity learning. The former generates changed regions of foreground and background separation, which activates the foreground from image-level supervision and constrains the foreground by maintaining spatial and semantic consistency in background regions. The latter introduces Complementary Fusion and Label Adaption (CFLA) strategies to further refine the foreground, which can mine complementary information from foreground sequences and suppress false activations. In addition, affinity learning is proposed to stabilize and supervise the above process. Complementary relationships between foreground and background are fully utilized. Tested on two popular CD datasets, the results demonstrate that our proposed BDFR-Net produces completely changed regions with clear boundaries and outperforms state-of-the-art weakly-supervised methods.
Fang Liu 0034, Jia Liu 0020, Jingxiang Yang, Xu Tang 0004, Liang Xiao 0001
IEEE Trans. Geosci. Remote. Sens.1
2025 Multiscale Sparse Cross-Attention Network for Remote Sensing Scene Classification
abstract
Remote sensing (RS) scene classification (RSSC) is a prominent research topic in the RS community. Multilevel feature fusion is an important way of addressing RS scene classification, and many methods have been proposed in recent years. Although they succeed, current methods can still be improved, particularly in distinguishing the contributions of different multilevel features and fully and effectively fusing them. To address the above issues and fully exploit the potential of multilevel features for RS scene classification tasks, we propose a new model named multiscale sparse cross-attention network (MSCN). It not only focuses on the effectiveness of feature learning but also emphasizes the rationality of feature fusion. In detail, MSCN first extracts multilevel features using a pre-trained ResNet50. Also, these features are divided into high- and low-level features according to the clues they involved. Then, a multiscale sparse cross-attention (MSC) module is developed to cross-fuse the high-level feature with various low-level features, thereby effectively mining helpful information from multilevel features. In the fusion process, MSC not only explores the multiscale messages in RS scenes but also mitigates the negative impact of irrelevant information by employing sparse operations. Third, a group convolutional block attention module (CBAM) enhancer (GCE) is presented to enhance the representation of classification features. GCE detects local salient information within classification features using grouped CBAM and further enhances crucial details by readjusting the CBAM attention weights. This way, the classification features’ discrimination can be improved. We conducted extensive experiments on three public RS scene classification datasets. The exceptional experimental results indicate that our proposed MSCN achieves superior classification accuracy, surpassing many existing methods. Our source codes are available at https://github.com/TangXu-Group/Remote-Sensing-Images-Classification/tree/main/MSCN.
Jingjing Ma 0001, Xu Tang 0004, Xiangrong Zhang, Fang Liu 0034, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.5
2025 Text-Driven Adaptive Semantic Alignment Network for Cross-Scene Hyperspectral Image Classification
abstract
Land cover in different scenes generally exhibits scene-invariant category semantic, typically represented and described consistently in a textual modality. Traditional cross-scene classification methods often treat categories as discrete class labels, neglecting their semantic information, or use category names merely as auxiliary textual modalities to enhance the discriminative representations of land cover. However, the cross-scene consistency of category semantic for land cover remains underexplored and underutilized. To address this issue, the text-driven adaptive semantic alignment network (TASA-Net) is proposed in this article for cross-scene hyperspectral image classification (HSIC). TASA-Net employs hand-crafted template prompts for stable category descriptions and vision-guided fine semantic prompts (VG-FSPs) for dynamic scene adaptation. Through a dual-gated adaptive mechanism, TASA-Net optimally weights coarse- and fine-grained semantics in a shared space, ensuring stable yet discriminative semantic representation. Additionally, cross-modal semantic alignment projects visual features into the shared semantic space, while a soft alignment strategy dynamically adjusts category correlations to enhance intraclass consistency and mitigate domain shifts. Ultimately, by leveraging text-driven semantic consistency representation, TASA-Net achieves zero-shot cross-scene transfer for unsupervised classification. Experiments demonstrate superior performance across multiple hyperspectral datasets, validating the critical role of textual modality in enhancing model robustness and cross-scene generalization ability.
Wenzhen Wang, Fang Liu 0034, Hongyuan Zhu 0002, Liang Xiao 0001
IEEE Trans. Geosci. Remote. Sens.2
2025 Cross-Modal Remote Sensing Image-Text Retrieval via Context and Uncertainty-Aware Prompt
abstract
The cross-modal remote sensing image-text retrieval (CMRSITR) is a lively research topic in the remote sensing (RS) community. Benefiting from the large pretrained image-text models, many successful CMRSITR methods have been proposed in recent years. Although their performance is attractive, there are still some challenges. First, fine-tuning large pretrained models requires a significant amount of computational resources. Second, most large models are pretrained by natural images, which reduces their effectiveness in processing RS images. To tackle these challenges, we propose a new CMRSITR network named context and uncertainty-aware prompt (CUP). First, prompt tuning theory is introduced into CUP to eliminate the burden of optimization resources. By training the prompt tokens rather than all parameters, the large model's knowledge can be transferred to CMRSITR tasks with small trainable parameters. Second, considering the differences between natural-image-based prior clues and RS images, apart from adopting the free-prompt tokens, we develop a prompt generation module (PGM) to produce the RS-oriented prompt tokens. The specific prompt tokens are rich in object-level messages of RS images, which help CUP narrow the gaps between natural large models and RS images. Third, we further design an uncertainty estimation module (UEM) to whittle down the uncertainties caused by the model and data. This way, can not only the semantic misalignment and intraclass diversity imbalance problems be mitigated but also the RS clues can be deeply explored. Competitive experimental results counted on three public benchmark datasets demonstrate that our CUP can achieve competitive performance in the CMRSITR task compared with many existing methods. Our source codes are available at: https://github.com/TangXu-Group/Cross-modal-remote-sensing-image-and-text-retrieval-models/tree/main/CUP.
Yijing Wang 0004, Xu Tang 0004, Jingjing Ma 0001, Xiangrong Zhang, Fang Liu 0034, Licheng Jiao
IEEE Trans. Neural Networks Learn. Syst.5
2024 A Cross-Modal Semantic Mapping Enhancement Model for Remote Sensing Visual Question Answering
abstract
Remote sensing visual question answering (RSVQA) aims to answer the questions based on the content in remote sensing (RS) images. Due to the complexity of RS images, it is challenging to focus on regions relevant to the questions in the RS images. To this end, we propose a channel-selective multi-scale cross-attention (CSCa) model for RSVQA tasks. Specifically, we design a text-driven multi-scale feature extractor to extract question-related features in RS images. To obtain the cross-attention map in this extractor, we design a novel channel selection mechanism to capture more accurate question-related regions in RS images and develop a channel-wise contrastive learning task to align the semantics between image and text features. We set up experiments on RSVQA-LR and RSIVQA datasets. Experiment results show that our CSCa achieves excellent performance.
Dabiao Huang, Xu Tang 0004, Xiangrong Zhang, Fang Liu 0034, Jingjing Ma 0001, Licheng Jiao
IGARSS4
2024 Stair Fusion Network With Context-Refined Attention for Remote Sensing Image Semantic Segmentation
abstract
Semantic segmentation of remote sensing images is essential in various fields, such as earth resource census, environmental pollution monitoring, and land use planning. The segmentation performance has been significantly improved recently with the development of deep learning. However, there are still some challenges in dealing with remote sensing images. One of the main issues is that features within the same category in remote sensing images could vary significantly, while features between different categories could be more similar, leading to confusion in segmentation. Moreover, the presence of large shadow areas narrows the feature differences between categories, making segmentation even more difficult. To address these challenges, one way is to leverage contextual and multi-scale information for accurate segmentation. As a consequence, in this paper, we propose a stair fusion network with context refined attention (SFCRNet). A context-based attention embedding module is proposed to enhance the representation of the processed features by utilizing the context to maximize information retention in the channel and spatial dimensions. It can retain the information on the original channel and the association between it and other channels. Furthermore, we present a stair fusion network where a stair shaped architecture and corresponding fusion module are designed to ensure that rich semantic information from high-level features is continuously transmitted to low-level layers. The experimental results on three datasets demonstrate the effectiveness of our proposed method.
Jia Liu 0020, Wenyi Hua, Fang Liu 0034, Liang Xiao 0001
IEEE Trans. Geosci. Remote. Sens.4
2024 Conjoint Cross-Attention Modeling and Joint Feature Calibrating for Remote Sensing Image Change Detection via a Triple-Double Network
abstract
Remote sensing (RS) image change detection (CD) based on deep learning (DL), has received increasing attention recently. However, the general independent learning of bi-temporal images ignores the relationship between them, falling short in learning of the change information. In this paper, a Triple-Double (TD) framework with ability of conjoint cross-attention modeling and joint feature calibrating is proposed for CD. Specifically, the TD framework composed of Triple-branch encoder and Double-branch decoder is constructed to extract diverse features and acquire changed maps with the guidance of original edge cues. To enhance the perception of the connection between the bi-temporal features, the multi-scale difference guidance (MDG) module and conjoint cross-attention (CCA) module are designed for the dual-branch encoder, wherein the CCA introduces a novel and efficient rule for modeling the affinity in spatial and channel dimension simultaneously. Furthermore, a joint feature calibration (JFC) module is introduced to enhance the expression of feature diversity in the joint features within the single-branch encoder. Experimental results on three public datasets demonstrate the superiority of the proposed method compared to the state-of-the-art (SOTA) methods.
Fang Liu 0034, Jia Liu 0020, Jingxiang Yang, Xu Tang 0004, Liang Xiao 0001
IEEE Trans. Geosci. Remote. Sens.1
2024 Candidate-Aware and Change-Guided Learning for Remote Sensing Change Detection
abstract
Change detection (CD) in remote sensing images aims at revealing earth surface changes between co-registered bitemporal images. A common way to reveal changed areas is to directly mix bitemporal features and generate CD results through supervised learning. However, a certain change usually corresponds to a real object in either of the two images, which exhibits coarse/fine shape in different scales. Therefore, a coarser-to-finer method called candidate-aware and change-guided network (CACG-Net) is proposed to effectively detect changes, where candidate objects are revealed and associated with interesting changes. Specifically, there are three key components. They are multistage change decoder (MCD), candidate-aware learning (CAL) and change guidance module (CGM). MCD reveals the most important changed objects in the coarse shape from the basic features extracted by the backbone (ResNet-18). To capture changes of interest, CAL is designed to select candidate objects in each temporal image, where a segmenter is utilized with variant change-losses. CGM intends to enrich the change details step-by-step through combining coarser change results and finer features, so that changed objects are gradually revealed in a coarser-to-finer way. Furthermore, deep supervision is employed throughout the layers of CACG-Net in the training procedure, which mitigates the learning difficulty in both deep and shallow layers. Test results on four popular datasets indicate that the proposed method outperforms several state-of-the-art CD algorithms in terms of accuracy and efficiency.
Fang Liu 0034, Yangguang Liu, Jia Liu 0020, Xu Tang 0004, Liang Xiao 0001
IEEE Trans. Geosci. Remote. Sens.1
2024 Difference Guidance Learning With Feature Alignment for Change Detection
abstract
Change detection (CD) in remote sensing aims at identifying changes of specific categories from multitemporal images acquired at different moments of a given scene. Due to seasonal alteration and light variation, there are always pseudo-changes hard to be recognized. To this end, we propose a difference guidance learning way to mitigate the effects of pseudo-change, which benefits capturing more discriminative information and identifying real changes. Specifically, it combines difference information with fused features in a guidance way and generates discriminative features in multiple scales. Besides that, feature alignment is conducted in the highest stage to learn feature correlations between bitemporal images, which benefits identifying semantic changes by information exchange. Therefore, the proposed method is named feature alignment and difference guidance network (FADG-Net). Furthermore, a set of convolutional layers with different receptive field sizes is also utilized to capture spatial information across different scales and enhance texture features accordingly. Tested on three public CD datasets, the effectiveness of the proposed FADG-Net is verified, where pseudo-change problem is mitigated and our method is superior to other comparison methods.
Yangguang Liu, Fang Liu 0034, Jia Liu 0020, Liang Xiao 0001
IEEE Trans. Geosci. Remote. Sens.2
2024 Content-Guided and Class-Oriented Learning for VHR Image Semantic Segmentation
abstract
With the flourishing of remote sensing (RS) platform techniques, very high-resolution (VHR) images have become more and more popular in recent years, which benefit the task of semantic segmentation but bring new challenges as well. Small objects, such as cars and trees, only occupy a few pixels in VHR images and are usually hard to segment. Moreover, the overlap problem about similar ground objects, such as low vegetation and trees, always results in underperformance. In this article, a content-guided and class-oriented network (CGCO-Net) for VHR image semantic segmentation is proposed to tackle this problem. Specifically, an adaptive content-guided fusion (ACGF) module with deformable convolution is introduced to capture long-distance dependencies and spatial aggregation effectively. With the guidance of the high-level features, the semantic content knowledge is gradually aggregated into low-level features and the details of the original features could be preserved. In addition, a multiscale channel alignment module is introduced into the encoder–decoder structure to further extract the long-range context information and reduce the calculation consumption. In order to improve the ability of pixel-level classification, a class-oriented representation learning (CORL) way is designed with transformer blocks by class embedding and deep supervision, which gradually enhance the discrimination and benefit the final segmentation. Furthermore, a weighted loss function and a threshold optimization strategy are employed to alleviate the sample imbalance problem. Tested on three public datasets and compared with several state-of-the-art methods, the proposed CGCO-net achieves good performance in both qualitative and quantitative analysis.
Fang Liu 0034, Keming Liu, Jia Liu 0020, Jingxiang Yang, Xu Tang 0004, Liang Xiao 0001
IEEE Trans. Geosci. Remote. Sens.1
2024 Spatial Pooling Transformer Network and Noise-Tolerant Learning for Noisy Hyperspectral Image Classification
abstract
Hyperspectral image (HSI) classification is a hot topic in remote sensing. A large number of studies have been proposed and achieved excellent performance. Most of them rely on accurate annotations. However, this requirement cannot always be met. Due to the complex contents within HSIs and the uncontrollable external interference factors, incorrect labels are inevitable. Thus, the study of noisy HSI classification is boomed. Some attempts have been made, and their central ideas are to filter the noisy samples from the training set. Although feasible, this would result in information loss, i.e., the contents covered by the removed samples are ignored. Besides, the characteristics of HSIs are not fully considered in many models. To overcome the above limitations, we develop a spatial pooling transformer network (SPTNet) and a noise-tolerant learning algorithm in this paper. SPTNet first uses a spectral feature extraction (SFE) module to capture the rich spectral information from HSI patches. Then, three spatial pooling transformers (SPTs) are constructed and stacked to explore the spatial knowledge and depress confusing clues caused by the HSI patch division. Finally, a standard transformer encoder is used to enhance the obtained spectral-spatial features for the downstream classification. To use SPTNet to handle noisy HSI classification, the noise-tolerant learning algorithm is designed. It encloses two parts, i.e., a data partition scheme and a label-independent similarity regularization. The data partition scheme divides the training data into clean and noisy sets. Then, the clean samples are used to train SPTNet with the classification loss function. At the same time, similarity regularization helps SPTNet to comprehensively understand HSIs by analyzing the resemblance between clean and noisy samples. Integrating two parts into a co-training framework, SPTNets can be trained under a noisy scenario. Four popular HSI datasets are selected to testify to our methods. The positive results demonstrate that the combination of SPTNet and the noise-tolerant learning algorithm is helpful to the noisy HSI classification. Our source codes are available at https://github.com/TangXu-Group/Hyperspectral-Images-Classification/tree/main/SPTNet-NTLA.
Jingjing Ma 0001, Yizhou Zou, Xu Tang 0004, Xiangrong Zhang, Fang Liu 0034, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.5
2024 Prior-Experience-Based Vision-Language Model for Remote Sensing Image-Text Retrieval
abstract
Remote sensing (RS) image-text retrieval (RSITR) aims to retrieve relevant texts (RS images) based on the content of a given RS image (text). Existing methods are used to employing the convolutional neural network (CNN) and recurrent neural network (RNN) as encoders to learn visual and textual features for retrieval. Although feasible, the global information hidden in different data does not receive the attention it deserves. To mitigate this problem, transformers have been introduced. Nevertheless, the complexity of RS images present challenges in directly introducing Transformer-based architectures to multimodal learning in RS scenes, particularly in visual feature extraction and cross-modal interaction. In addition, the textual captions are always simpler than the complex RS images, leading to a semantic description appearing in different images. This typical false-negative (FN) sample problem increases the difficulty of RSITR tasks. To address the above limitations, we propose a new RSITR model named prior-experience-based RS vision-language (PERSVL). First, the specific visual and text encoders are used to extract features from RS images and texts. Also, a high-level feature complement (HFC) module is developed based on the self-attention mechanism (SAM) for the visual encoder to explore the complex contents from RS images fully. Second, a dual-branch multimodal fusion encoder (DBMFE) is designed to complete the cross-modal learning. It comprises a dual-branch multimodal interaction (DBMI) module and a branch fusion module. DBMI is designed to fully explore the relationships between different modalities, enriching visual and textual features. The branch fusion module integrates the cross-modal features and utilizes a classification head to generate matching scores for retrieval. Finally, a learning from prior experiences (LPEs) module is designed to reduce the influence of FN samples by analyzing the historical data produced in the model training process. Experiments are conducted on three popular datasets, and the positive results show that our PERSVL model achieves superior performance compared with previous methods. By integrating the advantages of natural language and RS images, our PERSVL can be applied in various applications, such as environmental monitoring, disaster evaluation, and urban planning. Our source codes are available at:https://github.com/TangXu-Group/Cross-modal-remote-sensing-image-and-text-retrieval-models/tree/main/PERSVL.
Xu Tang 0004, Dabiao Huang, Jingjing Ma 0001, Xiangrong Zhang, Fang Liu 0034, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.5
2024 Multiple Information Collaborative Fusion Network for Joint Classification of Hyperspectral and LiDAR Data
abstract
Joint classification of hyperspectral image (HSI) and light detection and ranging (LiDAR) can simultaneously utilize rich spectral information and elevation information and has become a hot research topic in remote sensing (RS). Although many works have been proposed for this task, their performance cannot reach what we expected due to inadequate cross-modal feature learning and simple feature fusion. This article proposes a multiple information collaborative fusion network (MICF-Net) to overcome those limitations, which aims to leverage the essentially consistent spatial relationships and high-level semantic information in multimodal data to guide the extraction of multimodal fusion features. Specifically, MICF-Net first uses a simple two-branch convolutional neural network (CNN) for preliminary feature extraction. Then, a dual-branch cross-modal attention fusion transformer (CMAFT) is developed to mine global contextual content. By fusing the attention maps of two modalities and limiting their similarity, CMAFT can retain modality-specific information while achieving information interaction based on spatial relationships. Next, an adaptive mask modulation (AMM) module is designed to dynamically balance the learning rate of each modality to ensure the effectiveness of the features of all modalities. Finally, to mine the complementary information of HSI and LiDAR data, a semantic-guided feature fusion (SGFF) module is introduced. It achieves mutual guided learning by exchanging semantic information between two modalities. Positive experimental results counted on three popular HSI and LiDAR datasets demonstrate the effectiveness of the proposed MICF-Net. Our source codes are available athttps://github.com/TangXu-Group/Hyperspectral-Images-Classification/tree/main/MICF-Net.
Xu Tang 0004, Yizhou Zou, Jingjing Ma 0001, Xiangrong Zhang, Fang Liu 0034, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.5
2024 I2MEP-Net: Inter- and Intra-Modality Enhancing Prototypical Network for Few-Shot Hyperspectral Image Classification
abstract
Prototypical network, celebrated for its flexible network structure and metric computation capability, has become a prevailing strategy in addressing challenges associated with few-shot hyperspectral image classification. However, factors such as insufficient training samples, spectral mixing, and noise interference in complex scenarios severely impact the stability of its prototypes, ultimately leading to a degradation in classification performance. Therefore, this paper proposes a novel method called I2MEP-Net, which incorporates two auxiliary modalities to facilitate both inter- and intra-modality enhancing prototype learning with the base modality. Specifically, I2MEP-Net employs the auxiliary LiDAR modality with base HS modality from the same scene for inter-modality enhancing prototype learning, which offsets the sparsity of few-shot features through a cross-modal approach. In addition, it utilizes the target few-shot labeled data as the auxiliary HS modality for intra-modality enhancing prototype learning on the enhanced prototypes, in a way that adaptively generates diversity features, thereby further enriching the prototype embedding space and achieving more fine-grained and stable prototypes. Comprehensive experiments are conducted on the publicly available hyperspectral image datasets. These experiments indicate that the proposed I2MEP-Net outshines the existing state-of-the-art deep learning techniques and few-shot classification methodologies.
Wenzhen Wang, Fang Liu 0034, Liang Xiao 0001
IEEE Trans. Geosci. Remote. Sens.2
2024 Hyperspectral Reconstruction From RGB Images via Physically Guided Graph Deep Prior Learning
abstract
Recovering the latent hyperspectral image (HSI) from RGB or multispectral image (MSI), which is dubbed spectral super-resolution (SSR), has demonstrated outstanding performance owing to the advancements in convolutional neural networks (CNNs). However, most of the current algorithms concentrate on the pursuit of networks with more expensive or complex structures, while ignoring the significant role of physical degradation models in SSR. In addition, the inherent defects of CNN make these networks focus more on the local correlation, while their ability to model the long-range correlations in the spectral and spatial domains still has room to improve. To overcome this shortcoming, we propose a physical degradation-guided deep prior learning network (PGDL-Net) for SSR via unfolding the optimization process of the blind SSR model, in which the priors of unknown spectral response function (SRF) and latent HSI are learned explicitly and represented by proximal operators. To jointly extract the local and non-local information, we design a hybrid graph Transformer as the proximal operator to solve the latent HSI. Furthermore, to ensure efficient learning of SRF and HSI, we also propose a novel loss function constraining the reconstruction error, degradation consistency, and observation fidelity for the learned SRF and HSI. Experimental results on multiple datasets illustrate the improved performance and stability of our method in SSR.
Jingxiang Yang, Tian Lin 0001, Jia Liu 0020, Fang Liu 0034, Liang Xiao 0001
IEEE Trans. Geosci. Remote. Sens.5
2024 Semi-Supervised Multiscale Dynamic Graph Convolution Network for Hyperspectral Image Classification
abstract
In recent years, convolutional neural networks (CNNs)-based methods achieve cracking performance on hyperspectral image (HSI) classification tasks, due to its hierarchical structure and strong nonlinear fitting capacity. Most of them, however, are supervised approaches that need a large number of labeled data to train them. Conventional convolution kernels are fixed shape of rectangular with fixed sizes, which are good at capturing short-range relations between pixels within HSIs but ignore the long-range context within HSIs, limiting their performance. To overcome the limitations mentioned above, we present a dynamic multiscale graph convolutional network (GCN) classifier (DMSGer). DMSGer first constructs a relatively small graph at region-level based on a superpixel segmentation algorithm and metric-learning. A dynamic pixel-level feature update strategy is then applied to the region-level adjacency matrix, which can help DMSGer learn the pixel representation dynamically. Finally, to deeply understand the complex contents within HSIs, our model is expanded into a multiscale version. On the one hand, by introducing graph learning theory, DMSGer accomplishes HSI classification tasks in a semi-supervised manner, relieving the pressure of collecting abundant labeled samples. Superpixels are generally in irregular shapes and sizes which can group only similar pixels in a neighborhood. On the other hand, based on the proposed dynamic-GCN, the pixel-level and region-level information can be captured simultaneously in one graph convolution layer such that the classification results can be improved. Also, due to the proper multiscale expansion, more helpful information can be captured from HSIs. Extensive experiments were conducted on four public HSIs, and the promising results illustrate that our DMSGer is robust in classifying HSIs. Our source codes are available at https://github.com/TangXu-Group/DMSGer.
Yuqun Yang, Xu Tang 0004, Xiangrong Zhang, Jingjing Ma 0001, Fang Liu 0034, Xiuping Jia, Licheng Jiao
IEEE Trans. Neural Networks Learn. Syst.5
2023 Stair Fusion Network for Remote Sensing Image Semantic Segmentation
abstract
Semantic segmentation of very high spatial resolution remote sensing images plays a vital role in many fields, such as land resource management, urban planning, and biosphere monitoring. Due to the scare variance between different types of regions, it is important to fully utilize multi-scale features. Moreover, with the complexity of some ground objects, global semantic information should be specially considered. As a consequence, in this paper, we propose a stair fusion network to further refine and fuse low-level and high-level features. In addition, we propose a global information enhancement module (GIEM) to extract global semantic information from the high-level features and reduce the length of delivery chain from them to the final results via a skip connection. Experimental results demonstrate the effectiveness of our model.
Wenyi Hua, Jia Liu 0020, Fang Liu 0034
IGARSS3
2023 Spatial-Spectral Adaptive Learning With Pixelwise Filtering for Hyperspectral Image Classification
abstract
Hyperspectral image (HSI) classification is significant in remote sensing applications. However, most methods focus on the spectral and spatial correlation information in the neighborhood while ignoring the feature difference among global different pixels. In this article, we propose a spatial–spectral adaptive learning with pixelwise filtering (SSALPF) method to fully consider the discriminative information of pixels in different spatial locations, which mainly consists of a parallel spatial–spectral adaptive learning (SSAL) module and a pixelwise filtering (PF) module. Specifically, the former aims to obtain joint spatial–spectral discriminative features of each pixel point in a parallel manner and is used as a guide for adaptive selection of filter kernel. The latter uses the adaptive filter kernel to implement pixel-level filtering on HSI, in order to learn the discriminative features contained in different pixel points for classification. The adaptive filter kernel is generated by a linear combination of a predefined dictionary containing multiple filter bases. Experiments demonstrate that the proposed method is superior to other methods on popular hyperspectral datasets.
Wenfei Gao, Fang Liu 0034, Jia Liu 0020, Liang Xiao 0001, Xu Tang 0004
IEEE Trans. Geosci. Remote. Sens.2
2023 MANet: An Efficient Multidimensional Attention-Aggregated Network for Remote Sensing Image Change Detection
abstract
Deep learning has significantly advanced the change detection in remote sensing image with its excellent performance. For change detection tasks, there are two critical issues. First, with scale variance of different objects in remote sensing images, effectively aggregating multi-scale features helps to generate fine-grained change objects. Second, it is critical but challenging to fully exploit the variance information between bi-temporal images to avoid pseudo-variation and region blurring. To alleviate the above issues, this paper proposes an efficient multi-dimensional attention-aggregation network (MANet), which keeps better feature aggregation while maintaining excellent differential attention ability. This paper carries three main contributions. First, we propose a multiscale asymmetric convolutional attention (MACA) module. Due to the asymmetric convolution’s ability to focus on feature contours effectively, the MACA can not only aggregate multi-scale features effectively, but also refine the edge information of features. Second, we propose a dual-dimensional attention (DDA) module for adaptively fusing shallow and deep features, which is used to generate rich feature representations. Third, the difference guidance (DG) module is exploited for enhancing the attention of changed regions to mitigate the influence of uncorrelated changes on the change detection result. Experiments on four popular change detection datasets show that our network can accomplish higher detection accuracy than the state-of-the-art networks.
Kaixuan Jiang, Jia Liu 0020, Fang Liu 0034, Liang Xiao 0001
IEEE Trans. Geosci. Remote. Sens.4
2023 Adversarial Domain Alignment With Contrastive Learning for Hyperspectral Image Classification
abstract
Recently, deep learning-based hyperspectral image (HSI) classification techniques are flourishing and exhibit good performance, where cross domain information is usually utilized to reduce the dependency on large labeled samples. However, the gap between source domain and target domain makes it difficult to carry out knowledge transfer directly. In this paper, an adversarial domain alignment with contrastive learning method is designed for the HSI classification task to achieve feature consistency that benefits transferring knowledge. In details, spectral alignment and semantic alignment are conducted in local and global levels respectively in an adversarial learning way, and the adversarial loss acts on both source and target domains. In order to learn specific features for objects with different spatial scales, a multi-scale selection module is constructed in semantic alignment to select channel features adaptively. Moreover, contrastive learning is employed to increase both robustness and sensitiveness, where augmented data from the same/different samples are forced to be similar/dissimilar with each other. The training process is conducted in a few-shot learning way then the few-shot classification loss, the adversarial loss and the contrastive loss is optimized together. Tested on one source dataset and four target datasets, the experimental results show that the proposed method outperforms the other comparisons.
Fang Liu 0034, Wenfei Gao, Jia Liu 0020, Xu Tang 0004, Liang Xiao 0001
IEEE Trans. Geosci. Remote. Sens.1
2023 Interacting-Enhancing Feature Transformer for Cross-Modal Remote-Sensing Image and Text Retrieval
abstract
Cross-modal remote sensing image-text retrieval (CMRSITR) is a challenging topic in the remote sensing (RS) community. It has gained growing attention because it can be flexibly used in many practical applications. In the current deep era, with the help of deep convolutional neural networks (DCNNs), many successful CMRSITR methods have been proposed. Most of them first learn valuable features from RS images and texts respectively. Then, the obtained visual and textual features are mapped into a common space for the final retrieval. The above operations are feasible, however, two difficulties are still to be solved. One is that the semantics within the visual and textual features are misaligned due to the independent learning manner. The other one is that the deep links between RS images and texts cannot be fully explored by simple common space mapping. To overcome the above challenges, we propose a new model named interacting-enhancing feature transformer (IEFT) for CMRSITR, which regards the RS images and texts as a whole. First, a simple feature embedding module (FEM) is developed to map images and texts into the visual and textual feature spaces. Second, an information interacting-enhancing module (IIEM) is designed to simultaneously model the inner relationships between RS images and texts and enhance the visual features. IIEM consists of three feature interacting-enhancing (FIE) blocks, each of which contains an inter-modality relationship interacting (IMRI) sub-block and a visual feature enhancing (VFE) sub-block. The duty of IMRI is to exploit the hidden relations between cross-modal data, while the responsibility of VFE is to improve the visual features. By combining them, semantic bias can be mitigated, and the complex contents of RS images can be studied. Finally, the retrieval module (RM) is constructed to generate the matching scores for deciding the search results. Extensive experiments are conducted on four public RS data sets. The positive results demonstrate that our IEFT can achieve superior retrieval performance compared with many existing methods. Our source codes are available at https://github.com/TangXu-Group/Cross-modal-remote-sensing-image-and-text-retrieval-models/tree/main/IEFT.
Xu Tang 0004, Yijing Wang 0004, Jingjing Ma 0001, Xiangrong Zhang, Fang Liu 0034, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.5
2023 Cross-Modal Graph Knowledge Representation and Distillation Learning for Land Cover Classification
abstract
Complementary multimodal remote sensing (RS) data often leads to more robust and accurate classification performance. However, not all modal data can be available at the time of inference due to imaging conditions. To mitigate this issue, cross-modal knowledge distillation becomes an effective method, as it can leverage the complementary characteristics of multimodal data to guide cross-modal classification in cases with missing data. Therefore, this paper examines the shortcomings of traditional CNN cross-modal distillation methods in land cover classification: 1) insufficient knowledge representation; and 2) unstable knowledge transfer. Moreover, a novel cross-modal graph knowledge representation and distillation learning (CGKR-DL) framework is proposed to enhance land cover classification performance. The proposed CGKR-DL designs a single-stream joint feature learning network with convolutional neural network and graph convolutional network (CNN-GCN) to effectively construct the remote topology of data based on the strong correlation between land objects, thus enhancing the knowledge representation ability of the network. In addition, a multi-granularity graph distillation method is proposed to compensate for the inability of traditional CNN distillation in handling graph-structured information, where a feature distillation module based on graph discrimination (FD-GDM) is designed for stable graph feature distillation. We evaluate CGKR-DL on three publicly available multimodal RS datasets (HS-LiDAR, HS-SAR and HS-SAR-DSM) and achieve a significant improvement in comparison with several state-of-the-art methods.
Wenzhen Wang, Fang Liu 0034, Wenzi Liao, Liang Xiao 0001
IEEE Trans. Geosci. Remote. Sens.2
2023 Learning Degradation-Aware Deep Prior for Hyperspectral Image Reconstruction
abstract
Reconstructing the 3D hyperspectral image (HSI) from 2D snapshot measurements is a key task in spectral snapshot compressive imaging (SCI). Traditional model-based HSI reconstruction methods rely on hand-crafted priors. Recently, deep unfolding networks (DUNs) learn the priors using convolutional neural networks (CNNs) and have achieved satisfactory results. Most of DUNs assume the degradations of SCI are known. However, due to the phase aberration and distortion problems in real imaging process, there is a certain gap between the ideal and real degradation patterns, which may hinder the accurate HSI reconstruction. In this study, we propose a degradation-aware deep prior learning network (D2PL-Net), which tries to adaptively learn the practical degradation matrix during HSI reconstruction, thus bridges the gap between the ideal and real degradations. Specifically, we first propose a joint variational compressive reconstruction model, both of the latent HSI and unknown degradation can be explicitly solved. By unfolding the solutions into a deep network, D2PL-Net is built, which mainly consists of two parts, Degradation Matrix Learning (DML) mechanism and Degradation-guided Spectral-Spatial Transformer (DSST) in each stage. The former learns the degradation that approximates the real one; the latter represents the deep prior of latent HSI, it could exploit the spectral-wise and spatial-wise long-range dependencies of HSI under the guidance of learned degradation, and then reconstructs the HSI. To ensure an effective training of D2PL-Net, we propose a joint loss function constraining the HSI reconstruction errors, degradation-fidelity and degradation-consistency. Experiments on simulated and real-life datasets show that the proposed method is competitive with the state-of-the-art methods.
Jingxiang Yang, Tian Lin 0001, Fang Liu 0034, Liang Xiao 0001
IEEE Trans. Geosci. Remote. Sens.3
2022 Dual Unet: A Novel Siamese Network for Change Detection with Cascade Differential Fusion
abstract
Change detection (CD) of remote sensing images is to detect the change region by analyzing the difference between two bitemporal images. It is extensively used in land resource planning, natural hazards monitoring and other fields. In our study, we propose a novel Siamese neural network for change detection task, namely Dual-UNet. In contrast to previous individually encoded the bitemporal images, we design an encoder differential-attention module to focus on the spatial difference relationships of pixels. In order to improve the generalization of networks, it computes the attention weights between any pixels between bitemporal images and uses them to engender more discriminating features. In order to improve the feature fusion and avoid gradient vanishing, multi-scale weighted variance map fusion strategy is proposed in the decoding stage. Experiments demonstrate that the proposed approach consistently outperforms the most advanced methods on popular seasonal change detection datasets.
Kaixuan Jiang, Jia Liu 0020, Fang Liu 0034, Yangguang Liu, Jiao Shi
IGARSS3
2022 Spatial-Adaptive and Feature-Enhanced Siamese Network for Change Detection
abstract
Change detection (CD) plays an increasingly important role in earth observation and reveals surface changes according to multi-temporal images. Although deep learning-based CD methods work well for their excellent modeling ability, objects in different size and shape are generally processed by the same filter kernels in feature extraction, which leads to spatial blurring and degrades the CD performance. In this paper, a spatial adaptive and feature enhanced (SAFE) siamese network is proposed to tackle this problem, where the SAFE consists of a spatial-adaptive (SA) part and a feature-enhanced (FE) part. Specifically, pixel belonging to different objects possesses its own spatial knowledge, which is captured by a soft fusion of multi-scale difference images (DIs) called SA part. Changed and unchanged areas are strengthened or weakened by the FE, which combines object features with each DI accordingly. Moreover, since there are more unchanged pixels than changed pixels, a weight-pair is introduced to balance changed and unchanged objects in the training process. The experimental results verify that compared with four representative CD algorithms, our proposed method performs best on the Change Detection Dataset (CDD).
Yangguang Liu, Fang Liu 0034, Jia Liu 0020, Xu Tang 0004, Kaixuan Jiang, Liang Xiao 0001
IGARSS2
2022 Domain-Adaptive Few-Shot Learning for Hyperspectral Image Classification
abstract
Recently, hyperspectral image (HSI) classification by deep learning is flourishing. However, only a few labeled samples are available in practice since it is time-and-labor-consuming to label pixels in HSI (called target domain). This paper proposes a domain-adaptive few-shot learning (DAFSL) method to tackle this problem. Specifically, some other HSIs (called source domain) with large labeled samples are fully used as complementary information and a generative architecture is employed to adapt embedded features in source domain to that of target domain. We first perform domain adaptation with unsupervised learning. In details, the embedded features are generated by the encoder of an autoencoder, where both source and target samples could be well recovered and the reconstruction loss is used to measure the gap between source domain and target domain. At the same time, the embedded features are put into a metric space for classification in source domain and the encoder parameter is fine-tuned together with the classifier in target domain with few labels, so that both general and discriminative features are well captured. The experiment results show that DAFSL outperforms the other mainstream methods with limited labeled samples.
Andi Zhang 0003, Fang Liu 0034, Jia Liu 0020, Xu Tang 0004, Wenfei Gao, Liang Xiao 0001
IEEE Geosci. Remote. Sens. Lett.2
2022 Evolving Connections in Group of Neurons for Robust Learning
abstract
Artificial neural networks inspired from the learning mechanism of the brain have achieved great successes in machine learning, especially those with deep layers. The commonly used neural networks follow the hierarchical multilayer architecture with no connections between nodes in the same layer. In this article, we propose a new group architectures for neural-network learning. In the new architecture, the neurons are assigned irregularly in a group and a neuron may connect to any neurons in the group. The connections are assigned automatically by optimizing a novel connecting structure learning probabilistic model which is established based on the principle that more relevant input and output nodes deserve a denser connection between them. In order to efficiently evolve the connections, we propose to directly model the architecture without involving weights and biases which significantly reduce the computational complexity of the objective function. The model is optimized via an improved particle swarm optimization algorithm. After the architecture is optimized, the connecting weights and biases are then determined and we find the architecture is robust to corruptions. From experiments, the proposed architecture significantly outperforms existing popular architectures on noise-corrupted images when trained only by pure images.
Jia Liu 0020, Maoguo Gong, Liang Xiao 0001, Fang Liu 0034
IEEE Trans. Cybern.5
2022 Joint Variation Learning of Fusion and Difference Features for Change Detection in Remote Sensing Images
abstract
Remote sensing (RS) image change detection (CD) is an earth observation technique for detecting surface changes in the same area during a period. With the rapid development of deep learning, various deep neural networks especially Siamese ones have been widely used in the field of CD. However, they have the deficiency of insufficient contextual information aggregation, resulting in false and missed detections, and it is difficult to refine the detection of change edges. To alleviate these problems and obtain more accurate results, we propose an efficient self-weighted spatial-temporal attention network (SSANet). In contrast to the Siamese structure, our network is a novel joint learning framework composed of fusion sub-network, difference sub-network, and decoder. Fusion sub-network is used to extract multiscale object features where we propose a multi-core channel-aligning attention (MCA) module to capture the long-range semantic information for multi-scale context aggregation. Difference sub-network is used to extract the difference variation features, where we propose a feature differential reconfiguration (FDR) module to learn the temporal change information. FDR can effectively filter change information and reconstruct features to improve the perception of changed regions. To better balance the MCA and FDR modules, an asymmetric weighting (AW) module is proposed in the decoder to self-weight the multi-scale features and generate the change map. Experiments demonstrate the efficiency of proposed sub-networks and modules, and the state-of-the-art performance of SSANet.
Kaixuan Jiang, Jia Liu 0020, Fang Liu 0034, Liang Xiao 0001
IEEE Trans. Geosci. Remote. Sens.4
2022 Adaptive Graph Convolutional Network for PolSAR Image Classification
abstract
Polarimetric synthetic aperture radar (PolSAR) image classification is one of the hottest issues in remote sensing, where studies on pixel-level information and relationship are of great significance. In this article, graph convolutional network (GCN) is employed to accomplish this pixel-level task benefiting from its excellent capability in structure exploration and information propagation between different pixels. To reduce the communication burden between various PolSAR pixels and high computational cost for the whole PolSAR image, an adaptive GCN (AdapGCN) consisting of pixel-centered subgraphs is proposed in this article. In the AdapGCN, a data-adaptive kernel and a spatial-adaptive kernel are introduced to, respectively, model data structure and spatial structure for PolSAR image. Moreover, a multiscale learning structure is integrated to further explore complicated relations between pixels. Extensive comparative evaluations validate the superiority of our new AdapGCN model for PolSAR image classification over a wide range of state-of-the-art methods on three challenging benchmarks.
Fang Liu 0034, Jingya Wang 0001, Xu Tang 0004, Jia Liu 0020, Xiangrong Zhang, Liang Xiao 0001
IEEE Trans. Geosci. Remote. Sens.1
2022 A Probabilistic Model Based on Bipartite Convolutional Neural Network for Unsupervised Change Detection
abstract
This article presents a probabilistic model based on a bipartite convolutional architecture for unsupervised change detection. We aim to develop a robust change detection method that can adapt to different types of data and scenarios for multitemporal coregistered remote sensing images of the same spatial resolution. On the premise of coregistration, unsupervised change detection usually suffers from the distinct appearances (different intensities or data structures) of the same object in multitemporal images, such as images obtained in different climatic conditions (season, illumination, and so on), and by different and even heterogeneous sensors. Since change detection in heterogeneous images can also adapt to other scenarios, many methods have been proposed recently focusing on such data, but most of them are limited by the need for labeled data or by specific assumptions. With the excellent and flexible feature learning capability of neural networks, we model the change detection into a Gibbs probabilistic model based on a bipartite neural network. The model is driven by an energy function defined as the squared feature distance, which is the core of change detection. Via optimizing the model, the difference degree of each pixel is automatically obtained for further identification. The probabilistic model learns to capture the distribution in an unsupervised way. Therefore, the proposed method can adapt to various scenarios without being trained by labeled data. Experiments on different types of data and scenarios demonstrate the superiority of the proposed method.
Jia Liu 0020, Fang Liu 0034, Liang Xiao 0001
IEEE Trans. Geosci. Remote. Sens.3
2022 Class-Level Prototype Guided Multiscale Feature Learning for Remote Sensing Scene Classification With Limited Labels
abstract
Remote sensing scene classification (RSSC) is an open and challenging research topic in the remote sensing (RS) community. It aims to define semantic labels for RS scenes according to their contents. Recently, with the development of deep convolutional neural networks (DCNNs), the results of RSSC have been enhanced to a large extent. However, the cracking performance of these DCNN-based models depends on a large number of labeled data. Once the volume of the labeled data is decreased, their behavior would be weakened dramatically. In this article, we propose a new training algorithm that can work smoothly with a few labeled samples to address this limitation. Along with the introduced DCNN, the presented methods perform satisfactorily. In particular, we first construct a dual-branch network (DBNet) to mine the multiscale and multiangle information from RS scenes. Thus, the abundant land covers with diverse sizes, directions, and shapes can be captured simultaneously. Then, to train DBNet using scarce semantic labels, a class-level prototype guided learning (CPGL) algorithm is developed based on the meta-learning paradigm. Besides the usual episode training manner, a prototype refinement module (PRM) and a prototype discrimination module (PDM) are designed with the help of metric learning theory to ensure the effectiveness of our CPGL. The comprehensive experiments are conducted on four public RS scene datasets, and the encouraging results imply that our DBNet and CPGL can copy with RSSC tasks with small labeled data. Our source codes are available athttps://github.com/TangXu-Group/Remote-Sensing-Images-Classification/tree/main/CPGL.
Xu Tang 0004, Weiquan Lin, Jingjing Ma 0001, Xiangrong Zhang, Fang Liu 0034, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.5
2022 Meta-Hashing for Remote Sensing Image Retrieval
abstract
With the explosive growth of the volume and resolution of high-resolution remote-sensing (HRRS) images, the management of them becomes a challenging task. The traditional content-based remote-sensing image retrieval (CBRSIR) technologies cannot meet what we expect due to the large volume of image archives and complex contents within HRRS images. As a successful approximate nearest neighborhood (ANN) search technique, Hash learning has received wide attention, especially when deep convolutional neural networks (DCNNs) appear. Due to DCNNs’ strong capacity of feature learning, many DCNN-based hashing methods have been proposed and achieved good performance for large-scale CBRSIR tasks. Nevertheless, their limitation is that a large of labeled training samples should be collected for training the deep models. To overcome this limitation, this article, therefore, develops a new supervised hash learning method for the large-scale HRRS CBRSIR task based on meta-learning, which could achieve well-retrieval performance with a few labeled training samples. First, taking the characteristics of HRRS into account, we develop a self-adaptive convolution (SAP-Conv) block and design a hashing net based on the block. SAP-Conv can learn robust features from HRRS images by exploring their multiscale information. Second, to enhance the generalization of the hashing net under a few labeled training samples, the hash learning is formulated in a meta-way, and we name it meta-hashing. Meta-hashing can effectively preserve the similarities between support and query set, and the similarities between samples within support set by the developed loss function. To further improve the performance of meta-hashing, we expand it to a dynamic version named dynamic-meta-hashing, in which the numbers of support and query are changeable in the training phase. Experimental results counted on the three widely used HRRS datasets demonstrate our dynamic-meta-hashing and meta-hashing can achieve promising performance in large-scale HRRS CBRSIR tasks based on a few training samples. Our source codes are available athttps://github.com/TangXu-Group/Meta-hashing.
Xu Tang 0004, Yuqun Yang, Jingjing Ma 0001, Yiu-Ming Cheung, Chao Liu 0042, Fang Liu 0034, Xiangrong Zhang, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.6
2022 An Unsupervised Remote Sensing Change Detection Method Based on Multiscale Graph Convolutional Network and Metric Learning
abstract
As a fundamental application, change detection (CD) is widespread in the remote sensing (RS) community. With the increase in the spatial resolution of RS images, high-resolution remote sensing (HRRS) image CD tasks receive growing attention. The change information hidden in multitemporal HRRS images could help discover our planet comprehensively. In the current deep learning era, convolutional neural networks (CNNs) have become one of the most powerful tools for a wide range of RS tasks including HRRS image CD, due to their superb feature learning capacity. However, most of them need a large amount of labeled data to accomplish the CD process, which is challenging or even impractical in many RS applications. Also, given the limited valid receptive field, CNNs can only capture short-range context within HRRS images, which is probably not enough to fully explore change information from the images. To overcome these limitations, in this article, we propose an unsupervised CD method, termed GMCD, based on graph convolutional network (GCN) and metric learning. GMCD consists of a Siamese fully convolution network (FCN), a multiscale dynamic GCN (Mlt-GCN), and a pseudolabel generation mechanism based on metric learning. The Siamese FCN contains a Siamese encoder and a pyramid-shaped decoder, aiming to extract multiscale features and integrate them to generate reliable difference images (DIs). Mlt-GCN focuses on capturing the short- and long-range contextual patterns at feature map level to extract changed and unchanged areas completely. The pseudolabel generation mechanism aims to produce reliable pseudolabels (changed, unchanged, and uncertain) to help accomplish the model training in an unsupervised way. Experiments on four HRRS image CD datasets demonstrate that GMCD outperforms the existing state-of-the-art methods.
Xu Tang 0004, Lichao Mou, Fang Liu 0034, Xiangrong Zhang, Xiao Xiang Zhu 0001, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.4
2022 AR2Det: An Accurate and Real-Time Rotational One-Stage Ship Detector in Remote Sensing Images
abstract
Ship detection plays a significant role in the high-resolution remote sensing (HRRS) community, but it is a challenging task due to the complex contents within HRRS images and the diverse orientation of ships. Recently, with the development of deep learning, the performance of the HRRS ship detection model has been improved greatly. Most of them employ deep networks and complicate anchor mechanism to get well ship detection results. Nevertheless, this kind of combination limits the detection efficiency. To address this problem, a new approach named accurate and real-time rotational ship detector (AR2Det) is proposed in this article to detect ships without the anchor mechanism. Based on the extracted features by the feature extraction module (FEM) and the central information of ships, AR2Det adopts two simple modules, ship detector (SDet) and center detector (CDet), to generate and improve the detection results, respectively. AR2Det is efficient due to the simple postprocessing and the lightweight network. Also, AR2Det performs satisfactorily due to the effective generation and enhancement strategy of bounding boxes. The extensive experiments are conducted on a public HRRS image ship detection dataset HRSC2016. The promising results show that our method outperforms the state-of-the-art approaches in terms of both accuracy and speed.
Yuqun Yang, Xu Tang 0004, Yiu-Ming Cheung, Xiangrong Zhang, Fang Liu 0034, Jingjing Ma 0001, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.5
2021 Deep Hash Learning for Remote Sensing Image Retrieval
abstract
The content-based remote sensing image retrieval (CBRSIR) has attracted increasing attention with the number of remote sensing (RS) images growing explosively. Benefiting from the strong capacity of the deep convolutional neural network (DCNN), the performance of CBRSIR has been improved in recent years. Although great successes have been obtained, learning the RS images' representative features and enhancing the retrieval efficiency for the large-scale CBRSIR tasks are still two challenging problems. In this article, we propose a new CBRSIR method named feature and hash (FAH) learning, which consists of a deep feature learning model (DFLM) and an adversarial hash learning model (AHLM). The DFLM aims at learning the RS images' dense features to guarantee the retrieval precision. In the DFLM, the DCNN and the proposed feature aggregation are integrated to capture the multiscale features. Then, the discrimination of the obtained features can be highlighted by the attention map in the developed attention branch. The AHLM maps the dense features onto the compact hash codes so that the retrieval efficiency can be improved. The AHLM contains a hash learning submodel and an adversarial regularization submodel. In particular, the hash learning submodel learns the real-valued hash codes that are similarity preserved by semantic supervisions. The adversarial regularization submodel regularizes the real-valued hash codes to learn the discrete uniform distribution with possible values 0 and 1. In this way, the hash codes are coding-balanced and the quantization errors are reduced. Encouraging experimental results counted on three public benchmark data sets demonstrate that our FAH can achieve competitive performance in the CBRSIR task compared with many existing hash learning methods.
Chao Liu 0042, Jingjing Ma 0001, Xu Tang 0004, Fang Liu 0034, Xiangrong Zhang, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.4
2021 Large-Scope PolSAR Image Change Detection Based on Looking-Around-and-Into Mode
abstract
A new method based on the Looking-Around-and-Into (LAaI) mode is proposed for the task of change detection in large-scope Polarimetric Synthetic Aperture Radar (PolSAR) image. Specifically, the LAaI mode consists of two processes named Look-Around and Look-Into, which are accomplished by attention proposal network (APN) and recurrent convolutional neural network (CNN) (Recurrent CNN), respectively. The former provides certain subregions efficiently, and the latter detects changes in subregions accurately. In Look-Around, difference image (DI) of whole PolSAR images is calculated first to get global information; then, APN is established to locate the position of interested subregions intentionally by paying special attention to; next interested subregions that contain changed area in high probability are picked out as candidate-regions. Moreover, candidate-regions are sorted in importance descending order so that highly interested regions have priority to be detected. In Look-Into, candidate-regions of different scales are selected at first; then, Recurrent CNN is constructed and employed to deal with multiscale PolSAR subimages so that clearer and finer change detection results are generated. The process is repeated until all candidate-regions are detected. As a whole, the proposed algorithm based on the LAaI mode looks around whole images first to find out the possible position of changes (candidate-regions generation in Look-Around) and then reveal the exact shape of changes in different scales (multiscale change detection in Look-Into). The effect of APN and Recurrent CNN is verified in experiments, and it shows that the proposed method performs well in the task of change detection in the large-scope PolSAR image.
Fang Liu 0034, Xu Tang 0004, Xiangrong Zhang, Licheng Jiao, Jia Liu 0020
IEEE Trans. Geosci. Remote. Sens.1
2021 Hyperspectral Image Classification Based on 3-D Octave Convolution With Spatial-Spectral Attention Network
abstract
In recent years, with the development of deep learning (DL), the hyperspectral image (HSI) classification methods based on DL have shown superior performance. Although these DL-based methods have great successes, there is still room to improve their ability to explore spatial-spectral information. In this article, we propose a 3-D octave convolution with the spatial-spectral attention network (3DOC-SSAN) to capture discriminative spatial-spectral features for the classification of HSIs. Especially, we first extend the octave convolution model using 3-D convolution, namely, a 3-D octave convolution model (3D-OCM), in which four 3-D octave convolution blocks are combined to capture spatial-spectral features from HSIs. Not only the spatial information can be mined deeply from the high- and low-frequency aspects but also the spectral information can be taken into account by our 3D-OCM. Second, we introduce two attention models from spatial and spectral dimensions to highlight the important spatial areas and specific spectral bands that consist of significant information for the classification tasks. Finally, in order to integrate spatial and spectral information, we design an information complement model to transmit important information between spatial and spectral attention features. Through the information complement model, the beneficial parts of spatial and spectral attention features for the classification tasks can be fully utilized. Comparing with several existing popular classifiers, our proposed method can achieve competitive performance on four benchmark data sets.
Xu Tang 0004, Xiangrong Zhang, Yiu-Ming Cheung, Jingjing Ma 0001, Fang Liu 0034, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.6
2020 Hyperspectral Image Classification Based on Multiscale Spatial and Spectral Feature Network
abstract
With the development of deep learning, hyperspectral image (HSI) classification tasks have developed rapidly, the classification performance is improved in a big degree. Despite the great success of the existing methods, there is still room for improvement to extract features from spatial and spectral dimensions. In this paper, we propose a multiscale spatial and spectral feature network (MSSFN) to capture discriminative features for the classification of HSIs. Specifically, we first use three convolution layers to extract the features of original HSI data. Second, combining the spatial masks model and spectral attention model to build multiscale spatial and spectral model (MSSM). Through the MSSM model, the spatial information of different scales can be obtained and the useful spectral bands can be emphasized. Finally, in order to reduce the computation complexity and simplify network, the other three convolution layers with a small number of convolution kernels are adopted in our method. The experimental results demonstrate that our method is superior to most existing methods on two public HSI datasets.
Xu Tang 0004, Jingjing Ma 0001, Xiangrong Zhang, Fang Liu 0034, Qunnie Peng, Licheng Jiao
IGARSS5
2019 Task-Oriented GAN for PolSAR Image Classification and Clustering
abstract
Based on a generative adversarial network (GAN), a novel version named Task-Oriented GAN is proposed to tackle difficulties in PolSAR image interpretation, including PolSAR data analysis and small sample problem. Besides two typical parts in GAN, i.e., generator (G-Net) and discriminator (D-Net), there is a third part named TaskNet (T-Net) in the Task-Oriented GAN, where T-Net is employed to accomplish a certain task. Two tasks, PolSAR image classification and clustering, are studied in this paper, where T-Net acts as a Classifier and a Clusterer, respectively. The learning procedure of Task-Oriented GAN consists of two main stages. In the first stage, G-Net and D-Net vie with each other like that in a general GAN; in the second stage, G-Net is adjusted and oriented by T-Net so that more samples, which are benefit for the task and called fake data, are generated. As a result, Task-Oriented GAN not only has the advantage of GAN (no-assumption data modeling) but also overcomes the disadvantage of GAN (task-free). After learning, fake data are employed to enrich training set and avoid overfitting; so Task-Oriented GAN performs well even if the manual-labeled data are small. To verify the effectiveness of T-Net, a visualized comparison is provided, where some fake digits generated from Task-Oriented GAN are illustrated along with that from GAN. What is more, considering that there is a great difference between PolSAR data and general data, in our PolSAR image classification and clustering tasks, the specific PolSAR information is inserted into the structure of the Task-Oriented GAN. This enables researchers to mine inherent information in PolSAR data without any data hypothesis and find ways for small sample problem at the same time. Experiment results tested on three PolSAR images show that the proposed method performs well in dealing with PolSAR image classification and clustering.
Fang Liu 0034, Licheng Jiao, Xu Tang 0004
IEEE Trans. Neural Networks Learn. Syst.1
2019 Local Restricted Convolutional Neural Network for Change Detection in Polarimetric SAR Images
abstract
To detect changed areas in multitemporal polarimetric synthetic aperture radar (SAR) images, this paper presents a novel version of convolutional neural network (CNN), which is named local restricted CNN (LRCNN). CNN with only convolutional layers is employed for change detection first, and then LRCNN is formed by imposing a spatial constraint called local restriction on the output layer of CNN. In the training of CNN/LRCNN, the polarimetric property of SAR image is fully used instead of manual labeled pixels. As a preparation, a similarity measure for polarimetric SAR data is proposed, and several layered difference images (LDIs) of polarimetric SAR images are produced. Next, the LDIs are transformed into discriminative enhanced LDIs (DELDIs). CNN/LRCNN is trained to model these DELDIs by a regression pretraining, and then a classification fine-tuning is conducted with some pseudolabeled pixels obtained from DELDIs. Finally, the change detection result showing changed areas is directly generated from the output of the trained CNN/LRCNN. The relation of LRCNN to the traditional way for change detection is also discussed to illustrate our method from an overall point of view. Tested on one simulated data set and two real data sets, the effectiveness of LRCNN is certified and it outperforms various traditional algorithms. In fact, the experimental results demonstrate that the proposed LRCNN for change detection not only recognizes different types of changed/unchanged data, but also ensures noise insensitivity without losing details in changed areas.
Fang Liu 0034, Licheng Jiao, Xu Tang 0004, Shuyuan Yang 0001, Wenping Ma 0001, Biao Hou
IEEE Trans. Neural Networks Learn. Syst.1
2016 POL-SAR Image Classification Based on Wishart DBN and Local Spatial Information
abstract
Inspired by a popular deep neural network, i.e., deep belief network (DBN), a novel method for polarimetric synthetic aperture radar (POL-SAR) image classification is proposed in this paper. For the particularity of POL-SAR data, a new type of restricted Boltzmann machine (RBM) is specially defined, which we name the Wishart-Bernoulli RBM (WBRBM), and is used to form a deep network named as Wishart DBN (W-DBN). Numerous unlabeled POL-SAR pixels are made full use of in the modeling of POL-SAR pixels by W-DBN. In addition, the coherency matrix is used directly to represent a POL-SAR pixel without any manual feature extraction, which is simple and time saving. Local spatial information, together with the confusion matrix, is used in this paper to clean the preliminary classification result obtained by the method based on W-DBN. Making full use of the prior knowledge of POL-SAR data and local spatial information, the proposed method overcomes shortcomings of traditional methods, in which they are sensitive to extracted features and slow to execute. The experiments, tested on three POL-SAR data sets, show that the proposed method produces better results and is much faster than traditional methods.
Fang Liu 0034, Licheng Jiao, Biao Hou, Shuyuan Yang 0001
IEEE Trans. Geosci. Remote. Sens.1
2016 Wishart Deep Stacking Network for Fast POLSAR Image Classification
abstract
Inspired by the popular deep learning architecture, deep stacking network (DSN), a specific deep model for polarimetric synthetic aperture radar (POLSAR) image classification is proposed in this paper, which is named Wishart DSN (W-DSN). First of all, a fast implementation of Wishart distance is achieved by a special linear transformation, which speeds up the classification of POLSAR image and makes it possible to use this polarimetric information in the following neural network (NN). Then, a single-hidden-layer NN based on the fast Wishart distance is defined for POLSAR image classification, which is named Wishart network (WN) and improves the classification accuracy. Finally, a multi-layer NN is formed by stacking WNs, which is in fact the proposed deep learning architecture W-DSN for POLSAR image classification and improves the classification accuracy further. In addition, the structure of WN can be expanded in a straightforward way by adding hidden units if necessary, as well as the structure of the W-DSN. As a preliminary exploration on formulating specific deep learning architecture for POLSAR image classification, the proposed methods may establish a simple but clever connection between POLSAR image interpretation and deep learning. The experiment results tested on real POLSAR image show that the fast implementation of Wishart distance is very efficient (a POLSAR image with 768 000 pixels can be classified in 0.53 s), and both the single-hidden-layer architecture WN and the deep learning architecture W-DSN for POLSAR image classification perform well and work efficiently.
Licheng Jiao, Fang Liu 0034
IEEE Trans. Image Process.2