VLDB 2026 Research / reviewers in the wild / expert
Xianping Ma
dblp:321/3514
· DBLP profile ↗
23ranked-venue papers
9as first author
23since 2021 · last 2026
0000-0002-2180-2964ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 19 · 8 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Zero-shot domain adaptation for remote sensing image classification with vision-language models
Ziyao Wang 0001, Chengxuan Pei, Xianping Ma, Man-On Pun |
Neurocomputing | 3 |
| 2026 | LTTS-GAN: A long-term time series generative adversarial network
Xianping Ma, Man-On Pun, Zhimin Cheng |
Signal Process. | 1 |
| 2026 | Rule-Semantic Generative Calibration Blur Detection for UAV Imagery
Yihan Wen, Zhuo Zhang 0020, Xianping Ma, Peipei Zhu, Jinglei Li, Guanchong Niu, Qiguang Miao |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Auto-Prompting SAM for Weakly Supervised Landslide ExtractionabstractWeakly supervised landslide extraction aims to identify landslide regions from remote sensing data using models trained with weak labels, particularly image-level labels. However, it is often challenged by the imprecise boundaries of the extracted objects due to the lack of pixel-wise supervision and the properties of landslide objects. To tackle these issues, we propose a simple yet effective method by auto-prompting the Segment Anything Model (SAM), i.e., APSAM. Instead of depending on high-quality class activation maps (CAMs) for pseudo-labeling or fine-tuning SAM, our method directly yields fine-grained segmentation masks from SAM inference through prompt engineering. Specifically, it adaptively generates hybrid prompts from the CAMs obtained by an object localization network. To provide sufficient information for SAM prompting, an adaptive prompt generation (APG) algorithm is designed to fully leverage the visual patterns of CAMs, enabling the efficient generation of pseudo-masks for landslide extraction. These informative prompts are able to identify the extent of landslide areas (box prompts) and denote the centers of landslide objects (point prompts), guiding SAM in landslide segmentation. Experimental results on high-resolution aerial and satellite datasets demonstrate the effectiveness of our method, achieving improvements of at least 3.0% in F1 score and 3.69% in IoU compared to other state-of-the-art methods. The source codes and datasets will be available at https://github.com/zxk688. Xianping Ma, Weikang Yu, Pedram Ghamisi |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2025 | A Unified Framework With Multimodal Fine-Tuning for Remote Sensing Semantic SegmentationabstractMultimodal remote sensing data, acquired from diverse sensors, offer a comprehensive and integrated perspective of the Earth’s surface. Leveraging multimodal fusion techniques, semantic segmentation enables detailed and accurate analysis of geographic scenes, surpassing single-modality approaches. Building on advancements in vision foundation models, particularly the Segment Anything Model (SAM), this study proposes a unified framework incorporating a novel Multimodal Fine-tuning Network (MFNet) for remote sensing semantic segmentation. The proposed framework is designed to seamlessly integrate with various fine-tuning mechanisms, demonstrated through the inclusion of Adapter and Low-Rank Adaptation (LoRA) as representative examples. This extensibility ensures the framework’s adaptability to other emerging fine-tuning strategies, allowing models to retain SAM’s general knowledge while effectively leveraging multimodal data. Additionally, a pyramid-based Deep Fusion Module (DFM) is introduced to integrate high-level geographic features across multiple scales, enhancing feature representation prior to decoding. This work also highlights SAM’s robust generalization capabilities with Digital Surface Model (DSM) data, a novel application. Extensive experiments on three benchmark multimodal remote sensing datasets, ISPRS Vaihingen, ISPRS Potsdam and MMHunan, demonstrate that the proposed MFNet significantly outperforms existing methods in multimodal semantic segmentation, setting a new standard in the field while offering a versatile foundation for future research and applications. The source code for this work is accessible at https://github.com/sstary/SSRS. Xianping Ma, Man-On Pun, Bo Huang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2025 | Source-Free Multitarget Unsupervised Domain Adaptation for Cross-City Local Climate Zone ClassificationabstractLocal Climate Zones (LCZs) offer a standardized urban classification system critical for studying climate variations and the urban heat island effect. Large-scale LCZ mapping facilitates cross-city climate comparisons. Remote sensing (RS), particularly supervised learning, has become the primary LCZ classification method due to satellite imagery’s broad coverage and high resolution. However, current RS approaches face challenges, including high labeling demands and limited transferability. In this work, we aim to leverage limited labeled data from a source city (i.e., source domain) to improve LCZ classification performance for multiple unlabeled target cities (i.e., target domains). To address these challenges, this work proposes a novel Source-free Multi-target Unsupervised Domain Adaptation (SFMT-UDA) framework for cross-city LCZ classification. Our approach operates in two stages: first performing single-target adaptation between source and target domains using a self-supervised pseudo-labeling module enhanced by an LCZ class similarity matrix to reduce label noise, then conducting multi-target adaptation among target domains through a Multi-head Multi-target Domain Adaptation (MH-MTDA) network with an integrated style-transfer module for consistent self-training. Extensive experiments on the newly developed VHRLCZ dataset demonstrate that the proposed SFMT-UDA framework achieves superior performance compared to state-of-the-art methods, showing significant improvements in overall accuracy across multiple target domains. The related code and data are available at https://github.com/ctrlovefly/SFMT-UDA. Qianqian Wu 0004, Yinhe Liu, Yanfei Zhong, Kexin Lin, Xianping Ma, Man-On Pun |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | Multimodal Visual-Language Prompt Network for Remote Sensing Few-Shot SegmentationabstractFew-shot segmentation (FSS) aims to segment objects of interest in a query image using a limited set of support images. However, most existing FSS methods are designed for natural images. When extended to remote sensing scenes characterized by extreme intra-class variations and complex backgrounds, these methods struggle to provide robust segmentation guidance, leading to severe performance degradation. To address the aforementioned issues, we propose a multimodal visual-language prompt network (MVLPNet), which employs a collaborative optimization strategy for visual-textual features to tackle the remote sensing FSS task. Specifically, MVLPNet consists of a textual-visual consistency enhancement (TVCE) module and a prototype-guided semantic alignment (PGSA) module. To overcome the limited support set for better guiding the query segmentation, we propose a TVCE module that leverages the contrastive language-image pre-training model (CLIP) to capture category-specific text embeddings. An optimal transport (OT) plan is then established to tightly align these text embeddings with the visual features of query image, thereby extracting semantic information from the query image itself to mitigate the extreme intra-class variation in remote sensing images. Furthermore, a PGSA module is proposed to suppress interference caused by complex background regions. By aggregating lost foreground regions, more comprehensive support features are extracted. Then, the query and support features are precisely matched to activate consistent foreground regions, rather than ambiguously matching the query features via a single prototype or multiple prototypes. Extensive experiments on the iSAID-5i and LoveDA-2i datasets have demonstrated that our method achieves the state of the art. The code is available https://github.com/Gritiii/MVLPNet. Zhenhao Yang, Fukun Bi, Jianhong Han, Xianping Ma |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | A Novel Automated Urban Building Analysis Framework Based on GPT and SAMabstractRapid urban development necessitates advanced methodologies for efficiently acquiring and analyzing detailed building information. This study proposes an automated framework, named Urban Street Buildings Scanner (USBS), which combines geospatial data and deep learning techniques to analyze street-level imagery, thereby providing valuable insights into building features. Specifically, we leverage the OSMnx library for geospatial data retrieval before exploiting coordinate sorting and interpolation techniques to extract information about buildings on both roadsides from street images. After that, Segment Anything Model (SAM) and GPT are utilized to segment each building in street-view images for analysis. YOLOv7 is used to further identify the windows of individual buildings and to analyze the buildings in a more detailed way. Meanwhile, GPT can also estimate the numbers of floors and the heights of the buildings, providing a more comprehensive street view analysis. The experimental results demonstrate the efficacy of the proposed framework in effectively extracting and analyzing building characteristics from street-level images. Visual representations and statistical data illustrate the successful application of the framework, providing valuable information for urban planning and development. Yuchao Sun, Xianping Ma, Yizhen Yan, Man-On Pun, Bo Huang 0001 |
IGARSS | 2 |
| 2024 | A Sam-Empowered Dual-Stream Framework for Scene-Level Local Climate Zone Classification Using Google Earth and Sentinel ImagesabstractRecent advancements in remote sensing (RS)-based methods have shown remarkable effectiveness in large-scale local climate zone classification. However, conventional convolutional neural network (CNN)-based methods encounter limitations in effectively incorporating ground object priors. Additionally, commonly used medium-scale data sources, such as Sentinel-2, face challenges in capturing detailed ground object information. In light of these obstacles, we propose a data fusion method that integrates ground object priors extracted from high-resolution Google imagery with Sentinel-2 multi-spectral imagery. The proposed method introduces a novel dual-stream fusion framework (DF4LCZ-Net), which fully leverages instance-based location features from Google imagery and combines them with the scene-level spatial-spectral features extracted from Sentinel-2 images. For effective feature extraction of Google imagery, a SAM-based Graph Convolutional Network (GCN) branch is designed in the framework. We conducted experiments on a newly created multi-source remote sensing image dataset, and the final classification results show the superiority of our method. Qianqian Wu 0004, Xianping Ma, Jialu Sui, Man-On Pun |
IGARSS | 2 |
| 2024 | Extreme Low Bitrate Image Compression System for Mobile DeploymentabstractEnd-to-end image compression has achieved satisfactory results in recent studies. However, existing methods suffer from high complexity of complicated neural network computation and cannot be directly deployed on mobile devices due to the limitations of computing ability and storage. Therefore, considering the resource and computing ability constrains of the mobile devices, we make a trade-off in this paper between rate-distortion (R-D) performance, inference time, and model complexity. Then we design a novel lightweight perceptual image compression framework to alleviate the storage and complexity burden of mobile devices. Moreover, we design a hardware-friendly deployment scheme to apply the proposed compression framework on high-end mobile devices, which can achieve efficient image compression. Based on the above structures, we propose the first mobile system that achieves image compression on mobile devices. The supplementary material of our system demo is on https://sigport.org/documents/extreme-low-bitrate-Image-compression-system-mobile-deployment. Wenhong Duan, Xianping Ma, Jianhui Chang, Shanshe Wang, Siwei Ma 0001, Chuanmin Jia |
MMSP | 3 |
| 2024 | RS3Mamba: Visual State Space Model for Remote Sensing Image Semantic SegmentationabstractSemantic segmentation of remote sensing images is a fundamental task in geoscience research. However, convolutional neural networks (CNNs) and transformers have some significant shortcomings. The former are limited by insufficient long-range modeling capabilities, while the latter are hampered by computational complexity. Recently, a novel visual state space (VSS) model represented by Mamba has emerged, capable of modeling long-range relationships with linear computability. In this research, we propose a novel dual-branch network named remote sensing image semantic segmentation Mamba (RS3Mamba) designed specifically for remote sensing tasks. RS3Mamba uses VSS blocks to construct an auxiliary branch, providing additional global information to a convolution-based main branch. Moreover, considering the distinct characteristics of the two branches, we introduce a collaborative completion module (CCM) to refine and fuse features from the dual-encoder using a novel adaptive mechanism. Through experiments on two widely used datasets, the proposed RS3Mamba was found to outperform the state-of-the-art methods in terms of mIoU with 0.66% on ISPRS Vaihingen and 1.70% on LoveDA Urban, demonstrating its effectiveness and potential. The source code is available athttps://github.com/sstary/SSRS. Xianping Ma, Man-On Pun |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2024 | SAM-Assisted Remote Sensing Imagery Semantic Segmentation With Object and Boundary ConstraintsabstractSemantic segmentation of remote sensing imagery plays a pivotal role in extracting precise information for diverse downstream applications. Recent development of the segment anything model (SAM), an advanced general-purpose segmentation model, has revolutionized this field, presenting new avenues for accurate and efficient segmentation. However, SAM is limited to generating segmentation results without class information. Meanwhile, the segmentation map predicted by current methods generally exhibits excessive fragmentation and inaccuracy of boundary. This article introduces a streamlined framework designed to leverage the raw output of SAM by exploiting two novel concepts called SAM-generated object (SGO) and SAM-generated boundary (SGB). More specifically, we propose a novel object consistency loss and further introduce a boundary preservation loss in this work. Considering the content characteristics of SGO, we introduce the concept of object consistency to leverage segmented regions lacking semantic information. By imposing constraints on the consistency of predicted values within objects, the object consistency loss aims to enhance semantic segmentation performance. Furthermore, the boundary preservation loss capitalizes on the distinctive features of SGB by directing the model’s attention to the boundary information of the object. Experimental results on two well-known datasets, ISPRS Vaihingen and LoveDA Urban, demonstrate the effectiveness and broad applicability of the proposed method. The source code for this work is accessible athttps://github.com/sstary/SSRS. Xianping Ma, Qianqian Wu 0004, Man-On Pun, Bo Huang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | Decomposition-Based Unsupervised Domain Adaptation for Remote Sensing Image Semantic SegmentationabstractUnsupervised domain adaptation (UDA) techniques are vital for semantic segmentation in geosciences, effectively utilizing remote sensing imagery across diverse domains. However, most existing UDA methods, which focus on domain alignment at the high-level feature space, struggle to simultaneously retain local spatial details and global contextual semantics. To overcome these challenges, a novel decomposition scheme is proposed to guide domain-invariant representation learning. Specifically, multiscale high/low-frequency decomposition (HLFD) modules are proposed to decompose feature maps into high- and low-frequency components across different subspaces. This decomposition is integrated into a fully global-local generative adversarial network (GLGAN) that incorporates global-local transformer blocks (GLTBs) to enhance the alignment of decomposed features. By integrating the HLFD scheme and the GLGAN, a novel decomposition-based UDA framework called De-GLGAN is developed to improve the cross-domain transferability and generalization capability of semantic segmentation models. Extensive experiments on two UDA benchmarks, namely ISPRS Potsdam and Vaihingen, and LoveDA Rural and Urban, demonstrate the effectiveness and superiority of the proposed approach over existing state-of-the-art UDA methods. The source code for this work is accessible athttps://github.com/sstary/SSRS. Xianping Ma, Xingchen Ding, Man-On Pun, Siwei Ma 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | A Multilevel Multimodal Fusion Transformer for Remote Sensing Semantic SegmentationabstractAccurate semantic segmentation of remote sensing data plays a crucial role in the success of geoscience research and applications. Recently, multimodal fusion-based segmentation models have attracted much attention due to their outstanding performance as compared to conventional single-modal techniques. However, most of these models perform their fusion operation using convolutional neural networks (CNN) or the vision transformer (Vit), resulting in insufficient local-global contextual modeling and representative capabilities. In this work, a multilevel multimodal fusion scheme called FTransUNet is proposed to provide a robust and effective multimodal fusion backbone for semantic segmentation by integrating both CNN and Vit into one unified fusion framework. Firstly, the shallow-level features are first extracted and fused through convolutional layers and shallow-level feature fusion (SFF) modules. After that, deep-level features characterizing semantic information and spatial relationships are extracted and fused by a well-designed Fusion Vit (FVit). It applies Adaptively Mutually Boosted Attention (Ada-MBA) layers and Self-Attention (SA) layers alternately in a three-stage scheme to learn cross-modality representations of high inter-class separability and low intra-class variations. Specifically, the proposed Ada-MBA computes SA and Cross-Attention (CA) in parallel to enhance intra- and cross-modality contextual information simultaneously while steering attention distribution towards semantic-aware regions. As a result, FTransUNet can fuse shallow-level and deep-level features in a multilevel manner, taking full advantage of CNN and transformer to accurately characterize local details and global semantics, respectively. Extensive experiments confirm the superior performance of the proposed FTransUNet compared with other multimodal fusion approaches on two fine-resolution remote sensing datasets, namely ISPRS Vaihingen and Potsdam. The source code in this work is available at https://github.com/sstary/SSRS. Xianping Ma, Man-On Pun |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | GCD-DDPM: A Generative Change Detection Model Based on Difference-Feature-Guided DDPMabstractDeep learning (DL)-based methods have recently shown great promise in bitemporal change detection (CD). Existing discriminative methods based on convolutional neural networks (CNNs) and Transformers rely on discriminative representation learning for change recognition while struggling with exploring local and long-range contextual dependencies. As a result, it is still challenging to obtain fine-grained and robust CD maps in diverse ground scenes. To cope with this challenge, this work proposes a generative CD model called GCD-DDPM to directly generate CD maps by exploiting the denoising diffusion probabilistic model (DDPM), instead of classifying each pixel into changed or unchanged categories. Furthermore, the difference conditional encoder (DCE), is designed to guide the generation of CD maps by exploiting multilevel difference features. Leveraging the variational inference (VI) procedure, GCD-DDPM can adaptively recalibrate the CD results through an iterative inference process, while accurately distinguishing subtle and irregular changes in diverse scenes. Finally, a noise suppression-based semantic enhancer (NSSE) is specifically designed to mitigate noise in the current step’s change-aware feature representations from the CD Encoder. This refinement, serving as an attention map, can guide subsequent iterations while enhancing CD accuracy. Extensive experiments on four high-resolution CD datasets (CDD) confirm the superior performance of the proposed GCD-DDPM. The code for this work will be available athttps://github.com/udrs/GCD. Yihan Wen, Xianping Ma, Man-On Pun |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | DF4LCZ: A SAM-Empowered Data Fusion Framework for Scene-Level Local Climate Zone ClassificationabstractRecent advances in remote sensing technologies have highlighted their capability for accurate classification of local climate zones (LCZs). However, traditional methods using convolutional neural networks (CNNs) often fall short of effectively incorporating prior knowledge of ground objects. In addition, data sources such as Sentinel-2 struggle with capturing detailed information on ground objects. To address these issues, we introduce a novel data fusion approach that combines high-resolution Google imagery, which provides ground object priors, with Sentinel-2 multispectral imagery. Our method, the Dual-stream Fusion framework for LCZ classification (DF4LCZ), merges instance-based location features from Google imagery and spatial-spectral features from Sentinel-2. This framework is enhanced by a graph convolutional network (GCN) module, powered by the segment anything model (SAM), to improve feature extraction from Google imagery. Concurrently, a 3D-CNN architecture is utilized to process the spectral-spatial features of Sentinel-2 imagery. The effectiveness of DF4LCZ is demonstrated through experiments conducted on a specialized multisource remote sensing image dataset for LCZ classification. The related code and dataset are available athttps://github.com/ctrlovefly/DF4LCZ. Qianqian Wu 0004, Xianping Ma, Jialu Sui, Man-On Pun |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | CycleGAN-Based Cloud Removal from a Feature Enhancement Perspective by TransformerabstractCloud removal has attracted significant research attention in various remote sensing applications, such as object detection and semantic segmentation. In this study, the classical cycle-consistent generative adversarial network (CycleGAN) is adopted for suppressing clouds from a feature enhancement perspective by capitalizing on the transformer architecture. Specifically, a transformer-based feature enhancement (TFE) module is proposed to extract high-level cloud-clear features by leveraging the Swin transformer’s capability of building long-range dependencies. As a result, the proposed TFE module can eliminate clouds in remote sensing images while retaining cloud-free regions unchanged. Extensive simulation experiments on the RICE dataset are conducted to substantiate the impressive performance of the proposed TFE model as compared to several existing cloud-removal methods. Yiming Huang 0001, Xianping Ma, Man-On Pun |
IGARSS | 2 |
| 2023 | Machine Learning-Based Approach For Landslide Susceptibility Mapping Using Multimodal DataabstractIn this work, a practical Machine Learning (ML) approach is proposed to produce Landslide Susceptibility Mapping (LSM) by exploiting multiple data sources. In contrast to conventional ML-based methods that consider a large number of factors, the proposed method fuses only three carefully selected factors, namely inventory Landslide and Debris Flow (LDF) data, rainfall data and forest change information to model and predict landslides and debris flow. More specifically, a Global Deforestation Detection Algorithm (GDDA) based on Synthetic Aperture Radar (SAR) and Convolutional Neural Network (CNN) are first developed to generate a hazard map by identifying areas of substantial forest changes. Capitalizing on rainfall forecast data derived from the Doppler Weather Radar (DWR) over the areas of severe deforestation detected by GDDA, an SVM-based classifier is established to predict the likelihood of the occurrence of landslide and debris flow. Using two real-world landslide and debris flows in 2021, we demonstrate that the proposed LSM system was able to provide early warning in advance. Xianping Ma, Man-On Pun |
IGARSS | 1 |
| 2023 | DTRN: Dual Transformer Residual Network for Remote Sensing Super-ResolutionabstractThe synergy of the transformer and the convolutional neural network (CNN) has been well regarded as a promising technique for single image super-resolution (SISR) based on low-quality satellite remote sensing images. In this work, a Dual Transformer Residual Network (DTRN) consisting of one transformer branch and one CNN-based residual branch is proposed. More specifically, the transformer branch is designed to capture the global relationships of feature maps by exploiting three pairs of token embedding blocks and convolutional transformer blocks (CTB). Furthermore, the residual branch employs several residual blocks (resblocks) to effectively learn hierarchical features through global feature fusion. Extensive experiments on a large-scale remote sensing dataset called OLI2MSI confirm the superior performance of the proposed DTRN as compared to the existing SISR methods. Jialu Sui, Xianping Ma, Man-On Pun |
IGARSS | 2 |
| 2023 | MDAFNet: Monocular Depth-Assisted Fusion Networks for Semantic Segmentation of Complex Urban Remote Sensing DataabstractThis work proposes an end-to-end Monocular Depth-Assisted Fusion Network (MDAFNet) for semantic segmentation of complex urban remote sensing data. The proposed MDAFNet consists of a Monocular Depth Estimation Network (MDENet) and a Crossmodal Fusion Network (CFNet). More specifically, the MDENet first generates the earth surface depth information while the CFNet fuses the generated depth information and RGB images to address the segmentation task. In particular, the MDENet is capable of effectively extracting features of the ground surface while overcoming artifacts such as building shadows. Furthermore, the CFNet is designed to perform segmentation by extracting and fusing semantic information from generated depth information and Red-Green-Blue (RGB) images. Extensive experiments performed on a large-scale fine-resolution remote sensing dataset named the ISPRS Vaihingen confirm that the proposed MDAFNet outperforms conventional crossmodal models equipped with Digital Surface Model information. Xiaochen Xiu, Xianping Ma, Man-On Pun |
IGARSS | 2 |
| 2023 | Weakly Supervised Local-Global Anchor Guidance Network for Landslide Extraction With Image-Level AnnotationsabstractWeakly supervised learning using image-level annotations has become a popular choice for reducing labeling efforts of remote sensing object extraction. Existing methods exploit inter-pixel relations within an individual image patch for object localizations. When facing large-scale remote sensing images, it is still challenging to obtain global semantic contexts across image patches for feature representation, resulting in inaccurate object localizations. To remedy these issues, we propose a local-global anchor guidance network (LGAGNet) for weakly supervised landslide extraction. Specifically, a structure-aware object locating (SOL) module is developed to capture the spatial structure of landslide objects and extract local category anchors containing informative feature embeddings. Furthermore, we leverage a global anchor aggregation (GAA) module to excavate semantic patterns across image patches based on a memory bank, which is then used as additional context cues to enhance the feature presentation through a cross-attention mechanism. Finally, a hybrid loss function is designed to guide the network training, considering category-aware semantic contrasts and local activation consistency. Experimental results on high-resolution aerial and satellite image datasets verify the effectiveness of the proposed approach on landslide extraction. Weikang Yu, Xianping Ma, Xudong Kang |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2023 | Unsupervised Domain Adaptation Augmented by Mutually Boosted Attention for Semantic Segmentation of VHR Remote Sensing ImagesabstractThis work investigates unsupervised domain adaptation (UDA)-based semantic segmentation of very high-resolution (VHR) remote sensing (RS) images from different domains. Most existing UDA methods resort to generative adversarial networks (GANs) to cope with the domain shift problem caused by the discrepancies across different domains. However, these GAN-based UDA methods directly align two domains in the appearance, latent, or output space based on convolutional neural networks (CNNs), making them ineffective in exploiting long-range dependencies across the high-level feature maps derived from different domains. Unfortunately, such high-level features play an essential role in characterizing RS images with complex content. To circumvent this obstacle, a mutually boosted attention transformer (MBATrans) is proposed to capture cross-domain dependencies of semantic feature representations in this work. Compared with conventional UDA methods, MBATrans can significantly reduce domain discrepancies by capturing transferable features using global attention. More specifically, MBATrans utilizes a novel mutually boosted attention (MBA) module to align cross-domain feature maps while enhancing domain-general features. Furthermore, a novel GAN-based network with improved discriminative capability is devised by integrating an additional discriminator to learn domain-specific features. Extensive experiments on two large-scale VHR RS datasets, namely, International Society for Photogrammetry and Remote Sensing (ISPRS) Potsdam and Vaihingen, confirm the superior performance of the proposed MBATrans-augmented GAN (MBATA-GAN) architecture. The source code in this work is available athttps://github.com/sstary/SSRS. Xianping Ma, Zhiguo Wang 0005, Man-On Pun |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | MSFNET: Multi-Stage Fusion Network for Semantic Segmentation of Fine-Resolution Remote Sensing DataabstractThis work proposes a Multi-Stage Fusion Network (MSFNet) for semantic segmentation of fine-resolution remote sensing data by exploiting a multi -stage transformer architec-ture. The proposed MSFNet fuses information of different scales and modalities using a multi-stage scheme based on cross-attention mechanism. More specifically, the proposed MSFNet is composed of two Multi-Level Transformers (ML-Trans), one Crossmodal Fusion Transformer (CFTrans) and one Global-Context Augmented Transformer (GCATrans). DMLTrans and CFTrans are designed to fuse features in dif-ferent levels in each modality and high-level crossmodal ab-stract features, respectively, whereas GCATrans enhances the fusion feature of the main modal. Capitalizing on MSFNet, this work demonstrates the fusion of red-green-blue (RGB) remote sensing images and digital surface model (DSM) data. Extensive experiments on large-scale fine-resolution remote sensing data sets, namely the ISPRS Vaihingen, confirm the excellent performance of the proposed architecture as compared to conventional multimodal methods. Xianping Ma, Man-On Pun |
IGARSS | 1 |