Yongqiang Mao

dblp:318/1359 · DBLP profile ↗
← Back
17ranked-venue papers
5as first author
17since 2021 · last 2026
0000-0001-9256-3668ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 11 · 4 first-author · 11 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 RingMoE: Mixture-of-Modality-Experts Multi-Modal Foundation Models for Universal Remote Sensing Image Interpretation
abstract
The rapid advancement of foundation models has revolutionized visual representation learning in a self-supervised manner. However, their application in remote sensing (RS) remains constrained by a fundamental gap: existing models predominantly handle single or limited modalities, overlooking the inherently multi-modal nature of RS observations. Optical, synthetic aperture radar (SAR), and multi-spectral data offer complementary insights that significantly reduce the inherent ambiguity and uncertainty in single-source analysis. To bridge this gap, we introduce RingMoE, a unified multi-modal RS foundation model with 14.7 billion parameters, pre-trained on 400 million multi-modal RS images from nine satellites. RingMoE incorporates three key innovations: 1) A hierarchical Mixture-of-Experts (MoE) architecture comprising modal-specialized, collaborative, and shared experts, effectively modeling intra-modal knowledge while capturing cross-modal dependencies to mitigate conflicts between modal representations; 2) Physics-informed self-supervised learning, explicitly embedding sensor-specific radiometric characteristics into the pre-training objectives; 3) Dynamic expert pruning, enabling adaptive model compression from 14.7B to 1B parameters while maintaining performance, facilitating efficient deployment in Earth observation applications. Evaluated across 23 benchmarks spanning six key RS tasks (i.e., classification, detection, segmentation, tracking, change detection, and depth estimation), RingMoE outperforms existing foundation models and sets new SOTAs, demonstrating remarkable adaptability from single-modal to multi-modal scenarios. Beyond theoretical progress, it has been deployed and trialed in multiple sectors, including emergency response, land management, marine sciences, and urban planning.
Hanbo Bi, Yingchao Feng, Boyuan Tong, Haichen Yu, Yongqiang Mao, Wenhui Diao, Peijin Wang, Yue Yu 0001, Hanyang Peng, Yehong Zhang, Kun Fu 0001, Xian Sun 0001
IEEE Trans. Pattern Anal. Mach. Intell.6
2025 AgMTR: Agent Mining Transformer for Few-Shot Segmentation in Remote Sensing
Hanbo Bi, Yingchao Feng, Yongqiang Mao, Jianning Pei, Wenhui Diao, Xian Sun 0001
Int. J. Comput. Vis.3
2025 Prompt-and-Transfer: Dynamic Class-Aware Enhancement for Few-Shot Segmentation
abstract
For more efficient generalization to unseen domains (classes), most Few-shot Segmentation (FSS) would directly exploit pre-trained encoders and only fine-tune the decoder, especially in the current era of large models. However, such fixed feature encoders tend to be class-agnostic, inevitably activating objects that are irrelevant to the target class. In contrast, humans can effortlessly focus on specific objects in the line of sight. This paper mimics the visual perception pattern of human beings and proposes a novel and powerful prompt-driven scheme, called "Prompt and Transfer" (PAT), which constructs a dynamic class-aware prompting paradigm to tune the encoder for focusing on the interested object (target class) in the current task. Three key points are elaborated to enhance the prompting: 1) Cross-modal linguistic information is introduced to initialize prompts for each task. 2) Semantic Prompt Transfer (SPT) that precisely transfers the class-specific semantics within the images to prompts. 3) Part Mask Generator (PMG) that works in conjunction with SPT to adaptively generate different but complementary part prompts for different individuals. Surprisingly, PAT achieves competitive performance on 4 different tasks including standard FSS, Cross-domain FSS (e.g., CV, medical, and remote sensing domains), Weak-label FSS, and Zero-shot Segmentation, setting new state-of-the-arts on 11 benchmarks.
Hanbo Bi, Yingchao Feng, Wenhui Diao, Peijin Wang, Yongqiang Mao, Kun Fu 0001, Xian Sun 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2025 RingMo-Aerial: An Aerial Remote Sensing Foundation Model With Affine Transformation Contrastive Learning
abstract
Aerial Remote Sensing (ARS) vision tasks present significant challenges due to the unique viewing angle characteristics. Existing research has primarily focused on algorithms for specific tasks, which have limited applicability in a broad range of ARS vision applications. This paper proposes RingMo-Aerial, aiming to fill the gap in foundation model research in the field of ARS vision. A Frequency-Enhanced Multi-Head Self-Attention (FE-MSA) mechanism is introduced to strengthen the model's capacity for small-object representation. Complementarily, an affine transformation-based contrastive learning method improves its adaptability to the tilted viewing angles inherent in ARS tasks. Furthermore, the ARS-Adapter, an efficient parameter fine-tuning method, is proposed to improve the model's adaptability and performance in various ARS vision tasks. Experimental results demonstrate that RingMo-Aerial achieves SOTA performance on multiple downstream tasks. This indicates the practicality and efficacy of RingMo-Aerial in enhancing the performance of ARS vision tasks.
Wenhui Diao, Haichen Yu, Kaiyue Kang, Tong Ling, Yingchao Feng, Hanbo Bi, Libo Ren, Xuexue Li, Yongqiang Mao, Xian Sun 0001
IEEE Trans. Pattern Anal. Mach. Intell.10
2024 Body Joint Boundary Prototype Match for Few-Shot Remote Sensing Semantic Segmentation
abstract
Deep networks require a large number of samples for optimization, so few-shot segmentation in remote sensing scenes is still an open problem. However, this challenge is exacerbated by the feature blurring and aliasing of bodies (low frequency) and boundaries (high frequency). The existing methods usually only focus on the body part of the class, that is, the low-frequency part, and ignore the critical role of boundary information, that is, high-frequency details, on feature representation. In this letter, we propose a novel body joint boundary prototype match (B2PM) approach that aims to enable prior learning of low- and high-frequency information by explicitly modeling the body and boundary features of objects. First, body-aware prototype learning (BodyPL) realizes the adaptive modeling of the body part of the object through a precise farthest point sampling (FPS) initialization algorithm and an adaptive part shift (APS) strategy, which alleviates the feature ambiguity of the body. Second, boundary-aware prototype learning (BoundPL) explicitly models boundary prototypes by building a patch division and assignment strategy to alleviate feature aliasing at boundaries. Finally, prototype match performs prior knowledge aggregation by computing the affinity between query features and support prototypes. Extensive experiments on commonly used benchmarks (iSAID and PASCAL VOC) demonstrate that B2PM improves the state of the art by significant margins.
Yongqiang Mao, Zhizhuo Jiang, Yu Liu 0005, Yaowen Li, Chenggang Yan 0001, Bolun Zheng
IEEE Geosci. Remote. Sens. Lett.1
2024 SCLNet: A Scale-Robust Complementary Learning Network for Object Detection in UAV Images
abstract
Most recent unmanned aerial vehicle (UAV) detectors focus primarily on general challenges such as uneven distribution and occlusion. However, the neglect of scale challenges, which encompass scale variation and small objects, continues to hinder object detection in UAV images. Although existing works propose solutions, they are implicitly modeled and have redundant steps, so detection performance remains limited. One specific work addressing the above scale challenges can help improve the performance of UAV image detectors. Compared to natural scenes, scale challenges in UAV images happen with problems of limited perception in comprehensive scales and poor robustness to small objects. We found that complementary learning is beneficial for the detection model to address the scale challenges. Therefore, the article introduces it to form our scale-robust complementary learning network (SCLNet) in conjunction with the object detection model. The SCLNet consists of two implementations and a cooperation method. In detail, one implementation is based on our proposed scale-complementary decoder and scale-complementary loss function to explicitly extract complementary information as a complement, named comprehensive-scale complementary learning (CSCL). Another implementation is based on our proposed contrastive complement network and contrastive complement loss function to explicitly guide the learning of small objects with the rich texture detail information of the large objects, named interscale contrastive complementary learning (ICCL). In addition, an end-to-end cooperation (ECoop) between two implementations and with the detection model is proposed to exploit each potential. In short, SCLNet forms a more comprehensive representation through feature complementary and improves the representation of small objects through interscale contrast, which in turn comes to improve scale robustness and detection performance. Thorough experiments prove the effectiveness of our SCLNet on Visdrone and UAVDT datasets, including the fact that the novel components included in SCLNet are effective and competitive with many CNN-based and transformer-based methods, among other aspects. In general, our SCLNet can effectively address scale challenges and is a competitive model in UAV image object detection.
Xuexue Li, Wenhui Diao, Yongqiang Mao, Xian Sun 0001
IEEE Trans. Geosci. Remote. Sens.3
2024 SDL-MVS: View Space and Depth Deformable Learning Paradigm for Multiview Stereo Reconstruction in Remote Sensing
abstract
Research on multiview stereo (MVS) based on remote sensing images has promoted the development of large-scale urban 3-D reconstruction. However, remote sensing multiview image data suffer from the problems of occlusion and uneven brightness between views during acquisition, which leads to the problem of blurred details in depth estimation. To solve the above problem, we reexamine the deformable learning method in the MVS task and propose a novel paradigm based on view space and depth deformable learning (SDL-MVS), aiming to learn deformable interactions of features in different view spaces and deformably model the depth ranges and intervals to enable high accurate depth estimation. Specifically, to solve the problem of view noise caused by occlusion and uneven brightness, we propose a progressive space deformable sampling (PSS) mechanism, which performs deformable learning of sampling points in the 3-D frustum space and the 2-D image space in a progressive manner to embed source features to the reference feature adaptively. To further optimize the depth, we introduce depth hypothesis deformable discretization (DHD), which achieves precise positioning of the depth prior by adaptively adjusting the depth range hypothesis and performing deformable discretization of the depth interval hypothesis. Finally, our SDL-MVS achieves explicit modeling of occlusion and uneven brightness faced in MVS through the deformable learning paradigm of view space and depth, achieving accurate multiview depth estimation. Extensive experiments on LuoJia-MVS and WHU datasets show that our SDL-MVS reaches state-of-the-art performance. It is worth noting that our SDL-MVS achieves a mean absolute error (MAE) error of 0.086 and an accuracy of 98.9% for Acc$_{\lt 0.6\,\text {m}}$and 98.9% for Acc$_{\lt 3-\text {interval}}$on the LuoJia-MVS dataset under the premise of three views as input.
Yongqiang Mao, Hanbo Bi, Liangyu Xu, Kaiqiang Chen, Zhirui Wang 0003, Xian Sun 0001, Kun Fu 0001
IEEE Trans. Geosci. Remote. Sens.1
2024 TAFormer: A Unified Target-Aware Transformer for Video and Motion Joint Prediction in Aerial Scenes
abstract
As drone technology advances, using unmanned aerial vehicles for aerial surveys has become the dominant trend in modern low-altitude remote sensing. The surge in aerial video data necessitates accurate prediction for future scenarios and motion states of the interested target, particularly in applications like traffic management and disaster response. Existing video prediction methods focus solely on predicting future scenes (video frames), suffering from the neglect of explicitly modeling target’s motion states, which is crucial for aerial video interpretation. To address this issue, we introduce a novel task called Target-Aware Aerial Video Prediction, aiming to simultaneously predict future scenes and motion states of the target. Further, we design a model specifically for this task, named TAFormer, which provides a unified modeling approach for both video and target motion states. Specifically, we introduce Spatiotemporal Attention (STA), which decouples the learning of video dynamics into spatial static attention and temporal dynamic attention, effectively modeling the scene appearance and motion. Additionally, we design an Information Sharing Mechanism (ISM), which elegantly unifies the modeling of video and target motion by facilitating information interaction through two sets of messenger tokens. Moreover, to alleviate the difficulty of distinguishing targets in blurry predictions, we introduce Target-Sensitive Gaussian Loss (TSGL), enhancing the model’s sensitivity to both target’s position and content. Extensive experiments on UAV123VP and VisDroneVP (derived from single-object tracking datasets) demonstrate the exceptional performance of TAFormer in target-aware video prediction, showcasing its adaptability to the additional requirements of aerial video interpretation for target awareness.
Liangyu Xu, Wanxuan Lu, Yongqiang Mao, Hanbo Bi, Xian Sun 0001, Kun Fu 0001
IEEE Trans. Geosci. Remote. Sens.4
2023 Breaking Immutable: Information-Coupled Prototype Elaboration for Few-Shot Object Detection
abstract
Few-shot object detection, expecting detectors to detect novel classes with a few instances, has made conspicuous progress. However, the prototypes extracted by existing meta-learning based methods still suffer from insufficient representative information and lack awareness of query images, which cannot be adaptively tailored to different query images. Firstly, only the support images are involved for extracting prototypes, resulting in scarce perceptual information of query images. Secondly, all pixels of all support images are treated equally when aggregating features into prototype vectors, thus the salient objects are overwhelmed by the cluttered background. In this paper, we propose an Information-Coupled Prototype Elaboration (ICPE) method to generate specific and representative prototypes for each query image. Concretely, a conditional information coupling module is introduced to couple information from the query branch to the support branch, strengthening the query-perceptual information in support features. Besides, we design a prototype dynamic aggregation module that dynamically adjusts intra-image and inter-image aggregation weights to highlight the salient information useful for detecting query images. Experimental results on both Pascal VOC and MS COCO demonstrate that our method achieves state-of-the-art performance in almost all settings. Code will be available at: https://github.com/lxn96/ICPE.
Wenhui Diao, Yongqiang Mao, Junxi Li, Peijin Wang, Xian Sun 0001, Kun Fu 0001
AAAI3
2023 Light: Joint Individual Building Extraction and Height Estimation from Satellite Images Through a Unified Multitask Learning Network
abstract
Building extraction and height estimation are two important basic tasks in remote sensing image interpretation, which are widely used in urban planning, real-world 3D construction, and other fields. Most of the existing research regards the two tasks as independent studies. Therefore the height information cannot be fully used to improve the accuracy of building extraction and vice versa. In this work, we combine the individuaL buIlding extraction and heiGHt estimation through a unified multiTask learning network (LIGHT) for the first time, which simultaneously outputs a height map, bounding boxes, and a segmentation mask map of buildings. Specifically, LIGHT consists of an instance segmentation branch and a height estimation branch. In particular, so as to effectively unify multi-scale feature branches and alleviate feature spans between branches, we propose a Gated Cross Task Interaction (GCTI) module that can efficiently perform feature interaction between branches. Experiments on the DFC2023 dataset show that our LIGHT can achieve superior performance, and our GCTI module with ResNet 101 as the backbone can significantly improve the performance of multitask learning by 2.8% AP50 and 6.5% δ1, respectively.
Yongqiang Mao, Xian Sun 0001, Xingliang Huang, Kaiqiang Chen
IGARSS1
2023 Not Just Learning From Others but Relying on Yourself: A New Perspective on Few-Shot Segmentation in Remote Sensing
abstract
Few-shot segmentation (FSS) is proposed to segment unknown class targets with just a few annotated samples. Most current FSS methods follow the paradigm of mining the semantics from the support images to guide the query image segmentation. However, such a pattern of ‘learning from others’ struggles to handle the extreme intra-class variation, preventing FSS from being directly generalized to remote sensing scenes. To bridge the gap of intra-class variance, we develop a Dual-Mining network named DMNet for cross-image mining and self-mining, meaning that it no longer focuses solely on support images but pays more attention to the query image itself. Specifically, we propose a Class-public Region Mining (CPRM) module to effectively suppress irrelevant feature pollution by capturing the common semantics between the support-query image pair. The Class-specific Region Mining (CSRM) module is then proposed to continuously mine the class-specific semantics of the query image itself in a ‘filtering’ and ‘purifying’ manner. In addition, to prevent the co-existence of multiple classes in remote sensing scenes from exacerbating the collapse of FSS generalization, we also propose a new Known-class Meta Suppressor (KMS) module to suppress the activation of known-class objects in the sample. Extensive experiments on the iSAID and LoveDA remote sensing datasets have demonstrated that our method sets the state-of-the-art with a minimum number of model parameters. Significantly, our model with the backbone of Resnet-50 achieves the mIoU of 49.58% and 51.34% on iSAID under 1-shot and 5-shot settings, outperforming the state-of-the-art method by 1.8% and 1.12%, respectively. The code is publicly available at https://github.com/HanboBizl/DMNet/.
Hanbo Bi, Yingchao Feng, Yongqiang Mao, Wenhui Diao, Xian Sun 0001
IEEE Trans. Geosci. Remote. Sens.4
2023 Few-Shot Object Detection in Aerial Imagery Guided by Text-Modal Knowledge
abstract
Few-shot object detection (FSOD) has received numerous attention due to the difficulty and time-consuming of labeling objects. Recent researches achieve excellent performance in a natural scene by only using a few instances of novel classes to fine-tune the last prediction layer of the model well-trained on plentiful base data. However, compared with natural scene objects with a single direction and small size variety, the direction and size of the objects in remote sensing images (RSIs) vary greatly. The methods proposed for the natural scene cannot be directly applied to RSIs. In this article, we first propose a strong baseline for RSIs. It fine-tunes all detector components acting on high-level features and effectively improves the performance of novel classes. Further analyzing the results of the baseline, we find that the error for novel classes is mainly concentrated in classification. It misclassifies novel classes as confusable base classes or backgrounds due to the difficulty in extracting generalized information from limited instances. As is well-known, text-modal knowledge can highly summarize the generalized and unique characteristics of categories. Thus, we introduce text-modal descriptions for each category and propose an FSOD method guided by TExt-MOdal knowledge, called TEMO. Specifically, a text-modal knowledge extractor and a cross-modal assembly module are proposed to extract text features and fuse the text-modal features into visual-modal features. The fused features greatly reduce the classification confusion of novel classes. Furthermore, we introduce a mask strategy and a separation loss to avoid over-fitting and ambiguity of text-modal features. Experimental results on detection in optical remote sensing images (DIOR), Northwestern Polytechnical University (NWPU), and fine-grained object recognition in high-resolution remote sensing imagery (FAIR1M) illustrate that our TEMO achieves state-of-the-art performance in all settings.
Xian Sun 0001, Wenhui Diao, Yongqiang Mao, Junxi Li, Yidan Zhang 0002, Peijin Wang, Kun Fu 0001
IEEE Trans. Geosci. Remote. Sens.4
2023 Elevation Estimation-Driven Building 3-D Reconstruction From Single-View Remote Sensing Imagery
abstract
Building 3D reconstruction from remote sensing images has a wide range of applications in smart cities, photogrammetry and other fields. Methods for automatic 3D urban building modeling typically employ multi-view images as input to algorithms to recover point clouds and 3D models of buildings. However, such models rely heavily on multi-view images of buildings, which are time-intensive and limit the applicability and practicality of the models. To solve these issues, we focus on designing an efficient DSM estimation-driven reconstruction framework (Building3D), which aims to reconstruct 3D building models from the input single-view remote sensing image. Existing DSM estimation networks suffer from the imbalance between local features and global features, which leads to over-smooth DSM estimates at instance boundaries. To address this issue, we propose a Semantic Flow Field-guided DSM Estimation (SFFDE) network, which utilizes the proposed concept of elevation semantic flow to achieve the registration of local and global features. First, in order to make the network semantics globally aware, we propose an Elevation Semantic Globalization (ESG) module to realize the semantic globalization of instances. Further, in order to alleviate the semantic span of global features and original local features, we propose a Local-to-Global Elevation Semantic Registration (L2G-ESR) module based on elevation semantic flow. Our Building3D is rooted in the SFFDE network for building elevation prediction, synchronized with a building extraction network for building masks, and then sequentially performs point cloud reconstruction and surface reconstruction (or CityGML model reconstruction). On this basis, our Building3D can optionally generate CityGML models or surface mesh models of the buildings. Extensive experiments on ISPRS Vaihingen and DFC2019 datasets on the DSM estimation task show that our SFFDE significantly improves upon state-of-the-art and δ1, δ2and δ3metrics of our SFFDE are improved to 0.595, 0.897 and 0.970. Furthermore, our Building3D achieves impressive results in the 3D point cloud and 3D model reconstruction process.
Yongqiang Mao, Kaiqiang Chen, Liangjin Zhao, Deke Tang, Wenjie Liu 0016, Zhirui Wang 0003, Wenhui Diao, Xian Sun 0001, Kun Fu 0001
IEEE Trans. Geosci. Remote. Sens.1
2023 PICS: Paradigms Integration and Contrastive Selection for Semisupervised Remote Sensing Images Semantic Segmentation
abstract
Remote sensing images semantic segmentation is a fundamental yet challenging task, which has long relied heavily on sufficient pixelwise annotations. Semisupervised learning is proposed to address the problem of high dependence on labeled data by exploiting more learnable samples generated from the large amounts of accessible unlabeled data. However, affected by the complexity and diversity of remote sensing images, various misclassifications often occur and lead to errors accumulation during model training. Errors accumulation will destroy the consistency of model training and lead to degradation of final segmentation performance. In this article, in order to further alleviate the damage caused by the errors to the consistency of model training and improve final segmentation accuracy, we propose a novel semisupervised segmentation framework, paradigms integration and contrastive selection (PICS). First, multiple proven semisupervised paradigms are integrated to generate pseudolabeled samples with less noise. Second, a loss-based contrastive selection method is explored to distinguish generated samples that contain different degrees of inevitable misclassification, thereby further maintaining the approximation of the generated samples and the ground truth in the sample space. By generating and selecting high-quality pseudolabeled samples for selective self-training, we can better guarantee consistency during model training and obtain better segmentation results. Extensive experiments over the ISPRS Vaihingen, Potsdam, and the challenging iSAID benchmarks demonstrate that our method yields significant accuracy boosting on the segmentation results and achieves on-par performance with the state of the arts.
Xiyu Qi, Yongqiang Mao, Yidan Zhang 0002, Lei Wang 0077
IEEE Trans. Geosci. Remote. Sens.2
2023 Bridging the Gap Between Cumbersome and Light Detectors via Layer-Calibration and Task-Disentangle Distillation in Remote Sensing Imagery
abstract
With urgent application requirements, such as satellite in-orbit processing and unmanned aerial vehicle tracking, knowledge distillation (KD) following the teacher–student teaching mechanism has shown great potential to obtain lightweight detectors. However, compact students have limited accuracy due to the interference of large-scale variations and blurred boundaries in remote sensing objects. Specifically, previous methods mostly force teacher–student responses from the layer of the same depth and scale to align. Stereotyped manual interlayer associations may cause discriminative features of multiscale objects to be incorrectly bundled. Furthermore, the regression branch follows the identical distillation paradigm as the classification branch, resulting in ambiguous object bounding box deviations. To solve the above two issues, we propose an effective KD framework called layer-calibration and task-disentangle distillation (LTD). First, the cross-layer calibration distillation (CCD) structure is innovatively proposed. It adaptively binds a student layer with several related target layers, rather than a fixed layer in the teacher model. Appropriate and clear knowledge of large and small objects is transmitted. Since the CCD structure requires explicit global inner product computation between multiple layers, the local implicit calibration (LIC) module is further proposed to reduce distilled convergence difficulty. Second, the task-aware spatial disentangle distillation (TASD) structure is devised to transfer task-decoupled semantics and localization knowledge in a divide-and-conquer manner, alleviating objects’ localization imprecision. Experiments demonstrate that our LTD achieves state-of-the-art performance on several datasets and is a plug-and-play approach to most detectors. The code will be available soon.
Yidan Zhang 0002, Xian Sun 0001, Junxi Li, Yongqiang Mao, Lei Wang 0077
IEEE Trans. Geosci. Remote. Sens.6
2022 Bidirectional Feature Globalization for Few-shot Semantic Segmentation of 3D Point Cloud Scenes
abstract
Few-shot segmentation of point cloud remains a challenging task, as there is no effective way to convert local point cloud information to global representation, which hinders the generalization ability of point features. In this study, we propose a bidirectional feature globalization (BFG) approach, which leverages the similarity measurement between point features and prototype vectors to embed global perception to local point features in a bidirectional fashion. With point-to-prototype globalization (P02PrG), BFG aggregates local point features to prototypes according to similarity weights from dense point features to sparse prototypes. With prototype-to-point globalization (Pr2PoG), the global perception is embedded to local point features based on similarity weights from sparse prototypes to dense point features. The sparse prototypes of each class embedded with global perception are summarized to a single prototype for few-shot 3D segmentation based on the metric learning framework. Extensive experiments on S3DIS and ScanNet demonstrate that BFG significantly outperforms the state-of-the-art methods.
Yongqiang Mao, Zonghao Guo, Haowen Guo
3DV1
2022 Learning to Evaluate Performance of Multimodal Semantic Localization
abstract
Semantic localization (SeLo) refers to the task of obtaining the most relevant locations in large-scale remote sensing (RS) images using semantic information such as text. As an emerging task based on cross-modal retrieval, SeLo achieves semantic-level retrieval with only caption-level annotation, which demonstrates its great potential in unifying downstream tasks. Although SeLo has been carried out successively, but there is currently no work has systematically explores and analyzes this urgent direction. In this paper, we thoroughly study this field and provide a complete benchmark in terms of metrics and testdata to advance the SeLo task. Firstly, based on the characteristics of this task, we propose multiple discriminative evaluation metrics to quantify the performance of the SeLo task. The devised significant area proportion, attention shift distance, and discrete attention distance are utilized to evaluate the generated SeLo map from pixel-level and region-level. Next, to provide standard evaluation data for the SeLo task, we contribute a diverse, multi-semantic, multi-objective Semantic Localization Testset (AIR-SLT). AIR-SLT consists of 22 large-scale RS images and 59 test cases with different semantics, which aims to provide a comprehensive evaluations for retrieval models. Finally, we analyze the SeLo performance of RS cross-modal retrieval models in detail, explore the impact of different variables on this task, and provide a complete benchmark for the SeLo task. We have also established a new paradigm for RS referring expression comprehension, and demonstrated the great advantage of SeLo in semantics through combining it with tasks such as detection and road extraction. The proposed evaluation metrics, semantic localization testsets, and corresponding scripts have been open to access at https://github.com/xiaoyuan1996/SemanticLocalizationMetrics.
Wenkai Zhang 0002, Zhaoying Pan, Yongqiang Mao, Shuoke Li, Xian Sun 0001
IEEE Trans. Geosci. Remote. Sens.5