EDBT 2026 Demo / reviewers in the wild / expert
Peijin Wang
dblp:58/1713
· DBLP profile ↗
32ranked-venue papers
3as first author
28since 2021 · last 2026
0000-0002-0371-5584ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 21 · 3 first-author · 19 since 2021Artificial intelligence and machine learning · 9 · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CFVBench: A Comprehensive Video Benchmark for Fine-grained Multimodal Retrieval-Augmented GenerationabstractMultimodal Retrieval-Augmented Generation (MRAG) enables Multimodal Large Language Models (MLLMs) to generate responses with external multimodal evidence. Recently, numerous video-based MRAG benchmarks have been proposed to evaluate model capabilities across retrieval and generation stages in MRAG. However, existing benchmarks remain limited in modality coverage and format diversity, often focusing on single- or limited-modality tasks, or coarse-grained scene understanding. To address these gaps, we introduce CFVBench, a large-scale, manually verified benchmark constructed from 599 publicly available videos, yielding 5,360 open-ended QA pairs. CFVBench, spans high-density formats and domains such as chart-heavy reports, news broadcasts, and software tutorials, requiring models to retrieve and reason over long temporal video spans while maintaining fine-grained multimodal information. Using CFVBench, we systematically evaluate 7 retrieval methods and 14 widely-used MLLMs, revealing a critical bottleneck: current models (even GPT5 or Gemini) struggle to capture transient yet essential fine-grained multimodal details. To mitigate this, we propose Adaptive Visual Refinement (AVR), a plug-and-paly framework that adaptively increases frame sampling density and selectively invokes external tools when necessary. Experiments show that AVR consistently enhances fine-grained multimodal comprehension and improves performance across all evaluated MLLMs. Kaiwen Wei, Ruida Liu, Changzai Pan, Yidan Zhang 0002, Peijin Wang, Yingchao Feng |
WWW | 13 |
| 2026 | Open-Tag: A Generative Framework for Open-World Multimodal Tagging
Ziqi Zhang 0010, Zongyang Ma, Peijin Wang, Bing Li 0001, Chunfeng Yuan, Weiming Hu 0004 |
Int. J. Comput. Vis. | 3 |
| 2026 | RingMoE: Mixture-of-Modality-Experts Multi-Modal Foundation Models for Universal Remote Sensing Image InterpretationabstractThe rapid advancement of foundation models has revolutionized visual representation learning in a self-supervised manner. However, their application in remote sensing (RS) remains constrained by a fundamental gap: existing models predominantly handle single or limited modalities, overlooking the inherently multi-modal nature of RS observations. Optical, synthetic aperture radar (SAR), and multi-spectral data offer complementary insights that significantly reduce the inherent ambiguity and uncertainty in single-source analysis. To bridge this gap, we introduce RingMoE, a unified multi-modal RS foundation model with 14.7 billion parameters, pre-trained on 400 million multi-modal RS images from nine satellites. RingMoE incorporates three key innovations: 1) A hierarchical Mixture-of-Experts (MoE) architecture comprising modal-specialized, collaborative, and shared experts, effectively modeling intra-modal knowledge while capturing cross-modal dependencies to mitigate conflicts between modal representations; 2) Physics-informed self-supervised learning, explicitly embedding sensor-specific radiometric characteristics into the pre-training objectives; 3) Dynamic expert pruning, enabling adaptive model compression from 14.7B to 1B parameters while maintaining performance, facilitating efficient deployment in Earth observation applications. Evaluated across 23 benchmarks spanning six key RS tasks (i.e., classification, detection, segmentation, tracking, change detection, and depth estimation), RingMoE outperforms existing foundation models and sets new SOTAs, demonstrating remarkable adaptability from single-modal to multi-modal scenarios. Beyond theoretical progress, it has been deployed and trialed in multiple sectors, including emergency response, land management, marine sciences, and urban planning. Hanbo Bi, Yingchao Feng, Boyuan Tong, Haichen Yu, Yongqiang Mao, Wenhui Diao, Peijin Wang, Yue Yu 0001, Hanyang Peng, Yehong Zhang, Kun Fu 0001, Xian Sun 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 9 |
| 2026 | A Complex-Valued SAR Foundation Model Based on Physically Inspired Representation LearningabstractVision foundation models in remote sensing have been extensively studied due to their superior generalization on various downstream tasks. Synthetic Aperture Radar (SAR) offers all-day, all-weather imaging capabilities, providing significant advantages for Earth observation. However, establishing a foundation model for SAR image interpretation inevitably encounters the challenges of insufficient information utilization and poor interpretability. In this paper, we propose a remote sensing foundation model based on complex-valued SAR data, which simulates the polarimetric decomposition process for pre-training, i.e., characterizing pixel scattering intensity as a weighted combination of scattering bases and scattering coefficients, thereby endowing the foundation model with physical interpretability. Specifically, we construct a series of scattering queries, each representing an independent and meaningful scattering basis, which interact with SAR features in the scattering query decoder and output the corresponding scattering coefficient. To guide the pre-training process, polarimetric decomposition loss and power self-supervised loss are constructed. The former aligns the predicted coefficients with Yamaguchi coefficients, while the latter reconstructs power from the predicted coefficients and compares it to the input image's power. The performance of our foundation model is validated on nine typical downstream tasks, achieving state-of-the-art results. Notably, the foundation model can extract stable feature representations and exhibits strong generalization, even in data-scarce conditions. Hanbo Bi, Yingchao Feng, Linlin Xin, Shuo Gong, Peijin Wang, Wenhui Diao, Xian Sun 0001 |
IEEE Trans. Image Process. | 8 |
| 2025 | RS-vHeat: Heat Conduction Guided Efficient Remote Sensing Foundation ModelabstractRemote sensing foundation models largely break away from the traditional paradigm of designing task-specific models, offering greater scalability across multiple tasks. However, they face challenges such as low computational efficiency and limited interpretability, especially when dealing with large-scale remote sensing images. To overcome these, we draw inspiration from heat conduction, a physical process modeling local heat diffusion. Building on this idea, we are the first to explore the potential of using the parallel computing model of heat conduction to simulate the local region correlations in high-resolution remote sensing images, and introduce RS-vHeat, an efficient multi-modal remote sensing foundation model. Specifically, RS-vHeat 1) applies the Heat Conduction Operator (HCO) with a complexity of $O(N^{1.5})$ and a global receptive field, reducing computational overhead while capturing remote sensing object structure information to guide heat diffusion; 2) learns the frequency distribution representations of various scenes through a self-supervised strategy based on frequency domain hierarchical masking and multi-domain reconstruction; 3) significantly improves efficiency and performance over state-of-the-art techniques across 4 tasks and 10 datasets. Compared to attention-based remote sensing foundation models, we reduce memory usage by 84\%, FLOPs by 24\% and improves throughput by 2.7 times. The code will be made publicly available. Huiyang Hu, Peijin Wang, Hanbo Bi, Boyuan Tong, Zhaozhi Wang, Wenhui Diao, Yingchao Feng, Ziqi Zhang 0010, Yaowei Wang 0001, Qixiang Ye, Kun Fu 0001, Xian Sun 0001 |
ICCV | 2 |
| 2025 | Prompt-and-Transfer: Dynamic Class-Aware Enhancement for Few-Shot SegmentationabstractFor more efficient generalization to unseen domains (classes), most Few-shot Segmentation (FSS) would directly exploit pre-trained encoders and only fine-tune the decoder, especially in the current era of large models. However, such fixed feature encoders tend to be class-agnostic, inevitably activating objects that are irrelevant to the target class. In contrast, humans can effortlessly focus on specific objects in the line of sight. This paper mimics the visual perception pattern of human beings and proposes a novel and powerful prompt-driven scheme, called "Prompt and Transfer" (PAT), which constructs a dynamic class-aware prompting paradigm to tune the encoder for focusing on the interested object (target class) in the current task. Three key points are elaborated to enhance the prompting: 1) Cross-modal linguistic information is introduced to initialize prompts for each task. 2) Semantic Prompt Transfer (SPT) that precisely transfers the class-specific semantics within the images to prompts. 3) Part Mask Generator (PMG) that works in conjunction with SPT to adaptively generate different but complementary part prompts for different individuals. Surprisingly, PAT achieves competitive performance on 4 different tasks including standard FSS, Cross-domain FSS (e.g., CV, medical, and remote sensing domains), Weak-label FSS, and Zero-shot Segmentation, setting new state-of-the-arts on 11 benchmarks. Hanbo Bi, Yingchao Feng, Wenhui Diao, Peijin Wang, Yongqiang Mao, Kun Fu 0001, Xian Sun 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | RingMoGPT: A Unified Remote Sensing Foundation Model for Vision, Language, and Grounded TasksabstractRecently, multimodal large language models (MLLMs) have shown excellent reasoning capabilities in various fields. Most of the existing remote sensing (RS) MLLMs solve image-level text generation problems (e.g., image captioning), but ignore the core issues of object-level recognition, location, and multitemporal changes in the field of RS. In this article, we propose RingMoGPT, a multimodal foundation model that unifies vision, language, and localization. Based on the idea of domain adaption, RingMoGPT can complete training by fine-tuning only a few parameters. To make the model capable of object detection and change captioning, we further propose a location- and instruction-aware querying transformer (Q-Former) and a change detection module, respectively. To improve the performance of RingMoGPT, we carefully design the pretraining dataset and the instruction-tuning dataset. The pretraining dataset contains over a half million high-quality image and text pairs, which are generated through a low-cost and efficient data generation paradigm. The instruction-tuning dataset contains more than 1.6 million question-answer pairs, including six downstream tasks: scene classification, object detection, visual question answering (VQA), image captioning, grounded image captioning, and change captioning. Our experiments show that RingMoGPT performs well on six tasks, especially its ability to analyze multitemporal data changes and identify dense objects. We also verified the model under a zero-shot setting, and the results show that the proposed RingMoGPT also has good generalization ability in the face of new data. Peijin Wang, Huiyang Hu, Boyuan Tong, Ziqi Zhang 0010, Fanglong Yao, Yingchao Feng, Zining Zhu 0004, Wenhui Diao, Qixiang Ye, Xian Sun 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2025 | Balancing Attention to Base and Novel Categories for Few-Shot Object Detection in Remote Sensing ImageryabstractFew-shot object detection (FSOD) has garnered widespread attention in recent years, which makes it possible to learn novel classes with only a handful of labeled samples. Due to the obvious long-tail distribution of remote sensing (RS) data and serious challenges in data labeling, FSOD holds greater practical application value in RS. At present, the FSOD algorithms with fine-tuning of some parameters have attracted much attention due to their stronger incremental learning capacity. However, given large intraclass scale variations and small interclass feature differences of RS objects, we still need to place greater emphasis on the localization and classification of objects. In this article, we introduce the RoI feature refinement (RIFR) method for FSOD in RS imagery, which adopts a powerful training pipeline to better balance attention to the performance of base and novel classes. Aiming at large intraclass scale variations of RS objects, we design a scale-aware feature compensation module (SAFCM). By compensating for insufficient scale information, the model’s ability to perceive scale variations of the same class objects has been enhanced. Considering the high similarity among RS classes, we come up with a prototype trihard (PT) loss. It achieves the effect of interclass separability by constraining the relationship between samples and prototypes. Thus, the issue of interclass confusion has been resultfully solved. Comprehensive experiments on three datasets, DIOR, NWPU VHR-10.v2, and FAIR1M-Airplane, can showcase the efficacy of our RIFR method, and it can implement the most outstanding performance currently. The code will be available at:https://github.com/ningerhhh/RIFR. Zining Zhu 0004, Peijin Wang, Wenhui Diao, Jinze Yang, Lingyu Kong, Xian Sun 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | FAIR1M-GQA: Fine-Grained Grounded Question Answering Dataset in Remote SensingabstractWith the development of large language models (LLMs) and remote sensing technology, visual language (VL) tasks in the field of remote sensing have attracted more and more research attention. Commonly used VL datasets currently usually focus on the overall scene of the image, lacking the description of instance-level details such as location, size, category and so on. More importantly, users are usually not allowed to directly intercept regions in the image to ask questions through these datasets. However, the instance-level question answering based on these information is of great significance for target extraction in practical applications. In this manuscript, we build an innovative and challenging dataset FAIR1M-GQA. It unlocks the ability of the model to learn directly from text input and text output both with region coordinates, which are directly linked to fine-grained objects in remote sensing images. We experiment our dataset to verify the feasibility of the relevant task and provide the benchmark results. Huiyang Hu, Peijin Wang, Yingchao Feng, Wenhui Diao, Ziqi Zhang 0010, Xian Sun 0001, Kun Fu 0001 |
IGARSS | 2 |
| 2024 | Chinese Title Generation for Short Videos: Dataset, Metric and AlgorithmabstractPrevious work for video captioning aims to objectively describe the video content but the captions lack human interest and attractiveness, limiting its practical application scenarios. The intention of video title generation (video titling) is to produce attractive titles, but there is a lack of benchmarks. This work offers CREATE, the first large-scale Chinese shoRt vidEo retrievAl and Title gEneration dataset, to assist research and applications in video titling, video captioning, and video retrieval in Chinese. CREATE comprises a high-quality labeled 210 K dataset and two web-scale 3 M and 10 M pre-training datasets, covering 51 categories, 50K+ tags, 537K+ manually annotated titles and captions, and 10M+ short videos with original video information. This work presents ACTEr, a unique Attractiveness-Consensus-based Title Evaluation, to objectively evaluate the quality of video title generation. This metric measures the semantic correlation between the candidate (model-generated title) and references (manual-labeled titles) and introduces attractive consensus weights to assess the attractiveness and relevance of the video title. Accordingly, this work proposes a novel multi-modal ALignment WIth Generation model, ALWIG, as one strong baseline to aid future model development. With the help of a tag-driven video-text alignment module and a GPT-based generation module, this model achieves video titling, captioning, and retrieval simultaneously. We believe that the release of the CREATE dataset, ACTEr metric, and ALWIG model will encourage in-depth research on the analysis and creation of Chinese short videos. Ziqi Zhang 0010, Zongyang Ma, Chunfeng Yuan, Peijin Wang, Zhongang Qi, Chenglei Hao, Bing Li 0001, Ying Shan, Weiming Hu 0004, Stephen J. Maybank |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | A Triple-Branch Hybrid Attention Network With Bitemporal Feature Joint Refinement for Remote-Sensing Image Semantic Change DetectionabstractCompared with binary change detection (BCD), semantic change detection (SCD) further provides the category information of bitemporal changed regions which is significant for the practical application of Earth Observation. Although the recently proposed triple-branch structures including one BCD branch and two classification branches can effectively achieve the task balance, they still need to employ the carefully designed difference extraction module and branch interactions to capture the bitemporal correlations, which increases the complexity of the semantic information utilization. In this paper, we propose a new triple-branch network named JFRNet to tackle this challenge. From the perspective of the SCD process, because the category information and the change information are both derived from bitemporal images, we take the joint bitemporal features as the unified input, which can help each branch perceive the bitemporal semantic correlations without any additional interaction operations. From the perspective of the SCD structure, we introduce the convolutional attention fusion module (CAFM) and the convolutional attention refinement module (CARM) to unify the branch structure, which can help our model refine the unique semantic information without any specially designed difference extraction modules. Extensive experiment results on three available datasets indicate that compared with the baseline methods, our proposed JFRNet successfully simplifies the reasoning process and obtains the better SCD performance. Peijin Wang, Wenhui Diao, Guangluan Xu, Xian Sun 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Few-Shot Incremental Object Detection in Aerial Imagery via Dual-Frequency PromptabstractRecently, there has been a growing interest in few-shot incremental object detection (FSIOD). It learns new tasks with limited data while mitigating catastrophic forgetting on previous tasks. However, existing FSIOD methods experience parameter changes after training on new tasks, causing a parameter competition issue among tasks. Additionally, the background information differs among various tasks, and using a common background weight for all tasks results in the background shift. Constrained by these two issues, existing methods only alleviate catastrophic forgetting and cannot wholly prevent the performance decline on previous tasks. Especially for complex remote sensing images with messy background, the models trained on new tasks exhibit noticeable performance drops on previous tasks. In this paper, we propose a novel FSIOD method via dual-frequency prompt to address these challenges, named FSIOD-DFP. It can completely eliminate catastrophic forgetting while mitigating over-fitting. Specifically, a dual-frequency prompt generator is designed to tackle the parameter competition issue. It decouples the frequency components of images to produce prompts that modify the images to adapt to the base model trained on previous tasks. Compared to traditional prompts, our generator introduces fewer parameters to address over-fitting for limited data and allows freezing the base model to maintain the performance of previous data. Besides, a self-regularization loss is introduced to guide the prompt-modified images to leverage the knowledge of the base model effectively. Furthermore, we propose a task-decoupled detection head to address the background shift problem. It separates the detection heads for new and previous tasks to resolve the conflict in the background between different tasks. In FSIOD-DFP, only a prompt generator and a novel detection head are added and fine-tuned when learning a new task. Experiments on three remote sensing object detection datasets demonstrate that our method achieves state-of-the-art performance on both new and previous tasks in all few-shot incremental settings. Wenhui Diao, Junxi Li, Yidan Zhang 0002, Peijin Wang, Xian Sun 0001, Kun Fu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Remote Sensing Change Detection With Bitemporal and Differential Feature Interactive PerceptionabstractRecently, the transformer has achieved notable success in remote sensing (RS) change detection (CD). Its outstanding long-distance modeling ability can effectively recognize the change of interest (CoI). However, in order to obtain the precise pixel-level change regions, many methods directly integrate the stacked transformer blocks into the UNet-style structure, which causes the high computation costs. Besides, the existing methods generally consider bitemporal or differential features separately, which makes the utilization of ground semantic information still insufficient. In this paper, we propose the multiscale dual-space interactive perception network (MDIPNet) to fill these two gaps. On the one hand, we simplify the stacked multi-head transformer blocks into the single-layer single-head attention module and further introduce the lightweight parallel fusion module (LPFM) to perform the efficient information integration. On the other hand, based on the simplified attention mechanism, we propose the cross-space perception module (CSPM) to connect the bitemporal and differential feature spaces, which can help our model suppress the pseudo changes and mine the more abundant semantic consistency of CoI. Extensive experiment results on three challenging datasets and one urban expansion scene indicate that compared with the mainstream CD methods, our MDIPNet obtains the state-of-the-art (SOTA) performance while further controlling the computation costs. Peijin Wang, Wenhui Diao, Guangluan Xu, Xian Sun 0001 |
IEEE Trans. Image Process. | 2 |
| 2023 | Breaking Immutable: Information-Coupled Prototype Elaboration for Few-Shot Object DetectionabstractFew-shot object detection, expecting detectors to detect novel classes with a few instances, has made conspicuous progress. However, the prototypes extracted by existing meta-learning based methods still suffer from insufficient representative information and lack awareness of query images, which cannot be adaptively tailored to different query images. Firstly, only the support images are involved for extracting prototypes, resulting in scarce perceptual information of query images. Secondly, all pixels of all support images are treated equally when aggregating features into prototype vectors, thus the salient objects are overwhelmed by the cluttered background. In this paper, we propose an Information-Coupled Prototype Elaboration (ICPE) method to generate specific and representative prototypes for each query image. Concretely, a conditional information coupling module is introduced to couple information from the query branch to the support branch, strengthening the query-perceptual information in support features. Besides, we design a prototype dynamic aggregation module that dynamically adjusts intra-image and inter-image aggregation weights to highlight the salient information useful for detecting query images. Experimental results on both Pascal VOC and MS COCO demonstrate that our method achieves state-of-the-art performance in almost all settings. Code will be available at: https://github.com/lxn96/ICPE. Wenhui Diao, Yongqiang Mao, Junxi Li, Peijin Wang, Xian Sun 0001, Kun Fu 0001 |
AAAI | 5 |
| 2023 | A Self-Supervised Cross-Modal Remote Sensing Foundation Model with Multi-Domain Representation and Cross-Domain FusionabstractThe construction of a basic model to extract generalized features from a large number of multimodal data is a new challenge in the field of remote sensing. Compared with natural scene images, When faced with a complex application scenario of remote sensing of multi-sensor acquisition, models that are suitable for a specific task are difficult to generalize to new scenarios. In this paper, we propose a model architecture based on the concepts of multi-domain representation and cross-domain fusion. By extracting strong generalization features from massive multi-modal data, a single foundation model can accomplish generalization interpretation for multiple downstream tasks. Experimental results show that the proposed model performs well on multiple downstream tasks, which validates the feasibility of the remote sensing cross-modal foundation model in the interpretation task. Yingchao Feng, Peijin Wang, Wenhui Diao, Qibin He 0001, Huiyang Hu, Hanbo Bi, Xian Sun 0001, Kun Fu 0001 |
IGARSS | 2 |
| 2023 | From single- to multi-modal remote sensing imagery interpretation: a survey and taxonomy
Xian Sun 0001, Wanxuan Lu, Peijin Wang, Ruigang Niu, Kun Fu 0001 |
Sci. China Inf. Sci. | 4 |
| 2023 | AIR-PV: a benchmark dataset for photovoltaic panel extraction in optical remote sensing imagery
Peijin Wang, Feng Xu 0001, Xian Sun 0001, Wenhui Diao |
Sci. China Inf. Sci. | 2 |
| 2023 | Few-Shot Object Detection in Aerial Imagery Guided by Text-Modal KnowledgeabstractFew-shot object detection (FSOD) has received numerous attention due to the difficulty and time-consuming of labeling objects. Recent researches achieve excellent performance in a natural scene by only using a few instances of novel classes to fine-tune the last prediction layer of the model well-trained on plentiful base data. However, compared with natural scene objects with a single direction and small size variety, the direction and size of the objects in remote sensing images (RSIs) vary greatly. The methods proposed for the natural scene cannot be directly applied to RSIs. In this article, we first propose a strong baseline for RSIs. It fine-tunes all detector components acting on high-level features and effectively improves the performance of novel classes. Further analyzing the results of the baseline, we find that the error for novel classes is mainly concentrated in classification. It misclassifies novel classes as confusable base classes or backgrounds due to the difficulty in extracting generalized information from limited instances. As is well-known, text-modal knowledge can highly summarize the generalized and unique characteristics of categories. Thus, we introduce text-modal descriptions for each category and propose an FSOD method guided by TExt-MOdal knowledge, called TEMO. Specifically, a text-modal knowledge extractor and a cross-modal assembly module are proposed to extract text features and fuse the text-modal features into visual-modal features. The fused features greatly reduce the classification confusion of novel classes. Furthermore, we introduce a mask strategy and a separation loss to avoid over-fitting and ambiguity of text-modal features. Experimental results on detection in optical remote sensing images (DIOR), Northwestern Polytechnical University (NWPU), and fine-grained object recognition in high-resolution remote sensing imagery (FAIR1M) illustrate that our TEMO achieves state-of-the-art performance in all settings. Xian Sun 0001, Wenhui Diao, Yongqiang Mao, Junxi Li, Yidan Zhang 0002, Peijin Wang, Kun Fu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2023 | MiCro: Modeling Cross-Image Semantic Relationship Dependencies for Class-Incremental Semantic Segmentation in Remote Sensing ImagesabstractContinual learning is an effective way to overcome catastrophic forgetting (CF) in incremental learning for semantic segmentation. The existing continual semantic segmentation (CSS) methods of remote sensing (RS) ignore the semantic relationships among pixels across different images, which will lead to disappointing segmentation results, such as edge pixel misclassification and small object omission. In this paper, we propose a framework for modeling cross-image semantic relationship dependencies (MiCro), which aims to learn an inter-class separable and intra-class cohesive feature space from the pixel relationships across various images to ensure that learned categories can prevent CF in the incremental process. Specifically, we exploit the relationships among pixels of images in mini-batch to construct three losses: (a) Cross-image feature relationship distillation (CFRD) loss, which builds a well-structured feature space; (b) Cross-image intra-class feature cohesion (CIFC) loss, which is devised to make intra-class features more cohesive; and (c) Cross-image class-area weighted cross-entropy (CCWCE) loss, which is mainly employed to inversely weight the proportion of category area in mini-batch. The effectiveness of the proposed approach is demonstrated by extensive experiments on three RS semantic segmentation datasets from ISPRS Vaihingen, ISPRS Potsdam, and iSAID. MiCro is superior to the current most advanced methods in most incremental settings, especially improving mIoU by 11.59% on ISPRS Vaihingen, 13.17% on ISPRS Potsdam, and 15.01% on iSAID in the most difficult incremental settings, which promotes the CSS to a state-of-the-art (SOTA) level. The code will be available at https://github.com/RongXueE/MiCro. Xuee Rong, Peijin Wang, Wenhui Diao, Wenxin Yin, Xuan Zeng 0004, Xian Sun 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | RingMo: A Remote Sensing Foundation Model With Masked Image ModelingabstractDeep learning approaches have contributed to the rapid development of remote sensing (RS) image interpretation. The most widely used training paradigm is to use ImageNet pretrained models to process RS data for specified tasks. However, there are issues such as domain gap between natural and RS scenes and the poor generalization capacity of RS models. It makes sense to develop a foundation model with general RS feature representation. Since a large amount of unlabeled data is available, the self-supervised method has more development significance than the fully supervised method in RS. However, most of the current self-supervised methods use contrastive learning, whose performance is sensitive to data augmentation, additional information, and selection of positive and negative pairs. In this article, we leverage the benefits of generative self-supervised learning (SSL) for RS images and propose an RS foundationmodel framework called RingMo, which consists of two parts. First, a large-scale dataset is constructed by collecting two million RS images from satellite and aerial platforms, covering multiple scenes and objects around the world. Second, we propose an RS foundation model training method designed for dense and small objects in complicated RS scenes. We show that the foundation model trained on our dataset with RingMo method achieves state-of-the-art (SOTA) on eight datasets across four downstream tasks, demonstrating the effectiveness of the proposed framework. Through in-depth exploration, we believe it is time for RS researchers to embrace generative SSL and leverage its general representation capabilities to speed up the development of RS applications. Xian Sun 0001, Peijin Wang, Wanxuan Lu, Zicong Zhu, Qibin He 0001, Junxi Li, Xuee Rong, Zhujun Yang, Qinglin He, Ruiping Wang 0001, Jiwen Lu, Kun Fu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | MoCG: Modality Characteristics-Guided Semantic Segmentation in Multimodal Remote Sensing ImagesabstractThe rapid development of satellite platforms has yielded copious and diverse multi-source data for earth observation, greatly facilitating the growth of multimodal semantic segmentation (MSS) in remote sensing. However, MSS also suffers from numerous challenges: 1) Existing inherent defects in each modality due to the different imaging mechanisms. 2) Insufficient exploration of the intrinsic characteristics of modalities. 3) The existence of the huge semantic gap between heterogeneous data causes difficulties in feature fusion. The inability to effectively utilize the rich and diverse information provided by each modality and ignorance of the heterogeneity between modalities will hinder the feature enhancement, and further significantly impacts the semantic segmentation accuracy. Furthermore, neglecting the huge gap makes feature fusion challenging. In this study, we introduce a novel framework for multimodal semantic segmentation that effectively mitigates the aforementioned problems. Our approach employs a pseudo-siamese structure for feature extraction. Specifically, we propose a simple yet effective geometric topology structure modeling (GTSM) module to extract geometric relationships and texture information from optical data. Additionally, we present a modality intrinsic noise suppression (MINS) module to fully exploit radiation information and alleviate the effects of unique geometric distortions for SAR. Furthermore, we present an adaptive multimodal feature fusion (AMFF) module for fully fusing different modality features. Extensive experiments on both WHU-OPT-SAR and DFC23 datasets validate the robustness and effectiveness of the proposed Modality Characteristics-Guided Semantic Segmentation (MoCG) network compared to other state-of-the-art semantic segmentation methods, including multimodal and single-modal approaches. Our approach achieves the best performance on both datasets, resulting in mIoU/OA gains 69.1%/87.5% on WHU-OPT-SAR and 86.7%/97.3% on DFC23. Sining Xiao, Peijin Wang, Wenhui Diao, Xuee Rong, Xuexue Li, Kun Fu 0001, Xian Sun 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Hypertron: Explicit Social-Temporal Hypergraph Framework for Multi-Agent ForecastingabstractForecasting the future trajectories of multiple agents is a core technology for human-robot interaction systems. To predict multi-agent trajectories more accurately, it is inevitable that models need to improve interpretability and reduce redundancy. However, many methods adopt implicit weight calculation or black-box networks to learn the semantic interaction of agents, which obviously lack enough interpretation. In addition, most of the existing works model the relation among all agents in a one-to-one manner, which might lead to irrational trajectory predictions due to its redundancy and noise. To address the above issues, we present Hypertron, a human-understandable and lightweight hypergraph-based multi-agent forecasting framework, to explicitly estimate the motions of multiple agents and generate reasonable trajectories. The framework explicitly interacts among multiple agents and learns their latent intentions by our coarse-to-fine hypergraph convolution interaction module. Our experiments on several challenging real-world trajectory forecasting datasets show that Hypertron outperforms a wide array of state-of-the-art methods while saving over 60% parameters and reducing 30% inference time. Xingliang Huang, Ruigang Niu, Peijin Wang, Xian Sun 0001 |
IJCAI | 5 |
| 2022 | SIL-LAND: Segmentation Incremental Learning in Aerial Imagery via LAbel Number Distribution ConsistencyabstractSegmentation incremental learning has received a lot of attention in recent years due to the ability to overcome the problem of catastrophic forgetting. Our study found that differences in label number distribution affect the performance of segmentation incremental learning. Because the labels for pixels of the old category are marked as background when the model is trained on the new tasks, the label number distribution is inconsistent with static learning that is considered to be the upper bound on incremental learning, which hinders the mitigation of the catastrophic forgetting problem. In response to the above problems, we propose an incremental learning method named SIL-LAND, which improves the accuracy by making the label number distribution of our method close to that of static learning. From the perspective of high-level semantic labels, we propose the prototype update mechanism for the problem that non-adaptive representative prototypes ignore the sample diversity of semantic categories in remote sensing images. By compensating for the difference in label number distribution at the feature level, the distance between the prototype and the actual class center is reduced; Aiming at the lack of semantic consistency between feature vectors and prototypes, we propose a similarity measure module to increase the intra-class similarity between the prototype and corresponding feature vectors. From the perspective of one-hot labels, we propose label reconstruction, including foreground screening and background padding to make the number distribution of one-hot labels as close as possible to that of static learning. A series of experimental results demonstrate the effectiveness of our method. Junxi Li, Wenhui Diao, Peijin Wang, Yidan Zhang 0002, Zhujun Yang, Guangluan Xu, Xian Sun 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Class-Incremental Learning Network for Small Objects Enhancing of Semantic Segmentation in Aerial ImageryabstractDue to the differences in the feature distribution between classes, when the model learns in a continuous data stream, it will encounter catastrophic forgetting. The incremental learning methods have shown great potential to solve this problem. However, most existing methods based on task-incremental learning are difficult to adapt to characteristics of remote sensing scenes with few differences in appearance but large differences in features, which is not conducive to artificially distinguish task-identity document (ID). Thus, we propose a class-incremental learning (CIL) network for small objects enhancing semantic segmentation in aerial imagery. Specifically, considering the superior accuracy of the binary classifier, we propose a twin-auxiliary (TA) model that adds an auxiliary binary classification task. Then, for expansion and contraction at the edge and small object confusion problems, we introduce a diversity distillation loss, using the results of binary-classifier to constrain the multiclass segmentation results and strengthen the attention to the locations of the segmentation results that have changed. Finally, we design a conflict reduction mechanism for multihead classifier to achieve single-head prediction for CIL. Experiments demonstrate that our method has good performance on the Vaihingen and Potsdam datasets by the International Society for Photogrammetry and Remote Sensing (ISPRS), outperforming state-of-the-art (SOTA) incremental learning methods. The code will be available soon. Junxi Li, Xian Sun 0001, Wenhui Diao, Peijin Wang, Yingchao Feng, Guangluan Xu |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | LIL: Lightweight Incremental Learning Approach Through Feature Transfer for Remote Sensing Image Scene ClassificationabstractExisting deep learning models usually assume that all data obeys independent identically distribution, which is unreasonable in remote sensing. Due to the differences in camera parameters, spectral ranges, resolutions, and so on, the images acquired by remote sensing sensors may be greatly diverse, causing models to face catastrophic forgetting when they are trained on new data only. Thus, incremental learning is introduced. An ideal incremental learning model should be expanded as the number of tasks increases, so as to have enough ability to adapt to the changes in data. However, existing approaches normally expand heavy modules for each task, making the holistic models cumbersome. In this article, a lightweight incremental learning approach (LIL) is proposed for remote sensing image scene classification. We replace the role of the feature extractor with extracting features of a single task instead of task-sharing features of all tasks to lighten the backbone. In addition, we propose a light feature transfer module (FTM) to realize the alignment of data distributions between different tasks in the feature domain. Furthermore, dual-constraint loss with knowledge distillation and adversarial learning is introduced to promote the mapping and alignment of data distributions at both the feature level and the semantic level. In LIL, only a tiny FTM and a classifier are added to the model when the model learns a new task. Experimental results show that our approach with a small number of parameters outperforms state-of-the-art approaches for incremental learning on both a single dataset and a sequence of multiple datasets. Xian Sun 0001, Wenhui Diao, Yingchao Feng, Peijin Wang, Kun Fu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | Historical Information-Guided Class-Incremental Semantic Segmentation in Remote Sensing ImagesabstractDespite the extraordinary success of the deep architectures on semantic segmentation for remote sensing (RS) images, they have difficulties in learning new classes from a sequential data stream because of catastrophic forgetting. Continual learning for semantic segmentation (CSS) is an emerging trend for its capability to cope with the above problems effectively. However, old classes from previous steps are collapsed into the background, which further aggravates the challenge of CSS in the RS scene. In this article, we revisit the knowledge distillation (KD) strategy and the characteristics of class-incremental semantic segmentation (CISS) and then present a generalized and effective framework to learn new classes while preserving knowledge of the learned classes. In particular, we propose two novel historical information-guided modules: the feature global perception module and the label reconstruction (LR) module. The former enables the current model to pay more attention to the region related to the old categories identified by the historical information when learning new classes. Meanwhile, the latter retrieves pixels belonging to the learned classes from the background to handle the background shift problem and maintain the high performance of old classes. We have conducted comprehensive experiments on two RS semantic segmentation datasets of Instance Segmentation in Aerial Images Dataset (iSAID) and Gao Fen (GF) challenge semantic segmentation dataset (GCSS). The experimental results outperform the current state-of-the-art methods in most incremental settings, which demonstrates the effectiveness of the proposed framework. Xuee Rong, Xian Sun 0001, Wenhui Diao, Peijin Wang |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Range Sidelobe Suppression Approach for SAR Images Using Chaotic FM SignalsabstractRange sidelobe is very common in synthetic aperture radar (SAR) images, particularly when imaging scene includes strongly scattering targets such as ships or complex buildings. As a kind of interference, it may reduce the image quality and hinder the image interpretation. Hence, range sidelobe suppression is an important mission for SAR images. The main task of mitigating the sidelobe is how to achieve the most effective suppression with the minimal resolution loss and signal-to-noise ratio (SNR) loss. However, the widely recognized classic method, spatially variant apodization (SVA), still has a lot of residual sidelobe energy and other problems. This article proposes a novel suppression approach based on time-variant transmission of chaotic frequency modulation (CFM) signals. The key is to build an appropriate transmitted signal set, where the signals are generated by various chaotic initial states and the same special map with low mixing rate and uniform invariant probability density (IPD). Due to their beneficial autocorrelation properties, the proposed approach achieves superior performance in range sidelobe suppression and resolution preservation. More importantly, it maintains the energy of the signals and overcomes the SNR loss that occurs in some classic methods, such as spectral weighting (SW) and SVA. In addition, it is suitable for both vertical and squint side-looking mode and can well reconstruct the weakly scattering targets which are severely disturbed by range sidelobe. All of them are validated by comparative experiments. Youming Wu, Kun Fu 0001, Wenhui Diao, Peijin Wang, Xian Sun 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2021 | SRAF-Net: Shape Robust Anchor-Free Network for Garbage Dumps in Remote Sensing ImageryabstractThe detection of garbage dumps is of great significance for environmental protection. Recently, deep learning algorithms have brought impressive improvements for regular object detection. Different from conventional objects, garbage dumps are more inconspicuous and irregular and have the problem of blurred boundaries. To solve these problems, we propose a shape robust anchor-free network (SRAF-Net) that consists of feature extraction, multitask detection, and postprocessing. First, our network leverages the context-based deformable (CBD) module to combine context attention and deformable convolution. The contextual information obtained by context attention enables the network to focus on objects with inconspicuous appearance, while the deformable convolution enhances the feature representation. Then, we propose a multitask detection head to regress irregular garbage dumps in a more accurate and efficient way. The anchor-based methods need to define some anchors with a fixed shape. However, our detection method is anchor-free that learns the shapes of objects from training data. The detection head adaptively generates various shapes of bounding boxes with their classification confidences and localization confidences. Weighted by the localization confidences, we merge bounding boxes during postprocessing, which alleviates the blurred boundaries. In addition, we build a new public data set named garbage dumps data set (GDD) to verify the effectiveness of our method. Extensive experiments on GDD indicate that our method surpasses the existing detection methods in terms of speed and accuracy for the garbage dumps detection task. Xian Sun 0001, Yingfei Liu, Peijin Wang, Wenhui Diao, Kun Fu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2020 | Object Relational Graph With Teacher-Recommended Learning for Video CaptioningabstractTaking full advantage of the information from both vision and language is critical for the video captioning task. Existing models lack adequate visual representation due to the neglect of interaction between object, and sufficient training for content-related words due to long-tailed problems. In this paper, we propose a complete video captioning system including both a novel model and an effective training strategy. Specifically, we propose an object relational graph (ORG) based encoder, which captures more detailed interaction features to enrich visual representation. Meanwhile, we design a teacher-recommended learning (TRL) method to make full use of the successful external language model (ELM) to integrate the abundant linguistic knowledge into the caption model. The ELM generates more semantically similar word proposals which extend the groundtruth words used for training to deal with the long-tailed problem. Experimental evaluations on three benchmarks: MSVD, MSR-VTT and VATEX show the proposed ORG-TRL system achieves state-of-the-art performance. Extensive ablation studies and visualizations illustrate the effectiveness of our system. Ziqi Zhang 0010, Yaya Shi, Chunfeng Yuan, Bing Li 0001, Peijin Wang, Weiming Hu 0004, Zhengjun Zha |
CVPR | 5 |
| 2020 | FMSSD: Feature-Merged Single-Shot Detection for Multiscale Objects in Large-Scale Remote Sensing ImageryabstractRecently, the deep convolutional neural network has brought great improvements in object detection. However, the balance between high accuracy and high speed has always been a challenging task in multiclass object detection for large-scale remote sensing imagery. One-stage methods are more widely used because of their high efficiency but are limited by their performances on small object detection. In this article, we propose a unified framework called feature-merged single-shot detection (FMSSD) network, which aggregates the context information both in multiple scales and the same scale feature maps. First, our network leverages the atrous spatial feature pyramid (ASFP) module to fuse the context information in multiscale features by using feature pyramid and multiple atrous rates. Second, we propose a novel area-weighted loss function to pay more attention to small objects, while the replaced original loss treats all objects equally. We believe that small objects should be given more weight than large objects because they lose more information during training. Specifically, a monotonic decreasing function about the area is designed to add weights on the loss function. Extensive experiments on the DOTA data set and NWPU VHR-10 data set demonstrate that our method achieves state-of-the-art detection accuracy with high efficiency. We also build a new large-scale data set called AIR-OBJ data set from Google Earth and show the detection results of small objects, which validates the effectiveness on large-scale remote sensing imagery. Peijin Wang, Xian Sun 0001, Wenhui Diao, Kun Fu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2019 | Mergenet: Feature-Merged Network for Multi-Scale Object Detection in Remote Sensing ImagesabstractObject detection has been playing a significant role in the field of remote sensing for a long period while it is still full of challenges. The biggest one is how to detect multi-scale objects with high accuracy and fast speed in remote sensing images. One-stage object detectors have been achieving relatively high accuracy and efficiency with small memory footprint. However, they have a not very well performance on small objects. In this paper, we discuss the importance of the context information between feature maps in different scales which is helpful for detecting small objects. Especially, we propose a Feature-merged detection networks (MergeNet), which can be inserted into the one-stage detectors easily, to unify the multi-scale feature and context information effectively. Experiments on DOTA dataset demonstrate that our model can significantly improve the performance of the one-stage method. Peijin Wang, Xian Sun 0001, Wenhui Diao, Kun Fu 0001 |
IGARSS | 1 |
| 2004 | Complex Model Identification Based on RBF Neural Network
Yibin Song, Peijin Wang, Kaili Li |
ISNN (2) | 2 |