EDBT 2026 Demo / reviewers in the wild / expert
Yingchao Feng
dblp:239/6011
· DBLP profile ↗
24ranked-venue papers
5as first author
22since 2021 · last 2026
0000-0003-4017-8885ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 17 · 4 first-author · 15 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CFVBench: A Comprehensive Video Benchmark for Fine-grained Multimodal Retrieval-Augmented GenerationabstractMultimodal Retrieval-Augmented Generation (MRAG) enables Multimodal Large Language Models (MLLMs) to generate responses with external multimodal evidence. Recently, numerous video-based MRAG benchmarks have been proposed to evaluate model capabilities across retrieval and generation stages in MRAG. However, existing benchmarks remain limited in modality coverage and format diversity, often focusing on single- or limited-modality tasks, or coarse-grained scene understanding. To address these gaps, we introduce CFVBench, a large-scale, manually verified benchmark constructed from 599 publicly available videos, yielding 5,360 open-ended QA pairs. CFVBench, spans high-density formats and domains such as chart-heavy reports, news broadcasts, and software tutorials, requiring models to retrieve and reason over long temporal video spans while maintaining fine-grained multimodal information. Using CFVBench, we systematically evaluate 7 retrieval methods and 14 widely-used MLLMs, revealing a critical bottleneck: current models (even GPT5 or Gemini) struggle to capture transient yet essential fine-grained multimodal details. To mitigate this, we propose Adaptive Visual Refinement (AVR), a plug-and-paly framework that adaptively increases frame sampling density and selectively invokes external tools when necessary. Experiments show that AVR consistently enhances fine-grained multimodal comprehension and improves performance across all evaluated MLLMs. Kaiwen Wei, Ruida Liu, Changzai Pan, Yidan Zhang 0002, Peijin Wang, Yingchao Feng |
WWW | 14 |
| 2026 | RingMoE: Mixture-of-Modality-Experts Multi-Modal Foundation Models for Universal Remote Sensing Image InterpretationabstractThe rapid advancement of foundation models has revolutionized visual representation learning in a self-supervised manner. However, their application in remote sensing (RS) remains constrained by a fundamental gap: existing models predominantly handle single or limited modalities, overlooking the inherently multi-modal nature of RS observations. Optical, synthetic aperture radar (SAR), and multi-spectral data offer complementary insights that significantly reduce the inherent ambiguity and uncertainty in single-source analysis. To bridge this gap, we introduce RingMoE, a unified multi-modal RS foundation model with 14.7 billion parameters, pre-trained on 400 million multi-modal RS images from nine satellites. RingMoE incorporates three key innovations: 1) A hierarchical Mixture-of-Experts (MoE) architecture comprising modal-specialized, collaborative, and shared experts, effectively modeling intra-modal knowledge while capturing cross-modal dependencies to mitigate conflicts between modal representations; 2) Physics-informed self-supervised learning, explicitly embedding sensor-specific radiometric characteristics into the pre-training objectives; 3) Dynamic expert pruning, enabling adaptive model compression from 14.7B to 1B parameters while maintaining performance, facilitating efficient deployment in Earth observation applications. Evaluated across 23 benchmarks spanning six key RS tasks (i.e., classification, detection, segmentation, tracking, change detection, and depth estimation), RingMoE outperforms existing foundation models and sets new SOTAs, demonstrating remarkable adaptability from single-modal to multi-modal scenarios. Beyond theoretical progress, it has been deployed and trialed in multiple sectors, including emergency response, land management, marine sciences, and urban planning. Hanbo Bi, Yingchao Feng, Boyuan Tong, Haichen Yu, Yongqiang Mao, Wenhui Diao, Peijin Wang, Yue Yu 0001, Hanyang Peng, Yehong Zhang, Kun Fu 0001, Xian Sun 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2026 | A Complex-Valued SAR Foundation Model Based on Physically Inspired Representation LearningabstractVision foundation models in remote sensing have been extensively studied due to their superior generalization on various downstream tasks. Synthetic Aperture Radar (SAR) offers all-day, all-weather imaging capabilities, providing significant advantages for Earth observation. However, establishing a foundation model for SAR image interpretation inevitably encounters the challenges of insufficient information utilization and poor interpretability. In this paper, we propose a remote sensing foundation model based on complex-valued SAR data, which simulates the polarimetric decomposition process for pre-training, i.e., characterizing pixel scattering intensity as a weighted combination of scattering bases and scattering coefficients, thereby endowing the foundation model with physical interpretability. Specifically, we construct a series of scattering queries, each representing an independent and meaningful scattering basis, which interact with SAR features in the scattering query decoder and output the corresponding scattering coefficient. To guide the pre-training process, polarimetric decomposition loss and power self-supervised loss are constructed. The former aligns the predicted coefficients with Yamaguchi coefficients, while the latter reconstructs power from the predicted coefficients and compares it to the input image's power. The performance of our foundation model is validated on nine typical downstream tasks, achieving state-of-the-art results. Notably, the foundation model can extract stable feature representations and exhibits strong generalization, even in data-scarce conditions. Hanbo Bi, Yingchao Feng, Linlin Xin, Shuo Gong, Peijin Wang, Wenhui Diao, Xian Sun 0001 |
IEEE Trans. Image Process. | 3 |
| 2025 | RS-vHeat: Heat Conduction Guided Efficient Remote Sensing Foundation ModelabstractRemote sensing foundation models largely break away from the traditional paradigm of designing task-specific models, offering greater scalability across multiple tasks. However, they face challenges such as low computational efficiency and limited interpretability, especially when dealing with large-scale remote sensing images. To overcome these, we draw inspiration from heat conduction, a physical process modeling local heat diffusion. Building on this idea, we are the first to explore the potential of using the parallel computing model of heat conduction to simulate the local region correlations in high-resolution remote sensing images, and introduce RS-vHeat, an efficient multi-modal remote sensing foundation model. Specifically, RS-vHeat 1) applies the Heat Conduction Operator (HCO) with a complexity of $O(N^{1.5})$ and a global receptive field, reducing computational overhead while capturing remote sensing object structure information to guide heat diffusion; 2) learns the frequency distribution representations of various scenes through a self-supervised strategy based on frequency domain hierarchical masking and multi-domain reconstruction; 3) significantly improves efficiency and performance over state-of-the-art techniques across 4 tasks and 10 datasets. Compared to attention-based remote sensing foundation models, we reduce memory usage by 84\%, FLOPs by 24\% and improves throughput by 2.7 times. The code will be made publicly available. Huiyang Hu, Peijin Wang, Hanbo Bi, Boyuan Tong, Zhaozhi Wang, Wenhui Diao, Yingchao Feng, Ziqi Zhang 0010, Yaowei Wang 0001, Qixiang Ye, Kun Fu 0001, Xian Sun 0001 |
ICCV | 8 |
| 2025 | AgMTR: Agent Mining Transformer for Few-Shot Segmentation in Remote Sensing
Hanbo Bi, Yingchao Feng, Yongqiang Mao, Jianning Pei, Wenhui Diao, Xian Sun 0001 |
Int. J. Comput. Vis. | 2 |
| 2025 | Prompt-and-Transfer: Dynamic Class-Aware Enhancement for Few-Shot SegmentationabstractFor more efficient generalization to unseen domains (classes), most Few-shot Segmentation (FSS) would directly exploit pre-trained encoders and only fine-tune the decoder, especially in the current era of large models. However, such fixed feature encoders tend to be class-agnostic, inevitably activating objects that are irrelevant to the target class. In contrast, humans can effortlessly focus on specific objects in the line of sight. This paper mimics the visual perception pattern of human beings and proposes a novel and powerful prompt-driven scheme, called "Prompt and Transfer" (PAT), which constructs a dynamic class-aware prompting paradigm to tune the encoder for focusing on the interested object (target class) in the current task. Three key points are elaborated to enhance the prompting: 1) Cross-modal linguistic information is introduced to initialize prompts for each task. 2) Semantic Prompt Transfer (SPT) that precisely transfers the class-specific semantics within the images to prompts. 3) Part Mask Generator (PMG) that works in conjunction with SPT to adaptively generate different but complementary part prompts for different individuals. Surprisingly, PAT achieves competitive performance on 4 different tasks including standard FSS, Cross-domain FSS (e.g., CV, medical, and remote sensing domains), Weak-label FSS, and Zero-shot Segmentation, setting new state-of-the-arts on 11 benchmarks. Hanbo Bi, Yingchao Feng, Wenhui Diao, Peijin Wang, Yongqiang Mao, Kun Fu 0001, Xian Sun 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | RingMo-Aerial: An Aerial Remote Sensing Foundation Model With Affine Transformation Contrastive LearningabstractAerial Remote Sensing (ARS) vision tasks present significant challenges due to the unique viewing angle characteristics. Existing research has primarily focused on algorithms for specific tasks, which have limited applicability in a broad range of ARS vision applications. This paper proposes RingMo-Aerial, aiming to fill the gap in foundation model research in the field of ARS vision. A Frequency-Enhanced Multi-Head Self-Attention (FE-MSA) mechanism is introduced to strengthen the model's capacity for small-object representation. Complementarily, an affine transformation-based contrastive learning method improves its adaptability to the tilted viewing angles inherent in ARS tasks. Furthermore, the ARS-Adapter, an efficient parameter fine-tuning method, is proposed to improve the model's adaptability and performance in various ARS vision tasks. Experimental results demonstrate that RingMo-Aerial achieves SOTA performance on multiple downstream tasks. This indicates the practicality and efficacy of RingMo-Aerial in enhancing the performance of ARS vision tasks. Wenhui Diao, Haichen Yu, Kaiyue Kang, Tong Ling, Yingchao Feng, Hanbo Bi, Libo Ren, Xuexue Li, Yongqiang Mao, Xian Sun 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2025 | RingMoGPT: A Unified Remote Sensing Foundation Model for Vision, Language, and Grounded TasksabstractRecently, multimodal large language models (MLLMs) have shown excellent reasoning capabilities in various fields. Most of the existing remote sensing (RS) MLLMs solve image-level text generation problems (e.g., image captioning), but ignore the core issues of object-level recognition, location, and multitemporal changes in the field of RS. In this article, we propose RingMoGPT, a multimodal foundation model that unifies vision, language, and localization. Based on the idea of domain adaption, RingMoGPT can complete training by fine-tuning only a few parameters. To make the model capable of object detection and change captioning, we further propose a location- and instruction-aware querying transformer (Q-Former) and a change detection module, respectively. To improve the performance of RingMoGPT, we carefully design the pretraining dataset and the instruction-tuning dataset. The pretraining dataset contains over a half million high-quality image and text pairs, which are generated through a low-cost and efficient data generation paradigm. The instruction-tuning dataset contains more than 1.6 million question-answer pairs, including six downstream tasks: scene classification, object detection, visual question answering (VQA), image captioning, grounded image captioning, and change captioning. Our experiments show that RingMoGPT performs well on six tasks, especially its ability to analyze multitemporal data changes and identify dense objects. We also verified the model under a zero-shot setting, and the results show that the proposed RingMoGPT also has good generalization ability in the face of new data. Peijin Wang, Huiyang Hu, Boyuan Tong, Ziqi Zhang 0010, Fanglong Yao, Yingchao Feng, Zining Zhu 0004, Wenhui Diao, Qixiang Ye, Xian Sun 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2024 | FAIR1M-GQA: Fine-Grained Grounded Question Answering Dataset in Remote SensingabstractWith the development of large language models (LLMs) and remote sensing technology, visual language (VL) tasks in the field of remote sensing have attracted more and more research attention. Commonly used VL datasets currently usually focus on the overall scene of the image, lacking the description of instance-level details such as location, size, category and so on. More importantly, users are usually not allowed to directly intercept regions in the image to ask questions through these datasets. However, the instance-level question answering based on these information is of great significance for target extraction in practical applications. In this manuscript, we build an innovative and challenging dataset FAIR1M-GQA. It unlocks the ability of the model to learn directly from text input and text output both with region coordinates, which are directly linked to fine-grained objects in remote sensing images. We experiment our dataset to verify the feasibility of the relevant task and provide the benchmark results. Huiyang Hu, Peijin Wang, Yingchao Feng, Wenhui Diao, Ziqi Zhang 0010, Xian Sun 0001, Kun Fu 0001 |
IGARSS | 3 |
| 2024 | Self-Training-Based Semantic-Balanced Network for Weakly Supervised Object Detection in Remote-Sensing ImagesabstractA weakly supervised object detection (WSOD) task is to train a detector with only image-level labels provided. Except for the training difficulty introduced by weaker annotations, the inherent complexity of the remote-sensing images (RSIs) also adds to the challenge. To boost the detector’s localization accuracy, we aim to exploit more semantic information contained in images and help improve the general robustness of the model. Noticing previous methods tend to focus on the most discriminative part of an object, we design a self-training-based network that leverages local semantic features. To this end, we develop a semantic-balanced localization module (SBLM) that distinguishes foreground from background and accurate proposals from incomplete ones, by leveraging a balance of region of interest (ROI) and its context information. Moreover, we find that the self-training strategy highly relies on the quality of pseudo-ground-truth boxes. Motivated by this possible lack of robustness, we design a comprehensive clustering module (CCM) and saliency-based proposal filtering (SPF) module that select pseudo-ground truth more comprehensively under supervision. To be more specific, CCM aims to reduce the arbitrariness during assigning pseudo-labels by considering multiple categorical vectors simultaneously. Salient object detection (SOD) is applied in the SPF module to help evaluate the quality of the chosen pseudo-ground-truth boxes. The detection performance is significantly boosted with the proposed method. Extensive experiments conducted on the NWPU VHR-10.v2 dataset and the DIOR dataset validate that the proposed model outperforms the previous state-of-the-art methods favorably with an mAP of 64.9% and 28.1%, respectively. Xuanyi Du, Wenhui Diao, Yingchao Feng, Xian Sun 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | PKG-Net: Physical Knowledge Feature-Guided Learning for Aircraft Detection in Optical Remote Sensing ImagesabstractExisting aircraft detection methods primarily rely on loss function constraints to guide the learning of end-to-end detectors, which significantly diverges from the judgment logic of human experts. Inspired by how human experts make decisions based on cognitive features, this article introduces the concept of physical knowledge features. Leveraging three properties of physical knowledge features, we identify the circle grayscale (CG) feature of aircraft and propose a physical knowledge-guided network (PKG-Net). By embedding CG features into the supervised learning process, the network improves its proficiency in learning stable features, thereby enhancing detection accuracy. Within this network, a multiscale circular frequency filter module (MS-CFFM) is responsible for extracting and integrating aircraft CG features across different scales. Adaptive channel selection module (ACSM) selectively activates channels for learning CG features. The hybrid attention feature fusion module (HAFFM) focuses intensively on the central localization of aircraft and deep reinforcement of channels. Experimental results on the RSOD and UCAS-AOD datasets demonstrate that the proposed method surpasses existing techniques in accuracy, achieving state-of-the-art performance. Linlin Xin, Wenhui Diao, Yingchao Feng, Xian Sun 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | SIRS: Multitask Joint Learning for Remote Sensing Foreground-Entity Image-Text RetrievalabstractThe essence of improving the effect of cross-modal image-text retrieval (CIR) lies in the finer-grained modeling of homogeneous features between modalities. However, in remote sensing (RS) scenarios, existing methods usually apply the image-sentence granular feature alignment paradigm, bringing significant difficulties to the fine-grained representation of homogeneous features between modalities. Besides, more complex background noise and extreme scale ranges of foreground targets are hard to distinguish, causing the feature mottle problem. To address the above issues, we propose a novel Semantic-guided Image-text Retrieval framework with Segmentation (SIRS). It is a multi-task joint learning framework for plug-and-play and end-to-end training RS CIR models efficiently, including Semantic-guided Spatial Attention (SSA) and Adaptive Multi-scale Weighting (AMW) modules. First, SSA introduces a background reconstruction branch based on noise perception and a semantic segmentation branch based on pixel-level prediction. It explores a joint learning strategy that concisely filters background noise and refines foreground features considerably. Secondly, AMW performs multi-scale weighting on various layers of feature map output by the encoder, effectively improving the learning efficiency of foreground targets at different scales. It is worth mentioning that SIRS outputs combination results with image and segmentation mask, which is not available in other methods. Based on the RSITMD dataset, we complete the semantic segmentation annotation RSITMD-SS to verify the performance of the proposed method. Sufficient and complete experiments verify the effectiveness of the proposed method. With SIRS, the mainstream SVP and CLIP-based methods improve about 7 mR and derive segmentation prediction with acceptable computational cost optionally. The code and associated dataset will be available on https://github.com/StarBurstStream0/SIRS. Zicong Zhu, Jian Kang 0005, Wenhui Diao, Yingchao Feng, Junxi Li, Jingen Ni |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | A Self-Supervised Cross-Modal Remote Sensing Foundation Model with Multi-Domain Representation and Cross-Domain FusionabstractThe construction of a basic model to extract generalized features from a large number of multimodal data is a new challenge in the field of remote sensing. Compared with natural scene images, When faced with a complex application scenario of remote sensing of multi-sensor acquisition, models that are suitable for a specific task are difficult to generalize to new scenarios. In this paper, we propose a model architecture based on the concepts of multi-domain representation and cross-domain fusion. By extracting strong generalization features from massive multi-modal data, a single foundation model can accomplish generalization interpretation for multiple downstream tasks. Experimental results show that the proposed model performs well on multiple downstream tasks, which validates the feasibility of the remote sensing cross-modal foundation model in the interpretation task. Yingchao Feng, Peijin Wang, Wenhui Diao, Qibin He 0001, Huiyang Hu, Hanbo Bi, Xian Sun 0001, Kun Fu 0001 |
IGARSS | 1 |
| 2023 | Not Just Learning From Others but Relying on Yourself: A New Perspective on Few-Shot Segmentation in Remote SensingabstractFew-shot segmentation (FSS) is proposed to segment unknown class targets with just a few annotated samples. Most current FSS methods follow the paradigm of mining the semantics from the support images to guide the query image segmentation. However, such a pattern of ‘learning from others’ struggles to handle the extreme intra-class variation, preventing FSS from being directly generalized to remote sensing scenes. To bridge the gap of intra-class variance, we develop a Dual-Mining network named DMNet for cross-image mining and self-mining, meaning that it no longer focuses solely on support images but pays more attention to the query image itself. Specifically, we propose a Class-public Region Mining (CPRM) module to effectively suppress irrelevant feature pollution by capturing the common semantics between the support-query image pair. The Class-specific Region Mining (CSRM) module is then proposed to continuously mine the class-specific semantics of the query image itself in a ‘filtering’ and ‘purifying’ manner. In addition, to prevent the co-existence of multiple classes in remote sensing scenes from exacerbating the collapse of FSS generalization, we also propose a new Known-class Meta Suppressor (KMS) module to suppress the activation of known-class objects in the sample. Extensive experiments on the iSAID and LoveDA remote sensing datasets have demonstrated that our method sets the state-of-the-art with a minimum number of model parameters. Significantly, our model with the backbone of Resnet-50 achieves the mIoU of 49.58% and 51.34% on iSAID under 1-shot and 5-shot settings, outperforming the state-of-the-art method by 1.8% and 1.12%, respectively. The code is publicly available at https://github.com/HanboBizl/DMNet/. Hanbo Bi, Yingchao Feng, Yongqiang Mao, Wenhui Diao, Xian Sun 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | RingMo-SAM: A Foundation Model for Segment Anything in Multimodal Remote-Sensing ImagesabstractThe proposal of Segment Anything Model (SAM) has created a new paradigm for deep learning-based semantic segmentation field, and has shown amazing generalization performance. However, we find it may fail or perform poorly on multimodal remote sensing scenarios, especially the Synthetic Aperture Radar (SAR) images. Besides, SAM does not provide category information of objects. In this paper, we propose a foundation model for multimodal remote sensing image segmentation called RingMo-SAM, which can not only segment anything in optical and SAR remote sensing data, but also identify object categories. First, a large-scale dataset containing millions of segmentation instances is constructed by collecting multiple open-source datasets in this field to train the model. Then, by constructing an instance-type and terrain-type category-decoupling mask decoder, the category-wise segmentation of various objects is achieved. In addition, a prompt encoder embedded with the characteristics of multimodal remote sensing data is designed. It not only supports multi-box prompts to improve the segmentation accuracy of multi-objects in complicated remote sensing scenes, but also supports SAR characteristics prompts to improve the segmentation performance on SAR images. Extensive experimental results on several datasets including iSAID, ISPRS Vaihingen, ISPRS Potsdam, AIR-PolSAR-Seg, etc. have demonstrated the effectiveness of our method. Junxi Li, Xuexue Li, Ruixue Zhou, Wenkai Zhang 0002, Yingchao Feng, Wenhui Diao, Kun Fu 0001, Xian Sun 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2022 | Soft Weighted Ordinal Classification for Monocular Height Estimation in Remote Sensing ImageabstractEstimating height information from a single remote sensing image is a critical component for 3D perception. Recent methods formulate it as a dense height prediction task based on regression loss functions. However, the regression accuracy is limited by the infinite continuous solution space. In this paper, we propose the soft weighted ordinal (SWO) classification loss for height prediction model to convert the regression problem with infinite continuous values into the classification problem with finite discrete values. which greatly improves the accuracy of high estimation. Specifically, we first define the discrete height rule and introduce the distance penalty metric to transform the continuous ground truth height value to the soft probability distributions. This is then used as supervised information to optimize the pixel-wise classification model. Finally, we utilize soft weighted summation to generate continuous height values in the inference phase. The proposed SWO classification loss can be used directly with existing dense prediction structures whose performance can be strengthened by direct replacement of the loss functions. Comprehensive experiments on the IS-PRS Vaihingen dataset show that the proposed method has achieved promising results. Yingchao Feng, Xian Sun 0001, Wenhui Diao, Tao Xu 0053, Kun Fu 0001 |
IGARSS | 1 |
| 2022 | Continual Learning With Structured Inheritance for Semantic Segmentation in Aerial ImageryabstractWith the rapid update and iteration of current aerial image data, the continual learning scenarios and catastrophic forgetting problem attracted increased attention, especially in the semantic segmentation task. However, the existing methods mainly focus on the class continual learning in a single task and are not satisfactory when extended to multiple tasks. In this article, we consider more realistic and complicated settings, namely task continual learning. We revisit the characteristics of semantic segmentation and knowledge distillation (KD) strategy, then propose a general and effective framework, named structured inheritance, to learn new tasks while retaining high performance on old tasks. Specifically, we present two structure-preserving penalties: pixel affinity structure loss and representation consistency structure loss. The former breaks the isolation of pixels and retains the pixel interactive information learned by the old tasks. At the same time, the latter protects high-frequency stationary information between sequence semantic segmentation tasks. Our approach does not need to add extra parameters nor does it need to access the data stream of the old tasks. Therefore, it can be applied in practical applications with strict computational burden, memory cost, and storage budget. Extensive continual learning experiments on four semantic segmentation datasets of Vaihingen, Potsdam, DeepGlobe, and Gaofen challenge semantic segmentation dataset (GCSS) prove the effectiveness of our proposed framework, which outperforms the current state-of-the-art methods and even exceeds the theoretical upper-bound performance of multitask learning. The code and models will be made publicly available. Yingchao Feng, Xian Sun 0001, Wenhui Diao, Kun Fu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Class-Incremental Learning Network for Small Objects Enhancing of Semantic Segmentation in Aerial ImageryabstractDue to the differences in the feature distribution between classes, when the model learns in a continuous data stream, it will encounter catastrophic forgetting. The incremental learning methods have shown great potential to solve this problem. However, most existing methods based on task-incremental learning are difficult to adapt to characteristics of remote sensing scenes with few differences in appearance but large differences in features, which is not conducive to artificially distinguish task-identity document (ID). Thus, we propose a class-incremental learning (CIL) network for small objects enhancing semantic segmentation in aerial imagery. Specifically, considering the superior accuracy of the binary classifier, we propose a twin-auxiliary (TA) model that adds an auxiliary binary classification task. Then, for expansion and contraction at the edge and small object confusion problems, we introduce a diversity distillation loss, using the results of binary-classifier to constrain the multiclass segmentation results and strengthen the attention to the locations of the segmentation results that have changed. Finally, we design a conflict reduction mechanism for multihead classifier to achieve single-head prediction for CIL. Experiments demonstrate that our method has good performance on the Vaihingen and Potsdam datasets by the International Society for Photogrammetry and Remote Sensing (ISPRS), outperforming state-of-the-art (SOTA) incremental learning methods. The code will be available soon. Junxi Li, Xian Sun 0001, Wenhui Diao, Peijin Wang, Yingchao Feng, Guangluan Xu |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | Random Topology and Random Multiscale Mapping: An Automated Design of Multiscale and Lightweight Neural Network for Remote-Sensing Image RecognitionabstractWith the proposal of neural architecture search (NAS), automated network architecture design gradually becomes a new way in deep learning research. Due to its high capability regarding automated design, some pioneers have made an attempt to apply NAS in remote sensing and made some achievements, like 1-D/3-D Auto-convolutional neural network (CNN) and polarimetric synthetic aperture radar (PolSAR)-tailored Differentiable Architecture Search (PDAS). However, there are still some areas to be improved for existing NAS in remote-sensing field. In this article, we propose a random topology and random multiscale mapping (RTRMM) method to generate a multiscale and lightweight architecture for remote-sensing image recognition. First, a random topology generator generates the topology through random graph. Second, during the experiment, we find remote-sensing image features extracted by a multiscale network are more appropriate, compared with features extracted by a single-scale model. Nevertheless, the complexity inevitably increases with the introduction of a multiscale concept. Consequently, we design a variable search space consisting of decomposition convolution units under the guidance of mathematical analysis. The mapping of each neuron is then determined by a random multiscale mapping sampler. After that, we assemble the topology and mappings into blocks and construct three RTRMM models. Experiments on four scene classification datasets confirm the feature extraction capability and lightweight performance of RTRMM models. Moreover, we also observe that our approach achieves a better tradeoff between floating-point operations (FLOPs) and accuracy than some current well-behaved methods. Furthermore, the results on Vaihingen dataset verify the high feature-transfer capability. Martin Weinmann, Xian Sun 0001, Wenhui Diao, Yingchao Feng, Kun Fu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | LIL: Lightweight Incremental Learning Approach Through Feature Transfer for Remote Sensing Image Scene ClassificationabstractExisting deep learning models usually assume that all data obeys independent identically distribution, which is unreasonable in remote sensing. Due to the differences in camera parameters, spectral ranges, resolutions, and so on, the images acquired by remote sensing sensors may be greatly diverse, causing models to face catastrophic forgetting when they are trained on new data only. Thus, incremental learning is introduced. An ideal incremental learning model should be expanded as the number of tasks increases, so as to have enough ability to adapt to the changes in data. However, existing approaches normally expand heavy modules for each task, making the holistic models cumbersome. In this article, a lightweight incremental learning approach (LIL) is proposed for remote sensing image scene classification. We replace the role of the feature extractor with extracting features of a single task instead of task-sharing features of all tasks to lighten the backbone. In addition, we propose a light feature transfer module (FTM) to realize the alignment of data distributions between different tasks in the feature domain. Furthermore, dual-constraint loss with knowledge distillation and adversarial learning is introduced to promote the mapping and alignment of data distributions at both the feature level and the semantic level. In LIL, only a tiny FTM and a classifier are added to the model when the model learns a new task. Experimental results show that our approach with a small number of parameters outperforms state-of-the-art approaches for incremental learning on both a single dataset and a sequence of multiple datasets. Xian Sun 0001, Wenhui Diao, Yingchao Feng, Peijin Wang, Kun Fu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Improving Semantic Segmentation in Aerial Imagery via Graph Reasoning and Disentangled LearningabstractSemantic segmentation in aerial imagery is still an important, yet challenging task due to the complex characteristics of remote-sensing data. The critical issues consist of: 1) extreme foreground–background imbalance; 2) large intra-class variance; and 3) arbitrary-oriented, dense, and small objects. The above challenges make it unlikely to model the effective global interdependencies of semantic heterogeneous regions. Besides, general semantic segmentation methods suffer from feature ambiguity due to the joint feature learning paradigm, leading to inferior detail information. In this article, we propose an improved semantic segmentation framework to tackle these problems via graph reasoning (GR) and disentangled learning. On the one hand, a simple, yet effective GR unit is introduced to implement coordinate-interaction space mapping and perform relation reasoning over the graph. It can be deployed on the feature pyramid network (FPN) to exploit cross-stage multi-scale information. On the other hand, we propose a so- called disentangled learning paradigm to explicitly model the foreground and boundary objects, instantiated as foreground prior estimation (FPE) and boundary alignment (BA). The indication of the intermediate feature can be effectively emphasized to enhance the discriminative abilities of the network. Extensive experiments over iSAID, ISPRS Vaihingen, and the general Cityscapes datasets demonstrate the effectiveness and efficiency of the proposed framework over other state-of-the-art semantic segmentation methods. Ruigang Niu, Xian Sun 0001, Wenhui Diao, Yingchao Feng, Kun Fu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2021 | Double Similarity Distillation for Semantic Image SegmentationabstractThe balance between high accuracy and high speed has always been a challenging task in semantic image segmentation. Compact segmentation networks are more widely used in the case of limited resources, while their performances are constrained. In this paper, motivated by the residual learning and global aggregation, we propose a simple yet general and effective knowledge distillation framework called double similarity distillation (DSD) to improve the classification accuracy of all existing compact networks by capturing the similarity knowledge in pixel and category dimensions, respectively. Specifically, we propose a pixel-wise similarity distillation (PSD) module that utilizes residual attention maps to capture more detailed spatial dependencies across multiple layers. Compared with exiting methods, the PSD module greatly reduces the amount of calculation and is easy to expand. Furthermore, considering the differences in characteristics between semantic segmentation task and other computer vision tasks, we propose a category-wise similarity distillation (CSD) module, which can help the compact segmentation network strengthen the global category correlation by constructing the correlation matrix. Combining these two modules, DSD framework has no extra parameters and only a minimal increase in FLOPs. Extensive experiments on four challenging datasets, including Cityscapes, CamVid, ADE20K, and Pascal VOC 2012, show that DSD outperforms current state-of-the-art methods, proving its effectiveness and generality. The code and models will be publicly available. Yingchao Feng, Xian Sun 0001, Wenhui Diao |
IEEE Trans. Image Process. | 1 |
| 2019 | Ship Instance Segmentation from Remote Sensing Images Using Sequence Local Context ModuleabstractThe performance of object instance segmentation in remote sensing images has been greatly improved through the introduction of many landmark frameworks based on convolutional neural network. However, the object densely issue still affects the accuracy of such segmentation frameworks. Objects of the same class are easily confused, which is most likely due to the close docking between objects. We think context information is critical to address this issue. So, we propose a novel framework called SLCMASK-Net, in which a sequence local context module (SLC) is introduced to avoid confusion between objects of the same class. The SLC module applies a sequence of dilation convolution blocks to progressively learn multi-scale context information in the mask branch. Besides, we try to add SLC module to different locations in our framework and experiment with the effect of different parameter settings. Comparative experiments are conducted on remote sensing images acquired by QuickBird with a resolution of 0.5m - 1m and the results show that the proposed method achieves state-of-the-art performance. Yingchao Feng, Wenhui Diao, Yi Zhang 0026, Hao Li 0087, Zhonghan Chang, Menglong Yan, Xian Sun 0001 |
IGARSS | 1 |
| 2019 | Effective Classification of Local Climate Zones Based on Multi-Source Remote Sensing DataabstractThe local climate zone (LCZ) classification divides the urban areas into 17 categories, which are composed of 10 manmade structures and 7 natural landscapes. Though originally designed for temperature study, LCZ classification can be used for studies on economy and population. In this paper, we achieve a LCZ classification with convolutional neural networks based on the multi-source remote sensing data, including the polarimetric synthetic aperture radar (PolSAR) data and the corresponding multi-spectral imagery (MSI). Through experiments we attempt to reveal the contributions of the SAR data and the MSI to the classification performance. Furthermore, we emphasize the crucial importance of the preprocessing on the training data to derive a balanced dataset. We are ranked second in the Tianchi competition rankings when we submit our results. Yingchao Feng, Wenkai Zhang 0002, Yue Zhang 0016, Siyue Wang, Kun Fu 0001, Kaiqiang Chen |
IGARSS | 2 |