Junxi Li

dblp:314/0443 · DBLP profile ↗
← Back
11ranked-venue papers
2as first author
11since 2021 · last 2026
0000-0002-9428-9751ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 9 · 2 first-author · 9 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 VPAF-FSCIL: Virtual prototype calibration and parameter-adaptive freezing for few-shot class-incremental learning
abstract
Class Incremental Learning (CIL) enables models to learn new classes in images while retaining knowledge of previously learned classes. However, conventional CIL assumes abundant data for new classes, which is often unrealistic in real-world applications such as autonomous driving and medical diagnosis. Few-Shot Class Incremental Learning (FSCIL) addresses this issue by enabling the learning of new classes with minimal data. Nevertheless, it faces challenges like catastrophic forgetting, feature drift, and class imbalance, which can hinder effective adaptation. To tackle these challenges, we propose a method that integrates a feature-decoupled network architecture with a Parameter-Adaptive Freezing (PAF) training strategy, aiming to improve feature stability and adaptability to new classes. Inspired by Equiangular Tight Frame (ETF) theory, we decouple the model into three components: a Feature Extractor, a Feature Distributor, and an ETF Classifier. This structure integrates adaptive virtual class space allocation to ensure optimal embeddings for new classes. Moreover, PAF selectively freezes critical parameters based on the Fisher Information Matrix, which helps minimize forgetting of old classes while simultaneously improving generalization for new classes. Extensive experiments across CIFAR-100, miniImageNet, and CUB-200-2011 demonstrate that our proposed method outperforms state-of-the-art FSCIL methods in both classification accuracy and stability.
Junxi Li
Neural Networks2
2024 Few-Shot Incremental Object Detection in Aerial Imagery via Dual-Frequency Prompt
abstract
Recently, there has been a growing interest in few-shot incremental object detection (FSIOD). It learns new tasks with limited data while mitigating catastrophic forgetting on previous tasks. However, existing FSIOD methods experience parameter changes after training on new tasks, causing a parameter competition issue among tasks. Additionally, the background information differs among various tasks, and using a common background weight for all tasks results in the background shift. Constrained by these two issues, existing methods only alleviate catastrophic forgetting and cannot wholly prevent the performance decline on previous tasks. Especially for complex remote sensing images with messy background, the models trained on new tasks exhibit noticeable performance drops on previous tasks. In this paper, we propose a novel FSIOD method via dual-frequency prompt to address these challenges, named FSIOD-DFP. It can completely eliminate catastrophic forgetting while mitigating over-fitting. Specifically, a dual-frequency prompt generator is designed to tackle the parameter competition issue. It decouples the frequency components of images to produce prompts that modify the images to adapt to the base model trained on previous tasks. Compared to traditional prompts, our generator introduces fewer parameters to address over-fitting for limited data and allows freezing the base model to maintain the performance of previous data. Besides, a self-regularization loss is introduced to guide the prompt-modified images to leverage the knowledge of the base model effectively. Furthermore, we propose a task-decoupled detection head to address the background shift problem. It separates the detection heads for new and previous tasks to resolve the conflict in the background between different tasks. In FSIOD-DFP, only a prompt generator and a novel detection head are added and fine-tuned when learning a new task. Experiments on three remote sensing object detection datasets demonstrate that our method achieves state-of-the-art performance on both new and previous tasks in all few-shot incremental settings.
Wenhui Diao, Junxi Li, Yidan Zhang 0002, Peijin Wang, Xian Sun 0001, Kun Fu 0001
IEEE Trans. Geosci. Remote. Sens.3
2024 SIRS: Multitask Joint Learning for Remote Sensing Foreground-Entity Image-Text Retrieval
abstract
The essence of improving the effect of cross-modal image-text retrieval (CIR) lies in the finer-grained modeling of homogeneous features between modalities. However, in remote sensing (RS) scenarios, existing methods usually apply the image-sentence granular feature alignment paradigm, bringing significant difficulties to the fine-grained representation of homogeneous features between modalities. Besides, more complex background noise and extreme scale ranges of foreground targets are hard to distinguish, causing the feature mottle problem. To address the above issues, we propose a novel Semantic-guided Image-text Retrieval framework with Segmentation (SIRS). It is a multi-task joint learning framework for plug-and-play and end-to-end training RS CIR models efficiently, including Semantic-guided Spatial Attention (SSA) and Adaptive Multi-scale Weighting (AMW) modules. First, SSA introduces a background reconstruction branch based on noise perception and a semantic segmentation branch based on pixel-level prediction. It explores a joint learning strategy that concisely filters background noise and refines foreground features considerably. Secondly, AMW performs multi-scale weighting on various layers of feature map output by the encoder, effectively improving the learning efficiency of foreground targets at different scales. It is worth mentioning that SIRS outputs combination results with image and segmentation mask, which is not available in other methods. Based on the RSITMD dataset, we complete the semantic segmentation annotation RSITMD-SS to verify the performance of the proposed method. Sufficient and complete experiments verify the effectiveness of the proposed method. With SIRS, the mainstream SVP and CLIP-based methods improve about 7 mR and derive segmentation prediction with acceptable computational cost optionally. The code and associated dataset will be available on https://github.com/StarBurstStream0/SIRS.
Zicong Zhu, Jian Kang 0005, Wenhui Diao, Yingchao Feng, Junxi Li, Jingen Ni
IEEE Trans. Geosci. Remote. Sens.5
2023 Breaking Immutable: Information-Coupled Prototype Elaboration for Few-Shot Object Detection
abstract
Few-shot object detection, expecting detectors to detect novel classes with a few instances, has made conspicuous progress. However, the prototypes extracted by existing meta-learning based methods still suffer from insufficient representative information and lack awareness of query images, which cannot be adaptively tailored to different query images. Firstly, only the support images are involved for extracting prototypes, resulting in scarce perceptual information of query images. Secondly, all pixels of all support images are treated equally when aggregating features into prototype vectors, thus the salient objects are overwhelmed by the cluttered background. In this paper, we propose an Information-Coupled Prototype Elaboration (ICPE) method to generate specific and representative prototypes for each query image. Concretely, a conditional information coupling module is introduced to couple information from the query branch to the support branch, strengthening the query-perceptual information in support features. Besides, we design a prototype dynamic aggregation module that dynamically adjusts intra-image and inter-image aggregation weights to highlight the salient information useful for detecting query images. Experimental results on both Pascal VOC and MS COCO demonstrate that our method achieves state-of-the-art performance in almost all settings. Code will be available at: https://github.com/lxn96/ICPE.
Wenhui Diao, Yongqiang Mao, Junxi Li, Peijin Wang, Xian Sun 0001, Kun Fu 0001
AAAI4
2023 Few-Shot Object Detection in Aerial Imagery Guided by Text-Modal Knowledge
abstract
Few-shot object detection (FSOD) has received numerous attention due to the difficulty and time-consuming of labeling objects. Recent researches achieve excellent performance in a natural scene by only using a few instances of novel classes to fine-tune the last prediction layer of the model well-trained on plentiful base data. However, compared with natural scene objects with a single direction and small size variety, the direction and size of the objects in remote sensing images (RSIs) vary greatly. The methods proposed for the natural scene cannot be directly applied to RSIs. In this article, we first propose a strong baseline for RSIs. It fine-tunes all detector components acting on high-level features and effectively improves the performance of novel classes. Further analyzing the results of the baseline, we find that the error for novel classes is mainly concentrated in classification. It misclassifies novel classes as confusable base classes or backgrounds due to the difficulty in extracting generalized information from limited instances. As is well-known, text-modal knowledge can highly summarize the generalized and unique characteristics of categories. Thus, we introduce text-modal descriptions for each category and propose an FSOD method guided by TExt-MOdal knowledge, called TEMO. Specifically, a text-modal knowledge extractor and a cross-modal assembly module are proposed to extract text features and fuse the text-modal features into visual-modal features. The fused features greatly reduce the classification confusion of novel classes. Furthermore, we introduce a mask strategy and a separation loss to avoid over-fitting and ambiguity of text-modal features. Experimental results on detection in optical remote sensing images (DIOR), Northwestern Polytechnical University (NWPU), and fine-grained object recognition in high-resolution remote sensing imagery (FAIR1M) illustrate that our TEMO achieves state-of-the-art performance in all settings.
Xian Sun 0001, Wenhui Diao, Yongqiang Mao, Junxi Li, Yidan Zhang 0002, Peijin Wang, Kun Fu 0001
IEEE Trans. Geosci. Remote. Sens.5
2023 RingMo: A Remote Sensing Foundation Model With Masked Image Modeling
abstract
Deep learning approaches have contributed to the rapid development of remote sensing (RS) image interpretation. The most widely used training paradigm is to use ImageNet pretrained models to process RS data for specified tasks. However, there are issues such as domain gap between natural and RS scenes and the poor generalization capacity of RS models. It makes sense to develop a foundation model with general RS feature representation. Since a large amount of unlabeled data is available, the self-supervised method has more development significance than the fully supervised method in RS. However, most of the current self-supervised methods use contrastive learning, whose performance is sensitive to data augmentation, additional information, and selection of positive and negative pairs. In this article, we leverage the benefits of generative self-supervised learning (SSL) for RS images and propose an RS foundationmodel framework called RingMo, which consists of two parts. First, a large-scale dataset is constructed by collecting two million RS images from satellite and aerial platforms, covering multiple scenes and objects around the world. Second, we propose an RS foundation model training method designed for dense and small objects in complicated RS scenes. We show that the foundation model trained on our dataset with RingMo method achieves state-of-the-art (SOTA) on eight datasets across four downstream tasks, demonstrating the effectiveness of the proposed framework. Through in-depth exploration, we believe it is time for RS researchers to embrace generative SSL and leverage its general representation capabilities to speed up the development of RS applications.
Xian Sun 0001, Peijin Wang, Wanxuan Lu, Zicong Zhu, Qibin He 0001, Junxi Li, Xuee Rong, Zhujun Yang, Qinglin He, Ruiping Wang 0001, Jiwen Lu, Kun Fu 0001
IEEE Trans. Geosci. Remote. Sens.7
2023 RingMo-SAM: A Foundation Model for Segment Anything in Multimodal Remote-Sensing Images
abstract
The proposal of Segment Anything Model (SAM) has created a new paradigm for deep learning-based semantic segmentation field, and has shown amazing generalization performance. However, we find it may fail or perform poorly on multimodal remote sensing scenarios, especially the Synthetic Aperture Radar (SAR) images. Besides, SAM does not provide category information of objects. In this paper, we propose a foundation model for multimodal remote sensing image segmentation called RingMo-SAM, which can not only segment anything in optical and SAR remote sensing data, but also identify object categories. First, a large-scale dataset containing millions of segmentation instances is constructed by collecting multiple open-source datasets in this field to train the model. Then, by constructing an instance-type and terrain-type category-decoupling mask decoder, the category-wise segmentation of various objects is achieved. In addition, a prompt encoder embedded with the characteristics of multimodal remote sensing data is designed. It not only supports multi-box prompts to improve the segmentation accuracy of multi-objects in complicated remote sensing scenes, but also supports SAR characteristics prompts to improve the segmentation performance on SAR images. Extensive experimental results on several datasets including iSAID, ISPRS Vaihingen, ISPRS Potsdam, AIR-PolSAR-Seg, etc. have demonstrated the effectiveness of our method.
Junxi Li, Xuexue Li, Ruixue Zhou, Wenkai Zhang 0002, Yingchao Feng, Wenhui Diao, Kun Fu 0001, Xian Sun 0001
IEEE Trans. Geosci. Remote. Sens.2
2023 Label Propagation and Contrastive Regularization for Semisupervised Semantic Segmentation of Remote Sensing Images
abstract
Remarkable progress based on deep neural networks has been achieved on the semantic segmentation in remote sensing images. However, pixel-level labeling is expensive for remote sensing images. Semi-supervised semantic segmentation becomes an alternative approach to reduce the cost of annotation, and it is crucial to utilize efficiently a large number of unlabeled data. Nevertheless inevitably, there is the unbalanced class distribution between labeled and unlabeled data of remote sensing scene. Existing semi-supervised methods train unlabeled images in isolation from labeled images and only learn reliable pixel pseudo-labels, leading to underutilization of unlabeled images. This article proposes a novel semi-supervised semantic segmentation approach based on label propagation and contrastive regularization for remote sensing images. Specifically, the unlabeled images are augmented by randomly copy-pasting the class regions from labeled images. A prototype feature constraint module is used to enforce the constraint on the pixel features of unlabeled images relying on the prototype features from labeled images, achieving feature alignment on the entire dataset. Furthermore, we present the region contrastive learning module that guides the model to learn feature consistency under different perturbations and compact feature representations over class regions on unlabeled images. Extensive experimental results on multiple remote sensing datasets demonstrate that our proposed approach achieves superior performance compared with state-of-the-art semi-supervised semantic segmentation methods.
Zhujun Yang, Wenhui Diao, Yuzhuo Kang, Junxi Li, Xian Sun 0001
IEEE Trans. Geosci. Remote. Sens.6
2023 Bridging the Gap Between Cumbersome and Light Detectors via Layer-Calibration and Task-Disentangle Distillation in Remote Sensing Imagery
abstract
With urgent application requirements, such as satellite in-orbit processing and unmanned aerial vehicle tracking, knowledge distillation (KD) following the teacher–student teaching mechanism has shown great potential to obtain lightweight detectors. However, compact students have limited accuracy due to the interference of large-scale variations and blurred boundaries in remote sensing objects. Specifically, previous methods mostly force teacher–student responses from the layer of the same depth and scale to align. Stereotyped manual interlayer associations may cause discriminative features of multiscale objects to be incorrectly bundled. Furthermore, the regression branch follows the identical distillation paradigm as the classification branch, resulting in ambiguous object bounding box deviations. To solve the above two issues, we propose an effective KD framework called layer-calibration and task-disentangle distillation (LTD). First, the cross-layer calibration distillation (CCD) structure is innovatively proposed. It adaptively binds a student layer with several related target layers, rather than a fixed layer in the teacher model. Appropriate and clear knowledge of large and small objects is transmitted. Since the CCD structure requires explicit global inner product computation between multiple layers, the local implicit calibration (LIC) module is further proposed to reduce distilled convergence difficulty. Second, the task-aware spatial disentangle distillation (TASD) structure is devised to transfer task-decoupled semantics and localization knowledge in a divide-and-conquer manner, alleviating objects’ localization imprecision. Experiments demonstrate that our LTD achieves state-of-the-art performance on several datasets and is a plug-and-play approach to most detectors. The code will be available soon.
Yidan Zhang 0002, Xian Sun 0001, Junxi Li, Yongqiang Mao, Lei Wang 0077
IEEE Trans. Geosci. Remote. Sens.5
2022 SIL-LAND: Segmentation Incremental Learning in Aerial Imagery via LAbel Number Distribution Consistency
abstract
Segmentation incremental learning has received a lot of attention in recent years due to the ability to overcome the problem of catastrophic forgetting. Our study found that differences in label number distribution affect the performance of segmentation incremental learning. Because the labels for pixels of the old category are marked as background when the model is trained on the new tasks, the label number distribution is inconsistent with static learning that is considered to be the upper bound on incremental learning, which hinders the mitigation of the catastrophic forgetting problem. In response to the above problems, we propose an incremental learning method named SIL-LAND, which improves the accuracy by making the label number distribution of our method close to that of static learning. From the perspective of high-level semantic labels, we propose the prototype update mechanism for the problem that non-adaptive representative prototypes ignore the sample diversity of semantic categories in remote sensing images. By compensating for the difference in label number distribution at the feature level, the distance between the prototype and the actual class center is reduced; Aiming at the lack of semantic consistency between feature vectors and prototypes, we propose a similarity measure module to increase the intra-class similarity between the prototype and corresponding feature vectors. From the perspective of one-hot labels, we propose label reconstruction, including foreground screening and background padding to make the number distribution of one-hot labels as close as possible to that of static learning. A series of experimental results demonstrate the effectiveness of our method.
Junxi Li, Wenhui Diao, Peijin Wang, Yidan Zhang 0002, Zhujun Yang, Guangluan Xu, Xian Sun 0001
IEEE Trans. Geosci. Remote. Sens.1
2022 Class-Incremental Learning Network for Small Objects Enhancing of Semantic Segmentation in Aerial Imagery
abstract
Due to the differences in the feature distribution between classes, when the model learns in a continuous data stream, it will encounter catastrophic forgetting. The incremental learning methods have shown great potential to solve this problem. However, most existing methods based on task-incremental learning are difficult to adapt to characteristics of remote sensing scenes with few differences in appearance but large differences in features, which is not conducive to artificially distinguish task-identity document (ID). Thus, we propose a class-incremental learning (CIL) network for small objects enhancing semantic segmentation in aerial imagery. Specifically, considering the superior accuracy of the binary classifier, we propose a twin-auxiliary (TA) model that adds an auxiliary binary classification task. Then, for expansion and contraction at the edge and small object confusion problems, we introduce a diversity distillation loss, using the results of binary-classifier to constrain the multiclass segmentation results and strengthen the attention to the locations of the segmentation results that have changed. Finally, we design a conflict reduction mechanism for multihead classifier to achieve single-head prediction for CIL. Experiments demonstrate that our method has good performance on the Vaihingen and Potsdam datasets by the International Society for Photogrammetry and Remote Sensing (ISPRS), outperforming state-of-the-art (SOTA) incremental learning methods. The code will be available soon.
Junxi Li, Xian Sun 0001, Wenhui Diao, Peijin Wang, Yingchao Feng, Guangluan Xu
IEEE Trans. Geosci. Remote. Sens.1