VLDB 2026 Research / reviewers in the wild / expert
Xuee Rong
dblp:314/0295
· DBLP profile ↗
12ranked-venue papers
2as first author
12since 2021 · last 2024
0000-0002-7681-7020ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 10 · 2 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Dynamic and Adaptive Self-Training for Semi-Supervised Remote Sensing Image Semantic SegmentationabstractRemote sensing technology has made remarkable progress, providing a wealth of data for various applications, such as ecological conservation and urban planning. However, the meticulous annotation of this data is labor-intensive, leading to a shortage of labeled data, particularly in tasks like semantic segmentation. Semi-supervised methods, combining consistency regularization with self-training, offer a solution to efficiently utilize labeled and unlabeled data. However, these methods encounter challenges due to imbalanced data ratios. To tackle these challenges, we introduce a self-training approach namedDAST(Dynamic andAdaptiveSelf-Training), which is combined with dynamic pseudo-label sampling, distribution matching, and adaptive threshold updating. Dynamic pseudo-label sampling is tailored to address the issue of class distribution imbalance by giving priority to classes with fewer samples. Meanwhile, distribution matching and adaptive threshold updating aim to reduce distribution disparities by adjusting model predictions across augmented images within the framework of consistency regularization, ensuring they align with the actual data distribution. Experiment results on the Potsdam and iSAID datasets demonstrate that DAST effectively balances class distribution, aligns model predictions with data distribution, and stabilizes pseudo-labels, leading to state-of-the-art performance on both datasets. These findings highlight the potential of DAST in overcoming the challenges associated with significant disparities in labeled-to-unlabeled data ratios. Jidong Jin, Wanxuan Lu, Xuee Rong, Xian Sun 0001, Yirong Wu |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | SemiPSCN: Polarization Semantic Constraint Network for Semi-Supervised Segmentation in Large-Scale and Complex-Valued PolSAR ImagesabstractSince polarimetric synthetic aperture radar (PolSAR) terrain segmentation is a dense prediction task, the disadvantage of inadequate labeled samples greatly limits its performance. In this article, we present a semi-supervised segmentation network called SemiPSCN to reduce the data reliance on label annotation, which integrates semi-supervised learning (SSL) paradigm and the characteristics of PolSAR data into a unified architecture. First, considering the unreliability of pseudolabels caused by noise interference in PolSAR data, a pseudolabel error localization (PEL) module is designed. By mapping the pixels that have mispredictions in pseudolabels, PEL can greatly enhance the confidence of pseudolabels. Then, SemiPSCN introduces a category representation constraint (CRC) module to explicitly boost the category consistency between labeled and unlabeled PolSAR data. Via explicit intracategory and intercategory constraints, CRC can guarantee the invariant representations on the same category region between labeled and unlabeled data. Furthermore, a region consistency constraint (RCC) module is designed to enhance the regional consistency in PolSAR data. RCC leverages the conception of graph to model the understanding of spatial relationships among terrain targets, thereby facilitating consistent spatial region expression in semi-supervised process. Finally, we build a challenging large-scale dataset called LSPolSAR-Seg and conduct abundant experiments on LSPolSAR-Seg. SemiPSCN exhibits superior performance when compared with other advanced approaches, especially improving mean intersection over union (mIoU) by 3.44%–12.77% under 20% split setting, which promotes the performance to a state-of-the-art level. Xuan Zeng 0004, Zhirui Wang 0003, Yuelei Wang, Xuee Rong, Pengyu Guo, Xian Sun 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Beyond the limitation of monocular 3D detector via knowledge distillationabstractKnowledge distillation (KD) is a promising approach that facilitates the compact student model to learn dark knowledge from the huge teacher model for better results. Although KD methods are well explored in the 2D detection task, existing approaches are not suitable for 3D monocular detection without considering spatial cues. Motivated by the potential of depth information, we propose a novel distillation framework that validly improves the performance of the student model without extra depth labels. Specifically, we first put forward a perspective-induced feature imitation, which utilizes the perspective principle (the farther the smaller) to facilitate the student to imitate more features of farther objects from the teacher model. Moreover, we construct a depth-guided matrix by the predicted depth gap of teacher and student to facilitate the model to learn more knowledge of farther objects in prediction level distillation. The proposed method is available for advanced monocular detectors with various backbones, which also brings no extra inference time. Extensive experiments on the KITTI and nuScenes benchmarks with diverse settings demonstrate that the proposed method outperforms the state-of-the-art KD methods. Dongshuo Yin, Xuee Rong, Xian Sun 0001, Wenhui Diao |
ICCV | 3 |
| 2023 | MiCro: Modeling Cross-Image Semantic Relationship Dependencies for Class-Incremental Semantic Segmentation in Remote Sensing ImagesabstractContinual learning is an effective way to overcome catastrophic forgetting (CF) in incremental learning for semantic segmentation. The existing continual semantic segmentation (CSS) methods of remote sensing (RS) ignore the semantic relationships among pixels across different images, which will lead to disappointing segmentation results, such as edge pixel misclassification and small object omission. In this paper, we propose a framework for modeling cross-image semantic relationship dependencies (MiCro), which aims to learn an inter-class separable and intra-class cohesive feature space from the pixel relationships across various images to ensure that learned categories can prevent CF in the incremental process. Specifically, we exploit the relationships among pixels of images in mini-batch to construct three losses: (a) Cross-image feature relationship distillation (CFRD) loss, which builds a well-structured feature space; (b) Cross-image intra-class feature cohesion (CIFC) loss, which is devised to make intra-class features more cohesive; and (c) Cross-image class-area weighted cross-entropy (CCWCE) loss, which is mainly employed to inversely weight the proportion of category area in mini-batch. The effectiveness of the proposed approach is demonstrated by extensive experiments on three RS semantic segmentation datasets from ISPRS Vaihingen, ISPRS Potsdam, and iSAID. MiCro is superior to the current most advanced methods in most incremental settings, especially improving mIoU by 11.59% on ISPRS Vaihingen, 13.17% on ISPRS Potsdam, and 15.01% on iSAID in the most difficult incremental settings, which promotes the CSS to a state-of-the-art (SOTA) level. The code will be available at https://github.com/RongXueE/MiCro. Xuee Rong, Peijin Wang, Wenhui Diao, Wenxin Yin, Xuan Zeng 0004, Xian Sun 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | RingMo: A Remote Sensing Foundation Model With Masked Image ModelingabstractDeep learning approaches have contributed to the rapid development of remote sensing (RS) image interpretation. The most widely used training paradigm is to use ImageNet pretrained models to process RS data for specified tasks. However, there are issues such as domain gap between natural and RS scenes and the poor generalization capacity of RS models. It makes sense to develop a foundation model with general RS feature representation. Since a large amount of unlabeled data is available, the self-supervised method has more development significance than the fully supervised method in RS. However, most of the current self-supervised methods use contrastive learning, whose performance is sensitive to data augmentation, additional information, and selection of positive and negative pairs. In this article, we leverage the benefits of generative self-supervised learning (SSL) for RS images and propose an RS foundationmodel framework called RingMo, which consists of two parts. First, a large-scale dataset is constructed by collecting two million RS images from satellite and aerial platforms, covering multiple scenes and objects around the world. Second, we propose an RS foundation model training method designed for dense and small objects in complicated RS scenes. We show that the foundation model trained on our dataset with RingMo method achieves state-of-the-art (SOTA) on eight datasets across four downstream tasks, demonstrating the effectiveness of the proposed framework. Through in-depth exploration, we believe it is time for RS researchers to embrace generative SSL and leverage its general representation capabilities to speed up the development of RS applications. Xian Sun 0001, Peijin Wang, Wanxuan Lu, Zicong Zhu, Qibin He 0001, Junxi Li, Xuee Rong, Zhujun Yang, Qinglin He, Ruiping Wang 0001, Jiwen Lu, Kun Fu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 8 |
| 2023 | MoCG: Modality Characteristics-Guided Semantic Segmentation in Multimodal Remote Sensing ImagesabstractThe rapid development of satellite platforms has yielded copious and diverse multi-source data for earth observation, greatly facilitating the growth of multimodal semantic segmentation (MSS) in remote sensing. However, MSS also suffers from numerous challenges: 1) Existing inherent defects in each modality due to the different imaging mechanisms. 2) Insufficient exploration of the intrinsic characteristics of modalities. 3) The existence of the huge semantic gap between heterogeneous data causes difficulties in feature fusion. The inability to effectively utilize the rich and diverse information provided by each modality and ignorance of the heterogeneity between modalities will hinder the feature enhancement, and further significantly impacts the semantic segmentation accuracy. Furthermore, neglecting the huge gap makes feature fusion challenging. In this study, we introduce a novel framework for multimodal semantic segmentation that effectively mitigates the aforementioned problems. Our approach employs a pseudo-siamese structure for feature extraction. Specifically, we propose a simple yet effective geometric topology structure modeling (GTSM) module to extract geometric relationships and texture information from optical data. Additionally, we present a modality intrinsic noise suppression (MINS) module to fully exploit radiation information and alleviate the effects of unique geometric distortions for SAR. Furthermore, we present an adaptive multimodal feature fusion (AMFF) module for fully fusing different modality features. Extensive experiments on both WHU-OPT-SAR and DFC23 datasets validate the robustness and effectiveness of the proposed Modality Characteristics-Guided Semantic Segmentation (MoCG) network compared to other state-of-the-art semantic segmentation methods, including multimodal and single-modal approaches. Our approach achieves the best performance on both datasets, resulting in mIoU/OA gains 69.1%/87.5% on WHU-OPT-SAR and 86.7%/97.3% on DFC23. Sining Xiao, Peijin Wang, Wenhui Diao, Xuee Rong, Xuexue Li, Kun Fu 0001, Xian Sun 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Optimal Partition Assignment for Universal Object DetectionabstractThe label assignment problem is a core task in object detection, which mainly focuses on how to define the$positive/negative$samples during the training phase. Recent works have proved that label assignment is significant for performance improvement of the detector. In this article, we propose an exquisite strategy that can dynamically assign labels according samples' joint scores (classification and location). Moreover, our strategy can apply to both 2D and 3D monocular detectors. In our strategy, we formulate label assignment as an optimization problem. Concretely, we first calculate the classification and location costs of each sample, which are treated as points in a 2-D coordinate system. Then an optimal divider line that minimizes the sum of point-to-line distances is designed to separate the$positive/negative$samples. An iterative Genetic Algorithm is employed in acquiring the optimal solution. Furthermore, a GIoU auxiliary branch is devised to keep sample selection consistent during the training and testing phase. Benefitting from the non-maximum suppression (NMS) that utilizes the joint scores of classification and location, excellent detection performance is achieved. Extensive experiments conducted on MS COCO, PASCAL VOC (2D object detection), and KITTI (3D object detection) verify the effectiveness and universality of our proposed Optimal Partition Assignment (OPA). Xian Sun 0001, Wenhui Diao, Xuee Rong, Shiyao Yan, Dongshuo Yin |
IEEE Trans. Multim. | 4 |
| 2022 | Historical Information-Guided Class-Incremental Semantic Segmentation in Remote Sensing ImagesabstractDespite the extraordinary success of the deep architectures on semantic segmentation for remote sensing (RS) images, they have difficulties in learning new classes from a sequential data stream because of catastrophic forgetting. Continual learning for semantic segmentation (CSS) is an emerging trend for its capability to cope with the above problems effectively. However, old classes from previous steps are collapsed into the background, which further aggravates the challenge of CSS in the RS scene. In this article, we revisit the knowledge distillation (KD) strategy and the characteristics of class-incremental semantic segmentation (CISS) and then present a generalized and effective framework to learn new classes while preserving knowledge of the learned classes. In particular, we propose two novel historical information-guided modules: the feature global perception module and the label reconstruction (LR) module. The former enables the current model to pay more attention to the region related to the old categories identified by the historical information when learning new classes. Meanwhile, the latter retrieves pixels belonging to the learned classes from the background to handle the background shift problem and maintain the high performance of old classes. We have conducted comprehensive experiments on two RS semantic segmentation datasets of Instance Segmentation in Aerial Images Dataset (iSAID) and Gao Fen (GF) challenge semantic segmentation dataset (GCSS). The experimental results outperform the current state-of-the-art methods in most incremental settings, which demonstrates the effectiveness of the proposed framework. Xuee Rong, Xian Sun 0001, Wenhui Diao, Peijin Wang |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Dynamic Interactive Learning for Lightweight Detectors in Remote Sensing ImageryabstractThe lightweight model has played an important role in the remote sensing (RS) realm. The existing researchers have proposed many models with lightweight structures, but their performance still has a gap compared with the deep model. A promising approach to optimize the lightweight model is knowledge distillation (KD), which can be viewed as knowledge transfer from the teacher model. However, the existing KD approaches have some issues. On one hand, offline distillation methods usually ignore the interactive learning between the student model and the teacher model. On the other hand, knowledge transfer does not consider instance property. This offline distillation strategy without property perception may not suitable for multiscale, diverse, and complex RS instances and results in a suboptimum training status. In this article, we propose a dynamic interactive learning (DIL) framework for optimizing RS lightweight detectors. First, we propose an instance interaction learning module. It calculates the value of every instance in the batch of the teacher and student prediction by each model’s real-time state and instance property. Then according to the DIL thought, we facilitate the low-quality instance to learn from the high-quality one whether it is from the teacher or student model. Moreover, we also propose the instance property perception (IPP) strategy that weighs the distillation knowledge of instances according to their feature, category, and location property. In the proposed DIL framework, both the teacher and student models are trained together and it is cost-free in the testing phase. Extensive experiments on three RS datasets demonstrate the effectiveness of the DIL. Wenhui Diao, Xuee Rong, Xian Sun 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | A Lightweight Multi-Scale Crossmodal Text-Image Retrieval Method in Remote SensingabstractRemote sensing (RS) crossmodal text-image retrieval has become a research hotspot in recent years for its application in semantic localization. However, since multiple inferences on slices are demanded in semantic localization, designing a crossmodal retrieval model with less computation but well performance becomes an emergent and challenging task. In this article, considering the characteristics of multi-scale and target redundancy in RS, a concise but effective crossmodal retrieval model (LW-MCR) is designed. The proposed model incorporates multi-scale information and dynamically filters out redundant features when encoding RS image, while text features are obtained via lightweight group convolution. To improve the retrieval performance of LW-MCR, we come up with a novel hidden supervised optimization method based on knowledge distillation. This method enables the proposed model to acquire dark knowledge of the multi-level layers and representation layers in the teacher network, which significantly improves the accuracy of our lightweight model. Finally, on the basis of contrast learning, we present a method employing unlabeled data to boost the performance of RS retrieval model further. The experiment results on four RS image-text datasets demonstrate the efficiency of LW-MCR in RS crossmodal retrieval (RSCR) tasks. We have released some codes of the semantic localization and made it open to access athttps://github.com/xiaoyuan1996/retrievalSystem. Wenkai Zhang 0002, Xuee Rong, Kun Fu 0001, Xian Sun 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Remote Sensing Cross-Modal Text-Image Retrieval Based on Global and Local InformationabstractCross-modal remote sensing text-image retrieval (RSCTIR) has recently become an urgent research hotspot due to its ability of enabling fast and flexible information extraction on remote sensing (RS) images. However, current RSCTIR methods mainly focus on global features of RS images, which leads to the neglect of local features that reflect target relationships and saliency. In this article, we first propose a novel RSCTIR framework based on global and local information (GaLR), and design a multi-level information dynamic fusion (MIDF) module to efficaciously integrate features of different levels. MIDF leverages local information to correct global information, utilizes global information to supplement local information, and uses the dynamic addition of the two to generate prominent visual representation. To alleviate the pressure of the redundant targets on the graph convolution network (GCN) and to improve the model’s attention on salient instances during modeling local features, the denoised representation matrix and the enhanced adjacency matrix (DREA) are devised to assist GCN in producing superior local representations. DREA not only filters out redundant features with high similarity, but also obtains more powerful local features by enhancing the features of prominent objects. Finally, to make full use of the information in the similarity matrix during inference, we come up with a plug-and-play multivariate rerank (MR) algorithm. The algorithm utilizes the$k$nearest neighbors of the retrieval results to perform a reverse search, and improves the performance by combining multiple components of bidirectional retrieval. Extensive experiments on public datasets strongly demonstrate the state-of-the-art performance of GaLR methods on the RSCTIR task. The code of GaLR method, MR algorithm, and corresponding files have been made available at:https://github.com/xiaoyuan1996/GaLR. Wenkai Zhang 0002, Changyuan Tian 0001, Xuee Rong, Zhengyuan Zhang 0003, Kun Fu 0001, Xian Sun 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Weakly Supervised Semantic Segmentation in Aerial Imagery via Explicit Pixel-Level ConstraintsabstractIn recent years, image-level weakly supervised semantic segmentation (WSSS) has developed rapidly in natural scenes due to the easy availability of classification tags. However, limited to complex backgrounds, multi-category scenes, and dense small targets in remote sensing (RS) images, relatively little research has been conducted in this field. To alleviate the impact of the above problems in RS scenes, a self-supervised Siamese network based on an explicit pixel-level constraints framework is proposed, which greatly improves the quality of class activation maps and the positioning accuracy in multi-category RS scenes. Specifically, there are three novel devices in this paper to promote performance to a new level: (a) A pixel-soft classification loss is proposed, which realizes explicit constraints on pixels during the image-level training; (b) A pixel global awareness module, which captures high-level semantic context and low-level pixel spatial information, is constructed to improve the consistency and accuracy of RS object segmentation; (c) A dynamic multi-scale fusion module with a gating mechanism is devised, which enhances feature representation and improves the positioning accuracy of RS objects, particularly on small and dense objects. Experiments on two RS challenge datasets demonstrate that these proposed modules achieve new state-of-the-art results by only using image-level labels, which improve mIoU to 36.79% on iSAID and 45.43% on ISPRS in the WSSS task. To the best of our knowledge, this is the first work to perform image-level WSSS on multi-class RS scenes. Ruixue Zhou, Wenkai Zhang 0002, Xuee Rong, Wenjie Liu 0016, Kun Fu 0001, Xian Sun 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |