Xuan Zeng 0004

dblp:58/5418-4 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
7since 2021 · last 2024
0000-0001-5047-3488ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 7 · 3 first-author · 7 since 2021
YearPublicationVenuePosition
2024 RingMo-Lite: A Remote Sensing Lightweight Network With CNN-Transformer Hybrid Framework
abstract
In recent years, remote sensing (RS) vision foundation models such as RingMo have emerged and achieved excellent performance in various downstream tasks. However, the high demand for computing resources limits the application of these models on edge devices. It is necessary to design a more lightweight foundation model to support on-orbit RS image interpretation. Existing methods face challenges in achieving lightweight solutions while retaining generalization in RS image interpretation. This is due to the complex high and low-frequency spectral components in RS images, which make traditional single CNN or Vision Transformer methods unsuitable for the task. Therefore, this paper proposes RingMo-lite, a RS lightweight network with a CNN-Transformer hybrid framework, which effectively exploits the frequency-domain properties of RS to optimize the interpretation process on several tasks like classification, object detection, semantic segmentation, and change detection. It is combined by the Transformer module as a low-pass filter to extract global features of RS images through a dual-branch structure, and the CNN module as a stacked high-pass filter to extract fine-grained details effectively. Furthermore, a novelty-designed frequency-domain masked image modeling (FD-MIM) is employed during the pretraining stage for self-supervised learning, which combines the high-frequency and low-frequency characteristics of each image patch. This approach effectively captures the latent feature representation in RS data. As shown in Fig. 1, compared with RingMo, the proposed RingMo-lite reduces the parameters over 60% in various RS image interpretation tasks, the average accuracy drops by less than 2% in most of the scenes and achieves SOTA performance compared to models of the similar size. In addition, our work will be integrated into the MindSpore computing platform in the near future.
Yuelei Wang, Liangjin Zhao, Zhechao Wang, Ziqing Niu, Peirui Cheng, Kaiqiang Chen, Xuan Zeng 0004, Zhirui Wang 0003, Xian Sun 0001
IEEE Trans. Geosci. Remote. Sens.9
2024 SemiPSCN: Polarization Semantic Constraint Network for Semi-Supervised Segmentation in Large-Scale and Complex-Valued PolSAR Images
abstract
Since polarimetric synthetic aperture radar (PolSAR) terrain segmentation is a dense prediction task, the disadvantage of inadequate labeled samples greatly limits its performance. In this article, we present a semi-supervised segmentation network called SemiPSCN to reduce the data reliance on label annotation, which integrates semi-supervised learning (SSL) paradigm and the characteristics of PolSAR data into a unified architecture. First, considering the unreliability of pseudolabels caused by noise interference in PolSAR data, a pseudolabel error localization (PEL) module is designed. By mapping the pixels that have mispredictions in pseudolabels, PEL can greatly enhance the confidence of pseudolabels. Then, SemiPSCN introduces a category representation constraint (CRC) module to explicitly boost the category consistency between labeled and unlabeled PolSAR data. Via explicit intracategory and intercategory constraints, CRC can guarantee the invariant representations on the same category region between labeled and unlabeled data. Furthermore, a region consistency constraint (RCC) module is designed to enhance the regional consistency in PolSAR data. RCC leverages the conception of graph to model the understanding of spatial relationships among terrain targets, thereby facilitating consistent spatial region expression in semi-supervised process. Finally, we build a challenging large-scale dataset called LSPolSAR-Seg and conduct abundant experiments on LSPolSAR-Seg. SemiPSCN exhibits superior performance when compared with other advanced approaches, especially improving mean intersection over union (mIoU) by 3.44%–12.77% under 20% split setting, which promotes the performance to a state-of-the-art level.
Xuan Zeng 0004, Zhirui Wang 0003, Yuelei Wang, Xuee Rong, Pengyu Guo, Xian Sun 0001
IEEE Trans. Geosci. Remote. Sens.1
2023 DPFF-Net: Dual-Polarization Image Feature Fusion Network for SAR Ship Detection
abstract
Intelligent ship detection algorithms for synthetic aperture radar (SAR) images have achieved significant results in Earth observation applications. By learning features such as scale, shape and texture from samples, they can quickly locate and recognize ships in complex backgrounds. However, due to the lack of use of polarization features, the upper bound of detection performance is still limited, especially under poor image quality conditions such as ambiguous interference. To solve this, the dual-polarization image feature fusion network (DPFF-Net) is proposed. The key of it lies in adaptive mining, enhancement and fusion of polarization features through the designed siamese structure, polarization-aware enhancement block (PAEB) and dynamic gated fusion block (DGFB). With fully utilizing complementary information hidden between co-polarization and cross-polarization data, more comprehensive and accurate features are obtained and used as the detect head input. Thus, the proposed algorithm achieves state-of-the-art performance, and its effectiveness are validated by experiments on dual-polarization SAR datasets.
Jinyue Chen, Youming Wu, Xuan Zeng 0004, Wenhui Diao, Xian Sun 0001
IEEE Trans. Geosci. Remote. Sens.5
2023 MiCro: Modeling Cross-Image Semantic Relationship Dependencies for Class-Incremental Semantic Segmentation in Remote Sensing Images
abstract
Continual learning is an effective way to overcome catastrophic forgetting (CF) in incremental learning for semantic segmentation. The existing continual semantic segmentation (CSS) methods of remote sensing (RS) ignore the semantic relationships among pixels across different images, which will lead to disappointing segmentation results, such as edge pixel misclassification and small object omission. In this paper, we propose a framework for modeling cross-image semantic relationship dependencies (MiCro), which aims to learn an inter-class separable and intra-class cohesive feature space from the pixel relationships across various images to ensure that learned categories can prevent CF in the incremental process. Specifically, we exploit the relationships among pixels of images in mini-batch to construct three losses: (a) Cross-image feature relationship distillation (CFRD) loss, which builds a well-structured feature space; (b) Cross-image intra-class feature cohesion (CIFC) loss, which is devised to make intra-class features more cohesive; and (c) Cross-image class-area weighted cross-entropy (CCWCE) loss, which is mainly employed to inversely weight the proportion of category area in mini-batch. The effectiveness of the proposed approach is demonstrated by extensive experiments on three RS semantic segmentation datasets from ISPRS Vaihingen, ISPRS Potsdam, and iSAID. MiCro is superior to the current most advanced methods in most incremental settings, especially improving mIoU by 11.59% on ISPRS Vaihingen, 13.17% on ISPRS Potsdam, and 15.01% on iSAID in the most difficult incremental settings, which promotes the CSS to a state-of-the-art (SOTA) level. The code will be available at https://github.com/RongXueE/MiCro.
Xuee Rong, Peijin Wang, Wenhui Diao, Wenxin Yin, Xuan Zeng 0004, Xian Sun 0001
IEEE Trans. Geosci. Remote. Sens.6
2023 DCM: A Distributed Collaborative Training Method for the Remote Sensing Image Classification
abstract
As the number of aero and space remote sensing platforms increases, distributed observation and real-time terminal processing become mainstream in the future. However, most of the training methods for the multi-platform are still limited to centralized structures or independent training based on a single platform, which is inefficient or limited in accuracy. In order to solve this problem, we innovatively propose a distributed collaborative method (DCM) for remote sensing image classification training in this article. First, the proposed training method, which is based on one cloud and several terminals, can aggregate different parameters of the terminal network to the cloud to improve global accuracy. Second, a sample proximity network is designed to process the problem of data heterogeneity on different terminal networks, which further improves the accuracy during the model fusion on the cloud. Third, a multi-layer grouped concatenation module is applied after the model fusion to extract hierarchical features with different categories of remote sensing images. Experimental results on the challenging remote sensing image classification dataset FAIR1M show that the proposed training method has better collaborative learning ability than the centralized-based model or terminal-trained lightweight network under the heterogeneous data.
Yuelei Wang, Zhirui Wang 0003, Peirui Cheng, Xuan Zeng 0004, Xian Sun 0001, Kun Fu 0001
IEEE Trans. Geosci. Remote. Sens.4
2022 TS-SHES: Terrain Segmentation in Complex-Valued PolSAR Images Via Scattering Harmonization and Explicit Supervision
abstract
Convolutional neural network (CNN) has attracted extensive attention in the research field of polarimetric synthetic aperture radar (PolSAR) terrain segmentation. However, directly using CNN in PolSAR terrain segmentation while ignoring the characteristics of PolSAR images has become the main factor restricting the performance of algorithms. In this article, we propose an efficient PolSAR terrain segmentation algorithm called TS-SHES, which integrates the polarization scattering characteristics of PolSAR images and the CNN learning process into a unified architecture. First, considering the intrinsic structure of complex-valued PolSAR data, TS-SHES transforms the scattering matrix into the form of amplitude and phase components, which preserves the original information maximally. Then, TS-SHES introduces a scattering harmonized encoding method (SH-Enc) to balance the feature contributions of weak and strong scattering regions as well as map the two components into the same representation space. Through the above scattering harmonization operations, the segmentation performance of CNN on weak scattering regions can be improved, and the feature imbalance in amplitude and phase can be alleviated. Furthermore, in view of the implicit states of CNN feature construction, a scattering explicit learning network (SEL-Net) is presented to collect the scattering features of amplitude and phase. Via explicit supervision, SEL-Net avoids the incomplete collection of scattering information caused by implicit feature construction, thereby improving the segmentation accuracy. Abundant experiments are conducted on two PolSAR images acquired by the GaoFen-3 satellite, which demonstrates the superiority of our proposed algorithm.
Xuan Zeng 0004, Zhirui Wang 0003, Xian Sun 0001
IEEE Trans. Geosci. Remote. Sens.1
2022 DENet: Double-Encoder Network With Feature Refinement and Region Adaption for Terrain Segmentation in PolSAR Images
abstract
Recently, many studies exploit deep neural networks to promote terrain segmentation in polarimetric synthetic aperture radar (PolSAR) images. However, these works usually inherit the nature-scene approaches directly and may not be robust for the PolSAR image segmentation task. The main limitations include single-type feature construction, weak feature consistency, and geometry-agnostic collection of scattering information. In this article, we present the DENet, a double-encoder network with feature refinement and region adaption for the terrain segmentation in PolSAR images. First, a double-encoder architecture is proposed to leverage the multitype information of PolSAR images, which can provide more discriminative features than the previous methods using the single-type feature. Second, considering that the polarization information has strong consistency over the category-identical regions, a polarization-guided refinement module is proposed to maintain the feature consistency in the PolSAR segmentation model. This design alleviates the phenomenon of incomplete and fragmented segmentation results. Third, in view of the rich targets’ characteristics in the scattering information, a region-adaptive convolution module is developed to facilitate the scattering information collected over the geometry-irregular regions. This design can improve the segmentation accuracy on the geometry-irregular regions. Extensive experiments are conducted on six PolSAR images to verify the effectiveness of the DENet. Compared with the previous works, our method achieves competitive performance.
Xuan Zeng 0004, Zhirui Wang 0003, Xian Sun 0001, Zhonghan Chang
IEEE Trans. Geosci. Remote. Sens.1