Xinyang Pu

dblp:352/1403 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
8since 2021 · last 2025
0009-0002-0627-4603ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 8 · 2 first-author · 8 since 2021
YearPublicationVenuePosition
2025 Low-Rank Adaption on Transformer-Based Oriented Object Detector for Satellite Onboard Processing of Remote Sensing Images
abstract
Deep learning models in satellite onboard enable real-time interpretation of remote sensing images, reducing the need for data transmission to the ground and conserving communication resources. As satellite numbers and observation frequencies increase, the demand for satellite onboard real-time image interpretation grows, highlighting the expanding importance and development of this technology. However, updating the extensive parameters of models deployed on the satellites for spaceborne object detection model is challenging due to the limitations of uplink bandwidth in wireless satellite communications. To address this issue, this paper proposes a method based on parameter-efficient fine-tuning technology with low-rank adaptation module. It involves training low-rank matrix parameters and integrating them with the original model’s weight matrix through multiplication and summation, thereby fine-tuning the model parameters to adapt to new data distributions with minimal weight updates. The proposed method combines parameter-efficient fine-tuning with full fine-tuning in the parameter update strategy of the oriented object detection algorithm architecture. This strategy enables model performance improvements close to full fine-tuning effects with minimal parameter updates. In addition, low rank approximation is conducted to explore intrinsic dimensions of parameter matrices in normal-size models. Extensive experiments conducted on the DOTAv1.0, HRSC2016, and DIOR-R datasets verify the effectiveness of the proposed method. By fine-tuning and updating only 12.4% of the model’s total parameters, it is able to achieve 97% to 100% of the performance of full fine-tuning models. Additionally, the reduced number of trainable parameters accelerates model training iterations and enhances the generalization and robustness of the oriented object detection model. The source code is available at: https://github.com/fudanxu/LoRA-Det.
Xinyang Pu, Feng Xu 0001
IEEE Trans. Geosci. Remote. Sens.1
2024 A Siamese Network-Based Similarity Measure for Target in SAR Images
abstract
Synthetic Aperture Radar (SAR) has been widely applied due to its capability for continuous monitoring under all-weather and all-time conditions. Unlike optical images, targets in SAR images are highly sensitive to observation angles whose shapes and shadows change significantly under different viewing angles. Some current methods tackle the limited angle coverage problem in SAR targets through multi-view sample generation. However, for SAR image similarity assessment, existing techniques mainly concentrate on category and scene similarity, with insufficient attention to viewing angle similarity. This paper proposes a novel method for evaluating target viewing angle similarity in SAR images. Specifically, the method introduces a Siamese network to map SAR target features from high-dimensional space to a low-dimensional manifold. A contrastive loss function is constructed using Euclidean distance to measure the similarity of low-dimensional features of viewing angles. Additionally, a sample angle pairing method is designed for model training. Furthermore, a recognition module is incorporated into the model to assess the similarity of categories. Experiments based on the MSTAR dataset demonstrate the efficiency of the proposed method in evaluating SAR target viewing angle and category similarity.
Linghao Zheng, Xinyang Pu, Feng Xu 0001
IGARSS4
2024 Conditional Diffusion for SAR to Optical Image Translation
abstract
Synthetic aperture radar (SAR) offers all-weather and all-day high-resolution imaging, yet its unique imaging mechanism often necessitates expert interpretation, limiting its broader applicability. Addressing this challenge, this letter proposes a generative model that bridges SAR and optical imaging, facilitating the conversion of SAR images into more human-recognizable optical aerial images. This assists in the interpretation of SAR data, making it easier to recognize. Specifically, our model backbone is based on the recent diffusion models, which have powerful generative capabilities. We have innovatively tailored the diffusion model framework, incorporating SAR images as conditional constraints in the sampling process. This adaptation enables the effective translation from SAR to optical images. We conduct experiments on the satellite GF3 and SEN12 datasets and use structural similarity (SSIM) and Fréchet inception distance (FID) for quantitative evaluation. The results show that our model not only surpasses previous methods in quantitative evaluation but also significantly improves the visual quality of the generated images. This advancement underscores the model’s potential to enhance SAR image interpretation.
Xinyu Bai, Xinyang Pu, Feng Xu 0001
IEEE Geosci. Remote. Sens. Lett.2
2024 A Fast Progressive Ship Detection Method for Very Large Full-Scene SAR Images
abstract
Synthetic aperture radar (SAR) has emerged as a vital tool for ship monitoring due to its all-weather, all-day high-resolution imaging capabilities. In practical operations, the wide coverage and sparse ship distribution in very large full-scene SAR images pose challenges in terms of low efficiency and high false alarm rates. Traditional methods perform poorly in complex scenarios, while deep learning (DL) methods have high computation cost. This study proposes a fast progressive detection algorithm for ship targets in large SAR images, combining the advantages of traditional Non-DL methods and DL approaches. First, at a global scale, image preprocessing operations based on traditional methods are designed to quickly extract candidate regions. Then, at regional scale, an oriented ship detector is designed for refined ship detection within candidate regions. Finally, at individual-target scale, a false alarm discrimination network is constructed to further remove false alarms. Experimental results on GF-3 full-scene SAR images demonstrate that the proposed method can achieve minutes-level detection efficiency in images of billion-pixel-level size, while achieving high detection accuracy.
Hecheng Jia, Xinyang Pu, Qiaoyu Liu, Haipeng Wang 0002, Feng Xu 0001
IEEE Trans. Geosci. Remote. Sens.2
2024 ARNet: Prior Knowledge Reasoning Network for Aircraft Detection in Remote-Sensing Images
abstract
Amidst the landscape of contemporary remote sensing technology, the endeavor to detect and recognize aircraft within remote sensing images (RSIs) assumes pivotal strategic and practical significance. The complex nature of fine-grained aircraft recognition is a result of the intricate interplay between aircraft and their background environments, alongside category imbalance, which collectively lead to the emergence of a long-tail distribution within the dataset. However, experts proficient in RSIs interpretation can effectively address these challenges through the application of prior knowledge. This paper introduces the Aircraft Reasoning Network (ARNet), a framework tailored for aircraft detection and fine-grained recognition in RSIs, building upon prior knowledge employed in expert interpretation. Specifically, the Knowledge Reasoning Module (KRM) introduces a knowledge graph that incorporates both common and expert knowledge into the end-to-end network. Additionally, the network encompasses a Spatial Context Module (SCM) and an Airport Facility Relationship Module (AFRM). These components facilitate highly accurate detection and recognition of fine-grained aircraft in diverse environmental contexts by employing adaptive prior knowledge reasoning and optimizing target spatial location. Furthermore, an independent Aircraft Component Discrimination Module (ACDM) distinguishes aircraft based on their predominant component features, contributing to improved classification performance in both the few-shot and easily confused categories. Moreover, this paper introduces the AR-RSI dataset, a compilation of RSIs capturing fine-grained aircraft targets from diverse locations. The effectiveness and superiority of ARNet are exemplified on AR-RSI, achieving a minimum of 3.7 percentage higher mAP than the mainstream aircraft detection framework.
Yutong Qian, Xinyang Pu, Hecheng Jia, Haipeng Wang 0002, Feng Xu 0001
IEEE Trans. Geosci. Remote. Sens.2
2023 Cross-Domain SAR Ship Detection in Strong Interference Environment Based On Image-to-Image Translation
abstract
The model performance of object detection task may dramatically deteriorate when meeting the new dataset with discrepant data distribution compared with trained images. Especially for Synthetic Aperture Radar (SAR) images, the complicated imaging mechanism and diverse environments probably induce intense changes in image appearance and hurt the detection capability and robustness of models based on deep learning. In this paper, a method of learning strong interference characteristics of SAR images is proposed and conducted to generate artificial SAR images as extra training samples in the downstream task -- object detection to improve the detection accuracy and decrease the missing rate of models. Our approach utilized as a data augmentation strategy without annotation cost is confirmed to be efficacious and reliable by multiple experiments.
Xinyang Pu, Hecheng Jia, Feng Xu 0001
IGARSS1
2023 Tropical Cyclone Intensity Prediction by Spectral-Temporal Dislocation and Attention-Based Networks
abstract
Accurate prediction of Tropical cyclone (TC) intensity using multispectral images (MSIs) is critical to avoid economic loss and life casualty. Although existing methods have achieved good prediction results, they neglect changes in cloud patterns such as cyclone eyes and cloud spirals, which are closely related to TC intensity. How to leverage temporal-spatial-spectral features of MSIs to improve prediction accuracy is challenging task. In this paper, we propose a novel framework with Spectral-Temporal Dislocation and Attention-Based Networks (STD-AN) to predict MSW speed values near cyclone centers. The STD technique allows the framework to learn temporal-spatial-spectral features of TC. Meanwhile, the Self-Attention Modules (SAM) enable global attention feature extraction and Cross-Attention Modules (CAM) fuse different band features to improve prediction accuracy. Experimental results show that the proposed framework outperforms several state-of-the-art methods for TC intensity prediction.
Yahui Xiu, Xinyang Pu, Haixia Bi, Feng Xu 0001
IGARSS3
2023 Extension of Differentiable SAR Renderer for Ground Target Reconstruction From Multiview Images and Shadows
abstract
Three-dimensional (3D) reconstruction of complex targets on the ground from multi-view synthetic aperture radar (SAR) images is of great interests. The inherently-integrated forward-inverse architecture of the differentiable SAR renderer (DSR) provides a promising solution to the general inverse problem of SAR target reconstruction. In this context, the target’s shadow provides complementary information to its scattering image. Hence, this paper proposes a novel DSR-based target reconstruction approach using both the target image and its shadows. The capabilities of DSR are extended to generate not only target scattering images but also shadows. Furthermore, the gradients of the outputs, specifically illumination map and shadow map, with respect to the inputs, i.e., target geometry represented as a mesh, are derived. This enables us to develop a gradient-descent inverse approach for solving the general reconstruction problem. Extensive simulations and quantitative evaluations demonstrate that incorporating both the target scattering image and its shadows significantly improves the reconstruction performance. Moreover, our analyses indicate that achieving optimal reconstruction effects requires a minimum of 9 views with a relatively even distribution. Finally, the proposed algorithm is validated using real SAR images of vehicle targets.
Shilei Fu, Hecheng Jia, Xinyang Pu, Feng Xu 0001
IEEE Trans. Geosci. Remote. Sens.3