EDBT 2026 Demo / reviewers in the wild / expert
Zhirui Wang 0003
dblp:87/1453-3
· DBLP profile ↗
38ranked-venue papers
2as first author
32since 2021 · last 2025
0000-0003-2877-0384ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 34 · 2 first-author · 28 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SemStereo: Semantic-Constrained Stereo Matching Network for Remote SensingabstractSemantic segmentation and 3D reconstruction are two fundamental tasks in remote sensing, typically treated as separate or loosely coupled tasks. Despite attempts to integrate them into a unified network, the constraints between the two heterogeneous tasks are not explicitly modeled, since the pioneering studies either utilize a loosely coupled parallel structure or engage in only implicit interactions, failing to capture the inherent connections. In this work, we explore the connections between the two tasks and propose a new network that imposes semantic constraints on the stereo matching task, both implicitly and explicitly. Implicitly, we transform the traditional parallel structure to a new cascade structure termed Semantic-Guided Cascade structure, where the deep features enriched with semantic information are utilized for the computation of initial disparity maps, enhancing semantic guidance. Explicitly, we propose a Semantic Selective Refinement (SSR) module and a Left-Right Semantic Consistency (LRSC) module. The SSR refines the initial disparity map under the guidance of the semantic map. The LRSC ensures semantic consistency between two views via reducing the semantic divergence after transforming the semantic map from one view to the other using the disparity map. Experiments on the US3D and WHU datasets demonstrate that our method achieves state-of-the-art performance for both semantic segmentation and stereo matching. Chen Chen 0036, Liangjin Zhao, Yuanchun He, Yingxuan Long, Kaiqiang Chen, Zhirui Wang 0003, Yanfeng Hu, Xian Sun 0001 |
AAAI | 6 |
| 2025 | SA-Occ: Satellite-Assisted 3D Occupancy Prediction in Real WorldabstractExisting vision-based 3D occupancy prediction methods are inherently limited in accuracy due to their exclusive reliance on street-view imagery, neglecting the potential benefits of incorporating satellite views. We propose SA-Occ, the first Satellite-Assisted 3D occupancy prediction model, which leverages GPS & IMU to integrate historical yet readily available satellite imagery into real-time applications, effectively mitigating limitations of ego-vehicle perceptions, involving occlusions and degraded performance in distant regions. To address the core challenges of cross-view perception, we propose: 1) Dynamic-Decoupling Fusion, which resolves inconsistencies in dynamic regions caused by the temporal asynchrony between satellite and street views; 2) 3D-Proj Guidance, a module that enhances 3D feature extraction from inherently 2D satellite imagery; and 3) Uniform Sampling Alignment, which aligns the sampling density between street and satellite views. Evaluated on Occ3D-nuScenes, SA-Occ achieves state-of-the-art performance, especially among single-frame methods, with a 39.05% mIoU (a 6.97% improvement), while incurring only 6.93 ms of additional latency per frame. Our code and newly curated dataset are available at https://github.com/chenchen235/SA-Occ. Chen Chen 0036, Zhirui Wang 0003, Taowei Sheng, Yundu Li, Peirui Cheng, Luning Zhang, Kaiqiang Chen, Yanfeng Hu, Xue Yang 0005, Xian Sun 0001 |
ICCV | 2 |
| 2025 | Physics-Guided Deep Learning 3-D Inversion Based on Magnetic DataabstractThe 3-D inversion of magnetic data based on deep learning has achieved great success. This method relies on neural networks to extract features from a large amount of data and then generate structures, with high prediction accuracy and fast speed. However, this data-driven inversion method lacks a closed-form analytical expression and is likened to a “black box.” This feature makes it lack theoretical guidance and has poor interpretability. In addition, the training data often cannot fully cover all aspects of the target, so data-driven methods face the challenge of weak generalization ability in practical applications. Therefore, this letter proposes to integrate physical knowledge into the inversion method based on deep learning. The architecture and loss function of the deep learning model are designed based on the forward modeling of the magnetic field. This can not only improve the interpretability of the algorithm and provide a deeper understanding of the model but also help to make up for the lack of generalization capabilities of deep learning methods. Xiaoqing Shi, Zhirui Wang 0003, Xue Lu, Peirui Cheng, Luning Zhang, Liangjin Zhao |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2025 | Physics-Guided Detector for SAR AirplanesabstractThe disperse structure distributions (discreteness) and variant scattering characteristics (variability) of SAR airplane targets lead to special challenges of object detection and recognition. The current deep learning-based detectors encounter challenges in distinguishing fine-grained SAR airplanes against complex backgrounds. To address it, we propose a novel physics-guided detector (PGD) learning paradigm for SAR airplanes that comprehensively investigate their discreteness and variability to improve the detection performance. It is a general learning paradigm that can be extended to different existing deep learning-based detectors with ”backbone-neck-head” architectures. The main contributions of PGD include the physics-guided self-supervised learning, feature enhancement, and instance perception, denoted as PGSSL, PGFE, and PGIP, respectively. PGSSL aims to construct a self-supervised learning task based on a wide range of SAR airplane targets that encodes the prior knowledge of various discrete structure distributions into the embedded space. Then, PGFE enhances the multi-scale feature representation of a detector, guided by the physics-aware information learned from PGSSL. PGIP is constructed at the detection head to learn the refined and dominant scattering point of each SAR airplane instance, thus alleviating the interference from the complex background. We propose two implementations, denoted as PGD and PGD-Lite, and apply them to various existing detectors with different backbones and detection heads. The experiments demonstrate the flexibility and effectiveness of the proposed PGD, which can improve existing detectors on SAR airplane detection with fine-grained classification task (an improvement of 3.1% mAP most), and achieve the state-of-the-art performance (90.7% mAP) on SAR-AIRcraft-1.0 dataset. The project is open-source at https://github.com/XAI4SAR/PGD. Zhongling Huang, Shuxin Yang, Zhirui Wang 0003, Gong Cheng 0003, Junwei Han 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | UCDNet: Multi-UAV Collaborative 3-D Object Detection Network by Reliable Feature MappingabstractMulti-unmanned aerial vehicle (UAV) collaborative 3-D object detection can comprehend complex environments by integrating complementary information, with applications encompassing traffic monitoring, delivery services, and agricultural management. However, the extremely broad observations in aerial remote sensing and significant perspective differences across multiple UAVs make it challenging to achieve precise and consistent feature mapping from 2-D images to 3-D space in multi-UAV collaborative 3-D object detection paradigm. To address the problem, we propose an unparalleled camera-based multi-UAV collaborative 3-D object detection paradigm called UCDNet. Specifically, the depth information from the UAVs to the ground is explicitly utilized as a strong prior to provide a reference for more accurate and generalizable feature mapping. Additionally, we design a homologous point geometric consistency loss as an auxiliary self-supervision, which directly influences the feature mapping module, thereby strengthening the global consistency of multiview perception. Experiments on AeroCollab3D and CoPerception-UAVs datasets show that our method increases 4.7% and 10% mean Average Precision (mAP) respectively compared to the baseline, which demonstrates the superiority of UCDNet. Pengju Tian, Zhirui Wang 0003, Peirui Cheng, Zhechao Wang, Liangjin Zhao, Menglong Yan, Xue Yang 0005, Xian Sun 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | AVCPNet: An AAV-Vehicle Collaborative Perception Network for 3-D Object DetectionabstractWith the advancement of collaborative perception, the role of autonomous aerial vehicle (AAV)–vehicle collaborative perception has become increasingly significant. The demand for collaborative perception from various perspectives to construct comprehensive perceptual information is rising. However, challenges emerge due to differences in the field of view (FOV) between cross-domain agents and their varying sensitivities to image information. Furthermore, accurate depth information is essential for collaboration to transform image features into bird’s eye view (BEV) features. To address these challenges, we propose a framework specifically designed for aerial-ground collaboration. First, to address the deficiency of datasets for aerial-ground collaboration, we have developed a virtual dataset named V2U-COO for our research. Second, we design a cross-domain cross-adaptation (CDCA) module to align the target information obtained from different domains, thereby achieving more accurate perception results. Finally, we introduce a collaborative depth optimization (CDO) module to obtain more precise depth estimation results, leading to more accurate perception results. We conduct extensive experiments on both our virtual dataset and a public dataset to validate the effectiveness of our framework. Our method resolves the feature fusion issue under significant height differences, a challenge that previous BEV generation methods struggled to address effectively. Our experiments on the V2U-COO and DAIR-V2X datasets demonstrate improvements in detection accuracy of 6.1% and 2.7%, respectively. Our code will be released athttps://github.com/wyccoo/uvcp. Zhirui Wang 0003, Peirui Cheng, Pengju Tian, Ziyang Yuan, Liangjin Zhao |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | RingMo-Galaxy: A Remote Sensing Distributed Foundation Model for Diverse Downstream TasksabstractRemote sensing lightweight foundation models have successfully achieved online perception, providing real-time intelligent interpretation. However, their capabilities are restricted to inferences solely based on their respective observations and models, thus lacking a comprehensive understanding of large-scale remote sensing scenarios. To address this limitation, we propose RingMo-Galaxy, a remote sensing distributed foundation model based on generalized information mapping and interaction. RingMo-Galaxy can realize online collaborative perception across multiple platforms and diverse downstream tasks by mapping observations into a unified space and implementing a task-agnostic information interaction strategy. Specifically, we leverage the ground-based geometric prior of remote sensing oblique observations to change feature mapping from absolute to relative depth estimation, thereby enhancing the model’s ability to extract generalized features across diverse heights and perspectives. In addition, we present a dual-branch information compression module to decouple high-frequency and low-frequency features, achieving feature-level compression while preserving critical task-agnostic details. To support our research, we collect a multitask simulation dataset named AirCo-MultiTasks, specifically designed for multi-unmanned aerial vehicle (UAV) collaborative observation. We also conduct extensive experiments, including 3-D object detection, instance segmentation, and trajectory prediction. The numerous results demonstrate that our proposed RingMo-Galaxy achieves state-of-the-art performance across various downstream tasks. Zhechao Wang, Zhirui Wang 0003, Peirui Cheng, Liangjin Zhao, Pengju Tian, Mingxin Chen, Xian Sun 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | TeCCo: A Terminal-Cloud Cross-Domain Collaborative Framework for Remote Sensing Image ClassificationabstractThe terminal-cloud collaborative framework boosts precision and efficiency by integrating cloud computing power with low-latency terminal responsiveness, offering a suitable solution for the growing demands of multiplatform remote sensing (RS) image interpretation. However, the significant differences in data distribution across various RS platforms present a great challenge in balancing the cloud’s centralized processing capabilities with the local interpretation abilities of different terminals. To address this challenge, we propose a terminal-cloud cross-domain collaborative (TeCCo) framework that inherits the efficiency advantages of multiple platforms while ensuring high-accuracy interpretation of diverse data distributions from different terminals. First, the dual classifier co-learning (DCCL) module is designed to enhance cloud robustness. By combining a multilayer perceptron for instance-level classification and a graph convolutional network (GCN) for feature-level aggregation, it achieves mutual supervision and improves feature alignment across different data distributions. Second, the hypernetwork personalization (HNP) module is introduced to generate personalized classifier parameters for each terminal with little fine-tuning cost, allowing terminals to maintain their uniqueness while benefiting from the generalization advantages of collaborative training. Finally, a data-assisted progressive inference mechanism is proposed to enhance accuracy by jointly clustering the features transmitted from terminals and the features of supervised data in the cloud. Extensive experiments demonstrate that TeCCo effectively addresses data distribution challenges, enhancing both the generalization of the cloud model and the personalization of terminal models, achieving state-of-the-art (SOTA) performance in cross-domain and multiplatform RS image classification. Peirui Cheng, Liangjin Zhao, Zhirui Wang 0003, Lingyu Kong, Guangluan Xu, Xian Sun 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | AI-Powered Flood MapathonabstractFloods represent a pervasive natural hazard with global ramifications, impacting a vast population and resulting in substantial property damage and severe mortality. Particularly worrisome is their disproportionate effect on the least developed countries, which exacerbates developmental imbalances, posing a significant obstacle to the attainment of the United Nations Sustainable Development Goals (UN SDGs). This paper introduces the AI-powered Flood Mapathon activity, co-organized by the Aerospace Information Research Institute under the Chinese Academy of Sciences, in partnership with GEOVIS Technology Co., Ltd., GEOVIS Earth Technology Co., Ltd., and IEEE GRSS IADF. The activity seeks to mobilize individuals worldwide to address the most prevalent natural hazard-floods by collaboratively mapping inundated regions through the analysis of satellite imagery. Gaining widespread attention, the activity has garnered 30,755 submissions from 310 participants across 34 countries. Through collective efforts, participants have curated a semantic segmentation dataset focusing on floods, incorporating annotations of pertinent features related to both floods and human activities. Additionally, the paper elucidates the custom crowdsourcing mapping system, which seamlessly integrates cutting-edge AI technologies to alleviate mapping complexities. The activity contributes to sustainability by drawing extensive public attention, creating a public flood dataset for academic research, and establishing an efficient and intelligent mapping system. Kaiqiang Chen, Xue Lu, Taowei Sheng, Zhirui Wang 0003, Xian Sun 0001, Ronny Hänsch |
IGARSS | 6 |
| 2024 | Drones Help Drones: A Collaborative Framework for Multi-Drone Object Trajectory Prediction and BeyondabstractCollaborative trajectory prediction can comprehensively forecast the future motion of objects through multi-view complementary information. However, it encounters two main challenges in multi-drone collaboration settings. The expansive aerial observations make it difficult to generate precise Bird's Eye View (BEV) representations. Besides, excessive interactions can not meet real-time prediction requirements within the constrained drone-based communication bandwidth. To address these problems, we propose a novel framework named "Drones Help Drones" (DHD). Firstly, we incorporate the ground priors provided by the drone's inclined observation to estimate the distance between objects and drones, leading to more precise BEV generation. Secondly, we design a selective mechanism based on the local feature discrepancy to prioritize the critical information contributing to prediction tasks during inter-drone interactions. Additionally, we create the first dataset for multi-drone collaborative prediction, named "Air-Co-Pred", and conduct quantitative and qualitative experiments to validate the effectiveness of our DHD framework. The results demonstrate that compared to state-of-the-art approaches, DHD reduces position deviation in BEV representations by over 20\% and requires only a quarter of the transmission ratio for interactions while achieving comparable prediction performance. Moreover, DHD also shows promising generalization to the collaborative 3D object detection in CoPerception-UAVs. Zhechao Wang, Peirui Cheng, Minxing Chen, Pengju Tian, Zhirui Wang 0003, Xue Yang 0005, Xian Sun 0001 |
NeurIPS | 5 |
| 2024 | FS-DCL: Distributed Collaborative Learning for Few-Shot Remote Sensing Image ClassificationabstractWith the development of on-orbit hardware and distributed multiplatform observation systems in satellite remote sensing (RS) scenario, on-orbit collaborative model updating has become a promising trend. Due to restrictions of imaging conditions and storage resources, on-orbit updating is usually carried out with limited samples. However, existing collaborative learning methods rarely consider the few-shot problem. To address this issue, this letter innovatively proposes a distributed collaborative learning method for few-shot RS image classification (FS-DCL), which encourages the collaboration between satellites with similar data distribution to supplement useful information for each satellite, and design on-orbit models to extract more discriminative features. Specifically, a personalized parameter aggregation strategy (PPAS) is proposed to generate personalized parameters for each satellite based on information from satellites with similar data distributions, providing information gain to alleviate problems of insufficient samples. Besides, a feature enhancement method (FEM) is applied to on-orbit models to enhance the feature representation and produce a more discriminative feature space, thus improving the accuracy of few-shot metric classification. Extensive experiments on two RS datasets demonstrate the superiority of FS-DCL. Peirui Cheng, Yuelei Wang, Zhirui Wang 0003, Kaiqiang Chen, Xian Sun 0001, Daobing Zhang |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2024 | MDCNet: A Multiplatform Distributed Collaborative Network for Object Detection in Remote Sensing ImageryabstractWith the recent development of remote sensing (RS) technology, the amount of RS platforms has witnessed a substantial increase, and the capacity of Earth observation has been greatly enhanced. The interpretation of RS images has also gradually evolved from traditional centralized ground processing to on-orbit processing. However, the traditional single-platform on-orbit processing is limited to a single source of information, which results in the underutilization of the advantages of multiplatform observation in the current RS field, and restricts the accuracy of inference tasks. To tackle the aforementioned problem, we propose a multiplatform distributed collaborative inference network, which can combine the intermediate features from multiple platforms to improve the accuracy of inference tasks. First, we proposed the collaboration map generator, which generates the collaboration map for optimal collaborator selection autonomously. Second, a spatial feature compression (SFC) module is designed to compress the interplatform transmission features, adapting spatially sparse distribution characteristics of RS objects. Finally, a feature fusion module containing spatial priors is proposed to fuse the features collected from multiple platforms to obtain more precise inference results. We conducted extensive experiments on three public datasets and verified the effectiveness of the proposed framework. On the NWPU VHR-10 dataset, for example, the proposed method improves the detection accuracy by 13.7% and 10.3% under two experimental settings compared with a single platform and compresses the intermediate data transmission between platforms by more than 80%. Shujing Duan, Peirui Cheng, Zhechao Wang, Zhirui Wang 0003, Kaiqiang Chen, Xian Sun 0001, Kun Fu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | SDL-MVS: View Space and Depth Deformable Learning Paradigm for Multiview Stereo Reconstruction in Remote SensingabstractResearch on multiview stereo (MVS) based on remote sensing images has promoted the development of large-scale urban 3-D reconstruction. However, remote sensing multiview image data suffer from the problems of occlusion and uneven brightness between views during acquisition, which leads to the problem of blurred details in depth estimation. To solve the above problem, we reexamine the deformable learning method in the MVS task and propose a novel paradigm based on view space and depth deformable learning (SDL-MVS), aiming to learn deformable interactions of features in different view spaces and deformably model the depth ranges and intervals to enable high accurate depth estimation. Specifically, to solve the problem of view noise caused by occlusion and uneven brightness, we propose a progressive space deformable sampling (PSS) mechanism, which performs deformable learning of sampling points in the 3-D frustum space and the 2-D image space in a progressive manner to embed source features to the reference feature adaptively. To further optimize the depth, we introduce depth hypothesis deformable discretization (DHD), which achieves precise positioning of the depth prior by adaptively adjusting the depth range hypothesis and performing deformable discretization of the depth interval hypothesis. Finally, our SDL-MVS achieves explicit modeling of occlusion and uneven brightness faced in MVS through the deformable learning paradigm of view space and depth, achieving accurate multiview depth estimation. Extensive experiments on LuoJia-MVS and WHU datasets show that our SDL-MVS reaches state-of-the-art performance. It is worth noting that our SDL-MVS achieves a mean absolute error (MAE) error of 0.086 and an accuracy of 98.9% for Acc$_{\lt 0.6\,\text {m}}$and 98.9% for Acc$_{\lt 3-\text {interval}}$on the LuoJia-MVS dataset under the premise of three views as input. Yongqiang Mao, Hanbo Bi, Liangyu Xu, Kaiqiang Chen, Zhirui Wang 0003, Xian Sun 0001, Kun Fu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | FCIL-MSN: A Federated Class-Incremental Learning Method for Multisatellite NetworksabstractMulti-satellite networks have become the prevalent mode for remote sensing intelligent interpretation, with the onboard models requiring class-incremental updates to accommodate the new categories emerging in evolving data and tasks. Traditional model updating methods, which involve uploading models separately after ground-based updating, are inefficient due to limited uplink bandwidth and cumbersome ground update processes while underutilizing potential computing resources on satellites. To address the aforementioned problems, this paper innovatively proposes a collaborative in-orbit incremental update method termed FCIL-MSN, which leverages observational information and computing resources from multi-satellite networks. Firstly, FCIL-MSN achieves collaborative onboard model updates by introducing federated class-incremental learning into multi-satellite networks. Secondly, a bias calibration-guided relationship distillation module constructs a pseudo-feature set by collaborative multi-satellite networks, which alleviates the model bias caused by class imbalance from a global perspective, thereby enhancing model performance. Finally, a gradient information aggregation module is designed to facilitate the exclusion of unfavorable local updates by measuring the contribution of each terminal, thereby accelerating the convergence while obtaining the global model. We conduct extensive experiments on two datasets for scene classification tasks to verify the effectiveness of our proposed method. Experimental results demonstrate that FCIL-MSN outperforms existing general FCIL methods, improving average classification accuracy by 1.45% and decreasing the performance degradation rate by 6.40%. Ziqing Niu, Peirui Cheng, Zhirui Wang 0003, Liangjin Zhao, Xian Sun 0001, Zhi Guo |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | SRT-Net: Scattering Region Topology Network for Oriented Ship Detection in Large-Scale SAR ImagesabstractSynthetic aperture radar (SAR) ship detection plays an important role in the field of maritime security. However, certain unique imaging properties make it challenging to extract the shape features of ships, such as speckle noise and strong scattering interference from irrelevant objects. These factors result in inaccurate ship localization and obvious false alarms under complex large-scale inshore scenes. To address this issue, we propose the scattering region topology network (SRT-Net), which can dynamically capture the comprehensive global context and enhance the ship saliency. This is achieved through two key modules, namely the scattering region topological structure pyramid (SRTP) and the ship saliency enhancement (SSE) module. The former provides richer semantic information to distinguish the object from the background, while the latter offers an extra pixel-level classification task to guide accurate bounding box regression. Thanks to the guidance of richer information, the proposed method can achieve fewer false alarms and enhance location accuracy. Additionally, we introduce a scale feature adaptive (SFA) loss to balance the attention to ships with various scales, which improves the robustness of multiscale ship detection. The proposed method achieves state-of-the-art performance under complex inshore scenes, and its effectiveness is verified by experiments on a large-scale SAR ship detection dataset (LSSDD) and a public SAR ship detection dataset (SSDD+). Dece Pan, Jiamei Fu, Zhirui Wang 0003, Xian Sun 0001, Youming Wu |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | RingMo-Lite: A Remote Sensing Lightweight Network With CNN-Transformer Hybrid FrameworkabstractIn recent years, remote sensing (RS) vision foundation models such as RingMo have emerged and achieved excellent performance in various downstream tasks. However, the high demand for computing resources limits the application of these models on edge devices. It is necessary to design a more lightweight foundation model to support on-orbit RS image interpretation. Existing methods face challenges in achieving lightweight solutions while retaining generalization in RS image interpretation. This is due to the complex high and low-frequency spectral components in RS images, which make traditional single CNN or Vision Transformer methods unsuitable for the task. Therefore, this paper proposes RingMo-lite, a RS lightweight network with a CNN-Transformer hybrid framework, which effectively exploits the frequency-domain properties of RS to optimize the interpretation process on several tasks like classification, object detection, semantic segmentation, and change detection. It is combined by the Transformer module as a low-pass filter to extract global features of RS images through a dual-branch structure, and the CNN module as a stacked high-pass filter to extract fine-grained details effectively. Furthermore, a novelty-designed frequency-domain masked image modeling (FD-MIM) is employed during the pretraining stage for self-supervised learning, which combines the high-frequency and low-frequency characteristics of each image patch. This approach effectively captures the latent feature representation in RS data. As shown in Fig. 1, compared with RingMo, the proposed RingMo-lite reduces the parameters over 60% in various RS image interpretation tasks, the average accuracy drops by less than 2% in most of the scenes and achieves SOTA performance compared to models of the similar size. In addition, our work will be integrated into the MindSpore computing platform in the near future. Yuelei Wang, Liangjin Zhao, Zhechao Wang, Ziqing Niu, Peirui Cheng, Kaiqiang Chen, Xuan Zeng 0004, Zhirui Wang 0003, Xian Sun 0001 |
IEEE Trans. Geosci. Remote. Sens. | 10 |
| 2024 | SemiPSCN: Polarization Semantic Constraint Network for Semi-Supervised Segmentation in Large-Scale and Complex-Valued PolSAR ImagesabstractSince polarimetric synthetic aperture radar (PolSAR) terrain segmentation is a dense prediction task, the disadvantage of inadequate labeled samples greatly limits its performance. In this article, we present a semi-supervised segmentation network called SemiPSCN to reduce the data reliance on label annotation, which integrates semi-supervised learning (SSL) paradigm and the characteristics of PolSAR data into a unified architecture. First, considering the unreliability of pseudolabels caused by noise interference in PolSAR data, a pseudolabel error localization (PEL) module is designed. By mapping the pixels that have mispredictions in pseudolabels, PEL can greatly enhance the confidence of pseudolabels. Then, SemiPSCN introduces a category representation constraint (CRC) module to explicitly boost the category consistency between labeled and unlabeled PolSAR data. Via explicit intracategory and intercategory constraints, CRC can guarantee the invariant representations on the same category region between labeled and unlabeled data. Furthermore, a region consistency constraint (RCC) module is designed to enhance the regional consistency in PolSAR data. RCC leverages the conception of graph to model the understanding of spatial relationships among terrain targets, thereby facilitating consistent spatial region expression in semi-supervised process. Finally, we build a challenging large-scale dataset called LSPolSAR-Seg and conduct abundant experiments on LSPolSAR-Seg. SemiPSCN exhibits superior performance when compared with other advanced approaches, especially improving mean intersection over union (mIoU) by 3.44%–12.77% under 20% split setting, which promotes the performance to a state-of-the-art level. Xuan Zeng 0004, Zhirui Wang 0003, Yuelei Wang, Xuee Rong, Pengyu Guo, Xian Sun 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | ST-Net: Scattering Topology Network for Aircraft Classification in High-Resolution SAR ImagesabstractAircraft classification in synthetic aperture radar (SAR) images plays a considerable role in global region management and surveillance. Recently, deep learning has been applied to solve the classification problem and made significant progress. Due to the imaging variability at different angles and component scattering discreteness in SAR images, previous works have had difficulty in achieving desirable classification results. To address these issues, we study the positional and semantic relationship between the scattering points and propose an innovative scattering topology network (ST-Net) in this article. First, considering the diversity of imaging results caused by different target attitude angles, we extract and transform the scattering cluster centers to update the information of various categories. It can guide the model to strengthen the discriminative features and mitigate the impact of imaging variability on classification performance. Second, a novel scattering topology module (STM) is introduced to model the spatial relationships and semantic information interaction of discrete scattering points. In this process, the topology relations and scattering characteristics are enhanced for further accurate classification. Third, context attention excitation (CAE) is designed to capture significant global and semantic information, which is conducive to suppressing background interference and reducing category confusion. In conclusion, the ST-Net is presented with the SAR imaging mechanism and the topology geometric representation of aircraft. We construct the SAR aircraft category dataset (SAR-ACD) and conduct extensive experiments on it to show the effectiveness of ST-Net, which illustrates that our method achieves superior classification performance. Yuzhuo Kang, Zhirui Wang 0003, Haoyu Zuo, Yidan Zhang 0002, Zhujun Yang, Xian Sun 0001, Kun Fu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Elevation Estimation-Driven Building 3-D Reconstruction From Single-View Remote Sensing ImageryabstractBuilding 3D reconstruction from remote sensing images has a wide range of applications in smart cities, photogrammetry and other fields. Methods for automatic 3D urban building modeling typically employ multi-view images as input to algorithms to recover point clouds and 3D models of buildings. However, such models rely heavily on multi-view images of buildings, which are time-intensive and limit the applicability and practicality of the models. To solve these issues, we focus on designing an efficient DSM estimation-driven reconstruction framework (Building3D), which aims to reconstruct 3D building models from the input single-view remote sensing image. Existing DSM estimation networks suffer from the imbalance between local features and global features, which leads to over-smooth DSM estimates at instance boundaries. To address this issue, we propose a Semantic Flow Field-guided DSM Estimation (SFFDE) network, which utilizes the proposed concept of elevation semantic flow to achieve the registration of local and global features. First, in order to make the network semantics globally aware, we propose an Elevation Semantic Globalization (ESG) module to realize the semantic globalization of instances. Further, in order to alleviate the semantic span of global features and original local features, we propose a Local-to-Global Elevation Semantic Registration (L2G-ESR) module based on elevation semantic flow. Our Building3D is rooted in the SFFDE network for building elevation prediction, synchronized with a building extraction network for building masks, and then sequentially performs point cloud reconstruction and surface reconstruction (or CityGML model reconstruction). On this basis, our Building3D can optionally generate CityGML models or surface mesh models of the buildings. Extensive experiments on ISPRS Vaihingen and DFC2019 datasets on the DSM estimation task show that our SFFDE significantly improves upon state-of-the-art and δ1, δ2and δ3metrics of our SFFDE are improved to 0.595, 0.897 and 0.970. Furthermore, our Building3D achieves impressive results in the 3D point cloud and 3D model reconstruction process. Yongqiang Mao, Kaiqiang Chen, Liangjin Zhao, Deke Tang, Wenjie Liu 0016, Zhirui Wang 0003, Wenhui Diao, Xian Sun 0001, Kun Fu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2023 | DCM: A Distributed Collaborative Training Method for the Remote Sensing Image ClassificationabstractAs the number of aero and space remote sensing platforms increases, distributed observation and real-time terminal processing become mainstream in the future. However, most of the training methods for the multi-platform are still limited to centralized structures or independent training based on a single platform, which is inefficient or limited in accuracy. In order to solve this problem, we innovatively propose a distributed collaborative method (DCM) for remote sensing image classification training in this article. First, the proposed training method, which is based on one cloud and several terminals, can aggregate different parameters of the terminal network to the cloud to improve global accuracy. Second, a sample proximity network is designed to process the problem of data heterogeneity on different terminal networks, which further improves the accuracy during the model fusion on the cloud. Third, a multi-layer grouped concatenation module is applied after the model fusion to extract hierarchical features with different categories of remote sensing images. Experimental results on the challenging remote sensing image classification dataset FAIR1M show that the proposed training method has better collaborative learning ability than the centralized-based model or terminal-trained lightweight network under the heterogeneous data. Yuelei Wang, Zhirui Wang 0003, Peirui Cheng, Xuan Zeng 0004, Xian Sun 0001, Kun Fu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | DCNNet: A Distributed Convolutional Neural Network for Remote Sensing Image ClassificationabstractWith the development of information technology, multiplatform collaborative collection and processing of remote sensing (RS) images has become a significant trend. However, the existing models are challenging to achieve accurate and efficient image interpretation on RS multiplatform systems. To solve this problem, we propose a novel distributed convolutional neural network (DCNNet) and demonstrate the superiority of our method in RS image classification. First, a progressive inference mechanism is introduced to support most images to be classified in advance with satisfactory accuracy, which minimizes redundant cloud transmission and achieves higher inference acceleration. Meanwhile, a distributed self-distillation paradigm is designed to integrate and refine in-depth features, performing efficient knowledge transfer between the terminals and the cloud network. Second, a multiscale feature fusion (MSFF) module is presented to extract valid receptive fields and assign weights to crucial channel dimension features. Finally, a sampling augmentation (SA) attention is proposed to enhance the effective feature representation of RS images through a bottom-up and top-down feedforward structure. We conducted extensive experiments and visual analyses on three benchmark scene classification datasets and one fine-grained dataset. Compared with the existing methods, DCNNet consolidates several advantages in terms of accuracy, computation, transmission, and processing efficiency into a single framework for multiplatform RS image classification. Zhirui Wang 0003, Peirui Cheng, Guangluan Xu, Xian Sun 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Terrain Segmentation in Polarimetric SAR Images Using Dual-Attention Fusion NetworkabstractThe terrain segmentation in polarimetric synthetic aperture radar (PolSAR) images is an important task for image interpretation. Since the speckle noise and complex scattering mechanism exist in SAR images, the classification results achieved by traditional methods appear fragmented. Gradually, deep-learning-based methods are proposed to solve this problem. However, only the amplitude data in the SAR image is utilized, which limits the classification precision. In this letter, a novel method based on a dual-attention fusion network (DAFN) is presented. DAFN is mainly composed of a two-way structure encoder for feature extraction and the attention-based fusion module. Considering the terrain characteristic and the SAR imaging mechanism, the introduction of the polarization information in DAFN increases the discrimination of different categories, which contributes to the consistent and accurate fine-grained classification results. To demonstrate the effectiveness of the proposed method, the corresponding experiments are done based on a GaoFen-3 satellite full-polarization SAR data set, in which the superior performance in terrain segmentation is obtained. Daifeng Xiao, Zhirui Wang 0003, Youming Wu, Xian Sun 0001 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2022 | Few-Shot SAR Target Classification via MetalearningabstractThe state-of-the-art deep neural networks have made a great breakthrough in remote sensing image classification. However, the heavy dependence on large-scale data sets limits the application of the deep learning to synthetic aperture radar (SAR) automatic target recognition (ATR) field where the target sample set is generally small. In this work, a metalearning framework named MSAR, consisting of a metalearner and a base-learner, is proposed to solve the sample restriction problem, which can learn a good initialization as well as a proper update strategy. After training, MSAR can implement fast adaptation with a few training images on new tasks. To the best of our knowledge, this is the first study to solve a few-shot SAR target classification via metalearning. In particular, the few-task problem is defined by analyzing the effect of available training classes on the performance of metalearning models. In order to reduce the metalearning difficulties caused by the few-task problem, three transfer-learning methods are employed, which can leverage the prior knowledge from the pretraining phase. Besides, we design a hard task mining method for effective metalearning. Based on the Moving and Stationary Target Acquisition and Recognition (MSTAR) data set, a specialized data set named NIST-SAR is devised to train and evaluate the proposed method. The experiments on NIST-SAR have shown that the proposed method yields better performances with the largest absolute improvements of 1.7% and 2.3% for 1-shot and 5-shot, respectively, over the next best, which indicates that the proposed method is promising and metalearning is a feasible solution for few-shot SAR ATR. Kun Fu 0001, Tengfei Zhang 0004, Yue Zhang 0016, Zhirui Wang 0003, Xian Sun 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Rotation-Invariant Deep Embedding for Remote Sensing ImagesabstractEndowing convolutional neural networks (CNNs) with the rotation-invariant capability is important for characterizing the semantic contents of remote sensing (RS) images since they do not have typical orientations. Most of the existing deep methods for learning rotation-invariant CNN models are based on the design of proper convolutional or pooling layers, which aims at predicting the correct category labels of the rotated RS images equivalently. However, a few works have focused on learning rotation-invariant embeddings in the framework of deep metric learning for modeling the fine-grained semantic relationships among RS images in the embedding space. To fill this gap, we first propose a rule that the deep embeddings of rotated images should be closer to each other than those of any other images (including the images belonging to the same class). Then, we propose to maximize the joint probability of the leave-one-out image classification and rotational image identification. With the assumption of independence, such optimization leads to the minimization of a novel loss function composed of two terms: 1) a class-discrimination term and 2) a rotation-invariant term. Furthermore, we introduce a penalty parameter that balances these two terms and further propose a final loss to Rotation-invariant Deep embedding for RS images, termed RiDe. Extensive experiments conducted on two benchmark RS datasets validate the effectiveness of the proposed approach and demonstrate its superior performance when compared to other state-of-the-art methods. The codes of this article will be publicly available athttps://github.com/jiankang1991/TGRS_RiDe. Jian Kang 0005, Rubén Fernández-Beltran, Zhirui Wang 0003, Xian Sun 0001, Jingen Ni, Antonio Plaza |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | SFR-Net: Scattering Feature Relation Network for Aircraft Detection in Complex SAR ImagesabstractAircraft detection in synthetic aperture radar (SAR) images plays a significant role in dynamic monitoring and national security. Previous methods have difficulty in obtaining the desirable detection performance due to the interference of complex scenes and diversity of aircraft sizes. In order to solve these problems, we propose an innovative scattering feature relation network (SFR-Net) in this article. First, considering that the strong scattering points of the aircraft in SAR images are usually discrete, we leverage the proposed scattering point relation module to fulfill the analysis and correlation of scattering points. By enhancing the characteristics and relationships among the scattering points, this method is beneficial to guarantee the completeness of aircraft detection results. Second, we design a salient fusion module to adaptively aggregate the features from different layers of SFR-Net with rich semantic information and plentiful details, which can highlight the significant objects with different sizes and enhance the distinguishable features. Third, to reduce the false alarm and improve the localization accuracy, the contextual feature attention is presented to capture the global spatial and semantic information with a large receptive field. Overall, the SFR-Net is designed based on the SAR imaging mechanism and the scattering characteristics of aircrafts. The extensive experiments are conducted on the SAR aircraft detection dataset (AIRD) from the Gaofen-3 satellite to demonstrate the effectiveness of the SFR-Net and also illustrate that our method achieves state-of-the-art performance. Yuzhuo Kang, Zhirui Wang 0003, Jiamei Fu, Xian Sun 0001, Kun Fu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | DisOptNet: Distilling Semantic Knowledge From Optical Images for Weather-Independent Building SegmentationabstractSynthetic aperture radar (SAR) images provide all-weather and all-time capabilities for Earth observation, which becomes highly beneficial in the field of intelligent remote sensing (RS) image interpretation. Due to these advantages, SAR images have been widely exploited in automatic building segmentation tasks under poor weather conditions, especially when disasters happen. However, compared to optical images, the semantics inherent to SAR images are less rich and interpretable due to factors such as speckle noise and imaging geometry. In this scenario, most state-of-the-art methods are focused on designing advanced network architectures or loss functions for building footprint extraction. However, few works have been oriented toward improving segmentation performance through knowledge transfer from optical images. In this article, we propose a novel method based on theDisOptNetnetwork, which can distill the useful semantic knowledge from optical images into a network only trained with SAR data. Specifically, we first analyze the multilevel feature discrepancies between multiple stages of the networks pretrained on the two image modalities. We observe that feature discrepancies start to increase as the encoding stage gradually changes from low level to high level. Based on such observation, we reuse the early stage features and construct parallel convolutional neural network (CNN) branches that are responsible for capturing high-level domain-specific knowledge for each image modality. The optical branch is aimed at mimicking feature generation at the optical pretrained network given the input SAR images. Then, an aggregation module is introduced to calibrate and fuse the features from different modalities while generating the building segments. Extensive experiments were conducted on a large-scale multisensor all-weather building segmentation dataset with state-of-the-art methods used for comparison. Our experimental results validate the effectiveness ofDisOptNet, which demonstrates great potential in the task of weather-independent building footprint generation under real scenarios. The codes of this article will be made publicly available athttps://github.com/jiankang1991/TGRS_DisOptNet. Jian Kang 0005, Zhirui Wang 0003, Ruoxin Zhu, Junshi Xia, Xian Sun 0001, Rubén Fernández-Beltran, Antonio Plaza |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | SCAN: Scattering Characteristics Analysis Network for Few-Shot Aircraft Classification in High-Resolution SAR ImagesabstractRecently, deep learning in synthetic aperture radar (SAR) automatic target recognition (ATR) has made significant progress, but the sample limitation problem in the SAR field is still obvious. Compared with the optical remote sensing images, the SAR images are insufficient, especially those containing the geospatial targets with certain target attitude angles (TAAs). To solve these problems, a novel few-shot learning framework named scattering characteristics analysis network (SCAN) is proposed in this article. First, a scattering extraction module (SEM) is designed to combine the target imaging mechanism with the network, which learns the number and distribution of the scattering points for each target type via explicit supervision. Besides, considering the imaging variability of SAR targets, a TAA-guided metalearning network consisting of an angle self-adaption classifier (ASC) and a frequency embedded module (FEM) is designed. ASC guides the network to focus on the positive sample pairs with different TAAs. FEM combines pulse cosine transform (PCT) with the network training process effectively to enrich frequency-domain information. In addition, a new dataset named SAR aircraft category dataset is constructed for the experiments. Compared with other few-shot SAR target classification approaches, our model efficiently integrates the scattering characteristics with the learning process, and the test accuracy for 5-way 1-shot has been improved by 4.74%. Finally, the experimental results are provided to demonstrate the validity of the proposed method. Xian Sun 0001, Yixuan Lv, Zhirui Wang 0003, Kun Fu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Oriented Ship Detection Based on Strong Scattering Points Network in Large-Scale SAR ImagesabstractShip detection has broad applications in many areas, including fishery management, maritime rescue, and maritime monitoring. Recently, numerous detectors based on deep learning have been carried in ship detection in synthetic aperture radar (SAR) images. However, detecting the inshore ships faces enormous challenges because of the strong scattering interference of the inland area. In order to address such issues, a novel method named strong scattering points network for ship detection is proposed in this article. First, according to the SAR imaging mechanism, the ships usually appear strong scattering phenomenon in the SAR images. Therefore, the proposed method detects the strong scattering points on the ship and then aggregates their positions to obtain the ship’s arbitrary orientation box. Second, our method designs an embedding vector to cluster these points as an individual object to regress the oriented bounding box. Third, in order to distinguish the strong scattering points on land, a ship attention module is employed to extract the image texture features and representations of local features. It can suppress the false alarm caused by land interference in the detection process. Furthermore, to demonstrate the effectiveness of the proposed algorithm, this article introduces a new ship dataset for oriented ship detection named large-scale dataset for ship detection in SAR images (LDSD). Moreover, the public SAR ship detection dataset (SSDD) is utilized to verify the robustness and generalization ability of the detector. The experimental results on two datasets show that our method has a strong anti-interference ability in the inshore background and achieves state-of-the-art detection performance. Yuanrui Sun, Xian Sun 0001, Zhirui Wang 0003, Kun Fu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | DMML-Net: Deep Metametric Learning for Few-Shot Geographic Object Segmentation in Remote Sensing ImageryabstractGeographic object segmentation is a fundamental yet challenging problem for remote sensing image interpretation. The prevalent paradigm to solve this problem is to train deep neural networks on massive labeled samples. Although remarkable achievements have been attained, these methods suffer from the severe dependence on the large-scale dataset and require a long training process with high computation burden. To address these issues, a deep metametric learning framework, named DMML-Net, consisting of the metametric learner and the base-metric learner, is proposed for few-shot geographic object segmentation. First, DMML-Net formulates the segmentation as the metric-based pixel classification and develops a deep feature pyramid comparison network as the architecture of the metric learner for multiscale metric learning. Benefiting from this design, the segmentation can be efficiently solved, as well as being robust to deal with the scale variations of geographic objects. Second, an affinity-based fusion mechanism is introduced to adaptively reweight and fuse the semantic information across samples, effectively calibrating the deviation of prototypes induced by the intraclass variations. Third, considering the impact of the large interclass distribution divergences, DMML-Net presents a metametric training paradigm to provide the metric model with flexible scalability for fast adaptation to novel tasks. After metatraining, DMML-Net can be applied for the few-shot segmentation tasks of novel geographic objects with only a few gradient steps on the small training set. Experimental results on two benchmark remote sensing datasets demonstrate the validity and the superiority of our method in low-shot conditions where there are only one to ten labeled samples. Bing Wang 0015, Zhirui Wang 0003, Xian Sun 0001, Kun Fu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | TS-SHES: Terrain Segmentation in Complex-Valued PolSAR Images Via Scattering Harmonization and Explicit SupervisionabstractConvolutional neural network (CNN) has attracted extensive attention in the research field of polarimetric synthetic aperture radar (PolSAR) terrain segmentation. However, directly using CNN in PolSAR terrain segmentation while ignoring the characteristics of PolSAR images has become the main factor restricting the performance of algorithms. In this article, we propose an efficient PolSAR terrain segmentation algorithm called TS-SHES, which integrates the polarization scattering characteristics of PolSAR images and the CNN learning process into a unified architecture. First, considering the intrinsic structure of complex-valued PolSAR data, TS-SHES transforms the scattering matrix into the form of amplitude and phase components, which preserves the original information maximally. Then, TS-SHES introduces a scattering harmonized encoding method (SH-Enc) to balance the feature contributions of weak and strong scattering regions as well as map the two components into the same representation space. Through the above scattering harmonization operations, the segmentation performance of CNN on weak scattering regions can be improved, and the feature imbalance in amplitude and phase can be alleviated. Furthermore, in view of the implicit states of CNN feature construction, a scattering explicit learning network (SEL-Net) is presented to collect the scattering features of amplitude and phase. Via explicit supervision, SEL-Net avoids the incomplete collection of scattering information caused by implicit feature construction, thereby improving the segmentation accuracy. Abundant experiments are conducted on two PolSAR images acquired by the GaoFen-3 satellite, which demonstrates the superiority of our proposed algorithm. Xuan Zeng 0004, Zhirui Wang 0003, Xian Sun 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | DENet: Double-Encoder Network With Feature Refinement and Region Adaption for Terrain Segmentation in PolSAR ImagesabstractRecently, many studies exploit deep neural networks to promote terrain segmentation in polarimetric synthetic aperture radar (PolSAR) images. However, these works usually inherit the nature-scene approaches directly and may not be robust for the PolSAR image segmentation task. The main limitations include single-type feature construction, weak feature consistency, and geometry-agnostic collection of scattering information. In this article, we present the DENet, a double-encoder network with feature refinement and region adaption for the terrain segmentation in PolSAR images. First, a double-encoder architecture is proposed to leverage the multitype information of PolSAR images, which can provide more discriminative features than the previous methods using the single-type feature. Second, considering that the polarization information has strong consistency over the category-identical regions, a polarization-guided refinement module is proposed to maintain the feature consistency in the PolSAR segmentation model. This design alleviates the phenomenon of incomplete and fragmented segmentation results. Third, in view of the rich targets’ characteristics in the scattering information, a region-adaptive convolution module is developed to facilitate the scattering information collected over the geometry-irregular regions. This design can improve the segmentation accuracy on the geometry-irregular regions. Extensive experiments are conducted on six PolSAR images to verify the effectiveness of the DENet. Compared with the previous works, our method achieves competitive performance. Xuan Zeng 0004, Zhirui Wang 0003, Xian Sun 0001, Zhonghan Chang |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2021 | An Anchor-Free Method Based on Feature Balancing and Refinement Network for Multiscale Ship Detection in SAR ImagesabstractRecently, deep-learning methods have been successfully applied to the ship detection in the synthetic aperture radar (SAR) images. It is still a great challenge to detect multiscale SAR ships due to the broad diversity of the scales and the strong interference of the inshore background. Most prevalent approaches are based on the anchor mechanism that uses the predefined anchors to search the possible regions containing objects. However, the anchor settings have a great impact on their detection performance as well as the generalization ability. Furthermore, considering the sparsity of the ships, most anchors are redundant and will lead to the computation increase. In this article, a novel detection method named feature balancing and refinement network (FBR-Net) is proposed. First, our method eliminates the effect of anchors by adopting a general anchor-free strategy that directly learns the encoded bounding boxes. Second, we leverage the proposed attention-guided balanced pyramid to balance semantically the multiple features across different levels. It can help the detector learn more information about the small-scale ships in complex scenes. Third, considering the SAR imaging mechanism, the interference near the ship boundary with the similar scattering power probably affects the localization accuracy because of feature misalignment. To tackle the localization issue, a feature-refinement module is proposed to refine the object features and guide the semantic enhancement. Finally, extensive experiments are conducted to show the effectiveness of our FBR-Net compared with the general anchor-free baseline. The detection results on the SAR ship detection dataset (SSDD) and AIR-SARShip-1.0 dataset illustrate that our method achieves the state-of-the-art performance. Jiamei Fu, Xian Sun 0001, Zhirui Wang 0003, Kun Fu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2019 | Geometrical Model for the Layover of Gable-Roofed Buildings and its Application in Building ReconstructionabstractBuilding reconstruction from SAR images is a hot topic in recent years. Currently, related methods mainly deal with on flat-roofed buildings. In this paper, we extend the research scope to gable-roofed buildings, and try to present a parameterized geometrical model for the layover of gable-roofed buildings. Based on this model, a top-down building reconstruction technique based on MCMC method is proposed. Through representing the layover with parameterized geometrical models, building reconstruction is converted into an optimization problem under the Bayesian scheme. In order to obtain global optima, simulated annealing algorithm with MCMC is used in the optimization stage. Two groups of transmission kernels which are responsible for model updates are designed according to the model. Experiments show that the layover model is accurate and the reconstruction method is effective. Yue Zhang 0016, Zhirui Wang 0003, Liangjin Zhao, Wenkai Zhang 0002, Menglong Yan, Xian Sun 0001 |
IGARSS | 2 |
| 2019 | A Training-Free, One-Shot Detection Framework for Geospatial Objects in Remote Sensing ImagesabstractDeep learning based object detection has achieved great success. However, these supervised learning methods are data-hungry and time-consuming. This restriction makes them unsuitable for limited data and urgent tasks, especially in the applications of remote sensing. Inspired by the ability of humans to quickly learn new visual concepts from very few examples, we propose a training-free, one-shot geospatial object detection framework for remote sensing images. It consists of (1) a feature extractor with remote sensing domain knowledge, (2) a multi-level feature fusion method, (3) a novel similarity metric method, and (4) a 2-stage object detection pipeline. Experiments on sewage treatment plant and airport detections show that proposed method has achieved a certain effect. Our method can serve as a baseline for training-free, one-shot geospatial object detection. Tengfei Zhang 0004, Xian Sun 0001, Yue Zhang 0016, Menglong Yan, Yaoling Wang, Zhirui Wang 0003, Kun Fu 0001 |
IGARSS | 6 |
| 2019 | Ground Moving Target Indication Based on Optical Flow in Single-Channel SARabstractAn algorithm based on optical flow is proposed to detect a ground moving target via the single-channel synthetic aperture radar. First, the signal models of uniform moving targets are established and classified into three types. Next, the Doppler spectrum is divided to generate a multilook image sequence. Then, the motion feature of a moving target response is described in the image sequence, in which the optical flow is introduced to realize the moving target detection. The detection results of real moving targets are obtained after the false alarm elimination based on the response motion relevance. This algorithm has a large range of detectable velocity and can even be applied to detect the moving targets with acceleration. In addition, compared with constant false alarm rate method, the optical flow has a better anti-interference performance against the strong static scatters. Finally, some numerical experiments are provided to demonstrate the effectiveness of the proposed method. Zhirui Wang 0003, Xian Sun 0001, Wenhui Diao, Yue Zhang 0016, Menglong Yan, Lan Lan 0001 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2018 | Radial Velocity Retrieval for Multichannel SAR Moving Targets With Time-Space Doppler DeambiguityabstractIn this paper, with respect to multichannel synthetic aperture radar (SAR), we first formulate the problems of Doppler ambiguities on the radial velocity (RV) estimation of a ground moving target in the range-compressed domain, the range-Doppler domain, and the image domain, respectively. It is revealed that in these problems, the cascaded time–space Doppler ambiguity (CTSDA) may arise; that is, the time domain Doppler ambiguity in each channel arises first and then the spatial domain Doppler ambiguity among multichannels arises second. Accordingly, the multichannel SAR systems with different parameters are investigated in three cases with different Doppler ambiguity properties. Then, a multifrequency SAR is proposed for the RV estimation by solving the ambiguity problem based on the Chinese remainder theorem (CRT). In the first two cases, the ambiguity problem can be solved by the existing closed-form robust CRT. In the third case, it is found that the problem is different from the conventional CRT problem and we call it a double remaindering problem in this paper. We then propose a sufficient condition under which the double remaindering problem, i.e., the CTSDA, can also be solved by the closed-form robust CRT. When the sufficient condition is not satisfied, a searching-based method is proposed. Finally, some results of numerical experiments are provided to demonstrate the effectiveness of the proposed methods. Jia Xu 0001, Zu-Zhen Huang, Zhirui Wang 0003, Li Xiao 0002, Xiang-Gen Xia 0001, Teng Long 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2016 | DLSLA 3-D SAR Imaging Based on Reweighted Gridless Sparse Recovery MethodabstractDownward-looking sparse-linear-array 3-D synthetic aperture radar (DLSLA 3-D SAR) cross-track reconstruction usually suffers from incomplete observation and limited resolution. The incomplete observation is caused by the sparse and nonuniform distribution of the equivalent antenna phase centers (APCs) due to the array elements' installation location restriction, loss, or deviation. Sparse recovery methods provide a solution with improved resolution from the incomplete observation for the 3-D imaging scene that behaves with spatial sparsity. However, conventional grid-based sparse recovery (GB-SR) methods are under the assumption that the scatterers are located on the discretized grids; otherwise, the off-grid effect or basis mismatch problem will occur. In this letter, we propose a reweighted scheme-based gridless sparse recovery (GL-SR) method, i.e., reweighted gridless sparse iterative covariance-based estimation (RGLS), for DLSLA 3-D SAR cross-track imaging. The proposed method possesses the merits of gridless SPICE (GLS), i.e., free of off-grid effect and user parameters, and has a statistically more appealing property than GLS by adopting the reweighted scheme. As seen from the experiments that compare the performance of GB-SR and GL-SR methods for DLSLA 3-D SAR cross-track reconstruction, the proposed method performs outstandingly under the circumstance of sparse and nonuniform APCs' distribution. Qian Bao, Xueming Peng, Zhirui Wang 0003, Yun Lin 0002, Wen Hong |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2016 | Road-Aided Doppler Ambiguity Resolver for SAR Ground Moving Target in the Image DomainabstractA new Doppler ambiguity resolver (DAR) is proposed for ground moving targets of synthetic aperture radar (SAR) in the image domain. Based on the range-Doppler imaging of a static scene, the moving target's response is analyzed in the image domain, and the target's response slope is jointly determined by three motion parameters, i.e., azimuth velocity, ambiguous range velocity, and Doppler ambiguity number. A new DAR utilizing these three parameters is then proposed via the following steps. First, the moving target is detected after ground clutter cancelation between dual-channel images. Second, the ambiguous range velocity is estimated, and the road and its slope are extracted from the SAR image to establish the relationship between the 2-D velocities. Third, the Doppler ambiguity number is determined based on the response slope, the road slope, and the estimated ambiguous range velocity. Compared with the existing DARs in the 2-D time domain, the proposed algorithm works better in the most common signal-to-clutter-noise ratio scenarios. Finally, the results of the numerical experiments are provided to demonstrate the effectiveness of the proposed method. Zhirui Wang 0003, Jia Xu 0001, Zu-Zhen Huang, Xudong Zhang 0001, Xiang-Gen Xia 0001, Teng Long 0001 |
IEEE Geosci. Remote. Sens. Lett. | 1 |