Wenxian Yu

dblp:91/2159 · DBLP profile ↗
← Back
161ranked-venue papers
0as first author
72since 2021 · last 2026
0000-0002-8741-776XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 116 · 39 since 2021Artificial intelligence and machine learning · 22 · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 12 since 2021Systems, architecture and hardware · 9 · 6 since 2021Computer networks · 9 · 9 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 CoordAR: One-Reference 6D Pose Estimation of Novel Objects via Autoregressive Coordinate Map Generation
abstract
Object 6D pose estimation, a crucial task for robotics and augmented reality applications, becomes particularly challenging when dealing with novel objects whose 3D models are not readily available. To reduce dependency on 3D models, recent studies have explored one-reference-based pose estimation, which requires only a single reference view instead of a complete 3D model. However, existing methods that rely on real-valued coordinate regression suffer from limited global consistency due to the local nature of convolutional architectures and face challenges in symmetric or occluded scenarios owing to a lack of uncertainty modeling. We present CoordAR, a novel autoregressive framework for one-reference 6D pose estimation of unseen objects. CoordAR formulates 3D-3D correspondences between the reference and query views as a map of discrete tokens, which is obtained in an autoregressive and probabilistic manner. To enable accurate correspondence regression, CoordAR introduces 1) a novel coordinate map tokenization that enables probabilistic prediction over discretized 3D space; 2) a modality-decoupled encoding strategy that separately encodes RGB appearance and coordinate cues; and 3) an autoregressive transformer decoder conditioned on both position-aligned query features and the partially generated token sequence. With these novel mechanisms, CoordAR significantly outperforms existing methods on multiple benchmarks and demonstrates strong robustness to symmetry, occlusion, and other challenges in real-world tests.
Dexin Zuo, Wenxian Yu, Danping Zou
AAAI4
2026 Exploiting Unlabeled Data with Multiple Expert Teachers for Open Vocabulary Aerial Object Detection and Its Orientation Adaptation
Yan Li 0098, Weiwei Guo, Xue Yang 0005, Ning Liao, Shaofeng Zhang, Yi Yu 0010, Wenxian Yu, Junchi Yan
Int. J. Comput. Vis.7
2026 MT-FusionNet: Mamba-transformer-assisted feature fusion for visual place recognition
Muhammad Fahad 0017, Di He 0002, Wenxian Yu, Trieu-Kien Truong
Neurocomputing3
2026 Fractal-Domain Vision Graph Neural Network for Remote Sensing Ground Target Classification
abstract
To the best of our knowledge, this paper is the first to integrate fractal signal processing with vision graph neural networks, establishing a new graph representation learning paradigm consistent with fractal dynamics. Building on this foundation, we propose a Fractal-domain Vision Graph Neural Network (FD-ViG). Specifically, FD-ViG includes: (i) a Fractal-Domain Learning Module that maps images into the fractal-domain using local Hölder exponents and the Singularity Power Spectrum (SPS), enabling fractal-spatial feature fusion; (ii) a Fractal Graph Construction Module that adaptively generates a topology by combining semantic attention with fractal similarity in the fractal feature space; and (iii) a Graph Propagation Module with power-law multi-scale propagation to realize cross-scale diffusion and aggregation, enabling coupled texture-structure learning. Experiments on UCMerced, RSSCN7, and SIRI-WHU achieve overall accuracies of 91.75%, 89.52%, and 92.78%, respectively. Compared with representative vision graph models such as ViG, WiGNet, and ViHGNN, our method achieves consistent improvements over prior methods across all three datasets, while remaining lightweight (2.6 M parameters). Moreover, despite having far fewer parameters than ResNet-18, our model yields competitive or better performance on two datasets, and further demonstrates strong generalization ability in cross-dataset evaluation on SAR imagery. This work provides a principled and effective bridge between fractal theory and graph deep learning, benefiting interpretable remote sensing scene understanding under complex textures and structures.
Jiacheng Yin, Tao Zhen, Wenxian Yu
IEEE Trans. Pattern Anal. Mach. Intell.4
2026 Grid-Reg: Detector-Free Gridized Feature Learning and Matching for Large-Scale SAR-Optical Image Registration
abstract
It is highly challenging to register large-scale, heterogeneous SAR and optical images, particularly across platforms, due to significant geometric, radiometric, and temporal differences, which most existing methods struggle to address. To overcome these challenges, we propose Grid-Reg, a grid-based multimodal registration framework comprising a domain-robust descriptor extraction network, Hybrid Siamese Correlation Metric Learning Network (HSCMLNet), and a grid-based solver (Grid-Solver) for transformation parameter estimation. In heterogeneous imagery with large modality gaps and geometric differences, obtaining accurate correspondences is inherently difficult. To robustly measure similarity between gridded patches, HSCMLNet integrates a hybrid Siamese module with a correlation metric learning module (CMLModule) based on equiangular unit basis vectors (EUBVs), together with a manifold consistency loss to promote modality-invariant, discriminative feature learning. The Grid-Solver estimates transformation parameters by minimizing a global grid matching loss through a dual-loop search strategy to reliably find patch correspondences across entire images. Furthermore, we curate a challenging benchmark dataset for SAR-to-optical registration using UAV MiniSAR data and Google Earth optical imagery. Extensive experiments demonstrate that our proposed approach achieves superior performance over state-of-the-art methods.
Xiaochen Wei, Weiwei Guo, Zenghui Zhang, Wenxian Yu
IEEE Trans. Circuits Syst. Video Technol.4
2026 Point Cloud Analysis Under Slight Perturbations: A Manifold Distillation Approach Using Raw Coordinates
abstract
Point cloud is often regarded as a discrete sampling of Riemannian manifold and plays a pivotal role in the 3D image interpretation. Particularly, rotation perturbation, an unexpected small change in rotation caused by various factors (like equipment offset, system instability, measurement errors and so on), can easily lead to the inferior results in point cloud learning tasks. However, classical point cloud learning methods are sensitive to rotation perturbation, and the existing networks with rotation robustness also have much room for improvements in terms of performance and noise tolerance. Given these, this paper remodels the point cloud from the perspective of manifold as well as designs a manifold distillation method to achieve the robustness of rotation perturbation without any coordinate transformation. In brief, during the training phase, we introduce a teacher network to learn the rotation robustness information and transfer this information to the student network through online distillation. In the inference phase, the student network directly utilizes the raw 3D coordinate information to achieve the robustness of rotation perturbation. Experiments carried out on four different datasets verify the effectiveness of our method. On average, on the ModelNet40 and ScanObjectNN classification datasets with random rotation perturbations, our method improves classification accuracy by 4.41% and 3.65%, respectively, compared to popular rotation-robust networks. Similarly, on the ShapeNet and S3DIS segmentation datasets, our method achieves improvements in mIoU of 6.96% and 5.12%, respectively. Furthermore, the experimental results also demonstrate that our algorithm exhibits higher computational efficiency and stronger resistance to noise and outliers.
Tao Zhang 0027, Huazhen Liu, Feiming Wei, Huilin Xiong, Wenxian Yu
IEEE Trans. Circuits Syst. Video Technol.6
2025 An Enhanced 3-D Direct Positioning Method with Adaptive Grid Refinement in Indoor Environments
Di He 0002, Xuyu Gao, Wenxian Yu
GLOBECOM4
2025 MMCD: Memory-Based Multimodal Change Detection
abstract
Single-modal change detection methods based on optical or Synthetic Aperture Radar (SAR) images face challenges such as degradation due to adverse weather or noise interference. In contrast, multimodal change detection struggles with significant domain gaps between different modalities. Inspired by the SAM2 model’s temporal memory mechanism for video segmentation, this paper introduces the concept of memory into change detection and proposes a novel approach called Memory-based Multimodal Change Detection (MMCD). By treating change detection as a temporal problem and modeling remote sensing images as video sequences, the proposed method integrates historical optical images with current SAR images to enhance detection accuracy. Additionally, a difference map enhancement module is introduced to mitigate false changes caused by modality discrepancies. Experimental results show that this approach achieves state-of-the-art performance in multimodal change detection, demonstrating the effectiveness of the proposed method.
Limeng Zhang, Zenghui Zhang, Juanping Wu, Weiwei Guo, Tao Zhang 0027, Wenxian Yu
ICASSP6
2025 mmDEAR: mmWave Point Cloud Density Enhancement for Accurate Human Body Reconstruction
abstract
Millimeter-wave (mmWave) radar offers robust sensing capabilities in diverse environments, making it a highly promising solution for human body reconstruction due to its privacy-friendly and non-intrusive nature. However, the significant sparsity of mm Wave point clouds limits the estimation accuracy. To overcome this challenge, we propose a two-stage deep learning framework that enhances mm Wave point clouds and improves human body reconstruction accuracy. Our method includes a mm Wave point cloud enhancement module that densifies the raw data by leveraging temporal features and a multi-stage completion network, followed by a 2D-3D fusion module that extracts both 2D and 3D motion features to refine SMPL parameters. The mm Wave point cloud enhancement module learns the detailed shape and posture information from 2D human masks in single-view images. However, image-based supervision is involved only during the training phase, and the inference relies solely on sparse point clouds to maintain privacy. Experiments on multiple datasets demonstrate that our approach outperforms state-of-the-art methods, with the enhanced point clouds further improving performance when integrated into existing models.
Songpengcheng Xia, Zengyuan Lai, Qi Wu 0007, Wenxian Yu, Ling Pei
ICRA6
2025 PlanarGS: High-Fidelity Indoor 3D Gaussian Splatting Guided by Vision-Language Planar Priors
abstract
Three-dimensional Gaussian Splatting (3DGS) has recently emerged as an efficient representation for novel-view synthesis, achieving impressive visual quality. However, in scenes dominated by large and low-texture regions, common in indoor environments, the photometric loss used to optimize 3DGS yields ambiguous geometry and fails to recover high-fidelity 3D surfaces. To overcome this limitation, we introduce PlanarGS, a 3DGS-based framework tailored for indoor scene reconstruction. Specifically, we design a pipeline for Language-Prompted Planar Priors (LP3) that employs a pretrained vision-language segmentation model and refines its region proposals via cross-view fusion and inspection with geometric priors. 3D Gaussians in our framework are optimized with two additional terms: a planar prior supervision term that enforces planar consistency, and a geometric prior supervision term that steers the Gaussians toward the depth and normal cues. We have conducted extensive experiments on standard indoor benchmarks. The results show that PlanarGS reconstructs accurate and detailed 3D surfaces, consistently outperforming state-of-the-art methods by a large margin. Project page: https://planargs.github.io
Xirui Jin, Renbiao Jin, Boying Li, Danping Zou, Wenxian Yu
NeurIPS5
2025 An Indoor Direct Localization Method Utilizing Structural Sparsity and Low-Rankness With Grid Refinement
abstract
Indoor positioning technologies are pivotal for achieving high-precision localization in GPS-denied environments. Wireless positioning methods based on angle of arrival (AOA) models leverage spatial diversity to mitigate indoor multipath challenges. Recent approaches combine the extended sparse reconstruction framework with direct positioning (DP) techniques to jointly extract location parameters at the raw signal level. However, the disadvantage is that they have limited utilization for structure information inherent in a sparse solution and overlook the low-rankness in array manifolds induced by redundant grid points. To overcome these problems, a novel structure-aware direct localization method namely SaDPD is proposed that incorporates comprehensive structural constraints of joint row plus element sparsity and low-rankness in the designed sparse framework. A new sparse reconstruction problem is formulated by decomposing the position weight matrix into two distinct feature matrices, in which they are optimized for correlated line-of-sight (LOS) and inconsistent non-line-of-sight (NLOS) components, respectively. In particular, it employs the Alternating Direction Method of Multipliers (ADMM) to effectively address the extended high-dimensional optimization and utilizes grid refinement to avoid quantization constraints. Finally, extensive simulations verify that SaDPD tremendously improves sparse reconstruction accuracy and localization performance. Experimental result further demonstrates its effectiveness and practical applicability against real-world interference. These advancements establish SaDPD as a robust solution for asset positioning and target tracking where localization accuracy beyond half-meter level under multipath interference is critical.
Di He 0002, Longwei Tian, Wenxian Yu, Trieu-Kien Truong
IEEE Internet Things J.4
2025 SMART: Scene-Motion-Aware Human Action Recognition Framework for Mental Disorder Group
abstract
Patients with mental disorders often exhibit risky abnormal actions, such as climbing walls or hitting windows, necessitating intelligent video behavior monitoring for smart healthcare with the rising Internet of Things (IoT) technology. However, the development of vision-based human action recognition (HAR) for these actions is hindered by the lack of specialized algorithms and datasets. In this article, we innovatively propose to build a vision-based HAR dataset, including abnormal actions often occurring in the mental disorder group and then introduce a novel scene-motion-aware action recognition technology framework, named SMART, consisting of two technical modules. First, we propose a scene perception module to extract human motion trajectory and human-scene interaction features, which introduces additional scene information for a supplementary semantic representation of the above actions. Second, the multistage fusion module fuses the skeleton motion, motion trajectory, and human-scene interaction features, enhancing the semantic association between the skeleton motion and the above supplementary representation, thus generating a comprehensive representation with both human motion and scene information. The effectiveness of our proposed method has been validated on our self-collected HAR dataset (MentalHAD), achieving 94.9% and 93.1% accuracy in un-seen subjects and scenes and outperforming state-of-the-art approaches by 6.5% and 13.2%, respectively. The demonstrated subject- and scene- generalizability makes it possible for SMART’s migration to practical deployment in smart healthcare systems for mental disorder patients in medical settings. The code and dataset will be released publicly for further research:https://github.com/Inowlzy/SMART.git.
Zengyuan Lai, Songpengcheng Xia, Qi Wu 0007, Wenxian Yu, Ling Pei
IEEE Internet Things J.6
2025 Feature Adaptive Iteration and Multipath Pseudorange Correction for GNSS Positioning in Urban Environments
abstract
Global navigation satellite system (GNSS) has established strong connections with users worldwide. However, when propagating in complex urban environments, GNSS signals are susceptible to blockage and reflection from surrounding obstacles, resulting in multipath (MP) effects and non-line-of-sight (NLOS) reception, significantly degrading positioning precision. With the rapid development of artificial intelligence technology, numerous researchers have advocated employing machine learning algorithms to address this issue. To effectively mitigate MP and NLOS, this study firstly introduces a novel signal classification-based MP pseudorange error correction model by initially categorizing signals into line-of-sight (LOS), MP, and NLOS. Secondly, to enhance the prediction accuracy of the model, an innovative feature adaptive iteration algorithm is proposed to update feature values. Additionally, this study is the first to propose utilizing the more reliable differential receiver clock error non-excluded pseudorange error (DNEPE), obtained through differential techniques, as the label for pseudorange error prediction. The performance of the algorithm model was tested using dynamic GNSS raw data collected by the Google team in the San Francisco area of the United States. Experimental results demonstrate that the proposed method in this study significantly improves positioning precision compared to baseline and comparative methods. Specifically, in comparison to the baseline, the proposed method exhibits an enhancement of 54.50% and 66.10% in two-dimensional root mean square error (RMSE) of positioning errors on the testing sets 1 and 2, respectively. Overall, the methods proposed in this study offer novel insights and robust support for the mitigation of MP and NLOS in urban environments.
Di He 0002, Wenxian Yu
IEEE Internet Things J.4
2025 Sparse-to-Dense Hint Guided Stereo-LiDAR Fusion
abstract
One challenge in stereo-LiDAR fusion arises from the sparsity and non-uniform distribution of LiDAR data. Existing methods expand sparse LiDAR data to produce semi-dense hints as guidance for fusion. However, the absence of depth cues beyond the expanded areas may still limit performance. To address this challenge, we propose a novel sparse-to-dense hint guided stereo-LiDAR fusion method. The key idea is to use a dense hint map generated by a lightweight network as guidance, with sparse LiDAR points and a monocular image as inputs. The dense hints are then employed to construct and explicitly regularize a multi-modal cost volume via integrating the geometric cues from the hints and the visual information from the images to produce better stereo prediction. The construction and aggregation of cost volume follow a well-designed coarse-to-fine strategy along with a pixel-wise search range adjustment module, facilitating fast computation while preserving fine details. Finally, a confidence-based fusion module is performed to adaptively produce the ultimate prediction based on the monocular and stereo estimations. The experimental results show that our method significantly outperforms existing methods with high inference efficiency across multiple benchmark datasets. To contribute to the community, we will release the code at: https://github.com/LiAngLA66/DG-Fusion.
Ang Li 0029, Dexin Zuo, Anning Hu, Wenxian Yu, Danping Zou
IEEE Trans. Circuits Syst. Video Technol.4
2025 VLF-SAR: A Novel Vision-Language Framework for Few-Shot SAR Target Recognition
abstract
Due to the challenges of obtaining data from valuable targets, few-shot learning plays a critical role in synthetic aperture radar (SAR) target recognition. However, the high noise levels and complex backgrounds inherent in SAR data make this technology difficult to implement. To improve the recognition accuracy, in this paper, we propose a novel vision-language framework, VLF-SAR, with two specialized models: VLF-SAR-P for polarimetric SAR (PolSAR) data and VLF-SAR-T for traditional SAR data. Both models start with a frequency embedded module (FEM) to generate key structural features. For VLF-SAR-P, a polarimetric feature selector (PFS) is further introduced to identify the most relevant polarimetric features. Also, a novel adaptive multimodal triple attention mechanism (AMTAM) is designed to facilitate dynamic interactions between different kinds of features. For VLF-SAR-T, after FEM, a multimodal fusion attention mechanism (MFAM) is correspondingly proposed to fuse and adapt information extracted from frozen contrastive language-image pre-training (CLIP) encoders across different modalities. Extensive experiments on the OpenSARShip2.0, FUSAR-Ship, and SAR-AirCraft-1.0 datasets demonstrate the superiority of VLF-SAR over some state-of-the-art methods, offering a promising approach for few-shot SAR target recognition.
Nishang Xie, Tao Zhang 0027, Lanyu Zhang, Feiming Wei, Wenxian Yu
IEEE Trans. Circuits Syst. Video Technol.6
2025 CDPrompt: Multimodal Change Detection With In-Domain Prompt in Missing Modality Scenarios
abstract
The change detection aims to identify temporal changes in land cover. In emergency disaster scenarios, acquiring postchange optical images is often difficult due to factors such as adverse weather and illumination conditions. In contrast, the SAR-based change detection is robust to these environmental factors but is prone to speckle noise and often lacks clear semantic interpretation. These challenges highlight the importance of multimodal approaches that integrate the complementary information from different data sources. To address the domain gap between optical and SAR data, we propose change detection prompt (CDPrompt), an automatic prompt-learning framework that leverages in-domain change information as prompts to suppress fake changes caused by the domain gap between the two modalities. CDPrompt incorporates a modality-specific domain tuning module (DTM) to inject the domain knowledge into the segment anything model (SAM), enabling efficient adaptation to multimodal data with minimal labels and training costs. A low-level enhancement module (LwEM) further refines spatial details using historical optical images, while a consistency loss enhances the learning of domain-invariant features between prechange optical and SAR images. To support evaluation in disaster scenarios with missing modalities, we extend the DFC25 dataset and introduce the first disaster-oriented multimodal change detection dataset, DFC25-Extended, comprising DFC25-OS-S and DFC25-O-SO. Extensive experiments on the Onera Satellite Change Detection (OSCD) and DFC25-Extended datasets demonstrate the superior performance and practical value of CDPrompt. The code and dataset will be publicly available at:https://github.com/zhanglimeng13/CDPrompt
Limeng Zhang, Zenghui Zhang, Tao Zhang 0027, Gui Gao, Wenxian Yu
IEEE Trans. Geosci. Remote. Sens.5
2025 PreCM: The Padding-Based Rotation Equivariant Convolution Mode for Semantic Segmentation
abstract
Semantic segmentation is an important branch of image processing and computer vision. With the popularity of deep learning, various convolutional neural networks have been proposed for pixel-level classification and segmentation tasks. In practical scenarios, however, imaging angles are often arbitrary, encompassing instances such as water body images from remote sensing and capillary and polyp images in the medical domain, where prior orientation information is typically unavailable to guide these networks to extract more effective features. In this case, learning features from objects with diverse orientation information poses a significant challenge, as the majority of CNN-based semantic segmentation networks lack rotation equivariance to resist the disturbance from orientation information. To address this challenge, this paper first constructs a universal convolution-group framework aimed at more fully utilizing orientation information and equipping the network with rotation equivariance. Subsequently, we mathematically design a padding-based rotation equivariant convolution mode (PreCM), which is not only applicable to multi-scale images and convolutional kernels but can also serve as a replacement component for various types of convolutions, such as dilated convolutions, transposed convolutions, and asymmetric convolution. To quantitatively assess the impact of image rotation in semantic segmentation tasks, we also propose a new evaluation metric, Rotation Difference (RD). The replacement experiments related to six existing semantic segmentation networks on three datasets (i.e., Satellite Images of Water Bodies, DRIVE, and Floodnet) show that, the average Intersection Over Union (IOU) of their PreCM-based versions respectively improve 6.91%, 10.63%, 4.53%, 5.93%, 7.48%, 8.33% compared to their original versions in terms of random angle rotation. And the average RD values are decreased by 3.58%, 4.56%, 3.47%, 3.66%, 3.47%, 3.43% respectively. The code can be download from https://github.com/XinyuXu414.
Huazhen Liu, Tao Zhang 0027, Huilin Xiong, Wenxian Yu
IEEE Trans. Image Process.5
2025 In-P3VINS: Tightly-Coupled PPP/INS/Visual SLAM Based on Invariant Optimization Approach
abstract
The state estimation on$SE_{2}(3)$Lie Group has been proven to have the ability to improve the consistency of the estimated results. They are called invariant state estimation approaches, including filter-based ones and optimization-based ones. Precise Point Positioning (PPP) is a Global Navigation Satellite System (GNSS) positioning technology which can achieve high-precision positioning without commercial base stations. Visual-Inertial Odometry (VIO) combines Visual-SLAM and IMU, realizing a more robust local pose estimation than either of the two. In this paper, the invariant optimization approach has been applied to fuse PPP/INS/Visual-SLAM. The proposed positioning system in our paper is called In-P3VINS. All raw data of the In-P3VINS is modeled and optimized under an invariant factor graph framework. In particular, the carrier phase measurement is utilized by adding the phase ambiguity into the estimated states. Finally, In-P3VINS is evaluated in both simulation experiments and real-world experiments. In the simulation experiments, the accuracy and consistency of In-P3VINS are superior to the other compared methods. In the real-world experiments, In-P3VINS has the most accurate results.
Tao Li 0052, Tong Hua, Minglei Fu, Wen-An Zhang 0001, Ling Pei, Wenxian Yu, Trieu-Kien Truong
IEEE Trans. Intell. Transp. Syst.6
2024 Toward Open Vocabulary Aerial Object Detection with CLIP-Activated Student-Teacher Learning
Yan Li 0098, Weiwei Guo, Xue Yang 0005, Ning Liao, Dunyun He, Jiaqi Zhou 0017, Wenxian Yu
ECCV (86)7
2024 Stereo-LiDAR Depth Estimation with Deformable Propagation and Learned Disparity-Depth Conversion
abstract
Accurate and dense depth estimation with stereo cameras and LiDAR is an important task for automatic driving and robotic perception. While sparse hints from LiDAR points have improved cost aggregation in stereo matching, their effectiveness is limited by the low density and non-uniform distribution. To address this issue, we propose a novel stereo-LiDAR depth estimation network with Semi-Dense hint Guidance, named SDG-Depth. Our network includes a deformable propagation module for generating a semi-dense hint map and a confidence map by propagating sparse hints using a learned deformable window. These maps then guide cost aggregation in stereo matching. To reduce the triangulation error in depth recovery from disparity, especially in distant regions, we introduce a disparity-depth conversion module. Our method is both accurate and efficient. The experimental results on benchmark tests show its superior performance. Our code is available at https://github.com/SJTU-ViSYS/SDG-Depth.
Anning Hu, Wenxian Yu, Danping Zou
ICRA4
2024 Ground-Fusion: A Low-cost Ground SLAM System Robust to Corner Cases
abstract
We introduce Ground-Fusion, a low-cost sensor fusion simultaneous localization and mapping (SLAM) system for ground vehicles. Our system features efficient initialization, effective sensor anomaly detection and handling, real-time dense color mapping, and robust localization in diverse environments. We tightly integrate RGB-D images, inertial measurements, wheel odometer and GNSS signals within a factor graph to achieve accurate and reliable localization both indoors and outdoors. To ensure successful initialization, we propose an efficient strategy that comprises three different methods: stationary, visual, and dynamic, tailored to handle diverse cases. Furthermore, we develop mechanisms to detect sensor anomalies and degradation, handling them adeptly to maintain system accuracy. Our experimental results on both public and self-collected datasets demonstrate that Ground-Fusion outperforms existing low-cost SLAM systems in corner cases. We release the code and datasets at https://github.com/SJTU-ViSYS/Ground-Fusion.
Wenxian Yu, Danping Zou
ICRA4
2024 A Semantic Segmentation Method for SAR Image with Assistance of Self-Supervised Scene Classification
abstract
Unlike natural images, synthetic aperture radar (SAR) images exhibit a more scattered and uneven spatial distribution of objects, making semantic segmentation of SAR images a valuable topic of research. This paper presents a SAR image semantic segmentation method that incorporates the attention mechanism assisted by self-supervised scene classification. The self-supervised scene classification provides coarse scene classification at a higher semantic level, while the attention mechanism utilizes high-level semantic features to guide fine-grained classification of lower-level spatial structures. Overall, this approach improves the pixel-level classification performance of SAR images. We validate this method on the WHU-OPT-SAR dataset and compare its performance with previous works, providing a detailed analysis of its effectiveness.
Chen Li 0011, Zenghui Zhang, Wenxian Yu
IGARSS4
2024 ASC-RISE: Physical Information Guided Explanation of SAR ATR Models
abstract
Deep learning models have shown excellent performance in synthetic aperture radar (SAR) automatic target recognition (ATR) tasks. However, the opacity of the decision-making mechanisms within these models hamper their credibility in practical applications. Therefore, numerous explainable artificial intelligence (XAI) methods have been developed to interpret the models. Among them, the randomized input sampling for explanation (RISE) method introduces input random pixel perturbations to observe resulting changes in the output by generating saliency heatmaps. However, unlike optical images, SAR images have their unique physical properties. This paper introduces the attributed scattering center (ASC) into the RISE method, known as the ASC-RISE method guided by physical information, to explain the model. Experimental results demonstrate that the heatmaps generated by ASC-RISE effectively locate the model’s decision features and provide corresponding physical information.
Yuze Gao, Weiwei Guo, Dongying Li, Wenxian Yu
IGARSS4
2024 Few-Shot HRRP Recognition Based on The Statistical Prototypical Network
abstract
To mitigate the overfitting in the few-shot high-resolution range profile (HRRP) recognition, we introduce the Mahalanobis based statistical ProtoNet (MSP) with regularization, inspired by the prototypical network (ProtoNet). MSP leverages regularized feature covariance matrix to enhance the ProtoNet’s Euclidean distance metric based on the isotropic Gaussian distribution. Additionally, we propose a simplified MSP, the normalized statistical ProtoNet (NSP) for the faster inference of the statistical ProtoNet. Experiments demonstrate that statistical distance metrics enhance the few-shot recognition performance in scenarios with varying signal-to-noise ratios (SNR) and domain bias.
Jixi Li, Weiwei Guo, Dongying Li, Feiming Wei, Wenxian Yu
IGARSS5
2024 MAE-PixelSeg: Fine-Grained Urban Land Cover Classification with Self-Supervised Transformer
abstract
Land use land cover classification (LULC) plays a pivotal role in comprehending and managing the Earth’s surface. Numerous efforts have concentrated on deep learning methods utilizing supervised techniques, but their performance often relies on extensive well-annotated data, incurring significant time and labour expenses. This paper introduces MAE-PixelSeg, a novel self-supervised framework for LULC, leveraging a worldwide dataset of unlabeled Sentinel-2 satellite images from the Google Earth Engine (GEE) platform. MAE-PixelSeg employs the Vision Transformer (ViT) backbone, initialized by the MAE encoder, with a Shuffle Neck for high-resolution hierarchical feature map extraction and an ASPP Head for precise segmentation. Results demonstrate MAE-PixelSeg’s superiority over baseline methods, achieving a remarkable mean Intersection over Union (mIoU) of 71.07% on WorldCover and 77.43% on GID. Furthermore, we show that self-supervised pre-training yields robust feature representations, enabling commendable performance when transferred to other datasets, particularly in situations with severely limited annotated data.
Weiwei Guo, Wenxian Yu
IGARSS3
2024 MFSAF: A Plug-And-Play Module for SAR Ship Classification
abstract
This paper introduces a novel plug-and-play Multi-scale Feature Spatial Attention Fusion (MFSAF) module, aiming at enhancing the capabilities of convolutional neural networks (CNNs) in Synthetic Aperture Radar (SAR) ship classification tasks. The MFSAF module integrates spatial attention mechanisms and feature alignment strategies, providing a seamless integration into general CNNs to better capture ship features of different scales. The experimental results on the OpenSARShip2.0 and FUSARShip datasets demonstrate a significant improvement of the "baseline+MFSAF" model compared to baseline model, highlighting the effectiveness of the MFSAF module in capturing SAR ship features and its adaptability across different networks.
Nishang Xie, Mingkang Xiong, Feiming Wei, Tao Zhang 0027, Wenxian Yu
IGARSS6
2024 CA-LOSS: A Cosine Affinity Loss for Imbalanced SAR Ship Classification
abstract
To address the problem of imbalanced datasets in SAR ship classification, this paper presents a novel cosine affinity (CA) loss that enhances the Gaussian affinity (GA) loss. The CA loss focuses on the angular relationship between feature vectors, prioritizing their direction over their magnitude, which is advantageous for high-dimensional space analysis. In addition, class weights are incorporated to compute weighted distances. Importantly, the proposed CA loss does not increase the computational complexity of algorithm, nor does it lead to overfitting problems associated with data-level techniques. Through various experiments, its effectiveness has been demonstrated by achieving the highest F1 score and recall compared to other existing loss functions, highlighting its superior ability to classify minority classes in FUSARShip.
Nishang Xie, Mingkang Xiong, Feiming Wei, Tao Zhang 0027, Zhen Yang 0012, Wenxian Yu
IGARSS6
2024 Knowledge-Guided BiTCN Prototypical Network for Few-Shot Radar HRRP Target Recognition
abstract
The high-resolution range profile (HRRP) plays an important role in radar automatic target recognition (RATR) due to its rich target structure information and small data volume. However, it is difficult for HRRP to obtain sufficient data, especially for non-cooperative target, so few-shot HRRP target recognition methods have emerged. Inspired by human cognitive neuroscience, we compensate for the sample shortage by introducing semantic knowledge, focusing on key distinguishing features of categories while ignoring noise through knowledge guidance. To this end, we propose a knowledge-guided BiTCN prototypical network, where semantic knowledge is presented in the form of knowledge graph embeddings, HRRP feature extraction is completed using the temporal model BiTCN (bidirectional temporal convolutional network), and then knowledge-guided attention is used to enhance class prototypes. The experimental results demonstrate the effectiveness of our method.
Jiaqi Zhou 0017, Jixi Li, Weiwei Guo, Wenxian Yu
IGARSS5
2024 Improved Aligned Variational Autoencoders with Knowledge Graph for Generalized Zero-Shot Radar HRRP Target Recognition
abstract
With the rapid development of deep learning, significant progress has been made in radar high-resolution range profile (HRRP) target recognition methods based on deep neural networks. However, these closed set recognition methods assume that the categories of all targets are known during the model training phase, while in practice seen and unseen class targets coexist. Therefore, we propose a novel method for generalized zero-shot HRRP target recognition. Specifically, we use knowledge graph as auxiliary semantic information, utilizing two variational auto-encoders (VAEs) for cross-modal alignment and distribution alignment of semantic embeddings and HRRP features, and adopting a joint learning method of feature alignment and classification to enhance the separability of HRRP features for different categories. The experimental results demonstrate that our method is superior to existing methods.
Jiaqi Zhou 0017, Yan Li 0098, Weiwei Guo, Wenxian Yu
IGARSS5
2024 An Approach for Integrating SAR Imagery in Sea-Land Segmentation and Coastline Detection
abstract
Segmentation of Synthetic Aperture Radar (SAR) imagery constitutes the cornerstone of SAR image analysis, with sea-land segmentation in SAR images playing a crucial role in determining the precision of subsequent sea surface target detection. This study introduces an integrated approach for sea-land segmentation and coastline detection in SAR imagery, aiming to overcome the limitations posed by the traditional separation of these two tasks. In essence, the proposed method merges a segmentation module and an edge detection module, employing a hollow convolution and a global context mechanism. Additionally, the approach utilizes a cross-entropy loss function incorporating multiple losses with adaptive weighting, thereby enhancing the richness of the extracted feature information. To validate the algorithm’s efficacy, a specialized dataset for sea-land segmentation and coastline detection is constructed, utilizing the GRD data format from the Sentinel-1 satellite. Experimental outcomes show that the presented algorithm achieves scores of 0.988 and 0.981 on the Intersection over Union (IOU) metrics for sea-land segmentation, and 0.569 and 0.401 on the Optimal Dataset Scale (ODS) F1 and ODS IOU metrics for coastline detection.
Renke Zhu, Mingkang Xiong, Tao Zhang 0027, Feiming Wei, Sinong Quan, Wenxian Yu
IGARSS6
2024 Explicit Interaction for Fusion-Based Place Recognition
abstract
Fusion-based place recognition is an emerging technique jointly utilizing multi-modal perception data, to recognize previously visited places in GPS-denied scenarios for robots and autonomous vehicles. Recent fusion-based place recognition methods combine multi-modal features in implicit manners. While achieving remarkable results, they do not explicitly consider what the individual modality affords in the fusion system. Therefore, the benefit of multi-modal feature fusion may not be fully explored. In this paper, we propose a novel fusion-based network, dubbed EINet, to achieve explicit interaction of the two modalities. EINet uses LiDAR ranges to supervise more robust vision features for long time spans, and simultaneously uses camera RGB data to improve the discrimination of LiDAR point clouds. In addition, we develop a new benchmark for the place recognition task based on the nuScenes dataset. To establish this benchmark for future research with comprehensive comparisons, we introduce both supervised and self-supervised training schemes alongside evaluation protocols. We conduct extensive experiments on the proposed benchmark, and the experimental results show that our EINet exhibits better recognition performance as well as solid generalization ability compared to the state-of-the-art fusion-based place recognition approaches. Our open-source code and benchmark are released at: https://github.com/BIT-XJY/EINet.
Junyi Ma, Qi Wu 0007, Yue Wang 0020, Xieyuanli Chen, Wenxian Yu, Ling Pei
IROS7
2024 Thermal-NeRF: Neural Radiance Fields from an Infrared Camera
abstract
In recent years, Neural Radiance Fields (NeRFs) have demonstrated significant potential in encoding highly-detailed 3D geometry and environmental appearance, positioning themselves as a promising alternative to traditional explicit representation for 3D scene reconstruction. However, the predominant reliance on RGB imaging presupposes ideal lighting conditions—a premise frequently unmet in robotic applications plagued by poor lighting or visual obstructions. This limitation overlooks the capabilities of infrared (IR) cameras, which excel in low-light detection and present a robust alternative under such adverse scenarios. To tackle these issues, we introduce Thermal-NeRF, the first method that estimates a volumetric scene representation in the form of a NeRF solely from IR imaging. By leveraging a thermal mapping and structural thermal constraint derived from the thermal characteristics of IR imaging, our method showcases unparalleled proficiency in recovering NeRFs in visually degraded scenes where RGB-based methods fall short. We conduct extensive experiments to demonstrate that Thermal-NeRF can achieve superior quality compared to existing methods. Furthermore, we contribute a dataset for IR-based NeRF applications, paving the way for future research in IR NeRF reconstruction, see https://github.com/Cerf-Volant425/Thermal-NeRF.
Tianxiang Ye, Qi Wu 0007, Junyuan Deng, Liu Liu 0012, Songpengcheng Xia, Wenxian Yu, Ling Pei
IROS8
2024 Stochastic-Resonance-Networks-Enhanced Wireless Channel Parameter Estimation Approach
abstract
The estimation accuracy of conventional parameter estimation methods, including the maximum likelihood (ML) estimator and subspace-based estimation methods, diverges from the Cramer–Rao lower bound (CRLB) under low- signal-to-noise ratio (SNR) conditions. Conventional stochastic resonance (SR) technique has shown appealing weak signal improvement advantages under low SNR, but it still needs a priori information, such as the probability density functions (pdfs) of weak signal and channel noise. In this study, to address the channel parameter estimation for weak signal conditions, a novel channel parameter estimation algorithm based on dynamic stochastic resonance networks (SRNs) is introduced. Since the signal statistical properties are altered by the SRN processing, the CRLB of the wireless channel parameter estimation employing the SRN-enhanced signal is derived, and then the corresponding ML estimator is presented. Theoretical analyses show that the CRLB is lower than those from the original signal and SR-enhanced signal. Computer simulations are performed to verify the effectiveness of the theoretical CRLB expressions. Both simulation and real experimental results indicate that the proposed SRN processing approach outperforms the conventional SR processing through achieving the CRLB improvement, the ML estimation performance enhancement under low- SNR conditions, and the hardware complexity reduction.
Di He 0002, Pai Wang 0001, Wenxian Yu
IEEE Internet Things J.3
2024 Can We Trust Deep Learning Models in SAR ATR?
abstract
Deep learning has significantly enhanced the performance of automatic target recognition (ATR) in synthetic aperture radar (SAR). However, the concept of model overinterpretation, characterized by classifiers discerning strong class evidence within image regions that lack semantically salient features related to target (e.g., background clutter in SAR images), has undermined confidence in the reliability of deep learning models. Previous studies predominantly relied solely on one interpretability method to qualitatively identify the key input pixels, without assessing the efficiency of these features in decision-making process, posing a significant hurdle in evaluating the model overinterpretation. In this paper, we propose necessity-sufficiency index (NSI) to select models’ decision-making basis among all the key regions identified by multiple interpretability methods and segmentation algorithm. Furthermore, we propose weighted composition ratio statistic (WCRS) method to quantitatively analyze the model overinterpretation by incorporating the NSI as weighted average weights. The experimental results indicate that our methods are capable of accurately identifying the decision-making features and quantitatively analyzing the models’ tendency towards overinterpretation.
Yuze Gao, Weiwei Guo, Dongying Li, Wenxian Yu
IEEE Geosci. Remote. Sens. Lett.4
2024 PFDN: A Polarimetric Feature-Guided Deep Network for Dual-Polarized SAR Ship Classification
abstract
As one important application of synthetic aperture radar (SAR), ship classification attracts researchers’ attention in recent years. To improve the accuracy of ship classification in dual-polarized SAR images, we here put forward a novel polarimetric feature-guided deep network PFDN. Detailedly, a new polarimetric feature SPF (Smoothed-Polarimetric Information-Fusion) is first built through fusing both amplitudes of dual-polarized channels. Since only the amplitude information is used, SPF cannot be affected by the phase noise. Then, another two key components, i.e., the Multi-scale Feature Fusion Attention Module (MFFAM) and the Dynamic Gating Feature Fusion Mechanism (DGFFM), are further proposed to extract deeper classification features. Finally, via combing these three different modules together, PFDN is constructed. Experiments tested on the dataset OpenSARShip2.0 show that, PFDN can achieve higher accuracies (87.13% in three-class task and 65.97% in six-class task) than the other state-of-the-art (SOTA) methods.
Nishang Xie, Tao Zhang 0027, Feiming Wei, Wenxian Yu
IEEE Geosci. Remote. Sens. Lett.4
2024 RESeTNet: Residual Embedding Sequential Transformer Network for High-Repetition Pulse Deblurring in Passive Location
abstract
Pulse deblurring is a key technology for passive localization of reconnaissance signals, and the performance of pulse deblurring directly affects the accuracy of Time Difference of Arrival (TDOA) localization. Under the condition of high pulse repetition frequency, traditional pulse deblurring methods face significant challenges in terms of accuracy and efficiency. In response to the aforementioned issues, this paper proposes a high-repetition pulse deblurring method based on a residual embedding sequential transformer network (RESeTNet). Firstly, based on the ideal passive localization scenario and TDOA model, the pulse separation ambiguity problem under high-repetition frequency is analyzed. The TDOA deblurring problem is modeled as a pulse classification problem under a fixed category set. Secondly, by improving existing classification network models, the RESeTNet model and a deblurring method based on RESeTNet are proposed. Finally, model validation was conducted using simulated data from STK software. The results indicate that the proposed method can significantly improve the accuracy of pulse deblurring under high repetition frequency conditions, achieving a 97.5% unambiguous localization probability even under high repetition rate conditions of 600 kHz.
Tao Zhen, Wenxian Yu
IEEE Geosci. Remote. Sens. Lett.3
2024 A Graph Neural Network for Ship Link Prediction Based on Graph Attention Mechanism and Quaternion Embedding
abstract
In recent years, research on knowledge graphs (KGs) has exploded due to their capability of effective organization and representation of massive heterogeneous data. However, existing KGs are often incomplete and contain incorrect triples, which easily incurs a negative impact on the performance of downstream tasks. Besides, few works have been done on the construction of ship KGs. For the former, the mainstream solution is link prediction, also known as KG completion. To this end, we propose a novel method of KG embeddings (KGE) for link prediction, called QuatGAT, combining the graph attention mechanism with quaternion embeddings. Specifically, the multihead attention mechanism is employed first to obtain entity embeddings by capturing features of both entities and relations of the neighborhood. Then, to represent relations more sufficiently, we use quaternion embeddings to explicitly represent relations as well as entities. Experimental results on benchmark datasets FB15k and FB15k-237 demonstrate the superiority of QuatGAT over existing state-of-the-art methods. Moreover, in terms of ship KG construction, we also built a multimodal ship KG named MSKG. Likewise, experimental results on this dataset verify the effectiveness of QuatGAT.
Jiaqi Zhou 0017, Wenxian Yu, Siyuan Mu, Yan Li 0098
IEEE Geosci. Remote. Sens. Lett.2
2024 A Network for Merging SAR Image Sea-Land Segmentation and Coastline Detection Tasks
abstract
For the task of marine target detection in synthetic aperture radar (SAR) images, sea-land segmentation and coastline detection are often essential. Despite exciting results, many of them are still separately performed. Only a few studies have been done on the simultaneous realization of sea-land segmentation and coastline detection. To this end, this letter proposes a new network SAENet, wherein one edge enhancement module (EEM), one maximum fusion difference convolution (MaxFDC), and one multiscale spatial attention module (Multiscale SAM) are developed. In order to verify its effectiveness, we further construct one sea-land segmentation and coastline detection dataset with the Sentinel-1 ground range detected (GRD) data. The corresponding experimental results show that SAENet can reach 0.989 and 0.980 on the evaluation indexes$F1$and IOU for sea-land segmentation, and 0.577 and 0.391 on the evaluation indexes ODS$F1$and ODS IOU for coastline detection, which better accomplishes the task of sea-land segmentation and coastline detection simultaneously in comparison with other state-of-the-art (SOTA) methods.
Renke Zhu, Tao Zhang 0027, Feiming Wei, Wenxian Yu
IEEE Geosci. Remote. Sens. Lett.5
2024 TextSLAM: Visual SLAM With Semantic Planar Text Features
abstract
We propose a novel visual SLAM method that integrates text objects tightly by treating them as semantic features via fully exploring their geometric and semantic prior. The text object is modeled as a texture-rich planar patch whose semantic meaning is extracted and updated on the fly for better data association. With the full exploration of locally planar characteristics and semantic meaning of text objects, the SLAM system becomes more accurate and robust even under challenging conditions such as image blurring, large viewpoint changes, and significant illumination variations (day and night). We tested our method in various scenes with the ground truth data. The results show that integrating texture features leads to a more superior SLAM system that can match images across day and night. The reconstructed semantic 3D text map could be useful for navigation and scene understanding in robotic and mixed reality applications.
Boying Li, Danping Zou, Yuan Huang 0011, Xinghan Niu, Ling Pei, Wenxian Yu
IEEE Trans. Pattern Anal. Mach. Intell.6
2024 Dual Branch Deep Network for Ship Classification of Dual-Polarized SAR Images
abstract
Ship classification is usually a challenging task due to the small sizes of ship targets and the lack of significant differences between different categories. In terms of synthetic aperture radar (SAR) images, most existing deep learning-based methods are not designed from the angle of polarimetric characteristics to achieve ship classification. Thus, when facing the ship classification task of dual-polarized SAR images, these networks are often unsatisfactory. To cure this shortcoming, we here propose a novel dual branch deep network DBDN specifically designed for dual-polarized SAR ship classification. Our approach consists of three key modules: the image construction module ICM, the feature extraction module FEM, and the feature fusion and classifier module FFCM. In ICM, two novel pseudo RGB images are constructed for the first time, i.e., the polarimetric features-guided pseudo RGB image (PF-RGB) and the texture features-guided pseudo RGB image (TF-RGB), which can more accurately and comprehensively reflect ships’ characteristics. FEM enables the network to focus on important ship features and suppress irrelevant noise through transferred layers and designed ConvNeXt-Attention block (CNABlock), enhancing the discriminative capability of different ships. Finally, FFCM extracts and combines various ship features for classification, wherein the enhanced inverted residual block (EIRBlock) and the channel spatial attention module (CSAM) components are proposed as well. The performance of DBDN is evaluated on the OpenSARShip2.0 dataset, and experimental results show that DBDN achieves excellent performance in all evaluation metrics in comparison with some state-of-the-art (SOTA) algorithms. For example, compared to the recently proposed method DSN, DBDN further improves the accuracy by 4.74% and 4.07% in the three-class and six-class classification tasks, respectively.
Nishang Xie, Tao Zhang 0027, Weiwei Guo, Zenghui Zhang, Wenxian Yu
IEEE Trans. Geosci. Remote. Sens.5
2024 Timestamp-Supervised Wearable-Based Activity Segmentation and Recognition With Contrastive Learning and Order-Preserving Optimal Transport
abstract
Human activity recognition (HAR) with wearables is one of the serviceable technologies in ubiquitous and mobile computing applications. The sliding-window scheme is widely adopted while suffering from the multi-class windows problem. As a result, there is a growing focus on joint segmentation and recognition with deep-learning methods, aiming at simultaneously dealing with HAR and time-series segmentation issues. However, obtaining the full activity annotations of wearable data sequences is resource-intensive or time-consuming, while unsupervised methods yield poor performance. To address these challenges, we propose a novel method for joint activity segmentation and recognition with timestamp supervision, in which only a single annotated sample is needed in each activity segment. However, the limited information of sparse annotations exacerbates the gap between recognition and segmentation tasks, leading to sub-optimal model performance. Therefore, the prototypes are estimated by class-activation maps to form a sample-to-prototype contrast module for well-structured embeddings. Moreover, with the optimal transport theory, our approach generates the sample-level pseudo-labels that take advantage of unlabeled data between timestamp annotations for further performance improvement. Comprehensive experiments on four public HAR datasets demonstrate that our model trained with timestamp supervision is superior to the state-of-the-art weakly-supervised methods and achieves comparable performance to the fully-supervised approaches.
Songpengcheng Xia, Ling Pei, Wenxian Yu, Robert C. Qiu
IEEE Trans. Mob. Comput.5
2023 FeatureBooster: Boosting Feature Descriptors with a Lightweight Neural Network
abstract
We introduce a lightweight network to improve descriptors of keypoints within the same image. The network takes the original descriptors and the geometric properties of keypoints as the input, and uses an MLP-based self-boosting stage and a Transformer-based cross-boosting stage to enhance the descriptors. The boosted descriptors can be either real-valued or binary ones. We use the proposed network to boost both hand-crafted (ORB [34], SIFT [24]) and the state-of-the-art learning-based descriptors (SuperPoint [10], ALIKE [53]) and evaluate them on image matching, visual localization, and structure-from-motion tasks. The results show that our method significantly improves the performance of each task, particularly in challenging cases such as large illumination changes or repetitive patterns. Our method requires only 3.2ms on desktop GPU and 27ms on embedded GPU to process 2000 features, which is fast enough to be applied to a practical system. The code and trained weights are publicly available at github.com/SJTU-ViSYS/FeatureBooster.
Xinjiang Wang, Zeyu Liu 0001, Yu Hu 0019, Wenxian Yu, Danping Zou
CVPR5
2023 NeRF-LOAM: Neural Implicit Representation for Large-Scale Incremental LiDAR Odometry and Mapping
abstract
Simultaneously odometry and mapping using LiDAR data is an important task for mobile systems to achieve full autonomy in large-scale environments. However, most existing LiDAR-based methods prioritize tracking quality over reconstruction quality. Although the recently developed neural radiance fields (NeRF) have shown promising advances in implicit reconstruction for indoor environments, the problem of simultaneous odometry and mapping for large-scale scenarios using incremental LiDAR data remains unexplored. To bridge this gap, in this paper, we propose a novel NeRF-based LiDAR odometry and mapping approach, NeRF-LOAM, consisting of three modules neural odometry, neural mapping, and mesh reconstruction. All these modules utilize our proposed neural signed distance function, which separates LiDAR points into ground and non-ground points to reduce Z-axis drift, optimizes odometry and voxel embeddings concurrently, and in the end generates dense smooth mesh maps of the environment. Moreover, this joint optimization allows our NeRF-LOAM to be pre-trained free and exhibit strong generalization abilities when applied to different environments. Extensive evaluations on three publicly available datasets demonstrate that our approach achieves state-of-the-art odometry and mapping performance, as well as a strong generalization in large-scale environments utilizing LiDAR data. Furthermore, we perform multiple ablation studies to validate the effectiveness of our network design. The implementation of our approach will be made available at https://github.com/JunyuanDeng/NeRF-LOAM.
Junyuan Deng, Qi Wu 0007, Xieyuanli Chen, Songpengcheng Xia, Wenxian Yu, Ling Pei
ICCV7
2023 SAR Ship Detection in Range-Compressed Domain Based on LSTM Method
abstract
Most of the conventional ship detection methods based on synthetic aperture radar (SAR) intends to process the focused images, which does not take full advantages of the intermediate data in the SAR imaging process. In this paper, we introduce a new framework that treats a two-dimensional point target as multiple one-dimensional sequences in the range-compressed domain, and then employs a Long Short-Term Memory (LSTM)-based network to perform the ship detection, thus reducing the computational burden and improving efficiency significantly. To validate the effectiveness of our proposed method, we conduct experiments on real SAR data. The results demonstrate the superiority of our framework in ship detection tasks.
Yuze Gao, Dongying Li, Weiwei Guo, Wenxian Yu
IGARSS4
2023 Few-Shot Radar HRRP Recognition Based on Improved Prototypical Network
abstract
The high-resolution range profile (HRRP), with its compact vector expression and rich structural information, has become an essential part of radar automatic target recognition (RATR) systems. However, The limited training data strictly limits the recognition performance. In this paper, we proposed an HRRP recognition method for few-shot non-cooperative targets. The proposed method leverages the HRRP of simulated aerial targets as prior information and exploits the HRRP normalization and alignment to improve the generalization. In addition, we introduce an efficient feature extractor with squeeze-and-excitation attention to refine the feature map. The experiments based on simulated HRRP show that the proposed method achieves the best recognition performance under extremely few-shot conditions among conventional few-shot learning methods such as model-agnostic meta-learning (MAML) and prototypical networks (ProtoNet).
Jixi Li, Dongying Li, Wenxian Yu
IGARSS4
2023 Self-Supervised Learning Based Ship Target Classification Under Open Set Condition1
abstract
Deep learning techniques have shown promise in the field of remote sensing; however, they face challenges related to the availability of accurately labeled data and their performance in open-set cases. To overcome these limitations, this artic le presents a self-supervised learning approach for ship target classification in remote sensing images. The proposed method leverages a combination of self-supervised learning techniques applied to high-resolution optical images. Moreover, the article investigates the impact of unknown categories on classification performance and introduces a statistical extremum-based anomaly identification method to address the open-set problem. Experimental evaluations demonstrate that the proposed approach achieves state-of-the-art classification performance.
Dongying Li, Wenxian Yu
IGARSS2
2023 Ship Detection with the Nonlocal Information-Based Polarimetric Covariance Matrix
abstract
Ship plays an important role in human marine production and living activities at sea. In this paper, we design a ship detection method for polarimetric synthetic aperture radar (Pol-SAR) images. In brief, one nonlocal neighborhood polarimetric covariance matrix [NC] is first built by improving the neighborhood polarimetric covariance matrix [N] with a new similarity parameter rI. Then, the proposed method PWFNCis achieved through directly computing the polarimetric whiten filter (PWF) with [NC]. Experiments carried out on two real PolSAR datasets show that, compared to the recently proposed matrix [N], [NC] can better improve the ship detection performance of PWF.
Tao Zhang 0027, Wenxian Yu, Yonghu Zhang, Weiwei Guo
IGARSS2
2023 Occluded Target Recognition in SAR Imagery With Scattering Excitation Learning and Channel Dropout
abstract
Deep neural networks are widely used in SAR image classification and recognition, achieving state-of-the-art performance. But it remains a challenging task to recognize occluded targets. In this letter, we propose a novel robust SAR recognition method against occlusion. Specifically, we design a scattering excitation learning module that encourages the network to learn more robust features responding to the scattering centers of targets. In addition, we adopt a random feature channel dropout technique which can further improve robustness to occlusion. Our method makes the network more robust against occlusion but without any occlusion-simulated data for training. Experimental results on MSTAR dataset shows that our proposed method achieves remarkably improved robustness even under severe occlusions. Code is made available at https://github.com/koervcor/SEL-CD.
Dunyun He, Weiwei Guo, Tao Zhang 0027, Zenghui Zhang, Wenxian Yu
IEEE Geosci. Remote. Sens. Lett.5
2023 3DMAE: Joint SAR and Optical Representation Learning With Vertical Masking
abstract
The remote sensing community has shown increasingly interest in self-supervised learning for its ability to learn representations without labeled data. These representations can be easily adapted to downstream tasks through pre-training and fine-tuning. Recently, Masked Autoencoders (MAE) achieve better semantic representation by masking out a significant portion of the input image. However, the original design of MAE for RGB natural images may not be optimal for remote sensing (RS) images, which exhibit considerable variation between modalities like SAR and optical. To address this, we propose a 3D mask that enhances feature extraction along the vertical dimension. After fine-tuning, our 3DMAE model outperforms state-of-the-art contrastive and MAE-based models on BigEarthNet-MM classification and significantly reduces input data volume by at least 50% with the vertical mask, resulting in a more efficient model. Generalization experiments show a 5.9% F1-score improvement when applied to the SEN12MS dataset, which has diverse data distributions.
Limeng Zhang, Zenghui Zhang, Weiwei Guo, Tao Zhang 0027, Wenxian Yu
IEEE Geosci. Remote. Sens. Lett.5
2023 A Pose-Only Solution to Visual Reconstruction and Navigation
abstract
Visual navigation and three-dimensional (3D) scene reconstruction are essential for robotics to interact with the surrounding environment. Large-scale scenarios and computational robustness are great challenges facing the research community to achieve this goal. This paper raises a pose-only imaging geometry representation and algorithms that might help solve these challenges. The pose-only representation, equivalent to the classical multiple-view geometry, is discovered to be linearly related to camera global translations, which allows for efficient and robust camera motion estimation. As a result, the spatial feature coordinates can be analytically reconstructed and do not require nonlinear optimization. Comprehensive experiments demonstrate that the computational efficiency of recovering the scene and associated camera poses is significantly improved by 2-4 orders of magnitude.
Lilian Zhang, Yuanxin Wu, Wenxian Yu, Dewen Hu
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 Variable Length Sequential Iterable Convolutional Recurrent Network for UWB-IR Vehicle Target Recognition
abstract
A variable length sequential iterable convolutional recurrent network (VS-ICRN) is proposed in this paper, aiming at improving the vehicle target recognition ability for the Ultra-Wideband Impulse Radar (UWB-IR). Firstly, the array imaging technology is introduced into the UWB-IR, and thus a range-angle imaging method for the array UWB-IR is put forward, to simulate the array UWB-IR vehicles image under different observation conditions. Secondly, in order to make full use of both the deep features in the single image and the deep associated features between the sequence images, a VS-ICRN model is proposed, which includes three sub-modules: the image feature extraction based on the iterable convolution, the variable length sequential image associated feature extraction and the target classification, respectively. Finally, the experiment on the simulation dataset and the MSATAR is carried out to validate the effectiveness of the proposed method. The experimental results on the simulation dataset show that when SNR=-10dB, the proposed method is superior in the recognition rate to the GoogLeNet and AlexNet methods with 16% and 19%, respectively. Meanwhile, the proposed VS-ICRN method only needs 1.38% parameters quantity to achieve a comparable recognition rate as GoogLeNet on MSTAR dataset.
Lizhe Wang 0001, Wenxian Yu
IEEE Trans. Geosci. Remote. Sens.4
2023 Information Reconstruction-Based Polarimetric Covariance Matrix for PolSAR Ship Detection
abstract
In the last decades, how to detect ships with polarimetric synthetic aperture radar (PolSAR) has become one hot topic. Unfortunately, most of the existing ship detection methods cannot well detect small ships with weak backscattering. To deal with this issue, a ship detection matrix named complete polarimetric covariance matrix [CP] was recently proposed from the perspective of spatial information utilization. Although it is able to improve small ships’ target-to-clutter ratio (TCR) values, its calculation strategy still needs to be rethought due to the possible information loss of some ships. Besides, its mathematical characteristic (i.e., not positive semidefinite) also limits the successful applications of some existing polarimetric theories to it. To overcome these two drawbacks, we here develop an information reconstruction-based polarimetric covariance matrix [IC]. In brief, one new difference calculation strategy is first performed on the Sinclair matrix [$S$], so as to reconstruct its information, by which a feature vector$v$is subsequently extracted with the Lexicographic matrix basis. Then, via further performing an outer product operation on$v$, the matrix [IC] is proposed. Meanwhile, to demonstrate the effectiveness of [IC] in ship detection, two different [IC]-based intensity detectors, respectively, named SPANIC and PEDIC, are designed as well. Experiments carried out on three GF-3 PolSAR datasets show that: 1) the proposed matrix [IC] has a better performance than [CP] and the original polarimetric covariance matrix [$C$] in ship detection and 2) compared to the total power detector SPAN and geometrical perturbation-polarimetric notch filter (GP-PNF), both SPANIC and PEDIC can better detect ships, especially the small ships.
Tao Zhang 0027, Sinong Quan, Wei Wang 0099, Weiwei Guo, Zenghui Zhang, Wenxian Yu
IEEE Trans. Geosci. Remote. Sens.6
2023 A Boundary Consistency-Aware Multitask Learning Framework for Joint Activity Segmentation and Recognition With Wearable Sensors
abstract
With the development of industrial and sensing technology, sensor-based activity recognition has become a promising technology for informatics applications. However, in a typical activity recognition procedure, sensory data segmentation, usually considered a preprocess with sliding windows, rarely has been investigated and significantly affected the recognition performance. In this article, we propose a novel deep-learning method to jointly segment and recognize activities with wearable sensors. Our contributions are three-fold: First, we introduce a multistage temporal convolutional network for sample-level activity prediction to overcome the multiclass windows problem. Second, for alleviating oversegmentation errors, our model forms a multitask learning framework with a boundary prediction module to adjust the entire model’s gradients. Third, we innovatively propose a boundary consistency loss to enforce the consistency of the activity and boundary prediction. Our method shows impressive performance on three public datasets, especially achieving 16% improvement over very recently advanced competing methods with class-average F1-score on the Hospital dataset. The code of this work will be open source onhttps://github.com/xspc/Segmentation-Sensor-based-HAR.
Songpengcheng Xia, Ling Pei, Wenxian Yu, Robert C. Qiu
IEEE Trans. Ind. Informatics4
2022 Multi-level Contrast Network for Wearables-based Joint Activity Segmentation and Recognition
abstract
Human activity recognition (HAR) with wearables is promising research that can be widely adopted in many smart healthcare applications. In recent years, the deep learning-based HAR models have achieved impressive recognition performance. However, most HAR algorithms are susceptible to the multi-class windows problem that is essential yet rarely exploited. In this paper, we propose to relieve this challenging problem by introducing the segmentation technology into HAR, yielding joint activity segmentation and recognition. Especially, we introduce the Multi-Stage Temporal Convolutional Network (MS-TCN) architecture for sample-level activity prediction to joint segment and recognize the activity sequence. Furthermore, to enhance the robustness of HAR against the inter-class similarity and intra-class heterogeneity, a multi-level contrastive loss, containing the sample-level and segment-level contrast, has been proposed to learn a well-structured embedding space for better activity segmentation and recognition performance. Finally, with comprehensive experiments, we verify the effectiveness of the proposed method on two public HAR datasets, achieving significant improvements in the various evaluation metrics.
Songpengcheng Xia, Ling Pei, Wenxian Yu, Robert C. Qiu
GLOBECOM4
2022 Discovering Novel Categories in Sar Images in Open Set Conditions
abstract
In this paper, we deal with the issue of discovering data of novel categories for Synthetic Aperture Radar (SAR) images under open-set conditions. The traditional SAR image classification methods are trained under the closed-set setting where all categories in testing data are seen in training data. It does not always meet the requirements of the real SAR imagery interpretation applications. With a labelled SAR image dataset, we propose a multi-stage approach to effectively pick out images belonging to new classes in another unlabelled dataset and then cluster them into correct number of novel categories. To do so, our pipeline is composed of three major steps: (1) train a powerful feature extractor leveraging both the labelled and unlabelled dataset by semi-supervised inference; (2) identify the unknown data by openset detection; (3) cluster these unknown data based on the features generated by the extractor to discover novel categories. The proposed method is validated on a Sentinel-1 SAR image dataset OpenSARUrban [1].
Liu Dai, Weiwei Guo, Zenghui Zhang, Wenxian Yu
IGARSS4
2022 Heterogeneous Image Classification with Multi-Stage Conditional Adversarial Domain Adaptation Between SAR and Optical Imagery
abstract
In this paper, we deal with the problem of heterogeneous image classifiers transferring between SAR and optical im-agery through a novel multi-stage domain adaptation technique. The problem of transferring the optical image classifier to SAR and vice-versa is of practical importance because it allows us to leverage plenty of labelled data in the source do-main for the target domain task, but gains little attention. Be-cause there is a drastic distribution-gap between both the opti-cal and SAR imaging modalities, it is non-trivial to apply do-main adaption directly. We propose a multi-stage adversarial feature alignment procedure that firstly performs global ad-versarial feature alignment and then a class-conditional adver-sarial feature alignment is conducted to further enable class-discriminative feature adaption. The proposed method is val-idated on the SEN12MS dataset, and some discussions are provided about heterogeneous domain adaption between SAR and optical imagery.
Weiwei Guo, Zenghui Zhang, Wenxian Yu
IGARSS4
2022 Explainable Analysis of Deep Learning Methods for Sar Image Classification
abstract
Deep learning methods exhibit outstanding performance in synthetic aperture radar (SAR) image interpretation tasks. However, these are black box models that limit the com-prehension of their predictions. Therefore, to meet this challenge, we have utilized explainable artificial intelli-gence (XAI) methods for the SAR image classification task. Specifically, we trained state-of-the-art convolutional neural networks for each polarization format on OpenSARUrban dataset and then investigate eight explanation methods to analyze the predictions of the CNN classifiers of SAR images. These XAI methods are also evaluated qualitatively and quantitatively which shows that Occlusion achieves the most reliable interpretation performance in terms of Max-Sensitivity but with a low-resolution explanation heatmap. The explanation results provide some insights into the in-ternal mechanism of black-box decisions for SAR image classification.
Shenghan Su, Ziteng Cui, Weiwei Guo, Zenghui Zhang, Wenxian Yu
IGARSS5
2022 Exploring Similarity in Polarization: Contrastive Learning with Siamese Networks for Ship Classification in Sentinel-1 SAR Images
abstract
In this paper, we focus on synthetic aperture radar automatic target recognition for ships, and modify the Simple Siamese (SimSiam) framework, a contrastive self-supervised representation learning method, to improve ship classification accuracy. We design a novel sampling method that takes polarization information into account, in addition to the image augmentation based positive pair sampling method that is commonly used in contrastive learning approaches. The results of the experiments show that positive pairs of VH and VV polarized images can provide complementary information about ship targets to strengthen the classifiers. Besides, different similarity measurement functions are analyzed in the experiments.
Weiwei Guo, Zenghui Zhang, Wenxian Yu
IGARSS5
2022 Polsar Ship Detection with the Sub-Aperture Technology
abstract
Polarimetric synthetic aperture radar (PolSAR) designed to obtain the polarimetric information of scenes is a crucial tool for microwave remote sensing. Recently, a complete polarimetric covariance difference matrix [CP] was built to detect ships of PolSAR image. Along this work, this paper extends its application to the spectrum domain. Briefly speaking, four sub-aperture images are first separated from the original PolSAR data. Then, four different power values corresponding to the [CP] matrices of these sub-aperture images are respectively calculated. At last, via multiplying these values together, a PolSAR ship detector named MPS (Multiplicative Polarimetric SPAN) is proposed. The experiment carried out on one real PolSAR image demonstrates that, compared to traditional power detectors SPAN and$SPAN_{CP,}$MPS holds a better ability to detect small ships.
Tao Zhang 0027, Zenghui Zhang, Weiwei Guo, Huilin Xiong, Wenxian Yu
IGARSS5
2022 Multiple Embeddings Contrastive Pretraining for Remote Sensing Image Classification
abstract
This letter focuses on remote sensing image interpretation and aims to promote the use of contrastive self-supervised learning in varied applications of remote sensing image classification. The proposed method is a contrastive self-supervised pre-training framework that encourages the network to learn image representations by comparing image embeddings extracted by different encoders and predictors. Experiments were carried out on a variety of remote sensing image datasets to determine the efficacy of the proposed method for classification tasks. Results show that the proposed framework exploits the capabilities of encoders and outperforms the supervised learning method in terms of classification accuracy. Besides, it takes a few pre-training epochs to find a suboptimal initialization of network weights, and the pre-trained encoders use a little training data to get outstanding classification results, which shows the time and data efficiency of the proposed framework. Code is available at https://github.com/yinxu98/MECo.
Weiwei Guo, Zenghui Zhang, Wenxian Yu
IEEE Geosci. Remote. Sens. Lett.4
2022 Spatial Singularity-Exponent-Domain Multiresolution Imaging-Based SAR Ship Target Detection Method
abstract
A novel spatial singularity-exponent-domain multiresolution imaging (SSMRI) method is proposed in this article. It is aimed at improving the image processing ability for the single-channel synthetic aperture radar (SAR), thereby acquiring the robust image feature for SAR target detection at the very low SNR. Combining the 2-D singularity power spectrum (SPS) and the 2-D pseudo-Wigner–Ville distribution (PWVD), the SSMRI method is derived. Consequently, the single-channel SAR images are transformed to obtain spatial multiresolution SAR images with respect to the singularity exponent. Furthermore, the separability of the SAR image target in the space of the singularity exponent is analyzed, and the separability theorem in the sense of SSMRI is proved. In particular, under the background of GWN and fractal noise, the separability of SAR image targets based on SSMRI processing is studied. In addition, the stability and robustness of SSMRI-SAR feature extraction are demonstrated. On this basis, a maritime SAR ship target detection method based on SSMRI is put forward in the extremely low SNR condition. The experiment results on the SAR-Ship-Dataset indicate that the proposed method is superior in performance to the traditional CFAR or 2-D-SPS method. Especially when SNR = −30 dB, the detection performance with more than 99.4% can be achieved.
Wenxian Yu, Trieu-Kien Truong
IEEE Trans. Geosci. Remote. Sens.3
2022 Corrections to "Region-Based Polarimetric Covariance Difference Matrix for PolSAR Ship Detection"
abstract
In the above article[1], the average TCR values inTable IIwere incorrectly presented. The corrected table is given here:
Tao Zhang 0027, Wei Wang 0099, Sinong Quan, Huizhang Yang, Huilin Xiong, Zenghui Zhang, Wenxian Yu
IEEE Trans. Geosci. Remote. Sens.7
2022 Region-Based Polarimetric Covariance Difference Matrix for PolSAR Ship Detection
abstract
To more effectively detect small ships, in this article, a novel region-based polarimetric covariance difference matrix [RP] is put forward, which mainly consists of two stages. Briefly speaking, in the first stage, a new pixel representation way is proposed to depict the spatial characteristics of pixel, through which the difference information related to pixel’s local region is calculated as well. In the second stage, the global region difference information of pixel is computed. Finally, we construct [RP] via fusing these two different kinds of information together with a balance factor$c$. Meanwhile, considering that the backscattering energy of ships is useful for ship detection, a new intensity-driven polarimetric notch filter (ID-PNFRP) is also derived from [RP]. Three different datasets are adopted to evaluate the effectiveness of [RP] and ID-PNFRP. Experimental results show that: 1) compared with the polarimetric covariance matrix [$C$] and the polarimetric covariance difference matrix [$P$], [RP] is more suitable for ship detection and 2) compared with the original geometrical perturbation-polarimetric notch filter (GP-PNF) and the total power detector SPAN, the proposed method ID-PNFRPcan better detect small ships with greater figure of merit (FoM) and target-to-clutter ratio (TCR) values.
Tao Zhang 0027, Wei Wang 0099, Sinong Quan, Huizhang Yang, Huilin Xiong, Zenghui Zhang, Wenxian Yu
IEEE Trans. Geosci. Remote. Sens.7
2021 StructDepth: Leveraging the structural regularities for self-supervised indoor depth estimation
abstract
Self-supervised monocular depth estimation has achieved impressive performance on outdoor datasets. Its performance however degrades notably in indoor environments because of the lack of textures. Without rich textures, the photometric consistency is too weak to train a good depth network. Inspired by the early works on indoor modeling, we leverage the structural regularities exhibited in indoor scenes, to train a better depth network. Specifically, we adopt two extra supervisory signals for self-supervised training: 1) the Manhattan normal constraint and 2) the co-planar constraint. The Manhattan normal constraint enforces the major surfaces (the floor, ceiling, and walls) to be aligned with dominant directions. The co-planar constraint states that the 3D points be well fitted by a plane if they are located within the same planar region. To generate the supervisory signals, we adopt two components to classify the major surface normal into dominant directions and detect the planar regions on the fly during training. As the predicted depth becomes more accurate after more training epochs, the supervisory signals also improve and in turn feedback to obtain a better depth model. Through extensive experiments on indoor benchmark datasets, the results show that our network outperforms the state-of-the-art methods. The source code is available at https://github.com/SJTU-ViSYS/StructDepth.
Boying Li, Yuan Huang 0011, Zeyu Liu 0001, Danping Zou, Wenxian Yu
ICCV5
2021 Direct Oriented Ship Localization Regression in Remote Sensing Imagery with Curriculum Learning
abstract
Accurate and efficient ship detection in remote sensing images still remains a challenging task due to the large variations of scales, orientations and distributions. In this paper, we propose an anchor-free ship detector that directly regresses ship localization parameters, offering a simpler pipeline over the previous methods. The detection network is then trained in a multi-task fashion which contains not only the ship center-point maps and oriented bounding boxes but the ship masks. Instead of fixing the weights among the multiple task losses, we adopt a curriculum learning strategy which gradually adapts the loss weights during the training process so that the network can learn the discriminative ship features at the early stage and obtain more localization information while training continues. Experimental results on real dataset demonstrate the effectiveness and efficiency of our proposed method.
Weiwei Guo, Huiyuan Chen, Zenghui Zhang, Yanhua Zhang, Wenxian Yu
IGARSS5
2021 Can We Evaluate the Distinguishability of the Opensarurban Dataset?
abstract
In Synthetic Aperture Radar (SAR) image classification tasks, the performance depends on both the classifier and the dataset itself. However, in comparison with plenty of SAR classification methods, there is little work aimed at analyzing the distinguishability of the dataset. In the classification dataset, some classes are semantically different but their distinguishability is low, the classes are hard to be classified especially in some more practical cases that there are unknown classes without supervision exist. Referring to open set recognition (OSR), in this paper, we proposed the SAR Distinguishability Analysor (SAR-DA) to evaluate the distinguishability of the OpenSARUrban dataset. By modeling each class as a multivariate Gaussian distribution in latent space, SAR-DA can not only classify the classes having been seen in training phase, but also can recognize unknown samples if a test sample is out of each known distribution. Each class in OpenSARUr-ban is set unknown in turn, then we apply the SAR-DA on the split dataset in OSR and supervised setting. The distinguishability can be reflected by the unknown recognition recall rate. The experimental results show that the unknown recognition recall rate in OSR setting significantly decreased compared with those in supervised setting, indicating that even though the classes in OpenSARUrban are semantically different from each other, the latent distributions of some classes are quite similar and hard to be classified, thus these classes are of low distinguishability.
Ning Liao, Mihai Datcu, Zenghui Zhang, Weiwei Guo, Wenxian Yu
IGARSS5
2021 Self-Supervised Auto-Encoding Multi-Transformations for Airplane Classification
abstract
In this paper, we present a self-supervised learning method of Auto-Encoding Multi-Transformations (AEMT) for airplane classification. In this method, the image features are learned in an unsupervised way by simultaneously estimating multiple image transformations from the features of original and transformed images instead of reconstructing the input images. Besides, we propose two structure variants of the AEMT method: composite and parallel modes of which the former transforms the images in a composite fashion while the latter does it in parallel. The experimental results demonstrate that the proposed method outperforms the state-of-the-art self-supervised learning methods for the airplane classification task.
Ziteng Cui, Weiwei Guo, Zenghui Zhang, Wenxian Yu
IGARSS5
2021 Robust Initialization of Multi-camera SLAM with Limited View Overlaps and Inaccurate Extrinsic Calibration
abstract
This paper proposes a robust initialization method for a multi-camera visual SLAM system where cameras have only a limited common field of views and inaccurate extrinsic calibration. The limited common field of views leads to only a few common features that can be matched between cameras. Inaccurate extrinsic poses, caused by vibrations or misplacement of cameras after offline calibration, make it even harder for triangulating the seed 3D points to initialize the SLAM system successfully. Instead of taking the extrinsic parameters as constants for feature matching and 3D point triangulation as most multi-camera systems did, we propose to take the inaccurate extrinsic poses as soft constraints to accommodate the calibration errors. Our initialization method consists of two stages by matching across different cameras and between two key frames. Both stages involve optimizing the cost functions that contain the extrinsic pose priors from inaccurate calibration parameters. By incorporating those soft pose constraints, we may avoid false feature matching and triangulation caused by inaccurate extrinsic parameters, while keeping solution space limited when only a few feature correspondences exist. The results in real-world tests show that such a simple solution can improve the success rate of SLAM initialization notably, even when the pose priors from the offline calibration differ significantly from the real ones.
Ang Li 0029, Danping Zou, Wenxian Yu
IROS3
2021 MARS: Mixed Virtual and Real Wearable Sensors for Human Activity Recognition With Multidomain Deep Learning Model
abstract
Together with the rapid development of the Internet of Things, human activity recognition (HAR) using wearable inertial measurement units (IMUs) becomes a promising technology for many research areas. Recently, deep-learning-based methods pave a new way of understanding and performing analysis of the complex data in the HAR system. However, the performance of these methods is mostly based on the quality and quantity of the collected data. In this article, we innovatively propose to build a large data set based on virtual IMUs and then address technical issues by introducing a multiple-domain deep learning framework consisting of three technical parts. In the first part, we propose to learn the single-frame human activity from the noisy IMU data with hybrid convolutional neural networks in the semisupervised form. For the second part, the extracted data features are fused according to the principle of uncertainty-aware consistency, which reduces the uncertainty by weighting the importance of the features. The transfer learning is performed in the last part based on the newly released archive of motion capture as surface shapes data set, containing abundant synthetic human poses, which enhances the variety and diversity of the training data set and is beneficial for the process of training and feature transfer in the proposed method. The efficiency and effectiveness of the proposed method have been demonstrated in the real deep inertial poser data set. The experimental results show that the proposed methods can surprisingly converge within a few iterations and outperform all competing methods.
Ling Pei, Songpengcheng Xia, Fanyi Xiao, Qi Wu 0007, Wenxian Yu, Robert C. Qiu
IEEE Internet Things J.6
2021 A Novel Wireless Localization Approach Using Twice Receiving Array Spectra Fusions and ASSR Networks
abstract
Path fading and non-line-of-sight (NLOS) signals constitute serious problems in the wireless localization process. These problems cause unpredictable degradation in the localization precision. In this paper, a novel wireless localization approach, which is based on twice receiving array signal spectra fusions and asymmetric second-order stochastic resonance (ASSR) networks, is proposed. By combining and repetitively processing the above two techniques, the receiving signal-to-noise ratio (SNR) can be enhanced. Additionally, the receiving array signal without the line-of-sight (LOS) component can be determined and removed from the spectra fusion process. The theoretical analyses presented verify the unbiasedness and asymptotic efficiency of the proposed twice receiving spectra fusion approach. Computer simulations demonstrate that the fused spectra can significantly improve the wireless localization precision compared with conventional and up-to-date localization methods, especially under low SNR conditions.
Di He 0002, Ling Pei, Xin Chen 0017, Ling-ge Jiang, Jiaqing Qu, Wenxian Yu
IEEE Trans. Commun.7
2021 Radar Sea Clutter Reconstruction Based on Statistical Singularity Power Spectrum and Instantaneous Singularity Exponents Distribution
abstract
The modeling and reconstruction of radar sea clutter (RSC) based on fractal theory is studied in this article. With theoretical derivation and quantitative analysis, the radar clutter reconstruction model based on singularity power spectrum (SPS) and instantaneous singularity exponent (ISE) distribution is proposed. The proposed ISE-SPS model consists of two coupling channels: power measure channel and dimension measure channel. The former is controlled by SPS to characterize the power distribution of signals in the SE domain, while the latter is determined by ISE function to characterize the statistical distribution of ISE. The multiscale wavelet coefficients are estimated based on the coupling of the two channels to reconstruct the RSC. The proposed ISE-SPS method is tested on RSC data in different sea state, from the ice multiparameter imaging X-band (IPIX) radar. Experimental and simulation based on the IPIX radar data set indicates that the proposed method performs better than the N-partitioned random multiplicative model (NRMM) method and SPS-multifractal spectrum (MFS) method, and exceeds the traditional methods by nearly 35% in terms of correlation coefficient and normalized mean square error of MFS and SPS.
Wenxian Yu
IEEE Trans. Geosci. Remote. Sens.3
2021 Singularity-Exponent-Domain Image Feature Transform
abstract
Combining the generalized fractal theory and the time-frequency distribution, the image feature decomposition in the singularity exponent domain is studied in this paper. With the theoretical derivation and quantitative analysis, the singularity-exponent-domain image feature transform (SIFT) method is proposed to analyze and process images from new feature dimensions. If one derives from the generalized fractal characteristics of the image, the two-dimensional frequency variables of the 2D time-frequency transform of the image can be used to estimate the two-dimensional singularity power spectrum (SPS) in the space dimension. As a consequence, it leads to the SPS distribution of the original image in the spatial domain, i.e., SIFT images. Based on the SIFT, the feature transform images with different singularity exponent and feature curves of singularity power spectrum with respect to different physical regions can thus be obtained. The SIFT is rigorously derived from the 2D-SPS and the Pseudo Wigner-Ville distribution (PWVD). In addition, the feature images based on the SIFT is proved to be the SNR independence in the GWN background. In order to validate the effectiveness of feature extraction, the proposed methodology is tested on the breast ultrasound images, the visual images, and the synthetic aperture radar (SAR) images. Furthermore, the SAR target detection method based on the SIFT images is proposed, and the experiment results indicate that the proposed algorithm is superior in performance to the traditional CFAR or 2D-SPS method. In fact, this new SIFT is promising to provide a technical approach for image feature extraction, target detection, and recognition.
Wenxian Yu, Trieu-Kien Truong
IEEE Trans. Image Process.3
2020 TextSLAM: Visual SLAM with Planar Text Features
abstract
We propose to integrate text objects in man-made scenes tightly into the visual SLAM pipeline. The key idea of our novel text-based visual SLAM is to treat each detected text as a planar feature which is rich of textures and semantic meanings. The text feature is compactly represented by three parameters and integrated into visual SLAM by adopting the illumination-invariant photometric error. We also describe important details involved in implementing a full pipeline of text-based visual SLAM. To our best knowledge, this is the first visual SLAM method tightly coupled with the text features. We tested our method in both indoor and outdoor environments. The results show that with text features, the visual SLAM system becomes more robust and produces much more accurate 3D text maps that could be useful for navigation and scene understanding in robotic or augmented reality applications.
Boying Li, Danping Zou, Daniele Sartori, Ling Pei, Wenxian Yu
ICRA5
2020 Photovoltaic Panel Construction Change Monitoring Based on LSTM Models
abstract
Satellite and Aerial Image Time Series contain tremendous amounts of information of ground targets in space and time and have been commonly used in temporal pattern analysis tasks of ground targets. It has a wide range of applications including environment monitoring, urban planning, hazard assessment, etc. PV (photovoltaic) field construction monitoring is a new topic with increasing attentions. In this paper, we proposed a simple yet effective network based on LSTM (long short term memory) to detect PV field construction event. We build image time-series dataset from Sentinel-2 data of an Egypt PV field under construction to monitor the project progress. Compared with clustering model and CNNs (convolutional neural networks), our network achieves the better accuracy. Although low resolution and mislabeled pixels limit the accuracy cap, the simple network is still robust and easy to transfer for other applications.
Liuliang Chen, Weiwei Guo, Zeyu Liu 0001, Zenghui Zhang, Wenxian Yu
IGARSS5
2020 Ellipse-FCN: Oil Tanks Detection from Remote Sensing Images with Fully Convolution Network
abstract
Oil is an essential asset for every country, and plays a key role in world trade system. The detection of oil tanks is a very important task for both military and commerce. Recently, researchers have shown an increasing interest in oil tanks detection in remote sensing imagery. However, the previous works almost used the methods of circle detection, but the real oil tanks in remote sensing imagery are more close to ellipses. In this paper, we propose an oil tanks detector base on an U-shape Fully Convolutional Network(FCN) in optical remote sensing images. The structure of our network consists of three parts: feature extraction part, feature merge part and the output layer. The output layer consists of two output branches, one branch is a score map branch, which generates confidence score to indicate the region of oil tanks at pixel wise, and the other ends up with several channels which regress the ellipse geometric parameters (center, horizontal axis and vertical axis). In addition, we also design a novel loss function adapted to our network. The experimental results conducted on our dataset collected from Google Earth show that this method achieves promising performance on oil tanks detection in terms of both efficiency and accuracy in high-resolution optical remote sensing images.
Ziteng Cui, Weiwei Guo, Zenghui Zhang, Huiyuan Chen, Wenxian Yu
IGARSS5
2020 Reducing the Receiving Array Complexity By Using the Parallel Stochastic Resonance System
abstract
Nowadays the array signal processing has become a widely used technique in various applications. For example, it can be applied in the full-duplex jamming receiver and multiantenna jammer in wireless networks to improve the physical layer security 1, wireless signal localization in the near-field of an antenna array 2, phase enhancement for the low-angle estimation 3, direction-of-arrival (DoA) estimation in the nonuniform linear arrays 4, and so on. However, the theoretical and simulation performances in many applications using the array signal processing are seriously rely on the complexity or the antenna number of the signal receiving array. Here we show that the corresponding performance of the array signal processing can be enhanced by introducing the parallel stochastic resonance system (PSRS) in each branch or each antenna of the receiving array structure, especially under low signal-to-noise ratio (SNR) circumstance. We found that the output SNR of the receiving signal after the PSRS processing has been monotonically enhanced to a relatively high level with the increasing of number of parallel processing units in each antenna. And based on this result, it can be used to reduce the receiving array complexity. In other words, the array structure introducing the PSRS with small antenna number can also reach or even exceed the application performance of those array structures with more antenna number. Our results demonstrate an example of wireless signal DoA estimation error performance by using the proposed PSRS, which reveals that even a 4-antenna array structure with PSRS structure can outperformance a 48-antenna array structure without PSRS structure under very low SNR. We believe that this kind of technology could be applied very widely in many areas related to the array signal processing not only in reducing the array complexity, but also in signal quality or signal SNR enhancement, and so on.
Di He 0002, Fusheng Zhu, Wenxian Yu
IGARSS4
2020 Iron ORE Region Segmentation Using High-Resolution Remote Sensing Images Based on Res-U-Net
abstract
Deep learning has found many applications in high-resolution remote sensing image interpretation. In this study, an image analysis system is presented, consisting of image segmentation and mineral volume change estimation. In this system, a revised U-Net structure, called Res-U-Net, is proposed by combining U-Net and residual structure for image segmentation. Experiments are performed on the collected high-resolution remote sensing images, which were annotated with iron ore positive, iron ore negative, and background, and the results demonstrate the superiority of the proposed Res-U-Net over other image segmentation methods. Our proposed Res-U-Net outperforms the traditional U-Net by achieving pixel-wise accuracy of 92% and mean intersection over union (mIOU) 86%, as well as faster frame rate of 35 FPS on test dataset.
Noman Mustafa, Juanping Zhao, Zeyu Liu 0001, Zenghui Zhang, Wenxian Yu
IGARSS5
2020 Effective Moving Target Deception Jamming Against Multichannel SAR-GMTI Based on Multiple Jammers
abstract
This letter presents a novel scheme for moving target deception jamming (MTDJ) against multichannel synthetic aperture radar (SAR)-ground moving target indication (GMTI) based on multiple jammers. The multichannel SAR signal models for the real and false moving targets are established first. Then, the proposed MTDJ method based on multiple jammers is derived according to the principle that the interferometric phase matches the across-track velocity of the moving target. To ensure efficient deception jamming and to resolve the accurate scattering coefficient modulated in each jammer, the interferometric phase of the false moving target is decomposed into two parts, one is the term that matches the across-track velocity, while the other is the additional term generated by each jammer. Hence, the expected modulation coefficient in each jammer can be determined based on the additional interference phase cancellation. Finally, the proposed MTDJ is realized by generating the jamming signals in each jammer via the efficient deception jamming technique. The simulation results demonstrate the effectiveness of the proposed MTDJ method.
Qingyang Sun, Ting Shu 0003, Mang Tang, Kai-Bor Yu, Wenxian Yu
IEEE Geosci. Remote. Sens. Lett.5
2020 GIS-Supervised Building Extraction With Label Noise-Adaptive Fully Convolutional Neural Network
abstract
Automatic building extraction from aerial or satellite images is a dense pixel prediction task for many applications. It demands a large number of clean label data to train a deep neural network for building extraction. But it is labor expensive to collect such pixel-wise annotated data manually. Fortunately, the building footprint data of geographic information system (GIS) maps provide a cheap way of generating building label data, but these labels are imperfect due to misalignment between the GIS maps and images. In this letter, we consider the task of learning a deep neural network to label images pixel-wise from such noisy label data for building extraction. To this end, we propose a general label noise-adaptive (NA) neural network framework consisting of a base network followed by an additional probability transition modular (PTM) which is introduced to capture the relationship between the true label and the noisy label. The parameters of the PTM can be estimated as part of the training process of the whole network by the off-the-shelf backpropagation algorithm. We conduct experiments on real-world data set to demonstrate that our proposed PTM can better handle noisy labels and improve the performance of convolutional neural networks (CNNs) trained on the noisy label data generated by GIS maps for building extraction. The experimental results indicate that being armed with our proposed PTM for fully CNN, it provides a promising solution to reduce manual annotation effort for the labor-expensive object extraction tasks from remote sensing images.
Zenghui Zhang, Weiwei Guo, Mingjie Li 0006, Wenxian Yu
IEEE Geosci. Remote. Sens. Lett.4
2020 Sparse nested array with aperture extension for high accuracy angle estimation
Jin He 0001, Zenghui Zhang, Ting Shu 0003, Wenxian Yu
Signal Process.4
2020 Full-Aperture Azimuth Spatial-Variant Autofocus Based on Contrast Maximization for Highly Squinted Synthetic Aperture Radar
abstract
Generally, high-resolution imaging for highly squinted synthetic aperture radar (SAR) data is a difficult problem due to large range migration. Thus, when trying to solve this nontrivial problem, an azimuth-variant Doppler will arise, thereby leading to phase errors containing the azimuth spatial-variant (ASV) component. In this article, we analyze the characteristics of highly squinted SAR data and propose a new full-aperture ASV phase error autofocus algorithm. In this new algorithm, the accurate and suitable phase error signal model for highly squinted SAR data is derived. Moreover, the closed-form solution of the relationship between a distorted image and a focused image is also explicitly revealed. Furthermore, in this newly proposed algorithm, an accurate estimation of nonlinear ASV phase error is established based on the maximum contrast of the SAR imagery. In addition, an iterative gradient-based solver is introduced. The advantage of this new method provides a simple yet effective approach while being able to eliminate the ASV phase errors. More importantly, the accuracy of this new method using the full-aperture data is independent of SAR imaging algorithms. As a result, the proposed new method can be easily embedded in many existing imaging algorithms to produce focused imagery. Finally, two real highly squinted SAR data sets are provided to validate the advantages of our algorithm.
Darong Huang 0001, Xinrong Guo, Zenghui Zhang, Wenxian Yu, Trieu-Kien Truong
IEEE Trans. Geosci. Remote. Sens.4
2019 Sub-Image Blocks Based Joint Sparse Reconstruction Algorithm for Multi-Pass SAR Images Feature Enhancement
abstract
With the development of multi-pass synthetic aperture radar (SAR) imaging, feature enhancement with multiple SAR images has become an important research topic due to its distinct advantage of redundant scene information. The existing SAR image feature enhancement algorithms mainly focus on single SAR image processing by means of ℓq(02,1-norm minimization approach is introduced to achieve the purpose of feature enhancement. In addition, two dimensional fast iterative shrinkage thresholding algorithm (2D-FISTA) is utilized to increase the computational efficiency. Experimental results are illustrated to validate that the proposed method can improve SAR image performance significantly in terms of denoising and sidelobes suppression while maintaining higher computational efficiency.
Chunxiao Wu, Zenghui Zhang, Wenxian Yu
IGARSS4
2019 The Same Range Line Cells Based Fast Two-Dimensional Compressive Sensing For Airborne MIMO Array SAR 3-D Imaging
abstract
Airborne multiple-input multiple-output array synthetic aperture radar (MIMO array SAR) can be used to directly obtain the three-dimensional (3-D) imagery of the illuminated region with a single track. Different sparse reconstruction algorithms within the framework of compressive sensing (CS) have been conceived to reconstruct the cross-track signal because of its inherent spatial sparsity. However, the computational complexity and the number of two-dimensional (2-D) SAR images needed for sparse reconstruction of the existing algorithms are usually high. To overcome this problem, the same range line cells based 2-D real-valued CS is proposed in this paper. In this new algorithm, the original signal model is transformed from the complex domain to the real domain by means of unitary transformation. Finally, airborne MIMO array SAR real experimental results are illustrated to validate the prominent advantages of the proposed method.
Chunxiao Wu, Zenghui Zhang, Longyong Chen, Wenxian Yu
IGARSS4
2019 Learning Physical Scattering Patterns from PolSAR Images by Using Complex-Valued CNN
abstract
Full-polarimetric synthetic aperture radar (SAR) images have the ability to provide physical patterns of the earth observation, no more than geometric information. In order to learn physical patterns from non-full-polarimetric SAR images, a complex-valued CNN is leveraged to learn a model containing physical parameters. The parameters are learned from the original complex scattering matrix of full-polarimetric SAR images and they can be adopted to extract physical patterns from non-full-polarimetric SAR images. Cloude and Pottier's H-α division, as the annotation principle, is computed by way of coherence matrix. We perform experiments on (German Aerospace Center) DLR's full-polarimetric, airborne F-SAR data, demonstrating that extracting physical patterns from non-full-polarimetric images is feasible. The comparative results illustrate that: 1) The best physical categoric patterns can be extracted from HV and VH polarimetric images in general, while performance from HH and VV polarimetric images are limited; 2) Cross-polarimetric SAR images have greater ability for surface and volume scattering, while co-polarimetric ones are better for multiple scattering extraction.
Juanping Zhao, Mihai Datcu, Zenghui Zhang, Huilin Xiong, Wenxian Yu
IGARSS5
2019 A coupled convolutional neural network for small and densely clustered ship detection in SAR images
Juanping Zhao, Weiwei Guo, Zenghui Zhang, Wenxian Yu
Sci. China Inf. Sci.4
2019 Fast 3-D Imaging Algorithm Based on Unitary Transformation and Real-Valued Sparse Representation for MIMO Array SAR
abstract
Multiple-input multiple-output (MIMO) array synthetic aperture radar (SAR) with array antennas distributed along the cross-track direction can obtain 3-D scene information of the surveillance region. However, the cross-track resolution is unacceptable due to the length limitation of the MIMO antenna array. The superresolution algorithms within the framework of compressive sensing (CS) have been introduced to recover the cross-track signal because of its inherent spatial sparsity. The existing sparse recovery algorithms for 3-D SAR are attempted to find the sparse solution in the complex domain directly, which requires a very high computational complexity. To overcome this problem, a new fast 3-D imaging algorithm based on real-valued sparse representation is proposed in this paper. In this new algorithm, unitary transformation can be employed to transform the sparse signal recovery model of uniform/nonuniform MIMO array SAR from the complex domain to the real domain. Thus, a real-valued reweighted 12,1-norm minimization model is established. In addition, a modification of the fast iterative shrinkage-thresholding algorithm (FISTA) is used to reconstruct the 3-D image for further improving the computational efficiency. Moreover, the theoretical analysis of computational complexity of the proposed algorithm is derived when compared with an existing complex domain algorithm. Finally, numerical simulations and MIMO array SAR real experimental results are illustrated to validate that the proposed algorithm can reduce the computational complexity significantly in terms of CPU time while still maintaining the inherent advantages of superresolution and robustness against the noise.
Chunxiao Wu, Zenghui Zhang, Xingdong Liang, Longyong Chen, Wenxian Yu, Trieu-Kien Truong
IEEE Trans. Geosci. Remote. Sens.5
2019 SAR Target Detection in Complex Scene Based on 2-D Singularity Power Spectrum Analysis
abstract
The synthetic aperture radar (SAR) target detection method in the background of a complex scene and extremely low signal-to-noise ratio (EL-SNR) is studied in this paper. With theoretical derivation and analysis, the singularity power spectrum (SPS) is developed into two-dimensional SPS (2D-SPS). Developed from the 2-D Holder exponent and the SPS, the 2DSPS can be applied for the singularity power analysis of images. Furthermore, a novel SAR target detection method based on 2D-SPS is proposed. The proposed methodology is tested on sea clutters, both with and without SAR ship target, from the OpenSAR data sets. The experimental results indicate that the SAR target detection based on 2D-SPS performs better than conventional multifractal spectrum (MFS) and constant false alarm rate (CFAR) methods, and can achieve more than 95% and 98% detection probability, respectively, under EL-SNR with SNR = -30 dB, SNR = -10 dB, and Pf= 10-2, and almost 100% detection probability for weak and multiple targets within sea clutters. The proposed method can be applied to target detection under the background of fractal noise and provides a reference for target detection in related fields.
Liyang Zhu, Junye Li 0002, Wenxian Yu
IEEE Trans. Geosci. Remote. Sens.5
2019 Cross Correlation Singularity Power Spectrum Theory and Application in Radar Target Detection Within Sea Clutters
abstract
The cross correlation power spectrum of multiple signal sequences in the singularity domain is studied in this paper. With theoretical derivation and quantitative analysis, the cross correlation singularity power spectrum (CSPS) distribution theory is proposed. Developed from correlation function (CF), spectrum CF (SCF), singularity power spectrum (SPS), and multifractal cross correlation analysis, the CSPS can be applied for the correlation analysis of multiple fractal time series. In this paper, the CSPS is rigorously derived based on SPS and SCF, and is verified with classical multifractal time series. Furthermore, a target detection method based on the proposed CSPS method is also proposed. The proposed methodology is tested on sea clutters, both with and without target, from the Ice Multiparameter Imaging X-Band radar data set. The simulation results indicate that the target detection based on CSPS performs better than conventional multifractal spectrum methods, and can achieve almost 100% detection probability of detecting low-observable targets within sea clutters.
Caiping Xi, Dongying Li, Wenxian Yu
IEEE Trans. Geosci. Remote. Sens.4
2019 Ship Detection From PolSAR Imagery Using the Complete Polarimetric Covariance Difference Matrix
abstract
In this paper, we proposed a complete polarimetric covariance difference matrix [CP]-based algorithm for ship detection in polarimetric synthetic aperture radar (PolSAR) imagery. To calculate [C P], we first developed a scheme to reflect the polarimetric scattering differences between ship pixel (SP) and its neighboring pixels (ISPs) and, then, dividedly accumulated the amplitude and phase differences between SP and ISPs. Compared to the polarimetric covariance difference matrix [P] developed in our earlier work, [C P] effectively overcomes the drawback of the lack of the phase information in [P]. To demonstrate the effectiveness of the proposed algorithm, we applied the [CP]-based ship detection algorithm to four PolSAR data sets, including one UAVSAR L-band data set with 21 ships, two AIRSAR L-band data sets with 11 and 22 ships, respectively, and one Radarsat-2 C-band data set with 8 ships. Experimental results show that: (1) the proposed algorithm can effectively detect ships with high target-to-clutter ratio (TCR) values and (2) [C P] has a better performance than the traditional polarimetric covariance matrix [C] and [P] on ship detection. To be more specific, the average TCR value of the proposed algorithm (23.86 dB) is 6.07 and 7.47 dB higher than PNFC(i.e., the geometrical perturbation-polarimetric notch filter) and RSC(i.e., the reflection symmetry method), respectively.
Tao Zhang 0027, Jinsheng Ji, Xiaofeng Li 0001, Wenxian Yu, Huilin Xiong
IEEE Trans. Geosci. Remote. Sens.4
2019 Contrastive-Regulated CNN in the Complex Domain: A Method to Learn Physical Scattering Signatures From Flexible PolSAR Images
abstract
Single- and dual-polarimetric synthetic aperture radar (SAR) images provide very limited capabilities to interpret physical radar signatures. For generality and simplicity, we call single-polarimetric, dual-polarimetric, and fully polarimetric SAR (PolSAR) images flexible PolSAR images. In order to sufficiently extract physical scattering signatures from this kind of data and explore the potentials of different polarization modes on this task, this paper proposes a contrastive-regulated convolutional neural network (CNN) in the complex domain, attempting to learn a physically interpretable deep learning model directly from the original backscattered data. To achieve a better deep model containing physically interpretable parameters, the objective cost is compared to and selected from several commonly used loss functions in the complex form. The required ground-truth labels are generated automatically according to Cloude and Pottier's H-alpha division plane, which significantly reduces intensive labor cost and transfers this method to an unsupervised learning mechanism. The boundaries between different scattering signatures, however, sometimes show an erroneous separation. With the aim of aggregating intra-class instances and alienating inter-class instances, meanwhile, a complex-valued contrastive regularization term is computed mathematically and is added to the objective cost by a tradeoff factor. Moreover, data augmentation is applied to relieve the side effects caused by data imbalance. Finally, we performed experiments on German Aerospace Center's (DLR)'s L-band, high-resolution (HR), and airborne F-SAR data. Our results demonstrate the possibility of extracting physical scattering signatures from flexible PolSAR images. Physically interpretable potentials of SAR images with different polarization modes are analyzed, and we conclude with physical signature identification.
Juanping Zhao, Mihai Datcu, Zenghui Zhang, Huilin Xiong, Wenxian Yu
IEEE Trans. Geosci. Remote. Sens.5
2019 StructVIO: Visual-Inertial Odometry With Structural Regularity of Man-Made Environments
abstract
In this paper, we propose a novel visual-inertial odometry (VIO) approach that adopts structural regularity in man-made environments. Instead of using Manhattan world assumption, we use Atlanta world model to describe such regularity. An Atlanta world is a world that contains multiple local Manhattan worlds with different heading directions. Each local Manhattan world is detected on the fly, and their headings are gradually refined by the state estimator when new observations are received. With full exploration of structural lines that aligned with each local Manhattan worlds, our VIO method becomes more accurate and robust, as well as more flexible to different kinds of complex man-made environments. Through benchmark tests and real-world tests, the results show that the proposed approach outperforms existing visual-inertial systems in large-scale man-made environments.
Danping Zou, Yuanxin Wu, Ling Pei, Haibin Ling, Wenxian Yu
IEEE Trans. Robotics5
2019 Collaborative visual SLAM for multiple agents: A brief survey
abstract
This article presents a brief survey to visual simultaneous localization and mapping (SLAM) systems applied to multiple independently moving agents, such as a team of ground or aerial vehicles, a group of users holding augmented or virtual reality devices. Such visual SLAM system, name as collaborative visual SLAM, is different from a typical visual SLAM deployed on a single agent in that information is exchanged or shared among different agents to achieve better robustness, efficiency, and accuracy. We review the representative works on this topic proposed in the past ten years and describe the key components involved in designing such a system including collaborative pose estimation and mapping tasks, as well as the emerging topic of decentralized architecture. We believe this brief survey could be helpful to someone who are working on this topic or developing multi-agent applications, particularly micro-aerial vehicle swarm or collaborative augmented/virtual reality.
Danping Zou, Wenxian Yu
Virtual Real. Intell. Hardw.3
2018 A Ship Detector Based on the Improved Polarimetric Covariance Difference Matrix
abstract
Polarimetric Synthetic Aperture Radar data has been widely used for ship detection. In our earlier study, based on the differences between ship pixels and their surrounding background pixels, we designed a polarimetric covariance difference matrix (PCDM) to detect ships. Inadequately, the phase information of scattering differences is not included in PCD-M. Aiming at this deficiency, here, we present an improved PCDM matrix (IPCDM). Then an IPCDM-based ship detector is further proposed. To demonstrate the effectiveness of the method, two full polarimetric datasets are adopted. In comparing with other methods, we find that the result of our method is better.
Tao Zhang 0027, Yifang Ban, Huilin Xiong, Wenxian Yu
IGARSS4
2018 A Study of Boundary Layer Rolls Under Various Storm Conditions
abstract
The marine atmospheric boundary layer (MABL) roll plays an important role in the turbulent exchange of momentum, sensible heat, and moisture throughout the boundary layer of tropical cyclones. Hence, rolls are believed to be closely related to storm development and intensification. In this study, based on the RADARSAT-2 dataset including various tropical cyclone (TC) intensities, the roll characteristics are retrieved from synthetic aperture radar (SAR) images via fast Fourier transform (FFT). We investigate the roll wavelengths at variance of TC intensities and found that the roll wavelengths are related to the distance with respect to TC center and TC intensities. This study is promising to bring roll-induced effects into hurricane forecasting model.
Lanqing Huang, Xiaofeng Li 0001, Bin Liu 0019, Jun A. Zhang, Dongliang Shen, Zenghui Zhang, Wenxian Yu
IGARSS7
2018 A SAR Cross-Pol Correlation Sea Surface Wind Speed Study
abstract
A new study based on the complex and amplitude cross-pol correlation is made. It exploits Sentinel-1 Single Look Complex (SLC) SAR and reference HSCAT wind data. The behavior of these two polarimetric features vs wind speed is analyzed. Low-to-moderate and high wind speed regimes are considered. The study is focused on three subsets characterized by different incident angle ranges. The study shows that in all cases the two polarimetric features are sensitive to wind speed and better behave with respect to$S$vv.
Lanqing Huang, Maurizio Migliaccio, Ferdinando Nunziata, Valeria Corcione, Zenghui Zhang, Wenxian Yu
IGARSS6
2018 Rotated Region Based Fully Convolutional Network for Ship Detection
abstract
Ship detection from high-resolution optical remote sensing images has been a prevalent domain in recent years. Unlike objects in natural images, ships of interest can be anywhere in optical remote sensing images with multi-scale and multi-oriented which makes it more different to be detected. In this paper, we propose a novel method based on the fully convolutional network to detect ships. Our method has three important components: 1) we design a network merging different levels of feature map to fuse multi-scale information. Determining the existence of large ship require features from deep layers in the network, while predicting rotated bounding box enclosing small ships need shallow layers information; 2) The network can be trained end-to-end to generate score maps which indicates the confidence score for the ship region of interest in pixel-wise level through all locations and scaled of an image; 3) We design a rotated bounding box regression model to localize the ships. The experimental results on our dataset collected from Google Earth has demonstrated our proposed method achieves promising performance on ship detection in terms of both efficiency and accuracy in high-resolution optical remote sensing images.
Mingjie Li 0006, Weiwei Guo, Zenghui Zhang, Wenxian Yu, Tao Zhang 0027
IGARSS4
2018 A Revisited Approach to Lateral Acceleration Modeling for Quadrotor UAVs State Estimation
abstract
Quadrotor state estimation generally relies on the vehicle aerodynamics modeling to achieve improved performance. In this paper the effects of the rotors angular speeds on the quadrotor drag, and therefore on the lateral accelerations, are investigated. While these effects are usually disregarded, we analyze their modeling starting from the Blade Element Theory and flight test data. Two lateral acceleration formulations are proposed. They are adopted within a velocity and attitude state estimator and validated in real-world flights. The EKF-based estimator fuses measurements from low-cost sensors present in the majority of quadrotors (IMU, magnetometer, ultrasonic sensor, optical flow) with the accelerations of the vehicle predicted from the revisited models. Experimental results show the benefits of adopting these innovative models in the estimator when compared with the existing modeling approach.
Daniele Sartori, Danping Wou, Ling Pei, Wenxian Yu
IROS4
2018 Toward Arbitrary-Oriented Ship Detection With Rotated Region Proposal and Discrimination Networks
abstract
Ship detection from remote sensing images can provide important information for maritime reconnaissance and surveillance and is also a challenging task. Although previous detection methods including some advanced ones based on deep convolutional neural network expertize in detecting horizontal or nearly horizontal targets, they cannot give satisfying detection results for arbitrary-oriented ship detection. In this letter, we introduce a novel ship detection system that can detect arbitrary-oriented ships. In this method, a rotated region proposal networks (R2PN) is proposed to generate multiorientated proposals with ship orientation angle information. In R2PN, the orientation angles of bounding boxes are also regressed to make the inclined ship region proposals generated more accurately. For ship discrimination, a rotated region of interest pooling layer is adopted in the following classification subnetwork to extract discriminative features from such inclined candidate regions. The proposed whole ship detection system can be trained end to end. Experimental results conducted on our rotated ship data set and HRSD2016 benchmark demonstrate that our proposed method outperforms state-of-the-art approaches for the arbitrary-oriented ship detection task.
Zenghui Zhang, Weiwei Guo, Shengnan Zhu, Wenxian Yu
IEEE Geosci. Remote. Sens. Lett.4
2018 Ship Size Extraction for Sentinel-1 Images Based on Dual-Polarization Fusion and Nonlinear Regression: Push Error Under One Pixel
abstract
In this paper, we present a method of ship size extraction for Sentinel-1 synthetic aperture radar (SAR) images, which is composed of the image processing stage and the regression stage. In order to achieve extraction with high accuracy, considering the data characteristics of Sentinel-1 images, we propose to use the dual-polarization fusion and the nonlinear regression with the gradient boosting. The experiments and analyses on a relatively large data set show that: 1) compared with the existing and related studies, the proposed method achieves an improved performance. The extraction errors are pushed under one pixel, and they are 4.66% (8.80 m) and 7.01% (2.17 m) for length and width, respectively; 2) the dual-polarization information fusion does improve the size extraction accuracy; and 3) the nonlinear regression does exploit the relationship between the influential factors and the size parameters and provide a better performance than the linear regression. The experimental results verify that the proposed design is suitable for ship size extraction in Sentinel-1 SAR images.
Boying Li, Bin Liu 0019, Weiwei Guo, Zenghui Zhang, Wenxian Yu
IEEE Trans. Geosci. Remote. Sens.5
2017 Superpixel generation for SAR images based on DBSCAN clustering and probabilistic patch-based similarity
abstract
In this paper, we propose a superpixel generation method for synthetic aperture radar (SAR) images by using the density-based spatial clustering of applications with noise (DBSCAN) algorithm. The pixels is firstly grouped to generate initial superpixels by using probabilistic patch-based (PPB) dissimilarity. Then, small clusters are combined into their neighbor superpixels to get final results through a distance measurement defined by statistical models. Experiments on simulated data sets exhibit high boundary adherence of the generated superpixels and demonstrate the availability and efficiency of the proposed method.
Weiwei Guo, Zenghui Zhang, Wenxian Yu
IGARSS5
2017 Preliminary evaluation of vessel detectability for Sentinel-1 SAR data
abstract
Performance of ship detection is influenced by synthetic aperture radar (SAR) imaging characteristics and environmental conditions. In this paper, aiming at evaluating vessel detectability for Sentinel-1 SAR data, a model based on a large-scale Sentinel-1A vessel chips database is established. The model sensitivity is analyzed by simulation data. In the experiment, by inputting the parameters of imaging characteristics (incidence angle, polarization, spatial resolution) and environmental conditions (wind speed, wind direction, sea state) of a specific Sentinel-1 image, the minimum detectable vessel length can be estimated. Further validations demonstrate the availability of the estimated minimum detectable vessel length.
Lanqing Huang, Weiwei Guo, Zenghui Zhang, Wenxian Yu
IGARSS5
2017 Charaterization of densely arrayed targets patterns in high resolution SAR images: A study case in the Davis-Monthan air force base
abstract
This paper presents a new algorithm for recognizing patterns of densely arrayed targets in high resolution (HR) SAR images, serving the increasing demand for target detection and recognition in HR/VHR SAR images. The novelty of our work is to formulate the problem of pattern extarction as a jigsaw puzzle with similar target patches. Compared to existing work, our algorithm has multiple advantages, including: 1) extracting densely arrayed targets region from full scale SAR images automaticly; 2) predicting accurate displacement vector between neighboring similar patches; 3) synthetising an uniform target pattern with similar patches. These advantages are demonstratced by results of pattern extraction from a study case in the Davis-Monthan air force base.
Zeyu Liu 0001, Weiwei Guo, Zenghui Zhang, Wenxian Yu
IGARSS5
2017 Preliminary exploration of SAR image land cover classification with noisy labels
abstract
Synthetic Aperture Radar (SAR) image land cover classification is an important task in SAR image interpretation. Supervised learning, such as Convolutional Neural Network (CNN), demands instances which are accurately labeled. However, a large amount of accurately labeled SAR images are difficult to produce. In this paper, a Probability Transition CNN (PTCNN) is proposed for patch-level SAR image land cover classification with noisy labels. Firstly, deep features are extracted by a CNN model, followed by a probabilistic transition model, where true labels are treated as hidden variables and the posterior probabilities of true labels are transferred into their noisy versions. The whole network is trained with Caffe in a uniform fashion and a land cover database is used to produce noisy labels, which are randomly chosen with various proportions. Experimental results demonstrate that the proposed PTCNN model is robust to noise and gives a promising classification performance. Therefore, the PTCNN model may lower the standards for the quality of image labels, and shows its availability in practical applications.
Juanping Zhao, Weiwei Guo, Zenghui Zhang, Wenxian Yu, Shiyong Cui
IGARSS5
2017 An aerodynamic model-aided state estimator for multi-rotor UAVs
abstract
A robust state estimator is presented by fusing the aerodynamic model of multi-rotor UAVs with measurements from optical flow and other low-cost sensors such as IMU, magnetometer, and ultrasonic sensor. Due to the particular aerodynamics of multi-rotor UAVs, the body velocity in the rotor plane is able to be measured by the accelerometer. We therefore propose a novel state estimator by fully exploring the characteristic of aerodynamics of multi-rotor UAV. Our state estimator is fast and easy to be implemented. We have tested our estimator with different platforms in different scenes. Experimental results show that our estimator performs robustly in low light conditions where existing methods usually fail.
Rongzhi Wang, Danping Zou, Ling Pei, Wenxian Yu
IROS6
2017 Circulate shifted OFDM chirp waveform diversity design with digital beamforming for MIMO SAR
Shangwen Liu, Zenghui Zhang, Wenxian Yu
Sci. China Inf. Sci.3
2017 Learning the Conformal Transformation Kernel for Image Recognition
abstract
In this paper, we present a multiclass data classifier, denoted by optimal conformal transformation kernel (OCTK), based on learning a specific kernel model, the CTK, and utilize it in two types of image recognition tasks, namely, face recognition and object categorization. We show that the learned CTK can lead to a desirable spatial geometry change in mapping data from the input space to the feature space, so that the local spatial geometry of the heterogeneous regions is magnified to favor a more refined distinguishing, while that of the homogeneous regions is compressed to neglect or suppress the intraclass variations. This nature of the learned CTK is of great benefit in image recognition, since in image recognition we always have to face a challenge that the images to be classified are with a large intraclass diversity and interclass similarity. Experiments on face recognition and object categorization show that the proposed OCTK classifier achieves the best or second best recognition result compared with that of the state-of-the-art classifiers, no matter what kind of feature or feature representation is used. In computational efficiency, the OCTK classifier can perform significantly faster than the linear support vector machine classifier (linear LIBSVM) can.
Huilin Xiong, Wenxian Yu, Xin Yang 0007, M. N. S. Swamy 0001, Qiuze Yu
IEEE Trans. Neural Networks Learn. Syst.2
2016 Robust sparse recovery for compressive sensing in impulsive noise using ℓp-norm model fitting
abstract
This work considers the robust sparse recovery problem in compressive sensing (CS) in the presence of impulsive measurement noise. We propose a robust formulation for sparse recovery using the generalized lp-norm with 0 < p < 2 as the metric for the residual error under l1-norm regularization. An alternative direction method (ADM) has been proposed to solve this formulation efficiently. Moreover, a smoothing strategy has been used to derive a convergent method for the nonconvex case of p < 1. The convergence conditions of the proposed algorithm for both the convex and nonconvex cases have been provided. Numerical simulations demonstrated that the new algorithm can achieve state-of-the-art robust performance in highly impulsive noise.
Fei Wen 0005, Yipeng Liu 0001, Robert C. Qiu, Wenxian Yu
ICASSP5
2016 SAR image classification based on CRFs with object structure priors
abstract
Fine-scale classification in form of object extraction or segmentation for high resolution SAR images is a challenging task due to the existing local noises, object deformation and part missing. A novel SAR classification method based on CRFs which combines low-level features, label context and object structure priors is presented in this paper. Local label pattern is proposed in this paper to model the object structures by measuring the local label configuration on the grid layer of SAR images. We build a new CRFs model with label context and object structure priors for image classification. Besides, we adopt Mean Field approximation for efficient inference of our CRFs model. This work intends to implement an efficient classification framework by integrating high-level label context and object priors and apply it to fine-scale object extraction of SAR images. The framework demonstrates good performance in both accuracy and efficiency for object extraction or segmentation of simulated images and high resolution SAR images.
Yongke Ding, Weiwei Guo, Juanping Zhao, Weidong Xiang, Zenghui Zhang, Wenxian Yu
IGARSS7
2016 Fast topology preserving PolSAR image superpixel segmentation
abstract
In this paper, we propose a fast PolSAR image superpixel segmentation method. This method takes a simple coarse-to-fine optimization technique to minimize a Markov-Random-Field (MRF) like energy function which integrates the Pol- SAR image statistic, spatial position and boundary smoothing. It updates boundary of superpixels staring with a large block level and iterates down to the final pixel level. We demonstrate the performance of our approach both on the synthetic and real full polarimetric images , showing that our proposed approach can achieve significantly faster convergence than SLIC method, and make a good compromise between accuracy and computation speed.
Weiwei Guo, Zenghui Zhang, Juanping Zhao, Wenxian Yu
IGARSS4
2016 A damped Newton variational inversion for synthetic aperture radar wind retrieval
abstract
The variational inversion for synthetic aperture radar (SAR) wind retrieval can take errors of all sources involved into account, but the complexity of ascertaining errors of wind vectors is high and the iteration process is very time-consuming. In this paper, we modify the decomposition of wind vectors into speed and direction, and adopt a damped Newton method (DNVAR) to solve the cost function, which is based on inexact line search condition. Experimental results show that DNVAR can effectively reduce background wind vector errors, and the average number of iterations for DNVAR descends greatly. For practical applications, when the background wind speed is higher than 10 m/s, the accuracy of DNVAR is higher than direct SAR wind retrieval (DIRECT), otherwise, DIRECT performs better.
Zhuhui Jiang, Weidong Xiang, Fangjie Yu, Wenxian Yu
IGARSS5
2016 Object detection capability evaluation for SAR image
abstract
The existing SAR image quality assessment method could not be effectively used for assessing the performance of object detection. Thus, it is difficult to select SAR images and corresponding detection algorithms for SAR object detection. By analyzing the relationship between object detection results and basic image quality indicators, this paper studies the image quality assessment for image object detection. Based on the concept of “application suitability”, basic quality indicators including radiometric resolution, spatial resolution, PSLR and ISLR are integrated into a single indicator called Detection Index, which is able to comprehensively evaluate the degree to which SAR image is suitable for object detection tasks. Experimental results on aircraft detection with single scene show the effectiveness of the proposed model for SAR image capability evaluation in object detection applications.
Zheyuan Wang, Fangjie Yu, Wenxian Yu, Zhuhui Jiang, Yongke Ding
IGARSS4
2016 Ship detection based on the power of the Radarsat-2 polarimetric data
abstract
Ship target detection using PolSAR data has been an active research area and many algorithms have been developed in recent years. In this paper, we present a new method based on the difference between the ship pixels and background pixels, using a model similar to LBP (Local Binary Pattern). After that, the polarimetric signature method, namely, the SPAN (total power) detector, is used to detect ships. We adopt one Radarsat-2 data set with four-look processing for experiment, which was obtained in the Strait of Gibraltar ocean area. In comparing with other methods, we find that the result of our method is better than other detectors.
Tao Zhang 0027, Zhen Yang 0012, Huilin Xiong, Wenxian Yu
IGARSS4
2016 Convolutional Neural Network for SAR image classification at patch level
abstract
Convolutional Neural Network (CNN) has attracted much attention for feature learning and image classification, mostly related to close range photography. As a benchmark work, we trained a relatively large CNN to classify SAR image patches into five different categories, where the image patches tiled and annotated from a typical TerraSAR-X spotlight scene of Wuhan, China. The neural network designed in this paper consists of seven layers, including one input layer, two convolutional layers where each followed by a max-pooling layer, as well as two fully-connected layers with a final five-class softmax. Using the toolkit caffe, we achieved the training and testing accuracy of 85.7% and 85.6% respectively, which is considerably better than the traditional feature extraction and classification based SVM method and shows great potential of CNN used for SAR image interpretation. In order to accelerate the training process, a very efficient GPU implementation was employed.
Juanping Zhao, Weiwei Guo, Shiyong Cui, Zenghui Zhang, Wenxian Yu
IGARSS5
2016 Multidimensional Scaling-Based TDOA Localization Scheme Using an Auxiliary Line
abstract
This work deals with source localization with time-difference-of-arrival (TDOA) measurements in two-dimensional (2-D) scenarios. Although the celebrated two-step weighted least squares (2WLS) method is quite successful, its drawback lies in an ill-conditioning problem when the sensor array is quasi-linear. This work presents a multidimensional scaling (MDS)-based localization scheme. Based on the subspace analysis of the scalar product matrix, an auxiliary line is defined in the plane, close to the global minimizer of the cost function. Then, the minimizer on the auxiliary line is found as the estimation of the source position. Simulations show that the proposed scheme achieves high localization accuracy for all kinds of sensor arrays including quasi-linear arrays.
Wuyang Jiang, Ling Pei, Wenxian Yu
IEEE Signal Process. Lett.4
2015 Region-based L0 gradient minimization for PolSAR image segmentation
abstract
In order that global, dominant, and complete outlines of land covers are delineated, in this paper, we propose a regularized L0gradient minimization method which is specially developed for segmenting polarimetric synthetic aperture radar (PolSAR) images, and present a region-based stepwise design to implement it. The performance of the proposed method is tested and analyzed on two experimental data sets, with visual presentation as well as numerical evaluation. They both confirm that the proposed method achieves its principal goal and demonstrate its availability and advantage as a pre-processing step for PolSAR image interpretation chains.
Bin Liu 0019, Zenghui Zhang, Xingzhao Liu, Wenxian Yu
IGARSS4
2015 Distributed Compressed Sensing off the Grid
abstract
This letter investigates the joint recovery of a frequency-sparse signal ensemble sharing a common frequency-sparse component from the collection of their compressed measurements. Unlike conventional arts in compressed sensing, the frequencies follow an off-the-grid formulation and are continuously valued in [0, 1]. As an extension of atomic norm, the concatenated atomic norm minimization approach is proposed to handle the exact recovery of signals, which is reformulated as a computationally tractable semidefinite program. The optimality of the proposed approach is characterized using a dual certificate. Numerical experiments are performed to illustrate the effectiveness of the proposed approach and its advantage over separate recovery.
Zhenqi Lu, Rendong Ying, Sumxin Jiang, Wenxian Yu
IEEE Signal Process. Lett.5
2015 Representation and Spatially Adaptive Segmentation for PolSAR Images Based on Wedgelet Analysis
abstract
It is believed that it is essential to take the spatial adaptivity into the segmentation method for polarimetric synthetic aperture radar (PolSAR) images. The size and shape of each segment and the strength of the relationship of neighboring pixels need to depend on the local spatial complexity of the scene. The wedgelet framework provides a promising analysis tool for spatial information. The major advantage of the wedgelet analysis is that it captures the geometrical structure of images at multiple scales, with the local spatial complexity taken into consideration. Hence, in this paper, we propose a wedgelet approximation and analysis framework specially designed for PolSAR data. Based on this framework, a spatially adaptive representation and segmentation method is constructed and presented. It mainly consists of three parts: first, the multiscale wedgelet decomposition is applied to the PolSAR image, and the local geometrical information is captured in an optimal way; then, the image is segmented in a spatially adaptive manner by the multiscale wedgelet representation in the form of the regularized optimization, which keeps a balance between the approximation and parsimony of the representation; the final part is the spatial-complexity-adaptive segmentation refinement based on the Wishart Markov random field model. The performance of the proposed method is presented and analyzed on two experimental data sets, with visual presentation and numerical evaluation. It is also compared with an existing and theoretically well-founded segmentation method. The experiments and results demonstrate the availability and advantage of the proposed method.
Bin Liu 0019, Zenghui Zhang, Xingzhao Liu, Wenxian Yu
IEEE Trans. Geosci. Remote. Sens.4
2014 Spectral compressive sensing with model selection
abstract
The performance of existing approaches to the recovery of frequency-sparse signals from compressed measurements is limited by the coherence of required sparsity dictionaries and the discretization of frequency parameter space. In this paper, we adopt a parametric joint recovery-estimation method based on model selection in spectral compressive sensing. Numerical experiments show that our approach outperforms most state-of-the-art spectral CS recovery approaches in fidelity, tolerance to noise and computation efficiency.
Zhenqi Lu, Rendong Ying, Sumxin Jiang, Zenghui Zhang, Wenxian Yu
ICASSP6
2014 SAR target segmentation based on shape prior
abstract
It's very challenging to interpret the SAR image due to the speckle noise and target's heterogeneity, etc. And the use of image information alone often leads to poor segmentation. However, the introduction of shape prior into the active contours has proved to be very effective to segment the target. In this study, we define the shape of SAR target and introduce it into our active contours. To avoid the irregularity of level set evolution, a novel potential function is introduced to regularize the level set function. And the energy of gray intensity helps to locate the target while the energy of target's shape regularizes the target's contour. Experiments on both ship and building segmentation prove the validity of the proposal algorithm.
Yun Han, Wenxian Yu
IGARSS3
2014 Multi-temporal superpixel generation for high resolution SAR image analysis
abstract
In this paper, a multi-temporal (MT) superpixel generation method is proposed. MT superpixels are generated based on a novel edge extraction method designed for high resolution (HR) synthetic aperture radar (SAR) images and using MT spatial information in an optimization way. Numerical evaluation and comparison on simulated and real data sets demonstrate its availability for MT HR SAR image analysis.
Bin Liu 0019, Zenghui Zhang, Xingzhao Liu, Wenxian Yu
IGARSS5
2014 Discrete frequency-coding waveform design may not be ok for extended targets in MIMO SAR
abstract
As the need for radar applications has become more urgent, conventional radar performance does not meet the current demand for observation. In recent years, the MIMO radar is considered to have great potential. MIMO SAR can get more phase center for imaging, interference, GMTI, or any other application. Waveform design is the key point of MIMO SAR, and it's also the current biggest bottlenecks of MIMO SAR implementation. We prove that the DFCW/OFDM signal can hardly meet the required performance for MIMO SAR while the experiments focus on extended targets imaging.
Shangwen Liu, Zenghui Zhang, Wenxian Yu
IGARSS4
2014 Edge Extraction for Polarimetric SAR Images Using Degenerate Filter With Weighted Maximum Likelihood Estimation
abstract
The classic region-based filter for edge extraction for polarimetric synthetic aperture radar images is theoretically founded and efficient. However, in practical use, its performance is limited because the assumption of independence and identical distribution is often not met, particularly in heterogeneous areas. In this letter, we present a degenerate filter design integrated with the weighted maximum likelihood estimation to overcome this limitation. The performance of the proposed methodology is presented and analyzed on both simulated and real experimental data sets using visual presentation, as well as numerical evaluation and comparison with the classic method. They both demonstrate the availability and advantage of the proposed method.
Bin Liu 0019, Zenghui Zhang, Xingzhao Liu, Wenxian Yu
IEEE Geosci. Remote. Sens. Lett.4
2013 Scene scattering descriptor for urban classification in very high resolution SAR images
abstract
A novel urban land-use/land-cover classification framework for very high resolution SAR images which combines scene description and scene-based mapping is presented in this paper. We use textural features for scene region description, including the traditional GLCM feature and the new scene scattering descriptor. The scene scattering descriptor is proposed in this paper to model the scattering behavior and context appearance of urban scenes. This work intends to implement an efficient description of urban land covers from the urban backscattering behavior perspective and apply it to scene-based mapping of urban areas. The framework demonstrates good performance in both accuracy and efficiency for urban scene interpretation of very high resolution SAR images.
Yongke Ding, Lizhong Qiu, Pinglv Yang, Zeming Zhou, Wenxian Yu
IGARSS6
2013 Characterization and extraction of building layovers in urban areas using high resolution SAR imagery
abstract
In this paper, we present an image processing chain that can interpret high resolution synthetic aperture radar (SAR) imagery for building layover characterization and extraction in urban areas. It is composed of three main parts - generation of hint areas, generation of superpixels, and optimized cut of layovers via superpixel merging. The proposed framework is complete, and flexibly integrates necessary information, both area and boundary, for building layover extraction; the experimental results show that its performance is promising.
Bin Liu 0019, Florence Tupin, Xingzhao Liu, Wenxian Yu
IGARSS4
2013 Sparse imaging based on SAR complex image domain
abstract
In this paper, a novel synthetic aperture radar (SAR) imaging algorithm based on two-dimensional (2-D) compressed sensing (CS) is presented. First, a complex image is generated by applying a traditional imaging method to raw echoes. Then, the basis pursuit (BP) method is applied to a CS framework to obtain the scattering coefficients of target points from the complex image. Since CS is applied jointly in range and azimuth, this algorithm utilizes the sparsity of the target scene fully. Experimental results using the real field data demonstrate the effectiveness of our algorithm.
Wentao Lv, Junfeng Wang 0001, Lizhong Qiu, Wenxian Yu
IGARSS4
2013 Change detection method using a new difference image for remote sensing images
abstract
In this paper, a change detection method using a new difference image is proposed. The spatial position of pixels in neighborhood is considered to generate the new difference image, where the unchanged class pixels can be preferably modeled by Gaussian distribution. The Kittler and Illingworth threshold algorithm is used to perform the final change detection results. Experimental results on multi-temporal remote sensing images confirm the effectiveness of the proposed method.
Lizhong Qiu, Yongke Ding, Heping Lu, Wenxian Yu
IGARSS6
2013 Radiometric-spatial analysis for ship detection in high resolution synthetic aperture radar images
abstract
Ship detection is an important application in remote sensing and Earth observation since last century. Properties of ship targets changed greatly as observation accuracy of SAR sensors is strongly increased. Ship targets become objects detailed in structure by contrast to point targets in lower resolution SAR images. In this paper, we present a radiometric-spatial analysis (RSA) based method for ship detection in high resolution SAR images.
Bin Liu 0019, Wenxian Yu
IGARSS6
2013 Sparse representation based pan-sharpening
abstract
In this paper, we propose a novel pan-sharpening method which combines classical component substitution with recently developed sparse representation. We explore the sparse representations of multispectral and panchromatic images through two dictionaries which are trained to have the same sparse representations for each high-resolution and low-resolution image patch pair. The merging procedure is implemented in sparse domain. In order to avoid spectral distortion, partial replacement is used to extract details. At the same time, the introducing of dictionary pair also reduces the distortion caused by interpolating the MS at the initialization of the fusion process. As inherent characteristics and structure of signals are reflected better via sparse representation, the proposed method can well preserve spectral and spatial details of the source images. Experimental results on IKONOS and Quickbird images demonstrate our method's superiority in both the spatial resolution improvement and the spectral information preservation.
Wenxian Yu
IGARSS3
2013 An efficient geography registration method for InSAR coherent change detection
abstract
In this paper, an efficient geography registration method for INSAR coherent change detection is proposed. In the algorithm, considering Terra-SAR image, firstly we do geography registration on two images by oversampling the geography information in master image, then do coarse and fine registration. In coarse registration, find the range and azimuth shifts by calculating the cross correlation of the intensity sub-images that around ground control points (GCPs). In fine registration, calculate intensity cross correlation of sub-images that around GCPs to find the maximum of correlation map for getting range and azimuth shifts. In coarse and fine registration algorithms, time complexity is determined by the size of sub-image and number of GCPs. The proposed algorithm improves the efficiency while keeping the precision of registration. Experimental results obtained on acquired by Terra-SAR images confirm the effectiveness of the proposed approach.
Wanjun Zhang, Wenxian Yu
IGARSS6
2013 Superpixel-Based Classification With an Adaptive Number of Classes for Polarimetric SAR Images
abstract
Polarimetric synthetic aperture radar (PolSAR) image classification, an important technique in the remote sensing area, has been deeply studied for a couple of decades. In order to develop a robust automatic or semiautomatic classification system for PolSAR images, two important problems should be addressed: 1) incorporation of spatial relations between pixels; 2) estimation of the number of classes in the image. Therefore, in this paper, we present a novel superpixel-based classification framework with an adaptive number of classes for PolSAR images. The approach is mainly composed of three operations. First, the PolSAR image is partitioned into superpixels, which are local, coherent regions and preserve most of the characteristics necessary for image information extraction. Then, the number of classes and each class center within the data are estimated using the pairwise dissimilarity information between superpixels, followed by the final classification operation. The proposed framework takes the spatial relations between pixels into consideration and makes good use of the inherent statistical characteristics and contour information of PolSAR data. The framework is capable of improving the classification accuracy, making the results more understandable and easier for further analyses, and providing robust performance under various numbers of classes. The performance of the proposed classification framework on one synthetic and three real data sets is presented and analyzed; and the experimental results show that the framework provides a promising solution for unsupervised classification of PolSAR images.
Bin Liu 0019, Kaizhi Wang, Xingzhao Liu, Wenxian Yu
IEEE Trans. Geosci. Remote. Sens.6
2012 Land-cover classification of SAR images by combining low-level features and category context
abstract
A novel land-cover classification framework for HR SAR images which combines low-level features and category context is presented in this paper. We use patch-based features for low-level information extraction, including average intensity, texture within a patch and the super texture we proposed to model the texture similarity of neighboring patches. To represent the local category context of SAR images, we propose the label layout filter. This work resolves local ambiguities of low-level features from a category context perspective. The framework demonstrates good performance in both accuracy and visual appearance for HR SAR scene interpretation.
Yongke Ding, Lizhong Qiu, Qiuze Yu, Wenxian Yu, Xingzhao Liu
IGARSS4
2012 Context-aware information modeling for HR SAR image scene interpretation
abstract
In this paper, we improve the traditional bag-of-words-based image representation method in two aspects: preserving the semantics in vocabulary generation and incorporating spatial relations in image representation. Based on that, we present a novel context-aware information modeling method for high resolution synthetic aperture radar image scene interpretation. We compare the proposed method with traditional ones in scene interpretation on TerraSAR-X data sets.
Bin Liu 0019, Qiuze Yu, Xingzhao Liu, Wenxian Yu
IGARSS5
2012 A novel SAR imaging strategy based on compressed sensing
abstract
In this paper, we present a novel SAR (synthetic aperture radar) imaging algorithm based on compressed sensing (CS). We first obtain complex SAR images from raw SAR echoes, and then search all targets in complex-image domain by an improved MP (matching pursuit) algorithm. Unlike other CS-based imaging models, our algorithm has several advantages: 1) our algorithm is a two-dimensional (2-D) imaging system, rather than one-dimensional (1-D) imaging models introduced in most existing CS-based imaging algorithms; 2) unlike low efficiencies of most reconstruction algorithms, our method has an efficient reconstruction rate; 3) our imaging model applies existing SAR imaging systems, and inherits current SAR imaging results. The experimental results using real SAR echo data demonstrate the effectiveness of our algorithm.
Wentao Lv, Junfeng Wang 0001, Wenxian Yu
IGARSS3
2012 Bayesian change detection based on space contextual information and EM algorithm
abstract
In this paper, an unsupervised change detection method based on space contextual information and EM algorithm is proposed. In the algorithm, each pixel of the difference image is represented by a characteristic quantity constructed from the difference image values considering the space contextual information. EM algorithm is used to achieve the parameter estimation of each class pixels. Bayesian inference is then employed to perform the final change detection results. Experimental results obtained on multi-temporal optical images acquired by Landsat 5 TM confirm the effectiveness of the proposed approach.
Lizhong Qiu, Yongke Ding, Qiuze Yu, Wenxian Yu, Xingzhao Liu
IGARSS4
2012 An improved Normalized Cross Correlation algorithm for SAR image registration
abstract
This paper proposes a robust and fast matching method based on Normalized Cross Correlation (NCC) for Synthetic Aperture Radar (SAR) image matching. NCC is a robust algorithm in SAR image matching. Two main drawbacks of the NCC algorithm are the flatness of the similarity measure maxima, due to the self-similarity of the images, and the high computational complexity [1]. To tackle these two problems, we adopt the block partitioning strategy, texture feature analysis, and the Fast Fourier Transformation (FFT) algorithm and Integral Images to improve the performance of the conventional NCC algorithm. In the block partitioning strategy, we divide the template and the corresponding sub-window in the examined image into some sub-blocks, and there are several sub-blocks in the template, then we use texture features to increase the weight of sub-blocks which contain more terrain information in the template during the matching process, in this way we improve the flatness of the similarity measure maxima greatly. After that we use the FFT algorithm and Integral Images to speed up the proposed method, with the actual situation of our experiment we adopt the FFT and Integral Images based on the block partitioning strategy, thus we significantly reduce the number of computations required to carry out template matching based on the conventional NCC. Experimental results show that the proposed algorithm is more robust and faster than the conventional NCC algorithm.
Qiuze Yu, Wenxian Yu
IGARSS3
2012 Framework design and implementation for oil tank detection in optical satellite imagery
abstract
In this paper, we propose a coarse-to-fine framework design and implementation for oil tank detection in optical satellite imagery. The framework is mainly composed of two operations: 1) from the whole scene imagery, extraction of patches with oil tanks based on the probabilistic latent semantic analysis model; 2) in the relatively small size patches, detection of the oil tanks with Hough transform and template matching. Experiments show that the framework provides a promising solution for oil tank detection in optical satellite imagery.
Chenxian Zhu, Bin Liu 0019, Qiuze Yu, Xingzhao Liu, Wenxian Yu
IGARSS6
2012 Generation of a probabilistic fuzzy rule base by learning from examples
Weidong Hu, Wenxian Yu
Inf. Sci.4
2011 Combinational matching method of amplitude-scale and time-shift for radar HRRP recognition
abstract
Radar high-resolution range profiles (HRRPs) are very sensitive to amplitude-scale and time-shift, and their good recognition performance depends on precisely the matching methods of both sensitivities, thereby, the method to handle these sensitivities is a challenge in the field of Radar automatic targets recognition(RATR). However, most approaches available did not pay much attention to this problem. This paper proposes two algorithms to jointly matching the amplitude and time-shift for training and test phase, respectively, taking Automatic Gaussian Classifier (AGC) as an example. Experiments based on measured data show our algorithms have a remarkable larger average recognition rate than that of the regular method for AGC.
Wenxian Yu, Xingzhao Liu, Kaizhi Wang
IGARSS2
2011 Graph-based ship extraction scheme for optical satellite image
abstract
Automatic detection and recognition of ship in satellite images is very important and has a wide array of applications. This paper concentrates on optical satellite sensor, which provides an important approach for ship monitoring. Graph-based fore/background segmentation scheme is used to extract ship candidant from optical satellite image chip after the detection step, from course to fine. Shadows on the ship are extracted in a CFAR scheme. Because all the parameters in the graph-based algorithms and CFAR are adaptively determined by the algorithms, no parameter tuning problem exists in our method. Experiments based on measured optical satellite images shows our method achieved good balance between computation speed and ship extraction accuracy.
Wenxian Yu, Xingzhao Liu, Kaizhi Wang, Lin Gong, Wentao Lv
IGARSS2
2011 Towards a framework of algorithm management for remote sensing image analysis
abstract
This work is devoted to the proposal of the algorithm management framework for remote sensing image analysis. A hierarchical framework of algorithm management is first introduced. Then we establish an algorithm description model and corresponding reasoning mechanism based on the quotient space problem solving theory and graph theory. With this model the developers can avoid wasteful duplicate algorithm development and complicated parameter adjustment during the algorithm development process for remote sensing image. The effect and efficiency of the framework has been proved through our preliminary experiment results.
Yongke Ding, Kaizhi Wang, Wenxian Yu, Xingzhao Liu
IGARSS3
2011 Atomic decomposition-based SAR imaging technique
abstract
The conventional SAR imaging algorithms are developed based on specified imaging geometry models, and these algorithms become more complex considering finer resolution or more complicated geometry. The quality of the product is measured by resolution, PSLR and ISLR, etc. without regard to its application. This paper proposes a signal-model-based imaging scheme. Independent of imaging geometry, the new scheme focuses the backscattered echoes by estimating the signal parameters. Furthermore, two realizations of the scheme are presented. This scheme will also have great potential for target feature extraction and image interpretation.
Yesheng Gao, Kaizhi Wang, Xingzhao Liu, Wenxian Yu
IGARSS4
2011 SAR imaging based on Compressed Sensing
abstract
In this paper, we propose a CS-based SAR imaging algorithm after analyzing the model of echoes from point target. Our reconstruction algorithm is two-dimensional, unlike current one-dimensional CS-based SAR imaging algorithms. It can reconstruct targets with high resolution from relatively small number of echoes and effectively improves the efficiency of reconstruction. The processing results of simulated data demonstrate the effectiveness of our algorithm.
Yifeng Huan, Junfeng Wang 0001, Xingzhao Liu, Wenxian Yu
IGARSS5
2011 A number-of-classes-adaptive unsupervised classification framework for SAR images
abstract
In this paper, we present a number-of-classes-adaptive unsupervised classification framework for synthetic aperture radar (SAR) images. The framework aims at the provision of robust classification for SAR images even if the number of classes existing in the scene is unknown. It mainly consists of estimation of the number of classes, extraction of each class center, classification of image patches, and integration of spatial relations between patches. The experiment on a TerraSAR-X SAR image shows that the proposed framework presents a promising performance for SAR image classification.
Bin Liu 0019, Kaizhi Wang, Xingzhao Liu, Wenxian Yu
IGARSS5
2011 Scene interpretation for SAR images using supervised topic models
abstract
In this paper, we present a scene interpretation framework for Synthetic Aperture Radar (SAR) images, using keywords of the image contents provided by users. The framework consists of incorporation of prior knowledge with SAR iMage Annotation Tool (SARMAT), representation of SAR images, and prediction of scene labels based on the supervised Latent Dirichlet Allocation (sLDA) model. The experiment on a TerraSAR-X SAR image shows that the proposed framework provides a promising performance for SAR image scene interpretation.
Bin Liu 0019, Kaizhi Wang, Xingzhao Liu, Wenxian Yu
IGARSS5
2011 Concurrent SAR images denoising and segmentation based on a novel model of wavelet coefficients
abstract
A novel segmentation algorithm for Synthetic Aperture Radar (SAR) images is presented in this paper to improve performance. First, we design a model of wavelet coefficients based on the relativities of the coefficients at different scales to sup press noise. Furthermore, we employ a weight-variant graph cuts-based approach to extract objects from complex back ground. Finally, we compare our proposed algorithms with several segmentation measures on synthetic and real SAR images and the experimental results demonstrate that the pro posed strategies have better performances in speckle suppression and image segmentation compared with other methods.
Wentao Lv, Wenxian Yu, Qiuze Yu, Kaizhi Wang
IGARSS3
2011 Supper resolution radar imaging: A virtual array concept approach
abstract
A virtual array concept is proposed to get supper-resolution in synthetic aperture radar imagery. Two kinds of virtual arrays called virtual frequency array and virtual azimuth array are constructed, combining with phased array signal processing methods, to bring out better imaging performance than traditional range-Doppler and reciprocal spectrum algorithms used in previous literatures.
Gaohuan Lv, Kaizhi Wang, Xingzhao Liu, Wenxian Yu
IGARSS4
2011 A Preliminary study on imaging time difference among bands of WorldView-2 and its potential applications
abstract
WorldView-2 has 8 multispectral bands and 1 pan band. In its LIB level data files, these bands have the same camera model and orientation model. But based on the study of moving objects' characteristics in images from different bands, we found that, like the situation for QuickBird 2 sensor, different band has different imaging time. This paper introduced the calculation method for the imaging time difference of different bands of WorldView-2, reasoned out a possible CCD detector' layout of the focal plane, and pointed out some drawbacks of the current LIB level data of WorldView-2. Secondly, this paper introduced some potential usage for the imaging time difference of different bands of WorldView-2, including ground traffic animation, ground traffic data collection.
Jianwei Tao, Wenxian Yu
IGARSS2
2011 A Foreground/Background Separation Framework for Interpreting Polarimetric SAR Images
abstract
In this letter, we present a novel foreground/background separation (FBS) framework for interpreting polarimetric synthetic aperture radar (PolSAR) images. The FBS framework takes the spatial relations between pixels into consideration and incorporates the advantages of pairwise dissimilarity-based grouping schemes. The FBS method can separate specific targets and objects from the background, which is essential in an interpretation system. Multiple FBS operations can be integrated to interpret PolSAR images, flexibly fusing various inherent features of PolSAR data. Several PolSAR data sets are used to verify the proposed approach.
Bin Liu 0019, Kaizhi Wang, Xingzhao Liu, Wenxian Yu
IEEE Geosci. Remote. Sens. Lett.5
2010 Reciprocal spectrum algorithm for radar imaging with frequency sampling waveform
abstract
A radar imaging algorithm named reciprocal spectrum algorithm (RSA) is proposed in this paper to get higher resolution in azimuth direction with frequency sampling waveform. Theoretical analysis and simulation results show that the algorithm can give better performance than the tradition range Doppler algorithm (RDA) at the cost of peak value reduction at the object points in radar image.
Gaohuan Lv, Kaizhi Wang, Xingzhao Liu, Wenxian Yu, Guozhong Chen, Junli Chen
IGARSS4
2010 Progressive SAR imaging technique
abstract
A progressive SAR imaging technique is proposed as a novel SAR raw data processing scheme in this paper. Different from the classic SAR imaging algorithms which focus the SAR raw data as a whole regardless of the backscattering signal, the novel scheme discriminate the backscattering signal and choose the proper ones to be focused progressively according to some application context while the others will be discarded. The new scheme makes the SAR imaging algorithm not only focus the energy back to the scattering points but also form an image more suitable for some special applications. In this paper, a realization of the scheme present via Atomic Decomposition (AD)[1]–[3]. AD helps estimate the parameters of backscattered signal of a scattering point and detach it from the raw data. With the parameters, the backscattering-signal can be reconstructed and focused. Targets or scatter-points in the scene will be imaged progressively in a sequence of their energy due to the greedy natural of matching pursuit employed during AD. Therefore, the contents in the final SAR image can be controlled by an energy threshold. Besides the interested targets can be extracted easily from the background clutter and this may be helpful for the SAR image understanding
Kaizhi Wang, Xingzhao Liu, Wenxian Yu, Junli Chen, Guozhong Chen
IGARSS3
2010 Multi-channel radar array design method and algorithm
Yi Su 0003, Yutao Zhu 0005, Wenxian Yu, Wentai Lei
Sci. China Inf. Sci.3
2010 An ISAR Imaging Method Based on MIMO Technique
abstract
With the inverse synthetic aperture radar (ISAR) imaging model, targets should move smoothly during the coherent processing interval (CPI). Since the CPI is quite long, fluctuations of a target's velocity and gesture will deteriorate image quality. This paper presents a multiple-input-multiple-output (MIMO)-ISAR imaging method by combining MIMO techniques and ISAR imaging theory. By using a specialM-transmitterN-receiver linear array, a group ofMorthogonal phase-code modulation signals with identical bandwidth and center frequency is transmitted. With a matched filter set, every target response corresponding to the orthogonal signals can be isolated at each receiving channel, and range compression is completed simultaneously. Based on phase center approximation theory, the minimum entropy criterion is used to rearrange the echo data after the target's velocity has been estimated, and then, the azimuth imaging will finally finish. The analysis of imaging and simulation results show that the minimum CPI of the MIMO-ISAR imaging method is 1/MNof the conventional ISAR imaging method under the same azimuth-resolution condition. It means that most flying targets can satisfy the condition that targets should move smoothly during CPI; therefore, the applicability and the quality of ISAR imaging will be improved.
Yutao Zhu 0005, Yi Su 0003, Wenxian Yu
IEEE Trans. Geosci. Remote. Sens.3
2010 The BCGS-FFT Method Combined With an Improved Discrete Complex Image Method for EM Scattering From Electrically Large Objects in Multilayered Media
abstract
This paper presents an efficient algorithm combining the stabilized biconjugate gradient fast Fourier transform (BCGS-FFT) method with an improved discrete complex image method (DCIM) for electromagnetic scattering from electrically large objects in both lossless and lossy multilayered media. The required spatial Green's functions obtained by the improved DCIM are accurate both in the near- and far-field regions without any quasi-static and surface-wave extraction. Then, the scattering by buried objects is considered using the BCGS-FFT method combined with the improved DCIM. Numerical results show the improved DCIM can save tremendous CPU time in scattering involving buried objects.
Xingbin Ye, Weidong Hu, Wenxian Yu, Guoqiang Zhu
IEEE Trans. Geosci. Remote. Sens.5
2009 Antenna Pointing Measurement for Spaceborne SAR based on Sign-MLCC Algorithm
abstract
Dual-antenna single-pass synthetic aperture radar interferometry needs alignment of both antenna beams to achieve the best interferometric performance. Meanwhile, geosynchronous synthetic aperture radar, potentially used for global earthquake prediction and many other attractive applications, also requires antenna pointing control system to steer the radar antenna to illuminate desired territory, otherwise, even slight deviation from ideally boresight direction can cause a great variation of footprint position, because of the large slant range from the radar to mapped area. In principle, it is possible to measure antenna pointing information directly; however, measurement uncertainties will limit the accuracy. Thus it is feasible to resort to received radar data to measure the antenna pointing. This paper concentrates on the antenna pointing measurement using onboard Doppler centroid estimator, and furthermore, to drive pitch and yaw angles to steer the antenna pointing. In order to realize real-time onboard processing, a novel Doppler centroid estimation algorithm, called sign-MLCC, is presented here, utilizing the phase information of the received signal and the arcsine law by analyzing the sign alone, then evaluation of the algorithm is discussed. Finally, simulations are shown to prove the validity and reliability of the proposed method.
Yesheng Gao, Kaizhi Wang, Xingzhao Liu, Wenxian Yu
IGARSS (4)4
2009 A GPU based Time-domain Raw Signal Simulator for Interferometric SAR
abstract
A novel GPU based time-domain raw signal simulator for InSAR is proposed in this paper to exploit the parallel computation of GPU using CUDA language. This simulator combines the advantages of both time-domain and frequency-domain InSAR simulator, i.e., it considers the baseline oscillation and real orbit, and it is also very efficient. Experimental results show the effectiveness of the simulator in varieties of conditions.
Kaizhi Wang, Xingzhao Liu, Wenxian Yu
IGARSS (5)4
2009 Weights Updated Voting for Ensemble of Neural Networks Based Incremental Learning
Shengping Xia, Weidong Hu, Wenxian Yu
ISNN (1)4
2009 Local Degrees of Freedom of Airborne Array Radar Clutter for STAP
abstract
In this letter, the local degree-of-freedom (LDOF) theorem for reduced-dimension space-time adaptive processing (STAP) methods is presented, and a rigorous proof is provided. LDOF is more valuable for practical STAP methods than conventional full degrees of freedom. The effectiveness of the LDOF theorem is verified, and the influences of some operations in practice on the LDOF are analyzed by simulations.
Zenghui Zhang, Wenchong Xie, Weidong Hu, Wenxian Yu
IEEE Geosci. Remote. Sens. Lett.4
2008 A Study on Geophysical Model Function Modeling with Water Surface Temperature as One of the Input Parameters
abstract
Geophysical model function is the basis for the wind vector retrieval with scatterometer and a number of models have been developed to operationally retrieve the ocean surface wind in the past three decades. However, none of the operational models ever took the water surface temperature into account in its modeling, which is considered to have some effect on the ocean backscattering, and in turn on the model accuracy. Taking Sea Winds as an example, this paper attempts to develop new geophysical model functions with surface temperature to be taken into account by using its level 2A data and corresponding buoy data. For contrast, two independent models are established for the ocean water and fresh water respectively. The modeling results and analysis indicate that some effect of the surface temperature on backscatter were found for both types of water, but with a larger extent of the temperature effect for fresh water.
Xuetong Xie, Kehai Chen, Wenxian Yu, Weidong Hu, Qiming Zeng, Yu Fang 0001
IGARSS (1)3
2008 Validation of QSCAT-1 Geophysical Model Function Using Seawinds Level 2 and Buoy Data
abstract
Geophysical model function(GMF) is the basis and prerequisite for the ocean surface wind vector retrieval with scatterometers. Among many operational models, the Qscat-1 model was specifically developed for SeaWinds scatterometer and is being applied to its operational wind retrieval. This paper is to validate the accuracy of the Qscat-1 model by using some SeaWinds Level 2 data and corresponding buoy data. First, a comparison between L2B and co-located buoy wind speed was made to analyze the systematic bias between them, and then a new geophysical model function was established using the match-ups of the L2A and Buoy data to further valuate the accuracy of the Qscat-1 model. The analytical and modeling results indicate that there may be some systematic error in the Qscat-1 model.
Xuetong Xie, Qiming Zeng, Weidong Hu, Wenxian Yu, Kehai Chen, Yu Fang 0001
IGARSS (1)4
2008 Region-Based Classification of Polarimetric SAR Images Using Wishart MRF
abstract
The scattering measurements of individual pixels in polarimetric SAR images are affected by speckle; hence, the performance of classification approaches, taking individual pixels as elements, would be damaged. By introducing the spatial relation between adjacent pixels, a novel classification method, taking regions as elements, is proposed using a Markov random field (MRF). In this method, an image is oversegmented into a large amount of rectangular regions first. Then, to use fully the statisticalaprioriknowledge of the data and the spatial relation of neighboring pixels, a Wishart MRF model, combining the Wishart distribution with the MRF, is proposed, and an iterative conditional mode algorithm is adopted to adjust oversegmentation results so that the shapes of all regions match the ground truth better. Finally, a Wishart-based maximum likelihood, based on regions, is used to obtain a classification map. Real polarimetric images are used in experiments. Compared with the other three frequently used methods, higher accuracy is observed, and classification maps are in better agreement with the initial ground maps, using the proposed method.
Kefeng Ji, Wenxian Yu
IEEE Geosci. Remote. Sens. Lett.3
2005 Simulation of SAR image of ship
Kefeng Ji, Gangyao Kuang, Wenxian Yu
IGARSS4