EDBT 2026 Demo / reviewers in the wild / expert
Gangyao Kuang
dblp:03/2968 · also Gang-Yao Kuang
· DBLP profile ↗
133ranked-venue papers
0as first author
80since 2021 · last 2027
0000-0003-2620-889XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 96 · 59 since 2021Artificial intelligence and machine learning · 21 · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 10 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Multi-level heterogeneous knowledge transfer network on forward scattering center model for limited samples SAR ATR
Daochang Wang, Siqian Zhang, Wancong Li, Gangyao Kuang |
Expert Syst. Appl. | 6 |
| 2026 | Iterative Global Mapping-Local Searching for Heterogeneous Change Detection with Unregistered Images
Yuli Sun, Junzheng Wu, Han Zhang 0005, Lin Lei, Gangyao Kuang |
Int. J. Comput. Vis. | 6 |
| 2026 | Leveraging image transformation and optical flow for heterogeneous change detection under co-registration errors
Yuli Sun, Lin Lei, Gangyao Kuang |
Pattern Recognit. | 3 |
| 2026 | Change-Prior-Guided Unsupervised Change Detection of Heterogeneous Remote Sensing ImagesabstractHeterogeneous change detection (HeCD) enables the identification of land-cover changes using remote sensing imagery obtained from different sensors. Most existing methods overly emphasize modality transformation and shared feature extraction to bridge the gap between heterogeneous images. While these strategies facilitate comparable representations, they tend to neglect the intrinsic characteristics of the changes themselves, which limits their effectiveness in complex scenarios. To overcome this limitation, we propose a change prior-guided image transformation model (CPIT) for unsupervised HeCD. Specifically, starting from the definition of change detection, we analyze the connections among pairwise object relationships, change labels, and change semantics, and then derive change semantic consistency and inconsistency rules solely from the inherent nature of the change detection problem, without relying on data-specific assumptions. These rules are subsequently encoded as change semantic consistency and inconsistency constraints, which, from the perspective of graph signal processing, correspond to low-pass and high-pass spectral properties of the change signals. Finally, by integrating these semantic constraints with sparsity priors and image transformation constraints, we formulate a more precise transformation model for HeCD. Solving this model produces change detection results that conform to the change priors, thereby improving the detection performance. The derivation, formulation, and utilization of change priors in this work offer valuable insights for broader change detection research. Extensive experiments on five datasets validate the effectiveness of CPIT. The code will be released at https://github.com/yulisun/CPIT. Yuli Sun, Lin Lei, Gangyao Kuang |
IEEE Trans. Image Process. | 3 |
| 2025 | EOOD: End-to-end oriented object detection
Caiguang Zhang, Zilong Chen, Boli Xiong, Kefeng Ji, Gangyao Kuang |
Neurocomputing | 5 |
| 2025 | Temporal-Spatial Feature Interaction Network for Multi-Drone Multi-Object TrackingabstractMulti-drone multi-object tracking (MDMOT) aims to localize and identify targets from videos captured simultaneously by multiple drones. To accomplish this task, existing methods typically follow the strategy of associating localized targets to obtain identities. However, their localization and identification stages heavily rely on single-frame information, resulting in the localization being very sensitive to visual information decay and making it struggle to capture discriminative representations for target identification. Consequently, they usually exhibit unreliable performance in challenging scenarios, such as occlusion and high similarity among targets. To this end, we introduce a novel MDMOT framework to interact temporal-spatial features, exploring the guidance of tracklet information across time and space. Specifically, we introduce temporal-spatial feedback loops to enrich cues in our tracker. Meanwhile, a novel temporal-oriented target localization is proposed to enhance the response to difficult samples in feature space by utilizing prior knowledge from existing tracklets beyond the current frame for target localization. Moreover, a spatial-oriented target identification is designed to synergize cross-drone information of tracklets, thereby providing discriminative representations for target identification. It combines target and background information to extract identity representations and interacts features from multiple drones. To our best knowledge, this work reports the first MDMOT system that synergizes features across multiple drones to track targets. By incorporating these two elaborated networks, we develop a robust tracker (named TSMMT). Extensive experiments on the MDMT public dataset demonstrate the superiority of our proposed model. Specifically, TSMMT outperforms state-of-the-art methods by 2.76%~4.66% on MOTA and 2.06%~3.33% on IDF1. Hao Sun 0042, Kefeng Ji, Gangyao Kuang |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | DiffDual-AD: Diffusion-Based Dual-Stage Adversarial Defense Framework in Remote Sensing With Denoiser ConstraintabstractDeep neural networks (DNNs), though highly effective in various Earth observation tasks with remote sensing images (RSIs), are vulnerable to adversarial attacks, threatening their reliability. Each maliciously attacked RSI potentially contains unique and critical information, but current defenses lack a unified framework for rapidly detecting adversarial RSIs and accurately restoring them to their natural state. To bridge this research gap, we propose a diffusion-based dual-stage adversarial defense (DiffDual-AD) framework. In the first stage, we propose a novel adversarial detection method based on the score expectation (AD-SE), which is integrated into the forward process of diffusion model with an improved denoiser constrained for adversarial defense. During the reverse process, the second stage introduces the distance and label-guided adversarial purification (DL-GAP) to restore adversarial RSIs to natural ones. With the help of two proposed guidelines and the specific denoiser, the DL-GAP effectively smooths out the adversarial perturbations from the detected adversarial RSIs while preserving semantic information and local key features. Finally, DL-GAP yields nonadversarial purified RSIs based on the results of AD-SE. Both stages are integrated within the diffusion models, complementing each other to form a pipelined operational mode. Extensive experiments across three RSI scene classification datasets have proven its efficacy in resisting adversarial RSIs in both whitebox and black-box scenarios, achieving an average adversarial detection accuracy of 87.02% with AD-SE and an average final classification accuracy of 81.50% with the entire DiffDual-AD. The performances of our proposed AD-SE and DL-GAP have surpassed other advanced adversarial detection and purification methods when applied to RSIs. Therefore, DiffDual-AD provides a unified and universal solution to preserve the utility and security of processing the adversarial RSIs, which advances the adversarial defense research in remote sensing. Zihao Lu, Hao Sun 0042, Lin Lei, Yuli Sun, Gangyao Kuang |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | Signed Graph-Based Image Transformation for Heterogeneous Change DetectionabstractHeterogeneous change detection (HeCD) is a highly valuable yet challenging task in remote sensing. To enable the comparison of heterogeneous images with different imaging mechanisms, some structural consistency-based image transformation methods have been proposed, which utilize graph models to represent image structures and constrain the transformed images and original images to have the same structural characteristics on the graph model. Consequently, these graph-based methods face two challenges: adequately characterizing the image structure and effectively utilizing the change information. To address these challenges, this article proposes a signed graph-based image transformation (SGIT) method for unsupervised HeCD. First, we analyze the limitations of previous unsigned graph-based methods in capturing the image structure, which leads to the failure to detect changes in some scenes. In light of this, we construct signed graph models that utilize positive/negative weights to represent the similarity/dissimilarity relationships within the image, respectively, and employ adaptive weighting, negative sampling, and neighborhood expansion strategies to bolster the structure representation capability of signed graphs. Second, we analyze how the change would induce a bimodal distribution of vertex feature distances in original and transformed images. Subsequently, a distribution-induced reweighted graph Laplacian regularization (RGLR) is proposed to exploit this prior change information. Finally, a more accuracy image transformation model is obtained by incorporating three types of constraints: signed graph-based structural consistency term, bimodal distribution-induced RGLR, and change sparsity-based penalty term. Extensive comparative experiments on five real datasets have demonstrated the effectiveness of the proposed SGIT. Yuli Sun, Ming Li 0066, Lin Lei, Gangyao Kuang |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | Arbitrary-Direction SAR Ship Detection Method for Multiscale ImbalanceabstractArbitrary-oriented ship detection in SAR imagery remains especially challenging due to multi-scale imbalance and the characteristics of SAR imaging, a problem that is more pronounced than in optical ship detection. Unlike optical images, SAR data often lack rich textural and color cues, instead exhibiting non-uniform scattering, speckle noise, and non-standard elliptical ship shapes, all of which make robust feature extraction and bounding box regression significantly more difficult across different scales. To address these unique SAR-specific challenges, this paper proposes the Multi-Scale Dynamic Feature Fusion Network (MSDFF-Net) aims to alleviate multi-scale imbalance in three main ways. First, a Multi-Scale Large-Kernel Convolution Block (MSLK-Block) integrates large-kernel convolutions with partitioned heterogeneous operations to enhance multi-scale feature representation, tackling wide-ranging ship sizes under noisy conditions. Second, a Dynamic Feature Fusion Block (DFF-Block) handles scale-based feature utilization imbalance by adaptively balancing spatial and channel information, thereby reducing interference from clutter and strengthening discrimination for diverse-scale ships. Third, we propose the Gaussian Probability Distribution (GPD) loss function, which models ships’ elliptical scattering properties and mitigates regression loss imbalance for targets of varying scales and orientations. Experimental evaluations on the R-SSDD, R-HRSID, and CEMEE datasets demonstrate that MSDFF-Net reaches top-tier performance standards, outperforming 21 existing deep learning-based SAR ship detectors. Specifically, MSDFF-Net achieves 93.95% precision, 94.72% recall, 91.55% mAP, 94.33% F1-Score, and 135.79 FPS on the R-SSDD dataset, with a parameter size of only 8.94 M. Additionally, MSDFF-Net exhibits strong transferability across large-scale SAR images, making it suitable for real-world deployment. The code and datasets can be accessed publicly at https://github.com/SZZ-SXM/MSDFF-Net. Zhongzhen Sun, Xiangguang Leng, Boli Xiong, Kefeng Ji, Gangyao Kuang |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2025 | Corrections to "Arbitrary-Direction SAR Ship Detection Method for Multiscale Imbalance"abstractPresents corrections to the paper, (Corrections to “Arbitrary-Direction SAR Ship Detection Method for Multiscale Imbalance”). Zhongzhen Sun, Xiangguang Leng, Boli Xiong, Kefeng Ji, Gangyao Kuang |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2025 | Enhancing Geolocation Accuracy of High-Altitude Airborne SAR Through Tropospheric Delay CompensationabstractGeolocation is a crucial step in the processing of synthetic aperture radar (SAR) images. High-altitude airborne SAR systems present unique geolocation challenges due to travelling long distances through the troposphere. However, the impact of tropospheric delay on geolocation is often overlooked in existing airborne SAR studies, which can lead to inaccuracies. To address this issue, we propose a new positioning method, the tropospheric delay-compensated range-Doppler (TDC-RD) model. The TDC-RD model leverages reference atmospheric models to estimate and compensate for the tropospheric delay in SAR images. This model effectively mitigates the impact of tropospheric delay in SAR geolocation. To further optimize the TDC-RD model solution, a digital elevation model (DEM)-assisted dual iteration method is proposed. This method iteratively adjusts the target’s plane position and elevation in an alternating manner. The effectiveness of the TDC-RD model has been validated through both simulation experiments and actual flight experiments. The results show a significant improvement in geolocation accuracy compared to existing methods, with a maximum reduction of 11.39 m and 24.33% in the mean absolute error (MAE) of SAR geolocation. The TDC-RD model has a great advantage in long-range SAR geolocation. Our research enhances the accuracy and stability of high-altitude airborne SAR geolocation without requiring ground control points. Yaobing Xiang, Yuli Sun, Lin Lei, Kefeng Ji, Gangyao Kuang |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | Cross-Sensor SAR Image Target Detection Based on Dynamic Feature Discrimination and Center-Aware CalibrationabstractIn practical SAR target detection applications, it is often encountered that the training and testing data come from different SAR sensors, leading to a decline in SAR target detection performance. Although domain adaptation methods can achieve model generalization through feature transfer, the change of scattering characteristics for the same target and the difference of feature distribution, caused by cross-sensor, cannot be ignored in SAR images. It is inevitable to lead to the escalation of the offset in the bounding box regression and deviation of the feature alignment. To address these issues, a cross-sensor SAR image target detection method based on dynamic feature discrimination and center-aware calibration is proposed. Based on the domain adaptation framework, initially, a Dynamic Feature Discrimination Module (DFDM) is introduced to address the exacerbated offset in the regression. A bidirectional spatial feature aggregation mechanism is employed to aggregate features in both horizontal and vertical directions and a multi-scale structure is adopted to enhance the scattering and semantic features, which can dynamically constrain the target position while improving target discrimination capability. Then, the Center-Aware Calibration Module (CACM) is designed to address the alignment deviation in feature transfer. The target salience relationship is modeled based on the distance between different positions and the target center to suppress background clutter interference. The perception center of the target is focused by combining the centerness map and classification map, which can calibrate the domain-invariant features and alleviate misalignment. Finally, the proposed method is tested on two datasets, MiniSAR and FARAD, and compared with the latest domain adaption methods. Both mAP and F1 values have improved by more than 6%-20%, verifying the effectiveness of the proposed method. Siqian Zhang, Zhongzhen Sun, Chenfang Liu, Yuli Sun, Kefeng Ji, Gangyao Kuang |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2025 | Locality Preservation for Unsupervised Multimodal Change Detection in Remote Sensing ImageryabstractMultimodal change detection (MCD) is a topic of increasing interest in remote sensing. Due to different imaging mechanisms, the multimodal images cannot be directly compared to detect the changes. In this article, we explore the topological structure of multimodal images and construct the links between class relationships (same/different) and change labels (changed/unchanged) of pairwise superpixels, which are imaging modality-invariant. With these links, we formulate the MCD problem within a mathematical framework termed the locality-preserving energy model (LPEM), which is used to maintain the local consistency constraints embedded in the links: the structure consistency based on feature similarity and the label consistency based on spatial continuity. Because the foundation of LPEM, i.e., the links, is intuitively explainable and universal, the proposed method is very robust across different MCD situations. Noteworthy, LPEM is built directly on the label of each superpixel, so it is a paradigm that outputs the change map (CM) directly without the need to generate intermediate difference image (DI) as most previous algorithms have done. Experiments on different real datasets demonstrate the effectiveness of the proposed method. Source code of the proposed method is made available at https://github.com/yulisun/LPEM. Yuli Sun, Lin Lei, Dongdong Guan, Gangyao Kuang, Li Liu 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Few-Shot Class-Incremental SAR Target Recognition via Decoupled Scattering Augmentation ClassifierabstractDeep learning (DL) techniques have recently ignited remarkable prosperity in the Synthetic Aperture Radar Automatic Target Recognition (SAR ATR) field. Nevertheless, as targets of new categories are observed continually with few-shot examples in openly dynamic scenarios, endowing the DL-based SAR ATR systems with Few-Shot Class-Incremental Learning (FSCIL) ability is urgently demanded. In response, a Decoupled Scattering Augmentation Classifier (DSAC) is proposed to mitigate both intrinsic and domain-specific challenges of the FSCIL of SAR ATR. Specifically, as the significant partability of target structures in SAR imagery, virtual targets with potential scattering patterns are synthesized and pre-allocated by a Scattering Augmentation Module (SAM) to unleash the model’s forward compatibility for future categories. Once deployed, the DSAC is decoupled with dynamic worlds for prompt knowledge representation. Also, a prototypical Nearest-Class-Mean (NCM) classifier with cosine criterion is leveraged for stable and general identification. Extensive experiments conducted on an FSCIL of SAR ATR dataset verify the superiority of our method compared to various latest benchmarks. Yan Zhao 0026, Lingjun Zhao, Siqian Zhang, Kefeng Ji, Gangyao Kuang |
IGARSS | 5 |
| 2024 | SAR Target Open-Set Recognition Based on Joint Training of Class-Specific Sub-Dictionary LearningabstractSynthetic aperture radar (SAR) automatic target recognition (ATR) has attracted extensive attention and achieved satisfactory results. However, most SAR ATR methods follow the closed-set assumption, which assumes that all target classes in the test set have been contained by the training set. In actual scenarios, it may encounter the target classes that are not included in the training set, and it presents a challenge for SAR ATR. To tackle this issue, this letter proposes an open-set recognition method based on joint training of class-specific sub-dictionary learning. First, joint training is used to optimize the sub-dictionary learning process, and it could significantly enhance the discriminative ability of these sub-dictionaries. Second, the reconstruction errors of the targets on each sub-dictionary are calculated. These errors can be split into matched and no-matched errors. Third, extreme value theory (EVT) is employed to model the matched and no-matched errors of each class, which could determine the class boundaries. Finally, for the target to be recognized, its reconstruction errors on each sub-dictionary are calculated individually. The class of this target can be determined by comparing the errors with these class boundaries. Our method achieved an accuracy of 87.22–94.02 and an F1 score of 88.03–90.25 in multiple experiments on moving and stationary target automatic recognition (MSTAR) dataset. Compared with several state-of-the-art methods, it has better accuracy and robustness. Xiaojie Ma, Kefeng Ji, Linbin Zhang, Sijia Feng, Boli Xiong, Gangyao Kuang |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2024 | Ship Recognition for Complex SAR Images via Dual-Branch Transformer Fusion NetworkabstractShip recognition in synthetic aperture radar (SAR) is an essential challenge in SAR image interpretation. The measured SAR ship targets often contain complex background such as port facilities and neighboring ships, which are easy to interfere with the model and affect the recognition performance. To address this issue, a SAR ship recognition method with complex background based on dual-branch transformer fusion network is proposed in this paper. First of all, a dual-branch feature extraction and fusion architecture is designed in this paper, including significant feature extraction (SFE), global feature extraction (GFE), and dual-branch feature fusion (D-BFF). Specifically, the SFE effectively extracts the most discriminative local fine-grained features of ship target using multi-layer convolution of significant regions. The GFE capture global semantic information by residual module optimization. In addition, combined with the self-attention in the transformer block based on cross-attention and position encoding, the effective fusion of SFE and GFE is realized in D-BFF. Finally, extensive experiments are carried out based on Gaofen-3 seven-category dataset (anyone can get the dataset after sending the applying e-mail). The results reveal that the proposed method can achieve a recognition accuracy of 75.55%, which is significantly superior to other algorithms. Zhongzhen Sun, Xiangguang Leng, Boli Xiong, Kefeng Ji, Gangyao Kuang |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2024 | RCShip: A Dataset Dedicated to Ship Detection in Range-Compressed SAR DataabstractTimely monitoring of ships is imperative for ensuring the safety and security of maritime operations. Ship detection in synthetic aperture radar (SAR) is typically applicable to focused images. The time consumption of target detection primarily relies on the imaging process duration, encompassing intricate and time-intensive processing steps such as range migration correction and azimuth compression. Consequently, achieving real-time SAR ship detection poses a significant challenge. To address these issues, ship detection in the range-compressed domain of SAR has emerged as a viable approach. However, there is still a lack of reliable ship detection datasets that can satisfy the detection on the range-compressed domain. In this paper, we construct a dataset specifically designed for ship detection in range-compressed SAR data, called RCShip-1.0 (range-compressed ship dataset). The original data source is publicly available complex-valued data from the Sentinel-1 acquisition and the OpenSARShip-1.0 dataset, encompassing numerous ship targets. Subsequently, the inverse chirp scaling (ICS) algorithm is employed on the complex-valued data to acquire range-compressed SAR data. RCShip-1.0 encompasses training set, validation set, and test set acquired through two distinct approaches. It consists of 1580 large-scale SAR range-compressed images which are further divided into 18322 sub-images to facilitate subsequent display and analysis of detection results within large-scale SAR images. The experimental results demonstrate that each deep network achieves good performance on the dataset, with an F1-score exceeding 65%. The utilization of the RCShip-1.0 dataset in obtaining these experimental outcomes showcases its feasibility, standardization, and public availability. Xiangdong Tan, Xiangguang Leng, Kefeng Ji, Gangyao Kuang |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2024 | Data Distribution Loss for Imbalanced SAR Vehicle Target RecognitionabstractThe data distribution of synthetic aperture radar (SAR) vehicle targets in the actual missions is often imbalanced. However, the recent algorithms for SAR target recognition are designed either under abundant samples, or the situation of few labeled samples among all the categories. These cases all avoid facing the difficulties of imbalanced data distribution,i.e. the difference between the number of labeled samples among categories is huge. The samples in the majority classes will get more chances to be learnt by the deep neural network, which impedes the regular algorithms from achieving a high recognition rate. In this letter, a design guideline for imbalance loss and an example of data distribution (DD) loss based on the guideline is proposed, which provides an extremely effective way of handling the problem of imbalanced SAR target recognition. The DD loss takes the sample distribution and the data quantity of SAR vehicle targets into consideration. It can cause images with fewer samples in their categories to decrease more gradients proportionally. Moreover, the proposed DD loss adds no more burden to the networks and compared to other imbalanced algorithms with complex processes, the DD loss can be conducted easily. Plenty of experiments, which involve two various kinds of imbalanced datasets, are implemented and the proposed DD loss shows excellent performance among these imbalanced datasets. When there are only 40 labeled samples in minority categories, the DD loss can achieve over 95% in nine different cases, which exceeds other methods and losses of at least 7%. Linbin Zhang, Xiangguang Leng, Xiaojie Ma, Kefeng Ji, Gangyao Kuang, Li Liu 0002 |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2024 | Mitigating SAR Out-of-Distribution Overconfidence Based on Evidential UncertaintyabstractSynthetic aperture radar (SAR) automatic target recognition (ATR) is extensively applied in both military and civilian sectors. Nevertheless, test and training data distribution may differ in the open world. Therefore, SAR out-of-distribution (OOD) detection is important because it enhances the reliability and adaptability of SAR systems. However, most OOD detection models are based on maximum likelihood estimation (MLE) and overlook the impact of data uncertainty, leading to overconfidence output for both in-distribution (ID) and OOD data. To address this issue, we consider the effect of data uncertainty on prediction probabilities, treating these probabilities as random variables and modeling them using Dirichlet distribution. Building on this, we propose an evidential uncertainty aware mean squared error (UMSE) loss function to guide the model in learning highly distinguishable output between ID and OOD data. Furthermore, to comprehensively evaluate OOD detection performance, we have compiled and organized some publicly available data and constructed a new SAR OOD detection dataset named SAR-OOD. Experimental results on SAR-OOD demonstrate that the UMSE approach achieves state-of-the-art (SOTA) performance. The code and data are available at:https://github.com/Xiaoyan-Zhou/UMSE-SAR-OOD-Detection. Tao Tang 0006, Zhongzhen Sun, Gangyao Kuang, Janne Heikkilä, Li Liu 0002 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2024 | Deep Learning for Visual Speech Analysis: A SurveyabstractVisual speech, referring to the visual domain of speech, has attracted increasing attention due to its wide applications, such as public security, medical treatment, military defense, and film entertainment. As a powerful AI strategy, deep learning techniques have extensively promoted the development of visual speech learning. Over the past five years, numerous deep learning based methods have been proposed to address various problems in this area, especially automatic visual speech recognition and generation. To push forward future research on visual speech, this paper will present a comprehensive review of recent progress in deep learning methods on visual speech analysis. We cover different aspects of visual speech, including fundamental problems, challenges, benchmark datasets, a taxonomy of existing methods, and state-of-the-art performance. Besides, we also identify gaps in current research and discuss inspiring future research directions. Changchong Sheng, Gangyao Kuang, Liang Bai 0003, Chenping Hou, Yulan Guo, Xin Xu 0001, Matti Pietikäinen, Li Liu 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | TirSA: A Three Stage Approach for UAV-Satellite Cross-View Geo-Localization Based on Self-Supervised Feature EnhancementabstractCross-view geo-localization aims to associate geographical location with different view images shot from different platforms. One of the critical challenges is how to effectively emphasize architectural features and reducing background interference to achieve robust cross-view matching. Most of the existing methods fail to adequately address the features of buildings, treating foreground and background equally. Leveraging prior knowledge to enhance the features of crucial architectural foreground yields greater benefits in Geo-Localization. A comprehensive three stage approach (TirSA) is proposed in this paper, which consists of three components: Pre-processing, Generate Feature Embedding, and Post-processing. In the Pre-processing stage, we employ a self-supervised feature enhancement method (SFEM) to obtain the building aware mask. Without adding additional auxiliary information, the model is guided to learn from discriminative building regions. Besides, in the Generate Feature Embedding stage, we propose an adaptive feature integration module (AFIM) to enhance feature representation capability. We also train the Siamese network using a novel improved cross-domain triplet loss to reduce the impact of inter-view domain gap. Finally, in the Post-processing stage, we employ a re-ranking method to optimize the initial retrieval list, further enhancing the matching accuracy. Remarkably, extensive experiments show that our proposed TirSA exceeds state-of-the-art by a large margin and achieves optimality in both drone-view target localization and drone navigation. Especially in the drone navigation task, our method is superior to the existing methods, achieving an improvement of approximately 5%. Code will be released at https://github.com/SunJ1025/TirSA. Jian Sun 0038, Hao Sun 0042, Lin Lei, Kefeng Ji, Gangyao Kuang |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Geospatial Contextual Prior-Enabled Knowledge Reasoning Framework for Fine-Grained Aircraft Detection in Panoramic SAR ImageryabstractFine-grained aircraft detection from synthetic aperture radar (SAR) imagery is of significance in transportation and military domains. Based on the prior knowledge that aircraft are frequently found in airports, current research adopts airport-to-aircraft detection pipelines for aircraft detection and classification in panoramic SAR images. Geospatial information, including the approximate airport location and the spatial relationships between the airport and aircraft, represents valuable supplementary information that can enhance fine-grained aircraft detection performance. However, due to unreliable geographical information and the limited perceptual field, it is challenging to detect and classify aircraft in panoramic SAR imagery. To address this, a novel geospatial contextual prior-enabled knowledge reasoning framework is proposed. First, an unsupervised and lightweight geospatial-driven airport detection (AD) method is presented by combining geographic information matching and optical-to-SAR image registration, which can quickly locate airports and narrow the scope for fine detection. Then, an improved real-time model for object detection with a recursive-gated spatial interaction module (RTMDet-RSIM) is proposed for fine-grained aircraft detection. RTMDet-RSIM uses frequency-domain analysis to improve global information modeling ability without excessive computational cost. Finally, a relational prior-based reasoning strategy is proposed by modeling the geospatial category relationships from historical images to strengthen classification performance. Experiments on 61 panoramic SAR images covering 13 aircraft categories show that the proposed method achieves high detection accuracy with a mean average precision (mAP) of 81.2%, while the average test time for an image size of$8738\times 7636$is about 8.58 s. The source code and dataset will be released. Ru Luo, Qishan He, Lingjun Zhao, Siqian Zhang, Gangyao Kuang, Kefeng Ji |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Simulated Data Feature Guided Evolution and Distillation for Incremental SAR ATRabstractDeep neural network (DNN)-based synthetic aperture radar automatic target recognition (SAR ATR) methods have made great progress in recent years. However, the performance of DNN models relies on a large number of independent and identically distributed measured synthetic aperture radar (SAR) images, which is contrary to the SAR ATR in practice. Furthermore, DNN models also suffer from catastrophic forgetting when learning a sequence of new classes. To tackle these problems, we introduce simulated data into the class incremental learning of SAR ATR for the first time. Specifically, we aim to continuously learn a sequence of new classes with a small amount of measured data and a large amount of simulated data. We first investigate the properties of incremental learning using simulated data, and the main observation is that simulated data can achieve good performance in short-term incremental learning rather than long-term incremental learning. A novel class incremental learning method, namely, feature guided evolution and distillation (FGED), is then presented. On the one hand, FGED encourages simulated data to have the same feature relationship structure as the corresponding measured data to reduce their distribution discrepancy in short-term incremental learning. On the other hand, FGED adopts a feature distillation strategy to simultaneously reduce the distribution discrepancy accumulation of previous incremental classes and alleviate the catastrophic forgetting in long-term incremental learning. The experimental results obtained on the MSTAR benchmark dataset and two simulated datasets demonstrate the effectiveness of FGED. Hao Sun 0042, Yan Zhao 0026, Qishan He, Siqian Zhang, Gangyao Kuang |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2024 | Azimuth-Aware Subspace Classifier for Few-Shot Class-Incremental SAR ATRabstractWith the rapid acquisition of high-resolution Synthetic Aperture Radar(SAR) images, new categories are continually observed with few-shot instances in openly non-cooperative scenarios. Powering a SAR Automatic Target Recognition (SAR ATR) system with an ability of few-shot class-incremental learning (FSCIL) is nontrivial. Observing the pronounced azimuth-dependence and part-sparsity of targets in SAR images, an Azimuth-aware Subspace Classifier (AASC) on the Grassmannian manifold is proposed to tackle the FSCIL of SAR ATR stably and accurately. In the AASC, losses covering both semantic and manifold facets, which include Semantic Margin Separation (SMS), Deep Subspace Separation (DSS), and Structure Less Forgetting (SLF), are designed to strike both the intrinsic model’s stability and plasticity dilemma and domain-specific challenges. For plasticity, the novel-to-old semantic margins are enlarged by the SMS loss for knowledge transferring while avoiding inappropriate adaptions. The DSS loss derived from the Grassmannian geometry aims to regularize class subspaces orthogonality. For stability, semantic drifts of target spatial and global structures are punished by the SLF loss. As the periodicity and volatility of target azimuth-aware patterns, an Azimuth-aware Exemplar Selection (AES) strategy is designed to select representative and complementary exemplars. In experiments, the advantages of the subspace classifier and the designed losses and strategies are deeply verified. Comprehensive experiments on three FSCIL scenarios derived from both airborne and spaceborne datasets, including the MSTAR, the SAR-AIRcraft-1.0, and self-collected data sets, show that our method significantly outperforms various task-specific benchmarks, verifying its effectiveness for the FSCIL in real SAR ATR scenarios. Yan Zhao 0026, Lingjun Zhao, Siqian Zhang, Kefeng Ji, Gangyao Kuang, Li Liu 0002 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | A Domain-Adaptive Few-Shot SAR Ship Detection Algorithm Driven by the Latent Similarity Between Optical and SAR ImagesabstractDetecting ships in synthetic aperture radar (SAR) images poses a formidable challenge, primarily attributed to limited observation samples and complex environments. To address this problem, driven by latent similarity between optical and SAR images, we propose a domain-adaptive few-shot detection algorithm for SAR ship detection [single shot multibox detector (SSD)]. The algorithm requires only a few training samples of SAR images and effectively combines them with rich optical images to utilize domain information. First, we develop an efficient plug-and-play distance metric function. This function accurately measures the distances between features from the optical domain and the SAR domain. Second, we design a lossy branching mechanism to effectively utilize SAR domain knowledge. This branching mechanism is driven by the observed latent similarity in domain knowledge distribution between optical and SAR images. In addition, we introduce a dual-stream branching feature alignment extraction network with weight sharing. This network architecture enables better knowledge extraction and sharing between optical and SAR domains. To evaluate our method, we conducted experiments on a newly created dataset, DIOR2SSDD, which is designed for few-shot SAR image ship detections across optical and SAR domains. The experimental results show that under three-, five-, and ten-shot settings, the mean average precision (mAP) of our method can reach 59.2%, 61.2%, and 64.6%, and with only 10% SAR training data, the mAP can reach 89.3%. It indicates that our method can effectively transfer domain knowledge and achieve excellent ship detection performance in SAR images. Lingjun Zhao, Kefeng Ji, Gangyao Kuang |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Image Regression With Structure Cycle Consistency for Heterogeneous Change DetectionabstractChange detection (CD) between heterogeneous images is an increasingly interesting topic in remote sensing. The different imaging mechanisms lead to the failure of homogeneous CD methods on heterogeneous images. To address this challenge, we propose a structure cycle consistency-based image regression method, which consists of two components: the exploration of structure representation and the structure-based regression. We first construct a similarity relationship-based graph to capture the structure information of image; here, a k -selection strategy and an adaptive-weighted distance metric are employed to connect each node with its truly similar neighbors. Then, we conduct the structure-based regression with this adaptively learned graph. More specifically, we transform one image to the domain of the other image via the structure cycle consistency, which yields three types of constraints: forward transformation term, cycle transformation term, and sparse regularization term. Noteworthy, it is not a traditional pixel value-based image regression, but an image structure regression, i.e., it requires the transformed image to have the same structure as the original image. Finally, change extraction can be achieved accurately by directly comparing the transformed and original images. Experiments conducted on different real datasets show the excellent performance of the proposed method. The source code of the proposed method will be made available at https://github.com/yulisun/AGSCC. Yuli Sun, Lin Lei, Dongdong Guan, Junzheng Wu, Gangyao Kuang, Li Liu 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2023 | Change Alignment-Based Graph Structure Learning for Unsupervised Heterogeneous Change DetectionabstractHeterogeneous change detection (HCD) in remote sensing has gained significant attention. Heterogeneous images come from different sensors, which cannot be compared directly to detect changes. This letter proposes a change alignment-based graph structure learning method (CAGSL) for unsupervised HCD, which detects changes by calculating forward and backward structure differences. To achieve this objective, CAGSL incorporates two pivotal improvements. Firstly, CAGSL utilizes a graph auto-encoder (GAE) to optimize the graph structure, enabling a more accurate representation of the topological relationships between the real land covers. Secondly, CAGSL introduces a change alignment constraint based on the HCD task property that the forward and backward structural differences represent the same change event in order to enhance the optimization of the graph structure. Subsequently, the optimized graph structure is used to compute the structure difference images through graph mapping. Finally, the change map (CM) is obtained through Otsu segmentation. Experimental results demonstrate the effectiveness of the proposed CAGSL when compared to some state-of-the-art (SOTA) methods. Kuowei Xiao, Yuli Sun, Gangyao Kuang, Lin Lei |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2023 | Scattering Features Spatial-Structural Association Network for Aircraft Recognition in SAR ImagesabstractDue to the all-day, all-weather imaging characteristics of Synthetic Aperture Radar (SAR), aircraft recognition in SAR images is emerging as an vital issue. Electromagnetic scattering (ES) characteristics of target are the unique features of SAR systems. In particular, it is more pronounced for aircraft targets with discrete appearance. However, the existing deep learning methods effectively extract features in image domain, ignoring the potential ES features. To obtain more valuable features for SAR aircraft recognition, an innovative scattering features spatial-structural association network (SFSA) is proposed in this paper. In this work, the strong scattering points (SSPs) of aircraft are extracted and converted into graph structure data, which more directly represent the spatial-structural association among SSPs and clearly reflect the geometric structure of the aircraft target. Subsequently, the graph convolutional neural network (GCN) is employed to extract the structural and high-level semantic ES features. Then, a modified VGGNet is designed to effectively extract the image domain features of aircraft with extremely discrete appearances. In brief, the SFSA network is an end-to-end network that integrates structural ES features and more discriminative image domain features of the aircraft to achieve higher recognition performance. It is also demonstrated that structural ES features facilitate the recognition of aircraft. Experiments on the SAR aircraft dataset validate the effectiveness of SFSA network compared to traditional CNN-based SAR recognition networks and graph neural networks (GNN). Siqian Zhang, Ru Luo, Sijia Feng, Gangyao Kuang |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2023 | Attributed Scattering Center Guided Adversarial Attack for DCNN SAR Target RecognitionabstractRecently, deep learning has made significant progress in synthetic aperture radar automatic target recognition (SAR ATR). However, deep convolutional neural networks (DCNNs) are discovered to be susceptible to carefully crafted adversarial perturbations. Regarding the unique scattering mechanism in SAR imaging, the scattering feature such as attributed scattering centers (ASCs) should be deeply considered in the adversarial attack (AA) algorithms for DCNNs in SAR ATR. In this letter, an AA algorithm named ASC-STA is proposed to take advantage of the powerful SAR imaging property characterization capability of the ASC model. Considering the imaging characteristic that is presented in the ASC model, the spatial transformation module is specially designed to be performed on the strong backscattering structures guided by the reconstruction image of the ASC model. Spatial transformation shifts the robust scattering point features via altering the locally gray-scale relationships of the target microstructures in SAR images. Besides, a modified shape context evaluated metric is established to assess the validity of the feature alternation process. Experimental results on the moving and stationary target acquisition and recognition (MSTAR) database demonstrate that the proposed method achieves a high success rate without complex parameter settings and generates the perturbed image in high image quality. Compared with the norm-based AA algorithms, which are not involved with ASCs, the point feature of the original image is efficiently altered. At the same time, the perturbations are focused on the target regions accurately. Junfan Zhou, Sijia Feng, Hao Sun 0042, Linbin Zhang, Gangyao Kuang |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2023 | PAN: Part Attention Network Integrating Electromagnetic Characteristics for Interpretable SAR Vehicle Target RecognitionabstractMachine learning methods for synthetic aperture radar (SAR) image automatic target recognition (ATR) can be divided into two main types: traditional methods and deep learning methods. The deep learning methods can learn the high-dimensional features of the target directly, and usually obtain high target recognition accuracy. However, they lack full consideration of SAR targets’ inherent characteristics resulting in poor generalization and interpretation ability. Compared with the deep learning methods, traditional methods can get more interpretable and stable results with model-based features. In order to take full advantage of these two kinds of methods, we propose target part attention network based on the attributed scattering center (ASC) model to integrate the electromagnetic characteristics with the deep learning framework. Firstly, considering the importance of scattering structure for SAR ATR, we design a target part model based on ASC model. Then, a novel part attention module based on Scaled Dot-Product Attention mechanism is proposed, which directly associates the features of target parts with the classification results. Finally, we give the derivation method of the importance of each part, which is of great significance for practical application and the interpretation of SAR ATR. Experiments on the MSTAR data set demonstrate the effectiveness of the proposed part attention network. Compared with existing studies, it can achieve higher and more robust classification accuracy under different complex conditions. Furthermore, combined with the importance of parts, we constructed two effective interpretable analysis methods for deep learning network classification results. Sijia Feng, Kefeng Ji, Fulai Wang, Linbin Zhang, Xiaojie Ma, Gangyao Kuang |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2023 | Ship Detection From Raw SAR Echo DataabstractIn the context of ship monitoring in the ocean, targets are usually sparsely distributed. Thus, synthetic aperture radar (SAR) imaging of the whole scene is usually quite redundant and costly. However, raw SAR echo data were considered to be useless before focusing. Few studies have attempted to detect ships from raw SAR echo data. It seems to be an impossible task since the resolution of raw SAR echo data is too low. This article proposes a ship detection method for raw SAR echo data in view of a nonimaging target sensing paradigm. The core idea is that we can sense the existence of ships from raw SAR echo data without imaging. The underlying rationale is that the radar always speaks the same sentence, i.e., usually an exactly identical linear frequency modulated (LFM) signal, while target and clutter answer differently. The difference spread into each part of the whole echo sequence rather than only the focused energy after match filtering. Thus, the ships can be found by pattern analysis on one-dimension sequence data rather than two-dimension images. The experimental results based on simulation and typical real data validate our assumption. This study shows that SAR imaging is an unnecessary intermediate process and opens up new significant possibilities for ship detection in the vast ocean. Xiangguang Leng, Kefeng Ji, Gangyao Kuang |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Open Set Recognition With Incremental Learning for SAR Target ClassificationabstractSAR target classification is an important application in SAR image interpretation. In practical applications, the battlefield is open and dynamic, and the SAR target classification model often encounters the targets of unknown classes. However, most of the existing SAR target classification methods follow the close-set assumption. It makes them only classify several fixed classes of targets and can’t deal with the targets from unknown classes. To this end, this paper proposes a novel SAR target classification method. This method can not only classify the targets from known classes and search targets from unknown classes but also incrementally update the classification model with these unknown class targets. Specifically, an autoencoder improved by MS-SSIM (multi-scale structural similarity) loss is utilized to extract targets’ features, and it can better utilize the structural information in SAR images. Next, the classifier based on EVT (Extreme Value Theorem) is established, which can classify the known class targets and search the unknown class targets. Then, we perform improved model reduction on the established classifier. This operation could speed up the model and prepare for incremental learning. Finally, after manually labeling those unknown class targets, the classifier is updated with these data in incremental form. Experimental results on the MSTAR (Moving and Stationary Target Automatic Recognition) dataset indicate that, compared with the state-of-the-art methods, our proposed method has better performance in open set recognition and incremental learning. Xiaojie Ma, Kefeng Ji, Sijia Feng, Linbin Zhang, Boli Xiong, Gangyao Kuang |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2023 | Exploring Fine Polarimetric Decomposition Technique for Built-Up Area MonitoringabstractHighly variable polarimetric signatures caused by complex structures in built-up areas make interpretation of these scattering behaviors intractable for PolSAR remote sensing. This paper proposes a fine polarimetric decomposition method and derives several products to finely simulate the scattering mechanisms of urban buildings, thus fulfilling its use for effective surveillance. First, through theoretically establishing the roll-invariant condition for a completely general scatterer, a roll-invariant cross polarization (RICP) scattering model is constructed, which characterizes the cross polarization scattering in the manner of planar structure distribution. Second, by designing a root-discriminant-based parameter inversion strategy, a fine seven-component decomposition is proposed, which achieves the complete physical interpretation of matrix elements and reasonable inversion of model parameters. Third, by analyzing the external and internal scattering difference, the derivative products, i.e., scattering contribution synthesizers are derived for built-up area monitoring. Experimental results derived from real PolSAR data confirm the superiority and effectiveness of the constructed descriptors on the one hand. On the other hand, the extensibility of fine polarimetric decomposition in specific remote sensing is also explicitly demonstrated. Sinong Quan, Tao Zhang 0027, Wei Wang 0099, Gangyao Kuang, Xuesong Wang 0003, Bing Zeng 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Structural Regression Fusion for Unsupervised Multimodal Change DetectionabstractMultimodal change detection (MCD) is an increasingly interesting but very challenging topic in remote sensing, which is due to the unavailability of detecting changes by directly comparing multimodal images from different domains. In this paper, we first analyze the structural asymmetry between multitemporal images and show their negative impact on the previous MCD methods using image structures. Specifically, when there is a structural asymmetry, previous structure based methods can only complete a structure comparison or image regression in one direction and fails in the other direction, that is, they cannot transform or convert from complex structural images (with more categories) to simple structural images (with fewer categories). To reduce the influence of structural asymmetry, we propose a structural regression fusion based method (SRF) that simultaneously transforms the pre-event and post-event images into the image domain of each other, calculating the forward and backward changed images, respectively. Noteworthy, different from previous late fusion methods that fuse the forward and backward changed images in the post-processing stage, SRF incorporates fusion into the regression process, which can fully explore the connection between changed images, and thus improve image transformation performance and obtain better changed images. Specifically, SRF yields three types of constraints to perform the fused image transformation: structure consistency based regression term, change smoothness and alignment based fusion term, and prior sparsity based penalty term. Finally, the changes can be extracted by comparing the transformed and original images. The proposed SRF is verified on six real data sets by comparing with some state-of-the-art methods. Source code of the proposed method will be made available at https://github.com/yulisun/SRF. Yuli Sun, Lin Lei, Li Liu 0002, Gangyao Kuang |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | TCD: Task-Collaborated Detector for Oriented Objects in Remote Sensing ImagesabstractOriented object detection (OOD) in remote sensing image interpretation is challenging due to the difficulty of locating objects with arbitrary orientations. Existing methods have made considerable progress based on oriented heads or anchors. However, most of them follow the classical detection paradigm, such as assigning samples based on Intersection-over-Unions (IoU) and predicting through two independent tasks. These fixed strategies impair the consistency between classification and localization predictions, resulting in the prediction with optimal localization accuracy being suppressed by the nonoptimal ones during nonmaximum suppression (NMS). To address this problem, a task-collaborated detector (TCD) is proposed. Compared with current single-stage methods, its improvements include two aspects: task-collaborated assignment (TCA) and task-collaborated head (TCH). Specifically, to better pull closer the best anchors for two tasks, TCA introduces classification and localization confidence into sample assignment and tends to select the anchors with accurate and consistent predictions as positive during training. TCH provides a better balance for learning interactive and discriminative features. It can flexibly adjust the spatial feature distribution of classification and localization tasks by learning the joint features from the aggregation layer. Extensive experiments are conducted on HRSC2016, DOTA, and DIOR-R, and the proposed TCD achieves the state-of-the-art performance [90.60, 80.89, and 65.04 mean average precision (mAP), respectively]. Consistency analysis also demonstrates that TCD can significantly improve prediction consistency. Caiguang Zhang, Boli Xiong, Xiao Li 0017, Gangyao Kuang |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Few-Shot Class-Incremental SAR Target Recognition via Cosine Prototype LearningabstractRecent years have witnessed a remarkable breakthrough in Synthetic Aperture Radar Automatic Target Recognition (SAR ATR) with the development of deep learning (DL). Nonetheless, once deployed, the DL-based methods’ ability to incrementally learn new knowledge from few-shot samples without forgetting the old is fragile, hindering them from discriminating unseen targets in real-world situations. In this paper, we propose a Cosine Prototype Learning (CPL) framework to first unlock few-shot class-incremental learning (FSCIL) in the SAR ATR field inspired by the intrinsic relationships between target azimuth-aware knowledge and semantic features under the cosine criterion. By condensing class-specific characteristics into individual prototypes, stable profiles of targets are depicted without losing generalization. For the model’s plasticity, a pairwise structure separation (PSS) loss is introduced to separate old and new classes and compact intra-class features. Meanwhile, the model’s transferability on new classes is guaranteed by a prototype consistency (PC) loss. For the model’s stability, we propose a prototype-exemplar distillation (PED) loss and a prototype re-calibration (PR) strategy to penalize semantic drifts of old-class feature spaces and alleviate the misalignment of the learned prototypes successively. At inference, a nearest-class-mean (NCM) classifier is adopted for evaluation by comparing cosine similarity scores between testing samples and class-specific prototypes. In experiments, the proposed components of our method are explored by ablation studies. Strong baselines are established, and extensive experiments conducted on the MSTAR dataset show that our method outperforms state-of-the-art methods under various FSCIL conditions, verifying its effectiveness for the FSCIL of SAR ATR. Yan Zhao 0026, Lingjun Zhao, Dewen Hu, Gangyao Kuang, Li Liu 0002 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | Importance-Aware Information Bottleneck Learning Paradigm for Lip ReadingabstractLip reading is the task of decoding text from speakers' mouth movements. Numerous deep learning-based methods have been proposed to address this task. However, these existing deep lip reading models suffer from poor generalization due to overfitting the training data. To resolve this issue, we present a novel learning paradigm that aims to improve the interpretability and generalization of lip reading models. In specific, a Variational Temporal Mask (VTM) module is customized to automatically analyze the importance of frame-level features. Furthermore, the prediction consistency constraints of global information and local temporal important features are introduced to strengthen the model generalization. We evaluate the novel learning paradigm with multiple lip reading baseline models on the LRW and LRW-1000 datasets. Experiments show that the proposed framework significantly improves the generalization performance and interpretability of lip reading models. Changchong Sheng, Li Liu 0002, Wanxia Deng, Liang Bai 0003, Zhong Liu 0002, Songyang Lao, Gangyao Kuang, Matti Pietikäinen |
IEEE Trans. Multim. | 7 |
| 2022 | ASC-Parts Model Guided Multi-Level Fusion Network for SAR Target ClassificationabstractMost deep-learning based synthetic aperture radar (SAR) target classification methods directly apply models designed for natural scenes and do not consider the difference between SAR and optical images. To solve this problem, a parts model guided multi-level fusion network for synthetic aperture radar (SAR) target classification is proposed in this paper, which integrates the electromagnetic scattering features into the deep learning network. Firstly, attribute scattering center (ASC) model based target parts (ASC-Parts) are extracted, which provide scattering features of target from physical model. Then, under the guidance of the ASC-Parts model, the local features of target are further extracted based on the proposed SAR target parts segmentation network. Through this physics guided network, we can introduce the target scattering characteristics into deep learning. Finally, a multi-level fusion structure is developed, which adds the local features to the global features obtained from different layers of traditional deep learning network. The experimental results on Moving and Stationary Target Acquisition and Recognition (MSTAR) data set show the superiority of the novel method and illustrate the effectiveness of physics guided deep learning for SAR target classification. Sijia Feng, Kefeng Ji, Linbin Zhang, Xiaojie Ma, Gangyao Kuang |
IGARSS | 5 |
| 2022 | Ship Detection in Range-Compressed SAR DataabstractMost of synthetic aperture radar (SAR) based ship detection methods utilize two-dimension focused images. Ship detection in range-compressed data is promising since it needs no time-consuming azimuth focusing. This paper proposes a ship detection method in range-compressed SAR data, which employs the statistical characteristics and range trajectory of a ship target in the range-compressed time-domain. First, it employs complex signal kurtosis (CSK) to prescreen potential ship areas since CSK was demonstrated to be an reasonable indicator for SAR ship detection. Then, a convolutional neural networks (CNN) based discrimination is applied to the potential ship areas. The training samples comes from the simulation results of range trajectory based on the radar imaging parameters. Preliminary results show that the proposed method performs well in range-compressed SAR data. Xiangguang Leng, Kefeng Ji, Gangyao Kuang |
IGARSS | 4 |
| 2022 | Pixel-Level and Feature-Level Domain Adaptation for Heterogeneous SAR Target RecognitionabstractThe performance of synthetic aperture radar (SAR) target recognition has been substantially enhanced by the deep learning technology. It is still difficult to get strong recognition performance due to the distribution differences across heterogeneous SAR images. To enhance target recognition performance in heterogeneous SAR situations, a pixel-level and feature-level domain adaptation (PFDA) approach is proposed in this letter to deal with this problem. The pixel-level translation module in the first step generates images with high visual similarity with another distributed images from one distributed images. The second step involves introducing feature alignment into the network to lower the probability distribution divergence of heterologous SAR images in the feature space and enhance recognition performance. We evaluated our method on the Synthetic and Measured Paired Labeled Experiment (SAMPLE) dataset and a self-built airplane dataset to ensure its effectiveness. Compared to existing domain adaption approaches, experimental results reveal that our strategy greatly improves recognition performance and model stability for heterogeneous SAR target recognition. Zhuo Chen 0031, Lingjun Zhao, Qishan He, Gangyao Kuang |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2022 | Target Region Segmentation in SAR Vehicle Chip Image With ACM NetabstractTarget region segmentation of synthetic aperture radar (SAR) images is one of the challenging problems in SAR image interpretation. The existing conventional segmentation methods rely on parameter selection in different backgrounds. Compared with traditional methods, the deep-learning-based methods can reduce the dependency on parameters and achieve more accurate results. However, lacking annotation data limits the application of the deep-learning-based methods in SAR chip image segmentation aspect. To solve these problems, a refined network structure for SAR vehicle image semantic segmentation, namely, All-Convolutional networks (A-ConvNets)-based Mask (ACM) net, is proposed. The mask in the training dataset of the network is extracted from image reconstruction using the Attribute Scattering Center (ASC) model, which can solve the problem of the lack of manual annotation in the segmentation methods based on deep learning. The proposed ACM Net consists of a modified A-ConvNets-based backbone and two decoupled head branches which achieve target segmentation and label prediction results, respectively. Experiments on moving and stationary target acquisition and recognition (MSTAR) dataset show that the comprehensive segmentation performance of ACM Net is better than both traditional segmentation methods and deep-learning-based segmentation methods. The classification results outperform other instance or semantic segmentation methods with the state-of-the-art recognition accuracy. Sijia Feng, Kefeng Ji, Xiaojie Ma, Linbin Zhang, Gangyao Kuang |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2022 | SAR Target Recognition Based on Task-Driven Domain Adaptation Using Simulated DataabstractSynthetic aperture radar (SAR) images are highly susceptible to imaging conditions. However, the majority of deep learning (DL) models in SAR automatic target recognition (ATR) adopt enhanced network structures similar to those in dealing with optical image classification tasks, which is obviously unreasonable since the huge gap of the imaging conditions between training and testing data severely deteriorates the recognition performance. The main idea of the framework is to introduce SAR imaging condition information into the DL training stage to eliminate domain discrepancies between training and testing data. Based on this framework, we propose a task-driven domain adaptation (TDDA) transfer learning method, which can alleviate the degradation of recognition caused by the variance of depression angle between training and testing data. In order to introduce the prior imaging information into the method, simulated SAR data is first obtained by adding a simulated object radar reflectivity to a terrain model of individual point scatters using the known training and testing SAR imaging parameters. Then a domain confusion metric and a supervised classification loss are calculated on simulated data and source training data, respectively, to learn a representation that is semantically meaningful and domain invariant. Comparative experiments on the moving and stationary target acquisition and recognition (MSTAR) dataset demonstrate that the proposed method can obtain better recognition performance than the other methods. Qishan He, Lingjun Zhao, Kefeng Ji, Gangyao Kuang |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2022 | Adaptive Local Structure Consistency-Based Heterogeneous Remote Sensing Change DetectionabstractChange detection (CD) of heterogeneous remote sensing images is a challenging topic, which plays an important role in natural disaster emergency response. Due to the different imaging mechanisms of heterogeneous sensors, it is hard to directly compare the images. To address this challenge, we explore an unsupervised CD method based on adaptive local structure consistency (ALSC) between heterogeneous images in this letter, which constructs an adaptive graph representing the local structure for each patch in one image domain and then projects this graph to the other image domain to measure the change level. This local structure consistency exploits the fact that the heterogeneous images share the same structure information for the same ground object, which is imaging modality-invariant. To avoid heterogeneous data confusion, the pixelwise change image is calculated in the same image domain by graph projection. By comparing with some state-of-the-art methods, the experimental results show the effectiveness of the proposed ALSC-based CD method. Lin Lei, Yuli Sun, Gangyao Kuang |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2022 | Locality-Constrained Bilinear Network for Land Cover Classification Using Heterogeneous ImagesabstractOptical and SAR modalities provide complementary information of land properties, which can lead to outstanding classification performance. Recently, factorized bilinear coding (FBC) as an extension of bilinear pooling in respect of coding-pooling perspective, which extracted compact bilinear fusion features with second-order interaction information in the form of sparse representation, brought the performance improvements on multimodal learning tasks. However, it lost locality attributes among similar samples to be encoded. In this letter, we propose a novel locality-constrained bilinear network (LC-BNet) for land cover classification with heterogeneous remote sensing (RS) images. Specifically, the locality-constrained bilinear coding (LC-BC) introduces locality information to generate compact and discriminative fusion features for land cover classification. Extensive experimental results show superior performances of our work on two broad coregistered optical and SAR datasets. Xiao Li 0017, Lin Lei, Gangyao Kuang |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2022 | Multilevel Adaptive-Scale Context Aggregating Network for Semantic Segmentation in High-Resolution Remote Sensing ImagesabstractHigh-resolution remote sensing (HR2S) images contain complex land objects of difference sizes, and it is important for semantic segmentation of the HR2S images to extract multiscale information. In this letter, we introduce a novel multilevel adaptive-scale context aggregating network (MACANet) for semantic segmentation of the HR2S images, which mainly consists of two parts—adaptive-scale context extraction block (AS-CEB) and sequential aggregation block (SAB). In particular, the AS-CEB introduces an inflexible strategy to obtain the features with appropriate scale information based on different asymmetric convolutions and the gated mechanism. Meanwhile, the SAB progressively aggregates multilevel adaptive-scale features, which are used to relieve the semantic gap between different-level features and generate precise score maps. Experimental results on representative HR2S datasets show the advantages of our method. The code is available athttps://github.com/RSIP-NUDT/MACANet. Xiao Li 0017, Lin Lei, Gangyao Kuang |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2022 | A Cross-Layer Nonlocal Network for Remote Sensing Scene ClassificationabstractRemote sensing scene classification (RSSC) is a fundamental yet challenging task in the domain of remote sensing (RS). Currently, the methods based on deep features from convolutional neural networks (CNNs) have significantly improved the scene classification accuracy (ACC). However, the standard convolution operations have limited capacity to model the long-range correlations and cannot effectively obtain global contextual understanding ability. In this letter, we propose a novel scene classification framework, termed cross-layer nonlocal network (CL-NL-Net), consisting of a backbone network, a cross-layer nonlocal (CL-NL) module, and a classifier. Among them, the backbone network is used to obtain multilayer convolutional features. The CL-NL module is the core of the proposed method, which captures the long-range correlations between different layers, so as to achieve a better global scene understanding ability. To verify the effectiveness of the proposed CL-NL-Net, we conduct experiments on four benchmark datasets, and the results demonstrate that the proposed method achieves competitive classification ACC and outperforms some state-of-the-art methods. Ming Li 0066, Lin Lei, Yuli Sun, Xiao Li 0017, Gangyao Kuang |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2022 | An Open Set Recognition Method for SAR Targets Based on Multitask LearningabstractMost of the existing synthetic aperture radar (SAR) automatic target recognition (ATR) methods aim at the closed set situation, in which the classes of targets in the test set have appeared in the training set. However, in practice, the classifier is likely to encounter the targets from unseen categories and classify them incorrectly, which brings a huge challenge to current SAR ATR techniques. To overcome this problem, this letter proposes an open set recognition (OSR) method based on multitask learning, and the method is developed from generative adversarial network (GAN). Essentially, this method decomposes OSR into two tasks: classification and abnormal detection. The classification task is the same as that in the closed set situation, while the abnormal detection task is used to determine whether the targets belongs to the unseen categories. Correspondingly, the network structure of GAN is modified and the other full-connection network branch is added to the end of the discriminator, so it has the ability to accomplish the above two tasks. Finally, according to the results of two tasks, the OSR for SAR targets can be realized. The experimental results on moving and stationary target acquisition (MSTAR) dataset demonstrate that the proposed method has the better recall, precision,$F1$, and accuracy than other OSR methods. Xiaojie Ma, Kefeng Ji, Linbin Zhang, Sijia Feng, Boli Xiong, Gangyao Kuang |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2022 | Optical and SAR Image Matching Using Pixelwise Deep Dense FeaturesabstractImage matching is a primary technology to fuse the complementary information from optical and SAR images. Due to the high nonlinear radiometric and geometric relationship, the optical and SAR image matching task remains a widely unsolved challenge. In this study, we propose to use a Siamese convolutional neural network (CNN) architecture to learn pixelwise deep dense features. The proposed network is able to balance the learning of high-level semantic information and low-level fine-grained information, which is nonnegligible for feature matching task. Under the local searching framework, the loss function is defined based on the score map produced by the sum of squared differences (SSDs) between the learned pixelwise dense features of local optical and the SAR image patches, with a fast implementation in the frequency domain. The hardest negative mining strategy is adopted to increase the discrimination of the network. Extensive experiments are conducted on optical and SAR image pairs of different spatial resolution and different landcover types, verifying the superiority and robustness of the proposed method in terms of matching accuracy and matching precision. Han Zhang 0005, Lin Lei, Weiping Ni, Tao Tang 0006, Junzheng Wu, Deliang Xiang, Gangyao Kuang |
IEEE Geosci. Remote. Sens. Lett. | 7 |
| 2022 | Aspect-Ratio-Guided Detection for Oriented Objects in Remote Sensing ImagesabstractAlthough existing oriented object detection methods have made considerable progress based on oriented heads or anchors, the training process itself is not perfect. In this letter, we point out the inconsistency problem between the fixed network setting and varying aspect ratios, which greatly limits the performance. For example, the fixed parameters in label assignment and regression loss cannot fit the changes of aspect ratios and, thus, are harmful to the training process. Considering the prior information about objects’ aspect ratios, the aspect-ratio-guided (ARG) methods are proposed. Specifically, the ARG label assignment is used to adjust the label assignment criteria (intersection over union (IoU) threshold) automatically, and the ARG IoU loss can change the weights of angle regression dynamically. This ARG design makes better use of training samples and pushes the detector more robust to the change of aspect ratios. With no additional cost, our method improves upon the ResNet-50-feature pyramid network (FPN) baseline with 3.99% AP50 and 6.09% AP75 on HRSC2016. Caiguang Zhang, Boli Xiong, Xiao Li 0017, Gangyao Kuang |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2022 | Novel Loss Function in CNN for Small Sample Target Recognition in SAR ImagesabstractStudying synthetic aperture radar automatic target recognition (SAR-ATR) under small samples can get rid of the sample dependence and improve the practicality of the deep learning model. However, the deep learning model trained with small samples is prone to overfitting. In order to solve the above problem of SAR-ATR, a novel loss function called limited data loss function (LDLF) is proposed in this letter, which organically combines the cross-entropy loss function and the contrastive loss function. The LDLF supervises the convolutional neural network (CNN) to learn strong generalization performance features. Then, for simplifying the training and testing of CNN based on LDLF, a feature-combined module is proposed. This module makes up for the limitation that the model input must be image pairs and simplifies the testing process of CNN. Experiments on moving and stationary target acquisition and recognition (MSTAR) datasets show that the proposed loss function is better than the cross-entropy loss function and superior to the existing methods in synthetic aperture radar (SAR) image target recognition using small samples data. Tao Tang 0006, Yuting Cui, Linbin Zhang, Gangyao Kuang |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2022 | Iterative structure transformation and conditional random field based method for unsupervised multimodal change detection
Yuli Sun, Lin Lei, Dongdong Guan, Junzheng Wu, Gangyao Kuang |
Pattern Recognit. | 5 |
| 2022 | Deep Ladder-Suppression Network for Unsupervised Domain AdaptationabstractUnsupervised domain adaptation (UDA) aims at learning a classifier for an unlabeled target domain by transferring knowledge from a labeled source domain with a related but different distribution. Most existing approaches learn domain-invariant features by adapting the entire information of the images. However, forcing adaptation of domain-specific variations undermines the effectiveness of the learned features. To address this problem, we propose a novel, yet elegant module, called the deep ladder-suppression network (DLSN), which is designed to better learn the cross-domain shared content by suppressing domain-specific variations. Our proposed DLSN is an autoencoder with lateral connections from the encoder to the decoder. By this design, the domain-specific details, which are only necessary for reconstructing the unlabeled target data, are directly fed to the decoder to complete the reconstruction task, relieving the pressure of learning domain-specific variations at the later layers of the shared encoder. As a result, DLSN allows the shared encoder to focus on learning cross-domain shared content and ignores the domain-specific variations. Notably, the proposed DLSN can be used as a standard module to be integrated with various existing UDA frameworks to further boost performance. Without whistles and bells, extensive experimental results on four gold-standard domain adaptation datasets, for example: 1) Digits; 2) Office31; 3) Office-Home; and 4) VisDA-C, demonstrate that the proposed DLSN can consistently and significantly improve the performance of various popular UDA frameworks. Wanxia Deng, Lingjun Zhao, Gangyao Kuang, Dewen Hu, Matti Pietikäinen, Li Liu 0002 |
IEEE Trans. Cybern. | 3 |
| 2022 | Electromagnetic Scattering Feature (ESF) Module Embedded Network Based on ASC Model for Robust and Interpretable SAR ATRabstractDeep learning has been widely used in automatic target recognition (ATR) for synthetic aperture radar (SAR) recently. However, most of the studies are based on the network structure in optical images and lack full consideration of the inherent characteristics of SAR targets, which limits the improvement of recognition accuracy and makes poor generalization ability. In addition, due to the black-box characteristics, it is difficult to effectively interpret SAR ATR results. To conquer these problems, we propose an electromagnetic scattering feature (ESF) module embedded network based on attributed scattering center (ASC) model to incorporate the SAR targets’ characteristics into the deep learning framework. First, a novel convolutional neural network (CNN)-based algorithm for extracting ASC parameters is proposed, which makes the network focus on target features under the guidance of physical model. Then, the ESF module is designed based on a well-trained ASC parameters extractor to inject the learned target features into the classification network for more robust and interpretable results. Besides, two structures are proposed combined with the ESF module for single-view and multiview SAR target classification, which further illustrates the portability of the module. Experiments on the Moving and Stationary Target Acquisition and Recognition (MSTAR) dataset show the validity of the proposed CNN-based ASC extractor and the ESF module embedded classification network. Compared with ordinary networks, our method can achieve higher classification accuracy under complex conditions, which reflects the better generalization performance of the algorithm. Furthermore, through visualization analysis of the classification results, we show the interpretability of the network combined with the electromagnetic scattering characteristics. Sijia Feng, Kefeng Ji, Fulai Wang, Linbin Zhang, Xiaojie Ma, Gangyao Kuang |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2022 | Complex Signal Kurtosis - Indicator of Ship Target Signature in SAR ImagesabstractSynthetic aperture radar (SAR) signatures of ship targets are often degraded by various types of distortions due to the distinctive imaging mechanism. The negative impacts could affect the characterization and identification of ships. Recently, complex signal kurtosis (CSK) was found to be a vital indicator of ship detection in SAR images. Since CSK can be used as an indicator of ship detection, we consider it can also be used to indicate ship target signature. One of the basic rationales is that the larger the CSK is, the easier it is to be detected and thus the more obvious its characteristics are as a ship. Defocusing and sidelobes are two common problems that affect the quality of ships. Their impacts on ship target signature are studied and compared from the perspective of CSK. Specifically, a ship target signature improvement methodology based on the maximum CSK criterion is proposed for both refocusing and sidelobe suppression. The role the CSK plays in these improvements and the underlying rationales are elaborated. In addition, a preliminary evaluation of the global imaging quality is also provided. Experimental results based on real data demonstrate that CSK can be used to indicate and improve ship target signature. Xiangguang Leng, Kefeng Ji, Boli Xiong, Gangyao Kuang |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Dynamic-Hierarchical Attention Distillation With Synergetic Instance Selection for Land Cover Classification Using Missing Heterogeneity ImagesabstractOptical and SAR modalities can provide the complementary information on the land properties, which usually lead to more robust and better classification performance. However, due to the restriction of imaging condition, not all modalities included into the training data sets could be available in real testing samples. Therefore, it is important to explore how to learn discriminative representations using multimodal data during the training stage, while achieving fine land cover classification using missing modalities at test time. In this article, we propose a novel dynamic-hierarchical attention distillation network (DH-ADNet) with multimodal synergetic instance selection (MSIS) for land cover classification using missing data modalities. First, the MSIS realizes the selection of the most representative multimodal instances to enhance the DH-ADNet’s ability of discriminative feature extraction. Then, the DH-ADNet is training on the basis of the curriculum learning strategy and promotes the hallucination stream to learn the privileged information. In particular, a novel dynamic-hierarchical attention distillation module (DH-ADM) is introduced, which adaptively highlights different contributions of multilayer attention distillation by carefully exploring the classification losses of multilayer features over the training iterations. Comprehensive evaluations on two coregistered optical and SAR data sets and report state-of-the-art results in the privileged information scenario. Xiao Li 0017, Lin Lei, Yuli Sun, Gangyao Kuang |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Dense Adaptive Grouping Distillation Network for Multimodal Land Cover Classification With Privileged ModalityabstractMultimodal land cover classification (MLCC) is a fundamental problem in remote sensing interpretation, which can obtain excellent performance on account of the complementary information between the optical and SAR modalities. However, it is usually impossible to obtain multimodal data at the same time, due to the restriction of imaging conditions. When one of the modalities data is completely missing during test phase, classical multimodal learning methods might not be able to handle the MLCC task with privileged modality. In this paper, we propose an efficient Dense Adaptive Grouping Distillation Network (DAGDNet), which learns privileged information from available modalities in the train sets, and improves the classification performance in the test sets when one modality data is scarce. More specifically, to relieve the heterogeneous gaps between different modalities and then transfer the privileged information, we propose an Interactive Gated-based Feature Grouping Module (IG-FGM), which decomposes multimodal features into modalities-shared and modality-specific components to realize the decoupling of multimodal features and grouping distillation. Furthermore, the IG-FGM is inserted into different layers of the “teacher" network to implement progressive blending of multi-modalities. Then, to adaptively highlight the importance of hierarchical features distillation and grouping distillation, we propose a Multi-stage Adaptive Distillation Learning (MS-ADL) strategy so that the weights of different distillation losses are required to change continuously along with the training process. Finally, we evaluate the superior performances of our model on representative co-registered optical and SAR datasets. Xiao Li 0017, Lin Lei, Caiguang Zhang, Gangyao Kuang |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Multimodal Semantic Consistency-Based Fusion Architecture Search for Land Cover ClassificationabstractMultimodal Land Cover Classification (MLCC) using the optical and Synthetic Aperture Radar (SAR) modalities has resulted in outstanding performances over using only unimodal data due to their complementary information on land properties. Previous multimodal deep learning (MDL) methods have relied on handcrafted multi-branch convolutional neural networks (CNN) to extract the features of different modalities and merged them for land cover classification. However, natural images-oriented handcrafted CNN models may not the optimal strategies to handle Remote Sensing (RS) image interpretation problems, due to the huge difference in terms of imaging angles and imaging ways. Furthermore, few MDL methods have analyzed optimal combinations of hierarchical features from different modalities. In this article, we propose an efficient multimodal architecture search framework, namely Multimodal Semantic Consistency-Based Fusion Architecture Search (M2SC-FAS) in continuous search space with the gradient-based optimization method, which can not only discover optimal optical- and SAR-specific architectures according to the different characteristics of the optical and SAR images, respectively, but also realizes the search of optimal multimodal dense fusion architecture. Specifically, the semantic-consistency constraint is introduced to guarantee dense fusion between hierarchical optical and SAR features with high semantic consistency and then capture the complementary performance on land properties. Finally, the basis of curriculum learning strategy is adopted on the M2SC-FAS. Extensive experiments show superior performances of our work on three broad co-registered optical and SAR datasets. Xiao Li 0017, Lin Lei, Caiguang Zhang, Gangyao Kuang |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Graph Signal Processing for Heterogeneous Change DetectionabstractThis paper provides a new strategy for the heterogeneous change detection (HCD) problem: solving HCD from the perspective of graph signal processing (GSP). We construct a graph to represent the structure of each image, and treat each image as a graph signal defined on the graph. In this way, we convert the HCD into a GSP problem: a comparison of the responses of signals on systems defined on the graphs, which attempts to find structural differences and signal differences due to the changes between heterogeneous images. Firstly, we analyze the GSP for HCD from the vertex domain. We show that once a region has changed, the local structure of image changes,i.e. the connectivity of the vertex containing this region changes. Therefore, we can compare the output signals of the same input graph signal passing through filters defined on the two graphs to detect changes. We analyze the negative effects of changing regions on the change detection results from the viewpoint of signal propagation, and we also design different filters from the vertex domain to explore the high-order neighborhood information hidden in original graphs. Secondly, we analyze the GSP for HCD from the spectral domain. We explore the spectral properties of different images on the same graph, and show that their spectra exhibit commonalities and dissimilarities. Specifically, it is the change that leads to the dissimilarities of their spectra. With the help of graph spectral analysis, we propose a regression model for the HCD, which decomposes the source signal into the regressed signal and changed signal, and constrains the spectral property of the regressed signal. Experiments conducted on seven real data sets show the effectiveness of the vertex domain filtering based and spectral domain analysis based HCD methods. Source code will be made available at https://github.com/yulisun/HCD-GSP. Yuli Sun, Lin Lei, Dongdong Guan, Gangyao Kuang, Li Liu 0002 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Sparse-Constrained Adaptive Structure Consistency-Based Unsupervised Image Regression for Heterogeneous Remote-Sensing Change DetectionabstractChange detection of heterogeneous multitemporal satellite images is an important and challenging topic in remote sensing. Since the imaging mechanisms of heterogeneous sensors are different, it is not possible to directly compare heterogeneous images to detect changes as in the homogeneous images. To address this challenge, we propose an unsupervised image regression-based change detection method based on the structure consistency. The proposed method first adaptively constructs a similarity graph to represent the structure of a pre-event image, then uses the graph to translate the pre-event image to the domain of the post-event image, and then computes the difference image. Finally, a superpixel-based Markovian segmentation model is designed to segment the difference image into changed and unchanged classes. The proposed adaptive structure consistency-based image regression model can not only alleviate the impact of noise and changed pixels on the regression process by using the structure-based transformation, but also easily distinguish between changed and unchanged classes in the difference image by using the prior sparse knowledge of changes. Experimental results on six different datasets demonstrate the effectiveness of the proposed method by comparing with some state-of-the-art methods. Yuli Sun, Lin Lei, Dongdong Guan, Ming Li 0066, Gangyao Kuang |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | Structure Consistency-Based Graph for Unsupervised Change Detection With Homogeneous and Heterogeneous Remote Sensing ImagesabstractChange detection (CD) of remote sensing (RS) images is one of the important problems in earth observation, which has been extensively studied in recent years. However, with the development of RS technology, the specific characteristics of remotely sensed images, including sensor characteristics, resolutions, noises, and distortions in imagery, make the CD more complex. In this article, we propose a structure consistency-based method for CD, which detects changes by comparing the structures of two images, rather than comparing the pixel values of images. Because the image structure is imaging modality-invariant and not sensitive to noise, illumination, and other interference factors, the proposed method can be applied to a variety of CD scenarios and has strong robustness. Structural comparison is realized by constructing and mapping an improved nonlocal patch-based graph (NLPG) to avoid the data leakage of two images. First, we demonstrate the effectiveness of the method in homogeneous and heterogeneous CD, which shows that the proposed method can be used as a unified CD framework. Second, we extend the method to the heterogeneous CD with multichannel synthetic aperture radar (SAR) image, which can provide a reference for future research as the heterogeneous CD with multichannel SAR is rarely studied. Third, through the decomposition and in-depth analysis of NLPG, we modify the graph construction process, structure difference calculation, and the difference image fusion to make it more robust and accurate. Experiments on six scenarios 12 data sets demonstrate the effectiveness of the proposed method. Yuli Sun, Lin Lei, Xiao Li 0017, Xiang Tan, Gangyao Kuang |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | Domain Knowledge Powered Two-Stream Deep Network for Few-Shot SAR Vehicle RecognitionabstractSynthetic aperture radar (SAR) target recognition faces the challenge that there are very little labeled data. Although few-shot learning methods are developed to extract more information from a small amount of labeled data to avoid overfitting problems, recent few-shot or limited-data SAR target recognition algorithms overlook the unique SAR imaging mechanism. Domain knowledge-powered two-stream deep network (DKTS-N) is proposed in this study, which incorporates SAR domain knowledge related to the azimuth angle, the amplitude, and the phase data of vehicles, making it a pioneering work in few-shot SAR vehicle recognition. The two-stream deep network, extracting the features of the entire image and image patches, is proposed for more effective use of the SAR domain knowledge. To measure the structural information distance between the global and local features of vehicles, the deep Earth mover’s distance is improved to cope with the features from a two-stream deep network. Considering the sensitivity of the azimuth angle in SAR vehicle recognition, the nearest neighbor classifier replaces the structured fully connected layer for$K$-shot classification. All experiments are conducted under the configuration that the SARSIM and the Moving and Stationary Target Acquisition and Recognition (MSTAR) dataset work as a source and target task, respectively. Our proposed DKTS-N achieved 49.26% and 96.15% under ten-way one-shot and ten-way 25-shot, whose labeled samples are randomly selected from the training set. In standard operating condition (SOC) as well as three extended operating conditions (EOCs), DKTS-N demonstrated overwhelming advantages in accuracy and time consumption compared with other few-shot learning methods in$K$-shot recognition tasks. Linbin Zhang, Xiangguang Leng, Sijia Feng, Xiaojie Ma, Kefeng Ji, Gangyao Kuang, Li Liu 0002 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2022 | Explore Better Network Framework for High-Resolution Optical and SAR Image MatchingabstractTo fully explore the complementary information from optical and synthetic aperture radar (SAR) imageries, they need first to be coregistered with high accuracy. Due to the vast radiometric and geometric disparity, the problem to match high-resolution optical and SAR images is quite challenging. The present deep learning-based methods have shown advantages over the traditional approaches, but the performance increment is not significant. In this article, we explore a better network framework for high-resolution optical and SAR image matching from three aspects. First, we propose an effective multilevel feature fusion method, which helps to take advantage of both the low-level fine-grained features for precious feature location and the high-level semantic features for better discriminative ability. Second, a feature channel excitation procedure is conducted using a novel multifrequency channel attention module, which is able to make image features of different types and multiple levels effectively collaborate with each other and produce image matching features with high diversity. Third, the self-adaptive weighting loss is introduced, with which, each sample is assigned with an adaptive weighting factor, and therefore, information buried in all nearby samples can be better exploited. Under a pseudo-Siamese architecture, the proposed optical and SAR image matching network (OSMNet) is trained and tested on a large and diverse high-resolution optical and SAR dataset. Extensive experiments demonstrate that each component of the proposed deep framework helps to improve the matching accuracy. Also, the OSMNet shows overwhelming superior to the state-of-the-art handcrafted approaches on imageries of different land-cover types. Han Zhang 0005, Lin Lei, Weiping Ni, Tao Tang 0006, Junzheng Wu, Deliang Xiang, Gangyao Kuang |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2022 | Attentional Feature Refinement and Alignment Network for Aircraft Detection in SAR ImageryabstractAircraft detection in synthetic aperture radar (SAR) imagery is a challenging task in SAR automatic target recognition (SAR ATR) areas due to aircraft’s extremely discrete appearance, obvious intraclass variation, small size, and serious background’s interference. In this article, a single shot detector (SSD), namely, attentional feature refinement and alignment network (AFRAN), is proposed for detecting aircraft in SAR images with competitive accuracy and speed. Specifically, three significant components, including attention feature fusion module (AFFM), deformable lateral connection module (DLCM), and anchor-guided detection module (ADM), are carefully designed in our method for refining and aligning informative characteristics of aircraft. To represent the characteristics of aircraft with less interference, low-level textural and high-level semantic features of aircraft are fused and refined in AFFM thoroughly. The alignment between aircraft’s discrete backscatting points and convolutional sampling spots is promoted in DLCM. Eventually, the locations of aircraft are predicted precisely in ADM based on aligned features revised by refined anchors. To evaluate the performance of our method, a self-built SAR aircraft sliced dataset and a large scene SAR image are collected. Extensive quantitative and qualitative experiments with detailed analysis illustrate the effectiveness of the three proposed components. Furthermore, the topmost detection accuracy and competitive speed are achieved by our method compared with other domain-specific methods, e.g., dense attention pyramid network (DAPN) and pyramid attention dilated network (PADN), and general convolutional neural network (CNN)-based methods, e.g., Feature Pyramid Network (FPN), Cascade R-CNN, SSD, RefineDet, and RepPoints Detector (RPDet). Yan Zhao 0026, Lingjun Zhao, Zhong Liu 0002, Dewen Hu, Gangyao Kuang, Li Liu 0002 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | Informative Feature Disentanglement for Unsupervised Domain AdaptationabstractUnsupervised Domain Adaptation (UDA) aims at learning a classifier for an unlabeled target domain by transferring knowledge from a labeled source domain with a related but different distribution. The strategy of aligning the two domains in latent feature space via metric discrepancy or adversarial learning has achieved considerable progress. However, these existing approaches mainly focus on adapting the entire image and ignore the bottleneck that occurs when forced adaptation of uninformative domain-specific variations undermines the effectiveness of learned features. To address this problem, we propose a novel component called Informative Feature Disentanglement (IFD), which is equipped with the adversarial network or the metric discrepancy model, respectively. Accordingly, the new network architectures, named IFDAN and IFDMN, enable informative feature refinement before the adaptation. The proposed IFD is designed to disentangle informative features from the uninformative domain-specific variations, which are produced by a Variational Autoencoder (VAE) with lateral connections from the encoder to the decoder. We cooperatively apply the IFD to conduct supervised disentanglement for the source domain and unsupervised disentanglement for the target domain. In this way, informative features are disentangled from the domain-specific details before the adaptation. Extensive experimental results on three gold-standard domain adaptation datasets, e.g., Office31, Office-Home and VisDA-C, demonstrate the effectiveness of the proposed IFDAN and IFDMN models for UDA. Wanxia Deng, Lingjun Zhao, Qing Liao 0001, Deke Guo, Gangyao Kuang, Dewen Hu, Matti Pietikäinen, Li Liu 0002 |
IEEE Trans. Multim. | 5 |
| 2022 | Fine-grained visual classification via multilayer bilinear pooling with object localization
Ming Li 0066, Lin Lei, Hao Sun 0042, Xiao Li 0017, Gangyao Kuang |
Vis. Comput. | 5 |
| 2021 | Transferable Discriminative Feature Mining For Unsupervised Domain AdaptationabstractUnsupervised Domain Adaptation (UDA) aims to seek an effective model for unlabeled target domain by leveraging knowledge from a labeled source domain with a related but different distribution. Many existing approaches ignore the underlying discriminative features of the target data and the discrepancy of conditional distributions. To address these two issues simultaneously, the paper presents a Transferable Discriminative Feature Mining (TDFM) approach for UDA, which can naturally unify the mining of domain-invariant discriminative features and the alignment of class-wise features into one single framework. To be specific, to achieve the domain-invariant discriminative features, TDFM jointly learns a shared encoding representation for two tasks: supervised classification of labeled source data, and discriminative clustering of unlabeled target data. It then conducts the class-wise alignment by decreasing intra-class variations and increasing inter-class differences across domains, encouraging the emergence of transferable discriminative features. When combined, these two procedures are mutually beneficial. Comprehensive experiments verify that TDFM can obtain remarkable margins over state-of-the-art domain adaptation methods. Lingjun Zhao, Wanxia Deng, Gangyao Kuang, Dewen Hu, Li Liu 0002 |
ICIP | 3 |
| 2021 | Multi-Modal Fusion Architecture Search for Land Cover Classification Using Heterogeneous Remote Sensing ImagesabstractOptical and SAR modalities can provide the complementary information on land properties for better land cover classification. Most of existing multi-modal land cover classification methods based on two-streams convolutional neural networks (CNNs), which obtained fusion features by merging optical and SAR features that come from manually selective layer of different streams. However, they ignored different semantic between manually selective optical and SAR features, which might result in suboptimal fusion features. We tackle the problem of finding good fusion architectures for multimodal land cover classification inspired by the network architecture search (NAS), and introduces the multi-modal fusion architecture search network (M2PASNet). Extensive experimental results show superior performances of our work on a broad co-registered optical and SAR dataset. Xiao Li 0017, Lin Lei, Gangyao Kuang |
IGARSS | 3 |
| 2021 | Adversarial Robustness Evaluation of Deep Convolutional Neural Network Based SAR ATR AlgorithmabstractRobustness, both to accident and to malevolent perturbations, is a crucial determinant of the successful deployment of deep convolutional neural network based SAR ATR systems in various security-sensitive applications. This paper performs a detailed adversarial robustness evaluation of deep convolutional neural network based SAR ATR models across two public available SAR target recognition datasets. For each model, seven different adversarial perturbations, ranging from gradient based optimization to self-supervised feature distortion, are generated for each testing image. Besides adversarial average recognition accuracy, feature attribution techniques have also been adopted to analyze the feature diffusion effect of adversarial attacks, which promotes the understanding of vulnerability of deep learning models. Hao Sun 0042, Gangyao Kuang |
IGARSS | 3 |
| 2021 | Robust Remote Sensing Scene Classification by Adversarial Self-Supervised LearningabstractAdversarial training is an effective method to enhance adversarial robustness for deep neural networks. However, it qequires large amounts of labeled data, which are often difficult to acquire. Recent research has shown that self-supervised learning can help to improve model performance and model uncertainty using unlabeled data. In this paper, we introduce a new adversarial self-supervised learning framework to learn a robust pretrained model for remote sensing scene classification. The proposed method exploits the advantage of dual network structure, and it requires neither labeled data for adversarial example generation nor negative samples for contrastive learning. Specifically, it consists of three major steps. Firstly, we train the online model and the target model to extract deep image features. Secondly, we generate two kinds of instance-wise adversarial examples. Finally, we iteratively learn a robust model by implicit comparing the difference between clean data and their perturbed counterpart. Preliminary experimental results on remote sensing scene classification dataset shows that our method can obtain higher robust accuracy. Our method can also be combined with other adversarial defense techniques to further promote model robustness. Hao Sun 0042, Lin Lei, Gangyao Kuang, Kefeng Ji |
IGARSS | 5 |
| 2021 | Ship Detection and Recognition in Optical Remote Sensing Images Based on Scale Enhancement Rotating Cascade R-CNN NetworksabstractShip detection and recognition in remote sensing images have important significance in military and civilian applications. Traditional methods have insufficient generalization ability in complicated scenes. The Faster R-CNN-based methods cannot predict the orientation of the ship. The R2CNN-based methods can predict the orientation of the ship but not considerate the scale of object in classification. In order to solve the problems mentioned above, this article proposes a scale enhancement rotating Cascade R-CNN network (SER-Cascade). Using the multistage network of rotating Cascade R-CNN, the output of the previous stage is fed to current stage, which can effectively regress the orientation of the ship. To improve the recognition performance of multi-class ships, a novel RoI pooling method is proposed in this article, in which the scale information is enhanced and context information is reserved. To evaluate the proposed networks, a dataset named HR-SHIP-15 that currently contains 15 categories of ship targets has been produced for ship recognition. Experiments are conducted on HR-SHIP-15 dataset, and the results verify that the proposed method has state-of-the-art performance. Caiguang Zhang, Boli Xiong, Gangyao Kuang |
IGARSS | 3 |
| 2021 | Informative Class-Conditioned Feature Alignment for Unsupervised Domain AdaptationabstractThe goal of unsupervised domain adaptation is to learn a task classifier that performs well for the unlabeled target domain by borrowing rich knowledge from a well-labeled source domain. Although remarkable breakthroughs have been achieved in learning transferable representation across domains, two bottlenecks remain to be further explored. First, many existing approaches focus primarily on the adaptation of the entire image, ignoring the limitation that not all features are transferable and informative for the object classification task. Second, the features of the two domains are typically aligned without considering the class labels; this can lead the resulting representations to be domain-invariant but non-discriminative to the category. To overcome the two issues, we present a novel Informative Class-Conditioned Feature Alignment (IC2FA) approach for UDA, which utilizes a twofold method: informative feature disentanglement and class-conditioned feature alignment, designed to address the above two challenges, respectively. More specifically, to surmount the first drawback, we cooperatively disentangle the two domains to obtain informative transferable features; here, Variational Information Bottleneck (VIB) is employed to encourage the learning of task-related semantic representations and suppress task-unrelated information. With regard to the second bottleneck, we optimize a new metric, termed Conditional Sliced Wasserstein Distance (CSWD), which explicitly estimates the intra-class discrepancy and the inter-class margin. The intra-class and inter-class CSWDs are minimized and maximized, respectively, to yield the domain-invariant discriminative features. IC2FA equips class-conditioned feature alignment with informative feature disentanglement and causes the two procedures to work cooperatively, which facilitates informative discriminative features adaptation. Extensive experimental results on three domain adaptation datasets confirm the superiority of IC2FA. Wanxia Deng, Yawen Cui, Zhen Liu 0004, Gangyao Kuang, Dewen Hu, Matti Pietikäinen, Li Liu 0002 |
ACM Multimedia | 4 |
| 2021 | Pyramid Attention Dilated Network for Aircraft Detection in SAR ImagesabstractRecently, deep learning based methods have been successfully applied in synthetic aperture radar automatic target recognition (SAR ATR) fields. However, due to the effects of the special structures of aircrafts and the complexity of SAR imaging mechanism, detecting aircrafts accurately in SAR images is still challenging. To alleviate this problem, a novel network called pyramid attention dilated network (PADN) is proposed in this letter. The key component of PADN is the dilated attention block (DAB), which is composed of two submodules - multibranch dilated convolution module (MBDCM) and convolution block attention module (CBAM). In our method, MBDCM is used to enhance the relationship among discrete backscattering features of aircrafts. CBAM is employed to refine redundant information and highlight significant features of aircrafts. A well-designed fine-grained feature pyramid is established by combining the two modules reasonably into DAB when building lateral connections. To alleviate class imbalance, focal loss (FL) is employed to train our network. Experiments on a mixed SAR aircraft data set illustrate the efficiency of the proposed method for aircraft detection. Yan Zhao 0026, Lingjun Zhao, Chuyin Li, Gangyao Kuang |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2021 | Nonlocal patch similarity based heterogeneous remote sensing change detection
Yuli Sun, Lin Lei, Xiao Li 0017, Hao Sun 0042, Gangyao Kuang |
Pattern Recognit. | 5 |
| 2021 | Deep ladder reconstruction-classification network for unsupervised domain adaptation
Wanxia Deng, Zhuo Su 0002, Qiang Qiu 0001, Lingjun Zhao, Gangyao Kuang, Matti Pietikäinen, Huaxin Xiao, Li Liu 0002 |
Pattern Recognit. Lett. | 5 |
| 2021 | Radio Frequency Interference Detection and Localization in Sentinel-1 ImagesabstractThe C-band Sentinel-1 synthetic aperture radar (SAR) images are affected by radio frequency interference (RFI), and this article proposes an RFI detection and localization method for these images. First, a detailed analysis of RFI based on the generation mechanism is provided, which shows that RFI in the ground range detected (GRD) products differs from common backscattering not only in frequency spectrums but also in radiometric characteristics. Then, an RFI index (RFII) is proposed for RFI detection, which takes full advantage of unique RFI characteristics in dual-polarization GRD images. Finally, both detection results in ascending and descending passes are used to locate the ground RFI sources, assuming that the presence of RFI is persistent. Theoretical analyses show that the location area is a diamond of approximately 88.76 km2. Experimental results demonstrate that the proposed method performs quite well on Sentinel-1 GRD data. It provides a convenient and tractable tool for assessing Sentinel-1 data quality as well as for monitoring RFI. Thus, Sentinel-1 measurements can be used to monitor C-band RFI, which is an important task of electromagnetic spectrum management. Xiangguang Leng, Kefeng Ji, Gangyao Kuang |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2021 | Collaborative Attention-Based Heterogeneous Gated Fusion Network for Land Cover ClassificationabstractExisting land cover classification methods mostly rely on either the optical or synthetic aperture radar (SAR) features alone, which ignore the mutual complementary effects between optical and SAR sources. In this article, we compare the distribution histograms of deep semantic features extracted from optical and SAR modalities within land cover categories, which intuitively demonstrates that there are the large complementary potentials between the optical and SAR features. Therefore, we propose a novel collaborative attention-based heterogeneous gated fusion network (CHGFNet), which hierarchically fuses both optical and SAR features for land cover classification. More specifically, the CHGFNet consists of three main components: two-stream feature extractor, multimodal collaborative attention module (MCAM), and the gated heterogeneous fusion module (GHFM). Given optical and SAR patch pairs, two-stream feature extractor introduces multistage feature learning methodology to acquire discriminative optical and SAR features. Then, to explore the inherent complementarity between optical and SAR features, MCAM is embedded into CHGFNet, which provides an efficient stage to capture the correlation between optical and SAR features by jointly calculating the collaborative attention in joint feature space. Finally, to automatically learn the varying contributions of both optical and SAR features for classifying different land categories, GHFM is used to fuse both optical and SAR features. Extensive comparative evaluations demonstrate the advantages of CHGFNet within land cover classification over the state-of-the-art methods on three co-registered optical and SAR data sets. Xiao Li 0017, Lin Lei, Yuli Sun, Ming Li 0066, Gangyao Kuang |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2021 | SAR Image Speckle Reduction Based on Nonconvex Hybrid Total Variation ModelabstractSpeckle noise inherent in synthetic aperture radar (SAR) images seriously affects the visual effect and brings great difficulties to the postprocessing of the SAR image. Due to the edge-preserving feature, total variation (TV) regularization-based techniques have been extensively utilized to reduce the speckle. However, the strong scatters in SAR image with radiometry several orders of magnitude larger than their surrounding regions limit the effectiveness of TV regularization. Meanwhile, the ℓ1-norm first-order TV regularization sometimes causes staircase artifacts as it favors solutions that are piecewise constant, and it usually underestimates high-amplitude components of image gradient as the ℓ1-norm uniformly penalizes the amplitude. To overcome these shortcomings, a new hybrid variation model, called Fisher-Tippett (FT) distribution-ℓp-norm first-and second-order hybrid TVs (HTpVs), is proposed to reduce the speckle after removing the strong scatters. Especially, the FT-HTpV inherits the advantages of the distribution based data fidelity term, the nonconvex regularization, and the higher order TV regularization. Therefore, it can effectively remove the speckle while preserving point scatters and edges and reducing staircase artifacts well. To efficiently solve the nonconvex minimization problem, an iterative framework with a nonmonotone-accelerated proximal gradient (nmAPG) method and a matrix-vector acceleration strategy are used. Extensive experiments on both the simulated and real SAR images demonstrate the effectiveness of the proposed method. Yuli Sun, Lin Lei, Dongdong Guan, Xiao Li 0017, Gangyao Kuang |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2021 | Patch Similarity Graph Matrix-Based Unsupervised Remote Sensing Change Detection With Homogeneous and Heterogeneous SensorsabstractChange detection (CD) of remote sensing images is an important and challenging topic, which has found a wide range of applications in many fields. In particular, one of the main challenges is to detect changes between heterogeneous images, where the difference in imaging mechanism makes it difficult to carry out a direct comparison. In this article, we propose an unsupervised CD framework based on the patch similarity graph matrix (PSGM), which assumes that the patch similarity graph structure of each homogeneous or heterogeneous image is consistent if no change occurs. First, it learns the PSGM of one image based on the self-expressive property, which can be interpreted as containing the edges of the fully connected graphs with each image patch as a vertex. Then, the change level depends on how much one image still conforms to the similarity graph structure learned from the other image. Meanwhile, the change map can be further optimized by using the prior sparse knowledge that only a small part of the image changed and most areas remain unchanged. Experiments with both homogeneous and heterogeneous data sets demonstrate the effective performance of the proposed PSGM-based CD method. Yuli Sun, Lin Lei, Xiao Li 0017, Xiang Tan, Gangyao Kuang |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2021 | Joint Clustering and Discriminative Feature Alignment for Unsupervised Domain AdaptationabstractUnsupervised Domain Adaptation (UDA) aims to learn a classifier for the unlabeled target domain by leveraging knowledge from a labeled source domain with a different but related distribution. Many existing approaches typically learn a domain-invariant representation space by directly matching the marginal distributions of the two domains. However, they ignore exploring the underlying discriminative features of the target data and align the cross-domain discriminative features, which may lead to suboptimal performance. To tackle these two issues simultaneously, this paper presents a Joint Clustering and Discriminative Feature Alignment (JCDFA) approach for UDA, which is capable of naturally unifying the mining of discriminative features and the alignment of class-discriminative features into one single framework. Specifically, in order to mine the intrinsic discriminative information of the unlabeled target data, JCDFA jointly learns a shared encoding representation for two tasks: supervised classification of labeled source data, and discriminative clustering of unlabeled target data, where the classification of the source domain can guide the clustering learning of the target domain to locate the object category. We then conduct the cross-domain discriminative feature alignment by separately optimizing two new metrics: 1) an extended supervised contrastive learning, i.e., semi-supervised contrastive learning 2) an extended Maximum Mean Discrepancy (MMD), i.e., conditional MMD, explicitly minimizing the intra-class dispersion and maximizing the inter-class compactness. When these two procedures, i.e., discriminative features mining and alignment are integrated into one framework, they tend to benefit from each other to enhance the final performance from a cooperative learning perspective. Experiments are conducted on four real-world benchmarks (e.g., Office-31, ImageCLEF-DA, Office-Home and VisDA-C). All the results demonstrate that our JCDFA can obtain remarkable margins over state-of-the-art domain adaptation methods. Comprehensive ablation studies also verify the importance of each key component of our proposed algorithm and the effectiveness of combining two learning strategies into a framework. Wanxia Deng, Qing Liao 0001, Lingjun Zhao, Deke Guo, Gangyao Kuang, Dewen Hu, Li Liu 0002 |
IEEE Trans. Image Process. | 5 |
| 2021 | Iterative Robust Graph for Unsupervised Change Detection of Heterogeneous Remote Sensing ImagesabstractThis work presents a robust graph mapping approach for the unsupervised heterogeneous change detection problem in remote sensing imagery. To address the challenge that heterogeneous images cannot be directly compared due to different imaging mechanisms, we take advantage of the fact that the heterogeneous images share the same structure information for the same ground object, which is imaging modality-invariant. The proposed method first constructs a robust K -nearest neighbor graph to represent the structure of each image, and then compares the graphs within the same image domain by means of graph mapping to calculate the forward and backward difference images, which can avoid the confusion of heterogeneous data. Finally, it detects the changes through a Markovian co-segmentation model that can fuse the forward and backward difference images in the segmentation process, which can be solved by the co-graph cut. Once the changed areas are detected by the Markovian co-segmentation, they will be propagated back into the graph construction process to reduce the influence of changed neighbors. This iterative framework makes the graph more robust and thus improves the final detection performance. Experimental results on different data sets confirm the effectiveness of the proposed method. Source code of the proposed method is made available at https://github.com/yulisun/IRG-McS. Yuli Sun, Lin Lei, Dongdong Guan, Gangyao Kuang |
IEEE Trans. Image Process. | 4 |
| 2020 | Ship Target Signature Indication based on Complex Signal Kurtosis in SAR ImagesabstractRecently, complex signal kurtosis (CSK) is found to be a vital indicator of ship detection in synthetic aperture radar (SAR) images. Since CSK can be used as an indicator of ship detection, we consider it can also be used to indicate ship target signature. One of the basic rationales is that the higher the CSK is, the easier it is to be detected and thus the more obvious its characteristics are as a ship. Presence of sidelobes and defocusing are two common problems that affect the quality of ship targets. They are discussed in this paper from the perspective of CSK. Experimental results show that CSK can be used to indicate and improve ship quality. We believe that ship detection and recognition can benefit from the CSK indicator. Xiangguang Leng, Kefeng Ji, Boli Xiong, Gangyao Kuang |
IGARSS | 4 |
| 2020 | A robust recovery algorithm with smoothing strategies
Yuli Sun, Lin Lei, Xiao Li 0017, Ming Li 0066, Gangyao Kuang |
Neurocomputing | 5 |
| 2020 | A SAR Image Despeckling Method Using Multi-Scale Nonlocal Low-Rank ModelabstractSpeckle noise is an inherent nature of synthetic aperture radar (SAR) images, which degrades the quality of the images and makes the interpretation of SAR images difficult. In this letter, we propose a despeckling method by simultaneously exploring low-rank prior and multi-scale prior of SAR images. Especially, we propose a low-rank minimization model by considering a data fidelity term derived from the Fisher-Tippett distribution and a weighted nuclear norm regularization term. Furthermore, we explore the multi-scale prior by selecting similar patches from different scales of the SAR image. The resulting optimization problem is solved by the alternating direction method of multipliers (ADMM). Experiments conducted on both simulated and real SAR images demonstrate that the proposed method can provide promising despeckling results in terms of speckle reduction and texture and edge details preservation. Dongdong Guan, Deliang Xiang, Xiaoan Tang, Gangyao Kuang |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2020 | Edge Detection for PolSAR Images Integrating Scattering Characteristics and Optimal ContrastabstractSubject to the statistical distribution assumption and the fixed window shape, the classical edge detectors for polarimetric synthetic aperture radar (PolSAR) images generally generate inaccurate results in heterogeneous scenes. In this letter, a PolSAR image edge detector integrating the scattering characteristics and optimal contrast (OC) is proposed. Hierarchical model-based decomposition is first implemented for the scattering mechanism characterization. On this basis, a scattering mechanism-driven adaptive window is then designed, which contains pixels with uniform polarimetric scattering. Finally, to avoid making the assumption, the OC measurement is adopted for the edge strength calculation. Experimental results conducted on different PolSAR data confirm the effectiveness of the proposed method and its superiority over the classical edge detectors, especially in heterogeneous areas. Sinong Quan, Deliang Xiang, Boli Xiong, Gangyao Kuang |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2020 | Sparse optimization problem with s-difference regularization
Yuli Sun, Xiang Tan, Xiao Li 0017, Lin Lei, Gangyao Kuang |
Signal Process. | 5 |
| 2019 | Squeeze and Excitation Rank Faster R-CNN for Ship Detection in SAR ImagesabstractSynthetic aperture radar (SAR) ship detection is an important part of marine monitoring. With the development in computer vision, deep learning has been used for ship detection in SAR images such as the faster region-based convolutional neural network (R-CNN), single-shot multibox detector, and densely connected network. In SAR ship detection field, deep learning has much better detection performance than traditional methods on nearshore areas. This is because traditional methods need sea-land segmentation before detection, and inaccurate sea-land mask decreases its detection performance. Though current deep learning SAR ship detection methods still have many false detections in land areas, and some ships are missed in sea areas. In this letter, a new network architecture based on the faster R-CNN is proposed to further improve the detection performance by using squeeze and excitation mechanism. In order to improve performance, first, the feature maps are extracted and concatenated to obtain multiscale feature maps with ImageNet pretrained VGG network. After region of interest pooling, an encoding scale vector which has values between 0 and 1 is generated from subfeature maps. The scale vector is ranked, and only top K values will be preserved. Other values will be set to 0. Then, the subfeature maps are recalibrated by this scale vector. The redundant subfeature maps will be suppressed by this operation, and the detection performance of detector can be improved. The experimental results based on Sentinel-1 images show that the detection performance of the proposed method achieves 0.836 which is 9.7% better than the state-of-the-art method when using F1 as matric and executes 14% faster. Zhao Lin, Kefeng Ji, Xiangguang Leng, Gangyao Kuang |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2019 | Target recognition in SAR images via sparse representation in the frequency domain
Ganggang Dong, Hongwei Liu 0001, Gangyao Kuang, Jocelyn Chanussot |
Pattern Recognit. | 3 |
| 2019 | SAR Image Despeckling Based on Nonlocal Low-Rank RegularizationabstractIn this paper, we propose a new synthetic aperture radar (SAR) image despeckling method based on the nonlocal low-rank minimization model. First, some similar image patches are selected for each pixel to construct the patch group matrix (PGM). Then, a new low-rank minimization model, called Fisher-Tippett distribution (FT)-weighted nuclear norm minimization (WNNM), is proposed to recover the underlying low-rank component from the PGM. Specifically, the FT-WNNM is developed by reformulating the despeckling problem as the maximizing a posterior probability problem. The new model consists of a data fidelity term and a regularization term (also called prior term). The data fidelity term is derived from the statistical distribution of SAR images in the logarithm domain, which is known as the Fisher-Tippett distribution, and the regularization term is the recent weighted nuclear norm. Then, the alternating direction method of multipliers (ADMM) is introduced to solve the corresponding optimization problem. Under ADMM framework, the resulting subproblems can be solved efficiently and the convergence can be guaranteed. Extensive experiments on both simulated and real SAR images demonstrate that the proposed method can achieve comparable or even better despeckling performance than some state-of-the-art despeckling algorithms. Dongdong Guan, Deliang Xiang, Xiaoan Tang, Gangyao Kuang |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2018 | Locality-Constrained and Class-Specific Sparse Representation for Sar Target RecognitionabstractRecently, sparse representation has achieved the impressive performance on target recognition in synthetic aperture radar (SAR) image. However, the unstable and unsupervised optimization of the sparse representation may lead to undesired recognition result. In this paper, a locality-constrained and class-specific sparse representation (LCSR) framework is presented to alleviate these problems. Instead of the sparse constraint, the locality constraint is designed to utilize the local structure information of the training samples. It provides stable representation for the samples with minor variations, which is beneficial to classification. To further improve the recognition performance, the query sample is represented as a linear combination of class-specific galleries based on the supervision of class information. The inference is reached corresponding to the class with the minimum reconstruction error. The experimental results demonstrate the effectiveness and robustness of the proposed method. Meiting Yu, Lingjun Zhao, Siqian Zhang, Gangyao Kuang |
IGARSS | 4 |
| 2018 | SAR Image Classification by Exploiting Adaptive Contextual Information and Composite KernelsabstractFor synthetic aperture radar (SAR) image land cover classification, traditional feature-based methods are not always effective because of the heavy multiplicative noise. To solve this problem, we herein propose a new classification method for SAR images considering adaptive spatial contextual information. In contrast to preceding studies, the spatial contextual information of the SAR images is exploited via composite kernels (CKs). Additionally, an image superpixel strategy is employed to design an adaptive neighborhood, which enables the extraction of more accurate spatial information than a fixed-size neighborhood. Specifically, a modified superpixel map is first generated to produce the neighborhood. With this neighborhood, a context kernel is then defined by means of the Gaussian radial basis function. The resulting context kernel is combined with the conventional feature kernel via the designed CKs scheme. The relative proportion of these two kernels is controlled by a weight parameter. The label of each pixel is predicted by feeding the final CKs into a support vector machine classifier. Experiments on two real SAR images demonstrate that the proposed method can greatly improve the classification performance, both visually and quantitatively, in comparison to other traditional feature-based methods. Dongdong Guan, Deliang Xiang, Ganggang Dong, Tao Tang 0006, Xiaoan Tang, Gangyao Kuang |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2018 | Spatial-Spectral Total Variation Regularized Low-Rank Tensor Decomposition for Hyperspectral Image DenoisingabstractSeveral bandwise total variation (TV) regularized low-rank (LR)-based models have been proposed to remove mixed noise in hyperspectral images (HSIs). These methods convert high-dimensional HSI data into 2-D data based on LR matrix factorization. This strategy introduces the loss of useful multiway structure information. Moreover, these bandwise TV-based methods exploit the spatial information in a separate manner. To cope with these problems, we propose a spatial–spectral TV regularized LR tensor factorization (SSTV-LRTF) method to remove mixed noise in HSIs. From one aspect, the hyperspectral data are assumed to lie in an LR tensor, which can exploit the inherent tensorial structure of hyperspectral data. The LRTF-based method can effectively separate the LR clean image from sparse noise. From another aspect, HSIs are assumed to be piecewisely smooth in the spatial domain. The TV regularization is effective in preserving the spatial piecewise smoothness and removing Gaussian noise. These facts inspire the integration of the LRTF with TV regularization. To address the limitations of bandwise TV, we use the SSTV regularization to simultaneously consider local spatial structure and spectral correlation of neighboring bands. Both simulated and real data experiments demonstrate that the proposed SSTV-LRTF method achieves superior performance for HSI mixed-noise removal, as compared to the state-of-the-art TV regularized and LR-based methods. Haiyan Fan, Chang Li 0001, Yulan Guo, Gangyao Kuang, Jiayi Ma 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2018 | Derivation of the Orientation Parameters in Built-Up Areas: With Application to Model-Based DecompositionabstractThis paper concerns two polarization orientation parameters of built-up areas derived from the polarimetric synthetic aperture radar (PolSAR) data considering two modeling orthogonal dihedral structures. The first orientation parameter is derived from the well-known circular polarization algorithm with the enrichment of arc distance median filtering using an adaptive neighborhood. The derivation of the second orientation parameter is realized by combining the slope-induced changes in polarimetric orientation angle with the shape-from-shading technique. The combination provides the possibility to measure the incidence angle and the azimuth component of the terrain slopes from the cross-pol SAR intensity image. With reference to the cross scattering model, a doubled cross scattering model (DCSM) is introduced by incorporating the orientation parameters, thus serving to guide the model-based decomposition. Using the DCSM refines the estimation of the cross-pol component by enabling us to further reveal the scattering characteristics of built-up areas. Following a novel criterion, the decomposition is implemented at two layers: one for urban areas and one for nonurban areas. The performance of parameter derivation is demonstrated and evaluated with airborne synthetic aperture radar, uninhabited aerial vehicle synthetic aperture radar, and GF-3 fully PolSAR data over different test sites. The decomposed results are consistent with the reference information provided by the National Land Cover Database 2011 about the land cover classification of test sites and encourage the use of the proposed decomposition scheme for different applications. Sinong Quan, Boli Xiong, Deliang Xiang, Gangyao Kuang |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2018 | Matrix Completion for Downward-Looking 3-D SAR Imaging With a Random Sparse Linear ArrayabstractDownward-looking linear array 3-D synthetic aperture radar (SAR) has attracted increasing attention in the field of radar imaging. As widely reported, the volume of data can be significantly reduced by a random sparse linear array. However, the 2-D under-sampled azimuth-cross-track data brought by the sparse linear array will produce high-level side-lobes, as well as the aliasing and the false-alarm targets. To deal with those problems, this paper introduces a recently developed theory, matrix completion (MC). The new theory could recover a matrix with a small subset of known elements of the matrix. It is founded on the assumption that the matrix is essentially low rank. For downward-looking 3-D SAR with a random sparse linear array, the received 3-D data can be treated as a series of uncorrelated 2-D matrices by the separated channel process. First, range compression can be realized by means of pulse compression. Then, the sets of the 2-D under-sampled azimuth-cross-track matrix can be completed into a full-sampled one via MC trick. The resulting 3-D images can be focused by synthetic aperture technique along the azimuth direction and beamforming operation along the cross-track direction, with the recovered full-sampled matrix. The proposed algorithm achieves high resolution and low-level side-lobes with the acceptable computational cost and memory consumption. It is verified by several numerical simulations and multiple comparative studies on real data. The experimental results clearly demonstrate the imaging performance across different under-sampling rates and signal-noise rates. Siqian Zhang, Ganggang Dong, Gangyao Kuang |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2017 | Classification via Sparse Representation of Steerable Wavelet Frames on Grassmann Manifold: Application to Target Recognition in SAR ImageabstractAutomatic target recognition has been widely studied over the years, yet it is still an open problem. The main obstacle consists in extended operating conditions, e.g.., depression angle change, configuration variation, articulation, and occlusion. To deal with them, this paper proposes a new classification strategy. We develop a new representation model via the steerable wavelet frames. The proposed representation model is entirely viewed as an element on Grassmann manifolds. To achieve target classification, we embed Grassmann manifolds into an implicit reproducing Kernel Hilbert space (RKHS), where the kernel sparse learning can be applied. Specifically, the mappings of training sample in RKHS are concatenated to form an overcomplete dictionary. It is then used to encode the counterpart of query as a linear combination of its atoms. By designed Grassmann kernel function, it is capable to obtain the sparse representation, from which the inference can be reached. The novelty of this paper comes from: 1) the development of representation model by the set of directional components of Riesz transform; 2) the quantitative measure of similarity for proposed representation model by Grassmann metric; and 3) the generation of global kernel function by Grassmann kernel. Extensive comparative studies are performed to demonstrate the advantage of proposed strategy. Ganggang Dong, Gangyao Kuang, Na Wang 0002, Wei Wang 0099 |
IEEE Trans. Image Process. | 2 |
| 2016 | Registration for SAR and optical images based on straight line features and mutual informationabstractThis paper proposes a novel registration method for optical and SAR images which is based on straight line features and mutual information. Firstly, different edge detectors are employed to detect the line segments in both optical and SAR images respectively. Then, through the Hough transform and a straight line fitting and filtration method, the main straight lines of each image are extracted and their intersections are obtained and taken as the candidate matching points. With the RANSAC (RANdom SAmpling Consensus) method, corresponding point pairs (CPPs) are found with these candidate points and a coarse registration between the heterogeneous images is implemented. At last, by using the mutual information of the separated patches generated from the coarse registered images, a fine registration result is finally achieved. The experiment with a pair of X-band air-borne SAR and optical images validates the efficiency and precision of the proposed method. Boli Xiong, Wenchao Li 0002, Lingjun Zhao, Jun Lu 0008, Xiaoqiang Zhang 0005, Gangyao Kuang |
IGARSS | 6 |
| 2016 | A geometric parameter extraction method of ship target based on an improved snake modelabstractIt is the basis of realization ship recognition to accurately extract geometric parameters of the ship target in synthetic aperture radar (SAR) images. Due to the unique SAR imaging mechanism, speckle noise, azimuth ambiguity and side lobe effect seriously impact on geometric parameter estimation of the ship target. Therefore, a method is presented to extract geometric parameters of the ship based on ellipse fitting and the gradient vector flow (GVF) snake with shape priors. Xiaoqiang Zhang 0005, Boli Xiong, Gangyao Kuang |
IGARSS | 3 |
| 2016 | Affine invariant shape projection distribution for shape matching using relaxation labellingabstractShape is considered to be one of the most promising tools to represent and recognise an object. In this study, an effective and rigorous shape matching algorithm is developed based on a new descriptor and relaxation labelling technique. For each contour point, the descriptor captures the distribution of all points within the shape region along the vector perpendicular to that from the centroid to the point. In addition to stable affine invariance, the descriptor is robust to noise since it makes use of all points in the shape region. The descriptor distance is used to initialise the contour point matching probability, and relaxation labelling technique is utilised to update the matching probability using a new compatibility coefficient function, which is defined based on the shape projection preserving characteristic. The experiments on synthetic and real remote sensing data are provided to test the performance of the authors’ proposed algorithm. Compared to other four state‐of‐the‐art contour‐based shape matching algorithms, their algorithm is more robust and capable of shape matching under affine transformations and noise. Wei Wang 0099, Boli Xiong, Xingwei Yan, Yongmei Jiang, Gangyao Kuang |
IET Comput. Vis. | 5 |
| 2016 | A Soft Decision Rule for Sparse Signal Modeling via Dempster-Shafer Evidential ReasoningabstractRecently, the problem of recovering sparse linear representation of a query in terms of a redundant dictionary has received great interest. The query sample is represented as a linear combination of the atoms of the dictionary. The unique representation is obtained via sparsity constraint, and the decision is made in terms of the characteristics of representation on reconstruction. The decision rule for sparse signal modeling can be viewed as a typical application of Bayesian estimation, where the likelihood function is inversely proportional to the reconstruction error. Different from the conventional rule, where the decision is directly made according to the overall reconstruction error associated with each class, this letter proposes a soft decision via Dempster-Shafer theory of evidence. To model the imprecision on uncertainty measurement, we introduce the samplewise ambiguity and the classwise ambiguity during the quantification of probability mass. Each sample that participates in the recovery of the query is considered as an item of evidence that supports certain hypothesis in regard to the class membership of query. The amount of evidence is quantified by a function of the distance between the query and the weighted training sample, where the weights result from the sparse representation coefficient. Then, various pieces of evidence derived from the candidate samples are pooled by means of Dempster's rule of combination, from which a soft decision can be reached. Ganggang Dong, Gangyao Kuang |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2016 | Adaptive and Fast Prescreening for SAR ATR via Change Detection TechniqueabstractChange detection is a process of identifying changes in the state of objects between the reference and test images. This letter presents a target prescreening method that employs the change detection technique for automatic target recognition in synthetic aperture radar (SAR) images. First, four translated versions of an original SAR image are generated, and the corresponding four likelihood ratio images are computed. Then, a robust threshold is derived from the ratio of the histogram at two adjacent gray-level values of the likelihood ratio images. Finally, the threshold is applied to perform the prescreening. The proposed method implements the procedure without any prior knowledge and overcomes the weak adaptability of traditional algorithms. Two different real X-band airborne SAR images acquired over Beijing are used to quantitatively and qualitatively demonstrate the effectiveness of the proposed method. Sinong Quan, Boli Xiong, Siqian Zhang, Meiting Yu, Gangyao Kuang |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2015 | Analytic estimation performance bounds of downward-looking linear array 3-D SAR imaging based on compressive sensingabstractFor downward-looking linear array three-dimensional SAR, the resolution in cross-track direction is a curial problem. Hence, compressive sensing algorithm has been used to acquire the superresolution performance in cross-track direction. The limits of the proposed algorithm are investigated in this paper. How accurately can the scattering intensity of the scatterers be estimated? What is the closest separable distance of two scatterers at different levels of SNR? What is the influence of the acquisitions N on the resolution? For all of these questions, the theoretical analysis is given by Cramér-Rao Bound and numerical simulations are proven. The results can be considered as a fundamental bound on parameter estimates. Siqian Zhang, Yutao Zhu 0005, Gangyao Kuang, Lingjun Zhao |
IGARSS | 3 |
| 2015 | Target Recognition in SAR Images via Classification on Riemannian ManifoldsabstractIn this letter, synthetic aperture radar (SAR) target recognition via classification on Riemannian geometry is presented. To characterize SAR images, which have broad spectral information yet spatial localization, a 2-D analytic signal, i.e., the monogenic signal, is used. Then, the monogenic components are combined by computing a covariance matrix whose entries are the correlation of the components. Since the covariance matrix, a symmetric positive definite one, lies on the Riemannian manifold, it is unreasonable to be dealt with by the standard learning techniques. To address the problem, two classification schemes are proposed. The first maps the covariance matrix into the vector space and feeds the resulting descriptor into a recently developed framework, i.e., sparse representation-based classification. The other embeds the Riemannian manifold into an implicit reproducing kernel Hilbert space, followed by least square fitting technique to recover the test. The inference is reached by evaluating which class of samples could reconstruct the test as accurately as possible. Ganggang Dong, Gangyao Kuang |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2015 | Target Recognition via Information Aggregation Through Dempster-Shafer's Evidence TheoryabstractIn this letter, a novel classification via information aggregation through Dempster-Shafer's (DS) evidence theory has been presented to target recognition in a SAR image. Although the DS theory of evidence has been widely studied over the decades, less attention has been paid to its application for target recognition. To capture the characteristics of a SAR image, this letter exploits a new multidimensional analytic signal named monogenic signal. Since the components of the monogenic signal are of a high dimension, it is unrealistic to be directly used. To solve the problem, an intuitive idea is to derive a single feature by these components. However, this strategy usually results in some information loss. To boost the performance, this letter presents a classification framework via information aggregation. The monogenic components are individually fed into a recently developed algorithm, i.e., sparse representation-based classification, from which the residual with respect to each target class can be produced. Since the residual from a query sample reflects the distance to the manifold formed by the training samples of a certain class, it is reasonable to be used to define the probability mass. Then, the information provided by the monogenic signal can be aggregated via Dempster's rule; hence, the inference can be reached. Ganggang Dong, Gangyao Kuang |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2015 | Truncated SVD-Based Compressive Sensing for Downward-Looking Three-Dimensional SAR Imaging With Uniform/Nonuniform Linear ArrayabstractFor downward-looking linear array 3-D synthetic aperture radar, the resolution in cross-track direction is much lower than the ones in range and azimuth. Hence, superresolution reconstruction algorithms are desired. Since the cross-track signal to be reconstructed is sparse in the object domain, compressive sensing algorithm has been used. However, the imaging processing on the 3-D scene brings large computational loads, which renders challenges in both data acquisition and processing. To cover this shortage, truncated singular value decomposition is utilized to reconstruct a reduced-redundancy spatial measurement matrix. The proposed algorithm provides advantages in terms of computational time while maintaining the quality of the scene reconstructions. Moreover, our results on uniform linear array are generally applicable to sparse nonuniform linear array. Superresolution properties and reconstruction accuracies are demonstrated using simulations under the noise and clutter scenarios. Siqian Zhang, Yutao Zhu 0005, Ganggang Dong, Gangyao Kuang |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2015 | Imaging of Downward-Looking Linear Array Three-Dimensional SAR Based on FFT-MUSICabstractIn this letter, a novel imaging algorithm of downward-looking linear array 3-D synthetic aperture radar (SAR) is presented. To improve the resolution in cross-track direction, a multiple-signal-classification algorithm has been used. However, the computational cost is unattractive. To cover the shortage, fast Fourier transform (FFT) is utilized to roughly focus the targets in cross-track direction in this letter. Then, the range of peak searching can be considerably reduced. In addition, the backscattering coefficient can be directly obtained by FFT. On the other hand, since the scattering centers are always correlated in a real SAR system, the estimated covariance matrix is singular. To address the problem, the nearby spatial smoothing method is proposed. The array vectors of the nearby range and azimuth units are used to reconstruct the estimated correlation matrix. The effective aperture of the array can also be kept. Siqian Zhang, Yutao Zhu 0005, Gangyao Kuang |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2015 | Fusing Sorted Random Projections for Robust Texture and Material ClassificationabstractThis paper presents a conceptually simple, and robust, yet highly effective, approach to both texture classification and material categorization. The proposed system is composed of three components: 1) local, highly discriminative, and robust features based on sorted random projections (RPs), built on the universal and information-preserving properties of RPs; 2) an effective bag-of-words global model; and 3) a novel approach for combining multiple features in a support vector machine classifier. The proposed approach encompasses the simplicity, broad applicability, and efficiency of the three methods. We have tested the proposed approach on eight popular texture databases, including Flickr Materials Database, a highly challenging materials database. We compare our method with 13 recent state-of-the-art methods, and the experimental results show that our texture classification system yields the best classification rates of which we are aware of 99.37% for Columbia-Utrecht, 97.16% for Brodatz, 99.30% for University of Maryland Database, and 99.29% for Kungliga Tekniska högskolan-textures under varying illumination, pose, and scale. Moreover, the proposed approach significantly outperforms the current state-of-the-art approach in materials categorization, with an improvement to classification accuracy of 67%. Li Liu 0002, Paul W. Fieguth, Dewen Hu, Yingmei Wei, Gangyao Kuang |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2015 | Classification on the Monogenic Scale Space: Application to Target Recognition in SAR ImageabstractThis paper introduces a novel classification strategy based on the monogenic scale space for target recognition in Synthetic Aperture Radar (SAR) image. The proposed method exploits monogenic signal theory, a multidimensional generalization of the analytic signal, to capture the characteristics of SAR image, e.g., broad spectral information and simultaneous spatial localization. The components derived from the monogenic signal at different scales are then applied into a recently developed framework, sparse representation-based classification (SRC). Moreover, to deal with the data set, whose target classes are not linearly separable, the classification via kernel combination is proposed, where the multiple components of the monogenic signal are jointly considered into a unifying framework for target recognition. The novelty of this paper comes from: the development of monogenic feature via uniformly downsampling, normalization, and concatenation of the components at various scales; the development of score-level fusion for SRCs; and the development of composite kernel learning for classification. In particular, the comparative experimental studies under nonliteral operating conditions, e.g., structural modifications, random noise corruption, and variations in depression angle, are performed. The comparative experimental studies of various algorithms, including the linear support vector machine and the kernel version, the SRC and the variants, kernel SRC, kernel linear representation, and sparse representation of monogenic signal, are performed too. The feasibility of the proposed method has been successfully verified using Moving and Stationary Target Acquiration and Recognition database. The experimental results demonstrate that significant improvement for recognition accuracy can be achieved by the proposed method in comparison with the baseline algorithms. Ganggang Dong, Gangyao Kuang |
IEEE Trans. Image Process. | 2 |
| 2014 | Joint sparse representation of monogenic components: With application to automatic target recognition in SAR imageryabstractIn this paper, classification via joint sparse representation of the monogenic signal is presented for target recognition in SAR imagery. First, the monogenic signal is performed to capture the characteristics of SAR image. Since it is infeasible to directly apply the raw component to classification due to the high data dimension and redundancy, three augmented feature vectors are defined via uniform downampling of the real part, the imagery part, and the instantaneous phase. The monogenic features are then fed into a recently developed framework, sparse representation-based classification (SRC). Rather than produce individual sparse pattern, this paper generates the similar sparsity pattern for three feature vectors by imposing a mixed norm on the representation matrix. Extensive experiments on MSTAR database demonstrate that the proposed method could significantly improve the recognition accuracy. Ganggang Dong, Gangyao Kuang, Lingjun Zhao, Jun Lu 0008, Min Lu 0001 |
IGARSS | 2 |
| 2014 | Nonnegative and local linear regression for classification in SAR imageryabstractIn this paper, the classification via nonnegative and local linear regression model is proposed for SAR image-based target recognition. Recently, a simple yet effective method, linear regression for pattern recognition has been presented. By assuming that images from a single-object class lie on a linear subspace, it represents the test image as a linear combination of class-specific galleries. The representation is obtained by solving a typical inverse problem with least-square strategy. Since the negative weights play a counteractive role in reconstruction, it may be unreasonable to generate the negative weights. In addition, those elements close to the test sample should contribute much more than the ones far from the test. Thus this paper limits the feasible set of the representation by nonnegative and locality constraint. The decision is ruled in favor of the class with the minimum reconstruction error. Extensive experiments on MSTAR database demonstrate that the proposed methods significantly improve the accuracy than the standard one. Ganggang Dong, Gangyao Kuang, Lingjun Zhao, Jun Lu 0008, Min Lu 0001 |
IGARSS | 2 |
| 2014 | SAR Azimuth ambiguities removal for ship detection using time-frequency techniquesabstractIn this paper, a new azimuth ambiguities removal method is introduced for ship detection by Time-Frequency (TF) analysis. A TF coherence indicator is proposed to filter ghost echoes due to the different TF coherence characteristics between real ship target echoes and ambiguous ones. The effectiveness of this proposed TF coherence indicator for ship detection is demonstrated using single polarimetric spaceborne TerraSAR-X coherent data over the test sea/ocean site in Hongkong, China. Canbin Hu, Boli Xiong, Jun Lu 0008, Zhiyong Li 0008, Lingjun Zhao, Gangyao Kuang |
IGARSS | 6 |
| 2014 | Detecting floods accurately in SAR images without the interference of the reference imageabstractMost of methods for change detection of floods in SAR images are based on the difference image (DI), which is from comparing the flood image and the reference image straightforwardly. DI, however, is often distorted by the interferences of the complicated reference image. Meanwhile, the spatial contextual information of floods is difficult to be effectively used. In response to these problems, a novel change detection method by removing the interference of the reference image is proposed. Firstly, the difference image is segmented to derive the initial change mask, i.e. a part of flood areas. Secondly, the whole flood areas are gained by region growing directly in the flood image. Experimental results on the real ERS SAR dataset validate its effectiveness on both the quantitative and subjective aspects. Jun Lu 0008, Jonathan Li 0001, Min Lu 0001, Zhiyong Li 0008, Boli Xiong, Gangyao Kuang |
IGARSS | 6 |
| 2014 | Parameter estimation for SAR moving target detection using Fractional Fourier TransformabstractThis paper proposes an algorithm for multi-channel SAR ground moving target detection and estimation using the Fractional Fourier Transform(FrFT). To detect the moving target with low speed, the clutter is first suppressed by Displace Phase Center Antenna(DPCA), then the signal-to-clutter can be enhanced. Have suppressed the clutter, the echo of moving target remains and can be regarded as a chirp signal whose parameters can be estimated by FrFT. FrFT, one of the most widely used tools to time-frequency analysis, is utilized to estimate the Doppler parameters, from which the moving parameters, including the velocity and the acceleration can be obtained. The effectiveness of the proposed method is validated by the simulation. Yongmei Jiang, Gangyao Kuang, Jun Lu 0008, Zhiyong Li 0008 |
IGARSS | 3 |
| 2014 | Contour matching using the affine-invariant support point setabstractMoment has been widely used for contour matching. To use the moment to achieve contour matching under affine transformations, the affine‐invariant support point set (SPS) should be constructed first. Then, a novel method of acquiring SPS based on the contour projection (SPS‐CP) is proposed here. For an arbitrary selected contour point, the contour is projected onto the line vertical to the vector connecting the contour centroid and the selected point, and the contour points with the sampled projection values are picked up to form the SPS‐CP of the point. SPS‐CP which captures the global structure of the contour is stably affine‐invariant. Experiments on synthetic and real data demonstrate that moments generated from SPS‐CP outperform those generated from SPSs sampled by uniform spacing or affine length. Wei Wang 0099, Yongmei Jiang, Boli Xiong, Lingjun Zhao, Gangyao Kuang |
IET Comput. Vis. | 5 |
| 2014 | Sparse Representation of Monogenic Signal: With Application to Target Recognition in SAR ImagesabstractIn this letter, the classification via sparse representation of the monogenic signal is presented for target recognition in SAR images. To characterize SAR images, which have broad spectral information yet spatial localization, the monogenic signal is performed. Then an augmented monogenic feature vector is generated via uniform down-sampling, normalization and concatenation of the monogenic components. The resulting feature vector is fed into a recently developed framework, i.e., sparse representation based classification (SRC). Specifically, the feature vectors of the training samples are utilized as the basis vectors to code the feature vector of the test sample as a sparse linear combination of them. The representation is obtained via l1-norm minimization, and the inference is reached according to the characteristics of the representation on reconstruction. Extensive experiments on MSTAR database demonstrate that the proposed method is robust towards noise corruption, as well as configuration and depression variations. Ganggang Dong, Na Wang 0002, Gangyao Kuang |
IEEE Signal Process. Lett. | 3 |
| 2013 | Multi-dimensional coherent Time-Frequency analysis for ship detection in polsar imageryabstractThis paper proposes an algorithm for ship detection in complex scenes using multi-dimensional coherent Time-Frequency (TF) techniques and dual-polarization SAR data. The PolSAR multi-dimensional information is analysed by means of a linear TF decomposition approach which permits to describe the ship and background area polarimetric behaviour for different azimuth angles of observation and frequencies of illumination. A statistic descriptor related to the signal polarimetric coherence in the TF domain, is used for ship detection in different backgrounds, including the environments with presence of strong ghost ambiguities or small natural islands. Using polarimetric RADARSAT-2 data, experimental results demonstrate the efficiency of this method. Canbin Hu, Laurent Ferro-Famil, Gangyao Kuang |
IGARSS | 3 |
| 2013 | A new method of micro-motion parameters estimation based on cyclic autocorrelation function
Jie Niu, Kangle Li, Weidong Jiang, Xiang Li 0014, Gangyao Kuang |
Sci. China Inf. Sci. | 5 |
| 2013 | Unsupervised Land Cover/Land Use Classification Using PolSAR Imagery Based on Scattering SimilarityabstractThis paper presents a new unsupervised land cover/land use classification scheme using polarimetric synthetic aperture radar (PolSAR) imagery based on polarimetric scattering similarity. Compared with theH/alpha classification scheme based on a dominant “average” scattering mechanism, the proposed scheme has such advantages as the following: (1) The major scattering mechanism represents a target scattering in the low-entropy case; (2) it also represents both the major and minor scattering mechanisms in the medium-entropy case; and (3) all the scattering mechanisms in the high-entropy case can be represented. The major and minor scattering mechanisms have been identified automatically based on the relative magnitude of multiple-scattering similarities. The canonical scattering corresponding to maximum scattering similarity is regarded as the major scattering mechanism. The result obtained using the National Aeronautics and Space Administration/Jet Propulsion Laboratory's AIRSAR L-band PolSAR imagery reveals that the proposed scheme is more effective as compared to the existing models and promises to increase the accuracy of the classification and interpretation. Qiang Chen 0015, Gangyao Kuang, Jonathan Li 0001, Lichun Sui, Diangang Li |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2012 | Polarimetric SAR target detection based on polarization synthesisabstractThis paper addresses the Polarimetric synthetic aperture radar (PolSAR) CFAR target detection utilizing the PolSAR synthesis technique. The optimal polarization ellipticity and orientation angles, which maximizes the ratio of the antenna receiver power between target and clutter (SCR), are searched in the co-polarized and cross-polarized channels, and the obtained antenna receiver power is named as the polarization synthesis enhancement (PSE) metric. Using the PSE metric, a data fitting based target detection scheme is presented. First, the Fisher distribution, which has extensive modeling capacity in a large set of clutters, is utilized to fit the distribution of PSE metric. Then, the corresponding “Second Kind Statistics” (SKS) parameter estimator is presented and the numerical solution of the detection threshold is also derived. Afterward, the CFAR detection is implemented. The experimental results demonstrate the enhancement results of PSE metric are almost comparable with that of the Polarimetric Matched Filter (PMF) detector. However, the computation load of PSE metric is generally less than that of the PMF. Moreover, the target detection results show the CFAR detector based on the Fisher distribution can realize the accurate target detection in homogenous and heterogeneous clutter areas. Na Wang 0002, Canbin Hu, Lingjun Zhao, Yongmei Jiang, Gangyao Kuang |
IGARSS | 5 |
| 2012 | A method of acquiring tie points based on closed regions in SAR imagesabstractThis paper presents a method in finding tie points automatically in synthetic aperture radar (SAR) image pairs based on the extracted closed regions. There are mainly three steps during this process. The first step is to extract the closed regions in the SAR image with an image segmentation approach deducted by the geodesic active contour (GAC) model. Then, a polygonal approximation process is adopted to locate the feature points on the boundaries of these regions. With the obtained feature points, geometric hashing theory is employed to match these feature points as the tie points. A pair of simulated SAR images and a pair of high-resolution airborne SAR images are used to test and evaluate the proposed method. The experimental results show that the proposed method is effective and appropriate for the acquisitions of tie points in SAR image pairs. Boli Xiong, Zhiguo He, Canbin Hu, Yongmei Jiang, Gangyao Kuang |
IGARSS | 6 |
| 2012 | Extended local binary patterns for texture classification
Li Liu 0002, Lingjun Zhao, Yunli Long, Gangyao Kuang, Paul W. Fieguth |
Image Vis. Comput. | 4 |
| 2012 | Polarimetric SAR Target Detection Using the Reflection SymmetryabstractThis letter addresses the polarimetric synthetic aperture radar target detection using the magnitude of the (2, 3) term in the sample averaged coherency matrix. The theoretical analysis demonstrates that such term reveals the difference between the nonreflection symmetric targets and natural clutters. The statistical models for such term are derived within different degrees of homogeneity. Based on the statistical models, an automatic constant-false-alarm-rate detection scheme is completed. The parameter estimation and the solution for the detection threshold are given in detail. Experimental results demonstrate the capability of the proposed approach for detecting ships, oil stores, buildings, etc., in homogeneous and heterogeneous areas. Na Wang 0002, Gongtao Shi, Li Liu 0002, Lingjun Zhao, Gangyao Kuang |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2012 | A Threshold Selection Method Using Two SAR Change Detection Measures Based on the Markov Random Field ModelabstractThis letter presents a threshold selection method in change detection (CD) with synthetic aperture radar (SAR) images, which combines the characteristics of two different CD measures by using the Markov random field model. One is the well-known log-ratio CD measure, and the other is derived from the likelihood ratio and is based on the statistical properties of SAR intensity images. The proposed unsupervised CD algorithm overcomes the shortcomings and strengthens the advantages of these two measures. The experimental results with two pairs of SAR images show that the proposed algorithm is effective and better than the algorithms using the two aforementioned CD measures. Boli Xiong, Yongmei Jiang, Gangyao Kuang |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2012 | Sorted random projections for robust rotation-invariant texture classification
Li Liu 0002, Paul W. Fieguth, David A. Clausi, Gangyao Kuang |
Pattern Recognit. | 4 |
| 2012 | Estimation of the Repeat-Pass ALOS PALSAR Interferometric Baseline Through Direct Least-Square Ellipse FittingabstractThe precise estimation of the baseline is a crucial procedure in repeat-pass interferometric synthetic aperture radar (InSAR) applications. Using the ephemeris of the satellite, a polynomial regression algorithm can fit the satellite orbit at the third or higher order with a main shortcoming that the mutual constraints among the three dimensions defining the orbit are missed. In this paper, a new approach is presented to fit the satellite orbit based on the assumption that the satellite orbit is a 3-D ellipse, which retains the relations among the three dimensions. Considering the complexity of 3-D ellipse parameters estimation, the 3-D orbit is first transformed into three 2-D ellipses. Then, the parameters of these 2-D ellipses are estimated with a direct least-square ellipse fitting method (DLS-EFM). These two orbit fitting algorithms are tested with ten sets of advanced land observation satellite phased array L-band SAR data, which were acquired in north Toronto, Ontario, Canada, from September, 2008 to January, 2009. Moreover, two of them acquired with an adjacent period were chosen to form a repeat-pass InSAR, and the corresponding baseline is calculated with the proposed method as an example. The experimental results show that the error of the satellite position using DLS-EFM is at a submetric level, which is less than one-tenth of that of the polynomial regression algorithm. Consequently, the proposed method is appropriate for the baseline estimation in spaceborne InSAR applications. Boli Xiong, Jing M. Chen, Gangyao Kuang, Nobuhiko Kadowaki |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2011 | Generalized Local Binary Patterns for Texture ClassificationabstractThis paper presents a novel approach for texture classification, generalizing the wellknown local binary patterns (LBP). In the proposed approach, two different and complementary types of features are extracted from local patches, based on pixel intensities and differences. Inspired by the LBP approach, two intensity-based and two difference-based descriptors are developed. All four descriptors have the same form as the conventional LBP codes, thus they can be readily combined to form joint histograms to represent textured images. The proposed approach is computationally simple and is training-free: there is no need to learn a texton dictionary and no tuning of parameters. Extensive experimental results on two challenging texture databases (Outex and KTHTIPS2b) show that the proposed approach significantly outperforms the classical LBP approach and other state-of-the-art methods with a nearest neighbor classifier. Paul W. Fieguth, Gangyao Kuang |
BMVC | 3 |
| 2011 | Sorted Random Projections for robust texture classificationabstractThis paper presents a simple and highly effective system for robust texture classification, based on (1) random local features, (2) a simple global Bag-of-Words (BoW) representation, and (3) Support Vector Machines (SVMs) based classification. The key contribution in this work is to apply a sorting strategy to a universal yet information-preserving random projection (RP) technique, then comparing two different texture image representations (histograms and signatures) with various kernels in the SVMs. We have tested our texture classification system on six popular and challenging texture databases for exemplar based texture classification, comparing with 12 recent state-of-the-art methods. Experimental results show that our texture classification system yields the best classification rates of which we are aware of 99.37% for CUReT, 97.16% for Brodatz, 99.30% for UMD and 99.29% for KTH-TIPS. Moreover, combining random features significantly outperforms the state-of-the-art descriptors in material categorization. Li Liu 0002, Paul W. Fieguth, Gangyao Kuang, Hongbin Zha |
ICCV | 3 |
| 2011 | Combining sorted random features for texture classificationabstractThis paper explores the combining of powerful local texture descriptors and the advantages over single descriptors for texture classification. The proposed system is composed of three components: (i) highly discriminative and robust sorted random projections (SRP) features; (ii) a global Bag-of-Words (BoW) model; and (iii) the use of multiple kernel Support Vector Machines (SVMs) combining multiple features. The proposed system is also very simple, stemming from (1) the effortless extraction of the SRP features, (2) the simple orderless histogramming in the BoW model, (3) a strategy with low computational complexity for multiple kernel SVMs. We have tested our texture classification system on three popular and challenging texture databases and find that the SVMs combining of SRP features produces outstanding classification results, out-performing the state-of-the-art for CUReT (99.37%) and KTH-TIPS (99.29%), and with highly competitive results for UIUC (98.56%). Li Liu 0002, Paul W. Fieguth, Gangyao Kuang |
ICIP | 3 |
| 2010 | Compressed Sensing for Robust Texture Classification
Li Liu 0002, Paul W. Fieguth, Gangyao Kuang |
ACCV (1) | 3 |
| 2010 | Polarimetric Scattering Similarity Between a Random Scatterer and a Canonical ScattererabstractIn this letter, we propose a novel parameter to measure the scattering similarity between a random scatterer and a canonical scatterer. Compared with the similarity parameter proposed by Yang, the novel parameter not only has some advantages, such as its independence of the spans of a coherence matrix, but also can be applied directly in the case of a random scatterer made up of multiscattering centers. As an example, the novel parameter is adopted to extract some scattering characteristics of a target. With the full polarimetric L-band airborne synthetic aperture radar data, we illustrate the veracity of the novel parameter in measuring scattering similarity and its application in terrain classification. Yongmei Jiang, Lingjun Zhao, Gangyao Kuang |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2009 | An Optimization Procedure of the Lagrange Multiplier Method for Polarimetric Power OptimizationabstractThe Lagrange multiplier method is one of the basic optimization procedures to find the optimum polarizations for the incoherent scattering case. This letter proves for the first time that a fixed relationship exists between the optimum polarization and the Lagrange multiplier. Then, an optimization procedure is proposed to simplify the computational complexity of the Lagrange multiplier method. To speed up the convergence of the proposed procedure, the minimum search intervals are discussed and given theoretically. A numerical example is shown to demonstrate the effectiveness of the proposed procedure. Qiang Chen 0015, Yongmei Jiang, Lingjun Zhao, Gui Gao, Gangyao Kuang |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2009 | An Adaptive and Fast CFAR Algorithm Based on Automatic Censoring for Target Detection in High-Resolution SAR ImagesabstractAn adaptive and fast constant false alarm rate (CFAR) algorithm based on automatic censoring (AC) is proposed for target detection in high-resolution synthetic aperture radar (SAR) images. First, an adaptive global threshold is selected to obtain an index matrix which labels whether each pixel of the image is a potential target pixel or not. Second, by using the index matrix, the clutter environment can be determined adaptively to prescreen the clutter pixels in the sliding window used for detecting. The$G^{0}$distribution, which can model multilook SAR images within an extensive range of degree of homogeneity, is adopted as the statistical model of clutter in this paper. With the introduction of AC, the proposed algorithm gains good CFAR detection performance for homogeneous regions, clutter edge, and multitarget situations. Meanwhile, the corresponding fast algorithm greatly reduces the computational load. Finally, target clustering is implemented to obtain more accurate target regions. According to the theoretical performance analysis and the experiment results of typical real SAR images, the proposed algorithm is shown to be of good performance and strong practicability. Gui Gao, Li Liu 0002, Lingjun Zhao, Gongtao Shi, Gangyao Kuang |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2007 | Fast detecting and locating groups of targets in high-resolution SAR images
Gui Gao, Gangyao Kuang, DeRen Li |
Pattern Recognit. | 2 |
| 2005 | Simulation of SAR image of ship
Kefeng Ji, Gangyao Kuang, Wenxian Yu |
IGARSS | 2 |
| 2005 | Road extraction from high-resolution SAR imagery using Hough transformabstractThis paper presents a technique for the extraction of roads in a high resolution synthetic aperture radar (SAR) image using Hough transform. Roads in a high resolution SAR image can be modeled as a homogeneous dark area bounded by two parallel boundaries. Dark areas, which represent the candidate positions for roads, are extracted from the image using a Gaussian probability iteration segmentation, and the roads are accurately detected by Hough transform. For this purpose, we designed an average Hough transform, which is more reasonable than general Hough transform for the extraction of lines. We search the peak values in Hough space and try to reduce its overall computational cost by introducing a global CFAR detector. In this process, to detect roads more accurately, post-processing, including noisy dark regions removal and false roads removal, is performed. We applied our method to MSTAR clutter images of Redstone that have a resolution of about 1 ft /spl times/ 1 ft. The experimental results show that our method can accurately detect roads. Chengli Jia, Kefeng Ji, Yongmei Jiang, Gangyao Kuang |
IGARSS | 4 |