Tao Zhang 0027

dblp:15/4777-27 · DBLP profile ↗
← Back
82ranked-venue papers
20as first author
64since 2021 · last 2026
0000-0002-7192-5153ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 54 · 19 first-author · 43 since 2021Graphics, computer vision, multimedia, augmented reality and games · 20 · 14 since 2021Artificial intelligence and machine learning · 8 · 1 first-author · 6 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Point Cloud Analysis Under Slight Perturbations: A Manifold Distillation Approach Using Raw Coordinates
abstract
Point cloud is often regarded as a discrete sampling of Riemannian manifold and plays a pivotal role in the 3D image interpretation. Particularly, rotation perturbation, an unexpected small change in rotation caused by various factors (like equipment offset, system instability, measurement errors and so on), can easily lead to the inferior results in point cloud learning tasks. However, classical point cloud learning methods are sensitive to rotation perturbation, and the existing networks with rotation robustness also have much room for improvements in terms of performance and noise tolerance. Given these, this paper remodels the point cloud from the perspective of manifold as well as designs a manifold distillation method to achieve the robustness of rotation perturbation without any coordinate transformation. In brief, during the training phase, we introduce a teacher network to learn the rotation robustness information and transfer this information to the student network through online distillation. In the inference phase, the student network directly utilizes the raw 3D coordinate information to achieve the robustness of rotation perturbation. Experiments carried out on four different datasets verify the effectiveness of our method. On average, on the ModelNet40 and ScanObjectNN classification datasets with random rotation perturbations, our method improves classification accuracy by 4.41% and 3.65%, respectively, compared to popular rotation-robust networks. Similarly, on the ShapeNet and S3DIS segmentation datasets, our method achieves improvements in mIoU of 6.96% and 5.12%, respectively. Furthermore, the experimental results also demonstrate that our algorithm exhibits higher computational efficiency and stronger resistance to noise and outliers.
Tao Zhang 0027, Huazhen Liu, Feiming Wei, Huilin Xiong, Wenxian Yu
IEEE Trans. Circuits Syst. Video Technol.2
2025 MMCD: Memory-Based Multimodal Change Detection
abstract
Single-modal change detection methods based on optical or Synthetic Aperture Radar (SAR) images face challenges such as degradation due to adverse weather or noise interference. In contrast, multimodal change detection struggles with significant domain gaps between different modalities. Inspired by the SAM2 model’s temporal memory mechanism for video segmentation, this paper introduces the concept of memory into change detection and proposes a novel approach called Memory-based Multimodal Change Detection (MMCD). By treating change detection as a temporal problem and modeling remote sensing images as video sequences, the proposed method integrates historical optical images with current SAR images to enhance detection accuracy. Additionally, a difference map enhancement module is introduced to mitigate false changes caused by modality discrepancies. Experimental results show that this approach achieves state-of-the-art performance in multimodal change detection, demonstrating the effectiveness of the proposed method.
Limeng Zhang, Zenghui Zhang, Juanping Wu, Weiwei Guo, Tao Zhang 0027, Wenxian Yu
ICASSP5
2025 Scattering Enhancement and Feature Fusion Network for Aircraft Detection in SAR Images
abstract
Aircraft detection in synthetic aperture radar (SAR) images is one challenging task due to the discreteness of aircraft scattering, the diversity of aircraft size, and the interference of background. In order to deal with these problems, a novel method named scattering enhancement and feature fusion network (SEFFNet) is here proposed to detect aircraft via combining traditional image processing and deep learning together. At first, a scattering information extraction and enhancement module (SIEEM) is proposed to highlight the scattering points of aircraft targets. Then, to more effectively focus on the location of aircraft targets, a space-to-depth coordinate attention module (SDCAM) is further designed, following which an efficient multi-scale feature fusion pyramid (FFP) is also introduced to fuse the semantic information of different layers. At last, a contextual fusion head (CFH) is built to improve the receptive field for better detecting aircraft. The experiments carried out on the popular datasets SADD and SAR-AIRcraft-1.0 show that SEFFNet is more appropriate for aircraft detection, especially the small-size aircraft detection, in comparison with other state-of-the-art (SOTA) methods. Taking the dataset SADD for example, on average, the precision, recall, F1-score, and APs values are respectively 2.8%, 2.6%, 2.7%, and 2.0% higher than the baseline network YOLOv5.
Bocheng Huang, Tao Zhang 0027, Sinong Quan, Wei Wang 0099, Weiwei Guo, Zenghui Zhang
IEEE Trans. Circuits Syst. Video Technol.2
2025 VLF-SAR: A Novel Vision-Language Framework for Few-Shot SAR Target Recognition
abstract
Due to the challenges of obtaining data from valuable targets, few-shot learning plays a critical role in synthetic aperture radar (SAR) target recognition. However, the high noise levels and complex backgrounds inherent in SAR data make this technology difficult to implement. To improve the recognition accuracy, in this paper, we propose a novel vision-language framework, VLF-SAR, with two specialized models: VLF-SAR-P for polarimetric SAR (PolSAR) data and VLF-SAR-T for traditional SAR data. Both models start with a frequency embedded module (FEM) to generate key structural features. For VLF-SAR-P, a polarimetric feature selector (PFS) is further introduced to identify the most relevant polarimetric features. Also, a novel adaptive multimodal triple attention mechanism (AMTAM) is designed to facilitate dynamic interactions between different kinds of features. For VLF-SAR-T, after FEM, a multimodal fusion attention mechanism (MFAM) is correspondingly proposed to fuse and adapt information extracted from frozen contrastive language-image pre-training (CLIP) encoders across different modalities. Extensive experiments on the OpenSARShip2.0, FUSAR-Ship, and SAR-AirCraft-1.0 datasets demonstrate the superiority of VLF-SAR over some state-of-the-art methods, offering a promising approach for few-shot SAR target recognition.
Nishang Xie, Tao Zhang 0027, Lanyu Zhang, Feiming Wei, Wenxian Yu
IEEE Trans. Circuits Syst. Video Technol.2
2025 Remote Sensing Scene Classification via Pseudo-Category-Relationand Orthogonal Feature Learning
abstract
Remote sensing (RS) scene classification is a crucial component in the analysis of Earth observation data, aiding in a deeper understanding and monitoring of our dynamic planet. Its applications extend across various fields, including land management, urban analysis, and environmental monitoring. The complex semantic information in RS scene images and the relationships between different scene categories present significant challenges to improving scene classification tasks. Unlike previous methods that only focus on network structure or feature encoding, our approach emphasizes the association of scene categories, integrating feature learning and knowledge transfer together to enhance the analysis of scenes at a higher semantic level. To this end, we propose an RS scene classification scheme based on pseudo-scene category-relation reasoning and orthogonal feature (OF) learning modules, capturing the inherent semantic connections among diverse scene classes. Additionally, we introduce cascaded attention (CA) and selected separation modules to strategically optimize the network, targeting challenging classes with high feature similarities. Knowledge is then distilled across different branches, guiding to enhance the model’s robustness and prediction accuracy. Experiments are conducted on three challenging RS scene datasets of AID30, UCMerced21, and NWPU-RESISC45 to validate the effectiveness of the learned pseudo-category relationships. The results demonstrate that the proposed framework outperforms existing hierarchical approaches in leveraging the hierarchical structure of RS scene images.
Jinsheng Ji, Xiankai Lu, Tao Zhang 0027, Yiyou Guo, Gongping Yang 0001
IEEE Trans. Geosci. Remote. Sens.3
2025 CDPrompt: Multimodal Change Detection With In-Domain Prompt in Missing Modality Scenarios
abstract
The change detection aims to identify temporal changes in land cover. In emergency disaster scenarios, acquiring postchange optical images is often difficult due to factors such as adverse weather and illumination conditions. In contrast, the SAR-based change detection is robust to these environmental factors but is prone to speckle noise and often lacks clear semantic interpretation. These challenges highlight the importance of multimodal approaches that integrate the complementary information from different data sources. To address the domain gap between optical and SAR data, we propose change detection prompt (CDPrompt), an automatic prompt-learning framework that leverages in-domain change information as prompts to suppress fake changes caused by the domain gap between the two modalities. CDPrompt incorporates a modality-specific domain tuning module (DTM) to inject the domain knowledge into the segment anything model (SAM), enabling efficient adaptation to multimodal data with minimal labels and training costs. A low-level enhancement module (LwEM) further refines spatial details using historical optical images, while a consistency loss enhances the learning of domain-invariant features between prechange optical and SAR images. To support evaluation in disaster scenarios with missing modalities, we extend the DFC25 dataset and introduce the first disaster-oriented multimodal change detection dataset, DFC25-Extended, comprising DFC25-OS-S and DFC25-O-SO. Extensive experiments on the Onera Satellite Change Detection (OSCD) and DFC25-Extended datasets demonstrate the superior performance and practical value of CDPrompt. The code and dataset will be publicly available at:https://github.com/zhanglimeng13/CDPrompt
Limeng Zhang, Zenghui Zhang, Tao Zhang 0027, Gui Gao, Wenxian Yu
IEEE Trans. Geosci. Remote. Sens.3
2025 PreCM: The Padding-Based Rotation Equivariant Convolution Mode for Semantic Segmentation
abstract
Semantic segmentation is an important branch of image processing and computer vision. With the popularity of deep learning, various convolutional neural networks have been proposed for pixel-level classification and segmentation tasks. In practical scenarios, however, imaging angles are often arbitrary, encompassing instances such as water body images from remote sensing and capillary and polyp images in the medical domain, where prior orientation information is typically unavailable to guide these networks to extract more effective features. In this case, learning features from objects with diverse orientation information poses a significant challenge, as the majority of CNN-based semantic segmentation networks lack rotation equivariance to resist the disturbance from orientation information. To address this challenge, this paper first constructs a universal convolution-group framework aimed at more fully utilizing orientation information and equipping the network with rotation equivariance. Subsequently, we mathematically design a padding-based rotation equivariant convolution mode (PreCM), which is not only applicable to multi-scale images and convolutional kernels but can also serve as a replacement component for various types of convolutions, such as dilated convolutions, transposed convolutions, and asymmetric convolution. To quantitatively assess the impact of image rotation in semantic segmentation tasks, we also propose a new evaluation metric, Rotation Difference (RD). The replacement experiments related to six existing semantic segmentation networks on three datasets (i.e., Satellite Images of Water Bodies, DRIVE, and Floodnet) show that, the average Intersection Over Union (IOU) of their PreCM-based versions respectively improve 6.91%, 10.63%, 4.53%, 5.93%, 7.48%, 8.33% compared to their original versions in terms of random angle rotation. And the average RD values are decreased by 3.58%, 4.56%, 3.47%, 3.66%, 3.47%, 3.43% respectively. The code can be download from https://github.com/XinyuXu414.
Huazhen Liu, Tao Zhang 0027, Huilin Xiong, Wenxian Yu
IEEE Trans. Image Process.3
2024 Adversarial Attacks with Polarimetric Feature Constraints: A Focused Approach in Polsar Image Classification
abstract
Deep neural networks (DNNs) have been widely utilized in synthetic aperture radar (SAR) for automatic target recognition (ATR), demonstrating remarkable performance. Nevertheless, the vulnerability of DNNs to adversarial examples, particularly in SAR ATR tasks with high safety requirements, necessitates a critical examination. Existing adversarial attacks primarily concentrate on scenarios where classifier inputs consist solely of intensity information, leaving a significant research gap for attacks on classifiers that utilize polarimetric SAR (PolSAR) data. To address this gap, we propose a novel attack method aimed at PolSAR classifiers, where we manipulate the scattering matrix rather than its transformed real value images. Moreover, we incorporate a polarimetric feature constraint to enhance the stealth of the adversarial perturbations. This technique enables the generation of subtle yet effective perturbations concentrated on SAR object regions. Our experiments demonstrate high success rates in cheating state-of-the-art PolSAR classifiers and effectively evading advanced adversarial example detection methods.
Jiyuan Liu 0005, Mingkang Xiong, Zhenghong Zhang, Tao Zhang 0027, Huilin Xiong
IGARSS4
2024 MFSAF: A Plug-And-Play Module for SAR Ship Classification
abstract
This paper introduces a novel plug-and-play Multi-scale Feature Spatial Attention Fusion (MFSAF) module, aiming at enhancing the capabilities of convolutional neural networks (CNNs) in Synthetic Aperture Radar (SAR) ship classification tasks. The MFSAF module integrates spatial attention mechanisms and feature alignment strategies, providing a seamless integration into general CNNs to better capture ship features of different scales. The experimental results on the OpenSARShip2.0 and FUSARShip datasets demonstrate a significant improvement of the "baseline+MFSAF" model compared to baseline model, highlighting the effectiveness of the MFSAF module in capturing SAR ship features and its adaptability across different networks.
Nishang Xie, Mingkang Xiong, Feiming Wei, Tao Zhang 0027, Wenxian Yu
IGARSS5
2024 CA-LOSS: A Cosine Affinity Loss for Imbalanced SAR Ship Classification
abstract
To address the problem of imbalanced datasets in SAR ship classification, this paper presents a novel cosine affinity (CA) loss that enhances the Gaussian affinity (GA) loss. The CA loss focuses on the angular relationship between feature vectors, prioritizing their direction over their magnitude, which is advantageous for high-dimensional space analysis. In addition, class weights are incorporated to compute weighted distances. Importantly, the proposed CA loss does not increase the computational complexity of algorithm, nor does it lead to overfitting problems associated with data-level techniques. Through various experiments, its effectiveness has been demonstrated by achieving the highest F1 score and recall compared to other existing loss functions, highlighting its superior ability to classify minority classes in FUSARShip.
Nishang Xie, Mingkang Xiong, Feiming Wei, Tao Zhang 0027, Zhen Yang 0012, Wenxian Yu
IGARSS4
2024 Flood Change Detection Based on Prior Feature Estimation
abstract
Flood caused by torrential rain is one of the most influential meteorological disasters in the world. Currently, most of the flood change detection networks are based on homogeneous images. However, due to the influence of bad weather and satellite revisit cycle, the acquisition of homogeneous images is greatly limited. In this paper, a novel heterogeneous image change detection network based on prior feature estimation is proposed. In order to better guide the network to find the solution space related to change, we propose feature enhancement module to strengthen the water body features and introduce auxiliary information. At the same time, the fusion module is designed to solve the problem that the feature space of heterogeneous images is difficult to align while mapping water features into changing features. Experimental results on the CAU-Flood dataset demonstrate the effectiveness of our network.
Mingkang Xiong, Sinong Quan, Tao Zhang 0027, Feiming Wei
IGARSS5
2024 An Approach for Integrating SAR Imagery in Sea-Land Segmentation and Coastline Detection
abstract
Segmentation of Synthetic Aperture Radar (SAR) imagery constitutes the cornerstone of SAR image analysis, with sea-land segmentation in SAR images playing a crucial role in determining the precision of subsequent sea surface target detection. This study introduces an integrated approach for sea-land segmentation and coastline detection in SAR imagery, aiming to overcome the limitations posed by the traditional separation of these two tasks. In essence, the proposed method merges a segmentation module and an edge detection module, employing a hollow convolution and a global context mechanism. Additionally, the approach utilizes a cross-entropy loss function incorporating multiple losses with adaptive weighting, thereby enhancing the richness of the extracted feature information. To validate the algorithm’s efficacy, a specialized dataset for sea-land segmentation and coastline detection is constructed, utilizing the GRD data format from the Sentinel-1 satellite. Experimental outcomes show that the presented algorithm achieves scores of 0.988 and 0.981 on the Intersection over Union (IOU) metrics for sea-land segmentation, and 0.569 and 0.401 on the Optimal Dataset Scale (ODS) F1 and ODS IOU metrics for coastline detection.
Renke Zhu, Mingkang Xiong, Tao Zhang 0027, Feiming Wei, Sinong Quan, Wenxian Yu
IGARSS3
2024 Monocular depth estimation using self-supervised learning with more effective geometric constraints
Mingkang Xiong, Zhenghong Zhang, Jiyuan Liu 0005, Tao Zhang 0027, Huilin Xiong
Eng. Appl. Artif. Intell.4
2024 Low-rank preserving embedding regression for robust image feature extraction
abstract
Abstract Although low‐rank representation (LRR)‐based subspace learning has been widely applied for feature extraction in computer vision, how to enhance the discriminability of the low‐dimensional features extracted by LRR based subspace learning methods is still a problem that needs to be further investigated. Therefore, this paper proposes a novel low‐rank preserving embedding regression (LRPER) method by integrating LRR, linear regression, and projection learning into a unified framework. In LRPER, LRR can reveal the underlying structure information to strengthen the robustness of projection learning. The robust metric L 2,1 ‐norm is employed to measure the low‐rank reconstruction error and regression loss for moulding the noise and occlusions. An embedding regression is proposed to make full use of the prior information for improving the discriminability of the learned projection. In addition, an alternative iteration algorithm is designed to optimise the proposed model, and the computational complexity of the optimisation algorithm is briefly analysed. The convergence of the optimisation algorithm is theoretically and numerically studied. At last, extensive experiments on four types of image datasets are carried out to demonstrate the effectiveness of LRPER, and the experimental results demonstrate that LRPER performs better than some state‐of‐the‐art feature extraction methods.
Tao Zhang 0027, Chen-Feng Long, Yangjun Deng, Wei-Ye Wang, Siqiao Tan, Heng-Chao Li 0001
IET Comput. Vis.1
2024 An efficient feature pyramid attention network for person re-identification
Wanli Dang, Libo Cao, Tao Zhang 0027
Image Vis. Comput.6
2024 PolSAR Ship Detection Based on Superpixel-Level Contrast Enhancement
abstract
Ship detection in polarimetric synthetic aperture radar (PolSAR) images has attracted widespread attention in recent years. However, pixel level detection methods are heavily affected by inherent speckle noise. In this letter, we proposed a detection method that enhances the ship-sea contrast beforehand by combining local statistical saliency and scattering mechanism coherence in superpixel-level. Firstly, simple linear iterative clustering (SLIC) based segmentation method is adopted for PolSAR images to generate superpixels. Then, local saliency is calculated based on superpixel-level similarity from the perspective of statistical characteristics. Based on this, the superpixel-level modified polarimetric coherence metric is obtained from the perspective of physical scattering mechanisms, which can help distinguish small ships with low saliency and strong sea clutters with high saliency. Ship detection is achieved by combining the two features above. The experimental results based on real PolSAR data show that compared with other classic and state-of-the-art methods, the proposed method has improved the figure of merit by at least 4.28% and has increased the target clutter ratio by at least 8.43 decibel (dB) on average.
Jie Deng 0004, Wei Wang 0099, Huiqiang Zhang, Tao Zhang 0027, Jun Zhang 0044
IEEE Geosci. Remote. Sens. Lett.4
2024 PolSAR Ship Targets Generation via the Polarimetric Feature Guided Denoising Diffusion Probabilistic Model
abstract
The generation of realistic synthetic aperture radar (SAR) images holds notable significance due to their applicability across various crucial domains in remote sensing, such as automatic target recognition and electronic countermeasures. The majority of current SAR image synthesis methods only leverage amplitude, thereby lacking the phase information that also plays an important role in SAR image interpretation. In this letter, we introduce Polarimetric Feature Guided Denoising Diffusion Probabilistic Model (PFG-DDPM), to generate PolSAR images. The proposed method can effectively simulate the distribution of real PolSAR images, encompassing both their amplitude and phase components. Importantly, we introduce an innovative strategy that employs polarimetric features as supervised information to guide the generation process of PolSAR images. This approach allows PFG-DDPM to effectively utilize constraints among distinct polarimetric channels, resulting in generated PolSAR images whose distributions closely approximate real PolSAR data. Experiments underscore the ability of the proposed method to produce realistic PolSAR images valid for human visual perception. More significantly, these images exhibit a remarkable resemblance to real PolSAR images, evidenced by a 44.4% and 5.3% enhancement in alignment with polarimetric and statistical attributes, respectively, compared to the vanilla DDPM.
Jiyuan Liu 0005, Tao Zhang 0027, Huilin Xiong
IEEE Geosci. Remote. Sens. Lett.2
2024 Evaluating the Robustness of Polarimetric Features: A Case Study of PolSAR Ship Detection
abstract
Polarimetric features, no matter how they are extracted, play a crucial role in synthetic aperture radar (SAR) image interpretation. Since recently developed deep neural network (DNN)-based methods are recognized as very vulnerable to some specifically designed perturbations, it raises an interesting question: are the traditional polarimetric features, calculated according to the scattering mechanism of SAR, still robust to the specially crafted perturbation? In this letter, we investigate the robustness of several traditional polarimetric features, which are widely used in the application of ship target detection, to a particularly designed perturbation, which we call polarimetric interference method (PIM). Specifically, we first calculate the gradients of traditional polarimetric features, respecting each of the polarimetric channels, and then formulate the PIM perturbation as an optimization problem, and finally, we give the PIM algorithm to perturb polarimetric features. Experiments are carried out on two polarimetric SAR (PolSAR) datasets to demonstrate the following: 1) traditional polarimetric features are vulnerable to the PIM perturbation, even if it is imperceptible in the SAR images; 2) the PIM perturbation can remarkably degrade the performance of the traditional polarimetric features in the task of ship detection, leading to more than 50% decrease in detection rate; and 3) the proposed perturbation is transferable, which means that the PIM interference regarding one type of the traditional polarimetric features also works in most cases for other types of polarimetric features.
Jiyuan Liu 0005, Tao Zhang 0027, Zhenghong Zhang, Huilin Xiong
IEEE Geosci. Remote. Sens. Lett.2
2024 PFDN: A Polarimetric Feature-Guided Deep Network for Dual-Polarized SAR Ship Classification
abstract
As one important application of synthetic aperture radar (SAR), ship classification attracts researchers’ attention in recent years. To improve the accuracy of ship classification in dual-polarized SAR images, we here put forward a novel polarimetric feature-guided deep network PFDN. Detailedly, a new polarimetric feature SPF (Smoothed-Polarimetric Information-Fusion) is first built through fusing both amplitudes of dual-polarized channels. Since only the amplitude information is used, SPF cannot be affected by the phase noise. Then, another two key components, i.e., the Multi-scale Feature Fusion Attention Module (MFFAM) and the Dynamic Gating Feature Fusion Mechanism (DGFFM), are further proposed to extract deeper classification features. Finally, via combing these three different modules together, PFDN is constructed. Experiments tested on the dataset OpenSARShip2.0 show that, PFDN can achieve higher accuracies (87.13% in three-class task and 65.97% in six-class task) than the other state-of-the-art (SOTA) methods.
Nishang Xie, Tao Zhang 0027, Feiming Wei, Wenxian Yu
IEEE Geosci. Remote. Sens. Lett.2
2024 An Information-Expanding Network for Water Body Extraction Based on U-Net
abstract
Water body extraction is an important issue in flood surveillance and environmental protection. With the development of neural network, deep learning has been widely used in the water body extraction task because of its powerful feature extraction ability. Even so, most of existing deep learning networks only take into account the translation equivariance of convolution kernel for water body extraction. Actually, in terms of water body, its orientations imaged by optical sensor are usually various. So, when the orientated images are not well contained in the training set, the networks may yield some unsatisfactory extraction results. To solve this problem, in this paper, we propose an information-expanding network IE-Unet based on the traditional network U-net, where the rotation equivariant convolution, rotation-based channel attention mechanism, and the optimized Batchnorm are adopted jointly. To quantitatively evaluate its edge extraction capability, a new edge index AOD is proposed as well. The experimental results on one public dataset of water body demonstrate the effectiveness of IE-Unet. Compared with the original U-net, the IOU value of IE-Unet is increased by 7%, and the A0D value is reduced by 0.76.
Tao Zhang 0027, Huazhen Liu, Weiwei Guo, Zenghui Zhang
IEEE Geosci. Remote. Sens. Lett.2
2024 A Network for Merging SAR Image Sea-Land Segmentation and Coastline Detection Tasks
abstract
For the task of marine target detection in synthetic aperture radar (SAR) images, sea-land segmentation and coastline detection are often essential. Despite exciting results, many of them are still separately performed. Only a few studies have been done on the simultaneous realization of sea-land segmentation and coastline detection. To this end, this letter proposes a new network SAENet, wherein one edge enhancement module (EEM), one maximum fusion difference convolution (MaxFDC), and one multiscale spatial attention module (Multiscale SAM) are developed. In order to verify its effectiveness, we further construct one sea-land segmentation and coastline detection dataset with the Sentinel-1 ground range detected (GRD) data. The corresponding experimental results show that SAENet can reach 0.989 and 0.980 on the evaluation indexes$F1$and IOU for sea-land segmentation, and 0.577 and 0.391 on the evaluation indexes ODS$F1$and ODS IOU for coastline detection, which better accomplishes the task of sea-land segmentation and coastline detection simultaneously in comparison with other state-of-the-art (SOTA) methods.
Renke Zhu, Tao Zhang 0027, Feiming Wei, Wenxian Yu
IEEE Geosci. Remote. Sens. Lett.2
2024 Dual Branch Deep Network for Ship Classification of Dual-Polarized SAR Images
abstract
Ship classification is usually a challenging task due to the small sizes of ship targets and the lack of significant differences between different categories. In terms of synthetic aperture radar (SAR) images, most existing deep learning-based methods are not designed from the angle of polarimetric characteristics to achieve ship classification. Thus, when facing the ship classification task of dual-polarized SAR images, these networks are often unsatisfactory. To cure this shortcoming, we here propose a novel dual branch deep network DBDN specifically designed for dual-polarized SAR ship classification. Our approach consists of three key modules: the image construction module ICM, the feature extraction module FEM, and the feature fusion and classifier module FFCM. In ICM, two novel pseudo RGB images are constructed for the first time, i.e., the polarimetric features-guided pseudo RGB image (PF-RGB) and the texture features-guided pseudo RGB image (TF-RGB), which can more accurately and comprehensively reflect ships’ characteristics. FEM enables the network to focus on important ship features and suppress irrelevant noise through transferred layers and designed ConvNeXt-Attention block (CNABlock), enhancing the discriminative capability of different ships. Finally, FFCM extracts and combines various ship features for classification, wherein the enhanced inverted residual block (EIRBlock) and the channel spatial attention module (CSAM) components are proposed as well. The performance of DBDN is evaluated on the OpenSARShip2.0 dataset, and experimental results show that DBDN achieves excellent performance in all evaluation metrics in comparison with some state-of-the-art (SOTA) algorithms. For example, compared to the recently proposed method DSN, DBDN further improves the accuracy by 4.74% and 4.07% in the three-class and six-class classification tasks, respectively.
Nishang Xie, Tao Zhang 0027, Weiwei Guo, Zenghui Zhang, Wenxian Yu
IEEE Trans. Geosci. Remote. Sens.2
2023 From Coarse to Fine: Learning Semantic Relations for Hyperspectral Image Classification
abstract
Hyperspectral images (HSIs) is consisted of many narrow spectral bands which are capable of recording abundant features including both the spectral and spatial signatures information which have been widely used in various fields, such as urban planning, disaster monitoring. Due to the large number and high similarity of the collected spectral brands, many methods are developed to handle the problem of extracting effective features. Recently, many CNN-based methods have been proposed by exploiting the spectral-spatial signatures of the HSIs data and achieved promising results. Although many methods adopt patch-based input pattern to emphasize the importance of the spatial neighbor information of each pixel, the relations are still limited to a small area around the pixel and the latent relations among the pixels belonging to different semantic categories at the boundary are still not well exploited. To explore the relationships between pixels from a more global perspective, a neighbor-based relation mining framework is proposed to explore the long-range relations among different local regions. Experiments are conducted on two hyperspectral image classification datasets and the results demonstrate the effectiveness of the proposed long-range relations mining scheme by comparison with some state-of-the-art methods.
Jinsheng Ji, Xiankai Lu, Tao Zhang 0027, Yiyou Guo, Huan Xie 0001
IGARSS3
2023 Water Body Detection Based on an Improved U-Net
abstract
Nowadays, deep learning has been widely used for water body detection because of its high precision databased water segmentation ability. Although the networks based on deep learning have shown higher automation, applicability and extraction accuracy than the traditional threshold methods in water body detection, only the translation equivariance of the convolution kernel is considered in these networks. Actually, for the detection of water body, its rotation equivariance also needs to be considered. For this goal, we here propose a new convolutional neural network by improving the U-Net with the rotation equivariant convolution and attention mechanism, which is simplified as GACNN. Experimental results on optical water body images demonstrate the effectiveness of the improved network based on U-net.
Huazhen Liu, Tao Zhang 0027, Zenghui Zhang, Weiwei Guo
IGARSS3
2023 Ship Detection with the Nonlocal Information-Based Polarimetric Covariance Matrix
abstract
Ship plays an important role in human marine production and living activities at sea. In this paper, we design a ship detection method for polarimetric synthetic aperture radar (Pol-SAR) images. In brief, one nonlocal neighborhood polarimetric covariance matrix [NC] is first built by improving the neighborhood polarimetric covariance matrix [N] with a new similarity parameter rI. Then, the proposed method PWFNCis achieved through directly computing the polarimetric whiten filter (PWF) with [NC]. Experiments carried out on two real PolSAR datasets show that, compared to the recently proposed matrix [N], [NC] can better improve the ship detection performance of PWF.
Tao Zhang 0027, Wenxian Yu, Yonghu Zhang, Weiwei Guo
IGARSS1
2023 Low frequency sparse adversarial attack
Jiyuan Liu 0005, Bingyi Lu, Mingkang Xiong, Tao Zhang 0027, Huilin Xiong
Comput. Secur.4
2023 Learning-based padding: From connectivity on data borders to data padding
Chao Ning 0003, Hongping Gan, Minghe Shen, Tao Zhang 0027
Eng. Appl. Artif. Intell.4
2023 Self-supervised depth completion with multi-view geometric constraints
abstract
Abstract Self‐supervised learning‐based depth completion is a cost‐effective way for 3D environment perception. However, it is also a challenging task because sparse depth may deactivate neural networks. In this paper, a novel Sparse‐Dense Depth Consistency Loss (SDDCL) is proposed to penalize not only the estimated depth map with sparse input points but also consecutive completed dense depth maps. Combined with the pose consistency loss, a new self‐supervised learning scheme is developed, using multi‐view geometric constraints, to achieve more accurate depth completion results. Moreover, to tackle the sparsity issue of input depth, a Quasi Dense Representations (QDR) module with triplet branches for spatial pyramid pooling is proposed to produce more dense feature maps. Extensive experimental results on VOID, NYUv2, and KITTI datasets show that the method outperforms state‐of‐the‐art self‐supervised depth completion methods.
Mingkang Xiong, Zhenghong Zhang, Jiyuan Liu 0005, Tao Zhang 0027, Huilin Xiong
IET Image Process.4
2023 Dynamic representation-based tracker for long-term pedestrian tracking with occlusion
Zhen Yang 0012, Zhiyi Huang 0003, Dunyun He, Tao Zhang 0027, Fan Yang 0042
J. Vis. Commun. Image Represent.4
2023 Occluded Target Recognition in SAR Imagery With Scattering Excitation Learning and Channel Dropout
abstract
Deep neural networks are widely used in SAR image classification and recognition, achieving state-of-the-art performance. But it remains a challenging task to recognize occluded targets. In this letter, we propose a novel robust SAR recognition method against occlusion. Specifically, we design a scattering excitation learning module that encourages the network to learn more robust features responding to the scattering centers of targets. In addition, we adopt a random feature channel dropout technique which can further improve robustness to occlusion. Our method makes the network more robust against occlusion but without any occlusion-simulated data for training. Experimental results on MSTAR dataset shows that our proposed method achieves remarkably improved robustness even under severe occlusions. Code is made available at https://github.com/koervcor/SEL-CD.
Dunyun He, Weiwei Guo, Tao Zhang 0027, Zenghui Zhang, Wenxian Yu
IEEE Geosci. Remote. Sens. Lett.3
2023 Multidimensional Information Expansion and Processing Network for Hyperspectral Image Classification
abstract
In recent years, deep learning has been extensively used in hyperspectral image (HSI) classification. The representative method is the convolutional neural network (CNN). However, due to the limitations of its inherent network backbone, CNNs still easily fail to mine some important information of HSIs, such as the sequence attributes of spectral signatures. To deal with this problem and make full use of the spectral-spatial information of HSIs, we propose a novel network named Multi-dimensional Information Expansion and Processing Network (MIEPN) for HSI classification, which is mainly composed of one information expansion module (IEM), one feature information expansion and extraction module (FEEM), and one ViT module. Briefly speaking, IEM expands and fuses HSI information in a three-dimensional (3D) space, yet FEPM pays more attention to digging deeper information. After these, the extracted information is input into the ViT module for HSI classification. Experiments carried out on several typical datasets demonstrate that the proposed network MIEPN can provide competitive results compared to the other state-of-the-art CNN-based methods.
Zhen Yang 0012, Tao Zhang 0027, Weiwei Guo, Zenghui Zhang
IEEE Geosci. Remote. Sens. Lett.3
2023 3DMAE: Joint SAR and Optical Representation Learning With Vertical Masking
abstract
The remote sensing community has shown increasingly interest in self-supervised learning for its ability to learn representations without labeled data. These representations can be easily adapted to downstream tasks through pre-training and fine-tuning. Recently, Masked Autoencoders (MAE) achieve better semantic representation by masking out a significant portion of the input image. However, the original design of MAE for RGB natural images may not be optimal for remote sensing (RS) images, which exhibit considerable variation between modalities like SAR and optical. To address this, we propose a 3D mask that enhances feature extraction along the vertical dimension. After fine-tuning, our 3DMAE model outperforms state-of-the-art contrastive and MAE-based models on BigEarthNet-MM classification and significantly reduces input data volume by at least 50% with the vertical mask, resulting in a more efficient model. Generalization experiments show a 5.9% F1-score improvement when applied to the SEN12MS dataset, which has diverse data distributions.
Limeng Zhang, Zenghui Zhang, Weiwei Guo, Tao Zhang 0027, Wenxian Yu
IEEE Geosci. Remote. Sens. Lett.4
2023 AutoBCS: Block-Based Image Compressive Sensing With Data-Driven Acquisition and Noniterative Reconstruction
abstract
Block compressive sensing (CS) is a well-known signal acquisition and reconstruction paradigm with widespread application prospects in science, engineering, and cybernetic systems. However, state-of-the-art block-based image CS (BCS) methods generally suffer from two issues. The sparsifying domain and the sensing matrices widely used for image acquisition are not data driven and, thus, both the features of the image and the relationships among subblock images are ignored. Moreover, it requires to address a high-dimensional optimization problem with extensive computational complexity for image reconstruction. In this article, we provide a deep learning (DL) strategy for BCS, called AutoBCS, which automatically takes the prior knowledge of images into account in the acquisition step and establishes a reconstruction model for performing fast image reconstruction. More precisely, we present a learning-based sensing matrix to accomplish image acquisition, thereby capturing and preserving more image characteristics than those captured by the existing methods. In addition, we build a noniterative reconstruction network, which provides an end-to-end BCS reconstruction framework to maximize image reconstruction efficiency. Furthermore, we investigate comprehensive comparison studies with both traditional BCS approaches and newly developed DL methods. Compared with these approaches, our proposed AutoBCS can not only provide superior performance in terms of image quality metrics (SSIM and PSNR) and visual perception but also automatically benefit reconstruction speed.
Hongping Gan, Yang Gao 0030, Chunyi Liu, Haiwei Chen, Tao Zhang 0027, Feng Liu 0005
IEEE Trans. Cybern.5
2023 A System Optimization Scheme for Bias Correction of Polarimetric Phased-Array Radar
abstract
With the change in the spatial angle, the cross-polarization isolation (CPI) of polarimetric phased-array radar (PPAR) changes as well, destroying the estimation of the target polarization scattering matrix (PSM). To correct the bias in PPAR, this article comprehensively designs the transmitting antenna, receiving antenna, and signal waveform and proposes a bias correction method based on system optimization. First, for the transmitting antenna of PPAR, the second-order cone program (SOCP) model is proposed to optimize the weighting coefficient. With the SOCP-based beamforming method, not only beam pattern in any spatial angle can be achieved, but also arbitrary polarization state can be precisely configured. Then, for the wideband receiving signal with a certain beamwidth, an angle estimation method based on eigenvalue decomposition is proposed in this article, which can effectively cure the challenges introduced by the beamwidth and signal bandwidth. Subsequently, for the signal waveform, the phase code is designed to measure all the elements of PSM in simultaneous transmission and simultaneous reception (STSR) mode, which could eliminate the biases of the moving speed and the second-order cross-polarization error. Finally, this article compares with other methods based on differential reflectivity, and experiments show that the factors such as the spatial angle, array structure, antenna beamwidth, signal bandwidth, motion speed, and signal-to-noise ratio (SNR) have the least influence on the present method in this article.
Yaomin He, Tao Zhang 0027, Huafeng He, Junjun Yin 0001, Jian Yang 0011
IEEE Trans. Geosci. Remote. Sens.2
2023 Self-Supervised Classification of SAR Images With Optical Image Assistance
abstract
Supervised Deep Neural Networks (DNNs) have proven to be powerful tools for SAR image interpretation tasks. However, they present a formidable challenge in acquiring a substantial amount of labeled data. In this paper, we investigate the promising technique of contrastive self-supervised learning for SAR image classification. This approach allows us to take advantage of a large number of available unlabeled images to pre-train a SAR image classification model. Our novel contrastive learning framework conducts both instance-level and cluster-level pretext tasks, which not only enforce consistency between the images and their augmented "views" at the instance level but also their representation within clusters. Besides generating different views through random, low-level image transformations, we proposed two new strategies to construct positive sample pairs to improve contrastive SAR image feature learning: middle-level optical assistance and high-level graph searching. The middle-level optical assistance strategy is inspired by the observation that domain experts typically interpret SAR images with the aid of optical images. This insight spurs us to generate intermediate SAR images as positive samples from geographically matched optical data using CycleGAN. Furthermore, we augment the positive samples of the image with their KNN (K-Nearest Neighbor) counterparts, following the idea that the KNN samples should belong to the same cluster. Extensive experimental results conducted on the SEN12MS land cover classification benchmark dataset demonstrate that our method is competitive with state-of-the-art self-supervised methods for SAR image classification. Even with only a small amount of labeled data for fine-tuning the model, our method rapidly improves classification performance, surpassing models pre-trained on natural image datasets.
Chenxuan Li 0002, Weiwei Guo, Zenghui Zhang, Tao Zhang 0027
IEEE Trans. Geosci. Remote. Sens.4
2023 Residual in Residual Scaling Networks for Polarimetric SAR Image Despeckling
abstract
Speckle reduction is a longstanding topic for polarimetric synthetic aperture radar (PolSAR) images. In this paper, we propose a novel end-to-end PolSAR image despeckling framework for the first time, which predicts the weight matrices of neighboring pixels instead of the target pixel itself nor the nor the noise, to achieve image despeckling. It hardly relies on any assumptions on the speckle noise distribution. Within this framework, residual in residual scaling network (RIRSN) is developed by combining the advantages of residual connections and residual scaling. To reduce network redundancy further, a dynamic version of RIRSN (DRIRSN) is also proposed by adjusting the network structure dynamically based on noise level and image content. Specifically, in DRIRSN, we introduce a lightweight network called picture2vector to estimate noise level, and a well-designed loss function to estimate image information level and measure image denoising quality simultaneously. The proposed picture2vector and loss function guide DRIRSN to focus on image areas with rich content and information, enhancing the adaptability of the network. DRIRSN inherits the properties of RIRSN for adaptively selecting and weighting the pixels of the neighborhood, and dynamically adjusts the network structure according to the estimated noise level and image content. We compare the proposed networks with reference methods on both simulated images and real images. Experimental results demonstrate that the proposed networks can effectively reduce speckle noise with low time consumption and, meanwhile, better preserve the details and the repetitive structures such as textures and edges, and the polarimetric scattering characteristics, compared with the other methods.
Kan Jin, Junjun Yin 0001, Jian Yang 0011, Tao Zhang 0027, Feng Xu 0001, Ya-Qiu Jin
IEEE Trans. Geosci. Remote. Sens.5
2023 Exploring Fine Polarimetric Decomposition Technique for Built-Up Area Monitoring
abstract
Highly variable polarimetric signatures caused by complex structures in built-up areas make interpretation of these scattering behaviors intractable for PolSAR remote sensing. This paper proposes a fine polarimetric decomposition method and derives several products to finely simulate the scattering mechanisms of urban buildings, thus fulfilling its use for effective surveillance. First, through theoretically establishing the roll-invariant condition for a completely general scatterer, a roll-invariant cross polarization (RICP) scattering model is constructed, which characterizes the cross polarization scattering in the manner of planar structure distribution. Second, by designing a root-discriminant-based parameter inversion strategy, a fine seven-component decomposition is proposed, which achieves the complete physical interpretation of matrix elements and reasonable inversion of model parameters. Third, by analyzing the external and internal scattering difference, the derivative products, i.e., scattering contribution synthesizers are derived for built-up area monitoring. Experimental results derived from real PolSAR data confirm the superiority and effectiveness of the constructed descriptors on the one hand. On the other hand, the extensibility of fine polarimetric decomposition in specific remote sensing is also explicitly demonstrated.
Sinong Quan, Tao Zhang 0027, Wei Wang 0099, Gangyao Kuang, Xuesong Wang 0003, Bing Zeng 0001
IEEE Trans. Geosci. Remote. Sens.2
2023 Information Reconstruction-Based Polarimetric Covariance Matrix for PolSAR Ship Detection
abstract
In the last decades, how to detect ships with polarimetric synthetic aperture radar (PolSAR) has become one hot topic. Unfortunately, most of the existing ship detection methods cannot well detect small ships with weak backscattering. To deal with this issue, a ship detection matrix named complete polarimetric covariance matrix [CP] was recently proposed from the perspective of spatial information utilization. Although it is able to improve small ships’ target-to-clutter ratio (TCR) values, its calculation strategy still needs to be rethought due to the possible information loss of some ships. Besides, its mathematical characteristic (i.e., not positive semidefinite) also limits the successful applications of some existing polarimetric theories to it. To overcome these two drawbacks, we here develop an information reconstruction-based polarimetric covariance matrix [IC]. In brief, one new difference calculation strategy is first performed on the Sinclair matrix [$S$], so as to reconstruct its information, by which a feature vector$v$is subsequently extracted with the Lexicographic matrix basis. Then, via further performing an outer product operation on$v$, the matrix [IC] is proposed. Meanwhile, to demonstrate the effectiveness of [IC] in ship detection, two different [IC]-based intensity detectors, respectively, named SPANIC and PEDIC, are designed as well. Experiments carried out on three GF-3 PolSAR datasets show that: 1) the proposed matrix [IC] has a better performance than [CP] and the original polarimetric covariance matrix [$C$] in ship detection and 2) compared to the total power detector SPAN and geometrical perturbation-polarimetric notch filter (GP-PNF), both SPANIC and PEDIC can better detect ships, especially the small ships.
Tao Zhang 0027, Sinong Quan, Wei Wang 0099, Weiwei Guo, Zenghui Zhang, Wenxian Yu
IEEE Trans. Geosci. Remote. Sens.1
2023 Real-time instance segmentation with assembly parallel task
Zhen Yang 0012, Fan Yang 0042, Zhijian Yin, Tao Zhang 0027
Vis. Comput.5
2022 Polsar Ship Detection with the Sub-Aperture Technology
abstract
Polarimetric synthetic aperture radar (PolSAR) designed to obtain the polarimetric information of scenes is a crucial tool for microwave remote sensing. Recently, a complete polarimetric covariance difference matrix [CP] was built to detect ships of PolSAR image. Along this work, this paper extends its application to the spectrum domain. Briefly speaking, four sub-aperture images are first separated from the original PolSAR data. Then, four different power values corresponding to the [CP] matrices of these sub-aperture images are respectively calculated. At last, via multiplying these values together, a PolSAR ship detector named MPS (Multiplicative Polarimetric SPAN) is proposed. The experiment carried out on one real PolSAR image demonstrates that, compared to traditional power detectors SPAN and$SPAN_{CP,}$MPS holds a better ability to detect small ships.
Tao Zhang 0027, Zenghui Zhang, Weiwei Guo, Huilin Xiong, Wenxian Yu
IGARSS1
2022 ACSiam: Asymmetric convolution structures for visual tracking with Siamese network
Zhen Yang 0012, Chaohe Wen, Lingkun Luo, Hongping Gan, Tao Zhang 0027
J. Vis. Commun. Image Represent.5
2022 A New Form of the Polarimetric Notch Filter
abstract
Ship detection using polarimetric synthetic radar (PolSAR) imagery attracts a lot of attention in recent years. Most notably, the detector polarimetric notch filter (PNF) has been demonstrated to be effective for ship detection in PolSAR imagery, which gives excellent performances. In this work, a mathematical form of one new PNF (NPNF) based on physical mechanisms of targets and clutter is further developed for partial targets. The different mechanisms have been revealed based on the projection matrix. The experimental results including simulated and measured data demonstrate that the NPNF exhibits a better performance than the original PNF.
Tao Liu 0025, Ziyuan Yang 0002, Tao Zhang 0027, Yanlei Du, Armando Marino
IEEE Geosci. Remote. Sens. Lett.3
2022 Dual-Polarized SAR Ship Grained Classification Based on CNN With Hybrid Channel Feature Loss
abstract
This letter proposes a novel convolutional neural network (CNN) method for dual-polarized synthetic aperture radar (SAR) ship grained classification. The network employs hybrid channel feature loss that jointly utilizes the information contained in the polarized channels (VV and VH). It is demonstrated that, by adopting the proposed CNN framework and the novel loss function, the classification performance can be efficiently improved. First, instead of the prevalently used threefold or fourfold division (container ship, oil tanker, bulk carrier, and so on), the proposed method can further divide vessels into eight accurate categories. Second, this method can not only effectively classify targets into eight categories but also its accuracy in terms of fewer category classifications surpasses existing methods. Third, the method can achieve good performance on a small training data set. Experiments conducted on the OpenSARShip data sets indicate that the proposed classification method achieves state-of-the-art results.
Qingtao Zhu, Danwei Lu, Tao Zhang 0027, Hongmiao Wang, Junjun Yin 0001, Jian Yang 0011
IEEE Geosci. Remote. Sens. Lett.4
2022 PolSAR Ship Detection Using the Superpixel-Based Neighborhood Polarimetric Covariance Matrices
abstract
In order to detect ships from the imagery of polarimetric synthetic aperture radar (PolSAR), a neighborhood polarimetric covariance matrix (for simplicity, we call it [$N$] hereinafter) was recently constructed. However, its calculation process is time-consuming and the backscattering heterogeneity near ship edges is also not well considered. For curing these shortcomings, we here propose two novel superpixel-based neighborhood polarimetric covariance matrices. In brief, the first matrix denoted by [SN] uses the simple linear iterative clustering (SLIC) to yield superpixels, whereas in the second matrix denoted by [GN], the gradient operator Sobel is adopted to obtain superpixels. Based on these two different kinds of superpixels, then, two different feature vectors$v_{\text {SN}}$and$v_{\text {GN}}$are separately built to compute [SN] and [GN]. Experiments performed on the real PolSAR datasets show that, compared to [$N$], [SN] and [GN] can improve the performance of the polarimetric whitening filter (PWF) more significantly and the time consumptions of calculating [SN] and [GN] are both much less.
Tao Zhang 0027, Yanlei Du, Zhen Yang 0012, Sinong Quan, Tao Liu 0025, Fengtao Xue, Zhengzheng Chen, Jian Yang 0011
IEEE Geosci. Remote. Sens. Lett.1
2022 Ship Detection of Polarimetric SAR Images Using a Nonlocal Spatial Information-Guided Method
abstract
Ship detection of polarimetric synthetic aperture radar (PolSAR) plays an important role in marine monitoring and ocean protection. Over the past years, local spatial information around pixels has been successfully applied to this task. However, few works have been done on PolSAR ship detection using the nonlocal spatial information (NSI). Within this context, we here propose one NSI-guided ship detection method PMR. Briefly speaking, the feature power difference (PD) is first constructed by computing the total power difference between the center pixelcand its most similar nonlocal pixeliwithin a 7×7 window. Then, the polarimetric feature reflection symmetry (RS) is introduced into PD to construct the method PMR (i.e., PD Multiply RS) for further enhancing the target-to-clutter ratio (TCR) and improving the ship detection accuracy. Experiments carried out on three real PolSAR datasets show that, in comparison with some other methods, especially the recently proposed local neighborhood information-based ship detector PWFN, PMR is more apt for ship detection. On average, its figure of merit (FoM) and TCR values respectively surpass PWFN0.24 and 12.83 dB.
Tao Zhang 0027, Zenghui Zhang, Huizhang Yang, Weiwei Guo, Zhen Yang 0012
IEEE Geosci. Remote. Sens. Lett.1
2022 Active Learning SAR Image Classification Method Crossing Different Imaging Platforms
abstract
Synthetic aperture radar (SAR) image classification task when the training and test sets have different distributions can be initially solved using existing domain adaptation (DA) methods. However, considering that none of their classification accuracy is high, this letter proposes an active learning DA classification method to further solve this task. First, an adversarial learning-based DA pipeline is put forth, using labeled source and unlabeled target domains to conduct adversarial learning in order to narrow the domain gap. A prototype regularization process is then built, which further enhances the target domain data clusters’ ability to discriminate between them. In order to fully improve SAR image classification accuracy, we then propose a dynamic hard sample selection process to choose hard samples to supplement into the subsequent stage of training samples. This process involves moving the gradient direction of the query function closer to the gradient direction of the class margin objective function. Extensive experiments on SAR image datasets with different distributions from different imaging platforms and optical remote sensing datasets have verified the effectiveness and superiority of the proposed method.
Ying Luo 0001, Tao Zhang 0027, Weiwei Guo, Zenghui Zhang
IEEE Geosci. Remote. Sens. Lett.3
2022 Transferable SAR Image Classification Crossing Different Satellites Under Open Set Condition
abstract
For synthetic aperture radar (SAR) image classification problem, we need to take into account unlabeled datasets containing unknown classes crossing different satellites. In this letter, a spherical space domain adaptation (DA) network under open set condition is proposed to solve this problem. First, we transform the prior Euclidean feature space into the spherical space to construct a classification network such that features of the same class of SAR images are clustered together and features of different or unknown classes are separated on the hypersphere. Second, a correction module is designed to increase the accuracy of the pseudo-label obtained by the classifier. Then, based on the adversarial learning strategy, we introduce the gradient alignment module to achieve better alignment of the source and target domains. Finally, tests on two SAR benchmark datasets from distinct satellites show that the proposed network outperforms state-of-the-art (SOTA) approaches in terms of classification accuracy.
Zenghui Zhang, Tao Zhang 0027, Weiwei Guo, Ying Luo 0001
IEEE Geosci. Remote. Sens. Lett.3
2022 LD-Net: A Lightweight Network for Real-Time Self-Supervised Monocular Depth Estimation
abstract
Self-supervised monocular depth estimation from video sequences is promising for 3D environments perception. However, most existing methods use complicated depth networks to realize monocular depth estimation, which are often difficultly applied to resource-constrained devices. To solve this problem, in this letter, we propose a novel encoder-decoder-based lightweight depth network (LD-Net). Briefly speaking, the encoder is composed of six efficient downsampling units and the Atrous Spatial Pyramid Pooling (ASPP) module. The decoder consists of some novel upsampling units that adopt the sub-pixel convolutional layer (SP). Experiments tested on the KITTI dataset show that the proposed LD-Net can reach nearly 150 frames per second (FPS) on GPU, and remarkably decreases the model parameters while maintaining competitive accuracy compared with other state-of-the-art self-supervised monocular depth estimation methods.
Mingkang Xiong, Zhenghong Zhang, Tao Zhang 0027, Huilin Xiong
IEEE Signal Process. Lett.3
2022 Polarization Estimation With Vector Sensor Array in the Underdetermined Case
abstract
Ship target detection using radar is an important application in military and civilian fields. For the polarization estimation of scattering waves in the underdetermined case, i.e., the number of scattering waves from ships is larger than the number of sensors, this paper proposes two estimation methods with different measurement models. 1) For the single-vector-sensor model, this paper proposes thepolarization-invariantESPRIT-based method. This method can estimate the polarization of signals containing target echo, interference, and noise, which can cure the problem that the accuracy of existing method is poor under low interference signal ratio. 2) For the multi-vector-sensor model, this paper proposes an improved ESPRIT method based on thespatial-invariant and time-invariantsimultaneously, which can increase the degree of freedom without increasing hardware cost. As for another problem of multi-vector-sensor, i.e., almost all existing methods assume that the number of scattering waves is known, this paper introducesthe polarization spectrumfor the first time, which can estimate the polarization parameters when the number of scattering waves is unknown. Finally, we analyze the two proposed ESPRIT-based methods comparing with some existing methods through Monte Carlo simulation, which results demonstrate the efficience of the proposed methods.
Yaomin He, Tao Zhang 0027, Huafeng He, Jian Yang 0011
IEEE Trans. Geosci. Remote. Sens.2
2022 GPU-Oriented Designs of Constant False Alarm Rate Detectors for Fast Target Detection in Radar Images
abstract
Constant false alarm rate (CFAR) detector is a class of widely used methods for target detection in radar images. Classical CFAR detectors perform target detection on a pixel-by-pixel basis using certain sliding windows for estimating clutter statistics, which run fast for small images. However, as the image size gets large, the time cost of these detectors will increase significantly since the time complexity with respect toN×N-pixel image isO(N2). In practice, radar images, such as those in synthetic aperture radar (SAR), usually have very large numbers of pixels (which can be on the order of 10000 × 10000), making the classical CFAR detectors very time-consuming when applied to these images. In this paper, we present graphics processing unit (GPU)-oriented Designs for speeding up CFAR detectors, including smallest/greatest-of CFAR and order-statistic CFAR. The proposed designs implement CFAR detectors via tensor operations, including tensor convolution, shift, and boolean operation, which can be fast operated by GPU. Experiment results show that the proposed GPU-oriented CFAR detectors running on a high-performance Nvidia RTX 3090 GPU can be thousands of times faster than the classical CFAR detectors, and realize real-time target detection in large-size radar images. Examples using SAR and range-Doppler images are provided as illustrative applications of the proposed GPU CFAR detectors to target detection in radar images.
Huizhang Yang, Tao Zhang 0027, Yaomin He, Yihua Dan, Junjun Yin 0001, Benteng Ma, Jian Yang 0011
IEEE Trans. Geosci. Remote. Sens.2
2022 Two-Dimensional Spectral Analysis Filter for Removal of LFM Radar Interference in Spaceborne SAR Imagery
abstract
Radio spectrum bands allocated to spaceborne synthetic aperture radar (SAR) imagery are shared by multiple missions. In practical radio spectrum environments, these bands are also used by some ground radars, e.g., C-band weather radar. Due to this fact, radio frequency interference (RFI) may occur for a spaceborne SAR when its received signals contain the transmitted waveforms from another SAR or radar operating at the same frequency band. This particular class of RFI is usually linear-frequency-modulation (LFM) signals, which can cause bright radiometric artifacts in focused SAR images. Most existing signal processing approaches designed for addressing this problem belong to the class of preprocessing methods, which removes RFI in level-0 raw radar data before SAR focusing. In this article, we propose a postprocessing kernel—2-D SPECtral ANalysis (2-D SPECAN) filter, for removing the class of LFM RFI in level-1 SLC images. The filtering consists of three main steps: Step 1: focus LFM RFI artifacts in SLC images as point-like responses in the spectral domain via 2-D SPECAN; Step 2: perform 2-D notch filtering in the spectral domain to remove the most contribution of the RFI responses; and Step 3: transform the filtered spectrum back into the SLC image domain using the inverse operation of the 2-D SPECAN. For computation efficiency, we design a simplified processing flow and adopt a blockwise processing strategy. Experiments with several Sentinel-1 SLC images demonstrate that severe RFI artifacts in SLC images can be removed significantly by the proposed method.
Huizhang Yang, Yaomin He, Yanlei Du, Tao Zhang 0027, Junjun Yin 0001, Jian Yang 0011
IEEE Trans. Geosci. Remote. Sens.4
2022 A Two-Stage Method for Ship Detection Using PolSAR Image
abstract
Ship detection using polarimetric SAR (PolSAR) images has recently been an active topic in the Earth observation field. There, how to detect small ships is an open and challenging issue. Within this context, we put forward a two-stage ship detection model, by which a novel ship detection method is proposed as well. Briefly, in the first stage, a suppression manipulation is adopted to suppress sea clutter, where the feature SVVSOis built on the intensity information with the orientation angle compensation (OAC). In the second stage, an enhancement manipulation is further executed to highlight ships from the suppressed sea clutter, where the features PID (polarimetric intensity difference) and NsD (nonsurface degree) are first constructed with SVVSOand a series of theoretical derivations. Then, via fusing PID and NsD together, the two-stage-based method FPAN is proposed to detect ships. To demonstrate its performance, we apply FPAN to four different L-Band PolSAR datasets. Experimental results reveal that, compared to other state-of-the-art methods, especially the DBSPCPmethod, FPAN is more effective in detecting small ships. On average, its figure-of-merit (FoM) and target-to-clutter ratio (TCR) values are, respectively 9.40% and 25.18% greater than those of DBSPCP, while the time consumption is just 58.67% of the latter.
Tao Zhang 0027, Sinong Quan, Zhen Yang 0012, Weiwei Guo, Zenghui Zhang, Hongping Gan
IEEE Trans. Geosci. Remote. Sens.1
2022 Corrections to "Region-Based Polarimetric Covariance Difference Matrix for PolSAR Ship Detection"
abstract
In the above article[1], the average TCR values inTable IIwere incorrectly presented. The corrected table is given here:
Tao Zhang 0027, Wei Wang 0099, Sinong Quan, Huizhang Yang, Huilin Xiong, Zenghui Zhang, Wenxian Yu
IEEE Trans. Geosci. Remote. Sens.1
2022 Region-Based Polarimetric Covariance Difference Matrix for PolSAR Ship Detection
abstract
To more effectively detect small ships, in this article, a novel region-based polarimetric covariance difference matrix [RP] is put forward, which mainly consists of two stages. Briefly speaking, in the first stage, a new pixel representation way is proposed to depict the spatial characteristics of pixel, through which the difference information related to pixel’s local region is calculated as well. In the second stage, the global region difference information of pixel is computed. Finally, we construct [RP] via fusing these two different kinds of information together with a balance factor$c$. Meanwhile, considering that the backscattering energy of ships is useful for ship detection, a new intensity-driven polarimetric notch filter (ID-PNFRP) is also derived from [RP]. Three different datasets are adopted to evaluate the effectiveness of [RP] and ID-PNFRP. Experimental results show that: 1) compared with the polarimetric covariance matrix [$C$] and the polarimetric covariance difference matrix [$P$], [RP] is more suitable for ship detection and 2) compared with the original geometrical perturbation-polarimetric notch filter (GP-PNF) and the total power detector SPAN, the proposed method ID-PNFRPcan better detect small ships with greater figure of merit (FoM) and target-to-clutter ratio (TCR) values.
Tao Zhang 0027, Wei Wang 0099, Sinong Quan, Huizhang Yang, Huilin Xiong, Zenghui Zhang, Wenxian Yu
IEEE Trans. Geosci. Remote. Sens.1
2022 A Feature Decomposition-Based Method for Automatic Ship Detection Crossing Different Satellite SAR Images
abstract
In the face of Synthetic Aperture Radar (SAR) image object detection with different distributions of training and test data, traditional supervised learning methods cannot achieve good detection performance. Domain adaptation (DA) method has been shown to have the ability to solve this problem, but existing DA object detection algorithms all use adversarial DA theory for the detection task, which is ineffective in solving object regression localization in the detection task. In this article, to better solve the above problem, an automatic SAR image ship detection method based on feature decomposition crossing different satellites is proposed. The feature extraction layer of backbone network is divided into low level and high level, where domain-invariant feature extractors are designed for the local features extracted from the low level and the global features extracted from the high level, respectively. We argue that the local and global features extracted from source domain and target domain contain domain-specific features (DSF) for adversarial DA and domain-invariant features (DIF) that contribute to object regression localization. Then, we decompose the local features and global features into DSF and DIF via vector decomposition method. For DSF counterpart, we introduce adversarial DA attention for feature alignment. DIF from the local features are fused into the backbone network for high-level global feature extraction. Finally, by using region proposal network and adversarial domain classifier, we can get the accurate bounding box and object class of SAR image objects. Extensive experiments prove that the proposed method outperforms state-of-the-art methods in terms of detection performance.
Ying Luo 0001, Tao Zhang 0027, Weiwei Guo, Zenghui Zhang
IEEE Trans. Geosci. Remote. Sens.3
2022 TransCS: A Transformer-Based Hybrid Architecture for Image Compressed Sensing
abstract
Well-known compressed sensing (CS) is widely used in image acquisition and reconstruction. However, accurately reconstructing images from measurements at low sampling rates remains a considerable challenge. In this paper, we propose a novel Transformer-based hybrid architecture (dubbed TransCS) to achieve high-quality image CS. In the sampling module, TransCS adopts a trainable sensing matrix strategy that gains better image reconstruction by learning the structural information from the training images. In the reconstruction module, inspired by the powerful long-distance dependence modelling capacity of the Transformer, a customized iterative shrinkage-thresholding algorithm (ISTA)-based Transformer backbone that iteratively works with gradient descent and soft threshold operation is designed to model the global dependency among image subblocks. Moreover, the auxiliary convolutional neural network (CNN) is introduced to capture the local features of images. Therefore, the proposed hybrid architecture that integrates the customized ISTA-based Transformer backbone with CNN can gain high-performance reconstruction for image compressed sensing. The experimental results demonstrate that our proposed TransCS obtains superior reconstruction quality and noise robustness on several public benchmark datasets compared with other state-of-the-art methods. Our code is available on TransCS.
Minghe Shen, Hongping Gan, Chao Ning 0003, Tao Zhang 0027
IEEE Trans. Image Process.5
2021 Fine-Grained Classification of Neutrophils with Hybrid Loss
Qingtao Zhu, Danwei Lu, Tao Zhang 0027, Junjun Yin 0001, Jian Yang 0011
ICIG (1)3
2021 Effects of Ocean Wave Spectrum Truncation on Sea Clutter Distribution in Numerical Simulations
abstract
The effects of ocean spectrum truncation on the sea clutter distribution properties in numerical simulations are studied using a recently developed full-wave method, i.e., the multilevel steepest decent - sparse matrix canonical grid (MLSD-SMCG) method and the KHCC03 spectrum. Two types of ocean surface profiles are generated for Monte Carlo simulations based on the full and truncated spectra at the wind speed of 10 m/s. The surface profiles generated by the truncated spectrum have lengths about 1/6 of those using full spectrum. 1000 realizations are conducted for each type of profiles. The simulations are illustrated at L-band (1.4 GHz) and the incidence angle is 40°. For the simulated far-field scattering fields and normalized radar cross sections (NRCS), we use the K-distribution model to fit the probability density functions (PDF) of the amplitude and backscatter of clutters. It is found that spectrum truncation has non-negligible effects on the distribution characteristics of sea clutter in the numerical simulations, particularly for the amplitude distributions. The fitted PDF indicates that the simulated sea clutter using truncated spectrum has more small values compared with that using full spectrum.
Yanlei Du, Jian Yang 0011, Tao Liu 0025, Tao Zhang 0027, Xiaofeng Yang 0002
IGARSS5
2021 Simplified Power-Based Detectors for Ship Detection of PolSAR Imagery
abstract
Ship detection of polarimetric SAR (PolSAR) imagery has attracted lots of attentions in recent years. Also, it is known that among the polarimetric channels$HH, HV$, and$VV, VV$is the most sensitive to sea clutter. Following this guidance, in this paper, a novel ship detector SVVS is first proposed via subtracting the term$C_{33}$from the total power detector SPAN. And then, the complect polarimetric covariance difference matrix [$CP$] is utilized to calculate SVVS, leading to the construction of another novel ship detector$\text{SVVS}_{CP}$. Finally, we investigate the statistical distribution of sea clutter with$\text{SVVS}_{CP}$and further develop an adaptive$\text{SVVS}_{CP}$-based C-FAR detector for ship detection. The experiment carried out on one real PolSAR imagery shows that, compared to SPAN, both SVVS and$\text{SVVS}_{CP}$hold better ship detection performances.
Tao Zhang 0027, Hongping Gan, Zhen Yang 0012, Bing Zeng 0001, Jian Yang 0011
IGARSS1
2021 A Superpixel-Based Neighborhood Polarimetric Covariance Matrix for Polsar Ship Detection
abstract
In a recent work, a neighborhood polarimetric covariance matrix [N] was proposed to detect ships from polarimetric SAR (PolSAR) imagery. However, its computational complexity is extremely high. Besides, the heterogeneity surrounding ship edges is also not well considered in [N]. To cure these draw-backs, we construct a novel superpixels-based neighborhood polarimetric covariance matrix [SN] in this paper. Specifically, the simple linear iterative clustering (SLIC) is first used to obtain superpixels. Then, the vector vmeancorresponding to the mean value of superpixel is further computed so as to characterize the neighborhood information of each pixel in superpixel. Finally, by combining the original scattering vector v and vmeantogether, the vector t12is built to calculate [SN]. The experiment tested on one L-Band ALOS PolSAR imagery shows that i) the polarimetric whitening filter derived from [SN] (i.e., PWFSN) has a better detection performance than that derived from [N] (i.e., PWFN); ii) the calculation process of [SN] takes much less time than that of [N].
Tao Zhang 0027, Chengtao Ji, Yanlei Du, Tao Liu 0025, Jian Yang 0011
IGARSS1
2021 Multi-level dictionary learning for fine-grained images categorization with attention model
Jinsheng Ji, Yiyou Guo, Zhen Yang 0012, Tao Zhang 0027, Xiankai Lu
Neurocomputing4
2021 SWS-DAN: Subtler WS-DAN for fine-grained image classification
Zhen Yang 0012, Lingkun Luo, Hongping Gan, Tao Zhang 0027
J. Vis. Commun. Image Represent.5
2021 Ship Detection From PolSAR Imagery Using the Hybrid Polarimetric Covariance Matrix
abstract
In this letter, we first investigate the relationship between polarimetric covariance matrix [C] and complete polarimetric covariance difference matrix [CP], and then construct a scattering difference parameter SDP. Subsequently, a hybrid polarimetric covariance matrix [HC] is developed based on SDP for curing the disadvantage of [CP], that is the scattering difference information of small ships cannot be well contained in [CP]. By fusing the feature “1-SDP” and the power detector SPANHCderived from [HC] together, a novel ship detection method SPANSDPis finally proposed to detect ships. Experiments performed on the airborne SAR (AIRSAR) L-Band and GF-3 C-Band data verify that 1) SPANSDPcan detect small ships more accurately than other state-of-the-art methods and 2) [HC] is more effective in improving ship detectors' detection performances in comparison with [CP].
Tao Zhang 0027, Wei Wang 0099, Zhen Yang 0012, Junjun Yin 0001, Jian Yang 0011
IEEE Geosci. Remote. Sens. Lett.1
2021 Adversarial erasing attention for fine-grained image classification
Jinsheng Ji, Linfeng Jiang, Tao Zhang 0027, Weilin Zhong, Huilin Xiong
Multim. Tools Appl.3
2020 Ship Detection from Polsar Imagery Based on the Scattering Difference Parameter
abstract
In this paper, a new scattering difference parameter named as SDP is first constructed to characterize the relationship between polarimetric covariance matrix [C] and complete polarimetric covariance difference matrix [CP]. Then, by integrating ”1-SDP” and the power maximization synthesis detector (PMS) derived from [CP], a novel ship detection method OmSPcpis further developed. In order to demonstrate the performance of the proposed method, one AIRSAR L-Band Polarimetric SAR dataset with 22 ships is exploited. The experimental results show that, compared to other methods, OmSPcpcan hold a better ship detection accuracy.
Tao Zhang 0027, Zhen Yang 0012, Junjun Yin 0001, Jian Yang 0011
IGARSS1
2020 Combining Multilevel Features for Remote Sensing Image Scene Classification With Attention Model
abstract
Remote sensing (RS) image scene classification is a challenging task due to its intraclass variety and the interclass similarity. Recently, many convolutional neural network (CNN)-based methods explore the network to handle this task. However, RS images usually have confusing background in addition to the relevant objects, and features only derived from the whole RS images cannot achieve satisfying results. To solve the problem, this letter proposed a method of utilizing the attention network to localize multiscale discriminative regions of the RS scene images and combining features learned from the localized regions by a classification network. Specifically, the classification network is composed of three subnetworks, which are trained by certain scaled regions separately. To learn more discriminative feature representations, feature fusion module is introduced to fuse the features of the three subnetworks in a more effective way. Experiments conducted on the AID and NWPU-RESISC45 data sets evaluate the effectiveness of the proposed method.
Jinsheng Ji, Tao Zhang 0027, Linfeng Jiang, Weilin Zhong, Huilin Xiong
IEEE Geosci. Remote. Sens. Lett.2
2020 Exploiting context based on CNN and coding representations for pedestrian co-detection
Linfeng Jiang, Jinsheng Ji, Weilin Zhong, Tao Zhang 0027, Huilin Xiong
Multim. Tools Appl.4
2020 A part-based attention network for person re-identification
Weilin Zhong, Linfeng Jiang, Tao Zhang 0027, Jinsheng Ji, Huilin Xiong
Multim. Tools Appl.3
2020 PolSAR Ship Detection Using the Joint Polarimetric Information
abstract
In this article, we investigate the scattering components of ships and find that the surface scattering may be the primary scattering for some ships, especially small ships. Meanwhile, the drawbacks of the complete polarimetric covariance difference matrix [CP] are also pointed out in theory. Based on these analyses, two new methods are then constructed to detect the ships. More specifically, the first one RsP is constructed by directly combining the similarity parameter of surface scattering Rs and the power-maximization synthesis (PMS) detector. The second one RsDVH is designed by taking advantage of four different features (i.e., Rs, double-bounce scattering, volume scattering, and helix scattering), which are all derived from the joint polarimetric information that is developed by combing the information of the polarimetric covariance matrix [C] and [CP]. Subsequently, the generalized Gamma distribution (GΓD) is found suitable for characterizing the RsDVH values of the sea clutter. At last, an adaptive constant false-alarm-rate (CFAR) detector developed from RsDVH is proposed for ship detection. To verify the effectiveness of RsP and RsDVH, four polarization synthetic aperture radar (PolSAR) imageries are tested, including one L-band UAVSAR imagery with 19 ships, two L-band AIRSAR imageries with 22 and 53 ships, respectively, and one C-band GF-3 imagery with ten ships. The experimental results show that: 1) the surface scattering is beneficial to detecting ships, especially the ships with prominent surface scatterings; 2) compared with other state-of-the-art methods, RsDVH can more effectively enhance the target-to-clutter ratio (TCR) values of small ships in the case of rough sea surface; and 3) the joint polarimetric information that is put forward and exploited for the first time in this article has a greater potential to help ship detectors improve their detection performances than the traditional polarimetric information included in [C].
Tao Zhang 0027, Zhen Yang 0012, Hongping Gan, Deliang Xiang, Jian Yang 0011
IEEE Trans. Geosci. Remote. Sens.1
2019 Aircraft Detection from Remote Sensing Image Based on A Weakly Supervised Attention Model
abstract
Aircraft detection from high resolution remote sensing image is a challenging task due to the lack of annotation information, large-scale image size, and sparse distribution of aircraft. Recently, some convolutional neural network(CNN) based methods explore the attention based weakly supervised way to localize the aircraft without manual annotation information. However, the detection results are not satisfied with high false detection ratio. In this paper, a method of utilizing weakly supervised attention model to localize the multi-scale aircrafts is presented, in which the attention model is carried out in a weakly supervised way. Compared with other CNN based method, the proposed attention model can obtain more accurate attention map and localize the aircrafts more precisely. The experimental results on two challenging datasets demonstrate that the proposed method achieves higher detection accuracy and lower false detection ratio than other methods.
Jinsheng Ji, Tao Zhang 0027, Zhen Yang 0012, Linfeng Jiang, Weilin Zhong, Huilin Xiong
IGARSS2
2019 Ship Detection Using the Surface Scattering Similarity and Scattering Power
abstract
Sea surface and ship have different backscattering mechanisms, in which surface scattering is predominant for sea surface in the low sea state case. Based on this fact, many ship detectors have been developed by suppressing the surface scattering resulted from sea surface. Actually, small ship may also have strong surface scattering sometimes. In such a case, the methods of avoiding using surface scattering features may easily miss the detection of small ships. To verify this point, in this paper, we first analyze the shortcomings of An's method which is based on surface scattering similarity and the power maximization synthesis detector (PMS), and then improve it for detecting small ships more effectively. In order to demonstrate the performance of the proposed method, AIRSAR L-Band Polarimetric SAR dataset is exploited. Comparing to other methods, the new method shows a better ship detection performance.
Tao Zhang 0027, Zhen Yang 0012, Jian Yang 0011, Yifang Ban, Huilin Xiong
IGARSS1
2019 Combining multilevel feature extraction and multi-loss learning for person re-identification
Weilin Zhong, Linfeng Jiang, Tao Zhang 0027, Jinsheng Ji, Huilin Xiong
Neurocomputing3
2019 Discriminative representation learning for person re-identification via multi-loss training
Weilin Zhong, Tao Zhang 0027, Linfeng Jiang, Jinsheng Ji, Zenghui Zhang, Huilin Xiong
J. Vis. Commun. Image Represent.2
2019 Ship Detection From PolSAR Imagery Using the Complete Polarimetric Covariance Difference Matrix
abstract
In this paper, we proposed a complete polarimetric covariance difference matrix [CP]-based algorithm for ship detection in polarimetric synthetic aperture radar (PolSAR) imagery. To calculate [C P], we first developed a scheme to reflect the polarimetric scattering differences between ship pixel (SP) and its neighboring pixels (ISPs) and, then, dividedly accumulated the amplitude and phase differences between SP and ISPs. Compared to the polarimetric covariance difference matrix [P] developed in our earlier work, [C P] effectively overcomes the drawback of the lack of the phase information in [P]. To demonstrate the effectiveness of the proposed algorithm, we applied the [CP]-based ship detection algorithm to four PolSAR data sets, including one UAVSAR L-band data set with 21 ships, two AIRSAR L-band data sets with 11 and 22 ships, respectively, and one Radarsat-2 C-band data set with 8 ships. Experimental results show that: (1) the proposed algorithm can effectively detect ships with high target-to-clutter ratio (TCR) values and (2) [C P] has a better performance than the traditional polarimetric covariance matrix [C] and [P] on ship detection. To be more specific, the average TCR value of the proposed algorithm (23.86 dB) is 6.07 and 7.47 dB higher than PNFC(i.e., the geometrical perturbation-polarimetric notch filter) and RSC(i.e., the reflection symmetry method), respectively.
Tao Zhang 0027, Jinsheng Ji, Xiaofeng Li 0001, Wenxian Yu, Huilin Xiong
IEEE Trans. Geosci. Remote. Sens.1
2018 A Multi-part Convolutional Attention Network for Fine-Grained Image Recognition
abstract
The goal of fine-grained image recognition is to recognize hundreds of sub-categories affiliating to the same basic-level category (e.g., bird species). It is a highly challenging task due to the large intra-class variance and small inter-class variance. Existing approaches deal with the subtle difference among object classes via learning and localizing discriminative parts. However, most of the part localization methods follow a step-to-step manner that first localizes larger parts and then generates smaller parts from the larger ones, which is not efficient. In this paper, we present a Multi-part Convolutional Attention Network (M-CAN), which simultaneously focuses on the discriminative image parts at multiple scales. In specific, a convolutional attention based part localization network is presented to localize multi-scale parts from different layers of the deep Convolutional Neural Networks (CNN). Importantly, our part localization network requires no part annotations but only the image labels, which avoids the heavy labor of complex part labeling. We conduct comprehensive experiments and the experimental results show that, our method outperforms the state-of-the-art approaches on three challenging fine-grained datasets, including CUB-Birds, Stanford-Dogs and Stanford-Cars.
Weilin Zhong, Linfeng Jiang, Tao Zhang 0027, Jinsheng Ji, Huilin Xiong
ICPR3
2018 A Ship Detector Based on the Improved Polarimetric Covariance Difference Matrix
abstract
Polarimetric Synthetic Aperture Radar data has been widely used for ship detection. In our earlier study, based on the differences between ship pixels and their surrounding background pixels, we designed a polarimetric covariance difference matrix (PCDM) to detect ships. Inadequately, the phase information of scattering differences is not included in PCD-M. Aiming at this deficiency, here, we present an improved PCDM matrix (IPCDM). Then an IPCDM-based ship detector is further proposed. To demonstrate the effectiveness of the method, two full polarimetric datasets are adopted. In comparing with other methods, we find that the result of our method is better.
Tao Zhang 0027, Yifang Ban, Huilin Xiong, Wenxian Yu
IGARSS1
2018 Rotated Region Based Fully Convolutional Network for Ship Detection
abstract
Ship detection from high-resolution optical remote sensing images has been a prevalent domain in recent years. Unlike objects in natural images, ships of interest can be anywhere in optical remote sensing images with multi-scale and multi-oriented which makes it more different to be detected. In this paper, we propose a novel method based on the fully convolutional network to detect ships. Our method has three important components: 1) we design a network merging different levels of feature map to fuse multi-scale information. Determining the existence of large ship require features from deep layers in the network, while predicting rotated bounding box enclosing small ships need shallow layers information; 2) The network can be trained end-to-end to generate score maps which indicates the confidence score for the ship region of interest in pixel-wise level through all locations and scaled of an image; 3) We design a rotated bounding box regression model to localize the ships. The experimental results on our dataset collected from Google Earth has demonstrated our proposed method achieves promising performance on ship detection in terms of both efficiency and accuracy in high-resolution optical remote sensing images.
Mingjie Li 0006, Weiwei Guo, Zenghui Zhang, Wenxian Yu, Tao Zhang 0027
IGARSS5
2017 Multi-part compact bilinear CNN for person re-identification
abstract
In paper, we present a novel multi-part compact bilinear convolutional neural network (CNN) model, which consists of a bilinear CNN and two part-networks aiming to learn the global features and the finer local features simultaneously. The bilinear operation is simplified with recently proposed compact bilinear pooling method, and bilinear vectors are averagely pooled to keep more local spatial information. The proposed model is trained by using a histogram loss function in order to reduce the distribution overlap of positive pairs and negative pairs. Experiments show that, the combination of compact bilinear CNN and histogram loss can significantly improve the original models, and performs favorably compared to the state of the art.
Zhen Yang 0012, Tao Zhang 0027, Huilin Xiong
ICIP3
2017 Bi-directional long short-term memory architecture for person re-identification with modified triplet embedding
abstract
Matching a specific person across non-overlapping cameras, known as person re-identification, is an important yet challenging task owing to the intra-class variations of the images from the same person in pose, illumination, and occlusion. Most existing body-parts based deep methods simply concatenate the features or scores obtained from spatial parts and ignore the complex spatial correlation between them. In this paper, we present a bi-directional Long Short-Term Memory (Bi-LSTM) architecture that can process the spatial parts sequentially, and enable the messages of different parts to go through in a bi-directional manner. Therefore, the spatial and contextual visual information can be modeled efficiently by the bi-directional connections and the internal gating function in LSTM. Furthermore, we propose a modified triplet loss to learn more discriminative features to distinguish positive pairs from negative pairs. Experiments on CUHK01 and CUHK03 datasets are carried out to demonstrate the effectiveness of the proposed method.
Weilin Zhong, Huilin Xiong, Zhen Yang 0012, Tao Zhang 0027
ICIP4
2017 A ship detector applying principal component analysis to the Polarimetric Notch Filter
abstract
In this paper, a new algorithm for ship detection with Synthetic Aperture Radar (SAR) images is presented. We develop the proposed method by combing Principal Component Analysis (PCA) and the Geometrical Perturbation-Polarimetric Notch Filter (GP-PNF) method. In the first step, we replace the feature vector composed by the elements of the covariance matrix with more polarimetric features. Then, PCA is used to reduce the feature space. The new reduced feature vector is then used to detect ships by using the framework of the GP-PNF. In order to demonstrate the effectiveness of the proposed method, we exploited Sentinel-1 datasets. In this abstract, a dataset obtained in Gibraltar is considered. A comparison with other methods showed improvements in detection capability.
Tao Zhang 0027, Armando Marino, Huilin Xiong
IGARSS1
2017 An Azimuth ambiguities removal method based on Polarimetric Notch Filter
abstract
In this paper, a new algorithm for detecting ship and removing azimuth ambiguities is presented. The proposed method is developed by combing the third eigenvalue and the Geometrical Perturbation-Polarimetric Notch Filter (GP-PNF) methods. We firstly improve the GP-PNF feature vector with the third eigenvalue calculated by the eigenvalues-eigenvector decomposition method. Then, the new feature vector is used to remove azimuth ambiguities in the framework of the GP-PNF method. To demonstrate the effectiveness of the proposed method, we exploited one AIRSAR C-band dataset here. In comparing with the traditional GP-PNF method, we find our method has a better capability in removing azimuth ambiguities and detect real ships.
Tao Zhang 0027, Armando Marino, Weilin Zhong, Huilin Xiong
IGARSS1
2016 Ship detection based on the power of the Radarsat-2 polarimetric data
abstract
Ship target detection using PolSAR data has been an active research area and many algorithms have been developed in recent years. In this paper, we present a new method based on the difference between the ship pixels and background pixels, using a model similar to LBP (Local Binary Pattern). After that, the polarimetric signature method, namely, the SPAN (total power) detector, is used to detect ships. We adopt one Radarsat-2 data set with four-look processing for experiment, which was obtained in the Strait of Gibraltar ocean area. In comparing with other methods, we find that the result of our method is better than other detectors.
Tao Zhang 0027, Zhen Yang 0012, Huilin Xiong, Wenxian Yu
IGARSS1