Hao Sun 0042

dblp:82/2248-42 · DBLP profile ↗
← Back
19ranked-venue papers
3as first author
11since 2021 · last 2025
0009-0006-3450-8103ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 13 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2025 SAR-TinySNN: A Lightweight Spiking Neural Network for SAR Target Recognition
Hao Sun 0042, Yuli Sun, Tao Tang 0006, Lin Lei, Kefeng Ji
IEEE Geosci. Remote. Sens. Lett.2
2025 Temporal-Spatial Feature Interaction Network for Multi-Drone Multi-Object Tracking
abstract
Multi-drone multi-object tracking (MDMOT) aims to localize and identify targets from videos captured simultaneously by multiple drones. To accomplish this task, existing methods typically follow the strategy of associating localized targets to obtain identities. However, their localization and identification stages heavily rely on single-frame information, resulting in the localization being very sensitive to visual information decay and making it struggle to capture discriminative representations for target identification. Consequently, they usually exhibit unreliable performance in challenging scenarios, such as occlusion and high similarity among targets. To this end, we introduce a novel MDMOT framework to interact temporal-spatial features, exploring the guidance of tracklet information across time and space. Specifically, we introduce temporal-spatial feedback loops to enrich cues in our tracker. Meanwhile, a novel temporal-oriented target localization is proposed to enhance the response to difficult samples in feature space by utilizing prior knowledge from existing tracklets beyond the current frame for target localization. Moreover, a spatial-oriented target identification is designed to synergize cross-drone information of tracklets, thereby providing discriminative representations for target identification. It combines target and background information to extract identity representations and interacts features from multiple drones. To our best knowledge, this work reports the first MDMOT system that synergizes features across multiple drones to track targets. By incorporating these two elaborated networks, we develop a robust tracker (named TSMMT). Extensive experiments on the MDMT public dataset demonstrate the superiority of our proposed model. Specifically, TSMMT outperforms state-of-the-art methods by 2.76%~4.66% on MOTA and 2.06%~3.33% on IDF1.
Hao Sun 0042, Kefeng Ji, Gangyao Kuang
IEEE Trans. Circuits Syst. Video Technol.2
2025 DiffDual-AD: Diffusion-Based Dual-Stage Adversarial Defense Framework in Remote Sensing With Denoiser Constraint
abstract
Deep neural networks (DNNs), though highly effective in various Earth observation tasks with remote sensing images (RSIs), are vulnerable to adversarial attacks, threatening their reliability. Each maliciously attacked RSI potentially contains unique and critical information, but current defenses lack a unified framework for rapidly detecting adversarial RSIs and accurately restoring them to their natural state. To bridge this research gap, we propose a diffusion-based dual-stage adversarial defense (DiffDual-AD) framework. In the first stage, we propose a novel adversarial detection method based on the score expectation (AD-SE), which is integrated into the forward process of diffusion model with an improved denoiser constrained for adversarial defense. During the reverse process, the second stage introduces the distance and label-guided adversarial purification (DL-GAP) to restore adversarial RSIs to natural ones. With the help of two proposed guidelines and the specific denoiser, the DL-GAP effectively smooths out the adversarial perturbations from the detected adversarial RSIs while preserving semantic information and local key features. Finally, DL-GAP yields nonadversarial purified RSIs based on the results of AD-SE. Both stages are integrated within the diffusion models, complementing each other to form a pipelined operational mode. Extensive experiments across three RSI scene classification datasets have proven its efficacy in resisting adversarial RSIs in both whitebox and black-box scenarios, achieving an average adversarial detection accuracy of 87.02% with AD-SE and an average final classification accuracy of 81.50% with the entire DiffDual-AD. The performances of our proposed AD-SE and DL-GAP have surpassed other advanced adversarial detection and purification methods when applied to RSIs. Therefore, DiffDual-AD provides a unified and universal solution to preserve the utility and security of processing the adversarial RSIs, which advances the adversarial defense research in remote sensing.
Zihao Lu, Hao Sun 0042, Lin Lei, Yuli Sun, Gangyao Kuang
IEEE Trans. Geosci. Remote. Sens.2
2024 A Hybrid Feature Selection Method Based on Imbalanced Learning for Wave Prediction
abstract
Wave data mining and processing are important in ocean prediction. However, wave data often exhibits imbalance, resulting in low accuracy in predicting extreme phenomena. To alleviate this problem, we propose a hybrid feature selection method based on imbalance learning (HFS-IL) to improve accuracy of prediction models. Specifically, we first use a Long Short-Term Memory (LSTM) network to train an imbalance discriminator, which aims to classify input data into common and rare subsets. Secondly, we select the optimal feature subsets by a hybrid feature selection algorithm, which is innovatively designed by combining mutual information and forward selection. To verify the effectiveness of HFS-IL, we process an imbalanced wave dataset from ERA5 by HFS-IL and use the processed data as input for an intelligent prediction model. The experimental results demonstrate that HFS-IL can effectively alleviate the impact of data imbalance and improve the accuracy of prediction, especially at station 51000, the majority of metrics outperform GRU and OSP-FEAN.
Qinjie Lin, Xiaoli Ren, Hao Sun 0042, Jiaming Tan, Xiaoyong Li 0002, Jingze Lu
ISPA3
2024 TirSA: A Three Stage Approach for UAV-Satellite Cross-View Geo-Localization Based on Self-Supervised Feature Enhancement
abstract
Cross-view geo-localization aims to associate geographical location with different view images shot from different platforms. One of the critical challenges is how to effectively emphasize architectural features and reducing background interference to achieve robust cross-view matching. Most of the existing methods fail to adequately address the features of buildings, treating foreground and background equally. Leveraging prior knowledge to enhance the features of crucial architectural foreground yields greater benefits in Geo-Localization. A comprehensive three stage approach (TirSA) is proposed in this paper, which consists of three components: Pre-processing, Generate Feature Embedding, and Post-processing. In the Pre-processing stage, we employ a self-supervised feature enhancement method (SFEM) to obtain the building aware mask. Without adding additional auxiliary information, the model is guided to learn from discriminative building regions. Besides, in the Generate Feature Embedding stage, we propose an adaptive feature integration module (AFIM) to enhance feature representation capability. We also train the Siamese network using a novel improved cross-domain triplet loss to reduce the impact of inter-view domain gap. Finally, in the Post-processing stage, we employ a re-ranking method to optimize the initial retrieval list, further enhancing the matching accuracy. Remarkably, extensive experiments show that our proposed TirSA exceeds state-of-the-art by a large margin and achieves optimality in both drone-view target localization and drone navigation. Especially in the drone navigation task, our method is superior to the existing methods, achieving an improvement of approximately 5%. Code will be released at https://github.com/SunJ1025/TirSA.
Jian Sun 0038, Hao Sun 0042, Lin Lei, Kefeng Ji, Gangyao Kuang
IEEE Trans. Circuits Syst. Video Technol.2
2024 Simulated Data Feature Guided Evolution and Distillation for Incremental SAR ATR
abstract
Deep neural network (DNN)-based synthetic aperture radar automatic target recognition (SAR ATR) methods have made great progress in recent years. However, the performance of DNN models relies on a large number of independent and identically distributed measured synthetic aperture radar (SAR) images, which is contrary to the SAR ATR in practice. Furthermore, DNN models also suffer from catastrophic forgetting when learning a sequence of new classes. To tackle these problems, we introduce simulated data into the class incremental learning of SAR ATR for the first time. Specifically, we aim to continuously learn a sequence of new classes with a small amount of measured data and a large amount of simulated data. We first investigate the properties of incremental learning using simulated data, and the main observation is that simulated data can achieve good performance in short-term incremental learning rather than long-term incremental learning. A novel class incremental learning method, namely, feature guided evolution and distillation (FGED), is then presented. On the one hand, FGED encourages simulated data to have the same feature relationship structure as the corresponding measured data to reduce their distribution discrepancy in short-term incremental learning. On the other hand, FGED adopts a feature distillation strategy to simultaneously reduce the distribution discrepancy accumulation of previous incremental classes and alleviate the catastrophic forgetting in long-term incremental learning. The experimental results obtained on the MSTAR benchmark dataset and two simulated datasets demonstrate the effectiveness of FGED.
Hao Sun 0042, Yan Zhao 0026, Qishan He, Siqian Zhang, Gangyao Kuang
IEEE Trans. Geosci. Remote. Sens.2
2023 Attributed Scattering Center Guided Adversarial Attack for DCNN SAR Target Recognition
abstract
Recently, deep learning has made significant progress in synthetic aperture radar automatic target recognition (SAR ATR). However, deep convolutional neural networks (DCNNs) are discovered to be susceptible to carefully crafted adversarial perturbations. Regarding the unique scattering mechanism in SAR imaging, the scattering feature such as attributed scattering centers (ASCs) should be deeply considered in the adversarial attack (AA) algorithms for DCNNs in SAR ATR. In this letter, an AA algorithm named ASC-STA is proposed to take advantage of the powerful SAR imaging property characterization capability of the ASC model. Considering the imaging characteristic that is presented in the ASC model, the spatial transformation module is specially designed to be performed on the strong backscattering structures guided by the reconstruction image of the ASC model. Spatial transformation shifts the robust scattering point features via altering the locally gray-scale relationships of the target microstructures in SAR images. Besides, a modified shape context evaluated metric is established to assess the validity of the feature alternation process. Experimental results on the moving and stationary target acquisition and recognition (MSTAR) database demonstrate that the proposed method achieves a high success rate without complex parameter settings and generates the perturbed image in high image quality. Compared with the norm-based AA algorithms, which are not involved with ASCs, the point feature of the original image is efficiently altered. At the same time, the perturbations are focused on the target regions accurately.
Junfan Zhou, Sijia Feng, Hao Sun 0042, Linbin Zhang, Gangyao Kuang
IEEE Geosci. Remote. Sens. Lett.3
2022 Fine-grained visual classification via multilayer bilinear pooling with object localization
Ming Li 0066, Lin Lei, Hao Sun 0042, Xiao Li 0017, Gangyao Kuang
Vis. Comput.3
2021 Adversarial Robustness Evaluation of Deep Convolutional Neural Network Based SAR ATR Algorithm
abstract
Robustness, both to accident and to malevolent perturbations, is a crucial determinant of the successful deployment of deep convolutional neural network based SAR ATR systems in various security-sensitive applications. This paper performs a detailed adversarial robustness evaluation of deep convolutional neural network based SAR ATR models across two public available SAR target recognition datasets. For each model, seven different adversarial perturbations, ranging from gradient based optimization to self-supervised feature distortion, are generated for each testing image. Besides adversarial average recognition accuracy, feature attribution techniques have also been adopted to analyze the feature diffusion effect of adversarial attacks, which promotes the understanding of vulnerability of deep learning models.
Hao Sun 0042, Gangyao Kuang
IGARSS1
2021 Robust Remote Sensing Scene Classification by Adversarial Self-Supervised Learning
abstract
Adversarial training is an effective method to enhance adversarial robustness for deep neural networks. However, it qequires large amounts of labeled data, which are often difficult to acquire. Recent research has shown that self-supervised learning can help to improve model performance and model uncertainty using unlabeled data. In this paper, we introduce a new adversarial self-supervised learning framework to learn a robust pretrained model for remote sensing scene classification. The proposed method exploits the advantage of dual network structure, and it requires neither labeled data for adversarial example generation nor negative samples for contrastive learning. Specifically, it consists of three major steps. Firstly, we train the online model and the target model to extract deep image features. Secondly, we generate two kinds of instance-wise adversarial examples. Finally, we iteratively learn a robust model by implicit comparing the difference between clean data and their perturbed counterpart. Preliminary experimental results on remote sensing scene classification dataset shows that our method can obtain higher robust accuracy. Our method can also be combined with other adversarial defense techniques to further promote model robustness.
Hao Sun 0042, Lin Lei, Gangyao Kuang, Kefeng Ji
IGARSS2
2021 Nonlocal patch similarity based heterogeneous remote sensing change detection
Yuli Sun, Lin Lei, Xiao Li 0017, Hao Sun 0042, Gangyao Kuang
Pattern Recognit.4
2019 Learning Deep Ship Detector in SAR Images From Scratch
abstract
Recently, deep learning-based methods have brought new ideas for ship detection in synthetic aperture radar (SAR) images. However, several challenges still exist: 1) deep models contain millions of parameters, whereas the available annotated samples are not sufficient in number for training. Therefore, most deep detectors have to fine-tune networks pre-trained on ImageNet, which incurs learning bias due to the huge domain mismatch between SAR images and ImageNet images. Furthermore, it has a little flexibility to redesign the network structure; and 2) ships in SAR images are relatively small in size and densely clustered, whereas most deep detectors have poor performance with small objects due to the rough feature map used for detection and the extreme foreground–background imbalance. To address these problems, this paper proposes an effective approach to learn deep ship detector from scratch. First, we design a condensed backbone network, which consists of several dense blocks. Hence, earlier layers can receive additional supervision from the objective function through the dense connections, which makes it easy to train. In addition, feature reuse strategy is adopted to make it highly parameter efficient. Therefore, the backbone network could be freely designed and effectively trained from scratch without using a large amount of annotated samples. Second, we improve the cross-entropy loss to address the foreground–background imbalance and predict multi-scale ship proposals from several intermediate layers to improve the recall rate. Then, position-sensitive score maps are adopted to encode position information into each ship proposal for discrimination. The comparison results on the Sentinel-1 data set show that: 1) learning ship detector from scratch achieved better performance than ImageNet pre-trained model-based detectors and 2) our method is more effective than existing algorithms for detecting the small and densely clustered ships.
Zhipeng Deng, Hao Sun 0042, Shilin Zhou 0001, Juanping Zhao
IEEE Trans. Geosci. Remote. Sens.2
2018 CNN-Based Target Detection in Hyperspectral Imagery
abstract
This paper proposes a hyperspectral target detection framework with convolutional neural network (CNN). The number of training samples is first sufficiently enlarged by subtraction method to maximize the advantages of the multilayer CNN. Next, the CNN is given a target detection function by labelling the new pixels subtracted between target and background classes as 1, and the pixels subtracted between pixels within both the same and different background classes as 0. Finally, for each testing pixel, the difference between the central pixel and its adjacent pixels is input into the framework. If the testing pixel belongs to the target, the output score is close to the target label. Aircrafts and vehicles are selected as targets of interest in the experiment conducted to validate the proposed method. The experiment results show that the proposed method has an advantage over classic hyperspectral target detection algorithms in terms of precision and robustness.
Zhiyong Li 0008, Hao Sun 0042
IGARSS3
2017 Fast multiclass object detection in optical remote sensing images using region based convolutional neural networks
abstract
Fast multiclass object detection for remote sensing images plays an important role for a wide range of applications. Traditional methods based on a sliding window search lead to heavy computational costs and are unsuitable for multiclass detection. Recently, deep learning algorithms, especially faster region based convolutional neural networks (Faster R-CNN), which adopt a region proposal paradigm to avoid exhaustive search, has achieved state-of-the-art multiclass detection performance in computer vision. This paper investigates the use of Faster R-CNN in the earth observation community. We have three contributions: 1) It's the first time to successfully use Faster R-CNN for object detection in remote sensing images. It achieved faster speed (22 ×faster) and better performance (a mAP of 78% vs. 72%) than traditional methods; 2) we adopt data augmentation to train Faster R-CNN with limited samples; 3) we successfully tested our method on large-scale google earth images, which shows robustness of our method.
Zhipeng Deng, Hao Sun 0042, Shilin Zhou 0001, Juanping Zhao, Lin Lei, Huanxin Zou
IGARSS2
2016 Semi-supervised cross-view scene model adaptation for remote sensing image classification
abstract
In this paper, we address the problem of semi-supervised visual domain adaptation for transferring scene category models from ground view images to overhead view very high-resolution (VHR) remote sensing images. We introduce a multiple kernel learning domain adaptation algorithm to fuse the information from multiple features and cope with the considerable variation in feature distributions between images from two domains. For each image, we first extract eight state-of-art local features and use the pretrained scene attribute model from ground-level SUN attribute database to predict attribute labels. For each scene class we learn an adapted target classifier based on multiple feature kernels by minimizing both the structural risk functional and the mismatch between data distributions of two domains. Experimental results demonstrate that it is possible to use a scene category model learned on a set of ground view scenes for semi-supervised classification of VHR remote sensing images.
Zhipeng Deng, Hao Sun 0042, Shilin Zhou 0001, Kefeng Ji
IGARSS2
2016 Unsupervised Cross-View Semantic Transfer for Remote Sensing Image Classification
abstract
We address the problem of unsupervised visual domain adaptation for transferring scene category models and scene attribute models from ground view images to overhead view very high-resolution (VHR) remote sensing images. We introduce a discriminative cross-view subspace alignment algorithm where each view is represented by a subspace spanned by eigenvectors. The source subspace is created using partial least squares correlation, whereas the target subspace is constructed by principal component analysis. Then, a mapping that aligns the source subspace and the target subspace is learned by minimizing a Bregman matrix divergence function. Finally, we project the labeled source data into the target aligned source subspace and the unlabeled target data into the target subspace and perform classification. Experimental results demonstrate that it is possible to use a scene category model or a scene attribute model learned on a set of ground view scenes for classification of VHR remote sensing images. Furthermore, the transferred visual attribute-based representations are human understandable and the classification results are better or comparable with state-of-the-art methods.
Hao Sun 0042, Shilin Zhou 0001, Huanxin Zou
IEEE Geosci. Remote. Sens. Lett.1
2016 Discriminative Subspace Alignment for Unsupervised Visual Domain Adaptation
Hao Sun 0042, Shilin Zhou 0001
Neural Process. Lett.1
2014 Local Sparsity Divergence for Hyperspectral Anomaly Detection
abstract
Anomaly detection (AD) has increasingly become important in hyperspectral imagery (HSI) owing to its high spatial and spectral resolutions. Many anomaly detectors have been proposed, and most of them are based on a Reed–Xiaoli (RX) detector, which assumes that the spectrum signature of HSI pixels can be modeled with Gaussian distributions. However, recent studies show that the Gaussian and other unimodal distributions are not a good fit to the data and often lead to many false alarms. This letter proposes a novel hyperspectral AD algorithm based on local sparsity divergence (LSD) without any distribution hypothesis. Our algorithm exploits the fact that targets and background lie in different low-dimensional subspaces and that targets cannot be effectively represented by their local surrounding background. A sliding dual-window strategy is first adopted to construct local spectral and spatial dictionaries, which enable the extraction of the sparse coefficients of each HSI pixel. Then, a consistent sparsity divergence index is proposed to compute the LSD map at each spectral band separately. Finally, joint segmentation of LSD maps over different bands is performed for AD. Experimental results on both simulated data and recorded data demonstrate the effectiveness of the proposed algorithm.
Zongze Yuan, Hao Sun 0042, Kefeng Ji, Zhiyong Li 0008, Huanxin Zou
IEEE Geosci. Remote. Sens. Lett.2
2013 Multimodal Remote Sensing Data Fusion via Coherent Point Set Analysis
abstract
We present a novel fusion algorithm for electronic-reconnaissance (ER) satellite and optical imaging satellite data using coherent point set (CPS) analysis. This work is motivated by a large-scale maritime surveillance problem, where ship groups in the observations are of particular interest for tactical and strategic operations. Fusion of observations from ER satellite and optical imaging satellite is a challenging task. On the one hand, dense and continuous measurement is not available for optical imagery. On the other hand, it is difficult to extract robust features from ER measurements. Considering that the size of a ship is often less than the distance among different ships, we treat each ship as a mass point. The contributions of our work are threefold. First, multisensor data fusion is accomplished by CPS association. To the best of our knowledge, this letter is the first to investigate CPS for multimodal remote sensing data fusion. Second, a novel geometry descriptor, which encodes the topological characteristics of a point set, is presented. Third, we combine both topological features and attributive features within the framework of Dempster–Shafer theory for CPS analysis. The proposed method has been tested using different sets of simulated data and recorded data. Experimental results demonstrate the effectiveness of the proposed method.
Huanxin Zou, Hao Sun 0042, Kefeng Ji, Chun Du, Chunyan Lu
IEEE Geosci. Remote. Sens. Lett.2