Hao Shi 0006

dblp:67/6136-6 · DBLP profile ↗
← Back
15ranked-venue papers
3as first author
6since 2021 · last 2025
0000-0002-2013-6592ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 15 · 3 first-author · 6 since 2021
YearPublicationVenuePosition
2025 PGMNet: A Prototype-Guided Multimodal Network for Ship Recognition in SAR Images
abstract
Ship recognition in synthetic aperture radar (SAR) images has extensive applications across various fields. However, the substantial intra-class variability and inter-class similarity present inherent challenges to achieving high-precision recognition. Speckle noise in the background reduces the signal-to-noise ratio, complicating the extraction of discriminative characteristics. Additionally, traditional convolutional neural networks-based methods, which rely solely on image processing frameworks without leveraging additional modality information, struggle to accurately represent the features of SAR targets with bright scatters. To address these issues, we propose a prototype-guided multimodal network (PGMNet), marking a pioneering effort to introduce an image-text multimodal fusion processing paradigm into SAR ship recognition. First, a generation strategy of the prototype and the key area is designed to improve the distinguishability between targets and backgrounds. Besides, a prototype-guided alignment module (PGAM) is implemented to assist the network in characterizing key area information, enhancing intra-class feature consistency. Furthermore, a text feature processing branch is incorporated to precisely describe ship size information and effectively integrate image-text multimodal features, reducing intra-class feature distance while enlarging inter-class feature distance. Extensive experiments on the OpenSARShip and FUSARShip datasets demonstrate that the proposed PGMNet achieves state-of-the-art (SOTA) performance. Notably, the accuracy of PGMNet is at least 11% higher than the current SOTA algorithms on the OpenSARShip-VI dataset.
Liang Chen 0004, Honghu Zhong, Hao Shi 0006, Wei Li 0032
IEEE Trans. Geosci. Remote. Sens.4
2025 DOGAN: DINO-Based Optical-Prior-Driven GAN for SAR-to-Optical Image Translation
abstract
To leverage the complementary advantages of SAR’s all-weather and all-day imaging capability and optical imagery’s intuitive visualization, SAR-to-optical image translation (S2OIT) has emerged as a promising solution to mitigate the interpretability challenges posed by SAR’s speckle noise and geometric distortions. However, the scale of high-quality registered SAR-optical data is limited, where incorporating priors is a viable solution. What’s more, the digging out of optical prior is insufficient among the existing methods, leading to inadequate synthesis of optical-like texture in translated optical images. To address these challenges, we propose DOGAN, a DINO-based optical-prior-driven GAN framework that integrates ample optical priors extracted from a pretrained DINO model into the S2OIT process. Specifically, to fully exploit the tremendous optical prior preserved in pretrained DINO and extract multiscale optical prior, a DINO-based Optical-prior Extraction (DOE) module is proposed. Furthermore, to elevate the domain adaptability of optical prior, a lightweight Stacked Optical-Aware (SOA) adapter is proposed to fine-tune DINO for remote sensing data with minimal trainable parameters. To instill the extracted affluent optical prior into the S2OIT pipeline stably, the SAR-optical Multi-scale Domain Alignment (SO-MDA) module is proposed, which employs L1 and Multi-kernel Maximum Mean Discrepancy (MK-MMD) losses to align intermediate optical and S2O features. Extensive experiments on SAR2Opt and SEN1-2 datasets demonstrate that DOGAN achieves state-of-the-art performance in both translation fidelity and structural realism. To the best of our knowledge, this is the first work to leverage DINO-based optical priors for the S2OIT task.
Jingfei He, Liang Chen 0004, Hao Shi 0006, Wei Li 0032
IEEE Trans. Geosci. Remote. Sens.3
2024 SWDiff: Stage-Wise Hyperspectral Diffusion Model for Hyperspectral Image Classification
abstract
Hyperspectral image classification (HSIC) has been a popular task in recent years. Even benefiting from the rapid development of deep neural networks (DNNs), there are still remaining intrinsic problems, including inadequate utilization of spatial-spectral information and insufficient labeled samples. The recent emergency of diffusion models (DMs) came to the fore because of their impressive refined image generation performance. DMs have been proven to not only can capture the underlying information of data through training the decoder of DMs, but also have more stable training than GANs while retaining even better performance. To better perceive and utilize spectral-spatial information while alleviating insufficient labeled samples simultaneously, we introduce the DM into HSIC from a data generation perspective. Specifically, we propose a stage-wise DM framework (SWDiff), dividing the HSIC task into three stages, including: pretrain the diffusion decoder with the hyperspectral image (HSI); generate new HSI cubes through the well-trained decoder to extra supply the original HSI set; and utilize the supplied dataset to train varied classifiers to obtain a better classification performance. Suitable pretraining could enable the decoder to acquire spatial-spectral information of the HSIs sufficiently via modeling spectral-spatial relationships across samples, leading to better utilization of spectral and spatial information of HSIs. Furthermore, the DM could provide the inference stage with spatial-spectral prior knowledge to ensure the feasibility and plausibility of the dataset complement, which could alleviate the insufficient labeled samples problem. Eventually, the classification stage will benefit from the first two stages.
Liang Chen 0004, Jingfei He, Hao Shi 0006, Wei Li 0032
IEEE Trans. Geosci. Remote. Sens.3
2023 Domain Adaptive Remote Sensing Scene Recognition via Semantic Relationship Knowledge Transfer
abstract
Scene recognition has attracted rising attentions of many researchers in the remote sensing fields, owing to the rapidly advancing of remote sensing devices in recent years. However, images obtained from various sensors dominate diverse sensor-specific characteristics, which will dramatically weaken the model transferability trained on a source data domain to a different target domain on account of the domain shift issues. To mitigate the domain discrepancy, most existing methods attend to align the cross-domain distributions. While the valuable knowledge of semantic relationships between different scenes is generally overlooked, and the underlying correlation across scenes cannot be fully discovered. For the sake of tackling this challenge, we propose an adaptive remote sensing scene recognition network, which can successfully transfer both the discriminative knowledge and cross-scene relationship from source to target. Specifically, in this paper, we acquire sensor-invariant representations in an adversarial manner and realize fine-grained conditional distribution alignment contrastively. In such a way, the tremendous domain gap can be mitigated to a large extent, and the discriminative and well-matched representations will be derived favorably. In addition, we explicitly construct class-wise relationship distributions belonging to two domains respectively and minimize their divergence to conduct semantic relationship knowledge transfer (SRKT), for the purpose of sufficiently unearthing the intrinsic semantic relative structures that can prompt generality of the model in the target domain. Finally, we conduct multiple experiments on representative multi-domain remote sensing benchmarks, and the extensive experimental results demonstrate the superiority of our proposed approach.
Shuang Li 0008, Chi Harold Liu, Yuqi Han, Hao Shi 0006, Wei Li 0032
IEEE Trans. Geosci. Remote. Sens.5
2022 SAR-to-Optical Image Translating Through Generate-Validate Adversarial Networks
abstract
Synthetic aperture radar (SAR) has the advantages of high resolution in all-weather and all-day. However, SAR images are hard to be understood, due to their unique imaging mechanism. The SAR to optical image translation can assist in interpreting and has become a topic of growing interest in the field of remote sensing. In this letter, a SAR to optical image translation network is proposed, called generate-validate adversarial networks (GVANs). More specifically, there are two Pix2Pix networks form the cyclic structure. The validate module is employed to increase the training process and improve the edge retention ability. In order to improve multidomain images adaptability, the embedded layer is proposed. Additionally, the dilation convolution layer is employed in the generator, which is more suitable for the characteristics of SAR images. The proposed method has experimented on the SEN1-2 dataset. The result demonstrates the superiority of the proposed method over state-of-the-art methods.
Hao Shi 0006, Bocheng Zhang, Yupei Wang, Zihan Cui, Liang Chen 0004
IEEE Geosci. Remote. Sens. Lett.1
2022 Dual-Path Sparse Hierarchical Network for Semantic Segmentation of Remote Sensing Images
abstract
Semantic segmentation of remote sensing images aims to label every pixel with the correct semantic category. The core challenge of the current deep convolutional network (ConvNet)-based methods lies in the difficulty of effectively aggregating high-level categorical semantics and low-level local details along the hierarchy of backbone. Most current approaches consider only fusing adjacent feature layers gradually with short-range feature connections, which lack the diversity of feature interactions, such as long-range cross-scale connections. To this end, we propose a novel dual-path sparse hierarchical network that is characterized by rich cross-scale feature interactions. Multiscale features are first sparsely grouped with a predefined interval, which is then aggregated via both long-range and short-range cross-scale connections in a hierarchical manner. Moreover, in order to further enrich the diversity of feature interactions, we also introduce another fusion path in parallel but with different sparsity for feature grouping, forming a dual-path network. In this way, our model is able to effectively aggregate multilevel features by incorporating both long-range and short-range feature interactions in both parallel and hierarchical manner. Meanwhile, the semantic and resolution gap between multilevel features can also be bridged.
Yupei Wang, Hao Shi 0006, Shan Dong, Yin Zhuang, Liang Chen 0004
IEEE Geosci. Remote. Sens. Lett.2
2020 Feature Enhanced Centernet for Object Detection in Remote Sensing Images
abstract
Multi-scale object detection in optical remote sensing imagery is a challenging task due to the varied object scales. Existed state-of-art object detection methods have achieved significant growth. However, most of the methods are based on default anchors, which need to be predefined. The multi-scale object detection accuracy still needs to be improved, especially for small and dense objects. To improve the robustness of the detection algorithm and the performance of multi-scale object detection, a novel anchor-free multi-scale object detection method Feature Enhanced CenterNet is proposed in this paper. First, we use the “encoder-decoder” structure and introduce horizontal connections to enhance feature representation capabilities. Second, an context-aware up-sampling method is proposed to obtain feature maps with suitable scale. To demonstrate the performance of the proposed method, we perform abundant experiments on the public remote sensing datasets. The experimental results demonstrate the robustness and effectiveness of the proposed method.
Tong Zhang 0028, Guanqun Wang, Yin Zhuang, He Chen 0004, Hao Shi 0006, Liang Chen 0004
IGARSS5
2020 Marginal Center Loss for Deep Remote Sensing Image Scene Classification
abstract
Recently, remote sensing image scene classification technology has been widely applied in many applicable industries. As a result, several remote sensing image scene classification frameworks have been proposed; in particular, those based on deep convolutional neural networks have received considerable attention. However, most of these methods have performance limitations when analyzing images with large intraclass variations. To overcome this limitation, this letter presents the marginal center loss with an adaptive margin. The marginal center loss separates hard samples and enhances the contributions of hard samples to minimize the variations in features of the same class. Experimental results on public remote sensing image scene data sets demonstrate the effectiveness of our method. After the model is trained using the marginal center loss, the variations in the features of the same class are reduced. Furthermore, a comparison with state-of-the-art methods proves that our model has competitive performance in the field of remote sensing image scene classification.
Jue Wang 0011, Wenchao Liu 0001, He Chen 0004, Hao Shi 0006
IEEE Geosci. Remote. Sens. Lett.5
2019 Spatial Enhanced-SSD For Multiclass Object Detection in Remote Sensing Images
abstract
Accurate multiclass object detection in remote sensing images is a challenging task, especially for small objects. Since the scales of objects in remote sensing images have a great variance, almost all of the advanced detection methods have shortcomings. Consequently, improving the accuracy of multiclass objects detection has always been the direction of researchers' efforts. In this paper, a spatial enhanced-Single Shot MultiBox Detector (SE-SSD) is proposed. First, to enhance the spatial information, we enlarge the input image channels with embedding oriented-gradients feature maps. Second, the multiple output layers in the backbone network are changed to reduce one pooling operation. Finally, we design a context module to enhance the receptive field for feature layer description in SE-SSD framework. Experimental results on DOTA dataset demonstrate that Spatial Enhanced-SSD method reaches a much higher mean average precision (mAP) than Faster R-CNN, SSD and other classic detection network.
Guanqun Wang, Yin Zhuang, Zhiru Wang, He Chen 0004, Hao Shi 0006, Liang Chen 0004
IGARSS5
2017 A novel method of speckle reduction and enhancement for SAR image
abstract
A new speckle reduction algorithm is presented in this paper to improve the visualization of synthetic aperture radar (SAR) images. The algorithm contains two steps. First, we propose a speckle reduction filter. This filter use different smoothing effect according to scene of sub-window. So it can maintain the details while noise suppression. Second, we propose a method to enhance the details. We select a neighborhood of pixel and divide it into homogeneous and heterogeneous. Then we use different reconstruction strategy to enhance the texture. After these steps, the speckle noise of SAR images has been decreased and the texture details have been enhanced at the same time. Experiments validate the algorithm with good performance.
Hao Shi 0006, Liang Chen 0004, Yin Zhuang, Jian Yang 0011
IGARSS1
2017 Pyramid integral image reconstruction algorithm for infrared remote sensing sea-land segmentation
abstract
The middle wave infrared remote (MWIR) images has complex scene information, low contrast ratios, and bipolar problems. To solve these problems, we propose a method that uses pyramid integral image reconstruction algorithm achieving sea-land automation segmentation. First, we calculate a gradient feature map (GFM), which extracts the structural information from an MWIR scene. Then, the GFM uses for sum are table (SAT) generation. The pyramid integral image reconstruction technology uses different scale factor reconstruct MWIR images by using SAT. Then the adaptive threshold method is employed for the sea-land segmentation on the multi-scale integral reconstruction images. Finally, we get sea-land refine segmentation result of MWIR images by synthesis analysis the multi-scale reconstruction images. By using GFM and pyramid integral image reconstruction operation, are avoid with the complex gray scene information. The integral image reconstruction can enhance the structure information and improve the reconstruction image contrast. This paper proposed method is from structure and texture information view point for sea-land segmentation, so the bipolar problem is solve in our method for MWIR images sea-land segmentation.
Penglin Wang, Yin Zhuang, He Chen 0004, Liang Chen 0004, Hao Shi 0006, Fukun Bi
IGARSS5
2017 M-FCN: Effective Fully Convolutional Network-Based Airplane Detection Framework
abstract
Airplane detection is a challenging problem in complex remote sensing imaging. In this letter, an effective airplane detection framework called Markov random field-fully convolutional network (M-FCN) is proposed. The M-FCN uses a cascade strategy that consists of an FCN-based coarse candidate extraction stage, a multi-Markov random field (multi-MRF)-based region proposal (RP) generation stage, and a final classification stage. In the first stage, the FCN model is trained to be sensitive to airplanes, and a coarse candidate map is generated. This model is scale-, direction-, and color-invariant and does not require many training examples. After the first stage, the coarse candidate map is used as the initial labeling field for a multi-MRF algorithm, and RPs are generated according to the multi-MRF output. This RP-generating strategy can yield more accurate locations with fewer RPs. In the last stage, a convolutional neural network-based classifier is used to improve the precision of the entire framework. Experiments show that the M-FCN has high precision, recall, and location accuracy.
Yiding Yang, Yin Zhuang, Fukun Bi, Hao Shi 0006, Yizhuang Xie
IEEE Geosci. Remote. Sens. Lett.4
2017 Harbor Water Area Extraction From Pan-Sharpened Remotely Sensed Images Based on the Definition Circle Model
abstract
Harbor water area extraction is a key step in nearshore environment pollution surveillance using remote sensing image processing techniques. This letter proposes the definition circle (DC) model of color gradient to describe color fluctuations in harbor water surface areas based on pan-sharpened remote sensing images. The DC model includes two steps: center setting and radius tuning. In the center setting process, labeled training set pixels are selected in the red, green, and blue color space. Then, center setting is completed in the hue, saturation, and intensity color space using the perceptron model. In the radius tuning process, positive and negative sample pixels are used to tune the radius value. After these two steps, the DC model can describe the color gradient of a water surface area and provide accurate harbor water area extraction. A series of experiments shows that the proposed DC model is robust and performs better than other extraction methods based on pan-sharpened remote sensing images.
Yin Zhuang, Penglin Wang, Yiding Yang, Hao Shi 0006, He Chen 0004, Fukun Bi
IEEE Geosci. Remote. Sens. Lett.4
2016 Feature-Area Optimization: A Novel SAR Image Registration Method
abstract
This letter proposes a synthetic aperture radar (SAR) image registration method named feature-area optimization (FAO). First, the traditional area-based optimization model is reconstructed and decomposed into three key but uncertain factors: initialization, slice set, and regularization. Next, structural features are extracted by scale-invariant feature transform (SIFT) in dual-resolution space (SIFT-DRS), a novel SIFT-like method dedicated to FAO. Then, the three key factors are determined based on these features. Finally, solving the factor-determined optimization model can get the registration result. A series of experiments demonstrate that the proposed method can register multitemporal SAR images accurately and efficiently.
Fukun Bi, Liang Chen 0004, Hao Shi 0006
IEEE Geosci. Remote. Sens. Lett.4
2015 Accurate Urban Area Detection in Remote Sensing Images
abstract
Automatic urban area detection in remote sensing images is an important application in the field of earth observation. Most of the existing methods employ feature classifiers and thereby contain a data training process. Moreover, some methods cannot detect urban areas in complex scenes accurately. This letter proposes an automatic urban area detection method that uses multiple features that have different resolutions. First, a downsampled low-resolution image is used to segment the candidate area. After the corner points of the urban area are extracted, a weighted Gaussian voting matrix technique is employed to integrate the corner points into the candidate area. Then, the edge features and homogeneous region are extracted by using the original high-resolution image. Using these results as the input, the processes of guided filtering and contrast enhancement can finally detect accurately the urban areas. This method combines multiple features, such as corner, edge, and regional characteristics, to detect the urban areas. The experimental results show that the proposed method has better detection accuracy for urban areas than the existing algorithms.
Hao Shi 0006, Liang Chen 0004, Fukun Bi, He Chen 0004
IEEE Geosci. Remote. Sens. Lett.1