Shuang Wang 0001

dblp:86/220-1 · DBLP profile ↗
← Back
160ranked-venue papers
18as first author
81since 2021 · last 2026
0000-0003-4940-1211ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 98 · 9 first-author · 46 since 2021Artificial intelligence and machine learning · 40 · 3 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 30 · 6 first-author · 21 since 2021Computer networks · 4 · 1 first-author · 4 since 2021
YearPublicationVenuePosition
2026 Towards Stable Source-Free Domain Adaptive Semantic Segmentation
Dong Zhao 0007, Qi Zang, Nan Pu, Jinlong Li 0003, Shuang Wang 0001, Nicu Sebe, Zhun Zhong
Int. J. Comput. Vis.5
2026 KCI-Net: Knowledge-Based Contourlet Inference Network for Super-Resolution
abstract
Textural details are useful for image super-resolution, but massive CNN methods ignored the high-frequency components and generated over-smoothed outputs. The knowledge-based contourlet inference network is proposed in this paper. Different from other CNN-based methods that are directly infer high-resolution (HR) images, our model learns to reconstruct the HR image through the series of corresponding contourlet coefficients. Specifically, first, we consider the low-pass subbands of the contourlet as the corresponding low-resolution (LR) image. Then, feed it to the embedding net with residual blocks to provide adequate information for the contourlet coefficients prediction. Finally, we innovatively convert the estimation of contourlet coefficients into the estimation of the generalized gaussian distribution (GGD) parameters, and design the corresponding loss function to ensure training stability, which explores the smoothness of the contour effectively and guarantees the general structure and details of images. Experiments on four remote sensing datasets, four natural scenes and human-made content datasets, and the outdoor dataset demonstrate the superiority of the proposed model quantitatively and qualitatively.
Mengkun Liu, Licheng Jiao, Xu Liu 0006, Lingling Li 0002, Fang Liu 0001, Shuyuan Yang 0001, Shuang Wang 0001, Biao Hou
IEEE Trans. Circuits Syst. Video Technol.7
2025 ChangeDiff: A Multi-Temporal Change Detection Data Generator with Flexible Text Prompts via Diffusion Model
abstract
Data-driven deep learning models have enabled tremendous progress in change detection (CD) with the support of pixel-level annotations. However, collecting diverse data and manually annotating them is costly, laborious, and knowledge-intensive. Existing generative methods for CD data synthesis show competitive potential in addressing this issue but still face the following limitations: 1) difficulty in flexibly controlling change events, 2) dependence on additional data to train the data generators, 3) focus on specific change detection tasks. To this end, this paper focuses on the semantic CD (SCD) task and develops a multi-temporal SCD data generator ChangeDiff by exploring powerful diffusion models. ChangeDiff innovatively generates change data in two steps: first, it uses text prompts and a text-to-layout (T2L) model to create continuous layouts, and then it employs layout-to-image (L2I) to convert these layouts into images. Specifically, we propose multi-class distribution-guided text prompts (MCDG-TP), allowing for layouts to be generated flexibly through controllable classes and their corresponding ratios. Subsequently, to generalize the T2L model to the proposed MCDG-TP, a class distribution refinement loss is further designed as training supervision. Our generated data shows significant progress in temporal continuity, spatial diversity, and quality realism, empowering change detectors with accuracy and transferability.
Qi Zang, Shuang Wang 0001, Dong Zhao 0007, Wenjun Yi, Zhun Zhong
AAAI3
2025 Feature Spectrum Learning for Remote Sensing Change Detection
abstract
Change detection (CD) holds significant implications for Earth observation, in which pseudo-changes between bitemporal images induced by imaging environmental factors are key challenges. Existing methods mainly regard pseudo-changes as a kind of style shift and alleviate it by transforming bitemporal images into the same style using generative adversarial networks (GANs). Nevertheless, their efforts are limited by the complexity of optimizing GANs and the absence of guidance from physical properties. This paper finds that the spectrum transformation (ST) has the potential to mitigate pseudo-changes by aligning in the frequency domain carrying the style. However, the benefit of ST is largely constrained by two drawbacks: 1) limited transformation space and 2) inefficient parameter search. To address these limitations, we propose a Feature Spectrum learning (FeaSpect) that adaptively eliminate pseudo-changes in the latent space. For the drawback 1), FeaSpect directs the transformation towards stylealigned discriminative features via feature spectrum transformation (FST). For the drawback 2), FeaSpect allows FST to be trainable, efficiently discovering optimal parameters via extraction box with adaptive attention and extraction box with learnable strides. Extensive experiments on challenging datasets demonstrate that our method remarkably outperforms existing methods and achieves a commendable trade-off between accuracy and efficiency. Importantly, our method can be easily injected into other frameworks, achieving consistent improvements.
Qi Zang, Dong Zhao 0007, Shuang Wang 0001, Dou Quan, Zhun Zhong
CVPR3
2025 FisherTune: Fisher-Guided Robust Tuning of Vision Foundation Models for Domain Generalized Segmentation
abstract
Vision Foundation Models (VFMs) excel in generalization due to large-scale pretraining, but fine-tuning them for Domain Generalized Semantic Segmentation (DGSS) while maintaining this ability remains a challenge. Existing approaches either selectively fine-tune parameters or freeze the VFMs and update only the adapters, both of which may underutilize the VFMs’ full potential in DGSS tasks. We observe that domain-sensitive parameters in VFMs, arising from task and distribution differences, can hinder generalization. To address this, we propose FisherTune, a robust fine-tuning method guided by the Domain-Related Fisher Information Matrix (DR-FIM). DR-FIM measures parameter sensitivity across tasks and domains, enabling selective updates that preserve generalization and enhance DGSS adaptability. To stabilize DR-FIM estimation, FisherTune incorporates variational inference, treating parameters as Gaussian-Distributed variables and leveraging pre-trained priors. Extensive experiments show that Fisher-Tune achieves superior cross-domain segmentation while maintaining generalization, outperforming both selective-parameter and adapter-based methods.
Dong Zhao 0007, Jinlong Li 0003, Shuang Wang 0001, Qi Zang, Nicu Sebe, Zhun Zhong
CVPR3
2025 Domain-Aware Category-Level Geometry Learning Segmentation for 3D Point Clouds
Pei He, Lingling Li 0002, Licheng Jiao, Ronghua Shang, Fang Liu 0001, Shuang Wang 0001, Xu Liu 0006, Wenping Ma 0001
ICCV6
2025 Pseudo-SD: Pseudo Controlled Stable Diffusion for Semi-Supervised and Cross-Domain Semantic Segmentation
Dong Zhao 0007, Qi Zang, Shuang Wang 0001, Nicu Sebe, Zhun Zhong
ICCV3
2025 Predicting Spectral Information for Self-Supervised Signal Classification
abstract
Deep learning methods have demonstrated remarkable performance across various communication signal processing tasks. However, most signal classification methods require a substantial amount of labeled samples for training, posing significant challenges in the field of communication signals, as labeling necessitates expert knowledge. This paper proposes a novel self-supervised signal classification method called Spectral-Guided Self-Supervised Signal Classification (SGSSC). Specifically, to leverage frequency-domain information with modulation semantics as prior knowledge for the model, we design a previously unexplored pretext task tailored to the format of signal data. This task involves predicting spectral information from masked time-domain signals, enabling the model to learn implicit signal features through cross-domain pattern transformation. Furthermore, the pretext task in the SGSSC method is relevant to the downstream classification task, and using traditional fine-tuning strategies on the downstream task may lead to the loss of certain features associated with the pretext task. Therefore, we propose an attention mechanism-based fine-tuning strategy that adaptively integrates pre-trained features from different levels. Extensive experimental results validate the superiority of the SGSSC method. For instance, when the proportion of labeled samples is only 0.5%, our method achieves an average improvement of 2.3% in downstream classification tasks compared to the best-performing self-supervised training strategies.
Shuang Wang 0001, Hantong Xing, Chenxu Wang 0001, Dou Quan, Rui Yang 0038, Dong Zhao 0007, Luyang Mei
IJCAI2
2025 SeCoV2: Semantic Connectivity-Driven Pseudo-Labeling for Robust Cross-Domain Semantic Segmentation
abstract
Pseudo-labeling is a dominant strategy for cross-domain semantic segmentation (CDSS), yet its effectiveness is limited by fragmented and noisy pixel-level predictions under severe domain shifts. To address this, we propose a semantic connectivity-driven pseudo-labeling framework, SeCo, which constructs and refines pseudo-labels at the connectivity level by aggregating high-confidence pixels into coherent semantic regions. The framework includes two key components: Pixel Semantic Aggregation (PSA), which leverages a dual prompting strategy to preserve category-specific granularity, and Semantic Connectivity Correction with Loss Distribution (SCC-LD), which filters noisy regions based on early-loss statistics. Building upon this foundation, we further present SeCoV2, which introduces SCC-Unc, a novel uncertainty-aware correction module that constructs a connectivity graph and enforces relational consistency for robust refinement in ambiguous regions. SeCoV2 also broadens the applicability of SeCo by extending evaluation to more challenging scenarios, including open-set and multimodal adaptation, semi-supervised domain generalization, and by validating compatibility with different interactive foundation segmentation models such as SAM Kirillov et al. 2023, SEEM Zou et al. 2023, and Fast-SAM Zhao et al. 2023. Extensive experiments across six CDSS tasks demonstrate that SeCoV2 achieves consistent improvements over previous methods, with an average performance gain of up to +4.6%, establishing new state-of-the-art results. These findings highlight the effectiveness and generalization ability for robust adaptation in diverse real-world environments.
Dong Zhao 0007, Qi Zang, Nan Pu, Shuang Wang 0001, Nicu Sebe, Zhun Zhong
IEEE Trans. Pattern Anal. Mach. Intell.4
2025 AFLNet: Auxiliary Feature Learning-Guided Cross-Channel Automatic Modulation Classification
abstract
This paper conducted a thorough investigation into the primary difficulty of the cross-channel automatic modulation classification (AMC) task by examining data distribution and feature space of different channel conditions. We concluded that the disruption of the target channel feature space structure breakdown the mapping relationship across channels, serving as the main contributor to model performance degradation. Based on the above conclusion, in order to improve the performance of cross-channel AMC, we introduce the Auxiliary Feature Learning-Guided Network (AFLNet). This network improves the structure of the target feature space through two uniquely designed tasks and facilitates efficient cross-domain alignment via a collaborative alignment mechanism. Specifically, AFLNet integrates similarity-based and confidence-based auxiliary feature learning tasks to enhance the discriminability of the target feature space and maintain the correspondence of category structures across different channels, thereby reducing the difficulty of feature alignment. The collaborative alignment mechanism combines adversarial training-based and self-training-based feature alignment methods, leveraging their mutually reinforcing effect and complementary strengths in global alignment and class-level alignment to enhance overall alignment performance. We carried out extensive experiments across four scenarios characterized by substantial channel variations, verifying that AFLNet achieves state-of-the-art with accuracy improvement of up to 9.71%.
Hantong Xing, Shuang Wang 0001, Chenxu Wang 0004, Dou Quan, Hanlin Mo, Luyang Mei, Huaji Zhou, Licheng Jiao
IEEE Trans. Commun.2
2025 Joint Style and Layout Synthesizing: Toward Generalizable Remote Sensing Semantic Segmentation
abstract
This paper studies the domain generalized remote sensing semantic segmentation (RSSS), aiming to generalize a model trained only on the source domain to unseen domains. Existing methods in computer vision treat style information as domain characteristics to achieve domain-agnostic learning. Nevertheless, their generalizability to RSSS remains constrained, due to the incomplete consideration of domain characteristics. We argue that remote sensing scenes have layout differences beyond just style. Considering this, we devise a joint style and layout synthesizing framework, enabling the model to jointly learn out-of-domain samples synthesized from these two perspectives. For style, we estimate the variant intensities of per-class representations affected by domain shift and randomly sample within this modeled scope to reasonably expand the boundaries of style-carrying feature statistics. For layout, we explore potential scenes with diverse layouts in the source domain and propose granularity-fixed and granularity-learnable masks to perturb layouts, forcing the model to learn characteristics of objects rather than variable positions. The mask is designed to learn more context-robust representations by discovering difficult-to-recognize perturbation directions. Subsequently, we impose gradient angle constraints between the samples synthesized using the two ways to correct conflicting optimization directions. Extensive experiments demonstrate the superior generalization ability of our method over existing methods.
Qi Zang, Shuang Wang 0001, Dong Zhao 0007, Zhun Zhong, Biao Hou, Licheng Jiao
IEEE Trans. Circuits Syst. Video Technol.2
2025 FAFormer: Frequency-Analysis-Based Transformer Focusing on Correlation and Specificity for Pansharpening
abstract
Pan-sharpening refers to fusing remote sensing multispectral (MS) and panchromatic (PAN) images to generate high-resolution multispectral (HR-MS) images. Recent advancements in deep learning-based pan-sharpening techniques have shown promising results. However, they face the following two issues. On one hand, there is a modality gap between MS and PAN images. Directly fusing them can lead to spectral and spatial distortions. On the other hand, the fusion process is prone to information loss, which can lead to image blurriness. To tackle these issues, we develop a Transformer-based model: FAFormer, which incorporates frequency analysis and focuses on the correlation and specificity of the PAN and MS images. Focusing on correlation can reduce the spectral and spatial distortions while focusing on specificity can reflect the specific information from MS and PAN images in the fusion result. We utilize the Discrete Wavelet Transform (DWT) to obtain the correlate and specific features. We introduce bijective functions based on the Transformer to design an Integrated Attention Block (IAB). As a critical component of the model, it effectively utilizes the correlation and specificity of the two images. In designing the model’s overall framework, we employ a Correlative Feature Attention Module (CFAM) to leverage the correlation between MS and PAN. We utilize a Specific Feature Attention Module (SFAM) to integrate specific information into fused features gradually. Experimental results show that our method improves pan-sharpening performance and has practical value. Codes are available at https://github.com/Xidian-AIGroup190726/FAFormer.
Yifan Meng, Hao Zhu 0009, Xiaoyu Yi 0002, Biao Hou, Shuang Wang 0001, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.5
2025 Burden-Free Distillation From Foundation Model for Efficient Remote Sensing Change Detection
abstract
Applying vision foundation models to remote sensing change detection (CD) has attracted extensive research attention. These studies employ inherent general knowledge from vision foundation models to enhance CD performance. Existing methods explicitly employ the foundation model as a feature extractor while designing additional learnable modules to bridge the task gap. However, these methods substantially increase the computational burden and memory demand in the inference. This paper therefore focuses on addressing the core challenge of effectively leveraging the knowledge from vision foundation models to enhance CD performance while maintaining computational efficiency. Instead of explicitly utilizing the foundation model, we propose Burden-Free Distillation (BFD), an architecture-agnostic foundation model-based distillation framework for efficient CD. BFD transfers the general knowledge from foundation models to task-specific models, thereby eliminating the dependency on foundation models during inference. Specifically, BFD transfers the foundation model knowledge through Dual-temporal Feature Matching module (DFM). This module enables multi-level feature alignment by computing pixel-wise spatial similarity between the foundation models’ general features and the CD models’ task-specific features. Additionally, we leverage patch contrastive distillation, which transfers localized structural patterns to the CD model to further mitigate task discrepancies between foundation models and CD models. We conduct extensive experiments across multiple foundation models and CD architectures, experimental results demonstrate that BFD effectively adapts the knowledge of foundation models to CD tasks without additional computational burden. Compared to other foundation model-based CD methods, BFD reduces the model parameters by 80.3% and improves IoU by 1.78% on the S2Looking dataset. The code is available at https://github.com/Younger-hua/Burden-Free-Distillation.
Shuang Wang 0001, Chonghua Lv, Dou Quan, Ning Huyan, Xianwei Cao, Jingxi Sun, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.1
2025 CDFNet: Cross-Domain Feature Fusion Network for PolSAR Terrain Classification
abstract
The scarcity of labeled data and domain shift among polarimetric synthetic aperture radar (PolSAR) images degrades the performance of the supervised-learning-based algorithm. Some unsupervised domain adaptation (UDA) algorithms have been proposed to address this problem and achieve good performance. The existing UDA algorithms for PolSAR terrain classification focus on the feature distribution shift problem but ignore the label shift problem in UDA task. In addition, feature alignment-based algorithms generate pseudo labels for target domain which introduce label noise and compromising the UDA performance. To alleviate the problems above, we present a cross-domain feature fusion network (CDFNet) for PolSAR terrain classification. Specifically, a domain-balanced sampling (DBS) module is proposed to obtain a nearly balanced training dataset to alleviate the label shift problem. Then, a cross-domain feature fusion (CDF) module is presented to achieve class-wise feature alignment with no additional label noise introduction. Experimental results on four PolSAR datasets demonstrate that our algorithm outperforms state-of-the-art UDA algorithms in terms of target domain performance.
Shuang Wang 0001, Zhuangzhuang Sun, Tianquan Bian, Yuwei Guo 0001, Linwei Dai, Yanhe Guo, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.1
2025 Generalization-Aware Remote Sensing Change Detection via Domain-Agnostic Learning
abstract
Change detection has essential significance for the region's development, in which pseudo-changes between bitemporal images induced by imaging environmental factors are key challenges. Existing transformation-based methods regard pseudo-changes as a kind of style shift and alleviate it by transforming bitemporal images into the same style using generative adversarial networks (GANs). However, their efforts are limited by two drawbacks: 1) Transformed images suffer from distortion that reduces feature discrimination. 2) Alignment hampers the model from learning domain-agnostic representations that degrades performance on scenes with domain shifts from the training data. Therefore, oriented from pseudo-changes caused by style differences, we present a generalizable domain-agnostic difference learning network (DonaNet). For the drawback 1), we argue for local-level statistics as style proxies to assist against domain shifts. For the drawback 2), DonaNet learns domain-agnostic representations by removing domain-specific style of encoded features and highlighting the class characteristics of objects. In the removal, we propose a domain difference removal module to reduce feature variance while preserving discriminative properties and propose its enhanced version to provide possibilities for eliminating more style by decorrelating the correlation between features. In the highlighting, we propose a cross-temporal generalization learning strategy to imitate latent domain shifts, thus enabling the model to extract feature representations more robust to shifts actively. Extensive experiments conducted on three public datasets demonstrate that DonaNet outperforms existing state-of-the-art methods with a smaller model size and is more robust to domain shift.
Qi Zang, Shuang Wang 0001, Dong Zhao 0007, Dou Quan, Licheng Jiao
IEEE Trans. Multim.2
2025 Boosting Generalization of Semantic Segmentation With Unseen Style Seeking-Based Meta-Learning
abstract
This article considers a worst and most challenging scene in domain generalization (DG), where a model aims to generalize well on unseen domains while only one single domain is available for training. Existing randomization-based methods achieve this goal by enriching the style of the training data. However, they fail to guarantee the diversity of newly generated data required for generalization and thus lead to insufficient expansion of the training distribution. Thus, we propose a novel single DG (SDG) framework, unseen style seeking-based meta-learning (USSML). In USSML, multiple plausible domains with various styles are first constructed from a single source domain and the combination is performed across generated domains to emulate unseen images, extending the distribution boundaries of the source domain. The domain combination is performed at two levels, i.e., global and instance, to meet the generalization challenge in semantic segmentation. Then, the generated diverse domains are further exploited to force the model to optimize in an unbiased manner across all domains by relearning regions lacking domain-invariant representation capability, driving the model toward domain invariance. A point worth mentioning is that the proposed method is easily integrated into existing segmentation methods with little computational cost to improve their generalization. Extensive experiments are conducted on five popular segmentation datasets and the results have verified the effectiveness of USSML in improving the model's generalization and the superiority of USSML over existing works.
Qi Zang, Shuang Wang 0001, Dong Zhao 0007, Wanqing Li 0001, Dou Quan, Licheng Jiao
IEEE Trans. Neural Networks Learn. Syst.2
2025 PSRNet: Few-Shot Automatic Modulation Classification Under Potential Domain Differences
abstract
Learning from a limited number of samples in automatic modulation classification (AMC) has garnered considerable attention. However, existing few-shot AMC works solely focus on single-domain conditions where the training and testing data share the same data distribution, which overlook the potential domain differences. In practice, the complex and variable communication channels, along with different radio frequency (RF) devices, may result in significant data distribution differences, which can be defined as cross-domain conditions. The neglect of such cross-domain conditions may leads to a significant decline in the performance of existing few-shot AMC models. To consider a more general situation, this paper unifies single-domain and cross-domain few-shot AMC into one task, named SaC-FSL. We propose the Paired Samples Relationship Network (PSRNet) as a solution. PSRNet does not require additional network structure design for domain shifts. It distinguishes categories by learning the relationships between sample pairs rather than directly learning the features of samples. To achieve this, we randomly pair the samples to construct different relationships between different classes and domains, and learn these relationships through classification task. Extensive experiments conducted on multiple datasets have demonstrated the superiority of our PSRNet, which can achieve considerable improvements in both single-domain and cross-domain conditions.
Hantong Xing, Shuang Wang 0001, Luyang Mei, Huaji Zhou, Licheng Jiao
IEEE Trans. Wirel. Commun.2
2024 Stable Neighbor Denoising for Source-free Domain Adaptive Segmentation
abstract
We study source-free unsupervised domain adaptation (SFUDA) for semantic segmentation, which aims to adapt a source-trained model to the target domain without accessing the source data. Many works have been proposed to address this challenging problem, among which uncertainty-based self-training is a predominant approach. However, without comprehensive denoising mechanisms, they still largely fall into biased estimates when dealing with different domains and confirmation bias. In this paper, we observe that pseudo-label noise is mainly contained in unstable samples in which the predictions of most pixels undergo significant variations during self-training. Inspired by this, we propose a novel mechanism to denoise unstable samples with stable ones. Specifically, we introduce the Stable Neighbor Denoising (SND) approach, which effectively discovers highly correlated stable and unstable samples by nearest neighbor retrieval and guides the reliable optimization of unstable samples by bi-level learning. Moreover, we compensate for the stable set by object-level object paste, which can further eliminate the bias caused by less learned classes. Our SND enjoys two advantages. First, SND does not require a specific segmentor structure, endowing its universality. Second, SND simultaneously addresses the issues of class, domain, and confirmation biases during adaptation, ensuring its effectiveness. Extensive experiments show that SND consistently outperforms state-of-the-art methods in various SFUDA semantic segmentation settings. In addition, SND can be easily integrated with other approaches, obtaining further improvements. The source code is available at https://github.com/DZhaoXd/SND.
Dong Zhao 0007, Shuang Wang 0001, Qi Zang, Licheng Jiao, Nicu Sebe, Zhun Zhong
CVPR2
2024 A Dual-Branch Random Mask Alignment Framework for Semi-Supervised PolSAR Terrain Classification
abstract
Despite the recent success of deep learning based polarimetric synthetic aperture radar(PolSAR) classification algorithms, it remains challenging in the scenario of limited labeled samples. Existing semi-supervised PolSAR terrain classification methods focus on the exploitation of pseudolabels, which are unreliable with limited labeled samples. To solve this problem, we propose a dual-branch random mask alignment framework for PolSAR terrain classification task. First, we propose a balanced regional expansion algorithm for labeled sample expansion. Then, to fully exploit the massive unlabeled samples, we designed a dual-branch network using two different polarization decomposition features as inputs, and a random mask alignment loss is employed to achieve consistency constraints on the unlabeled samples. Experimental results on two PolSAR datasets demonstrate that the proposed method achieve excellent performance with limited labeled samples.
Tianquan Bian, Zhuangzhuang Sun, Shuang Wang 0001, Dou Quan, Yanhe Guo
IGARSS4
2024 TfNet: Building Detection in Remote Sensing Images Using Multi-Scale Feature Fusion
abstract
Building detection in remote sensing images is significant to urban land planning, battlefield environment perception, and illegal building monitoring. Existing methods excel in detecting small-scale building areas but struggle with remote sensing images with diverse and multi-scale characteristics. To solve this problem, this paper proposes Two-stage Feature fusion Network(TFNet). Specifically, we propose dense short connection module to enable internal interaction between features from different layers, achieving intra-stage feature fusion. Moreover, the holistically-nested edge detection network is used to integrate the effective information between different layers to achieve inter-stage feature fusion. In addition, the adaptive fusion weight is introduced to make the model adaptively select the weights of varying levels of features. Experimental results demonstrate the effectiveness of TFNet in improving the detection performance of multi-scale buildings in remote sensing images.
Chunlei Han, Luyang Mei, ZhongQian Jin, Shuang Wang 0001, Siyu Cao, Rui Yang 0038
IGARSS6
2024 Active Domain Adaptive Semantic Segmentation with Regional Relative Entropy for Remote Sensing Images
abstract
This paper presents a novel approach using active learning to tackle domain adaptation challenges in remote sensing semantic segmentation. Unsupervised Domain Adaptation for Semantic Segmentation (UDASS) aims to transfer a model trained on labeled source domain data to an unlabeled target domain. Existing UDASS methods struggle with the complexity of domain shift factors in remote sensing scenes, such as resolution, imaging mechanisms, geography, and species distribution, falling short of fully supervised performance. To address this, we propose integrating active learning, selecting a valuable (e.g. 2.2%) subset of pixel annotations from the target domain, and combining it with UDASS methods. Our method devises region-relative entropy metric to identify informative yet challenging pixels, facilitating better adaptation. Experimental results on two challenging domain adaptation tasks validate the efficacy of our technique, achieving performance comparable to fully supervised pixel labeling with only 2.2% annotated data.
Zhengyao Wang, Dong Zhao 0007, Shuang Wang 0001
IGARSS6
2024 Enhancing Change Detection Robustness: A Whitening Transformation Approach
abstract
The vast amount of remote sensing data has been instrumental in supporting research on change detection algorithms based on deep learning.However, factors such as geographical changes and variations in sensor parameters can result in significant style differences between remote sensing images at different time points, leading to a decline in model performance.To address this issue, this paper proposes a change detection algorithm based on whitening feature extraction, aiming to alleviate distribution differences by decoupling domain-invariant discriminative features from domain-specific style features.The effectiveness of the proposed method is demonstrated through transfer experiments from the SVCD dataset to the SZADA dataset and from the SYSU dataset to the SZADA dataset.
Qi Zang, Zhengyao Wang, Dou Quan, Shuang Wang 0001
IGARSS6
2024 ConDA: Continual Adaptation in Remote Sensing Via Visual Style Playback
abstract
This study focuses on continual adaptation in remote sensing semantic segmentation, addressing challenges posed by frequent data updates and model forgetting. Remote sensing images exhibit variations in visual styles due to factors like location, time, and weather conditions, creating distinct domains. To counter performance degradation in new domains, we introduce a new challenge task in remote sensing, termed Continual Domain Adaptation (ConDA). Our innovative Visual Style Replay method employs Variational Auto-Encoder (VAE) and knowledge distillation, enabling the model to continuously learn from historical domains without forgetting. The proposed approach, tested on the INRIA dataset under ConDA settings, outperforms existing methods in combating catastrophic forgetting in remote sensing segmentation tasks. This contributes to adapting deep learning models to real-world scenarios with continually evolving remote sensing data.
Ketao Zhong, Dong Zhao 0007, Shuang Wang 0001, Yanhe Guo
IGARSS5
2024 Fourier Domain Adaptive Multi-Modal Remote Sensing Image Template Matching Based on Siamese Network
abstract
Multi-modal remote sensing image template matching is a meaningful and crucial topic in remote sensing image processing. However, due to different imaging mechanisms, there are significant nonlinear radiometric variations among multi-modal remote sensing images, increasing the matching challenge and leading to poor matching performances. To tackle this issue, this paper proposes a Fourier Domain Adaptive Network (FDANet) for multi-modal remote sensing image matching. Firstly, FDANet randomly swaps the low-frequency spectrum information between multi-modal images through the Fourier transform to reduce differences among multi-modal images, enhancing network adaptability to different image modalities and improving the multi-modal image matching performance. Secondly, FDANet extracts domain-invariant features from the transformed images through a deep Siamese network. After that, FDANet performs template matching and achieves high-precision multi-modal remote sensing image matching. In addition, we adopt the contrastive learning loss to optimize the FDANet. Extensive experiments on multi-modal remote sensing image matching demonstrate the effectiveness and advantages of the proposed FDANet.
Chonghua Lv, Dou Quan, Shuang Wang 0001, Xiangming Jiang, Yu Gu 0015, Licheng Jiao
IGARSS4
2024 Mitigating Style Differences in Bitemporal Remote Sensing Images for Change Detection
abstract
Change detection has seen significant advancements with the development of deep learning. However, due to variations in sensors or atmospheric conditions, bitemporal images often exhibit visually significant style differences, posing challenges for the detection of changed regions. This paper presents a change detection network designed to effectively address the challenges posed by style differences in bitemporal images. The proposed network comprises a color difference unification module and a generalized feature extraction module, which focuses the network on really changed areas. The color difference unification module harmonizes the color space of bitemporal remote sensing images, thereby mitigating the impact of style differences attributed to objective conditions. The generalized feature extraction module, ensuring robust feature representation for image pairs and further reducing style differences between bitemporal images. Experimental results demonstrate the superiority of our proposed method compared to existing change detection algorithms, confirming its suitability for fulfilling the requirements of change detection tasks.
Qi Zang, Dong Zhao 0007, Shuang Wang 0001
IGARSS6
2024 Object-Level Change Detection via Siamese Detection Network
abstract
Traditional change detection methods often lack instance-specific analysis, resulting in inefficient resource allocation and response strategies. This paper introduces a novel network for instance-level change detection. Our approach utilizes a dual-stream encoder with shared-weight backbone to extract robust features, followed by a differential process to highlight changes while suppressing unchanged background. Integration of low-level and high-level semantic information using a Feature Pyramid Network (FPN) enhances the model’s ability to discern subtle changes. Our instance-level bounding box detection module isolates individual change instances, with a subsequent segmentation module delineating precise boundaries. Evaluation on diverse remote sensing datasets demonstrates superior accuracy and computational efficiency compared to existing techniques. This framework not only advances change detection but also offers insights into land cover and land use dynamics. The code is available at https://github.com/DZhaoXd/object-levelchange-detection.
Dong Zhao 0007, Hantong Xing, Shuang Wang 0001
IGARSS6
2024 MfrNet: A New Multi-Scale Feature Refining Method for Remote Sensing Image Change Captioning
abstract
Remote Sensing Image Change Captioning (RSICC) is an emerging multimodal field with promising prospects. This paper introduces a remote sensing image change caption model based on multi-scale and refined features. First, it extracts multi-scale features from dual-temporal images and then feeds them into the JointAtt and Dence Fusion (JADF) module for attention mutual guidance and feature refinement to eliminate noise. Next, the features are input into a transformer-based sentence generator for change statement generation. We conducted experiments on the Levir-CC dataset comparing our approach with existing methods, the results indicate that our MFRNet outperforms state-of-the-art methods in all metrics.
Kaiqi Xu, Yingping Han, Rui Yang 0038, Xiutiao Ye, Yanhe Guo, Hantong Xing, Shuang Wang 0001
IGARSS7
2024 Selection and Reconstruction of Key Locals: A Novel Specific Domain Image-Text Retrieval Method
abstract
In recent years, Vision-Language Pre-training (VLP) models have demonstrated rich prior knowledge for multimodal alignment, prompting investigations into their application in Specific Domain Image-Text Retrieval(SDITR) such as Text-Image Person Re-identification (TIReID) and Remote Sensing Image-Text Retrieval (RSITR). Due to the unique data characteristics in specific scenarios, the primary challenge is to leverage discriminative fine-grained local information for improved mapping of images and text into a shared space. Current approaches interact with all multimodal local features for alignment, implicitly focusing on discriminative local information to distinguish data differences, which may bring noise and uncertainty. Furthermore, their VLP feature extractors like CLIP often focus on instance-level representations, potentially reducing the discriminability of fine-grained local features. To alleviate these issues, we propose an Explicit Key Local information Selection and Reconstruction Framework (EKLSR), which explicitly selects key local information to enhance feature representation. Specifically, we introduce a Key Local information Selection and Fusion (KLSF) that utilizes hidden knowledge from the VLP model to select interpretably and fuse key local information. Secondly, we employ Key Local segment Reconstruction (KLR) based on multimodal interaction to reconstruct the key local segments of images (text), significantly enriching their discriminative information and enhancing both inter-modal and intra-modal interaction alignment. To demonstrate the effectiveness of our approach, we conducted experiments on five datasets across TIReID and RSITR. Notably, our EKLSR model achieves state-of-the-art performance on two RSITR datasets.
Yu Liao, Rui Yang 0038, Jianwei Tao, Bai Liu 0002, Zhipeng Hu, Shuang Wang 0001, Zeng Zhao
ACM Multimedia7
2024 Accurate and Lightweight Learning for Specific Domain Image-Text Retrieval
abstract
Recent advances in vision-language pre-trained models like CLIP have greatly enhanced general domain image-text retrieval performance. This success has led scholars to develop methods for applying CLIP to Specific Domain Image-Text Retrieval (SDITR) tasks such as Remote Sensing Image-Text Retrieval (RSITR) and Text-Image Person Re-identification (TIReID). However, these methods for SDITR often neglect two critical aspects: the enhancement of modal-level distribution consistency within the retrieval space and the reduction of CLIP's computational cost during inference. To address these issues, this paper presents a novel framework, Accurate and lightweight learning for specific domain Image-text Retrieval (AIR), based on the CLIP. AIR incorporates a Modal-Level distribution Consistency Enhancement regularization (MLCE) loss and a Self-Pruning Distillation Strategy (SPDS) to improve retrieval precision and computational efficiency. The MLCE loss harmonizes the sample distance distributions within image and text modalities, fostering a retrieval space closer to the ideal state. SPDS employs a strategic knowledge distillation process to transfer deep multimodal insights from CLIP to a shallower level, maintaining only the essential layers for inference, thus achieving model light-weighting. Comprehensive experiments across various datasets in RSITR and TIReID reveal that MLCE loss secures optimal retrieval, while SPDS achieves a favorable balance between accuracy and computational demand during testing.
Rui Yang 0038, Shuang Wang 0001, Jianwei Tao, Yingping Han, Qiaoling Lin, Yanhe Guo, Biao Hou, Licheng Jiao
ACM Multimedia2
2024 Generalized Source-Free Domain-adaptive Segmentation via Reliable Knowledge Propagation
Qi Zang, Shuang Wang 0001, Dong Zhao 0007, Dou Quan, Jinlong Li 0003, Nicu Sebe, Zhun Zhong
ACM Multimedia2
2024 Connectivity-Driven Pseudo-Labeling Makes Stronger Cross-Domain Segmenters
abstract
Presently, pseudo-labeling stands as a prevailing approach in cross-domain semantic segmentation, enhancing model efficacy by training with pixels assigned with reliable pseudo-labels. However, we identify two key limitations within this paradigm: (1) under relatively severe domain shifts, most selected reliable pixels appear speckled and remain noisy. (2) when dealing with wild data, some pixels belonging to the open-set class may exhibit high confidence and also appear speckled. These two points make it difficult for the pixel-level selection mechanism to identify and correct these speckled close- and open-set noises. As a result, error accumulation is continuously introduced into subsequent self-training, leading to inefficiencies in pseudo-labeling. To address these limitations, we propose a novel method called Semantic Connectivity-driven Pseudo-labeling (SeCo). SeCo formulates pseudo-labels at the connectivity level, which makes it easier to locate and correct closed and open set noise. Specifically, SeCo comprises two key components: Pixel Semantic Aggregation (PSA) and Semantic Connectivity Correction (SCC). Initially, PSA categorizes semantics into ``stuff'' and ``things'' categories and aggregates speckled pseudo-labels into semantic connectivity through efficient interaction with the Segment Anything Model (SAM). This enables us not only to obtain accurate boundaries but also simplifies noise localization. Subsequently, SCC introduces a simple connectivity classification task, which enables us to locate and correct connectivity noise with the guidance of loss distribution. Extensive experiments demonstrate that SeCo can be flexibly applied to various cross-domain semantic segmentation tasks, \textit{i.e.} domain generalization and domain adaptation, even including source-free, and black-box domain adaptation, significantly improving the performance of existing state-of-the-art methods. The code is provided in the appendix and will be open-source.
Dong Zhao 0007, Qi Zang, Shuang Wang 0001, Nicu Sebe, Zhun Zhong
NeurIPS3
2024 A deep learning-based approach for pseudo-satellite positioning
abstract
Abstract Traditional pseudo‐satellite‐based indoor positioning techniques are greatly affected by the presence of multipath effects, leading to a notable reduction in the positioning precision. In order to tackle this challenge, a pseudo‐satellite indoor positioning method based on deep learning is proposed. The method grids the localization region, thus transforming positioning from a regression problem to a classification problem in the gridded areas. 1D‐convolutional neural network is employed to extract the correlation between pseudo‐satellite data and the positioning of indoor areas. Data are collected and the method is validated in three types of areas of the experimental field, namely unobstructed area, semi‐unobstructed area and obstructed area. The experimental results demonstrate that the method exhibits superior positioning accuracy compared to traditional methods, enabling effective localization even in obstructed area.
Baoguo Yu, Hantong Xing, Shuang Wang 0001
IET Commun.5
2024 Continual learning for cross-modal image-text retrieval based on domain-selective attention
Rui Yang 0038, Shuang Wang 0001, Yu Gu 0015, Jihui Wang, Yingzhi Sun, Yu Liao, Licheng Jiao
Pattern Recognit.2
2024 LM-Net: A Lightweight Matching Network for Remote Sensing Image Matching and Registration
abstract
Deep feature learning methods have shown significant advantages over handcrafted feature-based methods in remote sensing image matching and registration. Existing deep learning methods usually introduce complex modules into the deep convolutional network for more robust feature learning. However, they usually require high computation and memory resources for the computing device and have expensive time costs for image registration. As a basic image-processing task, it is crucial to build a lightweight matching network (LM-Net) for fast and accurate image matching and registration. Unfortunately, the image-matching performance will decrease significantly when we directly compress the deep model to a lightweight one. This article proposes an LM-Net based on the knowledge distillation (KD) learning framework for remote sensing image matching and registration. We first build an LM-Net with three convolutional layers. Then, this article proposes an effective KD approach for network optimization, which transfers the effective knowledge from the deep matching network to LM-Net to improve image-matching performances. Specifically, this article considers the useful information in the instance samples and the relation information between samples. It designs the feature and feature relation distillation learning for LM-Net training. Extensive experimental results and analysis have shown the effectiveness and advantages of the proposed LM-Net. LM-Net can reduce the number of parameters and computational complexity of the matching network. Meanwhile, LM-Net can significantly decrease the time cost and achieve results comparable to those of the deep model. It reduces the average image registration time by 42% on remote sensing image matching and registration. Additionally, LM-Net generalizes well on other multimodal remote sensing images.
Dou Quan, Chonghua Lv, Shuang Wang 0001, Yi Li 0054, Bo Ren 0001, Jocelyn Chanussot, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.4
2024 F3Net: Adaptive Frequency Feature Filtering Network for Multimodal Remote Sensing Image Registration
abstract
Multimodal remote sensing image registration is crucial for multimodal information fusion and applications. The significant nonlinear appearance difference between multimodal images caused by the various imaging mechanisms dramatically increases the challenge of image registration. This article proposes an adaptive frequency feature filtering network (F3Net) for cross-modal remote sensing image registration. On the one hand, F3Net explicitly explores the useful frequency components across modal images based on multilevel deep features. On the other hand, F3Net can take advantage of the nonlocal receptive fields by frequency modulation for feature learning and boosting image registration performances. F3Net inserts frequency feature filtering (F3) modules in multilevel deep features. Specifically, F3Net first performs the fast Fourier transform (FFT) for deep features. Then, F3Net designs a frequency attention (FA) module to adaptive enhance the shared and discriminative frequency features between multimodal images while suppressing the frequency components that hinder the cross-modal image registration. In addition, F3Net adopts multiscale frequency filtering fusion to facilitate discriminative feature learning, including global frequency feature filtering (GF3) based on the global image spectrum and local frequency feature filtering (LF3) based on the spectrum of stacked image regions. Experimental results on many remote sensing images have demonstrated the efficiency of the F3Net on multimodal image registration.
Dou Quan, Shuang Wang 0001, Yunan Li 0001, Bo Ren 0001, Mengte Kang, Jocelyn Chanussot, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.3
2024 Multi-View Feature Fusion and Visual Prompt for Remote Sensing Image Captioning
abstract
Remote sensing image (RSI) captioning is a vision-language multimodal task concentrating on both image comprehension and sentence generation. Several studies suggest that encoder–decoder-based methods have achieved success in RSI captioning. However, existing encoder–decoder-based methods may not fully explore image representations for RSI captioning and suffer from a lack of additional prompt information for sentence generation. In this article, a novel multi-view feature fusion and prompt (MVP)-based model is proposed to obtain better RSI representations and enhance language model performance in RSI captioning. Specifically, we design an attention-based feature fusion module to dynamically fuse multi-view visual features, which are extracted from the fine-tuned vision-language pretraining (VLP) model and the vision-task pretraining (VP) model. Then, a flexible visual prefix mapping module is proposed to transform images into visual prefixes, providing semantic information for the subsequent sentence generation. Finally, a BERT-based caption generator is applied to generate accurate descriptions based on the fused visual features and the visual prefixes, which are both outputs from our designed modules. Extensive experiments are conducted on three well-known benchmark datasets, demonstrating that our method achieves state-of-the-art (SOTA) performance. The relevant code is available athttps://github.com/QiaoLing-Lin/MVP.
Shuang Wang 0001, Qiaoling Lin, Xiutiao Ye, Yu Liao, Dou Quan, ZhongQian Jin, Biao Hou, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.1
2024 Transcending Fusion: A Multiscale Alignment Method for Remote Sensing Image-Text Retrieval
abstract
Remote sensing image-text retrieval (RSITR) is pivotal for knowledge services and data mining in the remote sensing (RS) domain. Considering the multiscale representations in image content and text vocabulary can enable the models to learn richer representations and enhance retrieval. Current multiscale RSITR approaches typically align multiscale fused image features with text features but overlook aligning image-text pairs at distinct scales separately. This oversight restricts their ability to learn joint representations suitable for effective retrieval. We introduce a novel multiscale alignment (MSA) method to overcome this limitation. Our method comprises three key innovations: 1) a multiscale cross-modal alignment transformer (MSCMAT), which computes cross-attention between single-scale image features and localized text features, integrating global textual context to derive a matching score matrix within a mini-batch; 2) a multiscale cross-modal semantic alignment loss (MSCMA loss) that enforces semantic alignment across scales; and 3) a cross-scale multimodal semantic consistency loss (CSMMC loss) that uses the matching matrix from the largest scale to guide alignment at smaller scales. We evaluated our method across multiple datasets, demonstrating its efficacy with various visual backbones and establishing its superiority over existing state-of-the-art methods. The GitHub URL for our project ishttps://github.com/yr666666/MSA.
Rui Yang 0038, Shuang Wang 0001, Yingping Han, Yuanheng Li, Dong Zhao 0007, Dou Quan, Yanhe Guo, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.2
2024 A Unified Deep Learning Network for Remote Sensing Image Registration and Change Detection
abstract
Image registration and change detection are crucial for multitemporal remote sensing image analysis. The images should be registered before the change information detection. Existing deep learning methods have shown significant advantages in image registration and change detection tasks. They usually design two independent task-specific deep networks for image registration and change detection, respectively. These independent deep networks will learn from scratch and rely on many task-specific labeled training datasets. This article finds that image registration and change detection have similar learning mechanisms, which focus on extracting discriminative features. Inspired by this, we propose a Unified image Registration and Change detection Network (URCNet) that can perform image alignment and change information detection through a single network. Additionally, this article proposes various deep collaborative learning methods for URCNet optimization, which enforce that the URCNet can effectively support remote sensing image registration and change detection simultaneously. Extensive experiments demonstrate the effectiveness of the proposed URCNet for image registration and change detection, which can achieve comparable and better results with task-specific and more complex deep networks. The proposed URCNet can support multitasks based on the same scene images, different scene images, and even multimodal images. Moreover, URCNet shows significant advantages over other deep networks in change detection under limited labeled datasets.
Rufan Zhou, Dou Quan, Shuang Wang 0001, Chonghua Lv, Xianwei Cao, Jocelyn Chanussot, Yi Li 0054, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.3
2024 Few-Shot MS and PAN Joint Classification With Improved Cross-Source Contrastive Learning
abstract
The joint classification of multispectral (MS) and panchromatic (PAN) images aims to provide a more detailed and accurate interpretation of land features. Although deep-learning-based methods have achieved remarkable success in this task, the generalization performance of networks is compromised when labeled samples are insufficient. In this study, we explore the possibility of leveraging unlabeled remote sensing images (RSIs) through contrastive learning and demonstrate the challenges associated with directly applying contrastive learning to RSIs. To end this, we propose a cross-source contrastive learning method for few-shot MS and PAN joint classification (CrossCLMP), which aims to learn sufficient transferable representations in a self-supervised contrastive manner so as to provide a robust pretrained model for fine-tuning the downstream joint classification task. Specifically, we design: 1) intersource and intrasource alignment loss (ER-Align) to achieve self-supervised feature extraction and alignment; 2) the source-unique feature adaptive separation (SUAS) strategy to model source-unique information explicitly; and 3) the auxiliary contrastive learning (ACL) strategy to mitigate the adverse impact of numerous false-negative samples in the pretraining stage. The experimental results and the theoretical analyses on multiple popular datasets comprehensively demonstrate the effectiveness and robustness of the proposed method under few-shot. Our code is available at:https://github.com/Xidian-AIGroup190726/CrossCLMP.
Hao Zhu 0009, Pute Guo, Biao Hou, Changzhe Jiao, Bo Ren 0001, Licheng Jiao, Shuang Wang 0001
IEEE Trans. Geosci. Remote. Sens.8
2024 High-Low-Frequency Progressive-Guided Diffusion Model for PAN and MS Classification
abstract
With the rapid development of remote sensing technology, satellites can easily acquire multispectral (MS) and panchromatic (PAN) images. It is challenging to utilize their complementarity to effectively combine each other’s advantages and mitigate the differences between different modes. In this article, we propose a high-low-frequency progressive-guided diffusion model. It is used to generate an image with the advantages of both MS and PAN, which can be complementary to MS and PAN and, thus, can better reduce the modal differences between them. Therefore, we use the fusion image as an auxiliary mode and an intermediate bridge, which can better connect the characteristics between various sources. First, we design guidance information that contains the advantages of MS and PAN, and some operations can make this information better guide the generation stage. In addition, we design a high-low-frequency progressive guidance strategy; by using this strategy, we can first ensure the overall structure and layout of the image in the generation stage and then refine the local details and features of the image. This dramatically improves the quality of the generated image. Finally, we use mathematical knowledge to explain the rationality of the strategy. We validate our method on multiple datasets and achieve the best performance. Our code ishttps://github.com/Xidian-AIGroup190726/HLF-GDiffusion.
Hao Zhu 0009, Fengtao Yan, Pute Guo, Biao Hou, Shuang Wang 0001, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.7
2024 Cross-Domain Scene Unsupervised Learning Segmentation With Dynamic Subdomains
abstract
Unsupervised cross-domain scene segmentation approach adapts the source model to the target domain, which utilizes two-stage strategies to minimize the inter-domain and intra-domain gap. However, the accumulation of errors in the previous stages affects the training of the subsequent stages. In this paper, a framework called statistical and structural domain adaptation (SSDA) is proposed to optimize inter-domain and intra-domain adaptation jointly. Firstly, the statistical inter-domain adaptation (StaIA) is proposed to model dynamic subdomains, which continuously adjust seed samples during the process of domain adaptation to mitigate error accumulation. The dynamic subdomains are modeled by exploring Bayesian uncertainty statistics and global balance statistics, which alleviate the imbalance problem in uncertainty estimation. StaIA encourages the model to transfer comprehensive and genuine knowledge through the seed loss for inter-domain adaptation. Secondly, the structural intra-domain adaptation (StrIA) is proposed to align the intra-domain gap among dynamic subdomains by the structural priors. Specifically, the StrIA models structural priors by truncated conditional random field (TruCRF) loss within the neighborhood, which constrains intra-domain semantic consistency to reduce the intra-domain gap. Experimental results demonstrate the effectiveness of the proposed cross-domain scene segmentation approaches on two commonly-used unsupervised domain adaptation benchmarks. The code is available at https://github.com/ChicalH/SSDA.
Pei He, Licheng Jiao, Fang Liu 0001, Xu Liu 0006, Ronghua Shang, Shuang Wang 0001
IEEE Trans. Multim.6
2024 Multi-Scale Contourlet Knowledge Guide Learning Segmentation
abstract
For accurate segmentation, effective feature extraction has always been a challenging problem, since the variability of appearance and the fuzziness of object boundaries. Convolutional neural networks have recently gained recognition in feature representation learning. However, it is only conducted in the spatial domain, and lacks effective representation of directionality, singularity and regularity in the spectral domain for anomaly detection of images. This is the key to feature learning representation of high-order singularity. To solve this problem, a multi-scale contourlet knowledge guide learning network is proposed in this paper. It is novel in this sense that, different from the CNNs in the spatial domain, the proposed method learns the multi-scale contourlet sparse representation to obtain more effective and sparse features in multi-scales and multi-directions. Furthermore, the contourlet knowledge guide learning can enhance the representation of spectral domain features. It is shown that the proposed network can learn the multi-level discriminative features and capture the more accurate object boundaries. The segmentation ability in theoretical analysis and experiments on five polyp segmentation datasets (CVC-ColonDB, CVC-ClinicDB, Kvasir-SEG, ETIS-LaribPolypDB, EndoSceneStill) and two building datasets (Massachusetts, WHU) are compared with developed methods. It must be emphasized that there is potential in effective feature learning representation and the generalization capability of the proposed method in deep learning, recognition and interpretation.
Mengkun Liu, Licheng Jiao, Xu Liu 0006, Lingling Li 0002, Fang Liu 0001, Shuyuan Yang 0001, Shuang Wang 0001, Biao Hou
IEEE Trans. Multim.7
2024 A Patch Diversity Transformer for Domain Generalized Semantic Segmentation
abstract
Domain generalization (DG) is one of the critical issues for deep learning in unknown domains. How to effectively represent domain-invariant context (DIC) is a difficult problem that DG needs to solve. Transformers have shown the potential to learn generalized features, since the powerful ability to learn global context. In this article, a novel method named patch diversity Transformer (PDTrans) is proposed to improve the DG for scene segmentation by learning global multidomain semantic relations. Specifically, patch photometric perturbation (PPP) is proposed to improve the representation of multidomain in the global context information, which helps the Transformer learn the relationship between multiple domains. Besides, patch statistics perturbation (PSP) is proposed to model the feature statistics of patches under different domain shifts, which enables the model to encode domain-invariant semantic features and improve generalization. PPP and PSP can help to diversify the source domain at the patch level and feature level. PDTrans learns context across diverse patches and takes advantage of self-attention to improve DG. Extensive experiments demonstrate the tremendous performance advantages of the PDTrans over state-of-the-art DG methods.
Pei He, Licheng Jiao, Ronghua Shang, Xu Liu 0006, Fang Liu 0001, Shuyuan Yang 0001, Xiangrong Zhang, Shuang Wang 0001
IEEE Trans. Neural Networks Learn. Syst.8
2024 A Concurrent Multiscale Detector for End-to-End Image Matching
abstract
This article focuses on end-to-end image matching through joint key-point detection and descriptor extraction. To find repeatable and high discrimination key points, we improve the deep matching network from the perspectives of network structure and network optimization. First, we propose a concurrent multiscale detector (CS-det) network, which consists of several parallel convolutional networks to extract multiscale features and multilevel discriminative information for key-point detection. Moreover, we introduce an attention module to fuse the response maps of various features adaptively. Importantly, we propose two novel rank consistent losses (RC-losses) for network optimization, significantly improving image matching performances. On the one hand, we propose a score rank consistent loss (RC-S-loss) to ensure that the key points have high repeatability. Different from the score difference loss merely focusing on the absolute score of an individual key point, our proposed RC-S-loss pays more attention to the relative score of key points in the image. On the other hand, we propose a score-discrimination RC-loss to ensure that the key point has high discrimination, which can reduce the confusion from other key points in subsequent matching and then further enhance the accuracy of image matching. Extensive experimental results demonstrate that the proposed CS-det improves the mean matching result of deep detector by 1.4%-2.1%, and the proposed RC-losses can boost the matching performances by 2.7%-3.4% than score difference loss. Our source codes are available at https://github.com/iquandou/CS-Net.
Dou Quan, Shuang Wang 0001, Ning Huyan, Yi Li 0054, Ruiqi Lei, Jocelyn Chanussot, Biao Hou, Licheng Jiao
IEEE Trans. Neural Networks Learn. Syst.2
2024 Select, Purify, and Exchange: A Multisource Unsupervised Domain Adaptation Method for Building Extraction
abstract
Accurately extracting buildings from aerial images has essential research significance for timely understanding human intervention on the land. The distribution discrepancies between diversified unlabeled remote sensing images (changes in imaging sensor, location, and environment) and labeled historical images significantly degrade the generalization performance of deep learning algorithms. Unsupervised domain adaptation (UDA) algorithms have recently been proposed to eliminate the distribution discrepancies without re-annotating training data for new domains. Nevertheless, due to the limited information provided by a single-source domain, single-source UDA (SSUDA) is not an optimal choice when multitemporal and multiregion remote sensing images are available. We propose a multisource UDA (MSUDA) framework SPENet for building extraction, aiming at selecting, purifying, and exchanging information from multisource domains to better adapt the model to the target domain. Specifically, the framework effectively utilizes richer knowledge by extracting target-relevant information from multiple-source domains, purifying target domain information with low-level features of buildings, and exchanging target domain information in an interactive learning manner. Extensive experiments and ablation studies constructed on 12 city datasets prove the effectiveness of our method against existing state-of-the-art methods, e.g., our method achieves 59.1% intersection over union (IoU) on Austin and Kitsap → Potsdam, which surpasses the target domain supervised method by 2.2%. The code is available at https://github.com/QZangXDU/SPENet.
Shuang Wang 0001, Qi Zang, Dong Zhao 0007, Chaowei Fang, Dou Quan, Yutong Wan, Yanhe Guo, Licheng Jiao
IEEE Trans. Neural Networks Learn. Syst.1
2024 SigDA: A Superimposed Domain Adaptation Framework for Automatic Modulation Classification
abstract
Due to the uncertainty of non-cooperative communication channels, the received signals often contain various impairment factors, leading to a significant decline in the performance of existing deep learning (DL)-based automatic modulation classification (AMC) models. Several preliminary works utilize domain adaptation (DA) to alleviate this issue, however, they are constrained by singular domain difference factor, whereas in practice, these factors often manifest cumulatively. Therefore, this paper introduce a more realistic task named superimposed DA, where multiple domain difference factors are overlaid, reflecting the cumulative nature of them. We propose the SigDA as a solution framework, which adopts adversarial training to align the data distribution in different domains. Two technical modules, Multi-task based Masked Signal Feature Extractor (M2SFE) and Signal Feature Pyramid Aggregation (SFPA), are innovatively designed in SigDA. M2SFE utilizes mask and reconstruction task to enhance feature extraction and achieves discriminative feature selection through the design of feature mapping layers, while SFPA can solve the problem of inconsistent signal length in superimposed DA and can aggregate the features of signals into the same dimension. We consider and superimpose various typical signal domain difference factors, comprehensive experiments demonstrate that the proposed framework can achieve significant performance improvement in various communication channels.
Shuang Wang 0001, Hantong Xing, Chenxu Wang 0001, Huaji Zhou, Biao Hou, Licheng Jiao
IEEE Trans. Wirel. Commun.1
2023 Towards Better Stability and Adaptability: Improve Online Self-Training for Model Adaptation in Semantic Segmentation
abstract
Unsupervised domain adaptation (UDA) in semantic segmentation transfers the knowledge of the source domain to the target one to improve the adaptability of the segmentation model in the target domain. The need to access labeled source data makes UDA unable to handle adaptation scenarios involving privacy, property rights protection, and confidentiality. In this paper, we focus on unsupervised model adaptation (UMA), also called source-free domain adaptation, which adapts a source-trained model to the target domain without accessing source data. We find that the online self-training method has the potential to be deployed in UMA, but the lack of source domain loss will greatly weaken the stability and adaptability of the method. We analyze two reasons for the degradation of online self-training, i.e. inopportune updates of the teacher model and biased knowledge from the source-trained model. Based on this, we propose a dynamic teacher update mechanism and a training-consistency based resampling strategy to improve the stability and adaptability of online self-training. On multiple model adaptation benchmarks, our method obtains new state-of-the-art performance, which is comparable or even better than state-of-the-art UDA methods. The code is available at https://github.com/DZhaoXd/DT-ST.
Dong Zhao 0007, Shuang Wang 0001, Qi Zang, Dou Quan, Xiutiao Ye, Licheng Jiao
CVPR2
2023 Learning Pseudo-Relations for Cross-domain Semantic Segmentation
abstract
Domain adaptive semantic segmentation aims to adapt a model trained on labeled source domain to unlabeled target domain. Self-training shows competitive potential in this field. Existing methods along this stream mainly focus on selecting reliable predictions on target data as pseudo-labels for category learning, while ignoring the useful relations between pixels for relation learning. In this paper, we propose a pseudo-relation learning framework, Relation Teacher (RTea), which can exploitable pixel relations to efficiently use unreliable pixels and learn generalized representations. In this framework, we build reasonable pseudo-relations on local grids and fuse them with low-level relations in the image space, which are motivated by the reliable local relations prior and available low-level relations prior. Then, we design a pseudo-relation learning strategy and optimize the class probability to meet the relation consistency by finding the optimal sub-graph division. In this way, the model’s certainty and consistency of prediction are enhanced on the target domain, and the cross-domain inadaptation is further eliminated. Extensive experiments on three datasets demonstrate the effectiveness of the proposed method. The code will be available at https://github.com/DZhaoXd/RTea.
Dong Zhao 0007, Shuang Wang 0001, Qi Zang, Dou Quan, Xiutiao Ye, Rui Yang 0038, Licheng Jiao
ICCV2
2023 Relational Image Patch Matching for Remote Sensing
abstract
Feature descriptor-based methods have demonstrated remarkable performance in remote sensing image patch matching tasks and are usually optimized using contrastive loss and triplet loss. However, these optimization losses focus on calculating the distance between samples, ignoring the rich information of higher-order feature relationships between multiple image patches. The latter provides valuable information that can be used to improve task performance. Inspired by the superior performance of second-order relations in graph matching and clustering tasks, we aim to exploit the rich information available from high-order relations fully. This paper proposes a high-order relationship (HOR) learning method for remote sensing image patch matching. This method combines low-order feature relations between image patch pairs and high-order feature relations between multiple patches to enhance image matching performance. Extensive experimental results on a multimodel remote sensing image dataset, SEN 1-2, consisting of optical and SAR images, demonstrate that the proposed HOR learning method can improve the performance of remote sensing image patch matching.
Xianwei Cao, Dou Quan, Chonghua Lv, Yanhe Guo, Shuang Wang 0001, Biao Hou, Licheng Jiao
IGARSS5
2023 A Fast and Accurate Method for Remote Sensing Image-Text Retrieval Based On Large Model Knowledge Distillation
abstract
With the increasing development of remote sensing (RS) technology, remote sensing cross-modal image-text retrieval (RSCMITR) task has gradually attracted wide attention. At present, the large-scale pre-training model is brilliant in the field of natural images cross-modal retrieval, but the current RSCMITR models do not focus on it, resulting in less retrieval performance improvement. This paper proposes a lightweight network structure based on large-scale pre-training model and knowledge distillation, designing a lightweight model based on separable convolution and text convolution. Knowledge distillation technology is used to make the Light model learn the hidden knowledge of large-scale model CLIP-RS, which realizes fast and accurate retrieval. The proposed method achieves state-of-the-art performance on four commonly used RSCMITR datasets.
Yu Liao, Rui Yang 0038, Hantong Xing, Dou Quan, Shuang Wang 0001, Biao Hou
IGARSS6
2023 Domain Distribution Alignment for Boosting Multi-Modal Remote Sensing Image Matching
abstract
Multi-modal images can obtain complementary and rich information images, which are more widely used in various applications. However, due to the different imaging mechanisms of different sensors, there are significant domain distribution differences between multi-modal images. In multi-modal image matching, existing deep learning methods should deal with the image content difference caused by rotation transformation and the domain distribution difference caused by different sensors, which are very difficult for the deep network. To address this issue, we propose to combine an instance comparison and a batch comparison to deal with image content differences and domain distribution differences, respectively. We design a new domain distribution alignment method to explicitly constrain the sample domain distribution of the multi-modal images are consistent through the domain distribution alignment loss. Extensive multi-modal remote sensing image patch matching experiments have shown the effectiveness of the proposed method. Furthermore, the proposed multi-modal domain distribution alignment method has more obvious advantages when there are significant content differences and distribution differences.
Dou Quan, Chonghua Lv, Yanhe Guo, Shuang Wang 0001, Yu Gu 0015, Licheng Jiao
IGARSS5
2023 A Texture and Saliency Enhanced Image Learning Method For Cross-Modal Remote Sensing Image-Text Retrieval
abstract
Cross-modal remote sensing image-text retrieval (CMRSITR) can retrieve images of interest from a vast amount of remote sensing images and has received significant attention in recent years. However, existing methods do not consider saliency and texture information, which are essential for remote sensing images when extracting image features. Therefore, this paper proposes a novel texture and saliency enhanced image learning method for CMRSITR. We constructed a multi-task image feature extractor in this new method. A texture map and a saliency map are created by extracting texture and detecting the saliency of each RS image. Both maps are set as supervised information during training to make the extracted saliency and texture features gradually reconstructed to a saliency map and a texture map, respectively. At the same time, the retrieval features of each RS image are obtained from the retrieval feature branch of the image. Experiments conducted on two commonly used CMRSITR datasets, RSICD and UCM, showed that the proposed method is effective in improving retrieval performance and achieved state-of-the-art retrieval performance compared to existing methods.
Rui Yang 0038, Yanhe Guo, Shuang Wang 0001
IGARSS4
2023 Deep Continuous Matching Network for more Robust Multi-Modal Remote Sensing Image Patch Matching
abstract
Due to the powerful feature extraction capabilities of deep neural networks, traditional approaches are gradually replaced by deep learning approaches for image matching tasks. For multi-modal image patch matching, the deep model should mainly learn the modality-invariant features. For multi-modal images with rotation transformation (RT), the deep model should learn the modality-invariant features and rotation-invariant features simultaneously. However, the performance of the latter trained model is degraded for the former task. The main reason is that the modality invariance of the features degenerates. This paper proposes a deep multi-modal remote sensing image matching network (DCMNet) that combines descriptor learning and continuous learning to solve this problem. Firstly, DCMNet is trained for learning modality-invariant features in multi-modal image patch matching. Then, DCMNet is optimized for multi-modal image patch matching with RT. In the later learning process, we reduce the change of important parameters for the modality-invariant features learning. Experiments demonstrate the effectiveness and robustness of DCMNet in alleviating the modal invariance degradation problem of features.
Rufan Zhou, Dou Quan, Chonghua Lv, Yanhe Guo, Shuang Wang 0001, Yu Gu 0015, Licheng Jiao
IGARSS5
2023 Knowledge Decomposition and Replay: A Novel Cross-modal Image-Text Retrieval Continual Learning Method
abstract
To enable machines to mimic human cognitive abilities and alleviate the catastrophic forgetting problem in cross-modal image-text retrieval (CMITR), this paper proposes a novel continual learning method, Knowledge Decomposition and Replay (KDR), which emulates the process of knowledge decomposition and replay exhibited by humans in complex and changing environments. KDR has two components: a feature Decomposition-based CMITR Model (DCM) and a cross-task Generic Knowledge Replay strategy (GKR). DCM decomposes text and image features into task-specific and generic knowledge features, mimicking the human cognitive process of knowledge decomposition. Specifically, it employs a generic knowledge features extraction module for all tasks and a task-specific module for each task with a few trainable fully connected layers. Similarly, GKR emulates the human behavior of knowledge replay by utilizing the image-text similarity matrix output from the old task model with inputting the previous samples to induce the learning of the image-text similarity matrix output from the current task model with inputting the previous samples, using knowledge distillation technology. To demonstrate the effect of KDR, we adapted a continual learning dataset Seq-COCO from MSCOCO. Extensive experiments on Seq-COCO showed that KDR reduces catastrophic forgetting and consolidates general knowledge, improving the model's learning ability in CMITR.
Rui Yang 0038, Shuang Wang 0001, Yanhe Guo, Xiutiao Ye, Biao Hou, Licheng Jiao
ACM Multimedia2
2023 A Collaborative Learning Tracking Network for Remote Sensing Videos
abstract
With the increasing accessibility of remote sensing videos, remote sensing tracking is gradually becoming a hot issue. However, accurately detecting and tracking in complex remote sensing scenes is still a challenge. In this article, we propose a collaborative learning tracking network for remote sensing videos, including a consistent receptive field parallel fusion module (CRFPF), dual-branch spatial-channel co-attention (DSCA) module, and geometric constraint retrack strategy (GCRT). Considering the small-size objects of remote sensing scenes are difficult for general forward networks to extract effective features, we propose a CRFPF-module to establish parallel branches with consistent receptive fields to separately extract from shallow to deep features and then fuse hierarchical features adaptively. Since the objects and their background are difficult to distinguish, the proposed DSCA-module uses the spatial-channel co-attention mechanism to collaboratively learn the relevant information, which enhances the saliency of the objects and regresses to precise bounding boxes. Considering the interference of similar objects, we designed a GCRT-strategy to judge whether there is a false detection through the estimated motion trajectory and then recover the correct object by weakening the feature response of interference. The experimental results and theoretical analysis on multiple datasets demonstrate our proposed method's feasibility and effectiveness. Code and net are available at https://github.com/Dawn5786/CoCRF-TrackNet.
Licheng Jiao, Hao Zhu 0009, Fang Liu 0001, Shuyuan Yang 0001, Xiangrong Zhang, Shuang Wang 0001, Rong Qu
IEEE Trans. Cybern.7
2023 SSMU-Net: A Style Separation and Mode Unification Network for Multimodal Remote Sensing Image Classification
abstract
The rapid progress in remote sensing technology has made it convenient for satellites to capture both multispectral (MS) and panchromatic (PAN) images. MS has more spectral information, and PAN has higher spatial resolution. How to exploit the complementarity between MS and PAN images, and effectively combine their respective advantageous features while alleviating mode differences, has become a crucial research task. This paper designs a Style Separation and Mode Unification network (SSMU-Net) for MS and PAN image classification from a novel and effective perspective. The network can be divided into two stages: style separation and mode unification. In the style separation stage, we use wavelet decomposition and techniques similar to generative adversarial networks to preliminarily separate the information of MS and PAN into different components. These components better preserve complete information from the original data and have their own advantages in style and content. Then we propose a Symmetrical Triplet Traction module to perform style traction on different components, making style features more unique and content features more unified, achieving feature separation and purification. In the mode unification stage, we design an encoder-decoder model to reduce the impact of mode differences. The experimental results from multiple datasets validate the effectiveness of our proposed method. Our overall accuracy improved by approximately 4% on the Shanghai and Beijing datasets, and it has exceeded 99.28% on the Hohhot and Vancouver datasets. Our code is available at: https://github.com/proudpie/SSMU-Net.
Hao Zhu 0009, Licheng Jiao, Xiaoyu Yi 0002, Biao Hou, Wenping Ma 0001, Shuang Wang 0001
IEEE Trans. Geosci. Remote. Sens.8
2023 Contrastive Learning Based on Multiscale Hard Features for Remote-Sensing Image Scene Classification
abstract
The overwhelming majority of models for remote sensing image (RSI) scene classification generally require the weights pre-trained on natural images for initialization before formal training. However, differences in imaging mechanisms lead to huge discrepancies between natural images and RSIs, and the strong visual representation learned from massive natural images limits the performance of models when inferencing RSIs. To address this issue, the well-established self-supervised contrastive learning paradigm in the natural image field is introduced to the RSI field. We propose a contrastive learning method based on multi-scale hard features, MHCL, which aims to use finite RSIs to learn sufficient visual representations in an unsupervised contrastive manner, thus provide a powerful upstream pre-trained model for fine-tuning downstream scene classification task. Multi-level features extracted by intermediate layers of each encoder’s backbone are first gathered, and then a hard features transformation method is proposed to create hard positive features and diverse queues that save hard negatives, thereby enriching the finite scene information in small-scale RSIs. Furthermore, we redesign the multi-scale hard features joint contrastive loss to boost the model to explore sufficient invariant representations by additionally pulling hard positive pairs closer and pushing hard negative pairs farther away in the embedding space. Extensive experiments demonstrate that the upstream pre-training model generated by MHCL achieves competitive transferred performance on three popular scene classification datasets, outperforming the traditional model pre-trained on ImageNet and models pre-trained by other state-of-the-art contrastive learning methods. Our code will be released at: https://github.com/benesakitam/MHCL.
Zhihao Li 0005, Biao Hou, Xianpeng Guo, Siteng Ma, Yanyu Cui, Shuang Wang 0001, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.6
2023 A Novel Coarse-to-Fine Deep Learning Registration Framework for Multimodal Remote Sensing Images
abstract
Multi-modal remote sensing images with large rotation transformation (RT) are challenging to be registered. It needs to deal with the global geometric deformation caused by great RT and significant local appearance differences caused by different imaging mechanisms. Existing deep learning methods mainly use a single deep descriptor learning (DDL) network to extract invariant features for identifying matching samples and discriminative feature descriptors for separating non-matching samples. However, it is difficult to extract local invariant feature descriptors to RT and modality change through a single DDL network. This paper proposes a novel coarse-to-fine deep learning image registration framework for multi-modal remote sensing images based on two task-specific deep models. Specifically, in the coarse registration stage, this paper designs an effective deep ordinal regression (DOR) network for rotation correction, which can reduce the difficulty of multi-modal image registration and boost image registration. The proposed DOR network transforms the rotation correction task into a rotation ordinal regression problem, which can exploit the potential relationship between the rotation ordinals to improve the accuracy of rotation estimation. In the fine registration stage, we adopt the DDL network to deal with the image modality change based on the rotation-corrected images. Extensive experimental results on multi-modal image datasets demonstrate the significant advantages of the proposed coarse-to-fine deep learning registration framework. The DOR network achieves higher rotation correction accuracy, which can significantly improve the multi-modal image registration performances.
Dou Quan, Huiyuan Wei, Shuang Wang 0001, Yu Gu 0015, Biao Hou, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.3
2022 Deep Modality Independent Descriptor Learning for Optical and SAR Image Patch Matching
abstract
Due to the complementary information between multi-modal images, they are widely used in various applications. However, there are significant differences in appearance caused by different imaging mechanisms, which bring great challenges to multi-modal image patch matching. To solve this problem, this paper proposes a deep modality independent descriptor learning network (DMID-Net) for multi-modal image patch matching. DMID-Net computes the self-similarity of deep features as the structure descriptor for image patch matching, which is independent of image modality and shared between multi-modal images. Thus, the acquired deep modality independent descriptor(DMID) can reduce the influence of significant differences between multi-modal images, further improving the matching performances. Experimental results on a large number of optical and SAR image-pairs demonstrate the effectiveness of DMID-Net on multi-modal image patch matching.
Huiyuan Wei, Dou Quan, Ruiqi Lei, Baorui Duan, Shuang Wang 0001, Yi Li 0054, Biao Hou, Licheng Jiao
IGARSS5
2022 A Transformer-Based Cross-Modal Image-Text Retrieval Method using Feature Decoupling and Reconstruction
abstract
With the increasing application of remote sensing technology, the task of cross-modal retrieval of remote sensing images (CMRRS) has gradually attracted widespread attention. Ex-isting methods often completely map the features of different modalities to a shared space and do not decouple between the modal-invariant information and modal-heterogeneous in-formation, which leads to redundant information in feature mapping and usually gets sub-optimal retrieval performance. This paper proposes a Transformer-based CMRRS method using feature decoupling and reconstruction (TBFDR) to solve this problem. TBFDR achieves state-of-the-art performance in remote sensing image-text retrieval task on Sydney-Captions dataset.
Yingzhi Sun, Yu Liao, Rui Yang 0038, Shuang Wang 0001, Biao Hou, Licheng Jiao
IGARSS6
2022 Remote Sensing Object Tracking With Deep Reinforcement Learning Under Occlusion
abstract
Object tracking is an important research direction of space Earth observation in the field of remote sensing. Although the existing correlation filter-based and deep learning (DL)-based object tracking algorithms have achieved great success, they are still unsatisfactory for the problem of object occlusion. The occlusion caused by the complex change in background, and the deviation of the tracking lens, causes object information to go missing, which leads to the omission of detection. Traditionally, most methods for object tracking under occlusion adopt a complex network model, which redetects the occluded object. To address this issue, we propose a novel object tracking approach. First, an action decision-occlusion handling network (AD-OHNet) based on deep reinforcement learning (DRL) is built to achieve low computational complexity for object tracking under occlusion. Second, the temporal and spatial context, the object appearance model, and the motion vector are adopted to provide the occlusion information, which drives actions in reinforcement learning under complete occlusion and contributes to improving the accuracy of tracking while maintaining speed. Finally, the proposed AD-OHNet is evaluated on three remote sensing video datasets of Bogota, Hong Kong, and San Diego taken from Jilin-1 commercial remote sensing satellites. The video datasets all shared problems of low spatial resolution, background clutter, and small objects. Experimental results on the three video datasets validate the effectiveness and efficiency of the proposed tracker.
Yanyu Cui, Biao Hou, Bo Ren 0001, Shuang Wang 0001, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.5
2022 Adaptive Fuzzy Learning Superpixel Representation for PolSAR Image Classification
abstract
The increasing applications of polarimetric synthetic aperture radar (PolSAR) image classification demand for effective superpixels’ algorithms. Fuzzy superpixels’ algorithms reduce the misclassification rate by dividing pixels into superpixels, which are groups of pixels of homogenous appearance and undetermined pixels. However, two key issues remain to be addressed in designing a fuzzy superpixel algorithm for PolSAR image classification. First, the polarimetric scattering information, which is unique in PolSAR images, is not effectively used. Such information can be utilized to generate superpixels more suitable for PolSAR images. Second, the ratio of undetermined pixels is fixed for each image in the existing techniques, ignoring the fact that the difficulty of classifying different objects varies in an image. To address these two issues, we propose a polarimetric scattering information-based adaptive fuzzy superpixel (AFS) algorithm for PolSAR images classification. In AFS, the correlation between pixels’ polarimetric scattering information, for the first time, is considered through fuzzy rough set theory to generate superpixels. This correlation is further used to dynamically and adaptively update the ratio of undetermined pixels. AFS is evaluated extensively against different evaluation metrics and compared with the state-of-the-art superpixels’ algorithms on three PolSAR images. The experimental results demonstrate the superiority of AFS on PolSAR image classification problems.
Yuwei Guo 0001, Licheng Jiao, Rong Qu, Zhuangzhuang Sun, Shuang Wang 0001, Shuo Wang 0005, Fang Liu 0001
IEEE Trans. Geosci. Remote. Sens.5
2022 MANet: Multi-Scale Aware-Relation Network for Semantic Segmentation in Aerial Scenes
abstract
Semantic segmentation is an important yet unsolved problem in aerial scenes understanding. One of the major challenges is the intense variations of scenes and object scales. In this paper, we propose a novel multi-scale aware-relation network (MANet) to tackle this problem in remote sensing. Inspired by the process of human perception of multi-scale information, we explore discriminative and diverse multi-scale representations. For discriminative multi-scale representations, we propose an inter-class and intra-class region refinement method (IIRR) to reduce feature redundancy caused by fusion. IIRR utilizes the refinement maps with intra- and inter-class scale variation to guide multi-scale fine-grained features. Then, we propose multi-scale collaborative learning (MCL) to enhance the diversity of multi-scale feature representations. The MCL constrains the diversity of multi-scale feature network parameters to obtain diverse information. And the segmentation results are rectified according to the dispersion of the multi-level network predictions. In this way, MANet can learn multi-scale features by collaboratively exploiting the correlation among different scales. Extensive experiments on image and video datasets which have large scale variations have demonstrated the effectiveness of our proposed MANet.
Pei He, Licheng Jiao, Ronghua Shang, Shuang Wang 0001, Xu Liu 0006, Dou Quan, Dong Zhao 0007
IEEE Trans. Geosci. Remote. Sens.4
2022 A Neural Network Based on Consistency Learning and Adversarial Learning for Semisupervised Synthetic Aperture Radar Ship Detection
abstract
Ship detection in synthetic aperture radar (SAR) images has important application value. Sea clutter, complex scenes, a large size change in ships, and the arbitrary directionality of ships make ship detection challenging. With the development of deep learning, many deep learning algorithms have been applied to SAR images. These algorithms need a lot of labeled data for training. It is time-consuming to label SAR data, and the unlabeled data are easy to obtain. It is necessary to use the unlabeled data effectively to improve the performance of the algorithm. In this study, a semisupervised SAR ship detection network, named the semisupervised consistency learning adversarial network (SCLANet), is presented. SCLANet is a two-stage detection network. The local features around the ship can be extracted by the SCLANet, and the features generated from unlabeled data become closer to those generated from labeled data by using adversarial learning. There are two consistency learning modules in SCLANet: noise robustness consistency learning and output encoding consistency learning. Noise robustness consistency learning can increase the robustness of the SCLANet. Maintaining consistency between the noisy results and the original results can train the unlabeled data. In output encoding consistency learning, outputs are mapped to a picture that is fed into an encoder to obtain the intermediate representation embedding. Another embedding is a layer in the main network of the SCLANet. Reducing the error between two embeddings can train the SCLANet with unlabeled data. Two types of consistency learning can be used as pretext tasks for semisupervised learning. Experiments were conducted on two SAR ship datasets. Compared with other algorithms, the SCLANet achieved the highest detection accuracy, indicating that it is more advantageous to use in ship detection.
Biao Hou, Zitong Wu, Bo Ren 0001, Xianpeng Guo, Shuang Wang 0001, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.6
2022 Multicrop Fusion Strategy Based on Prototype Assignment for Remote Sensing Image Scene Classification
abstract
The gap between self-supervised visual representation learning and supervised learning is gradually closing. Self-supervised learning does not rely on a large amount of labeled data and reduces the loss of human labeled information. Compared with natural images, remote sensing images require rich samples and human annotation by experts. Moreover, many algorithms have poor interpretability and unconvincing results. Therefore, this paper proposes a self-supervised method based on prototype assignment by designing a pretext task so that the network maps features to prototypes in the process of learning, swaps the code corresponding to the obtained features, combines them with another data-enhancing feature, and then optimizes the network. The prototype is introduced to explain the clustering idea embodied in the whole process. Considering the existence of the scene information-rich characteristic of remote sensing images, we introduce multiple views with different resolutions to capture more detailed information on the images. Finally, if the data enhancement method is not powerful enough, the network can easily fall into an overfitting state, which prevents the network from learning subtle differences and detailed information. To address this shortcoming, we propose a fusion strategy to flatten the decision boundary of the framework so that the model can also learn the soft similarity between sample pairs. We name the whole framework MFPC. In extensive experiments conducted on three common remote sensing image datasets (i.e., UCMerced, AID, and NWPU45), MFPC achieves a maximum improvement of 4.3% over some existing self-supervised algorithms, indicating that it can achieve good results.
Siteng Ma, Biao Hou, Xianpeng Guo, Zhihao Li 0005, Zitong Wu, Shuang Wang 0001, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.6
2022 Deep Feature Correlation Learning for Multi-Modal Remote Sensing Image Registration
abstract
Deep descriptors have advantages over handcrafted descriptors on local image patch matching. However, due to the complex imaging mechanism of remote sensing images and the significant differences in appearance between multi-modal images, existing deep learning descriptors are unsuitable for multi-modal remote sensing image registration directly. To solve this problem, this paper proposes a deep feature correlation learning network (Cnet) for multi-modal remote sensing image registration. Firstly, Cnet builds a feature learning network based on the deep convolutional network with the attention learning module, to enhance the feature representation by focusing on meaningful features. Secondly, this paper designs a novel feature correlation loss function for Cnet optimization. It focuses on the relative feature correlation between matching and non-matching samples, which can improve the stability of network training and decrease the risk of overfitting. Additionally, the proposed feature correlation loss with a scale factor can further enhance the network training and accelerate the network convergence. Extensive experimental results on image patch matching (Brown, HPatches), cross-spectral image registration (VIS-NIR), multi-modal remote sensing image registration, and single-modal remote sensing image registration have demonstrated the effectiveness and robustness of the proposed method.
Dou Quan, Shuang Wang 0001, Yu Gu 0015, Ruiqi Lei, Bowu Yang, Shaowei Wei, Biao Hou, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.2
2022 Self-Distillation Feature Learning Network for Optical and SAR Image Registration
abstract
Optical and SAR image registration is important for multi-modal remote sensing image information fusion. Recently, deep matching networks have shown better performances than traditional methods on image matching. However, due to significant differences between optical and SAR images, the performances of existing deep learning methods still need to be further improved. This paper proposes a self-distillation feature learning network (SDNet) for optical and SAR image registration, improving performance from network structure and network optimization. Firstly, we explore the impact of different weight-sharing strategies on optical and SAR image matching. Then, we design a partially unshared feature learning network for multi-modal image feature learning. It has fewer parameters than the fully unshared network and has more flexibility than the fully shared network. Additionally, the limited binary supervised information (matching or non-matching) is insufficient to train the deep matching networks for optical-SAR image registration. Thus, we propose a self-distillation feature learning method to exploit more similarity information for deep network optimization enhancing, such as the similarity ordering between a series of non-matching patch-pairs. The exploited rich similarity information will significantly enhance network training and improve matching accuracy. Finally, considering that existing deep learning methods brute-force constrain the features of the matching optical and SAR image patches are similar, which will be lost many discriminative information, degenerating matching performances. Thus, we build an auxiliary task reconstruction learning to optimize the feature learning network to keep more discriminative information. Extensive experiments demonstrate the effectiveness of our proposed method on multi-modal image registration.
Dou Quan, Huiyuan Wei, Shuang Wang 0001, Ruiqi Lei, Baorui Duan, Yi Li 0054, Biao Hou, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.3
2022 N-Cluster Loss and Hard Sample Generative Deep Metric Learning for PolSAR Image Classification
abstract
Deep learning works normally in PolSAR image classification because the complex terrain scattering characteristic results in large intraclass differences and high interclass similarity. Deep metric learning (DML) aims to make the features keep a closer intraclass and a farther interclass distance. Therefore, we introduce DML and then propose an N-cluster generative adversarial net (N-cluster GAN) framework for PolSAR image classification. However, existing DML losses mainly focus on the relationship between individual samples in feature space. Hence, we propose N-cluster loss that pays more attention to the overall structure of all samples. Meanwhile, traditional hard negative sample mining methods occupy lots of computational resources. In addition, the hard level of the negative samples will affect the model’s performance. Therefore, we explore a new method based on a GAN framework to replace the sample mining. Positive N-cluster loss is added to the discriminator ($D$), and a negative one is added to the generator ($G$). In this way,$D$will possess better classification ability, and$G$can produce hard negative samples for$D$. Then, the hard level of the generated negative samples will change with the discrimination of$D$, which is appropriate for the proposed model. N-cluster loss can be directly calculated through the extracted features rather than redundant data preparation. The proposed model is verified on four PolSAR datasets from two aspects of the loss function and negative samples mining. Then, it achieves competitive performance compared with state-of-the-art algorithms.
Chen Yang 0027, Biao Hou, Jocelyn Chanussot, Bo Ren 0001, Shuang Wang 0001, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.6
2022 Reconstruction Error-Based Decomposition Feature Selection for PolSAR Image
abstract
Target decomposition features are the cornerstone of subsequent analyses for PolSAR images. Generally, adopting single or several decomposition algorithms limits the representation ability for original terrain characteristics. Using all the existing decomposition features, however, will definitely increase computational complexity. Besides, some features even have a negative effect on the following tasks. To address these problems, a sparse variational autoencoder feature selection framework (SVAE-FS) is proposed in this article. In detail, the encoder transforms the original feature set into latent space and then decoder reconstructs the corresponding pseudo set on this latent space. Similarly, a pseudo subset is subsequently obtained by the SVAE. The discrepancy, namely reconstruction error, between the pseudo set and the pseudo subset is taken as an evaluation criterion which reflects the feature representation ability of pseudo subset. Sparse constraint in the encoder makes the representative features stand out. Meanwhile, the linear feature transformation layer of the encoder enables the SVAE to evaluate different scale subsets without repeated training. Finally, a greedy selection approach with search scale$K$is proposed to find the suboptimal subset. This procedure not only reduces time consumption, but also ensures the performance of the subset. The selected features are analyzed on four real PolSAR datasets according to the terrain scattering characteristics. Furthermore, these features have achieved competitive performance on three PolSAR image tasks.
Chen Yang 0027, Biao Hou, Xianpeng Guo, Bo Ren 0001, Jocelyn Chanussot, Shuang Wang 0001, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.6
2022 PDFL: Polarimetric Decomposition Feature Learning via Deep Autoencoder
abstract
Model-based polarimetric target decomposition (TD) generally solves scattering components and parameters under pre-set decomposition base, then decomposition features are also obtained. However, pre-set base could not be adjusted according to different scenes. Furthermore, solving the polarimetric parameters needs to explore additional information or consider limiting conditions to build equations, which is hard and easily to bring negative effects into decomposition features. To this end, we regard the TD as a process of learning decomposition base and features by deep learning. Then, the polarimetric decomposition feature learning (PDFL) model is proposed in this paper. Strictly, this model is not an incoherent TD method but a learning-based method. It dose not need to construct the parameter solution equations or fixed base. Then, the decomposition base and feature can be adaptively learned according to scattering characteristics of current dataset. Due to the characteristics of unsupervised reconstruction, deep auto encoder (DAE) is used as the model foundation. Then, some adjustments and constraints are utilized to make the DAE fit closely with TD. The encoder extracts latent vector from PolSAR data, then the decoder reconstructs pseudo data on this latent vector. The reconstruction can be regarded as the inverse process of TD, so the base matrix of decoder and the latent vector indicate the learned decomposition base and features when the model converges. The effectiveness of PDFL is verified on simulated and real PolSAR datasets. Compared with representative algorithms, proposed model gains more discriminative features and achieves competitive performance on terrain classification and segmentation tasks.
Chen Yang 0027, Biao Hou, Bo Ren 0001, Jocelyn Chanussot, Shuang Wang 0001, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.6
2022 A Joint-Training Two-Stage Method For Remote Sensing Image Captioning
abstract
Compared with remote sensing image (RSI) captioning methods based on the traditional encoder-decoder model, two-stage RSI captioning methods include an auxiliary remote sensing task to provide prior information, which enables them to generate more accurate descriptions. In previous two-stage RSI captioning methods, however, the image captioning and the auxiliary remote sensing tasks are handled separately, which is time-consuming and ignores mutual interference between tasks. To solve this problem, we propose a novel joint-training two-stage (JTTS) RSI captioning method. We use multi-label classification to provide prior information, and we design a differentiable sampling operator to replace the traditional non-differentiable sampling operation to index the multi-label classification result. In contrast to previous two-stage RSI captioning methods, our method can implement joint-training, and the joint loss allows the error of the generated description to flow into the optimization of the multi-label classification via back-propagation. Specifically, we approximate the Heaviside step function with the steep logistic function to implement a differentiable sampling operator for the multi-label classification. We propose a dynamic contrast loss function for multi-label classification task to ensure that a certain margin is maintained between the probabilities of the positive label and the negative label during sampling. We design an attribute-guided decoder to filter the multi-label prior information obtained by the sampling operator to generate more accurate image captions. The results of extensive experiments show that the JTTS method achieves state-of-the-art performance on the RSICD, the UCM-Captions, and the Sydney-Captions datasets.
Xiutiao Ye, Shuang Wang 0001, Yu Gu 0015, Jihui Wang, Biao Hou, Fausto Giunchiglia, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.2
2022 Adaptive Dual-Path Collaborative Learning for PAN and MS Classification
abstract
Due to the limitation of sensor technology, researchers tend to obtain high-quality image information from panchromatic (PAN) images and multispectral (MS) images with different resolutions. Therefore, the classification of remote sensing images of PAN and MS have become a research hotspot. In this paper, we propose an adaptive dual-path collaborative learning method for PAN and MS classification. In the stage of sample generation and training, we propose an adaptive neighborhood sample grading (ANSG) strategy in the establishing sample stage so that each pixel to be classified can obtain neighborhood information suitable for itself. Further, to simulate biological cognitive mechanisms, we divide the samples into different levels, and design the self-paced progressive loss (SPL), thus allowing the network to do preference training in different stages. The network’s training can quickly reach the optimal of the current stage and the overall convergence is more thorough. In the network structure, we propose a dual-path module (DPM) to effectively alleviate the gradient degradation in theresidual path, while ensuring maximum gradient loss information flow between every two layers in thedensely connected path. This module can extract more robust features to cope with the complex characteristics of remote-sensing images. Moreover, using the characteristics of the dual path to better fuse the features by the gradual collaborative fusion (GCF) way. The experimental results and theoretical analysis have demonstrated the proposed approach’s effectiveness, feasibility, and robustness. Our model are available at https://github.com/AIpy-nan/DBFI-Net.
Hao Zhu 0009, Kenan Sun, Licheng Jiao, Fang Liu 0001, Biao Hou, Shuang Wang 0001
IEEE Trans. Geosci. Remote. Sens.7
2022 Cluster Alignment With Target Knowledge Mining for Unsupervised Domain Adaptation Semantic Segmentation
abstract
Unsupervised domain adaptation (UDA) carries out knowledge transfer from the labeled source domain to the unlabeled target domain. Existing feature alignment methods in UDA semantic segmentation achieve this goal by aligning the feature distribution between domains. However, these feature alignment methods ignore the domain-specific knowledge of the target domain. In consequence, 1) the correlation among pixels of the target domain is not explored; and 2) the classifier is not explicitly designed for the target domain distribution. To conquer these obstacles, we propose a novel cluster alignment framework, which mines the domain-specific knowledge when performing the alignment. Specifically, we design a multi-prototype clustering strategy to make the pixel features within the same class tightly distributed for the target domain. Subsequently, a contrastive strategy is developed to align the distributions between domains, with the clustered structure maintained. After that, a novel affinity-based normalized cut loss is devised to learn task-specific decision boundaries. Our method enhances the model's adaptability in the target domain, and can be used as a pre-adaptation for self-training to boost its performance. Sufficient experiments prove the effectiveness of our method against existing state-of-the-art methods on representative UDA benchmarks.
Shuang Wang 0001, Dong Zhao 0007, Yuwei Guo 0001, Qi Zang, Yu Gu 0015, Yi Li 0054, Licheng Jiao
IEEE Trans. Image Process.1
2022 Element-Wise Feature Relation Learning Network for Cross-Spectral Image Patch Matching
abstract
Recently, the majority of successful matching approaches are based on convolutional neural networks, which focus on learning the invariant and discriminative features for individual image patches based on image content. However, the image patch matching task is essentially to predict the matching relationship of patch pairs, that is, matching (similar) or non-matching (dissimilar). Therefore, we consider that the feature relation (FR) learning is more important than individual feature learning for image patch matching problem. Motivated by this, we propose an element-wise FR learning network for image patch matching, which transforms the image patch matching task into an image relationship-based pattern classification problem and dramatically improves generalization performances on image matching. Meanwhile, the proposed element-wise learning methods encourage full interaction between feature information and can naturally learn FR. Moreover, we propose to aggregate FR from multilevels, which integrates the multiscale FR for more precise matching. Experimental results demonstrate that our proposal achieves superior performances on cross-spectral image patch matching and single spectral image patch matching, and good generalization on image patch retrieval.
Dou Quan, Shuang Wang 0001, Ning Huyan, Jocelyn Chanussot, Ruojing Wang, Xuefeng Liang, Biao Hou, Licheng Jiao
IEEE Trans. Neural Networks Learn. Syst.2
2021 Cascade Attention Fusion for Fine-Grained Image Captioning Based on Multi-Layer LSTM
abstract
The conventional visual attention-based image captioning approaches typically use image information to guide caption generation. Results from these models tend to be coarse and ignore the details in the image, such as objects, attributes and the distinguishing aspects of each image. In this paper, we propose a visual and semantic fusion network with a margin-based training guidance mechanism to generate fine image descriptions that depict more objects, attributes and other distinguishing aspects of images. In our model, the visual attention layer introduces more low-level visual information, the semantic attention layer provides more high-level semantic attributes. Furthermore, the proposed margin-based loss encourages our model to produce more discriminative descriptions. Extensive experiments are conducted on COCO and Flickr30K image captioning datasets to validate our method, and the results show its superior performance at captioning. Our method achieves a state-of-the-art 70.6 CIDEr-D on the Flickr30K dataset, and a competitive 123.5 CIDEr-D on the MS-COCO dataset.
Shuang Wang 0001, Yun Meng, Yu Gu 0015, Xiutiao Ye, Jingxian Tian, Licheng Jiao
ICASSP1
2021 A Feature Decomposition Framework for Multi-Modal Image Patch Matching
abstract
Multi-modal remote sensing images have complementary information which is conducive to enhancing the performance of various applications. Image patch matching plays a crucial role in the combination of multi-modal images. However, there are great differences in appearance and texture of multi-modal images, which brings great difficulties to image patching matching. To solve this problem, we propose a novel feature decomposition framework for multi-modal image patch matching. It aims to eliminate the hinder caused by the significant difference in multi-modal images. Specifically, this paper proposes to decompose the feature of images into common feature and modal private feature. Then, only the common feature is used for image patch matching, so as to improve the matching accuracy. Experimental results on optical and SAR images demonstrate that our proposed feature decomposition framework can significantly improve the performance of multi-modal image patch matching.
Baorui Duan, Dou Quan, Yi Li 0054, Ruiqi Lei, Shuang Wang 0001, Biao Hou, Licheng Jiao
IGARSS5
2021 Graph Regular Loss for Semi-Supervised Polsar Terrain Classification
abstract
Amongst the utilizations of Polarimetric Synthetic Aperture Radar (PoISAR) data, semi-supervised terrain classification is much in demand. Samples of the same category are distributed in multiple regions of a PolSAR image, resulting in differences in the feature distributions of samples of the same category located in different regions. In addition, some samples of confusable categories, and samples of different categories located near the edges, have small feature differences in a PolSAR image. To address this problem, we introduce a graph regular loss constructed from pseudo-labels to constrain the intra-class similarity and inter-class similarity of features, and improve the discriminative property of features. In addition, since the quality of pseudo-labels affects the optimization of the model by the graph regular loss, we introduce the idea of clustering into the teacher-student model to improve the quality of pseudo-labels. Experiments on real PolSAR data show that our proposed method achieves an excellent performance.
Chunlei Han, Yuwei Guo 0001, Qi Zang, Baorui Duan, Dong Zhao 0007, Shuang Wang 0001
IGARSS8
2021 Deep Global Feature-Based Template Matching for Fast Multi-Modal Image Registration
abstract
Due to the different imaging mechanisms, there is a significant non-line difference between multi-modal images, which brings difficulties to multi-modal image registration. The traditional methods based on grayscale and handcraft features are difficult with obtain common features between different source images. The performances of deep local features matching methods rely on the quality and quantity of the detected keypoints, which can be quite time-consuming to register images. To achieve fast and accurate multi-modal image registration, we propose a deep global feature-based template matching method (GFTM) which uses a deep convolutional network to extract common global deep features from multi-modal images. Then, fast template matching is performed on global deep features to search the position with maximal similarity. Additionally, we build a similarity label map and design three losses to optimize our network, including contrast loss, error loss and peak loss. Extensive experimental results on optical and SAR images demonstrated that our proposed method is effective on multi-modal image registration.
Ruiqi Lei, Bowu Yang, Dou Quan, Yi Li 0054, Baorui Duan, Shuang Wang 0001, Huarong Jia, Biao Hou, Licheng Jiao
IGARSS6
2021 Multi-View Attention Network for Remote Sensing Image Captioning
abstract
In traditional remote sensing image captioning models, the attention mechanism plays a dominant role and has been used to integrate image features to infer the latent visual-semantic alignment. However, the scenes of remote sensing image are complex and diverse, using only one attention module to capture features often leads to insufficient semantic representation. In our work, we present a novel Multi-view Attention Network (MAN) model to realize feature integration from different views. With MAN, more semantically rich ensemble attended features can be obtained by different attention modules. Specifically, we enforce the weights of attention modules to be diverse through a cosine distance loss. This will provide the model with distinct views to make semantic predictions for each feature. Extensive experiments on benchmark datasets demonstrate the effectiveness of the proposed model for the task of remote sensing image captioning.
Yun Meng, Yu Gu 0015, Xiutiao Ye, Jingxian Tian, Shuang Wang 0001, Biao Hou, Licheng Jiao
IGARSS5
2021 Cross-Modal Feature Fusion Retrieval for Remote Sensing Image-Voice Retrieval
abstract
With the increasing popularity of remote sensing technology applications, some emergency scenarios require rapid retrieval of remote sensing images, such as earthquake rescue, etc. Due to the high efficiency of voice input, researchers have focused on cross-modal remote sensing image-voice retrieval methods. However, these methods have two major drawbacks: speech input lacks discrimination and the intra-modal semantic information is under used. To address these drawbacks, we propose a novel cross-modal feature fusion retrieval model. Our model provides a more optimized cross-modal common feature space than previous models and thus optimizes the retrieval performance. First, our model adds the extra textual keyword information to the audio feature for remote sensing image retrieval. Second, it introduces inter-modality adversarial learning and intra-modality semantic discrimination into the remote sensing image-voice retrieval task. We conducted experiments on two datasets modified from the UCM-Captions dataset and the Remote Sensing Image Caption Dataset. The experimental results show that our model outperforms state-of-the-art models in this task.
Rui Yang 0038, Yu Gu 0015, Yu Liao, Yingzhi Sun, Shuang Wang 0001, Biao Hou, Licheng Jiao
IGARSS6
2021 Multi-Relation Attention Network for Image Patch Matching
abstract
Deep convolutional neural networks attract increasing attention in image patch matching. However, most of them rely on a single similarity learning model, such as feature distance and the correlation of concatenated features. Their performances will degenerate due to the complex relation between matching patches caused by various imagery changes. To tackle this challenge, we propose a multi-relation attention learning network (MRAN) for image patch matching. Specifically, we propose to fuse multiple feature relations (MR) for matching, which can benefit from the complementary advantages between different feature relations and achieve significant improvements on matching tasks. Furthermore, we propose a relation attention learning module to learn the fused relation adaptively. With this module, meaningful feature relations are emphasized and the others are suppressed. Extensive experiments show that our MRAN achieves best matching performances, and has good generalization on multi-modal image patch matching, multi-modal remote sensing image patch matching and image retrieval tasks.
Dou Quan, Shuang Wang 0001, Yi Li 0054, Bowu Yang, Ning Huyan, Jocelyn Chanussot, Biao Hou, Licheng Jiao
IEEE Trans. Image Process.2
2020 PolSAR Scene Classification via Low-Rank Tensor-Based Multi-View Subspace Representation
abstract
In this paper, the polarimetric synthetic aperture radar (PolSAR) scene classification is solved by using a novel low -rank tensor-based multi-view subspace representation (LRT-MSR) method. PolSAR data can be described in multimodal feature spaces, such as PolSAR coherent/covariance/scattering matrices, or the various polarimetric decompositions. Different pseudo-color images from multiple spaces provide enough visual information for making a comprehensive representation. Our method applies a low-rank tensor-based subspace clustering way to explore the information from multi-view pseudo-color images. Tensor, as the high order matrix, is used to capture the correlations of underlying multi-view data. Furthermore, the method is constrained by a low-rank term that elegantly models the cross information from different views, and achieves a series of representation matrices from the redundant information. Finally, a spectral cluster method is used to make the final classification. The experimental results on PolSAR image dataset show the effectiveness of the applied method.
Mengqian Chen, Bo Ren 0001, Biao Hou, Jocelyn Chanussot, Shuang Wang 0001, Xiangrong Zhang
IGARSS5
2020 Panchromatic Image Land Cover Classification Via DCNN with Updating Iteration Strategy
abstract
Land cover classification is a critical research task in many significant remote sensing applications. There are emerging many powerful pixel-level classification methods based on deep convolutional neural network (DCNN) in the universal computer vision community. However, due to the complication of satellite image senses and the lack of high-quality labeled datasets, these computer vision techniques can not be applied to remote sensing applications directly. In this paper, we propose a novel DCNN method to extract abstract feature from the complicated remote scene. The proposed method fuses three level features from the encoder while the segmentation result is obtained by decoder. Furthermore, we propose an updating iteration strategy (UIS) with label smoothing on training set to reduce the impact of the incorrect labels. The proposed strategy employs the output of the network to update the low-confidence labels on training set, and utilizes the updated labels to continue training the network. In order to acquire a better segmentation result on a very high resolution (VHR) panchromatic image, we transfer the features trained on GID dataset to our dataset for training. Our experiments has demonstrated the oustanding performance of the proposed method in land cover classification compared to DeepLabv3 on the GID and our dataset.
Biao Hou, Yangfei Liu, Tuotuo Rong, Bo Ren 0001, Zijuan Xiang, Xiangrong Zhang, Shuang Wang 0001
IGARSS7
2020 Semi-Supervised PolSAR Image Classification Based on Improved Tri-Training With a Minimum Spanning Tree
abstract
In this article, the terrain classifications of polarimetric synthetic aperture radar (PolSAR) images are studied. A novel semi-supervised method based on improved Tri-training combined with a neighborhood minimum spanning tree (NMST) is proposed. Several strategies are included in the method: 1) a high-dimensional vector of polarimetric features that are obtained from the coherency matrix and diverse target decompositions is constructed; 2) this vector is divided into three subvectors and each subvector consists of one-third of the polarimetric features, randomly selected. The three subvectors are used to separately train the three different base classifiers in the Tri-training algorithm to increase the diversity of classification; and 3) a help-training sample selection with the improved NMST that uses both the coherency matrix and the spatial information is adopted to select highly reliable unlabeled samples to increase the training sets. Thus, the proposed method can effectively take advantage of unlabeled samples to improve the classification. Experimental results show that with a small number of labeled samples, the proposed method achieves a much better performance than existing classification methods.
Shuang Wang 0001, Yanhe Guo, Wenqiang Hua, Xinan Liu, Guoxin Song, Biao Hou, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.1
2020 POL-SAR Image Classification Based on Modified Stacked Autoencoder Network and Data Distribution
abstract
This article proposes a novel autoencoder (AE) network based on the distribution of polarimetric synthetic aperture radar (POL-SAR) data matrix, called a mixture autoencoder (MAE). Through a detailed analysis of the data distribution POL-SAR data matrix, a normalization method is also presented in succession. The proposed MAE defines the data error term in the loss function according to the data distribution. It can be regarded as a process of unsupervised feature extraction designed specifically for POL-SAR data matrix. Then, a softmax classifier is trained with the help of data features and the corresponding label information. Next, a stacked MAE (SMAE) network is reasonably constructed by considering the data distribution among different layers. Finally, this article also presents a classification network through discarding the decoder process of the proposed SMAE and connecting with a softmax classifier. The SMAE is trained layer by layer using the unlabeled data. The softmax classifier is also trained with a small number of labeled pixels. With parameters obtained from the above-mentioned procedures as the initial parameters, the whole classification network is trained by the labeled pixels to get a well-trained model, which is used for predicting the corresponding label of the pixel in the data set. Three real POL-SAR data sets, including the AIR-SAR L-band data of Flevoland, The Netherlands, are used in the experiments. Compared with one classical algorithm and two related models with the similar structure, both the proposed methods show improvements in overall accuracy and efficiency as well as possess better adaptability of the parameter and preferable consistency with the classification performance.
Jianlong Wang, Biao Hou, Licheng Jiao, Shuang Wang 0001
IEEE Trans. Geosci. Remote. Sens.4
2019 AFD-Net: Aggregated Feature Difference Learning for Cross-Spectral Image Patch Matching
abstract
Image patch matching across different spectral domains is more challenging than in a single spectral domain. We consider the reason is twofold: 1. the weaker discriminative feature learned by conventional methods; 2. the significant appearance difference between two images domains. To tackle these problems, we propose an aggregated feature difference learning network (AFD-Net). Unlike other methods that merely rely on the high-level features, we find the feature differences in other levels also provide useful learning information. Thus, the multi-level feature differences are aggregated to enhance the discrimination. To make features invariant across different domains, we introduce a domain invariant feature extraction network based on instance normalization (IN). In order to optimize the AFD-Net, we borrow the large margin cosine loss which can minimize intra-class distance and maximize inter-class distance between matching and non-matching samples. Extensive experiments show that AFD-Net largely outperforms the state-of-the-arts on the cross-spectral dataset, meanwhile, demonstrates a considerable generalizability on a single spectral dataset.
Dou Quan, Xuefeng Liang, Shuang Wang 0001, Shaowei Wei, Ning Huyan, Licheng Jiao
ICCV3
2019 Better and Faster: Exponential Loss for Image Patch Matching
abstract
Recent studies on image patch matching are paying more attention on hard sample learning, because easy samples do not contribute much to the network optimization. They have proposed various hard negative sample mining strategies, but very few addressed this problem from the perspective of loss functions. Our research shows that the conventional Siamese and triplet losses treat all samples linearly, thus make the training time consuming. Instead, we propose the exponential Siamese and triplet losses, which can naturally focus more on hard samples and put less emphasis on easy ones, meanwhile, speed up the optimization. To assist the exponential losses, we introduce the hard positive sample mining to further enhance the effectiveness. The extensive experiments demonstrate our proposal improves both metric and descriptor learning on several well accepted benchmarks, and outperforms the state-of-the-arts on the UBC dataset. Moreover, it also shows a better generalizability on cross-spectral image matching and image retrieval tasks.
Shuang Wang 0001, Xuefeng Liang, Dou Quan, Bowu Yang, Shaowei Wei, Licheng Jiao
ICCV1
2019 An Improved Fully Convolutional Network for Learning Rich Building Features
abstract
Many efficient approaches are proposed to detect building in remote sensing images. In this paper, in order to learning rich building features better, we propose a full convolutional network with dense connection. There contributions are made: 1) To strengthen feature propagation, an improved dense network is introduced to the full convolution network. 2) We have designed top-down short connections to facilitate the fusion of high and low feature information. 3) In addition, we add the weighted cross entropy edge loss function to make the network pay more attention to building edge in detail. Experiments show that the proposed method achieves excellent performance on the remote sensing image data taken by the QuickBird satellite.
Shuang Wang 0001, Pei He, Dou Quan, Xuefeng Liang, Biao Hou
IGARSS1
2019 Polsar Terrain Classification Based on Denoising-CNN
abstract
Terrain classification plays an important role in understanding Polari- metric Synthetic Aperture Radar (PolSAR) image intuitively. In the process of classification, feature extraction is critical. However, the preprocess of speckle noise filtering affects the effectiveness of the feature extractor which influences the accuracy of classification ultimately. Thus we integrate de-noising process and classification into an end-to-end framework based on CNN termed as Denoising-CNN, which improves the accuracy of classification. Experiments on real PolSAR data show that our proposed method offers an excellent performance.
Yanhe Guo, Shuang Wang 0001, Guoxin Song, Wenqiang Hua, Feihang Liu
IGARSS2
2019 Object Detection and Trcacking Based on Convolutional Neural Networks for High-Resolution Optical Remote Sensing Video
abstract
Object detection algorithms, from high-resolution optical remote sensing images, have been booming from the last few years. However, object tracking for high-resolution optical remote sensing video is a challenging task due to the large number and small size of objects. In this paper, we propose an object detection and tracking method based on deep convolutional neural networks for wide swath high-resolution optical remote sensing videos. The proposed method firstly segments each frame of a video into sub-samples using a sliding window of fixed size. In order to detect the objects appearing at the edge of the sliding window efficiently, we use an overlapping sliding window sampling method. Further, we design a network fusing region of interests (RoIs) of the previous and current frames to track the objects occurred in the previous frames of the video. RoIs of previous frame are applied directly to the feature layer of the current frame. Finally, for each frame, we merge the detection and tracking results of sub-samples by non-maximum suppression (NMS) method. The experimental results on our dataset demonstrate the validity and generality of the proposed detection algorithm.
Biao Hou, Jingliang Li, Xiangrong Zhang, Shuang Wang 0001, Licheng Jiao
IGARSS4
2019 Dual-Channel Convolutional Neural Network for Polarimetric SAR Images Classification
abstract
This paper presents a new dual-channel convolutional neural network (Dc-CNN) for Polarimetric synthetic aperture radar (PolSAR) image classification when labeled samples are small. First, a neighborhood minimum spanning tree (MST) is used to enlarge the labeled sample set. Then, in order to obtain the abundant spatial information, a new dual-channel CNN is designed to PolSAR image to acquire different spatial features. This network model contains two parallel CNN structures, which can extract different features used two multiscale convolution structure. Experiments results show that compared with other methods, the proposed method shows a satisfactory classification result.
Wenqiang Hua, Shuang Wang 0001, Yanhe Guo, Xiaomin Jin
IGARSS2
2018 Cross-Spectral Image Patch Matching by Learning Features of the Spatially Connected Patches in a Shared Space
Dou Quan, Shuai Fang, Xuefeng Liang, Shuang Wang 0001, Licheng Jiao
ACCV (2)4
2018 A Two-Branch Network with Semi-Supervised Learning for Hyperspectral Classification
abstract
In order to promote progress on fusion and analysis methodologies for multi-source remote sensing data, The Image Analysis and Data Fusion Technical Committee organized the 2018 IEEE GRSS Data Fusion contest. In this contest, we proposed a two-branch convolution network for hyperspectral image classification with a data re-sampling strategy and semi-supervised learning to address three existing problems, i.e. multi-scale feature learning, data imbalance, and small size of the dataset. The contest showed that our proposal achieved the best performance on two metrics: the overall accuracy of 77.39% and a kappa coefficient of 0.76 on the hyperspectral images provided by 2018 IEEE GRSS Data Fusion Contest.
Shuai Fang, Dou Quan, Shuang Wang 0001
IGARSS3
2018 PolSAR Image Classification Based on DBN and Tensor Dimensionality Reduction
abstract
This paper proposes a new semi-supervised PolSAR image classification method using deep belief network (DBN) and tensor dimensionality reduction, which uses multilinear principle component analysis (MPCA) to reduce the dimension of tensor form PolSAR data, and regards the multiple features of PolSAR data as the input of DBN. In order to take full advantage of neighborhood information of each pixel of PolSAR data, we take each pixel and its neighborhood as tensor form. For PolSAR data, simple feature has been proven not to be able to effectively classify complex terrains. Therefore, we combine multiple features of PolSAR data to obtain more abundant information, which can reflect some spatial structure of PolSAR data. The experimental results show that the overall classification accuracy based on the proposed method outperforms the traditional classification strategies.
Biao Hou, Xianpeng Guo, Weidan Hou, Shuang Wang 0001, Xiangrong Zhang, Licheng Jiao
IGARSS4
2018 Fully Convolutional Semi-Supervised Gan for Polsar Classification
abstract
We propose a novel semi -supervised fully convolutional network for Polarimetric synthetic aperture radar (PoISAR) terrain classification. First, by designing a fully convolutional structure, we can perform pixel-based classification tasks. Then, by applying semi -supervised generative adversarial networks (GANs), we utilize both labeled and unlabeled samples and aim to obtain higher classification accuracy. Through a mini-max two-player game, GAN has better performance than other “single-player” classifiers. Finally, we combine the fully convolutional structure with the semi-supervised GAN. Our fully convolutional semi-supervised GAN (FC-SGAN) has excellent spatial feature learning ability and can perform end-to-end pixel-based classification tasks. Experimental results show that compared with existing works, the proposed method has better performances. Even when the training set gets smaller, our method keeps high accuracy.
Mengchen Liu, Shuang Wang 0001, Yanhe Guo, Biao Hou, Licheng Jiao, Xiaojin Hou
IGARSS3
2018 Deep Generative Matching Network for Optical and SAR Image Registration
abstract
Multimodal remote sensing images contain complementary information, thus, could potentially benefit many remote sensing applications. To this end, the image registration is a common requirement for utilizing the multimodal images. However, due to the rather different imaging mechanisms, multimodal image registration becomes much more challenging than ordinary registration, particular for optical and synthetic aperture radar (SAR) images. In this work, we design a deep matching network to exploit the latent and coherent features between multimodal patch pairs for inferring their matching labels. But, the network requires immense data for training, which is not usually met. To address this issue, we propose a generative matching network (GMN) to generate the coupled optical and SAR images, hence, improve the quantity and diversity of the training data. The experimental results show that our proposal significantly improves the registration performance of optical and SAR image registration, and achieves subpixel or close to subpixel error.
Dou Quan, Shuang Wang 0001, Xuefeng Liang, Ruojing Wang, Shuai Fang, Biao Hou, Licheng Jiao
IGARSS2
2018 An external learning assisted self-examples learning for image super-resolution
Bo Yue, Shuang Wang 0001, Xuefeng Liang, Licheng Jiao
Neurocomputing2
2018 Combining ConvNets with hand-crafted features for action recognition based on an HMM-SVM classifier
Shuang Wang 0001, Yonghong Hou, Jiarong Dong, Chang Tang
Multim. Tools Appl.1
2018 Fuzzy Sparse Autoencoder Framework for Single Image Per Person Face Recognition
abstract
The issue of single sample per person (SSPP) face recognition has attracted more and more attention in recent years. Patch/local-based algorithm is one of the most popular categories to address the issue, as patch/local features are robust to face image variations. However, the global discriminative information is ignored in patch/local-based algorithm, which is crucial to recognize the nondiscriminative region of face images. To make the best of the advantage of both local information and global information, a novel two-layer local-to-global feature learning framework is proposed to address SSPP face recognition. In the first layer, the objective-oriented local features are learned by a patch-based fuzzy rough set feature selection strategy. The obtained local features are not only robust to the image variations, but also usable to preserve the discrimination ability of original patches. Global structural information is extracted from local features by a sparse autoencoder in the second layer, which reduces the negative effect of nondiscriminative regions. Besides, the proposed framework is a shallow network, which avoids the over-fitting caused by using multilayer network to address SSPP problem. The experimental results have shown that the proposed local-to-global feature learning framework can achieve superior performance than other state-of-the-art feature learning algorithms for SSPP face recognition.
Yuwei Guo 0001, Licheng Jiao, Shuang Wang 0001, Shuo Wang 0005, Fang Liu 0001
IEEE Trans. Cybern.3
2018 Global Low-Rank Image Restoration With Gaussian Mixture Model
abstract
Low-rank restoration has recently attracted a lot of attention in the research of computer vision. Empirical studies show that exploring the low-rank property of the patch groups can lead to superior restoration performance, however, there is limited achievement on the global low-rank restoration because the rank minimization at image level is too strong for the natural images which seldom match the low-rank condition. In this paper, we describe a flexible global low-rank restoration model which introduces the local statistical properties into the rank minimization. The proposed model can effectively recover the latent global low-rank structure via nuclear norm, as well as the fine details via Gaussian mixture model. An alternating scheme is developed to estimate the Gaussian parameters and the restored image, and it shows excellent convergence and stability. Besides, experiments on image and video sequence datasets show the effectiveness of the proposed method in image inpainting problems.
Licheng Jiao, Fang Liu 0001, Shuang Wang 0001
IEEE Trans. Cybern.4
2018 Fuzzy Superpixels for Polarimetric SAR Images Classification
abstract
Superpixels technique has drawn much attention in computer vision applications. Each superpixels algorithm has its own advantages. Selecting a more appropriate superpixels algorithm for a specific application can improve the performance of the application. In the last few years, superpixels are widely used in polarimetric synthetic aperture radar (PolSAR) image classification. However, no superpixel algorithm is especially designed for image classification. It is believed that both mixed superpixels and pure superpixels exist in an image. Nevertheless, mixed superpixels have negative effects on classification accuracy. Thus, it is necessary to generate superpixels containing as few mixed superpixels as possible for image classification. In this paper, first, a novel superpixels concept, named fuzzy superpixels, is proposed for reducing the generation of mixed superpixels. In fuzzy superpixels, not all pixels are assigned to a corresponding superpixel. We would rather ignore the pixels than assigning them to improper superpixels. Second, a new algorithm, named FuzzyS (FS), is proposed to generate fuzzy superpixels for PolSAR image classification. Three PolSAR images are used to verify the effect of the proposed FS algorithm. Experimental results demonstrate the superiority of the proposed FS algorithm over several state-of-the-art superpixels algorithms.
Yuwei Guo 0001, Licheng Jiao, Shuang Wang 0001, Shuo Wang 0005, Fang Liu 0001, Wenqiang Hua
IEEE Trans. Fuzzy Syst.3
2018 Mutual Learning Between Saliency and Similarity: Image Cosegmentation via Tree Structured Sparsity and Tree Graph Matching
abstract
This paper proposes a unified mutual learning framework based on image hierarchies, which integrates structured sparsity with tree-graph matching to conquer the problem of weakly supervised image cosegmentation. We focus on the interaction between two common-object properties: saliency and similarity. Most existing cosegmentation methods only pay emphasis on either of them. The proposed method realizes the learning of the prior knowledge for structured sparsity with the help of treegraph matching, which is capable of generating object-oriented salient regions. Meanwhile, it also reduces the searching space and computational complexity of tree-graph matching with the attendance of structured sparsity. We intend to thoughtfully exploit the hierarchically geometrical relationships of coherent objects. Experimental results compared with the state-of-thearts on benchmark datasets confirm that the mutual learning framework are capable of effectively delineating co-existing object patterns in multiple images.
Yan Ren 0002, Licheng Jiao, Shuyuan Yang 0001, Shuang Wang 0001
IEEE Trans. Image Process.4
2018 How Does the Low-Rank Matrix Decomposition Help Internal and External Learnings for Super-Resolution
abstract
Wisely utilizing the internal and external learning methods is a new challenge in super-resolution problem. To address this issue, we analyze the attributes of two methodologies and find two observations of their recovered details: 1) they are complementary in both feature space and image plane and 2) they distribute sparsely in the spatial space. These inspire us to propose a low-rank solution which effectively integrates two learning methods and then achieves a superior result. To fit this solution, the internal learning method and the external learning method are tailored to produce multiple preliminary results. Our theoretical analysis and experiment prove that the proposed low-rank solution does not require massive inputs to guarantee the performance, and thereby simplifying the design of two learning methods for the solution. Intensive experiments show the proposed solution improves the single learning method in both qualitative and quantitative assessments. Surprisingly, it shows more superior capability on noisy images and outperforms state-of-the-art methods.
Shuang Wang 0001, Bo Yue, Xuefeng Liang, Licheng Jiao
IEEE Trans. Image Process.1
2017 Fast graph-based SAR image segmentation via simple superpixels
abstract
Graph-based methods have been successfully applied in the field of computer vision for image segmentation. Unfortunately, most of them are not suitable to deal with large-scale SAR image segmentation due to their high computation complexity. A fast and efficient graph-based SAR image segmentation is proposed in this paper through using superpixels to reduce the computation complexity. Firstly, a SAR image is divided into several non-overlapped subdivisions with the same size. Each of the subdivision is processed as a single OpenMP parallel region, which can be processed at the single computing node with multi-core CPU. Secondly, the number of nodes and edges in the graph is reduced by extracting the superpixels other than single pixels based on global information of each subdivision. Finally, an effective rule is proposed to merge two adjacent sub-graphs from two different subdivisions into a new subgraph.
Biao Hou, Dezhao Gong, Shuang Wang 0001, Xiangrong Zhang, Licheng Jiao
IGARSS4
2017 Semi-supervised PolSAR Classification Based on Improved Tri-training
abstract
In this paper, we proposed a new semi-supervised method for polarimetric synthetic aperture radar (PolSAR) terrain classification based on improved tri-training. This method only needs a few numbers of labeled samples to achieve the results obtained by traditional supervised classification methods. First, it uses a variety of target decomposition methods to obtain high-dimensional feature. Second, a new feature selection method based on the ratio of between-class scatter and within-class scatter is proposed to reduce the redundant feature. Finally, an improved tri-training method is executed. A real PolSAR data is used to verify the proposed method. Experimental results show that the proposed method is efficient with a few labeled samples and effectively improve the classification accuracy compared with other traditional classification methods.
Wenqiang Hua, Shuang Wang 0001, Bo Yue, Yanhe Guo
IGARSS2
2017 SAR images super-resolution via cartoon-texture image decomposition and jointly optimized regressors
abstract
This paper presents a novel approach to enhance the spatial resolution of Synthetic Aperture Radar(SAR) images. SAR images super-resolution(SR) reconstruction is challenging since SAR images has more complex structures. Inspired by the recent advance on natural image SR techniques, we propose a joint learning based strategy[1], combined with the characteristics of SAR image, to reconstruct HR SAR images from LR SAR images. Our method has ability to handle the complicated structures of SAR images. Besides, SAR images are decomposed into cartoon components and texture components and processed respectively. The purpose of decomposing strategy is to reduce the influence of speckle noise of SAR images. The experimental results and comparative analyses verify the effectiveness of this algorithm.
Shuang Wang 0001, Caijin Xu, Bo Yue, Xuefeng Liang
IGARSS2
2017 Unsupervised classification of PolSAR data based on a novel polarization feature
abstract
In this paper, we present a new unsupervised classification method based on a novel polarization feature, which reflects the proportion of co-polarization component and cross-polarization component of scatters in PolSAR image. We combine this novel polarization feature with backscattering power and scattering power entropy to perform the initial classification. Then apply a merge criterion to merge clusters into the desired number of clusters. After each step, the complex Wishart clustering is performed to refine the classification results. Compared with the other three methods, the effectiveness of the proposed approach is demonstrated on NASA/JPL AIRSAR L-band data of San Francisco Bay.
Shuang Wang 0001, Wenqiang Hua
IGARSS2
2017 Unsupervised saliency-guided SAR image change detection
Yaoguo Zheng, Licheng Jiao, Hongying Liu 0001, Xiangrong Zhang, Biao Hou, Shuang Wang 0001
Pattern Recognit.6
2017 Robust coupled dictionary learning with ℓ1-norm coefficients transition constraint for noisy image super-resolution
Bo Yue, Shuang Wang 0001, Xuefeng Liang, Licheng Jiao
Signal Process.2
2017 A Salient Region Detection and Pattern Matching-Based Algorithm for Center Detection of a Partially Covered Tropical Cyclone in a SAR Image
abstract
Spaceborne microwave synthetic aperture radar (SAR), with its high spatial resolution, large area coverage, day/night imaging capability, and penetrating cloud capability, has been used as an important tool for tropical cyclone monitoring. The accuracy of locating tropical cyclone centers has a large impact on the accuracy of tropical cyclone track prediction. Usually, the center of a tropical cyclone can be accurately located if the tropical cyclone eye is fully covered by a SAR image. In some cases, due to the limited coverage of the SAR, only a part of a tropical cyclone can be imaged without the eye. From a SAR image processing point of view, these facts make the automatic center location of tropical cyclones a challenging work. This paper addresses the problem by proposing a semiautomatic center location method based on salient region detection and pattern matching. A salient region detection algorithm is proposed, in which the salient region map contains mainly the rain bands of a tropical cyclone in a SAR image. The pattern matching problem is transformed into an optimization problem solved by using the particle swarm optimization algorithm to search the best estimated center of a tropical cyclone. To estimate the accuracy of the located center, we compare the results with the NOAA National Hurricane Center's best track data. Experiments demonstrate that the proposed method achieves good accuracy for locating the centers of tropical cyclones from SAR images that do not contain a distinguishable eye signature.
Shaohui Jin, Shuang Wang 0001, Xiaofeng Li 0001, Licheng Jiao, Jun A. Zhang, Dongliang Shen
IEEE Trans. Geosci. Remote. Sens.2
2016 Unsupervised PolSAR image classification using boundary-preserving region division and region-based affinity propagation clustering
abstract
This paper presents a new method for polarimetric synthetic aperture radar (PolSAR) image classification. Firstly, to get a reasonable edge strength map, polarimetric information is used in edge strength calculation, and watershed algorithm is used to obtain the oversegmentation using the edge strength. Secondly, a searching table is used to determine the most suitable region to be merged. Finally, region-based affinity propagation clustering is employed to achieve an initial classification map, and the method provides an adjacent Wishart classifier with spatial relations to obtain the final classification result.
Biao Hou, Yuheng Jiang, Bo Ren 0001, Zaidao Wen, Shuang Wang 0001, Licheng Jiao
IGARSS5
2016 Using deep neural networks for synthetic aperture radar image registration
abstract
At present, the performance of image registration mainly depends on the extracted features in feature-based image registration. However, due to the speckle noise, synthetic aperture radar (SAR) image registration will have a lower accuracy and less robustness. For this purpose, we design a deep neural network (DNN) for SAR image registration, using the DNN to learn the image features, automatically. The deep learning could learn the more essential features of the images, which are make the image registration to achieve more robust features and accurate matching. Moreover, this paper proposed a new strategy to remove the wrong matching points based on the RANSAC. The experimental results on SAR image registration show that this image registration method based on DNN have a better performance, and the new RANSAC strategy could eliminate many wrong matching points and get a good transformational model.
Dou Quan, Shuang Wang 0001, Mengdan Ning, Licheng Jiao
IGARSS2
2016 Learning task-driven polarimetric target decomposition: A new perspective
abstract
Polarimetric target decomposition aims to decompose a polarimetric synthetic aperture (PolSAR) data on a base reflecting some scattering mechanisms. The corresponding coefficients will be further exploited as the feature vector for the subsequent interpretation task. Intuitively, its performance heavily depends on the choice of bases and many off-the-shelf ones have been constructed based on mathematical or physical model since last two decades. However, these fixed bases are generally insufficient to characterize all types of data in a PolSAR image so that the extracted features are not beneficial to the subsequent task. To address this issue, we propose a novel target decomposition framework to learn a set of task-desired bases as well as feature vectors from the input polarimetric data. Focusing on the classification task, involve a supervised regularizer is further involved in our framework to increase the discrimination of features. Experimental results demonstrate the effectiveness of proposed framework.
Zaidao Wen, Biao Hou, Shuang Wang 0001, Licheng Jiao
IGARSS3
2016 Local graph regularized sparse reconstruction for salient object detection
Lina Huo, Shuyuan Yang 0001, Licheng Jiao, Shigang Wang 0001, Shuang Wang 0001
Neurocomputing5
2016 Weighted multifeature hyperspectral image classification via kernel joint sparse representation
Erlei Zhang, Xiangrong Zhang, Licheng Jiao, Hongying Liu 0001, Shuang Wang 0001, Biao Hou
Neurocomputing5
2016 Locality-constraint discriminant feature learning for high-resolution SAR image classification
Licheng Jiao, Biao Hou, Shuang Wang 0001, Jiaqi Zhao 0001, Puhua Chen
Neurocomputing4
2016 Object-level saliency detection with color attributes
Lina Huo, Licheng Jiao, Shuang Wang 0001, Shuyuan Yang 0001
Pattern Recognit.3
2015 External and internal learning for single-image super-resolution
abstract
Super-resolution (SR) problem still faces a challenge of wisely utilizing diverse learned priors to recover the lost details in low resolution images. In this work, we propose a novel method using low rank decomposition which integrates diverse priors learned from external and internal learning to construct SR image. The proposed method first applies an external dictionary learning to get the meta-detail that is commonly shared among images, and then introduces an internal prior learning to learn the local self-similarity (local structure) that is shared in the image. Both are essential but different priors for SR image construction. With these priors, a bank of preliminary HR images are obtained but with estimation errors and noise. To restrain the errors and noise, we consider these HR images as a high dimension data in dimension reduction problem, and solve it using a low rank decomposition. Experimental results show the proposed method preserves image details effectively, also outperforms state-of-the-arts in both visual and quantitative assessments, especially in dealing with the noise.
Shuang Wang 0001, Chris S. Lin, Xuefeng Liang, Bo Yue, Licheng Jiao
ICIP1
2015 Wishart RBM based DBN for polarimetric synthetic radar data classification
abstract
Deep Belief Network (DBN) is a classic deep learning model, and it can learn higher feature and do better classification job. We combine DBN's basic component Restricted Boltzmann Machines (RBM) with the statistic distribution of Polarimetric SAR (PolSAR) data. Based on it, we develop a deep learning classification method that is suitable for PolSAR data. To verify the effectiveness of the method, a real PolSAR dataset is tested. Experiment result confirms that the proposed method provides fine improvements both in classification accuracy and visual effect.
Yanhe Guo, Shuang Wang 0001, Chenqiong Gao, Danrong Shi, Biao Hou
IGARSS2
2015 Polarimetric SAR images classification using deep belief networks with learning features
abstract
A novel polarimetric synthetic aperture radar (PolSAR) image classification method based on Deep Belief Networks (DBNs) is proposed in this paper. First, the coherency matrix data are converted to a 9-dimentional data. Second, many patches are randomly selected from each dimension in the 9-dimentional data, and many filters can be obtained from a Restricted Boltzmann Machine (RBM) trained by using these patches. Thus we can get the features for each pixel from each dimension in the 9-dimentional space. Finally, the learned features and the elements of coherent matrix are combined to train a 3-layers DBNs for PolSAR image classification. Experimental results show that the proposed method is efficient and effective for PolSAR image classification.
Biao Hou, Xiaohuan Luo, Shuang Wang 0001, Licheng Jiao, Xiangrong Zhang
IGARSS3
2015 Semi-supervised classification based on anchor-spatial graph for large polarimetric SAR data
abstract
Recently a few works of semi-supervised learning methods based on graph have been proposed for remote sensing. The common idea of these methods are that they build a graph using the samples of the image. Most of their time complexity is relatively large, and they ignore the spatial information of the image, which leads to unsatisfactory classification results. this paper proposes a novel semi-supervised classification method based on anchor-spatial graph for large PolSAR data. Firstly the unsupervised Wishart clustering is performed to select representative samples, which served as anchors according to the least distance between samples. Then an anchor graph is built using the selected anchors according to the multiple features of the samples. And it is further combined with the spatial information of the samples to construct an anchor-spatial graph. Finally the class information from small quantities of labeled samples propagates to the unlabeled ones. Experimental results show that the proposed method has a low time complexity compared with existing works and it could effectively cut down the processing time for large PolSAR data meanwhile keeps the classification accuracy.
Hongying Liu 0001, Dexiang Zhu, Shuyuan Yang 0001, Shuang Wang 0001, Biao Hou, Licheng Jiao
IGARSS5
2015 A Three-Component Fisher-Based Feature Weighting Method for Supervised PolSAR Image Classification
abstract
This letter presents a feature weighting method for polarimetric synthetic aperture radar (PolSAR) image classification. Appropriate feature weighting is essential for obtaining accurate classifications but so far has remained an open research problem. We propose in this letter a supervised three-component feature weighting method based on the Fisher linear discriminant. Fisher linear discriminant method is used to calculate a coefficient for each feature. Then, these coefficients are modified according to a three-component scattering power decomposition model, combining both physical and statistical scattering characteristics to adapt them for the particular scattering mechanisms inherent in PolSAR data and assigned to the coherency matrix to enhance the discriminating ability of the features. Freeman decomposition and Wishart classifier are used to classify the PolSAR image. The effectiveness of the proposed method is demonstrated by experiments NASA/JPL AIRSAR L-band and CSA Radarsat-2 C-band PolSAR images of the San Francisco area.
Bo Chen 0001, Shuang Wang 0001, Licheng Jiao, Rustam Stolkin, Hongying Liu 0001
IEEE Geosci. Remote. Sens. Lett.2
2015 A novel dynamic rough subspace based selective ensemble
Yuwei Guo 0001, Licheng Jiao, Shuang Wang 0001, Shuo Wang 0005, Fang Liu 0001, Kaixuan Rong
Pattern Recognit.3
2015 A new patch based change detector for polarimetric SAR data
Ganchao Liu, Licheng Jiao, Fang Liu 0001, Hua Zhong 0003, Shuang Wang 0001
Pattern Recognit.5
2015 Learning Interpolation via Regional Map for Pan-Sharpening
abstract
Although the bandwidth of the high-resolution panchromatic (HR PAN) image is wide, it is narrow in each band of the low-resolution multispectral (LR MS) image. Hence, the spatial resolution of the HR PAN image is much higher than that of the LR MS image. However, HR PAN image only has a single band. The purpose of the Pan-sharpening algorithm is to make the Pan-sharpened image with both high spatial resolution and good spectral information. In this paper, a novel learning interpolation method for Pan-sharpening is proposed by expanding the sketch information in the HR PAN image. The sketch information contains the edges and lines features of the image, and each segment of the sketch information has its own direction. According to the primal sketch graph of the HR PAN image, a regional map is obtained by a designed geometrical template. Since the size of the HR PAN image is different from that of the LR MS image, the LR MS image is interpolated into an interpolated multispectral (IMS) image by the nearest interpolation method. In addition, the IMS image can be mapped into the structure and the nonstructure regions by this regional map. The nonstructure regions are divided into the smooth and the texture regions by a variance value. For the structure and texture regions, the interpolated pixels in the IMS image are relearned and readjusted by the proposed structure and texture learning interpolation method, respectively. Experimental results show that the proposed Pan-sharpening method can provide superior performance in both visual effect and quality metrics, particularly for the images with a large spectral difference.
Cheng Shi 0002, Fang Liu 0001, Lingling Li 0002, Licheng Jiao, Yiping Duan, Shuang Wang 0001
IEEE Trans. Geosci. Remote. Sens.6
2015 A Resample-Based SVA Algorithm for Sidelobe Reduction of SAR/ISAR Imagery With Noninteger Nyquist Sampling Rate
abstract
A resample-based spatial variant apodization (SVA) algorithm for sidelobe reduction was studied for synthetic aperture radar (SAR) and inverse SAR (ISAR) imagery with a noninteger Nyquist sampling rate. The weighting function of every sample in the image domain was calculated with the sample and two adjacent noninteger samples. The noninteger samples were obtained by interpolation in the image domain using sinc function. With the proper selection of two noninteger samples, the monotonic property of the weighting function on each side of the sampling point was preserved. The unequivocal determination of sidelobe suppression was achieved for noninteger Nyquist sampled (NINS) SAR and ISAR imagery. In addition, the lower and upper boundaries of the weighting function under the cosine-on-pedestal condition were extended for further sidelobe suppression and main lobe sharpening. The algorithm was implemented and applied to NINS imagery that is simulated. The algorithm was then assessed for acquired SAR and ISAR images. Improved results have been qualitatively and quantitatively achieved in sidelobe suppression and main lobe sharping in comparison with an existing algorithm.
Shuang Wang 0001, Biao Hou, Yong Wang 0011, Hongying Liu 0001
IEEE Trans. Geosci. Remote. Sens.2
2014 Pol-SAR image classification using eigenvalue-based joint statistical framework
abstract
In this paper, a novel method based on joint statistical framework is proposed for classification of polarimetric SAR image. The Gaussian model of the maximum eigenvalue and volume scattering power for coherency matrix is estimated to describe their statistical distribution. And Bayesian classifier is used to classify the polarimetric SAR image. In order to make full use of the local context structure of image, the local statistical model is used based on the maximum posterior probability (MAP) rule. The method is tested with the NASA/JPL AIRSAR data.
Shuiping Gou, W. F. Wang, Licheng Jiao, Shuang Wang 0001, X. R. Zhang
IGARSS4
2014 SAR image segmentation based on random projection and signature frame
abstract
This paper proposes a new Synthetic Aperture Radar (SAR) image segmentation method based on the frame of Signature/Earth Mover's Distance (EMD). Firstly, Random Projection is used to extract features of SAR image, which has the abilities of preserving information and reducing dimensionality. Secondly, a signature is used to obtain the cluster center and the weight. Finally, by computing the distance between two signatures using Earth Mover's Distance, we can obtain the final segmentation result. The experimental results show that the proposed method is efficient and effective for SAR image segmentation.
Biao Hou, Shuang Wang 0001, Xiangrong Zhang
IGARSS3
2014 MSTAR image segmentation with multi-phase level set based on probability density model
abstract
Radar image segmentation is a fundamental problem in radar image interpretation. Radar images often contain a great deal of noise. Level set method, known as deformable model, is a powerful image segmentation technique. It can get accurate contours of clear-cut objects in image without noise, but has poor performance in getting contours of objects in a noisy image. In this paper, a new multi-phase level set based on probability density model is proposed. We use histogram, a non-parametric density estimation method, to describe the statistical information of each pixel in its neighborhood and the pixels in each subset in the image. The comparability between them computed by inner product function is used as the curve energy in multi-phase level set method. The statistical information is incorporated into the multi-phase level set framework, which can cope with the influence of noise on image segmentation. This new method is particularly well adapted to detection of objects of interesting in a noisy image. We illustrated the performance of the new method on MSTAR images. The experimental results show that incorporating statistical information into the multi-phase level set framework, consistent objects are obtained, and accurate and robust segmentations can be achieved.
Xiaojin Hou, Shuang Wang 0001, Biao Hou
IGARSS3
2014 Nonlocal filtering for Polarimetric SAR Data based on bilateral filtering
abstract
In this paper, we introduce a nonlocal filtering method for Polarimetric SAR Data based on bilateral filtering. This method is a combination of local and nonlocal filtering. In order to keep the spatial structure of image, we select similar blocks in a large area; and to adapt to the pixel similarities, the iterative bilateral filtering is applied to the selected similar blocks. To deal with polarimetric data, we propose a new similarities based on complex Wishart distance. We demonstrate the performance of the proposed algorithm by using the experimental data. The experiment shows the better performance of the proposed method. There are fewer speckle noise in homogeneous areas. And some detail of structures, i.e., edges, lines, points, and curves, are well protected.
Xiaozhen Lei, Shuang Wang 0001, Kun Liu 0011, Biao Hou
IGARSS2
2014 Unsupervised classification of polarimetric SAR images integrating color features
abstract
In conventional terrain classification for the polarimetric SAR (POLSAR) images, color features are rarely involved unless in one recent supervised work. Unlike that work, the color features are exploited for the unsupervised classification in this paper. Firstly, based on the polarimetric decomposition of the POLSAR data, the common color spaces, such as RGB, HSI, and CIELab are calculated. The color feature is quantitatively selected from these color spaces by introducing the color entropy. Then together with the spatial information, extended scattering power entropy and the copolarized ratio, the adaptive Mean-shift algorithm is used to segment the POLSAR image. Finally, the segments are merged according to the Wishart distance measurement. The experiments using AIRSAR L-band POLSAR data indicate that the proposed method has better discriminative ability for urban areas and for boundary preservation compared with existing works.
Hongying Liu 0001, Shuang Wang 0001, Biao Hou, Shuyuan Yang 0001, Junfei Shi, Licheng Jiao
IGARSS2
2014 A multiscale region-based approach to automatic SAR image registration using CLPSO
abstract
In this paper, a novel approach to automatic SAR image registration is proposed. First, the bi-temporal images are decomposed by wavelet transform. Then, the sensed image and its wavelet approximation component are divided to some saliency region based on spectrum residua approach. And the saliency regions are fitted into ellipses. Finally, the region matching is performed by searching the corresponding regions in the reference image at every layer by the comprehensive learning particle swarm optimization (CLPSO) twice. Thus, the registration in every layer is a coarse-to-fine process, which effectively improves the efficiency, accuracy and stability. The proposed technique is validated by testing on real SAR images and compared with existing state-of-the-art techniques.
Guiting Wang, Xiaoxian Liu, Licheng Jiao, Shuang Wang 0001
IGARSS4
2014 Multilayer feature learning for polarimetric synthetic radar data classification
abstract
Features are important for polarimetric synthetic aperture radar (PolSAR) image classification. Various methods focus on extracting feature artificially. Compared with them, we have developed a method to learn feature automatically. The method is based on deep learning which can learn multilayer features. In this paper, stacked sparse autoencoder (SAE) as one of the deep learning models is applied as a useful strategy to achieve the goal. For improving the classification result, we use a small amount of labels to fine-tuning the parameters of the proposed method. Finally, a real PolSAR dataset is used to verify the effectiveness. Experiment result confirms that the proposed method provides noteworthy improvements in classification accuracy and visual effect.
Huiming Xie, Shuang Wang 0001, Kun Liu 0011, Chris S. Lin, Biao Hou
IGARSS2
2014 Improve the performance of co-training by committee with refinement of class probability estimations
Shuang Wang 0001, Linsheng Wu, Licheng Jiao, Hongying Liu 0001
Neurocomputing1
2014 Classification Method for Fully PolSAR Data Based on Three Novel Parameters
abstract
In this letter, a new classification method for fully polarimetric synthetic aperture radar (PolSAR) data based on three novel parameters is presented. The three parameters are derived from the eigenspace of the coherency matrix as linear combinations of its three eigenvalues. In the proposed classification method, the maximum value out of the three parameters is determined to assign a label to each image pixel, and the PolSAR image is classified into three classes accordingly. Experimental results based on NASA/JPL AIRSAR L-band data and CSA RADARSAT-2 C-band data illustrate the validity and efficacy of the procedure.
Shuang Wang 0001, Bo Chen 0001, Shasha Mao
IEEE Geosci. Remote. Sens. Lett.2
2014 Improving Hyperspectral Image Classification Using Spectral Information Divergence
abstract
In order to improve the classification performance for hyperspectral image (HSI), a sparse representation classifier based on spectral information divergence (SID) is proposed. SID measures the discrepancy of probabilistic behaviors between the spectral signatures of two pixels from the aspect of information theory, which can be more effective in preserving spectral properties. Thus, the new method measures the similarity between the reconstructed pixel and the true pixel by SID instead of by the L2 norm used in traditional sparse model. Moreover, the spatial coherency across neighboring pixels sharing a common sparsity pattern is taken into account during the construction of SID-based joint sparse representation model. We propose a new version of the orthogonal matching pursuit method to solve SID-based recovery problems. The proposed SID-based algorithms are applied to real HSI for classification. Experimental results show that our algorithms outperform the classical sparse representation based classification algorithms in most cases.
Erlei Zhang, Xiangrong Zhang, Shuyuan Yang 0001, Shuang Wang 0001
IEEE Geosci. Remote. Sens. Lett.4
2014 A compressed sensing approach for efficient ensemble learning
Lin Li 0016, Rustam Stolkin, Licheng Jiao, Fang Liu 0001, Shuang Wang 0001
Pattern Recognit.5
2014 A Novel Coarse-to-Fine Scheme for Automatic Image Registration Based on SIFT and Mutual Information
abstract
Automatic image registration is a vital yet challenging task, particularly for remote sensing images. A fully automatic registration approach which is accurate, robust, and fast is required. For this purpose, a novel coarse-to-fine scheme for automatic image registration is proposed in this paper. This scheme consists of a preregistration process (coarse registration) and a fine-tuning process (fine registration). To begin with, the preregistration process is implemented by the scale-invariant feature transform approach equipped with a reliable outlier removal procedure. The coarse results provide a near-optimal initial solution for the optimizer in the fine-tuning process. Next, the fine-tuning process is implemented by the maximization of mutual information using a modified Marquardt-Levenberg search strategy in a multiresolution framework. The proposed algorithm is tested on various remote sensing optical and synthetic aperture radar images taken at different situations (multispectral, multisensor, and multitemporal) with the affine transformation model. The experimental results demonstrate the accuracy, robustness, and efficiency of the proposed algorithm.
Maoguo Gong, Shengmeng Zhao, Licheng Jiao, Dayong Tian, Shuang Wang 0001
IEEE Trans. Geosci. Remote. Sens.5
2014 Local Maximal Homogeneous Region Search for SAR Speckle Reduction With Sketch-Based Geometrical Kernel Function
abstract
With the flourish of the nonlocal mean method, the neighborwise similarity metric is widely applied in speckle reduction for its robust performance on the search of similar samples. In this metric, an isotropic kernel function is usually chosen to aggregate the corresponding pixels' distance between two neighborhoods. It means that the kernel function is considered as the explanation of the local spatial relationship at each pixel. However, for anisotropic features (such as edges and lines), a strong relationship exists along their directions rather than across them, so the isotropic kernel is not suitable to explain the spatial relationship around these features. Meanwhile, due to the inherent speckle in synthetic aperture radar (SAR) images, the discrimination and exploration of the geometrical properties of anisotropic features are important for the construction of adaptive kernel function. In this paper, the sketch map which is a representation of the sketch information of SAR images is extracted as the criterion for designing the kernel function. Meanwhile, due to the properties of symmetric and maximal self-similarity, a modified ratio distance is proposed and used jointly with the constructed kernel function as a similarity metric. Then, under the local stationary assumption, the local maximal homogeneous region of each pixel is searched by using the region growing method with the proposed metric. Moreover, maximal likelihood rule is used within the region for the estimation of true value. From the experiments on the synthetic and real SAR images, a promising performance in terms of speckle reduction and preservation of the details is achieved by our proposed method.
Jie Wu 0016, Fang Liu 0001, Licheng Jiao, Xiangrong Zhang, Hongxia Hao, Shuang Wang 0001
IEEE Trans. Geosci. Remote. Sens.6
2014 Minimum-Entropy-Based Autofocus Algorithm for SAR Data Using Chebyshev Approximation and Method of Series Reversion, and Its Implementation in a Data Processor
abstract
A novel autofocus method for synthetic aperture radar (SAR) image is studied. Based on a quadratic model for the phase error within each sub-area (narrow strip × sub-aperture) after a wide range swath is subdivided into narrow range strips and long azimuth aperture into sub-apertures, an objective function for estimation of the error is derived through the principle of minimum entropy. There is only one unknown variable in the function. With the Chebyshev approximation, the function is approximated as a polynomial, and the unknown is then solved using the method of series reversion. Curve-fitting methods are applied to estimate phase error for an entire scene of the full-swath by full-aperture. Through simulations, the proposed method is applied to restore the defocused SAR imagery that is well focused. The restored and original images are almost identical qualitatively and quantitatively. Next, the method is implemented into an existing SAR data processor. Two sets of SAR raw data at X- and Ku-bands are processed and two images are formed. Well-focused and high-resolution images from plain and rugged terrain are obtained even without the use of ancillary attitude data of the airborne SAR platform. Thus, the studied method is verified.
Mengdao Xing, Yong Wang 0011, Shuang Wang 0001, Jialian Sheng, Liang Guo 0002
IEEE Trans. Geosci. Remote. Sens.4
2014 A Novel Eye Localization Method With Rotation Invariance
abstract
This paper presents a novel learning method for precise eye localization, a challenge to be solved in order to improve the performance of face processing algorithms. Few existing approaches can directly detect and localize eyes with arbitrary angels in predicted eye regions, face images, and original portraits at the same time. To preserve rotation invariant property throughout the entire eye localization framework, a codebook of invariant local features is proposed for the representation of eye patterns. A heat map is then generated by integrating a 2-class sparse representation classifier with a pyramid-like detecting and locating strategy to fulfill the task of discriminative classification and precise localization. Furthermore, a series of prior information is adopted to improve the localization precision and accuracy. Experimental results on three different databases show that our method is capable of effectively locating eyes in arbitrary rotation situations (360° in plane).
Yan Ren 0002, Shuang Wang 0001, Biao Hou
IEEE Trans. Image Process.2
2013 SAR image ship detection based on visual attention model
abstract
This paper proposes a novel Synthetic Aperture Radar (SAR) image ship detection method based on human visual attention mechanism. Firstly, we obtain water segmentation image by combining the bottom-up and the top-down visual attention mechanisms. Secondly, we detect ship targets based on bottom-up the visual attention mechanism. The interested regions are extracted by measuring the visual conspicuity of each water regions. Then, the ships targets are detected in the interested regions by the k-means clustering algorithm. Finally, real SAR image is used to test our algorithm. Besides, we analysis the ship detection results using different band. The experiment results indicate that our algorithm can effectively detect ship targets from SAR images and C-band is superior to L-band in SAR image ship detection.
Biao Hou, Shuang Wang 0001, Xiaojin Hou
IGARSS3
2013 Hurricane eye extraction from SAR image using saliency-based visual attention algorithm
abstract
Automatic hurricane information extraction in synthetic aperture radar (SAR) images has been a research topic in development. In this study, using saliency-based visual attention model, we developed an image processing procedure to extract hurricane eyes from SAR images. Experiment results show that hurricane eyes can be well extracted even when it is not visually obvious in images.
Shaohui Jin, Xiaofeng Li 0001, Shuang Wang 0001
IGARSS3
2013 Unsupervised classification of POLSAR data based on the improved affinity propagation clustering
abstract
In this paper, the AP clustering algorithm is improved by defining a new similarity to be applied in the polarimetric SAR image classification. On this basis, a new unsupervised classification method is proposed which combines the Four-component decomposition and the improved AP clustering. The proposed method mainly consists of three steps: Firstly, Four-component decomposition is adopted to produce initial segmentation. Secondly, the improved affinity propagation clustering based on the Wishart distance measure is applied on the initial segmentation to merge clusters and obtain an appropriate number of categories. Finally, an iterative algorithm based on the complex wishart density function is applied. The effectiveness of this algorithm is demonstrated by the test with NASA/JPL AIRSAR L-band data of San Francisco and Flevoland.
Shuang Wang 0001, Yachao Liu, Kun Liu 0011, Xiaojin Hou, Biao Hou
IGARSS1
2013 High resolution SAR target reconstruction from compressive measurements with prior knowledge
abstract
In this paper, an effective prior knowledge based framework for target reconstruction from compressive measurements is proposed. In this framework, a traditional compressed imaging method is firstly introduced which indicates that for a range cell containing K strongest scattering points can be reconstructed based on the theory of compressive sensing. Secondly, a greedy iteration algorithm is modified which utilizes some prior knowledge of the target during the reconstruction step. The experiments are carried on the Moving and Stationary Target Acquisition and Recognition (MSTAR) database and the results show the effectiveness of our framework for target reconstruction.
Zaidao Wen, Biao Hou, Shuang Wang 0001
IGARSS3
2013 Nonlocal-Lee filter for SAR image despeckling based on hybrid patch similarity
abstract
This paper presents a novel nonlocal Lee (NL-Lee) filter based on both structure similarity and homogeneity similarity, namely hybrid patch similarity, which can effectively enhance the patch regularity assumption. The proposed NL-Lee filter shares the framework of the NLM filter but combines the traditional Lee filter in a distributive way. Structure similarity from the NLM filter and homogeneity similarity from the Lee filter are well balanced in the new filter. As a result, the new NL-Lee filter obtains good trade-off between speckle smoothing and detail preservation and leads to state-of-the-art results.
Hua Zhong 0003, Shuang Wang 0001, Xiaojing Hou
IGARSS4
2013 Unsupervised Classification of Fully Polarimetric SAR Images Based on Scattering Power Entropy and Copolarized Ratio
abstract
This letter presents a new unsupervised classification method for polarimetric synthetic aperture radar (POLSAR) images. Its novelties are reflected in three aspects: First, the scattering power entropy and the copolarized ratio are combined to produce initial segmentation. Second, an improved reduction technique is applied to the initial segmentation to obtain the desired number of categories. Finally, to improve the representation of each category, the data sets are classified by an iterative algorithm based on a complex Wishart density function. By using complementary information from the scattering power entropy and the copolarized ratio, the proposed method can increase the separability of terrains, which can be of benefit to POLSAR image processing. Three real POLSAR images, including the RADARSAT-2 C-band fully POLSAR image of western Xi'an, China, are used in the experiments. Compared with the other three state-of-the-art methods,$\hbox{H}/\alpha$-Wishart method, Lee category-preserving classification method, and Freeman decomposition combined with the scattering entropy method, the final classification map based on the proposed method shows improvements in the accuracy and efficiency of the classification. Moreover, high adaptability and better connectivity are observed.
Shuang Wang 0001, Kun Liu 0011, Jingjing Pei, Maoguo Gong, Yachao Liu
IEEE Geosci. Remote. Sens. Lett.1
2013 Fast Fisher Sparsity Preserving Projections
Licheng Jiao, Fanhua Shang, Shuang Wang 0001, Biao Hou
Neural Comput. Appl.4
2013 Selective multiple kernel learning for classification with ensemble strategy
Tao Sun 0007, Licheng Jiao, Fang Liu 0001, Shuang Wang 0001, Jie Feng 0003
Pattern Recognit.4
2013 Context-Based Hierarchical Unequal Merging for SAR Image Segmentation
abstract
This paper presents an image segmentation method named Context-based Hierarchical Unequal Merging for Synthetic aperture radar (SAR) Image Segmentation (CHUMSIS), which uses superpixels as the operation units instead of pixels. Based on the Gestalt laws, three rules that realize a new and natural way to manage different kinds of features extracted from SAR images are proposed to represent superpixel context. The rules are prior knowledge from cognitive science and serve as top-down constraints to globally guide the superpixel merging. The features, including brightness, texture, edges, and spatial information, locally describe the superpixels of SAR images and are bottom-up forces. While merging superpixels, a hierarchical unequal merging algorithm is designed, which includes two stages: 1) coarse merging stage and 2) fine merging stage. The merging algorithm unequally allocates computation resources so as to spend less running time in the superpixels without ambiguity and more running time in the superpixels with ambiguity. Experiments on synthetic and real SAR images indicate that this algorithm can make a balance between computation speed and segmentation accuracy. Compared with two state-of-the-art Markov random field models, CHUMSIS can obtain good segmentation results and successfully reduce running time.
Xiangrong Zhang, Shuang Wang 0001, Biao Hou
IEEE Trans. Geosci. Remote. Sens.3
2012 A visual attention model based on wavelet transform and its application on ship detection
abstract
Human visual system is very efficient and selective in scene analysis, which has been widely used in image processing. In this paper, a new visual attention model based on dyadic wavelet transform (DWT) used for ship detection is proposed. It is a bottom-up visual attention model driven by data rather than by task. First, the input image is converted from RGB color space to HIS color space. Second, the modulus of DWT is analyzed to obtain the conspicuity map of each feature. Third, the conspicuity maps are combined into the saliency map nonlinearly, different from Itti's method, the contribution rate of each conspicuity map to final saliency map is not equal. It is relevant to the difference between the level of the most active region and the average level of the other active regions in each conspicuity map. Finally, the detection result of ships based on saliency map is got by region growing method, where the seed is obtained from the saliency map and the growing process is implemented in intensity image. Experiments on natural ship images show that our method is robust and efficient compared with Itti's and Hou's method.
Biao Hou, Shuang Wang 0001
IGARSS3
2012 Low-rank and sparse matrix decomposition-based pan sharpening
abstract
This paper proposes a remote sensing image pan-sharpening method from the perspective of low-rank and sparse matrix decomposition. Based on the characteristic of multispectral (MS) images, the low spatial resolution information of MS images is modeled as low-rank, and the high spectral resolution information of MS images is modeled as sparse. First, the low-rank and sparse matrix decomposition algorithm is applied to the resampled MS images to extract the sparse component i.e. the high spectral resolution information. Second, the standard PCA fusion method is applied on the low-rank component to obtain the rough pan-sharpened MS images. Finally, adding the sparse MS images component on the rough result and one can get the final fused product. Experimental results demonstrate that the proposed method is competitive or even better than some other methods.
Kaixuan Rong, Shuang Wang 0001, Biao Hou
IGARSS2
2012 SAR image despeckling method using bivariate shrinkage based on dual-tree complex wavelet
abstract
In this paper, we propose a speckle suppression method for SAR image based on dual-tree complex wavelet. Non-Gaussian bivariate distribution model is proposed by considering the correlation of the real and imaginary parts of complex wavelet coefficients. Based on this model, the shrinkage function of real and imaginary parts of complex coefficients is obtained with the help of maximum a posteriori estimate. Compared with the state-of-the-art techniques through the visual effect and the equivalent number of looks (ENL), experimental results demonstrate that the proposed algorithm obtains good performance in smoothing speckles of homogeneous regions and preserving edges and details effectively.
Shuang Wang 0001, Biao Hou
IGARSS1
2011 Unsupervised classification of POLSAR data based on the polarimetric decomposition and the co-polarization ratio
abstract
In this paper, a new classification scheme for polarimetric SAR data sets is presented. The proposed method mainly involves two concepts: Freeman-Durden decomposition and co-polarization ratio. The core concept of Freeman-Durden decomposition is to decompose the covariance matrix into three scattering mechanisms: surface scattering, double bounce scattering and volume scattering; the co-polarization radio represents the proportion of horizontal polarization and vertical polarization. The combination of these two characters can distinguish different vegetation types effectively. The proposed method mainly consists of three steps: first, apply Freeman-Durden decomposition to get the scattering characters; second, combine the scattering powers and co-polarization radio to divide the images into corresponding initial clusters; finally, improve the representation of each class, the data sets of which are classified by an iterative algorithm based on a complex Wishart density function. The effectiveness of this algorithm is demonstrated by two sets of data: NASA/JPL AIRSAR L-band data of San Francisco and Flevoland in The Netherlands by NASA/JPL AIRSAR in 1989.
Shuang Wang 0001, Jingjing Pei, Kun Liu 0011, Bo Chen 0001
IGARSS1
2011 Water/land segmentation for sar images based on geodesic distance
abstract
In this paper, a novel method for water/land segmentation is proposed based on the framework of geodesic distance. The proposed method models the water/land according to the statistics of both the speckle and land covers, which leads to a fast point-wised coarse segmentation. Based on the water/land models, the boundary area between water and land can be localized with automatically generated class labels and adaptively determined bandwidth. Then the refined segmentation is implemented using an improved geodesic distance, combining the manifold idea to enlarge the inter-class differences. Experimental results on real synthetic aperture radar (SAR) images demonstrate the effectiveness and efficiency of the method. The bridges, harbors and coastline can be segmented correctly with very tiny details preserved.
Hua Zhong 0003, Qing Xie 0006, Licheng Jiao, Shuang Wang 0001
IGARSS4
2011 Shape-Adaptive Reversible Integer Lapped Transform for Lossy-to-Lossless ROI Coding of Remote Sensing Two-Dimensional Images
abstract
In this letter, we propose a shape-adaptive (SA) reversible integer lapped transform (SA-RLT) method. The new method can deal with arbitrarily shaped image areas while guaranteeing completely reversible integer-to-integer transform. Based on SA-RLT and object-based set partitioned embedded block coder, a new region-of-interest (ROI) compression scheme is designed for 2-D remote sensing images. Numerical experiments reveal that SA-RLT performs better than integer SA discrete wavelet transform, and the new ROI compression scheme performs comparably even better than the JPEG2000-ROI scheme. Advantages in hardware implementation have been preserved by SA-RLT, such as parallel processing and low memory requirement.
Licheng Jiao, Lei Wang 0018, Jiaji Wu, Jing Bai 0003, Shuang Wang 0001, Biao Hou
IEEE Geosci. Remote. Sens. Lett.5
2010 SAR Image Despeckling Using Edge Detection and Feature Clustering in Bandelet Domain
abstract
To effectively preserve the edges of a synthetic aperture radar (SAR) image when despeckling, an algorithm with edge detection and fuzzy clustering in the translation-invariant second-generation bandelet transform (TIBT) domain is proposed in this letter. A Canny operator is first utilized to detect and remove edges from the SAR image. Then, TIBT and fuzzy C-mean clustering are employed to decompose and despeckle the edge-removed image, respectively. Finally, the removed edges are added to the reconstructed image. The algorithm suggests each coefficient in high-frequency subbands as the clustering feature, proposes a calculation method of the best clustering number, and defines the signal and noise in the clustering results. Experimental results show that the visual quality and evaluation indexes outperform the other methods with no edge preservation. The proposed algorithm effectively realizes both despeckling and edge preservation and reaches the state-of-the-art performance.
Wenge Zhang, Fang Liu 0001, Licheng Jiao, Biao Hou, Shuang Wang 0001, Ronghua Shang
IEEE Geosci. Remote. Sens. Lett.5
2005 Radar target recognition using SVMs with a wrapper feature selection driven by immune clonal algorithm
Xiangrong Zhang, Shuang Wang 0001, Tan Shan, Licheng Jiao
ESANN2
2005 Robust Classification of Immunity Clonal Synergetic Network Inspired by Fuzzy Integral
Xiuli Ma, Shuang Wang 0001, Licheng Jiao
ISNN (2)2
2005 Image Representation in Visual Cortex and High Nonlinear Approximation
Tan Shan, Xiangrong Zhang, Shuang Wang 0001, Licheng Jiao
ISNN (1)3