Xiangrong Zhang

dblp:52/5237 · DBLP profile ↗
← Back
252ranked-venue papers
43as first author
137since 2021 · last 2026
0000-0003-0379-2042ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 175 · 29 first-author · 97 since 2021Artificial intelligence and machine learning · 51 · 9 first-author · 28 since 2021Graphics, computer vision, multimedia, augmented reality and games · 22 · 3 first-author · 15 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-authorSystems, architecture and hardware · 3 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 first-authorTheory of computation · 1
YearPublicationVenuePosition
2026 ID-Splat: Propagating Object Identities for Segmenting 3D Aerial-view Scenes
abstract
High-resolution Earth Observation technologies present unprecedented opportunities for geospatial analysis, yet traditional 2D aerial-view semantic segmentation remains limited by its inability to model spatial relationships and handle object occlusions. While 3D Aerial-view Segmentation (3DAS) has emerged to address these limitations, existing methods predominantly rely on 2D discriminative models pre-trained on natural scenes. These models struggle to accurately recognize aerial-view imagery, resulting in suboptimal performance due to significant domain discrepancies. This paper introduces ID-Splat, a novel object-centric framework that directly leverages multi-view object identities without discriminative information to enhance 3D semantic understanding. ID-Splat implements a two-stage process: first, Mask-object Tracking combines SAM and Point Tracking to establish robust and consistent object identities across multi-view aerial images; second, Object Integration & Propagation assigns these identities to 3D Gaussian Splatting (3DGS) points, enabling complete 3D segmentation through semantic propagation. Experimental results on the 3D-AS dataset demonstrate that ID-Splat significantly outperforms existing methods, particularly under sparse supervision conditions. ID-Splat also achieves state-of-the-art performance while reducing the need for extensive labeled data by effectively leveraging the inherent 3D structure.
Yijing Wang 0004, Xu Tang 0004, Xiangrong Zhang, Jingjing Ma 0001
AAAI3
2026 Softmatch distance: A novel distance for weakly-supervised trend change detection in bi-temporal images
Yuqun Yang, Xu Tang 0004, Xiangrong Zhang, Changzhe Jiao, Jingjing Ma 0001, Licheng Jiao
Pattern Recognit.3
2026 Tiny object detection based on dynamic scale-awareness label assignment and contextual enhancement
Tianyang Zhang 0002, Xiangrong Zhang, Chaozhuo Hua, Guanchun Wang, Xiao Han 0012, Licheng Jiao
Pattern Recognit.2
2026 Dual-Net: Dual Visual Spectral Affinity Monitoring Network for Hyperspectral Anomaly Detection
Xiangrong Zhang, Rongxia Qiu, Shiqi Wu, Guanchun Wang, Xiao Han 0012, Yifei Jiang, Licheng Jiao
IEEE Trans. Circuits Syst. Video Technol.1
2026 Cross-Image Federated Learning for Hyperspectral Image Classification
abstract
The contemporary research paradigm in remote sensing hyperspectral monitoring increasingly relies on multisatellite and multiplatform Earth observation. While the traditional hyperspectral research framework based on single-image processing (SIP) has facilitated the application of idealized scenarios and the development of standardized evaluation benchmarks, it inherently constrains the model's ability to generalize feature representations across varying spatial and temporal domains. As hyperspectral data applications grow in complexity and data requirements, the limitations of SIP in addressing the demands of modern remote sensing tasks become increasingly apparent. To overcome these research limitations, we utilize the decentralized nature and data security features of federated learning to propose a cross-image hyperspectral image (HSI) federated learning approach for classification tasks. We first develop a client-oriented self-guided knowledge-enhanced personalized learning method that enhances the personalization of the local learning process by leveraging relevant features from other clients, thereby improving the learning efficiency of each client. To address the issue of "bias" in global knowledge caused by uneven data distribution across the federated learning process, we introduce a multiscale semantic aligned dynamic aggregation method to ensure fairness in integrating global knowledge. To our knowledge, this article is the first to explore the joint learning of HSI classification using federated learning. Accordingly, we have constructed open-set and closed-set datasets tailored to this task and have demonstrated the effectiveness of our method on these datasets. The code is available at: https://github.com/Gallipaxi/FedHIC.
Xiangrong Zhang, Lijing Zheng, Guanchun Wang, Licheng Jiao
IEEE Trans. Neural Networks Learn. Syst.2
2026 PID: A Parameter-Efficient Isolation Domain-Incremental Learning Framework for Signal Modulation Classification
abstract
Deep neural networks have achieved promising progress in signal modulation classification (SMC), playing an essential role in a variety of applications such as cognitive radio networks, cyber defense, and electronic surveillance. However, most existing SMC methods still follow the traditional machine learning paradigm that trains on static closed datasets, lacking the ability to cope with the challenge of continuous data distribution shifts in real communication scenarios. Directly applying the model to a new environment may lead to severe degradation of classification performance on previous scenarios, i.e., catastrophic forgetting. To address this, this article proposes the first domain-incremental learning (DIL) paradigm for SMC and designs a parameter-efficient isolation DIL (PID) method, which enables SMC models to rapidly adjust to new scenarios by extending only a few parameters, while significantly retaining classification capabilities on previous scenarios. Specifically, we first propose a parameter space decomposition-based classifier (PSD), separating the model parameters into a set of bases and corresponding coefficients. By freezing the bases and fine-tuning the low-dimensional coefficients, the catastrophic forgetting problem can be efficiently eliminated. Furthermore, we design a scene-aware domain controller (SDC) to select the most suitable domain-specific coefficients for each sample, thereby maintaining the SMC model's classification capabilities across all domains. The extensive experimental results show the superiority of the proposed PID, which achieves state-of-the-art (SOTA) overall performance. The code will be available at: https://github.com/SMC-IL/PID.
Guanchun Wang, Xiangrong Zhang, Licheng Jiao
IEEE Trans. Neural Networks Learn. Syst.3
2025 Category-Specific Selective Feature Enhancement for Long-Tailed Multi-Label Image Classification
Ruiqi Du, Xu Tang 0004, Xiangrong Zhang, Jingjing Ma 0001
ICCV3
2025 Accelerating Paging via Novel User Prediction Mechanism with Spatial Indexed Correction
abstract
6G plans to deeply integrate Artificial Intelligence (AI) to support large-scale device deployment and optimize network management by predicting the UE-connected gNB. However, user prediction in 6 G mobile communication networks faces three major challenges, temporal complexity, spatial constraints, and latency sensitivity. This paper introduces a high-accuracy, high-efficient, and low-latency user prediction model, called RelNet, designed to address these challenges. RelNet consists of three modules. The UE Trajectory Patching (UTP) module first reduces computational complexity by transmitting UE trajectory data into patches. The Temporal Feature Extraction (TFE) module then captures both short-term fluctuations and long-term dependencies in massive UE trajectories. Finally, the Intelligent Positioning Refinement (IPR) module enables adaptive optimization and precise gNB positioning. Furthermore, this paper uses User Prediction Assisted Paging (UPAP) to validate the practicality of RelNet in 6 G, reducing signaling overhead by predicting the next-connected gNB and generate the secondary paging area (SPA). Experimental results demonstrate that RelNet achieves 69.98 % prediction accuracy, outperforming mainstream models. Moreover, the UPAP scheme effectively reduces signaling overhead by 69.72 % on a real-world 6 G dataset, confirming the feasibility of RelNet in 6 G networks.
Lijing Zheng, Xiangrong Zhang
ICPADS4
2025 RegionMatch: Pixel-Region Collaboration for Semi-Supervised Semantic Segmentation in Remote Sensing Images
abstract
Semi-supervised semantic segmentation (S4) has shown significant promise in reducing the burden of labor-intensive data annotation. However, existing methods mainly rely on pixel-level information, neglecting the strong region consistency inherent in remote sensing images (RSIs), which limits their effectiveness in handling the complex and diverse backgrounds of RSIs. To address this, we propose RegionMatch, a novel approach that leverages unlabeled data from a fresh object-level perspective, which is more tailored to the nature of semantic segmentation. We design the Pixel-Region Synergy Pseudo-Labeling strategy, which explicitly injects object-level contextual information into the S4 pipeline and promotes knowledge collaboration between pixel and region perspectives for generating high-quality pseudo-labels. In addition, we propose the Region Structure-Aware Correlation Consistency, which models object-level relationships by establishing inter-region correlations across images and pixel correlations within regions, providing more effective supervision signals for unlabeled data. Experimental results demonstrate that RegionMatch outperforms state-of-the-art methods on multiple authoritative remote sensing datasets, highlighting its superiority in the RSIs.
Xiaoqian Zhu, Xiangrong Zhang, Tianyang Zhang 0002, Chaowei Fang, Xu Tang 0004, Licheng Jiao
IJCAI2
2025 ACMamba: Fast Unsupervised Anomaly Detection via An Asymmetrical Consensus State Space Model
abstract
Unsupervised anomaly detection in hyperspectral images (HSI), aiming to detect unknown targets from backgrounds, is challenging for earth surface monitoring. However, current studies are hindered by steep computational costs due to the high-dimensional property of HSI and dense sampling-based training paradigm, constraining their rapid deployment. Our key observation is that, during training, not all samples within the same homogeneous area are indispensable, whereas ingenious sampling can provide a powerful substitute for reducing costs. Motivated by this, we propose an Asymmetrical Consensus State Space Model (ACMamba) to significantly reduce computational costs without compromising accuracy. Specifically, we design an asymmetrical anomaly detection paradigm that utilizes region-level instances as an efficient alternative to dense pixel-level samples. In this paradigm, a low-cost Mamba-based module is introduced to discover global contextual attributes of regions that are essential for HSI reconstruction. Additionally, we develop a consensus learning strategy from the optimization perspective to simultaneously facilitate background reconstruction and anomaly compression, further alleviating the negative impact of anomaly reconstruction. Theoretical analysis and extensive experiments across eight benchmarks verify the superiority of ACMamba, demonstrating a faster speed and stronger performance over the state-of-the-art. Code is released at https://github.com/PURE-melo/ACMamba.
Guanchun Wang, Xiangrong Zhang, Zelin Peng, Tianyang Zhang 0002, Xu Tang 0004, Licheng Jiao
ACM Multimedia2
2025 MIAF-Net: Multiscale interactive attention fusion network for hyperspectral image classification
Jinliang An, Longlong Dai, Weidong Zhang 0007, Xiangrong Zhang
Expert Syst. Appl.4
2025 Rural Tourist Attractions Recommendation Model Based on Multi-Feature Fusion Graph Neural Networks
abstract
With the rapid growth of the rural tourism industry, traditional tourism recommendation technologies can no longer meet the necessary requirements. To address the issue of rural tourist attraction recommendations, a rural tourist attraction recommendation model is constructed based on a multi-feature fusion graph neural network. First, construct a feature map based on the relationship between tourists’ preferences and tourist attractions, and incorporate the attention mechanism to enhance the model’s learning capabilities. Second, utilize a two-part graph model to extract positive and negative preference features of tourists, and a conversation graph model to extract tourists’ transfer preference features. Finally, various features are utilized to generate suggested content by computing scores for tourists’ travel preferences. To address the problem of recommending tourist groups, suitable features for random group matching are collected and the cosine function is employed to identify users with similar random group features. Finally, the multi-features are merged, and the tourists’ interest preferences are scored to arrive at content recommendations. In the experiment on individualized attraction recommendations, data from the Chengdu area were used to test the proposed model. The accuracy of the model’s recommendations was 0.822 for five recommendations which outperformed the other models. In the experiment for group-based attraction recommendations, this experiment tested the Chengdu dataset. The proposed model achieved the highest accuracy of 0.972 when the group size was 70, outperforming the other two models. Additionally, with regards to different numbers of recommendations, the proposed model’s accuracy was 0.5241, which was the best performance among the three models when the number of recommendations was set to five. The proposed recommendation model performs optimally in suggesting tourist attractions and meets the needs of rural tourism. The research content provides crucial technical references for tourist traveling and rural tourism development.
Xiangrong Zhang
Int. J. Comput. Intell. Appl.1
2025 MLMamba: A Mamba-Based Efficient Network for Multi-Label Remote Sensing Scene Classification
abstract
As a useful remote sensing (RS) scene interpretation technique, multi-label RS scene classification (RSSC) always attracts researchers’ attention and plays an important role in the RS community. To assign multiple semantic labels to a single RS image according to its complex contents, the existing methods focus on learning the valuable visual features and mining the latent semantic relationships from the RS images. This is a feasible and helpful solution. However, they are often associated with high computational costs due to the widespread use of Transformers. To alleviate this problem, we propose a Mamba-based efficient network based on the newly emerged state space model called MLMamba. In addition to the basic feature extractor (convolutional neural network and language model) and classifier (multiple perceptrons), MLMamba consists of two key components: a pyramid Mamba and a feature-guided semantic modeling (FGSM) Mamba. Pyramid Mamba uses multi-scale scanning to establish global relationships within and across different scales, improving MLMamba’s ability to explore RS images. Under the guidance of the obtained visual features, FGSM Mamba establishes associations between different land covers. Combining these two components can deeply mine local features, multi-scale information, and long-range dependencies from RS images and build semantic relationships between different surface covers. These superiorities guarantee that MLMamba can fully understand the complex contents within RS images and accurately determine which categories exist. Furthermore, the simple and effective structure and linear computational complexity of the state space model ensure that pyramid Mamba and FGSM Mamba will not impose too much computational burden on MLMamba. Extensive experiments counted on three benchmark multi-label RSSC data sets validate the effectiveness of MLMamba. The positive results demonstrate that MLMamba achieves state-of-the-art performance, surpassing existing methods in accuracy, model size, and computational efficiency. Our source codes are available athttps://github.com/TangXu-Group/ multilabelRSSC/tree/main/MLMamba.
Ruiqi Du, Xu Tang 0004, Jingjing Ma 0001, Xiangrong Zhang, Licheng Jiao
IEEE Trans. Circuits Syst. Video Technol.4
2025 OraL: An Observational Learning Paradigm for Unsupervised Hyperspectral Change Detection
abstract
Unsupervised hyperspectral change detection (UHCD), detecting subtle changes between bi-temporal images without manual annotations, is an essential but challenging task in the earth observation community. The current modus operandi often performs it in a feature comparison manner, which is limited by variations in imaging conditions. We observe that fully supervised paradigms using limited annotations are capable of overcoming this challenge. Based on this, we introduce a novel Observational Learning Paradigm (OraL) for UHCD by mimicking fully supervised paradigms. OraL comprises two sequential stages: Observation, which designs a spatial-temporal observation strategy (STO) that records the learning consistency of pixels under different training steps and views, to obtain reliable pseudo-labels. Reproduction, which retrains the model with these pseudo-labels and introduces a distribution-aware spectral learning strategy (DSL) to adaptively increase their learning difficulty according to spectral distributions, enhancing the robustness and generalization of the model. Extensive experiments on several public hyperspectral image datasets demonstrate its state-of-the-art performance and pluggability for previous unsupervised methods. Code will be made available.
Guanchun Wang, Xiangrong Zhang, Zelin Peng, Shunli Tian, Tianyang Zhang 0002, Xu Tang 0004, Licheng Jiao
IEEE Trans. Circuits Syst. Video Technol.2
2025 SPGFormer: Structure Perception Graph Transformer With Laplacian Position Encoding for Hyperspectral Image Classification
Jinliang An, Longlong Dai, Muzi Wang, Weidong Zhang 0007, Xiangrong Zhang
IEEE Trans. Geosci. Remote. Sens.5
2025 Semantic-Assisted Feature Integration Network for Multilabel Remote Sensing Scene Classification
abstract
With remote sensing (RS) images’ resolution increasing, a single scene label cannot adequately represent RS scenes’ contents. Therefore, multilabel RS scene classification (MLRSSC) is gradually attracting the researchers’ attention. Many methods have been proposed recently, and most use deep features or semantic connections to complete MLRSSC. However, they ignore the combination of these two aspects. In addition, the high interclass similarity and low intraclass similarity of RS images limit the robustness of these methods. In this article, we propose a semantic-assisted feature integration network (SFIN) to overcome the above limitations. It contains a dual-scale feature extractor module (DFEM), a local semantic enhance module (LSEM), a cross-scale interactive attention module (CIAM), and a classifier module (CM). DFEM utilizes the convolutional neural networks (CNNs) to extract multiscale features from RS images. LSEM extracts semantic information and establishes their relationships at different scales. CIAM enhances the feature representation by interacting with the clues across different scales. CM completes the prediction of classification (CLA) results. Integrating them into an end-to-end framework, SFIN can discover the diverse and complex land covers hidden in RS images. Furthermore, to ensure the accuracy of explored semantics and enhance the SFIN’s feature extraction ability, we design a semantic supervision (SS) loss and a semantic-based contrastive learning (SB-CL) loss. They are in charge of the correctness and discrimination of the mined semantics. Along with the typical CLA loss, SFIN can be adequately trained. Extensive experiments have been conducted on four MLRSSC datasets, and the positive results demonstrate that SFIN outperforms many existing methods in MLRSSC tasks. Our source codes are available at:https://github.com/TangXu-Group/multilabelRSSC/tree/main/SFIN.
Ruiqi Du, Xu Tang 0004, Jingjing Ma 0001, Xiangrong Zhang, Fang Liu 0034, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.4
2025 Multiscale Sparse Cross-Attention Network for Remote Sensing Scene Classification
abstract
Remote sensing (RS) scene classification (RSSC) is a prominent research topic in the RS community. Multilevel feature fusion is an important way of addressing RS scene classification, and many methods have been proposed in recent years. Although they succeed, current methods can still be improved, particularly in distinguishing the contributions of different multilevel features and fully and effectively fusing them. To address the above issues and fully exploit the potential of multilevel features for RS scene classification tasks, we propose a new model named multiscale sparse cross-attention network (MSCN). It not only focuses on the effectiveness of feature learning but also emphasizes the rationality of feature fusion. In detail, MSCN first extracts multilevel features using a pre-trained ResNet50. Also, these features are divided into high- and low-level features according to the clues they involved. Then, a multiscale sparse cross-attention (MSC) module is developed to cross-fuse the high-level feature with various low-level features, thereby effectively mining helpful information from multilevel features. In the fusion process, MSC not only explores the multiscale messages in RS scenes but also mitigates the negative impact of irrelevant information by employing sparse operations. Third, a group convolutional block attention module (CBAM) enhancer (GCE) is presented to enhance the representation of classification features. GCE detects local salient information within classification features using grouped CBAM and further enhances crucial details by readjusting the CBAM attention weights. This way, the classification features’ discrimination can be improved. We conducted extensive experiments on three public RS scene classification datasets. The exceptional experimental results indicate that our proposed MSCN achieves superior classification accuracy, surpassing many existing methods. Our source codes are available at https://github.com/TangXu-Group/Remote-Sensing-Images-Classification/tree/main/MSCN.
Jingjing Ma 0001, Xu Tang 0004, Xiangrong Zhang, Fang Liu 0034, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.4
2025 SpiralMamba: Spatial-Spectral Complementary Mamba With Spatial Spiral Scan for Hyperspectral Image Classification
abstract
Hyperspectral image (HSI) classification is crucial in the remote sensing (RS) community. In recent years, Transformers have been popular in this field due to their global information modeling capabilities. However, the quadratic complexity limits their performance under limited computational resources. Fortunately, a selective structured state space model named Mamba emerges. Like Transformer, it is good at modeling the long-distance relationships hidden in the pending data. Unlike Transformer, its complexity remains at a linear level. Therefore, a growing number of studies have been proposed to explore the usefulness of Mamba in HSI classification. Nevertheless, most of them only apply Mamba to HSIs directly but do not consider the inherent characteristics of HSIs properly. To exploit the potential of Mamba in HSI classification deeply, this paper presents a new spatial-spectral complementary Mamba with a spatial spiral scan named SpiralMamba. It mainly encloses three main components: a spatial Mamba encoder (SpaME), a spectral Mamba encoder (SpeME), and a spatial-spectral complementary fusion module (SSCFM). SpaME focuses on understanding the spatial context within HSIs. To this end, instead of the common scanning, a spatial spiral scan strategy is introduced to address the sequence transformation of non-causal HSIs. SpeME aims to comprehensively extract valuable spectral features from HSIs. To achieve this goal, besides developing a spectral bidirectional scan strategy, a multilayer convolution (MLC) is also incorporated to capture local variations within spectral tokens. SSCFM concentrates on building the complex connections between spatial and spectral features and fusing them. For this purpose, a relationship learning block (RLB) and a threshold enhancement mechanism (TEM) are developed. Positive experimental results counted on three public HSI datasets demonstrate the effectiveness of SpiralMamba. Our source codes are available at https://github.com/TangXu-Group/Hyperspectral-Images-Classification/tree/main/SpiralMamba.
Xu Tang 0004, Yuexi Yao, Jingjing Ma 0001, Xiangrong Zhang, Yuqun Yang, Bo Wang 0016, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.4
2025 Temperature-Aware Dynamic Fusion Network for Few-Shot Segmentation of Infrared Images
Bo Wang 0016, Xina Cheng, Yuan Li 0058, Xiangrong Zhang, Xu Tang 0004, Dingheng Wang, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.5
2025 S2Mamba: A Spatial-Spectral State Space Model for Hyperspectral Image Classification
abstract
The land cover analysis using hyperspectral images (HSIs) remains an open problem due to their low spatial resolution and complex spectral information. Recent studies are primarily dedicated to designing Transformer-based architectures for spatial-spectral long-range dependencies modeling, which is computationally expensive with quadratic complexity. Selective structured state space model (SSM; Mamba), which is efficient for modeling long-range dependencies with linear complexity, has recently shown promising progress. However, its potential in HSI processing that requires handling numerous spectral bands has not yet been explored. In this article, we innovatively propose S2Mamba, a spatial-spectral SSM for HSI classification, to excavate spatial-spectral contextual features, resulting in more efficient and accurate land cover analysis. In S2Mamba, two selective structured SSMs through different dimensions are designed for feature extraction, one for spatial, and the other for spectral, along with a spatial-spectral mixture gate (SMG) for optimal fusion. More specifically, S2Mamba first captures spatial contextual relations by interacting each pixel with its adjacent through a patch cross scanning (PCS) module and then explores semantic information from continuous spectral bands through a bidirectional spectral scanning (BSS) module. Considering the distinct expertise of the two attributes in homogenous and complicated texture scenes, we realize the SMG by a group of learnable matrices, allowing for the adaptive incorporation of representations learned across different dimensions. Extensive experiments conducted on HSI classification benchmarks demonstrate the superiority and prospect of S2Mamba. The code will be made available at:https://github.com/PURE-melo/S2Mamba.
Guanchun Wang, Xiangrong Zhang, Zelin Peng, Tianyang Zhang 0002, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.2
2025 Source-Free Cross-Domain Scene Classification of Remote Sensing Images via Statistics Matching and Noise Adaptation
abstract
In recent years, in order to alleviate the performance degradation problem caused by domain drift in scene classification tasks, some unsupervised domain adaptation methods have been introduced into the field of remote sensing images. Such methods require simultaneous access to both source and target domain data during training. However, the large storage and transmission costs of remote sensing images limit the further application of these methods. To address this challenge, we investigate the task of source-free cross-domain scene classification for remote sensing images. In the model adaptation process, only the target domain dataset and the trained source domain model are used. Our approach consists of two parts: a distribution alignment strategy based on source domain model statistics matching and a noise adaptation strategy. In order to fully utilize the knowledge of the pre-trained source domain model, we fix the classifier to get the feature distribution of the source domain, so that the target domain feature distribution is close to the feature distribution of the source domain. The noise adaptation layer is inserted after the classifier in order to improve the robustness of the model to noise-containing pseudo-labels, and the sample-wise noise transfer matrix is learned. Experimental results on 12 transfer tasks on the cross-scene dataset, and 2 transfer tasks on the cross-sensor dataset, to validate the effectiveness of our approach. Compared to traditional unsupervised domain adaptive methods, our method is able to achieve better performance under the condition of not accessing the source domain data.
Peng Zhu 0004, Xiangrong Zhang, Xiao Han 0012, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.2
2025 Beyond Single Pixel: Context Priors Guided Semi-Supervised Building Footprint Segmentation
abstract
Automated building footprint segmentation is crucial in remote sensing with widespread applications in various fields. Semi-supervised semantic segmentation methods are gaining traction in the remote sensing community as they significantly reduce the need for labor-intensive pixel-level annotations in training segmentation models. However, these methods typically rely on individual pixel-level supervision for unlabeled data, neglecting the contextual relationships between pixels. This limits their potential to exploit unlabeled data. To bridge this gap, this paper proposes a novel Context Priors Guided Semi-Supervised Building Footprint Segmentation method that leverages contextual relationships among numerous unlabeled pixels to build supervisory signals that extend beyond individual pixel-level guidance for learning on unlabeled data. The CPG comprises two main components: Spatial context priors-guided pseudo-label regularization and semantic context priors-guided representation learning. By integrating contextual knowledge from both pixel spatial locations and semantic representation spaces, these components capture comprehensive class semantic attributes and enable the model to learn complete shapes of building footprints. Our approach achieves state-of-the-art performance on three publicly available building footprint segmentation datasets, validating its effectiveness.
Xiaoqian Zhu, Xiangrong Zhang, Tianyang Zhang 0002, Xiao Han 0012, Licheng Jiao, Lianchao Zhang
IEEE Trans. Geosci. Remote. Sens.2
2025 Temporal-Feedback Self-Training for Semi-Supervised Object Detection in Remote Sensing Images
abstract
Although modern Remote Sensing Object Detection (RSOD) methods have achieved advanced performance, they heavily rely on a large amount of annotated data. This paper explores semi-supervised RSOD to mitigate annotation costs, leveraging recent extensive research in generic Semi-Supervised Object Detection (SSOD) based on the self-training paradigm. Current SSOD methods encounter challenges in adapting to remote sensing images due to the complexity and variability of RSIs. Two key issues remain underexplored: the noise in pseudo-labels caused by model instability and the difficulty in distinguishing similar categories. This paper introduces the Temporal-Feedback Self-Training (TST) framework, a novel approach to tackle these challenges in semi-supervised RSOD. TST consists of two components: Temporal Consistency Based Pseudo-labels Certainty Estimation (TCE) and Temporal Self-Feedback Feature Refinement (TSF). TCE addresses pseudo-label noise during training by evaluating the stability of pseudo-label classification and localization over time series to assess the quality of pseudo-labels. On the other hand, TSF enhances pseudo-label quality by dynamically identifying the models confusing categories as feedback for feature refinement. Both components facilitate the progression of the self-training-based RSOD during training. We conducted extensive experiments on two challenging public datasets, DOTA and DIOR. The results demonstrate that the proposed TST and TCE components significantly improve the baseline models performance, surpassing the state-of-the-art generic SSOD method. This suggests that our approach is more effective than generic SSOD methods in addressing the challenges posed by remote sensing images.
Xiaoqian Zhu, Xiangrong Zhang, Tianyang Zhang 0002, Xu Tang 0004, Puhua Chen, Huiyu Zhou 0001, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.2
2025 Cross-Modal Remote Sensing Image-Text Retrieval via Context and Uncertainty-Aware Prompt
abstract
The cross-modal remote sensing image-text retrieval (CMRSITR) is a lively research topic in the remote sensing (RS) community. Benefiting from the large pretrained image-text models, many successful CMRSITR methods have been proposed in recent years. Although their performance is attractive, there are still some challenges. First, fine-tuning large pretrained models requires a significant amount of computational resources. Second, most large models are pretrained by natural images, which reduces their effectiveness in processing RS images. To tackle these challenges, we propose a new CMRSITR network named context and uncertainty-aware prompt (CUP). First, prompt tuning theory is introduced into CUP to eliminate the burden of optimization resources. By training the prompt tokens rather than all parameters, the large model's knowledge can be transferred to CMRSITR tasks with small trainable parameters. Second, considering the differences between natural-image-based prior clues and RS images, apart from adopting the free-prompt tokens, we develop a prompt generation module (PGM) to produce the RS-oriented prompt tokens. The specific prompt tokens are rich in object-level messages of RS images, which help CUP narrow the gaps between natural large models and RS images. Third, we further design an uncertainty estimation module (UEM) to whittle down the uncertainties caused by the model and data. This way, can not only the semantic misalignment and intraclass diversity imbalance problems be mitigated but also the RS clues can be deeply explored. Competitive experimental results counted on three public benchmark datasets demonstrate that our CUP can achieve competitive performance in the CMRSITR task compared with many existing methods. Our source codes are available at: https://github.com/TangXu-Group/Cross-modal-remote-sensing-image-and-text-retrieval-models/tree/main/CUP.
Yijing Wang 0004, Xu Tang 0004, Jingjing Ma 0001, Xiangrong Zhang, Fang Liu 0034, Licheng Jiao
IEEE Trans. Neural Networks Learn. Syst.4
2025 Negative Deterministic Information-Based Multiple Instance Learning for Weakly Supervised Object Detection and Segmentation
abstract
Weakly supervised object detection (WSOD) and semantic segmentation with image-level annotations have attracted extensive attention due to their high label efficiency. Multiple instance learning (MIL) offers a feasible solution for the two tasks by treating each image as a bag with a series of instances (object regions or pixels) and identifying foreground instances that contribute to bag classification. However, conventional MIL paradigms often suffer from issues, e.g., discriminative instance domination and missing instances. In this article, we observe that negative instances usually contain valuable deterministic information, which is the key to solving the two issues. Motivated by this, we propose a novel MIL paradigm based on negative deterministic information (NDI), termed NDI-MIL, which is based on two core designs with a progressive relation: NDI collection and negative contrastive learning (NCL). In NDI collection, we identify and distill NDI from negative instances online by a dynamic feature bank. The collected NDI is then utilized in a NCL mechanism to locate and punish those discriminative regions, by which the discriminative instance domination and missing instances issues are effectively addressed, leading to improved object- and pixel-level localization accuracy and completeness. In addition, we design an NDI-guided instance selection (NGIS) strategy to further enhance the systematic performance. Experimental results on several public benchmarks, including PASCAL VOC 2007, PASCAL VOC 2012, and MS COCO, show that our method achieves satisfactory performance. The code is available at: https://github.com/GC-WSL/NDI.
Guanchun Wang, Xiangrong Zhang, Zelin Peng, Tianyang Zhang 0002, Xu Tang 0004, Huiyu Zhou 0001, Licheng Jiao
IEEE Trans. Neural Networks Learn. Syst.2
2024 Heterogeneous Open-Set Cross-Domain Manifold Embedding Aligned for HSI-MSI Collaborative Classification
abstract
Hyperspectral images (HSI) have higher spectral resolution than multispectral images (MSI), but due to limitations of imaging equipment, their width is narrower than MSI. When using partially overlapping HSI-MSI to improve the classification capabilities of MSI, there may be unknown classes that do not exist in HSI-MSI overlapping regions. To solve this problem, this paper proposes a heterogeneous open-set cross-domain manifold embedding aligned method for HSI-MSI collaborative classification. The method designs manifold embedding to align HSI-MSI features to map into subspaces, and gradually selects target domain samples for pseudo-labeling through the designed strategy while rejecting unknown class samples. The feature alignment and pseudo-labeled sample selection are continuously iterated to promote each other, reducing the intra-class distance while pushing the rejected target data away from known classes. The experimental results verify the superiority of our method.
Bin Guo 0015, Xiangrong Zhang, Tianzhu Liu, Yanfeng Gu
IGARSS2
2024 A Cross-Modal Semantic Mapping Enhancement Model for Remote Sensing Visual Question Answering
abstract
Remote sensing visual question answering (RSVQA) aims to answer the questions based on the content in remote sensing (RS) images. Due to the complexity of RS images, it is challenging to focus on regions relevant to the questions in the RS images. To this end, we propose a channel-selective multi-scale cross-attention (CSCa) model for RSVQA tasks. Specifically, we design a text-driven multi-scale feature extractor to extract question-related features in RS images. To obtain the cross-attention map in this extractor, we design a novel channel selection mechanism to capture more accurate question-related regions in RS images and develop a channel-wise contrastive learning task to align the semantics between image and text features. We set up experiments on RSVQA-LR and RSIVQA datasets. Experiment results show that our CSCa achieves excellent performance.
Dabiao Huang, Xu Tang 0004, Xiangrong Zhang, Fang Liu 0034, Jingjing Ma 0001, Licheng Jiao
IGARSS3
2024 Multi-Scale Sparse Transformer for Remote Sensing Scene Classification
abstract
Vision Transformer (ViT) has achieved great success in the field of computer vision since it was proposed, and there have been many works applying ViT based models to remote sensing scene classification (RSSC) tasks. The proposal of Pyramid Vision Transformer (PVT) greatly reduces the calculation amount of the ViT while maintaining accuracy. But PVT did not utilize multi-scale information in remote sensing (RS) scenes, which is crucial for RSSC. This paper proposes a multi-scale sparse transformer (MST) based on PVT. MST enables the network to learn multi-scale representations of RS scenes through spatial reduction implementations at different scales. In addition, we employ sparse operations to adaptively guide the model’s attention towards semantically relevant regions during self-attention computation, thereby reducing interference from semantically irrelevant areas. Experiments conducted on the UCM and AID datasets demonstrate the outstanding performance of the proposed MST.
Xu Tang 0004, Zhixi Feng, Yue Ma 0008, Jingjing Ma 0001, Xiangrong Zhang, Licheng Jiao
IGARSS6
2024 An Automated Workflow for Pixel-Level BRDF Extraction Using UAV-Based Multispectral Images
abstract
Bidirectional Reflectance Distribution Function (BRDF) plays a vital role in quantitative remote sensing. Recently, UAV has gradually emerged as the leading choice for BRDF acquirement. However, challenges remain in extracting usable BRDF data from UAV multispectral images (MSIs), including labor-intensive and limited accuracy. To tackle these challenges, an automated workflow for pixel-level BRDF extraction is proposed in this paper, comprising two main stages: 3D reconstruction and back projection. Pixel-level accuracy is achieved through 3D reconstruction, enhanced by the integration of commercial software for streamlined automation. Back projection is critical for precise location. Experiments were carried out for both selected ROIs and all pixel areas, validating the efficacy of our proposed approach.
Zhenqiang Qin, Xian Li 0001, Yanfeng Gu, Xiangrong Zhang
IGARSS4
2024 Pseudo-Viewpoint Regularized 3D Gaussian Splatting For Remote Sensing Few-Shot Novel View Synthesis
abstract
In remote sensing (RS), Few-Shot Novel View Synthesis (FS-NVS) focuses on creating images of unobserved viewpoints using limited training images. Recently, 3D Gaussian Splatting (3DGS) has drawn scholars’ attention by its increasing rendering speeds and providing an explicit neural representation for 3D scenes. However, 3DGS tends to overfit limited training data. To tackle this challenge, we propose a Pseudoview Regularized 3DGS (PR3DGS) FSNVS method for RS scenarios. Our PR3DGS method introduces a pseudo-views regularization module to discriminate synthetic RS images generated from training- or pseudo-viewpoints. Therefore, our PR3DGS method can effectively mitigate overfitting in seen views and enhance the model’s capability to generate more realistic RS images from novel viewpoints. Besides, the excellent experimental results on the LEVIR-NVS dataset demonstrate the effectiveness of our method in RS FSNVS.
Yijing Wang 0004, Xu Tang 0004, Jingjing Ma 0001, Xiangrong Zhang, Fang Liu 0001, Licheng Jiao
IGARSS4
2024 Application of Landweber with Optimization for Small Footprint Waveform Lidar Decomposition
abstract
Small-footprint waveform LiDAR requires waveform decomposition for accurate target structure characterization. To improve the ability for identifying close targets, this paper first introduces the Landweber (LW) deconvolution method to decompose the small-footprint LiDAR waveforms. Our study emphasizes the advantages of the deconvolution methods in capturing more targets of waveforms. Generally, the LW approach introduced with optimization excels in detecting more targets after false target removal. Experiments were conducted on datasets collected under various conditions using small-footprint waveform LiDAR system. The findings highlight an average target distance error of 0.083m, showcasing superior performance compared to direct decomposition methods. When compared with the GOLD and RL methods, the decomposition accuracy is nearly indistinguishable, but the success rates are higher. Our research establishes the LW method as a viable waveform decomposition method, contributing to the diversity of choices for waveform data processing.
Yanfeng Gu, Xian Li 0001, Xiangrong Zhang
IGARSS4
2024 A Dual-Branch Network for End-to-End Point-Supervised Object Detection on Remote Sensing Images
abstract
Learning object detectors for remote sensing images commonly requires for a huge number of annotated boundary boxes, which are not available without enormous manual efforts in annotating. Alternatively, points can indicate the existence of the objects of interests with reduced labeling cost. Existing Point-supervised object detection (PSOD) methods predominantly employ a two-stage training strategy, which involves propagating point annotations to pseudo boxes at the first stage then training an object detector with these pseudo boxes in a fully supervised manner. However, such paradigm substantially impedes the end-to-end flow of training gradients. In this work, we propose a novel dual-branch network (DBNet) for end-to-end weakly supervised object detection on remote sensing images. Firstly, a pseudo box generation network is attached to the object detector as a sibling branch, which produces semantic response maps for the objects of interest then extracts pseudo boxes by examining their spatial connectivity. Then, instead of training this pseudo box generation network separately, we jointly adjust the pseudo box generation network and the detection network through a multi-task loss. Experimental results on the DOTA-v1.0 dataset demonstrate the effectiveness of our proposed method, achieving an average precision (mAP50) of 32.3%.
Jie Feng 0003, Junpeng Zhang 0002, Ronghua Shang, Xiangrong Zhang, Licheng Jiao
IGARSS5
2024 Time-Guided Network for Remote Sensing Change Detection
abstract
In recent years, as a very important part of remote sensing interpretation, change detection (CD) has developed rapidly with deep learning and remote sensing interpretation. However, most of the existing CD methods focus on how to extract the spatial features of the bi-temporal image pairs, ignore the importance of the temporal features. To solve this problem, a new time-guided network (TG-Net) using temporal information to guide feature extraction is proposed in this paper. In TG-Net, we use a newly proposed time-guided feature fusion (TGF2) block that uses temporal information to guide spatial feature fusion to extract temporal and spatial information comprehensively. We conducted experiments on two publicly available remote sensing datasets, LEVIR-CD and WHU, and compared them with three common CD methods.
Xu Tang 0004, Xiangrong Zhang, Jingjing Ma 0001, Fang Liu 0001, Licheng Jiao
IGARSS3
2024 Center Mask Self-Attention Network for Hyperspectral Image Classification
abstract
Benefiting from the thousands of continuous band information in hyperspectral images (HSIs), the task of HSI classification has become an indispensable part of the field of remote sensing. With the development of deep learning, deep learning techniques such as convolutional neural networks have been widely introduced into HSI classification research. However, most of these methods do not fully consider the potential relationship between the central pixel and surrounding neighborhoods. Therefore, we introduce a novel center mask self-attention network (CMSAN) to enable the model to effectively capture the association between the central pixel and its neighbors for better feature extraction. We conduct experiments on two publicly available HSI datasets. The positive results on both datasets fully demonstrate the effectiveness of our proposed method.
Yizhou Zou, Xu Tang 0004, Yue Ma 0008, Jingjing Ma 0001, Xiangrong Zhang, Licheng Jiao
IGARSS6
2024 Multi-agent deep reinforcement learning for hyperspectral band selection with hybrid teacher guide
Jie Feng 0003, Qiyang Gao, Ronghua Shang, Xianghai Cao, Gaiqin Bai, Xiangrong Zhang, Licheng Jiao
Knowl. Based Syst.6
2024 Learning consensus-aware semantic knowledge for remote sensing image captioning
Yunpeng Li 0010, Xiangrong Zhang, Xina Cheng, Xu Tang 0004, Licheng Jiao
Pattern Recognit.2
2024 A Graph Association Motion-Aware Tracker for Tiny Object in Satellite Videos
abstract
Satellite video object tracking involves tracking a specified tiny object within a wide scene. The insufficient appearance features of these tiny objects pose significant challenges to appearance-based object trackers, particularly in situations involving occlusion, target blur, and similar interferences. In this paper, a novel Graph Association MOtion-aware tracker (GAMO) is proposed for tiny object in satellite videos, which integrates motion and spatial relationship information. First, a Gaussian motion estimator is proposed that decouples motion into velocity and direction, rather than using traditional x-y movement modeling. This estimator predicts the object’s position and estimates motion uncertainty with a directional motion probability map. Furthermore, the estimated motion serves as a prior to guide the proposal sampling. A probabilistic proposal sampling module is designed that samples candidate bounding boxes according to the directional motion probability map, focusing on the region where the target is most likely to appear. Additionally, we implement a graph association module to model and propagate the spatial relationships between the target and neighboring objects over time. This relationship information assists the appearance features in distinguishing the target from similar interferences. Experiments on the Skysat-1, SV248S, and VISO datasets demonstrate the superiority of the proposed tracker. GAMO leverages motion and surrounding information, resulting in significant improvements with minimal computational overhead. The code and results will be publicly available inhttps://github.com/Midkey/GAMO.
Zhongjian Huang, Licheng Jiao, Xu Liu 0006, Fang Liu 0001, Xiangrong Zhang, Lingling Li 0002, Puhua Chen
IEEE Trans. Circuits Syst. Video Technol.6
2024 Efficient LWPooling: Rethinking the Wavelet Pooling for Scene Parsing
abstract
Existing wavelet pooling methods discard the high-frequency sub-bands, which can improve the noise-robustness of convolutional neural networks (CNNs) but lose the essential detailed features. Besides, most of them depend on different wavelets, which is not adaptive. In this paper, a novel efficient lifting-based wavelet pooling (LWPooling) is proposed to alleviate the problems above. Firstly, wavelet pooling is rethought based on the equivalence of 2D discrete wavelet transform (DWT) and standard average pooling (SAP), which suggests the lack of detailed information on traditional wavelet pooling. Secondly, the efficient LWPooling module is proposed to adaptively capture and preserve the critical high-frequency features via lifting-based wavelets. It can constrain the features linear independence, which efficiently makes important features salient. Thirdly, the lifting-based wavelet collaborative network (LWCNet) is constructed for classification and segmentation tasks based on the efficient LWPooling module. Experiments are validated on Cifar10, Cifar100, and ADE20K datasets. It suggests that the efficient LWPooling can enhance CNN’s representation and achieve a particular performance advantage compared to average, maximum, and original wavelet pooling. Besides, the proposed LWCNet shows the potential for scene parsing. The code implementation will be available at https://github.com/yutinyang/LWCNet.
Yuting Yang 0008, Licheng Jiao, Xu Liu 0006, Lingling Li 0002, Fang Liu 0001, Shuyuan Yang 0001, Xiangrong Zhang
IEEE Trans. Circuits Syst. Video Technol.7
2024 Class-Aligned and Class-Balancing Generative Domain Adaptation for Hyperspectral Image Classification
abstract
The task of hyperspectral image (HSI) classification is fundamental and crucial in HSI processing. Currently, domain adaptive methods have become a research hotspot in HSI classification. However, most domain adaptive methods ignore the class alignment in different domains. Additionally, HSIs have the characteristics of category imbalance and complex spatial-spectral distribution, which restricts the adaptation performance in HSIs. To address these problems, a class-aligned and class-balancing generative domain adaptation (CCGDA) method is proposed for HSI classification. The architecture of CCGDA is designed by using the classifier, domain discriminator, sampler and two weight-sharing generators. In the classifier, split-level capsule network is constructed by extracting rich spatial information of shallow layer and spectral features of deep layer with equivariant characteristic. Then, the classifier provides the pseudo label of samples in the target domain. To prevent the generators from mode collapse caused by category imbalance, the sampler is designed. It samples and re-samples the samples of the target domain in an adaptive proportion according to the statistical calculation through confidence and distribution of pseudo labels. Finally, a novel class-aligned domain adversarial loss is defined to jointly optimize the generators and discriminator. It incorporates the class shift adjusting and adaptive sampling for the samples of the target domain to better adapt the discriminant boundary of the classifier to the target domain. Experiments on benchmark HSI datasets verify the superiority of the proposed method for domain adaptive classification.
Jie Feng 0003, Ziyu Zhou 0009, Ronghua Shang, Jinjian Wu, Xiangrong Zhang, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.6
2024 Few-Shot Multispectral-Hyperspectral Image Collaborative Classification With Feature Distribution Enhancement and Subdomain Alignment
abstract
With the development of observation technology, multispectral (MS) images of large scenes are easy to obtain, but the low spectral resolution limits their classification ability. Moreover, the collection of training samples is difficult and time-consuming, and limited labeled samples are a challenge for the precise classification of large-scene MS images. This article attempts to use hyperspectral (HS) images with limited labels to help classify MS images of large scenes, so as to achieve better classification results. To solve this problem, a few-shot MS-HS image collaborative classification method combining feature distribution enhancement (FDE) and subdomain alignment is proposed. Specifically, a residual 3-D convolution network embedded with a 3-D FDE module is designed to improve the diversity of the feature distribution extracted by the network and increase the generalization ability of the model under the few-shot condition. Furthermore, the local domain alignment between the source and target domains is achieved by subdomain alignment, which better aligns the categories in the source domain and the target domain, and achieves the distribution alignment of the subdomains. In addition, the feature bias adjustment (FBA) module is introduced in the test phase to correct the bias of the MS image feature representation, and to alleviate the cross-domain problem to some extent. The few-shot learning (FSL) is applied in the source and target domains to learn better feature mapping. The results of comparative experiments on three datasets show that the proposed method is superior to the most advanced method in the case of limited labeled samples.
Bin Guo 0015, Tianzhu Liu, Xiangrong Zhang, Yanfeng Gu
IEEE Trans. Geosci. Remote. Sens.3
2024 Few-Shot Open-Set Collaborative Classification of Multispectral and Hyperspectral Images With Adaptive Joint Similarity Metric
abstract
Hyperspectral images (HSIs) have higher spectral resolution than multispectral (MS) images, but they have a narrower swath than MS images. The limited spectral resolution of MS images constrains their classification capabilities, and annotating remote sensing data is time-consuming and laborious. In addition, large-scale MS images may contain unknown classes not present in the training data. This article attempts to use partially overlapping HS images with limited labels to assist in the classification of large-scene MS images. It can correctly distinguish known classes and simultaneously identify unknown classes, thereby achieving better classification results for MS images. To address this challenge, a few-shot open-set HS–MS image collaborative classification method is proposed. Specifically, a spectral–spatial feature interactive enhancement (SSFIE) module is designed for richer feature extraction and enhanced classification capabilities in the feature extraction stage. In the few-shot learning (FSL) stage, an adaptive joint similarity metric criterion is proposed to improve feature mapping between the source and target domains. Discriminative joint probability adaptation (DJPA) is used for domain adaptation and to enhance feature discriminability, while batch nuclear-norm maximization (BNM) is employed to increase the feature diversity. In the testing phase, the open-set classification module is designed to correctly classify samples of known classes while simultaneously distinguishing unknown classes. The experimental results on four cross-domain HS–MS data pairs demonstrate that our proposed method outperforms state-of-the-art methods.
Bin Guo 0015, Xiangrong Zhang, Tianzhu Liu, Yanfeng Gu
IEEE Trans. Geosci. Remote. Sens.2
2024 Change-Guided Similarity Pyramid Network for Semantic Change Detection
abstract
Semantic change detection (SCD) based on remote sensing images can provide an effective solution for large-scale land use monitoring. The existing change detection (CD) networks face limitations due to their limited receptive fields, which make it difficult to provide a comprehensive and consistent response to changes in similar large objects. Moreover, these limitations often cause the network to ignore weak changes localized in the images. To address these limitations, we propose change-guided similarity pyramid network (CG-SPNet), a decoupled multitask architecture that integrates three key components, including the multiscale positional similarity pyramid (MSSP), the symmetry-enhanced fusion CD unit (SFCD), and the change-guided feature interaction module (CGM). MSSP module captures positional correlation of semantic features at multiple scales by introducing an attention pyramid during pooling downsampling. SFCD utilizes a convolutional attention mechanism to enhance local features and highlight weak changes in dual images, and the symmetric fusion (SF) approach reduces the network’s sensitivity to timing. CGM is used to address the challenge of effective feature interaction, which combines cosine similarity loss to leverages cross-attention to compute the positional relevance of change features to embed a priori information for semantic features. Through rigorous experiments and analyses, we have successfully validated the essential components of CG-SPNet. It achieves the state-of-the-art performance on the SECOND dataset as well as our self-built CZWZ dataset.
Lei He 0006, Mingheng Zhang, Yuxia Li, Shiyu Luo, Shuguang Li 0002, Xiangrong Zhang
IEEE Trans. Geosci. Remote. Sens.7
2024 Intertemporal Interaction and Symmetric Difference Learning for Remote Sensing Image Change Captioning
abstract
Remote sensing image change captioning (RSICC) is more challenging than remote sensing change detection task, which requires extracting occurred changes in similar remote sensing image (RSI) pairs while generating change caption. However, few works have been investigated on RSICC, the main challenges come from how to learn abundant change clues and face the modality gap. To handle these problems, we rethink this task from the perspective of obtaining and aligning symmetrical change features for temporal RSIs. In this work, the proposed intertemporal interaction and symmetric difference learning network are cascaded through several multitemporal integration units to model differences from coarse to fine representations. Specifically, we design a cross-temporal attention (CTA) mechanism to probe direct interaction between bi-temporal RSIs for motivating information coupling between intralevel representations and suppressing irrelevant interferences. To learn robust change features, a symmetric difference transformer (SDT) module is devised to guarantee temporal symmetry between the “before-to-after” and “after-to-before” change representations. Besides, the bi-directional triplet ranking loss is adopted to guide the network to learn strongly discriminative and temporal-symmetric change representation. Extensive experimental results on Dubai-CC and LEVIR-CC datasets demonstrate that our framework with the proposed components can achieve excellent performance and surpass recent state-of-the-art methods.https://github.com/romanticLYP/TISDNet
Yunpeng Li 0010, Xiangrong Zhang, Xina Cheng, Puhua Chen, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.2
2024 EATDer: Edge-Assisted Adaptive Transformer Detector for Remote Sensing Change Detection
abstract
Change detection (CD) is one of the important research topics in remote sensing (RS) image processing. Recently, convolutional neural networks (CNNs) have dominated the RSCD community. Many successful CNN-based models have been proposed, and they achieved cracking performance. Nevertheless, influenced by the limited receptive field, the CNN-based models are not good at capturing long-distance context dependencies within RS images, negatively impacting their performance. With the appearance of the visual transformer, the above problems have been mitigated. However, the high time costs of the transformer-based models limit their applicability. In addition, previous CD networks (whether CNN-based or transform-based) do not pay attention to the edges of changed areas, reducing the quality of change maps. To overcome the shortcomings discussed above, we propose a new CD method named edge-assisted adaptive transformer detector (EATDer). EATDer consists of a Siamese encoder and an edge-aware decoder. Each branch in the Siamese encoder encloses three self-adaption vision transformer (SAVT) blocks, which aim to capture the local and global information within RS images. Also, two branches are connected by full-range fusion modules (FRFMs), which focus on mining the temporal clues among bi-temporal RS images and pointing out the changed/unchanged messages. The edge-aware decoder first integrates the multiscale features obtained by the encoder using a restoring block. Then, it enhances the combined features by a refining block. Finally, based on the refined features, both the change and edge detection results can be produced. Along with a joint loss function, we can get high-quality change maps in which the changed areas are correct and have clear and smooth edges. The usefulness of our EATDer is validated by extensive experiments conducted on three popular RSCD datasets. Our source codes are available athttps://github.com/TangXu-Group/Remote-Sensing-Image-Change-Detection/tree/main/EATDer
Jingjing Ma 0001, JunYi Duan, Xu Tang 0004, Xiangrong Zhang, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.4
2024 Spatial Pooling Transformer Network and Noise-Tolerant Learning for Noisy Hyperspectral Image Classification
abstract
Hyperspectral image (HSI) classification is a hot topic in remote sensing. A large number of studies have been proposed and achieved excellent performance. Most of them rely on accurate annotations. However, this requirement cannot always be met. Due to the complex contents within HSIs and the uncontrollable external interference factors, incorrect labels are inevitable. Thus, the study of noisy HSI classification is boomed. Some attempts have been made, and their central ideas are to filter the noisy samples from the training set. Although feasible, this would result in information loss, i.e., the contents covered by the removed samples are ignored. Besides, the characteristics of HSIs are not fully considered in many models. To overcome the above limitations, we develop a spatial pooling transformer network (SPTNet) and a noise-tolerant learning algorithm in this paper. SPTNet first uses a spectral feature extraction (SFE) module to capture the rich spectral information from HSI patches. Then, three spatial pooling transformers (SPTs) are constructed and stacked to explore the spatial knowledge and depress confusing clues caused by the HSI patch division. Finally, a standard transformer encoder is used to enhance the obtained spectral-spatial features for the downstream classification. To use SPTNet to handle noisy HSI classification, the noise-tolerant learning algorithm is designed. It encloses two parts, i.e., a data partition scheme and a label-independent similarity regularization. The data partition scheme divides the training data into clean and noisy sets. Then, the clean samples are used to train SPTNet with the classification loss function. At the same time, similarity regularization helps SPTNet to comprehensively understand HSIs by analyzing the resemblance between clean and noisy samples. Integrating two parts into a co-training framework, SPTNets can be trained under a noisy scenario. Four popular HSI datasets are selected to testify to our methods. The positive results demonstrate that the combination of SPTNet and the noise-tolerant learning algorithm is helpful to the noisy HSI classification. Our source codes are available at https://github.com/TangXu-Group/Hyperspectral-Images-Classification/tree/main/SPTNet-NTLA.
Jingjing Ma 0001, Yizhou Zou, Xu Tang 0004, Xiangrong Zhang, Fang Liu 0034, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.4
2024 Prior-Experience-Based Vision-Language Model for Remote Sensing Image-Text Retrieval
abstract
Remote sensing (RS) image-text retrieval (RSITR) aims to retrieve relevant texts (RS images) based on the content of a given RS image (text). Existing methods are used to employing the convolutional neural network (CNN) and recurrent neural network (RNN) as encoders to learn visual and textual features for retrieval. Although feasible, the global information hidden in different data does not receive the attention it deserves. To mitigate this problem, transformers have been introduced. Nevertheless, the complexity of RS images present challenges in directly introducing Transformer-based architectures to multimodal learning in RS scenes, particularly in visual feature extraction and cross-modal interaction. In addition, the textual captions are always simpler than the complex RS images, leading to a semantic description appearing in different images. This typical false-negative (FN) sample problem increases the difficulty of RSITR tasks. To address the above limitations, we propose a new RSITR model named prior-experience-based RS vision-language (PERSVL). First, the specific visual and text encoders are used to extract features from RS images and texts. Also, a high-level feature complement (HFC) module is developed based on the self-attention mechanism (SAM) for the visual encoder to explore the complex contents from RS images fully. Second, a dual-branch multimodal fusion encoder (DBMFE) is designed to complete the cross-modal learning. It comprises a dual-branch multimodal interaction (DBMI) module and a branch fusion module. DBMI is designed to fully explore the relationships between different modalities, enriching visual and textual features. The branch fusion module integrates the cross-modal features and utilizes a classification head to generate matching scores for retrieval. Finally, a learning from prior experiences (LPEs) module is designed to reduce the influence of FN samples by analyzing the historical data produced in the model training process. Experiments are conducted on three popular datasets, and the positive results show that our PERSVL model achieves superior performance compared with previous methods. By integrating the advantages of natural language and RS images, our PERSVL can be applied in various applications, such as environmental monitoring, disaster evaluation, and urban planning. Our source codes are available at:https://github.com/TangXu-Group/Cross-modal-remote-sensing-image-and-text-retrieval-models/tree/main/PERSVL.
Xu Tang 0004, Dabiao Huang, Jingjing Ma 0001, Xiangrong Zhang, Fang Liu 0034, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.4
2024 Multiple Information Collaborative Fusion Network for Joint Classification of Hyperspectral and LiDAR Data
abstract
Joint classification of hyperspectral image (HSI) and light detection and ranging (LiDAR) can simultaneously utilize rich spectral information and elevation information and has become a hot research topic in remote sensing (RS). Although many works have been proposed for this task, their performance cannot reach what we expected due to inadequate cross-modal feature learning and simple feature fusion. This article proposes a multiple information collaborative fusion network (MICF-Net) to overcome those limitations, which aims to leverage the essentially consistent spatial relationships and high-level semantic information in multimodal data to guide the extraction of multimodal fusion features. Specifically, MICF-Net first uses a simple two-branch convolutional neural network (CNN) for preliminary feature extraction. Then, a dual-branch cross-modal attention fusion transformer (CMAFT) is developed to mine global contextual content. By fusing the attention maps of two modalities and limiting their similarity, CMAFT can retain modality-specific information while achieving information interaction based on spatial relationships. Next, an adaptive mask modulation (AMM) module is designed to dynamically balance the learning rate of each modality to ensure the effectiveness of the features of all modalities. Finally, to mine the complementary information of HSI and LiDAR data, a semantic-guided feature fusion (SGFF) module is introduced. It achieves mutual guided learning by exchanging semantic information between two modalities. Positive experimental results counted on three popular HSI and LiDAR datasets demonstrate the effectiveness of the proposed MICF-Net. Our source codes are available athttps://github.com/TangXu-Group/Hyperspectral-Images-Classification/tree/main/MICF-Net.
Xu Tang 0004, Yizhou Zou, Jingjing Ma 0001, Xiangrong Zhang, Fang Liu 0034, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.4
2024 A Multitask Framework for Hyperspectral Change Detection and Band Reweighting With Unbalanced Contrastive Learning
abstract
Multitask learning has been widely applied in visual learning to significantly enhance the performance. The combination of hyperspectral change detection (HCD) and band reweighting can achieve discriminative feature enhancement for improving detection performance. However, existing multitask models for these two tasks are unidirectional, with band reweighting unable to learn from task guidance. To address this challenge, a multitask HCD (MHCD) framework with differential band reweighting and unbalanced contrastive learning is proposed. MHCD consists of a differential band reweighting network (DBRN) and a Siamese detection network. DBRN extracts discriminative information for HCD by analyzing the differential spatial-spectral information across time states, whose optimization is under the guidance of HCD. Furthermore, a multitemporal interaction module and multidomain fusion module are inserted into the Siamese detection network. They hierarchically connect cross-temporal features and fuse features from spatial, spectral, and temporal domains, providing complementary clues in these different domains. Considering the sample imbalance and enormous variation within a class in binary HCD, an unbalanced contrastive learning method based on multiple prototypes (UCLM) tailored has been considered. It estimates multiple prototypes to flexibly adjust the contribution of different classes of samples to the loss. The proposed method has been validated using three public benchmark datasets, demonstrating improvements in multiple metrics for change detection. The code of our paper is available at:https://github.com/jiefeng0109/MHCD.
Xiande Wu, Paolo Gamba, Jie Feng 0003, Ronghua Shang, Xiangrong Zhang, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.5
2024 ECPS: Cross Pseudo Supervision Based on Ensemble Learning for Semi-Supervised Remote Sensing Change Detection
abstract
Semi-supervised learning aims to exploit the potential of unlabeled data to enhance model performance, which makes it suitable for addressing the challenge of limited labeled data. As a popular technology, pseudo-label is widely applied in many semi-supervised remote sensing (RS) change detection methods. However, when facing limited labeled data, abundant low-quality pseudo-labels from a poorly-performing model hinder the effective enhancement of model performance. To address this issue, we propose a novel semi-supervised strategy, named ensemble cross pseudo supervision (ECPS). The utilization of ensemble learning to merge outputs from several change detection models enhances pseudo-label quality, leading to more accurate change information and a significant boost in model performance, even with limited labeled data. In this method, adopting crosswise supervision ensures that no additional inference costs caused by ensemble learning are consumed. This provides both high efficiency and effectiveness for identifying land-cover changes. On the other hand, a simple yet effective ensemble strategy is proposed, which allows to manually adjust the model’s tendency towards higher precision or recall for satisfying practical requirements. We conduct extensive experiments on four public RS change detection datasets, and the promising results demonstrate the superiority of the proposed method across various numbers of labeled samples. Our source codes are available at https://github.com/TangXu-Group/ECPS.
Yuqun Yang, Xu Tang 0004, Jingjing Ma 0001, Xiangrong Zhang, Shiji Pei, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.4
2024 FDLdet: A Change Detector Based on Forward Dictionary Learning for Remote Sensing Images
abstract
As an important topic in the remote sensing (RS) image processing community, change detection has attracted much attention from researchers, which aims to distinguish land-cover changes in a geographic position. This is a challenging task because the visual representations of land cover captured from RS images at different periods would vary widely and considerably, resulting in significant differences in feature representations. To alleviate this problem, many existing deep-based methods employ the parameter-shared strategy to map RS images into a common feature space for detecting the changes. Although they are feasible, the simple and single visual information learned by deep models is still not sophisticated enough for satisfactory results. To address this problem, we propose a forward dictionary learning (DL) model named forward DL detector (FDLdet) in this article. Besides the common visual features, our FDLdet takes into account the essential information, e.g., element composition and land-cover category, for change detection. FDLdet consists of a feature extractor, a coefficient generator, and a deep dictionary. Specifically, first, the feature extractor is used to extract shared deep features from RS images. Second, the coefficient generator transforms these deep features into word coefficients. Third, words within the deep dictionary are combined by word coefficients to generate the dictionary features with essential information. Finally, the dictionary features are used instead of deep features to detect land-cover changes. Extensive experiments are conducted on two public large-scale datasets, i.e., season-varying change detection (SVCD), Sun Yat-sen University change detection (SYSU-CD), and LEVIR change detection (LEVIR-CD). Experimental results demonstrate the effectiveness of the proposed FDLdet. Our source codes are available athttps://github.com/TangXu-Group/FDLdet.
Yuqun Yang, Xu Tang 0004, Xiangrong Zhang, Jingjing Ma 0001, Yiu-Ming Cheung, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.3
2024 Hierarchical Knowledge Graph for Multilabel Classification of Remote Sensing Images
abstract
Multilabel classification in remote sensing (RS) images aims to correctly predict multiple object labels in an RS image with the primary challenge of mining correlations among multiple labels. In this context, we argue that a scene can be treated as a high-level depiction of the interactions among multiple interconnected objects within the image. However, hierarchical relationships between the scene and local objects are often neglected in other state-of-the-art approaches. In this article, we consider multilabel classification as a global-to-local prediction process, whereas the scene of an image is first identified, followed by recognition of local objects in the image. To achieve this, we propose a novel hierarchical knowledge graph (HKG)-based framework for multilabel classification in RS images (ML-HKG). Specifically, we first construct a hierarchical KG to depict label correlations between scenes and objects and represent the hierarchical knowledge as interrelated scene- and object-level label embeddings. Subsequently, we generate a scene-aware enhanced feature map by recognizing scene categories in an image under the guidance of scene-level knowledge embeddings. Afterward, object-level embeddings are used to derive category-specific visual representations for final multilabel prediction. Extensive experiments on the UCM and AID datasets demonstrate the effectiveness of our framework.
Xiangrong Zhang, Xina Cheng, Xu Tang 0004, Huiyu Zhou 0001, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.1
2024 ESDINet: Efficient Shallow-Deep Interaction Network for Semantic Segmentation of High-Resolution Aerial Images
abstract
Semantic segmentation of high-resolution remote sensing images is essential in many fields. Nevertheless, in practical applications, constrained by limited computational resources and complex network structures, many advanced models on semantic segmentation often fail to show efficient performance, prompting research on lightweight models. For lightweight semantic segmentation models, the two-branch architecture has been shown to work well in speed and performance. However, such two-branch architectures usually do not utilize enough information for shallow structures to efficiently provide richer multiscale information for the two branches. The lightweight modules it uses are difficult to extract the global context information of the features effectively. Compared with the current advanced semantic segmentation models, lightweight models still have some differences in performance. In order to solve these problems, we propose a new lightweight dual-branch architecture efficient shallow-deep interaction network (ESDINet), which can quickly extract low-level spatial and high-level semantic information of images through the detail branch and semantic branch. Specifically, we have constructed an efficient double-branch structure with shallow and deep different interactions to achieve multiscale information interaction. At the same time, we optimize the semantic branch and propose a new linear attention block to effectively improve the global perception of the semantic branch. We performed extensive experiments and the results show that our model achieves a good balance between segmentation accuracy and inference speed. In particular, ESDINet achieves 82.03% mean intersection over union (mIoU) on the Vaihingen test set, while the proposed model achieves an inference speed of 116 frames/s (FPS) for$512\times512$inputs on a single NVIDIA GTX 2080Ti GPU.
Xiangrong Zhang, Zhenhang Weng, Peng Zhu 0004, Xiao Han 0012, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.1
2024 Multistage Enhancement Network for Tiny Object Detection in Remote Sensing Images
abstract
With the rapid advances in deep learning techniques, remote sensing object detection has achieved remarkable achievements in recent years. However, tiny object detection remains unsatisfactory and suffers from two main drawbacks, including (1) the high sensitivity of IoU for location deviation in tiny objects and (2) the poor-quality feature representations of tiny objects. To address the aforementioned problems, we propose a Multi-stage Enhancement Network (MENet) that achieves the instance-level and feature-level enhancement of tiny objects from different stages of the detector. Since the IoU-based label assignment drastically deteriorates the positive samples for tiny objects, we first propose a Central Region-based (CR) label assignment to substitute it in the Region Proposal Network (RPN). The CR label assignment regards the anchors that fall into the central region of ground-truth boxes as positive samples, which provides more positive samples for tiny objects. Then, we design a Gated Context Aggregation (GCA) module that selectively aggregates valuable context information to enhance the feature representation of tiny objects. Additionally, we devise a positive RoI feature (pRoI) generator in the Region Convolutional Neural Network (R-CNN) to generate a rich diversity of high-quality positive RoI features for tiny objects. We conduct extensive experiments on AI-TOD and SODA-A datasets, and the results demonstrate the effectiveness of our proposed method.
Tianyang Zhang 0002, Xiangrong Zhang, Xiaoqian Zhu, Guanchun Wang, Xiao Han 0012, Xu Tang 0004, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.2
2024 High-Resolution Remote Sensing Image Segmentation With Global-Guided Normalization and Local Affinity Distillation
abstract
In recent years, high-resolution (HR) remote sensing images (RSIs) segmentation has received growing attention. The huge number of pixels poses a challenge to the semantic segmentation algorithm, which is limited by the storage of GPUs, so the current methods for processing HR RSIs are categorized into two main categories, i.e., global methods and local methods. The former downsamples the original image and loses a lot of feature details. The latter crops the original image and fails to obtain global contextual information. Both types of methods lead to limited segmentation accuracy. In this article, we propose an end-to-end framework, called global injection network (GINet), which explores two levels of feature distribution and feature relationship to achieve tradeoff between global context and local details. In concrete terms, we propose the global-guided normalization (GGN) module, which injects global context information into local branch and modulates local features using global features to enhance the global perception of local branch. In addition, to constrain the spatial consistency of two branches, inspired by the knowledge distillation technique, we propose local affinity distillation (LAD) loss, which distills the relations in local features into global features to keep the similarity of the relationships corresponding to patches in the two branches. The comprehensive experimental results on three large-scale land-cover classification datasets, DeepGlobe ($2448 \times 2448$), Inria Aerial ($5000 \times 5000$), and GID-15 ($7200 \times 6800$), confirm the effectiveness and superiority of our method in HR semantic segmentation tasks.
Peng Zhu 0004, Xiangrong Zhang, Xiao Han 0012, Puhua Chen, Xu Tang 0004, Xina Cheng, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.2
2024 Bio-Inspired Multi-Scale Contourlet Attention Networks
abstract
Inspired by the sparse and hierarchical features representation in the ventral stream of the human visual system, the biologically inspired multi-scale contourlet attention network (BMCAnet) is proposed to extract robust discriminative features. First, we constructed the multi-scale contourlet filter banks as a population of neurons in the primary visual cortex (V1), and extracted sparse features in a multi-scale and multi-direction way. It simulated a simple cell in V1 that responds to stimuli in a specific direction. Second, in order to refine contourlet features adaptively, the Shannon block attention module (SBAM) is introduced by integrating Shannon entropy as the third branch of the channel attention module (CAM), thus the weights of contourlet coefficients can be learned adaptively. Third, the responses of the spatial and spectral features are pooled by the proposed contourlet pooling layer to obtain the invariant structure features with the specified rules, which roughly stimulate the pooling process of complex cells in the V1 area. Last, the combination of global average pooling (GAP) and full connection (FC) is used for classification. The competitive results on eight databases demonstrate that the BMCAnet can effectively extract sparse and effective features for the classification tasks.
Mengkun Liu, Licheng Jiao, Xu Liu 0006, Lingling Li 0002, Fang Liu 0001, Shuyuan Yang 0001, Xiangrong Zhang
IEEE Trans. Multim.7
2024 A Patch Diversity Transformer for Domain Generalized Semantic Segmentation
abstract
Domain generalization (DG) is one of the critical issues for deep learning in unknown domains. How to effectively represent domain-invariant context (DIC) is a difficult problem that DG needs to solve. Transformers have shown the potential to learn generalized features, since the powerful ability to learn global context. In this article, a novel method named patch diversity Transformer (PDTrans) is proposed to improve the DG for scene segmentation by learning global multidomain semantic relations. Specifically, patch photometric perturbation (PPP) is proposed to improve the representation of multidomain in the global context information, which helps the Transformer learn the relationship between multiple domains. Besides, patch statistics perturbation (PSP) is proposed to model the feature statistics of patches under different domain shifts, which enables the model to encode domain-invariant semantic features and improve generalization. PPP and PSP can help to diversify the source domain at the patch level and feature level. PDTrans learns context across diverse patches and takes advantage of self-attention to improve DG. Extensive experiments demonstrate the tremendous performance advantages of the PDTrans over state-of-the-art DG methods.
Pei He, Licheng Jiao, Ronghua Shang, Xu Liu 0006, Fang Liu 0001, Shuyuan Yang 0001, Xiangrong Zhang, Shuang Wang 0001
IEEE Trans. Neural Networks Learn. Syst.7
2024 AUD-Net: A Unified Deep Detector for Multiple Hyperspectral Image Anomaly Detection via Relation and Few-Shot Learning
abstract
This article addresses the problem of the building an out-of-the-box deep detector, motivated by the need to perform anomaly detection across multiple hyperspectral images (HSIs) without repeated training. To solve this challenging task, we propose a unified detector [anomaly detection network (AUD-Net)] inspired by few-shot learning. The crucial issues solved by AUD-Net include: how to improve the generalization of the model on various HSIs that contain different categories of land cover; and how to unify the different spectral sizes between HSIs. To achieve this, we first build a series of subtasks to classify the relations between the center and its surroundings in the dual window. Through relation learning, AUD-Net can be more easily generalized to unseen HSIs, as the relations of the pixel pairs are shared among different HSIs. Secondly, to handle different HSIs with various spectral sizes, we propose a pooling layer based on the vector of local aggregated descriptors, which maps the variable-sized features to the same space and acquires the fixed-sized relation embeddings. To determine whether the center of the dual window is an anomaly, we build a memory model by the transformer, which integrates the contextual relation embeddings in the dual window and estimates the relation embeddings of the center. By computing the feature difference between the estimated relation embeddings of the centers and the corresponding real ones, the centers with large differences will be detected as anomalies, as they are more difficult to be estimated by the corresponding surroundings. Extensive experiments on both the simulation dataset and 13 real HSIs demonstrate that this proposed AUD-Net has strong generalization for various HSIs and achieves significant advantages over the specific-trained detectors for each HSI.
Ning Huyan, Xiangrong Zhang, Dou Quan, Jocelyn Chanussot, Licheng Jiao
IEEE Trans. Neural Networks Learn. Syst.2
2024 Contrastive Learning-Based Dual Dynamic GCN for SAR Image Scene Classification
abstract
As a typical label-limited task, it is significant and valuable to explore networks that enable to utilize labeled and unlabeled samples simultaneously for synthetic aperture radar (SAR) image scene classification. Graph convolutional network (GCN) is a powerful semisupervised learning paradigm that helps to capture the topological relationships of scenes in SAR images. While the performance is not satisfactory when existing GCNs are directly used for SAR image scene classification with limited labels, because few methods to characterize the nodes and edges for SAR images. To tackle these issues, we propose a contrastive learning-based dual dynamic GCN (DDGCN) for SAR image scene classification. Specifically, we design a novel contrastive loss to capture the structures of views and scenes, and develop a clustering-based contrastive self-supervised learning model for mapping SAR images from pixel space to high-level embedding space, which facilitates the subsequent node representation and message passing in GCNs. Afterward, we propose a multiple features and parameter sharing dual network framework called DDGCN. One network is a dynamic GCN to keep the local consistency and nonlocal dependency of the same scene with the help of a node attention module and a dynamic correlation matrix learning algorithm. The other is a multiscale and multidirectional fully connected network (FCN) to enlarge the discrepancies between different scenes. Finally, the features obtained by the two branches are fused for classification. A series of experiments on synthetic and real SAR images demonstrate that the proposed method achieves consistently better classification performance than the existing methods.
Fang Liu 0001, Xiaoxue Qian, Licheng Jiao, Xiangrong Zhang, Lingling Li 0002, Yuanhao Cui
IEEE Trans. Neural Networks Learn. Syst.4
2024 Semi-Supervised Multiscale Dynamic Graph Convolution Network for Hyperspectral Image Classification
abstract
In recent years, convolutional neural networks (CNNs)-based methods achieve cracking performance on hyperspectral image (HSI) classification tasks, due to its hierarchical structure and strong nonlinear fitting capacity. Most of them, however, are supervised approaches that need a large number of labeled data to train them. Conventional convolution kernels are fixed shape of rectangular with fixed sizes, which are good at capturing short-range relations between pixels within HSIs but ignore the long-range context within HSIs, limiting their performance. To overcome the limitations mentioned above, we present a dynamic multiscale graph convolutional network (GCN) classifier (DMSGer). DMSGer first constructs a relatively small graph at region-level based on a superpixel segmentation algorithm and metric-learning. A dynamic pixel-level feature update strategy is then applied to the region-level adjacency matrix, which can help DMSGer learn the pixel representation dynamically. Finally, to deeply understand the complex contents within HSIs, our model is expanded into a multiscale version. On the one hand, by introducing graph learning theory, DMSGer accomplishes HSI classification tasks in a semi-supervised manner, relieving the pressure of collecting abundant labeled samples. Superpixels are generally in irregular shapes and sizes which can group only similar pixels in a neighborhood. On the other hand, based on the proposed dynamic-GCN, the pixel-level and region-level information can be captured simultaneously in one graph convolution layer such that the classification results can be improved. Also, due to the proper multiscale expansion, more helpful information can be captured from HSIs. Extensive experiments were conducted on four public HSIs, and the promising results illustrate that our DMSGer is robust in classifying HSIs. Our source codes are available at https://github.com/TangXu-Group/DMSGer.
Yuqun Yang, Xu Tang 0004, Xiangrong Zhang, Jingjing Ma 0001, Fang Liu 0034, Xiuping Jia, Licheng Jiao
IEEE Trans. Neural Networks Learn. Syst.3
2023 CTACL:Hyperspectral Image Change Detection Based on Adaptive Contrastive Learning
abstract
Hyperspectral image change detection (HSI-CD) can accurately identify changing regions by capturing subtle spectral differences and has become a research hotspot in the field of remote sensing (RS). Convolutional neural networks (CNNs) have excellent local context modeling capabilities and have been proven to be powerful feature extractors in HSI-CD. However, due to its inherent network structure limitation, CNN cannot well mine and represent the sequential properties of spectral features, especially the medium and long-term dependencies. In contrast, transformer-based network architecture shows a strong ability to model long-distance dependencies, which can fully mine and extract global features, but exhibits weak performance in extracting local information. To this end, we propose HSI-CD network based on adaptive contrastive learning (CTACL). Specifically, we first propose a parallel network of CNNs and transformers to mine local and global temporal-spatial-spectral features of HSI, respectively. Second, we propose adaptive contrastive learning to pre-train the network to learn the latent features of a large amount of unlabeled data and better mine and utilize local and global information. Experimental results on the farmland dataset show that the proposed method performs well.
Shunli Tian, Xiangrong Zhang, Guanchun Wang, Xiao Han 0012, Puhua Chen, Xina Cheng
IGARSS2
2023 Global-Local Representation Coupling Network for Remote Sensing Image Change Detection
abstract
Change detection is one of the important tasks in remote sensing image processing, and the powerful feature extraction ability of convolutional networks has achieved some success in change detection. However, the problem of the limited field size of pure convolutional networks makes the change detection accuracy of high-resolution remote sensing images limited. The introduction of transformers can link the concept of long-distance in space and time. Therefore, in order to maximize the respective advantages of transformers and CNNs, we propose a new parallel architecture. To accomplish the above goal, we propose a new network, which consists of a local detail branch and transformer global spatial-temporal feature branch and a feature fusion module. The experimental results reach the current sota level.
Fanghan Yang, Xiangrong Zhang, Peng Zhu 0004, Zhenhang Weng, Puhua Chen
IGARSS2
2023 Domain Adversarial Debiased Self-Training for Hyperspectral Image Classification
abstract
Unsupervised domain adaptation (UDA) has been widely used in hyperspectral image (HSI) classification. Domain adversarial learning methods and self-training methods are two major UDA methods. Most existing methods use these two methods independently, which limit their capacity to knowledge transferability and robustness of classification. Therefore, a novel domain adversarial debiased self-training (DADST) is proposed to combine these two methods, where domain adversarial learning reduces the domain discrepancy and debiased self-training updates model via a self-paced curriculum policy. To this end, we first apply debiased self-training into HSI classification. Then domain adversarial learning module is combined to align joint distribution between the source domain and the target domain. Experimental results on two cross-scene HSI datasets demonstrate that the proposed DADST method outperforms other domain adaptation approaches.
Jie Feng 0003, Ziyu Zhou 0009, Xiangrong Zhang, Licheng Jiao
IGARSS4
2023 3D-Mglnet: Moving Vehicle Detection in Satellite Videos with 3D Motion-Guided Lightweight Network
abstract
Object detectors based on convolutional neural networks have been widely-applied to detect moving vehicles in satellite videos. However, many detectors render superior detection accuracy at the expense of increased computational complexity and decreased inference speed. This prevents these detectors from being deployed into mobile devices. In this paper, an efficient 3D motion-guided lightweight network (3D-MGLNet) is proposed. Specifically, 3D-MGLNet constructs a motion-guided module based on 3D convolution to extract motion cues from spatial-temporal information. This module uses model compression strategies to detect moving vehicles in real-time while following the principle of "fewer channels, smaller convolution kernels," significantly reducing the number of parameters and computational complexity. Extensive experiments are conducted on the Jilin-1 and SkySat satellite video datasets. The results demonstrate that 3D-MGLNet gains strong performance by striking an excellent tradeoff between resource and accuracy, resulting in the fewest parameters (0.35M) and fastest speed (66.84 fps) compared to other popular models.
Jie Li 0001, Jie Feng 0003, Quanpeng Jiang, Xiangrong Zhang, Licheng Jiao
IGARSS5
2023 Unsupervised SAR Image Change Detection Based on Feature Fusion of Information Transfer
abstract
Synthetic aperture radar (SAR) image change detection is a hot but challenging task due to SAR images’ complex contents and inherent speckle noises. The expected change detection methods should reduce the influence of speckle noises, obtain the discriminative feature representations, and generate accurate change maps simultaneously. To these ends, we propose a new SAR image change detection method named feature fusion of information transfer network (FFITN). First, we develop a hybrid convolution block to depress the speckle noise impacts and explore the valuable information from SAR images. Thus, the feature extraction module (FEM) is constructed to obtain the multi-level features. Then, an information transfer module (ITM) is proposed to capture the salient regions from various aspects. Also, the salient knowledge is transferred among features at different levels to enhance their discrimination. Next, a self-attention-based feature fusion module (SAFFM) is introduced to fuse various features. Finally, a change map generation module (CMGM) with the clustering algorithm and specific loss functions is designed to produce the pseudo labels and change maps. Experimental results on three public SAR data sets demonstrate the model’s effectiveness. Our source codes are available at https://github.com/TangXu-Group/FFITN.
Jingjing Ma 0001, Xu Tang 0004, Yuqun Yang, Xiangrong Zhang, Licheng Jiao
IEEE Geosci. Remote. Sens. Lett.5
2023 A semi-supervised multi-task learning framework for cancer classification with weak annotation in whole-slide images
Zeyu Gao 0001, Bangyang Hong, Yang Li 0139, Xianli Zhang, Jialun Wu, Chunbao Wang 0002, Xiangrong Zhang, Tieliang Gong, Yefeng Zheng 0001, Deyu Meng, Chen Li 0011
Medical Image Anal.7
2023 Knowledge transfer evolutionary search for lightweight neural architecture with dynamic inference
Xiaoxue Qian, Fang Liu 0001, Licheng Jiao, Xiangrong Zhang, Shuo Li 0010, Puhua Chen, Xu Liu 0006
Pattern Recognit.4
2023 Cost-Sensitive Boosting Pruning Trees for Depression Detection on Twitter
abstract
Depression is one of the most common mental health disorders, and a large number of depressed people commit suicide each year. Potential depression sufferers usually do not consult psychological doctors because they feel ashamed or are unaware of any depression, which may result in severe delay of diagnosis and treatment. In the meantime, evidence shows that social media data provides valuable clues about physical and mental health conditions. In this paper, we argue that it is feasible to identify depression at an early stage by mining online social behaviours. Our approach, which is innovative to the practice of depression detection, does not rely on the extraction of numerous or complicated features to achieve accurate depression detection. Instead, we propose a novel classifier, namely, Cost-sensitive Boosting Pruning Trees (CBPT), which demonstrates a strong classification ability on two publicly accessible Twitter depression detection datasets. To comprehensively evaluate the classification capability of CBPT, we use additional three datasets from the UCI machine learning repository and CBPT obtains appealing classification results against several state of the arts boosting algorithms. Finally, we comprehensively explore the influence factors for the model prediction, and the results manifest that our proposed framework is promising for identifying Twitter users with depression.
Zheheng Jiang, Feixiang Zhou, Long Chen 0019, Jialin Lyu, Xiangrong Zhang, Qianni Zhang, Abdul Hamid Sadka, Yinhai Wang, Ling Li 0010, Huiyu Zhou 0001
IEEE Trans. Affect. Comput.7
2023 A Collaborative Learning Tracking Network for Remote Sensing Videos
abstract
With the increasing accessibility of remote sensing videos, remote sensing tracking is gradually becoming a hot issue. However, accurately detecting and tracking in complex remote sensing scenes is still a challenge. In this article, we propose a collaborative learning tracking network for remote sensing videos, including a consistent receptive field parallel fusion module (CRFPF), dual-branch spatial-channel co-attention (DSCA) module, and geometric constraint retrack strategy (GCRT). Considering the small-size objects of remote sensing scenes are difficult for general forward networks to extract effective features, we propose a CRFPF-module to establish parallel branches with consistent receptive fields to separately extract from shallow to deep features and then fuse hierarchical features adaptively. Since the objects and their background are difficult to distinguish, the proposed DSCA-module uses the spatial-channel co-attention mechanism to collaboratively learn the relevant information, which enhances the saliency of the objects and regresses to precise bounding boxes. Considering the interference of similar objects, we designed a GCRT-strategy to judge whether there is a false detection through the estimated motion trajectory and then recover the correct object by weakening the feature response of interference. The experimental results and theoretical analysis on multiple datasets demonstrate our proposed method's feasibility and effectiveness. Code and net are available at https://github.com/Dawn5786/CoCRF-TrackNet.
Licheng Jiao, Hao Zhu 0009, Fang Liu 0001, Shuyuan Yang 0001, Xiangrong Zhang, Shuang Wang 0001, Rong Qu
IEEE Trans. Cybern.6
2023 A Normalized Spatial-Spectral Supervoxel Segmentation Method for Multispectral Point Cloud Data
abstract
Airborne LiDAR point cloud segmentation (PCS) is often employed as a preprocessing step for the subsequent object recognition for scene interpretation. Current segmentation methods often aim at single-wavelength LiDAR data by fully exploiting the spatial information, which makes them unsuitable for multispectral point cloud (MPC) data due to ignoring the use of spectral signatures. In this article, a normalized spatial–spectral supervoxel segmentation method is proposed for MPC data. Specifically, a normalized spectral–spatial metric is developed to construct the${k}$-dimensional tree (KD tree) for MPC data clustering. Considering the uneven density distribution of MPC, an adaptive energy minimization principle based on the sum of the distance is devised to accurately select the seed points of voxels, solving the problem of undersegmentation. To reduce the cross-boundary points, the normalized spectral–spatial metric with the concave–convex judgment is extended to further optimize the edges between adjacent voxels. An important asset of our method is to segment MPC without the need for any manual annotation. Experiments on two MPC datasets show that the proposed method yields better performance compared to several comparative methods.
Likun Chen, Yanfeng Gu, Xian Li 0001, Xiangrong Zhang, Baisen Liu
IEEE Trans. Geosci. Remote. Sens.4
2023 Distilling Segmenters From CNNs and Transformers for Remote Sensing Images' Semantic Segmentation
abstract
Semantic segmentation is a crucial task in remote sensing and has been predominantly performed using convolutional neural networks (CNNs) for the past decade. Recently, transformers with self-attention mechanisms have demonstrated superior performance compared to CNNs. However, due to the locality of CNN and the high computational complexity and massive data resource requirements of transformer, neither of them can be well applied in resource-constrained practical remote sensing scenarios. Motivated by the limitations of using either convolutional neural networks (CNNs) or transformers alone in the task of semantic segmentation of remote sensing images, a novel cross-model knowledge distillation framework, named distilling segmenters from CNNs and transformers (DSCT), is proposed in this paper to harness the complementary advantages of both models. The framework utilizes a channel-weighted attention-guided feature distillation (CAFD) module to condense the feature from the teacher model and enhance the student model’s focus on the teacher-focused regions. Additionally, a target-nontarget knowledge distillation (TNKD) module is proposed that decouples logit distillation into target and nontarget knowledge distillation to guide the student model in learning the underlying representations and decision boundaries from the teacher model. By learning the complementary knowledge from the teacher, our proposed DSCT framework improves the student’s segmentation performance without adding trainable parameters. Experiments on four available remote sensing datasets (ISPRS Potsdam, Vaihingen, GID and LoveDA) indicate that the proposed DSCT outperforms the state-of-the-art knowledge distillation methods and demonstrates its effectiveness and robustness.
Guoming Gao, Tianzhu Liu, Yanfeng Gu, Xiangrong Zhang
IEEE Trans. Geosci. Remote. Sens.5
2023 MR-Selection: A Meta-Reinforcement Learning Approach for Zero-Shot Hyperspectral Band Selection
abstract
Band selection is an effective method to deal with the difficulties in image transmission, storage, and processing caused by redundant and noisy bands in hyperspectral images (HSIs). Existing band selection methods usually need to learn a specific model for each HSI dataset, which ignores the inherent correlation and common knowledge among different band selection tasks. Meanwhile, these methods lead to a huge waste of computation. In this article, a novel zero-shot band selection method, called MR-Selection, is proposed for HSI classification. It formalizes zero-shot band selection as a metalearning problem, where advantage actor–critic algorithm-based reinforcement learning (A2C-RL) is designed to extract the metaknowledge in the band selection tasks of various seen hyperspectral datasets through a shared agent. To learn a consistent representation among different tasks, a dynamic structure-aware graph convolutional network is constructed to build a shared agent in A2C-RL. In A2C-RL, the state is tailored in a feasible way and easy to adapt to various tasks. Meanwhile, the reward is defined according to an efficient evaluation network, which can evaluate each state effectively without any fine-tuning. Furthermore, a two-stage optimization strategy is designed to coordinate optimization directions of a shared agent from different tasks effectively. Once the shared agent is optimized, it can be directly applied to unseen HSI band selection tasks without any available samples. Experimental results demonstrate the effectiveness and efficiency of the MR-Selection on the band selection of unseen HSI datasets.
Jie Feng 0003, Gaiqin Bai, Xiangrong Zhang, Ronghua Shang, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.4
2023 Multi-Complementary Generative Adversarial Networks With Contrastive Learning for Hyperspectral Image Classification
abstract
In the last decade, generative adversarial network (GAN) and its variants provide a powerful training mechanism for hyperspectral image (HSI) classification. In HSIs, the distribution of samples is more complicated due to the existence of abundant spatial-spectral information and multi-scale information. The single generation pattern of GANs is prone to modal collapse for the sample generation of HSIs. Moreover, the promotion of the generator only relies on adversarial learning with the discriminator, which limits the generator’s performance. To address these problems, a multi-complementary GANs with contrastive learning (CMC-GAN) is proposed. CMC-GAN consists of two groups of GANs, where coarse-grained GAN adopts the structure in encoder-decoder form for hidden fine-scale and coarse-scale generation, and another fine-grained GAN is responsible for fine-scale generation. In fine-grained GAN, the discriminator is constructed to distinguish the fine-scale samples from different generators, which enforces the joint optimization of these two groups of GANs and makes GANs generate diverse multi-scale samples. Furthermore, a novel contrastive learning constraint is added into GANs, where a unidirectional contrastive loss guarantees the generators to extract intra-class invariant representation and a class-specific contrastive loss urges the discriminators to learn more discriminative features for classification. Finally, both discriminators are adaptively-fused to extract complementary multi-scale spatial-spectral features for classification under the guidance of diverse generated samples. The experimental results demonstrate CMC-GAN has superior classification performance, especially for small sample classification.
Jie Feng 0003, Zizhuo Gao, Ronghua Shang, Xiangrong Zhang, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.4
2023 Multisource Joint Representation Learning Fusion Classification for Remote Sensing Images
abstract
Multisource remote sensing images provide complementary multidimensional information for reliable and accurate classification. However, gaps in imaging mechanisms result in heterogeneity between multiple source images. During fusion, this heterogeneity causes the generated multisource representations may be redundant and ignore discriminative uni-source information, which significantly hampers the fusion classification performance. To address this challenge, we introduce a novel multisource joint representation learning method for remote sensing image fusion classification, termed Multisource Information Bottleneck Fusion Network (MIBF-Net). Based on the Information Bottleneck principle, MIBF-Net employs mutual information constraints to effectively integrate multisource information, generating a comprehensive and non-redundant multisource representation. Specifically, MIBF-Net first introduces an attribution-driven noise adaptation layer to dynamically balance the speed of feature learning across sources for extracting discriminative uni-source intrinsic information. Furthermore, a cross-source relationship encoding module is designed to fully explore cross-source complex dependencies for enhancing the richness of fused representations. Finally, we design an information bottleneck fusion module to fuse uni-source semantic information and cross-source information while reducing redundancy. In particular, we employ variational inference techniques to effectively address the mutual information optimization problem and provide theoretical derivations. Extensive experimental results on three heterogeneous multisource remote sensing data benchmarks show that the model significantly outperforms the state-of-the-art methods.
Xueli Geng, Licheng Jiao, Lingling Li 0002, Fang Liu 0001, Xu Liu 0006, Shuyuan Yang 0001, Xiangrong Zhang
IEEE Trans. Geosci. Remote. Sens.7
2023 MidNet: An Anchor-and-Angle-Free Detector for Oriented Ship Detection in Aerial Images
abstract
Ship detection in aerial images remains an active yet challenging task due to its arbitrary object orientation and various aspect ratios from the bird’s-eye perspective. Most existing oriented objection detection methods rely on angular prediction or predefined anchor boxes, making these methods highly sensitive to unstable angular regression and excessive hyper-parameter setting. To address these issues, we replace the angular-based object encoding with an anchor-and-angle-free paradigm, and propose a novel detector deploying a center and four midpoints for encoding each oriented object, namely MidNet. Moreover, MidNet designs a novel symmetrical deformable convolution for enhanceing the features of midpoints, then the center and midpoints for an identical ship are adaptively matched by predicting corresponding centripetal shift and matching radius. Finally, a concise analytical geometry algorithm is proposed to calculate the ship orientation and refine the keypoints step-wisely for building precise oriented bounding boxes. On two public ship detection datasets, HRSC2016 and FGSD2021, MidNet outperforms the state-of-the-art detectors by achieving APs of 90.52% and 86.50%.
Yuping Liang, Jie Feng 0003, Xiangrong Zhang, Junpeng Zhang 0002, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.3
2023 Multipretext-Task Prototypes Guided Dynamic Contrastive Learning Network for Few-Shot Remote Sensing Scene Classification
abstract
As a content management technique, remote sensing (RS) scene classification (RSSC) always attracts researchers’ attention. In the past decades, many successful methods have been proposed. Nevertheless, their prerequisite is that there are large labeled data sets, which is a strict demand in practice. To resolve this contradiction, developing RSSC models with the help of few-shot learning (FSL) has become popular. Due to lacking prior knowledge, most of the existing few-shot RSSC models pay attention to the learning algorithm. However, they do not attach importance to the complex contents within RS scenes and the intricate inter-/intra-class relations between RS scenes. This would influence their performance negatively. In this paper, we propose a new few-shot RSSC model named multi-pretext-task prototypes guided dynamic contrastive learning network (MPCL-Net). MPCL-Net consists of a multi-pretext tasks generation sub-module, a deep feature learning sub-module, and a joint optimization sub-module. First, two RS-oriented pretext tasks are constructed under the self-supervised learning (SSL) framework in the multi-pretext tasks generation sub-module, which aim to explore multi-scale and rotation-invariant information from RS scenes. Second, a simple convolutional neural network (CNN) is developed in the deep feature learning sub-module to transform the RS scenes into visual features. Third, three loss functions are formulated and integrated in the joint optimization sub-module. Their goals are to fully capture the diverse land covers within RS scenes and compact/separate the intra-/inter-class samples with limited supervision. Finally, our MPCL-Net can be trained in a meta way. The positive results counted on the three public RS scene data sets confirm that our MPCL-Net is helpful to RSSC tasks under the few-shot scenario. Our source codes are available at https://github.com/TangXu-Group/Remote-Sensing-Images-Classification/tree/main/MPCL.
Jingjing Ma 0001, Weiquan Lin, Xu Tang 0004, Xiangrong Zhang, Fang Liu 0001, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.4
2023 Interacting-Enhancing Feature Transformer for Cross-Modal Remote-Sensing Image and Text Retrieval
abstract
Cross-modal remote sensing image-text retrieval (CMRSITR) is a challenging topic in the remote sensing (RS) community. It has gained growing attention because it can be flexibly used in many practical applications. In the current deep era, with the help of deep convolutional neural networks (DCNNs), many successful CMRSITR methods have been proposed. Most of them first learn valuable features from RS images and texts respectively. Then, the obtained visual and textual features are mapped into a common space for the final retrieval. The above operations are feasible, however, two difficulties are still to be solved. One is that the semantics within the visual and textual features are misaligned due to the independent learning manner. The other one is that the deep links between RS images and texts cannot be fully explored by simple common space mapping. To overcome the above challenges, we propose a new model named interacting-enhancing feature transformer (IEFT) for CMRSITR, which regards the RS images and texts as a whole. First, a simple feature embedding module (FEM) is developed to map images and texts into the visual and textual feature spaces. Second, an information interacting-enhancing module (IIEM) is designed to simultaneously model the inner relationships between RS images and texts and enhance the visual features. IIEM consists of three feature interacting-enhancing (FIE) blocks, each of which contains an inter-modality relationship interacting (IMRI) sub-block and a visual feature enhancing (VFE) sub-block. The duty of IMRI is to exploit the hidden relations between cross-modal data, while the responsibility of VFE is to improve the visual features. By combining them, semantic bias can be mitigated, and the complex contents of RS images can be studied. Finally, the retrieval module (RM) is constructed to generate the matching scores for deciding the search results. Extensive experiments are conducted on four public RS data sets. The positive results demonstrate that our IEFT can achieve superior retrieval performance compared with many existing methods. Our source codes are available at https://github.com/TangXu-Group/Cross-modal-remote-sensing-image-and-text-retrieval-models/tree/main/IEFT.
Xu Tang 0004, Yijing Wang 0004, Jingjing Ma 0001, Xiangrong Zhang, Fang Liu 0034, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.4
2023 WNet: W-Shaped Hierarchical Network for Remote-Sensing Image Change Detection
abstract
Change detection (CD) is a hot research topic in the remote sensing (RS) community. With the increasing availability of high-resolution (HR) RS images, there is a growing demand for CD models with high detection accuracy and generalization ability. In other words, the CD models are expected to work well for various HRRS images. Convolutional neural networks (CNNs) have been dominated in HRRS image CD due to their excellent information extraction and nonlinear fitting capabilities. However, they are not skilled in modeling long-range contexts hidden in HRRS images, which limits their performance in CD tasks more or less. Recently, the Transformer, which is good at extracting global context dependencies, has become popular in the RS community. Nevertheless, detailed local knowledge receives insufficient emphasis in common Transformers. Considering the above discussion, we combine CNN and Transformer and propose a new W-shaped dual Siamese branch hierarchical network for HRRS image CD named WNet. WNet first incorporates a Siamese CNN and a Siamese Transformer into a dual-branch encoder to extract multi-level local fine-grained features and global long-range contextual dependencies. Also, we introduce deformable ideas into the Siamese CNN and Transformer to make WNet understand the critical and irregular areas within HRRS images. Second, the difference enhancement module (DEM) is developed and embedded into the encoder to produce the difference feature maps at different levels. Using simple pixel-wise subtraction and channel-wise concatenation, the changes of interest and irrelevant changes can be highlighted and suppressed in a learnable manner. Next, the multi-level difference feature maps are fused stage by stage by CNN-Transformer fusion modules (CTFMs), which are the basic units of the decoder in WNet. In CTFM, the local, global, and cross-scale clues are taken into account to ensure the integrity of information. Finally, a simple classifier is constructed and added at the top of the decoder to predict the change maps. Positive experimental results counted on four public datasets demonstrate that the proposed WNet is helpful in HRRS image CD tasks. Our source codes are available at https://github.com/TangXu-Group/Remote-Sensing-Image-Change-Detection/tree/main/WNet.
Xu Tang 0004, Jingjing Ma 0001, Xiangrong Zhang, Fang Liu 0001, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.4
2023 High-Quality Angle Prediction for Oriented Object Detection in Remote Sensing Images
abstract
Oriented object detection is a challenging task in remote sensing, where the detected objects can be represented by oriented bounding boxes (OBBs). Angle prediction in oriented object detection has been widely studied, due to its crucial role in object detection. However, the precision of angle prediction is severely limited by misalignments in most of the existing methods, including representation-, evaluation-, and optimization-based misalignments. To alleviate these misalignments, this paper presents a novel angle prediction method, called Angle Quality Estimation (AQE). Specifically, our proposed AQE transforms the angle prediction task into a distribution estimation task to address the representation misalignment problem and implicitly measure the quality of the predicted angles. Based on the estimated angle quality, we then propose a new metric to comprehensively evaluate the quality of OBBs. Then we propose an object aspect ratio based loss function to optimize angle prediction for addressing the optimization misalignment. Our proposed AQE is a plug-and-play method, which can be embedded on any existing oriented object detector. Experimental results on three public benchmarks, including DOTA, HRSC2016, and ICDAR2015 datasets, show that our method achieves better performance than the other state-of-the-art.
Guanchun Wang, Xiangrong Zhang, Peng Zhu 0004, Xu Tang 0004, Puhua Chen, Licheng Jiao, Huiyu Zhou 0001
IEEE Trans. Geosci. Remote. Sens.2
2023 Superpixel Consistency Saliency Map Generation for Weakly Supervised Semantic Segmentation of Remote Sensing Images
abstract
The weakly supervised semantic segmentation (WSSS) method aims to assign semantic labels to each image pixel from weak (image-level) instead of strong (pixel-level) labels, which can greatly reduce human labor costs. However, there are some problems in WSSS of remote sensing images such as how to locate labels accurately, and how to get precise segmentation edges. To address these issues, we propose a novel framework directly transferring the scene classification model to perform semantic segmentation. We first train a multi-label scene classification network as the encoder to obtain the pre-trained model, then the feature learned by the model is transferred to the decoder. Different from other methods, we propose a saliency map generator instead of the Class Activation Map for more accurate location information by making pixels belonging to the same class lie close together while different classes are separated in feature space. Meanwhile, we take the superpixel patch as processing unit to provide precise boundary inhibition for the saliency map. To assign semantic labels for each patch, combined with extracted salient region, we propose a module responsible for exploiting the consistency of spatial and semantic similarity between different patches. Finally, we incorporate the above two modules to supervise the training process of the decoder without generating pseudo labels as most methods do, thus simplifying the training process. Experimental results show that our method outperforms other weakly supervised approaches on DLRSD and WHDLD datasets with at least a 3% improvement on mean intersection over union.
Xiaopeng Zeng, Tengfei Wang 0001, Xiangrong Zhang, Yanfeng Gu
IEEE Trans. Geosci. Remote. Sens.4
2023 MFGNet: Multibranch Feature Generation Networks for Few-Shot Remote Sensing Scene Classification
abstract
Few-shot remote sensing scene classification aims to identify unseen classes using only a small number of labeled samples. Considering the large intra-class variances and inter-class similarity of remote sensing scenes, most existing methods focus on feature extraction, ignoring the overfitting problem caused by insufficient samples. To this end, we propose a novel few-shot learning framework, called multibranch feature generation networks (MFGNets), which solves the few-shot scene classification from the source by online sample generation at the representation space. Specifically, we first build a feature generation net to transform the few-shot classification into a regular classification problem, in which the generated samples are achieved by combining the class-specific features with the sampled intra-class features. Then, to ensure the quality of the generated samples, we introduce two novel regularization terms: the intra-class diversity loss (ID-Loss) and the inter-class consistency loss (IC-Loss), which aid the model in generating more diverse samples. Furthermore, we introduce a scale-angle aware self-supervised pretext to learn scale-invariant and rotation-invariant features, improving the model’s feature representation capability in remote sensing scenes. We evaluate the proposed method on three publicly available datasets, namely UC_Merced, NWPU-RESISC45, and AID. Our approach has achieved state-of-the-art performance, with an improvement of more than 3.31%, 2.64%, and 6.86% on the most challenging 1-shot tasks, respectively.
Xiangrong Zhang, Xiyu Fan, Guanchun Wang, Puhua Chen, Xu Tang 0004, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.1
2023 CAST: A Cascade Spectral-Aware Transformer for Hyperspectral Image Change Detection
abstract
Hyperspectral image change detection (HSI-CD) aims to detect subtle changes on the Earth’s surface through approximately continuous spectral information, which has gradually become a very important research hotspot in the field of remote sensing (RS). In recent years, convolutional neural networks (CNNs) based HSI-CD methods have shown strong feature extraction capabilities. However, due to the simple fusion of spectral information in the channel dimension by CNN, the medium and long-term sequence properties of spectral features cannot be well mined and represented. Most previous studies mainly extract semantic features from images at different times, ignoring the temporal correlation between features, which cannot fully extract and effectively utilize temporal-spatial-spectral features. To this end, this paper proposes a cascade spectral aware transformer (CAST) for HSI-CD. First, we propose a temporal-spatial transformer (TS-Former) to enhance the temporal correlation and spatial global relationship of extracted features, thereby addressing the insufficient consideration of temporal correlation. Second, a spectral awareness transformer (SA-Former) is designed to better mine and represent the sequence properties of spectral features, especially the medium and long-term dependencies. Finally, we observe a spectral distortion in the process of extracting temporal-spatial features and based on this present a spectral constraint module (SCM) to preserve the sequence properties of spectral features and reduce the distortion of the spectrum. Extensive experiments on three challenging hyperspectral datasets demonstrate that our method achieves state-of-the-art results. The code is available at: https://github.com/tianshunli/CAST.
Xiangrong Zhang, Shunli Tian, Guanchun Wang, Xu Tang 0004, Jie Feng 0003, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.1
2023 Bidirectional Multiple Object Tracking Based on Trajectory Criteria in Satellite Videos
abstract
Multiple object tracking (MOT) in satellite videos requires to detect all objects belonging to specified categories and identify each object, which plays a basic and necessary role in automatic driving, traffic surveillance, and smart city. The traditional MOT methods in satellite videos mostly follow the detection–association framework. However, the detection–association framework works under a strict assumption that all objects are correctly localized by the detector. In practice, MOT in satellite videos faces challenges such as low resolution, tiny objects, and the wide field of view, which leads to the degradation of detector performance. In order to reduce the impact of detector degradation, we propose a bidirectional MOT framework based on trajectory criteria (BMTC) in satellite videos. In BMTC, the single object tracking (SOT) tracker carries out locating the objects between consecutive frames and the detector is just used for finding new objects. Therefore, it is less dependent on the detector performance. According to the characteristics of satellite videos, the trajectory criteria are designed to control the state of the tracker, which includes trajectory density, the limit of consecutive virtual motion predictions, and trajectory similarity measurement. Invalid fragment trajectory backtracking is implemented to alleviate the misalignment caused by the above subsection trajectory criteria. The method is validated on the VISO benchmark and SkySat-1 dataset. The experimental results show the improvement of completeness and accuracy, and the proposed tracker achieves the state-of-the-art performance.
Xiangrong Zhang, Zhongjian Huang, Xina Cheng, Jie Feng 0003, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.2
2023 Spectral-Spatial Distribution Consistent Network Based on Meta-Learning for Cross-Domain Hyperspectral Image Classification
abstract
Cross-domain networks can solve the problem of insufficient labeled samples, especially for hyperspectral images (HSIs) where obtaining labeled samples is time-consuming and laborious. Most of the current methods rely on the spatial information to achieve domain alignment, without considering the rich spectral information of HSIs. Furthermore, the methods based on convolutional neural network (CNN) cannot get the spatial information of irregular image regions, resulting in poor classification results of object edges. Therefore, we design a spectral-spatial distribution consistent network (SSDC) based on meta-learning. Firstly, to improve the feature extraction ability of the cross-domain classification model, we introduce a feature pre-extraction module, which uses the spectral attention mechanism and the alternating meta-learning method to obtain the general features of the source domain and the discriminative features of the target domain, so as to obtain the spectral weight matrix for subsequent processing. Secondly, we propose a spectral consistent module based on singular value decomposition, which increases the difference between different classes of features by penalizing the singular values of the feature matrix to achieve data distribution alignment in the spectral dimension. Finally, aiming at the low classification accuracy of irregular image regions, we propose a spatial consistent module to obtain non-local spatial topological information through stacked cross modules and graph sample and aggregate networks, which can reduce domain shift. The experiments of SSDC on four classical HSI datasets show that the proposed method can obtain competitive results with other methods based on CNN and cross-domain.
Xiangrong Zhang, Qi Zhen, Xiao Han 0012, Puhua Chen, Xu Tang 0004, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.1
2023 Semantics and Contour Based Interactive Learning Network for Building Footprint Extraction
abstract
Building footprint extraction plays an important role in the analysis of remote sensing images and has an extensive range of applications. Obtaining precise boundaries of buildings remains a challenge in existing building extraction methods. Some previous works have made notable efforts to address this concern. However, most of these methods require cumbersome and expensive post-processing steps. Moreover, they ignored the correlation between building semantics and contours, which we believe is crucial for building footprint extraction. To mitigate this issue, our paper presents an intuitive and effective framework that explores semantic and contour cues of buildings and fully excavates their correlation. Specifically, we construct an interactive dual-stream decoder. The Intermediate connections within this decoder interactively transmit features between branches, contributing to learning correlations between semantics and contours. We propose the Semantic Collaboration Module (SCM) to strengthen the connection between the two branches. To further boost performance, we build the Multi-Scale Semantic Context Fusion Module (MSCF) to fuse semantic information from the higher and lower layers of the network, allowing the network to obtain superior feature representations. The experimental results on the WHU, INRIA, and Massachusetts building datasets demonstrate the superior performance of our method.
Xiaoqian Zhu, Xiangrong Zhang, Tianyang Zhang 0002, Xu Tang 0004, Puhua Chen, Huiyu Zhou 0001, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.2
2023 SDANet: Semantic-Embedded Density Adaptive Network for Moving Vehicle Detection in Satellite Videos
abstract
In satellite videos, moving vehicles are extremely small-sized and densely clustered in vast scenes. Anchor-free detectors offer great potential by predicting the keypoints and boundaries of objects directly. However, for dense small-sized vehicles, most anchor-free detectors miss the dense objects without considering the density distribution. Furthermore, weak appearance features and massive interference in the satellite videos limit the application of anchor-free detectors. To address these problems, a novel semantic-embedded density adaptive network (SDANet) is proposed. In SDANet, the cluster-proposals, including a variable number of objects, and centers are generated parallelly through pixel-wise prediction. Then, a novel density matching algorithm is designed to obtain each object via partitioning the cluster-proposals and matching the corresponding centers hierarchically and recursively. Meanwhile, the isolated cluster-proposals and centers are suppressed. In SDANet, the road is segmented in vast scenes and its semantic features are embedded into the network by weakly supervised learning, which guides the detector to emphasize the regions of interest. By this way, SDANet reduces the false detection caused by massive interference. To alleviate the lack of appearance information on small-sized vehicles, a customized bi-directional conv-RNN module extracts the temporal information from consecutive input frames by aligning the disturbed background. The experimental results on Jilin-1 and SkySat satellite videos demonstrate the effectiveness of SDANet, especially for dense objects.
Jie Feng 0003, Yuping Liang, Xiangrong Zhang, Junpeng Zhang 0002, Licheng Jiao
IEEE Trans. Image Process.3
2023 SAGN: Semantic-Aware Graph Network for Remote Sensing Scene Classification
abstract
The scene classification of remote sensing (RS) images plays an essential role in the RS community, aiming to assign the semantics to different RS scenes. With the increase of spatial resolution of RS images, high-resolution RS (HRRS) image scene classification becomes a challenging task because the contents within HRRS images are diverse in type, various in scale, and massive in volume. Recently, deep convolution neural networks (DCNNs) provide the promising results of the HRRS scene classification. Most of them regard HRRS scene classification tasks as single-label problems. In this way, the semantics represented by the manual annotation decide the final classification results directly. Although it is feasible, the various semantics hidden in HRRS images are ignored, thus resulting in inaccurate decision. To overcome this limitation, we propose a semantic-aware graph network (SAGN) for HRRS images. SAGN consists of a dense feature pyramid network (DFPN), an adaptive semantic analysis module (ASAM), a dynamic graph feature update module, and a scene decision module (SDM). Their function is to extract the multi-scale information, mine the various semantics, exploit the unstructured relations between diverse semantics, and make the decision for HRRS scenes, respectively. Instead of transforming single-label problems into multi-label issues, our SAGN elaborates the proper methods to make full use of diverse semantics hidden in HRRS images to accomplish scene classification tasks. The extensive experiments are conducted on three popular HRRS scene data sets. Experimental results show the effectiveness of the proposed SAGN. Our source codes are available at https://github.com/TangXu-Group/SAGN.
Yuqun Yang, Xu Tang 0004, Yiu-Ming Cheung, Xiangrong Zhang, Licheng Jiao
IEEE Trans. Image Process.4
2023 Detecting and Tracking of Multiple Mice Using Part Proposal Networks
abstract
The study of mouse social behaviors has been increasingly undertaken in neuroscience research. However, automated quantification of mouse behaviors from the videos of interacting mice is still a challenging problem, where object tracking plays a key role in locating mice in their living spaces. Artificial markers are often applied for multiple mice tracking, which are intrusive and consequently interfere with the movements of mice in a dynamic environment. In this article, we propose a novel method to continuously track several mice and individual parts without requiring any specific tagging. First, we propose an efficient and robust deep-learning-based mouse part detection scheme to generate part candidates. Subsequently, we propose a novel Bayesian-inference integer linear programming (BILP) model that jointly assigns the part candidates to individual targets with necessary geometric constraints while establishing pair-wise association between the detected parts. There is no publicly available dataset in the research community that provides a quantitative test bed for part detection and tracking of multiple mice, and we here introduce a new challenging Multi-Mice PartsTrack dataset that is made of complex behaviors. Finally, we evaluate our proposed approach against several baselines on our new datasets, where the results show that our method outperforms the other state-of-the-art approaches in terms of accuracy. We also demonstrate the generalization ability of the proposed approach on tracking zebra and locust.
Zheheng Jiang, Long Chen 0019, Xiangrong Zhang, Xiangyuan Lan, Danny Crookes, Ming-Hsuan Yang 0001, Huiyu Zhou 0001
IEEE Trans. Neural Networks Learn. Syst.5
2022 Semantic-Aware Context Modeling for Road Extraction in Remote Sensing Images
abstract
Road extraction faces the great challenges of occlusion, large span, and complex backgrounds in remote sensing images. Many existing methods receive context from regions near the road non-differently, and the context from irrelevant regions instead harms the semantics of features and leads to the mis-classification of the network. To address the above problem, we propose a Semantic-Aware Context Module (SACM) that encourages the network to model the context of different se-mantics supervised by a soft foreground map. And Strip Pooling Module (SPM) is introduced to match the fact that roads tend to be strip-shaped, contributing to the suppression of contamination information in irrelevant regions. Both SACM and SPM enable the network to obtain more specific semanti-cally relevant context. The experimental results on the Deep-Globe dataset show that the proposed method tremendously improves the performance of the network.
Xiangrong Zhang, Xiaoqian Zhu, Peng Zhu 0004, Xu Tang 0004, Licheng Jiao
IGARSS2
2022 NQ-Protonet: Noisy Query Prototypical Network for Few-Shot Remote Sensing Scene Classification
abstract
Few-shot remote sensing scene classification, which aims to recognize unseen classes given only a few labeled samples, is a challenge task due to the complex content contained in remote sensing scenes. In this end, we propose a noisy query prototypical network (NQ-ProtoNet), which uses query-mix module (QM) to produce extra query samples with inter-ference information for classification and thus implicitly enhance the feature learning ability of model. Our method alleviates the problem of large intraclass variances and inter-class similarity of remote sensing scenes to some extent, and the positive experimental results on UC Merced and NWPU data sets show that it outperforms several few-shot learning methods.
Weiquan Lin, Xu Tang 0004, Xiao Han 0012, Jingjing Ma 0001, Xiangrong Zhang, Licheng Jiao
IGARSS5
2022 Mask Decoupled Head for Instance Segmentation in Remote Sensing Images
abstract
Instance segmentation predicts the categories of all instances and locates them using pixel-level masks. Although existing methods have shown exemplary performance, the poor boundaries due to the lack of fine-grained information re-mains a challenge for RSIs instance segmentation. In this paper, to address the problem, we propose a novel instance segmentation branch, namely Mask Decoupled Head, which is mainly composed of a Feature Enhance Module (FEM) and a Feature Decoupled Module (FDM). FEM enhances the rep-resentation of the body features through the low-frequency component of images. FDM decouples the segmentation task by supervising body and edge separately and leverages fine-grained information to complement the boundary details. We performed comprehensive experiments on NWPU VHR -10 and HRSID datasets to evaluate the effectiveness of our pro-posed method and achieved good performance.
Xiangrong Zhang, Tianyang Zhang 0002, Xiaoqian Zhu, Xu Tang 0004, Licheng Jiao
IGARSS2
2022 Hyperspectral Image Classification Based on Spectrally-Enhanced and Densely Connected Transformer Model
abstract
Deep learning methods have been widely used in hyperspectral image classification. Most existing CNN-based deep learning methods can extract spatial features by the local receptive field. However, these methods do not pay enough attention to the long-term dependence on HSI data and cannot capture sequence attributes. To solve this problem, we design spectrally-enhanced and densely connected transformer model (SEDT). Due to the ability of transformers in obtaining global dependencies, dense connections are added on this basis to fuse the features from shallow layers to deep layers. Furthermore, a spectrally-enhanced module is constructed to enhance the network's ability to capture local contextual and semantic features. Extensive experiments show that the proposed method exhibits highly competitive classification performance on hyperspectral image datasets.
Yongen Wu, Jie Feng 0003, Gaiqin Bai, Qiyang Gao, Xiangrong Zhang
IGARSS5
2022 Few-Shot Hyperspectral Image Classification Based on Domain Adaptation of Class Balance
abstract
Hyperspectral image (HSI) classification has attracted ever-rising attention to better performance based on limited labeled data. In this paper, a domain adaptation method of class balance based on few-shot learning is proposed, which obtains the classification results of target HSI by training the dataset in the source domain containing sufficient labeled data. We use a random weighted sampling strategy in the source domain and the generative adversarial network (GAN) in the target domain to reduce the label distribution shift caused by unbalanced classes. Then, the conditional maximum mean discrepancy (CMMD) is presented for a more comprehensive domain alignment by considering the posterior data distribution. In addition, the double cross non-local block and multi-scale strategy are adopted in the feature extraction stage to get a refined classification result. Experimental results on public HSI datasets demonstrate that our method is efficient and outperforms other baselines.
Qi Zhen, Xiangrong Zhang, Biao Hou, Xu Tang 0004, Licheng Jiao
IGARSS2
2022 Spectral Constrained Residual Attention Network for Hyperspectral Pansharpening
abstract
Deep learning methods have been widely used in the task of hyperspectral pansharpening. However, most of these methods regard the Panchromatic (PAN) image as a kind of auxiliary information, which is mainly used as spatial details to add on the hyperspectral image (HSI) after processing. Obviously, this kind of methods utilize the PAN image insufficiently, resulting in the imbalance of spatial preservation and spatial preservation. In this paper, a spectral constrained residual attention network (SCRAN) is proposed by using the PAN image as the foundation of the pansharpening task and concerning on the spectral and spatial learning. The proposed SCRAN method consists of three parts: a spectral feature extraction net, an attention spatial residual net and a spectral reconstruction net. A spectral constrained loss function is designed to enhance the spectral learning ability of SCRAN. Additionally, in SCRAN, a deep back-projection network (DBPN) is operated to upsample the HSI, and the histogram matching is applied to the PAN image to make it closer to the HSI in terms of spectral bands.
Ziyu Zhou 0009, Jie Feng 0003, Xiande Wu, Jiao Shi, Xiangrong Zhang
IGARSS5
2022 Absolute Wrong Makes Better: Boosting Weakly Supervised Object Detection via Negative Deterministic Information
abstract
Weakly supervised object detection (WSOD) is a challenging task, in which image-level labels (e.g., categories of the instances in the whole image) are used to train an object detector. Many existing methods follow the standard multiple instance learning (MIL) paradigm and have achieved promising performance. However, the lack of deterministic information leads to part domination and missing instances. To address these issues, this paper focuses on identifying and fully exploiting the deterministic information in WSOD. We discover that negative instances (i.e. absolutely wrong instances), ignored in most of the previous studies, normally contain valuable deterministic information. Based on this observation, we here propose a negative deterministic information (NDI) based method for improving WSOD, namely NDI-WSOD. Specifically, our method consists of two stages: NDI collecting and exploiting. In the collecting stage, we design several processes to identify and distill the NDI from negative instances online. In the exploiting stage, we utilize the extracted NDI to construct a novel negative contrastive learning mechanism and a negative guided instance selection strategy for dealing with the issues of part domination and missing instances, respectively. Experimental results on several public benchmarks including VOC 2007, VOC 2012 and MS COCO show that our method achieves satisfactory performance.
Guanchun Wang, Xiangrong Zhang, Zelin Peng, Xu Tang 0004, Huiyu Zhou 0001, Licheng Jiao
IJCAI2
2022 CMNet: Classification-oriented multi-task network for hyperspectral pansharpening
Xiande Wu, Jie Feng 0003, Ronghua Shang, Xiangrong Zhang, Licheng Jiao
Knowl. Based Syst.4
2022 Background Representation Learning With Structural Constraint for Hyperspectral Anomaly Detection
abstract
Hyperspectral anomaly detection is a popular topic in remote sensing image intelligent interpretation. To detect anomaly, many methods for background representation have been proposed. However, the prior information of background and anomaly is not fully explored in these methods. To tackle this issue, we combine low-rank dictionary learning (LRDL) with total variation (TV) constraint for hyperspectral anomaly detection. To be specific, the LRDL is introduced for background representation to explore the low-rank priori of background. Considering the smooth structural characteristic of background in spatial, we introduce the TV constraint on coefficients matrix for better background representation learning. Then the residual part is used to discriminate anomaly. The experiments on three real data sets demonstrate the effectiveness of the proposed method compared with state-of-the-art algorithms in hyperspectral anomaly detection.
Xiaoxiao Ma 0003, Xiangrong Zhang, Ning Huyan, Xu Tang 0004, Licheng Jiao
IEEE Geosci. Remote. Sens. Lett.2
2022 Context Residual Attention Network for Remote Sensing Scene Classification
abstract
Remote sensing image scene classification has attracted much attention due to its wide application. In this letter, a new end-to-end attentive context network under the guidance of human visual system has been proposed. The network can focus on some critical regions of the image selectively and then extract high-level feature information so as to generalize the whole image. The contributions of this letter are as follows: 1) a novel attention structure which expresses the attention of image by stacking attention modules is designed and 2) the context information of the image is used to make a holistic analysis of the attention mechanism on the basis of the top–down feedforward structure. Experiments performed on several datasets demonstrate that the proposed framework can obtain outstanding performance compared with the state-of-the-art approaches.
Yuezhu Xu, Peiyuan Jiao, Xiangrong Zhang, Huanyu Cui
IEEE Geosci. Remote. Sens. Lett.5
2022 Graph label prediction based on local structure characteristics representation
Jingyi Ding, Ruohui Cheng, Jian Song 0003, Xiangrong Zhang, Licheng Jiao, Jianshe Wu
Pattern Recognit.4
2022 Dynamic Immunization Node Model for Complex Networks Based on Community Structure and Threshold
abstract
In the information age of big data, and increasingly large and complex networks, there is a growing challenge of understanding how best to restrain the spread of harmful information, for example, a computer virus. Establishing models of propagation and node immunity are important parts of this problem. In this article, a dynamic node immune model, based on the community structure and threshold (NICT), is proposed. First, a network model is established, which regards nodes carrying harmful information as new nodes in the network. The method of establishing the edge between the new node and the original node can be changed according to the needs of different networks. The propagation probability between nodes is determined by using community structure information and a similarity function between nodes. Second, an improved immune gain, based on the propagation probability of the community structure and node similarity, is proposed. The improved immune gain value is calculated for neighbors of the infected node at each time step, and the node is immunized according to the hand-coded parameter: immune threshold. This can effectively prevent invalid or insufficient immunization at each time step. Finally, an evaluation index, considering both the number of immune nodes and the number of infected nodes at each time step, is proposed. The immune effect of nodes can be evaluated more effectively. The results of network immunization experiments, on eight real networks, suggest that the proposed method can deliver better network immunization than several other well-known methods from the literature.
Ronghua Shang, Licheng Jiao, Xiangrong Zhang, Rustam Stolkin
IEEE Trans. Cybern.4
2022 Semantic Attention and Scale Complementary Network for Instance Segmentation in Remote Sensing Images
abstract
In this article, we focus on the challenging multicategory instance segmentation problem in remote sensing images (RSIs), which aims at predicting the categories of all instances and localizing them with pixel-level masks. Although many landmark frameworks have demonstrated promising performance in instance segmentation, the complexity in the background and scale variability instances still remain challenging, for instance, segmentation of RSIs. To address the above problems, we propose an end-to-end multicategory instance segmentation model, namely, the semantic attention (SEA) and scale complementary network, which mainly consists of a SEA module and a scale complementary mask branch (SCMB). The SEA module contains a simple fully convolutional semantic segmentation branch with extra supervision to strengthen the activation of interest instances on the feature map and reduce the background noise's interference. To handle the undersegmentation of geospatial instances with large varying scales, we design the SCMB that extends the original single mask branch to trident mask branches and introduces complementary mask supervision at different scales to sufficiently leverage the multiscale information. We conduct comprehensive experiments to evaluate the effectiveness of our proposed method on the iSAID dataset and the NWPU Instance Segmentation dataset and achieve promising performance.
Tianyang Zhang 0002, Xiangrong Zhang, Peng Zhu 0004, Xu Tang 0004, Chen Li 0011, Licheng Jiao, Huiyu Zhou 0001
IEEE Trans. Cybern.2
2022 Deep Reinforcement Learning for Semisupervised Hyperspectral Band Selection
abstract
Band selection is an important step in efficient processing of hyperspectral images (HSIs), which can be seen as the combination of powerful band search technique and effective evaluation criterion. The existing deep-learning-based methods make the network parameters sparse to search the spectral bands using threshold-based functions or regularization terms. These methods may lead to an intractable optimization problem. Furthermore, these methods need to repeatedly train deep networks for evaluating candidate band subsets. In this article, we formalize hyperspectral band selection as a reinforcement learning (RL) problem. Band search is regarded as a sequential decision-making process, where each state in the search space is a feasible band subset. To evaluate each state, a semisupervised convolutional neural network (CNN), called EvaluateNet, is constructed by adding the intraclass compactness constraint of both limited labeled and sufficient unlabeled samples. A simple stochastic band sampling method is designed to train EvaluateNet, making it possible to efficiently evaluate without any fine-tuning. In RL, new reward functions are defined by taking the EvaluateNet and the penalty of repeated selection into account. Finally, advantage actor–critic algorithms are designed to explore in the state space and select the band subset according to the expected accumulated reward. The experimental results on HSI data sets demonstrate the effectiveness and efficiency of the proposed algorithms for hyperspectral band selection.
Jie Feng 0003, Xianghai Cao, Ronghua Shang, Xiangrong Zhang, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.6
2022 Self-Supervised Divide-and-Conquer Generative Adversarial Network for Classification of Hyperspectral Images
abstract
Generative adversarial network (GAN) has been rapidly developed because of its powerful generating ability. However, imbalanced class distribution of hyperspectral images (HSIs) easily causes mode collapse in GAN. Moreover, limited training samples in HSIs restrict the generating ability of GAN. These issues may further deteriorate the classification performance of the discriminator. To conquer these issues, a novel self-supervised divide-and-conquer GAN (SDC-GAN) is proposed for HSI classification. In SDC-GAN, a pretext cluster task with an encoder-decoder architecture is designed by leveraging abundant unlabeled samples. By transferring the learned cluster representation from the cluster task, limited labeled samples are divided effectively in the downstream classification. According to the division of clustering, SDC-GAN constructs a generic and several specific branches for both the generator and discriminator. The generator generates all-class and specific-class samples by using the generic and specific branches separately and combines them adaptively. It can weaken the generation preference for the classes with large sample sizes and alleviate the mode collapse problem. Meanwhile, the classification ability of the discriminator is improved by integrating the judgment of specific branches into the generic branch. Experimental results show that SDC-GAN achieves competitive results for HSI classification compared with several state-of-the-art methods.
Jie Feng 0003, Ronghua Shang, Xiangrong Zhang, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.4
2022 Multitemporal Intrinsic Image Decomposition With Temporal-Spatial Energy Constraints for Remote Sensing Image Analysis
abstract
Due to interference with remote imaging by some natural factors, the multitemporal analysis ability is limited by the spectral drift between images. In this article, a new approach to optimize the existing multitemporal analysis system is proposed: multitemporal intrinsic image decomposition (MIID). The MIID method is designed to extract common spectral reflectance from multitemporal images. With MIID, multitemporal classification, changing detection, and index extracting will become extremely easy and more accurate. Firstly, without considering land cover change, the general MIID framework is proposed by adding local temporal–spatial energy constraints in traditional intrinsic images decomposition. On this basis, an improved MIID method with change detection (CD) (CD-MIID) capability is proposed to make the model adapt to the land cover change situation. Finally, specific steps of how to use MIID methods in the multitemporal analysis are given. Multitemporal multispectral/hyperspectral remote sensing images from GF-1, GF-2, GF-5, Landsat TM, and two groups of captured datasets with reflectance truth map are used to evaluate the performance. The experimental results show the following two points: first, the MIID methods achieve better extraction results of spectral reflectance. Second, the proposed MIID methods have better performance both on multitemporal classification and CD.
Guoming Gao, Baisen Liu, Xiangrong Zhang, Xudong Jin, Yanfeng Gu
IEEE Trans. Geosci. Remote. Sens.3
2022 Cluster-Memory Augmented Deep Autoencoder via Optimal Transportation for Hyperspectral Anomaly Detection
abstract
Hyperspectral anomaly detection aims to detect objects significantly different from their surrounding background. Recently, many detectors based on autoencoder (AE) exhibited promising performances in hyperspectral anomaly detection tasks. However, the fundamental hypothesis of the AE-based detector that anomaly is more challenging to be reconstructed than background may not always be true in practice. We demonstrate that an autoencoder could well reconstruct anomalies even without anomalies for training. Because AE models mainly focus on the quality of sample reconstruction and do not care if the encoded features solely represent the background rather than anomalies. If more information is preserved than needed to reconstruct the background, the anomalies will be well reconstructed. This paper proposes a cluster-memory augmented autoencoder via deep optimal transportation clustering (OTCMA) for hyperspectral anomaly detection to solve this problem. The deep clustering method based on optimal transportation is proposed to enhance the features consistency of samples within the same categories and features discrimination of samples in different categories. The memory module stores the background’s consistent features, which are the cluster centers for each category background. We retrieve more consistent features from the memory module instead of reconstructing a sample utilizing its own encoded features. The network focuses more on consistent feature reconstruction by training AE with a memory module. This effectively restricts the reconstruction ability of AE and prevents reconstructing anomalies. Extensive experiments on the benchmark datasets demonstrate that our proposed OTCMA achieves state-of-the-art results. Besides, this paper presents further discussions about the effectiveness of our proposed memory module and different criterion for better anomaly detection.
Ning Huyan, Xiangrong Zhang, Dou Quan, Jocelyn Chanussot, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.2
2022 Recurrent Attention and Semantic Gate for Remote Sensing Image Captioning
abstract
The remote sensing image captioning has attracted wide spread attention in remote sensing field due to its application potentiality. However, most existing approaches model limited interactions between image content and sentence and fail to exploit special characteristics of the remote sensing images. We introduce a novel recurrent attention and semantic gate (RASG) framework to facilitate the remote sensing image captioning in this article, which integrates competitive visual features and a recurrent attention mechanism to generate a better context vector for the images every time as well as enhances the representations of the current word state. Specifically, we first project each image into competitive visual features by taking the advantage of both static visual features and multiscale features. Then, a novel recurrent attention mechanism is developed to extract the high-level attentive maps from encoded features and nonvisual features, which can help the decoder recognize and focus on the effective information for understanding the complex content of the remote sensing images. Finally, the hidden states from the long short-term memory (LSTM) and other semantic references are incorporated into a semantic gate, which contributes to more comprehensive and precise semantic understanding. Comprehensive experiments on three widely used datasets, Sydney-Captions, UCM-Captions, and Remote Sensing Image Captioning Dataset, have demonstrated the superiority of the proposed RASG over a series of attentive models based on image captioning methods.
Yunpeng Li 0010, Xiangrong Zhang, Chen Li 0011, Xin Wang 0068, Xu Tang 0004, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.2
2022 Adaptive Graph Convolutional Network for PolSAR Image Classification
abstract
Polarimetric synthetic aperture radar (PolSAR) image classification is one of the hottest issues in remote sensing, where studies on pixel-level information and relationship are of great significance. In this article, graph convolutional network (GCN) is employed to accomplish this pixel-level task benefiting from its excellent capability in structure exploration and information propagation between different pixels. To reduce the communication burden between various PolSAR pixels and high computational cost for the whole PolSAR image, an adaptive GCN (AdapGCN) consisting of pixel-centered subgraphs is proposed in this article. In the AdapGCN, a data-adaptive kernel and a spatial-adaptive kernel are introduced to, respectively, model data structure and spatial structure for PolSAR image. Moreover, a multiscale learning structure is integrated to further explore complicated relations between pixels. Extensive comparative evaluations validate the superiority of our new AdapGCN model for PolSAR image classification over a wide range of state-of-the-art methods on three challenging benchmarks.
Fang Liu 0034, Jingya Wang 0001, Xu Tang 0004, Jia Liu 0020, Xiangrong Zhang, Liang Xiao 0001
IEEE Trans. Geosci. Remote. Sens.5
2022 A Hybrid Network With Structural Constraints for SAR Image Scene Classification
abstract
Data-based image classification methods, such as convolutional neural networks (CNNs), have achieved state-of-the-art performance. They usually leverage thousands of labeled samples to train the networks but ignore some prior knowledge. However, labeled samples are difficult to be obtained for synthetic aperture radar (SAR) images. Model-based methods are adept at utilizing the prior information of data, while they have to introduce some restrictions or assumptions during the realization of models. Consequently, to develop the advantages of both methods and improve their disadvantages, we propose a hybrid network by coupling the data-based with model-based methods for SAR image scene classification in this article. First, to fully use the prior information of SAR images and large amounts of unlabeled samples, we improve the$G^{0}$-based variational Bayesian inference model (GVBI) and construct a$G^{0}$-based convolutional variational auto-encoder (GCVAE) for unsupervised learning of the distributional characteristics of SAR images. After that, we further extend the GCVAE by combining it with CNN, resulting in a stronger hybrid network to classify SAR images with a few labeled samples. In addition, considering the abundant structural information is crucial for SAR image classification, we design a sketch fitter and two structural constraints on both pixel and sketch spaces to assist the hybrid network to improve its classification performance. Finally, we evaluate the performance of our method on real-SAR images, and the experimental results demonstrate that the proposed framework outperforms related methods on classification while reducing the manual annotation substantially.
Xiaoxue Qian, Fang Liu 0001, Licheng Jiao, Xiangrong Zhang, Puhua Chen, Lingling Li 0002, Yuanhao Cui
IEEE Trans. Geosci. Remote. Sens.4
2022 Class-Level Prototype Guided Multiscale Feature Learning for Remote Sensing Scene Classification With Limited Labels
abstract
Remote sensing scene classification (RSSC) is an open and challenging research topic in the remote sensing (RS) community. It aims to define semantic labels for RS scenes according to their contents. Recently, with the development of deep convolutional neural networks (DCNNs), the results of RSSC have been enhanced to a large extent. However, the cracking performance of these DCNN-based models depends on a large number of labeled data. Once the volume of the labeled data is decreased, their behavior would be weakened dramatically. In this article, we propose a new training algorithm that can work smoothly with a few labeled samples to address this limitation. Along with the introduced DCNN, the presented methods perform satisfactorily. In particular, we first construct a dual-branch network (DBNet) to mine the multiscale and multiangle information from RS scenes. Thus, the abundant land covers with diverse sizes, directions, and shapes can be captured simultaneously. Then, to train DBNet using scarce semantic labels, a class-level prototype guided learning (CPGL) algorithm is developed based on the meta-learning paradigm. Besides the usual episode training manner, a prototype refinement module (PRM) and a prototype discrimination module (PDM) are designed with the help of metric learning theory to ensure the effectiveness of our CPGL. The comprehensive experiments are conducted on four public RS scene datasets, and the encouraging results imply that our DBNet and CPGL can copy with RSSC tasks with small labeled data. Our source codes are available athttps://github.com/TangXu-Group/Remote-Sensing-Images-Classification/tree/main/CPGL.
Xu Tang 0004, Weiquan Lin, Jingjing Ma 0001, Xiangrong Zhang, Fang Liu 0034, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.4
2022 EMTCAL: Efficient Multiscale Transformer and Cross-Level Attention Learning for Remote Sensing Scene Classification
abstract
In recent years, convolutional neural network (CNN)-based methods have been widely used for remote sensing (RS) scene classification tasks and achieved excellent results. However, CNNs are not good at exploring contextual information, which is essential for fully understanding RS scenes. A new model named transformer attracts researchers’ attention to address this problem, which is skilled in mining the latent contextual information in RS scenes. Nevertheless, since the contents of RS scenes are diverse in type and various in scale, the performance of the original transformer in RS scene classification cannot reach what we expect. In addition, due to the specific self-attention mechanism, the time costs of the transformer are high, which hinders its practicability in the RS community. To overcome the above limitations, we propose a new model named efficient multi-scale transformer and cross-level attention learning (EMTCAL) for RS scene classification in this paper. EMTCAL combines the advantages of CNN and transformer to mine information within RS scenes fully. First, it uses a multi-layer feature extraction module (MFEM) to acquire global visual features and multi-level convolutional features from RS scenes. Second, a contextual information extraction module (CIEM) is proposed to capture rich contextual information from multi-level features. In CIEM, taking the characteristics of RS scenes and the computational complexity into account, we propose an efficient multi-scale transformer (EMST). EMST can mine the abundant knowledge with various scales hidden in RS scenes and model their inherent relations at small-time costs. Third, a cross-level attention module (CLAM) is developed to aggregate and explore correlations of multi-level features. Finally, a class score fusion module (CSFM) is designed to integrate the contributions of global and aggregated multi-level features for the discriminative scene representations. Extensive experiments are conducted on three public RS scene data sets. The positive results demonstrate that our EMTCAL can achieve superior classification performance and outperform many state-of-the-art methods.
Xu Tang 0004, Mingteng Li, Jingjing Ma 0001, Xiangrong Zhang, Fang Liu 0001, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.4
2022 Meta-Hashing for Remote Sensing Image Retrieval
abstract
With the explosive growth of the volume and resolution of high-resolution remote-sensing (HRRS) images, the management of them becomes a challenging task. The traditional content-based remote-sensing image retrieval (CBRSIR) technologies cannot meet what we expect due to the large volume of image archives and complex contents within HRRS images. As a successful approximate nearest neighborhood (ANN) search technique, Hash learning has received wide attention, especially when deep convolutional neural networks (DCNNs) appear. Due to DCNNs’ strong capacity of feature learning, many DCNN-based hashing methods have been proposed and achieved good performance for large-scale CBRSIR tasks. Nevertheless, their limitation is that a large of labeled training samples should be collected for training the deep models. To overcome this limitation, this article, therefore, develops a new supervised hash learning method for the large-scale HRRS CBRSIR task based on meta-learning, which could achieve well-retrieval performance with a few labeled training samples. First, taking the characteristics of HRRS into account, we develop a self-adaptive convolution (SAP-Conv) block and design a hashing net based on the block. SAP-Conv can learn robust features from HRRS images by exploring their multiscale information. Second, to enhance the generalization of the hashing net under a few labeled training samples, the hash learning is formulated in a meta-way, and we name it meta-hashing. Meta-hashing can effectively preserve the similarities between support and query set, and the similarities between samples within support set by the developed loss function. To further improve the performance of meta-hashing, we expand it to a dynamic version named dynamic-meta-hashing, in which the numbers of support and query are changeable in the training phase. Experimental results counted on the three widely used HRRS datasets demonstrate our dynamic-meta-hashing and meta-hashing can achieve promising performance in large-scale HRRS CBRSIR tasks based on a few training samples. Our source codes are available athttps://github.com/TangXu-Group/Meta-hashing.
Xu Tang 0004, Yuqun Yang, Jingjing Ma 0001, Yiu-Ming Cheung, Chao Liu 0042, Fang Liu 0034, Xiangrong Zhang, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.7
2022 An Unsupervised Remote Sensing Change Detection Method Based on Multiscale Graph Convolutional Network and Metric Learning
abstract
As a fundamental application, change detection (CD) is widespread in the remote sensing (RS) community. With the increase in the spatial resolution of RS images, high-resolution remote sensing (HRRS) image CD tasks receive growing attention. The change information hidden in multitemporal HRRS images could help discover our planet comprehensively. In the current deep learning era, convolutional neural networks (CNNs) have become one of the most powerful tools for a wide range of RS tasks including HRRS image CD, due to their superb feature learning capacity. However, most of them need a large amount of labeled data to accomplish the CD process, which is challenging or even impractical in many RS applications. Also, given the limited valid receptive field, CNNs can only capture short-range context within HRRS images, which is probably not enough to fully explore change information from the images. To overcome these limitations, in this article, we propose an unsupervised CD method, termed GMCD, based on graph convolutional network (GCN) and metric learning. GMCD consists of a Siamese fully convolution network (FCN), a multiscale dynamic GCN (Mlt-GCN), and a pseudolabel generation mechanism based on metric learning. The Siamese FCN contains a Siamese encoder and a pyramid-shaped decoder, aiming to extract multiscale features and integrate them to generate reliable difference images (DIs). Mlt-GCN focuses on capturing the short- and long-range contextual patterns at feature map level to extract changed and unchanged areas completely. The pseudolabel generation mechanism aims to produce reliable pseudolabels (changed, unchanged, and uncertain) to help accomplish the model training in an unsupervised way. Experiments on four HRRS image CD datasets demonstrate that GMCD outperforms the existing state-of-the-art methods.
Xu Tang 0004, Lichao Mou, Fang Liu 0034, Xiangrong Zhang, Xiao Xiang Zhu 0001, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.5
2022 Dual-Channel Capsule Generation Adversarial Network for Hyperspectral Image Classification
abstract
Deep learning-based methods have demonstrated significant breakthroughs in the application of hyperspectral image (HSI) classification. However, some challenging issues still exist, such as the overfitting problem caused by the limitation of training size with high-dimensional feature and the efficiency of spectral–spatial (SS) exploitation. Therefore, to efficiently model the relative position of samples within the generative adversarial network (GAN) setting, we proposed a dual-channel SS fusion capsule generative adversarial network (DcCapsGAN) for HSI classification. Dual channels (1-D-CapsGAN and 2-D-CapsGAN) are constructed by integrating the capsule network (CapsNet) with GAN for eliminating the mode collapse and gradient disappearance problem caused by traditional GAN. Meanwhile, octave convolution and multiscale convolution are integrated into the proposed model for further reducing the parameters of the CapsNet and extracting multiscale features. To further boost the classification performance, the SS channel fusion model is constructed to composite and switch the feature information of different channels, thereby facilitating the accuracy and robustness of the whole classification performance. Three commonly used HSI data sets are utilized to investigate the performance of the proposed DcCapsGAN model, and the performance of the experiment demonstrates that the proposed model can efficiently improve the classification accuracy and performance.
Jianing Wang 0003, Siying Guo, Runhu Huang, Linhao Li, Xiangrong Zhang, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.5
2022 Multiobjective Guided Divide-and-Conquer Network for Hyperspectral Pansharpening
abstract
Deep learning methods have gained rapid development in hyperspectral pansharpening (HP) due to powerful spatial–spectral feature extraction ability. However, most of these methods are optimized using a single reconstruction objective. It is difficult for these methods to find a balance between spectral preservation and spatial preservation. Furthermore, these methods adopt interpolation or convolution to upsample the hyperspectral images (HSIs), which tends to cause noticeable spectral distortion. To conquer these issues, a novel multiobjective guided divide-and-conquer network (MO-DCN) is proposed for HP. It consists of a deconvolution long short-term memories (LSTMs) network (DLSTM) and a divide-and-conquer network (DCN). DLSTM leverages bi-direction learning to upsample HSIs by considering 3-D spatiotemporal dependencies. Then, DCN designs a two-branch architecture to reconstruct spatial and spectral information from upsampled HSIs and panchromatic images (PANIs), respectively, where the spatial branch designs an attention-in-attention module (AIAM) to emphasize complementary attention in a coarse-to-fine way. Finally, co-improvement of spatial and spectral information is formulated as an Epsilon-constraint-based multiobjective optimization. The Epsilon constraint method transforms one objective into a constraint and regards it as a penalty bound to make an excellent tradeoff between different objectives. Experimental results demonstrated that the proposed method markedly improves pansharpening performance in both the spatial and spectral domains and has superior fusion performance than state-of-the-art methods.
Xiande Wu, Jie Feng 0003, Ronghua Shang, Xiangrong Zhang, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.4
2022 AR2Det: An Accurate and Real-Time Rotational One-Stage Ship Detector in Remote Sensing Images
abstract
Ship detection plays a significant role in the high-resolution remote sensing (HRRS) community, but it is a challenging task due to the complex contents within HRRS images and the diverse orientation of ships. Recently, with the development of deep learning, the performance of the HRRS ship detection model has been improved greatly. Most of them employ deep networks and complicate anchor mechanism to get well ship detection results. Nevertheless, this kind of combination limits the detection efficiency. To address this problem, a new approach named accurate and real-time rotational ship detector (AR2Det) is proposed in this article to detect ships without the anchor mechanism. Based on the extracted features by the feature extraction module (FEM) and the central information of ships, AR2Det adopts two simple modules, ship detector (SDet) and center detector (CDet), to generate and improve the detection results, respectively. AR2Det is efficient due to the simple postprocessing and the lightweight network. Also, AR2Det performs satisfactorily due to the effective generation and enhancement strategy of bounding boxes. The extensive experiments are conducted on a public HRRS image ship detection dataset HRSC2016. The promising results show that our method outperforms the state-of-the-art approaches in terms of both accuracy and speed.
Yuqun Yang, Xu Tang 0004, Yiu-Ming Cheung, Xiangrong Zhang, Fang Liu 0034, Jingjing Ma 0001, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.4
2022 Spatial Pooling Graph Convolutional Network for Hyperspectral Image Classification
abstract
Graph convolution networks (GCNs) have been applied in a variety of fields due to their powerful ability in processing graph-like data. However, the massive number of hyperspectral pixels makes it challenging to define general graph structures on hyperspectral images (HSIs). On the other hand, convolutional neural networks (CNNs) take in regular image regions with fixed square size, and have demonstrated impressive accuracy while being efficient in computation. Inspired by the classification framework of CNNs, we develop a GCN-based model that generates effective local spectral–spatial features for HSI classification. Specifically, graph convolutions are performed separately on every local region, which significantly limits the graph’s size. While graph convolution extracts features of every pixel, it does not reduce the number of them. To fuse suitable representations for the classification task, we develop a graph pooling operation to preserve classification-specific features and reduce redundant pixels. Based on local regions of HSIs, pooling in the graph domain is equivalent to spatial pooling in the spatial domain. The proposed method is thus named the spatial pooling graph convolutional network (SPGCN). Experimental results on several typical datasets demonstrated that the proposed SPGCN provides competitive results compared with other state-of-the-art CNN-based methods.
Xiangrong Zhang, Peng Zhu 0004, Xu Tang 0004, Jie Feng 0003, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.1
2022 Remote Sensing Image Super-Resolution via Dual-Resolution Network Based on Connected Attention Mechanism
abstract
Limited by hardware conditions and complex degradation processes, the obtained remote sensing images (RSIs) are often low-resolution (LR) data with insufficient high-frequency information. Image super-resolution (SR) aims to improve the spatial resolution of images and add reasonable detailed information. Although existing convolutional neural network (CNN)-based methods achieve good performance by adding residual structure and attention mechanism to the network, simply stacking the residual structure and embedding the attention module directly on the residual branch lead to localized use of features and information loss. To address the above problems, we propose a dual-resolution connected attention network (DRCAN). Specifically, a high-resolution (HR) learning branch is constructed to complement the mapping learning between LR images and HR images, and a connected attention module with residual learning is introduced to make full use of the different levels of intermediate layer features. Besides, we collect data at different resolutions from Google Earth to form a dataset named XD IPIU for RSIs SR. Extensive experiments demonstrate the effectiveness of the proposed model and DRCAN shows the state-of-the-art performance in terms of quantitative evaluation and visual quality.
Xiangrong Zhang, Tianyang Zhang 0002, Fengsheng Liu, Xu Tang 0004, Puhua Chen, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.1
2022 Spectral Partitioning Residual Network With Spatial Attention Mechanism for Hyperspectral Image Classification
abstract
Hyperspectral image (HSI) classification is one of the most important tasks in hyperspectral data analysis. Convolutional neural networks (CNN) have been introduced to HSI classification and achieved good performance. In this article, an effective and efficient CNN-based spectral partitioning residual network (SPRN) is proposed for HSI classification. The SPRN splits the input spectral bands into several nonoverlapping continuous subbands and uses cascaded parallel improved residual blocks to extract spectral–spatial features from these subbands, respectively. Finally, the features are fused and fed into a classifier. By equivalently using grouped convolutions, the spectral partition and feature extraction are embedded into an end-to-end network. Experimental results show that the proposed SPRN achieves state-of-the-art performance, meanwhile, with relatively fewer parameters and computational costs. Usually, the CNN takes a patch that contains continuous spatial information as the input and results in a class label of the center pixel. The large size of the input patch includes more spatial information, whereas also introduces interfering pixels that may lead to a degradation of classification accuracies. For that reason, we propose a novel spatial attention module named homogeneous pixel detection module (HPDM). The module alleviates the degradation of performance as the input patch size increases by capturing the homogeneous pixels in the input patch. The module can be integrated into any CNN-based HSI classification framework.
Xiangrong Zhang, Shouwang Shang, Xu Tang 0004, Jie Feng 0003, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.1
2022 Foreground Refinement Network for Rotated Object Detection in Remote Sensing Images
abstract
Object detection has been a fundamental task in the field of remote sensing and has made considerable progress in recent years. However, the high background complexity in remote sensing images (RSIs) remains challenging. In this article, we propose a refined rotation detector, namely, the Foreground Refinement Network (FoRDet), to alleviate the above problem by leveraging the information of foreground regions from the perspectives of feature and optimization. Specifically, we propose a foreground relation module (FRL) that aggregates the foreground-contextual representations from the coarse stage and improves the discrimination of foreground regions on feature maps in the refined stage. Besides, considering the risk of the potential foreground anchors being overwhelmed in the training phase, we design a foreground anchor reweighting (FRW) loss that integrates the classification confidence and localization accuracy of each foreground anchor from the coarse stage to dynamically regulate their contributions in the refined stage, which highlights the potential foreground anchors. The comprehensive experimental results on three public datasets for rotated object detection DOTA, HRSC2016, and UCAS-AOD demonstrate the effectiveness of our proposed method.
Tianyang Zhang 0002, Xiangrong Zhang, Peng Zhu 0004, Puhua Chen, Xu Tang 0004, Chen Li 0011, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.2
2022 Unsupervised Outlier Detection Using Memory and Contrastive Learning
abstract
Outlier detection is to separate anomalous data from inliers in the dataset. Recently, the most deep learning methods of outlier detection leverage an auxiliary reconstruction task by assuming that outliers are more difficult to recover than normal samples (inliers). However, it is not always true in deep auto-encoder (AE) based models. The auto-encoder based detectors may recover certain outliers even if outliers are not in the training data, because they do not constrain the feature learning. Instead, we think outlier detection can be done in the feature space by measuring the distance between outliers' features and the consistency feature of inliers. To achieve this, we propose an unsupervised outlier detection method using a memory module and a contrastive learning module (MCOD). The memory module constrains the consistency of features, which merely represent the normal data. The contrastive learning module learns more discriminative features, which boosts the distinction between outliers and inliers. Extensive experiments on four benchmark datasets show that our proposed MCOD performs well and outperforms eleven state-of-the-art methods.
Ning Huyan, Dou Quan, Xiangrong Zhang, Xuefeng Liang, Jocelyn Chanussot, Licheng Jiao
IEEE Trans. Image Process.3
2021 Automatic Design Recurrent Neural Network for Hyperspectral Image Classification
abstract
Hyperspectral image (HSI) classification is a hot research direction in remote sensing community. In HSIs, there are hundreds of narrow and continuous spectral bands. Due to powerful sequence data processing ability, recurrent neural network (RNN) has shown great potential in HSI classification in recent years. In order to ensure a good classification performance, how to design a proper RNN structure is of great importance. In this paper, an automatic RNN (Auto-RNN) for HSI classification is proposed. Firstly, a number of candidate modules, including ReLU, tanh, sigmoid, and identity, are provided. Then, a policy gradient-based reinforcement learning is us ed to search the feasible deep architecture. Finally, the best RNN structure is selected through evaluating on the validation set. The experimental results demonstrate that the proposed algorithm can yield promising classification performance compared with some existing methods.
Jie Feng 0003, Gaiqin Bai, Zizhuo Gao, Xiangrong Zhang, Xu Tang 0004
IGARSS4
2021 Improved SiamRPN++ with Clustering-Based Frame Differencing for Object Tracking of Remote Sensing Videos
abstract
Deep learning (DL) based object tracking methods have achieved encouraging results on natural videos. However, directly applying these DL-based methods to the vehicle tracking of optical remote sensing videos (ORSV) still faces many challenges. Different with the vehicles in nature videos, most vehicles in ORSV are blurry, small in size, and highly similar to other vehicles. Furthermore, blurred vehicles easily blend into the background and are difficult to distinguish. To solve these problems, an improved Siamrpn++ with clustering-based frame differencing (CFD-SiamRPN++) is proposed. In CFD-SiamRPN++, a clustering method is used to divide the differencing map among adjacent frames into two clusters. By using the statistics of the clustering results, each cluster is judged to the target or not. According to the judgment of each cluster, the differencing map is refined by reducing the background noises and retaining moving information of vehicles. Then, the refined differencing map is fused with the original frame as the input of the tracking network to enhance the discriminative ability of small-sized and blurred vehicles. In the tracking phase, SiamRPN++ is selected as the tracking network due to its feature extraction capability with multi-scale feature fusion and efficient tracking performance. Experiment results based on Jilin-1 ORSV dataset show that the proposed method provides a competitive tracking performance over state-of art deep learning methods.
Jie Feng 0003, Bingyu Hui, Yuping Liang, Quanhe Yao, Xiangrong Zhang
IGARSS5
2021 Cross-Source Image Retrieval Based on Ensemble Learning and Knowledge Distillation for Remote Sensing Images
abstract
As different kinds of high-resolution remote sensing (HRRS) image data sources increase, the cross-source content-based image retrieval (CS-CBRSIR) is becoming an important and urgent task to be solved. Most existing methods focus on optimizing the common space features for dual-source effectively. The source discrepancy in classifier level, however, has been ignored. To handle this problem, we propose teacher-ensemble learning with the knowledge distillation method in this paper. The ensemble of source-shared and source-specific classifiers could construct an effective teacher model. The useful information can be transferred back with the knowledge distillation. Besides, the feature pyramid network is introduced to learn the multi-scale features from HRRS images, which can describe the complex contents of HRRS images well. The positive experimental results conducted on DSRSID illustrates the effectiveness of the proposed method.
Jingjing Ma 0001, Duanpeng Shi, Xu Tang 0004, Xiangrong Zhang, Xiao Han 0012, Licheng Jiao
IGARSS4
2021 Hyperspectral Image Classification Based on Spectral Graph and Bidirectional LSTM Network
abstract
Convolutional neural networks (CNNs) have achieved cracking performance in the hyperspectral image (HSI) classification task. Nevertheless, most of them cannot meet what we expect when the numbers of labeled samples are small. Also, due to the rectangular convolution kernels, the long-range context information within HSIs is cannot fully be explored. To solve these problems, we propose a semi-supervised method based on the graph convolutional network (GCN) and bidirectional Long Short-Term Memory (Bi-LSTM). First, HSIs are over segmented into various superpixels and GCN is employed for mining the advanced spectral features. Second, we input the obtained spectral features to the Bi-LSTM model for exploring global spatial features. Due to the diverse receptive fields, the short-and long-range spatial relations can be discovered simultaneously. Finally, we map the features from region-level to pixel-level for classifying HSIs. The positive experimental results counted on two HSIs demonstrate that our method is superior to some popular methods.
Xu Tang 0004, Qionglin Zhou, Xiao Han 0012, Dalei Li, Xiangrong Zhang, Licheng Jiao
IGARSS6
2021 Remote Scene Image Scene Classification Based on Adaptive Segmentation and Dynamic Graph Convolution
abstract
As an important research topic in the remote sensing (RS) community, RS image scene classification is a challenging task due to the complex contents of RS images. In general, RS image scene classification is a single-label problem. Nevertheless, it is known that the contents within RS are huge in volume and diverse in type. Only a single semantic label cannot describe an RS scene completely, especially when the resolution of RS images is increased recently. The various semantics hidden in the high-resolution RS (HRRS) images are also important to the scene classification task. Taking the issues mentioned above into account, we develop a new scene classifier named graph scene classifier (GSCer) for HRRS images with the help of the deep convolution neural network (DCNN) and dynamic graph convolution (DGCN). Not only the global semantic but also the diverse hidden local semantics within an HRRS image can be fully explored. The encouraging experimental results counted on two public HRRS data sets demonstrate that our GSCer is effective in HRRS scene classification tasks.
Yuqun Yang, Xu Tang 0004, Xiao Han 0012, Jingjing Ma 0001, Xiangrong Zhang, Licheng Jiao
IGARSS5
2021 High-Resolution Remote Sensing Images Change Detection with Siamese Holistically-Guided FCN
abstract
Change Detection is an important and challenging task in the remote sensing (RS) field, especially when the resolution of RS images is getting higher. The appearance of the deep convolutional neural network (DCNN) provides new opportunities for the high-resolution RS (HRRS) image processing as well as the HRRS CD task. In this paper, we proposed a Siamese holistically-guided FCN (SHG-FCN) model to fully mine the low- and high-level features from HRRS images for completing the CD task. SHG-FCN consists of a Siamese encoder and a holistically-guided decoder. The Siamese encoder adopts five layers of convolution for feature extraction and generates multi-scale difference maps. The decoder employs the holistically-guided architecture, which uses the deep semantic feature as codewords to guide the up-sample of feature map and achieve multi-scale feature fusion. Our model is testified on two public HRRS datasets, and the obtained encouraging CD results illustrate that our method is effective in HRRS CD tasks.
Xu Tang 0004, Xiao Han 0012, Jingjing Ma 0001, Xiangrong Zhang, Licheng Jiao
IGARSS5
2021 Adaptive Affinity Loss and Erroneous Pseudo-Label Refinement for Weakly Supervised Semantic Segmentation
abstract
Semantic segmentation has been continuously investigated in the last ten years, and majority of the established technologies are based on supervised models. In recent years, image-level weakly supervised semantic segmentation (WSSS), including single- and multi-stage process, has attracted large attention due to data labeling efficiency. In this paper, we propose to embed affinity learning of multi-stage approaches in a single-stage model. To be specific, we introduce an adaptive affinity loss to thoroughly learn the local pairwise affinity. As such, a deep neural network is used to deliver comprehensive semantic information in the training phase, whilst improving the performance of the final prediction module. On the other hand, considering the existence of errors in the pseudo labels, we propose a novel label reassign loss to mitigate over-fitting. Extensive experiments are conducted on the PASCAL VOC 2012 dataset to evaluate the effectiveness of our proposed approach that outperforms other standard single-stage methods and achieves comparable performance against several multi-stage methods.
Xiangrong Zhang, Zelin Peng, Peng Zhu 0004, Tianyang Zhang 0002, Chen Li 0011, Huiyu Zhou 0001, Licheng Jiao
ACM Multimedia1
2021 Corrigendum to: Genetic mechanisms of COVID-19 and its association with smoking and alcohol consumption
abstract
When this paper was first published some text incorrectly appeared as ``Shuquan Ra is a professor of the Institute of Hematology & Blood Diseases Hospital, Chinese Academy of Medical Sciences & Peking Union Medical College''. This should have appeared as ``Shuquan Rao is a professor of State Key Laboratory of Experimental Hematology, National Clinical Research Center for Blood Diseases, Institute of Hematology & Blood Diseases Hospital, Chinese Academy of Medical Sciences & Peking Union Medical College.'' This has been corrected online.
Shuquan Rao, Ancha V. Baranova, Hongbao Cao, Jiu Chen, Xiangrong Zhang, Fuquan Zhang 0003
Briefings Bioinform.5
2021 Genetic mechanisms of COVID-19 and its association with smoking and alcohol consumption
abstract
We aimed to investigate the genetic mechanisms associated with coronavirus disease of 2019 (COVID-19) outcomes in the host and to evaluate the possible associations between smoking and drinking behavior and three COVID-19 outcomes: severe COVID-19, hospitalized COVID-19 and COVID-19 infection. We described the genomic loci and risk genes associated with the COVID-19 outcomes, followed by functional analyses of the risk genes. Then, a summary data-based Mendelian randomization (SMR) analysis, and a transcriptome-wide association study (TWAS) were performed for the severe COVID-19 dataset. A two-sample Mendelian randomization (MR) analysis was used to evaluate the causal associations between various measures of smoking and alcohol consumption and the COVID-19 outcomes. A total of 26 protein-coding genes, enriched in chemokine binding, cytokine binding and senescence-related functions, were associated with either severe COVID-19 or hospitalized COVID-19. The SMR and the TWAS analyses highlighted functional implications of some GWAS hits and identified seven novel genes for severe COVID-19, including CCR5, CCR5AS, IL10RB, TAC4, RMI1 and TNFSF15, some of which are targets of approved or experimental drugs. According to our studies, increasing consumption of cigarettes per day by 1 standard deviation is related to a 2.3-fold increase in susceptibility to severe COVID-19 and a 1.6-fold increase in COVID-19-induced hospitalization. Contrarily, no significant links were found between alcohol consumption or binary smoking status and COVID-19 outcomes. Our study revealed some novel COVID-19 related genes and suggested that genetic liability to smoking may quantitatively contribute to an increased risk for a severe course of COVID-19.
Shuquan Rao, Ancha V. Baranova, Hongbao Cao, Jiu Chen, Xiangrong Zhang, Fuquan Zhang 0003
Briefings Bioinform.5
2021 Dual-graph convolutional network based on band attention and sparse constraint for hyperspectral band selection
Jie Feng 0003, Zhanwei Ye, Shuai Liu 0016, Xiangrong Zhang, Jiantong Chen, Ronghua Shang, Licheng Jiao
Knowl. Based Syst.4
2021 Convolutional Neural Network Based on Bandwise-Independent Convolution and Hard Thresholding for Hyperspectral Band Selection
abstract
Band selection has been widely utilized in hyperspectral image (HSI) classification to reduce the dimensionality of HSIs. Recently, deep-learning-based band selection has become of great interest. However, existing deep-learning-based methods usually implement band selection and classification in isolation, or evaluate selected spectral bands by training the deep network repeatedly, which may lead to the loss of discriminative bands and increased computational cost. In this article, a novel convolutional neural network (CNN) based on bandwise-independent convolution and hard thresholding (BHCNN) is proposed, which combines band selection, feature extraction, and classification into an end-to-end trainable network. In BHCNN, a band selection layer is constructed by designing bandwise 1×1 convolutions, which perform for each spectral band of input HSIs independently. Then, hard thresholding is utilized to constrain the weights of convolution kernels with unselected spectral bands to zero. In this case, these weights are difficult to update. To optimize these weights, the straight-through estimator (STE) is devised by approximating the gradient. Furthermore, a novel coarse-to-fine loss calculated by full and selected spectral bands is defined to improve the interpretability of STE. In the subsequent layers of BHCNN, multiscale 3-D dilated convolutions are constructed to extract joint spatial-spectral features from HSIs with selected spectral bands. The experimental results on several HSI datasets demonstrate that the proposed method uses selected spectral bands to achieve more encouraging classification performance than current state-of-the-art band selection methods.
Jie Feng 0003, Jiantong Chen, Qigong Sun, Ronghua Shang, Xianghai Cao, Xiangrong Zhang, Licheng Jiao
IEEE Trans. Cybern.6
2021 Attention Multibranch Convolutional Neural Network for Hyperspectral Image Classification Based on Adaptive Region Search
abstract
Convolutional neural networks (CNNs) have demonstrated outstanding performance on image classification. To classify the hyperspectral images (HSIs), existing CNN-based approaches commonly adopt the architecture using single or several fixed spatial windows as inputs. This kind of architecture may lose contextual information or incorporate heterogeneous information due to the neglect of various land-cover distributions in HSIs. To deal with this problem, a novel attention multibranch CNN method based on adaptive region search (RS-AMCNN) is proposed for HSI classification. In RS-AMCNN, sizes and locations of spatial windows are searched in the nonlocal candidate region adaptively according to sample-specific distribution. These flexible spatial windows are input into several branches of RS-AMCNN. In each branch, convolutional long short-term memories (ConvLSTMs) are merged into CNN from shallow to deep layers, which not only extracts joint spatial-spectral features, but also exploits complementary information among different layers. Then, a branch attention mechanism is devised to emphasize more discriminative branches and suppress less useful ones. It forces RS-AMCNN to extract multiscale and multicontextual attention features for classification. Finally, RS-AMCNN is optimized end-to-end by combining the losses from the ramose classifiers of different branches and the main classifier. Experiments carried on several benchmark HSI data sets demonstrate that RS-AMCNN provides promising classification performance, especially in edge preservation and region uniformity.
Jie Feng 0003, Xiande Wu, Ronghua Shang, Chenhong Sui, Jie Li 0001, Licheng Jiao, Xiangrong Zhang
IEEE Trans. Geosci. Remote. Sens.7
2021 Deep Hash Learning for Remote Sensing Image Retrieval
abstract
The content-based remote sensing image retrieval (CBRSIR) has attracted increasing attention with the number of remote sensing (RS) images growing explosively. Benefiting from the strong capacity of the deep convolutional neural network (DCNN), the performance of CBRSIR has been improved in recent years. Although great successes have been obtained, learning the RS images' representative features and enhancing the retrieval efficiency for the large-scale CBRSIR tasks are still two challenging problems. In this article, we propose a new CBRSIR method named feature and hash (FAH) learning, which consists of a deep feature learning model (DFLM) and an adversarial hash learning model (AHLM). The DFLM aims at learning the RS images' dense features to guarantee the retrieval precision. In the DFLM, the DCNN and the proposed feature aggregation are integrated to capture the multiscale features. Then, the discrimination of the obtained features can be highlighted by the attention map in the developed attention branch. The AHLM maps the dense features onto the compact hash codes so that the retrieval efficiency can be improved. The AHLM contains a hash learning submodel and an adversarial regularization submodel. In particular, the hash learning submodel learns the real-valued hash codes that are similarity preserved by semantic supervisions. The adversarial regularization submodel regularizes the real-valued hash codes to learn the discrete uniform distribution with possible values 0 and 1. In this way, the hash codes are coding-balanced and the quantization errors are reduced. Encouraging experimental results counted on three public benchmark data sets demonstrate that our FAH can achieve competitive performance in the CBRSIR task compared with many existing hash learning methods.
Chao Liu 0042, Jingjing Ma 0001, Xu Tang 0004, Fang Liu 0034, Xiangrong Zhang, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.5
2021 Large-Scope PolSAR Image Change Detection Based on Looking-Around-and-Into Mode
abstract
A new method based on the Looking-Around-and-Into (LAaI) mode is proposed for the task of change detection in large-scope Polarimetric Synthetic Aperture Radar (PolSAR) image. Specifically, the LAaI mode consists of two processes named Look-Around and Look-Into, which are accomplished by attention proposal network (APN) and recurrent convolutional neural network (CNN) (Recurrent CNN), respectively. The former provides certain subregions efficiently, and the latter detects changes in subregions accurately. In Look-Around, difference image (DI) of whole PolSAR images is calculated first to get global information; then, APN is established to locate the position of interested subregions intentionally by paying special attention to; next interested subregions that contain changed area in high probability are picked out as candidate-regions. Moreover, candidate-regions are sorted in importance descending order so that highly interested regions have priority to be detected. In Look-Into, candidate-regions of different scales are selected at first; then, Recurrent CNN is constructed and employed to deal with multiscale PolSAR subimages so that clearer and finer change detection results are generated. The process is repeated until all candidate-regions are detected. As a whole, the proposed algorithm based on the LAaI mode looks around whole images first to find out the possible position of changes (candidate-regions generation in Look-Around) and then reveal the exact shape of changes in different scales (multiscale change detection in Look-Into). The effect of APN and Recurrent CNN is verified in experiments, and it shows that the proposed method performs well in the task of change detection in the large-scope PolSAR image.
Fang Liu 0034, Xu Tang 0004, Xiangrong Zhang, Licheng Jiao, Jia Liu 0020
IEEE Trans. Geosci. Remote. Sens.3
2021 Ridgelet-Nets With Speckle Reduction Regularization for SAR Image Scene Classification
abstract
With powerful feature representations, convolutional neural networks (CNNs) have produced tremendous achievements in image classification tasks and, typically, entail millions of labeled samples to train massive parameters. However, the sample labeling of synthetic aperture radar (SAR) images is extremely difficult, especially pixelwise labels, and has, sometimes, required field trips to accomplish labeling. Moreover, the inherent speckle noise may weaken the ability of networks to extract effective features from SAR images. In this article, we address these issues by labeling a few patchwise samples and propose Ridgelet-Nets with speckle reduction regularization for SAR image scene classification by combining deep learning with multiscale geometric analysis and statistical modeling of SAR images. First, we design Ridgelet-Nets with convolutional kernels constructed by ridgelet filters to reduce the training parameters and learn more discriminative features. Then, we embed speckle reduction regularization in the Ridgelet-Nets to restrain the influence of speckle noise and smooth the classification maps, in which the prior information of SAR image statistical modeling is introduced. Finally, we propose an adaptive SAR image scene classification framework based on an extended hierarchical visual semantic model, considering the differences in the structures and spatial relationships of different regions in the SAR images, particularly large-scale and complex scenes. Experimental results on real SAR images demonstrate that the proposed framework can achieve preferable classification performance using very limited labeled samples.
Xiaoxue Qian, Fang Liu 0001, Licheng Jiao, Xiangrong Zhang, Yuwei Guo 0001, Xu Liu 0006, Yuanhao Cui
IEEE Trans. Geosci. Remote. Sens.4
2021 Hyperspectral Image Classification Based on 3-D Octave Convolution With Spatial-Spectral Attention Network
abstract
In recent years, with the development of deep learning (DL), the hyperspectral image (HSI) classification methods based on DL have shown superior performance. Although these DL-based methods have great successes, there is still room to improve their ability to explore spatial-spectral information. In this article, we propose a 3-D octave convolution with the spatial-spectral attention network (3DOC-SSAN) to capture discriminative spatial-spectral features for the classification of HSIs. Especially, we first extend the octave convolution model using 3-D convolution, namely, a 3-D octave convolution model (3D-OCM), in which four 3-D octave convolution blocks are combined to capture spatial-spectral features from HSIs. Not only the spatial information can be mined deeply from the high- and low-frequency aspects but also the spectral information can be taken into account by our 3D-OCM. Second, we introduce two attention models from spatial and spectral dimensions to highlight the important spatial areas and specific spectral bands that consist of significant information for the classification tasks. Finally, in order to integrate spatial and spectral information, we design an information complement model to transmit important information between spatial and spectral attention features. Through the information complement model, the beneficial parts of spatial and spectral attention features for the classification tasks can be fully utilized. Comparing with several existing popular classifiers, our proposed method can achieve competitive performance on four benchmark data sets.
Xu Tang 0004, Xiangrong Zhang, Yiu-Ming Cheung, Jingjing Ma 0001, Fang Liu 0034, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.3
2021 Spectral-Difference Low-Rank Representation Learning for Hyperspectral Anomaly Detection
abstract
Anomaly detection of a hyperspectral image without any prior information has attracted much more attention in remote sensing image understanding and interpretation, which aims at determining whether a sample belongs to background or anomaly. Low-rank dictionary learning plays an important role in exploiting the low-rank prior of background for hyperspectral image (HSI) anomaly detection. In this article, the low-rank dictionary learning is introduced to learn a dictionary which can reconstruct the background positively, while anomaly cannot. Considering the high correlation of data especially between the adjacent bands, we resort to spectral-difference low-rank dictionary representation learning for global background modeling which can fully exploit the low-rank prior of background. Then, the residual matrix is used to distinguish anomaly. Different from the existing anomaly detection methods based on dictionary which is constructed or learned in a separated step, our proposed model can simultaneously learn the dictionary and separate anomaly by iterative learning. The experimental results on five real data sets demonstrate the superior performance of the proposed method for hyperspectral anomaly detection compared with other state-of-the-art algorithms.
Xiangrong Zhang, Xiaoxiao Ma 0003, Ning Huyan, Xu Tang 0004, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.1
2021 GRS-Det: An Anchor-Free Rotation Ship Detector Based on Gaussian-Mask in Remote Sensing Images
abstract
Ship detection is a significant and challenging task in remote sensing. Due to the arbitrary-oriented property and large aspect ratio of ships, most of the existing detectors adopt rotation boxes to represent ships. However, manual-designed rotation anchors are needed in these detectors, which causes multiplied computational cost and inaccurate box regression. To address the abovementioned problems, an anchor-free rotation ship detector, named GRS-Det, is proposed, which mainly consists of a feature extraction network with selective concatenation module (SCM), a rotation Gaussian-Mask model, and a fully convolutional network-based detection module. First, a U-shape network with SCM is used to extract multiscale feature maps. With the help of SCM, the channel unbalance problem between different-level features in feature fusion is solved. Then, a rotation Gaussian-Mask is designed to model the ship based on its geometry characteristics, which aims at solving the mislabeling problem of rotation bounding boxes. Meanwhile, the Gaussian-Mask leverages context information to strengthen the perception of ships. Finally, multiscale feature maps are fed to the detection module for classification and regression of each pixel. Our proposed method, evaluated on ship detection benchmarks, including HRSC2016 and DOTA Ship data sets, achieves state-of-the-art results.
Xiangrong Zhang, Guanchun Wang, Peng Zhu 0004, Tianyang Zhang 0002, Chen Li 0011, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.1
2020 Dual sentence representation model integrating prior knowledge for bio-text-mining
abstract
Data mining, especially the extraction of the relationship between genes and proteins, plays an important role in the biomedical field. Several related models have been proposed for data mining in the biomedical domain. Furthermore, manually curated biomedical knowledge bases, which could assist the task, have been used to enhance the data-mining model. However, due to the limitation of methods, much prior knowledge information is not be fully exploited. In this work, we propose a novel method that reasonably applied the curated prior knowledge for biomedical text mining by dual sentence representation models; one model is for the experimental data and the other one is for the prior knowledge information sentence. We evaluated our method on two community-supported datasets; BioNLP and BioCreative corpora. The experimental results demonstrate that the dual sentence representation model can successfully utilize external prior knowledge information to extract relationship from biomedical text. Our method can achieve state-of-art results and it could be an application of biomedical relation extraction in the future.
Zhijing Li 0005, YangYang Lan, Saikat Chatterjee, Pargorn Puttapirat, Xiangrong Zhang, Chen Li 0011
BIBM5
2020 Small Object Detection in Optical Remote Sensing Video with Motion Guided R-CNN
abstract
Deep learning (DL) based object detection methods have been making great achievements for natural images, which guides the vehicle detection of optical remote sensing videos (ORSV). Compared with natural images, objects in ORSV are smaller and blurrier, and most of vehicles are crowded. Thus, it is difficult for DL to detect these small objects only using the single-frame image. To address this problem, a motion guided R-CNN (MG-RCNN) is proposed. In MG-RCNN, motion information from consecutive frames is extracted by the mean differencing method and merged into apparent information to obtain motion-related discriminative features. Then, high-quality proposals are generated on the feature maps by mini-region proposal network (MRPN). For small targets, an improved loss function is defined by incorporating smooth factor, which makes the regression of shapes more stable. Experiments on ORSV demonstrate the proposed method shows superior detection performance over state-of-the-art deep learning methods.
Jie Feng 0003, Yuping Liang, Zhanwei Ye, Xiande Wu, Dening Zeng, Xiangrong Zhang, Xu Tang 0004
IGARSS6
2020 PolSAR Scene Classification via Low-Rank Tensor-Based Multi-View Subspace Representation
abstract
In this paper, the polarimetric synthetic aperture radar (PolSAR) scene classification is solved by using a novel low -rank tensor-based multi-view subspace representation (LRT-MSR) method. PolSAR data can be described in multimodal feature spaces, such as PolSAR coherent/covariance/scattering matrices, or the various polarimetric decompositions. Different pseudo-color images from multiple spaces provide enough visual information for making a comprehensive representation. Our method applies a low-rank tensor-based subspace clustering way to explore the information from multi-view pseudo-color images. Tensor, as the high order matrix, is used to capture the correlations of underlying multi-view data. Furthermore, the method is constrained by a low-rank term that elegantly models the cross information from different views, and achieves a series of representation matrices from the redundant information. Finally, a spectral cluster method is used to make the final classification. The experimental results on PolSAR image dataset show the effectiveness of the applied method.
Mengqian Chen, Bo Ren 0001, Biao Hou, Jocelyn Chanussot, Shuang Wang 0001, Xiangrong Zhang
IGARSS6
2020 Hyperspectral Image Classification Based on Semi-Supervised Dual-Branch Convolutional Autoencoder with Self-Attention
abstract
Deep learning method shows its powerful classification performance with sufficient available data. However, the labeled data is limited in hyperspectral images (HSIs). Semi-supervised algorithms have unique advantages on dealing with this problem. Therefore, a semi-supervised convolutional neural network is proposed in this paper. It consists of two branches, which use limited labeled samples and a large number of unlabeled samples, respectively. The first branch includes an encoder-decoder model to extract contextual information of unlabeled samples. The other one uses the similar construction except extra classification layers to extract discriminative features of labeled samples. In order to fuse contextual and discriminative information, we cascade the features of low-level layers from different branches. Furthermore, self-attention is added to the first branch, which focuses more on the global information for classification. The experiment results show that the proposed model provides a competitive result compared with state-of-the-art methods.
Jie Feng 0003, Zhanwei Ye, Yuping Liang, Xu Tang 0004, Xiangrong Zhang
IGARSS6
2020 Panchromatic Image Land Cover Classification Via DCNN with Updating Iteration Strategy
abstract
Land cover classification is a critical research task in many significant remote sensing applications. There are emerging many powerful pixel-level classification methods based on deep convolutional neural network (DCNN) in the universal computer vision community. However, due to the complication of satellite image senses and the lack of high-quality labeled datasets, these computer vision techniques can not be applied to remote sensing applications directly. In this paper, we propose a novel DCNN method to extract abstract feature from the complicated remote scene. The proposed method fuses three level features from the encoder while the segmentation result is obtained by decoder. Furthermore, we propose an updating iteration strategy (UIS) with label smoothing on training set to reduce the impact of the incorrect labels. The proposed strategy employs the output of the network to update the low-confidence labels on training set, and utilizes the updated labels to continue training the network. In order to acquire a better segmentation result on a very high resolution (VHR) panchromatic image, we transfer the features trained on GID dataset to our dataset for training. Our experiments has demonstrated the oustanding performance of the proposed method in land cover classification compared to DeepLabv3 on the GID and our dataset.
Biao Hou, Yangfei Liu, Tuotuo Rong, Bo Ren 0001, Zijuan Xiang, Xiangrong Zhang, Shuang Wang 0001
IGARSS6
2020 Remote Sensing Images Feature Learning based On Multi-Branch Networks
abstract
Remote sensing (RS) images feature learning, plays a crucial role in many RS images application, and attracts scholars' attention. However, since RS images contain complex contents, how to extract robust features that can fully represent RS images becomes an important and tough task. In this paper, we develop a feature learning method based on multi-branch networks, named M-Net, which consists of fine-grained branch and coarse branch. Considering the objects within RS images are diverse in type and resolution, the fine-grained branch is developed to capture rich object-level information. First, the RS images convolutional features are extracted by fine-grained branch. Second, through encoding the score maps which can highlight the important regions, the fine-grained structure mapping are obtained. Finally, the object-level features are generated by transforming the convolutional features through mapping. The coarse branch is developed to transform the obtained object-level features into global structure for representing images. The positive experimental results counted on RS benchmark data set demonstrate that the proposed M-Net can learn more powerful features.
Chao Liu 0042, Xu Tang 0004, Jingjing Ma 0001, Xiangrong Zhang, Fang Liu 0001, Junyong Ma, Licheng Jiao
IGARSS4
2020 Remote Sensing Scene Classification Based on Global and Local Consistent Network
abstract
Scene classification of remote sensing (RS) images has attracted increasing attention due to its wide applications. Recently, with the advances of deep learning models, especially convolutional neural networks (CNNs), the performance of remote sensing image scene classification has been significantly improved. In this paper, based on the popular CNN, we develop a new scene classification network, named the Global and Local Consistent (GLC) network, to deeply explore useful information from the RS images. First, we adopt a pre-trained CNN to learn the intermediate feature maps from the RS image pairs. Second, by introducing the visual attention mechanism, the global and local integration model is developed to mine the rich information from the obtained feature maps. Third, the attention consistent model is designed to eliminate the negative influence of the issue of attention inconsistency on the classification. To verify the effectiveness of the proposed method, we select two popular RS image data sets. Compared with some existing classification models, our network can achieve competitive results, which illustrates that our method is useul to the RS scene classification.
Jingjing Ma 0001, Qiushuo Ma, Xu Tang 0004, Xiangrong Zhang, Qunnie Peng, Licheng Jiao
IGARSS4
2020 Hyperspectral Image Classification Via Multi-Scale Encoder-Decoder Network
abstract
Hyperspectral image (HSI) classification is an important task in the remote sensing community. In general, many hyperspectral classification methods are based on pixel patch, which leads to information redundancy. In this paper, we propose a multi-scale encoder-decoder network for HSI classification. First, we adapt an encoder-decoder framework as the backbone network and use a skip connection between the encoder and decoder, the spatial information is obtained by this network. Second, we develop a multi-scale block to get the multi-scale information. Third, we retain complete spectral information through the constant number of spectral channels. Finally, an optimizer strategy is designed to achieve our model for the HSI classification task. We experiment with our method and other methods on two public datasets, and the results denote our model is useful for HSI classification task.
Jingjing Ma 0001, Linlin Wu, Xu Tang 0004, Xiangrong Zhang, Junyong Ma, Licheng Jiao
IGARSS4
2020 A Learnable Blur Kernel for Remote Sensing Image Retrieval
abstract
With the explosive increase of remote sensing images, content-based remote sensing image retrieval (CBRSIR) has aroused widespread attention. Convolutional Neural Network (CNN) based methods are widely used in CBRSIR due to the development of deep learning. However, common used CNN models have difficulties in holding shift-invariant property due to the widely used down-sampling method, which means a little shift of input may cause a mutation of feature representation. To mitigate the absence of shift-invariant in down-sampling, we propose the learnable blur kernel (LBK), that can enhance the feature extraction capability by leveraging more context information. We build on this concept without extra cost, which can be simply integrated with modern CNNs architecture. Our method is validated on the public remote sensing dataset and compared with other retrieval methods. The overall experimental results show that the proposed method achieves outstanding performance.
Zelin Peng, Guanchun Wang, Xiangrong Zhang, Xu Tang 0004, Licheng Jiao
IGARSS3
2020 Adaptive Feature Aggregation Network for Object Detection in Remote Sensing Images
abstract
Object detection in remote sensing images is a challenging task because of the large scale variations across the geospatial objects. The feature pyramid network (FPN) is widely used to alleviate the scale variations problem, however, it only fuses the features from adjacent levels and lacks the information of the entire feature hierarchy. In this paper, we propose a novel and effective feature pyramid aggregation network, called Adaptive Feature Aggregation Network (AFANet). Specifically, we propose the Adaptive Feature Aggregation (AFA) module to adaptively aggregate multi-level features of FPN and introduce the Bottom-up Path to enhance the location information of the entire feature levels. In addition, we use the Receptive Field Block (RFB) module to capture different receptive field features for each level feature map. We evaluate the effectiveness of our AFANet on the DOTA dataset and achieves noticeable performance compared with the baseline.
Wenliang Sun, Xiangrong Zhang, Tianyang Zhang 0002, Peng Zhu 0004, Xu Tang 0004
IGARSS2
2020 Hyperspectral Image Classification Based on Multiscale Spatial and Spectral Feature Network
abstract
With the development of deep learning, hyperspectral image (HSI) classification tasks have developed rapidly, the classification performance is improved in a big degree. Despite the great success of the existing methods, there is still room for improvement to extract features from spatial and spectral dimensions. In this paper, we propose a multiscale spatial and spectral feature network (MSSFN) to capture discriminative features for the classification of HSIs. Specifically, we first use three convolution layers to extract the features of original HSI data. Second, combining the spatial masks model and spectral attention model to build multiscale spatial and spectral model (MSSM). Through the MSSM model, the spatial information of different scales can be obtained and the useful spectral bands can be emphasized. Finally, in order to reduce the computation complexity and simplify network, the other three convolution layers with a small number of convolution kernels are adopted in our method. The experimental results demonstrate that our method is superior to most existing methods on two public HSI datasets.
Xu Tang 0004, Jingjing Ma 0001, Xiangrong Zhang, Fang Liu 0034, Qunnie Peng, Licheng Jiao
IGARSS4
2020 Supervised Adaptive-RPN Network for Object Detection in Remote Sensing Images
abstract
Object detection is one of the most important tasks in the field of very high resolution (VHR) remote sensing (RS) images understanding. Due to the characteristics of VHR RS images, the detection performance is always limited by class imbalance and intersection-over-union (IoU) distribution imbalance. To mitigate the adverse effects caused thereby, we propose a supervised adaptive-RPN (SA-RPN) model with the help of deep learning in this paper. First, we introduced a supervised multi-dimensional attention network to overcome the foreground-background class imbalance. It can help the network to highlight the foreground and suppress the background effectively. Second, we develop the adaptive-RPN to reduce the negative impact of the IoU distribution imbalance by adaptively selecting the size of the anchor. The positive experimental results on the public data set validate the usefulness of our SA-RPN model. Compared with the popular deep learning RS object detection methods, our method achieves improved performance.
Xu Tang 0004, Jingjing Ma 0001, Xiangrong Zhang, Licheng Jiao
IGARSS4
2020 Spatial-Spectral Smooth Graph Convolutional Network for Multispectral Point Cloud Classification
abstract
Multispectral point cloud, as a new type of data containing both spectrum and spatial geometry, opens the door to three-dimensional (3D) land cover classification at a finer scale. In this paper, we model the multispectral point cloud as a spatial-spectral graph and propose a smooth graph convolutional network for multispectral point cloud classification, abbreviated 3SGCN. We construct the spectral graph and spatial graph respectively to mine patterns in spectral and spatial geometric domains. Then, the multispectral point cloud graph is generated by combining the spatial and spectral graphs. For remote sensing scene classification tasks, it is usually desirable to make the classification map relatively smooth and avoid salt and pepper noise. Heat operator is introduced to enhance the low- frequency filters and enforce the smoothness in the graph signal. Further, a graph -based smoothness prior is deployed in our loss function. Experiments are conducted on real multispectral point cloud. The experimental results demonstrate that 3 SGCN can achieve significant improvements in comparison with several state-of-the art algori thms.
Qingwang Wang, Xiangrong Zhang, Yanfeng Gu
IGARSS2
2020 Scene Attention Mechanism for Remote Sensing Image Caption Generation
abstract
Remote sensing images play an important role in various applications. To make it easier for humans to understand remote sensing images, the task of remote sensing image captioning attracts more and more researchers' attention. Inspired from the way human receives visual information, attention mechanism has been widely used in remote sensing image understanding. To catch more scene information and improve the stability of the generated sentences, a new attention mechanism called scene attention is proposed. Except for the current attention via the current hidden state of the long shortterm memory network (LSTM), our proposed method simultaneously explores the global visual information from the mean feature of all convolutional features. The effectiveness of the proposed method is evaluated on UCM-captions, Sydney-captions and RSICD datasets. The results of our experiment show that comparing with some other captioning methods, our method is more stable and obtains a better performance.
Shiqi Wu, Xiangrong Zhang, Xin Wang 0068, Chen Li 0011, Licheng Jiao
IJCNN2
2020 Discriminative Feature Pyramid Network For Object Detection In Remote Sensing Images
abstract
Multi-class geospatial object detection in remote sensing images suffer great challenges, such as large scales variability and complex background. Although feature pyramid network (FPN) can alleviate the problem of scale variation to some extent, it causes the loss of spatial and semantic information which is not conducive to object location. To address the above problem, this paper proposes a discriminative feature pyramid network (DFPN) by introducing a global guidance module (GGM) and a feature aggregation module (FAM). Specifically, the global guidance module delivers the high-level semantic information to lower layers, so as to obtain feature maps with stronger semantic information to eliminate the interference caused by complex background. The feature aggregation module enhances the interflow of information between different layers and better captures the discrimination information at each layer. We validate the effectiveness of our method on the NWPU VHR-10 and RSOD datasets, the results outperform baseline by 2.06 and 3.88 points respectively.
Xiaoqian Zhu, Xiangrong Zhang, Tianyang Zhang 0002, Peng Zhu 0004, Xu Tang 0004, Chen Li 0011
IJCNN2
2020 BioNorm: deep learning-based event normalization for the curation of reaction databases
abstract
MOTIVATION: A biochemical reaction, bio-event, depicts the relationships between participating entities. Current text mining research has been focusing on identifying bio-events from scientific literature. However, rare efforts have been dedicated to normalize bio-events extracted from scientific literature with the entries in the curated reaction databases, which could disambiguate the events and further support interconnecting events into biologically meaningful and complete networks. RESULTS: In this paper, we propose BioNorm, a novel method of normalizing bio-events extracted from scientific literature to entries in the bio-molecular reaction database, e.g. IntAct. BioNorm considers event normalization as a paraphrase identification problem. It represents an entry as a natural language statement by combining multiple types of information contained in it. Then, it predicts the semantic similarity between the natural language statement and the statements mentioning events in scientific literature using a long short-term memory recurrent neural network (LSTM). An event will be normalized to the entry if the two statements are paraphrase. To the best of our knowledge, this is the first attempt of event normalization in the biomedical text mining. The experiments have been conducted using the molecular interaction data from IntAct. The results demonstrate that the method could achieve F-score of 0.87 in normalizing event-containing statements. AVAILABILITY AND IMPLEMENTATION: The source code is available at the gitlab repository https://gitlab.com/BioAI/leen and BioASQvec Plus is available on figshare https://figshare.com/s/45896c31d10c3f6d857a.
Peiliang Lou, Antonio Jimeno-Yepes, Zai Zhang 0002, Xiangrong Zhang, Chen Li 0011
Bioinform.5
2020 Bio-semantic relation extraction with attention-based external knowledge reinforcement
abstract
BACKGROUND: Semantic resources such as knowledge bases contains high-quality-structured knowledge and therefore require significant effort from domain experts. Using the resources to reinforce the information retrieval from the unstructured text may further exploit the potentials of such unstructured text resources and their curated knowledge. RESULTS: The paper proposes a novel method that uses a deep neural network model adopting the prior knowledge to improve performance in the automated extraction of biological semantic relations from the scientific literature. The model is based on a recurrent neural network combining the attention mechanism with the semantic resources, i.e., UniProt and BioModels. Our method is evaluated on the BioNLP and BioCreative corpus, a set of manually annotated biological text. The experiments demonstrate that the method outperforms the current state-of-the-art models, and the structured semantic information could improve the result of bio-text-mining. CONCLUSION: The experiment results show that our approach can effectively make use of the external prior knowledge information and improve the performance in the protein-protein interaction extraction task. The method should be able to be generalized for other types of data, although it is validated on biomedical texts.
Zhijing Li 0005, Yuchen Lian, Xiaoyong Ma, Xiangrong Zhang, Chen Li 0011
BMC Bioinform.4
2020 Fully Convolutional Network-Based Ensemble Method for Road Extraction From Aerial Images
abstract
This letter proposed a road extraction method based on fully convolutional networks (FCNs) with an ensemble strategy in order to solve the imbalance of road and background areas in aerial images. By utilizing the FCN, we consider road extraction as a semantic segmentation problem. In the network, the weight of the loss function is modified because of the imbalance between the roads and backgrounds, and there will be a larger punishment if roads are wrongly classified as background. Since it is difficult to determine an appropriate weight of the loss function for a given image, an ensemble method based on spatial consistency (SC) is proposed. The result maps that are obtained from the FCNs with different loss functions are fused in our proposed ensemble strategy, which also avoids the determination of weights. Our method is tested using the Massachusetts road data set, and it was proven to be effective compared with the base fully convolutional model according to our experimental result.
Xiangrong Zhang, Wenkang Ma, Chen Li 0011, Jie Wu 0016, Xu Tang 0004, Licheng Jiao
IEEE Geosci. Remote. Sens. Lett.1
2020 Patch Tensor-Based Multigraph Embedding Framework for Dimensionality Reduction of Hyperspectral Images
abstract
Graph-based dimensionality reduction (DR) techniques are of great interest in the field of image processing and especially on the analysis of hyperspectral images (HSIs). Considering the characteristics of hyperspectral data, many different types of graphs were designed to describe the structure of HSIs. Generally, the algorithms based on these graphs achieved promising performance. However, most of them only focus on how to improve the measurement of similarity between the data points by a single graph. Specifically, vector-based graph methods fail to capture the spatial information, while tensor-based graph methods assume that the pixels in each patch tensor belong to the same class, which is not exactly correct in practice. To overcome these shortcomings, this article proposes a patch tensor-based multigraph embedding (PTMGE) framework for the DR of HSIs, in which three different types of subgraphs are constructed to comprehensively describe the intrinsic geometrical structures of HSIs. First, a tensor subgraph is constructed to capture the spatial information and local geometrical structure. Second, for each two neighboring patch tensors in the tensor graph, a bipartite graph is designed to characterize the pixel-based relationships between the patch tensors. Then, considering that the diversity of pixels may be existed in each patch tensor, a pixel-based subgraph is built to describe the inner geometrical structures of every patch tensor. Finally, a novel graph fusion strategy is designed to calculate a final similarity matrix for projection learning. Experiments on three real hyperspectral data sets are conducted, and comparison with some state-of-the-art algorithms validated the effectiveness of our proposed PTMGE method.
Yangjun Deng, Heng-Chao Li 0001, Yong-Jian Sun, Xiangrong Zhang, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.5
2020 Multi-Feature Weighted Sparse Graph for SAR Image Analysis
abstract
Sparse representation (SR) method has the advantages of good category distinguishing performance, noise robustness, and data adaptiveness. In this article, a multi-feature weighted sparse graph (MWSG) is presented for synthetic aperture radar (SAR) image analysis. First, multiple types of features are extracted to fully describe the characteristics of SAR image. Then, multiple SRs of samples in multiple feature spaces are obtained by solving a weighted joint SR model, in which the weight is the Gaussian kernel distance among samples. Moreover, a new fusion mechanism is given to integrate multiple weighted SRs, which aims to eliminate the negative influence of the singular data, so the MWSG is obtained. Afterward, the brief steps of the SAR image segmentation and semisupervised classification based on MWSG are stated. A series of experiments on the simulated and real SAR images shows that the MWSG has better performance than other existing relevant methods.
Licheng Jiao, Fang Liu 0001, Xiangrong Zhang, Xu Tang 0004, Puhua Chen
IEEE Trans. Geosci. Remote. Sens.4
2020 Sketch-Based Region Adaptive Sparse Unmixing Applied to Hyperspectral Image
abstract
Hyperspectral image (HSI) unmixing is an important issue of research due to its effect on the subsequent processing of HSIs. Recently, the sparse regression method with spatial information has been successfully applied in hyperspectral unmixing (HU). However, most sparse regression methods ignore the difference in spatial structure handling with only one sparse constraint. In fact, the pixels in detail regions are more likely to be severely mixed with more endmembers participated, and the sparsity degree of its corresponding abundances is relatively low. Considering the sparsity difference of abundances, a sketchbased region adaptive sparse unmixing applied to HSI is proposed in this article. Inspired by the vision computing theory, we use the region generation algorithm based on a sketch map to differentiate the homogeneous regions and detail regions. Then, the abundances of these two kind regions in HSIs are separately constrained by sparse regularizers of L1/2and L1with a proposed manifold constraint. Our method not only makes full use of the spatial information in HSIs but also exploits the latent structure of data. The encouraging experimental results on three data sets validate the effectiveness of our method for HU.
Xiangrong Zhang, Xu Tang 0004, Puhua Chen, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.2
2019 OpenHI2 - Open source histopathological image platform
abstract
Transition from conventional to digital pathology requires a new category of biomedical informatic infrastructure which could facilitate delicate pathological routine. Pathological diagnoses are sensitive to many external factors and is known to be subjective. Only systems that can meet strict requirements in pathology would be able to run along pathological routines and eventually digitized the area, and the developed platform should comply with existing pathological routines and international standards. Currently, there are a number of available software tools which can perform histopathological tasks including virtual slide viewing, annotating, and basic image analysis, however, none of them can serve as a digital platform for pathology. Here we describe OpenHI2, an enhanced version Open Histopathological Image platform which is capable of supporting all basic pathological tasks and file formats; ready to be deployed in medical institutions on a standard server environment or cloud computing infrastructure. In this paper, we also describe the development decisions for the platform and propose solutions to overcome technical challenges including responsive region retrieval and viewing, virtual slide magnification, recording of diagnostic areas. These factors would promote OpenHI2 be used as a platform for histopathological images in real-world clinical settings. Furthermore, in research, OpenHI2 inherited the annotation functionality from the previous version, thus acquired annotations can be directly utilized by the newly added machine learning module which include popular machine learning models to perform tasks such as histology image classification and segmentation in the same environment. Addition can be made to the platform since each component is modularized and fully documented. OpenHI2 is free, open-source, and available at https://gitlab.com/BioAI/OpenHI.
Pargorn Puttapirat, Chen Li 0011, Haichuan Zhang 0001, Jingyi Deng, Yuxin Dong 0003, Jiangbo Shi, Zeyu Gao 0001, Chunbao Wang 0002, Xiangrong Zhang
BIBM10
2019 Effects of annotation granularity in deep learning models for histopathological images
abstract
Pathological is crucial to cancer diagnosis. Usually, Pathologists draw their conclusion based on observed cell and tissue structure on histology slides. Rapid development in machine learning, especially deep learning have established robust and accurate classifiers. They are being used to analyze histopathological slides and assist pathologists in diagnosis. Most machine learning systems rely heavily on annotated data sets to gain experiences and knowledge to correctly and accurately perform various tasks such as classification and segmentation. Generally, annotations made in pathology-related datasets have inherited annotation methods from natural scene images. This work investigates different granularity of annotations in histopathological data set including image-wise, bounding box, ellipse-wise, and pixel-wise to verify the influence of annotation in pathological slide on deep learning models. We design corresponding experiments to test classification and segmentation performance of deep learning models based on annotations with different annotation granularity. In classification, state-of-the-art deep learning-based classifiers perform better when trained by pixel-wise annotation dataset. On average, precision, recall and F1-score improves by 7.87%, 8.83% and 7.85% respectively. Thus, it is suggested that finer granularity annotations are better utilized by deep learning algorithms in classification tasks. Similarly, semantic segmentation algorithms can achieve 8.33% better segmentation accuracy when trained by pixel-wise annotations. Our study shows not only that finer-grained annotation can improve the performance of deep learning models, but also help they extract more accurate phenotypic information from histopathological slides. The accurate and spatially precise acquisitions of phenotypic information can improve the reliability of the model prediction. Intelligence systems trained on granular annotations may help pathologists inspecting certain regions and features in the slide that were mainly used to calculate the prediction. The compartmentalized prediction approach similar to this work may contribute to phenotype and genotype association studies.
Jiangbo Shi, Zeyu Gao 0001, Haichuan Zhang 0001, Pargorn Puttapirat, Chunbao Wang 0002, Xiangrong Zhang, Chen Li 0011
BIBM6
2019 Hyperspectral Band Selection Based On Ternary Weight Convolutional Neural Network
abstract
In this paper, a novel ternary weight convolution neural network (TWCNN) is proposed for band selection of hyperspectral images. TWCNN constructs deep-wise convolution layer with 1×1 filters as the first layer of the network, which is used for band selection. In the deep-wise convolution layer, weights are constrained to -1, 0, or 1. -1 and 1 represent that the corresponding band is selected, while 0 indicates it's not. TWCNN constructs subsequent layers to extract features and classify for selected spectral bands. It combines band selection, feature extraction and classification into a unified optimization procedure, which makes it to achieve end-to-end band selection and classification. Furthermore, the constraints of the number of spectral bands is added to the cost function of TWCNN. The specific number of spectral bands can be selected. The experiment results show that the proposed model provides a competitive result to state-of-the-art methods.
Jie Feng 0003, Jiantong Chen, Xiangrong Zhang, Xu Tang 0004, Xiande Wu
IGARSS4
2019 Joint Multilayer Spatial-Spectral Classification of Hyperspectral Images Based on CNN and Convlstm
abstract
The following topics are dealt with: remote sensing; geophysical image processing; synthetic aperture radar; radar imaging; remote sensing by radar; image classification; learning (artificial intelligence); feature extraction; vegetation; image resolution.
Jie Feng 0003, Xiande Wu, Jiantong Chen, Xiangrong Zhang, Xu Tang 0004
IGARSS4
2019 Object Detection and Trcacking Based on Convolutional Neural Networks for High-Resolution Optical Remote Sensing Video
abstract
Object detection algorithms, from high-resolution optical remote sensing images, have been booming from the last few years. However, object tracking for high-resolution optical remote sensing video is a challenging task due to the large number and small size of objects. In this paper, we propose an object detection and tracking method based on deep convolutional neural networks for wide swath high-resolution optical remote sensing videos. The proposed method firstly segments each frame of a video into sub-samples using a sliding window of fixed size. In order to detect the objects appearing at the edge of the sliding window efficiently, we use an overlapping sliding window sampling method. Further, we design a network fusing region of interests (RoIs) of the previous and current frames to track the objects occurred in the previous frames of the video. RoIs of previous frame are applied directly to the feature layer of the current frame. Finally, for each frame, we merge the detection and tracking results of sub-samples by non-maximum suppression (NMS) method. The experimental results on our dataset demonstrate the validity and generality of the proposed detection algorithm.
Biao Hou, Jingliang Li, Xiangrong Zhang, Shuang Wang 0001, Licheng Jiao
IGARSS3
2019 Weak Moving Object Detection In Optical Remote Sensing Video With Motion-Drive Fusion Network
abstract
Object detection in optical remote sensing video (ORSV) is a new trend which makes it possible for obtaining richer information in more complicated and diverse situations. However, the small objects are blurred in the videos captured from optical sensor assembled in satellite, limited by the devices and natural weather. The concept of weak object is defined in this situation that the objects are extremely small and hardly detected with only one static image. Therefore, we propose a simple but efficient method for weak moving object detection in ORSV by combining the temporal information from neighbor frames and spatial features from image pixels. First, we compute the difference map between two adjacent frames, and stack it with original RGB channel so that a (1+3)-channel input data is made. Then, a motion-drive D-RGB (difference map with RGB image) fusion network is developed to obtain the feature map of this (1+3)-channel data. To adapt unusual scale in ORSV images, based on statistical prior objects size, we change the size of anchor box in original Faster R-CNN. The proposed method is demonstrated to improve the mean average precision on detecting weak moving remote sensing objects.
Yuxuan Li 0004, Licheng Jiao, Xu Tang 0004, Xiangrong Zhang
IGARSS4
2019 Adversarial Hash-Code Learning for Remote Sensing Image Retrieval
abstract
Hashing, a useful solution for approximate nearest neighbor (ANN) search, is popular for large-scale image retrieval. In this paper, we presents a deep supervised hashing model for remote sensing image retrieval (RSIR) in the framework of generative adversarial networks (GAN), named GAN-assist Hashing (GAAH). First, to learn the compact and useful hash codes from the images, we define a novel loss function for the generator. The loss function mainly consists of classification, similarity, and bits entropy terms. The classification term makes the hash code is discriminative, the similarity term constrains the binary code is similarity preserving, and the bits entropy term assures the learned code is low-error in the quantization. Second, we construct the unique "true" matrix with the uniform distribution as the input of discriminator to limit the leaned hash codes are bit balanced. The final hash code is learned by a minimax optimization. The positive experimental results on a ground-truth remote sensing image archive validate the usefulness of our GAAH model. Compare with the popular deep hashing methods, our GAAH achieves improved performance.
Chao Liu 0042, Jingjing Ma 0001, Xu Tang 0004, Xiangrong Zhang, Licheng Jiao
IGARSS4
2019 Polsar Land Cover Classification via Tensorial Embedding Methods
abstract
In recent years, graph embedding has become a significant technique to deal with feature extraction and dimension reduction problems. Under the linearization and kernelization, it provides a unified framework in machine learning and other pattern recognition tasks. Polarimetric synthetic aperture (PolSAR) as a typical multi-channel sensor can obtain more geometrical and geophysical information. How to combine those polarimetric scattering signals and target decomposition features and explore the spatial information between pixels become a new research direction to address PolSAR data. In this paper, we utilize the tensorial embedding methods to extract the intrinsic features from a redundant feature space for the PolSAR land cover classification. The effectiveness of the proposed methods is demonstrated using AIRSAR Flevoland data set.
Bo Ren 0001, Biao Hou, Jocelyn Chanussot, Changzhe Jiao, Xiangrong Zhang
IGARSS5
2019 Remote Sensing Image Retrieval Based on Semi-Supervised Deep Hashing Learning
abstract
As an useful solution of the approximate nearest neighbor (ANN) search, hashing attracts growing attention in the topic of large-scale image retrieval. In this paper, we propose a semi-supervised deep hashing method based on the adversarial autoencoder (AAE) network for remote sensing image retrieval (RSIR), and we name it SSHAAE. Here, we assume the RS images have been represented by the visual features, and the target of our SSHAAE is mapping those features into the binary codes. First, a hashing layer is adopted to replace the part of original latent layer in AAE. In addition, the classical reconstruction loss function is selected to generate the hash code. Second, two discriminators are added simultaneously to make sure the hash code is bit balanced and the generated label variable is one-hot. Third, we design the hash loss function to guarantee the obtained hash code is discriminative, similarity persevering, and low quantization error. The presented SSHAAE model can be trained by the minimax optimization. The encouraging experimental results counted on a high-resolution RS image archive demonstrate our SSHAAE model is effective to RSIR.
Xu Tang 0004, Chao Liu 0042, Xiangrong Zhang, Jingjing Ma 0001, Changzhe Jiao, Licheng Jiao
IGARSS3
2019 Classification of Hyperspectral Images Based on Multiclass Spatial-Spectral Generative Adversarial Networks
abstract
Generative adversarial networks (GANs) are famous for generating samples by training a generator and a discriminator via an adversarial procedure. For hyperspectral image classification, the collection of samples is always difficult. However, directly applying GAN to hyperspectral image classification exists two problems. One is that the generated samples lack discriminative information. Meanwhile, the discriminator has no discriminative ability for multiclassification. Another is that spatial and spectral information requires to be considered in hyperspectral image classification simultaneously. To address these problems, a novel multiclass spatial-spectral GAN (MSGAN) method is proposed. In MSGAN, two generators are devised to generate the samples containing spatial and spectral information, respectively, and the discriminator is devised to extract joint spatial-spectral features and output multiclass probabilities. Moreover, novel adversarial objectives for multiclass are defined. The discriminator is devised to predict training samples belonging to true classes and generated samples belonging to all the classes with the same probability. The generators are devised to make the discriminator mistake. By adversarial learning between the discriminator and generators, the classification performance of the discriminator is promoted with the assistance of discriminative generated samples. Experimental results on hyperspectral images demonstrate that the proposed method achieves encouraging classification performance compared with several state-of-the-art methods, especially with the limited training samples.
Jie Feng 0003, Haipeng Yu, Xianghai Cao, Xiangrong Zhang, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.5
2019 Hyperspectral Anomaly Detection via Background and Potential Anomaly Dictionaries Construction
abstract
In this paper, we propose a new anomaly detection method for hyperspectral images based on two well-designed dictionaries: background dictionary and potential anomaly dictionary. In order to effectively detect an anomaly and eliminate the influence of noise, the original image is decomposed into three components: background, anomalies, and noise. In this way, the anomaly detection task is regarded as a problem of matrix decomposition. Considering the homogeneity of background and the sparsity of anomalies, the low-rank and sparse constraints are imposed in our model. Then, the background and potential anomaly dictionaries are constructed using the background and anomaly priors. For the background dictionary, a joint sparse representation (JSR)-based dictionary selection strategy is proposed, assuming that the frequently used atoms in the overcomplete dictionary tend to be the background. In order to make full use of the prior information of anomalies hidden in the scene, the potential anomaly dictionary is constructed. We define a criterion, i.e., the anomalous level of a pixel, by using the residual calculated in the JSR model within its local region. Then, it is combined with a weighted term to alleviate the influence of noise and background. Experiments show that our proposed anomaly detection method based on potential anomaly and background dictionaries construction can achieve superior results compared with other state-of-the-art methods.
Ning Huyan, Xiangrong Zhang, Huiyu Zhou 0001, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.2
2019 Unsupervised Cross-Temporal Classification of Hyperspectral Images With Multiple Geodesic Flow Kernel Learning
abstract
With the increasing acquisition ability of hyperspectral remote sensing images, unsupervised cross-temporal classification (UCTC) of hyperspectral images (HSIs) has attracted more and more attention. In this paper, we focus on cross-temporal HSI classification, i.e., using one labeled HSI to classify the other unlabeled HSI. A multiple geodesic flow kernel learning (MGFKL) framework is proposed to exploit both spatial and spectral features for UCTC with bitemporal HSIs and called S2-MGFKL. The proposed S2-MGFKL method first extracts extended multi-attribute profiles (EMAPs) from the original bitemporal HSIs. The spatial features of the bitemporal HSIs obtained by the same attribute filter are paired up, so are the original spectral features. Second, each pair of features from both source and target domains are used to construct multiple geodesic flows. According to the original definition of GFK, we can obtain the construction of Gaussian base GFKs. The base kernels consist of two parts, the spectral part is obtained base on the same geodesic flow (which is constructed on the bitemporal spectral features) by tuning the kernel scale, while the spatial part is obtained under the same kernel scale but different geodesic flows constructed on different spatial feature pairs. After that, the mean rule is adopted to acquire the combined kernel, which is fed into the supervised vector machine (SVM) to implement the cross-temporal classification task. Experiments are conducted on two real HSI data sets, and the results compared with several well-known methods demonstrate the effectiveness of the proposed method.
Tianzhu Liu, Xiangrong Zhang, Yanfeng Gu
IEEE Trans. Geosci. Remote. Sens.2
2018 OpenHI - An open source framework for annotating histopathological image
Pargorn Puttapirat, Haichuan Zhang 0001, Yuchen Lian, Chunbao Wang 0002, Xiangrong Zhang, Lixia Yao, Chen Li 0011
BIBM5
2018 PolSAR Image Classification Based on DBN and Tensor Dimensionality Reduction
abstract
This paper proposes a new semi-supervised PolSAR image classification method using deep belief network (DBN) and tensor dimensionality reduction, which uses multilinear principle component analysis (MPCA) to reduce the dimension of tensor form PolSAR data, and regards the multiple features of PolSAR data as the input of DBN. In order to take full advantage of neighborhood information of each pixel of PolSAR data, we take each pixel and its neighborhood as tensor form. For PolSAR data, simple feature has been proven not to be able to effectively classify complex terrains. Therefore, we combine multiple features of PolSAR data to obtain more abundant information, which can reflect some spatial structure of PolSAR data. The experimental results show that the overall classification accuracy based on the proposed method outperforms the traditional classification strategies.
Biao Hou, Xianpeng Guo, Weidan Hou, Shuang Wang 0001, Xiangrong Zhang, Licheng Jiao
IGARSS5
2018 Hyper-Laplacian Regularized Low-Rank Tensor Decomposition for Hyperspectral Anomaly Detection
abstract
This paper presents a novel method for hyperspectral anomaly detection considering the spectral redundancy and exploiting spectral-spatial information at the same time. We proposed a Hyper-Laplacian regularized low-rank tensor decomposition method combing with dimensionality reduction framework. Firstly, k-means++ algorithm is implemented to spectral bands and centers of each group are selected to reduce the HSI dimensionality in spectral direction. To jointly utilize spectral-spatial information, the cubic data (two spatial dimensions and one spectral dimension) is treated as a 3-order tensor. Then the non-local self-similarity is fully explored in our method. For the reason to reduce the ringing artifacts caused by over-lapped segmentation in exploring the non-local self-similarity, we introduce the hyper-Laplacian constrained low-rank tensor decomposition and we get the separated background and residual parts. Finally, to eliminate the effect of Gaussian noise, we use local-Rx basic detector to detect the residual matrix. Experimental results on two real hyperspectral data sets verified the effectiveness of the proposed algorithms for HSI anomaly detection.
Xiaoxiao Ma 0003, Xiangrong Zhang, Ning Huyan, Xu Tang 0004, Biao Hou, Licheng Jiao
IGARSS2
2018 Circular Relevance Feedback for Remote Sensing Image Retrieval
abstract
Relevance feedback (RF) is a popular reranking technique, which aims at improving the performance of image retrieval by taking the user's opinions into account. In this paper, we introduce a new RF method, named circular relevance feedback (CRF), to enhance the behavior of remote sensing image retrieval (RSIR). Instead of the manual selection used in the common RF method, we adopt the active learning (AL) algorithm to select the samples from the initial results automatically in each RF iteration. Moreover, to ensure the selected images are representative and informative enough, we choose different AL algorithms to complete the different RF processes. Finally, the contributions of all AL-driven RF methods are integrated using a circular fusion scheme. The encouraging experimental results on the ground truth RS image archive illustrate that our CRF is useful for enhancing the performance of RSIR. In addition, compared with many existing RF methods, our CRF achieves improved behavior.
Xu Tang 0004, Xiangrong Zhang, Fang Liu 0001, Licheng Jiao
IGARSS2
2018 Spatial-Spectral Graph-Based Nonlinear Embedding Dimensionality Reduction for Hyperspectral Image Classificaiton
abstract
Dimensionality reduction (DR) is one of the most important tasks to improve the performance of hyperspectral images classification. Recently, a sparse and low-rank graph embedding based method (SLGE) has been proposed to describe the intrinsic structure of data combined with the local and global constraint simultaneously, which is effective to reduce the dimension of hyperspectral data and obtain a better classification accuracy. However, SLGE is based on an assumption that low-dimensional feature can be obtained utilizing a linear projection. Its performance may degrade under nonlinearly distributed data. Moreover, spatial prior of HSI is not considered in the framework. In this paper, we proposed a novel dimensionality reduction method named spatial-spectral graph-based non-linear embedding (SSGNE). To generate a new graph-trained data, the segmentation strategy based on superpixel is adopted. The spatial-spectral graph is constructed by constraining the sparsity and low-rankness simultaneously on graph-trained data set. Finally, the kernel trick is adopted to extend the general graph embedding framework to nonlinearly space, which fully considers the complexity of real data. Experimental results show that the proposed method outperforms the state-of-the-art methods in terms of the classification accuracy.
Xiangrong Zhang, Yaru Han, Ning Huyan, Chen Li 0011, Jie Feng 0003, Xiaoxiao Ma 0003
IGARSS1
2018 Hyperspectral Unmixing via Deep Convolutional Neural Networks
abstract
Hyperspectral unmixing (HU) is a method used to estimate the fractional abundances corresponding to endmembers in each of the mixed pixels in the hyperspectral remote sensing image. In recent times, deep learning has been recognized as an effective technique for hyperspectral image classification. In this letter, an end-to-end HU method is proposed based on the convolutional neural network (CNN). The proposed method uses a CNN architecture that consists of two stages: the first stage extracts features and the second stage performs the mapping from the extracted features to obtain the abundance percentages. Furthermore, a pixel-based CNN and cube-based CNN, which can improve the accuracy of HU, are presented in this letter. More importantly, we also use dropout to avoid overfitting. The evaluation of the complete performance is carried out on two hyperspectral data sets: Jasper Ridge and Urban. Compared with that of the existing method, our results show significantly higher accuracy.
Xiangrong Zhang, Licheng Jiao
IEEE Geosci. Remote. Sens. Lett.1
2018 MMM: classification of schizophrenia using multi-modality multi-atlas feature representation and multi-kernel learning
Jin Liu 0012, Xiangrong Zhang, Yi Pan 0001, Jianxin Wang 0001
Multim. Tools Appl.3
2018 Tensor-Based Low-Rank Graph With Multimanifold Regularization for Dimensionality Reduction of Hyperspectral Images
abstract
Dimensionality reduction is an essential task in hyperspectral image processing. How to preserve the original intrinsic structure information and enhance the discriminant ability is still a challenge in this area. Recently, with the advantage of preserving global intrinsic structure information, low-rank representation has been applied to dimensionality reduction and achieved promising performance. By exploiting the submanifold information of the original data set, multimanifold learning is effective in enhancing the discriminant ability of the processed data set. In addition, due to the ability of preserving the spatial neighborhood structure information, the tensor analysis has become a popular technique for hyperspectral image processing. Motivated by the above-mentioned analysis, a novel tensor-based low-rank graph with multimanifold regularization (T-LGMR) for dimensionality reduction of hyperspectral images is proposed in this paper. In the T-LGMR, a low-rank constraint is employed to preserve the global data structure while multimanifold information is utilized to enhance the discriminant ability, and tensor representation is used to preserve the spatial neighborhood information. Finally, dimensionality reduction is achieved in the graph embedding framework. Experimental results on three real hyperspectral data sets demonstrate the superiority of the proposed method over several state-of-the-art approaches.
Jinliang An, Xiangrong Zhang, Huiyu Zhou 0001, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.2
2018 A Hybrid Method of SAR Speckle Reduction Based on Geometric-Structural Block and Adaptive Neighborhood
abstract
Given the improvement of synthetic aperture radar (SAR) imaging technologies, the resolution of SAR image is largely improved and the variation of backscatter amplitude should be considered in SAR image processing. In this paper, considering the spatial geometric properties of SAR image in gray pixel space and the sample selection in the estimation of true signal, local directional property of each pixel is explored with the help of SAR sketching method, and two specially designed filters are integrated for adaptive speckle reduction of SAR images. Specifically, based on the sketch map of a SAR image, the orientation of the sketch point lying at each sketch segment is assigned to the corresponding pixel, and thus all pixels of the SAR image are classified as the directional pixels and the nondirectional pixels. For the directional pixels, given the significant directionality of its neighborhood, a geometric-structural block (GB) is built to center on it and GB-wised nonlocal means filter is designed to estimate the true values of all pixels contained in the GB. Moreover, using the local orientation, the whole image is adopted as the searching range to search the similar GBs. For the nondirectional pixels, based on the locally estimated equivalent number of looks, a novel pixel-based metric is proposed to determine the local adaptive neighborhood (AN) with which an AN-based filter is developed to estimate its true value. Besides, since some nondirectional pixels are contained in GBs, a Bayesian-based fusion strategy is designed for the fusion of their estimated values. In the experiments, three synthetic speckled images and five real SAR images [obtained with different resolutions (e.g., 3, 1, and 0.1 m) and different bands (e.g., X-band, C-band, and Ka-band)] are used for evaluation and analysis. Owing to the usage of local spatial geometric property and the combination of two different filters, the proposed method shows a reasonable performance among the comparison methods, in terms of the speckle reduction and the details' preservation.
Fang Liu 0001, Jie Wu 0016, Lingling Li 0002, Licheng Jiao, Hongxia Hao, Xiangrong Zhang
IEEE Trans. Geosci. Remote. Sens.6
2018 Multifeature Hyperspectral Image Classification With Local and Nonlocal Spatial Information via Markov Random Field in Semantic Space
abstract
Hyperspectral images (HSIs) provide invaluable information in both spectral and spatial domains for image classification tasks. In this paper, we use semantic representation as a middle-level feature to describe image pixels' characteristics. Deriving effective semantic representation is critical for achieving good classification performance. Since different image descriptors depict characteristics from different perspectives, combining multiple features in the same semantic space makes semantic representation more meaningful. First, a probabilistic support vector machine is used to generate semantic representation-based multifeatures. In order to derive better semantic representation, we introduce a new adaptive spatial regularizer that well exploits the local spatial information, while a nonlocal regularizer is also used to search for global patch-pair similarities in the whole image. We combine multiple features with local and nonlocal spatial constraints using an extended Markov random field model in the semantic space. Experimental results on three hyperspectral data sets show that the proposed method provides better performance than several state-of-the-art techniques in terms of region uniformity, overall accuracy, average accuracy, and Kappa statistics.
Xiangrong Zhang, Zeyu Gao 0001, Licheng Jiao, Huiyu Zhou 0001
IEEE Trans. Geosci. Remote. Sens.1
2018 Hybrid Unmixing Based on Adaptive Region Segmentation for Hyperspectral Imagery
abstract
Unmixing is an important issue of hyperspectral images. Most unmixing methods adopt linear mixing models for simplicity. However, multiple scattering usually occurs between vegetation and soil in a bilinear scene. Thus, nonlinear mixing problems which are difficult to be solved should be taken into consideration under this circumstance. In practice, both linear and nonlinear spectral mixtures exist in hyperspectral scenes. Considering the characteristics of different regions in images, we propose a hybrid unmixing algorithm for hyperspectral images based on region adaptive segmentation. Our method uses a standard K-means clustering algorithm to obtain different regions, including homogeneous regions and detailed regions. The model of the homogeneous regions is assumed to be linear, which will be pursued using the method of sparse-constrained nonnegative matrix factorization (NMF), and the mixing in the detailed regions is assumed to be based on a nonlinear model. We also propose a new nonlinear unmixing method, called graph-regularized semi-NMF, which considers the manifold structure of hyperspectral data as the unmixing method to deal with the detailed regions. Finally, by combining the two regions, we obtain the abundance of the whole hyperspectral image. The proposed method can not only achieve more precise abundance but also be good at keeping the edge information of the bilinear abundance. The experimental results on both synthetic and real data also show that the proposed method is effective for improving the unmixing accuracy of hyperspectral remote-sensing images.
Xiangrong Zhang, Chen Li 0011, Cai Cheng, Licheng Jiao, Huiyu Zhou 0001
IEEE Trans. Geosci. Remote. Sens.1
2017 Hyperspectral image classification based on stacked marginal discriminative autoencoder
abstract
In this paper, a novel stacked marginal discriminative autoencoder (SMDAE) method is proposed for hyperspectral image classification. It uses a deep neural network to learn discriminative features from hyperspectral images automatically. In hyperspectral images, the collection of training samples is difficult. When the number of training samples is not enough, these training samples are difficult to estimate the statistical distribution of hyperspectral images accurately. In order to solve the small sample problem and improve the classification performance of the autoencoder, the marginal samples are selected through the distribution characteristics of samples. The marginal samples are searched based on k nearest neighbors between different classes. These samples are used to fine-tune the SMDAE network. The experimental results show that the proposed SMDAE method can achieve satisfying performance under small training set.
Jie Feng 0003, Liguo Liu, Xiangrong Zhang, Rongfang Wang, Hongying Liu 0001
IGARSS3
2017 Fast graph-based SAR image segmentation via simple superpixels
abstract
Graph-based methods have been successfully applied in the field of computer vision for image segmentation. Unfortunately, most of them are not suitable to deal with large-scale SAR image segmentation due to their high computation complexity. A fast and efficient graph-based SAR image segmentation is proposed in this paper through using superpixels to reduce the computation complexity. Firstly, a SAR image is divided into several non-overlapped subdivisions with the same size. Each of the subdivision is processed as a single OpenMP parallel region, which can be processed at the single computing node with multi-core CPU. Secondly, the number of nodes and edges in the graph is reduced by extracting the superpixels other than single pixels based on global information of each subdivision. Finally, an effective rule is proposed to merge two adjacent sub-graphs from two different subdivisions into a new subgraph.
Biao Hou, Dezhao Gong, Shuang Wang 0001, Xiangrong Zhang, Licheng Jiao
IGARSS5
2017 A weighted joint sparse of three channels method for full POL-SAR data classification
abstract
In recent years, both passive and active (i.e., Synthetic Aperture Radar or SAR) satellite remote sensing has proven to be valuable tools for mapping land cover. Most of the classification algorithms are based on image-intensity and they do not perform well in different coastal zone types, because these terrains have similar optical or radar backscattering signals. In this paper we propose a weighted joint sparse on the three-channel to mine the polarimetric features and texture information. The proposed method can update weights automatically according to the difference of the three channels' contribution and the least residual error. The similarity and distinctiveness of three channels are used to deal with the complex object classification. Hybrid sparse coefficients are input to the support vector machine for fully polarimetric image classification. The proposed algorithm performed well in distinguishing some coastal land-use types. A comparison study is also conducted to show that proposed algorithm outperforms two commonly classification methods.
Wenshuai Chen, Shuiping Gou, Xiangrong Zhang, Xiaofeng Li 0001, Licheng Jiao
IGARSS4
2017 Natural language description of remote sensing images based on deep learning
abstract
The semantic description of remote sensing image is a useful and meaningful task, which can help us to get a better understanding of the scene depicted in the remote sensing images and make better use of the remote sensing images. Nature language provides good solution for describing the semantic information of remote sensing images. Nature language description of a remote sensing image is to generate a meaningful sentence given a remote sensing image. This paper presents a novel method based on deep learning. First, a convolutional neural network is utilized to detect the main objects of the remote sensing images. Then a recurrent neural network language model is utilized to generate the natural language descriptions of the objects which are detected in the first step. Experimental results on a set of remote sensing images demonstrate that the proposed method is able to generate desirable description of the scene.
Xiangrong Zhang, Xiang Li 0013, Jinliang An, Biao Hou, Chen Li 0011
IGARSS1
2017 Recursive Autoencoders-Based Unsupervised Feature Learning for Hyperspectral Image Classification
abstract
For hyperspectral image (HSI) classification, it is very important to learn effective features for the discrimination purpose. Meanwhile, the ability to combine spectral and spatial information together in a deep level is also important for feature learning. In this letter, we propose an unsupervised feature learning method for HSI classification, which is based on recursive autoencoders (RAE) network. RAE utilizes the spatial and spectral information and produces high-level features from the original data. It learns features from the neighborhood of the investigated pixel to represent the whole local homogeneous area of the image. In addition, to obtain more accurate representation of the investigated pixel, a weighting scheme is adopted based on the neighboring pixels, where the weights are determined by the spectral similarity between the neighboring pixels and the investigated pixel. The effectiveness of our method is evaluated by the experiments on two hyperspectral data sets, and the results show that our proposed method has a better performance.
Xiangrong Zhang, Chen Li 0011, Ning Huyan, Licheng Jiao, Huiyu Zhou 0001
IEEE Geosci. Remote. Sens. Lett.1
2017 Unsupervised saliency-guided SAR image change detection
Yaoguo Zheng, Licheng Jiao, Hongying Liu 0001, Xiangrong Zhang, Biao Hou, Shuang Wang 0001
Pattern Recognit.4
2016 Automated Segmentation of MOOC Lectures towards Customized Learning
abstract
The sheer size of the student body for MOOC and the diversity of their learning styles and backgrounds demand that we develop alternatives to the one-size-fits-all pedagogy used in residential education. An important aspect of this endeavor is the segmentation of the video material, since it forms the omnipresent and central part of every course, and structuralized videos allow non-linear navigation as well as help learners with various needs find desired information efficiently. Here, we propose an automatic visual transition detection method to partition lecture videos into self-contained segments, which is the foundation to structuralize video and support non-linear navigation. Our method can be done at scale and has been proved being able to achieve reasonable quality.
Xiangrong Zhang, Chen Li 0011, Shang-Wen Li 0001, Victor Zue
ICALT1
2016 A weighted multi-task joint sparse representation method for hyperspectral image classification
abstract
In this paper, a novel weighted multi-task joint sparse representation method is proposed for hyperspectral image classification. It is assumed that the importance of atoms in a dictionary can be weighted when they are used in sparse representation according to the similarities between tasks and classes. We utilize tasks instead of classes in pre-classification to group all samples into several clusters, one cluster stands for one tasks. Then we calculate the correlations between tasks and labels based on the reconstruction errors of sparse representation acquired from training samples. The correlations are used for the consideration of heterogeneous neighborhood, which is the core of multi-task method. The weights of different tasks can be adjusted using training samples according to the reconstruction errors. At last, all samples can be classified more accurately via task correlations and reconstruction errors. Experimental results on real hyperspectral data sets exhibit its superiority to compared algorithm.
Jinliang An, Yu Mo, Zhi Guo, Xiangrong Zhang
IGARSS4
2016 Joint multi-feature hyperspectral image classification with spatial constraint in semantic manifold
abstract
This paper presents a novel method for hyperspectral classification combining multiple features and exploiting spatial information at the same time. We proposed a supervised classification method under the Markov random field (MRF)-based framework. Firstly using the probability SVM to map multiple features from different low-level subspace to the same semantic space (probability space), then integrating these features in semantic space with MRF-based model to enforce a smooth and accurate representation, in addition the manifold distance has been used in MRF-based model to measure the similarity of two point. To further improve the classification accuracy, a new approach of building the adaptive neighborhood has been proposed and used in our method. As our model is a derivable and convex problem, gradient descent can be used to solve this problem with less computational and time cost. Experimental results on real hyperspectral dataset shows that the proposed method provides improved classification accuracy in terms of the overall accuracy, average accuracy and kappa statistic.
Xiangrong Zhang, Zeyu Gao 0001, Jinliang An, Yanning Hu, Yangyang Li 0001, Biao Hou
IGARSS1
2016 Spatially constrained Bag-of-Visual-Words for hyperspectral image classification
abstract
This paper proposes a spatially constrained Bag-of-Visual-Words (BOV) method for hyperspectral image classification. We firstly extract the texture feature. The spectral and texture features are used as two types of low-level features, based on which, the high-level visual-words are constructed by the proposed method. We use the entropy rate superpixel segmentation method to segment the hyperspectral into patches that well keep the homogeneousness of regions. The patches are regarded as documents in BOV model. Then k-means clustering is implemented to cluster pixels to construct codebook. Finally, the BOV representation is constructed with the statistics of the occurrence of visual-words for each patch. Experiments on a real data show that the proposed method is comparable to several state of the art methods.
Xiangrong Zhang, Yaoguo Zheng, Jinliang An, Yanning Hu, Licheng Jiao
IGARSS1
2016 A nonlinear subspace multiple kernel learning for financial distress prediction of Chinese listed companies
Xiangrong Zhang, Longying Hu
Neurocomputing1
2016 Weighted multifeature hyperspectral image classification via kernel joint sparse representation
Erlei Zhang, Xiangrong Zhang, Licheng Jiao, Hongying Liu 0001, Shuang Wang 0001, Biao Hou
Neurocomputing2
2016 Single image super-resolution reconstruction based on genetic algorithm and regularization prior model
Yangyang Li 0001, Yang Wang 0075, Yaxiao Li, Licheng Jiao, Xiangrong Zhang, Rustam Stolkin
Inf. Sci.5
2016 Dimensionality Reduction Based on Group-Based Tensor Model for Hyperspectral Image Classification
abstract
Dimensionality reduction is one of the most important tasks in improving hyperspectral image (HSI) classification performance and has been widely studied. In traditional dimensionality reduction methods, the hyperspectral data cube is first converted to a matrix, and then, the matrix-algebra is employed, in which the spatial structure is usually not taken into consideration. To jointly utilize spectral and spatial information, researches considering HSI as a tensor have attracted more and more attention. In this letter, we propose a group-based tensor model for HIS dimensionality reduction. The local and nonlocal spatial information of HSI cubes is explored by segmenting the original HSI tensors into a lot of small tensors and grouping them into clusters. Finally, the clusters are projected into low-rank subspace to obtain a proper feature space. Experimental results on two real hyperspectral data sets exhibit the effectiveness of the proposed algorithm for HSI-dimensionality reduction.
Jinliang An, Xiangrong Zhang, Licheng Jiao
IEEE Geosci. Remote. Sens. Lett.2
2016 A Nonlocal Means for Speckle Reduction of SAR Image With Multiscale-Fusion-Based Steerable Kernel Function
abstract
For the robustness of a patch-based metric, the nonlocal means method is widely applied for speckle reduction of synthetic aperture radar (SAR) images, where the similarity computed by the patch-based metric is used as weight, and weighted averaging is used to obtain the true value. However, not knowing the local spatial property, a fixed kernel (e.g., Gaussian kernel or uniform kernel) is always used to compute the weight. This is not good for the preservation of geometrical features (e.g., edges, lines, and points). In this letter, considering the characteristics of SAR imagery, a multiscale-fusion-based steerable kernel function was formed to explore the local spatial property of SAR images. In addition, by combining the kernel function with a ratio-based similarity metric designed with the distribution of the speckle's ratio, a new patch-based metric was formed and used with the nonlocal scheme for speckle reduction. In the experiments, by comparing with two state-of-the-art methods, a reasonable performance was obtained by our method, in terms of speckle reduction and detail preservation.
Jie Wu 0016, Fang Liu 0001, Hongxia Hao, Lingling Li 0002, Licheng Jiao, Xiangrong Zhang
IEEE Geosci. Remote. Sens. Lett.6
2016 Hierarchical Discriminative Feature Learning for Hyperspectral Image Classification
abstract
Building effective image representations from hyperspectral data helps to improve the performance for classification. In this letter, we develop a hierarchical discriminative feature learning algorithm for hyperspectral image classification, which is a deformation of the spatial-pyramid-matching model based on the sparse codes learned from the discriminative dictionary in each layer of a two-layer hierarchical scheme. The pooling features achieved by the proposed method are more robust and discriminative for the classification. We evaluate the proposed method on two hyperspectral data sets: Indiana Pines and Salinas scene. The results show our method possessing state-of-the-art classification accuracy.
Xiangrong Zhang, Yunlong Liang, Yaoguo Zheng, Jinliang An, Licheng Jiao
IEEE Geosci. Remote. Sens. Lett.1
2016 Local Collaborative Representation With Adaptive Dictionary Selection for Hyperspectral Image Classification
abstract
Spectral-spatial representation based algorithms have been widely applied in hyperspectral image (HSI) classification, which exploit the fact that pixels in a local patch often have similar spectral reflectance values and probably belong to the same class. Collaborative representation (CR) is a typical supervised classification method for high-dimensional data, which has been widely used for spectral-spatial representation based HSI classification. However, it suffers from the degraded representation of redundant and irrelative pixels when all of the labeled pixels are used as a dictionary for representation. In this letter, a novel method, local CR with adaptive dictionary selection, is proposed to solve this problem, in which we first average the values of pixels from local patches to incorporate the contextual information of neighbors, and then, an adaptive dictionary selection method is presented to select the most similar pixels to each test pixel from the dictionary to reduce the influence of redundant and irrelevant pixels in representation. Experimental results on two HSIs show that the proposed method outperforms some spectral-spatial representation based algorithms in terms of classification accuracy.
Yaoguo Zheng, Licheng Jiao, Ronghua Shang, Biao Hou, Xiangrong Zhang
IEEE Geosci. Remote. Sens. Lett.5
2016 Unsupervised feature selection based on maximum information and minimum redundancy for hyperspectral images
Jie Feng 0003, Licheng Jiao, Fang Liu 0001, Tao Sun 0007, Xiangrong Zhang
Pattern Recognit.5
2016 Spectral-spatial hyperspectral image ensemble classification via joint sparse representation
Erlei Zhang, Xiangrong Zhang, Licheng Jiao, Lin Li 0016, Biao Hou
Pattern Recognit.2
2016 Multiple Kernel Learning Based on Discriminative Kernel Clustering for Hyperspectral Band Selection
abstract
In hyperspectral images, band selection plays a crucial role for land-cover classification. Multiple kernel learning (MKL) is a popular feature selection method by selecting the relevant features and classifying the images simultaneously. Unfortunately, a large number of spectral bands in hyperspectral images result in excessive kernels, which limit the application of MKL. To address this problem, a novel MKL method based on discriminative kernel clustering (DKC) is proposed. In the proposed method, a discriminative kernel alignment (KA) (DKA) is defined. Traditional KA measures kernel similarity independently of the current classification task. Compared with KA, DKA measures the similarity of discriminative information by introducing the comparison of intraclass and interclass similarities. It can evaluate both kernel redundancy and kernel synergy for classification. Then, DKA-based affinity-propagation clustering is devised to reduce the kernel scale and retain the kernels having high discrimination and low redundancy for classification. Additionally, an analysis of necessity for DKC in hyperspectral band selection is provided by empirical Rademacher complexity. Experimental results on several hyperspectral images demonstrate the effectiveness of the proposed band selection method in terms of classification performance and computation efficiency.
Jie Feng 0003, Licheng Jiao, Tao Sun 0007, Hongying Liu 0001, Xiangrong Zhang
IEEE Trans. Geosci. Remote. Sens.5
2016 SAR Image Segmentation Based on Hierarchical Visual Semantic and Adaptive Neighborhood Multinomial Latent Model
abstract
A synthetic aperture radar (SAR) imaging system usually produces pairs of bright area and dark area when depicting the ground objects, such as a building or tree and its shadow. Many buildings (trees) are aggregated together to form urban areas (forests). It means that the pairs of bright and dark areas often exist in the aggregated scenes. Conventional unsupervised segmentation approaches usually segment the scenes (e.g., urban areas and forests) into different regions simply according to the gray values of the image. However, a more convincing way is to regard them as the consistent regions. In this paper, we aim at addressing this issue and propose a new SAR image segmentation approach via a hierarchical visual semantic and adaptive neighborhood multinomial latent model. In this approach, the hierarchical visual semantic of SAR images is proposed, which divides SAR images into aggregated, structural, and homogeneous regions. Based on the division, different segmentation methods are chosen for these regions with different characteristics. For the aggregated region, locality-constrained linear coding-based hierarchical clustering is used for segmentation. For the structural region, visual semantic rules are designed for line object location, and a geometric structure window-based multinomial latent model is proposed for segmentation. For the homogeneous region, a multinomial latent model with adaptive window selection is proposed for segmentation. Finally, these results are integrated together to obtain the final segmentation. Experiments on both synthetic and real SAR images indicate that the proposed method achieves promising performances in terms of the consistencies of the regions and the preservations of the edges and line objects.
Fang Liu 0001, Yiping Duan, Lingling Li 0002, Licheng Jiao, Jie Wu 0016, Shuyuan Yang 0001, Xiangrong Zhang, Jialing Yuan
IEEE Trans. Geosci. Remote. Sens.7
2015 Polarimetric SAR images classification using deep belief networks with learning features
abstract
A novel polarimetric synthetic aperture radar (PolSAR) image classification method based on Deep Belief Networks (DBNs) is proposed in this paper. First, the coherency matrix data are converted to a 9-dimentional data. Second, many patches are randomly selected from each dimension in the 9-dimentional data, and many filters can be obtained from a Restricted Boltzmann Machine (RBM) trained by using these patches. Thus we can get the features for each pixel from each dimension in the 9-dimentional space. Finally, the learned features and the elements of coherent matrix are combined to train a 3-layers DBNs for PolSAR image classification. Experimental results show that the proposed method is efficient and effective for PolSAR image classification.
Biao Hou, Xiaohuan Luo, Shuang Wang 0001, Licheng Jiao, Xiangrong Zhang
IGARSS5
2015 Sparsity-constrained generalized bilinear model for hyperspectral unmixing
abstract
Generalized bilinear model (GBM) has been widely used for nonlinear hyperspectral image unmixing. However, it does not take the sparse information of abundance into account, which is a significant characteristic resulting from the correlation of hyperspectral data. This paper aims to extend the GBM by incorporating the sparsity constraint of abundance matrix with the semi-nonnegative matrix factorization, by dividing GBM into the linear part and the second-order part, which are optimized using an alternating optimization algorithm respectively. L1/2-norm is used to explore the sparse characteristic, and the L1/2-constrained semi-nonnegative matrix factorization (L1/2-semi-NMF) algorithm is presented, which leads to better results on both synthetic and real data.
Xiangrong Zhang, Cai Cheng, Jinliang An, Yaoguo Zheng, Erlei Zhang, Biao Hou
IGARSS1
2015 Scaling cut criterion-based discriminant analysis for supervised dimension reduction
Xiangrong Zhang, Yudi He, Licheng Jiao, Ruochen Liu 0006, Jie Feng 0003
Knowl. Inf. Syst.1
2015 Imbalanced Hyperspectral Image Classification Based on Maximum Margin
abstract
Hyperspectral remote sensing images own rich spectral information to distinguish different land-cover classes. Sometimes, it may encounter the case that some classes have much fewer pixels than other classes. In this case, traditional classification methods are not appropriate because they are prone to assign all the pixels to the classes with a large number of pixels. For such an imbalanced problem, ensemble learning is a good method by partitioning the majority classes into different groups with small sizes. However, the existing ensemble schemes are independent of classifiers, which will not get the best performance for a certain classifier. In this letter, the selected classifier, i.e., a support vector machine (SVM), is considered in an ensemble procedure to improve the classification accuracy. Specifically, the criterion of the SVM, i.e., the maximum margin, is adopted to guide the ensemble learning procedure for imbalanced hyperspectral image classification. Experiments state that our method obtains higher classification accuracy than the SVM and several representative imbalanced classification methods for hyperspectral images.
Tao Sun 0007, Licheng Jiao, Jie Feng 0003, Fang Liu 0001, Xiangrong Zhang
IEEE Geosci. Remote. Sens. Lett.5
2015 Fast Multifeature Joint Sparse Representation for Hyperspectral Image Classification
abstract
Since hyperspectral images (HSIs) usually have complex content and chaotic background, multiple kinds of features would be helpful for the classification task. Recently, representation-based methods with multifeature combination learning have been proposed. However, multifeature learning and the extended contextual information require much more computational burden, particularly for a large-scale dictionary case. In this letter, we propose a fast joint sparse representation classification method with multifeature combination learning for hyperspectral imagery. Once getting several complementary features (spectral, shape, and texture), the proposed model simultaneously acquires a representation vector for each kind of feature and imposes the joint sparsity ℓrow,0-norm regularization on the representation coefficients. The regularization can enforce the coefficients to share a common sparsity pattern, which preserves the crossfeature information. A new version of the simultaneous orthogonal matching pursuit is presented to solve the aforementioned problem because of its optimization with strong convergence guarantee and efficiency. Moreover, to further improve the classification performance, we incorporate contextual neighborhood information of the image into each kind of feature. Compared with state-of-the-art algorithms, it has been proved that the proposed algorithm with much less memory requirements performs tens to hundreds of times faster than those on real HSIs, while providing the same (or even better) accuracy.
Erlei Zhang, Xiangrong Zhang, Hongying Liu 0001, Licheng Jiao
IEEE Geosci. Remote. Sens. Lett.2
2015 Mutual-Information-Based Semi-Supervised Hyperspectral Band Selection With High Discrimination, High Information, and Low Redundancy
abstract
The large number of spectral bands in hyperspectral images provides abundant information to distinguish different land covers. However, these spectral bands have much redundancy and bring an extra computational burden. Thus, band selection is important for hyperspectral images. Since the labeled samples are difficult to obtain, a semi-supervised criterion based on maximum discrimination and information (MDI) is defined by using both limited labeled samples and sufficient unlabeled samples. This MDI criterion aims to select the most highly discriminative and informative bands, but it is hard to accurately calculate. Therefore, a novel criterion based on high discrimination, high information, and low redundancy (DIR) is proposed as its low-order approximation. Moreover, from an information theory perspective, a theoretical proof is given that many traditional semi-supervised feature selection criteria are the low-order approximations of this MDI criterion. Compared with them, the proposed criterion needs more relaxed approximation conditions. To search and optimize the proposed criterion, a novel clonal selection algorithm is proposed, where the adaptive clone and mutation operators are devised to speed up the convergence. Experimental results on hyperspectral images demonstrate the effectiveness of the proposed semi-supervised band selection method.
Jie Feng 0003, Licheng Jiao, Fang Liu 0001, Tao Sun 0007, Xiangrong Zhang
IEEE Trans. Geosci. Remote. Sens.5
2014 Biclustering of gene expression data using Particle Swarm Optimization integrated with pattern-driven local search
abstract
Biclustering is of great significance in the analysis of gene expression data and is proven to be a NP-hard problem. Among the existing intelligent optimization algorithms used in the gene expression data analysis, most concentrate on the global search ability but ignore the inherent trajectory information of gene expression data, so the search efficiency is low. In this paper, a pattern-driven local search operator is incorporated in the binary Particle Swarm Optimization (PSO) algorithm in order to improve the search efficiency. Experiments show that our approach is valid.
Yangyang Li 0001, Xiaolong Tian, Licheng Jiao, Xiangrong Zhang
IEEE Congress on Evolutionary Computation4
2014 A novel algorithm for many-objective dimension reductions: Pareto-PCA-NSGA-II
abstract
Many-objective problem has more than 3 objectives. Because of the extraordinary difficulty of acquiring their Pareto optimal solutions directly, traditional methods will be out of operation for such problems. In recent years, many researchers have turned their attention to the study of this area. They are interested in two areas: acquiring some part of Pareto front which is useful to the researchers (Preferred Solutions) and reducing redundant objectives. In this paper, we combine two dimension reduction methods: the method based on Pareto optimal solution analysis and the method based on correlation analysis, to form a novel algorithm for dimension reduction. Firstly, the Pareto optimal solutions are acquired through NSGA-II. Then the objectives who contribute little to the number of non-dominated solutions are removed. At last, the dimension of objectives is reduced further according to their contribution to the principal component in PCA analysis. In this way, we can acquire the right non-redundant objectives with low time complexity. Simulation results show that the proposed algorithm can effectively reduce redundant objectives and keep the non-redundant objectives with low time.
Ronghua Shang, Licheng Jiao, Wei Fang 0001, Xiangrong Zhang, Xiaolin Tian 0002
IEEE Congress on Evolutionary Computation5
2014 SAR image segmentation based on random projection and signature frame
abstract
This paper proposes a new Synthetic Aperture Radar (SAR) image segmentation method based on the frame of Signature/Earth Mover's Distance (EMD). Firstly, Random Projection is used to extract features of SAR image, which has the abilities of preserving information and reducing dimensionality. Secondly, a signature is used to obtain the cluster center and the weight. Finally, by computing the distance between two signatures using Earth Mover's Distance, we can obtain the final segmentation result. The experimental results show that the proposed method is efficient and effective for SAR image segmentation.
Biao Hou, Shuang Wang 0001, Xiangrong Zhang
IGARSS4
2014 Classification of imbalanced hyperspectral imagery data using support vector sampling
abstract
Due to the imbalance in obtaining labeled samples for different land-cover classes, hyperspectral image classification encounters the issue of imbalanced classification. In this paper, a novel and effective method is proposed to address the imbalanced learning problem in hyperspectral image classification, which combines support vector machine (SVM) and sampling strategy. The main novelty and contribution of our paper are that we propose to do sampling referring to the support vectors (SVs) rather than the training data to provide a balanced distribution during the model learning. Sampling among the training data may be time consuming, while sampling referring to the SVs is more efficient and representative with much lower complexity. Therefore, the proposed method is expected to be simple and effective for imbalanced learning problem. Experimental results on real hyperspectral image dataset show that our method can effectively improve the classification accuracy for the minority classes in the imbalanced dataset.
Xiangrong Zhang, Yaoguo Zheng, Biao Hou, Shuiping Gou
IGARSS1
2014 Low-rank representation based action recognition
abstract
Human action recognition is an important problem in computer vision, which has been applied to many applications. However, how to learn an accurate and discriminative representation of videos based on the features extracted from videos still remains to be a challenging problem. In this paper, we propose a novel method named low-rank representation based action recognition to recognize human actions. Given a dictionary, low-rank representation aims at finding the lowest-rank representation of all data, which can capture the global data structures. According to its characteristics, low-rank representation is robust against noises. Experimental results demonstrate the effectiveness of the proposed approach on several publicly available datasets.
Xiangrong Zhang, Yang Yang 0062, Hanghua Jia, Huiyu Zhou 0001, Licheng Jiao
IJCNN1
2014 SAR image segmentation based on quantum-inspired multiobjective evolutionary clustering algorithm
Yangyang Li 0001, Shixia Feng, Xiangrong Zhang, Licheng Jiao
Inf. Process. Lett.3
2014 Improving Hyperspectral Image Classification Using Spectral Information Divergence
abstract
In order to improve the classification performance for hyperspectral image (HSI), a sparse representation classifier based on spectral information divergence (SID) is proposed. SID measures the discrepancy of probabilistic behaviors between the spectral signatures of two pixels from the aspect of information theory, which can be more effective in preserving spectral properties. Thus, the new method measures the similarity between the reconstructed pixel and the true pixel by SID instead of by the L2 norm used in traditional sparse model. Moreover, the spatial coherency across neighboring pixels sharing a common sparsity pattern is taken into account during the construction of SID-based joint sparse representation model. We propose a new version of the orthogonal matching pursuit method to solve SID-based recovery problems. The proposed SID-based algorithms are applied to real HSI for classification. Experimental results show that our algorithms outperform the classical sparse representation based classification algorithms in most cases.
Erlei Zhang, Xiangrong Zhang, Shuyuan Yang 0001, Shuang Wang 0001
IEEE Geosci. Remote. Sens. Lett.2
2014 Using Combined Difference Image and k-Means Clustering for SAR Image Change Detection
abstract
In this letter, a simple and effective unsupervised approach based on the combined difference image and$k$-means clustering is proposed for the synthetic aperture radar (SAR) image change detection task. First, we use one of the most popular denoising methods, the probabilistic-patch-based algorithm, for speckle noise reduction of the two multitemporal SAR images, and the subtraction operator and the log ratio operator are applied to generate two kinds of simple change maps. Then, the mean filter and the median filter are used to the two change maps, respectively, where the mean filter focuses on making the change map smooth and the local area consistent, and the median filter is used to preserve the edge information. Second, a simple combination framework which uses the maps obtained by the mean filter and the median filter is proposed to generate a better change map. Finally, the$k$-means clustering algorithm with$k = 2$is used to cluster it into two classes, changed area and unchanged area. Local consistency and edge information of the difference image are considered in this method. Experimental results obtained on four real SAR image data sets confirm the effectiveness of the proposed approach.
Yaoguo Zheng, Xiangrong Zhang, Biao Hou, Ganchao Liu
IEEE Geosci. Remote. Sens. Lett.2
2014 Laplacian group sparse modeling of human actions
Xiangrong Zhang, Licheng Jiao, Yang Yang 0062, Feng Dong 0005
Pattern Recognit.1
2014 Hyperspectral Band Selection Based on Trivariate Mutual Information and Clonal Selection
abstract
Band selection is an important preprocessing step for hyperspectral data processing. It involves two crucial problems, i.e., suitable measure criterion and effective search strategy. Mutual information (MI) has been widely used as the measure criterion for its nonlinear and nonparametric characteristics. For efficient calculation, traditional MI-based criteria commonly use bivariate MI (BMI) to approximate the ideal MI-based criterion. However, these BMI-based criteria may miss the bands having discriminative information and do not give the condition of the approximation. In this paper, a novel criterion based on trivariate MI (TMI) is proposed to measure the redundancy for classification. From the multivariate MI perspective, the proposed TMI-based and traditional BMI-based criteria are proved as the low-order approximations of the ideal criterion under some assumptions. Compared with the BMI-based criteria, a more relaxed assumption condition is required for the TMI-based criterion. To alleviate the problem of few labeled samples existing in hyperspectral images, the TMI-based criterion is extended to the semisupervised TMI-based (STMI) method by adding a graph regulation term. Additionally, to search an appropriate band subset by the TMI- and STMI-based criteria, a new clonal selection algorithm (CSA) is proposed. In CSA, integer encoding and adaptive operators are devised to reduce space and time cost. Experimental results demonstrate the effectiveness of the proposed algorithms for hyperspectral band selection.
Jie Feng 0003, Licheng Jiao, Xiangrong Zhang, Tao Sun 0007
IEEE Trans. Geosci. Remote. Sens.3
2014 Eigenvalue Analysis-Based Approach for POL-SAR Image Classification
abstract
A novel polarimetric synthetic aperture radar (POL-SAR) image classification approach is proposed in this paper by exploiting coherency matrix eigenvalues for polarimetric information representation and understanding. The approach consists of two parts. Initially, the statistical distributions of eigenvalue for homogeneous areas are analyzed by taking eigenvalues as the features of polarimetric information. The Bayesian classification method is applied to verify the feasibility of distinguishing different homogeneous areas. As a result, this method can work well those pixels with the similar scatter mechanism by using different polarimetric intensity information from eigenvalues. But this process cannot adequately distinguish those pixels with similar eigenvalues distribution. So, an eigenvalues-based local operator is defined to overcome the insufficient of the similar pixels by introducing a similar measure and eigenvalues-based texture information. After all pixels are classified by Bayesian classification, if the similarity of the pixel is larger than the given threshold, this pixel will be further classified by support vector machine using texture information. The proposed method is tested on three POL-SAR datasets, in which the average classification accuracy of eight categories for the Flevoland data from our method reaches nearly 90%.
Shuiping Gou, Xiangrong Zhang, Weifang Wang, Fangfang Du
IEEE Trans. Geosci. Remote. Sens.3
2014 Local Maximal Homogeneous Region Search for SAR Speckle Reduction With Sketch-Based Geometrical Kernel Function
abstract
With the flourish of the nonlocal mean method, the neighborwise similarity metric is widely applied in speckle reduction for its robust performance on the search of similar samples. In this metric, an isotropic kernel function is usually chosen to aggregate the corresponding pixels' distance between two neighborhoods. It means that the kernel function is considered as the explanation of the local spatial relationship at each pixel. However, for anisotropic features (such as edges and lines), a strong relationship exists along their directions rather than across them, so the isotropic kernel is not suitable to explain the spatial relationship around these features. Meanwhile, due to the inherent speckle in synthetic aperture radar (SAR) images, the discrimination and exploration of the geometrical properties of anisotropic features are important for the construction of adaptive kernel function. In this paper, the sketch map which is a representation of the sketch information of SAR images is extracted as the criterion for designing the kernel function. Meanwhile, due to the properties of symmetric and maximal self-similarity, a modified ratio distance is proposed and used jointly with the constructed kernel function as a similarity metric. Then, under the local stationary assumption, the local maximal homogeneous region of each pixel is searched by using the region growing method with the proposed metric. Moreover, maximal likelihood rule is used within the region for the estimation of true value. From the experiments on the synthetic and real SAR images, a promising performance in terms of speckle reduction and preservation of the details is achieved by our proposed method.
Jie Wu 0016, Fang Liu 0001, Licheng Jiao, Xiangrong Zhang, Hongxia Hao, Shuang Wang 0001
IEEE Trans. Geosci. Remote. Sens.4
2013 Joint segmentation and classification of hyperspectral image using meanshift and sparse representation classifier
abstract
A novel spectral-spatial classification method based on mean shift and sparse representation classifier (SRC) for hyperspectral images is proposed in this paper. Firstly, the nonnegative matrix factorization, is used as a preprocessing for mean shift. Then, the mean shift algorithm is adopted to partition an image into amount of blocks and get the segmentation map. Through this way, many size-variable and close regions can be got while the boundary information is remained. Secondly, the classification map is obtained by using the SRC. Finally, the fusion of the segmentation map and the classification map is done by using the majority vote rule. Experimental results on two real hyperspectral images demonstrate the effectiveness and good performance of the proposed method.
Xiangrong Zhang, Yaoguo Zheng, Biao Hou, Xiaojin Hou
IGARSS1
2013 Spatial-spectral classification based on group sparse coding for hyperspectral image
abstract
In this paper, a novel hyperspectral image classification method is proposed, based on group sparse coding. The method is based on this acknowledgement that larger spatial variation exists in high spatial resolution hyperspectral image, which degrades the separability of hyperspectral image. In order to obtain a smooth representation, each pixel and its spatial neighbors are coded together by group sparse coding. Although nothing about class information is included, the neighbor pixels in a small spatial window are inclined to belong to the same class. Thus, that will reduce the within-class scatter and be favorable to the classification task. Then, the obtained sparse representation vectors are used for hyperspectral image classification with SVM. Experimental results show that our method exceeds the classical classification algorithms in accuracy and regional consistency.
Xiangrong Zhang, Peng Weng, Jie Feng 0003, Erlei Zhang, Biao Hou
IGARSS1
2013 Fully integrated passive UHF RFID transponder IC with a sensitivity of -12 dBm
abstract
This paper presents a fully integrated passive UHF RFID transponder, which includes an RF/analog front-end, a baseband processor and a 512-bit EEPROM memory. To improve power conversion efficiency, Schottky barrier diode based rectifier is adopted. The design of a novel voltage limiter is discussed in detail, which can provide a stable limiting voltage with low sensitivity to temperature variation and process dispersion. The whole transponder chip is implemented in 0.18 μm CMOS process with a die size of 800×800 μm2. Measurement results show that the total power consumption of the tag chip is only 7.4μW with a sensitivity of -12dBm.
Jinpeng Shen, Xin'an Wang, Bo Wang 0016, Shoucheng Li, Zhengkun Ruan, Xiangrong Zhang
ISCAS7
2013 A UWB mixer with a balanced wide band active balun using crossing centertaped inductor
abstract
A high gain 3-5GHz mixer merged with a very balanced wideband active balun is presented and demonstrated. It uses a Gilbert type folded structure with the active balun as the input trans-conductance stage and with a PMOS switch stage. The implemented circuit in the 0.18um CMOS technology exhibits less than 1dB amplitude imbalance in 2.06-5.2GHz, less than 1.5dB in 0.55-5.7GHz and less than 2 degree phase imbalance through 1.37GHz to 5.07GHz. The mixer exhibits a high IIP3 of -3dBm and a high isolation of LO-RF about -100dB in the wide bandwidth. It consumes only 6.2mW under a 1.8V power supply.
Xiangrong Zhang, Xiaole Cui, Bo Wang 0016, Chung-Len Lee 0001
ISCAS1
2013 An efficient multiple kernel computation method for regression analysis of economic data
Xiangrong Zhang, Longying Hu
Neurocomputing1
2013 Low-rank representation with local constraint for graph construction
Yaoguo Zheng, Xiangrong Zhang, Shuyuan Yang 0001, Licheng Jiao
Neurocomputing2
2013 Semisupervised Dimensionality Reduction of Hyperspectral Images via Local Scaling Cut Criterion
abstract
Hyperspectral images (HSIs) provide a vast amount of geometrical, radiation, and spectral information about a scene. However, high-dimensional data make HSI classification complex and time consuming. It is important to reduce the dimensionality and find a low-dimensional representation of the high-dimensional data. Since the labels of HSI data are really difficult to collect while the unlabeled data are abundant and easy to obtain, in this letter, a semisupervised dimensionality reduction method using both limited labeled samples and a large number of unlabeled samples based on a local scaling cut (LSC) criterion is proposed. LSC is similar to linear discriminant analysis (LDA), but it can handle the heteroscedastic and multimodal data for which LDA fails. The framework of our proposed method contains two terms: 1) a discrimination term based on the labeled samples and 2) a regularization term based on the prior knowledge provided by both labeled and unlabeled samples. Experimental results show that our proposed algorithm provides a relatively promising performance compared with other methods. Moreover, the algorithm is stable and insensitive to parameters.
Xiangrong Zhang, Yudi He, Yaoguo Zheng
IEEE Geosci. Remote. Sens. Lett.1
2013 Manifold-constrained coding and sparse representation for human action recognition
Xiangrong Zhang, Yang Yang 0062, Licheng Jiao, Feng Dong 0005
Pattern Recognit.1
2013 Robust non-local fuzzy c-means algorithm with edge preservation for SAR image segmentation
Jie Feng 0003, Licheng Jiao, Xiangrong Zhang, Maoguo Gong, Tao Sun 0007
Signal Process.3
2013 Sparse coding and classifier ensemble based multi-instance learning for image categorization
Xiangfa Song, Licheng Jiao, Shuyuan Yang 0001, Xiangrong Zhang, Fanhua Shang
Signal Process.4
2013 Arithmetic coding for image compression with adaptive weight-context classification
Jiaji Wu, Zhenzhen Xu, Gwanggil Jeon, Xiangrong Zhang, Licheng Jiao
Signal Process. Image Commun.4
2013 Context-Based Hierarchical Unequal Merging for SAR Image Segmentation
abstract
This paper presents an image segmentation method named Context-based Hierarchical Unequal Merging for Synthetic aperture radar (SAR) Image Segmentation (CHUMSIS), which uses superpixels as the operation units instead of pixels. Based on the Gestalt laws, three rules that realize a new and natural way to manage different kinds of features extracted from SAR images are proposed to represent superpixel context. The rules are prior knowledge from cognitive science and serve as top-down constraints to globally guide the superpixel merging. The features, including brightness, texture, edges, and spatial information, locally describe the superpixels of SAR images and are bottom-up forces. While merging superpixels, a hierarchical unequal merging algorithm is designed, which includes two stages: 1) coarse merging stage and 2) fine merging stage. The merging algorithm unequally allocates computation resources so as to spend less running time in the superpixels without ambiguity and more running time in the superpixels with ambiguity. Experiments on synthetic and real SAR images indicate that this algorithm can make a balance between computation speed and segmentation accuracy. Compared with two state-of-the-art Markov random field models, CHUMSIS can obtain good segmentation results and successfully reduce running time.
Xiangrong Zhang, Shuang Wang 0001, Biao Hou
IEEE Trans. Geosci. Remote. Sens.2
2012 Multi-objective Invasive Weed Optimization algortihm for clustering
abstract
In this paper, we proposed a new approach to solve the clustering problem in which the cluster number is uncertainty. It utilizes IWO (Invasive Weed Optimization) algorithm to optimize two fuzzy clustering objective function simultaneously, and a variable-length real-coded scheme has been adopted, the variable length weed encodes the cluster centers with variable numbers. In order to keep the diversity of the weeds, we introduce a new mechanism called feedback update mechanism to update the individuals which the corresponding number of cluster centers has been eliminated in one generation. Finally, the Silhouette index is used to select the best solution. The algorithm is used to cluster 15 artificial data sets and 4 real life data sets and shows good performance.
Ruochen Liu 0006, Yangyang Li 0001, Xiangrong Zhang
IEEE Congress on Evolutionary Computation4
2012 Optimized feature extraction by immune clonal selection algorithm
abstract
A new method of feature extraction based on immune clonal selection algorithm is proposed, in which the immune clonal selection algorithm is used to optimize the projection vector. Some orthogonal bases are randomly selected as the initial basis vector sets from the original feature space, and the direction of the basis vectors is optimized to generate the optimal projection vector using the immune clonal selection algorithm. This method provides a new scheme of applying the immune clonal algorithm to feature extraction. Experimental results on benchmark datasets and MSTAR dataset for SAR target recognition verify the effectiveness of the proposed method.
Xiangrong Zhang, Erlei Zhang, Runxin Li
IEEE Congress on Evolutionary Computation1
2012 SAR image change detection based on low rank matrix decomposition
abstract
In this paper we propose an unsupervised approach for SAR image change detection task. A new method based on compressed sensing is applied. First using the PPB method for the speckle reduction, and then the logarithm ratio method is applied to generate a simple change map, and then the compressed sensing-based method is used to part the change map into a low rank part and a sparse part, where the sparse part is correspond to the changed area, finally k-means algorithm is applied to cluster the sparse part into two clusters. Experiment results show the effectiveness and feasibility of the proposed method.
Xiangrong Zhang, Yaoguo Zheng, Jie Feng 0003, Shuiping Gou
IGARSS1
2012 A sparse kernel representation method for image classification
abstract
In this paper, we propose a sparse kernel representation classification algorithm (SKRC) for images classification and recognition. The training dictionary is composed by labeled samples directly, and both training dictionary and testing sample are mapped into feature space from original sample space by the sparse kernel which employs the “center” samples matrix constructed by a method similar to k-means clustering. Then in the feature space, the basic sparse representation based classification method is employed. We test our proposed algorithm on some different public database, and the results show that our proposed method can achieve higher classification accuracy without much time consumed.
Shuyuan Yang 0001, Xiangrong Zhang
IJCNN3
2012 Gene transposon based clone selection algorithm for automatic clustering
Ruochen Liu 0006, Licheng Jiao, Xiangrong Zhang, Yangyang Li 0001
Inf. Sci.3
2012 MPM SAR Image Segmentation Using Feature Extraction and Context Model
abstract
A new synthetic aperture radar (SAR) image segmentation method based on a maximization of posterior marginals (MPM) algorithm with feature extraction and context model is proposed in this letter. First, Gabor wavelet and texture descriptor are used to extract features, which enhance intraclass similarities and interclass differences. Second, the number of regions within the same class is reduced in order to improve the reliability of the regional statistical characteristics. Finally, the MPM of each region combined with the context model is calculated by considering both the intralayer correlation and interlayer correlation. The experimental results show that the proposed method is efficient and effective for SAR image segmentation.
Biao Hou, Xiangrong Zhang
IEEE Geosci. Remote. Sens. Lett.2
2011 MUlti information based Ground Control Points selection method
abstract
Ground Control Points (GCPs) are one of the most important data used in many fields of Remote Sensing. The number and distribution of GCPs are always the key factors for the success of some researches. A GCPs selection method by integrating the three dimensional spatial information (i.e. the earth coordinates (X, Y, Z) ) and the corresponding feature information underlying the data itself was proposed. To testify the performance of this method, a Rational Function Model Resolving experiment is conducted. Experiment results show that the accuracy and time-consuming performance are both improved using GCPs selected by the proposed method.
Yanfeng Gu, Zhimin Cao, Ye Zhang 0008, Xiangrong Zhang
IGARSS4
2011 Spectral clustering based unsupervised change detection in SAR images
abstract
An unsupervised change detection method based on spectral clustering and difference image methods for multitemporal single-channel single-polarization synthetic aperture radar (SAR) images is proposed. The difference image is generated by integrating the typical difference image method with Non-Local Filter, which exploits both the spatial neighborhood information and gray similarity information, and can well reduce the speckle noises of SAR images. The spectral clustering algorithm is employed to cluster the difference image into two clusters and get the change map. Compared with traditional clustering algorithms, such as A-means, SC can recognize the clusters of unusual shapes and obtain the globally optimal solutions. Experimental results confirm the effectiveness of the proposed techniques.
Xiangrong Zhang, Zemin Li, Biao Hou, Licheng Jiao
IGARSS1
2011 Bag-of-Visual-Words Based on Clonal Selection Algorithm for SAR Image Classification
abstract
Synthetic aperture radar (SAR) image classification involves two crucial issues: suitable feature representation technique and effective pattern classification methodology. Here, we concentrate on the first issue. By exploiting a famous image feature processing strategy, Bag-of-Visual-Words (BOV) in image semantic analysis and the artificial immune systems (AIS)'s abilities of learning and adaptability to solve complicated problems, we present a novel and effective image representation method for SAR image classification. In BOV, an effective fused feature sets for local feature representation are first formulated, which are viewed as the low-level features in it. After that, clonal selection algorithm (CSA) in AIS is introduced to optimize the prediction error of k-fold cross-validation for getting more suitable visual words from the low-level features. Finally, the BOV features are represented by the learned visual words for subsequent pattern classification. Compared with the other four algorithms, the proposed algorithm obtains more satisfactory and cogent classification experimental results.
Jie Feng 0003, Licheng Jiao, Xiangrong Zhang
IEEE Geosci. Remote. Sens. Lett.3
2010 Semi-supervised Tissue Segmentation of 3D Brain MR Images
abstract
Clustering algorithms have been popularly applied in tissue segmentation in MRI. However, traditional clustering algorithms could not take advantage of some prior knowledge of data even when it does exist. In this paper, we propose a new approach to tissue segmentation of 3D brain MRI using semi-supervised spectral clustering. Spectral clustering algorithm is more powerful than traditional clustering algorithms since it models the voxel-to-voxel relationship as opposed to voxel-to-cluster relationships. In the semi-supervised spectral clustering, two types of instance-level constraints: must-link and cannot-link as background prior knowledge are incorporated into spectral clustering, and the self-tuning parameter is applied to avoid the selection of the scaling parameter of spectral clustering. The semi-supervised spectral clustering is an effective tissue segmentation method because of its advantages in (1) better discovery of real data structure since there is no cluster shape restriction, (2) high quality segmentation results as it can obtain the global optimal solutions in the relaxed continuous domain by eigen-decomposition and combines the pairwise constraints information. Experimental results on simulated and real MRI data demonstrate its effectiveness.
Xiangrong Zhang, Feng Dong 0005, Gordon Clapworthy, Youbing Zhao, Licheng Jiao
IV1
2008 Unsupervised texture image segmentation using multiobjective evolutionary clustering ensemble algorithm
abstract
Multiobjective evolutionary clustering approach has been successfully utilized in data clustering. In this paper, we propose a novel unsupervised machine learning algorithm namely multiobjective evolutionary clustering ensemble algorithm (MECEA) to perform the texture image segmentation. MECEA comprises two main phases. In the first phase, MECEA uses a multiobjective evolutionary clustering algorithm to optimize two complementary clustering objectives: one based on compactness in the same cluster, and the other based on connectedness of different clusters. The output of the first phase is a set of Pareto solutions, which correspond to different tradeoffs between two clustering objectives, and different numbers of clusters. In the second phase, we make use of the meta-clustering algorithm (MCLA) to combine all the Pareto solutions to get the final segmentation. The segmentation results are evaluated by comparing with three known algorithms: K-means, fuzzy K-means (FCM), and evolutionary clustering algorithm (ECA). It is shown that MECEA is an adaptive clustering algorithm, which outperforms the three algorithms in the experiments we carried out.
Xiaoxue Qian, Xiangrong Zhang, Licheng Jiao, Wenping Ma 0001
IEEE Congress on Evolutionary Computation2
2008 A population-based artificial immune system for numerical optimization
Maoguo Gong, Licheng Jiao, Xiangrong Zhang
Neurocomputing3
2008 Spectral Clustering Ensemble Applied to SAR Image Segmentation
abstract
Spectral clustering (SC) has been used with success in the field of computer vision for data clustering. In this paper, a new algorithm named SC ensemble (SCE) is proposed for the segmentation of synthetic aperture radar (SAR) images. The gray-level cooccurrence matrix-based statistic features and the energy features from the undecimated wavelet decomposition extracted for each pixel being the input, our algorithm performs segmentation by combining multiple SC results as opposed to using outcomes of a single clustering process in the existing literature. The random subspace, random scaling parameter, and Nystrom approximation for component SC are applied to construct the SCE. This technique provides necessary diversity as well as high quality of component learners for an efficient ensemble. It also overcomes the shortcomings faced by the SC, such as the selection of scaling parameter, and the instability resulted from the Nystrom approximation method in image segmentation. Experimental results show that the proposed method is effective for SAR image segmentation and insensitive to the scaling parameter.
Xiangrong Zhang, Licheng Jiao, Fang Liu 0001, Liefeng Bo, Maoguo Gong
IEEE Trans. Geosci. Remote. Sens.1
2008 Quantum-Inspired Immune Clonal Algorithm for Global Optimization
abstract
Based on the concepts and principles of quantum computing, a novel immune clonal algorithm, called a quantum-inspired immune clonal algorithm (QICA), is proposed to deal with the problem of global optimization. In QICA, the antibody is proliferated and divided into a set of subpopulation groups. The antibodies in a subpopulation group are represented by multistate gene quantum bits. In the antibody's updating, the general quantum rotation gate strategy and the dynamic adjusting angle mechanism are applied to accelerate convergence. The quantum not gate is used to realize quantum mutation to avoid premature convergences. The proposed quantum recombination realizes the information communication between subpopulation groups to improve the search efficiency. Theoretical analysis proves that QICA converges to the global optimum. In the first part of the experiments, 10 unconstrained and 13 constrained benchmark functions are used to test the performance of QICA. The results show that QICA performs much better than the other improved genetic algorithms in terms of the quality of solution and computational cost. In the second part of the experiments, QICA is applied to a practical problem (i.e., multiuser detection in direct-sequence code-division multiple-access systems) with a satisfying result.
Licheng Jiao, Yangyang Li 0001, Maoguo Gong, Xiangrong Zhang
IEEE Trans. Syst. Man Cybern. Part B4
2007 SVMs ensemble for radar target recognition based on evolutionary feature selection
abstract
A novel radar target recognition method based on SVMs ensemble is presented, in which a set of suitable feature subsets are selected for component SVMs by Immune Clonal Algorithm, a new artificial immune system algorithm. With Immune Clonal Algorithm, high quality and high diversity of the components for SVMs ensemble are ensured. Experimental results on one-dimension radar high resolution range profiles demonstrate the validity and reliability of this new radar target recognition method.
Xiangrong Zhang, Licheng Jiao, Shuiping Gou
IEEE Congress on Evolutionary Computation1
2007 Recognition of SAR Occluded Targets Using SVM
Licheng Jiao, Weida Zhou, Xiangrong Zhang
MMM (2)5
2005 Radar target recognition using SVMs with a wrapper feature selection driven by immune clonal algorithm
Xiangrong Zhang, Shuang Wang 0001, Tan Shan, Licheng Jiao
ESANN1
2005 Dual ridgelet frame constructed using biorthogonal wavelet basis
abstract
A new system, called dual ridgelet frame, is introduced. The construction of the dual ridgelet frame starts with a dual frame constructed using a biorthogonal wavelet basis in the Radon domain, and, then, the image of the resulting dual frame under an isometric map from the Radon domain to the L/sup 2/(R/sup 2/) spatial domain is a dual frame again, and we call it a dual ridgelet frame. The dual ridgelet frame can be thought of as an extension of the notion of orthonormal ridgelet. It provides a more flexible and effective tool for image analysis and processing applications. The high performance of the dual ridgelet frame for image denoising is demonstrated experimentally.
Tan Shan, Xiangrong Zhang, Licheng Jiao
ICASSP (2)2
2005 A Review: Relationship Between Response Properties of Visual Neurons and Advances in Nonlinear Approximation Theory
Tan Shan, Xiuli Ma, Xiangrong Zhang, Licheng Jiao
ISNN (1)3
2005 Image Representation in Visual Cortex and High Nonlinear Approximation
Tan Shan, Xiangrong Zhang, Shuang Wang 0001, Licheng Jiao
ISNN (1)2