VLDB 2026 Research / reviewers in the wild / expert
Jingjing Ma 0001
dblp:48/3361-1
· DBLP profile ↗
68ranked-venue papers
10as first author
41since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 45 · 9 first-author · 34 since 2021Artificial intelligence and machine learning · 20 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ID-Splat: Propagating Object Identities for Segmenting 3D Aerial-view ScenesabstractHigh-resolution Earth Observation technologies present unprecedented opportunities for geospatial analysis, yet traditional 2D aerial-view semantic segmentation remains limited by its inability to model spatial relationships and handle object occlusions. While 3D Aerial-view Segmentation (3DAS) has emerged to address these limitations, existing methods predominantly rely on 2D discriminative models pre-trained on natural scenes. These models struggle to accurately recognize aerial-view imagery, resulting in suboptimal performance due to significant domain discrepancies. This paper introduces ID-Splat, a novel object-centric framework that directly leverages multi-view object identities without discriminative information to enhance 3D semantic understanding. ID-Splat implements a two-stage process: first, Mask-object Tracking combines SAM and Point Tracking to establish robust and consistent object identities across multi-view aerial images; second, Object Integration & Propagation assigns these identities to 3D Gaussian Splatting (3DGS) points, enabling complete 3D segmentation through semantic propagation. Experimental results on the 3D-AS dataset demonstrate that ID-Splat significantly outperforms existing methods, particularly under sparse supervision conditions. ID-Splat also achieves state-of-the-art performance while reducing the need for extensive labeled data by effectively leveraging the inherent 3D structure. Yijing Wang 0004, Xu Tang 0004, Xiangrong Zhang, Jingjing Ma 0001 |
AAAI | 4 |
| 2026 | Softmatch distance: A novel distance for weakly-supervised trend change detection in bi-temporal images
Yuqun Yang, Xu Tang 0004, Xiangrong Zhang, Changzhe Jiao, Jingjing Ma 0001, Licheng Jiao |
Pattern Recognit. | 5 |
| 2025 | Category-Specific Selective Feature Enhancement for Long-Tailed Multi-Label Image Classification
Ruiqi Du, Xu Tang 0004, Xiangrong Zhang, Jingjing Ma 0001 |
ICCV | 4 |
| 2025 | MLMamba: A Mamba-Based Efficient Network for Multi-Label Remote Sensing Scene ClassificationabstractAs a useful remote sensing (RS) scene interpretation technique, multi-label RS scene classification (RSSC) always attracts researchers’ attention and plays an important role in the RS community. To assign multiple semantic labels to a single RS image according to its complex contents, the existing methods focus on learning the valuable visual features and mining the latent semantic relationships from the RS images. This is a feasible and helpful solution. However, they are often associated with high computational costs due to the widespread use of Transformers. To alleviate this problem, we propose a Mamba-based efficient network based on the newly emerged state space model called MLMamba. In addition to the basic feature extractor (convolutional neural network and language model) and classifier (multiple perceptrons), MLMamba consists of two key components: a pyramid Mamba and a feature-guided semantic modeling (FGSM) Mamba. Pyramid Mamba uses multi-scale scanning to establish global relationships within and across different scales, improving MLMamba’s ability to explore RS images. Under the guidance of the obtained visual features, FGSM Mamba establishes associations between different land covers. Combining these two components can deeply mine local features, multi-scale information, and long-range dependencies from RS images and build semantic relationships between different surface covers. These superiorities guarantee that MLMamba can fully understand the complex contents within RS images and accurately determine which categories exist. Furthermore, the simple and effective structure and linear computational complexity of the state space model ensure that pyramid Mamba and FGSM Mamba will not impose too much computational burden on MLMamba. Extensive experiments counted on three benchmark multi-label RSSC data sets validate the effectiveness of MLMamba. The positive results demonstrate that MLMamba achieves state-of-the-art performance, surpassing existing methods in accuracy, model size, and computational efficiency. Our source codes are available athttps://github.com/TangXu-Group/ multilabelRSSC/tree/main/MLMamba. Ruiqi Du, Xu Tang 0004, Jingjing Ma 0001, Xiangrong Zhang, Licheng Jiao |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Semantic-Assisted Feature Integration Network for Multilabel Remote Sensing Scene ClassificationabstractWith remote sensing (RS) images’ resolution increasing, a single scene label cannot adequately represent RS scenes’ contents. Therefore, multilabel RS scene classification (MLRSSC) is gradually attracting the researchers’ attention. Many methods have been proposed recently, and most use deep features or semantic connections to complete MLRSSC. However, they ignore the combination of these two aspects. In addition, the high interclass similarity and low intraclass similarity of RS images limit the robustness of these methods. In this article, we propose a semantic-assisted feature integration network (SFIN) to overcome the above limitations. It contains a dual-scale feature extractor module (DFEM), a local semantic enhance module (LSEM), a cross-scale interactive attention module (CIAM), and a classifier module (CM). DFEM utilizes the convolutional neural networks (CNNs) to extract multiscale features from RS images. LSEM extracts semantic information and establishes their relationships at different scales. CIAM enhances the feature representation by interacting with the clues across different scales. CM completes the prediction of classification (CLA) results. Integrating them into an end-to-end framework, SFIN can discover the diverse and complex land covers hidden in RS images. Furthermore, to ensure the accuracy of explored semantics and enhance the SFIN’s feature extraction ability, we design a semantic supervision (SS) loss and a semantic-based contrastive learning (SB-CL) loss. They are in charge of the correctness and discrimination of the mined semantics. Along with the typical CLA loss, SFIN can be adequately trained. Extensive experiments have been conducted on four MLRSSC datasets, and the positive results demonstrate that SFIN outperforms many existing methods in MLRSSC tasks. Our source codes are available at:https://github.com/TangXu-Group/multilabelRSSC/tree/main/SFIN. Ruiqi Du, Xu Tang 0004, Jingjing Ma 0001, Xiangrong Zhang, Fang Liu 0034, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | Multiscale Sparse Cross-Attention Network for Remote Sensing Scene ClassificationabstractRemote sensing (RS) scene classification (RSSC) is a prominent research topic in the RS community. Multilevel feature fusion is an important way of addressing RS scene classification, and many methods have been proposed in recent years. Although they succeed, current methods can still be improved, particularly in distinguishing the contributions of different multilevel features and fully and effectively fusing them. To address the above issues and fully exploit the potential of multilevel features for RS scene classification tasks, we propose a new model named multiscale sparse cross-attention network (MSCN). It not only focuses on the effectiveness of feature learning but also emphasizes the rationality of feature fusion. In detail, MSCN first extracts multilevel features using a pre-trained ResNet50. Also, these features are divided into high- and low-level features according to the clues they involved. Then, a multiscale sparse cross-attention (MSC) module is developed to cross-fuse the high-level feature with various low-level features, thereby effectively mining helpful information from multilevel features. In the fusion process, MSC not only explores the multiscale messages in RS scenes but also mitigates the negative impact of irrelevant information by employing sparse operations. Third, a group convolutional block attention module (CBAM) enhancer (GCE) is presented to enhance the representation of classification features. GCE detects local salient information within classification features using grouped CBAM and further enhances crucial details by readjusting the CBAM attention weights. This way, the classification features’ discrimination can be improved. We conducted extensive experiments on three public RS scene classification datasets. The exceptional experimental results indicate that our proposed MSCN achieves superior classification accuracy, surpassing many existing methods. Our source codes are available at https://github.com/TangXu-Group/Remote-Sensing-Images-Classification/tree/main/MSCN. Jingjing Ma 0001, Xu Tang 0004, Xiangrong Zhang, Fang Liu 0034, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2025 | SpiralMamba: Spatial-Spectral Complementary Mamba With Spatial Spiral Scan for Hyperspectral Image ClassificationabstractHyperspectral image (HSI) classification is crucial in the remote sensing (RS) community. In recent years, Transformers have been popular in this field due to their global information modeling capabilities. However, the quadratic complexity limits their performance under limited computational resources. Fortunately, a selective structured state space model named Mamba emerges. Like Transformer, it is good at modeling the long-distance relationships hidden in the pending data. Unlike Transformer, its complexity remains at a linear level. Therefore, a growing number of studies have been proposed to explore the usefulness of Mamba in HSI classification. Nevertheless, most of them only apply Mamba to HSIs directly but do not consider the inherent characteristics of HSIs properly. To exploit the potential of Mamba in HSI classification deeply, this paper presents a new spatial-spectral complementary Mamba with a spatial spiral scan named SpiralMamba. It mainly encloses three main components: a spatial Mamba encoder (SpaME), a spectral Mamba encoder (SpeME), and a spatial-spectral complementary fusion module (SSCFM). SpaME focuses on understanding the spatial context within HSIs. To this end, instead of the common scanning, a spatial spiral scan strategy is introduced to address the sequence transformation of non-causal HSIs. SpeME aims to comprehensively extract valuable spectral features from HSIs. To achieve this goal, besides developing a spectral bidirectional scan strategy, a multilayer convolution (MLC) is also incorporated to capture local variations within spectral tokens. SSCFM concentrates on building the complex connections between spatial and spectral features and fusing them. For this purpose, a relationship learning block (RLB) and a threshold enhancement mechanism (TEM) are developed. Positive experimental results counted on three public HSI datasets demonstrate the effectiveness of SpiralMamba. Our source codes are available at https://github.com/TangXu-Group/Hyperspectral-Images-Classification/tree/main/SpiralMamba. Xu Tang 0004, Yuexi Yao, Jingjing Ma 0001, Xiangrong Zhang, Yuqun Yang, Bo Wang 0016, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | Cross-Modal Remote Sensing Image-Text Retrieval via Context and Uncertainty-Aware PromptabstractThe cross-modal remote sensing image-text retrieval (CMRSITR) is a lively research topic in the remote sensing (RS) community. Benefiting from the large pretrained image-text models, many successful CMRSITR methods have been proposed in recent years. Although their performance is attractive, there are still some challenges. First, fine-tuning large pretrained models requires a significant amount of computational resources. Second, most large models are pretrained by natural images, which reduces their effectiveness in processing RS images. To tackle these challenges, we propose a new CMRSITR network named context and uncertainty-aware prompt (CUP). First, prompt tuning theory is introduced into CUP to eliminate the burden of optimization resources. By training the prompt tokens rather than all parameters, the large model's knowledge can be transferred to CMRSITR tasks with small trainable parameters. Second, considering the differences between natural-image-based prior clues and RS images, apart from adopting the free-prompt tokens, we develop a prompt generation module (PGM) to produce the RS-oriented prompt tokens. The specific prompt tokens are rich in object-level messages of RS images, which help CUP narrow the gaps between natural large models and RS images. Third, we further design an uncertainty estimation module (UEM) to whittle down the uncertainties caused by the model and data. This way, can not only the semantic misalignment and intraclass diversity imbalance problems be mitigated but also the RS clues can be deeply explored. Competitive experimental results counted on three public benchmark datasets demonstrate that our CUP can achieve competitive performance in the CMRSITR task compared with many existing methods. Our source codes are available at: https://github.com/TangXu-Group/Cross-modal-remote-sensing-image-and-text-retrieval-models/tree/main/CUP. Yijing Wang 0004, Xu Tang 0004, Jingjing Ma 0001, Xiangrong Zhang, Fang Liu 0034, Licheng Jiao |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | A Cross-Modal Semantic Mapping Enhancement Model for Remote Sensing Visual Question AnsweringabstractRemote sensing visual question answering (RSVQA) aims to answer the questions based on the content in remote sensing (RS) images. Due to the complexity of RS images, it is challenging to focus on regions relevant to the questions in the RS images. To this end, we propose a channel-selective multi-scale cross-attention (CSCa) model for RSVQA tasks. Specifically, we design a text-driven multi-scale feature extractor to extract question-related features in RS images. To obtain the cross-attention map in this extractor, we design a novel channel selection mechanism to capture more accurate question-related regions in RS images and develop a channel-wise contrastive learning task to align the semantics between image and text features. We set up experiments on RSVQA-LR and RSIVQA datasets. Experiment results show that our CSCa achieves excellent performance. Dabiao Huang, Xu Tang 0004, Xiangrong Zhang, Fang Liu 0034, Jingjing Ma 0001, Licheng Jiao |
IGARSS | 5 |
| 2024 | Multi-Scale Sparse Transformer for Remote Sensing Scene ClassificationabstractVision Transformer (ViT) has achieved great success in the field of computer vision since it was proposed, and there have been many works applying ViT based models to remote sensing scene classification (RSSC) tasks. The proposal of Pyramid Vision Transformer (PVT) greatly reduces the calculation amount of the ViT while maintaining accuracy. But PVT did not utilize multi-scale information in remote sensing (RS) scenes, which is crucial for RSSC. This paper proposes a multi-scale sparse transformer (MST) based on PVT. MST enables the network to learn multi-scale representations of RS scenes through spatial reduction implementations at different scales. In addition, we employ sparse operations to adaptively guide the model’s attention towards semantically relevant regions during self-attention computation, thereby reducing interference from semantically irrelevant areas. Experiments conducted on the UCM and AID datasets demonstrate the outstanding performance of the proposed MST. Xu Tang 0004, Zhixi Feng, Yue Ma 0008, Jingjing Ma 0001, Xiangrong Zhang, Licheng Jiao |
IGARSS | 5 |
| 2024 | Pseudo-Viewpoint Regularized 3D Gaussian Splatting For Remote Sensing Few-Shot Novel View SynthesisabstractIn remote sensing (RS), Few-Shot Novel View Synthesis (FS-NVS) focuses on creating images of unobserved viewpoints using limited training images. Recently, 3D Gaussian Splatting (3DGS) has drawn scholars’ attention by its increasing rendering speeds and providing an explicit neural representation for 3D scenes. However, 3DGS tends to overfit limited training data. To tackle this challenge, we propose a Pseudoview Regularized 3DGS (PR3DGS) FSNVS method for RS scenarios. Our PR3DGS method introduces a pseudo-views regularization module to discriminate synthetic RS images generated from training- or pseudo-viewpoints. Therefore, our PR3DGS method can effectively mitigate overfitting in seen views and enhance the model’s capability to generate more realistic RS images from novel viewpoints. Besides, the excellent experimental results on the LEVIR-NVS dataset demonstrate the effectiveness of our method in RS FSNVS. Yijing Wang 0004, Xu Tang 0004, Jingjing Ma 0001, Xiangrong Zhang, Fang Liu 0001, Licheng Jiao |
IGARSS | 3 |
| 2024 | Time-Guided Network for Remote Sensing Change DetectionabstractIn recent years, as a very important part of remote sensing interpretation, change detection (CD) has developed rapidly with deep learning and remote sensing interpretation. However, most of the existing CD methods focus on how to extract the spatial features of the bi-temporal image pairs, ignore the importance of the temporal features. To solve this problem, a new time-guided network (TG-Net) using temporal information to guide feature extraction is proposed in this paper. In TG-Net, we use a newly proposed time-guided feature fusion (TGF2) block that uses temporal information to guide spatial feature fusion to extract temporal and spatial information comprehensively. We conducted experiments on two publicly available remote sensing datasets, LEVIR-CD and WHU, and compared them with three common CD methods. Xu Tang 0004, Xiangrong Zhang, Jingjing Ma 0001, Fang Liu 0001, Licheng Jiao |
IGARSS | 4 |
| 2024 | Center Mask Self-Attention Network for Hyperspectral Image ClassificationabstractBenefiting from the thousands of continuous band information in hyperspectral images (HSIs), the task of HSI classification has become an indispensable part of the field of remote sensing. With the development of deep learning, deep learning techniques such as convolutional neural networks have been widely introduced into HSI classification research. However, most of these methods do not fully consider the potential relationship between the central pixel and surrounding neighborhoods. Therefore, we introduce a novel center mask self-attention network (CMSAN) to enable the model to effectively capture the association between the central pixel and its neighbors for better feature extraction. We conduct experiments on two publicly available HSI datasets. The positive results on both datasets fully demonstrate the effectiveness of our proposed method. Yizhou Zou, Xu Tang 0004, Yue Ma 0008, Jingjing Ma 0001, Xiangrong Zhang, Licheng Jiao |
IGARSS | 4 |
| 2024 | EATDer: Edge-Assisted Adaptive Transformer Detector for Remote Sensing Change DetectionabstractChange detection (CD) is one of the important research topics in remote sensing (RS) image processing. Recently, convolutional neural networks (CNNs) have dominated the RSCD community. Many successful CNN-based models have been proposed, and they achieved cracking performance. Nevertheless, influenced by the limited receptive field, the CNN-based models are not good at capturing long-distance context dependencies within RS images, negatively impacting their performance. With the appearance of the visual transformer, the above problems have been mitigated. However, the high time costs of the transformer-based models limit their applicability. In addition, previous CD networks (whether CNN-based or transform-based) do not pay attention to the edges of changed areas, reducing the quality of change maps. To overcome the shortcomings discussed above, we propose a new CD method named edge-assisted adaptive transformer detector (EATDer). EATDer consists of a Siamese encoder and an edge-aware decoder. Each branch in the Siamese encoder encloses three self-adaption vision transformer (SAVT) blocks, which aim to capture the local and global information within RS images. Also, two branches are connected by full-range fusion modules (FRFMs), which focus on mining the temporal clues among bi-temporal RS images and pointing out the changed/unchanged messages. The edge-aware decoder first integrates the multiscale features obtained by the encoder using a restoring block. Then, it enhances the combined features by a refining block. Finally, based on the refined features, both the change and edge detection results can be produced. Along with a joint loss function, we can get high-quality change maps in which the changed areas are correct and have clear and smooth edges. The usefulness of our EATDer is validated by extensive experiments conducted on three popular RSCD datasets. Our source codes are available athttps://github.com/TangXu-Group/Remote-Sensing-Image-Change-Detection/tree/main/EATDer Jingjing Ma 0001, JunYi Duan, Xu Tang 0004, Xiangrong Zhang, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | Spatial Pooling Transformer Network and Noise-Tolerant Learning for Noisy Hyperspectral Image ClassificationabstractHyperspectral image (HSI) classification is a hot topic in remote sensing. A large number of studies have been proposed and achieved excellent performance. Most of them rely on accurate annotations. However, this requirement cannot always be met. Due to the complex contents within HSIs and the uncontrollable external interference factors, incorrect labels are inevitable. Thus, the study of noisy HSI classification is boomed. Some attempts have been made, and their central ideas are to filter the noisy samples from the training set. Although feasible, this would result in information loss, i.e., the contents covered by the removed samples are ignored. Besides, the characteristics of HSIs are not fully considered in many models. To overcome the above limitations, we develop a spatial pooling transformer network (SPTNet) and a noise-tolerant learning algorithm in this paper. SPTNet first uses a spectral feature extraction (SFE) module to capture the rich spectral information from HSI patches. Then, three spatial pooling transformers (SPTs) are constructed and stacked to explore the spatial knowledge and depress confusing clues caused by the HSI patch division. Finally, a standard transformer encoder is used to enhance the obtained spectral-spatial features for the downstream classification. To use SPTNet to handle noisy HSI classification, the noise-tolerant learning algorithm is designed. It encloses two parts, i.e., a data partition scheme and a label-independent similarity regularization. The data partition scheme divides the training data into clean and noisy sets. Then, the clean samples are used to train SPTNet with the classification loss function. At the same time, similarity regularization helps SPTNet to comprehensively understand HSIs by analyzing the resemblance between clean and noisy samples. Integrating two parts into a co-training framework, SPTNets can be trained under a noisy scenario. Four popular HSI datasets are selected to testify to our methods. The positive results demonstrate that the combination of SPTNet and the noise-tolerant learning algorithm is helpful to the noisy HSI classification. Our source codes are available at https://github.com/TangXu-Group/Hyperspectral-Images-Classification/tree/main/SPTNet-NTLA. Jingjing Ma 0001, Yizhou Zou, Xu Tang 0004, Xiangrong Zhang, Fang Liu 0034, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | Prior-Experience-Based Vision-Language Model for Remote Sensing Image-Text RetrievalabstractRemote sensing (RS) image-text retrieval (RSITR) aims to retrieve relevant texts (RS images) based on the content of a given RS image (text). Existing methods are used to employing the convolutional neural network (CNN) and recurrent neural network (RNN) as encoders to learn visual and textual features for retrieval. Although feasible, the global information hidden in different data does not receive the attention it deserves. To mitigate this problem, transformers have been introduced. Nevertheless, the complexity of RS images present challenges in directly introducing Transformer-based architectures to multimodal learning in RS scenes, particularly in visual feature extraction and cross-modal interaction. In addition, the textual captions are always simpler than the complex RS images, leading to a semantic description appearing in different images. This typical false-negative (FN) sample problem increases the difficulty of RSITR tasks. To address the above limitations, we propose a new RSITR model named prior-experience-based RS vision-language (PERSVL). First, the specific visual and text encoders are used to extract features from RS images and texts. Also, a high-level feature complement (HFC) module is developed based on the self-attention mechanism (SAM) for the visual encoder to explore the complex contents from RS images fully. Second, a dual-branch multimodal fusion encoder (DBMFE) is designed to complete the cross-modal learning. It comprises a dual-branch multimodal interaction (DBMI) module and a branch fusion module. DBMI is designed to fully explore the relationships between different modalities, enriching visual and textual features. The branch fusion module integrates the cross-modal features and utilizes a classification head to generate matching scores for retrieval. Finally, a learning from prior experiences (LPEs) module is designed to reduce the influence of FN samples by analyzing the historical data produced in the model training process. Experiments are conducted on three popular datasets, and the positive results show that our PERSVL model achieves superior performance compared with previous methods. By integrating the advantages of natural language and RS images, our PERSVL can be applied in various applications, such as environmental monitoring, disaster evaluation, and urban planning. Our source codes are available at:https://github.com/TangXu-Group/Cross-modal-remote-sensing-image-and-text-retrieval-models/tree/main/PERSVL. Xu Tang 0004, Dabiao Huang, Jingjing Ma 0001, Xiangrong Zhang, Fang Liu 0034, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Multiple Information Collaborative Fusion Network for Joint Classification of Hyperspectral and LiDAR DataabstractJoint classification of hyperspectral image (HSI) and light detection and ranging (LiDAR) can simultaneously utilize rich spectral information and elevation information and has become a hot research topic in remote sensing (RS). Although many works have been proposed for this task, their performance cannot reach what we expected due to inadequate cross-modal feature learning and simple feature fusion. This article proposes a multiple information collaborative fusion network (MICF-Net) to overcome those limitations, which aims to leverage the essentially consistent spatial relationships and high-level semantic information in multimodal data to guide the extraction of multimodal fusion features. Specifically, MICF-Net first uses a simple two-branch convolutional neural network (CNN) for preliminary feature extraction. Then, a dual-branch cross-modal attention fusion transformer (CMAFT) is developed to mine global contextual content. By fusing the attention maps of two modalities and limiting their similarity, CMAFT can retain modality-specific information while achieving information interaction based on spatial relationships. Next, an adaptive mask modulation (AMM) module is designed to dynamically balance the learning rate of each modality to ensure the effectiveness of the features of all modalities. Finally, to mine the complementary information of HSI and LiDAR data, a semantic-guided feature fusion (SGFF) module is introduced. It achieves mutual guided learning by exchanging semantic information between two modalities. Positive experimental results counted on three popular HSI and LiDAR datasets demonstrate the effectiveness of the proposed MICF-Net. Our source codes are available athttps://github.com/TangXu-Group/Hyperspectral-Images-Classification/tree/main/MICF-Net. Xu Tang 0004, Yizhou Zou, Jingjing Ma 0001, Xiangrong Zhang, Fang Liu 0034, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | ECPS: Cross Pseudo Supervision Based on Ensemble Learning for Semi-Supervised Remote Sensing Change DetectionabstractSemi-supervised learning aims to exploit the potential of unlabeled data to enhance model performance, which makes it suitable for addressing the challenge of limited labeled data. As a popular technology, pseudo-label is widely applied in many semi-supervised remote sensing (RS) change detection methods. However, when facing limited labeled data, abundant low-quality pseudo-labels from a poorly-performing model hinder the effective enhancement of model performance. To address this issue, we propose a novel semi-supervised strategy, named ensemble cross pseudo supervision (ECPS). The utilization of ensemble learning to merge outputs from several change detection models enhances pseudo-label quality, leading to more accurate change information and a significant boost in model performance, even with limited labeled data. In this method, adopting crosswise supervision ensures that no additional inference costs caused by ensemble learning are consumed. This provides both high efficiency and effectiveness for identifying land-cover changes. On the other hand, a simple yet effective ensemble strategy is proposed, which allows to manually adjust the model’s tendency towards higher precision or recall for satisfying practical requirements. We conduct extensive experiments on four public RS change detection datasets, and the promising results demonstrate the superiority of the proposed method across various numbers of labeled samples. Our source codes are available at https://github.com/TangXu-Group/ECPS. Yuqun Yang, Xu Tang 0004, Jingjing Ma 0001, Xiangrong Zhang, Shiji Pei, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | FDLdet: A Change Detector Based on Forward Dictionary Learning for Remote Sensing ImagesabstractAs an important topic in the remote sensing (RS) image processing community, change detection has attracted much attention from researchers, which aims to distinguish land-cover changes in a geographic position. This is a challenging task because the visual representations of land cover captured from RS images at different periods would vary widely and considerably, resulting in significant differences in feature representations. To alleviate this problem, many existing deep-based methods employ the parameter-shared strategy to map RS images into a common feature space for detecting the changes. Although they are feasible, the simple and single visual information learned by deep models is still not sophisticated enough for satisfactory results. To address this problem, we propose a forward dictionary learning (DL) model named forward DL detector (FDLdet) in this article. Besides the common visual features, our FDLdet takes into account the essential information, e.g., element composition and land-cover category, for change detection. FDLdet consists of a feature extractor, a coefficient generator, and a deep dictionary. Specifically, first, the feature extractor is used to extract shared deep features from RS images. Second, the coefficient generator transforms these deep features into word coefficients. Third, words within the deep dictionary are combined by word coefficients to generate the dictionary features with essential information. Finally, the dictionary features are used instead of deep features to detect land-cover changes. Extensive experiments are conducted on two public large-scale datasets, i.e., season-varying change detection (SVCD), Sun Yat-sen University change detection (SYSU-CD), and LEVIR change detection (LEVIR-CD). Experimental results demonstrate the effectiveness of the proposed FDLdet. Our source codes are available athttps://github.com/TangXu-Group/FDLdet. Yuqun Yang, Xu Tang 0004, Xiangrong Zhang, Jingjing Ma 0001, Yiu-Ming Cheung, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Semi-Supervised Multiscale Dynamic Graph Convolution Network for Hyperspectral Image ClassificationabstractIn recent years, convolutional neural networks (CNNs)-based methods achieve cracking performance on hyperspectral image (HSI) classification tasks, due to its hierarchical structure and strong nonlinear fitting capacity. Most of them, however, are supervised approaches that need a large number of labeled data to train them. Conventional convolution kernels are fixed shape of rectangular with fixed sizes, which are good at capturing short-range relations between pixels within HSIs but ignore the long-range context within HSIs, limiting their performance. To overcome the limitations mentioned above, we present a dynamic multiscale graph convolutional network (GCN) classifier (DMSGer). DMSGer first constructs a relatively small graph at region-level based on a superpixel segmentation algorithm and metric-learning. A dynamic pixel-level feature update strategy is then applied to the region-level adjacency matrix, which can help DMSGer learn the pixel representation dynamically. Finally, to deeply understand the complex contents within HSIs, our model is expanded into a multiscale version. On the one hand, by introducing graph learning theory, DMSGer accomplishes HSI classification tasks in a semi-supervised manner, relieving the pressure of collecting abundant labeled samples. Superpixels are generally in irregular shapes and sizes which can group only similar pixels in a neighborhood. On the other hand, based on the proposed dynamic-GCN, the pixel-level and region-level information can be captured simultaneously in one graph convolution layer such that the classification results can be improved. Also, due to the proper multiscale expansion, more helpful information can be captured from HSIs. Extensive experiments were conducted on four public HSIs, and the promising results illustrate that our DMSGer is robust in classifying HSIs. Our source codes are available at https://github.com/TangXu-Group/DMSGer. Yuqun Yang, Xu Tang 0004, Xiangrong Zhang, Jingjing Ma 0001, Fang Liu 0034, Xiuping Jia, Licheng Jiao |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2023 | Exchange Data Augmentation for Change DetectionabstractChange detection is developed to automatically identify semantic changes between remote sensing (RS) images captured at different points of time in a specific geographic location. Due to the limited availability of annotated data and the high complexity of the change detection problem, the performance of change detection models can not meet our expectations. To address this issue, many methods are proposed to solve this issue. However, ignoring the characteristics of the change detection task limits their performance. Therefore, we introduce a novel exchange data enhancement method (EDEM) strategy to generate image pairs to help the neural network to capture the temporal consistency in the data. We evaluate the proposed approach on two publicly available datasets and compare it with several state-of-the-art methods. The experimental results demonstrate that our proposed approach can effectively improve the performance of change detection models, achieving state-of-the-art performance on both datasets. JunYi Duan, Yijing Wang 0004, Xu Tang 0004, Jingjing Ma 0001, Yuqun Yang |
IGARSS | 4 |
| 2023 | Multi-Scale Interaction Prototypical Network For Few-Shot Remote Sensing Scene ClassificationabstractFew-shot remote sensing scene classification (FSRSSC) aims to make the model quickly adapt to new scenes with a small amount of annotation data. The large intra-class variance and high inter-class similarity in remote sensing (RS) scenes make this task more challenging. To this end, we propose a multi-scale interaction prototypical network, which pays attention to capturing multi-scale information of images during model learning, and then generates a prototype representation of mixed query information through a feature interaction module, thereby enhancing the rapid learning ability of the model, so as to reducing of the within-class and between-class variance ratio in RS scenes. The positive experimental results on UC-Merced and NWPU datasets demonstrate the effectiveness of our model in FSRSSC. Shiji Pei, Yijing Wang 0004, Jingjing Ma 0001, Xu Tang 0004, Yuqun Yang |
IGARSS | 3 |
| 2023 | Unsupervised SAR Image Change Detection Based on Feature Fusion of Information TransferabstractSynthetic aperture radar (SAR) image change detection is a hot but challenging task due to SAR images’ complex contents and inherent speckle noises. The expected change detection methods should reduce the influence of speckle noises, obtain the discriminative feature representations, and generate accurate change maps simultaneously. To these ends, we propose a new SAR image change detection method named feature fusion of information transfer network (FFITN). First, we develop a hybrid convolution block to depress the speckle noise impacts and explore the valuable information from SAR images. Thus, the feature extraction module (FEM) is constructed to obtain the multi-level features. Then, an information transfer module (ITM) is proposed to capture the salient regions from various aspects. Also, the salient knowledge is transferred among features at different levels to enhance their discrimination. Next, a self-attention-based feature fusion module (SAFFM) is introduced to fuse various features. Finally, a change map generation module (CMGM) with the clustering algorithm and specific loss functions is designed to produce the pseudo labels and change maps. Experimental results on three public SAR data sets demonstrate the model’s effectiveness. Our source codes are available at https://github.com/TangXu-Group/FFITN. Jingjing Ma 0001, Xu Tang 0004, Yuqun Yang, Xiangrong Zhang, Licheng Jiao |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2023 | Multipretext-Task Prototypes Guided Dynamic Contrastive Learning Network for Few-Shot Remote Sensing Scene ClassificationabstractAs a content management technique, remote sensing (RS) scene classification (RSSC) always attracts researchers’ attention. In the past decades, many successful methods have been proposed. Nevertheless, their prerequisite is that there are large labeled data sets, which is a strict demand in practice. To resolve this contradiction, developing RSSC models with the help of few-shot learning (FSL) has become popular. Due to lacking prior knowledge, most of the existing few-shot RSSC models pay attention to the learning algorithm. However, they do not attach importance to the complex contents within RS scenes and the intricate inter-/intra-class relations between RS scenes. This would influence their performance negatively. In this paper, we propose a new few-shot RSSC model named multi-pretext-task prototypes guided dynamic contrastive learning network (MPCL-Net). MPCL-Net consists of a multi-pretext tasks generation sub-module, a deep feature learning sub-module, and a joint optimization sub-module. First, two RS-oriented pretext tasks are constructed under the self-supervised learning (SSL) framework in the multi-pretext tasks generation sub-module, which aim to explore multi-scale and rotation-invariant information from RS scenes. Second, a simple convolutional neural network (CNN) is developed in the deep feature learning sub-module to transform the RS scenes into visual features. Third, three loss functions are formulated and integrated in the joint optimization sub-module. Their goals are to fully capture the diverse land covers within RS scenes and compact/separate the intra-/inter-class samples with limited supervision. Finally, our MPCL-Net can be trained in a meta way. The positive results counted on the three public RS scene data sets confirm that our MPCL-Net is helpful to RSSC tasks under the few-shot scenario. Our source codes are available at https://github.com/TangXu-Group/Remote-Sensing-Images-Classification/tree/main/MPCL. Jingjing Ma 0001, Weiquan Lin, Xu Tang 0004, Xiangrong Zhang, Fang Liu 0001, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | Interacting-Enhancing Feature Transformer for Cross-Modal Remote-Sensing Image and Text RetrievalabstractCross-modal remote sensing image-text retrieval (CMRSITR) is a challenging topic in the remote sensing (RS) community. It has gained growing attention because it can be flexibly used in many practical applications. In the current deep era, with the help of deep convolutional neural networks (DCNNs), many successful CMRSITR methods have been proposed. Most of them first learn valuable features from RS images and texts respectively. Then, the obtained visual and textual features are mapped into a common space for the final retrieval. The above operations are feasible, however, two difficulties are still to be solved. One is that the semantics within the visual and textual features are misaligned due to the independent learning manner. The other one is that the deep links between RS images and texts cannot be fully explored by simple common space mapping. To overcome the above challenges, we propose a new model named interacting-enhancing feature transformer (IEFT) for CMRSITR, which regards the RS images and texts as a whole. First, a simple feature embedding module (FEM) is developed to map images and texts into the visual and textual feature spaces. Second, an information interacting-enhancing module (IIEM) is designed to simultaneously model the inner relationships between RS images and texts and enhance the visual features. IIEM consists of three feature interacting-enhancing (FIE) blocks, each of which contains an inter-modality relationship interacting (IMRI) sub-block and a visual feature enhancing (VFE) sub-block. The duty of IMRI is to exploit the hidden relations between cross-modal data, while the responsibility of VFE is to improve the visual features. By combining them, semantic bias can be mitigated, and the complex contents of RS images can be studied. Finally, the retrieval module (RM) is constructed to generate the matching scores for deciding the search results. Extensive experiments are conducted on four public RS data sets. The positive results demonstrate that our IEFT can achieve superior retrieval performance compared with many existing methods. Our source codes are available at https://github.com/TangXu-Group/Cross-modal-remote-sensing-image-and-text-retrieval-models/tree/main/IEFT. Xu Tang 0004, Yijing Wang 0004, Jingjing Ma 0001, Xiangrong Zhang, Fang Liu 0034, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | WNet: W-Shaped Hierarchical Network for Remote-Sensing Image Change DetectionabstractChange detection (CD) is a hot research topic in the remote sensing (RS) community. With the increasing availability of high-resolution (HR) RS images, there is a growing demand for CD models with high detection accuracy and generalization ability. In other words, the CD models are expected to work well for various HRRS images. Convolutional neural networks (CNNs) have been dominated in HRRS image CD due to their excellent information extraction and nonlinear fitting capabilities. However, they are not skilled in modeling long-range contexts hidden in HRRS images, which limits their performance in CD tasks more or less. Recently, the Transformer, which is good at extracting global context dependencies, has become popular in the RS community. Nevertheless, detailed local knowledge receives insufficient emphasis in common Transformers. Considering the above discussion, we combine CNN and Transformer and propose a new W-shaped dual Siamese branch hierarchical network for HRRS image CD named WNet. WNet first incorporates a Siamese CNN and a Siamese Transformer into a dual-branch encoder to extract multi-level local fine-grained features and global long-range contextual dependencies. Also, we introduce deformable ideas into the Siamese CNN and Transformer to make WNet understand the critical and irregular areas within HRRS images. Second, the difference enhancement module (DEM) is developed and embedded into the encoder to produce the difference feature maps at different levels. Using simple pixel-wise subtraction and channel-wise concatenation, the changes of interest and irrelevant changes can be highlighted and suppressed in a learnable manner. Next, the multi-level difference feature maps are fused stage by stage by CNN-Transformer fusion modules (CTFMs), which are the basic units of the decoder in WNet. In CTFM, the local, global, and cross-scale clues are taken into account to ensure the integrity of information. Finally, a simple classifier is constructed and added at the top of the decoder to predict the change maps. Positive experimental results counted on four public datasets demonstrate that the proposed WNet is helpful in HRRS image CD tasks. Our source codes are available at https://github.com/TangXu-Group/Remote-Sensing-Image-Change-Detection/tree/main/WNet. Xu Tang 0004, Jingjing Ma 0001, Xiangrong Zhang, Fang Liu 0001, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Resformer: Bridging Residual Network and Transformer for Remote Sensing Scene ClassificationabstractRemote sensing (RS) scene classification is a crucial research topic in the RS community, and many convolutional neural networks (CNNs)-based methods have been proposed to improve classification performance. Due to the intrinsic locality of convolution operations, CNNs are good at extracting local information but are not easy to capture global contextual information which is also important to fully interpret RS scenes. Recently, transformer has shown the potential for learning global contextual information, but it pays less attention to local information. In this paper, we propose a new interactive dual-branch network for RS scene classification, named Resformer, which can use CNNs's efficiency in extracting local information as well as transformer's power in capturing global information. Besides, we propose a two-way feature interaction module (TFIM), which can not only efficiently fuse CNNs-based local features with transformer-based global fetures, but also extract multi-scale information from RS scenes. Finally, we use a class score fusion strategy to integrate the features extracted from the two branches. Encouraging experimental results counted on two public RS scene data sets demonstrate that our Resformer is effective in RS scene classification task. Mingteng Li, Jingjing Ma 0001, Xu Tang 0004, Xiao Han 0012, Licheng Jiao |
IGARSS | 2 |
| 2022 | NQ-Protonet: Noisy Query Prototypical Network for Few-Shot Remote Sensing Scene ClassificationabstractFew-shot remote sensing scene classification, which aims to recognize unseen classes given only a few labeled samples, is a challenge task due to the complex content contained in remote sensing scenes. In this end, we propose a noisy query prototypical network (NQ-ProtoNet), which uses query-mix module (QM) to produce extra query samples with inter-ference information for classification and thus implicitly enhance the feature learning ability of model. Our method alleviates the problem of large intraclass variances and inter-class similarity of remote sensing scenes to some extent, and the positive experimental results on UC Merced and NWPU data sets show that it outperforms several few-shot learning methods. Weiquan Lin, Xu Tang 0004, Xiao Han 0012, Jingjing Ma 0001, Xiangrong Zhang, Licheng Jiao |
IGARSS | 4 |
| 2022 | Multi-Scale Interactive Transformer for Remote Sensing Cross-Modal Image-Text RetrievalabstractCross-modal Remote sensing (RS) image-text retrieval (CMR-SITR) plays a crucial role in the RS community. A common way for CMRSITR is to extract texts and RS images' feature representations separately and then measure their similarities in the specific or common feature space. Recently, along with the booming of deep convolutional neural networks (DCNNs), these kinds of methods are vivid and achieve successes in their own applications. However, they neglect the inherent relationships between different features, and they are always heavy. To overcome the limitations mentioned above, we propose a new model for CMRSITR in this paper, named multi-scale interactive transformer (MSIT). MSIT first adopts simple feature learning models for texts and RS images which could ensure the whole model is not heavy. Then, MSIT introduces transformer encoders to enhance features' usefulness by considering the potential relations between different representations. Also, a lightweight multi-scale feature learning module is proposed to mine more plentiful contents from RS images. Finally, instead of outputting the features, MSIT produces matching scores for texts and RS images, which can be used to decide the retrieval results directly. The experimental results on two RS datasets indicate our modal is effective for CMRSITR. Yijing Wang 0004, Jingjing Ma 0001, Mingteng Li, Xu Tang 0004, Xiao Han 0012, Licheng Jiao |
IGARSS | 2 |
| 2022 | Remote Sensing Image Change Detection Based on Deep Dictionary LearningabstractAs a hot topic in the field of remote sensing (RS), change detection aims to identify the semantic change between bitemporal RS images. Due to the semantic complexity of RS images, how to accurately detect the semantic change has become a challenging problem. Recently, many deep-based methods are proposed to solve this issue. However, ignoring the representation difference of same semantics in different periods limits their performance, such as river is liquid in summer and solid in winter. Therefore, a new method is presented, named dictionary learning based change detector (DLCDet), which consists of feature pyramid network, deep dictionary learning and dual supervision modules. In DLCDet, the deep dictionary learning is proposed to reduce the representation difference so that DLCDet identifies the potential semantic change more accurately. Experiments are conducted on two public datasets change detection dataset (CDD) and building change detection dataset (BCDD), which demonstrates the effectiveness of the proposed method. Yuqun Yang, Xu Tang 0004, Fang Liu 0001, Jingjing Ma 0001, Licheng Jiao |
IGARSS | 4 |
| 2022 | Region NMS-based deep network for gigapixel level pedestrian detection with two-step cropping
Lingling Li 0002, Xiaohui Guo, Jingjing Ma 0001, Licheng Jiao, Fang Liu 0001, Xu Liu 0006 |
Neurocomputing | 4 |
| 2022 | Class-Level Prototype Guided Multiscale Feature Learning for Remote Sensing Scene Classification With Limited LabelsabstractRemote sensing scene classification (RSSC) is an open and challenging research topic in the remote sensing (RS) community. It aims to define semantic labels for RS scenes according to their contents. Recently, with the development of deep convolutional neural networks (DCNNs), the results of RSSC have been enhanced to a large extent. However, the cracking performance of these DCNN-based models depends on a large number of labeled data. Once the volume of the labeled data is decreased, their behavior would be weakened dramatically. In this article, we propose a new training algorithm that can work smoothly with a few labeled samples to address this limitation. Along with the introduced DCNN, the presented methods perform satisfactorily. In particular, we first construct a dual-branch network (DBNet) to mine the multiscale and multiangle information from RS scenes. Thus, the abundant land covers with diverse sizes, directions, and shapes can be captured simultaneously. Then, to train DBNet using scarce semantic labels, a class-level prototype guided learning (CPGL) algorithm is developed based on the meta-learning paradigm. Besides the usual episode training manner, a prototype refinement module (PRM) and a prototype discrimination module (PDM) are designed with the help of metric learning theory to ensure the effectiveness of our CPGL. The comprehensive experiments are conducted on four public RS scene datasets, and the encouraging results imply that our DBNet and CPGL can copy with RSSC tasks with small labeled data. Our source codes are available athttps://github.com/TangXu-Group/Remote-Sensing-Images-Classification/tree/main/CPGL. Xu Tang 0004, Weiquan Lin, Jingjing Ma 0001, Xiangrong Zhang, Fang Liu 0034, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | EMTCAL: Efficient Multiscale Transformer and Cross-Level Attention Learning for Remote Sensing Scene ClassificationabstractIn recent years, convolutional neural network (CNN)-based methods have been widely used for remote sensing (RS) scene classification tasks and achieved excellent results. However, CNNs are not good at exploring contextual information, which is essential for fully understanding RS scenes. A new model named transformer attracts researchers’ attention to address this problem, which is skilled in mining the latent contextual information in RS scenes. Nevertheless, since the contents of RS scenes are diverse in type and various in scale, the performance of the original transformer in RS scene classification cannot reach what we expect. In addition, due to the specific self-attention mechanism, the time costs of the transformer are high, which hinders its practicability in the RS community. To overcome the above limitations, we propose a new model named efficient multi-scale transformer and cross-level attention learning (EMTCAL) for RS scene classification in this paper. EMTCAL combines the advantages of CNN and transformer to mine information within RS scenes fully. First, it uses a multi-layer feature extraction module (MFEM) to acquire global visual features and multi-level convolutional features from RS scenes. Second, a contextual information extraction module (CIEM) is proposed to capture rich contextual information from multi-level features. In CIEM, taking the characteristics of RS scenes and the computational complexity into account, we propose an efficient multi-scale transformer (EMST). EMST can mine the abundant knowledge with various scales hidden in RS scenes and model their inherent relations at small-time costs. Third, a cross-level attention module (CLAM) is developed to aggregate and explore correlations of multi-level features. Finally, a class score fusion module (CSFM) is designed to integrate the contributions of global and aggregated multi-level features for the discriminative scene representations. Extensive experiments are conducted on three public RS scene data sets. The positive results demonstrate that our EMTCAL can achieve superior classification performance and outperform many state-of-the-art methods. Xu Tang 0004, Mingteng Li, Jingjing Ma 0001, Xiangrong Zhang, Fang Liu 0001, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Meta-Hashing for Remote Sensing Image RetrievalabstractWith the explosive growth of the volume and resolution of high-resolution remote-sensing (HRRS) images, the management of them becomes a challenging task. The traditional content-based remote-sensing image retrieval (CBRSIR) technologies cannot meet what we expect due to the large volume of image archives and complex contents within HRRS images. As a successful approximate nearest neighborhood (ANN) search technique, Hash learning has received wide attention, especially when deep convolutional neural networks (DCNNs) appear. Due to DCNNs’ strong capacity of feature learning, many DCNN-based hashing methods have been proposed and achieved good performance for large-scale CBRSIR tasks. Nevertheless, their limitation is that a large of labeled training samples should be collected for training the deep models. To overcome this limitation, this article, therefore, develops a new supervised hash learning method for the large-scale HRRS CBRSIR task based on meta-learning, which could achieve well-retrieval performance with a few labeled training samples. First, taking the characteristics of HRRS into account, we develop a self-adaptive convolution (SAP-Conv) block and design a hashing net based on the block. SAP-Conv can learn robust features from HRRS images by exploring their multiscale information. Second, to enhance the generalization of the hashing net under a few labeled training samples, the hash learning is formulated in a meta-way, and we name it meta-hashing. Meta-hashing can effectively preserve the similarities between support and query set, and the similarities between samples within support set by the developed loss function. To further improve the performance of meta-hashing, we expand it to a dynamic version named dynamic-meta-hashing, in which the numbers of support and query are changeable in the training phase. Experimental results counted on the three widely used HRRS datasets demonstrate our dynamic-meta-hashing and meta-hashing can achieve promising performance in large-scale HRRS CBRSIR tasks based on a few training samples. Our source codes are available athttps://github.com/TangXu-Group/Meta-hashing. Xu Tang 0004, Yuqun Yang, Jingjing Ma 0001, Yiu-Ming Cheung, Chao Liu 0042, Fang Liu 0034, Xiangrong Zhang, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | AR2Det: An Accurate and Real-Time Rotational One-Stage Ship Detector in Remote Sensing ImagesabstractShip detection plays a significant role in the high-resolution remote sensing (HRRS) community, but it is a challenging task due to the complex contents within HRRS images and the diverse orientation of ships. Recently, with the development of deep learning, the performance of the HRRS ship detection model has been improved greatly. Most of them employ deep networks and complicate anchor mechanism to get well ship detection results. Nevertheless, this kind of combination limits the detection efficiency. To address this problem, a new approach named accurate and real-time rotational ship detector (AR2Det) is proposed in this article to detect ships without the anchor mechanism. Based on the extracted features by the feature extraction module (FEM) and the central information of ships, AR2Det adopts two simple modules, ship detector (SDet) and center detector (CDet), to generate and improve the detection results, respectively. AR2Det is efficient due to the simple postprocessing and the lightweight network. Also, AR2Det performs satisfactorily due to the effective generation and enhancement strategy of bounding boxes. The extensive experiments are conducted on a public HRRS image ship detection dataset HRSC2016. The promising results show that our method outperforms the state-of-the-art approaches in terms of both accuracy and speed. Yuqun Yang, Xu Tang 0004, Yiu-Ming Cheung, Xiangrong Zhang, Fang Liu 0034, Jingjing Ma 0001, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2021 | Cross-Source Image Retrieval Based on Ensemble Learning and Knowledge Distillation for Remote Sensing ImagesabstractAs different kinds of high-resolution remote sensing (HRRS) image data sources increase, the cross-source content-based image retrieval (CS-CBRSIR) is becoming an important and urgent task to be solved. Most existing methods focus on optimizing the common space features for dual-source effectively. The source discrepancy in classifier level, however, has been ignored. To handle this problem, we propose teacher-ensemble learning with the knowledge distillation method in this paper. The ensemble of source-shared and source-specific classifiers could construct an effective teacher model. The useful information can be transferred back with the knowledge distillation. Besides, the feature pyramid network is introduced to learn the multi-scale features from HRRS images, which can describe the complex contents of HRRS images well. The positive experimental results conducted on DSRSID illustrates the effectiveness of the proposed method. Jingjing Ma 0001, Duanpeng Shi, Xu Tang 0004, Xiangrong Zhang, Xiao Han 0012, Licheng Jiao |
IGARSS | 1 |
| 2021 | Multi-Scale Meta-Learning-Based Networks for High-Resolution Remote Sensing Scene ClassificationabstractHigh-resolution remote sensing (HRRS) image scene classification based on limited data set is challenging in practical application. Although convolutional neural networks have shown powerful feature representation capability, they cannot perform well in the absence of rich label information in general. This paper proposes a multi-scale meta-learning-based (MSML) model to complete the HRRS scene classification with a little labeled data. First, we develop a multi-scale feature learning strategy to explore the rich information from HRRS scenes. Then, to use small data to train our network, we formulate the meta-learning as a regularization term and embed it into the classification loss function. By optimizing the proposed loss function, we can obtain a robust and generalized scene classification model. The positive experimental results counted on a public HRRS scene data set show that our MSML model is useful in HRRS scene classification tasks. Xu Tang 0004, Weiquan Lin, Chao Liu 0042, Xiao Han 0012, Jingjing Ma 0001, Licheng Jiao |
IGARSS | 6 |
| 2021 | Remote Scene Image Scene Classification Based on Adaptive Segmentation and Dynamic Graph ConvolutionabstractAs an important research topic in the remote sensing (RS) community, RS image scene classification is a challenging task due to the complex contents of RS images. In general, RS image scene classification is a single-label problem. Nevertheless, it is known that the contents within RS are huge in volume and diverse in type. Only a single semantic label cannot describe an RS scene completely, especially when the resolution of RS images is increased recently. The various semantics hidden in the high-resolution RS (HRRS) images are also important to the scene classification task. Taking the issues mentioned above into account, we develop a new scene classifier named graph scene classifier (GSCer) for HRRS images with the help of the deep convolution neural network (DCNN) and dynamic graph convolution (DGCN). Not only the global semantic but also the diverse hidden local semantics within an HRRS image can be fully explored. The encouraging experimental results counted on two public HRRS data sets demonstrate that our GSCer is effective in HRRS scene classification tasks. Yuqun Yang, Xu Tang 0004, Xiao Han 0012, Jingjing Ma 0001, Xiangrong Zhang, Licheng Jiao |
IGARSS | 4 |
| 2021 | High-Resolution Remote Sensing Images Change Detection with Siamese Holistically-Guided FCNabstractChange Detection is an important and challenging task in the remote sensing (RS) field, especially when the resolution of RS images is getting higher. The appearance of the deep convolutional neural network (DCNN) provides new opportunities for the high-resolution RS (HRRS) image processing as well as the HRRS CD task. In this paper, we proposed a Siamese holistically-guided FCN (SHG-FCN) model to fully mine the low- and high-level features from HRRS images for completing the CD task. SHG-FCN consists of a Siamese encoder and a holistically-guided decoder. The Siamese encoder adopts five layers of convolution for feature extraction and generates multi-scale difference maps. The decoder employs the holistically-guided architecture, which uses the deep semantic feature as codewords to guide the up-sample of feature map and achieve multi-scale feature fusion. Our model is testified on two public HRRS datasets, and the obtained encouraging CD results illustrate that our method is effective in HRRS CD tasks. Xu Tang 0004, Xiao Han 0012, Jingjing Ma 0001, Xiangrong Zhang, Licheng Jiao |
IGARSS | 4 |
| 2021 | Deep Hash Learning for Remote Sensing Image RetrievalabstractThe content-based remote sensing image retrieval (CBRSIR) has attracted increasing attention with the number of remote sensing (RS) images growing explosively. Benefiting from the strong capacity of the deep convolutional neural network (DCNN), the performance of CBRSIR has been improved in recent years. Although great successes have been obtained, learning the RS images' representative features and enhancing the retrieval efficiency for the large-scale CBRSIR tasks are still two challenging problems. In this article, we propose a new CBRSIR method named feature and hash (FAH) learning, which consists of a deep feature learning model (DFLM) and an adversarial hash learning model (AHLM). The DFLM aims at learning the RS images' dense features to guarantee the retrieval precision. In the DFLM, the DCNN and the proposed feature aggregation are integrated to capture the multiscale features. Then, the discrimination of the obtained features can be highlighted by the attention map in the developed attention branch. The AHLM maps the dense features onto the compact hash codes so that the retrieval efficiency can be improved. The AHLM contains a hash learning submodel and an adversarial regularization submodel. In particular, the hash learning submodel learns the real-valued hash codes that are similarity preserved by semantic supervisions. The adversarial regularization submodel regularizes the real-valued hash codes to learn the discrete uniform distribution with possible values 0 and 1. In this way, the hash codes are coding-balanced and the quantization errors are reduced. Encouraging experimental results counted on three public benchmark data sets demonstrate that our FAH can achieve competitive performance in the CBRSIR task compared with many existing hash learning methods. Chao Liu 0042, Jingjing Ma 0001, Xu Tang 0004, Fang Liu 0034, Xiangrong Zhang, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2021 | Hyperspectral Image Classification Based on 3-D Octave Convolution With Spatial-Spectral Attention NetworkabstractIn recent years, with the development of deep learning (DL), the hyperspectral image (HSI) classification methods based on DL have shown superior performance. Although these DL-based methods have great successes, there is still room to improve their ability to explore spatial-spectral information. In this article, we propose a 3-D octave convolution with the spatial-spectral attention network (3DOC-SSAN) to capture discriminative spatial-spectral features for the classification of HSIs. Especially, we first extend the octave convolution model using 3-D convolution, namely, a 3-D octave convolution model (3D-OCM), in which four 3-D octave convolution blocks are combined to capture spatial-spectral features from HSIs. Not only the spatial information can be mined deeply from the high- and low-frequency aspects but also the spectral information can be taken into account by our 3D-OCM. Second, we introduce two attention models from spatial and spectral dimensions to highlight the important spatial areas and specific spectral bands that consist of significant information for the classification tasks. Finally, in order to integrate spatial and spectral information, we design an information complement model to transmit important information between spatial and spectral attention features. Through the information complement model, the beneficial parts of spatial and spectral attention features for the classification tasks can be fully utilized. Comparing with several existing popular classifiers, our proposed method can achieve competitive performance on four benchmark data sets. Xu Tang 0004, Xiangrong Zhang, Yiu-Ming Cheung, Jingjing Ma 0001, Fang Liu 0034, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2020 | Remote Sensing Images Feature Learning based On Multi-Branch NetworksabstractRemote sensing (RS) images feature learning, plays a crucial role in many RS images application, and attracts scholars' attention. However, since RS images contain complex contents, how to extract robust features that can fully represent RS images becomes an important and tough task. In this paper, we develop a feature learning method based on multi-branch networks, named M-Net, which consists of fine-grained branch and coarse branch. Considering the objects within RS images are diverse in type and resolution, the fine-grained branch is developed to capture rich object-level information. First, the RS images convolutional features are extracted by fine-grained branch. Second, through encoding the score maps which can highlight the important regions, the fine-grained structure mapping are obtained. Finally, the object-level features are generated by transforming the convolutional features through mapping. The coarse branch is developed to transform the obtained object-level features into global structure for representing images. The positive experimental results counted on RS benchmark data set demonstrate that the proposed M-Net can learn more powerful features. Chao Liu 0042, Xu Tang 0004, Jingjing Ma 0001, Xiangrong Zhang, Fang Liu 0001, Junyong Ma, Licheng Jiao |
IGARSS | 3 |
| 2020 | Remote Sensing Scene Classification Based on Global and Local Consistent NetworkabstractScene classification of remote sensing (RS) images has attracted increasing attention due to its wide applications. Recently, with the advances of deep learning models, especially convolutional neural networks (CNNs), the performance of remote sensing image scene classification has been significantly improved. In this paper, based on the popular CNN, we develop a new scene classification network, named the Global and Local Consistent (GLC) network, to deeply explore useful information from the RS images. First, we adopt a pre-trained CNN to learn the intermediate feature maps from the RS image pairs. Second, by introducing the visual attention mechanism, the global and local integration model is developed to mine the rich information from the obtained feature maps. Third, the attention consistent model is designed to eliminate the negative influence of the issue of attention inconsistency on the classification. To verify the effectiveness of the proposed method, we select two popular RS image data sets. Compared with some existing classification models, our network can achieve competitive results, which illustrates that our method is useul to the RS scene classification. Jingjing Ma 0001, Qiushuo Ma, Xu Tang 0004, Xiangrong Zhang, Qunnie Peng, Licheng Jiao |
IGARSS | 1 |
| 2020 | Hyperspectral Image Classification Via Multi-Scale Encoder-Decoder NetworkabstractHyperspectral image (HSI) classification is an important task in the remote sensing community. In general, many hyperspectral classification methods are based on pixel patch, which leads to information redundancy. In this paper, we propose a multi-scale encoder-decoder network for HSI classification. First, we adapt an encoder-decoder framework as the backbone network and use a skip connection between the encoder and decoder, the spatial information is obtained by this network. Second, we develop a multi-scale block to get the multi-scale information. Third, we retain complete spectral information through the constant number of spectral channels. Finally, an optimizer strategy is designed to achieve our model for the HSI classification task. We experiment with our method and other methods on two public datasets, and the results denote our model is useful for HSI classification task. Jingjing Ma 0001, Linlin Wu, Xu Tang 0004, Xiangrong Zhang, Junyong Ma, Licheng Jiao |
IGARSS | 1 |
| 2020 | Hyperspectral Image Classification Based on Multiscale Spatial and Spectral Feature NetworkabstractWith the development of deep learning, hyperspectral image (HSI) classification tasks have developed rapidly, the classification performance is improved in a big degree. Despite the great success of the existing methods, there is still room for improvement to extract features from spatial and spectral dimensions. In this paper, we propose a multiscale spatial and spectral feature network (MSSFN) to capture discriminative features for the classification of HSIs. Specifically, we first use three convolution layers to extract the features of original HSI data. Second, combining the spatial masks model and spectral attention model to build multiscale spatial and spectral model (MSSM). Through the MSSM model, the spatial information of different scales can be obtained and the useful spectral bands can be emphasized. Finally, in order to reduce the computation complexity and simplify network, the other three convolution layers with a small number of convolution kernels are adopted in our method. The experimental results demonstrate that our method is superior to most existing methods on two public HSI datasets. Xu Tang 0004, Jingjing Ma 0001, Xiangrong Zhang, Fang Liu 0034, Qunnie Peng, Licheng Jiao |
IGARSS | 3 |
| 2020 | Supervised Adaptive-RPN Network for Object Detection in Remote Sensing ImagesabstractObject detection is one of the most important tasks in the field of very high resolution (VHR) remote sensing (RS) images understanding. Due to the characteristics of VHR RS images, the detection performance is always limited by class imbalance and intersection-over-union (IoU) distribution imbalance. To mitigate the adverse effects caused thereby, we propose a supervised adaptive-RPN (SA-RPN) model with the help of deep learning in this paper. First, we introduced a supervised multi-dimensional attention network to overcome the foreground-background class imbalance. It can help the network to highlight the foreground and suppress the background effectively. Second, we develop the adaptive-RPN to reduce the negative impact of the IoU distribution imbalance by adaptively selecting the size of the anchor. The positive experimental results on the public data set validate the usefulness of our SA-RPN model. Compared with the popular deep learning RS object detection methods, our method achieves improved performance. Xu Tang 0004, Jingjing Ma 0001, Xiangrong Zhang, Licheng Jiao |
IGARSS | 3 |
| 2019 | Evolutionary Multiobjective Change Detection via Self-paced Learning and Fuzzy ClusteringabstractFuzzy clustering algorithm based on multiobjective optimization can achieve accurate and comprehensive clustering results. However, the estimation of objective values for this multiobjective optimization problem (MOP) might be expensive. Offspring's selection driven by simple evaluation is time consuming. Therefore, we integrate regression techniques to determine the superiority of the offspring solutions in the evolution process. However, it suffers from an issue that it is hard to collect reliable samples to train such a robust regression model. In this paper, an evolutionary multiobjective fuzzy clustering method via self-paced learning is proposed for change detection. In the proposed method, the self-paced learning process is implemented to collect reliable training samples for training a robust regression model, which can help to select promising offspring solutions from the candidate solutions for MOP. Experiments on three remote sensing image datasets demonstrate that the proposed method can significantly outperform those state-of-art methods for change detection in terms of accuracy and robustness. Yingying Duan, Jingjing Ma 0001, Hao Li 0009, Mingyang Zhang 0002, Zedong Tang, Maoguo Gong |
CEC | 2 |
| 2019 | Adversarial Hash-Code Learning for Remote Sensing Image RetrievalabstractHashing, a useful solution for approximate nearest neighbor (ANN) search, is popular for large-scale image retrieval. In this paper, we presents a deep supervised hashing model for remote sensing image retrieval (RSIR) in the framework of generative adversarial networks (GAN), named GAN-assist Hashing (GAAH). First, to learn the compact and useful hash codes from the images, we define a novel loss function for the generator. The loss function mainly consists of classification, similarity, and bits entropy terms. The classification term makes the hash code is discriminative, the similarity term constrains the binary code is similarity preserving, and the bits entropy term assures the learned code is low-error in the quantization. Second, we construct the unique "true" matrix with the uniform distribution as the input of discriminator to limit the leaned hash codes are bit balanced. The final hash code is learned by a minimax optimization. The positive experimental results on a ground-truth remote sensing image archive validate the usefulness of our GAAH model. Compare with the popular deep hashing methods, our GAAH achieves improved performance. Chao Liu 0042, Jingjing Ma 0001, Xu Tang 0004, Xiangrong Zhang, Licheng Jiao |
IGARSS | 2 |
| 2019 | Remote Sensing Image Retrieval Based on Semi-Supervised Deep Hashing LearningabstractAs an useful solution of the approximate nearest neighbor (ANN) search, hashing attracts growing attention in the topic of large-scale image retrieval. In this paper, we propose a semi-supervised deep hashing method based on the adversarial autoencoder (AAE) network for remote sensing image retrieval (RSIR), and we name it SSHAAE. Here, we assume the RS images have been represented by the visual features, and the target of our SSHAAE is mapping those features into the binary codes. First, a hashing layer is adopted to replace the part of original latent layer in AAE. In addition, the classical reconstruction loss function is selected to generate the hash code. Second, two discriminators are added simultaneously to make sure the hash code is bit balanced and the generated label variable is one-hot. Third, we design the hash loss function to guarantee the obtained hash code is discriminative, similarity persevering, and low quantization error. The presented SSHAAE model can be trained by the minimax optimization. The encouraging experimental results counted on a high-resolution RS image archive demonstrate our SSHAAE model is effective to RSIR. Xu Tang 0004, Chao Liu 0042, Xiangrong Zhang, Jingjing Ma 0001, Changzhe Jiao, Licheng Jiao |
IGARSS | 4 |
| 2017 | Multi-objective endmember extraction for hyperspectral imagesabstractEndmember extraction is a critical step of spectral unmixing. In this paper, a novel endmember extraction algorithm based on evolutionary multi-objective optimization is proposed for hyperspectral remote sensing images. In the proposed method, endmember extraction is modeled as a multi-objective optimization problem. Then the root mean square error between the original image and its remixed image and the number of endmembers are chosen as two conflicting objective functions, which are simultaneously optimized by particle swarm optimization algorithm to find the best tradeoff solutions. In order to promote diversity and speed up the convergence of the algorithm, a new particle status updating strategy and a novel method for selecting leaders are designed. The experimental results on both simulated and real hyperspectral remote sensing images confirm the performance of the proposed approach over some existing methods. Hao Li 0009, Jingjing Ma 0001, Jia Liu 0020, Maoguo Gong, Mingyang Zhang 0002 |
CEC | 2 |
| 2017 | Memetic algorithm based feature selection for hyperspectral images classificationabstractBand selection is a crucial preprocessing step for hyperspectral image classification, which is a classic feature selection method. Feature selection is designed to select feature subsets to represent the whole feature space. For feature selection, two crucial issues need to be handled: preserving information and redundancy reducing. In this paper, a novel feature selection method for hyperspectral image classification is proposed, which is based on a newly designed memetic algorithm. In the proposed method, a suitable objective function is designed, which can measure the contained crucial information and redundancy information in the selected feature subsets. To optimize this objective function efficiently, a novel memetic algorithm is designed. The genetic operator and local search strategy are newly designed according to the characteristic of hyperspectral images. Experiments are implemented on three real data sets compared with some state of arts. The experimental results show that the proposed method can obtain stable and superior feature subsets for classification. Mingyang Zhang 0002, Jingjing Ma 0001, Maoguo Gong, Hao Li 0009, Jia Liu 0020 |
CEC | 2 |
| 2017 | Unsupervised Hyperspectral Band Selection by Fuzzy Clustering With Particle Swarm OptimizationabstractDue to the lack of label information and the intrinsic complexity of hyperspectral images (HSIs), unsupervised band selection is always one of the most challenging tasks in HSI processing. Fuzzy clustering is a promising technique for unsupervised band selection, which can partition unlabeled data into groups effectively. However, due to the limits of its optimization process, standard fuzzy clustering is sensitive to initialization and easy to be trapped in a local optimum. To address the limits, a novel unsupervised band selection method is proposed, combining fuzzy clustering with particle swarm optimization (PSO). A newly designed PSO algorithm is introduced to improve the performance of fuzzy clustering band selection. Moreover, a new strategy is designed to select representative cluster centers according to the characteristics of HSIs. The experimental results indicate that the proposed method has the ability to select high-quality band subsets with good and robust performance on HSI classification. Mingyang Zhang 0002, Jingjing Ma 0001, Maoguo Gong |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2016 | A new quantum-behaved particle swarm optimization based on cultural evolution mechanism for multiobjective problems
Licheng Jiao, Wenping Ma 0001, Jingjing Ma 0001, Ronghua Shang |
Knowl. Based Syst. | 4 |
| 2014 | A multi-swarm particle swarm optimization with orthogonal learning for locating and tracking multiple optimization in dynamic environmentsabstractDue to the specificity and complexity of the dynamic optimization problems (DOPs), those excellent static optimization algorithms cannot be applied in these problems directly. So some special algorithms only for DOPs are needed. There is a multi-swarm algorithm with a better performance than others in DOPs, which utilizes a parent swarm to explore the search space and some child swarms to exploit promising areas found by the parent swarm. In addition, a static optimization algorithm OLPSO is so attractive, which utilize an orthogonal learning (OL) strategy to utilize previous search information (experience) more efficiently to predict the positions of particles and improve the convergence speed. In this paper, we bring the essence of OLPSO called OL strategy to the multi-swarm algorithm to improve its performance further. The experimental results conducted on different dynamic environments modeled by moving peaks benchmark show that the efficiency of this algorithm for locating and tracking multiple optima in dynamic environments is outstanding in comparison with other particle swarm optimization models, including MPSO, a similar particle swarm algorithm for dynamic environments. Ruochen Liu 0006, Xu Niu, Licheng Jiao, Jingjing Ma 0001 |
IEEE Congress on Evolutionary Computation | 4 |
| 2014 | A memetic algorithm based on Immune multi-objective optimization for flexible job-shop scheduling problemsabstractThe flexible job-shop scheduling problem (FJSP) is an extension of the classical job scheduling which is concerned with the determination of a sequence of jobs, consisting of many operations, on different machines, satisfying parallel goals. This paper addresses the FJSP with two objectives: Minimize makespan, Minimize total operation cost. We introduce a memetic algorithm based on the Nondominated Neighbor Immune Algorithm (NNIA), to tackle this problem. The proposed algorithm adds, to NNIA, local search procedures including a rational combination of undirected simulated annealing (UDSA) operator, directed cost simulated annealing (DCSA) operator and directed makespan simulated annealing (DMSA) operator. We have validated its efficiency by evaluating the algorithm on multiple instances of the FJSPs. Experimental results show that the proposed algorithm is an efficient and effective algorithm for the FJSPs, and the combination of UDSA operator, DCSA operator and DMSA operator with NNIA is rational. Jingjing Ma 0001, Yu Lei 0002, Zhao Wang 0011, Licheng Jiao, Ruochen Liu 0006 |
IEEE Congress on Evolutionary Computation | 1 |
| 2014 | A compression optimization algorithm for community detectionabstractCommunity detection is important in understanding the structures and functions of complex networks. Many algorithms have been proposed. The most popular algorithms detect the communities through optimizing a criterion function known as modularity, which suffer from the resolution limit problem. Some algorithms require the number of communities as a prior. In this paper, a non-modularity based compression optimization algorithm for community detection is proposed without any prior knowledge, which is efficient and is suitable for large scale networks. Jianshe Wu, Qingliang Gong, Wenping Ma 0002, Jingjing Ma 0001, Yangyang Li 0001 |
IEEE Congress on Evolutionary Computation | 5 |
| 2013 | Fuzzy C-Means Clustering With Local Information and Kernel Metric for Image SegmentationabstractIn this paper, we present an improved fuzzy C-means (FCM) algorithm for image segmentation by introducing a tradeoff weighted fuzzy factor and a kernel metric. The tradeoff weighted fuzzy factor depends on the space distance of all neighboring pixels and their gray-level difference simultaneously. By using this factor, the new algorithm can accurately estimate the damping extent of neighboring pixels. In order to further enhance its robustness to noise and outliers, we introduce a kernel distance measure to its objective function. The new algorithm adaptively determines the kernel parameter by using a fast bandwidth selection rule based on the distance variance of all data points in the collection. Furthermore, the tradeoff weighted fuzzy factor and the kernel distance measure are both parameter free. Experimental results on synthetic and real images show that the new algorithm is effective and efficient, and is relatively independent of this type of noise. Maoguo Gong, Jiao Shi, Wenping Ma 0001, Jingjing Ma 0001 |
IEEE Trans. Image Process. | 5 |
| 2012 | An improved memetic algorithm for community detection in complex networksabstractThere is an increasing recognition on community detection in complex networks in recent years. In this study, we improve a recently proposed memetic algorithm for community detection in networks. By introducing a Population Generation via Label Propagation (PGLP) tactic, an Elitism Strategy (ES) and an Improved Simulated Annealing Combined Local Search (ISACLS) strategy, the improved memetic algorithm called (iMeme-Net) is put forward for solving community detection problems. Experiments on both computer-generated and real-world networks show the effectiveness and the multi-resolution ability of the proposed method. Maoguo Gong, Yangyang Li 0001, Jingjing Ma 0001 |
IEEE Congress on Evolutionary Computation | 4 |
| 2012 | Community Detection in Dynamic Social Networks Based on Multiobjective Immune Algorithm
Maoguo Gong, Ling-Jun Zhang, Jingjing Ma 0001, Licheng Jiao |
J. Comput. Sci. Technol. | 3 |
| 2012 | An efficient negative selection algorithm with further training for anomaly detection
Maoguo Gong, Jingjing Ma 0001, Licheng Jiao |
Knowl. Based Syst. | 3 |
| 2012 | Wavelet Fusion on Ratio Images for Change Detection in SAR ImagesabstractThis letter presents a novel method based on wavelet fusion for change detection in synthetic aperture radar (SAR) images. The proposed approach is applied to generate the difference image (DI) by using complementary information from mean-ratio and log-ratio images. To restrain the background (unchanged areas) information and enhance the information of changed regions in the fused DI, fusion rules based on weight averaging and minimum standard deviation are chosen to fuse the wavelet coefficients for low- and high-frequency bands, respectively. Experiments on real SAR images confirm that the proposed approach does better than the mean-ratio, log-ratio, and Rayleigh-distribution-ratio operators. Jingjing Ma 0001, Maoguo Gong |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2012 | Change Detection in Synthetic Aperture Radar Images based on Image Fusion and Fuzzy ClusteringabstractThis paper presents an unsupervised distribution-free change detection approach for synthetic aperture radar (SAR) images based on an image fusion strategy and a novel fuzzy clustering algorithm. The image fusion technique is introduced to generate a difference image by using complementary information from a mean-ratio image and a log-ratio image. In order to restrain the background information and enhance the information of changed regions in the fused difference image, wavelet fusion rules based on an average operator and minimum local area energy are chosen to fuse the wavelet coefficients for a low-frequency band and a high-frequency band, respectively. A reformulated fuzzy local-information C-means clustering algorithm is proposed for classifying changed and unchanged regions in the fused difference image. It incorporates the information about spatial context in a novel fuzzy way for the purpose of enhancing the changed information and of reducing the effect of speckle noise. Experiments on real SAR images show that the image fusion strategy integrates the advantages of the log-ratio operator and the mean-ratio operator and gains a better performance. The change detection results obtained by the improved fuzzy clustering algorithm exhibited lower error than its preexistences. Maoguo Gong, Jingjing Ma 0001 |
IEEE Trans. Image Process. | 3 |
| 2010 | Quantum-inspired immune clonal clustering algorithm based on watershedabstractBased on the concepts and principles of quantum computing, a novel clustering algorithm, called a quantum-inspired immune clonal clustering algorithm based on watershed (QICW), is proposed to deal with the problem of image segmentation. In QICW, antibody is proliferated and divided into a set of subpopulation groups. Antibodies in a subpopulation group are represented by multi-state gene quantum bits. In the antibody's updating, the quantum mutation operator is applied to accelerate convergence. The quantum recombination realizes the information communication between the subpopulation groups so as to avoid premature convergences. In this paper, the segmentation problem is viewed as a combinatorial optimization problem, the original image is partitioned into small blocks by watershed algorithm, and the quantum-inspired immune clonal algorithm is used to search the optimal clustering centre, and make the sequence of maximum affinity function as clustering result, and finally obtain the segmentation result. Experimental results show that the proposed method is effective for texture image and SAR image segmentation, compared with the genetic clustering algorithm based on watershed (W-GAC), and the k-means algorithm based on watershed (W-KM). Yangyang Li 0001, Nana Wu, Jingjing Ma 0001, Licheng Jiao |
IEEE Congress on Evolutionary Computation | 3 |
| 2010 | Optimizing detector distribution in V-detector negative selection using a constrained multiobjective immune algorithmabstractIn this paper, a novel constrained multiobjective immune algorithm for optimizing detector distribution in V-detector negative selection is proposed. The theory of artificial immune system (AIS) and the spirit of population evolution are introduced to generate detectors. By combining the constraint handling technique and AIS-based multiobjective optimization, the algorithm is able to steadily maximize the anomaly coverage with little extra cost, which means the distribution with maximized coverage of the non-self space and minimized overlapping among detectors with fixed size will be well realized. Furthermore, the new approach is tested on some benchmark problems. The experimental results show that compared with some state-of-the-art methods, our algorithm can remarkably outperform them in terms of enhancing the detection rate by optimizing distribution without increasing the number of detectors. Fang Liu 0001, Maoguo Gong, Jingjing Ma 0001, Licheng Jiao, Wei Zhang 0009 |
IEEE Congress on Evolutionary Computation | 3 |
| 2010 | Clonal Selection Algorithm for Image CompressionabstractVector Quantization (VQ) is a useful tool for data compression and can be applied to compress the data vectors in the database. The quality of the recovered data vector depends on a good codebook. Mean/residual vector quantization (M/RVQ) has been shown to be efficient in the encoding time and it only needs a little storage. In this paper, Clonal Selection Algorithm for Image Compression (CSAIC) is proposed. In CSAIC, Based on M/RVQ algorithm, an improved clonal selection algorithm is used to cluster the data of compressed images in order to obtain the optimal codebook. The proposed method has been extensively compared with Linde-Buzo-Gray(LBG), Self-Organizing Mapping (SOM) and Modified K-means(Mod-KM) over a test suit of seven natural images. The experimental results show that CSAIC outperforms other three algorithms in terms of image compression performance. Ruochen Liu 0006, Licheng Jiao, Wei Zhang 0009, Jingjing Ma 0001 |
IEEE Congress on Evolutionary Computation | 4 |
| 2010 | Unsupervised evolutionary clustering algorithm for mixed type dataabstractIn this paper, we propose a novel unsupervised evolutionary clustering algorithm for mixed type data, evolutionary k-prototype algorithm (EKP). As a partitional clustering algorithm, k-prototype (KP) algorithm is a well-known one for mixed type data. However, it is sensitive to initialization and converges to local optimum easily. Global searching ability is one of the most important advantages of evolutionary algorithm (EA), so an EA framework is introduced to help KP overcome its flaws. In this study, KP is applied as a local search strategy, and runs under the control of the EA framework. Experiments on synthetic and real-life datasets show that EKP is more robust and generates much better results than KP for mixed type data. Zhi Zheng 0002, Maoguo Gong, Jingjing Ma 0001, Licheng Jiao, Qiaodi Wu |
IEEE Congress on Evolutionary Computation | 3 |
| 2010 | A sphere-dominance based preference immune-inspired algorithm for dynamic multi-objective optimizationabstractReal-world optimization involving multiple objectives in changing environment known as dynamic multi-objective optimization (DMO) is a challenging task, especially special regions are preferred by decision maker (DM). Based on a novel preference dominance concept called sphere-dominance and the theory of artificial immune system (AIS), a sphere-dominance preference immune-inspired algorithm (SPIA) is proposed for DMO in this paper. The main contributions of SPIA are its preference mechanism and its sampling study, which are based on the novel sphere-dominance and probability statistics, respectively. Besides, SPIA introduces two hypermutation strategies based on history information and Gaussian mutation, respectively. In each generation, which way to do hypermutation is automatically determined by a sampling study for accelerating the search process. Furthermore, The interactive scheme of SPIA enables DM to include his/her preference without modifying the main structure of the algorithm. The results show that SPIA can obtain a well distributed solution set efficiently converging into the DM's preferred region for DMO. Ruochen Liu 0006, Wei Zhang 0009, Licheng Jiao, Fang Liu 0001, Jingjing Ma 0001 |
GECCO | 5 |
| 2009 | Intelligent multi-user detection using an artificial immune system
Maoguo Gong, Licheng Jiao, Wenping Ma 0001, Jingjing Ma 0001 |
Sci. China Ser. F Inf. Sci. | 4 |