Zhongle Ren

dblp:227/3301 · DBLP profile ↗
← Back
18ranked-venue papers
4as first author
15since 2021 · last 2026
0000-0002-5425-6437ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 16 · 3 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 DI3CL: Contrastive Learning With Dynamic Instances and Contour Consistency for SAR Land-Cover Classification Foundation Model
abstract
Although significant advances have been achieved in SAR land-cover classification, recent methods remain predominantly focused on supervised learning, which relies heavily on extensive labeled datasets. This dependency not only limits scalability and generalization but also restricts adaptability to diverse application scenarios. In this paper, a general-purpose foundation model for SAR land-cover classification is developed, serving as a robust cornerstone to accelerate the development and deployment of various downstream models. Specifically, a Dynamic Instance and Contour Consistency Contrastive Learning (DI3CL) pre-training framework is presented, which incorporates a Dynamic Instance (DI) module and a Contour Consistency (CC) module. DI module enhances global contextual awareness by enforcing local consistency across different views of the same region. CC module leverages shallow feature maps to guide the model to focus on the geometric contours of SAR land-cover objects, thereby improving structural discrimination. Additionally, to enhance robustness and generalization during pre-training, a large-scale and diverse dataset named SARSense, comprising 460,532 SAR images, is constructed to enable the model to capture comprehensive and representative features. To evaluate the generalization capability of our foundation model, we conducted extensive experiments across a variety of SAR land-cover classification tasks, including SAR land-cover mapping, water detection, and road extraction. The results consistently demonstrate that the proposed DI3CL outperforms existing methods. Our code and pre-trained weights are publicly available at: https://github.com/SARpre-train/DI3CL.
Zhongle Ren, Kai Wang 0053, Biao Hou, Xingyu Luo, Weibin Li 0002, Licheng Jiao
IEEE Trans. Image Process.1
2026 Scale-Aware Prompting With Optimal Transport for Remote Sensing Image Captioning
abstract
Remote sensing image captioning is a multimodal foundation task for fine-grained understanding of remote sensing images. However, remote sensing images contain complex scenes and rich objects, it is very challenging to accurately describe the objects in the scene with their attributes and dependencies. To address these issues, the article proposes a novel scale-aware prompting with optimal transport (SPOT) to learn effective multiscale features under diverse scenes, and to build fine-grained cross-modal alignment between semantic features and linguistic words during caption generation. Specifically, a scale-aware prompt extractor is constructed to explore feature integrations in complex scenes through learning prompts that query multi-scale features, and to enhance the representation of attributes and dependencies for objects by embedding positional relations. Besides, a fine-grained cross-modal alignment is designed to dynamically match image feature representations and textual semantics through optimal transport. Through the above manner, the model can learn effective language-aligned feature representations for caption generation. Finally, a caption Transformer with causal self-attention is introduced to generate accurate captions for remote sensing scenes. Extensive experiments demonstrate that the proposed method achieves state-of-the-art performance on three public datasets, with the superiority of the proposed method further demonstrated by ablating the role of each component.
Cheng Zhang 0028, Zhongle Ren, Biao Hou, Jiawei Ning, Kai Wang 0053, Weibin Li 0002, Licheng Jiao
IEEE Trans. Image Process.2
2025 Hierarchical Prototype Learning With Uncertainty-Aware Adaptation for Cross-Domain Semantic Segmentation of Remote Sensing Images
abstract
Prominent domain discrepancies in Remote Sensing Images (RSIs), such as sensor types, geographical patterns, and land usage, significantly hinder the research and practical applications of cross-scene land classification. Unsupervised Domain Adaptation (UDA) fully exploits domain invariance between labelled source and unlabelled target domain, which alleviates the challenge of inaccurate land classification due to lack of labels in RSIs. However, most existing UDA methods for RSIs semantic segmentation are insufficient in exploring cross-domain features and have difficulty in modelling fine-grained domain-invariant features between inter-class. In this paper, we propose a novel self-training UDA method named Hierarchical Prototype Learning (HPL), which learns the inherent domain-wise consistency and class-wise invariance through progressive exploration of the prototypy, significantly alleviating the ambiguity in uncertain regions. HPL mainly consists of Domain-wise Progressive Prototype Interaction (DPPI) and Class-wise Dynamic Prototype Collaboration (CDPC). DPPI and CDPC specialize in hierarchically building prototype interaction architectures tailored to domain-wise alignment and class-wise calibration, respectively. This design not only mitigates the sensitivity to cross-domain scenarios but also allows for precise correction of uncertain regions. Furthermore, CDPC exhibits the capacity for pixel-level category restoration and promotes the correct and fine-grained updating of pseudo-labels. Extensive Experiments on two public datasets and a private self-build dataset demonstrate the superiority of HPL over other state-of-the-art methods for UDA semantic segmentation of RSIs.
Jiawei Ning, Zhongle Ren, Biao Hou, Runnong Jiang, Weibin Li 0002, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.2
2025 Deep Geospatio-Semantic Guided Network With Pseudo-Label Consistency for Domain-Adaptive Remote Sensing Segmentation
abstract
Domain-adaptive Remote Sensing Images (RSIs) semantic segmentation mitigates the overfitting problem that affects the effectiveness of segmentation, which results from the scarcity of high-quality labels and the cross-domain styles of ground objects. The effectiveness of domain adaptive segmentation remains suboptimal in complex scenarios due to inadequate exploitation of latent geographic knowledge. Consequently, inter-class ambiguity and boundary agnostic are further exacerbated under cross-domain transfer scenarios. To address this issue, we first devise a deep geospatio-semantic guided network named DSSAL, which comprehensively investigates the potential spatial relationship and semantic correlation between classes of RSIs by geospatial aware interaction and geosemantic aware interaction, respectively. To mitigate class-wise cognitive deviation in the unlabeled domain, DSSAL-DA is developed to further enhance the segmentation effect with the spatio-semantic domain alignment module in manifold cross domain tasks. Furthermore, a pseudo-labels consistency filter is developed for DSSAL-DA to ensure reliability in self-training through cross-view consistency verification. Extensive experiments on two public datasets and a private dataset demonstrate the superiority of DSSAL and DSSAL-DA over the state-of-the-art methods for UDA semantic segmentation of RSIs.
Jiawei Ning, Zhongle Ren, Biao Hou, Weibin Li 0002, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.2
2025 Self-Supervised Learning of Contrast-Diffusion Models for Land Cover Classification in SAR Images
abstract
Deep learning methods has been widely applied to synthetic aperture radar (SAR) land cover classification. The complexity of SAR data and the limited availability of labeled samples greatly constrain the feature learning and the generalization ability of the model. Inspired by the excellent generative performance of Denoising Diffusion Probabilistic Models(DDPM) on complex data distributions, a self-supervised learning framework based on contrast-diffusion models (CDM) is proposed to expand the applicability to multiple broad scenarios with complex and varying imaging conditions under limited annotated data conditions. Specially, The proposed framework consists of the upstream CDM pre-training on all unlabeled samples and the downstream land cover classification with few labeled samples in each test scene. Concretely, in the upstream task, the features are captured through the generative learning of the DDPM. Following this, the Dimensionality Reduction and Resolution Expansion (DRRE) module is designed and embedded to reduce feature redundancy and align the feature granularity between layers and the input image. Finally, contrastive learning is employed to enforce semantic feature consistency across different steps. In the downstream task, the feature in pre-trained CDM is efficiently delivered in a single-step reverse diffusion process and then fine-tuned with few labeled samples from each test scene and finally output the predictions. Compared with several supervised and self-supervised methods, the proposed framework achieves superior classification and generalization performance on multiple broad scenes with complex and varying imaging conditions. For example, based on the average results from six test scenes, CDM shows improvements in overall accuracy (OA) of 41.42%, 34.40%, 7.70% ,8.53% and 9.60% compared to Deeplabv3+, CCNR, Segformer, MAE and DDPM respectively. The code is available at https://github.com/gosling123456/CDM.git.
Zhongle Ren, Zhe Du, Biao Hou, Weibin Li 0002, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.1
2025 Interactive Concept Network Enhanced Transformer for Remote Sensing Image Captioning
abstract
Remote sensing image captioning plays an important role in advancing remote sensing image understanding with natural language generation. However, it is difficult to generate accurate semantic descriptions of crucial objects and their relationships, due to large coverage and abundant information in remote sensing images. To address these issues, this article proposes a novel interactive concept network enhanced transformer (ICNET) for remote sensing image captioning. First, multilevel visual features are extracted within a local and global feature extraction module. To comprehensively capture key objects in the local features, a concept mapping network (CMN) is constructed to project multiscale local features onto high-level semantic concepts of the objects. This allows for the integration of the relevant feature vectors in the visual feature mapping into multiple relatively independent word features, thus bridging the gap between visual features and semantic concepts. Subsequently, a global feature enhancement (GFE) module is introduced to boost the discrimination of global relationships and filter irrelevant content. Finally, to aggregate semantic concepts and global features, a transformer equipped with a concept interaction module (CIM) is designed to facilitate feature alignment and generate captions with proper categories and relationships. The experimental results on three remote sensing image captioning datasets demonstrate the superiority of the proposed method.
Cheng Zhang 0028, Zhongle Ren, Biao Hou, Jianhua Meng, Weibin Li 0002, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.2
2025 Adaptive Scale-Aware Semantic Memory Network for Remote Sensing Image Captioning
abstract
Remote sensing image captioning between visual images and natural language remains a long-standing challenge in the remote sensing community. Due to the wide coverage and large amount of information in remote sensing images, existing methods struggle to effectively utilize the relevant semantic information about objects and their attributes at different scales across samples to generate descriptions. To address these issues, the article proposes a novel Adaptive Scale-aware Semantic Memory Network (ASSMN) for remote sensing image captioning. First, to fully extract the semantic information in remote sensing images, multilevel feature enhancement is constructed to improve the feature representation extracted from the CLIP pre-training model. Subsequently, a scale-aware attention aggregator is introduced to further integrate the enhanced multi-scale image features into the high-level semantics of remote sensing images. Then, to fully exploit the semantic information of the joint observed samples, a semantic memory reinforcement is designed to strengthen the semantic representation of the current scene through the relevant semantics obtained from other training samples. Finally, a captioning decoder is performed to generate a comprehensive scene caption with accurate objects and attributes. In the experiments, the performance of the proposed ASSMN on three remote sensing image captioning datasets is evaluated and compared with well-designed baselines and state-of-the-art methods, and the superiority of the proposed method is demonstrated by ablating the role of each proposed component. The code will be available at https://github.com/zcsisiyao/ASSMN.
Cheng Zhang 0028, Zhongle Ren, Biao Hou, Changhui Xu, Jianhua Meng, Weibin Li 0002, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.2
2024 Rebalancing Gaussian Location Loss for High-Precision Detection on Remote Sensing Images
abstract
Aerial image objects are usually orientated arbitrarily, with a large scale range, and densely distributed. Traditional horizontal bounding box (HBB) detectors tend to filter out densely distributed objects leading to missed detections, such as ship (SH) and vehicle. Therefore, oriented object detection has become a mainstream solution in recent years. The 2-D Gaussian distribution representation of the oriented bounding boxes (OBBs) solves the problem of angular discontinuity and boundary discontinuity and thus gets more attention. However, as the aspect ratio of the object gradually decreases, its predicted angular performance continues to decrease. We find that the angular gradient of an object decreases sharply as the aspect ratio decreases, resulting in a large gradient gap between a small aspect ratio object (SARO) and a large aspect ratio object (LARO). It makes the detector prefer to ignore SARO during training, which weakens the high precision performance of SARO. We call this phenomenon shape imbalance. To solve the problem, we proposed a simple gradient rebalancing strategy named shape balance. Since the shape imbalance is only related to the aspect ratio of the object, we designed a modulation function with an inverse aspect ratio to calculate the balance coefficient. The principle of the function is that the larger the aspect ratio, the smaller the balance coefficient; the smaller the aspect ratio, the larger the balance coefficient. We aim to get the balance coefficients for objects with different aspect ratios. Location loss multiplied by a balance coefficient can directly adjust the gradient gap between objects with different aspect ratios to achieve a rebalancing effect. Extensive experiments conducted on DOTA-v1.0 dataset and DIOR-R dataset verify the effectiveness of our proposed method. Our method improves the detection performance of Gaussian location loss by an average of 2.08%/1.01%(AP75/mAP) metrics on the DOTA-v1.0 dataset and 1.17%/0.82%(AP75/mAP) improvements for DIOR-R dataset.
Biao Hou, Zitong Wu, Xianpeng Guo, Bo Ren 0001, Zhongle Ren, Chen Yang 0027, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.6
2024 Self-Supervised Learning Guided by SAR Image Factors for Terrain Classification
abstract
Effective feature representation is the key to SAR image terrain classification. Limited by the abstract appearance and the scarcity of high-quality labeled data in this field, the features learned by current methods, especially deep learning models, do not have enough directivity and applicability, which hampers the performance. This paper proposes Multi-image Factor Self-Supervised Learning(MFSSL) to achieve directional feature learning and obtain generalized features with few patch-level labeled data. The framework consists of an upstream multi-factor image style transfer task and a downstream terrain classification task. In the upstream task, the goal of feature learning is first set up by multiple SAR image factors, including the observation region, the terrain category, and the imaging parameters. And then, different styles of SAR terrain images are generated and reconstructed under this goal. Through this bidirectional generative learning, the low-level external appearance of the terrain is removed, while the essential and discriminative feature representation is retained and shared across different factors. Finally, the downstream model inherits the general feature from the upstream model and implements the terrain classification task using a small amount of labeled data. Experiments conducted on three broad SAR scenes with different image factors demonstrate that the proposed framework can improve pixel-level terrain classification only with a few patch-level labeled data.
Zhongle Ren, Zhe Du, Biao Hou, Weibin Li 0002, Hao Zhu 0009, Bo Ren 0001, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.1
2024 A Semantically Nonredundant Continuous-Scale Feature Network for Panchromatic and Multispectral Classification
abstract
In recent years, panchromatic (PAN) images and multispectral (MS) images, as a type of multimodal remote sensing data, are attracting increasingly more attention to their classification problems. However, effectively representing size variations of targets in remote sensing images and reducing redundant representations of different modalities’ deep semantic features to enhance classification accuracy remains a challenge. In this article, we propose a semantically nonredundant continuous-scale feature network (SNCF-Net) for PAN and MS classification, consisting of two modules: the texture-enhanced continuous scale input generation module and the cross-modal feature Kernel interaction (CMKI) module. By simulating the human eye’s adjustment of distance to observe objects of different sizes, we employ 3-D convolution to extract continuous-scale images generated by the texture-enhanced continuous-scale input generation (TCIG) module, enabling optimal feature representation of objects in remote sensing images. Additionally, the texture enhancement (TE) strategy in the TCIG module alleviates texture diffusion in scale space, enhancing the network’s ability to represent texture features. Subsequently, the CMKI module utilizes the response differences between different features to generate convolution kernels from deep feature maps, enabling feature interaction between the PAN modal and MS modal. This reduces redundant representations of essential image content information in deep features of two modalities, facilitating a better mapping between dual-modal features and categories. Our results achieve state-of-the-art performance on multiple datasets. The code is available athttps://github.com/Xidian-AIGroup190726/SNCFNet.
Hao Zhu 0009, Wenhao Zhao, Biao Hou, Changzhe Jiao, Zhongle Ren, Wenping Ma 0001, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.6
2023 Gaussian Synthesis for High-Precision Location in Oriented Object Detection
abstract
In aerial image scenes, the objects have properties of arbitrary orientation, large-scale range, and dense distribution. Thus, the object detector uses oriented bounding box (OBB) to locate objects, which is more complex and challenging than horizontal bounding box (HBB) detector. Mainstream OBB detectors mostly use one-to-many label assignment strategy to predict multiple bounding boxes for the same object, and filter out repeat predictions by non-maximum suppression (NMS). NMS ranks with confidence and drops the detection box with IoU higher than the threshold, which is easy to get the local optimum result. The clustered synthesis method gets more accurate results than the original NMS, but applying it to the OBB detector leads to border shift, which arises from the angular discontinuity problem. Therefore, we use Gaussian OBB (G-OBB) to deal with the angular discontinuity and thus eliminate the offset generated by direct synthesis. G-OBB is not an easy to understand and describe representation. For this reason, we analyze the properties of G-OBB, and design a decoding method to convert a G-OBB to a rotated rectangular box, further discussing its conditions. Based on the decoding method, we propose a Gaussian synthesis algorithm (GauS), which transforms the OBB into Gaussian space, followed by synthesis, and finally transforms the synthesis result back into a new OBB. We have derived the synthesis and decoding methods, and further verified their effectiveness. The extensive experiments on several existing models show that GauS takes very little computation and improves detector’s high-precision performance. Extensive experiments verify the effectiveness, stability, and universality of the proposed algorithm. In addition, The RTMDet using GauS achieves a performance of 81.61 AP50and gains a 0.39% improvement in mAP, which achieves the SOTA performance. Our implementation is available at: https://github.com/lzh420202/GauS.
Biao Hou, Zitong Wu, Bo Ren 0001, Zhongle Ren, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.5
2023 A Dual-Stream Transformer With Diff-Attention for Multispectral and Panchromatic Classification
abstract
To minimize the feature redundancy of multispectral (MS) and panchromatic (PAN) images and maximize the complementary advantages of PAN and MS, a Dual-Stream Transformer with Diff-attention (DSTD)-Net is proposed for PAN and MS classification in this paper. Firstly, in terms of feature extraction, we use Self-attention and Co-attention (SCA) block to extract both specific advantageous features and common essential features. Based on that, a self-attention module strengthened by diff-attention (SSDA) that pays attention to the difference between two specific advantageous features is designed to reduce the essential redundancy in specific features. It can take advantage of the difference between two specific features and reduce the essential redundancy of the specific advantageous features, making them purer and better for classification. Finally, since the specific features and common features of multispectral (MS) and panchromatic (PAN) images make different contributions to classification, a Multi-stage Gated Fusion (MGF) strategy is used. The MGF strategy mainly uses Gated multisource units (GMU) to adapt the weight of different features and fuse them. So, our MGF strategy can strengthen the specific advantageous features beneficial for classification. Above all, the several experiment results verify our proposed networks’ effectiveness and robustness. Our code is available at: https://github.com/blackkiring/DSTD.
Lin Xu 0012, Hao Zhu 0009, Licheng Jiao, Wenhao Zhao, Biao Hou, Zhongle Ren, Wenping Ma 0001
IEEE Trans. Geosci. Remote. Sens.7
2022 Network Pruning for Remote Sensing Images Classification Based on Interpretable CNNs
abstract
Convolutional neural network (CNN)-based research has been successfully applied in remote sensing image classification due to its powerful feature representation ability. However, these high-capacity networks bring heavy inference costs and are easily overparameterized, especially for the deep CNNs pretrained on natural image datasets. Network pruning is regarded as a prevalent approach for compressing networks, but most existing research ignores model interpretability while formulating pruning criterion. To address these issues, a network pruning method for remote sensing image classification based on interpretable CNNs is proposed. More specifically, an original interpretable CNN with a predefined pruning ratio is trained at first. The filters, namely channels in the high convolutional layer, are able to learn specific semantic meanings in proportion to the predefined pruning ratio. The filters without interpretability are supposed to be removed. As for other convolutional layers, a sensitivity function is designed to assess the risk of pruning channels for each layer, and furthermore, the pruning ratio for each layer is corrected adaptively. The pruning method based on the proposed sensitivity function is effective and requires little computational costs to search abandoned channels without damaging classification performance. To demonstrate the effectiveness, the proposed method is implemented on different scales of modern CNN models, including VGG-VD and AlexNet. The experimental results, obtained on the UC Merced dataset and NWPU-RESISC45 dataset, prove that our method significantly reduces the inference costs and improves the interpretability of networks.
Xianpeng Guo, Biao Hou, Bo Ren 0001, Zhongle Ren, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.4
2022 A Collaborative Correlation-Matching Network for Multimodality Remote Sensing Image Classification
abstract
Recently, with the increasing availability of the high-quality panchromatic (PAN) and multispectral (MS) remote sensing (RS) images, the inherent complementarity between PAN and MS images provides a wide development prospect for the multimodality RS image classification task. However, how to cleverly relieve the modal differences and effectively integrate the single-modality PAN and MS features is still a challenge. In this article, we design a collaborative correlation-matching network (CCM-Net) for multimodality RS image classification. Concretely, we first propose a bidirectional dominant feature supervision (Bi-DFS) learning, it utilizes single-modality dominant features as supplementary supervision information to establish the joint optimization loss function, thereby adaptively narrowing the differences between modalities before the feature extraction. In the feature extraction stage, the interactive correlation feature matching (ICFM) learning, composing the spatial feature matching (Spa-FM) and spectral feature matching (Spe-FM) strategies, is proposed to establish interactive matching and enhancement between multimodality strong correlation features from the perspective of spatial and spectral, respectively, thereby effectively alleviating the semantic deviation of multimodality features. Finally, we aggregate finer multilevel multimodality features to obtain top-level features with high discrimination. The effectiveness of the proposed algorithm has been verified on multiple datasets. Our code is available at:https://github.com/Momuli/CCM-Net.git.
Wenping Ma 0001, Hao Zhu 0009, Kenan Sun, Zhongle Ren, Xu Tang 0004, Biao Hou, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.5
2021 Cost-Sensitive Latent Space Learning for Imbalanced PolSAR Image Classification
abstract
Land cover classification is an important application for polarimetric synthetic aperture radar (PolSAR) image interpretation. The classification performance of a promising parametric feature and classifier learning-based algorithm is limited when the amounts of pixels from different classes vary greatly. PolSAR data from minority classes is difficult to recognize correctly owing to a strong learning bias toward the majority classes, resulting in under-performing features for minority classes. To address this issue, a cost-sensitive latent space learning network based on the feature and classifier learning framework is proposed to reduce the learning bias for supporting the classification of imbalanced data in PolSAR images. First, a new cost-sensitive method is developed by adaptively computing the cost coefficient from predicted labels in the optimization process. Thus, the imbalanced distribution of PolSAR data can be obtained for both the labeled and unlabeled pixels rather than a predefined misclassifying matrix for labeled pixels. Second, latent space learning is used as an auxiliary task to assist the main task of classifier learning. By weighting the distance between the learned feature and the basis of the latent space with a different cost-sensitive coefficient, pixels in minority and majority classes are promoted to be more separable. Thus, the strong bias to majority classes is reduced from both the feature learning and classification process. Finally, the proposed method is studied through experiments on three different PolSAR images with several existing state-of-the-art methods. The experiments validate the effectiveness of the proposed method for balanced and imbalanced PolSAR land cover classification.
Biao Hou, Zaidao Wen, Zhongle Ren, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.4
2020 Object Detection in High-Resolution Panchromatic Images Using Deep Models and Spatial Template Matching
abstract
Automatic object detection from remote sensing images has attracted a significant attention due to its importance in both military and civilian fields. However, the low confidence of the candidates restricts the recognition of potential objects, and the unreasonable predicted boxes result in false positives (FPs). To address these issues, an accurate and fast object detection method called the refined single-shot multibox detector (RSSD) is proposed, consisting of a single-shot multibox detector (SSD), a refined network (RefinedNet), and a class-specific spatial template matching (STM) module. In the training stage, fed with augmented samples in diverse variation, the SSD can efficiently extract multiscale features for object classification and location. Meanwhile, RefinedNet is trained with cropped objects from the training set to further enhance the ability to distinguish each class of objects and the background. Class-specific spatial templates are also constructed from the statistics of objects of each class to provide reliable object templates. During the test phase, RefinedNet improves the confidence of potential objects from the predicted results of SSD and suppresses that of the background, which promotes the detection rate. Furthermore, several grotesque candidates are rejected by the well-designed class-specific spatial templates, thus reducing the false alarm rate. These three parts constitute a monolithic architecture, which contributes to the detection accuracy and maintains the speed. Experiments on high-resolution panchromatic (PAN) images of satellites GaoFen-2 and JiLin-1 demonstrate the effectiveness and efficiency of the proposed modules and the whole framework.
Biao Hou, Zhongle Ren, Wei Zhao 0014, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.2
2020 A Distribution and Structure Match Generative Adversarial Network for SAR Image Classification
abstract
Synthetic aperture radar (SAR) image classification is a fundamental research in the interpretation of SAR images. The previous methods are unilaterally based on statistical features or spatial features, which cannot capture features with complete SAR image characteristics and unavoidably limits the performance for classification. In this article, novel sample weighting and class adversarial training strategies are proposed to fuse complementary SAR characteristics. Based on these, a distribution and structure match auxiliary classifier generative adversarial network (DSM-ACGAN) is constructed for high-quality discriminative feature learning. Particularly, the characteristics of statistical distribution and spatial structure are jointly considered in class adversarial training of DSM-ACGAN. On the one hand, DSM-ACGAN sets the true SAR image characteristics as goals for the generator to learn generative models of each category. On the other hand, and more importantly, it guides the discriminator to simultaneously capture the desired statistical and structural features. Through the class adversarial processing, the discriminative feature learning progressively improves and contributes to classification. Additionally, class-balanced and plausible samples can be generated. Experimental results on three broad SAR images from different satellites confirm the effectiveness of class adversarial training and the superiority of discriminative feature learning in DSM-ACGAN. Visual performance and quantitative metrics also show the state-of-the-art performance of the novel model.
Zhongle Ren, Biao Hou, Zaidao Wen, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.1
2019 Pixel Dag-Recurrent Neural Network for Spectral-Spatial Hyperspectral Image Classification
abstract
Exploiting rich spatial and spectral features contributes to improve the classification accuracy of hyperspectral images (HSIs). In this paper, based on the mechanism of the population receptive field (pRF) in human visual cortex, we further utilize the spatial correlation of pixels in images and propose pixel directed acyclic graph recurrent neural network (Pixel DAG-RNN) to extract and apply spectral-spatial features for HSIs classification. In our model, an undirected cyclic graph (UCG) is used to represent the relevance connectivity of pixels in an image patch, and four DAGs are used to approximate the spatial relationship of UCGs. In order to avoid overfitting, weight sharing and dropout are adopted. The higher classification performance of our model on HSIs classification has been verified by experiments on three benchmark data sets.
Xiufang Li, Qigong Sun, Lingling Li 0002, Zhongle Ren, Fang Liu 0001, Licheng Jiao
IGARSS4