Bo Ren 0001

dblp:07/1796-1 · DBLP profile ↗
← Back
44ranked-venue papers
9as first author
34since 2021 · last 2026
0000-0002-0481-5069ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 41 · 9 first-author · 31 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 CSFE-Net: Cycle-consistency scattering feature extraction network for PolSAR image
Biqi Li, Chen Yang 0027, Biao Hou, Bo Ren 0001, Licheng Jiao
Neurocomputing5
2025 Interpretable Fine-Grained Aircraft Classification Network for Remote Sensing Image With Image Pair Interaction and Neural Tree
abstract
This paper proposes a novel interpretable framework for fine-grained aircraft classification in high-stakes remote sensing applications. Our approach addresses three key challenges: small inter-class variance, large intra-class variance, and the need for model interpretability. Specifically, our framework is built on the Swin Transformer (SwinT) backbone and includes three main modules. First, we present the Dynamic Attention Fusion Module (DAFM), which adaptively fuses multi-stage attention maps from the SwinT backbone. By leveraging a dispersion-based weighting mechanism, DAFM balances the contributions of coarse and fine-grained features, capturing both global structures and localized details. Second, we propose the Adaptive Image Pair Interaction Module (AIPI), which dynamically adjusts feature interaction strategies based on intra-class and inter-class similarity, effectively enhancing informative regions and improving robustness. To further optimize discriminative power, we incorporate an AIPI loss function that enforces intra-class consistency and inter-class separability. Finally, we develop a Binary Neural Tree Module (BNTM) to hierarchically select and propagate informative image patches, enhancing both feature refinement and interpretability through explicit path-based decision-making. Extensive experiments on benchmark datasets demonstrate that our framework significantly improves classification accuracy and interpretability, making it well-suited for applications requiring transparent and reliable decision-making. The Code can be found at https://github.com/StarmanGzx/BNTM.
Zhengxi Guo, Biao Hou, Xianpeng Guo, Chen Yang 0027, Zitong Wu, Bo Ren 0001, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.6
2025 DCIFNet: Cross-Modal Fusion With Correction and Interaction for Optical-SAR Land Cover Classification
abstract
Land cover classification (LCC) based on remote sensing image segmentation is a prominent task of remote sensing data interpretation. The commonly used optical data is susceptible to the weather, so it has the potential to utilize complementary features from the supplementary synthetic aperture radar (SAR) data to enhance segmentation performance. However, current multi-modal segmentation methods focus on the deep fusion of features, which usually ignores the significance of structural consistency information. In order to make use of the mutual correction and information exchange between multi-modal data, we propose DCIFNet, a dual-stream correction-interaction-fusion multi-modal LCC network. Specifically, we design a differential feature correction and enhancement module (DF-CEM) that leverages bidirectional differential features to correct multi-modal features. In addition, for corrected feature pairs, we deploy a parallel attention interaction module (PAIM) to focus on the pixel-level feature correlation and achieve effective information exchange in both channel and spatial dimensions. Through the expert fusion module (EFM), DCIFNet leverages the gate network to attain a flexible and compact feature fusion between multi-modal features. Experimental results show that our method achieves a superior performance compared with other multi-modal fusion segmentation methods on three optical-SAR datasets. The source code of DCIFNet is publicly available at https://gitee.com/asdwer2046/dcifnet.
Bo Ren 0001, Bo Liu 0009, Qianfang Wang, Biao Hou, Chen Yang 0027, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.1
2025 Incremental Land Cover Classification via Soft Label and Subregion Distillation
abstract
With the exponential growth of satellite remote sensing data, land cover classification models must adapt continuously to new classes. However, conventional incremental learning methods face critical challenges: catastrophic forgetting degrades recognition of old classes, and the softmax function further suppresses old-class probabilities due to ”class crowding.” Existing distillation techniques also struggle to transfer features in irregular geospatial regions. To address these issues, we propose Soft Labels and Subregion Distillation (SLSRD). SLSRD mitigates class crowding by employing soft labels instead of hard labels, derived from a hybrid of softmax and sigmoid outputs that preserve richer probabilistic information. Concretely, the soft label is a convex combination of softmax- and sigmoid-based probabilities that preserves inter-class relations while relaxing over-confident exclusivity for newly introduced categories, and it supervises all pixels across stages. In parallel, a breadth-first search identifies subregions within each image, which are weighted by probability and size, and similarity between corresponding subregions of the old and new models is maximized. This dual strategy effectively transfers fine-grained knowledge and overcomes the limitations of conventional distillation methods, particularly for large-scale remote sensing imagery. Experiments on three benchmark datasets-Vaihingen, GID, and FBP-demonstrate that SLSRD outperforms traditional methods, significantly improving incremental land cover classification.
Bo Ren 0001, Zhao Wang 0011, Hanyuan Ge, Biao Hou, Bo Liu 0009, Chen Yang 0027, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.1
2025 Contour-Aware Dynamic Low-High Frequency Integration for Pan-Sharpening
abstract
Pan-sharpening is the process of fusing panchromatic (PAN) and multispectral (MS) images. Its critical focus lies in accurately capturing the contour information from the PAN image during the fusion process and presenting it at a high resolution. However, existing deep learning methods lack the precise capture of delicate and smooth contour information, resulting in contour diffusion that affects the fusion results. Therefore, we introduce contourlet decomposition to capture multiscale directional delicate contour features and construct multiscale graph structures for semantic mining of dual-source contour features, continually updated through dynamic learning. By incorporating global features, we guide the multihead attention mechanism with directional decoding, enabling the network to pay more attention to high-resolution contour features, thereby gaining an advantage in image reconstruction. Cross-decoding between modalities provides strong representational capabilities for the advantageous features of both modalities, effectively enhancing the sharpening effect. Our algorithm achieves state-of-the-art results, and its effectiveness and advantages have been thoroughly validated across multiple datasets, including GaoFen-2, WorldView2, WorldView3, etc. Our code is available athttps://github.com/Xidian-AIGroup190726/CDFInet.
Xiaoyu Yi 0002, Hao Zhu 0009, Pute Guo, Biao Hou, Bo Ren 0001, Xiaoteng Wang, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.5
2024 Masked Angle-Aware Autoencoder for Remote Sensing Images
Zhihao Li 0005, Biao Hou, Siteng Ma, Zitong Wu, Xianpeng Guo, Bo Ren 0001, Licheng Jiao
ECCV (8)6
2024 Few-Shot Class Incremental Land Cover Classification with Masked Exemplar Set
abstract
Land cover categories and features change with time. It generates demand for developing incremental learning methods for land cover classification. Meanwhile, due to the expensive cost of sample annotation, it is very difficult to obtain a large amount of annotated data for training. Therefore, how to effectively develop a few-shot incremental semantic segmentation method for land cover classification has become a significant task for remote sensing data interpretation. In this paper, we propose a novel data replay method with masked exemplar set (RMES) to improve land cover classification performance under the condition of few samples. It maintains a masked sample queue for each class. In this method, at the end of each learning step, two operations need to be run, the threshold sample filtering operation and the sample masking storage operation. These two operations update the sample queue and make it part of the training set in the next incremental learning stage. This alleviates overfitting and catastrophic forgetting problems. As a result of the experiment, the proposed RMES had superior performance in the CCF dataset.
Junxi Guo, Bo Ren 0001, Zhao Wang 0011, Biao Hou
IGARSS2
2024 SwinTFNet: Dual-Stream Transformer With Cross Attention Fusion for Land Cover Classification
abstract
Land cover classification (LCC) is an important application in remote sensing data interpretation. As two common data sources, SAR images can be regarded as an effective complement to optical images, which will reduce the influence caused by single-modal data. But common LCC methods are focusing on designing advanced network architectures to process single-modal remote sensing data. Few works have been oriented toward improving segmentation performance through fusing multi-modal data. In order to deeply integrate SAR and optical features, we propose SwinTFNet, a dual-stream deep fusion network. Through the global context modeling capability of Transformer structure, SwinTFNet models teleconnections between pixels in other regions and pixels in cloud regions for better prediction in cloud regions. In addition, a Cross-Attention Fusion Module (CAFM) is proposed to fuse features from optical and SAR data. Experimental results show that our method improves greatly in the classification of clouded images compared with other excellent segmentation methods and achieves the best performance on multi-modal data.
Bo Ren 0001, Bo Liu 0009, Biao Hou, Zhao Wang 0011, Chen Yang 0027, Licheng Jiao
IEEE Geosci. Remote. Sens. Lett.1
2024 Unreliable Pixel Contrast Based on von Mises-Fisher Distribution for Semi-Supervised SAR Segmentation
abstract
The unique visual properties and the huge size of synthetic aperture radar (SAR) images pose challenges in labeling data. The scarcity of labeled data limits the training of SAR image segmentation networks. To address this issue, semi-supervised methods are used to train the network. However, traditional semi-supervised algorithms like Mean Teacher often generate erroneous pseudo-labels that carry over to subsequent training epochs, leading to overfitting and affecting network performance. This overfitting stems from an overreliance on unreliable pixels in Mean Teacher. In this study, an enhanced approach is proposed called unreliable pixel contrast (UPCo), where a von Mises-Fisher distribution is applied to constrain unreliable pixels in feature space. We augment the segmentation network with a feature output header for pixel-level contrastive learning in UPCo. Moreover, to minimize the computational effort during the training phase, hard sample selection and negative object non-uniform selection strategies are designed to facilitate contrastive learning. The proposed UPCo was evaluated on two large-scene SAR images and demonstrated its superiority over other comparative algorithms, achieving more optimal performance in semi-supervised segmentation of SAR images.
Zitong Wu, Biao Hou, Xianpeng Guo, Bo Ren 0001
IEEE Geosci. Remote. Sens. Lett.6
2024 Model-Based Decomposition Feature Learning With Adversarial Prior
abstract
Model-based target decomposition method has been widely applied due to its clear physical scattering significance. However, after establishing decomposition basis, the process of solving the scattering components and parameters is usually underdetermined, which will lead to the issues such as component negative power and overestimation. For this problem, this letter examines the target decomposition task from the perspective of deep learning and proposes an adversarial decomposition feature learning (ADFL) model. This model could learn decomposition features suitable for current terrain characteristics according to input data. At the same time, the model-based adversarial feature prior is embedded in ADFL to maintain the physical scattering meanings. On real PolSAR datasets, the learned features of proposed model are well correlated with real terrain scattering characteristics. Further, it avoids negative decomposition features and make more accurate fitting of scattering components, effectively alleviating the above problems.
Chen Yang 0027, Biao Hou, Bo Ren 0001, Jocelyn Chanussot, Licheng Jiao
IEEE Geosci. Remote. Sens. Lett.4
2024 MSRIP-Net: Addressing Interpretability and Accuracy Challenges in Aircraft Fine-Grained Recognition of Remote Sensing Images
abstract
The task of fine-grained aircraft recognition is crucial in the field of remote sensing. Despite some progress achieved by traditional deep learning methods in addressing this challenge, they are often perceived as a “black box,” lacking transparent explanations for model decisions. Current interpretable methods based on attention mechanisms, although providing some interpretability, do not align with human thought logic. Therefore, we propose a multiscale rotation-invariant prototype network (MSRIP-Net). Our approach simulates the intuitive reasoning process of humans in identifying objects by segmenting them into multiple components. Importantly, MSRIP-Net has the capability to automatically recognize rigid components on aircraft targets without relying on additional part annotations, using only image-level class labels. In addition, our approach effectively addresses challenges presented by noise, deformations, and multiscale variations in remote sensing targets and has been comprehensively evaluated on datasets FAIR1M1.0 and Rareplane. Our results demonstrate that MSRIP-Net achieves higher accuracy compared with existing fine-grained recognition methods. Furthermore, we provide insights into the model’s decision-making process to illustrate the interpretability of our approach.
Zhengxi Guo, Biao Hou, Xianpeng Guo, Zitong Wu, Chen Yang 0027, Bo Ren 0001, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.6
2024 MGC: MLP-Guided CNN Pretraining Using a Small-Scale Dataset for Remote Sensing Images
abstract
To overcome the inherent domain gap between natural images and remote sensing images (RSIs), it is highly desirable to develop pretraining methods specifically for RSIs. Considering the lack of widely recognized large-scale benchmarks like ImageNet in the RSI community and limited computational resources, this article proposes multilayer perceptron (MLP)-guided convolutional neural network (CNN) (MGC), a method that employs an MLP to guide the pretraining of a CNN from small-scale datasets for RSIs. MGC has two encoders, each consisting of a CNN branch and an MLP branch. We first contrast pairwise samples from the same type of branches or different types of branches across the encoders and employ a positive-pair guidance strategy to explore consistency. Due to the inherent locality issue of shallow layers in a CNN, the CNN branches often do not attend to correct foreground regions such as objects, regions of interest, and land coverage. Therefore, we further propose an attention guidance strategy to guide the CNN branches to focus on foreground regions and learn discriminative representations effectively. The proposed MGC method is validated by pretraining a CNN model using the MGC and applying it to different downstream tasks including scene classification, rotated object detection, semantic segmentation, and change detection on ten datasets. Results have confirmed the effectiveness of the proposed MGC. Our code will be released at:https://github.com/benesakitam/MGC.
Zhihao Li 0005, Biao Hou, Wanqing Li 0001, Zitong Wu, Bo Ren 0001, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.5
2024 Rebalancing Gaussian Location Loss for High-Precision Detection on Remote Sensing Images
abstract
Aerial image objects are usually orientated arbitrarily, with a large scale range, and densely distributed. Traditional horizontal bounding box (HBB) detectors tend to filter out densely distributed objects leading to missed detections, such as ship (SH) and vehicle. Therefore, oriented object detection has become a mainstream solution in recent years. The 2-D Gaussian distribution representation of the oriented bounding boxes (OBBs) solves the problem of angular discontinuity and boundary discontinuity and thus gets more attention. However, as the aspect ratio of the object gradually decreases, its predicted angular performance continues to decrease. We find that the angular gradient of an object decreases sharply as the aspect ratio decreases, resulting in a large gradient gap between a small aspect ratio object (SARO) and a large aspect ratio object (LARO). It makes the detector prefer to ignore SARO during training, which weakens the high precision performance of SARO. We call this phenomenon shape imbalance. To solve the problem, we proposed a simple gradient rebalancing strategy named shape balance. Since the shape imbalance is only related to the aspect ratio of the object, we designed a modulation function with an inverse aspect ratio to calculate the balance coefficient. The principle of the function is that the larger the aspect ratio, the smaller the balance coefficient; the smaller the aspect ratio, the larger the balance coefficient. We aim to get the balance coefficients for objects with different aspect ratios. Location loss multiplied by a balance coefficient can directly adjust the gradient gap between objects with different aspect ratios to achieve a rebalancing effect. Extensive experiments conducted on DOTA-v1.0 dataset and DIOR-R dataset verify the effectiveness of our proposed method. Our method improves the detection performance of Gaussian location loss by an average of 2.08%/1.01%(AP75/mAP) metrics on the DOTA-v1.0 dataset and 1.17%/0.82%(AP75/mAP) improvements for DIOR-R dataset.
Biao Hou, Zitong Wu, Xianpeng Guo, Bo Ren 0001, Zhongle Ren, Chen Yang 0027, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.5
2024 LM-Net: A Lightweight Matching Network for Remote Sensing Image Matching and Registration
abstract
Deep feature learning methods have shown significant advantages over handcrafted feature-based methods in remote sensing image matching and registration. Existing deep learning methods usually introduce complex modules into the deep convolutional network for more robust feature learning. However, they usually require high computation and memory resources for the computing device and have expensive time costs for image registration. As a basic image-processing task, it is crucial to build a lightweight matching network (LM-Net) for fast and accurate image matching and registration. Unfortunately, the image-matching performance will decrease significantly when we directly compress the deep model to a lightweight one. This article proposes an LM-Net based on the knowledge distillation (KD) learning framework for remote sensing image matching and registration. We first build an LM-Net with three convolutional layers. Then, this article proposes an effective KD approach for network optimization, which transfers the effective knowledge from the deep matching network to LM-Net to improve image-matching performances. Specifically, this article considers the useful information in the instance samples and the relation information between samples. It designs the feature and feature relation distillation learning for LM-Net training. Extensive experimental results and analysis have shown the effectiveness and advantages of the proposed LM-Net. LM-Net can reduce the number of parameters and computational complexity of the matching network. Meanwhile, LM-Net can significantly decrease the time cost and achieve results comparable to those of the deep model. It reduces the average image registration time by 42% on remote sensing image matching and registration. Additionally, LM-Net generalizes well on other multimodal remote sensing images.
Dou Quan, Chonghua Lv, Shuang Wang 0001, Yi Li 0054, Bo Ren 0001, Jocelyn Chanussot, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.6
2024 F3Net: Adaptive Frequency Feature Filtering Network for Multimodal Remote Sensing Image Registration
abstract
Multimodal remote sensing image registration is crucial for multimodal information fusion and applications. The significant nonlinear appearance difference between multimodal images caused by the various imaging mechanisms dramatically increases the challenge of image registration. This article proposes an adaptive frequency feature filtering network (F3Net) for cross-modal remote sensing image registration. On the one hand, F3Net explicitly explores the useful frequency components across modal images based on multilevel deep features. On the other hand, F3Net can take advantage of the nonlocal receptive fields by frequency modulation for feature learning and boosting image registration performances. F3Net inserts frequency feature filtering (F3) modules in multilevel deep features. Specifically, F3Net first performs the fast Fourier transform (FFT) for deep features. Then, F3Net designs a frequency attention (FA) module to adaptive enhance the shared and discriminative frequency features between multimodal images while suppressing the frequency components that hinder the cross-modal image registration. In addition, F3Net adopts multiscale frequency filtering fusion to facilitate discriminative feature learning, including global frequency feature filtering (GF3) based on the global image spectrum and local frequency feature filtering (LF3) based on the spectrum of stacked image regions. Experimental results on many remote sensing images have demonstrated the efficiency of the F3Net on multimodal image registration.
Dou Quan, Shuang Wang 0001, Yunan Li 0001, Bo Ren 0001, Mengte Kang, Jocelyn Chanussot, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.5
2024 Self-Supervised Learning Guided by SAR Image Factors for Terrain Classification
abstract
Effective feature representation is the key to SAR image terrain classification. Limited by the abstract appearance and the scarcity of high-quality labeled data in this field, the features learned by current methods, especially deep learning models, do not have enough directivity and applicability, which hampers the performance. This paper proposes Multi-image Factor Self-Supervised Learning(MFSSL) to achieve directional feature learning and obtain generalized features with few patch-level labeled data. The framework consists of an upstream multi-factor image style transfer task and a downstream terrain classification task. In the upstream task, the goal of feature learning is first set up by multiple SAR image factors, including the observation region, the terrain category, and the imaging parameters. And then, different styles of SAR terrain images are generated and reconstructed under this goal. Through this bidirectional generative learning, the low-level external appearance of the terrain is removed, while the essential and discriminative feature representation is retained and shared across different factors. Finally, the downstream model inherits the general feature from the upstream model and implements the terrain classification task using a small amount of labeled data. Experiments conducted on three broad SAR scenes with different image factors demonstrate that the proposed framework can improve pixel-level terrain classification only with a few patch-level labeled data.
Zhongle Ren, Zhe Du, Biao Hou, Weibin Li 0002, Hao Zhu 0009, Bo Ren 0001, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.7
2024 Few-Shot MS and PAN Joint Classification With Improved Cross-Source Contrastive Learning
abstract
The joint classification of multispectral (MS) and panchromatic (PAN) images aims to provide a more detailed and accurate interpretation of land features. Although deep-learning-based methods have achieved remarkable success in this task, the generalization performance of networks is compromised when labeled samples are insufficient. In this study, we explore the possibility of leveraging unlabeled remote sensing images (RSIs) through contrastive learning and demonstrate the challenges associated with directly applying contrastive learning to RSIs. To end this, we propose a cross-source contrastive learning method for few-shot MS and PAN joint classification (CrossCLMP), which aims to learn sufficient transferable representations in a self-supervised contrastive manner so as to provide a robust pretrained model for fine-tuning the downstream joint classification task. Specifically, we design: 1) intersource and intrasource alignment loss (ER-Align) to achieve self-supervised feature extraction and alignment; 2) the source-unique feature adaptive separation (SUAS) strategy to model source-unique information explicitly; and 3) the auxiliary contrastive learning (ACL) strategy to mitigate the adverse impact of numerous false-negative samples in the pretraining stage. The experimental results and the theoretical analyses on multiple popular datasets comprehensively demonstrate the effectiveness and robustness of the proposed method under few-shot. Our code is available at:https://github.com/Xidian-AIGroup190726/CrossCLMP.
Hao Zhu 0009, Pute Guo, Biao Hou, Changzhe Jiao, Bo Ren 0001, Licheng Jiao, Shuang Wang 0001
IEEE Trans. Geosci. Remote. Sens.6
2023 Multi-Source Fusion Network for Remote Sensing Image Segmentation with Hierarchical Transformer
abstract
Recently, due to the limitations of single sensor, it is hard to improve the performance of land cover classification. The traditional image segmentation methods can not process the optical remote sensing images effectively, especially when optical sensor is affected by complex weather conditions. However, as an active radar, synthetic aperture radar(SAR) has the advantage of not being restricted by weather conditions with the the penetrability of electromagnetic radiation. So multi-sensor data fusion provides a great potential for land cover classification. In this paper, a new fusion network called SegFusion is proposed to improve the performance of land cover classification. There are two main components in SegFusion which are hierarchical Transformer encoder and Swin-Fusion(SW-Fusion) module. First, a hierarchical Transformer encoder is used to extract multilevel feature of optical and SAR images. By integrating features from different layers, we can obtain powerful representation that combines both low-resolution fine-grained features and high-resolution coarse-grained features. Second, SW-Fusion module is used to fuse the features of optical and SAR data. In SW-Fusion, we use modified Swin Transformer [1] block with multi-head cross-attention mechanism to exchange information between features from different sources.
Bo Liu 0009, Bo Ren 0001, Biao Hou, Yu Gu 0015
IGARSS2
2023 RotPointNet: Keypoint based Oriented Object Detection for Aerial Image
abstract
Remote sensing object detection is characterized by arbitrary direction, dense targets and variable scale. Like common object detection methods predict horizontal rectangle boxes directly, most of the existing remote sensing oriented object detection methods predict oriented rectangle boxes directly. i.e. Center point position, length, width and rotation angle of the oriented rectangle boxes. These models often need to design complex rotation detection module to adapt to the rotating characteristics of targets, they lack good performance on targets with large aspect ratio changes as well. In this paper, we propose a keypoint based oriented object detection method for aerial images named RotPointNet to improve detection accuracy of rotating targets with large aspect ratio change. RotPointNet first locates the two endpoints of rotating targets using keypoint detection method, and then constructs the remote sensing oriented target based on the endpoints and additional width information. In addition, a new method of matching target key points is proposed in this paper. We conduct experiments to prove, the target detection model RotPoint-Net and key point matching strategy proposed in this paper can achieve good detection results on remote sensing images, especially for the detection of targets with large aspect ratio changes.
Chongyu Wang, Bo Ren 0001, Biao Hou
IGARSS2
2023 Incremental Land Cover Classification via Strategies for Edge Removal and Feature Point Aggregation
abstract
Convolutional neural networks will face the problem of catastrophic forgetting in the process of incremental learning. To solve this problem, we propose an incremental learning method called the strategy of edge removal and feature point aggregation, or ERFPA for short. In the cross-entropy loss, we perform edge detection and removal on the labels generated by the old model predictions, and then fuse them with the new labels. We calculate the mean point of different classes, and make the model learn features better by narrowing the distance with similar pixels. As demonstrated by the results of our experiment, on two remote sensing image datasets: CCF and Vaihingen, our method achieves state-of-the-art results.
Zhao Wang 0011, Bo Ren 0001, Biao Hou, Yu Gu 0015
IGARSS2
2023 Spatio-temporal segments attention for skeleton-based action recognition
Helei Qiu, Biao Hou, Bo Ren 0001
Neurocomputing3
2023 An Improved Neural Network Classification Algorithm by Expanding Training Samples for Polarimetric SAR Application
abstract
The idea of spatial correlation has been used in polarimetric synthetic aperture radar (PolSAR) classification for many years. It is common that the bigger the spatial correlation, the more information it contains. Though the recent advances in deep learning for PolSAR classification have achieved remarkable progress, the number of training samples is still a problem we have to face. Aiming to solve this problem, it is valuable to explore a greater spatial correlation of terrains as a priori for classification. Considering the correlated regions of the training sample as a prior, rather than a small neighborhood, the paper proposed Wishart locally constrained expansion algorithm. Based on this, a PolSAR classification algorithm is designed. The whole proposed PolSAR classification algorithm includes 3 parts: Wishart locally constrained expansion algorithm, convolution neural network training algorithm, and Markov random field post-processing algorithm. Supervised by the cluster knowledge from Wishart classifier, Wishart locally constrained expansion algorithm expands samples iteratively from the regions related to the training samples with a region-expanding technique, where very few labeled samples turn to a larger training set. Convolution neural network algorithm is designed to train convolution neural network with 3 types of training samples for a more accurate result. Convolution neural network is trained firstly with the numerous training samples generated through Wishart locally constrained expansion algorithm, and then the samples from consistency extraction and raw samples are involved to fine-tune the convolution neural network to get the improved result. Finally, Markov random field prior is used to smooth the result. Several benchmark datasets are adopted to evaluate the effectiveness of the proposed algorithm. The experimental results show that the new semi-supervised classification algorithm outperforms the state-of-the-art semi-supervised classification algorithms and classical supervised classification algorithms.
Biao Hou, Bo Ren 0001, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.3
2023 Complete Rotated Localization Loss Based on Super-Gaussian Distribution for Remote Sensing Images
abstract
Localization regression in oriented object detection tasks has long faced boundary discontinuity and angular discontinuity problems induced by periodic angles. These problems were successfully resolved by using a 2d Gaussian distribution to modelling the oriented bounding box (OBB). However, the angular information of square-like objects will be lost when they are converted to 2d Gaussian distribution, forming a systematic problem. Its fundamental reason is that when the aspect ratio of the object tends to 1, the equiprobability curve of 2d Gaussian distribution degenerates from an ellipse to a circle, thus losing the orientation information of the rotated object. This results in the bounding boxes of such square-like objects not being learned effectively. To resolve this problem, we used the Lamé curve (or superellipse) to modify the existing 2d Gaussian function and designed a super-Gaussian distribution. This distribution can maintain anisotropy at arbitrary aspect ratios, thus preserving the angular information of the oriented object. We used the Kullback-Leibler (KL) divergence to measure the distance between two super-Gaussian distributions and convert it into a localization loss (SGKLD) by a function. SGKLD is an improved version of KLD loss. By modifying the form of the probability distribution, we elegantly fix the angle missing problem of the traditional Gaussian distribution. We validated the effectiveness of the proposed algorithm on several datasets and obtained the performance of SOTA. Our algorithm achieves a mean average precision (mAP) of 80.07, 76.59, 62.27, and 90.55/98.13 on the DOTA-v1.0, DOTA-v1.5, DOTA-v2.0, and HRSC2016 datasets, respectively.
Biao Hou, Zitong Wu, Zhengxi Guo, Bo Ren 0001, Xianpeng Guo, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.5
2023 Gaussian Synthesis for High-Precision Location in Oriented Object Detection
abstract
In aerial image scenes, the objects have properties of arbitrary orientation, large-scale range, and dense distribution. Thus, the object detector uses oriented bounding box (OBB) to locate objects, which is more complex and challenging than horizontal bounding box (HBB) detector. Mainstream OBB detectors mostly use one-to-many label assignment strategy to predict multiple bounding boxes for the same object, and filter out repeat predictions by non-maximum suppression (NMS). NMS ranks with confidence and drops the detection box with IoU higher than the threshold, which is easy to get the local optimum result. The clustered synthesis method gets more accurate results than the original NMS, but applying it to the OBB detector leads to border shift, which arises from the angular discontinuity problem. Therefore, we use Gaussian OBB (G-OBB) to deal with the angular discontinuity and thus eliminate the offset generated by direct synthesis. G-OBB is not an easy to understand and describe representation. For this reason, we analyze the properties of G-OBB, and design a decoding method to convert a G-OBB to a rotated rectangular box, further discussing its conditions. Based on the decoding method, we propose a Gaussian synthesis algorithm (GauS), which transforms the OBB into Gaussian space, followed by synthesis, and finally transforms the synthesis result back into a new OBB. We have derived the synthesis and decoding methods, and further verified their effectiveness. The extensive experiments on several existing models show that GauS takes very little computation and improves detector’s high-precision performance. Extensive experiments verify the effectiveness, stability, and universality of the proposed algorithm. In addition, The RTMDet using GauS achieves a performance of 81.61 AP50and gains a 0.39% improvement in mAP, which achieves the SOTA performance. Our implementation is available at: https://github.com/lzh420202/GauS.
Biao Hou, Zitong Wu, Bo Ren 0001, Zhongle Ren, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.4
2023 Automatic Aug-Aware Contrastive Proposal Encoding for Few-Shot Object Detection of Remote Sensing Images
abstract
In the annotation of remote sensing images (RSIs), the effectiveness of common object detection methods trained on only a few samples decreases instantly, which has prompted increasing research on the few-shot problem in remote sensing. RSIs often exhibit suboptimal performance in few-shot scenarios due to the intricate nature of scene information interference and the high degree of cosine similarity, both of which present significant challenges to their effectiveness. In this paper, a two-stage detection framework based on fine-tuning is selected to deal with the common problems in the few-shot task of remote sensing domain. Considering the excessive scale variation of instances in remote sensing datasets, we introduce an automatically learned aug-aware search module to provide an intelligent data augmentation solution for Faster R-CNN using different optimal augmentation policies searched by the network to fit the current dataset. We introduce a contrastive RoI branch to better classify novel class proposal features that are easily confused by the base class. We named our work AACE and conducted extensive experiments on two common object detection datasets in remote sensing, NWPU VHR 10 and DIOR, on which AACE achieved about 2.30% and 2.61% improvement, respectively, in the number of shots listed in the paper, compared to other algorithms.
Siteng Ma, Biao Hou, Zitong Wu, Zhihao Li 0005, Xianpeng Guo, Bo Ren 0001, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.6
2023 Incremental Land Cover Classification via Label Strategy and Adaptive Weights
abstract
During incremental learning tasks, catastrophic forgetting occurs when old models are updated with new information. To address this issue, we propose a novel method called label strategy and adaptive weights (LSAW) that improves the incremental learning process. The label strategy introduces the old classes and solves the problem of how to reasonably use the wrong samples predicted by the old model. In the cross-entropy (CE) loss, we apply a threshold to filter the pseudolabels predicted by the old model. Subsequently, we merge the pixel samples with high probability with the current label. The probability here refers to the probability that the pixel belongs to the true class. This process enables the introduction of information from old classes that are not directly accessible in the current stage. Moreover, this information is relatively reliable, and the model exhibits confidence in its accuracy. For the remaining pixels, we retain all classes’ information through label smoothing. In the distillation function, the old class and background pixel samples are selected for distillation according to the prediction map of the old classes. The weights of the classes are adaptively updated and adjusted using specific label information from each batch and the different stages of incremental learning. As demonstrated by the results of our experiment, on three remote sensing image datasets: China Computer Federation (CCF), Potsdam, and Vaihingen, our method achieves the best results.
Bo Ren 0001, Zhao Wang 0011, Biao Hou, Bo Liu 0009, Zitong Wu, Jocelyn Chanussot, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.1
2022 Remote Sensing Object Tracking With Deep Reinforcement Learning Under Occlusion
abstract
Object tracking is an important research direction of space Earth observation in the field of remote sensing. Although the existing correlation filter-based and deep learning (DL)-based object tracking algorithms have achieved great success, they are still unsatisfactory for the problem of object occlusion. The occlusion caused by the complex change in background, and the deviation of the tracking lens, causes object information to go missing, which leads to the omission of detection. Traditionally, most methods for object tracking under occlusion adopt a complex network model, which redetects the occluded object. To address this issue, we propose a novel object tracking approach. First, an action decision-occlusion handling network (AD-OHNet) based on deep reinforcement learning (DRL) is built to achieve low computational complexity for object tracking under occlusion. Second, the temporal and spatial context, the object appearance model, and the motion vector are adopted to provide the occlusion information, which drives actions in reinforcement learning under complete occlusion and contributes to improving the accuracy of tracking while maintaining speed. Finally, the proposed AD-OHNet is evaluated on three remote sensing video datasets of Bogota, Hong Kong, and San Diego taken from Jilin-1 commercial remote sensing satellites. The video datasets all shared problems of low spatial resolution, background clutter, and small objects. Experimental results on the three video datasets validate the effectiveness and efficiency of the proposed tracker.
Yanyu Cui, Biao Hou, Bo Ren 0001, Shuang Wang 0001, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.4
2022 Network Pruning for Remote Sensing Images Classification Based on Interpretable CNNs
abstract
Convolutional neural network (CNN)-based research has been successfully applied in remote sensing image classification due to its powerful feature representation ability. However, these high-capacity networks bring heavy inference costs and are easily overparameterized, especially for the deep CNNs pretrained on natural image datasets. Network pruning is regarded as a prevalent approach for compressing networks, but most existing research ignores model interpretability while formulating pruning criterion. To address these issues, a network pruning method for remote sensing image classification based on interpretable CNNs is proposed. More specifically, an original interpretable CNN with a predefined pruning ratio is trained at first. The filters, namely channels in the high convolutional layer, are able to learn specific semantic meanings in proportion to the predefined pruning ratio. The filters without interpretability are supposed to be removed. As for other convolutional layers, a sensitivity function is designed to assess the risk of pruning channels for each layer, and furthermore, the pruning ratio for each layer is corrected adaptively. The pruning method based on the proposed sensitivity function is effective and requires little computational costs to search abandoned channels without damaging classification performance. To demonstrate the effectiveness, the proposed method is implemented on different scales of modern CNN models, including VGG-VD and AlexNet. The experimental results, obtained on the UC Merced dataset and NWPU-RESISC45 dataset, prove that our method significantly reduces the inference costs and improves the interpretability of networks.
Xianpeng Guo, Biao Hou, Bo Ren 0001, Zhongle Ren, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.3
2022 A Neural Network Based on Consistency Learning and Adversarial Learning for Semisupervised Synthetic Aperture Radar Ship Detection
abstract
Ship detection in synthetic aperture radar (SAR) images has important application value. Sea clutter, complex scenes, a large size change in ships, and the arbitrary directionality of ships make ship detection challenging. With the development of deep learning, many deep learning algorithms have been applied to SAR images. These algorithms need a lot of labeled data for training. It is time-consuming to label SAR data, and the unlabeled data are easy to obtain. It is necessary to use the unlabeled data effectively to improve the performance of the algorithm. In this study, a semisupervised SAR ship detection network, named the semisupervised consistency learning adversarial network (SCLANet), is presented. SCLANet is a two-stage detection network. The local features around the ship can be extracted by the SCLANet, and the features generated from unlabeled data become closer to those generated from labeled data by using adversarial learning. There are two consistency learning modules in SCLANet: noise robustness consistency learning and output encoding consistency learning. Noise robustness consistency learning can increase the robustness of the SCLANet. Maintaining consistency between the noisy results and the original results can train the unlabeled data. In output encoding consistency learning, outputs are mapped to a picture that is fed into an encoder to obtain the intermediate representation embedding. Another embedding is a layer in the main network of the SCLANet. Reducing the error between two embeddings can train the SCLANet with unlabeled data. Two types of consistency learning can be used as pretext tasks for semisupervised learning. Experiments were conducted on two SAR ship datasets. Compared with other algorithms, the SCLANet achieved the highest detection accuracy, indicating that it is more advantageous to use in ship detection.
Biao Hou, Zitong Wu, Bo Ren 0001, Xianpeng Guo, Shuang Wang 0001, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.3
2022 N-Cluster Loss and Hard Sample Generative Deep Metric Learning for PolSAR Image Classification
abstract
Deep learning works normally in PolSAR image classification because the complex terrain scattering characteristic results in large intraclass differences and high interclass similarity. Deep metric learning (DML) aims to make the features keep a closer intraclass and a farther interclass distance. Therefore, we introduce DML and then propose an N-cluster generative adversarial net (N-cluster GAN) framework for PolSAR image classification. However, existing DML losses mainly focus on the relationship between individual samples in feature space. Hence, we propose N-cluster loss that pays more attention to the overall structure of all samples. Meanwhile, traditional hard negative sample mining methods occupy lots of computational resources. In addition, the hard level of the negative samples will affect the model’s performance. Therefore, we explore a new method based on a GAN framework to replace the sample mining. Positive N-cluster loss is added to the discriminator ($D$), and a negative one is added to the generator ($G$). In this way,$D$will possess better classification ability, and$G$can produce hard negative samples for$D$. Then, the hard level of the generated negative samples will change with the discrimination of$D$, which is appropriate for the proposed model. N-cluster loss can be directly calculated through the extracted features rather than redundant data preparation. The proposed model is verified on four PolSAR datasets from two aspects of the loss function and negative samples mining. Then, it achieves competitive performance compared with state-of-the-art algorithms.
Chen Yang 0027, Biao Hou, Jocelyn Chanussot, Bo Ren 0001, Shuang Wang 0001, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.5
2022 Reconstruction Error-Based Decomposition Feature Selection for PolSAR Image
abstract
Target decomposition features are the cornerstone of subsequent analyses for PolSAR images. Generally, adopting single or several decomposition algorithms limits the representation ability for original terrain characteristics. Using all the existing decomposition features, however, will definitely increase computational complexity. Besides, some features even have a negative effect on the following tasks. To address these problems, a sparse variational autoencoder feature selection framework (SVAE-FS) is proposed in this article. In detail, the encoder transforms the original feature set into latent space and then decoder reconstructs the corresponding pseudo set on this latent space. Similarly, a pseudo subset is subsequently obtained by the SVAE. The discrepancy, namely reconstruction error, between the pseudo set and the pseudo subset is taken as an evaluation criterion which reflects the feature representation ability of pseudo subset. Sparse constraint in the encoder makes the representative features stand out. Meanwhile, the linear feature transformation layer of the encoder enables the SVAE to evaluate different scale subsets without repeated training. Finally, a greedy selection approach with search scale$K$is proposed to find the suboptimal subset. This procedure not only reduces time consumption, but also ensures the performance of the subset. The selected features are analyzed on four real PolSAR datasets according to the terrain scattering characteristics. Furthermore, these features have achieved competitive performance on three PolSAR image tasks.
Chen Yang 0027, Biao Hou, Xianpeng Guo, Bo Ren 0001, Jocelyn Chanussot, Shuang Wang 0001, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.4
2022 PDFL: Polarimetric Decomposition Feature Learning via Deep Autoencoder
abstract
Model-based polarimetric target decomposition (TD) generally solves scattering components and parameters under pre-set decomposition base, then decomposition features are also obtained. However, pre-set base could not be adjusted according to different scenes. Furthermore, solving the polarimetric parameters needs to explore additional information or consider limiting conditions to build equations, which is hard and easily to bring negative effects into decomposition features. To this end, we regard the TD as a process of learning decomposition base and features by deep learning. Then, the polarimetric decomposition feature learning (PDFL) model is proposed in this paper. Strictly, this model is not an incoherent TD method but a learning-based method. It dose not need to construct the parameter solution equations or fixed base. Then, the decomposition base and feature can be adaptively learned according to scattering characteristics of current dataset. Due to the characteristics of unsupervised reconstruction, deep auto encoder (DAE) is used as the model foundation. Then, some adjustments and constraints are utilized to make the DAE fit closely with TD. The encoder extracts latent vector from PolSAR data, then the decoder reconstructs pseudo data on this latent vector. The reconstruction can be regarded as the inverse process of TD, so the base matrix of decoder and the latent vector indicate the learned decomposition base and features when the model converges. The effectiveness of PDFL is verified on simulated and real PolSAR datasets. Compared with representative algorithms, proposed model gains more discriminative features and achieves competitive performance on terrain classification and segmentation tasks.
Chen Yang 0027, Biao Hou, Bo Ren 0001, Jocelyn Chanussot, Shuang Wang 0001, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.4
2021 Dual-Stream High Resolution Network for Multi-Source Remote Sensing Image Segmentation
abstract
Recently, the image segmentation has been a significant research direction in the field of optical remote sensing data processing. However, due to the limitation of the optical imaging mechanism, traditional image segmentation methods are not efficient for processing the optical remote sensing images, especially influencing by the complex weather conditions. In order to ensure the classification performance, synthetic aperture radar (SAR) data are employed as complementary to the data procedure for enhancing the capability of land cover interpretation. Then a dual-stream high-resolution network (HRNet) is proposed to combine two types of heterogeneous data (SAR and optical image), and a multi-modal squeeze-and-excitation (SE) module is exploited to make feature maps fused. Experiments show that the proposed method has excellent performance on the remote sensing data acquired by GF2 and GF3 satellites.
Bo Ren 0001, Shibin Ma, Biao Hou, Danfeng Hong
IGARSS1
2021 A Mutual Information-Based Self-Supervised Learning Model for PolSAR Land Cover Classification
abstract
Recently, deep learning methods have attracted much attention in the field of polarimetric synthetic aperture radar (PolSAR) data interpretation and understanding. However, for supervised methods, it requires large-scale labeled data to achieve better performance, and getting enough labeled data is a time-consuming and laborious task. Aiming to obtain a good classification result with limited labeled data, we focus on learning discriminative high-level features between multiple representations, which we call mutual information. As PolSAR data have multi-modal representations, there should have strong similarity between multi-modal features of the same pixel. In addition, each pixel has its own unique geocoding and scattering information. Hence, every pixel has great difference from other pixels in a specific representation space. Based on the above observations, this article proposes a mutual information-based self-supervised learning (MI-SSL) model to learn an implicit representation from unlabeled data. In this article, the self-supervised learning idea is first applied to PolSAR data processing. Furthermore, a reasonable pretext task, which is suitable for PolSAR data, is designed to extract mutual information for classification tasks. Compared with the state-of-the-art classification methods, experimental results on four PolSAR data sets demonstrate that our MI-SSL model produces impressive overall accuracy with fewer labeled data.
Bo Ren 0001, Biao Hou, Jocelyn Chanussot, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.1
2020 PolSAR Scene Classification via Low-Rank Tensor-Based Multi-View Subspace Representation
abstract
In this paper, the polarimetric synthetic aperture radar (PolSAR) scene classification is solved by using a novel low -rank tensor-based multi-view subspace representation (LRT-MSR) method. PolSAR data can be described in multimodal feature spaces, such as PolSAR coherent/covariance/scattering matrices, or the various polarimetric decompositions. Different pseudo-color images from multiple spaces provide enough visual information for making a comprehensive representation. Our method applies a low-rank tensor-based subspace clustering way to explore the information from multi-view pseudo-color images. Tensor, as the high order matrix, is used to capture the correlations of underlying multi-view data. Furthermore, the method is constrained by a low-rank term that elegantly models the cross information from different views, and achieves a series of representation matrices from the redundant information. Finally, a spectral cluster method is used to make the final classification. The experimental results on PolSAR image dataset show the effectiveness of the applied method.
Mengqian Chen, Bo Ren 0001, Biao Hou, Jocelyn Chanussot, Shuang Wang 0001, Xiangrong Zhang
IGARSS2
2020 Panchromatic Image Land Cover Classification Via DCNN with Updating Iteration Strategy
abstract
Land cover classification is a critical research task in many significant remote sensing applications. There are emerging many powerful pixel-level classification methods based on deep convolutional neural network (DCNN) in the universal computer vision community. However, due to the complication of satellite image senses and the lack of high-quality labeled datasets, these computer vision techniques can not be applied to remote sensing applications directly. In this paper, we propose a novel DCNN method to extract abstract feature from the complicated remote scene. The proposed method fuses three level features from the encoder while the segmentation result is obtained by decoder. Furthermore, we propose an updating iteration strategy (UIS) with label smoothing on training set to reduce the impact of the incorrect labels. The proposed strategy employs the output of the network to update the low-confidence labels on training set, and utilizes the updated labels to continue training the network. In order to acquire a better segmentation result on a very high resolution (VHR) panchromatic image, we transfer the features trained on GID dataset to our dataset for training. Our experiments has demonstrated the oustanding performance of the proposed method in land cover classification compared to DeepLabv3 on the GID and our dataset.
Biao Hou, Yangfei Liu, Tuotuo Rong, Bo Ren 0001, Zijuan Xiang, Xiangrong Zhang, Shuang Wang 0001
IGARSS4
2020 Modified Tensor Distance-Based Multiview Spectral Embedding for PolSAR Land Cover Classification
abstract
This letter proposes a novel method for combining multiview features in polarimetric synthetic aperture radar (PolSAR) for land cover classification. It is well-known that feature extraction and classifier design are two significant steps in machine learning methods for PolSAR data interpretation. Each PolSAR pixel can be represented in different feature spaces, such as polarimetric data scattering, or the polarimetric target decomposition spaces. In this letter, a tensor-based multiview embedding algorithm is proposed to fuse those features from different spaces in order to obtain a distinctive set of features for the subsequent classification. Based on the pixel-based classification tasks, a modified tensor distance (MTD) is designed to accurately calculate the distance between tensors. It emphasizes the importance of the central pixel, and decreases the influence of the neighbors in the feature patch when calculating tensor distance. Furthermore, the complementary properties of different views are exploited by an MTD measured tensor multiview spectral embedding method, so as to obtain relevant low-dimensional features. Compared with state-of-the-art methods, the validation and effectiveness of the proposed method is demonstrated on two real PolSAR data sets.
Bo Ren 0001, Biao Hou, Jocelyn Chanussot, Licheng Jiao
IEEE Geosci. Remote. Sens. Lett.1
2020 PolSAR Feature Extraction Via Tensor Embedding Framework for Land Cover Classification
abstract
Polarimetric synthetic aperture radar (PolSAR) as a typical multi-channel sensor can obtain refined geometrical and geophysical information. In the PolSAR land cover classification task, feature extraction is regarded as a critical step for the final classification. It can employ multi-modal features from the original scattering data, polarimetric target decomposition, and other transformation space. Then, how to efficiently combine multi-modal polarimetric information and extract discriminant features is an important challenge for PolSAR image processing. Graph embedding methods have become a significant technique to deal with feature extraction and dimensionality reduction (DR) problems in recent years. It provides a unified linearization framework in machine learning and other pattern recognition tasks. In this article, an extended tensor embedding framework is introduced to extract the intrinsic features for PolSAR land cover classification. First, each pixel is represented by a feature cube that is constructed by groups of polarimetric scattering signals and target decomposition features in a fixed size patch. Second, an intrinsic matrix is constructed to describe the original geometrical and statistical properties of the samples, and a penalty matrix is designed to represent some constraints. Third, the vector-based algorithms are transformed into tensor space in an unified framework and based on the pair of matrices to obtain the projection matrices in each mode by an iterative optimization process. The effectiveness of the proposed methods is demonstrated on three RADARSAT2 data sets covering the regions of Xi'an, San Francisco, and Flevoland, respectively. The visualization and quantification results show that the proposed method has superiority in land cover classification.
Bo Ren 0001, Biao Hou, Jocelyn Chanussot, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.1
2019 Polsar Land Cover Classification via Tensorial Embedding Methods
abstract
In recent years, graph embedding has become a significant technique to deal with feature extraction and dimension reduction problems. Under the linearization and kernelization, it provides a unified framework in machine learning and other pattern recognition tasks. Polarimetric synthetic aperture (PolSAR) as a typical multi-channel sensor can obtain more geometrical and geophysical information. How to combine those polarimetric scattering signals and target decomposition features and explore the spatial information between pixels become a new research direction to address PolSAR data. In this paper, we utilize the tensorial embedding methods to extract the intrinsic features from a redundant feature space for the PolSAR land cover classification. The effectiveness of the proposed methods is demonstrated using AIRSAR Flevoland data set.
Bo Ren 0001, Biao Hou, Jocelyn Chanussot, Changzhe Jiao, Xiangrong Zhang
IGARSS1
2019 A Novel Semicoupled Projective Dictionary Pair Learning Method for PolSAR Image Classification
abstract
Polarimetric synthetic aperture radar (PolSAR) image classification plays an important role in remote sensing image processing. In recent years, stacked auto-encoder (SAE) has obtained a series of excellent results in PolSAR image classification. The recently proposed projective dictionary pair learning (DPL) model takes both accuracy and time consumption into consideration, and another recently proposed semicoupled dictionary learning (SCDL) model gives a new way to fit different features. Based on the SAE, DPL, and SCDL models, we propose a novel semicoupled projective DPL method with SAE (SAE-SDPL) for PolSAR image classification. Our method can get the classification result efficiently and correctly and meanwhile giving a new method to fit different features. In this paper, three PolSAR images are used to test the performance of SAE-SDPL. Compared with some state-of-the-art methods, our method obtains excellent results in PolSAR image classification.
Yanqiao Chen, Licheng Jiao, Yangyang Li 0001, Lingling Li 0002, Bo Ren 0001, Naresh Marturi
IEEE Trans. Geosci. Remote. Sens.6
2019 CNN-Based Polarimetric Decomposition Feature Selection for PolSAR Image Classification
abstract
In order to better interpret polarimetric synthetic aperture radar (PolSAR) images, many scholars tend to do target decomposition for PolSAR images and utilize the obtained features to perform subsequent classification. These target decomposition features play an important role in terrain classification but completely utilizing them produces a high computational complexity. Furthermore, some features have a negative impact on the classification task. Therefore, selecting the appropriate amount of high-quality features is of great significance to the classification task. In this paper, we propose a convolutional neural network (CNN)-based feature selection algorithm for PolSAR image classification. First, we design a 1-D CNN for feature selection, then train the designed network with all the decomposition features to obtain a trained model. Second, the Kullback-Leibler distance (KLD) between different features is utilized as a standard to select feature subsets. Third, feature subsets with excellent performance form the final results. Due to the special structure of the 1-D CNN, repetitively training model is avoided when the input changes. Different from traditional feature selection methods, our method considers the performance of features combination rather than single feature contribution. To this end, the feature subsets selected by the proposed method are more useful to the classification task. Innovatively introducing KLD in the selection stage avoids random selection and improves the selection efficiency. Finally, we validate the performance of selected feature subsets in traditional and deep learning classification frameworks. Experiments demonstrate that features selected by the proposed method have a good performance comparing with others on three real PolSAR data sets.
Chen Yang 0027, Biao Hou, Bo Ren 0001, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.3
2018 Decomposition-Feature-Iterative-Clustering-Based Superpixel Segmentation for PolSAR Image Classification
abstract
Compared with traditional pixel-based polarimetric synthetic aperture radar (PolSAR) image classification methods, superpixel-based methods take advantages of the spatial information of pixels, so they can overcome the influence of speckle noise on the classification result. Since traditional superpixel methods do not utilize the scattering characteristics of a PolSAR image, the boundaries of the superpixels are poorly preserved. The inaccuracy of superpixel segmentation boundaries has a negative impact on the subsequent classification. In this letter, we propose a decomposition-feature-iterative-clustering (DFIC) superpixel segmentation method for PolSAR images. The DFIC method innovatively introduces the decomposition features in generating superpixels, so the superpixel segmentation boundaries are well preserved. Because we selectively utilize superpixel information to classify the PolSAR images by setting a threshold, the effect of superpixel segmentation inaccuracy on the classification results is reduced. Experiments on two real PolSAR images demonstrate that the proposed method outperforms several state-of-the-art superpixel methods, and that the DFIC superpixel-based classification obtains better results than the other pixel-based methods.
Biao Hou, Chen Yang 0027, Bo Ren 0001, Licheng Jiao
IEEE Geosci. Remote. Sens. Lett.3
2016 Unsupervised PolSAR image classification using boundary-preserving region division and region-based affinity propagation clustering
abstract
This paper presents a new method for polarimetric synthetic aperture radar (PolSAR) image classification. Firstly, to get a reasonable edge strength map, polarimetric information is used in edge strength calculation, and watershed algorithm is used to obtain the oversegmentation using the edge strength. Secondly, a searching table is used to determine the most suitable region to be merged. Finally, region-based affinity propagation clustering is employed to achieve an initial classification map, and the method provides an adjacent Wishart classifier with spatial relations to obtain the final classification result.
Biao Hou, Yuheng Jiang, Bo Ren 0001, Zaidao Wen, Shuang Wang 0001, Licheng Jiao
IGARSS3
2016 SAR Image Classification via Hierarchical Sparse Representation and Multisize Patch Features
abstract
In this letter, a novel hierarchical sparse representation-based classification (HSRC) for synthetic aperture radar (SAR) images is proposed. Features utilized in HSRC are extracted from the multisize patches around each pixel to precisely describe the complex terrains. Two thresholds are introduced in the sparse representation classifier to restrict the range of reconstruction residual, which classifies the reliable classified points, and the rest of the pixels are considered as the uncertain ones in the original SAR image. Then, a new dictionary is constructed by the reliable pixels, and the uncertain pixels will be reclassified in the next classification layer. The hierarchical structure is very reasonable and effective to employ simple features in each layer for describing the various topographic types. Compared with traditional sparse representation-based classification and support vector machines in several fixed-size patches, the proposed method can obtain better performance both in quantitative evaluation and visualization results.
Biao Hou, Bo Ren 0001, Guilin Ju, Huiyan Li, Licheng Jiao, Jin Zhao 0002
IEEE Geosci. Remote. Sens. Lett.2