VLDB 2026 Research / reviewers in the wild / expert
Biao Hou
dblp:93/5255
· DBLP profile ↗
218ranked-venue papers
27as first author
133since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 151 · 21 first-author · 84 since 2021Artificial intelligence and machine learning · 40 · 1 first-author · 30 since 2021Graphics, computer vision, multimedia, augmented reality and games · 22 · 14 since 2021Systems, architecture and hardware · 3 · 2 first-author · 2 since 2021Computer networks · 3 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Hierarchical Direction Perception via Atomic Dot-Product Operators for Rotation-Invariant Point Clouds LearningabstractPoint cloud processing has become a cornerstone technology in many 3D vision tasks. However, arbitrary rotations introduce variations in point cloud orientations, posing a long-standing challenge for effective representation learning. The core of this issue is the disruption of the point cloud's intrinsic directional characteristics caused by rotational perturbations. Recent methods attempt to implicitly model rotational equivariance and invariance, preserving directional information and propagating it into deep semantic spaces. Yet, they often fall short of fully exploiting the multiscale directional nature of point clouds to enhance feature representations. To address this, we propose the Direction-Perceptive Vector Network (DiPVNet). At its core is an atomic dot-product operator that simultaneously encodes directional selectivity and rotation invariance--endowing the network with both rotational symmetry modeling and adaptive directional perception. At the local level, we introduce a Learnable Local Dot-Product (L2DP) Operator, which enables interactions between a center point and its neighbors to adaptively capture the non-uniform local structures of point clouds. At the global level, we leverage generalized harmonic analysis to prove that the dot-product between point clouds and spherical sampling vectors is equivalent to a direction-aware spherical Fourier transform (DASFT). This leads to the construction of a global directional response spectrum for modeling holistic directional structures. We rigorously prove the rotation invariance of both operators. Extensive experiments on challenging scenarios involving noise and large-angle rotations demonstrate that DiPVNet achieves state-of-the-art performance on point cloud classification and segmentation tasks. Chenyu Hu, Hao Zhu 0009, Biao Hou |
AAAI | 4 |
| 2026 | Optimization Method for Surrogate Function in Spiking Neural Networks Based on Membrane Potential DistributionabstractSpiking Neural Networks (SNNs) offer promising energy efficiency and temporal sparsity for edge intelligence, but their training remains difficult due to gradient mismatch, membrane potential drift, and discretization errors. In this paper, we propose a membrane potential-guided surrogate optimization(MPO) framework that dynamically aligns the surrogate function with the membrane potential distribution to enhance the gradient propagation. Specifically, we introduce a KL-divergence-based regularization to stabilize membrane potential dynamics, and an adaptive width constraint to synchronize the surrogate gradient range with neural activity statistics. Additionally, we design a spike discretization error metric and a correction strategy to mitigate temporal discretization effects. Experiments on CIFAR-10, CIFAR-100, and ImageNet show our method achieves 94.76%, 74.20%, and 65.70% top-1 accuracy respectively, while improving gradient stability and energy efficiency. This work provides a principled optimization scheme for robust and scalable SNN training in practical neuromorphic systems. Kaige Geng, Biao Hou |
AAAI | 5 |
| 2026 | CSFE-Net: Cycle-consistency scattering feature extraction network for PolSAR image
Biqi Li, Chen Yang 0027, Biao Hou, Bo Ren 0001, Licheng Jiao |
Neurocomputing | 4 |
| 2026 | Recurrent progressive fusion-based learning for multi-source remote sensing image classification
Hao Zhu 0009, Biao Hou, Wenhao Zhao, Xiaoyu Yi 0002, Wenping Ma 0001, Licheng Jiao |
Pattern Recognit. | 4 |
| 2026 | KCI-Net: Knowledge-Based Contourlet Inference Network for Super-ResolutionabstractTextural details are useful for image super-resolution, but massive CNN methods ignored the high-frequency components and generated over-smoothed outputs. The knowledge-based contourlet inference network is proposed in this paper. Different from other CNN-based methods that are directly infer high-resolution (HR) images, our model learns to reconstruct the HR image through the series of corresponding contourlet coefficients. Specifically, first, we consider the low-pass subbands of the contourlet as the corresponding low-resolution (LR) image. Then, feed it to the embedding net with residual blocks to provide adequate information for the contourlet coefficients prediction. Finally, we innovatively convert the estimation of contourlet coefficients into the estimation of the generalized gaussian distribution (GGD) parameters, and design the corresponding loss function to ensure training stability, which explores the smoothness of the contour effectively and guarantees the general structure and details of images. Experiments on four remote sensing datasets, four natural scenes and human-made content datasets, and the outdoor dataset demonstrate the superiority of the proposed model quantitatively and qualitatively. Mengkun Liu, Licheng Jiao, Xu Liu 0006, Lingling Li 0002, Fang Liu 0001, Shuyuan Yang 0001, Shuang Wang 0001, Biao Hou |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2026 | A Progressive Semi-Distillation Model for Dual-Source Remote Sensing Image ClassificationabstractPanchromatic images (PANs) and multispectral (MS) images (MSs) are widely used for dual-source remote sensing image classification, gradually becoming a research hotspot. However, making the most of dual-source image information with insufficiently labeled samples is a significant challenge. This article proposes a progressive semi-distillation model (PSDM) to classify dual-source remote sensing images with insufficient samples. We design a framework of rookie teacher network (RTN)-teaching assistant system (TAS)-student grouping network (SGN) in the case of a traditional teacher network (TN) (i.e., rookie TN (RTN)) that does not provide excellent guidance to student network (SN) due to insufficient samples. The PSDM expands the samples and compresses the space through the RTN-SGN structure to cope with the dilemma of insufficient samples. To make RTN better guide the SGN, we design TAS, which can gradually guide SGN to learn the samples from easy to difficult. It can also further assist SGN training to improve the classification performance of SGN with insufficient samples. We design SGN and add cooperation and correction mechanism to better learn dual- source information. These strategies can eliminate SGN's over-dependence on the RTN, help SGN outperform the RTN, and achieve the effect of semi-distillation. Experimental results and theoretical analysis have sufficiently pointed out the proposed method's accuracy, efficiency, and robustness under insufficient sample situations. Our model is available at https://github.com/MarjordCpz/PSDM. Hao Zhu 0009, Peizhou Cao, Licheng Jiao, Biao Hou, Xiaoyu Yi 0002, Wenhao Zhao, Wenping Ma 0001 |
IEEE Trans. Cybern. | 5 |
| 2026 | DI3CL: Contrastive Learning With Dynamic Instances and Contour Consistency for SAR Land-Cover Classification Foundation ModelabstractAlthough significant advances have been achieved in SAR land-cover classification, recent methods remain predominantly focused on supervised learning, which relies heavily on extensive labeled datasets. This dependency not only limits scalability and generalization but also restricts adaptability to diverse application scenarios. In this paper, a general-purpose foundation model for SAR land-cover classification is developed, serving as a robust cornerstone to accelerate the development and deployment of various downstream models. Specifically, a Dynamic Instance and Contour Consistency Contrastive Learning (DI3CL) pre-training framework is presented, which incorporates a Dynamic Instance (DI) module and a Contour Consistency (CC) module. DI module enhances global contextual awareness by enforcing local consistency across different views of the same region. CC module leverages shallow feature maps to guide the model to focus on the geometric contours of SAR land-cover objects, thereby improving structural discrimination. Additionally, to enhance robustness and generalization during pre-training, a large-scale and diverse dataset named SARSense, comprising 460,532 SAR images, is constructed to enable the model to capture comprehensive and representative features. To evaluate the generalization capability of our foundation model, we conducted extensive experiments across a variety of SAR land-cover classification tasks, including SAR land-cover mapping, water detection, and road extraction. The results consistently demonstrate that the proposed DI3CL outperforms existing methods. Our code and pre-trained weights are publicly available at: https://github.com/SARpre-train/DI3CL. Zhongle Ren, Kai Wang 0053, Biao Hou, Xingyu Luo, Weibin Li 0002, Licheng Jiao |
IEEE Trans. Image Process. | 4 |
| 2026 | Scale-Aware Prompting With Optimal Transport for Remote Sensing Image CaptioningabstractRemote sensing image captioning is a multimodal foundation task for fine-grained understanding of remote sensing images. However, remote sensing images contain complex scenes and rich objects, it is very challenging to accurately describe the objects in the scene with their attributes and dependencies. To address these issues, the article proposes a novel scale-aware prompting with optimal transport (SPOT) to learn effective multiscale features under diverse scenes, and to build fine-grained cross-modal alignment between semantic features and linguistic words during caption generation. Specifically, a scale-aware prompt extractor is constructed to explore feature integrations in complex scenes through learning prompts that query multi-scale features, and to enhance the representation of attributes and dependencies for objects by embedding positional relations. Besides, a fine-grained cross-modal alignment is designed to dynamically match image feature representations and textual semantics through optimal transport. Through the above manner, the model can learn effective language-aligned feature representations for caption generation. Finally, a caption Transformer with causal self-attention is introduced to generate accurate captions for remote sensing scenes. Extensive experiments demonstrate that the proposed method achieves state-of-the-art performance on three public datasets, with the superiority of the proposed method further demonstrated by ablating the role of each component. Cheng Zhang 0028, Zhongle Ren, Biao Hou, Jiawei Ning, Kai Wang 0053, Weibin Li 0002, Licheng Jiao |
IEEE Trans. Image Process. | 3 |
| 2026 | Cetus: Online Context-Aware Cross-Layer Coordination for Efficient Live Volumetric Video StreamingabstractIn recent years, volumetric videos have gradually prospered as an intriguing video paradigm, offering users a fully immersive viewing experience with six Degrees of Freedom (DoF). However, most current live volumetric video streaming methods struggle to facilitate the real-time performance requirements due to the nature of frequent user interactions and the complexity of network environments during video playback. Inspired by the correlation between the human visual effects and adjacent frame motion features, we proposeCetus, a context-aware cross-layer coordination system for live volumetric videos. First, we present an application-layer Neural Radiance Fields (NeRF)-based codec framework that leverages spatio-temporal semantic information for optimizing the compression quality of each video frame. Second, we exploit a flexible cross-layer coordination framework that seamlessly integrates frame drop strategy with partially reliable transmission, orchestrating transport protocols and application-informed rates to enhance the Quality of Experience (QoE) for multiple users. Furthermore, we develop a lightweight branching decision tree algorithm that adaptively makes fine-grained frame drop decisions. Experimental evaluations of our implemented system prototype demonstrate that Cetus significantly outperforms existing baseline approaches. Compared to the state-of-the-art baselines, Cetus effectively improves video frame rate by at least 24.7% and video quality by an average of 32.6%. Biao Hou, Song Yang 0002, Youqi Li, Fan Li 0001, Liehuang Zhu, Xu Chen 0004, Ramin Yahyapour |
IEEE Trans. Netw. | 1 |
| 2025 | Partial Point Cloud Registration with Multi-view 2D Image LearningabstractLearning representations from numerous 2D image data has shown promising performance, yet very few works apply this representations to point cloud registration. In this paper, we explore how to leverage the 2D information to assist the point cloud registration, and propose IAPReg, an Image-Assisted Partial 3D point cloud Registration framework with the multi-view images generated by the input point cloud. It is expected to enrich 3D information with 2D knowledge, and leverage 2D knowledge to assist with point cloud registration. Specifically, we create multi-view depth maps by projecting the input point cloud from several specific views, and then extract 2D and 3D features using some well-established models. To fuse the information learned from 2D and 3D modalities, inter-modality multi-view learning module is proposed to enhance geometric information and complement semantic information. Weighted SVD is a common method to reduce the impact of inaccurate correspondences on registration. However, determining the correspondence weights is not trivial. Therefore, we design a 2D-weighted SVD method, where the 2D knowledge is employed to provide weight information of correspondences. Extensive experiments perform that our method outperform the state-of-the-art method without additional 2D training data. Yue Zhang 0040, Yue Wu 0004, Wenping Ma 0001, Maoguo Gong, Hao Li 0009, Biao Hou |
AAAI | 6 |
| 2025 | Ofl-Md: Exploring One-Shot Federated Learning for Melanoma Diagnosis With Pre-Trained Diffusion Model ClipabstractFederated learning enables a server to train a global model by leveraging data distributed across multiple clients. However, its inherently distributed and iterative nature introduces significant communication overhead as well as potential privacy concerns. To mitigate these issues, one-shot federated learning restricts communication between the server and clients to a single round. Nevertheless, this constraint often results in reduced accuracy. In the field of medical diagnosis, data security is particularly critical and has become a major research focus. In this paper, we propose a one-shot federated learning framework and conduct experiments on three popular datasets and a melanoma image dataset collected by ourselves. To further enhance performance, we introduce One-shot Federated Learning framework for Melanoma Diagnosis, OFL-MD, which incorporates a pre-trained diffusion model CLIP to assist clients in generating synthetic datasets and employs differential privacy to further enhance the privacy. The synthetic datasets are then transmitted to the global model for lightweight fine-tuning, thereby improving accuracy. We evaluate our approach on both widely used benchmark datasets and our own melanoma dataset. The results validate the effectiveness of our multimodal one-shot federated learning framework, which not only preserves data privacy and achieves high global model accuracy but also shows promise for future deployment in assisting dermatologists with melanoma detection of patients. Zegui Jiang, Liangxi Liu, Zelan Li, Yongyi Xie, Biao Hou, Jianning Chi |
BIBM | 6 |
| 2025 | Small dense Mini/Micro LED high-precision inspection based on instance segmentation with local detail enhancement
Jie Chu 0004, Jueping Cai, Biao Hou, Kailin Wen |
Adv. Eng. Informatics | 4 |
| 2025 | Semi-SNN: Biological-inspired semi-supervised image classification with spiking neural networks
Biao Hou, Chuanfeng Ma, Leida Li, Hao Zhu 0009, Licheng Jiao |
Neurocomputing | 3 |
| 2025 | Dual-model spiking neural network for remote sensing image classification using mutual knowledge distillation
Biao Hou, Hao Zhu 0009, Yifan Ge, Licheng Jiao |
Neurocomputing | 4 |
| 2025 | A two-stage strategy for brain-inspired unsupervised learning in spiking neural networks
Chuanfeng Ma, Biao Hou, Leida Li, Hao Zhu 0009, Dou Quan, Licheng Jiao |
Neurocomputing | 3 |
| 2025 | Learning opposite prompts for weakly supervised video anomaly detection
Helei Qiu, Biao Hou, Yanyu Cui |
Knowl. Based Syst. | 2 |
| 2025 | Joint Style and Layout Synthesizing: Toward Generalizable Remote Sensing Semantic SegmentationabstractThis paper studies the domain generalized remote sensing semantic segmentation (RSSS), aiming to generalize a model trained only on the source domain to unseen domains. Existing methods in computer vision treat style information as domain characteristics to achieve domain-agnostic learning. Nevertheless, their generalizability to RSSS remains constrained, due to the incomplete consideration of domain characteristics. We argue that remote sensing scenes have layout differences beyond just style. Considering this, we devise a joint style and layout synthesizing framework, enabling the model to jointly learn out-of-domain samples synthesized from these two perspectives. For style, we estimate the variant intensities of per-class representations affected by domain shift and randomly sample within this modeled scope to reasonably expand the boundaries of style-carrying feature statistics. For layout, we explore potential scenes with diverse layouts in the source domain and propose granularity-fixed and granularity-learnable masks to perturb layouts, forcing the model to learn characteristics of objects rather than variable positions. The mask is designed to learn more context-robust representations by discovering difficult-to-recognize perturbation directions. Subsequently, we impose gradient angle constraints between the samples synthesized using the two ways to correct conflicting optimization directions. Extensive experiments demonstrate the superior generalization ability of our method over existing methods. Qi Zang, Shuang Wang 0001, Dong Zhao 0007, Zhun Zhong, Biao Hou, Licheng Jiao |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Interpretable Fine-Grained Aircraft Classification Network for Remote Sensing Image With Image Pair Interaction and Neural TreeabstractThis paper proposes a novel interpretable framework for fine-grained aircraft classification in high-stakes remote sensing applications. Our approach addresses three key challenges: small inter-class variance, large intra-class variance, and the need for model interpretability. Specifically, our framework is built on the Swin Transformer (SwinT) backbone and includes three main modules. First, we present the Dynamic Attention Fusion Module (DAFM), which adaptively fuses multi-stage attention maps from the SwinT backbone. By leveraging a dispersion-based weighting mechanism, DAFM balances the contributions of coarse and fine-grained features, capturing both global structures and localized details. Second, we propose the Adaptive Image Pair Interaction Module (AIPI), which dynamically adjusts feature interaction strategies based on intra-class and inter-class similarity, effectively enhancing informative regions and improving robustness. To further optimize discriminative power, we incorporate an AIPI loss function that enforces intra-class consistency and inter-class separability. Finally, we develop a Binary Neural Tree Module (BNTM) to hierarchically select and propagate informative image patches, enhancing both feature refinement and interpretability through explicit path-based decision-making. Extensive experiments on benchmark datasets demonstrate that our framework significantly improves classification accuracy and interpretability, making it well-suited for applications requiring transparent and reliable decision-making. The Code can be found at https://github.com/StarmanGzx/BNTM. Zhengxi Guo, Biao Hou, Xianpeng Guo, Chen Yang 0027, Zitong Wu, Bo Ren 0001, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | Dense-Weak Ship Detection Based on Foreground-Guided Background Generation Network in SAR ImagesabstractCurrently, ship detection based on Synthetic Aperture Radar (SAR) images still faces significant challenges, particularly in detecting weak and densely distributed ships within complex backgrounds. In areas such as ports and land, the complex background features often resemble those of densely distributed ships, leading to reduced detection accuracy. Additionally, the overlapping and mutual interference of features among dense ships can cause the network to miss detections or produce false positives. Therefore, this paper proposes a Foreground-Guided Background Generation Network (FGBG-Net), which includes a Gaussian Foreground Localization (GFL) model and a Background Feature Removal (BFR) module. The GFL module identifies the approximate high-probability regions of ship foregrounds on the feature map, guiding the network to focus on these regions. The BFR module then progressively removes background interference features based on the positions provided by the GFL module, generating feature maps that are more suitable for detecting weak and dense ships. Our network has been validated on multiple SAR ship datasets, and the experimental results demonstrate noticeable performance improvements, with a mean Average Precision (mAP) increase of 3.4% on the SSDD and HRSID datasets. The relevant code is available at the following link: https://github.com/Xidian-AIGroup190726/FBGBNet/tree/master. Wenping Ma 0001, Xiaoting Yang, Hao Zhu 0009, Xiaoteng Wang, Biao Hou, Mengru Ma, Yue Wu 0004 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | FAFormer: Frequency-Analysis-Based Transformer Focusing on Correlation and Specificity for PansharpeningabstractPan-sharpening refers to fusing remote sensing multispectral (MS) and panchromatic (PAN) images to generate high-resolution multispectral (HR-MS) images. Recent advancements in deep learning-based pan-sharpening techniques have shown promising results. However, they face the following two issues. On one hand, there is a modality gap between MS and PAN images. Directly fusing them can lead to spectral and spatial distortions. On the other hand, the fusion process is prone to information loss, which can lead to image blurriness. To tackle these issues, we develop a Transformer-based model: FAFormer, which incorporates frequency analysis and focuses on the correlation and specificity of the PAN and MS images. Focusing on correlation can reduce the spectral and spatial distortions while focusing on specificity can reflect the specific information from MS and PAN images in the fusion result. We utilize the Discrete Wavelet Transform (DWT) to obtain the correlate and specific features. We introduce bijective functions based on the Transformer to design an Integrated Attention Block (IAB). As a critical component of the model, it effectively utilizes the correlation and specificity of the two images. In designing the model’s overall framework, we employ a Correlative Feature Attention Module (CFAM) to leverage the correlation between MS and PAN. We utilize a Specific Feature Attention Module (SFAM) to integrate specific information into fused features gradually. Experimental results show that our method improves pan-sharpening performance and has practical value. Codes are available at https://github.com/Xidian-AIGroup190726/FAFormer. Yifan Meng, Hao Zhu 0009, Xiaoyu Yi 0002, Biao Hou, Shuang Wang 0001, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | Hierarchical Prototype Learning With Uncertainty-Aware Adaptation for Cross-Domain Semantic Segmentation of Remote Sensing ImagesabstractProminent domain discrepancies in Remote Sensing Images (RSIs), such as sensor types, geographical patterns, and land usage, significantly hinder the research and practical applications of cross-scene land classification. Unsupervised Domain Adaptation (UDA) fully exploits domain invariance between labelled source and unlabelled target domain, which alleviates the challenge of inaccurate land classification due to lack of labels in RSIs. However, most existing UDA methods for RSIs semantic segmentation are insufficient in exploring cross-domain features and have difficulty in modelling fine-grained domain-invariant features between inter-class. In this paper, we propose a novel self-training UDA method named Hierarchical Prototype Learning (HPL), which learns the inherent domain-wise consistency and class-wise invariance through progressive exploration of the prototypy, significantly alleviating the ambiguity in uncertain regions. HPL mainly consists of Domain-wise Progressive Prototype Interaction (DPPI) and Class-wise Dynamic Prototype Collaboration (CDPC). DPPI and CDPC specialize in hierarchically building prototype interaction architectures tailored to domain-wise alignment and class-wise calibration, respectively. This design not only mitigates the sensitivity to cross-domain scenarios but also allows for precise correction of uncertain regions. Furthermore, CDPC exhibits the capacity for pixel-level category restoration and promotes the correct and fine-grained updating of pseudo-labels. Extensive Experiments on two public datasets and a private self-build dataset demonstrate the superiority of HPL over other state-of-the-art methods for UDA semantic segmentation of RSIs. Jiawei Ning, Zhongle Ren, Biao Hou, Runnong Jiang, Weibin Li 0002, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | Deep Geospatio-Semantic Guided Network With Pseudo-Label Consistency for Domain-Adaptive Remote Sensing SegmentationabstractDomain-adaptive Remote Sensing Images (RSIs) semantic segmentation mitigates the overfitting problem that affects the effectiveness of segmentation, which results from the scarcity of high-quality labels and the cross-domain styles of ground objects. The effectiveness of domain adaptive segmentation remains suboptimal in complex scenarios due to inadequate exploitation of latent geographic knowledge. Consequently, inter-class ambiguity and boundary agnostic are further exacerbated under cross-domain transfer scenarios. To address this issue, we first devise a deep geospatio-semantic guided network named DSSAL, which comprehensively investigates the potential spatial relationship and semantic correlation between classes of RSIs by geospatial aware interaction and geosemantic aware interaction, respectively. To mitigate class-wise cognitive deviation in the unlabeled domain, DSSAL-DA is developed to further enhance the segmentation effect with the spatio-semantic domain alignment module in manifold cross domain tasks. Furthermore, a pseudo-labels consistency filter is developed for DSSAL-DA to ensure reliability in self-training through cross-view consistency verification. Extensive experiments on two public datasets and a private dataset demonstrate the superiority of DSSAL and DSSAL-DA over the state-of-the-art methods for UDA semantic segmentation of RSIs. Jiawei Ning, Zhongle Ren, Biao Hou, Weibin Li 0002, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | Self-Supervised Learning of Contrast-Diffusion Models for Land Cover Classification in SAR ImagesabstractDeep learning methods has been widely applied to synthetic aperture radar (SAR) land cover classification. The complexity of SAR data and the limited availability of labeled samples greatly constrain the feature learning and the generalization ability of the model. Inspired by the excellent generative performance of Denoising Diffusion Probabilistic Models(DDPM) on complex data distributions, a self-supervised learning framework based on contrast-diffusion models (CDM) is proposed to expand the applicability to multiple broad scenarios with complex and varying imaging conditions under limited annotated data conditions. Specially, The proposed framework consists of the upstream CDM pre-training on all unlabeled samples and the downstream land cover classification with few labeled samples in each test scene. Concretely, in the upstream task, the features are captured through the generative learning of the DDPM. Following this, the Dimensionality Reduction and Resolution Expansion (DRRE) module is designed and embedded to reduce feature redundancy and align the feature granularity between layers and the input image. Finally, contrastive learning is employed to enforce semantic feature consistency across different steps. In the downstream task, the feature in pre-trained CDM is efficiently delivered in a single-step reverse diffusion process and then fine-tuned with few labeled samples from each test scene and finally output the predictions. Compared with several supervised and self-supervised methods, the proposed framework achieves superior classification and generalization performance on multiple broad scenes with complex and varying imaging conditions. For example, based on the average results from six test scenes, CDM shows improvements in overall accuracy (OA) of 41.42%, 34.40%, 7.70% ,8.53% and 9.60% compared to Deeplabv3+, CCNR, Segformer, MAE and DDPM respectively. The code is available at https://github.com/gosling123456/CDM.git. Zhongle Ren, Zhe Du, Biao Hou, Weibin Li 0002, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | DCIFNet: Cross-Modal Fusion With Correction and Interaction for Optical-SAR Land Cover ClassificationabstractLand cover classification (LCC) based on remote sensing image segmentation is a prominent task of remote sensing data interpretation. The commonly used optical data is susceptible to the weather, so it has the potential to utilize complementary features from the supplementary synthetic aperture radar (SAR) data to enhance segmentation performance. However, current multi-modal segmentation methods focus on the deep fusion of features, which usually ignores the significance of structural consistency information. In order to make use of the mutual correction and information exchange between multi-modal data, we propose DCIFNet, a dual-stream correction-interaction-fusion multi-modal LCC network. Specifically, we design a differential feature correction and enhancement module (DF-CEM) that leverages bidirectional differential features to correct multi-modal features. In addition, for corrected feature pairs, we deploy a parallel attention interaction module (PAIM) to focus on the pixel-level feature correlation and achieve effective information exchange in both channel and spatial dimensions. Through the expert fusion module (EFM), DCIFNet leverages the gate network to attain a flexible and compact feature fusion between multi-modal features. Experimental results show that our method achieves a superior performance compared with other multi-modal fusion segmentation methods on three optical-SAR datasets. The source code of DCIFNet is publicly available at https://gitee.com/asdwer2046/dcifnet. Bo Ren 0001, Bo Liu 0009, Qianfang Wang, Biao Hou, Chen Yang 0027, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | Incremental Land Cover Classification via Soft Label and Subregion DistillationabstractWith the exponential growth of satellite remote sensing data, land cover classification models must adapt continuously to new classes. However, conventional incremental learning methods face critical challenges: catastrophic forgetting degrades recognition of old classes, and the softmax function further suppresses old-class probabilities due to ”class crowding.” Existing distillation techniques also struggle to transfer features in irregular geospatial regions. To address these issues, we propose Soft Labels and Subregion Distillation (SLSRD). SLSRD mitigates class crowding by employing soft labels instead of hard labels, derived from a hybrid of softmax and sigmoid outputs that preserve richer probabilistic information. Concretely, the soft label is a convex combination of softmax- and sigmoid-based probabilities that preserves inter-class relations while relaxing over-confident exclusivity for newly introduced categories, and it supervises all pixels across stages. In parallel, a breadth-first search identifies subregions within each image, which are weighted by probability and size, and similarity between corresponding subregions of the old and new models is maximized. This dual strategy effectively transfers fine-grained knowledge and overcomes the limitations of conventional distillation methods, particularly for large-scale remote sensing imagery. Experiments on three benchmark datasets-Vaihingen, GID, and FBP-demonstrate that SLSRD outperforms traditional methods, significantly improving incremental land cover classification. Bo Ren 0001, Zhao Wang 0011, Hanyuan Ge, Biao Hou, Bo Liu 0009, Chen Yang 0027, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | Sample-Level Improved Cross-Source Contrastive Learning for PAN and MS Joint ClassificationabstractIn recent years, the number and ways of acquiring panchromatic images (PAN) and multispectral images (MS) have increased, and manual labeling costs have also increased. Processing these data efficiently has become a challenge. In this paper, we propose a sample-level improved cross-source contrastive learning method for PAN and MS joint classification (SLCL), which aims to provide a self-supervised pre-training model using unlabeled samples for downstream joint classification using a small quantity of labeled samples. First, we propose a sample weighting and screening (SWS) strategy, which enables the model to learn inter- and intra-source sample representations, while balancing the interference from false samples so that the model learns true samples. It solves the problems of homologous similar features embedded far away and false negative samples bringing the wrong learning direction, which exist in existing contrastive learning methods. In addition, we design a hard sample learning (HSL) module for the problem of mining and optimization of hard samples. The module efficiently mines hard samples and uses a new loss function to make the model more focused on hard sample optimization. It further improves the accuracy of pre-training models for downstream tasks. Our method performs best on multiple datasets, and it is experimentally validated and analyzed. The code is available at: https://github.com/Xidian-AIGroup190726/SLCL. Pengyu Tian, Hao Zhu 0009, Biao Hou, Pute Guo, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | EGPO: Enhanced Guidance and Pseudo-Label Optimization for Semi-Supervised Semantic Segmentation of Remote Sensing ImagesabstractOwing to complex remote sensing image features boosting manually labeled costs, semi-supervised semantic segmentation learns from limited labeled and abundant unlabeled data to alleviate this dilemma. However, there are still many challenges in the practical application of this technology, such as incorrect unsupervised information misleading the model and errors in pseudo-labels causing error accumulation. This paper proposes a semi-supervised semantic segmentation method for remote sensing images. The enhanced guided learning module we designed uses a label guided model to predict unlabeled data in the correct direction, enriching ground features and unsupervised information, and alleviating the problems of misprediction and consistency regularization failure caused by the lack of labels. At the same time, facing the noise generated during the data augmentation process, our designed unsupervised loss dynamic screening module aims to locate and suppress the noise adaptively. In addition, in the face of inevitable erroneous predictions in pseudo-labels, we design a pixel category selection module that produces high-quality and high-confidence pseudo-labels through multi-step filtering and dual model fusion. Ultimately, by conducting experiments on the DFC22, iSAID, MER, GID-15, and Vaihingen datasets, we successfully verified the effectiveness of our proposed method. The source code has been made public: https://github.com/Xidian-AIGroup190726/EGPO. Xifeng Xue, Hao Zhu 0009, Longsheng Qu, Biao Hou |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | Contour-Aware Dynamic Low-High Frequency Integration for Pan-SharpeningabstractPan-sharpening is the process of fusing panchromatic (PAN) and multispectral (MS) images. Its critical focus lies in accurately capturing the contour information from the PAN image during the fusion process and presenting it at a high resolution. However, existing deep learning methods lack the precise capture of delicate and smooth contour information, resulting in contour diffusion that affects the fusion results. Therefore, we introduce contourlet decomposition to capture multiscale directional delicate contour features and construct multiscale graph structures for semantic mining of dual-source contour features, continually updated through dynamic learning. By incorporating global features, we guide the multihead attention mechanism with directional decoding, enabling the network to pay more attention to high-resolution contour features, thereby gaining an advantage in image reconstruction. Cross-decoding between modalities provides strong representational capabilities for the advantageous features of both modalities, effectively enhancing the sharpening effect. Our algorithm achieves state-of-the-art results, and its effectiveness and advantages have been thoroughly validated across multiple datasets, including GaoFen-2, WorldView2, WorldView3, etc. Our code is available athttps://github.com/Xidian-AIGroup190726/CDFInet. Xiaoyu Yi 0002, Hao Zhu 0009, Pute Guo, Biao Hou, Bo Ren 0001, Xiaoteng Wang, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | Interactive Concept Network Enhanced Transformer for Remote Sensing Image CaptioningabstractRemote sensing image captioning plays an important role in advancing remote sensing image understanding with natural language generation. However, it is difficult to generate accurate semantic descriptions of crucial objects and their relationships, due to large coverage and abundant information in remote sensing images. To address these issues, this article proposes a novel interactive concept network enhanced transformer (ICNET) for remote sensing image captioning. First, multilevel visual features are extracted within a local and global feature extraction module. To comprehensively capture key objects in the local features, a concept mapping network (CMN) is constructed to project multiscale local features onto high-level semantic concepts of the objects. This allows for the integration of the relevant feature vectors in the visual feature mapping into multiple relatively independent word features, thus bridging the gap between visual features and semantic concepts. Subsequently, a global feature enhancement (GFE) module is introduced to boost the discrimination of global relationships and filter irrelevant content. Finally, to aggregate semantic concepts and global features, a transformer equipped with a concept interaction module (CIM) is designed to facilitate feature alignment and generate captions with proper categories and relationships. The experimental results on three remote sensing image captioning datasets demonstrate the superiority of the proposed method. Cheng Zhang 0028, Zhongle Ren, Biao Hou, Jianhua Meng, Weibin Li 0002, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | Adaptive Scale-Aware Semantic Memory Network for Remote Sensing Image CaptioningabstractRemote sensing image captioning between visual images and natural language remains a long-standing challenge in the remote sensing community. Due to the wide coverage and large amount of information in remote sensing images, existing methods struggle to effectively utilize the relevant semantic information about objects and their attributes at different scales across samples to generate descriptions. To address these issues, the article proposes a novel Adaptive Scale-aware Semantic Memory Network (ASSMN) for remote sensing image captioning. First, to fully extract the semantic information in remote sensing images, multilevel feature enhancement is constructed to improve the feature representation extracted from the CLIP pre-training model. Subsequently, a scale-aware attention aggregator is introduced to further integrate the enhanced multi-scale image features into the high-level semantics of remote sensing images. Then, to fully exploit the semantic information of the joint observed samples, a semantic memory reinforcement is designed to strengthen the semantic representation of the current scene through the relevant semantics obtained from other training samples. Finally, a captioning decoder is performed to generate a comprehensive scene caption with accurate objects and attributes. In the experiments, the performance of the proposed ASSMN on three remote sensing image captioning datasets is evaluated and compared with well-designed baselines and state-of-the-art methods, and the superiority of the proposed method is demonstrated by ablating the role of each proposed component. The code will be available at https://github.com/zcsisiyao/ASSMN. Cheng Zhang 0028, Zhongle Ren, Biao Hou, Changhui Xu, Jianhua Meng, Weibin Li 0002, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | Industrial-Mass Mini/Micro LED Sorting Using Patch Enhanced Lightweight Self-Attention Hybrid Neural NetworkabstractMini/Micro LEDs are becoming the next generation of displays, with sizes shrinking to less than 100μm/50μm and integration scaling increasing by more than 100-fold. Fast and precise sorting, which includes the identification and localization of mass mini/micro LED chips, has become an urgent need in the industry due to its high efficiency and reliability. However, the small size, high density, and weak defects of the mini/micro LED, combined with the industrial rapid sorting requirement, pose challenges to the sorting task. To address these issues, we propose a fast and precise visual sorting approach that includes accurate chip identification and pixel-level localization. In detail, a double-ended self-attention (SA) encoder–decoder hybrid convolutional encoder neural network framework is designed to accommodate both local features and global information. Then, two generic patch enhancement modules are constructed to compensate for the local feature modeling inefficiencies inherent in SA. Finally, the light SA encoder (Li-SA Co) and light SA covariance decoder (Li-SA-Cov Dec) basic blocks are proposed to speed up sorting. The effectiveness of the proposed approach is demonstrated by the actual mini/micro LED industrial images acquired, which have an mAP of 97.1%, a localization success of 98.6%, and a speed of 20.1FPS, all of which are higher than those of other state-of-the-art methods. Jie Chu 0004, Jueping Cai, Biao Hou, Kailin Wen |
IEEE Trans. Ind. Informatics | 3 |
| 2025 | Brain-Inspired Learning, Perception, and Cognition: A Comprehensive ReviewabstractThe progress of brain cognition and learning mechanisms has provided new inspiration for the next generation of artificial intelligence (AI) and provided the biological basis for the establishment of new models and methods. Brain science can effectively improve the intelligence of existing models and systems. Compared with other reviews, this article provides a comprehensive review of brain-inspired deep learning algorithms for learning, perception, and cognition from microscopic, mesoscopic, macroscopic, and super-macroscopic perspectives. First, this article introduces the brain cognition mechanism. Then, it summarizes the existing studies on brain-inspired learning and modeling from the perspectives of neural structure, cognitive module, learning mechanism, and behavioral characteristics. Next, this article introduces the potential learning directions of brain-inspired learning from four aspects: perception, cognition, understanding, and decision-making. Finally, the top-ten open problems that brain-inspired learning, perception, and cognition currently face are summarized, and the next generation of AI technology has been prospected. This work intends to provide a quick overview of the research on brain-inspired AI algorithms and to motivate future research by illuminating the latest developments in brain science. Licheng Jiao, Mengru Ma, Pei He, Xueli Geng, Xu Liu 0006, Fang Liu 0001, Wenping Ma 0001, Shuyuan Yang 0001, Biao Hou, Xu Tang 0004 |
IEEE Trans. Neural Networks Learn. Syst. | 9 |
| 2025 | Multiscale Deep Learning for Detection and Recognition: A Comprehensive SurveyabstractRecently, the multiscale problem in computer vision has gradually attracted people's attention. This article focuses on multiscale representation for object detection and recognition, comprehensively introduces the development of multiscale deep learning, and constructs an easy-to-understand, but powerful knowledge structure. First, we give the definition of scale, explain the multiscale mechanism of human vision, and then lead to the multiscale problem discussed in computer vision. Second, advanced multiscale representation methods are introduced, including pyramid representation, scale-space representation, and multiscale geometric representation. Third, the theory of multiscale deep learning is presented, which mainly discusses the multiscale modeling in convolutional neural networks (CNNs) and Vision Transformers (ViTs). Fourth, we compare the performance of multiple multiscale methods on different tasks, illustrating the effectiveness of different multiscale structural designs. Finally, based on the in-depth understanding of the existing methods, we point out several open issues and future directions for multiscale deep learning. Licheng Jiao, Xu Liu 0006, Lingling Li 0002, Fang Liu 0001, Zhixi Feng, Shuyuan Yang 0001, Biao Hou |
IEEE Trans. Neural Networks Learn. Syst. | 8 |
| 2025 | Pseudo Label Learning for Partial Point Cloud RegistrationabstractPartial point cloud registration plays a crucial role in computer vision and has widespread applications in 3D map construction, pose estimation, and high-precision localization. However, the collected point clouds often contain missing data due to hardware limitations and complex environments. Various partial registration algorithms have been proposed, most of which rely on estimating overlap regions. However, a significant proportion of these algorithms rely heavily on ground truth labels. Manual labeling is both time-consuming and labor-intensive, whereas algorithmic automatic labeling lacks sufficient accuracy. To tackle this issue, we present PSEudo Label learning for unsupervised partial point cloud registration (PSEL). This method utilizes complementary tasks to learn reliable pseudo labels for overlap regions and correspondences without depending on ground truth labels. The key idea is to use the complementarity between overlap estimation and registration to generate two types of pseudo labels based on the nearest points in pairs of aligned point clouds. These pseudo labels are then employed to supervise the learning of overlap regions and correspondences, gradually enhancing their accuracy throughout the learning process and ultimately establishing an unsupervised learning framework. PSEL consists of an overlap estimation module and a correspondence filtering module. The pseudo labels generated after registration are used to supervise both modules. Notably, the correspondence filtering module has two pipelines. The similarity and difference of the corresponding point features are used to eliminate false correspondences during the training and inference stages, respectively, with only the latter being optimized with pseudo labels. To validate the effectiveness of our registration method, we conducted experiments using the synthetic dataset ModelNet40, the indoor dataset 3DMatch, and the outdoor dataset KITTI. Wenping Ma 0001, Yue Wu 0004, Yue Zhang 0040, Hao Zhu 0009, Biao Hou, Licheng Jiao |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2024 | Masked Angle-Aware Autoencoder for Remote Sensing Images
Zhihao Li 0005, Biao Hou, Siteng Ma, Zitong Wu, Xianpeng Guo, Bo Ren 0001, Licheng Jiao |
ECCV (8) | 2 |
| 2024 | Few-Shot Class Incremental Land Cover Classification with Masked Exemplar SetabstractLand cover categories and features change with time. It generates demand for developing incremental learning methods for land cover classification. Meanwhile, due to the expensive cost of sample annotation, it is very difficult to obtain a large amount of annotated data for training. Therefore, how to effectively develop a few-shot incremental semantic segmentation method for land cover classification has become a significant task for remote sensing data interpretation. In this paper, we propose a novel data replay method with masked exemplar set (RMES) to improve land cover classification performance under the condition of few samples. It maintains a masked sample queue for each class. In this method, at the end of each learning step, two operations need to be run, the threshold sample filtering operation and the sample masking storage operation. These two operations update the sample queue and make it part of the training set in the next incremental learning stage. This alleviates overfitting and catastrophic forgetting problems. As a result of the experiment, the proposed RMES had superior performance in the CCF dataset. Junxi Guo, Bo Ren 0001, Zhao Wang 0011, Biao Hou |
IGARSS | 4 |
| 2024 | MSGFusion: Muti-scale Semantic Guided LiDAR-Camera Fusion for 3D Object Detectionabstract3D object detection is a key technology in automatic driving perception, which can provide the basis for safe and reliable autonomous driving. Aiming at the problem of false positive of low resolution object in point clouds, we present Multi-scale Semantic Guided LiDAR-Camera Fusion for 3D Object Detection(MSGFusion), which deeply fuses the features of image and LiDAR points. Specifically, we design multi-scale DenseFusion, which serially aggregate images features, point-wise features and voxel-wise feature volumes at different scales. At the same time, we design a new Image-based Predicted Keypoint Weighting(I-PKW). It predicts the object points based on the predicted foreground score map. Given the 3D proposals generated by the voxel CNN, we propose RoI-Pillar pooling. It abstracts the feature by aggregating the keypoints in the RoI by pillars. Compared with RoI-grid pooling, pillar-based feature encoding is more consistent with the distribution of fused feature keypoints to accurately regress the classification confidence and bounding box. Extensive experiments on the KITTI dataset show the superiority of MSGFusion. Huming Zhu, Yiyu Xue, Xinyue Cheng, Biao Hou |
IJCNN | 4 |
| 2024 | Accurate and Lightweight Learning for Specific Domain Image-Text RetrievalabstractRecent advances in vision-language pre-trained models like CLIP have greatly enhanced general domain image-text retrieval performance. This success has led scholars to develop methods for applying CLIP to Specific Domain Image-Text Retrieval (SDITR) tasks such as Remote Sensing Image-Text Retrieval (RSITR) and Text-Image Person Re-identification (TIReID). However, these methods for SDITR often neglect two critical aspects: the enhancement of modal-level distribution consistency within the retrieval space and the reduction of CLIP's computational cost during inference. To address these issues, this paper presents a novel framework, Accurate and lightweight learning for specific domain Image-text Retrieval (AIR), based on the CLIP. AIR incorporates a Modal-Level distribution Consistency Enhancement regularization (MLCE) loss and a Self-Pruning Distillation Strategy (SPDS) to improve retrieval precision and computational efficiency. The MLCE loss harmonizes the sample distance distributions within image and text modalities, fostering a retrieval space closer to the ideal state. SPDS employs a strategic knowledge distillation process to transfer deep multimodal insights from CLIP to a shallower level, maintaining only the essential layers for inference, thus achieving model light-weighting. Comprehensive experiments across various datasets in RSITR and TIReID reveal that MLCE loss secures optimal retrieval, while SPDS achieves a favorable balance between accuracy and computational demand during testing. Rui Yang 0038, Shuang Wang 0001, Jianwei Tao, Yingping Han, Qiaoling Lin, Yanhe Guo, Biao Hou, Licheng Jiao |
ACM Multimedia | 7 |
| 2024 | Deep convolutional encoder-decoder networks based on ensemble learning for semantic segmentation of high-resolution aerial imagery
Huming Zhu, Chendi Liu, Qiuming Li, Libing Wang, Licheng Jiao, Biao Hou |
CCF Trans. High Perform. Comput. | 8 |
| 2024 | Weakly supervised object localization via knowledge distillation based on foreground-background contrast
Siteng Ma, Biao Hou, Zhihao Li 0005, Zitong Wu, Xianpeng Guo, Chen Yang 0027, Licheng Jiao |
Neurocomputing | 2 |
| 2024 | A survey for table recognition based on deep learning
Weibin Li 0002, Wei Li 0318, Ruochen Liu 0006, Biao Hou, Licheng Jiao |
Neurocomputing | 6 |
| 2024 | MCDet: Multi-Content Collaboration Detector for Multiscale Remote Sensing ObjectabstractIn previous works, powerful CNN backbones are typically used for one- or two-stage detectors to facilitate multi-categories object classification. Unfortunately, continuous convolution and pooling operations tend to weaken the detailed information. We propose an end-to-end Multi-content Collaboration Detector (MCDet) to improve object recognition accuracy. First, we summarize the reasons for the disappearance of detailed features in traditional feature extraction backbone networks, and propose a Shallow Clue Refinement (SCR) module, which helps us to retain more critical local detail information in the downsampling process. Second, to receive more suitable contextual information, we design a Self-dilating Spatial Pooling (SSP) module, it adaptively learns a contextual reception field, thereby alleviating the mismatch between the theoretical receptive field of the network design and the practical requirements. Finally, extensive experiments on the NWPU VHR-10 and DIOR datasets have shown that the proposed MCDet significantly improves detection accuracy. Our code is available at https://github.com/Xidian-AIGroup190726/RS-objectdetection-MCDet. Wenping Ma 0001, Hao Zhu 0009, Yue Wu 0004, Biao Hou, Licheng Jiao |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2024 | SwinTFNet: Dual-Stream Transformer With Cross Attention Fusion for Land Cover ClassificationabstractLand cover classification (LCC) is an important application in remote sensing data interpretation. As two common data sources, SAR images can be regarded as an effective complement to optical images, which will reduce the influence caused by single-modal data. But common LCC methods are focusing on designing advanced network architectures to process single-modal remote sensing data. Few works have been oriented toward improving segmentation performance through fusing multi-modal data. In order to deeply integrate SAR and optical features, we propose SwinTFNet, a dual-stream deep fusion network. Through the global context modeling capability of Transformer structure, SwinTFNet models teleconnections between pixels in other regions and pixels in cloud regions for better prediction in cloud regions. In addition, a Cross-Attention Fusion Module (CAFM) is proposed to fuse features from optical and SAR data. Experimental results show that our method improves greatly in the classification of clouded images compared with other excellent segmentation methods and achieves the best performance on multi-modal data. Bo Ren 0001, Bo Liu 0009, Biao Hou, Zhao Wang 0011, Chen Yang 0027, Licheng Jiao |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2024 | Unreliable Pixel Contrast Based on von Mises-Fisher Distribution for Semi-Supervised SAR SegmentationabstractThe unique visual properties and the huge size of synthetic aperture radar (SAR) images pose challenges in labeling data. The scarcity of labeled data limits the training of SAR image segmentation networks. To address this issue, semi-supervised methods are used to train the network. However, traditional semi-supervised algorithms like Mean Teacher often generate erroneous pseudo-labels that carry over to subsequent training epochs, leading to overfitting and affecting network performance. This overfitting stems from an overreliance on unreliable pixels in Mean Teacher. In this study, an enhanced approach is proposed called unreliable pixel contrast (UPCo), where a von Mises-Fisher distribution is applied to constrain unreliable pixels in feature space. We augment the segmentation network with a feature output header for pixel-level contrastive learning in UPCo. Moreover, to minimize the computational effort during the training phase, hard sample selection and negative object non-uniform selection strategies are designed to facilitate contrastive learning. The proposed UPCo was evaluated on two large-scene SAR images and demonstrated its superiority over other comparative algorithms, achieving more optimal performance in semi-supervised segmentation of SAR images. Zitong Wu, Biao Hou, Xianpeng Guo, Bo Ren 0001 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2024 | Model-Based Decomposition Feature Learning With Adversarial PriorabstractModel-based target decomposition method has been widely applied due to its clear physical scattering significance. However, after establishing decomposition basis, the process of solving the scattering components and parameters is usually underdetermined, which will lead to the issues such as component negative power and overestimation. For this problem, this letter examines the target decomposition task from the perspective of deep learning and proposes an adversarial decomposition feature learning (ADFL) model. This model could learn decomposition features suitable for current terrain characteristics according to input data. At the same time, the model-based adversarial feature prior is embedded in ADFL to maintain the physical scattering meanings. On real PolSAR datasets, the learned features of proposed model are well correlated with real terrain scattering characteristics. Further, it avoids negative decomposition features and make more accurate fitting of scattering components, effectively alleviating the above problems. Chen Yang 0027, Biao Hou, Bo Ren 0001, Jocelyn Chanussot, Licheng Jiao |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2024 | LargeRSDet: A Large Mini-Batch Object Detector for Remote Sensing ImagesabstractDeep neural network models based on vision transformer (ViT) have shown unprecedented performance in the field of remote sensing image object detection. However, those models often require massive training data, which cost a lot of time to train and greatly prevent the research progress. Distributed training is a common way to accelerate the training period. In this letter, we propose a large batch object detector named LargeRSDet for remote sensing image object detection task, which can train with a batch size up to 1024 with only a little acceptable performance loss. Using the LargeRSDet, we can effectively utilize at most 1024 GPUs and greatly improve the training speed, which enables several benefits that not only help our model converge in a faster way but also provide the ability to reach a higher accuracy. Experimental results demonstrate that our method can finish training DIOR remote sensing image dataset in less than 5 min, and finally, the model achieves 75% mAP at 0.5. Huming Zhu, Qiuming Li, Kongmiao Miao, Biao Hou, Licheng Jiao |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2024 | Multi-grained clip focus for skeleton-based action recognition
Helei Qiu, Biao Hou |
Pattern Recognit. | 2 |
| 2024 | MSRIP-Net: Addressing Interpretability and Accuracy Challenges in Aircraft Fine-Grained Recognition of Remote Sensing ImagesabstractThe task of fine-grained aircraft recognition is crucial in the field of remote sensing. Despite some progress achieved by traditional deep learning methods in addressing this challenge, they are often perceived as a “black box,” lacking transparent explanations for model decisions. Current interpretable methods based on attention mechanisms, although providing some interpretability, do not align with human thought logic. Therefore, we propose a multiscale rotation-invariant prototype network (MSRIP-Net). Our approach simulates the intuitive reasoning process of humans in identifying objects by segmenting them into multiple components. Importantly, MSRIP-Net has the capability to automatically recognize rigid components on aircraft targets without relying on additional part annotations, using only image-level class labels. In addition, our approach effectively addresses challenges presented by noise, deformations, and multiscale variations in remote sensing targets and has been comprehensively evaluated on datasets FAIR1M1.0 and Rareplane. Our results demonstrate that MSRIP-Net achieves higher accuracy compared with existing fine-grained recognition methods. Furthermore, we provide insights into the model’s decision-making process to illustrate the interpretability of our approach. Zhengxi Guo, Biao Hou, Xianpeng Guo, Zitong Wu, Chen Yang 0027, Bo Ren 0001, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | MGC: MLP-Guided CNN Pretraining Using a Small-Scale Dataset for Remote Sensing ImagesabstractTo overcome the inherent domain gap between natural images and remote sensing images (RSIs), it is highly desirable to develop pretraining methods specifically for RSIs. Considering the lack of widely recognized large-scale benchmarks like ImageNet in the RSI community and limited computational resources, this article proposes multilayer perceptron (MLP)-guided convolutional neural network (CNN) (MGC), a method that employs an MLP to guide the pretraining of a CNN from small-scale datasets for RSIs. MGC has two encoders, each consisting of a CNN branch and an MLP branch. We first contrast pairwise samples from the same type of branches or different types of branches across the encoders and employ a positive-pair guidance strategy to explore consistency. Due to the inherent locality issue of shallow layers in a CNN, the CNN branches often do not attend to correct foreground regions such as objects, regions of interest, and land coverage. Therefore, we further propose an attention guidance strategy to guide the CNN branches to focus on foreground regions and learn discriminative representations effectively. The proposed MGC method is validated by pretraining a CNN model using the MGC and applying it to different downstream tasks including scene classification, rotated object detection, semantic segmentation, and change detection on ten datasets. Results have confirmed the effectiveness of the proposed MGC. Our code will be released at:https://github.com/benesakitam/MGC. Zhihao Li 0005, Biao Hou, Wanqing Li 0001, Zitong Wu, Bo Ren 0001, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Rebalancing Gaussian Location Loss for High-Precision Detection on Remote Sensing ImagesabstractAerial image objects are usually orientated arbitrarily, with a large scale range, and densely distributed. Traditional horizontal bounding box (HBB) detectors tend to filter out densely distributed objects leading to missed detections, such as ship (SH) and vehicle. Therefore, oriented object detection has become a mainstream solution in recent years. The 2-D Gaussian distribution representation of the oriented bounding boxes (OBBs) solves the problem of angular discontinuity and boundary discontinuity and thus gets more attention. However, as the aspect ratio of the object gradually decreases, its predicted angular performance continues to decrease. We find that the angular gradient of an object decreases sharply as the aspect ratio decreases, resulting in a large gradient gap between a small aspect ratio object (SARO) and a large aspect ratio object (LARO). It makes the detector prefer to ignore SARO during training, which weakens the high precision performance of SARO. We call this phenomenon shape imbalance. To solve the problem, we proposed a simple gradient rebalancing strategy named shape balance. Since the shape imbalance is only related to the aspect ratio of the object, we designed a modulation function with an inverse aspect ratio to calculate the balance coefficient. The principle of the function is that the larger the aspect ratio, the smaller the balance coefficient; the smaller the aspect ratio, the larger the balance coefficient. We aim to get the balance coefficients for objects with different aspect ratios. Location loss multiplied by a balance coefficient can directly adjust the gradient gap between objects with different aspect ratios to achieve a rebalancing effect. Extensive experiments conducted on DOTA-v1.0 dataset and DIOR-R dataset verify the effectiveness of our proposed method. Our method improves the detection performance of Gaussian location loss by an average of 2.08%/1.01%(AP75/mAP) metrics on the DOTA-v1.0 dataset and 1.17%/0.82%(AP75/mAP) improvements for DIOR-R dataset. Biao Hou, Zitong Wu, Xianpeng Guo, Bo Ren 0001, Zhongle Ren, Chen Yang 0027, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | MutSimNet: Mutually Reinforcing Similarity Learning for RS Image Change DetectionabstractChange detection involves analysis of discrepancies between two phases. However, when the unchanged elements are known, the changed features to be identified become straightforward. In addition, remote sensing image is constrained by limited spectral information, which leads to blurred boundaries between different semantics. Based on these two prior knowledge, in this artical, we introduce a novel change detection framework, named the mutually reinforcing similarity network (MutSimNet). This architecture aims to minimize false alarms along changing boundaries and reduce misjudgment rates among outliers. First, similarity learning is applied to change detection. The relationship between the two phases is considered when deriving the change feature maps. Second, we devise a mutually reinforcing loss function that integrates initial features with final features. Third, a self-attention module is connected in the feature pyramid network. This design mitigates information loss during the down-sampling process. Fourth, an attention feature fusion strategy is proposed for the integration of multi-layer features. This strategy takes into account the interaction between layer-by-layer features. Fifth, experimental results validate MutSimNet’s efficiency, particularly its ability to focus on edge contour learning. The MutSimNet also achieves superior performance on two benchmark datasets and predicts positive samples with higher probability. The codebase is accessible at https://github.com/ly-yu/MutSimNet. Xu Liu 0006, Yu Liu 0005, Licheng Jiao, Lingling Li 0002, Fang Liu 0001, Shuyuan Yang 0001, Biao Hou |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2024 | Intra- and Intersource Interactive Representation Learning Network for Remote Sensing Images ClassificationabstractRecently, remote sensing technology has developed faster and faster, and obtaining high-quality panchromatic (PAN) and multispectral (MS) images has become more accessible. The complementarity between them provides new opportunities in multisource remote sensing image classification. However, solving the problem of the semantic gap between multisource high-level features and, at the same time, utilizing the complementary properties between them to reduce intersource information redundancy is still a challenge. This article constructs an$I^{3}$RL-Net for the multisource remote sensing image classification task. Specifically, we design a cross-source interactive enhanced fusion module (CIEF-Module). For multilevel multisource features, by strengthening the dependencies of intrasource features and conducting intersource enhanced fusion, intrasource correlation features are refined, and the problem of the intersource semantic gap can be effectively alleviated. During the cross-source interaction process, we design a complementary representation supervised learning strategy (CRSL-Strategy). According to the similarities and differences of multisource features, it can adaptively promote complementary feature learning, thus generating a nonredundant multisource representation. The method has been verified to be effective on multiple RS datasets. The code is open source at:https://github.com/Xidian-AIGroup190726/Ping-Pie-I3RL-Net.git. Wenping Ma 0001, Yanshan Guo, Hao Zhu 0009, Xiaoyu Yi 0002, Wenhao Zhao, Yue Wu 0004, Biao Hou, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2024 | Adaptive Feature Separation Network for Remote Sensing Object DetectionabstractWith the development of remote sensing technology, remote sensing object detection has been widely applied in various fields, but it still faces some thorny challenges, such as the following: 1) the complexity of object scale changes in remote sensing images makes it difficult to improve the performance of small object detection and 2) remote sensing images have complex backgrounds and densely arranged small and weak objects, which pose a serious problem of feature interference. To alleviate these challenges, we propose an end-to-end adaptive feature separation network called AFSNet, which includes a scale-aware module (SAM) and a class-aware module (CAM). The SAM mainly enables feature maps of different resolutions to detect objects of different scales. Shallow feature maps mainly suppress the features of large objects they contain to focus on small object detection, while deep feature maps increase the detailed features of large objects they contain to focus on large object detection. The CAM is mainly used to distinguish the features in the feature map by category, separating the features of different categories into different channels, thus mitigating the problem of inter class feature interference, and blocking background interference. The effectiveness of this article has been proven on the NWPU VHR-10, IPIU-M, DIOR, and DOTA2.0 datasets. It can be widely applied in civilian, military, and other fields. Through experimental verification, our AFSNet achieved 97.70% mAP on the NWPU VHR-10 dataset, 78.9% mAP on the DIOR dataset, and 58.22% mAP on the DOTA2.0 dataset. Our code is available at:https://github.com/Xidian-AIGroup190726/AFSNet. Wenping Ma 0001, Yiting Wu, Hao Zhu 0009, Wenhao Zhao, Yue Wu 0004, Biao Hou, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2024 | ISSP-Net: An Interactive Spatial-Spectral Perception Network for Multimodal ClassificationabstractCoordinated and complementary spatial-spectral information is represented by the panchromatic (PAN) and multispectral (MS) images. The optimal utilization of the advantages of these images has become a subject of intense research interest. This article introduces the interactive spatial-spectral perception network (ISSP-Net) for multimodal remote sensing image classification, addressing the challenge of optimal utilization of complementary information from PAN and MS images. First, the pixel-guided spatial enhancement module (PGSE-Module) improves spatial location interaction using the spatial location enhancement learning strategy (SLEL-Strategy) and the cross-spatial aggregation learning strategy (CSAL-Strategy), integrating multiscale contextual information and emphasizing pixel-level features. Second, the time-frequency collaborative spectral enhancement module (TFCSE-Module) distinguishes useful frequency domain features through channel separation, lightweight convolutions, and adaptive Fourier transform learning. This approach enables comprehensive utilization of both primary and auxiliary information from multimodal data. Finally, experiments on four datasets demonstrate the ISSP-Net’s state-of-the-art performance in classifying MS and PAN images, with good generalization to hyperspectral (HS) and LiDAR data. The code is provided at:https://github.com/sun740936222/ISSP-Net. Wenping Ma 0001, Hekai Zhang, Mengru Ma, Biao Hou |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Self-Supervised Learning Guided by SAR Image Factors for Terrain ClassificationabstractEffective feature representation is the key to SAR image terrain classification. Limited by the abstract appearance and the scarcity of high-quality labeled data in this field, the features learned by current methods, especially deep learning models, do not have enough directivity and applicability, which hampers the performance. This paper proposes Multi-image Factor Self-Supervised Learning(MFSSL) to achieve directional feature learning and obtain generalized features with few patch-level labeled data. The framework consists of an upstream multi-factor image style transfer task and a downstream terrain classification task. In the upstream task, the goal of feature learning is first set up by multiple SAR image factors, including the observation region, the terrain category, and the imaging parameters. And then, different styles of SAR terrain images are generated and reconstructed under this goal. Through this bidirectional generative learning, the low-level external appearance of the terrain is removed, while the essential and discriminative feature representation is retained and shared across different factors. Finally, the downstream model inherits the general feature from the upstream model and implements the terrain classification task using a small amount of labeled data. Experiments conducted on three broad SAR scenes with different image factors demonstrate that the proposed framework can improve pixel-level terrain classification only with a few patch-level labeled data. Zhongle Ren, Zhe Du, Biao Hou, Weibin Li 0002, Hao Zhu 0009, Bo Ren 0001, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | MGPACNet: A Multiscale Geometric Prior Aware Cross-Modal Network for Images Fusion ClassificationabstractConvolutional neural networks (CNNs) and self-attention (SA) are highly effective techniques used for the fusion of multisource remote sensing (RS) data, and they have found extensive application in Earth observation (EO) tasks. Nevertheless, CNNs are insufficient for the comprehensive extraction of contextual information and the representation of the sequential properties of spectral features. Furthermore, the loss of edge geometry information is often a consequence of information mining, which limits its application in RS. To address the abovementioned limitations, we propose a method called “multiscale geometric prior aware cross-modal network (MGPACNet)” for RS image fusion classification. First, a geometric prior feature enhanced residual module (GPFEResM) is created to extract shallow multiscale geometric edge prior features and detailed information from multimodal RS data to enhance feature boundary information. Second, a multiscale global-local spatial-spectral feature extraction module (MG-LS2FEM) uses multiscale spatial modeling and global-local spectral modeling to perceive rich semantic information in the spatial-spectral domain. Finally, a dual attention fusion module (DAFM) is designed to use pixel-level SA and cross-attention between heterogeneous data to achieve deep aggregation and cross-focusing of cross-modal information in two branches, and enhance the complementarity of heterogeneous data. A comprehensive examination of public RS data (hyperspectral-synthetic aperture radar (HS-SAR) Augsuburg/Berlin, hyperspectral-light detection and ranging (HS-LiDAR) Trento/MUUFL) from four distinct modalities (HS/SAR/LiDAR) has revealed that our method outperforms alternative models. Licheng Jiao, Lingling Li 0002, Fang Liu 0001, Xu Liu 0006, Shuyuan Yang 0001, Biao Hou |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2024 | Lighter and Robust: A Rotation-Invariant Transformer for VHR Image Change DetectionabstractIn recent years, change detection (CD) has emerged as an increasingly intricate research domain. However, in natural images, the orientation of objects is often aligned with the image boundaries, whereas in RS images, the imaging angles are random. As a result, existing CD methods encounter limitations when effectively representing vector features. In this article, we propose a rotation-invariant CD architecture named RFormer. It effectively utilizes direction-sensitive position embedding (DSPE) to represent features in RS images. To address the challenge of the quadratic growth in attention mechanism complexity with sequence length, we introduce low-cost cross attention (LC2A) to reduce its complexity to$1/{C^{2}}$. Furthermore, we employ the implicit timing extraction process (TEP) to represent interframe bitemporal features. TEP plays a crucial role in mitigating prediction biases caused by seasonal changes in land cover and prevents overconfident discrimination by the classifier in CD tasks. Experimental results demonstrate that RFormer achieves competitive performance on WHU, deeply supervised image fusion network (DSIFN)-CD, CDD, and LEVIR-CD datasets. Licheng Jiao, Lingling Li 0002, Xu Liu 0006, Fang Liu 0001, Shuyuan Yang 0001, Biao Hou |
IEEE Trans. Geosci. Remote. Sens. | 8 |
| 2024 | Multi-View Feature Fusion and Visual Prompt for Remote Sensing Image CaptioningabstractRemote sensing image (RSI) captioning is a vision-language multimodal task concentrating on both image comprehension and sentence generation. Several studies suggest that encoder–decoder-based methods have achieved success in RSI captioning. However, existing encoder–decoder-based methods may not fully explore image representations for RSI captioning and suffer from a lack of additional prompt information for sentence generation. In this article, a novel multi-view feature fusion and prompt (MVP)-based model is proposed to obtain better RSI representations and enhance language model performance in RSI captioning. Specifically, we design an attention-based feature fusion module to dynamically fuse multi-view visual features, which are extracted from the fine-tuned vision-language pretraining (VLP) model and the vision-task pretraining (VP) model. Then, a flexible visual prefix mapping module is proposed to transform images into visual prefixes, providing semantic information for the subsequent sentence generation. Finally, a BERT-based caption generator is applied to generate accurate descriptions based on the fused visual features and the visual prefixes, which are both outputs from our designed modules. Extensive experiments are conducted on three well-known benchmark datasets, demonstrating that our method achieves state-of-the-art (SOTA) performance. The relevant code is available athttps://github.com/QiaoLing-Lin/MVP. Shuang Wang 0001, Qiaoling Lin, Xiutiao Ye, Yu Liao, Dou Quan, ZhongQian Jin, Biao Hou, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2024 | Renormalized Connection for Scale-Preferred Object Detection in Satellite ImageryabstractSatellite imagery, due to its long-range imaging, brings with it a variety of scale-preferred tasks, such as the detection of tiny/small objects, making the precise localization and detection of small objects of interest a challenging task. In this article, we design a knowledge discovery network (KDN) to implement the renormalization group theory in terms of efficient feature extraction (FE). Renormalized connection (RC) on the KDN enables “synergistic focusing” of multiscale features. Based on our observations of KDN, we abstract a class of RCs with different connection strengths, called$n21$C, and generalize it to feature pyramid network (FPN)-based multibranch detectors. In a series of FPN experiments on the scale-preferred tasks, we found that the “divide-and-conquer” idea of FPN severely hampers the detector’s learning in the right direction due to the large number of large-scale negative samples and interference from background noise. Moreover, these negative samples cannot be eliminated by the focal loss function. The RCs extends the multilevel feature’s “divide-and-conquer” mechanism of the FPN-based detectors to a wide range of scale-preferred tasks, and enables synergistic effects of multilevel features on the specific learning goal. In addition, interference activations in two aspects are greatly reduced and the detector learns in a more correct direction. Extensive experiments of 17 well-designed detection architectures embedded with$n21$Cs on five different levels of scale-preferred tasks validate the effectiveness and efficiency of the RCs. Especially the simplest linear form of RC—E421C performs well in all tasks, and it satisfies the scaling property of renormalization group theory. All experiments can be trained and tested on a graphics card with 8 GB of video memory, which greatly enhances the applicability of our methodology. We hope that our approach will transfer a large number of well-designed detectors from the computer vision community to the remote sensing community. Datasets and codes will be available at:https://github.com/rabbitme/ Fan Zhang 0041, Lingling Li 0002, Licheng Jiao, Xu Liu 0006, Fang Liu 0001, Shuyuan Yang 0001, Biao Hou |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2024 | Few-Shot MS and PAN Joint Classification With Improved Cross-Source Contrastive LearningabstractThe joint classification of multispectral (MS) and panchromatic (PAN) images aims to provide a more detailed and accurate interpretation of land features. Although deep-learning-based methods have achieved remarkable success in this task, the generalization performance of networks is compromised when labeled samples are insufficient. In this study, we explore the possibility of leveraging unlabeled remote sensing images (RSIs) through contrastive learning and demonstrate the challenges associated with directly applying contrastive learning to RSIs. To end this, we propose a cross-source contrastive learning method for few-shot MS and PAN joint classification (CrossCLMP), which aims to learn sufficient transferable representations in a self-supervised contrastive manner so as to provide a robust pretrained model for fine-tuning the downstream joint classification task. Specifically, we design: 1) intersource and intrasource alignment loss (ER-Align) to achieve self-supervised feature extraction and alignment; 2) the source-unique feature adaptive separation (SUAS) strategy to model source-unique information explicitly; and 3) the auxiliary contrastive learning (ACL) strategy to mitigate the adverse impact of numerous false-negative samples in the pretraining stage. The experimental results and the theoretical analyses on multiple popular datasets comprehensively demonstrate the effectiveness and robustness of the proposed method under few-shot. Our code is available at:https://github.com/Xidian-AIGroup190726/CrossCLMP. Hao Zhu 0009, Pute Guo, Biao Hou, Changzhe Jiao, Bo Ren 0001, Licheng Jiao, Shuang Wang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | High-Low-Frequency Progressive-Guided Diffusion Model for PAN and MS ClassificationabstractWith the rapid development of remote sensing technology, satellites can easily acquire multispectral (MS) and panchromatic (PAN) images. It is challenging to utilize their complementarity to effectively combine each other’s advantages and mitigate the differences between different modes. In this article, we propose a high-low-frequency progressive-guided diffusion model. It is used to generate an image with the advantages of both MS and PAN, which can be complementary to MS and PAN and, thus, can better reduce the modal differences between them. Therefore, we use the fusion image as an auxiliary mode and an intermediate bridge, which can better connect the characteristics between various sources. First, we design guidance information that contains the advantages of MS and PAN, and some operations can make this information better guide the generation stage. In addition, we design a high-low-frequency progressive guidance strategy; by using this strategy, we can first ensure the overall structure and layout of the image in the generation stage and then refine the local details and features of the image. This dramatically improves the quality of the generated image. Finally, we use mathematical knowledge to explain the rationality of the strategy. We validate our method on multiple datasets and achieve the best performance. Our code ishttps://github.com/Xidian-AIGroup190726/HLF-GDiffusion. Hao Zhu 0009, Fengtao Yan, Pute Guo, Biao Hou, Shuang Wang 0001, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | ConvGRU-Based Multiscale Frequency Fusion Network for PAN-MS Joint ClassificationabstractAs a hot research topic in remote sensing, effectively integrating the advantageous features of multispectral and panchromatic images is the main challenge for fusing these two remote sensing images. This article proposes a multiscale frequency fusion network based on ConvGRU. To address the underutilization of texture features, we extract multiscale bandpass and low-pass sub-bands representing texture and content features through Contourlet decomposition. Multiscale bandpass sub-bands contain more comprehensive and concentrated texture details. Then, by proposing a multiscale frequency feature extractor based on ConvGRU, we effectively integrate and enhance sub-bands of different scales and frequencies, fully utilizing the characteristics of multispectral and panchromatic images and scale transmission. With these enhanced sub-band features, we obtain more comprehensive scale-enhanced texture features. Simultaneously, content features are also preserved as dual-source image features. Moreover, to reduce redundancy between fused features and make more efficient use of the obtained enhanced features, we designed an Inver-band integrator (IBI) module. It can fuse enhanced features at different scales, improve the complementarity between features, and thus achieve effective fusion. Experimental results demonstrate the effectiveness and robustness of our model on multiple datasets. Our codes are available athttps://github.com/Xidian-AIGroup190726/GMFnet. Hao Zhu 0009, Xiaoyu Yi 0002, Biao Hou, Changzhe Jiao, Wenping Ma 0001, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | A Semantically Nonredundant Continuous-Scale Feature Network for Panchromatic and Multispectral ClassificationabstractIn recent years, panchromatic (PAN) images and multispectral (MS) images, as a type of multimodal remote sensing data, are attracting increasingly more attention to their classification problems. However, effectively representing size variations of targets in remote sensing images and reducing redundant representations of different modalities’ deep semantic features to enhance classification accuracy remains a challenge. In this article, we propose a semantically nonredundant continuous-scale feature network (SNCF-Net) for PAN and MS classification, consisting of two modules: the texture-enhanced continuous scale input generation module and the cross-modal feature Kernel interaction (CMKI) module. By simulating the human eye’s adjustment of distance to observe objects of different sizes, we employ 3-D convolution to extract continuous-scale images generated by the texture-enhanced continuous-scale input generation (TCIG) module, enabling optimal feature representation of objects in remote sensing images. Additionally, the texture enhancement (TE) strategy in the TCIG module alleviates texture diffusion in scale space, enhancing the network’s ability to represent texture features. Subsequently, the CMKI module utilizes the response differences between different features to generate convolution kernels from deep feature maps, enabling feature interaction between the PAN modal and MS modal. This reduces redundant representations of essential image content information in deep features of two modalities, facilitating a better mapping between dual-modal features and categories. Our results achieve state-of-the-art performance on multiple datasets. The code is available athttps://github.com/Xidian-AIGroup190726/SNCFNet. Hao Zhu 0009, Wenhao Zhao, Biao Hou, Changzhe Jiao, Zhongle Ren, Wenping Ma 0001, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Hierarchical Dynamic Graph Clustering NetworkabstractConnections between visual components are ubiquitous. Graphs, as a highly flexible data structure, not only allow imposing relational induction bias on data, but can provide a completely distinct learning perspective for regular image data. In this paper, we propose a hierarchical dynamic graph clustering network (HDGCN) for visual feature learning. We construct hierarchical graph representations in graph domain in an adaptive, data-adaptive and task-adaptive manner. First, the initial graph is constructed in high-dimensional feature domain of images. To mine the hierarchical geometric features in latent graph space, adaptive clustering network (ClusterNet) is performed to learn discriminative clusters and generates cluster-based coarse graph. Then, graph convolutional networks (GCNs) are used to diffuse, transform and aggregate information among clusters. So, the intra-class and inter-class information is fully explored to increase the discriminativity of graph representations. Next, coarsened graph representations are mapped to grid based on its affinity with linear projection features. To further improve the task adaptation of clusters and hierarchical graph representations, ClusterNet and GCNs are fused in the same framework for end-to-end training and clusters is updated dynamically. We have conducted extensive experiments on classification and segmentation tasks. The experimental results fully validate the robustness of the proposed algorithm. Jie Chen 0098, Licheng Jiao, Xu Liu 0006, Lingling Li 0002, Fang Liu 0001, Puhua Chen, Shuyuan Yang 0001, Biao Hou |
IEEE Trans. Knowl. Data Eng. | 8 |
| 2024 | Multi-Scale Contourlet Knowledge Guide Learning SegmentationabstractFor accurate segmentation, effective feature extraction has always been a challenging problem, since the variability of appearance and the fuzziness of object boundaries. Convolutional neural networks have recently gained recognition in feature representation learning. However, it is only conducted in the spatial domain, and lacks effective representation of directionality, singularity and regularity in the spectral domain for anomaly detection of images. This is the key to feature learning representation of high-order singularity. To solve this problem, a multi-scale contourlet knowledge guide learning network is proposed in this paper. It is novel in this sense that, different from the CNNs in the spatial domain, the proposed method learns the multi-scale contourlet sparse representation to obtain more effective and sparse features in multi-scales and multi-directions. Furthermore, the contourlet knowledge guide learning can enhance the representation of spectral domain features. It is shown that the proposed network can learn the multi-level discriminative features and capture the more accurate object boundaries. The segmentation ability in theoretical analysis and experiments on five polyp segmentation datasets (CVC-ColonDB, CVC-ClinicDB, Kvasir-SEG, ETIS-LaribPolypDB, EndoSceneStill) and two building datasets (Massachusetts, WHU) are compared with developed methods. It must be emphasized that there is potential in effective feature learning representation and the generalization capability of the proposed method in deep learning, recognition and interpretation. Mengkun Liu, Licheng Jiao, Xu Liu 0006, Lingling Li 0002, Fang Liu 0001, Shuyuan Yang 0001, Shuang Wang 0001, Biao Hou |
IEEE Trans. Multim. | 8 |
| 2024 | An Adaptive Migration Collaborative Network for Multimodal Image ClassificationabstractThe multispectral (MS) and the panchromatic (PAN) images belong to different modalities with specific advantageous properties. Therefore, there is a large representation gap between them. Moreover, the features extracted independently by the two branches belong to different feature spaces, which is not conducive to the subsequent collaborative classification. At the same time, different layers also have different representation capabilities for objects with large size differences. In order to dynamically and adaptively transfer the dominant attributes, reduce the gap between them, find the best shared layer representation, and fuse the features of different representation capabilities, this article proposes an adaptive migration collaborative network (AMC-Net) for multimodal remote-sensing (RS) images classification. First, for the input of the network, we combine principal component analysis (PCA) and nonsubsampled contourlet transformation (NSCT) to migrate the advantageous attributes of the PAN and the MS images to each other. This not only improves the quality of images themselves, but also increases the similarity between the two images, thereby reducing the representational gap between them and the pressure on the subsequent classification network. Second, for the interaction on the feature migrate branch, we design a feature progressive migration fusion unit (FPMF-Unit) based on the adaptive cross-stitch unit of correlation coefficient analysis (CCA), which can make the network automatically learn the features that need to be shared and migrated, aiming to find the best shared-layer representation for multifeature learning. And we design an adaptive layer fusion mechanism module (ALFM-Module), which can adaptively fuse features of different layers, aiming to clearly model the dependencies among multiple layers for different sized objects. Finally, for the output of the network, we add the calculation of the correlation coefficient to the loss function, which can make the network converge to the global optimum as much as possible. The experimental results indicate that AMC-Net can achieve competitive performance. And the code for the network framework is available at: https://github.com/ru-willow/A-AFM-ResNet. Wenping Ma 0001, Mengru Ma, Licheng Jiao, Fang Liu 0001, Hao Zhu 0009, Xu Liu 0006, Shuyuan Yang 0001, Biao Hou |
IEEE Trans. Neural Networks Learn. Syst. | 8 |
| 2024 | Explore the Influence of Shallow Information on Point Cloud RegistrationabstractFeature extraction is a key step for deep-learning-based point cloud registration. In the correspondence-free point cloud registration task, the previous work commonly aggregates deep information for global feature extraction and numerous shallow information which is positive to point cloud registration will be ignored with the deepening of the neural network. Shallow information tends to represent the structural information of the point cloud, while deep information tends to represent the semantic information of the point cloud. In addition, fusing information of different dimensions is conducive to making full use of shallow information. Inspired by this, we verify shallow information in the middle layers can bring a positive impact on the point cloud registration task. We design various architectures to combine shallow information and deep information to extract global features for point cloud registration. Experimental results on the ModelNet40 dataset illustrate that feature extractors that incorporate shallow information will bring positive performance. Wenping Ma 0001, Mingyu Yue, Yue Wu 0004, Yongzhe Yuan, Hao Zhu 0009, Biao Hou, Licheng Jiao |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2024 | A Concurrent Multiscale Detector for End-to-End Image MatchingabstractThis article focuses on end-to-end image matching through joint key-point detection and descriptor extraction. To find repeatable and high discrimination key points, we improve the deep matching network from the perspectives of network structure and network optimization. First, we propose a concurrent multiscale detector (CS-det) network, which consists of several parallel convolutional networks to extract multiscale features and multilevel discriminative information for key-point detection. Moreover, we introduce an attention module to fuse the response maps of various features adaptively. Importantly, we propose two novel rank consistent losses (RC-losses) for network optimization, significantly improving image matching performances. On the one hand, we propose a score rank consistent loss (RC-S-loss) to ensure that the key points have high repeatability. Different from the score difference loss merely focusing on the absolute score of an individual key point, our proposed RC-S-loss pays more attention to the relative score of key points in the image. On the other hand, we propose a score-discrimination RC-loss to ensure that the key point has high discrimination, which can reduce the confusion from other key points in subsequent matching and then further enhance the accuracy of image matching. Extensive experimental results demonstrate that the proposed CS-det improves the mean matching result of deep detector by 1.4%-2.1%, and the proposed RC-losses can boost the matching performances by 2.7%-3.4% than score difference loss. Our source codes are available at https://github.com/iquandou/CS-Net. Dou Quan, Shuang Wang 0001, Ning Huyan, Yi Li 0054, Ruiqi Lei, Jocelyn Chanussot, Biao Hou, Licheng Jiao |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2024 | Gamora: Learning-Based Buffer-Aware Preloading for Adaptive Short Video StreamingabstractNowadays, the emerging short video streaming applications have gained substantial attention. With the rapidly burgeoning demand for short video streaming services, maximizing their Quality of Experience (QoE) is an onerous challenge. Current video preloading algorithms cannot determine video preloading sequence decisions appropriately due to the impact of users’ swipes and bandwidth fluctuations. As a result, it is still ambiguous how to improve the overall QoE while mitigating bandwidth wastage to optimize short video streaming services. In this article, we devise Gamora, a buffer-aware short video streaming system to provide a high QoE of users. In Gamora, we first propose an unordered preloading algorithm that utilizes a Deep Reinforcement Learning (DRL) algorithm to make video preloading decisions. Then, we further devise an Asymmetric Imitation Learning (AIL) algorithm to guide the DRL-based preloading algorithm, which enables the agent to learn from expert demonstrations for fast convergence. Finally, we implement our proposed short video streaming system prototype and evaluate the performance of Gamora on various real-world network datasets. Our results demonstrate that Gamora significantly achieves QoE improvement by 28.7%–51.4% compared to state-of-the-art algorithms, while mitigating bandwidth wastage by 40.7%–83.2% without sacrificing video quality. Biao Hou, Song Yang 0002, Fan Li 0001, Liehuang Zhu, Lei Jiao 0002, Xu Chen 0004, Xiaoming Fu 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2024 | NOVA: Neural-Optimized Viewport Adaptive 360-Degree Video Streaming at the EdgeabstractThe 360-degree video streaming service provides a unique immersive viewing experience for users, who can freely change their Field-of-View (FoV) to view different portions of the videos. However, the demands for high throughput and low latency for 360-degree video pose substantial challenges to the current network infrastructure. Super Resolution (SR) is the procedure for reconstructing high-resolution images from low-resolution ones. Hence, caching video content on the network edge in advance, which is near end users, and applying the SR technique can significantly alleviate the transmission latency. In this article, we describeNOVA, an efficientNeural-OptimizedViewportAdaptive 360-degree video streaming system to improve the Quality of Experience (QoE) of users. In NOVA, we first design a foveated rendering SR approach to super-resolve video tiles utilizing computational resources at the edge. Subsequently, we present a meta-learning-based Multi-Agent Reinforcement Learning (MARL) algorithm to select SR depths and video tiles inside users’ viewports for agile video tile adaptation to optimize overall QoE under frequent network fluctuations. Finally, we implement the holistic prototype of NOVA and evaluate its performance on various real-world network datasets. Extensive experiments illustrate that compared to the state-of-the-art algorithms, NOVA improves average user-perceived QoE by up to 27%. Biao Hou, Song Yang 0002, Fan Li 0001, Liehuang Zhu, Xu Chen 0004, Yu Wang 0003, Xiaoming Fu 0001 |
IEEE Trans. Serv. Comput. | 1 |
| 2024 | SigDA: A Superimposed Domain Adaptation Framework for Automatic Modulation ClassificationabstractDue to the uncertainty of non-cooperative communication channels, the received signals often contain various impairment factors, leading to a significant decline in the performance of existing deep learning (DL)-based automatic modulation classification (AMC) models. Several preliminary works utilize domain adaptation (DA) to alleviate this issue, however, they are constrained by singular domain difference factor, whereas in practice, these factors often manifest cumulatively. Therefore, this paper introduce a more realistic task named superimposed DA, where multiple domain difference factors are overlaid, reflecting the cumulative nature of them. We propose the SigDA as a solution framework, which adopts adversarial training to align the data distribution in different domains. Two technical modules, Multi-task based Masked Signal Feature Extractor (M2SFE) and Signal Feature Pyramid Aggregation (SFPA), are innovatively designed in SigDA. M2SFE utilizes mask and reconstruction task to enhance feature extraction and achieves discriminative feature selection through the design of feature mapping layers, while SFPA can solve the problem of inconsistent signal length in superimposed DA and can aggregate the features of signals into the same dimension. We consider and superimpose various typical signal domain difference factors, comprehensive experiments demonstrate that the proposed framework can achieve significant performance improvement in various communication channels. Shuang Wang 0001, Hantong Xing, Chenxu Wang 0001, Huaji Zhou, Biao Hou, Licheng Jiao |
IEEE Trans. Wirel. Commun. | 5 |
| 2023 | Relational Image Patch Matching for Remote SensingabstractFeature descriptor-based methods have demonstrated remarkable performance in remote sensing image patch matching tasks and are usually optimized using contrastive loss and triplet loss. However, these optimization losses focus on calculating the distance between samples, ignoring the rich information of higher-order feature relationships between multiple image patches. The latter provides valuable information that can be used to improve task performance. Inspired by the superior performance of second-order relations in graph matching and clustering tasks, we aim to exploit the rich information available from high-order relations fully. This paper proposes a high-order relationship (HOR) learning method for remote sensing image patch matching. This method combines low-order feature relations between image patch pairs and high-order feature relations between multiple patches to enhance image matching performance. Extensive experimental results on a multimodel remote sensing image dataset, SEN 1-2, consisting of optical and SAR images, demonstrate that the proposed HOR learning method can improve the performance of remote sensing image patch matching. Xianwei Cao, Dou Quan, Chonghua Lv, Yanhe Guo, Shuang Wang 0001, Biao Hou, Licheng Jiao |
IGARSS | 6 |
| 2023 | A Fast and Accurate Method for Remote Sensing Image-Text Retrieval Based On Large Model Knowledge DistillationabstractWith the increasing development of remote sensing (RS) technology, remote sensing cross-modal image-text retrieval (RSCMITR) task has gradually attracted wide attention. At present, the large-scale pre-training model is brilliant in the field of natural images cross-modal retrieval, but the current RSCMITR models do not focus on it, resulting in less retrieval performance improvement. This paper proposes a lightweight network structure based on large-scale pre-training model and knowledge distillation, designing a lightweight model based on separable convolution and text convolution. Knowledge distillation technology is used to make the Light model learn the hidden knowledge of large-scale model CLIP-RS, which realizes fast and accurate retrieval. The proposed method achieves state-of-the-art performance on four commonly used RSCMITR datasets. Yu Liao, Rui Yang 0038, Hantong Xing, Dou Quan, Shuang Wang 0001, Biao Hou |
IGARSS | 7 |
| 2023 | Multi-Source Fusion Network for Remote Sensing Image Segmentation with Hierarchical TransformerabstractRecently, due to the limitations of single sensor, it is hard to improve the performance of land cover classification. The traditional image segmentation methods can not process the optical remote sensing images effectively, especially when optical sensor is affected by complex weather conditions. However, as an active radar, synthetic aperture radar(SAR) has the advantage of not being restricted by weather conditions with the the penetrability of electromagnetic radiation. So multi-sensor data fusion provides a great potential for land cover classification. In this paper, a new fusion network called SegFusion is proposed to improve the performance of land cover classification. There are two main components in SegFusion which are hierarchical Transformer encoder and Swin-Fusion(SW-Fusion) module. First, a hierarchical Transformer encoder is used to extract multilevel feature of optical and SAR images. By integrating features from different layers, we can obtain powerful representation that combines both low-resolution fine-grained features and high-resolution coarse-grained features. Second, SW-Fusion module is used to fuse the features of optical and SAR data. In SW-Fusion, we use modified Swin Transformer [1] block with multi-head cross-attention mechanism to exchange information between features from different sources. Bo Liu 0009, Bo Ren 0001, Biao Hou, Yu Gu 0015 |
IGARSS | 3 |
| 2023 | RotPointNet: Keypoint based Oriented Object Detection for Aerial ImageabstractRemote sensing object detection is characterized by arbitrary direction, dense targets and variable scale. Like common object detection methods predict horizontal rectangle boxes directly, most of the existing remote sensing oriented object detection methods predict oriented rectangle boxes directly. i.e. Center point position, length, width and rotation angle of the oriented rectangle boxes. These models often need to design complex rotation detection module to adapt to the rotating characteristics of targets, they lack good performance on targets with large aspect ratio changes as well. In this paper, we propose a keypoint based oriented object detection method for aerial images named RotPointNet to improve detection accuracy of rotating targets with large aspect ratio change. RotPointNet first locates the two endpoints of rotating targets using keypoint detection method, and then constructs the remote sensing oriented target based on the endpoints and additional width information. In addition, a new method of matching target key points is proposed in this paper. We conduct experiments to prove, the target detection model RotPoint-Net and key point matching strategy proposed in this paper can achieve good detection results on remote sensing images, especially for the detection of targets with large aspect ratio changes. Chongyu Wang, Bo Ren 0001, Biao Hou |
IGARSS | 3 |
| 2023 | Incremental Land Cover Classification via Strategies for Edge Removal and Feature Point AggregationabstractConvolutional neural networks will face the problem of catastrophic forgetting in the process of incremental learning. To solve this problem, we propose an incremental learning method called the strategy of edge removal and feature point aggregation, or ERFPA for short. In the cross-entropy loss, we perform edge detection and removal on the labels generated by the old model predictions, and then fuse them with the new labels. We calculate the mean point of different classes, and make the model learn features better by narrowing the distance with similar pixels. As demonstrated by the results of our experiment, on two remote sensing image datasets: CCF and Vaihingen, our method achieves state-of-the-art results. Zhao Wang 0011, Bo Ren 0001, Biao Hou, Yu Gu 0015 |
IGARSS | 3 |
| 2023 | EAVS: Edge-assisted Adaptive Video Streaming with Fine-grained Serverless PipelinesabstractRecent years have witnessed video streaming gradually evolve into one of the most popular Internet applications. With the rapidly growing personalized demand for real-time video streaming services, maximizing their Quality of Experience (QoE) is a long-standing challenge. The emergence of the serverless computing paradigm has potential to meet this challenge through its fine-grained management and highly parallel computing structures. However, it is still ambiguous how to implement and configure serverless components to optimize video streaming services. In this paper, we propose EAVS, an Edge-assisted Adaptive Video streaming system with Serverless pipelines, which facilitates fine-grained management for multiple concurrent video transmission pipelines. Then, we design a chunk-level optimization scheme to address video bitrate adaptation. We propose a Deep Reinforcement Learning (DRL) algorithm based on Proximal Policy Optimization (PPO) with a trinal-clip mechanism to make bitrate decisions efficiently for better QoE. Finally, we implement the serverless video streaming system prototype and evaluate the performance of EAVS on various real-world network traces. Our results show that EAVS significantly improves QoE and reduces the video stall rate, achieving over 9.1% QoE improvement and 60.2% latency reduction compared to state-of-the-art solutions. Biao Hou, Song Yang 0002, Fernando A. Kuipers, Lei Jiao 0002, Xiaoming Fu 0001 |
INFOCOM | 1 |
| 2023 | Knowledge Decomposition and Replay: A Novel Cross-modal Image-Text Retrieval Continual Learning MethodabstractTo enable machines to mimic human cognitive abilities and alleviate the catastrophic forgetting problem in cross-modal image-text retrieval (CMITR), this paper proposes a novel continual learning method, Knowledge Decomposition and Replay (KDR), which emulates the process of knowledge decomposition and replay exhibited by humans in complex and changing environments. KDR has two components: a feature Decomposition-based CMITR Model (DCM) and a cross-task Generic Knowledge Replay strategy (GKR). DCM decomposes text and image features into task-specific and generic knowledge features, mimicking the human cognitive process of knowledge decomposition. Specifically, it employs a generic knowledge features extraction module for all tasks and a task-specific module for each task with a few trainable fully connected layers. Similarly, GKR emulates the human behavior of knowledge replay by utilizing the image-text similarity matrix output from the old task model with inputting the previous samples to induce the learning of the image-text similarity matrix output from the current task model with inputting the previous samples, using knowledge distillation technology. To demonstrate the effect of KDR, we adapted a continual learning dataset Seq-COCO from MSCOCO. Extensive experiments on Seq-COCO showed that KDR reduces catastrophic forgetting and consolidates general knowledge, improving the model's learning ability in CMITR. Rui Yang 0038, Shuang Wang 0001, Yanhe Guo, Xiutiao Ye, Biao Hou, Licheng Jiao |
ACM Multimedia | 7 |
| 2023 | Spatio-temporal segments attention for skeleton-based action recognition
Helei Qiu, Biao Hou, Bo Ren 0001 |
Neurocomputing | 2 |
| 2023 | An Improved Neural Network Classification Algorithm by Expanding Training Samples for Polarimetric SAR ApplicationabstractThe idea of spatial correlation has been used in polarimetric synthetic aperture radar (PolSAR) classification for many years. It is common that the bigger the spatial correlation, the more information it contains. Though the recent advances in deep learning for PolSAR classification have achieved remarkable progress, the number of training samples is still a problem we have to face. Aiming to solve this problem, it is valuable to explore a greater spatial correlation of terrains as a priori for classification. Considering the correlated regions of the training sample as a prior, rather than a small neighborhood, the paper proposed Wishart locally constrained expansion algorithm. Based on this, a PolSAR classification algorithm is designed. The whole proposed PolSAR classification algorithm includes 3 parts: Wishart locally constrained expansion algorithm, convolution neural network training algorithm, and Markov random field post-processing algorithm. Supervised by the cluster knowledge from Wishart classifier, Wishart locally constrained expansion algorithm expands samples iteratively from the regions related to the training samples with a region-expanding technique, where very few labeled samples turn to a larger training set. Convolution neural network algorithm is designed to train convolution neural network with 3 types of training samples for a more accurate result. Convolution neural network is trained firstly with the numerous training samples generated through Wishart locally constrained expansion algorithm, and then the samples from consistency extraction and raw samples are involved to fine-tune the convolution neural network to get the improved result. Finally, Markov random field prior is used to smooth the result. Several benchmark datasets are adopted to evaluate the effectiveness of the proposed algorithm. The experimental results show that the new semi-supervised classification algorithm outperforms the state-of-the-art semi-supervised classification algorithms and classical supervised classification algorithms. Biao Hou, Bo Ren 0001, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | SSMU-Net: A Style Separation and Mode Unification Network for Multimodal Remote Sensing Image ClassificationabstractThe rapid progress in remote sensing technology has made it convenient for satellites to capture both multispectral (MS) and panchromatic (PAN) images. MS has more spectral information, and PAN has higher spatial resolution. How to exploit the complementarity between MS and PAN images, and effectively combine their respective advantageous features while alleviating mode differences, has become a crucial research task. This paper designs a Style Separation and Mode Unification network (SSMU-Net) for MS and PAN image classification from a novel and effective perspective. The network can be divided into two stages: style separation and mode unification. In the style separation stage, we use wavelet decomposition and techniques similar to generative adversarial networks to preliminarily separate the information of MS and PAN into different components. These components better preserve complete information from the original data and have their own advantages in style and content. Then we propose a Symmetrical Triplet Traction module to perform style traction on different components, making style features more unique and content features more unified, achieving feature separation and purification. In the mode unification stage, we design an encoder-decoder model to reduce the impact of mode differences. The experimental results from multiple datasets validate the effectiveness of our proposed method. Our overall accuracy improved by approximately 4% on the Shanghai and Beijing datasets, and it has exceeded 99.28% on the Hohhot and Vancouver datasets. Our code is available at: https://github.com/proudpie/SSMU-Net. Hao Zhu 0009, Licheng Jiao, Xiaoyu Yi 0002, Biao Hou, Wenping Ma 0001, Shuang Wang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2023 | Contrastive Learning Based on Multiscale Hard Features for Remote-Sensing Image Scene ClassificationabstractThe overwhelming majority of models for remote sensing image (RSI) scene classification generally require the weights pre-trained on natural images for initialization before formal training. However, differences in imaging mechanisms lead to huge discrepancies between natural images and RSIs, and the strong visual representation learned from massive natural images limits the performance of models when inferencing RSIs. To address this issue, the well-established self-supervised contrastive learning paradigm in the natural image field is introduced to the RSI field. We propose a contrastive learning method based on multi-scale hard features, MHCL, which aims to use finite RSIs to learn sufficient visual representations in an unsupervised contrastive manner, thus provide a powerful upstream pre-trained model for fine-tuning downstream scene classification task. Multi-level features extracted by intermediate layers of each encoder’s backbone are first gathered, and then a hard features transformation method is proposed to create hard positive features and diverse queues that save hard negatives, thereby enriching the finite scene information in small-scale RSIs. Furthermore, we redesign the multi-scale hard features joint contrastive loss to boost the model to explore sufficient invariant representations by additionally pulling hard positive pairs closer and pushing hard negative pairs farther away in the embedding space. Extensive experiments demonstrate that the upstream pre-training model generated by MHCL achieves competitive transferred performance on three popular scene classification datasets, outperforming the traditional model pre-trained on ImageNet and models pre-trained by other state-of-the-art contrastive learning methods. Our code will be released at: https://github.com/benesakitam/MHCL. Zhihao Li 0005, Biao Hou, Xianpeng Guo, Siteng Ma, Yanyu Cui, Shuang Wang 0001, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Complete Rotated Localization Loss Based on Super-Gaussian Distribution for Remote Sensing ImagesabstractLocalization regression in oriented object detection tasks has long faced boundary discontinuity and angular discontinuity problems induced by periodic angles. These problems were successfully resolved by using a 2d Gaussian distribution to modelling the oriented bounding box (OBB). However, the angular information of square-like objects will be lost when they are converted to 2d Gaussian distribution, forming a systematic problem. Its fundamental reason is that when the aspect ratio of the object tends to 1, the equiprobability curve of 2d Gaussian distribution degenerates from an ellipse to a circle, thus losing the orientation information of the rotated object. This results in the bounding boxes of such square-like objects not being learned effectively. To resolve this problem, we used the Lamé curve (or superellipse) to modify the existing 2d Gaussian function and designed a super-Gaussian distribution. This distribution can maintain anisotropy at arbitrary aspect ratios, thus preserving the angular information of the oriented object. We used the Kullback-Leibler (KL) divergence to measure the distance between two super-Gaussian distributions and convert it into a localization loss (SGKLD) by a function. SGKLD is an improved version of KLD loss. By modifying the form of the probability distribution, we elegantly fix the angle missing problem of the traditional Gaussian distribution. We validated the effectiveness of the proposed algorithm on several datasets and obtained the performance of SOTA. Our algorithm achieves a mean average precision (mAP) of 80.07, 76.59, 62.27, and 90.55/98.13 on the DOTA-v1.0, DOTA-v1.5, DOTA-v2.0, and HRSC2016 datasets, respectively. Biao Hou, Zitong Wu, Zhengxi Guo, Bo Ren 0001, Xianpeng Guo, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Gaussian Synthesis for High-Precision Location in Oriented Object DetectionabstractIn aerial image scenes, the objects have properties of arbitrary orientation, large-scale range, and dense distribution. Thus, the object detector uses oriented bounding box (OBB) to locate objects, which is more complex and challenging than horizontal bounding box (HBB) detector. Mainstream OBB detectors mostly use one-to-many label assignment strategy to predict multiple bounding boxes for the same object, and filter out repeat predictions by non-maximum suppression (NMS). NMS ranks with confidence and drops the detection box with IoU higher than the threshold, which is easy to get the local optimum result. The clustered synthesis method gets more accurate results than the original NMS, but applying it to the OBB detector leads to border shift, which arises from the angular discontinuity problem. Therefore, we use Gaussian OBB (G-OBB) to deal with the angular discontinuity and thus eliminate the offset generated by direct synthesis. G-OBB is not an easy to understand and describe representation. For this reason, we analyze the properties of G-OBB, and design a decoding method to convert a G-OBB to a rotated rectangular box, further discussing its conditions. Based on the decoding method, we propose a Gaussian synthesis algorithm (GauS), which transforms the OBB into Gaussian space, followed by synthesis, and finally transforms the synthesis result back into a new OBB. We have derived the synthesis and decoding methods, and further verified their effectiveness. The extensive experiments on several existing models show that GauS takes very little computation and improves detector’s high-precision performance. Extensive experiments verify the effectiveness, stability, and universality of the proposed algorithm. In addition, The RTMDet using GauS achieves a performance of 81.61 AP50and gains a 0.39% improvement in mAP, which achieves the SOTA performance. Our implementation is available at: https://github.com/lzh420202/GauS. Biao Hou, Zitong Wu, Bo Ren 0001, Zhongle Ren, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Unsupervised Prototype-Wise Contrastive Learning for Domain Adaptive Semantic Segmentation in Remote Sensing ImageabstractLabeling data in the field of remote sensing is time-consuming and labor-intensive, making domain adaptation between different domains an urgently needed solution. To address the domain gap between diverse datasets in the remote sensing domain, numerous methods tailored for domain adaptation in high-resolution remote sensing imagery have emerged. Some of the existing methods focus on reducing the domain gap at either the feature level or the pixel level, often overlooking their underlying connection. To tackle this issue, we introduce a prototype-wise contrastive feature alignment paradigm (PCFA) aimed at bridging the representations between the feature and pixel levels. By dynamically updating, we acquire prototype information encompassed by different mini-batches and employ an optimal transport mechanism to reasonably apply the prototype feature distribution in guiding the learning of target domain features. We conduct extensive domain adaptation semantic segmentation (DASS) experiments on the ISPRS Vaihingen and Potsdam datasets, achieving an improvement about 4%~5% in mIoU (mean Intersection over Union) compared to previous methods using the DeepLabV2 framework. Siteng Ma, Biao Hou, Xianpeng Guo, Zitong Wu, Zhihao Li 0005, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Automatic Aug-Aware Contrastive Proposal Encoding for Few-Shot Object Detection of Remote Sensing ImagesabstractIn the annotation of remote sensing images (RSIs), the effectiveness of common object detection methods trained on only a few samples decreases instantly, which has prompted increasing research on the few-shot problem in remote sensing. RSIs often exhibit suboptimal performance in few-shot scenarios due to the intricate nature of scene information interference and the high degree of cosine similarity, both of which present significant challenges to their effectiveness. In this paper, a two-stage detection framework based on fine-tuning is selected to deal with the common problems in the few-shot task of remote sensing domain. Considering the excessive scale variation of instances in remote sensing datasets, we introduce an automatically learned aug-aware search module to provide an intelligent data augmentation solution for Faster R-CNN using different optimal augmentation policies searched by the network to fit the current dataset. We introduce a contrastive RoI branch to better classify novel class proposal features that are easily confused by the base class. We named our work AACE and conducted extensive experiments on two common object detection datasets in remote sensing, NWPU VHR 10 and DIOR, on which AACE achieved about 2.30% and 2.61% improvement, respectively, in the number of shots listed in the paper, compared to other algorithms. Siteng Ma, Biao Hou, Zitong Wu, Zhihao Li 0005, Xianpeng Guo, Bo Ren 0001, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | A Novel Coarse-to-Fine Deep Learning Registration Framework for Multimodal Remote Sensing ImagesabstractMulti-modal remote sensing images with large rotation transformation (RT) are challenging to be registered. It needs to deal with the global geometric deformation caused by great RT and significant local appearance differences caused by different imaging mechanisms. Existing deep learning methods mainly use a single deep descriptor learning (DDL) network to extract invariant features for identifying matching samples and discriminative feature descriptors for separating non-matching samples. However, it is difficult to extract local invariant feature descriptors to RT and modality change through a single DDL network. This paper proposes a novel coarse-to-fine deep learning image registration framework for multi-modal remote sensing images based on two task-specific deep models. Specifically, in the coarse registration stage, this paper designs an effective deep ordinal regression (DOR) network for rotation correction, which can reduce the difficulty of multi-modal image registration and boost image registration. The proposed DOR network transforms the rotation correction task into a rotation ordinal regression problem, which can exploit the potential relationship between the rotation ordinals to improve the accuracy of rotation estimation. In the fine registration stage, we adopt the DDL network to deal with the image modality change based on the rotation-corrected images. Extensive experimental results on multi-modal image datasets demonstrate the significant advantages of the proposed coarse-to-fine deep learning registration framework. The DOR network achieves higher rotation correction accuracy, which can significantly improve the multi-modal image registration performances. Dou Quan, Huiyuan Wei, Shuang Wang 0001, Yu Gu 0015, Biao Hou, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | Incremental Land Cover Classification via Label Strategy and Adaptive WeightsabstractDuring incremental learning tasks, catastrophic forgetting occurs when old models are updated with new information. To address this issue, we propose a novel method called label strategy and adaptive weights (LSAW) that improves the incremental learning process. The label strategy introduces the old classes and solves the problem of how to reasonably use the wrong samples predicted by the old model. In the cross-entropy (CE) loss, we apply a threshold to filter the pseudolabels predicted by the old model. Subsequently, we merge the pixel samples with high probability with the current label. The probability here refers to the probability that the pixel belongs to the true class. This process enables the introduction of information from old classes that are not directly accessible in the current stage. Moreover, this information is relatively reliable, and the model exhibits confidence in its accuracy. For the remaining pixels, we retain all classes’ information through label smoothing. In the distillation function, the old class and background pixel samples are selected for distillation according to the prediction map of the old classes. The weights of the classes are adaptively updated and adjusted using specific label information from each batch and the different stages of incremental learning. As demonstrated by the results of our experiment, on three remote sensing image datasets: China Computer Federation (CCF), Potsdam, and Vaihingen, our method achieves the best results. Bo Ren 0001, Zhao Wang 0011, Biao Hou, Bo Liu 0009, Zitong Wu, Jocelyn Chanussot, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Which Target to Focus on: Class-Perception for Semantic Segmentation of Remote SensingabstractDeep Learning-based (DL) methods have dominated the task of semantic segmentation of remote sensing images. However, the sizes of different objects vary widely, and there is a great deal of label-noise due to the inevitable shadows. Therefore, there is an urgent need for a method that can precisely handle complex ground data. In this paper, we propose an Inter-Class Enhanced Network (ICEN) for representing features of varying sizes. It comprises two branches: Sparse Representation Network (SPN) and Feature Extraction Network (FEN). Then, a Class-Perception Block is inserted between the two branches to instruct the SPN’s low-level semantic features to be merged into the deeper network. Such a block can reduce label-noise in remote sensing image segmentation. In addition, the proposed EIRI provides a more precise classification process for target edges containing many misclassified points without requiring excessive computational overhead. The experimental results of our proposed Class-Perception Network (C-PNet) achieve competitive performance on the Vaihingen, Potsdam, LoveDA, and UAVid datasets. Lingling Li 0002, Yilin Shao, Licheng Jiao, Xu Liu 0006, Puhua Chen, Fang Liu 0001, Shuyuan Yang 0001, Biao Hou |
IEEE Trans. Geosci. Remote. Sens. | 9 |
| 2023 | A Dual-Stream Transformer With Diff-Attention for Multispectral and Panchromatic ClassificationabstractTo minimize the feature redundancy of multispectral (MS) and panchromatic (PAN) images and maximize the complementary advantages of PAN and MS, a Dual-Stream Transformer with Diff-attention (DSTD)-Net is proposed for PAN and MS classification in this paper. Firstly, in terms of feature extraction, we use Self-attention and Co-attention (SCA) block to extract both specific advantageous features and common essential features. Based on that, a self-attention module strengthened by diff-attention (SSDA) that pays attention to the difference between two specific advantageous features is designed to reduce the essential redundancy in specific features. It can take advantage of the difference between two specific features and reduce the essential redundancy of the specific advantageous features, making them purer and better for classification. Finally, since the specific features and common features of multispectral (MS) and panchromatic (PAN) images make different contributions to classification, a Multi-stage Gated Fusion (MGF) strategy is used. The MGF strategy mainly uses Gated multisource units (GMU) to adapt the weight of different features and fuse them. So, our MGF strategy can strengthen the specific advantageous features beneficial for classification. Above all, the several experiment results verify our proposed networks’ effectiveness and robustness. Our code is available at: https://github.com/blackkiring/DSTD. Lin Xu 0012, Hao Zhu 0009, Licheng Jiao, Wenhao Zhao, Biao Hou, Zhongle Ren, Wenping Ma 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2023 | Multicue Contrastive Self-Supervised Learning for Change Detection in Remote SensingabstractContrastive self-supervised learning (CSSL) is a promising method in extracting effective features from unlabeled data. It performs well in image-level tasks, such as image classification and retrieval. However, the existing CSSL methods are not suitable for pixel-level tasks, e.g., change detection (CD), since they ignore the correlation between local patches or pixels. In this paper, we firstly propose a multi-cue contrastive self-supervised learning (MC-CSSL) method to derive dense features for change detection. Besides data augmentation, the MC-CSSL takes advantage of more cues based on the semantic meaning and temporal correlation of local patches. Specially, the positive pair is built from local patches with the similar semantic meaning or temporal ones with the same geographic location. The assumption is that local patches belonging to the same kind of land-covering tend to share similar features. Secondly, the affinity matrix is truncated and introduced to extract change information between two temporal patches obtained from different types of sensors. As a result, some initial unchanged pixels are selected to serve as the supervision for mapping the dense features into a consistent space. Based on the distance between all bi-temporal pixels in the consistent space, a difference image (DI) is generated and more unchanged pixels can be available. The dense feature mapping and unchanged pixel updating proceed alternately. The proposed CD method is evaluated in both homogeneous and heterogeneous cases and the experimental results demonstrate its effectiveness and priority after comparison with some existing state-of-the-art methods. The source code will be available at https://github.com/Yang202308/ChangeDetection_CSSL. Meijuan Yang, Licheng Jiao, Fang Liu 0001, Biao Hou, Shuyuan Yang 0001, Yake Zhang, Jianlong Wang |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Multiscale Curvelet Scattering NetworkabstractFeature representation has received more and more attention in image classification. Existing methods always directly extract features via convolutional neural networks (CNNs). Recent studies have shown the potential of CNNs when dealing with images' edges and textures, and some methods have been explored to further improve the representation process of CNNs. In this article, we propose a novel classification framework called the multiscale curvelet scattering network (MSCCN). Using the multiscale curvelet-scattering module (CCM), image features can be effectively represented. There are two parts in MSCCN, which are the multiresolution scattering process and the multiscale curvelet module. According to multiscale geometric analysis, curvelet features are utilized to improve the scattering process with more effective multiscale directional information. Specifically, the scattering process and curvelet features are effectively formulated into a unified optimization structure, with features from different scale levels being efficiently aggregated and learned. Furthermore, a one-level CCM, which can essentially improve the quality of feature representation, is constructed to be embedded into other existing networks. Extensive experimental results illustrate that MSCCN achieves better classification accuracy when compared with state-of-the-art techniques. Eventually, the convergence, insight, and adaptability are evaluated by calculating the trend of loss function's values, visualizing some feature maps, and performing generalization analysis. Jie Gao 0013, Licheng Jiao, Fang Liu 0001, Shuyuan Yang 0001, Biao Hou, Xu Liu 0006 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2023 | A Multi-Scale Progressive Collaborative Attention Network for Remote Sensing Fusion ClassificationabstractWith the development of remote sensing technology, panchromatic images (PANs) and multispectral images (MSs) can be easily obtained. PAN has higher spatial resolution, while MS has more spectral information. So how to use the two kinds of images' characteristics to design a network has become a hot research field. In this article, a multi-scale progressive collaborative attention network (MPCA-Net) is proposed for PAN and MS's fusion classification. Compared to the traditional multi-scale convolution operations, we adopt an adaptive dilation rate selection strategy (ADR-SS) to adaptively select the dilation rate to deal with the problem of category area's excessive scale differences. For the traditional pixel-by-pixel sliding window sampling strategy, the patches which are generated by adjacent pixels but belonging to different categories contain a considerable overlap of information. So we change original sampling strategy and propose a center pixel migration (CPM) strategy. It migrates the center pixel to the most similar position of the neighborhood information for classification, which reduces network confusion and increases its stability. Moreover, due to the different spatial and spectral characteristics of PAN and MS, the same network structure for the two branches ignores their respective advantages. For a certain branch, as the network deepens, characteristic has different representations in different stages, so using the same module in multiple feature extraction stages is inappropriate. Thus we carefully design different modules for each feature extraction stage of the two branches. Between the two branches, because the strong mapping methods of directly cascading their features are too rough, we design collaborative progressive fusion modules to eliminate the differences. The experimental results verify that our proposed method can achieve competitive performance. Wenping Ma 0001, Hao Zhu 0009, Licheng Jiao, Jianchao Shen, Biao Hou |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2022 | Deep Modality Independent Descriptor Learning for Optical and SAR Image Patch MatchingabstractDue to the complementary information between multi-modal images, they are widely used in various applications. However, there are significant differences in appearance caused by different imaging mechanisms, which bring great challenges to multi-modal image patch matching. To solve this problem, this paper proposes a deep modality independent descriptor learning network (DMID-Net) for multi-modal image patch matching. DMID-Net computes the self-similarity of deep features as the structure descriptor for image patch matching, which is independent of image modality and shared between multi-modal images. Thus, the acquired deep modality independent descriptor(DMID) can reduce the influence of significant differences between multi-modal images, further improving the matching performances. Experimental results on a large number of optical and SAR image-pairs demonstrate the effectiveness of DMID-Net on multi-modal image patch matching. Huiyuan Wei, Dou Quan, Ruiqi Lei, Baorui Duan, Shuang Wang 0001, Yi Li 0054, Biao Hou, Licheng Jiao |
IGARSS | 7 |
| 2022 | A Transformer-Based Cross-Modal Image-Text Retrieval Method using Feature Decoupling and ReconstructionabstractWith the increasing application of remote sensing technology, the task of cross-modal retrieval of remote sensing images (CMRRS) has gradually attracted widespread attention. Ex-isting methods often completely map the features of different modalities to a shared space and do not decouple between the modal-invariant information and modal-heterogeneous in-formation, which leads to redundant information in feature mapping and usually gets sub-optimal retrieval performance. This paper proposes a Transformer-based CMRRS method using feature decoupling and reconstruction (TBFDR) to solve this problem. TBFDR achieves state-of-the-art performance in remote sensing image-text retrieval task on Sydney-Captions dataset. Yingzhi Sun, Yu Liao, Rui Yang 0038, Shuang Wang 0001, Biao Hou, Licheng Jiao |
IGARSS | 7 |
| 2022 | Few-Shot Hyperspectral Image Classification Based on Domain Adaptation of Class BalanceabstractHyperspectral image (HSI) classification has attracted ever-rising attention to better performance based on limited labeled data. In this paper, a domain adaptation method of class balance based on few-shot learning is proposed, which obtains the classification results of target HSI by training the dataset in the source domain containing sufficient labeled data. We use a random weighted sampling strategy in the source domain and the generative adversarial network (GAN) in the target domain to reduce the label distribution shift caused by unbalanced classes. Then, the conditional maximum mean discrepancy (CMMD) is presented for a more comprehensive domain alignment by considering the posterior data distribution. In addition, the double cross non-local block and multi-scale strategy are adopted in the feature extraction stage to get a refined classification result. Experimental results on public HSI datasets demonstrate that our method is efficient and outperforms other baselines. Qi Zhen, Xiangrong Zhang, Biao Hou, Xu Tang 0004, Licheng Jiao |
IGARSS | 4 |
| 2022 | Deep Multiview Union Learning Network for Multisource Image ClassificationabstractWith the development of the imaging technology of various sensors, multisource image classification has become a key challenge in the field of image interpretation. In this article, a novel classification method, called the deep multiview union learning network (DMULN), is proposed to classify multisensor data. First, an associated feature extractor is designed to process the multisource data by canonical correlation analysis (CCA) in the head of the network. Second, an improved deep learning architecture with two branches is presented to extract high-level view features from the associated features. Third, a novel pooling, called view union pooling, is proposed to fuse the multiview feature from the deep model. Finally, the fused feature is fed into the classifier. The proposed framework is easy to optimize since it is an end-to-end network. Extensive experiments and analysis on the datasets IEEE_grss_dfc_2017 and IEEE_grss_dfc_2018 show that the proposed method achieves comparable results. Our results demonstrate that abundant multisource information can improve the classification performance. Xu Liu 0006, Licheng Jiao, Lingling Li 0002, Fang Liu 0001, Shuyuan Yang 0001, Biao Hou |
IEEE Trans. Cybern. | 7 |
| 2022 | GAFnet: Group Attention Fusion Network for PAN and MS Image High-Resolution ClassificationabstractPanchromatic (PAN) and multispectral (MS) images have coordinated and paired spatial spectral information, which can complement each other and make up for their shortcomings for image interpretation. In this article, a novel classification method called the deep group spatial-spectral attention fusion network is proposed for PAN and MS images. First, the MS image is processed by unpooling to obtain the same resolution as that of the PAN image. Second, the group spatial attention and group spectral attention modules are proposed to extract image features. The PAN and the processed MS images are regarded as the input of the two modules, respectively. Third, the features from the previous step are fused by the attention fusion module, which aims to fully fuse multilevel features, take into account both the low-level features and the high-level features, and maintain the global abstract and local detailed information of the pixels. Finally, the fusion feature is fed into the classifier and the resulting map is obtained by pixel level. Extensive experiments and analysis on four datasets show that the proposed method achieves comparable results. Xu Liu 0006, Lingling Li 0002, Fang Liu 0001, Biao Hou, Shuyuan Yang 0001, Licheng Jiao |
IEEE Trans. Cybern. | 4 |
| 2022 | Remote Sensing Object Tracking With Deep Reinforcement Learning Under OcclusionabstractObject tracking is an important research direction of space Earth observation in the field of remote sensing. Although the existing correlation filter-based and deep learning (DL)-based object tracking algorithms have achieved great success, they are still unsatisfactory for the problem of object occlusion. The occlusion caused by the complex change in background, and the deviation of the tracking lens, causes object information to go missing, which leads to the omission of detection. Traditionally, most methods for object tracking under occlusion adopt a complex network model, which redetects the occluded object. To address this issue, we propose a novel object tracking approach. First, an action decision-occlusion handling network (AD-OHNet) based on deep reinforcement learning (DRL) is built to achieve low computational complexity for object tracking under occlusion. Second, the temporal and spatial context, the object appearance model, and the motion vector are adopted to provide the occlusion information, which drives actions in reinforcement learning under complete occlusion and contributes to improving the accuracy of tracking while maintaining speed. Finally, the proposed AD-OHNet is evaluated on three remote sensing video datasets of Bogota, Hong Kong, and San Diego taken from Jilin-1 commercial remote sensing satellites. The video datasets all shared problems of low spatial resolution, background clutter, and small objects. Experimental results on the three video datasets validate the effectiveness and efficiency of the proposed tracker. Yanyu Cui, Biao Hou, Bo Ren 0001, Shuang Wang 0001, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Network Pruning for Remote Sensing Images Classification Based on Interpretable CNNsabstractConvolutional neural network (CNN)-based research has been successfully applied in remote sensing image classification due to its powerful feature representation ability. However, these high-capacity networks bring heavy inference costs and are easily overparameterized, especially for the deep CNNs pretrained on natural image datasets. Network pruning is regarded as a prevalent approach for compressing networks, but most existing research ignores model interpretability while formulating pruning criterion. To address these issues, a network pruning method for remote sensing image classification based on interpretable CNNs is proposed. More specifically, an original interpretable CNN with a predefined pruning ratio is trained at first. The filters, namely channels in the high convolutional layer, are able to learn specific semantic meanings in proportion to the predefined pruning ratio. The filters without interpretability are supposed to be removed. As for other convolutional layers, a sensitivity function is designed to assess the risk of pruning channels for each layer, and furthermore, the pruning ratio for each layer is corrected adaptively. The pruning method based on the proposed sensitivity function is effective and requires little computational costs to search abandoned channels without damaging classification performance. To demonstrate the effectiveness, the proposed method is implemented on different scales of modern CNN models, including VGG-VD and AlexNet. The experimental results, obtained on the UC Merced dataset and NWPU-RESISC45 dataset, prove that our method significantly reduces the inference costs and improves the interpretability of networks. Xianpeng Guo, Biao Hou, Bo Ren 0001, Zhongle Ren, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | A Neural Network Based on Consistency Learning and Adversarial Learning for Semisupervised Synthetic Aperture Radar Ship DetectionabstractShip detection in synthetic aperture radar (SAR) images has important application value. Sea clutter, complex scenes, a large size change in ships, and the arbitrary directionality of ships make ship detection challenging. With the development of deep learning, many deep learning algorithms have been applied to SAR images. These algorithms need a lot of labeled data for training. It is time-consuming to label SAR data, and the unlabeled data are easy to obtain. It is necessary to use the unlabeled data effectively to improve the performance of the algorithm. In this study, a semisupervised SAR ship detection network, named the semisupervised consistency learning adversarial network (SCLANet), is presented. SCLANet is a two-stage detection network. The local features around the ship can be extracted by the SCLANet, and the features generated from unlabeled data become closer to those generated from labeled data by using adversarial learning. There are two consistency learning modules in SCLANet: noise robustness consistency learning and output encoding consistency learning. Noise robustness consistency learning can increase the robustness of the SCLANet. Maintaining consistency between the noisy results and the original results can train the unlabeled data. In output encoding consistency learning, outputs are mapped to a picture that is fed into an encoder to obtain the intermediate representation embedding. Another embedding is a layer in the main network of the SCLANet. Reducing the error between two embeddings can train the SCLANet with unlabeled data. Two types of consistency learning can be used as pretext tasks for semisupervised learning. Experiments were conducted on two SAR ship datasets. Compared with other algorithms, the SCLANet achieved the highest detection accuracy, indicating that it is more advantageous to use in ship detection. Biao Hou, Zitong Wu, Bo Ren 0001, Xianpeng Guo, Shuang Wang 0001, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | A Two-Stage Mutual Fusion Network for Multispectral and Panchromatic Image ClassificationabstractWith the rapid development of remote sensing technology, satellites can easily obtain multispectral (MS) and panchromatic (PAN) images. How to mine the essence and peculiarity of the MS and PAN images and utilize their complementary to improve classification performance is still a challenge. This paper designs a two-stage mutual fusion network (TSMF-Net) for MS and PAN image classification. The network can be divided into two stages: data fusion and feature fusion. In the data fusion stage, we propose an adaptive twin intensity-hue-saturation (ATIHS) strategy. It not only aligns the size and channels of the MS and PAN images by a novel q-Split operation, but also introduces an adaptive soft-average mask to reduce the differences between replacement components, effectively mitigating spectral distortion and paving the way for the next stage. In the feature fusion stage, we propose a feature graft block (FG-Block) in which we introduce triplet loss and design an interlaced channel addition (ICA) module. Under the supervision of triplet loss, the FG-Block separates and hauls each branch’s essential and peculiar features. With the help of the ICA module, it can effectively graft the essential feature between branches and retain the peculiar feature of each branch, improving the utilization and discrimination of features. Finally, composed of the ATIHS, FG-Blocks, and output layers, our TSMF-Net is proven to improve the accuracy of the remote sensing classification task. The experimental results on multiple datasets verify the effectiveness of our proposed algorithms. Our code is available at: https://github.com/liaoyinuo/TSMF-Net. Yinuo Liao, Hao Zhu 0009, Licheng Jiao, Kenan Sun, Xu Tang 0004, Biao Hou |
IEEE Trans. Geosci. Remote. Sens. | 8 |
| 2022 | Multicrop Fusion Strategy Based on Prototype Assignment for Remote Sensing Image Scene ClassificationabstractThe gap between self-supervised visual representation learning and supervised learning is gradually closing. Self-supervised learning does not rely on a large amount of labeled data and reduces the loss of human labeled information. Compared with natural images, remote sensing images require rich samples and human annotation by experts. Moreover, many algorithms have poor interpretability and unconvincing results. Therefore, this paper proposes a self-supervised method based on prototype assignment by designing a pretext task so that the network maps features to prototypes in the process of learning, swaps the code corresponding to the obtained features, combines them with another data-enhancing feature, and then optimizes the network. The prototype is introduced to explain the clustering idea embodied in the whole process. Considering the existence of the scene information-rich characteristic of remote sensing images, we introduce multiple views with different resolutions to capture more detailed information on the images. Finally, if the data enhancement method is not powerful enough, the network can easily fall into an overfitting state, which prevents the network from learning subtle differences and detailed information. To address this shortcoming, we propose a fusion strategy to flatten the decision boundary of the framework so that the model can also learn the soft similarity between sample pairs. We name the whole framework MFPC. In extensive experiments conducted on three common remote sensing image datasets (i.e., UCMerced, AID, and NWPU45), MFPC achieves a maximum improvement of 4.3% over some existing self-supervised algorithms, indicating that it can achieve good results. Siteng Ma, Biao Hou, Xianpeng Guo, Zhihao Li 0005, Zitong Wu, Shuang Wang 0001, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Feature Split-Merge-Enhancement Network for Remote Sensing Object DetectionabstractRecently, multicategory object detection in high-resolution remote sensing images is still a challenge. First, objects with significant scale differences exist in one scene simultaneously, so it is generally difficult for the detectors to balance the detection performance of large and small objects. Second, because of the complex background and the objects’ densely distributed characteristics in the remote sensing images, the extracted features usually have noise and blurred boundaries, which interfere with the detection performance of the object detectors. With this observation, we propose an end-to-end scale-aware network called feature split–merge–enhancement network (SME-Net) for remote sensing object detection, composed of the feature split-and-merge (FSM) module, the offset-error rectification (OER) module, and the object saliency enhancement (OSE) strategy. FSM eliminates salient information of large objects to highlight the features of small objects in the shallow feature maps. It also transmits the effective detailed features of large objects to the deep feature maps, alleviating feature confusion between multiscale objects. OER corrects the inconsistency of the features spatial layout among the multilayer feature maps by the proposed offset loss, so as to achieve supervised elimination and transmission in FSM. OSE enhances the features of interests and suppresses the background information by the proposed membership function, thus preventing false detection and missed detection caused by noise and blurred boundaries. The effectiveness of the proposed algorithm has been verified on multiple datasets. Our code is available at:https://github.com/Momuli/SMENet.git Wenping Ma 0001, Hao Zhu 0009, Licheng Jiao, Xu Tang 0004, Yuwei Guo 0001, Biao Hou |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2022 | A Collaborative Correlation-Matching Network for Multimodality Remote Sensing Image ClassificationabstractRecently, with the increasing availability of the high-quality panchromatic (PAN) and multispectral (MS) remote sensing (RS) images, the inherent complementarity between PAN and MS images provides a wide development prospect for the multimodality RS image classification task. However, how to cleverly relieve the modal differences and effectively integrate the single-modality PAN and MS features is still a challenge. In this article, we design a collaborative correlation-matching network (CCM-Net) for multimodality RS image classification. Concretely, we first propose a bidirectional dominant feature supervision (Bi-DFS) learning, it utilizes single-modality dominant features as supplementary supervision information to establish the joint optimization loss function, thereby adaptively narrowing the differences between modalities before the feature extraction. In the feature extraction stage, the interactive correlation feature matching (ICFM) learning, composing the spatial feature matching (Spa-FM) and spectral feature matching (Spe-FM) strategies, is proposed to establish interactive matching and enhancement between multimodality strong correlation features from the perspective of spatial and spectral, respectively, thereby effectively alleviating the semantic deviation of multimodality features. Finally, we aggregate finer multilevel multimodality features to obtain top-level features with high discrimination. The effectiveness of the proposed algorithm has been verified on multiple datasets. Our code is available at:https://github.com/Momuli/CCM-Net.git. Wenping Ma 0001, Hao Zhu 0009, Kenan Sun, Zhongle Ren, Xu Tang 0004, Biao Hou, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2022 | Transfer Representation Learning Meets Multimodal Fusion Classification for Remote Sensing ImagesabstractTo maximize the complementary advantages of synergistic multimodal, a transfer representation learning fusion network (TRLF-Net) is proposed for multisource remote sensing images collaborative classification in this article. First, with respect to the feature encoding, we design a dual-branch attention sparse transfer module (DAST-Module), which combines the spatial and channel attention (CA) masks to migrate the advantage attributes of the panchromatic (PAN) and the MS images mutually. This not only enhances their respective image advantages but also facilitates the sparse fusion of low-level features. Second, for the separation of multiscale information, a deep dual-scale decomposition module (DDSD-Module) is designed, which allows the decompose of high-frequency and low-frequency components. Then it uses the decomposed information to make the essential difference as small as possible, and the surrounding contour difference is as large as possible of the complementary multimodal image through the design of the loss function. Finally, to address the problem of large intraclass and small interclass differences, we develop a representation fusion of the global and local features’ module (RFGAL-Module). It mainly adopts global features to sort local features within classes, and then outputs them in a cascade. Thus, the characterization ability of features is improved, and the global and local features are used in a coordinated manner to accomplish the sample classification tasks. In particular, the experimental results demonstrate that TRLF-Net can obtain much improved accuracy and efficiency. The code is accessible in:https://github.com/ru-willow/SRLF-Net. Mengru Ma, Wenping Ma 0001, Licheng Jiao, Xu Liu 0006, Fang Liu 0001, Lingling Li 0002, Shuyuan Yang 0001, Biao Hou |
IEEE Trans. Geosci. Remote. Sens. | 8 |
| 2022 | A Novel Adaptive Hybrid Fusion Network for Multiresolution Remote Sensing Images ClassificationabstractWith the rapid development of earth observation technology, panchromatic (PAN) and multispectral (MS) images have also become easier to obtain. The multiresolution classification of PAN and MS images as a basic MS image analysis task has become a research hotspot. The main challenge in this field is how to process data and extract features to improve classification accuracy effectively. In this article, we design a novel adaptive hybrid fusion network (AHF-Net) for multiresolution remote sensing image classification. It includes two parts: data fusion and feature fusion. In the data fusion part, we propose an adaptive weighted intensity-hue-saturation (AWIHS) strategy, which can reduce the difference between MS and PAN images by adaptively adding each other’s unique information from the perspective of information sharing. In the feature fusion part, starting from the second-order correlation of features, we propose a correlation-based attention feature fusion (CAFF) module. It can improve the discrimination of fusion features by adaptively determining the fusion coefficient according to the importance of the input feature channel. Based on AWIHS and CAFF, inspired by the idea of feature pyramid, we combine the multilevel feature fusion and the dual-branch residual network as the backbone network of AHF-Net. By combining AWIHS and CAFF modules with the backbone network, our AHF-Net can effectively improve the classification accuracy of multiresolution remote sensing images. The effectiveness of the proposed algorithm has been verified on multiple data sets. Our code and model are available athttps://github.com/1826133674/AHF-Net. Wenping Ma 0001, Jianchao Shen, Hao Zhu 0009, Jun Zhang 0045, Jiliang Zhao, Biao Hou, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2022 | Very Low-Resolution Moving Vehicle Detection in Satellite VideosabstractThis paper proposes a practical end-to-end neural network framework to detect tiny moving vehicles in satellite videos with low imaging quality. Some instability factors such as illumination changes, motion blurs, and low contrast to the cluttered background make it difficult to distinguish true objects from noise and other point-shaped distractors. Moving vehicle detection in satellite videos can be carried out based on background subtraction or frame differencing. However, these methods are prone to produce lots of false alarms and miss many positive targets. Appearance-based detection can be an alternative but is not well-suited since classifier models are of weak discriminative power for the vehicles in top view at such low resolution. This article addresses these issues by integrating motion information from adjacent frames to facilitate the extraction of semantic features and incorporating the Transformer to refine the features for key points estimation and scale prediction. Our proposed model can well identify the actual moving targets and suppress interference from stationary targets or background. The experiments and evaluations using satellite videos show that the proposed approach can accurately locate the targets under weak feature attributes and improve the detection performance in complex scenarios. Zhaoliang Pi, Licheng Jiao, Fang Liu 0001, Xu Liu 0006, Lingling Li 0002, Biao Hou, Shuyuan Yang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2022 | Deep Feature Correlation Learning for Multi-Modal Remote Sensing Image RegistrationabstractDeep descriptors have advantages over handcrafted descriptors on local image patch matching. However, due to the complex imaging mechanism of remote sensing images and the significant differences in appearance between multi-modal images, existing deep learning descriptors are unsuitable for multi-modal remote sensing image registration directly. To solve this problem, this paper proposes a deep feature correlation learning network (Cnet) for multi-modal remote sensing image registration. Firstly, Cnet builds a feature learning network based on the deep convolutional network with the attention learning module, to enhance the feature representation by focusing on meaningful features. Secondly, this paper designs a novel feature correlation loss function for Cnet optimization. It focuses on the relative feature correlation between matching and non-matching samples, which can improve the stability of network training and decrease the risk of overfitting. Additionally, the proposed feature correlation loss with a scale factor can further enhance the network training and accelerate the network convergence. Extensive experimental results on image patch matching (Brown, HPatches), cross-spectral image registration (VIS-NIR), multi-modal remote sensing image registration, and single-modal remote sensing image registration have demonstrated the effectiveness and robustness of the proposed method. Dou Quan, Shuang Wang 0001, Yu Gu 0015, Ruiqi Lei, Bowu Yang, Shaowei Wei, Biao Hou, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2022 | Self-Distillation Feature Learning Network for Optical and SAR Image RegistrationabstractOptical and SAR image registration is important for multi-modal remote sensing image information fusion. Recently, deep matching networks have shown better performances than traditional methods on image matching. However, due to significant differences between optical and SAR images, the performances of existing deep learning methods still need to be further improved. This paper proposes a self-distillation feature learning network (SDNet) for optical and SAR image registration, improving performance from network structure and network optimization. Firstly, we explore the impact of different weight-sharing strategies on optical and SAR image matching. Then, we design a partially unshared feature learning network for multi-modal image feature learning. It has fewer parameters than the fully unshared network and has more flexibility than the fully shared network. Additionally, the limited binary supervised information (matching or non-matching) is insufficient to train the deep matching networks for optical-SAR image registration. Thus, we propose a self-distillation feature learning method to exploit more similarity information for deep network optimization enhancing, such as the similarity ordering between a series of non-matching patch-pairs. The exploited rich similarity information will significantly enhance network training and improve matching accuracy. Finally, considering that existing deep learning methods brute-force constrain the features of the matching optical and SAR image patches are similar, which will be lost many discriminative information, degenerating matching performances. Thus, we build an auxiliary task reconstruction learning to optimize the feature learning network to keep more discriminative information. Extensive experiments demonstrate the effectiveness of our proposed method on multi-modal image registration. Dou Quan, Huiyuan Wei, Shuang Wang 0001, Ruiqi Lei, Baorui Duan, Yi Li 0054, Biao Hou, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2022 | A Joint Siamese Attention-Aware Network for Vehicle Object Tracking in Satellite VideosabstractRemote sensing object tracking is a novel and challenging problem due to the negative effects of weak features and background noise. In this paper, from the perspective of attention-focus deep learning, we propose a Joint Siamese Attention-Aware Network (JSANet) for efficient remote sensing tracking which contains both self-attention and cross-attention modules. First, the self-attention modules we propose emphasize the interdependent channel-wise coefficient via channel attention and conduct corresponding space transformation of spatial domain information with spatial attention. Second, the cross-attention is designed to aggregate rich contextual interdependencies between the siamese branches via channel attention and excavate association produces reliable correspondence with spatial attention. In addition, a composite feature combine strategy is designed to fuse multiple attention features. Experimental results on the Jilin-1 satellite video datasets demonstrate that the proposed JSANet achieves state-of-the-art performance in terms of precision and success rate, demonstrate the effectiveness of the proposed methods. Licheng Jiao, Fang Liu 0001, Xu Liu 0006, Lingling Li 0002, Shuyuan Yang 0001, Biao Hou |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2022 | N-Cluster Loss and Hard Sample Generative Deep Metric Learning for PolSAR Image ClassificationabstractDeep learning works normally in PolSAR image classification because the complex terrain scattering characteristic results in large intraclass differences and high interclass similarity. Deep metric learning (DML) aims to make the features keep a closer intraclass and a farther interclass distance. Therefore, we introduce DML and then propose an N-cluster generative adversarial net (N-cluster GAN) framework for PolSAR image classification. However, existing DML losses mainly focus on the relationship between individual samples in feature space. Hence, we propose N-cluster loss that pays more attention to the overall structure of all samples. Meanwhile, traditional hard negative sample mining methods occupy lots of computational resources. In addition, the hard level of the negative samples will affect the model’s performance. Therefore, we explore a new method based on a GAN framework to replace the sample mining. Positive N-cluster loss is added to the discriminator ($D$), and a negative one is added to the generator ($G$). In this way,$D$will possess better classification ability, and$G$can produce hard negative samples for$D$. Then, the hard level of the generated negative samples will change with the discrimination of$D$, which is appropriate for the proposed model. N-cluster loss can be directly calculated through the extracted features rather than redundant data preparation. The proposed model is verified on four PolSAR datasets from two aspects of the loss function and negative samples mining. Then, it achieves competitive performance compared with state-of-the-art algorithms. Chen Yang 0027, Biao Hou, Jocelyn Chanussot, Bo Ren 0001, Shuang Wang 0001, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Reconstruction Error-Based Decomposition Feature Selection for PolSAR ImageabstractTarget decomposition features are the cornerstone of subsequent analyses for PolSAR images. Generally, adopting single or several decomposition algorithms limits the representation ability for original terrain characteristics. Using all the existing decomposition features, however, will definitely increase computational complexity. Besides, some features even have a negative effect on the following tasks. To address these problems, a sparse variational autoencoder feature selection framework (SVAE-FS) is proposed in this article. In detail, the encoder transforms the original feature set into latent space and then decoder reconstructs the corresponding pseudo set on this latent space. Similarly, a pseudo subset is subsequently obtained by the SVAE. The discrepancy, namely reconstruction error, between the pseudo set and the pseudo subset is taken as an evaluation criterion which reflects the feature representation ability of pseudo subset. Sparse constraint in the encoder makes the representative features stand out. Meanwhile, the linear feature transformation layer of the encoder enables the SVAE to evaluate different scale subsets without repeated training. Finally, a greedy selection approach with search scale$K$is proposed to find the suboptimal subset. This procedure not only reduces time consumption, but also ensures the performance of the subset. The selected features are analyzed on four real PolSAR datasets according to the terrain scattering characteristics. Furthermore, these features have achieved competitive performance on three PolSAR image tasks. Chen Yang 0027, Biao Hou, Xianpeng Guo, Bo Ren 0001, Jocelyn Chanussot, Shuang Wang 0001, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | PDFL: Polarimetric Decomposition Feature Learning via Deep AutoencoderabstractModel-based polarimetric target decomposition (TD) generally solves scattering components and parameters under pre-set decomposition base, then decomposition features are also obtained. However, pre-set base could not be adjusted according to different scenes. Furthermore, solving the polarimetric parameters needs to explore additional information or consider limiting conditions to build equations, which is hard and easily to bring negative effects into decomposition features. To this end, we regard the TD as a process of learning decomposition base and features by deep learning. Then, the polarimetric decomposition feature learning (PDFL) model is proposed in this paper. Strictly, this model is not an incoherent TD method but a learning-based method. It dose not need to construct the parameter solution equations or fixed base. Then, the decomposition base and feature can be adaptively learned according to scattering characteristics of current dataset. Due to the characteristics of unsupervised reconstruction, deep auto encoder (DAE) is used as the model foundation. Then, some adjustments and constraints are utilized to make the DAE fit closely with TD. The encoder extracts latent vector from PolSAR data, then the decoder reconstructs pseudo data on this latent vector. The reconstruction can be regarded as the inverse process of TD, so the base matrix of decoder and the latent vector indicate the learned decomposition base and features when the model converges. The effectiveness of PDFL is verified on simulated and real PolSAR datasets. Compared with representative algorithms, proposed model gains more discriminative features and achieves competitive performance on terrain classification and segmentation tasks. Chen Yang 0027, Biao Hou, Bo Ren 0001, Jocelyn Chanussot, Shuang Wang 0001, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | A Joint-Training Two-Stage Method For Remote Sensing Image CaptioningabstractCompared with remote sensing image (RSI) captioning methods based on the traditional encoder-decoder model, two-stage RSI captioning methods include an auxiliary remote sensing task to provide prior information, which enables them to generate more accurate descriptions. In previous two-stage RSI captioning methods, however, the image captioning and the auxiliary remote sensing tasks are handled separately, which is time-consuming and ignores mutual interference between tasks. To solve this problem, we propose a novel joint-training two-stage (JTTS) RSI captioning method. We use multi-label classification to provide prior information, and we design a differentiable sampling operator to replace the traditional non-differentiable sampling operation to index the multi-label classification result. In contrast to previous two-stage RSI captioning methods, our method can implement joint-training, and the joint loss allows the error of the generated description to flow into the optimization of the multi-label classification via back-propagation. Specifically, we approximate the Heaviside step function with the steep logistic function to implement a differentiable sampling operator for the multi-label classification. We propose a dynamic contrast loss function for multi-label classification task to ensure that a certain margin is maintained between the probabilities of the positive label and the negative label during sampling. We design an attribute-guided decoder to filter the multi-label prior information obtained by the sampling operator to generate more accurate image captions. The results of extensive experiments show that the JTTS method achieves state-of-the-art performance on the RSICD, the UCM-Captions, and the Sydney-Captions datasets. Xiutiao Ye, Shuang Wang 0001, Yu Gu 0015, Jihui Wang, Biao Hou, Fausto Giunchiglia, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2022 | Adaptive Dual-Path Collaborative Learning for PAN and MS ClassificationabstractDue to the limitation of sensor technology, researchers tend to obtain high-quality image information from panchromatic (PAN) images and multispectral (MS) images with different resolutions. Therefore, the classification of remote sensing images of PAN and MS have become a research hotspot. In this paper, we propose an adaptive dual-path collaborative learning method for PAN and MS classification. In the stage of sample generation and training, we propose an adaptive neighborhood sample grading (ANSG) strategy in the establishing sample stage so that each pixel to be classified can obtain neighborhood information suitable for itself. Further, to simulate biological cognitive mechanisms, we divide the samples into different levels, and design the self-paced progressive loss (SPL), thus allowing the network to do preference training in different stages. The network’s training can quickly reach the optimal of the current stage and the overall convergence is more thorough. In the network structure, we propose a dual-path module (DPM) to effectively alleviate the gradient degradation in theresidual path, while ensuring maximum gradient loss information flow between every two layers in thedensely connected path. This module can extract more robust features to cope with the complex characteristics of remote-sensing images. Moreover, using the characteristics of the dual path to better fuse the features by the gradual collaborative fusion (GCF) way. The experimental results and theoretical analysis have demonstrated the proposed approach’s effectiveness, feasibility, and robustness. Our model are available at https://github.com/AIpy-nan/DBFI-Net. Hao Zhu 0009, Kenan Sun, Licheng Jiao, Fang Liu 0001, Biao Hou, Shuang Wang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2022 | Coarse-to-Fine Contrastive Self-Supervised Feature Learning for Land-Cover Classification in SAR Images With Limited Labeled DataabstractContrastive self-supervised learning (CSSL) has achieved promising results in extracting visual features from unlabeled data. Most of the current CSSL methods are used to learn global image features with low-resolution that are not suitable or efficient for pixel-level tasks. In this paper, we propose a coarse-to-fine CSSL framework based on a novel contrasting strategy to address this problem. It consists of two stages, one for encoder pre-training to learn global features and the other for decoder pre-training to derive local features. Firstly, the novel contrasting strategy takes advantage of the spatial structure and semantic meaning of different regions and provides more cues to learn than that relying only on data augmentation. Specifically, a positive pair is built from two nearby patches sampled along the direction of the texture if they fall into the same cluster. A negative pair is generated from different clusters. When the novel contrasting strategy is applied to the coarse-to-fine CSSL framework, global and local features are learned successively by forcing the positive pair close to each other and the negative pair apart in an embedding space. Secondly, a discriminant constraint is incorporated into the per-pixel classification model to maximize the inter-class distance. It makes the classification model more competent at distinguishing between different categories that have similar appearance. Finally, the proposed method is validated on four SAR images for land-cover classification with limited labeled data and substantially improves the experimental results. The effectiveness of the proposed method is demonstrated in pixel-level tasks after comparison with the state-of-the-art methods. Meijuan Yang, Licheng Jiao, Fang Liu 0001, Biao Hou, Shuyuan Yang 0001, Yake Zhang, Jianlong Wang |
IEEE Trans. Image Process. | 4 |
| 2022 | New Generation Deep Learning for Video Object Detection: A SurveyabstractVideo object detection, a basic task in the computer vision field, is rapidly evolving and widely used. In recent years, deep learning methods have rapidly become widespread in the field of video object detection, achieving excellent results compared with those of traditional methods. However, the presence of duplicate information and abundant spatiotemporal information in video data poses a serious challenge to video object detection. Therefore, in recent years, many scholars have investigated deep learning detection algorithms in the context of video data and have achieved remarkable results. Considering the wide range of applications, a comprehensive review of the research related to video object detection is both a necessary and challenging task. This survey attempts to link and systematize the latest cutting-edge research on video object detection with the goal of classifying and analyzing video detection algorithms based on specific representative models. The differences and connections between video object detection and similar tasks are systematically demonstrated, and the evaluation metrics and video detection performance of nearly 40 models on two data sets are presented. Finally, the various applications and challenges facing video object detection are discussed. Licheng Jiao, Fang Liu 0001, Shuyuan Yang 0001, Biao Hou, Lingling Li 0002, Xu Tang 0004 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2022 | Element-Wise Feature Relation Learning Network for Cross-Spectral Image Patch MatchingabstractRecently, the majority of successful matching approaches are based on convolutional neural networks, which focus on learning the invariant and discriminative features for individual image patches based on image content. However, the image patch matching task is essentially to predict the matching relationship of patch pairs, that is, matching (similar) or non-matching (dissimilar). Therefore, we consider that the feature relation (FR) learning is more important than individual feature learning for image patch matching problem. Motivated by this, we propose an element-wise FR learning network for image patch matching, which transforms the image patch matching task into an image relationship-based pattern classification problem and dramatically improves generalization performances on image matching. Meanwhile, the proposed element-wise learning methods encourage full interaction between feature information and can naturally learn FR. Moreover, we propose to aggregate FR from multilevels, which integrates the multiscale FR for more precise matching. Experimental results demonstrate that our proposal achieves superior performances on cross-spectral image patch matching and single spectral image patch matching, and good generalization on image patch retrieval. Dou Quan, Shuang Wang 0001, Ning Huyan, Jocelyn Chanussot, Ruojing Wang, Xuefeng Liang, Biao Hou, Licheng Jiao |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2022 | DPFL-Nets: Deep Pyramid Feature Learning Networks for Multiscale Change DetectionabstractDue to the complementary properties of different types of sensors, change detection between heterogeneous images receives increasing attention from researchers. However, change detection cannot be handled by directly comparing two heterogeneous images since they demonstrate different image appearances and statistics. In this article, we propose a deep pyramid feature learning network (DPFL-Net) for change detection, especially between heterogeneous images. DPFL-Net can learn a series of hierarchical features in an unsupervised fashion, containing both spatial details and multiscale contextual information. The learned pyramid features from two input images make unchanged pixels matched exactly and changed ones dissimilar and after transformed into the same space for each scale successively. We further propose fusion blocks to aggregate multiscale difference images (DIs), generating an enhanced DI with strong separability. Based on the enhanced DI, unchanged areas are predicted and used to train DPFL-Net in the next iteration. In this article, pyramid features and unchanged areas are updated alternately, leading to an unsupervised change detection method. In the feature transformation process, local consistency is introduced to constrain the learned pyramid features, modeling the correlations between the neighboring pixels and reducing the false alarms. Experimental results demonstrate that the proposed approach achieves superior or at least comparable results to the existing state-of-the-art change detection methods in both homogeneous and heterogeneous cases. Meijuan Yang, Licheng Jiao, Fang Liu 0001, Biao Hou, Shuyuan Yang 0001, Meng Jian |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2021 | Dual-Stream High Resolution Network for Multi-Source Remote Sensing Image SegmentationabstractRecently, the image segmentation has been a significant research direction in the field of optical remote sensing data processing. However, due to the limitation of the optical imaging mechanism, traditional image segmentation methods are not efficient for processing the optical remote sensing images, especially influencing by the complex weather conditions. In order to ensure the classification performance, synthetic aperture radar (SAR) data are employed as complementary to the data procedure for enhancing the capability of land cover interpretation. Then a dual-stream high-resolution network (HRNet) is proposed to combine two types of heterogeneous data (SAR and optical image), and a multi-modal squeeze-and-excitation (SE) module is exploited to make feature maps fused. Experiments show that the proposed method has excellent performance on the remote sensing data acquired by GF2 and GF3 satellites. Bo Ren 0001, Shibin Ma, Biao Hou, Danfeng Hong |
IGARSS | 3 |
| 2021 | A Feature Decomposition Framework for Multi-Modal Image Patch MatchingabstractMulti-modal remote sensing images have complementary information which is conducive to enhancing the performance of various applications. Image patch matching plays a crucial role in the combination of multi-modal images. However, there are great differences in appearance and texture of multi-modal images, which brings great difficulties to image patching matching. To solve this problem, we propose a novel feature decomposition framework for multi-modal image patch matching. It aims to eliminate the hinder caused by the significant difference in multi-modal images. Specifically, this paper proposes to decompose the feature of images into common feature and modal private feature. Then, only the common feature is used for image patch matching, so as to improve the matching accuracy. Experimental results on optical and SAR images demonstrate that our proposed feature decomposition framework can significantly improve the performance of multi-modal image patch matching. Baorui Duan, Dou Quan, Yi Li 0054, Ruiqi Lei, Shuang Wang 0001, Biao Hou, Licheng Jiao |
IGARSS | 6 |
| 2021 | Deep Global Feature-Based Template Matching for Fast Multi-Modal Image RegistrationabstractDue to the different imaging mechanisms, there is a significant non-line difference between multi-modal images, which brings difficulties to multi-modal image registration. The traditional methods based on grayscale and handcraft features are difficult with obtain common features between different source images. The performances of deep local features matching methods rely on the quality and quantity of the detected keypoints, which can be quite time-consuming to register images. To achieve fast and accurate multi-modal image registration, we propose a deep global feature-based template matching method (GFTM) which uses a deep convolutional network to extract common global deep features from multi-modal images. Then, fast template matching is performed on global deep features to search the position with maximal similarity. Additionally, we build a similarity label map and design three losses to optimize our network, including contrast loss, error loss and peak loss. Extensive experimental results on optical and SAR images demonstrated that our proposed method is effective on multi-modal image registration. Ruiqi Lei, Bowu Yang, Dou Quan, Yi Li 0054, Baorui Duan, Shuang Wang 0001, Huarong Jia, Biao Hou, Licheng Jiao |
IGARSS | 8 |
| 2021 | Multi-View Attention Network for Remote Sensing Image CaptioningabstractIn traditional remote sensing image captioning models, the attention mechanism plays a dominant role and has been used to integrate image features to infer the latent visual-semantic alignment. However, the scenes of remote sensing image are complex and diverse, using only one attention module to capture features often leads to insufficient semantic representation. In our work, we present a novel Multi-view Attention Network (MAN) model to realize feature integration from different views. With MAN, more semantically rich ensemble attended features can be obtained by different attention modules. Specifically, we enforce the weights of attention modules to be diverse through a cosine distance loss. This will provide the model with distinct views to make semantic predictions for each feature. Extensive experiments on benchmark datasets demonstrate the effectiveness of the proposed model for the task of remote sensing image captioning. Yun Meng, Yu Gu 0015, Xiutiao Ye, Jingxian Tian, Shuang Wang 0001, Biao Hou, Licheng Jiao |
IGARSS | 7 |
| 2021 | Cross-Modal Feature Fusion Retrieval for Remote Sensing Image-Voice RetrievalabstractWith the increasing popularity of remote sensing technology applications, some emergency scenarios require rapid retrieval of remote sensing images, such as earthquake rescue, etc. Due to the high efficiency of voice input, researchers have focused on cross-modal remote sensing image-voice retrieval methods. However, these methods have two major drawbacks: speech input lacks discrimination and the intra-modal semantic information is under used. To address these drawbacks, we propose a novel cross-modal feature fusion retrieval model. Our model provides a more optimized cross-modal common feature space than previous models and thus optimizes the retrieval performance. First, our model adds the extra textual keyword information to the audio feature for remote sensing image retrieval. Second, it introduces inter-modality adversarial learning and intra-modality semantic discrimination into the remote sensing image-voice retrieval task. We conducted experiments on two datasets modified from the UCM-Captions dataset and the Remote Sensing Image Caption Dataset. The experimental results show that our model outperforms state-of-the-art models in this task. Rui Yang 0038, Yu Gu 0015, Yu Liao, Yingzhi Sun, Shuang Wang 0001, Biao Hou, Licheng Jiao |
IGARSS | 7 |
| 2021 | Real-time Surveillance Video Salient Object Detection Using Collaborative Cloud-Edge Deep Reinforcement LearningabstractIn recent years, with the advancement of cloud computing technology and the availability of cheaper hardware, surveillance systems have become more and more common. Unfortunately, most existing systems still face many limitations, such as latency and real-time analysis issues, etc. Edge computing effectively expands the boundaries of cloud computing, migrating some computing and analysis tasks to the edge devices for execution. Edge device could perform video analysis, which may be a good solution. In this paper, we adopt the collaborative Cloud-Edge architecture to analyze surveillance video and extract video keyframes for compressing video data at the edge. Then, we provide a residual U-net neural network to perform salient object detection on the extracted keyframes. Finally, we utilize the deep reinforcement learning Asynchronous Advantage Actor-Critic (A3C) algorithm to perform the residual U-net tasks scheduling, adaptive offloading in the cloud or edge, reducing system latency, and improving real-time performance. We verified the system performance using real road surveillance videos and other public datasets. The experiment results are inspiring. It proves that the real-time processing of the surveillance video system based on a collaborative cloud-edge mechanism could obtain the optimal result within the range of tolerable latency. Biao Hou, Junxing Zhang |
IJCNN | 1 |
| 2021 | Distance constraint between features for unsupervised domain adaptive person re-identification
Zhihao Li 0005, Bing Han 0003, Xinbo Gao 0001, Biao Hou, Zongyuan Liu 0003 |
Neurocomputing | 4 |
| 2021 | Hyperspectral image classification based on spatial and spectral kernels generation network
Wenping Ma 0001, Hao Zhu 0009, Licheng Jiao, Biao Hou |
Inf. Sci. | 7 |
| 2021 | A Mutual Information-Based Self-Supervised Learning Model for PolSAR Land Cover ClassificationabstractRecently, deep learning methods have attracted much attention in the field of polarimetric synthetic aperture radar (PolSAR) data interpretation and understanding. However, for supervised methods, it requires large-scale labeled data to achieve better performance, and getting enough labeled data is a time-consuming and laborious task. Aiming to obtain a good classification result with limited labeled data, we focus on learning discriminative high-level features between multiple representations, which we call mutual information. As PolSAR data have multi-modal representations, there should have strong similarity between multi-modal features of the same pixel. In addition, each pixel has its own unique geocoding and scattering information. Hence, every pixel has great difference from other pixels in a specific representation space. Based on the above observations, this article proposes a mutual information-based self-supervised learning (MI-SSL) model to learn an implicit representation from unlabeled data. In this article, the self-supervised learning idea is first applied to PolSAR data processing. Furthermore, a reasonable pretext task, which is suitable for PolSAR data, is designed to extract mutual information for classification tasks. Compared with the state-of-the-art classification methods, experimental results on four PolSAR data sets demonstrate that our MI-SSL model produces impressive overall accuracy with fewer labeled data. Bo Ren 0001, Biao Hou, Jocelyn Chanussot, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2021 | Multiscale CNN With Autoencoder Regularization Joint Contextual Attention Network for SAR Image ClassificationabstractSynthetic aperture radar (SAR) image classification is a fundamental research direction in image interpretation. With the development of various intelligent technologies, deep learning techniques are gradually being applied to SAR image classification. In this study, a new SAR classification algorithm known as the multiscale convolutional neural network with an autoencoder regularization joint contextual attention network (MCAR-CAN) is proposed. The MCAR-CAN has two branches: the autoencoder regularization branch and the context attention branch. First, autoencoder regularization is used for the reconstruction of the input to regularize the classification in the autoencoder regularization branch. Multiscale input and an asymmetric structure of the autoencoder branch cause the network more to be focused on classification than on reconstruction. Second, the attention mechanism is used to produce an attention map in which each attention weight corresponds to a context correlation in attention branch. The robust features are obtained by the attention mechanism. Finally, the features obtained by the two branches are spliced for classification. In addition, a new training strategy and a postprocessing method are designed to further improve the classification accuracy. Experiments performed on the data from three SAR images demonstrated the effectiveness and robustness of the proposed algorithm. Zitong Wu, Biao Hou, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2021 | Cost-Sensitive Latent Space Learning for Imbalanced PolSAR Image ClassificationabstractLand cover classification is an important application for polarimetric synthetic aperture radar (PolSAR) image interpretation. The classification performance of a promising parametric feature and classifier learning-based algorithm is limited when the amounts of pixels from different classes vary greatly. PolSAR data from minority classes is difficult to recognize correctly owing to a strong learning bias toward the majority classes, resulting in under-performing features for minority classes. To address this issue, a cost-sensitive latent space learning network based on the feature and classifier learning framework is proposed to reduce the learning bias for supporting the classification of imbalanced data in PolSAR images. First, a new cost-sensitive method is developed by adaptively computing the cost coefficient from predicted labels in the optimization process. Thus, the imbalanced distribution of PolSAR data can be obtained for both the labeled and unlabeled pixels rather than a predefined misclassifying matrix for labeled pixels. Second, latent space learning is used as an auxiliary task to assist the main task of classifier learning. By weighting the distance between the learned feature and the basis of the latent space with a different cost-sensitive coefficient, pixels in minority and majority classes are promoted to be more separable. Thus, the strong bias to majority classes is reduced from both the feature learning and classification process. Finally, the proposed method is studied through experiments on three different PolSAR images with several existing state-of-the-art methods. The experiments validate the effectiveness of the proposed method for balanced and imbalanced PolSAR land cover classification. Biao Hou, Zaidao Wen, Zhongle Ren, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2021 | Selective Adversarial Adaptation-Based Cross-Scene Change Detection Framework in Remote Sensing ImagesabstractSupervised change detection methods always face a big challenge that the current scene (target domain) is fully unlabeled. In remote sensing, it is common that we have sufficient labels in another scene (source domain) with a different but related data distribution. In this article, we try to detect changes in the target domain with the help of the prior knowledge learned from multiple source domains. To achieve this goal, we propose a change detection framework based on selective adversarial adaptation. The adaptation between multisource and target domains is fulfilled by two domain discriminators. First, the first domain discriminator regards each scene as an individual domain and is designed for identifying the domain to which each input sample belongs. According to the output of the first domain discriminator, a subset of important samples is selected from multisource domains to train a deep neural network (DNN)-based change detection model. As a result, not only the positive transfer is enhanced but also the negative transfer is alleviated. Second, as for the second domain discriminator, all the selected samples are thought from one domain. Adversarial learning is introduced to align the distributions of the selected source samples and the target ones. Consequently, it further adapts the knowledge of change from the source domain to the target one. At the fine-tuning stage, target samples with reliable labels and the selected source ones are used to jointly fine-tune the change detection model. As the target domain is fully unlabeled, homogeneity- and boundary-based strategies are exploited to make the pseudolabels from a preclassification map reliable. The proposed method is evaluated on three SAR and two optical data sets, and the experimental results have demonstrated its effectiveness and superiority. Meijuan Yang, Licheng Jiao, Biao Hou, Fang Liu 0001, Shuyuan Yang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2021 | Multi-Relation Attention Network for Image Patch MatchingabstractDeep convolutional neural networks attract increasing attention in image patch matching. However, most of them rely on a single similarity learning model, such as feature distance and the correlation of concatenated features. Their performances will degenerate due to the complex relation between matching patches caused by various imagery changes. To tackle this challenge, we propose a multi-relation attention learning network (MRAN) for image patch matching. Specifically, we propose to fuse multiple feature relations (MR) for matching, which can benefit from the complementary advantages between different feature relations and achieve significant improvements on matching tasks. Furthermore, we propose a relation attention learning module to learn the fused relation adaptively. With this module, meaningful feature relations are emphasized and the others are suppressed. Extensive experiments show that our MRAN achieves best matching performances, and has good generalization on multi-modal image patch matching, multi-modal remote sensing image patch matching and image retrieval tasks. Dou Quan, Shuang Wang 0001, Yi Li 0054, Bowu Yang, Ning Huyan, Jocelyn Chanussot, Biao Hou, Licheng Jiao |
IEEE Trans. Image Process. | 7 |
| 2020 | QoE Estimation of DASH-Based Mobile Video Application Using Deep Reinforcement Learning
Biao Hou, Junxing Zhang |
ICA3PP (2) | 1 |
| 2020 | PolSAR Scene Classification via Low-Rank Tensor-Based Multi-View Subspace RepresentationabstractIn this paper, the polarimetric synthetic aperture radar (PolSAR) scene classification is solved by using a novel low -rank tensor-based multi-view subspace representation (LRT-MSR) method. PolSAR data can be described in multimodal feature spaces, such as PolSAR coherent/covariance/scattering matrices, or the various polarimetric decompositions. Different pseudo-color images from multiple spaces provide enough visual information for making a comprehensive representation. Our method applies a low-rank tensor-based subspace clustering way to explore the information from multi-view pseudo-color images. Tensor, as the high order matrix, is used to capture the correlations of underlying multi-view data. Furthermore, the method is constrained by a low-rank term that elegantly models the cross information from different views, and achieves a series of representation matrices from the redundant information. Finally, a spectral cluster method is used to make the final classification. The experimental results on PolSAR image dataset show the effectiveness of the applied method. Mengqian Chen, Bo Ren 0001, Biao Hou, Jocelyn Chanussot, Shuang Wang 0001, Xiangrong Zhang |
IGARSS | 3 |
| 2020 | Panchromatic Image Land Cover Classification Via DCNN with Updating Iteration StrategyabstractLand cover classification is a critical research task in many significant remote sensing applications. There are emerging many powerful pixel-level classification methods based on deep convolutional neural network (DCNN) in the universal computer vision community. However, due to the complication of satellite image senses and the lack of high-quality labeled datasets, these computer vision techniques can not be applied to remote sensing applications directly. In this paper, we propose a novel DCNN method to extract abstract feature from the complicated remote scene. The proposed method fuses three level features from the encoder while the segmentation result is obtained by decoder. Furthermore, we propose an updating iteration strategy (UIS) with label smoothing on training set to reduce the impact of the incorrect labels. The proposed strategy employs the output of the network to update the low-confidence labels on training set, and utilizes the updated labels to continue training the network. In order to acquire a better segmentation result on a very high resolution (VHR) panchromatic image, we transfer the features trained on GID dataset to our dataset for training. Our experiments has demonstrated the oustanding performance of the proposed method in land cover classification compared to DeepLabv3 on the GID and our dataset. Biao Hou, Yangfei Liu, Tuotuo Rong, Bo Ren 0001, Zijuan Xiang, Xiangrong Zhang, Shuang Wang 0001 |
IGARSS | 1 |
| 2020 | Semisupervised Classification of PolSAR Image Incorporating Labels' Semantic PriorsabstractDeep learning techniques represented by deep convolutional neural networks (CNNs) have been widely used in polarimetric synthetic-aperture radar (PolSAR) image classification in recent years. One challenge is how to get a pleasant classification result with limited human-labeled samples. In this letter, a novel semisupervised classification method incorporating labels' semantic priors is proposed for PolSAR image classification with limited labeled samples. The core idea is that a good classification result map should have consistent regions and aligned boundaries. Thus, a cost function is proposed, which mainly contains three terms including a supervised term, a region consistency term, and a boundary kept term. The supervised term enforces the category label to be the same with the human-labeled labels. The region consistency term encourages the labels in one region to be consistent. The boundary-kept term constrains the region consistency term, preventing the classification map from being too smooth. An alternate iterative optimization method is proposed to solve this equation. First, a CNN is trained using the labeled samples and the classification probability map. Then, the classification probability map is updated by the prediction of the trained CNN and the labeled samples. Repeat these two procedures until the maximum number of iterations is met. Experiments on two real PolSAR images are conducted to validate the effectiveness of the proposed method compared with several state-of-the-art methods. Biao Hou, Jiaojiao Guan, Licheng Jiao |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2020 | Modified Tensor Distance-Based Multiview Spectral Embedding for PolSAR Land Cover ClassificationabstractThis letter proposes a novel method for combining multiview features in polarimetric synthetic aperture radar (PolSAR) for land cover classification. It is well-known that feature extraction and classifier design are two significant steps in machine learning methods for PolSAR data interpretation. Each PolSAR pixel can be represented in different feature spaces, such as polarimetric data scattering, or the polarimetric target decomposition spaces. In this letter, a tensor-based multiview embedding algorithm is proposed to fuse those features from different spaces in order to obtain a distinctive set of features for the subsequent classification. Based on the pixel-based classification tasks, a modified tensor distance (MTD) is designed to accurately calculate the distance between tensors. It emphasizes the importance of the central pixel, and decreases the influence of the neighbors in the feature patch when calculating tensor distance. Furthermore, the complementary properties of different views are exploited by an MTD measured tensor multiview spectral embedding method, so as to obtain relevant low-dimensional features. Compared with state-of-the-art methods, the validation and effectiveness of the proposed method is demonstrated on two real PolSAR data sets. Bo Ren 0001, Biao Hou, Jocelyn Chanussot, Licheng Jiao |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2020 | Object Detection in High-Resolution Panchromatic Images Using Deep Models and Spatial Template MatchingabstractAutomatic object detection from remote sensing images has attracted a significant attention due to its importance in both military and civilian fields. However, the low confidence of the candidates restricts the recognition of potential objects, and the unreasonable predicted boxes result in false positives (FPs). To address these issues, an accurate and fast object detection method called the refined single-shot multibox detector (RSSD) is proposed, consisting of a single-shot multibox detector (SSD), a refined network (RefinedNet), and a class-specific spatial template matching (STM) module. In the training stage, fed with augmented samples in diverse variation, the SSD can efficiently extract multiscale features for object classification and location. Meanwhile, RefinedNet is trained with cropped objects from the training set to further enhance the ability to distinguish each class of objects and the background. Class-specific spatial templates are also constructed from the statistics of objects of each class to provide reliable object templates. During the test phase, RefinedNet improves the confidence of potential objects from the predicted results of SSD and suppresses that of the background, which promotes the detection rate. Furthermore, several grotesque candidates are rejected by the well-designed class-specific spatial templates, thus reducing the false alarm rate. These three parts constitute a monolithic architecture, which contributes to the detection accuracy and maintains the speed. Experiments on high-resolution panchromatic (PAN) images of satellites GaoFen-2 and JiLin-1 demonstrate the effectiveness and efficiency of the proposed modules and the whole framework. Biao Hou, Zhongle Ren, Wei Zhao 0014, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2020 | PolSAR Feature Extraction Via Tensor Embedding Framework for Land Cover ClassificationabstractPolarimetric synthetic aperture radar (PolSAR) as a typical multi-channel sensor can obtain refined geometrical and geophysical information. In the PolSAR land cover classification task, feature extraction is regarded as a critical step for the final classification. It can employ multi-modal features from the original scattering data, polarimetric target decomposition, and other transformation space. Then, how to efficiently combine multi-modal polarimetric information and extract discriminant features is an important challenge for PolSAR image processing. Graph embedding methods have become a significant technique to deal with feature extraction and dimensionality reduction (DR) problems in recent years. It provides a unified linearization framework in machine learning and other pattern recognition tasks. In this article, an extended tensor embedding framework is introduced to extract the intrinsic features for PolSAR land cover classification. First, each pixel is represented by a feature cube that is constructed by groups of polarimetric scattering signals and target decomposition features in a fixed size patch. Second, an intrinsic matrix is constructed to describe the original geometrical and statistical properties of the samples, and a penalty matrix is designed to represent some constraints. Third, the vector-based algorithms are transformed into tensor space in an unified framework and based on the pair of matrices to obtain the projection matrices in each mode by an iterative optimization process. The effectiveness of the proposed methods is demonstrated on three RADARSAT2 data sets covering the regions of Xi'an, San Francisco, and Flevoland, respectively. The visualization and quantification results show that the proposed method has superiority in land cover classification. Bo Ren 0001, Biao Hou, Jocelyn Chanussot, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2020 | A Distribution and Structure Match Generative Adversarial Network for SAR Image ClassificationabstractSynthetic aperture radar (SAR) image classification is a fundamental research in the interpretation of SAR images. The previous methods are unilaterally based on statistical features or spatial features, which cannot capture features with complete SAR image characteristics and unavoidably limits the performance for classification. In this article, novel sample weighting and class adversarial training strategies are proposed to fuse complementary SAR characteristics. Based on these, a distribution and structure match auxiliary classifier generative adversarial network (DSM-ACGAN) is constructed for high-quality discriminative feature learning. Particularly, the characteristics of statistical distribution and spatial structure are jointly considered in class adversarial training of DSM-ACGAN. On the one hand, DSM-ACGAN sets the true SAR image characteristics as goals for the generator to learn generative models of each category. On the other hand, and more importantly, it guides the discriminator to simultaneously capture the desired statistical and structural features. Through the class adversarial processing, the discriminative feature learning progressively improves and contributes to classification. Additionally, class-balanced and plausible samples can be generated. Experimental results on three broad SAR images from different satellites confirm the effectiveness of class adversarial training and the superiority of discriminative feature learning in DSM-ACGAN. Visual performance and quantitative metrics also show the state-of-the-art performance of the novel model. Zhongle Ren, Biao Hou, Zaidao Wen, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2020 | Semi-Supervised PolSAR Image Classification Based on Improved Tri-Training With a Minimum Spanning TreeabstractIn this article, the terrain classifications of polarimetric synthetic aperture radar (PolSAR) images are studied. A novel semi-supervised method based on improved Tri-training combined with a neighborhood minimum spanning tree (NMST) is proposed. Several strategies are included in the method: 1) a high-dimensional vector of polarimetric features that are obtained from the coherency matrix and diverse target decompositions is constructed; 2) this vector is divided into three subvectors and each subvector consists of one-third of the polarimetric features, randomly selected. The three subvectors are used to separately train the three different base classifiers in the Tri-training algorithm to increase the diversity of classification; and 3) a help-training sample selection with the improved NMST that uses both the coherency matrix and the spatial information is adopted to select highly reliable unlabeled samples to increase the training sets. Thus, the proposed method can effectively take advantage of unlabeled samples to improve the classification. Experimental results show that with a small number of labeled samples, the proposed method achieves a much better performance than existing classification methods. Shuang Wang 0001, Yanhe Guo, Wenqiang Hua, Xinan Liu, Guoxin Song, Biao Hou, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2020 | POL-SAR Image Classification Based on Modified Stacked Autoencoder Network and Data DistributionabstractThis article proposes a novel autoencoder (AE) network based on the distribution of polarimetric synthetic aperture radar (POL-SAR) data matrix, called a mixture autoencoder (MAE). Through a detailed analysis of the data distribution POL-SAR data matrix, a normalization method is also presented in succession. The proposed MAE defines the data error term in the loss function according to the data distribution. It can be regarded as a process of unsupervised feature extraction designed specifically for POL-SAR data matrix. Then, a softmax classifier is trained with the help of data features and the corresponding label information. Next, a stacked MAE (SMAE) network is reasonably constructed by considering the data distribution among different layers. Finally, this article also presents a classification network through discarding the decoder process of the proposed SMAE and connecting with a softmax classifier. The SMAE is trained layer by layer using the unlabeled data. The softmax classifier is also trained with a small number of labeled pixels. With parameters obtained from the above-mentioned procedures as the initial parameters, the whole classification network is trained by the labeled pixels to get a well-trained model, which is used for predicting the corresponding label of the pixel in the data set. Three real POL-SAR data sets, including the AIR-SAR L-band data of Flevoland, The Netherlands, are used in the experiments. Compared with one classical algorithm and two related models with the similar structure, both the proposed methods show improvements in overall accuracy and efficiency as well as possess better adaptability of the parameter and preferable consistency with the classification performance. Jianlong Wang, Biao Hou, Licheng Jiao, Shuang Wang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2019 | An Improved Fully Convolutional Network for Learning Rich Building FeaturesabstractMany efficient approaches are proposed to detect building in remote sensing images. In this paper, in order to learning rich building features better, we propose a full convolutional network with dense connection. There contributions are made: 1) To strengthen feature propagation, an improved dense network is introduced to the full convolution network. 2) We have designed top-down short connections to facilitate the fusion of high and low feature information. 3) In addition, we add the weighted cross entropy edge loss function to make the network pay more attention to building edge in detail. Experiments show that the proposed method achieves excellent performance on the remote sensing image data taken by the QuickBird satellite. Shuang Wang 0001, Pei He, Dou Quan, Xuefeng Liang, Biao Hou |
IGARSS | 7 |
| 2019 | Object Detection and Trcacking Based on Convolutional Neural Networks for High-Resolution Optical Remote Sensing VideoabstractObject detection algorithms, from high-resolution optical remote sensing images, have been booming from the last few years. However, object tracking for high-resolution optical remote sensing video is a challenging task due to the large number and small size of objects. In this paper, we propose an object detection and tracking method based on deep convolutional neural networks for wide swath high-resolution optical remote sensing videos. The proposed method firstly segments each frame of a video into sub-samples using a sliding window of fixed size. In order to detect the objects appearing at the edge of the sliding window efficiently, we use an overlapping sliding window sampling method. Further, we design a network fusing region of interests (RoIs) of the previous and current frames to track the objects occurred in the previous frames of the video. RoIs of previous frame are applied directly to the feature layer of the current frame. Finally, for each frame, we merge the detection and tracking results of sub-samples by non-maximum suppression (NMS) method. The experimental results on our dataset demonstrate the validity and generality of the proposed detection algorithm. Biao Hou, Jingliang Li, Xiangrong Zhang, Shuang Wang 0001, Licheng Jiao |
IGARSS | 1 |
| 2019 | Polsar Land Cover Classification via Tensorial Embedding MethodsabstractIn recent years, graph embedding has become a significant technique to deal with feature extraction and dimension reduction problems. Under the linearization and kernelization, it provides a unified framework in machine learning and other pattern recognition tasks. Polarimetric synthetic aperture (PolSAR) as a typical multi-channel sensor can obtain more geometrical and geophysical information. How to combine those polarimetric scattering signals and target decomposition features and explore the spatial information between pixels become a new research direction to address PolSAR data. In this paper, we utilize the tensorial embedding methods to extract the intrinsic features from a redundant feature space for the PolSAR land cover classification. The effectiveness of the proposed methods is demonstrated using AIRSAR Flevoland data set. Bo Ren 0001, Biao Hou, Jocelyn Chanussot, Changzhe Jiao, Xiangrong Zhang |
IGARSS | 2 |
| 2019 | Fast unsupervised deep fusion network for change detection of multitemporal SAR images
Huan Chen 0006, Licheng Jiao, Miaomiao Liang, Fang Liu 0001, Shuyuan Yang 0001, Biao Hou |
Neurocomputing | 6 |
| 2019 | Fast Semisupervised Classification Using Histogram-Based Density Estimation for Large-Scale Polarimetric SAR DataabstractIn order to obtain high classification accuracy and reduce time consumption for large-scale polarimetric synthetic-aperture radar (PolSAR) data. In this letter, we propose a fast semisupervised classification algorithm using histogram-based density estimation (called FSHDE). First, a noniterative collaborative training using our proposed Wishart-clustering selection strategy is designed to expand the labeled sample set from unlabeled samples. Second, a fast feature mapping based on histogram density estimation is employed to reliably capture the interaction of nonlinear features. Third, submodular optimization is used to select optimal subspace features to reduce feature correlation. Experimental results on synthetic and real PolSAR data indicate that FSHDE greatly reduces the time consumption and improves the accuracy for terrain classification compared with the state-of-the-art methods. Hongying Liu 0001, Feixiang Wang, Shuyuan Yang 0001, Biao Hou, Licheng Jiao, Ri Yang |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2019 | Variational Learning of Mixture Wishart Model for PolSAR Image ClassificationabstractThe phase difference, amplitude product, and amplitude ratio between two polarizations are important discriminators for terrain classification, which derives a significant statistical-distribution-based polarimetric synthetic aperture radar (PolSAR) image classification. Traditionally, statistical-distribution-based PolSAR image classification models pay attention to two aspects: searching for a suitable distribution to model certain PolSAR image and a satisfactory solution for the corresponding distribution model with samples in every terrain. Usually, the described distribution form is too complicated to build. Besides, inaccurate parameter estimation may lead to poor classification performance for PolSAR image. In order to refrain from this phenomenon, a variational thought is adopted for the statistical-distribution-based PolSAR classification method in this paper. First, a mixture Wishart model is built to model the PolSAR image to replace the complicated distribution for the PolSAR image. Second, a learning-based method is suggested instead of inaccurate point estimation of parameters to determine the distribution for every class in the mixture Wishart model. Finally, the proposed learning-based mixture Wishart model will be built as a variational form to realize a parametric model for PolSAR image classification. In the experiments, it will be proved that the class centers are easier to distinguish among different terrains learned from the proposed variational model. In addition, a classification performance on the PolSAR image is superior to the original point estimation Wishart model on both visual classification result and accuracy. Biao Hou, Zaidao Wen, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2019 | CNN-Based Polarimetric Decomposition Feature Selection for PolSAR Image ClassificationabstractIn order to better interpret polarimetric synthetic aperture radar (PolSAR) images, many scholars tend to do target decomposition for PolSAR images and utilize the obtained features to perform subsequent classification. These target decomposition features play an important role in terrain classification but completely utilizing them produces a high computational complexity. Furthermore, some features have a negative impact on the classification task. Therefore, selecting the appropriate amount of high-quality features is of great significance to the classification task. In this paper, we propose a convolutional neural network (CNN)-based feature selection algorithm for PolSAR image classification. First, we design a 1-D CNN for feature selection, then train the designed network with all the decomposition features to obtain a trained model. Second, the Kullback-Leibler distance (KLD) between different features is utilized as a standard to select feature subsets. Third, feature subsets with excellent performance form the final results. Due to the special structure of the 1-D CNN, repetitively training model is avoided when the input changes. Different from traditional feature selection methods, our method considers the performance of features combination rather than single feature contribution. To this end, the feature subsets selected by the proposed method are more useful to the classification task. Innovatively introducing KLD in the selection stage avoids random selection and improves the selection efficiency. Finally, we validate the performance of selected feature subsets in traditional and deep learning classification frameworks. Experiments demonstrate that features selected by the proposed method have a good performance comparing with others on three real PolSAR data sets. Chen Yang 0027, Biao Hou, Bo Ren 0001, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2019 | Transferred Deep Learning-Based Change Detection in Remote Sensing ImagesabstractSupervised deep neural networks (DNNs) have been extensively used in diverse tasks. Generally, training such DNNs with superior performance requires a large amount of labeled data. However, it is time-consuming and expensive to manually label the data, especially for tasks in remote sensing, e.g., change detection. The situation motivates us to resort to the existing related images with labels, from which the concept of change can be adapted to new images. However, the distributions of the related labeled images (source domain) and unlabeled new images (target domain) are similar but not identical. It impedes a change detection model learned from source domains being well applied to the target domain. In this paper, we propose a transferred deep learning-based change detection framework to solve this problem. It consists of pretraining and fine-tuning stages. In the pretraining process, we propose two tasks to be learned simultaneously, namely, change detection for the source domain with labels and reconstruction of the unlabeled target data. The auxiliary task aims to reconstruct the difference image (DI) for the target domain. DI is an effective feature, such that the auxiliary task is of much relevance to change detection. The lower layers are shared between these two tasks in the training process. It mitigates the distribution discrepancy between the source and target domains and makes the concept of change from the source domain adapt to the target domain. In addition, we evaluate three modes of the U-net architecture to merge the information for a pair of patches. To fine-tune the change detection network (CDN) for the target domain, two strategies are exploited to select the pixels that have a high possibility of being correctly classified by an unsupervised approach. The proposed method demonstrates an excellent capacity for adapting the concept of change from the source domain to the target domain. It outperforms the state-of-the-art change detection methods via experimental results on real remote sensing data sets. Meijuan Yang, Licheng Jiao, Fang Liu 0001, Biao Hou, Shuyuan Yang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2019 | A Novel Segmentation Based Depth Map Up-SamplingabstractA novel color image segmentation-based depth map upsampling method is proposed in this paper. In this method, the color image is segmented into a certain number of connected regions first. Based on the segmentation result, the target pixels will be interpolated by the seed pixels11The seed pixels are directly from the low resolution depth maps, i.e., the ones that have depth values. The targets are those without depth and to be interpolated. regionally. In the segmentation part, simple linear iterative clustering is introduced to generate superpixels in the first place. Then, the obtained superpixels will be judged whether they are correct-clustered or not, and the incorrect-clustered ones will be subdivided with an adaptive region-growing strategy. Third, the regions that have no seed will be constantly merged into their nearest neighbors, until seed pixel can be found in each independent region. Finally, adjacent regions that have quite small depth gaps will be united as one. The proposed color image segmentation strictly follows the guidance of the depth; therefore, the segmented regions adhere to the depth boundary well. In the interpolation part, the targets will be interpolated with their surrounding seeds weighted by a joint trilateral filter (JTF). The JTF is constructed by three terms: the color term, the distance term, and the region term, which are driven by the previous segmentation result. Experimental results indicate that our method greatly reduces depth bleeding and depth confusion artifacts, and leads to clear depth boundary in the up-sampled image. Comparisons with the state of art verify the advantages of the proposed method in both visual experience and quantitative evaluations. Yiguo Qiao, Licheng Jiao, Shuyuan Yang 0001, Biao Hou |
IEEE Trans. Multim. | 4 |
| 2019 | Local Restricted Convolutional Neural Network for Change Detection in Polarimetric SAR ImagesabstractTo detect changed areas in multitemporal polarimetric synthetic aperture radar (SAR) images, this paper presents a novel version of convolutional neural network (CNN), which is named local restricted CNN (LRCNN). CNN with only convolutional layers is employed for change detection first, and then LRCNN is formed by imposing a spatial constraint called local restriction on the output layer of CNN. In the training of CNN/LRCNN, the polarimetric property of SAR image is fully used instead of manual labeled pixels. As a preparation, a similarity measure for polarimetric SAR data is proposed, and several layered difference images (LDIs) of polarimetric SAR images are produced. Next, the LDIs are transformed into discriminative enhanced LDIs (DELDIs). CNN/LRCNN is trained to model these DELDIs by a regression pretraining, and then a classification fine-tuning is conducted with some pseudolabeled pixels obtained from DELDIs. Finally, the change detection result showing changed areas is directly generated from the output of the trained CNN/LRCNN. The relation of LRCNN to the traditional way for change detection is also discussed to illustrate our method from an overall point of view. Tested on one simulated data set and two real data sets, the effectiveness of LRCNN is certified and it outperforms various traditional algorithms. In fact, the experimental results demonstrate that the proposed LRCNN for change detection not only recognizes different types of changed/unchanged data, but also ensures noise insensitivity without losing details in changed areas. Fang Liu 0034, Licheng Jiao, Xu Tang 0004, Shuyuan Yang 0001, Wenping Ma 0001, Biao Hou |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2018 | PolSAR Image Classification Based on DBN and Tensor Dimensionality ReductionabstractThis paper proposes a new semi-supervised PolSAR image classification method using deep belief network (DBN) and tensor dimensionality reduction, which uses multilinear principle component analysis (MPCA) to reduce the dimension of tensor form PolSAR data, and regards the multiple features of PolSAR data as the input of DBN. In order to take full advantage of neighborhood information of each pixel of PolSAR data, we take each pixel and its neighborhood as tensor form. For PolSAR data, simple feature has been proven not to be able to effectively classify complex terrains. Therefore, we combine multiple features of PolSAR data to obtain more abundant information, which can reflect some spatial structure of PolSAR data. The experimental results show that the overall classification accuracy based on the proposed method outperforms the traditional classification strategies. Biao Hou, Xianpeng Guo, Weidan Hou, Shuang Wang 0001, Xiangrong Zhang, Licheng Jiao |
IGARSS | 1 |
| 2018 | Fully Convolutional Semi-Supervised Gan for Polsar ClassificationabstractWe propose a novel semi -supervised fully convolutional network for Polarimetric synthetic aperture radar (PoISAR) terrain classification. First, by designing a fully convolutional structure, we can perform pixel-based classification tasks. Then, by applying semi -supervised generative adversarial networks (GANs), we utilize both labeled and unlabeled samples and aim to obtain higher classification accuracy. Through a mini-max two-player game, GAN has better performance than other “single-player” classifiers. Finally, we combine the fully convolutional structure with the semi-supervised GAN. Our fully convolutional semi-supervised GAN (FC-SGAN) has excellent spatial feature learning ability and can perform end-to-end pixel-based classification tasks. Experimental results show that compared with existing works, the proposed method has better performances. Even when the training set gets smaller, our method keeps high accuracy. Mengchen Liu, Shuang Wang 0001, Yanhe Guo, Biao Hou, Licheng Jiao, Xiaojin Hou |
IGARSS | 5 |
| 2018 | Hyper-Laplacian Regularized Low-Rank Tensor Decomposition for Hyperspectral Anomaly DetectionabstractThis paper presents a novel method for hyperspectral anomaly detection considering the spectral redundancy and exploiting spectral-spatial information at the same time. We proposed a Hyper-Laplacian regularized low-rank tensor decomposition method combing with dimensionality reduction framework. Firstly, k-means++ algorithm is implemented to spectral bands and centers of each group are selected to reduce the HSI dimensionality in spectral direction. To jointly utilize spectral-spatial information, the cubic data (two spatial dimensions and one spectral dimension) is treated as a 3-order tensor. Then the non-local self-similarity is fully explored in our method. For the reason to reduce the ringing artifacts caused by over-lapped segmentation in exploring the non-local self-similarity, we introduce the hyper-Laplacian constrained low-rank tensor decomposition and we get the separated background and residual parts. Finally, to eliminate the effect of Gaussian noise, we use local-Rx basic detector to detect the residual matrix. Experimental results on two real hyperspectral data sets verified the effectiveness of the proposed algorithms for HSI anomaly detection. Xiaoxiao Ma 0003, Xiangrong Zhang, Ning Huyan, Xu Tang 0004, Biao Hou, Licheng Jiao |
IGARSS | 5 |
| 2018 | Deep Generative Matching Network for Optical and SAR Image RegistrationabstractMultimodal remote sensing images contain complementary information, thus, could potentially benefit many remote sensing applications. To this end, the image registration is a common requirement for utilizing the multimodal images. However, due to the rather different imaging mechanisms, multimodal image registration becomes much more challenging than ordinary registration, particular for optical and synthetic aperture radar (SAR) images. In this work, we design a deep matching network to exploit the latent and coherent features between multimodal patch pairs for inferring their matching labels. But, the network requires immense data for training, which is not usually met. To address this issue, we propose a generative matching network (GMN) to generate the coupled optical and SAR images, hence, improve the quantity and diversity of the training data. The experimental results show that our proposal significantly improves the registration performance of optical and SAR image registration, and achieves subpixel or close to subpixel error. Dou Quan, Shuang Wang 0001, Xuefeng Liang, Ruojing Wang, Shuai Fang, Biao Hou, Licheng Jiao |
IGARSS | 6 |
| 2018 | Decomposition-Feature-Iterative-Clustering-Based Superpixel Segmentation for PolSAR Image ClassificationabstractCompared with traditional pixel-based polarimetric synthetic aperture radar (PolSAR) image classification methods, superpixel-based methods take advantages of the spatial information of pixels, so they can overcome the influence of speckle noise on the classification result. Since traditional superpixel methods do not utilize the scattering characteristics of a PolSAR image, the boundaries of the superpixels are poorly preserved. The inaccuracy of superpixel segmentation boundaries has a negative impact on the subsequent classification. In this letter, we propose a decomposition-feature-iterative-clustering (DFIC) superpixel segmentation method for PolSAR images. The DFIC method innovatively introduces the decomposition features in generating superpixels, so the superpixel segmentation boundaries are well preserved. Because we selectively utilize superpixel information to classify the PolSAR images by setting a threshold, the effect of superpixel segmentation inaccuracy on the classification results is reduced. Experiments on two real PolSAR images demonstrate that the proposed method outperforms several state-of-the-art superpixel methods, and that the DFIC superpixel-based classification obtains better results than the other pixel-based methods. Biao Hou, Chen Yang 0027, Bo Ren 0001, Licheng Jiao |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2018 | Discriminative Feature Learning for Real-Time SAR Automatic Target Recognition With the Nonlinear Analysis Cosparse ModelabstractThis letter presents an efficient application of the nonlinear analysis cosparse model (NACM) to the task of real-time synthetic aperture radar automatic target recognition (ATR). In contrast to the conventional synthesis sparse representation model, NACM enables efficient sparse feature extraction and selection using a feed-forward mechanism. Furthermore, NACM does not require a sparsity-inducing regularizer. This model uses a task-driven learning framework, in which a naive Bayes or a discriminative classifier is adaptively learned along with the regularized features. Experimental results with the moving and stationary target acquisition and recognition benchmark demonstrate the effectiveness and efficiency of our proposed approach. Compared with traditional classification algorithms using sparse representation, our approach not only achieves higher or comparable recognition accuracy but also dramatically reduces the execution time for real-time ATR. Zaidao Wen, Biao Hou, Licheng Jiao |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2018 | Discriminative Transformation Learning for Fuzzy Sparse Subspace ClusteringabstractThis paper develops a novel iterative framework for subspace clustering (SC) in a learned discriminative feature domain. This framework consists of two modules of fuzzy sparse SC and discriminative transformation learning. In the first module, fuzzy latent labels containing discriminative information and latent representations capturing the subspace structure will be simultaneously evaluated in a feature domain. Then the linear transforming operator with respect to the feature domain will be successively updated in the second module with the advantages of more discrimination, subspace structure preservation, and robustness to outliers. These two modules will be alternatively carried out and both theoretical analysis and empirical evaluations will demonstrate its effectiveness and superiorities. In particular, experimental results on three benchmark databases for SC clearly illustrate that the proposed framework can achieve significant improvements than other state-of-the-art approaches in terms of clustering accuracy. Zaidao Wen, Biao Hou, Licheng Jiao |
IEEE Trans. Cybern. | 2 |
| 2018 | Target-Oriented High-Resolution SAR Image Formation via Semantic Information Guided RegularizationsabstractSparsity-regularized synthetic aperture radar (SAR) imaging framework has shown its remarkable performance to generate a feature-enhanced high-resolution image, in which a sparsity-inducing regularizer is involved by exploiting the sparsity priors of some visual features in the underlying image. However, since the simple prior of low-level features is insufficient to describe different semantic contents in the image, this type of regularizer will be incapable of distinguishing between the target of interest and unconcerned background clutters. As a consequence, the features belonging to the target and clutters are simultaneously affected in the generated image without concerning their underlying semantic labels. To address this problem, we propose a novel semantic information guided generative framework for target-oriented SAR image formation, which aims at enhancing the interested target scatters while suppressing the background clutters. First, we develop a new semantics-specific regularizer for image formation by exploiting the statistical properties of different semantic categories in a target scene SAR image. In order to infer the semantic label for each pixel in an unsupervised way, we moreover induce a novel high-level prior-driven regularizer and some semantic causal rules from the prior knowledge. Finally, our regularized framework for image formation is further derived as a simple iteratively reweighted $\ell _{1}$ minimization problem that can be conveniently solved by many off-the-shelf solvers. Experimental results demonstrate the effectiveness and superiority of our framework for SAR image formation in terms of target enhancement and clutters suppression, compared with the state of the arts. Additionally, the proposed framework opens a new direction of devoting some machine learning strategies to image formation, which can benefit the subsequent decision-making tasks. Biao Hou, Zaidao Wen, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2018 | Adaptive Super-Resolution for Remote Sensing Images Based on Sparse Representation With Global Joint Dictionary ModelabstractSparse representation has been widely used in the field of remote sensing image super-resolution (SR) to restore a high-quality image from a low-resolution (LR) image, e.g., from the blurred and downsampled version of an LR image's high-resolution (HR) counterpart. It is well known that each image patch can be represented by a linear combination of the atoms of an overcomplete dictionary, and we can obtain an expression of sparse coefficients by l1norm regularization. Owing to the lack of an inner relationship between image patches and an image's global information, the traditional methods of jointly training two overcomplete dictionaries cannot obtain good SR results. Therefore, we propose an effective approach for remote sensing image SR based on sparse representation. More specifically, a novel global joint dictionary model (GJDM) is used to explore the prior knowledge of images, including local and global characteristics. First, we train two dictionaries for detail image patches and HR patches. Second, in order to enhance the inner relationship between image patches, we introduce a global self-compatibility model for global regularization. Finally, the sparse representation and the local and nonlocal constraints are integrated to improve the performance of the model, and the fast adaptive shrinkage-thresholding algorithm is employed to solve the convex optimization problem in the GJDM. Compared with other methods, the results of the proposed method show good SR performance in preserving details and texture information and significant improvement in a peak signal-to-noise ratio. Biao Hou, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2017 | Fast graph-based SAR image segmentation via simple superpixelsabstractGraph-based methods have been successfully applied in the field of computer vision for image segmentation. Unfortunately, most of them are not suitable to deal with large-scale SAR image segmentation due to their high computation complexity. A fast and efficient graph-based SAR image segmentation is proposed in this paper through using superpixels to reduce the computation complexity. Firstly, a SAR image is divided into several non-overlapped subdivisions with the same size. Each of the subdivision is processed as a single OpenMP parallel region, which can be processed at the single computing node with multi-core CPU. Secondly, the number of nodes and edges in the graph is reduced by extracting the superpixels other than single pixels based on global information of each subdivision. Finally, an effective rule is proposed to merge two adjacent sub-graphs from two different subdivisions into a new subgraph. Biao Hou, Dezhao Gong, Shuang Wang 0001, Xiangrong Zhang, Licheng Jiao |
IGARSS | 1 |
| 2017 | Change detection in synthetic aperture radar images based on log-mean operator and stacked auto-encoderabstractIn this paper, we recommend a novel method based on log-mean operator and stacked auto-encoder which is used in the change detection for synthetic aperture radar images. The approach detects the changed and unchanged areas by designing a stacked Auto-encoder. The main guideline is to produce a difference image (DI) through the log-mean operator, and then distinguish the changed and unchanged regions with the trained stacked auto-encoder. The log-mean operator can roughly classify the changed and unchanged regions, but there also some inaccuracy in the difference image. Then, the stacked Auto-encoder can repair the different image further. Experimental results compared with other methods show that the method is effective. Yangyang Li 0001, Linhao Zhou, Gao Lu, Biao Hou, Licheng Jiao |
IGARSS | 4 |
| 2017 | Natural language description of remote sensing images based on deep learningabstractThe semantic description of remote sensing image is a useful and meaningful task, which can help us to get a better understanding of the scene depicted in the remote sensing images and make better use of the remote sensing images. Nature language provides good solution for describing the semantic information of remote sensing images. Nature language description of a remote sensing image is to generate a meaningful sentence given a remote sensing image. This paper presents a novel method based on deep learning. First, a convolutional neural network is utilized to detect the main objects of the remote sensing images. Then a recurrent neural network language model is utilized to generate the natural language descriptions of the objects which are detected in the first step. Experimental results on a set of remote sensing images demonstrate that the proposed method is able to generate desirable description of the scene. Xiangrong Zhang, Xiang Li 0013, Jinliang An, Biao Hou, Chen Li 0011 |
IGARSS | 5 |
| 2017 | Unsupervised saliency-guided SAR image change detection
Yaoguo Zheng, Licheng Jiao, Hongying Liu 0001, Xiangrong Zhang, Biao Hou, Shuang Wang 0001 |
Pattern Recognit. | 5 |
| 2017 | Joint Sparse Recovery With Semisupervised MUSICabstractDiscrete multiple signal classification (MUSIC) with its low computational cost and mild condition requirement becomes a significant noniterative algorithm for joint sparse recovery (JSR). However, it fails in rank defective problem caused by coherent or limited amount of multiple measurement vectors (MMVs). In this letter, we provide a novel sight to address this problem by interpreting JSR as a binary classification problem with respect to atoms. Meanwhile, MUSIC essentially constructs a supervised classifier based on the labeled MMVs so that its performance will heavily depend on the quality and quantity of these training samples. From this viewpoint, we develop a semisupervised MUSIC (SS-MUSIC) in the spirit of machine learning, which declares that the insufficient supervised information in the training samples can be compensated from those unlabeled atoms. Instead of constructing a classifier in a fully supervised manner, we iteratively refine a semisupervised classifier by exploiting the labeled MMVs and some reliable unlabeled atoms simultaneously. Through this way, the required conditions and iterations can be greatly relaxed and reduced. Numerical experimental results demonstrate that SS-MUSIC can achieve much better recovery performances than other MUSIC extended algorithms as well as some typical greedy algorithms for JSR in terms of iterations and recovery probability. Zaidao Wen, Biao Hou, Licheng Jiao |
IEEE Signal Process. Lett. | 2 |
| 2017 | Robust Semisupervised Classification for PolSAR Image With Noisy LabelsabstractThe robustness of the supervised polarimetric synthetic aperture radar (PolSAR) image classification is severely affected by two main aspects, namely, the quantity and quality of the labeled training pixels. Specifically, limited manually labeled pixels with respect to the large scale of PolSAR image have limited the performance of the automatic classification methods, while manually labeled training pixels shall be unfaithful with the speckle and impure cell for their low qualities. In order to address the above two fundamental problems, we propose a robust semisupervised probability graphic-based classification framework. First, a semisupervised learning scheme is implemented to simultaneously exploit both labeled and unlabeled pixels for information compensation. Moreover, structural relationship among neighboring pixels inducing from the prior information is further benefit to reduce the influence of limited labeled pixels. Second, a robust classification loss function is added in the process of training classifier to enhance the robustness to the noisy labeled pixels. Third, unfaithful limited labeled data can be settled with a hybrid generative/discriminative classification framework, where labeled and unlabeled pixels are simultaneously exploited for learning high-level feature for the low-quality pixels. The effectiveness of the proposed framework on the specific aspect is validated in experiments on real PolSAR data sets, which reveal the superiority in both visual performance and classification accuracy compared with the state-of-the-art methods. Totally speaking, our model has improved the classification accuracy by at least 20% on data set Flevoland, 10% on Oberpfaffenhofen, and 5% on Weihe River than the compared ones. Biao Hou, Zaidao Wen, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2017 | Discriminative Nonlinear Analysis Operator Learning: When Cosparse Model Meets Image ClassificationabstractA linear synthesis model-based dictionary learning framework has achieved remarkable performances in image classification in the last decade. Behaved as a generative feature model, it, however, suffers from some intrinsic deficiencies. In this paper, we propose a novel parametric nonlinear analysis cosparse model (NACM) with which a unique feature vector will be much more efficiently extracted. Additionally, we derive a deep insight to demonstrate that NACM is capable of simultaneously learning the task-adapted feature transformation and regularization to encode our preferences, domain prior knowledge, and task-oriented supervised information into the features. The proposed NACM is devoted to the classification task as a discriminative feature model and yield a novel discriminative nonlinear analysis operator learning framework (DNAOL). The theoretical analysis and experimental performances clearly demonstrate that DNAOL will not only achieve the better or at least competitive classification accuracies than the state-of-the-art algorithms, but it can also dramatically reduce the time complexities in both training and testing phases. Zaidao Wen, Biao Hou, Licheng Jiao |
IEEE Trans. Image Process. | 2 |
| 2016 | Unsupervised PolSAR image classification using boundary-preserving region division and region-based affinity propagation clusteringabstractThis paper presents a new method for polarimetric synthetic aperture radar (PolSAR) image classification. Firstly, to get a reasonable edge strength map, polarimetric information is used in edge strength calculation, and watershed algorithm is used to obtain the oversegmentation using the edge strength. Secondly, a searching table is used to determine the most suitable region to be merged. Finally, region-based affinity propagation clustering is employed to achieve an initial classification map, and the method provides an adjacent Wishart classifier with spatial relations to obtain the final classification result. Biao Hou, Yuheng Jiang, Bo Ren 0001, Zaidao Wen, Shuang Wang 0001, Licheng Jiao |
IGARSS | 1 |
| 2016 | Terrain classification with Polarimetric SAR based on Deep Sparse Filtering NetworkabstractA new method for Polarimetric Synthetic Aperture Radar (PolSAR) terrain classification based on Deep Sparse Filtering Network (DSFN) is proposed in this paper. It uses a novel deep learning network to learn features from the input raw data automatically. And the spatial information between pixels on PolSAR image is combined into the input data. Moreover, unlike the conventional deep networks, the DSFN only needs to tune very few parameters during pre-training and fine-tuning. A real PolSAR data is used to verify the proposed method. Experimental results show that the proposed DSFN is efficient with less parameters and effectively improves the classification accuracy compared with conventional deep networks. Hongying Liu 0001, Qiang Min, Jin Zhao 0002, Shuyuan Yang 0001, Biao Hou, Jie Feng 0003, Licheng Jiao |
IGARSS | 6 |
| 2016 | Fast semi-supervised classification based on parallel auction graph for polarimetric SAR dataabstractAlthough the graph-based machine learning has received considerable attention in the remote sensing area and it has been widely used for terrain classification, the construction of graph in most existing algorithms still takes large memory and plenty of computational time especially for large Polarimetric Synthetic Aperture Radar (PolSAR) data. Addressing these issues, we propose a fast semi-supervised classification method based on parallel auction graph in this paper. The spatial relation between pixels is firstly preprocessed using the superpixel segmentation. Then we divide the PolSAR data into multiple groups, and each of them is used to construct a sparse auction graph. The semi-supervised classification is performed parallel on those graphs. Experimental results on simulated and real PolSAR data demonstrate its efficiency and effectiveness compared with existing methods. Hongying Liu 0001, Xing Xing, Shigang Wang 0001, Zhixi Feng, Erlei Zhang, Shuyuan Yang 0001, Biao Hou, Licheng Jiao |
IGARSS | 7 |
| 2016 | Learning task-driven polarimetric target decomposition: A new perspectiveabstractPolarimetric target decomposition aims to decompose a polarimetric synthetic aperture (PolSAR) data on a base reflecting some scattering mechanisms. The corresponding coefficients will be further exploited as the feature vector for the subsequent interpretation task. Intuitively, its performance heavily depends on the choice of bases and many off-the-shelf ones have been constructed based on mathematical or physical model since last two decades. However, these fixed bases are generally insufficient to characterize all types of data in a PolSAR image so that the extracted features are not beneficial to the subsequent task. To address this issue, we propose a novel target decomposition framework to learn a set of task-desired bases as well as feature vectors from the input polarimetric data. Focusing on the classification task, involve a supervised regularizer is further involved in our framework to increase the discrimination of features. Experimental results demonstrate the effectiveness of proposed framework. Zaidao Wen, Biao Hou, Shuang Wang 0001, Licheng Jiao |
IGARSS | 2 |
| 2016 | Joint multi-feature hyperspectral image classification with spatial constraint in semantic manifoldabstractThis paper presents a novel method for hyperspectral classification combining multiple features and exploiting spatial information at the same time. We proposed a supervised classification method under the Markov random field (MRF)-based framework. Firstly using the probability SVM to map multiple features from different low-level subspace to the same semantic space (probability space), then integrating these features in semantic space with MRF-based model to enforce a smooth and accurate representation, in addition the manifold distance has been used in MRF-based model to measure the similarity of two point. To further improve the classification accuracy, a new approach of building the adaptive neighborhood has been proposed and used in our method. As our model is a derivable and convex problem, gradient descent can be used to solve this problem with less computational and time cost. Experimental results on real hyperspectral dataset shows that the proposed method provides improved classification accuracy in terms of the overall accuracy, average accuracy and kappa statistic. Xiangrong Zhang, Zeyu Gao 0001, Jinliang An, Yanning Hu, Yangyang Li 0001, Biao Hou |
IGARSS | 6 |
| 2016 | Weighted multifeature hyperspectral image classification via kernel joint sparse representation
Erlei Zhang, Xiangrong Zhang, Licheng Jiao, Hongying Liu 0001, Shuang Wang 0001, Biao Hou |
Neurocomputing | 6 |
| 2016 | Locality-constraint discriminant feature learning for high-resolution SAR image classification
Licheng Jiao, Biao Hou, Shuang Wang 0001, Jiaqi Zhao 0001, Puhua Chen |
Neurocomputing | 3 |
| 2016 | SAR Image Classification via Hierarchical Sparse Representation and Multisize Patch FeaturesabstractIn this letter, a novel hierarchical sparse representation-based classification (HSRC) for synthetic aperture radar (SAR) images is proposed. Features utilized in HSRC are extracted from the multisize patches around each pixel to precisely describe the complex terrains. Two thresholds are introduced in the sparse representation classifier to restrict the range of reconstruction residual, which classifies the reliable classified points, and the rest of the pixels are considered as the uncertain ones in the original SAR image. Then, a new dictionary is constructed by the reliable pixels, and the uncertain pixels will be reclassified in the next classification layer. The hierarchical structure is very reasonable and effective to employ simple features in each layer for describing the various topographic types. Compared with traditional sparse representation-based classification and support vector machines in several fixed-size patches, the proposed method can obtain better performance both in quantitative evaluation and visualization results. Biao Hou, Bo Ren 0001, Guilin Ju, Huiyan Li, Licheng Jiao, Jin Zhao 0002 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2016 | Local Collaborative Representation With Adaptive Dictionary Selection for Hyperspectral Image ClassificationabstractSpectral-spatial representation based algorithms have been widely applied in hyperspectral image (HSI) classification, which exploit the fact that pixels in a local patch often have similar spectral reflectance values and probably belong to the same class. Collaborative representation (CR) is a typical supervised classification method for high-dimensional data, which has been widely used for spectral-spatial representation based HSI classification. However, it suffers from the degraded representation of redundant and irrelative pixels when all of the labeled pixels are used as a dictionary for representation. In this letter, a novel method, local CR with adaptive dictionary selection, is proposed to solve this problem, in which we first average the values of pixels from local patches to incorporate the contextual information of neighbors, and then, an adaptive dictionary selection method is presented to select the most similar pixels to each test pixel from the dictionary to reduce the influence of redundant and irrelevant pixels in representation. Experimental results on two HSIs show that the proposed method outperforms some spectral-spatial representation based algorithms in terms of classification accuracy. Yaoguo Zheng, Licheng Jiao, Ronghua Shang, Biao Hou, Xiangrong Zhang |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2016 | SAR Image Registration Based on Multifeature Detection and Arborescence Network MatchingabstractIn this letter, a novel synthetic aperture radar (SAR) image registration method, including two operators for feature detection and arborescence network matching (ANM) for feature matching, is proposed. The two operators, namely, SAR scale-invariant feature transform (SIFT) and R-SIFT, can detect corner points and texture points in SAR images, respectively. This process has an advantage of preserving two types of feature information in SAR images simultaneously. The ANM algorithm has a two-stage process for finding matching pairs. The backbone network and the branch network are successively built. This ANM algorithm combines feature constraints with spatial relations among feature points and possesses a larger number of matching pairs and higher subpixel matching precision than the original version. Experimental results on various SAR images show that the proposed method provides superior performance than other approaches investigated. Hao Zhu 0009, Wenping Ma 0001, Biao Hou, Licheng Jiao |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2016 | Spectral-spatial hyperspectral image ensemble classification via joint sparse representation
Erlei Zhang, Xiangrong Zhang, Licheng Jiao, Lin Li 0016, Biao Hou |
Pattern Recognit. | 5 |
| 2016 | POL-SAR Image Classification Based on Wishart DBN and Local Spatial InformationabstractInspired by a popular deep neural network, i.e., deep belief network (DBN), a novel method for polarimetric synthetic aperture radar (POL-SAR) image classification is proposed in this paper. For the particularity of POL-SAR data, a new type of restricted Boltzmann machine (RBM) is specially defined, which we name the Wishart-Bernoulli RBM (WBRBM), and is used to form a deep network named as Wishart DBN (W-DBN). Numerous unlabeled POL-SAR pixels are made full use of in the modeling of POL-SAR pixels by W-DBN. In addition, the coherency matrix is used directly to represent a POL-SAR pixel without any manual feature extraction, which is simple and time saving. Local spatial information, together with the confusion matrix, is used in this paper to clean the preliminary classification result obtained by the method based on W-DBN. Making full use of the prior knowledge of POL-SAR data and local spatial information, the proposed method overcomes shortcomings of traditional methods, in which they are sensitive to extracted features and slow to execute. The experiments, tested on three POL-SAR data sets, show that the proposed method produces better results and is much faster than traditional methods. Fang Liu 0034, Licheng Jiao, Biao Hou, Shuyuan Yang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2015 | Wishart RBM based DBN for polarimetric synthetic radar data classificationabstractDeep Belief Network (DBN) is a classic deep learning model, and it can learn higher feature and do better classification job. We combine DBN's basic component Restricted Boltzmann Machines (RBM) with the statistic distribution of Polarimetric SAR (PolSAR) data. Based on it, we develop a deep learning classification method that is suitable for PolSAR data. To verify the effectiveness of the method, a real PolSAR dataset is tested. Experiment result confirms that the proposed method provides fine improvements both in classification accuracy and visual effect. Yanhe Guo, Shuang Wang 0001, Chenqiong Gao, Danrong Shi, Biao Hou |
IGARSS | 6 |
| 2015 | Polarimetric SAR images classification using deep belief networks with learning featuresabstractA novel polarimetric synthetic aperture radar (PolSAR) image classification method based on Deep Belief Networks (DBNs) is proposed in this paper. First, the coherency matrix data are converted to a 9-dimentional data. Second, many patches are randomly selected from each dimension in the 9-dimentional data, and many filters can be obtained from a Restricted Boltzmann Machine (RBM) trained by using these patches. Thus we can get the features for each pixel from each dimension in the 9-dimentional space. Finally, the learned features and the elements of coherent matrix are combined to train a 3-layers DBNs for PolSAR image classification. Experimental results show that the proposed method is efficient and effective for PolSAR image classification. Biao Hou, Xiaohuan Luo, Shuang Wang 0001, Licheng Jiao, Xiangrong Zhang |
IGARSS | 1 |
| 2015 | Semi-supervised classification based on anchor-spatial graph for large polarimetric SAR dataabstractRecently a few works of semi-supervised learning methods based on graph have been proposed for remote sensing. The common idea of these methods are that they build a graph using the samples of the image. Most of their time complexity is relatively large, and they ignore the spatial information of the image, which leads to unsatisfactory classification results. this paper proposes a novel semi-supervised classification method based on anchor-spatial graph for large PolSAR data. Firstly the unsupervised Wishart clustering is performed to select representative samples, which served as anchors according to the least distance between samples. Then an anchor graph is built using the selected anchors according to the multiple features of the samples. And it is further combined with the spatial information of the samples to construct an anchor-spatial graph. Finally the class information from small quantities of labeled samples propagates to the unlabeled ones. Experimental results show that the proposed method has a low time complexity compared with existing works and it could effectively cut down the processing time for large PolSAR data meanwhile keeps the classification accuracy. Hongying Liu 0001, Dexiang Zhu, Shuyuan Yang 0001, Shuang Wang 0001, Biao Hou, Licheng Jiao |
IGARSS | 6 |
| 2015 | Sparsity-constrained generalized bilinear model for hyperspectral unmixingabstractGeneralized bilinear model (GBM) has been widely used for nonlinear hyperspectral image unmixing. However, it does not take the sparse information of abundance into account, which is a significant characteristic resulting from the correlation of hyperspectral data. This paper aims to extend the GBM by incorporating the sparsity constraint of abundance matrix with the semi-nonnegative matrix factorization, by dividing GBM into the linear part and the second-order part, which are optimized using an alternating optimization algorithm respectively. L1/2-norm is used to explore the sparse characteristic, and the L1/2-constrained semi-nonnegative matrix factorization (L1/2-semi-NMF) algorithm is presented, which leads to better results on both synthetic and real data. Xiangrong Zhang, Cai Cheng, Jinliang An, Yaoguo Zheng, Erlei Zhang, Biao Hou |
IGARSS | 6 |
| 2015 | Multilayer CFAR Detection of Ship Targets in Very High Resolution SAR ImagesabstractThis letter proposes a new ship target detection method for very high resolution (VHR) synthetic aperture radar (SAR) images based on multilayer constant false alarm rate (CFAR). First, combined with log-normal distribution, a multilayer CFAR method is designed to overcome the holes and the fracture in the traditional detected results. This method can retain more details of ships and takes much less time than the traditional CFAR method for VHR SAR images. Second, based on a priori knowledge of ships, we use the sliding window to remove the false alarm targets. Finally, In order to measure the size and shape of a ship, we extract the outline of a ship and fill it by a level set method. Experimental results, carried out on real SAR images, demonstrate that the proposed approach outperforms the previous one in terms of the detection ratio of pixels instead of the number of ships. Biao Hou, Xingzhong Chen, Licheng Jiao |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2015 | Spectral-Spatial Classification of Hyperspectral Data Using 3-D Morphological ProfileabstractA new spectral-spatial method based on a 3-D morphological profile (3D-MP) is proposed for hyperspectral data classification. As an extension of a previous approach, the proposed method uses both the spectral and spatial information for classification. First, random projection (RP) is used for dimensionality reduction of hyperspectral data. After RP in spectral domain, a novel 3D-MP method is proposed to exploit the dependence between data. Finally, the classification is performed by the widely used support vector machine classifier. Our experiments reveal that the proposed approach exploits the 3-D spectral-spatial feature to provide the state-of-the-art classification results for different hyperspectral data sets. Biao Hou, Taimin Huang, Licheng Jiao |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2015 | Nonconvex Compressed Sensing by Nature-Inspired Optimization AlgorithmsabstractThe l 0 regularized problem in compressed sensing reconstruction is nonconvex with NP-hard computational complexity. Methods available for such problems fall into one of two types: greedy pursuit methods and thresholding methods, which are characterized by suboptimal fast search strategies. Nature-inspired algorithms for combinatorial optimization are famous for their efficient global search strategies and superior performance for nonconvex and nonlinear problems. In this paper, we study and propose nonconvex compressed sensing for natural images by nature-inspired optimization algorithms. We get measurements by the block-based compressed sampling and introduce an overcomplete dictionary of Ridgelet for image blocks. An atom of this dictionary is identified by the parameters of direction, scale and shift. Of them, direction parameter is important for adapting to directional regularity. So we propose a two-stage reconstruction scheme (TS_RS) of nature-inspired optimization algorithms. In the first reconstruction stage, we design a genetic algorithm for a class of image blocks to acquire the estimation of atomic combinations in all directions; and in the second reconstruction stage, we adopt clonal selection algorithm to search better atomic combinations in the sub-dictionary resulted by the first stage for each image block further on scale and shift parameters. In TS_RS, to reduce the uncertainty and instability of the reconstruction problems, we adopt novel and flexible heuristic searching strategies, which include delicately designing the initialization, operators, evaluating methods, and so on. The experimental results show the efficiency and stability of the proposed TS_RS of nature-inspired algorithms, which outperforms classic greedy and thresholding methods. Fang Liu 0001, Leping Lin, Licheng Jiao, Lingling Li 0002, Shuyuan Yang 0001, Biao Hou, Hongmei Ma, Jinghuan Xu |
IEEE Trans. Cybern. | 6 |
| 2015 | A Resample-Based SVA Algorithm for Sidelobe Reduction of SAR/ISAR Imagery With Noninteger Nyquist Sampling RateabstractA resample-based spatial variant apodization (SVA) algorithm for sidelobe reduction was studied for synthetic aperture radar (SAR) and inverse SAR (ISAR) imagery with a noninteger Nyquist sampling rate. The weighting function of every sample in the image domain was calculated with the sample and two adjacent noninteger samples. The noninteger samples were obtained by interpolation in the image domain using sinc function. With the proper selection of two noninteger samples, the monotonic property of the weighting function on each side of the sampling point was preserved. The unequivocal determination of sidelobe suppression was achieved for noninteger Nyquist sampled (NINS) SAR and ISAR imagery. In addition, the lower and upper boundaries of the weighting function under the cosine-on-pedestal condition were extended for further sidelobe suppression and main lobe sharpening. The algorithm was implemented and applied to NINS imagery that is simulated. The algorithm was then assessed for acquired SAR and ISAR images. Improved results have been qualitatively and quantitatively achieved in sidelobe suppression and main lobe sharping in comparison with an existing algorithm. Shuang Wang 0001, Biao Hou, Yong Wang 0011, Hongying Liu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2014 | SAR image segmentation based on random projection and signature frameabstractThis paper proposes a new Synthetic Aperture Radar (SAR) image segmentation method based on the frame of Signature/Earth Mover's Distance (EMD). Firstly, Random Projection is used to extract features of SAR image, which has the abilities of preserving information and reducing dimensionality. Secondly, a signature is used to obtain the cluster center and the weight. Finally, by computing the distance between two signatures using Earth Mover's Distance, we can obtain the final segmentation result. The experimental results show that the proposed method is efficient and effective for SAR image segmentation. Biao Hou, Shuang Wang 0001, Xiangrong Zhang |
IGARSS | 1 |
| 2014 | MSTAR image segmentation with multi-phase level set based on probability density modelabstractRadar image segmentation is a fundamental problem in radar image interpretation. Radar images often contain a great deal of noise. Level set method, known as deformable model, is a powerful image segmentation technique. It can get accurate contours of clear-cut objects in image without noise, but has poor performance in getting contours of objects in a noisy image. In this paper, a new multi-phase level set based on probability density model is proposed. We use histogram, a non-parametric density estimation method, to describe the statistical information of each pixel in its neighborhood and the pixels in each subset in the image. The comparability between them computed by inner product function is used as the curve energy in multi-phase level set method. The statistical information is incorporated into the multi-phase level set framework, which can cope with the influence of noise on image segmentation. This new method is particularly well adapted to detection of objects of interesting in a noisy image. We illustrated the performance of the new method on MSTAR images. The experimental results show that incorporating statistical information into the multi-phase level set framework, consistent objects are obtained, and accurate and robust segmentations can be achieved. Xiaojin Hou, Shuang Wang 0001, Biao Hou |
IGARSS | 4 |
| 2014 | Nonlocal filtering for Polarimetric SAR Data based on bilateral filteringabstractIn this paper, we introduce a nonlocal filtering method for Polarimetric SAR Data based on bilateral filtering. This method is a combination of local and nonlocal filtering. In order to keep the spatial structure of image, we select similar blocks in a large area; and to adapt to the pixel similarities, the iterative bilateral filtering is applied to the selected similar blocks. To deal with polarimetric data, we propose a new similarities based on complex Wishart distance. We demonstrate the performance of the proposed algorithm by using the experimental data. The experiment shows the better performance of the proposed method. There are fewer speckle noise in homogeneous areas. And some detail of structures, i.e., edges, lines, points, and curves, are well protected. Xiaozhen Lei, Shuang Wang 0001, Kun Liu 0011, Biao Hou |
IGARSS | 4 |
| 2014 | Unsupervised classification of polarimetric SAR images integrating color featuresabstractIn conventional terrain classification for the polarimetric SAR (POLSAR) images, color features are rarely involved unless in one recent supervised work. Unlike that work, the color features are exploited for the unsupervised classification in this paper. Firstly, based on the polarimetric decomposition of the POLSAR data, the common color spaces, such as RGB, HSI, and CIELab are calculated. The color feature is quantitatively selected from these color spaces by introducing the color entropy. Then together with the spatial information, extended scattering power entropy and the copolarized ratio, the adaptive Mean-shift algorithm is used to segment the POLSAR image. Finally, the segments are merged according to the Wishart distance measurement. The experiments using AIRSAR L-band POLSAR data indicate that the proposed method has better discriminative ability for urban areas and for boundary preservation compared with existing works. Hongying Liu 0001, Shuang Wang 0001, Biao Hou, Shuyuan Yang 0001, Junfei Shi, Licheng Jiao |
IGARSS | 3 |
| 2014 | Multilayer feature learning for polarimetric synthetic radar data classificationabstractFeatures are important for polarimetric synthetic aperture radar (PolSAR) image classification. Various methods focus on extracting feature artificially. Compared with them, we have developed a method to learn feature automatically. The method is based on deep learning which can learn multilayer features. In this paper, stacked sparse autoencoder (SAE) as one of the deep learning models is applied as a useful strategy to achieve the goal. For improving the classification result, we use a small amount of labels to fine-tuning the parameters of the proposed method. Finally, a real PolSAR dataset is used to verify the effectiveness. Experiment result confirms that the proposed method provides noteworthy improvements in classification accuracy and visual effect. Huiming Xie, Shuang Wang 0001, Kun Liu 0011, Chris S. Lin, Biao Hou |
IGARSS | 5 |
| 2014 | Classification of imbalanced hyperspectral imagery data using support vector samplingabstractDue to the imbalance in obtaining labeled samples for different land-cover classes, hyperspectral image classification encounters the issue of imbalanced classification. In this paper, a novel and effective method is proposed to address the imbalanced learning problem in hyperspectral image classification, which combines support vector machine (SVM) and sampling strategy. The main novelty and contribution of our paper are that we propose to do sampling referring to the support vectors (SVs) rather than the training data to provide a balanced distribution during the model learning. Sampling among the training data may be time consuming, while sampling referring to the SVs is more efficient and representative with much lower complexity. Therefore, the proposed method is expected to be simple and effective for imbalanced learning problem. Experimental results on real hyperspectral image dataset show that our method can effectively improve the classification accuracy for the minority classes in the imbalanced dataset. Xiangrong Zhang, Yaoguo Zheng, Biao Hou, Shuiping Gou |
IGARSS | 4 |
| 2014 | Using Combined Difference Image and k-Means Clustering for SAR Image Change DetectionabstractIn this letter, a simple and effective unsupervised approach based on the combined difference image and$k$-means clustering is proposed for the synthetic aperture radar (SAR) image change detection task. First, we use one of the most popular denoising methods, the probabilistic-patch-based algorithm, for speckle noise reduction of the two multitemporal SAR images, and the subtraction operator and the log ratio operator are applied to generate two kinds of simple change maps. Then, the mean filter and the median filter are used to the two change maps, respectively, where the mean filter focuses on making the change map smooth and the local area consistent, and the median filter is used to preserve the edge information. Second, a simple combination framework which uses the maps obtained by the mean filter and the median filter is proposed to generate a better change map. Finally, the$k$-means clustering algorithm with$k = 2$is used to cluster it into two classes, changed area and unchanged area. Local consistency and edge information of the difference image are considered in this method. Experimental results obtained on four real SAR image data sets confirm the effectiveness of the proposed approach. Yaoguo Zheng, Xiangrong Zhang, Biao Hou, Ganchao Liu |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2014 | A Novel Eye Localization Method With Rotation InvarianceabstractThis paper presents a novel learning method for precise eye localization, a challenge to be solved in order to improve the performance of face processing algorithms. Few existing approaches can directly detect and localize eyes with arbitrary angels in predicted eye regions, face images, and original portraits at the same time. To preserve rotation invariant property throughout the entire eye localization framework, a codebook of invariant local features is proposed for the representation of eye patterns. A heat map is then generated by integrating a 2-class sparse representation classifier with a pyramid-like detecting and locating strategy to fulfill the task of discriminative classification and precise localization. Furthermore, a series of prior information is adopted to improve the localization precision and accuracy. Experimental results on three different databases show that our method is capable of effectively locating eyes in arbitrary rotation situations (360° in plane). Yan Ren 0002, Shuang Wang 0001, Biao Hou |
IEEE Trans. Image Process. | 3 |
| 2013 | SAR image ship detection based on visual attention modelabstractThis paper proposes a novel Synthetic Aperture Radar (SAR) image ship detection method based on human visual attention mechanism. Firstly, we obtain water segmentation image by combining the bottom-up and the top-down visual attention mechanisms. Secondly, we detect ship targets based on bottom-up the visual attention mechanism. The interested regions are extracted by measuring the visual conspicuity of each water regions. Then, the ships targets are detected in the interested regions by the k-means clustering algorithm. Finally, real SAR image is used to test our algorithm. Besides, we analysis the ship detection results using different band. The experiment results indicate that our algorithm can effectively detect ship targets from SAR images and C-band is superior to L-band in SAR image ship detection. Biao Hou, Shuang Wang 0001, Xiaojin Hou |
IGARSS | 1 |
| 2013 | Unsupervised classification of POLSAR data based on the improved affinity propagation clusteringabstractIn this paper, the AP clustering algorithm is improved by defining a new similarity to be applied in the polarimetric SAR image classification. On this basis, a new unsupervised classification method is proposed which combines the Four-component decomposition and the improved AP clustering. The proposed method mainly consists of three steps: Firstly, Four-component decomposition is adopted to produce initial segmentation. Secondly, the improved affinity propagation clustering based on the Wishart distance measure is applied on the initial segmentation to merge clusters and obtain an appropriate number of categories. Finally, an iterative algorithm based on the complex wishart density function is applied. The effectiveness of this algorithm is demonstrated by the test with NASA/JPL AIRSAR L-band data of San Francisco and Flevoland. Shuang Wang 0001, Yachao Liu, Kun Liu 0011, Xiaojin Hou, Biao Hou |
IGARSS | 5 |
| 2013 | High resolution SAR target reconstruction from compressive measurements with prior knowledgeabstractIn this paper, an effective prior knowledge based framework for target reconstruction from compressive measurements is proposed. In this framework, a traditional compressed imaging method is firstly introduced which indicates that for a range cell containing K strongest scattering points can be reconstructed based on the theory of compressive sensing. Secondly, a greedy iteration algorithm is modified which utilizes some prior knowledge of the target during the reconstruction step. The experiments are carried on the Moving and Stationary Target Acquisition and Recognition (MSTAR) database and the results show the effectiveness of our framework for target reconstruction. Zaidao Wen, Biao Hou, Shuang Wang 0001 |
IGARSS | 2 |
| 2013 | Joint segmentation and classification of hyperspectral image using meanshift and sparse representation classifierabstractA novel spectral-spatial classification method based on mean shift and sparse representation classifier (SRC) for hyperspectral images is proposed in this paper. Firstly, the nonnegative matrix factorization, is used as a preprocessing for mean shift. Then, the mean shift algorithm is adopted to partition an image into amount of blocks and get the segmentation map. Through this way, many size-variable and close regions can be got while the boundary information is remained. Secondly, the classification map is obtained by using the SRC. Finally, the fusion of the segmentation map and the classification map is done by using the majority vote rule. Experimental results on two real hyperspectral images demonstrate the effectiveness and good performance of the proposed method. Xiangrong Zhang, Yaoguo Zheng, Biao Hou, Xiaojin Hou |
IGARSS | 4 |
| 2013 | Spatial-spectral classification based on group sparse coding for hyperspectral imageabstractIn this paper, a novel hyperspectral image classification method is proposed, based on group sparse coding. The method is based on this acknowledgement that larger spatial variation exists in high spatial resolution hyperspectral image, which degrades the separability of hyperspectral image. In order to obtain a smooth representation, each pixel and its spatial neighbors are coded together by group sparse coding. Although nothing about class information is included, the neighbor pixels in a small spatial window are inclined to belong to the same class. Thus, that will reduce the within-class scatter and be favorable to the classification task. Then, the obtained sparse representation vectors are used for hyperspectral image classification with SVM. Experimental results show that our method exceeds the classical classification algorithms in accuracy and regional consistency. Xiangrong Zhang, Peng Weng, Jie Feng 0003, Erlei Zhang, Biao Hou |
IGARSS | 5 |
| 2013 | Novel Change Detection in SAR Imagery Using Local ConnectivityabstractMost change-detection techniques in synthetic aperture radar (SAR) imagery are based on the analysis of the difference image with a pixel-level decision approach. However, the pixel-level decision approach would cause a noisy change-detection map, with holes in connected regions and jagged boundaries. In this letter, we propose a novel change-detection method to deal with the problem of the pixel-level decision approach by considering local connectivity. We first get an initial change-detection result with an improved Gustafson–Kessel clustering algorithm using local spatial information and then refine the initial result through region-of-interest extraction and consideration of local connectivity of changed areas. Experimental results on real SAR image data sets demonstrate that the proposed method outperforms the related ones for change detection. Honglin Wan, C. Jung, Biao Hou, Guiting Wang, Q. X. Tang |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2013 | Fast Fisher Sparsity Preserving Projections
Licheng Jiao, Fanhua Shang, Shuang Wang 0001, Biao Hou |
Neural Comput. Appl. | 5 |
| 2013 | Context-Based Hierarchical Unequal Merging for SAR Image SegmentationabstractThis paper presents an image segmentation method named Context-based Hierarchical Unequal Merging for Synthetic aperture radar (SAR) Image Segmentation (CHUMSIS), which uses superpixels as the operation units instead of pixels. Based on the Gestalt laws, three rules that realize a new and natural way to manage different kinds of features extracted from SAR images are proposed to represent superpixel context. The rules are prior knowledge from cognitive science and serve as top-down constraints to globally guide the superpixel merging. The features, including brightness, texture, edges, and spatial information, locally describe the superpixels of SAR images and are bottom-up forces. While merging superpixels, a hierarchical unequal merging algorithm is designed, which includes two stages: 1) coarse merging stage and 2) fine merging stage. The merging algorithm unequally allocates computation resources so as to spend less running time in the superpixels without ambiguity and more running time in the superpixels with ambiguity. Experiments on synthetic and real SAR images indicate that this algorithm can make a balance between computation speed and segmentation accuracy. Compared with two state-of-the-art Markov random field models, CHUMSIS can obtain good segmentation results and successfully reduce running time. Xiangrong Zhang, Shuang Wang 0001, Biao Hou |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2012 | A visual attention model based on wavelet transform and its application on ship detectionabstractHuman visual system is very efficient and selective in scene analysis, which has been widely used in image processing. In this paper, a new visual attention model based on dyadic wavelet transform (DWT) used for ship detection is proposed. It is a bottom-up visual attention model driven by data rather than by task. First, the input image is converted from RGB color space to HIS color space. Second, the modulus of DWT is analyzed to obtain the conspicuity map of each feature. Third, the conspicuity maps are combined into the saliency map nonlinearly, different from Itti's method, the contribution rate of each conspicuity map to final saliency map is not equal. It is relevant to the difference between the level of the most active region and the average level of the other active regions in each conspicuity map. Finally, the detection result of ships based on saliency map is got by region growing method, where the seed is obtained from the saliency map and the growing process is implemented in intensity image. Experiments on natural ship images show that our method is robust and efficient compared with Itti's and Hou's method. Biao Hou, Shuang Wang 0001 |
IGARSS | 1 |
| 2012 | Low-rank and sparse matrix decomposition-based pan sharpeningabstractThis paper proposes a remote sensing image pan-sharpening method from the perspective of low-rank and sparse matrix decomposition. Based on the characteristic of multispectral (MS) images, the low spatial resolution information of MS images is modeled as low-rank, and the high spectral resolution information of MS images is modeled as sparse. First, the low-rank and sparse matrix decomposition algorithm is applied to the resampled MS images to extract the sparse component i.e. the high spectral resolution information. Second, the standard PCA fusion method is applied on the low-rank component to obtain the rough pan-sharpened MS images. Finally, adding the sparse MS images component on the rough result and one can get the final fused product. Experimental results demonstrate that the proposed method is competitive or even better than some other methods. Kaixuan Rong, Shuang Wang 0001, Biao Hou |
IGARSS | 4 |
| 2012 | SAR image despeckling method using bivariate shrinkage based on dual-tree complex waveletabstractIn this paper, we propose a speckle suppression method for SAR image based on dual-tree complex wavelet. Non-Gaussian bivariate distribution model is proposed by considering the correlation of the real and imaginary parts of complex wavelet coefficients. Based on this model, the shrinkage function of real and imaginary parts of complex coefficients is obtained with the help of maximum a posteriori estimate. Compared with the state-of-the-art techniques through the visual effect and the equivalent number of looks (ENL), experimental results demonstrate that the proposed algorithm obtains good performance in smoothing speckles of homogeneous regions and preserving edges and details effectively. Shuang Wang 0001, Biao Hou |
IGARSS | 4 |
| 2012 | MPM SAR Image Segmentation Using Feature Extraction and Context ModelabstractA new synthetic aperture radar (SAR) image segmentation method based on a maximization of posterior marginals (MPM) algorithm with feature extraction and context model is proposed in this letter. First, Gabor wavelet and texture descriptor are used to extract features, which enhance intraclass similarities and interclass differences. Second, the number of regions within the same class is reduced in order to improve the reliability of the regional statistical characteristics. Finally, the MPM of each region combined with the context model is calculated by considering both the intralayer correlation and interlayer correlation. The experimental results show that the proposed method is efficient and effective for SAR image segmentation. Biao Hou, Xiangrong Zhang |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2011 | SAR image despeckling based on improved Directionlet domain Gaussian Mixture ModelabstractIn this paper, a new SAR image despeckling method based on the improved Directionlet domain Gaussian Mixture Model (GMM) is proposed. Firstly, the cartoon texture model is used to decompose the SAR image to a cartoon part and a texture part. Secondly, the cartoon part is kept unchanged, the coefficients of the texture part in the improved Directionlet domain are modeled by the Gaussian Mixture Model. Thirdly, the Bayesian minimum mean square error estimation is used to evaluate each of coefficients. Finally, the two parts are added to obtain the despeckled image. Experimental results show that the proposed method outperforms the spatial filters and other methods based on wavelets, stationary wavelet and non-subsampled contourlets in terms of speckle reduction as well as detail and edge preservation. Biao Hou, H. Guan, J. G. Jiang, Licheng Jiao |
IGARSS | 1 |
| 2011 | Spectral clustering based unsupervised change detection in SAR imagesabstractAn unsupervised change detection method based on spectral clustering and difference image methods for multitemporal single-channel single-polarization synthetic aperture radar (SAR) images is proposed. The difference image is generated by integrating the typical difference image method with Non-Local Filter, which exploits both the spatial neighborhood information and gray similarity information, and can well reduce the speckle noises of SAR images. The spectral clustering algorithm is employed to cluster the difference image into two clusters and get the change map. Compared with traditional clustering algorithms, such as A-means, SC can recognize the clusters of unusual shapes and obtain the globally optimal solutions. Experimental results confirm the effectiveness of the proposed techniques. Xiangrong Zhang, Zemin Li, Biao Hou, Licheng Jiao |
IGARSS | 3 |
| 2011 | Shape-Adaptive Reversible Integer Lapped Transform for Lossy-to-Lossless ROI Coding of Remote Sensing Two-Dimensional ImagesabstractIn this letter, we propose a shape-adaptive (SA) reversible integer lapped transform (SA-RLT) method. The new method can deal with arbitrarily shaped image areas while guaranteeing completely reversible integer-to-integer transform. Based on SA-RLT and object-based set partitioned embedded block coder, a new region-of-interest (ROI) compression scheme is designed for 2-D remote sensing images. Numerical experiments reveal that SA-RLT performs better than integer SA discrete wavelet transform, and the new ROI compression scheme performs comparably even better than the JPEG2000-ROI scheme. Advantages in hardware implementation have been preserved by SA-RLT, such as parallel processing and low memory requirement. Licheng Jiao, Lei Wang 0018, Jiaji Wu, Jing Bai 0003, Shuang Wang 0001, Biao Hou |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2011 | SAR Image Despeckling Based on Local Homogeneous-Region Segmentation by Using Pixel-Relativity MeasurementabstractThis paper provides a novel pointwise-adaptive speckle filter based on local homogeneous-region segmentation with pixel-relativity measurement. A ratio distance is proposed to measure the distance between two speckled-image patches. The theoretical proofs indicate that the ratio distance is valid for multiplicative speckle, while the traditional Euclidean distance failed in this case. The probability density function of the ratio distance is deduced to map the distance into a relativity value. This new relativity-measurement method is free of parameter setting and more functional compared with the Gaussian kernel-projection-based ones. The new measurement method is successfully applied to segment a local shape-adaptive homogeneous region for each pixel, and a simplified strategy for the segmentation implementation is given in this paper. After segmentation, the maximum likelihood rule is introduced to estimate the true signal within every homogeneous region. A novel evaluation metric of edge-preservation degree based on ratio of average is also provided for more precise quantitative assessment. The visual and numerical experimental results show that the proposed filter outperforms the existing state-of-the-art despeckling filters. Hongxiao Feng, Biao Hou, Maoguo Gong |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2011 | Multivariate Compressive Sensing for Image Reconstruction in the Wavelet Domain: Using Scale Mixture ModelsabstractMost wavelet-based reconstruction methods of compressive sensing (CS) are developed under the independence assumption of the wavelet coefficients. However, the wavelet coefficients of images have significant statistical dependencies. Lots of multivariate prior models for the wavelet coefficients of images have been proposed and successfully applied to the image estimation problems. In this paper, the statistical structures of the wavelet coefficients are considered for CS reconstruction of images that are sparse or compressive in wavelet domain. A multivariate pursuit algorithm (MPA) based on the multivariate models is developed. Several multivariate scale mixture models are used as the prior distributions of MPA. Our method reconstructs the images by means of modeling the statistical dependencies of the wavelet coefficients in a neighborhood. The proposed algorithm based on these scale mixture models provides superior performance compared with many state-of-the-art compressive sensing reconstruction algorithms. Jiao Wu 0002, Fang Liu 0001, Licheng Jiao, Xiaodong Wang 0011, Biao Hou |
IEEE Trans. Image Process. | 5 |
| 2010 | SAR Image Despeckling Using Edge Detection and Feature Clustering in Bandelet DomainabstractTo effectively preserve the edges of a synthetic aperture radar (SAR) image when despeckling, an algorithm with edge detection and fuzzy clustering in the translation-invariant second-generation bandelet transform (TIBT) domain is proposed in this letter. A Canny operator is first utilized to detect and remove edges from the SAR image. Then, TIBT and fuzzy C-mean clustering are employed to decompose and despeckle the edge-removed image, respectively. Finally, the removed edges are added to the reconstructed image. The algorithm suggests each coefficient in high-frequency subbands as the clustering feature, proposes a calculation method of the best clustering number, and defines the signal and noise in the clustering results. Experimental results show that the visual quality and evaluation indexes outperform the other methods with no edge preservation. The proposed algorithm effectively realizes both despeckling and edge preservation and reaches the state-of-the-art performance. Wenge Zhang, Fang Liu 0001, Licheng Jiao, Biao Hou, Shuang Wang 0001, Ronghua Shang |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2005 | A Fully Unsupervised Image Segmentation Algorithm Based on Wavelet-Domain Hidden Markov Tree Models
Yuheng Sha, Xinbo Gao 0001, Biao Hou, Licheng Jiao |
ACIVS | 4 |
| 2005 | Automatic Texture Segmentation Based on Wavelet-Domain Hidden Markov Tree
Biao Hou, Licheng Jiao |
CIARP | 2 |
| 2004 | Unsupervised image segmentation based on the anisotropic texture informationabstractBased on the anisotropic characteristic of brushlet, a new feature named directional texture histogram in brushlet domain is presented, which represents the local anisotropic information in image. A novel unsupervised image segmentation method via local directional texture histogram in the brushlet domain is developed. The segmentation results of synthetic mosaics, aerial photo and synthetic aperture radar (SAR) image show that our method represents better performance and its error probability of the synthetic mosaics is lower than the wavelet-based method. Yuheng Sha, Biao Hou, Licheng Jiao |
ICIG | 3 |