Heng-Chao Li 0001

dblp:21/4241-1 · also Hengchao Li 0001 · DBLP profile ↗
← Back
131ranked-venue papers
14as first author
80since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 104 · 12 first-author · 60 since 2021Artificial intelligence and machine learning · 16 · 1 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 1 first-author · 7 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 STFAR: Test-time adaptive object detection through self-training and feature alignment regularization
Nanqing Liu, Yongyi Su, Lile Cai, Heng-Chao Li 0001, Kui Jia, Tianrui Li 0001, Xun Xu 0002, Chuan-Sheng Foo
Expert Syst. Appl.6
2026 Semisupervised graph U-Net with G-ConvLSTM for hyperspectral image classification
Jin-Yu Yang, Heng-Chao Li 0001, Xin-Ru Feng, Feng Gao 0005, Qian Du 0001, Antonio Plaza
Expert Syst. Appl.2
2026 Blueprint Multiscale Aware and Linearized Feature Enhancement Network for Efficient Remote Sensing Image Super-Resolution
abstract
Remote-Sensing image super-resolution (RSISR) technology aims to enhance the spatial resolution of low-resolution remote sensing images. Although deep learning-based super-resolution methods offer considerable theoretical advantages, their practical applications are severely limited by high memory consumption and computational costs. Existing lightweight RSISR methods typically reduce complexity by employing grouped or separated feature processing, but this can limit effective exploitation of feature interdependencies. To address this issue, this paper proposes a Blueprint Multiscale Aware and Linearized Feature Enhancement Network (BLNet). Specifically, the lightweight Blueprint MultiScale Aware Block (BMAB) is designed to extract and fuse multiscale information, tailored to the characteristics of remote sensing images. In addition, the Linearized Feature Enhancement Module (LFEM) is introduced to capture global spatial-channel features. Extensive experiments demonstrate that the proposed method significantly improves RSISR reconstruction performance while maintaining high computational efficiency. Our code will be available at https://github.com/crcherry/BLNet.
Nanqing Liu, Yun-Cheng Li, Sen Lei, Heng-Chao Li 0001
IEEE Geosci. Remote. Sens. Lett.5
2026 DNN-aided low-rank and sparse decomposition model for infrared small target detection
Jia-Jie Yin, Heng-Chao Li 0001, Yu-Bang Zheng, Xiongfei Geng
Pattern Recognit.2
2026 Unsupervised Feature Dimensionality Reduction via Latent Low-Rank Embedding Projection for Classification of Hyperspectral Images
abstract
To deal with the curse of dimensionality in hyperspectral images, numerous feature dimensionality reduction (FDR) methods have been proposed to map high-dimensional data into a low-dimensional subspace. However, most of existing FDR methods lack robustness against noise corruption. To this end, the representation-based subspace learning has been developed to find a robust projection matrix for FDR. Nevertheless, most of them only consider a single direction of the matrix, which ignore the information from other directions. Moreover, the majority of existing methods fail to account for both global structure and feature correlations effectively. To address the above problems, we propose a novel robust projection learning method called latent low-rank embedding (LatLRE), which integrates the latent low-rank representation (LatLRR) with projection learning. In particular, the proposed model can maintain the strong robustness of LatLRR and simultaneously learn a projection for FDR. Moreover, the nuclear norm and logarithmic norm are employed to approximate the two underlying rank functions and provide a more accurate measure of correlation. In addition, LatLRE is optimized using the alternating direction method of multipliers (ADMM) algorithm with the theoretical convergence guarantee. To verify the FDR performance of LatLRE, extensive experiments are conducted on three benchmark hyperspectral datasets. The experimental results demonstrate that LatLRE outperforms other FDR methods considered in this paper.
Heng-Chao Li 0001, Jun-Qiu Wang, Si-Jia Xiang, Qian Du 0001
IEEE Trans. Circuits Syst. Video Technol.1
2026 Multidimensional Image Reconstruction via Deep Nonlinear Low-Rank Tensor Decomposition
abstract
Low-rank tensor decomposition (LRTD) has demonstrated significant efficacy in multidimensional image reconstruction. Indeed, LRTD driven by nonlinear relationship can capture the underlying low-rank structure more accurately, since real-world data often exhibits complex nonlinear interactions. However, the existing nonlinear LRTD methods do not to investigate the inherent nonlinear interactions in spatial neighborhoods and spectral or temporal models. To address these challenges, we propose a novel deep nonlinear low-rank tensor decomposition (DNLRTD). Specifically, we design a deep nonlinear transform network (DNTN) using multiple convolutional layers and channel attention modules to form a deep nonlinear transform (DNT). The custom-designed DNT effectively captures nonlinear interactions within spatial neighborhoods while paying attention to the nonlinear interactions of spectral or temporal dimensions, consequently achieving a lower-rank representation. By integrating DNT into the low-tubal-rank decomposition framework, we induce the deep tubal-rank and form the DNLRTD. Also, we design a customized DNLRTD optimization strategy to make it flexible for different multidimensional image reconstruction tasks. Based on DNLRTD, we construct two multidimensional image reconstruction models and develop corresponding algorithms based on the alternating direction method of multipliers (ADMM) to solve them. Extensive experimental results on spectral compressive imaging and dynamic magnetic resonance image (MRI) reconstruction verify the superior performance of the proposed method.
Yu-Bang Zheng, Heng-Chao Li 0001, Antonio Plaza
IEEE Trans. Circuits Syst. Video Technol.3
2026 MLAgg-UNet: Advancing Medical Image Segmentation With Efficient Transformer and Mamba-Inspired Multi-Scale Sequence
abstract
Transformers and state space sequence models (SSMs) have attracted interest in biomedical image segmentation for their ability to capture longrange dependency. However, traditional visual state space (VSS) methods suffer from the incompatibility of image tokens with autoregressive assumption. Although Transformer attention does not require this assumption, its high computational cost limits effective channelwise information utilization. To overcome these limitations, we propose the Mamba-Like Aggregated UNet (MLAgg-UNet), which introduces Mamba-inspired mechanism to enrich Transformer channel representation and exploit implicit autoregressive characteristic within U-shaped architecture. For establishing dependencies among image tokens in single scale, the Mamba-Like Aggregated Attention (MLAgg) block is designed to balance representational ability and computational efficiency. Inspired by the human foveal vision system, Mamba macro-structure, and differential attention, MLAgg block can slide its focus over each image token, suppress irrelevant tokens, and simultaneously strengthen channel-wise information utilization. Moreover, leveraging causal relationships between consecutive low-level and high-level features in U-shaped architecture, we propose the Multi-Scale Mamba Module with Implicit Causality (MSMM) to optimize complementary information across scales. Embedded within skip connections, this module enhances semantic consistency between encoder and decoder features. Extensive experiments on four benchmark datasets, including AbdomenMRI, ACDC, BTCV, and EndoVis17, which cover MRI, CT, and endoscopy modalities, demonstrate that the proposed MLAgg-UNet consistently outperforms state-of-the-art CNN-based, Transformer-based, and Mamba-based methods. Specifically, it achieves improvements of at least 1.24%, 0.20%, 0.33%, and 0.39% in DSC scores on these datasets, respectively. These results highlight the model's ability to effectively capture feature correlations and integrate complementary multi-scale information, providing a robust solution for medical image segmentation.
Jiaxu Jiang, Sen Lei, Heng-Chao Li 0001
IEEE J. Biomed. Health Informatics3
2026 Multimodal Quaternion Representation Network for Multisource Remote Sensing Data Classification
abstract
The effective integration and classification of hyperspectral images (HSIs) and light detection and ranging (LiDAR) data is of great significance in Earth observation missions, which are confronted with challenges such as insufficient information utilization and feature heterogeneity. This article proposes a multimodal quaternion representation network (MMQRN) for multisource remote sensing (RS) data classification. Specifically, we first propose the multimodal quaternion representation (MMQR), which employs the orthogonal imaginary components of quaternions to model the complex nonlinear interactions among complementary features, thereby enabling their comprehensive fusion and utilization. Subsequently, we design a multimodal feature cross-fusion (MFCF) framework to integrate multisource, multimodal, and multilevel features adequately. Finally, we leverage the ability to capture long-term dependencies of transformers to design a quaternion convolutional transformer network (QCTN) for modeling global and local spatial-spectral information, respectively. Experiments conducted on three multisource RS datasets demonstrate the superior performance of the proposed MMQRN relative to other state-of-the-art classification methods.
Yu-Le Wei, Heng-Chao Li 0001, Jian-Li Wang, Yu-Bang Zheng, Qian Du 0001
IEEE Trans. Neural Networks Learn. Syst.2
2026 Tensor Multi-Subspace Representation for Remote Sensing Image Mixed Noise Removal
abstract
Remote sensing image (RSI) denoising is an important and fundamental task in RSI processing. Existing denoising methods usually assume that RSI lies in a single matrix or tensor subspace. However, due to the wavelength difference or/and temporal variability, the assumption of a single subspace may not be suitable for RSI. To address this, we propose a tensor multi-subspace representation (TenMSR) for RSI mixed noise removal. To be specific, in this work, we introduce TenMSR to finely characterize the intrinsic tensor multi-subspace structure of RSI. Compared with the single matrix/tensor subspace-based methods, the proposed method can not only precisely describe the wavelength difference or/and temporal variability of RSI but also produce a more compact image distribution in tensor multi-subspace. To mine and preserve the multi-subspace structure, we introduce a nonlinear transform-based 3-D tensor nuclear norm to characterize the tensor low rankness of the multi-subspace representation coefficient. An effective algorithm based on the proximal alternating minimization (PAM) framework is developed to solve the proposed model with theoretical convergence analysis. Extensive experiments show the effectiveness and superiority of the proposed method over existing state-of-the-art single matrix/tensor subspace RSI denoising methods.
Heng-Chao Li 0001, Meng Ding 0002, Xi-Le Zhao, Wen-Yu Hu
IEEE Trans. Neural Networks Learn. Syst.2
2025 Memory-Augmented Differential Network for Infrared Small Target Detection
abstract
Traditional U-Net-based methods in infrared small target detection (IRSTD) have demonstrated good performance. However, they often struggle with challenges such as blurred contour and strong interference in complex backgrounds. To overcome these issues, we propose a memory-augmented differential network (MAD-Net), which integrates two key modules: the adaptive differential convolution module (AdaDCM) and the memory-augmented attention module (MemA2M). AdaDCM leverages multiple differential convolutions to capture detailed edge information, with an adaptive fusion mechanism to weight and aggregate these features. In the deeper layers, by introducing the dataset-level representations through a learnable memory bank (LMB), MemA2M can enhance current features and effectively mitigate background interference. Extensive experiments on four public IRSTD datasets demonstrate that MAD-Net outperforms state-of-the-art methods, showcasing its superior capability in handling complex scenarios. The code is available at:https://github.com/joan2joan/MAD-Net.
Yanqiong Liu, Sen Lei, Nanqing Liu, Heng-Chao Li 0001
IEEE Geosci. Remote. Sens. Lett.5
2025 Tensor network decomposition for data recovery: Recent advancements and future prospects
Yu-Bang Zheng, Xi-Le Zhao, Heng-Chao Li 0001, Chao Li 0013, Ting-Zhu Huang, Qibin Zhao
Neural Networks3
2025 Multidimensional nonlinear transform-based tensor representation for high-dimensional image reconstruction
Yu-Bang Zheng, Heng-Chao Li 0001
Pattern Recognit.4
2025 Fully-connected tensor network decomposition with gradient factors regularization for robust tensor completion
Heng-Chao Li 0001, Rui Wang 0090, Yu-Bang Zheng
Signal Process.2
2025 Tensor Decomposition-Based Relaxed Linear Regression for Hyperspectral Image Classification
abstract
Linear regression and its variants have achieved considerable success in image classification. However, those methods still encounter two challenges when dealing with hyperspectral image (HSI) classification. On the one hand, the existing ones focus on mining the relationship between the label space and original data space during the classifier training, which is generally sensitive to noise corruptions. On the other hand, transforming the training samples into a strict binary label matrix makes the generalization ability of the classifier limited. To address these challenges, this paper constructs a novel integrative model called tensor decomposition-based relaxed linear regression (TDRLR) for HSI classification. Firstly, the model adopts tensor canonical polyadic (CP) decomposition to learn two dictionaries from spatial and spectral directions respectively, which can help to generate a double dictionary representation for HSI data. Then, the linear regression classifier is integrated to learn a transformation that reveals the mapping relation between the double dictionary representation and label space rather than the original data for enhancing robustness. Meanwhile, a more flexible way, label relaxation, is employed to enlarge the margins between different classes. More importantly, the learned double dictionary representation and classifier can be fine-tuned in tandem to enhance performance through the designed alternate iterative jointly learning algorithm. Experiments conducted on four real-world HSI datasets demonstrate that the proposed method achieves significant improvements in classification performance with a small size training set, when compared with state-of-the-art HSI classification methods.
Yangjun Deng, Lv-Wei Zhang, Longfei Ren, Heng-Chao Li 0001, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.5
2025 An Oriented Ship Detection Method of Remote Sensing Image With Contextual Global Attention Mechanism and Lightweight Task-Specific Context Decoupling
abstract
Ship detection in remote sensing images has been attracting a lot of attention due to its great application value in both military and civilian fields. However, ships in high-resolution remote sensing images are characterized by the remarkable features of multiscale, arbitrary orientation, and dense arrangement, which is a great challenge for fast and accurate target detection. In order to solve problems, we propose a YOLOV5-based oriented ship detection method of remote sensing images with contextual global attention mechanism and lightweight task-specific context decoupling (CGTC-RYOLO) in this article. First, a cross-stage partial context transformer (CSP-COT) module is introduced to capture global contextual spatial relations using multihead self-attention (MHSA) to verify their implications in implicit dependencies. Second, we propose an angle classification prediction branch in the YOLOV5 head network for detecting targets in any direction and design probability and distribution loss function (PrfoIoU) to optimize the regression effect. Third, the lightweight task-specific context decoupling (LTSCODE) for target detection is employed to replace the original head in the YOLOV5 model, which is used to solve the accuracy problem caused by YOLOV5’s hybridization of classification and localization. Ablation experiments demonstrate the importance and effectiveness of each module. Compared with the benchmark model, the CGTC-RYOLO has the 5.9%, 3.7%, and 4.3% mAP improvements on the DOTA-ship dataset, the HRSC2016 dataset, and the UCAS-AOD dataset, respectively. Moreover, the model’s generalization is also validated. Compared with state-of-the-artmethods, the CGTC-RYOLO can achieve better accuracy and fewer parameters.
Gui Gao, Gang Yang 0006, Libo Yao, Xi Zhang 0028, Heng-Chao Li 0001, Gaosheng Li
IEEE Trans. Geosci. Remote. Sens.7
2025 A Multibranch Embedding Network With Bi-Classifier for Few-Shot Ship Classification of SAR Images
abstract
Ship classification in synthetic aperture radar (SAR) images is a challenge in the field of ocean monitoring. On the one hand, there are few labeled samples in SAR remote sensing ship datasets, and a commonly used single classification criterion cannot effectively represent the distribution of categories. On the other hand, the small size of the SAR ship and the inconspicuous appearance characteristics lead to the fact that the SAR ship samples are with less discriminative information; therefore, the rich feature space of a ship cannot be effectively obtained, which increases the difficulty of target distinguishability. A multibranch embedding network with bi-classifier (MBEN-BC) model was proposed to address these problems and for few-shot SAR ship classification. First, the MBEN module was utilized to extract the multiscale feature map spatial information of the input image at multiple levels and establish cross-channel information interaction so as to obtain discriminative features at the local and global levels, which effectively enriched the feature space. Then, the BC module was constructed to represent the image features from the image level and descriptor level, respectively, and the two classification criteria were presented to promote a more compact distribution of similar samples in the feature space in order to effectively represent the distribution of categories with a small number of labeled samples. Experimental validation was carried out using the FUSAR-Ship, Moving and Stationary Target Acquisition and Recognition (MSTAR) dataset, and OPENSAR-Ship dataset, and the MBEN-BC method achieved superior performance and good generalization ability compared to the current popular and state-of-the-art few-shot methods.
Gui Gao, Meixiang Wang, Libo Yao, Xi Zhang 0028, Heng-Chao Li 0001, Gaosheng Li
IEEE Trans. Geosci. Remote. Sens.6
2025 Unsupervised Domain Adaptation With Hierarchical Masked Dual-Adversarial Network for End-to-End Classification of Multisource Remote Sensing Data
abstract
Although unsupervised domain adaptation (UDA) has been successfully applied for cross-scene classification of multisource remote sensing (MSRS) data, there are still some tough issues: 1) The vast majority of them are patch-based, requiring pixel by pixel processing at high complexity and ignoring the roles of unlabeled data between different domains. 2) Traditional masked autoencoder (MAE)-based methods lack effective multiscale analysis and require pre-training, ignoring the roles of low-level representations. As such, a hierarchical masked dual-adversarial DA network (HMDA-DANet) is proposed for cross-domain end-to-end classification of MSRS data. Firstly, a hierarchical asymmetric MAE (HAMAE) without pre-training is designed, containing a frequency dynamic large-scale convolutional (FDLConv) block to enhance important structural information in the frequency domain, and an intramodality enhancement and intermodality interaction (IAEIEI) block to embed some additional information beyond the domain distribution by expanding the cross-modal reconstruction space. Representative multimodal multiscale features can be extracted, while to some extent improving their generalization to the target domain. Then, a multimodal multiscale feature fusion (MMFF) block is built to model the spatial and scale dependencies for feature fusion and reduce the layer by layer transmission of redundancy or interference information. Finally, a dual-discriminator-based DA (DDA) block is designed for class-specific semantic feature and global structural alignments in both spatial and prediction spaces. It will enable HAMAE to model the cross-modal, cross-scale, and cross-domain associations, yielding more representative domain-invariant multimodal fusion features. Extensive experiments on five cross-domain MSRS datasets verify the superiority of the proposed HMDA-DANet over other state-of-the-art methods.
Wen-Shuai Hu, Wei Li 0032, Heng-Chao Li 0001, Xudong Zhao 0003, Mengmeng Zhang 0005, Ran Tao 0003
IEEE Trans. Geosci. Remote. Sens.3
2025 Exploring Fine-Grained Image-Text Alignment for Referring Remote Sensing Image Segmentation
abstract
Given a language expression, referring remote sensing image segmentation (RRSIS) aims to identify ground objects and assign pixelwise labels within the imagery. One of the key challenges for this task is to capture discriminative multimodal features via image-text alignment. However, the existing RRSIS methods use one vanilla and coarse alignment, where the language expression is directly extracted to be fused with the visual features. In this article, we argue that a “fine-grained image-text alignment” can improve the extraction of multimodal information. To this point, we propose a new RRSIS method to fully exploit the visual and linguistic representations. Specifically, the original referring expression is regarded as context text, which is further decoupled into the ground object and spatial position texts. The proposed fine-grained image-text alignment module (FIAM) would simultaneously leverage the features of the input image and the corresponding texts, obtaining better discriminative multimodal representation. Meanwhile, to handle the various scales of ground objects in remote sensing, we introduce a text-aware multiscale enhancement module (TMEM) to adaptively perform cross-scale fusion and intersections. We evaluate the effectiveness of the proposed method on two public referring remote sensing datasets including RefSegRS and RRSIS-D, and our method obtains superior performance over several state-of-the-art methods. The code will be publicly available athttps://github.com/Shaosifan/FIANet.
Sen Lei, Xinyu Xiao, Heng-Chao Li 0001, Zhenwei Shi 0001, Qing Zhu 0012
IEEE Trans. Geosci. Remote. Sens.4
2025 SAM-Based Building Change Detection With Distribution-Aware Fourier Adaptation and Edge-Constrained Warping
abstract
Building change detection remains challenging for urban development, disaster assessment, and military reconnaissance. While foundation models like Segment Anything Model (SAM) show strong segmentation capability, they exhibit limited performance in building change detection due to the domain gap between natural and remote sensing images. Existing adapter-based fine-tuning methods struggle with the imbalanced distribution of changed buildings, resulting in suboptimal detection of building changes exhibiting sparse distribution, small scale, or weak contrast. Additionally, bi-temporal alignment methods, including optical flow, are vulnerable to background noise interference. To address these limitations, we propose the SAM-based Network with Distribution-Aware Fourier Adaptation and Edge-Constrained Warping (FAEWNet) for building change detection. Unlike previous adapters overlook the imbalanced distribution of changed buildings, our proposed Distribution-Aware Fourier Aggregation Adapter not only addresses the domain gap issue, but also models the distribution of changed buildings. Furthermore, to mitigate noise interference and misalignment caused by registration errors, we design Multiscale Aware Flow Aggregation module that refines building edge extraction and enhances the perception of changed buildings. The results on the LEVIR-CD, S2Looking and WHU-CD datasets highlight the effectiveness of FAEWNet. Specifically, FAEWNet achieves the best performance on the LEVIR-CD with 91.29%Rc, 92.41%F1, and 85.89%IoU. On the S2Looking dataset, our method achieves improvements of 3.09% inRc, 1.07% inF1, and 1.23% inIoUcompared to the second-best method. On the WHU-CD dataset, it achieves the highest overall performance, with aRcof 94.20%, anF1 of 94.99%, and anIoUof 90.45%. The code is available at https://github.com/SUPERMAN123000/FAEWNet.
Yun-Cheng Li, Sen Lei, Heng-Chao Li 0001, Jun Li 0009, Antonio Plaza
IEEE Trans. Geosci. Remote. Sens.4
2025 PointSAM: Pointly-Supervised Segment Anything Model for Remote Sensing Images
abstract
Segment anything model (SAM) is an advanced foundational model for image segmentation, which is gradually being applied to remote sensing images (RSIs). Due to the domain gap between RSIs and natural images, traditional methods typically use SAM as a source pretrained model and fine-tune it with fully supervised masks. Unlike these methods, our work focuses on fine-tuning SAM using more convenient and challenging point annotations. Leveraging SAM’s zero-shot capability, we adopt a self-training framework that iteratively generates pseudolabels. However, noisy labels in pseudolabels can cause error accumulation. To address this, we introduce prototype-based regularization (PBR), where target prototypes are extracted from the dataset and matched to predicted prototypes using the Hungarian algorithm to guide learning in the correct direction. In addition, RSIs have complex backgrounds and densely packed objects, making it possible for point prompts to mistakenly group multiple objects as one. To resolve this, we propose a negative prompt calibration (NPC) method based on the nonoverlapping nature of instance masks, where overlapping masks are used as negative signals to refine segmentation. Combining these techniques, we present a novel pointly-supervised SAM (PointSAM). We conduct experiments on three RSI datasets, including WHU, HRSID, and NWPU VHR-10, showing that our method significantly outperforms direct testing with SAM, SAM2, and other comparison methods. In addition, PointSAM can act as a point-to-box converter for oriented object detection, achieving promising results and indicating its potential for other point-supervised tasks. The code is available athttps://github.com/Lans1ng/PointSAM.
Nanqing Liu, Xun Xu 0002, Yongyi Su, Heng-Chao Li 0001
IEEE Trans. Geosci. Remote. Sens.5
2025 ULADiff: Unmixing-Guided Learnable Abundance-Latent Diffusion for Hyperspectral Image Denoising
abstract
Hyperspectral image (HSI) denoising is a critical preprocessing step in remote sensing. Recently, the denoising diffusion probabilistic models (DDPMs) have emerged as the powerful generative models. However, applying DDPM to HSI denoising task remains challenging owing to the scarcity and acquisition difficulty of HSIs. Thus, how to effectively incorporate physical priors into the DDPM to improve denoising performance remains an underexplored issue. To this end, we propose an Unmixing-Guided Learnable Abundance-Latent Diffusion for HSI Denoising (ULADiff), which is a from-scratch, task-specific diffusion framework that incorporates physically interpretable priors and conditional information into the DDPM. ULADiff comprises three key components, including a Spectral Unmixing Transformer (SUT) network, an abundance-based diffusion model, and a reconstruction module. Specifically, we employ a learnable block-based SUT module in a self-supervised manner to decompose noisy HSIs into the abundance maps and endmembers. The SUT module enables the diffusion model to operate in a lower-dimensional abundance domain that better captures the underlying structure of HSIs. Then, we incorporate the first eigenimage, the reconstructed image via Singular Value Decomposition, as a physically meaningful condition to facilitate controllable generation. Furthermore, we propose a reconstruction module that enforces a spatial-spectral consistency prior by simultaneously imposing total variation regularization on the endmembers and a sparsity constraint on the abundance maps. This design preserves the intrinsic structures of the HSI and improves reconstruction quality. Comprehensive evaluations on synthetic and real-world datasets demonstrate that ULADiff outperforms state-of-the-art methods in both quantitative performance and visual fidelity.
Zhemin Wei, Heng-Chao Li 0001, Yu-Bang Zheng, Jian-Li Wang, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.2
2025 Tensorized High-Order Hypergraph Convolutional Network for Hyperspectral Image Classification
abstract
In recent years, graph convolutional networks (GCNs) have gained increasing attention in hyperspectral image (HSI) classification due to their good ability to model the pairwise relationships between two pixels. However, it is difficult to effectively model more complex relationships among multiple pixels with simple graphs. To solve this problem, we propose a novel tensorized high-order hypergraph convolutional network (TH2GCN) for HSI classification. Specifically, the hypergraph structure is employed to effectively model complex spatial relationships between pixels in HSIs, and we propose a new tensor-based algebraic representation of hypergraphs as a powerful strategy for describing the high-order interaction structures of the hypergraph. Besides, by extending the adjacency matrix-based GCN to the tensor domain and exploiting the tensor decomposition, the TH2GCN method is designed to efficiently extract high-order discriminative information from the hypergraph at low complexity for improving HSI classification performance. Furthermore, the construction of the adjacency tensor on all the data requires a huge amount of memory, especially for large-scale remote sensing images. To this end, the TH2GCN is trained and tested for HSI data in a minibatch fashion. Experimental results on three HSI datasets prove that the performance of the proposed method outperforms the comparison methods.
Jin-Yu Yang, Heng-Chao Li 0001, Shaohui Mei, Qian Du 0001, Antonio Plaza
IEEE Trans. Geosci. Remote. Sens.3
2025 Spectral-Temporal Consistency Prior for Cloud Removal From Remote Sensing Images
abstract
Thick cloud removal for multitemporal remote sensing images (MTRSIs) is a necessary preprocessing step for subsequent applications. Existing methods for cloud removal ignore spectral–temporal consistency prior (STCP), such as smooth regions of different time and bands existing in the same spatial location. To address this problem, we propose a factor-based group sparsity regularization within the low-rank tensor factorization (LRTF) framework and theoretically prove that it can characterize the STCP in MTRSIs. Based on this regularization, we construct a cloud removal model for MTRSIs. On one hand, the introduction of STCP enables the model to achieve superior cloud removal performance. On the other hand, regularization on small-sized factors rather than on the original data enables the model to have extremely low computational complexity. To solve this model, we develop a proximal alternating minimization (PAM)-based algorithm, in which we integrate a mask acquisition method based on separated cloud and shadow components. Comparative experiments using both simulated and real data demonstrate that the proposed method outperforms recent mask-unknown and mask-known methods in terms of performance and efficiency.
Shi-Jun Yang, Yu-Bang Zheng, Heng-Chao Li 0001, Yong Chen 0013, Qing Zhu 0012
IEEE Trans. Geosci. Remote. Sens.3
2025 Exploring Generalizable Pretraining for Real-World Change Detection via Geometric Estimation
abstract
As an essential procedure in earth observation system, change detection (CD) aims to reveal the spatial-temporal evolution of the observation regions. A key prerequisite for existing change detection algorithms is aligned geo-references between multi-temporal images by fine-grained registration. However, in the majority of real-world scenarios, a prior manual registration is required between the original images, which significantly increases the complexity of the CD workflow. In this paper, we proposed a self-supervision motivated CD framework with geometric estimation, called “MatchCD”. Specifically, the proposed MatchCD framework utilizes the zero-shot capability to optimize the encoder with self-supervised contrastive representation, which is reused in the downstream image registration and change detection to simultaneously handle the bi-temporal unalignment and object change issues. Moreover, unlike the conventional change detection requiring segmenting the full-frame image into small patches, our MatchCD framework can directly process the original large-scale image (e.g., 6K× 4Kresolutions) with promising performance. The performance in multiple complex scenarios with significant geometric distortion demonstrates the effectiveness of our proposed framework.
Sen Lei, Nanqing Liu, Heng-Chao Li 0001, Turgay Çelik 0001, Qing Zhu 0012
IEEE Trans. Geosci. Remote. Sens.4
2025 Bilateral Tensor Ring Decomposition for Thick Cloud Removal in Multitemporal Remote Sensing Images
abstract
Cloud removal is crucial for enhancing the quality of remote sensing images (RSIs) and broadening their applicability. Tensor decomposition, which extracts latent correlations across multidimensions, has driven the development of various methods for thick cloud removal in multitemporal RSI (MTRSI). However, current tensor decomposition methods are not specifically tailored for MTRSIs, resulting in insufficient correlation representation and requiring high computational costs in processing MTRSIs. In this article, we construct a novel bilateral tensor ring (BTR) decomposition, the first method specifically designed for MTRSIs, which enables a customized representation of spatial, spectral, and temporal correlations with lower computational complexity. The fundamental idea behind BTR decomposition is to effectively distinguish between the weak spatial correlation and the strong spectral–temporal correlation while simultaneously capturing the interaction between these two components. With the support of BTR decomposition, we propose an MTRSI cloud removal model and develop an efficient proximal alternating minimization (PAM)-based algorithm to solve it. In theory, we prove a convergence guarantee for the algorithm. Extensive experimental results verify that our method offers superior cloud removal performance and delivers a 10- to 100-fold acceleration in computational efficiency compared to the state-of-the-art tensor-based methods. The code is available at:https://yubangzheng.github.io.
Yu-Bang Zheng, Jia-Le Ma, Heng-Chao Li 0001, Qing Zhu 0012
IEEE Trans. Geosci. Remote. Sens.3
2025 Global Clue-Guided Cross-Memory Quaternion Transformer Network for Multisource Remote Sensing Data Classification
abstract
Multisource remote sensing data classification is a challenging research topic, and how to address the inherent heterogeneity between multimodal data while exploring their complementarity is crucial. Existing deep learning models usually directly adopt feature-level fusion designs, most of which, however, fail to overcome the impact of heterogeneity, limiting their performance. As such, a multimodal joint classification framework, called global clue-guided cross-memory quaternion transformer network (GCCQTNet), is proposed for multisource data [i.e., hyperspectral image (HSI) and synthetic aperture radar (SAR)/light detection and ranging (LiDAR)] classification. First, a three-branch structure is built to extract the local and global features, where an independent squeeze-expansion-like fusion (ISEF) structure is designed to update the local and global representations by considering the global information as an agent, suppressing the negative impact of multimodal heterogeneity layer by layer. A cross-memory quaternion transformer (CMQT) structure is further constructed to model the complex inner relationships between the intramodality and intermodality features to capture more discriminative fusion features that fully characterize multimodal complementarity. Finally, a cross-modality comparative learning (CMCL) structure is developed to impose the consistency constraint on global information learning, which, in conjunction with a classification head, is used to guide the end-to-end training of GCCQTNet. Extensive experiments on three public multisource remote sensing datasets illustrate the superiority of our GCCQTNet with regards to other state-of-the-art methods.
Wen-Shuai Hu, Wei Li 0032, Heng-Chao Li 0001, Ran Tao 0003
IEEE Trans. Neural Networks Learn. Syst.3
2025 Fully Tensorized Lightweight ConvLSTM Neural Networks for Hyperspectral Image Classification
abstract
Convolutional long short-term memory (ConvLSTM) possesses a remarkable capability of encoding spatial information and capturing long-range dependencies in sequential data. As a result, ConvLSTM has garnered success in hyperspectral image (HSI) classification. Nonetheless, the design of the special gate structures and convolution operations contributes to a high model complexity, making it challenging to deploy in resource-constrained environments. In this article, we propose a fully tensorized ConvLSTM model for HSI spatial-spectral classification under the premise of low complexity. First, we devise a novel and efficient tensor-sequenced convolution in the tensor train (TT) format, called ETTConv. ETTConv can reduce the number of parameters and computations in the standard convolutional layer by tensorizing the convolution kernels and mapping them to a series of smaller ones. Building upon this innovation, we present a novel ETTConvLSTM unit, formed by jointly compressing all weight tensors within the recurrent units. Using it as the fundamental unit, we construct the lightweight a efficient tensor train ConvLSTM 2-D neural network (ETTCL2DNN) model, characterized by reduced complexity without compromised classification performance. Furthermore, to better preserve the joint spatial-spectral structure of HSI data, we extend the ETTConv layer and the ETTConvLSTM unit to their 3-D versions, resulting in a new lightweight a efficient tensor train ConvLSTM 3-D neural network (ETTCL3DNN) model. Extensive quantitative experimental results on three widely used HSI datasets demonstrate the superiority of the proposed methods, exhibiting enhanced classification performance with reduced model complexity.
Tian-Yu Ma, Heng-Chao Li 0001, Yu-Bang Zheng, Qian Du 0001, Antonio Plaza
IEEE Trans. Neural Networks Learn. Syst.2
2024 SVDinsTN: A Tensor Network Paradigm for Efficient Structure Search from Regularized Modeling Perspective
abstract
Tensor network (TN) representation is a powerful technique for computer vision and machine learning. TN structure search (TN-SS) aims to search for a customized structure to achieve a compact representation, which is a chal-lenging NP-hard problem. Recent “sampling-evaluation”-based methods require sampling an extensive collection of structures and evaluating them one by one, resulting in pro-hibitively high computational costs. To address this issue, we propose a novel TN paradigm, named SVD-inspired TN decomposition (SVDinsTN), which allows us to efficiently solve the TN-SS problem from a regularized modeling per-spective, eliminating the repeated structure evaluations. To be specific, by inserting a diagonal factor for each edge of the fully-connected TN, SVDinsTN allows us to calculate TN cores and diagonal factors simultaneously, with the factor sparsity revealing a compact TN structure. In theory, we prove a convergence guarantee for the proposed method. Experimental results demonstrate that the proposed method achieves approximately 100 ~ 1000 times acceleration compared to the state-of-the-art TN-SS methods while maintaining a comparable level of representation ability.
Yu-Bang Zheng, Xi-Le Zhao, Junhua Zeng, Chao Li 0013, Qibin Zhao, Heng-Chao Li 0001, Ting-Zhu Huang
CVPR6
2024 Clip-Guided Source-Free Object Detection in Aerial Images
abstract
Domain adaptation is crucial in aerial imagery, as the visual representation of these images can significantly vary based on factors such as geographic location, time, and weather conditions. Additionally, high-resolution aerial images often require substantial storage space and may not be readily accessible to the public. To address these challenges, we propose a novel Source-Free Object Detection (SFOD) method. Specifically, our approach begins with a self-training framework, which significantly enhances the performance of baseline methods. To alleviate the noisy labels in self-training, we utilize Contrastive Language-Image Pre-training (CLIP) to guide the generation of pseudo-labels, termed CLIP-guided Aggregation (CGA). By leveraging CLIP’s zero-shot classification capability, we aggregate its scores with the original predicted bounding boxes, enabling us to obtain refined scores for the pseudo-labels. To validate the effectiveness of our method, we constructed two new datasets from different domains based on the DIOR dataset, named DIOR-C and DIOR-Cloudy. Experimental results demonstrate that our method outperforms other comparative algorithms. The code is available at https://github.com/Lans1ng/SFOD-RS.
Nanqing Liu, Xun Xu 0002, Yongyi Su, Peiliang Gong, Heng-Chao Li 0001
IGARSS6
2024 Infrared Small Target Detection Based on Weighted Tensor Average Rank Minimization and Directional Structure Tensor
abstract
To address the problems of poor robustness and target over-shrinkage in complex backgrounds of infrared small target detection (ISTD) algorithms, we propose a novel model by two steps. Firstly, we introduce a weighted tensor average nuclear norm with lpfunction (WTANN-lp) as the constraint term of the infrared background patch tensor to eliminate the issue of information’s underutilization due to transposition variability. This improves the robustness of the algorithm under complex background. Secondly, we design a directional structure tensor (DST) for extracting the target’s local prior information, which can distinguish the target from sparse residuals in multiple directions and overcome the challenge of target over-shrinkage. We use the alternating direction multiplier method (ADMM) to solve the proposed model, and extensive experiments demonstrate the superior target detection and shape reproduction performance of the proposed model compared to seven baseline detection methods.
Xi-Hu Yang, Yu-Bang Zheng, Jia-Jie Yin, Shi-Jun Yang, Heng-Chao Li 0001
IGARSS5
2024 Ship Detection Based on Feature Erasure and Channel-Wise Global Pooling Attention in SAR Images
abstract
Ship detection in synthetic aperture radar (SAR) images is an essential research direction in the field of remote sensing image interpretation. However, the performance of SAR image ship detection is inevitably susceptible to the influence of complex background and surrounding environment. To alleviate these issues, we propose a ship detection method based on feature erasure and attention mechanism. The Attention-based Feature Erasure Layer (AFELayer) is utilized to capture background information around the ship as a sub-significant feature to assist network localization. In addition, a Channel-wise Global Pooling Attention module (CGPA) is employed to correlate the significance of channels with each other on a global scale. Compared with existing algorithms, the experimental results show that the detection accuracy of proposed method reaches 75.5% on SSDD dataset, which shows superior performance in SAR image ship detection.
Heng-Chao Li 0001
IGARSS2
2024 Toward Distortion-Aware Change Detection in Realistic Scenarios
abstract
In the conventional change detection (CD) pipeline, two manually registered and labeled remote sensing datasets serve as the input of the model for training and prediction. However, in realistic scenarios, data from different periods or sensors could fail to be aligned as a result of various coordinate systems. Geometric distortion caused by coordinate shifting remains a thorny issue for CD algorithms. In this paper, we propose a reusable self-supervised framework for bitemporal geometric distortion in CD tasks. The whole framework is composed of Pretext Representation Pre-training, Bitemporal Image Alignment, and Down-stream Decoder Fine-Tuning. With only single-stage pre-training, the key components of the framework can be reused for assistance in the bitemporal image alignment, while simultaneously enhancing the performance of the CD decoder. Experimental results in 2 large-scale realistic scenarios demonstrate that our proposed method can alleviate the bitemporal geometric distortion in CD tasks.
Heng-Chao Li 0001, Nanqing Liu, Rui Wang 0090
IGARSS2
2024 Low-rank preserving embedding regression for robust image feature extraction
abstract
Abstract Although low‐rank representation (LRR)‐based subspace learning has been widely applied for feature extraction in computer vision, how to enhance the discriminability of the low‐dimensional features extracted by LRR based subspace learning methods is still a problem that needs to be further investigated. Therefore, this paper proposes a novel low‐rank preserving embedding regression (LRPER) method by integrating LRR, linear regression, and projection learning into a unified framework. In LRPER, LRR can reveal the underlying structure information to strengthen the robustness of projection learning. The robust metric L 2,1 ‐norm is employed to measure the low‐rank reconstruction error and regression loss for moulding the noise and occlusions. An embedding regression is proposed to make full use of the prior information for improving the discriminability of the learned projection. In addition, an alternative iteration algorithm is designed to optimise the proposed model, and the computational complexity of the optimisation algorithm is briefly analysed. The convergence of the optimisation algorithm is theoretically and numerically studied. At last, extensive experiments on four types of image datasets are carried out to demonstrate the effectiveness of LRPER, and the experimental results demonstrate that LRPER performs better than some state‐of‐the‐art feature extraction methods.
Tao Zhang 0027, Chen-Feng Long, Yangjun Deng, Wei-Ye Wang, Siqiao Tan, Heng-Chao Li 0001
IET Comput. Vis.6
2024 Dual auto-weighted multi-view clustering via autoencoder-like nonnegative matrix factorization
Si-Jia Xiang, Heng-Chao Li 0001, Xin-Ru Feng
Inf. Sci.2
2024 Feature Dimensionality Reduction With L2,p-Norm-Based Robust Embedding Regression for Classification of Hyperspectral Images
abstract
The curse of dimensionality and noise corruption are two tough problems that need to be solved in hyperspectral image (HSI) classification. However, the current feature dimensionality reduction methods, including both feature extraction and feature selection ones, cannot simultaneously solve the above two problems well. To address this issue, this paper proposes a novel method calledL2,p-norm-based robust embedding regression (L2,p-RER) for robust feature dimensionality reduction of HSI, which can effectively suppress the impact of noises and reduce the feature dimensions. Specifically,L2,p-RER first integrates projection learning with robust principle component analysis (RPCA) to remove noise in a low-dimensional space. Secondly, an embedding regression regularization is proposed to improve the discriminability of the extracted low-dimensional features. Thirdly, aL2,1-norm constraint is imposed to improve the interpretability of the learned projection matrix, which can jointly extract the key features from all bands with their physical meanings certainly preserved. Last but most important, theL2,p-norm that can adaptively balance the sparsity and the convexity is employed to model the noise and regression residual in the embedded low-dimensional space, which can further enhance the robustness and generalization of the proposed method. In addition, extensive experiments conducted on three benchmark HSI datasets validated the effectiveness of the proposed method.
Yangjun Deng, Menglong Yang, Heng-Chao Li 0001, Chen-Feng Long, Kui Fang, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.3
2024 Unsupervised Classification for Multilook Polarimetric SAR Images via Double Dirichlet Process Mixture Model
abstract
This paper proposes a hierarchical double Dirichlet process mixture model (DDPMM) for multilook polarimetric synthetic aperture radar (PolSAR) data unsupervised classification. Specifically, within the framework of product model (PM), an observed PolSAR data point can be factorized as the multiplication representation of a positive-scalar texture variable and a complex-Wishart-distributed speckle component. Based on this assumption, the polarization DPMM and texture DPMM in the proposed model are hierarchically established to characterize the polarimetric matrix and texture variable, respectively, thus yielding the generation procedure of the observation data to be learned sufficiently. Meanwhile, instead of sharing the same texture vector in many existing PM-based methods, each data point in DDPMM is associated with its own texture vector, which can be characterized as the weighted summation of several densities via texture DPMM rather than following a single distribution, such that the texture information can be fully and flexibly captured. In particular, dual local spatial constraints based on the statistical representations of polarization and texture spaces are also explored on the two DPMMs, allowing the local correlation to be adequately and dynamically incorporated. Moreover, all closed-form updates are derived with the variational Bayesian inference algorithm and the cluster number of the proposed model can be determined automatically. Experimental results on four real PolSAR datasets demonstrate the superiority of the proposed DDPMM to some state-of-the-art methods.
Heng-Chao Li 0001, Gui Gao, Wen Hong, William J. Emery
IEEE Trans. Geosci. Remote. Sens.2
2024 IDA-SiamNet: Interactive- and Dynamic-Aware Siamese Network for Building Change Detection
abstract
Building change detection (BCD) is a critical task in remote sensing which aims to identify the building changes within the same geographical area over time. The complexity of BCD is heightened when utilizing very high-resolution (VHR) remote sensing images, leading to two primary challenges: distinguishing between building and nonbuilding changes and accommodating the diverse range of building shapes and sizes. The existing mainstream methods neglect interactions between encoders, thereby compromising the ability to recognize building and nonbuilding changes. Additionally, most BCD methods overlook feature alignment and fusion which hinders the precise extraction of buildings with varying shapes and sizes. To address these limitations, we propose an interactive- and dynamic-aware Siamese network (IDA-SiamNet) for BCD. Our method comprises the spatial exchange feature interaction (SEFI) module, the channel exchange feature interaction (CEFI) module, and the dynamic-deformable dual-alignment fusion (D3AF) module. The SEFI and CEFI modules play a pivotal role in facilitating mutual information exchange between Siamese encoders, enhancing discrimination between building and nonbuilding changes. Furthermore, the D3AF module dynamically aggregates multiple parallel convolutional kernels to improve feature alignment and fusion for accurate building outline extraction. D3AF adapts its receptive field (RF) based on object size and covers diverse building shapes without introducing excessive background information. Experimental evaluations on three widely used BCD datasets, learning, vision, and remote sensing change detection (LEVIR-CD), satellite side-looking (S2Looking), and WHU BCD (WHU-CD), demonstrate the superior performance of our proposed method over state-of-the-art alternatives. Code and weights are made available athttps://github.com/SUPERMAN123000/IDA-SiamNet.
Yun-Cheng Li, Sen Lei, Nanqing Liu, Heng-Chao Li 0001, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.4
2024 HiReNet: Hierarchical-Relation Network for Few-Shot Remote Sensing Image Scene Classification
abstract
Few-shot scene classification aims to develop models that can quickly adapt to new scenes with only a few labeled samples that are not present in training sets. In recent years, convolutional neural networks (CNNs) have made significant advancements in few-shot remote sensing image scene classification tasks. However, most existing approaches focus solely on utilizing high-level embeddings of remote sensing images to learn similarity relations, while neglecting intrinsic hierarchical representations that could be crucial in distinguishing scenes with substantial interclass similarities. To address this limitation, we propose a novel few-shot scene classification method for remote sensing images called hierarchical-relation network (HiReNet). This approach leverages the hierarchical features of a query sample and its corresponding support sample to learn discriminative representations. HiReNet consists of an embedding network and a relation network. The embedding network employs a Siamese architecture to extract representations, while the relation network utilizes these representations for classification. Within the relation network, we introduce a hierarchical relation learning (HRL) structure to capture the hierarchical relations among query and support samples. Additionally, to extract stronger features, we introduce a feature aggregation module that concatenates multilevel features and employs channel attention to re- weight these features. Experimental results demonstrate the superior performance of our HiReNet compared to several state-of-the-art few-shot scene classification methods.
Sen Lei, Yingbo Zhou 0001, Jialin Cheng, Guohao Liang, Zhengxia Zou, Heng-Chao Li 0001, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.7
2024 Multifrequency Graph Convolutional Network With Cross-Modality Mutual Enhancement for Multisource Remote Sensing Data Classification
abstract
The mining of meaningful features and effective fusion of multisource remote sensing (RS) data have always been the challenging research problems in the joint classification of hyperspectral image (HSI) and light detection and ranging (LiDAR) data. In this paper, we propose a Multi-Frequency Graph Convolutional Network with Cross-modality Mutual Enhancement (MFGCN-CME) for multisource RS data classification. Specifically, we design an adaptive multi-frequency graph feature learning module to capture the low- and high-frequency multiscale features of HSI and LiDAR in parallel and further adaptively aggregate them. Then, we propose a bipartite graph enhancement learning module to obtain the spatial-enhanced HSI features and spectral-enhanced LiDAR features by propagating inter-modality information. To the best of our knowledge, the bipartite graph is first used to multisource RS data classification task. Furthermore, compared with traditional fusion methods, a gated fusion module is used to fully explore the complementarity of two data sources. Finally, a joint loss function combing a classification loss and a semi-supervised contrastive loss is developed to improve the model robustness. Comprehensive experiments on different HSI and LiDAR datasets demonstrate that our proposed method can yield better performance compared with several state-of-the-art multisource RS data classification methods.
Jin-Yu Yang, Heng-Chao Li 0001, Lei Pan 0003, Qian Du 0001, Antonio Plaza
IEEE Trans. Geosci. Remote. Sens.2
2024 Spatial-Temporal Weighted and Regularized Tensor Model for Infrared Dim and Small Target Detection
abstract
Due to the confusion of target-like sparse structures and the interference of linear features in complex scenarios, many infrared small target detection methods struggle to effectively detect dim and small targets. In response to this challenge, we propose a new 3-D paradigm framework, which combines spatial-temporal weighting and regularization within a low-rank sparse tensor decomposition model. First, we design a novel spatial-temporal local prior structure tensor, named 3DST, which can significantly distinguish between targets and target-like sparse structures. Second, we introduce a three-directional log-based tensor nuclear norm (3DLogTNN) to provide a full characterization of the low-rankness of the background tensor. Third, we suggest a weighted three-directional total variation (3DTV) regularization to constrain smoothness features in background images. Finally, we develop an efficient alternating direction method of multipliers (ADMMs) to solve the proposed model. In particular, we devise a fast and accurate Sylvester tensor equation for accelerated subproblem solving. Extensive experimental results demonstrate that the proposed model has superior target detection and background suppression performance in complex scenarios compared with other detection methods.
Jia-Jie Yin, Heng-Chao Li 0001, Yu-Bang Zheng, Gui Gao, Ran Tao 0003
IEEE Trans. Geosci. Remote. Sens.2
2024 SSLChange: A Self-Supervised Change Detection Framework Based on Domain Adaptation
abstract
In conventional remote sensing change detection (RSCD) procedures, extensive manual labeling for bi-temporal images is first required to maintain the performance of subsequent fully supervised training. However, pixel-level labeling for change detection (CD) tasks is very complex and time-consuming. In this article, we explore a novel self-supervised contrastive framework applicable to the RSCD task, which promotes the model to accurately capture spatial, structural, and semantic information through the domain adapter (DA) and the hierarchical contrastive head. The proposed SSLChange framework accomplishes self-learning only by taking a single-temporal sample and can be flexibly transferred to mainstream CD baselines. With self-supervised contrastive learning, feature representation pretraining can be performed directly based on the original data even without labeling. After a certain number of labels are subsequently obtained, the pretrained features will be aligned with the labels for fully supervised fine-tuning. Without introducing any additional data or labels, the performance of downstream baselines will experience a significant enhancement. Experimental results on two entire datasets and six diluted datasets show that our proposed SSLChange improves the performance and stability of CD baseline in data-limited situations. The code of SSLChange is available athttps://github.com/MarsZhaoYT/SSLChange
Turgay Çelik 0001, Nanqing Liu, Feng Gao 0005, Heng-Chao Li 0001
IEEE Trans. Geosci. Remote. Sens.5
2023 An optimized deep supervised hashing model for fast image retrieval
Abid Hussain 0002, Heng-Chao Li 0001, Muqadar Ali, Fakhar Abbas, Mehboob Hussain
Image Vis. Comput.2
2023 Break Through the Border Restriction of Horizontal Bounding Box for Arbitrary-Oriented Ship Detection in SAR Images
abstract
Substantial progress has been made in detecting ships of arbitrary orientation in synthetic aperture radar (SAR) images. However, the mainstream method is still limited by the horizontal bounding box (HBB) boundary, which cannot provide scaling information for the length and width of the oriented bounding box (OBB) in an intuitive way. In this study, we propose a novel encode representation to describe the OBB by breaking through the border restriction of the HBB. Specifically, we derive an inclination factor from two left-top point offsets (LTPO), which enables us to directly infer the coordinates of the four OBB vertices and obtain an oriented rectangular proposal. To obtain high-quality oriented semantic features, we utilize a feature adaptive module (FAM) to learn the shape and orientation implied by arbitrary-oriented ships through spatial transformation. Our comparative experiments demonstrate that our proposed method achieves superior performance and detection accuracy on two commonly-used benchmark datasets for oriented SAR ship detection, namely SSDD and HRSID.
Turgay Çelik 0001, Nanqing Liu, Heng-Chao Li 0001
IEEE Geosci. Remote. Sens. Lett.4
2023 Quaternion Convolutional Neural Network With EMAP Representation for Multisource Remote-Sensing Data Classification
abstract
The fusion and classification of hyperspectral images (HSIs) and light detection and ranging (LiDAR) data have been extensively studied using deep learning. However, traditional real-valued deep learning methods have limitations in distinguishing internal and external relations and capturing fine spatial characteristics. To break through these limitations, this letter proposes a quaternion convolutional neural network (QCNN) with extended morphological attribute profile (EMAP) quaternion representation (called EQR) for multisource remote sensing (RS) data classification by utilizing quaternion properties. Specifically, we first propose the EQR for each single-source data, which encodes the multi-attribute features in a compact yet comprehensive manner, highlighting the internal relations. Secondly, we embed EQR into QCNN to preserve the internal relations and enable the interaction of multi-attribute features. Then we develop the 3-D quaternion convolution to better exploit the 3-D characteristic of HSI. Finally, we design different attention mechanisms and a two-level fusion strategy for multisource data to learn enhanced features. Experiments on two multisource RS data sets show that the proposed method achieved better performance than other state-of-the-art classification methods.
Yu-Le Wei, Yu-Bang Zheng, Rui Wang 0090, Heng-Chao Li 0001
IEEE Geosci. Remote. Sens. Lett.4
2023 Convolution and Attention Mixer for Synthetic Aperture Radar Image Change Detection
abstract
Synthetic aperture radar (SAR) image change detection is a critical task and has received increasing attentions in the remote sensing community. However, existing SAR change detection methods are mainly based on convolutional neural networks (CNNs), with limited consideration of global attention mechanism. In this letter, we explore Transformer-like architecture for SAR change detection to incorporate global attention. To this end, we propose a convolution and attention mixer (CAMixer). First, to compensate the inductive bias for Transformer, we combine self-attention with shift convolution in a parallel way. The parallel design effectively captures the global semantic information via the self-attention and performs local feature extraction through shift convolution simultaneously. Second, we adopt a gating mechanism in the feed-forward network to enhance the non-linear feature transformation. The gating mechanism is formulated as the element-wise multiplication of two parallel linear layers. Important features can be highlighted, leading to high-quality representations against speckle noise. Extensive experiments conducted on three SAR datasets verify the superior performance of the proposed CAMixer. The source codes will be publicly available at https://github.com/summitgao/CAMixer.
Haopeng Zhang 0016, Zijing Lin, Feng Gao 0005, Junyu Dong, Qian Du 0001, Heng-Chao Li 0001
IEEE Geosci. Remote. Sens. Lett.6
2023 t-Linear Tensor Subspace Learning for Robust Feature Extraction of Hyperspectral Images
abstract
Subspace learning has been widely applied for feature extraction of hyperspectral images (HSIs) and achieved great success. However, the current methods still leave two problems that need to be further investigated. First, those methods mainly focus on finding one or multiple projection matrices for mapping the high-dimensional data into a low-dimensional subspace, which can only capture the information from each direction of high-order hyperspectral data separately. Second, the performance of feature extraction is barely satisfactory when the hyperspectral data is severely corrupted by noise. To address these issues, this article presents a t-linear tensor subspace learning (tLTSL) model for robust feature extraction of HSIs based on t-product projection. In the model, t-product projection is a new defined tensor transformation way similar to linear transformation in vector space, which can maximally capture the intrinsic structure of tensor data. The integrated tensor low-rank and sparse decomposition can effectively remove the noise corruption and the learned t-product projection can directly transform the high-order hyperspectral data into a subspace with information from all modes comprehensively considered. Moreover, a proposition related to tensor rank is proofed for interpreting the meaning of the tLTSL model. Extensive experiments are conducted on two different kinds of noise (i.e., simulated and real noise) corrupted HSI data, which validate the effectiveness of tLTSL.
Yangjun Deng, Heng-Chao Li 0001, Siqiao Tan, Junhui Hou, Qian Du 0001, Antonio Plaza
IEEE Trans. Geosci. Remote. Sens.2
2023 Hybrid Fully Connected Tensorized Compression Network for Hyperspectral Image Classification
abstract
Deep learning models, such as convolutional neural networks (CNNs), have made significant progress in hyperspectral image (HSI) classification. However, these models require a large number of parameters, which occupy a lot of storage space and suffer from overfitting, thus resulting in performance loss. To solve the above problems, in this article, we propose a new compression network [namely, a Hybrid Fully Connected Tensorized Compression Network (HybridFCTCN)] by considering the high dimensionality of HSI data. First, using the low-rank fully connected tensor network decomposition (FCTND), three novel units, i.e., FCTN-FC, FCTNConv2D, and FCTNConv3D, are designed to compress the weight tensor of standard fully connected (FC) layer and kernel tensor of convolutional layer, reducing their parameters. In the novel units, the intrinsic correlation of the decomposed factors is adequately exploited by the FC structures, which enhances their feature extraction and classification abilities. Then, benefiting from the hybrid network backbone composed of the FCTNConv3D and FCTNConv2D units, HybridFCTCN can extract more discriminative features with fewer parameters, while it has great generalization capability and robustness, enabling better HSI classification. Finally, the rank of above-designed units is defined, and its determination is discussed to facilitate the application of the proposed model. Extensive experiments on three widely used HSI datasets reveal that the proposed model achieves state-of-the-art classification performance for different training sample sizes with a very small number of parameters.
Heng-Chao Li 0001, Zhi-Xin Lin, Tian-Yu Ma, Xi-Le Zhao, Antonio Plaza, William J. Emery
IEEE Trans. Geosci. Remote. Sens.1
2023 Transformation-Invariant Network for Few-Shot Object Detection in Remote-Sensing Images
abstract
Object detection in remote sensing images relies on a large amount of labeled data for training. However, the increasing number of new categories and class imbalance make exhaustive annotation impractical. Few-shot object detection (FSOD) addresses this issue by leveraging meta-learning on seen base classes and fine-tuning on novel classes with limited labeled samples. Nonetheless, the substantial scale and orientation variations of objects in remote sensing images pose significant challenges to existing few-shot object detection methods. To overcome these challenges, we propose integrating a feature pyramid network and utilizing prototype features to enhance query features, thereby improving existing FSOD methods. We refer to this modified FSOD approach as a Strong Baseline, which has demonstrated significant performance improvements compared to the original baselines. Furthermore, we tackle the issue of spatial misalignment caused by orientation variations between the query and support images by introducing a Transformation-Invariant Network (TINet). TINet ensures geometric invariance and explicitly aligns the features of the query and support branches, resulting in additional performance gains while maintaining the same inference speed as the Strong Baseline. Extensive experiments on three widely used remote sensing object detection datasets, i.e., NWPU VHR-10.v2, DIOR, and HRRSD demonstrated the effectiveness of the proposed method.
Nanqing Liu, Xun Xu 0002, Turgay Çelik 0001, Zongxin Gan, Heng-Chao Li 0001
IEEE Trans. Geosci. Remote. Sens.5
2023 What Catch Your Attention in SAR Images: Saliency Detection Based on Soft-Superpixel Lacunarity Cue
abstract
In existing superpixel-wise saliency detection algorithms, superpixel generation often is an isolated preprocessing step. The performance of saliency maps is determined by the accuracy of superpixels to a certain extent. However, it is still a challenge to develop a stable superpixel generation method. In this article, we attempt to incorporate the superpixel generation and saliency calculation steps into an end-to-end trainable deep network. First, we employ a recently proposed differentiable superpixel generation method to over-segment the synthetic aperture radar (SAR) images, which outputs the possibility that the pixels assigned to neighbor superpixels (soft superpixel). In saliency calculation part, as one of our main contributions, we propose a differentiable and computationally simple saliency model, i.e., lacunarity cue. It is inspired by the fact that generally the backscattering intensity of regions of interest (ROIs) in SAR images irregularly fluctuates, while the areas with consistent pixels are often ignored as the clusters. We improve the pixelwise box differential dimension algorithm to measure the irregularity of scattering points in a superpixel. The superpixel generation and saliency calculation can be implemented under a unified deep network. Hence, the shapes of the superpixels can be iteratively adjusted according to the saliency maps until the ROIs are correctly detected. Experiments on real SAR images with different sizes and scenes show that the saliency maps can effectively highlight the target areas, thus outperforming the state-of-the-art saliency detection models.
Fei Ma 0001, Xuejiao Sun, Fan Zhang 0007, Yongsheng Zhou, Heng-Chao Li 0001
IEEE Trans. Geosci. Remote. Sens.5
2023 Nearest Neighbor-Based Contrastive Learning for Hyperspectral and LiDAR Data Classification
abstract
The joint hyperspectral image (HSI) and light detection and ranging (LiDAR) data classification aims to interpret ground objects at more detailed and precise level. Although deep learning methods have shown remarkable success in the multisource data classification task, self-supervised learning has rarely been explored. It is commonly nontrivial to build a robust self-supervised learning model for multisource data classification, due to the fact that the semantic similarities of neighborhood regions are not exploited in the existing contrastive learning framework. Furthermore, the heterogeneous gap induced by the inconsistent distribution of multisource data impedes the classification performance. To overcome these disadvantages, we propose a nearest neighbor-based contrastive learning network (NNCNet), which takes full advantage of large amounts of unlabeled data to learn discriminative feature representations. Specifically, we propose a nearest neighbor-based data augmentation scheme to use enhanced semantic relationships among nearby regions. The intermodal semantic alignments can be captured more accurately. In addition, we design a bilinear attention module to exploit the second-order and even high-order feature interactions between the HSI and LiDAR data. Extensive experiments on four public datasets demonstrate the superiority of our NNCNet over state-of-the-art methods. The source codes are available athttps://github.com/summitgao/NNCNet.
Feng Gao 0005, Junyu Dong, Heng-Chao Li 0001, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.4
2023 Unsupervised Domain Factorization Network for Thick Cloud Removal of Multitemporal Remotely Sensed Images
abstract
Cloud removal is an important task in the remotely sensed images (RSIs) processing, which is beneficial for downstream applications, such as unmixing, fusion, and target detection. Multi-temporal remotely sensed images (MRSIs), which contains the abundant spatial-spectral-temporal (SST) information, potentially bring the new opportunities for cloud removal. However, how to effectively and efficiently explore the rich information of MRSIs remains a challenge. Inspired by the low-rankness of MRSIs, we propose an Unsupervised Domain Factorization Network (UnDFN) for thick cloud removal, which allows us to effectively and efficiently exploit the rich SST information of MRSIs. In UnDFN framework, we first factorize RSI for each time node of MRSIs into its corresponding spatial factor and spectral factor. Due to the powerful expressive ability, the untrained neural networks are leveraged to faithfully capture the spatial and spectral factors. Especially, motivated by the low-rankness of the concatenated spatial factors of all time nodes, a low-rank spatial factor module is elaborately designed to effectively and efficiently capture the spatial factors of all time nodes as compared with separately using networks to capture spatial factors for each time node. Extensive experiments on simulated and real MRSIs of different satellites (including Sentinel-2 and Landsat-8) substantiate that the proposed UnDFN achieves state-of-the-art performance in thick cloud removal compared to other methods.
Jian-Li Wang, Xi-Le Zhao, Heng-Chao Li 0001, Ke-Xiang Cao, Jiaqing Miao, Ting-Zhu Huang
IEEE Trans. Geosci. Remote. Sens.3
2022 Robust Patch Tensor-based Multigraph Embedding for Dimensionality Reduction of Hyperspectral Images
abstract
Since hyperspectral image (HSI) is naturally presented as 3D data cube, patch tensor-based graph embedding methods have been widely applied for dimensionality reduction (DR) of HSI. However, these methods are usually developed with the patch tensors generated by the raw data which is inevitably contaminated by noise. To eliminate the negative affects of noise, this paper introduces region covariance descriptor (R-CD) to characterize the HSI data and proposes a robust patch tensor-based multigraph embedding (RPTMGE) method for DR of HSI. RPTMGE constructs three types of subgraphs to comprehensively describe the intrinsic structure of HSI. Specifically, the manifold subgraph in RPTMGE is constructed with the RCD of HSI, which can significantly enhance the robustness of RPTMGE. Finally, experiments on real HSI data are conducted and the results demonstrated the effectiveness of the proposed method.
Yangjun Deng, Wei-Ye Wang, Heng-Chao Li 0001
IGARSS5
2022 SAR Image Data Augmentation via Residual and Attention-Based Generative Adversarial Network for Ship Detection
abstract
In recent years, generative adversarial networks (GANs) have been successfully applied to generate the SAR images. However, due to the fact that it is more difficult to generate the images than to distinguish the real or fake, GANs usually suffer from the problems of unstable training and mode collapse. As such, a residual and attention-based generative adversarial network (RAGAN) is proposed for SAR data augmentation. Firstly, the directional bounding box is used as a constraint in the RAGAN to limit the position of ship in the generated SAR image, which can be further set as the annotation of the SAR image for ship detection directly. After that, inspired by the residual and attention learning, a residual and attention block (RABlock) and a transposed RABlock (TRABlock) are designed to improve the generator of the RAGAN, thus preventing the whole model from gradient vanishing and suppressing the effects of speckle noise and background to enhance the quality of the generated SAR images. Experimental results on the HRSID data set demonstrate the effectiveness of our RAGAN model in SAR data augmentation for ship detection.
Yu-Shi Guo, Heng-Chao Li 0001, Wen-Shuai Hu, Wei-Ye Wang
IGARSS2
2022 Spatially Variant Gamma-WMM with Extended Variational Inference for Unsupervised PolSAR Classification
abstract
The Wishart mixture model (WMM) has been widely used for classification of polarimetric synthetic aperture radar (PolSAR) images; however, the WMM-based models usually fail to provide reliable classification results and explore the spatial information effectively in the heterogeneous areas. As such, an unsupervised spatially variant Gamma-WMM with extended variational inference algorithm (SVGaWMM-EVI) is proposed for classification of PolSAR images. Firstly, the Gamma prior distribution is imposed on the texture variable of the proposed model, which associates a set of unique texture variables with each data point to utilize the spatial information in the heterogeneous areas. Then, since the existing expectation maximization-based WMM algorithms usually fall into local optimal and update slowly, an extended variational inference algorithm is developed to improve the parameter estimation of our model, where a help function is designed to solve the intractable term. Experimental results on the real-world PolSAR data set demonstrate that our model can obtain better performance than some widely used unsupervised methods.
Heng-Chao Li 0001, Wen-Shuai Hu, Lei Pan 0003
IGARSS2
2022 Weakly supervised building semantic segmentation via superpixel-CRF with initial deep seeds guiding
abstract
Abstract The segmentation of building from satellite and airborne images is necessary for high‐resolution buildings maps generation and it is still challenging. On annotated pixel‐level images, trained deep convolutional neural networks (CNNs) were used to improve segmentation of building. The cost of labelling training data is high, which reduces their usage. Human labelling efforts can be significantly reduced using weakly supervised segmentation techniques. Here, a novel weakly supervised framework is introduced for building semantic segmenting that relies on deep seeds to construct a superpixels‐CRF model over superpixels segmentation in order to generate high‐quality initial pixel‐level annotations, as the initialization step. Then, the segmentation network is trained using the initial pixel‐level annotations. Next, the CRF model is used to refine the segmentation masks, and the segmentation network is retrained to achieve accurate pixel‐level annotations while iteratively optimizing the segmentation. The experimental results on three public building datasets demonstrate that the proposed framework significantly improved the quality of building semantic segmentation while remaining computationally efficient.
Khaled Moghalles, Heng-Chao Li 0001, Zaid Al-Huda, Asad Malik 0002
IET Image Process.2
2022 Fully Convolutional Lightweight Pyramid Network for Vehicle Detection in Aerial Images
abstract
Vehicle detection in aerial images has a wide range of applications, and the majority of vehicle detection methods use the bounding-box approach for localization. But the bounding-box approach usually yields low precision and recall rates, especially under the dense vehicle density situations where the close spatial proximity of the vehicles confuses the bounding-box detectors. This letter proposes a Fully Convolutional Lightweight Pyramid Network (FCLPN) to detect vehicles in visible-spectrum aerial images. Unlike the bounding-box approach, FCLPN performs pixel-level localization and classification. FCLPN trained on the DLR-3K dataset is directly tested on DLR-3K, VEDAI, COWC, and VAID datasets to validate its generalization strength. Experimental results show that FCLPN performs better than state-of-the-art methods in aerial vehicle detection in terms of Precision, F1 score, and mAP.
Qingsong Du, Turgay Çelik 0001, Heng-Chao Li 0001
IEEE Geosci. Remote. Sens. Lett.4
2022 Correntropy-Based Autoencoder-Like NMF With Total Variation for Hyperspectral Unmixing
abstract
In order to unmix the hyperspectral imagery (HSI) with better performance, this letter proposes a correntropy-based autoencoder-like nonnegative matrix factorization (NMF) (CANMF) with total variation (CANMF-TV) method. NMF is extensively applied to unmix the mixed pixels. However, it only reconstructs the original data from the abundances in endmember space. To directly project the original data space into the endmember space, and then achieve the abundance matrix, we first exploit an autoencoder-like NMF for hyperspectral unmixing, which integrates bothdecoderandencoder. Considering that HSI is typically degraded by noise, the correntropy-induced metric (CIM) is introduced to construct a CANMF model. In addition, TV regularizer is imposed into the CANMF model so as to preserve the spatial-contextual information by promoting the piecewise smoothness of abundances. Finally, a series of experiments are conducted on both synthetic and real data sets, demonstrating the effectiveness of the proposed CANMF-TV method over comparison.
Xin-Ru Feng, Heng-Chao Li 0001, Shuang Liu 0016, Hongyan Zhang 0001
IEEE Geosci. Remote. Sens. Lett.2
2022 Correction to "Correntropy-Based Autoencoder-Like NMF With Total Variation for Hyperspectral Unmixing"
abstract
In the above article[1], it should be noted that the second value of each column in the last row ofTable I(i.e., 0.52%, 0.35%, 0.51%, and 0.72%) is calculated using average deviation rather than standard deviation. In order to be consistent with the title ofTable Iin[1], the corresponding standard deviations are provided here.
Xin-Ru Feng, Heng-Chao Li 0001, Shuang Liu 0016, Hongyan Zhang 0001
IEEE Geosci. Remote. Sens. Lett.2
2022 Recurrent Feedback Convolutional Neural Network for Hyperspectral Image Classification
abstract
Deep neural networks have achieved promising performance for hyperspectral image (HSI) classification. However, due to the limitation of the available labeled samples, the traditional deeper and wider neural networks usually cause the overfitting problem and lose the detailed information. To solve this problem, a brain-like structure, namely spatial attention-driven recurrent feedback convolutional neural network (SARFNN), is proposed by utilizing the recurrent feedback and attention mechanism structures, from which two deep models are further developed for HSI classification. First, a 2-D SARFNN (SARF2DNN) model is developed to learn the spatial features from HSI data. After that, to better exploit the 3-D characteristic, the 3-D version is extended from SARF2DNN, thus constructing an SARF3DNN model to extract joint spatial-spectral features. Moreover, with the help of the idea of brain-likeness, the recurrent feedback module is designed to recover information loss caused by deeper structure and the dimension reduction operation. The experimental results conducted on two HSI data sets show that our SARFNN architecture can achieve more competitive performance than other state-of-the-art algorithms.
Heng-Chao Li 0001, Shuang-Shuang Li, Wen-Shuai Hu, Jun-Huan Feng, Weiwei Sun 0005, Qian Du 0001
IEEE Geosci. Remote. Sens. Lett.1
2022 Gated Ladder-Shaped Feature Pyramid Network for Object Detection in Optical Remote Sensing Images
abstract
This letter presents a new feature pyramid network (FPN) called the gated ladder-shaped FPN (GLFPN) to construct more representative feature pyramids for detecting objects of different sizes in optical remote sensing images. We first use convolution and concatenation operations to fuse three base features extracted by a ResNet backbone. We then obtain multilevel features from these base features. Finally, we use a selective gate to fuse features from multiple levels with equivalent sizes. To evaluate the effectiveness of the proposed GLFPN, we integrate it into the RetinaNet architecture by replacing the conventional FPN. The experimental results on two optical remote sensing image data sets show that the proposed method outperforms the methods compared in this letter.
Nanqing Liu, Turgay Çelik 0001, Heng-Chao Li 0001
IEEE Geosci. Remote. Sens. Lett.3
2022 MSNet: A Multiple Supervision Network for Remote Sensing Scene Classification
abstract
Remote sensing scene classification is a complex task due to large intraclass variations in object appearances with a small number of samples per class and high interclass similarities due to shared objects in different classes, which usually cause model overfitting and high interclass confusion. To address these challenges, a multiple supervision approach, called multiple supervision network (MSNet), consisting of the ResNet-50 backbone, a feature discriminative branch (FDB), and a feature confusion branch (FCB) is proposed in this letter. The FDB selects discriminative features per class and suppresses peaks in feature maps to examine more informative regions with lower feature magnitudes. Meanwhile, the FCB reduces overfitting by introducing confusion to the input of a fully connected layer which also enhances the robust features. The FDB and FCB are only used in training of the backbone and not used in inference. Thus, the proposed method does not introduce additional computing time on the backbone while it significantly boosts its performance in scene classification. The experimental results show that MSNet outperforms the methods considered in this letter.
Nanqing Liu, Turgay Çelik 0001, Heng-Chao Li 0001
IEEE Geosci. Remote. Sens. Lett.3
2022 Synthetic Aperture Radar Image Change Detection via Layer Attention-Based Noise-Tolerant Network
abstract
Recently, change detection methods for synthetic aperture radar (SAR) images based on convolutional neural networks (CNN) have gained increasing research attention. However, existing CNN-based methods neglect the interactions among multilayer convolutions, and errors involved in the preclassification restrict the network optimization. To this end, we proposed a layer attention-based noise-tolerant network, termed LANTNet. In particular, we design a layer attention module that adaptively weights the feature of different convolution layers. In addition, we design a noise-tolerant loss function that effectively suppresses the impact of noisy labels. Therefore, the model is insensitive to noisy labels in the preclassification results. The experimental results on three SAR datasets show that the proposed LANTNet performs better compared to several state-of-the-art methods. The source codes are available at https://github.com/summitgao/LANTNet.
Desen Meng, Feng Gao 0005, Junyu Dong, Qian Du 0001, Heng-Chao Li 0001
IEEE Geosci. Remote. Sens. Lett.5
2022 Change Detection in Synthetic Aperture Radar Images Using a Dual-Domain Network
abstract
Change detection from synthetic aperture radar (SAR) imagery is a critical yet challenging task. Existing methods mainly focus on feature extraction in the spatial domain, and little attention has been paid to the frequency domain. Furthermore, in patch-wise feature analysis, some noisy features in the marginal region may be introduced. To tackle the above two challenges, we propose a dual-domain network (DDNet). Specifically, we take features from the discrete cosine transform (DCT) domain into consideration and the reshaped DCT coefficients are integrated into the proposed model as the frequency domain branch. Feature representations from both frequency and spatial domain are exploited to alleviate the speckle noise. In addition, we further propose a multi-region convolution (MRC) module, which emphasizes the central region of each patch. The contextual information and central region features are modeled adaptively. The experimental results on three SAR data sets demonstrate the effectiveness of the proposed model. Our codes are available athttps://github.com/summitgao/SAR_CD_DDNet.
Xiaofan Qu, Feng Gao 0005, Junyu Dong, Qian Du 0001, Heng-Chao Li 0001
IEEE Geosci. Remote. Sens. Lett.5
2022 Unsupervised Robust Projection Learning by Low-Rank and Sparse Decomposition for Hyperspectral Feature Extraction
abstract
Owing to the strong correlation between the spectral bands of hyperspectral images (HSIs), many feature extraction (FE) methods have been proposed to reduce the redundancy of hyperspectral data. However, Euclidean distance-based FE methods are sensitive to noise. To address this issue, this letter proposed a new unsupervised FE method called robust projection learning (RPL) by integrating the low-rank and sparse decomposition with projection learning. Specifically, in order to enhance the discrimination of traditional robust principal component analysis (RPCA), discriminative RPCA (DRPCA) is first proposed by decomposing the raw data into a low-rank part, a discriminative sparse part, and a structured noise. Moreover, for the purpose of redundancy reduction, projection learning is integrated into DRPCA to obtain a projection matrix with robustness and discrimination. To verify the validity of RPL, two real hyperspectral data sets are used for basic comparison and robust analysis. The corresponding experimental results demonstrate that RPL outperforms the comparative FE methods.
Heng-Chao Li 0001, Lei Pan 0003, Yangjun Deng, Qian Du 0001
IEEE Geosci. Remote. Sens. Lett.2
2022 Adaptive Cross-Attention-Driven Spatial-Spectral Graph Convolutional Network for Hyperspectral Image Classification
abstract
Recently, graph convolutional networks (GCNs) have been developed to explore the spatial relationship between pixels, achieving better classification performance of hyperspectral images (HSIs). However, these methods fail to sufficiently leverage the relationship between spectral bands in HSI data. As such, we propose an adaptive cross-attention-driven spatial–spectral graph convolutional network (ACSS-GCN), which is composed of a spatial GCN (Sa-GCN) subnetwork, a spectral GCN (Se-GCN) subnetwork, and a graph cross-attention fusion module (GCAFM). Specifically, Sa-GCN and Se-GCN are proposed to extract the spatial and spectral features by modeling the correlations between spatial pixels and between spectral bands, respectively. Then, by integrating attention mechanism into information aggregation of the graph, the GCAFM, including three parts, i.e., the spatial graph attention block, the spectral graph attention block, and the fusion block, is designed to fuse the spatial and spectral features, and suppress noise interference in Sa-GCN and Se-GCN. Moreover, the idea of the adaptive graph is introduced to explore an optimal graph through backpropagation during the training process. Experiments on two HSI datasets show that the proposed method achieves better performance than other classification methods.
Jin-Yu Yang, Heng-Chao Li 0001, Wen-Shuai Hu, Lei Pan 0003, Qian Du 0001
IEEE Geosci. Remote. Sens. Lett.2
2022 A Comparative Analysis of GAN-Based Methods for SAR-to-Optical Image Translation
abstract
Unlike optical sensors, Synthetic Aperture Radar (SAR) sensors acquire images of the Earth’s surface with all-weather and all-time capabilities, which is vital in a situation such as a disaster assessment. However, SAR sensors do not offer as rich visual information as optical sensors. SAR-to-Optical image-to-image translation generates optical images from SAR images to benefit from what both imaging modalities have to offer. It also enables multi-sensor image analysis of the same scene for applications such as heterogeneous change detection. Various architectures of Generative Adversarial Networks (GANs) have achieved remarkable image-to-image translation results in different domains. Still, their performances in SAR-to-Optical image translation have not been analyzed in the remote sensing domain. This paper compares and analyses the state-of-the-art GAN-based translation methods with open-source implementations for SAR-to-Optical image translation. The results show that GAN-based SAR-to-Optical image translation methods achieve satisfactory results. However, their performances depend on the structural complexity of the observed scene and the spatial resolution of the data. We also introduce a new dataset with a higher resolution than the existing SAR-to-Optical image datasets and release implementations of GAN-based methods considered in this paper to support the reproducible research in remote sensing.
Turgay Çelik 0001, Nanqing Liu, Heng-Chao Li 0001
IEEE Geosci. Remote. Sens. Lett.4
2022 An Arbitrary-Oriented Object Detector Based on Variant Gaussian Label in Remote Sensing Images
abstract
Arbitrary-oriented remote sensing object detection is a challenging task because of the periodicity of the angle and large variety of aspect ratios. To address these issues, this letter proposes two simple yet powerful methods: a variant Gaussian label (VGL) generating method and a channel-wise pixel attention (CPA) module. Specifically, VGL is generated to handle the periodicity of the angle and produce various labels to adapt diverse aspect ratios of objects. Then, CPA is utilized to fuse features between channels and pixels, to obtain a global receptive field and extract more robust features. In addition, we collect a dataset, namely, Northwestern Polytechnical University Very-High-Resolution (NWPU VHR-10-R), which is relabeled with oriented bounding boxes based on NWPU VHR-10. To evaluate our proposed method, experiments are conducted on several publicly available benchmark datasets, including the High Resolution Ship Collections 2016 (HRSC2016) and NWPU VHR-10-R. Experimental results show that our method achieves substantial gains compared with baseline approaches.
Tingyu Zhao, Nanqing Liu, Turgay Çelik 0001, Heng-Chao Li 0001
IEEE Geosci. Remote. Sens. Lett.4
2022 Pseudo Complex-Valued Deformable ConvLSTM Neural Network With Mutual Attention Learning for Hyperspectral Image Classification
abstract
Convolutional long short-term memory (ConvLSTM) has received much attention for hyperspectral image (HSI) classification due to its ability of modeling long-range correlations, which, however, is vulnerable to too many parameters and insufficient training, limiting its classification accuracy, especially for small samples. Different from it, traditional hand-crafted methods extract the features with basic attributes of HSIs, which can provide the lack of details and interpretability of deep semantic features. However, existing methods fail to incorporate their complementarity for HSI classification. As such, a Pseudo complex-valued (CV) Deformable ConvLSTM Neural Network with mutual Attention learning (APDCLNN) is proposed, providing a new way to realize the collaborative learning of hand-crafted and deep features for HSI classification. First, a 2-D pseudo CV deformable ConvLSTM (PDConvLSTM2D) cell is designed using deformable convolution and complex operations, with which a spatial–spectral PDConvLSTM2D neural network (SSPDCL2DNN) is built to extract scale- and spectral-enhanced deep spatial–spectral features. Then, 3-D Gabor filter is used to extract hand-crafted features, and a mutual attention-based multimodality feature learning and fusion (MAMLF) module is designed to integrate them into deep features for training and optimization of SSPDCL2DNN. Finally, an attention loss subnetwork is designed to refine the classification results. As we know, this is the first attempt to apply the idea of mutual attention learning to fuse hand-crafted and deep features for HSI classification. Extensive experiments on three widely used HSI datasets show the advantages of our model over other deep methods in terms of both quantitative and visual quality.
Wen-Shuai Hu, Heng-Chao Li 0001, Rui Wang 0090, Feng Gao 0005, Qian Du 0001, Antonio Plaza
IEEE Trans. Geosci. Remote. Sens.2
2022 Self-Supervised Robust Deep Matrix Factorization for Hyperspectral Unmixing
abstract
Hyperspectral unmixing is a critical step to process hyperspectral images (HSIs). Nonnegative matrix factorization (NMF) has drawn extensive attention in remotely sensed hyperspectral unmixing since it does not require prior knowledge about the pure spectral constituents (endmembers) in the scene. However, this approach is normally implemented as a single-layer procedure, which does not allow for a refinement of the obtained endmember abundances. In addition, HSIs suffer from the interference of sparse noise (besides Gaussian noise), which brings challenges when pursuing efficient hyperspectral unmixing. To address these issues, we propose a new self-supervised robust deep matrix factorization (SSRDMF) model for hyperspectral unmixing, which consists of two parts:encoderanddecoder. In theencoder, a multilayer nonlinear structure is designed to directly map the observed HSI data to the corresponding abundances. The abundances are then decoded by thedecoder, in which the connected weights are treated as the extracted endmembers. By modeling the sparse noise explicitly, the proposed method can reduce the effect caused by both Gaussian and sparse noise. Furthermore, a self-supervised constraint is included for exploring the spectral information, which is beneficial to further improve unmixing performance. To validate our method, we have conducted extensive experiments on both synthetic and real datasets. Our experiments reveal that our newly developed SSRDMF achieves superior unmixing performance compared to other state-of-the-art methods.
Heng-Chao Li 0001, Xin-Ru Feng, Donghai Zhai, Qian Du 0001, Antonio Plaza
IEEE Trans. Geosci. Remote. Sens.1
2022 Lightweight Tensorized Neural Networks for Hyperspectral Image Classification
abstract
Deep learning methods have demonstrated excellent performance in hyperspectral image (HSI) classification. However, these methods mainly focus on improving the classification accuracy while ignoring their high complexity. By considering that the data formats of both HSIs and network weights can be represented in the form of tensors, we develop a new lightweight tensorized neural network for HSI classification that takes advantage of low-rank tensor decomposition techniques to reduce complexity. Firstly, inspired by tensor train (TT)-based tensorized convolutional layers, a new tensorized 2D convolutional layer based on chain calculation (with better expression ability) is introduced. Based on this innovation, a new lightweight 2D tensorized neural network (2D-TNN) is designed for HSI classification. Furthermore, to better preserve the intrinsic structure of HSI data, a new lightweight 3D tensorized neural network (3D-TNN) is proposed by extending the tensorized 2D convolutional layers to their 3D versions. Quantitative and comparative experiments on three widely used data sets show that the proposed models are able to achieve state-of-the-art performance (with a low number of model parameters) for different training sample sizes, especially for very small training sets.
Tian-Yu Ma, Heng-Chao Li 0001, Rui Wang 0090, Qian Du 0001, Xiuping Jia, Antonio Plaza
IEEE Trans. Geosci. Remote. Sens.2
2022 A Multiscale Spectral Features Graph Fusion Method for Hyperspectral Band Selection
abstract
This article proposes a multiscale spectral features graph fusion (MSFGF) method for selecting proper hyperspectral bands. The MSFGF regards that the selected bands should reflect diagnostic spectral information of ground objects at different scales, and it explores band selection from the aspect of multiple spatial scales. First, it adopts the multiscale low-rank decomposition (MSLRD) model to find multiscale spectral features of different ground objects. The model considers divergent spatial structures or spatial correlations of ground objects at different scales, and factorizes the hyperspectral data cube into a series of low-rank block-wise data cubes, where the blocks take spatial structures of different ground objects at increasing scales. Second, the MSFGF presents the multiscale sparse spectral clustering (MSSC) model to fuse the separate connected graphs of multiscale spectral features into a consensus graph. The consensus graph combines the complementary information of multiscale spectral features and helps to reveal the intrinsic clustering structure of all spectral bands. Finally, the MSFGF utilizes spectral clustering to find clusters from the consensus graph and selects representative bands. Experimental results on three widely used hyperspectral data prove the superiority of MSFGF in selecting bands, where it outperforms other seven state-of-the-art methods in classification with an acceptable computational cost.
Weiwei Sun 0005, Gang Yang 0006, Jiangtao Peng, Xiangchao Meng, Wei Li 0032, Heng-Chao Li 0001, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.7
2022 Information Fusion for Classification of Hyperspectral and LiDAR Data Using IP-CNN
abstract
Joint use of multisensor information has attracted considerable attention in the remote sensing community. While applications in land-cover observation benefit from information diversity, multisensor integration technique is confronted with many challenges, including inconsistent size of data, different data structures, uncorrelated physical properties, and scarcity of training data. In this article, an information fusion network, named interleaving perception convolutional neural network (IP-CNN), is proposed for integrating heterogeneous information and improving joint classification performance of hyperspectral image (HSI) and light detection and ranging (LiDAR) data. Specifically, a bidirectional autoencoder is designed to reconstruct hyperspectral and LiDAR data together, and the reconstruction process is trained with no dependence upon annotated information. Both HSI-perception constraint and LiDAR-perception constraint are imposed on multisource structural information integration. Accordingly, fused data are fed into a two-branch CNN for final classification. To validate the effectiveness of the model, the experiments were conducted using three datasets (i.e., Muufl Gulfport data, Trento data, and Houston data). The final results demonstrate that the proposed framework can significantly outperform state-of-the-art methods even with small-size training samples.
Mengmeng Zhang 0005, Wei Li 0032, Ran Tao 0003, Heng-Chao Li 0001, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.4
2022 A3 CLNN: Spatial, Spectral and Multiscale Attention ConvLSTM Neural Network for Multisource Remote Sensing Data Classification
abstract
The problem of effectively exploiting the information multiple data sources has become a relevant but challenging research topic in remote sensing. In this article, we propose a new approach to exploit the complementarity of two data sources: hyperspectral images (HSIs) and light detection and ranging (LiDAR) data. Specifically, we develop a new dual-channel spatial, spectral and multiscale attention convolutional long short-term memory neural network (called dual-channel$A^{3}$CLNN) for feature extraction and classification of multisource remote sensing data. Spatial, spectral, and multiscale attention mechanisms are first designed for HSI and LiDAR data in order to learn spectral- and spatial-enhanced feature representations and to represent multiscale information for different classes. In the designed fusion network, a novel composite attention learning mechanism (combined with a three-level fusion strategy) is used to fully integrate the features in these two data sources. Finally, inspired by the idea of transfer learning, a novel stepwise training strategy is designed to yield a final classification result. Our experimental results, conducted on several multisource remote sensing data sets, demonstrate that the newly proposed dual-channel$A^{\,3}$CLNN exhibits better feature representation ability (leading to more competitive classification performance) than other state-of-the-art methods.
Heng-Chao Li 0001, Wen-Shuai Hu, Wei Li 0032, Jun Li 0009, Qian Du 0001, Antonio Plaza
IEEE Trans. Neural Networks Learn. Syst.1
2021 FER-YOLO: Detection and Classification Based on Facial Expressions
Hui Ma 0014, Turgay Çelik 0001, Heng-Chao Li 0001
ICIG (1)3
2021 Spatial-Spectral Tensor Graph Convolutional Network for Hyperspectral Image Classification
abstract
Graph convolutional networks (GCNs) have come necessarily into being with the increase of graph-structure data in real world, which recently, have been applied for feature extraction and classification of hyperspectral image (HSI). However, only spatial relationship of pixels being considered by GCN-based models, the relationship between spectral bands are underutilized. To solve above issue, benefit from introducing tensor theory to GCN, a novel spatial-spectral tensor graph convolutional network (SSTGCN) is proposed to learn tensor representation of spatial-spectral feature. Firstly, from the perspective of the spatial-spectral characteristic in HSI, spatial graph tensor and spectral graph tensor are constructed to model the spatial and spectral relationships, respectively. Then, two types of modules, i.e., the spatial tensor graph convolution (SATGC) module and spectral tensor graph convolution (SETGC) module, are designed to extract discriminative features. Experimental results on the Indian Pines data set demonstrate that the proposed method outperforms other HSI classification methods.
Jin-Yu Yang, Heng-Chao Li 0001, Tian-Yu Ma
IGARSS2
2021 SeqFace: Learning discriminative features by using face sequences
abstract
Abstract Deep convolutional neural networks (CNNs) have greatly improved the Face Recognition (FR) performance in recent years. Almost all CNNs in FR are trained on the carefully labeled datasets containing plenty of identities. However, such high‐quality datasets are very expensive to collect, which restricts many researchers to achieve state‐of‐the‐art performance. In this paper, a framework, called SeqFace, for learning discriminative face features is proposed. Besides a traditional identity training dataset, the designed SeqFace can train CNNs by using an additional dataset which includes a large number of face sequences collected from videos. Moreover, the label smoothing regularization (LSR) and a new proposed discriminative sequence agent (DSA) loss are employed to enhance the discrimination power of deep face features via making full use of the sequence data. Only with a single ResNet model, the method achieves very competitive performance on several face recognition benchmarks, including LFW, YTF, CFP, AgeDB, and MegaFace. The code and model are publicly available at the website https://github.com/huangyangyu/SeqFace .
Wei Hu 0004, Yangyu Huang, Fan Zhang 0007, Ruirui Li 0001, Heng-Chao Li 0001
IET Image Process.5
2021 SAR Image Change Detection Based on Multiscale Capsule Network
abstract
Traditional synthetic-aperture radar (SAR) image change detection methods based on convolutional neural networks (CNNs) face the challenges of speckle noise and deformation sensitivity. To mitigate these issues, we proposed a multiscale capsule network (Ms-CapsNet) to extract the discriminative information between the changed and unchanged pixels. On the one hand, the multiscale capsule module is employed to exploit the spatial relationship of features. Therefore, equivariant properties can be achieved by aggregating the features from different positions. On the other hand, an adaptive fusion convolution (AFC) module is designed for the proposed Ms-CapsNet. The higher semantic features can be captured for the primary capsules. Feature extracted by the AFC module significantly improves the robustness to speckle noise. The effectiveness of the proposed Ms-CapsNet is verified on three real SAR data sets. The comparison experiments with four state-of-the-art methods demonstrate the efficiency of the proposed method. Our codes are available at https://github.com/summitgao/SAR_CD_MS_CapsNet.
Yunhao Gao, Feng Gao 0005, Junyu Dong, Heng-Chao Li 0001
IEEE Geosci. Remote. Sens. Lett.4
2021 Generative Adversarial Capsule Network With ConvLSTM for Hyperspectral Image Classification
abstract
Recently, deep learning has been widely applied in hyperspectral image (HSI) classification since it can extract high-level spatial-spectral features. However, deep learning methods are restricted due to the lack of sufficient annotated samples. To address this problem, this letter proposes a novel generative adversarial network (GAN) for HSI classification that can generate artificial samples for data augmentation to improve the HSI classification performance with few training samples. In the proposed network, a new discriminator is designed by exploiting capsule network (CapsNet) and convolutional long short-term memory (ConvLSTM), which extracts the low-level features and combines them together with local space sequence information to form the high-level contextual features. In addition, a structured sparse L2,1constraint is imposed on sample generation to control the modes of data being generated and achieve more stable training. The experimental results on two real HSI data sets show that the proposed method can obtain better classification performance than the several state-of-the-art deep classification methods.
Wei-Ye Wang, Heng-Chao Li 0001, Yangjun Deng, Li-Yang Shao, Xiaoqiang Lu, Qian Du 0001
IEEE Geosci. Remote. Sens. Lett.2
2021 Random Neighbor Pixel-Block-Based Deep Recurrent Learning for Polarimetric SAR Image Classification
abstract
Polarimetric synthetic aperture radar (PolSAR) image classification is an important part of SAR data interpretation and provides more intuitive and detailed SAR polarization information. To bridge the PolSAR data and applications, it is necessary to design a comprehensive PolSAR classification framework to achieve satisfactory results. The deep neural network (DNN) appears to be a solution for the classification issue, in which it outperforms the classical supervised classifiers under the condition of sufficient training data. However, the volume of training data will greatly limit the effectiveness of practical applications. In this article, we try to solve the dependence issue on training data in three different ways: recurrent learning, data augmentation, and postprocessing. First, the long short-term memory (LSTM) network is introduced to achieve pixel sequence learning by taking into account the spatial and polarimetric features. Second, the random neighbor pixel-block (RNPB) method is proposed to increase the number of training samples for sequence learning. Third, the conditional random field (CRF) model is employed to further improve the classification accuracy. In the experiments, three sets of PolSAR data are used to evaluate the small sample performance of the proposed classification method. With only 0.5% labeled pixels for training, the proposed RNPB-LSTM-CRF method can approach 99% overall classification accuracy for all the data sets. Compared with the existing methods, the proposed method can achieve state-of-the-art results for PolSAR image classification under the condition of 1% training samples.
Fan Zhang 0007, Qiang Yin 0001, Yongsheng Zhou, Heng-Chao Li 0001, Wen Hong
IEEE Trans. Geosci. Remote. Sens.5
2021 Learning Center Probability Map for Detecting Objects in Aerial Images
abstract
One fundamental problem in Earth Vision is to accurately find the locations and identify the categories of the interesting objects in the aerial images, for which oriented bounding boxes (OBBs) are usually employed to depict better the objects emerging with arbitrary orientations. However, the regression of the OBBs always suffers from the ambiguous problem in the definition of the regression targets, which often reduces the convergency efficiency and decreases the detection accuracy. Although there are some methods like the binary segmentation map that can handle this problem, it brings a new problem of ambiguous background pixels in the OBBs. In this article, we propose to cast the OBB regression as a center-probability-map (CenterMap)-prediction problem, thus largely eliminating the ambiguities on the target definitions and the background pixels. The predicted CenterMaps are then used to generate the OBBs. The CenterMap OBB representation is simple, yet effective. Furthermore, to distinguish better the interesting objects from the cluttered background, a weighted pseudosegmentation-guided attention network is adopted to provide the object-level features for predicting the horizontal bounding boxes and the OBBs. The experimental results on three widely used data sets, i.e., DOTA, HRSC2016, and UCAS-AOD, demonstrate the effectiveness of our proposed method.
Jinwang Wang, Wen Yang 0001, Heng-Chao Li 0001, Gui-Song Xia
IEEE Trans. Geosci. Remote. Sens.3
2020 Change Detection of Polarimetric SAR Images Using Minkowski Log-Ratio Distance
abstract
The Minkowski log-ratio (MLR) distance admits closed-form formula for mixture model of exponential families (Gaussian family, Wishart family, etc.), and is suitable for measuring the dissimilarity of polarimetric SAR (PolSAR) data. In this paper, MLR distance is introduced for PolSAR image change detection. Specifically, the PolSAR images are estimated by Wishart mixture models and over-segmented into superpixels first. After that, the statistical distribution differences between two corresponding superpixels are measured by MLR distance. Finally, the change detection map is obtained by the Kittler-Illingworth thresholding method based on the generated difference map. Qualitative and quantitative experimental analysis shows the MLR distance can provide a desirable difference map which facilitates the following thresholding stage.
Shuailin Chen, Xiangli Yang, Tongyuan Zou, Dong Peng, Wen Yang 0001, Heng-Chao Li 0001
IGARSS6
2020 Hyperspectral Image Classification Based on Tensor-Train Convolutional Long Short-Term Memory
abstract
In recent years, deep learning models have shown great advantages for hyperspectral images (HSIs) classification, in which long short-term memory (LSTM) has attracted plenty of attentions for its characteristic of modeling long-range dependencies. However, for the 2-D extended architecture of it (namely 2-D convolutional LSTM, ConvLSTM2D), it is the special gate structures of ConvLSTM2D that leads to a large number of training parameters and high requirements for device storage. To address this shortcoming, in this paper, a lightweight ConvLSTM2D cell is developed by using tensor-train decomposition (TTD) for the compression of training parameters, which is named TT-ConvLSTM2D and further applied to two state-of-the-art ConvLSTM2D-based HSI classification models for verifying its superiority. Experiments on a widely-used Indian Pines HSI data set are conducted, whose results demonstrate that the proposed TT-ConvLSTM2D cell can effectively reduce the number of the parameters and memory requirements of the whole models within a small range of accuracy degradation.
Wen-Shuai Hu, Heng-Chao Li 0001, Tian-Yu Ma, Qian Du 0001, Antonio Plaza, William J. Emery
IGARSS2
2020 Edge-Driven Object Matching for UAV Images and Satellite SAR Images
abstract
The task of matching between images acquired by terminal equipment and satellites is important and challenging due to the dramatic viewpoint changes and unknown orientations, especially with different imaging sensors. In this paper, we firstly present the task to match the optical/infrared images acquired by UAVs with satellite SAR images. Many previous works mainly focused on matching with the images of the same modality, and may not perform well to our task. To overcome the difficulties caused by the diversity among the three modalities of data, we mine the common features of them and propose a novel edge-driven matching framework to find the correspondence between the UAV images and SAR images. Experimental results demonstrate the effectiveness and superiority of our method.
Ruixiang Zhang, Huai Yu, Wen Yang 0001, Heng-Chao Li 0001
IGARSS5
2020 Patch Tensor-Based Multigraph Embedding Framework for Dimensionality Reduction of Hyperspectral Images
abstract
Graph-based dimensionality reduction (DR) techniques are of great interest in the field of image processing and especially on the analysis of hyperspectral images (HSIs). Considering the characteristics of hyperspectral data, many different types of graphs were designed to describe the structure of HSIs. Generally, the algorithms based on these graphs achieved promising performance. However, most of them only focus on how to improve the measurement of similarity between the data points by a single graph. Specifically, vector-based graph methods fail to capture the spatial information, while tensor-based graph methods assume that the pixels in each patch tensor belong to the same class, which is not exactly correct in practice. To overcome these shortcomings, this article proposes a patch tensor-based multigraph embedding (PTMGE) framework for the DR of HSIs, in which three different types of subgraphs are constructed to comprehensively describe the intrinsic geometrical structures of HSIs. First, a tensor subgraph is constructed to capture the spatial information and local geometrical structure. Second, for each two neighboring patch tensors in the tensor graph, a bipartite graph is designed to characterize the pixel-based relationships between the patch tensors. Then, considering that the diversity of pixels may be existed in each patch tensor, a pixel-based subgraph is built to describe the inner geometrical structures of every patch tensor. Finally, a novel graph fusion strategy is designed to calculate a final similarity matrix for projection learning. Experiments on three real hyperspectral data sets are conducted, and comparison with some state-of-the-art algorithms validated the effectiveness of our proposed PTMGE method.
Yangjun Deng, Heng-Chao Li 0001, Yong-Jian Sun, Xiangrong Zhang, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.2
2020 Spatial-Spectral Feature Extraction via Deep ConvLSTM Neural Networks for Hyperspectral Image Classification
abstract
In recent years, deep learning has presented a great advance in the hyperspectral image (HSI) classification. Particularly, long short-term memory (LSTM), as a special deep learning structure, has shown great ability in modeling long-term dependencies in the time dimension of video or the spectral dimension of HSIs. However, the loss of spatial information makes it quite difficult to obtain better performance. In order to address this problem, two novel deep models are proposed to extract more discriminative spatial-spectral features by exploiting the convolutional LSTM (ConvLSTM). By taking the data patch in a local sliding window as the input of each memory cell band by band, the 2-D extended architecture of LSTM is considered for building the spatial-spectral ConvLSTM 2-D neural network (SSCL2DNN) to model long-range dependencies in the spectral domain. To better preserve the intrinsic structure information of the hyperspectral data, the spatial-spectral ConvLSTM 3-D neural network (SSCL3DNN) is proposed by extending LSTM to the 3-D version for further improving the classification performance. The experiments, conducted on three commonly used HSI data sets, demonstrate that the proposed deep models have certain competitive advantages and can provide better classification performance than the other state-of-the-art approaches.
Wen-Shuai Hu, Heng-Chao Li 0001, Lei Pan 0003, Wei Li 0032, Ran Tao 0003, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.2
2020 Discriminative Marginalized Least-Squares Regression for Hyperspectral Image Classification
abstract
Least-squares regression (LSR)-based classifiers are effective in multiclassification tasks. However, most existing methods use limited projections, resulting in loss of much discriminant information; furthermore, they focus only on exactly fitting samples to target matrix while ignoring overfitting issue. To solve these drawbacks, discriminative marginalized LSR (DMLSR) is proposed to learn a more discriminative projection matrix with consideration of class separability and data-reconstruction ability simultaneously. In the proposed framework, an intraclass compactness graph is employed to avoid the overfitting problem and enhance class separability, and a data-reconstruction constraint is imposed to preserve discriminant information on limited projections. Experimental results on several hyperspectral data sets demonstrate that the proposed method significantly outperforms some state-of-the-art classifiers.
Yuxiang Zhang 0005, Wei Li 0032, Heng-Chao Li 0001, Ran Tao 0003, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.3
2020 Joint Classification of Hyperspectral and LiDAR Data Using Hierarchical Random Walk and Deep CNN Architecture
abstract
Earth observation using multisensor data is drawing increasing attention. Fusing remotely sensed hyperspectral imagery and light detection and ranging (LiDAR) data helps to increase application performance. In this article, joint classification of hyperspectral imagery and LiDAR data is investigated using an effective hierarchical random walk network (HRWN). In the proposed HRWN, a dual-tunnel convolutional neural network (CNN) architecture is first developed to capture spectral and spatial features. A pixelwise affinity branch is proposed to capture the relationships between classes with different elevation information from LiDAR data and confirm the spatial contrast of classification. Then in the designed hierarchical random walk layer, the predicted distribution of dual-tunnel CNN serves as global prior while pixelwise affinity reflects the local similarity of pixel pairs, which enforce spatial consistency in the deeper layers of networks. Finally, a classification map is obtained by calculating the probability distribution. Experimental results validated with three real multisensor remote sensing data demonstrate that the proposed HRWN significantly outperforms other state-of-the-art methods. For example, the two branches CNN classifier achieves an accuracy of 88.91% on the University of Houston campus data set, while the proposed HRWN classifier obtains an accuracy of 93.61%, resulting in an improvement of approximately 5%.
Xudong Zhao 0003, Ran Tao 0003, Wei Li 0032, Heng-Chao Li 0001, Qian Du 0001, Wenzi Liao, Wilfried Philips
IEEE Trans. Geosci. Remote. Sens.4
2019 Hyperspectral Unmixing Based on Sparsity-Constrained Nonnegative Matrix Factorization with Adaptive Total Variation
abstract
Hyperspectral unmixing is a critical processing step for many remote sensing applications. Nonnegative matrix factorization (NMF) has drawn extensive attention in hyperspectral image analysis recently. Considering that the abundance matrix is generally sparse and smooth, we propose a sparsity-constrained NMF with adaptive total variation (SNMF-ATV) algorithm for hyperspectral unmixing. Specifically, the ATV could promote the smoothness of the estimated abundances while avoid the staircase effect caused by TV model. The comparison with other unmixing methods on both synthetic and real data sets demonstrates the effectiveness and superiority of the proposed SNMF-ATV algorithm with regard to the other considered methods.
Xin-Ru Feng, Heng-Chao Li 0001, Rui Wang 0090
IGARSS2
2019 Semi-Supervised Classification of Polarimetric SAR Images Using Markov Random Field and Two-Level Wishart Mixture Model
abstract
In this work, we propose a semi-supervised method for classification of polarimetric synthetic aperture radar (PolSAR) images. In the proposed method, a 2-level mixture model is constructed by associating each component density with a unique Wishart mixture model (instead of a single Wishart distribution as that in the conventional Wishart mixture model). This modeling scheme facilitates the accurate description of data for the categories, each of which includes multiple subcategories. The learning algorithm for the proposed model is developed based on variational inference and all the update equations are obtained in closed form. In the learning algorithm, the spatial interdependencies are incorporated by imposing a Markov random field prior on the indicator variable to alleviate the speckle effect on the classification results. The experimental results demonstrate the improved performance of the proposed method compared with the unsupervised version and supervised version of the proposed model as well as an existing method for semi-supervised classification.
Wenzi Liao, Heng-Chao Li 0001, Rui Wang 0090, Wilfried Philips
IGARSS3
2019 Data Augmentation for Hyperspectral Image Classification With Deep CNN
abstract
Convolutional neural network (CNN) has been widely used in hyperspectral imagery (HSI) classification. Data augmentation is proven to be quite effective when training data size is relatively small. In this letter, extensive comparison experiments are conducted with common data augmentation methods, which draw an observation that common methods can produce a limited and up-bounded performance. To address this problem, a new data augmentation method, named as pixel-block pair (PBP), is proposed to greatly increase the number of training samples. The proposed method takes advantage of deep CNN to extract PBP features, and decision fusion is utilized for final label assignment. Experimental results demonstrate that the proposed method can outperform the existing ones.
Wei Li 0032, Chen Chen 0001, Mengmeng Zhang 0005, Heng-Chao Li 0001, Qian Du 0001
IEEE Geosci. Remote. Sens. Lett.4
2019 Unsupervised Change Detection of SAR Images Based on Variational Multivariate Gaussian Mixture Model and Shannon Entropy
abstract
In this letter, we propose an unsupervised change detection method for synthetic aperture radar (SAR) images based on variational multivariate Gaussian mixture model (MGMM) and Shannon entropy. First, the difference features are generated from the Gabor wavelet transform of two SAR images. In variational inference framework, the variational MGMM is first introduced to implement accurate modeling for the data distribution of difference features and to output responsibilities. Subsequently, spatial information is explored on the responsibilities to yield thecontextual responsibilitiesfor improving the accuracy and reliability of change detection. Then,a posterioriprobabilities of the changed and unchanged classes are derived from thecontextual responsibilities, and Shannon entropy, being directly related to the classification error rate, is proposed to determine the optimal index integer. Finally, the binary change mask is achieved by separating the pixels into the changed and unchanged classes. The experiments on three pairs of SAR images for describing urban sprawl and water bodies demonstrate the effectiveness of the proposed method.
Gang Yang 0006, Heng-Chao Li 0001, Wen Yang 0001, Kun Fu 0001, Yong-Jian Sun, William J. Emery
IEEE Geosci. Remote. Sens. Lett.2
2019 Bayesian estimation of generalized Gamma mixture model based on variational EM algorithm
Heng-Chao Li 0001, Kun Fu 0001, Fan Zhang 0007, Mihai Datcu, William J. Emery
Pattern Recognit.2
2019 Robust Weighting Nearest Regularized Subspace Classifier for PolSAR Imagery
abstract
Polarimetric synthetic aperture radar (PolSAR) imagery classification is an important part of SAR data interpretation. The number of available labeled samples limits the applications of supervised classifiers. In order to solve this issue, the representation based classification algorithms have been widely used. Usually, PolSAR image features are extracted by various methods, and their divergence is very significant. In the data representation based methods, the feature divergence is ignored in the distance metric, thus the different features have the same metric contributions. In this letter, we propose a robust weighting nearest regularized subspace (NRS) method, which introduces the robust statistics to construct the weights of distance metric according to the feature divergence. This method can increase the representation ability of the training samples by the weighted calculation of the biasing Tikhonov matrix. The experimental results show that the weighted distance metric can boost the original NRS classifier by 1.5%, and prove that the feature divergence should be taken into account in the data representation process.
Fan Zhang 0007, Qiang Yin 0001, Heng-Chao Li 0001
IEEE Signal Process. Lett.4
2019 Unsupervised Change Detection Based on a Unified Framework for Weighted Collaborative Representation With RDDL and Fuzzy Clustering
abstract
In this paper, we propose a novel unsupervised change detection method of remote sensing (RS) images based on a unified framework for weighted collaborative representation (WCR) with robust deep dictionary learning (RDDL) and fuzzy clustering. Specifically, WCR is employed to collaboratively represent neighborhood features with lower computational complexity, for which the RDDL model is built to learn more effective and representative overcomplete dictionary and enhance the robustness against the noise and outliers. Meanwhile, in order to make the resulting collaborative coefficients more beneficial for clustering, the unified framework for WCR with RDDL and fuzzy clustering is designed. By doing so, our framework not only precludes the utilization of third-party clustering algorithm, but also achieves better detection performance. Subsequently, the spatial constraint is enforced on the membership matrix to yield the updated one for further improving the accuracy of change detection. Finally, a binary change mask (CM) is achieved by assigning the pixels into the changed and unchanged classes. Experiments are performed on five pairs of RS images, and experimental results demonstrate the effectiveness of the proposed method.
Gang Yang 0006, Heng-Chao Li 0001, Wei-Ye Wang, Wen Yang 0001, William J. Emery
IEEE Trans. Geosci. Remote. Sens.2
2019 Variational Textured Dirichlet Process Mixture Model With Pairwise Constraint for Unsupervised Classification of Polarimetric SAR Images
abstract
This paper proposes an unsupervised classification method for multilook polarimetric synthetic aperture radar (Pol-SAR) data. The proposed method simultaneously deals with the heterogeneity and incorporates the local correlation in PolSAR images. Specifically, within the probabilistic framework of the Dirichlet process mixture model (DPMM), an observed PolSAR data point is described by the multiplication of a Wishartdistributed component and a class-dependent random variable (i.e., the textual variable). This modeling scheme leads to the proposed textured DPMM (tDPMM), which possesses more flexibility in characterizing PolSAR data in heterogeneous areas and from high-resolution images due to the introduction of the classdependent texture variable. The proposed tDPMM is learned by solving an optimization problem to achieve its Bayesian inference. With the knowledge of this optimization-based learning, the local correlation is incorporated through the pairwise constraint, which integrates an appropriate penalty term into the objective function so as to encourage the neighboring pixels to fall into the same category and to alleviate the "salt-and-pepper" classification appearance.We develop the learning algorithm with all the closed-form updates. The performance of the proposed method is evaluated with both low-resolution and high-resolution PolSAR images, which involve homogeneous, heterogeneous, and extremely heterogeneous areas. The experimental results reveal that the class-dependent texture variable is beneficial to PolSAR image classification and the pairwise constraint can effectively incorporate the local correlation in PolSAR images.
Heng-Chao Li 0001, Wenzi Liao, Wilfried Philips, William J. Emery
IEEE Trans. Image Process.2
2018 Hyperspectral Image Classification Based on Capsule Network
abstract
In this paper, we propose two novel classification frameworks for hyperspectral image (HSI) based on capsule network (CapsNet), which could address the drawbacks of convolutional neural network (CNN) and problem of limited training samples by introducing affine transformation matrix. Specifically, the proposed framework first performs the classification of HSI based on spectral information. Second, considering the importance of spatial information for HSI processing, we integrate the spatial and spectral information into the proposed framework to further improve the classification performance. Experimental results on real HSI data demonstrate the effectiveness of the proposed framework.
Wei-Ye Wang, Heng-Chao Li 0001, Lei Pan 0003, Gang Yang 0006, Qian Du 0001
IGARSS2
2018 Deep Semi-Nonnegative Matrix Factorization Based Unsupervised Change Detection of Remote Sensing Images
abstract
In the paper, an unsupervised change detection method for remote sensing (RS) images based on deep semi-nonnegative matrix factorization (semi-NMF) is proposed. Firstly, the difference image is generated in different ways, depending on the types of input images. Then principal component analysis (PCA) is applied on the difference image to form the feature matrix X for improving the capability against various noise. In order to exploit more useful information from the resulting feature matrix, deep semi-NMF is introduced to factorize X into L+1 factors consisting of L nonrestricted matrices {Fl}l=1Land nonnegative cluster indicator matrix GL. Finally, the binary change mask (CM) is generated by assigning the pixels into changed and unchanged classes according to maximum criterion. The experimental results on two pairs of multitemporal RS images demonstrate the effectiveness of the proposed method.
Gang Yang 0006, Heng-Chao Li 0001, Wen Yang 0001, William J. Emery
IGARSS2
2018 Nuclear norm-based matrix regression preserving embedding for face recognition
Yangjun Deng, Heng-Chao Li 0001, Qi Wang 0009, Qian Du 0001
Neurocomputing2
2018 Modified Tensor Locality Preserving Projection for Dimensionality Reduction of Hyperspectral Images
abstract
By considering the cubic nature of hyperspectral image (HSI) to address the issue of the curse of dimensionality, we have introduced a tensor locality preserving projection (TLPP) algorithm for HSI dimensionality reduction and classification. The TLPP algorithm reveals the local structure of the original data through constructing an adjacency graph. However, the hyperspectral data are often susceptible to noise, which may lead to inaccurate graph construction. To resolve this issue, we propose a modified TLPP (MTLPP) via building an adjacency graph on a dual feature space rather than the original space. To this end, the region covariance descriptor is exploited to characterize a region of interest around each hyperspectral pixel. The resulting covariances are the symmetric positive definite matrices lying on a Riemannian manifold such that the Log-Euclidean metric is utilized as the similarity measure for the search of the nearest neighbors. Since the defined covariance feature is more robust against noise, the constructed graph can preserve the intrinsic geometric structure of data and enhance the discriminative ability of features in the low-dimensional space. The experimental results on two real HSI data sets validate the effectiveness of our proposed MTLPP method.
Yangjun Deng, Heng-Chao Li 0001, Lei Pan 0003, Li-Yang Shao, Qian Du 0001, William J. Emery
IEEE Geosci. Remote. Sens. Lett.2
2018 Hyperspectral Image Reconstruction by Latent Low-Rank Representation for Classification
abstract
To effectively reduce the spectral variation that degrades classification performance, a novel low-rank subspace recovery method based on latent low-rank representation (LatLRR) is proposed for hyperspectral images in this letter. Different from the robust principal component analysis, LatLRR focuses on exploring the low-rank property from the perspective of row space and column space simultaneously through the low-rank regularization on their corresponding coefficient matrix. Following that, the self-expressiveness-based reconstruction is adopted to recover the intrinsic data from row and column spaces. More accurate subspace structure can be successfully preserved both in spectral domain and spatial domain; meanwhile, the robustness to noise is improved. Experimental results on two hyperspectral data sets demonstrate the effectiveness of the proposed method.
Lei Pan 0003, Heng-Chao Li 0001, Yong-Jian Sun, Qian Du 0001
IEEE Geosci. Remote. Sens. Lett.2
2018 Tensor Low-Rank Discriminant Embedding for Hyperspectral Image Dimensionality Reduction
abstract
Recently, low-rank embedding (LRE) has yielded satisfactory results in dimensionality reduction (DR), for which low-rank representation and projection learning are integrated into one model to generate robust low-dimensional features. However, LRE requires to convert samples into vectors even if the data naturally appear in high-order form. Furthermore, LRE fails to take the label information into consideration. To address these problems, this paper proposes a novel supervised DR method based on multilinear algebra, i.e., the algebra of tensors. By the motivation of extending LRE into tensor space and simultaneously combining the tensor discriminant analysis, we establish tensor low-rank discriminant embedding (TLRDE) model for hyperspectral image (HSI) DR. The model of TLRDE is solved by an alternative iteration algorithm, whose convergence is also mathematically proven. The proposed TLRDE method employs the tensor representation to preserve the intrinsic geometrical structure, uses low-rank reconstruction to uncover the potential relationship among the data points, and combines label information to enhance the discriminability of features. Moreover, the proposed TLRDE does not suffer from the small sample size problem. The experimental results on three real HSI data sets validate the effectiveness of our proposed TLRDE method.
Yangjun Deng, Heng-Chao Li 0001, Kun Fu 0001, Qian Du 0001, William J. Emery
IEEE Trans. Geosci. Remote. Sens.2
2018 Hyperspectral Unmixing Using Sparsity-Constrained Deep Nonnegative Matrix Factorization With Total Variation
abstract
Hyperspectral unmixing is an important processing step for many hyperspectral applications, mainly including: 1) estimation of pure spectral signatures (endmembers) and 2) estimation of the abundance of each endmember in each pixel of the image. In recent years, nonnegative matrix factorization (NMF) has been highly attractive for this purpose due to the nonnegativity constraint that is often imposed in the abundance estimation step. However, most of the existing NMF-based methods only consider the information in a single layer while neglecting the hierarchical features with hidden information. To alleviate such limitation, in this paper, we propose a new sparsity-constrained deep NMF with total variation (SDNMF-TV) technique for hyperspectral unmixing. First, by adopting the concept of deep learning, the NMF algorithm is extended to deep NMF model. The proposed model consists ofpretraining stageandfine-tuning stage, where the former pretrains all factors layer by layer and the latter is used to reduce the total reconstruction error. Second, in order to exploit adequately the spectral and spatial information included in the original hyperspectral image, we enforce two constraints on the abundance matrix. Specifically, the$L_{1/2}$constraint is adopted, since the distribution of each endmember is sparse in the 2-D space. The TV regularizer is further introduced to promote piecewise smoothness in abundance maps. For the optimization of the proposed model, multiplicative update rules are derived using the gradient descent method. The effectiveness and superiority of the SDNMF-TV algorithm are demonstrated by comparing with other unmixing methods on both synthetic and real data sets.
Xin-Ru Feng, Heng-Chao Li 0001, Jun Li 0009, Qian Du 0001, Antonio Plaza, William J. Emery
IEEE Trans. Geosci. Remote. Sens.2
2018 Unsupervised Classification of Multilook Polarimetric SAR Data Using Spatially Variant Wishart Mixture Model with Double Constraints
abstract
This paper addresses the unsupervised classification problems for multilook Polarimetric synthetic aperture radar (PolSAR) images by proposing a patch-level spatially variant Wishart mixture model (SVWMM) with double constraints. We construct this model by jointly modeling the pixels in a patch (rather than an individual pixel) so as to effectively capture the local correlation in the PolSAR images. More importantly, a responsibility parameter is introduced to the proposed model, providing not only the possibility to represent the importance of different pixels within a patch but also the additional flexibility for incorporating the spatial information. As such, double constraints are further imposed by simultaneously utilizing the similarities of the neighboring pixels, respectively, defined on two different parameter spaces (i.e., the hyperparameter in the posterior distribution of mixing coefficients and the responsibility parameter). Furthermore, the variational inference algorithm is developed to achieve effective learning of the proposed SVWMM with the closed-form updates, facilitating the automatic determination of the cluster number. Experimental results on several PolSAR data sets from both airborne and spaceborne sensors demonstrate that the proposed method is effective and it enables better performances on unsupervised classification than the conventional methods.
Wenzi Liao, Heng-Chao Li 0001, Kun Fu 0001, Wilfried Philips
IEEE Trans. Geosci. Remote. Sens.3
2018 Spectral-Spatial Weighted Sparse Regression for Hyperspectral Image Unmixing
abstract
Spectral unmixing aims at estimating the fractional abundances of a set of pure spectral materials (endmembers) in each pixel of a hyperspectral image. The wide availability of large spectral libraries has fostered the role of sparse regression techniques in the task of characterizing mixed pixels in remotely sensed hyperspectral images. A general solution for sparse unmixing methods consists of using the l2regularizer to control the sparsity, resulting in a very promising performance but also suffering from sensitivity to large and small sparse coefficients. A recent trend to address this issue is to introduce weighting factors to penalize the nonzero coefficients in the unmixing solution. While most methods for this purpose focus on analyzing the hyperspectral data by considering the pixels as independent entities, it is known that there exists a strong spatial correlation among features in hyperspectral images. This information can be naturally exploited in order to improve the representation of pixels in the scene. In order to take advantage of the spatial information for hyperspectral unmixing, in this paper, we develop a new spectral-spatial weighted sparse unmixing (S2WSU) framework, which uses both spectral and spatial weighting factors, further imposing sparsity on the solution. Our experimental results, conducted using both simulated and real hyperspectral data sets, illustrate the good potential of the proposed S2WSU, which can greatly improve the abundance estimation results when compared with other advanced spectral unmixing methods.
Shaoquan Zhang, Jun Li 0009, Heng-Chao Li 0001, Chengzhi Deng, Antonio Plaza
IEEE Trans. Geosci. Remote. Sens.3
2017 Tensor locality preserving projection for hyperspectral image classification
abstract
By considering the cubic nature of hyperspectral image (HSI) and to address the issue of the curse of dimensionality, we introduce a tensor locality preserving projection (TLPP) algorithm for HSI classification. TLPP has been proved to be effective in preserving the geometrical structure of data for dimensionality reduction. More importantly, data can be taken directly in the form of a tensor of arbitrary order as input, such that the damage to sample's geometrical structure is avoided during vectorizing. For the HSI classification, TLPP can effectively embed both spatial structure and spectral information into low-dimensional space simultaneously by a series of projection matrices trained for each mode of input samples. The experimental results on the AVIRIS hyperspectral image confirm the effectiveness of TLPP.
Yangjun Deng, Heng-Chao Li 0001, Lei Pan 0003, William J. Emery
IGARSS2
2017 Non-negative matrix factorization with mixture of Itakura-Saito divergence for SAR images
abstract
Synthetic aperture radar (SAR) data are becoming more and more accessible and have been widely used in many applications. To effectively and efficiently represent multiple SAR images, we propose the mixture of Itakura-Saito (IS) divergence for non-negative matrix factorization (NMF) to perform the dimension reduction. Our proposed method incorporates the unit-mean Gamma mixture model into the NMF to model the multiplicative noise. To obtain the closed-form update equations as much as possible, we approximate the log-likelihood function with its lower bound. Finally, we apply Expectation-Maximization (EM) algorithm to estimate the parameters, resulting in the closed-form multiplicative update rules for the two matrix factors. Experimental results on real SAR dataset demonstrate the effectiveness of the proposed method and its applicability to post applications (e.g., classification) with improved performances over the conventional dimension reduction methods.
Wenzi Liao, Heng-Chao Li 0001, Wilfried Philips
IGARSS3
2017 Spatial weighted sparse regression for hyperspectral image unmixing
abstract
Sparse unmixing of hyperspectral data is an important technique which aims at estimating the fractional abundances of endmembers (pure spectral components). It is well known that enforcing sparseness becomes a necessary process in sparse unmixing methods. To better exploit the sparsity in hyperspectral imagery, a double reweighted sparse unmixing algorithm has been proposed. However, it focusses on analyzing the hyperspectral data without fully incorporating the spatial information. To address this limitation, a spatial weighted sparse unmixing (SWSU) algorithm is proposed in this paper, which can take full advantage of the spatial information and further enhance the sparsity of the abundances. This is done by incorporating local neighborhood weights into the double reweighted sparse unmixing formulation. Experimental results on simulated hyperspectral data sets illustrate the good potential of the spatial weighted strategy for sparse unmixing introduced in this paper, which can greatly improve abundance estimation results.
Shaoquan Zhang, Jun Li 0009, Javier Plaza, Heng-Chao Li 0001, Antonio Plaza
IGARSS4
2017 ℋ Distribution for Multilook Polarimetric SAR Data
abstract
Polarimetric synthetic aperture radar (PolSAR) is an advanced imaging radar system, for which the acquired data provide not only the information of each channel but also the correlation between channels. To fully utilize and accurately model the multilook PolSAR data, a novel compound distribution, named the H distribution, is proposed based on the generalized Fisher distribution (GFD). Specifically, the GFD introduces a power parameter to the ordinary Fisher distribution. With one more free parameter, the GFD is flexible and versatile enough to characterize different kinds of texture. Then, by assuming the generalized-Fisher-distributed texture and the Wishart-distributed speckle, the H distribution is derived, whose closed-form expression is obtained with the help of Fox's H-function. As such, the H distribution has a compact form and is conveniently applied to practical problems, such as modeling and classification of PolSAR data. The effectiveness of this method is tested by modeling the multilook PolSAR data and performing image classification. The experimental results demonstrate that the H distribution is a flexible and effective way to model multilook PolSAR data.
Heng-Chao Li 0001, Xian Sun 0001, William J. Emery
IEEE Geosci. Remote. Sens. Lett.2
2017 Hyperspectral Image Classification via Low-Rank and Sparse Representation With Spectral Consistency Constraint
abstract
In this letter, a low-rank and sparse representation classifier with a spectral consistency constraint (LRSRC-SCC) is proposed. Different from the SRC that represents samples individually, LRSRC-SCC reconstructs samples jointly and is able to capture the local and global structures simultaneously. In this proposed classifier, an adaptive spectral constraint is imposed on both the low-rank and sparse terms so as to better reveal the data structure and enhance its discriminative power. In addition, the alternating direction method is introduced to solve the underlying minimization problem, in which, more importantly, the subobjective function associated with the low-rank term is optimized based on the rank equivalence between a matrix and its Gram matrix, resulting in a closed-form solution. Finally, LRSRC-SCC is extended to LRSRC-SCCE for fully exploiting the spatial information. Experimental results on two hyperspectral data sets demonstrate that the proposed LRSRC-SCC and LRSRC-SCCE methods outperform some state-of-the-art methods.
Lei Pan 0003, Heng-Chao Li 0001, Hua Meng 0001, Wei Li 0032, Qian Du 0001, William J. Emery
IEEE Geosci. Remote. Sens. Lett.2
2017 Hyperspectral Unmixing Using Double Reweighted Sparse Regression and Total Variation
abstract
Spectral unmixing is an important technique in hyperspectral image applications. Recently, sparse regression has been widely used in hyperspectral unmixing, but its performance is limited by the high mutual coherence of spectral libraries. To address this issue, a new sparse unmixing algorithm, called double reweighted sparse unmixing and total variation (TV), is proposed in this letter. Specifically, the proposed algorithm enhances the sparsity of fractional abundances in both spectral and spatial domains through the use of double weights, where one is used to enhance the sparsity of endmembers in spectral library, and the other is introduced to improve the sparsity of fractional abundances. Moreover, a TV-based regularization is further adopted to explore the spatial-contextual information. As such, the simultaneous utilization of both double reweighted l1minimization and TV regularizer can significantly improve the sparse unmixing performance. Experimental results on both synthetic and real hyperspectral data sets demonstrate the effectiveness of the proposed algorithm both visually and quantitatively.
Rui Wang 0090, Heng-Chao Li 0001, Aleksandra Pizurica, Jun Li 0009, Antonio Plaza, William J. Emery
IEEE Geosci. Remote. Sens. Lett.2
2017 Single Image Super-Resolution Reconstruction Technique based on A Single Hybrid Dictionary
Chanzi Liu, Qingchun Chen, Heng-Chao Li 0001
Multim. Tools Appl.3
2017 Discriminant Analysis of Hyperspectral Imagery Using Fast Kernel Sparse and Low-Rank Graph
abstract
Due to the high-dimensional characteristic of hyperspectral images, dimensionality reduction (DR) is an important preprocessing step for classification. Recently, sparse and low-rank graph-based discriminant analysis (SLGDA) has been developed for DR of hyperspectral images, for which the properties of sparsity and low-rankness are simultaneously exploited to capture both local and global structures. However, SLGDA may not achieve satisfactory results when handling complex data with nonlinear nature. To address this problem, this paper presents two kernel extensions of SLGDA. In the first proposed classical kernel SLGDA (cKSLGDA), the kernel trick is exploited to implicitly map the original data into a high-dimensional space. With a totally different perspective, we further propose a Nyström-based kernel SLGDA (nKSLGDA) by constructing a virtual kernel space by the Nyström method, in which virtual samples can be explicitly obtained from the original data. Both cKSLGDA and nKSLGDA can achieve more informative graphs than SLGDA, and offer superiority over other state-of-the-art DR methods. More importantly, the nKSLGDA can outperform cKSLGDA with much lower computational cost.
Lei Pan 0003, Heng-Chao Li 0001, Wei Li 0032, Guangning Wu, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.2
2016 A multitexture model of multilook polarimetric SAR data based on generalized gamma distribution
abstract
In this paper, we extend the product model from common texture to the case of multitexture. With the generalized Gamma distribution (GΓD) being the texture distribution, a multitexture model for multilook polarimetric SAR data is proposed and the estimation of parameters is performed based on the method of matrix log-cumulant. As far as the diagonal random texture matrix is concerned, the textures for the co-polarized and cross-polarized channels are treated as mutually independent random variables instead of a common variable. Based on this treatment and with the help of the Fox's H-function, the distribution of the proposed model is evaluated in closed form. Moreover, it is also capable of dealing with data covering various scenes due to the introduction of the highly flexible GΓD. The classification result is illustrated to verify the effectiveness of the proposed model.
Heng-Chao Li 0001
IGARSS1
2016 Locality constrained low-rank representation for hyperspectral image classification
abstract
This paper addresses the problem of hyperspectral image classification with the low-rank representation (LRR) which has been widely applied in computer vision and pattern recognition. As is known, it has been proved to be effective in subspace segmentation under the assumption that all the subspaces are mutually independent. Nevertheless, in practical applications, this assumption could hardly be guaranteed. In this paper, to sidestep this limitation, we simultaneously exploit the spectral similarity and spatial information of pixels to design a local constraint as the regularizer of LRR, which is referred to as the locality constrained LRR (LCLRR). The experimental results on the AVIRIS hyperspectral image confirm the effectiveness of our proposed method.
Lei Pan 0003, Heng-Chao Li 0001
IGARSS2
2016 Double reweighted sparse regression for hyperspectral unmixing
abstract
Spectral unmixing is an important technology in hyperspectral image applications. Recently, sparse regression is widely used in hyperspectral unmixing. This paper proposes a double reweighted sparse regression method for hyperspectral unmixing. The proposed method enhances the sparsity of abundance fraction in both spectral and spatial domains through double weights, in which one is used to enhance the sparsity of endmembers in the spectral library, and the other to improve the sparseness of abundance fraction of every material. Experimental results on both synthetic and real hyperspectral data sets demonstrate effectiveness of the proposed method both visually and quantitatively.
Rui Wang 0090, Heng-Chao Li 0001, Wenzi Liao, Aleksandra Pizurica
IGARSS2
2016 An mmWave Wireless Communication and Radar Detection Integrated Network for Railways
abstract
With large available continuous bandwidth, millimeter wave (mmWave) bands hold promise as a carrier frequency for fifth generation (5G) wireless communications. Moreover, mmWave bands also play an important role in radar detections. Based on this observation, we propose an mmWave wireless communication and radar detection integrated network architecture for railways, not only to increase the capacity of railway wireless communication systems, but also to realize train operation environment detection to enhance train operation safety. To overcome the aggravated path loss in mmWave bands, directional beamforming is generally used to concentrate signal radiation energy both in wireless communications and radar detections. Nevertheless, for wireless communications, the signaling blind zone of directional beamforming makes it less effective in wireless link establishment and maintenance. Therefore, in this proposed integrated network, two frequency bands are employed, where omnidirectionally radiated licensed lower frequency bands carry critical signaling and data, and mmWave bands are time-division multiplexed to transmit large-volume communication data for trains, or to perform environment detection for enhancing train operation safety. To mitigate handovers and to increase baseband processing resource utilization in railway scenarios, the proposed integrated network is deployed based on the cloud radio access network (C-RAN) architecture, where both licensed lower frequency band radio remote units (RRUs) and mmWave band RRUs are connected to a building baseband unit (BBU) pool through high-speed backhauls. Besides, physical (PHY) frame structures and up/downlink communication signaling procedures are designed for this proposed integrated network. Performance analysis results have demonstrated that the proposed integrated network can highly increase the capacity for railway wireless communication systems and achieve high distance and angular resolution for radar detections.
Li Yan 0002, Xuming Fang, Heng-Chao Li 0001, Chao Li 0013
VTC Spring3
2016 Digital Watermarking Processing Technique Based on Overcomplete Dictionary
abstract
A novel sparse domain-based information hiding framework is proposed in this paper to attach the watermarking signal to the most significant sparse components of the host signal over the pre-defined overcomplete dictionary. The adaptive sparse domain can be utilized to embed watermarking logo with better security and robustness. This can be realized owing to the fact that, not only the sparse domain can be customized from the given samples, but also the sparse transform coefficients of the original watermarking signal can be embedded, which provides inherent privacy. This paper provides two kinds of methods that embed watermark directly and embed the sparse representation coefficients of watermarking logo, and analyzes the condition of uniqueness of the sparse solution. Experimental results demonstrate the superiority of the proposed sparse domain digital watermarking technique over the traditional frequency domain or spatial domain schemes.
Chanzi Liu, Qingchun Chen, Hongbin Liang, Heng-Chao Li 0001
Int. J. Pattern Recognit. Artif. Intell.4
2016 Unsupervised Learning of Generalized Gamma Mixture Model With Application in Statistical Modeling of High-Resolution SAR Images
abstract
The accurate statistical modeling of synthetic aperture radar (SAR) images is a crucial problem in the context of effective SAR image processing, interpretation, and application. In this paper, a semi-parametric approach is designed within the framework of finite mixture models based on the generalized Gamma distribution in view of its flexibility and compact form. Specifically, we develop a generalized Gamma mixture model to implement an effective statistical analysis of high-resolution SAR images and prove the identifiability of such mixtures. A low-complexity unsupervised estimation method is derived by combining the proposed histogram-based expectation-conditional maximization algorithm and the Figueiredo-Jain algorithm. This results in a numerical maximum-likelihood (ML) estimator that can simultaneously determine the ML estimates of component parameters and the optimal number of mixture components. Finally, the state-of-the-art performance of this proposed method is verified by experiments with a wide range of high-resolution SAR images.
Heng-Chao Li 0001, Vladimir A. Krylov, Pingzhi Fan, Josiane Zerubia, William J. Emery
IEEE Trans. Geosci. Remote. Sens.1
2015 Hyperspectral image classification by sparse representation with nonlocal adaptive dictionary
abstract
In this paper, a novel nonlocal dictionary learning method is proposed for sparse-representation-based classification (SRC) to label high-dimensional hyperspectral imagery (HSI). In SRC, the conventional dictionary is constructed using all of the training pixels, which is inefficient due to the high-dimension low-sample-size classification problem. In this paper, we construct the dictionary by adding more appropriate pixels into the dictionary. Specifically, we select the supplementaries from the neighboring pixels of the original training pixels based on the assumption that the adjacent pixels belong to the same class with a high probability, and propose an estimative function for the selection. Furthermore, this estimative function is adopted again to select the components of signal matrix in joint sparsity model (JSM) to improve classification accuracy. Experimental results have shown that the dictionary optimized using our method can achieve better classification results with substantially expanded dictionary size than only using the training pixels.
Heng-Chao Li 0001
IGARSS2
2015 Key techniques for 5G wireless communications: network architecture, physical layer, and MAC layer perspectives
Zheng Ma 0001, Zhengquan Zhang, Zhiguo Ding 0001, Pingzhi Fan, Heng-Chao Li 0001
Sci. China Inf. Sci.5
2015 Gabor Feature Based Unsupervised Change Detection of Multitemporal SAR Images Based on Two-Level Clustering
abstract
In this letter, we propose a simple yet effective unsupervised change detection approach for multitemporal synthetic aperture radar images from the perspective of clustering. This approach jointly exploits the robust Gabor wavelet representation and the advanced cascade clustering. First, a log-ratio image is generated from the multitemporal images. Then, to integrate contextual information in the feature extraction process, Gabor wavelets are employed to yield the representation of the log-ratio image at multiple scales and orientations, whose maximum magnitude over all orientations in each scale is concatenated to form the Gabor feature vector. Next, a cascade clustering algorithm is designed in this discriminative feature space by successively combining the first-level fuzzy c-means clustering with the second-level nearest neighbor rule. Finally, the two-level combination of the changed and unchanged results generates the final change map. Experimental results are presented to demonstrate the effectiveness of the proposed approach.
Heng-Chao Li 0001, Turgay Çelik 0001, Nathan Longbotham, William J. Emery
IEEE Geosci. Remote. Sens. Lett.1
2014 Unsupervised change detection of remote sensing images based on semi-nonnegative matrix factorization
abstract
In this paper, we propose an unsupervised change detection approach for the multitemporal remote sensing images based on semi-nonnegative matrix factorization (semi-NMF). Specifically, the multitemporal source images, acquired at the same geographical area but at two different time instances, are first utilized to generate the difference image. Then, feature vector is created for each pixel of the difference image in such a way that its corresponding h × h block data is projected on the generated eigenvector space by principal component analysis (PCA), which is further arranged as a column vector to form a feature-by-item data matrix X. Next, we implement semi-NMF to factorize X into two nonnegative factors (i.e., the basis matrix F and the coefficient matrix G). Finally, the change detection is achieved by discriminating each column of GTaccording to the maximum criterion. Experimental results verify the feasibility and effectiveness of the proposed approach.
Heng-Chao Li 0001, Nathan Longbotham, William J. Emery
IGARSS1
2014 Nonlocal similarity regularization for sparse hyperspectral unmixing
abstract
This paper is concerned with semisupervised hyperspectral unmixing using a nonlocal similarity prior on the abundance images. To this end, the nonlocal self-similarity regularization is incorporated into the classical sparse regression formula to propose a new model for hyperspectral sparse unmixing. The rationale is the idea that there are many nonlocal similar patches to the given patch in the abundance images. The effectiveness of the proposed algorithm is illustrated using the synthetic and real data sets.
Rui Wang 0090, Heng-Chao Li 0001
IGARSS2
2013 FRFT-based improved algorithm of unsupervised change detection in SAR images via PCA and K-means clustering
abstract
This paper presents an improved algorithm of unsupervised change detection technique by taking the same low-order fractional Fourier transform (FRFT) on multitemporal images acquired on the same geographical area but at different time instances, then generates the difference image by the absolute log-ratio operator. In order to acquire the eigenvector space, we perform principal component analysis (PCA) on m × m nonoverlapping difference image blocks. The feature vectors are extracted using m × m data blocks projection onto eigenvector space. The change detection map is generated by clustering the feature vectors using k-means algorithm into two disjoint classes: changed and unchanged. The final results obtained by the improved algorithm exhibited lower error than its preexistence.
Yongqiang Cheng 0004, Heng-Chao Li 0001, Turgay Çelik 0001, Fan Zhang 0007
IGARSS2
2013 Bayesian Wavelet Shrinkage With Heterogeneity-Adaptive Threshold for SAR Image Despeckling Based on Generalized Gamma Distribution
abstract
Synthetic aperture radar (SAR) images are inherently affected by multiplicative speckle noise, which will degrade the human interpretation and computer-aided scene analysis. In this paper, we propose a novel Bayesian multiscale method for SAR image despeckling in the non-homomorphic framework. To address the multiplicative nature, we first make the speckle contribution additive by a linear decomposition. Then, in the stationary wavelet transform domain, a two-sided generalized Gamma distribution (GTD) is introduced as a prior to capture the heavy-tailed nature of wavelet coefficients of the noise-free reflectivity. By exploiting this prior together with a Gaussian likelihood, an analytical wavelet shrinkage function is derived based on maximum a posteriori criteria, which further adopts heterogeneity-adaptive thresholding technique to achieve better estimates of noise-free wavelet coefficients. Moreover, a pilot-signal-assisted strategy is proposed to estimate the parameters of two-sided GTD with the estimator based on second-kind cumulants. Finally, experimental results, carried out on the synthetic and actual SAR images, are given to demonstrate the validity of the proposed despeckling method.
Heng-Chao Li 0001, Wen Hong, Yirong Wu, Pingzhi Fan
IEEE Trans. Geosci. Remote. Sens.1
2012 Relaxed generalized minimum-error thresholoding for unsupervised change detection from SAR amplitude images
abstract
Generalized Kittler and Illingworth minimum-error thresholding (GKIT) algorithm was proposed by G.Moser for change detection in synthetic aperture radar (SAR) images with non-Gussion distribution. In this paper, we present an improved GKIT approach for unsupervised change detection from synthetic aperture radar (SAR) amplitude images by relaxing the demand of the same equivalent number of looks (ENL) in the GKIT approach based on Nakagami model. Experimental results on actual SAR images are given to demonstrate the validity of our proposed method.
Heng-Chao Li 0001
IGARSS2
2012 MCMC estimation of finite generalized gamma mixture model
abstract
Recently, the generalized Gamma distribution (GGD) has proved to be a very efficient model for SAR image processing. In this paper, a fully Bayesian framework is presented for the finite generalized gamma mixture model (GGMM). It considers the cases of known mixture size, as opposed to most previous work on mixture models, the model is estimated using Markov chain Monte Carlo (MCMC) algorithm, this algorithm uses a Gibbs and Metropolis-Hastings sampling, relying on the missing data structure of the mixture model. A Monte Carlo simulation study carried out with the synthetic and real data is performed to demonstrate the algorithm excellent performance.
Yan-Hui Zou, Heng-Chao Li 0001
IGARSS2
2010 An Efficient and Flexible Statistical Model Based on Generalized Gamma Distribution for Amplitude SAR Images
abstract
In the context of synthetic aperture radar (SAR) image processing and applications, the precise modeling of statistical knowledge is a crucial problem. In this paper, an efficient and flexible statistical model, called generalized Gamma Rayleigh (G¿R) distribution, for amplitude SAR images is proposed by assuming a two-sided generalized Gamma distribution for the real and imaginary parts of the complex SAR backscattered signal. It is shown that the Rayleigh and recently proposed generalized Gaussian Rayleigh distributions can be regarded as special cases of G¿R distribution. Considering that the probability density function estimation problem is formulated as a parameter estimation one for the parametric statistical analysis of SAR images, a two-stage estimator based on second-kind cumulants is derived for the parameters of G¿R distribution. Furthermore, experimental results on several actual SAR images are given to demonstrate the validity and flexibility of the proposed model.
Heng-Chao Li 0001, Wen Hong, Yirong Wu, Pingzhi Fan
IEEE Trans. Geosci. Remote. Sens.1
2009 Optimal variable-weight optical orthogonal codes via cyclic difference families
abstract
Variable-weight Optical orthogonal code (OOC) was introduced by G-C Yang for multimedia optical CDMA systems with multiple quality of service (QoS) requirement. In this paper, a construction for optimal variable-weight OOCs via cyclic difference families is given. Several new constructions for cyclic difference families are also given. By using these constructions, new optimal (n,W, 1,Q)-OOCs for 2 ≤ |W| ≤ 4 are constructed.
Heng-Chao Li 0001, Pingzhi Fan, Dianhua Wu, Parampalli Udaya
ISIT1
2007 Texture-Preserving Despeckling of SAR Images Using Evidence Framework
abstract
In this letter, a texture-preserving despeckling algorithm for synthetic aperture radar images using an evidence framework is proposed. The salient aspects of this approach are given as follows. (1) The maximuma posterioriestimate can be guaranteed to converge to the optima by selecting the Gaussian distribution and Gaussian Markov random field model as the likelihood function and prior model, respectively. (2) MacKay's evidence framework can automatically sustain the balance between speckle reduction and texture preservation. (3) We use the Jeffreys prior to perform the second-level inference of the evidence framework. Experimental results are given to demonstrate the validity of the proposed despeckling method.
Heng-Chao Li 0001, Wen Hong, Yirong Wu, Heng-Ming Tai
IEEE Geosci. Remote. Sens. Lett.1
2006 Research of Chaos Theory and Local Support Vector Machine in Effective Prediction of VBR MPEG Video Traffic
Heng-Chao Li 0001, Wen Hong, Yirong Wu, Si-Jie Xu
ICIC (1)1