Wei He 0003

dblp:20/6417-3 · DBLP profile ↗
← Back
83ranked-venue papers
14as first author
60since 2021 · last 2026
0000-0003-3410-0643ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 55 · 8 first-author · 40 since 2021Artificial intelligence and machine learning · 17 · 4 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 4 first-author · 11 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 FreedomDiVe: Task-free image fusion via marginal distribution-based diffusion variational estimation
Zihan Gui, Wei He 0003
Pattern Recognit.6
2026 Learning a Self-Supervised Low-Rank Decomposition Network for Hyperspectral Image Super-Resolution
abstract
Hyperspectral image (HSI) super-resolution, which reconstructs a high-resolution HSI (HR-HSI) through hyperspectral and multispectral image fusion (HMIF) tasks that integrate a low-resolution HSI (LR-HSI) with a high-resolution multispectral image (HR-MSI), has emerged as a promising technique for enhancing spatial–spectral quality. Recently, low-rank representations have demonstrated significant advances in various hyperspectral-related applications, providing an effective solution to HMIF tasks. However, most existing methods rely on model priors to learn the low-rank representation of HSIs, which restricts their adaptability to low-rank variations across different datasets. To address this issue, this paper introduces a self-supervised low-rank decomposition network (SSLRDN) framework specifically designed for HMIF, inspired by the observation that the HR-MSI and HR-HSI of the same scene share highly similar spatial features, whereas different hyperspectral scenes exhibit variations in both spectral and spatial features. In SSLRDN, we develop a self-supervised network to adaptively learn the low-rank decomposition (spectral subspace and spatial coefficients) across different HR-HSIs, overcoming the inefficiency of conventional alternating optimization methods where factor updates fail to mutually promote each other. Given the spatial feature consistency between HR-MSI and HR-HSI, we leverage the rich spatial information from HR-MSI to guide the learning of spatial coefficients in HR-HSI. To enhance the self-supervised learning of spatial coefficient images, we further integrate an externally pre-trained denoiser to improve their estimation accuracy, effectively fusing and mutually promoting both self-supervised and pre-trained learning paradigms. Experimental results show that the proposed method achieves superior performance in both visual quality and quantitative metrics, without requiring pretraining on external datasets.
Yong Chen 0013, Xinfeng Gui, Feiwang Yuan, Wei He 0003, Jinshan Zeng
IEEE Trans. Circuits Syst. Video Technol.5
2026 Toward Complex Backgrounds: A Unified Difference-Aware Decoder for Binary Segmentation
abstract
Binary segmentation is used to distinguish objects of interest from background, and is an active area of convolutional encoder-decoder network research. The current decoders are designed for specific objects based on the common backbones as the encoders, but cannot deal with complex backgrounds. Inspired by the way human eyes detect objects, we propose a new unified dual-branch decoder paradigm, termed the difference-aware decoder, to better explore the differences between foreground and background and to separate objects of interest in optical images. This decoder operates in two stages, leveraging multi-level features from the encoder. In the first stage, coarse detection of foreground objects is achieved by directly utilizing high-level semantic features, mimicking the initial rough observation of human vision. In the second stage, the decoder refines segmentation by exploring differences in low-level features, guided by the coarse map from the first stage. To enhance this process, we introduce two key innovations. First, a difference-aware prototype generation strategy leverages the guide map to extract foreground and background prototypes from high-level features, and calculates the similarity between these prototypes and corresponding representations in low-level feature spaces. Second, an overlapped window cross-level semantic guidance mechanism integrates high-level semantic information into low-level features through channel grouping and multi-scale aligned window pairs, guided by the similarities computed in the first strategy. Together, these innovations significantly enhance the DAD’s ability to discern subtle differences, enabling precise foreground extraction and effectively addressing the challenges of complex and varied backgrounds. To verify the performance of the proposed difference-aware decoder, we choose three well known backbones including ResNet, Res2Net, PVT, and two binary segmentation tasks,i.e., salient object detection, and camouflaged object detection, for comparative experiments. The results demonstrate that the difference-aware decoder can achieve higher accuracy than the other state-of-the-art binary segmentation methods for these tasks. The source code will be available on https://github.com/Henryjiepanli/DAD.
Jiepan Li, Wei He 0003, Fangxiao Lu, Hongyan Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.2
2026 Unsupervised High-Order Implicit Neural Representation With Line Attention for Metal Artifact Reduction
abstract
The presence of metallic implants introduces bright and dark streaks that appear in computed tomography (CT) images, degrading image quality and interfering with medical diagnosis. To reduce these artifacts, deep learning approaches have been applied for metal-corrupted restoration, which usually requires a large amount of simulated degraded-clean pairs for training. To achieve metal artifact reduction (MAR) without reference images, implicit neural representation (INR) has emerged and shown capabilities for image restoration in an unsupervised manner. However, existing INR methods for MAR usually treat the spatial coordinates independently and ignore their correlation, resulting in detail loss and artifacts remaining. In this paper, we propose an INR-based unsupervised MAR framework and design a High-order Line Attention Network to capture local contextual and geometric representations from X-rays, which maps the spatial coordinates into discrete linear attenuation coefficients of imaged objects for artifact-free CT image reconstruction. The second-order feature interaction can effectively improve the spectral bias problems and fit low and high-frequency details of real signals well. The proposed line-attention module with linear complexity can establish global relationships among spatial point tokens from sampled rays. To provide more local contextual information, a multiple local adjacent ray sampling strategy is adopted to compose several sub-fan beams with more context as a training batch. With the help of these components, the unsupervised MAR framework can approximate the implicit continuous function to estimate measurements and generate artifact-free CT images. Simulated and real experiments indicated that the proposed approach achieved superior MAR performance compared with other state-of-the-art methods.
Hongyu Chen 0003, Shaoguang Huang, Wei He 0003, Hongyan Zhang 0001
IEEE Trans. Medical Imaging3
2025 Advancing Chinese Lip Reading through Contextual Enhancement
abstract
Lip reading, also known as visual speech recognition, aims to extract spoken words from silent videos by analyzing lip movements. Most existing methods have been developed primarily for English-language videos and are not optimized for Chinese lip reading. Given that Chinese is rich in homophones, which demand nuanced contextual cues for accurate disambiguation, conventional lip reading techniques tend to experience a significant drop in performance when applied to Chinese. To address these challenges, we propose the Context-Enhanced Visual Speech Recognition Network (ConVSR), specifically designed to improve Chinese lip reading. ConVSR incorporates intermediate connectionist temporal classification (inter CTC) residual modules to provide enhanced CTC supervision, improving the model’s contextual comprehension. Additionally, a Bi-Transformer Decoder is employed to simultaneously capture both past and future contextual information, thereby strengthening the model’s semantic understanding. Our approach achieves a $74.40 \%$ character error rate (CER) on the CAS-VSR-MOV20 dataset in track 1 of the Mandarin Audio-Visual Speech Recognition Challenge (MAVSR) 2025, securing first place and underscoring the effectiveness of ConVSR in advancing Chinese lip reading accuracy.
Ruoyao Xue, Jiepan Li, Zhehui Wu, Wei He 0003
FG4
2025 MP-HSIR: A Multi-Prompt Framework for Universal Hyperspectral Image Restoration
abstract
Hyperspectral images (HSIs) often suffer from diverse and unknown degradations during imaging, leading to severe spectral and spatial distortions. Existing HSI restoration methods typically rely on specific degradation assumptions, limiting their effectiveness in complex scenarios. In this paper, we propose \textbf{MP-HSIR}, a novel multi-prompt framework that effectively integrates spectral, textual, and visual prompts to achieve universal HSI restoration across diverse degradation types and intensities. Specifically, we develop a prompt-guided spatial-spectral transformer, which incorporates spatial self-attention and a prompt-guided dual-branch spectral self-attention. Since degradations affect spectral features differently, we introduce spectral prompts in the local spectral branch to provide universal low-rank spectral patterns as prior knowledge for enhancing spectral reconstruction. Furthermore, the text-visual synergistic prompt fuses high-level semantic representations with fine-grained visual features to encode degradation information, thereby guiding the restoration process. Extensive experiments on 9 HSI restoration tasks, including all-in-one scenarios, generalization tests, and real-world cases, demonstrate that MP-HSIR not only consistently outperforms existing all-in-one methods but also surpasses state-of-the-art task-specific approaches across multiple tasks. The code and models are available at https://github.com/ZhehuiWu/MP-HSIR.
Zhehui Wu, Yong Chen 0013, Naoto Yokoya, Wei He 0003
ICCV4
2025 CHOICE: Benchmarking the Remote Sensing Capabilities of Large Vision-Language Models
abstract
The rapid advancement of Large Vision-Language Models (VLMs), both general-domain models and those specifically tailored for remote sensing, has demonstrated exceptional perception and reasoning capabilities in Earth observation tasks. However, a benchmark for systematically evaluating their capabilities in this domain is still lacking. To bridge this gap, we propose CHOICE, an extensive benchmark designed to objectively evaluate the hierarchical remote sensing capabilities of VLMs. Focusing on 2 primary capability dimensions essential to remote sensing: perception and reasoning, we further categorize 6 secondary dimensions and 23 leaf tasks to ensure a well-rounded assessment coverage. CHOICE guarantees the quality of all 10,507 problems through a rigorous process of data collection from 50 globally distributed cities, question construction, and quality control. The newly curated data and the format of multiple-choice questions with definitive answers allow for an objective and straightforward performance assessment. Our evaluation of 3 proprietary and 21 open-source VLMs highlights their critical limitations within this specialized context. We hope that CHOICE will serve as a valuable resource and offer deeper insights into the challenges and potential of VLMs in the field of remote sensing. Code and dataset are available at this https URL.
Xiao An, Jiaxing Sun 0001, Zihan Gui, Wei He 0003
NeurIPS4
2025 Wavelength- and Depth-Aware Deep Image Prior for Blind Hyperspectral Imagery Deblurring with Coarse Depth Guidance
abstract
Hyperspectral imagery (HSI) provides detailed spectral information, enabling precise analysis of materials. However, HSI imaging suffers from blurring degradation which results in the loss of fine details and hinders subsequent applications. The degree of blurriness is highly related to wavelength and depth, existing deblurring methods either lack the utilization of spectral correlation or ignore the depth variation since paired HSI and depth data are difficult to acquire and less discussed, leading to degraded performance when encountering wide-range HSIs of non-planar scenes. To address these challenges in both data acquisition and algorithm design, we propose a novel approach that simultaneously collects both modalities and integrates depth refinement into a blind HSI deblurring model with wavelength- and depth-aware deep image prior. Specifically, we capture blurred HSI and coarse depth map with separate devices, followed by registration. Our method performs depth-guided deblurring through depth-variant multi-channel kernel estimation and soft-weight map-based layer composition, while simultaneously refining the depth. The proposed approach effectively restores fine details with fewer artifacts, showing superior performance for both simulated blurred HSIs and real captured HSIs.
Jiahuan Li, Wei He 0003, Naoto Yokoya
WACV3
2025 Human Tide, Clear Sight: Semantically Enhanced Visual Localization in High-Crowd Scenarios
abstract
Accurate visual localization is essential in IoT applications, particularly for robotics, autonomous systems, and augmented reality. Traditional feature-based methods struggle with efficiency and robustness against environmental variations. To enhance the robustness of visual localization algorithms against these variations, state-of-the-art (SOTA) methods have incorporated semantic information as an advanced dimension into their models, but still suffer from several shortcomings. These methods often embed semantic information implicitly, which limits their extensibility and interpretability. Moreover, the introduction of some unstable semantic labels may, on the contrary, degrade the localization accuracy. Therefore, modularity, quantization, and filtering semantic labels by their stability become critical. To address these gaps, this article proposes a method that explicitly and quantitatively integrates semantic information through a plug-and-play module. This module scores image-to-image and feature-to-feature correspondences based on semantic similarity and stability, with a particular focus on improving smartphone-based visual localization in high-crowd indoor scenarios. This module is introduced into two key stages of visual hierarchical localization: 1) visual place recognition (coarse localization) and 2) 6-Degree-of-Freedom pose estimation (fine localization). Specifically, correspondences with low scores imply a higher probability of matching errors and are therefore suppressed. To validate the proposed approach, a novel dataset designed for semantic visual localization tasks is collected, rich with dynamic objects and scene variations. The method demonstrates superior accuracy and robustness, particularly in environments with significant scene appearance changes, with 13.6% and 5.4% improvement in localization accuracy in Cafds and Libds datasets, respectively, compared to the SOTA approach. The code and dataset are available athttps://github.com/1da1da/SEVL.
Yida Wei, Sikang Liu 0001, Wei He 0003, You Li 0001
IEEE Internet Things J.5
2025 Low-Rank Tensor Meets Deep Prior: Coupling Model-Driven and Data-Driven Methods for Hyperspectral Image Reconstruction
abstract
Snapshot compressive imaging (SCI) captures a 3D hyperspectral image (HSI) using a 2D compressive measurement and reconstructs the desired 3D HSI from that 2D measurement. The effective reconstruction method thus is crucial in SCI. Despite recent successes of deep learning (DL)-based methods over traditional approaches, they often ignore the intrinsic characteristics of HSI and are trained for a specific imaging system using sufficient paired datasets. To address this, we propose a novel self-supervised HSI reconstruction framework called low-rank tensor meets deep prior (LDMeet), which couples model-driven and data-driven methods. The design of LDMeet is inspired by the traditional model-driven low-rank tensor prior constructed based on domain knowledge, which can explore the intrinsic global spatial-spectral correlation of HSI and make the reconstruction method interpretable. To further utilize the powerful learning ability of DL-based approaches, we introduce a self-supervised spatial-spectral guided network (SSG-Net) into LDMeet to learn the implicit deep spatial-spectral prior of HSI without requiring training data, making it adaptable to various imaging systems. An efficient alternating direction method of multiplier (ADMM) is designed to solve the LDMeet model. Comprehensive experiments confirm that our LDMeet achieves superior results compared to self-supervised HSI reconstruction methods, while also yielding competitive results with supervised learning methods.
Yong Chen 0013, Feiwang Yuan, Wenzhen Lai, Jinshan Zeng, Wei He 0003
IEEE Trans. Circuits Syst. Video Technol.5
2025 KDGraph: A Keypoint Detection Method for Road Graph Extraction From Remote Sensing Images
abstract
Road graph extraction from remote sensing images is essential in navigation and urban planning. However, shadows and occlusions in these images frequently disrupt the continuity of road representations, resulting in fragmented and poorly connected road graphs extracted by existing methods. To overcome these challenges, we introduce KDGraph, a novel keypoint detection method for road graph extraction from remote sensing images. Specifically, keypoints are defined as vertices situated at the endpoints, corners, or junctions of roads, characterized by their positions and multiple directional attributes. In addition, we develop a greedy parsing algorithm to connect these keypoints and construct the road graph based on their directional information. The key innovation of our method lies in reformulating road mask extraction as a keypoint detection task. By positioning keypoints in areas less affected by shadows and occlusions, KDGraph effectively reduces their impact on road graph connectivity. Extensive experiments are conducted using the SpaceNet3 dataset and a newly constructed shadow-occluded road graph extraction (SoR) dataset, which includes road segments from 15 cities with varying degrees of shadows and occlusions. A patch expansion strategy is introduced for large-scale inference on the SoR dataset. Results show that KDGraph outperforms all comparative methods in both generalization and the capability of handling shadows and occlusions. Code and SoR dataset available at:https://github.com/ruoyxue/KDGraph.
Wei He 0003, Ruoyao Xue, Fangxiao Lu, Jinjun Xu, Hongyan Zhang 0001
IEEE Trans. Geosci. Remote. Sens.1
2025 MGNet: A Remote Sensing Oriented Object Detector Based on Multicascaded Feature Selection and Geometric Constraints
abstract
High inter-class similarity and diverse geometric configurations in high spatial resolution remote sensing images pose significant challenges for remote sensing object detection. To address these challenges, this study proposed a novel remote sensing oriented object detector MGNet based on multi-cascaded feature selection and geometric constraints, designed to enhance discriminative feature representation and optimize oriented bounding box regression. First, we designed a novel backbone Multi-Cascaded Feature Selection Network (MCFSNet) to improve inter-class separability by jointly leveraging and fusing contextual information and multi-scale features. Second, we designed a Geometric-Constrained Probabilistic Loss, which introduced an angular offset factor and a geometric perception factor to refine angular prediction accuracy of square targets and shape sensitivity. To systematically evaluate angular localization performance, we constructed a Multi-scale Plane Remote Sensing Object Detection dataset (MPRSOD) with arbitrary-orientated annotations and proposed the Angle Error Rate (AER), the first dedicated evaluation metric for quantifying angular prediction errors. Extensive experiments on DIOR-R, DOTA V1.5, MAR20 and MPRSOD datasets demonstrate that MGNet obtains competitive mAP results while maintaining computational efficiency.
Yujie Lei, Jie Zhang 0123, Wei He 0003, Jiasong Zhu, Qingquan Li 0001
IEEE Trans. Geosci. Remote. Sens.4
2025 A Comprehensive Deep-Learning Framework for Fine-Grained Farmland Mapping From High-Resolution Images
abstract
The extraction of large-scale farmland is essential for optimizing agricultural production and advancing sustainable development. To meet the urgent need for efficient farmland extraction and overcome existing technical challenges, we have developed a comprehensive farmland mapping framework that integrates advanced data, methodology, and cartographic techniques. Regarding data, we present the fine-grained farmland dataset (FGFD), which compiles high-quality, meticulously annotated very high-resolution (VHR) satellite images and captures distinct regional characteristics across eastern, southern, western, northern, and central China. Building on the FGFD, we propose the dual-branch boundary-aware network (DBBANet), which employs ResNet-50 as the encoder to extract multilayer encoded features and introduces two parallel decoding branches: a spatial-aware branch and a boundary-aware branch. The dual-branch architecture leverages both unique semantic information relevant to farmland and detailed boundary information, facilitating a more comprehensive and accurate representation of farmland areas. By combining this dataset with our innovative methodology, we further propose a farmland mapping framework designed for large-scale applications. The proposed framework enables the direct generation of high-precision vector maps from VHR images, providing crucial technical support for farmland management, resource assessment, and agricultural planning. Extensive experiments conducted on the FGFD have established benchmarks for 13 segmentation methods, demonstrating the state-of-the-art (SOTA) performance of our approach. In practical large-scale applications, our mapping framework produces high-precision vector maps with clear boundaries, bridging the gap in fine-grained farmland mapping and paving the way for further research and applications in this field. The source code of the proposed DBBANet and FGFD is available at:https://github.com/Henryjiepanli/DBBANet.
Jiepan Li, Yipan Wei, Tiangao Wei, Wei He 0003
IEEE Trans. Geosci. Remote. Sens.4
2025 Toward Faithful Scene-Adaptive Knowledge for Semantic Segmentation of Remote Sensing Images
abstract
Semantic segmentation is a critical procedure in remote sensing image analysis that backs up various applications. High resolution remote sensing images contain a wealth of ground object features, which are organized into various describable scenes. The visual content offset towards each scene is an intuitive sense for understanding geospatial objects. However, existing semantic segmentation methods for remote sensing images generally neglect this intuition and lack the ability to adjust their perception preference on different images. To address this problem, we propose a paradigm for collecting scene information and dynamically adjusting the model inference process to be scene-aware. Specifically, our method leverages the class feature from the image to enhance the fixed class representation from the model. The interaction of these information is facilitated by a neighbor-friendly embedding space, making it more faithful to associate the image features and model parameters. For the model to better understand the complex scenes, a manifold mixup method is proposed to expand the effective embedding space on intra-class and inter-class regions, forcing the model to challenge the ambiguous instances in remote sensing images. Extensive experiments on four publicly available datasets demonstrated that our proposed improved the accuracy of semantic segmentation models on remote sensing images, overcoming the state-of-the-art methods.
Yue Liao, Wei He 0003, Hongyan Zhang 0001
IEEE Trans. Geosci. Remote. Sens.2
2025 Fusing Global Structural and Local Deep Features for Thick Cloud Removal in Multitemporal Remote Sensing Images
abstract
Optical remote sensing images are inevitably affected by thick cloud cover, leading to information missing, which seriously hinders subsequent Earth observation tasks. Existing cloud removal methods typically focus on extracting either global or local features, but lack effective fusion of these two aspects, resulting in deficiencies in recovering global structure and fine details. However, fusing multiple features is challenging for single-driven methods due to the diversity of images and clouds. To this end, this paper proposes a novel joint-driven cloud removal framework that can simultaneously exploit the global structural and local deep features of both the image and cloud components. Specifically, in our joint-driven scheme, we devise a flexible model-driven low-rank group sparse decomposition method, in which the global temporal-spectral correlations and shared sparse features of multitemporal remote sensing images and clouds from different scenes are effectively captured. To fully account for the differences in the local features of images and clouds from different scenes, we devise effective data-driven dual self-supervised networks to adaptively learn their local deep features. Additionally, we design an efficient half-quadratic splitting algorithm to iteratively optimize the joint-driven model, fostering mutual enhancement between the two and achieving further performance improvements. Experimental results show that the proposed method outperforms existing thick cloud removal approaches in both simulated and real-world datasets, under varying factors such as spectral bands, cloud coverage, time nodes, and spatial resolutions.
Yeqi Xu, Yong Chen 0013, Wei He 0003, Min Huang 0005, Jinshan Zeng
IEEE Trans. Geosci. Remote. Sens.4
2025 TANet: Thin Cloud-Aware Network for Cloud Detection in Optical Remote Sensing Image
abstract
Accurate cloud detection is essential for subsequent optical remote sensing imagery processing. Although many deep learning (DL)-based cloud detection methods have been proposed, the accurate detection of thin clouds still remains a challenge. To solve this issue, this article introduces a thin cloud-aware network (TANet). TANet tackles the problem from the aspects of color, texture, spatial distribution, and feature by constructing unique strategies to enhance the sensitivity of the network to thin cloud regions, thereby improving the overall accuracy of cloud detection. On the one hand, the TANet utilizes a color prior guidance module (CPGM) to incorporate robust dehaze priors, guiding the network to pay more attention to thin cloud areas. On the other hand, the global information aggregation module (GIAM) is employed to deeply extract long-distance dependence between pixels, mining potential correlations between thin and thick clouds, and addressing the challenge of recognizing thin clouds in a local perspective. In addition, we construct a plug-and-play cloud feature difference (CFD) loss, encouraging the network to learn more distinctive features between pixels of thin clouds and cloud-free regions, thereby strengthening the network’s ability to distinguish highly similar samples of different classes. The experimental results substantiate that our proposed method attains the lowest omission error and the highest detection accuracy. This affirms the superior capability of TANet in thin cloud detection, thereby yielding more dependable cloud detection results.
Wei He 0003, Yu Xia 0032, Hongyan Zhang 0001, Ting Hu 0003
IEEE Trans. Geosci. Remote. Sens.2
2025 Feature Fusion-Guided Network With Sparse Prior Constraints for Unsupervised Hyperspectral Image Quality Improvement
abstract
Due to imaging hardware limitations and atmospheric interference, hyperspectral image (HSI) often suffers from low spatial resolution or mixed noise degradation. HSI fusion and denoising are two key strategies to improve HSI quality. Traditional model-based methods rely on data-specific manual priors, while supervised deep learning methods are typically developed specifically for a single task and require a large number of training datasets. To address these limitations, we propose a novel unsupervised feature fusion-guided network (UFFGNet) as a general prior that effectively leverages multi-scale semantic features from a guidance image while incorporating sparse prior constraints to suppress outliers. Specifically, UFFGNet comprises a deep feature extraction network to capture multiscale semantic features from a guidance image, and an attentionbased feature generation network that generates an output image from random noise. These two networks are connected by a feature refinement module to embed the refined features from the feature aggregation module into the generation network. Furthermore, the sparse prior constraint is incorporated to model sparse noise (including impulse noise, stripe artifacts, and deadlines) in HSI, thereby improving the robustness of UFFGNet. The proposed network optimizes network parameters in an unsupervised manner without external additional training data, using the fidelity term in the degradation model as a loss function to learn the prior information of the original image. Extensive experiments demonstrate that the proposed UFFGNet outperforms state-of-the-art methods in both HSI fusion and denoising tasks, significantly improving HSI quality.
Feiwang Yuan, Yong Chen 0013, Wei He 0003, Jinshan Zeng
IEEE Trans. Geosci. Remote. Sens.3
2024 Learning without Exact Guidance: Updating Large-Scale High-Resolution Land Cover Maps from Low-Resolution Historical Labels
abstract
Large-scale high-resolution (HR) land-cover mapping is a vital task to survey the Earth's surface and resolve many challenges facing humanity. However, it is still a nontrivial task hindered by complex ground details, various landforms, and the scarcity of accurate training labels over a wide-span geographic area. In this paper, we propose an efficient, weakly supervised framework (Paraformer) to guide large-scale HR land-cover mapping with easy-access historical land-cover data of low resolution (LR). Specifically, existing land-cover mapping approaches reveal the dominance of CNNs in preserving local ground details but still suffer from insufficient global modeling in various landforms. Therefore, we design a parallel CNN-Transformer feature extractor in Paraformer, consisting of a downsampling-free CNN branch and a Transformer branch, to jointly capture local and global contextual information. Besides, facing the spatial mismatch of training data, a pseudo-label-assisted training (PLAT) module is adopted to reasonably refine LR labels for weakly supervised semantic segmentation of HR images. Experiments on two large-scale datasets demonstrate the superiority of Paraformer over other state-of-the-art methods for automatically up-dating HR land-cover maps from LR historical labels.
Zhuohong Li, Wei He 0003, Jiepan Li, Fangxiao Lu, Hongyan Zhang 0001
CVPR2
2024 Cross Modal Few Shot Learning for Tree Species Classification Using Airborne Hyperspectral Images
abstract
Tree species classification is essential for forest resource surveys and monitoring activities. Although airborne hyperspectral images (HSIs) can provide rich spatial and spectral information, the lack of labeled samples and the high similarity between spectra remain challenges for achieving fine-grained tree species classification mapping. In this article, a cross modal few shot learning framework is presented for multiple tree species classification (CMTSC). Notably, we innovatively introduce language prior knowledge to guide the generation of discriminative visual features, aiming to use additional modalities to improve the uni-modal classification task. Firstly, an improved three-dimensional ghost attention network (TGAN) with strong learning capability without massive parameters is constructed. Secondly, we use linguistic features from class names to optimize the decision boundary of the visual classifier by cross-modal adaptation. Thirdly, the cross domian few shot learning (FSL) strategy is employed to overcome the dilemma of sparse labeled samples and fixed application scenarios. Experiments on Gaofeng Forest Farm B (GFF-B) in Nanning City demonstrate the effectiveness of the proposed method compared to other state-of-the-art methods. The codes will be available at: https://github.com/HlEvag/CMTSC.
Lei Hu 0001, Wei He 0003, Hongyan Zhang 0001
IGARSS2
2024 Overcoming the Uncertainty Challenges in Flood Rapid Mapping with SAR Data
abstract
The escalating intensity and frequency of floods, exacerbated by global climate change, emphasize the urgent need to address the growing risks of floods. Rapid and precise flood detection is paramount for efficiently responding to emergencies and executing disaster relief measures, enabling swift reactions to flood disasters and minimizing the losses incurred thus. The 2024 IEEE GRSS Data Fusion Contest Track 1 is centered on leveraging multi-source remote sensing data, particularly synthetic aperture radar (SAR) data, to classify flood and non-flood areas. In this contest, we acknowledge the significance of managing uncertain predictions and present an efficient Uncertainty-Aware Fusion Network (UAFNet). Specifically, we build on the traditional encoder-decoder architecture, initially employing the pyramid visual transformer (PVT) as a feature extractor. Subsequently, we apply a typical decoding strategy, namely the feature pyramid network, to obtain a flood extraction map with relatively high uncertainty. Furthermore, leveraging the uncertain extraction map, we introduce an Uncertainty Rank Algorithm to quantify the uncertainty level of each pixel of the foreground and background. We seamlessly integrate this algorithm with our proposed Uncertainty-Aware Fusion Module, enabling level-by-level feature refinement and ultimately yielding a refined extraction map with minimal uncertainty. Employing the proposed UAFNet, we utilize diverse versions of PVT as encoders to train multiple UAFNets. Additionally, we enhance our approach with online testing augmentation and multi-model fusion strategy, aiming to enhance the final flood extraction accuracy. Our technical solution has exhibited outstanding performance, earning the first-place ranking in the 2024 IEEE GRSS Data Fusion Contest Track 1 and achieving an impressive F1 score of 82.985% on the official test set.
Jiepan Li, Wei He 0003, Hongyan Zhang 0001, Liangpei Zhang 0001
IGARSS3
2024 Overcoming the Uncertainty Challenges in Flood Rapid Mapping with Multi-Source Optical Data
abstract
As global climate change worsens, floods are becoming more severe and frequent, urgently demanding effective flood risk mitigation strategies. Timely and precise flood inundation mapping is crucial for emergency response and relief. The 2024 IEEE GRSS Data Fusion Contest Track 2 aims to pioneer innovative algorithms for accurate flood extraction using multi-source optical remote sensing (RS) data. However, data diversity introduces aleatoric uncertainty, especially with synthetic, non-real data. Meanwhile, the vast coverage of RS imagery and the small proportion of flood areas cause a significant class imbalance, leading to epistemic uncertainty. In this paper, we propose an Uncertainty-aware Detail-Preserving Network (UADPNet) for rapid flood mapping of multi-source optical data. Firstly, we design an Aleatoric Uncertainty Estimator to model aleatoric uncertainty in multi-source data. Secondly, we introduce a Multi-Scale Convolution Block to extract multi-scale information without downsampling. Thirdly, we utilize a multi-level supervised strategy to quantify epistemic uncertainty and highlight uncertain pixels via the Uncertainty-Aware Fusion Module. With UADPNet, we adopt a multi-model fusion and post-processing strategy to enhance the final flood extraction accuracy. Our outstanding experimental results in the official test set showcase the superiority of our method, which secured the first-place ranking in the 2024 IEEE GRSS Data Fusion Contest Track 2, boasting an impressive F1 score of 89.843%.
Jiepan Li, Wei He 0003, Hongyan Zhang 0001, Liangpei Zhang 0001
IGARSS3
2024 Identifying Every Building's Function in Large-Scale Urban areas with Multi-Modality Remote-Sensing Data
abstract
Buildings, as fundamental man-made structures in urban environments, serve as crucial indicators for understanding various city function zones. Rapid urbanization has raised an urgent need for efficiently surveying building footprints and functions. In this study, we proposed a semi-supervised framework to identify every building’s function in large-scale urban areas with multi-modality remote-sensing data. In detail, optical images, building height, and nighttime-light data are collected to describe the morphological attributes of buildings. Then, the area of interest (AOI) and building masks from the volunteered geographic information (VGI) data are collected to form sparsely labeled samples. Furthermore, the multi-modality data and weak labels are utilized to train a segmentation model with a semi-supervised strategy. Finally, results are evaluated by 20,000 validation points and statistical survey reports from the government. The evaluations reveal that the produced function maps achieve an OA of 82% and Kappa of 71% among 1,616,796 buildings in Shanghai, China. This study has the potential to support large-scale urban management and sustainable urban development. All collected data and produced maps are open access at https://github.com/LiZhuoHong/BuildingMap.
Zhuohong Li, Wei He 0003, Jiepan Li, Hongyan Zhang 0001
IGARSS2
2024 Detector-Free Multimodal Image Matching
abstract
Multimodal image matching is essential in image stitching, image fusion, change detection, and land cover mapping. However, the severe nonlinear radiometric distortion and geometric distortion of multimodal images severely limit the accuracy of multimodal image matching. To solve these problems, we propose a detector-free multimodal image matching approach to establish pixel-level dense correspondences. We mitigate the impact of modality differences on feature point extraction by establishing robust reference points. Specifically, we design a phase congruency module to keep the location of the reference point centered on the image edge structures. Simultaneously, a guiding correction module exploits the geometric relationships between pixels and reference points to establish accurate pixel correspondences. Finally, refined correspondences are obtained by finely positioning highly correlated pixel matches. Experiments show that our method can obtain sufficient and robust correspondences on multimodal images.
Wei He 0003, Hongyan Zhang 0001, Mbulisi Sibanda, Elhadi Adam
IGARSS2
2024 Unidirectional Spatial and Spectral Smoothed Tensor Ring Decomposition for Hyperspectral Image Denoising and Destriping
abstract
In this letter, we propose a novel unidirectional spatial and spectral smoothed tensor ring (U3STR) decomposition for hyperspectral image (HSI) denoising and destriping. The powerful tensor ring (TR) decomposition is introduced to explore the global spatial-spectral correlation of HSI, which transforms the restoration of HSI into estimating three TR factors. To address the local spatial-spectral smoothness of HSI and the directional characteristic of stripe noise, unidirectional spatial and spectral smoothed constraints are applied to the horizontal spatial and spectral TR factors, respectively. Moreover, considering that the stripe noise shares spatial correlation and local smoothness with the image component, we strategically utilize band-by-band low rank and unidirectional total variation (TV) regularization, effectively disentangling stripe noise from the image content without conflicting the image regularization. The proposed U3STR model is solved by the alternating direction method of multipliers (ADMM) algorithm effectively. Experimental results demonstrate that our method outperforms other HSI restoration methods in denoising and destriping, notably enhancing the quality of the restored image by an average of 3 dB over existing methods.
Yong Chen 0013, Jinshan Zeng, Wei He 0003, Min Huang 0005
IEEE Geosci. Remote. Sens. Lett.4
2024 Thick Cloud Removal in Multitemporal Remote Sensing Images via Low-Rank Regularized Self-Supervised Network
abstract
The existence of thick clouds covers the comprehensive Earth observation of optical remote sensing images (RSIs). Cloud removal is an effective and economical preprocessing step to improve the subsequent applications of RSIs. Deep learning (DL)-based methods have attracted much attention and achieved state-of-the-art results. However, most of these methods suffer from the following issues: 1) ignore the physical characteristics of RSIs; 2) require paired images with/without cloud or extra auxiliary images (such as SAR); and 3) demand the cloud mask. These issues might have limited the flexibility of existing networks. In this paper, we propose a novel low-rank regularized self-supervised network (LRRSSN) that couples model-driven and data-driven methods to remove the thick cloud from multitemporal remote sensing images (MRSIs). First, motivated by the equal importance of image and cloud components as well as their intrinsic characteristics, we decompose the observed image into low-rank image and structural sparse cloud components. In this way, we obtain a model-driven thick cloud removal method where the spectral-temporal low-rank correlation of the image component and the spectral structural sparsity of the cloud component are effectively exploited. Second, to capture the complex nonlinear features of different scenarios, the data-driven self-supervised network that does not require external training datasets is designed to explore the deep prior of the image component. Third, the coupled model-driven and data-driven LRRSSN is optimized by an efficient half-quadratic splitting algorithm. Finally, without knowing the exact cloud mask, we estimate the cloud mask to preserve information in cloud-free areas as much as possible. Experiments conducted in synthetic and real-world scenarios demonstrate the effectiveness of the proposed approach.
Yong Chen 0013, Wei He 0003, Jinshan Zeng, Min Huang 0005, Yu-Bang Zheng
IEEE Trans. Geosci. Remote. Sens.3
2024 Pretrain a Remote Sensing Foundation Model by Promoting Intra-Instance Similarity
abstract
Self-supervised learning (SSL) has gained significant traction within the remote sensing community, with pretraining a foundation model on large-scale unlabeled datasets for the interpretation of remote sensing images (RSIs) emerging as a trending direction. This approach aims to supplant the conventional practice of loading ImageNet pretrained weights, offering a more versatile and potentially more effective solution for handling RSIs. Among SSL techniques, contrastive learning excels in extracting general representations in the field of remote sensing. However, its excessive focus on inter-instance discrimination hinders the effectiveness of pretraining due to the diverse and complex geographical information present in RSIs. Moreover, the typical two-variations-as-one-pair pattern may be suboptimal, particularly given the temporal information specific to RSIs. In this article, we propose a novel method called promoting intra-instance similarity (PIS) for short, which leverages the temporal information specific to RSIs and increases the intra-instance variations to expand the positive representation space. Additionally, by PIS within this space, our foundation models develop the ability to extract more general and instance-invariant features that prove beneficial for various downstream tasks. Experiments show that our PIS method achieves state-of-the-art (SOTA) performance on ten datasets across four downstream remote sensing tasks, demonstrating the generalizability and efficacy of the proposed method. Through our preliminary investigation into intra-instance characteristics, we believe there exists substantial potential in this aspect, holding considerable promise for further exploration. The codes are available on the website:https://github.com/ShawnAn-WHU/PIS.git.
Xiao An, Wei He 0003, Jiaqi Zou, Hongyan Zhang 0001
IEEE Trans. Geosci. Remote. Sens.2
2024 Fast Large-Scale Hyperspectral Image Denoising via Noniterative Low-Rank Subspace Representation
abstract
Denoising of hyperspectral image (HSI) is challenging, especially when dealing with large-scale data. Model-based methods show promise in HSI denoising due to their good generalization, but they suffer from computational complexity due to complex priors [like nonlocal self-similarity (NSS)] and iterations, resulting in low efficiency for large-scale HSI processing. To address these challenges, we propose a fast large-scale HSI denoising (FallHyDe) method based on noniterative low-rank (LR) subspace representation to enjoy high denoising efficiency, effectiveness, and flexibility simultaneously. By leveraging the global spectral property of HSI, FallHyDe efficiently estimates spectral subspace and spatial representation coefficients (SRCs) from the observed noisy HSI, reducing computation complexity caused by the high spectral dimension during processing. In addition, we innovatively explore the presence of high signal-to-noise ratio bands (HSNRBs) in real HSI, enabling fast SRC estimation through a least squares problem without relying on complex priors and iterations. FallHyDe requires neither iteration nor parameter tuning, enabling our method to process large-scale HSI denoising quickly and flexibly. Experimental results on both simulated and real HSI datasets demonstrate that our proposed method not only achieves competitive results in quality but also speeds up the restoration by more than ten times than the representative fast HSI denoising methods. The code is available athttps://chenyong1993.github.io/yongchen.github.io/.
Yong Chen 0013, Jinshan Zeng, Wei He 0003, Xi-Le Zhao, Tai-Xiang Jiang
IEEE Trans. Geosci. Remote. Sens.3
2024 An Unsupervised Dehazing Network With Hybrid Prior Constraints for Hyperspectral Image
abstract
Haze pollution in hyperspectral images (HSIs) leads to surface information lack and image clarity degradation, which seriously affects the performance of subsequent image interpretation. Existing model-based hyperspectral haze removal methods enjoy good interpretability and generalization, but they can only process images in a specific wavelength range due to the principle limitation. Deep learning-based dehazing methods have good feature extraction capability, but the cost of acquiring sufficient training data is high in practical applications. At the same time, taking into account that HSIs have spectral low-rank structures, fully utilizing the low-rank property will facilitate the reconstruction of HSIs. In order to combine the complementary benefits of deep learning-based and physical model-based approaches, we decide to formulate HSI dehazing reconstruction as an unsupervised DIP framework. Specifically, we propose an unsupervised dehazing network with hybrid prior constraints (HPC-UDN) for HSI haze removal, which effectively integrates low-rank prior, deep priors, and physical haze prior. First, the low-rank prior of hyperspectral data is characterized by matrix decomposition, where the decomposition factors are learned through two generative networks. Then, multiple spectral groups are divided based on the correlation and complementarity between spectral bands. In order to exchange information between adjacent spectral groups, a novel spectral grouping feature fusion module is designed, which connects neighboring spectral groups to transfer spectral and spatial features. Finally, high-quality HSI is recovered by merging the features extracted from each spectral group. Extensive simulated and real-data experiments certify the effectiveness and robustness of the presented unsupervised approach and potential applications in the GF-5 image dehazing task.
Wei He 0003, Yong Chen 0013, Hongyan Zhang 0001
IEEE Trans. Geosci. Remote. Sens.1
2024 Adaptive Regularized Low-Rank Tensor Decomposition for Hyperspectral Image Denoising and Destriping
abstract
Hyperspectral images (HSIs) are inevitably degraded by a mixture of various types of noise, such as Gaussian noise, impulse noise, stripe noise, and dead pixels, which greatly limits the subsequent applications. Although various denoising methods have already been developed, accurately recovering the spatial-spectral structure of HSIs remains a challenging problem to be addressed. Furthermore, serious stripe noise, which is common in real HSIs, is still not fully separated by the previous models. In this paper, we propose an adaptive hyper-Laplacian regularized low-rank tensor decomposition (LRTDAHL) method for HSI denoising and destriping. On the one hand, the stripe noise is separately modeled by the tensor decomposition, which can effectively encode the spatial-spectral correlation of the stripe noise. On the other hand, adaptive hyper-Laplacian spatial-spectral regularization is introduced to represent the distribution structure of different HSI gradient data by adaptively estimating the optimal hyper-Laplacian parameter, which can reduce the spatial information loss and over-smoothing caused by the previous total variation regularization. The proposed model is solved using the alternating direction method of multipliers (ADMM) algorithm. Extensive simulation and real-data experiments all demonstrate the effectiveness and superiority of the proposed method.
Dong Chu, Xiaobin Guan, Wei He 0003, Huanfeng Shen
IEEE Trans. Geosci. Remote. Sens.4
2024 UANet: An Uncertainty-Aware Network for Building Extraction From Remote Sensing Images
abstract
Building extraction aims to segment building pixels from remote sensing images and plays an essential role in many applications, such as city planning and urban dynamic monitoring. Over the past few years, deep learning methods with encoder–decoder architectures have achieved remarkable performance due to their powerful feature representation capability. Nevertheless, due to the varying scales and styles of buildings, conventional deep learning models always suffer from uncertain predictions and cannot accurately distinguish the complete footprints of the building from the complex distribution of ground objects, leading to a large degree of omission and commission. In this paper, we realize the importance of uncertain prediction and propose a novel and straightforward Uncertainty-Aware Network (UANet) to alleviate this problem. Specifically, we first apply a general encoder–decoder network to obtain a building extraction map with relatively high uncertainty. Second, in order to aggregate the useful information in the highest-level features, we design a Prior Information Guide Module to guide the highest-level features in learning the prior information from the conventional extraction map. Third, based on the uncertain extraction map, we introduce an Uncertainty Rank Algorithm to measure the uncertainty level of each pixel belonging to the foreground and the background. We further combine this algorithm with the proposed Uncertainty-Aware Fusion Module to facilitate level-by-level feature refinement and obtain the final refined extraction map with low uncertainty. To verify the performance of our proposed UANet, we conduct extensive experiments on three public building datasets, including the WHU building dataset, the Massachusetts building dataset, and the Inria aerial image dataset. Results demonstrate that the proposed UANet outperforms other state-of-the-art algorithms by a large margin. The source code of the proposed UANet is available at https://github.com/Henryjiepanli/Uncertainty-aware-Network.
Jiepan Li, Wei He 0003, Weinan Cao, Liangpei Zhang 0001, Hongyan Zhang 0001
IEEE Trans. Geosci. Remote. Sens.2
2024 G2LDIE: Global-to-Local Dynamic Information Enhancement Framework for Weakly Supervised Building Extraction From Remote Sensing Images
abstract
Image-level weakly supervised semantic segmentation (WSSS) methods have gained prominence in remote sensing image building extraction tasks, primarily due to their cost-effectiveness in manual annotation. However, owing to the intricate details present in remote sensing building images, the pseudolabels generated from existing image-level weakly supervised methods often encounter issues of incorrect activation and unclear boundaries. In this article, we propose a global-to-local dynamic information enhancement (G2LDIE) framework. This framework effectively extracts global information from remote sensing building images and supplements local details through a local information enhancement (LIE) module, generating more accurate pseudolabels. Additionally, we propose a dynamic label guide strategy (DLGS) to enhance model consistency in category representation across various scale images. To address the challenge of unclear building boundaries issue in pseudolabels, we introduce a segment anything model (SAM) postprocessing (SPP) method, which can better correct the boundaries of building images at different resolutions while reducing computational costs. Extensive and detailed experiments on three datasets confirm that our framework can generate refined pseudolabels and outperform other image-level weakly supervised methods in terms of accuracy and generalization performance in building extraction.
Jiaxing Sun 0001, Wei He 0003, Hongyan Zhang 0001
IEEE Trans. Geosci. Remote. Sens.2
2024 SOSSF: Landsat-8 Image Synthesis on the Blending of Sentinel-1 and MODIS Data
abstract
Landsat optical sensor is crucial for the long-term observations of the Earth’s surface with a 30 m spatial resolution. However, the 16-day revisit cycle and severe atmospheric interference have impeded the monitoring of rapid surface changes. Spatiotemporal fusion (STF) is a classic method of predicting Landsat surface reflectance with multi-temporal and multi-source data, but it is limited by unpredictable temporal changes and cloudy Landsat-MODIS image pairs. Another emerging solution is synthetic aperture radar (SAR)-to-optical image translation (S2OIT), which always produces spectral distortions. To tackle these defects, we propose a new data-driven solution, SAR-optical data-based spatial–spectral fusion (SOSSF), which combines the high-spatial and cloud-free advantages of Sentinel-1 data and the high-spectral and high-temporal advantages of MODIS images to synthesize high-spatial and high-temporal Landsat-8 images. To achieve this solution, we first establish a worldwide benchmark dataset, namely SMILE, with various land cover types and all meteorological seasons, satisfying the big data requirements of deep learning. Second, we design an attention-based dual-path fusion network (ADFNet) to respectively extract and fully fuse spatial and spectral information from SAR-optical data. Extensive experiments suggest that the proposed SOSSF solution outperforms the state-of-the-art STF and S2OIT solutions, robustly performing in the continuously changing and frequently cloudy regions. The proposed ADFNet model achieves the best visual effect and the highest accuracy in different scenes, seasons, and bands. Furthermore, the proposed SOSSF solution is proven to be a practical way to simulate time-series and large-scale Landsat-8 surface reflectance, considerably enriching raw Landsat-8 products.
Yu Xia 0032, Wei He 0003, Hongyu Chen 0003, Hongyan Zhang 0001
IEEE Trans. Geosci. Remote. Sens.2
2024 PSFormer: Pyramid Superpixel Transformer for Hyperspectral Image Classification
abstract
Hyperspectral image (HSI) classification is a core processing procedure in the remote sensing community, which has been recently well studied using vision transformers (ViTs). However, due to the high computational and memory complexities, existing transformer-based classification methods tend to restrict the spatial extent of the transformer to small cropped HSI patches instead of the whole HSI data, thus sacrificing the essential strength of transformers in long-range interaction modeling and overlooking the beneficial multiscale features in HSI data. Inspiringly, here we propose PSFormer, a novel pyramid superpixel transformer (PSFormer) method specifically for HSI classification, in order to make full use of the transformer to excavate multiscale local-global features in HSI data. Specifically, a progressive superpixel merging strategy is introduced to flexibly control the scale of feature maps. Furthermore, a unique transformer backbone design based on a spectral attention layer and a classification head with a gate mechanism are developed, to adaptively exploit valuable local-global information at different scales with low computational cost. Extensive experimental results on five widely used datasets demonstrate the superiority of PSFormer over other state-of-the-art networks. For the sake of reproducibility, the related code of the PSFormer method will be open-sourced at:https://github.com/immortal13.
Jiaqi Zou, Wei He 0003, Hongyan Zhang 0001
IEEE Trans. Geosci. Remote. Sens.2
2024 Hyperspectral Compressive Snapshot Reconstruction via Coupled Low-Rank Subspace Representation and Self-Supervised Deep Network
abstract
Coded aperture snapshot spectral imaging (CASSI) is an important technique for capturing three-dimensional (3D) hyperspectral images (HSIs), and involves an inverse problem of reconstructing the 3D HSI from its corresponding coded 2D measurements. Existing model-based and learning-based methods either could not explore the implicit feature of different HSIs or require a large amount of paired data for training, resulting in low reconstruction accuracy or poor generalization performance as well as interpretability. To remedy these deficiencies, this paper proposes a novel HSI reconstruction method, which exploits the global spectral correlation from the HSI itself through a formulation of model-driven low-rank subspace representation and learns the deep prior by a data-driven self-supervised deep learning scheme. Specifically, we firstly develop a model-driven low-rank subspace representation to decompose the HSI as the product of an orthogonal basis and a spatial representation coefficient, then propose a data-driven deep guided spatial-attention network (called DGSAN) to adaptively reconstruct the implicit spatial feature of HSI by learning the deep coefficient prior (DCP), and finally embed these implicit priors into an iterative optimization framework through a self-supervised training way without requiring any training data. Thus, the proposed method shall enhance the reconstruction accuracy, generalization ability, and interpretability. Extensive experiments on several datasets and imaging systems validate the superiority of our method. The source code and data of this article will be made publicly available at https://github.com/ChenYong1993/LRSDN.
Yong Chen 0013, Wenzhen Lai, Wei He 0003, Xi-Le Zhao, Jinshan Zeng
IEEE Trans. Image Process.3
2024 GRiD: Guided Refinement for Detector-Free Multimodal Image Matching
abstract
Multimodal image matching is essential in image stitching, image fusion, change detection, and land cover mapping. However, the severe nonlinear radiometric distortion (NRD) and geometric distortions in multimodal images severely limit the accuracy of multimodal image matching, posing significant challenges to existing methods. Additionally, detector-based methods are prone to feature point offset issues in regions with substantial modal differences, which also hinder the subsequent fine registration and fusion of images. To address these challenges, we propose a guided refinement for detector-free multimodal image matching (GRiD) method, which weakens feature point offset issues by establishing pixel-level correspondences and utilizes reference points to guide and correct matches affected by NRD and geometric distortions. Specifically, we first introduce a detector-free framework to alleviate the feature point offset problem by directly finding corresponding pixels between images. Subsequently, to tackle NRD and geometric distortion in multimodal images, we design a guided correction module that establishes robust reference points (RPs) to guide the search for corresponding pixels in regions with significant modality differences. Moreover, to enhance RPs reliability, we incorporate a phase congruency module during the RPs confirmation stage to concentrate RPs around image edge structures. Finally, we perform finer localization on highly correlated corresponding pixels to obtain the optimized matches. We conduct extensive experiments on four multimodal image datasets to validate the effectiveness of the proposed approach. Experimental results demonstrate that our method can achieve sufficient and robust matches across various modality images and effectively suppress the feature point offset problem.
Wei He 0003, Hongyan Zhang 0001
IEEE Trans. Image Process.2
2024 Combining Low-Rank and Deep Plug-and-Play Priors for Snapshot Compressive Imaging
abstract
Snapshot compressive imaging (SCI) is a promising technique that captures a 3-D hyperspectral image (HSI) by a 2-D detector in a compressed manner. The ill-posed inverse process of reconstructing the HSI from their corresponding 2-D measurements is challenging. However, current approaches either neglect the underlying characteristics, such as high spectral correlation, or demand abundant training datasets, resulting in an inadequate balance among performance, generalizability, and interpretability. To address these challenges, in this article, we propose a novel approach called LR2DP that integrates the model-driven low-rank prior and data-driven deep priors for SCI reconstruction. This approach not only captures the spectral correlation and deep spatial features of HSI but also takes advantage of both model-based and learning-based methods without requiring any extra training datasets. Specifically, to preserve the strong spectral correlation of the HSI effectively, we propose that the HSI lies in a low-rank subspace, thereby transforming the problem of reconstructing the HSI into estimating the spectral basis and spatial representation coefficient. Inspired by the mutual promotion of unsupervised deep image prior (DIP) and trained deep denoising prior (DDP), we integrate the unsupervised network and pre-trained deep denoiser into the plug-and-play (PnP) regime to estimate the representation coefficient together, aiming to explore the internal target image prior (learned by DIP) and the external training image prior (depicted by pre-trained DDP) of the HSI. An effective half-quadratic splitting (HQS) technique is employed to optimize the proposed HSI reconstruction model. Extensive experiments on both simulated and real datasets demonstrate the superiority of the proposed method over the state-of-the-art approaches.
Yong Chen 0013, Xinfeng Gui, Jinshan Zeng, Xi-Le Zhao, Wei He 0003
IEEE Trans. Neural Networks Learn. Syst.5
2024 Spectral Super-Resolution via Model-Guided Cross-Fusion Network
abstract
Spectral super-resolution, which reconstructs a hyperspectral image (HSI) from a single red-green-blue (RGB) image, has acquired more and more attention. Recently, convolution neural networks (CNNs) have achieved promising performance. However, they often fail to simultaneously exploit the imaging model of the spectral super-resolution and complex spatial and spectral characteristics of the HSI. To tackle the above problems, we build a novel cross fusion (CF)-based model-guided network (called SSRNet) for spectral super-resolution. In specific, based on the imaging model, we unfold the spectral super-resolution into the HSI prior learning (HPL) module and imaging model guiding (IMG) module. Instead of just modeling one kind of image prior, the HPL module is composed of two subnetworks with different structures, which can effectively learn the complex spatial and spectral priors of the HSI, respectively. Furthermore, a CF strategy is used to establish the connection between the two subnetworks, which further improves the learning performance of the CNN. The IMG module results in solving a strong convex optimization problem, which adaptively optimizes and merges the two features learned by the HPL module by exploiting the imaging model. The two modules are alternately connected to achieve optimal HSI reconstruction performance. Experiments on both the simulated and real data demonstrate that the proposed method can achieve superior spectral reconstruction results with relatively small model size. The code will be available at https://github.com/renweidian.
Renwei Dian, Tianci Shan, Wei He 0003
IEEE Trans. Neural Networks Learn. Syst.3
2023 A New Cloud Feature Difference Loss For Enhancing The Detection Of Clouds In Remote Sensing Images
abstract
Research has shown that most of the Earth’s surface is covered by clouds, which have reduced the usability of optical remote sensing images. Therefore, it is critical to detect clouds quickly and accurately. Cloud detection methods based on deep learning have been widely studied in recent years. However, the existing methods still face a significant challenge that thin clouds are often translucent and easily missed. To enhance the detection of thin clouds by convolutional neural networks, we propose a cloud feature difference (CFD) loss, which gathers the samples in thin clouds, thick clouds and subsurface as a sample group, namely Cloud-Triplet. By measuring the features difference between samples, the CFD loss forces the network to model discriminative features during training, thus improving the ability of detecting thin clouds. Experiments show that our proposed CFD loss is effective in enhancing the detection of clouds.
Wei He 0003, Yu Xia 0032, Hongyan Zhang 0001
IGARSS2
2023 A Spatial-Spectral Transformer Network With Total Variation Loss for Hyperspectral Image Denoising
abstract
Hyperspectral image (HSI) denoising has an essential effect on HSI analysis and interpretation. In recent years, denoising methods based on convolutional neural networks (CNNs) have made great progress. However, the convolution kernel in the CNN model is content-independent, and the ability to capture long-distance correlation is weak, which leads to spectral distortion and edge blurring. To address this problem, we propose a spatial–spectral transformer network for HSI denoising, which introduces the shifted window-based transformer method to denoise HSIs by modeling image content correlation while preserving the local inductive bias. In detail, to jointly explore the spatial–spectral features, we first formulate the spatial–spectral cubes as network input, which are composed of the current band and its adjacent fixed K-bands. Second, these spatial–spectral cubes are forwarded to cascaded hyperspectral transformer blocks (HTBs) with skip connection for deep feature extraction. The HTB contains multiple transformer layers based on different window partitioning, which not only reduces memory cost, but also enhances the feature extraction capability. Finally, for the denoised${B}$spatial–spectral cubes, we average the pixels of overlapping spectral bands to generate a complete HSI. Furthermore, we introduce the total variation (TV) to preserve the smoothness structures of HSIs. The experimental results on simulated and real data indicate that the proposed spatial-spectral transformer denoising (SSTD) is superior to other mainstream learning-based HSI denoising algorithms.
Wei He 0003, Hongyan Zhang 0001
IEEE Geosci. Remote. Sens. Lett.2
2023 Accurate Multiobjective Low-Rank and Sparse Model for Hyperspectral Image Denoising Method
abstract
Due to the unavoidable influence of sparse and Gaussian noise during the process of data acquisition, the quality of hyperspectral images (HSIs) is degraded and their applications are greatly limited. It is therefore necessary to restore clean HSIs. In the traditional methods, low-rank and sparse matrix decomposition methods are usually applied to restore the pure data matrix from the observed data matrix. However, due to the fact that the optimization of the${l}_{0}$-norm for the sparse modeling is a nonconvex and NP-hard problem, convex relaxation and regularization parameters are usually introduced. However, convex relaxation often leads to inaccurate sparse modeling results, and the sensitive regularization parameters can lead to unstable results. Thus, in this article, to address these issues, an accurate multiobjective low-rank and sparse denoising framework is proposed for HSIs to achieve accurate modeling. The${l}_{0}$-norm is directly modeled as the sparse noise and is optimized by an evolutionary algorithm, and the denoising problem is converted into a multiobjective optimization problem through simultaneously optimizing the low-rank term, the sparse term, and the data fidelity term, without sensitive regularization parameters. However, since the low-rank clean image and sparse noise of the HSI are encoded into a solution, the length of the solution is too long to be optimized. In this article, a subfitness strategy is constructed to achieve effective optimization by comparing the objective function values corresponding to each band for each solution. The experiments undertaken with simulated images in 11 noise cases and four real noisy images confirm the effectiveness of the proposed method.
Yuting Wan, Ailong Ma, Wei He 0003, Yanfei Zhong
IEEE Trans. Evol. Comput.3
2023 Cross-Domain Meta-Learning Under Dual-Adjustment Mode for Few-Shot Hyperspectral Image Classification
abstract
Hyperspectral image (HSI) classification with limited training samples has been well studied in recent years. Among them, the few-shot learning (FSL) technique demonstrates excellent processing capability under limited labeled samples. Nevertheless, the current FSL-based works provide scarce attention to effective class prototypes and metric types, resulting in high generalization error and poor interpretation during the cross-domain testing phase. A dual-adjustment mode-based cross-domain meta-learning (DMCM) method for few-shot HSI classification is proposed to tackle this issue. Specifically, a three-dimensional ghost attention network (TGAN) with strong learning capability without massive parameters is first constructed. Meanwhile, a dual-adjustment mode comprising intra-correction (IC) and inter-alignment (IA) learning strategies is then adopted to solve domain shift issue via episode-level meta tasks, where IC and IA focus on effective class prototypes and data distribution differences between domains, respectively. Afterward, considering that the traditional Euclidean distance metric is insensitive to the distribution of within-class samples, the class-covariance metric is employed to account for the distribution in feature space of each class to optimize decision boundary and alleviate the misclassification problem. Extensive experiments on three publicly available target hyperspectral datasets demonstrate the effectiveness of the proposed method in comparison with other state-of-the-art methods. The codes will be available on the website: https://github.com/HlEvag/DMCM.git.
Lei Hu 0001, Wei He 0003, Liangpei Zhang 0001, Hongyan Zhang 0001
IEEE Trans. Geosci. Remote. Sens.2
2023 GLoCNet: Robust Feature Matching With Global-Local Consistency Network for Remote Sensing Image Registration
abstract
Feature matching is a fundamental and critical task for remote sensing image registration. However, numerous outliers (false matches) harm the feature point neighborhood structure due to the view transformation from the camera. Meanwhile, unknown local distortion obscures the distinction between inliers (correct matches) and outliers. To solve these problems, we propose a global-local consistency network (GLoCNet) for feature matching to exclude the interference of outliers under various transformation patterns and provide stable neighborhood support for the similarity metric of feature points. Specifically, a global transformation consistency module is proposed to obtain a neighbor pool by exploiting the compact nature of the inlier distribution under different transformation patterns. In addition, feature points interact with information through the local neighborhood consistency module in center-based graph construction. Finally, the difference between inliers and outliers is increased by dynamically adjusting the upper limitation distance of the outlier and suppressing its effect in the neighborhood. We conducted rich experiments on extensive datasets to verify the effectiveness of the proposed method. The experimental results illustrate that the proposed GLoCNet can effectively handle numerous outliers and achieve satisfactory registration results under various transformation patterns.
Wei He 0003, Hongyan Zhang 0001
IEEE Trans. Geosci. Remote. Sens.2
2023 Content-Aware Subspace Low-Rank Tensor Recovery for Hyperspectral Image Restoration
abstract
The low-rank tensor model has made great progress for hyperspectral image (HSI) restoration. Recently, the low-rank tensor methods have further been boosted with subspace learning by transforming original HSI into a low-dimensional subspace with reduced computational burden and discriminative feature representation. However, existing subspace-based methods consistently employ a fixed subspace dimension for all patches, which may violate the intrinsic dimension discrepancy of different image content, leading to information loss or redundancy. In this work, our key observation is that the intrinsic subspace of different image patches along different dimensions is different, which should be adaptively modelled for compact feature extraction. Therefore, we propose a content-aware subspace low-rank tensor recovery (CSLRTR) methods by leveraging both deep network and low-rank tensor model. Specifically, we first analyze the intrinsic discrepancy of different HSI patches among both spatial and spectral dimensions, and design a simple network to adaptively learn the optimal subspace dimension. The adaptive subspace learning and low-rank tensor recovery are iteratively performed and mutually promote each other. On one hand, the learned subspace would contribute to more compact low-rank representation for better restoration; on the other hand, the low rank tensor recovery with less degradations would definitely ease the difficulty of the subspace estimation. Note that, the adaptive content-aware subspace strategy has been simultaneously employed on both spectral and nonlocal dimensions, where the spectral-spatial relationship has been further strengthened with better restoration. We have performed extensive experiments on different datasets and restoration tasks, and extended the content-aware subspace strategy to previous methods.
Xueyao Xiao, Wei Zhang 0161, Yi Chang 0002, Shuning Cao, Wei He 0003, Houzhang Fang, Luxin Yan
IEEE Trans. Geosci. Remote. Sens.5
2023 Hyperspectral Image Denoising: Reconciling Sparse and Low-Tensor-Ring-Rank Priors in the Transformed Domain
abstract
Recently, the transform-based tensor nuclear norm (TNN) framework has yielded promising results for hyperspectral image (HSI) denoising as compared with previous original-domain tensor-based models. However, the TNN framework only exploits the low-rankness of each band of HSIs (tensors) under a single spectral transform. The correlation between all bands under the transform (i.e., the global low-rankness of the transformed tensor) and the sparsity of the transformed HSI, which are beneficial for HSI denoising, is usually neglected in the TNN framework. In this article, we propose to reconcile sparse and low-tensor-ring (TR)-rank priors in the learned transformed domain (called T-RSTR model) for HSI denoising. In T-RSTR, the transform-based low-TR-rank and sparse regularizers are designed to characterize the global low-rankness and sparsity of the transformed tensors, respectively, and then the transform-based low-TR-rank and sparse regularizers are organically integrated and benefit from each other for substantially boosting denoising performance. To tackle the T-RSTR model, we elaborately design a proximal alternating minimization-based algorithm with the theoretical convergence. Extensive numerical results demonstrate that T-RSTR is superior to the competing methods.
Hao Zhang 0103, Ting-Zhu Huang, Xi-Le Zhao, Wei He 0003, Jae Kyu Choi, Yu-Bang Zheng
IEEE Trans. Geosci. Remote. Sens.4
2022 Spectrum-Aware and Transferable Architecture Search for Hyperspectral Image Restoration
Wei He 0003, Quanming Yao, Naoto Yokoya, Tatsumi Uezato, Hongyan Zhang 0001, Liangpei Zhang 0001
ECCV (19)1
2022 Autoencoder in Autoencoder Network Based on Low-Rank Embedding for Anomaly Detection in Hyperspectral Images
abstract
The main purpose of anomaly detection in hyperspectral images is to detect targets that are different from their surroundings. With the development of deep learning technology, anomaly detection in hyperspectral images using deep neural networks has drawn great attention in recent years. However, most of the existing deep learning-based anomaly detection algorithms fail to consider the low-rank properties of the background and underutilize the rich spectral information of the image. In this paper, we propose a novel autoencoder in autoencoder network based on low-rank module embedding for anomaly detection in hyperspectral images. Firstly, the background is purified by using the low-rank module (LRM), and then the image background is reconstructed by using autoencoder in autoencoder network (AiANet), which is a spatial-spectral dual encoding-decoding network. AiANet fully considers the differences between anomalies and backgrounds in spatial and spectral dimensions to better reconstruct the background. Finally, the anomaly appears in images as reconstruction errors. Our proposed method effectively exploits the low-rank property of the backgrounds and makes full use of the spectral information to extract pure backgrounds to separate anomalies. Experiments on two real hyperspectral images demonstrate that the proposed method outperforms the other competitors.
Weinan Cao, Hongyan Zhang 0001, Wei He 0003, Hongyu Chen 0003, Ewe Hong Tat
IGARSS3
2022 MSBRNet: Multi-Scale Background Reconstruction Network with Low-Rank Embedding for Anomaly Detection in Hyperspectral Images
abstract
The primary purpose of anomaly detection in hyperspectral images (HSI) is to detect different anomaly targets from their surrounding backgrounds. Recently, anomaly detection has been well developed by deep learning technology. However, previous works utilize deep neural networks as a feature extractor followed by a traditional detector to detect anomalies, which cannot separate background and anomaly effectively. In this paper, we try to extract low-rank background features using neural networks to make full use of the low-rank properties of the background and then reconstruct the background with these features to directly separate the anomaly from the background, so we propose an multi-scale background reconstruction network with low-rank embedding (MSBRNet) for anomaly detection in HSI. Firstly, we use a low-rank background features extraction module (LBM) to extract low-rank background features. Then the background is reconstructed using a multi-scale background reconstruction module (MBRM). Finally, we calculate the mean square error of the input image and the output background to measure the effect of background reconstruction, and anomalies appear as reconstruction errors. Experiments on two publicly available experimental datasets demonstrate the significant advantage of the proposed method over other competitors.
Weinan Cao, Hongyan Zhang 0001, Wei He 0003, Hongyu Chen 0003, Ewe Hong Tat
IGARSS3
2022 Non-Local Meets Global: An Iterative Paradigm for Hyperspectral Image Restoration
abstract
Non-local low-rank tensor approximation has been developed as a state-of-the-art method for hyperspectral image (HSI) restoration, which includes the tasks of denoising, compressed HSI reconstruction and inpainting. Unfortunately, while its restoration performance benefits from more spectral bands, its runtime also substantially increases. In this paper, we claim that the HSI lies in a global spectral low-rank subspace, and the spectral subspaces of each full band patch group should lie in this global low-rank subspace. This motivates us to propose a unified paradigm combining the spatial and spectral properties for HSI restoration. The proposed paradigm enjoys performance superiority from the non-local spatial denoising and light computation complexity from the low-rank orthogonal basis exploration. An efficient alternating minimization algorithm with rank adaptation is developed. It is done by first solving a fidelity term-related problem for the update of a latent input image, and then learning a low-dimensional orthogonal basis and the related reduced image from the latent input image. Subsequently, non-local low-rank denoising is developed to refine the reduced image and orthogonal basis iteratively. Finally, the experiments on HSI denoising, compressed reconstruction, and inpainting tasks, with both simulated and real datasets, demonstrate its superiority with respect to state-of-the-art HSI restoration methods.
Wei He 0003, Quanming Yao, Chao Li 0013, Naoto Yokoya, Qibin Zhao, Hongyan Zhang 0001, Liangpei Zhang 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2022 Unmixing Convolutional Features for Crisp Edge Detection
abstract
This article presents a context-aware tracing strategy (CATS) for crisp edge detection with deep edge detectors, based on an observation that the localization ambiguity of deep edge detectors is mainly caused by the mixing phenomenon of convolutional neural networks: Feature mixing in edge classification and side mixing during fusing side predictions. The CATS consists of two modules: A novel tracing loss that performs feature unmixing by tracing boundaries for better side edge learning, and a context-aware fusion block that tackles the side mixing by aggregating the complementary merits of learned side edges. Experiments demonstrate that the proposed CATS can be integrated into modern deep edge detectors to improve localization accuracy. With the vanilla VGG16 backbone, in terms of BSDS500 dataset, our CATS improves the F-measure (ODS) of the RCF and BDCN deep edge detectors by 12 and 6 percent, respectively when evaluating without using the morphological non-maximal suppression scheme for edge detection.
Linxi Huan, Nan Xue 0001, Xianwei Zheng, Wei He 0003, Jianya Gong, Gui-Song Xia
IEEE Trans. Pattern Anal. Mach. Intell.4
2022 Hyperspectral super-resolution via coupled tensor ring factorization
Wei He 0003, Yong Chen 0013, Naoto Yokoya, Chao Li 0013, Qibin Zhao
Pattern Recognit.1
2022 Anisotropic Spatial-Spectral Total Variation Regularized Double Low-Rank Approximation for HSI Denoising and Destriping
abstract
Hyperspectral images (HSIs) can finely discriminate distinct objects with a high spectral resolution, and they are widely employed in various applications. However, mixed noise severely degrades the quality of HSI and restricts the performance of subsequent tasks. As one of the critical pre-processing steps, HSI denoising has been developed rapidly, among which low-rank (LR) prior-based methods have achieved superior performance. Nevertheless, the existing approaches frequently fail to completely remove noise and reconstruct high-quality HSIs when tackling complicated mixed noise with multiple types of high-intensity stripe noise. To solve this problem, we propose an HSI denoising and destriping method based on anisotropic spatial and spectral total variation regularized double low-rank approximation (ATVDLR). The double low-rank approximation framework is devoted to separating the clean image from the mixed noise by exploiting both the global correlations of HSI tensor and the LR structure of stripe noise. Furthermore, the anisotropic spatial and spectral total variation regularization is introduced to preserve the spatial–spectral smoothness of HSI and the directional feature of stripes, thereby further suppressing high-level stripes and Gaussian noise. Finally, the alternating direction method of multipliers (ADMM) technique is designed to solve the proposed ATVDLR model. Extensive experimental results indicate that the proposed method outperforms other state-of-the-art techniques in multi-type high-intensity mixed noise reduction and image structural information protection, and has superior performance in complex mixed noise removal of real Gaofen-5 HSIs.
Jingyi Cai, Wei He 0003, Hongyan Zhang 0001
IEEE Trans. Geosci. Remote. Sens.2
2022 Hyperspectral Image Denoising Using Factor Group Sparsity-Regularized Nonconvex Low-Rank Approximation
abstract
Hyperspectral image (HSI) mixed noise removal is a fundamental problem and an important preprocessing step in remote sensing fields. The low-rank approximation-based methods have been verified effective to encode the global spectral correlation for HSI denoising. However, due to the large scale and complexity of real HSI, previous low-rank HSI denoising techniques encounter several problems, including coarse rank approximation (such as nuclear norm), the high computational cost of singular value decomposition (SVD) (such as Schatten$p$-norm), and adaptive rank selection (such as low-rank factorization). In this article, two novel factor group sparsity-regularized nonconvex low-rank approximation (FGSLR) methods are introduced for HSI denoising, which can simultaneously overcome the mentioned issues of previous works. The FGSLR methods capture the spectral correlation via low-rank factorization, meanwhile utilizing factor group sparsity regularization to further enhance the low-rank property. It is SVD-free and robust to rank selection. Moreover, FGSLR is equivalent to Schatten$p$-norm approximation (Theorem 1), and thus FGSLR is tighter than the nuclear norm in terms of rank approximation. To preserve the spatial information of HSI in the denoising process, the total variation regularization is also incorporated into the proposed FGSLR models. Specifically, the proximal alternating minimization is designed to solve the proposed FGSLR models. Experimental results have demonstrated that the proposed FGSLR methods significantly outperform existing low-rank approximation-based HSI denoising methods.
Yong Chen 0013, Ting-Zhu Huang, Wei He 0003, Xi-Le Zhao, Hongyan Zhang 0001, Jinshan Zeng
IEEE Trans. Geosci. Remote. Sens.3
2022 Exploring Nonlocal Group Sparsity Under Transform Learning for Hyperspectral Image Denoising
abstract
Hyperspectral image (HSI) denoising has been regarded as an effective and economical preprocessing step in data subsequent applications. Recent nonlocal low-rank approximation on each full band patch group has demonstrated their superiority for HSI denoising. These methods, however, directly design the low-rank regularization to the grouped patch image itself (i.e., original domain), which ignores the spatial information of the grouped patch image and cannot explores the potential structure. To address these issues, this paper proposes a nonlocal group sparsifying transform learning method (dubbed TLNLGS) for HSI denoising. Motivated by the global spectral correlation in the HSI, we firstly impose a certain low-dimensional subspace hypothesis over the HSI to prevent the heavy computation burden with the spectral band increases, and then explore a discriminatively intrinsic nonlocal group sparse prior of the reduced image by transform model. The learned group sparse prior can not only excavate the nonlocal self-similarity as recent nonlocal low-rank approximation methods but also preserve the local spatial smooth structure of the image. Moreover, compared with the fixed transform domain (e.g., gradient and discrete cosine transformation domains), the transform learning scheme can improve the sparse representation ability. An efficient block coordinate descent (BCD) algorithm is developed to solve the proposed model. Extensive experiments, including simulated and real HSI datasets, indicate the superiority of the proposed TLNLGS method over the state-of-the-art HSI denoising approaches.
Yong Chen 0013, Wei He 0003, Xi-Le Zhao, Ting-Zhu Huang, Jinshan Zeng, Hui Lin 0002
IEEE Trans. Geosci. Remote. Sens.2
2022 Hyperspectral and Multispectral Image Fusion Using Factor Smoothed Tensor Ring Decomposition
abstract
Fusing a pair of low-spatial-resolution hyperspectral image (LR-HSI) and high-spatial-resolution multispectral image (HR-MSI) has been regarded as an effective and economical strategy to achieve HR-HSI, which is essential to many applications. Among existing fusion models, the tensor ring (TR) decomposition-based model has attracted rising attention due to its superiority in approximating high-dimensional data compared to other traditional matrix/tensor decomposition models. Unlike directly estimating HR-HSI in traditional models, the TR fusion model translates the fusion procedure into an estimate of the TR factor of HR-HSI, which can efficiently capture the spatial–spectral correlation of HR-HSI. Although the spatial–spectral correlation has been preserved well by TR decomposition, the spatial–spectral continuity of HR-HSI is ignored in existing TR decomposition models, sometimes resulting in poor quality of reconstructed images. In this article, we introduce a factor smoothed regularization for TR decomposition to capture the spatial–spectral continuity of HR-HSI. As a result, our proposed model is calledfactor smoothed TR decompositionmodel, dubbedFSTRD. In order to solve the suggested model, we develop an efficient proximal alternating minimization algorithm. A series of experiments on four synthetic datasets and one real-world dataset show that the quality of reconstructed images can be significantly improved by the introduced factor smoothed regularization, and thus, the suggested method yields the best performance by comparing it to state-of-the-art methods.
Yong Chen 0013, Jinshan Zeng, Wei He 0003, Xi-Le Zhao, Ting-Zhu Huang
IEEE Trans. Geosci. Remote. Sens.3
2022 Unsupervised Spectral-Spatial Semantic Feature Learning for Hyperspectral Image Classification
abstract
Can we automatically learn meaningful semantic feature representations when training labels are absent? Several recent unsupervised deep learning approaches have attempted to tackle this problem by solving the data reconstruction task. However, these methods can easily latch on low-level features. To solve this problem, we propose an end-to-end spectral–spatial semantic feature learning network (S3FN) for unsupervised deep semantic feature extraction (FE) from hyperspectral images (HSIs). Our main idea is to learn spectral-spatial features from high-level semantic perspective. First, we utilize the feature transformation to obtain two feature descriptions of the same source data from different views. Then, we propose the spectral–spatial feature learning network to project the two feature descriptions into the deep embedding space. Subsequently, a contrastive loss function is introduced to align the two projected features, which should have the same implied semantic meaning. The proposed S3FN learns the spectral and spatial features separately, and then merges them. Finally, the learned spectral–spatial features by S3FN are processed by a classifier to evaluate their effectiveness. Experimental results on three publicly available HSI datasets show that our proposed S3FN can produce promising classification results with a lower time cost than other state-of-the-art (SOTA) deep learning-based unsupervised FE methods.
Huilin Xu, Wei He 0003, Liangpei Zhang 0001, Hongyan Zhang 0001
IEEE Trans. Geosci. Remote. Sens.2
2022 Breaking Limits of Remote Sensing by Deep Learning From Simulated Data for Flood and Debris-Flow Mapping
abstract
We propose a framework that estimates the inundation depth (maximum water level) and debris-flow-induced topographic deformation from remote sensing imagery by integrating deep learning and numerical simulation. A water and debris-flow simulator generates training data for various artificial disaster scenarios. We show that regression models based on Attention U-Net and LinkNet architectures trained on such synthetic data can predict the maximum water level and topographic deformation from a remote sensing-derived change detection map and a digital elevation model. The proposed framework has an inpainting capability, thus mitigating the false negatives that are inevitable in remote sensing image analysis. Our framework breaks limits of remote sensing and enables rapid estimation of inundation depth and topographic deformation, essential information for emergency response, including rescue and relief activities. We conduct experiments with both synthetic and real data for two disaster events that caused simultaneous flooding and debris flows and demonstrate the effectiveness of our approach quantitatively and qualitatively. Our code and data sets are available athttps://github.com/nyokoya/dlsim.
Naoto Yokoya, Kazuki Yamanoi, Wei He 0003, Gerald Baier, Bruno Adriano, Hiroyuki Miura, Satoru Oishi
IEEE Trans. Geosci. Remote. Sens.3
2022 Double Low-Rank Matrix Decomposition for Hyperspectral Image Denoising and Destriping
abstract
Hyperspectral images (HSIs) have a wealth of applications in many areas, due to their fine spectral discrimination ability. However, in the practical imaging process, HSIs are often degraded by a mixture of various types of noise, for example, Gaussian noise, impulse noise, dead pixels, dead lines, and stripe noise. Low-rank matrix decomposition theory has been widely used in HSI denoising, and has achieved competitive results by modeling the impulse noise, dead pixels, dead lines, and stripe noise as sparse components. However, the existing low-rank-based methods for HSI denoising cannot completely remove stripe noise when the stripe noise is no longer sparse. In this article, we extend the HSI observation model and propose a double low-rank (DLR) matrix decomposition method for HSI denoising and destriping. By simultaneously exploring the low-rank characteristic of the lexicographically ordered noise-free HSI and the low-rank structure of the stripe noise on each band of the HSI, the two low-rank constraints are formulated into one unified framework, to achieve separation of the noise-free HSI, stripe noise, and other mixed noise. The proposed DLR model is then solved by the augmented Lagrange multiplier (ALM) algorithm efficiently. Both simulation and real HSI data experiments were carried out to verify the superiority of the proposed DLR method.
Hongyan Zhang 0001, Jingyi Cai, Wei He 0003, Huanfeng Shen, Liangpei Zhang 0001
IEEE Trans. Geosci. Remote. Sens.3
2022 A Gather-to-Guide Network for Remote Sensing Semantic Segmentation of RGB and Auxiliary Image
abstract
Convolutional neural network (CNN)-based feature fusion of RGB and auxiliary remote sensing data is known to enable improved semantic segmentation. However, such fusion is challengeable because of the substantial variance in data characteristics and quality (e.g., data uncertainties and misalignment) between two modality data. In this article, we propose a unified gather-to-guide network (G2GNet) for remote sensing semantic segmentation of RGB and auxiliary data. The key aspect of the proposed architecture is a novel gather-to-guide module (G2GM) that consists of a feature gatherer and a feature guider. The feature gatherer generates a set of cross-modal descriptors by absorbing the complementary merits of RGB and auxiliary modality data. The feature guider calibrates the RGB feature response by using the channel-wise guide weights extracted from the cross-modal descriptors. In this way, the G2GM can perform RGB feature calibration with different modality data in a gather-to-guide fashion, thus preserving the informative features while suppressing redundant and noisy information. Extensive experiments conducted on two benchmark datasets show that the proposed G2GNet is robust to data uncertainties while also improving the semantic segmentation performance of RGB and auxiliary remote sensing data.
Xianwei Zheng, Xiujie Wu, Linxi Huan, Wei He 0003, Hongyan Zhang 0001
IEEE Trans. Geosci. Remote. Sens.4
2022 LESSFormer: Local-Enhanced Spectral-Spatial Transformer for Hyperspectral Image Classification
abstract
Currently, the convolutional neural networks (CNNs) have become the mainstream methods for hyperspectral image (HSI) classification, due to their powerful ability to extract local features. However, CNNs fail to effectively and efficiently capture the long-range contextual information and diagnostic spectral information of HSI. In contrast, the leading-edge vision transformers are capable of capturing long-range dependencies and processing sequential data such as spectral signatures. Nevertheless, pre-existing transformer-based classification methods generally generate inaccurate token embeddings from a single spectral or spatial dimension of raw HSIs and encounter difficulty modeling locality with insufficient training data. To mitigate these limitations, we propose a novel local-enhanced spectral-spatial transformer method (i.e., LESSFormer) specifically devised for HSI classification. Two effective and efficient modules are designed in LESSFormer, i.e., the HSI2Token module and the local-enhanced transformer encoder. The former is devised to transform HSI into the adaptive spectral-spatial tokens, and the latter is built to further enhance the representation ability of tokens by reinforcing the local information explicitly with a simple attention mask as well as retaining the long-range information in the meantime. Extensive experimental results on the new Xiong’an dataset and the widely used Pavia University and Houston University datasets have shown the superiority of LESSFormer over other state-of-the-art networks.
Jiaqi Zou, Wei He 0003, Hongyan Zhang 0001
IEEE Trans. Geosci. Remote. Sens.2
2021 Fast Hyperspectral Image Recovery of Dual-Camera Compressive Hyperspectral Imaging via Non-Iterative Subspace-Based Fusion
abstract
Coded aperture snapshot spectral imaging (CASSI) is a promising technique for capturing three-dimensional hyperspectral images (HSIs), in which algorithms are used to perform the inverse problem of HSI reconstruction from a single coded two-dimensional (2D) measurement. Due to the ill-posed nature of this problem, various regularizers have been exploited to reconstruct 3D data from 2D measurements. Unfortunately, the accuracy and computational complexity are unsatisfactory. One feasible solution is to utilize additional information such as the RGB measurement in CASSI. Considering the combined CASSI and RGB measurements, in this paper, we propose a fusion model for HSI reconstruction. Specifically, we investigate the low-dimensional spectral subspace property of HSIs composed of a spectral basis and spatial coefficients. In particular, the RGB measurement is utilized to estimate the coefficients, while the CASSI measurement is adopted to provide the spectral basis. We further propose a patch processing strategy to enhance the spectral low-rank property of HSIs. The optimization of the proposed model requires neither iteration nor the spectral sensing matrix of the RGB detector. Extensive experiments on both simulated and real HSI datasets demonstrate that our proposed method not only outperforms previous state-of-the-art (iterative algorithms) methods in quality but also speeds up the reconstruction by more than 5000 times.
Wei He 0003, Naoto Yokoya, Xin Yuan 0002
IEEE Trans. Image Process.1
2020 Guided Deep Decoder: Unsupervised Image Pair Fusion
Tatsumi Uezato, Danfeng Hong, Naoto Yokoya, Wei He 0003
ECCV (6)4
2020 Hyperspectral Image Restoration Using Weighted Group Sparsity-Regularized Low-Rank Tensor Decomposition
abstract
Mixed noise (such as Gaussian, impulse, stripe, and deadline noises) contamination is a common phenomenon in hyperspectral imagery (HSI), greatly degrading visual quality and affecting subsequent processing accuracy. By encoding sparse prior to the spatial or spectral difference images, total variation (TV) regularization is an efficient tool for removing the noises. However, the previous TV term cannot maintain the shared group sparsity pattern of the spatial difference images of different spectral bands. To address this issue, this article proposes a group sparsity regularization of the spatial difference images for HSI restoration. Instead of using ℓ1or ℓ2-norm (sparsity) on the difference image itself, we introduce a weighted ℓ2,1norm to constrain the spatial difference image cube, efficiently exploring the shared group sparse pattern. Moreover, we employ the well-known low-rank Tucker decomposition to capture the global spatial-spectral correlation from three HSI dimensions. To summarize, a weighted group sparsity-regularized low-rank tensor decomposition (LRTDGS) method is presented for HSI restoration. An efficient augmented Lagrange multiplier algorithm is employed to solve the LRTDGS model. The superiority of this method for HSI restoration is demonstrated by a series of experimental results from both simulated and real data, as compared with the other state-of-the-art TV-regularized low-rank matrix/tensor decomposition methods.
Yong Chen 0013, Wei He 0003, Naoto Yokoya, Ting-Zhu Huang
IEEE Trans. Cybern.2
2020 Robust Nonlocal Low-Rank SAR Time Series Despeckling Considering Speckle Correlation by Total Variation Regularization
abstract
Outliers and speckle both corrupt time series of synthetic aperture radar (SAR) acquisitions. Owing to the coherence between SAR acquisitions, their speckle can no longer be regarded as independent. In this study, we propose an algorithm for nonlocal low-rank time series despeckling, which is robust against outliers and also specifically addresses speckle correlation between acquisitions. By imposing total variation regularization on the signal's speckle component, the correlation between acquisitions can be identified, facilitating the extraction of outliers from unfiltered signals and the correlated speckle. This robustness against outliers also addresses matching errors and inaccuracies in the nonlocal similarity search. Such errors include mismatched data in the nonlocal estimation process, which degrade the denoising performance of conventional similarity-based filtering approaches. Multiple experiments on real and synthetic data assess the performance of the approach by comparing it with state-of-the-art methods. It provides filtering results of comparable quality but is not adversely affected by outliers. The source code is available at https://github.com/gbaier/nllrtv.
Gerald Baier, Wei He 0003, Naoto Yokoya
IEEE Trans. Geosci. Remote. Sens.2
2020 Nonlocal Tensor-Ring Decomposition for Hyperspectral Image Denoising
abstract
Hyperspectral image (HSI) denoising is a fundamental problem in remote sensing and image processing. Recently, nonlocal low-rank tensor approximation-based denoising methods have attracted much attention due to their advantage of being capable of fully exploiting the nonlocal self-similarity and global spectral correlation. Existing nonlocal low-rank tensor approximation methods were mainly based on two common decomposition [Tucker or CANDECOMP/PARAFAC (CP)] methods and achieved the state-of-the-art results, but they are subject to certain issues and do not produce the best approximation for a tensor. For example, the number of parameters for Tucker decomposition increases exponentially according to its dimensions, and CP decomposition cannot better preserve the intrinsic correlation of the HSI. In this article, a novel nonlocal tensor-ring (TR) approximation is proposed for HSI denoising by using TR decomposition to explore the nonlocal self-similarity and global spectral correlation simultaneously. TR decomposition approximates a high-order tensor as a sequence of cyclically contracted third-order tensors, which has strong ability to explore these two intrinsic priors and to improve the HSI denoising results. Moreover, an efficient proximal alternating minimization algorithm is developed to optimize the proposed TR decomposition model efficiently. Extensive experiments on three simulated data sets under several noise levels and two real data sets verify that the proposed TR model provides better HSI denoising results than several state-of-the-art methods in terms of quantitative and visual performance evaluations.
Yong Chen 0013, Wei He 0003, Naoto Yokoya, Ting-Zhu Huang, Xi-Le Zhao
IEEE Trans. Geosci. Remote. Sens.2
2020 Hyperspectral Image Denoising With Total Variation Regularization and Nonlocal Low-Rank Tensor Decomposition
abstract
Hyperspectral images (HSIs) are normally corrupted by a mixture of various noise types, which degrades the quality of the acquired image and limits the subsequent application. In this article, we propose a novel denoising method for the HSI restoration task by combining nonlocal low-rank tensor decomposition and total variation regularization, which we refer to as TV-NLRTD. To simultaneously capture the nonlocal similarity and high spectral correlation, the HSI is first segmented into overlapping 3-D cubes that are grouped into several clusters by the k-means++ algorithm and exploited by low-rank tensor approximation. Spatial-spectral total variation (SSTV) regularization is then investigated to restore the clean HSI from the denoised overlapping cubes. Meanwhile, the ℓ1-norm facilitates the separation of the clean nonlocal low-rank tensor groups and the sparse noise. The proposed TV-NLRTD method is optimized by employing the efficient alternating direction method of multipliers (ADMM) algorithm. The experimental results obtained with both simulated and real hyperspectral data sets confirm the validity and superiority of the proposed method compared with the current state-of-the-art HSI denoising algorithms.
Hongyan Zhang 0001, Wei He 0003, Liangpei Zhang 0001
IEEE Trans. Geosci. Remote. Sens.3
2020 Double-Factor-Regularized Low-Rank Tensor Factorization for Mixed Noise Removal in Hyperspectral Image
abstract
As a preprocessing step, hyperspectral image (HSI) restoration plays a critical role in many subsequent applications. Recently, based on the framework of subspace representation and low-rank matrix/tensor factorization (LRMF/LRTF), many single-factor-regularized methods add various regularizations on the spatial factor to characterize its spatial prior knowledge. However, these methods neglect the common characteristics among different bands and the spectral continuity of HSIs. To tackle this issue, this article establishes a bridge between the factor-based regularization and the HSI priors and proposes a double-factor-regularized LRTF model for HSI mixed noise removal. The proposed model employs LRTF to characterize the spectral global low rankness, introduces a weighted group sparsity constraint on the spatial difference images (SpatDIs) of the spatial factor to promote the group sparsity in the SpatDIs of HSIs, and suggests a continuity constraint on the spectral factor to promote the spectral continuity of HSIs. Moreover, we develop a proximal alternating minimization-based algorithm to solve the proposed model. Extensive experiments conducted on the simulated and real HSIs demonstrate that the proposed method has superior performance on mixed noise removal compared with the state-of-the-art methods based on subspace representation, noise modeling, and LRMF/LRTF.
Yu-Bang Zheng, Ting-Zhu Huang, Xi-Le Zhao, Yong Chen 0013, Wei He 0003
IEEE Trans. Geosci. Remote. Sens.5
2020 Hyperspectral Image Compressive Sensing Reconstruction Using Subspace-Based Nonlocal Tensor Ring Decomposition
abstract
Hyperspectral image compressive sensing reconstruction (HSI-CSR) can largely reduce the high expense and low efficiency of transmitting HSI to ground stations by storing a few compressive measurements, but how to precisely reconstruct the HSI from a few compressive measurements is a challenging issue. It has been proven that considering the global spectral correlation, spatial structure, and nonlocal self-similarity priors of HSI can achieve satisfactory reconstruction performances. However, most of the existing methods cannot simultaneously capture the mentioned priors and directly design the regularization term to the HSI. In this article, we propose a novel subspace-based nonlocal tensor ring decomposition method (SNLTR) for HSI-CSR. Instead of designing the regularization of the low-rank approximation to the HSI, we assume that the HSI lies in a low-dimensional subspace. Moreover, to explore the nonlocal self-similarity and preserve the spatial structure of HSI, we introduce a nonlocal tensor ring decomposition strategy to constrain the related coefficient image, which can decrease the computational cost compared to the methods that directly employ the nonlocal regularization to HSI. Finally, a well-known alternating minimization method is designed to efficiently solve the proposed SNLTR. Extensive experimental results demonstrate that our SNLTR method can significantly outperform existing approaches for HSI-CSR.
Yong Chen 0013, Ting-Zhu Huang, Wei He 0003, Naoto Yokoya, Xi-Le Zhao
IEEE Trans. Image Process.3
2020 Illumination Invariant Hyperspectral Image Unmixing Based on a Digital Surface Model
abstract
Although many spectral unmixing models have been developed to address spectral variability caused by variable incident illuminations, the mechanism of the spectral variability is still unclear. This paper proposes an unmixing model, named illumination invariant spectral unmixing (IISU). IISU makes the first attempt to use the radiance hyperspectral data and a LiDAR-derived digital surface model (DSM) in order to physically explain variable illuminations and shadows in the unmixing framework. Incident angles, sky factors, visibility from the sun derived from the LiDAR-derived DSM support the explicit explanation of endmember variability in the unmixing process from radiance perspective. The proposed model was efficiently solved by a straightforward optimization procedure. The unmixing results showed that the other state-of-the-art unmixing models did not work well especially in the shaded pixels. On the other hand, the proposed model estimated more accurate abundances and shadow compensated reflectance than the existing models.
Tatsumi Uezato, Naoto Yokoya, Wei He 0003
IEEE Trans. Image Process.3
2019 Non-Local Meets Global: An Integrated Paradigm for Hyperspectral Denoising
abstract
Non-local low-rank tensor approximation has been developed as a state-of-the-art method for hyperspectral image (HSI) denoising. Unfortunately, while their denoising performance benefits little from more spectral bands, the running time of these methods significantly increases. In this paper, we claim that the HSI lies in a global spectral low-rank subspace, and the spectral subspaces of each full band patch groups should lie in this global low-rank subspace. This motivates us to propose a unified spatial-spectral paradigm for HSI denoising. As the new model is hard to optimize, An efficient algorithm motivated by alternating minimization is developed. This is done by first learning a low-dimensional orthogonal basis and the related reduced image from the noisy HSI. Then, the non-local low-rank denoising and iterative regularization are developed to refine the reduced image and orthogonal basis, respectively. Finally, the experiments on synthetic and both real datasets demonstrate the superiority against the stateof-the-art HSI denoising methods.
Wei He 0003, Quanming Yao, Chao Li 0013, Naoto Yokoya, Qibin Zhao
CVPR1
2019 Guaranteed Matrix Completion Under Multiple Linear Transformations
abstract
Low-rank matrix completion (LRMC) is a classical model in both computer vision (CV) and machine learning, and has been successfully applied to various real applications. In the recent CV tasks, the completion is usually employed on the variants of data, such as "non-local" or filtered, rather than their original forms. This fact makes that the theoretical analysis of the conventional LRMC is no longer suitable in these applications. To tackle this problem, we propose a more general framework for LRMC, in which the linear transformations of the data are taken into account. We rigorously prove the identifiability of the proposed model and show an upper bound of the reconstruction error. Furthermore, we derive an efficient completion algorithm by using augmented Lagrangian multipliers and the sketching trick. In the experiments, we apply the proposed method to the classical image inpainting problem and achieve the state-of-the-art results.
Chao Li 0013, Wei He 0003, Longhao Yuan, Zhun Sun, Qibin Zhao
CVPR2
2019 Total-variation-regularized Tensor Ring Completion for Remote Sensing Image Reconstruction
abstract
In recent studies, tensor ring (TR) decomposition has shown to be effective in data compression and representation. However, the existing TR-based completion methods only exploit the global low-rank property of the visual data. When applying them to remote sensing (RS) image processing, the spatial information in the RS image is ignored. In this paper, we introduce the TR decomposition to RS image processing and propose a tensor completion method for RS image reconstruction. We incorporate the total-variation regularization into the TR completion model to exploit the low-rank property and spatial continuity of the RS image simultaneously. The proposed algorithm is solved by the augmented Lagrange multiplier method and has shown the superior performance in hyperspectral image reconstruction and multi-temporal RS image cloud removal against the state-of-the-art algorithms.
Wei He 0003, Longhao Yuan, Naoto Yokoya
ICASSP1
2019 Weighted Group Sparsity Regularized Low-Rank Tensor Decomposition for Hyperspectral Image Restoration
abstract
Total variation (TV) regularization has been widely used in the hyperspectral image (HSI) mixed noise removal problem by utilizing ℓ1-norm to constrain the spatial difference image and promote the piecewise smooth structure. Unfortunately, it cannot depict the group sparse structure of spatial difference image along the spectral dimension. This paper proposes a new HSI restoration method using weighted group sparsity regularized low-rank tensor decomposition (LRTDGS). Specifically, we use a weighted group sparsity regularization which is denoted by ℓ2,1-norm to explore the group structure of spatial difference image along the spectral dimension. Moreover, the spatial-spectral correlation from three directions of HSI is depicted by low-rank Tucker decomposition. We use efficient augmented Lagrange multiplier method to optimize the proposed LRTDGS model, and a series of experimental results are presented to demonstrate the effectiveness of the proposed method.
Yong Chen 0013, Wei He 0003, Naoto Yokoya, Ting-Zhu Huang
IGARSS2
2019 Total Variation Regularized Low-Rank Sparsity Decomposition for Blind Cloud and Cloud Shadow Removal from Multitemporal Imagery
abstract
This paper proposes a spatial-spectral total variation (TV) regularized low-rank sparsity decomposition model for blind cloud and cloud shadow (cloud/shadow) detection and removal of multitemporal remote sensing imagery. Our concept is to decompose the contaminated image into the surface-reflected component and the cloud/shadow component. Low-rank regularization is utilized to model the spectral-temporal correlation of the surface-reflected component, meanwhile, the `1-norm and spatial-spectral total variation regularization is employed to describe the sparse prior and spatial-spectral continuity of the cloud/shadow component. To better preserve the information in cloud/shadow-free areas, the cloud/shadow detection results obtained as a by-product of our method are used to guide the information compensation from the original contaminated images. Several experiments are presented to demonstrate the effectiveness of the proposed method.
Yong Chen 0013, Wei He 0003, Naoto Yokoya, Ting-Zhu Huang
IGARSS2
2019 Robust Nonlocal Low-Rank Sar Stack Despeckling With Application To Change Detection
abstract
We present a nonlocal low-rank denoising algorithm for synthetic aperture radar (SAR) image stacks. The method extends the widely known DespecKS algorithm by integrating low-rank approximation, outlier removal, and total variation (TV) regularization into the estimation process. Preliminary experiments shows increased robustness against outliers and comparable performance to state-of-the-art stack despeckling algorithms.
Gerald Baier, Wei He 0003, Bruno Adriano, Junshi Xia, Naoto Yokoya
IGARSS2
2019 Superpixel-based spatial-spectral dimension reduction for hyperspectral imagery classification
Huilin Xu, Hongyan Zhang 0001, Wei He 0003, Liangpei Zhang 0001
Neurocomputing3
2019 Remote Sensing Image Reconstruction Using Tensor Ring Completion and Total Variation
abstract
Time-series remote sensing (RS) images are often corrupted by various types of missing information such as dead pixels, clouds, and cloud shadows that significantly influence the subsequent applications. In this paper, we introduce a new low-rank tensor decomposition model, termed tensor ring (TR) decomposition, to the analysis of RS data sets and propose a TR completion method for the missing information reconstruction. The proposed TR completion model has the ability to utilize the low-rank property of time-series RS images from different dimensions. To further explore the smoothness of the RS image spatial information, total-variation regularization is also incorporated into the TR completion model. The proposed model is efficiently solved using two algorithms, the augmented Lagrange multiplier (ALM) and the alternating least square (ALS) methods. The simulated and real-data experiments show superior performance compared to other state-of-the-art low-rank related algorithms.
Wei He 0003, Naoto Yokoya, Longhao Yuan, Qibin Zhao
IEEE Trans. Geosci. Remote. Sens.1
2018 Superpixel Based Dimension Reduction for Hyperspectral Imagery
abstract
This paper focuses on dimension reduction (DR) technique for hyperspectral image (HSI). In this paper, we proposed a superpixel-based linear discriminant analysis (SP-LDA) dimension reduction method for HSI classification. Pixels within a local spatial neighborhood are expected to have similar spectral curves and share the same class label. To fully exploit the spatial structure, superpixel segmentation is firstly introduced to generate the superpixel map, which can adaptively explore the neighborhood structure information. Moreover, we extend the SP-LDA algorithm by combining the extracted feature from spectral and spatial dimensions, which can fully exploit complementary and consistent information from both dimensions. The experimental results on two standard hyperspectral datasets confirm the superiority of the proposed algorithms.
Huilin Xu, Hongyan Zhang 0001, Wei He 0003, Liangpei Zhang 0001
IGARSS3
2017 Total Variation Regularized Reweighted Sparse Nonnegative Matrix Factorization for Hyperspectral Unmixing
abstract
Blind hyperspectral unmixing (HU), which includes the estimation of endmembers and their corresponding fractional abundances, is an important task for hyperspectral analysis. Recently, nonnegative matrix factorization (NMF) and its extensions have been widely used in HU. Unfortunately, most of the NMF-based methods can easily lead to an unsuitable solution, due to the nonconvexity of the NMF model and the influence of noise. To overcome this limitation, we make the best use of the structure of the abundance maps, and propose a new blind HU method named total variation regularized reweighted sparse NMF (TV-RSNMF). First, the abundance matrix is assumed to be sparse, and a weighted sparse regularizer is incorporated into the NMF model. The weights of the weighted sparse regularizer are adaptively updated related to the abundance matrix. Second, the abundance map corresponding to a single fixed endmember should be piecewise smooth. Therefore, the TV regularizer is adopted to capture the piecewise smooth structure of each abundance map. In our multiplicative iterative solution to the proposed TV-RSNMF model, the TV regularizer can be regarded as an abundance map denoising procedure, which improves the robustness of TV-RSNMF to noise. A number of experiments were conducted in both simulated and real-data conditions to illustrate the advantage of the proposed TV-RSNMF method for blind HU.
Wei He 0003, Hongyan Zhang 0001, Liangpei Zhang 0001
IEEE Trans. Geosci. Remote. Sens.1
2016 Hyperspectral unmixing using total variation regularized reweighted sparse non-negative matrix factorization
abstract
Recently, non-negative matrix factorization (NMF) model has been widely used in hyperspectral unmixing (HU). In this paper, based on NMF, we explore the properties of abundance maps, and propose a new blind HU algorithm named total variation regularized reweighted sparse NMF (TV-RSNMF). Typically, only a subset of endmembers are assumed to generate the fixed pixel. As a result, the abundance maps are assumed to be sparse. So we introduce a weighted sparse regularization to explore the sparsity of abundance maps in the NMF model. In addition, the abundance maps related to fixed material are assumed to be piecewise smooth and we adopt a total variation (TV) regularizer to promote the piecewise smooth property. TV regularizer can be regarded as an abundance maps denoising procedure, which significantly improves the robustness of the proposed method to noise. Several experiments were conducted to illustrate the performance of the proposed algorithm.
Wei He 0003, Hongyan Zhang 0001, Liangpei Zhang 0001
IGARSS1
2016 Weighted Sparse Graph Based Dimensionality Reduction for Hyperspectral Images
abstract
Dimensionality reduction (DR) is an important and helpful preprocessing step for hyperspectral image (HSI) classification. Recently, sparse graph embedding (SGE) has been widely used in the DR of HSIs. SGE explores the sparsity of the HSI data and can achieve good results. However, in most cases, locality is more important than sparsity when learning the features of the data. In this letter, we propose an extended SGE method: the weighted sparse graph based DR (WSGDR) method for HSIs. WSGDR explicitly encourages the sparse coding to be local and pays more attention to those training pixels that are more similar to the test pixel in representing the test pixel. Furthermore, WSGDR can offer data-adaptive neighborhoods, which results in the proposed method being more robust to noise. The proposed method was tested on two widely used HSI data sets, and the results suggest that WSGDR obtains sparser representation results. Furthermore, the experimental results also confirm the superiority of the proposed WSGDR method over the other state-of-the-art DR methods.
Wei He 0003, Hongyan Zhang 0001, Liangpei Zhang 0001, Wilfried Philips, Wenzi Liao
IEEE Geosci. Remote. Sens. Lett.1
2016 Total-Variation-Regularized Low-Rank Matrix Factorization for Hyperspectral Image Restoration
abstract
In this paper, we present a spatial spectral hyperspectral image (HSI) mixed-noise removal method named total variation (TV)-regularized low-rank matrix factorization (LRTV). In general, HSIs are not only assumed to lie in a low-rank subspace from the spectral perspective but also assumed to be piecewise smooth in the spatial dimension. The proposed method integrates the nuclear norm, TV regularization, and L1-norm together in a unified framework. The nuclear norm is used to exploit the spectral low-rank property, and the TV regularization is adopted to explore the spatial piecewise smooth structure of the HSI. At the same time, the sparse noise, which includes stripes, impulse noise, and dead pixels, is detected by the L1-norm regularization. To tradeoff the nuclear norm and TV regularization and to further remove the Gaussian noise of the HSI, we also restrict the rank of the clean image to be no larger than the number of endmembers. A number of experiments were conducted in both simulated and real data conditions to illustrate the performance of the proposed LRTV method for HSI restoration.
Wei He 0003, Hongyan Zhang 0001, Liangpei Zhang 0001, Huanfeng Shen
IEEE Trans. Geosci. Remote. Sens.1
2014 A noise-adjusted iterative randomized singular value decomposition method for hyperspectral image denoising
abstract
In this paper, a new denoising algorithm is proposed for hyperspectral image data cubes. With the strong correlations of the image bands, the low-rank structure of the hyperspectral image is explored by lexicographically ordering the 3-D data cube into 2-D matrix. Based on this property, the traditional principal component analysis (PCA) denoising model is established. For hyperspectral images (HSIs), the noise intensity in different bands is different. Therefore, a noise-adjusted iterative randomized singular value decomposition (NAIRSVD) algorithm is proposed to solve this PCA model. Combined with adaptive noise estimation and upper bound rank estimation, the proposed NAIRSVD algorithm is free from manual parameter determination. Several experiments were conducted to illustrate the performance of the proposed algorithm.
Wei He 0003, Hongyan Zhang 0001, Liangpei Zhang 0001, Huanfeng Shen
IGARSS1
2014 Hyperspectral Image Restoration Using Low-Rank Matrix Recovery
abstract
Hyperspectral images (HSIs) are often degraded by a mixture of various kinds of noise in the acquisition process, which can include Gaussian noise, impulse noise, dead lines, stripes, and so on. This paper introduces a new HSI restoration method based on low-rank matrix recovery (LRMR), which can simultaneously remove the Gaussian noise, impulse noise, dead lines, and stripes. By lexicographically ordering a patch of the HSI into a 2-D matrix, the low-rank property of the hyperspectral imagery is explored, which suggests that a clean HSI patch can be regarded as a low-rank matrix. We then formulate the HSI restoration problem into an LRMR framework. To further remove the mixed noise, the “Go Decomposition” algorithm is applied to solve the LRMR problem. Several experiments were conducted in both simulated and real data conditions to verify the performance of the proposed LRMR-based HSI restoration method.
Hongyan Zhang 0001, Wei He 0003, Liangpei Zhang 0001, Huanfeng Shen, Qiangqiang Yuan
IEEE Trans. Geosci. Remote. Sens.2