Xiangchao Meng

dblp:170/9965 · DBLP profile ↗
← Back
67ranked-venue papers
11as first author
59since 2021 · last 2026
0000-0002-7405-3143ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 49 · 10 first-author · 43 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 1 first-author · 13 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021
YearPublicationVenuePosition
2026 Interactive feature fusion for camera-radar-based vehicle segmentation in bird's-eye view
Chenyang Lu 0002, Liang Li 0010, Xiangchao Meng, Qiuping Jiang, Feng Shao 0001
Pattern Recognit.4
2025 A Progressive Spatial-Spectral Interactive Network for Integrated Fusion of Panchromatic, Multispectral, and Hyperspectral Images
abstract
Satellite-based hyperspectral (HS) imagery holds great potential in remote sensing applications due to its fine spectral resolution. However, the low spatial resolution limits its practical utility. Combining ancillary high resolution panchromatic (PAN) or multispectral (MS) images has become a common practice to improve the spatial quality of HS images. Most present approaches, however, are based on dual-sensor fusion (e.g., MS-HS or PAN-HS), which generally falls short of comprehensively integrating their complementary spatial and spectral information of PAN, MS, and HS images. Meanwhile, existing few integrated fusion methods suffer from two key limitations: modality mismatch due to inconsistent spatial-spectral characteristics among PAN, MS, and HS data, and shallow and redundant cross-modal coupling caused by inadequate modeling of inter-modal relationships. In this paper, we propose a progressive spatial-spectral interactive network (PSSNet) for the integrated fusion of panchromatic, multispectral, and hyperspectral images. Specifically, a context-aware fusion block is introduced to extract and enhance contextual spatial and spectral information across different modalities. To ensure an effective integration of spatial and spectral details, the entire network is structured progressively, allowing for a smooth transition and fusion of features from HS, MS, and PAN images. Additionally, a spatial-spectral feature recombination module is designed to dynamically adjust the contribution of spectral features at various levels. This module, in combination with a spatial enhancement component, facilitates the optimal fusion of spatial and spectral information by enhancing their interactions. Extensive experiments on simulated and real datasets, both qualitatively and quantitatively, demonstrate the superiority of PSSNet compared to other state-of-the-art methods.
Yufu Bai, Minchao Luo, Shenfu Zhang, Qiang Liu 0035, Weiwei Sun 0005, Xiangchao Meng
IEEE Trans. Geosci. Remote. Sens.7
2025 Adversarial Robust Salient Object Detection in Optical Remote Sensing Images With Implicit Feature Enhancement
abstract
Deep neural networks (DNNs) have achieved significant progress in optical remote sensing images salient object detection (ORSI-SOD) and are widely applied to various remote sensing image analysis tasks. However, few SOD models demonstrate robust performance under adversarial perturbations, which ultimately leads to a decline in detection accuracy. Moreover, most existing defense methods inject fixed Gaussian noise globally into the image. Although such approaches are easy to implement, they have several limitations in inaccurate uncertainty estimation and neglecting the unique characteristics of local salient regions. Furthermore, existing adversarial defense research rarely addresses the challenges specific to the ORSI-SOD task, leaving a gap in effective defense strategies. To tackle these issues, we propose a novel defense method, which enhances the adversarial robustness of ORSI-SOD models through implicit feature enhancement. The algorithm first proposes a two-stage strategy of reverse local noise search and forward global noise optimization, enhancing generalization ability by implicitly enhancing features to better simulate network uncertainty. Then, the algorithm proposes a global-guided texture information enhancement (GTIE) module for low-level features and a global-guided semantics information enhancement (GSIE) module for high-level features, focusing on strengthening low-level texture information and enhancing the model’s understanding of high-level contextual semantic features, respectively. This dual-module design effectively weakens the impact of adversarial noise, significantly improving the robustness and accuracy of object detection. Extensive experiments on three ORSI-SOD datasets demonstrate that our defense strategy better estimates the uncertainty, resulting in an average performance improvement of 23.2% in$F_{\beta } ^{\mathrm { max}}$and 34.1% in$E_{\xi } ^{\mathrm { max}}$across six ORSI-SOD models under five different adversarial attack methods. Our code will be released in the public repository athttps://github.com/kexi0714/IFe.
Feng Shao 0001, Xiangchao Meng, Hangwei Chen, Xiongli Chai, Zhiyi Mo
IEEE Trans. Geosci. Remote. Sens.3
2025 Dual-Branch Cross-Resolution Interaction Learning Network for Change Detection at Different Resolutions
abstract
Change detection (CD) plays a critical role in remote sensing (RS) image analysis. However, the detection accuracy is often compromised due to differences in imaging conditions between bitemporal images, especially in scenarios where images have varying spatial resolutions from multisource RS satellites. To address this challenge, we propose an innovative dual-branch cross-resolution interaction learning network (DCILNet). This network employs strategies for image spatial resolution alignment and feature space correction to achieve efficient CD. To fully leverage the information from images with different resolutions, we design two cross-resolution CD branches. These branches interact through a cross-resolution feature correction module (CRFCM) and a multiresolution feature fusion module (MRFM), facilitating feature interaction and learning across branches to maximize feature representation. We conduct both qualitative and quantitative experiments on three public datasets. The experimental results show that, compared to other comparative methods, the proposed DCILNet exhibits stronger competitiveness. Our research shows that by integrating image spatial resolution alignment and feature space correction strategies and adopting dual-branch interactive learning, the model effectively addresses the challenges posed by resolution discrepancies in CD tasks. Our code will be available athttps://github.com/Li738/DCILNet.
Feng Shao 0001, Xiangchao Meng
IEEE Trans. Geosci. Remote. Sens.3
2025 IM-CMDet: An Intramodal Enhancement and Cross-Modal Fusion Network for Small Object Detection in UAV Aerial Visible-Infrared Imagery
abstract
UAV aerial Visible-Infrared (RGBT) object detection has been widely applied in fields such as military operations and rescue missions. However, although numerous UAV aerial RGBT object detection methods exist, several challenges remain in this field. On the one hand, drones typically operate at high altitudes, and objects only occupy a small number of pixels in imaging, posing a significant challenge to object detection. On the other hand, spatial misalignment between modalities remains a major obstacle in cross-modal fusion—especially given the small size of the objects. To address the above issues, this paper proposes IM-CMDet, an intra-modal enhancement and cross-modal fusion network for small object detection in UAV-based RGBT imagery, which comprises three effective modules: the Detail-Semantics Joint Enhancement module (DSJE), the Differential-based Fusion Weight Generation module (DFWG) and the Feature Reconstruction Network (FRN). The DSJE module prevents small object features from being overwhelmed by background noise through optimizing feature representations across different levels. The FRN module is designed to overcome modality differences and build inter-modality information correlation via swin-Transformer architecture. To further enhance the network’s sensitivity to small objects, the DFWG combines differential and spatial attention to generate the final fusion weights while reducing the impact of background noise on detection performance. Extensive experiments on RGBTDronePerson and two additional benchmarks demonstrate that IM-CMDet achieves state-of-the-art performance through effective cross-modal fusion, significantly advancing small-object detection in complex aerial scenarios. The code is available at https://github.com/RS-Minchao/IM-CMDet.
Minchao Luo, Rui Zhao 0003, Shenfu Zhang, Feng Shao 0001, Xiangchao Meng
IEEE Trans. Geosci. Remote. Sens.6
2025 Dual-Task Cascaded Network for Spatial-Temporal-Spectral Remote Sensing Image Fusion
abstract
Spatial-temporal-spectral fusion is dedicated to integrating the complementary advantages of multisource images to obtain fused image with all high spatial, high temporal and high spectral resolutions, which is promising but more challenging. On the one hand, traditional studies deployed on MODIS and Landsat data cannot be transferred to most spaceborne hyperspectral (HS) data with lower temporal resolution; on the other hand, the rigid time relation modeling in most existing studies exhibits weakness orienting to non-linear land-cover changes. In this paper, we propose a dual-task cascaded network for spatial-temporal-spectral fusion, with collaborative modeling on spatialspectral joint enhancement and temporal variation estimation in a unified framework. The spatial-spectral joint enhancement task was designed with an iterative alternating projection, meticulously crafted to address the scale variance among observation. Additionally, the spatial enhancement unit and error correction unit were coupled modeling to enhance the spatial and spectral fidelity. The temporal variation estimation on spectral fine tuning network was developed, to further enhance the temporal and spectral fidelity. Extensive experiments were implemented on Ziyuan(ZY)-1 02D HS data and Sentinel-2 multispectral (MS) data. Both qualitative and quantitative results demonstrated the competitive performance of the proposed method.
Xiangchao Meng, Xu Chen 0041, Mengjing Zhang, Feng Shao 0001, Gang Yang 0006, Weiwei Sun 0005
IEEE Trans. Geosci. Remote. Sens.1
2025 Integrated Fusion for Panchromatic, Multispectral, Hyperspectral Remote Sensing Images: Insights From Multispectral Images
abstract
The integrated fusion of the high-spatial-resolution (HR) panchromatic image (PAN), the relative “moderate”-spatial-resolution (MR) multispectral image (MSI), and the low-spatial-resolution (LR) hyperspectral image (HSI), to generate the optimal HR HS fused image, is promising but challenging. On the one hand, existing mainstream fusion models mostly focus on the “pairwise fusion” between HR PAN, MR MSI, and LR HSI, which cannot sufficiently integrate their complementary spatial and spectral advantages. On the other hand, one of the few integrated fusion methods roughly introduced the MR MSI as a simple intermediate medium; however, the role of MSIs as a spatial and spectral “bridge” between HR PAN and LR HSI, generally characterized by significant scale differences, remains largely unexplored. To solve these problems, we proposed an integrated PAN–MSI–HSI fusion method from the perspective of MSIs, by comprehensively considering the scale difference among the multisource observations. In the proposed method, a spatial–spectral feature transfer network was designed by comprehensively exploring the spatial–spectral variations and connections among the HR PAN, MR MSI, and LR HSI. Then, a spatial–spectral joint reconstruction module was constructed to reconstruct the HR HSI with optimal spatial and spectral fidelity. Experiments were conducted on simulated and real datasets from qualitative and quantitative aspects. The experimental results demonstrated the competitive effectiveness over other state-of-the-art methods.
Xiangchao Meng, Xiangjun Meng, Yufu Bai, Shenfu Zhang, Qiang Liu 0035, Gang Yang 0006, Weiwei Sun 0005
IEEE Trans. Geosci. Remote. Sens.1
2025 Dual-Domain Aligned Temporal-Spatial-Spectral Fusion Networks for No-Paired Hyperspectral and Multispectral Images
Jiawen Weng, Weiwei Sun 0005, Kai Ren 0003, Gang Yang 0006, Xiangchao Meng, Jiangtao Peng
IEEE Trans. Geosci. Remote. Sens.5
2025 Bidirectional Spectral Attention Multiscale Aggregation Network for Spectral Super-Resolution
abstract
Spectral super-resolution (SSR) is the computational process of generating a high-dimensional hyperspectral image from a low-dimensional image through spectral reconstruction techniques. Recently, deep learning has demonstrated remarkable potential in the field of SSR, achieving impressive results. However, existing deep learning-based approaches often fail to deliver high-fidelity SSR outcomes. These methods tend to focus primarily on spectral information while paying insufficient attention to the critical role of spatial features. Furthermore, they lack effective strategies for capturing inter-band relationships, resulting in suboptimal spectral information modeling. To address these limitations, we propose a novel network for SSR, termed Bidirectional Spectral Attention Multi-Scale Aggregation Network (BiSANet). BiSANet features three U-Net-like branches and integrates two advanced attention mechanisms. The bidirectional spectral attention modules dynamically model inter-spectral dependencies through forward and reverse spectral feature extraction, enhanced by a weight-sharing strategy. Specifically, we reverse the spectral order of feature maps to activate complementary global trends and local details, overcoming the limitations of unidirectional modeling in traditional methods. Additionally, an independent spatial reconstruction branch with a dedicated loss function ensures precise spatial detail preservation. Experimental results demonstrate that BiSANet outperforms state-of-the-art methods across three benchmarks. For instance, on the DFC2018 Houston dataset, it achieves a 4.26% PSNR improvement and an 11.52% SAM reduction, highlighting its robustness and accuracy in spectral-spatial reconstruction.
Xintao Zhong, Shenfu Zhang, Gang Yang 0006, Weiwei Sun 0005, Feng Shao 0001, Xiangchao Meng
IEEE Trans. Geosci. Remote. Sens.7
2025 GCM-PDA: A Generative Compensation Model for Progressive Difference Attenuation in Spatiotemporal Fusion of Remote Sensing Images
abstract
High-resolution satellite imagery with dense temporal series is crucial for long-term surface change monitoring. Spatiotemporal fusion seeks to reconstruct remote sensing image sequences with both high spatial and temporal resolutions by leveraging prior information from multiple satellite platforms. However, significant radiometric discrepancies and large spatial resolution variations between images acquired from different satellite sensors, coupled with the limited availability of prior data, present major challenges to accurately reconstructing missing data using existing methods. To address these challenges, this paper introduces GCM-PDA, a novel generative compensation model with progressive difference attenuation for spatiotemporal fusion of remote sensing images. The proposed model integrates multi-scale image decomposition within a progressive fusion framework, enabling the efficient extraction and integration of information across scales. Additionally, GCM-PDA employs domain adaptation techniques to mitigate radiometric inconsistencies between heterogeneous images. Notably, this study pioneers the use of style transformation in spatiotemporal fusion to achieve spatial-spectral compensation, effectively overcoming the constraints of limited prior image information. Experimental results demonstrate that GCM-PDA not only achieves competitive fusion performance but also exhibits strong robustness across diverse conditions.
Kai Ren 0003, Weiwei Sun 0005, Xiangchao Meng, Gang Yang 0006
IEEE Trans. Image Process.3
2025 Spatial-Spectral Heterogeneity-Aware Network for Hyperspectral and LiDAR Joint Classification
abstract
The integration of hyperspectral (HS) imagery and light detection and ranging (LiDAR) data for land cover classification has emerged as a prominent research focus. Despite the satisfactory classification accuracies achieved by existing methodologies, several unaddressed issues that remain warrant consideration. First, current approaches overlook the pronounced spectral and spatial heterogeneities in remote sensing (RS) images designated for multiclassification tasks, limiting the performance of classification models. Moreover, most existing studies amalgamate elevation features with other characteristics through simple addition and interaction operations, and they do not delve deeply into exploiting elevation height information, leading to an imbalance in the representation of elevation height. In light of the aforementioned issues, this article introduces a spatial-spectral heterogeneity-aware network (S2HANet) for the joint classification of HS and LiDAR data. Specifically, a shared spectral correction module (SSCM) is designed in the spectral branch to preliminarily alleviate the problem of large intraclass variance, followed by the use of a contrastive learning framework to enhance the intraclass compactness and interclass separability of spectral features. A multichannel signed distance discrimination module (MCSDDM) is developed to learn the distance relationships between intra- and interclass pixels and boundaries, and using prior boundary information to improve spatial boundary information. In addition, an elevation boost module (EBM) and an elevation injection module (EIM) are meticulously designed to phase-in elevation height information, further enhancing the utilization of elevation data and better facilitating the fusion of the two modalities. The proposed S2HANet has demonstrated exceptional classification performance across three opening benchmark datasets.
Shenfu Zhang, Qiang Liu 0035, Rui Zhao 0003, Feng Shao 0001, Xiangchao Meng
IEEE Trans. Neural Networks Learn. Syst.7
2024 A Lightweight and Enhanced Semantic Segmentation Network for Mapping of Retrogressive Thaw Slumps from Sentinel-2 Images
abstract
Fine mapping of retrogressive thaw slumps (RTSs) holds paramount significance in the study of permafrost degradation and carbon exchange. We propose a lightweight and enhanced semantic segmentation network (LessNet) for automatically mapping the RTSs from Sentinel-2 images. LessNet is constructed on the encoder-decoder framework with innovative incorporation of attention mechanism and dual-level semantic features fusion. The lightweight architecture of LessNet eliminates the need for pre-training, and the network hyperparameters are automatically updated based on the training dataset, which allows for fast convergence of supervised learning. Experiments conducted in the Beiluhe region of the Tibetan Plateau highlight the robustness and competitive performance of the model.
Guiyun Zhou, Zhonghua Su, Weiwei Sun 0005, Xiangchao Meng
IGARSS6
2024 Multi-domain pseudo-reference quality evaluation for infrared and visible image fusion
abstract
Abstract Infrared and visible image fusion involves merging the advantages of infrared and visible images to generate a composite image that encompasses thermal radiation as well as intricate texture details. Infrared and visible image fusion has garnered increasing attention, with numerous fusion methods proposed. However, how to fairly perceive the performance of fused image remains a contentious topic. This paper is dedicated to solving this problem from two perspectives (e.g., subjective and objective aspects). Firstly, an infrared and visible fusion image quality assessment dataset was constructed, including 60 pairs of infrared and visible images captured in various scenes, along with 540 fusion images with different types and degrees of distortions. Additionally, a subjective evaluation dataset of 16,200 subjective scores by 30 participants was further provided for the fused image. Secondly, to overcome the challenging assessment for infrared and visible fusion images without a real reference image, an interesting multi‐domain pseudo‐reference image quality assessment model (MPIQAM) is proposed, by comprehensively considering the thermal radiation information distortion, texture information distortion, and overall naturalness of the fused image. The proposed MPIQAM was compared with 18 mainstream objective metrics, and the experimental findings showcased a commendable level of competitiveness.
Xiangchao Meng, Chaoqi Chen, Qiang Liu 0035, Feng Shao 0001
IET Image Process.1
2024 Integrating Multitemporal SAR and Optical Information for Missing Optical Imagery Generation
abstract
Cloud cover and long revisit cycle of satellites can cause gaps in optical images and pose a significant obstacle to the consistency of Earth observation missions. Recently, synthetic aperture radar (SAR)-to-optical image translation (S2OIT) has become an emerging approach to reconstruct the missing information of optical remote sensing images. However, the previous studies ignored the mechanism difference between SAR and optical data and produced color distortion, image blurriness, and texture detail loss in the generated optical images. To tackle these challenges, we propose a multitemporal S2OIT network (MTS2ONet) for high-quality optical image generation. The proposed model comprises two subnetworks: change feature extraction subnetwork (Change_Extractor) and the S2OIT subnetwork (S2O_Translator). The first subnetwork is tasked with extracting change features from SAR images captured at dates T and${T} +1$, and then translating them from the SAR domain to the optical domain. Subsequently, the S2O_Translator integrates the optical image at date${T} +1$with the change features extracted by the Change_Extractor to generate the optical image at date T. In addition, we produce a dual-temporal SAR-optical dataset called DTSEN1-2 for model evaluation. Experiments on the DTSEN1-2 dataset reveal that our method is superior to the state-of-the-art (SOTA) methods with the metrics peak-signal-to-noise ratio (PSNR; 36.0435), structural similarity index measure (SSIM; 0.9896), learned perceptual image patch similarity (LPIPS; 0.0443), and root mean square error (RMSE; 0.0174) and exhibits preferable results in visual effects. Our dataset and codes can be accessed via the following link:https://github.com/hopeupup/MTS2ONet.
Chunyu Dong, Gang Yang 0006, Weiwei Sun 0005, Xiangchao Meng, Binjie Chen
IEEE Trans. Geosci. Remote. Sens.5
2024 Multiscale Spatial-Spectral Invertible Compensation Network for Hyperspectral Remote Sensing Image Denoising
abstract
Hyperspectral image (HSI) has fine spectral resolution and abundant spatial information to detect subtle differences between targets. However, it is heavily contaminated with noise due to sensor design and atmospheric radiative transfer, resulting in spectral shifts and spatial discontinuities. Current denoising methods usually establish constraints directly on the ground truth and denoised image, lacking supervision of intermediate parameters of the network, resulting in insufficient model constraints and poor convergence. In addition, existing methods do not consider spatial-spectral compensation, so the denoising results have obvious spatial-spectral distortion. To this end, we propose a novel multiscale spatial-spectral invertible compensation network (MSIC-Net) for HSI denoising. The method constructs an invertible spatial-spectral compensation (ISSC) module, which supervises intermediate features through inverse constraints, realizes the circulation of multiscale information, and improves the stability of the model. At the same time, we also introduce style transfer for spatial-spectral compensation, which uses its superior fine feature control ability to precisely compensate for the lost spatial and spectral detail features. The method is extensively validated experimentally and categorically on simulated and real datasets. The experimental results show that MSIC-Net outperforms other state-of-the-art denoising methods in quantitative and qualitative evaluations.
Huiyang Li, Kai Ren 0003, Weiwei Sun 0005, Gang Yang 0006, Xiangchao Meng
IEEE Trans. Geosci. Remote. Sens.5
2024 Uncertain Category-Aware Fusion Network for Hyperspectral and LiDAR Joint Classification
abstract
The integration of hyperspectral (HS) imagery and light detection and ranging (LiDAR) for land cover classification has become a significant research topic. Numerous existing methods aim to interactively fuse the complementary features of HS and LiDAR to enhance the classification accuracy. However, most existing studies overlook the fact that spectral, spatial, and elevation features of HS and LiDAR possess significant discriminative information for specific categories. The rough and simple interacting or stacking these features may hinder the effective expression of this significant discriminative information. Moreover, existing approaches neglect the shared spatial characteristics between HS and LiDAR. In this article, an uncertain category-aware fusion network (UCAFNet) is proposed to tackle the above challenges. Specifically, we proposed an uncertain category-aware fusion strategy (UCAFS) that dynamically weights the spectral, spatial, and elevation branches based on their respective capabilities in identifying different categories to achieve targeted information aggregation. Moreover, we introduce the spatial information purification module (SIPM) and adaptive weighted fusion module (AWFM), to extract and enhance shared spatial features from HS and LiDAR for effective integration. The experimental results on three public benchmark datasets demonstrate the superior performance of the proposed UCAFNet.
Xiangchao Meng, Shenfu Zhang, Qiang Liu 0035, Gang Yang 0006, Weiwei Sun 0005
IEEE Trans. Geosci. Remote. Sens.1
2024 Multistage Hybrid Denoising Network for Satellite Hyperspectral Images
abstract
The hyperspectral imaging instrument makes a trade-off by sacrificing spatial resolution to achieve high spectral resolution. This compromise leads to a low signal-to-noise ratio, and hyperspectral images (HSIs) are often heavily contaminated with mixed noise, which is an inherent challenge. Previous research has achieved satisfactory results for natural image denoising; hyperspectral denoising has remained a formidable task. In this article, we introduce an innovative method called the multistage hybrid-denoising network for satellite hyperspectral images (SUC-MSDN). SUC-MSDN initially decomposes the noisy HSI into multiple scales and constructs a multistage denoising network by analyzing the spatial spectrum texture distribution characteristics of noise signals. Instead of simply stacking the output results from each scale, SUC-MSDN uses the denoising results from the low-scale network as prior knowledge for the high-scale denoising network to more accurately remove the final noise components. Extensive experimental datasets are used to validate the performance of SUC-MSDN. Experimental results show that SUC-MSDN outperforms benchmark methods and significantly enhances the accuracy of land cover mapping.
Kai Ren 0003, Weiwei Sun 0005, Gang Yang 0006, Xiangchao Meng, Jiangtao Peng, Huiyang Li
IEEE Trans. Geosci. Remote. Sens.4
2024 Domain Transform Model Driven by Deep Learning for Anti-Noise Hyperspectral and Multispectral Image Fusion
abstract
While fusion of hyperspectral images (HSIs) with low spatial resolution and multispectral images (MSIs) with high spatial resolution has achieved significant success, high-quality fusion between noisy images has always been challenging. In this article, we propose a domain transform model driven by deep learning for anti-noise hyperspectral and multispectral image fusion (DTAFN). This marks the first time that wavelet decomposition theory is combined with deep learning for noise reduction in hyperspectral and MSI fusion. DTAFN initially decomposes hyperspectral and MSIs into frequency components and constructs a novel feature interaction fusion module (FIFM). This module, while using MSIs to guide the removal of noise from HSIs, also achieves the fusion of spatial and spectral information. Furthermore, it maps the fused features to a lower dimensional subspace to enhance computational efficiency. Additionally, we introduce a spatial-spectral self-attention mechanism to optimize the reconstructed frequency components using the subspace features. In the end, the wavelet inverse transform is used to reconstruct the clean fused image. It is worth noting that the extraction of the subspace is considered a process of nonlinear low-rank component extraction, which, to a certain extent, suppresses noise signals. Numerous experiments of mixed noise image fusion are carried out, and the experimental results show that DTAFN can obtain high-quality fusion results, is robust, and superior to the state-of-the-art methods.
Weiwei Sun 0005, Kai Ren 0003, Xiangchao Meng, Gang Yang 0006, Jiancheng Li, Jingfeng Huang
IEEE Trans. Geosci. Remote. Sens.3
2024 CIG-STF: Change Information Guided Spatiotemporal Fusion for Remote Sensing Images
abstract
Spatiotemporal fusion has been attracting increasing attention in remote sensing applications, such as environmental monitoring and land cover change detection, due to its excellent ability to obtain high spatial and temporal resolution images. The land cover change has always been a great challenge in spatiotemporal fusion. Although most spatiotemporal fusion methods have demonstrated satisfactory performance in addressing phenological changes, the performance in terms of abrupt land cover type changes, such as floods or mudslides, falls short. To alleviate this issue, we propose a change information guided spatiotemporal fusion (CIG-STF) method. The proposed CIG-STF integrates change detection and spatiotemporal fusion in a unified framework, by taking advantage of change detection in capturing land cover changes to assist spatiotemporal fusion. Specifically, the CIG-STF comprises three modules: multiscale dilated feature extractor module (MDFE), spatiotemporal fusion-change detection integrated module (STF-CD), and reconstruction module. The MDFE employs multiscale dilated convolutions to comprehensively extract features, to prevent crucial information loss by increasing the convolutional receptive field. In the STF-CD, a change detection module on attention strategy is integrated into the spatiotemporal fusion task, by excavating land cover changes to further enhance the fusion performance. In addition, we design a dynamic decay loss function to further leverage change information, ensuring the accuracy of both change information and prediction results. The experiments were verified on the publicly available LGC and Daxing datasets with manual change labels. The experimental results demonstrate the superior performance of the proposed CIG-STF in both phenological variations and land cover type changes.
Mingzhu You, Xiangchao Meng, Qiang Liu 0035, Feng Shao 0001, Randi Fu
IEEE Trans. Geosci. Remote. Sens.2
2024 Collaborative Learning and Style-Adaptive Pooling Network for Perceptual Evaluation of Arbitrary Style Transfer
abstract
Although the research of arbitrary style transfer (AST) has achieved great progress in recent years, few studies pay special attention to the perceptual evaluation of AST images that are usually influenced by complicated factors, such as structure-preserving, style similarity, and overall vision (OV). Existing methods rely on elaborately designed hand-crafted features to obtain quality factors and apply a rough pooling strategy to evaluate the final quality. However, the importance weights between the factors and the final quality will lead to unsatisfactory performances by simple quality pooling. In this article, we propose a learnable network, named collaborative learning and style-adaptive pooling network (CLSAP-Net) to better address this issue. The CLSAP-Net contains three parts, i.e., content preservation estimation network (CPE-Net), style resemblance estimation network (SRE-Net), and OV target network (OVT-Net). Specifically, CPE-Net and SRE-Net use the self-attention mechanism and a joint regression strategy to generate reliable quality factors for fusion and weighting vectors for manipulating the importance weights. Then, grounded on the observation that style type can influence human judgment of the importance of different factors, our OVT-Net utilizes a novel style-adaptive pooling strategy guiding the importance weights of factors to collaboratively learn the final quality based on the trained CPE-Net and SRE-Net parameters. In our model, the quality pooling process can be conducted in a self-adaptive manner because the weights are generated after understanding the style type. The effectiveness and robustness of the proposed CLSAP-Net are well validated by extensive experiments on the existing AST image quality assessment (IQA) databases. Our code will be released at https://github.com/Hangwei-Chen/CLSAP-Net.
Hangwei Chen, Feng Shao 0001, Xiongli Chai, Qiuping Jiang, Xiangchao Meng, Yo-Sung Ho
IEEE Trans. Neural Networks Learn. Syst.5
2023 Multidiscriminator Supervision-Based Dual-Stream Interactive Network for High-Fidelity Cloud Removal on Multitemporal SAR and Optical Images
abstract
Optical remote sensing images have the advantages in clear visual characteristics and strong interpretability. Unfortunately, cloud coverage limits the quality and availability of optical images in practical applications. In contrast, Synthetic Aperture Radar (SAR) images provide all-day and all-weather imaging, which can serve as effective auxiliary information for cloud removal. Existing cloud removal methods are difficult to obtain high-fidelity cloud-free results due to the insufficient spectral and spatial information exploration in the multitemporal SAR and optical images. In this paper, we propose a multi-discriminator supervision-based dual-stream interactive network (MDS-DIN) for cloud removal. Specifically, we first design a dual-stream interactive learning module to take full advantage of the complementary information between multitemporal SAR and optical images. Moreover, we specially design an adaptive weight fusion module to adaptively allocate fusion weights to the dual-stream results by considering the discriminative features in spectral and spatial levels. In addition, multi-discriminator is employed to jointly optimize overall networks for high-fidelity cloud removal. Experiments on simulated and real data sets demonstrate the competitive performance of our proposed method.
Zhenfei Wang, Qiang Liu 0035, Xiangchao Meng, Wei Jin 0003
IEEE Geosci. Remote. Sens. Lett.3
2023 Modality-Induced Transfer-Fusion Network for RGB-D and RGB-T Salient Object Detection
abstract
The ability of capturing the complementary information of multi-modality data is critical to the development of multi-modality salient object detection (SOD). Most of existing studies attempt to integrate multi-modality information through various fusion strategies. However, most of these methods ignore the inherent differences in multi-modality data, resulting in poor performance when dealing with some challenging scenarios. In this paper, we propose a novel Modality-Induced Transfer-Fusion Network (MITF-Net) for RGB-D and RGB-T SOD by fully exploring the complementarity in multi-modality data. Specifically, we first deploy a modality transfer fusion (MTF) module to bridge the semantic gap between single and multi-modality data, and then mine the cross-modality complementarity based on point-to-point structural similarity information. Then, we design a cycle-separated attention (CSA) module to optimize the cross-layer information recurrently, and measure the effectiveness of cross-layer features through point-wise convolution-based multi-scale channel attention. Furthermore, we refine the boundaries in the decoding stage to obtain high-quality saliency maps with sharp boundaries. Extensive experiments on 13 RGB-D and RGB-T SOD datasets show that the proposed MITF-Net achieves a competitive and excellent performance.
Feng Shao 0001, Xiongli Chai, Hangwei Chen, Qiuping Jiang, Xiangchao Meng, Yo-Sung Ho
IEEE Trans. Circuits Syst. Video Technol.6
2023 Quality Evaluation of Arbitrary Style Transfer: Subjective Study and Objective Metric
abstract
Arbitrary neural style transfer is a vital topic with great research value and wide industrial application, which strives to render the structure of one image using the style of another. Recent researches have devoted great efforts on the task of arbitrary style transfer (AST) for improving the stylization quality. However, there are very few explorations about the quality evaluation of AST images, even it can potentially guide the design of different algorithms. In this paper, we first construct a new AST images quality assessment database (AST-IQAD), which consists 150 content-style image pairs and the corresponding 1200 stylized images produced by eight typical AST algorithms. Then, a subjective study is conducted on our AST-IQAD database, which obtains the subjective rating scores of all stylized images on the three subjective evaluations, i.e., content preservation (CP), style resemblance (SR), and overall vision (OV). To quantitatively measure the quality of AST image, we propose a new sparse representation-based method, which computes the quality according to the sparse feature similarity. Experimental results on our AST-IQAD have demonstrated the superiority of the proposed method. The dataset and source code will be released athttps://github.com/Hangwei-Chen/AST-IQAD-SRQE
Hangwei Chen, Feng Shao 0001, Xiongli Chai, Yuese Gu, Qiuping Jiang, Xiangchao Meng, Yo-Sung Ho
IEEE Trans. Circuits Syst. Video Technol.6
2023 Cross-Modality Double Bidirectional Interaction and Fusion Network for RGB-T Salient Object Detection
abstract
RGB-T salient object detection (SOD) aims to detect and segment saliency regions on RGB images and the corresponding thermal maps. The ability of alleviating the modality difference between RGB and thermal modality plays a vital role in the development of RGB-T SOD. However, most of the existing methods try to integrate multi-modal information through various fusion strategies, or reduce the modality difference via unidirectional or undifferentiated bidirectional interaction, but failing in some challenging scenes. To deal with the above question, a novel Cross-Modality Double Bidirectional Interaction and Fusion Network (CMDBIF-Net) for RGB-T SOD is proposed. Specifically, we construct an interactive branch to indirectly bridge the RGB and thermal modalities. In addition, we propose a double bidirectional interaction (DBI) module composed of a forward interaction block (FIB) and a backward interaction block (BIB) to reduce the cross-modality differences. Moreover, a multi-scale feature enhancement and fusion (MSFEF) module is introduced to integrate the multi-modal features with considering the internal gap of different modality. Finally, we use a cascaded decoder and a cross-level feature enhancement (CLFE) module to generate high-quality saliency map. Extensive experiments are conducted on three publicly available RGB-T SOD datasets shows that the proposed CMDBIF-Net achieves outstanding performance against the state-of-the-art (SOTA) RGB-T SOD methods.
Zhengxuan Xie, Feng Shao 0001, Hangwei Chen, Qiuping Jiang, Xiangchao Meng, Yo-Sung Ho
IEEE Trans. Circuits Syst. Video Technol.6
2023 FTDN: Multispectral and Hyperspectral Image Fusion With Diverse Temporal Difference Spans
abstract
Multispectral (MS)-hyperspectral (HS) image fusion, which aims to enhance the spatial resolution of low spatial resolution HS images with a high spatial resolution MS has provided a wide range of applications in remote sensing. However, relatively long revisit cycles of HS satellites and irresistible weather factors cause the acquisition of HS and MS images at the same time difficult. Most of the existing approaches neglect the temporal difference between MS and HS images, and perform weakness in the challenging case with diverse temporal difference spans. In this paper, we propose a novel image fusion strategy with embedding a stage of feature matching before interaction. On the one hand, we explore the role of spectral correlation modeling between HS and MS images, which accounts for the utilization of available spatial information from MS images. On the other hand, we design a feature aggregation module to fully exploit the nonlinear gaps and dependencies of heterogeneous data and utilize adaptive gains to realize complementary information projection and fusion. We build Dongying (DY) and Yellow River Estuary (YRE) remote sensing datasets based on Sentinel-2 and ZiYuan(ZY)-1 02D satellites with diverse temporal difference spans. The extensive experiments demonstrate that our method is robust to the span of temporal difference and shows superior performance over the existing methods visually and quantitatively.
Xu Chen 0041, Xiangchao Meng, Qiang Liu 0035, Huiping Jiang, Gang Yang 0006, Weiwei Sun 0005, Feng Shao 0001
IEEE Trans. Geosci. Remote. Sens.2
2023 PSSTFN: A Progressive Spatial-Temporal-Spectral Fusion Network for Remote Sensing Images
abstract
Spatial-temporal-spectral fusion (STSF) is highly desirable to generate dense-time image series with high spatial and spectral resolution by integrating the complementary advantages of multi-source and multi-temporal observations. However, most existing STSF methods are still limited to the assumption of linear temporal, spatial and spectral relations. In addition, the STSF methods on Landsat and MODIS data are insufficient to characterize the inherent properties of the current spaceborne hyperspectral images with low spatial and temporal resolutions. For these, we propose a progressive STSF network (PSSTFN) by interestingly integrating spatial-spectral fusion and spectral-temporal fusion into a unified end-to-end STSF framework. Specifically, in the spatial-spectral fusion stage, we obtain the hierarchical features with different receptive fields and propose a multi-attention guided module for joint learning and refinement of spatial-spectral features. In the spectral-temporal fusion stage, a feature insertion module is presented to embed the difference images into the resulting spatial-spectral features, and the estimation from deeper layers is cascaded for more reliable spatial information. We build Dongying (DY) and Yellow River Estuary (YRE) remote sensing datasets based on Sentinel-2 and ZiYuan(ZY)-1 02D satellites for verification, and the experimental results on reduce- and full-resolution data demonstrate the superior performance of our method over the existing methods visually and quantitatively.
Xu Chen 0041, Xiangchao Meng, Feng Shao 0001, Weiwei Sun 0005
IEEE Trans. Geosci. Remote. Sens.2
2023 Dual-Task Interactive Learning for Unsupervised Spatio-Temporal-Spectral Fusion of Remote Sensing Images
abstract
Spatio-temporal-spectral fusion aims to produce high spatio-temporal-spectral resolution images by integrating the complementary spatial, temporal, and spectral advantages of multi-source remote sensing images. However, on one hand, existing spatio-temporal-spectral fusion methods are insufficient to exploit the inherent complex nonlinear spatial, temporal, and spectral relationship among multisource and multitemporal observations. On the other hand, since the unavailability of real high spatio-temporal-spectral resolution images, it is difficult to adopt deep learning methods with supervised training. In this paper, we propose an effective Unsupervised Spatio-Temporal-Spectral Fusion Model (USTSFM) with dual-task interactive learning to alleviate these problems. The proposed USTSFM has two branches: the Spatio-Temporal-Spectral Mapping (STSM) branch is to describe the temporal relationship, and the Spectral Super Resolution (SSR) branch is to model the spectral relationship. Moreover, the spatial-spectral interaction compensation block is designed to make the two branches compensate and benefited from each other. This intrinsically related and mutually facilitated strategy allows the USTSFM to sufficiently exploit the inherent spatial, temporal, and spectral relationship. In addition, a shared reconstruction module is meticulously designed for the two tasks, which not only reduces the parameters but also allows the supervised task to guide the convergence of the unsupervised task, boosting the stability of unsupervised training. The qualitative and quantitative results demonstrated the proposed USTSFM has richer spatial details and more accurate predictions than the other state-of-the-art methods.
Qiang Liu 0035, Xu Chen 0041, Xiangchao Meng, Hangwei Chen, Feng Shao 0001, Weiwei Sun 0005
IEEE Trans. Geosci. Remote. Sens.3
2023 Self-Supervised Feature Learning Based on Spectral Masking for Hyperspectral Image Classification
abstract
Deep learning has emerged as a powerful method for hyperspectral image (HSI) classification. However, a significant prerequisite for HSI classification using deep learning is enough labeled samples, which is both time-consuming and labor-intensive. Yet, labeled samples are essential for training deep learning models. This paper proposes an HSI classification method based on the self-supervised learning of spectral masking (SSLSM). The method mainly includes two steps: self-supervised pre-training and fine-tuning. First, considering the rich spectral information of HSI, we propose masked spectral reconstruction as the pretext task. The unmasked data is input into the encoder and decoder sequentially, which are composed of a multi-layer transformer, for feature learning for masked spectral reconstruction. Second, we use reference samples to fine-tune the network, and the encoder and decoder are innovatively cascaded for deep semantic feature extraction, which can further improve the ability of feature extraction in the downstream classification tasks. Experiment results show that, compared with other methods, the SSLSM obtains the highest classification accuracy of 96.52%, 97.03%, and 96.70% on the Indian Pines dataset, Pavia University dataset, and Yancheng Wetlands dataset, respectively. Our method can also be applied to other HSI datasets, and the codes will be available from https://github.com/CIRSM-GRoup/2023-TGRS-SSLSM.
Weiwei Liu 0009, Weiwei Sun 0005, Gang Yang 0006, Kai Ren 0003, Xiangchao Meng, Jiangtao Peng
IEEE Trans. Geosci. Remote. Sens.6
2023 Detail Injection-Based Spatio-Temporal Fusion for Remote Sensing Images With Land Cover Changes
abstract
Spatio-temporal fusion can generate time-series images with high spatial resolution, and it is highly desirable in various applications, especially in monitoring fine dynamic changes of surface features on remote sensing images. Currently, most spatio-temporal fusion methods predict the target fine image by employing the auxiliary fine images on neighboring phases; however, they are generally limited in abrupt land cover changes between the target and the neighboring auxiliary images. In this paper, we propose a novel Detail Injection-based Spatio-Temporal Fusion (DISTF) model to alleviate this problem, by exploring the inherent relationship between the spatio-temporal fusion and spatio-spectral fusion. The proposed DISTF consists of three modules: a Three-branch Detail Injection (TDI) module, a Fine Detail Prediction (FDP) module, and a reconstruction module. The interpretable TDI module is inspired by spatio-spectral fusion, aiming to inject the non-changed detail information extracted from the neighboring fine images into the target coarse image, which can preserve the abrupt change information captured in the target coarse image. The FDP module is designed to further integrate the correlated information from the outputs of TDI and refine the spatial-spectral information to boost the fusion accuracy. Finally, the reconstruction module and the hybrid loss function are designed to more effective reconstruct the high-quality target fine image. The qualitative and quantitative experiment results on two datasets with different types of changes demonstrated that the proposed DISTF method achieves richer spatial detail and more accurate prediction than the eight existing methods.
Qiang Liu 0035, Xiangchao Meng, Xinghua Li 0002, Feng Shao 0001
IEEE Trans. Geosci. Remote. Sens.2
2023 Domain Adaptive Cross Reconstruction for Change Detection of Heterogeneous Remote Sensing Images via a Feedback Guidance Mechanism
abstract
Change detection on heterogeneous optical and synthetic aperture radar (SAR) images is soaring and plays a crucial role in monitoring land cover changes, such as disaster emergencies and natural resource monitoring. This is commonly recognized as a promising but challenging work due to the intrinsic differences in imaging mechanisms between the optical and SAR images. Recently, deep learning-based change detection methods based on two-step processing have attracted attention, i.e., first image translation between optical and SAR images to alleviate their modality differences and then change detection based on the translated images. However, image translation itself is a trouble task for the heterogeneous optical and SAR images. The unreliable image translation results further limit the accuracy of change detection. In this paper, to mitigate this problem, we propose a change detection model on domain adaptation by novelty integrating change detection and image reconstruction into a unified framework. Specifically, we first transform the optical and SAR images into an intermediate common domain for comparison. Moreover, cross reconstruction for optical and SAR images is designed to maintain the characteristics of the images and improve the performance of domain adaptation. In addition, a feedback guidance mechanism is circumspectly designed to co-optimize change detection and image reconstruction tasks. Extensive experiments were conducted on four publicly available datasets, the results demonstrate the effectiveness of our proposed method.
Qiang Liu 0035, Kai Ren 0003, Xiangchao Meng, Feng Shao 0001
IEEE Trans. Geosci. Remote. Sens.3
2023 MGFEI-Net: Multiscale Grouping Feedback Embedded Integrated Network for Panchromatic, Multispectral, and Hyperspectral Image Fusion
abstract
The spaceborne hyperspectral (HS) imagery with fine spectral information has broad application aspects; however, the low spatial resolution has limited the potential application values. Over the past few decades, a general strategy to improve the spatial resolution of the HS is to fuse the low spatial resolution (LR) HS with an auxiliary moderate spatial resolution (MR) multispectral (MS) or a high spatial resolution (HR) panchromatic (PAN) image. However, most of the existing methods mainly focus on two-sensor fusion with the LR HS and MR MS images (i.e., MS-HS fusion) or the LR HS and HR PAN images (i.e., the PAN-HS fusion). How to comprehensively combine the complementary spatial and spectral advantages of the LR HS, MR MS, and HR PAN observations, to obtain the optimal high-fidelity HR HS image is interesting and challenging. In this paper, we propose a multi-scale grouping feedback embedded integrated fusion network (MGFEI-Net) for the LR HS, MR MS, and HR PAN images. Specifically, an attention-based hybrid-scale integrated module is designed by considering the spatial scale diversity of the HR PAN, MR MS, and LR HS images. Moreover, a multi-scale grouping feedback embedded module with a top-to-bottom manner is proposed to capture more usual spatial-spectral features. Experiments were performed on the simulated and real datasets. Moreover, the robustness of the proposed PAN-MS-HS fusion under different large spatial resolution ratios (such as 8, 16, 32, 64) was analyzed. The experimental results demonstrated the competitive performance of the proposed method.
Xiangchao Meng, Xiangjun Meng, Qiang Liu 0035, Shutao Li 0001
IEEE Trans. Geosci. Remote. Sens.1
2023 CDFSL: Image Registration for Spaceborne Hyperspectral and Multispectral Data Having Large Spatial-Resolution Difference
abstract
Image registration aims to eliminate the geometric deviation between multi-source data with the same range, and to promote the collaborative application of data. In recent years, spaceborne hyperspectral (HS) and multispectral (MS) data have been widely used in Earth observation. However, the difference in the number of bands, spatial resolution, and spectral resolution puts forward higher requirements on the registration algorithm. The key to HS and MS image registration is to extract more common key points, weaken and eliminate the difference of radiation and spatial texture information to build superior descriptors, and achieve high-precision matching of key points. This paper introduces a new robust HS and MS registration method based on common deep feature subspaces. We first construct the common deep feature subspaces extraction network to extract consistent edge features and common subspace images of the image pair. Then, Harris algorithm is used to extract key points from consistent edge features between images, which reduces the impact of spatial resolution differences between images. Besides, the SIFT descriptor and subspace images are used to describe key points, which reduces the impact of radiation differences between images. Finally, Euclidean distance is used for the initial matching of key points, and the affine matrix is calculated after the outliers are eliminated, and image registration is performed. We perform experiments on spaceborne HS and MS datasets of different spatial resolutions and comparisons with state-of-the-art methods. Experimental results show that our method can obtain satisfactory registration results and is robust.
Kai Ren 0003, Weiwei Sun 0005, Xiangchao Meng, Gang Yang 0006, Jiangtao Peng, Jingfeng Huang, Jiancheng Li
IEEE Trans. Geosci. Remote. Sens.3
2023 Unsupervised 3-D Tensor Subspace Decomposition Network for Spatial-Temporal-Spectral Fusion of Hyperspectral and Multispectral Images
abstract
Due to sensor design limitations and the influence of weather factors, it is currently challenging to obtain remote sensing images with high temporal, spatial, and spectral resolution. Spatial-temporal-spectral fusion aims to integrate the temporal, spatial, and spectral information from multiple sources of remote sensing images to reconstruct a remote sensing image with high temporal, spatial, and spectral resolution. Existing methods typically require at least three types of data to achieve spatial-temporal-spectral fusion. However, acquiring remote sensing data observed at the same time poses significant difficulties. The major challenge lies in effectively utilizing hyperspectral images with low spatial and temporal resolution and multispectral images with high temporal and spatial resolution to reconstruct remote sensing images with high temporal, spatial, and spectral resolution. To address the aforementioned issues, we propose a novel unsupervised 3D tensor subspace decomposition network. Our method incorporates the theory of 3D tensor subspace decomposition, utilizing a 3D hyperspectral/multispectral tensor subspace extraction network to predict the hyperspectral tensor subspace features with low spatial resolution missing at other times (To better understand, the missing moment is defined as time 2). Subsequently, the 3D hyperspectral tensor subspace reconstruction network is employed along with the time 2 hyperspectral tensor subspace features with low spatial resolution and the time 2 multispectral image to reconstruct the time 2 hyperspectral image with high spatial resolution. In the experiment, we utilize three simulated datasets and two real datasets to evaluate the fusion performance of our proposed method. The results demonstrate that our method achieves high-quality fusion results and exhibits comparable performance, and has robustness and practicality.
Weiwei Sun 0005, Kai Ren 0003, Xiangchao Meng, Gang Yang 0006, Jiangtao Peng, Jiancheng Li
IEEE Trans. Geosci. Remote. Sens.3
2023 Coupled Temporal Variation Information Estimation and Resolution Enhancement for Remote Sensing Spatial-Temporal-Spectral Fusion
abstract
Spatial-temporal-spectral fusion (STSF) of remote sensing imagery can produce data with the highest spatial and spectral resolution, only as well as fine temporal resolution, by integrating images with complementary information in both the temporal and spectral domains. Accuracy of temporal variation is an important guarantee for achieving fidelity fusion in STSF. However, current STSF methods estimate the temporal variation only by utilizing the temporal variation between observed multispectral image (MSI) and the relationship between MSI and hyperspectral image (HSI), which is difficult to obtain accurate temporal variation. To address this problem, this paper proposes a coupled temporal variation information estimation and resolution enhancement for remote sensing image spatial-temporal-spectral fusion (CTVRE-STSF). The temporal variation information estimation model estimates the temporal variation of the target image, while the resolution enhancement model provides additional constraints for estimating the temporal variation. For the temporal variation information reconstruction model, we build a temporal variation information estimation based on a generalized linear mixed model and use the temporal variation between MSIs. In addition, a resolution enhancement model is constructed to estimate the temporal variation of the target image by incorporating relevant prior knowledge. The introduction of the resolution enhancement model in the prior provides additional constraints on the estimation of the temporal variation high-dimensional information, thus facilitating the resolution improvement. Experimental results on two real datasets demonstrate the effectiveness and superiority of our proposed method over current state-of-the-art methods, especially in terms of spectral fidelity.
Weiwei Sun 0005, Xiangchao Meng, Gang Yang 0006, Kai Ren 0003, Jiangtao Peng
IEEE Trans. Geosci. Remote. Sens.3
2023 A Progressive Feature Enhancement Deep Network for Large-Scale Remote Sensing Image Superresolution
abstract
The pursuit of super-resolution (SR) with large upscaling factors such as 8×, for enhancing the spatial resolution of low-resolution (LR) remote sensing images is a persistent and challenging problem. To address this issue, we propose the Progressive Feature Enhancement SR (PFESR) network with an 8× upscaling factor. Given the limited high-frequency information provided by a single LR image, we propose an improved style transfer technology to generate auxiliary details that aid in the recovery of high-resolution (HR) images. Additionally, multi-scale texture features are extracted through the Visual Geometry Group (VGG) feature extraction (VFE) block. To efficiently fuse various features, we combine hard and soft attention mechanisms. Finally, we use a hierarchical fusion block to address the progressive fusion problem of multiple scale features. Experiments on three datasets demonstrate that our method achieves state-of-the-art performance and exhibits good robustness in 8× and higher scale SR tasks.
Weiwei Liu 0009, Weiwei Sun 0005, Xiangchao Meng, Gang Yang 0006, Kai Ren 0003
IEEE Trans. Geosci. Remote. Sens.4
2023 Perceptual Quality Assessment of Cartoon Images
abstract
In the animation industry, automatically predicting the quality of cartoon images based on the inputs of general distortions and color change is an urgent task, while the existing no-reference (NR) methods usually measure the perceptual quality of the natural images. In this paper, based on the observation that structure and color are the main factors affecting cartoon images quality, we proposed a new NR quality prediction metric for cartoon images, which fully takes gradient and color information into account. The experimental results on our newly constructed NBU-CIQAD dataset with color change and other existing cartoon image dataset demonstrate that the proposed method significantly outperforms existing no-references methods for the task of cartoon image quality assessment. The database and code will be released athttps://github.com/1010075746/NBU-CIQAD.
Hangwei Chen, Xiongli Chai, Feng Shao 0001, Xuejin Wang, Qiuping Jiang, Xiangchao Meng, Yo-Sung Ho
IEEE Trans. Multim.6
2022 A Temporal-Spectral Generative Adversarial Fusion Network for Improving Satellite Hyperspectral Temporal Resolution
abstract
The improvement of temporal resolution of hyperspectral (HS) data is a fundamental and challenging problem. In this paper, we propose a Temporal-Spectral fusion method based on Generative Adversarial Network (TSF-GAN). First, the generator is used to train the nonlinear relationship between multispectral (MS) and HS data pairs at time T1 and T3, and we map the relationship to the MS data at T2 to obtain the HS data. Second, the discriminator is used to identify whether the differential image of HS data at different times is consistent with that of MS data, and whether the HS data at time T2 after spectral down-sampling is consistent with that of MS data at time T2. Preliminary experimental results demonstrate that the proposed TSF-GAN achieves comparative fidelity and has strong practicability.
Kai Ren 0003, Weiwei Sun 0005, Xiangchao Meng, Gang Yang 0006, Jiangtao Peng
IGARSS4
2022 Integrated Fusion for Panchromatic, Multispectral, Hyperspectral Remote Sensing Images With Different Swath Widths
abstract
Zi Yuan (ZY)-1 02D satellite simultaneously provides the low spatial resolution (LR) and narrow swath-width hyperspectral (HS) image, the moderate spatial resolution (MR) multispectral (MS) image with a wider swath width, and the high spatial resolution (HR) panchromatic (PAN) image with the same wide swath width to the MR MS. How to comprehensively integrate their complementary advantages to obtain the wide swath-width and high-fidelity HR HS image is interesting but challenging. In this paper, we propose an integrated fusion method for the HR PAN, MR MS, and LR HS images with different swath widths, to generate the optimal wide swath-width HR HS image. The proposed method is based on the encoder-decoder learning framework. In the proposed fusion framework, a novel multi-branch encoder structure with an enhanced HS-encoder module and the multilevel spatial-spectral aggregation block is designed, by considering the difference in the spatial and spectral resolution among the multi-sensor images. The experiments on synthetic and real datasets from both qualitative and quantitative aspects demonstrated the competitive performance of the proposed method.
Xiangjun Meng, Xiangchao Meng, Qiang Liu 0035, Jinfang Shu, Feng Shao 0001, Gang Yang 0006, Weiwei Sun 0005
IEEE Geosci. Remote. Sens. Lett.2
2022 SARF: A Simple, Adjustable, and Robust Fusion Method
abstract
Pansharpening aims to sharpen a low spatial resolution (LR) multispectral (MS) image using a high spatial resolution (HR) panchromatic (PAN) image to obtain the HR MS image. Though large numbers of pansharpening methods have been proposed, and many advanced methods have shown high quantitative results, few of them are widely used in real applications. This may be attributed to their instability for different images with different ground surface features, or the complexity to be implemented and the time-consuming process for some state-of-the-art methods. In this letter, we proposed a simple, adjustable, and robust fusion (SARF) method. In the proposed method, a spatial-spectral coenhanced strategy was proposed, and several details of the proposed fusion model were specifically designed for the “simple, adjustable, robust” features. It was tested and verified by four-band and eight-band MS images based on reduced resolution (RR) and full resolution (FR) experiments. The experimental results demonstrated the promising spatial visuality of the proposed method, and the spectral fidelity was more robust than most of component substitution (CS)-based and multiresolution analysis (MRA)-based methods.
Xiangchao Meng, Gang Yang 0006, Feng Shao 0001, Weiwei Sun 0005, Huanfeng Shen, Shutao Li 0001
IEEE Geosci. Remote. Sens. Lett.1
2022 Multiscale Low-Rank Spatial Features for Hyperspectral Image Classification
abstract
This letter presents a multiscale low-rank decomposition (MSLRD) method to extract multiscale spatial structures from hyperspectral images. The MSLRD assumes that ground objects have divergent characteristics in changing spatial scales. It decomposes each band image into a series of block-wise matrices, where these low-rank blocks take detailed spatial structures at multiple scales. It formulates the low-rank matrix decomposition problem into minimizing the ranks of all block matrices and adopts the alternative direction of the multiplier method to optimize it. Experiments on Indian Pines and Pavia University data sets show that the MSLRD can greatly improve the classification performance of regular classification on spectral features (i.e., all bands) and perform better than five state-of-the-art spatial feature extraction methods.
Weiwei Sun 0005, Wenjing Shao, Jiangtao Peng, Gang Yang 0006, Xiangchao Meng, Qian Du 0001
IEEE Geosci. Remote. Sens. Lett.5
2022 Monocular and Binocular Interactions Oriented Deformable Convolutional Networks for Blind Quality Assessment of Stereoscopic Omnidirectional Images
abstract
Stereoscopic omnidirectional content, as a novel visual media, has drawn wide attention in recent years due to its ability in providing strong immersive experience. Since Stereoscopic Omnidirectional Images (SOIs) involve the properties from panoramic and stereoscopic visual perception, it is very challenging to establish an efficient and effective visual quality evaluation model for SOIs. To better measure the user’s experience in virtual reality, we put forward a novel deep learning framework to assess the quality of SOIs in this paper. Firstly, the deformable convolutions instead of standard convolutions are adopted to ensure the invariant receptive fields of convolutional kernels on Equi-Rectangular Projection (ERP). Secondly, according to the stereoscopic property, we use binocular-difference information and a coarse-to-fine mechanism to construct the binocular feature extraction network. Thirdly, a three-channel network involving left-view, right-view and binocular-difference channels is presented to simulate the process of monocular and binocular interactions, in which independent quality labels are provided for each channel to reflect the individual effect of monocular and binocular visions on the whole visual quality. Finally, experimental results on two available benchmark databases demonstrate the superiority of the proposed metric over the state-of-the-art blind quality assessment models in predicting the quality of SOIs. Moreover, our model is efficient in computational cost as the feature extraction is directly applied on ERP images.
Xiongli Chai, Feng Shao 0001, Qiuping Jiang, Xiangchao Meng, Yo-Sung Ho
IEEE Trans. Circuits Syst. Video Technol.4
2022 CGMDRNet: Cross-Guided Modality Difference Reduction Network for RGB-T Salient Object Detection
abstract
How to explore the interaction between the RGB and thermal modalities is the key success of the RGB-T saliency object detection (SOD). Most of the existing methods integrate multi-modality information by designing various fusion strategies. However, the modality gap between the RGB and thermal features will lead to unsatisfactory performances by simple feature concatenation. To solve this problem, we innovatively propose a cross-guided modality difference reduction network (CGMDRNet) to achieve intrinsic consistency feature fusion via reducing the modality differences. Specifically, we design a modality difference reduction (MDR) module, which is embedded in each layer of the backbone network. The module uses a cross-guided strategy to reduce the modality difference between the RGB and thermal features. Then, a cross-attention fusion (CAF) module is designed to fuse cross-modality features with small modality differences. In addition, we use a transformer-based feature enhancement (TFE) module to enhance the high-level feature representation that contributes more to performance. Finally, the high-level features guide the fusion of low-level features to obtain a saliency map with clear boundaries. Extensive experiments on three public RGB-T datasets show that the proposed CGMDRNet achieves competitive performance compared with state-of-the-art (SOTA) RGB-T SOD models.
Feng Shao 0001, Xiongli Chai, Hangwei Chen, Qiuping Jiang, Xiangchao Meng, Yo-Sung Ho
IEEE Trans. Circuits Syst. Video Technol.6
2022 VSOIQE: A Novel Viewport-Based Stitched 360° Omnidirectional Image Quality Evaluator
abstract
With the rapid development of virtual reality (VR), 360° omnidirectional images and videos have drawn wide attention. However, the quality assessment of 360° omnidirectional images is a challenging task, especially when the panoramic image contains multiple stitching distortions. We propose a viewport-based stitched 360° omnidirectional image quality evaluator (VSOIQE), by first extracting the features of salient and stitching viewports, and then inferring the overall perceptual quality via multiple linear regression (MLR). Comprehensive image attributes including edge, color, shape and information entropy are considered in the framework. Experimental results on two benchmark databases demonstrate the superiority of the proposed metric over both the state-of-the-art quality models designed for 2D images and the quality models developed for 360° omnidirectional images.
Chongzhen Tian, Xiongli Chai, Feng Shao 0001, Qiuping Jiang, Xiangchao Meng, Long Xu 0001, Yo-Sung Ho
IEEE Trans. Circuits Syst. Video Technol.6
2022 A Locally Optimized Model for Hyperspectral and Multispectral Images Fusion
abstract
The maintenance of spectral variability between subclass objects and the relationship between hyperspectral (HS) bands have been a fundamental but challenging problem for fusing low spatial resolution (LR) HS and high spatial resolution (HR) multispectral (MS) images. This article presents a locally optimized image segmentation fusion (LOISF) framework for HS super-resolution reconstruction. First, LR HS and HR MS are clustered and segmented, and the label attributes of the segmented objects are identified by the prior information. Then, a novel joint fusion model for different typical ground objects is constructed based on spectral unmixing. The fusion problem is formulated mathematically as a convex optimization of a Frobenius norm, which includes spatial, spectral, and index constraints, with an alternating-directions’ optimization featuring linearization providing the solution. Experimental results demonstrate that the proposed LOISF preserves both spatial details and texture, achieving high spectral fidelity, and yielding significantly improved image quality compared to other state-of-the-art fusion methods.
Kai Ren 0003, Weiwei Sun 0005, Xiangchao Meng, Gang Yang 0006, Jiangtao Peng, Jingfeng Huang
IEEE Trans. Geosci. Remote. Sens.3
2022 A Blind Full Resolution Assessment Method for Pansharpened Images Based on Multistream Collaborative Learning
abstract
Pansharpening aims to fuse a high spatial resolution (HR) panchromatic (PAN) image and a low spatial resolution (LR) multispectral (MS) image to obtain an HR-MS image. However, due to the lack of the real HR-MS reference image, determining pansharpened image quality at full resolution has always been a contentious issue in the community. We propose a blind full resolution assessment method for pansharpened images based on multi-stream collaborative learning. The proposed method designs a Siamese framework to collaboratively learn the spatial, spectral, and overall quality of the fused image. The parameters of the feature extraction layer in the spatial and spectral evaluation models are frozen for the overall evaluation model, improving accuracy and convergence speed. The proposed method was comprehensively tested and verified based on a large-scale data set consisting of 13620 fused images obtained by six pansharpening methods with four different thematic data sets. Furthermore, a large-scale subjective evaluation data set in which each of the 13620 fused images was assessed by 28 participants, was utilized to comprehensively valid the proposed method. The experimental results demonstrated the superior performance of the proposed method to other state-of-the-art quality assessments.
Kedi Bao, Xiangchao Meng, Xiongli Chai, Feng Shao 0001
IEEE Trans. Geosci. Remote. Sens.2
2022 A Dual Global-Local Attention Network for Hyperspectral Band Selection
abstract
This article proposes a dual global–local attention network (DGLAnet), which is an end-to-end unsupervised band selection (UBS) method that fully utilizes spatial and spectral information in both global and local aspects. The DGLAnet assumes that BS can be realized using the hyperspectral image (HSI) reconstruction process. First, the DGLAnet implements a dual attention module to obtain spatial–spectral and global–local features to reweight the HSI data. It adopts bi-directional relations to grasp spatial and spectral features from a global perspective. Meanwhile, the DGLAnet extracts local features through max-pooling and mean-pooling and then merges them via the convolution operation. Global–local features are utilized to learn attention to recalibrate the original data, and the reconstruction module is adopted to restore the original image from the reweighted HSI data. Finally, a proper band subset is selected by the constructed band evaluation index. Experiments on three hyperspectral data show that the DGLAnet outperforms other state-of-the-art methods and uses all bands with a lower computational cost.
Weiwei Sun 0005, Gang Yang 0006, Xiangchao Meng, Kai Ren 0003, Jiangtao Peng, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.4
2022 PSTAF-GAN: Progressive Spatio-Temporal Attention Fusion Method Based on Generative Adversarial Network
abstract
Spatio-temporal fusion aims to integrate multisource remote sensing images with complementary high spatial and temporal resolutions, so as to obtain time-series high spatial resolution fused images. Currently, deep learning (DL)-based spatio-temporal fusion methods have received broad attention. However, on one hand, most of the existing DL-based methods train the model in a band-by-band manner, ignoring the correlations among bands. On the other hand, the general coarse spatio-temporal changes in low spatial resolution images (e.g., MODIS) calculated at the pixel domain cannot completely cover the fine spatio-temporal changes in high spatial resolution images (e.g., Landsat), due to complex surface features and the general large spatial resolution ratio between fine and coarse images. Besides, the existing DL-based spatio-temporal fusion methods are insufficient in exploring multiscale information by only stacking convolutional kernels with different sizes. To alleviate the above challenges, we propose a progressive spatio-temporal attention fusion model in a multiband training manner based on generative adversarial network (PSTAF-GAN). Specifically, we design a flexible multiscale feature extraction architecture to extract multiscale feature hierarchies. Then, spatio-temporal changes are calculated on the feature domain in different feature hierarchies. Besides, a spatio-temporal attention fusion architecture is proposed to fuse the spatio-temporal changes and ground details in a coarse-to-fine manner, which can explore multiscale information more sufficient and gradually recover the target image. The results of quantitative and qualitative experiments on two publicly available benchmark datasets show that the proposed PSTAF-GAN can achieve the best performance compared with the state-of-the-art methods.
Qiang Liu 0035, Xiangchao Meng, Feng Shao 0001, Shutao Li 0001
IEEE Trans. Geosci. Remote. Sens.2
2022 A Blind Full-Resolution Quality Evaluation Method for Pansharpening
abstract
Pansharpening methods have been developed for nearly 40 years; however, how to quantitatively evaluate the quality of pansharpened images at full resolution (FR) is probably the most debated topic in this field due to the inherent unavailable of the real HR MS reference image. In this article, a novel blind FR quality evaluation method for pansharpening is proposed. In the proposed method, spatial and spectral features that are sensitive to spatial and spectral distortions of fused images are comprehensively considered and jointly learned based on online multivariate Gaussian (MVG) to construct the evaluation model. It directly outputs the quality of fused images, rather than the stepwise evaluation of spectral score, spatial score, and final overall quality score by the weighted combination of them, which may introduce contradictory results. First, a pristine benchmark evaluation model is established on the spatial features from the original high-spatial-resolution (HR) panchromatic (PAN) image and the spectral invariant assumption between ideal fused and original multispectral (MS) images. Second, a testing evaluation model for the fused image is founded. Finally, the quality of the fused image is measured based on the distance between the testing and benchmark models. The experimental results demonstrated the superior performance of the proposed method. Furthermore, the proposed method can be generalized to other interesting tasks, such as the nonreference evaluation for pansharpening with missing information and the nonreference evaluation for hyperspectral image fusion. The source code is available onhttps://github.com/yyxhpkq/MQNR.
Xiangchao Meng, Kedi Bao, Jinfang Shu, Bingzhong Zhou, Feng Shao 0001, Weiwei Sun 0005, Shutao Li 0001
IEEE Trans. Geosci. Remote. Sens.1
2022 Spatio-Temporal-Spectral Collaborative Learning for Spatio-Temporal Fusion with Land Cover Changes
abstract
Spatio-temporal fusion by combining the complementary spatial and temporal advantages of multi-source remote sensing images to obtain time-series high spatial resolution images is highly desirable in monitoring surface dynamics. Currently, deep learning (DL)-based fusion methods have received extensive attention. However, existing DL-based spatio-temporal fusion methods are generally limited in fusing the images with land cover changes. In this paper, we propose a spatio-temporal-spectral collaborative learning framework for spatio-temporal fusion to alleviate this problem. Specifically, the proposed method integrates the convolutional neural network and recurrent neural network into a unified framework, consisting of three sub-networks: multi-scale siamese convolutional neural network, multi-layer convolutional recurrent neural network, and adaptive weighting fusion network. The multi-scale siamese convolutional neural network has a flexible weight-sharing network to extract multi-scale spatial-spectral features from multi-source remote sensing images. The multi-layer convolutional recurrent neural network is constructed on the convolutional long-short term memory units to comprehensively learn the land cover changes by spatial, spectral, and temporal joint features. The adaptive weighting fusion network with a spatio-temporal-spectral change loss is proposed to further improve the interpretability and robustness. The experiments were performed on the publicly available benchmark datasets featured by phenology and land cover type changes, respectively. The experimental results demonstrated the competitive performance of the proposed method than other state-of-the-art fusion methods.
Xiangchao Meng, Qiang Liu 0035, Feng Shao 0001, Shutao Li 0001
IEEE Trans. Geosci. Remote. Sens.1
2022 Vision Transformer for Pansharpening
abstract
Pansharpening is a fundamental and hot-spot research topic in remote sensing image fusion. In recent years, self-attention-based transformer has attracted considerable attention in natural language processing (NLP) and introduced to attend to computer vision (CV) tasks. Inspired by great success of the vision transformer (ViT) in image classification, we propose an improved and advanced purely transformer-based model for pansharpening. In the proposed method, stacked multispectral (MS) and panchromatic (PAN) images are cropped into patches (i.e., tokens), and after a three-layer self-attention-based encoder, these tokens contain rich information. After upsampled and stitched, a high spatial resolution (HR) MS image is finally obtained. Instead of convolutional neural networks (CNNs) pursuing a short-distance dependency, our proposed method aims to build up a long-distance dependency, to make full use of more useful features. The experiments were conducted on an opening benchmark dataset, including IKONOS with four-band MS/PAN images and WorldView-2 MS images featured by eight bands. In addition, the experiments were performed on reduced and full-resolution datasets from both qualitative and quantitative evaluation aspects. The experimental results indicate the competitive performance of the proposed model than other pansharpening methods, including the state-of-the-art pansharpening algorithms based on CNN.
Xiangchao Meng, Feng Shao 0001, Shutao Li 0001
IEEE Trans. Geosci. Remote. Sens.1
2022 A Band Divide-and-Conquer Multispectral and Hyperspectral Image Fusion Method
abstract
The nonoverlapped spectrum range between low spatial resolution (LR) hyperspectral (HS) and high spatial resolution (HR) multispectral (MS) images has been a fundamental but challenging problem for MS/HS fusion. The spectrum of HS data is generally 400–2500 nm, and the spectrum of MS data is generally 400–900 nm; how to obtain the high-fidelity HR HS fused image within the whole spectrum of 400–2500 nm? In this article, we proposed a band divide-and-conquer framework (BDCF) to solve the problem, by comprehensively considering spectral fidelity, spatial enhancement, and computational efficiency. First, the spectral bands of HS were divided into overlapped and nonoverlapped bands according to the spectral response between HS and MS. Then, a novel improved component substitution (CS)-based method by combing neural network was proposed to fuse the overlapped bands of LR HS. Then, a mapping-based method with the neural network was presented to construct the complicated nonlinear relationship between overlapped and nonoverlapped bands of the original LR HS data. The trained network was mapped to the fused overlapped HR HS bands to estimate the nonoverlapped HR HS bands. Experimental results on two simulated data sets and two realistic data sets of Gaofen (GF)-5 LR HS, GF-1 MS, and Sentinel-2A MS show that the proposed BDCF has superior performance in both high spectral fidelity and sharp spatial details, and it obtained competitive fusion behaviors compared with other state-of-the-art methods. Moreover, BDCF has relatively higher computational efficiency than optimal solution-based methods and deep learning-based fusion methods.
Weiwei Sun 0005, Kai Ren 0003, Xiangchao Meng, Chenchao Xiao, Gang Yang 0006, Jiangtao Peng
IEEE Trans. Geosci. Remote. Sens.3
2022 MLR-DBPFN: A Multi-Scale Low Rank Deep Back Projection Fusion Network for Anti-Noise Hyperspectral and Multispectral Image Fusion
abstract
Fusing low spatial resolution (LR) hyperspectral (HS) data and high spatial resolution (HR) multispectral (MS) data aims to obtain HR HS data. However, due to bad weather and the aging of sensor equipment, HS images usually contain a lot of noise, e.g., Gaussian noise, strip noise, and mixed noise, which would make the fused image have low quality. To solve this problem, we propose the multiscale low-rank deep back projection fusion network (MLR-DBPFN). First, HS and MS are superimposed, and multiscale spectral features of the stacked image are extracted through multiscale low-rank decomposition and convolution operation, which effectively removes noisy spectral features. Second, the upsampling and downsampling network mechanisms are used to extract the multiscale spatial features from each layer of spectral features. Finally, the multiscale spectral features and multiscale spatial features are combined for network training, and the weight of the noisy spectrum features is reduced through the network feedback mechanism, which suppresses the noisy spectrum and improves the noisy HS fusion performance. Experimental results on datasets of different noise demonstrate that MLR-DBPFN has superior spatial and spectral fidelity, comparative fusion quality, and robust antinoise performance compared with state-of-the-art methods.
Weiwei Sun 0005, Kai Ren 0003, Xiangchao Meng, Gang Yang 0006, Chenchao Xiao, Jiangtao Peng, Jingfeng Huang
IEEE Trans. Geosci. Remote. Sens.3
2022 A Multiscale Spectral Features Graph Fusion Method for Hyperspectral Band Selection
abstract
This article proposes a multiscale spectral features graph fusion (MSFGF) method for selecting proper hyperspectral bands. The MSFGF regards that the selected bands should reflect diagnostic spectral information of ground objects at different scales, and it explores band selection from the aspect of multiple spatial scales. First, it adopts the multiscale low-rank decomposition (MSLRD) model to find multiscale spectral features of different ground objects. The model considers divergent spatial structures or spatial correlations of ground objects at different scales, and factorizes the hyperspectral data cube into a series of low-rank block-wise data cubes, where the blocks take spatial structures of different ground objects at increasing scales. Second, the MSFGF presents the multiscale sparse spectral clustering (MSSC) model to fuse the separate connected graphs of multiscale spectral features into a consensus graph. The consensus graph combines the complementary information of multiscale spectral features and helps to reveal the intrinsic clustering structure of all spectral bands. Finally, the MSFGF utilizes spectral clustering to find clusters from the consensus graph and selects representative bands. Experimental results on three widely used hyperspectral data prove the superiority of MSFGF in selecting bands, where it outperforms other seven state-of-the-art methods in classification with an acceptable computational cost.
Weiwei Sun 0005, Gang Yang 0006, Jiangtao Peng, Xiangchao Meng, Wei Li 0032, Heng-Chao Li 0001, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.4
2022 Convolution-Embedded Vision Transformer With Elastic Positional Encoding for Pansharpening
abstract
Transformer, especially vision transformer (ViT), is attracting increasing attention in various computer vision (CV) tasks. However, two urgent problems exist for the ViT: 1) Owing to its attending to an image in the patch level, the vision transformer seems to have a better performance in fetching global representations but limited in extracting local features, which is an inherent advantage for the convolutional neural network (CNN); 2) the learnable positional encoding plays a positive role, but limits the cross-resolution ability of the network. Specifically, the pre-trained model could only generate images with the same size during training. To conquer the two problems, we propose a novel convolution-embedded vision transformer with elastic positional encoding in this paper. On one hand, we propose a joint CNN and self-attention network to collaboratively extract local and global features. On the other hand, we propose to integrate the elastic CNN-based positional encoder into the framework to solve the rigid limitation of the ViT in cross resolution issues and improve the performance. Extensive experiments were conducted on IKONOS and WorldView-2 with 4-band and 8-band multispectral images, respectively. The visual and numerical results show the competitive performance of the proposed method.
Xiangjun Meng, Xiangchao Meng, Feng Shao 0001
IEEE Trans. Geosci. Remote. Sens.3
2022 Generalized Linear Spectral Mixing Model for Spatial-Temporal-Spectral Fusion
abstract
Image fusion effectively solves the trade-off between spatial resolution, temporal resolution, and spectral resolution of remote sensing sensors. However, most of existing methods focus on the fusion of two of the spatial, temporal, and spectral metrics of remote sensing images. The few spatial-temporal-spectral fusion (STSF) methods available are mainly for fusing MODIS and Landsat images, which are not suitable for the characteristics of the spaceborne hyperspectral images with low temporal resolution, such as Hyperion, ZY-1 02D, and PRISMA. For this purpose, we proposed a novel generalized linear spectral mixing model for spatial-temporal-spectral fusion (GLMM-STSF). In the method, the GLMM is introduced into the STSF problem, and the temporal variations of images at different times are transferred to the endmember and abundance matrix variations of images for estimation. To the best of our knowledge, for the first time, the STSF task of remote sensing images is handled from the perspective of spectral unmixing. Compared with existing STSF fusion methods, our method targets the task of fusing spaceborne HSI with low temporal and spatial resolutions with multispectral image featured by high temporal and spatial resolutions. Taking the STSF of ZY-1 02D hyperspectral and Sentinel-2 multispectral real datasets as an example, comparisons with related state-of-the-art methods demonstrate that our proposed method achieves superior fusion performance.
Weiwei Sun 0005, Xiangchao Meng, Gang Yang 0006, Kai Ren 0003, Jiangtao Peng
IEEE Trans. Geosci. Remote. Sens.3
2022 List-Wise Rank Learning for Stereoscopic Image Retargeting Quality Assessment
abstract
Stereoscopic imageretargeting (SIR) techniques attempt to display stereoscopic images on stereoscopic devices of various resolutions and aspect ratios to provide the users with better viewing experience. However, new quality perceptual problems emerge in the retargeted stereoscopic images generated by current SIR operators are quite different from those in the retargeted 2D images. In this paper, we dedicate to exploring the perceptual quality-related factors (e.g., shape preservation, object preservation and visual comfort.) of retargeted stereoscopic images, and propose a novel quality evaluation metric for SIR to achieve a more consistent evaluation with 3D perception and image degradation mechanism in the SIR process. Moreover, image quality features and 3D perceptual features are integrated into one representation for an overall perceptual quality prediction using a list-wise ranking approach, which gives priority to the ranking among the SIR results generated from the same stereoscopic source. Experimental results demonstrate that the proposed method outperforms most quality models developed for retargeted 2D/stereoscopic images.
Xuejin Wang, Feng Shao 0001, Qiuping Jiang, Xiongli Chai, Xiangchao Meng, Yo-Sung Ho
IEEE Trans. Multim.5
2022 Combining Retargeting Quality and Depth Perception Measures for Quality Evaluation of Retargeted Stereopairs
abstract
Stereoscopic Image Retargeting (SIR) aims to adapt stereoscopic images and videos to 3D display devices with various aspect ratios by emphasizing the important content while retaining surrounding context with minimal visual distortion. To address the issue of SIR evaluation, this paper presents a new objective quality assessment method for retargeted stereopairs by combining image quality and depth perception measures. Specifically, the image quality measure is conducted between the source and retargeted intermediate views generated by the view synthesis method to characterize the geometric distortion and content loss of the retargeted stereopair, while several depth-aware features are extracted to measure the visual comfort/discomfort and depth sensation when human views a 3D scene. Then, the extracted features are integrated into an overall perceptual quality prediction. Experiment results on NBU SIRQA and SIRD databases verify the superiority of our method.
Xuejin Wang, Feng Shao 0001, Qiuping Jiang, Zhenqi Fu, Xiangchao Meng, Ke Gu 0001, Yo-Sung Ho
IEEE Trans. Multim.5
2021 Subjective and Objective Quality Assessment for Stereoscopic Image Retargeting
abstract
Binocular stereoscopic image retargeting (SIR) aims to adjust 3D images into target aspect ratios. In recent years, various SIR methods have been proposed, but there are few researches on visual quality assessment. As a consequence, we construct a benchmark stereoscopic image retargeting quality assessment database (NBU-SIRQA), which contains 720 stereoscopic retargeted images generated by eight representative SIR operators. Subjective test is conducted to obtain the mean opinion score (MOS) for each stereoscopic retargeted image. Additionally, we propose an objective SIRQA metric based on grid deformation and information loss (GDIL). The main idea of GDIL is to decompose the SIR operator into two transformations: monocular image retargeting transformation and viewpoint transformation. In each transformation, grid deformation and information loss are extracted simultaneously to represent image quality and 3D perception quality. Experimental results validated on our established NBU-SIRQA database show the superiority of our metric in measuring the quality of stereoscopic retargeted images over the existing approaches.
Zhenqi Fu, Feng Shao 0001, Qiuping Jiang, Xiangchao Meng, Yo-Sung Ho
IEEE Trans. Multim.4
2021 Measuring Coarse-to-Fine Texture and Geometric Distortions for Quality Assessment of DIBR-Synthesized Images
abstract
A synthesized view can be generated via Depth-Image-Based Rendering (DIBR) technique using one (or more) color images and the associated depth maps. However, several artifacts may occur in the synthesized views due to the imperfect color images, depth maps or texture inpainting techniques, which cannot be effectively estimated by the conventional quality metrics designed for natural images. In this paper, a new quality metric is proposed to evaluate DIBR-synthesized images by measuring texture and geometric distortions. The artifacts are first analyzed on different phases of the synthesis process, and the associated features are extracted to estimate the degree of texture and geometric distortions from both coarse and fine scales. Finally, individual quality scores are aggregated into an overall quality via regression. Experimental results on three publicly available DIBR datasets demonstrate the superiority of the proposed method over the state-of-the-art quality models.
Xuejin Wang, Feng Shao 0001, Qiuping Jiang, Xiangchao Meng, Yo-Sung Ho
IEEE Trans. Multim.4
2020 Blind quality assessment for multiply distorted stereoscopic images towards IoT-based 3D capture systems
Xuejin Wang, Meiling Qi, Feng Shao 0001, Qiuping Jiang, Xiangchao Meng
J. Vis. Commun. Image Represent.5
2020 A large-scale remote sensing database for subjective and objective quality assessment of pansharpened images
Yiming Xiong, Feng Shao 0001, Xiangchao Meng, Qiuping Jiang, Weiwei Sun 0005, Randi Fu, Yo-Sung Ho
J. Vis. Commun. Image Represent.3
2019 Fine Classification Comparsion of GF-1 GF-5 and Landsat-8 Remote Sensing Data Based on Optimized Sample Selection Method
abstract
This paper aims to compare the performance of GaoFen-1 (GF-1), GaoFen-5 (GF-5), Landsat-8 data in fine classification. An optimized sample selection method (OSSM) is developed to ensure the high quality of samples. This method adopts different band combination strategies to realize optimal selection of training samples under the aid of normalized vegetation index (NDVI), normalized water index (NDWI) and the components of Kauth-Thomas (KT) Transformation. After that, support vector machine (SVM) is implemented on these three types of data. Experimental results on China Dunhuang calibration field, Gansu Ying-mao-tuo exploration area and Gan River lower reaches datasets show that GF-5 data performs best in both qualitative and quantitative evaluation of fine classification thanks to its hyperspectral properties.
Gang Yang 0006, Leilei Jiao, Weiwei Sun 0005, Huimin Lu 0009, Xiangchao Meng, Yinnian Liu
IGARSS5
2019 Pansharpening for Cloud-Contaminated Very High-Resolution Remote Sensing Images
abstract
The optical remote sensing images not only have to make a fundamental tradeoff between the spatial and spectral resolutions, but also are inevitable to be polluted by the clouds; however, the existing pansharpening methods mainly focus on the resolution enhancement of the optical remote sensing images without cloud contamination. How to fuse the cloud-contaminated images to achieve the joint resolution enhancement and cloud removal is a promising and challenging work. In this paper, a pansharpening method for the challenging cloud-contaminated very high-resolution remote sensing images is proposed. Furthermore, the cloud-contaminated conditions for the practical observations with all the thick clouds, the thin clouds, the haze, and the cloud shadows are comprehensively considered. In the proposed methods, a two-step fusion framework based on multisource and multitemporal observations is presented: 1) the thin clouds, the haze, and the light cloud shadows are proposed to be first jointly removed and 2) a variational-based integrated fusion model is then proposed to achieve the joint resolution enhancement and missing information reconstruction for the thick clouds and dark cloud shadows. Through the proposed fusion method, a promising cloud-free fused image with both high spatial and high spectral resolutions can be obtained. To comprehensively test and verify the proposed method, the experiments were implemented based on both the cloud-free and cloud-contaminated images, and a number of different remote sensing satellites including the IKONOS, the QuickBird, the Jilin (JL)-1, and the Deimos-2 images were utilized. The experimental results confirm the effectiveness of the proposed method.
Xiangchao Meng, Huanfeng Shen, Qiangqiang Yuan, Huifang Li 0001, Liangpei Zhang 0001, Weiwei Sun 0005
IEEE Trans. Geosci. Remote. Sens.1
2017 Multi-scale-and-depth convolutional neural network for remote sensed imagery pan-sharpening
abstract
Pan-sharpening is a fundamental and significant task in the field of remote sensed imagery fusion, which demands fusion of panchromatic and multi-spectral images with the rich information accurately preserved in both spatial and spectral domains. In this paper, to overcome the drawbacks of traditional pan-sharpening methodologies, we employed the advanced concept of deep learning to propose a Multi-Scale-and-Depth Convolutional Neural Network (MSDCNN) as an end-to-end pan-sharpening model. By the results of a large number of quantitative and visual assessments, the qualities of images fused by the proposed network have been confirmed superior to compared state-of-the-art methods.
Yancong Wei, Qiangqiang Yuan, Xiangchao Meng, Huanfeng Shen, Liangpei Zhang 0001, Michael Kwok-Po Ng
IGARSS3
2016 Hyperspectral Image Super-Resolution by Spectral Mixture Analysis and Spatial-Spectral Group Sparsity
abstract
Due to the limitation of hyperspectral sensors and optical imaging systems, there are several irreconcilable conflicts between high spatial resolution and high spectral resolution of hyperspectral images (HSIs). Therefore, HSI super-resolution (SR) is regarded as an important preprocessing task for subsequent applications. In this letter, we use sparse representation to analyze the spectral and spatial feature of HSIs. Considering the sparse characteristic of spectral unmixing and high pattern repeatability of spatial-spectral blocks, we proposed a novel HSI SR framework utilizing spectral mixture analysis and spatial-spectral group sparsity. By simultaneously combining the sparsity and the nonlocal self-similarity of the images in the spatial and spectral domains, the method not only maintains the spectral consistency but also produces plenty of image details. Experiments on three hyperspectral data sets confirm that the proposed method is robust to noise and achieves better results than traditional methods.
Jie Li 0022, Qiangqiang Yuan, Huanfeng Shen, Xiangchao Meng, Liangpei Zhang 0001
IEEE Geosci. Remote. Sens. Lett.4
2016 An Integrated Framework for the Spatio-Temporal-Spectral Fusion of Remote Sensing Images
abstract
Remote sensing satellite sensors feature a tradeoff between the spatial, temporal, and spectral resolutions. In this paper, we propose an integrated framework for the spatio-temporal-spectral fusion of remote sensing images. There are two main advantages of the proposed integrated fusion framework: it can accomplish different kinds of fusion tasks, such as multiview spatial fusion, spatio-spectral fusion, and spatio-temporal fusion, based on a single unified model, and it can achieve the integrated fusion of multisource observations to obtain high spatio-temporal-spectral resolution images, without limitations on the number of remote sensing sensors. The proposed integrated fusion framework was comprehensively tested and verified in a variety of image fusion experiments. In the experiments, a number of different remote sensing satellites were utilized, including IKONOS, the Enhanced Thematic Mapper Plus (ETM+), the Moderate Resolution Imaging Spectroradiometer (MODIS), the Hyperspectral Digital Imagery Collection Experiment (HYDICE), and Système Pour l' Observation de la Terre-5 (SPOT-5). The experimental results confirm the effectiveness of the proposed method.
Huanfeng Shen, Xiangchao Meng, Liangpei Zhang 0001
IEEE Trans. Geosci. Remote. Sens.2
2015 A unified framework for spatio-temporal-spectral fusion of remote sensing images
abstract
In this paper, a unified framework for the spatio-temporal-spectral fusion of remote sensing images is proposed. The relationships between the observed images and the desired image are first established based on general image observation models. Maximum a posteriori (MAP) theory is then employed to formulate the unified fusion framework. The proposed method is able to fuse images from an arbitrary number of optical sensors with different spatial, temporal, and spectral resolutions. The experimental results verify the effectiveness of the proposed method.
Xiangchao Meng, Huanfeng Shen, Liangpei Zhang 0001, Qiangqiang Yuan, Huifang Li 0001
IGARSS1