Kai Zhang 0010

dblp:55/957-10 · DBLP profile ↗
← Back
48ranked-venue papers
13as first author
37since 2021 · last 2026
0000-0002-9218-5916ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 23 · 9 first-author · 17 since 2021Artificial intelligence and machine learning · 12 · 2 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 first-author · 9 since 2021Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Progressive dynamic Taylor unfolding network for multi-modal image fusion
Xuquan Wang, Shengka Shi, Yingjie Kong, Chunan Guan, Jiande Sun 0001, Kai Zhang 0010, Jianfei Cao
Eng. Appl. Artif. Intell.6
2026 Task-driven infrared and visible image fusion via detail and semantic dual injection
Kai Zhang 0010, Ludan Sun, Feng Zhang 0028, Wenbo Wan, Jiande Sun 0001
Neural Networks1
2026 A universal pansharpening network via spatial-spectral contrastive learning
Kai Zhang 0010, Yunlong Liu 0005, Feng Zhang 0028, Wenbo Wan, Lingchen Gu, Jiande Sun 0001
Pattern Recognit.1
2025 Orientation-Aware Reversible Data Hiding With Brainstorming Optimization for UAV Aerial Images
abstract
In recent years, with the rapid development of unmanned aerial vehicle (UAV), aerial images have extended across various industries such as intelligent building, agriculture, transportation, and Industry 4.0. Notably, the security of UAV‐assisted data acquisition during transmission has become a critical concern. The reversible data hiding (RDH) method can hide data in aerial images for transmission and ensure secure communication. In general, an aerial image may exhibit substantially different orientation regularity from a natural scene image. This casts major challenges to the RDH method, for which existing approaches lack effective mechanisms to capture such content type variations, and thus are difficult to generalize from one type to another. In this paper, the orientation‐aware selectivity mechanism is introduced to achieve an accurate orientation‐aware prediction along different directions in local regions with different structure regularity. Furthermore, we propose a progressive brainstorming optimization algorithm (BSO)‐guided optimal PSNR value strategy, which can obtain a superior perceptual performance and the corresponding thresholds by further exploring the pixel correlations within the UAV aerial images. Experimental results on the USC‐SIPI Miscellaneous dataset and two challenging aerial datasets, including the USC‐SIPI High Altitude Aerial Imagery dataset and the Kaggle dataset, demonstrate that the proposed framework enhances the imperceptibility powerfully in marked UAV aerial images and ensures sufficient embedding capacity effectively. The average PSNR of the marked image obtained by the proposed method is 63.85 dB when embedded with 30,000 bits of data, which is an improvement of 0.59 dB compared to the current state‐of‐the‐art RDH methods.
Xiaodan Tai, Yannan Ren, Jing Li 0046, Jiande Sun 0001, Kai Zhang 0010, Wenbo Wan
Int. J. Intell. Syst.5
2025 Multiscale Integration Network With Quaternion Convolution for Pansharpening
abstract
In this letter, we proposed a multiscale integration network with quaternion convolution (MQ-Net) for the fusion of low spatial resolution multispectral (LRMS) and panchromatic (PAN) images. In this network, LRMS and PAN images are resampled at different scales and fed into feature fusion modules (FFMs) to merge the spatial and spectral information among them. Then, multiscale feature enhancement modules (MFEMs) are designed to sufficiently learn the spatial and spectral information at different scales. Meanwhile, we employ a quaternion convolution module (QCM) to better capture the dependencies within spectral bands of LRMS images. Then, the quaternion features are introduced into MFEMs for efficient feature enhancement. Finally, all information from different scales is integrated for the reconstruction of high LRMS images. Reduced- and full-resolution experiments are performed on GeoEye-1 and WorldView-2 satellite datasets. Compared to some state-of-the-art pansharpening methods, the proposed MQ-Net obtains better results in terms of qualitative and quantitative evaluations. The code is available athttps://github.com/RSMagneto/MQ-Net.
Yingjie Kong, Xuquan Wang, Kai Zhang 0010, Hong Li 0005, Wenbo Wan, Jiande Sun 0001
IEEE Geosci. Remote. Sens. Lett.3
2025 Spatial-spectral unfolding network with mutual guidance for multispectral and hyperspectral image fusion
Kai Zhang 0010, Qinzhu Sun, Chiru Ge, Wenbo Wan, Jiande Sun 0001, Huaxiang Zhang 0001
Pattern Recognit.2
2025 Dual-Conditionally Guided Diffusion Models for Fusion of Unregistered Multisource Remote Sensing Images
abstract
Remote sensing image fusion is a critical technique for enhancing the quality of remote sensing images. Typically, it is presumed that images have been accurately registered to facilitate effective fusion. Traditional methods involve a sequential process of registration followed by fusion, or a parallel approach to registration and fusion, which do not fully leverage the interplay between these two stages. To overcome this limitation, we propose a novel joint learning method termed the Dual-Conditionally Guided Registration and Fusion Diffusion Network (RFDifNet) for unregistered remote sensing images. This method integrates image registration and fusion into a unified framework. The RFDifNet comprises two conditional diffusion-based subnetworks: the Registration Diffusion Module (RegDM) and the Fusion Diffusion Module (FusDM). In this architecture, the RegDM corrects the misalignment of the unregistered image and provides it as a conditional input to the FusDM to generate the fused image. Conversely, the fused image is also fed back into the RegDM as a conditional input, enabling a closed-loop iteration of the registration and fusion processes. To further refine the image registration by reducing noise interference and preserving edge details within the RegDM, we introduce a joint learning loss function based on fractional-order derivatives, demonstrating superior performance in geometric and detail preservation compared to traditional methods. Experimental results validate the outstanding performance of the proposed RFDifNet in both image registration and fusion tasks. The source code is available at: https://github.com/DDXNJUST/RFDifNet.
Wenxiu Diao, Ling Hu 0003, Kai Zhang 0010, Liang Xiao 0001
IEEE Trans. Geosci. Remote. Sens.3
2025 Texture-Content Dual Guided Network for Visible and Infrared Image Fusion
abstract
The preservation and enhancement of texture information is crucial for the fusion of visible and infrared images. However, most current deep neural network (DNN)-based methods ignore the differences between texture and content, leading to unsatisfactory fusion results. To further enhance the quality of fused images, we propose a texture-content dual guided (TCDG-Net) network, which produces the fused image by the guidance inferred from source images. Specifically, a texture map is first estimated jointly by combining the gradient information of visible and infrared images. Then, the features learned by the shallow feature extraction (SFE) module are enhanced with the guidance of the texture map. To effectively model the texture information in the long-range dependencies, we design the texture-guided enhancement (TGE) module, in which the texture-guided attention mechanism is utilized to capture the global similarity of the texture regions in source images. Meanwhile, we employ the content-guided enhancement (CGE) module to refine the content regions in the fused result by utilizing the complement of the texture map. Finally, the fused image is generated by adaptively integrating the enhanced texture and content information. Extensive experiments on three benchmark datasets demonstrate the effectiveness of the proposed TCDG-Net in terms of qualitative and quantitative evaluations. Besides, the fused images generated by our proposed TCDG-Net also show better performance in downstream tasks, such as objection detection and semantic segmentation.
Kai Zhang 0010, Ludan Sun, Wenbo Wan, Jiande Sun 0001, Shuyuan Yang 0001, Huaxiang Zhang 0001
IEEE Trans. Multim.1
2024 Complementary Fusion Network Based on Frequency Hybrid Attention for Pansharpening
abstract
Pansharpening is a feasible way to obtain the high-resolution (HR) multispectral (MS) images by using panchromatic (PAN) images to sharpen low-resolution MS images. Despite its great advances, most existing pansharpening methods neglect the importance of integrating local and non-local characteristics of images, resulting in the imbalance of spatial and spectral distribution. In this paper, we propose a complementary fusion network (CFNet) based on frequency hybrid attention mechanism for pansharpening. By introducing the frequency transformation and the deformable cross-attention, our model takes image-wide receptive field into consideration to explore global feature learning. Combined with the convolutional layers with local receptive field, CFNet can well capture local and non-local features. Experimental results demonstrate that the proposed method outperforms the comparison methods in terms of visual and quantitative qualities.
Yinghui Xing, Litao Qu, Kai Zhang 0010, Yan Zhang 0127, Xiuwei Zhang 0001, Yanning Zhang 0001
ICASSP3
2024 Efficient Image Harmonization via RGB Transformation
abstract
Image harmonization aims to adjust the appearance of the foreground to make it harmonious with the background, thereby maintaining visual consistency in composite images. Previous deep learning-based methods have mainly focused on reconstructing harmonized images with the same size as the input composite images, often leading to complex network structures and a large number of parameters. In this paper, we propose a simple yet effective lightweight image harmonization network architecture. First, we generate a low-resolution 3-channel feature map to represent the rough variations in the RGB channels of a composite image, which is then upsampled and added to this composite image to obtain a preliminary harmonization result. Then, a refinement module is applied to refine the preliminary result and output the final harmonized image. Additionally, we design a dynamic data generation and training strategy to pre-train our model on another dataset. Experimental results on the iHarmony4 dataset show that our method indicates a significant reduction in the number of parameters compared to other methods, yet it still achieved competitive performance.
Jiande Sun 0001, Wenbo Wan, Kai Zhang 0010, Jian Wang 0004
MMSP5
2024 Triple disentangled network with dual attention for remote sensing image fusion
Feng Zhang 0028, Guishuo Yang, Jiande Sun 0001, Wenbo Wan, Kai Zhang 0010
Expert Syst. Appl.5
2024 Learning spatial-spectral dual adaptive graph embedding for multispectral and hyperspectral image fusion
Xuquan Wang, Feng Zhang 0028, Kai Zhang 0010, Weijie Wang 0002, Xiong Dun, Jiande Sun 0001
Pattern Recognit.3
2024 Joint Spatio-Temporal Modeling for Semantic Change Detection in Remote Sensing Images
abstract
Semantic Change Detection (SCD) refers to the task of simultaneously extracting the changed areas and the semantic categories (before and after the changes) in Remote Sensing Images (RSIs). This is more meaningful than Binary Change Detection (BCD) since it enables detailed change analysis in the observed areas. Previous works established triple-branch Convolutional Neural Network (CNN) architectures as the paradigm for SCD. However, it remains challenging to exploit semantic information with a limited amount of change samples. In this work, we investigate to jointly consider the spatio-temporal dependencies to improve the accuracy of SCD. First, we propose a Semantic Change Transformer (SCanFormer) to explicitly model the ’from-to’ semantic transitions between the bi-temporal RSIs. Then, we introduce a semantic learning scheme to leverage the spatio-temporal constraints, which are coherent to the SCD task, to guide the learning of semantic changes. The resulting network (SCanNet) significantly outperforms the baseline method in terms of both detection of critical semantic changes and semantic consistency in the obtained bi-temporal results. It achieves the SOTA accuracy on two benchmark datasets for the SCD.
Lei Ding 0008, Jing Zhang 0023, Haitao Guo, Kai Zhang 0010, Bing Liu 0018, Lorenzo Bruzzone
IEEE Trans. Geosci. Remote. Sens.4
2024 Building Change Detection in Earthquake: A Multiscale Interaction Network With Offset Calibration and a Dataset
abstract
As one of the most destructive natural disasters, earthquakes have struck many countries around the world in recent years, causing serious economic losses. Change detection (CD) can be applied to postearthquake building CD as it can infer interested change regions from multitemporal remote sensing (RS) images. Furthermore, the CD with short imaging intervals will better satisfy the needs of the emergency rescues after earthquakes. However, the capability of current methods built on deep neural networks (DNNs) is limited because the dataset with short imaging intervals is absent. To meet postdisaster immediate relief, we create a CD dataset, the Turkey earthquake CD dataset (TUE-CD), for the detection of building collapse in the short term after an earthquake. Due to the high requirement for timeliness of postevent images, the orbit of the satellite during postevent imaging deviates from that during preevent imaging, which leads to a side-looking problem between bitemporal images. To deal with these challenges, we present a multiscale feature interaction network (MSI-Net) for efficient interaction between bitemporal features, as well as mitigating the effect of side-looking problems. Specifically, the proposed MSI-Net consists of joint cross-attention (JCA) modules, multiscale offset calibration (MOC) modules, and feature integration (FeI) modules. The JCA module unifies channel cross-attention (CCA) and spatial joint attention (SJA) for sufficient feature interaction. The MOC module further estimates the offsets to align the bitemporal image with the multiscale features. Finally, calibrated features and multiscale features are fused by FeI modules for the prediction of changed areas. The best mF1 and mIoU scores are achieved on two public datasets and the constructed TUE-CD dataset: WHU-CD (95.58%, 91.81%), CLCD (82.96%, 73.53%), and TUE-CD (78.02%, 68.48%). Experimental results demonstrate that the proposed MSI-Net provides competitive performance compared to the state-of-the-art CD methods. The TUE-CD dataset and the code of MSI-Net will be available athttps://github.com/RSMagneto/MSI-Net.
Yunlong Liu 0005, Kai Zhang 0010, Chunan Guan, Shanxin Zhang, Hong Li 0005, Wenbo Wan, Jiande Sun 0001
IEEE Trans. Geosci. Remote. Sens.2
2024 Content-Guided Spatial-Spectral Integration Network for Change Detection in HR Remote Sensing Images
abstract
The integration of spatial and spectral information is beneficial to the improvement of change detection (CD) performance. However, existing methods cannot efficiently suppress the influences of spatial and spectral differences (SDs) in unchanged areas. To address these issues, in this article, we propose a content-guided spatial–spectral integration network (CSI-Net) for the fusion of global spatial details and SD information. Specifically, the proposed CSI-Net is composed of a spatial reasoning (SR) module, an SD module, and a content-guided integration (CGI) module. In the SR module, the spatial information is learned by cascaded graph convolution (GC) blocks for global modeling. The SD module is responsible for the extraction of spectral features, by calculating the means and variances of features to reduce the impact of SDs in unchanged regions. In addition, in order to integrate the spatial–spectral features efficiently, we design a CGI module to further take advantage of their complementary information. In this module, high-level content information is introduced as a guide for proper interaction. Due to the efficient spatial–spectral fusion, the proposed CSI-Net can learn the changed features better while achieving suppression of SDs. Experimental results on LEVIR-CD, WHU-CD, and CLCD datasets demonstrate that the proposed CSI-Net produces better performance compared to state-of-the-art methods, and is applicable to different scenarios. The code of CSI-Net is available athttps://github.com/RSMagneto/CSI-Net.
Yunlong Liu 0005, Feng Zhang 0028, Shanxin Zhang, Kai Zhang 0010, Jiande Sun 0001, Lorenzo Bruzzone
IEEE Trans. Geosci. Remote. Sens.4
2024 Spectral-Spatial Dual Graph Unfolding Network for Multispectral and Hyperspectral Image Fusion
abstract
Recently, deep neural network (DNN)-based methods have achieved good results in terms of the fusion of low spatial resolution hyperspectral (LR HS) and high spatial resolution multispectral (HR MS) images. However, the spectral band correlation (SBC) and the spatial nonlocal similarity (SNS) in hyperspectral (HS) images are not sufficiently exploited by them. To model the two priors efficiently, we propose a spectral-spatial dual graph unfolding network (SDGU-Net), which is derived from the optimization of graph regularized restoration models. Specifically, we introduce spectral and spatial graphs to regularize the reconstruction of the desired high spatial resolution hyperspectral (HR HS) image. To explore the SBC and SNS priors of HS images in feature space and utilize the powerful learning ability of DNNs simultaneously, the iterative optimization of the spectral and spatial graph regularized models is unfolded as a network, which is composed of spectral and spatial graph unfolding modules. The two kinds of modules are designed according to the solutions of the spectral and spatial graph regularized models. In these modules, we employ graph convolution networks (GCNs) to capture the SBC and SNS in the fused image. Then, the learned features are integrated by the corresponding feature fusion modules and fed into the feature condense module to generate the HR HS image. We conduct extensive experiments on three benchmark datasets and the results demonstrate the effectiveness of our proposed SDGU-Net.
Kai Zhang 0010, Feng Zhang 0028, Chiru Ge, Wenbo Wan, Jiande Sun 0001, Huaxiang Zhang 0001
IEEE Trans. Geosci. Remote. Sens.1
2024 DRFormer: Learning Disentangled Representation for Pan-Sharpening via Mutual Information- Based Transformer
abstract
In this article, we propose a new pan-sharpening method that disentangles low spatial resolution multispectral (LRMS) and panchromatic (PAN) images in terms of sensor-specific features and common features. These features are obtained by defining mutual information (MI)-based transformers designed to achieve disentangled learning. In the proposed method, LRMS and PAN images are cross-reconstructed by cross-coupled transformers to facilitate the disentanglement of the common features and sensor-specific features. To ensure compatibility among the disentangled features, self-reconstructions of LRMS and PAN images are imposed on them, and source images are reconstructed by self-coupled transformers. In addition to the reconstruction-guided disentangled learning, we maximize the MI between the common features of LRMS and PAN images to improve the correlation of the common features from different images. We also minimize the MI between the common features and sensor-specific features from the same image to reduce the redundancy among them. Through the reconstruction and disentangled representation of source images, sensor-specific features and common features can be decomposed efficiently. Finally, all disentangled features are integrated by a fusion transformer to generate the high spatial resolution multispectral (HRMS) image. Experiments on different datasets demonstrate that the proposed method produces competitive fusion results. The code is available athttps://github.com/RSMagneto/DRFormer.
Feng Zhang 0028, Kai Zhang 0010, Jiande Sun 0001, Jian Wang 0004, Lorenzo Bruzzone
IEEE Trans. Geosci. Remote. Sens.2
2024 CrossDiff: Exploring Self-SupervisedRepresentation of Pansharpening via Cross-Predictive Diffusion Model
abstract
Fusion of a panchromatic (PAN) image and corresponding multispectral (MS) image is also known as pansharpening, which aims to combine abundant spatial details of PAN and spectral information of MS images. Due to the absence of high-resolution MS images, available deep-learning-based methods usually follow the paradigm of training at reduced resolution and testing at both reduced and full resolution. When taking original MS and PAN images as inputs, they always obtain sub-optimal results due to the scale variation. In this paper, we propose to explore the self-supervised representation for pansharpening by designing a cross-predictive diffusion model, named CrossDiff. It has two-stage training. In the first stage, we introduce a cross-predictive pretext task to pre-train the UNet structure based on conditional Denoising Diffusion Probabilistic Model (DDPM). While in the second stage, the encoders of the UNets are frozen to directly extract spatial and spectral features from PAN and MS images, and only the fusion head is trained to adapt for pansharpening task. Extensive experiments show the effectiveness and superiority of the proposed model compared with state-of-the-art supervised and unsupervised methods. Besides, the cross-sensor experiments also verify the generalization ability of proposed self-supervised representation learners for other satellite datasets. Code is available at https://github.com/codgodtao/CrossDiff.
Yinghui Xing, Litao Qu, Shizhou Zhang, Kai Zhang 0010, Yanning Zhang 0001, Lorenzo Bruzzone
IEEE Trans. Image Process.4
2024 Deep Rank-N Decomposition Network for Image Fusion
abstract
Existing deep neural network (DNN)-based image fusion methods seldom consider low-rank priors for the decomposition of source images, which cannot efficiently model base and detail components in images. To exploit the low-rank priors better, we propose a deep rank-Ndecomposition network (DRDec-Net) according to the rank-Ndecomposition of source images. Specifically, a rank-Ndecomposition model is first established by imposing low-rank priors on the base component of source images. Then, based on the decomposition model, we construct DRDec-Net, which is composed of low-rank decomposition (LRD) modules, a detail fusion (DetailF) module, and a low-rank fusion (LRF) module. In DRDec-Net, it is assumed that source images share the same base component, which is expressed as the sum of rank-1 components. We employNcascaded LRD modules to extract these rank-1 components from source images. Meanwhile, detail components are obtained by subtracting the base component from source images. Next, the extracted rank-1 components and detail components are integrated by LRF and DetailF modules to produce the base component and detail component of the fused image. Finally, the sum of the two obtained components is regarded as the fused image. Compared to some state-of-the-art methods, experimental results demonstrate that the proposed DRDec-Net can produce a better performance on three image fusion tasks, including infrared and visible images, multi-exposure images, and multi-focus images.
Ludan Sun, Kai Zhang 0010, Feng Zhang 0028, Wenbo Wan, Jiande Sun 0001
IEEE Trans. Multim.2
2023 S-Feature Pyramid Network and Attention Model for Drone Detection
abstract
The issue of aviation safety has always received a great of attention and focus, and birds are also an important issue in aviation safety. Nowadays, drones have emerged and share the same airspace with birds at low altitudes. The problems associated with drones should also be taken into account. For example, small drones can be misused for illegal activities and the threat from them is on the rise. Driven by this situation, we used data provided by the ICASSP Drone-vs-Bird detection Grand Challenge for drone detection and used the method of adding shallow feature pyramid network and attention model on SSD [1] (SFA-SSD) to solve the problem of drone detection in competition. Out of 30 test videos, our method was able to detect drones in 11 videos, with 8 videos scoring above 0.1 and only 3 videos scoring above 0.7.
Pengcheng Dong, Chuntao Wang, Zhenyong Lu, Kai Zhang 0010, Wenbo Wan, Jiande Sun 0001
ICASSP4
2023 Long-Short Attention Network For The Spectral Super-Resolution Of Multispectral Images
abstract
Owing to the efficiency in terms of the modeling of long-range dependencies, transformer-based spectral reconstruction methods have produced satisfactory hyperspectral (HS) images from multispectral (MS) images. Some transformer-based methods applied self-attention to all bands in the HS image to model the relationships among them, which ignore high correlations between adjacent bands and low correlations among nonadjacent ones. To learn the global relationships among all bands and the correlations between adjacent bands simultaneously, this paper proposes a long-short attention network (LSA-Net) for the spectral super-resolution of MS images. Specifically, LSA-Net is composed of cascaded long-range attention blocks and short-range attention blocks. In long-range attention blocks, the transformer is imposed on all channels by modeling each channel as a token. Then, grouped channels are fed into short-range attention blocks for correlation learning, which is inferred from the similarities among neighboring channels. With the introduction of long- and short-range attention, the relationships among spectral bands can be preserved better. Experiments on the CAVE dataset demonstrate the effectiveness of the proposed LSA-Net. The code is available at https://github.com/RSMagneto/LSA-Net.
Kai Zhang 0010, Feng Zhang 0028, Jiande Sun 0001
ICASSP1
2023 Hierarchical Feature Fusion and Selection for Hyperspectral Image Classification
abstract
Most existing classification methods design complicated and large deep neural network (DNN) model to deal with the ubiquitous spectral variability and nonlinearity of hyperspectral images (HSIs). However, their application is blocked by limited training samples and considerable computational costs in real scenes. To solve these problems, we propose a simple spectral hierarchical feature fusion and selection network (HFFSNet). Specifically, we apply 1-D grouped convolution for dimensionality reduction and multilevel feature extraction, then the multilevel features are fused to assist the adaptive feature selection of different layer features via the soft attention mechanism, and finally the selected features are fused to further enhance the feature representation. Extensive experimental results on three hyperspectral datasets demonstrate the effectiveness of the proposed network.
Zhixi Feng, Xuehu Liu, Shuyuan Yang 0001, Kai Zhang 0010, Licheng Jiao
IEEE Geosci. Remote. Sens. Lett.4
2023 Multispectral and hyperspectral image fusion based on low-rank unfolding network
Kai Zhang 0010, Feng Zhang 0028, Chiru Ge, Wenbo Wan, Jiande Sun 0001
Signal Process.2
2023 3D geometrical total variation regularized low-rank matrix factorization for hyperspectral image denoising
Feng Zhang 0028, Kai Zhang 0010, Wenbo Wan, Jiande Sun 0001
Signal Process.2
2023 A Unified Two-Stage Spatial and Spectral Network With Few-Shot Learning for Pansharpening
abstract
Recently, pan-sharpening methods based on deep learning (DL) have achieved state-of-the-art results. However, current existing DL-based pan-sharpening methods need to be trained repetitively for different satellite sensors to obtain satisfactory fusion performance and therefore require a large number of training images for each satellite. To deal with these issues, in this paper we propose a unified two-stage spatial and spectral network (UTSN) for pan-sharpening. A branch of networks is constructed for each different satellite, in which the spatial enhancement network (SEN) is shared to improve the spatial details in the fused images from different satellites. A spectral adjustment network (SAN) is employed to capture the spectral characteristics of the specific satellite. Through SAN, the spectral information in the intermediate image from SEN is refined to produce the final fusion results. Such a framework can integrate the datasets from different satellites together for sufficient training of SEN. The proposed method is able to achieve promising pan-sharpening results also for a new satellite with limited training images by only learning a new SAN on the few-shot datasets due to the simple but efficient structure of SAN. The experimental results show that the proposed method can produce state-of-the-art fusion results in both the standard and few-shot cases. The source code is publicly available at https://github.com/RSMagneto/UTSN.
Zhi Sheng, Feng Zhang 0028, Jiande Sun 0001, Yanyan Tan, Kai Zhang 0010, Lorenzo Bruzzone
IEEE Trans. Geosci. Remote. Sens.5
2023 Cross-Resolution Semi-Supervised Adversarial Learning for Pansharpening
abstract
Existing deep neural network (DNN)-based methods have produced good pansharpened images. However, supervised DNN-based pansharpening methods suffer from performance degradation when fusing low spatial resolution multispectral (LR MS) and panchromatic (PAN) images at full resolution. Unsupervised DNN-based methods alleviate the issue, but their training becomes difficult owing to the absence of reference images. This article establishes a novel semi-supervised framework to jointly learn the reconstructions of fused images from reduced- and full-resolution datasets. Specifically, we propose a cross-resolution semi-supervised adversarial learning network (CrossNet), which is composed of a supervised module and an unsupervised module. In these two modules, reduced- or full-resolution source images are disentangled as resolution-invariant components and resolution-aware components by reconstructing the fused images. Moreover, cross-resolution fused images are synthesized to enhance the disentanglement of the two kinds of components. Through the reconstruction of cross-resolution fused images, supervised and unsupervised modules are also coupled efficiently. Then, the semi-supervised framework can simultaneously make use of the supervised information in reduced-resolution datasets and mitigate the performance degradation via full-resolution datasets. Besides, adversarial learning is employed to improve the consistency between resolution-invariant components of source images at different resolutions. Finally, extensive experiments on QuickBird, GeoEye-1, WorldView-2, and WorldView-3 datasets demonstrate that the proposed CrossNet can produce state-of-the-art fusion results in terms of qualitative and quantitative evaluations. The source code is available athttps://github.com/RSMagneto/CrossNet.
Guishuo Yang, Kai Zhang 0010, Feng Zhang 0028, Jian Wang 0004, Jiande Sun 0001
IEEE Trans. Geosci. Remote. Sens.2
2023 Spatial-Spectral Dual Back-Projection Network for Pansharpening
abstract
Deep unfolding networks have obtained satisfactory performance in the pansharpening task owing to their sufficient interpretability. Inspired by the back-projection (BP) mechanism, we propose a BP-driven model, spatial-spectral dual back-project network (S2DBPN), to fuse the low spatial resolution multispectral (LR MS) and the high spatial resolution panchromatic (PAN) images by exploiting the BP in spatial and spectral domains. Specifically, the proposed S2DBPN is made up of a spatial BP network, a spectral BP network, and a reconstruction network. In the spatial BP network, spatial down- and up-projection modules are derived from BP, which is responsible for the projection of the LR MS image into the spatial domain. By analogy with the spatial BP, we reformulate the degradation between high spatial resolution multispectral (HR MS) and PAN images as spectral down- and up-projections. Then, the spectral BP network is constructed for the projection of the PAN image along the channel dimension. Finally, the features from spatial and spectral BP networks are integrated to produce the desired HR MS image through the reconstruction network. Compared to the state-of-the-art methods, extensive experiments on QuickBird, GeoEye-1, and WorldView-2 datasets demonstrate that our S2DBPN produces better HR MS images in terms of qualitative and quantitative evaluation metrics. The code of S2DBPN is released at: https://github.com/RSMagneto/S2DBPN.
Kai Zhang 0010, Anfei Wang, Feng Zhang 0028, Wenbo Wan, Jiande Sun 0001, Lorenzo Bruzzone
IEEE Trans. Geosci. Remote. Sens.1
2023 Learning Deep Multiscale Local Dissimilarity Prior for Pansharpening
abstract
Various deep neural networks (DNNs) have been constructed to inject the spatial information of the panchromatic (PAN) image into the low spatial resolution multispectral (LR MS) image. However, most of them ignore the local dissimilarity (LD) prior between MS and PAN images, which has a negative influence on the fused image. Considering the above-mentioned issues, we propose a deep multiscale local dissimilarity network (DMLD-Net) to learn the LD prior at different scales and enhance the spatial and spectral information in the fused image better. Specifically, we first synthesize a downsampled PAN image from the original PAN image to match the scale of the LR MS image. Then, a LD metric is designed to calculate the dissimilarity map between the two images in feature space. According to the learned dissimilarity map, we utilize a LD-guided attention block (LDGAB) to suppress the impact of LD, which filters out the dissimilar information in the features of the PAN image. To learn the LD prior between MS and PAN images sufficiently, the multiscale architecture is considered and we infer the dissimilar maps hierarchically and inject filtered features into the LR MS image progressively. Finally, the fused image is generated by a reconstruction block. Through the LD learning at different scales, reasonable spatial information is extracted from the PAN image, by which the distortions in the fused image caused by LD can be reduced efficiently. Extensive experiments are conducted on GeoEye-1 and WorldView-2 datasets and the results demonstrate the effectiveness of the proposed DMLD-Net in terms of spatial and spectral preservation. The code is available at https://github.com/RSMagneto/DMLD-Net.
Kai Zhang 0010, Guishuo Yang, Feng Zhang 0028, Wenbo Wan, Man Zhou 0003, Jiande Sun 0001, Huaxiang Zhang 0001
IEEE Trans. Geosci. Remote. Sens.1
2023 Relation Changes Matter: Cross-Temporal Difference Transformer for Change Detection in Remote Sensing Images
abstract
Thanks to their capability of modeling global information, transformers have been recently applied to change detection in remote sensing images. Generally, the changes in terms of shape and appearance of objects lead to relation changes among these objects in multi-temporal images. However, in this context, the attention mechanism in transformers has not been fully explored yet to learn relation changes in the observed scenes. In this paper, we analyze the relation changes in multi-temporal images and propose a cross-temporal difference (CTD) attention to capture these changes efficiently. Through the CTD attention, the changed areas are distinguished better from the unchanged areas. Based on the CTD attention, two CTD-transformer encoders are constructed to extract the features of changed areas from the embedded tokens of multi-temporal images in a cross manner. Then, the extracted features at the coarse scale are further improved to the fine-scale by the corresponding CTD-transformer decoders. In addition, consistency-perception blocks (CPBs) are designed to preserve the structures and contours of changed areas. Finally, all extracted features from multi-temporal images are concatenated to produce the desired change map. Compared to state-of-the-art methods, experimental results on LEVIR-CD, WHU-CD, and CLCD datasets demonstrate that the proposed method produces better performance. The source code is available at https://github.com/RSMagneto/CTD-Former.
Kai Zhang 0010, Feng Zhang 0028, Lei Ding 0008, Jiande Sun 0001, Lorenzo Bruzzone
IEEE Trans. Geosci. Remote. Sens.1
2023 ZeRGAN: Zero-Reference GAN for Fusion of Multispectral and Panchromatic Images
abstract
In this article, we present a new pansharpening method, a zero-reference generative adversarial network (ZeRGAN), which fuses low spatial resolution multispectral (LR MS) and high spatial resolution panchromatic (PAN) images. In the proposed method, zero-reference indicates that it does not require paired reduced-scale images or unpaired full-scale images for training. To obtain accurate fusion results, we establish an adversarial game between a set of multiscale generators and their corresponding discriminators. Through multiscale generators, the fused high spatial resolution MS (HR MS) images are progressively produced from LR MS and PAN images, while the discriminators aim to distinguish the differences of spatial information between the HR MS images and the PAN images. In other words, the HR MS images are generated from LR MS and PAN images after the optimization of ZeRGAN. Furthermore, we construct a nonreference loss function, including an adversarial loss, spatial and spectral reconstruction losses, a spatial enhancement loss, and an average constancy loss. Through the minimization of the total loss, the spatial details in the HR MS images can be enhanced efficiently. Extensive experiments are implemented on datasets acquired by different satellites. The results demonstrate that the effectiveness of the proposed method compared with the state-of-the-art methods. The source code is publicly available at https://github.com/RSMagneto/ZeRGAN.
Wenxiu Diao, Feng Zhang 0028, Jiande Sun 0001, Yinghui Xing, Kai Zhang 0010, Lorenzo Bruzzone
IEEE Trans. Neural Networks Learn. Syst.5
2022 TAGAN: Texture and Attention Guided Generative Adversarial Network for Image Super Resolution
abstract
Super Resolution (SR) methods based on Generative Adversarial Networks (GANs) accomplish predominant execution in visual perception and image quality. These methods are mainly generated by traditional Peak-Signal-to-Noise-Ratio (PSNR)oriented or perceptual-driven. As the reconstruction process usually loses high frequency information, various methods aim to preserve more details. To make the details of the generated image richer, the Gradient Weight (GW) loss is introduced in the proposed method, because the gradient can reflect the texture of the image to a certain extent. The GW loss function is helpful to improve the edge and detailed texture of the generated image. Furthermore, we introduce attention mechanism to the image reconstruction block via Squeeze and Excitation Net (SENet). Attention mechanism can effectively aggregate the global features obtained by the nonlinear mapping network, and improve the channel sensitivity of the model. With the help of GW and attention mechanism, the proposed method can achieve better performance and visual quality in image texture detail restoration. The performance comparison between the state-of-the-art methods and our proposed method verifies the feasibility and reliability of the proposed method.
Haitao Wang 0023, Jiande Sun 0001, Wenxiu Diao, Jing Li 0046, Kai Zhang 0010
ISCAS5
2022 HLF-Net: Pansharpening Based on High- and Low-Frequency Fusion Networks
abstract
Many deep neural networks have been constructed for the pansharpening task. However, the differences between the high and low frequencies in images are not considered in some DNN-based pansharpening methods. As high and low frequencies have different information of images, it is difficult for the same network to learn and reconcile the two kinds of frequencies. Considering the aforementioned differences, we propose a new pansharpening network to fuse the high and low frequencies in low spatial resolution multispectral and panchromatic images separately. Specifically, a high and low frequency fusion network is constructed, which is composed of a high-frequency fusion network and a low-frequency fusion network. In the high-frequency fusion network, skip attention is introduced into U-Net to better retain the high frequencies in feature maps. The low-frequency fusion network uses the involution to capture the dependency among the channels of feature maps. Experiments on the GeoEye-1 dataset reveal that the proposed network outperforms some state-of-the-art methods. The code can be accessed at https://github.com/RSMagneto/HLF-Net.
Wenxiu Diao, Feng Zhang 0028, Haitao Wang 0023, Wenbo Wan, Jiande Sun 0001, Kai Zhang 0010
IEEE Geosci. Remote. Sens. Lett.6
2022 Unsupervised Change Detection of Multispectral Images Based on PCA and Low-Rank Prior
abstract
In this letter, we propose a new unsupervised change detection method based on low-rank prior for multispectral images. It is assumed that the changed and unchanged pixels are from different subspaces due to different appearance and statistical properties. So, low-rank representation (LRR) is employed to find informative pixels from the superpixels of the difference image (DI). Besides, taking the sparsity of changed pixels in the observed scenes into consideration, the selection rule is designed to distinguish these pixels. Then, principal component analysis (PCA) is used for the training of changed and unchanged dictionaries from these pixels. Finally, the change map is estimated by comparing the reconstruction error of each pixel in DI on changed and unchanged dictionaries. By LRR, more representative pixels are found for subsequent dictionary learning, which can efficiently improve the performance of the proposed method. Experiments on multitemporal images from the Landsat satellite demonstrate the effectiveness of the proposed method.
Jing Li 0046, Feng Zhang 0028, Jiande Sun 0001, Kai Zhang 0010
IEEE Geosci. Remote. Sens. Lett.5
2022 Pan-Sharpening Based on Transformer With Redundancy Reduction
abstract
Pan-sharpening methods based on deep neural network (DNN) have produced the state-of-the-art results. However, the common information in the panchromatic (PAN) image and the low spatial resolution multispectral (LRMS) image is not sufficiently explored. As PAN and LRMS images are collected from the same scene, there exists some common information among them, in addition to their respective unique information. The direct concatenation of extracted features leads to some redundancy in the feature space. To reduce the redundancy among features and exploit the global information in source images, we proposed a novel pan-sharpening method by combining the convolution neural network and transformer. Specifically, PAN and LRMS images are encoded as unique features and common features by the subnetworks consisting of convolution blocks and transformer blocks. Then, the common features are averaged and combined with unique features from source images for the reconstruction of the fused image. To extract accurate common features, the equality constraint is imposed on them. Experimental results show that the proposed method outperforms the state-of-the-art methods on both reduced-scale and full-scale datasets. The source code is available athttps://github.com/RSMagneto/TRRNet.
Kai Zhang 0010, Feng Zhang 0028, Wenbo Wan, Jiande Sun 0001
IEEE Geosci. Remote. Sens. Lett.1
2022 Spatial and Spectral Extraction Network With Adaptive Feature Fusion for Pansharpening
abstract
Pansharpening methods based on deep neural networks (DNNs) have been attracting great attention due to their powerful representation capabilities. In this article, to combine the feature maps from different subnetworks efficiently, we propose a novel pansharpening method based on a spatial and spectral extraction network (SSE-Net). Different from the other methods based on DNNs that directly concatenate the features from different subnetworks, we design adaptive feature fusion modules (AFFMs) to merge these features according to their information content. First, the spatial and spectral features are extracted by the subnetworks from low spatial resolution multispectral (LR MS) and panchromatic (PAN) images. Then, by fusing the features at different levels, the desired high spatial resolution MS (HR MS) images are generated by the fusion network consisting of AFFMs. In the fusion network, the features from different subnetworks are integrated adaptively, and the redundancy among them is reduced. Moreover, the spectral ratio loss and the gradient loss are defined to ensure the effective learning of spatial and spectral features. The spectral ratio loss captures the nonlinear relationships among the bands in the MS image to reduce the spectral distortions in the fusion result. Extensive experiments were conducted on QuickBird and GeoEye-1 satellite datasets. Visual and numerical results demonstrate that the proposed method produces better fusion results compared with literature techniques. The source code is available athttps://github.com/RSMagneto/SSE-Net.
Kai Zhang 0010, Anfei Wang, Feng Zhang 0028, Wenxiu Diao, Jiande Sun 0001, Lorenzo Bruzzone
IEEE Trans. Geosci. Remote. Sens.1
2021 Sparse flow adversarial model for robust image compression
Shihui Zhao, Shuyuan Yang 0001, Zhi Liu 0010, Zhixi Feng, Kai Zhang 0010
Knowl. Based Syst.5
2021 Visual Security Assessment via Saliency-Weighted Structure and Orientation Similarity for Selective Encrypted Images
abstract
Selective encryption has been widely used in image privacy protection. Visual security assessment is necessary for the effectiveness and practicability of image encryption methods, and there have been a series of research studies on this aspect. However, these methods do not take into account perceptual factors. In this paper, we propose a new visual security assessment (VSA) by saliency-weighted structure and orientation similarity. Considering that the human visual perception is sensitive to the characteristics of selective encrypted images, we extract the structure and orientation feature maps, and then similarity measurements are conducted on these feature maps to generate the structure and orientation similarity maps. Next, we compute the saliency map of the original image. Then, a simple saliency-based pooling strategy is subsequently used to combine these measurements and generate the final visual security score. Extensive experiments are conducted on two public encryption databases, and the results demonstrate the superiority and robustness of our proposed VSA compared with the existing most advanced work.
Zhengguo Wu, Kai Zhang 0010, Yannan Ren, Jing Li 0046, Jiande Sun 0001, Wenbo Wan
Secur. Commun. Networks2
2020 Multispectral and Panchromatic Image Fusion Via Convolution Sparse Coding with Joint Sparsity
abstract
In this paper, a low spatial resolution multispectral (LR MS) and panchromatic (PAN) image fusion method based on convolution sparse coding (CSC) is proposed to model the global structures existing in source images. In the proposed method, CSC is adopted to decompose the high frequency (HF) component properly and joint sparse prior is also used to capture the correlation in the bands of MS images. By joint sparsity, the correlation is further inherited into their corresponding feature maps. Then, the spatial information in LR MS image is enhanced well after detailed fusion rule for spatial details. Finally, the fusion image is reconstructed by the fused low frequency and HF. The experimental results on real datasets from QuickBird and Geoeye-1 satellites verify that the proposed method can better preserve the spatial and spectral information in the fused images.
Feng Zhang 0028, Kai Zhang 0010
IGARSS2
2020 Superpixel guided structure sparsity for multispectral and hyperspectral image fusion over couple dictionary
Feng Zhang 0028, Kai Zhang 0010
Multim. Tools Appl.2
2019 Patch Based Pansharpening Using Weighted Nuclear Norm Minimization
abstract
This paper proposed a multispectral (MS) and panchromatic (PAN) image fusion method based on low-rank assumption captured by weighted nuclear norm minimization (WNNM). In this method, low-rank matrix factorization is considered to model the relationship between low spatial resolution (LR) and high spatial resolution (HR) MS images. In the formulation, MS and PAN images are partitioned into patches and then clustered to further guarantee the low-rank property. Besides, WNNM is used to capture the prior about singular values, in which larger singular values are shrunked with smaller weights. By WNNM, the spatial details in MS images can be well enhanced. Finally, the fusion model is established by combining the low-rank matrix factorization with the fidelity term about PAN image. The experimental results on degraded and real datasets demonstrate the effectiveness of the proposed method.
Kai Zhang 0010, Feng Zhang 0028
IGARSS1
2019 Convolution Structure Sparse Coding for Fusion of Panchromatic and Multispectral Images
abstract
Recently, sparse coding-based image fusion methods have been developed extensively. Although most of them can produce competitive fusion results, three issues need to be addressed: 1) these methods divide the image into overlapped patches and process them independently, which ignore the consistency of pixels in overlapped patches; 2) the partition strategy results in the loss of spatial structures for the entire image; and 3) the correlation in the bands of multispectral (MS) image is ignored. In this paper, we propose a novel image fusion method based on convolution structure sparse coding (CSSC) to deal with these issues. First, the proposed method combines convolution sparse coding with the degradation relationship of MS and panchromatic (PAN) images to establish a restoration model. Then, CSSC is elaborated to depict the correlation in the MS bands by introducing structural sparsity. Finally, feature maps over the constructed high-spatial-resolution (HR) and low-spatial-resolution (LR) filters are computed by alternative optimization to reconstruct the fused images. Besides, a joint HR/LR filter learning framework is also described in detail to ensure consistency and compatibility of HR/LR filters. Owing to the direct convolution on the entire image, the proposed CSSC fusion method avoids the partition of the image, which can efficiently exploit the global correlation and preserve the spatial structures in the image. The experimental results on QuickBird and Geoeye-1 satellite images show that the proposed method can produce better results by visual and numerical evaluation when compared with several well-known fusion methods.
Kai Zhang 0010, Min Wang 0007, Shuyuan Yang 0001, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.1
2019 Classification of Coral Reefs in the South China Sea by Combining Airborne LiDAR Bathymetry Bottom Waveforms and Bathymetric Features
abstract
Geographic information describing coral reefs plays an important role in constructing electronic chart systems and protecting the ecological environment of the ocean. To derive geographic information of coral reefs more effectively, this paper proposes a methodology to detect coral reefs by combining airborne LiDAR bathymetry (ALB) bottom waveform and bathymetric feature data. A feature vector was established by deriving bottom waveform variables (the peak amplitude, pulsewidth, area, skewness, kurtosis, and backscatter cross section) and bathymetric variables (the depth standard deviation, slope, bathymetric position index, Gaussian curvature, mean curvature, and roughness). Using a support vector machine classifier, coral reefs were detected by distinguishing two classes (coral reefs and others) on the seafloor. To evaluate the classification performance of coral reefs, the developed method was applied to Yuanzhi Island, South China Sea surveys, and verified by field data (aerial digital camera images and underwater video images). The results showed that the classification overall accuracy of coral reefs can be greatly improved from 80.59%/90.31% when ALB bottom waveform or bathymetric variables features were used separately to 93.57% when using a combination of ALB bottom waveform and bathymetric features. In addition, the kappa coefficient can also be greatly improved from approximately 0.61/0.80 to 0.87. And the new proposed method performs better compared to the current classification method using ALB data to detect coral reefs with an overall accuracy of 90.92% and Kappa of 0.81. This highlights the potential of ALB data, combining waveform data and bathymetric data, for precisely detecting coral reefs in shallow water areas.
Dianpeng Su, Fanlin Yang, Yue Ma 0002, Kai Zhang 0010, Jue Huang, Mingwei Wang 0004
IEEE Trans. Geosci. Remote. Sens.4
2019 Self-Paced Learning-Based Probability Subspace Projection for Hyperspectral Image Classification
abstract
In this paper a self-paced learning-based probability subspace projection (SL-PSP) method is proposed for hyperspectral image classification. First, a probability label is assigned for each pixel, and a risk is assigned for each labeled pixel. Then, two regularizers are developed from a self-paced maximum margin and a probability label graph, respectively. The first regularizer can increase the discriminant ability of features by gradually involving the most confident pixels into the projection to simultaneously push away heterogeneous neighbors and pull inhomogeneous neighbors. The second regularizer adopts a relaxed clustering assumption to make avail of unlabeled samples, thus accurately revealing the affinity between mixed pixels and achieving accurate classification with very few labeled samples. Several hyperspectral data sets are used to verify the effectiveness of SL-PSP, and the experimental results show that it can achieve the state-of-the-art results in terms of accuracy and stability.
Shuyuan Yang 0001, Zhixi Feng, Min Wang 0007, Kai Zhang 0010
IEEE Trans. Neural Networks Learn. Syst.4
2018 Sparse tensor neighbor embedding based pan-sharpening via N-way block pursuit
Min Wang 0007, Kai Zhang 0010, Xi Pan, Shuyuan Yang 0001
Knowl. Based Syst.2
2018 Salient Region Detection via Discriminative Dictionary Learning and Joint Bayesian Inference
abstract
In past decades, saliency detection has received increasing attention from computer vision communities, for its potential usage in many vision-related tasks. However, finding representative and discriminative features to accurately locate salient regions from complex scenes remains a challenging problem. Recent research on primary visual cortex (V1) shows that vision neurons are sparsely connected to form a compact representation of natural scenes and different visual stimuli are processed separately according to their semantic importance. Inspired by the above characteristics of visual perception, in this paper we advance a novel saliency detection method via representative and discriminative dictionary learning. An assumption that salient and nonsalient information are sparsely coded under two separate dictionaries is cast on the problem and we propose to learn a compact background dictionary from the image itself for saliency estimation. Different from previous methods, our saliency cues are obtained via active learning strategies rather than artificially designed rules, and thus is more adaptive. Followed by this, a probabilistic inference model is deduced to fully excavate multisource information about the scenes for high-quality saliency map generation. This joint inference scheme takes both spatial and color space information into consideration and is proved to be quite effective in practice. Finally, to investigate the performance of the proposed model, some experiments are conducted on two benchmark data sets along with other 20 state-of-the-art saliency detection approaches. The experimental results show that our method outperforms its counterparts and can correctly detect salient regions, even when other methods fail. Besides, the usability of the proposed method in real application-based cases is verified by applying it to content-based image resizing and promising results are obtained.
Shigang Wang 0001, Min Wang 0007, Shuyuan Yang 0001, Kai Zhang 0010
IEEE Trans. Circuits Syst. Video Technol.4
2018 Pansharpening With Multiscale Geometric Support Tensor Machine
abstract
In this paper, a new pansharpening method is proposed by constructing a set of multiscale geometric support tensor filters (MGSTFs). First, a least-square ridgelet support tensor machine is developed to derive a series of MGSTFs. Then the source images are formulated as tensors and filtered by MGSTFs to capture geometric and salient features of images. These features are then fused at each scale and direction to obtain the fused products. The distortions can be reduced by exploring the tensor formulation of multispectral data and endowing the filters’ directionality to capture the geometric details of images. Some experiments are carried out on several groups of QuickBird and GeoEye-1 images, and the results show that our proposed method can simultaneously reduce spectral distortions and preserve spatial details in the fused image.
Yinghui Xing, Min Wang 0007, Shuyuan Yang 0001, Kai Zhang 0010
IEEE Trans. Geosci. Remote. Sens.4
2018 Learning Low-Rank Decomposition for Pan-Sharpening With Spatial-Spectral Offsets
abstract
Finding accurate injection components is the key issue in pan-sharpening methods. In this paper, a low-rank pan-sharpening (LRP) model is developed from a new perspective of offset learning. Two offsets are defined to represent the spatial and spectral differences between low-resolution multispectral and high-resolution multispectral (HRMS) images, respectively. In order to reduce spatial and spectral distortions, spatial equalization and spectral proportion constraints are designed and cast on the offsets, to develop a spatial and spectral constrained stable low-rank decomposition algorithm via augmented Lagrange multiplier. By fine modeling and heuristic learning, our method can simultaneously reduce spatial and spectral distortions in the fused HRMS images. Moreover, our method can efficiently deal with noises and outliers in source images, for exploring low-rank and sparse characteristics of data. Extensive experiments are taken on several image data sets, and the results demonstrate the efficiency of the proposed LRP.
Shuyuan Yang 0001, Kai Zhang 0010, Min Wang 0007
IEEE Trans. Neural Networks Learn. Syst.2
2017 Multispectral and Hyperspectral Image Fusion Based on Group Spectral Embedding and Low-Rank Factorization
abstract
Fusing low spatial resolution hyperspectral (LRHS) images and high spatial resolution multispectral (HRMS) images to obtain high spatial resolution hyperspectral images (HRHS) has received increasing interests in recent years. In this paper, a new group spectral embedding (GSE)-based LRHS and HRMS image fusion method is proposed by exploring the multiple manifold structures of spectral bands and the low-rank structure of HRHS data. First, a low-rank factorization fusion (LRFF)-based robust recovery model is developed for HRHS images, by regarding HRMS images as the spectral degradation of HRHS images and exploring the group sparse prior of difference images. Then, an assumption that grouped spectral bands share the similar local geometry is cast on LRHS and HRHS images, to formulate a GSE regularizer in the LRFF model. Finally, an iterative optimization algorithm based on augmented Lagrangian multiplier is advanced to recover HRHS images. Experimental results on several data sets show the effectiveness of the proposed method on visual and numerical comparison.
Kai Zhang 0010, Min Wang 0007, Shuyuan Yang 0001
IEEE Trans. Geosci. Remote. Sens.1