Renwei Dian

dblp:191/2688 · DBLP profile ↗
← Back
46ranked-venue papers
14as first author
36since 2021 · last 2026
0000-0001-9197-6292ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 10 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 3 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 2 first-author · 7 since 2021
YearPublicationVenuePosition
2026 Equivariant Bayesian Hyperspectral Imaging via Mosaiced and PAN Image Fusion
abstract
A high-resolution (HR) hyperspectral imaging at video rates can be achieved by fusing a multi-band low-resolution (LR) mosaiced image and a single-band HR panchromatic (PAN) image in a single shot within a short time. However, the current fusion methods always suffer from a spatial or spectral distortion. To alleviate this, we propose an equivariant Bayesian variational inference framework. Specifically, we decompose an HR hyperspectral image (HSI) into the principal component and sparsity residual, which are modeled as latent variables with Gaussian priors. Each component is estimated via a shared deep neural network (DNN) under a variational inference framework, leveraging the shared spatial structures to enhance parameter efficiency and reconstruction accuracy. Additionally, to tackle the challenge of unavailable ground truth in real-world scenarios, we integrate the equivariant imaging (EI) prior with the Bayesian framework. By enforcing the consistency between the transformed fusion result and the re-inference output, this strategy enables the network to learn beyond the range space. Furthermore, we propose to utilize the learnable degradation functions derived from the physical imaging model to enable the proposed framework, which ensures an enhanced performance by posing plausible constraints on parameters of the degradation functions. Specifically, we explicitly model the point spread function (PSF) and spectral response function (SRF) with learnable parameters and impose non-negativity and sum-to-one constraints. Extensive experiments conducted on both simulated and real-world datasets demonstrate the effectiveness of the proposed framework, paving the way for HSI computational imaging.
Renwei Dian, Anjing Guo, Shutao Li 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2026 SSSMN: Spatial-spectral sparse Mamba network for efficient hyperspectral fusion super-resolution
Chenguo Feng, Renwei Dian, Shutao Li 0001
Pattern Recognit.3
2026 Frequency-Guided Lightweight Network for Unsupervised Spectral Demosaicing
abstract
Multispectral filter array (MSFA) cameras enable snapshot spectral data acquisition but require effective demosaicing algorithms to recover spatial-spectral information. Existing methods either rely on iterative optimization with limited high-frequency recovery, or require large-scale paired data and sophisticated network architectures that introduce considerable computational overhead. Moreover, insufficient frequency awareness in both spectral image initialization and the training process often leads to biased representations and unstable optimization under unsupervised settings. To address these issues, we propose a lightweight unsupervised spectral demosaicing framework, termed Frequency-Guided Lightweight Network (FGLN). First, we propose the Frequency-Parameterized Convolution (FPC) as the spectral initialization strategy, enhancing frequency coverage by explicitly embedding frequency-aware parameterization into convolutional weights. Second, we develop a frequency-aware unsupervised training framework based on equivariant imaging to regulate the optimization dynamics and suppress degenerate solutions. Finally, the Lightweight Spectral Demosaicing Unit (LSDU) improves reconstruction quality while maintaining low computational complexity. Extensive experiments on both synthetic and real-world datasets demonstrate that FGLN achieves competitive performance compared with state-of-the-art methods. The code will be uploaded at https://github.com/Matsuri247/FGLN.
Yaohang Wu, Renwei Dian, Shutao Li 0001
IEEE Trans. Circuits Syst. Video Technol.2
2026 Equivariant High-Resolution Hyperspectral Imaging via Mosaiced and PAN Image Fusion
abstract
Existing mosaic-based snapshot hyperspectral imaging systems struggle to capture high resolution (HR) hyperspectral image (HSI), limiting its application. Fusing a low resolution (LR) mosaiced image with an HR panchromatic (PAN) image serves as a feasible solution to obtain the HR HSI. Therefore, we propose a dual-sensor based HSI imaging system, combining a $4\times 4$ spectral filter array (SFA) mosaiced image sensor with a co-aligned PAN image sensor to provide complementary spatial-spectral information. To reconstruct HR HSI, we propose an unsupervised equivariant imaging (EI)-based training framework with a learnable degradation function, overcoming the inaccessibility of ground truth and spectral response function (SRF). Specifically, we formulate the degradation process as a combination of $8\times 8$ mosaicing and $2\times 2$ average downsampling for the LR mosaiced image, while modeling the PAN image as a linear projection of the HR HSI using SRF. Since parameters of SRF are inaccessible, we propose to make them learnable to have an accurate estimation. By enforcing transformation equivariance between the input-output pair of the fusion network, the proposed framework ensures the reconstructed HSI preserves spatial-spectral consistency without relying on paired supervision. Furthermore, we instantiate the proposed HSI imaging system and collect a real-world dataset of 60 paired mosaiced / PAN images. The mosaiced image exhibits 16 spectral bands ranging from 722 to 896 nm and $1020\times 1104$ spatial pixels while the PAN image exhibits $2040\times 2208$ spatial pixels. Comprehensive experiments demonstrate that the proposed method exhibits high spatial consistency and spectral fidelity while maintaining computational efficiency.
Anjing Guo, Renwei Dian, Shutao Li 0001
IEEE Trans. Image Process.3
2026 Deep Error-Aware Iterative Optimization Network for Broadband Mosaiced Hyperspectral Imaging
abstract
Snapshot hyperspectral imaging based on narrowband mosaic array encoding suffers from limitations such as low signal-to-noise ratio, limited spectral range, and low spatial resolution. To address these challenges, we propose a novel hyperspectral imaging system that integrates broadband mosaic image with high-resolution (HR) panchromatic (PAN) image of the same scene, establishing a new paradigm for HR hyperspectral image (HSI) acquisition. To fully leverage the complementary information from multi-source images, we introduce a Deep Error-aware Iterative Optimization Network (EIONet), which iteratively reduces reconstruction errors to successfully reconstruct images with both high spatial resolution and high spectral quality. Specifically, we design a Hierarchical Error-aware Cube Updating Mechanism (HECUM) that dynamically partitions image regions based on their reconstruction difficulty during iterations. By prioritizing the enhancement of feature representation in high-difficulty areas, it effectively suppresses the accumulation and propagation of errors. Meanwhile, we employ a Physics-Based Spectral Degradation Modeling approach, constructing a spectral response function with well-defined physical meaning to accurately model the degradation process from the target domain to the observation domain. Experimental results on two public datasets demonstrate that EIONet achieves state-of-the-art performance across multiple evaluation metrics. The related code is available at: https://github.com/Xiexieiii/EIONet.
Yunyu Xie, Renwei Dian, Lishan Tan, Shutao Li 0001
IEEE Trans. Image Process.3
2025 A Selective Re-learning Mechanism for Hyperspectral Fusion Imaging
abstract
Hyperspectral fusion imaging is challenged by high computational cost due to the abundant spectral information. We find that pixels in regions with smooth spatial-spectral structure can be reconstructed well using a shallow network, while only those in regions with complex spatial-spectral structure require a deeper network. However, existing methods process all pixels uniformly, which ignores this property. To leverage this property, we propose a Selective Re-Learning Fusion Network (SRLF) that initially extracts features from all pixels uniformly and then selectively refines distorted feature points. Specifically, SRLF first employs a Preliminary Fusion Module with robust global modeling capability to generate a preliminary fusion feature. Afterward, it applies a Selective Re-Learning Module to focus on improving distorted feature points in the preliminary fusion feature. To achieve targeted learning, we present a novel Spatial-Spectral Structure-Guided Selective Re-Learning Mechanism (SSG-SRL) that integrates the observation model to identify the feature points with spatial or spectral distortions. Only these distorted points are sent to the corresponding re-learning blocks, reducing both computational cost and the risk of overfitting. Finally, we develop an SRLF-Net, composed of multiple cascaded SRLFs, which surpasses multiple state-of-the-art methods on several datasets with minimal computational cost.
Yuanye Liu, Jinyang Liu 0004, Renwei Dian, Shutao Li 0001
CVPR3
2025 Low-Rank Transformer for High-Resolution Hyperspectral Computational Imaging
Yuanye Liu, Renwei Dian, Shutao Li 0001
Int. J. Comput. Vis.2
2025 Building Non-Uniform Degradation Model for Position-Aware Hyperspectral Image Fusion
abstract
The fusion of low-spatial-resolution hyperspectral image (LR-HSI) with high-spatial-resolution multispectral image (HR-MSI) has become an effective way to obtain the high-spatial-resolution hyperspectral image (HR-HSI). Currently, learning-based methods have emerged as the mainstream solution in this field. However, these methods typically rely on predefined or simplified degradation models during fusion training, resulting in inaccurate supervision of the fusion networks. Meanwhile, most methods overlook the degradation characteristics in designing the fusion networks, leading to a mismatch between the degradation and fusion processes. These limitations ultimately result in unsatisfactory fusion performance on real data. To enhance the practicality of learning-based methods, accurate degradation modeling and effective network design have become the critical priorities. We observe that, in practical scenarios, the degree of pixel degradation varies across different positions due to the unforeseen factors such as illumination variations and imaging system fluctuations. Considering this, we propose a non-uniform degradation model (NUD), which introduces non-uniformity into the degradation processes of LR-HSI and HR-MSI. In addition, we emphasize that the essence of fusion is to reverse the degradation process. Therefore, to align with the non-uniform degradation process, the fusion process should exhibit similar positional specificity. For this purpose, we propose a position-aware fusion network (PAF), which employs positional encoding to endow the fusion process with the position-aware attribute. Experimental results show that our proposed methods provide an effective solution for HSI fusion in practical scenarios.
Lizhi Wang 0001, Lin Zhu 0012, Renwei Dian, Zhiwei Xiong, Hua Huang 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2025 Asymptotic Spectral Mapping for Hyperspectral Image Fusion
abstract
The fusion of low-resolution hyperspectral images (LR HSI) and high-resolution multispectral images (HR MSI) is a crucial approach for generating hyperspectral images (HSI). However, existing hyperspectral image fusion methods often rely on a single feature mapping process, which makes it difficult to accommodate the significant differences in features between the source images and the real images. Consequently, the generated images frequently exhibit varying degrees of information loss across different spectral bands and limit the overall performance of the fusion. To address this issue, we propose a novel hyperspectral image fusion network. Specifically, an asymptotic spectral mapping module is designed to enhance the detail information fitting capabilities of the fusion network. This module transforms the fitting process of missing information into multiple sets of fitting processes with varying degrees, which can map features with different spectral fidelity to various scales and gradually fit spectral information, thereby reducing spectral distortion. Additionally, we introduce an adaptive defect optimization loss that guides the network to focus on reconstructing regions with substantial spectral differences between LR HSI and HSI, optimizing the network’s constraints regarding the similarity between predicted and real images. Experimental results demonstrate that the proposed fusion network outperforms existing state-of-the-art methods across diverse datasets.
Jinyang Liu 0004, Shutao Li 0001, Renwei Dian, Lishan Tan
IEEE Trans. Circuits Syst. Video Technol.3
2025 MCFNet: Multiscale Cross-Domain Fusion Network for HSI and LiDAR Data Joint Classification
abstract
Hyperspectral image (HSI) encompasses abundant spatial and spectral details, while Light Detection and Ranging (LiDAR) delivers precise elevation data. The amalgamation of HSI and LiDAR data significantly improves the precision of image classification. However, most methods focus solely on spatial features while neglecting frequency domain information, limiting the ability of deep models to characterize land cover. Furthermore, how to establish a sufficient interaction between different modalities is also an important issue. In this paper, we propose a novel multiscale cross-domain fusion network (MCFNet) for joint classification of HSI and LiDAR data. The main idea is that the wavelet transform can provide details at different resolutions simultaneously, supplementing spatial domain information and enriching feature representation. In addition, the multimodal fusion module (MFM) guided by HSI and the cross-domain fusion module (CDFM) strategy are developed to integrate features from diverse modalities and domains, respectively. Specifically, frequency domain features are extracted by discrete wavelet transform, and spatial domain features of the image are captured through a set of convolution operations. Then interactive fusion is performed by MFM and CDFM, and finally the integrated features are categorized using a classification module. Extensive experiments on three widely-used HSI and LiDAR datasets indicate that MCFNet outperforms the SOTA methods. The code will be available at https://github.com/MSFLabX/MCFNet.
Qiya Song, Feng Mo, Kexing Ding, Lin Xiao 0002, Renwei Dian, Xudong Kang, Shutao Li 0001
IEEE Trans. Geosci. Remote. Sens.5
2025 Detail-Aware Network for Infrared Image Enhancement
abstract
Infrared (IR) images inherently face the dual challenges of noise contamination and reduced contrast. However, existing image enhancement methods often overlook the intrinsic correlations between these factors—noise and low contrast—during multistage enhancement processes. Consequently, this oversight leads to a significant reduction in the fidelity of intricate details in IR images. In this article, we present a synergistic IR image enhancement network that simultaneously achieves denoising, contrast improvement, and detail preservation (DCDNet), which breaks down the overall enhancement process into more manageable steps. DCDNet is comprised of a detail awareness unit (DAU), a deep denoising prior (DDP), and a contrast improvement module (CIM). To maintain the details in the IR image, DAU is developed to extract the original detail feature information in DDP and integrate them into the CIM during contrast improvement to improve the final result. The detail information is derived from the encoder of the DDP, which focuses on denoising. The preserved detail features are subsequently incorporated into the decoder of the CIM, which is dedicated to enhancing contrast. Experimental results validate that our proposed approach surpasses other state-of-the-art methods for enhancing IR images in terms of the peak signal-to-noise ratio (PSNR), structural similarity index (SSIM), visual information fidelity (VIF), and performance in downstream tasks. The code and dataset are publicly available athttps://github.com/ChickenEating/IR-Enhancement.
Ruiheng Zhang 0001, Guanyu Liu, Qi Zhang 0004, Xiankai Lu, Renwei Dian, Yang Yang 0074, Lixin Xu 0001
IEEE Trans. Geosci. Remote. Sens.5
2025 Better Image Filter for Pansharpening
abstract
The modulation transfer function tailored image filter (MTF-TIF) has long been regarded as the optimal filter for multispectral image pansharpening. It excels at simulating the camera's frequency response, thereby capturing finer image details and significantly improving pansharpening performance. However, we are skeptical about whether the pre-measured MTF is sufficient to describe the characteristics of actually acquired panchromatic image (PAN) and multispectral image (MSI). For example, any image resampling operations in geometric correction or image registration inevitably change the sharpness of acquired PAN and MSI, and the processed images no longer conform to the camera's MTF. Further, following the Wald protocol, in deep learning (DL) methods using MTF-TIF for downsampling images to construct training data does not satisfy the generalization consistency of training and testing. To prove our point, we propose a pair of symmetric frameworks based on DL in this paper, to find better image filters suitable for both traditional and DL pansharpening methods. We embed two learnable filters into the frameworks to simulate the optimal image filter, namely anisotropic Gaussian image filter and arbitrary image filter. Further, the proposed frameworks can capture subtle offsets between images and maintain the smoothness of the global deformation field. Extensive experiments on various satellite datasets demonstrate that the proposed frameworks can find better image filters than MTF-TIFs, which can achieve better pansharpening performance with stronger generalization ability.
Anjing Guo, Renwei Dian, Shutao Li 0001
IEEE Trans. Image Process.2
2025 Heterospectral Structure Compensation Sampling for Hyperspectral Fusion Computational Imaging
abstract
Existing hyperspectral fusion computational imaging methods primarily rely on using high-resolution multispectral images (HRMSI) to provide spatial details for low-resolution hyperspectral images (LRHSI), thereby enabling the reconstruction of hyperspectral images. However, these methods are often limited by the low spectral resolution of the HRMSI, making the sampled tensors unable to provide effective information for the LRHSI in a finer spectral range. To achieve more accurate computational imaging results, we propose a Heterospectral Structure Compensation Sampling (HSC-sampling) mechanism. Unlike traditional spatial sampling methods, which directly calculate the interpolation between adjacent pixels, this mechanism analyzes the structural complementarity among different bands in LRHSI. It utilizes the information from other bands to compensate for the missing details in the current band. Additionally, a novel Multi-phase Mixed Modeling (M2M) approach is designed, expanding the model's analytical capabilities into multiple phases to accommodate the high-dimensional nature of HSI data. Specifically, it extracts fusion features from three phases and organizes the generated features along with the input features into a multi-variate mixed cube based on phase relationships, thereby capturing feature correlations across different phases. Based on the HSC-sampling mechanism and the M2M approach, we construct a Merging Residual Concatenation (MRC) hyperspectral fusion computational imaging network. Compared to other state-of-the-art methods, this network achieves significant improvements in fusion performance across multiple datasets. Moreover, the effectiveness of the HSC-sampling mechanism has been demonstrated in various hyperspectral imaging tasks. Code is available at: https://github.com/1318133/HSC-Sampling.
Jinyang Liu 0004, Shutao Li 0001, Renwei Dian, Yuanye Liu
IEEE Trans. Image Process.4
2025 Mosaic Pattern Excavation Transformer for Spectral Imaging
abstract
Single spectral image demosaicing for multispectral filter array (MSFA) is an essential task in spectral imaging, aiming to recover a mosaic-free spectral image from its mosaic raw counterpart. Existing deep learning-based methods typically improve the reconstruction performance by indiscriminately stacking CNN-based blocks, failing to effectively handle the intertwined spatio-spectral correlations caused by spatial sub-sampling and spectral aliasing. In this paper, we propose Mosaic Pattern Excavation Transformer (MPEFormer) to achieve better reconstruction by effectively modelling the intertwined spatio-spectral correlations. Specifically, the proposed three-branch model integrates low-frequency information, edge information, and fine high-frequency details essential for spectral image reconstruction, with the third branch serving as the core component. In this branch, we design the Dual Fusion Self-attention Block (DFSAB) and the Mosaic Pattern-guided Spectral Modulation Module (MPSM). DFSAB incorporates the Mosaic Pattern Excavation Self-attention (MPESA) mechanism, which effectively captures non-local spatio-spectral correlations induced by the MSFA pattern distributed across the whole image, thereby enhancing the expressive capability of the model. By dynamically integrating various MSFA pattern-related dependencies, MPSM enables adaptive recalibration of spectral information. Extensive experimental results demonstrate the effectiveness of our MPEFormer, highlighting its greater potential over the state-of-the-art MSFA demosaicing methods. The code will be uploaded at https://github.com/Matsuri247/MPEFormer.
Yaohang Wu, Jinyang Liu 0004, Renwei Dian, Shutao Li 0001, Yining Yang
IEEE Trans. Image Process.3
2025 Multi-Granularity Context Perception Network for Open Set Recognition of Camouflaged Objects
abstract
Open set recognition (OSR) aims to identify whether a test sample belongs to a semantic class in the classifier training set. Existing OSR methods exhibit prominent performance on various image datasets. However, they are primarily designed for general object recognition rather than more complex camouflaged object recognition. When an object is camouflaged, i.e., it exhibits a similar pattern to the background, it is difficult to finely identify it and differentiate between known and unknown categories. To address this problem, we propose a novel multi-granularity context perception network (MCPNet) for OSR of camouflaged objects, which can accurately identify camouflaged objects by fusing coarse-grained and fine-grained context features. In MCPNet, the vision transformer is first utilized to extract coarse-grained context features to locate the approximate location of camouflaged objects. Then, an adaptive local focus module (ALFM) is proposed to pick out the most discriminative regions and learn the fine-grained context of these regions. Finally, multi-granular context features are fused to obtain recognition results. During the training, a contrastive clustering module (CCM) is introduced to guide the network to effectively utilize multi-granularity context to generate high-confidence decision boundaries. We also built two camouflaged object classification datasets named ACOC and NCOC which mainly consist of artificial camouflage and natural camouflage respectively to facilitate research in OSR of camouflaged objects. Experimental results on two datasets show that MCPNet outperforms state-of-the art methods.
Xudong Kang, Xiaohui Wei 0001, Renwei Dian, Jinyang Liu 0004, Shutao Li 0001
IEEE Trans. Multim.4
2025 Spectral Super-Resolution via Deep Low-Rank Tensor Representation
abstract
Spectral super-resolution has attracted the attention of more researchers for obtaining hyperspectral images (HSIs) in a simpler and cheaper way. Although many convolutional neural network (CNN)-based approaches have yielded impressive results, most of them ignore the low-rank prior of HSIs resulting in huge computational and storage costs. In addition, the ability of CNN-based methods to capture the correlation of global information is limited by the receptive field. To surmount the problem, we design a novel low-rank tensor reconstruction network (LTRN) for spectral super-resolution. Specifically, we treat the features of HSIs as 3-D tensors with low-rank properties due to their spectral similarity and spatial sparsity. Then, we combine canonical-polyadic (CP) decomposition with neural networks to design an adaptive low-rank prior learning (ALPL) module that enables feature learning in a 1-D space. In this module, there are two core modules: the adaptive vector learning (AVL) module and the multidimensionwise multihead self-attention (MMSA) module. The AVL module is designed to compress an HSI into a 1-D space by using a vector to represent its information. The MMSA module is introduced to improve the ability to capture the long-range dependencies in the row, column, and spectral dimensions, respectively. Finally, our LTRN, mainly cascaded by several ALPL modules and feedforward networks (FFNs), achieves high-quality spectral super-resolution with fewer parameters. To test the effect of our method, we conduct experiments on two datasets: the CAVE dataset and the Harvard dataset. Experimental results show that our LTRN not only is as effective as state-of-the-art methods but also has fewer parameters. The code is available at https://github.com/renweidian/LTRN.
Renwei Dian, Yuanye Liu, Shutao Li 0001
IEEE Trans. Neural Networks Learn. Syst.1
2025 Hyperspectral Image Fusion via a Novel Generalized Tensor Nuclear Norm Regularization
abstract
Recently, low-rank tensor regularization has received more and more attention in hyperspectral and multispectral fusion (HMF). However, these methods often suffer from inflexible low-rank tensor definition and are highly sensitive to the permutation of tensor modes, which hinder their performance. To tackle this problem, we propose a novel generalized tensor nuclear norm (GTNN)-based approach for the HMF. First, we define a novel GTNN by extending the existing third-mode-based tensor nuclear norm (TNN) to arbitrary mode, which conducts the Fourier transform on an arbitrary single mode and then computes the TNN for each mode. In this way, we can not only capture more extensive correlations for the three modes of a tensor, and also omit the adverse effect of permutation of tensor modes. To utilize the correlations among spectral bands, the high-resolution hyperspectral image (HSI) is approximated as low-rank spectral basis multiplication by coefficients, and we estimate the spectral basis by conducting singular-value decomposition (SVD) on HSI. Then, the coefficients are estimated by addressing the proposed GTNN regularized optimization. In specific, to exploit the non-local similarities of the HSI, we first cluster the patches of the coefficient into a 3-D, which contains spatial, spectral, and non-local modes. Since the collected tensor contains the strong non-local spatial-spectral similarities of the HSI, the proposed low-rank tensor regularization is imposed on these collected tensors, which fully model the non-local self-similarities. Fusion experiments on both simulated and real datasets prove the advantages of this approach. The code is available at https://github.com/renweidian/GTNN.
Renwei Dian, Yuanye Liu, Shutao Li 0001
IEEE Trans. Neural Networks Learn. Syst.1
2025 Denoiser Learning for Infrared and Visible Image Fusion
abstract
Infrared image (IR) and visible image (VI) fusion creates fusion images that contain richer information and gain improved visual effects. Existing methods generally use the operators of manual design, such as intensity and gradient operators, to mine the image information. However, it is hard for them to achieve a complete and accurate description of information, which limits the image fusion performance. To this end, a novel information measurement method is proposed to achieve IR and VI fusion. Its core idea is to guide a generator in achieving image fusion by learning the denoisers. Specifically, by using denoisers to restore fusion images with different noise interference to source images, a mutual competition relationship is formed between denoisers, which helps the generator thoroughly explore the data specificity of the source images and guide it to achieve more accurate feature representation. In addition, a semantic adaptive measurement loss function is proposed to constrain the generator, which fuses semantic information adaptively by considering the semantic information density of different source images. The results of quantitative and qualitative experiments have shown that the proposed method can achieve a higher quality information fusion and has a faster fusion speed on three public datasets when compared with advanced methods.
Jinyang Liu 0004, Shutao Li 0001, Lishan Tan, Renwei Dian
IEEE Trans. Neural Networks Learn. Syst.4
2025 Multistage Spatial-Spectral Fusion Network for Spectral Super-Resolution
abstract
Spectral super-resolution (SSR) aims to restore a hyperspectral image (HSI) from a single RGB image, in which deep learning has shown impressive performance. However, the majority of the existing deep-learning-based SSR methods inadequately address the modeling of spatial-spectral features in HSI. That is to say, they only sufficiently capture either the spatial correlations or the spectral self-similarity, which results in a loss of discriminative spatial-spectral features and hence limits the fidelity of the reconstructed HSI. To solve this issue, we propose a novel SSR network dubbed multistage spatial-spectral fusion network (MSFN). From the perspective of network design, we build a multistage Unet-like architecture that differentially captures the multiscale features of HSI both spatialwisely and spectralwisely. It consists of two types of the self-attention mechanism, which enables the proposed network to achieve global modeling of HSI comprehensively. From the perspective of feature alignment, we innovatively design the spatial fusion module (SpatialFM) and spectral fusion module (SpectralFM), aiming to preserve the comprehensively captured spatial correlations and spectral self-similarity. In this manner, the multiscale features can be better fused and the accuracy of reconstructed HSI can be significantly enhanced. Quantitative and qualitative experiments on the two largest SSR datasets (i.e., NTIRE2022 and NTIRE2020) demonstrate that our MSFN outperforms the state-of-the-art SSR methods. The code implementation will be uploaded at https://github.com/Matsuri247/MSFN-for-Spectral-Super-Resolution.
Yaohang Wu, Renwei Dian, Shutao Li 0001
IEEE Trans. Neural Networks Learn. Syst.2
2024 Completing Saliency from Details
Jin Zhang 0021, Lingxiang Wu, Renwei Dian, Yiheng Yao, Shihao Huang, Yang Yang 0074, Ruiheng Zhang 0001
PRCV (13)4
2024 MDENet: Multidomain Differential Excavating Network for Remote Sensing Image Change Detection
abstract
Remote sensing image change detection can analyze alterations on the Earth’s surface within a specific region. However, the accuracy of change detection has consistently been hindered by the style differences in captured images caused by seasonal or lighting variations, as well as the challenge of distinguishing similar features between the background and foreground in the scene. To this end, a multidomain differential excavating network (MDENet) for change detection is introduced. Using the novel multidomain differential collaboration module (MDCM) to precisely capture object features on the frequency and spatial domains across diverse temporal domains, it enables simultaneous querying of global and local change information. Moreover, the multineighborhood frequency gate attention (MFGatt) is devised to eliminate the impact of image style relevance information and consolidate attention toward object localization, thereby enhancing the adaptability of the network to variations in image style. Extensive experiments have illustrated that our proposed network achieves better detection accuracy compared with current state-of-the-art (SOTA) methods on various datasets.
Jinyang Liu 0004, Shutao Li 0001, Renwei Dian, Xudong Kang
IEEE Trans. Geosci. Remote. Sens.3
2024 Click-Pixel Cognition Fusion Network With Balanced Cut for Interactive Image Segmentation
abstract
Interactive image segmentation (IIS) has been widely used in various fields, such as medicine, industry, etc. However, some core issues, such as pixel imbalance, remain unresolved so far. Different from existing methods based on pre-processing or post-processing, we analyze the cause of pixel imbalance in depth from the two perspectives of pixel number and pixel difficulty. Based on this, a novel and unified Click-pixel Cognition Fusion network with Balanced Cut (CCF-BC) is proposed in this paper. On the one hand, the Click-pixel Cognition Fusion (CCF) module, inspired by the human cognition mechanism, is designed to increase the number of click-related pixels (namely, positive pixels) being correctly segmented, where the click and visual information are fully fused by using a progressive three-tier interaction strategy. On the other hand, a general loss, Balanced Normalized Focal Loss (BNFL), is proposed. Its core is to use a group of control coefficients related to sample gradients and forces the network to pay more attention to positive and hard-to-segment pixels during training. As a result, BNFL always tends to obtain a balanced cut of positive and negative samples in the decision space. Theoretical analysis shows that the commonly used Focal and BCE losses can be regarded as special cases of BNFL. Experiment results of five well-recognized datasets have shown the superiority of the proposed CCF-BC method compared to other state-of-the-art methods. The source code is publicly available at https://github.com/lab206/CCF-BC.
Jiacheng Lin, Xiaohui Wei 0001, Puhong Duan, Renwei Dian, Zhiyong Li 0001, Shutao Li 0001
IEEE Trans. Image Process.6
2024 Focus Relationship Perception for Unsupervised Multi-Focus Image Fusion
abstract
Multi-focus image fusion can extract the focus regions from different source images and combine them into a fully clear image. Existing unsupervised methods typically use gradient information to measure the focus regions in images and generate a fusion weight map, but ordinary gradient operators are difficult to measure information accurately in regions with weaker textures. In addition, using only gradient information as a constraint cannot make the model fully distinguish all the focus regions in the image, which seriously restricts the clarity of the fusion image. To address these issues, a novel unsupervised multi-focus image fusion method is proposed in this paper. Specifically, a neighborhood information fusion network is designed to generate an initial fusion weight map. It can capture features within different neighborhood ranges at once, which enhances the information association between different regions. In addition, to further improve the feature extraction ability of the model in the regions with low texture information, a local difference evaluation loss function is proposed. It is combined with the gradient measure loss function to constrain the network. Finally, a fusion weight optimization module is proposed to improve the clarity of the fusion image in the repeated defocusing regions and overexposed regions of different source images, which redistributes the weights of different source images. The proposed fusion method is compared with advanced methods on three public multi-focus datasets. Experimental results indicate that the proposed method has achieved better performance in qualitative and quantitative aspects.
Jinyang Liu 0004, Shutao Li 0001, Renwei Dian
IEEE Trans. Multim.3
2024 Spectral Super-Resolution via Model-Guided Cross-Fusion Network
abstract
Spectral super-resolution, which reconstructs a hyperspectral image (HSI) from a single red-green-blue (RGB) image, has acquired more and more attention. Recently, convolution neural networks (CNNs) have achieved promising performance. However, they often fail to simultaneously exploit the imaging model of the spectral super-resolution and complex spatial and spectral characteristics of the HSI. To tackle the above problems, we build a novel cross fusion (CF)-based model-guided network (called SSRNet) for spectral super-resolution. In specific, based on the imaging model, we unfold the spectral super-resolution into the HSI prior learning (HPL) module and imaging model guiding (IMG) module. Instead of just modeling one kind of image prior, the HPL module is composed of two subnetworks with different structures, which can effectively learn the complex spatial and spectral priors of the HSI, respectively. Furthermore, a CF strategy is used to establish the connection between the two subnetworks, which further improves the learning performance of the CNN. The IMG module results in solving a strong convex optimization problem, which adaptively optimizes and merges the two features learned by the HPL module by exploiting the imaging model. The two modules are alternately connected to achieve optimal HSI reconstruction performance. Experiments on both the simulated and real data demonstrate that the proposed method can achieve superior spectral reconstruction results with relatively small model size. The code will be available at https://github.com/renweidian.
Renwei Dian, Tianci Shan, Wei He 0003
IEEE Trans. Neural Networks Learn. Syst.1
2024 LRAF-Net: Long-Range Attention Fusion Network for Visible-Infrared Object Detection
abstract
Visible-infrared object detection aims to improve the detector performance by fusing the complementarity of visible and infrared images. However, most existing methods only use local intramodality information to enhance the feature representation while ignoring the efficient latent interaction of long-range dependence between different modalities, which leads to unsatisfactory detection performance under complex scenes. To solve these problems, we propose a feature-enhanced long-range attention fusion network (LRAF-Net), which improves detection performance by fusing the long-range dependence of the enhanced visible and infrared features. First, a two-stream CSPDarknet53 network is used to extract the deep features from visible and infrared images, in which a novel data augmentation (DA) method is designed to reduce the bias toward a single modality through asymmetric complementary masks. Then, we propose a cross-feature enhancement (CFE) module to improve the intramodality feature representation by exploiting the discrepancy between visible and infrared images. Next, we propose a long-range dependence fusion (LDF) module to fuse the enhanced features by associating the positional encoding of multimodality features. Finally, the fused features are fed into a detection head to obtain the final detection results. Experiments on several public datasets, i.e., VEDAI, FLIR, and LLVIP, show that the proposed method obtains state-of-the-art performance compared with other methods.
Haolong Fu, Shixun Wang, Puhong Duan, Changyan Xiao, Renwei Dian, Shutao Li 0001, Zhiyong Li 0001
IEEE Trans. Neural Networks Learn. Syst.5
2024 SSTF-Unet: Spatial-Spectral Transformer-Based U-Net for High-Resolution Hyperspectral Image Acquisition
abstract
To obtain a high-resolution hyperspectral image (HR-HSI), fusing a low-resolution hyperspectral image (LR-HSI) and a high-resolution multispectral image (HR-MSI) is a prominent approach. Numerous approaches based on convolutional neural networks (CNNs) have been presented for hyperspectral image (HSI) and multispectral image (MSI) fusion. Nevertheless, these CNN-based methods may ignore the global relevant features from the input image due to the geometric limitations of convolutional kernels. To obtain more accurate fusion results, we provide a spatial-spectral transformer-based U-net (SSTF-Unet). Our SSTF-Unet can capture the association between distant features and explore the intrinsic information of images. More specifically, we use the spatial transformer block (SATB) and spectral transformer block (SETB) to calculate the spatial and spectral self-attention, respectively. Then, SATB and SETB are connected in parallel to form the spatial-spectral fusion block (SSFB). Inspired by the U-net architecture, we build up our SSTF-Unet through stacking several SSFBs for multiscale spatial-spectral feature fusion. Experimental results on public HSI datasets demonstrate that the designed SSTF-Unet achieves better performance than other existing HSI and MSI fusion approaches.
Chenguo Feng, Renwei Dian, Shutao Li 0001
IEEE Trans. Neural Networks Learn. Syst.3
2024 A Lightweight Pixel-Level Unified Image Fusion Network
abstract
In recent years, deep-learning-based pixel-level unified image fusion methods have received more and more attention due to their practicality and robustness. However, they usually require a complex network to achieve more effective fusion, leading to high computational cost. To achieve more efficient and accurate image fusion, a lightweight pixel-level unified image fusion (L-PUIF) network is proposed. Specifically, the information refinement and measurement process are used to extract the gradient and intensity information and enhance the feature extraction capability of the network. In addition, these information are converted into weights to guide the loss function adaptively. Thus, more effective image fusion can be achieved while ensuring the lightweight of the network. Extensive experiments have been conducted on four public image fusion datasets across multimodal fusion, multifocus fusion, and multiexposure fusion. Experimental results show that L-PUIF can achieve better fusion efficiency and has a greater visual effect compared with state-of-the-art methods. In addition, the practicability of L-PUIF in high-level computer vision tasks, i.e., object detection and image segmentation, has been verified.
Jinyang Liu 0004, Shutao Li 0001, Renwei Dian, Xiaohui Wei 0001
IEEE Trans. Neural Networks Learn. Syst.4
2023 Multi-scale Conformer Fusion Network for Multi-participant Behavior Analysis
abstract
Understanding and elucidating human behavior across diverse scenarios represents a pivotal research challenge in pursuing seamless human-computer interaction. However, previous research on multi-participant dialogues has mostly relied on proprietary datasets, which are not standardized and openly accessible. To propel advancements in this domain, the MultiMediate'23 Challenge presents two sub-challenges: Eye contact detection and Next speaker prediction, aiming to foster a comprehensive understanding of multi-participant behavior. To tackle these challenges, we propose a multi-scale conformer fusion network (MSCFN) for enhancing the perception of multi-participant group behaviors. The conformer block combines the strengths of transformers and convolution networks to facilitate the establishment of global and local contextual relationships between sequences. Then the output features from all Conformer blocks are concatenated to fusion multi-scale representations. Our proposed method was evaluated using the officially provided dataset, and it achieves the best and second best performance in next speaker prediction and gaze detection tasks of MultiMediate'23, respectively.
Qiya Song, Renwei Dian, Bin Sun 0001, Jie Xie 0002, Shutao Li 0001
ACM Multimedia2
2023 Learning the external and internal priors for multispectral and hyperspectral image fusion
Shutao Li 0001, Renwei Dian
Sci. China Inf. Sci.2
2023 Zero-Shot Hyperspectral Sharpening
abstract
Fusing hyperspectral images (HSIs) with multispectral images (MSIs) of higher spatial resolution has become an effective way to sharpen HSIs. Recently, deep convolutional neural networks (CNNs) have achieved promising fusion performance. However, these methods often suffer from the lack of training data and limited generalization ability. To address the above problems, we present a zero-shot learning (ZSL) method for HSI sharpening. Specifically, we first propose a novel method to quantitatively estimate the spectral and spatial responses of imaging sensors with high accuracy. In the training procedure, we spatially subsample the MSI and HSI based on the estimated spatial response and use the downsampled HSI and MSI to infer the original HSI. In this way, we can not only exploit the inherent information in the HSI and MSI, but the trained CNN can also be well generalized to the test data. In addition, we take the dimension reduction on the HSI, which reduces the model size and storage usage without sacrificing fusion accuracy. Furthermore, we design an imaging model-based loss function for CNN, which further boosts the fusion performance. The experimental results show the significantly high efficiency and accuracy of our approach.
Renwei Dian, Anjing Guo, Shutao Li 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2023 A Deep Framework for Hyperspectral Image Fusion Between Different Satellites
abstract
Recently, fusing a low-resolution hyperspectral image (LR-HSI) with a high-resolution multispectral image (HR-MSI) of different satellites has become an effective way to improve the resolution of an HSI. However, due to different imaging satellites, different illumination, and adjacent imaging time, the LR-HSI and HR-MSI may not satisfy the observation models established by existing works, and the LR-HSI and HR-MSI are hard to be registered. To solve the above problems, we establish new observation models for LR-HSIs and HR-MSIs from different satellites, then a deep-learning-based framework is proposed to solve the key steps in multi-satellite HSI fusion, including image registration, blur kernel learning, and image fusion. Specifically, we first construct a convolutional neural network (CNN), called RegNet, to produce pixel-wise offsets between LR-HSI and HR-MSI, which are utilized to register the LR-HSI. Next, according to the new observation models, a tiny network, called BKLNet, is built to learn the spectral and spatial blur kernels, where the BKLNet and RegNet can be trained jointly. In the fusion part, we further train a FusNet by downsampling the registered data with the learned spatial blur kernel. Extensive experiments demonstrate the superiority of the proposed framework in HSI registration and fusion accuracy.
Anjing Guo, Renwei Dian, Shutao Li 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 Multispectral Image Pan-Sharpening Guided by Component Substitution Model
abstract
Multispectral image pan-sharpening aims to increase the spatial details of multispectral images by fusing multispectral and panchromatic images. Existing component substitution-based deep learning pan-sharpening is generally regarded as a black box and fails to mine the image interaction relation with physical significance in each step of pan-sharpening, which not only limits the improvement of image resolution, but also ignores the physical interpretability of the models. To improve this situation, according to the traditional component substitution-based detail injection pan-sharpening model, we consider the matrix calculation in each step as the transformation between image pixel values and carry out linear transformations, and therefore the pan-sharpened multispectral image is represented as the sum of two multispectral images. Then given the spatial and spectral heterogeneity, the two summed images are decomposed based on the fact that any real number can be expressed as the product of two real numbers. Ultimately, the multispectral image pan-sharpening model can be constructed as the sum of two Hadamard products. We design a dual-branch network with attention mechanisms that merges the sum and the Hadamard products into a concise formulation. This method not only enhances physical interpretability but also improves spatial resolution. Experiments on five real-world datasets validate that the proposed multispectral image pan-sharpening model can improve performance.
Huiling Gao, Shutao Li 0001, Jun Li 0009, Renwei Dian
IEEE Trans. Geosci. Remote. Sens.4
2023 FSNet: Focus Scanning Network for Camouflaged Object Detection
abstract
Camouflaged object detection (COD) aims to discover objects that blend in with the background due to similar colors or textures, etc. Existing deep learning methods do not systematically illustrate the key tasks in COD, which seriously hinders the improvement of its performance. In this paper, we introduce the concept of focus areas that represent some regions containing discernable colors or textures, and develop a two-stage focus scanning network for camouflaged object detection. Specifically, a novel encoder-decoder module is first designed to determine a region where the focus areas may appear. In this process, a multi-layer Swin transformer is deployed to encode global context information between the object and the background, and a novel cross-connection decoder is proposed to fuse cross-layer textures or semantics. Then, we utilize the multi-scale dilated convolution to obtain discriminative features with different scales in focus areas. Meanwhile, the dynamic difficulty aware loss is designed to guide the network paying more attention to structural details. Extensive experimental results on the benchmarks, including CAMO, CHAMELEON, COD10K, and NC4K, illustrate that the proposed method performs favorably against other state-of-the-art methods.
Xudong Kang, Xiaohui Wei 0001, Renwei Dian, Shutao Li 0001
IEEE Trans. Image Process.5
2022 Hyperspectral and Multispectral Image Fusion Via Self-Supervised Loss and Separable Loss
abstract
Fusion of hyperspectral images with low-spatial and high-spectral resolution and multispectral images with high-spatial and low-spectral resolution is an important method to improve spatial resolution. Existing deep learning-based image fusion technologies usually neglect the ability of neural networks to understand differential features. In addition, the loss constraints do not stem from the physical characteristics of the hyperspectral imaging sensors. We propose the self-supervised loss and the spatially and spectrally separable loss, respectively. 1) The self-supervised loss: Different from the previous way of directly stacking the upsampled hyperspectral images and multispectral images as input, we expect the potentially processed hyperspectral images to ensure not only the integrity of hyperspectral image information, but also the most reasonable balance between overall spatial and spectral features. Firstly, the pre-interpolated hyperspectral images are decomposed into subspaces as self-supervised labels. Then, a network is designed to learn subspace information and obtain the most discriminative features. 2) The separable loss: According to the physical characteristics of hyperspectral images, the pixel-based mean square error loss is first divided into the domain loss and spectral domain loss, and then the similarity score of the images is calculated and used to construct the weighting coefficients of the two domain losses. At last, the separable loss is jointly expressed by the weights. Experiments on public benchmark datasets indicate that the self-supervised loss and separable loss can improve fusion performance.
Huiling Gao, Shutao Li 0001, Renwei Dian
IEEE Trans. Geosci. Remote. Sens.3
2022 Asymmetric Hash Code Learning for Remote Sensing Image Retrieval
abstract
Remote sensing image retrieval (RSIR), aiming at searching for a set of similar items to a given query image, is a very important task in remote sensing applications. Deep hashing learning as the current mainstream method has achieved satisfactory retrieval performance. On one hand, various deep neural networks are used to extract semantic features of remote sensing images. On the other hand, the hashing techniques are subsequently adopted to map the high-dimensional deep features to the low-dimensional binary codes. This kind of method attempts to learn one hash function for both the query and database samples in a symmetric way. However, with the number of database samples increasing, it is typically time-consuming to generate the hash codes of large-scale database images. In this article, we propose a novel deep hashing method, named asymmetric hash code learning (AHCL), for RSIR. The proposed AHCL generates the hash codes of query and database images in an asymmetric way. In more detail, the hash codes of query images are obtained by binarizing the output of the network, while the hash codes of database images are directly learned by solving the designed objective function. In addition, we combine the semantic information of each image and the similarity information of pairs of images as supervised information to train a deep hashing network, which improves the representation ability of deep features and hash codes. The experimental results on three public datasets demonstrate that the proposed method outperforms symmetric methods in terms of retrieval accuracy and efficiency. The source code is available athttps://github.com/weiweisong415/Demo_AHCL_for_TGRS2022.
Zhi Gao 0005, Renwei Dian, Pedram Ghamisi, Yongjun Zhang 0002, Jón Atli Benediktsson
IEEE Trans. Geosci. Remote. Sens.3
2021 Regularizing Hyperspectral and Multispectral Image Fusion by CNN Denoiser
abstract
Hyperspectral image (HSI) and multispectral image (MSI) fusion, which fuses a low-spatial-resolution HSI (LR-HSI) with a higher resolution multispectral image (MSI), has become a common scheme to obtain high-resolution HSI (HR-HSI). This article presents a novel HSI and MSI fusion method (called as CNN-Fus), which is based on the subspace representation and convolutional neural network (CNN) denoiser, i.e., a well-trained CNN for gray image denoising. Our method only needs to train the CNN on the more accessible gray images and can be directly used for any HSI and MSI data sets without retraining. First, to exploit the high correlations among the spectral bands, we approximate the desired HR-HSI with the low-dimensional subspace multiplied by the coefficients, which can not only speed up the algorithm but also lead to more accurate recovery. Since the spectral information mainly exists in the LR-HSI, we learn the subspace from it via singular value decomposition. Due to the powerful learning performance and high speed of CNN, we use the well-trained CNN for gray image denoising to regularize the estimation of coefficients. Specifically, we plug the CNN denoiser into the alternating direction method of multipliers (ADMM) algorithm to estimate the coefficients. Experiments demonstrate that our method has superior performance over the state-of-the-art fusion methods.
Renwei Dian, Shutao Li 0001, Xudong Kang
IEEE Trans. Neural Networks Learn. Syst.1
2020 Unsupervised Blur Kernel Learning for Pansharpening
abstract
Deep learning (DL) for pansharpening has recently attracted considerable attentions. To construct training data, DL based pansharpening approaches often downsample the original multispectral image (MSI) and panchromatic image (PAN) with fixed blur kernel, which can be different from the real point spread functions (PSF) of the satellites. And a mismatched blur kernel will cause the pansharpening performance to drop dramatically. In this paper, we propose a novel blur kernel learning method for pansharpening, which can learn the spatial and spectral blur kernels between PAN and MSI in an unsupervised way. Specifically, we analyze the relationship between PAN and MSI, and then construct a mini net for blur kernel learning. Once the spatial blur kernel is found, a convolutional neural network (CNN) for pansharpening is trained on the downsampled dataset using the learned spatial blur kernel. Experimental results on GF-2 images demonstrate the superiority of the proposed method.
Anjing Guo, Renwei Dian, Shutao Li 0001
IGARSS2
2020 Nonlocal Sparse Tensor Factorization for Semiblind Hyperspectral and Multispectral Image Fusion
abstract
Combining a high-spatial-resolution multispectral image (HR-MSI) with a low-spatial-resolution hyperspectral image (LR-HSI) has become a common way to enhance the spatial resolution of the HSI. The existing state-of-the-art LR-HSI and HR-MSI fusion methods are mostly based on the matrix factorization, where the matrix data representation may be hard to fully make use of the inherent structures of 3-D HSI. We propose a nonlocal sparse tensor factorization approach, called the NLSTF_SMBF, for the semiblind fusion of HSI and MSI. The proposed method decomposes the HSI into smaller full-band patches (FBPs), which, in turn, are factored as dictionaries of the three HSI modes and a sparse core tensor. This decomposition allows to solve the fusion problem as estimating a sparse core tensor and three dictionaries for each FBP. Similar FBPs are clustered together, and they are assumed to share the same dictionaries to make use of the nonlocal self-similarities of the HSI. For each group, we learn the dictionaries from the observed HR-MSI and LR-HSI. The corresponding sparse core tensor of each FBP is computed via tensor sparse coding. Two distinctive features of NLSTF_SMBF are that: 1) it is blind with respect to the point spread function (PSF) of the hyperspectral sensor and 2) it copes with spatially variant PSFs. The experimental results provide the evidence of the advantages of the NLSTF_SMBF method over the existing state-of-the-art methods, namely, in semiblind scenarios.
Renwei Dian, Shutao Li 0001, Leyuan Fang, Ting Lu 0002, José M. Bioucas-Dias
IEEE Trans. Cybern.1
2019 Hyperspectral and Multispectral Image Fusion Based on Spectral Low Rank and Non-Local Spatial Similarities
abstract
Fusing a hyperspectral image (HSI) with a multispectral image (MSI) of the same scene has become a popular way to increase the spatial resolution of HSI. In this paper, we propose a novel HSI and MSI fusion method (termed as the SSS), which is based on spectral low rank and non-local spatial similarities. Firstly, to exploit the high spectral correlations of the desired high spatial resolution HSI, we formulate the fusion problem as the estimation of low-dimensional spectral subspace and coefficients. Since the HSI preserves most of spectral information, the spectral subspace is estimated from HSI via singular value decomposition. With the spectral subspace known, we plug a state-of-the-art denoising algorithm, weighted nuclear norm minimization, into the alternating direction method of multipliers to estimate the coefficients, which can effectively promote the non-local similarities of desired high spatial resolution HSI. Experiments demonstrate that our method is competitive to the state-of-the-art approaches.
Renwei Dian, Shutao Li 0001
IGARSS1
2019 Hyperspectral Image Super-Resolution via Subspace-Based Low Tensor Multi-Rank Regularization
abstract
Recently, combining a low spatial resolution hyperspectral image (LR-HSI) with a high spatial resolution multispectral image (HR-MSI) into an HR-HSI has become a popular scheme to enhance the spatial resolution of HSI. We propose a novel subspace-based low tensor multi-rank regularization method for the fusion, which fully exploits the spectral correlations and non-local similarities in the HR-HSI. To make use of high spectral correlations, the HR-HSI is approximated by spectral subspace and coefficients. We first learn the spectral subspace from the LR-HSI via singular value decomposition, and then estimate the coefficients via the low tensor multi-rank prior. More specifically, based on the learned cluster structure in the HR-MSI, the patches in coefficients are grouped. We collect the coefficients in the same cluster into a three-dimensional tensor and impose the low tensor multi-rank prior on these collected tensors, which fully model the non-local self-similarities in the HR-HSI. The coefficients optimization is solved by the alternating direction method of multipliers. Experiments on two public HSI datasets demonstrate the advantages of tour method.
Renwei Dian, Shutao Li 0001
IEEE Trans. Image Process.1
2019 Learning a Low Tensor-Train Rank Representation for Hyperspectral Image Super-Resolution
abstract
Hyperspectral images (HSIs) with high spectral resolution only have the low spatial resolution. On the contrary, multispectral images (MSIs) with much lower spectral resolution can be obtained with higher spatial resolution. Therefore, fusing the high-spatial-resolution MSI (HR-MSI) with low-spatial-resolution HSI of the same scene has become the very popular HSI super-resolution scheme. In this paper, a novel low tensor-train (TT) rank (LTTR)-based HSI super-resolution method is proposed, where an LTTR prior is designed to learn the correlations among the spatial, spectral, and nonlocal modes of the nonlocal similar high-spatial-resolution HSI (HR-HSI) cubes. First, we cluster the HR-MSI cubes as many groups based on their similarities, and the HR-HSI cubes are also clustered according to the learned cluster structure in the HR-MSI cubes. The HR-HSI cubes in each group are much similar to each other and can constitute a 4-D tensor, whose four modes are highly correlated. Therefore, we impose the LTTR constraint on these 4-D tensors, which can effectively learn the correlations among the spatial, spectral, and nonlocal modes because of the well-balanced matricization scheme of TT rank. We formulate the super-resolution problem as TT rank regularized optimization problem, which is solved via the scheme of alternating direction method of multipliers. Experiments on HSI data sets indicate the effectiveness of the LTTR-based method.
Renwei Dian, Shutao Li 0001, Leyuan Fang
IEEE Trans. Neural Networks Learn. Syst.1
2018 Hyperspectral Image Super-Resolution via Local Low-Rank and Sparse Representations
abstract
Remotely sensed hyperspectral images (HSIs) usually have high spectral resolution but low spatial resolution. A way to increase the spatial resolution of HSIs is to solve a fusion inverse problem, which fuses a low spatial resolution HSI (LR-HSI) with a high spatial resolution multispectral image (HR-MSI) of the same scene. In this paper, we propose a novel HSI super-resolution approach (called LRSR), which formulates the fusion problem as the estimation of a spectral dictionary from the LR-HSI and the respective regression coefficients from both images. The regression coefficients are estimated by formulating a variational regularization problem which promotes local (in the spatial sense) low-rank and sparse regression coefficients. The local regions, where the spectral vectors are low-rank, are estimated by segmenting the HR-MSI. The formulated convex optimization is solved with SALSA. Experiments provide evidence that LRSR is competitive with respect to the state-of-the-art methods.
Renwei Dian, Shutao Li 0001, Leyuan Fang, José M. Bioucas-Dias
IGARSS1
2018 Fusing Hyperspectral and Multispectral Images via Coupled Sparse Tensor Factorization
abstract
Fusing a low spatial resolution hyperspectral image (LR-HSI) with a high spatial resolution multispectral image (HR-MSI) to obtain a high spatial resolution hyperspectral image (HR-HSI) has attracted increasing interest in recent years. In this paper, we propose a coupled sparse tensor factorization (CSTF) based approach for fusing such images. In the proposed CSTF method, we consider an HR-HSI as a three-dimensional tensor and redefine the fusion problem as the estimation of a core tensor and dictionaries of the three modes. The high spatial-spectral correlations in the HR-HSI are modeled by incorporating a regularizer which promotes sparse core tensors. The estimation of the dictionaries and the core tensor are formulated as a coupled tensor factorization of the LR-HSI and of the HR-MSI. Experiments on two remotely sensed HSIs demonstrate the superiority of the proposed CSTF algorithm over current state-of-the-art HSI-MSI fusion approaches.
Shutao Li 0001, Renwei Dian, Leyuan Fang, José M. Bioucas-Dias
IEEE Trans. Image Process.2
2018 Deep Hyperspectral Image Sharpening
abstract
Hyperspectral image (HSI) sharpening, which aims at fusing an observable low spatial resolution (LR) HSI (LR-HSI) with a high spatial resolution (HR) multispectral image (HR-MSI) of the same scene to acquire an HR-HSI, has recently attracted much attention. Most of the recent HSI sharpening approaches are based on image priors modeling, which are usually sensitive to the parameters selection and time-consuming. This paper presents a deep HSI sharpening method (named DHSIS) for the fusion of an LR-HSI with an HR-MSI, which directly learns the image priors via deep convolutional neural network-based residual learning. The DHSIS method incorporates the learned deep priors into the LR-HSI and HR-MSI fusion framework. Specifically, we first initialize the HR-HSI from the fusion framework via solving a Sylvester equation. Then, we map the initialized HR-HSI to the reference HR-HSI via deep residual learning to learn the image priors. Finally, the learned image priors are returned to the fusion framework to reconstruct the final HR-HSI. Experimental results demonstrate the superiority of the DHSIS approach over existing state-of-the-art HSI sharpening approaches in terms of reconstruction accuracy and running time.
Renwei Dian, Shutao Li 0001, Anjing Guo, Leyuan Fang
IEEE Trans. Neural Networks Learn. Syst.1
2017 Hyperspectral Image Super-Resolution via Non-local Sparse Tensor Factorization
abstract
Hyperspectral image (HSI) super-resolution, which fuses a low-resolution (LR) HSI with a high-resolution (HR) multispectral image (MSI), has recently attracted much attention. Most of the current HSI super-resolution approaches are based on matrix factorization, which unfolds the three-dimensional HSI as a matrix before processing. In general, the matrix data representation obtained after the matrix unfolding operation makes it hard to fully exploit the inherent HSI spatial-spectral structures. In this paper, a novel HSI super-resolution method based on non-local sparse tensor factorization (called as the NLSTF) is proposed. The sparse tensor factorization can directly decompose each cube of the HSI as a sparse core tensor and dictionaries of three modes, which reformulates the HSI super-resolution problem as the estimation of sparse core tensor and dictionaries for each cube. To further exploit the non-local spatial self-similarities of the HSI, similar cubes are grouped together, and they are assumed to share the same dictionaries. The dictionaries are learned from the LR-HSI and HR-MSI for each group, and corresponding sparse core tensors are estimated by spare coding on the learned dictionaries for each cube. Experimental results demonstrate the superiority of the proposed NLSTF approach over several state-of-the-art HSI super-resolution approaches.
Renwei Dian, Leyuan Fang, Shutao Li 0001
CVPR1
2016 Non-local sparse representation for hyperspectral image super-resolution
abstract
In this paper, a non-local based sparse representation (called as the NLSR) is proposed for the super-resolution of hyperspectral image. Specifically, the NLSR firstly uses the non-local Kmeans to partition pixels of low spatial resolution hyperspectral image into several classes. The non-local Kmeans can exploit the similar patterns and structures of the low spatial resolution image to enhance the efficiency of the sparse solution. Then, the sparse representation is independently applied on each class of the low resolution hyperspectral image and high spatial resolution multispectral image to obtain the high resolution hyperspectral image. Experimental results demonstrate the superiority of the proposed NLSR method over several well-known super-resolution methods.
Renwei Dian, Shutao Li 0001, Leyuan Fang
ICIP1