Jingxiang Yang

dblp:150/2286 · DBLP profile ↗
← Back
35ranked-venue papers
12as first author
24since 2021 · last 2026
0000-0002-1234-0614ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 28 · 9 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 first-author
YearPublicationVenuePosition
2026 D&D-Net: A diffusion and deep priors regularized network for hyperspectral reconstruction
Jingxiang Yang, Tian Lin 0001, Wenxiu Diao, Fang Liu 0034, Jia Liu 0020, Hongyi Liu 0001, Liang Xiao 0001
Signal Process.1
2025 Multi-scale Feature Interaction and Adaptive Experts for Panoptic Segmentation in Remote Sensing Images
abstract
Panoptic segmentation unifies the traditional tasks of instance and semantic segmentation. It plays a crucial role in the field of remote sensing; however, it encounters challenges in recognizing small objects and in the model’s ability to generalize across complex scenes. In this paper, we introduce the MFIAE framework to address two specific challenges: a multi-scale interactive attention fusion (MSIAF) module and an adaptive disturbance sparse mixture-of-experts (ADSMoE) module based on Transformer. The MSIAF module is designed to fully utilize the rich contextual information captured by low-resolution features while simultaneously utilizing the advantages of high-resolution features to enhance small object segmentation. In the ADSMoE module, adaptive noise is introduced to disturb the expert selection in order to enhance the randomness and exploration of the model when it comes to selecting experts, thereby improving its capacity for generalization and robustness. Additionally, it also reduces the computational overhead and the model complexity. The experimental results demonstrate that our approach achieves state-of-the-art performance on the BSB Aerial dataset.
Zhenkun Sun, Jia Liu 0020, Jingxiang Yang, Liang Xiao 0001
ICASSP5
2025 Mask-guided Multi-scale Spatial-Spectral Transformer for Snapshot Compressive Imaging
abstract
Effectively reconstructing 3D hyperspectral images (HSIs) from 2D measurements presents a significant challenge in Coded Aperture Snapshot Spectral Imaging (CASSI) systems. While recent transformers exhibit potential in HSI reconstruction, they often suffer from inadequate exploration of multi-scale spatial-spectral self-similarity, leading to mean effects and information loss. Additionally, these methods struggle with insufficient modeling of the degradation inherent in the compressive imaging process. To address these issues, we propose a novel Mask-guided Multi-scale Spatial-Spectral Transformer (MMSST). Specifically, we introduce a Degradation Aware Mask Attention (DAMA) module to incorporate degradation information of the compressive imaging process. Furthermore, MMSST leverages Local-Regional SpAtial attention (LRSA) and Global-Regional SpEctral attention (GRSE) to effectively exploit multi-scale self-similarity across spatial and spectral dimensions. Extensive experimental results demonstrate the effectiveness of our MMSST.
Heyuan Yin, Jingxiang Yang, Jia Liu 0020, Liang Xiao 0001
ICASSP2
2025 Deep one-class probability learning for end-to-end image classification
Jia Liu 0020, Jingxiang Yang, Liang Xiao 0001
Neural Networks4
2025 Box2Change: A Novel Weakly Supervised Way for Change Detection via Consistency Instance Segmentation
abstract
Change detection in remote sensing images aims at revealing interesting changes about the earth surface and has been one of the most important issues in earth observation. In recent years, lots of fully-supervised change detection methods have achieved good performance with the help of deep learning architectures, which rely on large amounts of pixel-level labels. However, obtaining high-quality pixel-level labels is laborious and expensive. To alleviate this problem, we propose a novel weakly-supervised change detection way via consistency instance segmentation called Box2Change, which requires only box-level labels and achieves competitive results to fully-supervised change detection method. Compared with pixel-level label, it is much more efficient to get box-level label, which locates the potential changed area by a rectangle box. There are two key components in the proposed method, the Changed Instance Segmentation (CIS) and the Self-Supervised Consistency Learning (SSCL) in affine space. The former generates multi-scale changed instances, which learns positional information from box-level labels and segments the instance boundaries within a given bounded region. The latter introduces affine transform and employs consistency constraints in a self-supervised manner to increases the robustness to pseudo-change situations caused by light or noise. In experiments, three popular public change detection datasets are tested and both visual and numerical assessment are discussed, where the proposed method exhibits competitive performance to fully-supervised methods and achieves the state-of-the-art results compared with the other weakly-supervised change detection methods.
Fang Liu 0034, Kanghua Yin, Jia Liu 0020, Jingxiang Yang, Xu Tang 0004, Liang Xiao 0001
IEEE Trans. Geosci. Remote. Sens.4
2025 Background-Driven and Foreground-Refined Network for Weakly Supervised Change Detection
abstract
Change detection (CD) in remote sensing aims to reveal meaningful surface changes and has been flourishing in recent years. Compared with fully-supervised methods based on pixel-level labels, image-level labels are easy to acquire, which reduces manual labor to a large extent. However, image-level labels lack spatial-and-shape information while containing the least semantic information, which poses a great challenge to the weakly-supervised CD task. Motivated by the prior that bi-temporal images have background semantic consistency, we propose Background-Driven and Foreground-Refined (BDFR-Net) to ameliorate the above problem. Specifically, there are two key components in the proposed method: the Background-Driven Reconstruction (BDR) with image-level supervision and the Foreground-Refined Learning (FRL) with affinity learning. The former generates changed regions of foreground and background separation, which activates the foreground from image-level supervision and constrains the foreground by maintaining spatial and semantic consistency in background regions. The latter introduces Complementary Fusion and Label Adaption (CFLA) strategies to further refine the foreground, which can mine complementary information from foreground sequences and suppress false activations. In addition, affinity learning is proposed to stabilize and supervise the above process. Complementary relationships between foreground and background are fully utilized. Tested on two popular CD datasets, the results demonstrate that our proposed BDFR-Net produces completely changed regions with clear boundaries and outperforms state-of-the-art weakly-supervised methods.
Fang Liu 0034, Jia Liu 0020, Jingxiang Yang, Xu Tang 0004, Liang Xiao 0001
IEEE Trans. Geosci. Remote. Sens.4
2024 Spectral-Spatial Attentions and Deep Supervision for Change Detection in Remote Sensing Images
abstract
The task of remote sensing image change detection involves identifying differences between images captured in the same geographical area but at different times. When dealing with dual-time-series images, lighting and seasonal variations often make recognition challenging. To address the challenges, based on Unet++, we innovatively introduce the Spectral-Spatial Attention Module (SSAM) to better focus on fine-grained details. SSAM uses different frequency components to allocate differential weights to channels, allowing the network to pay more attention to the features relevant to the current task. Moreover, to better capture the change details, a multi-level deep supervision strategy is introduced to enhance the discriminative ability and robustness of early features. Our proposed method is named as SSUNet and has been validated on the CDD and LEVIRE-CD datasets, demonstrating significant advantages in detail recognition.
Jia Liu 0020, Fang Liu 0001, Jingxiang Yang, Liang Xiao 0001
IGARSS5
2024 MIMO-SST: Multi-Input Multi-Output Spatial-Spectral Transformer for Hyperspectral and Multispectral Image Fusion
abstract
The current advanced hyperspectral super-resolution methods utilize Convolutional Neural Networks (CNNs) that are either deeper or wider. These networks are designed to acquire end-to-end mapping capability, facilitating the transformation from Low-Resolution Hyperspectral Images (LR-HSI) and High-Resolution Multispectral Images (HR-MSI) to High-Resolution Hyperspectral Images (HR-HSI). The existing methods lack the capability to capture details and structures in the image effectively, while multi-input and multi-output methods can address this issue efficiently. Therefore, this paper proposes a novel network architecture named Multi-Input Multi-Output Spatial-Spectral Transformer (MIMO-SST). To apply the multi-input and multi-output methods in HSI fusion, specifically integrating the spatial-spectral information of LR-HSI and HR-MSI, we introduce multi-head feature map attention, multi-head feature channel attention, and a multi-scale convolutional gated feedforward network, constructing the proposed Mixture spatial-spectral Transformer. Moreover, to enhance the expressive power of image edges and recover the sharpened structure details, this study incorporates a novel wavelet-based high-frequency loss into the ultimate comprehensive loss, with the objective of refining the reconstruction of high-frequency details. Experimental studies on three simulated datasets and one real-world dataset demonstrate that the proposed method in this study outperforms contemporary state-of-the-art methods in terms of performance. It is noteworthy that our method exhibits a 0.85 dB improvement in terms of the PSNR metric on the CAVE dataset compared to state-of-the-art methods. Our code is publicly available at https://github.com/Freelancefangjian/MIMO-SST.
Jingxiang Yang, Abdolraheem Khader, Liang Xiao 0001
IEEE Trans. Geosci. Remote. Sens.2
2024 Deep Unfolding Network Enhanced by Transformer Priors for Unregistered Hyperspectral and Multispectral Image Fusion
abstract
In satellite remote sensing, the complementary nature of hyperspectral (HSI) and multispectral (MSI) imagery necessitates their fusion to enhance both spatial and spectral resolution. However, the inherent misalignment between these datasets, due to differences in acquisition conditions, poses a significant challenge. This study presents a novel approach, the deep unfolding network enhanced by Transformer priors (DUNET), to address the simultaneous registration and fusion of HSI and MSI. Unlike conventional deep fusion methods, which are often treated as opaque “black boxes,” DUNET incorporates the deep unfolding method, leveraging mutual information and deep priors to facilitate a better degradation model-informed fusion process. The proposed network incorporates hybrid attention Transformers (HATs) and spatial-frequency modules to fully exploit the spatial-spectral information of HSI, resulting in a more accurate and detailed representation of the scene. We conducted extensive quantitative and visual experiments on three standard HSI datasets. The results demonstrate that our proposed DUNET method outperforms the existing mainstream algorithms in the field of remote sensing image fusion, showcasing its effectiveness. Specifically, our proposed method achieves the improvements of 3.4, 5.1, and 8.2 dB in terms of peak signal-to-noise ratio (PSNR) compared with the latest methods on the ICVL, Chikusei, and Houston datasets, respectively.
Jingxiang Yang, Abdolraheem Khader, Liang Xiao 0001
IEEE Trans. Geosci. Remote. Sens.2
2024 Conjoint Cross-Attention Modeling and Joint Feature Calibrating for Remote Sensing Image Change Detection via a Triple-Double Network
abstract
Remote sensing (RS) image change detection (CD) based on deep learning (DL), has received increasing attention recently. However, the general independent learning of bi-temporal images ignores the relationship between them, falling short in learning of the change information. In this paper, a Triple-Double (TD) framework with ability of conjoint cross-attention modeling and joint feature calibrating is proposed for CD. Specifically, the TD framework composed of Triple-branch encoder and Double-branch decoder is constructed to extract diverse features and acquire changed maps with the guidance of original edge cues. To enhance the perception of the connection between the bi-temporal features, the multi-scale difference guidance (MDG) module and conjoint cross-attention (CCA) module are designed for the dual-branch encoder, wherein the CCA introduces a novel and efficient rule for modeling the affinity in spatial and channel dimension simultaneously. Furthermore, a joint feature calibration (JFC) module is introduced to enhance the expression of feature diversity in the joint features within the single-branch encoder. Experimental results on three public datasets demonstrate the superiority of the proposed method compared to the state-of-the-art (SOTA) methods.
Fang Liu 0034, Jia Liu 0020, Jingxiang Yang, Xu Tang 0004, Liang Xiao 0001
IEEE Trans. Geosci. Remote. Sens.4
2024 Content-Guided and Class-Oriented Learning for VHR Image Semantic Segmentation
abstract
With the flourishing of remote sensing (RS) platform techniques, very high-resolution (VHR) images have become more and more popular in recent years, which benefit the task of semantic segmentation but bring new challenges as well. Small objects, such as cars and trees, only occupy a few pixels in VHR images and are usually hard to segment. Moreover, the overlap problem about similar ground objects, such as low vegetation and trees, always results in underperformance. In this article, a content-guided and class-oriented network (CGCO-Net) for VHR image semantic segmentation is proposed to tackle this problem. Specifically, an adaptive content-guided fusion (ACGF) module with deformable convolution is introduced to capture long-distance dependencies and spatial aggregation effectively. With the guidance of the high-level features, the semantic content knowledge is gradually aggregated into low-level features and the details of the original features could be preserved. In addition, a multiscale channel alignment module is introduced into the encoder–decoder structure to further extract the long-range context information and reduce the calculation consumption. In order to improve the ability of pixel-level classification, a class-oriented representation learning (CORL) way is designed with transformer blocks by class embedding and deep supervision, which gradually enhance the discrimination and benefit the final segmentation. Furthermore, a weighted loss function and a threshold optimization strategy are employed to alleviate the sample imbalance problem. Tested on three public datasets and compared with several state-of-the-art methods, the proposed CGCO-net achieves good performance in both qualitative and quantitative analysis.
Fang Liu 0034, Keming Liu, Jia Liu 0020, Jingxiang Yang, Xu Tang 0004, Liang Xiao 0001
IEEE Trans. Geosci. Remote. Sens.4
2024 Hyperspectral Reconstruction From RGB Images via Physically Guided Graph Deep Prior Learning
abstract
Recovering the latent hyperspectral image (HSI) from RGB or multispectral image (MSI), which is dubbed spectral super-resolution (SSR), has demonstrated outstanding performance owing to the advancements in convolutional neural networks (CNNs). However, most of the current algorithms concentrate on the pursuit of networks with more expensive or complex structures, while ignoring the significant role of physical degradation models in SSR. In addition, the inherent defects of CNN make these networks focus more on the local correlation, while their ability to model the long-range correlations in the spectral and spatial domains still has room to improve. To overcome this shortcoming, we propose a physical degradation-guided deep prior learning network (PGDL-Net) for SSR via unfolding the optimization process of the blind SSR model, in which the priors of unknown spectral response function (SRF) and latent HSI are learned explicitly and represented by proximal operators. To jointly extract the local and non-local information, we design a hybrid graph Transformer as the proximal operator to solve the latent HSI. Furthermore, to ensure efficient learning of SRF and HSI, we also propose a novel loss function constraining the reconstruction error, degradation consistency, and observation fidelity for the learned SRF and HSI. Experimental results on multiple datasets illustrate the improved performance and stability of our method in SSR.
Jingxiang Yang, Tian Lin 0001, Jia Liu 0020, Fang Liu 0034, Liang Xiao 0001
IEEE Trans. Geosci. Remote. Sens.2
2024 Unsupervised Deep Tensor Network for Hyperspectral-Multispectral Image Fusion
abstract
Fusing low-resolution (LR) hyperspectral images (HSIs) with high-resolution (HR) multispectral images (MSIs) is a significant technology to enhance the resolution of HSIs. Despite the encouraging results from deep learning (DL) in HSI-MSI fusion, there are still some issues. First, the HSI is a multidimensional signal, and the representability of current DL networks for multidimensional features has not been thoroughly investigated. Second, most DL HSI-MSI fusion networks need HR HSI ground truth for training, but it is often unavailable in reality. In this study, we integrate tensor theory with DL and propose an unsupervised deep tensor network (UDTN) for HSI-MSI fusion. We first propose a tensor filtering layer prototype and further build a coupled tensor filtering module. It jointly represents the LR HSI and HR MSI as several features revealing the principal components of spectral and spatial modes and a sharing code tensor describing the interaction among different modes. Specifically, the features on different modes are represented by the learnable filters of tensor filtering layers, the sharing code tensor is learned by a projection module, in which a co-attention is proposed to encode the LR HSI and HR MSI and then project them onto the sharing code tensor. The coupled tensor filtering module and projection module are jointly trained from the LR HSI and HR MSI in an unsupervised and end-to-end way. The latent HR HSI is inferred with the sharing code tensor, the features on spatial modes of HR MSIs, and the spectral mode of LR HSIs. Experiments on simulated and real remote-sensing datasets demonstrate the effectiveness of the proposed method.
Jingxiang Yang, Liang Xiao 0001, Yongqiang Zhao 0001, Jonathan Cheung-Wai Chan
IEEE Trans. Neural Networks Learn. Syst.1
2023 Edge-Guided Feature Dense Fusion Network for Remote Sensing Image Change Detection
abstract
Remote sensing change detection (CD) is of great importance to Earth observation. Recently, Deep Learning (DL) has been increasingly used to extract useful features and make accurate decisions in a large number of remote sensing images, due to its ability to automatically learn semantic features. However, insufficient fusion of bitemporal images and the lack of prior knowledge of edge structures in current DL methods will result in inaccurate CD results, especially for building boundaries. To alleviate these problems, an edge-guided feature-densely-fused network (EGFDFN) is proposed in this paper. In contrast to conventional Siamese networks, EGFDFN extracts bitemporal features from an extra dual decoder instead of a dual encoder to obtain more accurate change features. In addition, an attention and dense fusion module (ADFM) and an edge guidance module (EGM) are used to enhance features and make full use of edge information. Experimental results demonstrate that the proposed method outperforms on LEVIR-CD dataset among other representative methods.
Hejun Luo, Jia Liu 0020, Fang Liu 0001, Jingxiang Yang, Liang Xiao 0001
IGARSS5
2023 Multiple Deep Proximal Learning for Hyperspectral-Multispectral Image Fusion
abstract
Fusing low resolution (LR) hyperspectral image (HSI) with a high resolution (HR) multispectral image (MSI) could enhance the spatial resolution and quality of HSI. Current deep learning (DL) HSI-MSI fusion networks have achieved encouraging results, but their performance relies on large number of training images with known degradations consistent with the testing data. The trained DL model may fail on data with unseen degradations during inference. In this study, we propose a multiple deep proximal learning network (MDPro-Net) for HSI-MSI fusion, the unknown spatial-spectral degradations and latent HR HSI can be adaptively inferred. We first propose a joint variational fusion model with both the degradations and HR HSI as to-be-solved variables, which are regularized by multiple deep priors. Then we optimize the fusion model using quadratic splitting and alternative optimization strategy. The unknown blurring kernel, spectral degradation, and HR HSI are explicitly solved by three deep proximal operators. Through unrolling the solutions into a DL network, we build MDPro-Net, in which the deep proximal operators for degradations and HR HSI are learned in an end-to-end manner. Furthermore, in the deep proximal operator for latent HR HSI, a multi-scale transformer is designed to exploit the local and non-local dependencies. Experiments demonstrate that the proposed MDPro-Net is competitive with state-of-the-art fusion methods, in particular, it is robust in inferring the unseen degradations.
Jingxiang Yang, Tian Lin 0001, Xiaoyang Chen 0005, Liang Xiao 0001
IEEE Trans. Geosci. Remote. Sens.1
2023 Learning Degradation-Aware Deep Prior for Hyperspectral Image Reconstruction
abstract
Reconstructing the 3D hyperspectral image (HSI) from 2D snapshot measurements is a key task in spectral snapshot compressive imaging (SCI). Traditional model-based HSI reconstruction methods rely on hand-crafted priors. Recently, deep unfolding networks (DUNs) learn the priors using convolutional neural networks (CNNs) and have achieved satisfactory results. Most of DUNs assume the degradations of SCI are known. However, due to the phase aberration and distortion problems in real imaging process, there is a certain gap between the ideal and real degradation patterns, which may hinder the accurate HSI reconstruction. In this study, we propose a degradation-aware deep prior learning network (D2PL-Net), which tries to adaptively learn the practical degradation matrix during HSI reconstruction, thus bridges the gap between the ideal and real degradations. Specifically, we first propose a joint variational compressive reconstruction model, both of the latent HSI and unknown degradation can be explicitly solved. By unfolding the solutions into a deep network, D2PL-Net is built, which mainly consists of two parts, Degradation Matrix Learning (DML) mechanism and Degradation-guided Spectral-Spatial Transformer (DSST) in each stage. The former learns the degradation that approximates the real one; the latter represents the deep prior of latent HSI, it could exploit the spectral-wise and spatial-wise long-range dependencies of HSI under the guidance of learned degradation, and then reconstructs the HSI. To ensure an effective training of D2PL-Net, we propose a joint loss function constraining the HSI reconstruction errors, degradation-fidelity and degradation-consistency. Experiments on simulated and real-life datasets show that the proposed method is competitive with the state-of-the-art methods.
Jingxiang Yang, Tian Lin 0001, Fang Liu 0034, Liang Xiao 0001
IEEE Trans. Geosci. Remote. Sens.1
2022 Learning a Coupled Multilinear Network for Unsupervised Hyperspectral-Multispectral Image Fusion
abstract
Fusing low resolution (LR) HSI with high resolution (HR) multispectral image (MSI) is an important technology to obtain HR hypersepctral image (HSI), which is hard to directly acquire due to the hardware limitation. Deep learning (DL) has been applied in HSI-MSI fusion, but the representability of DL networks for multidimensional (i.e., spectral-spatial) features still need improvement. And most DL HSI-MSI fusion networks are in supervised fashion, HR ground truth HSI is required for training, which is unavailable in reality. In this work, we investigate tensor theory, and propose a coupled multilinear network (CMuNet) for unsupervised HSI-MSI fusion, where deep image prior and degradation model can be jointly learned. CMuNet consists of coupled multilinear filtering subnets, it jointly represents the LR HSI and HR MSI as a random code and multidimensional features on spatial and spectral modes. The HR HSI is inferred with the random code, features on spatial modes of HR MSI and features on spectral mode of LR HSI. Experiments on several HSIs demonstrate the effectiveness of the proposed method.
Jingxiang Yang, Liang Xiao 0001
IGARSS1
2022 Learning Deep Subspace Projection Prior for Dual-Camera Compressive Hyperspectral Imaging
abstract
Coded aperture snapshot spectral imaging (CASSI) captures the 3-D hyperspectral images (HSI) in the form of 2-D coded images. The dual-camera compressive hyperspectral imaging (DCCHI) can effectively improve the reconstruction quality by adding a parallel complementary panchromatic camera. Several regularization-based methods have been proposed for dual-camera reconstruction. However, the handcrafted priors of these methods are limited in representing the complex intrinsic structure of HSI. In this letter, we propose to learn deep subspace projection prior for dual-camera compressive reconstruction. We first design a deep subspace projection prior regularized dual-camera compressive reconstruction model and minimize it with alternative optimization. Then, we unfold the optimization process into a network. Specifically, the deep subspace projection prior learning leads to features with low-rank characteristics, which could efficiently exploit the spectral correlation of HSI. The dual-camera compressive reconstruction network is learned in an end-to-end manner. Extensive experiments substantiate the performance and efficiency of other start-of-the-art algorithms.
Xiaoyang Chen 0005, Jingxiang Yang, Liang Xiao 0001
IEEE Geosci. Remote. Sens. Lett.2
2022 Multilevel Superpixel Structured Graph U-Nets for Hyperspectral Image Classification
abstract
Limited by the shape-fixed kernels, convolutional neural networks (CNNs) are usually difficult to model difform land covers in hyperspectral images (HSIs), leading to inadequate land use. Recently, benefiting from the ability to conduct shape-adaptive convolutions and model complex patterns in graph-structured data, graph convolutional networks (GCNs) have been applied to HSI classification. However, due to the massive computation in GCNs, HSI is usually pretreated into a graph based on a specific superpixel segmentation, which limits the modeling of spatial topologies to the same scale. To break this limitation, we propose a multilevel superpixel structured graph U-Net (MSSGU) to learn multiscale features on multilevel graphs. Specifically, we construct several hierarchical segmentations from fine to coarse by progressively merging adjacent superpixels and then convert them into multilevel graphs. Meanwhile, based on the merging relations between hierarchical superpixels, we establish the pooling and unpooling functions to transfer features from one graph to another, thereby enabling different-level graphs to collaborate in a single network. Different from concatenating different-scale features straightforwardly in the feature fusion stage, MSSGU fuses them in a coarse-to-fine progressive manner, which can generate subtler fusion features adaptive to the pixelwise classification task. Moreover, we use a CNN instead of GCN to extract and fuse the pixel-level features, which greatly reduces the computation. Such a hybrid U-Net can exploit features of HSIs from a multiscale hierarchical perspective, and its performance has been proven competitive with other deep-learning-based methods by extensive experiments on three benchmark datasets.
Qichao Liu, Liang Xiao 0001, Jingxiang Yang, Zhihui Wei
IEEE Trans. Geosci. Remote. Sens.3
2022 Detail-Injection-Model-Inspired Deep Fusion Network for Pansharpening
abstract
Pansharpening is an image fusion procedure, which aims to produce a high spatial resolution multispectral image by combining a low spatial resolution multispectral image and a high spatial resolution panchromatic image. The most popular and successful paradigm for pansharpening is the framework known as detail injection, while it cannot fully exploit complex and non-linear complementary features of both images. In this paper, we propose a detail injection model inspired deep fusion network for pansharpening (DIM-FuNet). Firstly, by treating pansharpening as a complicated and non-linear details learning and injection problem, we establish a unified optimizing detail-injection model with triple detail fidelity terms: 1) a band-dependent spatial detail fidelity term, 2) a local detail fidelity term and 3) a complicated details synthesis term. Secondly, the model is optimized via the iterative gradient descent and unfolded into a deep convolutional neural network. Subsequently, the unrolling network has triple branches, in which, a point-wise convolutional sub-network, a depth-wise convolutional sub-network are corresponding to the former two detail constrained terms, and an adaptive weighted reconstruction module with a fusion sub-network to aggregate details of two branches and synthesis the final complicated details. Finally, the deep unrolling network is trained in end-to-end manners. Different from traditional deep fusion networks, the architecture design of DIM-FuNet is guided by the optimizing model and thus promotes better interpretability. Experimental results on reduced and full-resolution demonstrate the effectiveness of the proposed DIM-FuNet which achieves the best performance compared with the state-of-the-art pansharpening method.
Zhikang Xiang, Liang Xiao 0001, Jingxiang Yang, Wenzi Liao, Wilfried Philips
IEEE Trans. Geosci. Remote. Sens.3
2022 Variational Regularization Network With Attentive Deep Prior for Hyperspectral-Multispectral Image Fusion
abstract
Hyperspectral–multispectral image (HSI-MSI) fusion relies on a robust degradation model and data prior, where the former describes the degeneration of HSI in the spectral and spatial domains, and the latter reveals the latent statistics of the expected high-resolution (HR) HSI. In practice, the degradation model is often unknown, and the data prior is usually too complicated to be expressed analytically. In this study, we propose a variational network for HSI-MSI fusion (VaFuNet), in which the degradation model and data prior are implicitly represented by a deep learning network and jointly learned from the training data. A variational fusion model regularized by deep prior is first proposed, and then, it is optimized via a half-quadratic splitting and unfolded into a deep network. The deep prior is implicitly represented by a proximity operator. Due to the structural self-similarity, HSI possesses structural recurrences across different scales. To exploit such nonlocal prior and enhance the representability of network, we also propose a multiscale nonlocal attention and embed it into the deep prior proximity. The degradation model and deep prior proximity are jointly learned via end-to-end training. Experimental results on simulated and real-life HSI datasets demonstrate the effectiveness of the proposed VaFuNet HSI-MSI fusion method.
Jingxiang Yang, Liang Xiao 0001, Yongqiang Zhao 0001, Jonathan Cheung-Wai Chan
IEEE Trans. Geosci. Remote. Sens.1
2021 Multi-Supervised Recursive-CNN for Hyperspectral and Multispectral Image Fusion
abstract
Deep learning has been widely used in remote sensing images fusion in recent years. However, many deep learning based methods use interpolation or a deconvolutional layer to upsample the low-resolution hyperspectral(LRHS) image, then the upsampled image is concated with the high-resolution multispectral(HRMS) image to fed the network, which lead to the spatial-spectral information loss. In this study, we propose a multi-supervised recursive convolutional neural network(MRCNN) for HS/MS images fusion. Specifically, we use the upsampling recursive sub-net(URSN) to upsample the LRHS image, which can effectively avoid the spatial-spectral information loss. In addition, to make our deep network much lighter, we introduce recursive learning to the network by using a residual block as the recursive unit for several recursions. Finally, a multi-supervised learning strategy is adopted for enhancing the gradient propagation and avoiding vanishing and exploding gradients. Simulated experiments on Cave and Moffett Field datasets show that the proposed network outperforms many state-of-the-art ones.
Yuda Lu, Jingxiang Yang, Liang Xiao 0001
IGARSS2
2021 Hybrid Local and Nonlocal 3-D Attentive CNN for Hyperspectral Image Super-Resolution
abstract
A deep convolutional neural network (CNN) has shown its great potential in hyperspectral image (HSI) super-resolution (SR). Integrating CNN with attention mechanism is expected to boost the SR performance. However, how to learn attention along the spectral, spatial, and channel dimensions of HSI is still an open issue, and the current attention mechanism is not efficient in capturing long-range interdependency in HSI. In this letter, we first design a local 3-D attention module to learn the spectral-spatial-channel attention by exploiting local contextual information in HSI. Then, we propose a nonlocal 3-D attention module, in which the long-range interdependency in HSI can be exploited for attention learning. By jointly embedding the local and nonlocal attention in a residual 3-D CNN, a hybrid local and nonlocal 3-D attentive CNN can be built for HSI SR. The experimental results show that local and nonlocal attention formulation leads to competitive SR performance.
Jingxiang Yang, Liang Xiao 0001, Yongqiang Zhao 0001, Jonathan Cheung-Wai Chan
IEEE Geosci. Remote. Sens. Lett.1
2021 CNN-Enhanced Graph Convolutional Network With Pixel- and Superpixel-Level Feature Fusion for Hyperspectral Image Classification
abstract
Recently, the graph convolutional network (GCN) has drawn increasing attention in the hyperspectral image (HSI) classification. Compared with the convolutional neural network (CNN) with fixed square kernels, GCN can explicitly utilize the correlation between adjacent land covers and conduct flexible convolution on arbitrarily irregular image regions; hence, the HSI spatial contextual structure can be better modeled. However, to reduce the computational complexity and promote the semantic structure learning of land covers, GCN usually works on superpixel-based nodes rather than pixel-based nodes; thus, the pixel-level spectral–spatial features cannot be captured. To fully leverage the advantages of the CNN and GCN, we propose a heterogeneous deep network called CNN-enhanced GCN (CEGCN), in which CNN and GCN branches perform feature learning on small-scale regular regions and large-scale irregular regions, and generate complementary spectral–spatial features at pixel and superpixel levels, respectively. To alleviate the structural incompatibility of the data representation between the Euclidean data-oriented CNN and non-Euclidean data-oriented GCN, we propose the graph encoder and decoder to propagate features between image pixels and graph nodes, thus enabling the CNN and GCN to collaborate in a single network. In contrast to other GCN-based methods that encode HSI into a graph during preprocessing, we integrate the graph encoding process into the network and learn edge weights from training data, which can promote the node feature learning and make the graph more adaptive to HSI content. Extensive experiments on three data sets demonstrate that the proposed CEGCN is both qualitatively and quantitatively competitive compared with other state-of-the-art methods.
Qichao Liu, Liang Xiao 0001, Jingxiang Yang, Zhihui Wei
IEEE Trans. Geosci. Remote. Sens.3
2020 Video Deblurring Via 3d CNN and Fourier Accumulation Learning
abstract
Camera shake and target movement often leads to undesirable image blurring in videos. How to exploit spatial-temporal information of adjacent frames and reduce the processing time of deblurring are two major issues in video deblurring. In this paper, we propose a simple yet effective Fourier accumulation embedded 3D convolutional encoder-decoder network for video deblurring. Firstly, a 3D convolutional encoder-decoder module is constructed to extract multiscale spatial-temporal deep features and generate intermediate deblurred frames with complementary information which is beneficial for the deblurring of each frame. Then we embed a Fourier accumulation module following the 3D convolutional encoder-decoder, the Fourier accumulation module could fuse intermediate deblurred frames with learned weights in Fourier domain and then produce shaper deblurred frames. Experimental results show that our method has competitive performance compared with other state-of-the-art methods.
Liang Xiao 0001, Jingxiang Yang
ICASSP3
2020 A Blind CSI Prediction Method Based on Deep Learning for V2I Millimeter-Wave Channel
abstract
With the development of the Internet of vehicles and 5G, there emerge more and more challenging application scenarios with fast time-varying channels and high mobility nodes, such as high speed trains environment and vehicle-to-infrastructure (V2I) communication in highway. To support the reliable vehicular communication and mobile edge computing (MEC), it is important to obtain the future channel state information (CSI), which can help optimize system transmission scheme. In this paper, we propose an efficient blind CSI prediction model, called BCPMN. We first reshape the sampled signal into a specific 2-dimensional matrix. Then we propose a learning framework contains of convolutional neural network (CNN), long short-term memory (LSTM) network and fully connected layers. To validate the proposed model, we conduct extensive experiment in three modulation modes. The results show that the BCPMN achieves highly accurate signal-to-noise ratio (SNR) prediction in the fast changing channel model with different modulation modes. In particular, the proposed model can obtain better performance than other methods, and can achieve better performance than other methods without the payload cost of pilot.
Jingxiang Yang, Liyan Li, Minjian Zhao
ICNP1
2020 Content-Guided Convolutional Neural Network for Hyperspectral Image Classification
abstract
Convolutional neural networks (CNNs) are of great interest and have demonstrated remarkable performance in hyperspectral images (HSIs) classification. However, due to the current configuration of the convolution layers with a fixed kernel shape, regular CNNs are inherently limited in modeling the diverse land-cover structures, particularly in the cross-classes edge regions, where irregular class boundaries would lead to high classification errors. To address this issue, we propose a content-guided CNN (CGCNN) for HSI classification. Compared with the shape-fixed kernel in the traditional CNN, the proposed content-guided convolution adaptively adjusts its kernel shape according to the spatial distribution of land covers. The content pattern is reflected by a latent guide map automatically learned from HSI. Such content-adaptive kernel with CGCNN could suppress the irregularity and unexpected features in class boundaries and, thus, improve the feature learning in cross-classes regions. Based on the content-guided convolution, a novel guided feature extraction unit (GFEU) is constructed for spectral-spatial feature learning of HSI. Finally, the CGCNN classification framework is established by stacking multiple GFEUs with dense connection, which is helpful for mitigating the gradient vanishing and increasing the robustness to overfitting. Extensive experiments on several HSIs demonstrate that the proposed approach possesses great details' preserving ability and its performance outperforms other state-of-the-art methods.
Qichao Liu, Liang Xiao 0001, Jingxiang Yang, Jonathan Cheung-Wai Chan
IEEE Trans. Geosci. Remote. Sens.3
2019 Hyperspectral Image Super-Resolution Based on Multi-Scale Wavelet 3D Convolutional Neural Network
abstract
Super-resolution (SR) of hyperspectral image (HSI) is of significance for its applications. Wavelet decomposition can be used to capture textures and structures in the HSI. In this study, we propose a multi-scale wavelet 3D convolutional neural network (MW-3D-CNN) for HSI SR. Instead of reconstructing the high resolution (HR) HSI directly, we predict the wavelet coefficients of HR HSI with the proposed network, which is composed of an embedding subnet and a predicting subnet. Both of them are built with 3D convolutional layers. The embedding subnet extracts deep spatial-spectral features from the low resolution (LR) HSI and represents the LR HSI as a set of feature cubes. The feature cubes are then fed to the predicting subnet. There are multiple output branches in the predicting subnet, each of which corresponds to a wavelet sub-band and predicts the wavelet coefficients of HR HSI. By applying inverse wavelet transform to the predicted wavelet coefficients, the HR HSI can be obtained. In the training stage, we propose to train MW-3D-CNN with L1 norm loss, which is more suitable than the conventional L2 norm loss for penalizing the errors in different wavelet sub-bands. In the experiment, the performance is tested on several HSI datasets.
Jingxiang Yang, Yongqiang Zhao 0001, Jonathan Cheung-Wai Chan
IGARSS1
2017 Learning and Transferring Deep Joint Spectral-Spatial Features for Hyperspectral Classification
abstract
Feature extraction is of significance for hyperspectral image (HSI) classification. Compared with conventional hand-crafted feature extraction, deep learning can automatically learn features with discriminative information. However, two issues exist in applying deep learning to HSIs. One issue is how to jointly extract spectral features and spatial features, and the other one is how to train the deep model when training samples are scarce. In this paper, a deep convolutional neural network with two-branch architecture is proposed to extract the joint spectral-spatial features from HSIs. The two branches of the proposed network are devoted to features from the spectral domain as well as the spatial domain. The learned spectral features and spatial features are then concatenated and fed to fully connected layers to extract the joint spectral-spatial features for classification. When the training samples are limited, we investigate the transfer learning to improve the performance. Low and mid-layers of the network are pretrained and transferred from other data sources; only top layers are trained with limited training samples extracted from the target scene. Experiments on Airborne Visible/Infrared Imaging Spectrometer and Reflective Optics System Imaging Spectrometer data demonstrate that the learned deep joint spectral-spatial features are discriminative, and competitive classification results can be achieved when compared with state-of-the-art methods. The experiments also reveal that the transferred features boost the classification performance.
Jingxiang Yang, Yongqiang Zhao 0001, Jonathan Cheung-Wai Chan
IEEE Trans. Geosci. Remote. Sens.1
2017 Joint Hyperspectral Superresolution and Unmixing With Interactive Feedback
abstract
This paper presents an interactive feedback scheme of spatial resolution enhancement and spectral unmixing in hyperspectral imaging. Traditionally spatial resolution enhancement and spectral unmixing operations have been carried out separately, often in series. In such sequential processing, spatially enhanced hyperspectral images (HSIs) may introduce distortion in spectral fidelity making spectral unmixing results unreliable, or vice versa. Since both high- and low-resolution HSIs have the same endmembers, the deviation in spectral unmixing between targets and estimated high-resolution HSIs can be used as feedback to control spatial resolution enhancement. The spatial difference before and after unmixing can also be used as feedback to enhance spectral unmixing. Therefore, spectral unmixing is utilized as a constraint to spatial resolution enhancement, while spatial resolution enhancement helps improve spectral unmixing results. The performance of spatial resolution enhancement and spectral unmixing can be improved since one behaves like a prior to the other. Experimental results on both simulated and real HSI data sets demonstrate that the proposed interactive feedback scheme simultaneously achieved spatial resolution enhancement and spectral unmixing fidelity. This paper is an extended version of the previous work.
Yongqiang Zhao 0001, Jingxiang Yang, Jonathan Cheung-Wai Chan, Seong G. Kong
IEEE Trans. Geosci. Remote. Sens.3
2016 Hyperspectral image classification using two-channel deep convolutional neural network
abstract
Performance of hyperspectral image classification depends on feature extraction. Compared with conventional hand-crafted feature extraction, deep learning can learn feature with more discriminative information. In this paper, a two-channel deep convolutional neural network (Two-CNN) is proposed to learn jointly spectral-spatial feature from hyperspectral image. The proposed model is composed of two channels of CNN, each of which learns feature from spectral domain and spatial domain respectively. The learned spectral feature and spatial feature are then concatenated and fed to fully connected layer to extract joint spectral-spatial feature for classification. When number of training samples is limited, we propose to train the deep model using transfer learning to improve the performance. Low-layer and mid-layer features of the deep model are learned and transferred from other scenes, only top-layer feature is learned using the limited training samples of the current scene. Experiment results on real data demonstrate the effectiveness of the proposed method.
Jingxiang Yang, Yongqiang Zhao 0001, Jonathan Cheung-Wai Chan
IGARSS1
2016 A hyperspectral spatial-spectral enhancement algorithm
abstract
Low spatial and spectral resolution hyperspectral image will always degrade the performance of the subsequent applications, such as classification and object detection. The desired hyperspectral image is assumed to be reconstructed based on both high spatial and spectral features, which are always represented using endmembers and their abundances. In this paper, we propose a hyperspectral spatial and spectral resolution enhancement algorithm based on spectral unmixing and spatial constraints to simultaneously obtain high spatial-spectral resolution result. An intermediate high spatial but low spectral resolution HSI is introduced to establish mapping scheme of abundances and endmembers between low resolution input and desired high spatial-spectral resolution result. Experiments on the Sandigo dataset have illustrated that the proposed method is comparable or superior to other state-of-art methods.
Yongqiang Zhao 0001, Jingxiang Yang
IGARSS3
2016 Coupled Sparse Denoising and Unmixing With Low-Rank Constraint for Hyperspectral Image
abstract
Hyperspectral image (HSI) denoising is significant for correct interpretation. In this paper, a sparse representation framework that unifies denoising and spectral unmixing in a closed-loop manner is proposed. While conventional approaches treat denoising and unmixing separately, the proposed scheme utilizes spectral information from unmixing as feedback to correct spectral distortion. Both denoising and spectral unmixing act as constraints to the others and are solved iteratively. Noise is suppressed via sparse coding, and fractional abundance in spectral unmixing is estimated using the sparsity prior of endmembers from a spectral library. The abundance of endmembers is used as a spectral regularizer for denoising based on the hypothesis that spectral signatures obtained from a denoising process result are close to those of unmixing. Unmixing restrains spectral distortion and results in better denoising, which reciprocally leads to further improvements in unmixing. The strength of our proposed method is illustrated by simulated and real HSIs with performance competitive to the state-of-the-art denoising and unmixing methods.
Jingxiang Yang, Yongqiang Zhao 0001, Jonathan Cheung-Wai Chan, Seong G. Kong
IEEE Trans. Geosci. Remote. Sens.1
2015 Hyperspectral Image Denoising via Sparse Representation and Low-Rank Constraint
abstract
Hyperspectral image (HSI) denoising is an essential preprocess step to improve the performance of subsequent applications. For HSI, there is much global and local redundancy and correlation (RAC) in spatial/spectral dimensions. In addition, denoising performance can be improved greatly if RAC is utilized efficiently in the denoising process. In this paper, an HSI denoising method is proposed by jointly utilizing the global and local RAC in spatial/spectral domains. First, sparse coding is exploited to model the global RAC in the spatial domain and local RAC in the spectral domain. Noise can be removed by sparse approximated data with learned dictionary. At this stage, only local RAC in the spectral domain is employed. It will cause spectral distortion. To compensate the shortcoming of local spectral RAC, low-rank constraint is used to deal with the global RAC in the spectral domain. Different hyperspectral data sets are used to test the performance of the proposed method. The denoising results by the proposed method are superior to results obtained by other state-of-the-art hyperspectral denoising methods.
Yongqiang Zhao 0001, Jingxiang Yang
IEEE Trans. Geosci. Remote. Sens.2
2014 Coupled hyperspectral super-resolution and unmixing
abstract
The acquired hyperspectral data are always in low resolution in both spatial and spectral domains, which will result in lots of mixed pixels and degrade the detection and recognition performance in civil and military applications. So many super resolution techniques are applied to overcome this limit. In this paper, we propose a coupled hyperspectral spatial super-resolution and spectral unmixing method based on sparse representation. Combing spatial super-resolution and spectral unmixing can precisely conserve both spatial information and spectral correlation among different bands. Spectral unmixing is taken as a regularization term in spatial super-resolution to test spectral consistency and avoid spectral distortion, while spatial super-resolution is used to enhance the resolution of abundance map after spectral unmixing.
Yongqiang Zhao 0001, Jingxiang Yang, Jonathan Cheung-Wai Chan
IGARSS3