Song Xiao 0001

dblp:x/SongXiao-1 · DBLP profile ↗
← Back
67ranked-venue papers
7as first author
44since 2021 · last 2026
0000-0001-8988-4233ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 23 · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 21 · 3 first-author · 13 since 2021Artificial intelligence and machine learning · 15 · 13 since 2021Computer networks · 3 · 1 since 2021Security and privacy · 1Databases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2026 Cross-scene adversarial learning with Gaussian mixture model for hyperspectral anomaly detection
Shaoxiong Hou, Song Xiao 0001, Jiahui Qu, Wenqian Dong
Expert Syst. Appl.2
2026 CrossNeXt: Interactive siamese ConvNeXt with contrastive learning and edge-aware recalibration for constrained image splicing detection and localization
Song Xiao 0001, Wenqian Yue, Shengwei Xu
Expert Syst. Appl.3
2026 RCDIFO: A registration-change detection iterative feedback optimization network for unwell registered hyperspectral images
Jiahui Qu, Song Xiao 0001, Wenqian Dong, Yunsong Li 0001
Pattern Recognit.3
2026 TSFE-Net: Document Image Forgery Localization via Text Structural Feature Enhancement
Yiyuan Sun, Hanhan Wang, Song Xiao 0001
IEEE Signal Process. Lett.4
2025 MUN: Image Forgery Localization Based on M³ Encoder and UN Decoder
abstract
Image forgeries can entirely change the semantic information of an image, and can be used for unscrupulous purposes. In this paper, we propose a novel image forgery localization network named as MUN, which consists of an M^3 encoder and a UN decoder. Firstly, the M^3 encoder is constructed based on a Multi-scale Max-pooling query module to extract Multi-clue forged features. Noiseprint++ is adopted to assist the RGB clue, and its deployment methodology is discussed. A Multi-scale Max-pooling Query (MMQ) module is proposed to integrate RGB and noise features. Secondly, a novel UN decoder is proposed to extract hierarchical features from both top-down and bottom-up directions, reconstructing both high-level and low-level features at the same time. Thirdly, we formulate an IoU-recalibrated Dynamic Cross-Entropy (IoUDCE) loss to dynamically adjust the weights on forged regions according to IoU which can adaptively balance the influence of authentic and forged regions. Last but not least, we propose a data augmentation method, i.e., Deviation Noise Augmentation (DNA), which acquires accessible prior knowledge of RGB distribution to improve the generalization ability. Extensive experiments on publicly available datasets show that MUN outperforms the state-of-the-art works.
Shuhuan Chen, Haichao Shi, Xiaoyu Zhang 0002, Song Xiao 0001, Qiang Cai 0001
AAAI5
2025 Do You Steal My Model? Signature Diffusion Embedded Dual-Verification Watermarking for Protecting Intellectual Property of Hyperspectral Image Classification Models
abstract
Due to the high cost of data collection and training, the well-performed hyperspectral image (HSI) classification models are of great value and vulnerable to piracy threat during transmission and use. Model watermarking is a promising technology for intellectual property (IP) protection of models. However, the existing model watermarking methods for RGB image classification models ignore the complexity of ground objects and high dimension of HSIs, which makes trigger samples easy to be detected and forged. To address this problem, we propose a signature diffusion embedded dual-verification watermarking method, which generates imperceptible trigger samples with explicit owner information to achieve dual verification of both model ownership and legality of trigger set. Specifically, the subpixel-space owner signature diffusion incorporated imperceptible trigger set generation method is proposed to manipulate owner signature incorporated to the abundance matrix of seeds via diffusion model in subpixel space, thus balancing the perceptual quality of trigger samples and signature extraction capability. To resist ownership confusion, dual-stamp ownership verification is proposed to query the suspicious model with trigger samples for ownership verification, and further extracts signature from trigger samples to guarantee their legality. Extensive experiments demonstrate the proposed method can effectively protect IP of HSI classification models.
Song Xiao 0001, Lixiang Li 0001, Wenqian Dong, Jiahui Qu
IJCAI2
2025 Cycle-Consistent Mamba-Based Registration-Fusion Joint Network for Unregistered Hyperspectral Image Super-Resolution
abstract
Hyperspectral image super-resolution (HSI-SR) has attracted significant attention in high-resolution HSI reconstruction. Most existing fusion-based HSI-SR methods assume that multi-source images are perfectly registered, which is impractical due to varying imaging conditions. Furthermore, methods that consider the registration issue typically treat registration and fusion as two separate steps, resulting in the accumulation of registration errors during the fusion process. To address these issues, we propose a Cycle-Consistent Mamba-Based Registration-Fusion Joint Network (CCM-RFJN), which step-wise optimizes the Registration-Fusion Unified Module (RFUM) through multiple cyclic iterative SR processes. Specifically, in each SR iteration, we map the super-resolved HR-HSI obtained through the RFUM back to the unregistered LR-HSI for the next SR, with cycle-consistency constraints imposed on both LR-HSI and HR-HSI to adaptively optimize the RFUM based on the reciprocal training strategy. In RFUM, we integrate the proposed Interactive Mamba Registration Module (IMR) and Dual-attention Mamba Fusion Module (DAMF), thereby achieving registration-fusion joint optimization. Specifically, IMR is developed to incorporate the interactive Mamba encoder into a pyramid architecture to facilitate multi-level information interactions, generating the deformation field to correct non-rigid misalignments. DAMF is designed to utilize the dual-attention Mamba mechanism to highlight and aggregate key features, thereby enhancing fusion performance. Experiments on three public datasets demonstrate that CCM-RFJN achieves the state-of-the-art performance. The code is available at https://github.com/Jiahuiqu/CCM-RFJN.
Quangui He, Jiahui Qu, Wenqian Dong, Song Xiao 0001, Qinghao Gao
ACM Multimedia4
2025 Multiscale common-private feature adversarial decoupling network for hyperspectral pansharpening
Shaoxiong Hou, Song Xiao 0001, Jiahui Qu, Wenqian Dong
Knowl. Based Syst.2
2025 Registration-fusion binocular diffusion model: Exploring continuous fusion of unregistered hyperspectral and multispectral images
Jiahui Qu, Wenqian Dong, Hongxiang Li 0002, Song Xiao 0001, Yunsong Li 0001
Knowl. Based Syst.5
2025 RHN: RoI Restricted Hybrid Network for Instance-Aware Image-to-Image Translation
abstract
Image-to-image translation recently attracts a lot of attention, while it also encounters challenges when processing images which contain numerous instances. In this letter, a novel RoI restricted Hybrid Network (RHN) is proposed for instance-aware image-to-image translation. Firstly, we propose a batch-wise RoI restricted module which combines region restriction with batch-wise attention to recalibrate instance-level features. Secondly, we design a Transformer block with local Xception which can supplement local information on global features extracted from attention, to provide discriminative multi-scale features. Finally, we formulate a multi-task loss which can enforce RHN to achieve good instance-ware image-to-image translation performance. Experiments on publicly available datasets demonstrate the effectiveness of RHN.
Hanhan Wang, Song Xiao 0001, Qiang Cai 0001
IEEE Signal Process. Lett.4
2025 MambaMTL: Progressive Mutual-Guided Mamba Multitask Learning for Hyperspectral Image Pansharpening and Classification
abstract
Multi-task learning (MTL) serves as a effective technology to improving both the performance of hyperspectral image (HSI) pansharpening and that of downstream classification tasks. However, most of the existing MTL frameworks overlook the close relationship between the two tasks, resulting in suboptimal performance. To solve the problem, we propose a progressive mutual-guided Mamba MTL framework (MambaMTL), which achieves mutual enhancement between the two tasks by interacting the refined spatial-spectral features and the rich class semantic information generated by each task. Specifically, we design a dual-branch Mamba-Unet subnetwork with class-aware refinement (MTL-CRMNet) for HSI pansharpening. It incorporates parallel multi-head class-aware refinement layers (MCRLs) to enhance object details in the reconstructed HSI by enriching class-specific features with semantic information from the classification network, thereby improving pansharpening performance. Besides, we introduce a Mamba-based pyramid subnetwork with hierarchical guided fusion (MTL-MGFNet) for HSI classification. It utilizes the pansharpened HSI to guide multi-scale feature fusion in a feature pyramid manner, enhancing classification accuracy. Notably, both MTL-MGFNet and MTL-CRMNet share the same multi-scale feature encoder, enabling the classification subnetwork to generate spatial-spectral features rich in texture and semantics. The interaction between tasks significantly enhances their respective performance. The evaluation results conducted on relevant datasets illustrate the superiority of the proposed MambaMTL in simultaneously achieving better HSI classification accuracy and higher reconstructed image quality.
Shaoxiong Hou, Song Xiao 0001, Jiahui Qu, Wenqian Dong
IEEE Trans. Geosci. Remote. Sens.2
2025 A Progressive Registration-Fusion Co-Optimization A-Mamba Network: Toward Deep Unregistered Hyperspectral and Multispectral Fusion
abstract
The existing methods of hyperspectral image (HSI) and multispectral image (MSI) fusion usually overlook the fact that multi-source images acquired under different imaging conditions are generally not perfectly registered. Despite the many such methods that have begun to address registration issues, it is still a challenge that most works perform registration and fusion as two separate steps, resulting in a cumulative error. To address this challenge, we propose a progressive registration-fusion co-optimization A-Mamba network (PRFCoAM), which iteratively optimizes the modal-aligned progressively registration-fusion (MAPRF) module to adaptively corrects the deformation from an extensive to a detailed level and refines the fusion results at each level to achieve progressive registration-fusion co-optimization. The proposed MAPRF module integrates the modal unified local aware registration (MULAR) block and interactive attention Mamba fusion (IAMF) block, which facilitates the network comprehensively and efficiently capture features of different levels. Specifically, MULAR adaptively learns spectral and spatial degradation functions to transform the input images into a unified modality and progressively repairs non-rigid pixel offsets by capturing the correlations and differences between corresponding regions of images. IAMF multi-directionally scans the spatial and spectral global dependent features of the well-registered images, which can stimulate the potential of Mamba in fusion and achieve a win-win situation of computational efficiency and selectivity advantages in the global acceptance domain. Extensive experiments demonstrate PRFCoAM can flexibly deal with different degrees and kinds of non-rigid deformation and achieves state-of-the-art performance. The code will be available at https://github.com/Jiahuiqu/PRFCoAM-for-HSI-MSI-Registration-Fusion.
Zan Li 0001, Yue Wen, Song Xiao 0001, Jiahui Qu, Wenqian Dong
IEEE Trans. Geosci. Remote. Sens.3
2025 Cycle Translation-Based Collaborative Training for Hyperspectral-RGB Multimodal Change Detection
abstract
Hyperspectral image change detection (HSI-CD) benefits from HSIs with continuous spectral bands, which uniquely enables the analysis of more subtle changes. Existing methods have achieved desirable performance relying on multi-temporal homogenous HSIs over the same region, which is generally difficult to obtain in real scenes. HSI-RGB multimodal CD overcomes the constraint of limited HSI availability by incorporating another temporal RGB data, and the combination of advantages within different modalities enhances the robustness of detection results. Nevertheless, due to the different imaging mechanisms between two modalities, existing HSI CD methods cannot be directly applied. In this paper, we propose a cycle translation-based collaborative training (co-training) for HSI-RGB multimodal CD, which achieves cross-modal mutual guidance to collaboratively learn complementary difference information from diverse modalities for identifying changes. Specifically, a cross-modal guided CycleGAN-based image translation module is designed to implement bi-directional image translation, which mitigates modal difference and enables the extraction of information related to land cover changes. Then, a spatial-spectral interactive co-training CD module is proposed to achieve iterative interaction between cross-modal information, which jointly extracts the multimodal difference features to generate the final results. The proposed method outperforms several leading CD methods in extensive experiments carried out on both real and synthetic datasets. In addition, a new public HSI-RGB multimodal dataset along with our code are available at https://github.com/Jiahuiqu/CT2Net.
Wenqian Dong, Junying Ren, Song Xiao 0001, Leyuan Fang, Jiahui Qu, Yunsong Li 0001
IEEE Trans. Image Process.3
2025 Cycle-Based Frequency Disentanglement Diffusion Model With Self-Training for Cross-Domain Hyperspectral-RGB Change Detection
abstract
Hyperspectral images (HSI) change detection (CD) has become a powerful tool to analyze the sublte surface changes. However, the application of HSI CD is constrained by the limited availability of homogeneous HSIs. HSI-RGB multimodal CD address these limitations by collaboratively utilizing multi-source data. Although multimodal CD methods have achieved encouraging results, their performance often relies on the assumption that the training and test samples have similar distributions. Recently, some domain adaptive CD methods have been introduced. However, the additional modality differences in cross-domain multimodal CD pose challenges to existing domain adaptation techniques. To address these challenges, we propose a cycle-based frequency disentanglement diffusion model with self-training for cross-domain HSI-RGB multimodal CD, which explores a frequency-domain diffusion-driven self-training mechanism to enhance consistency in change representations across different modalities and domains. Specifically, a cyclic frequency domain disentanglement-based modality-domain alignment diffusion network is proposed to achieve modality and domain alignment within a unified diffusion framework. Subsequently, a curriculum-learning based self-training dual-domain CD network is designed to process the aligned images, which leverages pseudo-label reliability to ensure stable transfer of prior knowledge while exploits complementary features across modalities for collaborative CD. Extensive experiments demonstrate that our method significantly outperforms state-of-the-art approaches in cross-domain multimodal CD tasks.
Jiahui Qu, Junying Ren, Wenqian Dong, Song Xiao 0001, Yunsong Li 0001
IEEE Trans. Image Process.4
2025 Learning Generalization From Various Unaware Degradations for Blind Hyperspectral Image Super-Resolution via Transparent Diffusion Model
abstract
Hyperspectral image (HSI) super-resolution through the fusion of low-resolution HSI (LrHSI) and high-resolution multispectral image (HrMSI) has emerged as a critical technique for enhancing the quality of HSIs. The recent progress in this field predominantly assume a known mapping relationships between high-resolution HSI (HrHSI) and low-resolution version, relying on networks to learn this mapping to generate HrHSI. However, this assumption is often unrealistic in practical applications. To address this limitation, we propose the Spatial-Spectral-Integrated Transparent Diffusion Model (S2TD) for blind HSI-SR, which is more adaptive to scene-variant degradations with a universal framework for both spatial and spectral reconstruction. Specifically, we design a multi-order degradation pool to generate diverse samples, thereby reducing the distribution gap between low-resolution images in real scenes. Additionally, we develop a spatial-spectral consistent degradation model, which is iteratively solved using an optimization algorithm and unrolled into neural networks for separate restoration in spatial and spectral aspects. Furthermore, the capabilty of progressive reconstruction in the diffusion model is involved to fit various degradations in different dimensions using similar network architectures, thereby enhancing the overall robustness of the network to various and complex scenarios. Comprehensive experiments conducted on three publicly synthetic datasets and one real-world dataset validate the superior performance of the proposed method under the condition that the degradation remains unknown.
Jiahui Qu, Song Xiao 0001, Wenqian Dong, Yunsong Li 0001
IEEE Trans. Multim.3
2025 A Principle Design of Registration-Fusion Consistency: Toward Interpretable Deep Unregistered Hyperspectral Image Fusion
abstract
For hyperspectral image (HSI) and multispectral image (MSI) fusion, it is often overlooked that multisource images acquired under different imaging conditions are difficult to be perfectly registered. Although some works attempt to fuse unregistered images, two thorny challenges remain. One is that registration and fusion are usually modeled as two independent tasks, and there is no yet a unified physical model to tightly couple them. Another is that deep learning (DL)-based methods may lack sufficient interpretability and generalization. In response to the above challenges, we propose an unregistered HSI fusion framework energized by a unified model of registration and fusion. First, a novel registration-fusion consistency physical perception model (RFCM) is designed, which uniformly models the image registration and fusion problem to greatly reduce the sensitivity of fusion performance to registration accuracy. Then, an HSI fusion framework (MoE-PNP) is proposed to learn the knowledge reasoning process for solving RFCM. Each basic module of MoE-PNP one-to-one corresponds to the operation in the optimization algorithm of RFCM, which can ensure clear interpretability of the network. Moreover, MoE-PNP captures the general fusion principle for different unregistered images and therefore has good generalization. Extensive experiments demonstrate that MoE-PNP achieves state-of-the-art performance for unregistered HSI and MSI fusion. The code is available at https://github.com/Jiahuiqu/MoE-PNP.
Jiahui Qu, Jizhou Cui, Wenqian Dong, Qian Du 0001, Song Xiao 0001, Yunsong Li 0001
IEEE Trans. Neural Networks Learn. Syst.6
2024 Fusion from a Distributional Perspective: A Unified Symbiotic Diffusion Framework for Any Multisource Remote Sensing Data Classification
Song Xiao 0001, Wenqian Dong, Jiahui Qu, Yueguang Yang
IJCAI2
2024 Language-Guided Visual Prompt Compensation for Multi-Modal Remote Sensing Image Classification with Modality Absence
Ling Huang 0009, Wenqian Dong, Song Xiao 0001, Jiahui Qu, Yunsong Li 0001
ACM Multimedia3
2024 Adaptive Communication Resource Allocation for Federated Learning with UEP Strategies
abstract
This paper proposes an adaptive communication resource allocation algorithm to address the model transfer challenge in federated learning (FL). Our approach dynamically adjusts transmission power and channel coding rates based on FL training states, utilizing a bit-position-aware transmission scheme rooted in unequal error protection (UEP) principles. The system's performance is analytically assessed through a derived upper bound on the FL convergence rate, considering model transmission accuracy and client gradient diversities in heterogeneous data distributions. Then a resource allocation problem, aimed at minimizing the proposed upper bound, is decomposed into two sub-problems and solved. The first sub-problem, minimizing a derived upper bound on model transmission error, is tackled using a proposed algorithm for adaptive bit power allocation. The second sub-problem, minimizing cumulative gradient diversities, is formulated as a Markov decision process (MDP) and solved using deep Q-learning. Numerical evaluations show that our method outperforms vanilla and existing UEP-coded algorithms.
Muhang Lan, Song Xiao 0001, Wenyi Zhang 0001
VTC Spring2
2024 CFMDM: Coarse-to-Fine Meta-Diffusion Model for Scale-Arbitrary Hyperspectral Super-Resolution
abstract
Hyperspectral image super-resolution (HSISR) has shown very promising potential for earth observation and deep space exploration tasks. However, most existing HSISR methods formulate HSISR tasks with different scale factors as independent tasks, and train a specific model for each scale factor. In this letter, we propose a coarse-to-fine meta diffusion HSISR method, termed as CFMDM, which is capable of solving the problem of HSISR with scale-arbitrary factors in a unified model. The proposed CFMDM is composed of a coarse-to-fine upsampling module. The module encompasses two pivotal units: a coarse meta upsampling unit that utilizes meta-learning to map features of arbitrary scales to the corresponding scales, and a gradual refinement diffusion unit, which is designed to refine the details of the reconstructed HSI. In addition, we develop an imaging model-driven downsampling algorithm for generating training samples tailored to practical applications. The proposed method performs well in both quantitative and qualitative evaluation on benchmark datasets, achieving the average PSNR of 41.45dB at 1.5x super-resolution for the CAVE dataset.
Jizhou Cui, Wenqian Dong, Jiahui Qu, Song Xiao 0001, Yunsong Li 0001
IEEE Geosci. Remote. Sens. Lett.5
2024 Deep Spatial - Spectral Joint-Sparse Prior Encoding Network for Hyperspectral Target Detection
abstract
Hyperspectral target detection aims to locate targets of interest in the scene, and deep learning-based detection methods have achieved the best results. However, black box network architectures are usually designed to directly learn the mapping between the original image and the discriminative features in a single data-driven manner, a choice that lacks sufficient interpretability. On the contrary, this article proposes a novel deep spatial-spectral joint-sparse prior encoding network (JSPEN), which reasonably embeds the domain knowledge of hyperspectral target detection into the neural network, and has explicit interpretability. In JSPEN, the sparse encoded prior information with spatial-spectral constraints is learned end-to-end from hyperspectral images (HSIs). Specifically, an adaptive joint spatial-spectral sparse model (AS2JSM) is developed to mine the spatial-spectral correlation of HSIs and improves the accuracy of data representation. An optimization algorithm is designed for iteratively solving AS2JSM, and JSPEN is proposed to simulate the iterative optimization process in the algorithm. Each basic module of JSPEN one-to-one corresponds to the operation in the optimization algorithm so that each intermediate result in the network has a clear explanation, which is convenient for intuitive analysis of the operation of the network. With end-to-end training, JSPEN can automatically capture the general sparse properties of HSIs and faithfully characterize the features of background and target. Experimental results verify the effectiveness and accuracy of the proposed method. Code is available at https://github.com/Jiahuiqu/JSPEN.
Wenqian Dong, Jiahui Qu, Paolo Gamba, Song Xiao 0001, Anna Vizziello, Yunsong Li 0001
IEEE Trans. Cybern.5
2024 Incremental Detection of Hyperspectral Targets in Consistent Scenes With Continuous Learning
abstract
Hyperspectral target detection is a binary classification problem of detecting targets by utilizing the spectral characteristics of hyperspectral images (HSIs). However, in real-world applications, there is more than one class of interest in the same scene. The accuracy of detection for the previously learned target classes may be decreased when the model is retrained for detecting new target classes in scenes, which is called catastrophic forgetting. Consequently, how to ensure that the model has high detection performance for previously learned targets while learning new ones has become a key challenge for hyperspectral multitarget detection tasks. In this article, we propose an incremental detection of hyperspectral targets (IDHTs) method based on continual learning. IDHT decomposes the multitarget detection task into a series of independent subtasks and learns them sequentially. Within each subtask, our proposed incremental spectral detector (ISD) enables the training and learning of new class targets. Simultaneously, we introduce the previous label replay strategy (PLRS), which synthesizes fused labels for the current task training by combining detection outcomes from the previous model with pseudo-labels of the current target. PLRS effectively bridges the knowledge gap across various subtasks in hyperspectral multitarget detection tasks. The proposed IDHT can flexibly and dynamically adapt to new categories and overcome the limitations of fixed-category feature learning. In addition, two hyperspectral datasets are disclosed to evaluate the proposed method. Our method demonstrates significant effectiveness and superiority on both public datasets and two self-collected datasets. Our code and dataset are available athttps://github.com/Jiahuiqu/IDHT.
Wenqian Dong, Song Xiao 0001, Jiahui Qu, Yunsong Li 0001
IEEE Trans. Geosci. Remote. Sens.3
2024 ISPDiff: Interpretable Scale-Propelled Diffusion Model for Hyperspectral Image Super-Resolution
abstract
Hyperspectral image (HSI) super-resolution (SR) employing the denoising diffusion probabilistic model (DDPM) holds significant promise with its remarkable performance. However, existing relevant works exhibit two limitations: i) Directly applying DDPM to fusion-based HSI SR (HSI-SR) ignores the physical mechanism of HSI-SR and unique characteristics of HSI, resulting in less interpretability; ii) Scale-invariant DDPM suffers from a time-consuming inference. To tackle these issues, we propose an interpretable scale-propelled diffusion model (ISPDiff) for HSI-SR, which combines the underlying principles of HSI-SR with DDPM for progressively unrolling reconstruction by learning its distribution at various scales, enhancing the transparency significantly and reducing the inference time prominently. Concretely, we destroy and downsample HSI into Gaussian noise in the forward process of ISPDiff. Then we design a unified scale-flexible model in the backward process to iteratively refine HSI in a coarse-to-fine manner through scale-matched reconstruction and cross-scale upsampling, which can be unfolded with optimization algorithms. These solved equations are one-to-one corresponding unrolled into two deep neural networks, called progressive perceptual model-driven scale-matched restoration network (P2MSRN) and cross-scale model-driven upsampling network (CMUN). Through end-to-end training, the proposed ISPDiff implements HSI-SR with a scale-propelled unrolling diffusion characterized by enhanced interpretability, stronger task orientation, and reduced time consumption. Systematic experiments have been conducted on three public datasets, demonstrating that ISPDiff outperforms state-of-the-art methods. Code is available at https://github.com/Jiahuiqu/ISPDiff.
Wenqian Dong, Song Xiao 0001, Jiahui Qu, Yunsong Li 0001
IEEE Trans. Geosci. Remote. Sens.3
2024 Contrastive Constrained Cross-Scene Model- Informed Interpretable Classification Strategy for Hyperspectral and LiDAR Data
abstract
Domain adaptation (DA) aims to transfer knowledge from a labeled source domain (SD) to an unlabeled target domain (TD), and its effectiveness has been demonstrated in unsupervised multisource remote sensing image classification. Existing DA frameworks simultaneously learn the mapping between data in SD and category labels, as well as minimize the distribution discrepancy between different domains. However, significant computational resources are needed to optimize the dual objectives in a data-driven DA network. In addition, the lack of interpretability of deep learning (DL)-based methods results in unpredictable feature distributions, thereby impeding the smooth update of the network in the desired direction. To address these issues, we propose a contrastive constrained cross-scene model-informed interpretable classification strategy (C3MI-C) for hyperspectral image (HSI) and light detection and ranging (LiDAR), which achieves a model-interpretable decoupling of domain adaptive task and classification task. The proposed C3MI-C optimizes the classification network interpretably in the same subspace and further aligns deep-adapted features extracted from two domains to accomplish high-precision unsupervised cross-scene classification. Comparative experiment results and ablation studies show that C3MI-C performs better than other advanced methods.
Wenqian Dong, Jiahui Qu, Tian Zhang 0017, Song Xiao 0001, Yunsong Li 0001
IEEE Trans. Geosci. Remote. Sens.4
2024 Feature Mutual Representation-Based Graph Domain Adaptive Network for Unsupervised Hyperspectral Change Detection
abstract
Recently, deep neural networks (DNNs) have been widely used in hyperspectral image change detection (HSI-CD). Generally, training such a DNN-based HSI-CD network often requires a large number of labeled training samples. However, it is time-consuming, labor-intensive, or even infeasible to label training samples in practice. In this article, we propose a feature mutual representation-based graph domain adaptive network (FGDANet) for unsupervised HSI-CD. This method constructs a pseudosiamese backbone consisting of two customized unsupervised learning domains, which can make full use of the information from different domains through the graph domain adaptation strategy to improve the feature expression capability and generalization. There are three key characteristics: first, in each customized unsupervised learning domain, a graph convolutional network (GCN)-based difference feature extraction architecture is designed to model the local and global dependence among the features of multitemporal HSIs; second, a progressive graph-to-pixel joint constraint strategy (PJCS) is proposed to provide the high-confidence training sample labels for the unsupervised learning of the network in each domain; and third, the homogeneous mutual representation joint graph feature alignment (HJGFA) module of the graph domain adaptation strategy can make full use of the difference features from the two domains through the information interaction to facilitate the model to capture the changed and unchanged essential characteristics. The experimental results on four HSI datasets demonstrate the superiority of the proposed FGDANet. Code is available athttps://github.com/Jiahuiqu/FGDANet.
Jiahui Qu, Jingyu Zhao 0011, Wenqian Dong, Song Xiao 0001, Yunsong Li 0001, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.4
2024 Parallel Compared-and-Stacked Pyramid Transformer Network for Unsupervised Hyperspectral Change Detection
abstract
Convolutional neural networks (CNNs) with good feature learning capabilities are widely used in hyperspectral image change detection (HSI-CD) tasks. However, most existing CNN-based HSI-CD methods face two inherent challenges: 1) the lack of available labeled datasets and 2) the limited receptive field that cannot capture the long-distance dependence between the spectral sequences of HSIs. In this article, we propose a parallel compared-and-stacked pyramid transformer network (PCPTNet) for unsupervised HSI-CD, which can model the context information of spectral sequences in the input multi-temporal HSI patch without real labeled data. Specifically, a superpixel-level joint decision-based training samples selection strategy is presented that fully considers the correlation between pixels to improve the reliability of training samples. Then, taking advantage of transformer in context information modeling, PCPTNet is proposed to capture sufficient difference features and stacked features with different scales for CD, which can effectively reduce missed and false detection. The multiscale features containing sufficient low-level detail information and high-level semantic features are fused hierarchically to classify changed and unchanged pixels. Extensive experiments on three real HSI datasets demonstrate that the PCPTNet outperforms other state-of-the-art HSI-CD methods in both visual and quantitative results.
Yunshuang Xu, Song Xiao 0001, Jiahui Qu, Wenqian Dong, Yunsong Li 0001, Haoming Xia
IEEE Trans. Geosci. Remote. Sens.2
2024 TMCFN: Text-Supervised Multidimensional Contrastive Fusion Network for Hyperspectral and LiDAR Classification
abstract
The joint classification of hyperspectral images (HSIs) and LiDAR data plays a crucial role in earth observation missions. Most advanced methods are based on discrete label supervision. However, since discrete labels only convey limited information that a sample belongs to a single definite class and lack of prior information, it is difficult to supervise the model to capture rich inherent semantic information in complex data distributions, hindering the classification performance. To this end, we propose a text-supervised multidimensional contrastive fusion network, termed as TMCFN, which leverages class text information to guide the learning of visual representations while establishing a semantic association of text and visual features for classification by using multidimensionally incorporated contrastive learning (CL) paradigms. Specifically, TMCFN is composed of text information encoding (TIE), visual features representation (VFR) and text-visual features alignment and classification (TVFAC). TIE is employed to extract semantic information from class text extended from class names, intrinsic attributes and inter-class relationships. VFR mainly comprises a new fusion-based contrastive feature learning module (FCFLM) to extract discriminative visual features and a text-guided attention feature fusion module (TAF2M) to fuse visual features under the guidance of text information. TVFAC optimizes the learning of visual features under the supervision of text information while using a CL paradigm to align text and visual features for establishing the semantic association, and achieves the classification by directly computing the similarity between the visual features and each text feature without an additional classifier. Experiments with three standard datasets verify the effectiveness of TMCFN.
Yueguang Yang, Jiahui Qu, Wenqian Dong, Tongzhen Zhang, Song Xiao 0001, Yunsong Li 0001
IEEE Trans. Geosci. Remote. Sens.5
2024 Graph Embedding Interclass Relation-Aware Adaptive Network for Cross-Scene Classification of Multisource Remote Sensing Data
abstract
The unsupervised domain adaptation (UDA) based cross-scene remote sensing image classification has recently become an appealing research topic, since it is a valid solution to unsupervised scene classification by exploiting well-labeled data from another scene. Despite its good performance in reducing domain shifts, UDA in multisource data scenarios is hindered by several critical challenges. The first one is the heterogeneity inherent in multisource data complicates domain alignment. The second challenge is the incomplete representation of feature distribution caused by the neglect of the contribution from global information. The third challenge is the inaccuracies in alignment due to errors in establishing target domain conditional distributions. Since UDA does not guarantee the complete consistency of the distribution of the two domains, networks using simple classifiers are still affected by domain shifts, resulting in poor performance. In this paper, we propose a graph embedding interclass relation-aware adaptive network (GeIraA-Net) for unsupervised classification of multi-source remote sensing data, which facilitates knowledge transfer at the class level for two domains by leveraging aligned features to perceive inter-class relation. More specifically, a graph-based progressive hierarchical feature extraction network is constructed, capable of capturing both local and global features of multisource data, thereby consolidating comprehensive domain information within a unified feature space. To deal with the imprecise alignment of data distribution, a joint de-scrambling alignment strategy is designed to utilize the features obtained by a three-step pseudo-label generation module for more delicate domain calibration. Moreover, an adaptive inter-class topology based classifier is constructed to further improve the classification accuracy by making the classifier domain adaptive at the category level. The experimental results show that GeIraA-Net has significant advantages over the current state-of-the-art cross-scene classification methods.
Song Xiao 0001, Jiahui Qu, Wenqian Dong, Qian Du 0001, Yunsong Li 0001
IEEE Trans. Image Process.2
2024 A Spatio-Spectral Fusion Method for Hyperspectral Images Using Residual Hyper-Dense Network
abstract
Spatio-spectral fusion of panchromatic (PAN) and hyperspectral (HS) images is of great importance in improving spatial resolution of images acquired by many commercial HS sensors. DenseNets have recently achieved great success for image super-resolution because they facilitate gradient flow by concatenating all the feature outputs in a feedforward manner. In this article, we propose a residual hyper-dense network (RHDN) that extends the DenseNet to solve the spatio-spectral fusion problem. The overall structure of the proposed RHDN method is a two-branch network, which allows the network to capture the features of HS images within and outside the visible range separately. At each branch of the network, a two-stream strategy of feature extraction is designed to process PAN and HS images individually. A convolutional neural network (CNN) with cascade residual hyper-dense blocks (RHDBs), which allows direct connections between the pairs of layers within the same stream and those across different streams, is proposed to learn more complex combinations between the HS and PAN images. The residual learning is adopted to make the network efficient. Extensive benchmark evaluations well demonstrate that the proposed RHDN fusion method yields significant improvements over many widely accepted state-of-the-art approaches.
Jiahui Qu, Zhangchun Xu, Wenqian Dong, Song Xiao 0001, Yunsong Li 0001, Qian Du 0001
IEEE Trans. Neural Networks Learn. Syst.4
2023 Component Substitution Model-Guided Deep Spatial-Spectral Fusion of Hyperspectral Imagery
abstract
Hyperspectral (HS) spatial-spectral fusion technology is an important means to obtain high spatial and spectral resolution data for major strategic missions such as manned spaceflight and earth observation. The existing parameter-tuning black-box fusion model lacks the guidance of mathematical theory and is not explicable, so it is difficult to apply in practice. In this letter, we design an interpretable deep network for spatial-spectral fusion tasks, called S2Fusion, which embedded the mature component substitution (CS) fusion model into a deep neural network. As a result of the careful design of the method, each module in the S2Fusion network corresponds to a specific operation in the CS fusion model, which is easily interpretable. Compared with the traditional fusion models, the S2Fusion makes it easier to ensure the implementation mechanism of spatial-spectral fusion throughout the network flow. Extensive experiments demonstrate the superiority of S2Fusion both quantitatively and visually over state-of-the-art methods.
Song Xiao 0001, Wenqian Dong, Jiahui Qu
IEEE Geosci. Remote. Sens. Lett.2
2023 Dictionary Learning-Guided Deep Interpretable Network for Hyperspectral Change Detection
abstract
Hyperspectral image (HSI) change detection is a technique to observe the change information between the multitemporal HSIs, which is currently considered a major focus of research in the filed of remote sensing intelligent interpretation. Most existing deep learning-based methods have created satisfactory performance, but these methods lack transparency and have poor generalization. To tackle the problems outlined above, we propose a dictionary learning-guided deep interpretable network for hyperspectral change detection, which unfolds a dictionary learning-based change detection model into an interpretable deep neural network. Specifically, we first design a dictionary learning-based change detection model, whose solution process can be decomposed into two iterative subproblems. Then, the mathematical model can be unfolded into a dual-branch deep neural network with two modules iterating with each other. Finally, the difference map of the coefficients output from the ultimate stage is classified to obtain the change detection result. Experimental results prove that the proposed method has comparable or even better performance than state-of-the-art methods.
Jingyu Zhao 0011, Song Xiao 0001, Wenqian Dong, Jiahui Qu, Yunsong Li 0001
IEEE Geosci. Remote. Sens. Lett.2
2023 Joint Contextual Representation Model-Informed Interpretable Network With Dictionary Aligning for Hyperspectral and LiDAR Classification
abstract
The effective utilization of hyperspectral image (HSI) and light detection and ranging (LiDAR) data is essential for land cover classification. Recently, deep learning-based classification approaches have achieved remarkable success. However, most deep learning classification methods are data-driven and designed in a black-box architecture, lacking sufficient interpretability, and ignoring the potential correlation of heterogeneous complementary information between multisource data. To address these issues, we propose an interpretable deep neural network, namely multisource aligning joint contextual representation model-informed interpretable classification network (MACRMoI-N), which fully exploits correlation of multisource data by aligning complementary spectral-spatial-elevation information during end-to-end training. We first present a multimodal aligning joint contextual representation classification model (MACR-M), which incorporates local spatial-spectral prior information into representation. MACR-M is optimized by an iterative algorithm to solve dictionaries of HSI and LiDAR and their corresponding sparse coefficients, in which the dictionary distribution are aligned to enable the complementary information of multisource data to guide a more accurate classification. We further propose the unfolded MACRMoI-N, where each module corresponds to a specific operation of the optimization algorithm, and the parameters are optimized in an end-to-end manner. Comparative experiment results and ablation studies show that MACRMoI-N performs better than other advanced methods.
Wenqian Dong, Jiahui Qu, Tian Zhang 0017, Song Xiao 0001, Yunsong Li 0001
IEEE Trans. Circuits Syst. Video Technol.5
2023 Local Information-Enhanced Graph-Transformer for Hyperspectral Image Change Detection With Limited Training Samples
abstract
Hyperspectral image (HSI) change detection is a challenging task that focuses on identifying the differences between multi-temporal HSIs. The recent advancement of convolutional neural network (CNN) has made great progress on HSIs change detection. However, due to the limited receptive field, most CNN based change detection models trained with sufficient labeled samples cannot flexibly model the global information that is essential for distinguishing complex objects, thereby achieving relatively-low performance. In this paper, we propose a dual-branch local information enhanced graph-transformer change detection network to fully exploit the local-global spectral-spatial features of the multi-temporal HSIs with limited training samples for change recognition. Specifically, the proposed network is composed of a cascaded of local information enhanced graph-transformer (LIEG) blocks, which jointly extracts local-global features by learning local information representation to enhance the information of graph-transformer. A novel graph-transformer is developed to model global spectral–spatial correlation between graph nodes, enabling the spectral information preservation of HSIs and accurate change detection of areas with various sizes. Extensive experiments have proved that our method achieves significant performance improvement than other state-of-the-art methods on four commonly used HSI datasets.
Wenqian Dong, Jiahui Qu, Song Xiao 0001, Yunsong Li 0001
IEEE Trans. Geosci. Remote. Sens.4
2023 Abundance Matrix Correlation Analysis Network Based on Hierarchical Multihead Self-Cross-Hybrid Attention for Hyperspectral Change Detection
abstract
Hyperspectral image (HSI) change detection is a technique for detecting the changes between the multitemporal HSIs of the same scene. Many existing change detection methods have achieved good results, but there still exist problems as follows: 1) mixed pixels exist in HSI due to the low spatial resolution of hyperspectral sensor and other external interference and 2) many existing deep learning-based networks cannot make full use of the correlation difference information between the bitemporal images. These problems are not conducive to further improving the accuracy of change detection. In this article, we propose an abundance matrix correlation analysis network based on hierarchical multihead self-cross-hybrid attention (AMCAN-HMSchA) for HSI change detection, which hierarchically highlights the correlation difference information at the subpixel level to detect the subtle changes. The endmember sharing-based abundance matrix learning module (AMLM) maps the changed information between bitemporal HSIs to the corresponding abundance matrices. The hierarchical MSchA extracts the enhanced difference features by constantly comparing the self-correlation with cross correlation between the abundance matrices of the HSIs. Then, the difference features are concatenated and fed into the fully connected layers to obtain the change map. Experiments on three widely used datasets show that the proposed method has superior performance compared with other state-of-the-art methods.
Wenqian Dong, Jingyu Zhao 0011, Jiahui Qu, Song Xiao 0001, Shaoxiong Hou, Yunsong Li 0001
IEEE Trans. Geosci. Remote. Sens.4
2023 Noise Prior Knowledge Informed Bayesian Inference Network for Hyperspectral Super-Resolution
abstract
Well-known deep learning (DL) is widely used in fusion based hyperspectral image super-resolution (HS-SR). However, DL-based HS-SR models have been designed mostly using off-the-shelf components from current deep learning toolkits, which lead to two inherent challenges: i) they have largely ignored the prior information contained in the observed images, which may cause the output of the network to deviate from the general prior configuration; ii) they are not specifically designed for HS-SR, making it hard to intuitively understand its implementation mechanism and therefore uninterpretable. In this paper, we propose a noise prior knowledge informed Bayesian inference network for HS-SR. Instead of designing a "black-box" deep model, our proposed network, termed as BayeSR, reasonably embeds the Bayesian inference with the Gaussian noise prior assumption to the deep neural network. In particular, we first construct a Bayesian inference model with the Gaussian noise prior assumption that can be solved iteratively by the proximal gradient algorithm, and then convert each operator involved in the iterative algorithm into a specific form of network connection to construct an unfolding network. In the process of network unfolding, based on the characteristics of the noise matrix, we ingeniously convert the diagonal noise matrix operation which represents the noise variance of each band into the channel attention. As a result, the proposed BayeSR explicitly encodes the prior knowledge possessed by the observed images and considers the intrinsic generation mechanism of HS-SR through the whole network flow. Qualitative and quantitative experimental results demonstrate the superiority of the proposed BayeSR against some state-of-the-art methods.
Wenqian Dong, Jiahui Qu, Song Xiao 0001, Tongzhen Zhang, Yunsong Li 0001, Xiuping Jia
IEEE Trans. Image Process.3
2023 Quantization Bits Allocation for Wireless Federated Learning
abstract
Federated learning (FL) enables multiple clients to collaborate on a common learning task via only exchanging model updates. With the progressive improvements in deep learning models, communication is becoming a primary bottleneck of FL. Quantization of model updates before transmitting is an effective technique to reduce communication overhead. Most prior literature assumes lossless transmission, but in practice, quantized model updates are distorted by wireless channels due to the variation of client locations. Therefore, this paper focuses on analysis and design of personalized model update quantization with explicitly incorporating channel diversity in wireless FL. We present a novel convergence analysis of quantized FL, which encompasses full and partial client participation, single and multiple local training iterations, and convex and non-convex loss functions. This analysis explicitly embodies the impact of personalized quantization error, channel diversity and model aggregation in FL, and also elucidates their tradeoff on tightening a convergence rate upper bound. An optimization framework, which seeks an optimal allocation scheme given a total budget of quantization bits, is proposed by minimizing an upper bound with respect to channel quality. A nearly optimal solution is derived for this non-convex integer programming problem via analytically solving Karush–Kuhn–Tucker (KKT) optimality conditions and linear search. From a perspective of outlier detection, this channel-aware allocation scheme is also extended to robust model aggregation against client dropouts. Comprehensive numerical evaluation demonstrates the performance enhancement of the proposed scheme over the vanilla allocation scheme with equal quantization bits, particularly in terms of training stability, test accuracy, and robustness.
Muhang Lan, Qing Ling 0001, Song Xiao 0001, Wenyi Zhang 0001
IEEE Trans. Wirel. Commun.3
2022 Multi-level features fusion via cross-layer guided attention for hyperspectral pansharpening
Shaoxiong Hou, Song Xiao 0001, Wenqian Dong, Jiahui Qu
Neurocomputing2
2022 Hyperspectral Pansharpening via Local Intensity Component and Local Injection Gain Estimation
abstract
Hyperspectral (HS) pansharpening is an attractive topic in the field of remote sensing, which has attracted the attention of many researchers. Component substitution (CS)-based HS pansharpening algorithms are of great interest due to their simplicity and high spatial quality, and they mainly consist of two phases: detail extraction and detail injection. Detail extraction is performed by estimating the intensity component, whereas detail injection depends on the definition of injection gain. In the classic CS-based pansharpening methods, the intensity component is estimated through a global synthesis scheme, and injection gains can be obtained by a context-adaptive or a global approach. In this letter, we propose an improved CS-based HS pansharpening method in which the intensity component and the injection gain are estimated locally achieved by the binary partition tree (BPT) image segmentation algorithm. The proposed method is applied to two credible CS-based HS pansharpening algorithms, including the Gram–Schmidt adaptive (GSA) and the Brovey transform (Brovey). The experimental results show that the proposed method improves the performance of GSA and Brovey and creates promising results perceptually and quantitatively.
Wenqian Dong, Jiahui Qu, Song Xiao 0001, Qian Du 0001
IEEE Geosci. Remote. Sens. Lett.4
2022 A Mutual Guidance Attention-Based Multi-Level Fusion Network for Hyperspectral and LiDAR Classification
abstract
Hyperspectral image (HSI) and light detection and ranging (LiDAR) data classification has attracted more and more attention in remote sensing. Convolution neural network (CNN) has been proven to be effective for HSI and LiDAR data classification. In this letter, a novel three-branch CNN is designed to learn spectral, spatial, and elevation features, each of which adopts the multi-level feature fusion (MLF) module to fuse the shallow and deep features. Furthermore, in order to fully fuse the spatial and elevation information, we propose a mutual guidance attention (MGA) module. The MGA module increases the information flow between spatial and elevation branches, highlights the features of interest, and weakens useless features. The proposed method is evaluated on public datasets Houston and Trento. Experimental results demonstrate that our proposed method can provide higher classification accuracy than some existing methods.
Tongzhen Zhang, Song Xiao 0001, Wenqian Dong, Jiahui Qu
IEEE Geosci. Remote. Sens. Lett.2
2022 Laplacian Pyramid Dense Network for Hyperspectral Pansharpening
abstract
Hyperspectral (HS) pansharpening aims to create a pansharpened image that integrates the spatial details of the panchromatic (PAN) image and the spectral content of the HS image. In this article, we present a deep convolutional network within the mature Gaussian–Laplacian pyramid for pansharpening (LPPNet). The overall structure of LPPNet is a cascade of the Laplacian pyramid dense network with a similar structure at each pyramid level. Following the general idea of multiresolution analysis (MRA), the subband residuals of the desired HS images are extracted from the PAN image and injected into the upsampled HS image to reconstruct the high-resolution HS images level by level. Applying the mature Laplace pyramid decomposition technique to the convolution neural network (CNN) can simplify the pansharpening problem into several pyramid-level learning problems so that the pansharpening problem can be solved with a shallow CNN with fewer parameters. Specifically, the Laplacian pyramid technology is used to decompose the image into different levels that can differentiate large- and small-scale details, and each level is handled by a spatial subnetwork in a divide-and-conquer way to make the network more efficient. Experimental results show that the proposed LPPNet method performs favorably against some state-of-the-art pansharpening methods in terms of objective indexes and subjective visual appearance.
Wenqian Dong, Tongzhen Zhang, Jiahui Qu, Song Xiao 0001, Jie Liang 0001, Yunsong Li 0001
IEEE Trans. Geosci. Remote. Sens.4
2022 Multibranch Feature Fusion Network With Self- and Cross-Guided Attention for Hyperspectral and LiDAR Classification
abstract
The effective fusion of multi-source data helps to improve performance of land cover classification. Most existing convolutional neural network (CNN) based methods adopt an early/late fusion strategy to fuse the low-level/high-level features for classification, which still has two inherent challenges: i) the conventional convolution operation performs a weighted average operation on each pixel in the receptive field, which will reduce the discriminability of the center pixel due to the influence of the interference pixels, and ii) the spatial-spectral features of the hyperspectral image (HSI), the elevation features of light detection and ranging (LiDAR), and the complementary features between the multimodal data are not fully exploited, which results in the reduction of classification accuracy. In this paper, an effective multi-branch feature fusion network with self- and cross-guided attention (MB2FscgaNet) is proposed for joint classification of LiDAR and HSI. The main concern of this paper is how to accurately estimate more effective spectral-spatial-elevation features and yield more effective transfer in network. Specifically, MB2FscgaNet adopts a multi-branch feature fusion architecture to fully exploit the hierarchical features from LiDAR and HSI level by level. At each level of the network, a self- and cross-guided attention (SCGA) is developed to assign higher weight to interesting areas and channels of LiDAR and HSI feature maps to obtain refined spectral-spatial-elevation features and provide complementary information cross guidance between LiDAR and HS. We further designed a spectral supplement module (SeSuM) to improve the discriminative ability of the center pixel. Comparative classification results and ablation studies demonstrate that the proposed MB2FscgaNet achieves competitive performance against state-of-the-art methods.
Wenqian Dong, Tian Zhang 0017, Jiahui Qu, Song Xiao 0001, Tongzhen Zhang, Yunsong Li 0001
IEEE Trans. Geosci. Remote. Sens.4
2022 A Dual-Branch Detail Extraction Network for Hyperspectral Pansharpening
abstract
Hyperspectral (HS) pansharpening aims at creating a high-resolution hyperspectral (HR-HS) image by integrating a high spatial resolution panchromatic (HR-PAN) image with a low-resolution hyperspectral (LR-HS) image. It is an important preprocessing procedure in many remote sensing tasks. Most of the existing pansharpening methods train a specific convolutional neural network (CNN) model for each type of dataset with the same number of spectral bands. The main contribution of this study is to propose a new dual-branch detail extraction pansharpening network (called DBDENet) that can sharpen HS images with any number of spectral bands using a single pre-trained model by fine-tuning the parameters of a small module in the network. Specifically, DBDENet extracts spatial details from LR-HS and HR-PAN images by two bidirectional branches of the dual-branch detail extraction network level by level. For each level, the spatial details captured from the HR-PAN and those of the LR-HS images are fused by a spatial cross attention fusion module (SCAFM). The spatial details fused by the last SCAFM module are injected into the upsampled HS image to obtain an HR-HS image. Experimental results prove to show the proposed DBDENet is superior to other widely accepted state-of-the-art methods in terms of objective indicators and visual appearance.
Jiahui Qu, Shaoxiong Hou, Wenqian Dong, Song Xiao 0001, Qian Du 0001, Yunsong Li 0001
IEEE Trans. Geosci. Remote. Sens.4
2022 A Deep Multiscale Pyramid Network Enhanced With Spatial-Spectral Residual Attention for Hyperspectral Image Change Detection
abstract
Change detection plays an important role in Earth surface observation and has been extensively investigated over recent decades. A hyperspectral image (HSI) with high spectral resolution provides abundant ground object information, which is expected by finer change detection. The existing convolutional neural network (CNN)-based methods extract image features with a fixed kernel, which is incompetent to cope with complicated object details at diverse scales in HSI. In this article, we propose a deep multiscale pyramid network enhanced with spatial–spectral residual attention (DMP$\text {s}^{2} $raN) for HSI change detection, which has strong capability to mine multilevel and multiscale spatial–spectral features, improving the performance in complex changed regions. There are two key characteristics: 1) the multiscale spatial–spectral features are extracted by the multiscale pyramid convolution and enhanced by spatial–spectral residual attention module ($\text {S}^{2} $RAM) of each scale and 2) the multilevel features are obtained by aggregating the multiscale features level by level. As a result of this design, the proposed DMP$\text {s}^{2} $raN learns more discriminative features with both strong semantic information and rich spatial–spectral information. Experiments carried out on three datasets demonstrate the competitive performance of the proposed method in both qualitative and quantitative analyses.
Jiahui Qu, Song Xiao 0001, Wenqian Dong, Yunsong Li 0001, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.3
2022 Generative Dual-Adversarial Network With Spectral Fidelity and Spatial Enhancement for Hyperspectral Pansharpening
abstract
Hyperspectral (HS) pansharpening is of great importance in improving the spatial resolution of HS images for remote sensing tasks. HS image comprises abundant spectral contents, whereas panchromatic (PAN) image provides spatial information. HS pansharpening constitutes the possibility for providing the pansharpened image with both high spatial and spectral resolution. This article develops a specific pansharpening framework based on a generative dual-adversarial network (called PS-GDANet). Specifically, the pansharpening problem is formulated as a dual task that can be solved by a generative adversarial network (GAN) with two discriminators. The spatial discriminator forces the intensity component of the pansharpened image to be as consistent as possible with the PAN image, and the spectral discriminator helps to preserve spectral information of the original HS image. Instead of designing a deep network, PS-GDANet extends GANs to two discriminators and provides a high-resolution pansharpened image in a fraction of iterations. The experimental results demonstrate that PS-GDANet outperforms several widely accepted state-of-the-art pansharpening methods in terms of qualitative and quantitative assessment.
Wenqian Dong, Shaoxiong Hou, Song Xiao 0001, Jiahui Qu, Qian Du 0001, Yunsong Li 0001
IEEE Trans. Neural Networks Learn. Syst.3
2020 Fusion of hyperspectral and panchromatic images using structure tensor and matting model
Wenqian Dong, Song Xiao 0001, Jie Liang 0001, Jiahui Qu
Neurocomputing2
2020 Joint group and residual sparse coding for image compressive sensing
Lizhao Li, Song Xiao 0001
Neurocomputing2
2020 Saliency Analysis and Gaussian Mixture Model-Based Detail Extraction Algorithm for Hyperspectral Pansharpening
abstract
The purpose of hyperspectral (HS) pansharpening is to improve the spatial resolution of HS images using panchromatic (PAN) images so as to obtain the pansharpened images with both high spectral diversity and high spatial resolution. The classical component substitution (CS)-based pansharpening approaches can be decomposed into two sequential phases: detail extraction and detail injection. In general, the detail extraction is performed by computing the difference between the PAN image and a weighted average of the HS bands, whereas the detail injection depends on the injection gain which is defined locally or globally. In this article, we introduce a novel pansharpening algorithm in which the extracted details and the injection gain are estimated over salient and nonsalient regions achieved via saliency analysis and Gaussian mixture model. The proposed method is applied to four credible CS-based pansharpening methods and also compared to other state-of-the-art methods. Experimental results show that the modified CS methods have better performance than the original methods. In addition, our method also achieves comparable or better performance than the other state-of-the-art methods.
Wenqian Dong, Jie Liang 0001, Song Xiao 0001
IEEE Trans. Geosci. Remote. Sens.3
2019 An Energy Efficient Spatio-Temporal Compression Clustering Data Collection Scheme for Wireless Sensor Networks
abstract
Energy efficiency is a critical issue in data gathering, which involves conserving sensor nodes energy and maximizing network lifetime. In this paper, we propose a distributed clustering scheme with spatio-temporal compressed sensing to organize the sensors into clusters. In the process of inter- cluster transmission, a energy-aware transmission method based on network coding is used to further balance the network load and energy. Simulation results show that, compared with the algorithms have been proposed, our scheme can minimize the total energy spent in the network, make the remaining energy of the node more balanced and extend the lifetime of the system.
Lina Xiao, Song Xiao 0001
VTC Spring2
2018 Fusion of Hyperspectral and Panchromatic Images Based on Matting Model
abstract
In this paper, a novel hyperspectral (HS) image fusion method using matting model is presented. Matting model refers to each band of an HS image that can be decomposed into three components, i.e., alpha channel, spectral foreground, and spectral background. First, panchromatic (PAN) image is sharpened to enhance details, and the spatial information of each band of HS image is obtained by weighted least squares filtering. Different from traditional matting model based methods that PAN image is served as the alpha channel, we do the PCA transformation to PAN image and spatial information of each band to obtain the first principal component channel which is selected for the alpha channel. This processing reduces spatial distortion. Finally, HS foreground and HS background are estimated by the alpha channel, and the fused HS image is reconstructed nearly perfectly. Experiments reveal that the proposed method is superior to the state-of-the-art methods.
Wenqian Dong, Song Xiao 0001, Jiahui Qu, Hongping Gan
IGARSS2
2018 A New Hyperspectral Pansharpening Method With Intrisic Image Decomposition
abstract
The component substitution (CS) and multiresolution analysis (MRA) based methods have been well adopted in hyperspectral pansharpening. The major contribution of this paper is a novel MRA and CS hybrid framework based on the intrinsic image decomposition. First, the weighted least squares (WLS) filter is performed on the sharpened panchromatic (P) image to extract the high-frequency component. Then, the intrinsic image decomposition (IID) is adopted to decompose the interpolated hyperspectral (H) image into the illumination and reflectance components. Finally, the detail map is generated by making a proper compromise between the high-frequency component of the P image and the illumination component of the H image. The detail map further refined by the information ratio of different bands of the H image is injected into each band of the interpolated H image. Experimental results indicate that the proposed method achieves a better fusion result than several state-of-the-art hyperspectral pansharpening methods.
Wenqian Dong, Song Xiao 0001, Jiahui Qu
IGARSS2
2018 Broadcast Cost Reduction in Wireless Sensor Networks with Instantly Decodable Network Codes
abstract
Consider the cluster-based wireless sensor networks (WSNs) within a set of sensor nodes which need a reliable data broadcast, most recent works ignore the broadcast cost and the cache of sensor nodes when using instantly decodable network codes (IDNC). In this paper, we firstly propose a 2C-IDNC graph by taking advantage of Cache-IDNC graph and Cost-IDNC graph. To find a maximum weight clique in 2C-IDNC graph with smaller broadcast cost, we utilize the dynamic selection to achieve a proper broadcast cost, which can ensure the decode opportunities of encoded packets in WSNs. To this end, a heuristic algorithm is proposed to reduce the complexity of computation which operates on the novel 2C-IDNC graph. The simulation results show that the broadcast cost can be indeed reduced by the 2C- IDNC scheme.
Song Xiao 0001, Hongping Gan
VTC Spring2
2018 Hyperspectral pansharpening based on guided filter and Gaussian filter
Wenqian Dong, Song Xiao 0001, Yongxu Li
J. Vis. Commun. Image Represent.2
2018 A large class of chaotic sensing matrices for compressed sensing
Hongping Gan, Song Xiao 0001
Signal Process.2
2018 Construction of efficient and structural chaotic sensing matrix for compressive sensing
Hongping Gan, Song Xiao 0001, Xiao Xue 0003
Signal Process. Image Commun.2
2014 Comment on 'Sparse block circulant matrices for compressed sensing'
abstract
In ‘Sparse block circulant matrices for compressed sensing’, in order to apply Lemma 4, every off‐diagonal element of Gram matrix for the sparse block circulant matrix was separated into two component sums to make sure the terms in each sum are independent. In this comment, however, the authors show that separating every element into two sums is not sufficient to guarantee the independency of the terms in each sum. The authors also prove that the entries should be split into three parts instead of two to satisfy the requirements of Lemma 4. Finally, the authors modify the deduction and the result.
Lei Quan, Song Xiao 0001, Mengsi Wang
IET Commun.2
2011 A network coding based hybrid ARQ algorithm for wireless video broadcast
Ji Lu, Chengke Wu 0001, Song Xiao 0001, Jianchao Du
Sci. China Inf. Sci.3
2011 Random linear network coding with ladder-shaped global coding matrix for robust video transmission
Hui Wang 0005, Song Xiao 0001, C.-C. Jay Kuo
J. Vis. Commun. Image Represent.2
2010 A practical bit stream organization algorithm for robust H.264/SVC transmission
Song Xiao 0001, Hui Wang 0005, C.-C. Jay Kuo
J. Vis. Commun. Image Represent.1
2009 Efficient Protection Scheme for SVC Content Based on Network Coding
abstract
How to protect the video streaming for reasonable usage is a necessary technology in multimedia services over network. However, most the previous schemes are too complicated and time-consuming to apply in real multimedia content protection system. Network Coding (NC) is a new technology which codes the input information at intermediate nodes and can make the network information flow achieve the upper bound of max-flow min-cut theory. Random Linear Network Coding (RLNC) is a practical approach of NC with advantage of decentralized operation. In this paper, we proposed a novel scalable video streaming protection algorithm with an efficient packetization scheme and a distribution scheme according to authority levels of different users based on RLNC approach. The algorithm classified the packetized packets into different levels and employed three kinds of permission matrix (PM) to control the distribution of video streaming. The implementation of the proposed algorithm was described in detail to show the efficiency and low complexity of the method.
Ji Lu, Song Xiao 0001, Chengke Wu 0001
IAS2
2008 Priority Ordering Algorithm for Scalable Video Coding Transmission over Heterogeneous Network
abstract
Scalable video representation of the state-of-art scalable extension of the H.264/MPEG-4 AVC (SVC) and its combined 3d scalability feature offer an excellent solution to flexible video multicast over IP network. In order to relieve the impact of the severe bandwidth fluctuations and packet loss to the reconstructed quality in heterogeneous networks like the internet, a GOP-adaptive layer-based priority ordering algorithm for SVC is proposed in this paper. The method ordered the SVC bit stream according to the rate distortion contribution of different layer to the whole performance within a GOP, which makes the transmission more efficient and robust under the same bandwidth condition. Simulation results of different sequences are given to demonstrate that the proposed algorithms offer better performance in video quality as compared the default SVC ordering method and the SNR-based ordering method.
Song Xiao 0001, Chengke Wu 0001, Yunsong Li 0001, Jianchao Du, C.-C. Jay Kuo
AINA1
2007 Robust and Flexible Wireless Video Multicast with Network Coding
abstract
A robust and flexible wireless video multicast system with the H.264 scalable video coding (SVC) coding format and network coding (NC) is proposed in this work. The system enables efficient video streaming to heterogeneous receivers in packet erasure channels. NC is used to simplify erasure protection as well as the complexity of multicast tree construction and maintenance. It is shown by simulation results that the proposed solution improves PSNR over the traditional video multicast with error correction codes (ECC) in store-and-forward networks.
Hui Wang 0005, Song Xiao 0001, C.-C. Jay Kuo
GLOBECOM2
2007 Wyner-Ziv Video Coding with Spatio-Temporal Side Information
abstract
In this paper, we propose a Wyner-Ziv video coding scheme with spatio-temporal side information. In the scheme, temporal side information is generated through motion compensated temporal interpolation, using frames that are already decoded. While spatial side information is created through spatial prediction, using the temporal side information. The spatio-temporal side information are both incorporated into a turbo code based Slepian-Wolf decoder to decode the quantized source, but only temporal side information is used for reconstruction. Extensive simulations over various test video sequences show, compared with a similar codec with only temporal side information, the proposed scheme can achieve more than 20% bit rate savings without sacrifice in PSNR, and hence significantly improves the coding efficiency.
Yangli Wang, Jechang Jeong, Chengke Wu 0001, Song Xiao 0001
ICME4
2006 Reliable Transmission of H.264 Video over Wireless Network
abstract
A new error resilient method based on data partitioning and unequal error protection for H.264 video transmission over wireless network is proposed in this paper. By introducing the Impact factor of Inter MB and /spl rho/-curve, the method further divides the C type partition into several subtypes according to their effects to error propagation, then unequal error protection are applied to each partition including that of A, B and C partition. The theoretical model of rate allocation and a fast bidirectional local search algorithm based on iteratively improvement are also presented. The simulation results show that our proposed approach can achieve better results than conventional H.264 coding method (about 0.2-0.6 dB's gain in PSNR) when transmitting over wireless channel and can obtain more graceful degradation with the increasing of packet loss rate.
Song Xiao 0001, Chengke Wu 0001, Jianchao Du, Yadong Yang
AINA (2)1
2006 Error Resilient Transmission of H.264 Video over Wireless Network
abstract
Summary form only given. This paper proposes a new partition method for reliable H.264 video transmission over wireless channels wherein the impact of each inter MB to error propagation is analyzed. A relationship between image quality and the correctly received MB is established and this helps the encoder find the truncation point to satisfy the quality requirement. The inter partition (or C type) data are further divided into several partitions, then unequal error protection (UEP) can be applied to them with the use of the priority encoding transmission (PET) method. The use of a new fast iteratively improved bidirectional local search algorithm is suggested. Simulation results show that the proposed approach can achieve better results than conventional H.264 coding method when transmitting over wireless channel and can obtain more graceful degradation with the increasing of packet loss rate.
Song Xiao 0001, Chengke Wu 0001, Jianchao Du, Yadong Yang
DCC1
2005 Robust Video Communication Based on Source Modeling and Network Congestion Control
abstract
A new method combining source channel coding with source modeling and network congestion control for reliable video transmission over wireless network is proposed in this paper. Based on scene modeling and characteristic analyzing, all layers generated by scalable coding are classified into several types and two queuing are made respectively according to their contribution to network congestion and to the quality of reconstructed video. Then the source rate, protection level and congestion control strategy are dynamically adjusted according to different error status in which the packet loss is caused by network congestion or by unreliable transmission in wireless channel. The simulation results show that our proposed approach can achieve better results than MPEG-4 video source coding with fix rate Turbo coding and than the one which dynamically adjusting source coding and channel coding rates only.
Song Xiao 0001, Chengke Wu 0001, Jechang Jeong
AINA1
2004 A new multiple description layered coding method over ad-hoc network
abstract
A new combined multiple description and layered coding method based on adaptive frame insertion (AFl-MDLC) for transmission of video bit stream over wireless ad-hoc network is proposed in this paper. The main idea of the method is to adaptively insert the transition frames according to the relative motion between two neighboring frames, and then divide the video sequence into two descriptions with independent prediction loops. Afterwards, the base layer and the enhancement layer bit stream for each description are generated and transmitted across wireless network, which is modeled with finite state Markov process. A multiple-path transmitting strategy is presented for the AFI-MDLC data packets according to the characteristics of ad hoc multi-hop wireless network. The experimental results show that the method can help the decoder recover from packet loss more quickly as compared with the previous methods, and provides more stable and better quality for the reconstruction of video sequence.
Song Xiao 0001, Chengke Wu 0001
ICIP1
2003 A New Robust Multiple Description Coding Method Based on Region of Interest
abstract
A new revised multiple description SPIHT coding is proposed to combat packet loss. According to the region of interest of human eyes, the method reorders the zero trees of wavelets and assigns different coding rates to redundant trees. Simulation results show that the method can improve the image quality both objectively and subjectively compared to other multiple description coding methods in the case of packet loss.
Song Xiao 0001, Chengke Wu 0001, Yunsong Li 0001, Yaoping Yan
AINA1