Xiao-Yang Zhao 0003

dblp:223/4321 · also Xiaoyang Zhao 0003 · DBLP profile ↗
← Back
17ranked-venue papers
4as first author
13since 2021 · last 2026
0000-0002-8707-3768ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 10 · 4 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4Artificial intelligence and machine learning · 3 · 3 since 2021
YearPublicationVenuePosition
2026 Perceptive scale and selective attention few-shot learning network for hyperspectral and light detection and ranging fusion classification
Xiang-Hai Wang 0001, Tingting Geng, Xiaohan Xie, Xiao-Yang Zhao 0003, Siyao Li
Eng. Appl. Artif. Intell.5
2026 HiF2-FSLF: Hierarchical frequency fusion few-Shot learning framework for hyperspectral and lidar classification
Xiang-Hai Wang 0001, Xiaohan Xie, Xiao-Yang Zhao 0003, Siyao Li
Expert Syst. Appl.4
2025 Patch- and Class-Wise Hyperspectral Knowledge Learning: A Composite Consistency-Constrained Self-Ensemble Framework for Change Detection
abstract
Obtaining fine land surface change information from multitemporal hyperspectral images (HSIs) is a key goal pursued in remote sensing image processing. Recently, HSI change detection (HSI-CD) methods based on convolutional neural networks (CNNs) have achieved surprising detection results. One of the reasons is the support of large-scale labeled samples for network learning. However, the existence of mixed pixels greatly increases the difficulty of HSI interpretation, resulting in accurate pixel-level labeling work with a heavy burden and unable to meet the needs of time-sensitive applications. For this reason, achieving stable and high-precision CD with fewer samples is a difficult issue in this field. To address the above problems, a composite consistency-constrained self-ensemble framework (C3SelF) for HSI-CD is proposed, to alleviate the problems of low detection accuracy and instability caused by small samples. The framework mainly comprises two lightweight networks with the same structure aiming at accelerating the model inference process and thus improving the processing timeliness. The composite learning mode implements patch-wise classification loss, class-wise consistency loss on labeled samples, and patch-wise consistency loss on unlabeled samples under a multilevel noise perturbation strategy, which improves the classification results and reduces the labeling cost. Moreover, to exploit the multidimensional features contained in HSIs, a lightweight selective spatial-spectral feature joint network (S3Net) is designed to overcome over-fitting, and to deeply mine the discriminative information in unlabeled samples, a new sample screening strategy is designed to ensure the stability of the network during training unlabeled samples. Extensive experiments prove that the proposed C3SelF outperforms the state-of-the-art (SOTA) methods at a sampling rate of 0.1%, reaching 93.44% Kappa and 97.26% overall accuracy (OA) on the Farmland dataset. The source code of the proposed framework will be released athttps://github.com/zxylnnu/C3SelF.
Xiao-Yang Zhao 0003, Siyao Li, Chuanming Song 0001, Xiang-Hai Wang 0001
IEEE Trans. Geosci. Remote. Sens.1
2024 Interdomain Collaboration Between Hyperspectral and VHR Remote Sensing Images: A Cross-Scene Few-Shot Learning Framework for Change Detection
abstract
Hyperspectral image change detection (HSI-CD) based on deep learning (DL) has made significant progress. However, these methods rely significantly on the number of labeled data. Annotating HSI is a highly complex task that requires professional knowledge for guidance, resulting in a scarcity of high-quality labeled samples. The emergence of few-shot learning (FSL), which supports model learning from limited labeled samples, can address this issue. However, FSL-based methods still face some challenges: 1) existing methods mainly rely on single-source or homogenous cross-domain HSI data, which is difficult to adequately cope with the problem of scarcity of HSI labeled data; 2) most existing methods usually only focus on local features within patches and neglect interrelationships between patches, which is also important for model learning; and 3) transformers modeling long-range relationships rely on extensive labeled data, making it difficult to perform well in few-shot scenarios. Therefore we propose a cross-scene FSL framework based on interdomain collaboration (CSIDC-FSL) for HSI-CD. Specifically, the following is proposed: 1) FSL is performed on very high-resolution image (VHRI) and HSI, aiming to use the learnable information in VHRI with low annotation cost to help HSI-CD, reducing the dependence of the model on HSI annotation data while enabling multilevel feature hybrid perceptual CD; 2) a dual-information integrated mapping module (DI2M) is proposed, which designs a CNN and transformer integrated structure that can simultaneously focus on local features and class-wise long-range relationships to break the constraints of local perception of CNN while optimizing the performance of transformer under few-shot situations; and 3) the interdomain joint information allocation module (IDM) is designed to capture cross-scene domain-wise distribution features, and mitigate the impact of distribution differences in cross-scene data (VHRI and HSI) on knowledge learning and migration through the collaboratively consistent interdomain features. Under the condition of five samples per class, the CD results of CSIDC-FSL are better than those of recently advanced algorithms, with average improvements of 1.46%–1.5% for overall accuracy (OA) and average accuracy (AA), respectively. The code will be made available athttps://github.com/lsylnnu/CSIDC-FSL.
Xiang-Hai Wang 0001, Siyao Li, Xiao-Yang Zhao 0003, Yuetong Zhao
IEEE Trans. Geosci. Remote. Sens.3
2024 GTransCD: Graph Transformer-Guided Multitemporal Information United Framework for Hyperspectral Image Change Detection
abstract
Using multitemporal hyperspectral images (HSIs) to obtain fine-grained land cover change information is an essential task in remote sensing (RS) image processing. Convolutional neural networks (CNNs), which have strong feature extraction and nonlinear regression capabilities, have recently aided in the advancement of this subject. However, the performance of these supervised methods is usually limited by small receptive field and less labeled samples. To this end, how to break through the aforementioned bottleneck and build a more suitable change detection (CD) framework for HSI is a crucial and challenging issue. To this end, a graph transformer-guided multitemporal information united framework for HSI-CD (GTransCD) is proposed, which mainly consists of the following three components: 1) applying transformer to the graph structure, a salient relationship strengthening graph transformer (GTrans) module is created, making it possible for the network to capture distant change information, and on this basis, local- and global-range information are aggregated simultaneously; 2) a gated change information fusion (GCF) unit is designed to inject the GTrans-guided change features into the original bitemporal concatenated features to further enhance the representation of change information in the network; and 3) a general HSI-CD framework that can organically blend change features guided by GTrans module with original features is proposed, with the intention of reducing the reliance on training samples by utilizing the semi-supervised learning mode of graph neural networks. numerous experiments demonstrate the proposed GTransCD surpasses the state-of-the-art methods and has a high level even at low sampling rates. The source code of the proposed framework will be released athttps://github.com/zxylnnu/GTransCD.
Xiao-Yang Zhao 0003, Siyao Li, Tingting Geng, Xiang-Hai Wang 0001
IEEE Trans. Geosci. Remote. Sens.1
2024 GaMPF: A Full-Scale Gated Message Passing Framework Based on Collaborative Estimation for VHR Remote Sensing Image Change Detection
abstract
With the maturity and popularization of high-performance sensor technology, it is now possible to acquire huge amounts of very high-resolution (VHR) remote sensing images. The change detection (CD) for VHR images is currently receiving special attention for remote sensing earth observation applications, however, as a hot research field, it needs to be studied in depth to improve the detection accuracy of fine changes. To this end, a full-scale gated message passing framework (GaMPF) based on collaborative estimation for VHR remote sensing image change detection is proposed in this paper. On one hand, the key embedding representation is generated for each feature map by means of the collaborative estimation (CE) strategy; On the other hand, grounded in timing analysis, bitemporal features are sent selectively on dual paths according to the full-scale gated (FsG) mechanism. Specifically, this framework consists of the following four components: 1) Taking shared-weights Siamese network as an encoder to extract multi-scale features; 2) Generate a set of shared compact bases under the CE strategy and infer the key embedding representations on the basis of the shared bases for feature maps at the same level, considering the representations as the gated switches; 3) FsG mechanism is used as the mode of message passing between bitemporal images, which guides the information can be transmitted simultaneously on both within-and cross-temporal paths. 4) Creating a stepwise dense fusion module (DFM) as a decoder for predicting the change map. Experimental results show that the GaMPF proposed in this paper outperforms existing SOTA methods, and is particularly good at detecting edges and small objects. The source code will be released at https://github.com/zxylnnu/GaMPF.
Xiao-Yang Zhao 0003, Keyun Zhao, Siyao Li, Chuanming Song 0001, Xiang-Hai Wang 0001
IEEE Trans. Geosci. Remote. Sens.1
2023 MCT-Net: Multi-hierarchical cross transformer for hyperspectral and multispectral image fusion
Xiang-Hai Wang 0001, Xinying Wang 0005, Ruoxi Song, Xiao-Yang Zhao 0003, Keyun Zhao
Knowl. Based Syst.4
2023 GTMSiam: Gated Transmitting-Based Multiscale Siamese Network for Hyperspectral Image Change Detection
abstract
Hyperspectral image change detection (HSI-CD) is a technique that detects changes in land cover occurring in a specific area within a closed time. At present, most existing methods for HSI-CD employ exceedingly intricate network architectures, leading to a high model complexity that hampers the achievement of a favorable trade-off between change detection accuracy and timeliness. Furthermore, existing methods often confine the feature extraction process to a single scale rather than multiple diverse scales. However, employing a multiscale approach for feature extraction allows for capturing finer-grained features encompassing more intricate details, as well as coarser-grained features that aggregate local information over a larger range. On the other hand, most existing methods overemphasize the complexity of the feature extraction process and underestimate the importance of the conversion process from bi-temporal features to valuable change features. To this end, a gated transmitting based multiscale siamese network (GTMSiam) is proposed, which mainly contains the following two portions: 1) dual branches with the siamese structure, which capture spatial features of the HSIs at multiple scales while preserving rich spectral information. Moreover, the siamese design effectively reduces the network parameters, thereby alleviating the computational complexity of the model. 2) gated change information transmitting module (GTM), which utilizes gated neural units to transform bi-temporal image features into land cover change information, while progressively transmitting change information at different scales. This enables the network to leverage diverse scale change information for comprehensive discrimination of land object changes. Experimental results on three publicly available datasets demonstrate the superior performance of the proposed GTMSiam. Simultaneously, the complexity analysis experiment proves that the GTMSiam can give consideration to both detection performance and timeliness. The source code of this letter will be released at https://github.com/zkylnnu/GTMSiam.
Xiang-Hai Wang 0001, Keyun Zhao, Xiao-Yang Zhao 0003, Siyao Li
IEEE Geosci. Remote. Sens. Lett.3
2023 BiG-FSLF: A Cross Heterogeneous Domain Few-Shot Learning Framework Based on Bidirectional Generation for Hyperspectral Image Change Detection
abstract
In recent years, hyperspectral image change detection (HSI-CD) based on deep learning has achieved high detection accuracy, but these methods obtain excellent detection results usually rely on having sufficient labeled samples to train the network. However, the production of HSI label is difficult, costly and inefficient. In practical tasks, often only a limited number of labeled samples can be obtained due to the limitation of timeliness. To address this problem, a cross heterogeneous domain few-shot learning framework based on bidirectional generation (BiG-FSLF) is proposed for HSI-CD, which aims to solve the few-shot problem of HSI-CD by few-shot learning (FSL), and to assist HSI-FSL perform better by obtaining learnable changed information (i.e., empirical knowledge) from another remote sensing data. Specifically, a multitask generation encoder (MLGenE) is designed to take on both the tasks of FSL and domain adaptation to achieve HSI-CD under the condition of cross heterogeneous domain few-shot. First, we take any pair of image data in a very high resolution image (VHRI) CD dataset as the source domain and HSI is used as the target domain, using sufficient labeled samples in source domain and a small number of labeled samples in target domain for FSL. Meanwhile, a bidirectional generation domain adaptation (BiGDA) method based on generative adversarial strategy is proposed to achieve adaptive alignment of the two heterogeneous domains (source and target domains) feature distributions, to mitigate the impact of the domain shift problem inherent to cross domain data on FSL. Abundant experiments with only five training samples on the publicly available popular HSI-CD datasets confirm that the proposed method can show great detection performance. The source code of the proposed framework will be released at https://github.com/lsylnnu/BiG-FSLF.
Xiang-Hai Wang 0001, Siyao Li, Xiao-Yang Zhao 0003, Keyun Zhao
IEEE Trans. Geosci. Remote. Sens.3
2023 TriTF: A Triplet Transformer Framework Based on Parents and Brother Attention for Hyperspectral Image Change Detection
abstract
Hyperspectral image (HSI) change detection (CD) is a technique to accurately detect land cover changes by using HSIs with rich spatial-spectral information. In recent years, the HSI-CD methods based on convolutional neural networks (CNNs) have achieved great success because of their flexible and effective feature extraction ability. However, these methods often take the HSI patches as the input of the networks, which undoubtedly hinders the overall perception of the HSIs. Meanwhile, the valuable temporal information in HSIs is often underutilized. For this end, a triplet transformer framework (TriTF) based on parents-temporal attention and brother-spatial attention is proposed for HSI-CD. The proposed framework mainly contains the following three parts: 1) Transformer-based network backbone, which uses the self-attention to capture the correlation between arbitrarily two pixels in the same patch and extracts the global spatial correlation in the unit of encoded input patches; 2) parents-temporal attention (PTA) branch. Unlike the previous cross-temporal attention mechanisms of the “T1↔T2” mode which only consider the interaction between bi-temporal HSIs, this paper constructs a novel PTA of the “T1→T3←T2” mode which takes the difference-temporal image T3 as the core. The impact of bi-temporal HSIs on the land cover changes is more concerned in the PTA; 3) brother-spatial attention (BSA) branch. The most similar patch in the current training batch of each patch is defined as its brother patch. Furthermore, cross-spatial attention is applied to propagate the features of the brother patch to the current patch. Thus, the middle- and long-range dependencies can be utilized and the scope of feature propagation can be extended. In this paper, the experiments under low and high sampling rates are conducted and proved the outstanding change detection performance of the proposed TriTF when compared with abundant state-of-the-art (SOTA) CD algorithms. The source code of this paper will be released at https://github.com/zkylnnu/TriTF.
Xiang-Hai Wang 0001, Keyun Zhao, Xiao-Yang Zhao 0003, Siyao Li
IEEE Trans. Geosci. Remote. Sens.3
2023 GeSANet: Geospatial-Awareness Network for VHR Remote Sensing Image Change Detection
abstract
The characteristics of very high resolution (VHR) remote sensing images (RSIs) have higher spatial resolution inherently, and are easier to obtain globally compared with hyperspectral images (HSIs), making it possible to detect small-scale land cover changes in multiple applications. RSI change detection (RSI-CD) based on deep learning has been paid attention to and become a frontier research field in recent years, and is currently facing two challenging problems: The first is high dependence on registration between bi-temporal images caused by high spatial resolution; The other is high pseudo-change information response caused by low spectral resolution. In order to address the above-mentioned two problems, a novel RSI-CD framework called Geospatial-Awareness Network (GeSANet) based on the geospatial Position Matching Mechanism (PMM) with multi-level adjustment and the geo-spatial Content Reasoning Mechanism (CRM) with diverse pseudo-change information filtering is proposed. First of all, the PMM assigns independent two-dimensional offset coordinates to each position in the previous temporal image, afterwards, bilinear interpolation is employed to obtain the subpixel feature value after the offset, and the sparse results based on the difference are transmitted to the next level prediction to realize multi-level geospatial correction. The CRM extracts global features from the corrected sparse feature map in terms of dimensions, implementing effective discriminant feature extraction on basis of the original feature map in a stepwise refinement manner through the cross-dimension exchange mechanism, to filter out various pseudo-change information as well as maintain real change information. Comparison experiments with five recent SOTA methods are carried out on two popular datasets with diverse changes, the results show that the proposed method has good robustness and validity for multi-temporal RSI-CD. In particular, it has a strong comparative advantage in detecting small entity changes and edge details. The source code of the proposed framework can be downloaded from https://github.com/zxylnnu/GeSANet.
Xiao-Yang Zhao 0003, Keyun Zhao, Siyao Li, Xiang-Hai Wang 0001
IEEE Trans. Geosci. Remote. Sens.1
2022 CSDBF: Dual-Branch Framework Based on Temporal-Spatial Joint Graph Attention With Complement Strategy for Hyperspectral Image Change Detection
abstract
Hyperspectral image (HSI) change detection (CD) aims at obtaining internal components’ change information of land cover and land use. In recent years, the development of convolutional neural networks (CNNs) has greatly promoted the research progress in this field. However, the fixed small-size convolution kernels used by CNNs have severely limited the receptive field of information. Another defect of most CNN-based models is their strong dependence on samples, and they are not competent for tasks with a small number of samples. Besides, the traditional CNN-based models can only perform convolution to learn the spatial–spectral features in the Euclidean space, which is not conducive to capturing the geometric changes in land covers in the HSIs. Differently, the graph attention network (GAT) has come into prominence due to its ability to capture the holistic topology structure of images flexibly, and the attention coefficients can be used to effectively model the long-range correlations between land covers. The semi-supervised nature of GAT is also well-suited to handle HSI-CD tasks with limited samples. Nevertheless, the pixel-level topology structure often generates expensive computational costs. To this end, a dual-branch framework based on temporal–spatial joint graph attention (TSJGAT) with complement strategy (CSDBF) is proposed for HSI-CD, which extracts superpixel- and pixel-level features from bitemporal HSIs in parallel and enables them to complement each other. The proposed CSDBF mainly consists of two branches: superpixel-level feature extraction branch (S-branch) and pixel-level feature extraction branch (P-branch). In the S-branch, we introduce the idea of GAT into HSI-CD for the first time and propose a novel TSJGAT module. Thus, the temporal–spatial features of HSIs are propagated and aggregated on the nonlinear graph structure, which makes the changed regions more discriminable. In the P-branch, pixel-level features are obtained by CNNs to correct uncertain factors caused by superpixel segmentation in the S-branch, which is complementary to the S-branch and lays a foundation for more accurate CD. Abundant experiments show that compared with other pioneer methods, the proposed CSDBF can improve the Kappa coefficient by more than 1.9% and 2.5% on average in general sampling rate situations and a low sampling rate situation, respectively, which shows better robustness and better detection accuracy than most existing state-of-the-art methods. The source code of this article can be downloaded fromhttps://github.com/zkylnnu/CSDBF.
Xiang-Hai Wang 0001, Keyun Zhao, Xiao-Yang Zhao 0003, Siyao Li
IEEE Trans. Geosci. Remote. Sens.3
2022 FSL-Unet: Full-Scale Linked Unet With Spatial-Spectral Joint Perceptual Attention for Hyperspectral and Multispectral Image Fusion
abstract
The application of hyperspectral image (HSI) is more and more extensive, but the lower spatial resolution seriously affects its application effect. Using low-resolution hyperspectral image (LR-HSI) and high-resolution multispectral image (MSI) fusion technology to achieve super-resolution reconstruction of HSI has become a mainstream method. However, most of the existing fusion methods do not make full use of the large-scale range of remote sensing images, and neglect the preservation of spatial-spectral information in the fusion process. Considering that the spectral information in fused high-resolution hyperspectral image (HR-HSI) mainly depends on HSI, and the spatial information mainly depends on MSI, this paper proposes a full-scale linked Unet with spatial-spectral joint perceptual attention for hyperspectral and multispectral image fusion (FSL-Unet). The FSL-Unet consists of two modules, the first is spatial-spectral attention extraction module (SSAE), which is used to calculate the spectral attention of LR-HSI and the spatial attention of HR-MSI at different scales. The second is the full-scale link U-shaped fusion module (FLUF), which adopts a multi-level feature extraction strategy, using denser full-scale skip connections to explore feature information in a finer-grained range, enabling flexible combination of multi-scale and multi-path features. At the same time, we propose spatial-spectral joint peceptual attention (SSJPA) on the encoder side of FLUF. SSJPA can make full use of the attention maps computed by the SSAE, and then effectively embed spatial and spectral information into the fused image, enabling uninterrupted information transfer and aggregation. To demonstrate the effectiveness of FSL-Unet, we selected five public hyperspectral datasets for experiments. Compared with other eight state-of-the-art fusion methods, the experimental results show that the FSL-Unet achieves competitive results. The source code for FSL-Unet can be downloaded from https://github.com/wxy11-27/FSL-Unet.
Xiang-Hai Wang 0001, Xinying Wang 0005, Keyun Zhao, Xiao-Yang Zhao 0003, Chuanming Song 0001
IEEE Trans. Geosci. Remote. Sens.4
2020 NSST and vector-valued C-V model based image segmentation algorithm
abstract
Image segmentation is a process of partitioning an image into non‐overlapping regions. Existing unsupervised image segmentation methods include level set, automatic thresholding and region‐based CV mode and so on. However, image segmentation as a key technology in the field of image processing has not been solved indeed, especially for images with complex texture. For this reason, the authors proposed a novel image segmentation algorithm based on NSST and the vector‐valued Chan–Vese (C–V) model. First, they obtained a multi‐scale representation by exploiting the non‐subsampled shearlet transform (NSST) to extract multi‐dimensional data in the image. Afterwards, they gave the vector‐valued C–V model, and applied it to all subbands of NSST, which are treated as a vector‐valued image. By comparing with other class methods, the experimental results show that the proposed method has better visual effects and lower error rates. But at the same time, it is a little time consuming. The proposed method is reasonable and effective, by taking full advantages of each subband's directional information during its diffusion process, compared with traditional C–V model.
Xiang-Hai Wang 0001, Xiao-Yang Zhao 0003, Yihuan Zhu
IET Image Process.2
2019 Patch-based contour prior image denoising for salt and pepper noise
Bo Fu 0001, Xiao-Yang Zhao 0003, Xiang-Hai Wang 0001
Multim. Tools Appl.2
2019 A convolutional neural networks denoising approach for salt and pepper noise
Bo Fu 0001, Xiao-Yang Zhao 0003, Xiang-Hai Wang 0001, Yonggong Ren
Multim. Tools Appl.2
2019 A salt and pepper noise image denoising method based on the generative classification
Bo Fu 0001, Xiao-Yang Zhao 0003, Chuanming Song 0001, Ximing Li 0002, Xiang-Hai Wang 0001
Multim. Tools Appl.2