EDBT 2026 Demo / reviewers in the wild / expert
Yuxiang Zhang 0005
dblp:73/7697-5
· DBLP profile ↗
27ranked-venue papers
10as first author
25since 2021 · last 2026
0009-0002-9262-3135ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 15 · 4 first-author · 13 since 2021Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Block-Wise Contrastive Few-Shot Learning With Multiview Mixing for Hyperspectral Image Cross-Scene ClassificationabstractThe cross-scene classification of hyperspectral images (HSI) faces challenges posed by different sensors and various land cover categories, and methods based on the cross-domain Few-shot Learning (FSL) framework are commonly used to address the challenge. However, current few-shot learning (FSL) methods struggle with the issue of intra-domain inductive bias, where models tend to learn simplistic and potentially erroneous correlations within samples (classes with similar backgrounds are always confused due to simple correlations with the background), leading to poor generalization ability to target domains. Furthermore, most approaches employ a patch-wise representation learning mechanism, which provides the model with all information within a sample, but inherently limits its representation capacity when dealing with the limited prior information in FSL tasks. To address these issues, we propose a Block-wise Contrastive Few-shot Learning (BCFSL) framework. First, a multi-view cross-domain mixing strategy enhances spatial diversity and constructs harder samples to mitigate inductive bias. Second, a novel block-wise representation mechanism performs fine-grained feature extraction from local regions, improving generalization. Finally, a multi-view spectral reconstruction module preserves essential high-dimensional spectral information during training and assists the block-wise representation in fully leveraging prior knowledge. Extensive experiments on three benchmark HSI datasets validate the effectiveness and superiority of the proposed method. Yingze Xie, Zhengyi Lv, Yuxiang Zhang 0005, Mengmeng Zhang 0005 |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2026 | Alliance: All-in-One Spectral-Spatial-Frequency Awareness Foundation ModelabstractFrequency domain analysis reveals fundamental image patterns difficult to observe in raw pixel values, while avoiding redundant information in original image processing. Although recent remote sensing foundation models (FMs) have made progress in leveraging spatial and spectral information, they have limitations in fully utilizing frequency characteristics that capture hidden features. Existing FMs that incorporate frequency properties often struggle to maintain connections with the original image content, creating a semantic gap that affects downstream performance. To address these challenges, we propose the All-in-One Spectral-Spatial-Frequency Awareness Foundation Model (Alliance), a framework that effectively integrates information across all three domains. Alliance introduces several key innovations: (1) a progressive frequency decoding mechanism inspired by human visual cognition that minimizes multi-domain information gaps while preserving connections between general image information and frequency characteristics, progressively reconstructing from low to mid to high frequencies to extract patterns difficult to observe in raw pixel values; (2) a triple-domain fusion attention module that separately processes amplitude, phase, and spectral-spatial relationships for comprehensive feature integration; and (3) frequency embedding with frequency-aware Cls token initialization and frequency-specific mask token initialization that achieves fine-grained modeling of different frequency band information. Additionally, to evaluate FMs generalizability, we construct the Yellow River dataset, a large-scale multi-temporal collection that introduces challenging cross-domain tasks and establishes more rigorous standards for FMs assessment. Extensive experiments across six downstream tasks demonstrate Alliance's superior performance. Wei Li 0032, Yuxiang Zhang 0005, Ran Tao 0003, Qian Du 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2026 | LiteMFT: Lightweight Multi-Modal Fine-Tuning for Semantic SegmentationabstractMulti-modal image segmentation has recently attracted considerable attention due to its ability to integrate complementary information from diverse sensors, thereby enabling more accurate semantic predictions in complex or specialized scenarios. However, as data volume and model capacity continue to grow, many existing methods suffer substantial increases in parameters and computational costs, particularly with the widespread adoption of Vision Foundation Models (VFMs). To address these challenges, we introduce a Lightweight Multi-modal Fine-Tuning framework (LiteMFT) designed for efficient and generalizable adaptation of RGB-pretrained VFMs to multi-modal semantic segmentation. By incorporating only a small number of trainable parameters, LiteMFT enables effective extension of existing models to handle multi-modal image fusion tasks. The framework centers around two key components: the Modality Local Competition (MLC) module, which dynamically and efficiently fuses complementary features across modalities, and the Gated Low-Rank Adapter (GLR), which improves the backbone's adaptability to multi-modal data through content-aware low-rank transformation. Extensive experiments on both bi-modal and tri-modal segmentation tasks demonstrate that LiteMFT not only achieves competitive or superior performance but also exhibits strong scalability for additional modalities, underscoring its practicality and broad applicability in multi-modal semantic segmentation. Chengwang Guo, Yuxiang Zhang 0005, Mengmeng Zhang 0005, Huan Liu 0015, Wei Li 0032 |
IEEE Trans. Image Process. | 2 |
| 2026 | Cross-Scene Hyperspectral Image Classification via Bidirectional Mamba and Domain Mixing NetworkabstractTo overcome the challenges posed by domain shift in hyperspectral image (HSI) classification, methods based on domain adaptation (DA) have been widely used. Currently, most HSI DA methods focus on designing complex strategies to align the distributions of the source domain (SD) and the target domain (TD) in the feature space after feature extraction, yielding promising results. However, when there exists a large domain shift between SD and TD, it becomes challenging to map them into the same feature space. In this article, we propose the bidirectional mamba and domain mixing network (BMDMnet). Since pure CNN architectures are constrained in local feature extraction, while transformer-based models improve global feature capturing capability at the cost of high computational complexity, we propose the bidirectional mamba module (BMM) as an efficient solution for capturing long-range dependencies. In addition, a self-distillation strategy is employed during training. By utilizing a more stable teacher model, reliable predictions can be obtained in the TD. Subsequently, a domain mixing supervised learning (DMSL) module is designed, which creates a mixed domain by selecting low-entropy sample-pseudo-label pairs from the TD and randomly combining them with sample-label pairs from the SD. DMSL aims to introduce mixed domain to mitigate the inter-domain gap in the data space, thereby enabling the model to learn TD representations more effectively. Experiments demonstrate that BMDMnet outperforms state-of-the-art algorithms across three cross-scene datasets. Junzhe Dang, Chengwang Guo, Mengmeng Zhang 0005, Yuxiang Zhang 0005, Wen Jia, Wei Li 0032 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2025 | Cross-Domain Hyperspectral Image Classification Based on Bi-Directional Domain AdaptationabstractUtilizing hyperspectral remote sensing technology enables the extraction of fine-grained land cover classes. Typically, satellite or airborne images used for training and testing are acquired from different regions or times, where the same class has significant spectral shifts in different scenes. In this paper, we propose a Bi-directional Domain Adaptation (BiDA) framework for cross-domain hyperspectral image (HSI) classification, which focuses on extracting both domain-invariant features and domain-specific information in the independent adaptive space, thereby enhancing the adaptability and separability to the target scene. In the proposed BiDA, a triple-branch transformer architecture (the source branch, target branch, and coupled branch) with semantic tokenizer is designed as the backbone. Specifically, the source branch and target branch independently learn the adaptive space of source and target domains, a Coupled Multi-head Cross-attention (CMCA) mechanism is developed in coupled branch for feature interaction and inter-domain correlation mining. Furthermore, a bi-directional distillation loss is designed to guide adaptive space learning using inter-domain correlation. Finally, we propose an Adaptive Reinforcement Strategy (ARS) to encourage the model to focus on specific generalized feature extraction within both source and target scenes in noise condition. Experimental results on cross-temporal/scene airborne and satellite datasets demonstrate that the proposed BiDA performs significantly better than some state-of-the-art domain adaptation approaches. In the cross-temporal tree species classification task, the proposed BiDA is more than 3%∼5% higher than the most advanced method. The codes will be available from the website: https://github.com/YuxiangZhang-BIT/IEEE TCSVT BiDA. Yuxiang Zhang 0005, Wei Li 0032, Wen Jia, Mengmeng Zhang 0005, Ran Tao 0003, Shunlin Liang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | MGCD: Change Detection Focuses on Multiscale Style Domain GeneralizationabstractChange detection (CD) plays a crucial role in remote sensing (RS) applications. Although deep learning (DL)-based CD methods have achieved impressive performance, they typically rely on large amounts of well-annotated data, which is often scarce in real-world scenarios. This data scarcity leads to overfitting and limited generalization ability in conventional CD models. To address these challenges, this article proposes a novel few-sample CD framework named multiscale style DG method for change detection (MGCD). The core idea is to enhance the model’s ability to learn domain-invariant features by increasing data diversity across multiple style scales. Specifically, MGCD introduces global style diversity by incorporating out-of-domain natural images, and enriches local structural styles through unsupervised clustering and randomization. Additionally, a semantic consistency supervision (SCS) strategy is designed to guide multitemporal feature learning, enabling the network to better capture changes across diverse styles and scales. Extensive experiments conducted on three benchmark datasets demonstrate the effectiveness and robustness of the proposed MGCD framework in few-sample CD tasks. Chengwang Guo, Mengmeng Zhang 0005, Yuxiang Zhang 0005, Huan Liu 0015, Wei Li 0032 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | Attention Multiscale Network for Semantic Segmentation of Multimodal Remote Sensing ImagesabstractDue to recent advancements in deep learning, techniques for urban structure extraction and semantic segmentation of multimodal remote sensing images have significant improvements. However, the challenge arises from the variable color intensity and complex texture of urban structures in optical images, particularly in buildings and roads. Fortunately, the light detection and ranging (LiDAR) images promote the task of developing an optimal multimodal fusion network that effectively leverages information from different modalities. In this article, we propose an attention multiscale network (AMSNet) for binary semantic segmentation tasks focused on building extraction, as well as multiclass semantic segmentation tasks, by integrating optical and LiDAR remote sensing images. AMSNet introduces two feature fusion modules—spatial scale adaptive fusion (S2AF) and semantic guided fusion (SGF). S2AF facilitates feature fusion between optical and LiDAR images within the same layer. This module contains a spatial scale selection strategy and an adaptive weight learning strategy, which enables the network to adaptively extract and intentionally select multiscale features from multimodal data. SGF addresses the semantic gap between different layered block features through semantic feature guidance strategy while achieving feature fusion. Furthermore, we introduce robust feature learning (RFL) to ensure the network robustness in rotation and variation in objects, making it resilient to images captured from different viewpoints and sensors. RFL incorporates point-to-point similarity learning strategy and multiscale feature reuse strategy. Experimental results on publicly available datasets demonstrate that AMSNet outperforms other state-of-the-art models. Extensive ablation studies further confirm the significance of all key components in the proposed approach. The source code of this method is available athttps://github.com/B-LG-J/AMSNet.git. Zhen Ye 0007, Yuan Li 0037, Zhen Li 0063, Huan Liu 0015, Yuxiang Zhang 0005, Wei Li 0032 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Locality Robust Domain Adaptation for cross-scene hyperspectral image classification
Wei Li 0032, Yuxiang Zhang 0005, Ran Tao 0003 |
Expert Syst. Appl. | 4 |
| 2024 | Unbalanced Class Learning Network With Scale-Adaptive Perception for Complicated Scene in Remote Sensing Images SegmentationabstractThe semantic segmentation of wide-field remote sensing images plays a significant role in many fields. However, due to the complexity of the content of remote sensing images, the dataset often has an uneven distribution of land type between different classes and large gaps in the scales of different objects. This often creates great problems for fine segmentation. To solve the issues, an unbalanced class learning network with Scale-adaptive perception (UCSANet) is proposed, which can adaptively cope with Multi-scale objects and unbalanced classes. The design can be inserted in any convolution network easily and can enrich features without increasing too many parameters. The network groups feature and uses atrous convolutions with different dilated rates on different groups to extract Multi-scale features while separable convolutions reduce the amount of network parameters. Then, the fusion of features between different scales is achieved through the self-attention mechanism. Furthermore, a weight map is designed to adaptively combine the predictions of two segmentation heads with Cross-Entropy loss and Lovasz-Softmax loss respectively, which enable the network to focus on learning low-frequency classes without affecting high-frequency classes. Experimental results on GF-6 MSI datasets demonstrate that the proposed UCSANet performs significantly better than others and achieves multi-class segmentation more accurately. Mengmeng Zhang 0005, Wei Li 0032, Yunhao Gao, Yuanyuan Gui, Yuxiang Zhang 0005 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2024 | GCCD: A Generative Cross-Domain Change Detection NetworkabstractChange detection (CD) in hyperspectral image (HSI) is of great importance in the remote sensing area. The HSI-CD method based on deep learning (DL) has shown significant progress in achieving precise detection performance. However, many existing methods overlook cross-domain challenges in the CD task. In addition, the scarcity of annotated samples makes the DL models prone to overfitting. To address these issues, a generative cross-domain CD (GCCD) network based on a domain generalization (DG) technique is proposed. GCCD consists of a generator and a discriminator. The generator, with a Morph encoder (ME) and a Semantic encoder (SE), preserves fundamental structural information while introducing randomization to style and content. The discriminator extracts change information for discrimination through dual-temporal images and their difference map. Supervised adversarial learning between the generator and discriminator enhances the model’s ability to extract domain-invariant information. Extensive experiments on various datasets demonstrate the superior performance of the proposed method. Mengmeng Zhang 0005, Chengwang Guo, Yuxiang Zhang 0005, Huan Liu 0015, Wei Li 0032 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Intermediate Domain Prototype Contrastive Adaptation for Spartina alterniflora Segmentation Using Multitemporal Remote Sensing ImagesabstractAs an invasive plant in wetlands, Spartina alterniflora (S. alterniflora) causes immeasurable damage to wetland ecosystems. Observing S.alterniflora using multitemporal remote sensing data helps us better understand its further development and facilitates effective containment of its invasion trend. However, inconsistent representation across remote sensing data from different time periods poses a challenge. Fortunately, the utilization of unsupervised domain adaptation (UDA) techniques helps in addressing such issues and enables the exploration of rich temporal dimension information in multitemporal remote sensing data, revealing the spatio-temporal distribution characteristics of S.alterniflora. However, existing UDA methods mostly focus on directly aligning the global or intraclass distribution representations across domains, which overlooks the issue of significant differences between extreme domains and lacks exploration of interclass relationships. To address these limitations, an intermediate domain prototype class-level learning network (IDPNet) is proposed. IDPNet utilizes dynamically generated intermediate domain (ID) features to construct class prototypes while incorporating interclass information into the prototype construction, achieving the class-centered distribution alignment for adaptation. Moreover, intermediate domain feature generation module (IFM) is employed in IDPNet to blend the latent representations from various domains and generate ID features in real time. Additionally, the hierarchical feature fusion module (HFM) is designed to enable IDPNet to learn more discriminative and robust spatio-temporal distribution features, thereby reducing the loss of information from patches. Experimental results on two cross-year multispectral datasets demonstrate that the proposed IDPNet outperforms several state-of-the-art UDA methods. Mengmeng Zhang 0005, Wei Li 0032, Xiukai Song, Yunhao Gao, Yuxiang Zhang 0005 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2024 | Graph Information Aggregation Cross-Domain Few-Shot Learning for Hyperspectral Image ClassificationabstractMost domain adaptation (DA) methods in cross-scene hyperspectral image classification focus on cases where source data (SD) and target data (TD) with the same classes are obtained by the same sensor. However, the classification performance is significantly reduced when there are new classes in TD. In addition, domain alignment, as one of the main approaches in DA, is carried out based on local spatial information, rarely taking into account nonlocal spatial information (nonlocal relationships) with strong correspondence. A graph information aggregation cross-domain few-shot learning (Gia-CFSL) framework is proposed, intending to make up for the above-mentioned shortcomings by combining FSL with domain alignment based on graph information aggregation. SD with all label samples and TD with a few label samples are implemented for FSL episodic training. Meanwhile, intradomain distribution extraction block (IDE-block) and cross-domain similarity aware block (CSA-block) are designed. The IDE-block is used to characterize and aggregate the intradomain nonlocal relationships and the interdomain feature and distribution similarities are captured in the CSA-block. Furthermore, feature-level and distribution-level cross-domain graph alignments are used to mitigate the impact of domain shift on FSL. Experimental results on three public HSI datasets demonstrate the superiority of the proposed method. The codes will be available from the website: https://github.com/YuxiangZhang-BIT/IEEE_TNNLS_Gia-CFSL. Yuxiang Zhang 0005, Wei Li 0032, Mengmeng Zhang 0005, Ran Tao 0003, Qian Du 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | Cross-Scene Joint Classification of Multisource Data With Multilevel Domain Adaption NetworkabstractDomain adaption (DA) is a challenging task that integrates knowledge from source domain (SD) to perform data analysis for target domain. Most of the existing DA approaches only focus on single-source-single-target setting. In contrast, multisource (MS) data collaborative utilization has been extensively used in various applications, while how to integrate DA with MS collaboration still faces great challenges. In this article, we propose a multilevel DA network (MDA-NET) for promoting information collaboration and cross-scene (CS) classification based on hyperspectral image (HSI) and light detection and ranging (LiDAR) data. In this framework, modality-related adapters are built, and then a mutual-aid classifier is used to aggregate all the discriminative information captured from different modalities for boosting CS classification performance. Experimental results on two cross-domain datasets show that the proposed method consistently provides better performance than other state-of-the-art DA approaches. Mengmeng Zhang 0005, Xudong Zhao 0003, Wei Li 0032, Yuxiang Zhang 0005, Ran Tao 0003, Qian Du 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2023 | Multi-Modal Domain Generalization for Cross-Scene Hyperspectral Image ClassificationabstractThe large-scale pre-training image-text foundation models have excelled in a number of downstream applications. The majority of domain generalization techniques, however, have never focused on mining linguistic modal knowledge to enhance model generalization performance. Additionally, text information has been ignored in hyperspectral image classification (HSI) tasks. To address the aforementioned shortcomings, a Multi-modal Domain Generalization Network (MDG) is proposed to learn cross-domain invariant representation from cross-domain shared semantic space. Only the source domain (SD) is used for training in the proposed method, after which the model is directly transferred to the target domain (TD). Visual and linguistic features are extracted using the dual-stream architecture, which consists of an image encoder and a text encoder. A generator is designed to obtain extended domain (ED) samples that are different from SD. Furthermore, linguistic features are used to construct a cross-domain shared semantic space, where visual-linguistic alignment is accomplished by supervised contrastive learning. Extensive experiments on two datasets show that the proposed method outperforms state-of-the-art approaches. Yuxiang Zhang 0005, Mengmeng Zhang 0005, Wei Li 0032, Ran Tao 0003 |
ICASSP | 1 |
| 2023 | Hyperspectral and LiDAR Data Classification Based on Structural Optimization TransmissionabstractWith the development of the sensor technology, complementary data of different sources can be easily obtained for various applications. Despite the availability of adequate multisource observation data, for example, hyperspectral image (HSI) and light detection and ranging (LiDAR) data, existing methods may lack effective processing on structural information transmission and physical properties alignment, weakening the complementary ability of multiple sources in the collaborative classification task. The complementary information collaboration manner and the redundancy exclusion operator need to be redesigned for strengthening the semantic relatedness of multisources. As a remedy, we propose a structural optimization transmission framework, namely, structural optimization transmission network (SOT-Net), for collaborative land-cover classification of HSI and LiDAR data. Specifically, the SOT-Net is developed with three key modules: 1) cross-attention module; 2) dual-modes propagation module; and 3) dynamic structure optimization module. Based on above designs, SOT-Net can take full advantage of the reflectance-specific information of HSI and the detailed edge (structure) representations of multisource data. The inferred transmission plan, which integrates a self-alignment regularizer into the classification task, enhances the robustness of the feature extraction and classification process. Experiments show consistent outperformance of SOT-Net over baselines across three benchmark remote sensing datasets, and the results also demonstrate that the proposed framework can yield satisfying classification result even with small-size training samples. Mengmeng Zhang 0005, Wei Li 0032, Yuxiang Zhang 0005, Ran Tao 0003, Qian Du 0001 |
IEEE Trans. Cybern. | 3 |
| 2023 | Mask-Reconstruction-Based Decoupled Convolution Network for Hyperspectral Imagery ClassificationabstractDeep learning has attracted much attention in hyperspectral image(HSI) classification. However, most deep learning methods ignore the information loss during spatial-spectral feature extraction, which potentially affects the classification performance. In this article, Mask Reconstruction-Based Decoupled Convolution Network(MrDCN) is proposed, which including the decoupled feature extraction module (DFEM) to extract spectral information and spatial information of target HSI patch respectively. The reconstruction modules are designed to maintain the feature extraction ability of DFEM and ensure that discriminative information in high-dimensional features and low-dimensional features is preserved. MrDCN outperforms state-of-the-art methods in classification on three datasets of various scenarios, which indicates its effectiveness, and experiments on embedded devices are executed to affirm the efficiency of MrDCN. Lujie Song, Mengmeng Zhang 0005, Wei Li 0032, Daguang Jiang, Huan Liu 0015, Yuxiang Zhang 0005 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2023 | Language-Aware Domain Generalization Network for Cross-Scene Hyperspectral Image ClassificationabstractText information including extensive prior knowledge about land cover classes has been ignored in hyperspectral image (HSI) classification tasks. It is necessary to explore the effectiveness of linguistic mode in assisting HSI classification. In addition, the large-scale pretraining image–text foundation models have demonstrated great performance in a variety of downstream applications, including zero-shot transfer. However, most domain generalization methods have never addressed mining linguistic modal knowledge to improve the generalization performance of model. To compensate for the inadequacies listed above, a language-aware domain generalization network (LDGnet) is proposed to learn cross-domain-invariant representation from cross-domain shared prior knowledge. The proposed method only trains on the source domain (SD) and then transfers the model to the target domain (TD). The dual-stream architecture including the image encoder and text encoder is used to extract visual and linguistic features, in which coarse-grained and fine-grained text representations are designed to extract two levels of linguistic features. Furthermore, linguistic features are used as cross-domain shared semantic space, and visual–linguistic alignment is completed by supervised contrastive learning in semantic space. Extensive experiments on three datasets demonstrate the superiority of the proposed method when compared with the state-of-the-art techniques. The codes will be available from the website:https://github.com/YuxiangZhang-BIT/IEEE_TGRS_LDGnet. Yuxiang Zhang 0005, Mengmeng Zhang 0005, Wei Li 0032, Ran Tao 0003 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | Multiple Attention Network for Spartina alterniflora Segmentation Using Multitemporal Remote Sensing ImagesabstractThe semantic segmentation of multi-temporal remote sensing images to construct wetland land surface coverage is the basis for the perception and dynamic modeling of geographic scenes. However, the segmentation of Spartina alterniflora (S.alterniflora) in remote sensing images on wetlands faces the problems such as low level for cooperative interpretation in multi-temporal images and high fragmentation in the distribution of S.alterniflora. To solve the issues, a multiple attention network (MARNet) based on transfer learning is proposed. The method is designed with a plug-and-play attention module to enhance the learning of vegetation features and improve the network’s ability to focus on small areas of S.alterniflora. At the same time, MARNet designs the transfer learning architecture from both inter-domain alignment and intra-domain adaptation perspectives,aligning the statistical distribution by using the maximum mean difference (MMD) between the source and target domains, and entropy minimization within the domain of the target domain to enhance the high confidence prediction of this domain. In addition, since the samples have a serious imbalance problem, redundant cutting and splicing steps are employed for the prediction results to prevent the poor edge prediction of some image blocks. Experimental results on three cross-year RSIs datasets demonstrate that the proposed MARNet performs significantly better than other networks and is able to extract S.alterniflora in wetlands more accurately. Mengmeng Zhang 0005, Jianbu Wang, Xiukai Song, Yuanyuan Gui, Yuxiang Zhang 0005, Wei Li 0032 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2023 | Single-Source Domain Expansion Network for Cross-Scene Hyperspectral Image ClassificationabstractCurrently, cross-scene hyperspectral image (HSI) classification has drawn increasing attention. It is necessary to train a model only on source domain (SD) and directly transferring the model to target domain (TD), when TD needs to be processed in real time and cannot be reused for training. Based on the idea of domain generalization, a Single-source Domain Expansion Network (SDEnet) is developed to ensure the reliability and effectiveness of domain extension. The method uses generative adversarial learning to train in SD and test in TD. A generator including semantic encoder and morph encoder is designed to generate the extended domain (ED) based on encoder-randomization-decoder architecture, where spatial randomization and spectral randomization are specifically used to generate variable spatial and spectral information, and the morphological knowledge is implicitly applied as domain invariant information during domain expansion. Furthermore, the supervised contrastive learning is employed in the discriminator to learn class-wise domain invariant representation, which drives intra-class samples of SD and ED. Meanwhile, adversarial training is designed to optimize the generator to drive intra-class samples of SD and ED to be separated. Extensive experiments on two public HSI datasets and one additional multispectral image (MSI) dataset demonstrate the superiority of the proposed method when compared with state-of-the-art techniques. The codes will be available from the website:https://github.com/YuxiangZhang-BIT/IEEE_TIP_SDEnet. Yuxiang Zhang 0005, Wei Li 0032, Ran Tao 0003, Qian Du 0001 |
IEEE Trans. Image Process. | 1 |
| 2023 | Topological Structure and Semantic Information Transfer Network for Cross-Scene Hyperspectral Image ClassificationabstractDomain adaptation techniques have been widely applied to the problem of cross-scene hyperspectral image (HSI) classification. Most existing methods use convolutional neural networks (CNNs) to extract statistical features from data and often neglect the potential topological structure information between different land cover classes. CNN-based approaches generally only model the local spatial relationships of the samples, which largely limits their ability to capture the nonlocal topological relationship that would better represent the underlying data structure of HSI. In order to make up for the above shortcomings, a Topological structure and Semantic information Transfer network (TSTnet) is developed. The method employs the graph structure to characterize topological relationships and the graph convolutional network (GCN) that is good at processing for cross-scene HSI classification. In the proposed TSTnet, graph optimal transmission (GOT) is used to align topological relationships to assist distribution alignment between the source domain and the target domain based on the maximum mean difference (MMD). Furthermore, subgraphs from the source domain and the target domain are dynamically constructed based on CNN features to take advantage of the discriminative capacity of CNN models that, in turn, improve the robustness of classification. In addition, to better characterize the correlation between distribution alignment and topological relationship alignment, a consistency constraint is enforced to integrate the output of CNN and GCN. Experimental results on three cross-scene HSI datasets demonstrate that the proposed TSTnet performs significantly better than some state-of-the-art domain-adaptive approaches. The codes will be available from the website: https://github.com/YuxiangZhang-BIT/IEEE_TNNLS_TSTnet. Yuxiang Zhang 0005, Wei Li 0032, Mengmeng Zhang 0005, Ying Qu 0001, Ran Tao 0003, Hairong Qi 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2022 | Dual Graph Cross-Domain Few-Shot Learning for Hyperspectral Image ClassificationabstractMost domain adaptation (DA) methods focus on the case where the source data (SD) and target data (TD) with the same classes are obtained by the same sensor in cross-scene hyperspectral image (HSI) classification tasks. However, the classification performance is significantly reduced when there are new classes in TD. In addition, domain alignment is carried out based on local spatial information in most methods, rarely taking into account the non-local spatial information (non-local relationships) with strong correspondence. A Dual Graph Cross-domain Few-shot Learning (DG-CFSL) framework is proposed, trying to make up for the above shortcomings by combining Few-shot Learning (FSL) with domain alignment. Both SD with all label samples and TD with a few label samples are implemented for FSL episodic training. Meanwhile, Intra-domain Distribution Extraction block (IDE-block) is designed to characterize and aggregate the intra-domain non-local relationships. Furthermore, feature- and distribution-level cross-domain graph alignments are used to mitigate the impact of domain shift on FSL. Experimental results on two public HSI data sets demonstrate the effectiveness of the proposed method. Yuxiang Zhang 0005, Wei Li 0032, Mengmeng Zhang 0005, Ran Tao 0003 |
ICASSP | 1 |
| 2022 | Multi-Source Remote Sensing Data Cross Scene Classification Based on Multi-Graph MatchingabstractMulti-source joint classification has been extensively investigated in single scenario setting; however, for cross scene (CS) classification, few studies have been conducted for evaluating the collaborative performance of multi-sources. In this paper, using hyperspectral image (HSI) and light detection and ranging (LiDAR) data, we propose a multi-source CS classification method, and build source-related alignment to reduce statistical shift. Both geometrical and statistical alignments are considered to learn common-subspaces of each source with preserving discrimination information. Finally, the aligned features from both sources are integrated for final classification. Experimental results demonstrate the superior of the proposed method over other state-of-the-art CS approaches. Mengmeng Zhang 0005, Xudong Zhao 0003, Wei Li 0032, Yuxiang Zhang 0005 |
IGARSS | 4 |
| 2021 | Domain Adaptation Based on Graph and Statistical Features for Cross-Scene Hyperspectral Image ClassificationabstractCross-scene hyperspectral image (HSI) classification has gradually received widespread attention, because most models perform unsatisfactory classification performance on training and testing samples from two different scenes. At present, the domain adaptation technique is used to solve this problem, most of which only design models from the level of data statistical features, while ignore the potential topological relationships between the land cover classes. In order to make up for the above shortcoming, a domain adaptation based on graph and statistical features is proposed in the papaer. This method uses convolutional neural network (CNN) extracting features with rich semantic information to dynamically construct graphs, and further introduces graph optimal transport (GOT) to align topological relations to assist distribution alignment based on maximum mean discrepancy (MMD). The experimental results on two cross-scene HSI datasets demonstrate the effectiveness of the proposed method. Yuxiang Zhang 0005, Wei Li 0032, Ran Tao 0003 |
IGARSS | 1 |
| 2021 | Physically Constrained Transfer Learning Through Shared Abundance Space for Hyperspectral Image ClassificationabstractHyperspectral image (HSI) classification is one of the most active research topics and has achieved promising results boosted by the recent development of deep learning. However, most state-of-the-art approaches tend to perform poorly when the training and testing images are on different domains, e.g., the source domain and target domain, respectively, due to the spectral variability caused by different acquisition conditions. Transfer learning-based methods address this problem by pretraining in the source domain and fine-tuning on the target domain. Nonetheless, a considerable amount of data on the target domain has to be labeled and nonnegligible computational resources are required to retrain the whole network. In this article, we propose a new transfer learning scheme to bridge the gap between the source and target domains by projecting the HSI data from the source and target domains into a shared abundance space based on their own physical characteristics. In this way, the domain discrepancy would be largely reduced such that the model trained on the source domain could be applied to the target domain without extra efforts for data labeling or network retraining. The proposed method is referred to as physically constrained transfer learning through shared abundance space (PCTL-SAS). Extensive experimental results demonstrate the superiority of the proposed method as compared to the state of the art. The success of this endeavor would largely facilitate the deployment of HSI classification for real-world sensing scenarios. Ying Qu 0001, Razieh Kaviani Baghbaderani, Wei Li 0032, Lianru Gao, Yuxiang Zhang 0005, Hairong Qi 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2021 | Cross-Scene Hyperspectral Image Classification With Discriminative Cooperative AlignmentabstractCross-scene classification is one of the major challenges for hyperspectral image (HSI) classification, especially for target scenes without label samples. Most traditional domain adaptive methods learn a domain invariant subspace to reduce statistical shift while ignoring the fact that there may not exist a shared subspace when marginal distributions of source and target domains are very different. In addition, it is important for HSI classification to preserve discriminant information in the original space. To solve this issue, discriminative cooperative alignment (DCA) of subspace and distribution is proposed to cooperatively reduce the geometric and statistical shift. In the proposed framework, both geometrical and statistical alignments are considered to learn subspaces of the two domains with preserving discrimination information. Furthermore, a reconstruction constraint is imposed to enhance the robustness of subspace projection. Experimental results on three cross-scene HSI data sets demonstrate that the proposed DCA is significantly better than some state-of-the-art domain-adaptive approaches. Yuxiang Zhang 0005, Wei Li 0032, Ran Tao 0003, Jiangtao Peng, Qian Du 0001, Zhaoquan Cai 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2020 | Discriminative Marginalized Least-Squares Regression for Hyperspectral Image ClassificationabstractLeast-squares regression (LSR)-based classifiers are effective in multiclassification tasks. However, most existing methods use limited projections, resulting in loss of much discriminant information; furthermore, they focus only on exactly fitting samples to target matrix while ignoring overfitting issue. To solve these drawbacks, discriminative marginalized LSR (DMLSR) is proposed to learn a more discriminative projection matrix with consideration of class separability and data-reconstruction ability simultaneously. In the proposed framework, an intraclass compactness graph is employed to avoid the overfitting problem and enhance class separability, and a data-reconstruction constraint is imposed to preserve discriminant information on limited projections. Experimental results on several hyperspectral data sets demonstrate that the proposed method significantly outperforms some state-of-the-art classifiers. Yuxiang Zhang 0005, Wei Li 0032, Heng-Chao Li 0001, Ran Tao 0003, Qian Du 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2019 | Structure-Aware Collaborative Representation for Hyperspectral Image ClassificationabstractRecently, collaborative representation (CR) has drawn increasing attention in hyperspectral image classification due to its simplicity and effectiveness. However, existing representation-based classifiers do not explicitly utilize class label information of training samples in estimating representation coefficients. To solve this issue, a structure-aware CR with Tikhonov regularization (SaCRT) method is proposed to consider both class label information of training samples and spectral signatures of testing pixels to estimate more discriminative representation coefficients. In the proposed framework, marginal regression is employed; furthermore, an interclass row-sparsity structure is designed to preserve the compact relationship among intraclass pixels and more separable interclass pixels, thereby enhancing class separability. The experimental results evaluated using three hyperspectral data sets demonstrate that the proposed method significantly outperforms some state-of-the-art classifiers. Wei Li 0032, Yuxiang Zhang 0005, Na Liu 0014, Qian Du 0001, Ran Tao 0003 |
IEEE Trans. Geosci. Remote. Sens. | 2 |