EDBT 2026 Demo / reviewers in the wild / expert
Bobo Xi
dblp:213/5876
· DBLP profile ↗
31ranked-venue papers
13as first author
30since 2021 · last 2026
0000-0002-0309-1556ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 22 · 7 first-author · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HyperPRET: Few-shot class incremental learning with precognition and retrospection for hyperspectral imagery
Bobo Xi, Tie Zheng, Jiaojiao Li 0001, Shou Feng, Yunsong Li 0001 |
Pattern Recognit. | 2 |
| 2025 | DSNet: Dynamic Stitchable Neural Network for Hyperspectral Image ClassificationabstractHyperspectral image classification (HSIC) aims to identify land cover categories by leveraging the spectral and spatial information contained in hyperspectral images (HSI). Currently, many deep learning approaches utilize dual-branch networks to process spectral and spatial data separately, followed by the application of specialized modules to facilitate feature interaction or fusion. However, the design of these modules demands considerable time and effort from researchers and may not adequately capture the inherent relationships between independent spatial and spectral features in a dynamic manner. To address these issues, we propose the dynamic stitchable neural network (DSNet) for HSIC. While the DSNet maintains a dual-branch structure, it operates without traditional feature fusion or interaction. Instead, it employs a stitching network approach to integrate the two branches. Specifically, a spatial-spectral stitching module is presented to incorporates multiple stitching layers at various positions between the two network branches, creating new stitched networks that retain the strengths of both original networks. Additionally, a reinforcement learning-based strategy is designed for dynamically selecting stitching positions tailored to specific datasets, enabling the model to adaptively optimize the integration of spatial and spectral features. Recognizing the effectiveness of vision transformer (ViT) in learning spatial information and the capability of 1D convolutional neural network (1DCNN) in capturing spectral details, the DSNet directly stitches these two networks together. This fusion maximizes the utilization of both foundational networks, yielding a new hybrid network that delivers exceptional performance while also alleviating the burden on researchers to develop new architectures from scratch. Extensive experiments and analyses conducted on three public HSI datasets demonstrate the superiority of the proposed method, validating the effectiveness of our innovative modules. The codes of this work will be available from the website: https://github.com/ZZC/IEEE-TGRS-DSNet. Shou Feng, Zicheng Zhao, Bobo Xi, Chunhui Zhao 0003, Wei Li 0032, Ran Tao 0003, Yunsong Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | SemiBaCon: Semi-Supervised Balanced Contrastive Learning for Multimodal Remote Sensing Image ClassificationabstractThe limited availability of annotated training data significantly constrains the classification accuracy of hyperspectral image (HSI) and LiDAR fusion approaches. Although contrastive learning has emerged as a potential solution, current implementations frequently neglect the critical class imbalance issues during unlabeled sample selection. To address the issue, we introduce a novel semi-supervised balanced contrastive learning (SemiBaCon) framework for multi-modal remote sensing image classification. First, we propose a superpixel-based balanced sampling (SPBS) mechanism that fundamentally addresses class imbalance through intelligent pseudo-label generation. By segmenting the HSI data into homogeneous superpixels and implementing intra-region label propagation, the method ensures statistically balanced pseudo-label selection across categories, effectively overcoming the bias introduced by conventional random sampling strategies. Second, our architecture integrates a dual-stream encoder combining convolutional neural networks (CNNs) with Transformers, enabling hierarchical feature extraction from spectral-spatial characteristics of HSI and elevation patterns of LiDAR. This design facilitates the construction of multi-modal positive sample pairs, achieving enhanced representation learning through inter-modal consistency constraints. Third, we develop a pseudo-label guided contrastive learning (PLCL) paradigm that synergistically combines pseudo-label confidence with feature similarity metrics, which effectively reduces intra-class variance and improves decision boundaries in the latent space. Comprehensive evaluations on three benchmark datasets demonstrate the framework’s superior performance compared to the state-of-the-art methods. Bobo Xi, Tie Zheng, Yunsong Li 0001, Changbin Xue, Ming Shen 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | A Multimodal Cross-Domain Segmentation Network of Remote Sensing Imagery With Multilevel Deep Cross-Fusion and Adversarial Domain AdaptationabstractDeep learning techniques have recently achieved significant success in semantic segmentation for Earth observation tasks using single-modality data within a specific region. However, such optimal conditions are seldom found in real-world applications, leading to subpar performance and limited generalization of existing models when faced with multimodal cross-domain scenarios. To tackle this issue, we introduce a novel segmentation network for remote sensing imagery (RSI) called MC-Seg, which incorporates multilevel deep cross-fusion and adversarial domain adaptation. The MC-Seg framework begins by performing multilevel deep cross-fusion of multimodal data through a complete feature extraction module and three fusion levels: feature, modal, and layer. This approach ensures the thorough utilization of features from various modalities. Following this, the framework integrates a local-global adversarial domain adaptation module to minimize the domain gap, effectively aligning the source and target domains despite the differences in RSI data representations. Experimental validation on the C2Seg dataset shows that our proposed method outperforms existing state-of-the-art techniques. Bobo Xi, Tie Zheng, Yunsong Li 0001, Changbin Xue |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | DIFTransNet: Dual-Branch Interactive Fusion Network With CNN and Multiscale Transformer for Infrared Small Target DetectionabstractRecent advances in hybrid architectures combining convolutional neural networks (CNNs) and Transformers have demonstrated significant potential in infrared small target detection. However, existing methods suffer from limited interaction depth, leading to weak cross-layer feature coordination. To address this, we propose the DIFTransNet, a dual-branch interactive fusion network that implements hierarchical mutual refinement between CNN and Transformer pathways. Specifically, the DIFTransNet incorporates an interactive fusion module (IFM) at each hybrid encoding stage that bridges local details and global semantics through shared latent space projection and cross-attention mechanisms, enabling aligned deep fusion. In the Transformer branch, we introduce the hierarchical decoupled sparse attention (HDSA) with resolution-progressive window partitioning, which preserves dim targets through dense local sampling while suppressing noise via cross-region sparse correlations. Moreover, rethinking the information gap in conventional skip connections between shallow and deep layers, we propose a reorganized skip augmentation module (RSAM). It introduces a cross-level bidirectional fusion strategy, where the closed-loop architecture enables dual compensation between spatial details and semantic contexts. The experimental results on the NUDT-SIRST, SIRST, and IRSTD-1K datasets demonstrate that the proposed DIFTransNet surpasses current state-of-the-art methods in accurate detection of infrared small targets. Shenao Liu, Bobo Xi, Tie Zheng, Jiaojiao Li 0001, Yunsong Li 0001, Changbin Xue |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | HyLiOSR: Staged Progressive Learning for Joint Open-Set Recognition of Hyperspectral and LiDAR DataabstractThe joint classification of hyperspectral images (HSIs) and light detection and ranging (LiDAR) data have seen significant advancements in recent research. However, it would be more practical if we could simultaneously detect the unknown classes in a more realistic open-set scenario. In this article, we introduce a novel open-set recognition (OSR) method for HSI and LiDAR data, termed HyLiOSR, which devises a staged progressive learning strategy to effectively bridge the gap between closed-set and open-set feature distributions within an autoencoder framework. Specifically, for the first stage, the reconstruction-based network is dedicated to accurately modeling each known category by learning multiple Gaussian prototypes, which facilitates OSR by disentangling the distribution of known classes. In the second stage, we actively synthesize samples of unknown classes during the feature extraction phase and create a virtual unknown classifier, enabling the network to effectively differentiate between known and unknown class samples. This approach establishes a distinct separation between known and unknown classes in the latent feature space, thereby enhancing the capability of the frameworks to distinguish between them. Comprehensive experiments conducted on three benchmark datasets demonstrate that the proposed HyLiOSR outperforms existing state-of-the-art methods. The source code will be accessible athttps://github.com/B-Xi/TGRS_2025_HyLiOSR. Bobo Xi, Mingshuo Cai, Jiaojiao Li 0001, Zhengjue Wang, Shou Feng, Yunsong Li 0001, Jocelyn Chanussot |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2025 | MCTGCL: Mixed CNN-Transformer for Mars Hyperspectral Image Classification With Graph Contrastive LearningabstractHyperspectral image (HSI) classification has been extensively studied in the context of Earth observation. However, its application in Mars exploration remains limited. Although convolutional neural networks (CNNs) have proven effective in HSI processing, their local receptive fields hinder their ability to capture long-range features. Transformers excel in global modeling and perform well in HSI classification (HSIC), but they often neglect the effective representation of local spectral and spatial features and tend to be more complex. To address these challenges, we propose a mixed CNN-transformer network for Mars HSI classification with graph contrastive learning to enhance classification performance. Specifically, we introduce an information-enhanced attention module (IEAM) designed to aggregate attention features from multiple perspectives. Additionally, we develop a lightweight dual-branch CNN-transformer (LDCT) network that efficiently extracts both local and global spectral-spatial features with lower complexity. To improve the discrimination of inter-class features, we apply graph contrastive learning to the topological structure of labeled samples. Furthermore, we annotated three Mars HSI datasets, referred to as HyMars, to validate the effectiveness of our proposed mixed CNN–“transformer network for Mars HSIC with graph contrastive learning (MCTGCL). Comprehensive experimental results across different amounts of labeled samples consistently demonstrate the superiority of the method. The source code is available athttps://github.com/B-Xi/TGRS_2025_MCTGCL. Bobo Xi, Jiaojiao Li 0001, Tie Zheng, Xunfeng Zhao, Changbin Xue, Yunsong Li 0001, Jocelyn Chanussot |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2025 | Transductive Few-Shot Learning With Enhanced Spectral-Spatial Embedding for Hyperspectral Image ClassificationabstractFew-shot learning (FSL) has been rapidly developed in the hyperspectral image (HSI) classification, potentially eliminating time-consuming and costly labeled data acquisition requirements. Effective feature embedding is empirically significant in FSL methods, which is still challenging for the HSI with rich spectral-spatial information. In addition, compared with inductive FSL, transductive models typically perform better as they explicitly leverage the statistics in the query set. To this end, we devise a transductive FSL framework with enhanced spectral-spatial embedding (TEFSL) to fully exploit the limited prior information available. First, to improve the informative features and suppress the redundant ones contained in the HSI, we devise an attentive feature embedding network (AFEN) comprising a channel calibration module (CCM). Next, a meta-feature interaction module (MFIM) is designed to optimize the support and query features by learning adaptive co-attention using convolutional filters. During inference, we propose an iterative graph-based prototype refinement scheme (iGPRS) to achieve test-time adaptation, making the class centers more representative in a transductive learning manner. Extensive experimental results on four standard benchmarks demonstrate the superiority of our model with various handfuls (i.e., from 1 to 5) labeled samples. The code will be available online at https://github.com/B-Xi/TIP_2025_TEFSL. Bobo Xi, Jiaojiao Li 0001, Yan Huang 0018, Yunsong Li 0001, Zan Li 0001, Jocelyn Chanussot |
IEEE Trans. Image Process. | 1 |
| 2025 | HyperTaFOR: Task-Adaptive Few-Shot Open-Set Recognition With Spatial-Spectral Selective Transformer for Hyperspectral ImageryabstractOpen-set recognition (OSR) aims to accurately classify known categories while effectively rejecting unknown negative samples. Existing methods for OSR in hyperspectral images (HSI) can be generally divided into two categories: reconstruction-based and distance-based methods. Reconstruction-based approaches focus on analyzing reconstruction errors during inference, whereas distance-based methods determine the rejection of unknown samples by measuring their distance to each prototype. However, these techniques often require a substantial amount of training data, which can be both time-consuming and expensive to gather, and they require manual threshold setting, which can be difficult for different tasks. Furthermore, effectively utilizing spectral-spatial information in HSI remains a significant challenge, particularly in open-set scenarios. To tackle these challenges, we introduce a few-shot OSR framework for HSI named HyperTaFOR, which incorporates a novel spatial-spectral selective transformer (S3Former). This framework employs a meta-learning strategy to implement a negative prototype generation module (NPGM) that generates task-adaptive rejection scores, allowing flexible categorization of samples into various known classes and anomalies for each task. Additionally, the S3Former is designed to extract spectral-spatial features, optimizing the use of central pixel information while reducing the impact of irrelevant spatial data. Comprehensive experiments conducted on three benchmark hyperspectral datasets show that our proposed method delivers competitive classification and detection performance in open-set environments when compared to state-of-the-art methods. The code is available online at https://github.com/B-Xi/TIP_2025_HyperTaFOR. Bobo Xi, Jiaojiao Li 0001, Rui Song 0003, Yunsong Li 0001 |
IEEE Trans. Image Process. | 1 |
| 2025 | HyperCASR: Spectral-Spatial Open-Set Recognition With Category-Aware Semantic Reconstruction for Hyperspectral ImageryabstractOpen-set recognition (OSR) in hyperspectral imagery (HSI) focuses on accurately classifying known classes while effectively rejecting unknown negative samples. Most existing reconstruction-based approaches are susceptible to noise interference in the input images, and known classes can easily lead to inter-class confusion during the reconstruction process. Moreover, effectively utilizing the abundant spectral-spatial information in HSI within an open-set context presents significant challenges. To address these issues, we propose HyperCASR, an innovative framework for HSI OSR that integrates a grouped spectral-spatial retentive transformer (GSSRT) and a class-aware semantic reconstruction (CASR) module. This method begins by designing the GSSRT to extract features from HSI, enhancing the extraction capability of spatial-spectral information by introducing a grouped pixel embedding (GPE) module and a novel spatial retentive attention (SRA) mechanism. Subsequently, an independent autoencoder (AE) is assigned to each known class to reconstruct semantic features, which helps to mitigate noise interference and inter-class confusion. Additionally, by minimizing reconstruction errors to estimate class affiliation, the framework effectively identifies unknown classes. Experimental results across three benchmark datasets indicate that the HyperCASR framework significantly enhances classification performance for both known and unknown classes when compared to existing state-of-the-art methods. The code is available at https://github.com/B-Xi/TIP_2025_HyperCASR. Bobo Xi, Jiaojiao Li 0001, Rui Song 0003, Yunsong Li 0001 |
IEEE Trans. Image Process. | 1 |
| 2025 | FDGNet: Frequency Disentanglement and Data Geometry for Domain Generalization in Cross-Scene Hyperspectral Image ClassificationabstractCross-scene hyperspectral image classification (HSIC) poses a significant challenge in recognizing hyperspectral images (HSIs) from different domains. The current mainstream approaches based on domain adaptation (DA) methods need to access target data when aligning distributions between domains, limiting the applicability of the model. In contrast, recent domain generalization (DG) methods aim to directly generalize to unseen domains, eliminating the requirements for target data during training. Nonetheless, most DG-based methods overly focus on randomizing sample styles, leading to semantically compromised samples. In addition, broadening the source distribution without ensuring reasonable support may result in undesired extended distributions. To address these issues, we propose a novel DG network with frequency disentanglement and data geometry (FDGNet) for cross-scene HSIC. Specifically, we first develop a spectral-spatial encoder based on frequency disentanglement (FDSS encoder), which facilitates synthesized domains to preserve their semantic consistency while simulating interdomain gaps with the source domain. Second, to avoid the generation of unrealistic samples, we incorporate data geometry into adversarial training. This helps diversify new domains while keeping the data geometry of extended domains in an explainable support. To improve the learning of domain-invariant representation, we propose an intermediate domain sampling strategy based on the class-wise perceptual manifold. This strategy synthesizes reliable intermediate domains by sampling from class-wise manifold flows estimated over the source and extended domains. Extensive experiments and analysis on three public HSI datasets yield the superiority of our proposed FDGNet. The codes will be available from the website: https://github.com/Qba-heu/FDGNet. Boao Qin, Shou Feng, Chunhui Zhao 0003, Bobo Xi, Wei Li 0032, Ran Tao 0003 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Diamond-Unet: A Novel Semantic Segmentation Network Based on U-Net Network and Transformer for Deep Space Rock ImagesabstractExtracting rock objects from the surface of celestial bodies in deep space exploration environments is crucial for self-service path planning, navigation of detectors, and regional information evaluation. Most existing image saemantic segmentation frameworks decrease the spatial resolution of the feature maps as networks deepen, resulting in limitations in detecting small targets and the inability to accurately segment boundary regions. In this letter, we propose a novel semantic segmentation network based on U-Net network and Transformer for deep space rock images, referred to as Diamond-Unet. This model integrates overcomplete and undercomplete branches and incorporates a global-local feature extraction (GLFE) module based on Transformer and CNN technologies to effectively capture discriminative information. Furthermore, an innovative feature cross-fusion path (FCFP) is introduced to enhance information exchange between the dual-branch networks, enabling the capture of both fine-grained details and coarse-grained semantics in the full-scale image segmentation architecture. Experimental results demonstrate that the Diamond-Unet achievesMIoUscores of 79.32% and 93.43% on two public datasets, which are superior to the compared methods. Bobo Xi, Tie Zheng, Yunsong Li 0001, Changbin Xue, Jocelyn Chanussot |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2024 | Small Object-Aware Video Coding for Machines via Feature-Motion SynergyabstractVideo coding for machines (VCM) is a rapidly growing field dedicated to bridging the gap between video and feature coding. For storage-intensive aerial videos, VCM offers valuable insights into a more efficient coding paradigm. However, the frequent occurrence of small objects poses a challenge to VCM, with limited distinctive features and inherent distortion in the reconstructed videos. To address this issue, we propose small object-aware VCM (SOAVCM), a joint video and feature coding approach that handles small objects. Particularly, the video coding incorporates a feature-guided residual (FGR) codec to preserve the small objects, utilizing features obtained from feature coding. Simultaneously, feature coding employs the motion vector (MV) estimated in video coding to generate compact high-level features. By leveraging the inherent synergy between features and MVs, SOAVCM significantly enhances overall coding efficiency. Experimental results demonstrate that SOAVCM outperforms several deep-learning-based methods and traditional coding standards in video coding. Moreover, the encoded feature representation improves detection accuracy and achieves substantial bitrate savings. Qihan Xu, Bobo Xi, Yunsong Li 0001, Jocelyn Chanussot |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2024 | Mind the Gap: Multilevel Unsupervised Domain Adaptation for Cross-Scene Hyperspectral Image ClassificationabstractRecently, cross-scene hyperspectral image classification (HSIC) has attracted increasing attention, alleviating the dilemma of no labeled samples in the target domain. Although collaborative source and target training has dominated this field, training effective feature extractors and overcoming intractable domain gaps remains challenging. To cope with this issue, we propose a multi-level unsupervised domain adaptation (MLUDA) framework, which comprises image-, feature-, and logic-level alignment between domains to fully investigate the comprehensive spectral-spatial information. Specifically, at the image level, we propose an innovative domain adaptation method named GuidedPGC based on classic image matching techniques and the guided filter. The adaptation results are physically explainable with intuitive visual observations. Regarding the feature level, we design a multi-branch cross attention structure (MBCA) specifically for HSIC, which enhances the interaction between the features from the source and target domains through dot-product attention. Finally, at the logic level, we adopt a supervised contrastive learning (SCL) approach that incorporates a pseudo-label strategy and local maximum mean discrepancy loss, increasing inter-class distance across diverse domains and further improving the classification performance. Experimental results on three benchmark cross-scene datasets demonstrate that our proposed method consistently outperforms the compared approaches. The source code is available at https://github.com/cfcys/MLUDA. Mingshuo Cai, Bobo Xi, Jiaojiao Li 0001, Shou Feng, Yunsong Li 0001, Zan Li 0001, Jocelyn Chanussot |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Cross-Domain Few-Shot Learning Based on Decoupled Knowledge Distillation for Hyperspectral Image ClassificationabstractExisting cross-domain few-shot learning (FSL) methods for hyperspectral image (HSI) classification have garnered widespread attention due to their excellent performance in recognizing novel classes. To mitigate domain shift, researchers focus on designing sophisticated domain adaptation (DA) modules to directly apply biased metaknowledge in the target domain (TD). However, this paradigm proves somewhat inadequate in the face of significant differences in distribution. To cope with this dilemma, we adopted a new mindset of treating metaknowledge extraction and debiasing from the source domain (SD) as a synergistic process and proposed a cross-domain FSL framework based on decoupled knowledge distillation for HSI classification (HSIC). In general, to efficiently acquire and utilize unbiased metaknowledge, this framework centralizes on a knowledge distillation (KD) strategy. Through the effective information transfer process, the extraction and debiasing of metaknowledge were integrated into a comprehensive and productive process. Simultaneously, to release the constraints imposed by the coupled logits in the KD process on the knowledge interaction, the decoupled logit interaction (DLI) module is employed in the framework. This module decouples the traditional KD into two controllable components, making a more balanced and comprehensive interaction of task-related knowledge and data-intrinsic knowledge between models. Moreover, to facilitate the extraction of critical discriminative metaknowledge from the abundant redundant information in HSI, the discriminative information refinement (DIR) module is designed to develop distinctive features for similar bands. Extensive experiments on three public HSI datasets exhibited the superior performance of the proposed cross-domain few-shot learning method based on decoupled knowledge distillation for HSIC (DKD-FSL) method in comparison with seven state-of-the-art approaches. Shou Feng, Hongzhe Zhang, Bobo Xi, Chunhui Zhao 0003, Yunsong Li 0001, Jocelyn Chanussot |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Multilevel Attention Dynamic-Scale Network for HSI and LiDAR Data Fusion ClassificationabstractLand use/land cover classification with multimodal data has attracted increasing attention. For hyperspectral images (HSIs) and light detection and ranging (LiDAR) data, the combination of them can make the classification more accurate and robust. However, how to effectively utilize their respective strengths and integrate them with the classification task is still a challenging problem. In this article, a multilevel attention dynamic-scale network (MADNet) is proposed. First, in the feature extraction stage, the two modalities are divided into two branches with different scales, which are then fed into the convolutional neural networks (CNNs) to learn shallow features. Then, considering the characteristics of the HSI, a spectral angle attention module (SAAM) with low-level attention is designed to highlight surrounding pixels that have similar spectra to the central pixel of the patch. After that, a dynamic-scale selection module (DSSM) is proposed to screen an appropriate scale for the patches by pixel similarity analysis. Next, combining the Transformer and the CNN, a global-local cross-attention module (GLCAM) is devised to investigate the fused deep-level multimodal features. Distinct from the vanilla Transformer, the GLCAM deploys a distance-weight operator to decrease the redundancies at long distances and effectively reduce misclassifications. Extensive experiments on three paired HSI and LiDAR datasets demonstrate that the proposed MADNet has certain advantages over the existing methods. Bobo Xi, Tie Zheng, Yunsong Li 0001, Changbin Xue, Jocelyn Chanussot |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Hyperspherical Structural-Aware Distillation Enhanced Spatial-Spectral Bidirectional Interaction Network for Hyperspectral Image ClassificationabstractThe existing methods for hyperspectral image classification (HSIC) mainly focus on the extraction of spectral and spatial features while paying less attention to the interaction of each other. Besides, most of them directly use a parameterized classifier as the final layer of the network. While this design is convenient for end-to-end optimization with the backbone, it overlooks the utilization of the metric space. In this article, a novel hyperspherical structural-aware distillation enhanced spatial–spectral bidirectional interaction network (HSDBIN) is proposed for HSIC. HSDBIN uses a dual-branch design combining the 1-D CNN and transformer to separately learn the detailed spectral correlations and global spatial relationships in parallel. Then, by interacting and aggregating the independent information between two parallel branches, a bidirectional interaction block across branches is designed to explore complementary clues between spectral and spatial pipelines. Finally, to enhance the utilization of metric space and keep compact intraclass relationship, we propose a hyperspherical structural-aware distillation (HSD) to transfer the geometric relationship of hyperspherical space into the metric space of output logits. Extensive experiments and analysis on three public HSI datasets suggest the superiority of the proposed method and verify the effectiveness of the proposed modules. Boao Qin, Shou Feng, Chunhui Zhao 0003, Bobo Xi, Wei Li 0032, Ran Tao 0003, Yunsong Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | A Lightweight Framework With Knowledge Distillation for Zero-Shot Mars Scene ClassificationabstractGathering extensive labeled data during Mars missions is costly and unrealistic, especially considering the complex and unpredictable Martian environment where new and unfamiliar scenes may emerge. Traditional Mars scene classification (MSC) methods depend heavily on large amounts of labeled data, which makes it impractical to recognize previously unseen scene classes without the necessary labeled examples. In addition, the significant computational demands and parameter requirements of modern models also pose challenges for their integration into resource-constrained systems used in Mars exploration. To address these issues, we propose a zero-shot MSC (ZSMSC) framework, which is able to categorize unseen Martian image scenes without the prior acquisition of vast visual examples. Specifically, the framework combines lightweight model design with knowledge distillation (KD) techniques, known as KDMSC, to streamline complex zero-shot learning (ZSL) models. It employs a KD loss that captures essential knowledge through the training of the teacher model from scratch, thereby improving the zero-shot classification performance of the student model. Consequently, the lightweight student model is tailored for deployment on devices with limited resources while fulfilling the requirements of the ZSMSC tasks. Moreover, to support the ZSMSC initiative, we developed a dataset named ZSMars to further advance this field. Experimental results indicate that our model excels in the ZSMSC tasks while maintaining low computational complexity and storage requirements. Xiaomeng Tan, Bobo Xi, Jiaojiao Li 0001, Yunsong Li 0001, Changbin Xue, Jocelyn Chanussot |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | CTF-SSCL: CNN-Transformer for Few-Shot Hyperspectral Image Classification Assisted by Semisupervised Contrastive LearningabstractFew-shot learning (FSL) has rapidly advanced in the hyperspectral image classification (HSIC), potentially reducing the need for laborious and expensive labeled data collection. Due to the limited receptive field, the convolutional neural network (CNN) struggles to capture long-range dependencies for extracting global features. Additionally, the transformer focuses on global correlation while overlooking the effective representation of local spatial and spectral features. Moreover, contrastive learning (CL) has emerged as a powerful technique for improving consistency across different augmented views of samples of the same category. To this end, we devise a novel CNN-Transformer for few-shot HSIC assisted by semisupervised contrastive learning, named CTF-SSCL, to boost the classification performance. Specifically, the cascaded CNN-Transformer incorporates a lightweight spatial-spectral interactive convolution module (LSSICM) and a multiscale transformer (MSFormer) to exploit local features from submaps and global information from the entire patch. Subsequently, the semisupervised contrastive loss, comprising unsupervised and supervised components, serves as an auxiliary to optimize the model with the classification loss. Wherein, recognizing the unified spectral-spatial information in HSI, we propose a spectral feature shift strategy (SFSS) to create sample pairs for the unsupervised CL, utilizing unsupervised contrastive loss among groups of samples with identical labels. Extensive experiments on four standard benchmarks demonstrate the effectiveness of the proposed CTF-SSCL with varying amounts of labeled samples. The code will be available online athttps://github.com/B-Xi/CTF-SSCL. Bobo Xi, Jiaojiao Li 0001, Yunsong Li 0001, Zan Li 0001, Jocelyn Chanussot |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | DGSSC: A Deep Generative Spectral-Spatial Classifier for Imbalanced Hyperspectral ImageryabstractIn recent years, hyperspectral image classification (HSIC) has achieved impressive progress with emerging studies on deep learning models. However, the classification performance downgrades due to the limited number of annotated samples, especially for minority classes. Notably, the imbalanced data dilemma is familiar in remote sensing hyperspectral image because the ground objects are commonly distributed without evenness. Therefore, this paper proposes a novel deep generative spectral-spatial classifier (DGSSC) for addressing the issues of imbalanced HSIC. Specifically, the DGSSC comprises three components, a two-stage encoder, a decoder, and a classifier, which are trained in an end-to-end manner. In particular, to exploit the abundant spectral-spatial features with relatively low computational complexity, the first stage of the encoder comprises successive three-dimensional (3D) and two-dimensional (2D) convolutions, exploring the spectral-spatial and deep spatial information. In addition, the second stage involves the deep latent variable model to achieve minority-class data augmentation. Furthermore, a patch distance-based reconstruction loss function is designed to facilitate the outputs of the decoder being more similar to the input 3D patch samples. The proposed DGSSC can outperform the state-of-the-art methods on three benchmark datasets, especially with its more robust prediction results. For instance, the DGSSC achieves a remarkable 97.85% mean overall accuracy with 0.24% standard deviation over ten independent runs with randomly selected imbalanced 1% training samples on the University of Pavia dataset. Bobo Xi, Jiaojiao Li 0001, Yan Diao, Yunsong Li 0001, Zan Li 0001, Yan Huang 0018, Jocelyn Chanussot |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | Class-Specific Autoaugment Architecture Based on Schmidt Mathematical Theory for Imbalanced Hyperspectral ClassificationabstractHyperspectral image classification (HSIC) often suffers from severe imbalanced category distribution in real applications, which causes bias toward the dominated categories. As an effective method, the deep generative model (DGM) can be used to augment the features of imbalanced data through a learnable method to achieve superior classification performance. However, the features extracted by DGM are preset as a standard Gaussian distribution which results in low interclass difference. Besides, the generated features are too consistent with the original ones, which cannot play a positive role in the discriminability of minority categories (MCs). To conquer these drawbacks, we propose a class-specific autoaugment architecture based on Schmidt mathematical theory (CACS) for the challenging of imbalanced data which consists of two stages: one is training a superior features extractor, and the other one is augmenting features. The class-specific features of the whole HSI are extracted in stage one that supports the following feature augmented module. Specifically, we weighted the classifier in the first phase according to cost-sensitive learning, to prevent the classifier from overfitting. To expand the dispersion between categories, we construct feature prototypes obeying different Gaussian distributions for each class, respectively, and generate class-specific features. Then, the features are augmented in the second phase based on Schmidt’s mathematical theory, which enhances the discriminability of minority class features, thus further improving the classification accuracy with interpretability. Extensive experimental results on three benchmarking datasets demonstrate that CACS is outstanding in comparison algorithms, especially in MCs. Jiaojiao Li 0001, Yan Diao, Rui Song 0003, Bobo Xi, Yunsong Li 0001, Qian Du 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Semisupervised Cross-Scale Graph Prototypical Network for Hyperspectral Image ClassificationabstractIn practice, the acquirement of labeled samples for hyperspectral image (HSI) is time-consuming and labor-intensive. It frequently induces the trouble of model overfitting and performance degradation for the supervised methodologies in HSI classification (HSIC). Fortunately, semisupervised learning can alleviate this deficiency, and graph convolutional network (GCN) is one of the most effective semisupervised approaches, which propagates the node information from each other in a transductive manner. In this study, we propose a cross-scale graph prototypical network (X-GPN) to achieve semisupervised high-quality HSIC. Specifically, considering the multiscale appearance of the land covers in the same remotely captured scene, we involve the neighborhoods of different scales to construct the adjacency matrices and simultaneously design a multibranch framework to investigate the abundant spectral-spatial features through graph convolutions. Furthermore, to exploit the complementary information between different scales, we simply employ the standard 1-D convolution to excavate the dependence of the intranode and concatenate the output with the features generated from other scales. Intuitively, different branches for various samples should have different importance to predict their categories. Thus, we develop a self-branch attentional addition (SBAA) module to adaptively highlight the most critical features produced by multiple branches. In addition, different from previous GCN for HSIC, we devise an innovative prototypical layer comprising a distance-based cross-entropy (DCE) loss function and a novel temporal entropy-based regularizer (TER), which can enhance the discrimination and representativeness of the node features and prototypes actively. Extensive experiments demonstrate that the proposed X-GPN is superior to the classic and state-of-the-art (SOTA) methods in terms of the classification performance. Bobo Xi, Jiaojiao Li 0001, Yunsong Li 0001, Rui Song 0003, Yuchao Xiao, Qian Du 0001, Jocelyn Chanussot |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2022 | A Triplet Semisupervised Deep Network for Fusion Classification of Hyperspectral and LiDAR DataabstractData fusion of hyperspectral and light detection and ranging (LiDAR) is conducive to obtain more comprehensive surface information and thereby achieve better classification result in Earth Monitoring Systems. However, lack of labeled samples usually limits the performance of supervised classifiers, and the heterogeneity of multi-source data also brings great challenges to data fusion. Aiming to address these issues, we propose a triplet semi-supervised deep convolutional neural network (TSDN) for fusion classification of hyperspectral and LiDAR. Specifically, we utilize three basic pathways to extract deep learning features: 1D-CNN for spectral features in hyperspectral, 2D-CNN for spatial features in hyperspectral and Cascade Net for elevation features in LiDAR data. Furthermore, a novel label calibration module (LCM) is proposed to generate effective pseudo labels with high confidence based on the superpixel segmentation by comparing the multi-view classification results for assisting semi-supervised model training. In addition, we design a novel 3D-Cross Attention Block to enhance the complementary spatial features of multi-source data. Experiments on three public HSI-LiDAR benchmarks: Houston, Trento, and MUUFL Gulfport have demonstrated the effectiveness and superiority of our proposed method. Jiaojiao Li 0001, Yinle Ma, Rui Song 0003, Bobo Xi, Danfeng Hong, Qian Du 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Corrections to "Multiscale Context-Aware Ensemble Deep KELM for Efficient Hyperspectral Image Classification"abstractIn the above article[1],Fig. 19was incorrectly placed. The correct image and caption are provided here: Bobo Xi, Jiaojiao Li 0001, Yunsong Li 0001, Rui Song 0003, Weiwei Sun 0005, Qian Du 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Multi-Direction Networks With Attentional Spectral Prior for Hyperspectral Image ClassificationabstractConvolutional neural networks (CNNs) have achieved prominent progress in recent years and demonstrated remarkable properties in spectral–spatial hyperspectral image (HSI) classification. However, conventional spatial-context-based CNNs commonly adopt the single patchwise scheme to represent the to-be-classified samples, which often fails to completely investigate the wealthy spectral–spatial information in complicated situations. For instance, it has great probability to cause misclassifications on the irregular or inhomogeneous areas, especially for the borders across different classes. To counteract this deficiency, we propose a unified multi-direction network (MDN) for HSI Classification (HSIC), which can exhaustively explore the abundant spectral and detailed spatial-context information through multi-direction samples. Additionally, considering the image-spectrum merged structure of the HSI, 3-D Squeeze-and-Excitation residual (3DSERes) blocks are devised in each stream of the framework to consecutively learn the spectral and spatial from low-level to high-level features. Specifically, 3DSERes can not only facilitate fluent gradient in backpropagation through skip connections, but also emphasize the significant spectral–spatial features and constrain the futile ones. This characteristic is beneficial to enhance the model’s generalization capability even with limited training samples. Furthermore, for properly aggregating the multi-direction deep features, we exploit the simple, yet effective attentional spectral prior (ASP) creatively through leveraging the original spectral correlations. Extensive experimental results on three benchmark data sets indicate that the proposed MDN-ASP can achieve promising classification performance compared to the state-of-the-art methods. Bobo Xi, Jiaojiao Li 0001, Yunsong Li 0001, Rui Song 0003, Yuchao Xiao, Yanzi Shi, Qian Du 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Target Detection With Unconstrained Linear Mixture Model and Hierarchical Denoising Autoencoder in Hyperspectral ImageryabstractHyperspectral imagery with very high spectral resolution provides a new insight for subtle nuances identification of similar substances. However, hyperspectral target detection faces significant challenges of intraclass dissimilarity and interclass similarity due to the unavoidable interference caused by atmosphere, illumination, and sensor noise. In order to effectively alleviate these spectral inconsistencies, this paper proposes a novel target detection method without strict assumptions on data distribution based on an unconstrained linear mixture model and deep learning. Our proposed detector firstly reduces interference via a specifically designed deep-learning-based hierarchical denoising autoencoder, and then carries out accurate detection with a two-step subspace projection, aiming at background suppression and target enhancement. Additionally, to generate representative background and reliable target samples required in the detection procedure, an efficient spatial-spectral unified endmember extraction method has been developed. Performance comparison with several state-of-the-art detection methods and further analysis on four real-world hyperspectral images demonstrate the effectiveness and efficiency of our proposed target detector. Yunsong Li 0001, Yanzi Shi, Bobo Xi, Jiaojiao Li 0001, Paolo Gamba |
IEEE Trans. Image Process. | 4 |
| 2022 | Few-Shot Learning With Class-Covariance Metric for Hyperspectral Image ClassificationabstractRecently, embedding and metric-based few-shot learning (FSL) has been introduced into hyperspectral image classification (HSIC) and achieved impressive progress. To further enhance the performance with few labeled samples, we in this paper propose a novel FSL framework for HSIC with a class-covariance metric (CMFSL). Overall, the CMFSL learns global class representations for each training episode by interactively using training samples from the base and novel classes, and a synthesis strategy is employed on the novel classes to avoid overfitting. During the meta-training and meta-testing, the class labels are determined directly using the Mahalanobis distance measurement rather than an extra classifier. Benefiting from the task-adapted class-covariance estimations, the CMFSL can construct more flexible decision boundaries than the commonly used Euclidean metric. Additionally, a lightweight cross-scale convolutional network (LXConvNet) consisting of 3D and 2D convolutions is designed to thoroughly exploit the spectral-spatial information in the high-frequency and low-frequency scales with low computational complexity. Furthermore, we devise a spectral-prior-based refinement module (SPRM) in the initial stage of feature extraction, which cannot only force the network to emphasize the most informative bands while suppressing the useless ones, but also alleviate the effects of the domain shift between the base and novel categories to learn a collaborative embedding mapping. Extensive experiment results on four benchmark data sets demonstrate that the proposed CMFSL can outperform the state-of-the-art methods with few-shot annotated samples. Bobo Xi, Jiaojiao Li 0001, Yunsong Li 0001, Rui Song 0003, Danfeng Hong, Jocelyn Chanussot |
IEEE Trans. Image Process. | 1 |
| 2021 | Semi-Supervised Graph Prototypical Networks for Hyperspectral Image ClassificationabstractGraph convolutional network (GCN) is one of the most favorable semi-supervised approaches, which demonstrates encouraging performance for hyperspectral image classification (HSIC), especially under the condition of small sample sizes. In this paper, we propose a novel semi-supervised graph prototypical network (SSGPN) for high-precise HSIC. Different from prevenient GCN, we devise a prototypical layer comprising a distance-based cross-entropy (DCE) loss function and a novel temporal entropy-based regularizer (TER) in the frameworks of SSGPN. This effective layer can facilitate to generate more discriminative embedding features along with the representative prototypes to each class, so as to achieve accurate identification of various land-cover categories. Additionally, to promote computational efficiency, we present a graph normalization (G-Norm) to accelerate the convergence speed and boost the training procedure. Experimental results demonstrate that our proposed SSGPN can obtain promising performance compared with the state-of-the-art methods. Bobo Xi, Jiaojiao Li 0001, Yunsong Li 0001, Qian Du 0001 |
IGARSS | 1 |
| 2021 | Hyperspectral Target Detection With RoI Feature Transformation and Multiscale Spectral AttentionabstractTarget detection plays a core issue in hyperspectral remote sensing, but faces serious challenges of how to deal with the spatial and spectral redundancies and spectral variations. In this article, a novel network block is developed, called RFT-MSA block (abbreviated as RM), which includes the region-of-interest (RoI) feature transformation (RFT) and the multiscale-spectral-attention (MSA) module as to reduce the spatial and spectral redundancies simultaneously and provide strong discrimination. Furthermore, a deep spatial-spectral network (DSSN) is presented by stacking several RM and deconvolutional (DC) blocks for hyperspectral target detection in an unsupervised manner, and a feature loss term is investigated to simultaneously restrict the target to be sparse and minimize the energy of the background. The proposed algorithm mainly consists of three steps. First, an RoI map is detected using a classical detector (no statistic assumption is needed) with an edge-preserving filter. Then, the hyperspectral image (HSI) and the corresponding RoI map are considered as inputs to the DSSN for extracting the spatial and spectral feature of interest (SSFI). Finally, we apply the nearest neighbors (NNs) to the SSFI for detection-map refinement. The experimental results on one synthetic and three real HSIs demonstrate that the proposed algorithm outperforms other benchmark approaches in detection performance and robustness. In addition, further analysis also demonstrates the effectiveness of the proposed RM block. Yanzi Shi, Jiaojiao Li 0001, Yuxuan Zheng, Bobo Xi, Yunsong Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2021 | Multiscale Context-Aware Ensemble Deep KELM for Efficient Hyperspectral Image ClassificationabstractRecently, multiscale spatial features have been widely utilized to improve the hyperspectral image (HSI) classification performance. However, fixed-size neighborhood involving the contextual information probably leads to misclassifications, especially for the boundary pixels. Additionally, it has been demonstrated that deep neural network (DNN) is practical to extract representative features for the classification tasks. Nevertheless, under the condition of high dimensionality versus small sample sizes, DNN tends to be over-fitting and it is generally time-consuming due to the deep-level feature learning process. To alleviate the aforementioned issues, we propose a multiscale context-aware ensemble deep kernel extreme learning machine (MSC-EDKELM) for efficient HSI classification. First, the scene of the HSI data set is over-segmented in multiscale via using the adaptive superpixel segmentation technique. Second, superpixel pattern (SP) and attentional neighboring superpixel pattern (ANSP) are generated by leveraging the superpixel maps, which can automatically comprise local and global contextual information, respectively. Afterward, an ensemble deep kernel extreme learning machine (EDKELM) is presented to investigate the deep-level characteristics in the SP and ANSP. Finally, the category of each pixel is accurately determined by the decision fusion and weighted output layer fusion strategy. Experimental results on four real-world HSI data sets demonstrate that the proposed frameworks outperform some classic and state-of-the-art methods with high computational efficiency, which can be employed to serve real-time applications. Bobo Xi, Jiaojiao Li 0001, Yunsong Li 0001, Rui Song 0003, Weiwei Sun 0005, Qian Du 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2018 | Classification of Hyperspectral Imagery Using a New Fully Convolutional Neural NetworkabstractWith success of convolutional neural networks (CNNs) in computer vision, the CNN has attracted great attention in hyperspectral classification. Many deep learning-based algorithms have been focused on deep feature extraction for classification improvement. In this letter, a novel deep learning framework for hyperspectral classification based on a fully CNN is proposed. Through convolution, deconvolution, and pooling layers, the deep features of hyperspectral data are enhanced. After feature enhancement, the optimized extreme learning machine (ELM) is utilized for classification. The proposed framework outperforms the existing CNN and other traditional classification algorithms by including deconvolution layers and an optimized ELM. Experimental results demonstrate that it can achieve outstanding hyperspectral classification performance. Jiaojiao Li 0001, Yunsong Li 0001, Qian Du 0001, Bobo Xi, Jing Hu 0005 |
IEEE Geosci. Remote. Sens. Lett. | 5 |