EDBT 2026 Demo / reviewers in the wild / expert
Bing Liu 0018
dblp:l/BingLiu18
· DBLP profile ↗
23ranked-venue papers
5as first author
21since 2021 · last 2026
0000-0003-0848-8453ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 17 · 4 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards structured knowledge extraction from lecture videos via multi-model collaboration
Mingqi Zheng, Bing Liu 0018, Yiwen Ye, Sinian Lin |
Expert Syst. Appl. | 3 |
| 2025 | Semantic Segmentation of Remote Sensing Images With Deep Information EnhancementabstractUsing semantic segmentation networks to intelligently categorize remote sensing images (RSIs) is essential for urban planning, land use, and environmental monitoring. However, the complexity of the foreground and background in RSIs, along with the multiscale characteristics of segmented objects, poses a significant challenge for the accurate segmentation of multiclass objects. The multi-source data fusion strategy can improve the segmentation accuracy of RSIs by incorporating complementary information. Inspired by this approach and the robust generalization capabilities of foundation models, we propose a novel method that combines depth maps of foundation models reasoning to improve the segmentation accuracy of RSIs. Specifically, we first utilize Depth Anything (DAM) to extract depth information. Next,we employ two lightweight convolutional layers to fuse depth information at the feature level. Finally, we implement U-Net for end-to-end training and prediction. We conducted numerous semantic segmentation experiments on the Vaihingen dataset. The experimental results demonstrate that our method achieves 73.23% a mean cross-union ratio (mIoU) on the Vaihingen dataset, which is 2.54% higher than the baseline. This performance improvement validates the effectiveness of the proposed method. Bing Liu 0018, Anzhu Yu, Xuefeng Cao, Guozheng Si |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2025 | Multimodal self-supervised learning for remote sensing data land cover classification
Zhixiang Xue, Guopeng Yang, Xuchu Yu, Anzhu Yu, Yinggang Guo, Bing Liu 0018, Jianan Zhou 0005 |
Pattern Recognit. | 6 |
| 2025 | Depth Feature Extraction for Hyperspectral Image Small Sample ClassificationabstractThe problem of insufficient labeled samples has restricted the application of deep learning method in hyperspectral image classification tasks. Fusion of remote sensing images from different sources such as hyperspectral image, Lidar is a common strategy to improve the classification accuracy. However, obtaining multi-source registered remote sensing images of the same area is time-consuming, which limits the application of multi-source strategy in practice. Motivated by the recent success of large models in different fields, we propose to extract depth information from large models and fuse it with hyperspectral images to improve the small sample classification accuracy. Specifically, we use the pre-trained foundation large model to estimate the depth information of hyperspectral images as the depth features, and then input the original spectral features and depth features into the support vector machine to complete the classification. In order to further improve the classification accuracy, we propose to use the sliding window method to extract the depth features of different bands, so as to obtain more rich depth features. A large number of classification experiments on six benchmark datasets verify the effectiveness of the proposed method. Bing Liu 0018, Zhixiang Xue, Pengqiang Zhang, Jiaying Yue |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2025 | Rethinking Semantic Segmentation With Multi-Grained Logical PrototypeabstractThe last decade has witnessed significant advances in semantic segmentation brought about by deep learning. However, existing methods only fit the data-label correspondence in a data-driven manner and do not fully conform to the abstraction and structuralization characteristics of the human visual cognition process, which limits the upper bounds of their performance. To this end, a multi-grained logical prototype (MGLP) method is proposed to rethink semantic segmentation based on these two key characteristics. Its novel design can be summarized as follows. 1) For abstraction, prototypes of the same class at different grain levels are established: a label generation method is proposed to automatically generate a multi-grained label space, which can guide the learning of the multi-grained prototypes for each class. 2) For structuralization, the intrinsic logical structure across different semantic levels is explicitly modeled: the horizontal metric relationships are established via metric relation operations on prototypes at the same grain level, to improve the discriminability between classes while taking the vertical semantic hierarchy into account. Moveover, the vertical logical relationships are established as the sub-to-super positive and super-to-sub negative constraints, to strengthen the semantic dependencies among prototypes at different grain levels. 3)MGLP is plug-and-play and can be directly combined with existing segmentation methods. Extensive experimental results indicate that MGLP can significantly improve the segmentation performance of existing methods, which opens up a new avenue for future research. Anzhu Yu, Kuiliang Gao, Xiong You, Yanfei Zhong, Bing Liu 0018, Chunping Qiu |
IEEE Trans. Image Process. | 6 |
| 2024 | Joint Spatio-Temporal Modeling for Semantic Change Detection in Remote Sensing ImagesabstractSemantic Change Detection (SCD) refers to the task of simultaneously extracting the changed areas and the semantic categories (before and after the changes) in Remote Sensing Images (RSIs). This is more meaningful than Binary Change Detection (BCD) since it enables detailed change analysis in the observed areas. Previous works established triple-branch Convolutional Neural Network (CNN) architectures as the paradigm for SCD. However, it remains challenging to exploit semantic information with a limited amount of change samples. In this work, we investigate to jointly consider the spatio-temporal dependencies to improve the accuracy of SCD. First, we propose a Semantic Change Transformer (SCanFormer) to explicitly model the ’from-to’ semantic transitions between the bi-temporal RSIs. Then, we introduce a semantic learning scheme to leverage the spatio-temporal constraints, which are coherent to the SCD task, to guide the learning of semantic changes. The resulting network (SCanNet) significantly outperforms the baseline method in terms of both detection of critical semantic changes and semantic consistency in the obtained bi-temporal results. It achieves the SOTA accuracy on two benchmark datasets for the SCD. Lei Ding 0008, Jing Zhang 0023, Haitao Guo, Kai Zhang 0010, Bing Liu 0018, Lorenzo Bruzzone |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | Exploiting Discriminative Advantage of Spectrum for Hyperspectral Image Classification: SpectralFormer Enhanced by Spectrum Motion FeatureabstractAs for hyperspectral images (HSIs), the discrepancy of contiguous spectral information should be the main basis for the identification of ground objects. Due to the difficulty of spectral sequence coding and the spectrum similarity between categories, successful deep-learning-based classification methods always attempt to capture the spatial information to improve the accuracy by convolutional neural networks (CNNs) or other excellent spatial feature extractors. However, extracting spatial features is generally accompanied by the distortion of ground objects distribution and categories boundary. To effectively represent spectral features, the SpectralFormer based on transformer backbone can better capture the long-term dependence of the spectrum, which improves the performance of spectral feature methods significantly. However, it is still unable to compete with advanced spectral–spatial feature methods. To exploit the discriminative advantage of the spectrum fully, this letter introduces an efficient sparse-to-dense optical flow estimation method to track the spectrum variation in the HSI. Then, we take such a variation as a spectrum motion feature to enhance the original spectrum. At last, we continue to use the SpectralFormer to encode the concatenated spectrum sequence for classification. Extensive experiments show that the SpectralFormer enhanced by the spectrum motion feature (SF-SMF) significantly improves the performance of spectral feature methods, even surpassing advanced spectral–spatial feature methods. SF-SMF can avoid interference with additional spatial information to obtain exquisite whole-domain classification maps, showing its practical value. The codes will be public athttps://github.com/sssssyf/SF-SMF. Yifan Sun 0008, Bing Liu 0018, Xuchu Yu, Anzhu Yu, Pengqiang Zhang, Zhixiang Xue |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2023 | Prototype and Context-Enhanced Learning for Unsupervised Domain Adaptation Semantic Segmentation of Remote Sensing ImagesabstractIn unsupervised domain adaptation (UDA) of remote sensing images (RSIs), the huge inter-domain discrepancies and intra-domain variances lead to complicated class-level relations. Specifically, the instances of the same class differ greatly while instances of different classes are similar, whether across different RSIs domains or within the same RSIs domain. However, existing methods cannot fully consider these problems, limiting the performance of UDA semantic segmentation of RSIs. To this end, this paper proposes a novel cross-domain multi-prototypes learning method, the core idea of which is to abstract the cross-and intra-domain class-level relations into multiple prototypes. Specifically, the multiple prototypes belonging to different classes can detailedly describe complex inter-class relations, and the multiple prototypes within the same class can better model rich intra-class relations. Further, the source and target samples are jointly used for prototypes calculation, to fully fuse the feature information of different RSIs. In a nutshell, utilizing the samples from different RSIs domains to learn multiple prototypes for each class can achieve better domain alignment at the class level. In addition, considering that RSIs simultaneously contain large targets with wide coverage and important small targets, two masked consistency learning strategies are designed to better explore the contextual structure of target RSIs and improve the quality of pseudo labels for prototype updating. The global consistency strategy can strengthen the utilization of global context relations, while the local consistency strategy can further improve the learning of local context details. Therefore, the proposed method is actually a prototype and context enhanced learning method for UDA semantic segmentation of RSIs. Extensive experiments demonstrate that the proposed method can achieve better performance than existing state-of-the-art UDA methods. Kuiliang Gao, Anzhu Yu, Xiong You, Chunping Qiu, Bing Liu 0018 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | Spectral-Spatial MLP-Like Network With Reciprocal Points Learning for Open-Set Hyperspectral Image ClassificationabstractIn recent years, deep-learning-based hyperspectral image (HSI) classification methods have achieved significant development and gradually become widely applied. The existing advanced methods can achieve near-saturation performance with sufficient labels in a closed-set environment (CSE), i.e., training set and test set are all known categories of ground objects. However, the real world is usually open because of the diversity of land covers, i.e., test-set exists unknown categories that are not labeled in the training set. Therefore, the prevalent advanced CSE methods still cannot effectively and robustly handle unknown categories of ground objects in an open-set environment (OSE). Therefore, we propose a spectral-spatial MLP-like network with reciprocal points learning (SSMLP-RPL) to improve the performance of open-set HSI classification. First, a feature learning framework based on reciprocal points learning (RPL) is constructed to model the extra-category space and reduce the risk of open space. The learned feature space enables to enlarge the distance between the known and unknown categories. Besides, we further propose to utilize a learnable dynamic threshold of each known category to effectively distinguish the unknown categories and improve open performance of the model. Second, to enhance the capacity of feature learning, a spectral-spatial MLP-like network (SSMLP) is designed to capture the spectral-spatial feature merely with a series of fully-connected (FC) layers, which mainly involve SpeFC and SpaFC two modules. Among them, the SpaFC module enables to model spacial semantics, and the SpeFC module enables to model long-distance spectral dependence. Extensive experiments on three benchmark HSIs show that SSMLP-RPL has a competitive performance both in CSE and OSE and even surpasses currently advanced closed-set and open-set HSI classification methods. As an end-to-end HSI classification framework of MLP-backbone, SSMLP network can compete with the advanced works based on CNN and transformer. The code will be open at: https://github.com/sssssyf/SSMLP-RPL. Yifan Sun 0008, Bing Liu 0018, Ruirui Wang, Pengqiang Zhang, Mofan Dai |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Hyperspectral Meets Optical Flow: Spectral Flow Extraction for Hyperspectral Image ClassificationabstractHyperspectral image (HSI) classification has always been recognised as a difficult task. It is therefore a research hotspot in remote sensing image processing and analysis, and a number of studies have been conducted to better extract spectral and spatial features. This study aimed to track the variation of the spectrum in hyperspectral images from a sequential data perspective to obtain more distinguishable features. Based on the characteristics of optical flow, this study introduces an optical flow technique for the extraction of spectral flow that denotes the spectral variation and implements a dense optical flow extraction method based on deep matching. Lastly, the extracted spectral flow are combined with the original spectral features and input into a commonly used support vector machine (SVM) classifier to complete the classification. Extensive classification experiments on three benchmark HSI test sets show that the classification accuracy obtained by the spectral flow extracted in this study (SpectralFlow) is higher than traditional spatial feature extraction methods, texture feature extraction methods, and the latest deep-learning-based methods. Furthermore, the proposed method can produce finer classification thematic maps, thereby demonstrating strong practical application potential. Bing Liu 0018, Yifan Sun 0008, Anzhu Yu, Zhixiang Xue, Xibing Zuo |
IEEE Trans. Image Process. | 1 |
| 2022 | Self-Supervised Feature Learning and Few-Shot Land Cover Classification for Cross-Modal Remote Sensing ImagesabstractWith the rapid development of remote sensing data acquisition technology, there are multimodal images over the same observed scenes. These multimodal remote sensing images could provide complementary valuable information for land cover classification. In this article, we propose a novel self-supervised feature learning and few-shot classification model for multimodal remote sensing images, called S2FL. Specifically, a contrastive learning architecture is investigated to learn spatial feature representations from very high resolution (VHR) image. And the spectral features from hyperspectral data are integrated with learned spatial features for few-shot land cover classification. Classification experiments are conducted on a widely-used dataset, i.e., Houston 2018, to verify the effectiveness and superiority of the proposed S2FL model compared with several state-of-the-art baseline approaches. Zhixiang Xue, Xuchu Yu, Pengqiang Zhang, Xiong Tan, Anzhu Yu, Bing Liu 0018 |
IGARSS | 6 |
| 2022 | MP-ResNet: Multipath Residual Network for the Semantic Segmentation of High-Resolution PolSAR ImagesabstractThere are limited studies on the semantic segmentation of high-resolution polarimetric synthetic aperture radar (PolSAR) images due to the scarcity of training data and the complexity of managing speckle noise. The Gaofen contest has provided open access a high-quality PolSAR semantic segmentation dataset. Taking this opportunity, we propose a multipath residual network (MP-ResNet) architecture for the semantic segmentation of high-resolution PolSAR images. Compared to conventional U-shape encoder–decoder convolutional neural network (CNN) architectures, the MP-ResNet learns semantic context with its parallel multiscale branches, which greatly enlarges its valid receptive fields and improves the embedding of local discriminative features. In addition, MP-ResNet adopts a multilevel feature fusion design in its decoder to effectively exploit the features learned from its different branches. Comparisons with the baseline method of fully connected network (FCN with ResNet34) show that the MP-ResNet has achieved significant accuracy improvements. It also surpasses several state-of-the-art methods in terms of overall accuracy (OA),$\text{m}F_{1}$and frequency weighted intersection over union (fwIoU), with only a limited increase of computational costs. This CNN architecture can be used as a baseline method for future studies on the semantic segmentation of PolSAR images. The code is available at:https://github.com/ggsDing/SARSeg. Lei Ding 0008, Dong Lin, Yuxing Chen 0002, Bing Liu 0018, Jiansheng Li, Lorenzo Bruzzone |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2022 | Pixel-Level Self-Supervised Learning for Semi-Supervised Building Extraction From Remote Sensing ImagesabstractThe building extraction from remote sensed images ash been a challenging yet vital task for applicable purposes such as urban monitoring and cartography. Most of the existing learning based approaches focus on the supervised building extraction methods, of which the models should be trained with images and the corresponding labels. This research exploits a self-supervised approach for building extraction, which could train the backbone within a building extraction network without annotations. Specifically, the backbone is initially trained with a pixel-level self-supervised module instead of commonly used supervised approaches or instance-level self-supervised modules. Next, the pretrained backbone is embedded into a task-specific network followed by tuning with limited annotations. The experiments were conducted on three popular datasets and the results show that our method achieves improvements regarding both intersection over union (IoU) and F1-score compared to supervised approach and instance-level self-supervised methods. Our study thus confirms the potential of pixel-level self-supervised approach for semantic segmentation for remote sensing images. Anzhu Yu, Bing Liu 0018, Xuefeng Cao, Chunping Qiu, Wenyue Guo, Yujun Quan |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2022 | Perceiving Spectral Variation: Unsupervised Spectrum Motion Feature Learning for Hyperspectral Image ClassificationabstractIn recent years, deep-learning-based hyperspectral image (HSI) classification methods have achieved significant development. The superior capability of feature extraction from these data-driven methods dramatically improves the classification performance. However, the previous methods usually require to retrain the network from scratch to obtain the capability of feature extraction adaptive for the target image when facing a new HSI to be classified, which is a time-consuming and redundant process. In this paper, we consider putting this process ahead and making the network have a robust capability of feature extraction with generalization through pre-training. Therefore, the network enables to directly extract features of the target HSI without re-training. For this purpose, we rethink the three-dimension (3D) HSI data from a perspective of spectral sequence, and we attempt to extract the spectral variation information as the spectrum motion feature. Then, we construct an unsupervised spectrum motion feature learning framework (SMF-UL), which can be pre-trained on mass unlabeled HSI data to learn the knowledge about perceiving spectral variation. Furthermore, to achieve the expansion of source data for pre-training, we develop an extendable training dataset construction method, which can integrate HSIs of different sizes, number of bands and sensors into a unified training set to utilize the rapidly growing mass unlabeled HSI data effectively. Finally, we use the trained network to directly extract the spectrum motion feature of the target HSI for classification, so the laborious re-training of the network can be avoided. Extensive experiments show that the proposed SMF-UL acquires the robust capability of feature extraction with generalization through unsupervised learning on mass unlabeled HSI data, and the classification performance of extracted spectrum motion feature is competitive to advanced in-domain and cross-domain methods, which shows its flexibility and superiority. The code of SMF-UL will be open at: https://github.com/sssssyf/SMF-UL. Yifan Sun 0008, Bing Liu 0018, Xuchu Yu, Anzhu Yu, Kuiliang Gao, Lei Ding 0008 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Self-Supervised Feature Representation and Few-Shot Land Cover Classification of Multimodal Remote Sensing ImagesabstractAlthough deep learning-based approaches have made significant progress in remote sensing image classification, the supervised learning paradigm has shortcomings under a limited number of labeled samples, which restricts the classification performance to a great extent. In this article, we investigate an effective self-supervised feature representation architecture (SSFR) for multimodal remote sensing images few-shot land cover classification. Specifically, we exploit multiview learning strategy to construct multiple views from multimodal remote sensing images. This method builds several complementary views of the same observed scenes from hyperspectral images or different modalities of remote sensing data. Then we build the deep feature extractor to learn high-level feature representations from each view via contrastive learning. The contrastive learning aggregates the samples of the same scene while separating samples of different scenes in the latent space, and this process does not require any labeled information. What’s more, to learn more robust features from different views, we utilize multitask learning strategy to train the feature extraction network. Finally, a lightweight machine learning method is employed to classify the learned features using a few annotated samples. To further demonstrate the self-supervised feature learning capability of the proposed model, we train the feature representation network in multiple source datasets. Comprehensive feature learning and classification experiments have certified the effectiveness and superiority of the proposed method. Zhixiang Xue, Bing Liu 0018, Anzhu Yu, Xuchu Yu, Pengqiang Zhang, Xiong Tan |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Multiscale Deep Learning Network With Self-Calibrated Convolution for Hyperspectral and LiDAR Data Collaborative ClassificationabstractIn this article, we propose a novel multiscale deep learning network with self-calibrated convolution (MSNetSC) for hyperspectral and light detection and ranging (LiDAR) data collaborative classification. Conventional deep learning methods have limitations in extracting multiscale features at a granular level from multimodality data and fusing these features in a context-awareness way, which will severely restrict the performance of hyperspectral and LiDAR data joint classification. The proposed multiscale deep learning network utilizes a hierarchical residual structure combined with self-calibrated convolution to extract features with different receptive fields, and this can enhance the model’s capability to represent the multimodality data. Besides, we employ spectral and spatial self-attention modules to adaptively calibrate weights of features with different scales, thereby enhancing the discriminative ability of extracted multiscale features. Furthermore, the attentional feature fusion module can dynamically and adaptively fuse the features from multimodality data in a contextual scale-aware way, and this attention-based feature fusion method will further improve the collaborative classification performance of hyperspectral and LiDAR data. Four benchmark multimodality data (i.e., hyperspectral and LiDAR data) sets collected by different sensors and at different acquisition times are employed for joint classification experiments. These comparative classification results and ablation study sufficiently certify the superiority of the proposed model in terms of collaborative classification accuracy when compared with other state-of-the-art methods. Zhixiang Xue, Xuchu Yu, Xiong Tan, Bing Liu 0018, Anzhu Yu, Xiangpo Wei |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Self-Supervised Feature Learning for Multimodal Remote Sensing Image Land Cover ClassificationabstractDeep learning models have shown great potential in remote sensing image processing and analysis. Nevertheless, there are insufficient labeled samples to train deep networks, which seriously affects the performance of these models. To resolve this contradiction, we propose a generative self-supervised feature learning (S2FL) architecture for multimodal remote sensing image land cover classification. Specifically, multiple complementary observed views are constructed from multimodal remote sensing images, which are employed for following generative self-supervised learning. The proposed S2FL architecture is capable of extracting high-level meaningful feature representations from multiview data, and this process does not require any labeled information, providing a feasible solution to relieve the urgent need for annotated samples. The learned features are normalized and merged with corresponding spectral information to further improve the discriminative capability of feature representations, and we utilize these fused features for land cover classification. Compared with existing supervised, semi-supervised, and self-supervised approaches, the proposed generative self-supervised model achieves superior performance in terms of feature learning and land cover classification, especially in the small sample classification case. Zhixiang Xue, Xuchu Yu, Anzhu Yu, Bing Liu 0018, Pengqiang Zhang, Shentong Wu |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | FSL-EGNN: Edge-Labeling Graph Neural Network for Hyperspectral Image Few-Shot ClassificationabstractThe existing hyperspectral image (HSI) classification encounters the obstacle of improving the classification accuracy with limited labeled samples. In this context, as a typical implementation of meta-learning, few-shot learning (FSL) makes the model learn by episodic training on source HSI, which has achieved significant improvements in small sample classification of target HSI. However, the existing FSL methods lack explicit consideration and exploration of the association between pixels, especially the intraclass association and interclass association between pixels in the support set and query set. To mitigate these issues, an FSL method based on edge-labeling graph neural network (FSL-EGNN) is proposed for small sample classification of HSI, which is the first attempt to explicitly quantify the associations between pixels by exploiting EGNN in HSI few-shot classification (FSC). Specifically, based on graph construction of HSI, episodic training is performed on the existing source HSI. During training, EGNN is used to predict the edge labels on the graph, thereby explicitly modeling the intraclass similarity and interclass dissimilarity between pixels of HSI. After the trained model is fine-tuned, it can realize FSC on the unseen target HSI. Experiments conducted on three benchmark HSI datasets demonstrate that the proposed FSL-EGNN outperforms the existing state-of-the-art methods with limited labeled samples. Xibing Zuo, Xuchu Yu, Bing Liu 0018, Pengqiang Zhang, Xiong Tan |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Unsupervised Meta Learning With Multiview Constraints for Hyperspectral Image Small Sample set ClassificationabstractThe difficulties of obtaining sufficient labeled samples have always been one of the factors hindering deep learning models from obtaining high accuracy in hyperspectral image (HSI) classification. To reduce the dependence of deep learning models on training samples, meta learning methods have been introduced, effectively improving the classification accuracy in small sample set scenarios. However, the existing methods based on meta learning still need to construct a labeled source data set with several pre-collected HSIs, and must utilize a large number of labeled samples for meta-training, which is actually time-consuming and labor-intensive. To solve this problem, this paper proposes a novel unsupervised meta learning method with multiview constraints for HSI small sample set classification. Specifically, the proposed method first builds an unlabeled source data set using unlabeled HSIs. Then, multiple spatial-spectral multiview features of each unlabeled sample are generated to construct tasks for unsupervised meta learning. Finally, the designed residual relation network is used for meta-training and small sample set classification based on the voting strategy. Compared with existing supervised meta learning methods for HSI classification, our method can only utilize HSIs without any label for unsupervised meta learning, which significantly reduces the number of requisite labeled samples in the whole classification process. To verify the effectiveness of the proposed method, extensive experiments are carried out on 8 public HSIs in the cross-domain and in-domain classification scenarios. The statistical results demonstrate that, compared with existing supervised meta learning methods and other advanced classification models, the proposed method can achieve competitive or better classification performance in small sample set scenarios. Kuiliang Gao, Bing Liu 0018, Xuchu Yu, Anzhu Yu |
IEEE Trans. Image Process. | 2 |
| 2022 | Deep Hierarchical Vision Transformer for Hyperspectral and LiDAR Data ClassificationabstractIn this study, we develop a novel deep hierarchical vision transformer (DHViT) architecture for hyperspectral and light detection and ranging (LiDAR) data joint classification. Current classification methods have limitations in heterogeneous feature representation and information fusion of multi-modality remote sensing data (e.g., hyperspectral and LiDAR data), these shortcomings restrict the collaborative classification accuracy of remote sensing data. The proposed deep hierarchical vision transformer architecture utilizes both the powerful modeling capability of long-range dependencies and strong generalization ability across different domains of the transformer network, which is based exclusively on the self-attention mechanism. Specifically, the spectral sequence transformer is exploited to handle the long-range dependencies along the spectral dimension from hyperspectral images, because all diagnostic spectral bands contribute to the land cover classification. Thereafter, we utilize the spatial hierarchical transformer structure to extract hierarchical spatial features from hyperspectral and LiDAR data, which are also crucial for classification. Furthermore, the cross attention (CA) feature fusion pattern could adaptively and dynamically fuse heterogeneous features from multi-modality data, and this contextual aware fusion mode further improves the collaborative classification performance. Comparative experiments and ablation studies are conducted on three benchmark hyperspectral and LiDAR datasets, and the DHViT model could yield an average overall classification accuracy of 99.58%, 99.55%, and 96.40% on three datasets, respectively, which sufficiently certify the effectiveness and superior performance of the proposed method. Zhixiang Xue, Xiong Tan, Xuchu Yu, Bing Liu 0018, Anzhu Yu, Pengqiang Zhang |
IEEE Trans. Image Process. | 4 |
| 2021 | Deep Multiview Learning for Hyperspectral Image ClassificationabstractRecently, the field of hyperspectral image (HSI) classification is dominated by deep learning-based methods. However, training deep learning models usually needs a large number of labeled samples to optimize thousands of parameters. In this article, a deep multiview learning method is proposed to deal with the small sample problem of HSI. First, two views of an HSI scene are constructed by applying principal component analysis to different bands. Second, a deep residual network is designed to embed the different views of a sample to a latent space. The designed deep residual network is trained by maximizing agreement between differently augmented views of the same data sample via a contrastive loss in the latent space. Note that the training procedure of the designed deep residual network does not use labeled information. Therefore, the proposed method belongs to the category of unsupervised learning, which could alleviate the lack of labeled training samples. Finally, a conventional machine learning method (e.g., support vector machine) is used to complete the classification task in the learned latent space. To demonstrate the effectiveness of the proposed method, extensive experiments are carried on four widely used hyperspectral data sets. The experimental results demonstrate that the proposed method could improve the classification accuracy with small samples. Bing Liu 0018, Anzhu Yu, Xuchu Yu, Ruirui Wang, Kuiliang Gao, Wenyue Guo |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2019 | Deep Few-Shot Learning for Hyperspectral Image ClassificationabstractDeep learning methods have recently been successfully explored for hyperspectral image (HSI) classification. However, training a deep-learning classifier notoriously requires hundreds or thousands of labeled samples. In this paper, a deep few-shot learning method is proposed to address the small sample size problem of HSI classification. There are three novel strategies in the proposed algorithm. First, spectral–spatial features are extracted to reduce the labeling uncertainty via a deep residual 3-D convolutional neural network. Second, the network is trained by episodes to learn a metric space where samples from the same class are close and those from different classes are far. Finally, the testing samples are classified by a nearest neighbor classifier in the learned metric space. The key idea is that the designed network learns a metric space from the training data set. Furthermore, such metric space could generalize to the classes of the testing data set. Note that the classes of the testing data set are not seen in the training data set. Four widely used HSI data sets were used to assess the performance of the proposed algorithm. The experimental results indicate that the proposed method can achieve better classification accuracy than the conventional semisupervised methods with only a few labeled samples. Bing Liu 0018, Xuchu Yu, Anzhu Yu, Pengqiang Zhang, Ruirui Wang |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2018 | Supervised Deep Feature Extraction for Hyperspectral Image ClassificationabstractHyperspectral image classification has become a research focus in recent literature. However, well-designed features are still open issues that impact on the performance of classifiers. In this paper, a novel supervised deep feature extraction method based on siamese convolutional neural network (S-CNN) is proposed to improve the performance of hyperspectral image classification. First, a CNN with five layers is designed to directly extract deep features from hyperspectral cube, where the CNN can be intended as a nonlinear transformation function. Then, the siamese network composed by two CNNs is trained to learn features that show a low intraclass and high interclass variability. The important characteristic of the presented approach is that the S-CNN is supervised with a margin ranking loss function, which can extract more discriminative features for classification tasks. To demonstrate the effectiveness of the proposed feature extraction method, the features extracted from three widely used hyperspectral data sets are fed into a linear support vector machine (SVM) classifier. The experimental results demonstrate that the proposed feature extraction method in conjunction with a linear SVM classifier can obtain better classification performance than that of the conventional methods. Bing Liu 0018, Xuchu Yu, Pengqiang Zhang, Anzhu Yu, Qiongying Fu, Xiangpo Wei |
IEEE Trans. Geosci. Remote. Sens. | 1 |