EDBT 2026 Demo / reviewers in the wild / expert
Anzhu Yu
dblp:197/8015
· DBLP profile ↗
30ranked-venue papers
3as first author
28since 2021 · last 2026
0000-0002-3332-9668ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 20 · 2 first-author · 18 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Weakly supervised urban change monitoring via generative training and the knowledge of typical geographic features fusion
Beibei Wu, Qing Xu 0005, Longhao Wang, Xin Chen 0088, Anzhu Yu |
Expert Syst. Appl. | 6 |
| 2026 | P3D: Plug-and-play prompt-driven framework for RGB-thermal semantic segmentationabstract• A plug-and-play prompt-driven framework for RGB-thermal image semantic segmentation. • LoRA-based fine-tuning strategy for SAM series model integration. • A model-agnostic encoder to generate statistical distributed prompts for training. The semantic segmentation of RGB-thermal images is critical for applications with low-light conditions. Existing works primarily focus on feature fusion strategies and model design to enhance performance. While Visual Foundation Models (VFMs) have been introduced in previous studies to improve generalization and segmentation accuracy, they suffer from poor compatibility with other models thus requiring full model retraining. Additionally, the domain gap and modality gap between VFM pre-training datasets and RGB-thermal semantic segmentation datasets pose significant challenges to VFM adaptation for downstream tasks. To address these issues, in this paper a plug-and-play prompt driven framework P 3 D is proposed. Unlike existing VFM-based methods that require complete retraining for each specific architecture, P 3 D is designed with a model-agnostic training strategy that enables one-time training and seamless integration with various existing methods without requiring retraining. First, a dual-branch LoRA (Low-Rank Adaptation) fine-tuned (DBLF) image encoder for the RGB and thermal image branches is proposed to narrow the domain gap and modality gap when incorporating SAM series models into our task. Second, a unified prompt generation and representation (UPGR) encoder is proposed. It generates diverse prompts using semantic labels during the training stage, ensuring the generated prompts are model-agnostic and compatible with existing methods. Finally, a cross-modality spatial-channel attention (CM-SCA) decoder is developed to fuse the embeddings from two-modality images and prompts for the final prediction. Extensive experiments are conducted on three popular benchmarks. Results demonstrate that P 3 D not only improves the performance of existing models but also outperforms current state-of-the-art (SOTA) methods leveraging < 1% trainable parameters. More importantly, by simply plugging P 3 D into existing methods, we consistently achieve significant performance improvements without retraining these base models, demonstrating the practical value of our plug-and-play design. Yongqi Sun, Chenguang Dai, Hanyun Wang, Longguang Wang, Wenke Li, Anzhu Yu |
Pattern Recognit. | 8 |
| 2026 | Survey of automated 3D reconstruction from optical satellite imagery
Anzhu Yu, Danyang Hong, Chunping Qiu, Junyi Fan, Song Ji |
Pattern Recognit. | 2 |
| 2025 | Semantic Segmentation of Remote Sensing Images With Deep Information EnhancementabstractUsing semantic segmentation networks to intelligently categorize remote sensing images (RSIs) is essential for urban planning, land use, and environmental monitoring. However, the complexity of the foreground and background in RSIs, along with the multiscale characteristics of segmented objects, poses a significant challenge for the accurate segmentation of multiclass objects. The multi-source data fusion strategy can improve the segmentation accuracy of RSIs by incorporating complementary information. Inspired by this approach and the robust generalization capabilities of foundation models, we propose a novel method that combines depth maps of foundation models reasoning to improve the segmentation accuracy of RSIs. Specifically, we first utilize Depth Anything (DAM) to extract depth information. Next,we employ two lightweight convolutional layers to fuse depth information at the feature level. Finally, we implement U-Net for end-to-end training and prediction. We conducted numerous semantic segmentation experiments on the Vaihingen dataset. The experimental results demonstrate that our method achieves 73.23% a mean cross-union ratio (mIoU) on the Vaihingen dataset, which is 2.54% higher than the baseline. This performance improvement validates the effectiveness of the proposed method. Bing Liu 0018, Anzhu Yu, Xuefeng Cao, Guozheng Si |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2025 | Learning-Based Multiview Stereo for Remote Sensed Imagery With Relative DepthabstractThe depth map estimation from a set of remote sensed images has been a challenging task due to the complexity of the real-world scenes, yet it is of great importance for applications such as 3D reconstruction and digital surface model (DSM). Most of the existing learning based multi-view stereo (MVS) approaches reconstruct depth map through supervised models that are trained with large amounts of data without any prior knowledge about the areas of interest. A recent foundational model Depth Anything, provides a robust method to estimate the relative depth map (RDM) based on single optical image. Leveraging its excellent performance in depth estimation and fine generalization capability, we propose a RDM fusion module that can be integrated with most state-of-the-art (SOTA) learning based MVS frameworks to improve their performance in 3D reconstruction. Extensive experiments are conducted to verify the effectiveness of the proposed module and positive results indicate that the integration of the proposed module leads to better accuracy and completeness compared to the benchmark models. The codes are available at https://github.com/2022hong/RDM-MVS/tree/main. Anzhu Yu, Danyang Hong, Xuanbei Lu, Song Ji, Junyi Fan |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2025 | Multimodal self-supervised learning for remote sensing data land cover classification
Zhixiang Xue, Guopeng Yang, Xuchu Yu, Anzhu Yu, Yinggang Guo, Bing Liu 0018, Jianan Zhou 0005 |
Pattern Recognit. | 4 |
| 2025 | A Frequency Decoupling Network for Semantic Segmentation of Remote Sensing ImagesabstractSemantic segmentation of remote sensing images (RSIs) is vital for numerous geospatial applications, including land-use mapping, urban planning, and environmental monitoring. Traditional neural networks for semantic segmentation primarily focus on learning in the spatial domain, which often results in suboptimal performance due to the complexity of RSIs that exhibit diverse and intricate structures. To address this problem, we propose a novel frequency decoupling network (FDNet) that enhances feature representation by independently refining high-frequency and low-frequency components in the frequency domain. FDNet introduces three core components: a sparse-aware spectral enhancement module (SSEM) that optimizes spectral feature learning by compressing redundant information while highlighting informative spectral bands, a frequency decoupling attention module (FDAM) that precisely distinguishes and enhances high-frequency and low-frequency features and an attentive frequency context module (AFCM) that integrates SSEM and FDAM into a cohesive framework for enriched spectral context modeling. Extensive experiments conducted on four benchmark datasets demonstrate that FDNet outperforms several state-of-the-art methods, achieving superior segmentation accuracy and robustness across various terrains and imaging conditions. Ablation experiments further confirm the impacts of SSEM, FDAM, and AFCM. Xin Li 0090, Feng Xu 0008, Anzhu Yu, Xin Lyu 0001, Hongmin Gao 0001, Jun Zhou 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | Elevation Correction of a Large-Scale DEM Using ICESat-2 Laser Altimetry DataabstractThe digital elevation model (DEM) is a digital representation of the surface elevation, but it contains elevation errors arising from vegetation coverage and terrain undulation. Spaceborne LiDAR, with its large-scale and high-precision elevation measurement capabilities, can effectively correct these errors and has become an important means to improve the elevation accuracy of DEM. In this study, a hybrid incremental regression (HIR) model based on the Ice, Cloud, and land Elevation Satellite-2 (ICESat-2) laser altimetry data is designed to correct large-scale DEM elevation errors. The model combines a neural network and a decision tree model to efficiently fit the elevation error through stepwise regression, which can be applied to different terrains and landforms over large areas. In the experiment, first, the proposed model was compared with four classic models [multiple linear regression (MLR), random forest (RF), backpropagation neural network (BPNN) and light gradient-boosting machine (LightGBM)] in seven different landform areas, and the results showed that our model was better than other models in terms of accuracy improvement and was applicable to various terrains. Then, the model was applied to China, and the higher precision, large-scale DEM datasets were produced. Compared with the original DEM, the root mean squared error (RMSE) and mean absolute error (MAE) of the optimized DEM were reduced by 1.603 and 1.453 m, respectively. The differences in accuracy under different slopes, land cover types, and vegetation heights were also analyzed, and it was found that the enhancement effect was the best in the high-vegetation cover area. Finally, it is validated using three regions of high-precision validation data, and the results showed that the RMSE had improved to varying degrees, ranging from 0.193 to 2.853 m. Weiqi Lian, Ziqi Nie, Guo Zhang 0001, Ke Li 0005, Xuefeng Cao, Anzhu Yu, Xin Li 0103 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | Rethinking Semantic Segmentation With Multi-Grained Logical PrototypeabstractThe last decade has witnessed significant advances in semantic segmentation brought about by deep learning. However, existing methods only fit the data-label correspondence in a data-driven manner and do not fully conform to the abstraction and structuralization characteristics of the human visual cognition process, which limits the upper bounds of their performance. To this end, a multi-grained logical prototype (MGLP) method is proposed to rethink semantic segmentation based on these two key characteristics. Its novel design can be summarized as follows. 1) For abstraction, prototypes of the same class at different grain levels are established: a label generation method is proposed to automatically generate a multi-grained label space, which can guide the learning of the multi-grained prototypes for each class. 2) For structuralization, the intrinsic logical structure across different semantic levels is explicitly modeled: the horizontal metric relationships are established via metric relation operations on prototypes at the same grain level, to improve the discriminability between classes while taking the vertical semantic hierarchy into account. Moveover, the vertical logical relationships are established as the sub-to-super positive and super-to-sub negative constraints, to strengthen the semantic dependencies among prototypes at different grain levels. 3)MGLP is plug-and-play and can be directly combined with existing segmentation methods. Extensive experimental results indicate that MGLP can significantly improve the segmentation performance of existing methods, which opens up a new avenue for future research. Anzhu Yu, Kuiliang Gao, Xiong You, Yanfei Zhong, Bing Liu 0018, Chunping Qiu |
IEEE Trans. Image Process. | 1 |
| 2025 | Supervised Contrastive Learning for Indoor Point Cloud OversegmentationabstractPoint cloud oversegmentation method can obtain a series of superpoints by grouping points that are semantically and geometrically consistent. The generated superpoints can be treated as the basic processing units in various downstream tasks to improve task performance and processing efficiency. However, due to the high semantic and geometric complexity of point cloud scenes, obtaining high-quality superpoints is still challenging. Aiming to generate high-quality indoor superpoints, we propose an end-to-end supervised contrastive learning framework SCL-OverSeg for indoor point cloud oversegmentation. Firstly, to solve the challenge of balancing the importance of geometric similarity and spatial proximity constraint between points and superpoints in indoor scenes, we integrate the geometric similarity and spatial proximity constraint into the supervision signal by generating the superpoint ground truth. To solve the challenge of superpoints crossing objects, we propose to utilize instance labels rather than semantic labels to generate the ideal superpoint ground truth as the object-level supervision signal. Secondly, to construct the distinguishable embedding space facilitating to the assignments of points to superpoints, we propose point-superpoint contrastive learning to compel the network to project each point to be closer to the reasonable superpoint in embedding space. Besides, with the instance labels, to improve the superpoint performance on object boundaries, we propose the object boundary contrastive learning to enhance the feature distinguishability between tough points across the object boundaries. Extensive experiments demonstrate that SCL-OverSeg can effectively improve indoor oversegmentation performance, especially on object boundaries. The relevant codes will be available onhttps://github.com/sssssyf/SCL-OverSeg. Yifan Sun 0008, Chenguang Dai, Wenke Li, Song Ji, Anzhu Yu, Yiping Chen 0002, Hanyun Wang |
IEEE Trans. Multim. | 6 |
| 2024 | Updating Road Maps at City Scale With Remote Sensed Images and Existing Vector MapsabstractCurrently, many countries have built geo-information databases and gathered large amounts of geographic data. However, with the extensive construction of infrastructure and rapid expansion of cities, road updating process is imperative to maintain the high quality of current basic geographic information. Currently, road extraction and change detection are two commonly used methods to solve road updating problems. Most of the existing methods rely on a large number of accurate road labels to generate road information, while ignoring the use of quantities of available but incomplete road maps. In our work, we proposed a semi-supervised road extraction method specifically for road-updating applications (SRUNet). In this approach, historical road maps are fused with the latest remote sensing images, and state of the roads are updated directly. A multi-branch network is the core of the method, which consists of three noteworthy parts: Map Encoding Branch (MEB) proposed for representation learning, Boundary Enhancement Module (BEM) for improving the accuracy of boundary prediction, and Residual Refinement Module (RRM) for further optimizing the prediction results. We applied our method to two datasets: the DeepGlobe public dataset and our self-constructed dataset from Zhengzhou and Nanjing. Experimental results shows that our method achievs an improvement of 14.37% over the baseline approach. Notably, the addition of historical maps improved the model’s performance by 12.4%. Promising results were obtained on two cities’ large-scale road networks. With the reliable prediction results and improved performance, we believe SRUNet is meaningful for a wide range of road renewal applications. Xin Chen 0088, Anzhu Yu, Wenyue Guo, Qing Xu 0005, Bowei Wen |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Integrating Multiple Sources Knowledge for Class Asymmetry Domain Adaptation Segmentation of Remote Sensing ImagesabstractIn the existing unsupervised domain adaptation (UDA) methods for remote sensing images (RSIs) semantic segmentation, class symmetry is a widely followed ideal assumption, where the source and target RSIs have exactly the same class space. In practice, however, it is often very difficult to find a source RSI with exactly the same classes as the target RSI. More commonly, there are multiple source RSIs available. And there is always an intersection or inclusion relationship between the class spaces of each source–target pair, which can be referred to as class asymmetry. Nevertheless, the class asymmetry domain adaptation segmentation of RSIs with multiple sources has not yet been explored. To this end, a novel class asymmetry RSIs domain adaptation method is proposed for the first time in this article, which consists of four key components. First, a multibranch segmentation network is built to learn an expert for each source RSI. Second, a novel collaborative learning method with the cross-domain mixing strategy is proposed, to supplement the class information for each source while achieving the domain adaptation of each source–target pair. Third, a pseudolabel generation strategy is proposed to effectively combine the strengths of different experts, which can be flexibly applied to two cases where the source class union is equal to or includes the target class set. Fourth, a multiview-enhanced knowledge integration module is developed for high-level knowledge routing and transfer from multiple domains to target predictions. The experimental results of six different class settings on airborne and spaceborne RSIs show that the proposed method can effectively perform the multisource domain adaptation in the case of class asymmetry, and the obtained segmentation performance of target RSIs is significantly better than the existing relevant methods. Kuiliang Gao, Anzhu Yu, Xiong You, Wenyue Guo, Ke Li 0005, Ningbo Huang |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | A Fusion Framework for Producing an Accurate PWV Map With Spatiotemporal Continuity Based on GNSS, ERA5, and MODIS DataabstractSpatiotemporally seamless precipitable water vapor (PWV) maps with high accuracy, spatiotemporal resolution, and continuity are of significance for climatical research. The current frameworks for the PWV maps have predominantly concentrated on fusing PWV derived from ERA5 reanalysis (ERA-PWV) and Moderate Resolution Imaging Spectroradiometer (MODIS) near-infrared data (MOD-NIR-PWV), falling short in accuracy. In this study, PWV derived from global navigation satellite systems (GNSS-PWV) is introduced to produce spatiotemporally PWV maps with improved accuracy. The fusion framework involves two main steps: 1) spatial fusion for producing an initial PWV map with spatiotemporal continuity through spherical cap harmonic (SCH) analysis and 2) temporal fusion for producing a refined PWV map with high accuracy through residual correction. GNSS-PWV over 188 stations,$0.25^{\circ } \times 0.25^{\circ }$ERA-PWV, and$0.05^{\circ } \times 0.05^{\circ }$MOD-NIR-PWV over China from 2013 to 2018 are used to produce the daily$0.05^{\circ } \times 0.05^{\circ }$PWV maps. The performance is evaluated using out-of-sample data, containing 18 GNSS-PWV and 90 ERA-PWV, and independent reference data, containing radiosonde-derived PWV over 72 stations. When compared to the out-of-sample GNSS-PWV and ERA-PWV, the PWV maps exhibit the mean biases of −1.04 and −0.55 mm and the rms of 1.75 and 0.97 mm, respectively. These are equivalent to 8.7% and 32.1% reductions in bias and 25.5% and 49.5% reductions in RMSE relative to MOD-NIR-PWV. When radiosonde-derived PWV is used as the reference, the PWV maps have the mean bias and the RMSE of −0.53 and 2.21 mm, respectively, which outperforms ERA-PWV (−0.75 and 2.69 mm). These results indicate the effectiveness of the novel fusion framework in producing seamless PWV maps. Dantong Zhu, Qingfeng Hu, Kefei Zhang 0003, Suqin Wu, Peipei He, Anzhu Yu, Weibo Yin |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2023 | Open Self-Supervised Features for Remote-Sensing Image Scene Classification Using Very Few SamplesabstractBig models, large datasets, and self-supervised learning (SSL) have recently gained substantial research interest due to their potential to alleviate our reliance on annotations. Considering the current high generalization ability of self-supervised models in literature, we explore in the letter how helpful SSL can be for a crucial task in remote sensing (RS), image scene classification, when forced to rely on only a few labeled samples. We proposed a simple prototype-based classification procedure without training and fine-tuning, which uses open self-supervised features from the contrastive language-image pre-training (CLIP). We test our method by exploiting ready-to-use open features on four diversified benchmark datasets, including red-green-blue (RGB) and multispectral (MS) images. Highly competitive accuracy has been obtained compared to work with similar settings, i.e., based on an exceedingly small number of labels. To the best of our knowledge, our model is the first to achieve such high accuracy in austere label conditions. We further analyze our approach from different perspectives, including its advantages and limitations, reasons for its astonishing performance, potential applications, and future improvements. Chunping Qiu, Anzhu Yu, Xiaodong Yi 0002, Naiyang Guan, Dian-xi Shi, Xiaochong Tong |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2023 | Exploiting Discriminative Advantage of Spectrum for Hyperspectral Image Classification: SpectralFormer Enhanced by Spectrum Motion FeatureabstractAs for hyperspectral images (HSIs), the discrepancy of contiguous spectral information should be the main basis for the identification of ground objects. Due to the difficulty of spectral sequence coding and the spectrum similarity between categories, successful deep-learning-based classification methods always attempt to capture the spatial information to improve the accuracy by convolutional neural networks (CNNs) or other excellent spatial feature extractors. However, extracting spatial features is generally accompanied by the distortion of ground objects distribution and categories boundary. To effectively represent spectral features, the SpectralFormer based on transformer backbone can better capture the long-term dependence of the spectrum, which improves the performance of spectral feature methods significantly. However, it is still unable to compete with advanced spectral–spatial feature methods. To exploit the discriminative advantage of the spectrum fully, this letter introduces an efficient sparse-to-dense optical flow estimation method to track the spectrum variation in the HSI. Then, we take such a variation as a spectrum motion feature to enhance the original spectrum. At last, we continue to use the SpectralFormer to encode the concatenated spectrum sequence for classification. Extensive experiments show that the SpectralFormer enhanced by the spectrum motion feature (SF-SMF) significantly improves the performance of spectral feature methods, even surpassing advanced spectral–spatial feature methods. SF-SMF can avoid interference with additional spatial information to obtain exquisite whole-domain classification maps, showing its practical value. The codes will be public athttps://github.com/sssssyf/SF-SMF. Yifan Sun 0008, Bing Liu 0018, Xuchu Yu, Anzhu Yu, Pengqiang Zhang, Zhixiang Xue |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2023 | Prototype and Context-Enhanced Learning for Unsupervised Domain Adaptation Semantic Segmentation of Remote Sensing ImagesabstractIn unsupervised domain adaptation (UDA) of remote sensing images (RSIs), the huge inter-domain discrepancies and intra-domain variances lead to complicated class-level relations. Specifically, the instances of the same class differ greatly while instances of different classes are similar, whether across different RSIs domains or within the same RSIs domain. However, existing methods cannot fully consider these problems, limiting the performance of UDA semantic segmentation of RSIs. To this end, this paper proposes a novel cross-domain multi-prototypes learning method, the core idea of which is to abstract the cross-and intra-domain class-level relations into multiple prototypes. Specifically, the multiple prototypes belonging to different classes can detailedly describe complex inter-class relations, and the multiple prototypes within the same class can better model rich intra-class relations. Further, the source and target samples are jointly used for prototypes calculation, to fully fuse the feature information of different RSIs. In a nutshell, utilizing the samples from different RSIs domains to learn multiple prototypes for each class can achieve better domain alignment at the class level. In addition, considering that RSIs simultaneously contain large targets with wide coverage and important small targets, two masked consistency learning strategies are designed to better explore the contextual structure of target RSIs and improve the quality of pseudo labels for prototype updating. The global consistency strategy can strengthen the utilization of global context relations, while the local consistency strategy can further improve the learning of local context details. Therefore, the proposed method is actually a prototype and context enhanced learning method for UDA semantic segmentation of RSIs. Extensive experiments demonstrate that the proposed method can achieve better performance than existing state-of-the-art UDA methods. Kuiliang Gao, Anzhu Yu, Xiong You, Chunping Qiu, Bing Liu 0018 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Hyperspectral Meets Optical Flow: Spectral Flow Extraction for Hyperspectral Image ClassificationabstractHyperspectral image (HSI) classification has always been recognised as a difficult task. It is therefore a research hotspot in remote sensing image processing and analysis, and a number of studies have been conducted to better extract spectral and spatial features. This study aimed to track the variation of the spectrum in hyperspectral images from a sequential data perspective to obtain more distinguishable features. Based on the characteristics of optical flow, this study introduces an optical flow technique for the extraction of spectral flow that denotes the spectral variation and implements a dense optical flow extraction method based on deep matching. Lastly, the extracted spectral flow are combined with the original spectral features and input into a commonly used support vector machine (SVM) classifier to complete the classification. Extensive classification experiments on three benchmark HSI test sets show that the classification accuracy obtained by the spectral flow extracted in this study (SpectralFlow) is higher than traditional spatial feature extraction methods, texture feature extraction methods, and the latest deep-learning-based methods. Furthermore, the proposed method can produce finer classification thematic maps, thereby demonstrating strong practical application potential. Bing Liu 0018, Yifan Sun 0008, Anzhu Yu, Zhixiang Xue, Xibing Zuo |
IEEE Trans. Image Process. | 3 |
| 2022 | Self-Supervised Feature Learning and Few-Shot Land Cover Classification for Cross-Modal Remote Sensing ImagesabstractWith the rapid development of remote sensing data acquisition technology, there are multimodal images over the same observed scenes. These multimodal remote sensing images could provide complementary valuable information for land cover classification. In this article, we propose a novel self-supervised feature learning and few-shot classification model for multimodal remote sensing images, called S2FL. Specifically, a contrastive learning architecture is investigated to learn spatial feature representations from very high resolution (VHR) image. And the spectral features from hyperspectral data are integrated with learned spatial features for few-shot land cover classification. Classification experiments are conducted on a widely-used dataset, i.e., Houston 2018, to verify the effectiveness and superiority of the proposed S2FL model compared with several state-of-the-art baseline approaches. Zhixiang Xue, Xuchu Yu, Pengqiang Zhang, Xiong Tan, Anzhu Yu, Bing Liu 0018 |
IGARSS | 5 |
| 2022 | Multiscale Feature Learning by Transformer for Building Extraction From Satellite ImagesabstractExtracting buildings from very high-resolution satellite images is a challenging yet important task for applications such as urban monitoring. Multiscale feature learning proves to be a potential solution toward accurate extraction of buildings. This study exploits a powerful multiscale feature learning module, a hierarchical vision transformer by shifted windows (swin), as a backbone within a building extraction network. To this end, we first designed a general structure for building extraction, consisting of a backbone to extract multiscale features and a head network to fuse and refine features. Then, we integrated swin into the structure as a backbone and utilized channel-wise and spatial-wise enhancement in a head network. Experimental results show that our method achieves improvements regarding both F1-score and intersection over union (IoU) compared to the multiple attending path neural network (MAP-Net), which is the current state-of-the-art (SOTA) algorithm for building extraction from remote sensing images. Our study thus confirms the potential of swin transformers as backbones for semantic segmentation tasks based on satellite images. Xin Chen 0088, Chunping Qiu, Wenyue Guo, Anzhu Yu, Xiaochong Tong, Michael Schmitt 0003 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2022 | Pixel-Level Self-Supervised Learning for Semi-Supervised Building Extraction From Remote Sensing ImagesabstractThe building extraction from remote sensed images ash been a challenging yet vital task for applicable purposes such as urban monitoring and cartography. Most of the existing learning based approaches focus on the supervised building extraction methods, of which the models should be trained with images and the corresponding labels. This research exploits a self-supervised approach for building extraction, which could train the backbone within a building extraction network without annotations. Specifically, the backbone is initially trained with a pixel-level self-supervised module instead of commonly used supervised approaches or instance-level self-supervised modules. Next, the pretrained backbone is embedded into a task-specific network followed by tuning with limited annotations. The experiments were conducted on three popular datasets and the results show that our method achieves improvements regarding both intersection over union (IoU) and F1-score compared to supervised approach and instance-level self-supervised methods. Our study thus confirms the potential of pixel-level self-supervised approach for semantic segmentation for remote sensing images. Anzhu Yu, Bing Liu 0018, Xuefeng Cao, Chunping Qiu, Wenyue Guo, Yujun Quan |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2022 | Perceiving Spectral Variation: Unsupervised Spectrum Motion Feature Learning for Hyperspectral Image ClassificationabstractIn recent years, deep-learning-based hyperspectral image (HSI) classification methods have achieved significant development. The superior capability of feature extraction from these data-driven methods dramatically improves the classification performance. However, the previous methods usually require to retrain the network from scratch to obtain the capability of feature extraction adaptive for the target image when facing a new HSI to be classified, which is a time-consuming and redundant process. In this paper, we consider putting this process ahead and making the network have a robust capability of feature extraction with generalization through pre-training. Therefore, the network enables to directly extract features of the target HSI without re-training. For this purpose, we rethink the three-dimension (3D) HSI data from a perspective of spectral sequence, and we attempt to extract the spectral variation information as the spectrum motion feature. Then, we construct an unsupervised spectrum motion feature learning framework (SMF-UL), which can be pre-trained on mass unlabeled HSI data to learn the knowledge about perceiving spectral variation. Furthermore, to achieve the expansion of source data for pre-training, we develop an extendable training dataset construction method, which can integrate HSIs of different sizes, number of bands and sensors into a unified training set to utilize the rapidly growing mass unlabeled HSI data effectively. Finally, we use the trained network to directly extract the spectrum motion feature of the target HSI for classification, so the laborious re-training of the network can be avoided. Extensive experiments show that the proposed SMF-UL acquires the robust capability of feature extraction with generalization through unsupervised learning on mass unlabeled HSI data, and the classification performance of extracted spectrum motion feature is competitive to advanced in-domain and cross-domain methods, which shows its flexibility and superiority. The code of SMF-UL will be open at: https://github.com/sssssyf/SMF-UL. Yifan Sun 0008, Bing Liu 0018, Xuchu Yu, Anzhu Yu, Kuiliang Gao, Lei Ding 0008 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Self-Supervised Feature Representation and Few-Shot Land Cover Classification of Multimodal Remote Sensing ImagesabstractAlthough deep learning-based approaches have made significant progress in remote sensing image classification, the supervised learning paradigm has shortcomings under a limited number of labeled samples, which restricts the classification performance to a great extent. In this article, we investigate an effective self-supervised feature representation architecture (SSFR) for multimodal remote sensing images few-shot land cover classification. Specifically, we exploit multiview learning strategy to construct multiple views from multimodal remote sensing images. This method builds several complementary views of the same observed scenes from hyperspectral images or different modalities of remote sensing data. Then we build the deep feature extractor to learn high-level feature representations from each view via contrastive learning. The contrastive learning aggregates the samples of the same scene while separating samples of different scenes in the latent space, and this process does not require any labeled information. What’s more, to learn more robust features from different views, we utilize multitask learning strategy to train the feature extraction network. Finally, a lightweight machine learning method is employed to classify the learned features using a few annotated samples. To further demonstrate the self-supervised feature learning capability of the proposed model, we train the feature representation network in multiple source datasets. Comprehensive feature learning and classification experiments have certified the effectiveness and superiority of the proposed method. Zhixiang Xue, Bing Liu 0018, Anzhu Yu, Xuchu Yu, Pengqiang Zhang, Xiong Tan |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Multiscale Deep Learning Network With Self-Calibrated Convolution for Hyperspectral and LiDAR Data Collaborative ClassificationabstractIn this article, we propose a novel multiscale deep learning network with self-calibrated convolution (MSNetSC) for hyperspectral and light detection and ranging (LiDAR) data collaborative classification. Conventional deep learning methods have limitations in extracting multiscale features at a granular level from multimodality data and fusing these features in a context-awareness way, which will severely restrict the performance of hyperspectral and LiDAR data joint classification. The proposed multiscale deep learning network utilizes a hierarchical residual structure combined with self-calibrated convolution to extract features with different receptive fields, and this can enhance the model’s capability to represent the multimodality data. Besides, we employ spectral and spatial self-attention modules to adaptively calibrate weights of features with different scales, thereby enhancing the discriminative ability of extracted multiscale features. Furthermore, the attentional feature fusion module can dynamically and adaptively fuse the features from multimodality data in a contextual scale-aware way, and this attention-based feature fusion method will further improve the collaborative classification performance of hyperspectral and LiDAR data. Four benchmark multimodality data (i.e., hyperspectral and LiDAR data) sets collected by different sensors and at different acquisition times are employed for joint classification experiments. These comparative classification results and ablation study sufficiently certify the superiority of the proposed model in terms of collaborative classification accuracy when compared with other state-of-the-art methods. Zhixiang Xue, Xuchu Yu, Xiong Tan, Bing Liu 0018, Anzhu Yu, Xiangpo Wei |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | Self-Supervised Feature Learning for Multimodal Remote Sensing Image Land Cover ClassificationabstractDeep learning models have shown great potential in remote sensing image processing and analysis. Nevertheless, there are insufficient labeled samples to train deep networks, which seriously affects the performance of these models. To resolve this contradiction, we propose a generative self-supervised feature learning (S2FL) architecture for multimodal remote sensing image land cover classification. Specifically, multiple complementary observed views are constructed from multimodal remote sensing images, which are employed for following generative self-supervised learning. The proposed S2FL architecture is capable of extracting high-level meaningful feature representations from multiview data, and this process does not require any labeled information, providing a feasible solution to relieve the urgent need for annotated samples. The learned features are normalized and merged with corresponding spectral information to further improve the discriminative capability of feature representations, and we utilize these fused features for land cover classification. Compared with existing supervised, semi-supervised, and self-supervised approaches, the proposed generative self-supervised model achieves superior performance in terms of feature learning and land cover classification, especially in the small sample classification case. Zhixiang Xue, Xuchu Yu, Anzhu Yu, Bing Liu 0018, Pengqiang Zhang, Shentong Wu |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Unsupervised Meta Learning With Multiview Constraints for Hyperspectral Image Small Sample set ClassificationabstractThe difficulties of obtaining sufficient labeled samples have always been one of the factors hindering deep learning models from obtaining high accuracy in hyperspectral image (HSI) classification. To reduce the dependence of deep learning models on training samples, meta learning methods have been introduced, effectively improving the classification accuracy in small sample set scenarios. However, the existing methods based on meta learning still need to construct a labeled source data set with several pre-collected HSIs, and must utilize a large number of labeled samples for meta-training, which is actually time-consuming and labor-intensive. To solve this problem, this paper proposes a novel unsupervised meta learning method with multiview constraints for HSI small sample set classification. Specifically, the proposed method first builds an unlabeled source data set using unlabeled HSIs. Then, multiple spatial-spectral multiview features of each unlabeled sample are generated to construct tasks for unsupervised meta learning. Finally, the designed residual relation network is used for meta-training and small sample set classification based on the voting strategy. Compared with existing supervised meta learning methods for HSI classification, our method can only utilize HSIs without any label for unsupervised meta learning, which significantly reduces the number of requisite labeled samples in the whole classification process. To verify the effectiveness of the proposed method, extensive experiments are carried out on 8 public HSIs in the cross-domain and in-domain classification scenarios. The statistical results demonstrate that, compared with existing supervised meta learning methods and other advanced classification models, the proposed method can achieve competitive or better classification performance in small sample set scenarios. Kuiliang Gao, Bing Liu 0018, Xuchu Yu, Anzhu Yu |
IEEE Trans. Image Process. | 4 |
| 2022 | Deep Hierarchical Vision Transformer for Hyperspectral and LiDAR Data ClassificationabstractIn this study, we develop a novel deep hierarchical vision transformer (DHViT) architecture for hyperspectral and light detection and ranging (LiDAR) data joint classification. Current classification methods have limitations in heterogeneous feature representation and information fusion of multi-modality remote sensing data (e.g., hyperspectral and LiDAR data), these shortcomings restrict the collaborative classification accuracy of remote sensing data. The proposed deep hierarchical vision transformer architecture utilizes both the powerful modeling capability of long-range dependencies and strong generalization ability across different domains of the transformer network, which is based exclusively on the self-attention mechanism. Specifically, the spectral sequence transformer is exploited to handle the long-range dependencies along the spectral dimension from hyperspectral images, because all diagnostic spectral bands contribute to the land cover classification. Thereafter, we utilize the spatial hierarchical transformer structure to extract hierarchical spatial features from hyperspectral and LiDAR data, which are also crucial for classification. Furthermore, the cross attention (CA) feature fusion pattern could adaptively and dynamically fuse heterogeneous features from multi-modality data, and this contextual aware fusion mode further improves the collaborative classification performance. Comparative experiments and ablation studies are conducted on three benchmark hyperspectral and LiDAR datasets, and the DHViT model could yield an average overall classification accuracy of 99.58%, 99.55%, and 96.40% on three datasets, respectively, which sufficiently certify the effectiveness and superior performance of the proposed method. Zhixiang Xue, Xiong Tan, Xuchu Yu, Bing Liu 0018, Anzhu Yu, Pengqiang Zhang |
IEEE Trans. Image Process. | 5 |
| 2021 | PL-VSCN: Patch-level vision similarity compares network for image matchingabstractAbstract Image matching plays an important role in various computer vision tasks, such as image retrieval and loop closure detection in Simultaneous Localization and Mapping. The authors propose a discriminative patch‐based image matching method that converts the problem of whole image matching to that of local patch matching. To construct the patch representation, the Patch‐Level Vision Similarity Compare Network (PL‐VSCN) is proposed to produce the patch feature. In the image matching process, local patches that potentially contain objects within images are initially detected, and the discriminative feature of each patch is extracted based on the pre‐trained PL‐VSCN. Then, the similarities between the patch pairs are calculated to construct the similarity matrix, and the corresponding patch pairs are detected based on the mutual matching mechanism on the similarity matrix. Experimental results indicate that the proposed PL‐VSCN can generate the discriminative patch feature, which can accurately match the patch pairs with the corresponding content and distinguish those with non‐corresponding content. In addition, the comparison experiments demonstrate that the proposed image matching method outperforms existing approaches on most datasets and effectively completes the image matching task. Xiong You, Qin Li 0005, Ke Li 0005, Anzhu Yu, Shuhui Bu |
IET Comput. Vis. | 4 |
| 2021 | Deep Multiview Learning for Hyperspectral Image ClassificationabstractRecently, the field of hyperspectral image (HSI) classification is dominated by deep learning-based methods. However, training deep learning models usually needs a large number of labeled samples to optimize thousands of parameters. In this article, a deep multiview learning method is proposed to deal with the small sample problem of HSI. First, two views of an HSI scene are constructed by applying principal component analysis to different bands. Second, a deep residual network is designed to embed the different views of a sample to a latent space. The designed deep residual network is trained by maximizing agreement between differently augmented views of the same data sample via a contrastive loss in the latent space. Note that the training procedure of the designed deep residual network does not use labeled information. Therefore, the proposed method belongs to the category of unsupervised learning, which could alleviate the lack of labeled training samples. Finally, a conventional machine learning method (e.g., support vector machine) is used to complete the classification task in the learned latent space. To demonstrate the effectiveness of the proposed method, extensive experiments are carried on four widely used hyperspectral data sets. The experimental results demonstrate that the proposed method could improve the classification accuracy with small samples. Bing Liu 0018, Anzhu Yu, Xuchu Yu, Ruirui Wang, Kuiliang Gao, Wenyue Guo |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2019 | Deep Few-Shot Learning for Hyperspectral Image ClassificationabstractDeep learning methods have recently been successfully explored for hyperspectral image (HSI) classification. However, training a deep-learning classifier notoriously requires hundreds or thousands of labeled samples. In this paper, a deep few-shot learning method is proposed to address the small sample size problem of HSI classification. There are three novel strategies in the proposed algorithm. First, spectral–spatial features are extracted to reduce the labeling uncertainty via a deep residual 3-D convolutional neural network. Second, the network is trained by episodes to learn a metric space where samples from the same class are close and those from different classes are far. Finally, the testing samples are classified by a nearest neighbor classifier in the learned metric space. The key idea is that the designed network learns a metric space from the training data set. Furthermore, such metric space could generalize to the classes of the testing data set. Note that the classes of the testing data set are not seen in the training data set. Four widely used HSI data sets were used to assess the performance of the proposed algorithm. The experimental results indicate that the proposed method can achieve better classification accuracy than the conventional semisupervised methods with only a few labeled samples. Bing Liu 0018, Xuchu Yu, Anzhu Yu, Pengqiang Zhang, Ruirui Wang |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2018 | Supervised Deep Feature Extraction for Hyperspectral Image ClassificationabstractHyperspectral image classification has become a research focus in recent literature. However, well-designed features are still open issues that impact on the performance of classifiers. In this paper, a novel supervised deep feature extraction method based on siamese convolutional neural network (S-CNN) is proposed to improve the performance of hyperspectral image classification. First, a CNN with five layers is designed to directly extract deep features from hyperspectral cube, where the CNN can be intended as a nonlinear transformation function. Then, the siamese network composed by two CNNs is trained to learn features that show a low intraclass and high interclass variability. The important characteristic of the presented approach is that the S-CNN is supervised with a margin ranking loss function, which can extract more discriminative features for classification tasks. To demonstrate the effectiveness of the proposed feature extraction method, the features extracted from three widely used hyperspectral data sets are fed into a linear support vector machine (SVM) classifier. The experimental results demonstrate that the proposed feature extraction method in conjunction with a linear SVM classifier can obtain better classification performance than that of the conventional methods. Bing Liu 0018, Xuchu Yu, Pengqiang Zhang, Anzhu Yu, Qiongying Fu, Xiangpo Wei |
IEEE Trans. Geosci. Remote. Sens. | 4 |