VLDB 2026 Research / reviewers in the wild / expert
Zhe Xu 0016
dblp:97/3701-16
· DBLP profile ↗
8ranked-venue papers
3as first author
8since 2021 · last 2026
0000-0001-5264-0235ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Unlocking Pseudolabel Potential and Alignment for Unpaired Cross-Modality Adaptation in Remote Sensing Image SegmentationabstractWith the growth of multisource sensor technology, multimodal learning has become pivotal in remote sensing (RS) image segmentation. Despite its potential, current methods face challenges in acquiring large-scale paired samples. When annotated optical images are available, but synthetic aperture radar (SAR) images lack annotations, learning discriminative features for SAR images from optical images becomes difficult. Unsupervised domain adaptation (UDA) offers a potential solution to this challenge, which we refer to as unpaired cross-modality UDA. In this article, we propose unlocking pseudolabel potential and alignment (ULPA) for unpaired cross-modality adaptation in RS image segmentation, a novel one-stage adaptation framework designed to enhance cross-modality knowledge transfer. Our approach employs a prototypical multidomain alignment (PMDA) strategy, which reduces the modality gap through contrastive learning between features and prototypes of identical classes across different modalities. In addition, we introduce the unreliable-sample-guided feature contrast (UFC) loss to address the underutilization of unreliable pixels during training. This strategy separates reliable and unreliable pixels based on prediction confidence, assigning unreliable pixels to a category-wise queue of negative samples, thus ensuring all candidate pixels contribute to the training process. Extensive experiments show that the integration of PMDA and UFC loss can lead to more effective cross-modality domain alignment and substantially boost the model's generalization capability. Zhe Xu 0016, Jie Geng 0005, Wen Jiang 0002, Shuai Song |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2025 | Unsupervised Remote Sensing Image Semantic Segmentation Based on Multiscale Contrastive Domain AdaptationabstractUnsupervised domain adaptation for remote sensing image semantic segmentation aims to train a deep model on the labeled source domain and apply it to the unlabeled target domain. However, resolution and scene inconsistencies of cross-domain remote sensing images lead to great distribution differences, which weakens the semantic segmentation effect. To solve the above issues, an unsupervised remote sensing image semantic segmentation method is proposed based on multi-scale contrastive domain adaptation. Firstly, the mean teacher model is introduced into the unsupervised domain adaptation paradigm to generate pseudo-labels for target domain data, thereby achieving the cross-domain segmentation capability. A dynamic class balance sampling method is proposed to mitigate the class imbalance problem in cross-domain data by increasing the sampling frequency of the categories with fewer samples. Then, a data augmentation method called cross-domain mixup is developed to reduce the gap between the source and target domains. Finally, a multi-scale cross-domain contrastive loss is developed, which introduces the contrastive learning to learn domain-consistent features across the source and target domains, resulting in a more coherent and discriminative feature representation. Experimental results show that the proposed method can yield superior performance for unsupervised remote sensing image semantic segmentation. Jie Geng 0005, Shuai Song, Zhe Xu 0016, Wen Jiang 0002 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Pseudo-label meta-learner in semi-supervised few-shot learning for remote sensing image scene classification
Wang Miao, Zhe Xu 0016, Jie Geng 0005, Wen Jiang 0002 |
Appl. Intell. | 3 |
| 2023 | ECAE: Edge-Aware Class Activation Enhancement for Semisupervised Remote Sensing Image Semantic SegmentationabstractRemote sensing image semantic segmentation (RSISS) remains challenging due to the scarcity of labeled data. Semi-supervised learning can leverage pseudo-labels to enhance the model’s ability to learn from unlabeled data. However, accurately generating pseudo-labels for RSISS remains a significant challenge that severely affects the model’s performance, especially for the edges of different classes. In order to overcome these issues, we propose a semi-supervised semantic segmentation framework for remote sensing images based on edge-aware class activation enhancement (ECAE). Firstly, the baseline network is constructed based on the average teacher model, which separates the training of labeled and unlabeled data using student and teacher networks. Secondly, considering local continuity and global discreteness of object distribution in remote sensing images, the class activation mapping enhancement (CAME) network is designed to predict local areas more remarkably. Finally, the edge-aware network (EAN) is proposed to improve the performance of edge segmentation in remote sensing images. The combination of the CAME with the EAN further heightens the generation of high-confidence pseudo-labels. Experiments were performed on two publicly available remote sensing semantic segmentation datasets, Potsdam and ISPRS Vaihingen, which verify the superiorities of the proposed ECAE model. Wang Miao, Zhe Xu 0016, Jie Geng 0005, Wen Jiang 0002 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | MMT: Mixed-Mask Transformer for Remote Sensing Image Semantic SegmentationabstractRemote sensing image semantic segmentation is a crucial step in the intelligent interpretation of remote sensing. Most of the current approaches are based on the attention mechanism to enhance long-range representations. However, these works ignore the key problem of foreground-background imbalance, and their performances encounter a bottleneck. In this paper, we introduce mask classification into remote sensing image interpretation for the first time, and propose a novel mixed-mask Transformer (MMT) for remote sensing image semantic segmentation. Specifically, we propose a mixed-mask attention mechanism, a simple but effective module, which assists the network to learn more explicit intraclass and interclass correlations by capturing long-range interdependent representations. In addition, a progressive multi-scale learning strategy is proposed to solve the problem of large scale-varied targets in remote sensing images, which integrates semantic and visual representations of different scale targets by efficiently utilizing large scale feature maps in Transformer. Experimental results show that the proposed MMT exceeds the existing alternative approaches and achieves state-of-the-art performance on three semantic segmentation datasets. Zhe Xu 0016, Jie Geng 0005, Wen Jiang 0002 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Rotated Object Detection of Remote Sensing Image Based on Binary Smooth Encoding and Ellipse-Like Focus LossabstractRemote sensing image object detection has been widely developed in many applications. Objects in remote sensing data have the characteristic of arbitrary directions, which leads to poor detection performance based on horizontal box detectors. To address this issue, a novel rotated object detection model based on binary smooth encoding and ellipse-like focus loss is proposed in this paper. Firstly, a multi-layer feature fusion network with attention mechanism is developed to extract features of multi-scale objects. Then, an anchor free detection module with binary smooth encoding is proposed, which aims to predict the rotated angles of objects. Moreover, an ellipse-like focus loss is proposed to obtain high-quality bounding boxes drawing near the object center. Experimental results on two public remote sensing datasets verify that the proposed method can yield superior detection performance than other related rotated object detection models. Jie Geng 0005, Zhe Xu 0016, Zihao Zhao 0009, Wen Jiang 0002 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2021 | Pseudo-loss Confidence Metric for Semi-supervised Few-shot LearningabstractSemi-supervised few-shot learning is developed to train a classifier that can adapt to new tasks with limited labeled data and a fixed quantity of unlabeled data. Most semi-supervised few-shot learning methods select pseudo-labeled data of unlabeled set by task-specific confidence estimation. This work presents a task-unified confidence estimation approach for semi-supervised few-shot learning, named pseudo-loss confidence metric (PLCM). It measures the data credibility by the loss distribution of pseudo-labels, which is synthetical considered multi-tasks. Specifically, pseudo-labeled data of different tasks are mapped to a unified metric space by mean of the pseudo-loss model, making it possible to learn the prior pseudo-loss distribution. Then, confidence of pseudo-labeled data is estimated according to the distribution component confidence of its pseudo-loss. Thus highly reliable pseudo-labeled data are selected to strengthen the classifier. Moreover, to overcome the pseudo-loss distribution shift and improve the effectiveness of classifier, we advance the multi-step training strategy coordinated with the class balance measures of class-apart selection and class weight. Experimental results on four popular benchmark datasets demonstrate that the proposed approach can effectively select pseudo-labeled data and achieve the state-of-the-art performance. Jie Geng 0005, Wen Jiang 0002, Xinyang Deng, Zhe Xu 0016 |
ICCV | 5 |
| 2021 | Triplet Attention Feature Fusion Network for SAR and Optical Image Land Cover ClassificationabstractWith recent advances in remote sensing, abundant multimodal data are available for applications. However, considering the redundancy and the huge domain differences among multimodal data, how to effectively integrate these data is becoming important and challenging. In this paper, we proposed a triplet attention feature fusion network (TAFFN) for SAR and optical image fusion classification. Specifically, spatial attention module and spectral attention module based on self-attention mechanism are developed to extract spatial and spectral long-range information from the SAR image and optical image respectively, at the same time, cross-attention mechanism is proposed to capture the long-range interactive representation. Triplet attentions are concatenated to further integrate the complementary information of SAR and optical images. Experiments on a SAR and optical multimodal dataset demonstrate that the proposed method can achieve the state-of-the-arts performance. Zhe Xu 0016, Jinbiao Zhu, Jie Geng 0005, Xinyang Deng, Wen Jiang 0002 |
IGARSS | 1 |