EDBT 2026 Demo / reviewers in the wild / expert
Zhuojun Xie
dblp:89/8664
· DBLP profile ↗
12ranked-venue papers
5as first author
12since 2021 · last 2025
0009-0005-2087-5167ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 9 · 4 first-author · 9 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Progressive joint distribution alignment network for cross-scene hyperspectral image classification
Zhuojun Xie, Puhong Duan, Xudong Kang, Wang Liu 0001, Shutao Li 0001 |
Sci. China Inf. Sci. | 1 |
| 2025 | SSFNet: Spectral-Spatial Fusion Network for Hyperspectral Remote Sensing Scene ClassificationabstractRemote sensing scene classification (RSSC) plays a vital role in a variety of applications and has attracted much more attention. In recent years, much progress has been made to release diverse datasets or develop all kinds of techniques for scene classification of multispectral remote sensing images. Nevertheless, very few studies have focused on hyperspectral image scene classification. Moreover, the existing scene classification approaches fail to fully employ the rich spectral information of the input images, which cannot achieve satisfactory performance for hyperspectral images. To alleviate these issues, this work proposes a spectral-spatial fusion network (SSFNet) for hyperspectral RSSC (HRSSC). First, a multiscale regional growth search (MSRGS) method is designed to extract salient object regions from the hyperspectral remote sensing scene. Then, a three-stream network architecture is proposed to extract the global spatial, local spatial, and spectral features, respectively. Finally, the fully connected layer is performed on the extracted features to obtain a class score followed by a decision fusion scheme to generate the final classification result. To evaluate the effectiveness of the proposed SSFNet, we created a publicly available benchmark for the HRSSC dataset, which contains 1445 hyperspectral images, covering 11 scene classes. Experiments on the HRSSC database claim that the proposed SSFNet can attain superior classification performance with respect to other state-of-the-art scene classification techniques. The code of the proposed SSFNet will be available athttps://github.com/PuhongDuan/SSFNet. Puhong Duan, Jialin Zheng, Zhuojun Xie, Xudong Kang, Jianwei Yin, Shutao Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | Learning From Vision Foundation Models for Cross-Domain Remote Sensing Image SegmentationabstractCross-domain image segmentation plays a crucial role in the field of remote sensing. Current approaches often rely on a mean-teacher model that is integrated from student models to guide the training of the student model itself. However, the feature space of the mean-teacher model exhibits significant domain discrepancy and considerable class overlap, which results in suboptimal performance. Motivated by the idea of learning from stronger teachers, we introduce a robust domain adaptation method called LFMDA. This novel approach is the first to explicitly enhance cross-domain semantic segmentation performance by leveraging vision foundation models (VFMs) within remote sensing applications. Specifically, we propose a prototypical contrastive knowledge distillation loss (PCD) that enables the student model to produce domain-invariant yet category-discriminative features by distilling knowledge from a domain-generalized VFM teacher. Additionally, we introduce a local region homogenization strategy (LRH) to generate high-quality and high-quantity pseudo-labels by incorporating a Segment Anything Model (SAM). Extensive empirical evaluations demonstrate that our method outperforms existing approaches, setting a new state-of-the-art (SOTA) method in domain-adaptive remote sensing image segmentation. The code is available at https://github.com/StuLiu/LFMDA. Wang Liu 0001, Puhong Duan, Zhuojun Xie, Xudong Kang, Shutao Li 0001 |
IEEE Trans. Image Process. | 3 |
| 2025 | SOSNet: Real-Time Small Object Segmentation via Hierarchical Decoding and Example MiningabstractReal-time semantic segmentation plays an important role in auto vehicles. However, most real-time small object segmentation methods fail to obtain satisfactory performance on small objects, such as cars and sign symbols, since the large objects usually tend to devote more to the segmentation result. To solve this issue, we propose an efficient and effective architecture, termed small objects segmentation network (SOSNet), to improve the segmentation performance of small objects. The SOSNet works from two perspectives: methodology and data. Specifically, with the former, we propose a dual-branch hierarchical decoder (DBHD) which is viewed as a small-object sensitive segmentation head. The DBHD consists of a top segmentation head that predicts whether the pixels belong to a small object class and a bottom one that estimates the pixel class. In this situation, the latent correlation among small objects can be fully explored. With the latter, we propose a small object example mining (SOEM) algorithm for balancing examples between small objects and large objects automatically. The core idea of the proposed SOEM is that most of the hard examples on small-object classes are reserved for training while most of the easy examples on large-object classes are banned. Experiments on three commonly used datasets show that the proposed SOSNet architecture greatly improves the accuracy compared to the existing real-time semantic segmentation methods while keeping efficiency. The code will be available at https://github.com/StuLiu/SOSNet. Wang Liu 0001, Xudong Kang, Puhong Duan, Zhuojun Xie, Xiaohui Wei 0001, Shutao Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Learn From Segment Anything Model: Local Region Homogenizing for Cross-Domain Remote Sensing Image SegmentationabstractUnsupervised domain adaption (UDA) has gained popularity in narrowing performance gaps across domains in remote sensing image semantic segmentation (RSISS). However, current UDA methods suffer from serious noisy pseudo-labels, adversely affecting domain adaptation performance. In this work, a local region homogenizing domain adaptation method (RegDA) is proposed to tackle this issue. Specifically, a generalized segment anything model (SAM) is utilized to obtain the semantic-consistent regions for the images in the target domain. Furthermore, a pixel-level voting scheme is proposed to get the semantic label for each local region and assign it to each pixel within this region. In this way, more reliable pseudo-labels are obtained and domain adaptation performance is improved. Experiment results on ISPRS datasets demonstrate that the proposed RegDA outperforms previous UDA approaches for RSISS. The code will be available at https://github.com/StuLiu/RegDA. Wang Liu 0001, Puhong Duan, Zhuojun Xie, Xudong Kang, Shutao Li 0001 |
IGARSS | 3 |
| 2024 | CTSFFNet: Cross-Temporal Symmetric Feature Fusion Network for Hyperspectral Image Change DetectionabstractHyperspectral change detection (HCD) aims to identify the changed and unchanged pixels in bitemporal images, which has been applied in various aspects. Currently, many deep learning-based change detection methods have been developed. However, existing change detection methods only focus on changed or temporal information while neglecting the complementary information between them. To solve this issue, a novel cross-temporal symmetric feature fusion network (CTSFFNet) is proposed for change detection of hyperspectral images. First, we perform pixel-wise subtraction and concatenation on the multi-temporal hyperspectral images to obtain the difference data and temporal data, respectively. Then, a three-layer convolutional neural network is performed on the difference data and temporal data to yield the difference and temporal features. Finally, a cross-temporal symmetric feature fusion (CTSSF) module is designed to merge the extracted features followed by a fully connected layer to obtain the final change regions. Experiments on two popular datasets demonstrate that the proposed CTSSFNet achieves superior detection performance compared to other state-of-the-art methods. Xukun Lu, Puhong Duan, Zhuojun Xie, Xudong Kang |
IGARSS | 4 |
| 2024 | Prototype-based Inter-Intra Domain Alignment Network for Unsupervised Cross-Scene Hyperspectral Image ClassificationabstractUnsupervised cross-scene hyperspectral image classification transfers the learnable knowledge from a labeled source scene to an unlabeled target scene. Currently, many statistical distribution alignment methods are introduced to mitigate domain discrepancy. However, these methods ignore the finer class specific structure which may cause negative transfer. To solve this issue, a prototype-based inter-intra domain alignment network is proposed for unsupervised cross-scene hyperspectral image classification. Specifically, a prototype-based inter-intra alignment method is proposed to narrow the feature distribution gap. Furthermore, an uncertainty estimation is developed to obtain highly reliable pseudo-labels in the target scene. Experiment results on several datasets imply that the proposed method outperform several cutting-edge unsupervised classification methods. Zhuojun Xie, Puhong Duan, Wang Liu 0001, Xudong Kang, Shutao Li 0001 |
IGARSS | 1 |
| 2024 | FAA-Det: Feature Augmentation and Alignment for Anchor-Free Oriented Object DetectionabstractOriented object detection with remote sensing scenes has made excellent progress in recent years, especially using anchor-free detectors. Without the limitation of inherent prior spatial information, anchor-free detectors regress the detection boxes from the object center or edge in an elegant way. However, anchor-free detectors suffer severe feature misalignment and inconsistency between classification and regression. Especially in remote sensing scenes, there are densely arranged instances and multi-scale representations, which will affect the detection accuracy. Therefore, a feature augmentation module (FAM) and an oriented feature alignment (OFA) module are proposed for oriented object detection called FAA-Det. More specifically, we first introduce a FAM to enhance the object representation. After that, the augmented feature maps will be fed into OFA for feature alignment and accurate detection. OFA has two independent branches for classification and regression, and their separate structures can alleviate the inconsistency in detection. FAM and OFA comprise the FAA-Head in our detector. Extensive evaluation demonstrates the effectiveness of our proposed FAA-Det that performs the state-of-the-art (SOTA) mean average precision (mAP) on the DOTA and HRSC2016 datasets without bells and whistles. Our code will be available athttps://github.com/jimuIee/FAA-Det. Zikang Li, Wang Liu 0001, Zhuojun Xie, Xudong Kang, Puhong Duan, Shutao Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Uncertain Example Mining Network for Domain Adaptive Segmentation of Remote Sensing ImagesabstractDomain adaptive segmentation has recently gained more and more attention in the remote sensing field. However, current methods often generate a significant number of uncertain examples, i.e., noisy pseudo-labels, in the target domain, which adversely affects model convergence. To solve this issue, an uncertain example mining network is proposed for domain adaptive segmentation of remote sensing images. Specifically, a novel strategy called multilevel pseudo-label correcting (MPC) is proposed to correct the pseudo-labels in class, pixel, and superpixel levels. In this way, more reliable pseudo-labels can be selected for the subsequent training stage. Furthermore, a noise-robust example mining strategy, termed uncertainty-based valuable example mining (UVEM), is proposed to prioritize confident examples with significant gradients for training effectively. Extensive empirical evaluations on IsprsDA and LoveDA datasets demonstrate that the proposed method outperforms previous approaches, establishing state-of-the-art results in domain adaptive remote sensing image segmentation (RSIS). The code will be available athttps://github.com/StuLiu/UemDA. Wang Liu 0001, Puhong Duan, Zhuojun Xie, Xudong Kang, Shutao Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Classwise Prototype-Guided Alignment Network for Cross-Scene Hyperspectral Image ClassificationabstractIn the past few years, there has been significant progress in hyperspectral image classification (HSIC). However, when the trained classifier on the source scene is directly applied to a new scene, the classification performance tends to dramatically decrease because of the spectral shift phenomenon. Most existing techniques use feature alignment to learn knowledge from labeled scenes to unlabeled scenes, often overlooking the impact of noisy samples and outliers. To tackle this issue, the classwise prototype-guided alignment network (CPGAN) is proposed for cross-scene HSIC. The core idea is that classwise prototypes across scenes are employed as alignment intermediaries to guide cross-scene feature alignment. Specifically, first, spectral-spatial features from different scenes are extracted with a common feature extractor. Then, an uncertainty-aware pseudolabel selection (UPS) is designed to obtain high-confidence pseudolabels for unlabeled target scenes. Finally, a novel classwise prototype-guided alignment method is proposed to simultaneously achieve interdomain and intradomain alignment (IntraDA). The experimental results conducted on three datasets show that our method achieves superior performance compared to other cutting-edge classification algorithms. Zhuojun Xie, Puhong Duan, Xudong Kang, Wang Liu 0001, Shutao Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | Feature Consistency-Based Prototype Network for Open-Set Hyperspectral Image ClassificationabstractHyperspectral image (HSI) classification methods have made great progress in recent years. However, most of these methods are rooted in the closed-set assumption that the class distribution in the training and testing stages is consistent, which cannot handle the unknown class in open-world scenes. In this work, we propose a feature consistency-based prototype network (FCPN) for open-set HSI classification, which is composed of three steps. First, a three-layer convolutional network is designed to extract the discriminative features, where a contrastive clustering module is introduced to enhance the discrimination. Then, the extracted features are used to construct a scalable prototype set. Finally, a prototype-guided open-set module (POSM) is proposed to identify the known samples and unknown samples. Extensive experiments reveal that our method achieves remarkable classification performance over other state-of-the-art classification techniques. Zhuojun Xie, Puhong Duan, Wang Liu 0001, Xudong Kang, Xiaohui Wei 0001, Shutao Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2022 | Multilayer Global Spectral-Spatial Attention Network for Wetland Hyperspectral Image ClassificationabstractCoastal wetland monitoring plays an important role in the protection and restoration of ecosystems in this world. UAV-hyperspectral imaging, as an emerging technique for Earth observation and space exploration, provides the huge potential ability to identify different wetland species. In this work, a multilayer global spectral–spatial attention network (MGSSAN) is proposed for mapping coastal wetlands, which mainly consists of two major steps. First, a two-branch convolutional neural network (CNN) framework with residual connection is developed to obtain an initial classification probability map, in which one branch is used to capture the spectral information, the other branch is used to extract spatial information, and a global spectral–spatial attention module is designed to guide networks focusing on those features that are more discriminative. Second, an extended random walker method is utilized to optimize the initial classification probabilities, so as to yield the final map. Experiments performed on three wetland HSI datasets created by ourselves verify that the proposed method can obtain superior performance with respect to several state-of-the-art hyperspectral image classification methods. Zhuojun Xie, Jianwen Hu, Xudong Kang, Puhong Duan, Shutao Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |