EDBT 2026 Demo / reviewers in the wild / expert
Xiao Han 0012
dblp:01/2095-12
· DBLP profile ↗
17ranked-venue papers
0as first author
17since 2021 · last 2026
0000-0002-1953-8658ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 15 · 15 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Tiny object detection based on dynamic scale-awareness label assignment and contextual enhancement
Tianyang Zhang 0002, Xiangrong Zhang, Chaozhuo Hua, Guanchun Wang, Xiao Han 0012, Licheng Jiao |
Pattern Recognit. | 5 |
| 2026 | Dual-Net: Dual Visual Spectral Affinity Monitoring Network for Hyperspectral Anomaly Detection
Xiangrong Zhang, Rongxia Qiu, Shiqi Wu, Guanchun Wang, Xiao Han 0012, Yifei Jiang, Licheng Jiao |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Source-Free Cross-Domain Scene Classification of Remote Sensing Images via Statistics Matching and Noise AdaptationabstractIn recent years, in order to alleviate the performance degradation problem caused by domain drift in scene classification tasks, some unsupervised domain adaptation methods have been introduced into the field of remote sensing images. Such methods require simultaneous access to both source and target domain data during training. However, the large storage and transmission costs of remote sensing images limit the further application of these methods. To address this challenge, we investigate the task of source-free cross-domain scene classification for remote sensing images. In the model adaptation process, only the target domain dataset and the trained source domain model are used. Our approach consists of two parts: a distribution alignment strategy based on source domain model statistics matching and a noise adaptation strategy. In order to fully utilize the knowledge of the pre-trained source domain model, we fix the classifier to get the feature distribution of the source domain, so that the target domain feature distribution is close to the feature distribution of the source domain. The noise adaptation layer is inserted after the classifier in order to improve the robustness of the model to noise-containing pseudo-labels, and the sample-wise noise transfer matrix is learned. Experimental results on 12 transfer tasks on the cross-scene dataset, and 2 transfer tasks on the cross-sensor dataset, to validate the effectiveness of our approach. Compared to traditional unsupervised domain adaptive methods, our method is able to achieve better performance under the condition of not accessing the source domain data. Peng Zhu 0004, Xiangrong Zhang, Xiao Han 0012, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | Beyond Single Pixel: Context Priors Guided Semi-Supervised Building Footprint SegmentationabstractAutomated building footprint segmentation is crucial in remote sensing with widespread applications in various fields. Semi-supervised semantic segmentation methods are gaining traction in the remote sensing community as they significantly reduce the need for labor-intensive pixel-level annotations in training segmentation models. However, these methods typically rely on individual pixel-level supervision for unlabeled data, neglecting the contextual relationships between pixels. This limits their potential to exploit unlabeled data. To bridge this gap, this paper proposes a novel Context Priors Guided Semi-Supervised Building Footprint Segmentation method that leverages contextual relationships among numerous unlabeled pixels to build supervisory signals that extend beyond individual pixel-level guidance for learning on unlabeled data. The CPG comprises two main components: Spatial context priors-guided pseudo-label regularization and semantic context priors-guided representation learning. By integrating contextual knowledge from both pixel spatial locations and semantic representation spaces, these components capture comprehensive class semantic attributes and enable the model to learn complete shapes of building footprints. Our approach achieves state-of-the-art performance on three publicly available building footprint segmentation datasets, validating its effectiveness. Xiaoqian Zhu, Xiangrong Zhang, Tianyang Zhang 0002, Xiao Han 0012, Licheng Jiao, Lianchao Zhang |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2024 | ESDINet: Efficient Shallow-Deep Interaction Network for Semantic Segmentation of High-Resolution Aerial ImagesabstractSemantic segmentation of high-resolution remote sensing images is essential in many fields. Nevertheless, in practical applications, constrained by limited computational resources and complex network structures, many advanced models on semantic segmentation often fail to show efficient performance, prompting research on lightweight models. For lightweight semantic segmentation models, the two-branch architecture has been shown to work well in speed and performance. However, such two-branch architectures usually do not utilize enough information for shallow structures to efficiently provide richer multiscale information for the two branches. The lightweight modules it uses are difficult to extract the global context information of the features effectively. Compared with the current advanced semantic segmentation models, lightweight models still have some differences in performance. In order to solve these problems, we propose a new lightweight dual-branch architecture efficient shallow-deep interaction network (ESDINet), which can quickly extract low-level spatial and high-level semantic information of images through the detail branch and semantic branch. Specifically, we have constructed an efficient double-branch structure with shallow and deep different interactions to achieve multiscale information interaction. At the same time, we optimize the semantic branch and propose a new linear attention block to effectively improve the global perception of the semantic branch. We performed extensive experiments and the results show that our model achieves a good balance between segmentation accuracy and inference speed. In particular, ESDINet achieves 82.03% mean intersection over union (mIoU) on the Vaihingen test set, while the proposed model achieves an inference speed of 116 frames/s (FPS) for$512\times512$inputs on a single NVIDIA GTX 2080Ti GPU. Xiangrong Zhang, Zhenhang Weng, Peng Zhu 0004, Xiao Han 0012, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Multistage Enhancement Network for Tiny Object Detection in Remote Sensing ImagesabstractWith the rapid advances in deep learning techniques, remote sensing object detection has achieved remarkable achievements in recent years. However, tiny object detection remains unsatisfactory and suffers from two main drawbacks, including (1) the high sensitivity of IoU for location deviation in tiny objects and (2) the poor-quality feature representations of tiny objects. To address the aforementioned problems, we propose a Multi-stage Enhancement Network (MENet) that achieves the instance-level and feature-level enhancement of tiny objects from different stages of the detector. Since the IoU-based label assignment drastically deteriorates the positive samples for tiny objects, we first propose a Central Region-based (CR) label assignment to substitute it in the Region Proposal Network (RPN). The CR label assignment regards the anchors that fall into the central region of ground-truth boxes as positive samples, which provides more positive samples for tiny objects. Then, we design a Gated Context Aggregation (GCA) module that selectively aggregates valuable context information to enhance the feature representation of tiny objects. Additionally, we devise a positive RoI feature (pRoI) generator in the Region Convolutional Neural Network (R-CNN) to generate a rich diversity of high-quality positive RoI features for tiny objects. We conduct extensive experiments on AI-TOD and SODA-A datasets, and the results demonstrate the effectiveness of our proposed method. Tianyang Zhang 0002, Xiangrong Zhang, Xiaoqian Zhu, Guanchun Wang, Xiao Han 0012, Xu Tang 0004, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | High-Resolution Remote Sensing Image Segmentation With Global-Guided Normalization and Local Affinity DistillationabstractIn recent years, high-resolution (HR) remote sensing images (RSIs) segmentation has received growing attention. The huge number of pixels poses a challenge to the semantic segmentation algorithm, which is limited by the storage of GPUs, so the current methods for processing HR RSIs are categorized into two main categories, i.e., global methods and local methods. The former downsamples the original image and loses a lot of feature details. The latter crops the original image and fails to obtain global contextual information. Both types of methods lead to limited segmentation accuracy. In this article, we propose an end-to-end framework, called global injection network (GINet), which explores two levels of feature distribution and feature relationship to achieve tradeoff between global context and local details. In concrete terms, we propose the global-guided normalization (GGN) module, which injects global context information into local branch and modulates local features using global features to enhance the global perception of local branch. In addition, to constrain the spatial consistency of two branches, inspired by the knowledge distillation technique, we propose local affinity distillation (LAD) loss, which distills the relations in local features into global features to keep the similarity of the relationships corresponding to patches in the two branches. The comprehensive experimental results on three large-scale land-cover classification datasets, DeepGlobe ($2448 \times 2448$), Inria Aerial ($5000 \times 5000$), and GID-15 ($7200 \times 6800$), confirm the effectiveness and superiority of our method in HR semantic segmentation tasks. Peng Zhu 0004, Xiangrong Zhang, Xiao Han 0012, Puhua Chen, Xu Tang 0004, Xina Cheng, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | CTACL:Hyperspectral Image Change Detection Based on Adaptive Contrastive LearningabstractHyperspectral image change detection (HSI-CD) can accurately identify changing regions by capturing subtle spectral differences and has become a research hotspot in the field of remote sensing (RS). Convolutional neural networks (CNNs) have excellent local context modeling capabilities and have been proven to be powerful feature extractors in HSI-CD. However, due to its inherent network structure limitation, CNN cannot well mine and represent the sequential properties of spectral features, especially the medium and long-term dependencies. In contrast, transformer-based network architecture shows a strong ability to model long-distance dependencies, which can fully mine and extract global features, but exhibits weak performance in extracting local information. To this end, we propose HSI-CD network based on adaptive contrastive learning (CTACL). Specifically, we first propose a parallel network of CNNs and transformers to mine local and global temporal-spatial-spectral features of HSI, respectively. Second, we propose adaptive contrastive learning to pre-train the network to learn the latent features of a large amount of unlabeled data and better mine and utilize local and global information. Experimental results on the farmland dataset show that the proposed method performs well. Shunli Tian, Xiangrong Zhang, Guanchun Wang, Xiao Han 0012, Puhua Chen, Xina Cheng |
IGARSS | 4 |
| 2023 | Spectral-Spatial Distribution Consistent Network Based on Meta-Learning for Cross-Domain Hyperspectral Image ClassificationabstractCross-domain networks can solve the problem of insufficient labeled samples, especially for hyperspectral images (HSIs) where obtaining labeled samples is time-consuming and laborious. Most of the current methods rely on the spatial information to achieve domain alignment, without considering the rich spectral information of HSIs. Furthermore, the methods based on convolutional neural network (CNN) cannot get the spatial information of irregular image regions, resulting in poor classification results of object edges. Therefore, we design a spectral-spatial distribution consistent network (SSDC) based on meta-learning. Firstly, to improve the feature extraction ability of the cross-domain classification model, we introduce a feature pre-extraction module, which uses the spectral attention mechanism and the alternating meta-learning method to obtain the general features of the source domain and the discriminative features of the target domain, so as to obtain the spectral weight matrix for subsequent processing. Secondly, we propose a spectral consistent module based on singular value decomposition, which increases the difference between different classes of features by penalizing the singular values of the feature matrix to achieve data distribution alignment in the spectral dimension. Finally, aiming at the low classification accuracy of irregular image regions, we propose a spatial consistent module to obtain non-local spatial topological information through stacked cross modules and graph sample and aggregate networks, which can reduce domain shift. The experiments of SSDC on four classical HSI datasets show that the proposed method can obtain competitive results with other methods based on CNN and cross-domain. Xiangrong Zhang, Qi Zhen, Xiao Han 0012, Puhua Chen, Xu Tang 0004, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Resformer: Bridging Residual Network and Transformer for Remote Sensing Scene ClassificationabstractRemote sensing (RS) scene classification is a crucial research topic in the RS community, and many convolutional neural networks (CNNs)-based methods have been proposed to improve classification performance. Due to the intrinsic locality of convolution operations, CNNs are good at extracting local information but are not easy to capture global contextual information which is also important to fully interpret RS scenes. Recently, transformer has shown the potential for learning global contextual information, but it pays less attention to local information. In this paper, we propose a new interactive dual-branch network for RS scene classification, named Resformer, which can use CNNs's efficiency in extracting local information as well as transformer's power in capturing global information. Besides, we propose a two-way feature interaction module (TFIM), which can not only efficiently fuse CNNs-based local features with transformer-based global fetures, but also extract multi-scale information from RS scenes. Finally, we use a class score fusion strategy to integrate the features extracted from the two branches. Encouraging experimental results counted on two public RS scene data sets demonstrate that our Resformer is effective in RS scene classification task. Mingteng Li, Jingjing Ma 0001, Xu Tang 0004, Xiao Han 0012, Licheng Jiao |
IGARSS | 4 |
| 2022 | NQ-Protonet: Noisy Query Prototypical Network for Few-Shot Remote Sensing Scene ClassificationabstractFew-shot remote sensing scene classification, which aims to recognize unseen classes given only a few labeled samples, is a challenge task due to the complex content contained in remote sensing scenes. In this end, we propose a noisy query prototypical network (NQ-ProtoNet), which uses query-mix module (QM) to produce extra query samples with inter-ference information for classification and thus implicitly enhance the feature learning ability of model. Our method alleviates the problem of large intraclass variances and inter-class similarity of remote sensing scenes to some extent, and the positive experimental results on UC Merced and NWPU data sets show that it outperforms several few-shot learning methods. Weiquan Lin, Xu Tang 0004, Xiao Han 0012, Jingjing Ma 0001, Xiangrong Zhang, Licheng Jiao |
IGARSS | 3 |
| 2022 | Multi-Scale Interactive Transformer for Remote Sensing Cross-Modal Image-Text RetrievalabstractCross-modal Remote sensing (RS) image-text retrieval (CMR-SITR) plays a crucial role in the RS community. A common way for CMRSITR is to extract texts and RS images' feature representations separately and then measure their similarities in the specific or common feature space. Recently, along with the booming of deep convolutional neural networks (DCNNs), these kinds of methods are vivid and achieve successes in their own applications. However, they neglect the inherent relationships between different features, and they are always heavy. To overcome the limitations mentioned above, we propose a new model for CMRSITR in this paper, named multi-scale interactive transformer (MSIT). MSIT first adopts simple feature learning models for texts and RS images which could ensure the whole model is not heavy. Then, MSIT introduces transformer encoders to enhance features' usefulness by considering the potential relations between different representations. Also, a lightweight multi-scale feature learning module is proposed to mine more plentiful contents from RS images. Finally, instead of outputting the features, MSIT produces matching scores for texts and RS images, which can be used to decide the retrieval results directly. The experimental results on two RS datasets indicate our modal is effective for CMRSITR. Yijing Wang 0004, Jingjing Ma 0001, Mingteng Li, Xu Tang 0004, Xiao Han 0012, Licheng Jiao |
IGARSS | 5 |
| 2021 | Cross-Source Image Retrieval Based on Ensemble Learning and Knowledge Distillation for Remote Sensing ImagesabstractAs different kinds of high-resolution remote sensing (HRRS) image data sources increase, the cross-source content-based image retrieval (CS-CBRSIR) is becoming an important and urgent task to be solved. Most existing methods focus on optimizing the common space features for dual-source effectively. The source discrepancy in classifier level, however, has been ignored. To handle this problem, we propose teacher-ensemble learning with the knowledge distillation method in this paper. The ensemble of source-shared and source-specific classifiers could construct an effective teacher model. The useful information can be transferred back with the knowledge distillation. Besides, the feature pyramid network is introduced to learn the multi-scale features from HRRS images, which can describe the complex contents of HRRS images well. The positive experimental results conducted on DSRSID illustrates the effectiveness of the proposed method. Jingjing Ma 0001, Duanpeng Shi, Xu Tang 0004, Xiangrong Zhang, Xiao Han 0012, Licheng Jiao |
IGARSS | 5 |
| 2021 | Multi-Scale Meta-Learning-Based Networks for High-Resolution Remote Sensing Scene ClassificationabstractHigh-resolution remote sensing (HRRS) image scene classification based on limited data set is challenging in practical application. Although convolutional neural networks have shown powerful feature representation capability, they cannot perform well in the absence of rich label information in general. This paper proposes a multi-scale meta-learning-based (MSML) model to complete the HRRS scene classification with a little labeled data. First, we develop a multi-scale feature learning strategy to explore the rich information from HRRS scenes. Then, to use small data to train our network, we formulate the meta-learning as a regularization term and embed it into the classification loss function. By optimizing the proposed loss function, we can obtain a robust and generalized scene classification model. The positive experimental results counted on a public HRRS scene data set show that our MSML model is useful in HRRS scene classification tasks. Xu Tang 0004, Weiquan Lin, Chao Liu 0042, Xiao Han 0012, Jingjing Ma 0001, Licheng Jiao |
IGARSS | 4 |
| 2021 | Hyperspectral Image Classification Based on Spectral Graph and Bidirectional LSTM NetworkabstractConvolutional neural networks (CNNs) have achieved cracking performance in the hyperspectral image (HSI) classification task. Nevertheless, most of them cannot meet what we expect when the numbers of labeled samples are small. Also, due to the rectangular convolution kernels, the long-range context information within HSIs is cannot fully be explored. To solve these problems, we propose a semi-supervised method based on the graph convolutional network (GCN) and bidirectional Long Short-Term Memory (Bi-LSTM). First, HSIs are over segmented into various superpixels and GCN is employed for mining the advanced spectral features. Second, we input the obtained spectral features to the Bi-LSTM model for exploring global spatial features. Due to the diverse receptive fields, the short-and long-range spatial relations can be discovered simultaneously. Finally, we map the features from region-level to pixel-level for classifying HSIs. The positive experimental results counted on two HSIs demonstrate that our method is superior to some popular methods. Xu Tang 0004, Qionglin Zhou, Xiao Han 0012, Dalei Li, Xiangrong Zhang, Licheng Jiao |
IGARSS | 4 |
| 2021 | Remote Scene Image Scene Classification Based on Adaptive Segmentation and Dynamic Graph ConvolutionabstractAs an important research topic in the remote sensing (RS) community, RS image scene classification is a challenging task due to the complex contents of RS images. In general, RS image scene classification is a single-label problem. Nevertheless, it is known that the contents within RS are huge in volume and diverse in type. Only a single semantic label cannot describe an RS scene completely, especially when the resolution of RS images is increased recently. The various semantics hidden in the high-resolution RS (HRRS) images are also important to the scene classification task. Taking the issues mentioned above into account, we develop a new scene classifier named graph scene classifier (GSCer) for HRRS images with the help of the deep convolution neural network (DCNN) and dynamic graph convolution (DGCN). Not only the global semantic but also the diverse hidden local semantics within an HRRS image can be fully explored. The encouraging experimental results counted on two public HRRS data sets demonstrate that our GSCer is effective in HRRS scene classification tasks. Yuqun Yang, Xu Tang 0004, Xiao Han 0012, Jingjing Ma 0001, Xiangrong Zhang, Licheng Jiao |
IGARSS | 3 |
| 2021 | High-Resolution Remote Sensing Images Change Detection with Siamese Holistically-Guided FCNabstractChange Detection is an important and challenging task in the remote sensing (RS) field, especially when the resolution of RS images is getting higher. The appearance of the deep convolutional neural network (DCNN) provides new opportunities for the high-resolution RS (HRRS) image processing as well as the HRRS CD task. In this paper, we proposed a Siamese holistically-guided FCN (SHG-FCN) model to fully mine the low- and high-level features from HRRS images for completing the CD task. SHG-FCN consists of a Siamese encoder and a holistically-guided decoder. The Siamese encoder adopts five layers of convolution for feature extraction and generates multi-scale difference maps. The decoder employs the holistically-guided architecture, which uses the deep semantic feature as codewords to guide the up-sample of feature map and achieve multi-scale feature fusion. Our model is testified on two public HRRS datasets, and the obtained encouraging CD results illustrate that our method is effective in HRRS CD tasks. Xu Tang 0004, Xiao Han 0012, Jingjing Ma 0001, Xiangrong Zhang, Licheng Jiao |
IGARSS | 3 |