VLDB 2026 Research / reviewers in the wild / expert
Yanhe Guo
dblp:171/0087
· DBLP profile ↗
18ranked-venue papers
2as first author
12since 2021 · last 2025
0009-0003-6144-5070ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 15 · 2 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | CDFNet: Cross-Domain Feature Fusion Network for PolSAR Terrain ClassificationabstractThe scarcity of labeled data and domain shift among polarimetric synthetic aperture radar (PolSAR) images degrades the performance of the supervised-learning-based algorithm. Some unsupervised domain adaptation (UDA) algorithms have been proposed to address this problem and achieve good performance. The existing UDA algorithms for PolSAR terrain classification focus on the feature distribution shift problem but ignore the label shift problem in UDA task. In addition, feature alignment-based algorithms generate pseudo labels for target domain which introduce label noise and compromising the UDA performance. To alleviate the problems above, we present a cross-domain feature fusion network (CDFNet) for PolSAR terrain classification. Specifically, a domain-balanced sampling (DBS) module is proposed to obtain a nearly balanced training dataset to alleviate the label shift problem. Then, a cross-domain feature fusion (CDF) module is presented to achieve class-wise feature alignment with no additional label noise introduction. Experimental results on four PolSAR datasets demonstrate that our algorithm outperforms state-of-the-art UDA algorithms in terms of target domain performance. Shuang Wang 0001, Zhuangzhuang Sun, Tianquan Bian, Yuwei Guo 0001, Linwei Dai, Yanhe Guo, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2024 | A Dual-Branch Random Mask Alignment Framework for Semi-Supervised PolSAR Terrain ClassificationabstractDespite the recent success of deep learning based polarimetric synthetic aperture radar(PolSAR) classification algorithms, it remains challenging in the scenario of limited labeled samples. Existing semi-supervised PolSAR terrain classification methods focus on the exploitation of pseudolabels, which are unreliable with limited labeled samples. To solve this problem, we propose a dual-branch random mask alignment framework for PolSAR terrain classification task. First, we propose a balanced regional expansion algorithm for labeled sample expansion. Then, to fully exploit the massive unlabeled samples, we designed a dual-branch network using two different polarization decomposition features as inputs, and a random mask alignment loss is employed to achieve consistency constraints on the unlabeled samples. Experimental results on two PolSAR datasets demonstrate that the proposed method achieve excellent performance with limited labeled samples. Tianquan Bian, Zhuangzhuang Sun, Shuang Wang 0001, Dou Quan, Yanhe Guo |
IGARSS | 6 |
| 2024 | ConDA: Continual Adaptation in Remote Sensing Via Visual Style PlaybackabstractThis study focuses on continual adaptation in remote sensing semantic segmentation, addressing challenges posed by frequent data updates and model forgetting. Remote sensing images exhibit variations in visual styles due to factors like location, time, and weather conditions, creating distinct domains. To counter performance degradation in new domains, we introduce a new challenge task in remote sensing, termed Continual Domain Adaptation (ConDA). Our innovative Visual Style Replay method employs Variational Auto-Encoder (VAE) and knowledge distillation, enabling the model to continuously learn from historical domains without forgetting. The proposed approach, tested on the INRIA dataset under ConDA settings, outperforms existing methods in combating catastrophic forgetting in remote sensing segmentation tasks. This contributes to adapting deep learning models to real-world scenarios with continually evolving remote sensing data. Ketao Zhong, Dong Zhao 0007, Shuang Wang 0001, Yanhe Guo |
IGARSS | 6 |
| 2024 | MfrNet: A New Multi-Scale Feature Refining Method for Remote Sensing Image Change CaptioningabstractRemote Sensing Image Change Captioning (RSICC) is an emerging multimodal field with promising prospects. This paper introduces a remote sensing image change caption model based on multi-scale and refined features. First, it extracts multi-scale features from dual-temporal images and then feeds them into the JointAtt and Dence Fusion (JADF) module for attention mutual guidance and feature refinement to eliminate noise. Next, the features are input into a transformer-based sentence generator for change statement generation. We conducted experiments on the Levir-CC dataset comparing our approach with existing methods, the results indicate that our MFRNet outperforms state-of-the-art methods in all metrics. Kaiqi Xu, Yingping Han, Rui Yang 0038, Xiutiao Ye, Yanhe Guo, Hantong Xing, Shuang Wang 0001 |
IGARSS | 5 |
| 2024 | Accurate and Lightweight Learning for Specific Domain Image-Text RetrievalabstractRecent advances in vision-language pre-trained models like CLIP have greatly enhanced general domain image-text retrieval performance. This success has led scholars to develop methods for applying CLIP to Specific Domain Image-Text Retrieval (SDITR) tasks such as Remote Sensing Image-Text Retrieval (RSITR) and Text-Image Person Re-identification (TIReID). However, these methods for SDITR often neglect two critical aspects: the enhancement of modal-level distribution consistency within the retrieval space and the reduction of CLIP's computational cost during inference. To address these issues, this paper presents a novel framework, Accurate and lightweight learning for specific domain Image-text Retrieval (AIR), based on the CLIP. AIR incorporates a Modal-Level distribution Consistency Enhancement regularization (MLCE) loss and a Self-Pruning Distillation Strategy (SPDS) to improve retrieval precision and computational efficiency. The MLCE loss harmonizes the sample distance distributions within image and text modalities, fostering a retrieval space closer to the ideal state. SPDS employs a strategic knowledge distillation process to transfer deep multimodal insights from CLIP to a shallower level, maintaining only the essential layers for inference, thus achieving model light-weighting. Comprehensive experiments across various datasets in RSITR and TIReID reveal that MLCE loss secures optimal retrieval, while SPDS achieves a favorable balance between accuracy and computational demand during testing. Rui Yang 0038, Shuang Wang 0001, Jianwei Tao, Yingping Han, Qiaoling Lin, Yanhe Guo, Biao Hou, Licheng Jiao |
ACM Multimedia | 6 |
| 2024 | Transcending Fusion: A Multiscale Alignment Method for Remote Sensing Image-Text RetrievalabstractRemote sensing image-text retrieval (RSITR) is pivotal for knowledge services and data mining in the remote sensing (RS) domain. Considering the multiscale representations in image content and text vocabulary can enable the models to learn richer representations and enhance retrieval. Current multiscale RSITR approaches typically align multiscale fused image features with text features but overlook aligning image-text pairs at distinct scales separately. This oversight restricts their ability to learn joint representations suitable for effective retrieval. We introduce a novel multiscale alignment (MSA) method to overcome this limitation. Our method comprises three key innovations: 1) a multiscale cross-modal alignment transformer (MSCMAT), which computes cross-attention between single-scale image features and localized text features, integrating global textual context to derive a matching score matrix within a mini-batch; 2) a multiscale cross-modal semantic alignment loss (MSCMA loss) that enforces semantic alignment across scales; and 3) a cross-scale multimodal semantic consistency loss (CSMMC loss) that uses the matching matrix from the largest scale to guide alignment at smaller scales. We evaluated our method across multiple datasets, demonstrating its efficacy with various visual backbones and establishing its superiority over existing state-of-the-art methods. The GitHub URL for our project ishttps://github.com/yr666666/MSA. Rui Yang 0038, Shuang Wang 0001, Yingping Han, Yuanheng Li, Dong Zhao 0007, Dou Quan, Yanhe Guo, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2024 | Select, Purify, and Exchange: A Multisource Unsupervised Domain Adaptation Method for Building ExtractionabstractAccurately extracting buildings from aerial images has essential research significance for timely understanding human intervention on the land. The distribution discrepancies between diversified unlabeled remote sensing images (changes in imaging sensor, location, and environment) and labeled historical images significantly degrade the generalization performance of deep learning algorithms. Unsupervised domain adaptation (UDA) algorithms have recently been proposed to eliminate the distribution discrepancies without re-annotating training data for new domains. Nevertheless, due to the limited information provided by a single-source domain, single-source UDA (SSUDA) is not an optimal choice when multitemporal and multiregion remote sensing images are available. We propose a multisource UDA (MSUDA) framework SPENet for building extraction, aiming at selecting, purifying, and exchanging information from multisource domains to better adapt the model to the target domain. Specifically, the framework effectively utilizes richer knowledge by extracting target-relevant information from multiple-source domains, purifying target domain information with low-level features of buildings, and exchanging target domain information in an interactive learning manner. Extensive experiments and ablation studies constructed on 12 city datasets prove the effectiveness of our method against existing state-of-the-art methods, e.g., our method achieves 59.1% intersection over union (IoU) on Austin and Kitsap → Potsdam, which surpasses the target domain supervised method by 2.2%. The code is available at https://github.com/QZangXDU/SPENet. Shuang Wang 0001, Qi Zang, Dong Zhao 0007, Chaowei Fang, Dou Quan, Yutong Wan, Yanhe Guo, Licheng Jiao |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2023 | Relational Image Patch Matching for Remote SensingabstractFeature descriptor-based methods have demonstrated remarkable performance in remote sensing image patch matching tasks and are usually optimized using contrastive loss and triplet loss. However, these optimization losses focus on calculating the distance between samples, ignoring the rich information of higher-order feature relationships between multiple image patches. The latter provides valuable information that can be used to improve task performance. Inspired by the superior performance of second-order relations in graph matching and clustering tasks, we aim to exploit the rich information available from high-order relations fully. This paper proposes a high-order relationship (HOR) learning method for remote sensing image patch matching. This method combines low-order feature relations between image patch pairs and high-order feature relations between multiple patches to enhance image matching performance. Extensive experimental results on a multimodel remote sensing image dataset, SEN 1-2, consisting of optical and SAR images, demonstrate that the proposed HOR learning method can improve the performance of remote sensing image patch matching. Xianwei Cao, Dou Quan, Chonghua Lv, Yanhe Guo, Shuang Wang 0001, Biao Hou, Licheng Jiao |
IGARSS | 4 |
| 2023 | Domain Distribution Alignment for Boosting Multi-Modal Remote Sensing Image MatchingabstractMulti-modal images can obtain complementary and rich information images, which are more widely used in various applications. However, due to the different imaging mechanisms of different sensors, there are significant domain distribution differences between multi-modal images. In multi-modal image matching, existing deep learning methods should deal with the image content difference caused by rotation transformation and the domain distribution difference caused by different sensors, which are very difficult for the deep network. To address this issue, we propose to combine an instance comparison and a batch comparison to deal with image content differences and domain distribution differences, respectively. We design a new domain distribution alignment method to explicitly constrain the sample domain distribution of the multi-modal images are consistent through the domain distribution alignment loss. Extensive multi-modal remote sensing image patch matching experiments have shown the effectiveness of the proposed method. Furthermore, the proposed multi-modal domain distribution alignment method has more obvious advantages when there are significant content differences and distribution differences. Dou Quan, Chonghua Lv, Yanhe Guo, Shuang Wang 0001, Yu Gu 0015, Licheng Jiao |
IGARSS | 4 |
| 2023 | A Texture and Saliency Enhanced Image Learning Method For Cross-Modal Remote Sensing Image-Text RetrievalabstractCross-modal remote sensing image-text retrieval (CMRSITR) can retrieve images of interest from a vast amount of remote sensing images and has received significant attention in recent years. However, existing methods do not consider saliency and texture information, which are essential for remote sensing images when extracting image features. Therefore, this paper proposes a novel texture and saliency enhanced image learning method for CMRSITR. We constructed a multi-task image feature extractor in this new method. A texture map and a saliency map are created by extracting texture and detecting the saliency of each RS image. Both maps are set as supervised information during training to make the extracted saliency and texture features gradually reconstructed to a saliency map and a texture map, respectively. At the same time, the retrieval features of each RS image are obtained from the retrieval feature branch of the image. Experiments conducted on two commonly used CMRSITR datasets, RSICD and UCM, showed that the proposed method is effective in improving retrieval performance and achieved state-of-the-art retrieval performance compared to existing methods. Rui Yang 0038, Yanhe Guo, Shuang Wang 0001 |
IGARSS | 3 |
| 2023 | Deep Continuous Matching Network for more Robust Multi-Modal Remote Sensing Image Patch MatchingabstractDue to the powerful feature extraction capabilities of deep neural networks, traditional approaches are gradually replaced by deep learning approaches for image matching tasks. For multi-modal image patch matching, the deep model should mainly learn the modality-invariant features. For multi-modal images with rotation transformation (RT), the deep model should learn the modality-invariant features and rotation-invariant features simultaneously. However, the performance of the latter trained model is degraded for the former task. The main reason is that the modality invariance of the features degenerates. This paper proposes a deep multi-modal remote sensing image matching network (DCMNet) that combines descriptor learning and continuous learning to solve this problem. Firstly, DCMNet is trained for learning modality-invariant features in multi-modal image patch matching. Then, DCMNet is optimized for multi-modal image patch matching with RT. In the later learning process, we reduce the change of important parameters for the modality-invariant features learning. Experiments demonstrate the effectiveness and robustness of DCMNet in alleviating the modal invariance degradation problem of features. Rufan Zhou, Dou Quan, Chonghua Lv, Yanhe Guo, Shuang Wang 0001, Yu Gu 0015, Licheng Jiao |
IGARSS | 4 |
| 2023 | Knowledge Decomposition and Replay: A Novel Cross-modal Image-Text Retrieval Continual Learning MethodabstractTo enable machines to mimic human cognitive abilities and alleviate the catastrophic forgetting problem in cross-modal image-text retrieval (CMITR), this paper proposes a novel continual learning method, Knowledge Decomposition and Replay (KDR), which emulates the process of knowledge decomposition and replay exhibited by humans in complex and changing environments. KDR has two components: a feature Decomposition-based CMITR Model (DCM) and a cross-task Generic Knowledge Replay strategy (GKR). DCM decomposes text and image features into task-specific and generic knowledge features, mimicking the human cognitive process of knowledge decomposition. Specifically, it employs a generic knowledge features extraction module for all tasks and a task-specific module for each task with a few trainable fully connected layers. Similarly, GKR emulates the human behavior of knowledge replay by utilizing the image-text similarity matrix output from the old task model with inputting the previous samples to induce the learning of the image-text similarity matrix output from the current task model with inputting the previous samples, using knowledge distillation technology. To demonstrate the effect of KDR, we adapted a continual learning dataset Seq-COCO from MSCOCO. Extensive experiments on Seq-COCO showed that KDR reduces catastrophic forgetting and consolidates general knowledge, improving the model's learning ability in CMITR. Rui Yang 0038, Shuang Wang 0001, Yanhe Guo, Xiutiao Ye, Biao Hou, Licheng Jiao |
ACM Multimedia | 5 |
| 2020 | Semi-Supervised PolSAR Image Classification Based on Improved Tri-Training With a Minimum Spanning TreeabstractIn this article, the terrain classifications of polarimetric synthetic aperture radar (PolSAR) images are studied. A novel semi-supervised method based on improved Tri-training combined with a neighborhood minimum spanning tree (NMST) is proposed. Several strategies are included in the method: 1) a high-dimensional vector of polarimetric features that are obtained from the coherency matrix and diverse target decompositions is constructed; 2) this vector is divided into three subvectors and each subvector consists of one-third of the polarimetric features, randomly selected. The three subvectors are used to separately train the three different base classifiers in the Tri-training algorithm to increase the diversity of classification; and 3) a help-training sample selection with the improved NMST that uses both the coherency matrix and the spatial information is adopted to select highly reliable unlabeled samples to increase the training sets. Thus, the proposed method can effectively take advantage of unlabeled samples to improve the classification. Experimental results show that with a small number of labeled samples, the proposed method achieves a much better performance than existing classification methods. Shuang Wang 0001, Yanhe Guo, Wenqiang Hua, Xinan Liu, Guoxin Song, Biao Hou, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2019 | Polsar Terrain Classification Based on Denoising-CNNabstractTerrain classification plays an important role in understanding Polari- metric Synthetic Aperture Radar (PolSAR) image intuitively. In the process of classification, feature extraction is critical. However, the preprocess of speckle noise filtering affects the effectiveness of the feature extractor which influences the accuracy of classification ultimately. Thus we integrate de-noising process and classification into an end-to-end framework based on CNN termed as Denoising-CNN, which improves the accuracy of classification. Experiments on real PolSAR data show that our proposed method offers an excellent performance. Yanhe Guo, Shuang Wang 0001, Guoxin Song, Wenqiang Hua, Feihang Liu |
IGARSS | 1 |
| 2019 | Dual-Channel Convolutional Neural Network for Polarimetric SAR Images ClassificationabstractThis paper presents a new dual-channel convolutional neural network (Dc-CNN) for Polarimetric synthetic aperture radar (PolSAR) image classification when labeled samples are small. First, a neighborhood minimum spanning tree (MST) is used to enlarge the labeled sample set. Then, in order to obtain the abundant spatial information, a new dual-channel CNN is designed to PolSAR image to acquire different spatial features. This network model contains two parallel CNN structures, which can extract different features used two multiscale convolution structure. Experiments results show that compared with other methods, the proposed method shows a satisfactory classification result. Wenqiang Hua, Shuang Wang 0001, Yanhe Guo, Xiaomin Jin |
IGARSS | 4 |
| 2018 | Fully Convolutional Semi-Supervised Gan for Polsar ClassificationabstractWe propose a novel semi -supervised fully convolutional network for Polarimetric synthetic aperture radar (PoISAR) terrain classification. First, by designing a fully convolutional structure, we can perform pixel-based classification tasks. Then, by applying semi -supervised generative adversarial networks (GANs), we utilize both labeled and unlabeled samples and aim to obtain higher classification accuracy. Through a mini-max two-player game, GAN has better performance than other “single-player” classifiers. Finally, we combine the fully convolutional structure with the semi-supervised GAN. Our fully convolutional semi-supervised GAN (FC-SGAN) has excellent spatial feature learning ability and can perform end-to-end pixel-based classification tasks. Experimental results show that compared with existing works, the proposed method has better performances. Even when the training set gets smaller, our method keeps high accuracy. Mengchen Liu, Shuang Wang 0001, Yanhe Guo, Biao Hou, Licheng Jiao, Xiaojin Hou |
IGARSS | 4 |
| 2017 | Semi-supervised PolSAR Classification Based on Improved Tri-trainingabstractIn this paper, we proposed a new semi-supervised method for polarimetric synthetic aperture radar (PolSAR) terrain classification based on improved tri-training. This method only needs a few numbers of labeled samples to achieve the results obtained by traditional supervised classification methods. First, it uses a variety of target decomposition methods to obtain high-dimensional feature. Second, a new feature selection method based on the ratio of between-class scatter and within-class scatter is proposed to reduce the redundant feature. Finally, an improved tri-training method is executed. A real PolSAR data is used to verify the proposed method. Experimental results show that the proposed method is efficient with a few labeled samples and effectively improve the classification accuracy compared with other traditional classification methods. Wenqiang Hua, Shuang Wang 0001, Bo Yue, Yanhe Guo |
IGARSS | 5 |
| 2015 | Wishart RBM based DBN for polarimetric synthetic radar data classificationabstractDeep Belief Network (DBN) is a classic deep learning model, and it can learn higher feature and do better classification job. We combine DBN's basic component Restricted Boltzmann Machines (RBM) with the statistic distribution of Polarimetric SAR (PolSAR) data. Based on it, we develop a deep learning classification method that is suitable for PolSAR data. To verify the effectiveness of the method, a real PolSAR dataset is tested. Experiment result confirms that the proposed method provides fine improvements both in classification accuracy and visual effect. Yanhe Guo, Shuang Wang 0001, Chenqiong Gao, Danrong Shi, Biao Hou |
IGARSS | 1 |