Chao Tao 0001

dblp:08/3778-1 · DBLP profile ↗
← Back
37ranked-venue papers
13as first author
22since 2021 · last 2026
0000-0003-0071-310XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 33 · 13 first-author · 18 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021
YearPublicationVenuePosition
2026 A Gift From the Integration of Discriminative and Diffusion-Based Generative Learning: Boundary Refinement Remote Sensing Semantic Segmentation
abstract
Remote sensing semantic segmentation must address both what the ground objects are within an image and where they are located. Consequently, segmentation models must ensure not only the semantic correctness of large-scale patches (low-frequency information) but also the precise localization of boundaries between patches (high-frequency information related to boundary components). However, most existing approaches rely heavily on discriminative learning, which excels at capturing low-frequency features, while overlooking its inherent limitations in learning high-frequency features for semantic segmentation. Recent studies have revealed that diffusion generative models excel at generating high-frequency details. Our theoretical analysis confirms that the diffusion denoising process significantly enhances the model's ability to learn high-frequency features; however, we also observe that these models exhibit insufficient semantic inference for low-frequency features when guided solely by the original image. Therefore, we integrate the strengths of both discriminative and generative learning, proposing the Integration of Discriminative and diffusion-based Generative learning for Boundary Refinement (IDGBR) framework. The framework first generates a coarse segmentation map using a discriminative backbone model. This map and the original image are fed into a conditioning guidance network to jointly learn a guidance representation subsequently leveraged by an iterative denoising diffusion process refining the coarse segmentation. Extensive experiments across five remote sensing semantic segmentation datasets (binary and multi-class segmentation) confirm our framework's capability of consistent boundary refinement for coarse results from diverse discriminative architectures. The source code is available at https://github.com/KeyanHu-git/IDGBR.
Hao Wang 0069, Keyan Hu, Haifeng Li 0007, Chao Tao 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2026 Multi-modality transfer learning for cloudy remote sensing images: Addressing modality imbalance with knowledge distillation
Yuze Wang 0005, Haifeng Li 0007, Mariana Belgiu, Chao Tao 0001
Pattern Recognit.5
2025 IFShip: Interpretable fine-grained ship classification with domain knowledge-enhanced vision-language models
Mingning Guo, Mengwei Wu, Yuxiang Shen, Haifeng Li 0007, Chao Tao 0001
Pattern Recognit.5
2025 HASNet: A Foreground Association-Driven Siamese Network With Hard Sample Optimization for Remote Sensing Image Change Detection
abstract
Remote sensing change detection (RS-CD) relies on the model’s ability to learn features of marked change objects, known as foreground targets. Beyond foreground targets, the background targets are more valuable samples for change detection, such as unlabeled ones, semantically ambiguous ones, pseudo-changes, and non-interesting changes, referred to as hard case samples (HCSs) in this article. There are two additional challenges to learning HCSs: 1) the loss function focusing on the foreground targets with rich labels and ignoring the HCSs in the background, called the imbalance problem and 2) it is difficult for a model to learn the change information of HCSs directly, which is called HCSs missingness. This article proposed a foreground association-driven Siamese network with hard sample optimization (HASNet). To deal with the imbalance problem, we propose an equilibrium optimization loss (EO-loss) function to regulate the optimization focus of the foreground and background, determine the HCSs through the distribution of the loss values, and introduce dynamic weights in the loss term to gradually shift the optimization focus of the loss from the foreground to the background hard cases as the training progresses. To address the HCSs missingness, we propose the scene-foreground association module by using potential remote sensing spatial scene information to model the association between the target of interest in the foreground and the related context to obtain scene embedding to reinforce the feature of hard cases. Experiments on four public datasets with 11 baselines show that HASNet outperforms current state-of-the-art CD methods, particularly in detecting HCSs. The source code is available athttps://github.com/GeoX-Lab/HASNet.
Chao Tao 0001, Dongsheng Kuang, Zhenyang Huang, Chengli Peng, Haifeng Li 0007
IEEE Trans. Geosci. Remote. Sens.1
2025 Weakly Supervised Framework Considering Multi-Temporal Information for Large-Scale Cropland Mapping With Satellite Imagery
abstract
Accurately mapping large-scale cropland is crucial for agricultural production management and planning. Currently, the combination of remote sensing data and deep learning techniques has shown outstanding performance in cropland mapping. However, those approaches require massive precise labels, which are labor-intensive. To reduce the label cost, this study presented a weakly supervised framework considering multi-temporal information for large-scale cropland mapping. Specifically, we extract high-quality labels according to their consistency among global land cover (GLC) products to construct the supervised learning signal. On the one hand, to alleviate the over-fitting problem caused by the model’s over-trust of remaining errors in high-quality labels, we encode the similarity/ aggregation of cropland in the visual/spatial domain to construct the unsupervised learning signal, and take it as the regularization term to constrain the supervised part. On the other hand, to sufficiently leverage the plentiful information in the samples without high-quality labels, we also incorporate the unsupervised learning signal in these samples, enriching the diversity of the feature space. After that, to capture the phenological features of croplands, we introduce dense satellite image time series (SITS) to extend the proposed framework in the temporal dimension. We also visualized the high-dimensional phenological features to uncover how multi-temporal information benefits cropland extraction, and assessed the method’s robustness under conditions of data scarcity. The proposed framework has been experimentally validated for strong adaptability across three study areas (Hunan Province, Southeast France, and Kansas) in large-scale cropland mapping, and the internal mechanism and temporal generalizability are also investigated. The source codes are available at https://github.com/wangyuze-csu/WSFCMI.
Yuze Wang 0005, Aoran Hu, Ji Qi 0001, Chao Tao 0001
IEEE Trans. Geosci. Remote. Sens.5
2025 FSVLM: A Vision-Language Model for Remote Sensing Farmland Segmentation
abstract
Existing visual deep learning paradigms, which are based on labels, struggle to capture the intricate interrelationships between farmland and its surrounding environment and fail to account for temporal variations associated with the phenological cycle. These limitations lead to omissions and confusion in the recognition process, greatly impacting the accuracy and efficiency of farmland recognition. Language can accurately depict the spatial attributes of farmland, profoundly express the unique phenological landscapes of farmland that change with seasons and growth stages, and express the intricate interactions between these changes and environmental factors. This capability can address the deficiencies of label-based visual deep learning in understanding the complex features of farmland. This study explored, for the first time, the application of language-guided vision-language models (VLMs) for farmland segmentation. First, as current VLM research lacks a dedicated farmland image text (FIT)pair dataset, this study constructed an FIT dataset in two steps. Step 1, designed a semi-automatic text description annotation framework for farmland images based on 12 key factors influencing farmland segmentation. Step 2, used the framework to construct the FIT dataset. Then, a VLM for farmland segmentation (FSVLM) was designed by combining a semantic segmentation model with a multimodal large language model (LLM). Comparative experiments demonstrated that the proposed method outperforms existing farmland segmentation methods in both generalization and segmentation accuracy. In addition, a series of ablation experiments were conducted to examine the impacts of language descriptions of different semantic levels on the model’s farmland information extraction performance.
Zhuofei Du, Dandan Zhong, Yuze Wang 0005, Chao Tao 0001
IEEE Trans. Geosci. Remote. Sens.5
2024 Global-Local Coupled Style Transfer for Semantic Segmentation of Bitemporal Remote Sensing Images
abstract
Due to the different acquisition conditions, large variations in the feature distributions of two temporal domains generally exist, known as temporal domain shift. The temporal domain shift is primarily influenced by coupled dual-factor: global style variations (such as illumination and weather conditions) and local style variations (such as the inherent phenological properties of land cover classes). In this article, we first formulate the temporal domain shift problem as an issue of dual-factor coupled interference on feature distributions in remote sensing (RS) community. To address this issue, we propose a semantic-guided style transfer (SGST) framework seamlessly integrating global feature alignment with local feature semantic matching. We use an adaptive segmentation model to provide pseudosegmentation maps and feed them into the style transfer model as semantic guidance. Under semantic guidance, a semantic-constrained style normalization (SCSN) module is designed to achieve style transfer at both global and local levels. Furthermore, a dual learning approach is introduced to make the style transfer model and the adaptive segmentation model promote each other. As a result, the style transfer model generates high-quality style-transferred images, and the adaptive segmentation model progressively predicts more accurate pseudosegmentation maps. Extensive experiments demonstrate the superiority of our proposed framework over state-of-the-art methods in terms of both perceptual quality and quantitative performance.
Hao Wang 0069, Mingning Guo, Shaoxian Li, Haifeng Li 0007, Chao Tao 0001
IEEE Trans. Geosci. Remote. Sens.5
2023 MSINet: Mining scale information from digital surface models for semantic segmentation of aerial images
Chengli Peng, Haifeng Li 0007, Chao Tao 0001, Yansheng Li 0001, Jiayi Ma 0001
Pattern Recognit.3
2023 Self-Supervised Remote Sensing Feature Learning: Learning Paradigms, Challenges, and Future Works
abstract
Deep learning has achieved great success in learning features from massive remote sensing images (RSIs). To better understand the connection between three feature learning paradigms, which are unsupervised feature learning (USFL), supervised feature learning (SFL), and self-supervised feature learning (SSFL), this paper analyzes and compares them from the perspective of feature learning signals, and gives a unified feature learning framework. Under this unified framework, we analyze the advantages of SSFL over the other two learning paradigms in RSI understanding tasks and give a comprehensive review of existing SSFL works in RS, including the pre-training dataset, self-supervised feature learning signals, and the evaluation methods. We further analyze the effects of SSFL signals and pre-training data on the learned features to provide insights into RSI feature learning. Finally, we briefly discuss some open problems and possible research directions.
Chao Tao 0001, Ji Qi 0001, Mingning Guo, Qing Zhu 0012, Haifeng Li 0007
IEEE Trans. Geosci. Remote. Sens.1
2023 GraSS: Contrastive Learning With Gradient-Guided Sampling Strategy for Remote Sensing Image Semantic Segmentation
abstract
Self-supervised contrastive learning (SSCL) has achieved significant milestones in remote sensing image (RSI) understanding. Its essence lies in designing an unsupervised instance discrimination pretext task to extract image features from a large number of unlabeled images that are beneficial for downstream tasks. However, existing instance discrimination based SSCL suffers from two limitations when applied to the RSI semantic segmentation task: 1) Positive sample confounding issue, SSCL treats different augmentations of the same RSI as positive samples, but the richness, complexity, and imbalance of RSI ground objects lead to the model actually pulling a variety of different ground objects closer while pulling positive samples closer, which confuse the feature of different ground objects. 2) Feature adaptation bias, SSCL treats RSI patches containing various ground objects as individual instances for discrimination and obtains instance-level features, which are not fully adapted to pixel-level or object-level semantic segmentation tasks. To address the above limitations, we consider constructing samples containing single ground objects to alleviate positive sample confounding issue, and make the model obtain object-level features from the contrastive between single ground objects. Meanwhile, we observed that the discrimination information can be mapped to specific regions in RSI through the gradient of unsupervised contrastive loss, these specific regions tend to contain single ground objects. Based on this, we propose contrastive learning with Gradient guided Sampling Strategy (GraSS) for RSI semantic segmentation. GraSS consists of two stages: 1) the instance discrimination warm-up stage to provide initial discrimination information to the contrastive loss gradients, 2) the gradient guided sampling contrastive training stage to adaptively construct samples containing more singular ground objects using the discrimination information. Experimental results on three open datasets demonstrate that GraSS effectively enhances the performance of SSCL in high-resolution RSI semantic segmentation. Compared to eight baseline methods from six different types of SSCL, GraSS achieves an average improvement of 1.57% and a maximum improvement of 3.58% in terms of mean intersection over the union. Additionally, we discovered that the unsupervised contrastive loss gradients contain rich feature information, which inspires us to utilize gradient information more extensively during model training to attain additional model capacity. The source code is available at https://github.com/GeoX-Lab/GraSS.
Chao Tao 0001, Yunsheng Zhang 0001, Chengli Peng, Haifeng Li 0007
IEEE Trans. Geosci. Remote. Sens.3
2022 MDANet: Unsupervised, Mixed-Domain Adaptation for Semantic Segmentation of Remote Sensing Images
abstract
The imaging process of optical remote sensing images are easily affected by external conditions. Therefore, remote sensing images under different imaging conditions often show color differences, resulting in feature distribution differences between the source and target domain, hindering the migration of semantic segmentation models between domains. Currently, most domain adaptation methods are for single-source and single-target domains. Here, we proposed a novel and concise method, coined MDANet, for the adaptation of patch images of multi-source and multi-target domains and for reducing the distribution differences of different patch images by projecting them onto the virtual center of a mixed-domain. MDANet is a lightweight and self-supervised network that can be grafted with any semantic segmentation model. Our method significantly improved the segmentation accuracy of semantic segmentation models and showed higher stability and competitiveness than existing methods.
Hao Cui 0002, Guo Zhang 0001, Ji Qi 0001, Haifeng Li 0007, Chao Tao 0001, Shasha Hou, DeRen Li
IEEE Geosci. Remote. Sens. Lett.5
2022 Remote Sensing Image Scene Classification With Self-Supervised Paradigm Under Limited Labeled Samples
abstract
With the development of deep learning, supervised learning methods perform well in remote sensing image (RSI) scene classification. However, supervised learning requires a huge number of annotated data for training. When labeled samples are not sufficient, the most common solution is to fine-tune the pretraining models using a large natural image data set (e.g., ImageNet). However, this learning paradigm is not a panacea, especially when the target RSIs (e.g., multispectral and hyperspectral data) have different imaging mechanisms from RGB natural images. To solve this problem, we introduce a new self-supervised learning (SSL) mechanism to obtain the high-performance pretraining model for RSI scene classification from large unlabeled data. Experiments on three commonly used RSI scene classification data sets demonstrated that this new learning paradigm outperforms the traditional dominant ImageNet pretrained model. Moreover, we analyze the impacts of several factors in SSL on RSI scene classification, including the choice of self-supervised signals, the domain difference between the source and target data sets, and the amount of pretraining data. The insights distilled from this work can help to foster the development of SSL in the remote sensing community. Since SSL could learn from unlabeled massive RSIs, which are extremely easy to obtain, it will be a promising way to alleviate dependence on labeled samples and thus efficiently solve many problems, such as global mapping.
Chao Tao 0001, Ji Qi 0001, Weipeng Lu, Hao Wang 0069, Haifeng Li 0007
IEEE Geosci. Remote. Sens. Lett.1
2022 Cross-Sensor Remote-Sensing Images Scene Understanding Based on Transfer Learning Between Heterogeneous Networks
abstract
Over the past decades, the successful invention and employment of multiple sensors have marked the advent of a new era in multisensor remote-sensing (RS) images acquisition. To effectively utilize the massive multisensor images for RS scene understanding, we expect that a scene classification model learned with particular sensor data can generalize well to other sensor data. However, this is a very challenging task due to the cross-sensor data differences. In the deep learning (DL) pipeline, a common way to handle this challenging task is to fine-tune the models pretrained on source sensor data with limited labeled data from the target sensor. Unfortunately, fine-tune technique is usually applied between homogeneous networks, which may not be the best choice if the source and target data are largely different. To address these issues, we formulate the cross-sensor RS scene understanding problem as a heterogeneous network-oriented transfer learning problem, in which the source and the target networks are different and data-oriented selected. Afterward, the knowledge between heterogeneous networks is transferred using the pseudo-label recursive propagation mechanism inspired by the concept of knowledge distillation. To the best of our knowledge, this is the first time to investigate the cross-sensor scene classification problem by constructing such a heterogeneous networks’ transfer scheme in RS fields. Our experiments using two cross-sensor RS datasets [aerial images$\rightarrow $multispectral images (MSIs) and aerial images$\rightarrow $hyper-spectral images (HSIs)] demonstrated that the proposed transfer learning strategy based on heterogeneous networks outperforms the supervised learning (SL) and fine-tune scheme for cross-sensor scene classification.
Yuze Wang 0005, Ji Qi 0001, Chao Tao 0001
IEEE Geosci. Remote. Sens. Lett.4
2022 Fine-Grained Road Scene Understanding From Aerial Images Based on Semisupervised Semantic Segmentation Networks
abstract
High-precision electronic maps are required to provide more detailed and accurate information than traditional maps. With the rapid development of high-resolution remote sensing technology, it has become possible to extract fine-grained road scene information such as vehicles, road lines, zebra crossings, ground signs, and lane widths of roads from unmanned aerial vehicle (UAV) remote sensing images, which opens up opportunities for automatic mapping high-precision maps. The traditional method of deciphering remote sensing images is often obtained through manual visual interpretation. Due to the high cost and long lead time of this method, it leads to inefficiencies in updating large amounts of information. To address this problem, this letter models the fine-grained road scene understanding task as an image semantic segmentation problem and innovatively proposes a semisupervised fully convolutional neural network to extract the information efficiently at a low cost. Compared with the traditional supervised full convolutional neural network, this method can simultaneously optimize the standard supervised classification loss on labeled samples and the unsupervised consistency loss on unlabeled samples by using an integrated prediction technology and then input them to the end-to-end semantic segmentation network for training. This method is designed to effectively improve the classification accuracy of the semantic segmentation network and validly alleviates overfitting problems in the case of small numbers of labeled samples. In order to verify the effectiveness of this method, we constructed a data set for experimental, which is used to verify the effect of a variable number of unlabeled samples on model performance. Experimental results show that our method can efficiently complete the extraction of fine-grained road scene information such as vehicles, road lines, zebra crossings, ground signs, and lane widths of roads with a small number of labeled samples.
Yuze Wang 0005, Chao Tao 0001
IEEE Geosci. Remote. Sens. Lett.3
2022 FALSE: False Negative Samples Aware Contrastive Learning for Semantic Segmentation of High-Resolution Remote Sensing Image
abstract
Self-supervised contrastive learning (SSCL) is a potential learning paradigm for learning remote sensing image (RSI)-invariant features through the label-free method. The existing SSCL of RSI is built based on constructing positive and negative sample pairs. However, due to the richness of RSI ground objects and the complexity of the RSI contextual semantics, the same RSI patches have the coexistence and imbalance of positive and negative samples, which causing the SSCL pushing negative samples far away while pushing positive samples far away, and vice versa. We call this the sample confounding issue (SCI). To solve this problem, we propose a False negAtive sampLes aware contraStive lEarning model (FALSE) for the semantic segmentation of high-resolution RSIs. Since the SSCL pretraining is unsupervised, the lack of definable criteria for false negative sample (FNS) leads to theoretical undecidability, we designed two steps to implement the FNS approximation determination: coarse determination of FNS and precise calibration of FNS. We achieve coarse determination of FNS by the FNS self-determination (FNSD) strategy and achieve calibration of FNS by the FNS confidence calibration (FNCC) loss function. Experimental results on three RSI semantic segmentation datasets demonstrated that the FALSE effectively improves the accuracy of the downstream RSI semantic segmentation task compared with the current three models, which represent three different types of SSCL models. The mean Intersection-over-Union on ISPRS Potsdam dataset is improved by 0.7% on average; on CVPR DGLC dataset is improved by 12.28% on average; and on Xiangtan dataset this is improved by 1.17% on average. This indicates that the SSCL model has the ability to self-differentiate FNS and that the FALSE effectively mitigates the SCI in self-supervised contrastive learning.
Xuying Wang, Xiaoming Mei, Chao Tao 0001, Haifeng Li 0007
IEEE Geosci. Remote. Sens. Lett.4
2022 Contextual Information-Preserved Architecture Learning for Remote-Sensing Scene Classification
abstract
Convolutional neural networks (CNNs) have recently been widely used in remote-sensing scene classification. Additionally, it is becoming very popular to automatically learn specific CNN architectures for specific data sets. The rich contextual information in high-resolution remote-sensing images (RSIs) is critical to remote-sensing intelligent understanding tasks. However, architecture learning approaches tend to simplify the original data (i.e., resizing images to smaller resolution) for efficiency, yet result in contextual information loss of RSIs. In this article, we proposed a contextual information-preserved architecture learning (CIPAL) framework for remote-sensing scene classification to utilize the contextual information in RSIs as much as possible during the architecture learning process. We introduce channel compression into CIPAL, which can reduce the memory and time consumption of architecture learning and make it possible to construct a larger architecture space. We add potential operators that are rarely used for scene classification tasks (i.e., atrous convolution) into the architecture space to explore unknown architectures that are more suitable for remote-sensing scenes. The experimental results on four remote-sensing scene classification benchmarks indicate that CIPAL learns architectures with less time consumption than similar works, and the newly found architectures outperform popular hand-designed architectures for better use of contextual information in RSIs. Different architectures are good at learning different representations, and our proposed architecture learning method potentially helps us understand which types of representations are crucial for RSI intelligent understanding.
Jie Chen 0048, Haozhe Huang, Jian Peng 0009, Li Chen 0025, Chao Tao 0001, Haifeng Li 0007
IEEE Trans. Geosci. Remote. Sens.6
2022 Thick Cloud Removal in Optical Remote Sensing Images Using a Texture Complexity Guided Self-Paced Learning Method
abstract
Thick clouds seriously impact the quality of optical remote sensing images (RSIs) and limit their application. For removing the cloud, some learning-based methods have been proposed and attracted considerable attention. However, these methods need to train paired multitemporal images with/without cloud, which are difficult and costly to collect. To solve this problem, we propose a novel texture complexity-guided self-paced learning (SPL) framework to remove the thick cloud from single RSIs. The framework does not need paired images and it exploits a texture complexity-guided mechanism to rank the self-generated cloud-corrupted training samples by texture complexity from low to high and then trains the generative adversarial cloud removal network using the SPL technique. In this way, the cloud removal network learns to restore the cloud-corrupted areas from easy to hard and thus to realize the image reconstruction for different difficulty levels. In addition, we introduce a structural similarity (SSIM) loss function to optimize the training network and improve the coherence of the image structure. Simulated and real experiments are performed on the single images acquired by Gao Fen-1 (GF-1) and Sentinel-2 satellites to validate the effectiveness of the proposed method. The results show that the proposed method has a better performance in cloud removal than other state-of-the-art methods, especially for the images of the areas with complex textures. The source codes are available athttps://github.com/GeoX-Lab/TPL.
Chao Tao 0001, Siyang Fu, Ji Qi 0001, Haifeng Li 0007
IEEE Trans. Geosci. Remote. Sens.1
2022 Avoiding Negative Transfer for Semantic Segmentation of Remote Sensing Images
abstract
Reducing the feature distribution shift caused by the factor of visual-environment changes, namely as VE-changes, is a hot issue in domain adaptation learning. However, in the semantic segmentation task of remote sensing imageries, besides VE-changes, the change of semantic-scenes (SS-changes) is another factor raising domain gap, which brings the label distribution shift. For example, although urban and rural share the same landcover label, there is still a gap in label distribution. If there is little relation that can be found in neither feature nor label space, forcibly adapting to a new domain could have a high risk of negative transfer. Hence, we propose a new Transitive Domain Adaptation method for Remote Sensing images (TDARS). Firstly, we introduce an intermediate domain to enlarge the relation between the given source and target domains. Secondly, we learn from primary and non-primary confident classes to increase the likelihood of transferring valuable information. As a result, TDARS enables the given source and target domains to be connected through the selected intermediate domain and performs effective knowledge transfer among all domains. The proposed method is evaluated on three domain adaptation datasets of remote sensing images. Extensive experiments show the approach can effectively handle the domain shift problem from remote sensing images compared to other state-of-the-art domain adaptation methods.
Hao Wang 0069, Chao Tao 0001, Ji Qi 0001, Haifeng Li 0007
IEEE Trans. Geosci. Remote. Sens.2
2022 KST-GCN: A Knowledge-Driven Spatial-Temporal Graph Convolutional Network for Traffic Forecasting
abstract
While considering the spatial and temporal features of traffic, capturing the impacts of various external factors on travel is an essential step towards achieving accurate traffic forecasting. However, existing studies seldom consider external factors or neglect the effect of the complex correlations among external factors on traffic. Intuitively, knowledge graphs can naturally describe these correlations. Since knowledge graphs and traffic networks are essentially heterogeneous networks, it is challenging to integrate the information in both networks. On this background, this study presents a knowledge representation-driven traffic forecasting method based on spatial-temporal graph convolutional networks. We first construct a knowledge graph for traffic forecasting and derive knowledge representations by a knowledge representation learning method named KR-EAR. Then, we propose the Knowledge Fusion Cell (KF-Cell) to combine the knowledge and traffic features as the input of a spatial-temporal graph convolutional backbone network. Experimental results on the real-world dataset show that our strategy enhances the forecasting performances of backbones at various prediction horizons. The ablation and perturbation analysis further verify the effectiveness and robustness of the proposed method. To the best of our knowledge, this is the first study that constructs and utilizes a knowledge graph to facilitate traffic forecasting; it also offers a promising direction to integrate external information and spatial-temporal information for traffic forecasting. The source code is available athttps://github.com/lehaifeng/T-GCN/tree/master/KST-GCN.
Xing Han, Hanhan Deng, Chao Tao 0001, Ling Zhao 0005, Pu Wang 0005, Tao Lin 0008, Haifeng Li 0007
IEEE Trans. Intell. Transp. Syst.4
2021 SCAttNet: Semantic Segmentation Network With Spatial and Channel Attention Mechanism for High-Resolution Remote Sensing Images
abstract
High-resolution remote sensing images (HRRSIs) contain substantial ground object information, such as texture, shape, and spatial location. Semantic segmentation, which is an important task for element extraction, has been widely used in processing mass HRRSIs. However, HRRSIs often exhibit large intraclass variance and small interclass variance due to the diversity and complexity of ground objects, thereby bringing great challenges to a semantic segmentation task. In this letter, we propose a new end-to-end semantic segmentation network, which integrates lightweight spatial and channel attention modules that can refine features adaptively. We compare our method with several classic methods on the ISPRS Vaihingen and Potsdam data sets. Experimental results show that our method can achieve better semantic segmentation results. The source codes are available at https://github.com/lehaifeng/SCAttNet.
Haifeng Li 0007, Kaijian Qiu, Li Chen 0025, Xiaoming Mei, Chao Tao 0001
IEEE Geosci. Remote. Sens. Lett.6
2021 Spatial Information Considered Network for Scene Classification
abstract
Remote sensing image (RSI) scene classification (RSISC) is a fundamental problem for understanding high-resolution RSIs. More recently, deep learning methods, especially convolutional neural networks (CNNs), and large data sets have greatly promoted the RSISC. However, deep learning methods rely heavily on the visual features extracted from the patches cropped from original RSIs, so the intraclass diversity and interclass similarity are two big challenges. To address these problems, in this letter, we propose a spatial information considered model to learn more discriminative features. By combining CNN and recurrent neural network, the proposed method can exploit both local and long-range spatial relation information to enhance the representational ability of the learned features. As the initial visual features of a single patch are transformed into higher-level features with spatial information, the proposed method achieves more accurate scene classification. Besides, we present an RSISC data set named as CSU-RSISC10 data set to preserve the spatial information between scenes in a new way of organization. Experiments demonstrate that the proposed method outperforms other three state-of-the-art methods in scene classification using CSU-RSISC10 data set.
Chao Tao 0001, Weipeng Lu, Ji Qi 0001, Hao Wang 0069
IEEE Geosci. Remote. Sens. Lett.1
2021 RS-MetaNet: Deep Metametric Learning for Few-Shot Remote Sensing Scene Classification
abstract
Training a modern deep neural network on massive labeled samples is the main paradigm in solving the scene classification problem for remote sensing, but learning from only a few data points remains a challenge. Existing methods for a few-shot remote sensing scene classification are performed in a sample-level manner, resulting in easy overfitting of learned features to individual samples and inadequate generalization of learned category segmentation surfaces. To solve this problem, learning should be organized at the task level rather than the sample level. Learning on tasks sampled from a task family can help tune learning algorithms to perform well on new tasks sampled in that family. Therefore, we propose a simple but effective method, called RS-MetaNet, to resolve the issues related to few-shot remote sensing scene classification in the real world. On the one hand, RS-MetaNet raises the level of learning from the sample to the task by organizing training in a metaway, and it learns to learn a metric space that can well classify remote sensing scenes from a series of tasks. We also propose a new loss function, called balance loss, which maximizes the generalization ability of the model to new samples by maximizing the distance between different categories, providing the scenes in different categories with better linear segmentation planes while ensuring model fit. The experimental results on three open and challenging remote sensing data sets, UCMerced_LandUse, NWPU-RESISC45, and Aerial Image Data, demonstrate that our proposed RS-MetaNet method achieves state-of-the-art results in cases where there are only 1 ~ 20 labeled samples.
Haifeng Li 0007, Zhenqi Cui, Zhiqiang Zhu, Li Chen 0025, Haozhe Huang, Chao Tao 0001
IEEE Trans. Geosci. Remote. Sens.7
2019 Spatial Information Inference Net: Road Extraction Using Road-Specific Contextual Information
abstract
For road extraction tasks in VHR satellite imagery, a deep neural network may perform well. But a network with certain reasoning ability as human will get a more satisfying result. To this end, we focus on how to effectively model the context information of the road and propose a well-designed spatial information inference structure (SIIS) which can add into any typical semantic segmentation network. The network with SIIS called SII-Net can not only learn the local visual characteristic of the road but also the global spatial structure information (such as the continuity and trend of the road). So, it can effectively solve the challenging occlusion problem in road detection and well preserve the continuity of the extracted road. The experimental results of two datasets show that the proposed method can improve the comprehensive performance of road extraction.
Ji Qi 0001, Chao Tao 0001, Hao Wang 0069, Zhenqi Cui
IGARSS2
2019 Semi-Supervised Variational Generative Adversarial Networks for Hyperspectral Image Classification
abstract
Though Hyperspectral Image (HSI) Classification has been extensively investigated over recent decades, it is still a challenge task especially when the number of labeled samples is extremely limited. In this paper, we overcome this challenge by using synthetic samples, and proposed a semi-supervised variational Generative Adversarial Networks(GANs) for this purpose. Compared to the conditional GAN which is recently used for generating samples for HSI classification, the proposed approach has two novel aspects. First, we extend the classic variational generative adversarial network to the semi-supervised context through an ensemble prediction technique. By this way, our model can be trained using limited labeled samples (only 5 samples per class) with a large number of unlabeled samples. Second, we adopt an encoder-decoder network to explicitly learn the relationship between the latent space and the real image space. This property enables our model producing diverse samples by simply varying some latent parameters, which is desirable for enriching the training dataset. We have shown that the proposed model can achieve better and robust performance for HSI classification compared to conditional GAN, especially when the labeled data is limited.
Hao Wang 0069, Chao Tao 0001, Ji Qi 0001, Haifeng Li 0007
IGARSS2
2019 Scene Context-Driven Vehicle Detection in High-Resolution Aerial Images
abstract
As the spatial resolution of remote sensing images is improving gradually, it is feasible to realize “scene-object” collaborative image interpretation. Unfortunately, this idea is not fully utilized in vehicle detection from high-resolution aerial images, and most of the existing methods may be promoted by considering the variability of vehicle spatial distribution in different image scenes and treating vehicle detection tasks scene-specific. With this motivation, a scene context-driven vehicle detection method is proposed in this paper. At first, we perform scene classification using the deep learning method and, then, detect vehicles in roads and parking lots separately through different vehicle detectors. Afterward, we further optimize the detection results using different postprocessing rules according to different scene types. Experimental results show that the proposed approach outperforms the state-of-the-art algorithms in terms of higher detection accuracy rate and lower false alarm rate.
Chao Tao 0001, Li Mi, Yansheng Li 0001, Ji Qi 0001
IEEE Trans. Geosci. Remote. Sens.1
2017 An On-Road Vehicle Detection Method for High-Resolution Aerial Images Based on Local and Global Structure Learning
abstract
With the continuous improvement of image resolution, details on aerial images provide abundant available information for vehicle detection. Nevertheless, traditional works mainly exploited the overall information of the vehicles ignoring the local details, such as front and rear windshields, and thus, there were usually more than 15% false alarms in the final vehicle detection results. In this letter, we propose a vehicle detection method making full use of high level details on aerial images. In the training stage, we choose front windshield samples to train a part detector and whole vehicle samples to train a root detector. In the matching stage, we first use the root detector to define an entire vehicle obtaining the root response, then the part detector is scanned in the root bounding box to decide a front windshield and get the part response. Afterward, the part response is transformed by setting weight w based on the part position offset. More importantly, contextual information is appropriately used in the process of determining the part position offset. Final detection score is the combination of root response and the transformed part response. We have demonstrated that the proposed method has achieved better performance with more than 6.43% increase of correct detection rate and more than 5.63% decrease of false detection rate compared with the state-of-the-art approaches.
Chao Tao 0001, Zhengrong Zou
IEEE Geosci. Remote. Sens. Lett.2
2016 A vehicle detection method taking shadow areas into account for high resolution aerial imagery
abstract
Detecting cars from high-resolution remote sensing images is vulnerable to the effect of shadows in the image, as a result, cars in shadow areas are hard to be detected due to their weak visual features. In order to solve this problem, a vehicle detection method which takes the shaded areas into account is proposed in this letter. Firstly, we extract shadows in road areas using a shadow detection algorithm which is based on color features, and then we enhance the shaded areas by histogram equalization to improve cars' visual characteristics; Secondly, vehicle detection models M1,M2are trained in shaded and non-shaded regions respectively using HOG features combined with SVM classification method. Finally, we use M1to extract cars in shadows and the ones which are not in shadows are detected by M2, then the final result is obtained through the combination of the two outcomes. Experiments using multiple sets of test images show that: in contrast with traditional car counting method, the proposed one has the ability to improve the performance of vehicle detection, the probability of wrong detection is decreased as well.
Chao Tao 0001, Zhengrong Zou
IGARSS2
2016 Unsupervised Multilayer Feature Learning for Satellite Image Scene Classification
abstract
This letter proposes a simple but effective approach to automatically learn a multilayer image feature for satellite image scene classification. Different from the hand-crafted features which are empirically designed but lack high generalization ability, the proposed approach can autonomously extract the data-dependent feature. The presented feature extraction algorithm is composed of two layers, and the bases of these two layers are uniformly learned by a plain $K$-means clustering algorithm. Coincidentally, the feature extraction performance of the aforementioned two layers is consistent with visual processing of human visual cortex. More specifically, the first layer can generate edgelike bases, which are analogous to the neuron responses of primary visual cortex (V1), and the second layer can produce cornerlike bases, which resemble the neuron responses of visual extrastriate cortical area two (V2). The proposed feature extraction approach can automatically extract not only simple structure features (e.g., edges) but also complex structure features (e.g., corners and junctions). The learned feature is further discriminated by the linear support vector machine classifier for scene classification. In order to fairly demonstrate the validity of the proposed feature extraction approach, its satellite image scene classification performance is evaluated on the public UCM-21 data set. Experimental results show that the proposed approach can outperform several recent state-of-the-art approaches.
Yansheng Li 0001, Chao Tao 0001, Yihua Tan, Jinwen Tian
IEEE Geosci. Remote. Sens. Lett.2
2015 Unsupervised Spectral-Spatial Feature Learning With Stacked Sparse Autoencoder for Hyperspectral Imagery Classification
abstract
In this letter, different from traditional methods using original spectral features or handcraft spectral-spatial features, we propose to adaptively learn a suitable feature representation from unlabeled data. This is achieved by learning a feature mapping function based on stacked sparse autoencoder. Considering that hyperspectral imagery (HSI) is intrinsically defined in both the spectral and spatial domains, we further establish two variants of feature learning procedures for sparse spectral feature learning and multiscale spatial feature learning. Finally, we embed the learned spectral-spatial feature into a linear support vector machine for classification. Experiments on two hyperspectral images indicate the following: 1) the learned spectral-spatial feature representation is more discriminative for HSI classification compared to previously hand-engineered spectral-spatial features, especially when the training data are limited and 2) the learned features appear not to be specific to a particular image but general in that they are applicable to multiple related images (e.g., images acquired by the same sensor but varying with location or time).
Chao Tao 0001, Yansheng Li 0001, Zhengrou Zou
IEEE Geosci. Remote. Sens. Lett.1
2014 Hyperspectral Imagery Classification Based on Rotation-Invariant Spectral-Spatial Feature
abstract
In this letter, we present a novel approach for spectral-spatial classification in hyperspectral imagery. After applying principal component (PC) analysis for dimensionality reduction, we extract the spectral-spatial information by first reorganizing the local image patch with the first d PCs into a vector representation, followed by a sorting scheme to make the vector invariant to local image rotation. Since no additional operation except sorting the pixels is required, this step is performed efficiently. Afterward, the resulting feature descriptors are embedded into a linear support vector machine for classification. To evaluate the proposed method, experiments are preformed on two hyperspectral images with high spatial resolution. The experimental results confirm that the proposed method outperforms the existing algorithms on classification accuracy.
Chao Tao 0001, Chong Fan, Zhengrong Zou
IEEE Geosci. Remote. Sens. Lett.1
2013 Compressed texton based sorted visual words co-occurrence matrix for high resolution remote sensing imagery classification
abstract
A novel, simple, yet effective texture extraction method for high resolution remote sensing imagery classification based on visual words co-occurrence matrix is proposed in this paper. First, Local texture is represented by compressed texton learned from raw image patch with a sorting scheme and random projection. Then the sorted visual words co-occurrence matrix obtained with dictionary learning and nearest neighbor encoding is used for representing global texture. Finally, the support vector machine is applied for classification. Two imagery from Pavia city of Italy with public ground truth dataset are used in our experiments. The results show that the proposed method is effective and outperforms other existing methods.
Chao Tao 0001, Huiyun Ma, Zhengrong Zou
IGARSS2
2013 Hyperspectral imagery classification based on rotation invariant spectral-spatial feature
abstract
In this letter, we present a novel approach for spectral-spatial classification in hyperspectral imagery. To this end, after applying principal component analysis (PCA) for dimensionality reduction, we extract the spectral-spatial information by first reorganizing the local image patch with the first d principal components(PCs) into a vector representation, followed by a sorting scheme to make it invariant to local image rotation.. Since no additional operation except sorting the pixels is required, this step is performed efficiently. Afterwards, the resulting feature descriptors are embedded into a linear support vector machine (SVM) for classification. To evaluate the proposed method, experiments were preformed on two hyperspectral images with high spatial resolution. The experimental results confirmed that the proposed method outperforms the existing algorithms on classification accuracy.
Chao Tao 0001, Zhengrong Zou
IGARSS1
2013 A Robust Directional Saliency-Based Method for Infrared Small-Target Detection Under Various Complex Backgrounds
abstract
Infrared small-target detection plays an important role in image processing for infrared remote sensing. In this letter, different from traditional algorithms, we formulate this problem as salient region detection, which is inspired by the fact that a small target can often attract attention of human eyes in infrared images. This visual effect arises from the discrepancy that a small target resembles isotropic Gaussian-like shape due to the optics point spread function of the thermal imaging system at a long distance, whereas background clutters are generally local orientational. Based on this observation, a new robust directional saliency-based method is proposed incorporating with visual attention theory for infrared small-target detection. Experimental results demonstrate that the proposed algorithm outperforms the state-of-the-art methods for real infrared images with various typical complex backgrounds.
Shengxiang Qi, Jie Ma 0003, Chao Tao 0001, Changcai Yang, Jinwen Tian
IEEE Geosci. Remote. Sens. Lett.3
2013 Unsupervised Detection of Built-Up Areas From Multiple High-Resolution Remote Sensing Images
abstract
Given a set of high-resolution remote sensing images covering different scenes, we propose an unsupervised approach to simultaneously detect possible built-up areas from them. The motivation behind is that the frequently recurring appearance patterns or repeated textures corresponding to common objects of interest (e.g., built-up areas) in the input image data set can help us discriminate built-up areas from others. With this inspiration, our method consists of two steps. First, we extract a large set of corners from each input image by an improved Harris corner detector. Afterward, we incorporate the extracted corners into a likelihood function to locate candidate regions in each input image. Given a set of candidate build-up regions, in the second stage, we formulate the problem of build-up area detection as an unsupervised grouping problem. The candidate regions are modeled through texture histogram, and the grouping problem is solved by spectrum clustering and graph cuts. Experimental results show that the proposed approach outperforms the existing algorithms in terms of detection accuracy.
Chao Tao 0001, Yihua Tan, Zhengrong Zou, Jinwen Tian
IEEE Geosci. Remote. Sens. Lett.1
2012 Urban area detection using multiple Kernel Learning and graph cut
abstract
This paper presents a new method for urban detection from high-spatial-resolution satellite images. Unlike traditional approaches using only texture information for urban detection, we integrate several complementary image features through multiple Kernel Learning framework, and demonstrate that fusing multiple features can help improving urban detection accuracy rate. Furthermore, since that most of supervised urban classification approaches are mainly based on block-based image interpretation, the resulting urban boundary is very coarse. To handle this, we formulate the urban boundary refinement as a binary labeling problem, and propose a graph cut based approach to solve it. Experimental results show that the proposed approach outperforms the existing algorithm in terms of detection accuracy.
Chao Tao 0001, Yihua Tan, Jin-Gang Yu, Jin-Wen Tian
IGARSS1
2012 High-resolution satellite image registration using local feature and contour fragment
abstract
Recently, image registration approaches based on local feature matching have revealed promising results in the fields of natural and medical image processing. However, the performance becomes worse when it is applied to high-resolution satellite image. This arises because the complex nature of this kind of image, which results in a lot of outliers in the initial matches. Therefore, how to select reliable matches from low-precision initial matches becomes very crucial. In this paper, we modify the traditional SIFT descriptor and propose a novel algorithm to identify outliers from initial SIFT matches, which is based on the consistence of local contour fragments between two matched regions. The experimental results show that the presented approach can extract high-precision matches from low-precision initial SIFT matches, and is efficient for high-resolution satellite image registration.
Chao Tao 0001, Zhengrong Zou, Hanqiu Sun
IGARSS1
2011 Airport Detection From Large IKONOS Images Using Clustered SIFT Keypoints and Region Information
abstract
This letter presents a new method for airport detection from large high-spatial-resolution IKONOS images. To this end, we describe airport by a set of scale-invariant feature transform (SIFT) keypoints and detect it using an improved SIFT matching strategy. After obtaining SIFT matched keypoints, to both discard the redundant matched points and locate the possible regions of candidates that contain the target, a novel region-location algorithm is proposed, which exploits the clustering information from matched SIFT keypoints, as well as the region information extracted through the image segmentation. Finally, airport recognition is achieved by applying the prior knowledge to the candidate regions. Experimental results show that the proposed approach outperforms the existing algorithms in terms of detection accuracy.
Chao Tao 0001, Yihua Tan, Huajie Cai, Jin-Wen Tian
IEEE Geosci. Remote. Sens. Lett.1