Haifeng Li 0007

dblp:17/246-7 · DBLP profile ↗
← Back
62ranked-venue papers
6as first author
55since 2021 · last 2026
0000-0003-1173-6593ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 36 · 3 first-author · 34 since 2021Artificial intelligence and machine learning · 13 · 1 first-author · 12 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author · 4 since 2021Systems, architecture and hardware · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author
YearPublicationVenuePosition
2026 A Gift From the Integration of Discriminative and Diffusion-Based Generative Learning: Boundary Refinement Remote Sensing Semantic Segmentation
abstract
Remote sensing semantic segmentation must address both what the ground objects are within an image and where they are located. Consequently, segmentation models must ensure not only the semantic correctness of large-scale patches (low-frequency information) but also the precise localization of boundaries between patches (high-frequency information related to boundary components). However, most existing approaches rely heavily on discriminative learning, which excels at capturing low-frequency features, while overlooking its inherent limitations in learning high-frequency features for semantic segmentation. Recent studies have revealed that diffusion generative models excel at generating high-frequency details. Our theoretical analysis confirms that the diffusion denoising process significantly enhances the model's ability to learn high-frequency features; however, we also observe that these models exhibit insufficient semantic inference for low-frequency features when guided solely by the original image. Therefore, we integrate the strengths of both discriminative and generative learning, proposing the Integration of Discriminative and diffusion-based Generative learning for Boundary Refinement (IDGBR) framework. The framework first generates a coarse segmentation map using a discriminative backbone model. This map and the original image are fed into a conditioning guidance network to jointly learn a guidance representation subsequently leveraged by an iterative denoising diffusion process refining the coarse segmentation. Extensive experiments across five remote sensing semantic segmentation datasets (binary and multi-class segmentation) confirm our framework's capability of consistent boundary refinement for coarse results from diverse discriminative architectures. The source code is available at https://github.com/KeyanHu-git/IDGBR.
Hao Wang 0069, Keyan Hu, Haifeng Li 0007, Chao Tao 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2026 Multi-modality transfer learning for cloudy remote sensing images: Addressing modality imbalance with knowledge distillation
Yuze Wang 0005, Haifeng Li 0007, Mariana Belgiu, Chao Tao 0001
Pattern Recognit.3
2025 Causal invariant geographic network representations with feature and structural distribution shifts
abstract
Relationships between geographic entities, including human-land and human-people relationships, can be naturally modelled by graph structures, and geographic network representation is an important theoretical issue. The existing methods learn geographic network representations through deep graph neural networks (GNNs) based on the i.i.d. assumption. However, the spatial heterogeneity and temporal dynamics of geographic data make the out-of-distribution (OOD) generalisation problem particularly salient. We classify geographic network representations into invariant representations that always stabilise the predicted labels under distribution shifts and background representations that vary with different distributions. The latter are particularly sensitive to distribution shifts (feature and structural shifts) between testing and training data and are the main causes of the out-of-distribution generalisation (OOD) problem. Spurious correlations are present between invariant and background representations due to selection biases/environmental effects, resulting in the model extremes being more likely to learn background representations. The existing approaches focus on background representation changes that are determined by shifts in the feature distributions of nodes in the training and test data while ignoring changes in the proportional distributions of heterogeneous and homogeneous neighbour nodes, which we refer to as structural distribution shifts. We propose a feature-structure mixed invariant representation learning (FSM-IRL) model that accounts for both feature distribution shifts and structural distribution shifts. To address structural distribution shifts, we introduce a sampling method based on causal attention, encouraging the model to identify nodes possessing strong causal relationships with labels or nodes that are more similar to the target node. This approach significantly enhances the invariance of the representations between the source and target domains while reducing the dependence on background representations that arise by chance or in specific patterns. Inspired by the Hilbert–Schmidt independence criterion, we implement a reweighting strategy to maximise the orthogonality of the node representations, thereby mitigating the spurious correlations among the node representations and suppressing the learning of background representations. In addition, we construct an educational-level geographic network dataset under out-of-distribution (OOD) conditions. Our experiments demonstrate that FSM-IRL exhibits strong learning capabilities on both geographic and social network datasets in OOD scenarios.
Silu He, Qinyao Luo, Hongyuan Yuan, Ling Zhao 0005, Haifeng Li 0007
Future Gener. Comput. Syst.7
2025 Task Knowledge Injection: Training-Free Adaptation of Multimodal Large Language Models for Remote Sensing Image Understanding
abstract
Parameter fine-tuning is the mainstream approach for adapting Multimodal Large Langauge Models (MLLMs) to downstream remote sensing tasks. However, such method risks degrading pre-trained knowledge and also incur significant costs. This paper argues that downstream adaptation of MLLMs essentially involves effective injection of task-specific knowledge, which does not necessarily require parameter updates. Based on this perspective, we propose a training-free knowledge injection method. By constructing a multi-task knowledge base (MTKB), the model can dynamically retrieve task-related knowledge to serve as context during inference, thereby enhancing its understanding. Specifically, we design a three-part framework. (1) Task knowledge construction: Diverse texts are unified into a key and value structure for image-text matching, forming the MTKB. (2) Two-stage retrieval: A coarse-to-fine process is employed to match query images with task knowledge in the MTKB, utilizing both unimodal and cross-modal similarity. (3) Knowledge injection: Matched knowledge is integrated into the MLLM via extended embeddings, without altering parameters. Experimental results across multiple datasets demonstrate that our method significantly enhances the image understanding capabilities of the model. Our approach achieves about 5% improvement in accuracy and related metrics across several datasets, with performance on the RSVQ-LR dataset comparable to specialized models.
Haifeng Li 0007, Qiujun Li, Wang Guo, Hongyuan Yuan, Run Shao, Chengli Peng
IEEE Geosci. Remote. Sens. Lett.1
2025 Causal Invariant Representation Learning Based on Style Intervention Identity Regularization for Remote Sensing Image
abstract
An intelligent understanding model of a remote sensing image will present different visual representations of the same object in the remote sensing image, under the interference of offset factors, such as weather and season. This variability adversely affects the generalization ability of the model; therefore, an open challenge is how to learn invariant features. These features can maintain the model’s generalization ability under various imaging conditions and given the spatiotemporal heterogeneity of the object. From the causal point of view, each object of the remote sensing image is decomposed into two parts, content and style; content is the causal variable of the target task, which is robust and stable, and style is the noncausal variable, which is more susceptible to the interference of bias factors and is volatile. This paper proposes a remote sensing imagery invariance representation method based on style intervention identity regularization (ICRNet). Within the contrastive learning framework, style transformation based on a diffusion model is introduced as an intervention factor to modify the style of remote sensing images. In addition, style intervention identity regularization is designed, employing symmetric JS divergence as the similarity measure. This approach forces remote sensing images with identical content but different styles to have a similar impact on the target task, thereby enabling learning invariant content representations while ignoring style representations. The experiments were conducted on three remote sensing semantic segmentation datasets. The results reveal that ICRNet learns invariant content representations better than alternative representation learning approaches based on contrast learning. Additionally, ICRNet exhibits higher generalizability under seasonal bias. The source code is available at https://github.com/GeoX-Lab/ICRNet.
Yunsheng Zhang 0001, Fanfan Liu, Haifeng Li 0007
IEEE Geosci. Remote. Sens. Lett.4
2025 STAR: A First-Ever Dataset and a Large-Scale Benchmark for Scene Graph Generation in Large-Size Satellite Imagery
abstract
Scene graph generation (SGG) in satellite imagery (SAI) benefits promoting understanding of geospatial scenarios from perception to cognition. In SAI, objects exhibit great variations in scales and aspect ratios, and there exist rich relationships between objects (even between spatially disjoint objects), which makes it attractive to holistically conduct SGG in large-size very-high-resolution (VHR) SAI. However, there lack such SGG datasets. Due to the complexity of large-size SAI, mining triplets subject, relationship, object heavily relies on long-range contextual reasoning. Consequently, SGG models designed for small-size natural imagery are not directly applicable to large-size SAI. This paper constructs a large-scale dataset for SGG in large-size VHR SAI with image sizes ranging from 512 × 768 to 27,860 × 31,096 pixels, named STAR (Scene graph generaTion in lArge-size satellite imageRy), encompassing over 210K objects and over 400K triplets. To realize SGG in large-size SAI, we propose a context-aware cascade cognition (CAC) framework to understand SAI regarding object detection (OBD), pair pruning and relationship prediction for SGG. We also release a SAI-oriented SGG toolkit with about 30 OBD and 10 SGG methods which need further adaptation by our devised modules on our challenging STAR dataset. The dataset and toolkit are available at: https://linlin-dev.github.io/project/STAR.
Yansheng Li 0001, Tingzhu Wang, Xue Yang 0005, Qi Wang 0009, Youming Deng, Xian Sun 0001, Haifeng Li 0007, Bo Dang 0002, Yongjun Zhang 0002, Yi Yu 0010, Junchi Yan
IEEE Trans. Pattern Anal. Mach. Intell.10
2025 IFShip: Interpretable fine-grained ship classification with domain knowledge-enhanced vision-language models
Mingning Guo, Mengwei Wu, Yuxiang Shen, Haifeng Li 0007, Chao Tao 0001
Pattern Recognit.4
2025 SparseFormer: A Credible Dual-CNN Expert-Guided Transformer for Remote Sensing Image Segmentation With Sparse Point Annotation
abstract
Although significant advances have been made in the semantic segmentation of high-resolution remote sensing (RS) images, obtaining accurate pixelwise annotations remains resource-intensive. We propose SparseFormer, a credible dual-convolutional neural network (CNN) expert-guided Transformer model designed for semantic segmentation using point-level annotations to reduce this annotation burden. SparseFormer comprises three branches, where two CNN branches employ different attention mechanisms to encourage diverse outputs. To enhance the local consistency of pseudolabels, we introduce a pixel-adaptive refinement (PAR) module that dynamically refines CNN output probabilities by incorporating image information during training. A credible assessment is then performed to combine the CNN outputs, producing high-quality pseudolabels that supervise the CNN-Transformer hybrid branch. This hybrid branch integrates global representations with local features, achieving precise segmentation. To further strengthen the CNN branches, we introduce a knowledge distillation strategy that steadily feeds back information from the hybrid branch to CNN branches, mitigating overfitting risks caused by sparse supervision. SparseFormer employs credible assessment to reduce pseudolabel uncertainty, followed by continuous interaction and dynamic information enhancement among the three branches in an end-to-end training process. Extensive experiments on two benchmark datasets demonstrate that SparseFormer significantly outperforms state-of-the-art methods. Our code is available at:https://github.com/Yujia73/SparseFormer.
Hao Cui 0002, Guo Zhang 0001, Zhigang Xie, Haifeng Li 0007, DeRen Li
IEEE Trans. Geosci. Remote. Sens.6
2025 AllSpark: A Multimodal Spatiotemporal General Intelligence Model With Ten Modalities via Language as a Reference Framework
abstract
RGB, multispectral, point and other spatio-temporal modal data fundamentally represent different observational approaches for the same geographic object. Therefore, leveraging multimodal data is an inherent requirement for comprehending geographic objects. However, due to the high heterogeneity in structure and semantics among various spatio-temporal modalities, the joint interpretation of multimodal spatio-temporal data has long been an extremely challenging problem. The primary challenge resides in striking a trade-off between the cohesion and autonomy of diverse modalities. This trade-off becomes progressively nonlinear as the number of modalities expands. Inspired by the human cognitive system and linguistic philosophy, where perceptual signals from the five senses converge into language, we introduce the Language as Reference Framework (LaRF), a fundamental principle for constructing a multimodal unified model. Building upon this, we propose AllSpark, a multimodal spatio-temporal general artificial intelligence model. Our model integrates ten different modalities into a unified framework, including one-dimensional (language, code, table), two-dimensional (RGB, SAR, multispectral, hyperspectral, graph, trajectory), and three-dimensional (point cloud) modalities. To achieve modal cohesion, AllSpark introduces a modal bridge and multimodal large language model (LLM) to map diverse modal features into the language feature space. To maintain modality autonomy, AllSpark uses modality-specific encoders to extract the tokens of various spatio-temporal modalities. Finally, observing a gap between the model’s interpretability and downstream tasks, we designed modality-specific prompts and task heads, enhancing the model’s generalization capability across specific tasks. Experiments indicate that the incorporation of language enables AllSpark to excel in few-shot classification tasks for RGB and point cloud modalities without additional training, surpassing baseline performance by up to 41.82%. Additionally, AllSpark, despite lacking expert knowledge in most spatio-temporal modalities and utilizing a unified structure, demonstrates strong adaptability across ten modalities. LaRF and AllSpark contribute to the shift in the research paradigm in spatio-temporal intelligence, transitioning from a modality-specific and task-specific paradigm to a general paradigm. The source code is available at https://github.com/GeoX-Lab/AllSpark.
Run Shao, Qiujun Li, Linrui Xu, Qing Zhu 0012, Yongjun Zhang 0002, Yansheng Li 0001, Yu Liu 0003, Shizhong Yang, Haifeng Li 0007
IEEE Trans. Geosci. Remote. Sens.15
2025 HASNet: A Foreground Association-Driven Siamese Network With Hard Sample Optimization for Remote Sensing Image Change Detection
abstract
Remote sensing change detection (RS-CD) relies on the model’s ability to learn features of marked change objects, known as foreground targets. Beyond foreground targets, the background targets are more valuable samples for change detection, such as unlabeled ones, semantically ambiguous ones, pseudo-changes, and non-interesting changes, referred to as hard case samples (HCSs) in this article. There are two additional challenges to learning HCSs: 1) the loss function focusing on the foreground targets with rich labels and ignoring the HCSs in the background, called the imbalance problem and 2) it is difficult for a model to learn the change information of HCSs directly, which is called HCSs missingness. This article proposed a foreground association-driven Siamese network with hard sample optimization (HASNet). To deal with the imbalance problem, we propose an equilibrium optimization loss (EO-loss) function to regulate the optimization focus of the foreground and background, determine the HCSs through the distribution of the loss values, and introduce dynamic weights in the loss term to gradually shift the optimization focus of the loss from the foreground to the background hard cases as the training progresses. To address the HCSs missingness, we propose the scene-foreground association module by using potential remote sensing spatial scene information to model the association between the target of interest in the foreground and the related context to obtain scene embedding to reinforce the feature of hard cases. Experiments on four public datasets with 11 baselines show that HASNet outperforms current state-of-the-art CD methods, particularly in detecting HCSs. The source code is available athttps://github.com/GeoX-Lab/HASNet.
Chao Tao 0001, Dongsheng Kuang, Zhenyang Huang, Chengli Peng, Haifeng Li 0007
IEEE Trans. Geosci. Remote. Sens.5
2025 RSD-BiasEval: A Framework for Remote Sensing Dataset Bias Analysis and Evaluation
Xinrui Xie, ZiYue Lin, Haifeng Li 0007, Ji Qi 0001, Yibo Wang 0014, Yiping Chen 0002, Xinchang Zhang 0002
IEEE Trans. Geosci. Remote. Sens.4
2024 AdaER: An adaptive experience replay approach for continual lifelong learning
Bo Tang 0011, Haifeng Li 0007
Neurocomputing3
2024 CAT: A causal graph attention network for trimming heterophilic graphs
Silu He, Qinyao Luo, Xinsha Fu, Ling Zhao 0005, RongHua Du, Haifeng Li 0007
Inf. Sci.6
2024 LSTTN: A Long-Short Term Transformer-based spatiotemporal neural network for traffic flow forecasting
Qinyao Luo, Silu He, Xing Han, Haifeng Li 0007
Knowl. Based Syst.5
2024 Adversarial Examples for Vehicle Detection With Projection Transformation
abstract
Unmanned aerial vehicle (UAV) imaging object detection systems based on deep neural networks are vulnerable to adversarial patch attacks. However, existing UAV image adversarial patch generation methods mainly target flat digital images, neglecting the adjustments to the adversarial patch morphology brought about by changes in the imaging projection matrix in a 3-D physical environment. This leads to adaptability issues with the attack patches, where changes in the UAV’s observation angle might cause dynamic variations in the vehicle target’s position and size, such geometric instability means that preset attack patches may no longer be applicable. To address these issues, this article proposes an adversarial patch generation method based on a projection transformation model, which we call projective-patch attack. To tackle the adaptability problem of the attack patches, we employ an adversarial patch relocation strategy, adjusting the preset adversarial patch’s morphology, size, and positioning on the target vehicle through the projection transformation model to ensure omnidirectional adaptability at different angles and altitudes. We then train an adversarial patch using data captured from specific vehicle targets in multiple scenarios to enhance its attack universality across various real-world aerial photography scenarios. Results from the vanishing attack experiments show that our method enhanced the attack success rate (ASR) against YOLOv3 by 37.62% and 19.95% on the UAV and VisDrone datasets, respectively, compared to the baseline. For the Faster R-CNN detector, the success rates increased by 11.63% and 14.76%. Additionally, we investigated the transferability of projection patches and their attack performance at different sizes and angles, with our adversarial examples consistently performing well.
Jiahao Cui 0004, Wang Guo, Haikuo Huang, Xun Lv, Haifeng Li 0007
IEEE Trans. Geosci. Remote. Sens.6
2024 Global-Local Coupled Style Transfer for Semantic Segmentation of Bitemporal Remote Sensing Images
abstract
Due to the different acquisition conditions, large variations in the feature distributions of two temporal domains generally exist, known as temporal domain shift. The temporal domain shift is primarily influenced by coupled dual-factor: global style variations (such as illumination and weather conditions) and local style variations (such as the inherent phenological properties of land cover classes). In this article, we first formulate the temporal domain shift problem as an issue of dual-factor coupled interference on feature distributions in remote sensing (RS) community. To address this issue, we propose a semantic-guided style transfer (SGST) framework seamlessly integrating global feature alignment with local feature semantic matching. We use an adaptive segmentation model to provide pseudosegmentation maps and feed them into the style transfer model as semantic guidance. Under semantic guidance, a semantic-constrained style normalization (SCSN) module is designed to achieve style transfer at both global and local levels. Furthermore, a dual learning approach is introduced to make the style transfer model and the adaptive segmentation model promote each other. As a result, the style transfer model generates high-quality style-transferred images, and the adaptive segmentation model progressively predicts more accurate pseudosegmentation maps. Extensive experiments demonstrate the superiority of our proposed framework over state-of-the-art methods in terms of both perceptual quality and quantitative performance.
Hao Wang 0069, Mingning Guo, Shaoxian Li, Haifeng Li 0007, Chao Tao 0001
IEEE Trans. Geosci. Remote. Sens.4
2024 Adaptive Multiscale Slimming Network Learning for Remote Sensing Image Feature Extraction
abstract
Effective feature representation is pivotal in numerous remote sensing image (RSI) interpretation tasks. Notably, a distinct attribute of RSIs is their inclination toward multiscale feature dependence. Previous research predominantly focuses on designing intricate and complex networks or modules to encapsulate rich multiscale features. However, these approaches compromise on either the model’s compactness or its representational efficacy, thereby constraining the practical deployment of remote sensing technologies, particularly in limited-capacity environments like small-scale devices or on-orbit satellites. In this study, we explore the problem of how to augment the diversity of encoded features while avoiding heavy parameter scale growth in deep convolutional neural networks (CNNs). We proposed an adaptive multiscale framework RISV which presents two key features: 1) rich scale information: during training, RISV decomposes each convolutional layer into various-sized convolutions, extracting multiscale characteristics; and 2) small model volume: RISV incorporates a differentiable elect layer after each convolutional layer, adaptively calculating and polarizing channel importance during learning. After training, the added convolution kernel and the significant channels selected by the elect layer will be linearly equivalent merged, minimizing the impact of pruning on the model’s feature extraction capability. Different from traditional model slimming, it focused on a slimmed-down network while enhancing the representation of multiscale features in RSIs. Versatile and adaptable across various model frameworks like VGG and ResNet. Experimental results demonstrate that our methodology not only preserves accuracy across standard skeletal frameworks but also attains a compression ratio exceeding 80%, surpassing the baseline by an average of 40%. Furthermore, the application of GradCAM on the NWPU dataset reveals our method’s proficiency in acquiring detailed and accurate subject information from RSIs. The source code can be available athttps://github.com/GeoX-Lab/RISV.
Dingqi Ye, Jian Peng 0009, Wang Guo, Haifeng Li 0007
IEEE Trans. Geosci. Remote. Sens.4
2024 Augmentation-Free Graph Contrastive Learning of Invariant-Discriminative Representations
abstract
Graph contrastive learning (GCL) is a promising direction toward alleviating the label dependence, poor generalization and weak robustness of graph neural networks, learning representations with invariance, and discriminability by solving pretasks. The pretasks are mainly built on mutual information estimation, which requires data augmentation to construct positive samples with similar semantics to learn invariant signals and negative samples with dissimilar semantics to empower representation discriminability. However, an appropriate data augmentation configuration depends heavily on lots of empirical trials such as choosing the compositions of data augmentation techniques and the corresponding hyperparameter settings. We propose an augmentation-free GCL method, invariant-discriminative GCL (iGCL), that does not intrinsically require negative samples. iGCL designs the invariant-discriminative loss (ID loss) to learn invariant and discriminative representations. On the one hand, ID loss learns invariant signals by directly minimizing the mean square error (MSE) between the target samples and positive samples in the representation space. On the other hand, ID loss ensures that the representations are discriminative by an orthonormal constraint forcing the different dimensions of representations to be independent of each other. This prevents representations from collapsing to a point or subspace. Our theoretical analysis explains the effectiveness of ID loss from the perspectives of the redundancy reduction criterion, canonical correlation analysis (CCA), and information bottleneck (IB) principle. The experimental results demonstrate that iGCL outperforms all baselines on five node classification benchmark datasets. iGCL also shows superior performance for different label ratios and is capable of resisting graph attacks, which indicates that iGCL has excellent generalization and robustness. The source code is available at https://github.com/lehaifeng/ T-GCN/tree/master/iGCL.
Haifeng Li 0007, Qinyao Luo, Silu He, Xuying Wang
IEEE Trans. Neural Networks Learn. Syst.1
2024 Lifelong Learning With Cycle Memory Networks
abstract
Learning from a sequence of tasks for a lifetime is essential for an agent toward artificial general intelligence. Despite the explosion of this research field in recent years, most work focuses on the well-known catastrophic forgetting issue. In contrast, this work aims to explore knowledge-transferable lifelong learning without storing historical data and significant additional computational overhead. We demonstrate that existing data-free frameworks, including regularization-based single-network and structure-based multinetwork frameworks, face a fundamental issue of lifelong learning, named anterograde forgetting, i.e., preserving and transferring memory may inhibit the learning of new knowledge. We attribute it to the fact that the learning network capacity decreases while memorizing historical knowledge and conceptual confusion between the irrelevant old knowledge and the current task. Inspired by the complementary learning theory in neuroscience, we endow artificial neural networks with the ability to continuously learn without forgetting while recalling historical knowledge to facilitate learning new knowledge. Specifically, this work proposes a general framework named cycle memory networks (CMNs). The CMN consists of two individual memory networks to store short- and long-term memories separately to avoid capacity shrinkage and a transfer cell between them. It enables knowledge transfer from the long-term to the short-term memory network to mitigate conceptual confusion. In addition, the memory consolidation mechanism integrates short-term knowledge into the long-term memory network for knowledge accumulation. We demonstrate that the CMN can effectively address the anterograde forgetting on several task-related, task-conflict, class-incremental, and cross-domain benchmarks. Furthermore, we provide extensive ablation studies to verify each framework component. The source codes are available at: https://github.com/GeoX-Lab/CMN.
Jian Peng 0009, Dingqi Ye, Bo Tang 0011, Yinjie Lei, Yu Liu 0003, Haifeng Li 0007
IEEE Trans. Neural Networks Learn. Syst.6
2023 Continual Learning for Remote Sensing Image Scene Classification With Prompt Learning
abstract
Overcoming catastrophic forgetting is a key difficulty for remote sensing image (RSI) classification in open world applications. The core of this problem lies in the ability of RSI scene classification models to adapt to the changing environment and maintain the learned knowledge while continually learning new knowledge. Mainstream replay-based approaches overcome catastrophic forgetting by reenacting and retracing past experiences in the process of learning new data. However, such approaches rely heavily on the storage of historical data, and the recent rise of new paradigms based on prompt learning offers a new perspective of using only task-related “instructions” (i.e., prompts) to guide the model’s continual learning and reasoning. Therein, the task knowledge encoded by the prompt improves the model’s ability to overcome forgetting while reducing the amount of data and model parameters required by traditional data-driven approaches. Therefore, we propose a continual learning method based on prompt learning for RSI classification. We systematically analyze and reveal the potential of prompt learning for continual learning of RSI classification. Experiments on three publicly available remote sensing datasets show that prompt learning significantly outperforms two comparable methods on 3, 6, and 9 tasks, with an average accuracy (ACC) improvement of approximately 43%. Performance improvements of 4% to 6% were achieved when compared to advanced prototype network methods. We found that prompt-generation strategies and prompt-related components significantly affect performance: (1) prompt-generation strategies are strongly correlated with the model’s performance in overcoming catastrophic forgetting; (2) prompt-related components are correlated with remote sensing images of different scales. The new paradigm of prompt learning potentially provides a new idea for the continual learning problem of RSI classification.
Ling Zhao 0005, Linrui Xu, Dingqi Ye, Jian Peng 0009, Haifeng Li 0007
IEEE Geosci. Remote. Sens. Lett.8
2023 MSINet: Mining scale information from digital surface models for semantic segmentation of aerial images
Chengli Peng, Haifeng Li 0007, Chao Tao 0001, Yansheng Li 0001, Jiayi Ma 0001
Pattern Recognit.2
2023 Memory-Contrastive Unsupervised Domain Adaptation for Building Extraction of High-Resolution Remote Sensing Imagery
abstract
Deep learning-based semantic segmentation has been widely applied for building extraction. However, due to the domain gap, the extraction of building in high-resolution remote sensing imagery is difficult when the model trained on a source dataset is directly used to test on a target data. Considering that humans can retrieve memory to deal with correlative tasks in different domains, memory mechanisms have been developed effectively to assist cross-domain feature extraction. However, whether the memory mechanisms can achieve satisfactory result or not highly depends on the premise that the memory is relevant to the task. Therefore, the domain-invariant memory is crucial in cross-domain building extraction task. To this end, a memory-contrastive unsupervised domain adaptation method is proposed on the basis of a novel memory mechanism. Specifically, to facilitate the model to memorize domain-invariant features, we first conduct a normalization-based image style transfer strategy and a discriminator-based adversarial method at the image level and feature level, respectively. Subsequently, we carry out a memory-contrastive module to obtain domain-invariant features. Especially, a teacher–student network is exploited to help knowledge transferring by knowledge distillation to enhance the performance of the memory-contracted module. To narrow the distance between the two domains, a memory bank is designed to store and update category features obtained from the source domain, and then the similarity between category features in the target domain and memory bank is calculated. Results of the cross-domain experiments show that the proposed method can achieve optimal building extraction. (GitHub: https://github.com/RS-CSU/MDANet).
Jie Chen 0048, Peien He, Jingru Zhu, Ya Guo 0002, Geng Sun 0005, Haifeng Li 0007
IEEE Trans. Geosci. Remote. Sens.7
2023 Stable Prototype-Guided Single-Temporal Supervised Learning for Change Detection and Extraction of Building
abstract
Change detection and extraction of buildings based on convolutional neural networks (CNNs) have made encouraging progress in the remote sensing community. Although these two tasks are different in objective and application scenarios, both focus on building objects. However, previous methods were accustomed to considering these two tasks separately, and the change detection task suffered from the precondition that bitemporal labeled images were used as paired supervision signals. In this study, we propose a stable prototype guided single-temporal supervised learning framework (PGLF) as a joint solution for building change detection and cross-temporal extraction by exploring two cores: knowledge commonality and task specificity. For knowledge commonality, we introduced a multi-prototype representation module (MPRM) to generate stable building prototypes from support foreground features and designed a prototype to query feature adaptive fusion (PQAF) module to suppress background noise and extract discriminative building features in a way that support prototypes-guided query feature enhancement. For task specificity, we designed a multiscale spatiotemporal interaction module (MSTI) to capture bidirectional change features with strong spatiotemporal correlations. Besides, we developed a pseudo bitemporal image pair construction method to improve the performance of building change detection and cross-temporal extraction under single-temporal supervision signals. PGLF was trained on pseudo bitemporal labeled image pairs and tested on the public aerial WHU building dataset and proposed satellite SPO building dataset. The comprehensive experimental results demonstrate the superiority of the proposed method.
Shasha Hou, Guo Zhang 0001, Hao Cui 0002, Haifeng Li 0007
IEEE Trans. Geosci. Remote. Sens.6
2023 Self-Supervised Remote Sensing Feature Learning: Learning Paradigms, Challenges, and Future Works
abstract
Deep learning has achieved great success in learning features from massive remote sensing images (RSIs). To better understand the connection between three feature learning paradigms, which are unsupervised feature learning (USFL), supervised feature learning (SFL), and self-supervised feature learning (SSFL), this paper analyzes and compares them from the perspective of feature learning signals, and gives a unified feature learning framework. Under this unified framework, we analyze the advantages of SSFL over the other two learning paradigms in RSI understanding tasks and give a comprehensive review of existing SSFL works in RS, including the pre-training dataset, self-supervised feature learning signals, and the evaluation methods. We further analyze the effects of SSFL signals and pre-training data on the learned features to provide insights into RSI feature learning. Finally, we briefly discuss some open problems and possible research directions.
Chao Tao 0001, Ji Qi 0001, Mingning Guo, Qing Zhu 0012, Haifeng Li 0007
IEEE Trans. Geosci. Remote. Sens.5
2023 GraSS: Contrastive Learning With Gradient-Guided Sampling Strategy for Remote Sensing Image Semantic Segmentation
abstract
Self-supervised contrastive learning (SSCL) has achieved significant milestones in remote sensing image (RSI) understanding. Its essence lies in designing an unsupervised instance discrimination pretext task to extract image features from a large number of unlabeled images that are beneficial for downstream tasks. However, existing instance discrimination based SSCL suffers from two limitations when applied to the RSI semantic segmentation task: 1) Positive sample confounding issue, SSCL treats different augmentations of the same RSI as positive samples, but the richness, complexity, and imbalance of RSI ground objects lead to the model actually pulling a variety of different ground objects closer while pulling positive samples closer, which confuse the feature of different ground objects. 2) Feature adaptation bias, SSCL treats RSI patches containing various ground objects as individual instances for discrimination and obtains instance-level features, which are not fully adapted to pixel-level or object-level semantic segmentation tasks. To address the above limitations, we consider constructing samples containing single ground objects to alleviate positive sample confounding issue, and make the model obtain object-level features from the contrastive between single ground objects. Meanwhile, we observed that the discrimination information can be mapped to specific regions in RSI through the gradient of unsupervised contrastive loss, these specific regions tend to contain single ground objects. Based on this, we propose contrastive learning with Gradient guided Sampling Strategy (GraSS) for RSI semantic segmentation. GraSS consists of two stages: 1) the instance discrimination warm-up stage to provide initial discrimination information to the contrastive loss gradients, 2) the gradient guided sampling contrastive training stage to adaptively construct samples containing more singular ground objects using the discrimination information. Experimental results on three open datasets demonstrate that GraSS effectively enhances the performance of SSCL in high-resolution RSI semantic segmentation. Compared to eight baseline methods from six different types of SSCL, GraSS achieves an average improvement of 1.57% and a maximum improvement of 3.58% in terms of mean intersection over the union. Additionally, we discovered that the unsupervised contrastive loss gradients contain rich feature information, which inspires us to utilize gradient information more extensively during model training to attain additional model capacity. The source code is available at https://github.com/GeoX-Lab/GraSS.
Chao Tao 0001, Yunsheng Zhang 0001, Chengli Peng, Haifeng Li 0007
IEEE Trans. Geosci. Remote. Sens.6
2022 Deformation and Correspondence Aware Unsupervised Synthetic-to-Real Scene Flow Estimation for Point Clouds
abstract
Point cloud scene flow estimation is of practical importance for dynamic scene navigation in autonomous driving. Since scene flow labels are hard to obtain, current methods train their models on synthetic data and transfer them to real scenes. However, large disparities between existing synthetic datasets and real scenes lead to poor model transfer. We make two major contributions to address that. First, we develop a point cloud collector and scene flow annotator for GTA-V engine to automatically obtain diverse realistic training samples without human intervention. With that, we develop a large-scale synthetic scene flow dataset GTA-SF. Second, we propose a mean-teacher-based domain adaptation framework that leverages self-generated pseudo-labels of the target domain. It also explicitly incorporates shape deformation regularization and surface correspondence refinement to address distortions and misalignments in domain transfer. Through extensive experiments, we show that our GTA-SF dataset leads to a consistent boost in model generalization to three real datasets (i.e., Waymo, Lyft and KITTI) as compared to the most widely used FT3D dataset. Moreover, our framework achieves superior adaptation performance on six source-target dataset pairs, remarkably closing the average domain gap by 60%. Data and codes are available at https://github.com/leolyj/DCA-SRSFE
Yinjie Lei, Naveed Akhtar, Haifeng Li 0007, Munawar Hayat
CVPR4
2022 A data-driven adversarial examples recognition framework via adversarial feature genomes
abstract
Adversarial examples pose many security threats to convolutional neural networks (CNNs). Most defense algorithms prevent these threats by finding differences between the original images and adversarial examples. However, the found differences do not contain features about the classes, so these defense algorithms can only detect adversarial examples without recovering the correct labels. In this regard, we propose the Adversarial Feature Genome (AFG), a novel type of data that contain both the differences and features about classes. This method is inspired by an observed phenomenon, namely, the Adversarial Feature Separability, where the difference between the feature maps of the original images and adversarial examples becomes larger with deeper layers. On top of that, we further develop an adversarial example recognition framework that detects adversarial examples and can recover the correct labels. In the experiments, the detection and classification of adversarial examples by AFGs has an accuracy of more than 90.01% in various attack scenarios. To the best of our knowledge, our method is the first method that focuses on both attack detecting and recovering. AFG gives a new data-driven perspective to improve the robustness of CNNs.
Li Chen 0025, Qi Li 0031, Weiye Chen, Haifeng Li 0007
Int. J. Intell. Syst.5
2022 Curvature graph neural network
Haifeng Li 0007, Yu Liu 0003, Qing Zhu 0012, Guohua Wu 0001
Inf. Sci.1
2022 Feature pyramid-based graph convolutional neural network for graph classification
Mingming Lu, Zhixiang Xiao, Haifeng Li 0007, Naixue Xiong
J. Syst. Archit.3
2022 EFCNet: Ensemble Full Convolutional Network for Semantic Segmentation of High-Resolution Remote Sensing Images
abstract
Convolutional neural networks (CNNs) have achieved remarkable results in semantic segmentation of high-resolution remote sensing images (HRRSIs). However, the scales and textures of HRRSIs are diverse, which makes it difficult for a fixed-layer CNN to obtain rich features. In this regard, we propose an end-to-end ensemble fully convolutional network (EFCNet), which mainly includes two modules: the adaptive fusion module (AFM) and the separable convolutional module (SCM). The AFM can fuse features of different scales based on ensemble learning, whereas the SCM can reduce the complexity of the model under multifeature fusion. In the experiment, we use UNet and PSPNet to verify the framework on the ISPRS Vaihingen and Potsdam datasets. The experimental results show that the EFCNet can effectively improve the final segmentation performance and reduce the complexity of the ensemble model.
Li Chen 0025, Xin Dou, Jian Peng 0009, Wenbo Li 0004, Bingyu Sun, Haifeng Li 0007
IEEE Geosci. Remote. Sens. Lett.6
2022 Lie to Me: A Soft Threshold Defense Method for Adversarial Examples of Remote Sensing Images
abstract
Adversarial examples fool the models into predicting wrong results through generated perturbations, demonstrating the vulnerability of convolutional neural networks (CNNs). Recent studies also show that many CNNs applied to remote sensing image (RSI) scene classification are still subject to adversarial example attacks. Through further analysis of adversarial examples of RSIs, it is found that the misclassified classes are not random, and these adversarial examples have demonstrated attack selectivity. Based on this finding, we propose a soft threshold defense method. First, we take the images with the correct prediction of each class as positive samples and adversarial examples as negative samples. Then, we take their output confidence as input and get the decision boundaries by the logistic regression algorithm. Finally, the confidence threshold of each class can be further obtained based on the decision boundary. It is the soft threshold used for defense, which can determine whether the image is an adversarial example or not. When the model predicts the new RSI, the input is an original image if the output confidence is higher than the soft threshold of the corresponding class, and the opposite is an adversarial example. Our proposed algorithm does not require modification of the model structure and is computationally uncomplicated, and it is simple and effective. For the FGSM, BIM, Deepfool, and C&W attack algorithms, their fooling rates are reduced by an average of 97.76%, 99.77%, 68.18%, and 97.95% in several scenarios. The soft threshold defense method can effectively defend against adversarial examples.
Li Chen 0025, Pu Zou, Haifeng Li 0007
IEEE Geosci. Remote. Sens. Lett.4
2022 Generating Multiscale Maps From Satellite Images via Series Generative Adversarial Networks
abstract
Considering the success of generative adversarial networks (GANs) for image-to-image translation, researchers have attempted to translate satellite images to maps (si2map) through GAN for cartography. However, these studies involved limited scales, which hinders multiscale map creation. By extending their method, high-resolution satellite images can be trivially translated to multiscale maps through scale-wise si2map generators trained for certain scales. However, this strategy has two theoretical limitations. First, inconsistency between high-resolution satellite images and object generalization on multiscale maps (SI-M inconsistency) increasingly complicates the extraction of geographical information from satellite images for generators with decreasing scale. Second, as si2map translation is cross-domain, generators incur high computation costs to transform the pixel distribution on satellite images to that on maps. Thus, we designed a series strategy of generators for multiscale si2map translation to address these limitations. In this strategy, high-resolution satellite images are inputted to an si2map generator to output large-scale maps, which are translated to multiscale maps through series multiscale map generators. The series strategy avoids SI-M inconsistency as high-resolution satellite images are only translated to large-scale maps and transforms cross-domain translation to approximately intradomain translation when generating multiscale maps. Our experimental results showed better quality multiscale map generation with the series strategy, as shown by average increases of 11.69%, 53.78%, 55.42%, and 72.34% in the structural similarity index (SSIM), edge structural similarity index (ESSI), intersection over union (road), and intersection over union (water) for data from Mexico City and Tokyo at zoom levels 17–13.
Xu Chen 0042, Bangguo Yin, Songqiang Chen, Haifeng Li 0007
IEEE Geosci. Remote. Sens. Lett.4
2022 Message-Passing-Driven Triplet Representation for Geo-Object Relational Inference in HRSI
abstract
A high-resolution remote sensing image (HRSI) scene typically contains multiple geo-objects, and geospatial relations among these geo-objects are obvious. As the important information conveyed by HRSI, the intelligent expression of geospatial relation is helpful in understanding HRSI scenes. Previous HRSI semantic understanding was mainly based on image captions that only generate one sentence to describe image content, thereby resulting in insufficient understanding of the scene. Thus, the present letter proposes an approach to represent geospatial relations in an HRSI scene with structured form of$\langle $subject, geospatial relation, object$\rangle $. A geospatial relation triplet representation data set that contains visual and semantic information, such as category, location, and geospatial relations of the geo-objects, is constructed first. An “object-relation” message-passing mechanism is adopted to enhance the information exchange between the geo-objects and geospatial relations to predict triplets accurately. The experimental results show that the proposed method can effectively predict the geospatial relation in a HRSI scene.
Jie Chen 0048, Yi Zhang 0064, Geng Sun 0005, Haifeng Li 0007
IEEE Geosci. Remote. Sens. Lett.6
2022 MDANet: Unsupervised, Mixed-Domain Adaptation for Semantic Segmentation of Remote Sensing Images
abstract
The imaging process of optical remote sensing images are easily affected by external conditions. Therefore, remote sensing images under different imaging conditions often show color differences, resulting in feature distribution differences between the source and target domain, hindering the migration of semantic segmentation models between domains. Currently, most domain adaptation methods are for single-source and single-target domains. Here, we proposed a novel and concise method, coined MDANet, for the adaptation of patch images of multi-source and multi-target domains and for reducing the distribution differences of different patch images by projecting them onto the virtual center of a mixed-domain. MDANet is a lightweight and self-supervised network that can be grafted with any semantic segmentation model. Our method significantly improved the segmentation accuracy of semantic segmentation models and showed higher stability and competitiveness than existing methods.
Hao Cui 0002, Guo Zhang 0001, Ji Qi 0001, Haifeng Li 0007, Chao Tao 0001, Shasha Hou, DeRen Li
IEEE Geosci. Remote. Sens. Lett.4
2022 Spatial-Temporal Invariant Contrastive Learning for Remote Sensing Scene Classification
abstract
Self-supervised learning achieves close to supervised learning results on remote sensing image (RSI) scene classification. This is due to the current popular self-supervised learning methods that learn representations by applying different augmentations to images and completing the instance discrimination task which enables convolutional neural networks (CNNs) to learn representations invariant to augmentation. However, RSIs are spatial-temporal heterogeneous, which means that similar features may exhibit different characteristics in different spatial-temporal scenes. Therefore, the performance of CNNs that learn only representations invariant to augmentation still degrades for unseen spatial-temporal scenes due to the lack of spatial-temporal invariant representations. We propose a spatial-temporal invariant contrastive learning (STICL) framework to learn spatial-temporal invariant representations from unlabeled images containing a large number of spatial-temporal scenes. We use optimal transport to transfer an arbitrary unlabeled RSI into multiple other spatial-temporal scenes and then use STICL to make CNNs produce similar representations for the views of the same RSI in different spatial-temporal scenes. We analyze the performance of our proposed STICL on four commonly used RSI scene classification datasets, and the results show that our method achieves better performance on RSIs in unseen spatial-temporal scenes compared to popular self-supervised learning methods. Based on our findings, it can be inferred that spatial-temporal invariance is an indispensable property for a remote sensing model that can be applied to a wider range of remote sensing tasks, which also inspires the study of more general remote sensing models. The source code is available athttps://github.com/GeoX-Lab/G-RSIM/tree/main/T-SC-STICL.
Haozhe Huang, Zhongfeng Mou, Yunying Li, Qiujun Li, Jie Chen 0048, Haifeng Li 0007
IEEE Geosci. Remote. Sens. Lett.6
2022 Remote Sensing Image Scene Classification With Self-Supervised Paradigm Under Limited Labeled Samples
abstract
With the development of deep learning, supervised learning methods perform well in remote sensing image (RSI) scene classification. However, supervised learning requires a huge number of annotated data for training. When labeled samples are not sufficient, the most common solution is to fine-tune the pretraining models using a large natural image data set (e.g., ImageNet). However, this learning paradigm is not a panacea, especially when the target RSIs (e.g., multispectral and hyperspectral data) have different imaging mechanisms from RGB natural images. To solve this problem, we introduce a new self-supervised learning (SSL) mechanism to obtain the high-performance pretraining model for RSI scene classification from large unlabeled data. Experiments on three commonly used RSI scene classification data sets demonstrated that this new learning paradigm outperforms the traditional dominant ImageNet pretrained model. Moreover, we analyze the impacts of several factors in SSL on RSI scene classification, including the choice of self-supervised signals, the domain difference between the source and target data sets, and the amount of pretraining data. The insights distilled from this work can help to foster the development of SSL in the remote sensing community. Since SSL could learn from unlabeled massive RSIs, which are extremely easy to obtain, it will be a promising way to alleviate dependence on labeled samples and thus efficiently solve many problems, such as global mapping.
Chao Tao 0001, Ji Qi 0001, Weipeng Lu, Hao Wang 0069, Haifeng Li 0007
IEEE Geosci. Remote. Sens. Lett.5
2022 LaST: Label-Free Self-Distillation Contrastive Learning With Transformer Architecture for Remote Sensing Image Scene Classification
abstract
The increase in self-supervised learning (SSL), especially contrastive learning, has enabled one to train deep neural network models with unlabeled data for remote sensing image (RSI) scene classification. Nevertheless, it still suffers from the following issues. 1. The performance of the contrastive learning method is significantly impacted by the hard negative sample (HNS) issue, since the RSI scenario is complex in semantics and rich in surface features. 2. The multiscale characteristic of RSI is missed in the existing contrastive learning methods. 3. As the backbone of a deep learning model, especially in the case of limited annotation, a CNN does not include the adequate receptive field of convolutional kernels to capture the broad contextual information of RSI. In this regard, we propose label-free self-distillation contrastive learning with a transformer architecture (LaST). We introduce the self-distillation contrastive learning mechanism to address the HNS issue. Specifically, the LaST architecture comprises two modules: scale alignment with a multicrop module and a long-range dependency capture backbone module. In the former, we present global-local crop and scale alignment to encourage local-to-global correspondence and acquire multiscale relations. Then, the distorted views are fed into a transformer as a backbone, which is good at capturing the long-range-dependent contextual information of the RSI while maintaining the spatial smoothness of the learned features. Experiments on public datasets show that in the downstream scene classification task, LaST improves the performance of the self-supervised trained model by a maximum of 2.18% compared to the HNS-impacted contrastive learning approaches, and only 1.5% of labeled data can achieve the performance of supervised training CNNs with 10% labeled data. Moreover, this letter supports the integration of a transformer architecture and self-supervised paradigms in RSI interpretation.
Xuying Wang, Zhengliang Yan, Yunsheng Zhang 0001, Yansheng Chen, Haifeng Li 0007
IEEE Geosci. Remote. Sens. Lett.7
2022 FALSE: False Negative Samples Aware Contrastive Learning for Semantic Segmentation of High-Resolution Remote Sensing Image
abstract
Self-supervised contrastive learning (SSCL) is a potential learning paradigm for learning remote sensing image (RSI)-invariant features through the label-free method. The existing SSCL of RSI is built based on constructing positive and negative sample pairs. However, due to the richness of RSI ground objects and the complexity of the RSI contextual semantics, the same RSI patches have the coexistence and imbalance of positive and negative samples, which causing the SSCL pushing negative samples far away while pushing positive samples far away, and vice versa. We call this the sample confounding issue (SCI). To solve this problem, we propose a False negAtive sampLes aware contraStive lEarning model (FALSE) for the semantic segmentation of high-resolution RSIs. Since the SSCL pretraining is unsupervised, the lack of definable criteria for false negative sample (FNS) leads to theoretical undecidability, we designed two steps to implement the FNS approximation determination: coarse determination of FNS and precise calibration of FNS. We achieve coarse determination of FNS by the FNS self-determination (FNSD) strategy and achieve calibration of FNS by the FNS confidence calibration (FNCC) loss function. Experimental results on three RSI semantic segmentation datasets demonstrated that the FALSE effectively improves the accuracy of the downstream RSI semantic segmentation task compared with the current three models, which represent three different types of SSCL models. The mean Intersection-over-Union on ISPRS Potsdam dataset is improved by 0.7% on average; on CVPR DGLC dataset is improved by 12.28% on average; and on Xiangtan dataset this is improved by 1.17% on average. This indicates that the SSCL model has the ability to self-differentiate FNS and that the FALSE effectively mitigates the SCI in self-supervised contrastive learning.
Xuying Wang, Xiaoming Mei, Chao Tao 0001, Haifeng Li 0007
IEEE Geosci. Remote. Sens. Lett.5
2022 MFS: A Brain-Inspired Memory Formation System for GAN
abstract
Generative adversarial networks (GANs) are subject to catastrophic forgetting when learning stream of data. Inspired by the knowledge of neuroscience, this article develops a memory formation system (MFS) to establish memory for GANs. MFS is composed of three modules, including the identifier, weights upgrade (WU), and weights reactivate (WR). These modules simulate the function of encoding, consolidating, and retrieving in memory formation of human. Identifier creates indexes to label continuous tasks and these indexes are used as a cue when the corresponding tasks are recalled. WU and WR work as synaptic consolidation and system consolidation, respectively. In WU, a novel method, weight saliency measure (WSM) is proposed to measure the saliency of weight. Valuable weights are protected and invaluable weights are released for update when GAN is trained for the new task. In WR, traditional data-replay methods are improved by selecting${k}$representatives in each class and their feature vectors are restored. When pseudo data of this class are regenerated, only the samples with distance less than a threshold can be used in future training. Experimental results based on the testing of continual image generation and continual 3-D reconstruction show that MFS can establish a memory system to handle the catastrophic forgetting problem effectively.
Yifan Chang, Jian Peng 0009, Ziyi Dong, Haifeng Li 0007, Wenbo Li 0004
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2022 Contextual Information-Preserved Architecture Learning for Remote-Sensing Scene Classification
abstract
Convolutional neural networks (CNNs) have recently been widely used in remote-sensing scene classification. Additionally, it is becoming very popular to automatically learn specific CNN architectures for specific data sets. The rich contextual information in high-resolution remote-sensing images (RSIs) is critical to remote-sensing intelligent understanding tasks. However, architecture learning approaches tend to simplify the original data (i.e., resizing images to smaller resolution) for efficiency, yet result in contextual information loss of RSIs. In this article, we proposed a contextual information-preserved architecture learning (CIPAL) framework for remote-sensing scene classification to utilize the contextual information in RSIs as much as possible during the architecture learning process. We introduce channel compression into CIPAL, which can reduce the memory and time consumption of architecture learning and make it possible to construct a larger architecture space. We add potential operators that are rarely used for scene classification tasks (i.e., atrous convolution) into the architecture space to explore unknown architectures that are more suitable for remote-sensing scenes. The experimental results on four remote-sensing scene classification benchmarks indicate that CIPAL learns architectures with less time consumption than similar works, and the newly found architectures outperform popular hand-designed architectures for better use of contextual information in RSIs. Different architectures are good at learning different representations, and our proposed architecture learning method potentially helps us understand which types of representations are crucial for RSI intelligent understanding.
Jie Chen 0048, Haozhe Huang, Jian Peng 0009, Li Chen 0025, Chao Tao 0001, Haifeng Li 0007
IEEE Trans. Geosci. Remote. Sens.7
2022 MKN: Metakernel Networks for Few Shot Remote Sensing Scene Classification
abstract
Few-shot remote sensing scene classification tries to make a model quickly adapt to new scenes with only a few samples that do not appear in the closed training set. Since limited samples can hardly describe the distribution of data, it is a challenge for a model to learn good generalized features. Since limited samples are rarely representative, it is another challenge for a model to learn classification boundaries that depend on sample bias. Therefore, we propose a method called metakernel networks (MKNs) to solve the challenges via integrating a parametric linear classifier (PLC) into the metalearning framework to address the former problem assembling a metakernel strategy (MKS) and a stretching loss (Stloss) to address the later problem. The PLC learns prior knowledge from different tasks sampled from the same task family to learn rich features. The MKS is designed to remap low-dimensional indistinguishable features to a high-dimensional space to solve the low-dimensional feature entanglement caused by large intraclass differences between remote sensing images. The Stloss reduces the dependence of the hyperplane formed by a few points in each class on sample selection by reducing the intraclass and interclass variance ratios, thus solving the boundary fragility problem caused by interclass similarity. Experiments on three public datasets, UC_Merced, NWPU-RESISC45, and AID, show the state-of-the-art performance by the accuracy improvement of 5.04%, 4.81%, and 4.69% respectively. Our findings provide a new perspective by suggesting that the previously neglected issue of classification boundaries may be a key factor for few-shot remote sensing scene classification.
Zhenqi Cui, Li Chen 0025, Haifeng Li 0007
IEEE Trans. Geosci. Remote. Sens.4
2022 Thick Cloud Removal in Optical Remote Sensing Images Using a Texture Complexity Guided Self-Paced Learning Method
abstract
Thick clouds seriously impact the quality of optical remote sensing images (RSIs) and limit their application. For removing the cloud, some learning-based methods have been proposed and attracted considerable attention. However, these methods need to train paired multitemporal images with/without cloud, which are difficult and costly to collect. To solve this problem, we propose a novel texture complexity-guided self-paced learning (SPL) framework to remove the thick cloud from single RSIs. The framework does not need paired images and it exploits a texture complexity-guided mechanism to rank the self-generated cloud-corrupted training samples by texture complexity from low to high and then trains the generative adversarial cloud removal network using the SPL technique. In this way, the cloud removal network learns to restore the cloud-corrupted areas from easy to hard and thus to realize the image reconstruction for different difficulty levels. In addition, we introduce a structural similarity (SSIM) loss function to optimize the training network and improve the coherence of the image structure. Simulated and real experiments are performed on the single images acquired by Gao Fen-1 (GF-1) and Sentinel-2 satellites to validate the effectiveness of the proposed method. The results show that the proposed method has a better performance in cloud removal than other state-of-the-art methods, especially for the images of the areas with complex textures. The source codes are available athttps://github.com/GeoX-Lab/TPL.
Chao Tao 0001, Siyang Fu, Ji Qi 0001, Haifeng Li 0007
IEEE Trans. Geosci. Remote. Sens.4
2022 Avoiding Negative Transfer for Semantic Segmentation of Remote Sensing Images
abstract
Reducing the feature distribution shift caused by the factor of visual-environment changes, namely as VE-changes, is a hot issue in domain adaptation learning. However, in the semantic segmentation task of remote sensing imageries, besides VE-changes, the change of semantic-scenes (SS-changes) is another factor raising domain gap, which brings the label distribution shift. For example, although urban and rural share the same landcover label, there is still a gap in label distribution. If there is little relation that can be found in neither feature nor label space, forcibly adapting to a new domain could have a high risk of negative transfer. Hence, we propose a new Transitive Domain Adaptation method for Remote Sensing images (TDARS). Firstly, we introduce an intermediate domain to enlarge the relation between the given source and target domains. Secondly, we learn from primary and non-primary confident classes to increase the likelihood of transferring valuable information. As a result, TDARS enables the given source and target domains to be connected through the selected intermediate domain and performs effective knowledge transfer among all domains. The proposed method is evaluated on three domain adaptation datasets of remote sensing images. Extensive experiments show the approach can effectively handle the domain shift problem from remote sensing images compared to other state-of-the-art domain adaptation methods.
Hao Wang 0069, Chao Tao 0001, Ji Qi 0001, Haifeng Li 0007
IEEE Trans. Geosci. Remote. Sens.5
2022 Better Memorization, Better Recall: A Lifelong Learning Framework for Remote Sensing Image Scene Classification
abstract
To infer unknown remote sensing scenarios, most existing technologies use a supervised learning paradigm to train deep neural network (DNN) models on closed datasets. This paradigm faces challenges such as highly spatiotemporal variants and ever-changing scale-heterogeneous remote sensing scenarios. Additionally, DNN models cannot scale to new scenarios. Lifelong learning is an effective solution to these problems. Current lifelong learning approaches focus on overcoming thecatastrophic forgettingissue (i.e., a successive increase in heterogeneous remote sensing scenes causes models to forget historical scenes) while ignoring theknowledge recallissue (i.e., how to facilitate the learning of new scenes by recalling historical experiences), which is a significant problem. This paper proposes a lifelong learning framework called asymmetric collaborative network (SCN) for lifelong remote sensing image classification. This framework consists of two structurally distinct networks: a preserving network (Pres-Net) and a transient network (Trans-Net), which imitates the long- and short-term memory processes in the brain, respectively. Moreover, this framework is based on two synergistic knowledge transfer mechanisms: triple distillation and prior feature fusion. The triple distillation mechanism enables knowledge persistence from Trans-Net to Pres-Net to achieve better memorization; the prior feature fusion mechanism enables knowledge transfer from Pres-Net to Trans-Net to achieve better recall. Experiments on three open datasets demonstrate the effectiveness of SCN for 3-, 6-, and 9-task-length learning. The idea of asymmetric separation networks and the synergistic strategy proposed in this paper are expected to provide new solutions to the translatability of the classification of remote sensing images in real world scenarios.
Dingqi Ye, Jian Peng 0009, Haifeng Li 0007, Lorenzo Bruzzone
IEEE Trans. Geosci. Remote. Sens.3
2022 KST-GCN: A Knowledge-Driven Spatial-Temporal Graph Convolutional Network for Traffic Forecasting
abstract
While considering the spatial and temporal features of traffic, capturing the impacts of various external factors on travel is an essential step towards achieving accurate traffic forecasting. However, existing studies seldom consider external factors or neglect the effect of the complex correlations among external factors on traffic. Intuitively, knowledge graphs can naturally describe these correlations. Since knowledge graphs and traffic networks are essentially heterogeneous networks, it is challenging to integrate the information in both networks. On this background, this study presents a knowledge representation-driven traffic forecasting method based on spatial-temporal graph convolutional networks. We first construct a knowledge graph for traffic forecasting and derive knowledge representations by a knowledge representation learning method named KR-EAR. Then, we propose the Knowledge Fusion Cell (KF-Cell) to combine the knowledge and traffic features as the input of a spatial-temporal graph convolutional backbone network. Experimental results on the real-world dataset show that our strategy enhances the forecasting performances of backbones at various prediction horizons. The ablation and perturbation analysis further verify the effectiveness and robustness of the proposed method. To the best of our knowledge, this is the first study that constructs and utilizes a knowledge graph to facilitate traffic forecasting; it also offers a promising direction to integrate external information and spatial-temporal information for traffic forecasting. The source code is available athttps://github.com/lehaifeng/T-GCN/tree/master/KST-GCN.
Xing Han, Hanhan Deng, Chao Tao 0001, Ling Zhao 0005, Pu Wang 0005, Tao Lin 0008, Haifeng Li 0007
IEEE Trans. Intell. Transp. Syst.8
2022 Overcoming Long-Term Catastrophic Forgetting Through Adversarial Neural Pruning and Synaptic Consolidation
abstract
Enabling a neural network to sequentially learn multiple tasks is of great significance for expanding the applicability of neural networks in real-world applications. However, artificial neural networks face the well-known problem of catastrophic forgetting. What is worse, the degradation of previously learned skills becomes more severe as the task sequence increases, known as the long-term catastrophic forgetting. It is due to two facts: first, as the model learns more tasks, the intersection of the low-error parameter subspace satisfying for these tasks becomes smaller or even does not exist; second, when the model learns a new task, the cumulative error keeps increasing as the model tries to protect the parameter configuration of previous tasks from interference. Inspired by the memory consolidation mechanism in mammalian brains with synaptic plasticity, we propose a confrontation mechanism in which Adversarial Neural Pruning and synaptic Consolidation (ANPyC) is used to overcome the long-term catastrophic forgetting issue. The neural pruning acts as long-term depression to prune task-irrelevant parameters, while the novel synaptic consolidation acts as long-term potentiation to strengthen task-relevant parameters. During the training, this confrontation achieves a balance in that only crucial parameters remain, and non-significant parameters are freed to learn subsequent tasks. ANPyC avoids forgetting important information and makes the model efficient to learn a large number of tasks. Specifically, the neural pruning iteratively relaxes the current task's parameter conditions to expand the common parameter subspace of the task; the synaptic consolidation strategy, which consists of a structure-aware parameter-importance measurement and an element-wise parameter updating strategy, decreases the cumulative error when learning new tasks. Our approach encourages the synapse to be sparse and polarized, which enables long-term learning and memory. ANPyC exhibits effectiveness and generalization on both image classification and generation tasks with multiple layer perceptron, convolutional neural networks, and generative adversarial networks, and variational autoencoder. The full source code is available at https://github.com/GeoX-Lab/ANPyC.
Jian Peng 0009, Bo Tang 0011, Hao Jiang 0020, Yinjie Lei, Tao Lin 0008, Haifeng Li 0007
IEEE Trans. Neural Networks Learn. Syst.7
2022 Bottom-Up Mechanism and Improved Contract Net Protocol for Dynamic Task Planning of Heterogeneous Earth Observation Resources
abstract
Earth observation resources are becoming increasingly indispensable in disaster relief, damage assessment, and other related domains. Many unpredictable factors, such as changes in observation task requirements, bad weather, and resource malfunctions, may cause the scheduled observation scheme to become infeasible. In these cases, it is crucial to promptly reformulate high-quality observation schemes while exerting minimal negative effects on the previously scheduled tasks. Accordingly, in this study, a bottom-up distributed coordination framework together with an improved contract net is proposed, aiming to facilitate dynamic task replanning for heterogeneous Earth observation resources. This hierarchical framework consists of three levels: 1) neighboring resource coordination; 2) single planning center coordination; and 3) multiple planning center coordination. The observation tasks affected by unpredicted factors are managed along with a bottom-up route from resources to planning centers. This bottom-up distributed coordination framework transfers part of the computing load to various nodes of the observation systems to plan tasks more efficiently and robustly. To support the prompt replanning of multiple tasks to proper Earth observation resources in dynamic environments, we propose a multiround combinatorial allocation (MCA) method. Moreover, a new float interval-based local search algorithm is proposed to quickly obtain a promising replanning scheme. The simulation results demonstrate that the MCA method can achieve a better task completion rate for large-scale tasks with satisfactory time efficiency. In addition, this method can efficiently obtain replanning schemes based on original schemes in dynamic environments.
Baoju Liu, Guohua Wu 0001, Xinyu Pei, Haifeng Li 0007, Witold Pedrycz
IEEE Trans. Syst. Man Cybern. Syst.5
2021 A method to evaluate task-specific importance of spatio-temporal units based on explainable artificial intelligence
abstract
Big geo-data are often aggregated according to spatio-temporal units for analyzing human activities and urban environments. Many applications categorize such data into groups and compare the characteristics across groups. The intergroup differences vary with spatio-temporal units, and the essential is to identify the spatio-temporal units with apparently different data characteristics. However, spatio-temporal dependence, data variety, and the complexity of tasks impede an effective unit assessment. Inspired by the applications to extract critical image components based on explainable artificial intelligence (XAI), we propose a spatio-temporal layer-wise relevance propagation method to assess spatio-temporal units as a general solution. The method organizes input data into an extensible three-dimensional tensor form. We provide two means of labeling the spatio-temporal tensor data for typical geographical applications, using temporally or spatially relevant information. Neural network training proceeds to extract the global and local characteristics of data for corresponding analytical tasks. Then the method propagates classification results backward into units as obtained task-specific importance. A case study with taxi trajectory data in Beijing validates the method. The results prove that the proposed method can evaluate the task-specific importance of spatio-temporal units with dependence. This study also attempts to discover task-related knowledge using XAI.
Ximeng Cheng, Haifeng Li 0007, Yi Zhang 0064, Lun Wu, Yu Liu 0003
Int. J. Geogr. Inf. Sci.3
2021 SCAttNet: Semantic Segmentation Network With Spatial and Channel Attention Mechanism for High-Resolution Remote Sensing Images
abstract
High-resolution remote sensing images (HRRSIs) contain substantial ground object information, such as texture, shape, and spatial location. Semantic segmentation, which is an important task for element extraction, has been widely used in processing mass HRRSIs. However, HRRSIs often exhibit large intraclass variance and small interclass variance due to the diversity and complexity of ground objects, thereby bringing great challenges to a semantic segmentation task. In this letter, we propose a new end-to-end semantic segmentation network, which integrates lightweight spatial and channel attention modules that can refine features adaptively. We compare our method with several classic methods on the ISPRS Vaihingen and Potsdam data sets. Experimental results show that our method can achieve better semantic segmentation results. The source codes are available at https://github.com/lehaifeng/SCAttNet.
Haifeng Li 0007, Kaijian Qiu, Li Chen 0025, Xiaoming Mei, Chao Tao 0001
IEEE Geosci. Remote. Sens. Lett.1
2021 SMAPGAN: Generative Adversarial Network-Based Semisupervised Styled Map Tile Generation Method
abstract
Traditional online map tiles, which are widely used on the Internet, such as by Google Maps and Baidu Maps, are rendered from vector data. The timely updating of online map tiles from vector data, for which generation is time-consuming, is a difficult mission. Generating map tiles over time from remote sensing images is relatively simple and can be performed quickly without vector data. However, this approach used to be challenging or even impossible. Inspired by image-to-image translation (img2img) techniques based on generative adversarial networks (GANs), we proposed a semisupervised generation of styled map tiles based on the GANs (SMAPGAN) model to generate styled map tiles directly from remote sensing images. In this model, we designed a semisupervised learning strategy to pretrain SMAPGAN on rich unpaired samples and fine-tune it on limited paired samples in reality. We also designed the image gradient L1 loss and the image gradient structure loss to generate a styled map tile with global topological relationships and detailed edge curves for objects, which are important in cartography. Moreover, we proposed the edge structural similarity index (ESSI) as a metric to evaluate the quality of the topological consistency between the generated map tiles and ground truth. The experimental results show that SMAPGAN outperforms state-of-the-art (SOTA) works according to the mean squared error, the structural similarity index, and the ESSI. Also, SMAPGAN gained higher approval than SOTA in a human perceptual test on the visual realism of cartography. Our work shows that SMAPGAN is a new tool with excellent potential for producing styled map tiles. Our implementation of SMAPGAN is available at https://github.com/imcsq/SMAPGAN.
Xu Chen 0042, Songqiang Chen, Bangguo Yin, Jian Peng 0009, Xiaoming Mei, Haifeng Li 0007
IEEE Trans. Geosci. Remote. Sens.7
2021 An Empirical Study of Adversarial Examples on Remote Sensing Image Scene Classification
abstract
Deep neural networks (DNNs), which learn a hierarchical representation of features, have shown remarkable performance in big data analytics of remote sensing. However, previous research indicates that DNNs are easily spoofed by adversarial examples, which are crafted images with artificial perturbations that fool DNN models toward wrong predictions. To comprehensively evaluate the impact of adversarial examples on the remote sensing image (RSI) scene classification, this study tests eight state-of-the-art classification DNNs on six RSI benchmarks. These data sets include both optical and synthetic-aperture radar (SAR) images of different spectral and spatial resolutions. In the experiment, we create 48 classification scenarios and use four cutting-edge attack algorithms to investigate the influence of the adversarial example on the classification of RSIs. The experimental result shows that the fooling rates of the attacks are all over 98% across the 48 scenarios. We also find that, for the optical data, the seriousness of the adversarial problem has a negative relationship with the richness of the feature information. Besides, adversarial examples generated from SAR images are used easily for fooling the models with an average fooling rate of 76.01%. By analyzing the class distribution of these adversarial examples, we find that the distribution of the misclassifications is not affected by the types of models and attack algorithms-adversarial examples of RSIs of the same class cluster on fixed several classes. The analysis of classes of adversarial examples not only helps us explore the relationships between data set classes but also provides insights for further designing defensive algorithms.
Li Chen 0025, Zewei Xu, Qi Li 0031, Jian Peng 0009, Shaowen Wang 0001, Haifeng Li 0007
IEEE Trans. Geosci. Remote. Sens.6
2021 RS-MetaNet: Deep Metametric Learning for Few-Shot Remote Sensing Scene Classification
abstract
Training a modern deep neural network on massive labeled samples is the main paradigm in solving the scene classification problem for remote sensing, but learning from only a few data points remains a challenge. Existing methods for a few-shot remote sensing scene classification are performed in a sample-level manner, resulting in easy overfitting of learned features to individual samples and inadequate generalization of learned category segmentation surfaces. To solve this problem, learning should be organized at the task level rather than the sample level. Learning on tasks sampled from a task family can help tune learning algorithms to perform well on new tasks sampled in that family. Therefore, we propose a simple but effective method, called RS-MetaNet, to resolve the issues related to few-shot remote sensing scene classification in the real world. On the one hand, RS-MetaNet raises the level of learning from the sample to the task by organizing training in a metaway, and it learns to learn a metric space that can well classify remote sensing scenes from a series of tasks. We also propose a new loss function, called balance loss, which maximizes the generalization ability of the model to new samples by maximizing the distance between different categories, providing the scenes in different categories with better linear segmentation planes while ensuring model fit. The experimental results on three open and challenging remote sensing data sets, UCMerced_LandUse, NWPU-RESISC45, and Aerial Image Data, demonstrate that our proposed RS-MetaNet method achieves state-of-the-art results in cases where there are only 1 ~ 20 labeled samples.
Haifeng Li 0007, Zhenqi Cui, Zhiqiang Zhu, Li Chen 0025, Haozhe Huang, Chao Tao 0001
IEEE Trans. Geosci. Remote. Sens.1
2021 MAP-Net: Multiple Attending Path Neural Network for Building Footprint Extraction From Remote Sensed Imagery
abstract
Building footprint extraction is a basic task in the fields of mapping, image understanding, computer vision, and so on. Accurately and efficiently extracting building footprints from a wide range of remote sensed imagery remains a challenge due to the complex structures, variety of scales, and diverse appearances of buildings. Existing convolutional neural network (CNN)-based building extraction methods are criticized for their inability to detect tiny buildings because the spatial information of CNN feature maps is lost during repeated pooling operations of the CNN. In addition, large buildings still have inaccurate segmentation edges. Moreover, features extracted by a CNN are always partially restricted by the size of the receptive field, and large-scale buildings with low texture are always discontinuous and holey when extracted. To alleviate these problems, multiscale strategies are introduced in the latest research works to extract buildings with different scales. The features with higher resolution generally extracted from shallow layers, which extracted insufficient semantic information for tiny buildings. This article proposes a novel multiple attending path neural network (MAP-Net) for accurately extracting multiscale building footprints and precise boundaries. Unlike existing multiscale feature extraction strategies, MAP-Net learns spatial localization-preserved multiscale features through a multiparallel path in which each stage is gradually generated to extract high-level semantic features with fixed resolution. Then, an attention module adaptively squeezes the channel-wise features extracted from each path for optimized multiscale fusion, and a pyramid spatial pooling module captures global dependence for refining discontinuous building footprints. Experimental results show that our method achieved 0.88%, 0.93%, and 0.45% F1-score and 1.53%, 1.50%, and 0.82% intersection over union (IoU) score improvements without increasing computational complexity compared with the latest HRNetv2 on the Urban 3-D, Deep Globe, and WHU data sets, respectively. Specifically, MAP-Net outperforms multiscale aggregation fully convolutional network (MA-FCN), which is the state-of-the-art (SOTA) algorithms with postprocessing and model voting strategies, on the WHU data set without pretraining and postprocessing. The TensorFlow implementation is available at https://github.com/lehaifeng/MAPNet.
Qing Zhu 0012, Han Hu 0005, Xiaoming Mei, Haifeng Li 0007
IEEE Trans. Geosci. Remote. Sens.5
2021 Hierarchical Paired Channel Fusion Network for Street Scene Change Detection
abstract
Street Scene Change Detection (SSCD) aims to locate the changed regions between a given street-view image pair captured at different times, which is an important yet challenging task in the computer vision community. The intuitive way to solve the SSCD task is to fuse the extracted image feature pairs, and then directly measure the dissimilarity parts for producing a change map. Therefore, the key for the SSCD task is to design an effective feature fusion method that can improve the accuracy of the corresponding change maps. To this end, we present a novel Hierarchical Paired Channel Fusion Network (HPCFNet), which utilizes the adaptive fusion of paired feature channels. Specifically, the features of a given image pair are jointly extracted by a Siamese Convolutional Neural Network (SCNN) and hierarchically combined by exploring the fusion of channel pairs at multiple feature levels. In addition, based on the observation that the distribution of scene changes is diverse, we further propose a Multi-Part Feature Learning (MPFL) strategy to detect diverse changes. Based on the MPFL strategy, our framework achieves a novel approach to adapt to the scale and location diversities of the scene change regions. Extensive experiments on three public datasets (i.e., PCD, VL-CMU-CD and CDnet2014) demonstrate that the proposed framework achieves superior performance which outperforms other state-of-the-art methods with a considerable margin.
Yinjie Lei, Duo Peng, Qiuhong Ke, Haifeng Li 0007
IEEE Trans. Image Process.5
2021 A Two-Phase Coordinated Planning Approach for Heterogeneous Earth-Observation Resources to Monitor Area Targets
abstract
Monitoring various types of disasters involves diversified requirements, such as the spectral band, resolution, and timeliness. However, at present, different types of observation platforms are separately operated. This isolated resource organization model is insufficient to meet the requirements of various Earth-observation tasks, especially when disasters occur. As a result, it is necessary to construct an Earth-observation network that contains space-air-ground observation resources and makes unified task planning for the included heterogeneous resources, such that the efficiency of the entire observation system is maximized. In this article, an architecture with two planning phases is proposed for the coordinated planning of heterogeneous Earth-observation resources, in which area targets and four types of space-air-ground observation resources [i.e., satellites, unmanned aerial vehicles (UAVs), airships, and ground monitoring vehicles] are considered. The two-phase approach in this architecture includes an area target decomposition phase and a task allocation phase. In the first phase, an area target hierarchical decomposition (ATHD) method is proposed to decompose the area targets into subtasks. In the second phase, a task conflict heuristic allocation (TCHA) method is proposed to allocate the decomposed subtasks to subplanning centers. Extensive experiments on simulated and realistic scenarios are conducted to verify the effectiveness of the proposed ATHD and TCHA methods. The computational results show that the ATHD method substantially improves the efficiency of the coordinated task planning process. Moreover, compared with traditional task allocation methods, the TCHA method could produce high-quality observation plans for the Earth-observation network, as it brings complementary benefits via the comprehensive usage of heterogeneous space-air-ground resources.
Baoju Liu, Sumin Li, RongHua Du, Guohua Wu 0001, Haifeng Li 0007, Ling Wang 0001
IEEE Trans. Syst. Man Cybern. Syst.6
2020 Solving large-scale many-objective optimization problems by covariance matrix adaptation evolution strategy with scalable small subpopulations
Huangke Chen, Ran Cheng 0004, Jinming Wen, Haifeng Li 0007, Jian Weng 0001
Inf. Sci.4
2020 T-GCN: A Temporal Graph Convolutional Network for Traffic Prediction
abstract
Accurate and real-time traffic forecasting plays an important role in the intelligent traffic system and is of great significance for urban traffic planning, traffic management, and traffic control. However, traffic forecasting has always been considered an “open” scientific issue, owing to the constraints of urban road network topological structure and the law of dynamic change with time. To capture the spatial and temporal dependences simultaneously, we propose a novel neural network-based traffic forecasting method, the temporal graph convolutional network (T-GCN) model, which is combined with the graph convolutional network (GCN) and the gated recurrent unit (GRU). Specifically, the GCN is used to learn complex topological structures for capturing spatial dependence and the gated recurrent unit is used to learn dynamic changes of traffic data for capturing temporal dependence. Then, the T-GCN model is employed to traffic forecasting based on the urban road network. Experiments demonstrate that our T-GCN model can obtain the spatio-temporal correlation from traffic data and the predictions outperform state-of-art baselines on real-world traffic datasets. Our tensorflow implementation of the T-GCN is available at https://www.github.com/lehaifeng/T-GCN.
Ling Zhao 0005, Yujiao Song, Yu Liu 0003, Pu Wang 0005, Tao Lin 0008, Haifeng Li 0007
IEEE Trans. Intell. Transp. Syst.8
2019 Semi-Supervised Variational Generative Adversarial Networks for Hyperspectral Image Classification
abstract
Though Hyperspectral Image (HSI) Classification has been extensively investigated over recent decades, it is still a challenge task especially when the number of labeled samples is extremely limited. In this paper, we overcome this challenge by using synthetic samples, and proposed a semi-supervised variational Generative Adversarial Networks(GANs) for this purpose. Compared to the conditional GAN which is recently used for generating samples for HSI classification, the proposed approach has two novel aspects. First, we extend the classic variational generative adversarial network to the semi-supervised context through an ensemble prediction technique. By this way, our model can be trained using limited labeled samples (only 5 samples per class) with a large number of unlabeled samples. Second, we adopt an encoder-decoder network to explicitly learn the relationship between the latent space and the real image space. This property enables our model producing diverse samples by simply varying some latent parameters, which is desirable for enriching the training dataset. We have shown that the proposed model can achieve better and robust performance for HSI classification compared to conditional GAN, especially when the labeled data is limited.
Hao Wang 0069, Chao Tao 0001, Ji Qi 0001, Haifeng Li 0007
IGARSS4
2018 Ensemble of differential evolution variants
Guohua Wu 0001, Xin Shen 0001, Haifeng Li 0007, Huangke Chen, Anping Lin, Ponnuthurai N. Suganthan
Inf. Sci.3
2016 Coordinated Planning of Heterogeneous Earth Observation Resources
abstract
Different Earth observation resources (EORs) [e.g., satellites, airships, and unmanned aerial vehicles (UAVs)] are usually managed by different organization sub-planners, which lack interactions and cooperation among one another. Such independent resource operations are no longer efficient to meet diverse and vast observation requests, especially in emergency situations, such as earthquakes, flooding, and forest fire disasters. This paper addresses the issue of coordinated planning of heterogeneous EORs, including satellites, airships, and UAVs. A hierarchical coordinated planning architecture is proposed to integrate heterogeneous EORs for the construction of a distributed and loosely coupled Earth observation system. The architecture comprises four component categories, namely, observation resource, sub-planner, coordination, and information management. Moreover, we propose two task assignment algorithms to coordinate and allocate observation tasks to sub-planners. The first algorithm is a highest-weight-first-allocated algorithm, and the second is a tabu-list-based simulated annealing (SA-TL) algorithm. Experiments and comparative studies demonstrate the efficiency of the coordinated planning architecture and SA-TL algorithm. We also show that the system responds dynamically to unexpected situations through effective disturbance-handling mechanisms.
Guohua Wu 0001, Witold Pedrycz, Haifeng Li 0007, Manhao Ma
IEEE Trans. Syst. Man Cybern. Syst.3
2014 Superior solution guided particle swarm optimization combined with local search techniques
Guohua Wu 0001, Dishan Qiu, Witold Pedrycz, Manhao Ma, Haifeng Li 0007
Expert Syst. Appl.6
2011 Non-cooperative Game Based QoS-Aware Web Services Composition Approach for Concurrent Tasks
abstract
Web services make tools which used to be merely accessible to the specialist available to all, and permitting previous manual data processing and analysis tasks to be automated. One of key problem is Web services composition in terms of Quality of Service (QoS). There are many task concurrencies, such as remote sensing image processing, in computation-intensive scientific applications. However, existing Web service optimal combination approaches are mainly focused on single tasks by using "selfish" behavior to pursue optimal solutions. This causes conflicts because many concurrent tasks are competing for limited optimal resources, and the reducing of service quality in services. Based on the best reply function of quantified task conflicts and game theory, this paper establishes a mathematical model to depict the competitive relationship between multitasks and Web service under QoS constraints and it guarantees that every task can obtain optimal utility services considering other task combination strategies. Moreover, an iterative algorithm to reach the Nash equilibrium is also proposed. Theory and experimental analysis show the approach has a fine convergence property, and can considerably enhance the actual utility of all tasks when compared with existing Web services combinatorial methods. The proposed approach provides a new path for QoS-aware Web service with optimal combinations for concurrent tasks.
Haifeng Li 0007, Qing Zhu 0012, Yiqiang Ouyang
ICWS1