Weikang Yu

dblp:303/8734 · DBLP profile ↗
← Back
12ranked-venue papers
5as first author
12since 2021 · last 2025
0000-0003-1111-572XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 11 · 5 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Auto-Prompting SAM for Weakly Supervised Landslide Extraction
abstract
Weakly supervised landslide extraction aims to identify landslide regions from remote sensing data using models trained with weak labels, particularly image-level labels. However, it is often challenged by the imprecise boundaries of the extracted objects due to the lack of pixel-wise supervision and the properties of landslide objects. To tackle these issues, we propose a simple yet effective method by auto-prompting the Segment Anything Model (SAM), i.e., APSAM. Instead of depending on high-quality class activation maps (CAMs) for pseudo-labeling or fine-tuning SAM, our method directly yields fine-grained segmentation masks from SAM inference through prompt engineering. Specifically, it adaptively generates hybrid prompts from the CAMs obtained by an object localization network. To provide sufficient information for SAM prompting, an adaptive prompt generation (APG) algorithm is designed to fully leverage the visual patterns of CAMs, enabling the efficient generation of pseudo-masks for landslide extraction. These informative prompts are able to identify the extent of landslide areas (box prompts) and denote the centers of landslide objects (point prompts), guiding SAM in landslide segmentation. Experimental results on high-resolution aerial and satellite datasets demonstrate the effectiveness of our method, achieving improvements of at least 3.0% in F1 score and 3.69% in IoU compared to other state-of-the-art methods. The source codes and datasets will be available at https://github.com/zxk688.
Xianping Ma, Weikang Yu, Pedram Ghamisi
IEEE Geosci. Remote. Sens. Lett.4
2024 MineNet-CD: Global Mining Change Detection Dataset
abstract
Mining change detection requires dedicated datasets because it includes unique objects such as pits or quarries, tailings dams, overburden, processing plants, haul roads, access roads, buildings/sheds, mining and blasting equipment, and plants, among others. This paper introduces a benchmark, large dataset for mining change detection, termed the MineNet-CD, to facilitate large-scale change detection. The proposed dataset contains a total of 100 high-resolution bi-temporal mining images from all over the world with corresponding ground truth. Unlike existing datasets, the images of MineNet-CD exhibit significant background, topological variation, and well-defined ground truth that ignores insignificant change maps. The work analyzes the efficacy of state-of-the-art deep learning methods for change detection. The results demonstrate that more efficient and advanced networks are required to accurately predict the change maps. The dataset and code are available at https://github.com/EricYu97/MineNet-CD.
Weikang Yu, Samiran Das, Aldino Rizaldy, Richard Gloaguen, Pedram Ghamisi
IGARSS1
2024 MaskCD: A Remote Sensing Change Detection Network Based on Mask Classification
abstract
Change detection (CD) from remote sensing (RS) images using deep learning has been widely investigated in the literature. It is typically regarded as a pixelwise labeling task that aims to classify each pixel as changed or unchanged. Although per-pixel classification networks in encoder-decoder structures have shown dominance, they still suffer from imprecise boundaries and incomplete object delineation at various scenes. For high-resolution RS images, partly or totally changed objects are more worthy of attention rather than a single pixel. Therefore, we revisit the CD task from the mask prediction and classification perspective and propose mask classification-based CD (MaskCD) to detect changed areas by adaptively generating categorized masks from input image pairs. Specifically, it utilizes a cross-level change representation perceiver (CLCRP) to learn multiscale change-aware representations and capture spatiotemporal relations from encoded features by exploiting deformable multihead self-attention (DeformMHSA). Subsequently, a masked cross-attention-based detection transformers (MCA-DETRs) decoder is developed to accurately locate and identify changed objects based on masked cross-attention and self-attention (SA) mechanisms. It reconstructs the desired changed objects by decoding the pixelwise representations into learnable mask proposals and making final predictions from these candidates. Experimental results on five benchmark datasets demonstrate the proposed approach outperforms other state-of-the-art models. Codes and pretrained models are available online at:https://github.com/EricYu97/MaskCD.
Weikang Yu, Samiran Das, Xiao Xiang Zhu 0001, Pedram Ghamisi
IEEE Trans. Geosci. Remote. Sens.1
2024 MineNetCD: A Benchmark for Global Mining Change Detection on Remote Sensing Imagery
abstract
Monitoring land changes triggered by mining activities is crucial for industrial control, environmental management, and regulatory compliance, yet it poses significant challenges due to the vast and often remote locations of mining sites. Remote sensing technologies have increasingly become indispensable to detect and analyze these changes over time. We thus introduce MineNetCD, a comprehensive benchmark designed for global mining change detection using remote sensing imagery. The benchmark comprises three key contributions. First, we establish a global mining change detection dataset featuring more than 70k paired patches of bitemporal high-resolution remote sensing images and pixel-level annotations from 100 mining sites worldwide. Second, we develop a novel baseline model based on a change-aware fast Fourier transform (ChangeFFT) module, which enhances various backbones by leveraging essential spectrum components within features in the frequency domain and capturing the channelwise correlation of bitemporal feature differences to learn change-aware representations. Third, we construct a unified change detection (UCD) framework that currently integrates 20 change detection methods. This framework is designed for streamlined and efficient processing, using the cloud platform hosted by HuggingFace. Extensive experiments have been conducted to demonstrate the superiority of the proposed baseline model compared with 19 state-of-the-art change detection approaches. Empirical studies on modularized backbones comprehensively confirm the efficacy of different representation learners on change detection. This benchmark represents significant advancements in the field of remote sensing and change detection, providing a robust resource for future research and applications in global mining monitoring. Dataset and Codes are available via the link.
Weikang Yu, Richard Gloaguen, Xiao Xiang Zhu 0001, Pedram Ghamisi
IEEE Trans. Geosci. Remote. Sens.1
2024 Prototypical Unknown-Aware Multiview Consistency Learning for Open-Set Cross-Domain Remote Sensing Image Classification
abstract
Developing a cross-domain classification model for remote sensing images has drawn significant attention in the literature. By leveraging the open-set unsupervised domain adaptation (UDA) technique, the generalization performance of deep learning models has been improved with the capability to recognize unknown categories. However, it remains challenging to explore distribution patterns in the target domain using uncertain category-wise supervision from unlabeled datasets while reducing negative transfer caused by unknown samples. To develop a robust open-set UDA framework, this article presents prototypical unknown-aware multiview consistency learning (PUMCL) designed for remote sensing scene classification across heterogeneous domains. Specifically, it employs a consistency learning scheme with multiview and multilevel perturbations to improve feature learning from unlabeled target samples. An entropy separation strategy is utilized to facilitate open-set detection and recognition during adaptation, enabling unknown-aware feature alignment. Furthermore, the introduction of prototypical constraints optimizes pseudo-label generation through online denoising and promotes a compact category-wise feature subspace for improved class separation across domains. Experiments conducted on six cross-domain scenarios using AID, NWPU, and UCMD datasets demonstrate the method’s superior performance compared to nine state-of-the-art approaches, achieving a gain of 4.5% to 21.2% in mIoU. More importantly, it shows promising class separability with clear boundaries between different classes and compact clustering of unknown samples in the feature space. The source code will be available athttps://github.com/zxk688.
Wanjing Wu, Mi Zhang 0004, Weikang Yu, Pedram Ghamisi
IEEE Trans. Geosci. Remote. Sens.4
2023 Weakly Supervised Local-Global Anchor Guidance Network for Landslide Extraction With Image-Level Annotations
abstract
Weakly supervised learning using image-level annotations has become a popular choice for reducing labeling efforts of remote sensing object extraction. Existing methods exploit inter-pixel relations within an individual image patch for object localizations. When facing large-scale remote sensing images, it is still challenging to obtain global semantic contexts across image patches for feature representation, resulting in inaccurate object localizations. To remedy these issues, we propose a local-global anchor guidance network (LGAGNet) for weakly supervised landslide extraction. Specifically, a structure-aware object locating (SOL) module is developed to capture the spatial structure of landslide objects and extract local category anchors containing informative feature embeddings. Furthermore, we leverage a global anchor aggregation (GAA) module to excavate semantic patterns across image patches based on a memory bank, which is then used as additional context cues to enhance the feature presentation through a cross-attention mechanism. Finally, a hybrid loss function is designed to guide the network training, considering category-aware semantic contrasts and local activation consistency. Experimental results on high-resolution aerial and satellite image datasets verify the effectiveness of the proposed approach on landslide extraction.
Weikang Yu, Xianping Ma, Xudong Kang
IEEE Geosci. Remote. Sens. Lett.2
2023 Federated Deep Learning With Prototype Matching for Object Extraction From Very-High-Resolution Remote Sensing Images
abstract
Deep convolutional neural networks (DCNNs) have become the leading tools for object extraction from very-high-resolution (VHR) remote sensing images. However, the label scarcity problem of local datasets hinders the prediction performances of DCNNs, and privacy concerns regarding remote sensing data often arise in the traditional deep learning schemes. To cope with these problems, we propose a novel federated learning scheme with prototype matching (FedPM) to collaboratively learn a richer DCNN model by leveraging remote sensing data distributed among multiple clients. This scheme conducts the federated optimization of DCNNs by aggregating clients’ knowledge in the gradient space without compromising data privacy. Specifically, the prototype matching method is developed to regularize the local training using prototypical representations while reducing the distribution divergence across heterogeneous image data. Furthermore, the derived deviations across local and global prototypes are applied to quantify the effects of local models on the decision boundary and optimize the global model updating via the attention-weighted aggregation scheme. Finally, the sparse ternary compression (STC) method is used to alleviate communication costs. Extensive experimental results derived from VHR aerial and satellite image datasets verify that the FedPM can dramatically improve the prediction performance of DCNNs on object extraction with lower communication costs. To the best of our knowledge, this is the first time that federated learning has been applied for remote sensing visual tasks.
Weikang Yu, Xudong Kang
IEEE Trans. Geosci. Remote. Sens.3
2023 Txt2Img-MHN: Remote Sensing Image Generation From Text Using Modern Hopfield Networks
abstract
The synthesis of high-resolution remote sensing images based on text descriptions has great potential in many practical application scenarios. Although deep neural networks have achieved great success in many important remote sensing tasks, generating realistic remote sensing images from text descriptions is still very difficult. To address this challenge, we propose a novel text-to-image modern Hopfield network (Txt2Img-MHN). The main idea of Txt2Img-MHN is to conduct hierarchical prototype learning on both text and image embeddings with modern Hopfield layers. Instead of directly learning concrete but highly diverse text-image joint feature representations for different semantics, Txt2Img-MHN aims to learn the most representative prototypes from text-image embeddings, achieving a coarse-to-fine learning strategy. These learned prototypes can then be utilized to represent more complex semantics in the text-to-image generation task. To better evaluate the realism and semantic consistency of the generated images, we further conduct zero-shot classification on real remote sensing data using the classification model trained on synthesized images. Despite its simplicity, we find that the overall accuracy in the zero-shot classification may serve as a good metric to evaluate the ability to generate an image from text. Extensive experiments on the benchmark remote sensing text-image dataset demonstrate that the proposed Txt2Img-MHN can generate more realistic remote sensing images than existing methods. Code and pre-trained models are available online (https://github.com/YonghaoXu/Txt2Img-MHN).
Yonghao Xu, Weikang Yu, Pedram Ghamisi, Michael Kopp 0001, Sepp Hochreiter
IEEE Trans. Image Process.2
2022 Cloud Removal in Optical Remote Sensing Imagery Using Multiscale Distortion-Aware Networks
abstract
Cloud layer contamination is a common problem in optical remote sensing (RS) images. Deep-learning-based cloud removal from RS imagery has attracted increasing attention in recent years. However, it remains challenging to exploit useful multiscale cloud-aware representations from cloud imagery due to the lack of effective modeling of cloud distortion effects and the weak feature representation capabilities of networks. To circumvent these challenges, we propose a multiscale distortion-aware cloud removal (MSDA-CR) network consisting of multiple cloud-distortion-aware representation learning (CDARL) modules combined in a multiscale grid architecture. Specifically, cloud distortion control functions (CDCFs) are defined and incorporated into the CDARL modules to adaptively model the distortion effects induced by cloud interference in the imaging process, with learnable parameters for the exploitation of distortion-restored representations. These representations are further distilled across different scales in the MSDA-CR network and integrated based on an attention mechanism to restore cloud-free images while retaining the spatial structures of ground objects. Extensive experiments on visible and multispectral RS datasets confirm the effectiveness of the proposed MSDA-CR network.
Weikang Yu, Man-On Pun
IEEE Geosci. Remote. Sens. Lett.1
2022 Multilevel Deformable Attention-Aggregated Networks for Change Detection in Bitemporal Remote Sensing Imagery
abstract
Deep learning (DL) approaches based on convolutional encoder–decoder networks have shown promising results in bitemporal change detection. However, their performance is limited by insufficient contextual information aggregation because they cannot fully capture the implicit contextual dependency relationships among feature maps at different levels. Moreover, harvesting long-range contextual information typically incurs high computational complexity. To circumvent these challenges, we propose multilevel deformable attention-aggregated networks (MLDANets) to effectively learn long-range dependencies across multiple levels of bitemporal convolutional features for multiscale context aggregation. Specifically, a multilevel change-aware deformable attention (MCDA) module consisting of linear projections with learnable parameters is built based on multihead self-attention (SA) with a deformable sampling strategy. It is applied in the skip connections of an encoder–decoder network taking a bitemporal deep feature hypersequence (BDFH) as input. MCDA can progressively address a set of informative sampling locations in multilevel feature maps for each query element in the BDFH. Simultaneously, MCDA learns to characterize beneficial information from different spatial and feature subspaces of BDFH using multiple attention heads for change perception. As a result, contextual dependencies across multiple levels of bitemporal feature maps can be adaptively aggregated via attention weights to generate multilevel discriminative change-aware representations. Experiments on very-high-resolution (VHR) datasets verify that MLDANets outperform state-of-the-art change detection approaches with dramatically faster training convergence and high computational efficiency.
Weikang Yu, Man-On Pun
IEEE Trans. Geosci. Remote. Sens.2
2021 A Hybrid Model-Based and Data-Driven Approach for Cloud Removal in Satellite Imagery Using Multi-Scale Distortion-Aware Networks
abstract
Cloud layer contamination is a common problem in optical remote sensing images. Cloud removal from remote sensing images has attracted increasing attention in recent years. To this end, we propose a multi-scale distortion-aware network for cloud removal from remote sensing images. A novel Cloud Aware and Feature Extraction (CAFE) module is developed by incorporating the physical model of cloud distortion considering atmosphere light, cloud reflectance light and cloud transmission. These distortion factors in CAFE module are encoded into trainabile parameters for feature extraction from contaminated images. The network is trained in an end-to-end manner with cloud contaminated images and ground truth data. Finally, experimentl results on the Remote sensing Image Cloud rEmoving (RICE) dataset demonstrate the effectiveness of the proposed approach.
Weikang Yu, Man-On Pun
IGARSS1
2021 Style Transformation-Based Change Detection Using Adversarial Learning with Object Boundary Constraints
abstract
Deep learning has shown promising results on change detection (CD) from bi-temporal remote sensing imagery in recent years. However, it still remains challenging to cope with the pseudo-changes caused by seasonal differences and style variations of bi-temporal images. In this paper, an object-level boundary-preserving generative adversarial network (BPGAN) is developed for style transformation-based CD of bi-temporal images. To achieve this purpose, image objects derived in the spectral domain are incorporated into the image translation to generate object-level target-style-like images. In particular, constraints on object boundary consistency and object homogeneity are established in the adversarial learning to maintain the style and content consistency while regularizing the network training. Furthermore, the Superpixel-Based Fast Fuzzy c-Means (SF-FCM) algorithm is utilized for efficient CD from the object-level style-transformed images. Extensive experiments on SPOT5 and GF1 data confirm the effectiveness of the proposed approach.
Weikang Yu, Man-On Pun
IGARSS2