Keyan Chen 0001

dblp:256/2434 · DBLP profile ↗
← Back
33ranked-venue papers
6as first author
33since 2021 · last 2025
0000-0003-0483-1306ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 28 · 5 first-author · 28 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2025 Open-CD: A Comprehensive Toolbox for Change Detection
abstract
We present Open-CD, a change detection toolbox that contains a rich set of change detection methods as well as related components and modules. The toolbox started from a series of open source general vision task tools, including OpenMMLab Toolkits, PyTorch Image Models (Timm), etc. It gradually evolves into a unified platform that covers many popular change detection methods and contemporary modules. It not only includes training and inference codes, but also provides some useful scripts for data analysis. We believe this toolbox is by far the most comprehensive change detection toolbox. In this report, we introduce the features, supported methods and applications of Open-CD. In addition, we also conduct a benchmarking study on different methods and components. We wish that the toolbox and benchmark could serve the growing research community by providing a flexible toolkit to re-implement existing methods and develop their own new change detectors. Code and models are available at https://github.com/likyoo/open-cd.
Kaiyu Li 0001, Chengxi Han, Yupeng Deng 0001, Keyan Chen 0001, Zhuo Zheng, Hao Chen 0045, Ziyuan Liu 0006, Yuantao Gu, Zhengxia Zou, Zhenwei Shi 0001, Sheng Fang 0001, Deyu Meng, Zhi Wang 0002, Xiangyong Cao
ACM Multimedia5
2025 Heterogeneous Mixture of Experts for Remote Sensing Image Super-Resolution
abstract
Remote sensing image super-resolution (SR) aims to reconstruct high-resolution remote sensing images from low-resolution inputs, thereby addressing limitations imposed by sensors and imaging conditions. However, the inherent characteristics of remote sensing images, including diverse ground object types and complex details, pose significant challenges to achieving high-quality reconstruction. Existing methods typically employ a uniform structure to process various types of ground objects without distinction, making it difficult to adapt to the complex characteristics of remote sensing images. To address this issue, we introduce a Mixture of Experts (MoE) model and design a set of heterogeneous experts. These experts are organized into multiple expert groups, where experts within each group are homogeneous while being heterogeneous across groups. This design ensures that specialized activation parameters can be employed to handle the diverse and intricate details of ground objects effectively. To better accommodate the heterogeneous experts, we propose a multi-level feature aggregation strategy to guide the routing process. Additionally, we develop a dual-routing mechanism to adaptively select the optimal expert for each pixel. Experiments conducted on the UCMerced and AID datasets demonstrate that our proposed method achieves superior SR reconstruction accuracy compared to state-of-the-art methods. The code will be available at https://github.com/Mr-Bamboo/MFG-HMoE.
Bowen Chen 0002, Keyan Chen 0001, Mohan Yang, Zhengxia Zou, Zhenwei Shi 0001
IEEE Geosci. Remote. Sens. Lett.2
2025 Structural Representation-Guided GAN for Remote Sensing Image Cloud Removal
abstract
Optical remote sensing imagery is often compromised by cloud cover, making effective cloud-removal techniques essential for enhancing the usability of such data. We designed a novel structural representation-guided generative adversarial network (GAN) framework for cloud removal, in which structure and gradient branches are integrated into the network, helping the model focus on the structural representations of ground objects during image reconstruction. Different from previous methods that concentrate on recovering pixel information, we emphasize learning the structural information of remote sensing images. We then utilize error feedback to fuse features from the structural auxiliary branch, guiding the image reconstruction process. During the training phase, synthetic cloud images are used to supervise the optimization of the cloud-removal network, while real cloud images are employed in an adversarial training manner for unsupervised learning to improve the generalization ability of the network. Additionally, multitemporal revisit images from remote sensing satellites are employed as auxiliary inputs, aiding the network to remove thick clouds reliably. We evaluated our framework on a dataset derived from SEN12MS-CR, and the proposed method outperformed classical cloud-removal methods in both objective performance and subjective visual quality. Furthermore, compared to other methods, our approach achieved superior cloud-removal results on real images.
Keyan Chen 0001, Liqin Liu, Zhengxia Zou, Zhenwei Shi 0001
IEEE Geosci. Remote. Sens. Lett.3
2025 Diffusion Models for Imperceptible and Transferable Adversarial Attack
abstract
Many existing adversarial attacks generate -norm perturbations on image RGB space. Despite some achievements in transferability and attack success rate, the crafted adversarial examples are easily perceived by human eyes. Towards visual imperceptibility, some recent works explore unrestricted attacks without -norm constraints, yet lacking transferability of attacking black-box models. In this work, we propose a novel imperceptible and transferable attack by leveraging both the generative and discriminative power of diffusion models. Specifically, instead of direct manipulation in pixel space, we craft perturbations in the latent space of diffusion models. Combined with well-designed content-preserving structures, we can generate human-insensitive perturbations embedded with semantic clues. For better transferability, we further "deceive" the diffusion model which can be viewed as an implicit recognition surrogate, by distracting its attention away from the target regions. To our knowledge, our proposed method, DiffAttack, is the first that introduces diffusion models into the adversarial attack field. Extensive experiments conducted across diverse model architectures (CNNs, Transformers, and MLPs), datasets (ImageNet, CUB-200, and Standford Cars), and defense mechanisms underscore the superiority of our attack over existing methods such as iterative attacks, GAN-based attacks, and ensemble attacks. Furthermore, we provide a comprehensive discussion on future research avenues in diffusion-based adversarial attacks, aiming to chart a course for this burgeoning field.
Jianqi Chen, Hao Chen 0045, Keyan Chen 0001, Yilan Zhang, Zhengxia Zou, Zhenwei Shi 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 SeG-SR: Integrating Semantic Knowledge Into Remote Sensing Image Super-Resolution via Vision-Language Model
abstract
High-resolution (HR) remote sensing imagery plays a vital role in a wide range of applications, including urban planning and environmental monitoring. However, due to limitations in sensors and data transmission links, the images acquired in practice often suffer from resolution degradation. Remote Sensing Image Super-Resolution (RSISR) aims to reconstruct HR images from low-resolution (LR) inputs, providing a cost-effective and efficient alternative to direct HR image acquisition. Existing RSISR methods primarily focus on low-level characteristics in pixel space, while neglecting the high-level understanding of remote sensing scenes. This may lead to semantically inconsistent artifacts in the reconstructed results. Motivated by this observation, our work aims to explore the role of high-level semantic knowledge in improving RSISR performance. We propose a Semantic-Guided Super-Resolution framework, SeG-SR, which leverages Vision-Language Models (VLMs) to extract semantic knowledge from input images and uses it to guide the super resolution (SR) process. Specifically, we first design a Semantic Feature Extraction Module (SFEM) that utilizes a pretrained VLM to extract semantic knowledge from remote sensing images. Next, we propose a Semantic Localization Module (SLM), which derives a series of semantic guidance from the extracted semantic knowledge. Finally, we develop a Learnable Modulation Module (LMM) that uses semantic guidance to modulate the features extracted by the SR network, effectively incorporating high-level scene understanding into the SR pipeline. We validate the effectiveness and generalizability of SeG-SR through extensive experiments: SeG-SR achieves state-of-the-art performance on three datasets, and consistently improves performance across various SR architectures. Notably, for the ×4 SR task on the UCMerced dataset, it attained a PSNR of 29.3042 dB and an SSIM of 0.7961. Codes can be found at https://github.com/Mr-Bamboo/SeG-SR.
Bowen Chen 0002, Keyan Chen 0001, Mohan Yang, Zhengxia Zou, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.2
2025 TriDF: Triplane-Accelerated Density Fields for Few-Shot Remote Sensing Novel View Synthesis
abstract
Remote sensing novel view synthesis (NVS) offers significant potential for 3D interpretation of remote sensing scenes, with important applications in urban planning and environmental monitoring. However, remote sensing scenes frequently lack sufficient multi-view images due to acquisition constraints. While existing NVS methods tend to overfit when processing limited input views, advanced few-shot NVS methods are computationally intensive and perform sub-optimally in remote sensing scenes. This paper presents TriDF, an efficient hybrid 3D representation for fast remote sensing NVS from as few as 3 input views. Our approach decouples color and volume density information, modeling them independently to reduce the computational burden on implicit radiance fields and accelerate reconstruction. We explore the potential of the triplane representation in few-shot NVS tasks by mapping high-frequency color information onto this compact structure, and the direct optimization of feature planes significantly speeds up convergence. Volume density is modeled as continuous density fields, incorporating reference features from neighboring views through image-based rendering to compensate for limited input data. Additionally, we introduce depth-guided optimization based on point clouds, which effectively mitigates the overfitting problem in few-shot NVS. Comprehensive experiments across multiple remote sensing scenes demonstrate that our hybrid representation achieves a 30× speed increase compared to NeRF-based methods, while simultaneously improving rendering quality metrics over advanced few-shot methods (7.4% increase in PSNR and 3.4% in SSIM). The code is publicly available at https://github.com/kanehub/TriDF.
Jiaming Kang, Keyan Chen 0001, Zhengxia Zou, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.2
2025 MarsSeg: Mars Surface Semantic Segmentation With Multilevel Extractor and Connector
abstract
The segmentation and interpretation of the Martian surface play a pivotal role in Mars exploration, providing essential data for the trajectory planning and obstacle avoidance of rovers. However, the complex topography, self-similar surface features, and the lack of extensive annotated data pose significant challenges to the high-precision semantic segmentation of the Martian surface. To address these challenges, we propose a novel encoder-decoder-based Mars segmentation network, termed MarsSeg. To facilitate a high-level semantic understanding across the multi-level feature maps, we introduce a feature enhancement module, which incorporates Multi-scale Feature Pyramid (MFP) and Strip Attention Pyramid Pooling Module (SAPPM). The MFP is specifically designed for shallow feature enhancement, thereby enabling the expression of local details and small objects. Conversely, the SAPPM is employed for deep feature enhancement, facilitating the extraction of high-level semantic category-related information. To effectively fuse features from different levels, we propose a feature fusion module, which contains Mars Polarized Self Attention (Mars-PSA) and Pixel Attention Head (PA-Head). Mars-PSA enables the fusion of multi-level information while directing the model’s attention to salient features. The PA-Head focuses on detailed information at the pixel level. Experimental results derived from the MarsSeg and AI4Mars datasets prove that the proposed MarsSeg outperforms other state-of-the-art methods in segmentation performance, validating the efficacy of each proposed component.
Keyan Chen 0001, Gengju Tian, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.2
2025 CDMamba: Incorporating Local Clues Into Mamba for Remote Sensing Image Binary Change Detection
abstract
Recently, the Mamba architecture based on state-space models has demonstrated remarkable performance in a series of natural language processing tasks and has been rapidly applied to remote sensing change detection (CD) tasks. However, most methods enhance the global receptive field by directly modifying the scanning mode of Mamba, neglecting the crucial role that local information plays in dense prediction tasks (e.g., binary CD). In this article, we propose a model called CDMamba, which effectively combines global and local features for handling binary CD tasks. Specifically, the scaled residual ConvMamba (SRCM) block is proposed to utilize the ability of Mamba to extract global features and convolution to enhance the local details, to alleviate the issue that current Mamba-based methods lack detailed clues and are difficult to achieve fine detection in dense prediction tasks. Furthermore, considering the characteristics of bi-temporal feature interaction required for CD, the adaptive global–local guided fusion (AGLGF) block is proposed to dynamically facilitate the bi-temporal interaction guided by other temporal global/local features. Our intuition is that more discriminative change features can be acquired with the guidance of other temporal features. Extensive experiments on five datasets demonstrate that our proposed CDMamba is comparable to the current methods (such as the F1/intersection over union (IoU) scores are improved by 2.10%/3.00%, 2.44%/2.91%, on LEVIR+CD and CLCD, respectively). Our code is open-sourced athttps://github.com/zmoka-zht/CDMamba.
Haotian Zhang 0010, Keyan Chen 0001, Hao Chen 0045, Zhengxia Zou, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.2
2025 FoBa: A Foreground-Background Co-Guided Method and New Benchmark for Remote Sensing Semantic Change Detection
abstract
Despite the remarkable progress achieved in remote sensing semantic change detection (SCD), two major challenges remain. At the data level, existing SCD datasets suffer from limited change categories, insufficient change types, and a lack of fine-grained class definitions, making them inadequate to fully support practical applications. At the methodological level, most current approaches underutilize change information, typically treating it as a post-processing step to enhance spatial consistency, which constrains further improvements in model performance. To address these issues, we construct a new benchmark for remote sensing SCD, LevirSCD. Focused on the Beijing area, the dataset covers 16 change categories and 210 specific change types, with more fine-grained class definitions (e.g., roads are divided into unpaved and paved roads). Furthermore, we propose a foreground-background co-guided SCD (FoBa) method, which leverages foregrounds that focus on regions of interest and backgrounds enriched with contextual information to guide the model collaboratively, thereby alleviating semantic ambiguity while enhancing its ability to detect subtle changes. Considering the requirements of bi-temporal interaction and spatial consistency in SCD, we introduce a gated interaction fusion (GIF) module along with a simple consistency loss to further enhance the model’s detection performance. Extensive experiments on three datasets (SECOND, JL1, and the proposed LevirSCD) demonstrate that FoBa achieves competitive results compared to current SOTA methods, with improvements of 1.48%, 3.61%, and 2.81% in the SeK metric, respectively. Our code and dataset are available at https://github.com/zmoka-zht/FoBa.
Haotian Zhang 0010, Keyan Chen 0001, Hao Chen 0045, Zhengxia Zou, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.3
2025 Zero-Shot Image Harmonization With Generative Model Prior
abstract
We propose a zero-shot approach to image harmonization, aiming to overcome the reliance on large amounts of synthetic composite images in existing methods. These methods, while showing promising results, involve significant training expenses and often struggle with generalization to unseen images. To this end, we introduce a fully modularized framework inspired by human behavior. Leveraging the reasoning capabilities of recent foundation models in language and vision, our approach comprises three main stages. Initially, we employ a pretrained vision-language model (VLM) to generate descriptions for the composite image. Subsequently, these descriptions guide the foreground harmonization direction of a text-to-image generative model (T2I). We refine text embeddings for enhanced representation of imaging conditions and employ self-attention and edge maps for structure preservation. Following each harmonization iteration, an evaluator determines whether to conclude or modify the harmonization direction. The resulting framework, mirroring human behavior, achieves harmonious results without the need for extensive training. We present compelling visual results across diverse scenes and objects, along with quantitative comparisons validating the effectiveness of our approach.
Jianqi Chen, Yilan Zhang, Zhengxia Zou, Keyan Chen 0001, Zhenwei Shi 0001
IEEE Trans. Multim.4
2024 Time Travelling Pixels: Bitemporal Features Integration with Foundation Model for Remote Sensing Image Change Detection
abstract
Change detection, a prominent research area in remote sensing, is pivotal in observing and analyzing surface transformations. Despite significant advancements achieved through deep learning-based methods, executing high-precision change detection in spatiotemporally complex remote sensing scenarios still presents a substantial challenge. The recent emergence of foundation models, with their powerful universality and generalization capabilities, offers potential solutions. However, bridging the gap of data and tasks remains a significant obstacle. In this paper, we introduce Time Travelling Pixels (TTP), a novel approach that integrates the latent knowledge of the SAM foundation model into change detection. TTP can effectively address the domain shift in general knowledge transfer and the challenge of expressing homogeneous and heterogeneous characteristics of multi-temporal images. The state-of-the-art results obtained on the LEVIR-CD underscore the efficacy of the TTP. The code has been made publicly available at https://github.com/KyanChen/TTP.
Keyan Chen 0001, Chengyang Liu, Wenyuan Li 0002, Hao Chen 0045, Haotian Zhang 0010, Zhengxia Zou, Zhenwei Shi 0001
IGARSS1
2024 Learning to Detect Cloud and Snow in Remote Sensing Images from Noisy Labels
abstract
Detecting clouds and snow in remote sensing images is an essential preprocessing task for remote sensing imagery. Previous works draw inspiration from semantic segmentation models in computer vision, with most research focusing on improving model architectures to enhance detection performance. However, unlike natural images, the complexity of scenes and the diversity of cloud types in remote sensing images result in many inaccurate labels in cloud and snow detection datasets, introducing unnecessary noises into the training and testing processes. By constructing a new dataset and proposing a novel training strategy with the curriculum learning paradigm, we guide the model in reducing overfitting to noisy labels. Additionally, we design a more appropriate model performance evaluation method, that alleviates the performance assessment bias caused by noisy labels. By conducting experiments on models with UNet and Segformer, we have validated the effectiveness of our proposed method. This paper is the first to consider the impact of label noise on the detection of clouds and snow in remote sensing images.
Hao Chen 0045, Wenyuan Li 0002, Keyan Chen 0001, Zipeng Qi, Zhengxia Zou, Zhenwei Shi 0001
IGARSS4
2024 Pixel-Level Change Detection Pseudo-Label Learning For Remote Sensing Change Captioning
abstract
The existing Remote Sensing Image Change Captioning (RSICC) methods perform well in simple scenes but exhibit poorer performance in complex scenes. This limitation is primarily attributed to the model’s constrained visual ability to distinguish and locate changes. Acknowledging the inherent correlation between change detection (CD) and RSICC tasks, we believe pixel-level CD is significant for describing the differences between images through language. Regrettably, the current RSICC dataset lacks readily available pixel-level CD labels. To address this deficiency, we leverage a model trained on existing CD datasets to derive CD pseudo-labels. We propose an innovative network with an auxiliary CD branch, supervised by pseudo-labels. Furthermore, a semantic fusion augment (SFA) module is proposed to fuse the feature information extracted by the CD branch, thereby facilitating the nuanced description of changes. Experiments demonstrate that our method achieves state-of-the-art performance and validate that learning pixel-level CD pseudo-labels significantly contributes to change captioning.
Keyan Chen 0001, Zipeng Qi, Haotian Zhang 0010, Zhengxia Zou, Zhenwei Shi 0001
IGARSS2
2024 Residual Group Enhanced GAN for Remote Sensing Image Cloud Shadow Removal
abstract
Cloud shadow is an unignorable factor affecting the quality of remote sensing images, but there are few researches specifically focusing on cloud shadow removal and its impact on downstream remote sensing tasks. In this paper, we propose a residual group enhanced generative adversarial network (RGE-GAN) for cloud shadow removal. We design an encoder-decoder with residual group enhancement (RGE) module to remove cloud shadows from remote sensing images. RGE module can effectively enhance the deep features extracted by encoder. We further introduce a discriminator network and employ adversarial training strategy to constrain the generator to reconstruct high-quality cloud shadow removed images conforming to the distribution of remote sensing images. The joint experiments of cloud shadow removal and building extraction on real remote sensing dataset show that our cloud shadow removal method can effectively enhance the quality of remote sensing images and improve the performance of downstream remote sensing processing tasks.
Keyan Chen 0001, Zhengxia Zou, Zhenwei Shi 0001
IGARSS3
2024 RSMamba: Remote Sensing Image Classification With State Space Model
abstract
Remote sensing image classification forms the foundation of various understanding tasks, serving a crucial function in remote sensing image interpretation. The recent advancements of Convolutional Neural Networks (CNNs) and Transformers have markedly enhanced classification accuracy. Nonetheless, remote sensing scene classification remains a significant challenge, especially given the complexity and diversity of remote sensing scenarios and the variability of spatiotemporal resolutions. The capacity for whole-image understanding can provide more precise semantic cues for scene discrimination. In this paper, we introduce RSMamba, a novel architecture for remote sensing image classification. RSMamba is based on the State Space Model (SSM) and incorporates an efficient, hardware-aware design known as the Mamba. It integrates the advantages of both a global receptive field and linear modeling complexity. To overcome the limitation of the vanilla Mamba, which can only model causal sequences and is not adaptable to two-dimensional image data, we propose a dynamic multi-path activation mechanism to augment Mamba’s capacity to model non-causal data. Notably, RSMamba maintains the inherent modeling mechanism of the vanilla Mamba, yet exhibits superior performance across multiple remote sensing image classification datasets,e.g., F1 scores of 95.25, 92.63, and 95.18 on the UC Merced, AID, and RESISC45 classification datasets respectively, exceeding those of concurrent Vim and VMamba. This indicates that RSMamba holds significant potential to function as the backbone of future visual foundation models. The code is available at https://github.com/KyanChen/RSMamba.
Keyan Chen 0001, Bowen Chen 0002, Wenyuan Li 0002, Zhengxia Zou, Zhenwei Shi 0001
IEEE Geosci. Remote. Sens. Lett.1
2024 RSCaMa: Remote Sensing Image Change Captioning With State Space Model
abstract
Remote Sensing Image Change Captioning (RSICC) aims to describe surface changes between multi-temporal remote sensing images in language, including the changed object categories, locations, and dynamics of changing objects (e.g., added or disappeared). This poses challenges to spatial and temporal modeling of bi-temporal features. Despite previous methods progressing in the spatial change perception, there are still weaknesses in joint spatial-temporal modeling. To address this, in this paper, we propose a novel RSCaMa model, which achieves efficient joint spatial-temporal modeling through multiple CaMa layers, enabling iterative refinement of bi-temporal features. To achieve efficient spatial modeling, we introduce the recently popular Mamba (a state space model) with a global receptive field and linear complexity into the RSICC task and propose the Spatial Difference-aware SSM (SD-SSM), overcoming limitations of previous CNN- and Transformer-based methods in the receptive field and computational complexity. SD-SSM enhances the model’s ability to capture spatial changes sharply. In terms of efficient temporal modeling, considering the potential correlation between the temporal scanning characteristics of Mamba and the temporality of the RSICC, we propose the Temporal-Traversing SSM (TT-SSM), which scans bi-temporal features in a temporal cross-wise manner, enhancing the model’s temporal understanding and information interaction. Experiments validate the effectiveness of the efficient joint spatial-temporal modeling and demonstrate the outstanding performance of RSCaMa and the potential of the Mamba in the RSICC task. Additionally, we systematically compare three different language decoders, including Mamba, GPT-style decoder, and Transformer decoder, providing valuable insights for future RSICC research. The code will be available at https://github.com/Chen-Yang-Liu/RSCaMa.
Keyan Chen 0001, Bowen Chen 0002, Haotian Zhang 0010, Zhengxia Zou, Zhenwei Shi 0001
IEEE Geosci. Remote. Sens. Lett.2
2024 Dense Pixel-to-Pixel Harmonization via Continuous Image Representation
abstract
High-resolution (HR) image harmonization is of great significance in real-world applications such as image synthesis and image editing. However, due to the high memory costs, existing dense pixel-to-pixel harmonization methods are mainly focusing on processing low-resolution (LR) images. Some recent works resort to combining with color-to-color transformations but are either limited to certain resolutions or heavily depend on hand-crafted image filters. In this work, we explore leveraging the implicit neural representation (INR) and propose a novel image Harmonization method based on Implicit neural Networks (HINet), which to the best of our knowledge, is the first dense pixel-to-pixel method applicable to HR images without any hand-crafted filter design. Inspired by the Retinex theory, we decouple the MLPs into two parts to respectively capture the content and environment of composite images. A Low-Resolution Image Prior (LRIP) network is designed to alleviate the Boundary Inconsistency problem, and we also propose new designs for the training and inference process. Extensive experiments have demonstrated the effectiveness of our method compared with state-of-the-art methods. Furthermore, some interesting and practical applications of the proposed method are explored. Our code is available at https://github.com/WindVChen/INR-Harmonization.
Jianqi Chen, Yilan Zhang, Zhengxia Zou, Keyan Chen 0001, Zhenwei Shi 0001
IEEE Trans. Circuits Syst. Video Technol.4
2024 RSPrompter: Learning to Prompt for Remote Sensing Instance Segmentation Based on Visual Foundation Model
abstract
Leveraging the extensive training data from SA-1B, the Segment Anything Model (SAM) demonstrates remarkable generalization and zero-shot capabilities. However, as a category-agnostic instance segmentation method, SAM heavily relies on prior manual guidance, including points, boxes, and coarse-grained masks. Furthermore, its performance in remote sensing image segmentation tasks remains largely unexplored and unproven. In this paper, we aim to develop an automated instance segmentation approach for remote sensing images, based on the foundational SAM model and incorporating semantic category information. Drawing inspiration from prompt learning, we propose a method to learn the generation of appropriate prompts for SAM. This enables SAM to produce semantically discernible segmentation results for remote sensing images, a concept we have termed RSPrompter. We also propose several ongoing derivatives for instance segmentation tasks, drawing on recent advancements within the SAM community, and compare their performance with RSPrompter. Extensive experimental results, derived from the WHU building, NWPU VHR-10, and SSDD datasets, validate the effectiveness of our proposed method. The code for our method is publicly available at https://kychen.me/RSPrompter.
Keyan Chen 0001, Hao Chen 0045, Haotian Zhang 0010, Wenyuan Li 0002, Zhengxia Zou, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.1
2024 Digital-to-Physical Visual Consistency Optimization for Adversarial Patch Generation in Remote Sensing Scenes
abstract
In contrast to digital image adversarial attacks, adversarial patch attacks involve physical operations that project crafted perturbations into real-world scenarios. During the digital-to-physical transition, adversarial patches inevitably undergo information distortion. Existing approaches focus on data augmentation and printer color gamut regularization to improve the generalization of adversarial patches to the physical world. However, these efforts overlook a critical issue within the adversarial patch crafting pipeline—namely, the significant disparity between the appearance of adversarial patches during the digital optimization phase and their manifestation in the physical world. This unexplored concern, termed “Digital-to-Physical Visual Inconsistency", introduces inconsistent objectives between the digital and physical realms, potentially skewing optimization directions for adversarial patches. To tackle this challenge, we propose a novel harmonization-based adversarial patch attack. Our approach involves the design of a self-supervised harmonization method, seamlessly integrated into the adversarial patch generation pipeline. This integration aligns the appearance of adversarial patches overlaid on digital images with the imaging environment of the background, ensuring a consistent optimization direction with the primary physical attack goal. We validate our method through extensive testing on the aerial object detection task. To enhance the controllability of environmental factors for method evaluation, we construct a dataset of 3D simulated scenarios using a graphics rendering engine. Extensive experiments on these scenarios demonstrate the efficacy of our approach. Our code and dataset are publicly accessible at https://github.com/WindVChen/VCO-AP.
Jianqi Chen, Yilan Zhang, Keyan Chen 0001, Zhengxia Zou, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.4
2024 Deriving Accurate Surface Meteorological States at Arbitrary Locations via Observation-Guided Continuous Neural Field Modeling
abstract
Accurately retrieving surface meteorological states at arbitrary locations is of great application significance in weather forecasting and climate modeling. Since meteorological variables are typically provided as coarse-resolution gridded fields, common methods that obtain the states at a specific location directly through spatial interpolation can lead to significant accuracy deviations compared to actual observations. Traditional downscaling, the process of obtaining fixed-scale high-resolution meteorological fields from low-resolution inputs, has been proposed as a way to indirectly improve the accuracy of retrieving states at arbitrary locations by providing more detailed subgrid-scale information. However, for arbitrary locations at the station scale, their states are influenced by subgrid information, resulting in systematic biases between the downscaled results after interpolation and the actual observations at specific station locations. To address this issue, in this article, we propose a new task called station-scale downscaling, which aims to directly derive accurate meteorological states at any given station location from a coarse-resolution meteorological field. To achieve this, we propose a new downscaling model based on hypernetwork architecture, namely, HyperDS, which efficiently integrates the multiscale observational information to guide the continuous neural field modeling of the meteorological variables, enabling accurate sampling of the states at any target location. Through extensive experiments, our proposed method outperforms other specially designed baseline models on multiple surface variables. Notably, the mean squared error (mse) for wind speed and surface pressure improved by 67% and 19.5% compared with other methods, respectively.
Hao Chen 0045, Lei Bai 0001, Wenyuan Li 0002, Keyan Chen 0001, Wanli Ouyang, Zhengxia Zou, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.5
2024 Change-Agent: Toward Interactive Comprehensive Remote Sensing Change Interpretation and Analysis
abstract
Monitoring changes in the Earth’s surface is crucial for understanding natural processes and human impacts, necessitating precise and comprehensive interpretation methodologies. Remote sensing (RS) satellite imagery offers a unique perspective for monitoring these changes, leading to the emergence of RS image change interpretation (RSICI) as a significant research focus. Current RSICI technology encompasses change detection and change captioning, each with its limitations in providing comprehensive interpretation. To address this, we propose an interactive Change-Agent, which can follow user instructions to achieve comprehensive change interpretation and insightful analysis, such as change detection and change captioning, change object counting, and change cause analysis. The Change-Agent integrates a multilevel change interpretation (MCI) model as the eyes and a large language model (LLM) as the brain. The MCI model contains two branches of pixel-level change detection and semantic-level change captioning, in which the BI-temporal iterative interaction (BI3) layer is proposed to enhance the model’s discriminative feature representation capabilities. To support the training of the MCI model, we build the LEVIR-MCI dataset with a large number of change masks and captions of changes. Experiments demonstrate the state-of-the-art (SOTA) performance of the MCI model in achieving both change detection and change description simultaneously and highlight the promising application value of our Change-Agent in facilitating comprehensive interpretation of surface changes, which opens up a new avenue for intelligent RS applications. To facilitate future research, we will make our dataset and codebase publicly available athttps://github.com/Chen-Yang-Liu/Change-Agent.
Keyan Chen 0001, Haotian Zhang 0010, Zipeng Qi, Zhengxia Zou, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.2
2024 Generating Imperceptible and Cross-Resolution Remote Sensing Adversarial Examples Based on Implicit Neural Representations
abstract
Deep neural networks (DNNs) have been widely applied in remote sensing, and the research on its adversarial attack algorithm is the key to evaluating its robustness. Current adversarial attack methods primarily prioritize maximizing the attack success rate, disregarding the imperceptibility of the generated adversarial noise to human visual perception. Moreover, research on adversarial sample transferability has mostly focused on cross-model and cross-dataset scenarios, overlooking the investigation of adversarial attacks across different resolutions, while the rarely studied cross-resolution adversarial attacks are critical for remote sensing with different resolutions. In this article, we propose a novel method for generating imperceptible adversarial samples for cross-resolution remote sensing images based on implicit neural representations (INRs). By mapping the discrete images to a continuous neural functional space, we explicitly guarantee the visual quality of adversarial samples and decouple the model input from the image resolution. To enhance the visual fidelity of the generated adversarial samples, a multiscale discriminative learning scheme is proposed for the optimization process. For cross-resolution adversarial attacks, we align with images of different resolutions and generate cross-resolution adversarial perturbation by benefiting from the natural properties of the continuous resolution of INRs. To validate the effectiveness of our method, we compare it with the existing adversarial attacking methods using four evaluation metrics. Experiments show that our method achieves the best results in terms of attack success rate, imperceptibility, and cross-resolution attack transferability. Our code will be made publicly available.
Jianqi Chen, Liqin Liu, Keyan Chen 0001, Zhenwei Shi 0001, Zhengxia Zou
IEEE Trans. Geosci. Remote. Sens.4
2024 BiFA: Remote Sensing Image Change Detection With Bitemporal Feature Alignment
abstract
Despite the success of deep learning-based change detection methods, their existing insufficiency in temporal (channel, spatial) and multi-scale alignment have rendered them insufficient capability in mitigating external factors (illumination changes and perspective differences, etc.) arising from different imaging conditions during change detection. In this paper, a Bi-temporal Feature Alignment (BiFA) model is proposed to produce a precise change detection map in a lightweight manner by reducing the impact of irrelevant factors. Specifically, for the temporal alignment, the Bi-temporal Interaction (BI) module is proposed to realize the alignment of the bi-temporal image channel level. Our intuition is introducing the bi-temporal interaction in the feature extraction stage may benefit suppressing the interference, such as illumination changes. Simultaneously, the Alignment module based on Differential Flow Field (ADFF) is proposed to explicitly estimate the offset of the bi-temporal image and realize their spatial level alignment to mitigate the inadequate registration resulting from different perspectives. Furthermore, for the multi-scale alignment, we introduce the Implicit Neural alignment Decoder (IND) to produce more refined prediction maps achieving precise alignment of multi-scale features by learning continuous image representations in coordinate space. Our BiFA outperforms other state-of-the-art methods on six datasets (such as the F1/IoU scores are improved by 2.70%/3.91%, 2.01%/2.94% on LEVIR+-CD and SYSU-CD, respectively) and displays greater robustness in cross-resolutions change detection. Our code is available at https://github.com/zmoka-zht/BiFA.
Haotian Zhang 0010, Hao Chen 0045, Chenyao Zhou, Keyan Chen 0001, Zhengxia Zou, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.4
2024 Semantic-CC: Boosting Remote Sensing Image Change Captioning via Foundational Knowledge and Semantic Guidance
abstract
Remote sensing image change captioning (RSICC) aims to articulate the changes in objects of interest within bitemporal remote sensing images using natural language. Given the limitations of current RSICC methods in expressing general features across multitemporal and spatial scenarios, and their deficiency in providing granular, robust, and precise change descriptions, we introduce a novel change captioning (CC) method based on the foundational knowledge and semantic guidance, which we term Semantic-CC. Semantic-CC alleviates the dependency of high-generalization algorithms on extensive annotations by harnessing the latent knowledge of foundation models, and it generates more comprehensive and accurate change descriptions guided by pixel-level semantics from change detection (CD). Specifically, we propose a bitemporal SAM-based encoder for dual-image feature extraction; a multitask semantic aggregation neck for facilitating information interaction between heterogeneous tasks; a straightforward multiscale CD decoder to provide pixel-level semantic guidance; and a change caption decoder based on the large language model (LLM) to generate change description sentences. Moreover, to ensure the stability of the joint training of CD and CC, we propose a three-stage training strategy that supervises different tasks at various stages. We validate the proposed method on the LEVIR-CC and LEVIR-CD datasets. The experimental results corroborate the complementarity of CD and CC, demonstrating that Semantic-CC can generate more accurate change descriptions and achieve optimal performance across both tasks.
Yongshuo Zhu, Keyan Chen 0001, Fugen Zhou, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.3
2023 OvarNet: Towards Open-Vocabulary Object Attribute Recognition
abstract
In this paper, we consider the problem of simultaneously detecting objects and inferring their visual attributes in an image, even for those with no manual annotations provided at the training stage, resembling an open-vocabulary scenario. To achieve this goal, we make the following contributions: (i) we start with a naive two-stage approach for open-vocabulary object detection and attribute classification, termed CLIP-Attr. The candidate objects are first proposed with an offline RPN and later classified for semantic category and attributes; (ii) we combine all available datasets and train with a federated strategy to finetune the CLIP model, aligning the visual representation with attributes, additionally, we investigate the efficacy of leveraging freely available online image-caption pairs under weakly supervised learning; (iii) in pursuit of efficiency, we train a Faster-RCNN type model end-to-end with knowledge distillation, that performs class-agnostic object proposals and classification on semantic categories and attributes with classifiers generated from a text encoder; Finally, (iv) we conduct extensive experiments on VAW, MS-COCO, LSA, and OVAD datasets, and show that recognition of semantic category and attributes is complementary for visual scene understanding, i.e., jointly training object detection and attributes prediction largely outperform existing approaches that treat the two tasks independently, demonstrating strong generalization ability to novel attributes and categories.
Keyan Chen 0001, Yao Hu 0002, Xu Tang 0007, Yan Gao 0017, Jianqi Chen, Weidi Xie
CVPR1
2023 Resolution-Agnostic Remote Sensing Scene Classification With Implicit Neural Representations
abstract
Remote sensing scene classification is an important yet challenging task. In recent years, the excellent feature representation ability of convolutional neural networks (CNNs) has led to substantial improvements in scene classification accuracy. However, handling resolution variations of remote sensing images is still challenging because CNNs are not inherently capable of modeling multiresolution input images. In this letter, we propose a novel scene classification method with scale and resolution adaptation ability by leveraging the recent advances in implicit neural representations (INRs). Unlike previous CNN-based methods that make predictions based on rasterized image inputs, the proposed method converts the images as continuous functions with INRs optimization and then performs classification within the function space. When the image is represented as a function, the image resolution can be decoupled from the pixel values so that the resolution does not have much impact on the classification performance. Our method also shows great potential for multiresolution remote sensing scene classification. Using only a simple multilayer perceptron (MLP) classifier in the proposed function space, our method achieves classification accuracy comparable to deep CNNs but exhibits better adaptability to image scale and resolution changes.
Keyan Chen 0001, Wenyuan Li 0002, Jianqi Chen, Zhengxia Zou, Zhenwei Shi 0001
IEEE Geosci. Remote. Sens. Lett.1
2023 Object Detection in 20 Years: A Survey
abstract
Object detection, as of one the most fundamental and challenging problems in computer vision, has received great attention in recent years. Over the past two decades, we have seen a rapid technological evolution of object detection and its profound impact on the entire computer vision field. If we consider today’s object detection technique as a revolution driven by deep learning, then, back in the 1990s, we would see the ingenious thinking and long-term perspective design of early computer vision. This article extensively reviews this fast-moving research field in the light of technical evolution, spanning over a quarter-century’s time (from the 1990s to 2022). A number of topics have been covered in this article, including the milestone detectors in history, detection datasets, metrics, fundamental building blocks of the detection system, speedup techniques, and recent state-of-the-art detection methods.
Zhengxia Zou, Keyan Chen 0001, Zhenwei Shi 0001, Yuhong Guo, Jieping Ye
Proc. IEEE2
2023 Continuous Remote Sensing Image Super-Resolution Based on Context Interaction in Implicit Function Space
abstract
Despite its fruitful applications in remote sensing, image super-resolution is troublesome to train and deploy as it handles different resolution magnifications with separate models. Accordingly, we propose a highly-applicable super-resolution framework called FunSR, which settles different magnifications with a unified model by exploiting context interaction within implicit function space. FunSR composes a functional representor, a functional interactor, and a functional parser. Specifically, the representor transforms the low-resolution image from Euclidean space to multi-scale pixel-wise function maps; the interactor enables pixel-wise function expression with global dependencies; and the parser, which is parameterized by the interactor’s output, converts the discrete coordinates with additional attributes to RGB values. Extensive experimental results demonstrate that FunSR reports state-of-the-art performance on both fixed-magnification and continuous-magnification settings, meanwhile, it provides many friendly applications thanks to its unified nature. Our code is available at https://github.com/KyanChen/FunSR.
Keyan Chen 0001, Wenyuan Li 0002, Sen Lei, Jianqi Chen, Zhengxia Zou, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.1
2023 Continuous Cross-Resolution Remote Sensing Image Change Detection
abstract
Most contemporary supervised Remote Sensing (RS) image Change Detection (CD) approaches are customized for equal-resolution bitemporal images. Real-world applications raise the need for cross-resolution change detection, aka, CD based on bitemporal images with different spatial resolutions. Given training samples of a fixed bitemporal resolution difference (ratio) between the high-resolution (HR) image and the low-resolution (LR) one, current cross-resolution methods may fit a certain ratio but lack adaptation to other resolution differences. Toward continuous cross-resolution CD, we propose scale-invariant learning to enforce the model consistently predicting HR results given synthesized samples of varying resolution differences. Concretely, we synthesize blurred versions of the HR image by random downsampled reconstructions to reduce the gap between HR and LR images. We introduce coordinate-based representations to decode per-pixel predictions by feeding the coordinate query and corresponding multi-level embedding features into an MLP that implicitly learns the shape of land cover changes, therefore benefiting recognizing blurred objects in the LR image. Moreover, considering that spatial resolution mainly affects the local textures, we apply local-window self-attention to align bitemporal features during the early stages of the encoder. Extensive experiments on two synthesized and one real-world different-resolution CD datasets verify the effectiveness of the proposed method. Our method significantly outperforms several vanilla CD methods and two cross-resolution CD methods on the three datasets both in in-distribution and out-of-distribution settings. The empirical results suggest that our method could yield relatively consistent HR change predictions regardless of varying bitemporal resolution ratios. Our code will be public.
Hao Chen 0045, Haotian Zhang 0010, Keyan Chen 0001, Chenyao Zhou, Zhengxia Zou, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.3
2022 Contrastive Learning for Fine-Grained Ship Classification in Remote Sensing Images
abstract
Fine-grained image classification can be considered as a discriminative learning process where images of different subclasses are separated from each other while the same subclass images are clustered. Most existing methods perform synchronous discriminative learning in their approaches. Although achieving promising results in fine-grained visual classification (FGVC) in natural images, these methods may fail in fine-grained ship classification (FGSC) problem in remote sensing (RS) images due to the highly “imbalanced fineness" and “imbalanced appearances" of ships among subclasses. To tackle the issue, we propose an asynchronous contrastive learning-based method for effective FGSC. The proposed method, which we refer to as “Push-and-Pull Network (P2Net)", includes a “push-out stage” and a “pull-in stage”, where the first stage forces all the instances to be de-correlated and then the second one groups them into each subclass. A dual-branch network is designed to separate/de-correlate the images with each other, while an Integration Module is designed to aggregate the de-correlated images into their corresponding subclass together with a Proxy-based Module designed for acceleration. In this way, the correlation between subclasses can be decoupled, which in turn makes the final classification much easier. Our method can be trained end-to-end and requires no additional annotations other than category information. Extensive experiments are conducted on two large-scale FGSC datasets (FGSC-23 and FGSCR-42). Our method outperforms other state-of-the-art approaches. Ablation experiments also suggest the effectiveness of our design. Our code is available at https://github.com/WindVChen/Push-and-Pull-Network.
Jianqi Chen, Keyan Chen 0001, Hao Chen 0045, Wenyuan Li 0002, Zhengxia Zou, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.2
2022 A Degraded Reconstruction Enhancement-Based Method for Tiny Ship Detection in Remote Sensing Images With a New Large-Scale Dataset
abstract
The rapid detection of ships within the wide sea area is essential for intelligence acquisition. Most modern deep learning-based ship detection methods focus on locating ships in high-resolution (HR) remote sensing (RS) images. Seldom efforts have been made on ship detection in medium-resolution (MR) RS images. An MR image covers a much wider area than an HR one of the same size, thus facilitating quick ship detection. To this end, we propose a tiny ship detection method namely, Degraded Reconstruction Enhancement Network (DRENet), for MR RS images. Different from previous methods that mainly focus on feature fusion strategies to improve the expression ability of the detector, we design an additional network branch, i.e., degraded reconstruction enhancer, to learn to regress an object-aware blurred version of the input image in the training phase. Our intuition is that the proposed reconstruction branch may guide the backbone to focus more on tiny ship targets instead of the vast background. Moreover, we incorporate a CRoss-stage Multi-head Attention module in the detector to further improve the feature discrimination by leveraging the self-attention mechanism. To fill the gap of lacking a large-scale MR ship detection dataset, we introduce Levir-Ship, which contains 3876 GF-1/GF-6 multi-spectral images and over 3K tiny ship instances. Experiments on Levir-Ship validate the effectiveness and efficiency of the proposed method. Our method achieves 82.4 AP with 85 FPS, which outperforms many state-of-the-art ship detection methods. Our code and dataset will be made public.
Jianqi Chen, Keyan Chen 0001, Hao Chen 0045, Zhengxia Zou, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.2
2022 Geographical Knowledge-Driven Representation Learning for Remote Sensing Images
abstract
The proliferation of remote sensing satellites has resulted in a massive amount of remote sensing images. However, due to human and material resource constraints, the vast majority of remote sensing images remain unlabeled. As a result, it cannot be applied to currently available deep learning methods. To fully utilize the remaining unlabeled images, we propose a Geographical Knowledge-driven Representation learning method for remote sensing images (GeoKR), improving network performance and reduce the demand for annotated data. The global land cover products and geographical location associated with each remote sensing image are regarded as geographical knowledge to provide supervision for representation learning and network pre-training. An efficient pre-training framework is proposed to eliminate the supervision noises caused by imaging times and resolutions difference between remote sensing images and geographical knowledge. A large scale pre-training dataset Levir-KR is proposed to support network pre-training. It contains 1,431,950 remote sensing images from Gaofen series satellites with various resolutions. Experimental results demonstrate that our proposed method outperforms ImageNet pre-training and self-supervised representation learning methods and significantly reduces the burden of data annotation on downstream tasks such as scene classification, semantic segmentation, object detection, and cloud / snow detection. It demonstrates that our proposed method can be used as a novel paradigm for pre-training neural networks. Codes will be available on https://github.com/flyakon/Geographical-Knowledge-driven-Representaion-Learning.
Wenyuan Li 0002, Keyan Chen 0001, Hao Chen 0045, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.2
2022 Geographical Supervision Correction for Remote Sensing Representation Learning
abstract
Global land cover (GLC) products can be utilized to provide geographical supervision for remote sensing representation learning, which has significantly improved downstream tasks’ performance and decreased the demand of manual annotations. However, the time differences between remote sensing images and GLC products may introduce deviations in geographical supervision. In this paper, we propose a Geographical supervision Correction method (GeCo) for remote sensing representation learning. Deviated geographical supervision generated by GLC products can be corrected adaptively using the correction matrix during network pre-training and joint optimization process is designed to simultaneously update the correction matrix and network parameters. Additionally, we identify prior knowledge on geographical supervision to guide representation learning and restrict the correction process. The prior knowledge named “minor changes” implies that the geographical supervision may not change significantly, whereas the prior knowledge named “spatial aggregation” implies that land covers are aggregated in their spatial distribution. According to the prior knowledge, corresponding regularization terms are proposed to prevent abrupt changes in geographical supervision correction process and excessive smoothing of network outputs, thereby ensuring the adaptive correction process’s correctness. Experimental results demonstrate that our proposed method outperforms random initialization, ImageNet pre-training, and other representation learning methods on a variety of downstream tasks. In particular, when compared to the method that learns representations directly from deviated geographical supervision, it is proved that our method can eliminate the influence of deviations and further improve the effect of representation learning.
Wenyuan Li 0002, Keyan Chen 0001, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.2