Jie Chen 0048

dblp:92/6289-48 · DBLP profile ↗
← Back
14ranked-venue papers
10as first author
13since 2021 · last 2025
0000-0002-3864-9265ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 14 · 10 first-author · 13 since 2021
YearPublicationVenuePosition
2025 Multiscale Feature Knowledge Distillation and Implicit Object Discovery for Few-Shot Object Detection in Remote Sensing Images
abstract
Dynamic or sudden changes in various scenes may give rise to new objects. These new objects with limited annotated samples are susceptible to overfitting in deep learning. While few-shot object detection (FSOD) is effective with limited samples, current FSOD methods for remote sensing images still face specific challenges. The “pretraining-transfer” paradigm tends to forget the feature representations of base classes, impacting the learning process for novel classes during few-shot training. Furthermore, the presence of implicit objects in sparsely labeled instances of remote sensing images introduces erroneous supervisory information. To address these challenges, we propose an FSOD method that incorporates multiscale feature knowledge distillation and implicit object discovery, named MFKDIOD, which preserves the performance of base classes and mitigates the impact of implicit objects. Specifically, we first design a multiscale feature knowledge distillation (MFKD) module, which transfers the knowledge of base classes from a teacher network to a student network, enabling the student network to better retain the base class feature representations. Second, we design an implicit object discovery (IOD) module that utilizes both the teacher and student networks to discover implicit objects within the few-shot training data and generate pseudolabels. The code will be available athttps://github.com/RS-CSU/MFKDIOD.
Jie Chen 0048, Ya Guo 0002, Dengda Qin, Jingru Zhu, Zhenbo Gou, Geng Sun 0005
IEEE Trans. Geosci. Remote. Sens.1
2025 A Dual-Contrast Adaptation Network Coupling Global Context and Geometry Information for Cross-Domain Building Extraction
abstract
Semantic segmentation of high-resolution remote sensing imagery (HRSI) suffers from the domain shift, resulting in poor performance of the model in another unseen domain. Unsupervised domain adaptation (UDA) semantic segmentation aims to adapt the semantic segmentation model trained on the labeled source domain to an unlabeled target domain. However, the existing UDA semantic segmentation models tend to focus only on the semantic mask information of the building and neglect modeling the structural and boundary information of the building, which leads to irregularities and inaccuracies of prediction results. We propose a dual-contrast adaptation network coupling global context and geometry information for cross-domain building extraction of HRSIs. It first uses a hierarchical transformer to extract multilevel global context information of buildings from HRSI and adopts a multilevel context fusion module to perform feature fusion on the context information extracted by the transformer at different levels to improve the model’s ability to learn domain-invariant knowledge. Then, it employs hierarchical supervised signals of attraction field map (AFM) and masks to guide the transformer-based network to migrate the building knowledge in a shape-aware manner. Finally, it implements prototype contrast learning (CL) on high-level mask prediction and AFM prediction to bridge the discrepancy between source and target domain on mask semantic information and geometry information, respectively. Extensive experiments under four cross-domain tasks indicate that the proposed method is remarkably superior to the state-of-the-art methods.
Jingru Zhu, Geng Sun 0005, Zhenbo Gou, Ya Guo 0002, Jie Chen 0048
IEEE Trans. Geosci. Remote. Sens.7
2024 Unsupervised Domain Adaptation for Building Extraction of High-Resolution Remote Sensing Imagery Based on Decoupling Style and Semantic Features
abstract
When the buildings themselves or the environments they are located in changes, it poses a cross-domain problem for building extraction. Since the semantics of same category should be consistent across domains, a memory mechanism can be used to drive models to capture it. A key prerequisite for a satisfactory memory mechanism is that the memory content is highly task-relevant, i.e., the memory mechanism represents domain-invariant semantic features of buildings as much as possible. It is worth noting that images from different domains have diverse styles, which can interfere with the representation of domain-invariant semantic features of buildings. To maximize the effectiveness of memory mechanisms in cross-domain building extraction tasks, it is necessary to reduce the interference of style information. Therefore, we propose a cross-domain extraction method for buildings based on decoupling style-semantic feature. It expresses the image style features and building semantic features separately and optimizes the domain-invariant features of buildings by constraining their relationship through an orthogonal loss. Specifically, on the one hand, style features within the same domain as well as semantic features within and across domains are narrowed down through contrastive learning to ensure the consistency of style features extracted within the same domain and the consistency of semantics across domains. On the other hand, the style features and domain-invariant semantic features in memory mechanism can be stored and reused to fully exploit self-supervised information. Results of the cross-domain experiments show that the proposed method can achieve optimal building extraction.
Jie Chen 0048, Jingru Zhu, Peien He, Ya Guo 0002, Yin Yang 0003, Geng Sun 0005
IEEE Trans. Geosci. Remote. Sens.1
2024 Causal Prototype-Inspired Contrast Adaptation for Unsupervised Domain Adaptive Semantic Segmentation of High-Resolution Remote Sensing Imagery
abstract
Semantic segmentation of high-resolution remote sensing imagery (HRSI) suffers from the domain shift, resulting in poor performance of the model in another unseen domain. Unsupervised domain adaptive (UDA) semantic segmentation aims to adapt the semantic segmentation model trained on the labeled source domain to an unlabeled target domain. However, the existing UDA semantic segmentation models tend to align pixels or features based on statistical information related to labels in source and target domain data and make predictions accordingly, which leads to uncertainty and fragility of prediction results. In this article, we propose a causal prototype-inspired contrast adaptation (CPCA) method to explore the invariant causal mechanisms between different HRSIs domains and their semantic labels. It first disentangles causal features and bias features from the source and target domain images through a causal feature disentanglement (CFD) module. Then, a causal prototypical contrast (CPC) module is used to learn domain invariant causal features. To further de-correlate causal and bias features, a causal intervention (CI) module is introduced to intervene on the bias features to generate counterfactual unbiased samples. By forcing the causal features to meet the principles of separability, invariance, and intervention, CPCA can simulate the causal factors of source and target domains, and make decisions on the target domain based on the causal features, which can observe improved generalization ability. Extensive experiments under three cross-domain tasks indicate that CPCA is remarkably superior to the state-of-the-art methods.
Jingru Zhu, Ya Guo 0002, Geng Sun 0005, Jie Chen 0048
IEEE Trans. Geosci. Remote. Sens.5
2023 Memory-Contrastive Unsupervised Domain Adaptation for Building Extraction of High-Resolution Remote Sensing Imagery
abstract
Deep learning-based semantic segmentation has been widely applied for building extraction. However, due to the domain gap, the extraction of building in high-resolution remote sensing imagery is difficult when the model trained on a source dataset is directly used to test on a target data. Considering that humans can retrieve memory to deal with correlative tasks in different domains, memory mechanisms have been developed effectively to assist cross-domain feature extraction. However, whether the memory mechanisms can achieve satisfactory result or not highly depends on the premise that the memory is relevant to the task. Therefore, the domain-invariant memory is crucial in cross-domain building extraction task. To this end, a memory-contrastive unsupervised domain adaptation method is proposed on the basis of a novel memory mechanism. Specifically, to facilitate the model to memorize domain-invariant features, we first conduct a normalization-based image style transfer strategy and a discriminator-based adversarial method at the image level and feature level, respectively. Subsequently, we carry out a memory-contrastive module to obtain domain-invariant features. Especially, a teacher–student network is exploited to help knowledge transferring by knowledge distillation to enhance the performance of the memory-contracted module. To narrow the distance between the two domains, a memory bank is designed to store and update category features obtained from the source domain, and then the similarity between category features in the target domain and memory bank is calculated. Results of the cross-domain experiments show that the proposed method can achieve optimal building extraction. (GitHub: https://github.com/RS-CSU/MDANet).
Jie Chen 0048, Peien He, Jingru Zhu, Ya Guo 0002, Geng Sun 0005, Haifeng Li 0007
IEEE Trans. Geosci. Remote. Sens.1
2023 Unsupervised Domain Adaptation Semantic Segmentation of High-Resolution Remote Sensing Imagery With Invariant Domain-Level Prototype Memory
abstract
Semantic segmentation is a key technique involved in automatic interpretation of high-resolution remote sensing (HRS) imagery and has drawn much attention in the remote sensing community. Deep convolutional neural networks (DCNNs) have been successfully applied to the HRS imagery semantic segmentation task due to their hierarchical representation ability. However, the heavy dependence on a large number of training data with dense annotation and the sensitiveness to the variation of data distribution severely restrict the potential application of DCNNs for the semantic segmentation of HRS imagery. This study proposes a novel unsupervised domain adaptation semantic segmentation network (MemoryAdaptNet) for the semantic segmentation of HRS imagery. MemoryAdaptNet constructs an output space adversarial learning scheme to bridge the domain distribution discrepancy between the source domain and the target domain and to narrow the influence of domain shift. Specifically, we embed an invariant feature memory module to store invariant domain-level prototype information because the features obtained from adversarial learning only tend to represent the variant feature of current limited inputs. This module is integrated by a category attention-driven invariant domain-level memory aggregation module to current pseudo-invariant feature for further augmenting the representations. An entropy-based pseudo label filtering strategy is used to update the memory module with high-confident pseudo-invariant feature of current target images. Extensive experiments under three cross-domain tasks indicate that our proposed MemoryAdaptNet is remarkably superior to the state-of-the-art methods. Our code is available athttps://github.com/RS-CSU/MemoryAdaptNet-master.
Jingru Zhu, Ya Guo 0002, Geng Sun 0005, Libo Yang, Jie Chen 0048
IEEE Trans. Geosci. Remote. Sens.6
2022 Message-Passing-Driven Triplet Representation for Geo-Object Relational Inference in HRSI
abstract
A high-resolution remote sensing image (HRSI) scene typically contains multiple geo-objects, and geospatial relations among these geo-objects are obvious. As the important information conveyed by HRSI, the intelligent expression of geospatial relation is helpful in understanding HRSI scenes. Previous HRSI semantic understanding was mainly based on image captions that only generate one sentence to describe image content, thereby resulting in insufficient understanding of the scene. Thus, the present letter proposes an approach to represent geospatial relations in an HRSI scene with structured form of$\langle $subject, geospatial relation, object$\rangle $. A geospatial relation triplet representation data set that contains visual and semantic information, such as category, location, and geospatial relations of the geo-objects, is constructed first. An “object-relation” message-passing mechanism is adopted to enhance the information exchange between the geo-objects and geospatial relations to predict triplets accurately. The experimental results show that the proposed method can effectively predict the geospatial relation in a HRSI scene.
Jie Chen 0048, Yi Zhang 0064, Geng Sun 0005, Haifeng Li 0007
IEEE Geosci. Remote. Sens. Lett.1
2022 Spatial-Temporal Invariant Contrastive Learning for Remote Sensing Scene Classification
abstract
Self-supervised learning achieves close to supervised learning results on remote sensing image (RSI) scene classification. This is due to the current popular self-supervised learning methods that learn representations by applying different augmentations to images and completing the instance discrimination task which enables convolutional neural networks (CNNs) to learn representations invariant to augmentation. However, RSIs are spatial-temporal heterogeneous, which means that similar features may exhibit different characteristics in different spatial-temporal scenes. Therefore, the performance of CNNs that learn only representations invariant to augmentation still degrades for unseen spatial-temporal scenes due to the lack of spatial-temporal invariant representations. We propose a spatial-temporal invariant contrastive learning (STICL) framework to learn spatial-temporal invariant representations from unlabeled images containing a large number of spatial-temporal scenes. We use optimal transport to transfer an arbitrary unlabeled RSI into multiple other spatial-temporal scenes and then use STICL to make CNNs produce similar representations for the views of the same RSI in different spatial-temporal scenes. We analyze the performance of our proposed STICL on four commonly used RSI scene classification datasets, and the results show that our method achieves better performance on RSIs in unseen spatial-temporal scenes compared to popular self-supervised learning methods. Based on our findings, it can be inferred that spatial-temporal invariance is an indispensable property for a remote sensing model that can be applied to a wider range of remote sensing tasks, which also inspires the study of more general remote sensing models. The source code is available athttps://github.com/GeoX-Lab/G-RSIM/tree/main/T-SC-STICL.
Haozhe Huang, Zhongfeng Mou, Yunying Li, Qiujun Li, Jie Chen 0048, Haifeng Li 0007
IEEE Geosci. Remote. Sens. Lett.5
2022 Improving Few-Shot Remote Sensing Scene Classification With Class Name Semantics
abstract
Few-shot remote sensing scene classification (FSRSSC) has been used for new class recognition in the presence of a limited number of labeled samples. The representation vector (prototype) of categories obtained using images only confronts some challenges, such as insufficient generalization when the number of samples is too small. To address this problem, we propose a new FSRSSC method based on prototype networks, named CNSPN, which combines semantic information of class names (name of the scene categories, such as aircraft, harbor, and bridge). First, CNSPN extracts semantics for class names using a pre-trained word-embedding model, which enriches the feature representation ability of the category at the source. Then, an enhanced fusion prototype is generated by fusing the semantic information of text and visual information in the image through a multimodal prototype fusion module (MPFM). Finally, the query image is classified by measuring the distance between the query sample and the visual prototype, and between the query sample and the fusion prototype. Comparative experiments on the NWPU-RESISC45 and RSD46-WHU datasets show that the proposed method significantly improves FSRSSC performance. Code is available at https://github.com/RS-CSU/CNSPN.git.
Jie Chen 0048, Ya Guo 0002, Jingru Zhu, Geng Sun 0005, Dengda Qin
IEEE Trans. Geosci. Remote. Sens.1
2022 Contextual Information-Preserved Architecture Learning for Remote-Sensing Scene Classification
abstract
Convolutional neural networks (CNNs) have recently been widely used in remote-sensing scene classification. Additionally, it is becoming very popular to automatically learn specific CNN architectures for specific data sets. The rich contextual information in high-resolution remote-sensing images (RSIs) is critical to remote-sensing intelligent understanding tasks. However, architecture learning approaches tend to simplify the original data (i.e., resizing images to smaller resolution) for efficiency, yet result in contextual information loss of RSIs. In this article, we proposed a contextual information-preserved architecture learning (CIPAL) framework for remote-sensing scene classification to utilize the contextual information in RSIs as much as possible during the architecture learning process. We introduce channel compression into CIPAL, which can reduce the memory and time consumption of architecture learning and make it possible to construct a larger architecture space. We add potential operators that are rarely used for scene classification tasks (i.e., atrous convolution) into the architecture space to explore unknown architectures that are more suitable for remote-sensing scenes. The experimental results on four remote-sensing scene classification benchmarks indicate that CIPAL learns architectures with less time consumption than similar works, and the newly found architectures outperform popular hand-designed architectures for better use of contextual information in RSIs. Different architectures are good at learning different representations, and our proposed architecture learning method potentially helps us understand which types of representations are crucial for RSI intelligent understanding.
Jie Chen 0048, Haozhe Huang, Jian Peng 0009, Li Chen 0025, Chao Tao 0001, Haifeng Li 0007
IEEE Trans. Geosci. Remote. Sens.1
2022 Multiscale Object Contrastive Learning-Derived Few-Shot Object Detection in VHR Imagery
abstract
The feature representation capability of the object detection model is considerably reduced if the available training samples are few-shot. The challenges of arbitrary orientation and complex background of ground objects are universal in very high spatial resolution remote sensing imageries, resulting in massive difficulty on few-shot object detection task. However, existing methods for few-shot object detection are not explored in terms of the capabilities of feature representation in remote sensing images. To solve these issues, we propose a few-shot object detection method incorporating a multiscale object contrastive learning. First, our method performs contrastive learning to represent the object feature fully by adopting Siamese network structure in the few-shot training. On the one hand, the Siamese structure’s lower branch is embedded in the contrastive learning process to cope with the challenge of the complexity of images; on the other hand, a multiscale instance feature module is designed to obtain multiscale contrastive information. Second, we leverage the proposed contrastive multiscale proposal loss (CMSP Loss) to make full use of multiscale information. It can promote our method to fit the data better. The experimental results show that the proposed method has good feature representation capabilities in the few-shot object detection. Moreover, the performance of the proposed method is better than that of related methods.
Jie Chen 0048, Dengda Qin, Dongyang Hou, Geng Sun 0005
IEEE Trans. Geosci. Remote. Sens.1
2022 Unsupervised Domain Adaptation for Semantic Segmentation of High-Resolution Remote Sensing Imagery Driven by Category-Certainty Attention
abstract
Semantic segmentation is an important task of analysis and understanding of high-resolution remote sensing images (HRSIs). The deep convolutional neural network (DCNN)-based model shows their excellent performance in remote sensing image semantic segmentation. Most of the existing HRSI semantic segmentation methods are only designed for a very limited data domain, that is, the training and test images are from the same dataset. The accuracy drops sharply once a model trained on a certain dataset is used for cross-domain prediction due to the difference in feature distribution of the dataset. To this end, this article proposes an unsupervised domain adaptation framework based on adversarial learning for HRSI semantic segmentation. This framework uses high-level feature alignment to narrow the difference between the source and target domains at the semantic level. It uses the category-certainty attention module to reduce the attention of the classifier on category-level aligned features and increase the attention on category-level unaligned features. Experimental results show that the proposed method performs favorably against the state-of-the-art methods in cross-domain segmentation.
Jie Chen 0048, Jingru Zhu, Ya Guo 0002, Geng Sun 0005, Yi Zhang 0064
IEEE Trans. Geosci. Remote. Sens.1
2021 SMAF-Net: Sharing Multiscale Adversarial Feature for High-Resolution Remote Sensing Imagery Semantic Segmentation
abstract
Semantic segmentation of high-resolution remote sensing imagery (HRSI) is a major task in remote sensing analysis. Although deep convolutional neural network (DCNN)-based semantic segmentation models have powerful capacity in pixel-wise classification, they still face challenge in obtaining intersemantic continuity and extraboundary accuracy because of the geo-object’s characteristic feature of diverse scales and various distributions in HRSI. Inspired by the transfer learning, in this study, we propose an efficient semantic segmentation framework named SMAF-Net, which shares multiscale adversarial features into a U-shaped semantic segmentation model. Specifically, it uses multiscale adversarial feature representation obtained from a well-trained generative adversarial network to grasp the pixel correlation and further improve the boundary accuracy of multiscale geo-objects. Comparison experiments on the Potsdam and Vaihingen data sets demonstrate that the proposed framework can achieve considerable improvement in the semantic segmentation of HRSI.
Jie Chen 0048, Jingru Zhu, Geng Sun 0005
IEEE Geosci. Remote. Sens. Lett.1
2020 Multi-Scale Spatial and Channel-wise Attention for Improving Object Detection in Remote Sensing Imagery
abstract
The spatial resolution of remote sensing images is continuously improved by the development of remote sensing satellite and sensor technology. Hence, background information in an image becomes increasingly complex and causes considerable interference to the object detection task. Can we pay as much attention to the object in an image as human vision does? This letter proposes a multi-scale spatial and channel-wise attention (MSCA) mechanism to answer this question. MSCA has two advantages that help improve object detection performance. First, attention is paid to the spatial area related to the foreground, and compared with other channels, more attention is given to the feature channel with a greater response to the foreground region. Second, for objects with different scales, MSCA can generate an attention distribution map that integrates multi-scale information and applies it to the feature map of the deep network. MSCA is a flexible module that can be easily embedded into any object detection model based on deep learning. With the attention exerted by MSCA, the deep neural network can efficiently focus on objects of different backgrounds and sizes in remote sensing images. Experiments show that the mean average precision of object detection is improved after the addition of MSCA to the current object detection model.
Jie Chen 0048, Jingru Zhu
IEEE Geosci. Remote. Sens. Lett.1