Jiaqi Feng 0001

dblp:94/7627-1 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
8since 2021 · last 2024
0009-0005-4185-3856ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 8 · 3 first-author · 8 since 2021
YearPublicationVenuePosition
2024 Prototype-Guided Structural Learning from Visual Foundation Model for Few-Shot Aerial Image Semantic Segmentation
abstract
Few-shot aerial image semantic segmentation aims to segment query images with few annotated support samples. It is challenging due to intra-class variations and complex object details in remote aeiral images. However, these two issues are inadequately addressed in existing few-shot segmentation methods. In this paper, we propose a novel Prototype-Guided structural learning (PGSL) framework based on recently proposed segment anything model (SAM). Specifically, to accommodate intra-class variation in aerial image, a novel Prototype-Guided transformer is designed to interact the multiple prototypes from support images with query images, yielding initial segmentation map. Moreover, to improve the performance on object contours, we propose a refine branch based on the SAM, which adopts initial segmentation maps as prompt. This integrates the structural knowledge inherent in SAM into our model. Experiment on iSAID-5i dataset demonstrates the proposed PGSL framework outperforms other state-of-the-art methods.
Qixiong Wang, Hongxiang Jiang, Jiaqi Feng 0001, Guangyun Zhang, Jihao Yin
IGARSS3
2024 S2JO: Spatial-Spectral Joint Optimization for Hyperspectral Image Classification With Noisy Labels
abstract
Hyperspectral image (HSI) annotation often suffers from noisy labels, which brings challenges in classification models training. Existing methods typically address this issue through two separate stages: noisy labels cleaning and robust model design. However, model performance is constrained by insufficient noisy label filtering. To explore an end-to-end joint optimization framework for HSI classification with noisy labels, we propose a spatial–spectral joint optimization (S2JO) network. The S2JO consists of a spatial nonuniform sampling (SNS) module and a spectral prototypes learning (SPL) module. The SNS module filters noisy labels dynamically by picking out the top-N confident points, purifying the input samples for the network. Meanwhile, the SPL module uses multicenter spectral prototypes to extract discriminate features accurately even if the input contains some noisy labels. Extensive experiments on the C2Seg-Beijing HSI subdataset demonstrate the superiority of the proposed S2JO over other state-of-the-art methods.
Jiaqi Feng 0001, Qixiong Wang, Hongxiang Jiang, Guangyun Zhang, Jihao Yin
IEEE Geosci. Remote. Sens. Lett.1
2024 Scene-Object Holistic Relation Network for Fine-Grained Airplane Detection
abstract
The airplane detection and fine-grained recognition in the remote sensing images are challenging due to high interclass indistinction. The subtle distinctions between classes make it difficult to accurately classify objects based purely on bounding box features without considering the broader context. However, recent studies on remote sensing object detection focuses on refining the representation of bounding boxes while ignoring holistic context knowledge in remote sensing scenarios. This letter addresses this gap by introducing the scene-object holistic relation (SOHR) network for fine-grained airplane detection. Specifically, the SOHR network distinctively exploits global scene-object context information through a novel lightweight scene context attention (SCA) module, which aggregates scene context feature and object position information. Furthermore, the object relation transformer (ORT) is designed to model interactions among all objects within the scene explicitly, thereby increasing the model performance for ambiguous hard samples. The experimental results obtained from the FAIR1M dataset demonstrate that the proposed SOHR-Net achieves a state-of-the-art detection accuracy of 56.110% mean average precision (mAP). Compared with the baseline, SOHR-Net exhibits an increase of 2.517%.
Weiyu Ning, Qixiong Wang, Jiaqi Feng 0001, Hongxiang Jiang, Guangyun Zhang, Jihao Yin
IEEE Geosci. Remote. Sens. Lett.3
2024 CAT: Center Attention Transformer With Stratified Spatial-Spectral Token for Hyperspectral Image Classification
abstract
Most hyperspectral image (HSI) classification methods rely on square patch sampling to incorporate spatial information, thereby facilitating the label prediction of the center pixel. However, square patch sampling introduces numerous heterogeneous pixels, which could distort the label prediction of center pixel. Moreover, it generates fixed training patch sample for each center pixel, hampering the performance of transformer-based models requiring a large number of training data. To address the above problems, we proposed Center Attention Transformer (CAT) with stratified spatial-spectral token generated by superpixel sampling for HSI classification. Firstly, to mitigate the inference of heterogeneous pixels, we propose Sampling From Superpixel Region mechanism to generate purer image cubes than traditional square neighborhood. Secondly, to expand the training data for transformer, we propose Multiple Stratified Random Sampling mechanism, which generates ample training samples without introducing additional labels. Finally, to more effectively extract information from the sampled patch tokens, we propose Spatial Spectral Token Generation mechanism and Center Attention Transformer structure with Gaussian Positional Embedding. This framework can extract long-range correlations of spectral information and pay more attention on the center pixel in spatial dimension. Experimental results on three HSI datasets demonstrate the performance of our proposed method CAT outperforms several state-of-the-art methods. The code of this work is available at https://github.com/fengjiaqi927/CAT-Center_Attention_Transformer.
Jiaqi Feng 0001, Qixiong Wang, Guangyun Zhang, Xiuping Jia, Jihao Yin
IEEE Trans. Geosci. Remote. Sens.1
2024 Balanced Orthogonal Subspace Separation Detector for Few-Shot Object Detection in Aerial Imagery
abstract
Few-shot object detection (FSOD) in remote sensing images (RSIs) aims to achieve object location and classification with only a few training samples. Currently, mainstream transfer-learning methods employ a two-stage approach: pretraining on data-abundant base classes and fine-tuning on few-shot novel classes. However, existing approaches suffer notable degradation in both base and novel classes during fine-tuning, because of gradient conflict and class imbalance. To address this, we construct the balanced orthogonal subspace separation (BOSS) detector, a novel two-stage framework for FSOD. Specifically, to avoid contradictory gradients, BOSS distinctly isolates the training of base and novel classes at both structural and feature levels. For structural separation, a low-rank subspace adapter (LoSA) is introduced to ensure network optimization for novel classes without hampering base classes’ pretraining performance, effectively addressing over-fitting in few-shot scenarios. For feature disentanglement, an orthogonal subspace extractor (OSE) is presented, enhancing class separability by learning class-specific, orthogonal basis-spanned subspace. Finally, a balanced classifier (BC) is proposed to equalize the imbalanced loss, with its dual-component design mitigating bias toward predicting background or base classes. Comparative evaluations on diverse remote sensing datasets demonstrate BOSS’s superiority, outperforming state-of-the-art methods in mean average precision (mAP). These results underscore BOSS’s effectiveness in FSOD, particularly in challenging remote sensing contexts.
Hongxiang Jiang, Qixiong Wang, Jiaqi Feng 0001, Guangyun Zhang, Jihao Yin
IEEE Trans. Geosci. Remote. Sens.3
2024 Disentangled Foreground-Semantic Adapter Network for Generalized Aerial Image Few-Shot Semantic Segmentation
abstract
Semantic segmentation of remote sensing imagery requires extensive annotated samples for training, facing challenges in adapting to novel classes with few annotations. Few-shot semantic segmentation (FS-Seg) employs a support-to-query paradigm, which encounters many practical constraints. Recently, generalized few-shot semantic segmentation (GFS-Seg) has been proposed to align with general semantic segmentation paradigms, enabling the segmentation of all classes (both base and novel classes) in the image. However, existing GFS-Seg methods struggle with a large intra-class variance of background, degradation on base classes, and overfitting on novel classes during fine-tuning in aerial imagery. To address the above issues, we propose the disentangled foreground-semantic adapter network (DFSA-Net) for generalized aerial image FS-Seg. Specifically, to reduce the interference from background features, DFSA-Net employs a foreground-semantic decoder (FSD) to decompose semantic segmentation into foreground aggregation and multiclass refinement. To mitigate the base classes degradation and novel class overfitting during fine-tuning, we propose disentangled low-rank adapter (DLA) for fine-tuning phase, designed to preserve the base parameters while ensuring efficient adaptation to novel classes. Finally, we introduce an inference ensemble strategy that merges base and novel decoder prediction to achieve final output. Experimental results on NWPU and iSAID datasets demonstrate the superiority of our DFSA-Net over other compared methods.
Qixiong Wang, Jihao Yin, Hongxiang Jiang, Jiaqi Feng 0001, Guangyun Zhang
IEEE Trans. Geosci. Remote. Sens.4
2023 Multiscale Prototype Contrast Network for High-Resolution Aerial Imagery Semantic Segmentation
abstract
Semantic segmentation of high-resolution aerial images is a challenging task on account of complex scene-variation and large scale-difference. However, these two issues are inadequately addressed in general semantic segmentation methods. In this paper, we propose a Multi-scale Prototype Contrast Network (MPCNet) to improve the adaptive capability for different scenes and scales. Specifically, a novel multi-scale prototype transformer decoder (MPTD) is designed to extract dynamic scene-specific prototypes as pixel classifier by fusing information of feature maps and learnable class tokens. To exploit cross-scene context information and accommodate the large scale-difference in aerial image, we build a multi-scale prototype memory queue to store these multi-scale prototypes during training. Upon the multi-scale prototype memory queue, a novel multi-scale prototype contrastive loss is proposed to increase object feature discriminability across multiple scale, which brings better consistency of intermediate feature and boosts the convergence of network. Extensive experimental results on three publicly available datasets demonstrate the effectiveness and efficiency of our MPCNet over other state-of-the-art methods. The code is available at https://github.com/qixiong-wang/mmsegmentation-mpcnet.
Qixiong Wang, Xiaoyan Luo, Jiaqi Feng 0001, Guangyun Zhang, Xiuping Jia, Jihao Yin
IEEE Trans. Geosci. Remote. Sens.3
2022 Spectral Transformer with Dynamic Spatial Sampling and Gaussian Positional Embedding for Hyperspectral Image Classification
abstract
Owing to the global information extraction ability, transformers have been tentatively applied to hyperspectral image(HSI) classification. However, the existing transformer-based methods have not made full use of the flexible characteristics of spatial sampling nor considered the importance of the central pixel to the classification of HSI cubes. In order to enhance adaptability of transformers for HSI classification, we have proposed a novel spectral transformer with dynamic spatial sampling and gaussian positional embedding. To improve the effectiveness of spatial neighborhood information, Spatial Sample Selection(3S) mechanism generates image cube from super pixel region, making image cube more pure for classification. To extract long-range information in spectral dimension, Spectral Feature Extraction(SFE) network splits spectral bands into several slices and calculates the attention between them. To stress the importance of the central pixel to the classification of image cube, Gaussian Positional Embedding(GPE) reduces the weight of surrounding pixels during feature embedding stage. Experimental results demonstrate the performance of our proposed method. The code of this work is available at https://github.com/fengjiaqi927/HSI_transformer.
Jiaqi Feng 0001, Xiaoyan Luo, Qixiong Wang, Jihao Yin
IGARSS1