EDBT 2026 Demo / reviewers in the wild / expert
Qixiong Wang
dblp:285/7078
· DBLP profile ↗
14ranked-venue papers
4as first author
13since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 12 · 4 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Prototype-Guided Structural Learning from Visual Foundation Model for Few-Shot Aerial Image Semantic SegmentationabstractFew-shot aerial image semantic segmentation aims to segment query images with few annotated support samples. It is challenging due to intra-class variations and complex object details in remote aeiral images. However, these two issues are inadequately addressed in existing few-shot segmentation methods. In this paper, we propose a novel Prototype-Guided structural learning (PGSL) framework based on recently proposed segment anything model (SAM). Specifically, to accommodate intra-class variation in aerial image, a novel Prototype-Guided transformer is designed to interact the multiple prototypes from support images with query images, yielding initial segmentation map. Moreover, to improve the performance on object contours, we propose a refine branch based on the SAM, which adopts initial segmentation maps as prompt. This integrates the structural knowledge inherent in SAM into our model. Experiment on iSAID-5i dataset demonstrates the proposed PGSL framework outperforms other state-of-the-art methods. Qixiong Wang, Hongxiang Jiang, Jiaqi Feng 0001, Guangyun Zhang, Jihao Yin |
IGARSS | 1 |
| 2024 | S2JO: Spatial-Spectral Joint Optimization for Hyperspectral Image Classification With Noisy LabelsabstractHyperspectral image (HSI) annotation often suffers from noisy labels, which brings challenges in classification models training. Existing methods typically address this issue through two separate stages: noisy labels cleaning and robust model design. However, model performance is constrained by insufficient noisy label filtering. To explore an end-to-end joint optimization framework for HSI classification with noisy labels, we propose a spatial–spectral joint optimization (S2JO) network. The S2JO consists of a spatial nonuniform sampling (SNS) module and a spectral prototypes learning (SPL) module. The SNS module filters noisy labels dynamically by picking out the top-N confident points, purifying the input samples for the network. Meanwhile, the SPL module uses multicenter spectral prototypes to extract discriminate features accurately even if the input contains some noisy labels. Extensive experiments on the C2Seg-Beijing HSI subdataset demonstrate the superiority of the proposed S2JO over other state-of-the-art methods. Jiaqi Feng 0001, Qixiong Wang, Hongxiang Jiang, Guangyun Zhang, Jihao Yin |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2024 | Scene-Object Holistic Relation Network for Fine-Grained Airplane DetectionabstractThe airplane detection and fine-grained recognition in the remote sensing images are challenging due to high interclass indistinction. The subtle distinctions between classes make it difficult to accurately classify objects based purely on bounding box features without considering the broader context. However, recent studies on remote sensing object detection focuses on refining the representation of bounding boxes while ignoring holistic context knowledge in remote sensing scenarios. This letter addresses this gap by introducing the scene-object holistic relation (SOHR) network for fine-grained airplane detection. Specifically, the SOHR network distinctively exploits global scene-object context information through a novel lightweight scene context attention (SCA) module, which aggregates scene context feature and object position information. Furthermore, the object relation transformer (ORT) is designed to model interactions among all objects within the scene explicitly, thereby increasing the model performance for ambiguous hard samples. The experimental results obtained from the FAIR1M dataset demonstrate that the proposed SOHR-Net achieves a state-of-the-art detection accuracy of 56.110% mean average precision (mAP). Compared with the baseline, SOHR-Net exhibits an increase of 2.517%. Weiyu Ning, Qixiong Wang, Jiaqi Feng 0001, Hongxiang Jiang, Guangyun Zhang, Jihao Yin |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2024 | CAT: Center Attention Transformer With Stratified Spatial-Spectral Token for Hyperspectral Image ClassificationabstractMost hyperspectral image (HSI) classification methods rely on square patch sampling to incorporate spatial information, thereby facilitating the label prediction of the center pixel. However, square patch sampling introduces numerous heterogeneous pixels, which could distort the label prediction of center pixel. Moreover, it generates fixed training patch sample for each center pixel, hampering the performance of transformer-based models requiring a large number of training data. To address the above problems, we proposed Center Attention Transformer (CAT) with stratified spatial-spectral token generated by superpixel sampling for HSI classification. Firstly, to mitigate the inference of heterogeneous pixels, we propose Sampling From Superpixel Region mechanism to generate purer image cubes than traditional square neighborhood. Secondly, to expand the training data for transformer, we propose Multiple Stratified Random Sampling mechanism, which generates ample training samples without introducing additional labels. Finally, to more effectively extract information from the sampled patch tokens, we propose Spatial Spectral Token Generation mechanism and Center Attention Transformer structure with Gaussian Positional Embedding. This framework can extract long-range correlations of spectral information and pay more attention on the center pixel in spatial dimension. Experimental results on three HSI datasets demonstrate the performance of our proposed method CAT outperforms several state-of-the-art methods. The code of this work is available at https://github.com/fengjiaqi927/CAT-Center_Attention_Transformer. Jiaqi Feng 0001, Qixiong Wang, Guangyun Zhang, Xiuping Jia, Jihao Yin |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Balanced Orthogonal Subspace Separation Detector for Few-Shot Object Detection in Aerial ImageryabstractFew-shot object detection (FSOD) in remote sensing images (RSIs) aims to achieve object location and classification with only a few training samples. Currently, mainstream transfer-learning methods employ a two-stage approach: pretraining on data-abundant base classes and fine-tuning on few-shot novel classes. However, existing approaches suffer notable degradation in both base and novel classes during fine-tuning, because of gradient conflict and class imbalance. To address this, we construct the balanced orthogonal subspace separation (BOSS) detector, a novel two-stage framework for FSOD. Specifically, to avoid contradictory gradients, BOSS distinctly isolates the training of base and novel classes at both structural and feature levels. For structural separation, a low-rank subspace adapter (LoSA) is introduced to ensure network optimization for novel classes without hampering base classes’ pretraining performance, effectively addressing over-fitting in few-shot scenarios. For feature disentanglement, an orthogonal subspace extractor (OSE) is presented, enhancing class separability by learning class-specific, orthogonal basis-spanned subspace. Finally, a balanced classifier (BC) is proposed to equalize the imbalanced loss, with its dual-component design mitigating bias toward predicting background or base classes. Comparative evaluations on diverse remote sensing datasets demonstrate BOSS’s superiority, outperforming state-of-the-art methods in mean average precision (mAP). These results underscore BOSS’s effectiveness in FSOD, particularly in challenging remote sensing contexts. Hongxiang Jiang, Qixiong Wang, Jiaqi Feng 0001, Guangyun Zhang, Jihao Yin |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Disentangled Foreground-Semantic Adapter Network for Generalized Aerial Image Few-Shot Semantic SegmentationabstractSemantic segmentation of remote sensing imagery requires extensive annotated samples for training, facing challenges in adapting to novel classes with few annotations. Few-shot semantic segmentation (FS-Seg) employs a support-to-query paradigm, which encounters many practical constraints. Recently, generalized few-shot semantic segmentation (GFS-Seg) has been proposed to align with general semantic segmentation paradigms, enabling the segmentation of all classes (both base and novel classes) in the image. However, existing GFS-Seg methods struggle with a large intra-class variance of background, degradation on base classes, and overfitting on novel classes during fine-tuning in aerial imagery. To address the above issues, we propose the disentangled foreground-semantic adapter network (DFSA-Net) for generalized aerial image FS-Seg. Specifically, to reduce the interference from background features, DFSA-Net employs a foreground-semantic decoder (FSD) to decompose semantic segmentation into foreground aggregation and multiclass refinement. To mitigate the base classes degradation and novel class overfitting during fine-tuning, we propose disentangled low-rank adapter (DLA) for fine-tuning phase, designed to preserve the base parameters while ensuring efficient adaptation to novel classes. Finally, we introduce an inference ensemble strategy that merges base and novel decoder prediction to achieve final output. Experimental results on NWPU and iSAID datasets demonstrate the superiority of our DFSA-Net over other compared methods. Qixiong Wang, Jihao Yin, Hongxiang Jiang, Jiaqi Feng 0001, Guangyun Zhang |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | Beyond One-to-One: Rethinking the Referring Image SegmentationabstractReferring image segmentation aims to segment the target object referred by a natural language expression. However, previous methods rely on the strong assumption that one sentence must describe one target in the image, which is often not the case in real-world applications. As a result, such methods fail when the expressions refer to either no objects or multiple objects. In this paper, we address this issue from two perspectives. First, we propose a Dual Multi-Modal Interaction (DMMI) Network, which contains two decoder branches and enables information flow in two directions. In the text-to-image decoder, text embedding is utilized to query the visual feature and localize the corresponding target. Meanwhile, the image-to-text decoder is implemented to reconstruct the erased entity-phrase conditioned on the visual feature. In this way, visual features are encouraged to contain the critical semantic information about target entity, which supports the accurate segmentation in the text-to-image decoder in turn. Secondly, we collect a new challenging but realistic dataset called Ref-ZOM, which includes image-text pairs under different settings. Extensive experiments demonstrate our method achieves state-of-the-art performance on different datasets, and the Ref-ZOM-trained model performs well on various types of text inputs. Codes and datasets are available at https://github.com/toggle1995/RIS-DMMI. Yutao Hu 0002, Qixiong Wang, Wenqi Shao, Enze Xie, Zhenguo Li, Jungong Han, Ping Luo 0002 |
ICCV | 2 |
| 2023 | Multiscale Prototype Contrast Network for High-Resolution Aerial Imagery Semantic SegmentationabstractSemantic segmentation of high-resolution aerial images is a challenging task on account of complex scene-variation and large scale-difference. However, these two issues are inadequately addressed in general semantic segmentation methods. In this paper, we propose a Multi-scale Prototype Contrast Network (MPCNet) to improve the adaptive capability for different scenes and scales. Specifically, a novel multi-scale prototype transformer decoder (MPTD) is designed to extract dynamic scene-specific prototypes as pixel classifier by fusing information of feature maps and learnable class tokens. To exploit cross-scene context information and accommodate the large scale-difference in aerial image, we build a multi-scale prototype memory queue to store these multi-scale prototypes during training. Upon the multi-scale prototype memory queue, a novel multi-scale prototype contrastive loss is proposed to increase object feature discriminability across multiple scale, which brings better consistency of intermediate feature and boosts the convergence of network. Extensive experimental results on three publicly available datasets demonstrate the effectiveness and efficiency of our MPCNet over other state-of-the-art methods. The code is available at https://github.com/qixiong-wang/mmsegmentation-mpcnet. Qixiong Wang, Xiaoyan Luo, Jiaqi Feng 0001, Guangyun Zhang, Xiuping Jia, Jihao Yin |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Spectral Transformer with Dynamic Spatial Sampling and Gaussian Positional Embedding for Hyperspectral Image ClassificationabstractOwing to the global information extraction ability, transformers have been tentatively applied to hyperspectral image(HSI) classification. However, the existing transformer-based methods have not made full use of the flexible characteristics of spatial sampling nor considered the importance of the central pixel to the classification of HSI cubes. In order to enhance adaptability of transformers for HSI classification, we have proposed a novel spectral transformer with dynamic spatial sampling and gaussian positional embedding. To improve the effectiveness of spatial neighborhood information, Spatial Sample Selection(3S) mechanism generates image cube from super pixel region, making image cube more pure for classification. To extract long-range information in spectral dimension, Spectral Feature Extraction(SFE) network splits spectral bands into several slices and calculates the attention between them. To stress the importance of the central pixel to the classification of image cube, Gaussian Positional Embedding(GPE) reduces the weight of surrounding pixels during feature embedding stage. Experimental results demonstrate the performance of our proposed method. The code of this work is available at https://github.com/fengjiaqi927/HSI_transformer. Jiaqi Feng 0001, Xiaoyan Luo, Qixiong Wang, Jihao Yin |
IGARSS | 4 |
| 2022 | Class-Balanced Contrastive Learning for Fine-Grained Airplane DetectionabstractAirplane detection and fine-grained recognition in remote sensing images are challenging due to class imbalance and high inter-class indistinction. To alleviate these issues, we propose class-balanced contrastive learning (CBCL) approach for airplane detection to exploit the correlation between samples in different images, which is rarely explored in previous research. Specifically, we first dynamically build class-balanced memory queues during training, which mitigates class imbalance by memorizing training samples. Upon class-balanced memory queues, hard triplet contrastive learning is introduced to increase the inter-class discriminability, which enforces the maximum distance of the positive sample pair to be smaller than the minimum distance of the negative sample pair. We integrate the proposed CBCL strategy into oriented object detection frameworks for fine-grained airplane detection. The experimental results on the FAIR1M dataset reveal that several state-of-the-art algorithms with CBCL achieve significantly improvements. Qixiong Wang, Xiaoyan Luo, Jihao Yin |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2022 | H2AN: Hierarchical Homogeneity-Attention Network for Hyperspectral Image ClassificationabstractRecently, a self-attention network (SAN) is developed as an effective strategy to extract features from attention areas for image classification. However, for hyperspectral image (HSI), the lack of position supervision of object regions and inefficient similarity computation lead to unsatisfactory classification performance on mixed pixels. To alleviate the above two problems for HSI image classification, we propose a novel hierarchical homogeneity-attention network (H2AN) in this article. First, we design a homogeneity-attention block (HAB) to depict the feature correlation with the homogeneous mask. Using the supervision of homogeneity mask, we can calculate the attention guided by the predefined number of homogeneity embeddings, which can reduce the heavy computation instead of the global search in self-attention block (SAB) of SAN. Second, we propose a hierarchical convolutional neural network (HCNN) inserting the HAB into different levels of network cells for highly efficient feature extraction of target regions, named H2AN. Because of the transferring of homogeneity property from shallow layer to deep layer, our H2AN outperforms the state-of-the-art methods in qualitative and quantitative experiments on three typical datasets. Xiaoyan Luo, Qixiong Wang, Jihao Yin |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2021 | Hyperspectral Classification Using Cooperative Spatial-Spectral Attention Network With Tensor Low-Rank ReconstructionabstractSpatial and spectral attention networks have been both well introduced to Hyperspectral image (HSI) classification. However, in previous works, they are seldom considered jointly. To obtain a 3D spatial-spectral attention map, which is beneficial for extracting discriminative spatial-spectral features, we propose a novel cooperative spatial-spectral attention network with tensor low-rank reconstruction. Firstly, a tensor low-rank reconstruction (TLRR) block is designed to learn a spatial-spectral attention map tensor, which adaptively emphasizes the attention features of the salient spatial positions and informative spectral bands simultaneously. Secondly, these attention features are merged into simple convolutional features which are more discriminative for classification. Finally, the experimental results demonstrate that our proposed method outperforms some state-of-the-art methods on two typical HSI datasets. Xiaoyan Luo, Qixiong Wang, Weifa Shen, Jihao Yin |
ICIP | 3 |
| 2021 | Unsupervised Domain Adaptation for Semantic Segmentation via Self-SupervisionabstractRecently, deep learning (DL) methods have been widely used for semantic segmentation of remote sensing and achieved significant progress. However, DL-based methods are time-consuming and labor intensive for the networks requiring abundant data with accurate labeling. To solve this issue, unsupervised domain adaption (UDA) has recently been used to transfer the information from labeled source domain to unlabeled target domain. In this paper, we propose a novel UDA approach based on the self-supervised theory for remote sensing image. Specifically, we firstly utilize the inter-domain adaptation to reduce the gap between the source and target domain. Secondly, based on our proposed spatial-frequency (SF) index, we detach the target domain into an easy and hard split. Ultimately, we adopt the intra-domain adaptation by self-supervised adaptation to improve the performance of hard split. Experimental results on ISPRS Vaihingen and Potsdam datasets demonstrate the effectiveness and rationality of our methods against the other state-of-the-art approaches. Weifa Shen, Qixiong Wang, Hongxiang Jiang, Jihao Yin |
IGARSS | 2 |
| 2020 | Neural Network Pruning for Hyperspectral Image Band SelectionabstractNeural network pruning attempts to reduce parameters without hurting original performance by inducing connection matrix sparsity of network. Inspired by this idea, we proposed an effective pruning-based band selection strategy, which is a potent feature extraction tool in hyperspectral image (HSI) classification. At first, we take the whole HSI bands as input to train original network parameters. For each band, all parameters in network are integrated to measure the band importance. With the novel band signification factor constraining, then the convolutional neural network (CNN) is pruned and remains some representative weights to retrain the compact sub-network, which can finally deal with the hyperspectral band selection problem. Experimental results on the real HSI dataset demonstrate that network pruning-based method can outperform the original CNN in classification accuracy. Also, it can achieve the superiority over filter-based and other CNN-based band selection algorithms in classification accuracy. Our code is available at https://github.com/qixiong-wang/Network-pruning-for-HSI-band-selection. Qixiong Wang, Xiaoyan Luo, Jihao Yin |
IGARSS | 1 |