Guangyun Zhang

dblp:52/10343 · DBLP profile ↗
← Back
20ranked-venue papers
5as first author
13since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 18 · 5 first-author · 12 since 2021Artificial intelligence and machine learning · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2024 Prototype-Guided Structural Learning from Visual Foundation Model for Few-Shot Aerial Image Semantic Segmentation
abstract
Few-shot aerial image semantic segmentation aims to segment query images with few annotated support samples. It is challenging due to intra-class variations and complex object details in remote aeiral images. However, these two issues are inadequately addressed in existing few-shot segmentation methods. In this paper, we propose a novel Prototype-Guided structural learning (PGSL) framework based on recently proposed segment anything model (SAM). Specifically, to accommodate intra-class variation in aerial image, a novel Prototype-Guided transformer is designed to interact the multiple prototypes from support images with query images, yielding initial segmentation map. Moreover, to improve the performance on object contours, we propose a refine branch based on the SAM, which adopts initial segmentation maps as prompt. This integrates the structural knowledge inherent in SAM into our model. Experiment on iSAID-5i dataset demonstrates the proposed PGSL framework outperforms other state-of-the-art methods.
Qixiong Wang, Hongxiang Jiang, Jiaqi Feng 0001, Guangyun Zhang, Jihao Yin
IGARSS4
2024 S2JO: Spatial-Spectral Joint Optimization for Hyperspectral Image Classification With Noisy Labels
abstract
Hyperspectral image (HSI) annotation often suffers from noisy labels, which brings challenges in classification models training. Existing methods typically address this issue through two separate stages: noisy labels cleaning and robust model design. However, model performance is constrained by insufficient noisy label filtering. To explore an end-to-end joint optimization framework for HSI classification with noisy labels, we propose a spatial–spectral joint optimization (S2JO) network. The S2JO consists of a spatial nonuniform sampling (SNS) module and a spectral prototypes learning (SPL) module. The SNS module filters noisy labels dynamically by picking out the top-N confident points, purifying the input samples for the network. Meanwhile, the SPL module uses multicenter spectral prototypes to extract discriminate features accurately even if the input contains some noisy labels. Extensive experiments on the C2Seg-Beijing HSI subdataset demonstrate the superiority of the proposed S2JO over other state-of-the-art methods.
Jiaqi Feng 0001, Qixiong Wang, Hongxiang Jiang, Guangyun Zhang, Jihao Yin
IEEE Geosci. Remote. Sens. Lett.4
2024 Scene-Object Holistic Relation Network for Fine-Grained Airplane Detection
abstract
The airplane detection and fine-grained recognition in the remote sensing images are challenging due to high interclass indistinction. The subtle distinctions between classes make it difficult to accurately classify objects based purely on bounding box features without considering the broader context. However, recent studies on remote sensing object detection focuses on refining the representation of bounding boxes while ignoring holistic context knowledge in remote sensing scenarios. This letter addresses this gap by introducing the scene-object holistic relation (SOHR) network for fine-grained airplane detection. Specifically, the SOHR network distinctively exploits global scene-object context information through a novel lightweight scene context attention (SCA) module, which aggregates scene context feature and object position information. Furthermore, the object relation transformer (ORT) is designed to model interactions among all objects within the scene explicitly, thereby increasing the model performance for ambiguous hard samples. The experimental results obtained from the FAIR1M dataset demonstrate that the proposed SOHR-Net achieves a state-of-the-art detection accuracy of 56.110% mean average precision (mAP). Compared with the baseline, SOHR-Net exhibits an increase of 2.517%.
Weiyu Ning, Qixiong Wang, Jiaqi Feng 0001, Hongxiang Jiang, Guangyun Zhang, Jihao Yin
IEEE Geosci. Remote. Sens. Lett.5
2024 CAT: Center Attention Transformer With Stratified Spatial-Spectral Token for Hyperspectral Image Classification
abstract
Most hyperspectral image (HSI) classification methods rely on square patch sampling to incorporate spatial information, thereby facilitating the label prediction of the center pixel. However, square patch sampling introduces numerous heterogeneous pixels, which could distort the label prediction of center pixel. Moreover, it generates fixed training patch sample for each center pixel, hampering the performance of transformer-based models requiring a large number of training data. To address the above problems, we proposed Center Attention Transformer (CAT) with stratified spatial-spectral token generated by superpixel sampling for HSI classification. Firstly, to mitigate the inference of heterogeneous pixels, we propose Sampling From Superpixel Region mechanism to generate purer image cubes than traditional square neighborhood. Secondly, to expand the training data for transformer, we propose Multiple Stratified Random Sampling mechanism, which generates ample training samples without introducing additional labels. Finally, to more effectively extract information from the sampled patch tokens, we propose Spatial Spectral Token Generation mechanism and Center Attention Transformer structure with Gaussian Positional Embedding. This framework can extract long-range correlations of spectral information and pay more attention on the center pixel in spatial dimension. Experimental results on three HSI datasets demonstrate the performance of our proposed method CAT outperforms several state-of-the-art methods. The code of this work is available at https://github.com/fengjiaqi927/CAT-Center_Attention_Transformer.
Jiaqi Feng 0001, Qixiong Wang, Guangyun Zhang, Xiuping Jia, Jihao Yin
IEEE Trans. Geosci. Remote. Sens.3
2024 Balanced Orthogonal Subspace Separation Detector for Few-Shot Object Detection in Aerial Imagery
abstract
Few-shot object detection (FSOD) in remote sensing images (RSIs) aims to achieve object location and classification with only a few training samples. Currently, mainstream transfer-learning methods employ a two-stage approach: pretraining on data-abundant base classes and fine-tuning on few-shot novel classes. However, existing approaches suffer notable degradation in both base and novel classes during fine-tuning, because of gradient conflict and class imbalance. To address this, we construct the balanced orthogonal subspace separation (BOSS) detector, a novel two-stage framework for FSOD. Specifically, to avoid contradictory gradients, BOSS distinctly isolates the training of base and novel classes at both structural and feature levels. For structural separation, a low-rank subspace adapter (LoSA) is introduced to ensure network optimization for novel classes without hampering base classes’ pretraining performance, effectively addressing over-fitting in few-shot scenarios. For feature disentanglement, an orthogonal subspace extractor (OSE) is presented, enhancing class separability by learning class-specific, orthogonal basis-spanned subspace. Finally, a balanced classifier (BC) is proposed to equalize the imbalanced loss, with its dual-component design mitigating bias toward predicting background or base classes. Comparative evaluations on diverse remote sensing datasets demonstrate BOSS’s superiority, outperforming state-of-the-art methods in mean average precision (mAP). These results underscore BOSS’s effectiveness in FSOD, particularly in challenging remote sensing contexts.
Hongxiang Jiang, Qixiong Wang, Jiaqi Feng 0001, Guangyun Zhang, Jihao Yin
IEEE Trans. Geosci. Remote. Sens.4
2024 Disentangled Foreground-Semantic Adapter Network for Generalized Aerial Image Few-Shot Semantic Segmentation
abstract
Semantic segmentation of remote sensing imagery requires extensive annotated samples for training, facing challenges in adapting to novel classes with few annotations. Few-shot semantic segmentation (FS-Seg) employs a support-to-query paradigm, which encounters many practical constraints. Recently, generalized few-shot semantic segmentation (GFS-Seg) has been proposed to align with general semantic segmentation paradigms, enabling the segmentation of all classes (both base and novel classes) in the image. However, existing GFS-Seg methods struggle with a large intra-class variance of background, degradation on base classes, and overfitting on novel classes during fine-tuning in aerial imagery. To address the above issues, we propose the disentangled foreground-semantic adapter network (DFSA-Net) for generalized aerial image FS-Seg. Specifically, to reduce the interference from background features, DFSA-Net employs a foreground-semantic decoder (FSD) to decompose semantic segmentation into foreground aggregation and multiclass refinement. To mitigate the base classes degradation and novel class overfitting during fine-tuning, we propose disentangled low-rank adapter (DLA) for fine-tuning phase, designed to preserve the base parameters while ensuring efficient adaptation to novel classes. Finally, we introduce an inference ensemble strategy that merges base and novel decoder prediction to achieve final output. Experimental results on NWPU and iSAID datasets demonstrate the superiority of our DFSA-Net over other compared methods.
Qixiong Wang, Jihao Yin, Hongxiang Jiang, Jiaqi Feng 0001, Guangyun Zhang
IEEE Trans. Geosci. Remote. Sens.5
2024 SuperNeRF: High-Precision 3-D Reconstruction for Large-Scale Scenes
abstract
Recent approaches based on neural radiance field (NeRF) showcase remarkable results in the 3-D reconstruction of small-scale scenes by encoding volume density and color observations using implicit functions. However, when confronted with complex and diverse large-scale scenes, it invariably experiences issues such as blurry textures and missing details. In this work, we present a superpixel-based neural radiance field (named SuperNeRF), an additional loss for learning radiance fields that takes advantage of superpixels texture constraints. Building upon NeRF, we leverage superpixels to establish spatial consistency constraints, enabling the precise extraction of 3-D geometry and appearance for large-scale scenes. SuperNeRF is capable of guiding locally adjacent and similar pixels to form nearly consistent ray termination distributions, and it is compatible with the state-of-the-art NeRF-based methods. Comprehensive experiments conducted on representative aviation and aerospace datasets demonstrate that our SuperNeRF exhibits a significant superiority in accuracy over state-of-the-art methods. Code will be available athttps://github.com/xczbecalm/supernerf.
Guangyun Zhang, Chaozhong Xue
IEEE Trans. Geosci. Remote. Sens.1
2023 RADANet: Road Augmented Deformable Attention Network for Road Extraction From Complex High-Resolution Remote-Sensing Images
abstract
Extracting roads from complex high-resolution remote sensing images to update road networks has become a recent research focus. How to apply the contextual spatial correlation and topological structure of the roads properly to improve the extraction accuracy becomes a challenge in the increasingly complex road environment. In this article, inspired by the prior knowledge of the road shape and the progress in deformable convolution, we proposed a road augmented deformable attention network (RADANet) to learn the long-range dependencies for specific road pixels. We developed a road augmentation module (RAM) to capture the semantic shape information of the road from four strip convolutions. Deformable attention module (DAM) combines the sparse sampling capability of deformable convolution with the spatial self-attention mechanism. The integration of RAM enables DAM to extract road features more specifically. Furthermore, RAM is placed behind the fourth stage of encoder, and DAM is placed between last four stages of encoder and decoder in RADANet to extract multiscale road semantic information. Comprehensive experiments on representative public datasets (DeepGlobe and CHN6-CUG road datasets) demonstrate that our RADANet achieves advanced results compared with the state-of-the-art methods.
Guangyun Zhang
IEEE Trans. Geosci. Remote. Sens.2
2023 Multiscale Prototype Contrast Network for High-Resolution Aerial Imagery Semantic Segmentation
abstract
Semantic segmentation of high-resolution aerial images is a challenging task on account of complex scene-variation and large scale-difference. However, these two issues are inadequately addressed in general semantic segmentation methods. In this paper, we propose a Multi-scale Prototype Contrast Network (MPCNet) to improve the adaptive capability for different scenes and scales. Specifically, a novel multi-scale prototype transformer decoder (MPTD) is designed to extract dynamic scene-specific prototypes as pixel classifier by fusing information of feature maps and learnable class tokens. To exploit cross-scene context information and accommodate the large scale-difference in aerial image, we build a multi-scale prototype memory queue to store these multi-scale prototypes during training. Upon the multi-scale prototype memory queue, a novel multi-scale prototype contrastive loss is proposed to increase object feature discriminability across multiple scale, which brings better consistency of intermediate feature and boosts the convergence of network. Extensive experimental results on three publicly available datasets demonstrate the effectiveness and efficiency of our MPCNet over other state-of-the-art methods. The code is available at https://github.com/qixiong-wang/mmsegmentation-mpcnet.
Qixiong Wang, Xiaoyan Luo, Jiaqi Feng 0001, Guangyun Zhang, Xiuping Jia, Jihao Yin
IEEE Trans. Geosci. Remote. Sens.4
2023 Mesh-Based DGCNN: Semantic Segmentation of Textured 3-D Urban Scenes
abstract
Textured 3D mesh is one of the final user products in photogrammetry and remote sensing. However, research on the semantic segmentation of complex urban scenes represented by textured 3D meshes is in its infancy. We present a mesh-based dynamic graph CNN (DGCNN) for the semantic segmentation of textured 3D meshes. To represent each mesh facet, composite input feature vectors are constructed by concatenating the face-inherent features, i.e., XYZ coordinates of the center of gravity (CoG), texture values, and normal vectors. A texture fusion module is embedded into the proposed mesh-based DGCNN to generate high-level semantic features of the high-resolution texture information, which is useful for semantic segmentation. We achieve competitive accuracies when the proposed method is applied to the SUM mesh datasets. The overall accuracy (OA), Kappa coefficient (Kap), mean precision (mP), mean recall (mR), mean F1 score (mF1), and mean intersection over union (mIoU) are 93.3%, 88.7%, 79.6%, 83.0%, 80.7%, and 69.6%, respectively. In particular, the OA, mean class accuracy (mAcc), mIoU, and mF1 increase by 0.3%, 12.4%, 3.4%, and 6.9%, respectively, compared to the state-of-the-art method.
Guangyun Zhang, Jihao Yin, Xiuping Jia, Ajmal Mian
IEEE Trans. Geosci. Remote. Sens.2
2022 On the Convolutions of Sea Wave Spectrum in Radar Backscattering From Ocean Surfaces
abstract
By expanding the first-order small slope approximation scattering model in the polynomial terms containing convolutions of the sea wave spectrum, this letter quantitatively investigated how the convolutions of the sea wave spectrum affect the normalized ocean backscattering cross-section (NBCS) under various wind speeds, radar frequencies, and incident angles. First, numerical results show that only a portion of convolution powers of the sea wave spectrum contribute significantly to the NBCS, and such a portion of convolution powers are defined as the “effective convolutions”. Based on this, an empirical model related to the radar frequency, incident angle, and wind speed is established to determine the effective convolutions to reduce the computation cost. Finally, the NBCSs using the effective convolutions by the empirical model and the reference data calculated by the direct numerical integration are compared. The comparison results indicate that compared to the direct numerical integration, implementing the first-order small slope approximation with the effective convolutions has high accuracy and relatively low computational cost, notably with small Rayleigh parameters.
Dengfeng Xie, Rui Jiang 0002, Zhen Xu 0001, Guangyun Zhang
IEEE Geosci. Remote. Sens. Lett.4
2022 Performance analysis of inverting optical properties based on quasi-analytical algorithms
Jie Zhan, Dianjun Zhang, Lifeng Tan, Guangyun Zhang, Robert Zupan
Multim. Tools Appl.4
2022 A Deformable Attention Network for High-Resolution Remote Sensing Images Semantic Segmentation
abstract
Deformable convolutional networks (DCNs) can mitigate the inherent limited geometric transformation. We reformulate the spatialwise attention mechanism using DCNs in this article for semantic segmentation of high-resolution remote sensing (HRRS) images. It combines the sparse spatial sampling strategy and the long-range relationship modeling capability, namely, deformable attention module (DAM). Such locality awareness, more adaptable to HRRS image structures, can capture each pixel’s neighboring structural information. A reasonable multiscale deformable attention net (MDANet) is designed for the HRRS image semantic segmentation with a slightly increased computational cost based on the proposed DAM. Specifically, standard convolutional layers in the raw ResNet50 are equipped with a DAM to control sampling over a broader range of feature levels and aggregate multiscale context information. The experimental results evaluated on Vaihingen and DeepGlobe Land Cover Classification datasets show that the performance accuracy of MDANet is improved by 7.77% and 8.45% compared with the backbone network (ResNet50) in terms of Miou evaluation, respectively. Furthermore, a DAM can perform better than a global spatial attention mechanism with less computation on the$3 \times 64 \times 64$feature map. In addition, the added ablation studies demonstrate the effectiveness and efficiency of the DAM and multiscale strategy, respectively. Moreover, the sensitivity of critical hyperparameters is analyzed.
Renxiang Zuo, Guangyun Zhang, Xiuping Jia
IEEE Trans. Geosci. Remote. Sens.2
2016 Dem-based shadow detection and removal for lunar craters
abstract
In this paper, we focus on the shadow problem of the lunar surface, which hinders the implementation of visual tasks and visual processing for moon exploration projects. A random walker model is firstly applied to detect the shadowed pixels in lunar craters. Then, the detected shadows are removed by rectifying the illumination coefficients and detail coefficients, which are obtained by using the multi-scale decomposition technique. The proposed algorithm is evaluated on three groups of CCD images and the corresponding DEM data. Satisfactory results, which achieve comparative illumination yet preserve enough details, demonstrate the effectiveness of the proposed algorithm.
Hongmei Zhu, Jihao Yin, Ding Yuan 0001, Guangyun Zhang
IGARSS5
2015 The texture extraction and mapping of buildings with occlusion detection
abstract
Texture extraction and mapping is a key step for 3D reconstruction. The major methods are depending on experiences, scene photograph, close-range photogrammetry and oblique photogrammetry. They all require collecting additional images, which it is costly, time consuming and laborious. Meantime, the texture information acquired in the original aerial photos with high resolution is lost. Then, it is more economical and closer to reality if we can recover texture information from original aerial photographs. Thus, this paper presents one texture extraction and mapping method, with consideration of occlusion detection. There are two major characteristics: 1) Based on visibility judgment, we add the occlusion detection, which avoid the false texture extraction when the current building texture is occluded by other buildings. 2) We realize the real texture extraction and mapping from original aerial images. The location of the texture in original images is calculated with the collinearity equation and rectified texture coordinates.
Qingli Luo, Guoqing Zhou 0001, Guangyun Zhang, Jingjin Huang
IGARSS3
2015 Superpixel-Based Graphical Model for Remote Sensing Image Mapping
abstract
Object-oriented remote sensing image classification is becoming more and more popular because it can integrate spatial information from neighboring regions of different shapes and sizes into the classification procedure to improve the mapping accuracy. However, object identification itself is difficult and challenging. Superpixels, which are groups of spatially connected similar pixels, have the scale between the pixel level and the object level and can be generated from oversegmentation. In this paper, we establish a new classification framework using a superpixel-based graphical model. Superpixels instead of pixels are applied as the basic unit to the graphical model to capture the contextual information and the spatial dependence between the superpixels. The advantage of this treatment is that it makes the classification less sensitive to noise and segmentation scale. The contribution of this paper is the application of a graphical model to remote sensing image semantic segmentation. It is threefold. 1) Gradient fusion is applied to multispectral images before the watershed segmentation algorithm is used for superpixel generation. 2) A probabilistic fusion method is designed to derive node potential in the superpixel-based graphical model to address the problem of insufficient training samples at the superpixel level. 3) A boundary penalty between the superpixels is introduced in the edge potential evaluation. Experiments on three real data sets were conducted. The results show that the proposed method performs better than the related state-of-the-art methods tested.
Guangyun Zhang, Xiuping Jia, Jiankun Hu
IEEE Trans. Geosci. Remote. Sens.1
2014 Global attractive sets of a novel bounded chaotic system
Fuchen Zhang, Guangyun Zhang
Neural Comput. Appl.2
2012 Super pixel based remote sensing image classification with histogram descriptors on spectral and spatial data
abstract
Categorization based on objects is an effective way to integrate spectral and spatial information into remote sensing image classification. In this paper, we establish a classification framework which represents objects by super pixels. The non-parametric k-NN approach is chosen for this super pixel based method, as it is simple and free of class data distribution. A new descriptor for the features distribution of each super pixel, called 4-D color histograms, is used for both spectral and texture information. This descriptor provides a better tolerance for value fluctuations inside the super pixel. Furthermore, the Ç2distance is used as the measure of the similarity between color histograms of the super pixels. Experiments are conducted to illustrate the application of the proposed method.
Guangyun Zhang, Xiuping Jia, Ngai Ming Kwok
IGARSS1
2012 Simplified Conditional Random Fields With Class Boundary Constraint for Spectral-Spatial Based Remote Sensing Image Classification
abstract
Conditional random fields (CRF) have been introduced to remote sensing image classification recently to integrate contextual information into remote sensing classification. It employs the spatial property on both pixel's spectral data and labels. However, this leads to a large number of model parameters to train. In this letter, the training efficiency is improved by modifying the conventional CRF model. At the same time, a class boundary constraint is imposed into this framework to avoid over correction. The advantages of the developed method are demonstrated in the experimental results using real remotely sensed data.
Guangyun Zhang, Xiuping Jia
IEEE Geosci. Remote. Sens. Lett.1
2011 Feature selection using Kernel based Local Fisher Discriminant Analysis for hyperspectral image classification
abstract
Feature extraction is an important research aspect for hyperspectral remote sensing image classification to reduce the complexity and improve the classification accuracy. In this paper, a new feature extraction method, Kernel based Local Fisher Discriminative Analysis (KLFDA), is applied to hyperspectral remote sensing processing. This method integrates the advantages of conventional supervised Fisher Discriminative Analysis and unsupervised Locality Preserving Projection methods. Several experiments using the real images have been conducted, which indicate a high efficiency of this algorithm for hyperspectral image classification.
Guangyun Zhang, Xiuping Jia
IGARSS1