Shuguo Jiang

dblp:294/7759 · DBLP profile ↗
← Back
13ranked-venue papers
1as first author
13since 2021 · last 2025
0009-0001-9850-1043ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 9 · 1 first-author · 9 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Exploring Scene Affinity for Semi-Supervised LiDAR Semantic Segmentation
abstract
This paper explores scene affinity (AIScene), namely intra-scene consistency and inter-scene correlation, for semi-supervised LiDAR semantic segmentation in driving scenes. Adopting teacher-student training, AIScene employs a teacher network to generate pseudo-labeled scenes from unlabeled data, which then supervise the student network’s learning. Unlike most methods that include all points in pseudo-labeled scenes for forward propagation but only pseudo-labeled points for backpropagation, AIScene removes points without pseudo-labels, ensuring consistency in both forward and backward propagation within the scene. This simple point erasure strategy effectively prevents unsupervised, semantically ambiguous points (excluded in backpropagation) from affecting the learning of pseudo-labeled points. Moreover, AIScene incorporates patch-based data augmentation, mixing multiple scenes at both scene and instance levels. Compared to existing augmentation techniques that typically perform scene-level mixing between two scenes, our method enhances the semantic diversity of labeled (or pseudo-labeled) scenes, thereby improving the semi-supervised performance of segmentation models. Experiments show that AIScene outperforms previous methods on two popular benchmarks across four settings, achieving notable improvements of 1.9% and 2.1% in the most challenging 1% labeled data. The code will be released at https://github.com/azhuantou/AIScene.
Chuandong Liu, Xingxing Weng, Shuguo Jiang, Pengcheng Li 0017, Lei Yu 0006, Gui-Song Xia
CVPR3
2025 Fuzzy Boundary-Aware Network for Hyperspectral Individual Tree Fine Recognition
abstract
Different tree species have different carbon storage and growth rates. Therefore, accurate segmentation and identification of individual trees can provide more detailed carbon storage data, which is the basis for accurately estimating forest carbon storage. However, individual tree segmentation and recognition in dense forest areas face challenges such as crown overlap, complex terrain, and species diversity. To address these challenges and improve recognition accuracy, this paper proposes a fuzzy boundary-aware network (FBAN) for hyperspectral individual tree segmentation and recognition in dense forests. The proposed FBAN inclues a boundary-aware module (BAM) that explores channel boundaries between trees and non-trees, spatial boundaries of trees, and spectral boundaries between different trees by intergrating channel attention, spaital attention, and spectral attention. This enhances the separability of individual trees, especially those of the same species that are contiguous in dense forest areas. Additionally, an adaptive crown-aware module (ACAM) is constructed to adapt diverse-size crown features by coupling Transformer layers with dialted convolution layers. Experimental results on different hyperspectral datasets show that the proposed FBAN network outperforms existing methods in dense forest areas and different tree canopy areas, e.g., the AP on the SZU-South dataset is 4.5 points higher than that of Mask2former. It not only improves the accuracy of individual tree segmentation and recognition but also exhibits high generalization and robustness.
Nanying Li, Shuguo Jiang, Wangquan He, Sen Jia 0001
IEEE Trans. Geosci. Remote. Sens.2
2025 SAGT: Structure-Adaptive Graph Transformer for Hyperspectral Image Classification
abstract
Hyperspectral images (HSIs) are vital for scene analysis, as they capture detailed spatial and spectral information to characterize surface materials. However, accurate HSI classification is challenged by significant intra-class spectral variability and spatial complexity. To address this, we leverage the fact that pixels of the same class typically form irregular local regions. We propose a structure-adaptive graph transformer (SAGT) that dynamically captures irregular spatial topologies and homogeneous spectral information to achieve adaptive HSI representation and precise classification. Specifically, a structure-aware self-attention (SASA) module is developed to embed graph structures into the self-attention mechanism as a robust positional indicator, which can be extended easily and effectively. SASA comprehensively accounts for the spatial structures and spectral autocorrelation of ground objects, facilitating the aggregation of homogeneous spectral information for noise-robust spectral representations. Additionally, a structure-adaptive pooling (SAP) module is designed to dynamically adjust graph structures by discarding irrelevant edges, thus better indicating spatial relationships. By coupling the SASA and SAP modules, our proposed SAGT model significantly alleviates spectral variability and tolerates prior noise. Furthermore, data augmentation techniques of random discard and random offset are built, which randomly drop and shift graph nodes to generate more diverse samples during preprocessing. In postprocessing, multiview decision-making integrates results from multiple contextual views to provide more robust predictions. Experimental results on three benchmark datasets consistently demonstrate that SAGT is more effective and reliable than other state-of-the-art methods. To facilitate reproduction, we will release the source code for SAGT at https://github.com/ShuGuoJ/SAGT.git.
Shuyu Zhang 0002, Shuguo Jiang, Wenlong Yin, Weixi Wang, Meng Xu 0002, Jiasong Zhu, Sen Jia 0001
IEEE Trans. Geosci. Remote. Sens.2
2024 DMTG: One-Shot Differentiable Multi-Task Grouping
abstract
We aim to address Multi-Task Learning (MTL) with a large number of tasks by Multi-Task Grouping (MTG). Given $N$ tasks, we propose to simultaneously identify the best task groups from $2^N$ candidates and train the model weights simultaneously in one-shot, with the high-order task-affinity fully exploited. This is distinct from the pioneering methods which sequentially identify the groups and train the model weights, where the group identification often relies on heuristics. As a result, our method not only improves the training efficiency, but also mitigates the objective bias introduced by the sequential procedures that potentially leads to a suboptimal solution. Specifically, we formulate MTG as a fully differentiable pruning problem on an adaptive network architecture determined by an unknown Categorical distribution. To categorize $N$ tasks into $K$ groups (represented by $K$ encoder branches), we initially set up $KN$ task heads, where each branch connects to all $N$ task heads to exploit the high-order task-affinity. Then, we gradually prune the $KN$ heads down to $N$ by learning a relaxed differentiable Categorical distribution, ensuring that each task is exclusively and uniquely categorized into only one branch. Extensive experiments on CelebA and Taskonomy datasets with detailed ablations show the promising performance and efficiency of our method. The codes are available at https://github.com/ethanygao/DMTG.
Yuan Gao 0015, Shuguo Jiang, Moran Li, Jin-Gang Yu, Gui-Song Xia
ICML2
2024 A Center-Masked Transformer for Hyperspectral Image Classification
abstract
Convolutional neural networks (CNNs) are widely used in hyperspectral image (HSI) classification. However, the fixed receptive field of CNN-based methods limits their capability to extract global features. In recent years, transformer has been introduced into networks to tackle this limitation, but it brings other challenges, including a significant increase in model size, the number of labeled training samples required, and the limited effectiveness of sample encoding-reconstruction pretraining methods for HSI classification. To address these issues, a center-masked transformer (CMT) approach is proposed to improve the HSI classification accuracy from two perspectives. On one hand, a local-to-global token embedding (L2GTE) framework coupled with a multiscale convolutional token embedding (MCTE) module is used, which is well-designed to obtain local and global embedding tokens. This effectively reduces the number of model parameters. On the other hand, a regularized center-masked pretraining (RCPT) task is proposed and first introduced into the transformer-based network, which enables the network to learn the dependencies between central ground objects and neighboring objects without labels during the pretraining process. The experimental results conducted on five public HSI datasets demonstrate that our CMT approach outperforms other state-of-the-art methods for HSI classification when training samples are insufficient.
Sen Jia 0001, Shuguo Jiang, Ruyan He
IEEE Trans. Geosci. Remote. Sens.3
2024 SQformer: Spectral-Query Transformer for Hyperspectral Image Arbitrary-Scale Super-Resolution
abstract
Super-resolution is vital for the quality improvement of hyperspectral images (HSIs) under the spatial and spectral resolution trade-off. However, deep learning HSI super-resolution approaches typically adopt the “one model and one scale” scheme that is inefficient in training and storing. This is difficult in maximizing orbit equipment performance and aligning multiple spatial resolution data in remote sensing. Therefore, this article intends to address HSI arbitrary-scale super-resolution, enabling the scaling of HSIs to arbitrary sizes using a single model. To do this end, we treat HSI arbitrary-scale super-resolution as a retrieval problem. It conceptualizes the HSI as a dictionary of pixelwise tokens with spatial-spectral features, position information, and scale information. Its objective is to employ a set of initialized tokens related to the high-resolution (HR) HSI as queries to retrieve matched spectral features from low-resolution (LR) one, which is so-called token-based query-to-spectrum. Since these query tokens can be constructed flexibly (e.g., through random initialization), we can generate a desired number of them to reconstruct our HR HSI, thus achieving arbitrary-scale super-resolution. This process considers not only position information but also spectral features so that it can decrease spectral distortion. With the above idea, we developed an HSI arbitrary-scale super-resolution method, dubbed as spectral-query transformer (SQformer). Specifically, it begins by converting the LR HSI into a dictionary of LR tokens and then constructs a desired number of HR tokens. To enable flexible token construction, we design an implicit spectral token (particularly a learnable vector) and replicate it$\alpha H \times \alpha W$times to form the HR tokens. Next, the HR and LR tokens are passed into a transformer decoder to find the most matched spectral response for the former by soft-weighting the LR tokens. Finally, the HR tokens are spatially rearranged in order, forming an HR HSI. Extensive experiments have demonstrated its effectiveness on remote sensing data. The code will be released at:https://github.com/ShuGuoJ/SQformer.git.
Shuguo Jiang, Nanying Li, Meng Xu 0002, Shuyu Zhang 0002, Sen Jia 0001
IEEE Trans. Geosci. Remote. Sens.1
2024 Texture-Aware Self-Attention Model for Hyperspectral Tree Species Classification
abstract
Forests play an irreplaceable role in carbon sinks. However, there are obvious differences in the carbon sink capacity of different tree species, so the scientific and accurate identification of surface forest vegetation is the key to achieving the double carbon goal. Due to the disordered distribution of trees, varied crown geometry, and high difficulty in labeling tree species, traditional methods have a poor ability to represent complex spatial–spectral structures. Therefore, how to quickly and accurately obtain key and subtle features of tree species to finely identify tree species is an urgent problem to be solved in current research. To address these issues, a texture-aware self-attention model (TASAM) is proposed to improve spatial contrast and overcome spectral variance, achieving accurate classification of tree species hyperspectral images (HSIs). In our model, a nested spatial pyramid module is first constructed to accurately extract the multiview and multiscale features that highlight the distinction between tree species and surrounding backgrounds. In addition, a cross-spectral–spatial attention module is designed, which can capture spatial–spectral joint features over the entire image domain. The Gabor feature is introduced as an auxiliary function to guide self-attention to autonomously focus on latent space texture features, further extract more appropriate and accurate information, and enhance the distinction between the target and the background. Verification experiments on three tree species hyperspectral datasets prove that the proposed method can obtain finer and more accurate tree species classification under the condition of limited labeled samples. This method can effectively solve the problem of tree species classification in complex forest structures and can meet the application requirements of tree species diversity monitoring, forestry resource investigation, and forestry carbon sink analysis based on HSIs.
Nanying Li, Shuguo Jiang, Songxin Ye, Sen Jia 0001
IEEE Trans. Geosci. Remote. Sens.2
2024 Graph-in-Graph Convolutional Network for Hyperspectral Image Classification
abstract
With the development of hyperspectral sensors, accessible hyperspectral images (HSIs) are increasing, and pixel-oriented classification has attracted much attention. Recently, graph convolutional networks (GCNs) have been proposed to process graph-structured data in non-Euclidean domains and have been employed in HSI classification. But most methods based on GCN are hard to sufficiently exploit information of ground objects due to feature aggregation. To solve this issue, in this article, we proposed a graph-in-graph (GiG) model and a related GiG convolutional network (GiGCN) for HSI classification from a superpixel viewpoint. The GiG representation covers information inside and outside superpixels, respectively, corresponding to the local and global characteristics of ground objects. Concretely, after segmenting HSI into disjoint superpixels, each one is converted to an internal graph. Meanwhile, an external graph is constructed according to the spatial adjacent relationships among superpixels. Significantly, each node in the external graph embeds a corresponding internal graph, forming the so-called GiG structure. Then, GiGCN composed of internal and External graph convolution (EGC) is designed to extract hierarchical features and integrate them into multiple scales, improving the discriminability of GiGCN. Ensemble learning is incorporated to further boost the robustness of GiGCN. It is worth noting that we are the first to propose the GiG framework from the superpixel point and the GiGCN scheme for HSI classification. Experiment results on four benchmark datasets demonstrate that our proposed method is effective and feasible for HSI classification with limited labeled samples. For study replication, the code developed for this study is available at https://github.com/ShuGuoJ/GiGCN.git.
Sen Jia 0001, Shuguo Jiang, Shuyu Zhang 0002, Meng Xu 0002, Xiuping Jia
IEEE Trans. Neural Networks Learn. Syst.2
2023 Structure-Adaptive Convolutional Neural Network for Hyperspectral Image Classification
abstract
Hyperspectral image (HSI) classification based on deep learning is a hot research topic. The convolutional model employs a single rectangular window to interpret the sample neighborhood features, whereas effective characterization of the complex spatial structure of HSI is still an unsolved problem. In this article, we propose a structure-adaptive convolutional neural network (SACNN) for HSI classification, which efficiently exploits the intrinsic spatial geometry information. Four novel strategies are designed to construct the proposed SACNN network. First, superpixel homogeneous region (SHR) sample generation is introduced to achieve neighborhood features within the intercepted rectangular window of the superpixel. Second, online batch-wise standardization uses zero padding to unify the size of inputs in the same batch, thereby realizing parallel processing of irregular inputs. Third, structure-adaptive convolution (SConv) and structure-adaptive average pooling (SAP) are correspondingly constructed to extract deep spectral, spatial, and geometric features from the effective mapping area of superpixels, and further aggregate the information within irregular boundaries. Finally, a sample-adaptive loss weight (SLW) scheme is designed to adjust the influence of different labels on the same input. Experimental results show that the overall classification accuracy of SACNN reaches 93.11%, 90.96%, and 85.04% for 15 randomly selected training samples per class on three HSI datasets, respectively, obtaining an improvement of 0.97%–2.97% with respect to the best-compared method.
Sen Jia 0001, Dongsheng Bi, Jianhui Liao, Shuguo Jiang, Meng Xu 0002, Shuyu Zhang 0002
IEEE Trans. Geosci. Remote. Sens.4
2023 Collaborative Contrastive Learning for Hyperspectral and LiDAR Classification
abstract
Using single-source remote sensing (RS) data for classification of ground objects has certain limitations, however, multi-modal RS data contain different types of features, such as spectral features and spatial features of hyperspectral image (HSI) and elevation information of light detection and ranging (LiDAR) data, which can be used to extract and fuse high-quality features to improve the classification accuracy. Nevertheless, the existing fusion techniques are mostly limited by the number of labeled samples due to the difficulty of label collection in the multi-modal RS data. In this article, a fusion method of collaborative contrastive learning (CCL) is proposed to tackle the abovementioned issues for HSI and LiDAR data classification. The proposed CCL approach includes two stages of pre-training (CCL-PT) and fine-tuning (CCL-FT). In the CCL-PT stage, a collaborative strategy is introduced into contrastive learning (CL), which can extract features from HSI and LiDAR data separately, and achieve the coordinated feature representation and matching between the two-modal RS data without labeled samples. In the CCL-FT stage, a multi-level fusion network is designed to optimize and fuse the unsupervised collaborative features which are extracted in the CCL-PT stage for the classification tasks. Experimental results on three real-world data sets show that the developed CCL approach can perform excellently on the small sample classification tasks and CL is feasible for the fusion of multi-modal RS data.
Sen Jia 0001, Shuguo Jiang, Ruyan He
IEEE Trans. Geosci. Remote. Sens.3
2023 Dual Self-Attention Swin Transformer for Hyperspectral Image Super-Resolution
abstract
Spatial resolution is a crucial indicator for measuring the quality of hyperspectral imaging (HSI) and obtaining high-resolution (HR) hyperspectral images without any auxiliary information has become increasingly challenging. One promising approach is to use deep-learning (DL) techniques to reconstruct HR hyperspectral images from low-resolution (LR) images, namely super-resolution (SR). While convolutional neural networks are commonly used for hyperspectral image SR (HSI-SR), they often lead to unavoidable performance degradation due to the lack of long-range dependence learning ability. In this article, we propose a dual self-attention Swin transformer SR (DSSTSR) network that utilizes the ability of the shifted windows (Swin) transformer in the spatial representation of both global and local features and learns spectral sequence information from adjacent bands of HSI. Additionally, DSSTSR incorporates an image denoising module using the wavelet transformation method to mitigate the impact of stripe noise on HSI-SR. Our extensive experiments using publicly close-range datasets demonstrate that DSSTSR outperforms other state-of-art HSI-SR methods in terms of three image quality metrics. Furthermore, we applied DSSTSR to the SR of satellite hyperspectral images and achieved improved classification results. Compared to its competitors, DSSTSR exhibits superior performance in enhancing spatial resolution while preserving spectral information. These results suggest that the DSSTSR network has great potential for standardization in remote-sensing image processing and practical applications.
Yaqian Long, Meng Xu 0002, Shuyu Zhang 0002, Shuguo Jiang, Sen Jia 0001
IEEE Trans. Geosci. Remote. Sens.5
2022 A Semisupervised Siamese Network for Hyperspectral Image Classification
abstract
With the development of hyperspectral imaging technology, hyperspectral images (HSIs) have become important when analyzing the class of ground objects. In recent years, benefiting from the massive labeled data, deep learning has achieved a series of breakthroughs in many fields of research. However, labeling HSIs requires sufficient domain knowledge and is time-consuming and laborious. Thus, how to apply deep learning effectively to small labeled samples is an important topic of research in HSI classification. To solve this problem, we propose a semisupervised Siamese network that embeds Siamese network into a semisupervised learning scheme. It integrates an autoencoder module and a Siamese network to, respectively, investigate information in a large amount of unlabeled data and rectify it with a limited labeled sample set, which is called 3DAES. First, the autoencoder method is trained on the massive unlabeled data to learn the refinement representation, creating an unsupervised feature. Second, based on this unsupervised feature, limited labeled samples are used to train a Siamese network to rectify the unsupervised feature to improve feature separability among various classes. Furthermore, by training the Siamese network, a random sampling scheme is used to accelerate training and avoid imbalance among various sample classes. Experiments on three benchmark HSI datasets consistently demonstrate the effectiveness and robustness of the proposed 3DAES approach with limited labeled samples. For study replication, the code developed for this study is available athttps://github.com/ShuGuoJ/3DAES.git.
Sen Jia 0001, Shuguo Jiang, Meng Xu 0002, Weiwei Sun 0005, Jiasong Zhu, Xiuping Jia
IEEE Trans. Geosci. Remote. Sens.2
2021 A survey: Deep learning for hyperspectral image classification with few labeled samples
abstract
With the rapid development of deep learning technology and improvement in computing capability, deep learning has been widely used in the field of hyperspectral image (HSI) classification. In general, deep learning models often contain many trainable parameters and require a massive number of labeled samples to achieve optimal performance. However, in regard to HSI classification, a large number of labeled samples is generally difficult to acquire due to the difficulty and time-consuming nature of manual labeling. Therefore, many research works focus on building a deep learning model for HSI classification with few labeled samples. In this article, we concentrate on this topic and provide a systematic review of the relevant literature. Specifically, the contributions of this paper are twofold. First, the research progress of related methods is categorized according to the learning paradigm, including transfer learning, active learning and few-shot learning. Second, a number of experiments with various state-of-the-art approaches has been carried out, and the results are summarized to reveal the potential research directions. More importantly, it is notable that although there is a vast gap between deep learning models (that usually need sufficient labeled samples) and the HSI scenario with few labeled samples, the issues of small-sample sets can be well characterized by fusion of deep learning methods and related techniques, such as transfer learning and a lightweight model. For reproducibility, the source codes of the methods assessed in the paper can be found at https://github.com/ShuGuoJ/HSI-Classification.git.
Sen Jia 0001, Shuguo Jiang, Nanying Li, Meng Xu 0002, Shiqi Yu 0001
Neurocomputing2