VLDB 2026 Research / reviewers in the wild / expert
Shuyu Zhang 0002
dblp:86/2981-2
· DBLP profile ↗
11ranked-venue papers
7as first author
11since 2021 · last 2025
0000-0003-2038-0349ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 10 · 7 first-author · 10 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SAGT: Structure-Adaptive Graph Transformer for Hyperspectral Image ClassificationabstractHyperspectral images (HSIs) are vital for scene analysis, as they capture detailed spatial and spectral information to characterize surface materials. However, accurate HSI classification is challenged by significant intra-class spectral variability and spatial complexity. To address this, we leverage the fact that pixels of the same class typically form irregular local regions. We propose a structure-adaptive graph transformer (SAGT) that dynamically captures irregular spatial topologies and homogeneous spectral information to achieve adaptive HSI representation and precise classification. Specifically, a structure-aware self-attention (SASA) module is developed to embed graph structures into the self-attention mechanism as a robust positional indicator, which can be extended easily and effectively. SASA comprehensively accounts for the spatial structures and spectral autocorrelation of ground objects, facilitating the aggregation of homogeneous spectral information for noise-robust spectral representations. Additionally, a structure-adaptive pooling (SAP) module is designed to dynamically adjust graph structures by discarding irrelevant edges, thus better indicating spatial relationships. By coupling the SASA and SAP modules, our proposed SAGT model significantly alleviates spectral variability and tolerates prior noise. Furthermore, data augmentation techniques of random discard and random offset are built, which randomly drop and shift graph nodes to generate more diverse samples during preprocessing. In postprocessing, multiview decision-making integrates results from multiple contextual views to provide more robust predictions. Experimental results on three benchmark datasets consistently demonstrate that SAGT is more effective and reliable than other state-of-the-art methods. To facilitate reproduction, we will release the source code for SAGT at https://github.com/ShuGuoJ/SAGT.git. Shuyu Zhang 0002, Shuguo Jiang, Wenlong Yin, Weixi Wang, Meng Xu 0002, Jiasong Zhu, Sen Jia 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2025 | Global-Local Residual Fusion Network for Hyperspectral Image ClassificationabstractAs hyperspectral images (HSIs) continue to increase in data resolution and information richness, current deep learning models need to enhance their feature extraction and understanding capabilities for classification tasks. The complementarity between convolution and attention mechanisms in deep learning enables the capture of both local and global information. However, it faces the problems of intrinsic coupling of different operators and precise fusion of different features. In this study, a novel global-local residual fusion network (GLRFNet) is proposed to improve the HSI classification. Firstly, a feature projection with multiple kernels is designed to generate the feature pool before deep extraction and enhance the information connection between operators. Then, a global-local residual (GLR) feature extraction network is built to capture both fine-grained details and large-scale dependencies, improving the feature perception in various classification scenes. It consists of local convolution, global attention, and residual construction branches in a separate and coupled manner. Finally, an optimized inverted bottleneck fusion (IBF) module is built to perform nonlinear and comprehensive feature fusion for mining the high-level semantics and understanding the category relationships. The experiments on four HSI datasets demonstrate the superiority of GLRFNet compared to other state-of-the-art methods, especially its classification performance under small sample conditions, with higher recognition accuracy and better class boundaries. In addition, parameter analysis and ablation experiments are also conducted to determine the optimal parameters and verify the module effectiveness. Shuyu Zhang 0002, Wenlong Yin, Yawen Fu, Sen Jia 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | SQformer: Spectral-Query Transformer for Hyperspectral Image Arbitrary-Scale Super-ResolutionabstractSuper-resolution is vital for the quality improvement of hyperspectral images (HSIs) under the spatial and spectral resolution trade-off. However, deep learning HSI super-resolution approaches typically adopt the “one model and one scale” scheme that is inefficient in training and storing. This is difficult in maximizing orbit equipment performance and aligning multiple spatial resolution data in remote sensing. Therefore, this article intends to address HSI arbitrary-scale super-resolution, enabling the scaling of HSIs to arbitrary sizes using a single model. To do this end, we treat HSI arbitrary-scale super-resolution as a retrieval problem. It conceptualizes the HSI as a dictionary of pixelwise tokens with spatial-spectral features, position information, and scale information. Its objective is to employ a set of initialized tokens related to the high-resolution (HR) HSI as queries to retrieve matched spectral features from low-resolution (LR) one, which is so-called token-based query-to-spectrum. Since these query tokens can be constructed flexibly (e.g., through random initialization), we can generate a desired number of them to reconstruct our HR HSI, thus achieving arbitrary-scale super-resolution. This process considers not only position information but also spectral features so that it can decrease spectral distortion. With the above idea, we developed an HSI arbitrary-scale super-resolution method, dubbed as spectral-query transformer (SQformer). Specifically, it begins by converting the LR HSI into a dictionary of LR tokens and then constructs a desired number of HR tokens. To enable flexible token construction, we design an implicit spectral token (particularly a learnable vector) and replicate it$\alpha H \times \alpha W$times to form the HR tokens. Next, the HR and LR tokens are passed into a transformer decoder to find the most matched spectral response for the former by soft-weighting the LR tokens. Finally, the HR tokens are spatially rearranged in order, forming an HR HSI. Extensive experiments have demonstrated its effectiveness on remote sensing data. The code will be released at:https://github.com/ShuGuoJ/SQformer.git. Shuguo Jiang, Nanying Li, Meng Xu 0002, Shuyu Zhang 0002, Sen Jia 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Vertical Attention-Based Siamese ConvLSTM Network for Argo Data Error DetectionabstractThe international array for real-time geostrophic oceanography (Argo) project is committed to rapidly and precisely acquiring comprehensive 3-D data on ocean temperature and salinity, which is crucial for monitoring ocean climate change and natural phenomena. During the buoy observation, environmental factors, human mistakes, and equipment malfunctions can cause abnormalities such as density inversion and spike, and thus detecting the errors in Argo data is significant to ensure its reliability and applicability. Traditional methods mainly rely on the knowledge and judgment of marine experts, ensuring high accuracy but requiring large amounts of effort. Machine-learning methods are used for automatic Argo data error detection, while they still struggle with extracting deep and discriminative features from profiles. Recently, deep-learning methods have received increasing attention in this field, yet their effectiveness have not been widely explored, faced with challenges of imbalanced samples, joint detection, and complicated patterns. In this article, a novel vertical attention-based siamese ConvLSTM (VAS-CLSTM) network is proposed for the accurate error detection of Argo data. First, an oversampling approach with optimized deep clustering based on inheritance theory and Mahalanobis distance is designed to effectively augment the error samples. Second, a siamese convolutional long–short-term memory (ConvLSTM) network with contextual connection and spatial–temporal adjacent profile search is built to learn interactively from temperature and salinity profiles. Third, a depth-based vertical attention mechanism with grouped weights and vertical trends is proposed for adaptive modeling and flexible learning. Experimental results of North and South Atlantic datasets show that the proposed VAS-CLSTM method effectively improves the accuracy and reliability of error detection in Argo observation data. Shuyu Zhang 0002, Zhaoji Shi, Chuhong Wu, Yan Li 0066, Xiaomei Liao, Sen Jia 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | Spatial-Temporal Siamese Convolutional Neural Network for Subsurface Temperature ReconstructionabstractThe reconstruction of subsurface ocean temperature using sea surface observations and in situ Argo measurements is an important yet challenging task. The availability of long-term and high-resolution sea surface remote sensing, combined with advancements in deep learning technology, has opened new opportunities for studying subsurface temperature (ST) reconstruction. In this study, a novel spatial–temporal Siamese convolutional neural network (SSCNN) is proposed to improve the accuracy of ST reconstruction in the Indian Ocean. First, considering the distinctions of temperature characteristics among different sea areas, a multiscale division scheme based on the correlation coefficient of integral ST is designed for refined reconstruction modeling. Second, since ocean heat is significantly affected by solar radiation, asymmetric convolutional operation with rectangular patches and kernels is designed to capture the information characteristics in longitude and latitude directions, respectively. Third, given the temporal changes and correlations of ocean temperature, an SSCNN with shared parameters is proposed for multiview feature mining and accurate temperature structure reconstruction. The reconstructed results provide a precise depiction of the subsurface Indian Ocean dipole (sub-IOD)’s evolution, including the spatial distribution of positive and negative anomaly signals and its temporal changes. It demonstrates that the subsurface dipole index series obtained from SSCNN reconstruction is consistent with that from International Pacific Research Center (IPRC) observation, remaining within a reasonable error range. Comparative experiments indicate that the SSCNN model surpasses other existing methods in terms of higher accuracy and smaller error. Overall, this study provides a promising approach for effectively reconstructing the ST using deep learning methods and offers valuable insights for analyzing the evolution of subsurface positive dipole in Indian Ocean. Shuyu Zhang 0002, Yizhou Yang, Kangwen Xie, Jiahao Gao, Qianru Niu, Gongjie Wang, Zhihui Che, Sen Jia 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | Graph-in-Graph Convolutional Network for Hyperspectral Image ClassificationabstractWith the development of hyperspectral sensors, accessible hyperspectral images (HSIs) are increasing, and pixel-oriented classification has attracted much attention. Recently, graph convolutional networks (GCNs) have been proposed to process graph-structured data in non-Euclidean domains and have been employed in HSI classification. But most methods based on GCN are hard to sufficiently exploit information of ground objects due to feature aggregation. To solve this issue, in this article, we proposed a graph-in-graph (GiG) model and a related GiG convolutional network (GiGCN) for HSI classification from a superpixel viewpoint. The GiG representation covers information inside and outside superpixels, respectively, corresponding to the local and global characteristics of ground objects. Concretely, after segmenting HSI into disjoint superpixels, each one is converted to an internal graph. Meanwhile, an external graph is constructed according to the spatial adjacent relationships among superpixels. Significantly, each node in the external graph embeds a corresponding internal graph, forming the so-called GiG structure. Then, GiGCN composed of internal and External graph convolution (EGC) is designed to extract hierarchical features and integrate them into multiple scales, improving the discriminability of GiGCN. Ensemble learning is incorporated to further boost the robustness of GiGCN. It is worth noting that we are the first to propose the GiG framework from the superpixel point and the GiGCN scheme for HSI classification. Experiment results on four benchmark datasets demonstrate that our proposed method is effective and feasible for HSI classification with limited labeled samples. For study replication, the code developed for this study is available at https://github.com/ShuGuoJ/GiGCN.git. Sen Jia 0001, Shuguo Jiang, Shuyu Zhang 0002, Meng Xu 0002, Xiuping Jia |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | Structure-Adaptive Convolutional Neural Network for Hyperspectral Image ClassificationabstractHyperspectral image (HSI) classification based on deep learning is a hot research topic. The convolutional model employs a single rectangular window to interpret the sample neighborhood features, whereas effective characterization of the complex spatial structure of HSI is still an unsolved problem. In this article, we propose a structure-adaptive convolutional neural network (SACNN) for HSI classification, which efficiently exploits the intrinsic spatial geometry information. Four novel strategies are designed to construct the proposed SACNN network. First, superpixel homogeneous region (SHR) sample generation is introduced to achieve neighborhood features within the intercepted rectangular window of the superpixel. Second, online batch-wise standardization uses zero padding to unify the size of inputs in the same batch, thereby realizing parallel processing of irregular inputs. Third, structure-adaptive convolution (SConv) and structure-adaptive average pooling (SAP) are correspondingly constructed to extract deep spectral, spatial, and geometric features from the effective mapping area of superpixels, and further aggregate the information within irregular boundaries. Finally, a sample-adaptive loss weight (SLW) scheme is designed to adjust the influence of different labels on the same input. Experimental results show that the overall classification accuracy of SACNN reaches 93.11%, 90.96%, and 85.04% for 15 randomly selected training samples per class on three HSI datasets, respectively, obtaining an improvement of 0.97%–2.97% with respect to the best-compared method. Sen Jia 0001, Dongsheng Bi, Jianhui Liao, Shuguo Jiang, Meng Xu 0002, Shuyu Zhang 0002 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2023 | Dual Self-Attention Swin Transformer for Hyperspectral Image Super-ResolutionabstractSpatial resolution is a crucial indicator for measuring the quality of hyperspectral imaging (HSI) and obtaining high-resolution (HR) hyperspectral images without any auxiliary information has become increasingly challenging. One promising approach is to use deep-learning (DL) techniques to reconstruct HR hyperspectral images from low-resolution (LR) images, namely super-resolution (SR). While convolutional neural networks are commonly used for hyperspectral image SR (HSI-SR), they often lead to unavoidable performance degradation due to the lack of long-range dependence learning ability. In this article, we propose a dual self-attention Swin transformer SR (DSSTSR) network that utilizes the ability of the shifted windows (Swin) transformer in the spatial representation of both global and local features and learns spectral sequence information from adjacent bands of HSI. Additionally, DSSTSR incorporates an image denoising module using the wavelet transformation method to mitigate the impact of stripe noise on HSI-SR. Our extensive experiments using publicly close-range datasets demonstrate that DSSTSR outperforms other state-of-art HSI-SR methods in terms of three image quality metrics. Furthermore, we applied DSSTSR to the SR of satellite hyperspectral images and achieved improved classification results. Compared to its competitors, DSSTSR exhibits superior performance in enhancing spatial resolution while preserving spectral information. These results suggest that the DSSTSR network has great potential for standardization in remote-sensing image processing and practical applications. Yaqian Long, Meng Xu 0002, Shuyu Zhang 0002, Shuguo Jiang, Sen Jia 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Multivariate Temporal Self-Attention Network for Subsurface Thermohaline Structure ReconstructionabstractArgo observations are spatially sparse and temporally uneven, whereas satellites can provide high-resolution and continuous observations at the sea surface. The reconstruction of subsurface thermohaline structure using multi-source remote sensing data is thus of great significance for investigating the ocean interior dynamics. Aiming at the existing problems of temporal feature extraction and nonlinear relationship fitting, this paper proposes a multivariate temporal self-attention network (MTSAN) to effectively reconstruct the subsurface temperature anomaly (STA) and subsurface salinity anomaly (SSA) in the Pacific Ocean. The model integrates multi-source remote sensing data, including sea surface temperature and salinity, wind speed, absolute dynamic topography, and significant wave height. In order to better extract the complex small- and medium-scale signals, a two-branch asymmetric residual module based on dilation causal convolution is designed to enhance the representation ability. Moreover, zonal weighted loss function with comprehensive indicators is proposed, in order to minimize the real error of grids and raise the accuracy of self-attention network. MTSAN reconstructs the STA and SSA during the El Niño event, and the results show that it has good performance for spatial distribution, vertical variation, and temporal extension. The overallR2and RMSE of STA are 0.536 and 0.241° C, respectively, and the overallR2and RMSE of SSA are 0.645 and 0.037psu, respectively. In addition, the results of comparison experiments illustrate the superiority of MTSAN over other machine learning and deep learning based methods. Overall, we provide a new temporal self-attention approach to accurately reconstruct the three-dimensional thermohaline structure using high-resolution quasi-real-time satellite observations. Shuyu Zhang 0002, Yuesen Deng, Qianru Niu, Zhihui Che, Sen Jia 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Superpixel-Guided Variable Gabor Phase Coding Fusion for Hyperspectral Image Classificationabstract3-D Gabor, as a typical filter, plays a critical role in extracting discriminative spectral–spatial features from hyperspectral images (HSIs). However, the performance of traditional 3-D Gabor is limited by the uniform response to each direction, which is inconsistent with the complexity of land cover distribution. It has been a continuing concern for researchers to investigate the anisotropic 3-D Gabor filters. In addition, the 3-D Gabor wavelets do not make full use of spatial distribution information, thus reducing the accuracy. This article proposes a superpixel-guided variable 3-D Gabor phase coding fusion (SuVGF) framework for HSI classification with limited training samples. First, the variable 3-D Gabor filters are created based on various asymmetric sinusoidal waves and spatial kernel sizes to achieve multidirectional features. Second, the local Gabor phase ternary pattern is adopted to encode the Gabor phases and improve the feature discrimination. Meanwhile, a scale map is produced by the majority voting of multiscale simple noniterative clustering (SNIC) and entropy rate superpixel (ERS) segmentation, which contains sufficient and complementary spatial distribution information. Then, geometric optimization is employed on the scale map to reduce noise disturbances. Finally, all Gabor features are modified by the filter with the guidance of a scale map and fused together as a confidence cube, and the random forest algorithm is exploited for classification. TheSuVGF is applied to three real hyperspectral datasets to demonstrate the superiority of higher accuracy, stronger robustness, and less computational complexity in comparison with several state-of-the-art ones. Shuyu Zhang 0002, Dingding Tang, Nanying Li, Xiuping Jia, Sen Jia 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Unsupervised Spatial-Spectral CNN-Based Feature Learning for Hyperspectral Image ClassificationabstractThe rapid development of remote sensing sensors makes the acquisition, analysis, and application of hyperspectral images (HSIs) more and more extensive. However, the limited sample sets, high-dimensional features, highly correlated bands, and mixing spectral information make the classification of HSIs a great challenge. In this article, an unsupervised multiscale and diverse feature learning (UMsDFL) method is proposed for HSI classification, which deeply considers the spatial–spectral features via convolutional neural networks (CNNs). Specifically, after employing the simple noniterative clustering (SNIC) algorithm with the heuristic calculation of superpixel size, the HSIs are segmented into superpixels for feature learning. The unsupervised network is designed with the convolutional encoder and decoder, the additional clustering branch, and the multilayer feature fusion to enhance the distinguishability of feature learning and the reusability of feature maps. Then, the spatial relationships and object attributes in large- and small-scale contexts are learned collaboratively through the unsupervised network to utilize the complementary multiscale characteristics. Moreover, the diverse features of hyperspectral information and nonsubsampled contourlet transform (NSCT) textures are learned simultaneously via the unsupervised network to alleviate the insufficiency of geometric representation. Finally, the random forest (RF) is adopted as the comprehensive classifier for land cover mapping based on the UMsDFL, and superpixel regularization is adopted to optimize the classification results. A series of experiments are performed on three real-world HSI datasets to demonstrate the effectiveness of our UMsDFL approach. The experimental results show that the proposed UMsDFL can achieve the overall accuracy of 79.23%, 96.49%, and 77.26% for Houston, Pavia, and Dioni datasets, respectively, when there are only five samples per class for training. Shuyu Zhang 0002, Meng Xu 0002, Jun Zhou 0001, Sen Jia 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |