EDBT 2026 Demo / reviewers in the wild / expert
Jiaqi Yang 0005
dblp:131/7234-5
· DBLP profile ↗
12ranked-venue papers
8as first author
12since 2021 · last 2025
0000-0001-8322-9270ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 8 · 5 first-author · 8 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Spectral Structure-Aware Initialization and Probability-Consistent Self-Training for Cross-Scene Hyperspectral Image ClassificationabstractCross-scene classification of hyperspectral images (HSI) aims to classify target domain (TD) data using only labeled source domain (SD) data and unlabeled TD data during training. However, challenges such as spectral shifts across scenes and semantic discrepancies between domains significantly degrade classification performance. To address these issues, domain adaptation (DA) has gained increasing attention in the hyperspectral remote sensing community. This paper proposes a novel framework for cross-scene HSI classification, termed Data Structure-Aware Initialization and Probability-Consistent Self-Training (S2PST) framework. The framework employs batch nuclear-norm maximization to constrain the probability responses of TD outputs, implicitly aligning feature distributions between SD and TD. To enhance the model’s robustness and spectral feature representation ability, we introduce a spectral structure-aware initialization method that integrates the strengths of traditional machine learning and deep learning. Furthermore, to mitigate the model’s bias toward SD training data, we propose a self-supervised training strategy that dynamically incorporates pseudo-labeled TD samples into the training process by comparing the similarity of high-confidence samples in the probability space between SD and TD. Extensive experiments are conducted on the Houston, HyRANK, and Pavia datasets, and compared with several state-of-the-art DA methods. The experiment results demonstrate the effectiveness of the proposed framework. Our code will be available at https://github.com/liurongwhm. Junye Liang, Jiaqi Yang 0005, Quanwei Liu, Peng Zhu 0001 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2025 | HyperSIGMA: Hyperspectral Intelligence Comprehension Foundation ModelabstractAccurate hyperspectral image (HSI) interpretation is critical for providing valuable insights into various earth observation-related applications such as urban planning, precision agriculture, and environmental monitoring. However, existing HSI processing methods are predominantly task-specific and scene-dependent, which severely limits their ability to transfer knowledge across tasks and scenes, thereby reducing the practicality in real-world applications. To address these challenges, we present HyperSIGMA, a vision transformer-based foundation model that unifies HSI interpretation across tasks and scenes, scalable to over one billion parameters. To overcome the spectral and spatial redundancy inherent in HSIs, we introduce a novel sparse sampling attention (SSA) mechanism, which effectively promotes the learning of diverse contextual features and serves as the basic block of HyperSIGMA. HyperSIGMA integrates spatial and spectral features using a specially designed spectral enhancement module. In addition, we construct a large-scale hyperspectral dataset, HyperGlobal-450K, for pre-training, which contains about 450 K hyperspectral images, significantly surpassing existing datasets in scale. Extensive experiments on various high-level and low-level HSI tasks demonstrate HyperSIGMA's versatility and superior representational capability compared to current state-of-the-art methods. Moreover, HyperSIGMA shows significant advantages in scalability, robustness, cross-modal transferring capability, real-world applicability, and computational efficiency. Di Wang 0023, Meiqi Hu, Yuchun Miao, Jiaqi Yang 0005, Yichu Xu, Xiaolei Qin, Jiaqi Ma 0002, Chenxing Li, Chuan Fu, Hongruixuan Chen, Chengxi Han, Naoto Yokoya, Jing Zhang 0037, Minqiang Xu, Lefei Zhang, Chen Wu 0003, Bo Du 0001, Dacheng Tao, Liangpei Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | DHSNet: Dual Classification Head Self-Training Network for Cross-Scene Hyperspectral Image ClassificationabstractDue to the difficulty of obtaining labeled data for hyperspectral images (HSIs), cross-scene classification has emerged as a widely adopted approach in the remote sensing community. It involves training a model using labeled data from a source domain (SD) and unlabeled data from a target domain (TD), followed by inference on the TD. However, variations in the reflectance spectrum of the same object between the SD and the TD, as well as differences in the feature distribution of the same land cover class, pose significant challenges to the performance of cross-scene classification. To address this issue, we propose a dual classification head self-training network (DHSNet). This method aligns class-wise features across domains, ensuring that the trained classifier can accurately classify TD data of different classes. We introduce a dual classification head self-training strategy for the first time in the cross-scene HSI classification field and design a self-training loss based on the prediction of the two classification heads. The proposed approach mitigates the domain gap while preventing the accumulation of incorrect pseudo-labels in the model. Additionally, we incorporate a novel central feature attention mechanism to enhance the model’s capacity to learn scene-invariant features across domains. DHSNet significantly outperforms state-of-the-art methods on three cross-scene HSI datasets, achieving 80.23±1.92% OA on the Houston dataset. The code for DHSNet will be available at https://github.com/liurongwhm. Junye Liang, Jiaqi Yang 0005, Meiqi Hu, Peng Zhu 0001, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | Adaptive Smooth Adversarial Learning for Cross-Domain Hyperspectral Image ClassificationabstractCross-domain classification has emerged as a widely employed method in remote sensing due to the challenges associated with obtaining labels for hyperspectral image (HSI). To address the impact of inter-domain differences on classification performance, domain-adaptive deep learning (DL) methods have proven effective in recent years. However, many existing methods struggle to effectively integrate high-dimensional information in hyperspectral cross-scene classification, are insufficient for inter-class discriminability, and are sensitive to small perturbations. These limitations hinder the accurate extraction of meaningful features across different domains. To tackle these issues, we propose a cross-scene HSI classification method named Adaptive Smooth Adversarial Learning (ASAL). This approach utilizes a dynamic attention mechanism to capture region semantic information, where the spectral attention deformable block (SADB) is designed to achieve dense pixel-level attention. Starting from the concept of class confusion, combining category minimization with multi-scale analysis, we design Entropy-weighted Minimum-class Confusion (EMC) loss, which improves the ability to capture inter-class relationships. The sharpness of the loss function is also reduced using a double smoothing regularization strategy, which enhances the generalization ability of the model. Experimental results show the proposed method improves accuracy by about 3.00 % or higher over state-of-the-art methods across four cross-scene HSI benchmarks. The code for the proposed network will be available at https://github.com/liurongwhm. Mengqing Zhou, Jiaqi Yang 0005 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Hyperspectral Image-Text Coupling Network For Hyperspectral Image ClassificationabstractDeep learning-based approaches have blossomed in the field of hyperspectral image (HSI) classification. However, most of the existing methods focus only on visual information and ignore textual clues. Textual properties can describe the shape, color, and size of ground objects with a wealth of information and are a great complement to visual features. Therefore, the HSI-Text (HSI-T) coupling network is proposed in this study. Specifically, a geocoder is first designed to encode the geospatial information of the image. Moreover, a textual transformer is introduced to extract the attributes of the language. On the basis of the above structures, an HSI-T coupling network is built to explicitly model visual-textual dependencies. Experimental results on benchmark HSI datasets can verify that the proposed approach can gain better classification accuracy and more fine-grained land cover maps. Jiaqi Yang 0005, Bo Du 0001 |
IGARSS | 1 |
| 2024 | ITER: Image-to-Pixel Representation for Weakly Supervised HSI ClassificationabstractRecent years have witnessed the superiority of deep learning-based algorithms in the field of HSI classification. However, a prerequisite for the favorable performance of these methods is a large number of refined pixel-level annotations. Due to atmospheric changes, sensor differences, and complex land cover distribution, pixel-level labeling of high-dimensional hyperspectral image (HSI) is extremely difficult, time-consuming, and laborious. To overcome the above hurdle, an Image-To-pixEl Representation (ITER) approach is proposed in this paper. To the best of our knowledge, this is the first time that image-level annotation is introduced to predict pixel-level classification maps for HSI. The proposed model is along the lines of subject modeling to boundary refinement, corresponding to pseudo-label generation and pixel-level prediction. Concretely, in the pseudo-label generation part, the spectral/spatial activation, spectral-spatial alignment loss, and geographic element enhancement are sequentially designed to locate discriminate regions of each category, optimize multi-domain class activation map (CAM) collaborative training, and refine labels, respectively. For the pixel-level prediction portion, a high frequency-aware self-attention in a high-enhanced transformer is put forward to achieve detailed feature representation. With the two-stage pipeline, ITER explores weakly supervised HSI classification with image-level tags, bridging the gap between image-level annotation and dense prediction. Extensive experiments in three benchmark datasets with state-of-the-art (SOTA) works show the performance of the proposed approach. Jiaqi Yang 0005, Bo Du 0001, Di Wang 0023, Liangpei Zhang 0001 |
IEEE Trans. Image Process. | 1 |
| 2024 | Overcoming the Barrier of Incompleteness: A Hyperspectral Image Classification Full ModelabstractDeep learning-based methods have shown promising outcomes in many fields. However, the performance gain is always limited to a large extent in classifying hyperspectral image (HSI). We discover that the reason behind this phenomenon lies in the incomplete classification of HSI, i.e., existing works only focus on a certain stage that contributes to the classification, while ignoring other equally or even more significant phases. To address the above issue, we creatively put forward three elements needed for complete classification: the extensive exploration of available features, adequate reuse of representative features, and differential fusion of multidomain features. To the best of our knowledge, these three elements are being established for the first time, providing a fresh perspective on designing HSI-tailored models. On this basis, an HSI classification full model (HSIC-FM) is proposed to overcome the barrier of incompleteness. Specifically, a recurrent transformer corresponding to Element 1 is presented to comprehensively extract short-term details and long-term semantics for local-to-global geographical representation. Afterward, a feature reuse strategy matching Element 2 is designed to sufficiently recycle valuable information aimed at refined classification using few annotations. Eventually, a discriminant optimization is formulized in accordance with Element 3 to distinctly integrate multidomain features for the purpose of constraining the contribution of different domains. Numerous experiments on four datasets at small-, medium-, and large-scale demonstrate that the proposed method outperforms the state-of-the-art (SOTA) methods, such as convolutional neural network (CNN)-, fully convolutional network (FCN)-, recurrent neural network (RNN)-, graph convolutional network (GCN)-, and transformer-based models (e.g., accuracy improvement of more than 9% with only five training samples per class). The code will be available soon at https://github.com/jqyang22/ HSIC-FM. Jiaqi Yang 0005, Bo Du 0001, Liangpei Zhang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | LGFormer: Local-to-Global Transformer for Hyperspectral Image ClassificationabstractRecently, many transformer-based approaches have emerged in the field of hyperspectral image (HSI) classification. However, existing transformer-based works either model the information of spectral vectors using a transformer or introduce a vision transformer (ViT) for feature expression of the spatial patch after principal component analysis (PCA), neglecting the spectral-spatial correlation of HSI. Besides, local details and global distributions are always challenging to simultaneously extract in these methods. To address the above issues, a local-to-global transformer (LGFormer) is proposed in this paper. In detail, the proposed approach directly extracts inherent features on the originally spectral-spatial patch, which not only takes the spatial distribution and spectral continuity into account but also preserves the intrinsically spectral-spatial correlation of the HSI cube. Moreover, a local-to-global self-attention (LGSA) including 3-D convolutional neural network (CNN) and ViT is designed in the presented LGFormer. With the above HSI-tailored structure, both fine- and coarse-grained features can be captured by progressively spectral-spatial feature learning. Experimental results on benchmark HSI datasets demonstrate that the proposed LGFormer can outperform other methods in terms of higher accuracy and finer classification maps. Jiaqi Yang 0005, Bo Du 0001, Chen Wu 0003 |
IGARSS | 1 |
| 2023 | Can Spectral Information Work While Extracting Spatial Distribution? - An Online Spectral Information Compensation Network for HSI ClassificationabstractIn the past few years, deep learning-based methods have shown commendable performance for hyperspectral image (HSI) classification. Many works focus on designing independent spectral and spatial branches and then fusing the output features from two branches for category prediction. In this way, the correlation that exists between spectral and spatial information is not completely explored, and spectral information extracted from one branch is always not sufficient. Some studies also try to directly extract spectral-spatial features using 3D convolutions but are accompanied by the severe over-smoothing phenomenon and poor representation ability of spectral signatures. Unlike the above-mentioned approaches, in this paper, we propose a novel online spectral information compensation network (OSICN) for HSI classification, which consists of a candidate spectral vector mechanism, progressive filling process, and multi-branch network. To the best of our knowledge, this paper is the first to online supplement spectral information into the network when spatial features are extracted. The proposed OSICN makes the spectral information participate in network learning in advance to guide spatial information extraction, which truly processes spectral and spatial features in HSI as a whole. Accordingly, OSICN is more reasonable and more effective for complex HSI data. Experimental results on three benchmark datasets demonstrate that the proposed approach has more outstanding classification performance compared with the state-of-the-art methods, even with a limited number of training samples. Jiaqi Yang 0005, Bo Du 0001, Yonghao Xu, Liangpei Zhang 0001 |
IEEE Trans. Image Process. | 1 |
| 2022 | Hybrid Vision Transformer Model for Hyperspectral Image ClassificationabstractDue to the local connectivity property, convolutional neural network (CNN) can effectively extract contextual detailed information. Therefore, a large number of CNN-based methods are introduced to hyperspectral image (HSI) classification. However, receptive fields of these methods are greatly limited, and information extraction process is usually inadequate. Recently, transformer structure has attracted extensive attention owing to its ability to capture global dependency. With a self-attention mechanism, transformer can extract long-tail distribution and model global features to enhance the representation of data. Consequently, it is a natural idea to combine CNN and transformer to obtain both local detail and global distribution. In this paper, we propose a hybrid vision transformer model (Hybrid ViT) to jointly learn global and local information of HSI, including a convolution block and a vision transformer block. With the unified architecture, Hybrid ViT model can not only access detailed features of narrow targets but also extract the global distribution of large objects. Experimental results on benchmark HSI datasets demonstrate that the proposed Hybrid ViT can outperform other methods with higher classification accuracy and finer classification maps. Jiaqi Yang 0005, Bo Du 0001, Chen Wu 0003 |
IGARSS | 1 |
| 2021 | Automatically Adjustable Multi-Scale Feature Extraction Framework for Hyperspectral Image ClassificationabstractRecently, deep learning-based methods have shown the great potential in hyperspectral image (HSI) classification. Nevertheless, feature extraction by convolutional neural network (CNN) is often performed on only one scale, resulting in multi-scale information loss. To address this problem, in this paper, we propose an automatically adjustable multi-scale feature extraction framework (A2MFE-Framework) for hyperspectral classification, including a scale reference network and two scale transformation networks. With the well-designed architecture, A2MFE-Framework can not only extract multiscale features, but also automatically change the network structure to match input features of different scales. Experimental results on two benchmark HSI datasets demonstrate that the A2MFE-Framework can better capture multi-scale features of different objects via an automatically adjustable feature extraction framework with higher classification accuracy compared with previous methods. Jiaqi Yang 0005, Bo Du 0001, Chen Wu 0003, Liangpei Zhang 0001 |
IGARSS | 1 |
| 2021 | Enhanced Multiscale Feature Fusion Network for HSI ClassificationabstractDeep learning-based hyperspectral image (HSI) classification methods have recently attracted significant attention. However, features captured by convolutional neural network (CNN) are always partial due to the restrictions of the respective fields and the loss of multiscale information, which lead to features being discontinuous when extracted. In a departure from existing approaches, in this article, we propose a novel Enhanced Multiscale Feature Fusion Network (EMFFN). As a deeper and wider network, EMFFN can extract sufficiently multiscale features from the parallel multipath of three stages for HSI classification purposes. There are two subnetworks for multiscale spectral and spatial information in EMFFN, respectively. First, we propose a spectral Cascaded Dilated Convolutional Network (CDCN) designed to obtain a larger respective field for long-ranged information and extract multiscale features. Subsequently, a Parallel Multipath Network (PMN) is proposed to capture large-scale, middle-scale, and small-scale spatial features in parallel during all three stages. In the next step, hierarchical features are fused successively, and shallower feature maps can achieve better learning performance when guided by deeper semantic information. As PMN deepens in different stages, more multiscale information flows into the network, enabling finer classification results. To incorporate abundant spectral and spatial features, moreover, we combine features collected from two subnetworks into EMFFN using the designed consolidated loss function. As a result, the network facilitates the learning of not only localization-preserved features, but also high-level semantic features. In our experiments, three benchmark HSIs are utilized to evaluate the performance of the proposed method. Our results demonstrate that the proposed EMFFN can outperform state-of-the-art methods. Jiaqi Yang 0005, Chen Wu 0003, Bo Du 0001, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |