VLDB 2026 Research / reviewers in the wild / expert
Aitao Yang
dblp:347/1776
· DBLP profile ↗
11ranked-venue papers
6as first author
11since 2021 · last 2026
0000-0003-2535-2371ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Diff-Mamba: A diffusion-Mamba framework for hyperspectral image classification
Shuaibing Shi, Min Li 0030, Yongqi Yin, Yujie He 0001, Aitao Yang |
Neurocomputing | 5 |
| 2025 | CTFN: Multi-scale CNN and transformer with graph encodings fusion network for hyperspectral image classification
Aitao Yang, Min Li 0030, Yao Ding 0010, Meiqiao Bi, Qinghe Zheng |
Expert Syst. Appl. | 1 |
| 2025 | A robust low-pass filtering graph diffusion clustering framework for hyperspectral images
Aitao Yang, Min Li 0030, Yao Ding 0010, Yaoming Cai, Yuanchao Su |
Knowl. Based Syst. | 1 |
| 2025 | Adaptive Homophily Clustering: Structure Homophily Graph Learning With Adaptive Filter for Hyperspectral ImageabstractHyperspectral image (HSI) clustering is a fundamental yet challenging task that typically operates without training labels. Recent advancements in deep graph clustering methods have shown promise for HSI due to their ability to effectively encode spatial structural information. However, limitations such as inadequate utilization of structural information, poor feature representation, and weak graph update capabilities hinder their performance. In this article, we propose an adaptive homophily structure graph clustering (AHSGC) method for HSI. Our approach begins with the generation of homogeneous regions to process HSI and construct the initial graph. Next, we design an adaptive filter graph encoder that captures both high and low-frequency features for subsequent processing. We then develop a graph embedding clustering self-training decoder using KL Divergence to generate pseudo-labels for network training. To enhance graph learning, we introduce homophily-enhanced structure learning, which updates the graph based on the clustering task. This involves estimating node connections through orient correlation estimation and dynamically adjusting graph edges via graph edge sparsification. Finally, we implement joint network optimization to facilitate self-training and graph updates, with K-means used to express latent features. The clustering accuracy on three datasets is 83.60%, 63.65%, and 86.03%, the FLOPs are 3.57G, 30.62G, and 2.95G. The source code will be available athttps://github.com/DY-HYX. Yao Ding 0010, Weijie Kang, Aitao Yang, Junyang Zhao, Jie Feng 0003, Danfeng Hong, Qinghe Zheng |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | SLCGC: A lightweight Self-supervised Low-Pass Contrastive Graph Clustering Network for Hyperspectral ImagesabstractSelf-supervised hyperspectral image (HSI) clustering remains a fundamental yet challenging task due to the absence of labeled data and the inherent complexity of spatial-spectral interactions. While recent advancements have explored innovative approaches, existing methods face critical limitations in clustering accuracy, feature discriminability, computational efficiency, and robustness to noise, hindering their practical deployment. In this paper, a self-supervised efficient low-pass contrastive graph clustering (SLCGC) is introduced for HSIs. Our approach begins with homogeneous region generation, which aggregates pixels into spectrally consistent regions to preserve local spatial-spectral coherence while drastically reducing graph complexity. We then construct a structural graph using an adjacency matrix A and introduce a low-pass graph denoising mechanism to suppress high-frequency noise in the graph topology, ensuring stable feature propagation. A dual-branch graph contrastive learning module is developed, where Gaussian noise perturbations generate augmented views through two multilayer perceptrons (MLPs), and a cross-view contrastive loss enforces structural consistency between views to learn noise-invariant representations. Finally, latent embeddings optimized by this process are clustered via K-means. Extensive experiments and repeated comparative analysis have verified that our SLCGC contains high clustering accuracy, low computational complexity, and strong robustness. The code source will be available athttps://github.com/DY-HYX. Yao Ding 0010, Aitao Yang, Yaoming Cai, Xiongwu Xiao, Danfeng Hong, Junsong Yuan 0001 |
IEEE Trans. Multim. | 3 |
| 2024 | S²GFormer: A Transformer and Graph Convolution Combining Framework for Hyperspectral Image ClassificationabstractTransformer-based methods have a great ability to model nonlocal interactions between spectral and spatial information, while the local features are easily ignored. Graph convolutional neural networks (GCNs) tend to do well in exploiting neighborhood vertex interactions based on their unique aggregation mechanism, while the ability to extract global information is limited. In this article, we study to comprehensively utilize the advantages of transformer and graph convolution by combining the two structures into a unified Transformer (Graphormer) to construct both local and global interactions for hyperspectral image (HSI) classification, and spatial–spectral features enhanced Graphormer framework (S2GFormer) is proposed. Specifically, a follow patch mechanism is first proposed to transform the pixel in HSI to patches while preserving the local spatial features and reducing the computational cost. Moreover, a patchwise spectral embedding block is designed to extract the spectral features of the patch, in which a neighborhood convolution is inserted for comprehensive spectral information extraction. Finally, a multilayer Graphormer Encoder module is proposed to extract the representative spatial–spectral features from the patch for HSI classification. In our network, we jointly integrate the three aforementioned parts into a unified network, and each component benefits the other. The experimental results demonstrate its suitability for HSI classification when compared with other state-of-the-art (SOTA) classifiers, particularly in scenarios with very limited labeled samples. The code of S2GFormer will be made publicly available at:https://github.com/DY-HYX. Yao Ding 0010, Aitao Yang, Shujun Yang, Yaoming Cai, Weiwei Cai 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | GraphMamba: An Efficient Graph Structure Learning Vision Mamba for Hyperspectral Image ClassificationabstractEfficient extraction of spectral sequences and geospatial information is crucial in hyperspectral image (HSI) classification. Recurrent neural networks (RNNs) and Transformers excel in capturing long-range spectral features, while convolutional neural networks (CNNs) excel in aggregating spatial information through convolutional kernels. However, RNNs and Transformers suffer from low-computational efficiency, and CNNs have limitations in perceiving global contextual information. To address these issues, this article proposes GraphMamba—an efficient graph structure learning vision Mamba for HSI classification. Specifically, GraphMamba is a novel hyperspectral information processing paradigm that preserves spatial-spectral features by constructing spatial-spectral cubes and employs a linear spectral encoder to enhance the operability of subsequent tasks. The core components of GraphMamba include the HyperMamba module, which enhances computational efficiency, and the SpatialGCN module, designed for adaptive spatial context awareness. The HyperMamba mitigates clutter interference by employing a global mask (GM) and introduces a parallel training and inference architecture to alleviate computational bottlenecks. Meanwhile, the SpatialGCN utilizes weighted multihop aggregation (WMA) for spatial encoding, emphasizing highly correlated spatial structural features. This approach enables flexible aggregation of contextual information while minimizing spatial noise interference. Notably, the encoding modules of the proposed GraphMamba architecture are both flexible and scalable, providing a novel approach for the joint mining of spatial-spectral information in hyperspectral images. Extensive experiments were conducted on three different scales of real HSI datasets. When compared with state-of-the-art classification methods, GraphMamba demonstrated superior performance. The core code will be released athttps://github.com/ahappyyang/GraphMamba. Aitao Yang, Min Li 0030, Yao Ding 0010, Leyuan Fang, Yaoming Cai, Yujie He 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | An Efficient and Lightweight Spectral-Spatial Feature Graph Contrastive Learning Framework for Hyperspectral Image ClusteringabstractDue to the scarcity of prior information and the high complexity of spectral data, hyperspectral image (HSI) clustering presents a significant challenge. Although recent deep clustering methods have demonstrated remarkable performance, their intricate network structures and poor robustness hinder their practical application. To address this issue, we propose an efficient and lightweight spectral-spatial feature graph contrastive learning (S2GCL) framework for robust HSI clustering. Specifically, we have designed a novel spectral-spatial feature encoder that fully leverages the information in HSI by incorporating both spatial structure and spectral similarity matrices. To establish a lightweight model, we implement several effective designs: First, S2GCL eliminates the commonly used data augmentation and discriminator in GCL during the generation of positive embeddings. Second, we use a multilayer perceptron (MLP) to produce low-dimensional embeddings instead of relying on graph convolutional networks (GCNs). Third, negative embeddings are generated through row-shuffling, avoiding the use of neural networks. Finally, we propose a multiple boundary loss function to extract complementary information from spatial structures and neighboring nodes, while also constraining the interclass differences between positive and negative examples. We conducted extensive experiments on four publicly available datasets and compared S2GCL with state-of-the-art clustering methods. The results indicate that S2GCL achieves satisfactory performance. The code for S2GCL will be released athttps://github.com/ahappyyang/S2GCL. Aitao Yang, Min Li 0030, Yao Ding 0010, Xiongwu Xiao, Yujie He 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | Anchor-Intermediate Detector: Decoupling and Coupling Bounding Boxes for Accurate Object DetectionabstractAnchor-based detectors have been continuously developed for object detection. However, the individual anchor box makes it difficult to predict the boundary’s offset accurately. Instead of taking each bounding box as a closed individual, we consider using multiple boxes together to get prediction boxes. To this end, this paper proposes the Box Decouple-Couple(BDC) strategy in the inference, which no longer discards the overlapping boxes, but decouples the corner points of these boxes. Then, according to each corner’s score, we couple the corner points to select the most accurate corner pairs. To meet the BDC strategy, a simple but novel model is designed named the Anchor-Intermediate Detector(AID), which contains two head networks, i.e., an anchor-based head and an anchor-free Corner-aware head. The corner-aware head is able to score the corners of each bounding box to facilitate the coupling between corner points. Extensive experiments on MS COCO show that the proposed anchor-intermediate detector respectively outperforms their baseline RetinaNet and GFL method by ∼2.4 and ∼1.2 AP on the MS COCO test-dev dataset without any bells and whistles. Yilong Lv, Min Li 0030, Yujie He 0001, Zhuzhen He, Shao-peng Li 0002, Aitao Yang |
ICCV | 6 |
| 2023 | CDF-net: A convolutional neural network fusing frequency domain and spatial domain featuresabstractAbstract Convolutional neural network (CNN), as a classic deep learning algorithm, has been applied to various computer vision tasks. However, most classic CNN models focus on the extraction and utilisation of spatial domain features, while ignoring the potential ability of frequency domain feature extraction. In this study, the mechanism in the backbone design is explored. Firstly, the traditional DCT formula is converted into a convolution form through mathematical derivation. On the basis of a a new type of convolution, namely the DCT Convolution is designed. It is more applicable to deep learning network architectures. Secondly, based on the DCT Convolution, a new cross‐domain fusion network named CDF‐Net is designed. The frequency domain and spatial domain features of the input sample are extracted and fused by the network. CDF‐Net is a general network framework which can be applied to most existing prevalent networks. Finally, various experiments are conducted. On image classification task, for Imagenet2012 dataset, the method proposed was applied to ResNet50, and the accuracy of Top1 was increased by 3.684%. On object detection task, for COCO2017 dataset, the method proposed in this study was applied to ResNet50 and ResNeXt50, mAP were improved by 0.5% and 1.2% respectively. Aitao Yang, Min Li 0030, Zhaoqing Wu, Yujie He 0001, Xiaohua Qiu, Weidong Du, Yao Gou |
IET Comput. Vis. | 1 |
| 2023 | GTFN: GCN and Transformer Fusion Network With Spatial-Spectral Features for Hyperspectral Image ClassificationabstractTransformer has been widely used in classification tasks for hyperspectral images (HSI) in recent years. Because it can mine spectral sequence information to establish long-range dependence, its classification performance can be comparable with the convolutional neural network (CNN). However, both CNN and Transformer focus excessively on spatial or spectral domain features, resulting in an insufficient combination of spatial-spectral domain information from HSI for modeling. To solve this problem, we propose a new end-to-end graph convolutional network (GCN) and Transformer fusion network with the spatial-spectral feature extraction (GTFN) in this paper, which combines the strengths of GCN and Transformer in both spatial and spectral domain feature extraction, taking full advantage of the contextual information of classified pixels while establishing remote dependencies in the spectral domain compared with previous approaches. In addition, GTFN uses Follow Patch as an input to the GCN and effectively solves the problem of high model complexity while mining the relationship between pixels. It is worth noting that the spectral attention module is introduced in the process of GCN feature extraction, focusing on the contribution of different spectral bands to the classification. More importantly, to overcome the problem that Transformer is too scattered in the frequency domain feature extraction, a neighborhood convolution module is designed to fuse the local spectral domain features. On Indian Pines, Salinas, and Pavia University datasets, the overall accuracies (OAs) of our GTFN are 94.00%, 96.81%, and 95.14%, respectively. The core code of GTFN is released at https://github.com/1useryang/GTFN. Aitao Yang, Min Li 0030, Yao Ding 0010, Danfeng Hong, Yilong Lv, Yujie He 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |