Anyong Qin

dblp:203/2050 · DBLP profile ↗
← Back
28ranked-venue papers
8as first author
21since 2021 · last 2026
0000-0002-2538-822XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 11 · 6 first-author · 8 since 2021Artificial intelligence and machine learning · 8 · 1 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1
YearPublicationVenuePosition
2026 Ground-to-Aerial Scene Adaptation: Unsupervised drone video action recognition via domain adaptation
Feng Yang 0015, Zhijia Li, Fulin Luo, Anyong Qin, Tiecheng Song, Yue Zhao 0012, Chenqiang Gao
Eng. Appl. Artif. Intell.5
2026 Progressive spectral-frequency-spatial guidance network for hyperspectral image dehazing
Qianru Liu, Tiecheng Song, Kaizhao Zhang, Anyong Qin, Feng Yang 0015, Chenqiang Gao
Expert Syst. Appl.5
2026 RotCLIP: Tuning CLIP with visual adapter and textual prompts for rotation robust remote sensing image classification
Tiecheng Song, Anyong Qin
Signal Process. Image Commun.3
2025 Enhanced facial image essence transfer via semantic guidance
Ailin Li, Anyong Qin, Ying Huang 0004
Eng. Appl. Artif. Intell.2
2025 Frequency-prompt guided spectral-spatial transformer for hyperspectral image classification
Tiecheng Song, Longlong Zhang, Anyong Qin, Feng Yang 0015, Chenqiang Gao
Eng. Appl. Artif. Intell.4
2025 Aerial video classification with Window Semantic Enhanced Video Transformers
Feng Yang 0015, Botong Zhou, Xuehua Guan, Anyong Qin, Tiecheng Song, Yue Zhao 0012, Chenqiang Gao
Expert Syst. Appl.5
2025 Global-local prompts guided image-text embedding, alignment and aggregation for multi-label zero-shot learning
Tiecheng Song, Feng Yang 0015, Anyong Qin, Yue Zhao 0012, Chenqiang Gao
J. Vis. Commun. Image Represent.4
2025 Occlusion-aware multi-person pose estimation with keypoint grouping and dual-prompt guidance in crowded scenes
Tiecheng Song, Anyong Qin, Yue Zhao 0012, Feng Yang 0015, Chenqiang Gao
J. Vis. Commun. Image Represent.4
2025 Dual-Branch Residual Network for Cross-Domain Few-Shot Hyperspectral Image Classification With Refined Prototype
Anyong Qin, Chaoqi Yuan, Feng Yang 0015, Tiecheng Song, Chenqiang Gao
IEEE Geosci. Remote. Sens. Lett.1
2025 ConvFormer-CD: Hybrid CNN-Transformer With Temporal Attention for Detecting Changes in Remote Sensing Imagery
abstract
Recently, the combination of Transformers and convolutional neural networks (CNNs) has witnessed significant advancements in change detection (CD) tasks. However, it remains unexplored how to interactively integrate long-range dependency and local information to enhance the model’s global-local context awareness for effectively mitigating pseudo-changes. In addition, accurate identification and distinction of building changes from complex backgrounds still pose challenges due to the insufficient semantic context modeling across time between bi-temporal images. To address these issues, we propose a hybrid model ConvFormer-CD with parallel convolution and multihead self-attention (MSA). This combination enables better interaction of global and local information, thereby enhancing the adaptability to complex scenarios. Moreover, we introduce a novel module called Temporal Attention to establish cross-temporal semantic relationships between image pairs, effectively highlighting change regions by learning shared and nonshared semantics. This enables our model to accurately detect changed targets even in scenarios characterized by intricate geo-spatial arrangements and distributions. To further refine the differences in bi-temporal images, we propose a difference integration module (DIM) that connects the encoder and the decoder to fuse high-level semantic features across channels. We conduct extensive experiments on four benchmark datasets, including LEVIR-CD, LEVIR-CD+, WHU-CD, and S2Looking-CD, which demonstrates that the proposed ConvFormer-CD outperforms other state-of-the-art (SOTA) methods. Our codes will be available athttps://github.com/taomi-lab/ConvFormer-CD.
Feng Yang 0015, Mengtao Li, Wenqiang Shu, Anyong Qin, Tiecheng Song, Chenqiang Gao, Gui-Song Xia
IEEE Trans. Geosci. Remote. Sens.4
2025 Towards Student Actions in Classroom Scenes: New Dataset and Baseline
abstract
Analyzing student actions is an important and challenging task in educational research. Existing efforts have been hampered by the lack of accessible datasets to capture the nuanced action dynamics in classrooms. In this paper, we present a new multi-labelStudent Action Video(SAV) dataset, specifically designed for action detection in classroom settings. The SAV dataset consists of 4,324 carefully trimmed video clips from 758 different classrooms, annotated with 15 distinct student actions. Compared to existing action detection datasets, the SAV dataset stands out by providing a wide range of real classroom scenarios, high-quality video data, and unique challenges, including subtle movement differences, dense object engagement, significant scale differences, varied shooting angles, and visual occlusion. These complexities introduce new opportunities and challenges to advance action detection methods. To benchmark this, we propose a novel baseline method based on a visual transformer, designed to enhance attention to key local details within small and dense object regions. Our method demonstrates excellent performance with a mean Average Precision (mAP) of 67.9% and 27.4% on the SAV and AVA datasets, respectively. This paper not only provides the dataset but also calls for further research into AI-driven educational tools that may transform teaching methodologies and learning outcomes. The code and dataset are released athttps://github.com/Ritatanz/SAV.
Zhuolin Tan, Chenqiang Gao, Anyong Qin, Ruixin Chen, Tiecheng Song, Feng Yang 0015, Deyu Meng
IEEE Trans. Multim.3
2025 Layer-Wise Mutual Information Meta-Learning Network for Few-Shot Segmentation
abstract
The goal of few-shot segmentation (FSS) is to segment unlabeled images belonging to previously unseen classes using only a limited number of labeled images. The main objective is to transfer label information effectively from support images to query images. In this study, we introduce a novel meta-learning framework called layer-wise mutual information (LayerMI), which enhances the propagation of label information by maximizing the mutual information (MI) between support and query features at each layer. Our approach involves the utilization of a LayerMI Block based on information-theoretic co-clustering. This block performs online co-clustering on the joint probability distribution obtained from each layer, generating a target-specific attention map. The LayerMI Block can be seamlessly integrated into the meta-learning framework and applied to all convolutional neural network (CNN) layers without altering the training objectives. Notably, the LayerMI Block not only maximizes MI between support and query features but also facilitates internal clustering within the image. Extensive experiments demonstrate that LayerMI significantly enhances the performance of baseline and achieves competitive performance compared to state-of-the-art methods on three challenging benchmarks: PASCAL- $5^{i}$ , COCO- $20^{i}$ , and FSS-1000.
Xiaoliu Luo, Zhao Duan, Anyong Qin, Zhuotao Tian, Ting Xie 0004, Taiping Zhang, Yuan Yan Tang
IEEE Trans. Neural Networks Learn. Syst.3
2024 Deep Updated Subspace Networks for Few-Shot Remote Sensing Scene Classification
abstract
Due to the difficulty of manually labeling remote sensing scene images and the demand for the ability to recognize new scene classes, few-shot remote sensing scene classification (FSRSSC) has attracted more and more attention. At present, metric-based FSRSSC methods have made promising progress, especially the prototypical networks-based methods. However, due to the complexity of the background of remote sensing scene images, the prototype classifier, which takes the average features of support samples as the metric benchmark, retains the features of category-irrelevant objects and other background information in the image. This leads to a bad classification result. Therefore, in this work, we propose a FSRSSC method based on the deep updated subspace network (DUSN), which uses class subspace as a metric benchmark to represent the commonality of a category and can effectively mitigate the negative impact of irrelevant objects on the classifier. In addition, for the higher inter-class similarity and larger intra-class variance of remote sensing scene images, we further propose an inter-class constraint and an intra-class constraint to mitigate the classification confusion. We leverage the inter-class constraint to make the images of different classes as far apart as possible, and the intra-class constraint to keep the images of the same class clustered as closely together as possible. Experimental results on three public benchmark datasets demonstrate that our method performs better than the state-of-the-art methods for FSRSSC.
Anyong Qin, Fuyang Chen, Lingyun Tang, Feng Yang 0015, Yue Zhao 0012, Chenqiang Gao
IEEE Trans. Geosci. Remote. Sens.1
2024 Few-Shot Learning With Prototype Rectification for Cross-Domain Hyperspectral Image Classification
abstract
Deep learning has been extensively applied to hyperspectral image (HSI) classification and has achieved significant success. However, the number of labeled samples available for HSI classification tasks is typically limited in practical applications, which makes the high-accuracy of HSI small-sample classification still a challenging research task. Therefore, metric-based prototypical networks for few-shot learning (FSL) have become increasingly popular. However, the majority of existing FSL methods typically have problems with biased prototypes and domain shifts. To address these issues, this article proposed a prototype rectification network framework for cross-domain few-shot HSI classification. Specifically, to obtain more representative prototypes, we designed a query-guided prototype rectification module, which can rectify the feature distribution of the support set prototype and obtain a more representative prototype for subsequent training tasks. Then, we introduced a prototype-based interclass loss function to alleviate the interclass confusion that may result from prototype rectification. Furthermore, we construct an intermediate domain between the source domain and the target domain to alleviate domain shift, which helps mitigate the difficulties of domain transfer and achieve a more comprehensive domain alignment. The experimental results on four publicly available HSI datasets demonstrate that our proposed method outperforms the existing FSL methods.
Anyong Qin, Chaoqi Yuan, Xiaoliu Luo, Feng Yang 0015, Tiecheng Song, Chenqiang Gao
IEEE Trans. Geosci. Remote. Sens.1
2024 Change-Aware Cascaded Dual-Decoder Network for Remote Sensing Image Change Detection
abstract
Change detection aims to detect changes of objects or scenes in remote sensing images, which is critical for observing the Earth’s surface. However, due to the insufficient correlation and aggregation of bitemporal features, the existing deep learning methods are still impacted by varied imaging conditions and complicated boundaries of ground objects in high-resolution remote sensing images. To tackle these challenges, we propose a change-aware cascaded dual-decoder network (CACD2Net), which integrates bitemporal features at different levels to facilitate learning change maps from coarse to fine, thus empowering the network to effectively identify changes and refine pixelwise boundaries in a progressive manner. Within the cascaded dual-decoder architecture, the change location decoder utilizes high-level features to generate a coarse change map, which approximates changes’ localization, while the mask refinement decoder further leverages low-level features to create a texture-aware map that captures more texture and structural information about the change regions. By using the coarse change map as guidance and directing the texture-aware map to focus on the details of changes, the boundaries can be gradually refined, ultimately resulting in an accurate change detection mask. We test our model on the season-varying change detection (SVCD) dataset and the Sun Yat-sen University change detection (SYSU-CD) dataset, and the experimental results show that our model surpasses other state-of-the-art change detection methods. Our codes will be available athttps://github.com/Moonquakes0/CACD2Net.
Feng Yang 0015, Yifeng Yuan, Anyong Qin, Yue Zhao 0012, Tiecheng Song, Chenqiang Gao
IEEE Trans. Geosci. Remote. Sens.3
2023 Distribution preserving-based deep semi-NMF for data representation
Anyong Qin, Zhuolin Tan, Xingli Tan, Cheng Jing, Yuan Yan Tang
Neurocomputing1
2023 Distance Constraint-Based Generative Adversarial Networks for Hyperspectral Image Classification
abstract
Hyperspectral image (HSI) classification suffers from two serious problems, one is the limited labeled pixels, and the other is the class imbalance problem. As a result, the number of labeled pixels in many categories is not sufficient to characterize the spectral-spatial information, and train a satisfying deep model. By making full use of the information of unlabeled pixels, semi-supervised methods can provide better classification performance in the case of limited labeled pixels. However, they do not take into account the imbalance in the HSI data. As a method of data enhancement, generative adversarial networks focus on the above two problems and have also been widely used for the task of the HSI classification. In this work, we propose a distance constraints-based generative adversarial networks (DGAN) method for HSI classification to address these two problems. The DGAN employs the convolution autoencoder (AE) to extract the latent features of the HSI samples, and considers the reconstructed samples from the AE as the real samples for the later classifier and discriminator. In addition, the DGAN uses two distance constraints to solve the problems of the few labeled samples and class imbalance, the one latent-data distance constraint enforcing the generator to generate HSI samples for each class (especially the minority class), another discriminator-score distance constraint guiding the generator to synthesize samples that resemble the real HSI samples. Finally, the generated samples are combined classwise with the reconstructed samples and the real HSI samples to learn the parameters of the classifier and discriminator. Experimental results show that our method achieves state-of-the-art performance in terms of overall accuracy (OA) when trained with only 0.5%-4% of data sets from Indian Pines, Pavia University, and Botswana. Specifically, our method demonstrates improvements of 5.48%, 8.79%, and 0.91% on these three datasets, respectively. It reveals the great potential of the DGAN model in generating the HSI samples for each class, which contributes to improving the classification performance of the HSI data.
Anyong Qin, Zhuolin Tan, Yongqing Sun, Feng Yang 0015, Yue Zhao 0012, Chenqiang Gao
IEEE Trans. Geosci. Remote. Sens.1
2022 Active Learning for Hyperspectral Image Classification via Hypergraph Neural Network
abstract
Graph convolution network (GCN) has been extensively applied to the area of hyperspectral image (HSI) classification. However, the graph can not effectively describe the complex relationships between HSI pixels and the GCN still faces the challenge of insufficient labeled pixels. In order to alleviate the above two issues faced by the GCN in HSI classification, we propose a novel framework that integrates the active learning and the hypergraph neural network. First, we construct a hypergraph that can reveal the complex non-pairwise relationships embedded in the hyperspectral images. Next, we train a semi-supervised hypergraph neural network (GNN) with the fewer labeled training set. Then, exploiting the local structural properties of the hypergraph, the most useful HSI pixels are actively selected for labeling. Finally, we fine-tune the GNN with original training set along with the newly labeled pixels. And the last three steps are iteratively carried on. Compared with the other traditional and active learning approaches of HSI classification, the proposed active hypergraph neural network (ACGNN) can achieve better performance on the three HSI datasets.
Yongqing Sun, Anyong Qin, Yukihiro Bandoh, Chenqiang Gao, Yusuke Hiwasaki
ICIP2
2022 Multiscale Spatio-Temporal Network for Aerial Video Event Recognition
abstract
Unmanned aerial vehicles (UAVs) are widely used in the field of remote sensing because of their advantages of providing real-time and high-resolution videos at a low cost. Compared with generic video understanding, aerial video event recognition is faced with emerging challenges: 1) aerial videos contain richer scene information; 2) the scale variations between different videos are large. To address these issues, we propose a Multiscale Spatio-Temporal Network (MSTN) in this paper. More precisely, the MSTN consists of a Pyramid Spatio-Temporal (PST) module and a Multi-Time Scale Decision (MTSD) module, which learn multi-scale spatio-temporal features together. The two modules can better learn spatio-temporal characteristics and boost the performance by 3.7% compared with the baseline method. In ERA, an aerial event recognition dataset, our method achieves the state-of-the-art results.
Feng Yang 0015, Yue Zhao 0012, Anyong Qin, Chenqiang Gao
IGARSS4
2021 Distribution Preserving Deep Semi-Nonnegative Matrix Factorization
abstract
Deep semi-nonnegative matrix factorization can obtain the hidden hierarchical representations according to the unknown attributes of the given data. On the other hand, the inherent structure of the each data cluster can be described by the distribution of the intra-class data. Then one hopes to learn a new low dimensional representation which can preserve the intrinsic structure embedded in the original high dimensional data space perfectly. Here we propose a novel distribution preserving deep semi-nonnegative matrix factorization method (DPNMF) to achieve this goal. As a result, the manifold structures in the raw data are well preserved in the feature space being from the top layer. The experimental results on the real-world datasets show that the proposed algorithm has good performance in terms of cluster accuracy and normalized mutual information (NMI).
Zhuolin Tan, Anyong Qin, Yongqing Sun, Yuan Yan Tang
SMC2
2021 Infrared and Visible Cross-Modal Image Retrieval Through Shared Features
abstract
Image retrieval is one of the key techniques of computer vision, and has been studied for a long time. Nevertheless, little attention is paid to infrared and visible cross-modal retrieval which can be widely used in various applications, e.g., infrared and visible surveillance systems. In this paper, we propose a shared features based infrared-visible cross-modal image retrieval method. The similar visual features are extracted from infrared and visible images as the shared features, and the Euclidean distance is used to measure the similarity between these features. The core of the proposed method comes from three aspects: 1) Feature separation network can separate image features into shared features and exclusive features; 2) Maximum Mean Discrepancy (MMD) loss is employed to constrain the distribution of shared features, which can reduce the retrieval error caused by different imaging angles and similarity of infrared images. 3) The cross-layer fusion encoder compensates for the context loss in the convolution of infrared images. Experimental results on the Infrared-Visible dataset demonstrate the proposed method is effective and outperforms the state-of-the-art approaches.
Fangcen Liu, Chenqiang Gao, Yongqing Sun, Yue Zhao 0012, Feng Yang 0015, Anyong Qin, Deyu Meng
IEEE Trans. Circuits Syst. Video Technol.6
2020 Adaptive Fusion and Mask Refinement Instance Segmentation Network for High Resolution Remote Sensing Images
abstract
Instance segmentation of remote sensing images (RSIs) is an active yet challenging task because of the huge scale variation and arbitrary complex shapes of objects. To address these issues, we propose an adaptive fusion and mask refinement (AFMR) instance segmentation network for RSIs in this paper. More precisely, AFMR consists of an adaptive fusion module to learn multi-scale complementary spatial features in an unsupervised manner, and a content-aware module for segmentation mask refinement. These two modules enable a better feature learning of convolutional neural network and boost the performance by 1.5% compared with the baseline method. In iSAID, a large-scale dataset for RSIs instance segmentation, our AFMR framework achieves the state-of-the-art accuracy, which verifies the superiority of the proposed method.
Jie Ran, Feng Yang 0015, Chenqiang Gao, Yue Zhao 0012, Anyong Qin
IGARSS5
2019 Distribution Preserving Network Embedding
abstract
The deep autoencoder network which is based on constraining non-negative weights, can learn a low dimensional part-based representation. On the other hand, the inherent structure of the each data cluster can be described by the distribution of the intraclass sample. Then one hopes to learn a new low dimensional feature which can preserve the intrinsic structure embedded in the high dimensional data space perfectly. In this paper, by preserving data distribution, a deep part-based representation can be learned, and the novel algorithm is called Distribution Preserving Network Embedding (DPNE). In DPNE, we first need to estimate the distribution of the original data, and then we seek a part-based representation which respects the distribution. The experimental results on real-world data sets show that the proposed algorithm has good performance in terms of cluster accuracy and adjusted mutual information (AMI).
Anyong Qin, Zhaowei Shang, Taiping Zhang, Yuan Yan Tang
ICASSP1
2019 A cost-sensitive meta-learning classifier: SPFCNN-Miner
Linchang Zhao, Zhaowei Shang, Anyong Qin, Taiping Zhang, Yuan Yan Tang
Future Gener. Comput. Syst.3
2019 Spectral-Spatial Sparse Subspace Clustering Based on Three-Dimensional Edge-Preserving Filtering for Hyperspectral Image
abstract
Integrating spatial information into the sparse subspace clustering (SSC) models for hyperspectral images (HSIs) is an effective way to improve clustering accuracy. Since HSI is a three-dimensional (3D) cube datum, 3D spectral-spatial filtering becomes a simple method for extracting the spectral-spatial information. In this paper, a novel spectral-spatial SSC framework based on 3D edge-preserving filtering (EPF) is proposed to improve the clustering accuracy of HSI. First, the initial sparse coefficient matrix is obtained in the sparse representation process of the classical SSC model. Then, a 3D EPF is conducted on the initial sparse coefficient matrix to obtain a more accurate coefficient matrix by solving an optimization problem based on ADMM, which is used to build the similarity graph. Finally, the clustering result of HSI data is achieved by applying the spectral clustering algorithm to the similarity graph. Specifically, the filtered matrix can not only capture the spectral-spatial information but the intensity differences. The experimental results on three real-world HSI datasets demonstrated that the potential of including the proposed 3D EPF into the SSC framework can improve the clustering accuracy.
Ailin Li, Anyong Qin, Zhaowei Shang, Yuan Yan Tang
Int. J. Pattern Recognit. Artif. Intell.2
2019 Spectral-Spatial Graph Convolutional Networks for Semisupervised Hyperspectral Image Classification
abstract
Collecting labeled samples is quite costly and time-consuming for hyperspectral image (HSI) classification task. Semisupervised learning framework, which combines the intrinsic information of labeled and unlabeled samples, can alleviate the deficient labeled samples and increase the accuracy of HSI classification. In this letter, we propose a novel semisupervised learning framework that is based on spectral-spatial graph convolutional networks (S2GCNs). It explicitly utilizes the adjacency nodes in graph to approximate the convolution. In the process of approximate convolution on graph, the proposed method makes full use of the spatial information of the current pixel. The experimental results on three real-life HSI data sets, i.e., Botswana Hyperion, Kennedy Space Center, and Indian Pines, show that the proposed S2GCN can significantly improve the classification accuracy. For instance, the overall accuracy on Indian data is increased from 66.8% (GCN) to 91.6%.
Anyong Qin, Zhaowei Shang, Jinyu Tian 0001, Yulong Wang 0002, Taiping Zhang, Yuan Yan Tang
IEEE Geosci. Remote. Sens. Lett.1
2017 Maximum correntropy criterion for convex anc semi-nonnegative matrix factorization
abstract
Matrix factorization is a popular low dimensional representation approach that plays an important role in many pattern recognition and computer vision domains. Among them, convex and semi-nonnegative matrix factorizations have attracted considerable interest, owing to its clustering interpretation. On the other hand, the generalized correlation function (correntropy) as the error measure does not depend on the assumption of Gaussianity, which the mean square error (MSE) heavily depends on. In this paper, we propose two novel algorithms, called Maximum Correntropy Criterion based Convex and Semi-Nonnegative Matrix Factorization (MCC-ConvexNMF, MCC-SemiNMF). Compared with the mean square error based convex and semi-nonnegative matrix factorization, the proposed methods can extract more information from the data and produce more accurate solutions. Experimental results on both synthetic dataset and the popular face database illustrate the effectiveness of our methods.
Anyong Qin, Zhaowei Shang, Jinyu Tian 0001, Ailin Li, Yulong Wang 0002, Yuan Yan Tang
SMC1
2017 Learning the Distribution Preserving Semantic Subspace for Clustering
abstract
This paper proposes a new clustering method for images called distribution preserving indexing (DPI). It aims to find a lower dimensional semantic space approximating the original image space in the sense of preserving the distribution of the data. In the theory, the intrinsic structure of the data clusters can be described by the distribution of the data effectively. Therefore, the cluster structure of the data in a lower dimensional semantic space derived by the DPI becomes clear. Unlike these distance-based clustering methods, which reveal the intrinsic Euclidean structure of data, our method attempts to discover the intrinsic cluster structure of the data space that actually is the union of some sub-manifolds. Moreover, we propose a revised kernel density estimator for the case of high-dimensional data, which is a crucial step in DPI. In addition, we provide a theoretical analysis of the bound of our method. Finally, the extensive experiments compared with other algorithms, on COIL20, CBCL, and MNIST demonstrate the effectiveness of our proposed approach.
Jinyu Tian 0001, Taiping Zhang, Anyong Qin, Zhaowei Shang, Yuan Yan Tang
IEEE Trans. Image Process.3