Zhaokui Li

dblp:143/9268 · DBLP profile ↗
← Back
51ranked-venue papers
16as first author
40since 2021 · last 2026
0000-0002-5331-7664ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 26 · 9 first-author · 26 since 2021Artificial intelligence and machine learning · 13 · 4 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 2 since 2021Systems, architecture and hardware · 2 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Semantic anchor enhanced prototype learning for few-shot hyperspectral image classification
Jinrong He, Xiangqing Zhang, Zhaokui Li
Knowl. Based Syst.4
2025 A Review of Computer Vision-based Hyperspectral Seed Quality Detection
abstract
The integration of hyperspectral imaging (HSI) and computer vision (CV) technologies enables rapid, accurate, and non-destructive seed quality detection methods, providing technical support for accurate seed classification and rational utilization of seed resources. However, in practical applications, challenges such as high computational costs and complex feature extraction remain, leading to limited modeling capability. This paper summarizes applications of CV-based HSI technology in seed quality detection, focusing on the current research status of seed quality-related models and methods, such as 1) variety classification; 2) vigor assessment; 3) moisture content determination; 4) health determination; along with the future development trends.
Chengcheng Chen, Liya Yao, Tiantian Pang, XianChang Wang, HeLong Yu, Jiehong Wu, Zhaokui Li
INDIN7
2025 Few-shot tire tracks recognition based on metric learning
Longlong Zhai, Jinrong He, Zhaokui Li, Yingzhou Bi
Eng. Appl. Artif. Intell.3
2025 Semantic Guided prototype learning for Cross-Domain Few-Shot hyperspectral image classification
Yuhang Li 0014, Jinrong He, Hanchi Liu, Zhaokui Li
Expert Syst. Appl.5
2025 Multimodal prototypical networks with Co-metric fusion for few-shot hyperspectral image classification
Yuhang Li 0014, Jinrong He, Hanchi Liu, Zhaokui Li
Neurocomputing5
2025 Spectral context-aware frequency alignment for few-shot hyperspectral image classification
Jinrong He, Zhaokui Li
Knowl. Based Syst.4
2025 Two-Stage Domain Adaptation for Hyperspectral Image Classification Based on Self-Distillation and Test-Time Adaptation
abstract
Unsupervised domain adaptation methods can effectively mitigate the spectral drift in cross-scene hyperspectral image classification. Among them, adversarial training methods are particularly noteworthy due to their outstanding performance. However, due to the inherent mechanism of adversarial training, these methods suffer from continuing training instability and limited classifier generalizability. To overcome these limitations, this letter proposes a two-stage domain adaptation (TSDA) framework that incorporates self-distillation and test-time adaptation. The self-distillation strategy promotes stability during adversarial training by improving training consistency across iterations. Specifically, each batch includes a subset of data from the previous iteration, and self-distillation ensures that the output of this subset in the current iteration is consistent with the previous one. This mechanism stabilizes gradient computations during the training process, facilitating more robust parameter updates. Subsequently, the test-time adaptation module utilizes a limited set of unlabeled target domain samples to refine the classifier. During this stage, a confident learning module identifies and selects high-confidence pseudo-labels to optimize the classifier, enhancing its generalizability in the target domain. Thus, TSDA facilitates domain adaptation during both the training and testing stages. Experimental results on two cross-domain datasets demonstrate the effectiveness of the proposed method. The code is available at https://github.com/Li-ZK/TSDA-2025.
Zhuoqun Fang, Zhaokui Li, Xuewei Gong
IEEE Geosci. Remote. Sens. Lett.3
2025 An Open-Set Domain Adaptation Framework for Hyperspectral Image Classification With Pixel-Aware Weighting and Decoupled Alignment
abstract
Recent studies have shown that deep domain adaptation techniques perform excellently in cross-domain hyperspectral image classification. However, these methods typically assume that the source domain and the target domain share the same class set, while in practice, the target domain may include unknown classes, and direct alignment can result in negative transfer. Moreover, in hyperspectral image classification based on deep learning, using the label of the central pixel to represent the label of the image patch may lead to feature bias due to the uncertainty of the labels of neighboring pixels, thereby reducing the generalization performance of the model. To address this, this paper proposes an open-set domain adaptation framework, including a Pixel-Aware Weight Learning (PAWL) module and a Decoupled Dual Alignment (DDA) strategy. The PAWL module effectively reduces the feature bias caused by inconsistency in neighboring pixel labels by analyzing the uncertainty of neighboring pixel labels and utilizing adaptive weight learning, thereby improving recognition performance in open-set environments. The DDA strategy decouples the features of the source domain and target domain into known and unknown classes and aligns them separately to mitigate negative transfer. Experiments on two cross-scene hyperspectral datasets validated the effectiveness of the method.
Zhaokui Li, Mingtai Qi, Yan Wang 0087, Xuewei Gong, Cuiwei Liu, Jinjun Wang
IEEE Geosci. Remote. Sens. Lett.1
2025 Entropy-Guided Weighted Adversarial Open-Set Domain Adaptation Method for Hyperspectral Image Classification
abstract
Closed Set Domain Adaptation (CSDA) assumes identical class sets between source and target domains and is an important solution for reducing domain bias. Compared to CSDA, Open Set Domain Adaptation (OSDA) is closer to realworld applications by allowing unknown class samples in the target domain. In addition, previous OSDA methods mainly rely on similarity detection between the target and source domains to identify unknown classes, which does not fully capture the characteristics of the target domain. To address these limitations, this letter proposes an open set domain adaptation method integrating entropy-guided weighted adversarial networks and contrastive self-supervised learning for hyperspectral image (HSI) classification. The approach introduces an entropy-guided weighted adversarial network to distinguish between known and unknown classes in the target domain, while weighing their importance for aligning the feature distributions. Contrastive self-supervised learning is introduced to learn the intrinsic structure and discriminative features of the target domain from unlabeled target domain data. Experimental validation on two HSI cross-domain datasets demonstrates significant performance improvements over existing methods.
Zhaokui Li, Linlin Zeng, Yan Wang 0087, Xuewei Gong, Jiaxu Guo, Mingtai Qi
IEEE Geosci. Remote. Sens. Lett.1
2025 Hyperspectral Target Detection Based on Unsupervised Contrastive Learning and Spatial-Spectral Knowledge Distillation
abstract
Hyperspectral target detection faces challenges due to limited labeled data and high spectral similarity across material samples, making feature extraction difficult. Existing methods for spatial-spectral feature learning are often computationally constrained. To address these issues, we propose a framework combining unsupervised spectral contrastive learning and spatial-spectral knowledge distillation (UCLKD). Our approach uses an unsupervised contrastive learning module to enhance spectral feature representation while preventing model collapse. A spatial-spectral knowledge distillation strategy then leverages features from a hyperspectral Foundation Model to guide the spectral feature extraction network, effectively capturing discriminative spectral features. Extensive experiments demonstrate that UCLKD outperforms existing methods in target detection performance. The code is available at https://github.com/Li-ZK/UCLKD.
Zhaokui Li, Xuewei Gong, Chuanyun Wang, Bo Yuan 0013
IEEE Geosci. Remote. Sens. Lett.2
2025 Incremental Classification of Cross-Scene Hyperspectral Images Based on Dual Constraints and Knowledge Transfer
abstract
With the rapid advancement of hyperspectral imaging technology, there has been a dramatic surge in the volume of hyperspectral data, presenting unprecedented challenges for the incremental classification of hyperspectral images (HSI). In cross-scene HSI incremental classification, two critical challenges emerge: (1) the effectiveness of traditional sample replay methods is limited by high-dimensional data storage and privacy protection constraints; and (2) diverse scenes lead to catastrophic forgetting during new task learning. To address these challenges, this paper proposes an Incremental Classification algorithm based on Dual Constraints and Knowledge Transfer (IC-DCKT) for cross-scene hyperspectral images. First, to overcome the limitations of traditional replay methods in data storage and privacy protection, IC-DCKT innovatively introduces a sample-free storage approach. By combining regularization with knowledge distillation techniques, it achieves efficient knowledge transfer and retention. Second, to effectively mitigate catastrophic forgetting, the algorithm implements a dual-constraint mechanism: the approximate Null Space Projection Constraint (NSPC) restricts gradient update directions to preserve historical task feature distributions, while the Cosine Similarity Distillation Constraint (CSDC) enforces feature alignment between old and new models, significantly enhancing the model’s ability to retain knowledge of old tasks. Finally, this paper pioneers the integration of the pretrained hyperspectral large model HyperSIGMA into an incremental learning framework. The Feature Alignment Loss (FAL) not only improves the speed of new task learning but also compensates for the loss of critical information in new tasks caused by gradient update direction constraints imposed by NSPC. Experimental results on three hyperspectral datasets demonstrate that the proposed IC-DCKT method outperforms existing state-of-the-art incremental learning approaches.
Zhaokui Li, Jiaxu Guo, Yan Wang 0087
IEEE Geosci. Remote. Sens. Lett.2
2025 A Cooperative Meta-Learning and Spectral Diversity Adaptation Framework for Hyperspectral Target Detection
abstract
Meta-learning has demonstrated significant potential in addressing the limited annotation data challenge in hyperspectral target detection. However, existing meta-learningbased methods face two major challenges: 1) weak inter-task correlation leading to unstable optimization directions, and 2) meta-knowledge adaptation based on single prior spectral information fails to effectively characterize spectral variation properties of targets in new scenarios. To overcome these challenges, this letter proposes a cooperative meta-learning framework with spectral diversity adaptation for hyperspectral target detection. The framework introduces a co-learner through cooperative learning to dynamically capture cross-task knowledge and stabilize optimization directions, while designing a spectral diversity-based meta-knowledge adaptation strategy to enhance the model’s ability to understand the spectral variation characteristics of targets in new scenarios and precisely distinguish the spectral features between targets and backgrounds. Experimental results on two public datasets demonstrate that the proposed method outperforms state-of-the-art hyperspectral target detection algorithms. The code is available at https://github.com/Li-ZK/CMLSDA.
Yan Wang 0087, Bo Yuan 0013, Zhaokui Li, Xiaobin Zhao, Jiaxu Guo
IEEE Geosci. Remote. Sens. Lett.3
2025 An Entropy-Driven Clustering and Semantic Association Framework for Cross-Domain Few-Shot Hyperspectral Image Classification
abstract
Recently, few-shot learning (FSL) has shown promising results in hyperspectral image (HSI) classification. However, in practical applications, insufficient labeled training data makes it difficult to capture the intra-class variation of novel classes, making it challenging for the model to learn inaccurate feature distributions, which in turn leads to inaccurate decision boundaries. To solve this problem, we propose an entropy-driven clustering and semantic association framework (ECSA-FSL). We design a deep semantic association feature enhancement module (FEA), which first explores the potential semantic relationship between the source and target domains, and then constructs a cross-domain feature enhancement strategy to generate more discriminative features. In addition, we employ an entropy-driven clustering mechanism (EDC) to optimize the feature space distribution of the target domain. Our approach achieves remarkable classification accuracy with a small number of samples, particularly excelling in scenarios with high intra-class variability and limited training data. Experiments on two publicly available HSI datasets confirm that ECSA-FSL significantly outperforms existing few-shot learning methods under similar conditions. The code is available at https://github.com/Li-ZK/ECSA-FSL-2025.
Yan Wang 0087, Jing Tian 0003, Xuewei Gong, Zhaokui Li
IEEE Geosci. Remote. Sens. Lett.5
2025 RoGLSNet: An Efficient Global-Local Scene Awareness Network With Rotary Position Embedding for Remote Image Segmentation
abstract
Accurate segmentation of very high-resolution remote sensing images is vital for downstream tasks. Most semantic segmentation methods fail to fully consider the inherent characteristics of the images, such as intricate backgrounds, significant intraclass variance, and spatial interdependence of geographic object distribution. To address these challenges, we propose an efficient global–local scene awareness network with rotary position embedding (RoGLSNet). Specifically, we introduce the dynamic global filter (DGF) module to adaptively select frequency components, thereby mitigating interference from background noise. For high intraclass variance, the class center aware block (CCAB) performs class-level contextual modeling with spatial information integration. Additionally, the rotary position embedding (RoPE) is incorporated into vanilla attention to indirectly model the positional and distance relationships of geographic target objects. Extensive experimental results on two widely used datasets demonstrate that RoGLSNet outperforms the state-of-the-art (SOTA) segmentation methods. The code is available athttps://github.com/bai101315/RoGLSNet
Xiaosheng Yu 0001, Weiqi Bai, Jubo Chen, Zhuoqun Fang, Zhaokui Li
IEEE Geosci. Remote. Sens. Lett.6
2025 Multilevel Feature Score Learning for Few-Shot Open-Set Recognition of Hyperspectral Images
Zhaokui Li, Yan Wang 0087, Xuewei Gong, Jiaxu Guo, Jing Tian 0003
IEEE Geosci. Remote. Sens. Lett.2
2025 Open-Set Domain Adaptation for Hyperspectral Image Classification Based on Weighted Generative Adversarial Networks and Dynamic Thresholding
abstract
Recent studies have shown that the deep domain adaptation (DA) technique has achieved remarkable results in cross-domain hyperspectral image (HSI) classification task. However, these DA methods assume that the source and target domains share the same classes, which may not hold true in real-world applications. Under open-set conditions, since the target domain may contain classes unseen in the source domain, direct domain alignment can lead to negative transfer phenomena. Moreover, the presence of multiple unknown classes in the target domain makes it difficult to learn more discriminative classification boundaries between known and unknown classes. To address these issues, we propose an open-set DA (OSDA) method for HSI classification based on weighted generative adversarial networks and dynamic thresholding (WGDT). First, we introduce a class anchor (CA) strategy to learn the metric space of known classes in the source domain. By calculating the similarity between the target-domain samples and the CA, we compute the reliability weights of the samples belonging to known classes. Then, based on these weights, we design an instance-level weighted-domain adversarial learning strategy to better align samples that are more likely to belong to known classes, avoiding negative transfer phenomena. Finally, we propose a dynamic thresholding method to learn the classification boundaries between known and unknown classes in the feature space and reject unknown class samples, thereby separating known class samples in the target domain. The experimental results on four cross-scene HSI classification tasks demonstrate that our proposed method outperforms some existing methods. The code is available athttps://github.com/Li-ZK/WGDT.
Ke Bi, Zhaokui Li, Yushi Chen 0002, Qian Du 0001, Li Ma 0005, Yan Wang 0087, Zhuoqun Fang, Mingtai Qi
IEEE Trans. Geosci. Remote. Sens.2
2025 HZSCM: Hyperspectral Image Zero-Shot Classification via Vision-Language Models
abstract
Most hyperspectral image (HSI) classification methods assume that all classes in the test set are present during training. However, in real-world applications, acquiring labeled training samples is challenging. As a result, it is difficult for the training dataset to cover all possible land cover types, leading to the generalized zero-shot learning (GZSL) problem. Recently, vision-language models (VLMs) have provided rich semantic priors for land cover classes, offering promising potential for GZSL. However, two fundamental gaps hinder their application to HSI classification: the task paradigm gap, arising from the difference between image-level VLMs and the pixel-level HSI classification task; and the knowledge gap, due to the inconsistency between VLM features and HSI spectral–spatial representations. To bridge both gaps, a novel framework leveraging VLM semantic priors for GZSL in HSI classification is proposed, primarily using pseudo-labeling technique to provide knowledge for unseen classes. Specifically, a pseudo-label generation and enhancement module enables a paradigm transition from image-level understanding to pixel-level classification by incorporating HSI’s spatial information. A pseudo-label correction module then refines noisy labels using spectral cues to address the knowledge gap. Finally, a global learning strategy integrates pseudo-label distillation, supervised learning, and feature regularization to classify seen classes while enabling generalization to unseen ones. Experiments on benchmark HSI datasets demonstrate the proposed method’s superiority in generalized zero-shot classification. This work highlights the potential of VLMs in advancing HSI classification in practical applications.
Lingbo Huang, Yushi Chen 0002, Zhaokui Li, Pedram Ghamisi, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.3
2025 Lifelong Learning With Adaptive Knowledge Fusion and Class Margin Dynamic Adjustment for Hyperspectral Image Classification
abstract
With the rapid growth in satellite imagery acquisition and decreasing revisit intervals, efficient on-orbit processing of hyperspectral data has become critical due to limited onboard computing resources. In this context, lifelong learning (LLL) offers a promising solution to enable continuous learning from new data without storing all previous data or retraining from scratch. However, the plasticity-stability dilemma remains a significant challenge, particularly in hyperspectral image (HSI) classification under class-incremental scenarios. To address this, we propose a novel network architecture that integrates contrastive learning and an angular penalty loss. The contrastive learning module facilitates adaptive knowledge fusion, enabling the model to effectively incorporate new information while preserving prior knowledge. The angular penalty loss allows the classifier to dynamically expand for new classes while maintaining discrimination between old and new categories. Together, these components ensure robust knowledge retention, transfer, and adaptability. Experimental results on three benchmark hyperspectral datasets demonstrate that our method significantly outperforms existing approaches, highlighting its efficacy in addressing LLL challenges in HSI classification. The code is available athttps://github.com/Li-ZK/LLL-AFCA.
Zihui Jiang, Zhaokui Li, Yan Wang 0087, Wei Li 0032, Jing Tian 0003, Chuanyun Wang, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.2
2025 Hyperspectral Target Detection Using Diffusion Model and Convolutional Gated Linear Unit
abstract
Deep learning can effectively extract latent information from data to enhance target-background separation in hyperspectral target detection (HTD). However, these models typically require extensive labeled samples, while available target spectra in hyperspectral images (HSI) are scarce. Additionally, existing deep models struggle with target detection in complex backgrounds due to subtle spectral differences. To address these issues, we propose a novel HTD method based on diffusion model and convolutional gated linear unit (HTD-DMCG). First, the diffusion model is integrated with MixUp for data augmentation to generate a diverse and sufficiently large sample set. Next, a Transformer architecture utilizing a convolutional gated linear unit is designed to effectively capture global dependencies and local feature correlations, leading to more discriminative feature representations. Additionally, a new target aggregation and background separation loss is introduced, which emphasizes target sample aggregation while increasing the distance between targets and background samples to enhance separability. The HTD-DMCG method is compared against classical and state-of-the-art HTD methods on four real HSI datasets. Extensive experiments show that it can effectively outperform existing methods in target detection performance. The code is available at https://github.com/Li-ZK/HTD-DMCG.
Zhaokui Li, Xiaobin Zhao, Cuiwei Liu, Xuewei Gong, Wei Li 0032, Qian Du 0001, Bo Yuan 0013
IEEE Trans. Geosci. Remote. Sens.1
2025 Momentum-Enhanced Dual-Prototype Learning Framework for Robust Few-Shot Hyperspectral Image Classification
abstract
Prototypical network-based few-shot learning (FSL) has demonstrated promising performance for hyperspectral image (HSI) classification tasks under scarce sample conditions. However, existing prototype-based FSL methods suffer from data distribution variations among randomly sampled tasks, leading to unstable class prototype representations and weak cross-task generalization with limited samples. To address this issue, we propose a momentum-enhanced dual-prototype learning (MEDPL) framework for robust few-shot HSI classification. Firstly, a momentum-updated prototype mechanism constructs an iteratively optimized prototype memory bank. It obtains accumulated prototypes by exponentially decaying weighted fusion of historical and current prototypes, significantly suppressing noise from randomly sampled data and class center shifts caused by distribution bias. Simultaneously, a class-conditioned perturbation-augmentation strategy is introduced. It generates adaptive noise perturbations for support set features based on learnable covariance matrices to obtain enhanced prototypes, thereby improving the generalization representation capability of class prototypes across tasks. Secondly, a dual-prototype metric learning framework is designed, jointly utilizing accumulated prototypes and enhanced prototypes to synergistically enhance the model’s classification stability and cross-task generalization, thus significantly improving the robustness of few-shot classification. Experimental results demonstrate that MEDPL outperforms other few-shot hyperspectral image classification methods. Our source code is available at https://github.com/hejinrong/MEDPL.
Hanchi Liu, Jinrong He, Xiangqing Zhang, Zhaokui Li
IEEE Trans. Geosci. Remote. Sens.4
2025 A Novel EAGLe Framework for Robust UAV-View Geo-Localization
abstract
This paper addresses the UAV-view geo-localization task, which focuses on bi-directional retrieval between UAV-view and satellite-view images. Generally, existing methods aim to learn image representations that can distinguish between different locations while effectively mitigating the cross-view domain gap. However, these methods often struggle in noisy UAV flight environments, as they fail to account for environmental domain shifts caused by varying weather and lighting conditions. To this end, we propose a novel Environment-Agnostic Geo-Localization (EAGLe) framework, which integrates a dual-objective discriminator and a style mixture module into diverse UAV-view geo-localization networks to enhance their robustness in dynamic environments. Specifically, the dual-objective discriminator not only distinguishes between UAV and satellite views but also identifies various environmental styles in UAV-view images. Through adversarial learning, the dual-objective discriminator encourages the feature encoder to produce features that remain invariant to both viewpoint and environmental variations. Furthermore, the style mixture module is integrated into the feature encoder to extend diversity at the feature level, allowing EAGLe to learn a broader range of environmental styles beyond the training data. Extensive experiments on the University-1652 and SUES-200 datasets demonstrate that the proposed EAGLe significantly improves the reliability of UAV-view geo-localization networks under dynamic and unpredictable environmental conditions, while maintaining inference efficiency.
Cuiwei Liu, Shiting Peng, Shishen Li, Huaijun Qiu, Yuhao Xia, Zhaokui Li, Liang Zhao 0004
IEEE Trans. Geosci. Remote. Sens.6
2024 Enhancing action recognition from low-quality skeleton data via part-level knowledge distillation
Cuiwei Liu, Youzhi Jiang, Chong Du, Zhaokui Li
Signal Process.4
2024 Masked Self-Distillation Domain Adaptation for Hyperspectral Image Classification
abstract
Deep learning-based unsupervised domain adaptation (UDA) has shown potential in cross-scene hyperspectral image (HSI) classification. However, existing methods often experience reduced feature discriminability during domain alignment due to the difficulty of extracting semantic information from unlabeled target domain data. This challenge is exacerbated by ambiguous categories with similar material compositions and the underutilization of target domain samples. To address these issues, we propose a novel masked self-distillation domain adaptation (MSDA) framework, which enhances feature discriminability by integrating masked self-distillation (MSD) into domain adaptation. A class-separable adversarial training (CSAT) module is introduced to prevent misclassification between ambiguous categories by decreasing class correlation. Simultaneously, CSAT reduces the discrepancy between source and target domains through biclassifier adversarial training. Furthermore, the MSD module performs a pretext task on target domain samples to extract class-relevant knowledge. Specifically, MSD enforces consistency between outputs generated from masked target images, where spatial-spectral portions of an HSI patch are randomly obscured, and predictions are produced based on the complete patches by an exponential moving average (EMA) teacher. By minimizing consistency loss, the network learns to associate categorical semantics with unmasked regions. Notably, MSD is tailored for HSI data by preserving the samples’ central pixel and the object to be classified, thus maintaining class information. Consequently, MSDA extracts highly discriminative features by improving class separability and learning class-relevant knowledge, ultimately enhancing UDA performance. Experimental results on four datasets demonstrate that MSDA surpasses the existing state-of-the-art UDA methods for HSI classification. The code is available athttps://github.com/Li-ZK/MSDA-2024.
Zhuoqun Fang, Wenqiang He, Zhaokui Li, Qian Du 0001, Qiusheng Chen
IEEE Trans. Geosci. Remote. Sens.3
2024 Cross-Domain Few-Shot Hyperspectral Image Classification With Cross-Modal Alignment and Supervised Contrastive Learning
abstract
Recently, metric-based few-shot learning (FSL) methods have achieved good performance in hyperspectral image (HSI) classification. However, existing methods suffer from two problems: over-reliance on image modality information leads to inaccurate prototype representation, where a prototype refers to the centroid of each class in the dataset, and the impact of redundant and noisy pixels on model discriminability is rarely considered. These problems result in insufficient discriminability of the model for the target domain. To address the above issues, we propose a cross-domain few-shot HSI classification framework with cross-modal alignment and supervised contrastive learning (CDFS-CASCL). It is well known that human visual learning greatly benefits from the input of various modal information such as vision, language and video. Inspired by the way humans abstract image class concepts in language form and understand the essence of classes, we perform cross-modal alignment (CA) between similar image and text prototypes, and use abstract text semantics to guide the model to learn semantic related features with good generalization ability in images, so as to improve the accuracy of image prototypes representation of the prototypes. In addition, through supervised contrastive learning (SCL) based on neighborhood pixel mask in the target domain, the enhanced sample features belonging to the same class are closer, while the enhanced sample features belonging to different classes are pulled further, enabling the model to learn mask-robust discriminative feature representations, suppressing the negative impact of redundant and noisy pixels, and improving the model’s discriminability. The experimental results demonstrate the superiority of the proposed CDFS-CASCL. The code is available at https://github.com/Li-ZK/CDFS-CASCL-2024.
Zhaokui Li, Yan Wang 0087, Wei Li 0032, Qian Du 0001, Zhuoqun Fang, Yushi Chen 0002
IEEE Trans. Geosci. Remote. Sens.1
2023 A Quantitative Spectra Analysis Framework Combining Mixup and Band Attention for Predicting Soluble Solid Content of Blueberries
Zhaokui Li, Jinen Zhang, Wei Li 0032, Fei Li 0018, Ke Bi
KSEM (4)1
2023 View Distribution Alignment with Progressive Adversarial Learning for UAV Visual Geo-Localization
Cuiwei Liu, Huaijun Qiu, Zhaokui Li, Xiangbin Shi
KSEM (2)4
2023 A Transformer-Based Adaptive Semantic Aggregation Method for UAV Visual Geo-Localization
Shishen Li, Cuiwei Liu, Huaijun Qiu, Zhaokui Li
PRCV (4)4
2023 Supervised Contrastive Learning for Open-Set Hyperspectral Image Classification
abstract
Although hyperspectral image (HSI) classification has made great progress, most classification methods assume that the training and test data have the same class, and that there are no classes in the test data that are not present in the training data. As a result, unknown classes are ignored during model building, which requires the use of open-set classification (OSC) methods to reject unknown classes. However, the current OSC methods do not consider the constraints during feature learning, which can lead to the problem that the feature spaces of known and unknown classes may tend to be consistent. To ensure the discriminability of the feature space and improve the accuracy of the OSC, we propose a novel open-set HSI classification framework based on supervised contrastive learning (OSC-SCL). By adding SCL to spectral and spatial feature learning respectively, not only samples in the same class can be pulled closer, but also unknown classes can be distinguished from known classes. We also introduce a class anchor-based clustering strategy, which can effectively reject unknown classes while ensuring that known classes are correctly classified. Our method is validated on two HSI datasets and outperforms existing state-of-the-art methods.
Zhaokui Li, Ke Bi, Yan Wang 0087, Zhuoqun Fang, Jinen Zhang
IEEE Geosci. Remote. Sens. Lett.1
2023 Few-Shot Hyperspectral Image Classification With Self-Supervised Learning
abstract
Recently, few-shot learning (FSL) has been introduced for hyperspectral image (HSI) classification with few labeled samples. However, existing FSL-based HSI classification methods mainly focus on the meta-knowledge transfer between HSIs. Compared with HSIs, natural images have sufficient annotated data. To utilize natural images (base class data) to achieve accurate classification of HSIs (novel class data), we propose a novel few-shot classification framework with SSL (FSCF-SSL) for HSIs in this article. The orientation of objects in natural images is relatively unitary, whereas the objects of image patches for each pixel in HSIs have diverse orientations in the spatial domain. To make better use of base classes, we design an SSL with geometric transformations (SSLGTs), which sets rotation labels as supervision to extract low-level features that can better represent diverse orientations, and then conduct SSLGT and FSL on base classes to learn transferable spatial meta-knowledge. Next, a spectral-spatial feature extraction network is carefully designed to better utilize the spatial and spectral information of HSIs, where the weights of the first seven layers of the spatial part are initialized by the weights of the corresponding layers trained on base classes. Finally, to fully explore the few annotated data from novel classes, we design an SSL with contrastive learning (SSLCL) that can mine the category-invariant features contained in the novel class data itself, and then perform SSLCL and FSL on novel classes to learn more discriminative individual knowledge. Experimental results on four HSI datasets show that FSCF-SSL offers a significant improvement over state-of-the-art methods. The code is available athttps://github.com/Li-ZK/FSCF-SSL-2023.
Zhaokui Li, Yushi Chen 0002, Cuiwei Liu, Qian Du 0001, Zhuoqun Fang, Yan Wang 0087
IEEE Trans. Geosci. Remote. Sens.1
2023 Supervised Contrastive Learning-Based Unsupervised Domain Adaptation for Hyperspectral Image Classification
abstract
Deep domain adaptation has achieved promising results in cross-domain hyperspectral image (HSI) classification. However, existing methods often focus on aligning data distributions without sufficient consideration of separability of source and target domain data themselves. In addition, current adversarial domain adaptation methods aim to achieve similar distributions between domains by confusing the discriminator, rather than obtaining a more compact distribution. In particular, existing methods are not discriminative enough for the target domain due to the difficulty of obtaining high-confidence labeled samples of the target domain. To address the above challenges, we propose a supervised contrastive learning-based unsupervised domain adaptation for HSI classification. A supervised contrastive learning strategy is then performed in both the source and target domains, which allows samples from the same category to be pulled closer together and samples from different categories to be pushed further apart, thus enhancing the separability of the data within the domain. The domain adaptation task is treated as a one-class classification (OCC) task, and a novel domain similarity loss based on OCC is introduced to reduce the discrepancy between domains. Finally, a confidence learning-based sample selection strategy is designed to select high-confidence labeled samples from the target domain to fine-tune the domain adaptation model, which can enhance the discrimination of the model to the target domain. Experimental results on three cross-domain datasets demonstrate that our proposed method outperforms existing domain adaptation methods. Our source code is available at https://github.com/Li-ZK/SCLUDA-2023.
Zhaokui Li, Li Ma 0005, Zhuoqun Fang, Yan Wang 0087, Wenqiang He, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.1
2022 Hyperspectral Image Classification With Multiattention Fusion Network
abstract
Hyperspectral image (HSI) has hundreds of continuous bands that contain a lot of redundant information. Besides, a spatial patch of a hyperspectral cube often contains some pixels different from the center pixel category, which are usually called interference pixels. The existence of such interference pixels has a negative effect on extracting more discriminative information. Therefore, in this letter, a multiattention fusion network (MAFN) for HSI classification is proposed. Compared with the current state-of-the-art methods, MAFN uses band attention module (BAM) and spatial attention module (SAM), respectively, to alleviate the influence of redundant bands and interfering pixels. In this way, MAFN realizes feature reuse and obtains complementary information from different levels by combining multiattention and multilevel fusion mechanisms, which can extract more representative features. Experiments were conducted on two public HSI data sets to demonstrate the effectiveness of MAFN. Our source code is available athttps://github.com/Li-ZK/MAFN-2021.
Zhaokui Li, Xiaodan Zhao, Yimin Xu, Wei Li 0032, Lin Zhai, Zhuoqun Fang, Xiangbin Shi
IEEE Geosci. Remote. Sens. Lett.1
2022 Soft Augmentation-Based Siamese CNN for Hyperspectral Image Classification With Limited Training Samples
abstract
The lack of training samples remains one of the major obstacles in applying convolutional neural networks (CNNs) to the hyperspectral image (HSI) classification. In this letter, the accurate classification of HSI with limited training samples is investigated. Due to the advantages of minimizing the distance between samples in the same class and maximizing the distance between samples in different classes, siamese CNN is used for HSI classification with limited training samples. After that, to improve the classification performance, data augmentation is investigated for siamese CNN-based HSI classification. Specifically, pair data augmentation based on CutMix is proposed to generate the training pairs of the same or different classes and the new generated training pairs are used to train siamese CNN. Traditional data augmentation methods simply generate new training samples. It is not proper when data augmentation methods keep the loss function and change the content of the input at the same time. Therefore, a soft-loss-based siamese CNN, which changes its loss according to the coupled replacement data augmentation, is proposed to further address the HSI classification with limited training samples. In the experimental part, on two widely used hyperspectral datasets, the influences of different training samples and window sizes are discussed, and the experimental results reveal that the proposed soft-augmentation-based siamese CNN provides competitive results with limited training samples compared with state-of-the-art methods.
Yushi Chen 0002, Xin He 0004, Zhaokui Li
IEEE Geosci. Remote. Sens. Lett.4
2022 Heterogeneous Few-Shot Learning for Hyperspectral Image Classification
abstract
Deep learning has achieved great success in hyperspectral image (HSI) classification. However, its success relies on the availability of sufficient training samples. Unfortunately, the collection of training samples is expensive, time-consuming, and even impossible in some cases. Natural image datasets that are different from HSI, such as Image Net and mini-ImageNet, have abundant texture and structure information. Effective knowledge transfer between two heterogeneous datasets can significantly improve the accuracy of HSI classification. In this letter, heterogeneous few-shot learning (HFSL) for HSI classification is proposed with only a few labeled samples per class. First, few-shot learning is performed on the mini-ImageNet datasets to learn the transferable knowledge. Then, to make full use of the spatial and spectral information, a spectral–spatial fusion network is devised. Spectral information is obtained by the residual network with pure 1-D operators. Spatial information is extracted by a convolution network with pure 2-D operators, and the weights of the spatial network are initialized by the weights of the model trained on the mini-ImageNet datasets. Finally, few-shot learning is fine-tuned on HSI to extract discriminative spectral–spatial features and individual knowledge, which can improve the classification performance of the new classification task. Experiments conducted on two public HSI datasets demonstrate that the HFSL outperforms the existing few-shot learning methods and supervised learning methods for HSI classification with only a few labeled samples. Our source code is available athttps://github.com/Li-ZK/HFSL.
Yan Wang 0087, Zhaokui Li, Qian Du 0001, Yushi Chen 0002, Fei Li 0018, Haibo Yang 0003
IEEE Geosci. Remote. Sens. Lett.4
2022 A Novel Two-Stage Knowledge Distillation Framework for Skeleton-Based Action Prediction
abstract
This letter addresses the challenging problem of action prediction with partially observed sequences of skeletons. Towards this goal, we propose a novel two-stage knowledge distillation framework, which transfers prior knowledge to assist the early prediction of ongoing actions. In the first stage, the action prediction model (also referred to as the student) learns from a couple of teachers to adaptively distill action knowledge at different progress levels for partial sequences. Then the learned student acts as a teacher in the next stage, with the objective of optimizing a better action prediction model in a self-training manner. We design an adaptive self-training strategy from the perspective of undermining the supervision from the annotated labels, since this hard supervision is actually too strict for partial sequences without enough discriminative information. Finally, the action prediction models trained in the two stages jointly constitute a two-stream architecture for action prediction. Extensive experiments on the large-scale NTU RGB+D dataset validate the effectiveness of the proposed method.
Cuiwei Liu, Zhaokui Li, Zhuo Yan, Chong Du
IEEE Signal Process. Lett.3
2022 Confident Learning-Based Domain Adaptation for Hyperspectral Image Classification
abstract
Cross-domain hyperspectral image classification is one of the major challenges in remote sensing, especially for target domain data without labels. Recently, deep learning approaches have demonstrated effectiveness in domain adaptation. However, most of them leverage unlabeled target data only from a statistical perspective but neglect the analysis at the instance level. For better statistical alignment, existing approaches employ the entire unevaluated target data in an unsupervised manner, which may introduce noise and limit the discriminability of the neural networks. In this article, we propose confident learning-based domain adaptation (CLDA) to address the problem from a new perspective of data manipulation. To this end, a novel framework is presented to combine domain adaptation with confident learning (CL), where the former reduces the interdomain discrepancy and generates pseudo-labels for the target instances, from which the latter selects high-confidence target samples. Specifically, the confident learning part evaluates the confidence of each pseudo-labeled target sample based on the assigned labels and the predicted probabilities. Then, high-confidence target samples are selected as training data to increase the discriminative capacity of the neural networks. In addition, the domain adaptation part and the confident learning part are trained alternately to progressively increase the proportion of high-confidence labels in the target domain, thus further improving the accuracy of classification. Experimental results on four datasets demonstrate that the proposed CLDA method outperforms the state-of-the-art domain adaptation approaches. Our source code is available athttps://github.com/Li-ZK/CLDA-2022.
Zhuoqun Fang, Zhaokui Li, Wei Li 0032, Yushi Chen 0002, Li Ma 0005, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.3
2022 Deep Cross-Domain Few-Shot Learning for Hyperspectral Image Classification
abstract
One of the challenges in hyperspectral image (HSI) classification is that there are limited labeled samples to train a classifier for very high-dimensional data. In practical applications, we often encounter an HSI domain (called target domain) with very few labeled data, while another HSI domain (called source domain) may have enough labeled data. Classes between the two domains may not be the same. This article attempts to use source class data to help classify the target classes, including the same and new unseen classes. To address this classification paradigm, a meta-learning paradigm for few-shot learning (FSL) is usually adopted. However, existing FSL methods do not account for domain shift between source and target domain. To solve the FSL problem under domain shift, a novel deep cross-domain few-shot learning (DCFSL) method is proposed. For the first time, DCFSL tackles FSL and domain adaptation issues in a unified framework. Specifically, a conditional adversarial domain adaptation strategy is utilized to overcome domain shift, which can achieve domain distribution alignment. In addition, FSL is executed in source and target classes at the same time, which can not only discover transferable knowledge in the source classes but also learn a discriminative embedding model to the target classes. Experiments conducted on four public HSI data sets demonstrate that DCFSL outperforms the existing FSL methods and deep learning methods for HSI classification. Our source code is available athttps://github.com/Li-ZK/DCFSL-2021.
Zhaokui Li, Yushi Chen 0002, Yimin Xu, Wei Li 0032, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.1
2022 Dual-Channel Residual Network for Hyperspectral Image Classification With Noisy Labels
abstract
Hyperspectral image (HSI) classification has drawn increasing attention recently. However, it suffers from noisy labels that may occur during field surveys due to a lack of prior information or human mistakes. To address this issue, this article proposes a novel dual-channel residual network (DCRN) to resolve HSI classification with noisy labels. Currently, the influence of noisy labels is reduced by simply detecting and removing those anomalous samples. Different from such a specifically designed noise cleansing method, DCRN is easy to implement but highly effective. It enhances its model robustness to noisy labels to a great extent by employing a novel dual-channel structure and a noise-robust loss function. In this way, DCRN can mitigate influence from noisy labels while fully utilizing useful information from mislabeled samples for augmented training. Experiments are conducted on several hyperspectral data sets with manually generated noisy labels to demonstrate its excellent performance. The code is available athttps://github.com/Li-ZK/DCRN-2021.
Yimin Xu, Zhaokui Li, Wei Li 0032, Qian Du 0001, Cuiwei Liu, Zhuoqun Fang, Lin Zhai
IEEE Trans. Geosci. Remote. Sens.2
2021 A Novel Key Point Trajectory Model for Fall Detection from RGB-D Videos
abstract
This paper aims to address the problem of fall detection from RGB-D image sequences. Towards this goal, we propose a novel Key Point Trajectory Model which represents a fall action as a series of trajectory descriptors. In the proposed model, 16 key points including 14 skeleton points and 2 centers of body parts are extracted from each pair of RGB and depth images. Then a global trajectory descriptor is constructed on 16 trajectories that are obtained by connecting the key points across several frames in the RGB-D sequence. The trajectory descriptor incorporates the spatial, depth, and temporal context of key points and characterizes the global motion of human over a short period of time. A random forest is employed to learn the classifier of trajectory descriptors, and an integration rule is developed for detecting falls according to the classification results of all trajectory descriptors within a video. Experiments conducted on two fall detection datasets demonstrate that our method achieves better performance in comparison with state-of-the-art methods.
Cuiwei Liu, Jianxiong Lv, Zhaokui Li, Zhuo Yan, Xiangbin Shi
CSCWD4
2021 Action Prediction Network with Auxiliary Observation Ratio Regression
abstract
This paper focuses on predicting the category of an ongoing action with incomplete observations. We propose a novel Action Prediction Network with Auxiliary Observation Ratio regression (AORAP Net), which enhances the discriminative power of partial video features by encoding the prior knowledge of complete actions. The proposed AORAP Net consists of an encoder, an observation ratio regression module, and an action classifier. The encoder transfers global action information from full videos to construct enhanced partial video features. The observation ratio regression module is developed to achieve an auxiliary task of inferring the progress level of an ongoing action. We demonstrate that this module can guide the encoder to learn enhanced features similar to full video features and discriminative for action prediction. The classifier makes the final decision of the action category in terms of the enhanced features. Considering the fact that there is not enough discriminative information at the early stage of actions, a new classification loss is designed to adapt to the action prediction task and alleviate the over-fitting of training videos. Extensive experiments on the BIT-Interaction dataset and the UT-Interaction dataset validate the effectiveness of the proposed method.
Cuiwei Liu, Yiming Gao 0001, Zhaokui Li, Chong Du, Xiangbin Shi
ICME3
2021 Anti-occlusion and Scale Adaptive Target Tracking Algorithm Based on Kernel Correlation Filter
abstract
The Kernel Correlation Filter Tracking Algorithm(KCF)is a lightweight tracking algorithm with the advantages of fast tracking speed and good effect. However, when the target is occluded and the scale changes, the algorithm will have tracking drift and tracking loss. Aiming at the kernel-related filter tracking algorithm that cannot solve the tracking failure caused by occlusion, target scale changes and other factors in the tracking process, an anti-occlusion and scale-adaptive kernel-related filtering algorithm is proposed. We build a scale pool and use the scale pool to train a one-dimensional fast scale filter to solve the problem of target scale changes. This paper uses the average occlusion distance metric and the size of the context occlusion factor to determine the occlusion state of the target, and dynamically select the learning update rate of the target model according to the target occlusion state. When it is judged that the target is severely occluded, the target position is predicted according to the previous motion state of the target, and the small-range re-detection positioning mechanism proposed in this paper is used to re-detect the target within a certain range of the predicted position. At the same time, the re-detected target is occluded again Judgment to determine whether the target is out of the occlusion. If it is determined that the target is still severely occluded, it means that the target is not out of the occlusion area, the re-detection of the target position is inaccurate, and the predicted position is output. Experimental results show that the accuracy and success rate of the algorithm in this paper are 0.819 and 0.669, which are 8.33% and 7.04% higher than the KCF algorithm. The tracking effect of this algorithm is better than that of KCF algorithm.
Chuanyun Wang, Zhongrui Shi, Keyi Si, Zhaokui Li, Ershen Wang
TrustCom5
2018 A Deep Network Based on Multiscale Spectral-Spatial Fusion for Hyperspectral Classification
Zhaokui Li, Deyuan Zhang, Cuiwei Liu, Yan Wang 0087, Xiangbin Shi
KSEM (2)1
2017 Action Prediction Using Unsupervised Semantic Reasoning
Cuiwei Liu, Yaguang Lu, Xiangbin Shi, Zhaokui Li, Liang Zhao 0004
ICONIP (3)4
2017 Study for ELM-based recognition of fold structure aiming at remote sensing image
abstract
The recognition and classification of the remote sensing image is a key technology in the application of remote sensing image, which has been one of the research hotspots. Extreme learning machine (EML), is a kind of machine learning method, which has advantages of fast learning speed and good generalization capability, and has been getting more and more attention from the researchers. Based on the two research hotspots, in this paper, a method for recognition of fold structure information in the remote sensing image based on ELM is proposed. Through the comparison of results obtained by this method and artificial interpretation, it shows that the method has good accuracy and strong applicability, and also, has important theoretical basis and practical application value in guiding the planning and construction and operation maintenance of special regional highway.
Jiehong Wu, Liangkai Zou, Zhaokui Li
IJCNN4
2017 Collaborator recommendation in heterogeneous bibliographic networks using random walks
Lixin Ding, Zhaokui Li, Runze Wan
Inf. Retr. J.3
2017 Mean Laplacian mappings-based difference LDA for face recognition
Zhaokui Li, Yan Wang 0087, Xiangbin Shi, Runze Wan
Multim. Tools Appl.1
2017 Unsupervised feature selection based on decision graph
Jinrong He, Yingzhou Bi, Lixin Ding, Zhaokui Li, Shenwen Wang
Neural Comput. Appl.4
2017 Image preprocessing method based on local approximation gradient with application to face recognition
Zhaokui Li, Yan Wang 0087, Chunlong Fan, Jinrong He
Pattern Anal. Appl.1
2016 Streaming Applications on Heterogeneous Platforms
Zhaokui Li, Jianbin Fang, Tao Tang 0001, Xuhao Chen 0001, Canqun Yang
NPC1
2016 Score level fusion method based on multiple oblique gradient operators for face recognition
Zhaokui Li, Lixin Ding, Yan Wang 0087
Multim. Tools Appl.1
2014 Face Representation with Gradient orientations and Euler Mapping: Application to Face Recognition
abstract
This paper proposes a simple, yet very powerful local face representation, called the Gradient Orientations and Euler Mapping (GOEM). GOEM consists of two stages: gradient orientations and Euler mapping. In the first stage, we calculate gradient orientations of a central pixel and get the corresponding orientation representations by performing convolution operator. These representation results display spatial locality and orientation properties. To encompass different spatial localities and orientations, we concatenate all these representation results and derive a concatenated orientation feature vector. In the second stage, we define an explicit Euler mapping which maps the space of the concatenated orientation into a complex space. For a mapping image, we find that the imaginary part and the real part characterize the high frequency and the low frequency components, respectively. To encompass different frequencies, we concatenate the imaginary part and the real part and derive a concatenated mapping feature vector. For a given image, we use the two stages to construct a GOEM image and derive an augmented feature vector which resides in a space of very high dimensionality. In order to derive low-dimensional feature vector, we present a class of GOEM-based kernel subspace learning methods for face recognition. These methods, which are robust to changes in occlusion and illumination, apply the kernel subspace learning model with explicit Euler mapping to an augmented feature vector derived from the GOEM representation of face images. Experimental results show that our methods significantly outperform popular methods and achieve state-of-the-art performance for difficult problems such as illumination and occlusion-robust face recognition.
Zhaokui Li, Lixin Ding, Yan Wang 0087, Jinrong He
Int. J. Pattern Recognit. Artif. Intell.1
2014 Intrinsic dimensionality estimation based on manifold assumption
Jinrong He, Lixin Ding, Lei Jiang 0007, Zhaokui Li, Qinghui Hu
J. Vis. Commun. Image Represent.4