Shuai Shao 0006

dblp:71/8201-6 · DBLP profile ↗
← Back
28ranked-venue papers
10as first author
26since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 4 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 8 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 6 since 2021
YearPublicationVenuePosition
2026 CO3+: Improved Collaborative Consortium of Foundation Models for Open-World Few-Shot Learning
abstract
Open-World Few-Shot Learning (OFSL) is a critical research domain focused on accurately identifying target samples under conditions where data is scarce and labels are unreliable. This field is highly relevant to real-world scenarios, holding significant practical implications. Currently, the field has only a few solutions, primarily relying on conventional methods such as metric learning and feature aggregation. However, these methods often struggle in more complex scenarios. Recent breakthroughs in foundation models such as CLIP and DINO have demonstrated their strong representational capabilities, even in resource-limited environments. These advancements have led to a shift from “training model from scratch” towards “exploiting the extensive capabilities and expertise of these pre-trained foundation models for OFSL”. Inspired by this shift, we introduce the Improved Collaborative Consortium of Foundation Models (CO+3), an extension of CO3, first presented in AAAI 2024. CO+3significantly improves the accuracy of OFSL by integrating the strengths of four foundational models. It includes three decoupled blocks: (1) The Label Correction Block (LC-Block) rectifies unreliable labels, (2) the Data Augmentation Block (DA-Block) enriches the available data, and (3) the Text-guided Fusion Adapter (TeFu-Adapter) merges various features and reduces the impact of noisy labels through semantic constraints. We evaluate CO+3across eleven benchmark datasets, comparing it against recent state-of-the-art methods. Our thorough evaluations demonstrate that the proposed CO+3consistently surpasses existing methods by a substantial margin, particularly in high-noise scenarios.
Shuai Shao 0006, Rui Xu 0012, Bingfeng Zhang, Baodi Liu, Weifeng Liu 0001, Yicong Zhou
IEEE Trans. Circuits Syst. Video Technol.1
2025 Excluding the Impossible for Open Vocabulary Semantic Segmentation
abstract
Open vocabulary semantic segmentation is a hot topic in research, focusing on segmenting and recognizing a diverse array of categories in varied environments, including those previously unknown, thereby holding significant practical value. Mainstream studies utilize the CLIP model for direct semantic segmentation (denoted as “forward methods”), which often struggles to represent underrepresented categories effectively. To address this issue, this paper introduces a novel approach Excluding the ImpossibLe Semantic Segmentation Network (ELSE-Net) based on reverse thinking. By excluding improbable categories, ELSE-Net narrows the selection range for forward methods, significantly reducing the risk of misclassification. In implementation, we initially draw on leading research to design the General Processing Block (GP-Block), which generates inclusion probabilities (the likelihood of belonging to a category) by using the CLIP model cooperated with a Mask Proposal Network (MPN). We then present the EXcluding the ImPossible Block (EXP-Block), which computes exclusion probabilities (the likelihood of not belonging to a category) through the CLIPN model and a custom-designed Reverse Retrieval Adapter (R2-Adapter). These exclusion probabilities are subsequently used to refine the inclusion probabilities, which are ultimately employed to annotate class-agnostic masks. Moreover, the core component of our EXP-Block is model-agnostic, enabling it to enhance the capabilities of existing frameworks. Experimental results from four benchmark datasets validate the effectiveness of ELSE-Net and underscore the seamless model-agnostic functionality of the EXP-Block.
Shiyuan Zhao, Baodi Liu, Weifeng Liu 0001, Shuai Shao 0006
AAAI5
2025 Feature aggregation and connectivity for object re-identification
Dongchen Han, Baodi Liu, Shuai Shao 0006, Weifeng Liu 0001, Yicong Zhou
Pattern Recognit.3
2024 Collaborative Consortium of Foundation Models for Open-World Few-Shot Learning
abstract
Open-World Few-Shot Learning (OFSL) is a crucial research field dedicated to accurately identifying target samples in scenarios where data is limited and labels are unreliable. This research holds significant practical implications and is highly relevant to real-world applications. Recently, the advancements in foundation models like CLIP and DINO have showcased their robust representation capabilities even in resource-constrained settings with scarce data. This realization has brought about a transformative shift in focus, moving away from “building models from scratch” towards “effectively harnessing the potential of foundation models to extract pertinent prior knowledge suitable for OFSL and utilizing it sensibly”. Motivated by this perspective, we introduce the Collaborative Consortium of Foundation Models (CO3), which leverages CLIP, DINO, GPT-3, and DALL-E to collectively address the OFSL problem. CO3 comprises four key blocks: (1) the Label Correction Block (LC-Block) corrects unreliable labels, (2) the Data Augmentation Block (DA-Block) enhances available data, (3) the Feature Extraction Block (FE-Block) extracts multi-modal features, and (4) the Text-guided Fusion Adapter (TeFu-Adapter) integrates multiple features while mitigating the impact of noisy labels through semantic constraints. Only the adapter's parameters are adjustable, while the others remain frozen. Through collaboration among these foundation models, CO3 effectively unlocks their potential and unifies their capabilities to achieve state-of-the-art performance on multiple benchmark datasets. https://github.com/The-Shuai/CO3.
Shuai Shao 0006, Yan Wang 0076, Baodi Liu, Bin Liu 0021
AAAI1
2024 DeIL: Direct-and-Inverse CLIP for Open-World Few-Shot Learning
abstract
Open-World Few-Shot Learning (OFSL) is a critical field of research, concentrating on the precise identification of target samples in environments with scarce data and unre-liable labels, thus possessing substantial practical signif-icance. Recently, the evolution of foundation models like CLIP has revealed their strong capacity for representation, even in settings with restricted resources and data. This development has led to a significant shift in focus, tran-sitioning from the traditional method of “building models from scratch” to a strategy centered on “efficiently utilizing the capabilities of foundation models to extract rele-vant prior knowledge tailored for OFSL and apply it judi-ciously”. Amidst this backdrop, we unveil the Direct-and-Inverse CLIP (DeIL), an innovative method leveraging our proposed “Direct-and-Inverse” concept to activate CLIP-based methods for addressing OFSL. This concept transforms conventional single-step classification into a nuanced two-stage process: initially filtering out less probable cate-gories, followed by accurately determining the specific cat-egory of samples. DeIL comprises two key components: a pretrainer (frozen) for data denoising, and an adapter (tun-able) for achieving precise final classification. In experiments, DeIL achieves SOTA performance on 11 datasets. https://github.com/The-Shuai/DeIL.
Shuai Shao 0006, Yan Wang 0076, Baodi Liu, Yicong Zhou
CVPR1
2024 Ensembling Multi-View Discriminative Semantic Feature for Few-Shot Classification
Rui Xu 0012, Shuai Shao 0006, Lei Xing 0005, Yanjiang Wang 0001, Baodi Liu, Weifeng Liu 0001
Eng. Appl. Artif. Intell.2
2024 Feedback-Irrelevant Mapping: An evaluation method for decoupled few-shot classification
Rui Xu 0012, Shuai Shao 0006, Lei Xing 0005, Yanjiang Wang 0001, Baodi Liu, Weifeng Liu 0001
Eng. Appl. Artif. Intell.2
2024 SELM: Self-Motivated Ensemble Learning Model for Cross-Domain Few-Shot Classification in Hyperspectral Images
abstract
Hyperspectral image (HSI) classification is a common task in remote sensing that often faces challenges due to limited samples and cross-domain discrepancies between training and test data. This particular problem is termed as HSI Cross-domain Few-Shot Classification (HSI-CFSC). To solve this problem, we propose a Self-motivated Ensemble Learning Model (SELM). Building upon source pre-trained representations, our end-to-end approach comprises a self-training paradigm to iteratively refine target representations independent of direct source supervision. Moreover, an ensemble classifier suite leveraging diverse decision boundaries is optimized to excavate comprehensive classification cues from limited labeled target data. The OA, AA and Kappa of SELM in UP, PC and Salinas data sets are respectively 86.55%, 82.27%, 82.10%, 98.07%, 94.20%, 97.30% and 91.33%, 94.96%, 90.37%, which achieve the state-of-art performance compared with other classical methods.
Shiyuan Zhao, Shuai Shao 0006, Weifeng Liu 0001, Xinmin Ge, Baodi Liu
IEEE Geosci. Remote. Sens. Lett.3
2024 Few-shot image classification via hybrid representation
Baodi Liu, Shuai Shao 0006, Lei Xing 0005, Weifeng Liu 0001, Weijia Cao, Yicong Zhou
Pattern Recognit.2
2024 FADS: Fourier-Augmentation Based Data-Shunting for Few-Shot Classification
abstract
Collecting a substantial number of labeled samples is infeasible in many real-world scenarios, thereby bringing out challenges for supervised classification. The research on Few-Shot Classification (FSC) aims to address this issue. Current FSC methods mainly leverage ideas such as meta-learning, self-supervised learning, and data augmentation. Among them, data augmentation appears to be an extremely efficient approach to alleviate the aforementioned data-deficiency problem. Here, we propose a novel data augmentation based FSC method termed Fourier-Augmentation based Data-Shunting (FADS). FADS mainly contains two operations, namely Fourier-based data augmentation (FDA) and data shunting. (i) Fourier transform has a desirable property for classification tasks: the image’s phase and amplitude components in the frequency domain correspond to its high-level structure (i.e., semantic) and low-level style (i.e., statistic) information, which do not interfere with each other. Inspired by this observation, we design the FDA operation, which changes the amplitude spectrum of the to-be-augmented images to obtain new images of the same category. (ii) Then we design the data shunting operation to cooperate with the FDA to accomplish FSC. Specifically, it splits the augmented data into different groups to get independent, weak decisions and then fuses them to obtain a unified, strong decision. We conduct experiments on four benchmark datasets. Results show that utilizing our method brings a performance gain of 0.3%-2% in terms of classification accuracy, compared with the classical methods.
Shuai Shao 0006, Yan Wang 0076, Bin Liu 0021, Weifeng Liu 0001, Yanjiang Wang 0001, Baodi Liu
IEEE Trans. Circuits Syst. Video Technol.1
2024 EME: Energy-Based Multiexpert Model for Long-Tailed Remote Sensing Image Classification
abstract
The distribution of remote sensing scene images often follows a long-tailed pattern, where there is an abundance of samples in a few dominant classes and a scarcity of samples in most other classes. This presents two major challenges when it comes to identifying this type of data: Head-Dominance: Models trained on such data tend to prioritize the dominant classes, overlooking the tail classes and resulting in poor performance when it comes to recognizing them. Tail-Interference: The presence of tail classes disrupts the learned representations for the head classes, acting as noise that negatively impacts the recognition accuracy of the head data. To address these challenges, we propose an innovative solution called the energy-based multiexpert (EME) model. The core concept behind this approach is to utilize energy-based discriminators (EDors) to separate the data into head and tail categories. Subsequently, we design multiple experts to classify the head and tail data separately, ensuring that the significant differences in data volume between these categories do not interfere with each other. Experimental results obtained by applying the EME model to three remote sensing datasets demonstrate its efficiency, outperforming current state-of-the-art (SOTA) methods. These findings underscore the effectiveness of our proposed approach in addressing the challenges posed by the long-tailed distribution in remote sensing scene images.
Shuai Shao 0006, Shiyuan Zhao, Weifeng Liu 0001, Dapeng Tao, Baodi Liu
IEEE Trans. Geosci. Remote. Sens.2
2023 CSN: Component supervised network for few-shot classification
Rui Xu 0012, Shuai Shao 0006, Lei Xing 0005, Yujun Wei, Weifeng Liu 0001, Baodi Liu, Yanjiang Wang 0001
Eng. Appl. Artif. Intell.2
2023 Attention-Based Multi-View Feature Collaboration for Decoupled Few-Shot Learning
abstract
Decoupled Few-shot learning (FSL) is an effective methodology that deals with the problem of data-scarce. Its standard paradigm includes two phases: (1) Pre-train. Generating a CNN-based feature extraction model (FEM) via base data. (2) Meta-test. Employing the frozen FEM to obtain the novel data features, then classifying them. Obviously, one crucial factor, the category gap, prevents the development of FSL, i.e., it is challenging for the pre-trained FEM to adapt to the novel class flawlessly. Inspired by a common-sense theory: the FEMs based on different strategies focus on different priorities, we attempt to address this problem from the multi-view feature collaboration (MVFC) perspective. Specifically, we first denoise the multi-view features by subspace learning method, then design three attention blocks (loss-attention block, self-attention block and graph-attention block) to balance the representation between different views. The proposed method is evaluated on four benchmark datasets and achieves significant improvements of 0.9%-5.6% compared with SOTAs.
Shuai Shao 0006, Lei Xing 0005, Yanjiang Wang 0001, Baodi Liu, Weifeng Liu 0001, Yicong Zhou
IEEE Trans. Circuits Syst. Video Technol.1
2023 RAN: Region-Aware Network for Remote Sensing Image Super-Resolution
abstract
The remote sensing (RS) image super-resolution (SR) algorithm aims to reconstruct a high-resolution (HR) image with rich texture details from a given low-resolution (LR) image, improving the spatial resolution. It has been widely concerned in remote sensing image processing and application. Most current deep learning-based methods rely on paired training datasets. However, most datasets are often based on bicubic degradation. This single construction way limits the performance of the pre-trained network. Moreover, SR is an ill-posed problem in that multiple SR images are constructed from a single LR input. This paper proposes a Region-Aware Network (RAN) for remote sensing image super-resolution to alleviate the above issues. First, we introduce the contrastive learning strategy to mine the latent degraded representation of the image and serve as the prior knowledge of the network. Considering the RS images are acquired in specific scenes that have apparent self-similarity. Then, we propose a Region-Aware Module (RAM) based on attention mechanisms and the graph neural network to explore region information and cross-patch self-similarity. Extensive experiments have demonstrated that the proposed RAN adapts to RS image super-resolution tasks with various degradations and performs better in constructing texture information.
Baodi Liu, Lifei Zhao, Shuai Shao 0006, Weifeng Liu 0001, Dapeng Tao, Weijia Cao, Yicong Zhou
IEEE Trans. Geosci. Remote. Sens.3
2023 FVP: Fourier Visual Prompting for Source-Free Unsupervised Domain Adaptation of Medical Image Segmentation
abstract
Medical image segmentation methods normally perform poorly when there is a domain shift between training and testing data. Unsupervised Domain Adaptation (UDA) addresses the domain shift problem by training the model using both labeled data from the source domain and unlabeled data from the target domain. Source-Free UDA (SFUDA) was recently proposed for UDA without requiring the source data during the adaptation, due to data privacy or data transmission issues, which normally adapts the pre-trained deep model in the testing stage. However, in real clinical scenarios of medical image segmentation, the trained model is normally frozen in the testing stage. In this paper, we propose Fourier Visual Prompting (FVP) for SFUDA of medical image segmentation. Inspired by prompting learning in natural language processing, FVP steers the frozen pre-trained model to perform well in the target domain by adding a visual prompt to the input target data. In FVP, the visual prompt is parameterized using only a small amount of low-frequency learnable parameters in the input frequency space, and is learned by minimizing the segmentation loss between the predicted segmentation of the prompted target image and reliable pseudo segmentation label of the target image under the frozen model. To our knowledge, FVP is the first work to apply visual prompts to SFUDA for medical image segmentation. The proposed FVP is validated using three public datasets, and experiments demonstrate that FVP yields better segmentation results, compared with various existing methods.
Yan Wang 0076, Jian Cheng 0002, Shuai Shao 0006, Lanyun Zhu, Zhenzhou Wu, Tao Liu 0067, Haogang Zhu
IEEE Trans. Medical Imaging4
2022 Object re-identification with distribution corrected ranking list
Dongchen Han, Shuai Shao 0006, Weifeng Liu 0001, Baodi Liu
Neurocomputing2
2022 DLDL: Dynamic label dictionary learning via hypergraph regularization
Shuai Shao 0006, Rui Xu 0012, Zhenfang Wang, Weifeng Liu 0001, Yanjiang Wang 0001, Baodi Liu
Neurocomputing1
2022 Learning task-specific discriminative embeddings for few-shot image classification
Lei Xing 0005, Shuai Shao 0006, Weifeng Liu 0001, Anxun Han, Xiangshuai Pan, Baodi Liu
Neurocomputing2
2022 Rethinking Few-Shot Remote Sensing Scene Classification: A Good Embedding Is All You Need?
abstract
In recent years, few-shot remote sensing scene classification (FSRSSC) has attracted more and more attention. For FSRSSC, most methods currently focus on designing a meta-learning algorithm, which obtains meta-knowledge from limited samples and then applies it to novel tasks. In this work, on the one hand, we optimize the training pipeline of the feature extractor; on the other hand, we apply a novel model fusion method further to optimize the feature extractor capability of the feature extractor. We show a novel few-shot remote sensing scene classification baseline: learning two feature representations through using two self-supervised methods on the meta-training set and then fusing the two representations into one. Then, training a linear classifier on this representation achieves state-of-the-art performance. It shows that training a good feature extractor can be more efficient than complex meta-learning algorithms for FSRSSC. We believe that our results can inspire a rethinking of few-shot remote sensing scene classification benchmarks.
Lei Xing 0005, Yuteng Ma, Weijia Cao, Shuai Shao 0006, Weifeng Liu 0001, Baodi Liu
IEEE Geosci. Remote. Sens. Lett.4
2022 Learning to Cooperate: Decision Fusion Method for Few-Shot Remote-Sensing Scene Classification
abstract
Recently, remote-sensing scene classification has become an essential primary research topic. Nowadays, scholars have proposed various few-shot remote-sensing scene classification methods to achieve superior performance with few labeled data. Most of the prior work utilized a meta-learning strategy, which suffered from too little data affecting performance. In this letter, we apply the pre-trained feature extractor for image embedding. Meanwhile, because of the negative transfer problem caused by the inadaptability of the pre-trained feature extractor to remote-sensing data, we propose to exploit two pre-trained models to classify the remote-sensing scene, respectively. Then we fuse the decision to obtain the final classification category. We design a decision attention module to automatically update combination weights for each decision. It comprehensively considers the contribution of various decisions and further improves the discrimination of features. We conduct comprehensive experiments to validate the method and achieve state-of-the-art performance on two benchmark remote-sensing scene datasets, namely NWPU-RESISC45 and UC Merced.
Lei Xing 0005, Shuai Shao 0006, Yuteng Ma, Yanjiang Wang 0001, Weifeng Liu 0001, Baodi Liu
IEEE Geosci. Remote. Sens. Lett.2
2022 Co-Learning for Few-Shot Learning
Rui Xu 0012, Lei Xing 0005, Shuai Shao 0006, Baodi Liu, Kai Zhang 0029, Weifeng Liu 0001
Neural Process. Lett.3
2022 MDFM: Multi-Decision Fusing Model for Few-Shot Learning
abstract
In recent years, researchers pay growing attention to the few-shot learning (FSL) task to address the data-scarce problem. A standard FSL framework is composed of two components: i) Pre-train. Employ the base data to generate a CNN-based feature extraction model (FEM). ii) Meta-test. Apply the trained FEM to the novel data (category is different from base data) to acquire the feature embeddings and recognize them. Although researchers have made remarkable breakthroughs in FSL, there still exists a fundamental problem. Since the trained FEM with base data usually cannot adapt to the novel class flawlessly, the novel data’s feature may lead to the distribution shift problem. To address this challenge, we hypothesize that even if most of the decisions based on different FEMs are viewed asweak decisions, which are not available for all classes, they still perform decent in some specific categories. Inspired by this assumption, we propose a novel method Multi-Decision Fusing Model (MDFM), which comprehensively considers the decisions based on multiple FEMs to enhance the efficacy and robustness of the model. MDFM is a simple, flexible, non-parametric method that can directly apply to the existing FEMs. Besides, we extend the proposed MDFM to two FSL settings (e.g., supervised and semi-supervised settings). We evaluate the proposed method on five benchmark datasets and achieve significant improvements of 3.4%-7.3% compared with state-of-the-arts.
Shuai Shao 0006, Lei Xing 0005, Rui Xu 0012, Weifeng Liu 0001, Yanjiang Wang 0001, Baodi Liu
IEEE Trans. Circuits Syst. Video Technol.1
2022 GCT: Graph Co-Training for Semi-Supervised Few-Shot Learning
abstract
Few-shot learning (FSL), purposing to resolve the problem of data-scarce, has attracted considerable attention in recent years. A popular FSL framework contains two phases: (i) the pre-train phase employs the base data to train a CNN-based feature extractor. (ii) the meta-test phase applies the frozen feature extractor to novel data (novel data has different categories from base data) and designs a classifier for recognition. To correct few-shot data distribution, researchers propose Semi-Supervised Few-Shot Learning (SSFSL) by introducing unlabeled data. Although SSFSL has been proved to achieve outstanding performances in the FSL community, there still exists a fundamental problem: the pre-trained feature extractor cannot adapt to the novel data flawlessly due to the cross-category setting. Usually, large amounts of noises are introduced to the novel feature. We dub it as Feature-Extractor-Maladaptive (FEM) problem. To tackle FEM, we make two efforts in this paper. First, we propose a novel label prediction method, Isolated Graph Learning (IGL). IGL introduces the Laplacian operator to encode the raw data to graph space, which helps reduce the dependence on features when classifying, and then project graph representation to label space for prediction. The key point is that: IGL can weaken the negative influence of noise from the feature representation perspective, and is also flexible to independently complete training and testing procedures, which is suitable for SSFSL. Second, we propose Graph Co-Training (GCT) to tackle this challenge from a multi-modal fusion perspective by extending the proposed IGL to the co-training framework. GCT is a semi-supervised method that exploits the unlabeled samples with two modal features to crossly strengthen the IGL classifier. We estimate our method on five benchmark few-shot learning datasets and achieve outstanding performances compared with other state-of-the-art methods. It demonstrates the effectiveness of our GCT.
Rui Xu 0012, Lei Xing 0005, Shuai Shao 0006, Lifei Zhao, Baodi Liu, Weifeng Liu 0001, Yicong Zhou
IEEE Trans. Circuits Syst. Video Technol.3
2021 OPS-Net: Over-Parameterized Sharing Networks for Video Frame Interpolation
abstract
The video frame interpolation algorithm can improve temporal resolution by inserting non-existent frames in the video sequence. With the help of skip connections, many kernel-based methods train deep neural networks to accurately establish the complicated spatiotemporal relationship among pixels in adjacent frames. Still, these connections are only performed in the feature dimension. To this end, we introduce the Over-Parameterized Sharing Networks (OPS-Net) to implement weight sharing under different layers, capable of integrating deep and shallow features more directly. Specifically, we over-parameterize each convolutional layer to capture movement information efficiently, where the additional trainable weights from distinct ones will be shared. After the training, the additional weights will be fused into the conventional convolutional layer and do not increase the test phase’s computation. Experimental results show that the proposed method can generate favorable frames compared with several state-of-the-art approaches.
Zhenfang Wang, Yanjiang Wang 0001, Shuai Shao 0006, Baodi Liu
ICIP3
2021 SSDL: Self-Supervised Dictionary Learning
abstract
The label-embedded dictionary learning (DL) algorithms generate influential dictionaries by introducing discriminative information. However, there exists a limitation: All the label-embedded DL methods rely on the labels due that this way merely achieves ideal performances in supervised learning. While in semi-supervised and unsupervised learning, it is no longer sufficient to be effective. Inspired by the concept of self-supervised learning (e.g., setting the pretext task to generate a universal model for the downstream task), we propose a Self-Supervised Dictionary Learning (SSDL) framework to address this challenge. Specifically, we first design a p-Laplacian Attention Hypergraph Learning (pAHL) block as the pretext task to generate pseudo soft labels for DL. Then, we adopt the pseudo labels to train a dictionary from a primary label-embedded DL method. We evaluate our SSDL on two human activity recognition datasets. The comparison results with other state-of-the-art methods have demonstrated the efficiency of SSDL.
Shuai Shao 0006, Lei Xing 0005, Wei Yu 0004, Rui Xu 0012, Yanjiang Wang 0001, Baodi Liu
ICME1
2021 MHFC: Multi-Head Feature Collaboration for Few-Shot Learning
abstract
Few-shot learning (FSL) aims to address the data-scarce problem. A standard FSL framework is composed of two components: (1) Pre-train. Employ the base data to generate a CNN-based feature extraction model (FEM). (2) Meta-test. Apply the trained FEM to acquire the novel data's features and recognize them. FSL relies heavily on the design of the FEM. However, various FEMs have distinct emphases. For example, several may focus more attention on the contour information, whereas others may lay particular emphasis on the texture information. The single-head feature is only a one-sided representation of the sample. Besides the negative influence of cross-domain (e.g., the trained FEM can not adapt to the novel class flawlessly), the distribution of novel data may have a certain degree of deviation compared with the ground truth distribution, which is dubbed as distribution-shift-problem (DSP). To address the DSP, we propose Multi-Head Feature Collaboration (MHFC) algorithm, which attempts to project the multi-head features (e.g., multiple features extracted from a variety of FEMs) to a unified space and fuse them to capture more discriminative information. Typically, first, we introduce a subspace learning method to transform the multi-head features to aligned low-dimensional representations. It corrects the DSP via learning the feature with more powerful discrimination and overcomes the problem of inconsistent measurement scales from different head features. Then, we design an attention block to update combination weights for each head feature automatically. It comprehensively considers the contribution of various perspectives and further improves the discrimination of features. We evaluate the proposed method on five benchmark datasets (including cross-domain experiments) and achieve significant improvements of 2.1%-7.8% compared with state-of-the-arts.
Shuai Shao 0006, Lei Xing 0005, Yan Wang 0076, Rui Xu 0012, Yanjiang Wang 0001, Baodi Liu
ACM Multimedia1
2020 Label embedded dictionary learning for image classification
Shuai Shao 0006, Rui Xu 0012, Weifeng Liu 0001, Baodi Liu, Yanjiang Wang 0001
Neurocomputing1
2020 Class specific or shared? A cascaded dictionary learning framework for image classification
Yanjiang Wang 0001, Shuai Shao 0006, Rui Xu 0012, Weifeng Liu 0001, Baodi Liu
Signal Process.2