Shiming Chen 0002

dblp:63/3682-2 · DBLP profile ↗
← Back
38ranked-venue papers
16as first author
37since 2021 · last 2026
0000-0001-9633-3392ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 26 · 14 first-author · 26 since 2021Graphics, computer vision, multimedia, augmented reality and games · 20 · 5 first-author · 20 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Mutually Causal Semantic Distillation Network for Zero-Shot Learning
Shiming Chen 0002, Shuhuang Chen, Guosen Xie, Xinge You
Int. J. Comput. Vis.1
2026 From Small to Large: In-Context Learning as a New Paradigm for Domain Generalization
Guanglin Zhou, Zhongyi Han, Shaoan Xie, Shiming Chen 0002, Biwei Huang, Liming Zhu 0001, Xinbo Gao 0001, Lina Yao 0001, Salman Khan 0001
Int. J. Comput. Vis.4
2026 Concept Drift and Long-Tailed Distribution in Fine-Grained Visual Categorization: Benchmark and Method
abstract
Data is the foundation for the development of computer vision, and the establishment of datasets plays an important role in advancing the techniques of fine-grained visual categorization (FGVC). In the existing FGVC datasets used in computer vision, it is generally assumed that each collected instance has fixed characteristics and the distribution of different categories is relatively balanced. In contrast, the real world scenario reveals the fact that the characteristics of instances tend to vary with time and exhibit a long-tailed distribution. Hence, the collected datasets may mislead the optimization of the fine-grained classifiers, resulting in unpleasant performance in real applications. Starting from the real-world conditions and to promote the practical progress of fine-grained visual categorization, we present a Concept Drift and Long-Tailed Distribution (CDLT) dataset. Specifically, the dataset is collected by gathering 11195 images of 250 instances in different species for 47 consecutive months in their natural contexts. The collection process involves dozens of crowd workers for photographing and domain experts for labeling. Meanwhile, we propose a feature recombination framework to address the learning challenges associated with CDLT. Experimental results validate the efficacy of our method while also highlighting the limitations of popular large vision-language models (e.g., CLIP) in the context of long-tailed distributions. This emphasizes the significance of CDLT as a benchmark for investigating these challenges.
Shuo Ye, Shiming Chen 0002, Ruxin Wang 0002, Tianxu Wu, Salman Khan 0001, Fahad Shahbaz Khan, Ling Shao 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2026 Dual Adversarial Perturbations for Zero-Shot Learning
abstract
In Zero-Shot Learning (ZSL), embedding-based methods learn a visual–semantic mapping that leverages the attribute knowledge of seen classes to predict the attributes of unseen classes, enabling knowledge transfer from seen to unseen classes. However, distributional discrepancies between seen and unseen classes introduce an inherent domain shift, and inter-class variations cause the same attribute to be expressed differently across categories. As a result, models trained on seen classes often struggle to accurately recognize attributes in unseen classes, limiting their generalization ability. To address these challenges, we propose DAPZSL, a dual adversarial perturbation framework that enhances the robustness of visual–semantic mappings through Feature-Level Adversarial Perturbation (FAP) and improves the model’s generalization ability via Weight-Level Adversarial Perturbation (WAP). Specifically, FAP generates semantically perturbed adversarial samples at the feature level, and incorporating these samples during training encourages the model to learn more robust visual–semantic mappings that are resilient to semantic variations, which improves attribute recognition on unseen classes. Meanwhile, WAP introduces adversarial perturbations into the model’s weight space, promoting a flatter loss landscape that alleviates overfitting to seen classes and enhances generalization. Extensive experiments on multiple benchmark datasets—including AWA2, SUN, and CUB—demonstrate that DAPZSL significantly outperforms existing ZSL models.
Shiming Chen 0002, Guosen Xie, Chaojian Yu, Xinhua You, Qinmu Peng, Xinge You
IEEE Trans. Circuits Syst. Video Technol.2
2026 Few-Shot Object Detection via Spatial-Channel State Space Model
abstract
Due to the limited training samples in few-shot object detection (FSOD), we observe that current methods may struggle to accurately extract effective features from each channel. Specifically, this issue manifests in two aspects: i) channels with high weights may not necessarily be effective, and ii) channels with low weights may still hold significant value. To handle this problem, we consider utilizing inter-channel correlation to ensure that the novel model can effectively highlight relevant channels and rectify incorrect ones, thereby strengthening channel quality. Since the channel sequence is also 1-dimensional, its similarity with the temporal sequence inspires us to take Mamba for modeling the correlation in the channel sequence Based on this concept, we propose the Spatial-Channel State Space Modeling (SCSM) module for spatial-channel-sequence modeling to accurately extract effective features from each channel. In SCSM, we design the Spatial Feature Modeling (SFM) module to ensure the quality of spatial feature representations. We then introduce the Channel State Modeling (CSM) module, which treats channels as a 1-dimensional sequence and take mamba to capture the correlation between channels. Extensive experiments on the VOC and COCO datasets show that the SCSM module enables the novel detector to improve the quality of channel feature representations and achieve state-of-the-art performance. Code is released at https://github.com/zhimengXin/SCSM.
Zhimeng Xin, Tianxu Wu, Yixiong Zou, Shiming Chen 0002, Dingjie Fu, Xinge You
IEEE Trans. Circuits Syst. Video Technol.4
2025 FSL-Rectifier: Rectify Outliers in Few-Shot Learning via Test-Time Augmentation
abstract
Few-shot learning (FSL) commonly requires a model to identify images (queries) that belong to classes unseen during training, based on a few labelled samples of the new classes (support set) as reference. So far, plenty of algorithms involve training data augmentation to improve the generalization capability of FSL models, but outlier queries or support images during inference can still pose great generalization challenges. In this work, to reduce the bias caused by the outlier samples, we generate additional test-class samples by combining original samples with suitable train-class samples via a generative image combiner. Then, we obtain averaged features via an augmentor, which leads to more typical representations through the averaging. We experimentally and theoretically demonstrate the effectiveness of our method, obtaining a test accuracy improvement proportion of around 10% (e.g., from 46.86% to 53.28%) for trained FSL models. Importantly, given a pretrained image combiner, our method is training-free for off-the-shelf FSL models, whose performance can be improved without extra datasets nor further training of the models themselves.
Yunwei Bai, Ying Kiat Tan, Shiming Chen 0002, Yao Shu, Tsuhan Chen
AAAI3
2025 ZeroMamba: Exploring Visual State Space Model for Zero-Shot Learning
abstract
Zero-shot learning (ZSL) aims to recognize unseen classes by transferring semantic knowledge from seen classes to unseen ones, guided by semantic information. To this end, existing works have demonstrated remarkable performance by utilizing global visual features from Convolutional Neural Networks (CNNs) or Vision Transformers (ViTs) for visual-semantic interactions. Due to the limited receptive fields of CNNs and the quadratic complexity of ViTs, however, these visual backbones achieve suboptimal visual-semantic interactions. In this paper, motivated by the visual state space model (i.e., Vision Mamba), which is capable of capturing long-range dependencies and modeling complex visual dynamics, we propose a parameter-efficient ZSL framework called ZeroMamba to advance ZSL. Our ZeroMamba comprises three key components: Semantic-aware Local Projection (SLP), Global Representation Learning (GRL), and Semantic Fusion (SeF). Specifically, SLP integrates semantic embeddings to map visual features to local semantic-related representations, while GRL encourages the model to learn global semantic representations. SeF combines these two semantic representations to enhance the discriminability of semantic features. We incorporate these designs into Vision Mamba, forming an end-to-end ZSL framework. As a result, the learned semantic representations are better suited for classification. Through extensive experiments on four prominent ZSL benchmarks, ZeroMamba demonstrates superior performance, significantly outperforming the state-of-the-art (i.e., CNN-based and ViT-based) methods under both conventional ZSL (CZSL) and generalized ZSL (GZSL) settings.
Wenjin Hou, Dingjie Fu, Kun Li 0008, Shiming Chen 0002, Hehe Fan, Yi Yang 0001
AAAI4
2025 Interpretable Zero-Shot Learning with Locally-Aligned Vision-Language Model
abstract
Large-scale vision-language models (VLMs), such as CLIP, have achieved remarkable success in zero-shot learning (ZSL) by leveraging large-scale visual-text pair datasets. However, these methods often lack interpretability, as they compute the similarity between an entire query image and the embedded category words, making it difficult to explain their predictions. One approach to address this issue is to develop interpretable models by integrating language, where classifiers are built using discrete attributes, similar to human perception. This introduces a new challenge: how to effectively align local visual features with corresponding attributes based on pre-trained VLMs. To tackle this, we propose LaZSL, a locally-aligned vision-language model for interpretable ZSL. LaZSL employs local visual-semantic alignment via optimal transport to perform interaction between visual regions and their associated attributes, facilitating effective alignment and providing interpretable similarity without the need for additional training. Extensive experiments demonstrate that our method offers several advantages, including enhanced interpretability, improved accuracy, and strong domain generalization. Codes available at: https://github.com/shiming-chen/LaZSL.
Shiming Chen 0002, Bowen Duan 0001, Salman Khan 0001, Fahad Shahbaz Khan
ICCV1
2025 ZeroDiff: Solidified Visual-semantic Correlation in Zero-Shot Learning
abstract
Zero-shot Learning (ZSL) aims to enable classifiers to identify unseen classes. This is typically achieved by generating visual features for unseen classes based on learned visual-semantic correlations from seen classes. However, most current generative approaches heavily rely on having a sufficient number of samples from seen classes. Our study reveals that a scarcity of seen class samples results in a marked decrease in performance across many generative ZSL techniques. We argue, quantify, and empirically demonstrate that this decline is largely attributable to spurious visual-semantic correlations. To address this issue, we introduce ZeroDiff, an innovative generative framework for ZSL that incorporates diffusion mechanisms and contrastive representations to enhance visual-semantic correlations. ZeroDiff comprises three key components: (1) Diffusion augmentation, which naturally transforms limited data into an expanded set of noised data to mitigate generative model overfitting; (2) Supervised-contrastive (SC)-based representations that dynamically characterize each limited sample to support visual feature generation; and (3) Multiple feature discriminators employing a Wasserstein-distance-based mutual learning approach, evaluating generated features from various perspectives, including pre-defined semantics, SC-based representations, and the diffusion process. Extensive experiments on three popular ZSL benchmarks demonstrate that ZeroDiff not only achieves significant improvements over existing ZSL methods but also maintains robust performance even with scarce training data. Our codes are available at https://github.com/FouriYe/ZeroDiff_ICLR25.
Zihan Ye, Shreyank N. Gowda, Shiming Chen 0002, Xiaowei Huang 0001, Fahad Shahbaz Khan, Yaochu Jin, Kaizhu Huang, Xiao-Bo Jin
ICLR3
2025 GenZSL: Generative Zero-Shot Learning Via Inductive Variational Autoencoder
abstract
Remarkable progress in zero-shot learning (ZSL) has been achieved using generative models. However, existing generative ZSL methods merely generate (imagine) the visual features from scratch guided by the strong class semantic vectors annotated by experts, resulting in suboptimal generative performance and limited scene generalization. To address these and advance ZSL, we propose an inductive variational autoencoder for generative zero-shot learning, dubbed GenZSL. Mimicking human-level concept learning, GenZSL operates by inducting new class samples from similar seen classes using weak class semantic vectors derived from target class names (i.e., CLIP text embedding). To ensure the generation of informative samples for training an effective ZSL classifier, our GenZSL incorporates two key strategies. Firstly, it employs class diversity promotion to enhance the diversity of class semantic vectors. Secondly, it utilizes target class-guided information boosting criteria to optimize the model. Extensive experiments conducted on three popular benchmark datasets showcase the superiority and potential of our GenZSL with significant efficacy and efficiency over f-VAEGAN, e.g., 24.7% performance gains and more than $60\times$ faster training speed on AWA2. Codes are available at https://github.com/shiming-chen/GenZSL.
Shiming Chen 0002, Dingjie Fu, Salman Khan 0001, Fahad Shahbaz Khan
ICML1
2025 Semantics-Conditioned Generative Zero-Shot Learning via Feature Refinement
Shiming Chen 0002, Ziming Hong, Xinge You, Ling Shao 0001
Int. J. Comput. Vis.1
2025 Multi-Modal Prompts With Primitives Enhancement for Compositional Zero-Shot Learning
abstract
Compositional zero-shot learning (CZSL) aims to recognize novel compositions of known attributes and objects without requiring additional training data. Recent CZSL methods based on vision-language models(e.g., CLIP) suffer from relying solely on text prompts and neglecting the crucial primitive features within compositions, which limits generalization to unseen compositions. To overcome these limitations, we propose a Multi-modal Prompt and Primitives Enhancement method, termed MPPE, which incorporates two key aspects. First, MPPE introduces both text and visual prompts. The text prompts consist of the composition and its corresponding attribute and object prompts, while the visual prompts leverage image masks generated by the segment anything model (SAM). These masks are integrated via an additional Alpha branch to strengthen the CLIP visual encoder to focus on regions of interest within the image. Second, we design a primitives enhancement (PE) module based on cross-attention, which refines attribute and object features obtained from the CLIP text encoder, thereby enriching the representation of novel composition features. Extensive experiments demonstrate the effectiveness of our approach, achieving state-of-the-art performance on three widely-used CZSL benchmarks in both closed-world and open-world CZSL scenarios. Codes are available at https://github.com/YtJin-git/MPPE.
Yutang Jin, Shiming Chen 0002, Tianle Tong, Weiping Ding 0001, Yisong Wang 0004
IEEE Trans. Circuits Syst. Video Technol.2
2025 DRC: Discrete Representation Classifier With Salient Features via Fixed-Prototype
abstract
Image classification models including convolutional neural networks (CNN) and vision transformers (ViT) commonly employ a fully connected (FC) layer as the classifier. However, the fully connected nature of FC brings large amounts of weight parameters, limits the efficiency of inference, tends to over-fit the training data, and struggles to learn distinct class weights. To solve these problems, we propose a discrete representation classifier (DRC), a generic parameter-free classifier that offers efficiency, robustness, and more discriminative categorization. Specifically, the DRC discards numerous unimportant features and focuses solely on the salient features which are reinforced during training and presented in short discrete form during inference. Unlike the way of learning pseudo-prototypes (weights) from data laden with complex patterns and noises in FC, the DRC introducing discriminative fixed-prototypes which are almost uniformly distributed across the high-dimensional feature space, thus helps the model to learn more distinct boundaries between categories. Further leveraging the advantage of DRC’s focus on salient features, we propose Salient-CAM, which is able to locate the most important region in image without the need for weighting feature maps. The experiments demonstrate that simply replacing the model’s classifier from FC to DRC can lead to a significant acceleration in the whole model’s inference and a more robust classification. Additionally, the proposed Salient-CAM exhibits excellent object localization ability in complex natural scenes.
Qinglei Li, Qi Wang 0079, Yongbin Qin, Xingcai Wu, Shiming Chen 0002, Wu Liu 0005, Yong-Jin Liu 0001, Jiebo Luo 0001
IEEE Trans. Circuits Syst. Video Technol.6
2025 Adversarial Feature Training for Few-Shot Object Detection
abstract
Currently, most few-shot object detection (FSOD) methods apply the two-stage training strategy, which first requires training in abundant base classes and transfers the learned prior knowledge to the novel stage. However, due to the inherent imbalance between the base and novel classes, the trained model tends to have a bias toward recognizing novel classes as base ones when they are similar. To address this problem, we propose an adversarial feature training (AFT) strategy aimed at effectively calibrating the decision boundary between novel and base classes to alleviate classification confusion in FSOD. Specifically, we introduce the Classification Level Fast Gradient Sign Method (CL-FGSM), which leverages gradient information from the classifier module to generate adversarial samples with extra feature attention. By attacking the high-level features, we can create adversarial feature samples that are combined with clean high-level features in a suitable range of proportions. Such adversarial feature samples, generated by CL-FGSM, are then combined with clean high-level features in a suitable range of proportions to train the few-shot detector. By this, the novel model is forced to learn extra class-specific features that improve the robustness of the classifier to establish a correct decision boundary, which avoids confusion between base and novel classes in FSOD. Extensive experiments demonstrate that our proposed AFT strategy effectively calibrates the classification decision boundary to avoid classification confusion between base and novel classes and significantly improves the performance of FSOD. Our code is available athttps://github.com/wutianxu/AFT.
Tianxu Wu, Zhimeng Xin, Shiming Chen 0002, Yixiong Zou, Xinge You
IEEE Trans. Circuits Syst. Video Technol.3
2025 HCVP: Leveraging Hierarchical Contrastive Visual Prompt for Domain Generalization
abstract
Domain Generalization (DG) endeavors to create machine learning models that excel in unseen scenarios by learning invariant features. In DG, the prevalent practice of constraining models to a fixed structure or uniform parameterization to encapsulate invariant features can inadvertently blend specific aspects. Such an approach struggles with nuanced differentiation of inter-domain variations and may exhibit bias towards certain domains, hindering the precise learning of domain-invariant features. Recognizing this, we introduce a novel method designed to supplement the model with domain-level and task-specific characteristics. This approach aims to guide the model in more effectively separating invariant features from specific characteristics, thereby boosting the generalization. Building on the emerging trend of visual prompts in the DG paradigm, our work introduces the novelHierarchicalContrastiveVisualPrompt (HCVP) methodology. This represents a significant advancement in the field, setting itself apart with a unique generative approach to prompts, alongside an explicit model structure and specialized loss functions. Differing from traditional visual prompts that are often shared across entire datasets, HCVP utilizes a hierarchical prompt generation network enhanced by prompt contrastive learning. These generative prompts are instance-dependent, catering to the unique characteristics inherent to different domains and tasks. Additionally, we devise a prompt modulation network that serves as a bridge, effectively incorporating the generated visual prompts into the vision transformer backbone. Experiments conducted on five DG datasets demonstrate the effectiveness of HCVP, outperforming both established DG algorithms and adaptation protocols.
Guanglin Zhou, Zhongyi Han, Shiming Chen 0002, Biwei Huang, Liming Zhu 0001, Tongliang Liu, Lina Yao 0001, Kun Zhang 0001
IEEE Trans. Multim.3
2025 Toward Disentangled and Controllable Deep Metric Learning With Human-Like Concept Decomposition
abstract
Deep metric learning (DML) has shown significant advancements in learning discriminative embeddings for images, playing a crucial role in various vision tasks. However, existing methods typically rely on deep neural networks to extract holistic embeddings, which are challenging to disentangle and interpret. To address this issue, we take inspiration from human cognition, where objects are decomposed into distinct concepts for better understanding. Specifically, we propose the concept metrics network (CMNs) to achieve disentangled and controllable DML. CMN begins by initializing learnable concept vectors to represent various visual concepts. These vectors are then associated with regional visual features via cross-attention mechanism, ensuring each vector corresponds to specific visual properties. Finally, the concept values, determined by their presence in the image, form the output embedding. Comprehensive experiments demonstrate that CMN effectively disentangles visual concepts, with each embedding dimension corresponding to a specific concept. Our method not only outperforms existing state-of-the-art methods in conventional DML application (i.e., image retrieval), but also enables more flexible and controllable application. The code is available at https://github.com/shchen0001/CMN.
Shuhuang Chen, Shiming Chen 0002, Shuo Ye, Yuetian Wang, Xinge You
IEEE Trans. Neural Networks Learn. Syst.2
2025 Visual-Semantic Graph Matching Net for Zero-Shot Learning
abstract
Zero-shot learning (ZSL) aims to leverage additional semantic information to recognize unseen classes. To transfer knowledge from seen to unseen classes, most ZSL methods often learn a shared embedding space by simply aligning visual embeddings with semantic prototypes. However, methods trained under this paradigm often struggle to learn robust embedding space because they align the two modalities in an isolated manner among classes, which ignore the crucial class relationship during the alignment process. To address the aforementioned challenges, this article proposes a visual-semantic graph matching net (VSGMN), which leverages semantic relationships among classes to aid in visual-semantic embedding. VSGMN uses a graph build net (GBN) and a graph matching net (GMN) to achieve two-stage visual-semantic alignment. Specifically, GBN first uses an embedding-based approach to build visual and semantic graphs in the semantic space and align the embedding with its prototype for first-stage alignment. In addition, to supplement unseen class relationships in these graphs, GBN also builds the unseen class nodes based on semantic relationships. In the second stage, GMN continuously integrates neighbor and cross-graph information into the constructed graph nodes and aligns the node relationships between the two graphs under the class relationship constraint. Extensive experiments on three benchmark datasets demonstrate that VSGMN achieves superior performance in both conventional and generalized ZSL (GZSL) scenarios. The implementation of our VSGMN and experimental results are available at github: https://github.com/dbwfd/VSGMN.
Bowen Duan 0001, Shiming Chen 0002, Yufei Guo 0001, Guosen Xie, Weiping Ding 0001, Yisong Wang 0004
IEEE Trans. Neural Networks Learn. Syst.2
2024 Progressive Semantic-Guided Vision Transformer for Zero-Shot Learning
abstract
Zero-shot learning (ZSL) recognizes the unseen classes by conducting visual-semantic interactions to transfer se-mantic knowledge from seen classes to unseen ones, sup-ported by semantic information (e.g., attributes). However, existing ZSL methods simply extract visual features using a pre-trained network backbone (i.e., CNN or ViT), which fail to learn matched visual-semantic correspondences for rep-resenting semantic-related visual features as lacking of the guidance of semantic information, resulting in undesirable visual-semantic interactions. To tackle this issue, we pro-pose a progressive semantic-guided vision transformer for zero-shot learning (dubbed ZSLViT). ZSLViT mainly consid-ers two properties in the whole network: i) discover the semantic-related visual representations explicitly, and ii) discard the semantic-unrelated visual information. Specif-ically, we first introduce semantic-embedded token learning to improve the visual-semantic correspondences via semantic enhancement and discover the semantic-related visual tokens explicitly with semantic-guided token attention. Then, we fuse low semantic-visual correspondence visual tokens to discard the semantic-unrelated visual in-formation for visual enhancement. These two operations are integrated into various encoders to progressively learn semantic-related visual representations for accurate visual-semantic interactions in ZSL. The extensive experiments show that our ZSLViT achieves significant performance gains on three popular benchmark datasets, i.e., CUB, SUN, and AWA2.
Shiming Chen 0002, Wenjin Hou, Salman Khan 0001, Fahad Shahbaz Khan
CVPR1
2024 Visual-Augmented Dynamic Semantic Prototype for Generative Zero-Shot Learning
abstract
Generative Zero-shot learning (ZSL) learns a generator to synthesize visual samples for unseen classes, which is an effective way to advance ZSL. However, existing generative methods rely on the conditions of Gaussian noise and the predefined semantic prototype, which limit the generator only optimized on specific seen classes rather than characterizing each visual instance, resulting in poor generalizations (e.g., overfitting to seen classes). To address this issue, we propose a novel Visual-Augmented Dynamic Semantic prototype method (termed VADS) to boost the generator to learn accurate semantic-visual mapping by fully exploiting the visual-augmented knowledge into semantic conditions. In detail, VADS consists of two modules: (1) Visual-aware Domain Knowledge Learning module (VDKL) learns the local bias and global prior of the visual features (referred to as domain visual knowledge), which replace pure Gaussian noise to provide richer prior noise information; (2) VisionOriented Semantic Updation module (VOSU) updates the semantic prototype according to the visual representations of the samples. Ultimately, we concatenate their output as a dynamic semantic prototype, which serves as the condition of the generator. Extensive experiments demonstrate that our VADS achieves superior CZSL and GZSL performances on three prominent datasets and outperforms other state-of-the-art methods with averaging increases by 6.4%, 5.9% and 4.2% on SUN, CUB and AWA2, respectively.
Wenjin Hou, Shiming Chen 0002, Shuhuang Chen, Ziming Hong, Xuetao Feng, Salman Khan 0001, Fahad Shahbaz Khan, Xinge You
CVPR2
2024 Improving Non-Transferable Representation Learning by Harnessing Content and Style
abstract
Non-transferable learning (NTL) aims to restrict the generalization of models toward the target domain(s). To this end, existing works learn non-transferable representations by reducing statistical dependence between the source and target domain. However, such statistical methods essentially neglect to distinguish between *styles* and *contents*, leading them to inadvertently fit (i) spurious correlation between *styles* and *labels*, and (ii) fake independence between *contents* and *labels*. Consequently, their performance will be limited when natural distribution shifts occur or malicious intervention is imposed. In this paper, we propose a novel method (dubbed as H-NTL) to understand and advance the NTL problem by introducing a causal model to separately model *content* and *style* as two latent factors, based on which we disentangle and harness them as guidances for learning non-transferable representations with intrinsically causal relationships. Specifically, to avoid fitting spurious correlation and fake independence, we propose a variational inference framework to disentangle the naturally mixed *content factors* and *style factors* under our causal model. Subsequently, based on dual-path knowledge distillation, we harness the disentangled two *factors* as guidances for non-transferable representation learning: (i) we constraint the source domain representations to fit *content factors* (which are the intrinsic cause of *labels*), and (ii) we enforce that the target domain representations fit *style factors* which barely can predict labels. As a result, the learned feature representations follow optimal untransferability toward the target domain and minimal negative influence on the source domain, thus enabling better NTL performance. Empirically, the proposed H-NTL significantly outperforms competing methods by a large margin.
Ziming Hong, Zhenyi Wang 0001, Li Shen 0008, Yu Yao 0005, Shiming Chen 0002, Chuanwu Yang, Mingming Gong, Tongliang Liu
ICLR6
2024 FSL-QuickBoost: Minimal-Cost Ensemble for Few-Shot Learning
abstract
Few-shot learning (FSL) usually trains models on data from one set of classes, but tests them on data from a different set of classes, providing a few labeled support samples of the unseen classes as a reference for the trained model. Due to the lack of target-relevant training data, there is usually high generalization error with respect to the test classes. In this work, we conduct empirical explorations and propose an ensemble method (namely QuickBoost), which is efficient and effective for improving the generalization of FSL. Specifically, QuickBoost includes an alternative-architecture pretrained encoder with a one-vs-all binary classifier (namely FSL-Forest) based on random forest algorithm, and is ensembled with the off-the-shelf FSL models via logit-level averaging. Experiments on three benchmarks demonstrate that our method achieves state-of-the-art performance with good efficiency. Codes are available at https://github.com/WendyBaiYunwei/FSL-QuickBoost.
Yunwei Bai, Bill Yang Cai, Ying Kiat Tan, Zangwei Zheng, Shiming Chen 0002, Tsuhan Chen
ACM Multimedia5
2024 Causal Visual-semantic Correlation for Zero-shot Learning
Shuhuang Chen, Dingjie Fu, Shiming Chen 0002, Shuo Ye, Wenjin Hou, Xinge You
ACM Multimedia3
2024 Rethinking attribute localization for zero-shot learning
Shuhuang Chen, Shiming Chen 0002, Guosen Xie, Xiangbo Shu, Xinge You, Xuelong Li 0001
Sci. China Inf. Sci.2
2024 EGANS: Evolutionary Generative Adversarial Network Search for Zero-Shot Learning
abstract
Zero-shot learning (ZSL) aims to recognize the novel classes which cannot be collected for training a prediction model. Accordingly, generative models (e.g., generative adversarial network (GAN)) are typically used to synthesize the visual samples conditioned by the class semantic vectors and achieve remarkable progress for ZSL. However, existing GAN-based generative ZSL methods are based on hand-crafted models, which cannot adapt to various datasets/scenarios and fails to model instability. To alleviate these challenges, we propose evolutionary generative adversarial network search (termed EGANS) to automatically design the generative network with good adaptation and stability, enabling reliable visual feature sample synthesis for advancing ZSL. Specifically, we adopt cooperative dual evolution to conduct a neural architecture search for both generator and discriminator under a unified evolutionary adversarial framework. EGANS is learned by two stages: evolution generator architecture search and evolution discriminator architecture search. During the evolution generator architecture search, we adopt a many-to-one adversarial training strategy to evolutionarily search for the optimal generator. Then the optimal generator is further applied to search for the optimal discriminator in the evolution discriminator architecture search with a similar evolution search algorithm. Once the optimal generator and discriminator are searched, we entail them into various generative ZSL baselines for ZSL classification. Extensive experiments show that EGANS consistently improve existing generative ZSL methods on the standard CUB, SUN, AWA2 and FLO datasets. The significant performance gains indicate that the evolutionary neural architecture search explores a virgin field in ZSL.
Shiming Chen 0002, Shuhuang Chen, Wenjin Hou, Weiping Ding 0001, Xinge You
IEEE Trans. Evol. Comput.1
2024 ECEA: Extensible Co-Existing Attention for Few-Shot Object Detection
abstract
Few-shot object detection (FSOD) identifies objects from extremely few annotated samples. Most existing FSOD methods, recently, apply the two-stage learning paradigm, which transfers the knowledge learned from abundant base classes to assist the few-shot detectors by learning the global features. However, such existing FSOD approaches seldom consider the localization of objects from local to global. Limited by the scarce training data in FSOD, the training samples of novel classes typically capture part of objects, resulting in such FSOD methods being unable to detect the completely unseen object during testing. To tackle this problem, we propose an Extensible Co-Existing Attention (ECEA) module to enable the model to infer the global object according to the local parts. Specifically, we first devise an extensible attention mechanism that starts with a local region and extends attention to co-existing regions that are similar and adjacent to the given local region. We then implement the extensible attention mechanism in different feature scales to progressively discover the full object in various receptive fields. In the training process, the model learns the extensible ability on the base stage with abundant samples and transfers it to the novel stage of continuous extensible learning, which can assist the few-shot model to quickly adapt in extending local regions to co-existing regions. Extensive experiments on the PASCAL VOC and COCO datasets show that our ECEA module can assist the few-shot detector to completely predict the object despite some regions failing to appear in the training samples and achieve the new state-of-the-art compared with existing FSOD methods. Code is released at https://github.com/zhimengXin/ECEA.
Zhimeng Xin, Tianxu Wu, Shiming Chen 0002, Yixiong Zou, Ling Shao 0001, Xinge You
IEEE Trans. Image Process.3
2024 GNDAN: Graph Navigated Dual Attention Network for Zero-Shot Learning
abstract
Zero-shot learning (ZSL) tackles the unseen class recognition problem by transferring semantic knowledge from seen classes to unseen ones. Typically, to guarantee desirable knowledge transfer, a direct embedding is adopted for associating the visual and semantic domains in ZSL. However, most existing ZSL methods focus on learning the embedding from implicit global features or image regions to the semantic space. Thus, they fail to: 1) exploit the appearance relationship priors between various local regions in a single image, which corresponds to the semantic information and 2) learn cooperative global and local features jointly for discriminative feature representations. In this article, we propose the novel graph navigated dual attention network (GNDAN) for ZSL to address these drawbacks. GNDAN employs a region-guided attention network (RAN) and a region-guided graph attention network (RGAT) to jointly learn a discriminative local embedding and incorporate global context for exploiting explicit global embeddings under the guidance of a graph. Specifically, RAN uses soft spatial attention to discover discriminative regions for generating local embeddings. Meanwhile, RGAT employs an attribute-based attention to obtain attribute-based region features, where each attribute focuses on the most relevant image regions. Motivated by the graph neural network (GNN), which is beneficial for structural relationship representations, RGAT further leverages a graph attention network to exploit the relationships between the attribute-based region features for explicit global embedding representations. Based on the self-calibration mechanism, the joint visual embedding learned is matched with the semantic embedding to form the final prediction. Extensive experiments on three benchmark datasets demonstrate that the proposed GNDAN achieves superior performances to the state-of-the-art methods. Our code and trained models are available at https://github.com/shiming-chen/GNDAN.
Shiming Chen 0002, Ziming Hong, Guosen Xie, Qinmu Peng, Xinge You, Weiping Ding 0001, Ling Shao 0001
IEEE Trans. Neural Networks Learn. Syst.1
2023 Both Diverse and Realism Matter: Physical Attribute and Style Alignment for Rainy Image Generation
abstract
Although considerable progress has been made in image deraining under synthetic data, real rain removal is still a tough problem due to the huge domain gap between synthetic and real data. Besides, difficulties in collecting and labeling diverse real rain images hinder the progress of this field. Consequently, we attempt to promote real rain removal from rain image generation (RIG) perspective. Existing RIG methods mainly focus on diversity but miss realistic, or the realistic but neglect diversity of the generation. To solve this dilemma, we propose a physical alignment and controllable generation network (PCGNet) for diverse and realistic rain generation. Our key idea is to simultaneously utilize the controllability of attributes from synthetic and the realism of appearance from real data. Specifically, we devise a unified framework to disentangle background, rain attributes, and appearance style from synthetic and real data. Then we collaboratively align the factors with a novel semi-supervised weight moving strategy for attribute, an explicit distribution modeling method for real rain style. Furthermore, we pack these aligned factors into the generation model, achieving physical controllable mapping from the attributes to real rain with image-level and attribute-level consistency loss. Extensive experiments show that PCGNet can effectively generate appealing rainy results, which significantly improve the performance under synthetic and real scenes for all existing deraining methods.
Changfeng Yu, Shiming Chen 0002, Yi Chang 0002, Yibing Song, Luxin Yan
ICCV2
2023 Evolving Semantic Prototype Improves Generative Zero-Shot Learning
abstract
In zero-shot learning (ZSL), generative methods synthesize class-related sample features based on predefined semantic prototypes. They advance the ZSL performance by synthesizing unseen class sample features for better training the classifier. We observe that each class’s predefined semantic prototype (also referred to as semantic embedding or condition) does not accurately match its real semantic prototype. So the synthesized visual sample features do not faithfully represent the real sample features, limiting the classifier training and existing ZSL performance. In this paper, we formulate this mismatch phenomenon as the visual-semantic domain shift problem. We propose a dynamic semantic prototype evolving (DSP) method to align the empirically predefined semantic prototypes and the real prototypes for class-related feature synthesis. The alignment is learned by refining sample features and semantic prototypes in a unified framework and making the synthesized visual sample features approach real sample features. After alignment, synthesized sample features from unseen classes are closer to the real sample features and benefit DSP to improve existing generative ZSL methods by 8.5%, 8.0%, and 9.7% on the standard CUB, SUN AWA2 datasets, the significant performance improvement indicates that evolving semantic prototype explores a virgin field in ZSL.
Shiming Chen 0002, Wenjin Hou, Ziming Hong, Xiaohan Ding, Yibing Song, Xinge You, Tongliang Liu, Kun Zhang 0001
ICML1
2023 TransZero++: Cross Attribute-Guided Transformer for Zero-Shot Learning
abstract
Zero-shot learning (ZSL) tackles the novel class recognition problem by transferring semantic knowledge from seen classes to unseen ones. Semantic knowledge is typically represented by attribute descriptions shared between different classes, which act as strong priors for localizing object attributes that represent discriminative region features, enabling significant and sufficient visual-semantic interaction for advancing ZSL. Existing attention-based models have struggled to learn inferior region features in a single image by solely using unidirectional attention, which ignore the transferable and discriminative attribute localization of visual features for representing the key semantic knowledge for effective knowledge transfer in ZSL. In this paper, we propose a cross attribute-guided Transformer network, termed TransZero++, to refine visual features and learn accurate attribute localization for key semantic knowledge representations in ZSL. Specifically, TransZero++ employs an attribute → visual Transformer sub-net (AVT) and a visual → attribute Transformer sub-net (VAT) to learn attribute-based visual features and visual-based attribute features, respectively. By further introducing feature-level and prediction-level semantical collaborative losses, the two attribute-guided transformers teach each other to learn semantic-augmented visual embeddings for key semantic knowledge representations via semantical collaborative learning. Finally, the semantic-augmented visual embeddings learned by AVT and VAT are fused to conduct desirable visual-semantic interaction cooperated with class semantic vectors for ZSL classification. Extensive experiments show that TransZero++ achieves the new state-of-the-art results on three golden ZSL benchmarks and on the large-scale ImageNet dataset. The project website is available at: https://shiming-chen.github.io/TransZero-pp/TransZero-pp.html.
Shiming Chen 0002, Ziming Hong, Wenjin Hou, Guosen Xie, Yibing Song, Jian Zhao 0006, Xinge You, Shuicheng Yan, Ling Shao 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2023 Kernelized Similarity Learning and Embedding for Dynamic Texture Synthesis
abstract
Dynamic texture (DT) exhibits statistical stationarity in the spatial domain and stochastic repetitiveness in the temporal dimension, indicating that different frames of DT possess a high similarity correlation that is critical prior knowledge. However, existing methods cannot effectively learn a synthesis model for high-dimensional DT from a small number of training samples. In this article, we propose a novel DT synthesis method, which makes full use of similarity as prior knowledge to address this issue. Our method is based on the proposed kernel similarity embedding, which can not only mitigate the high dimensionality and small sample issues, but also has the advantage of modeling nonlinear feature relationships. Specifically, we first put forward two hypotheses that are essential for the DT model to generate new frames using similarity correlations. Then, we integrate kernel learning and the extreme learning machine into a unified synthesis model to learn kernel similarity embeddings for representing DTs. Extensive experiments on DT videos collected from the Internet and two benchmark datasets, i.e., Gatech Graphcut Textures and Dyntex, demonstrate that the learned kernel similarity embeddings can provide discriminative representations for DTs. Further, our method can preserve the long-term temporal continuity of the synthesized DT sequences with excellent sustainability and generalization. Meanwhile, it effectively generates realistic DT videos with higher speed and lower computation than the current state-of-the-art methods. The code and more synthesis videos are available at our project pagehttps://shiming-chen.github.io/Similarity-page/Similarit.html.
Shiming Chen 0002, Peng Zhang 0040, Guosen Xie, Qinmu Peng, Zehong Cao, Wei Yuan 0001, Xinge You
IEEE Trans. Syst. Man Cybern. Syst.1
2022 TransZero: Attribute-Guided Transformer for Zero-Shot Learning
abstract
Zero-shot learning (ZSL) aims to recognize novel classes by transferring semantic knowledge from seen classes to unseen ones. Semantic knowledge is learned from attribute descriptions shared between different classes, which are strong prior for localization of object attribute for representing discriminative region features enabling significant visual-semantic interaction. Although few attention-based models have attempted to learn such region features in a single image, the transferability and discriminative attribute localization of visual features are typically neglected. In this paper, we propose an attribute-guided Transformer network to learn the attribute localization for discriminative visual-semantic embedding representations in ZSL, termed TransZero. Specifically, TransZero takes a feature augmentation encoder to alleviate the cross-dataset bias between ImageNet and ZSL benchmarks and improve the transferability of visual features by reducing the entangled relative geometry relationships among region features. To learn locality-augmented visual features, TransZero employs a visual-semantic decoder to localize the most relevant image regions to each attributes from a given image under the guidance of attribute semantic information. Then, the locality-augmented visual features and semantic vectors are used for conducting effective visual-semantic interaction in a visual-semantic embedding network. Extensive experiments show that TransZero achieves a new state-of-the-art on three ZSL benchmarks. The codes are available at: https://github.com/shiming-chen/TransZero.
Shiming Chen 0002, Ziming Hong, Yang Liu 0069, Guosen Xie, Baigui Sun, Hao Li 0030, Qinmu Peng, Ke Lu 0002, Xinge You
AAAI1
2022 MSDN: Mutually Semantic Distillation Network for Zero-Shot Learning
abstract
The key challenge of zero-shot learning (ZSL) is how to infer the latent semantic knowledge between visual and attribute features on seen classes, and thus achieving a desirable knowledge transfer to unseen classes. Prior works either simply align the global features of an image with its associated class semantic vector or utilize unidirectional attention to learn the limited latent semantic representations, which could not effectively discover the intrinsic semantic knowledge (e.g., attribute semantics) between visual and attribute features. To solve the above dilemma, we propose a Mutually Semantic Distillation Network (MSDN), which progressively distills the intrinsic semantic representations between visual and attribute features for ZSL. MSDN incorporates an attribute→visual attention sub-net that learns attribute-based visual features, and a visual→attribute attention sub-net that learns visual-based attribute features. By further introducing a semantic distillation loss, the two mutual attention sub-nets are capable of learning collaboratively and teaching each other throughout the training process. The proposed MSDN yields significant improvements over the strong baselines, leading to new state-of-the-art performances on three popular challenging benchmarks. Our codes have been available at: https://github.com/shiming-chen/MSDN.
Shiming Chen 0002, Ziming Hong, Guosen Xie, Wenhan Yang, Qinmu Peng, Kai Wang 0036, Jian Zhao 0006, Xinge You
CVPR1
2022 Semantic Compression Embedding for Generative Zero-Shot Learning
abstract
Generative methods have been successfully applied in zero-shot learning (ZSL) by learning an implicit mapping to alleviate the visual-semantic domain gaps and synthesizing unseen samples to handle the data imbalance between seen and unseen classes. However, existing generative methods simply use visual features extracted by the pre-trained CNN backbone. These visual features lack attribute-level semantic information. Consequently, seen classes are indistinguishable, and the knowledge transfer from seen to unseen classes is limited. To tackle this issue, we propose a novel Semantic Compression Embedding Guided Generation (SC-EGG) model, which cascades a semantic compression embedding network (SCEN) and an embedding guided generative network (EGGN). The SCEN extracts a group of attribute-level local features for each sample and further compresses them into the new low-dimension visual feature. Thus, a dense-semantic visual space is obtained. The EGGN learns a mapping from the class-level semantic space to the dense-semantic visual space, thus improving the discriminability of the synthesized dense-semantic unseen visual features. Extensive experiments on three benchmark datasets, i.e., CUB, SUN and AWA2, demonstrate the significant performance gains of SC-EGG over current state-of-the-art methods and its baselines.
Ziming Hong, Shiming Chen 0002, Guosen Xie, Wenhan Yang, Jian Zhao 0006, Yuanjie Shao, Qinmu Peng, Xinge You
IJCAI2
2021 FREE: Feature Refinement for Generalized Zero-Shot Learning
abstract
Generalized zero-shot learning (GZSL) has achieved significant progress, with many efforts dedicated to over-coming the problems of visual-semantic domain gap and seen-unseen bias. However, most existing methods directly use feature extraction models trained on ImageNet alone, ignoring the cross-dataset bias between ImageNet and GZSL benchmarks. Such a bias inevitably results in poor-quality visual features for GZSL tasks, which potentially limits the recognition performance on both seen and unseen classes. In this paper, we propose a simple yet effective GZSL method, termed feature refinement for generalized zero-shot learning (FREE), to tackle the above problem. FREE employs a feature refinement (FR) module that in-corporates semantic→visual mapping into a unified generative model to refine the visual features of seen and unseen class samples. Furthermore, we propose a self-adaptive margin center loss (SAMC-loss) that cooperates with a semantic cycle-consistency loss to guide FR to learn class- and semantically-relevant representations, and concatenate the features in FR to extract the fully refined features. Extensive experiments on five benchmark datasets demonstrate the significant performance gain of FREE over its baseline and current state-of-the-art methods. The code is available at https://github.com/shiming-chen/FREE.
Shiming Chen 0002, Beihao Xia, Qinmu Peng, Xinge You, Feng Zheng 0001, Ling Shao 0001
ICCV1
2021 Norm-guided Adaptive Visual Embedding for Zero-Shot Sketch-Based Image Retrieval
abstract
Zero-shot sketch-based image retrieval (ZS-SBIR), which aims to retrieve photos with sketches under the zero-shot scenario, has shown extraordinary talents in real-world applications. Most existing methods leverage language models to generate class-prototypes and use them to arrange the locations of all categories in the common space for photos and sketches. Although great progress has been made, few of them consider whether such pre-defined prototypes are necessary for ZS-SBIR, where locations of unseen class samples in the embedding space are actually determined by visual appearance and a visual embedding actually performs better. To this end, we propose a novel Norm-guided Adaptive Visual Embedding (NAVE) model, for adaptively building the common space based on visual similarity instead of language-based pre-defined prototypes. To further enhance the representation quality of unseen classes for both photo and sketch modality, modality norm discrepancy and noisy label regularizer are jointly employed to measure and repair the modality bias of the learned common embedding. Experiments on two challenging datasets demonstrate the superiority of our NAVE over state-of-the-art competitors.
Yufeng Shi 0003, Shiming Chen 0002, Qinmu Peng, Feng Zheng 0001, Xinge You
IJCAI3
2021 HSVA: Hierarchical Semantic-Visual Adaptation for Zero-Shot Learning
abstract
Zero-shot learning (ZSL) tackles the unseen class recognition problem, transferring semantic knowledge from seen classes to unseen ones. Typically, to guarantee desirable knowledge transfer, a common (latent) space is adopted for associating the visual and semantic domains in ZSL. However, existing common space learning methods align the semantic and visual domains by merely mitigating distribution disagreement through one-step adaptation. This strategy is usually ineffective due to the heterogeneous nature of the feature representations in the two domains, which intrinsically contain both distribution and structure variations. To address this and advance ZSL, we propose a novel hierarchical semantic-visual adaptation (HSVA) framework. Specifically, HSVA aligns the semantic and visual domains by adopting a hierarchical two-step adaptation, i.e., structure adaptation and distribution adaptation. In the structure adaptation step, we take two task-specific encoders to encode the source data (visual domain) and the target data (semantic domain) into a structure-aligned common space. To this end, a supervised adversarial discrepancy (SAD) module is proposed to adversarially minimize the discrepancy between the predictions of two task-specific classifiers, thus making the visual and semantic feature manifolds more closely aligned. In the distribution adaptation step, we directly minimize the Wasserstein distance between the latent multivariate Gaussian distributions to align the visual and semantic distributions using a common encoder. Finally, the structure and distribution adaptation are derived in a unified framework under two partially-aligned variational autoencoders. Extensive experiments on four benchmark datasets demonstrate that HSVA achieves superior performance on both conventional and generalized ZSL. The code is available at \url{https://github.com/shiming-chen/HSVA}.
Shiming Chen 0002, Guosen Xie, Yang Liu 0069, Qinmu Peng, Baigui Sun, Hao Li 0030, Xinge You, Ling Shao 0001
NeurIPS1
2021 CDE-GAN: Cooperative Dual Evolution-Based Generative Adversarial Network
abstract
Generative adversarial networks (GANs) have been a popular deep generative model for real-world applications. Despite many recent efforts on GANs that have been contributed, mode collapse and instability of GANs are still open problems caused by their adversarial optimization difficulties. In this article, motivated by the cooperative co-evolutionary algorithm, we propose a cooperative dual evolution-based GAN (CDE-GAN) to circumvent these drawbacks. In essence, CDE-GAN incorporates dual evolution with respect to the generator(s) and discriminators into a unified evolutionary adversarial framework to conduct effective adversarial multiobjective optimization. Thus, it exploits the complementary properties and injects dual mutation diversity into the training, to steadily diversify the estimated density in capturing multimodes and improve generative performance. Specifically, CDE-GAN decomposes the complex adversarial optimization problem into two subproblems (generation and discrimination), and each subproblem is solved with a separated subpopulation (E-GeneratorsandE-Discriminators), evolved by its own evolutionary algorithm. Additionally, we further propose aSoft Mechanismto balance the tradeoff between E-Generators and E-Discriminators to conduct steady training for CDE-GAN. Extensive experiments on one synthetic dataset and three real-world benchmark image datasets demonstrate that the proposed CDE-GAN achieves a competitive and superior performance in generating good quality and diverse samples over baselines. The code and more generated results are available at our project homepagehttps://shiming-chen.github.io/CDE-GAN-website/CDE-GAN.html.
Shiming Chen 0002, Beihao Xia, Xinge You, Qinmu Peng, Zehong Cao, Weiping Ding 0001
IEEE Trans. Evol. Comput.1
2019 Semi-supervised feature learning for improving writer identification
Shiming Chen 0002, Yisong Wang 0004, Chin-Teng Lin, Weiping Ding 0001, Zehong Cao
Inf. Sci.1