Yunlong Yu 0001

dblp:45/7404-1 · DBLP profile ↗
← Back
51ranked-venue papers
10as first author
37since 2021 · last 2026
0000-0002-0294-2099ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 38 · 9 first-author · 26 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 3 first-author · 17 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Chat-MedGen: omni-adaptation of multi-modal large language models for diverse biomedical tasks
Qinyue Tong, Ziqian Lu, Zheming Lu 0001, Yunlong Yu 0001, Yang-Ming Zheng
Knowl. Based Syst.4
2026 Improving anomaly detection with foundation-model synthesis and wavelet-domain attention
Wensheng Wu, Zheming Lu 0001, Ziqian Lu, Zewei He, Xuecheng Sun, Jungong Han, Yunlong Yu 0001
Neural Networks8
2026 Parameter-Efficient Fine-Tuning for Continual Learning: A Neural Tangent Kernel Perspective
abstract
Parameter-efficient fine-tuning for continual learning (PEFT-CL) has shown promise in adapting pre-trained models to sequential tasks while mitigating catastrophic forgetting problem. However, understanding the mechanisms that dictate continual performance in this paradigm remains elusive. To unravel this mystery, we undertake a rigorous analysis of PEFT-CL dynamics to derive relevant metrics for continual scenarios using Neural Tangent Kernel (NTK) theory. With the aid of NTK as a mathematical analysis tool, we recast the challenge of test-time forgetting into the quantifiable generalization gaps during training, identifying three key factors that influence these gaps and the performance of PEFT-CL: training sample size, task-level feature orthogonality, and regularization. To address these challenges, we introduce NTK-CL, a novel framework that eliminates task-specific parameter storage while adaptively generating task-relevant features. Aligning with theoretical guidance, NTK-CL triples the feature representation of each sample, theoretically and empirically reducing the magnitude of both task-interplay and task-specific generalization gaps. Grounded in NTK analysis, our framework imposes an adaptive exponential moving average mechanism and constraints on task-level feature orthogonality, maintaining intra-task NTK forms while attenuating inter-task NTK forms. Ultimately, by fine-tuning optimizable parameters with appropriate regularization, NTK-CL achieves state-of-the-art performance on established PEFT-CL benchmarks. This work provides a theoretical foundation for understanding and improving PEFT-CL models, offering insights into the interplay between feature representation, task orthogonality, and generalization, contributing to the development of more efficient continual learning systems.
Jingren Liu, Zhong Ji, Yunlong Yu 0001, Jiale Cao, Yanwei Pang, Jungong Han, Xuelong Li 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 Mask-free Iterative Refinement Network for weakly-supervised Few-shot Semantic Segmentation
Shanjuan Chen, Yunlong Yu 0001, Yingming Li, Ziqian Lu
Neurocomputing2
2025 Hybrid mask generation for infrared small target detection with single-point supervision
Mushui Liu, Yunlong Yu 0001
Neurocomputing3
2025 Synth-CLIP: Synthetic data make CLIP generalize better in data-limited scenarios
Mushui Liu, Ziqian Lu, Jun Dan, Yunlong Yu 0001, Yingming Li, Xi Li 0001, Jungong Han
Neural Networks5
2025 Variational Adapter: Improving CLIP in Data-Imbalanced Scenarios
abstract
In this paper, we propose the Prompt-based Variational Adapter (PVA), a novel approach designed to fine-tune the pre-trained Vision-Language Models (VLMs) in data-imbalanced scenarios. Unlike existing methods that focus primarily on pairwise alignment of visual-text relationships during fine-tuning, PVA relaxes pairwise explicit constrains and emphasizes the harmonization of visual and text modality distributions, enhancing generalization and cross-modal understanding. To realize this harmonization, we develop two variational adapters, which are appended separately to the visual and text encoders. These adapters transform the feature embeddings into latent spaces that implicitly align with the corresponding modality distributions. We then adopt a divide-and-conquer strategy, dividing classes into data-abundant and data-limited sets to reduce prediction bias. Within each set, we independently fine-tune the models by incorporating both the model’s original general knowledge and specialized knowledge gained from training samples. Extensive experiments across two data-imbalanced scenarios validate the superiority of our approach, establishing a new state-of-the-art on popular benchmarks.
Ziqian Lu, Mushui Liu, Yunlong Yu 0001, Xi Li 0001, Jungong Han
IEEE Trans. Circuits Syst. Video Technol.3
2025 Frequency-Spatial Complementation: Unified Channel-Specific Style Attack for Cross-Domain Few-Shot Learning
abstract
Cross-Domain Few-Shot Learning (CD-FSL) addresses the challenges of recognizing targets with out-of-domain data when only a few instances are available. Many current CD-FSL approaches primarily focus on enhancing the generalization capabilities of models in spatial domain, which neglects the role of the frequency domain in domain generalization. To take advantage of frequency domain in processing global information, we propose a Frequency-Spatial Complementation (FSC) model, which combines frequency domain information with spatial domain information to learn domain-invariant information from attacked data style. Specifically, we design a Frequency and Spatial Fusion (FusionFS) module to enhance the ability of the model to capture style-related information. Besides, we propose two attack strategies, i.e., the Gradient-guided Unified Style Attack (GUSA) strategy and the Channel-specific Attack Intensity Calculation (CAIC) strategy, which conduct targeted attacks on different channels to provide more diversified style data during the training phase, especially in single-source domain scenarios where the source domain data style is homogeneous. Extensive experiments across eight target domains demonstrate that our method significantly improves the model's performance under various styles.
Zhong Ji, Zhilong Wang 0001, Xiyao Liu 0002, Yunlong Yu 0001, Yanwei Pang, Jungong Han
IEEE Trans. Image Process.4
2025 Relationship-Incremental Scene Graph Generation by a Divide-and-Conquer Pipeline With Feature Adapter
abstract
As a challenging computer vision task, Scene Graph Generation (SGG) finds the latent semantic relationships among objects from a given image, which may be limited by the datasets and real-world scenarios. In this paper, we consider a novel incremental learning task called Relationship-Incremental Scene Graph Generation (RISGG) that learns the semantic relationships among objects in an incremental way. Compared with classic Class-Incremental Learning (CIL) problem, RISGG suffers from its special issues: 1) Old class shift - the relationship-labeled object pair may have different labels during different learning sessions; 2) Background shift - the relationship-unlabeled object pair may not be a real unlabeled one. In this work, we address the above issues from the following aspects. First, we present a Divide-and-Conquer (DaC) pipeline to deal with the old class shift via decoupling the recognition of relationship classes and recognizing relationships individually. In this way, label confusion and interaction among different relationships are eliminated during training. Second, we propose a Feature Adapter (FA) to bridge the feature space gap between the current session and the previous one and use our extra supervision to mine old relationship information in the current session. Our proposed network combined DaC and FA, abbreviated DaCFA-Net, for RISGG. Experimental results on the benchmark dataset demonstrate the significant performance gain of DaCFA-Net in RISGG. It gains about 20% improvement against the SGG baselines on the popular VG dataset.
Xuewei Li 0003, Guangcong Zheng, Yunlong Yu 0001, Naye Ji, Xi Li 0001
IEEE Trans. Image Process.3
2025 Balancing Feature Alignment and Uniformity for Few-Shot Classification
abstract
In Few-Shot Learning (FSL), the objective is to correctly recognize new samples from novel classes with only a few available samples per class. Existing methods in FSL primarily focus on learning transferable knowledge from base classes by maximizing the information between feature representations and their corresponding labels. However, this approach may suffer from the "supervision collapse" issue, which arises due to a bias towards the base classes. In this paper, we propose a solution to address this issue by preserving the intrinsic structure of the data and enabling the learning of a generalized model for the novel classes. Following the InfoMax principle, our approach maximizes two types of mutual information (MI): between the samples and their feature representations, and between the feature representations and their class labels. This allows us to strike a balance between discrimination (capturing class-specific information) and generalization (capturing common characteristics across different classes) in the feature representations. To achieve this, we adopt a unified framework that perturbs the feature embedding space using two low-bias estimators. The first estimator maximizes the MI between a pair of intra-class samples, while the second estimator maximizes the MI between a sample and its augmented views. This framework effectively combines knowledge distillation between class-wise pairs and enlarges the diversity in feature representations. By conducting extensive experiments on popular FSL benchmarks, our proposed approach achieves comparable performances with state-of-the-art competitors. For example, we achieved an accuracy of 69.53% on the miniImageNet dataset and 77.06% on the CIFAR-FS dataset for the 5-way 1-shot task.
Yunlong Yu 0001, Dingyi Zhang, Zhong Ji, Xi Li 0001, Jungong Han, Zhongfei Zhang
IEEE Trans. Image Process.1
2024 Improving Zero-Shot Generalization for CLIP with Variational Adapter
Ziqian Lu, Mushui Liu, Yunlong Yu 0001, Xi Li 0001
ECCV (20)4
2024 RCS-Prompt: Learning Prompt to Rearrange Class Space for Prompt-Based Continual Learning
Longrong Yang, Hanbin Zhao, Yunlong Yu 0001, Xiaodong Zeng, Xi Li 0001
ECCV (47)3
2024 CoMoFusion: Fast and High-Quality Fusion of Infrared and Visible Image with Consistency Model
Zhiming Meng, Hui Li 0037, Zeyang Zhang 0002, Yunlong Yu 0001, Xiaoning Song, Xiaojun Wu 0001
PRCV (8)5
2024 Pixel Matching Network for Cross-Domain Few-Shot Segmentation
abstract
Few-Shot Segmentation (FSS) aims to segment the novel class images with a few annotated samples. In the past, numerous studies have concentrated on cross-category tasks, where the training and testing sets are derived from the same dataset, while these methods face significant difficulties in domain-shift scenarios. To better tackle the cross-domain tasks, we propose a pixel matching network (PMNet) to extract the domain-agnostic pixel-level affinity matching with a frozen backbone and capture both the pixel-to-pixel and pixel-to-patch relations in each support-query pair with the bidirectional 3D convolutions. Different from the existing methods that remove the support background, we design a hysteretic spatial filtering module (HSFM) to filter the background-related query features and retain the foreground-related query features with the assistance of the support background, which is beneficial for eliminating interference objects in the query background. We comprehensively evaluate our PMNet on ten benchmarks under cross-category, cross-dataset, and cross-domain FSS tasks. Experimental results demonstrate that PMNet performs very competitively under different settings with only 0.68M parameters, especially under cross-domain FSS tasks, showing its effectiveness and efficiency. Code will be released at: https://github.com/chenhao-zju/PMNet
Hao Chen 0107, Yonghan Dong, Zheming Lu 0001, Yunlong Yu 0001, Jungong Han
WACV4
2024 Dense affinity matching for Few-Shot Segmentation
Hao Chen 0107, Yonghan Dong, Zheming Lu 0001, Yunlong Yu 0001, Yingming Li, Jungong Han, Zhongfei Zhang
Neurocomputing4
2024 A cognition-driven framework for few-shot class-incremental learning
Xuan Wang 0016, Zhong Ji, Yanwei Pang, Yunlong Yu 0001
Neurocomputing4
2024 Learning Multiple Criteria Calibration for Generalized Zero-shot Learning
Ziqian Lu, Zheming Lu 0001, Yunlong Yu 0001, Zewei He, Hao Luo 0001, Yangming Zheng
Knowl. Based Syst.3
2024 Diversity-Infused Network for Unsupervised Few-Shot Remote Sensing Scene Classification
abstract
Few-shot Remote Sensing Scene Classification (RSSC) confronts challenges due to its dependence on extensive labeled datasets. Addressing this, we propose the Diversity-Infused Network (DIN), an unsupervised paradigm for few-shot RSSC, utilizing unlabeled data in training and adapting to novel classes with limited labeled samples. Within an augmentation-based framework, DIN includes a Random Augmentation Sampling (RAS) strategy for task diversity in the meta-training stage, and a Channel-Driven Metric Learning (CDML) module to decode complex channel interactions, enhancing information diversity. Additionally, DIN presents a Multi-Mutual Information (Multi-MI) objective function to balance the architecture and reduce the unreliability and potential biases from pseudo-training. Experimental results demonstrate that DIN surpasses other unsupervised approaches by over 11% in 1-shot and nearly 9% in 5-shot settings on WHU-RS19, and closely approaches supervised methods, with less than 1% difference in both settings.
Liyuan Hou 0002, Zhong Ji, Xuan Wang 0016, Yunlong Yu 0001, Yanwei Pang
IEEE Geosci. Remote. Sens. Lett.4
2024 Tolerant Self-Distillation for image classification
Mushui Liu, Yunlong Yu 0001, Zhong Ji, Jungong Han, Zhongfei Zhang
Neural Networks2
2024 Self-Prompting Perceptual Edge Learning for Dense Prediction
abstract
Numerous studies have employed prompt learning structures to enhance dense prediction tasks by integrating additional semantic or geometric information. While the inclusion of extra information has shown improvements in performance, it also poses challenges for applications that cannot provide extra input. To address this issue, this study evaluates the performance of different prompts and introduces an additional-input-free method, called self-prompting perceptual edge learning (SPPEL), which extracts edge-embedded semantic prompts directly from the image feature itself using trainable handcrafted edge operators within a plug-and-play module. To obtain the edge features, our approach incorporates an adversarial structure that compares the similarity between two edge features generated by the Hog and Kirsch operators, where the edge features are measured using multiplication, finetuned through a trainable all-one embedding, and enhanced with channel-to-channel attention. We conduct extensive evaluations of SPPEL on 7 tasks, utilizing 7 different backbones and applying 5 distinct methods. Our experimental results demonstrate that SPPEL achieves strong competitiveness in various settings with an average improvement of 1.7% across all 7 tasks, including ADE20K, COCO (Instance Segmentation), COCO (Object Detection), Pascal VOC2012, STARE, CHASE DB1, and HRF, while incurring a parameter increase of less than 3% (the detailed computation analysis of parameters and Gflops are shown in different experimental tables). Code will be released at: https://github.com/chenhao-zju/sppel.
Hao Chen 0107, Yonghan Dong, Zheming Lu 0001, Yunlong Yu 0001, Jungong Han
IEEE Trans. Circuits Syst. Video Technol.4
2024 NTK-Guided Few-Shot Class Incremental Learning
abstract
The proliferation of Few-Shot Class Incremental Learning (FSCIL) methodologies has highlighted the critical challenge of maintaining robust anti-amnesia capabilities in FSCIL learners. In this paper, we present a novel conceptualization of anti-amnesia in terms of mathematical generalization, leveraging the Neural Tangent Kernel (NTK) perspective. Our method focuses on two key aspects: ensuring optimal NTK convergence and minimizing NTK-related generalization loss, which serve as the theoretical foundation for cross-task generalization. To achieve global NTK convergence, we introduce a principled meta-learning mechanism that guides optimization within an expanded network architecture. Concurrently, to reduce the NTK-related generalization loss, we systematically optimize its constituent factors. Specifically, we initiate self-supervised pre-training on the base session to enhance NTK-related generalization potential. These self-supervised weights are then carefully refined through curricular alignment, followed by the application of dual NTK regularization tailored specifically for both convolutional and linear layers. Through the combined effects of these measures, our network acquires robust NTK properties, ensuring optimal convergence and stability of the NTK matrix and minimizing the NTK-related generalization loss, significantly enhancing its theoretical generalization. On popular FSCIL benchmark datasets, our NTK-FSCIL surpasses contemporary state-of-the-art approaches, elevating end-session accuracy by 2.9% to 9.3%.
Jingren Liu, Zhong Ji, Yanwei Pang, Yunlong Yu 0001
IEEE Trans. Image Process.4
2024 Model Attention Expansion for Few-Shot Class-Incremental Learning
abstract
Few-Shot Class-Incremental Learning (FSCIL) aims at incrementally learning new knowledge from limited training examples without forgetting previous knowledge. However, we observe that existing methods face a challenge known as supervision collapse, where the model disproportionately emphasizes class-specific features of base classes at the detriment of novel class representations, leading to restricted cognitive capabilities. To alleviate this issue, we propose a new framework, Model aTtention Expansion for Few-Shot Class-Incremental Learning (MTE-FSCIL), aimed at expanding the model attention fields to improve transferability without compromising the discriminative capability for base classes. Specifically, the framework adopts a dual-stage training strategy, comprising pre-training and meta-training stages. In the pre-training stage, we present a new regularization technique, named the Reserver (RS) loss, to expand the global perception and reduce over-reliance on class-specific features by amplifying feature map activations. During the meta-training stage, we introduce the Repeller (RP) loss, a novel pair-based loss that promotes variation in representations and improves the model's recognition of sample uniqueness by scattering intra-class samples within the embedding space. Furthermore, we propose a Transformational Adaptation (TA) strategy to enable continuous incorporation of new knowledge from downstream tasks, thus facilitating cross-task knowledge transfer. Extensive experimental results on mini-ImageNet, CIFAR100, and CUB200 datasets demonstrate that our proposed framework consistently outperforms the state-of-the-art methods.
Xuan Wang 0016, Zhong Ji, Yunlong Yu 0001, Yanwei Pang, Jungong Han
IEEE Trans. Image Process.3
2024 Multi-Content Interaction Network for Few-Shot Segmentation
abstract
Few-Shot Segmentation (FSS) poses significant challenges due to limited support images and large intra-class appearance discrepancies. Most existing approaches focus on aligning the support-query correlations from the same layer of the frozen backbone while neglecting the bias between different tasks and different layers. In this article, we propose a Multi-Content Interaction Network (MCINet) to remedy these issues by fully exploiting and interacting with the different contextual information contained in distinct branches. Specifically, MCINet improves FSS from three perspectives: (1) boosting the query representations through incorporating the independent information from another learnable branch into the features from the frozen backbone, (2) enhancing the support-query correlations by exploiting both the same-layer and adjacent-layer features, and (3) refining the predicted results with a multi-scale mask prediction strategy. Experiments on three benchmarks demonstrate that our approach reaches state-of-the-art performances and outperforms the best competitors with many desirable advantages, especially on the challenging COCO dataset. Code will be released on GitHub ( https://github.com/chenhao-zju/mcinet ).
Hao Chen 0107, Yunlong Yu 0001, Yonghan Dong, Zheming Lu 0001, Yingming Li, Zhongfei Zhang
ACM Trans. Multim. Comput. Commun. Appl.2
2023 GaitGCI: Generative Counterfactual Intervention for Gait Recognition
abstract
Gait is one of the most promising biometrics that aims to identify pedestrians from their walking patterns. However, prevailing methods are susceptible to confounders, resulting in the networks hardly focusing on the regions that re-flect effective walking patterns. To address this fundamen-tal problem in gait recognition, we propose a Generative Counterfactual Intervention framework, dubbed GaitGCI, consisting of Counterfactual Intervention Learning (CIL) and Diversity-Constrained Dynamic Convolution (DCDC). CIL eliminates the impacts of confounders by maximizing the likelihood difference between factual/counterfactual attention while DCDC adaptively generates sample-wise factual/counterfactual attention to efficiently perceive the sample-wise properties. With matrix decomposition and diversity constraint, DCDC guarantees the model to be efficient and effective. Extensive experiments indicate that proposed GaitGCI.· 1) could effectively focus on the discrimi-native and interpretable regions that reflect gait pattern; 2) is model-agnostic and could be plugged into existing models to improve performance with nearly no extra cost; 3) efficiently achieves state-of-the-art performance on arbitrary scenarios (in-the-lab and in-the-wild).
Huanzhang Dou, Wei Su 0009, Yunlong Yu 0001, Yining Lin, Xi Li 0001
CVPR4
2023 DenseDINO: Boosting Dense Self-Supervised Learning with Token-Based Point-Level Consistency
abstract
In this paper, we propose a simple yet effective transformer framework for self-supervised learning called DenseDINO to learn dense visual representations. To exploit the spatial information that the dense prediction tasks require but neglected by the existing self-supervised transformers, we introduce point-level supervision across views in a novel token-based way. Specifically, DenseDINO introduces some extra input tokens called reference tokens to match the point-level features with the position prior. With the reference token, the model could maintain spatial consistency and deal with multi-object complex scene images, thus generalizing better on dense prediction tasks. Compared with the vanilla DINO, our approach obtains competitive performance when evaluated on classification in ImageNet and achieves a large margin (+7.2% mIoU) improvement in semantic segmentation on PascalVOC under the linear probing protocol for segmentation.
Yike Yuan, Xinghe Fu, Yunlong Yu 0001, Xi Li 0001
IJCAI3
2023 Zero-shot classification with unseen prototype learning
Zhong Ji, Biying Cui, Yunlong Yu 0001, Yanwei Pang, Zhongfei Zhang
Neural Comput. Appl.3
2023 Uncertainty-Aware Scene Graph Generation
Xuewei Li 0003, Guangcong Zheng, Yunlong Yu 0001, Xi Li 0001
Pattern Recognit. Lett.4
2023 Knowledge Distillation Classifier Generation Network for Zero-Shot Learning
abstract
In this article, we present a conceptually simple but effective framework called knowledge distillation classifier generation network (KDCGN) for zero-shot learning (ZSL), where the learning agent requires recognizing unseen classes that have no visual data for training. Different from the existing generative approaches that synthesize visual features for unseen classifiers' learning, the proposed framework directly generates classifiers for unseen classes conditioned on the corresponding class-level semantics. To ensure the generated classifiers to be discriminative to the visual features, we borrow the knowledge distillation idea to both supervise the classifier generation and distill the knowledge with, respectively, the visual classifiers and soft targets trained from a traditional classification network. Under this framework, we develop two, respectively, strategies, i.e., class augmentation and semantics guidance, to facilitate the supervision process from the perspectives of improving visual classifiers. Specifically, the class augmentation strategy incorporates some additional categories to train the visual classifiers, which regularizes the visual classifier weights to be compact, under supervision of which the generated classifiers will be more discriminative. The semantics-guidance strategy encodes the class semantics into the visual classifiers, which would facilitate the supervision process by minimizing the differences between the generated and the real-visual classifiers. To evaluate the effectiveness of the proposed framework, we have conducted extensive experiments on five datasets in image classification, i.e., AwA1, AwA2, CUB, FLO, and APY. Experimental results show that the proposed approach performs best in the traditional ZSL task and achieves a significant performance improvement on four out of the five datasets in the generalized ZSL task.
Yunlong Yu 0001, Bin Li 0038, Zhong Ji, Jungong Han, Zhongfei Zhang
IEEE Trans. Neural Networks Learn. Syst.1
2022 MetaGait: Learning to Learn an Omni Sample Adaptive Representation for Gait Recognition
Huanzhang Dou, Wei Su 0009, Yunlong Yu 0001, Xi Li 0001
ECCV (5)4
2022 Adaptive Cross-domain Learning for Generalizable Person Re-identification
Huanzhang Dou, Yunlong Yu 0001, Xi Li 0001
ECCV (14)3
2022 Multi-Proxy Learning from an Entropy Optimization Perspective
abstract
Deep Metric Learning, a task that learns a feature embedding space where semantically similar samples are located closer than dissimilar samples, is a cornerstone of many computer vision applications. Most of the existing proxy-based approaches usually exploit the global context via learning a single proxy for each training class, which struggles in capturing the complex non-uniform data distribution with different patterns. In this work, we present an easy-to-implement framework to effectively capture the local neighbor relationships via learning multiple proxies for each class that collectively approximate the intra-class distribution. In the context of large intra-class visual diversity, we revisit the entropy learning under the multi-proxy learning framework and provide a training routine that both minimizes the entropy of intra-class probability distribution and maximizes the entropy of inter-class probability distribution. In this way, our model is able to better capture the intra-class variations and smooth the inter-class differences and thus facilitates to extract more semantic feature representations for the downstream tasks. Extensive experimental results demonstrate that the proposed approach achieves competitive performances. Codes and an appendix are provided.
Yunlong Yu 0001, Dingyi Zhang, Yingming Li, Zhongfei Zhang
IJCAI1
2022 Training a Lightweight ViT Network for Image Retrieval
Yunlong Yu 0001, Yingming Li, Zhongfei Zhang
PRICAI (3)2
2022 Learn more from less: Generalized zero-shot learning with severely limited labeled data
Ziqian Lu, Zheming Lu 0001, Yunlong Yu 0001, Zonghui Wang
Neurocomputing3
2022 Local spatial alignment network for few-shot learning
Yunlong Yu 0001, Dingyi Zhang, Sidi Wang, Zhong Ji, Zhongfei Zhang
Neurocomputing1
2022 Reparameterized attention for convolutional neural networks
Yiming Wu 0006, Yunlong Yu 0001, Xi Li 0001
Pattern Recognit. Lett.3
2022 Semantic-Guided Class-Imbalance Learning Model for Zero-Shot Image Classification
abstract
In this article, we focus on the task of zero-shot image classification (ZSIC) that equips a learning system with the ability to recognize visual images from unseen classes. In contrast to the traditional image classification, ZSIC more easily suffers from the class-imbalance issue since it is more concerned with the class-level knowledge transferring capability. In the real world, the sample numbers of different categories generally follow a long-tailed distribution, and the discriminative information in the sample-scarce seen classes is hard to transfer to the related unseen classes in the traditional batch-based training manner, which degrades the overall generalization ability a lot. To alleviate the class-imbalance issue in ZSIC, we propose a sample-balanced training process to encourage all training classes to contribute equally to the learned model. Specifically, we randomly select the same number of images from each class across all training classes to form a training batch to ensure that the sample-scarce classes contribute equally as those classes with sufficient samples during each iteration. Considering that the instances from the same class differ in class representativeness, we further develop an efficient semantic-guided feature fusion model to obtain the discriminative class visual prototype for the following visual-semantic interaction process via distributing different weights to the selected samples based on their class representativeness. Extensive experiments on three imbalanced ZSIC benchmark datasets for both traditional ZSIC and generalized ZSIC tasks demonstrate that our approach achieves promising results, especially for the unseen categories that are closely related to the sample-scarce seen categories. Besides, the experimental results on two class-balanced datasets show that the proposed approach also improves the classification performance against the baseline model.
Zhong Ji, Xuejie Yu, Yunlong Yu 0001, Yanwei Pang, Zhongfei Zhang
IEEE Trans. Cybern.3
2021 Reweighting and information-guidance networks for Few-Shot Learning
Zhong Ji, Xingliang Chai, Yunlong Yu 0001, Zhongfei Zhang
Neurocomputing3
2020 Episode-Based Prototype Generating Network for Zero-Shot Learning
abstract
We introduce a simple yet effective episode-based training framework for zero-shot learning (ZSL), where the learning system requires to recognize unseen classes given only the corresponding class semantics. During training, the model is trained within a collection of episodes, each of which is designed to simulate a zero-shot classification task. Through training multiple episodes, the model progressively accumulates ensemble experiences on predicting the mimetic unseen classes, which will generalize well on the real unseen classes. Based on this training framework, we propose a novel generative model that synthesizes visual prototypes conditioned on the class semantic prototypes. The proposed model aligns the visual-semantic interactions by formulating both the visual prototype generation and the class semantic inference into an adversarial framework paired with a parameter-economic Multi-modal Cross-Entropy Loss to capture the discriminative information. Extensive experiments on four datasets under both traditional ZSL and generalized ZSL tasks show that our model outperforms the state-of-the-art approaches by large margins.
Yunlong Yu 0001, Zhong Ji, Jungong Han, Zhongfei Zhang
CVPR1
2020 Multi-modal generative adversarial network for zero-shot learning
Zhong Ji, Junyue Wang, Yunlong Yu 0001, Zhongfei Zhang
Knowl. Based Syst.4
2020 Improved prototypical networks for few-Shot learning
Zhong Ji, Xingliang Chai, Yunlong Yu 0001, Yanwei Pang, Zhongfei Zhang
Pattern Recognit. Lett.3
2020 Attribute-Guided Network for Cross-Modal Zero-Shot Hashing
abstract
Zero-shot hashing (ZSH) aims at learning a hashing model that is trained only by instances from seen categories but can generate well to those of unseen categories. Typically, it is achieved by utilizing a semantic embedding space to transfer knowledge from seen domain to unseen domain. Existing efforts mainly focus on single-modal retrieval task, especially image-based image retrieval (IBIR). However, as a highlighted research topic in the field of hashing, cross-modal retrieval is more common in real-world applications. To address the cross-modal ZSH (CMZSH) retrieval task, we propose a novel attribute-guided network (AgNet), which can perform not only IBIR but also text-based image retrieval (TBIR). In particular, AgNet aligns different modal data into a semantically rich attribute space, which bridges the gap caused by modality heterogeneity and zero-shot setting. We also design an effective strategy that exploits the attribute to guide the generation of hash codes for image and text within the same network. Extensive experimental results on three benchmark data sets (AwA, SUN, and ImageNet) demonstrate the superiority of AgNet on both cross-modal and single-modal zero-shot image retrieval tasks.
Zhong Ji, Yunlong Yu 0001, Yanwei Pang, Jungong Han
IEEE Trans. Neural Networks Learn. Syst.3
2019 Class-specific synthesized dictionary model for Zero-Shot Learning
Zhong Ji, Junyue Wang, Yunlong Yu 0001, Yanwei Pang, Jungong Han
Neurocomputing3
2019 Zero-Shot Learning via Latent Space Encoding
abstract
Zero-shot learning (ZSL) is typically achieved by resorting to a class semantic embedding space to transfer the knowledge from the seen classes to unseen ones. Capturing the common semantic characteristics between the visual modality and the class semantic modality (e.g., attributes or word vector) is a key to the success of ZSL. In this paper, we propose a novel encoder-decoder approach, namely latent space encoding (LSE), to connect the semantic relations of different modalities. Instead of requiring a projection function to transfer information across different modalities like most previous work, LSE performs the interactions of different modalities via a feature aware latent space, which is learned in an implicit way. Specifically, different modalities are modeled separately but optimized jointly. For each modality, an encoder-decoder framework is performed to learn a feature aware latent space via jointly maximizing the recoverability of the original space from the latent space and the predictability of the latent space from the original space. To relate different modalities together, their features referring to the same concept are enforced to share the same latent codings. In this way, the common semantic characteristics of different modalities are generalized with the latent representations. Another property of the proposed approach is that it is easily extended to more modalities. Extensive experimental results on four benchmark datasets [animal with attribute, Caltech UCSD birds, aPY, and ImageNet] clearly demonstrate the superiority of the proposed approach on several ZSL tasks, including traditional ZSL, generalized ZSL, and zero-shot retrieval.
Yunlong Yu 0001, Zhong Ji, Jichang Guo, Zhongfei Zhang
IEEE Trans. Cybern.1
2018 Stacked Semantics-Guided Attention Model for Fine-Grained Zero-Shot Learning
abstract
Zero-Shot Learning (ZSL) is generally achieved via aligning the semantic relationships between the visual features and the corresponding class semantic descriptions. However, using the global features to represent fine-grained images may lead to sub-optimal results since they neglect the discriminative differences of local regions. Besides, different regions contain distinct discriminative information. The important regions should contribute more to the prediction. To this end, we propose a novel stacked semantics-guided attention (S2GA) model to obtain semantic relevant features by using individual class semantic features to progressively guide the visual features to generate an attention map for weighting the importance of different local regions. Feeding both the integrated visual features and the class semantic features into a multi-class classification architecture, the proposed framework can be trained end-to-end. Extensive experimental results on CUB and NABird datasets show that the proposed approach has a consistent improvement on both fine-grained zero-shot classification and retrieval tasks.
Yunlong Yu 0001, Zhong Ji, Yanwei Fu 0001, Jichang Guo, Yanwei Pang, Zhongfei Zhang
NeurIPS1
2018 Semantic softmax loss for zero-shot learning
Zhong Ji, Yunlong Yu 0001, Jichang Guo, Yanwei Pang
Neurocomputing3
2018 Transductive Zero-Shot Learning With a Self-Training Dictionary Approach
abstract
As an important and challenging problem in computer vision, zero-shot learning (ZSL) aims at automatically recognizing the instances from unseen object classes without training data. To address this problem, ZSL is usually carried out in the following two aspects: 1) capturing the domain distribution connections between seen classes data and unseen classes data and 2) modeling the semantic interactions between the image feature space and the label embedding space. Motivated by these observations, we propose a bidirectional mapping-based semantic relationship modeling scheme that seeks for cross-modal knowledge transfer by simultaneously projecting the image features and label embeddings into a common latent space. Namely, we have a bidirectional connection relationship that takes place from the image feature space to the latent space as well as from the label embedding space to the latent space. To deal with the domain shift problem, we further present a transductive learning approach that formulates the class prediction problem in an iterative refining process, where the object classification capacity is progressively reinforced through bootstrapping-based model updating over highly reliable instances. Experimental results on four benchmark datasets (animal with attribute, Caltech-UCSD Bird2011, aPascal-aYahoo, and SUN) demonstrate the effectiveness of the proposed approach against the state-of-the-art approaches.
Yunlong Yu 0001, Zhong Ji, Xi Li 0001, Jichang Guo, Zhongfei Zhang, Haibin Ling, Fei Wu 0001
IEEE Trans. Cybern.1
2018 Transductive Zero-Shot Learning With Adaptive Structural Embedding
abstract
Zero-shot learning (ZSL) endows the computer vision system with the inferential capability to recognize new categories that have never seen before. Two fundamental challenges in it are visual-semantic embedding and domain adaptation in cross-modality learning and unseen class prediction steps, respectively. This paper presents two corresponding methods named Adaptive STructural Embedding (ASTE) and Self-PAced Selective Strategy (SPASS) for both challenges. Specifically, ASTE formulates the visual-semantic interactions in a latent structural support vector machine framework by adaptively adjusting the slack variables to embody different reliablenesses among training instances. To alleviate the domain shift problem in ZSL, SPASS borrows the idea from self-paced learning by iteratively selecting the unseen instances from reliable to less reliable to gradually adapt the knowledge from the seen domain to the unseen domain. Consequently, by combining SPASS and ASTE, we present a self-paced Transductive ASTE (TASTE) method to progressively reinforce the classification capacity. Extensive experiments on three benchmark data sets (i.e., AwA, CUB, and aPY) demonstrate the superiorities of ASTE and TASTE. Furthermore, we also propose a fast training (FT) strategy to improve the efficiency of most existing ZSL methods. The FT strategy is surprisingly simple and general enough, which speeds up the training time of most existing ZSL methods by 4~300 times while holding the previous performance.
Yunlong Yu 0001, Zhong Ji, Jichang Guo, Yanwei Pang
IEEE Trans. Neural Networks Learn. Syst.1
2017 Zero-shot learning with regularized cross-modality ranking
Yunlong Yu 0001, Zhong Ji, Jichang Guo, Yanwei Pang
Neurocomputing1
2017 Manifold regularized cross-modal embedding for zero-shot learning
Zhong Ji, Yunlong Yu 0001, Yanwei Pang, Jichang Guo, Zhongfei Zhang
Inf. Sci.2
2017 Zero-shot learning with Multi-Battery Factor Analysis
Zhong Ji, Yunlong Yu 0001, Yanwei Pang, Zhongfei Zhang
Signal Process.2
2016 Visual search reranking with RElevant Local Discriminant Analysis
Peiguang Jing, Zhong Ji, Yunlong Yu 0001, Zhongfei Zhang
Neurocomputing3