Ji Zhang 0012

dblp:86/1953-12 · DBLP profile ↗
← Back
18ranked-venue papers
10as first author
15since 2021 · last 2026
0000-0001-6949-3673ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 6 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 6 first-author · 9 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 A Closer Look at Conditional Prompt Tuning for Vision-Language Models
Ji Zhang 0012, Shihan Wu 0001, Lianli Gao, Jingkuan Song, Nicu Sebe, Heng Tao Shen
Int. J. Comput. Vis.1
2026 From Channel Bias to Feature Redundancy: Uncovering the "Less is More" Principle in Few-Shot Learning
abstract
Deep neural networks often fail to adapt representations to novel tasks under distribution shifts, especially when only a few examples are available. This paper identifies a core obstacle behind this failure: Channel Bias, where networks develop a rigid emphasis on feature dimensions that were discriminative for the source task, but this emphasis is misaligned and fails to adapt to the distinct needs of a novel task. This bias leads to a striking and detrimental consequence: Feature Redundancy. We demonstrate that for few-shot tasks, classification accuracy is significantly improved by using as few as 1-5% of the most discriminative feature dimensions, revealing that the vast majority are actively harmful. Our theoretical analysis confirms that this redundancy originates from confounding feature dimensions-those with high intra-class variance but low inter-class separability-which are especially problematic in low-data regimes. This "Less is More" phenomenon is a defining characteristic of the few-shot setting, diminishing as more samples become available. To address this, we propose a simple yet effective soft-masking method, Augmented Feature Importance Adjustment (AFIA), which estimates feature importance from augmented data to mitigate the issue. By establishing the cohesive link from channel bias to its consequence of extreme feature redundancy, this work provides a foundational principle for few-shot representation transfer and a practical method for developing more robust few-shot learning algorithms.
Ji Zhang 0012, Xu Luo 0003, Lianli Gao, Difan Zou, Heng Tao Shen, Jingkuan Song
IEEE Trans. Pattern Anal. Mach. Intell.1
2026 OmniCharacter++: Toward Comprehensive Benchmark for Realistic Role-Playing Agents
Haonan Zhang 0003, Pengpeng Zeng, Ji Zhang 0012, Jingkuan Song, Nicu Sebe, Heng Tao Shen, Lianli Gao
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 Skip Tuning: Pre-trained Vision-Language Models are Effective and Efficient Adapters Themselves
abstract
Prompt tuning (PT) has long been recognized as an effective and efficient paradigm for transferring large pre-trained vision-language models (VLMs) to downstream tasks by learning a tiny set of context vectors. Nevertheless, in this work, we reveal that freezing the parameters of VLMs during learning the context vectors neither facilitates the transferability of pre-trained knowledge nor improves the memory and time efficiency significantly. Upon further investigation, we find that reducing both the length and width of the feature-gradient propagation flows of the full fine-tuning (FT) baseline is key to achieving effective and efficient knowledge transfer. Motivated by this, we propose Skip Tuning, a novel paradigm for adapting VLMs to downstream tasks. Unlike existing PT or adapter-based methods, Skip Tuning applies Layer-wise Skipping (LSkip) and Classwise Skipping (CSkip) upon the FT baseline without introducing extra context vectors or adapter modules. Extensive experiments across a wide spectrum of benchmarks demonstrate the superior effectiveness and efficiency of our Skip Tuning over both PT and adapter-based methods. Code: https://github.com/Koorye/SkipTuning.
Shihan Wu 0001, Ji Zhang 0012, Pengpeng Zeng, Lianli Gao, Jingkuan Song, Heng Tao Shen
CVPR2
2025 Reliable Few-Shot Learning Under Dual Noises
abstract
Recent advances in model pre-training give rise to task adaptation-based few-shot learning (FSL), where the goal is to adapt a pre-trained task-agnostic model for capturing task-specific knowledge with a few-labeled support samples of the target task. Nevertheless, existing approaches may still fail in the open world due to the inevitable in-distribution (ID) and out-of-distribution (OOD) noise from both support and query samples of the target task. With limited support samples available, i) the adverse effect of the dual noises can be severely amplified during task adaptation, and ii) the adapted model can produce unreliable predictions on query samples in the presence of the dual noises. In this work, we propose DEnoised Task Adaptation (DETA++) for reliable FSL. DETA++ uses a Contrastive Relevance Aggregation (CoRA) module to calculate image and region weights for support samples, based on which a clean prototype loss and a noise entropy maximization loss are proposed to achieve noise-robust task adaptation. Additionally, DETA++ employs a memory bank to store and refine clean regions for each inner-task class, based on which a Local Nearest Centroid Classifier (LocalNCC) is devised to yield noise-robust predictions on query samples. Moreover, DETA++ utilizes an Intra-class Region Swapping (IntraSwap) strategy to rectify ID class prototypes during task adaptation, enhancing the model's robustness to the dual noises. Extensive experiments demonstrate the effectiveness and flexibility of DETA++.
Ji Zhang 0012, Jingkuan Song, Lianli Gao, Nicu Sebe, Heng Tao Shen
IEEE Trans. Pattern Anal. Mach. Intell.1
2025 Modeling Temporal Dependencies Within the Target for Long-Term Time Series Forecasting
abstract
Long-term time series forecasting (LTSF) is a critical task across diverse domains. Despite significant advancements in LTSF research, we identify a performance bottleneck in existing LTSF methods caused by the inadequate modeling of Temporal Dependencies within the Target (TDT). To address this issue, we propose a novel and generic temporal modeling framework, Temporal Dependency Alignment (TDAlign), that equips existing LTSF methods with TDT learning capabilities. TDAlign introduces two key innovations: 1) a loss function that aligns the change values between adjacent time steps in the predictions with those in the target, ensuring consistency with variation patterns, and 2) an adaptive loss balancing strategy that seamlessly integrates the new loss function with existing LTSF methods without introducing additional learnable parameters. As a plug-and-play framework, TDAlign enhances existing methods with minimal computational overhead, featuring only linear time complexity and constant space complexity relative to the prediction length. Extensive experiments on six strong LTSF baselines across seven real-world datasets demonstrate the effectiveness and flexibility of TDAlign. On average, TDAlign reduces baseline prediction errors by1.47%to9.19%and change value errors by4.57%to15.78%, highlighting its substantial performance improvements.
Minbo Ma, Ji Zhang 0012, Jie Xu 0007, Tianrui Li 0001
IEEE Trans. Knowl. Data Eng.4
2024 DePT: Decoupled Prompt Tuning
abstract
This work breaks through the Base-New Tradeoff (BNT) dilemma in prompt tuning, i.e., the better the tuned model generalizes to the base (or target) task, the worse it generalizes to new tasks, and vice versa. Specifically, through an in-depth analysis of the learned features of the base and new tasks, we observe that the BNT stems from a channel bias issue - the vast majority of feature channels are occupied by base-specific knowledge, leading to the collapse of task-shared knowledge important to new tasks. To address this, we propose the Decoupled Prompt Tuning (DePT) framework, which decouples base-specific knowledge from feature channels into an isolated feature space during prompt tuning, so as to maximally preserve task-shared knowl-edge in the original feature space for achieving better zero-shot generalization on new tasks. Importantly, our DePT is orthogonal to existing prompt tuning approaches, and can enhance them with negligible additional computational cost. Extensive experiments on several datasets show the flexibility and effectiveness of DePT. Code is available at https://github.com/Koorye/DePT.
Ji Zhang 0012, Shihan Wu 0001, Lianli Gao, Heng Tao Shen, Jingkuan Song
CVPR1
2024 Effective and Efficient Few-shot Fine-tuning for Vision Transformers
abstract
Parameter-efficient fine-tuning (PEFT), updating only a small set of parameters either inherently in the model or additionally introduced, reduces the cost of adaptation of large vision models (e.g. Vision Transformers) and avoids overfitting to few-shot samples. However, the selection of parameters to update often follows heuristic criteria, thus lacking systematic analysis and may lead to suboptimal results. In this work, we adopt the concept of skilled parameter localization (SPL) from the NLP community, which can identify the location of task-specific parameters in a fine-tuned model automatically given any task. By applying this technique to ViTs, we observe that while the task-specific (skilled) parameters scatter in the parameter space across different tasks, the out-projection bias of attention and MLP layers are often concentrated with these skilled parameters. Inspired by this, we propose Out-projection Bias Fine-Tuning, or OBFT, a simple yet effective PEFT method that conducts few-shot adaptation solely relying on the out-projection bias of attention and MLP modules in pre-trained ViTs. We demonstrate the effectiveness and efficiency of our OBFT over 10 diverse datasets: 1) OBFT achieves superior parameter efficiency than a broad spectrum of PEFT strategies; 2) by updating only 0.01% parameters of ViTs, OBFT attains comparable performance with full fine-tuning, while significantly reducing training costs, as it does not need to maintain optimizer states for most parameters.
Hao Wu 0070, Ji Zhang 0012, Lianli Gao, Jingkuan Song
ICME3
2023 DETA: Denoised Task Adaptation for Few-Shot Learning
abstract
Test-time task adaptation in few-shot learning aims to adapt a pre-trained task-agnostic model for capturing task-specific knowledge of the test task, rely only on few-labeled support samples. Previous approaches generally focus on developing advanced algorithms to achieve the goal, while neglecting the inherent problems of the given support samples. In fact, with only a handful of samples available, the adverse effect of either the image noise (a.k.a. X-noise) or the label noise (a.k.a. Y-noise) from support samples can be severely amplified. To address this challenge, in this work we propose DEnoised Task Adaptation (DETA), a first, unified image- and label-denoising framework orthogonal to existing task adaptation approaches. Without extra supervision, DETA filters out task-irrelevant, noisy representations by taking advantage of both global visual information and local region details of support samples. On the challenging Meta-Dataset, DETA consistently improves the performance of a broad spectrum of baseline methods applied on various pre-trained models. Notably, by tackling the overlooked image noise in Meta-Dataset, DETA establishes new state-of-the-art results. Code is released at https://github.com/JimZAI/DETA.
Ji Zhang 0012, Lianli Gao, Xu Luo 0003, Heng Tao Shen, Jingkuan Song
ICCV1
2023 A Closer Look at Few-shot Classification Again
abstract
Few-shot classification consists of a training phase where a model is learned on a relatively large dataset and an adaptation phase where the learned model is adapted to previously-unseen tasks with limited labeled samples. In this paper, we empirically prove that the training algorithm and the adaptation algorithm can be completely disentangled, which allows algorithm analysis and design to be done individually for each phase. Our meta-analysis for each phase reveals several interesting insights that may help better understand key aspects of few-shot classification and connections with other fields such as visual representation learning and transfer learning. We hope the insights and research challenges revealed in this paper can inspire future work in related directions. Code and pre-trained models (in PyTorch) are available at https://github.com/Frankluox/CloserLookAgainFewShot.
Xu Luo 0003, Hao Wu 0070, Ji Zhang 0012, Lianli Gao, Jingkuan Song
ICML3
2023 From Global to Local: Multi-Scale Out-of-Distribution Detection
abstract
Out-of-distribution (OOD) detection aims to detect "unknown" data whose labels have not been seen during the in-distribution (ID) training process. Recent progress in representation learning gives rise to distance-based OOD detection that recognizes inputs as ID/OOD according to their relative distances to the training data of ID classes. Previous approaches calculate pairwise distances relying only on global image representations, which can be sub-optimal as the inevitable background clutter and intra-class variation may drive image-level representations from the same ID class far apart in a given representation space. In this work, we overcome this challenge by proposing Multi-scale OOD DEtection (MODE), a first framework leveraging both global visual information and local region details of images to maximally benefit OOD detection. Specifically, we first find that existing models pretrained by off-the-shelf cross-entropy or contrastive losses are incompetent to capture valuable local representations for MODE, due to the scale-discrepancy between the ID training and OOD detection processes. To mitigate this issue and encourage locally discriminative representations in ID training, we propose Attention-based Local PropAgation (ALPA), a trainable objective that exploits a cross-attention mechanism to align and highlight the local regions of the target objects for pairwise examples. During test-time OOD detection, a Cross-Scale Decision (CSD) function is further devised on the most discriminative multi-scale representations to distinguish ID/OOD data more faithfully. We demonstrate the effectiveness and flexibility of MODE on several benchmarks - on average, MODE outperforms the previous state-of-the-art by up to 19.24% in FPR, 2.77% in AUROC. Code is available at https://github.com/JimZAI/MODE-OOD.
Ji Zhang 0012, Lianli Gao, Bingguang Hao, Jingkuan Song, Heng Tao Shen
IEEE Trans. Image Process.1
2022 Class Gradient Projection For Continual Learning
abstract
Catastrophic forgetting is one of the most critical challenges in Continual Learning (CL). Recent approaches tackle this problem by projecting the gradient update orthogonal to the gradient subspace of existing tasks. While the results are remarkable, those approaches ignore the fact that these calculated gradients are not guaranteed to be orthogonal to the gradient subspace of each class due to the class deviation in tasks, e.g., distinguishing "Man" from "Sea" v.s. differentiating "Boy" from "Girl". Therefore, this strategy may still cause catastrophic forgetting for some classes. In this paper, we propose Class Gradient Projection (CGP), which calculates the gradient subspace from individual classes rather than tasks. Gradient update orthogonal to the gradient subspace of existing classes can be effectively utilized to minimize interference from other classes. To improve the generalization and efficiency, we further design a Base Refining (BR) algorithm to combine similar classes and refine class bases dynamically. Moreover, we leverage a contrastive learning method to improve the model's ability to handle unseen tasks. Extensive experiments on benchmark datasets demonstrate the effectiveness of our proposed approach. It improves the previous methods by 2.0% on the CIFAR-100 dataset. The code is available at https://github.com/zackschen/CGP.
Ji Zhang 0012, Jingkuan Song, Lianli Gao
ACM Multimedia2
2022 Free-Lunch for Cross-Domain Few-Shot Learning: Style-Aware Episodic Training with Robust Contrastive Learning
abstract
Cross-Domain Few-Shot Learning (CDFSL) aims for training an adaptable model that can learn out-of-domain classes with a handful of samples. Compared to the well-studied few-shot learning problem, the difficulty for CDFSL lies in that the available training data from test tasks is not only extremely limited but also presents severe class differences from training tasks. To tackle this challenge, we propose Style-aware Episodic Training with Robust Contrastive Learning (SET-RCL), which is motivated by the key observation that a remarkable style-shift between tasks from source and target domains plays a negative role in cross-domain generalization. SET-RCL addresses the style-shift from two perspectives: 1) simulating the style distributions of unknown target domains (data perspective); and 2) learning a style-invariant representation (model perspective). Specifically, Style-aware Episodic Training (SET) focuses on manipulating the styl distribution of training tasks in the source domain, such that the learned model can achieve better adaption on test tasks with domain-specific styles. To further improve cross-domain generalization under style-shift, we develop Robust Contrastive Learning (RCL) to capture style-invariant and discriminative representations from the manipulated tasks. Notably,our SET-RCL is orthogonal to existing FSL approaches, thus can be adopted as a "free-lunch" for boosting their CDFSL performance. Extensive experiments on nine benchmark datasets and six baseline methods demonstrate the effectiveness of our method.
Ji Zhang 0012, Jingkuan Song, Lianli Gao, Heng Tao Shen
ACM Multimedia1
2022 Progressive Meta-Learning With Curriculum
abstract
Meta-learning offers an effective solution to learn new concepts under scarce supervision through an episodic-training scheme: a series of target-like tasks sampled from base classes are sequentially fed into a meta-learner to extract cross-task knowledge, which can facilitate the quick acquisition of task-specific knowledge of the target task with few samples. Despite its noticeable improvements, the episodic-training strategy samples tasks randomly and uniformly, without considering their hardness and quality, which may not progressively improve the meta-leaner’s generalization. In this paper, we propose Progressive Meta-learning using tasks from easy to hard. First, based on a predefined curriculum, we develop a Curriculum-Based Meta-learning (CubMeta) method. CubMeta is in a stepwise manner, and in each step, we design a BrotherNet module to establish harder tasks and an effective learning scheme for obtaining an ensemble of stronger meta-learners. Then we move a step further to propose an end-to-end Self-Paced Meta-learning (SepMeta) method. The curriculum in SepMeta is effectively integrated as a regularization term into the objective so that the meta-learner can measure the hardness of tasks adaptively, according to what the model has already learned. Extensive experiments on benchmark datasets demonstrate the effectiveness of the proposed methods. Our code is available athttps://github.com/nobody-777.
Ji Zhang 0012, Jingkuan Song, Lianli Gao, Heng Tao Shen
IEEE Trans. Circuits Syst. Video Technol.1
2021 Curriculum-Based Meta-learning
abstract
Meta-learning offers an effective solution to learn new concepts with scarce supervision through an episodic training scheme: a series of target-like tasks sampled from base classes are sequentially fed into a meta-learner to extract common knowledge across tasks, which can facilitate the quick acquisition of task-specific knowledge of the target task with few samples. Despite its noticeable improvements, the episodic training strategy samples tasks randomly and uniformly, without considering their hardness and quality, which may not progressively improve the meta-leaner's generalization ability. In this paper, we present a Curriculum-Based Meta-learning (CubMeta) method to train the meta-learner using tasks from easy to hard. Specifically, the framework of CubMeta is in a progressive way, and in each step, we design a module named BrotherNet to establish harder tasks and an effective learning scheme for obtaining an ensemble of stronger meta-learners. In this way, the meta-learner's generalization ability can be progressively improved, and better performance can be obtained even with fewer training tasks. We evaluate our method for few-shot classification on two benchmarks - mini-ImageNet and tiered-ImageNet, where it achieves consistent performance improvements on various meta-learning paradigms.
Ji Zhang 0012, Jingkuan Song, Yazhou Yao, Lianli Gao
ACM Multimedia1
2019 A factor graph model for unsupervised feature selection
Hongjun Wang 0002, Yinghui Zhang 0005, Ji Zhang 0012, Tianrui Li 0001, Lingxi Peng
Inf. Sci.3
2019 Improved Gaussian-Bernoulli restricted Boltzmann machine for learning discriminative representations
Ji Zhang 0012, Hongjun Wang 0002, Jielei Chu, Shudong Huang, Tianrui Li 0001, Qigang Zhao
Knowl. Based Syst.1
2019 Particle Subswarms Collaborative Clustering
abstract
Collaborative clustering aims to find a common data structure between several distributed data sets governed by different privacy constraints and technical limitations that prohibit a central collection of data for processing. Therefore, it is required to process the data sets separately using collaboration, which allows clustering algorithms to work locally on an individual data set while exchanging information about the finding with algorithms in other data locations. Thus, the different data locations share information to improve individual clustering result amidst technical and privacy limitations but without breaching privacy. In this article, we present a framework of collaborative clustering that does not require interaction coefficients to regulate the effect of collaboration. We further adapt the framework to cluster distributed data using crisp and fuzzy clustering algorithms. We use particle swarm optimization techniques to inference the framework and, therefore, call it particle subswarms. Moreover, the collaboration increases the number of particles in the swarm without increasing the number of clusters in the data set. This article, therefore, provides the theoretical foundations of particle subswarms and some experimental results on several data sets.
Collins Census, Hongjun Wang 0002, Ji Zhang 0012, Ping Deng 0002, Tianrui Li 0001
IEEE Trans. Comput. Soc. Syst.3