Linglan Zhao

dblp:266/9556 · DBLP profile ↗
← Back
16ranked-venue papers
6as first author
15since 2021 · last 2026
0000-0002-2241-6977ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 10 since 2021Artificial intelligence and machine learning · 7 · 4 first-author · 6 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Mitigating Catastrophic Forgetting in Online Continual Learning With Dual-Margin Contrastive Replay
abstract
Online Continual Learning (OCL) enables machine learning models to learn from a stream of non-stationary tasks, making it more aligned with real-world scenarios. However, OCL faces a significant challenge: catastrophic forgetting, wherein the model learned in previous tasks is substantially overwritten upon encountering new tasks, leading to a biased forgetting of prior knowledge. Among various OCL strategies, replay-based methods have proven particularly effective in mitigating catastrophic forgetting by maintaining a small buffer of past samples and retraining them alongside new data. However, due to strict memory constraints, these replay buffers often fail to adequately represent the true data distribution of previous tasks. This leads to distributional shifts in the feature space, amplifying forgetting and degrading model performance. To address the problem, in this paper, we propose a novel replay strategy, termed Dual-Margin Contrastive Replay (DMCR), to anchor the distribution of old tasks and reduce the negative transfer effects. First, we propose to select memory for more representative samples guided by constructed centroids in a data stream. Then, to keep the model from distribution chaos in biased replay, a two-level angular cross-task Contrastive Margin Loss (CML) is proposed, to encourage the intra-class and intra-task compactness, and increase the inter-class and inter-task discrepancy. Finally, to further suppress the distributional drift, we present an optional Centroid Distillation Loss (CDL) on the replay memory to anchor the knowledge in feature space for each previous old task. Extensive experimental results on five benchmark datasets validate that the proposed DMCR can effectively mitigate the catastrophic forgetting and achieve state-of-the-art (SOTA) performance in OCL.
Fan Lyu, Gongbo Cheng, Daofeng Liu, Linglan Zhao, Zhang Zhang 0001, Fuyuan Hu, Liang Wang 0001
IEEE Trans. Circuits Syst. Video Technol.4
2026 GAIN: Global-Atomic INteraction Graph for Few-Shot Class-Incremental Learning
Fan Lyu, Linglan Zhao, Chengyan Liu, Yinying Mei, Zhang Zhang 0001, Baoqing Yu, Fuyuan Hu, Liang Wang 0001
IEEE Trans. Circuits Syst. Video Technol.2
2025 Generalized Category Discovery via Reciprocal Learning and Class-Wise Distribution Regularization
abstract
Generalized Category Discovery (GCD) aims to identify unlabeled samples by leveraging the base knowledge from labeled ones, where the unlabeled set consists of both base and novel classes. Since clustering methods are time-consuming at inference, parametric-based approaches have become more popular. However, recent parametric-based methods suffer from inferior base discrimination due to unreliable self-supervision. To address this issue, we propose a Reciprocal Learning Framework (RLF) that introduces an auxiliary branch devoted to base classification. During training, the main branch filters the pseudo-base samples to the auxiliary branch. In response, the auxiliary branch provides more reliable soft labels for the main branch, leading to a virtuous cycle. Furthermore, we introduce Class-wise Distribution Regularization (CDR) to mitigate the learning bias towards base classes. CDR essentially increases the prediction confidence of the unlabeled data and boosts the novel class performance. Combined with both components, our proposed method, RLCD, achieves superior performance in all classes with negligible extra computation. Comprehensive experiments across seven GCD datasets validate its superiority. Our codes are available at https://github.com/APORduo/RLCD.
Zhiquan Tan, Linglan Zhao, Xiangzhong Fang, Weiran Huang 0001
ICML3
2025 VISA: Group-wise Visual Token Selection and Aggregation via Graph Summarization for Efficient MLLMs Inference
abstract
In this study, we introduce a novel method called group-wise VI sual token Selection and Aggregation (VISA) to address the issue of inefficient inference stemming from excessive visual tokens in multimoal large language models (MLLMs). Compared with previous token pruning approaches, our method can preserve more visual information while compressing visual tokens. We first propose a graph-based visual token aggregation (VTA) module. VTA treats each visual token as a node, forming a graph based on semantic similarity among visual tokens. It then aggregates information from removed tokens into kept tokens based on this graph, producing a more compact visual token representation. Additionally, we introduce a group-wise token selection strategy (GTS) to divide visual tokens into kept and removed ones, guided by text tokens from the final layers of each group. This strategy progressively aggregates visual information, enhancing the stability of the visual information extraction process. We conduct comprehensive experiments on LLaVA-1.5, LLaVA-NeXT, and Video-LLaVA across various benchmarks to validate the efficacy of VISA. Our method consistently outperforms previous methods, achieving a superior trade-off between model performance and inference speed.
Hanjun Li 0002, Linglan Zhao, Fei Chao 0001, Shouhong Ding, Rongrong Ji
ACM Multimedia3
2025 Few-Shot Class-Incremental Learning via Asymmetric Supervised Contrastive Learning
abstract
Few-Shot Class-Incremental Learning (FSCIL) is to continuously learn novel classes from a few samples without forgetting previous knowledge. Adapting directly to limited novel data typically results in significant forgetting of base class knowledge. Consequently, prevailing FSCIL methods are devoted to training a strong initial model that can be frozen in incremental sessions. However, these works face a dilemma in poor generalization: they benefit mainly from base class performance, yet underperform in novel classes. To alleviate this issue, we design a two-stage training framework to simultaneously enhance generalization for novel classes and maintain base class discrimination. In the first stage, an asymmetric supervised contrastive learning (AsyCon) algorithm is proposed. AsyCon introduces a predicted feature to achieve an asymmetric alignment of positive pairs. It alleviates over-similarity within positive features, allowing the model to better transfer to new classes in incremental sessions. In the second stage, the model is finetuned for promoting its performance on base classes. To maintain the generalization obtained in the first stage, we employ an L2 normalized regularization (LR) to keep the feature consistent with the model in the first stage. The finetuned model, termed AsyCLR, effectively balances generalization and discrimination, significantly outperforming existing FSCIL works especially in novel class accuracy. Experiments on CUB200, CIFAR100, and mini-ImageNet verify the effectiveness of our method. Additionally, our method also performs well in the standard few-shot recognition scenario due to its strong generalization ability. Our codes are available at https://github.com/APORduo/AsyCLR.
Duo Liu 0001, Linglan Zhao, Fan Lyu, Xiangzhong Fang, Liang Wang 0001
IEEE Trans. Circuits Syst. Video Technol.2
2025 A Three-Dimensional Complete Complementary Coded Spread Spectrum System Designed for Multi-User-Multi-Target ISAC Scenarios
abstract
Integrated sensing and communication (ISAC) has been widely recognized as an effective solution to achieving robust performances in both communication and sensing within the same spectrum, but interference poses a critical challenge in waveform design. To address this issue, we propose a code-domain waveform approach applicable to multi-user and multi-target ISAC scenarios. Specifically, this work describes the scheme to mitigate multipath, multi-user, multi-antenna, and mutual interferences between communication and sensing by applying three-dimensional complete complementary codes (3D-CCC). This study provides a comprehensive overview of the codebook structure and details the signal processing workflow. The impact of codebook parameters on system performance is evaluated through simulations considering bit error rate (BER), data rate, radar detection probability, and Kullback-Leibler divergence (KLD). The simulation results show that, in terms of communication performance, the code-domain-based spread spectrum technique enhances the robustness to interference and ensures reliable signal transmission. In terms of sensing performance, 3D-CCC achieves higher range resolution and lower angular mean square error. Under certain conditions, the proposed scheme outperforms existing systems in terms of detection probability.
Xiqing Liu, Linglan Zhao, Mugen Peng
IEEE Trans. Wirel. Commun.3
2024 Distillation Excluding Positives for Few-Shot Class-Incremental Learning
abstract
Few-Shot Class-Incremental Learning (FSCIL) defines a challenging task to continually recognize novel classes with few training data without forgetting old classes. Considering the catastrophic forgetting and overfitting issues, mainstream FSCIL methods resort to obtaining a strong model in the base session and freezing it in incremental sessions. Although prevailing methods perform well in the base classes, they often struggle with poor novel class generalization. To strengthen the representation of these models, this paper focuses on Knowledge Distillation (KD). Since existing KD methods are incompetent for FSCIL and introduce limited improvement, we propose the Distillation Excluding Positives (DEP) method for boosting performance on FSCIL tasks. Specifically, DEP consists of negative relationship distillation and asymmetric self-feature distillation. It can mitigate the over-similarity of intra-class features, leading to a more generalized model. Extensive experiments on three FSCIL benchmarks validate the superiority of DEP over current SOTAs.
Linglan Zhao, Fuhan Cai, Xiangzhong Fang
ICME2
2024 SAFE: Slow and Fast Parameter-Efficient Tuning for Continual Learning with Pre-Trained Models
abstract
Continual learning aims to incrementally acquire new concepts in data streams while resisting forgetting previous knowledge. With the rise of powerful pre-trained models (PTMs), there is a growing interest in training incremental learning systems using these foundation models, rather than learning from scratch. Existing works often view PTMs as a strong initial point and directly apply parameter-efficient tuning (PET) in the first session for adapting to downstream tasks. In the following sessions, most methods freeze model parameters for tackling forgetting issues. However, applying PET directly to downstream data cannot fully explore the inherent knowledge in PTMs. Additionally, freezing the parameters in incremental sessions hinders models' plasticity to novel concepts not covered in the first session. To solve the above issues, we propose a Slow And Fast parameter-Efficient tuning (SAFE) framework. In particular, to inherit general knowledge from foundation models, we include a transfer loss function by measuring the correlation between the PTM and the PET-applied model. After calibrating in the first session, the slow efficient tuning parameters can capture more informative features, improving generalization to incoming classes. Moreover, to further incorporate novel concepts, we strike a balance between stability and plasticity by fixing slow efficient tuning parameters and continuously updating the fast ones. Specifically, a cross-classification loss with feature alignment is proposed to circumvent catastrophic forgetting. During inference, we introduce an entropy-based aggregation strategy to dynamically utilize the complementarity in the slow and fast learners. Extensive experiments on seven benchmark datasets verify the effectiveness of our method by significantly surpassing the state-of-the-art.
Linglan Zhao, Xuerui Zhang, Shouhong Ding, Weiran Huang 0001
NeurIPS1
2023 Few-Shot Class-Incremental Learning via Class-Aware Bilateral Distillation
abstract
Few-Shot Class-Incremental Learning (FSCIL) aims to continually learn novel classes based on only few training samples, which poses a more challenging task than the well-studied Class-Incremental Learning (CIL) due to data scarcity. While knowledge distillation, a prevailing technique in CIL, can alleviate the catastrophic forgetting of older classes by regularizing outputs between current and previous model, it fails to consider the overfitting risk of novel classes in FSCIL. To adapt the powerful distillation technique for FSCIL, we propose a novel distillation structure, by taking the unique challenge of overfitting into account. Concretely, we draw knowledge from two complementary teachers. One is the model trained on abundant data from base classes that carries rich general knowledge, which can be leveraged for easing the overfitting of current novel classes. The other is the updated model from last incremental session that contains the adapted knowledge of previous novel classes, which is used for alleviating their forgetting. To combine the guidances, an adaptive strategy conditioned on the class-wise semantic similarities is introduced. Besides, for better preserving base class knowledge when accommodating novel concepts, we adopt a two-branch network with an attention-based aggregation module to dynamically merge predictions from two complementary branches. Extensive experiments on 3 popular FSCIL datasets: mini-ImageNet, CIFAR100 and CUB200 validate the effectiveness of our method by surpassing existing works by a significant margin. Code is available at https://github.com/LinglanZhao/BiDistFSCIL.
Linglan Zhao, Jing Lu 0004, Yunlu Xu, Zhanzhan Cheng, Dashan Guo, Xiangzhong Fang
CVPR1
2023 Rethinking Self-Supervision for Few-Shot Class-Incremental Learning
abstract
Few-Shot Class-Incremental Learning (FSCIL) focuses on progressively absorbing new concepts given only limited training data. For tackling this challenge, several recent FS-CIL works resort to pre-training models with Self-Supervised Learning (SSL) to obtain features that can generalize well to new classes. However, to avoid overfitting and catastrophic forgetting, previous works only leverage SSL in the base session and keep all or most parameters fixed in incremental sessions, resulting in inadequate adaptation to novel classes. Thus, in this paper, we explore the setting where more parameters can be updated for adapting to novel concepts, and discover that the model pre-trained with SSL leads to degraded performance even compared to that without SSL. It can be attributed to the severer forgetting of base class knowledge. To address this issue, we propose an imprinting-based distillation module for effectively regularizing the adaption process, and a mathematically provable routing strategy for further improved results. The effectiveness of our approach is verified on 3 popular FSCIL benchmarks by significantly outperforming previous methods.
Linglan Zhao, Jing Lu 0004, Zhanzhan Cheng, Xiangzhong Fang
ICME1
2022 Boosting Few-shot visual recognition via saliency-guided complementary attention
Linglan Zhao, Dashan Guo, Wei Li 0084, Xiangzhong Fang
Neurocomputing1
2021 A Strong Baseline for Semi-Supervised Incremental Few-Shot Learning
Linglan Zhao, Dashan Guo, Yunlu Xu, Liang Qiao 0001, Zhanzhan Cheng, Shiliang Pu, Xiangzhong Fang
BMVC1
2021 Saliency-Guided Complementary Attention for Improved Few-Shot Learning
abstract
Despite significant progress in recent deep neural networks, most deep learning algorithms rely heavily on abundant training samples. To address this problem, we propose an effective and interpretable few-shot classification model using Saliency-Guided Complementary Attention (SGCA), which aims to learn transferable representations and to build a robust classification module simultaneously. Concretely, we propose to train our feature extractor using an auxiliary task to separate object regions from background clutter guided by saliency detection signals. In addition, to make the separation beneficial to the downstream tasks, we introduce a complementary attention mechanism to force the classification module to focus on various informative parts of the image. Extensive experiments on few-shot learning tasks demonstrate the effectiveness of our proposed method, e.g., we achieve 68.81% and 84.60% for 5-way 1-shot and 5-shot settings on mini-ImageNet, respectively.
Linglan Zhao, Dashan Guo, Wei Li 0084, Xiangzhong Fang
ICME1
2021 Class-wise Metric Scaling for Improved Few-Shot Classification
abstract
Few-shot classification aims to generalize basic knowledge to recognize novel categories from a few samples. Recent centroid-based methods achieve promising classification performance with the nearest neighbor rule. However, we consider that those methods intrinsically ignore per-class distribution, as the decision boundaries are biased due to the diversity of intra-class variances. Hence, we propose a class-wise metric scaling (CMS) mechanism, which can be applied to both training and testing stages. Concretely, metric scalars are set as learnable parameters in the training stage, helping to learn a more discriminative and transferable feature representation. As for testing, we construct a convex optimization problem to generate an optimal scalar vector for refining the nearest neighbor decisions. Besides, we also involve a low-ranking bilinear pooling layer for improved representation capacity, which further provides significant performance gains. Extensive experiments are conducted on a series of feature extractor backbones, datasets, and testing modes, which have shown consistent improvements compared to prior SOTA methods, e.g., we achieve accuracies of 66.64 % and 83.63 % for 5-way 1-shot and 5-shot settings on the mini-ImageNet, respectively. Under the semi-supervised inductive mode, results are further up to 78.34 % and 87.53 %, respectively.
Linglan Zhao, Wei Li 0084, Dashan Guo, Xiangzhong Fang
WACV2
2021 PDA: Proxy-based domain adaptation for few-shot image recognition
Linglan Zhao, Xiangzhong Fang
Image Vis. Comput.2
2020 Visual question answering with attention transfer and a cross-modal gating mechanism
Wei Li 0084, Jianhui Sun, Linglan Zhao, Xiangzhong Fang
Pattern Recognit. Lett.4