Fan Lyu

dblp:227/4526 · DBLP profile ↗
← Back
43ranked-venue papers
6as first author
36since 2021 · last 2026
0000-0002-0878-5485ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 32 · 5 first-author · 26 since 2021Artificial intelligence and machine learning · 22 · 3 first-author · 20 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Sparse Tuning Enhances Plasticity in PTM-based Continual Learning
abstract
Continual Learning with Pre-trained Models holds great promise for efficient adaptation across sequential tasks. However, most existing approaches freeze PTMs and rely on auxiliary modules like prompts or adapters, limiting model plasticity and leading to suboptimal generalization when facing significant distribution shifts. While full fine-tuning can improve adaptability, it risks disrupting crucial pre-trained knowledge. In this paper, we propose Mutual Information-guided Sparse Tuning (MIST), a plug-and-play method that selectively updates a small subset of PTM parameters, less than 5%, based on sensitivity to mutual information objectives. MIST enables effective task-specific adaptation while preserving generalization. To further reduce interference, we introduce strong sparsity regularization by randomly dropping gradients during tuning, resulting in fewer than 0.5% of parameters being updated per step. Applied before standard freeze-based methods, MIST consistently boosts performance across diverse continual learning benchmarks. Experiments show that integrating our method into multiple baselines yields significant performance gains.
Shenghua Fan, Shuyu Dong, Yujin Zheng, Dingwen Wang, Fan Lyu
AAAI6
2026 A negative-anchored self-relabeling strategy for multi-label class-incremental learning
Kaile Du, Junzhou Xie, Fan Lyu, Zihan Ye, Guangcan Liu
Neurocomputing3
2026 Negative-weighted knowledge distillation regularized graph convolutional network for multi-label class-incremental learning
Kaile Du, Junzhou Xie, Fan Lyu, Zihan Ye, Yuyang Li 0005, Guangcan Liu
Pattern Recognit.3
2026 Mitigating Catastrophic Forgetting in Online Continual Learning With Dual-Margin Contrastive Replay
abstract
Online Continual Learning (OCL) enables machine learning models to learn from a stream of non-stationary tasks, making it more aligned with real-world scenarios. However, OCL faces a significant challenge: catastrophic forgetting, wherein the model learned in previous tasks is substantially overwritten upon encountering new tasks, leading to a biased forgetting of prior knowledge. Among various OCL strategies, replay-based methods have proven particularly effective in mitigating catastrophic forgetting by maintaining a small buffer of past samples and retraining them alongside new data. However, due to strict memory constraints, these replay buffers often fail to adequately represent the true data distribution of previous tasks. This leads to distributional shifts in the feature space, amplifying forgetting and degrading model performance. To address the problem, in this paper, we propose a novel replay strategy, termed Dual-Margin Contrastive Replay (DMCR), to anchor the distribution of old tasks and reduce the negative transfer effects. First, we propose to select memory for more representative samples guided by constructed centroids in a data stream. Then, to keep the model from distribution chaos in biased replay, a two-level angular cross-task Contrastive Margin Loss (CML) is proposed, to encourage the intra-class and intra-task compactness, and increase the inter-class and inter-task discrepancy. Finally, to further suppress the distributional drift, we present an optional Centroid Distillation Loss (CDL) on the replay memory to anchor the knowledge in feature space for each previous old task. Extensive experimental results on five benchmark datasets validate that the proposed DMCR can effectively mitigate the catastrophic forgetting and achieve state-of-the-art (SOTA) performance in OCL.
Fan Lyu, Gongbo Cheng, Daofeng Liu, Linglan Zhao, Zhang Zhang 0001, Fuyuan Hu, Liang Wang 0001
IEEE Trans. Circuits Syst. Video Technol.1
2026 GAIN: Global-Atomic INteraction Graph for Few-Shot Class-Incremental Learning
Fan Lyu, Linglan Zhao, Chengyan Liu, Yinying Mei, Zhang Zhang 0001, Baoqing Yu, Fuyuan Hu, Liang Wang 0001
IEEE Trans. Circuits Syst. Video Technol.1
2026 MambaPTP: Exploring the Potential of Mamba for Pedestrian Trajectory Prediction
abstract
Pedestrian Trajectory Prediction (PTP) aims to predict the future trajectory of pedestrians based on a historical trajectory. Transformer-based approaches have demonstrated unparalleled performance for PTP tasks, encoding long-term temporal dependencies and heterogeneous spatial interactions of pedestrians. However, Transformer often involves redundant information and noisy interactions from irrelevant regions by considering all available trajectory features. Recently, the structured state space model, Mamba has been proposed, which captures long-range dependency in sequences with a selective mechanism to filter out redundant information. To further tap into the potential of the novel Mamba architecture for the PTP task, in this paper, we presentMambaPTP, which predicts future trajectories based purely on Mamba mechanisms, to mitigate the noisy interactions of irrelevant trajectory features and avoid repetitive trajectory modeling, while maintaining high-performance trajectory prediction. Specifically, we propose a new Bidirectional Gating Mamba (BGM) module with bidirectional state space models, which leverages the sparse gate mechanism to select informative temporal patterns and spatial interactions. Moreover, we design a Bidirectional Trajectory Alignment (BTA) module towards aligning the predicted trajectory to the ground truth, ensuring that the model to learn the effective sparse feature representation of trajectories. We conduct extensive experiments on several mainstream pedestrian trajectory prediction datasets. The results demonstrate that the proposed MambaPTP achieves competitive performance compared to advanced Transformer-based models. We hope this paper can further inspire research in Mamba for the PTP task, leading to a tighter integration of the Mamba and PTP communities.
Shuangqing Zhang, Gangming Zhao, Fan Lyu, Songping Wang, Zhang Zhang 0001, Fang Zhao 0006, Caifeng Shan, Liang Wang 0001
IEEE Trans. Circuits Syst. Video Technol.3
2026 MultiHuman: Leverage Multimodal Prompts for Controllable Multi-Person Image Synthesizing
Liuqing Zhao, Baicheng Chen, Fan Lyu, Richang Hong
IEEE Trans. Circuits Syst. Video Technol.4
2026 Constructing Enhanced Mutual Information for Online Class-Incremental Learning
Fan Lyu, Shenghua Fan, Yujin Zheng, Dingwen Wang
IEEE Trans. Multim.2
2025 Rebalancing Multi-Label Class-Incremental Learning
abstract
Multi-label class-incremental learning (MLCIL) is essential for real-world multi-label applications, allowing models to learn new labels while retaining previously learned knowledge continuously. However, recent MLCIL approaches can only achieve suboptimal performance due to the oversight of the positive-negative imbalance problem, which manifests at both the label and loss levels because of the task-level partial label issue. The imbalance at the label level arises from the substantial absence of negative labels, while the imbalance at the loss level stems from the asymmetric contributions of the positive and negative loss parts to the optimization. To address the issue above, we propose a Rebalance framework for both the Loss and Label levels (RebLL), which integrates two key modules: asymmetric knowledge distillation (AKD) and online relabeling (OR). AKD is proposed to rebalance at the loss level by emphasizing the negative label learning in classification loss and down-weighting the contribution of overconfident predictions in distillation loss. OR is designed for label rebalance, which restores the original class distribution in memory by online relabeling the missing classes. Our comprehensive experiments on the PASCAL VOC and MS-COCO datasets demonstrate that this rebalancing strategy significantly improves performance, achieving new state-of-the-art results even with a vanilla CNN backbone.
Kaile Du, Fan Lyu, Yuyang Li 0005, Junzhou Xie, Yixi Shen, Fuyuan Hu, Guangcan Liu
AAAI3
2025 Maintaining Consistent Inter-Class Topology in Continual Test-Time Adaptation
abstract
This paper introduces Topological Consistency Adaptation (TCA), a novel approach to Continual Test-time Adaptation (CTTA) that addresses the challenges of domain shifts and error accumulation in testing scenarios. TCA ensures the stability of inter-class relationships by enforcing a class topological consistency constraint, which minimizes the distortion of class centroids and preserves the topological structure during continuous adaptation. Additionally, we propose an intra-class compactness loss to maintain compactness within classes, indirectly supporting inter-class stability. To further enhance model adaptation, we introduce a batch imbalance topology weighting mechanism that accounts for class distribution imbalances within each batch, optimizing centroid distances and stabilizing the inter-class topology. Experiments show that our method demonstrates improvements in handling continuous domain shifts, ensuring stable feature distributions and boosting predictive performance. Our code is available at: https://github.com/Successybbdwm/TCA.
Chenggong Ni, Fan Lyu, Jiayao Tan, Fuyuan Hu, Tao Zhou 0009
CVPR2
2025 Dual Semantic Guidance for Open Vocabulary Semantic Segmentation
abstract
Open-vocabulary semantic segmentation aims to enable models to segment arbitrary categories. Currently, though pre-trained Vision-Language Models (VLMs) like CLIP have established a robust foundation for this task by learning to match text and image representations from large-scale data, their lack of pixel-level recognition necessitates further fine-tuning. Most existing methods leverage text as a guide to achieve pixel-level recognition. However, the inherent biases in text semantic descriptions and the lack of pixel-level supervisory information make it challenging to fine-tune CLIP-based models effectively. This paper considers leveraging image-text data to simultaneously capture the semantic information contained in both image and text, thereby constructing Dual Semantic Guidance and corresponding pixel-level pseudo annotations. Particularly, the visual semantic guidance is enhanced via explicitly exploring foreground regions and minimizing the influence of background. The dual semantic guidance is then jointly utilized to fine-tune CLIP-based segmentation models, achieving decent fine-grained recognition capabilities. As the comprehensive evaluation shows, our method outperforms state-of-art results with large margins, on eight commonly used datasets with/without background.
Tingliang Feng, Fan Lyu, Fanhua Shang, Wei Feng 0005
CVPR3
2025 Beyond Background Shift: Rethinking Instance Replay in Continual Semantic Segmentation
abstract
In this work, we focus on continual semantic segmentation (CSS), where segmentation networks are required to continuously learn new classes without erasing knowledge of previously learned ones. Although storing images of old classes and directly incorporating them into the training of new models has proven effective in mitigating catastrophic forgetting in classification tasks, this strategy presents notable limitations in CSS. Specifically, the stored and new images with partial category annotations leads to confusion between unannotated categories and the background, complicating model fitting. To tackle this issue, this paper proposes a novel Enhanced Instance Replay (EIR) method, which not only preserves knowledge of old classes while simultaneously eliminating background confusion by instance storage of old classes, but also mitigates background shifts in the new images by integrating stored instances with new images. By effectively resolving background shifts in both stored and new images, EIR alleviates catastrophic forgetting in the CSS task, thereby enhancing the model’s capacity for CSS. Experimental results validate the efficacy of our approach, which significantly outperforms state-of-the-art CSS methods. The code is available at https://github.com/YikeYin97/EIR.
Hongmei Yin, Tingliang Feng, Fan Lyu, Fanhua Shang, Hongying Liu 0001, Wei Feng 0005
CVPR3
2025 Controllable Continual Test-Time Adaptation
abstract
Continual Test-Time Adaptation (CTTA) is an emerging and challenging task where a model trained in a source domain must adapt to continuously changing conditions during testing, without access to the original source data. CTTA is prone to error accumulation due to uncontrollable domain shifts, leading to blurred decision boundaries between categories. Existing CTTA methods primarily focus on suppressing domain shifts, which proves inadequate during the unsupervised test phase. In contrast, we introduce a novel approach that guides rather than suppresses these shifts. Specifically, we propose Controllable Continual Test-Time Adaptation (C-CoTTA), which explicitly prevents any single category from encroaching on others, thereby mitigating the mutual influence between categories caused by uncontrollable shifts. Moreover, our method reduces the sensitivity of model to domain transformations, thereby minimizing the magnitude of category shifts. Extensive quantitative experiments demonstrate the effectiveness of our method, while qualitative analyses, such as t-SNE plots, confirm the theoretical validity of our approach. Our code is available at https://github.com/RenshengJi/C-CoTTA.
Ziqi Shi, Fan Lyu, Fanhua Shang, Fuyuan Hu, Wei Feng 0005, Zhang Zhang 0001, Liang Wang 0001
ICME2
2025 DAA: Amplifying Unknown Discrepancy for Test-Time Discovery
abstract
Test-Time Discovery (TTD) addresses the critical challenge of identifying and adapting to novel classes during inference while maintaining performance on known classes, which is a capability essential for dynamic real-world environments such as healthcare and autonomous driving. Recent TTD methods adopt training-free, memory-based strategies but rely on frozen models and static representations, resulting in poor generalization. In this paper, we propose a Discrepancy-Amplifying Adapter (DAA), a trainable module that enables real-time adaptation by amplifying feature-level discrepancies between known and unknown classes. During training, DAA is optimized using simulated unknowns and a novel warm-up strategy to enhance its discriminative capacity. To ensure continual adaptation at test time, we introduce a Short-Term Memory Renewal (STMR) mechanism, which maintains a queue-based memory for unknown classes and selectively refreshes prototypes using recent, reliable samples. DAA is further updated through self-supervised learning, promoting knowledge retention for known classes while improving discrimination of emerging categories. Extensive experiments show that our method maintains high adaptability and stability, and significantly improves novel class discovery performance. Our code will be available.
Fan Lyu, Chenggong Ni, Zhang Zhang 0001, Fuyuan Hu, Liang Wang 0001
NeurIPS2
2025 Partition-Then-Adapt: Combating Prediction Bias for Reliable Multi-Modal Test-Time Adaptation
abstract
Existing test-time adaptation (TTA) methods primarily focus on scenarios involving domain shifts in a single modality. However, they often prove ineffective when multiple modalities simultaneously undergo domain shifts, as they struggle to identify and utilize reliable samples within testing batches amid severe prediction bias. To address this problem, we propose Partition-Then-Adapt (PTA), a novel approach combating prediction bias for TTA with multi-modal domain shifts. PTA comprises two key components: Partition and Debiased Reweighting (PDR) and multi-modal Attention-Guided Alignment (AGA). Specifically, PDR evaluates each sample’s predicted label frequency relative to the batch average, partitioning the batch into potential reliable and unreliable subsets. It then reweights each sample by jointly assessing its bias and confidence levels through a quantile-based approach. By applying weighted entropy loss, PTA simultaneously promotes learning from reliable subsets and discourages reliance on unreliable ones. Moreover, AGA regularizes PDR to focus on semantically meaningful multi-modal cues. Extensive experiments validate the effectiveness of PTA, surpassing state-of-the-art method by 6.1\% on Kinetics50-MC and 5.8\% on VGGSound-MC, respectively. Code of this paper is available at https://github.com/MPI-Lab/PTA.
Fan Lyu, Changxing Ding
NeurIPS2
2025 Few-Shot Class-Incremental Learning via Asymmetric Supervised Contrastive Learning
abstract
Few-Shot Class-Incremental Learning (FSCIL) is to continuously learn novel classes from a few samples without forgetting previous knowledge. Adapting directly to limited novel data typically results in significant forgetting of base class knowledge. Consequently, prevailing FSCIL methods are devoted to training a strong initial model that can be frozen in incremental sessions. However, these works face a dilemma in poor generalization: they benefit mainly from base class performance, yet underperform in novel classes. To alleviate this issue, we design a two-stage training framework to simultaneously enhance generalization for novel classes and maintain base class discrimination. In the first stage, an asymmetric supervised contrastive learning (AsyCon) algorithm is proposed. AsyCon introduces a predicted feature to achieve an asymmetric alignment of positive pairs. It alleviates over-similarity within positive features, allowing the model to better transfer to new classes in incremental sessions. In the second stage, the model is finetuned for promoting its performance on base classes. To maintain the generalization obtained in the first stage, we employ an L2 normalized regularization (LR) to keep the feature consistent with the model in the first stage. The finetuned model, termed AsyCLR, effectively balances generalization and discrimination, significantly outperforming existing FSCIL works especially in novel class accuracy. Experiments on CUB200, CIFAR100, and mini-ImageNet verify the effectiveness of our method. Additionally, our method also performs well in the standard few-shot recognition scenario due to its strong generalization ability. Our codes are available at https://github.com/APORduo/AsyCLR.
Duo Liu 0001, Linglan Zhao, Fan Lyu, Xiangzhong Fang, Liang Wang 0001
IEEE Trans. Circuits Syst. Video Technol.4
2024 Long-Tailed Learning as Multi-Objective Optimization
abstract
Real-world data is extremely imbalanced and presents a long-tailed distribution, resulting in models biased towards classes with sufficient samples and performing poorly on rare classes. Recent methods propose to rebalance classes but they undertake the seesaw dilemma (what is increasing performance on tail classes may decrease that of head classes, and vice versa). In this paper, we argue that the seesaw dilemma is derived from the gradient imbalance of different classes, in which gradients of inappropriate classes are set to important for updating, thus prone to overcompensation or undercompensation on tail classes. To achieve ideal compensation, we formulate long-tailed recognition as a multi-objective optimization problem, which fairly respects the contributions of head and tail classes simultaneously. For efficiency, we propose a Gradient-Balancing Grouping (GBG) strategy to gather the classes with similar gradient directions, thus approximately making every update under a Pareto descent direction. Our GBG method drives classes with similar gradient directions to form a more representative gradient and provides ideal compensation to the tail classes. Moreover, we conduct extensive experiments on commonly used benchmarks in long-tailed learning and demonstrate the superiority of our method over existing SOTA methods. Our code is released at https://github.com/WickyLee1998/GBG_v1.
Fan Lyu, Fanhua Shang, Wei Feng 0005
AAAI2
2024 Parameter-Selective Continual Test-Time Adaptation
Jiaxu Tian, Fan Lyu
ACCV (8)2
2024 Confidence Self-calibration for Multi-label Class-Incremental Learning
Kaile Du, Fan Lyu, Yuyang Li 0005, Chen Lu 0004, Guangcan Liu
ECCV (31)3
2024 MetaMask: Improving Few-Shot Semantic Segmentation via Multi-Mask Calibriation
abstract
Few-shot Semantic Segmentation (FSS) aims to develop models that can segment previously unseen classes with only a few annotations. Recent approaches employ a "multi-mask" framework, which initially generates various mask proposals from query images and then matches related mask proposals to get the final output guided by support images. Despite its promise, this framework is limited by the quality of mask proposals for unseen classes and a naive mask matching process. To address such limitations, in this paper, we propose a meta-learning-based method called MetaMask. First, MetaMask builds a Support-Guided Latent Object Segmenter (SG-LOS) module, which incorporates unseen class information into mask proposal generation for query images, where episodic training is used to enhance mask generation for latent unseen classes. Second, MetaMask improves the mask-matching mechanism through our proposed Contrastive Mask Matching (CMM) module with a cross-image multi-level contrastive learning strategy, bolstering feature embedding spaces. Our method shows competitive results on two main benchmarks: 69.9% mIoU on Pascal-5ione-shot setting and 49.6% mIoU COCO-20ione-shot setting, marginally outperforming our baseline by 6.6% and 5.4%, setting a new state-of-the-art on the both Pascal-5iand COCO-20idatasets.
Li Dinghang, Zongqing Lu 0001, Weiliang Zheng, Qingmin Liao, Fan Lyu
IJCNN5
2024 Understanding Driving Risks via Prompt Learning
abstract
Understanding driving risks is crucial for enhancing driving safety. It is a challenging task to evaluate driving risks in various complex driving scenarios. Inspired by prompt-based learning, we propose an end-to-end approach for identifying the highest-risk object in the current driving scenario based on a learnable risk pool. Specifically, a method based on key-value pair matching is designed to build a memory system for learning a collection of risk prototypes. Extensive experiments on the DRAMA dataset show that the proposed method achieves an improvement of 18.6% in Mean-IOU and 3.0% in B4 score compared to the state-of-the-art (SOTA) methods, which indicates that our method can effectively localize risky objects and accurately describe the driving scenes.
Yubo Chang, Fan Lyu, Zhang Zhang 0001, Liang Wang 0001
SMC2
2024 Combinational sign language recognition
Liqing Gao, Wei Feng 0005, Fan Lyu, Liang Wang 0001
Comput. Vis. Image Underst.3
2024 Overcoming Modality Bias in Question-Driven Sign Language Video Translation
abstract
Question-Driven Sign Language Translation (QSLT) addresses the challenge of translating sign language using pertinent questions in question-answering contexts. However, the pronounced modality complexity between question text and sign video poses a predicament: the model tends to overly depend on questions to generate translations, thereby neglecting the value of visual cues. To tackle this issue, the paper presents a Gloss-Bridged Translator (GBT), which introduces sign gloss as an intermediary conduit to establish semantic connections between questions and videos. By leveraging gloss, visual features are transformed into textual counterparts, mitigating the modality imbalance between these representations. Moreover, a cross-modal contrastive learning strategy is implemented, bolstering the global contextual relevance and local semantic alignment between questions and sign language. The proposed methodology is validated through extensive experiments on the proposed QSL dataset and other public sign language datasets. The results show the efficacy of integrating questions into sign language translation. The GBT yields remarkable improvements over prevailing SLT methods, attesting to its effectiveness and rationale. Our code and dataset is available athttps://github.com/glq-1992/QSL.
Liqing Gao, Fan Lyu, Lei Zhu 0003, Junfu Pu, Liang Wang 0001, Wei Feng 0005
IEEE Trans. Circuits Syst. Video Technol.2
2024 Multi-Label Continual Learning Using Augmented Graph Convolutional Network
abstract
Multi-Label Continual Learning (MLCL) is a framework designed for class-incremental multi-label image recognition. However, MLCL faces two critical challenges: the construction of label relationships onpast-missing and future-missing partial labelsof training data, and the problem ofcatastrophic forgetting, which leads to poor generalization. To address these challenges, this study proposes an enhanced version of the Augmented Graph Convolutional Network (AGCN++), capable of constructing cross-task label relationships and mitigating catastrophic forgetting. First, an Augmented Correlation Matrix (ACM) is constructed across all observed classes, incorporating intra-task relationships derived from hard label statistics. Additionally, inter-task relationships are established by leveraging both hard and soft labels obtained from the data, as well as a constructed expert network. Next, a novel partial label encoder (PLE) is introduced for MLCL, enabling the extraction of dynamic class representations for each partial label image as graph nodes. This PLE also facilitates the generation of soft labels, which contribute to the creation of a more persuasive ACM and effectively mitigate forgetting. Lastly, a relationship-preserving constrainter is proposed to address the issue of forgetting label dependencies across old tasks. In the AGCN++, the label relationships topology can be augmented automatically, thereby generating efficient class representations. The effectiveness of the proposed method is evaluated using two multi-label image benchmarks. The experimental results demonstrate that the proposed approach is highly effective in the context of MLCL image recognition. It can establish compelling correlations across tasks, even in scenarios where the old task labels are missing.
Kaile Du, Fan Lyu, Fuyuan Hu, Wei Feng 0005, Fenglei Xu, Hanjing Cheng
IEEE Trans. Multim.2
2023 Centroid Distance Distillation for Effective Rehearsal in Continual Learning
abstract
Rehearsal, retraining on a stored small data subset of old tasks, has been proven effective in solving catastrophic forgetting in continual learning. However, due to the sampled data may have a large bias towards the original dataset, retraining them is susceptible to driving continual domain drift of old tasks in feature space, resulting in forgetting. In this paper, we focus on tackling the continual domain drift problem with centroid distance distillation. First, we propose a centroid caching mechanism for sampling data points based on constructed centroids to reduce the sample bias in rehearsal. Then, we present a centroid distance distillation that only stores the centroid distance to reduce the continual domain drift. The experiments on four continual learning datasets show the superiority of the proposed method, and the continual domain drift can be reduced. Our code is available at https://github.com/Daofeng-liu/CDD-R.
Daofeng Liu, Fan Lyu, Zhenping Xia, Fuyuan Hu
ICASSP2
2023 Measuring Asymmetric Gradient Discrepancy in Parallel Continual Learning
abstract
In Parallel Continual Learning (PCL), the parallel multiple tasks start and end training unpredictably, thus suffering from both training conflict and catastrophic forgetting issues. The two issues are raised because the gradients from parallel tasks differ in directions and magnitudes. Thus, in this paper, we formulate the PCL into a minimum distance optimization problem among gradients and propose an explicit Asymmetric Gradient Distance (AGD) to evaluate the gradient discrepancy in PCL. AGD considers both gradient magnitude ratios and directions, and has a tolerance when updating with a small gradient of inverse direction, which reduces the imbalanced influence of gradients on parallel task training. Moreover, we present a novel Maximum Discrepancy Optimization (MaxDO) strategy to minimize the maximum discrepancy among multiple gradients. Solving by MaxDO with AGD, parallel training reduces the influence of the training conflict and suppresses the catastrophic forgetting of finished tasks. Extensive experiments validate the effectiveness of our approach on three image recognition datasets in task-incremental and class-incremental PCL. Our code is available at https://github.com/fanlyu/maxdo.
Fan Lyu, Fanhua Shang, Wei Feng 0005
ICCV1
2023 Multi-semantic hypergraph neural network for effective few-shot learning
Hao Chen 0011, Fuyuan Hu, Fan Lyu, Liuqing Zhao, Kaizhu Huang, Wei Feng 0005, Zhenping Xia
Pattern Recognit.4
2023 Difference-guided multi-scale spatial-temporal representation for sign language recognition
Liqing Gao, Lianyu Hu 0003, Fan Lyu, Lei Zhu 0003, Chi-Man Pun, Wei Feng 0005
Vis. Comput.3
2022 AGCN: Augmented Graph Convolutional Network for Lifelong Multi-Label Image Recognition
abstract
The Lifelong Multi-Label (LML) image recognition builds an online class-incremental classifier in a sequential multilabel image recognition data stream. However, training on the data with different Partial Labels may result in more serious Catastrophic Forgetting in old classes. To solve the problem, the study proposes an Augmented Graph Convolutional Network (AGCN)to build an Augmented Correlation Matrix (ACM) across the sequential partial-label tasks and sustain the catastrophic forgetting. First, in ACM, the intra-task relations derive from the hard label statistics, while the inter-task relations further leverage the soft labels from a stored expert network. Then, based on the ACM, AGCN captures label dependencies with dynamic augmented structure and yields effective class representations. Our method is evaluated on two multi-label image benchmarks and the results show that the proposed method is effective for LML image recognition.
Kaile Du, Fan Lyu, Fuyuan Hu, Wei Feng 0005, Fenglei Xu, Qiming Fu 0001
ICME2
2022 Exploring Example Influence in Continual Learning
abstract
Continual Learning (CL) sequentially learns new tasks like human beings, with the goal to achieve better Stability (S, remembering past tasks) and Plasticity (P, adapting to new tasks). Due to the fact that past training data is not available, it is valuable to explore the influence difference on S and P among training examples, which may improve the learning pattern towards better SP. Inspired by Influence Function (IF), we first study example influence via adding perturbation to example weight and computing the influence derivation. To avoid the storage and calculation burden of Hessian inverse in neural networks, we propose a simple yet effective MetaSP algorithm to simulate the two key steps in the computation of IF and obtain the S- and P-aware example influence. Moreover, we propose to fuse two kinds of example influence by solving a dual-objective optimization problem, and obtain a fused influence towards SP Pareto optimality. The fused influence can be used to control the update of model and optimize the storage of rehearsal. Empirical results show that our algorithm significantly outperforms state-of-the-art methods on both task- and class-incremental benchmark CL datasets.
Fan Lyu, Fanhua Shang, Wei Feng 0005
NeurIPS2
2022 Harnessing Multi-Semantic Hypergraph for Few-Shot Learning
Hao Chen 0011, Zhenping Xia, Fan Lyu, Liuqing Zhao, Kaizhu Huang, Wei Feng 0005, Fuyuan Hu
PRCV (1)4
2022 Visual Grounding Via Accumulated Attention
abstract
Visual grounding (VG) aims to locate the most relevant object or region in an image, based on a natural language query. Generally, it requires the machine to first understand the query, identify the key concepts in the image, and then locate the target object by specifying its bounding box. However, in many real-world visual grounding applications, we have to face with ambiguous queries and images with complicated scene structures. Identifying the target based on highly redundant and correlated information can be very challenging, and often leading to unsatisfactory performance. To tackle this, in this paper, we exploit an attention module for each kind of information to reduce internal redundancies. We then propose an accumulated attention (A-ATT) mechanism to reason among all the attention modules jointly. In this way, the relation among different kinds of information can be explicitly captured. Moreover, to improve the performance and robustness of our VG models, we additionally introduce some noises into the training procedure to bridge the distribution gap between the human-labeled training data and the real-world poor quality data. With this "noised" training strategy, we can further learn a bounding box regressor, which can be used to refine the bounding box of the target object. We evaluate the proposed methods on four popular datasets (namely ReferCOCO, ReferCOCO+, ReferCOCOg, and GuessWhat?!). The experimental results show that our methods significantly outperform all previous works on every dataset in terms of accuracy.
Chaorui Deng, Qi Wu 0001, Qingyao Wu, Fuyuan Hu, Fan Lyu, Mingkui Tan
IEEE Trans. Pattern Anal. Mach. Intell.5
2022 Disentangling Semantic-to-Visual Confusion for Zero-Shot Learning
abstract
Using generative models to synthesize visual features from semantic distribution is one of the most popular solutions to ZSL image classification in recent years. The triplet loss (TL) is popularly used to generate realistic visual distributions from semantics by automatically searching discriminative representations. However, the traditional TL cannot search reliable unseen disentangled representations due to the unavailability of unseen classes in ZSL. To alleviate this drawback, we propose in this work a multi-modal triplet loss (MMTL) which utilizes multi-modal information to search adisentangledrepresentation space. As such, all classes can interplay which can benefit learning disentangled class representations in the searched space. Furthermore, we develop a novel model called Disentangling Class Representation Generative Adversarial Network (DCR-GAN) focusing on exploiting the disentangled representations in training, feature synthesis, and final recognition stages. Benefiting from the disentangled representations, DCR-GAN could fit a more realistic distribution over both seen and unseen features. Extensive experiments show that our proposed model can lead to superior performance to the state-of-the-arts on four benchmark datasets.
Zihan Ye, Fuyuan Hu, Fan Lyu, Kaizhu Huang
IEEE Trans. Multim.3
2022 Fine-grained scale space learning for single image super-resolution
Fan Lyu, Wei Feng 0005
Vis. Comput.3
2021 Multi-Domain Multi-Task Rehearsal for Lifelong Learning
abstract
Rehearsal, seeking to remind the model by storing old knowledge in lifelong learning, is one of the most effective ways to mitigate catastrophic forgetting, i.e., biased forgetting of previous knowledge when moving to new tasks. However, the old tasks of the most previous rehearsal-based methods suffer from the unpredictable domain shift when training the new task. This is because these methods always ignore two significant factors. First, the Data Imbalance between the new task and old tasks that makes the domain of old tasks prone to shift. Second, the Task Isolation among all tasks will make the domain shift toward unpredictable directions; To address the unpredictable domain shift, in this paper, we propose Multi-Domain Multi-Task (MDMT) rehearsal to train the old tasks and new task parallelly and equally to break the isolation among tasks. Specifically, a two-level angular margin loss is proposed to encourage the intra-class/task compactness and inter-class/task discrepancy, which keeps the model from domain chaos. In addition, to further address domain shift of the old tasks, we propose an optional episodic distillation loss on the memory to anchor the knowledge for each old task. Experiments on benchmark datasets validate the proposed approach can effectively mitigate the unpredictable domain shift.
Fan Lyu, Wei Feng 0005, Zihan Ye, Fuyuan Hu, Song Wang 0002
AAAI1
2021 Each Attribute Matters: Contrastive Attention for Sentence-based Image Editing
Liuqing Zhao, Fan Lyu, Fuyuan Hu, Kaizhu Huang, Fenglei Xu
BMVC2
2020 Associating Multi-Scale Receptive Fields For Fine-Grained Recognition
abstract
Extracting and fusing part features have become the key of fined-grained image recognition. Recently, Non-local (NL) module has shown excellent improvement in image recognition. However, it lacks the mechanism to model the interactions between multi-scale part features, which is vital for fine-grained recognition. In this paper, we propose a novel cross-layer non-local (CNL) module to associate multi-scale receptive fields by two operations. First, CNL computes correlations between features of a query layer and all response layers. Second, all response features are weighted according to the correlations and are added to the query features. Due to the interactions of cross-layer features, our model builds spatial dependencies among multi-level layers and learns more discriminative features. In addition, we can reduce the aggregation cost if we set low-dimensional deep layer as query layer. Experiments are conducted to show our model achieves or surpasses state-of-the-art results on three benchmark datasets of fine-grained classification. Our codes can be found at github.com/FouriYe/CNL-ICIP2020.
Zihan Ye, Fuyuan Hu, Zhenping Xia, Fan Lyu, Pengqing Liu
ICIP5
2020 Modeling Cross-View Interaction Consistency for Paired Egocentric Interaction Recognition
abstract
With the development of Augmented Reality (AR), egocentric action recognition (EAR) plays an important role in accurately understanding demands from the user. However, EAR is designed to help recognize human-machine interaction in single egocentric view, thus difficult to capture interactions between two face-to-face AR users. Paired egocentric interaction recognition (PEIR) is the task to collaboratively recognize the interactions between two persons with the videos in their corresponding views. Unfortunately, existing PEIR methods always directly use linear decision function to fuse the features extracted from two corresponding egocentric videos, which ignore the consistency of interaction in paired egocentric videos. The consistency of interactions in paired videos, and features extracted from them, are correlated to each other. On top of that, we propose to derive the relevance between two views using bilinear pooling, which captures the consistency of two views in feature-level. Specifically, each neuron in the feature maps from one view connects to the neurons from the other view, which enforces the compact consistency between two views and then all possible paired neurons are used for PEIR. To be efficient, we use compact bilinear pooling with Count Sketch to avoid directly computing outer product. Experimental results on the PEV dataset shows the superiority of the proposed methods on the task PEIR.
Zhongguo Li, Fan Lyu, Wei Feng 0005, Song Wang 0002
ICME2
2020 Mutatt: Visual-Textual Mutual Guidance For Referring Expression Comprehension
abstract
Referring expression comprehension (REC) aims to localize a text-related region in a given image by a referring expression in natural language. Existing methods focus on how to build convincing visual and language representations independently, which may significantly isolate visual and language information. In this paper, we argue that for REC the referring expression and the target region are semantically correlated and subject, location and relationship consistency exist between vision and language. On top of this, we propose a novel approach called MutAtt to construct mutual guidance between vision and language, which treats vision and language equally thus yields compact information matching. Specifically, for each module of subject, location and relationship, MutAtt builds two kinds of attention-based mutual guidance strategies. One strategy is to generate vision-guided language embedding for the sake of matching relevant visual features. The other reversely generates language-guided visual features to match relevant language embedding. This mutual guidance strategy can effectively enforce the vision-language consistency in three modules. Experiments on three popular REC datasets demonstrate that the proposed approach outperforms the current state-of-the-art methods.
Fan Lyu, Wei Feng 0005, Song Wang 0002
ICME2
2020 vtGraphNet: Learning weakly-supervised scene graph for complex visual grounding
Fan Lyu, Wei Feng 0005, Song Wang 0002
Neurocomputing1
2019 SR-GAN: Semantic Rectifying Generative Adversarial Network for Zero-shot Learning
abstract
The existing Zero-Shot learning (ZSL) methods may suffer from the vague class attributes that are highly overlapped for different classes. Unlike these methods that ignore the discrimination among classes, in this paper, we propose to classify unseen image by rectifying the semantic space guided by the visual space. First, we pre-train a Semantic Rectifying Network (SRN) to rectify semantic space with a semantic loss and a rectifying loss. Then, a Semantic Rectifying Generative Adversarial Network (SR-GAN) is built to generate plausible visual feature of unseen class from both semantic feature and rectified semantic feature. To guarantee the effectiveness of rectified semantic features and synthetic visual features, a pre-reconstruction and a post reconstruction networks are proposed, which keep the consistency between visual feature and semantic feature. Experimental results demonstrate that our approach significantly outperforms the state-of-the-arts on four benchmark datasets.
Zihan Ye, Fan Lyu, Qiming Fu 0001, Jinchang Ren, Fuyuan Hu
ICME2
2019 Attend and Imagine: Multi-Label Image Classification With Visual Attention and Recurrent Neural Networks
abstract
Real images often have multiple labels, i.e., each image is associated with multiple objects or attributes. Compared to single-label image classification, the multilabel classification problem is much more challenging due to several issues. At first, multiple objects can be anywhere in the image. Second, the importance of different regions in an image is different, and the regions of interest in a multilabel image can be very different from another one. Finally, multiple labels of an image can have label dependencies due to complex image structures. To address these challenges, in this paper, we propose to predict the labels sequentially by applying the recurrent neural networks (RNNs), which are used to encode the label dependencies. When predicting a specific label, we introduce a dynamic attention mechanism to enable the model to focus on only regions of interest in the image. Two benchmark datasets (i.e., Pascal VOC and MS-COCO) are adopted to demonstrate the effectiveness of our work. Moreover, we construct a new dataset, which includes many semantic dependent labels in each image, to verify the effectiveness of our model. Experimental results show that our method outperforms several state-of-the-arts, especially when predicting some semantic relative labels.
Fan Lyu, Qi Wu 0001, Fuyuan Hu, Qingyao Wu, Mingkui Tan
IEEE Trans. Multim.1
2018 Visual Grounding via Accumulated Attention
abstract
Visual Grounding (VG) aims to locate the most relevant object or region in an image, based on a natural language query. The query can be a phrase, a sentence or even a multi-round dialogue. There are three main challenges in VG: 1) what is the main focus in a query; 2) how to understand an image; 3) how to locate an object. Most existing methods combine all the information curtly, which may suffer from the problem of information redundancy (i.e. ambiguous query, complicated image and a large number of objects). In this paper, we formulate these challenges as three attention problems and propose an accumulated attention (A-ATT) mechanism to reason among them jointly. Our A-ATT mechanism can circularly accumulate the attention for useful information in image, query, and objects, while the noises are ignored gradually. We evaluate the performance of A-ATT on four popular datasets (namely Refer-COCO, ReferCOCO+, ReferCOCOg, and Guesswhat?!), and the experimental results show the superiority of the proposed method in term of accuracy.
Chaorui Deng, Qi Wu 0001, Qingyao Wu, Fuyuan Hu, Fan Lyu, Mingkui Tan
CVPR5