Zihuan Qiu

dblp:310/1853 · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
10since 2021 · last 2026
0000-0001-7281-3885ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021
YearPublicationVenuePosition
2026 DesCLIP: Robust Continual Learning via General Attribute Descriptions for VLM-Based Visual Recognition
abstract
Continual learning of vision-language models (VLMs) focuses on leveraging cross-modal pretrained knowledge to incrementally adapt to expanding downstream tasks and datasets, while tackling the challenge of knowledge forgetting. Existing research often focuses on connecting visual features with specific class text in downstream tasks, overlooking the latent relationships between general and specialized knowledge. Our findings reveal that forcing models to optimize inappropriate visual-text matches exacerbates forgetting of VLM's recognition ability. To tackle this issue, we propose DesCLIP, which leverages general attribute (GA) descriptions to guide the understanding of specific class objects, enabling VLMs to establish robust <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">vision-GA-class</i> trilateral associations rather than relying solely on <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">vision-class</i> connections. Specifically, we introduce a language assistant to generate concrete GA description candidates via proper request prompts. Then, an anchor-based embedding filter is designed to obtain highly relevant GA description embeddings, which are leveraged as the paired text embeddings for visual-textual instance matching, thereby tuning the visual encoder. Correspondingly, the class text embeddings are gradually calibrated to align with these shared GA description embeddings. Extensive experiments demonstrate the advancements and efficacy of our proposed method, with comprehensive empirical evaluations highlighting its superior performance in VLM-based recognition compared to existing continual learning methods.
Chiyuan He, Zihuan Qiu, Fanman Meng, Linfeng Xu 0001, Qingbo Wu 0001, Hongliang Li 0001
IEEE Trans. Multim.2
2025 Leveraging Pre-Trained Models for Multimodal Class-Incremental Learning under Adaptive Fusion
abstract
Unlike traditional Multimodal Class-Incremental Learning (MCIL) methods that focus only on vision and text, this paper explores MCIL across vision, audio and text modalities, addressing challenges in integrating complementary information and mitigating catastrophic forgetting. To tackle these issues, we propose an MCIL method based on multimodal pre-trained models. Firstly, a Multimodal Incremental Feature Extractor (MIFE) based on Mixture-of-Experts (MoE) structure is introduced to achieve effective incremental fine-tuning for AudioCLIP. Secondly, to enhance feature discriminability and generalization, we propose an Adaptive Audio-Visual Fusion Module (AAVFM) that includes a masking threshold mechanism and a dynamic feature fusion mechanism, along with a strategy to enhance text diversity. Thirdly, a novel multimodal class-incremental contrastive training loss is proposed to optimize cross-modal alignment in MCIL. Finally, two MCIL-specific evaluation metrics are introduced for comprehensive assessment. Extensive experiments on three multimodal datasets validate the effectiveness of our method.
Zihuan Qiu, Fanman Meng, Hongliang Li 0001, Linfeng Xu 0001, Qingbo Wu 0001
ICASSP2
2025 DPM-CLIP: Zero-Shot Multimodal Egocentric Activity Recognition based on Dual-Prediction Mechanism
abstract
Advancements in Zero-shot Multimodal Egocentric Activity Recognition (ZS-MM-EAR) largely rely on Vision-Language Model (VLM). However, existing methods struggle with VLM’s inadequate representation of egocentric activities, including challenges in capturing egocentric-specific features, adapting to domain shifts between egocentric video and pre-training data, and effectively leveraging complementary data such as Inertial Measurement Unit (IMU). To address these issues, we propose DPM-CLIP, a ZS-MM-EAR method tailored for vision, text and IMU modalities. Firstly, we design an attribute-driven text augmentation module that leverages a Large Language Model (LLM) to generate fine-grained textual descriptions of activities. Secondly, we construct an Instance-feature Repository (IFR) to store base class features and generate pseudo-features for novel classes through feature center migration. Finally, we introduce a dual-prediction mechanism with a prediction correction module to enhance generalization and recognition accuracy. Extensive experiments on the UESTC-MMEA-CL dataset validate the effectiveness of the proposed method.
Zihuan Qiu, Mingzhou He, Fanman Meng, Linfeng Xu 0001, Qingbo Wu 0001, Hongliang Li 0001
ICIP3
2025 MINGLE: Mixture of Null-Space Gated Low-Rank Experts for Test-Time Continual Model Merging
abstract
Continual model merging integrates independently fine-tuned models sequentially without access to the original training data, offering a scalable and efficient solution for continual learning. However, existing methods face two critical challenges: parameter interference among tasks, which leads to catastrophic forgetting, and limited adaptability to evolving test distributions. To address these issues, we introduce the task of Test-Time Continual Model Merging (TTCMM), which leverages a small set of unlabeled test samples during inference to alleviate parameter conflicts and handle distribution shifts. We propose MINGLE, a novel framework for TTCMM. MINGLE employs a mixture-of-experts architecture with parameter-efficient, low-rank experts, which enhances adaptability to evolving test distributions while dynamically merging models to mitigate conflicts. To further reduce forgetting, we propose Null-Space Constrained Gating, which restricts gating updates to subspaces orthogonal to prior task representations, thereby suppressing activations on old tasks and preserving past knowledge. We further introduce an Adaptive Relaxation Strategy that adjusts constraint strength dynamically based on interference signals observed during test-time adaptation, striking a balance between stability and adaptability. Extensive experiments on standard continual merging benchmarks demonstrate that MINGLE achieves robust generalization, significantly reduces forgetting, and consistently surpasses previous state-of-the-art methods by 7–9% on average across diverse task orders. Our code is available at: https://github.com/zihuanqiu/MINGLE
Zihuan Qiu, Yi Xu 0008, Chiyuan He, Fanman Meng, Linfeng Xu 0001, Qingbo Wu 0001, Hongliang Li 0001
NeurIPS1
2025 Distribution-Level Memory Recall for Continual Learning: Preserving Knowledge and Avoiding Confusion
abstract
Continual learning (CL) aims to enable deep neural networks (DNNs) to learn new data without forgetting previously learned knowledge. The key to achieving this goal is to avoid confusion at the feature level, i.e., to avoid confusion within old tasks and between new and old tasks. Existing prototype-based CL methods generate pseudo features for old knowledge replay by adding Gaussian noise to the centroids of old classes. However, the distribution in the feature space exhibits anisotropy during the incremental process, which prevents the pseudo features from faithfully reproducing the distribution of old knowledge in the feature space, leading to confusion at the classification boundaries within old tasks. To address this issue, we propose the distribution-level memory recall (DMR) method, which uses a Gaussian mixture model to precisely fit the feature distribution of old knowledge at the distribution level and generate pseudo features in the next stage. Furthermore, resistance to confusion at the distribution level is crucial for multimodal learning. Multimodal imbalance, which refers to uneven optimization processes among encoders of different modalities, results in significant differences in feature responses between modalities; this exacerbates confusion within old tasks in prototype-based CL methods. Therefore, we mitigate the multimodal imbalance problem by using the intermodal guidance and intramodal mining (IGIM) method to guide weaker modalities with prior information from dominant modalities and further explore useful information within modalities. To avoid confusion between new and old tasks, we propose using the confusion index to quantitatively describe a model's ability to distinguish between new and old tasks, and we use the incremental mixup feature enhancement (IMFE) method to enhance pseudo features with new sample features, alleviating classification confusion between new and old knowledge. We conduct extensive experiments on the CIFAR100, ImageNet100, TinyImageNet, ImageNet-1K and UESTC-MMEA-CL datasets and achieve state-of-the-art results.
Shaoxu Cheng, Kanglei Geng, Chiyuan He, Zihuan Qiu, Linfeng Xu 0001, Heqian Qiu, Lanxiao Wang, Qingbo Wu 0001, Fanman Meng, Hongliang Li 0001
IEEE Trans. Multim.4
2024 Dual-Consistency Model Inversion for Non-Exemplar Class Incremental Learning
abstract
Non-exemplar class incremental learning (NECIL) aims to continuously assimilate new knowledge without forgetting previously acquired ones when historical data are un-available. One of the generative NECIL methods is to in-vert the images of old classes for joint training. However, these synthetic images suffer significant domain shifts compared with real data, hampering the recognition of old classes. In this paper, we present a novel method termed Dual-Consistency Model Inversion (DCMI) to generate better synthetic samples of old classes through two pivotal consistency alignments: (1) the semantic consistency between the synthetic images and the corresponding prototypes, and (2) domain consistency between synthetic and real images of new classes. Besides, we introduce Prototypical Routing (PR) to provide task-prior information and generate unbi-ased and accurate predictions. Our comprehensive experiments across diverse datasets consistently showcase the superiority of our method over previous state-of-the-art approaches.
Zihuan Qiu, Yi Xu 0008, Fanman Meng, Hongliang Li 0001, Linfeng Xu 0001, Qingbo Wu 0001
CVPR1
2023 MFAT: A Multi-Level Feature Aggregated Transformer for Person Re-Identification
abstract
Recently, with the development of the Transformer, re-identification (ReID) has great success in various applications. Existing works prefer to utilize the Transformer’s highest-level information as its discriminative feature, which focuses on a few concentrated parts or areas. However, in ReID filed, under such various scenes and camera views, only using a few concentrated parts to distinguish the query person is insufficient. Meanwhile, we find that Transformer’s lower-level information is also helpful for the recognition accuracy of the query person, especially, when the scene changes greatly. Therefore, we propose a Multi-level Feature Aggregated Transformer for person re-identification (MFAT) with high performance. To aggregate multi-level information, two novel modules are carefully designed. (i) The Global Content and Structure Aggregation (GCSA) module is proposed to aggregate multi-level information in a global manner. (ii) The Local Convolution Aggregation (LCA) module which consists of a series of convolutional blocks, is introduced to aggregate multi-level features with local operations. To the best of our knowledge, this is the first work to aggregate multi-level features with a Transformer backbone for person ReID task. Experiment results show that our method has achieved state-of-the-art on three person ReID benchmarks, with both Pyramid Vision Transformer (PVT) and Vision Transformer (ViT) backbones.
Bowen Tan, Linfeng Xu 0001, Zihuan Qiu, Qingbo Wu 0001, Fanman Meng
ICASSP3
2023 ISM-Net: Mining incremental semantics for class incremental learning
Zihuan Qiu, Linfeng Xu 0001, Qingbo Wu 0001, Fanman Meng, Hongliang Li 0001
Neurocomputing1
2023 GFR: Generic feature representations for class incremental learning
abstract
Class incremental learning (CIL) aims to continuously learn new classes while maintaining discrimination for old classes with sequentially coming data. Due to the lack of old-class samples, existing CIL methods fail to learn discriminative representations for both old and new classes simultaneously, resulting in a severe performance drop in old classes, which is the well-known catastrophic forgetting phenomenon. Different from most existing works, we facilitate CIL by learning generic feature representations that perform well in seen and unseen classes. Specifically, we prove that representations with a substantial number of significant singular values benefit CIL via better old knowledge reservation. However, the overly uniform singular value spectrum will hurt the discrimination of current tasks. Furthermore, we propose that increasing the embedding dimension can enhance the number of significant singular values and validate this assumption from two perspectives: adopting different pooling techniques and devising a wider network. Meanwhile, we also prove that satisfactory current task accuracy and old knowledge reservation can be achieved simultaneously. Finally, the simple yet effective generic feature representation regulation (GFR) is devised and incorporated into two baselines. Extensive experiments are conducted on CIFAR100, ImageNet-Subset, and ImageNet. The results show that the proposed method boosts the performance of both baselines with a large margin (2.00%-9.58% on CIFAR100, 0.68%-7.10% on ImageNet-Subset and 1.18%-5.04% on ImageNet) which outperforms existing SOTAs.
Linfeng Xu 0001, Zihuan Qiu, Qingbo Wu 0001, Fanman Meng, Hongliang Li 0001
Neurocomputing3
2022 Eldnet: Establishment and Refinement of Edge Likelihood Distributions for Camouflaged Object Detection
abstract
Camouflaged object detection (COD) focuses on detecting objects assimilated into surroundings. Distractions from the background make it extremely challenging to find the correct semantics and border details. Some of the current methods aim to explore edge details directly which has the risk of leading to boundary errors under strong background interference. While other conservative methods find rough semantic locations, then implicitly infer the edge against the background which cannot fully pay attention to the spatial details manifesting as partial semantic loss and blurring. Different from the existence, we propose a framework that generates progressively refined edge likelihood maps to guide the feature fusion of camouflaged objects. To get robust and refined edge likelihood maps, a module is first trained to fit the reasonable edge likelihood distribution based on semantic position clues and then a progressive architecture is designed to refine the possible edge area which gradually approaches the real edge. Our method can avoid catastrophic errors during exploring the fine edge, but also improves border details, solving the border blurring. Experiments on four datasets demonstrate the significant improvements and generalization ability of ours compared with state-of-the-art methods.
Chiyuan He, Linfeng Xu 0001, Zihuan Qiu
ICIP3