Xiang Song 0005

dblp:71/6574-5 · DBLP profile ↗
← Back
14ranked-venue papers
4as first author
14since 2021 · last 2026
0000-0003-1740-9104ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 4 first-author · 11 since 2021Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Multi-Task Unified Domain Incremental Learning With Domain Difference Adapters
abstract
This paper focuses on domain incremental learning (DIL) for multiple vision tasks, including object detection, instance segmentation, and image classification. DIL aims to adapt a model to new domains over time without forgetting previously acquired knowledge. Recent DIL methods append learnable prompts to input embeddings of a frozen base model to learn from new domains. However, due to prompts' limited representation ability, they struggle to adapt the feature space to new domain data distributions. To overcome this limitation, we propose a novel DIL method named Domain Difference Adapters (DD-Adapters). Through feature visualization and singular value analysis, we identify the cross-domain clustering ability of the base model and the low-rank property of domain difference. Based on these insights, our method imposes low-rank constraints on the base model to capture the principal components of domain differences, while freezing the base model to maintain its cross-domain clustering ability, thereby adapting to new domains effectively. Additionally, we introduce a prototype-guided domain selector (PDS) to dynamically select the appropriate DD-Adapters during inference, mitigating catastrophic forgetting in DIL. Extensive experimental evaluations on eight benchmark datasets demonstrate the performance superiority of the proposed method on three vision tasks, with minimal extra parameter usage.
Xiang Song 0005, Yuhang He 0001, Lin Peng 0003, Yihong Gong
IEEE Trans. Image Process.1
2025 DualCP: Rehearsal-Free Domain-Incremental Learning via Dual-Level Concept Prototype
abstract
Domain-Incremental Learning (DIL) enables vision models to adapt to changing conditions in real-world environments while maintaining the knowledge acquired from previous domains. Given privacy concerns and training time, Rehearsal-Free DIL (RFDIL) is more practical. Inspired by the incremental cognitive process of the human brain, we design Dual-level Concept Prototypes (DualCP) for each class to address the conflict between learning new knowledge and retaining old knowledge in RFDIL. To construct DualCP, we propose a Concept Prototype Generator (CPG) that generates both coarse-grained and fine-grained prototypes for each class. Additionally, we introduce a Coarse-to-Fine calibrator (C2F) to align image features with DualCP. Finally, we propose a Dual Dot-Regression (DDR) loss function to optimize our C2F module. Extensive experiments on the DomainNet, CDDB, and CORe50 datasets demonstrate the effectiveness of our method.
Yuhang He 0001, Songlin Dong, Xiang Song 0005, Jizhou Han, Haoyu Luo, Yihong Gong
AAAI4
2025 Learning Endogenous Attention for Incremental Object Detection
abstract
In this paper, we focus on a challenging Incremental Object Detection (IOD) problem. Existing IOD methods adopt an image-to-annotation alignment paradigm, which attempts to complete the absent old category annotations and learns both new and old categories concurrently in new tasks. This paradigm inherently introduces missing/redundant/inaccurate annotations of old categories, resulting in a suboptimal performance. Instead, we propose a novel annotation-to-instance alignment IOD paradigm and develop a corresponding method named Learning Endogenous Attention (LEA). Inspired by the human brain, LEA enables the model to focus on annotated task-specific objects, while ignoring irrelevant ones, thus solving the annotation incomplete problem in IOD. Concretely, our LEA consists of Endogenous Attention Modules (EAMs) and an Energy-Based Task Modulator (ETM). During training, we add the dedicated EAMs for each new task and train them to focus on the new categories. During testing, ETM predicts task IDs using energy functions, directing the model to detect task-specific objects. The detection results corresponding to all task IDs are combined as the final output, thereby alleviating the catastrophic forgetting of old knowledge. Extensive experiments on COCO 2017 and Pascal VOC 2007 demonstrate the effectiveness of our method1.
Xiang Song 0005, Yuhang He 0001, Yihong Gong
CVPR1
2025 Boosting Domain Incremental Learning: Selecting the Optimal Parameters is All You Need
abstract
Deep neural networks (DNNs) often underperform in real-world, dynamic settings where data distributions change over time. Domain Incremental Learning (DIL) offers a solution by enabling continuous model adaptation, with Parameter-Isolation DIL (PIDIL) emerging as a promising paradigm to reduce knowledge conflicts. However, existing PIDIL methods struggle with parameter selection accuracy, especially as the number of domains and corresponding classes grows. To address this, we propose SOYO, a lightweight framework that improves domain selection in PIDIL. SOYO introduces a Gaussian Mixture Compressor (GMC) and Domain Feature Resampler (DFR) to store and balance prior domain data efficiently, while a Multi-level Domain Feature Fusion Network (MDFN) enhances domain feature extraction. Our framework supports multiple Parameter-Efficient Fine-Tuning (PEFT) methods and is validated across tasks such as image classification, object detection, and speech enhancement. Experimental results on six benchmarks demonstrate SOYO’s consistent superiority over existing baselines, showcasing its robustness and adaptability in complex, evolving environments.
Xiang Song 0005, Yuhang He 0001, Jizhou Han, Chenhao Ding, Xinyuan Gao, Yihong Gong
CVPR2
2025 CIA: Class- and Instance-aware Adaptation for Vision-Language Models
abstract
Few-shot parameter-efficient tuning methods demonstrate promising potential for Vision-Language (V-L) models in downstream tasks. However, existing approaches primarily focus on class-level alignment between image and text features, overlooking crucial instance-specific semantic information. This limitation leads to suboptimal performance on challenging tasks and restricted generalization capability to unseen data. To address these issues, we propose Class- and Instance-aware Adaptation (CIA), a novel framework that simultaneously optimizes both class-level and instance-level alignments. Specifically, CIA introduces a novel instance encoder that leverages cross-modal self-attention to generate instance-specific text features, accompanied by a carefully designed regularization mechanism to maintain consistency between class-level and instance-level representations. Extensive experiments across 15 benchmark datasets demonstrate that CIA significantly improves the downstream adaptation of V-L models.
Lin Peng 0003, Cong Wan, Shaokun Wang, Xiang Song 0005, Yuhang He 0001, Yihong Gong
ACM Multimedia4
2025 Cross-Domain Invariant Feature Absorption and Domain-Specific Feature Retention for Domain Incremental Chest X-Ray Classification
abstract
Chest X-ray (CXR) images have been widely adopted in clinical care and pathological diagnosis in recent years. Some advanced methods on CXR classification task achieve impressive performance by training the model statically. However, in the real clinical environment, the model needs to learn continually and this can be viewed as a domain incremental learning (DIL) problem. Due to large domain gaps, DIL is faced with catastrophic forgetting. Therefore, in this paper, we propose a Cross-domain invariant feature absorption and Domain-specific feature retention (CaD) framework. To be specific, we adopt a Cross-domain Invariant Feature Absorption (CIFA) module to learn the domain invariant knowledge and a Domain-Specific Feature Retention (DSFR) module to learn the domain-specific knowledge. The CIFA module contains the C(lass)-adapter and an absorbing strategy is used to fuse the common features among different domains. The DSFR module contains the D(omain)-adapter for each domain and it connects to the network in parallel independently to prevent forgetting. A multi-label contrastive loss (MLCL) is used in the training process and improves the class distinctiveness within each domain. We leverage publicly available large-scale datasets to simulate domain incremental learning scenarios, extensive experimental results substantiate the effectiveness of our proposed methods and it has reached state-of-the-art performance.
Mengchu Wang, Yuhang He 0001, Lin Peng 0003, Xiang Song 0005, Songlin Dong, Yihong Gong
IEEE Trans. Medical Imaging4
2024 Non-exemplar Domain Incremental Object Detection via Learning Domain Bias
abstract
Domain incremental object detection (DIOD) aims to gradually learn a unified object detection model from a dataset stream composed of different domains, achieving good performance in all encountered domains. The most critical obstacle to this goal is the catastrophic forgetting problem, where the performance of the model improves rapidly in new domains but deteriorates sharply in old ones after a few sessions. To address this problem, we propose a non-exemplar DIOD method named learning domain bias (LDB), which learns domain bias independently at each new session, avoiding saving examples from old domains. Concretely, a base model is first obtained through training during session 1. Then, LDB freezes the weights of the base model and trains individual domain bias for each new incoming domain, adapting the base model to the distribution of new domains. At test time, since the domain ID is unknown, we propose a domain selector based on nearest mean classifier (NMC), which selects the most appropriate domain bias for a test image. Extensive experimental evaluations on two series of datasets demonstrate the effectiveness of the proposed LDB method in achieving high accuracy on new and old domain datasets. The code is available at https://github.com/SONGX1997/LDB.
Xiang Song 0005, Yuhang He 0001, Songlin Dong, Yihong Gong
AAAI1
2024 Evolving Parameterized Prompt Memory for Continual Learning
abstract
Recent studies have demonstrated the potency of leveraging prompts in Transformers for continual learning (CL). Nevertheless, employing a discrete key-prompt bottleneck can lead to selection mismatches and inappropriate prompt associations during testing. Furthermore, this approach hinders adaptive prompting due to the lack of shareability among nearly identical instances at more granular level. To address these challenges, we introduce the Evolving Parameterized Prompt Memory (EvoPrompt), a novel method involving adaptive and continuous prompting attached to pre-trained Vision Transformer (ViT), conditioned on specific instance. We formulate a continuous prompt function as a neural bottleneck and encode the collection of prompts on network weights. We establish a paired prompt memory system consisting of a stable reference and a flexible working prompt memory. Inspired by linear mode connectivity, we progressively fuse the working prompt memory and reference prompt memory during inter-task periods, resulting in continually evolved prompt memory. This fusion involves aligning functionally equivalent prompts using optimal transport and aggregating them in parameter space with an adjustable bias based on prompt node attribution. Additionally, to enhance backward compatibility, we propose compositional classifier initialization, which leverages prior prototypes from pre-trained models to guide the initialization of new classifiers in a subspace-aware manner. Comprehensive experiments validate that our approach achieves state-of-the-art performance in both class and domain incremental learning scenarios.
Muhammad Rifki Kurniawan, Xiang Song 0005, Zhiheng Ma, Yuhang He 0001, Yihong Gong, Xing Wei 0001
AAAI2
2024 Prompt-Agnostic Adversarial Perturbation for Customized Diffusion Models
abstract
Diffusion models have revolutionized customized text-to-image generation, allowing for efficient synthesis of photos from personal data with textual descriptions. However, these advancements bring forth risks including privacy breaches and unauthorized replication of artworks. Previous researches primarily center around using “prompt-specific methods” to generate adversarial examples to protect personal images, yet the effectiveness of existing methods is hindered by constrained adaptability to different prompts. In this paper, we introduce a Prompt-Agnostic Adversarial Perturbation (PAP) method for customized diffusion models. PAP first models the prompt distribution using a Laplace Approximation, and then produces prompt-agnostic perturbations by maximizing a disturbance expectation based on the modeled distribution. This approach effectively tackles the prompt-agnostic attacks, leading to improved defense stability. Extensive experiments in face privacy and artistic style protection, demonstrate the superior generalization of our method in comparison to existing techniques.
Cong Wan, Yuhang He 0001, Xiang Song 0005, Yihong Gong
NeurIPS3
2024 Overcoming Catastrophic Forgetting for Multi-Label Class-Incremental Learning
abstract
Despite the recent progress of class-incremental learning (CIL) methods, their capabilities in real-world scenarios such as multi-label settings remain unexplored. This paper focuses on a more practical CIL problem named multi-label class-incremental learning (MLCIL). MLCIL requires the vision models to overcome catastrophic forgetting of old knowledge while learning new classes from multi-label samples. Direct application of existing CIL methods to MLCIL leads to label absence, representative sample selection, and feature dilution problems. To address these problems, we present a novel AdaPtive Pseudo-Label-drivEn (APPLE) framework consisting of three components. First, the adaptive pseudo-label strategy is proposed to solve the label absence problem, which leverages the old model to annotate old classes for new samples. Second, a cluster sampling strategy is proposed to obtain more diverse samples to alleviate catastrophic forgetting under the MLCIL setting better. Finally, a class attention decoder is designed to mitigate the object feature dilution problem in multi-label samples. The extensive experiments on PASCAL VOC 2007 and MS-COCO demonstrate that our proposed method significantly outperforms other representative state-of-the-art CIL methods.
Xiang Song 0005, Kuang Shu, Songlin Dong, Xing Wei 0001, Yihong Gong
WACV1
2024 Global self-sustaining and local inheritance for source-free unsupervised domain adaptation
Lin Peng 0003, Yuhang He 0001, Shaokun Wang, Xiang Song 0005, Songlin Dong, Xing Wei 0001, Yihong Gong
Pattern Recognit.4
2024 Learning Noise Adapters for Incremental Speech Enhancement
abstract
Incremental speech enhancement (ISE), with the ability to incrementally adapt to new noise domains, represents a critical yet comparatively under-investigated topic. While the regularization-based method has been proposed to solve the ISE task, it usually suffers from the dilemma wherein the gain of one domain directly entails the loss of another. To solve this issue, we propose an effective paradigm, termed Learning Noise Adapters (LNA), which significantly mitigates the catastrophic domain forgetting phenomenon in the ISE task. In our methodology, we employ a frozen pre-trained model to train and retain a domain-specific adapter for each newly encountered domain, enabling the capture of variations in feature distributions within these domains. Subsequently, our approach involves the development of an unsupervised, training-free noise selector for the inference stage, which is responsible for identifying the domains of test speech samples. A comprehensive experimental validation has substantiated the effectiveness of our approach.
Ziye Yang, Xiang Song 0005, Jie Chen 0022, Cédric Richard, Israel Cohen
IEEE Signal Process. Lett.2
2024 Domain Incremental Object Detection Based on Feature Space Topology Preserving Strategy
abstract
Object detection with the capacity to incrementally adapt to new domains is a crucial yet relatively under-explored research topic. The catastrophic forgetting problem presents a significant challenge to achieve this goal, where the model’s performance improves quickly in new conditions but deteriorates sharply in old ones after several incremental learning sessions. Drawing on recent discoveries in visual memories of the human brain, we introduce the Topology-Preserving Domain Incremental Object Detection (TP-DIOD) approach, which aims to address the catastrophic forgetting problem by extracting the topological structure of the feature space learned by the Convolutional Neural Network (CNN) model and preserving this topology during the subsequent incremental learning sessions. Specifically, we model the feature space topology using the self-organizing map (SOM) and construct an anchor image set based on the centroid vectors of the SOM nodes to memorize the feature space topology. We then develop the anchor loss function to penalize the topological changes of the feature space during the subsequent incremental learning sessions. Experimental evaluations on two sets of datasets demonstrate the effectiveness of the proposed TP-DIOD method in mitigating the catastrophic forgetting problem and achieving high accuracy on both old and new domain datasets.
Xiang Song 0005, Yuhang He 0001, Changxin Wang, Songlin Dong, Xing Wei 0001, Yihong Gong
IEEE Trans. Circuits Syst. Video Technol.2
2023 Semantic Knowledge Guided Class-Incremental Learning
abstract
Driven by practical needs, research on Class-Incremental Learning (CIL) has received more and more attentions in recent years. A technical challenge to be conquered by CIL methods is the catastrophic forgetting problem, where the model’s performance improves rapidly on new classes while deteriorates drastically on old ones. The main causes behind catastrophic forgetting include network drifts, inter-class confusions, etc. In this paper, we propose a novel CIL method that solves the catastrophic forgetting problem from two aspects. First, to solve the inter-class confusion problem, we propose a novel Semantic knOwledge gUided ciL framework (SOUL) that consists of a CNN feature extractor and a Bi-GCN (Graph Convolutional Network) classifier. In each CIL session, we use the semantic knowledge extracted from the class labels to build two inter-class relation graphs among all the encountered old and new classes. Using these two relation graphs, we develop a Bi-GCN classifier to fuse two kinds of semantic relations in a balanced way, and then to transfer the inter-class relations from semantic modality to image classification weights. The entire SOUL framework is trained end-to-end by the standard BP algorithm, which optimizes the Bi-GCN classifier and the CNN feature extractor jointly. Second, to prevent the network drift, we develop the local topology preserving strategy that divides the global topological structure of the learned feature space into a set of local topological relations, and maintains these local relations at CIL session. Experimental evaluations demonstrate the state-of-the-art performance accuracies on benchmark image classification datasets.
Shaokun Wang, Weiwei Shi 0003, Songlin Dong, Xinyuan Gao, Xiang Song 0005, Yihong Gong
IEEE Trans. Circuits Syst. Video Technol.5