Zijian Gao

dblp:250/3021 · DBLP profile ↗
← Back
31ranked-venue papers
9as first author
30since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 16 · 4 first-author · 16 since 2021Artificial intelligence and machine learning · 11 · 6 first-author · 11 since 2021Systems, architecture and hardware · 3 · 3 since 2021Security and privacy · 2 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Perturbing to Preserve: Defending Fragile Knowledge in Online Continual Learning
abstract
Online continual learning requires models to learn from non‑stationary data streams while retaining prior knowledge. We identify an overlooked phenomenon—knowledge fragility—where correctly learned instances are rapidly forgotten after minor parameter updates. Our analysis attributes this fragility to a temporal–spatial dual mechanism: temporal instability, high-frequency parameter oscillations cause forgetting to outpace adaptation; and spatial vulnerability, fragile instances lie in sharp, high‑curvature regions of the loss landscape that are extremely sensitive to optimization noise. These insights motivate PDFK (Perturbing to Defend Fragile Knowledge), a unified framework that defends fragile knowledge along both dimensions. Temporally, we apply exponential moving averaging to smooth parameter evolution and stabilize long‑term memory. Spatially, we inject minimal structured perturbations with a consistency constraint to flatten sharp regions and enhance robustness. PDFK requires no task‑boundary annotations. Extensive experiments demonstrate that PDFK substantially improves knowledge retention and outperforms strong baselines under diverse and challenging continual learning settings.
Dulan Zhou, Zijian Gao, Kele Xu
AAAI2
2026 Detect and Act: Automated Dynamic Optimizer through Meta-Black-Box Optimization
abstract
Dynamic Optimization Problems (DOPs) are challenging to address due to their complex nature, i.e., dynamic environment variation. Evolutionary Computation methods are generally advantaged in solving DOPs since they resemble dynamic biological evolution. However, existing evolutionary dynamic optimization methods rely heavily on human-crafted adaptive strategy to detect environment variation in DOPs, and then adapt the searching strategy accordingly. These hand-crafted strategies may perform ineffectively at out-of-box scenarios. In this paper, we propose a reinforcement learning-assisted approach to enable automated variation detection and self-adaption in evolutionary algorithms. This is achieved by borrowing the bi-level learning-to-optimize idea from recent Meta-Black-Box Optimization works. We use a deep Q-network as optimization dynamics detector and searching strategy adapter: It is fed as input with current-step optimization state and then dictates desired control parameters to underlying evolutionary algorithms for next-step optimization. The learning objective is to maximize the expected performance gain across a problem distribution. Once trained, our approach could generalize toward unseen DOPs with automated environment variation detection and self-adaption. To facilitate comprehensive validation, we further construct an easy-to-difficult DOPs testbed with diverse synthetic instances. Extensive benchmark results demonstrate flexible searching behavior and superior performance of our approach in solving DOPs, compared to state-of-the-art baselines. Our code is publicly available here: https://github.com/MetaEvo/Meta-DO
Zijian Gao, Zeyuan Ma, Yuanting Zhong, Yue-Jiao Gong, Hongshu Guo
GECCO1
2026 Multimodal Deepfake Detection with Quantum State Inspired Analytic Incremental Adaptability Learning
abstract
Multimodal deepfake technologies have emerged rapidly in recent years, with wide application prospects in various fields. The conventional single-training paradigm with inherent limited generalization illustrates inadequate for addressing the continuous evolution of multimodal deepfakes. However, fine-tuning a model with new deepfake data faces past forgery patterns loss and the significant domain shift in diverse novel multimodal deepfake technologies. To address these issues, we propose a novel Quantum State Analytic Incremental Adaptability Learning method (Qsaint) for multimodal deepfake detection. To stabilize prior deepfake memory, Qsaint recursively learns detection-label mapping relations for the new deepfakes artifact with a closed-form solution, preserving the distribution memory from the historical deepfake domains without accessing previous videos. During incremental learning stages, we propose a deepfake quantum state adaptability module inspired by quantum information science. It adapts to the new forgery states and aligns them with the historical deepfake knowledge through cooling and evolution operations, eliminating deepfake domain shift issues. Comprehensive experiments demonstrate that Qsaint significantly mitigates the memory interference of historical deepfakes, effectively balancing the adaptability for new forgery tasks with the memorization of known deepfake patterns.
Jianbin Ye, Bo Liu 0014, Huaping Hu, Zijian Gao, Shaojing Fu, Kele Xu, Huaimin Wang 0001
ICMR5
2026 Rethinking Obscured Sub-Optimality in Analytic Learning for Exemplar-Free Class-Incremental Learning
abstract
Exemplar-free Class-Incremental Learning (EFCIL) poses a significant challenge in mitigating catastrophic forgetting, due to the absence of exemplars. Recently, analytic learning-based methods propose a recursive alignment procedure to execute EFCIL in a phase-invariant manner and show state-of-the-art performance. However, they heavily rely on a frozen feature extractor trained with the initial dataset to avoid the misalignment between feature and label spaces, ignoring the importance of acquiring generalizable features across incremental tasks for performance improvement. To tackle this, we rethink the obscured sub-optimality of analytic learning-based methods, particularly through empirical reevaluation, and then introduce the Multi-head analytic learning (Muheal) approach. Muheal forms the multi-head model with a delicate feature extractor, thereby introducing a feature optimization procedure and a forgetting compensation module to balance the learning and forgetting. Specifically, within the feature optimization procedure, the feature extractor seeks to learn more generalizable features in a self-supervised manner using the fully-connected classification head. An analytic learning-based classification head follows to align the feature-label space. Additionally, we employ the compensation module to generate and align pseudo-features with a replicated analytic head, thus preventing overfitting and testing. Comprehensive experiments on several benchmark datasets have demonstrated that Muheal significantly outperforms existing state-of-the-art EFCIL methods and is comparable, if not superior, to methods that use replay techniques.
Zijian Gao, Kele Xu, Xingxing Zhang 0001, Huiping Zhuang, Tianjiao Wan, Bo Ding 0001, Xinjun Mao, Huaimin Wang 0001
IEEE Trans. Circuits Syst. Video Technol.1
2026 Dynamic Confidence Variance for Generalized Coreset in Active Learning
abstract
Active Learning (AL) aims to reduce data annotation costs by selecting the most informative samples from an unlabeled data pool. Traditional AL methods often rely on a single snapshot to identify uncertain or representative samples, often overlooking the poor generalization of a single model. Recent AL studies have attempted to address this issue by tracking a broader range of training dynamics for data selection, typically using averaging or accumulating manner. However, both our theoretical and experimental analyses reveal that these methods obscure the variability inherent in the training process, potentially prioritizing hard-to-learn samples that result in poor generalization. In this paper, we propose a novel AL method termed as Dynamic Confidence Variance (DCoV), that seamlessly integrates variability with the training dynamic to effectively identify a well-generalized Coreset. DCoV leverages the variance of the model’s prediction confidence throughout the training process for active sampling and model training. Our theoretical analysis demonstrates that DCoV provides a lower bound on the population risk of the model learned from selected labeled subset, spanning the entire training process. Extensive experiments demonstrate that our approach significantly outperforms existing state-of-the-art AL methods on various balanced and imbalanced benchmark datasets across various modalities.
Tianjiao Wan, Zijian Gao, Xudong Gong, Bo Ding 0001, Huaimin Wang 0001, Kele Xu
IEEE Trans. Circuits Syst. Video Technol.2
2026 Adaptive Affinity Memorization With Layer Mutation for Multimodal Deepfake Continual Detection
Jianbin Ye, Bo Liu 0014, Zijian Gao, Wuyang Chen 0002, Tao Li 0008, Huaimin Wang 0001, Kele Xu
IEEE Trans. Inf. Forensics Secur.4
2026 Few-Shot Pulmonary Vessel Segmentation Based on Tubular-Aware Prompt-Tuning
abstract
Segmentation of the pulmonary vessel from computed tomography (CT) images plays a crucial role in the diagnosis and treatment of various lung diseases. Although deep learning-based approaches have shown remarkable progress in recent years, their performance is often hindered by the lack of high-quality annotated datasets, in which the complex anatomy and morphology of pulmonary vessels make manual annotation challenging, time-consuming, and prone to errors. To address this, we propose PV25, the first dataset that features finely paired annotations of both pulmonary vessels and airways. Moreover, we propose TPNet, a novel tubular-aware prompt-tuning framework for pulmonary vessel segmentation under few-shot training with limited annotations. Specifically, based on an advanced and frozen segmentation backbone, TPNet proposes tunable encoding and decoding networks that learn tubular structures as transfer learning priors, bridging the gap between the source and target pulmonary vessel domains. Specifically, TPNet is built in an encoder-decoder manner, including the fixed segmentation backbone, tunable encoding and decoding networks. In encoding stage, the Morphology-Driven Region Growing (MDRG) module is developed to leverage the tubular connectivity of vessels to guide the network in capturing fine-grained features of pulmonary vessels. In decoding stage, the Cross-Correlation Guidance (CCG) module is introduced to integrate multi-scale correlations between airway and vessel structures in a coarse-to-fine manner. Extensive experiments conducted on multiple datasets demonstrate that TPNet achieves state-of-the-art performance in pulmonary vessel segmentation under limited training data. Besides, TPNet shows strong performance in related tasks such as airway segmentation and artery-vein classification, highlighting its robustness and versatility.
Zijian Gao, Lai Jiang 0004, Sukun Tian, Yuchun Sun, Mai Xu, Liyuan Tao
IEEE Trans. Medical Imaging1
2025 Maintaining Fairness in Logit-based Knowledge Distillation for Class-Incremental Learning
abstract
Logit-based knowledge distillation (KD) is commonly used to mitigate catastrophic forgetting in class-incremental learning (CIL) caused by data distribution shifts. However, the strict match of logit values between student and teacher models conflicts with the cross-entropy (CE) loss objective of learning new classes, leading to significant recency bias (i.e. unfairness). To address this issue, we rethink the overlooked limitations of KD-based methods through empirical analysis. Inspired by our findings, we introduce a plug-and-play pre-process method that normalizes the logits of both the student and teacher across all classes, rather than just the old classes, before distillation. This approach allows the student to focus on both old and new classes, capturing intrinsic inter-class relations from the teacher. By doing so, our method avoids the inherent conflict between KD and CE, maintaining fairness between old and new classes. Additionally, recognizing that overconfident teacher predictions can hinder the transfer of inter-class relations (i.e., dark knowledge), we extend our method to capture intra-class relations among different instances, ensuring fairness within old classes. Our method integrates seamlessly with existing logit-based KD approaches, consistently enhancing their performance across multiple CIL benchmarks without incurring additional training costs.
Zijian Gao, Shanhao Han, Xingxing Zhang 0001, Kele Xu, Dulan Zhou, Xinjun Mao, Yong Dou, Huaimin Wang 0001
AAAI1
2025 Knowledge Memorization and Rumination for Pre-trained Model-based Class-Incremental Learning
abstract
Class-Incremental Learning (CIL) enables models to continuously learn new classes while mitigating catastrophic forgetting. Recently, Pre-Trained Models (PTMs) have greatly enhanced CIL performance, even when fine-tuning is limited to the first task. This advantage is particularly beneficial for CIL methods that freeze the feature extractor after first-task fine-tuning, such as analytic learning-based approaches using a least squares solution-based classification head to acquire knowledge recursively. In this work, we revisit the analytical learning approach combined with PTMs and identify its limitations in adapting to new classes, leading to sub-optimal performance. To address this, we propose the Momentum-based Analytical Learning (MoAL) approach. MoAL achieves robust knowledge memorization via an analytical classification head and improves adaptivity to new classes through momentum-based adapter weight interpolation, leading to forgetting outdated knowledge. Importantly, we introduce a knowledge rumination mechanism that leverages refined adaptivity, allowing the model to revisit and reinforce old knowledge, thereby improving performance on old classes. MoAL facilitates the acquisition of new knowledge and consolidates old knowledge, achieving a win-win outcome between plasticity and stability. Extensive experiments on various incremental settings show MoAL’s state-of-the-art performance1.
Zijian Gao, Wangwang Jia, Xingxing Zhang 0001, Dulan Zhou, Kele Xu, Yong Dou, Xinjun Mao, Huaimin Wang 0001
CVPR1
2025 Complementary Learning System Theory-based Active Learning for Audio Classification
abstract
Deep learning has significantly advanced the audio classification, achieving remarkable results. However, these successes often rely on extensive manual annotation of audio, a labor-intensive and costly process. Active Learning (AL) presents a promising solution by minimizing the required amount of annotation through the iterative selection of the most informative audio samples. Current AL methods for audio classification typically depend solely on the latest model checkpoint, overlooking the dynamics of the entire training process. The Complementary Learning Systems (CLS) theory posits that the interplay between short-term and long-term memory systems can effectively measure sample uncertainty, offering a means to capture training dynamics. In this work, we introduce a novel AL framework for audio classification, termed CLS-AL, which addresses the limitations of existing methods by simultaneously maintaining both short-term and long-term memory models. This dual-memory approach allows for a more comprehensive consideration of training dynamics. The divergence in predictions between these memory models provides a new metric for evaluating the uncertainty of unlabeled samples, enhancing the effectiveness of the AL sample selection process. We demonstrate the effectiveness and generalizability of CLS-AL through extensive experiments on a diverse set of audio datasets, showing that CLS-AL obviously outperforms existing state-of-the-art methods.
Hui Geng, Zijian Gao, Tianjiao Wan, Kele Xu
ICASSP2
2025 Text-guided Multimodal Fusion for the Multimodal Emotion and Intent Joint Understanding
abstract
Emotion and Intent Joint Understanding in Multi-modal Conversation is a challenging task in the field of affective computing, aiming to decode the semantic information manifested in the multimodal conversational while simultaneously inferring the emotions and intents of the utterance. To address this challenge, we propose the Text-guided Multimodal Emotion-Intent Joint Recognition method. By leveraging the text modality to guide the fusion process, it effectively reduces the noise introduced by other modalities. To strengthen the text modality’s guiding role, we use large language models (LLMs) for multi-turn targeted data augmentation and oversampling strategies to address data imbalance. Our approach achieved first place in Track 1 (English) of the ICASSP 2025 MEIJU Challenge, demonstrating its effectiveness in practical applications.
Yu Zhang 0133, Bin Chen 0006, Hongfei Ye, Zijian Gao, Tianjiao Wan, Long Lan, Kele Xu
ICASSP4
2025 SSAST-Adapter: A Parameter-efficient Incremental Learning Algorithm for Underwater Acoustic Target Recognition
abstract
Underwater acoustic target recognition involves identifying and classifying targets in underwater environments using acoustic signals. In recent years, deep learning has made significant progress in this field. However, the models require the entire dataset to be available upfront, and classification categories must be predefined. In practical scenarios, new objects or species may appear in underwater environments over time, making it impractical to retrain a model from scratch each time new data is introduced. At the same time, training on new data inevitably leads to catastrophic forgetting of past data. To address these challenges, we propose a large-scale pre-training strategy combined with adapters to enable incremental learning without the need for complete retraining. To minimize the number of parameters during model fine-tuning, we employ a adapter structure at various network layers, reducing the number of trainable parameters to less than 2%. Experimental results demonstrate that our proposed method effectively reduces the need for full retraining by allowing the model to update in a resource-efficient manner using only new data, and the model has achieved high recognition accuracy in underwater acoustic target recognition tasks.
Qisheng Xu, Boqing Zhu, Zijian Gao, Lingbin Zeng, Kele Xu
ICASSP4
2025 Self-supervised Bidirectional Synchronization Estimation for Multimodal Deepfake Detection with Short-term Dependency
abstract
Deepfake technology induces substantial societal challenges, establishing deepfake detection as an important area of research. However, existing research mainly relies on target deepfake datasets, which limits its generalizability across out-of-distribution tasks to some extent. Also, it often emphasizes visual modalities while neglecting the complementary information of the auditory data. Their autoregressive-based strategies also introduce long-term information interference, further constraining the detection performance. Consequently, the potential to exploit complementary relations between visual and auditory modalities and to leverage strongly correlated short-range information remains underexplored for the detection task. To address these challenges, this paper introduces Self-BiSterm, a novel self-supervised learning framework for deepfake detection. First, we propose a bidirectional synchronization distribution modeling mechanism, which calculates inconsistent distributions for video-to-audio and audio-to-video scenarios. This mechanism effectively measures audio-visual inconsistencies, improving the model's generalization performance in practical applications. Second, to mitigate the issue of long-term information distortion, we develop a short-term temporal dependency module to estimate the adjacent local receptive fields. This module facilitates the estimation of subsequent distributions by capturing short-term temporal dependencies with high precision. The effectiveness of the proposed Self-BiSterm framework is validated on various benchmarks, demonstrating superior performance compared to existing methods.
Jianbin Ye, Bo Liu 0014, Zijian Gao, Kele Xu, Xiaodong Wang 0002
ICMR4
2025 Analytic Synaptic Dynamic Scaling Balancer for Multimodal Deepfake Continual Detection
abstract
Multimodal deepfakes pose growing security threats across diverse domains, driven by rapid advancements in generative models. This demands effective Multimodal Deepfake Continual Detection (MDCD) methods capable of adapting to evolving and heterogeneous deepfake techniques. However, MDCD remains underexplored, facing two major challenges: (1) modality-specific feature disparities limit the effectiveness of simple feature fusion, exacerbating the forgetting of previous forgery-relevant knowledge; and (2) newly introduced deepfake videos initially exhibit limited scale that gradually expand, causing class imbalance dominated by forged samples, undermines authentic content understanding in comming tasks. To address these issues, we propose the Analytic Synaptic Dynamic Scaling Balancer (ADanser) that adapts to modality-specific biases and class imbalance while employing a closed-form update to preserve prior multimodal deepfake knowledge in an evolving data stream. Inspired by synaptic scaling in neuroscience, ADanser introduces a modality synaptic scaling mechanism that applies modality-aware attention to extract discriminative and complementary forgery patterns, improving cross-modal knowledge retention. Additionally, a class-wise contribution balancer dynamically reweights learning signals to reduce class bias and enhance authentic video representation. Extensive experiments on benchmark multimodal deepfake datasets demonstrate that ADanser significantly outperforms state-of-the-art continual learning methods, effectively coordinating adaptation and retention in imbalanced, cross-modal scenarios.
Jianbin Ye, Bo Liu 0014, Zijian Gao, Kele Xu, Xiaodong Wang 0002
ACM Multimedia4
2025 A missing multimodal imputation diffusion model for 2D X-ray and 3D CT in COVID-19 diagnosis
Zijian Gao, Yiqing Shen 0003
Expert Syst. Appl.1
2025 Hardware Architecture Design for Iterative Reconstruction Algorithms Toward Palm-Size Photoacoustic Tomography
abstract
Photoacoustic (PA) imaging technology integrates the deep penetration depth of ultrasound imaging with the high resolution of optical imaging, demonstrating significant potential in biomedical applications. Many preclinical studies and clinical applications urgently require a portable, high-quality, low-cost, and fast imaging system. Thus, translating advanced image reconstruction algorithms into hardware implementations is highly desired. However, existing iterative PA image reconstructions, although exhibit higher accuracy than the delay-and-sum algorithm, suffer from high computational cost. In this paper, we introduce a novel model-based hardware architecture for palm-size PA tomography (palm-PAT), aiming at enhancing both the speed and performance of image reconstruction at a much lower system cost. To achieve this, we propose an innovative data reuse method that significantly reduces the consumption of hardware storage resources, achieving a reduction of approximately 75% in hardware storage requirements. We conducted experiments utilizing the FPGA implementation of the algorithm, using phantom, human finger data in vivo and ex vivo breast tumor data to verify the feasibility of the proposed method. The results demonstrate that our proposed architecture can substantially reduce system cost while maintaining high imaging performance. The novel hardware architecture design of the model-based algorithm achieves a speedup of up to approximately 270 times compared to the CPU, while the corresponding energy efficiency ratio is improved by more than 2700 times.
Yuwei Zheng, Zijian Gao, Yuting Shen, Ruixi Sun, Daohuai Jiang, Xiran Cai, Feng Gao 0021, Yuan Gao 0002, Fei Gao 0010
IEEE Trans. Circuits Syst. I Regul. Pap.2
2025 Dual Temporal Masked Modeling for KPI Anomaly Detection via Similarity Aggregation
abstract
With the expanding scale of current industries, monitoring systems centered around Key Performance Indicators (KPIs) play an increasingly crucial role. KPI anomaly detection can monitor the potential risks according to KPI data and has garnered widespread attention due to its rapid responsiveness and adaptability to dynamic changes. Considering the absence of labels and the high cost of manual annotation of KPI data, the self-supervised approaches are proposed. Among them, mask modeling methods draw great attention and can learn the intrinsic distribution of data without relying on prior assumptions. However, conventional mask modeling often overlooks the examination of relationships between unsynchronized variables, treating them with equal importance, and inducing inaccurate detection results. To address this, this paper proposes a Dual Masked modeling Approach combined with Similarity Aggregation, named DMASA. Starting from a self-supervised approach based on mask modeling, DMASA incorporates spectral residual techniques to explore inter-variable dependencies and aggregates information from similar data to eliminate interference from irrelevant variables in anomaly detection. Extensive experiments on eight datasets and state-of-the-art results demonstrate the effectiveness of our approach. Our code is available athttps://github.com/colaudiolab/GT-DMASA.
Zijian Gao, Kele Xu, Xu Wang 0064, Peichang Shi, Bo Ding 0001
IEEE Trans. Netw. Serv. Manag.2
2024 Optimistic Model Rollouts for Pessimistic Offline Policy Optimization
abstract
Model-based offline reinforcement learning (RL) has made remarkable progress, offering a promising avenue for improving generalization with synthetic model rollouts. Existing works primarily focus on incorporating pessimism for policy optimization, usually via constructing a Pessimistic Markov Decision Process (P-MDP). However, the P-MDP discourages the policies from learning in out-of-distribution (OOD) regions beyond the support of offline datasets, which can under-utilize the generalization ability of dynamics models. In contrast, we propose constructing an Optimistic MDP (O-MDP). We initially observed the potential benefits of optimism brought by encouraging more OOD rollouts. Motivated by this observation, we present ORPO, a simple yet effective model-based offline RL framework. ORPO generates Optimistic model Rollouts for Pessimistic offline policy Optimization. Specifically, we train an optimistic rollout policy in the O-MDP to sample more OOD model rollouts. Then we relabel the sampled state-action pairs with penalized rewards, and optimize the output policy in the P-MDP. Theoretically, we demonstrate that the performance of policies trained with ORPO can be lower-bounded in linear MDPs. Experimental results show that our framework significantly outperforms P-MDP baselines by a margin of 30%, achieving state-of-the-art performance on the widely-used benchmark. Moreover, ORPO exhibits notable advantages in problems that require generalization.
Yuanzhao Zhai, Yiying Li, Zijian Gao, Xudong Gong, Kele Xu, Bo Ding 0001, Huaimin Wang 0001
AAAI3
2024 Temporal Inconsistency-Based Active Learning
abstract
Deep supervised learning has demonstrated strong capabilities; however, such progress relies on massive and expensive data annotation. Active Learning (AL) has been introduced to selectively annotate samples, thus reducing the human labeling effort. Previous AL research has focused on employing recently trained models to design sampling strategies, based on uncertainty or representativeness. Drawing inspiration from the issue of model forgetting, we propose a novel AL framework called Temporal Inconsistency-Based Active Learning (TIR-AL). In this framework, multiple snapshots of the models across consecutive cycles are jointly utilized to select samples with higher temporal inconsistency, by computing the proposed self-weighted nuclear norm metric. Furthermore, we introduce a consistency regularization term to mitigate the issue of forgetting. Together, these components make full use of the potential of data and facilitate effective interaction within the AL loop. To demonstrate the efficacy of TIR-AL, we conducted a set of experiments illustrating how our approach outperforms state-of-the-art methods without incurring any additional training costs.
Tianjiao Wan, Yutao Dou, Kele Xu, Zijian Gao, Bo Ding 0001, Huaimin Wang 0001
ICASSP4
2024 Tracing Training Progress: Dynamic Influence Based Selection for Active Learning
abstract
Active learning (AL) aims to select highly informative data points from an unlabeled dataset for annotation, mitigating the need for extensive human labeling effort. However, classical AL methods heavily rely on human expertise to design the sampling strategy, inducing limited scalability and generalizability. Many efforts have sought to address this limitation by directly connecting sample selection with model performance improvement, typically through influence function. Nevertheless, these approaches often ignore the dynamic nature of model behavior during training optimization, despite empirical evidence highlights the importance of dynamic influence to track the sample contribution. This oversight can lead to suboptimal selection, hindering the generalizability of model. In this study, we explore the dynamic influence based data selection strategy by tracing the impact of unlabeled instances on model performance throughout the training process. Our theoretical analyses suggest that selecting samples with higher projected gradients along the accumulated optimization direction at each checkpoint leads to improved performance. Furthermore, to capture a wider range of training dynamics without incurring excessive computational or memory costs, we introduce an additional dynamic loss term designed to encapsulate more generalized training progress information. These insights are integrated into a universal and task-agnostic AL framework termed Dynamic Influence Scoring for Active Learning (DISAL). Comprehensive experiments across various tasks have demonstrated that DISAL significantly surpasses existing state-of-the-art AL methods, demonstrating its ability to facilitate more efficient and effective learning in different domains.
Tianjiao Wan, Kele Xu, Long Lan, Zijian Gao, Bo Ding 0001, Huaimin Wang 0001
ACM Multimedia4
2024 Stabilizing Zero-Shot Prediction: A Novel Antidote to Forgetting in Continual Vision-Language Tasks
abstract
Continual learning (CL) empowers pre-trained vision-language (VL) models to efficiently adapt to a sequence of downstream tasks. However, these models often encounter challenges in retaining previously acquired skills due to parameter shifts and limited access to historical data. In response, recent efforts focus on devising specific frameworks and various replay strategies, striving for a typical learning-forgetting trade-off. Surprisingly, both our empirical research and theoretical analysis demonstrate that the stability of the model in consecutive zero-shot predictions serves as a reliable indicator of its anti-forgetting capabilities for previously learned tasks. Motivated by these insights, we develop a novel replay-free CL method named ZAF (Zero-shot Antidote to Forgetting), which preserves acquired knowledge through a zero-shot stability regularization applied to wild data in a plug-and-play manner. To enhance efficiency in adapting to new tasks and seamlessly access historical models, we introduce a parameter-efficient EMA-LoRA neural architecture based on the Exponential Moving Average (EMA). ZAF utilizes new data for low-rank adaptation (LoRA), complemented by a zero-shot antidote on wild data, effectively decoupling learning from forgetting. Our extensive experiments demonstrate ZAF's superior performance and robustness in pre-trained models across various continual VL concept learning tasks, achieving leads of up to 3.70\%, 4.82\%, and 4.38\%, along with at least a 10x acceleration in training speed on three benchmarks, respectively. Additionally, our zero-shot antidote significantly reduces forgetting in existing models by at least 6.37\%. Our code is available at https://github.com/Zi-Jian-Gao/Stabilizing-Zero-Shot-Prediction-ZAF.
Zijian Gao, Xingxing Zhang 0001, Kele Xu, Xinjun Mao, Huaimin Wang 0001
NeurIPS1
2024 Less confidence, less forgetting: Learning with a humbler teacher in exemplar-free Class-Incremental learning
Zijian Gao, Kele Xu, Huiping Zhuang, Li Liu 0036, Xinjun Mao, Bo Ding 0001, Huaimin Wang 0001
Neural Networks1
2023 Complementary Learning System Based Intrinsic Reward in Reinforcement Learning
abstract
Deep reinforcement learning has achieved encouraging performance in many realms. However, one of its primary challenges is the sparsity of extrinsic rewards, which is still far from solved. Complementary learning system theory suggests that effective human learning relies on two complementary learning systems utilizing short-term and long-term memories. Inspired by the fact that humans evaluate curiosity by comparing current observations with historical information, we propose a novel intrinsic reward, namely CLS-IR, which aims to address the problems caused by sparse extrinsic rewards. Specifically, we train a self-supervised predictive model with short-term and long-term memories via exponential moving averages. We employ the information gain between the two memories as the intrinsic reward, which does not incur additional training costs but leads to better exploration. To investigate the effectiveness of CLS-IR, we conduct extensive experimental evaluations; the results demonstrate that CLS-IR can achieve state-of-the-art performance on Atari games and DeepMind Control Suite.
Zijian Gao, Kele Xu, Hongda Jia, Tianjiao Wan, Bo Ding 0001, Xinjun Mao, Huaimin Wang 0001
ICASSP1
2023 Diversifying Message Aggregation in Multi-Agent Communication Via Normalized Tensor Nuclear Norm Regularization
abstract
The use of graph attention networks (GAT) in communication-enhanced multi-agent reinforcement learning (Comm-MARL) has become prevalent. While successful, GAT can lead to homogeneity in the strategies of message aggregation, which can severely limit multi-agent coordination. To address this challenge, we study the adjacency tensor of the communication graph. Then we define a new nuclear tensor rank and its convex surrogate, the normalized tensor nuclear norm to measure the homogeneity of message aggregation. Leveraging the norm, we further propose a plug-and-play regularizer on the adjacency tensor, named Normalized Tensor Nuclear Norm Regularization (NTNNR), to actively enrich the diversity of message aggregation during the training stage. NTNNR is agnostic to specific Comm-MARL algorithms and can be flexibly integrated with different graph-attention methods. Empirical results demonstrate that aggregating messages using NTNNR-enhanced GAT can improve the efficiency of the training and achieve higher asymptotic performance than existing message aggregation methods.
Yuanzhao Zhai, Kele Xu, Bo Ding 0001, Zijian Gao, Huaimin Wang 0001
ICASSP5
2023 Bi-level Multi-Agent Actor-Critic Methods with ransformers
abstract
Recently, deep multi-agent reinforcement learning methods have witnessed great progress, including multi-agent actor-critic methods. However, it’s worth noticing there is a performance gap between multi-agent actor-critic methods and state-of-the-art value-based methods. In this paper, we investigate the causes and attribute inferior performance to issues of contribution-mismatch and indiscriminate guidance. To overcome these problems, we introduce a novel bi-level multi-agent actorcritic reinforcement learning approach with transformers, called BMT. Specifically, we propose a simple but efficient bi-level optimization mechanism to learn both global critic and agentspecific critic, thus jointly guiding the policy update. In addition, we adopt the transformer-based model as the policy network to decouple complicated relationships and generate flexible policy. BMT is also general enough to be plugged into any actor-critic multi-agent reinforcement learning approach, such as MAPPO, and equips it with strong expression. On multiple benchmarks including multi-agent particle environments and a challenging set of StarCraft II micromanagement tasks, large-scale empirical experiments demonstrate that BMT-based multi-agent reinforcement learning methods achieve superior performance over both state-of-the-art actor-critic and value-based approaches.
Tianjiao Wan, Haibo Mi, Zijian Gao, Yuanzhao Zhai, Bo Ding 0001
JCC3
2023 Multi-scale confusion and filling mechanism for pressure footprint recognition
Yan Zhang 0106, Yongsheng Sun, Nian Wang 0002, Zijian Gao, Jun Tang 0007
Neural Comput. Appl.4
2022 Uncertainty Estimation based Intrinsic Reward For Efficient Reinforcement Learning
abstract
For reinforcement learning, the extrinsic reward is a core factor for the learning process which however can be very sparse or completely missing. In response, researchers have proposed the idea of intrinsic reward, such as encouraging the agent to visit novel states through prediction error. However, the deep prediction model can provide over-confident and miscalibrated predictions. To mitigate the impact of inaccurate prediction, previous research applied deep ensembles and achieved superior results, despite the increased computation and storage space. In this paper, inspired by the uncertainty estimation, we leverage Monte Carlo Dropout to generate intrinsic reward from the perspective of uncertainty estimation with the goal to decrease the demands for computing resources while retaining superior performance. Utilizing the simple yet effective approach, we conduct extensive experiments across a variety of benchmark environments. The experimental results suggest that our method provides a competitive performance in final score and is faster in running speed, while requiring much fewer computing resources and storage space.
Tianjiao Wan, Peichang Shi, Bo Ding 0001, Zijian Gao
JCC5
2021 Multi-Actor-Attention-Critic Reinforcement Learning for Central Place Foraging Swarms
abstract
Multiple agents with relatively low cost, decentralized control, and robustness have the advantages of completing a foraging task more efficiently than a single advanced robot. Despite many foraging algorithms are efficient in multiple robot systems, most are pre-designed or not very adaptive to different environments since they have to evolve the parameters of the foraging algorithm in each different environment. Besides, designing an efficient collision avoidance strategy for multiple agents is a challenge. Addressing these issues, we introduce the multi-actor-attention-critic(MAAC) reinforcement learning method into the multiple foraging agents. We train the foraging strategy for multiple simulated agents. We compare our approach with existing foraging algorithms for multiple robots, the Central Place Foraging Algorithm (CPFA) and the Distributed Deterministic Spiral Algorithm (DDSA). Experimental results demonstrate that our approach outperforms the two algorithms. Also, we illustrate that our approach has a better performance in avoiding obstacles and adapting to different environments.
Kele Xu, Bo Ding 0001, Zijian Gao
IJCNN5
2021 Batch Weighted Nuclear-Norm Minimization for Medical Image Sequence Segmentation
Kele Xu, Zijian Gao, Jilong Wang 0007, Ming Feng
ISBRA2
2021 MSEC: Multi-Scale Erasure and Confusion for fine-grained image classification
Yan Zhang 0106, Yongsheng Sun, Nian Wang 0002, Zijian Gao, Jun Tang 0007
Neurocomputing4
2020 Security-Driven hybrid collaborative recommendation method for cloud-based iot services
Shunmei Meng, Zijian Gao, Qianmu Li, Hao Wang 0003, Hongning Dai, Lianyong Qi
Comput. Secur.2