Jiangmeng Li

dblp:293/0997 · DBLP profile ↗
← Back
57ranked-venue papers
11as first author
57since 2021 · last 2026
0000-0002-3376-1522ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 47 · 10 first-author · 47 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 2 first-author · 18 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author · 6 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 HTG-GCL: Leveraging Hierarchical Topological Granularity from Cellular Complexes for Graph Contrastive Learning
abstract
Graph contrastive learning (GCL) aims to learn discriminative semantic invariance by contrasting different views of the same graph that share critical topological patterns. However, existing GCL approaches with structural augmentations often struggle to identify task-relevant topological structures, let alone adapt to the varying coarse-to-fine topological granularities required across different downstream tasks. To remedy this issue, we introduce Hierarchical Topological Granularity Graph Contrastive Learning (HTG-GCL), a novel framework that leverages transformations of the same graph to generate multi-scale ring-based cellular complexes, embodying the concept of topological granularity, thereby generating diverse topological views. Recognizing that a certain granularity may contain misleading semantics, we propose a multi-granularity decoupled contrast and apply a granularity-specific weighting mechanism based on uncertainty estimation. Comprehensive experiments on various benchmarks demonstrate the effectiveness of HTG-GCL, highlighting its superior performance in capturing meaningful graph representations through hierarchical topological information.
Qirui Ji, Bin Qin 0001, Yunze Zhao, Chuxiong Sun, Changwen Zheng, Jianwen Cao 0001, Jiangmeng Li
AAAI8
2026 Doubly Debiased Test-Time Prompt Tuning for Vision-Language Models
abstract
Test-time prompt tuning for vision-language models has demonstrated impressive generalization capabilities under zero-shot settings. However, tuning the learnable prompts solely based on unlabeled test data may induce prompt optimization bias, ultimately leading to suboptimal performance on downstream tasks. In this work, we analyze the underlying causes of prompt optimization bias from both the model and data perspectives. In terms of the model, the entropy minimization objective typically focuses on reducing the entropy of model predictions while overlooking their correctness. This can result in overconfident yet incorrect outputs, thereby compromising the quality of prompt optimization. On the data side, prompts affected by optimization bias can introduce misalignment between visual and textual modalities, which further aggravates the prompt optimization bias. To this end, we propose a Doubly Debiased Test-Time Prompt Tuning method, abbreviated as D2TPT. Specifically, we first introduce a dynamic retrieval-augmented modulation module that retrieves high-confidence knowledge from a dynamic knowledge base using the test image feature as a query, and uses the retrieved knowledge to modulate the predictions. Guided by the refined predictions, we further develop a reliability-aware prompt optimization module that incorporates a confidence-based weighted ensemble and cross-modal consistency distillation to impose regularization constraints during prompt tuning. Extensive experiments across 15 benchmark datasets involving both natural distribution shifts and cross-datasets generalization demonstrate that D2TPT outperforms baselines, validating its effectiveness in mitigating prompt optimization bias.
Rui Wang 0079, Jiahuan Zhou, Changwen Zheng, Jiangmeng Li
AAAI6
2026 M2I2: Learning Efficient Multi-Agent Communication via Masked State Modeling and Intention Inference
abstract
Communication is essential in coordinating the behaviors of multiple agents. However, existing methods primarily emphasize content, timing, and partners for information sharing, often neglecting the critical aspect of integrating shared information. This gap can significantly impact agents' ability to understand and respond to complex, uncertain interactions, thus affecting overall communication efficiency. To address this issue, we introduce M2I2, a novel framework designed to enhance the agents' capabilities to assimilate and utilize received information effectively. M2I2 equips agents with advanced capabilities for masked state modeling and joint-action prediction, enriching their perception of environmental uncertainties and facilitating the anticipation of teammates' intentions. This approach ensures that agents are furnished with both comprehensive and relevant information, bolstering more informed and synergistic behaviors. Moreover, we propose a Dimensional Rational Network, innovatively trained via a meta-learning paradigm, to identify the importance of dimensional pieces of information, evaluating their contributions to decision-making and auxiliary tasks. Then, we implement an importance-based heuristic for selective information masking and sharing. This strategy optimizes the efficiency of masked state modeling and the rationale behind information sharing. We evaluate M2I2 across diverse multi-agent tasks, the results demonstrate its superior performance, efficiency, and generalization capabilities, over existing state-of-the-art methods in various complex scenarios.
Chuxiong Sun, Qirui Ji, Zehua Zang, Jiangmeng Li, Rui Wang 0079, Wei Wang 0353
AAAI5
2026 TMAE: Learning Targeted Multi-Agent Exploration via Causal Inference
abstract
Exploration in sparse-reward tasks remains a fundamental challenge in multi-agent reinforcement learning (MARL) due to complex inter-agent interactions and the expansive exploration space. To address this issue, we propose Targeted Multi-Agent Exploration (TMAE), a novel framework that uncovers the causal relationships between the state space and the reward function, thereby reducing the exploration space and enabling more targeted exploration. Specifically, we construct a structural causal model (SCM) to model the causality between sub-state variables and sparse rewards, providing a robust analytical foundation for subsequent causal inference. Through counterfactual causal intervention, TMAE identifies the most critical subspaces for discovering rare but pivotal events while filtering out confounders. By incorporating these causal insights into the exploration process, TMAE prioritizes subspaces with stronger causal effects on sparse rewards, significantly enhancing exploration efficiency. We evaluate TMAE on a range of MARL benchmarks featuring sparse rewards, consistently demonstrating superior exploration efficiency compared to state-of-the-art methods. Furthermore, visualized causal insights derived from TMAE reveal its ability to effectively capture intricate dependencies and priorities in targeted exploration, showcasing strong alignment with prior domain knowledge.
Chuxiong Sun, Dunqi Yao, Rui Wang 0079, Wenwen Qiang, Changwen Zheng, Jiangmeng Li
AAAI6
2026 Uniformity Preserving Transfer for Visual Prompt Tuning under Long-tailed Distribution
Hao Chen 0102, Bin Qin 0001, Jiangmeng Li, Jindong Wang 0001, Bing Su 0001
Int. J. Comput. Vis.4
2026 AmPLe: Supporting Vision-Language Models via Adaptive-Debiased Ensemble Multi-Prompt Learning
Jiangmeng Li, Rui Wang 0079, Changwen Zheng, Fanjiang Xu, Hui Xiong 0001
Int. J. Comput. Vis.3
2026 Towards continual low-light image enhancement through causal inference
Fan Ji, Jiangmeng Li, Xiongxin Tang, Fanjiang Xu
Neural Networks3
2026 Visual reinforcement learning via sequential consistency preserved policy contrast from optimal transport view
Zehua Zang, Jiangmeng Li, Chuxiong Sun, Rui Wang 0079, Fuchun Sun 0001
Neural Networks2
2026 On the Transferability and Discriminability of Representation Learning in Unsupervised Domain Adaptation
abstract
In this paper, we addressed the limitation of relying solely on distribution alignment and source-domain empirical risk minimization in Unsupervised Domain Adaptation (UDA). Our information-theoretic analysis showed that this standard adversarial-based framework neglects the discriminability of target-domain features, leading to suboptimal performance. To bridge this theoretical-practical gap, we defined "good representation learning" as guaranteeing both transferability and discriminability, and proved that an additional loss term targeting target-domain discriminability is necessary. Building on these insights, we proposed a novel adversarial-based UDA framework that explicitly integrates a domain alignment objective with a discriminability-enhancing constraint. Instantiated as Domain-Invariant Representation Learning with Global and Local Consistency (RLGLC), our method leverages Asymmetrically-Relaxed Wasserstein of Wasserstein Distance (AR-WWD) to address class imbalance and semantic dimension weighting, and employs a local consistency mechanism to preserve fine-grained target-domain discriminative information. Extensive experiments across multiple benchmark datasets demonstrate that RLGLC consistently surpasses state-of-the-art methods, confirming the value of our theoretical perspective and underscoring the necessity of enforcing both transferability and discriminability in adversarial-based UDA.
Wenwen Qiang, Ziyin Gu, Lingyu Si, Jiangmeng Li, Changwen Zheng, Fuchun Sun 0001, Hui Xiong 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2026 All-in-One Image Restoration via Causal-Deconfounding Wavelet-Disentangled Prompt Network
abstract
Image restoration represents a promising approach for addressing the inherent defects of image content distortion. Standard image restoration approaches suffer from high storage cost and the requirement towards the known degradation pattern, including type and degree, which can barely be satisfied in dynamic practical scenarios. In contrast, all-in-one image restoration (AiOIR) eliminates multiple degradations within a unified model to circumvent the aforementioned issues. However, according to our causal analysis, we disclose that two significant defects still exacerbate the effectiveness and generalization of AiOIR models: 1) the spurious correlation between non-degradation semantic features and degradation patterns; 2) the biased estimation of degradation patterns. To obtain the true causation between degraded images and restored images, we propose Causal-deconfounding Wavelet-disentangled Prompt Network (CWP-Net) to perform effective AiOIR. CWP-Net introduces two modules for decoupling, i.e., wavelet attention module of encoder and wavelet attention module of decoder. These modules explicitly disentangle the degradation and semantic features to tackle the issue of spurious correlation. To address the issue stemming from the biased estimation of degradation patterns, CWP-Net leverages a wavelet prompt block to generate the alternative variable for causal deconfounding. Extensive experiments on two all-in-one settings prove the effectiveness and superior performance of our proposed CWP-Net over the state-of-the-art AiOIR methods.
Bin Qin 0001, Jiangmeng Li, Fanjiang Xu, Fuchun Sun 0001, Hui Xiong 0001
IEEE Trans. Image Process.3
2025 MAP: Supporting Multimodal Knowledge Graph Completion via Augmented Modality Alignment and Instance Preserving
abstract
Multimodal knowledge graphs (KGs) have found widespread applications in data integration and processing, yet existing multimodal knowledge graphs are often highly incomplete, which impedes their wide adoption. Thereby multimodal knowledge graph completion (MKGC) has attracted widespread attention. However, the heterogeneity of multiple modalities degenerates the representations’ capacity to model modalityshared discriminative knowledge. The state-of-the-art approach addresses this challenge by aligning the modality distributions by adopting a Sinkhorn-based approach, but such an approach is computationally expensive and the practical sampling strategy largely degrades the model performance. Therefore we propose the augmented modality distribution alignment module, which imposes the generalized Radon transform-based approach to perform efficient and accurate distribution alignment. Yet the alignment may result in undesirable instance-level feature structure disorder. We thus propose the relation-aware instance preserving module. Empirical comparisons on well-established MKGC benchmarks demonstrate the effectiveness of proposed method.
Qingmeng Zhu, Changwen Zheng, Jiangmeng Li
ICASSP5
2025 Self-Reinforcing Prototype Evolution with Dual-Knowledge Cooperation for Semi-Supervised Lifelong Person Re-Identification
abstract
Current lifelong person re-identification (LReID) methods predominantly rely on fully labeled data streams. However, in real-world scenarios where annotation resources are limited, a vast amount of unlabeled data coexists with scarce labeled samples, leading to the Semi-Supervised LReID (Semi-LReID) problem where LReID methods suffer severe performance degradation. Existing LReID methods, even when combined with semi-supervised strategies, suffer from limited long-term adaptation performance due to struggling with the noisy knowledge occurring during unlabeled data utilization. In this paper, we pioneer the investigation of Semi-LReID, introducing a novel Self-Reinforcing Prototype Evolution with Dual-Knowledge Cooperation framework (SPRED). Our key innovation lies in establishing a self-reinforcing cycle between dynamic prototype-guided pseudo-label generation and new-old knowledge collaborative purification to enhance the utilization of unlabeled data. Specifically, learnable identity prototypes are introduced to dynamically capture the identity distributions and generate high-quality pseudo-labels. Then, the dual-knowledge cooperation scheme integrates current model specialization and historical model generalization, refining noisy pseudo-labels. Through this cyclic design, reliable pseudo-labels are progressively mined to improve current-stage learning and ensure positive knowledge propagation over long-term learning. Experiments on the established Semi-LReID benchmarks show that our SPRED achieves state-of-the-art performance. Our source code is available at https://github.com/zhoujiahuan1991/ICCV2025-SPRED
Kunlun Xu, Fan Zhuo, Jiangmeng Li, Xu Zou 0002, Jiahuan Zhou
ICCV3
2025 DenoiseVAE: Learning Molecule-Adaptive Noise Distributions for Denoising-based 3D Molecular Pre-training
abstract
Denoising learning of 3D molecules learns molecular representations by imposing noises into the equilibrium conformation and predicting the added noises to recover the equilibrium conformation, which essentially captures the information of molecular force fields. Due to the specificity of Potential Energy Surfaces, the probabilities of physically reasonable noises for each atom in different molecules are different. However, existing methods apply the shared heuristic hand-crafted noise sampling strategy to all molecules, resulting in inaccurate force field learning. In this paper, we propose a novel 3D molecular pre-training method, namely DenoiseVAE, which employs a Noise Generator to acquire atom-specific noise distributions for different molecules. It utilizes the stochastic reparameterization technique to sample noisy conformations from the generated distributions, which are inputted into a Denoising Module for denoising. The Noise Generator and the Denoising Module are jointly learned in a manner conforming with the paradigm of Variational Auto Encoder. Consequently, the sampled noisy conformations can be more diverse, adaptive, and informative, and thus DenoiseVAE can learn representations that better reveal the molecular force fields. Extensive experiments show that DenoiseVAE outperforms the current state-of-the-art methods on various molecular property prediction tasks, demonstrating the effectiveness of it.
Yurou Liu, Jiangmeng Li, Wenbing Huang 0001, Bing Su 0001
ICLR4
2025 Learning Adaptive High-Frequency Semantic Guidance for Low-light Image Enhancement
abstract
The low-light image enhancement has always been an important yet challenging task, which attracts significant attention in many fields. However, prior methods either ignore integrating semantic priors or depend on the masks generated by the pre-trained segmentation model. This way is complex and inevitably leads to inaccurate masks when facing unseen scenarios, which may be incompatible with the original feature and result in suboptimal performance. To address this issue, we first consider the high-frequency physical prior is more related to structural and textural properties, which embrace the rich semantic clues and can adaptively assist the learning process under various scenarios. Inspired by this, we propose the high-frequency semantic-aware guidance framework (HighFreNet) to leverage the guidance of semantic information tailored for enhancing low-light images. Specifically, the core parts are the novel Frequency-based Semantic Embedding Module (FSEM) and the Spatial-based Semantic Embedding Module (SSEM), which are separately designed to fully exploit the structure knowledge to modulate the original representation from frequency and spatial perspectives. Extensive experiments showcase that our method significantly outperforms the state-of-the-art methods on five benchmark datasets both in natural and remote sensing environments.
Jingxuan Zhou, Jiangmeng Li, Xiongxin Tang, Fanjiang Xu
ICME4
2025 RBDN: A Robust Background Denoising Network for Weakly Supervised Temporal Language Grounding
abstract
Temporal Language Grounding (TLG) with weak supervision aims to retrieve events in untrimmed videos corresponding to text queries using only video-text pairs as annotations. However, noisy background semantics in videos lead to inconsistencies in cross-modal alignment, which further hinder the grounding of query events. To address this, we introduce a Robust Background Denoising Network (RBDN), which refines backgrounds and mitigates their negative impact on events. RBDN leverages robust PCA and Frequency Augmentation techniques to filter out spatially static and temporally regular background information, thereby denoising the undesired gradient influence. Extensive experiments on well-known datasets, such as Charades-Sta and ActivityNet-Captions, demonstrate that our approach significantly outperforms current benchmarks.
Zehua Zang, Hongzhou Wu, Jiangmeng Li
ICME5
2025 Causal Deconfounding for Spurious Correlation in Domain Generalization
abstract
Existing machine learning techniques often blindly exploit data correlation, which may result in learning domain-dependent spurious correlation, thereby exacerbating the challenges of model generalization in out-of-distribution (OOD) scenarios. However, the essential definition of spurious correlation lacks causal mechanism analysis, leaving the manner in which they degrade models’ OOD generalization capabilities unclear. To this end, we construct a structural causal model (SCM) for the representation learning process and uncover a key insight: the deficiency of models in OOD generalization stems from confounding bias induced by spurious correlation. This bias misguides the model to rely on domain-sensitive correlation between spurious features and labels. In this regard, we leverage the modeled SCM graph to guide the adjustment of spurious features and implement backdoor adjustment. Then we further propose to control confounding bias by introducing a propensity score weighted estimator, which can be integrated into any existing OOD method as a plug-and-play module. The empirical results comprehensively demonstrate the effectiveness of our method on synthetic and large-scale real OOD datasets.
Bin Qin 0001, Jiangmeng Li, Xuesong Wu 0005, Yupeng Wang 0005, Jianwen Cao 0001
ICME3
2025 Rethinking the Bias of Foundation Model under Long-tailed Distribution
abstract
Long-tailed learning has garnered increasing attention due to its practical significance. Among the various approaches, the fine-tuning paradigm has gained considerable interest with the advent of foundation models. However, most existing methods primarily focus on leveraging knowledge from these models, overlooking the inherent biases introduced by the imbalanced training data they rely on. In this paper, we examine how such imbalances from pre-training affect long-tailed downstream tasks. Specifically, we find the imbalance biases inherited in foundation models on downstream task as parameter imbalance and data imbalance. During fine-tuning, we observe that parameter imbalance plays a more critical role, while data imbalance can be mitigated using existing re-balancing strategies. Moreover, we find that parameter imbalance cannot be effectively addressed by current re-balancing techniques, such as adjusting the logits, during training, unlike data imbalance. To tackle both imbalances simultaneously, we build our method on causal learning and view the incomplete semantic factor as the confounder, which brings spurious correlations between input samples and labels. To resolve the negative effects of this, we propose a novel backdoor adjustment method that learns the true causal effect between input samples and labels, rather than merely fitting the correlations in the data. Notably, we achieve an average performance increase of about 1.67% on each dataset.
Bin Qin 0001, Jiangmeng Li, Hao Chen 0102, Bing Su 0001
ICML3
2025 On the Out-of-Distribution Generalization of Self-Supervised Learning
abstract
In this paper, we focus on the out-of-distribution (OOD) generalization of self-supervised learning (SSL). By analyzing the mini-batch construction during the SSL training phase, we first give one plausible explanation for SSL having OOD generalization. Then, from the perspective of data generation and causal inference, we analyze and conclude that SSL learns spurious correlations during the training process, which leads to a reduction in OOD generalization. To address this issue, we propose a post-intervention distribution (PID) grounded in the Structural Causal Model. PID offers a scenario where the spurious variable and label variable is mutually independent. Besides, we demonstrate that if each mini-batch during SSL training satisfies PID, the resulting SSL model can achieve optimal worst-case OOD performance. This motivates us to develop a batch sampling strategy that enforces PID constraints through the learning of a latent variable model. Through theoretical analysis, we demonstrate the identifiability of the latent variable model and validate the effectiveness of the proposed sampling strategy. Experiments conducted on various downstream OOD tasks demonstrate the effectiveness of the proposed sampling strategy.
Wenwen Qiang, Zeen Song, Jiangmeng Li, Changwen Zheng
ICML4
2025 Learning Invariant Causal Mechanism from Vision-Language Models
abstract
Contrastive Language-Image Pretraining (CLIP) has achieved remarkable success, but its performance can degrade when fine-tuned in out-of-distribution (OOD) scenarios. We model the prediction process using a Structural Causal Model (SCM) and show that the causal mechanism involving both invariant and variant factors in training environments differs from that in test environments. In contrast, the causal mechanism with solely invariant factors remains consistent across environments. We theoretically prove the existence of a linear mapping from CLIP embeddings to invariant factors, which can be estimated using interventional data. Additionally, we provide a condition to guarantee low OOD risk of the invariant predictor. Based on these insights, we propose the Invariant Causal Mechanism of CLIP (CLIP-ICM) framework. CLIP-ICM involves collecting interventional data, estimating a linear projection matrix, and making predictions within the invariant subspace. Experiments on several OOD datasets show that CLIP-ICM significantly improves the performance of CLIP. Our method offers a simple but powerful enhancement, boosting the reliability of CLIP in real-world applications.
Zeen Song, Jiangmeng Li, Changwen Zheng, Wenwen Qiang
ICML4
2025 Towards the Causal Complete Cause of Multi-Modal Representation Learning
abstract
Multi-Modal Learning (MML) aims to learn effective representations across modalities for accurate predictions. Existing methods typically focus on modality consistency and specificity to learn effective representations. However, from a causal perspective, they may lead to representations that contain insufficient and unnecessary information. To address this, we propose that effective MML representations should be causally sufficient and necessary. Considering practical issues like spurious correlations and modality conflicts, we relax the exogeneity and monotonicity assumptions prevalent in prior works and explore the concepts specific to MML, i.e., Causal Complete Cause ($C^3$). We begin by defining $C^3$, which quantifies the probability of representations being causally sufficient and necessary. We then discuss the identifiability of $C^3$ and introduce an instrumental variable to support identifying $C^3$ with non-exogeneity and non-monotonicity. Building on this, we conduct the $C^3$ measurement, i.e., $C^3$ risk. We propose a twin network to estimate it through (i) the real-world branch: utilizing the instrumental variable for sufficiency, and (ii) the hypothetical-world branch: applying gradient-based counterfactual modeling for necessity. Theoretical analyses confirm its reliability. Based on these results, we propose $C^3$ Regularization, a plug-and-play method that enforces the causal completeness of the learned representations by minimizing $C^3$ risk. Extensive experiments demonstrate its effectiveness.
Wenwen Qiang, Jiangmeng Li, Changwen Zheng, Fuchun Sun 0001, Hui Xiong 0001
ICML4
2025 CellCLAT: Preserving Topology and Trimming Redundancy in Self-Supervised Cellular Contrastive Learning
abstract
Self-supervised topological deep learning (TDL) represents a nascent but underexplored area with significant potential for modeling higher-order interactions in simplicial complexes and cellular complexes to derive representations of unlabeled graphs. Compared to simplicial complexes, cellular complexes exhibit greater expressive power. However, the advancement in self-supervised learning for cellular TDL is largely hindered by two core challenges: extrinsic structural constraints inherent to cellular complexes, and intrinsic semantic redundancy in cellular representations. The first challenge highlights that traditional graph augmentation techniques may compromise the integrity of higher-order cellular interactions, while the second underscores that topological redundancy in cellular complexes potentially diminish task-relevant information. To address these issues, we introduce Cellular Complex Contrastive Learning with Adaptive Trimming (CellCLAT), a twofold framework designed to adhere to the combinatorial constraints of cellular complexes while mitigating informational redundancy. Specifically, we propose a parameter perturbation-based augmentation method that injects controlled noise into cellular interactions without altering the underlying cellular structures, thereby preserving cellular topology during contrastive learning. Additionally, a cellular trimming scheduler is employed to mask gradient contributions from task-irrelevant cells through a bi-level meta-learning approach, effectively removing redundant topological elements while maintaining critical higher-order semantics. We provide theoretical justification and empirical validation to demonstrate that CellCLAT achieves substantial improvements over existing self-supervised graph learning methods, marking a significant attempt in this domain.
Bin Qin 0001, Qirui Ji, Jiangmeng Li, Yupeng Wang 0005, Xuesong Wu 0005, Jianwen Cao 0001, Fanjiang Xu
KDD (2)3
2025 C2Prompt: Class-aware Client Knowledge Interaction for Federated Continual Learning
abstract
Federated continual learning (FCL) tackles scenarios of learning from continuously emerging task data across distributed clients, where the key challenge lies in addressing both temporal forgetting over time and spatial forgetting simultaneously. Recently, prompt-based FCL methods have shown advanced performance through task-wise prompt communication. In this study, we underscore that the existing prompt-based FCL methods are prone to class-wise knowledge coherence between prompts across clients. The class-wise knowledge coherence includes two aspects: (1) intra-class distribution gap across clients, which degrades the learned semantics across prompts, (2) inter-prompt class-wise relevance, which highlights cross-class knowledge confusion. During prompt communication, insufficient class-wise coherence exacerbates knowledge conflicts among new prompts and induces interference with old prompts, intensifying both spatial and temporal forgetting. To address these issues, we propose a novel Class-aware Client Knowledge Interaction (C$^2$Prompt) method that explicitly enhances class-wise knowledge coherence during prompt communication. Specifically, a local class distribution compensation mechanism (LCDC) is introduced to reduce intra-class distribution disparities across clients, thereby reinforcing intra-class knowledge consistency. Additionally, a class-aware prompt aggregation scheme (CPA) is designed to alleviate inter-class knowledge confusion by selectively strengthening class-relevant knowledge aggregation. Extensive experiments on multiple FCL benchmarks demonstrate that C$^2$Prompt achieves state-of-the-art performance. Our code will be released.
Kunlun Xu, Yibo Feng, Jiangmeng Li, Yongsheng Qi, Jiahuan Zhou
NeurIPS3
2025 Continual Test-Time Adaptation for Single Image Defocus Deblurring via Causal Siamese Networks
Jiangmeng Li, Xiongxin Tang, Bing Su 0001, Fanjiang Xu, Hui Xiong 0001
Int. J. Comput. Vis.3
2025 Rethinking Generalizability and Discriminability of Self-Supervised Learning from Evolutionary Game Theory Perspective
Jiangmeng Li, Zehua Zang, Qirui Ji, Chuxiong Sun, Wenwen Qiang, Junge Zhang, Changwen Zheng, Fuchun Sun 0001, Hui Xiong 0001
Int. J. Comput. Vis.1
2025 On the Generalization and Causal Explanation in Self-Supervised Learning
Wenwen Qiang, Zeen Song, Ziyin Gu, Jiangmeng Li, Changwen Zheng, Fuchun Sun 0001, Hui Xiong 0001
Int. J. Comput. Vis.4
2025 Learning Complementary Knowledge via Trusted Multi-view Space Decomposition for Self-Supervised Contrastive Learning
Jiangmeng Li, Yunze Zhao, Changwen Zheng, Wenwen Qiang
Mach. Learn.1
2025 Supporting vision-language model few-shot inference with confounder-pruned knowledge prompt
Jiangmeng Li, Wenyi Mo, Chuxiong Sun, Wenwen Qiang, Bing Su 0001, Changwen Zheng
Neural Networks1
2025 Intervening on few-shot object detection based on the front-door criterion
Jiangmeng Li, Qirui Ji, Changwen Zheng, Wenwen Qiang
Neural Networks2
2024 Rethinking Causal Relationships Learning in Graph Neural Networks
abstract
Graph Neural Networks (GNNs) demonstrate their significance by effectively modeling complex interrelationships within graph-structured data. To enhance the credibility and robustness of GNNs, it becomes exceptionally crucial to bolster their ability to capture causal relationships. However, despite recent advancements that have indeed strengthened GNNs from a causal learning perspective, conducting an in-depth analysis specifically targeting the causal modeling prowess of GNNs remains an unresolved issue. In order to comprehensively analyze various GNN models from a causal learning perspective, we constructed an artificially synthesized dataset with known and controllable causal relationships between data and labels. The rationality of the generated data is further ensured through theoretical foundations. Drawing insights from analyses conducted using our dataset, we introduce a lightweight and highly adaptable GNN module designed to strengthen GNNs' causal learning capabilities across a diverse range of tasks. Through a series of experiments conducted on both synthetic datasets and other real-world datasets, we empirically validate the effectiveness of the proposed module. The codes are available at https://github.com/yaoyao-yaoyao-cell/CRCG.
Hang Gao 0004, Chengyu Yao, Jiangmeng Li, Lingyu Si, Fengge Wu, Changwen Zheng, Huaping Liu 0001
AAAI3
2024 Rethinking Dimensional Rationale in Graph Contrastive Learning from Causal Perspective
abstract
Graph contrastive learning is a general learning paradigm excelling at capturing invariant information from diverse perturbations in graphs. Recent works focus on exploring the structural rationale from graphs, thereby increasing the discriminability of the invariant information. However, such methods may incur in the mis-learning of graph models towards the interpretability of graphs, and thus the learned noisy and task-agnostic information interferes with the prediction of graphs. To this end, with the purpose of exploring the intrinsic rationale of graphs, we accordingly propose to capture the dimensional rationale from graphs, which has not received sufficient attention in the literature. The conducted exploratory experiments attest to the feasibility of the aforementioned roadmap. To elucidate the innate mechanism behind the performance improvement arising from the dimensional rationale, we rethink the dimensional rationale in graph contrastive learning from a causal perspective and further formalize the causality among the variables in the pre-training stage to build the corresponding structural causal model. On the basis of the understanding of the structural causal model, we propose the dimensional rationale-aware graph contrastive learning approach, which introduces a learnable dimensional rationale acquiring network and a redundancy reduction constraint. The learnable dimensional rationale acquiring network is updated by leveraging a bi-level meta-learning technique, and the redundancy reduction constraint disentangles the redundant features through a decorrelation process during learning. Empirically, compared with state-of-the-art methods, our method can yield significant performance boosts on various benchmarks with respect to discriminability and transferability. The code implementation of our method is available at https://github.com/ByronJi/DRGCL.
Qirui Ji, Jiangmeng Li, Jie Hu 0019, Rui Wang 0079, Changwen Zheng, Fanjiang Xu
AAAI2
2024 Hierarchical Topology Isomorphism Expertise Embedded Graph Contrastive Learning
abstract
Graph contrastive learning (GCL) aims to align the positive features while differentiating the negative features in the latent space by minimizing a pair-wise contrastive loss. As the embodiment of an outstanding discriminative unsupervised graph representation learning approach, GCL achieves impressive successes in various graph benchmarks. However, such an approach falls short of recognizing the topology isomorphism of graphs, resulting in that graphs with relatively homogeneous node features cannot be sufficiently discriminated. By revisiting classic graph topology recognition works, we disclose that the corresponding expertise intuitively complements GCL methods. To this end, we propose a novel hierarchical topology isomorphism expertise embedded graph contrastive learning, which introduces knowledge distillations to empower GCL models to learn the hierarchical topology isomorphism expertise, including the graph-tier and subgraph-tier. On top of this, the proposed method holds the feature of plug-and-play, and we empirically demonstrate that the proposed method is universal to multiple state-of-the-art GCL models. The solid theoretical analyses are further provided to prove that compared with conventional GCL methods, our method acquires the tighter upper bound of Bayes classification error. We conduct extensive experiments on real-world benchmarks to exhibit the performance superiority of our method over candidate GCL methods, e.g., for the real-world graph representation learning experiments, the proposed method beats the state-of-the-art method by 0.23% on unsupervised representation learning setting, 0.43% on transfer learning setting. Our code is available at https://github.com/jyf123/HTML.
Jiangmeng Li, Hang Gao 0004, Wenwen Qiang, Changwen Zheng, Fuchun Sun 0001
AAAI1
2024 T2MAC: Targeted and Trusted Multi-Agent Communication through Selective Engagement and Evidence-Driven Integration
abstract
Communication stands as a potent mechanism to harmonize the behaviors of multiple agents. However, existing work primarily concentrates on broadcast communication, which not only lacks practicality, but also leads to information redundancy. This surplus, one-fits-all information could adversely impact the communication efficiency. Furthermore, existing works often resort to basic mechanisms to integrate observed and received information, impairing the learning process. To tackle these difficulties, we propose Targeted and Trusted Multi-Agent Communication (T2MAC), a straightforward yet effective method that enables agents to learn selective engagement and evidence-driven integration. With T2MAC, agents have the capability to craft individualized messages, pinpoint ideal communication windows, and engage with reliable partners, thereby refining communication efficiency. Following the reception of messages, the agents integrate information observed and received from different sources at an evidence level. This process enables agents to collectively use evidence garnered from multiple perspectives, fostering trusted and cooperative behaviors. We evaluate our method on a diverse set of cooperative multi-agent tasks, with varying difficulties, involving different scales and ranging from Hallway, MPE to SMAC. The experiments indicate that the proposed model not only surpasses the state-of-the-art methods in terms of cooperative performance and communication efficiency, but also exhibits impressive generalization.
Chuxiong Sun, Zehua Zang, Jiangmeng Li, Rui Wang 0079, Changwen Zheng
AAAI4
2024 Demo:SCDRL: Scalable and Customized Distributed Reinforcement Learning System
abstract
Reinforcement Learning (RL) has marked significant achievements across a variety of complex tasks in real-world scenarios. However, the efficacy of RL predominantly relies on the availability of extensive datasets and considerable training resources. Hence, there is the critical need for a distributed system capable of generating and processing vast amounts of data with efficiency. In this work, we introduce a Scalable and Customized Distributed Reinforcement Learning system (SCDRL). Concretely, we analyze the paradigm of RL and decouple the major RL computations into three main aspects, i.e. environment simulation, policy inference and policy training. Such decouple enables SCDRL to efficiently allocate computing resources (be it CPUs or GPUs of varying computational capabilities) tailored to the specific needs of each component. We demonstrate the effectiveness of our method across several key RL environments, demonstrating that our system not only achieves significant learning outcomes and enhanced throughput but also utilizes computing resources with greater efficiency. Notably, our findings reveal SCDRL's proficiency in optimizing resource use not just in single-machine setups but also in multi-machine configurations, all the while maintaining data efficiency and resource utilization.
Chuxiong Sun, Wenwen Qiang, Jiangmeng Li
ICDCS3
2024 BayesPrompt: Prompting Large-Scale Pre-Trained Language Models on Few-shot Inference via Debiased Domain Abstraction
abstract
As a novel and effective fine-tuning paradigm based on large-scale pre-trained language models (PLMs), prompt-tuning aims to reduce the gap between downstream tasks and pre-training objectives. While prompt-tuning has yielded continuous advancements in various tasks, such an approach still remains a persistent defect: prompt-tuning methods fail to generalize to specific few-shot patterns. From the perspective of distribution analyses, we disclose that the intrinsic issues behind the phenomenon are the over-multitudinous conceptual knowledge contained in PLMs and the abridged knowledge for target downstream domains, which jointly result in that PLMs mis-locate the knowledge distributions corresponding to the target domains in the universal knowledge embedding space. To this end, we intuitively explore to approximate the unabridged target domains of downstream tasks in a debiased manner, and then abstract such domains to generate discriminative prompts, thereby providing the de-ambiguous guidance for PLMs. Guided by such an intuition, we propose a simple yet effective approach, namely BayesPrompt, to learn prompts that contain the domain discriminative information against the interference from domain-irrelevant knowledge. BayesPrompt primitively leverages known distributions to approximate the debiased factual distributions of target domains and further uniformly samples certain representative features from the approximated distributions to generate the ultimate prompts for PLMs. We provide theoretical insights with the connection to domain adaptation. Empirically, our method achieves state-of-the-art performance on benchmarks.
Jiangmeng Li, Wenwen Qiang, Changwen Zheng, Fuchun Sun 0001, Hui Xiong 0001
ICLR1
2024 MSI: Multi-modal Recommendation via Superfluous Semantics Discarding and Interaction Preserving
abstract
Multi-modal recommendation aims at leveraging data of auxiliary modalities (e.g., linguistic descriptions and images) to enhance the representations of items, thereby accurately recommending items that users prefer from the vast expanse of Web-based data. Current multi-modal recommendation methods typically utilize multi-modal features to assist in learning item representations in a direct manner. However, the superfluous semantics in multi-modal features are ignored, resulting in the inclusion of excessive redundancy within the representations of items. Moreover, we disclose that multi-modal features of items rarely contain user-item interaction information. Hence, during the interaction among different item features, the user-item interaction information in ID-based representations diminishes, leading to the degeneration of recommendation performance. To this end, we propose a novel multi-modal recommendation approach, which compresses representations of extra modalities under the guidance of solid theoretical analysis and leverages two auxiliary multi-modal graphs to integrate user-item interaction information into multi-modal features. Empirical experiments on three multi-modal recommendation datasets demonstrate that our method outperforms benchmarks.
Qingmeng Zhu, Changwen Zheng, Jiangmeng Li
ICMR4
2024 Rethinking Misalignment in Vision-Language Model Adaptation from a Causal Perspective
abstract
Foundational Vision-Language models such as CLIP have exhibited impressive generalization in downstream tasks. However, CLIP suffers from a two-level misalignment issue, i.e., task misalignment and data misalignment, when adapting to specific tasks. Soft prompt tuning has mitigated the task misalignment, yet the data misalignment remains a challenge. To analyze the impacts of the data misalignment, we revisit the pre-training and adaptation processes of CLIP and develop a structural causal model. We discover that while we expect to capture task-relevant information for downstream tasks accurately, the task-irrelevant knowledge impacts the prediction results and hampers the modeling of the true relationships between the images and the predicted classes. As task-irrelevant knowledge is unobservable, we leverage the front-door adjustment and propose Causality-Guided Semantic Decoupling and Classification (CDC) to mitigate the interference of task-irrelevant knowledge. Specifically, we decouple semantics contained in the data of downstream tasks and perform classification based on each semantic. Furthermore, we employ the Dempster-Shafer evidence theory to evaluate the uncertainty of each prediction generated by diverse semantics. Experiments conducted in multiple different settings have consistently demonstrated the effectiveness of CDC.
Jiangmeng Li, Wenwen Qiang
NeurIPS2
2024 Introducing diminutive causal structure into graph representation learning
abstract
When engaging in end-to-end graph representation learning with Graph Neural Networks (GNNs), the intricate causal relationships and rules inherent in graph data pose a formidable challenge for the model in accurately capturing authentic data relationships. A proposed mitigating strategy involves the direct integration of rules or relationships corresponding to the graph data into the model. However, within the domain of graph representation learning, the inherent complexity of graph data obstructs the derivation of a comprehensive causal structure that encapsulates universal rules or relationships governing the entire dataset. Instead, only specialized diminutive causal structures, delineating specific causal relationships within constrained subsets of graph data, emerge as discernible. Motivated by empirical insights, it is observed that GNN models exhibit a tendency to converge towards such specialized causal structures during the training process. Consequently, we posit that the introduction of these specific causal structures is advantageous for the training of GNN models. Building upon this proposition, we introduce a novel method that enables GNN models to glean insights from these specialized diminutive causal structures, thereby enhancing overall performance. Our method specifically extracts causal knowledge from the model representation of these diminutive causal structures and incorporates interchange intervention to optimize the learning process. Theoretical analysis serves to corroborate the efficacy of our proposed method. Furthermore, empirical experiments consistently demonstrate significant performance improvements across diverse datasets.
Hang Gao 0004, Peng Qiao, Fengge Wu, Jiangmeng Li, Changwen Zheng
Knowl. Based Syst.5
2024 Unsupervised social event detection via hybrid graph contrastive learning and reinforced incremental clustering
abstract
Detecting events from social media data streams is gradually attracting researchers. The innate challenge for detecting events is to extract discriminative information from social media data thereby assigning the data into different events. Due to the excessive diversity and high updating frequency of social data, using supervised approaches to detect events from social messages is hardly achieved. To this end, recent works explore learning discriminative information from social messages by leveraging graph contrastive learning (GCL) and embedding clustering in an unsupervised manner. However, two intrinsic issues exist in benchmark methods: conventional GCL can only roughly explore partial attributes, thereby insufficiently learning the discriminative information of social messages; for benchmark methods, the learned embeddings are clustered in the latent space by taking advantage of certain specific prior knowledge , which conflicts with the principle of unsupervised learning paradigm . In this paper, we propose a novel unsupervised social media event detection method via hybrid graph contrastive learning and reinforced incremental clustering (HCRC), which uses hybrid graph contrastive learning to comprehensively learn semantic and structural discriminative information from social messages and reinforced incremental clustering to perform efficient clustering in a solidly unsupervised manner. We conduct comprehensive experiments to evaluate HCRC on the Twitter and Maven datasets. The experimental results demonstrate that our approach yields consistent significant performance boosts. In traditional incremental setting, semi-supervised incremental setting and solidly unsupervised setting, the model performance has achieved maximum improvements of 53%, 45%, and 37%, respectively.
Zehua Zang, Hang Gao 0004, Rui Wang 0086, Jiangmeng Li
Knowl. Based Syst.7
2024 Physics-Guided Optical Simulation and PSF Analysis for Remote Sensing Images Deblurring
abstract
The presence of blur is prevalent in satellite remote sensing images (RSIs), and its detrimental impact on downstream applications cannot be overlooked. Current deep learning approaches for image deblurring have gained substantial attention due to their effectiveness and fast inference speed. However, these methods often heavily rely on extensive paired training datasets and lack interpretability. Existing deblurring datasets primarily include regular scenes while remote sensing images exhibit distinct blurring mechanisms. Consequently, deep learning methods lacking prior physical knowledge can only tackle the image deblurring problem in specific scenarios, but hard to achieve satisfactory results on remote sensing images. To address these problems, it is essential to construct a remote sensing image dataset that incorporates the realistic causes of blurriness and integrate prior knowledge into the methods. In this work, we first analyze the satellite imaging system and use Zernike polynomials to approximate the optical aberrations to simulate the RSI blurring process which ensures the proposed dataset adhering solid physical principles. Moreover, we propose a novel physics-guided RSI deblurring (PGRSID) network that integrates an explicit Wiener deconvolution process in both spatial and deep feature space. This integration better leverages the physical interpretation to facilitate effective learning for the RSI deblurring network. We further incorporate denoise loss and cycle consistency loss in the objective function to facilitate the model’s learning process for RSI deblurring. Extensive experiments are conducted on both our synthetic dataset and real GF-1A/PMS data. Qualitative and quantitative experiment results highlight the effectiveness and superiority of our physics-guided deblurring network for satellite RSI.
Fan Ji, Jiangmeng Li, Xiongxin Tang, Fanjiang Xu
IEEE Trans. Geosci. Remote. Sens.4
2023 Robust Causal Graph Representation Learning against Confounding Effects
abstract
The prevailing graph neural network models have achieved significant progress in graph representation learning. However, in this paper, we uncover an ever-overlooked phenomenon: the pre-trained graph representation learning model tested with full graphs underperforms the model tested with well-pruned graphs. This observation reveals that there exist confounders in graphs, which may interfere with the model learning semantic information, and current graph representation learning methods have not eliminated their influence. To tackle this issue, we propose Robust Causal Graph Representation Learning (RCGRL) to learn robust graph representations against confounding effects. RCGRL introduces an active approach to generate instrumental variables under unconditional moment restrictions, which empowers the graph representation learning model to eliminate confounders, thereby capturing discriminative information that is causally related to downstream predictions. We offer theorems and proofs to guarantee the theoretical effectiveness of the proposed approach. Empirically, we conduct extensive experiments on a synthetic dataset and multiple benchmark datasets. Experimental results demonstrate the effectiveness and generalization ability of RCGRL. Our codes are available at https://github.com/hang53/RCGRL.
Hang Gao 0004, Jiangmeng Li, Wenwen Qiang, Lingyu Si, Changwen Zheng, Fuchun Sun 0001
AAAI2
2023 Disentangle and Remerge: Interventional Knowledge Distillation for Few-Shot Object Detection from a Conditional Causal Perspective
abstract
Few-shot learning models learn representations with limited human annotations, and such a learning paradigm demonstrates practicability in various tasks, e.g., image classification, object detection, etc. However, few-shot object detection methods suffer from an intrinsic defect that the limited training data makes the model cannot sufficiently explore semantic information. To tackle this, we introduce knowledge distillation to the few-shot object detection learning paradigm. We further run a motivating experiment, which demonstrates that in the process of knowledge distillation, the empirical error of the teacher model degenerates the prediction performance of the few-shot object detection model as the student. To understand the reasons behind this phenomenon, we revisit the learning paradigm of knowledge distillation on the few-shot object detection task from the causal theoretic standpoint, and accordingly, develop a Structural Causal Model. Following the theoretical guidance, we propose a backdoor adjustment-based knowledge distillation method for the few-shot object detection task, namely Disentangle and Remerge (D&R), to perform conditional causal intervention toward the corresponding Structural Causal Model. Empirically, the experiments on benchmarks demonstrate that D&R can yield significant performance boosts in few-shot object detection. Code is available at https://github.com/ZYN-1101/DandR.git.
Jiangmeng Li, Wenwen Qiang, Lingyu Si, Chengbo Jiao, Changwen Zheng, Fuchun Sun 0001
AAAI1
2023 M2HGCL: Multi-scale Meta-path Integrated Heterogeneous Graph Contrastive Learning
Rongcheng Duan, Jiangmeng Li
ADMA (3)6
2023 Meta Attention-Generation Network for Cross-Granularity Few-Shot Learning
Wenwen Qiang, Jiangmeng Li, Bing Su 0001, Jianlong Fu, Hui Xiong 0001, Ji-Rong Wen
Int. J. Comput. Vis.2
2023 Information theory-guided heuristic progressive multi-view coding
Jiangmeng Li, Hang Gao 0004, Wenwen Qiang, Changwen Zheng
Neural Networks1
2023 Modeling Multiple Views via Implicitly Preserving Global Consistency and Local Complementarity
abstract
While self-supervised learning techniques are often used to mine hidden knowledge from unlabeled data via modeling multiple views, it is unclear how to perform effective representation learning in a complex and inconsistent context. To this end, we propose a new multi-view self-supervised learning method, namelyconsistency and complementarity network(CoCoNet), to comprehensively learn global inter-view consistent and local cross-view complementarity-preserving representations from multiple views. To capture crucial common knowledge which is implicitly shared among views, CoCoNet employs a global consistency module that aligns the probabilistic distribution of views by utilizing an efficient discrepancy metric based on the generalized sliced Wasserstein distance. To incorporate cross-view complementary information, CoCoNet proposes a heuristic complementarity-aware contrastive learning approach, which extracts a complementarity-factor jointing cross-view discriminative knowledge and uses it as the contrast to guide the learning of view-specific encoders. Theoretically, the superiority of CoCoNet is verified by our information-theoretical-based analyses. Empirically, our thorough experimental results show that CoCoNet outperforms the state-of-the-art self-supervised methods by a significant margin, for instance, CoCoNet beats the best benchmark method by an average margin of 1.1% on ImageNet.
Jiangmeng Li, Wenwen Qiang, Changwen Zheng, Bing Su 0001, Farid Razzak, Ji-Rong Wen, Hui Xiong 0001
IEEE Trans. Knowl. Data Eng.1
2023 Robust Local Preserving and Global Aligning Network for Adversarial Domain Adaptation
abstract
Unsupervised domain adaptation (UDA) requires source domain samples with clean ground truth labels during training. Accurately labeling a large number of source domain samples is time-consuming and laborious. An alternative is to utilize samples with noisy labels for training. However, training with noisy labels can greatly reduce the performance of UDA. In this paper, we address the problem that learning UDA models only with access to noisy labels and propose a novel method called robust local preserving and global aligning network (RLPGA). RLPGA improves the robustness of the label noise from two aspects. One is learning a classifier by a robust informative-theoretic-based loss function. The other is constructing two adjacency weight matrices and two negative weight matrices by the proposed local preserving module to preserve the local topology structures of input data. We conduct theoretical analysis on the robustness of the proposed RLPGA and prove that the robust informative-theoretic-based loss and the local preserving module are beneficial to reduce the empirical risk of the target domain. A series of empirical studies show the effectiveness of our proposed RLPGA.
Wenwen Qiang, Jiangmeng Li, Changwen Zheng, Bing Su 0001, Hui Xiong 0001
IEEE Trans. Knowl. Data Eng.2
2022 Supporting Medical Relation Extraction via Causality-Pruned Semantic Dependency Forest
abstract
Medical Relation Extraction (MRE) task aims to extract relations between entities in medical texts. Traditional relation extraction methods achieve impressive success by exploring the syntactic information, e.g., dependency tree. However, the quality of the 1-best dependency tree for medical texts produced by an out-of-domain parser is relatively limited so that the performance of medical relation extraction method may degenerate. To this end, we propose a method to jointly model semantic and syntactic information from medical texts based on causal explanation theory. We generate dependency forests consisting of the semantic-embedded 1-best dependency tree. Then, a task-specific causal explainer is adopted to prune the dependency forests, which are further fed into a designed graph convolutional network to learn the corresponding representation for downstream task. Empirically, the various comparisons on benchmark medical datasets demonstrate the effectiveness of our model.
Jiangmeng Li, Chengbo Jiao
COLING2
2022 Weight-Aware Graph Contrastive Learning
Hang Gao 0004, Jiangmeng Li, Peng Qiao, Changwen Zheng
ICANN (2)2
2022 MetAug: Contrastive Learning via Meta Feature Augmentation
abstract
What matters for contrastive learning? We argue that contrastive learning heavily relies on informative features, or “hard” (positive or negative) features. Early works include more informative features by applying complex data augmentations and large batch size or memory bank, and recent works design elaborate sampling approaches to explore informative features. The key challenge toward exploring such features is that the source multi-view data is generated by applying random data augmentations, making it infeasible to always add useful information in the augmented data. Consequently, the informativeness of features learned from such augmented data is limited. In response, we propose to directly augment the features in latent space, thereby learning discriminative representations without a large amount of input data. We perform a meta learning technique to build the augmentation generator that updates its network parameters by considering the performance of the encoder. However, insufficient input data may lead the encoder to learn collapsed features and therefore malfunction the augmentation generator. A new margin-injected regularization is further added in the objective function to avoid the encoder learning a degenerate mapping. To contrast all features in one gradient back-propagation step, we adopt the proposed optimization-driven unified contrastive loss instead of the conventional contrastive loss. Empirically, our method achieves state-of-the-art results on several benchmark datasets.
Jiangmeng Li, Wenwen Qiang, Changwen Zheng, Bing Su 0001, Hui Xiong 0001
ICML1
2022 Interventional Contrastive Learning with Meta Semantic Regularizer
abstract
Contrastive learning (CL)-based self-supervised learning models learn visual representations in a pairwise manner. Although the prevailing CL model has achieved great progress, in this paper, we uncover an ever-overlooked phenomenon: When the CL model is trained with full images, the performance tested in full images is better than that in foreground areas; when the CL model is trained with foreground areas, the performance tested in full images is worse than that in foreground areas. This observation reveals that backgrounds in images may interfere with the model learning semantic information and their influence has not been fully eliminated. To tackle this issue, we build a Structural Causal Model (SCM) to model the background as a confounder. We propose a backdoor adjustment-based regularization method, namely Interventional Contrastive Learning with Meta Semantic Regularizer (ICL-MSR), to perform causal intervention towards the proposed SCM. ICL-MSR can be incorporated into any existing CL methods to alleviate background distractions from representation learning. Theoretically, we prove that ICL-MSR achieves a tighter error bound. Empirically, our experiments on multiple benchmark datasets demonstrate that ICL-MSR is able to improve the performances of different state-of-the-art CL methods.
Wenwen Qiang, Jiangmeng Li, Changwen Zheng, Bing Su 0001, Hui Xiong 0001
ICML2
2022 Bootstrapping Informative Graph Augmentation via A Meta Learning Approach
abstract
Recent works explore learning graph representations in a self-supervised manner. In graph contrastive learning, benchmark methods apply various graph augmentation approaches. However, most of the augmentation methods are non-learnable, which causes the issue of generating unbeneficial augmented graphs. Such augmentation may degenerate the representation ability of graph contrastive learning methods. Therefore, we motivate our method to generate augmented graph with a learnable graph augmenter, called MEta Graph Augmentation (MEGA). We then clarify that a "good" graph augmentation must have uniformity at the instance-level and informativeness at the feature-level. To this end, we propose a novel approach to learning a graph augmenter that can generate an augmentation with uniformity and informativeness. The objective of the graph augmenter is to promote our feature extraction network to learn a more discriminative feature representation, which motivates us to propose a meta-learning paradigm. Empirically, the experiments across multiple benchmark datasets demonstrate that MEGA outperforms the state-of-the-art methods in graph self-supervised learning tasks. Further experimental studies prove the effectiveness of different terms of MEGA. Our codes are available at https://github.com/hang53/MEGA.
Hang Gao 0004, Jiangmeng Li, Wenwen Qiang, Lingyu Si, Fuchun Sun 0001, Changwen Zheng
IJCAI2
2022 MetaMask: Revisiting Dimensional Confounder for Self-Supervised Learning
abstract
As a successful approach to self-supervised learning, contrastive learning aims to learn invariant information shared among distortions of the input sample. While contrastive learning has yielded continuous advancements in sampling strategy and architecture design, it still remains two persistent defects: the interference of task-irrelevant information and sample inefficiency, which are related to the recurring existence of trivial constant solutions. From the perspective of dimensional analysis, we find out that the dimensional redundancy and dimensional confounder are the intrinsic issues behind the phenomena, and provide experimental evidence to support our viewpoint. We further propose a simple yet effective approach MetaMask, short for the dimensional Mask learned by Meta-learning, to learn representations against dimensional redundancy and confounder. MetaMask adopts the redundancy-reduction technique to tackle the dimensional redundancy issue and innovatively introduces a dimensional mask to reduce the gradient effects of specific dimensions containing the confounder, which is trained by employing a meta-learning paradigm with the objective of improving the performance of masked representations on a typical self-supervised task. We provide solid theoretical analyses to prove MetaMask can obtain tighter risk bounds for downstream classification compared to typical contrastive methods. Empirically, our method achieves state-of-the-art performance on various benchmarks.
Jiangmeng Li, Wenwen Qiang, Wenyi Mo, Changwen Zheng, Bing Su 0001, Hui Xiong 0001
NeurIPS1
2022 Self-supervised Graph Learning with Segmented Graph Channels
Hang Gao 0004, Jiangmeng Li, Changwen Zheng
ECML/PKDD (2)2
2022 Multi-view representation learning from local consistency and global alignment
Lingyu Si, Wenwen Qiang, Jiangmeng Li, Fanjiang Xu, Funchun Sun
Neurocomputing3
2022 RHMC: Modeling consistent information from deep multiple views via Regularized and Hybrid Multiview Coding
Jiangmeng Li, Wenwen Qiang, Changwen Zheng, Bing Su 0001
Knowl. Based Syst.1
2021 Short Text Clustering with a Deep Multi-embedded Self-supervised Model
Jiangmeng Li, Haichang Li
ICANN (5)3
2021 Auxiliary task guided mean and covariance alignment network for adversarial domain adaptation
Wenwen Qiang, Jiangmeng Li, Changwen Zheng, Bing Su 0001
Knowl. Based Syst.2