EDBT 2026 Demo / reviewers in the wild / expert
Wenwen Qiang
dblp:261/6913
· DBLP profile ↗
53ranked-venue papers
8as first author
52since 2021 · last 2026
0000-0002-7985-5743ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 39 · 6 first-author · 39 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 1 first-author · 15 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Group Causal Policy Optimization for Post-Training Large Language ModelsabstractRecent advances in large language models (LLMs) have broadened their applicability across diverse tasks, yet specialized domains still require targeted post-training. Among existing methods, Group Relative Policy Optimization (GRPO) stands out for its efficiency, leveraging groupwise relative rewards while avoiding costly value function learning. However, GRPO treats candidate responses as independent, overlooking semantic interactions such as complementarity and contradiction. To address this challenge, we first introduce a Structural Causal Model (SCM) that reveals hidden dependencies among candidate responses induced by conditioning on a final integrated output, forming a collider structure. Then, our causal analysis leads to two insights: (1) projecting responses onto a causally-informed subspace improves prediction quality, and (2) this projection yields a better baseline than query-only conditioning. Building on these insights, we propose Group Causal Policy Optimization (GCPO), which integrates causal structure into optimization through two key components: a causally-informed reward adjustment and a novel KL-regularization term that aligns the policy with a causally-projected reference distribution. Comprehensive experimental evaluations on various benchmarks demonstrate that GCPO consistently surpasses existing methods. Ziyin Gu, Ran Zuo, Chuxiong Sun, Zeen Song, Changwen Zheng, Wenwen Qiang |
AAAI | 7 |
| 2026 | Exploring Transferability of Self-Supervised Learning by Task Conflict CalibrationabstractIn this paper, we explore the transferability of SSL by addressing two central questions: (i) what is the representation transferability of SSL, and (ii) how can we effectively model this transferability? Transferability is defined as the ability of a representation learned from one task to support the objective of another. Inspired by the meta-learning paradigm, we construct multiple SSL tasks within each training batch to support explicitly modeling transferability. Based on empirical evidence and causal analysis, we find that although introducing task-level information improves transferability, it is still hindered by task conflict. To address this issue, we propose a Task Conflict Calibration method to alleviate the impact of task conflict. Specifically, it first splits batches to create multiple SSL tasks, infusing task-level information. Next, it uses a factor extraction network to produce causal generative factors for all tasks and a weight extraction network to assign dedicated weights to each sample, employing data reconstruction, orthogonality, and sparsity to ensure effectiveness. Finally, the method calibrates sample representations during SSL training and integrates into the pipeline via a two-stage bi-level optimization framework to boost the transferability of learned representations. Experimental results on multiple downstream tasks demonstrate that our method consistently improves the transferability of SSL models. Huijie Guo, Peizheng Guo, Xingchen Shen, Changwen Zheng, Wenwen Qiang |
AAAI | 6 |
| 2026 | Causal Reward Adjustment: Mitigating Reward Hacking in External Reasoning via Backdoor CorrectionabstractExternal reasoning systems combine language models with process reward models (PRMs) to select high-quality reasoning paths for complex tasks such as mathematical problem solving. However, these systems are prone to reward hacking, where high-scoring but logically incorrect paths are assigned high scores by the PRMs, leading to incorrect answers. From a causal inference perspective, we attribute this phenomenon primarily to the presence of confounding semantic features. To address it, we propose Causal Reward Adjustment (CRA), a method that mitigates reward hacking by estimating the true reward of a reasoning path. CRA trains sparse autoencoders on the PRM’s internal activations to recover interpretable features, then corrects confounding by using backdoor adjustment. Experiments on math solving datasets demonstrate that CRA mitigates reward hacking and improves final accuracy, without modifying the policy model or retraining PRM. Ruike Song, Zeen Song, Huijie Guo, Wenwen Qiang |
AAAI | 4 |
| 2026 | TMAE: Learning Targeted Multi-Agent Exploration via Causal InferenceabstractExploration in sparse-reward tasks remains a fundamental challenge in multi-agent reinforcement learning (MARL) due to complex inter-agent interactions and the expansive exploration space. To address this issue, we propose Targeted Multi-Agent Exploration (TMAE), a novel framework that uncovers the causal relationships between the state space and the reward function, thereby reducing the exploration space and enabling more targeted exploration. Specifically, we construct a structural causal model (SCM) to model the causality between sub-state variables and sparse rewards, providing a robust analytical foundation for subsequent causal inference. Through counterfactual causal intervention, TMAE identifies the most critical subspaces for discovering rare but pivotal events while filtering out confounders. By incorporating these causal insights into the exploration process, TMAE prioritizes subspaces with stronger causal effects on sparse rewards, significantly enhancing exploration efficiency. We evaluate TMAE on a range of MARL benchmarks featuring sparse rewards, consistently demonstrating superior exploration efficiency compared to state-of-the-art methods. Furthermore, visualized causal insights derived from TMAE reveal its ability to effectively capture intricate dependencies and priorities in targeted exploration, showcasing strong alignment with prior domain knowledge. Chuxiong Sun, Dunqi Yao, Rui Wang 0079, Wenwen Qiang, Changwen Zheng, Jiangmeng Li |
AAAI | 4 |
| 2026 | Enhancing Large Language Models for Time-Series Forecasting via Vector-Injected In-Context LearningabstractThe World Wide Web needs reliable predictive capabilities to respond to changes in user behavior and usage patterns. Time series forecasting (TSF) is a key means to achieve this goal. In recent years, the large language models (LLMs) for TSF (LLM4TSF) have achieved good performance. However, there is a significant difference between pretraining corpora and time series data, making it hard to guarantee forecasting quality when directly applying LLMs to TSF; fine-tuning LLMs can mitigate this issue, but often incurs substantial computational overhead. Thus, LLM4TSF faces a dual challenge of prediction performance and compute overhead. To address this, we aim to explore a method for improving the forecasting performance of LLM4TSF while freezing all LLM parameters to reduce computational overhead. Inspired by in-context learning (ICL), we propose LVICL. LVICL uses our vector-injected ICL to inject example information into a frozen LLM, eliciting its in-context learning ability and thereby enhancing its performance on the example-related task (i.e., TSF). Specifically, we first use the LLM together with a learnable context vector adapter to extract a context vector from multiple examples adaptively. This vector contains compressed, example-related information. Subsequently, during the forward pass, we inject this vector into every layer of the LLM to improve forecasting performance. Compared with conventional ICL that adds examples into the prompt, our vector-injected ICL does not increase prompt length; moreover, adaptively deriving a context vector from examples suppresses components harmful to forecasting, thereby improving model performance. Extensive experiments demonstrate the effectiveness of our approach. Jianqi Zhang, Wenwen Qiang, Fanjiang Xu, Changwen Zheng |
WWW | 3 |
| 2026 | Self-Supervised Video Representation Learning in a Heuristic Decoupled Perspective
Zeen Song, Wenwen Qiang, Changwen Zheng, Hui Xiong 0001, Gang Hua 0001 |
Int. J. Comput. Vis. | 2 |
| 2026 | On the Transferability and Discriminability of Representation Learning in Unsupervised Domain AdaptationabstractIn this paper, we addressed the limitation of relying solely on distribution alignment and source-domain empirical risk minimization in Unsupervised Domain Adaptation (UDA). Our information-theoretic analysis showed that this standard adversarial-based framework neglects the discriminability of target-domain features, leading to suboptimal performance. To bridge this theoretical-practical gap, we defined "good representation learning" as guaranteeing both transferability and discriminability, and proved that an additional loss term targeting target-domain discriminability is necessary. Building on these insights, we proposed a novel adversarial-based UDA framework that explicitly integrates a domain alignment objective with a discriminability-enhancing constraint. Instantiated as Domain-Invariant Representation Learning with Global and Local Consistency (RLGLC), our method leverages Asymmetrically-Relaxed Wasserstein of Wasserstein Distance (AR-WWD) to address class imbalance and semantic dimension weighting, and employs a local consistency mechanism to preserve fine-grained target-domain discriminative information. Extensive experiments across multiple benchmark datasets demonstrate that RLGLC consistently surpasses state-of-the-art methods, confirming the value of our theoretical perspective and underscoring the necessity of enforcing both transferability and discriminability in adversarial-based UDA. Wenwen Qiang, Ziyin Gu, Lingyu Si, Jiangmeng Li, Changwen Zheng, Fuchun Sun 0001, Hui Xiong 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | Enhancing Human Motion Prediction via Multi-range Decoupling Decoding with Gating-adjusting AggregationabstractExpressive representation of pose sequences is crucial for accurate motion modeling in human motion prediction (HMP). While recent deep learning-based methods have shown promise in learning motion representations, these methods tend to overlook the varying relevance and dependencies between historical information and future moments, with a stronger correlation for short-term predictions and weaker for distant future predictions. This limits the learning of motion representation and then hampers prediction performance. In this paper, we propose a novel approach called multi-range decoupling decoding with gating-adjusting aggregation (MD2GA), which leverages the temporal correlations to refine motion representation learning. This approach employs a two-stage strategy for HMP. In the first stage, a multi-range decoupling decoding adeptly adjusts feature learning by decoding the shared features into distinct future lengths, where different decoders offer diverse insights into motion patterns. In the second stage, a gating-adjusting aggregation dynamically combines the diverse insights guided by input motion data. Extensive experiments demonstrate that the proposed method can be easily integrated into other motion prediction methods and enhance their prediction performance. Jiexin Wang 0003, Wenwen Qiang, Zhao Yang 0006, Bing Su 0001 |
ICME | 2 |
| 2025 | On the Out-of-Distribution Generalization of Self-Supervised LearningabstractIn this paper, we focus on the out-of-distribution (OOD) generalization of self-supervised learning (SSL). By analyzing the mini-batch construction during the SSL training phase, we first give one plausible explanation for SSL having OOD generalization. Then, from the perspective of data generation and causal inference, we analyze and conclude that SSL learns spurious correlations during the training process, which leads to a reduction in OOD generalization. To address this issue, we propose a post-intervention distribution (PID) grounded in the Structural Causal Model. PID offers a scenario where the spurious variable and label variable is mutually independent. Besides, we demonstrate that if each mini-batch during SSL training satisfies PID, the resulting SSL model can achieve optimal worst-case OOD performance. This motivates us to develop a batch sampling strategy that enforces PID constraints through the learning of a latent variable model. Through theoretical analysis, we demonstrate the identifiability of the latent variable model and validate the effectiveness of the proposed sampling strategy. Experiments conducted on various downstream OOD tasks demonstrate the effectiveness of the proposed sampling strategy. Wenwen Qiang, Zeen Song, Jiangmeng Li, Changwen Zheng |
ICML | 1 |
| 2025 | Learning Invariant Causal Mechanism from Vision-Language ModelsabstractContrastive Language-Image Pretraining (CLIP) has achieved remarkable success, but its performance can degrade when fine-tuned in out-of-distribution (OOD) scenarios. We model the prediction process using a Structural Causal Model (SCM) and show that the causal mechanism involving both invariant and variant factors in training environments differs from that in test environments. In contrast, the causal mechanism with solely invariant factors remains consistent across environments. We theoretically prove the existence of a linear mapping from CLIP embeddings to invariant factors, which can be estimated using interventional data. Additionally, we provide a condition to guarantee low OOD risk of the invariant predictor. Based on these insights, we propose the Invariant Causal Mechanism of CLIP (CLIP-ICM) framework. CLIP-ICM involves collecting interventional data, estimating a linear projection matrix, and making predictions within the invariant subspace. Experiments on several OOD datasets show that CLIP-ICM significantly improves the performance of CLIP. Our method offers a simple but powerful enhancement, boosting the reliability of CLIP in real-world applications. Zeen Song, Jiangmeng Li, Changwen Zheng, Wenwen Qiang |
ICML | 6 |
| 2025 | Towards the Causal Complete Cause of Multi-Modal Representation LearningabstractMulti-Modal Learning (MML) aims to learn effective representations across modalities for accurate predictions. Existing methods typically focus on modality consistency and specificity to learn effective representations. However, from a causal perspective, they may lead to representations that contain insufficient and unnecessary information. To address this, we propose that effective MML representations should be causally sufficient and necessary. Considering practical issues like spurious correlations and modality conflicts, we relax the exogeneity and monotonicity assumptions prevalent in prior works and explore the concepts specific to MML, i.e., Causal Complete Cause ($C^3$). We begin by defining $C^3$, which quantifies the probability of representations being causally sufficient and necessary. We then discuss the identifiability of $C^3$ and introduce an instrumental variable to support identifying $C^3$ with non-exogeneity and non-monotonicity. Building on this, we conduct the $C^3$ measurement, i.e., $C^3$ risk. We propose a twin network to estimate it through (i) the real-world branch: utilizing the instrumental variable for sufficiency, and (ii) the hypothetical-world branch: applying gradient-based counterfactual modeling for necessity. Theoretical analyses confirm its reliability. Based on these results, we propose $C^3$ Regularization, a plug-and-play method that enforces the causal completeness of the learned representations by minimizing $C^3$ risk. Extensive experiments demonstrate its effectiveness. Wenwen Qiang, Jiangmeng Li, Changwen Zheng, Fuchun Sun 0001, Hui Xiong 0001 |
ICML | 3 |
| 2025 | Learning to Think: Information-Theoretic Reinforcement Fine-Tuning for LLMsabstractLarge language models (LLMs) excel at complex tasks thanks to advances in their reasoning abilities. However, existing methods overlook the trade-off between reasoning effectiveness and efficiency, often encouraging unnecessarily long reasoning chains and wasting tokens. To address this, we propose Learning to Think (L2T), an information-theoretic reinforcement fine-tuning framework for LLMs to make the models achieve optimal reasoning with fewer tokens. Specifically, L2T treats each query-response interaction as a hierarchical session of multiple episodes and proposes a universal dense process reward, i.e., quantifies the episode-wise information gain in parameters, requiring no extra annotations or task-specific evaluators. We propose a method to quickly estimate this reward based on PAC-Bayes bounds and the Fisher information matrix. Theoretical analyses show that it significantly reduces computational complexity with high estimation accuracy. By immediately rewarding each episode's contribution and penalizing excessive updates, L2T optimizes the model via reinforcement learning to maximize the use of each episode and achieve effective updates. Empirical results on various reasoning benchmarks and base models demonstrate the advantage of L2T across different tasks, boosting both reasoning effectiveness and efficiency. Wenwen Qiang, Zeen Song, Changwen Zheng, Hui Xiong 0001 |
NeurIPS | 2 |
| 2025 | Rethinking Generalizability and Discriminability of Self-Supervised Learning from Evolutionary Game Theory Perspective
Jiangmeng Li, Zehua Zang, Qirui Ji, Chuxiong Sun, Wenwen Qiang, Junge Zhang, Changwen Zheng, Fuchun Sun 0001, Hui Xiong 0001 |
Int. J. Comput. Vis. | 5 |
| 2025 | On the Generalization and Causal Explanation in Self-Supervised Learning
Wenwen Qiang, Zeen Song, Ziyin Gu, Jiangmeng Li, Changwen Zheng, Fuchun Sun 0001, Hui Xiong 0001 |
Int. J. Comput. Vis. | 1 |
| 2025 | On the discriminability of self-supervised representation learning
Zeen Song, Wenwen Qiang, Changwen Zheng, Fuchun Sun 0001, Hui Xiong 0001 |
Inf. Sci. | 2 |
| 2025 | Learning Complementary Knowledge via Trusted Multi-view Space Decomposition for Self-Supervised Contrastive Learning
Jiangmeng Li, Yunze Zhao, Changwen Zheng, Wenwen Qiang |
Mach. Learn. | 5 |
| 2025 | Supporting vision-language model few-shot inference with confounder-pruned knowledge prompt
Jiangmeng Li, Wenyi Mo, Chuxiong Sun, Wenwen Qiang, Bing Su 0001, Changwen Zheng |
Neural Networks | 5 |
| 2025 | Intervening on few-shot object detection based on the front-door criterion
Jiangmeng Li, Qirui Ji, Changwen Zheng, Wenwen Qiang |
Neural Networks | 7 |
| 2024 | Self-Supervised Representation Learning with Meta Comprehensive RegularizationabstractSelf-Supervised Learning (SSL) methods harness the concept of semantic invariance by utilizing data augmentation strategies to produce similar representations for different deformations of the same input. Essentially, the model captures the shared information among multiple augmented views of samples, while disregarding the non-shared information that may be beneficial for downstream tasks. To address this issue, we introduce a module called CompMod with Meta Comprehensive Regularization (MCR), embedded into existing self-supervised frameworks, to make the learned representations more comprehensive. Specifically, we update our proposed model through a bi-level optimization mechanism, enabling it to capture comprehensive features. Additionally, guided by the constrained extraction of features using maximum entropy coding, the self-supervised learning model learns more comprehensive features on top of learning consistent features. In addition, we provide theoretical support for our proposed method from information theory and causal counterfactual perspective. Experimental results show that our method achieves significant improvement in classification, object detection and semantic segmentation tasks on multiple benchmark datasets. Huijie Guo, Ying Ba, Jie Hu 0019, Lingyu Si, Wenwen Qiang, Lei Shi 0002 |
AAAI | 5 |
| 2024 | Hierarchical Topology Isomorphism Expertise Embedded Graph Contrastive LearningabstractGraph contrastive learning (GCL) aims to align the positive features while differentiating the negative features in the latent space by minimizing a pair-wise contrastive loss. As the embodiment of an outstanding discriminative unsupervised graph representation learning approach, GCL achieves impressive successes in various graph benchmarks. However, such an approach falls short of recognizing the topology isomorphism of graphs, resulting in that graphs with relatively homogeneous node features cannot be sufficiently discriminated. By revisiting classic graph topology recognition works, we disclose that the corresponding expertise intuitively complements GCL methods. To this end, we propose a novel hierarchical topology isomorphism expertise embedded graph contrastive learning, which introduces knowledge distillations to empower GCL models to learn the hierarchical topology isomorphism expertise, including the graph-tier and subgraph-tier. On top of this, the proposed method holds the feature of plug-and-play, and we empirically demonstrate that the proposed method is universal to multiple state-of-the-art GCL models. The solid theoretical analyses are further provided to prove that compared with conventional GCL methods, our method acquires the tighter upper bound of Bayes classification error. We conduct extensive experiments on real-world benchmarks to exhibit the performance superiority of our method over candidate GCL methods, e.g., for the real-world graph representation learning experiments, the proposed method beats the state-of-the-art method by 0.23% on unsupervised representation learning setting, 0.43% on transfer learning setting. Our code is available at https://github.com/jyf123/HTML. Jiangmeng Li, Hang Gao 0004, Wenwen Qiang, Changwen Zheng, Fuchun Sun 0001 |
AAAI | 4 |
| 2024 | Demo:SCDRL: Scalable and Customized Distributed Reinforcement Learning SystemabstractReinforcement Learning (RL) has marked significant achievements across a variety of complex tasks in real-world scenarios. However, the efficacy of RL predominantly relies on the availability of extensive datasets and considerable training resources. Hence, there is the critical need for a distributed system capable of generating and processing vast amounts of data with efficiency. In this work, we introduce a Scalable and Customized Distributed Reinforcement Learning system (SCDRL). Concretely, we analyze the paradigm of RL and decouple the major RL computations into three main aspects, i.e. environment simulation, policy inference and policy training. Such decouple enables SCDRL to efficiently allocate computing resources (be it CPUs or GPUs of varying computational capabilities) tailored to the specific needs of each component. We demonstrate the effectiveness of our method across several key RL environments, demonstrating that our system not only achieves significant learning outcomes and enhanced throughput but also utilizes computing resources with greater efficiency. Notably, our findings reveal SCDRL's proficiency in optimizing resource use not just in single-machine setups but also in multi-machine configurations, all the while maintaining data efficiency and resource utilization. Chuxiong Sun, Wenwen Qiang, Jiangmeng Li |
ICDCS | 2 |
| 2024 | BayesPrompt: Prompting Large-Scale Pre-Trained Language Models on Few-shot Inference via Debiased Domain AbstractionabstractAs a novel and effective fine-tuning paradigm based on large-scale pre-trained language models (PLMs), prompt-tuning aims to reduce the gap between downstream tasks and pre-training objectives. While prompt-tuning has yielded continuous advancements in various tasks, such an approach still remains a persistent defect: prompt-tuning methods fail to generalize to specific few-shot patterns. From the perspective of distribution analyses, we disclose that the intrinsic issues behind the phenomenon are the over-multitudinous conceptual knowledge contained in PLMs and the abridged knowledge for target downstream domains, which jointly result in that PLMs mis-locate the knowledge distributions corresponding to the target domains in the universal knowledge embedding space. To this end, we intuitively explore to approximate the unabridged target domains of downstream tasks in a debiased manner, and then abstract such domains to generate discriminative prompts, thereby providing the de-ambiguous guidance for PLMs. Guided by such an intuition, we propose a simple yet effective approach, namely BayesPrompt, to learn prompts that contain the domain discriminative information against the interference from domain-irrelevant knowledge. BayesPrompt primitively leverages known distributions to approximate the debiased factual distributions of target domains and further uniformly samples certain representative features from the approximated distributions to generate the ultimate prompts for PLMs. We provide theoretical insights with the connection to domain adaptation. Empirically, our method achieves state-of-the-art performance on benchmarks. Jiangmeng Li, Wenwen Qiang, Changwen Zheng, Fuchun Sun 0001, Hui Xiong 0001 |
ICLR | 4 |
| 2024 | Unbiased Image Synthesis via Manifold Guidance in Diffusion ModelsabstractDiffusion Models are a potent class of generative models capable of producing high-quality images. However, they often inadvertently favor certain data attributes, undermining the diversity of generated images. This issue is starkly apparent in skewed datasets like CelebA, where the initial dataset disproportionately favors females over males by 57.9%, this bias amplified in generated data where female representation outstrips males by 148%. In response, we propose a plug-and-play method named Manifold Guidance Sampling, which is also the first unsupervised method to mitigate bias issue in DDPMs. Leveraging the inherent structure of the data manifold, this method steers the sampling process towards a more uniform distribution, effectively dispersing the clustering of biased data. Without the need for modifying the existing model or additional training, it significantly mitigates data bias and enhances the quality and unbiasedness of the generated images. Xingzhe Su, Daixi Jia, Fengge Wu, Junsuo Zhao, Changwen Zheng, Wenwen Qiang |
ICME | 6 |
| 2024 | Hacking Task Confounder in Meta-Learning
Zeen Song, Jianqi Zhang, Changwen Zheng, Wenwen Qiang |
IJCAI | 6 |
| 2024 | Is Encoded Popularity Always Harmful? Explicit Debiasing with Augmentation for Contrastive Collaborative FilteringabstractCollaborative Filtering (CF) models based on Graph Contrastive Learning (GCL) have effectively improved the performance of long-tail recommendation. However, the popularity bias still presents a challenge in further enhancing their effectiveness. Some studies suggest that achieving better recommendations, particularly for the long-tail, requires learning representations with a more uniform distribution to implicitly mitigate popularity bias. Nevertheless, our analysis of various CF models reveals that different models exhibit varying abilities in capturing popularity and those with superior performance might encode more popularity information in item representations. This raises a question: Does encoding popularity always lead to harmful bias? We speculate that superior recommendations may emerge from leveraging the encoded popularity information to optimize the representations of users and items rather than eliminating its existence in representations. This motivates a data augmentation approach, wherein we generate augmented samples by mixing representations of items with different popularity levels and explicitly debias using the encoded popularity information which is often neglected. Additionally, we propose an adaptive contrastive loss, leveraging structural information and unifying the recommendation and contrastive learning objectives, which adaptively re-weights positive samples and ensures the capture of item popularity. Our proposed framework remains scalable without requiring multiple forward computations throughout the entire graph. Extensive experiments demonstrate improvements in both overall and long-tail recommendation performance. Guanming Chen, Wenwen Qiang, Yuanxin Ouyang, Chuantao Yin, Zhang Xiong 0001 |
IJCNN | 2 |
| 2024 | Not All Frequencies Are Created Equal: Towards a Dynamic Fusion of Frequencies in Time-Series ForecastingabstractLong-term time series forecasting is a long-standing challenge in various applications. A central issue in time series forecasting is that methods should expressively capture long-term dependency. Furthermore, time series forecasting methods should be flexible when applied to different scenarios. Although Fourier analysis offers an alternative to effectively capture reusable and periodic patterns to achieve long-term forecasting in different scenarios, existing methods often assume high-frequency components represent noise and should be discarded in time series forecasting. However, we conduct a series of motivation experiments and discover that the role of certain frequencies varies depending on the scenarios. In some scenarios, removing high-frequency components from the original time series can improve the forecasting performance, while in others scenarios, removing them is harmful to forecasting performance. Therefore, it is necessary to treat the frequencies differently according to specific scenarios. To achieve this, we first reformulate the time series forecasting problem as learning a transfer function of each frequency in the Fourier domain. Further, we design Frequency Dynamic Fusion (FreDF), which individually predicts each Fourier component, and dynamically fuses the output of different frequencies. Moreover, we provide a novel insight into the generalization ability of time series forecasting and propose the generalization bound of time series forecasting. Then we prove FreDF has a lower bound, indicating that FreDF has better generalization ability. Extensive experiments conducted on multiple benchmark datasets and ablation studies demonstrate the effectiveness of FreDF. Zeen Song, Huijie Guo, Jianqi Zhang, Changwen Zheng, Wenwen Qiang |
ACM Multimedia | 7 |
| 2024 | Rethinking Misalignment in Vision-Language Model Adaptation from a Causal PerspectiveabstractFoundational Vision-Language models such as CLIP have exhibited impressive generalization in downstream tasks. However, CLIP suffers from a two-level misalignment issue, i.e., task misalignment and data misalignment, when adapting to specific tasks. Soft prompt tuning has mitigated the task misalignment, yet the data misalignment remains a challenge. To analyze the impacts of the data misalignment, we revisit the pre-training and adaptation processes of CLIP and develop a structural causal model. We discover that while we expect to capture task-relevant information for downstream tasks accurately, the task-irrelevant knowledge impacts the prediction results and hampers the modeling of the true relationships between the images and the predicted classes. As task-irrelevant knowledge is unobservable, we leverage the front-door adjustment and propose Causality-Guided Semantic Decoupling and Classification (CDC) to mitigate the interference of task-irrelevant knowledge. Specifically, we decouple semantics contained in the data of downstream tasks and perform classification based on each semantic. Furthermore, we employ the Dempster-Shafer evidence theory to evaluate the uncertainty of each prediction generated by diverse semantics. Experiments conducted in multiple different settings have consistently demonstrated the effectiveness of CDC. Jiangmeng Li, Wenwen Qiang |
NeurIPS | 4 |
| 2024 | Intriguing Property and Counterfactual Explanation of GAN for Remote Sensing Image Generation
Xingzhe Su, Wenwen Qiang, Jie Hu 0019, Changwen Zheng, Fengge Wu, Fuchun Sun 0001 |
Int. J. Comput. Vis. | 2 |
| 2024 | Towards Task Sampler Learning for Meta-Learning
Wenwen Qiang, Xingzhe Su, Changwen Zheng, Fuchun Sun 0001, Hui Xiong 0001 |
Int. J. Comput. Vis. | 2 |
| 2024 | Regularized Hypothesis-Induced Wasserstein Divergence for unsupervised domain adaptation
Lingyu Si, Wenwen Qiang, Changwen Zheng, Junzhi Yu 0001, Fuchun Sun 0001 |
Knowl. Based Syst. | 3 |
| 2024 | A Novel Causal Inference-Guided Feature Enhancement Framework for PolSAR Image ClassificationabstractIn recent years, there has been a prominent focus on enhancing the quality of features derived from convolutional neural networks (CNNs) within the field of polarimetric synthetic aperture radar (PolSAR) image classification. Targeting this challenge, this article first visualizes the lack of discriminability and generalizability in CNN features through several empirical observations. Subsequently, we explain why these problems arise from a causal perspective, accomplished by means of a structural causal model (SCM) constructed according to the training and testing process of CNNs. This SCM facilitates the identification of variables that affect the quality of PolSAR image feature learning, as well as an intervention on those variables using backdoor adjustment. Building upon this groundwork, a novel causal inference-guided feature enhancement framework is constructed. It can be seamlessly integrated into any CNN-based PolSAR image classifier in a plug-and-play manner, enabling the enhanced classifier to filter out interference information and prevent model overfitting. These two aspects bring better feature discriminability and generalizability, respectively, leading to improved classification performance. Experimental results on four widely-used PolSAR image datasets demonstrate the effectiveness of our proposed framework. We integrate it into several mainstream methods in the field and show that the accuracy of the enhanced classifier is improved compared to the original model. Lingyu Si, Wenwen Qiang, Lamei Zhang, Junzhi Yu 0001, Yuquan Wu, Changwen Zheng, Fuchun Sun 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | A Trusted Generative-Discriminative Joint Feature Learning Framework for Remote Sensing Image ClassificationabstractRemote sensing image (RSI) classification is a popular research topic that aims to assign semantic labels to images acquired from aerial or maritime platforms. Existing deep feature learning methods for this task can be divided into two paradigms: generative and discriminative. The former methods are good at capturing every local detail of images, while the later approaches focus on the most salient area. The significant differences between the two types of methods, both in terms of their underlying mechanisms and practical implementation, motivate us to integrate information acquired by both paradigms by exploiting their complementary strengths. However, this idea faces a challenge that local information in the extracted features, especially those from generative methods, may not be reliable for RSI classification. The reason for this challenge is that, due to the characteristics of the ground observation perspective, some RSIs, while semantically different, exhibit a significant degree of similarity in local details. This phenomenon leads to insufficient discriminability of local features to separate multiple RSI categories, which implies that the classification results overly focused on local information may be unreliable. To address this issue, in this article, we propose a novel framework that integrates generative and discriminative feature learning methods with evidential learning for RSI classification. Our framework uses the Dirichlet distribution to model the predicted probabilities to be integrated, thereby collecting evidence about their reliability. This enables us to integrate multiple features at an evidence level and make reliable decisions, overcoming the unreliabilities of generative-discriminative joint feature learning induced by RSI characteristics. We evaluate the proposed framework on several satellite and shipborne RSI classification datasets. The experimental results show that our method outperforms the state-of-the-art baselines in terms of accuracy and robustness. Lingyu Si, Wenwen Qiang, Zeen Song, Bo Du 0001, Junzhi Yu 0001, Fuchun Sun 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Manifold Constraint Regularization for Remote Sensing Image GenerationabstractGenerative adversarial networks (GANs) have shown notable accomplishments in remote sensing (RS) domain. However, this article reveals that their performance on RS images falls short when compared to their impressive results with natural images. This study identifies a previously overlooked issue: GANs exhibit a heightened susceptibility to overfitting on RS images. To address this challenge, this article analyzes the characteristics of RS images and proposes manifold constraint regularization (MCR), a novel approach that tackles overfitting of GANs on RS images for the first time. Our method includes a new measure for evaluating the structure of the data manifold. Leveraging this measure, we propose the MCR term, which not only alleviates the overfitting problem, but also promotes alignment between the generated and real data manifolds, leading to enhanced quality in the generated images. The effectiveness and versatility of this method have been corroborated through extensive validation on various RS datasets and GAN models. The proposed method not only enhances the quality of the generated images, reflected in a 3.13% improvement in Fréchet inception distance (FID) score, but also boosts the performance of the GANs on downstream tasks, evidenced by a 3.76% increase in classification accuracy. The source code is available athttps://github.com/rootSue/Manifold-RSGAN. Xingzhe Su, Changwen Zheng, Wenwen Qiang, Fengge Wu, Junsuo Zhao, Fuchun Sun 0001, Hui Xiong 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Robust Causal Graph Representation Learning against Confounding EffectsabstractThe prevailing graph neural network models have achieved significant progress in graph representation learning. However, in this paper, we uncover an ever-overlooked phenomenon: the pre-trained graph representation learning model tested with full graphs underperforms the model tested with well-pruned graphs. This observation reveals that there exist confounders in graphs, which may interfere with the model learning semantic information, and current graph representation learning methods have not eliminated their influence. To tackle this issue, we propose Robust Causal Graph Representation Learning (RCGRL) to learn robust graph representations against confounding effects. RCGRL introduces an active approach to generate instrumental variables under unconditional moment restrictions, which empowers the graph representation learning model to eliminate confounders, thereby capturing discriminative information that is causally related to downstream predictions. We offer theorems and proofs to guarantee the theoretical effectiveness of the proposed approach. Empirically, we conduct extensive experiments on a synthetic dataset and multiple benchmark datasets. Experimental results demonstrate the effectiveness and generalization ability of RCGRL. Our codes are available at https://github.com/hang53/RCGRL. Hang Gao 0004, Jiangmeng Li, Wenwen Qiang, Lingyu Si, Changwen Zheng, Fuchun Sun 0001 |
AAAI | 3 |
| 2023 | Disentangle and Remerge: Interventional Knowledge Distillation for Few-Shot Object Detection from a Conditional Causal PerspectiveabstractFew-shot learning models learn representations with limited human annotations, and such a learning paradigm demonstrates practicability in various tasks, e.g., image classification, object detection, etc. However, few-shot object detection methods suffer from an intrinsic defect that the limited training data makes the model cannot sufficiently explore semantic information. To tackle this, we introduce knowledge distillation to the few-shot object detection learning paradigm. We further run a motivating experiment, which demonstrates that in the process of knowledge distillation, the empirical error of the teacher model degenerates the prediction performance of the few-shot object detection model as the student. To understand the reasons behind this phenomenon, we revisit the learning paradigm of knowledge distillation on the few-shot object detection task from the causal theoretic standpoint, and accordingly, develop a Structural Causal Model. Following the theoretical guidance, we propose a backdoor adjustment-based knowledge distillation method for the few-shot object detection task, namely Disentangle and Remerge (D&R), to perform conditional causal intervention toward the corresponding Structural Causal Model. Empirically, the experiments on benchmarks demonstrate that D&R can yield significant performance boosts in few-shot object detection. Code is available at https://github.com/ZYN-1101/DandR.git. Jiangmeng Li, Wenwen Qiang, Lingyu Si, Chengbo Jiao, Changwen Zheng, Fuchun Sun 0001 |
AAAI | 3 |
| 2023 | Rethinking skip connection model as a learnable Markov chain
Dengsheng Chen, Jie Hu 0019, Wenwen Qiang, Xiaoming Wei, Enhua Wu |
ICLR | 3 |
| 2023 | Spatio-Temporal Branching for Motion Prediction using Motion IncrementsabstractHuman motion prediction (HMP) has emerged as a popular research topic due to its diverse applications. Traditional methods rely on hand-crafted features and machine learning techniques, which often struggle to model the complex dynamics of human motion. Recent deep learning-based methods have achieved success by learning spatio-temporal representations of motion, but these models often overlook the reliability of motion data. Additionally, the temporal and spatial dependencies of skeleton nodes are distinct. The temporal relationship captures motion information over time, while the spatial relationship describes body structure and the relationships between different nodes. In this paper, we propose a novel spatio-temporal branching network using incremental information for HMP, which decouples the learning of temporal-domain and spatial-domain features, extracts more motion information, and achieves complementary cross-domain knowledge learning through knowledge distillation. Our approach effectively reduces noise interference and provides more expressive information for characterizing motion by separately extracting temporal and spatial features. We evaluate our approach on standard HMP benchmarks and outperform state-of-the-art methods in terms of prediction accuracy. Code is available at https://github.com/JasonWang959/STPMP. Jiexin Wang 0003, Wenwen Qiang, Ying Ba, Bing Su 0001, Ji-Rong Wen |
ACM Multimedia | 3 |
| 2023 | Zero-shot Skeleton-based Action Recognition via Mutual Information Estimation and MaximizationabstractZero-shot skeleton-based action recognition aims to recognize actions of unseen categories after training on data of seen categories. The key is to build the connection between visual and semantic space from seen to unseen classes. Previous studies have primarily focused on encoding sequences into a singular feature vector, with subsequent mapping the features to an identical anchor point within the embedded space. Their performance is hindered by 1) the ignorance of the global visual/semantic distribution alignment, which results in a limitation to capture the true interdependence between the two spaces. 2) the negligence of temporal information since the frame-wise features with rich action clues are directly pooled into a single feature vector. We propose a new zero-shot skeleton-based action recognition method via mutual information (MI) estimation and maximization. Specifically, 1) we maximize the MI between visual and semantic space for distribution alignment; 2) we leverage the temporal information for estimating the MI by encouraging MI to increase as more frames are observed. Extensive experiments on three large-scale skeleton action datasets confirm the effectiveness of our method. Wenwen Qiang, Anyi Rao, Ning Lin, Bing Su 0001, Jiaqi Wang 0003 |
ACM Multimedia | 2 |
| 2023 | Meta Attention-Generation Network for Cross-Granularity Few-Shot Learning
Wenwen Qiang, Jiangmeng Li, Bing Su 0001, Jianlong Fu, Hui Xiong 0001, Ji-Rong Wen |
Int. J. Comput. Vis. | 1 |
| 2023 | Information theory-guided heuristic progressive multi-view coding
Jiangmeng Li, Hang Gao 0004, Wenwen Qiang, Changwen Zheng |
Neural Networks | 3 |
| 2023 | Modeling Multiple Views via Implicitly Preserving Global Consistency and Local ComplementarityabstractWhile self-supervised learning techniques are often used to mine hidden knowledge from unlabeled data via modeling multiple views, it is unclear how to perform effective representation learning in a complex and inconsistent context. To this end, we propose a new multi-view self-supervised learning method, namelyconsistency and complementarity network(CoCoNet), to comprehensively learn global inter-view consistent and local cross-view complementarity-preserving representations from multiple views. To capture crucial common knowledge which is implicitly shared among views, CoCoNet employs a global consistency module that aligns the probabilistic distribution of views by utilizing an efficient discrepancy metric based on the generalized sliced Wasserstein distance. To incorporate cross-view complementary information, CoCoNet proposes a heuristic complementarity-aware contrastive learning approach, which extracts a complementarity-factor jointing cross-view discriminative knowledge and uses it as the contrast to guide the learning of view-specific encoders. Theoretically, the superiority of CoCoNet is verified by our information-theoretical-based analyses. Empirically, our thorough experimental results show that CoCoNet outperforms the state-of-the-art self-supervised methods by a significant margin, for instance, CoCoNet beats the best benchmark method by an average margin of 1.1% on ImageNet. Jiangmeng Li, Wenwen Qiang, Changwen Zheng, Bing Su 0001, Farid Razzak, Ji-Rong Wen, Hui Xiong 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2023 | Robust Local Preserving and Global Aligning Network for Adversarial Domain AdaptationabstractUnsupervised domain adaptation (UDA) requires source domain samples with clean ground truth labels during training. Accurately labeling a large number of source domain samples is time-consuming and laborious. An alternative is to utilize samples with noisy labels for training. However, training with noisy labels can greatly reduce the performance of UDA. In this paper, we address the problem that learning UDA models only with access to noisy labels and propose a novel method called robust local preserving and global aligning network (RLPGA). RLPGA improves the robustness of the label noise from two aspects. One is learning a classifier by a robust informative-theoretic-based loss function. The other is constructing two adjacency weight matrices and two negative weight matrices by the proposed local preserving module to preserve the local topology structures of input data. We conduct theoretical analysis on the robustness of the proposed RLPGA and prove that the robust informative-theoretic-based loss and the local preserving module are beneficial to reduce the empirical risk of the target domain. A series of empirical studies show the effectiveness of our proposed RLPGA. Wenwen Qiang, Jiangmeng Li, Changwen Zheng, Bing Su 0001, Hui Xiong 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2022 | MetAug: Contrastive Learning via Meta Feature AugmentationabstractWhat matters for contrastive learning? We argue that contrastive learning heavily relies on informative features, or “hard” (positive or negative) features. Early works include more informative features by applying complex data augmentations and large batch size or memory bank, and recent works design elaborate sampling approaches to explore informative features. The key challenge toward exploring such features is that the source multi-view data is generated by applying random data augmentations, making it infeasible to always add useful information in the augmented data. Consequently, the informativeness of features learned from such augmented data is limited. In response, we propose to directly augment the features in latent space, thereby learning discriminative representations without a large amount of input data. We perform a meta learning technique to build the augmentation generator that updates its network parameters by considering the performance of the encoder. However, insufficient input data may lead the encoder to learn collapsed features and therefore malfunction the augmentation generator. A new margin-injected regularization is further added in the objective function to avoid the encoder learning a degenerate mapping. To contrast all features in one gradient back-propagation step, we adopt the proposed optimization-driven unified contrastive loss instead of the conventional contrastive loss. Empirically, our method achieves state-of-the-art results on several benchmark datasets. Jiangmeng Li, Wenwen Qiang, Changwen Zheng, Bing Su 0001, Hui Xiong 0001 |
ICML | 2 |
| 2022 | Interventional Contrastive Learning with Meta Semantic RegularizerabstractContrastive learning (CL)-based self-supervised learning models learn visual representations in a pairwise manner. Although the prevailing CL model has achieved great progress, in this paper, we uncover an ever-overlooked phenomenon: When the CL model is trained with full images, the performance tested in full images is better than that in foreground areas; when the CL model is trained with foreground areas, the performance tested in full images is worse than that in foreground areas. This observation reveals that backgrounds in images may interfere with the model learning semantic information and their influence has not been fully eliminated. To tackle this issue, we build a Structural Causal Model (SCM) to model the background as a confounder. We propose a backdoor adjustment-based regularization method, namely Interventional Contrastive Learning with Meta Semantic Regularizer (ICL-MSR), to perform causal intervention towards the proposed SCM. ICL-MSR can be incorporated into any existing CL methods to alleviate background distractions from representation learning. Theoretically, we prove that ICL-MSR achieves a tighter error bound. Empirically, our experiments on multiple benchmark datasets demonstrate that ICL-MSR is able to improve the performances of different state-of-the-art CL methods. Wenwen Qiang, Jiangmeng Li, Changwen Zheng, Bing Su 0001, Hui Xiong 0001 |
ICML | 1 |
| 2022 | Bootstrapping Informative Graph Augmentation via A Meta Learning ApproachabstractRecent works explore learning graph representations in a self-supervised manner. In graph contrastive learning, benchmark methods apply various graph augmentation approaches. However, most of the augmentation methods are non-learnable, which causes the issue of generating unbeneficial augmented graphs. Such augmentation may degenerate the representation ability of graph contrastive learning methods. Therefore, we motivate our method to generate augmented graph with a learnable graph augmenter, called MEta Graph Augmentation (MEGA). We then clarify that a "good" graph augmentation must have uniformity at the instance-level and informativeness at the feature-level. To this end, we propose a novel approach to learning a graph augmenter that can generate an augmentation with uniformity and informativeness. The objective of the graph augmenter is to promote our feature extraction network to learn a more discriminative feature representation, which motivates us to propose a meta-learning paradigm. Empirically, the experiments across multiple benchmark datasets demonstrate that MEGA outperforms the state-of-the-art methods in graph self-supervised learning tasks. Further experimental studies prove the effectiveness of different terms of MEGA. Our codes are available at https://github.com/hang53/MEGA. Hang Gao 0004, Jiangmeng Li, Wenwen Qiang, Lingyu Si, Fuchun Sun 0001, Changwen Zheng |
IJCAI | 3 |
| 2022 | MetaMask: Revisiting Dimensional Confounder for Self-Supervised LearningabstractAs a successful approach to self-supervised learning, contrastive learning aims to learn invariant information shared among distortions of the input sample. While contrastive learning has yielded continuous advancements in sampling strategy and architecture design, it still remains two persistent defects: the interference of task-irrelevant information and sample inefficiency, which are related to the recurring existence of trivial constant solutions. From the perspective of dimensional analysis, we find out that the dimensional redundancy and dimensional confounder are the intrinsic issues behind the phenomena, and provide experimental evidence to support our viewpoint. We further propose a simple yet effective approach MetaMask, short for the dimensional Mask learned by Meta-learning, to learn representations against dimensional redundancy and confounder. MetaMask adopts the redundancy-reduction technique to tackle the dimensional redundancy issue and innovatively introduces a dimensional mask to reduce the gradient effects of specific dimensions containing the confounder, which is trained by employing a meta-learning paradigm with the objective of improving the performance of masked representations on a typical self-supervised task. We provide solid theoretical analyses to prove MetaMask can obtain tighter risk bounds for downstream classification compared to typical contrastive methods. Empirically, our method achieves state-of-the-art performance on various benchmarks. Jiangmeng Li, Wenwen Qiang, Wenyi Mo, Changwen Zheng, Bing Su 0001, Hui Xiong 0001 |
NeurIPS | 2 |
| 2022 | Multi-view representation learning from local consistency and global alignment
Lingyu Si, Wenwen Qiang, Jiangmeng Li, Fanjiang Xu, Funchun Sun |
Neurocomputing | 2 |
| 2022 | RHMC: Modeling consistent information from deep multiple views via Regularized and Hybrid Multiview Coding
Jiangmeng Li, Wenwen Qiang, Changwen Zheng, Bing Su 0001 |
Knowl. Based Syst. | 2 |
| 2022 | Unified feature extraction framework based on contrastive learning
Wenwen Qiang, Yingyi Chen, Ling Jing |
Knowl. Based Syst. | 2 |
| 2022 | Feature extraction framework based on contrastive learning with adaptive positive and negative samples
Wenwen Qiang, Yingyi Chen, Ling Jing |
Neural Networks | 3 |
| 2021 | Locality cross-view regression for feature extraction
Wenwen Qiang, Naiyang Deng, Ling Jing |
Eng. Appl. Artif. Intell. | 3 |
| 2021 | Auxiliary task guided mean and covariance alignment network for adversarial domain adaptation
Wenwen Qiang, Jiangmeng Li, Changwen Zheng, Bing Su 0001 |
Knowl. Based Syst. | 1 |
| 2020 | Robust weighted linear loss twin multi-class support vector regression for large-scale classification
Wenwen Qiang, Ling Zhen, Ling Jing |
Signal Process. | 1 |