VLDB 2026 Research / reviewers in the wild / expert
Wei Li 0233
dblp:64/6025-233
· DBLP profile ↗
15ranked-venue papers
10as first author
15since 2021 · last 2026
0000-0003-2548-3486ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 7 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards explainable visual question answering via cross-modal causal reasoningabstractExplainable Visual Question Answering (EVQA) aims to not only predict accurate answers to visual questions but also generate human-friendly multimodal explanations that reveal the underlying reasoning process. Despite significant progress, existing EVQA methods suffer from two critical limitations: (1) they often rely on spurious cross-modal correlations (e.g., linguistic biases or visual shortcuts) rather than genuine causal relations, leading to unreliable reasoning; (2) the consistency between predicted answers and generated explanations is compromised due to the lack of explicit modeling of their causal dependencies. To address these issues, we propose a Cross-Modal Causal Reasoning (CMCR) framework that integrates causal inference with multimodal learning to disentangle causal effects from spurious correlations and enforce answer-explanation consistency. Specifically, CMCR incorporates three key innovations: (1) Causal Intervention, which employs backdoor adjustment to eliminate linguistic biases and frontdoor adjustment to mitigate visual shortcut biases; (2) a Neural-Symbolic Explanation Generator designed to translate symbolic reasoning processes into natural language explanations, thereby enhancing process explainability; and (3) Variational Causal Inference, which enforces causal consistency between answers and explanations. Experiments on benchmark datasets demonstrate that CMCR outperforms state-of-the-art methods, achieving a 1.19% higher accuracy, a 1.05% higher grounding for explanation quality, and a 0.42% higher answer-explanation consistency. Wei Li 0233, Fuyun Deng, Zhixin Li 0001 |
Eng. Appl. Artif. Intell. | 1 |
| 2026 | Context-guided deep interaction networks for image captioning
Fuyun Deng, Zhixin Li 0001, Wei Li 0233, Jie Yang 0034 |
Expert Syst. Appl. | 3 |
| 2026 | CMGR: Cross-modal graph routing for knowledge-based visual question answering
Wei Li 0233, Zhixin Li 0001 |
Neurocomputing | 1 |
| 2026 | Question-guided attention and cross-modal alignment for knowledge-based visual question answering
Wei Li 0233, Fuyun Deng, Zhixin Li 0001 |
Inf. Process. Manag. | 1 |
| 2026 | Bidirectional causal learning for visual question answering
Faning Long, Peiyi Wei, Peiyun Li, Wei Li 0233 |
Image Vis. Comput. | 5 |
| 2026 | Counterfactual causal inference for robust visual question answering
Wei Li 0233, Zhixin Li 0001, Fuyun Deng, Canlong Zhang |
Neural Networks | 1 |
| 2025 | QGCMA: A Framework for Knowledge-Based Visual Question Answering
Wei Li 0233, Zhixin Li 0001 |
CIKM | 1 |
| 2025 | Towards Robust Visual Question Answering via Causal Intervention and Contrastive LearningabstractVisual question answering systems are designed to integrate visual and linguistic information to provide accurate answers to questions. However, current VQA models are susceptible to shortcut learning, where they exploit spurious correlations rather than genuine multimodal interactions. This shortcut bias leads to an over-reliance on a single modality, resulting in poor generalization performance and vulnerability to distributional shifts between training and testing sets. While existing solutions primarily address shortcut learning within the linguistic modality, they often overlook other types of shortcut biases. This paper introduces a novel approach based on causal intervention to mitigate various shortcut biases in VQA. By explicitly addressing these biases, our method achieves state-of-the-art performance on the VQA-CP v2 dataset, demonstrating its effectiveness and superiority, and offering a significant advancement in the improvement of VQA systems. Wei Li 0233, Zhixin Li 0001 |
ICME | 1 |
| 2025 | Key region Semantic information Augmented Transformer for Image Captioning
Fuyun Deng, Wei Li 0233, Zhixin Li 0001 |
Eng. Appl. Artif. Intell. | 2 |
| 2023 | Causal-ViT: Robust Vision Transformer by causal intervention
Wei Li 0233, Zhixin Li 0001, Xiwei Yang, Huifang Ma |
Eng. Appl. Artif. Intell. | 1 |
| 2023 | MLCB-Net: a multi-level class balancing network for domain adaptive semantic segmentation
Wei Li 0233, Xiwei Yang, Zhixin Li 0001 |
Multim. Syst. | 1 |
| 2022 | Causal-SETR: A SEgmentation TRansformer Variant Based on Causal Intervention
Wei Li 0233, Zhixin Li 0001 |
ACCV (7) | 1 |
| 2022 | Domain Adaptative Semantic Segmentation by Fine-Grained Alignment
Zhixin Li 0001, Wei Li 0233 |
ICANN (4) | 2 |
| 2022 | Distinguishing foreground and background alignment for unsupervised domain adaptative semantic segmentation
Wei Li 0233, Zhixin Li 0001 |
Image Vis. Comput. | 2 |
| 2021 | Domain Adaptative Semantic Segmentation by alleviating Long-tail ProblemabstractThe domain adaptive method based on the adversarial network can be effectively applied to unsupervised semantic segmentation tasks. State-of-the-art approaches have proved that domain alignment at the semantic level can improve segmentation networks' performance. Based on data observation between different domains, we find that the long-tail problem exists in these datasets. We propose a two-level class balancing model to alleviate the semantic class imbalance of data to address this problem. Specifically, we count the category frequencies in the datasets and treat this frequency information as mutual information. Then, we feed this mutual information to the cross-entropy method for fitting so that the model can alleviate the long-tail problem globally. Besides, we resample the data in the model's training process by using two classifiers to balance the head class and the tail class locally. Finally, we use self-supervised learning to supervise the target domain's alignment and the source domain, thus achieving further improvement. We conduct experiments on the benchmark of mainstream unsupervised domain adaptive semantic segmentation tasks, and the experimental results show that our proposed method is effective. Wei Li 0233, Zhixin Li 0001 |
IJCNN | 1 |