VLDB 2026 Research / reviewers in the wild / expert
Xigang Bao
dblp:344/7338
· DBLP profile ↗
7ranked-venue papers
3as first author
7since 2021 · last 2024
0009-0002-3250-2403ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 3 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Pseudo-Label Calibration Semi-supervised Multi-Modal Entity AlignmentabstractMulti-modal entity alignment (MMEA) aims to identify equivalent entities between two multi-modal knowledge graphs for integration. Unfortunately, prior arts have attempted to improve the interaction and fusion of multi-modal information, which have overlooked the influence of modal-specific noise and the usage of labeled and unlabeled data in semi-supervised settings. In this work, we introduce a Pseudo-label Calibration Multi-modal Entity Alignment (PCMEA) in a semi-supervised way. Specifically, in order to generate holistic entity representations, we first devise various embedding modules and attention mechanisms to extract visual, structural, relational, and attribute features. Different from the prior direct fusion methods, we next propose to exploit mutual information maximization to filter the modal-specific noise and to augment modal-invariant commonality. Then, we combine pseudo-label calibration with momentum-based contrastive learning to make full use of the labeled and unlabeled data, which improves the quality of pseudo-label and pulls aligned entities closer. Finally, extensive experiments on two MMEA datasets demonstrate the effectiveness of our PCMEA, which yields state-of-the-art performance. Pengnian Qi, Xigang Bao, Chunlai Zhou, Biao Qin |
AAAI | 3 |
| 2024 | M3TQA: Multi-View, Multi-Hop and Multi-Stage Reasoning for Temporal Question AnsweringabstractKnowledge Graph (KG) have attained notable triumph over Question Answering (QA) tasks. However, the presence of temporal constraints on numerous facts within the real world has sparked heightened interest towards Temporal KGQA (TKGQA). Although previous methods have achieved great progress, they still have the following limitations: 1)PLMs cannot capture the entity drift caused by time constraints in the question. 2) Complex questions require multi-hop reasoning between entities. 3) Fusion strategies (addition or concatenation) of PLMs and KG information ignore feature differences, resulting in suboptimal solutions. To alleviate the above problems, we propose a novel Multi-view, Multi-hop and Multi-stage reasoning paradigm for TKGQA (M3TQA). Specifically, we first design a multi-view calibration module for fusing KG information to calibrate question representation. We next construct graph neural network in a multi-hop modeling module to capture multi-hop message passing between entities. Finally, we design multi-stage aggregation that facilitates the adaptive fusion of heterogeneous information with a two-stage interaction alignment process. The performance on two mainstream benchmark datasets verifies the effectiveness of our proposed model. Zhiyuan Zha, Pengnian Qi, Xigang Bao, Mengyuan Tian, Biao Qin |
ICASSP | 3 |
| 2024 | Contrastive Pre-training with Multi-level Alignment for Grounded Multimodal Named Entity RecognitionabstractRecently, Grounded Multimodal Named Entity Recognition (GM-NER) task has been introduced to refine the Multimodal Named Entity Recognition (MNER) task.Existing MNER studies fall short in that they merely focus on extracting text-based entity-type pairs, often leading to entity ambiguities and failing to contribute to multimodal knowledge graph construction.In the GMNER task, the objective becomes more challenging: identifying named entities in text, determining their entity types, and locating their corresponding bounding boxes in linked images, necessitating precise alignment between the textual and visual information.We introduce a novel multi-level alignment pre-training method, engaging with both text-image and entity-object dimensions to foster deeper congruence between multimodal data.Specifically, we innovatively harness potential objects identified within images, aligning them with textual entity prompts, thereby generating refined soft pseudolabels.These labels serve as self-supervised signals that pre-train the model to more accurately extract entities from textual input.To address misalignments that often plague modality integration, our method employs a sophisticated diffusion model that performs back-translation on the text to generate a corresponding visual representation, thus refining the model's multimodal interpretative accuracy.Empirical evidence from the GMNER dataset validates that our approach significantly outperforms existing state-of-theart models.Moreover, the versatility of our pre-training process complements virtually all extant models, offering an additional avenue for augmenting their multimodal entity recognition acumen. Xigang Bao, Mengyuan Tian, Zhiyuan Zha, Biao Qin |
ICMR | 1 |
| 2024 | A Sample-driven Selection Framework: Towards Graph Contrastive Networks with Reinforcement Learning
Xiangping Zheng 0002, Xiuxin Hao, Bo Wu 0026, Xigang Bao, Xuan Zhang 0009, Wei Li 0109, Xun Liang 0001 |
ACM Multimedia | 4 |
| 2023 | MPMRC-MNER: A Unified MRC framework for Multimodal Named Entity Recognition based Multimodal PromptabstractMultimodal named entity recognition (MNER) is a vision-language task, which aims to detect entity spans and classify them to corresponding entity types given a sentence-image pair. Existing methods often regard an image as a set of visual objects, trying to explicitly capture the relations between visual objects and entities. However, since visual objects are often not identical to entities in quantity and type, they may suffer the bias introduced by visual objects rather than aid. Inspired by the success of textual prompt-based fine-tuning (PF) approaches in many methods, in this paper, we propose a Multimodal Prompt-based Machine Reading Comprehension based framework to implicit alignment between text and image for improving MNER, namely MPMRC-MNER. Specifically, we transform text-only query in MRC into multimodal prompt containing image tokens and text tokens. To better integrate image tokens and text tokens, we design a prompt-aware attention mechanism for better cross-modal fusion. At last, contrastive learning with two types of contrastive losses is designed to learn more consistent representation of two modalities and reduce noise. Extensive experiments and analyses on two public MNER datasets, Twitter2015 and Twitter2017, demonstrate the better performance of our model against the state-of-the-art methods. Xigang Bao, Mengyuan Tian, Zhiyuan Zha, Biao Qin |
CIKM | 1 |
| 2023 | Wukong-CMNER: A Large-Scale Chinese Multimodal NER Dataset with Images Modality
Xigang Bao, Shouhui Wang, Pengnian Qi, Biao Qin |
DASFAA (3) | 1 |
| 2023 | bfE3-MG: End-to-End Expert Linking via Multi-Granularity Representation Learning
Zhiyuan Zha, Pengnian Qi, Xigang Bao, Biao Qin |
ICONIP (13) | 3 |