Xigang Bao

dblp:344/7338 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
7since 2021 · last 2024
0009-0002-3250-2403ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 3 first-author · 3 since 2021
YearPublicationVenuePosition
2024 Pseudo-Label Calibration Semi-supervised Multi-Modal Entity Alignment
abstract
Multi-modal entity alignment (MMEA) aims to identify equivalent entities between two multi-modal knowledge graphs for integration. Unfortunately, prior arts have attempted to improve the interaction and fusion of multi-modal information, which have overlooked the influence of modal-specific noise and the usage of labeled and unlabeled data in semi-supervised settings. In this work, we introduce a Pseudo-label Calibration Multi-modal Entity Alignment (PCMEA) in a semi-supervised way. Specifically, in order to generate holistic entity representations, we first devise various embedding modules and attention mechanisms to extract visual, structural, relational, and attribute features. Different from the prior direct fusion methods, we next propose to exploit mutual information maximization to filter the modal-specific noise and to augment modal-invariant commonality. Then, we combine pseudo-label calibration with momentum-based contrastive learning to make full use of the labeled and unlabeled data, which improves the quality of pseudo-label and pulls aligned entities closer. Finally, extensive experiments on two MMEA datasets demonstrate the effectiveness of our PCMEA, which yields state-of-the-art performance.
Pengnian Qi, Xigang Bao, Chunlai Zhou, Biao Qin
AAAI3
2024 M3TQA: Multi-View, Multi-Hop and Multi-Stage Reasoning for Temporal Question Answering
abstract
Knowledge Graph (KG) have attained notable triumph over Question Answering (QA) tasks. However, the presence of temporal constraints on numerous facts within the real world has sparked heightened interest towards Temporal KGQA (TKGQA). Although previous methods have achieved great progress, they still have the following limitations: 1)PLMs cannot capture the entity drift caused by time constraints in the question. 2) Complex questions require multi-hop reasoning between entities. 3) Fusion strategies (addition or concatenation) of PLMs and KG information ignore feature differences, resulting in suboptimal solutions. To alleviate the above problems, we propose a novel Multi-view, Multi-hop and Multi-stage reasoning paradigm for TKGQA (M3TQA). Specifically, we first design a multi-view calibration module for fusing KG information to calibrate question representation. We next construct graph neural network in a multi-hop modeling module to capture multi-hop message passing between entities. Finally, we design multi-stage aggregation that facilitates the adaptive fusion of heterogeneous information with a two-stage interaction alignment process. The performance on two mainstream benchmark datasets verifies the effectiveness of our proposed model.
Zhiyuan Zha, Pengnian Qi, Xigang Bao, Mengyuan Tian, Biao Qin
ICASSP3
2024 Contrastive Pre-training with Multi-level Alignment for Grounded Multimodal Named Entity Recognition
abstract
Recently, Grounded Multimodal Named Entity Recognition (GM-NER) task has been introduced to refine the Multimodal Named Entity Recognition (MNER) task.Existing MNER studies fall short in that they merely focus on extracting text-based entity-type pairs, often leading to entity ambiguities and failing to contribute to multimodal knowledge graph construction.In the GMNER task, the objective becomes more challenging: identifying named entities in text, determining their entity types, and locating their corresponding bounding boxes in linked images, necessitating precise alignment between the textual and visual information.We introduce a novel multi-level alignment pre-training method, engaging with both text-image and entity-object dimensions to foster deeper congruence between multimodal data.Specifically, we innovatively harness potential objects identified within images, aligning them with textual entity prompts, thereby generating refined soft pseudolabels.These labels serve as self-supervised signals that pre-train the model to more accurately extract entities from textual input.To address misalignments that often plague modality integration, our method employs a sophisticated diffusion model that performs back-translation on the text to generate a corresponding visual representation, thus refining the model's multimodal interpretative accuracy.Empirical evidence from the GMNER dataset validates that our approach significantly outperforms existing state-of-theart models.Moreover, the versatility of our pre-training process complements virtually all extant models, offering an additional avenue for augmenting their multimodal entity recognition acumen.
Xigang Bao, Mengyuan Tian, Zhiyuan Zha, Biao Qin
ICMR1
2024 A Sample-driven Selection Framework: Towards Graph Contrastive Networks with Reinforcement Learning
Xiangping Zheng 0002, Xiuxin Hao, Bo Wu 0026, Xigang Bao, Xuan Zhang 0009, Wei Li 0109, Xun Liang 0001
ACM Multimedia4
2023 MPMRC-MNER: A Unified MRC framework for Multimodal Named Entity Recognition based Multimodal Prompt
abstract
Multimodal named entity recognition (MNER) is a vision-language task, which aims to detect entity spans and classify them to corresponding entity types given a sentence-image pair. Existing methods often regard an image as a set of visual objects, trying to explicitly capture the relations between visual objects and entities. However, since visual objects are often not identical to entities in quantity and type, they may suffer the bias introduced by visual objects rather than aid. Inspired by the success of textual prompt-based fine-tuning (PF) approaches in many methods, in this paper, we propose a Multimodal Prompt-based Machine Reading Comprehension based framework to implicit alignment between text and image for improving MNER, namely MPMRC-MNER. Specifically, we transform text-only query in MRC into multimodal prompt containing image tokens and text tokens. To better integrate image tokens and text tokens, we design a prompt-aware attention mechanism for better cross-modal fusion. At last, contrastive learning with two types of contrastive losses is designed to learn more consistent representation of two modalities and reduce noise. Extensive experiments and analyses on two public MNER datasets, Twitter2015 and Twitter2017, demonstrate the better performance of our model against the state-of-the-art methods.
Xigang Bao, Mengyuan Tian, Zhiyuan Zha, Biao Qin
CIKM1
2023 Wukong-CMNER: A Large-Scale Chinese Multimodal NER Dataset with Images Modality
Xigang Bao, Shouhui Wang, Pengnian Qi, Biao Qin
DASFAA (3)1
2023 bfE3-MG: End-to-End Expert Linking via Multi-Granularity Representation Learning
Zhiyuan Zha, Pengnian Qi, Xigang Bao, Biao Qin
ICONIP (13)3