EDBT 2026 Demo / reviewers in the wild / expert
Yongxiu Xu
dblp:294/1202
· DBLP profile ↗
20ranked-venue papers
4as first author
20since 2021 · last 2026
0000-0003-2963-0914ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 3 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 8 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Information-Theoretic Minimal Sufficient Representation for Multi-Domain Knowledge Graph CompletionabstractMulti-domain knowledge graph completion (MKGC) seeks to predict missing triples in a target KG by leveraging triples from multiple KGs in different domains (e.g., languages or sources). Existing studies typically learn and fuse multi-domain KG representations solely with alignments or fusion modules, which can be affected by redundant information within KGs. This issue can conceal task-relevant information in representations, impeding further improvements when scaling to numerous KGs. To this end, we propose IMKGC, an information-theoretic MKGC framework to learn minimal sufficient representations. In particular, IMKGC learns entity representations by explicitly preserving endogenous contextual information within each KG, exogenous complementary information from other KGs, and consistent information of equivalent entities, while suppressing redundant information through variational constraints. Furthermore, we achieve compressed relation representations with a devised relation reasoning decoder that captures relatedness among relations, also improving triple prediction. Extensive experiments on 14 KGs in three benchmark datasets demonstrate that IMKGC significantly outperforms previous state-of-the-art methods, especially in redundant scenarios. Jiawei Sheng, Taoyu Su, Linghui Wang, Yongxiu Xu, Tingwen Liu |
AAAI | 5 |
| 2026 | Event-Centric Structural Modeling for Zero-Shot Video Moment RetrievalabstractZero-Shot Video Moment Retrieval (VMR) aims to localize a specific temporal segment in an untrimmed video that corresponds to a natural language query by leveraging frozen vision-language models, eliminating the need for costly temporal annotations. Fundamentally, an untrimmed video consists of a sequence of atomic events with varying lengths. A critical challenge in VMR is thus to disentangle these events into coherent candidates and accurately identify the one that best aligns with the query. However, conventional methods often overlook this inherent event structure: rigid sliding windows tend to fragment semantically coherent events, while uniform pooling across frames allows ambiguous boundary noise to dilute the relevance of the discriminative core. To mitigate these issues, we propose an Event-Centric Structural Modeling (ECSM) framework. Specifically, our approach replaces rigid windowing with adaptive global segmentation to preserve event integrity. Furthermore, we introduce Gaussian-based weighting to highlight the discriminative core of events while suppressing boundary interference. Extensive experiments on Charades-STA and ActivityNet Captions demonstrate that our method achieves state-of-the-art performance in the training-free setting and exhibits robust generalization across diverse out-of-distribution scenarios. Code is available at https://github.com/youziizii/ECSM. Yongxiu Xu, Yuyao Kong, Gaopeng Gou |
ICMR | 2 |
| 2025 | MAKAR: a Multi-Agent framework based Knowledge-Augmented Reasoning for Grounded Multimodal Named Entity RecognitionabstractXinkui Lin, Yuhui Zhang, Yongxiu Xu, Kun Huang, Hongzhang Mu, Yubin Wang, Gaopeng Gou, Li Qian, Li Peng, Wei Liu, Jian Luan, Hongbo Xu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Xinkui Lin, Yongxiu Xu, Hongzhang Mu, Gaopeng Gou, Wei Liu 0302, Jian Luan 0001 |
EMNLP | 3 |
| 2025 | RDA: Regularized Domain Adaptation for Multimedia Event Extraction
Yongxiu Xu, Xinkui Lin, Gaopeng Gou |
ICIC (9) | 2 |
| 2025 | EilMoB: Emotion-aware Incongruity Learning and Modality Bridging Network for Multi-modal Sarcasm Detection
Yongxiu Xu, Xinkui Lin, Jiarui Lu |
ICMR | 2 |
| 2025 | REMOTE: A Unified Multimodal Relation Extraction Framework with Multilevel Optimal Transport and Mixture-of-ExpertsabstractMultimodal relation extraction (MRE) is a crucial task in the fields of Knowledge Graph and Multimedia, playing a pivotal role in multimodal knowledge graph construction. However, existing methods are typically limited to extracting a single type of relational triplet, which restricts their ability to extract triplets beyond the specified types. Directly combining these methods fails to capture dynamic cross-modal interactions and introduces significant computational redundancy. Therefore, we propose a novel \textit{unified multimodal Relation Extraction framework with Multilevel Optimal Transport and mixture-of-Experts}, termed REMOTE, which can simultaneously extract intra-modal and inter-modal relations between textual entities and visual objects. To dynamically select optimal interaction features for different types of relational triplets, we introduce mixture-of-experts mechanism, ensuring the most relevant modality information is utilized. Additionally, considering that the inherent property of multilayer sequential encoding in existing encoders often leads to the loss of low-level information, we adopt a multilevel optimal transport fusion module to preserve low-level features while maintaining multilayer encoding, yielding more expressive representations. Correspondingly, we also create a Unified Multimodal Relation Extraction (UMRE) dataset to evaluate the effectiveness of our framework, encompassing diverse cases where the head and tail entities can originate from either text or image. Extensive experiments show that REMOTE effectively extracts various types of relational triplets and achieves state-of-the-art performanc on almost all metrics across two other public MRE datasets. We release our resources at https://github.com/Nikol-coder/REMOTE. Xinkui Lin, Yongxiu Xu |
ACM Multimedia | 2 |
| 2025 | An Adaptive Semantic-Aware Fusion Method for Multimodal Entity Linking
Yongxiu Xu, Xinkui Lin, Gaopeng Gou |
NLPCC (1) | 2 |
| 2024 | An Effective Span-based Multimodal Named Entity Recognition with Consistent Cross-Modal AlignmentabstractWith the increasing availability of multimodal content on social media, consisting primarily of text and images, multimodal named entity recognition (MNER) has gained a wide-spread attention. A fundamental challenge of MNER lies in effectively aligning different modalities. However, the majority of current approaches rely on word-based sequence labeling framework and align the image and text at inconsistent semantic levels (whole image-words or regions-words). This misalignment may lead to inferior entity recognition performance. To address this issue, we propose an effective span-based method, named SMNER, which achieves a more consistent multimodal alignment from the perspectives of information-theoretic and cross-modal interaction, respectively. Specifically, we first introduce a cross-modal information bottleneck module for the global-level multimodal alignment (whole image-whole text). This module aims to encourage the semantic distribution of the image to be closer to the semantic distribution of the text, which can enable the filtering out of visual noise. Next, we introduce a cross-modal attention module for the local-level multimodal alignment (regions-spans), which captures the correlations between regions in the image and spans in the text, enabling a more precise alignment of the two modalities. Extensive ex- periments conducted on two benchmark datasets demonstrate that SMNER outperforms the state-of-the-art baselines. Yongxiu Xu, Heyan Huang, Shiyao Cui, Longzheng Wang |
LREC/COLING | 1 |
| 2024 | Position and Type Aware Anchor Link Prediction Across Social Networks
Dongwei Zhu, Yongxiu Xu |
ICANN (9) | 2 |
| 2023 | Broaden Your Horizons: Inter-news Relation Mining for Fake News Detection
Fengzhao Shi, Chaoyang Yan, Yongxiu Xu |
DASFAA (4) | 5 |
| 2023 | PTSTEP: Prompt Tuning for Semantic Typing of Event Processes
Yongxiu Xu, Dongwei Zhu |
ICANN (3) | 2 |
| 2023 | Anchor Link Prediction Based on Trusted Anchor Re-identification
Dongwei Zhu, Yongxiu Xu |
ICANN (8) | 2 |
| 2023 | Multi-hop Reading Comprehension Learning Method Based on Answer Contrastive Learning
Hao You, Heyan Huang, Yue Hu 0002, Yongxiu Xu |
KSEM (4) | 4 |
| 2023 | Cross-modal Contrastive Learning for Multimodal Fake News DetectionabstractAutomatic detection of multimodal fake news has gained a widespread attention recently. Many existing approaches seek to fuse unimodal features to produce multimodal news representations. However, the potential of powerful cross-modal contrastive learning methods for fake news detection has not been well exploited. Besides, how to aggregate features from different modalities to boost the performance of the decision-making process is still an open question. To address that, we propose COOLANT, a cross-modal contrastive learning framework for multimodal fake news detection, aiming to achieve more accurate image-text alignment. To further capture the fine-grained alignment between vision and language, we leverage an auxiliary task to soften the loss term of negative samples during the contrast process. A cross-modal fusion module is developed to learn the cross-modality correlations. An attention mechanism with an attention guidance module is implemented to help effectively and interpretably aggregate the aligned unimodal representations and the cross-modality correlations. Finally, we evaluate the COOLANT and conduct a comparative study on two widely used datasets, Twitter and Weibo. The experimental results demonstrate that our COOLANT outperforms previous approaches by a large margin and achieves new state-of-the-art results on the two datasets. Longzheng Wang, Yongxiu Xu |
ACM Multimedia | 4 |
| 2023 | Exploring Cross-Modal Inconsistency in Entities and Emotions for Multimodal Fake News Detection
Longzheng Wang, Yongxiu Xu |
PRCV (1) | 4 |
| 2023 | Prompting Generative Language Model with Guiding Augmentation for Aspect Sentiment Triplet Extraction
Yongxiu Xu, Xinghua Zhang 0001, Wenyuan Zhang 0002 |
PRICAI (2) | 2 |
| 2022 | DoSEA: A Domain-specific Entity-aware Framework for Cross-Domain Named Entity RecogitionabstractCross-domain named entity recognition aims to improve performance in a target domain with shared knowledge from a well-studied source domain. The previous sequence-labeling based method focuses on promoting model parameter sharing among domains. However, such a paradigm essentially ignores the domain-specific information and suffers from entity type conflicts. To address these issues, we propose a novel machine reading comprehension based framework, named DoSEA, which can identify domain-specific semantic differences and mitigate the subtype conflicts between domains. Concretely, we introduce an entity existence discrimination task and an entity-aware training setting, to recognize inconsistent entity annotations in the source domain and bring additional reference to better share information across domains. Experiments on six datasets prove the effectiveness of our DoSEA. Our source code can be obtained from https://github.com/mhtang1995/DoSEA. Yongquan He, Yongxiu Xu, Chengpeng Chao |
COLING | 4 |
| 2022 | Wlinker: Modeling Relational Triplet Extraction As Word LinkingabstractRelational triplet extraction (RTE) is a fundamental task for automatically extracting information from unstructured text, which has attracted growing interest in recent years. However, it remains challenging due to the difficulty in extracting the overlapping relational triplets. Existing approaches for overlapping RTE, either suffer from exposure bias or designing complex tagging scheme. In light of these limitations, we take an innovative perspective on RTE by modeling it as a word linking problem that learns to link from subject words to object words for each relation type. To this end, we propose a simple but effective multi-task learning model, WLinker, which can extract overlapping relational triplets in an end-to-end fashion. Specifically, we perform word link prediction based on multi-level biaffine attention for leaning the word-level correlations under each relation type. Additionally, our model joint entity detection and word link prediction tasks by a multi-task framework, which combines the local sequential and global dependency structures of words in sentence and captures the implicit interactions between the two tasks. Extensive experiments are conducted on two benchmark datasets NYT and WebNLG. The results demonstrate the effectiveness of WLinker, in comparison with a range of previous state-of-the-art baselines. Yongxiu Xu, Chuan Zhou 0001, Heyan Huang, Jing Yu 0007, Yue Hu 0002 |
ICASSP | 1 |
| 2021 | A Supervised Multi-Head Self-Attention Network for Nested Named Entity RecognitionabstractIn recent years, researchers have shown an increased interest in recognizing the overlapping entities that have nested structures. However, most existing models ignore the semantic correlation between words under different entity types. Considering words in sentence play different roles under different entity types, we argue that the correlation intensities of pairwise words in sentence for each entity type should be considered. In this paper, we treat named entity recognition as a multi-class classification of word pairs and design a simple neural model to handle this issue. Our model applies a supervised multi-head self-attention mechanism, where each head corresponds to one entity type, to construct the word-level correlations for each type. Our model can flexibly predict the span type by the correlation intensities of its head and tail under the corresponding type. In addition, we fuse entity boundary detection and entity classification by a multitask learning framework, which can capture the dependencies between these two tasks. To verify the performance of our model, we conduct extensive experiments on both nested and flat datasets. The experimental results show that our model can outperform the previous state-of-the-art methods on multiple tasks without any extra NLP tools or human annotations. Yongxiu Xu, Heyan Huang, Chong Feng 0001, Yue Hu 0002 |
AAAI | 1 |
| 2021 | A Relation-aware Attention Neural Network for Modeling the Usage of Scientific Online ResourcesabstractMore and more online resources for computer science are introduced, used and released in scientific literature in recent years. Knowledge about the usage of these online resources can help researchers easily find the applicable resources for their works. However, most existing methods ignore the importance of the content of the online resource citations. To this end, we manually create SciR, a dataset that contains 3,012 annotation sentences for this task, and introduce a multi-task learning framework to automatically extract the entities and relations from the context of online resource citations in scientific papers. Furthermore, considering the words in a sentence usually play different roles under different relations. In this paper, we treat different relations as distinctive sub-spaces and model the correlations between words in sentence for each relation type by a supervised biaffine attention network. Based on this relation-aware attention network, our model can not only effectively obtain the word-level correlations under each relation, but also naturally avoid the problem of overlapping relations. To evaluate the effectiveness of our model, we conduct comprehensive experiments on three datasets and the experimental results demonstrate that our model outperforms other state-of-the-art methods on the two tasks of entity recognition and relation extraction. Yongxiu Xu, Heyan Huang, Chong Feng 0001, Chuan Zhou 0001, Jiarui Zhang 0003, Yue Hu 0002 |
IJCNN | 1 |