VLDB 2026 Research / reviewers in the wild / expert
Xiaoyu Li 0004
dblp:18/6855-4
· DBLP profile ↗
13ranked-venue papers
0as first author
13since 2021 · last 2025
0000-0003-0286-6660ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Flexible Optimal Transport With Contrastive Graphical Modeling for Multimodal Hate DetectionabstractMultimodal hate detection plays a crucial role in maintaining harmonious online environments by identifying harmful content, such as hateful memes. Although previous research has made significant progress in detecting explicit hate speech, there remains a critical gap in analyzing implicit hate, which is particularly challenging due to the absence of explicit harmful text claims or demographic visual cues. Despite the promising results based on cross-modal attention, previous methods may suffer from the distributional modality gap caused by the non-literal associations between multimodal elements, which lacks apparent alignment in implicit hateful contents. In this work, we propose a novel framework: Flexible Optimal Transport (FLOT) to capture the non-literal cross-modal alignment for multimodal hate in the context of memes. FLOT formulates the problem of cross-modal alignment as finding optimal transportation plans, which leverages a kernel method to capture complementary information from multiple modalities. The kernel embeddings reproduce a kernel Hilbert space (RKHS) to serve as a non-linear transformation of alignment, which effectively reduces the distributional modality gap with more interpretability. Moreover, we established topological structures with contrastive modeling for the aligned representations, which are optimized to achieve comprehensive alignment between different modalities, and facilitate local reasoning based on multimodal elements. Experimental results have demonstrated that our FLOT achieved state-of-the-art performance on three publicly available benchmark datasets. Furthermore, extensive qualitative analysis confirms the superior ability of FLOT in capturing implicit cross-modal alignment. Linhao Zhang, Li Jin 0001, Xiaoyu Li 0004, Xian Sun 0001, Xin Wang 0117, Zequn Zhang, Jian Liu 0032, Zhicong Lu, Guangluan Xu |
IEEE Trans. Multim. | 3 |
| 2024 | CAMEL: Capturing Metaphorical Alignment with Context Disentangling for Multimodal Emotion RecognitionabstractUnderstanding the emotional polarity of multimodal content with metaphorical characteristics, such as memes, poses a significant challenge in Multimodal Emotion Recognition (MER). Previous MER researches have overlooked the phenomenon of metaphorical alignment in multimedia content, which involves non-literal associations between concepts to convey implicit emotional tones. Metaphor-agnostic MER methods may be misinformed by the isolated unimodal emotions, which are distinct from the real emotions blended in multimodal metaphors. Moreover, contextual semantics can further affect the emotions associated with similar metaphors, leading to the challenge of maintaining contextual compatibility. To address the issue of metaphorical alignment in MER, we propose to leverage a conditional generative approach for capturing metaphorical analogies. Our approach formulates schematic prompts and corresponding references based on theoretical foundations, which allows the model to better grasp metaphorical nuances. In order to maintain contextual sensitivity, we incorporate a disentangled contrastive matching mechanism, which undergoes curricular adjustment to regulate its intensity during the learning process. The automatic and human evaluation experiments on two benchmarks prove that, our model provides considerable and stable improvements in recognizing multimodal emotion with metaphor attributes. Linhao Zhang, Li Jin 0001, Guangluan Xu, Xiaoyu Li 0004, Kaiwen Wei, Nayu Liu |
AAAI | 4 |
| 2024 | More Than Syntaxes: Investigating Semantics to Zero-shot Cross-lingual Relation Extraction and Event Argument Role LabellingabstractSyntactic dependency structures are commonly utilized as language-agnostic features to solve the word order difference issues in zero-shot cross-lingual relation and event extraction tasks. However, while sentences in multiple forms can be employed to express the same meaning, the syntactic structure may vary considerably in specific scenarios. To fix this problem, we find semantics are rarely considered, which could provide a more consistent semantic analysis of sentences and be served as another bridge between different languages. Therefore, in this article, we introduce Syntax and Semantic Driven Network (SSDN) to equip syntax and semantic knowledge across languages simultaneously. Specifically, predicate–argument structures from semantic role labelling are explicitly incorporated into word representations. Then, a semantic-aware relational graph convolutional network and a transformer-based encoder are utilized to model both semantic dependency and syntactic dependency structures, respectively. Finally, a fusion module is introduced to integrate output representations adaptively. We conduct experiments on the widely used Automatic Content Extraction 2005 English, Chinese, and Arabic datasets. The evaluation results demonstrate that the proposed method achieves the state-of-the-art performance. Further study also indicates SSDN could produce robust representations that facilitate the transfer operations across languages. Kaiwen Wei, Li Jin 0001, Zequn Zhang, Zhi Guo, Xiaoyu Li 0004, Qing Liu 0021, Weimiao Feng |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 5 |
| 2024 | Relation-Aware Multi-Pass Comparison Deconfounded Network for Change CaptioningabstractChange captioning aims to describe the semantic change between a pair of images with natural language while remaining immune to viewpoint change. Based on the encoder-decoder architecture, most existing methods primarily focus on encoding effective change representations for transmission to the decoder. However, they suffer from an insufficient understanding of visual semantics, inadequate single-pass feature comparison, and a confounding bias caused by imbalanced viewpoint change data. These impair change representations and hinder unbiased caption generation. In this paper, we analyze and identify the confounding bias from a causality perspective and propose a Relation-aware Multi-pass Comparison Deconfounded (RMCD) network for change captioning, which elevates the encoding of change representations and mitigates the bias. Specifically, in the encoding stage, to sufficiently understand visual semantics, a position-guided context aggregating module is presented to capture the positional and contextual relations among objects in the image. Then, to achieve comprehensive change representations, we present a multi-pass feature comparison module to recognize semantic differences at various feature levels and progressively integrate them. In the decoding stage, to generate de-biased captions, the causal intervention is employed to remove the confounding bias which introduces spurious correlations between encoded change representations and captions. The newly achieved state-of-the-art performance on four publicly available benchmark datasets and further visual analysis demonstrate the superiority of our method. Zhicong Lu, Li Jin 0001, Changyuan Tian 0001, Xian Sun 0001, Xiaoyu Li 0004, Yi Zhang 0083, Qi Li 0051, Guangluan Xu |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2023 | TOT:Topology-Aware Optimal Transport for Multimodal Hate DetectionabstractMultimodal hate detection, which aims to identify the harmful content online such as memes, is crucial for building a wholesome internet environment. Previous work has made enlightening exploration in detecting explicit hate remarks. However, most of their approaches neglect the analysis of implicit harm, which is particularly challenging as explicit text markers and demographic visual cues are often twisted or missing. The leveraged cross-modal attention mechanisms also suffer from the distributional modality gap and lack logical interpretability. To address these semantic gap issues, we propose TOT: a topology-aware optimal transport framework to decipher the implicit harm in memes scenario, which formulates the cross-modal aligning problem as solutions for optimal transportation plans. Specifically, we leverage an optimal transport kernel method to capture complementary information from multiple modalities. The kernel embedding provides a non-linear transformation ability to reproduce a kernel Hilbert space (RKHS), which reflects significance for eliminating the distributional modality gap. Moreover, we perceive the topology information based on aligned representations to conduct bipartite graph path reasoning. The newly achieved state-of-the-art performance on two publicly available benchmark datasets, together with further visual analysis, demonstrate the superiority of TOT in capturing implicit cross-modal alignment. Linhao Zhang, Li Jin 0001, Xian Sun 0001, Guangluan Xu, Zequn Zhang, Xiaoyu Li 0004, Nayu Liu, Qing Liu 0021, Shiyao Yan |
AAAI | 6 |
| 2023 | Event Causality Extraction via Implicit Cause-Effect InteractionsabstractEvent Causality Extraction (ECE) aims to extract the cause-effect event pairs from the given text, which requires the model to possess a strong reasoning ability to capture event causalities.However, existing works have not adequately exploited the interactions between the cause and effect event that could provide crucial clues for causality reasoning.To this end, we propose an Implicit Cause-Effect interaction (ICE) framework, which formulates ECE as a template-based conditional generation problem.The proposed method captures the implicit intra-and inter-event interactions by incorporating the privileged information (ground truth event types and arguments) for reasoning, and a knowledge distillation mechanism is introduced to alleviate the unavailability of privileged information in the test stage.Furthermore, to facilitate knowledge transfer from teacher to student, we design an event-level alignment strategy named Cause-Effect Optimal Transport (CEOT) to strengthen the semantic interactions of cause-effect event types and arguments.Experimental results indicate that ICE achieves state-of-the-art performance on the ECE-CCKS dataset. Zequn Zhang, Kaiwen Wei, Zhi Guo, Xian Sun 0001, Li Jin 0001, Xiaoyu Li 0004 |
EMNLP | 7 |
| 2023 | Emotion-cause pair extraction with bidirectional multi-label sequence tagging
Zequn Zhang, Zhi Guo, Li Jin 0001, Xiaoyu Li 0004, Kaiwen Wei, Xian Sun 0001 |
Appl. Intell. | 5 |
| 2023 | KEPT: Knowledge Enhanced Prompt Tuning for event causality identification
Zequn Zhang, Zhi Guo, Li Jin 0001, Xiaoyu Li 0004, Kaiwen Wei, Xian Sun 0001 |
Knowl. Based Syst. | 5 |
| 2022 | DPNet: domain-aware prototypical network for interdisciplinary few-shot relation classification
Li Jin 0001, Xiaoyu Li 0004, Xian Sun 0001, Zhi Guo, Zequn Zhang, Shuchao Li |
Appl. Intell. | 3 |
| 2022 | SF-ANN: leveraging structural features with an attention neural network for candidate fact ranking
Li Jin 0001, Zequn Zhang, Xiaoyu Li 0004, Qing Liu 0021 |
Appl. Intell. | 4 |
| 2022 | TSPNet: Translation supervised prototype network via residual learning for multimodal social relation extraction
Hankun Kang, Xiaoyu Li 0004, Li Jin 0001, Zequn Zhang, Shuchao Li |
Neurocomputing | 2 |
| 2022 | Trigger is Non-central: Jointly event extraction via label-aware representations with multi-task learning
Jianwei Lv, Zequn Zhang, Li Jin 0001, Shuchao Li, Xiaoyu Li 0004, Guangluan Xu, Xian Sun 0001 |
Knowl. Based Syst. | 5 |
| 2021 | HGEED: Hierarchical graph enhanced event detection
Jianwei Lv, Zequn Zhang, Li Jin 0001, Shuchao Li, Xiaoyu Li 0004, Guangluan Xu, Xian Sun 0001 |
Neurocomputing | 5 |