Rui Dong 0002

dblp:127/0871-2 · DBLP profile ↗
← Back
17ranked-venue papers
0as first author
15since 2021 · last 2026
0000-0002-4110-3976ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 M3UCD: A Multi-task Multimodal Metaphor Understanding Challenge Dataset for LLMs
abstract
Understanding multimodal metaphors represents a crucial pathway for machines to comprehend human cognition. However, current research remains constrained by superficial dataset annotations, insufficient systematic evaluation of large language models, and fragmented task frameworks. To bridge these gaps, the paper proposes a systematic solution featuring: (I) We present the largest fine-grained Multi-task Multimodal Metaphor Understanding Challenge Dataset (M3UCD) built via multi-perspective collaborative annotation. It contains 15,345 samples, each annotated with 12 manual attribute labels. (II) Systematic benchmarking of LLMs' capacity boundaries in metaphor understanding. Evaluation results reveal the persistent challenges LLMs face in this domain while validating M3UCD's effectiveness and potential. (III) A concise and unified multi-task baseline framework was developed and demonstrated its effectiveness in enhancing the metaphor understanding capabilities of MLLMs.
Tianlong Zheng, Yating Yang, Rui Dong 0002, Bo Ma 0004, Lei Wang 0065, Xi Zhou 0007, Siru Miao, Osman Turghun
AAAI3
2026 Flickr-ABS: A Benchmark for Abstract Intent in Image-Text Retrieval
abstract
In image–text retrieval, users often express abstract, incomplete, or polysemous intents rather than fully concrete visual descriptions. In contrast, existing retrieval benchmarks predominantly rely on object- and action-centric captions that presuppose detailed and visually grounded descriptions, leaving model behavior under semantically fuzzy queries largely underexplored. To address this gap, we introduce Flickr-ABS, an abstraction-aware image–text retrieval benchmark built upon Flickr30k. For each image, Flickr-ABS associates retrieval queries at three explicitly defined semantic levels, ranging from concrete scene descriptions to high-level conceptual interpretations. Extensive evaluations on a wide range of vision–language models reveal a substantial degradation in retrieval performance as query abstraction increases, highlighting the fragility of current retrieval systems under abstract semantic intent. We further demonstrate that jointly leveraging supervision from multiple abstraction levels can improve retrieval robustness for abstract queries while preserving performance on concrete ones. We further propose a semantic recall protocol that accounts for conceptual equivalence, enabling more faithful evaluation under high-level abstraction. Together, Flickr-ABS and the proposed evaluation protocol offer a practical benchmark for analyzing and improving abstraction-aware image–text retrieval. Our code is available at https://github.com/LT1anY/Flickr-abs.
Tianyuan Li, Lei Wang 0065, Yating Yang, Rui Dong 0002, Ahtamjan Ahmat, Bangju Han
ICMR4
2025 Low-Resource Language Expansion and Translation Capacity Enhancement for LLM: A Study on the Uyghur
abstract
Although large language models have significantly advanced natural language generation, their potential in low-resource machine translation has not yet been fully explored, especially for languages that translation models have not been trained on. In this study, we provide a detailed demonstration of how to efficiently expand low-resource languages for large language models and significantly enhance the model’s translation ability, using Uyghur as an example. The process involves four stages: collecting and pre-processing monolingual data, conducting continuous pre-training with extensive monolingual data, fine-tuning with less parallel corpora using translation supervision, and proposing a direct preference optimization based on translation self-evolution (DPOSE) on this basis. Extensive experiments have shown that our strategy effectively expands the low-resource languages supported by large language models and significantly enhances the model’s translation ability in Uyghur with less parallel data. Our research provides detailed insights for expanding other low-resource languages into large language models.
Kaiwen Lu, Yating Yang, Fengyi Yang, Rui Dong 0002, Bo Ma 0004, Aihetamujiang Aihemaiti, Abibulla Atawulla, Lei Wang 0065, Xi Zhou 0007
COLING4
2025 OpenForecast: A Large-Scale Open-Ended Event Forecasting Dataset
abstract
Complex events generally exhibit unforeseen, multifaceted, and multi-step developments, and cannot be well handled by existing closed-ended event forecasting methods, which are constrained by a limited answer space. In order to accelerate the research on complex event forecasting, we introduce OpenForecast, a large-scale open-ended dataset with two features: (1) OpenForecast defines three open-ended event forecasting tasks, enabling unforeseen, multifaceted, and multi-step forecasting. (2) OpenForecast collects and annotates a large-scale dataset from Wikipedia and news, including 43,419 complex events spanning from 1950 to 2024. Particularly, this annotation can be completed automatically without any manual annotation cost. Meanwhile, we introduce an automatic LLM-based Retrieval-Augmented Evaluation method (LRAE) for complex events, enabling OpenForecast to evaluate the ability of complex event forecasting of large language models. Finally, we conduct comprehensive human evaluations to verify the quality and challenges of OpenForecast, and the consistency between LEAE metric and human evaluation. OpenForecast and related codes will be publicly released.
Zhen Wang 0062, Xi Zhou 0007, Yating Yang, Bo Ma 0004, Lei Wang 0065, Rui Dong 0002, Azmat Anwar
COLING6
2025 Mining the Past with Dual Criteria: Integrating Three types of Historical Information for Context-aware Event Forecasting
abstract
Event forecasting requires modeling historical event data to predict future events, and achieving accurate predictions depends on effectively capturing the relevant historical information that aids forecasting.Most existing methods focus on entities and structural dependencies to capture historical clues but often overlook implicitly relevant information.This limitation arises from overlooking event semantics and deeper factual associations that are not explicitly connected in the graph structure but are nonetheless critical for accurate forecasting.To address this, we propose a dual-criteria constraint strategy that leverages event semantics for relevance modeling and incorporates a self-supervised semantic filter based on factual event associations to capture implicitly relevant historical information.Building on this strategy, our method, termed ITHI (Integrating Three types of Historical Information), combines sequential event information, periodically repeated event information, and relevant historical information to achieve context-aware event forecasting.We evaluated the proposed ITHI method on three public benchmark datasets, achieving state-of-the-art performance and significantly outperforming existing approaches.Additionally, we validated its effectiveness on two structured temporal knowledge graph forecasting dataset 1 .
Lei Wang 0065, Yating Yang, Bo Ma 0004, Rui Dong 0002, Fengyi Yang, Ahtamjan Ahmat, Kaiwen Lu
EMNLP5
2025 Bidirectional Feature Fusion and Adaptive Decision Network for Multimodal Fake News Detection
abstract
The proliferation of multimedia content on social media platforms has rendered the detection of fake news an increasingly critical challenge. Despite the progress made in multimodal fake news detection methods, significant challenges remain: information loss during the cross-modal feature fusion process and restricted analysis due to single-perspective decision mechanisms. We propose the Bidirectional Feature Fusion and Adaptive Decision Network (BFF-ADN), which introduces two key innovations: a bidirectional token-level feature fusion mechanism that captures key correlations between visual and textual modalities while reducing information loss, and an Adaptive Multi-perspective Decision Integration Network (AMDI-Net) that integrates multiple specialized detectors for comprehensive analysis. Experiments on three real-world datasets (Weibo, GossipCop, and PolitiFact) demonstrate that BFF-ADN achieves state-of-the-art performance. Ablation studies and visualizations confirm the effectiveness of our proposed components, particularly highlighting the advantages of the bidirectional feature fusion mechanism and Adaptive Multi-perspective Decision Integration Network.
Dilxat Abdureyim, Bo Ma 0004, Yating Yang, Rui Dong 0002, YiDu Chen, Azmat Anwar, Lei Wang 0065
ICME4
2025 LyRE: Learning Varying Fusion Degrees with Hierarchical Aggregation to Improve Multimodal Misinformation Detection
YiDu Chen, Bo Ma 0004, Yating Yang, Dilxat Abdureyim, Rui Dong 0002, Zhen Wang 0062, Lei Wang 0065, Xi Zhou 0007
NLPCC (2)5
2025 Self-Distillation Across Modalities: Enhancing Cross-Modal Correlation Perception for Multimodal Fake News Detection
YiDu Chen, Bo Ma 0004, Yating Yang, Dilxat Abdureyim, Chirui Zhang, Rui Dong 0002, Lei Wang 0065, Xi Zhou 0007
NLPCC (3)6
2025 ISIPN: Intention-Semantic Incongruity Perception Network for Multimodal Metaphor Detection
Tianlong Zheng, Yating Yang, Rui Dong 0002, Bo Ma 0004, Lei Wang 0065, Xi Zhou 0007
NLPCC (3)3
2025 M2-VLP: Enhancing Multilingual Vision-Language Pre-Training via Multi-Grained Alignment
abstract
Recently, multilingual Vision-Language Pre-training (mVLP) has shown remarkable progress in learning joint representations across different modalities and languages. However, most existing methods learn semantic alignment at a coarse-grained level and fail to capture fine-grained correlations between different languages and modalities. To address this, we propose a Multi-grained Multilingual Vision-Language Pre-training (M2-VLP) model, which aims to learn cross-lingual cross-modal alignment at different semantic granular levels. In cross-lingual interaction, the model learns the global alignment of parallel sentence pairs and the word-level correlations. In cross-modal interaction, the model aligns images with captions and image regions with corresponding words. To integrate the cross-lingual and cross-modal alignment above, we propose a unified multi-grained contrastive learning paradigm. Under zero-shot cross-lingual and fine-tuned multilingual settings, extensive experiments on vision-language downstream tasks across twenty languages demonstrate the effectiveness of M2-VLP over competitive contrastive models. Code and models are available at https://github.com/ahtamjan/M2-VLP.
Ahtamjan Ahmat, Lei Wang 0065, Yating Yang, Bo Ma 0004, Rui Dong 0002, Kaiwen Lu
WWW5
2025 Cross-lingual prompting method with semantic-based answer space clustering
Ahtamjan Ahmat, Yating Yang, Bo Ma 0004, Rui Dong 0002, Lei Wang 0065
Appl. Intell.4
2025 SGEU: enhancing LLM reasoning via backward exemplar generation and verification
Zhen Wang 0062, Xi Zhou 0007, Yating Yang, Bo Ma 0004, Lei Wang 0065, Rui Dong 0002
Appl. Intell.6
2024 Pruning Residual Networks in Multilingual Neural Machine Translation to Improve Zero-Shot Translation
Kaiwen Lu, Yating Yang, Rui Dong 0002, Bo Ma 0004, Lei Wang 0065, Xi Zhou 0007, Ahtamjan Ahmat
NLPCC (3)3
2023 A Domain-Transfer Meta Task Design Paradigm for Few-Shot Slot Tagging
abstract
Few-shot slot tagging is an important task in dialogue systems and attracts much attention of researchers. Most previous few-shot slot tagging methods utilize meta-learning procedure for training and strive to construct a large number of different meta tasks to simulate the testing situation of insufficient data. However, there is a widespread phenomenon of overlap slot between two domains in slot tagging. Traditional meta tasks ignore this special phenomenon and cannot simulate such realistic few-shot slot tagging scenarios. It violates the basic principle of meta-learning which the meta task is consistent with the real testing task, leading to historical information forgetting problem. In this paper, we introduce a novel domain-transfer meta task design paradigm to tackle this problem. We distribute a basic domain to each target domain based on the coincidence degree of slot labels between these two domains. Unlike classic meta tasks which only rely on small samples of target domain, our meta tasks aim to correctly infer the class of target domain query samples based on both abundant data in basic domain and scarce data in target domain. To accomplish our meta task, we propose a Task Adaptation Network to effectively transfer the historical information from the basic domain to the target domain. We carry out sufficient experiments on the benchmark slot tagging dataset SNIPS and the name entity recognition dataset NER. Results demonstrate that our proposed model outperforms previous methods and achieves the state-of-the-art performance.
Fengyi Yang, Xi Zhou 0007, Yating Yang, Bo Ma 0004, Rui Dong 0002, Abibulla Atawulla
AAAI5
2023 WAD-X: Improving Zero-shot Cross-lingual Transfer via Adapter-based Word Alignment
abstract
Multilingual pre-trained language models (mPLMs) have achieved remarkable performance on zero-shot cross-lingual transfer learning. However, most mPLMs implicitly encourage cross-lingual alignment in pre-training stage, making it hard to capture accurate word alignment across languages. In this paper, we propose Word-align ADapters for Cross-lingual transfer (WAD-X) to explicitly align word representations of mPLMs via language-specific subspace. Taking a mPLM as the backbone model, WAD-X constructs subspace for each source-target language pair via adapters. The adapters use statistical alignment as the prior knowledge to guide word-level aligning in the corresponding bilingual semantic subspace. We evaluate our model across a set of target languages on three zero-shot cross-lingual transfer tasks: part-of-speech tagging (POS), dependency parsing (DP), and sentiment analysis (SA). Experimental results demonstrate that our proposed model improves zero-shot cross-lingual transfer on three benchmarks, with improvements of 2.19, 2.50, and 1.61 points in POS, DP, and SA tasks over strong baselines.
Ahtamjan Ahmat, Yating Yang, Bo Ma 0004, Rui Dong 0002, Kaiwen Lu, Lei Wang 0065
ACM Trans. Asian Low Resour. Lang. Inf. Process.4
2019 Research of Uyghur-Chinese Machine Translation System Combination Based on Semantic Information
Xiao Li 0007, Yating Yang, Azmat Anwar, Rui Dong 0002
NLPCC (2)5
2015 Optimized Uyghur Segmentation for Statistical Machine Translation
Chenggang Mi 0002, Yating Yang, Rui Dong 0002, Xi Zhou 0007, Lei Wang 0065, Xiao Li 0007, Tonghai Jiang, Osman Turghun
NLDB3