Lei Wang 0065

dblp:w/LeiWang65 · DBLP profile ↗
← Back
24ranked-venue papers
0as first author
16since 2021 · last 2026
0000-0001-6685-0858ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 12 since 2021Databases, data management, data science and information retrieval · 4 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 M3UCD: A Multi-task Multimodal Metaphor Understanding Challenge Dataset for LLMs
abstract
Understanding multimodal metaphors represents a crucial pathway for machines to comprehend human cognition. However, current research remains constrained by superficial dataset annotations, insufficient systematic evaluation of large language models, and fragmented task frameworks. To bridge these gaps, the paper proposes a systematic solution featuring: (I) We present the largest fine-grained Multi-task Multimodal Metaphor Understanding Challenge Dataset (M3UCD) built via multi-perspective collaborative annotation. It contains 15,345 samples, each annotated with 12 manual attribute labels. (II) Systematic benchmarking of LLMs' capacity boundaries in metaphor understanding. Evaluation results reveal the persistent challenges LLMs face in this domain while validating M3UCD's effectiveness and potential. (III) A concise and unified multi-task baseline framework was developed and demonstrated its effectiveness in enhancing the metaphor understanding capabilities of MLLMs.
Tianlong Zheng, Yating Yang, Rui Dong 0002, Bo Ma 0004, Lei Wang 0065, Xi Zhou 0007, Siru Miao, Osman Turghun
AAAI5
2026 Flickr-ABS: A Benchmark for Abstract Intent in Image-Text Retrieval
abstract
In image–text retrieval, users often express abstract, incomplete, or polysemous intents rather than fully concrete visual descriptions. In contrast, existing retrieval benchmarks predominantly rely on object- and action-centric captions that presuppose detailed and visually grounded descriptions, leaving model behavior under semantically fuzzy queries largely underexplored. To address this gap, we introduce Flickr-ABS, an abstraction-aware image–text retrieval benchmark built upon Flickr30k. For each image, Flickr-ABS associates retrieval queries at three explicitly defined semantic levels, ranging from concrete scene descriptions to high-level conceptual interpretations. Extensive evaluations on a wide range of vision–language models reveal a substantial degradation in retrieval performance as query abstraction increases, highlighting the fragility of current retrieval systems under abstract semantic intent. We further demonstrate that jointly leveraging supervision from multiple abstraction levels can improve retrieval robustness for abstract queries while preserving performance on concrete ones. We further propose a semantic recall protocol that accounts for conceptual equivalence, enabling more faithful evaluation under high-level abstraction. Together, Flickr-ABS and the proposed evaluation protocol offer a practical benchmark for analyzing and improving abstraction-aware image–text retrieval. Our code is available at https://github.com/LT1anY/Flickr-abs.
Tianyuan Li, Lei Wang 0065, Yating Yang, Rui Dong 0002, Ahtamjan Ahmat, Bangju Han
ICMR2
2025 Low-Resource Language Expansion and Translation Capacity Enhancement for LLM: A Study on the Uyghur
abstract
Although large language models have significantly advanced natural language generation, their potential in low-resource machine translation has not yet been fully explored, especially for languages that translation models have not been trained on. In this study, we provide a detailed demonstration of how to efficiently expand low-resource languages for large language models and significantly enhance the model’s translation ability, using Uyghur as an example. The process involves four stages: collecting and pre-processing monolingual data, conducting continuous pre-training with extensive monolingual data, fine-tuning with less parallel corpora using translation supervision, and proposing a direct preference optimization based on translation self-evolution (DPOSE) on this basis. Extensive experiments have shown that our strategy effectively expands the low-resource languages supported by large language models and significantly enhances the model’s translation ability in Uyghur with less parallel data. Our research provides detailed insights for expanding other low-resource languages into large language models.
Kaiwen Lu, Yating Yang, Fengyi Yang, Rui Dong 0002, Bo Ma 0004, Aihetamujiang Aihemaiti, Abibulla Atawulla, Lei Wang 0065, Xi Zhou 0007
COLING8
2025 OpenForecast: A Large-Scale Open-Ended Event Forecasting Dataset
abstract
Complex events generally exhibit unforeseen, multifaceted, and multi-step developments, and cannot be well handled by existing closed-ended event forecasting methods, which are constrained by a limited answer space. In order to accelerate the research on complex event forecasting, we introduce OpenForecast, a large-scale open-ended dataset with two features: (1) OpenForecast defines three open-ended event forecasting tasks, enabling unforeseen, multifaceted, and multi-step forecasting. (2) OpenForecast collects and annotates a large-scale dataset from Wikipedia and news, including 43,419 complex events spanning from 1950 to 2024. Particularly, this annotation can be completed automatically without any manual annotation cost. Meanwhile, we introduce an automatic LLM-based Retrieval-Augmented Evaluation method (LRAE) for complex events, enabling OpenForecast to evaluate the ability of complex event forecasting of large language models. Finally, we conduct comprehensive human evaluations to verify the quality and challenges of OpenForecast, and the consistency between LEAE metric and human evaluation. OpenForecast and related codes will be publicly released.
Zhen Wang 0062, Xi Zhou 0007, Yating Yang, Bo Ma 0004, Lei Wang 0065, Rui Dong 0002, Azmat Anwar
COLING5
2025 Mining the Past with Dual Criteria: Integrating Three types of Historical Information for Context-aware Event Forecasting
abstract
Event forecasting requires modeling historical event data to predict future events, and achieving accurate predictions depends on effectively capturing the relevant historical information that aids forecasting.Most existing methods focus on entities and structural dependencies to capture historical clues but often overlook implicitly relevant information.This limitation arises from overlooking event semantics and deeper factual associations that are not explicitly connected in the graph structure but are nonetheless critical for accurate forecasting.To address this, we propose a dual-criteria constraint strategy that leverages event semantics for relevance modeling and incorporates a self-supervised semantic filter based on factual event associations to capture implicitly relevant historical information.Building on this strategy, our method, termed ITHI (Integrating Three types of Historical Information), combines sequential event information, periodically repeated event information, and relevant historical information to achieve context-aware event forecasting.We evaluated the proposed ITHI method on three public benchmark datasets, achieving state-of-the-art performance and significantly outperforming existing approaches.Additionally, we validated its effectiveness on two structured temporal knowledge graph forecasting dataset 1 .
Lei Wang 0065, Yating Yang, Bo Ma 0004, Rui Dong 0002, Fengyi Yang, Ahtamjan Ahmat, Kaiwen Lu
EMNLP2
2025 Bidirectional Feature Fusion and Adaptive Decision Network for Multimodal Fake News Detection
abstract
The proliferation of multimedia content on social media platforms has rendered the detection of fake news an increasingly critical challenge. Despite the progress made in multimodal fake news detection methods, significant challenges remain: information loss during the cross-modal feature fusion process and restricted analysis due to single-perspective decision mechanisms. We propose the Bidirectional Feature Fusion and Adaptive Decision Network (BFF-ADN), which introduces two key innovations: a bidirectional token-level feature fusion mechanism that captures key correlations between visual and textual modalities while reducing information loss, and an Adaptive Multi-perspective Decision Integration Network (AMDI-Net) that integrates multiple specialized detectors for comprehensive analysis. Experiments on three real-world datasets (Weibo, GossipCop, and PolitiFact) demonstrate that BFF-ADN achieves state-of-the-art performance. Ablation studies and visualizations confirm the effectiveness of our proposed components, particularly highlighting the advantages of the bidirectional feature fusion mechanism and Adaptive Multi-perspective Decision Integration Network.
Dilxat Abdureyim, Bo Ma 0004, Yating Yang, Rui Dong 0002, YiDu Chen, Azmat Anwar, Lei Wang 0065
ICME7
2025 LyRE: Learning Varying Fusion Degrees with Hierarchical Aggregation to Improve Multimodal Misinformation Detection
YiDu Chen, Bo Ma 0004, Yating Yang, Dilxat Abdureyim, Rui Dong 0002, Zhen Wang 0062, Lei Wang 0065, Xi Zhou 0007
NLPCC (2)7
2025 Self-Distillation Across Modalities: Enhancing Cross-Modal Correlation Perception for Multimodal Fake News Detection
YiDu Chen, Bo Ma 0004, Yating Yang, Dilxat Abdureyim, Chirui Zhang, Rui Dong 0002, Lei Wang 0065, Xi Zhou 0007
NLPCC (3)7
2025 ISIPN: Intention-Semantic Incongruity Perception Network for Multimodal Metaphor Detection
Tianlong Zheng, Yating Yang, Rui Dong 0002, Bo Ma 0004, Lei Wang 0065, Xi Zhou 0007
NLPCC (3)5
2025 M2-VLP: Enhancing Multilingual Vision-Language Pre-Training via Multi-Grained Alignment
abstract
Recently, multilingual Vision-Language Pre-training (mVLP) has shown remarkable progress in learning joint representations across different modalities and languages. However, most existing methods learn semantic alignment at a coarse-grained level and fail to capture fine-grained correlations between different languages and modalities. To address this, we propose a Multi-grained Multilingual Vision-Language Pre-training (M2-VLP) model, which aims to learn cross-lingual cross-modal alignment at different semantic granular levels. In cross-lingual interaction, the model learns the global alignment of parallel sentence pairs and the word-level correlations. In cross-modal interaction, the model aligns images with captions and image regions with corresponding words. To integrate the cross-lingual and cross-modal alignment above, we propose a unified multi-grained contrastive learning paradigm. Under zero-shot cross-lingual and fine-tuned multilingual settings, extensive experiments on vision-language downstream tasks across twenty languages demonstrate the effectiveness of M2-VLP over competitive contrastive models. Code and models are available at https://github.com/ahtamjan/M2-VLP.
Ahtamjan Ahmat, Lei Wang 0065, Yating Yang, Bo Ma 0004, Rui Dong 0002, Kaiwen Lu
WWW2
2025 Cross-lingual prompting method with semantic-based answer space clustering
Ahtamjan Ahmat, Yating Yang, Bo Ma 0004, Rui Dong 0002, Lei Wang 0065
Appl. Intell.6
2025 SGEU: enhancing LLM reasoning via backward exemplar generation and verification
Zhen Wang 0062, Xi Zhou 0007, Yating Yang, Bo Ma 0004, Lei Wang 0065, Rui Dong 0002
Appl. Intell.5
2024 Pruning Residual Networks in Multilingual Neural Machine Translation to Improve Zero-Shot Translation
Kaiwen Lu, Yating Yang, Rui Dong 0002, Bo Ma 0004, Lei Wang 0065, Xi Zhou 0007, Ahtamjan Ahmat
NLPCC (3)5
2024 Relational concept enhanced prototypical network for incremental few-shot relation classification
Bo Ma 0004, Lei Wang 0065, Xi Zhou 0007, Zhen Wang 0062, Yating Yang
Knowl. Based Syst.3
2023 WAD-X: Improving Zero-shot Cross-lingual Transfer via Adapter-based Word Alignment
abstract
Multilingual pre-trained language models (mPLMs) have achieved remarkable performance on zero-shot cross-lingual transfer learning. However, most mPLMs implicitly encourage cross-lingual alignment in pre-training stage, making it hard to capture accurate word alignment across languages. In this paper, we propose Word-align ADapters for Cross-lingual transfer (WAD-X) to explicitly align word representations of mPLMs via language-specific subspace. Taking a mPLM as the backbone model, WAD-X constructs subspace for each source-target language pair via adapters. The adapters use statistical alignment as the prior knowledge to guide word-level aligning in the corresponding bilingual semantic subspace. We evaluate our model across a set of target languages on three zero-shot cross-lingual transfer tasks: part-of-speech tagging (POS), dependency parsing (DP), and sentiment analysis (SA). Experimental results demonstrate that our proposed model improves zero-shot cross-lingual transfer on three benchmarks, with improvements of 2.19, 2.50, and 1.61 points in POS, DP, and SA tasks over strong baselines.
Ahtamjan Ahmat, Yating Yang, Bo Ma 0004, Rui Dong 0002, Kaiwen Lu, Lei Wang 0065
ACM Trans. Asian Low Resour. Lang. Inf. Process.6
2021 In silico drug repositioning using deep learning and comprehensive similarity measures
abstract
BACKGROUND: Drug repositioning, meanings finding new uses for existing drugs, which can accelerate the processing of new drugs research and development. Various computational methods have been presented to predict novel drug-disease associations for drug repositioning based on similarity measures among drugs and diseases. However, there are some known associations between drugs and diseases that previous studies not utilized. METHODS: In this work, we develop a deep gated recurrent units model to predict potential drug-disease interactions using comprehensive similarity measures and Gaussian interaction profile kernel. More specifically, the similarity measure is used to exploit discriminative feature for drugs based on their chemical fingerprints. Meanwhile, the Gaussian interactions profile kernel is employed to obtain efficient feature of diseases based on known disease-disease associations. Then, a deep gated recurrent units model is developed to predict potential drug-disease interactions. RESULTS: The performance of the proposed model is evaluated on two benchmark datasets under tenfold cross-validation. And to further verify the predictive ability, case studies for predicting new potential indications of drugs were carried out. CONCLUSION: The experimental results proved the proposed model is a useful tool for predicting new indications for drugs or new treatments for diseases, and can accelerate drug repositioning and related drug research and discovery.
Zhu-Hong You, Lei Wang 0065, Xiao-Rui Su 0001, Xi Zhou 0007, Tonghai Jiang
BMC Bioinform.3
2020 Topological Graph Representation Learning on Property Graph
Yishuo Zhang, Daniel Gao, Aswani Kumar Cherukuri, Lei Wang 0065, Shaowei Pan
KSEM (1)4
2018 Toward Better Loanword Identification in Uyghur Using Cross-lingual Word Embeddings
abstract
To enrich vocabulary of low resource settings, we proposed a novel method which identify loanwords in monolingual corpora. More specifically, we first use cross-lingual word embeddings as the core feature to generate semantically related candidates based on comparable corpora and a small bilingual lexicon; then, a log-linear model which combines several shallow features such as pronunciation similarity and hybrid language model features to predict the final results. In this paper, we use Uyghur as the receipt language and try to detect loanwords in four donor languages: Arabic, Chinese, Persian and Russian. We conduct two groups of experiments to evaluate the effectiveness of our proposed approach: loanword identification and OOV translation in four language pairs and eight translation directions (Uyghur-Arabic, Arabic-Uyghur, Uyghur-Chinese, Chinese-Uyghur, Uyghur-Persian, Persian-Uyghur, Uyghur-Russian, and Russian-Uyghur). Experimental results on loanword identification show that our method outperforms other baseline models significantly. Neural machine translation models integrating results of loanword identification experiments achieve the best results on OOV translation(with 0.5-0.9 BLEU improvements)
Chenggang Mi 0002, Yating Yang, Lei Wang 0065, Xi Zhou 0007, Tonghai Jiang
COLING3
2018 Improved Spoken Uyghur Segmentation for Neural Machine Translation
abstract
To increase vocabulary overlap in spoken Uyghur neural machine translation (NMT), we propose a novel method to enhance the common used subword units based segmentation method. In particular, we apply a log-linear model as the main framework and integrate several features such as subword, morphological information, bilingual word alignment and monolingual language model into it. Experimental results show that spoken Uyghur segmentation with our proposed method improves the performance of the spoken Uyghur-Chinese NMT significantly (yield up to 1.52 BLEU improvements).
Chenggang Mi 0002, Yating Yang, Xi Zhou 0007, Lei Wang 0065, Tonghai Jiang
ICTAI4
2018 A Neural Network Based Model for Loanword Identification in Uyghur
Chenggang Mi 0002, Yating Yang, Lei Wang 0065, Xi Zhou 0007, Tonghai Jiang
LREC3
2017 Learning Bilingual Lexicon for Low-Resource Language Pairs
ShaoLin Zhu, Xiao Li 0007, Yating Yang, Lei Wang 0065, Chenggang Mi 0002
NLPCC4
2016 Recurrent Neural Network Based Loanwords Identification in Uyghur
Chenggang Mi 0002, Yating Yang, Xi Zhou 0007, Lei Wang 0065, Xiao Li 0007, Tonghai Jiang
PACLIC4
2015 Optimized Uyghur Segmentation for Statistical Machine Translation
Chenggang Mi 0002, Yating Yang, Rui Dong 0002, Xi Zhou 0007, Lei Wang 0065, Xiao Li 0007, Tonghai Jiang, Osman Turghun
NLDB5
2014 Detection of Loan Words in Uyghur Texts
Chenggang Mi 0002, Yating Yang, Lei Wang 0065, Xiao Li 0007, Kamali Dalielihan
NLPCC3