EDBT 2026 Demo / reviewers in the wild / expert
Bo Ma 0004
dblp:26/4179-4
· DBLP profile ↗
19ranked-venue papers
1as first author
17since 2021 · last 2026
0000-0003-3082-648XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 1 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | M3UCD: A Multi-task Multimodal Metaphor Understanding Challenge Dataset for LLMsabstractUnderstanding multimodal metaphors represents a crucial pathway for machines to comprehend human cognition. However, current research remains constrained by superficial dataset annotations, insufficient systematic evaluation of large language models, and fragmented task frameworks. To bridge these gaps, the paper proposes a systematic solution featuring: (I) We present the largest fine-grained Multi-task Multimodal Metaphor Understanding Challenge Dataset (M3UCD) built via multi-perspective collaborative annotation. It contains 15,345 samples, each annotated with 12 manual attribute labels. (II) Systematic benchmarking of LLMs' capacity boundaries in metaphor understanding. Evaluation results reveal the persistent challenges LLMs face in this domain while validating M3UCD's effectiveness and potential. (III) A concise and unified multi-task baseline framework was developed and demonstrated its effectiveness in enhancing the metaphor understanding capabilities of MLLMs. Tianlong Zheng, Yating Yang, Rui Dong 0002, Bo Ma 0004, Lei Wang 0065, Xi Zhou 0007, Siru Miao, Osman Turghun |
AAAI | 4 |
| 2026 | Ask-to-Retrieve: VQA-Guided Information-Gain Question Selection for Interactive Text-to-Image RetrievalabstractInteractive image retrieval iteratively interacts with users to capture retrieval intent and refine textual queries, effectively mitigating the ambiguity and performance degradation inherent in single-round text-to-image retrieval. However, existing LLM-based methods largely rely on image captions or dialogue history to extract features within a limited candidate space. Consequently, they often generate questions that are either irrelevant to the target’s distinctive characteristics or fail to capture fine-grained visual differences. Furthermore, an effective question should efficiently distinguish the target from non-target samples—a critical aspect overlooked by prior works, resulting in suboptimal retrieval efficiency. To bridge this gap, we present A2R (Ask-to-Retrieve), a visually grounded framework designed to maximize information gain. Specifically, a Diversity Candidate Perception (DCP) module constructs a compact visual grid to capture diverse visual semantics. Subsequently, a Discriminative Question Generation (DQG) module observes this grid to propose discriminative questions. Finally, a Maximum Information Gain Decision (IGD) module evaluates these questions and selects the one maximizing entropy reduction. Experiments on VisDial, COCO, and Flickr30k demonstrate that our method achieves superior retrieval accuracy and efficiency compared to strong baselines. Shuaiwei Xie, Yating Yang, Bo Ma 0004, Xi Zhou 0007, Zhen Wang 0062, Ahtamjan Ahmat |
ICMR | 3 |
| 2025 | Low-Resource Language Expansion and Translation Capacity Enhancement for LLM: A Study on the UyghurabstractAlthough large language models have significantly advanced natural language generation, their potential in low-resource machine translation has not yet been fully explored, especially for languages that translation models have not been trained on. In this study, we provide a detailed demonstration of how to efficiently expand low-resource languages for large language models and significantly enhance the model’s translation ability, using Uyghur as an example. The process involves four stages: collecting and pre-processing monolingual data, conducting continuous pre-training with extensive monolingual data, fine-tuning with less parallel corpora using translation supervision, and proposing a direct preference optimization based on translation self-evolution (DPOSE) on this basis. Extensive experiments have shown that our strategy effectively expands the low-resource languages supported by large language models and significantly enhances the model’s translation ability in Uyghur with less parallel data. Our research provides detailed insights for expanding other low-resource languages into large language models. Kaiwen Lu, Yating Yang, Fengyi Yang, Rui Dong 0002, Bo Ma 0004, Aihetamujiang Aihemaiti, Abibulla Atawulla, Lei Wang 0065, Xi Zhou 0007 |
COLING | 5 |
| 2025 | OpenForecast: A Large-Scale Open-Ended Event Forecasting DatasetabstractComplex events generally exhibit unforeseen, multifaceted, and multi-step developments, and cannot be well handled by existing closed-ended event forecasting methods, which are constrained by a limited answer space. In order to accelerate the research on complex event forecasting, we introduce OpenForecast, a large-scale open-ended dataset with two features: (1) OpenForecast defines three open-ended event forecasting tasks, enabling unforeseen, multifaceted, and multi-step forecasting. (2) OpenForecast collects and annotates a large-scale dataset from Wikipedia and news, including 43,419 complex events spanning from 1950 to 2024. Particularly, this annotation can be completed automatically without any manual annotation cost. Meanwhile, we introduce an automatic LLM-based Retrieval-Augmented Evaluation method (LRAE) for complex events, enabling OpenForecast to evaluate the ability of complex event forecasting of large language models. Finally, we conduct comprehensive human evaluations to verify the quality and challenges of OpenForecast, and the consistency between LEAE metric and human evaluation. OpenForecast and related codes will be publicly released. Zhen Wang 0062, Xi Zhou 0007, Yating Yang, Bo Ma 0004, Lei Wang 0065, Rui Dong 0002, Azmat Anwar |
COLING | 4 |
| 2025 | Mining the Past with Dual Criteria: Integrating Three types of Historical Information for Context-aware Event ForecastingabstractEvent forecasting requires modeling historical event data to predict future events, and achieving accurate predictions depends on effectively capturing the relevant historical information that aids forecasting.Most existing methods focus on entities and structural dependencies to capture historical clues but often overlook implicitly relevant information.This limitation arises from overlooking event semantics and deeper factual associations that are not explicitly connected in the graph structure but are nonetheless critical for accurate forecasting.To address this, we propose a dual-criteria constraint strategy that leverages event semantics for relevance modeling and incorporates a self-supervised semantic filter based on factual event associations to capture implicitly relevant historical information.Building on this strategy, our method, termed ITHI (Integrating Three types of Historical Information), combines sequential event information, periodically repeated event information, and relevant historical information to achieve context-aware event forecasting.We evaluated the proposed ITHI method on three public benchmark datasets, achieving state-of-the-art performance and significantly outperforming existing approaches.Additionally, we validated its effectiveness on two structured temporal knowledge graph forecasting dataset 1 . Lei Wang 0065, Yating Yang, Bo Ma 0004, Rui Dong 0002, Fengyi Yang, Ahtamjan Ahmat, Kaiwen Lu |
EMNLP | 4 |
| 2025 | Bidirectional Feature Fusion and Adaptive Decision Network for Multimodal Fake News DetectionabstractThe proliferation of multimedia content on social media platforms has rendered the detection of fake news an increasingly critical challenge. Despite the progress made in multimodal fake news detection methods, significant challenges remain: information loss during the cross-modal feature fusion process and restricted analysis due to single-perspective decision mechanisms. We propose the Bidirectional Feature Fusion and Adaptive Decision Network (BFF-ADN), which introduces two key innovations: a bidirectional token-level feature fusion mechanism that captures key correlations between visual and textual modalities while reducing information loss, and an Adaptive Multi-perspective Decision Integration Network (AMDI-Net) that integrates multiple specialized detectors for comprehensive analysis. Experiments on three real-world datasets (Weibo, GossipCop, and PolitiFact) demonstrate that BFF-ADN achieves state-of-the-art performance. Ablation studies and visualizations confirm the effectiveness of our proposed components, particularly highlighting the advantages of the bidirectional feature fusion mechanism and Adaptive Multi-perspective Decision Integration Network. Dilxat Abdureyim, Bo Ma 0004, Yating Yang, Rui Dong 0002, YiDu Chen, Azmat Anwar, Lei Wang 0065 |
ICME | 2 |
| 2025 | LyRE: Learning Varying Fusion Degrees with Hierarchical Aggregation to Improve Multimodal Misinformation Detection
YiDu Chen, Bo Ma 0004, Yating Yang, Dilxat Abdureyim, Rui Dong 0002, Zhen Wang 0062, Lei Wang 0065, Xi Zhou 0007 |
NLPCC (2) | 2 |
| 2025 | Self-Distillation Across Modalities: Enhancing Cross-Modal Correlation Perception for Multimodal Fake News Detection
YiDu Chen, Bo Ma 0004, Yating Yang, Dilxat Abdureyim, Chirui Zhang, Rui Dong 0002, Lei Wang 0065, Xi Zhou 0007 |
NLPCC (3) | 2 |
| 2025 | ISIPN: Intention-Semantic Incongruity Perception Network for Multimodal Metaphor Detection
Tianlong Zheng, Yating Yang, Rui Dong 0002, Bo Ma 0004, Lei Wang 0065, Xi Zhou 0007 |
NLPCC (3) | 4 |
| 2025 | M2-VLP: Enhancing Multilingual Vision-Language Pre-Training via Multi-Grained AlignmentabstractRecently, multilingual Vision-Language Pre-training (mVLP) has shown remarkable progress in learning joint representations across different modalities and languages. However, most existing methods learn semantic alignment at a coarse-grained level and fail to capture fine-grained correlations between different languages and modalities. To address this, we propose a Multi-grained Multilingual Vision-Language Pre-training (M2-VLP) model, which aims to learn cross-lingual cross-modal alignment at different semantic granular levels. In cross-lingual interaction, the model learns the global alignment of parallel sentence pairs and the word-level correlations. In cross-modal interaction, the model aligns images with captions and image regions with corresponding words. To integrate the cross-lingual and cross-modal alignment above, we propose a unified multi-grained contrastive learning paradigm. Under zero-shot cross-lingual and fine-tuned multilingual settings, extensive experiments on vision-language downstream tasks across twenty languages demonstrate the effectiveness of M2-VLP over competitive contrastive models. Code and models are available at https://github.com/ahtamjan/M2-VLP. Ahtamjan Ahmat, Lei Wang 0065, Yating Yang, Bo Ma 0004, Rui Dong 0002, Kaiwen Lu |
WWW | 4 |
| 2025 | Cross-lingual prompting method with semantic-based answer space clustering
Ahtamjan Ahmat, Yating Yang, Bo Ma 0004, Rui Dong 0002, Lei Wang 0065 |
Appl. Intell. | 3 |
| 2025 | SGEU: enhancing LLM reasoning via backward exemplar generation and verification
Zhen Wang 0062, Xi Zhou 0007, Yating Yang, Bo Ma 0004, Lei Wang 0065, Rui Dong 0002 |
Appl. Intell. | 4 |
| 2024 | Pruning Residual Networks in Multilingual Neural Machine Translation to Improve Zero-Shot Translation
Kaiwen Lu, Yating Yang, Rui Dong 0002, Bo Ma 0004, Lei Wang 0065, Xi Zhou 0007, Ahtamjan Ahmat |
NLPCC (3) | 4 |
| 2024 | Relational concept enhanced prototypical network for incremental few-shot relation classification
Bo Ma 0004, Lei Wang 0065, Xi Zhou 0007, Zhen Wang 0062, Yating Yang |
Knowl. Based Syst. | 2 |
| 2023 | A Domain-Transfer Meta Task Design Paradigm for Few-Shot Slot TaggingabstractFew-shot slot tagging is an important task in dialogue systems and attracts much attention of researchers. Most previous few-shot slot tagging methods utilize meta-learning procedure for training and strive to construct a large number of different meta tasks to simulate the testing situation of insufficient data. However, there is a widespread phenomenon of overlap slot between two domains in slot tagging. Traditional meta tasks ignore this special phenomenon and cannot simulate such realistic few-shot slot tagging scenarios. It violates the basic principle of meta-learning which the meta task is consistent with the real testing task, leading to historical information forgetting problem. In this paper, we introduce a novel domain-transfer meta task design paradigm to tackle this problem. We distribute a basic domain to each target domain based on the coincidence degree of slot labels between these two domains. Unlike classic meta tasks which only rely on small samples of target domain, our meta tasks aim to correctly infer the class of target domain query samples based on both abundant data in basic domain and scarce data in target domain. To accomplish our meta task, we propose a Task Adaptation Network to effectively transfer the historical information from the basic domain to the target domain. We carry out sufficient experiments on the benchmark slot tagging dataset SNIPS and the name entity recognition dataset NER. Results demonstrate that our proposed model outperforms previous methods and achieves the state-of-the-art performance. Fengyi Yang, Xi Zhou 0007, Yating Yang, Bo Ma 0004, Rui Dong 0002, Abibulla Atawulla |
AAAI | 4 |
| 2023 | A Slot-Shared Span Prediction-Based Neural Network for Multi-Domain Dialogue State TrackingabstractThere are a large number of candidate values shared among slots in multi-domain dialogue state tracking (DST). The existing span prediction-based DST methods generally adopt slot-independent value extraction architecture, which ignore the value sharing. Besides, the slot-independent design leads to poor scalability. In this paper, we propose a Slot-shared Span Prediction based Network (SSNet) with a general value extraction module for all slots to tackle these problems. To ensure that the value extraction module is able to distinguish different slots, we introduce a Dynamic Fusion Mechanism (DFM) to extract different slot-aware features. DFM plays the routing role, highlighting different dialogue context tokens for different slots. Specifically, DFM firstly calculates similarity matrixes between the dialogue context and different slots, and then determines important dialogue context token with respect to each slot. Experimental results demonstrate that SSNet outperforms the existing start-of-the-art models on both MultiWOZ 2.1 and MultiWOZ 2.2 datasets. Abibulla Atawulla, Xi Zhou 0007, Yating Yang, Bo Ma 0004, Fengyi Yang |
ICASSP | 4 |
| 2023 | WAD-X: Improving Zero-shot Cross-lingual Transfer via Adapter-based Word AlignmentabstractMultilingual pre-trained language models (mPLMs) have achieved remarkable performance on zero-shot cross-lingual transfer learning. However, most mPLMs implicitly encourage cross-lingual alignment in pre-training stage, making it hard to capture accurate word alignment across languages. In this paper, we propose Word-align ADapters for Cross-lingual transfer (WAD-X) to explicitly align word representations of mPLMs via language-specific subspace. Taking a mPLM as the backbone model, WAD-X constructs subspace for each source-target language pair via adapters. The adapters use statistical alignment as the prior knowledge to guide word-level aligning in the corresponding bilingual semantic subspace. We evaluate our model across a set of target languages on three zero-shot cross-lingual transfer tasks: part-of-speech tagging (POS), dependency parsing (DP), and sentiment analysis (SA). Experimental results demonstrate that our proposed model improves zero-shot cross-lingual transfer on three benchmarks, with improvements of 2.19, 2.50, and 1.61 points in POS, DP, and SA tasks over strong baselines. Ahtamjan Ahmat, Yating Yang, Bo Ma 0004, Rui Dong 0002, Kaiwen Lu, Lei Wang 0065 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 3 |
| 2019 | A Multi-View Spatial-Temporal Network for Vehicle Refueling Demand Inference
Bo Ma 0004, Yating Yang |
KSEM (2) | 1 |
| 2018 | HEMD: a highly efficient random forest-based malware detection framework for Android
Tonghai Jiang, Bo Ma 0004, Zhu-Hong You, Wei-Lei Shi |
Neural Comput. Appl. | 3 |