Xi Zhou 0007

dblp:42/5705-7 · DBLP profile ↗
← Back
31ranked-venue papers
0as first author
23since 2021 · last 2026
0000-0003-2867-1765ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 11 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021
YearPublicationVenuePosition
2026 M3UCD: A Multi-task Multimodal Metaphor Understanding Challenge Dataset for LLMs
abstract
Understanding multimodal metaphors represents a crucial pathway for machines to comprehend human cognition. However, current research remains constrained by superficial dataset annotations, insufficient systematic evaluation of large language models, and fragmented task frameworks. To bridge these gaps, the paper proposes a systematic solution featuring: (I) We present the largest fine-grained Multi-task Multimodal Metaphor Understanding Challenge Dataset (M3UCD) built via multi-perspective collaborative annotation. It contains 15,345 samples, each annotated with 12 manual attribute labels. (II) Systematic benchmarking of LLMs' capacity boundaries in metaphor understanding. Evaluation results reveal the persistent challenges LLMs face in this domain while validating M3UCD's effectiveness and potential. (III) A concise and unified multi-task baseline framework was developed and demonstrated its effectiveness in enhancing the metaphor understanding capabilities of MLLMs.
Tianlong Zheng, Yating Yang, Rui Dong 0002, Bo Ma 0004, Lei Wang 0065, Xi Zhou 0007, Siru Miao, Osman Turghun
AAAI6
2026 MMC-Det: Structure-Preserving and Morphology-Aware Detection for Organoids in Bright-Field Microscopy
Xi Zhou 0007, Jun Zhang 0003, Le Tong, Lun Hu, Pengwei Hu 0001
ICIC (30)4
2026 From Self-supervised Pre-training to Confounder-Aware Refinement: A New Approach for miRNA-Drug Association Prediction
Runzhou Tang, Xi Zhou 0007, Jun Zhang 0003, Lun Hu, Pengwei Hu 0001
ICIC (30)3
2026 Ask-to-Retrieve: VQA-Guided Information-Gain Question Selection for Interactive Text-to-Image Retrieval
abstract
Interactive image retrieval iteratively interacts with users to capture retrieval intent and refine textual queries, effectively mitigating the ambiguity and performance degradation inherent in single-round text-to-image retrieval. However, existing LLM-based methods largely rely on image captions or dialogue history to extract features within a limited candidate space. Consequently, they often generate questions that are either irrelevant to the target’s distinctive characteristics or fail to capture fine-grained visual differences. Furthermore, an effective question should efficiently distinguish the target from non-target samples—a critical aspect overlooked by prior works, resulting in suboptimal retrieval efficiency. To bridge this gap, we present A2R (Ask-to-Retrieve), a visually grounded framework designed to maximize information gain. Specifically, a Diversity Candidate Perception (DCP) module constructs a compact visual grid to capture diverse visual semantics. Subsequently, a Discriminative Question Generation (DQG) module observes this grid to propose discriminative questions. Finally, a Maximum Information Gain Decision (IGD) module evaluates these questions and selects the one maximizing entropy reduction. Experiments on VisDial, COCO, and Flickr30k demonstrate that our method achieves superior retrieval accuracy and efficiency compared to strong baselines.
Shuaiwei Xie, Yating Yang, Bo Ma 0004, Xi Zhou 0007, Zhen Wang 0062, Ahtamjan Ahmat
ICMR4
2025 Low-Resource Language Expansion and Translation Capacity Enhancement for LLM: A Study on the Uyghur
abstract
Although large language models have significantly advanced natural language generation, their potential in low-resource machine translation has not yet been fully explored, especially for languages that translation models have not been trained on. In this study, we provide a detailed demonstration of how to efficiently expand low-resource languages for large language models and significantly enhance the model’s translation ability, using Uyghur as an example. The process involves four stages: collecting and pre-processing monolingual data, conducting continuous pre-training with extensive monolingual data, fine-tuning with less parallel corpora using translation supervision, and proposing a direct preference optimization based on translation self-evolution (DPOSE) on this basis. Extensive experiments have shown that our strategy effectively expands the low-resource languages supported by large language models and significantly enhances the model’s translation ability in Uyghur with less parallel data. Our research provides detailed insights for expanding other low-resource languages into large language models.
Kaiwen Lu, Yating Yang, Fengyi Yang, Rui Dong 0002, Bo Ma 0004, Aihetamujiang Aihemaiti, Abibulla Atawulla, Lei Wang 0065, Xi Zhou 0007
COLING9
2025 OpenForecast: A Large-Scale Open-Ended Event Forecasting Dataset
abstract
Complex events generally exhibit unforeseen, multifaceted, and multi-step developments, and cannot be well handled by existing closed-ended event forecasting methods, which are constrained by a limited answer space. In order to accelerate the research on complex event forecasting, we introduce OpenForecast, a large-scale open-ended dataset with two features: (1) OpenForecast defines three open-ended event forecasting tasks, enabling unforeseen, multifaceted, and multi-step forecasting. (2) OpenForecast collects and annotates a large-scale dataset from Wikipedia and news, including 43,419 complex events spanning from 1950 to 2024. Particularly, this annotation can be completed automatically without any manual annotation cost. Meanwhile, we introduce an automatic LLM-based Retrieval-Augmented Evaluation method (LRAE) for complex events, enabling OpenForecast to evaluate the ability of complex event forecasting of large language models. Finally, we conduct comprehensive human evaluations to verify the quality and challenges of OpenForecast, and the consistency between LEAE metric and human evaluation. OpenForecast and related codes will be publicly released.
Zhen Wang 0062, Xi Zhou 0007, Yating Yang, Bo Ma 0004, Lei Wang 0065, Rui Dong 0002, Azmat Anwar
COLING2
2025 Dual-Channel MiRNA Drug Resistance Prediction Model Based on Multimodal Feature Alignment
Runzhou Tang, Zimai Zhang, Jun Zhang 0003, Lun Hu, Xi Zhou 0007, Pengwei Hu 0001
ICIC (26)5
2025 LyRE: Learning Varying Fusion Degrees with Hierarchical Aggregation to Improve Multimodal Misinformation Detection
YiDu Chen, Bo Ma 0004, Yating Yang, Dilxat Abdureyim, Rui Dong 0002, Zhen Wang 0062, Lei Wang 0065, Xi Zhou 0007
NLPCC (2)8
2025 Self-Distillation Across Modalities: Enhancing Cross-Modal Correlation Perception for Multimodal Fake News Detection
YiDu Chen, Bo Ma 0004, Yating Yang, Dilxat Abdureyim, Chirui Zhang, Rui Dong 0002, Lei Wang 0065, Xi Zhou 0007
NLPCC (3)8
2025 ISIPN: Intention-Semantic Incongruity Perception Network for Multimodal Metaphor Detection
Tianlong Zheng, Yating Yang, Rui Dong 0002, Bo Ma 0004, Lei Wang 0065, Xi Zhou 0007
NLPCC (3)6
2025 SGEU: enhancing LLM reasoning via backward exemplar generation and verification
Zhen Wang 0062, Xi Zhou 0007, Yating Yang, Bo Ma 0004, Lei Wang 0065, Rui Dong 0002
Appl. Intell.2
2024 DNMDA: Deep Non-negative Matrix Factorization with Multi-level Integration for MiRNA-Drug Interaction Prediction
abstract
Numerous studies have demonstrated that the interaction between miRNAs and drugs plays a pivotal role in regulating gene expression and cellular function. Therefore, predicting these interactions is crucial for the development of novel drugs and personalized therapies. Existing methods for predicting miRNA-drug interactions often fail to leverage the full spectrum of molecular and biological features and overlook complex high-dimensional patterns. Deep non-negative matrix factorization (DNMF) addresses these limitations by extracting higher-level representations, thereby enhancing prediction accuracy and robustness. Building on this, this paper proposes a model called DNMDA. In this model, we integrate multiple similarity networks for both miRNAs and drugs and then extract their features through three key modules. Moreover, autoencoders are used to combine various feature sets, allowing for the capture of complementary information and enhancing the model’s capacity for making precise and reliable predictions. The resulting features are consolidated into a unified feature vector for each miRNA-drug pair. Ultimately, these feature vectors and their associated labels are provided to the classifier for training. To verify the predictions, a five-fold cross-validation was conducted. The five-fold cross-validation demonstrated a clear advantage in DNMDA’s metrics, underscoring its reliability in predicting potential miRNA-drug interactions. This claim is further supported by the predictive results section in the paper, offering concrete evidence of DNMDA’s efficacy in this field.
Yujie Qi, Zhu-Hong You, Zimai Zhang, Lun Hu, Xi Zhou 0007, Pengwei Hu 0001
BIBM7
2024 Pruning Residual Networks in Multilingual Neural Machine Translation to Improve Zero-Shot Translation
Kaiwen Lu, Yating Yang, Rui Dong 0002, Bo Ma 0004, Lei Wang 0065, Xi Zhou 0007, Ahtamjan Ahmat
NLPCC (3)6
2024 Relational concept enhanced prototypical network for incremental few-shot relation classification
Bo Ma 0004, Lei Wang 0065, Xi Zhou 0007, Zhen Wang 0062, Yating Yang
Knowl. Based Syst.4
2024 Discovering Consensus Regions for Interpretable Identification of RNA N6-Methyladenosine Modification Sites via Graph Contrastive Clustering
abstract
As a pivotal post-transcriptional modification of RNA, N6-methyladenosine (m6A) has a substantial influence on gene expression modulation and cellular fate determination. Although a variety of computational models have been developed to accurately identify potential m6A modification sites, few of them are capable of interpreting the identification process with insights gained from consensus knowledge. To overcome this problem, we propose a deep learning model, namely M6A-DCR, by discovering consensus regions for interpretable identification of m6A modification sites. In particular, M6A-DCR first constructs an instance graph for each RNA sequence by integrating specific positions and types of nucleotides. The discovery of consensus regions is then formulated as a graph clustering problem in light of aggregating all instance graphs. After that, M6A-DCR adopts a motif-aware graph reconstruction optimization process to learn high-quality embeddings of input RNA sequences, thus achieving the identification of m6A modification sites in an end-to-end manner. Experimental results demonstrate the superior performance of M6A-DCR by comparing it with several state-of-the-art identification models. The consideration of consensus regions empowers our model to make interpretable predictions at the motif level. The analysis of cross validation through different species and tissues further verifies the consistency between the identification results of M6A-DCR and the evolutionary relationships among species.
Bo-Wei Zhao, Xiao-Rui Su 0001, Yue Yang 0035, Pengwei Hu 0001, Xi Zhou 0007, Lun Hu
IEEE J. Biomed. Health Informatics6
2023 A Domain-Transfer Meta Task Design Paradigm for Few-Shot Slot Tagging
abstract
Few-shot slot tagging is an important task in dialogue systems and attracts much attention of researchers. Most previous few-shot slot tagging methods utilize meta-learning procedure for training and strive to construct a large number of different meta tasks to simulate the testing situation of insufficient data. However, there is a widespread phenomenon of overlap slot between two domains in slot tagging. Traditional meta tasks ignore this special phenomenon and cannot simulate such realistic few-shot slot tagging scenarios. It violates the basic principle of meta-learning which the meta task is consistent with the real testing task, leading to historical information forgetting problem. In this paper, we introduce a novel domain-transfer meta task design paradigm to tackle this problem. We distribute a basic domain to each target domain based on the coincidence degree of slot labels between these two domains. Unlike classic meta tasks which only rely on small samples of target domain, our meta tasks aim to correctly infer the class of target domain query samples based on both abundant data in basic domain and scarce data in target domain. To accomplish our meta task, we propose a Task Adaptation Network to effectively transfer the historical information from the basic domain to the target domain. We carry out sufficient experiments on the benchmark slot tagging dataset SNIPS and the name entity recognition dataset NER. Results demonstrate that our proposed model outperforms previous methods and achieves the state-of-the-art performance.
Fengyi Yang, Xi Zhou 0007, Yating Yang, Bo Ma 0004, Rui Dong 0002, Abibulla Atawulla
AAAI2
2023 Making the Implicit Explicit: Depression Detection in Web across Posted Texts and Images
abstract
The utilization of web social media for depression detection has been proven effective in recent years since the multimedia signal on web can reflect users’ emotions, feelings, and personality traits in advance. However, most earlier studies simply used users’ submitted words or user profiles to predict depression risk. The implicit information accessible in users’ posted images, which can be effective in depression detection, still remains unexplored. In this paper, an implicit and explicit multi-modal feature fusion (IEMFF) model is proposed for depression detection. We successfully make the implicit information inherent in users’ posted images explicit and further incorporate such explicit features with the textual features directly extracted from user-posted texts. A multi-modal feature fusion approach is applied for depression detection. Extensive experiments have been conducted on public Twitter datasets. Experimental results show that our approach has achieved state-of-the-art performance for depression detection.
Pengwei Hu 0001, Chenhao Lin, Jiajia Li 0004, Feng Tan 0002, Xue Han 0018, Xi Zhou 0007, Lun Hu
BIBM6
2023 A Slot-Shared Span Prediction-Based Neural Network for Multi-Domain Dialogue State Tracking
abstract
There are a large number of candidate values shared among slots in multi-domain dialogue state tracking (DST). The existing span prediction-based DST methods generally adopt slot-independent value extraction architecture, which ignore the value sharing. Besides, the slot-independent design leads to poor scalability. In this paper, we propose a Slot-shared Span Prediction based Network (SSNet) with a general value extraction module for all slots to tackle these problems. To ensure that the value extraction module is able to distinguish different slots, we introduce a Dynamic Fusion Mechanism (DFM) to extract different slot-aware features. DFM plays the routing role, highlighting different dialogue context tokens for different slots. Specifically, DFM firstly calculates similarity matrixes between the dialogue context and different slots, and then determines important dialogue context token with respect to each slot. Experimental results demonstrate that SSNet outperforms the existing start-of-the-art models on both MultiWOZ 2.1 and MultiWOZ 2.2 datasets.
Abibulla Atawulla, Xi Zhou 0007, Yating Yang, Bo Ma 0004, Fengyi Yang
ICASSP2
2023 Artificial intelligence accelerates multi-modal biomedical process: A Survey
Jiajia Li 0004, Xue Han 0018, Feng Tan 0002, Xi Zhou 0007, Lun Hu, Pengwei Hu 0001
Neurocomputing8
2022 A geometric deep learning framework for drug repositioning over heterogeneous information networks
abstract
Drug repositioning (DR) is a promising strategy to discover new indicators of approved drugs with artificial intelligence techniques, thus improving traditional drug discovery and development. However, most of DR computational methods fall short of taking into account the non-Euclidean nature of biomedical network data. To overcome this problem, a deep learning framework, namely DDAGDL, is proposed to predict drug-drug associations (DDAs) by using geometric deep learning (GDL) over heterogeneous information network (HIN). Incorporating complex biological information into the topological structure of HIN, DDAGDL effectively learns the smoothed representations of drugs and diseases with an attention mechanism. Experiment results demonstrate the superior performance of DDAGDL on three real-world datasets under 10-fold cross-validation when compared with state-of-the-art DR methods in terms of several evaluation metrics. Our case studies and molecular docking experiments indicate that DDAGDL is a promising DR tool that gains new insights into exploiting the geometric prior knowledge for improved efficacy.
Bo-Wei Zhao, Xiao-Rui Su 0001, Pengwei Hu 0001, Yu-Peng Ma, Xi Zhou 0007, Lun Hu
Briefings Bioinform.5
2022 Effectively predicting HIV-1 protease cleavage sites by using an ensemble learning approach
abstract
BACKGROUND: The site information of substrates that can be cleaved by human immunodeficiency virus 1 proteases (HIV-1 PRs) is of great significance for designing effective inhibitors against HIV-1 viruses. A variety of machine learning-based algorithms have been developed to predict HIV-1 PR cleavage sites by extracting relevant features from substrate sequences. However, only relying on the sequence information is not sufficient to ensure a promising performance due to the uncertainty in the way of separating the datasets used for training and testing. Moreover, the existence of noisy data, i.e., false positive and false negative cleavage sites, could negatively influence the accuracy performance. RESULTS: In this work, an ensemble learning algorithm for predicting HIV-1 PR cleavage sites, namely EM-HIV, is proposed by training a set of weak learners, i.e., biased support vector machine classifiers, with the asymmetric bagging strategy. By doing so, the impact of data imbalance and noisy data can thus be alleviated. Besides, in order to make full use of substrate sequences, the features used by EM-HIV are collected from three different coding schemes, including amino acid identities, chemical properties and variable-length coevolutionary patterns, for the purpose of constructing more relevant feature vectors of octamers. Experiment results on three independent benchmark datasets demonstrate that EM-HIV outperforms state-of-the-art prediction algorithm in terms of several evaluation metrics. Hence, EM-HIV can be regarded as a useful tool to accurately predict HIV-1 PR cleavage sites.
Lun Hu, Zhenfeng Li, Zehai Tang, Xi Zhou 0007, Pengwei Hu 0001
BMC Bioinform.5
2021 SGANRDA: semi-supervised generative adversarial networks for predicting circRNA-disease associations
abstract
Emerging research shows that circular RNA (circRNA) plays a crucial role in the diagnosis, occurrence and prognosis of complex human diseases. Compared with traditional biological experiments, the computational method of fusing multi-source biological data to identify the association between circRNA and disease can effectively reduce cost and save time. Considering the limitations of existing computational models, we propose a semi-supervised generative adversarial network (GAN) model SGANRDA for predicting circRNA-disease association. This model first fused the natural language features of the circRNA sequence and the features of disease semantics, circRNA and disease Gaussian interaction profile kernel, and then used all circRNA-disease pairs to pre-train the GAN network, and fine-tune the network parameters through labeled samples. Finally, the extreme learning machine classifier is employed to obtain the prediction result. Compared with the previous supervision model, SGANRDA innovatively introduced circRNA sequences and utilized all the information of circRNA-disease pairs during the pre-training process. This step can increase the information content of the feature to some extent and reduce the impact of too few known associations on the model performance. SGANRDA obtained AUC scores of 0.9411 and 0.9223 in leave-one-out cross-validation and 5-fold cross-validation, respectively. Prediction results on the benchmark dataset show that SGANRDA outperforms other existing models. In addition, 25 of the top 30 circRNA-disease pairs with the highest scores of SGANRDA in case studies were verified by recent literature. These experimental results demonstrate that SGANRDA is a useful model to predict the circRNA-disease association and can provide reliable candidates for biological experiments.
Lei Wang 0121, Zhu-Hong You, Xi Zhou 0007
Briefings Bioinform.4
2021 In silico drug repositioning using deep learning and comprehensive similarity measures
abstract
BACKGROUND: Drug repositioning, meanings finding new uses for existing drugs, which can accelerate the processing of new drugs research and development. Various computational methods have been presented to predict novel drug-disease associations for drug repositioning based on similarity measures among drugs and diseases. However, there are some known associations between drugs and diseases that previous studies not utilized. METHODS: In this work, we develop a deep gated recurrent units model to predict potential drug-disease interactions using comprehensive similarity measures and Gaussian interaction profile kernel. More specifically, the similarity measure is used to exploit discriminative feature for drugs based on their chemical fingerprints. Meanwhile, the Gaussian interactions profile kernel is employed to obtain efficient feature of diseases based on known disease-disease associations. Then, a deep gated recurrent units model is developed to predict potential drug-disease interactions. RESULTS: The performance of the proposed model is evaluated on two benchmark datasets under tenfold cross-validation. And to further verify the predictive ability, case studies for predicting new potential indications of drugs were carried out. CONCLUSION: The experimental results proved the proposed model is a useful tool for predicting new indications for drugs or new treatments for diseases, and can accelerate drug repositioning and related drug research and discovery.
Zhu-Hong You, Lei Wang 0065, Xiao-Rui Su 0001, Xi Zhou 0007, Tonghai Jiang
BMC Bioinform.5
2020 A Gaussian Kernel Similarity-Based Linear Optimization Model for Predicting miRNA-lncRNA Interactions
Leon Wong, Zhu-Hong You, Xi Zhou 0007, Mei-Yuan Cao
ICIC (2)4
2020 Prediction of lncRNA-miRNA Interactions via an Embedding Learning Graph Factorize Over Heterogeneous Information Network
Ji-Ren Zhou, Zhu-Hong You, Xi Zhou 0007
ICIC (2)4
2018 Toward Better Loanword Identification in Uyghur Using Cross-lingual Word Embeddings
abstract
To enrich vocabulary of low resource settings, we proposed a novel method which identify loanwords in monolingual corpora. More specifically, we first use cross-lingual word embeddings as the core feature to generate semantically related candidates based on comparable corpora and a small bilingual lexicon; then, a log-linear model which combines several shallow features such as pronunciation similarity and hybrid language model features to predict the final results. In this paper, we use Uyghur as the receipt language and try to detect loanwords in four donor languages: Arabic, Chinese, Persian and Russian. We conduct two groups of experiments to evaluate the effectiveness of our proposed approach: loanword identification and OOV translation in four language pairs and eight translation directions (Uyghur-Arabic, Arabic-Uyghur, Uyghur-Chinese, Chinese-Uyghur, Uyghur-Persian, Persian-Uyghur, Uyghur-Russian, and Russian-Uyghur). Experimental results on loanword identification show that our method outperforms other baseline models significantly. Neural machine translation models integrating results of loanword identification experiments achieve the best results on OOV translation(with 0.5-0.9 BLEU improvements)
Chenggang Mi 0002, Yating Yang, Lei Wang 0065, Xi Zhou 0007, Tonghai Jiang
COLING4
2018 Improved Spoken Uyghur Segmentation for Neural Machine Translation
abstract
To increase vocabulary overlap in spoken Uyghur neural machine translation (NMT), we propose a novel method to enhance the common used subword units based segmentation method. In particular, we apply a log-linear model as the main framework and integrate several features such as subword, morphological information, bilingual word alignment and monolingual language model into it. Experimental results show that spoken Uyghur segmentation with our proposed method improves the performance of the spoken Uyghur-Chinese NMT significantly (yield up to 1.52 BLEU improvements).
Chenggang Mi 0002, Yating Yang, Xi Zhou 0007, Lei Wang 0065, Tonghai Jiang
ICTAI3
2018 A Neural Network Based Model for Loanword Identification in Uyghur
Chenggang Mi 0002, Yating Yang, Lei Wang 0065, Xi Zhou 0007, Tonghai Jiang
LREC4
2016 A Bilingual Discourse Corpus and Its Applications
Yang Liu 0085, Jiajun Zhang 0001, Chengqing Zong, Yating Yang, Xi Zhou 0007
LREC5
2016 Recurrent Neural Network Based Loanwords Identification in Uyghur
Chenggang Mi 0002, Yating Yang, Xi Zhou 0007, Lei Wang 0065, Xiao Li 0007, Tonghai Jiang
PACLIC3
2015 Optimized Uyghur Segmentation for Statistical Machine Translation
Chenggang Mi 0002, Yating Yang, Rui Dong 0002, Xi Zhou 0007, Lei Wang 0065, Xiao Li 0007, Tonghai Jiang, Osman Turghun
NLDB4