Chenggang Mi 0001

dblp:151/1355-1 · also Cheng-Gang Mi 0001 · DBLP profile ↗
← Back
17ranked-venue papers
10as first author
16since 2021 · last 2026
0000-0002-6903-8774ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 9 first-author · 14 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 KA-DDI: A knowledge-adaptive contrastive learning framework for drug-drug interaction prediction
Yu Li 0030, Jia-Ming Liu, Yue-Chao Li, Hai-Ru You, Zhu-Hong You, Chenggang Mi 0001
Expert Syst. Appl.7
2026 Large language model enhanced image caption for Chinese cultural relics
Chenggang Mi 0001, Shaoliang Xie
Expert Syst. Appl.1
2026 MRGCDDI: Multi-Relation Graph Contrastive Learning Without Data Augmentation for Drug-Drug Interaction Events Prediction
abstract
Predicting drug-drug interactions (DDIs) is a significant concern in the field of deep learning. It can effectively reduce potential adverse consequences and improve therapeutic safety. Graph neural network (GNN)-based models have made satisfactory progress in DDI event prediction. However, most existing models overlook crucial drug structure and interaction information, which is necessary for accurate DDI event prediction. To tackle this issue, we introduce a new method called MRGCDDI. This approach employs contrastive learning, but unlike conventional methods, it does not require data augmentation, thereby avoiding additional noise. MRGCDDI maintains the semantics of the graphical data during encoder perturbation through a simple yet effective contrastive learning approach, without the need for manual trial and error, tedious searching, or expensive domain knowledge to select enhancements. The approach presented in this study effectively integrates drug features extracted from drug molecular graphs and information from multi-relational drug-drug interaction (DDI) networks. Extensive experimental results demonstrate that MRGCDDI outperforms state-of-the-art methods on both datasets. Specifically, on Deng's dataset, MRGCDDI achieves an average increase of 4.33% in accuracy, 11.57% in Macro-F1, 10.97% in Macro-Recall, and 10.64% in Macro-Precision. Similarly, on Ryu's dataset, the model shows improvements with an average increase of 2.42% in accuracy, 3.86% in Macro-F1, 3.49% in Macro-Recall, and 2.75% in Macro-Precision.
Yu Li 0030, Lin-Xuan Hou, Zhu-Hong You, Chenggang Mi 0001
IEEE J. Biomed. Health Informatics5
2025 Multi-source knowledge fusion for multilingual loanword identification
Chenggang Mi 0001, Shaolin Zhu
Expert Syst. Appl.1
2025 Improving Ancient Chinese Word Segmentation With Knowledge-Enhanced Prompting for Large Language Models
abstract
This paper introduces a cost‐effective prompt optimization strategy for ancient Chinese word segmentation using large language models, aiming to mitigate the substantial computational resources and training expenses of fine‐tuning. We developed two knowledge‐enhanced frameworks, a General Knowledge Prompt framework and a Domain‐Specific Knowledge Prompt framework, and evaluated their effectiveness across various ancient Chinese corpora using seven mainstream LLMs, including ERNIE Bot, Qwen, SparkDesk, DeepSeek, ChatGPT, Gemini, and Copilot. Our findings confirm that both prompt frameworks enhance the segmentation capability of LLMs to varying extents, with the Domain‐Specific Knowledge Prompt framework yielding the most significant improvements. Notably, the DeepSeek model achieves 94.01% F 1 score (94.24% precision, 93.79% recall) on the test set, while the Qwen model demonstrates a remarkable 15.73% increase in the F 1 score with the Domain‐Specific Knowledge Prompt framework. Our ablation studies indicate that the entries Rules and Examples are the most crucial to the success of prompt frameworks, effectively addressing the challenges of rule inconsistency and insufficient annotated data.
Meng-Tian Tang, Chenggang Mi 0001
Int. J. Intell. Syst.2
2025 Bilingual and monolingual features enhanced morphological segmentation for spoken language neural machine translation
Chenggang Mi 0001, Shaoliang Xie
Multim. Tools Appl.1
2025 Loanword Identification in Social Media Texts with Extended Code-Switching Datasets
abstract
As a new data augmentation method for cross-lingual natural language processing task, loanword identification has attracted more and more attentions in recent years. However, previous studies on this topic mainly focus on loanwords that were borrowed through long time language contact. In recent years, many new loanwords appear in social media texts such as X (formerly known as Twitter), Weixin, Weibo et al. Due to that most of these new loanwords do not follow the rules of lexical borrowing, it is very difficult to achieve promising performance by the existing loanword identification methods. In this study, we propose a novel loanword identification method in social media texts based on a BERT-BLSTM-CRF model with extended code-switching data. To allievate the data sparsity exising in model training, three code-switching data generation strategies are introduced, such as high frequency phrase replacement, machine translation generation, and multi-criteria searching. We also describe a multilingual loanword identification model, which can identify loanwords from different donor languages with one model. Experimental results on several datasets show that our model outperforms other baseline systems significantly.
Chenggang Mi 0001, Shaoliang Xie, Yu Li 0030, Zhenghan He
ACM Trans. Asian Low Resour. Lang. Inf. Process.1
2024 Attention-Based Learning for Predicting Drug-Drug Interactions in Knowledge Graph Embedding Based on Multisource Fusion Information
abstract
Drug combinations can reduce drug resistance and side effects and enable the improvement of disease treatment efficacy. Therefore, how to effectively identify drug-drug interactions (DDIs) is a challenging problem. Currently, there exist several approaches that leverage advanced representation learning and graph-based techniques for DDIs prediction. While these methods have demonstrated promising results, a limited number of approaches effectively utilize the potential of knowledge graphs (KGs), which provide information on drug attributes and multirelation among entities. In this work, we introduce a novel attention-based KGs representation learning framework. To encode drug SMILES sequence, a pretrained model is used, while molecular structure information is mapped as the initialization of nodes within the KG using a message-passing neural network. Additionally, the knowledge-aware graph attention network is employed to capture the drug and its topological neighbor representation in the KG representation module. To prevent the oversmoothing problem, the residual layer is used in the DDI prediction module. Comprehensive experiments on several datasets have demonstrated that the proposed method outperforms the state-of-the-art algorithms on the DDI prediction task across a range of evaluation metrics. It achieves an accuracy of 0.924 and an AUC of 0.9705 on the KEGG dataset and attains an ACC of 0.9777 and an AUC of 0.9959 on the OGB-biokg dataset. These experimental findings affirm that our approach is a dependable model for predicting the association of drugs.
Yu Li 0030, Zhu-Hong You, Shu-Min Wang, Chenggang Mi 0001, Meineng Wang
Int. J. Intell. Syst.4
2024 Language relatedness evaluation for multilingual neural machine translation
Chenggang Mi 0001, Shaoliang Xie
Neurocomputing1
2024 Multi-granularity Knowledge Sharing in Low-resource Neural Machine Translation
abstract
As the rapid development of deep learning methods, neural machine translation (NMT) has attracted more and more attention in recent years. However, lack of bilingual resources decreases the performance of the low-resource NMT model seriously. To overcome this problem, several studies put their efforts on knowledge transfer from high-resource language pairs to low-resource language pairs. However, these methods usually focus on one single granularity of language and the parameter sharing among different granularities in NMT is not well studied. In this article, we propose to improve the parameter sharing in low-resource NMT by introducing multi-granularity knowledge such as word, phrase and sentence. This knowledge can be monolingual and bilingual. We build the knowledge sharing model for low-resource NMT based on a multi-task learning framework, three auxiliary tasks such as syntax parsing, cross-lingual named entity recognition, and natural language generation are selected for the low-resource NMT. Experimental results show that the proposed method consistently outperforms six strong baseline systems on several low-resource language pairs.
Chenggang Mi 0001, Shaoliang Xie
ACM Trans. Asian Low Resour. Lang. Inf. Process.1
2023 Loanword identification based on web resources: A case study on wikipedia
Chenggang Mi 0001
Comput. Speech Lang.1
2023 Improving the Robustness of Loanword Identification in Social Media Texts
abstract
As a potential bilingual resource, loanwords play a very important role in many natural language processing tasks. If loanwords in a low-resource language can be identified effectively, the generated donor-receipt word pairs will benefit many cross-lingual natural language processing tasks. However, most studies on loanword identification mainly focus on formal texts such as news and government documents. Loanword identification in social media texts is still an under-studied field. Since it faces many challenges and can be widely used in several downstream tasks, more efforts should be put on loanword identification in social media texts. In this study, we present a multi-task learning architecture with deep bi-directional recurrent neural networks for loanword identification in social media texts, where different task supervision can happen at different layers. The multi-task neural network architecture learns higher-order feature representations from word and character sequences along with basic spell error checking, part-of-speech tagging, and named entity recognition information. Experimental results on Uyghur loanword identification in social media texts in five donor languages (Chinese, Arabic, Russian, Turkish, and Farsi) show that our method achieves the best performance compared with several strong baseline systems. We also combine the loanword detection results into the training data of neural machine translation for low-resource language pairs. Experiments show that models trained on the extended datasets achieve significant improvements compared with the baseline models in all language pairs.
Chenggang Mi 0001
ACM Trans. Asian Low Resour. Lang. Inf. Process.1
2023 Unsupervised Parallel Sentences of Machine Translation for Asian Language Pairs
abstract
Parallel sentence pairs play a very important role in many natural language processing tasks, especially cross-lingual tasks such as machine translation. So far, many Asian language pairs lack bilingual parallel sentences. As collecting bilingual parallel data is very time-consuming and difficult, it is very important for many low-resource Asian language pairs. While existing methods have shown encouraging results, they rely on bilingual data seriously or have some drawbacks in an unsupervised situation. To address these issues, we propose a new unsupervised similarity calculation and dynamic selection metric to obtain parallel sentence pairs in an unsupervised situation. First, our method maps bilingual word embedding by postdoc adversarial training, which rotates the source space to match the target without parallel data. Then, we introduce a new cross-domain similarity adaption to obtain parallel sentence pairs. Experimental results on real-world datasets show that our model can obtain better accuracy and recall on mining parallel sentence pairs. We also show that the extracted bilingual sentence corpora can significantly improve the performance of neural machine translation.
Shaolin Zhu, Chenggang Mi 0001
ACM Trans. Asian Low Resour. Lang. Inf. Process.2
2022 Improving data augmentation for low resource speech-to-text translation with diverse paraphrasing
Chenggang Mi 0001, Lei Xie 0001, Yanning Zhang 0001
Neural Networks1
2021 Inducing Bilingual Word Representations for Non-isomorphic Spaces by an Unsupervised Way
ShaoLin Zhu, Chenggang Mi 0001
KSEM2
2021 Improving bilingual word embeddings mapping with monolingual context information
ShaoLin Zhu, Chenggang Mi 0001, FuHua Zhang
Mach. Transl.2
2020 Loanword Identification in Low-Resource Languages with Minimal Supervision
abstract
Bilingual resources play a very important role in many natural language processing tasks, especially the tasks in cross-lingual scenarios. However, it is expensive and time consuming to build such resources. Lexical borrowing happens in almost every language. This inspires us to detect these loanwords effectively, and to use the “loanword (in receipt language)”-“donor word (in donor language)” to extend the bilingual resource for NLP tasks in low-resource languages. In this article, we propose a novel method to identify loanwords in Uyghur. The most important advantage of this method is that the model only relies on large amount of monolingual corpora and only a small scale of annotated data. Our loanword identification model includes two parts: loanword candidate generation and loanword prediction. In the first part, we use two large-scale monolingual corpora and a small bilingual dictionary to train a cross-lingual embedding model. Since semantic unrelated words often cannot be treated as loanword pairs, a loanword candidate list will be generated according to this model and a word list in Uyghur. In the second part, we predict from the preceding candidates based on a log-linear model that integrates several features such as pronunciation similarity, part-of-speech tags, and hybrid language modeling. To evaluate the effectiveness of our proposed method, we conduct two types of experiments: loanword identification and OOV translation. Experimental results showed that (1) our proposed method achieved significant F1 improvements compared to other models in all four loanword identification tasks in Uyghur, and (2) after extending the existing translation models with loanword identification results, OOV rates in several language pairs reduced significantly and the translation performance improved.
Chenggang Mi 0001, Lei Xie 0001, Yanning Zhang 0001
ACM Trans. Asian Low Resour. Lang. Inf. Process.1