Kuo Yang 0001

dblp:55/10445-1 · DBLP profile ↗
← Back
28ranked-venue papers
4as first author
22since 2021 · last 2026
0000-0003-0736-4512ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 24 · 4 first-author · 19 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Meta Relation-Aware Metric Learning Framework for Few-Shot Uncertain Knowledge Graph Completion
abstract
Uncertainty is a fundamental aspect of human knowledge, widely present in various domains such as medicine, biology, and finance. Uncertain Knowledge Graph Completion (UKGC) tasks encompass both link prediction and confidence prediction. However, the inaccessibility of high-quality knowledge results in sparse knowledge graphs, where a large number of entities exhibit relatively sparse neighborhoods. This sparsity severely hinders the performance of few-shot UKGC models, which generally rely on neighborhood aggregation to enhance entity representations. In this study, we propose a Meta relation-aware Metric learning framework designed for few-shot UKGC, named MMUC. To mitigate neighbor sparsity, MMUC introduces a relation-centric paradigm that decouples representation enhancement from background graphs, integrating an uncertainty-driven similarity measure through a learnable confidence learner to achieve robust, non-linear fusion for joint link and confidence prediction. Extensive evaluations across multiple benchmarks demonstrate MMUC's superior performance through rigorous link and confidence prediction, ablation studies, and sparsity analysis, highlighting its particular resilience to neighbor sparsity in few-shot uncertain KGs. The source code can be accessed inhttps://github.com/sienna-wxy/MMUC.
Xinyan Wang 0002, Kuo Yang 0001, Xianan Li, Zixin Shu, Xuezhong Zhou
IEEE Trans. Knowl. Data Eng.2
2025 KGCRR: An Effective Metric-Driven Knowledge Graph Completion Framework by Designing a Novel Upper Bound Function with Adaptive Approximation to Reciprocal Rank
abstract
Knowledge Graph Embedding (KGE) methods have achieved great success in predicting missing links in knowledge graphs, a task also known as Knowledge Graph Completion (KGC). Under this task, the Reciprocal Rank (RR) of ground-truth items serve as a key indicator for evaluating the method’s performance. However, most existing studies have overlooked the inconsistency between the ranking metric, RR, and the optimization objective functions, resulting in sub-optimal KGC performance. To address this issue, we propose a KGC framework called KGCRR by designing a novel upper bound function named CRR. By introducing the parameter-pressure ρ to shift the sigmoid function, CRR achieves a better approximation to RR compared with existing objective functions. We theoretically proved that by adjusting ρ, CRR can achieve a more effective approximation to RR. By narrowing the discrepancy with RR and alleviating the gradient vanishing issue associated with the direct optimization of RR loss, CRR demonstrates an advantage in optimizing RR. CRR serves as a plug-and-play objective, capable of seamless integration into various KGE methods. Through extensive experiments conducted on FB15k-237 and WN18RR datasets, we have obtained promising results, with an average improvement of 19.06% in MRR, indicating that CRR significantly enhances the performance of existing methods.
Kuo Yang 0001, Xiangkui Lu, Xuezhong Zhou
AAAI2
2025 PresRecSKG: Enhancing Herbal Prescription Recommendation via Knowledge Graph-Guided Data Augmentation
abstract
With the rapid advancement of medical artificial intelligence, Traditional Chinese Medicine (TCM) has demonstrated significant potential in intelligent prescription recommendation. However, the scarcity of structured and high-quality clinical data often leads to issues such as data sparsity and terminology inconsistency, which limit the generalization capability and practical performance of existing models. To tackle these challenges, this study proposes PresRecSKG, a knowledge graph-guided data augmentation framework that utilizes a herbsymptom knowledge graph to enhance both multi-label and multi-class prescription recommendation tasks. The core of the framework is the proposed Sampling-by-Knowledge-Graph (SabKG) strategy, which effectively incorporates external knowledge for data synthesis, thereby mitigating data sparsity and improving data quality control. Experimental results show that PresRecSKG achieves the best performance in the multi-label task, with an average improvement of approximately 2.5% across all evaluation metrics. In the multi-class task, the framework elevates performance from 0.94 to over 0.97 across key metrics. These findings indicate that PresRecSKG offers an effective data augmentation paradigm for TCM prescription recommendation, enhancing the accuracy, compatibility, and interpretability of intelligent TCM systems.
Xin Dong 0017, Jiahe Liu, Kuo Yang 0001, Xuezhong Zhou
BIBM3
2025 Cross-Domain Multi-Modal Transfer Learning for Target Identification of TCM's Compounds
Huan Gu, Xunpeng Xiao, Chongyu Wang, Kuo Yang 0001, Xuezhong Zhou
BIBM5
2025 CPOne: Enhancing Prediction of Compound-Protein Interactions Through One-Shot Meta Learning
abstract
Predicting the interactions between compounds and their potential target proteins is crucial in drug discovery. Existing methods often assume that each compound has an adequate number of target proteins available for model training. However, in practice, the number of target proteins associated with compounds is often limited, making it difficult to gather a sufficient number of training examples. This results in learning bias in the model, leading to poor performance on these compounds. Moreover, this issue is evident in widely used datasets, such as GPCR and kinase, where 77.51% and 12.71% of compounds, respectively, have only one target protein available for model training. However, the issue has not been fully explored, presenting a challenge for model development. In this research, we propose CPOne, a framework designed to address the aforementioned issue. CPOne is a meta-learning based approach for compound-protein interaction prediction in one-shot scenario where each compound has only one target protein available for model training. By utilizing a meta compound learner, CPOne extracts the meta representation of compound from each task. Through fast gradient updates, this representation is quickly adapted to generate a compound-specific representation for the current task, thereby improving performance in one-shot scenario. Through comprehensive experiments, we empirically validate the superiority of CPOne, which demonstrates a promising performance improvement over established methods.
Kuo Yang 0001, Chongyu Wang, Xinyan Wang 0002, Hanyu Yuan, Jian Yu 0001, Xuezhong Zhou
IEEE Trans. Comput. Biol. Bioinform.2
2024 LMDTA: Molecular Pre-trained and Interaction Fine-tuned Attention Neural Network for Drug-Target Affinity Prediction
abstract
Accurately predicting drug-target binding affinity is crucial for advancing drug discovery. Recent molecular pre-training models and biological large models provide general molecular representations, but effectively leveraging these features for specific tasks like Drug-Target Affinity (DTA) prediction remains challenging. To address this, we propose LMDTA, a novel attention-based neural network combining molecular pre-training with interaction fine-tuning for affinity prediction. LMDTA learns complex drug and protein representations through pre-trained models, while fine-tuning an interaction-specific model for the DTA task. The two-side attention mechanism integrates pre-trained and fine-tuned features, capturing key drug-target interactions. Experiments on benchmark datasets show LMDTA achieves state-of-the-art performance, with ablation studies validating the model’s design.
Minjie Hu, Kuo Yang 0001, Xuezhong Zhou
BIBM2
2024 PresRecCD: A Novel Herbal Prescription Recommendation Framework with Cross-Domain Learning and Neural Collaborative Filtering
abstract
Herbal prescriptions hold significant importance in Traditional Chinese Medicine (TCM) diagnosis and treatment, embodying millennia of clinical case summaries and wisdom. Despite numerous proposed methods for herbal prescription recommendation (HPR), significant challenges persist due to the lack of comprehensive clinical data, particularly regarding the relationships between symptoms and herbs. This scarcity poses considerable hurdles for effective HPR modeling. In this study, we introduced a novel herbal prescription recommendation framework with cross-domain learning and neural collaborative filtering (termed PresRecCD). The cross-domain learning mechanism is introduced to learn the noise-reduced cross-domain features of herbs and symptoms in the unified space, alleviating the sparsity of data, and neural collaborative filtering is utilized to carry out prescription recommendations. Comprehensive experiments demonstrate the superiority of the proposed PresRecCD model over the SOTA model. This study contributes to enhancing the performance of the HPR model, ultimately benefiting the efficiency and precision of clinical treatment.
Wansong Zhang, Xin Dong 0017, Kuo Yang 0001, Rouye Huang, Runshun Zhang, Xuezhong Zhou
BIBM3
2024 Benchmarking Biomedical Relation Knowledge in Large Language Models
Kuo Yang 0001, Chenqian Zhao, Haixu Li, Xin Dong 0016, Xuezhong Zhou
ISBRA (2)2
2024 HTINet2: herb-target prediction via knowledge graph embedding and residual-like graph neural network
abstract
Target identification is one of the crucial tasks in drug research and development, as it aids in uncovering the action mechanism of herbs/drugs and discovering new therapeutic targets. Although multiple algorithms of herb target prediction have been proposed, due to the incompleteness of clinical knowledge and the limitation of unsupervised models, accurate identification for herb targets still faces huge challenges of data and models. To address this, we proposed a deep learning-based target prediction framework termed HTINet2, which designed three key modules, namely, traditional Chinese medicine (TCM) and clinical knowledge graph embedding, residual graph representation learning, and supervised target prediction. In the first module, we constructed a large-scale knowledge graph that covers the TCM properties and clinical treatment knowledge of herbs, and designed a component of deep knowledge embedding to learn the deep knowledge embedding of herbs and targets. In the remaining two modules, we designed a residual-like graph convolution network to capture the deep interactions among herbs and targets, and a Bayesian personalized ranking loss to conduct supervised training and target prediction. Finally, we designed comprehensive experiments, of which comparison with baselines indicated the excellent performance of HTINet2 (HR@10 increased by 122.7% and NDCG@10 by 35.7%), ablation experiments illustrated the positive effect of our designed modules of HTINet2, and case study demonstrated the reliability of the predicted targets of Artemisia annua and Coptis chinensis based on the knowledge base, literature, and molecular docking.
Pengbo Duan, Kuo Yang 0001, Xin Su 0011, Shuyue Fan, Xin Dong 0017, Xianan Li, Xiaoyan Xing, Jian Yu 0001, Xuezhong Zhou
Briefings Bioinform.2
2024 KDGene: knowledge graph completion for disease gene prediction using interactional tensor decomposition
abstract
The accurate identification of disease-associated genes is crucial for understanding the molecular mechanisms underlying various diseases. Most current methods focus on constructing biological networks and utilizing machine learning, particularly deep learning, to identify disease genes. However, these methods overlook complex relations among entities in biological knowledge graphs. Such information has been successfully applied in other areas of life science research, demonstrating their effectiveness. Knowledge graph embedding methods can learn the semantic information of different relations within the knowledge graphs. Nonetheless, the performance of existing representation learning techniques, when applied to domain-specific biological data, remains suboptimal. To solve these problems, we construct a biological knowledge graph centered on diseases and genes, and develop an end-to-end knowledge graph completion framework for disease gene prediction using interactional tensor decomposition named KDGene. KDGene incorporates an interaction module that bridges entity and relation embeddings within tensor decomposition, aiming to improve the representation of semantically similar concepts in specific domains and enhance the ability to accurately predict disease genes. Experimental results show that KDGene significantly outperforms state-of-the-art algorithms, whether existing disease gene prediction methods or knowledge graph embedding methods for general domains. Moreover, the comprehensive biological analysis of the predicted results further validates KDGene's capability to accurately identify new candidate genes. This work proposes a scalable knowledge graph completion framework to identify disease candidate genes, from which the results are promising to provide valuable references for further wet experiments. Data and source codes are available at https://github.com/2020MEAI/KDGene.
Xinyan Wang 0002, Kuo Yang 0001, Ting Jia, Fanghui Gu, Chongyu Wang, Zixin Shu, Jianan Xia, Xuezhong Zhou
Briefings Bioinform.2
2024 DrugRepPT: a deep pretraining and fine-tuning framework for drug repositioning based on drug's expression perturbation and treatment effectiveness
abstract
MOTIVATION: Drug repositioning (DR), identifying novel indications for approved drugs, is a cost-effective strategy in drug discovery. Despite numerous proposed DR models, integrating network-based features, differential gene expression, and chemical structures for high-performance DR remains challenging. RESULTS: We propose a comprehensive deep pretraining and fine-tuning framework for DR, termed DrugRepPT. Initially, we design a graph pretraining module employing model-augmented contrastive learning on a vast drug-disease heterogeneous graph to capture nuanced interactions and expression perturbations after intervention. Subsequently, we introduce a fine-tuning module leveraging a graph residual-like convolution network to elucidate intricate interactions between diseases and drugs. Moreover, a Bayesian multiloss approach is introduced to balance the existence and effectiveness of drug treatment effectively. Extensive experiments showcase the efficacy of our framework, with DrugRepPT exhibiting remarkable performance improvements compared to SOTA (state of the arts) baseline methods (improvement 106.13% on Hit@1 and 54.45% on mean reciprocal rank). The reliability of predicted results is further validated through two case studies, i.e. gastritis and fatty liver, via literature validation, network medicine analysis, and docking screening. AVAILABILITY AND IMPLEMENTATION: The code and results are available at https://github.com/2020MEAI/DrugRepPT.
Shuyue Fan, Kuo Yang 0001, Kezhi Lu, Xin Dong 0017, Xianan Li, Shao Li, Jianyang Zeng 0001, Xuezhong Zhou
Bioinform.2
2024 PresRecST: a novel herbal prescription recommendation algorithm for real-world patients with integration of syndrome differentiation and treatment planning
abstract
OBJECTIVES: Herbal prescription recommendation (HPR) is a hot topic and challenging issue in field of clinical decision support of traditional Chinese medicine (TCM). However, almost all previous HPR methods have not adhered to the clinical principles of syndrome differentiation and treatment planning of TCM, which has resulted in suboptimal performance and difficulties in application to real-world clinical scenarios. MATERIALS AND METHODS: We emphasize the synergy among diagnosis and treatment procedure in real-world TCM clinical settings to propose the PresRecST model, which effectively combines the key components of symptom collection, syndrome differentiation, treatment method determination, and herb recommendation. This model integrates a self-curated TCM knowledge graph to learn the high-quality representations of TCM biomedical entities and performs 3 stages of clinical predictions to meet the principle of systematic sequential procedure of TCM decision making. RESULTS: To address the limitations of previous datasets, we constructed the TCM-Lung dataset, which is suitable for the simultaneous training of the syndrome differentiation, treatment method determination, and herb recommendation. Overall experimental results on 2 datasets demonstrate that the proposed PresRecST outperforms the state-of-the-art algorithm by significant improvements (eg, improvements of P@5 by 4.70%, P@10 by 5.37%, P@20 by 3.08% compared with the best baseline). DISCUSSION: The workflow of PresRecST effectively integrates the embedding vectors of the knowledge graph for progressive recommendation tasks, and it closely aligns with the actual diagnostic and treatment procedures followed by TCM doctors. A series of ablation experiments and case study show the availability and interpretability of PresRecST, indicating the proposed PresRecST can be beneficial for assisting the diagnosis and treatment in real-world TCM clinical settings. CONCLUSION: Our technology can be applied in a progressive recommendation scenario, providing recommendations for related items in a progressive manner, which can assist in providing more reliable diagnoses and herbal therapies for TCM clinical task.
Xin Dong 0016, Xinpeng Song, Kuo Yang 0001, Xuezhong Zhou
J. Am. Medical Informatics Assoc.11
2023 Analysis of the effect of the interaction between medical and comorbidities on hospital readmission in patients with rheumatoid arthritis based on case data
abstract
This study utilized electronic medical record data of RA inpatients to ascertain potential risk factors associated with patient readmission. Specifically, the investigation concentrated on the impact of various Traditional Chinese Medicine (TCM) syndromes, comorbidities, interactions between TCM syndromes and comorbidities, as well as interactions among comorbidities themselves, on the likelihood of hospital readmission among RA patients. The present study employed logistic regression analysis to examine the interaction between dichotomous variables. The findings revealed that individuals diagnosed with rheumatoid arthritis (RA) and experiencing arthralgia caused by wind-cold-dampness, in addition to hyperlipidemia, exhibited a heightened likelihood of hospital readmission. Similarly, RA patients with hyperlipidemia and the syndrome of dampness-heat blocking collaterals were found to be at an even greater risk of readmission. Furthermore, individuals with RA and comorbidities such as atherosclerosis and chronic gastritis were identified as being at a high risk of readmission. Lastly, RA patients with hyperlipidemia and fatty liver were also found to have an increased likelihood of being readmitted to the hospital. Hence, it is advisable to closely monitor the disease progression among the aforementioned population and implement essential interventions in clinical practice to mitigate the readmission rate of patients with rheumatoid arthritis (RA), consequently alleviating the disease and economic burdens experienced by patients.
Lifeng Fa, Kuo Yang 0001, Jinlong Yu, Zhenzhen Han, Hongtao Guo
BIBM2
2023 Application of constraint-based frequent closed itemsets Mining in TCM Clinical data Analysis
abstract
Objective: To propose a frequent closed itemsets mining method based on item constraints to study the combination rules between symptoms, diagnoses, and herbs in clinical data of TCM. Methods: Based on the Charm algorithm, prune the itemsets which do not meet the constraint conditions firstly, design a constraint based frequent closed itemsets mining algorithm, which was used on the clinical dataset of TCM to confirm the effectiveness of the method and the clinical significance of the mining results. Results: Based on the Charm algorithm, a constraint based frequent closed itemsets mining algorithm was proposed. The results of the method for mining frequent closed itemsets from data of symptom, diagnosis and herbs were consistent with the relevant theories of TCM.Conclusion: The combination of constraint conditions and frequent closed itemsets mining methods can effectively reduce the mining space. Using constraint based frequent closed itemsets mining methods to mine the combination rules between symptoms, diagnosis, and herbs in TCM clinical data maybe helpful for assisting clinical diagnosis and treatment.
Jinlong Yu, Lifeng Fa, Kuo Yang 0001
BIBM5
2023 DRONet: effectiveness-driven drug repositioning framework using network embedding and ranking learning
abstract
As one of the most vital methods in drug development, drug repositioning emphasizes further analysis and research of approved drugs based on the existing large amount of clinical and experimental data to identify new indications of drugs. However, the existing drug repositioning methods didn't achieve enough prediction performance, and these methods do not consider the effectiveness information of drugs, which make it difficult to obtain reliable and valuable results. In this study, we proposed a drug repositioning framework termed DRONet, which make full use of effectiveness comparative relationships (ECR) among drugs as prior information by combining network embedding and ranking learning. We utilized network embedding methods to learn the deep features of drugs from a heterogeneous drug-disease network, and constructed a high-quality drug-indication data set including effectiveness-based drug contrast relationships. The embedding features and ECR of drugs are combined effectively through a designed ranking learning model to prioritize candidate drugs. Comprehensive experiments show that DRONet has higher prediction accuracy (improving 87.4% on Hit@1 and 37.9% on mean reciprocal rank) than state of the art. The case analysis also demonstrates high reliability of predicted results, which has potential to guide clinical drug development.
Kuo Yang 0001, Yuxia Yang, Shuyue Fan, Jianan Xia, Qiguang Zheng, Xin Dong 0017, Zhuye Gao, Runshun Zhang, Baoyan Liu, Xuezhong Zhou
Briefings Bioinform.1
2023 DrugAI: a multi-view deep learning model for predicting drug-target activating/inhibiting mechanisms
abstract
Understanding the mechanisms of candidate drugs play an important role in drug discovery. The activating/inhibiting mechanisms between drugs and targets are major types of mechanisms of drugs. Owing to the complexity of drug-target (DT) mechanisms and data scarcity, modelling this problem based on deep learning methods to accurately predict DT activating/inhibiting mechanisms remains a considerable challenge. Here, by considering network pharmacology, we propose a multi-view deep learning model, DrugAI, which combines four modules, i.e. a graph neural network for drugs, a convolutional neural network for targets, a network embedding module for drugs and targets and a deep neural network for predicting activating/inhibiting mechanisms between drugs and targets. Computational experiments show that DrugAI performs better than state-of-the-art methods and has good robustness and generalization. To demonstrate the reliability of the predictive results of DrugAI, bioassay experiments are conducted to validate two drugs (notopterol and alpha-asarone) predicted to activate TRPV1. Moreover, external validation bears out 61 pairs of mechanism relationships between natural products and their targets predicted by DrugAI based on independent literatures and PubChem bioassays. DrugAI, for the first time, provides a powerful multi-view deep learning framework for robust prediction of DT activating/inhibiting mechanisms.
Siqin Zhang, Kuo Yang 0001, Xinxing Lai, Jianyang Zeng 0001, Shao Li
Briefings Bioinform.2
2023 SympGAN: A systematic knowledge integration system for symptom-gene associations network
Kezhi Lu, Kuo Yang 0001, Hailong Sun 0007, Qiguang Zheng, Xuezhong Zhou
Knowl. Based Syst.2
2022 RESurv: A Deep Survival Analysis Model to Reveal Population Heterogeneity by Individual Risk
abstract
Surviva1 analysis is a widely used statistical approach to model and analysis time-to-event data, which consists of event occurrence and duration before it occurs. Traditional statistical methods have good interpretability but performance is limited by strong assumptions. In recent years, methods based on machine learning, especially deep learning, have achieved success, but the lack of interpretability limits their real-world applications. This paper proposes a novel interpretable deep learning survival analysis method, RESurv, which predicts the hazard directly without any priori assumptions. Experiments on multiple public datasets validate that RESurv outperforms or approaches state-of-the-art methods. We further estimate the inflence of exposure factors on prognosis at individual level by using counterfactual approaches. Patients are then divided into subgroups by hierarchical clustering. Significant differences in risk response patterns and prognosis suggest heterogeneity in their pathological mechanisms. We explore the risk factors between subpopulations identified by different characteristics and find distinct risk patterns among them, some of which are consistent with medical research findings. The results suggest the potential of RESurv to estimate individual risk, and thereby revealing the heterogeneity of risk response patterns across the populations.
Qiguang Zheng, Qifan Shen, Xin Su 0011, Kuo Yang 0001, Zixin Shu, Xuezhong Zhou
BIBM4
2022 Decoding multilevel relationships with the human tissue-cell-molecule network
abstract
Understanding the biological functions of molecules in specific human tissues or cell types is crucial for gaining insights into human physiology and disease. To address this issue, it is essential to systematically uncover associations among multilevel elements consisting of disease phenotypes, tissues, cell types and molecules, which could pose a challenge because of their heterogeneity and incompleteness. To address this challenge, we describe a new methodological framework, called Graph Local InfoMax (GLIM), based on a human multilevel network (HMLN) that we established by introducing multiple tissues and cell types on top of molecular networks. GLIM can systematically mine the potential relationships between multilevel elements by embedding the features of the HMLN through contrastive learning. Our simulation results demonstrated that GLIM consistently outperforms other state-of-the-art algorithms in disease gene prediction. Moreover, GLIM was also successfully used to infer cell markers and rewire intercellular and molecular interactions in the context of specific tissues or diseases. As a typical case, the tissue-cell-molecule network underlying gastritis and gastric cancer was first uncovered by GLIM, providing systematic insights into the mechanism underlying the occurrence and development of gastric cancer. Overall, our constructed methodological framework has the potential to systematically uncover complex disease mechanisms and mine high-quality relationships among phenotypical, tissue, cellular and molecular elements.
Siyu Hou, Peng Zhang 0149, Kuo Yang 0001, Changzheng Ma, Yanda Li, Shao Li
Briefings Bioinform.3
2022 PDGNet: Predicting Disease Genes Using a Deep Neural Network With Multi-View Features
abstract
The knowledge of phenotype-genotype associations is crucial for the understanding of disease mechanisms. Numerous studies have focused on developing efficient and accurate computing approaches to predict disease genes. However, owing to the sparseness and complexity of medical data, developing an efficient deep neural network model to identify disease genes remains a huge challenge. Therefore, we develop a novel deep neural network model that fuses the multi-view features of phenotypes and genotypes to identify disease genes (termed PDGNet). Our model integrated the multi-view features of diseases and genes and leveraged the feedback information of training samples to optimize the parameters of deep neural network and obtain the deep vector features of diseases and genes. The evaluation experiments on a large data set indicated that PDGNet obtained higher performance than the state-of-the-art method (precision and recall improved by 9.55 and 9.63 percent). The analysis results for the candidate genes indicated that the predicted genes have strong functional homogeneity and dense interactions with known genes. We validated the top predicted genes of Parkinson's disease based on external curated data and published medical literatures, which indicated that the candidate genes have a huge potential to guide the selection of causal genes in the 'wet experiment'. The source codes and the data of PDGNet are available at https://github.com/yangkuoone/PDGNet.
Kuo Yang 0001, Kezhi Lu, Kai Chang, Ning Wang 0048, Zixin Shu, Jian Yu 0001, Baoyan Liu, Zhuye Gao, Xuezhong Zhou
IEEE ACM Trans. Comput. Biol. Bioinform.1
2021 TCMPR: TCM Prescription recommendation based on subnetwork term mapping and deep learning
abstract
Traditional Chinese medicine (TCM) has played an indispensable role in clinical diagnose and treatment. Based on patient’s symptom phenotypes, computation-based prescription recommendation methods can recommend personalized TCM prescription using machine learning and artificial intelligence technologies. However, owing to the complexity and individuation of patient’s clinical phenotypes, current prescription recommendation methods cannot obtain good performance. Meanwhile, it’s very difficult to conduct effective representation for unrecorded symptom terms in existing knowledge base. In this study, we proposed a subnetwork-based symptom term mapping method (SSTM), and constructed a SSTM-based TCM prescription recommendation method (termed TCMPR). Our SSTM can extract the subnetwork structure between symptoms from knowledge network to effectively represent the embedding features of clinical symptom terms (especially, the unrecorded terms). The experimental results showed that our method performs better than state-of-the-art methods. In addition, the comprehensive experiments of TCMPR with different hyper parameters (i.e., feature embedding, feature dimension and feature fusion) that demonstrates that our method has high performance on TCM prescription recommendation and potentially promote clinical diagnosis and treatment of TCM precision medicine.
Xin Dong 0017, Zixin Shu, Kai Chang, Dengying Yan, Jianan Xia, Kunyu Zhong, Xinyan Wang 0002, Kuo Yang 0001, Xuezhong Zhou
BIBM10
2021 Phenonizer: A fine-grained phenotypic named entity recognizer for Chinese clinical texts
abstract
Biomedical named entity recognition from clinical texts is a fundamental task for clinical data analysis due to the availability of large volume of electronic medical record data, which are mostly in free text format, in real-world clinical settings. Clinical text data incorporates significant phenotypic medical entities, which could be used for profiling the clinical characteristics of patients in specific disease conditions. However, general approaches mostly rely on the coarse-grained annotations (e.g. mentions of symptom terms) of phenotypic entities in benchmark text dataset. Owing to the numerous negation expressions of phenotypic entities (e.g. “no fever”, “no cough” and “no hypertension”) in clinical texts, this could not feed the subsequent data analysis process with well-prepared structured clinical data. Thus, we constructed a fine-grained Chinese clinical corpus. Thereafter, we proposed a phenotypic named entity recognizer (Phenonizer). The results on the test set show that Phenonizer outperform those methods based on Word2Vec with Fl-score of 0.896. By comparing character embeddings from different data, it is found that character embeddings trained by clinical corpora can improve F-score by 0.0103. Furthermore, the fine-grained dataset enables methods to distinguish between negated symptoms and presented symptoms, and avoids the interference of negated symptoms. Finally, we tested the generalization performance of Phenonier, achieving a superior F1-score of 0.8389. In summary, together with fine-grained annotated benchmark dataset, Phenonier proposes a feasible approach to effectively extract symptom information from Chinese clinical texts with acceptable performance.
Qunsheng Zou, Kuo Yang 0001, Kai Chang, Xuezhong Zhou
BIBM2
2020 Network-based gene prediction for TCM symptoms
abstract
The diagnosis and treatment of traditional Chinese medicine (TCM) are formed based on the differentiation of syndromes and symptoms. Symptom management is always the core task of nursing science. Connotation between TCM symptoms and Modern medicine (MM) symptoms are obvious different, especially tongue and pulse symptoms of TCM. However, the underlying molecular mechanisms of most TCM symptoms remain unclear. Here, we developed a network-based framework to predict candidate genes of TCM symptoms (called PTsGene) and construction a high-quality set of TCM symptom-gene associations. Experimental results indicated that PTsGene performed significantly better than the baseline algorithms. The reliability of the candidate genes of symptoms (containing one of typical symptoms of COVID-19, fever) were validated by the analysis of functional homogeneity, molecular co-expression, and recently published literatures. Finally, a high-quality set of TCM symptom-gene associations is constructed to promote the mechanism developments of TCM symptoms. Prediction and construction for reliable TCM symptom-gene associations are valuable for uncovering the underlying molecular mechanisms of TCM symptoms. Our TCM symptom-gene associations deliver a highly insightful data sources for researchers both from basic and clinical settings of precision healthcare.
Yinyan Wang, Kuo Yang 0001, Zixin Shu, Dengying Yan, Xuezhong Zhou
BIBM2
2020 Disease phenotype synonymous prediction through network representation learning from PubMed database
abstract
Synonym mapping between phenotype concepts from different terminologies is difficult because terminology databases have been developed largely independently. Existing maps of synonymous phenotype concepts from different terminology databases are highly incomplete, and manually mapping is time consuming and laborious. Therefore, building an automatic method for predictive mapping of synonymous phenotypes is of special importance. We propose a classifier-based phenotype mapping prediction model (CPM) to predict synonymous relationships between phenotype concepts from different terminology databases. The model takes network semantic representations of phenotypes as input and predicts synonymous relationships by training binary classifiers with a voting strategy. We compared the performance of the CPM with a similarity-based phenotype mapping prediction model (SPM), which predicts mapping based on the ranked cosine similarity of candidate mapping concepts. Based on a network representation N2V-TFIDF, with a majority voting strategy method MV, the CPM achieved accuracy of 0.943, which was 15.4% higher than that of the SPM using the cosine similarity method (0.789) and 23.8% higher than that of the SSDTM method (0.724) proposed in our previous work.
Shiwen Ma, Kuo Yang 0001, Ning Wang 0048, Zhuye Gao, Runshun Zhang, Baoyan Liu, Xuezhong Zhou
Artif. Intell. Medicine2
2020 Integrated network analysis of symptom clusters across disease conditions
Kezhi Lu, Kuo Yang 0001, Edouard Niyongabo, Zixin Shu, Kai Chang, Qunsheng Zou, Jiyue Jiang, Caiyan Jia, Baoyan Liu, Xuezhong Zhou
J. Biomed. Informatics2
2019 HerGePred: Heterogeneous Network Embedding Representation for Disease Gene Prediction
abstract
The discovery of disease-causing genes is a critical step towards understanding the nature of a disease and determining a possible cure for it. In recent years, many computational methods to identify disease genes have been proposed. However, making full use of disease-related (e.g., symptoms) and gene-related (e.g., gene ontology and protein-protein interactions) information to improve the performance of disease gene prediction is still an issue. Here, we develop a heterogeneous disease-gene-related network (HDGN) embedding representation framework for disease gene prediction (called HerGePred). Based on this framework, a low-dimensional vector representation (LVR) of the nodes in the HDGN can be obtained. Then, we propose two specific algorithms, namely, an LVR-based similarity prediction and a random walk with restart on a reconstructed heterogeneous disease-gene network (RW-RDGN), to predict disease genes with high performance. First, to validate the rationality of the framework, we analyze the similarity-based overlap distribution of disease pairs and design an experiment for disease-gene association recovery, the results of which revealed that the LVR of nodes performs well at preserving the local and global network structure of the HDGN. Then, we apply tenfold cross validation and external validation to compare our methods with other well-known disease gene prediction algorithms. The experimental results show that the RW-RDGN performs better than the state-of-the-art algorithm. The prediction results of disease candidate genes are essential for molecular mechanism investigation and experimental validation. The source codes of HerGePred and experimental data are available at https://github.com/yangkuoone/HerGePred.
Kuo Yang 0001, Ruyu Wang, Zixin Shu, Ning Wang 0048, Runshun Zhang, Jian Yu 0001, Xuezhong Zhou
IEEE J. Biomed. Health Informatics1
2018 Heterogeneous network embedding for identifying symptom candidate genes
abstract
Objective: Investigating the molecular mechanisms of symptoms is a vital task in precision medicine to refine disease taxonomy and improve the personalized management of chronic diseases. Although there are abundant experimental studies and computational efforts to obtain the candidate genes of diseases, the identification of symptom genes is rarely addressed. We curated a high-quality benchmark dataset of symptom-gene associations and proposed a heterogeneous network embedding for identifying symptom genes. Methods: We proposed a heterogeneous network embedding representation algorithm, which constructed a heterogeneous symptom-related network that integrated symptom-related associations and applied an embedding representation algorithm to obtain the low-dimensional vector representation of nodes. By measuring the relevance between symptoms and genes via calculating the similarities of their vectors, the candidate genes of given symptoms can be obtained. Results: A benchmark dataset of 18 270 symptom-gene associations between 505 symptoms and 4549 genes was curated. We compared our method to baseline algorithms (FSGER and PRINCE). The experimental results indicated our algorithm achieved a significant improvement over the state-of-the-art method, with precision and recall improved by 66.80% (0.844 vs 0.506) and 53.96% (0.311 vs 0.202), respectively, for TOP@3 and association precision improved by 37.71% (0.723 vs 0.525) over the PRINCE. Conclusions: The experimental validation of the algorithms and the literature validation of typical symptoms indicated our method achieved excellent performance. Hence, we curated a prediction dataset of 17 479 symptom-candidate genes. The benchmark and prediction datasets have the potential to promote investigations of the molecular mechanisms of symptoms and provide candidate genes for validation in experimental settings.
Kuo Yang 0001, Ning Wang 0048, Ruyu Wang, Jian Yu 0001, Runshun Zhang, Xuezhong Zhou
J. Am. Medical Informatics Assoc.1
2016 Similarity-based algorithms for Disease Terminology Mapping
abstract
Classification of diseases and their related terms are important data resources for basic medical research. However, disease terms in different terminological databases are largely developed independent of each other and the mapping relationships between them are not complete. The purpose of this paper is to propose similarity-based disease terminology mapping methods to map disease terms with same or similar semantic concepts in different terminological databases. By integrating the bibliographic medical records from PubMed and the manually curated associations of disease-gene and disease-phenotype, we proposed two methods, namely Semantic-based Similarity for Disease Terminology Mapping (SSDTM) and Information Recommend-based Disease Terminology Mapping (IRDTM) to predict the disease terms in OMIM to MeSH. The experimental results show that both methods can support predicting mapping between disease databases. From leave one out cross validation, the prediction performance of SSDTM (Hits@10: 87.3%) is better than IRDTM; From manual evaluation, the hits rate in top 10 of SSDTM is 94.4%. The similarity-based disease terminology mapping method can be applied with the supplement of manual review for different mainstream disease terminology databases and help improve the efficiency of integrated medical ontology development and translational bioinformatics that need incorporate multiple data sources from different disciplines.
Shiwen Ma, Kuo Yang 0001, Xuezhong Zhou
BIBM2