Xuezhong Zhou

dblp:44/1540 · DBLP profile ↗
← Back
58ranked-venue papers
7as first author
26since 2021 · last 2026
0000-0002-4713-3594ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 45 · 3 first-author · 21 since 2021Artificial intelligence and machine learning · 12 · 5 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 Distribution-aware neural network: A novel patient representation learning algorithm for Traditional Chinese medicine diagnosis prediction
abstract
Traditional Chinese medicine (TCM) diagnosis involves complex and implicit associations between heterogeneous symptoms and diagnostic patterns, and distributional heterogeneity across diseases and patients. Existing intelligent diagnostic models focus primarily on architectural optimization but lack explicit modeling of underlying symptom-diagnosis distributions, resulting in limited robustness, cross-disease generalization, and interpretability. To address these challenges, we propose a distribution-aware neural network (DANN) for diagnostic representation learning. The proposed framework incorporates explicit representations of both global and class-conditional feature distributions and integrates discriminative pruning and latent structure decomposition to capture population-level diagnostic regularities and fine-grained differential variations. In addition, we introduce a cross-disease clinical dataset (TCM-Chronic) covering 15 chronic diseases, 5876 clinical cases, and 97 diagnostic labels to simulate real-world comorbidity scenarios. Experiments on both a public multilabel dataset (TCM-Lung) and a cross-disease dataset demonstrate that the DANN consistently outperforms state-of-the-art machine learning, deep learning, and large language model baselines. With respect to TCM-Lung, the F1-score of the DANN is 0.5112, which is 3.6 percentage points greater than that of the strongest baseline. With respect to TCM-Chronic, the DANN achieves an F1-score of 0.7146, outperforming the random forest by 7.21 percentage points. Ablation and expert evaluations further confirm that distribution-aware modeling contributes to increased diagnostic robustness and better interpretability. These results indicate that explicitly modeling diagnostic feature distributions provides an effective paradigm for intelligent diagnosis, with potential applicability beyond TCM to broader clinical decision-support tasks. • A cross-disease diagnostic dataset with 15 diseases and 97 labels is constructed. • A distribution-aware neural network models symptom-diagnosis distributions. • State-of-the-art performance is shown across multilabel and small-sample settings. • Ablation studies and expert evaluation confirm enhanced interpretability.
Zongyao Zhao, Xin Dong 0016, Xinpeng Song, Zuoyuan Luo, Geyan Pan, Sicen Wang, Xuezhong Zhou
Eng. Appl. Artif. Intell.11
2026 Meta Relation-Aware Metric Learning Framework for Few-Shot Uncertain Knowledge Graph Completion
abstract
Uncertainty is a fundamental aspect of human knowledge, widely present in various domains such as medicine, biology, and finance. Uncertain Knowledge Graph Completion (UKGC) tasks encompass both link prediction and confidence prediction. However, the inaccessibility of high-quality knowledge results in sparse knowledge graphs, where a large number of entities exhibit relatively sparse neighborhoods. This sparsity severely hinders the performance of few-shot UKGC models, which generally rely on neighborhood aggregation to enhance entity representations. In this study, we propose a Meta relation-aware Metric learning framework designed for few-shot UKGC, named MMUC. To mitigate neighbor sparsity, MMUC introduces a relation-centric paradigm that decouples representation enhancement from background graphs, integrating an uncertainty-driven similarity measure through a learnable confidence learner to achieve robust, non-linear fusion for joint link and confidence prediction. Extensive evaluations across multiple benchmarks demonstrate MMUC's superior performance through rigorous link and confidence prediction, ablation studies, and sparsity analysis, highlighting its particular resilience to neighbor sparsity in few-shot uncertain KGs. The source code can be accessed inhttps://github.com/sienna-wxy/MMUC.
Xinyan Wang 0002, Kuo Yang 0001, Xianan Li, Zixin Shu, Xuezhong Zhou
IEEE Trans. Knowl. Data Eng.8
2025 KGCRR: An Effective Metric-Driven Knowledge Graph Completion Framework by Designing a Novel Upper Bound Function with Adaptive Approximation to Reciprocal Rank
abstract
Knowledge Graph Embedding (KGE) methods have achieved great success in predicting missing links in knowledge graphs, a task also known as Knowledge Graph Completion (KGC). Under this task, the Reciprocal Rank (RR) of ground-truth items serve as a key indicator for evaluating the method’s performance. However, most existing studies have overlooked the inconsistency between the ranking metric, RR, and the optimization objective functions, resulting in sub-optimal KGC performance. To address this issue, we propose a KGC framework called KGCRR by designing a novel upper bound function named CRR. By introducing the parameter-pressure ρ to shift the sigmoid function, CRR achieves a better approximation to RR compared with existing objective functions. We theoretically proved that by adjusting ρ, CRR can achieve a more effective approximation to RR. By narrowing the discrepancy with RR and alleviating the gradient vanishing issue associated with the direct optimization of RR loss, CRR demonstrates an advantage in optimizing RR. CRR serves as a plug-and-play objective, capable of seamless integration into various KGE methods. Through extensive experiments conducted on FB15k-237 and WN18RR datasets, we have obtained promising results, with an average improvement of 19.06% in MRR, indicating that CRR significantly enhances the performance of existing methods.
Kuo Yang 0001, Xiangkui Lu, Xuezhong Zhou
AAAI6
2025 PresRecSKG: Enhancing Herbal Prescription Recommendation via Knowledge Graph-Guided Data Augmentation
abstract
With the rapid advancement of medical artificial intelligence, Traditional Chinese Medicine (TCM) has demonstrated significant potential in intelligent prescription recommendation. However, the scarcity of structured and high-quality clinical data often leads to issues such as data sparsity and terminology inconsistency, which limit the generalization capability and practical performance of existing models. To tackle these challenges, this study proposes PresRecSKG, a knowledge graph-guided data augmentation framework that utilizes a herbsymptom knowledge graph to enhance both multi-label and multi-class prescription recommendation tasks. The core of the framework is the proposed Sampling-by-Knowledge-Graph (SabKG) strategy, which effectively incorporates external knowledge for data synthesis, thereby mitigating data sparsity and improving data quality control. Experimental results show that PresRecSKG achieves the best performance in the multi-label task, with an average improvement of approximately 2.5% across all evaluation metrics. In the multi-class task, the framework elevates performance from 0.94 to over 0.97 across key metrics. These findings indicate that PresRecSKG offers an effective data augmentation paradigm for TCM prescription recommendation, enhancing the accuracy, compatibility, and interpretability of intelligent TCM systems.
Xin Dong 0017, Jiahe Liu, Kuo Yang 0001, Xuezhong Zhou
BIBM4
2025 Cross-Domain Multi-Modal Transfer Learning for Target Identification of TCM's Compounds
Huan Gu, Xunpeng Xiao, Chongyu Wang, Kuo Yang 0001, Xuezhong Zhou
BIBM6
2025 CPOne: Enhancing Prediction of Compound-Protein Interactions Through One-Shot Meta Learning
abstract
Predicting the interactions between compounds and their potential target proteins is crucial in drug discovery. Existing methods often assume that each compound has an adequate number of target proteins available for model training. However, in practice, the number of target proteins associated with compounds is often limited, making it difficult to gather a sufficient number of training examples. This results in learning bias in the model, leading to poor performance on these compounds. Moreover, this issue is evident in widely used datasets, such as GPCR and kinase, where 77.51% and 12.71% of compounds, respectively, have only one target protein available for model training. However, the issue has not been fully explored, presenting a challenge for model development. In this research, we propose CPOne, a framework designed to address the aforementioned issue. CPOne is a meta-learning based approach for compound-protein interaction prediction in one-shot scenario where each compound has only one target protein available for model training. By utilizing a meta compound learner, CPOne extracts the meta representation of compound from each task. Through fast gradient updates, this representation is quickly adapted to generate a compound-specific representation for the current task, thereby improving performance in one-shot scenario. Through comprehensive experiments, we empirically validate the superiority of CPOne, which demonstrates a promising performance improvement over established methods.
Kuo Yang 0001, Chongyu Wang, Xinyan Wang 0002, Hanyu Yuan, Jian Yu 0001, Xuezhong Zhou
IEEE Trans. Comput. Biol. Bioinform.8
2024 TCM-FTP: Fine-Tuning Large Language Models for Herbal Prescription Prediction
abstract
Traditional Chinese medicine (TCM) has relied on specific combinations of herbs in prescriptions to treat various symptoms and signs for thousands of years. Predicting TCM prescriptions poses a fascinating technical challenge with significant practical implications. However, this task faces limitations due to the scarcity of high-quality clinical datasets and the complex relationship between symptoms and herbs. To address these issues, we introduce DigestDS, a novel dataset comprising practical medical records from experienced experts in digestive system diseases. We also propose a method, TCM-FTP (TCM Fine-Tuning Pre-trained), to leverage pre-trained large language models (LLMs) via supervised fine-tuning on DigestDS. Additionally, we enhance computational efficiency using a low-rank adaptation technique. Moreover, TCM-FTP incorporates data augmentation by permuting herbs within prescriptions, exploiting their order-agnostic nature. Impressively, TCM-FTP achieves an F1-score of 0.8031, significantly outperforming previous methods. Furthermore, it demonstrates remarkable accuracy in dosage prediction, achieving a normalized mean square error of 0.0604. In contrast, LLMs without fine-tuning exhibit poor performance. Although LLMs have demonstrated wide-ranging capabilities, our work underscores the necessity of fine-tuning for TCM prescription prediction and presents an effective way to accomplish this.
Xingzhi Zhou 0002, Xin Dong 0017, Chunhao Li, Yuning Bai, Ka Chun Cheung, Simon See, Xinpeng Song, Runshun Zhang, Xuezhong Zhou, Nevin Lianwen Zhang
BIBM10
2024 LMDTA: Molecular Pre-trained and Interaction Fine-tuned Attention Neural Network for Drug-Target Affinity Prediction
abstract
Accurately predicting drug-target binding affinity is crucial for advancing drug discovery. Recent molecular pre-training models and biological large models provide general molecular representations, but effectively leveraging these features for specific tasks like Drug-Target Affinity (DTA) prediction remains challenging. To address this, we propose LMDTA, a novel attention-based neural network combining molecular pre-training with interaction fine-tuning for affinity prediction. LMDTA learns complex drug and protein representations through pre-trained models, while fine-tuning an interaction-specific model for the DTA task. The two-side attention mechanism integrates pre-trained and fine-tuned features, capturing key drug-target interactions. Experiments on benchmark datasets show LMDTA achieves state-of-the-art performance, with ablation studies validating the model’s design.
Minjie Hu, Kuo Yang 0001, Xuezhong Zhou
BIBM4
2024 PresRecCD: A Novel Herbal Prescription Recommendation Framework with Cross-Domain Learning and Neural Collaborative Filtering
abstract
Herbal prescriptions hold significant importance in Traditional Chinese Medicine (TCM) diagnosis and treatment, embodying millennia of clinical case summaries and wisdom. Despite numerous proposed methods for herbal prescription recommendation (HPR), significant challenges persist due to the lack of comprehensive clinical data, particularly regarding the relationships between symptoms and herbs. This scarcity poses considerable hurdles for effective HPR modeling. In this study, we introduced a novel herbal prescription recommendation framework with cross-domain learning and neural collaborative filtering (termed PresRecCD). The cross-domain learning mechanism is introduced to learn the noise-reduced cross-domain features of herbs and symptoms in the unified space, alleviating the sparsity of data, and neural collaborative filtering is utilized to carry out prescription recommendations. Comprehensive experiments demonstrate the superiority of the proposed PresRecCD model over the SOTA model. This study contributes to enhancing the performance of the HPR model, ultimately benefiting the efficiency and precision of clinical treatment.
Wansong Zhang, Xin Dong 0017, Kuo Yang 0001, Rouye Huang, Runshun Zhang, Xuezhong Zhou
BIBM8
2024 Benchmarking Biomedical Relation Knowledge in Large Language Models
Kuo Yang 0001, Chenqian Zhao, Haixu Li, Xin Dong 0016, Xuezhong Zhou
ISBRA (2)7
2024 Dynamic Feature Fusion Based on Consistency and Complementarity of Brain Atlases
Qiye Lin, Ruiwen Fan, Xuezhong Zhou, Jianan Xia
PRCV (15)4
2024 HTINet2: herb-target prediction via knowledge graph embedding and residual-like graph neural network
abstract
Target identification is one of the crucial tasks in drug research and development, as it aids in uncovering the action mechanism of herbs/drugs and discovering new therapeutic targets. Although multiple algorithms of herb target prediction have been proposed, due to the incompleteness of clinical knowledge and the limitation of unsupervised models, accurate identification for herb targets still faces huge challenges of data and models. To address this, we proposed a deep learning-based target prediction framework termed HTINet2, which designed three key modules, namely, traditional Chinese medicine (TCM) and clinical knowledge graph embedding, residual graph representation learning, and supervised target prediction. In the first module, we constructed a large-scale knowledge graph that covers the TCM properties and clinical treatment knowledge of herbs, and designed a component of deep knowledge embedding to learn the deep knowledge embedding of herbs and targets. In the remaining two modules, we designed a residual-like graph convolution network to capture the deep interactions among herbs and targets, and a Bayesian personalized ranking loss to conduct supervised training and target prediction. Finally, we designed comprehensive experiments, of which comparison with baselines indicated the excellent performance of HTINet2 (HR@10 increased by 122.7% and NDCG@10 by 35.7%), ablation experiments illustrated the positive effect of our designed modules of HTINet2, and case study demonstrated the reliability of the predicted targets of Artemisia annua and Coptis chinensis based on the knowledge base, literature, and molecular docking.
Pengbo Duan, Kuo Yang 0001, Xin Su 0011, Shuyue Fan, Xin Dong 0017, Xianan Li, Xiaoyan Xing, Jian Yu 0001, Xuezhong Zhou
Briefings Bioinform.11
2024 KDGene: knowledge graph completion for disease gene prediction using interactional tensor decomposition
abstract
The accurate identification of disease-associated genes is crucial for understanding the molecular mechanisms underlying various diseases. Most current methods focus on constructing biological networks and utilizing machine learning, particularly deep learning, to identify disease genes. However, these methods overlook complex relations among entities in biological knowledge graphs. Such information has been successfully applied in other areas of life science research, demonstrating their effectiveness. Knowledge graph embedding methods can learn the semantic information of different relations within the knowledge graphs. Nonetheless, the performance of existing representation learning techniques, when applied to domain-specific biological data, remains suboptimal. To solve these problems, we construct a biological knowledge graph centered on diseases and genes, and develop an end-to-end knowledge graph completion framework for disease gene prediction using interactional tensor decomposition named KDGene. KDGene incorporates an interaction module that bridges entity and relation embeddings within tensor decomposition, aiming to improve the representation of semantically similar concepts in specific domains and enhance the ability to accurately predict disease genes. Experimental results show that KDGene significantly outperforms state-of-the-art algorithms, whether existing disease gene prediction methods or knowledge graph embedding methods for general domains. Moreover, the comprehensive biological analysis of the predicted results further validates KDGene's capability to accurately identify new candidate genes. This work proposes a scalable knowledge graph completion framework to identify disease candidate genes, from which the results are promising to provide valuable references for further wet experiments. Data and source codes are available at https://github.com/2020MEAI/KDGene.
Xinyan Wang 0002, Kuo Yang 0001, Ting Jia, Fanghui Gu, Chongyu Wang, Zixin Shu, Jianan Xia, Xuezhong Zhou
Briefings Bioinform.10
2024 DrugRepPT: a deep pretraining and fine-tuning framework for drug repositioning based on drug's expression perturbation and treatment effectiveness
abstract
MOTIVATION: Drug repositioning (DR), identifying novel indications for approved drugs, is a cost-effective strategy in drug discovery. Despite numerous proposed DR models, integrating network-based features, differential gene expression, and chemical structures for high-performance DR remains challenging. RESULTS: We propose a comprehensive deep pretraining and fine-tuning framework for DR, termed DrugRepPT. Initially, we design a graph pretraining module employing model-augmented contrastive learning on a vast drug-disease heterogeneous graph to capture nuanced interactions and expression perturbations after intervention. Subsequently, we introduce a fine-tuning module leveraging a graph residual-like convolution network to elucidate intricate interactions between diseases and drugs. Moreover, a Bayesian multiloss approach is introduced to balance the existence and effectiveness of drug treatment effectively. Extensive experiments showcase the efficacy of our framework, with DrugRepPT exhibiting remarkable performance improvements compared to SOTA (state of the arts) baseline methods (improvement 106.13% on Hit@1 and 54.45% on mean reciprocal rank). The reliability of predicted results is further validated through two case studies, i.e. gastritis and fatty liver, via literature validation, network medicine analysis, and docking screening. AVAILABILITY AND IMPLEMENTATION: The code and results are available at https://github.com/2020MEAI/DrugRepPT.
Shuyue Fan, Kuo Yang 0001, Kezhi Lu, Xin Dong 0017, Xianan Li, Shao Li, Jianyang Zeng 0001, Xuezhong Zhou
Bioinform.9
2024 PresRecST: a novel herbal prescription recommendation algorithm for real-world patients with integration of syndrome differentiation and treatment planning
abstract
OBJECTIVES: Herbal prescription recommendation (HPR) is a hot topic and challenging issue in field of clinical decision support of traditional Chinese medicine (TCM). However, almost all previous HPR methods have not adhered to the clinical principles of syndrome differentiation and treatment planning of TCM, which has resulted in suboptimal performance and difficulties in application to real-world clinical scenarios. MATERIALS AND METHODS: We emphasize the synergy among diagnosis and treatment procedure in real-world TCM clinical settings to propose the PresRecST model, which effectively combines the key components of symptom collection, syndrome differentiation, treatment method determination, and herb recommendation. This model integrates a self-curated TCM knowledge graph to learn the high-quality representations of TCM biomedical entities and performs 3 stages of clinical predictions to meet the principle of systematic sequential procedure of TCM decision making. RESULTS: To address the limitations of previous datasets, we constructed the TCM-Lung dataset, which is suitable for the simultaneous training of the syndrome differentiation, treatment method determination, and herb recommendation. Overall experimental results on 2 datasets demonstrate that the proposed PresRecST outperforms the state-of-the-art algorithm by significant improvements (eg, improvements of P@5 by 4.70%, P@10 by 5.37%, P@20 by 3.08% compared with the best baseline). DISCUSSION: The workflow of PresRecST effectively integrates the embedding vectors of the knowledge graph for progressive recommendation tasks, and it closely aligns with the actual diagnostic and treatment procedures followed by TCM doctors. A series of ablation experiments and case study show the availability and interpretability of PresRecST, indicating the proposed PresRecST can be beneficial for assisting the diagnosis and treatment in real-world TCM clinical settings. CONCLUSION: Our technology can be applied in a progressive recommendation scenario, providing recommendations for related items in a progressive manner, which can assist in providing more reliable diagnoses and herbal therapies for TCM clinical task.
Xin Dong 0016, Xinpeng Song, Kuo Yang 0001, Xuezhong Zhou
J. Am. Medical Informatics Assoc.12
2024 Lingdan: enhancing encoding of traditional Chinese medicine knowledge for clinical reasoning tasks with large language models
abstract
OBJECTIVE: The recent surge in large language models (LLMs) across various fields has yet to be fully realized in traditional Chinese medicine (TCM). This study aims to bridge this gap by developing a large language model tailored to TCM knowledge, enhancing its performance and accuracy in clinical reasoning tasks such as diagnosis, treatment, and prescription recommendations. MATERIALS AND METHODS: This study harnessed a wide array of TCM data resources, including TCM ancient books, textbooks, and clinical data, to create 3 key datasets: the TCM Pre-trained Dataset, the Traditional Chinese Patent Medicine (TCPM) Question Answering Dataset, and the Spleen and Stomach Herbal Prescription Recommendation Dataset. These datasets underpinned the development of the Lingdan Pre-trained LLM and 2 specialized models: the Lingdan-TCPM-Chat Model, which uses a Chain-of-Thought process for symptom analysis and TCPM recommendation, and a Lingdan Prescription Recommendation model (Lingdan-PR) that proposes herbal prescriptions based on electronic medical records. RESULTS: The Lingdan-TCPM-Chat and the Lingdan-PR Model, fine-tuned on the Lingdan Pre-trained LLM, demonstrated state-of-the art performances for the tasks of TCM clinical knowledge answering and herbal prescription recommendation. Notably, Lingdan-PR outperformed all state-of-the-art baseline models, achieving an improvement of 18.39% in the Top@20 F1-score compared with the best baseline. CONCLUSION: This study marks a pivotal step in merging advanced LLMs with TCM, showcasing the potential of artificial intelligence to help improve clinical decision-making of medical diagnostics and treatment strategies. The success of the Lingdan Pre-trained LLM and its derivative models, Lingdan-TCPM-Chat and Lingdan-PR, not only revolutionizes TCM practices but also opens new avenues for the application of artificial intelligence in other specialized medical fields. Our project is available at https://github.com/TCMAI-BJTU/LingdanLLM.
Xin Dong 0016, Zixin Shu, Yunhui Hu, Shuiping Zhou, Kaijing Yan, Xijun Yan, Kai Chang, Yuning Bai, Runshun Zhang, Xuezhong Zhou
J. Am. Medical Informatics Assoc.16
2023 Classification Characteristics of COPD Based on Combination of Disease and Syndrome in Real World
abstract
Objective: To study the phenotypic characteristics of the combination of disease and syndrome of COPD. Methods: Structured clinical EMRs of inpatients in the respiratory department of the First Affiliated Hospital of Henan University of Chinese Medicine for a total of 10 years from 2010 to 2019 were collected. Patients with COPD in Western medicine diagnosis on admission or discharge were used as the study subjects. A Heterogeneous Medical Record network based on COPD patients was constructed. Then, the K-means clustering algorithm was used to divide the supplemented patient medical record feature matrix into different modules (ie, subtypes), and enrichment analysis was performed on the phenotypes characteristics of different modules to determine the typical clinical phenotypes characteristics of different modules. Results: After the analysis of different phenotypic characteristics of COPD patients, it was found that the phenotypic characteristics of COPD patients were diverse. The prominent manifestation is that COPD patients often have diseases and symptoms other than respiratory system, and more than 70% of them have 2-6 different diseases. Through cluster study and analysis of each subtype of COPD patients revealed that each module has its own unique core characteristics. Conclusions: After the analysis of different phenotypic characteristics of COPD patients, it was found that the phenotypic characteristics of COPD patients were diverse, and each subtype has its own unique core characteristics.
Kunyu Zhong, Kai Chang, Lifeng Fa, Jinlong Yu, Xuezhong Zhou
BIBM7
2023 DRONet: effectiveness-driven drug repositioning framework using network embedding and ranking learning
abstract
As one of the most vital methods in drug development, drug repositioning emphasizes further analysis and research of approved drugs based on the existing large amount of clinical and experimental data to identify new indications of drugs. However, the existing drug repositioning methods didn't achieve enough prediction performance, and these methods do not consider the effectiveness information of drugs, which make it difficult to obtain reliable and valuable results. In this study, we proposed a drug repositioning framework termed DRONet, which make full use of effectiveness comparative relationships (ECR) among drugs as prior information by combining network embedding and ranking learning. We utilized network embedding methods to learn the deep features of drugs from a heterogeneous drug-disease network, and constructed a high-quality drug-indication data set including effectiveness-based drug contrast relationships. The embedding features and ECR of drugs are combined effectively through a designed ranking learning model to prioritize candidate drugs. Comprehensive experiments show that DRONet has higher prediction accuracy (improving 87.4% on Hit@1 and 37.9% on mean reciprocal rank) than state of the art. The case analysis also demonstrates high reliability of predicted results, which has potential to guide clinical drug development.
Kuo Yang 0001, Yuxia Yang, Shuyue Fan, Jianan Xia, Qiguang Zheng, Xin Dong 0017, Zhuye Gao, Runshun Zhang, Baoyan Liu, Xuezhong Zhou
Briefings Bioinform.16
2023 SympGAN: A systematic knowledge integration system for symptom-gene associations network
Kezhi Lu, Kuo Yang 0001, Hailong Sun 0007, Qiguang Zheng, Xuezhong Zhou
Knowl. Based Syst.8
2022 Research and Implementation of Real World Traditional Chinese Medicine Clinical Scientific Research Information Electronic Medical Record Sharing System
abstract
The standardization degree of traditional Chinese medicine clinical data in the real world is low and heterogeneous data aggregation among institutions is difficult, which leads to the difficulty of sharing clinical data and scientific research data of traditional Chinese medicine. This paper designs and implements a real world traditional Chinese medicine clinical scientific research information electronic medical record sharing system. The system consists mainly of two subsystems, namely electronic medical record collection system and electronic medical record integration system. The collection system can collect, normalize and structured storage inpatient electronic medical records, outpatient electronic medical records and cloud platform electronic medical records. The integration system can integrate heterogeneous data from different traditional Chinese medicine diagnostic and treatment institutions to realize the sharing of traditional Chinese medicine clinical and research data.
Qi Xie 0005, Runshun Zhang, Xuezhong Zhou, Tiancai Wen, Xingping Zhang, Xu Miao, Baoyan Liu
BIBM4
2022 RESurv: A Deep Survival Analysis Model to Reveal Population Heterogeneity by Individual Risk
abstract
Surviva1 analysis is a widely used statistical approach to model and analysis time-to-event data, which consists of event occurrence and duration before it occurs. Traditional statistical methods have good interpretability but performance is limited by strong assumptions. In recent years, methods based on machine learning, especially deep learning, have achieved success, but the lack of interpretability limits their real-world applications. This paper proposes a novel interpretable deep learning survival analysis method, RESurv, which predicts the hazard directly without any priori assumptions. Experiments on multiple public datasets validate that RESurv outperforms or approaches state-of-the-art methods. We further estimate the inflence of exposure factors on prognosis at individual level by using counterfactual approaches. Patients are then divided into subgroups by hierarchical clustering. Significant differences in risk response patterns and prognosis suggest heterogeneity in their pathological mechanisms. We explore the risk factors between subpopulations identified by different characteristics and find distinct risk patterns among them, some of which are consistent with medical research findings. The results suggest the potential of RESurv to estimate individual risk, and thereby revealing the heterogeneity of risk response patterns across the populations.
Qiguang Zheng, Qifan Shen, Xin Su 0011, Kuo Yang 0001, Zixin Shu, Xuezhong Zhou
BIBM6
2022 TMNP: a transcriptome-based multi-scale network pharmacology platform for herbal medicine
abstract
One of the most difficult problems that hinder the development and application of herbal medicine is how to illuminate the global effects of herbs on the human body. Currently, the chemo-centric network pharmacology methodology regards herbs as a mixture of chemical ingredients and constructs the 'herb-compound-target-disease' connections based on bioinformatics methods, to explore the pharmacological effects of herbal medicine. However, this approach is severely affected by the complexity of the herbal composition. Alternatively, gene-expression profiles induced by herbal treatment reflect the overall biological effects of herbs and are suitable for studying the global effects of herbal medicine. Here, we develop an online transcriptome-based multi-scale network pharmacology platform (TMNP) for exploring the global effects of herbal medicine. Firstly, we build specific functional gene signatures for different biological scales from molecular to higher tissue levels. Then, specific algorithms are designed to measure the correlations of transcriptional profiles and types of gene signatures. Finally, TMNP uses pharmacotranscriptomics of herbal medicine as input and builds associations between herbs and different biological scales to explore the multi-scale effects of herb medicine. We applied TMNP to a single herb Astragalus membranaceus and Xuesaitong injection to demonstrate the power to reveal the multi-scale effects of herbal medicine. TMNP integrating herbal medicine and multiple biological scales into the same framework, will greatly extend the conventional network pharmacology model centering on the chemical components, and provide a window for systematically observing the complex interactions between herbal medicine and the human body. TMNP is available at http://www.bcxnfz.top/TMNP.
Peng Li 0004, Wuxia Zhang, Lingmin Zhan, Ning Wang 0048, Caiping Chen, Bangze Fu, Jinzhong Zhao, Xuezhong Zhou, Shuzhen Guo
Briefings Bioinform.10
2022 PDGNet: Predicting Disease Genes Using a Deep Neural Network With Multi-View Features
abstract
The knowledge of phenotype-genotype associations is crucial for the understanding of disease mechanisms. Numerous studies have focused on developing efficient and accurate computing approaches to predict disease genes. However, owing to the sparseness and complexity of medical data, developing an efficient deep neural network model to identify disease genes remains a huge challenge. Therefore, we develop a novel deep neural network model that fuses the multi-view features of phenotypes and genotypes to identify disease genes (termed PDGNet). Our model integrated the multi-view features of diseases and genes and leveraged the feedback information of training samples to optimize the parameters of deep neural network and obtain the deep vector features of diseases and genes. The evaluation experiments on a large data set indicated that PDGNet obtained higher performance than the state-of-the-art method (precision and recall improved by 9.55 and 9.63 percent). The analysis results for the candidate genes indicated that the predicted genes have strong functional homogeneity and dense interactions with known genes. We validated the top predicted genes of Parkinson's disease based on external curated data and published medical literatures, which indicated that the candidate genes have a huge potential to guide the selection of causal genes in the 'wet experiment'. The source codes and the data of PDGNet are available at https://github.com/yangkuoone/PDGNet.
Kuo Yang 0001, Kezhi Lu, Kai Chang, Ning Wang 0048, Zixin Shu, Jian Yu 0001, Baoyan Liu, Zhuye Gao, Xuezhong Zhou
IEEE ACM Trans. Comput. Biol. Bioinform.10
2021 TCMPR: TCM Prescription recommendation based on subnetwork term mapping and deep learning
abstract
Traditional Chinese medicine (TCM) has played an indispensable role in clinical diagnose and treatment. Based on patient’s symptom phenotypes, computation-based prescription recommendation methods can recommend personalized TCM prescription using machine learning and artificial intelligence technologies. However, owing to the complexity and individuation of patient’s clinical phenotypes, current prescription recommendation methods cannot obtain good performance. Meanwhile, it’s very difficult to conduct effective representation for unrecorded symptom terms in existing knowledge base. In this study, we proposed a subnetwork-based symptom term mapping method (SSTM), and constructed a SSTM-based TCM prescription recommendation method (termed TCMPR). Our SSTM can extract the subnetwork structure between symptoms from knowledge network to effectively represent the embedding features of clinical symptom terms (especially, the unrecorded terms). The experimental results showed that our method performs better than state-of-the-art methods. In addition, the comprehensive experiments of TCMPR with different hyper parameters (i.e., feature embedding, feature dimension and feature fusion) that demonstrates that our method has high performance on TCM prescription recommendation and potentially promote clinical diagnosis and treatment of TCM precision medicine.
Xin Dong 0017, Zixin Shu, Kai Chang, Dengying Yan, Jianan Xia, Kunyu Zhong, Xinyan Wang 0002, Kuo Yang 0001, Xuezhong Zhou
BIBM11
2021 COX-2 as the key target of LianXia NingXin Formula for treating Sympathetic Remodeling after Myocardial Infarction: pharmacological network prediction with experimental validation
abstract
Ethnopharmacological relevance: Previous studies have shown that LianXia NingXin (LXNX) formula, a Chinese herbal prescription, can prevent the sympathetic remodeling after myocardial infarction (MI). It is absolutely necessary to further investigate the pharmacological mechanism of action for LXNX formula to treat sympathetic remodeling after MI occurs. Material and methods: In this study, the most important protein target of LXNX formula to prevent sympathetic remodeling was identified by ranking the number of influencing ingredients (degree) of those targets in this network and how strong those targets had influence on sympathetic remodeling. Next, preclinical research was conducted on MI rat models treated with LXNX formula and celecoxib to verify this predicted pharmacological results. Celecoxib is a modern drug known to inhibit the target predicted through the method of the network analysis. After thirty day treatment, the protein and mRNA expression of this predicted protein in the peri-infarct zone of rats was measured. Results: The results of pharmacological network analysis showed that Cyclooxygenase-2 (COX-2) was predicted as the most important drug target of LXNX formula to treat sympathetic remodeling. Preclinical research indicated that both LXNX formula and celecoxib significantly reduced distribution of sympathetic remodeling around the peri-infarct zone of rats in pre-clinical research; also, the levels of COX-2 protein and its gene expression in the myocardium of rats post-MI were significantly decreased by LXNX formula and celecoxib. Related literature evidences demonstrated that LXNX formula is less likely to damage the gastrointestinal mucous membrane because the effects of its ingredients like curcumin and hesperidin to reduce cell apoptosis in the gastrointestinal tract. Conclusion: This study indicated that COX-2 is the key drug target of LXNX formula through which it inhibits adverse sympathetic nerve sprouting after MI.
Fangyuan Dai, Xuezhong Zhou, Ping Li 0063
BIBM5
2021 Phenonizer: A fine-grained phenotypic named entity recognizer for Chinese clinical texts
abstract
Biomedical named entity recognition from clinical texts is a fundamental task for clinical data analysis due to the availability of large volume of electronic medical record data, which are mostly in free text format, in real-world clinical settings. Clinical text data incorporates significant phenotypic medical entities, which could be used for profiling the clinical characteristics of patients in specific disease conditions. However, general approaches mostly rely on the coarse-grained annotations (e.g. mentions of symptom terms) of phenotypic entities in benchmark text dataset. Owing to the numerous negation expressions of phenotypic entities (e.g. “no fever”, “no cough” and “no hypertension”) in clinical texts, this could not feed the subsequent data analysis process with well-prepared structured clinical data. Thus, we constructed a fine-grained Chinese clinical corpus. Thereafter, we proposed a phenotypic named entity recognizer (Phenonizer). The results on the test set show that Phenonizer outperform those methods based on Word2Vec with Fl-score of 0.896. By comparing character embeddings from different data, it is found that character embeddings trained by clinical corpora can improve F-score by 0.0103. Furthermore, the fine-grained dataset enables methods to distinguish between negated symptoms and presented symptoms, and avoids the interference of negated symptoms. Finally, we tested the generalization performance of Phenonier, achieving a superior F1-score of 0.8389. In summary, together with fine-grained annotated benchmark dataset, Phenonier proposes a feasible approach to effectively extract symptom information from Chinese clinical texts with acceptable performance.
Qunsheng Zou, Kuo Yang 0001, Kai Chang, Xuezhong Zhou
BIBM6
2020 Analysis of diabetic comorbidities and their interrelationships in 5227 Chinese patients with type 2 diabetes
abstract
Aim: To explore the distribution and correlation of comorbidities in Chinese adult patients with type 2 diabetes mellitus(T2DM) by using real-world clinical medical record data. Methods: This study is a retrospective study, screening data from the previous medical records using integrated traditional Chinese and western medicine in the treatment of diabetes. In this study, descriptive statistic, association rules, complex network and other data mining methods were used to extract and analyze the associated data of diabetes mellitus comorbidities. Results: A total of 5227 clinical records of patients with type 2 diabetes were included in this study, and the top 10 comorbidities were identified. It was found that there was a correlation among concomitant diseases, and hypertension was most strongly associated with other concomitant diseases (degree > 248). It was also found that the distribution of comorbidities was closely related to age, and the distribution showed a certain rule, which is consistent with the understanding of the pathogenesis evolution of diabetes in TCM. Conclusion: Comorbidities of diabetes show associations with different intensity, and there are obviously characteristic disease groups in different age stages. Based on the results of this study, patients with diabetes can purposefully improve their awareness of preventing the related diseases. And it also provides useful reference for the prevention and treatment of diabetes and its associated diseases.
Min Pi, Runshun Zhang, Xiong He, Xuezhong Zhou, Motan Qiu, Tiancai Wen
BIBM6
2020 Network-based gene prediction for TCM symptoms
abstract
The diagnosis and treatment of traditional Chinese medicine (TCM) are formed based on the differentiation of syndromes and symptoms. Symptom management is always the core task of nursing science. Connotation between TCM symptoms and Modern medicine (MM) symptoms are obvious different, especially tongue and pulse symptoms of TCM. However, the underlying molecular mechanisms of most TCM symptoms remain unclear. Here, we developed a network-based framework to predict candidate genes of TCM symptoms (called PTsGene) and construction a high-quality set of TCM symptom-gene associations. Experimental results indicated that PTsGene performed significantly better than the baseline algorithms. The reliability of the candidate genes of symptoms (containing one of typical symptoms of COVID-19, fever) were validated by the analysis of functional homogeneity, molecular co-expression, and recently published literatures. Finally, a high-quality set of TCM symptom-gene associations is constructed to promote the mechanism developments of TCM symptoms. Prediction and construction for reliable TCM symptom-gene associations are valuable for uncovering the underlying molecular mechanisms of TCM symptoms. Our TCM symptom-gene associations deliver a highly insightful data sources for researchers both from basic and clinical settings of precision healthcare.
Yinyan Wang, Kuo Yang 0001, Zixin Shu, Dengying Yan, Xuezhong Zhou
BIBM5
2020 Disease phenotype synonymous prediction through network representation learning from PubMed database
abstract
Synonym mapping between phenotype concepts from different terminologies is difficult because terminology databases have been developed largely independently. Existing maps of synonymous phenotype concepts from different terminology databases are highly incomplete, and manually mapping is time consuming and laborious. Therefore, building an automatic method for predictive mapping of synonymous phenotypes is of special importance. We propose a classifier-based phenotype mapping prediction model (CPM) to predict synonymous relationships between phenotype concepts from different terminology databases. The model takes network semantic representations of phenotypes as input and predicts synonymous relationships by training binary classifiers with a voting strategy. We compared the performance of the CPM with a similarity-based phenotype mapping prediction model (SPM), which predicts mapping based on the ranked cosine similarity of candidate mapping concepts. Based on a network representation N2V-TFIDF, with a majority voting strategy method MV, the CPM achieved accuracy of 0.943, which was 15.4% higher than that of the SPM using the cosine similarity method (0.789) and 23.8% higher than that of the SSDTM method (0.724) proposed in our previous work.
Shiwen Ma, Kuo Yang 0001, Ning Wang 0048, Zhuye Gao, Runshun Zhang, Baoyan Liu, Xuezhong Zhou
Artif. Intell. Medicine8
2020 Integrated network analysis of symptom clusters across disease conditions
Kezhi Lu, Kuo Yang 0001, Edouard Niyongabo, Zixin Shu, Kai Chang, Qunsheng Zou, Jiyue Jiang, Caiyan Jia, Baoyan Liu, Xuezhong Zhou
J. Biomed. Informatics11
2019 HerGePred: Heterogeneous Network Embedding Representation for Disease Gene Prediction
abstract
The discovery of disease-causing genes is a critical step towards understanding the nature of a disease and determining a possible cure for it. In recent years, many computational methods to identify disease genes have been proposed. However, making full use of disease-related (e.g., symptoms) and gene-related (e.g., gene ontology and protein-protein interactions) information to improve the performance of disease gene prediction is still an issue. Here, we develop a heterogeneous disease-gene-related network (HDGN) embedding representation framework for disease gene prediction (called HerGePred). Based on this framework, a low-dimensional vector representation (LVR) of the nodes in the HDGN can be obtained. Then, we propose two specific algorithms, namely, an LVR-based similarity prediction and a random walk with restart on a reconstructed heterogeneous disease-gene network (RW-RDGN), to predict disease genes with high performance. First, to validate the rationality of the framework, we analyze the similarity-based overlap distribution of disease pairs and design an experiment for disease-gene association recovery, the results of which revealed that the LVR of nodes performs well at preserving the local and global network structure of the HDGN. Then, we apply tenfold cross validation and external validation to compare our methods with other well-known disease gene prediction algorithms. The experimental results show that the RW-RDGN performs better than the state-of-the-art algorithm. The prediction results of disease candidate genes are essential for molecular mechanism investigation and experimental validation. The source codes of HerGePred and experimental data are available at https://github.com/yangkuoone/HerGePred.
Kuo Yang 0001, Ruyu Wang, Zixin Shu, Ning Wang 0048, Runshun Zhang, Jian Yu 0001, Xuezhong Zhou
IEEE J. Biomed. Health Informatics10
2018 Discovery of Xuantoujiedu Decoction and its Molecular Mechanisms Using Integrated Network Analysis
Weilian Kong, Xuezhong Zhou, Runshun Zhang, Yanxing Xue
BIBM5
2018 Analysis of Disease Comorbidity Patterns in a Large-Scale China Population
Mengfei Guo, Tiancai Wen, Baoyan Liu, Jin Zhang 0044, Runshun Zhang, Yanning Zhang 0001, Xuezhong Zhou
ICIC (2)9
2018 Heterogeneous network embedding for identifying symptom candidate genes
abstract
Objective: Investigating the molecular mechanisms of symptoms is a vital task in precision medicine to refine disease taxonomy and improve the personalized management of chronic diseases. Although there are abundant experimental studies and computational efforts to obtain the candidate genes of diseases, the identification of symptom genes is rarely addressed. We curated a high-quality benchmark dataset of symptom-gene associations and proposed a heterogeneous network embedding for identifying symptom genes. Methods: We proposed a heterogeneous network embedding representation algorithm, which constructed a heterogeneous symptom-related network that integrated symptom-related associations and applied an embedding representation algorithm to obtain the low-dimensional vector representation of nodes. By measuring the relevance between symptoms and genes via calculating the similarities of their vectors, the candidate genes of given symptoms can be obtained. Results: A benchmark dataset of 18 270 symptom-gene associations between 505 symptoms and 4549 genes was curated. We compared our method to baseline algorithms (FSGER and PRINCE). The experimental results indicated our algorithm achieved a significant improvement over the state-of-the-art method, with precision and recall improved by 66.80% (0.844 vs 0.506) and 53.96% (0.311 vs 0.202), respectively, for TOP@3 and association precision improved by 37.71% (0.723 vs 0.525) over the PRINCE. Conclusions: The experimental validation of the algorithms and the literature validation of typical symptoms indicated our method achieved excellent performance. Hence, we curated a prediction dataset of 17 479 symptom-candidate genes. The benchmark and prediction datasets have the potential to promote investigations of the molecular mechanisms of symptoms and provide candidate genes for validation in experimental settings.
Kuo Yang 0001, Ning Wang 0048, Ruyu Wang, Jian Yu 0001, Runshun Zhang, Xuezhong Zhou
J. Am. Medical Informatics Assoc.8
2017 Framing Electronic Medical Records as Polylingual Documents in Query Expansion
Edward W. Huang, Sheng Wang 0012, Doris J. Lee, Runshun Zhang, Baoyan Liu, Xuezhong Zhou, ChengXiang Zhai
AMIA6
2017 An analysis of human microbe-disease associations
abstract
The microbiota living in the human body has critical impacts on our health and disease, but a systems understanding of its relationships with disease remains limited. Here, we use a large-scale text mining-based manually curated microbe-disease association data set to construct a microbe-based human disease network and investigate the relationships between microbes and disease genes, symptoms, chemical fragments and drugs. We reveal that microbe-based disease loops are significantly coherent. Microbe-based disease connections have strong overlaps with those constructed by disease genes, symptoms, chemical fragments and drugs. Moreover, we confirm that the microbe-based disease analysis is able to predict novel connections and mechanisms for disease, microbes, genes and drugs. The presented network, methods and findings can be a resource helpful for addressing some issues in medicine, for example, the discovery of bench knowledge and bedside clinical solutions for disease mechanism understanding, diagnosis and therapy.
Wei Ma 0010, Pan Zeng, Chuanbo Huang, Bin Geng, Jichun Yang, Xuezhong Zhou, Qinghua Cui
Briefings Bioinform.9
2016 Similarity-based algorithms for Disease Terminology Mapping
abstract
Classification of diseases and their related terms are important data resources for basic medical research. However, disease terms in different terminological databases are largely developed independent of each other and the mapping relationships between them are not complete. The purpose of this paper is to propose similarity-based disease terminology mapping methods to map disease terms with same or similar semantic concepts in different terminological databases. By integrating the bibliographic medical records from PubMed and the manually curated associations of disease-gene and disease-phenotype, we proposed two methods, namely Semantic-based Similarity for Disease Terminology Mapping (SSDTM) and Information Recommend-based Disease Terminology Mapping (IRDTM) to predict the disease terms in OMIM to MeSH. The experimental results show that both methods can support predicting mapping between disease databases. From leave one out cross validation, the prediction performance of SSDTM (Hits@10: 87.3%) is better than IRDTM; From manual evaluation, the hits rate in top 10 of SSDTM is 94.4%. The similarity-based disease terminology mapping method can be applied with the supplement of manual review for different mainstream disease terminology databases and help improve the efficiency of integrated medical ontology development and translational bioinformatics that need incorporate multiple data sources from different disciplines.
Shiwen Ma, Kuo Yang 0001, Xuezhong Zhou
BIBM3
2016 A conditional probabilistic model for joint analysis of symptoms, diseases, and herbs in traditional Chinese medicine patient records
abstract
Traditional Chinese medicine (TCM) can provide important complementary medical care to modern medicine, and is widely practiced in China and many other countries. Unfortunately, due to its empirical nature and history of trial and error, effective diagnosis and prescription methods are not well-defined. This setback results in a significant challenge in retaining, sharing, and inheriting knowledge among physicians. In this paper, we propose a new asymmetric probabilistic model for the joint analysis of symptoms, diseases, and herbs in patient records to discover and extract latent TCM knowledge. We base our model on the comprehensive evaluation of modern medicine and TCM-specific symptoms in addition to herb prescriptions for particular diseases. Experimental results on a large dataset demonstrate the effectiveness of the proposed model for discovering useful knowledge and its potential clinical applications.
Sheng Wang 0012, Edward W. Huang, Runshun Zhang, Baoyan Liu, Xuezhong Zhou, ChengXiang Zhai
BIBM6
2016 Extracting relations from traditional Chinese medicine literature via heterogeneous entity networks
abstract
OBJECTIVE: Traditional Chinese medicine (TCM) is a unique and complex medical system that has developed over thousands of years. This article studies the problem of automatically extracting meaningful relations of entities from TCM literature, for the purposes of assisting clinical treatment or poly-pharmacology research and promoting the understanding of TCM in Western countries. METHODS: Instead of separately extracting each relation from a single sentence or document, we propose to collectively and globally extract multiple types of relations (eg, herb-syndrome, herb-disease, formula-syndrome, formula-disease, and syndrome-disease relations) from the entire corpus of TCM literature, from the perspective of network mining. In our analysis, we first constructed heterogeneous entity networks from the TCM literature, in which each edge is a candidate relation, then used a heterogeneous factor graph model (HFGM) to simultaneously infer the existence of all the edges. We also employed a semi-supervised learning algorithm estimate the model's parameters. RESULTS: We performed our method to extract relations from a large dataset consisting of more than 100,000 TCM article abstracts. Our results show that the performance of the HFGM at extracting all types of relations from TCM literature was significantly better than a traditional support vector machine (SVM) classifier (increasing the average precision by 11.09%, the recall by 13.83%, and the F1-measure by 12.47% for different types of relations, compared with a traditional SVM classifier). CONCLUSION: This study exploits the power of collective inference and proposes an HFGM based on heterogeneous entity networks, which significantly improved our ability to extract relations from TCM literature.
Huaiyu Wan, Marie-Francine Moens, Walter Luyten, Xuezhong Zhou, Qiaozhu Mei, Lu Liu 0005, Jie Tang 0001
J. Am. Medical Informatics Assoc.4
2013 Integrating phenotype-genotype data for prioritization of candidate symptom genes
abstract
Symptoms and signs (symptoms in brief) are the essential clinical manifestations for traditional Chinese medicine (TCM) diagnosis and treatments. To gain insights into the molecular mechanism of symptoms, this paper presents a network-based data mining method to integrate multiple phenotype-genotype data sources and predict the prioritizing gene rank list of symptoms. The result of this pilot study suggested some insights on the molecular mechanism of symptoms.
Xuezhong Zhou, Yonghong Peng, Runshun Zhang, Jingqing Hu, Jian Yu 0001, Baoyan Liu
BIBM2
2013 Complex network approach for analyzing TCM clinical herb-symptom relationships
abstract
Traditional Chinese Medicine (TCM) is a discipline of clinical medicine, which focuses on individualized diagnosis and treatment based on observation of the clinical manifestations of real-world patients. The complicated interactions between different medical entities play significant role for individualized treatment. In this paper, we aim to find out the meaningful herb-symptom relationship from large number of clinical data with using complex network approach. We construct two different patient networks to verify the positive correlations between herbs and symptoms in TCM clinical treatment.
Xuezhong Zhou, Runshun Zhang, Jingqing Hu, Qi Xie 0005, Baoyan Liu
BIBM2
2013 Frequent itemsets compressing based on minimum cover: An efficient method for mining medication law of Chinese herbs
abstract
Frequent itemsets mining is often used to find medication law from dataset of Chinese herb prescriptions. Threshold of support count is difficult to set for traditional algorithm of frequent itemsets mining. In the meantime, the number of frequent itemsets is always so big that the result is hard to understand. Some algorithms were proposed to find significant and redundant-aware itemsets. However, the itemsets obtained could not reflect all the information in the dataset. In this paper, a new method was proposed to obtain a collection of itemsets which had the feature of significant, redundant-aware and comprehensive. Firstly, closed frequent itemsets were mined from the dataset of Chinese herbs prescriptions using CHARM algorithm. Then, the itemsets were compressed by FICMC (Frequent Itemsets Compressing based on Minimum Cover) algorithm. Medication law of Chinese herbs could be fully mined from the dataset using this method.
Yi-guo Wang, Qiming Zhang 0004, Xuezhong Zhou, Jian Yu 0001, Xiuhua Guo
BIBM4
2012 Multidimensional analysis for Traditional Chinese Medicine diagnosis and treatment on hepatitis diseases
abstract
Traditional Chinese Medicine (TCM) has been widely used to treat various diseases like infectious diseases. Treatment Based on Syndrome Differentiation (TBSD) is the main principle in TCM clinical practice. So exploring the relationships between diagnoses and treatments from successful cases is important and valuable for better treating hepatitis diseases, which is a decision support problem based on data warehouse in nature. In this paper, using the multidimensional analysis techniques of BusinessObjects (BO) platform, we introduce a series of online analytical processing (OLAP) reports which cover different subjects of TCM on hepatitis diseases and a corresponding system used to manage the reports. It has been found that these OLAP reports are useful in experience sharing of making diagnoses and giving treatments on hepatitis diseases.
Lanxin Bi, Xuezhong Zhou, Runshun Zhang
Healthcom2
2012 Using link topic model to analyze traditional Chinese Medicine Clinical symptom-herb regularities
abstract
Traditional Chinese Medicine (TCM) is a clinical medicine, which focuses on human physiology, pathology, diagnosis and treatment of diseases. Numerous clinical practice and theory research in the TCM field have accumulated huge amount of data. These data include TCM basic databases, TCM literature, as well as a large number of databases or data warehouse on TCM clinical diagnoses and treatment. More and more people pay attention to the discovery of hidden regularities of TCM clinical data. In recent years, topic model has been popularly used for text analysis and information retrieval by extracting latent and significant topics from corpus. In this paper, we apply the Link Latent Dirichlet Allocation (LinkLDA), to automatically extract the latent topic structures which contain the information of both symptoms and their corresponding herbs. By experimental results, the latent topic with symptoms and their corresponding herbs show clinical meaningful results. Furthermore, the model is also compared with other topic models, such as author-topic model, and the result of LinkLDA got better results.
Zaixing Jiang, Xuezhong Zhou
Healthcom2
2012 Clinical data preprocessing and case studies of POMDP for TCM treatment knowledge discovery
abstract
Partially Observable Markov Decision Processes (POMDP) has been applied to induce sequential treatment scheme from Traditional Chinese Medicinal (TCM) clinical data. The data required by POMDP should be of rich structure and with heterogeneous variables. But sometimes there is large number of missing values in the real-world TCM clinical data set. This makes it difficult for data preprocessing. This paper designs a data preprocessing framework of TCM clinical data for POMDP applications. It significantly facilitates the process of sequential treatment scheme discovery through POMDP when applying the framework on TCM clinical cases of coronary heart disease and lung cancer. We also systematically analyze the sequential treatment scheme.
Xuezhong Zhou
Healthcom2
2012 Enhanced data extraction, transforming and loading processing for Traditional Chinese Medicine clinical data warehouse
abstract
Clinical data warehouse has been developed as a fundamental data infrastructure for large scale TCM clinical data management and decision support services. However, as a key component, data extraction, transforming and loading (ETL) is a complicated and labor intensive task to ensure high data quality before all kinds of data analyses. This paper introduces an enhanced ETL technique framework, which includes operational data store (ODS) model and two step data preprocessing subcomponents, to perform the ETL tasks. The ODS data model was designed to integrate the heterogeneous clinical data sources and support the direct copy from these data sources to ODS database by ETL. Therefore, ETL task has been separated into two core steps in enhanced ETL component: (1) dynamic filter and copy of the original operational data sources to ODS; (2) specialized transforming the ODS data to detailed clinical data warehouse. This enhanced technique framework improves the ETL performance to be used in clinical data center since there would have various kinds of operational data sources that need be integrated in this data environments. This paper has a description of the related enhanced ETL framework and proposes some key procedures to accomplish the tasks.
Xishui Pan, Xuezhong Zhou, Hongmei Song, Runshun Zhang
Healthcom2
2012 Constructing ideas of health service platform for the elderly
abstract
The construction of health service platform for the elderly must attach great importance to the health service demands of the elderly ,basing on innovation and whole process of health service by taking full advantage of IOT technology, data warehousing, data mining analysis technology, cloud computing technology and other modern information technologies, widely applying the modern trans-regional remote health information collection and transmission equipment, and setting up health service technology platform with the close connection between production and research so as to enhance the ability and level of health service of the elderly.
Huaxin Shi, Qi Xie 0005, Baoyan Liu, Shusong Mao, Xuezhong Zhou
Healthcom6
2012 Real-world clinical data mining on TCM clinical diagnosis and treatment: A survey
abstract
This paper provides a survey of data mining methods that have been commonly applied to real-world TCM clinical data in recent years, and sets forth the requirements of data mining on real-world TCM clinical diagnosis and treatment data, in order to provide reference for better analyzing the syndrome differentiation and treatment principle hidden in the massive TCM clinical data in the future.
Xuezhong Zhou, Runshun Zhang, Baoyan Liu, Qi Xie 0005
Healthcom2
2010 Development of traditional Chinese medicine clinical data warehouse for medical knowledge discovery and decision support
Xuezhong Zhou, Baoyan Liu, Runsun Zhang, Ping Li 0063, Zhuye Gao, Xiufeng Yan
Artif. Intell. Medicine1
2010 Text mining for traditional Chinese medical knowledge discovery: A survey
Xuezhong Zhou, Yonghong Peng, Baoyan Liu
J. Biomed. Informatics1
2009 An Uncertainty-Based Belief Selection Method for POMDP Value Iteration
Xuezhong Zhou, Houkuan Huang
ECSQARU2
2007 Integrative mining of traditional Chinese medicine literature and MEDLINE for functional gene networks
Xuezhong Zhou, Baoyan Liu, Zhaohui Wu 0001, Yi Feng 0004
Artif. Intell. Medicine1
2006 Knowledge discovery in traditional Chinese medicine: State of the art and perspectives
Yi Feng 0004, Zhaohui Wu 0001, Xuezhong Zhou, Zhongmei Zhou, Weiyu Fan
Artif. Intell. Medicine3
2005 Text Mining for Clinical Chinese Herbal Medical Knowledge Discovery
Xuezhong Zhou, Baoyan Liu, Zhaohui Wu 0001
Discovery Science1
2004 Text Mining for Finding Functional Community of Related Genes Using TCM Knowledge
Zhaohui Wu 0001, Xuezhong Zhou, Baoyan Liu, Junli Chen
PKDD2
2004 Distributional Character Clustering for Chinese Text Categorization
Xuezhong Zhou, Zhaohui Wu 0001
PRICAI1
2004 Ontology development for unified traditional Chinese medical language system
Xuezhong Zhou, Zhaohui Wu 0001, Aining Yin, Lancheng Wu, Weiyu Fan, Ruen Zhang
Artif. Intell. Medicine1
2001 TCMMDB: a distributed multidatabase query system and its key technique implemention
abstract
With the amount of information exploding, the number of databases and information stored in the database increases very quickly. Those databases are more likely to be heterogeneous. There are many distributed databases systems to resolve this problem, but the main drawback is that they don't do well in database cooperation. We provide a general architecture of a distributed multidatabase query system: TCMMDB (traditional Chinese medicine multidatabase). We discuss key techniques such as distributed query plan generation, distributed query synchronization and distributed error handling in TCMMDB in relation to distributed databases.
Xuezhong Zhou, Zhaohui Wu 0001
SMC1