VLDB 2026 Research / reviewers in the wild / expert
Runshun Zhang
dblp:87/10614
· DBLP profile ↗
20ranked-venue papers
0as first author
5since 2021 · last 2024
0000-0002-7127-8865ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 18 · 5 since 2021Artificial intelligence and machine learning · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | TCM-FTP: Fine-Tuning Large Language Models for Herbal Prescription PredictionabstractTraditional Chinese medicine (TCM) has relied on specific combinations of herbs in prescriptions to treat various symptoms and signs for thousands of years. Predicting TCM prescriptions poses a fascinating technical challenge with significant practical implications. However, this task faces limitations due to the scarcity of high-quality clinical datasets and the complex relationship between symptoms and herbs. To address these issues, we introduce DigestDS, a novel dataset comprising practical medical records from experienced experts in digestive system diseases. We also propose a method, TCM-FTP (TCM Fine-Tuning Pre-trained), to leverage pre-trained large language models (LLMs) via supervised fine-tuning on DigestDS. Additionally, we enhance computational efficiency using a low-rank adaptation technique. Moreover, TCM-FTP incorporates data augmentation by permuting herbs within prescriptions, exploiting their order-agnostic nature. Impressively, TCM-FTP achieves an F1-score of 0.8031, significantly outperforming previous methods. Furthermore, it demonstrates remarkable accuracy in dosage prediction, achieving a normalized mean square error of 0.0604. In contrast, LLMs without fine-tuning exhibit poor performance. Although LLMs have demonstrated wide-ranging capabilities, our work underscores the necessity of fine-tuning for TCM prescription prediction and presents an effective way to accomplish this. Xingzhi Zhou 0002, Xin Dong 0017, Chunhao Li, Yuning Bai, Ka Chun Cheung, Simon See, Xinpeng Song, Runshun Zhang, Xuezhong Zhou, Nevin Lianwen Zhang |
BIBM | 9 |
| 2024 | PresRecCD: A Novel Herbal Prescription Recommendation Framework with Cross-Domain Learning and Neural Collaborative FilteringabstractHerbal prescriptions hold significant importance in Traditional Chinese Medicine (TCM) diagnosis and treatment, embodying millennia of clinical case summaries and wisdom. Despite numerous proposed methods for herbal prescription recommendation (HPR), significant challenges persist due to the lack of comprehensive clinical data, particularly regarding the relationships between symptoms and herbs. This scarcity poses considerable hurdles for effective HPR modeling. In this study, we introduced a novel herbal prescription recommendation framework with cross-domain learning and neural collaborative filtering (termed PresRecCD). The cross-domain learning mechanism is introduced to learn the noise-reduced cross-domain features of herbs and symptoms in the unified space, alleviating the sparsity of data, and neural collaborative filtering is utilized to carry out prescription recommendations. Comprehensive experiments demonstrate the superiority of the proposed PresRecCD model over the SOTA model. This study contributes to enhancing the performance of the HPR model, ultimately benefiting the efficiency and precision of clinical treatment. Wansong Zhang, Xin Dong 0017, Kuo Yang 0001, Rouye Huang, Runshun Zhang, Xuezhong Zhou |
BIBM | 7 |
| 2024 | Lingdan: enhancing encoding of traditional Chinese medicine knowledge for clinical reasoning tasks with large language modelsabstractOBJECTIVE: The recent surge in large language models (LLMs) across various fields has yet to be fully realized in traditional Chinese medicine (TCM). This study aims to bridge this gap by developing a large language model tailored to TCM knowledge, enhancing its performance and accuracy in clinical reasoning tasks such as diagnosis, treatment, and prescription recommendations. MATERIALS AND METHODS: This study harnessed a wide array of TCM data resources, including TCM ancient books, textbooks, and clinical data, to create 3 key datasets: the TCM Pre-trained Dataset, the Traditional Chinese Patent Medicine (TCPM) Question Answering Dataset, and the Spleen and Stomach Herbal Prescription Recommendation Dataset. These datasets underpinned the development of the Lingdan Pre-trained LLM and 2 specialized models: the Lingdan-TCPM-Chat Model, which uses a Chain-of-Thought process for symptom analysis and TCPM recommendation, and a Lingdan Prescription Recommendation model (Lingdan-PR) that proposes herbal prescriptions based on electronic medical records. RESULTS: The Lingdan-TCPM-Chat and the Lingdan-PR Model, fine-tuned on the Lingdan Pre-trained LLM, demonstrated state-of-the art performances for the tasks of TCM clinical knowledge answering and herbal prescription recommendation. Notably, Lingdan-PR outperformed all state-of-the-art baseline models, achieving an improvement of 18.39% in the Top@20 F1-score compared with the best baseline. CONCLUSION: This study marks a pivotal step in merging advanced LLMs with TCM, showcasing the potential of artificial intelligence to help improve clinical decision-making of medical diagnostics and treatment strategies. The success of the Lingdan Pre-trained LLM and its derivative models, Lingdan-TCPM-Chat and Lingdan-PR, not only revolutionizes TCM practices but also opens new avenues for the application of artificial intelligence in other specialized medical fields. Our project is available at https://github.com/TCMAI-BJTU/LingdanLLM. Xin Dong 0016, Zixin Shu, Yunhui Hu, Shuiping Zhou, Kaijing Yan, Xijun Yan, Kai Chang, Yuning Bai, Runshun Zhang, Xuezhong Zhou |
J. Am. Medical Informatics Assoc. | 14 |
| 2023 | DRONet: effectiveness-driven drug repositioning framework using network embedding and ranking learningabstractAs one of the most vital methods in drug development, drug repositioning emphasizes further analysis and research of approved drugs based on the existing large amount of clinical and experimental data to identify new indications of drugs. However, the existing drug repositioning methods didn't achieve enough prediction performance, and these methods do not consider the effectiveness information of drugs, which make it difficult to obtain reliable and valuable results. In this study, we proposed a drug repositioning framework termed DRONet, which make full use of effectiveness comparative relationships (ECR) among drugs as prior information by combining network embedding and ranking learning. We utilized network embedding methods to learn the deep features of drugs from a heterogeneous drug-disease network, and constructed a high-quality drug-indication data set including effectiveness-based drug contrast relationships. The embedding features and ECR of drugs are combined effectively through a designed ranking learning model to prioritize candidate drugs. Comprehensive experiments show that DRONet has higher prediction accuracy (improving 87.4% on Hit@1 and 37.9% on mean reciprocal rank) than state of the art. The case analysis also demonstrates high reliability of predicted results, which has potential to guide clinical drug development. Kuo Yang 0001, Yuxia Yang, Shuyue Fan, Jianan Xia, Qiguang Zheng, Xin Dong 0017, Zhuye Gao, Runshun Zhang, Baoyan Liu, Xuezhong Zhou |
Briefings Bioinform. | 13 |
| 2022 | Research and Implementation of Real World Traditional Chinese Medicine Clinical Scientific Research Information Electronic Medical Record Sharing SystemabstractThe standardization degree of traditional Chinese medicine clinical data in the real world is low and heterogeneous data aggregation among institutions is difficult, which leads to the difficulty of sharing clinical data and scientific research data of traditional Chinese medicine. This paper designs and implements a real world traditional Chinese medicine clinical scientific research information electronic medical record sharing system. The system consists mainly of two subsystems, namely electronic medical record collection system and electronic medical record integration system. The collection system can collect, normalize and structured storage inpatient electronic medical records, outpatient electronic medical records and cloud platform electronic medical records. The integration system can integrate heterogeneous data from different traditional Chinese medicine diagnostic and treatment institutions to realize the sharing of traditional Chinese medicine clinical and research data. Qi Xie 0005, Runshun Zhang, Xuezhong Zhou, Tiancai Wen, Xingping Zhang, Xu Miao, Baoyan Liu |
BIBM | 3 |
| 2020 | Analysis of diabetic comorbidities and their interrelationships in 5227 Chinese patients with type 2 diabetesabstractAim: To explore the distribution and correlation of comorbidities in Chinese adult patients with type 2 diabetes mellitus(T2DM) by using real-world clinical medical record data. Methods: This study is a retrospective study, screening data from the previous medical records using integrated traditional Chinese and western medicine in the treatment of diabetes. In this study, descriptive statistic, association rules, complex network and other data mining methods were used to extract and analyze the associated data of diabetes mellitus comorbidities. Results: A total of 5227 clinical records of patients with type 2 diabetes were included in this study, and the top 10 comorbidities were identified. It was found that there was a correlation among concomitant diseases, and hypertension was most strongly associated with other concomitant diseases (degree > 248). It was also found that the distribution of comorbidities was closely related to age, and the distribution showed a certain rule, which is consistent with the understanding of the pathogenesis evolution of diabetes in TCM. Conclusion: Comorbidities of diabetes show associations with different intensity, and there are obviously characteristic disease groups in different age stages. Based on the results of this study, patients with diabetes can purposefully improve their awareness of preventing the related diseases. And it also provides useful reference for the prevention and treatment of diabetes and its associated diseases. Min Pi, Runshun Zhang, Xiong He, Xuezhong Zhou, Motan Qiu, Tiancai Wen |
BIBM | 4 |
| 2020 | Syndrome Evolution and Chinese Herb Formula Regularity of TCM Heat Syndrome in Type 2 Diabetes Mellitus: Complex Network Community Discovery Algorithm and Sankey Diagram VisualizationabstractObjective: Analyzed the evolution of heat syndrome of type 2 diabetes mellitus (T2DM) and the characteristics of Chinese herb formula, in order to provide reference for guiding the clinical practice of T2DM. Methods: Based on the electronic medical record of 2826 patients with T2DM, complex network community discovery algorithm, Sankey diagram and multi-scale backbone network analysis were used to mine the syndrome evolution of different visitand the Chinese herb formula different syndrome groups. Results: In this paper, the complex network of T2DM syndrome was divided into seven main communities (represented by A~G). Among them, three heat syndrome communities were spleen deficiency and stomach heat syndrome (community B, 16.38% nodes of network), phlegm-heat syndrome (community C, 16.1%) and damp-heat syndrome (community D, 10.17%), accounting for 47.04%. Most of the patients diagnosed with group B, C and D had no change in syndrome for a long-term visit, followed by conversion to group A. Specifically, patients in group B were easier to transform to syndrome community A and C, while group C were likely to transform to syndrome community A and B. Community D had potential to turn into syndrome community A, C, B. The core herb of T2DM heat syndrome was Coptidis Rhizoma, and the mostly herb compatible with it were Zingiberis Rhizoma, Pinelliae Rhizoma, Rehmanniae Radix, Scutellariae Radix, Phellodendri Chinensis Cortex, Anemarrhenae Rhizoma, et al. Conclusion: Qi-yin deficiency syndrome is the common evolution direction of T2DM heat syndrome, and The traditional Chinese medicine commonly used in T2DM heat syndrome is Coptidis Rhizoma. Min Pi, Runshun Zhang, Tiancai Wen |
BIBM | 4 |
| 2020 | Disease phenotype synonymous prediction through network representation learning from PubMed databaseabstractSynonym mapping between phenotype concepts from different terminologies is difficult because terminology databases have been developed largely independently. Existing maps of synonymous phenotype concepts from different terminology databases are highly incomplete, and manually mapping is time consuming and laborious. Therefore, building an automatic method for predictive mapping of synonymous phenotypes is of special importance. We propose a classifier-based phenotype mapping prediction model (CPM) to predict synonymous relationships between phenotype concepts from different terminology databases. The model takes network semantic representations of phenotypes as input and predicts synonymous relationships by training binary classifiers with a voting strategy. We compared the performance of the CPM with a similarity-based phenotype mapping prediction model (SPM), which predicts mapping based on the ranked cosine similarity of candidate mapping concepts. Based on a network representation N2V-TFIDF, with a majority voting strategy method MV, the CPM achieved accuracy of 0.943, which was 15.4% higher than that of the SPM using the cosine similarity method (0.789) and 23.8% higher than that of the SSDTM method (0.724) proposed in our previous work. Shiwen Ma, Kuo Yang 0001, Ning Wang 0048, Zhuye Gao, Runshun Zhang, Baoyan Liu, Xuezhong Zhou |
Artif. Intell. Medicine | 6 |
| 2019 | HerGePred: Heterogeneous Network Embedding Representation for Disease Gene PredictionabstractThe discovery of disease-causing genes is a critical step towards understanding the nature of a disease and determining a possible cure for it. In recent years, many computational methods to identify disease genes have been proposed. However, making full use of disease-related (e.g., symptoms) and gene-related (e.g., gene ontology and protein-protein interactions) information to improve the performance of disease gene prediction is still an issue. Here, we develop a heterogeneous disease-gene-related network (HDGN) embedding representation framework for disease gene prediction (called HerGePred). Based on this framework, a low-dimensional vector representation (LVR) of the nodes in the HDGN can be obtained. Then, we propose two specific algorithms, namely, an LVR-based similarity prediction and a random walk with restart on a reconstructed heterogeneous disease-gene network (RW-RDGN), to predict disease genes with high performance. First, to validate the rationality of the framework, we analyze the similarity-based overlap distribution of disease pairs and design an experiment for disease-gene association recovery, the results of which revealed that the LVR of nodes performs well at preserving the local and global network structure of the HDGN. Then, we apply tenfold cross validation and external validation to compare our methods with other well-known disease gene prediction algorithms. The experimental results show that the RW-RDGN performs better than the state-of-the-art algorithm. The prediction results of disease candidate genes are essential for molecular mechanism investigation and experimental validation. The source codes of HerGePred and experimental data are available at https://github.com/yangkuoone/HerGePred. Kuo Yang 0001, Ruyu Wang, Zixin Shu, Ning Wang 0048, Runshun Zhang, Jian Yu 0001, Xuezhong Zhou |
IEEE J. Biomed. Health Informatics | 6 |
| 2018 | Discovery of Xuantoujiedu Decoction and its Molecular Mechanisms Using Integrated Network Analysis
Weilian Kong, Xuezhong Zhou, Runshun Zhang, Yanxing Xue |
BIBM | 6 |
| 2018 | Analysis of Disease Comorbidity Patterns in a Large-Scale China Population
Mengfei Guo, Tiancai Wen, Baoyan Liu, Jin Zhang 0044, Runshun Zhang, Yanning Zhang 0001, Xuezhong Zhou |
ICIC (2) | 7 |
| 2018 | Heterogeneous network embedding for identifying symptom candidate genesabstractObjective: Investigating the molecular mechanisms of symptoms is a vital task in precision medicine to refine disease taxonomy and improve the personalized management of chronic diseases. Although there are abundant experimental studies and computational efforts to obtain the candidate genes of diseases, the identification of symptom genes is rarely addressed. We curated a high-quality benchmark dataset of symptom-gene associations and proposed a heterogeneous network embedding for identifying symptom genes. Methods: We proposed a heterogeneous network embedding representation algorithm, which constructed a heterogeneous symptom-related network that integrated symptom-related associations and applied an embedding representation algorithm to obtain the low-dimensional vector representation of nodes. By measuring the relevance between symptoms and genes via calculating the similarities of their vectors, the candidate genes of given symptoms can be obtained. Results: A benchmark dataset of 18 270 symptom-gene associations between 505 symptoms and 4549 genes was curated. We compared our method to baseline algorithms (FSGER and PRINCE). The experimental results indicated our algorithm achieved a significant improvement over the state-of-the-art method, with precision and recall improved by 66.80% (0.844 vs 0.506) and 53.96% (0.311 vs 0.202), respectively, for TOP@3 and association precision improved by 37.71% (0.723 vs 0.525) over the PRINCE. Conclusions: The experimental validation of the algorithms and the literature validation of typical symptoms indicated our method achieved excellent performance. Hence, we curated a prediction dataset of 17 479 symptom-candidate genes. The benchmark and prediction datasets have the potential to promote investigations of the molecular mechanisms of symptoms and provide candidate genes for validation in experimental settings. Kuo Yang 0001, Ning Wang 0048, Ruyu Wang, Jian Yu 0001, Runshun Zhang, Xuezhong Zhou |
J. Am. Medical Informatics Assoc. | 6 |
| 2017 | Framing Electronic Medical Records as Polylingual Documents in Query Expansion
Edward W. Huang, Sheng Wang 0012, Doris J. Lee, Runshun Zhang, Baoyan Liu, Xuezhong Zhou, ChengXiang Zhai |
AMIA | 4 |
| 2016 | A conditional probabilistic model for joint analysis of symptoms, diseases, and herbs in traditional Chinese medicine patient recordsabstractTraditional Chinese medicine (TCM) can provide important complementary medical care to modern medicine, and is widely practiced in China and many other countries. Unfortunately, due to its empirical nature and history of trial and error, effective diagnosis and prescription methods are not well-defined. This setback results in a significant challenge in retaining, sharing, and inheriting knowledge among physicians. In this paper, we propose a new asymmetric probabilistic model for the joint analysis of symptoms, diseases, and herbs in patient records to discover and extract latent TCM knowledge. We base our model on the comprehensive evaluation of modern medicine and TCM-specific symptoms in addition to herb prescriptions for particular diseases. Experimental results on a large dataset demonstrate the effectiveness of the proposed model for discovering useful knowledge and its potential clinical applications. Sheng Wang 0012, Edward W. Huang, Runshun Zhang, Baoyan Liu, Xuezhong Zhou, ChengXiang Zhai |
BIBM | 3 |
| 2013 | Integrating phenotype-genotype data for prioritization of candidate symptom genesabstractSymptoms and signs (symptoms in brief) are the essential clinical manifestations for traditional Chinese medicine (TCM) diagnosis and treatments. To gain insights into the molecular mechanism of symptoms, this paper presents a network-based data mining method to integrate multiple phenotype-genotype data sources and predict the prioritizing gene rank list of symptoms. The result of this pilot study suggested some insights on the molecular mechanism of symptoms. Xuezhong Zhou, Yonghong Peng, Runshun Zhang, Jingqing Hu, Jian Yu 0001, Baoyan Liu |
BIBM | 4 |
| 2013 | Complex network approach for analyzing TCM clinical herb-symptom relationshipsabstractTraditional Chinese Medicine (TCM) is a discipline of clinical medicine, which focuses on individualized diagnosis and treatment based on observation of the clinical manifestations of real-world patients. The complicated interactions between different medical entities play significant role for individualized treatment. In this paper, we aim to find out the meaningful herb-symptom relationship from large number of clinical data with using complex network approach. We construct two different patient networks to verify the positive correlations between herbs and symptoms in TCM clinical treatment. Xuezhong Zhou, Runshun Zhang, Jingqing Hu, Qi Xie 0005, Baoyan Liu |
BIBM | 3 |
| 2012 | Co-evolution of symptom-herb relationshipabstractTraditional Chinese Medicine (TCM) is a complementary alternative medical approach. Its holistic approach is drastically different from the western medicine (WM). Upon the gathering of various symptoms in a diagnosis, a TCM practitioner prescribes treatment methods, of which herbal medicine is still one of the most popular. Each formula consists of multiple herbs. Since it is not a one-to-one mapping between symptom and herb, overlapping subsets of herbs are meant to address sets of overlapping symptoms. As a result, the discovery of the symptoms-herbs relationship is a crucial step to the research of the underlying TCM principle. The discovery of many existing formulas took a long time to stabilize to the current configurations. In this paper, the relationship discovery is argued to be more than just an evolutionary process, but a coevolutionary process, i.e. a set of symptoms searches for candidate sets of herbs, while a given set of herbs are appropriate for multiple sets of symptoms. In other words, a well recognized symptoms-herbs relationship is the result of a dynamic equilibrium of two inter-related evolutionary processes. This model of discovery was implemented using a Combined Gene Genetic Algorithm (CoGA1) where the symptoms and herbs are encoded in the same chromosome to evolve over time. The algorithm was tested with an insomnia dataset from a TCM hospital. The algorithm was able to find the symptoms-herbs relationships that are consistent with TCM principles and have better fitness from Simple GA. Josiah Poon, Dawei Yin 0002, Simon K. Poon, Runshun Zhang, Baoyan Liu, Daniel Man-yuen Sze |
IEEE Congress on Evolutionary Computation | 4 |
| 2012 | Multidimensional analysis for Traditional Chinese Medicine diagnosis and treatment on hepatitis diseasesabstractTraditional Chinese Medicine (TCM) has been widely used to treat various diseases like infectious diseases. Treatment Based on Syndrome Differentiation (TBSD) is the main principle in TCM clinical practice. So exploring the relationships between diagnoses and treatments from successful cases is important and valuable for better treating hepatitis diseases, which is a decision support problem based on data warehouse in nature. In this paper, using the multidimensional analysis techniques of BusinessObjects (BO) platform, we introduce a series of online analytical processing (OLAP) reports which cover different subjects of TCM on hepatitis diseases and a corresponding system used to manage the reports. It has been found that these OLAP reports are useful in experience sharing of making diagnoses and giving treatments on hepatitis diseases. Lanxin Bi, Xuezhong Zhou, Runshun Zhang |
Healthcom | 4 |
| 2012 | Enhanced data extraction, transforming and loading processing for Traditional Chinese Medicine clinical data warehouseabstractClinical data warehouse has been developed as a fundamental data infrastructure for large scale TCM clinical data management and decision support services. However, as a key component, data extraction, transforming and loading (ETL) is a complicated and labor intensive task to ensure high data quality before all kinds of data analyses. This paper introduces an enhanced ETL technique framework, which includes operational data store (ODS) model and two step data preprocessing subcomponents, to perform the ETL tasks. The ODS data model was designed to integrate the heterogeneous clinical data sources and support the direct copy from these data sources to ODS database by ETL. Therefore, ETL task has been separated into two core steps in enhanced ETL component: (1) dynamic filter and copy of the original operational data sources to ODS; (2) specialized transforming the ODS data to detailed clinical data warehouse. This enhanced technique framework improves the ETL performance to be used in clinical data center since there would have various kinds of operational data sources that need be integrated in this data environments. This paper has a description of the related enhanced ETL framework and proposes some key procedures to accomplish the tasks. Xishui Pan, Xuezhong Zhou, Hongmei Song, Runshun Zhang |
Healthcom | 4 |
| 2012 | Real-world clinical data mining on TCM clinical diagnosis and treatment: A surveyabstractThis paper provides a survey of data mining methods that have been commonly applied to real-world TCM clinical data in recent years, and sets forth the requirements of data mining on real-world TCM clinical diagnosis and treatment data, in order to provide reference for better analyzing the syndrome differentiation and treatment principle hidden in the massive TCM clinical data in the future. Xuezhong Zhou, Runshun Zhang, Baoyan Liu, Qi Xie 0005 |
Healthcom | 3 |