VLDB 2026 Research / reviewers in the wild / expert
Zixin Shu
dblp:242/2116
· DBLP profile ↗
9ranked-venue papers
0as first author
6since 2021 · last 2026
0000-0001-6078-5645ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 8 · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Meta Relation-Aware Metric Learning Framework for Few-Shot Uncertain Knowledge Graph CompletionabstractUncertainty is a fundamental aspect of human knowledge, widely present in various domains such as medicine, biology, and finance. Uncertain Knowledge Graph Completion (UKGC) tasks encompass both link prediction and confidence prediction. However, the inaccessibility of high-quality knowledge results in sparse knowledge graphs, where a large number of entities exhibit relatively sparse neighborhoods. This sparsity severely hinders the performance of few-shot UKGC models, which generally rely on neighborhood aggregation to enhance entity representations. In this study, we propose a Meta relation-aware Metric learning framework designed for few-shot UKGC, named MMUC. To mitigate neighbor sparsity, MMUC introduces a relation-centric paradigm that decouples representation enhancement from background graphs, integrating an uncertainty-driven similarity measure through a learnable confidence learner to achieve robust, non-linear fusion for joint link and confidence prediction. Extensive evaluations across multiple benchmarks demonstrate MMUC's superior performance through rigorous link and confidence prediction, ablation studies, and sparsity analysis, highlighting its particular resilience to neighbor sparsity in few-shot uncertain KGs. The source code can be accessed inhttps://github.com/sienna-wxy/MMUC. Xinyan Wang 0002, Kuo Yang 0001, Xianan Li, Zixin Shu, Xuezhong Zhou |
IEEE Trans. Knowl. Data Eng. | 7 |
| 2024 | KDGene: knowledge graph completion for disease gene prediction using interactional tensor decompositionabstractThe accurate identification of disease-associated genes is crucial for understanding the molecular mechanisms underlying various diseases. Most current methods focus on constructing biological networks and utilizing machine learning, particularly deep learning, to identify disease genes. However, these methods overlook complex relations among entities in biological knowledge graphs. Such information has been successfully applied in other areas of life science research, demonstrating their effectiveness. Knowledge graph embedding methods can learn the semantic information of different relations within the knowledge graphs. Nonetheless, the performance of existing representation learning techniques, when applied to domain-specific biological data, remains suboptimal. To solve these problems, we construct a biological knowledge graph centered on diseases and genes, and develop an end-to-end knowledge graph completion framework for disease gene prediction using interactional tensor decomposition named KDGene. KDGene incorporates an interaction module that bridges entity and relation embeddings within tensor decomposition, aiming to improve the representation of semantically similar concepts in specific domains and enhance the ability to accurately predict disease genes. Experimental results show that KDGene significantly outperforms state-of-the-art algorithms, whether existing disease gene prediction methods or knowledge graph embedding methods for general domains. Moreover, the comprehensive biological analysis of the predicted results further validates KDGene's capability to accurately identify new candidate genes. This work proposes a scalable knowledge graph completion framework to identify disease candidate genes, from which the results are promising to provide valuable references for further wet experiments. Data and source codes are available at https://github.com/2020MEAI/KDGene. Xinyan Wang 0002, Kuo Yang 0001, Ting Jia, Fanghui Gu, Chongyu Wang, Zixin Shu, Jianan Xia, Xuezhong Zhou |
Briefings Bioinform. | 7 |
| 2024 | Lingdan: enhancing encoding of traditional Chinese medicine knowledge for clinical reasoning tasks with large language modelsabstractOBJECTIVE: The recent surge in large language models (LLMs) across various fields has yet to be fully realized in traditional Chinese medicine (TCM). This study aims to bridge this gap by developing a large language model tailored to TCM knowledge, enhancing its performance and accuracy in clinical reasoning tasks such as diagnosis, treatment, and prescription recommendations. MATERIALS AND METHODS: This study harnessed a wide array of TCM data resources, including TCM ancient books, textbooks, and clinical data, to create 3 key datasets: the TCM Pre-trained Dataset, the Traditional Chinese Patent Medicine (TCPM) Question Answering Dataset, and the Spleen and Stomach Herbal Prescription Recommendation Dataset. These datasets underpinned the development of the Lingdan Pre-trained LLM and 2 specialized models: the Lingdan-TCPM-Chat Model, which uses a Chain-of-Thought process for symptom analysis and TCPM recommendation, and a Lingdan Prescription Recommendation model (Lingdan-PR) that proposes herbal prescriptions based on electronic medical records. RESULTS: The Lingdan-TCPM-Chat and the Lingdan-PR Model, fine-tuned on the Lingdan Pre-trained LLM, demonstrated state-of-the art performances for the tasks of TCM clinical knowledge answering and herbal prescription recommendation. Notably, Lingdan-PR outperformed all state-of-the-art baseline models, achieving an improvement of 18.39% in the Top@20 F1-score compared with the best baseline. CONCLUSION: This study marks a pivotal step in merging advanced LLMs with TCM, showcasing the potential of artificial intelligence to help improve clinical decision-making of medical diagnostics and treatment strategies. The success of the Lingdan Pre-trained LLM and its derivative models, Lingdan-TCPM-Chat and Lingdan-PR, not only revolutionizes TCM practices but also opens new avenues for the application of artificial intelligence in other specialized medical fields. Our project is available at https://github.com/TCMAI-BJTU/LingdanLLM. Xin Dong 0016, Zixin Shu, Yunhui Hu, Shuiping Zhou, Kaijing Yan, Xijun Yan, Kai Chang, Yuning Bai, Runshun Zhang, Xuezhong Zhou |
J. Am. Medical Informatics Assoc. | 4 |
| 2022 | RESurv: A Deep Survival Analysis Model to Reveal Population Heterogeneity by Individual RiskabstractSurviva1 analysis is a widely used statistical approach to model and analysis time-to-event data, which consists of event occurrence and duration before it occurs. Traditional statistical methods have good interpretability but performance is limited by strong assumptions. In recent years, methods based on machine learning, especially deep learning, have achieved success, but the lack of interpretability limits their real-world applications. This paper proposes a novel interpretable deep learning survival analysis method, RESurv, which predicts the hazard directly without any priori assumptions. Experiments on multiple public datasets validate that RESurv outperforms or approaches state-of-the-art methods. We further estimate the inflence of exposure factors on prognosis at individual level by using counterfactual approaches. Patients are then divided into subgroups by hierarchical clustering. Significant differences in risk response patterns and prognosis suggest heterogeneity in their pathological mechanisms. We explore the risk factors between subpopulations identified by different characteristics and find distinct risk patterns among them, some of which are consistent with medical research findings. The results suggest the potential of RESurv to estimate individual risk, and thereby revealing the heterogeneity of risk response patterns across the populations. Qiguang Zheng, Qifan Shen, Xin Su 0011, Kuo Yang 0001, Zixin Shu, Xuezhong Zhou |
BIBM | 5 |
| 2022 | PDGNet: Predicting Disease Genes Using a Deep Neural Network With Multi-View FeaturesabstractThe knowledge of phenotype-genotype associations is crucial for the understanding of disease mechanisms. Numerous studies have focused on developing efficient and accurate computing approaches to predict disease genes. However, owing to the sparseness and complexity of medical data, developing an efficient deep neural network model to identify disease genes remains a huge challenge. Therefore, we develop a novel deep neural network model that fuses the multi-view features of phenotypes and genotypes to identify disease genes (termed PDGNet). Our model integrated the multi-view features of diseases and genes and leveraged the feedback information of training samples to optimize the parameters of deep neural network and obtain the deep vector features of diseases and genes. The evaluation experiments on a large data set indicated that PDGNet obtained higher performance than the state-of-the-art method (precision and recall improved by 9.55 and 9.63 percent). The analysis results for the candidate genes indicated that the predicted genes have strong functional homogeneity and dense interactions with known genes. We validated the top predicted genes of Parkinson's disease based on external curated data and published medical literatures, which indicated that the candidate genes have a huge potential to guide the selection of causal genes in the 'wet experiment'. The source codes and the data of PDGNet are available at https://github.com/yangkuoone/PDGNet. Kuo Yang 0001, Kezhi Lu, Kai Chang, Ning Wang 0048, Zixin Shu, Jian Yu 0001, Baoyan Liu, Zhuye Gao, Xuezhong Zhou |
IEEE ACM Trans. Comput. Biol. Bioinform. | 6 |
| 2021 | TCMPR: TCM Prescription recommendation based on subnetwork term mapping and deep learningabstractTraditional Chinese medicine (TCM) has played an indispensable role in clinical diagnose and treatment. Based on patient’s symptom phenotypes, computation-based prescription recommendation methods can recommend personalized TCM prescription using machine learning and artificial intelligence technologies. However, owing to the complexity and individuation of patient’s clinical phenotypes, current prescription recommendation methods cannot obtain good performance. Meanwhile, it’s very difficult to conduct effective representation for unrecorded symptom terms in existing knowledge base. In this study, we proposed a subnetwork-based symptom term mapping method (SSTM), and constructed a SSTM-based TCM prescription recommendation method (termed TCMPR). Our SSTM can extract the subnetwork structure between symptoms from knowledge network to effectively represent the embedding features of clinical symptom terms (especially, the unrecorded terms). The experimental results showed that our method performs better than state-of-the-art methods. In addition, the comprehensive experiments of TCMPR with different hyper parameters (i.e., feature embedding, feature dimension and feature fusion) that demonstrates that our method has high performance on TCM prescription recommendation and potentially promote clinical diagnosis and treatment of TCM precision medicine. Xin Dong 0017, Zixin Shu, Kai Chang, Dengying Yan, Jianan Xia, Kunyu Zhong, Xinyan Wang 0002, Kuo Yang 0001, Xuezhong Zhou |
BIBM | 3 |
| 2020 | Network-based gene prediction for TCM symptomsabstractThe diagnosis and treatment of traditional Chinese medicine (TCM) are formed based on the differentiation of syndromes and symptoms. Symptom management is always the core task of nursing science. Connotation between TCM symptoms and Modern medicine (MM) symptoms are obvious different, especially tongue and pulse symptoms of TCM. However, the underlying molecular mechanisms of most TCM symptoms remain unclear. Here, we developed a network-based framework to predict candidate genes of TCM symptoms (called PTsGene) and construction a high-quality set of TCM symptom-gene associations. Experimental results indicated that PTsGene performed significantly better than the baseline algorithms. The reliability of the candidate genes of symptoms (containing one of typical symptoms of COVID-19, fever) were validated by the analysis of functional homogeneity, molecular co-expression, and recently published literatures. Finally, a high-quality set of TCM symptom-gene associations is constructed to promote the mechanism developments of TCM symptoms. Prediction and construction for reliable TCM symptom-gene associations are valuable for uncovering the underlying molecular mechanisms of TCM symptoms. Our TCM symptom-gene associations deliver a highly insightful data sources for researchers both from basic and clinical settings of precision healthcare. Yinyan Wang, Kuo Yang 0001, Zixin Shu, Dengying Yan, Xuezhong Zhou |
BIBM | 3 |
| 2020 | Integrated network analysis of symptom clusters across disease conditions
Kezhi Lu, Kuo Yang 0001, Edouard Niyongabo, Zixin Shu, Kai Chang, Qunsheng Zou, Jiyue Jiang, Caiyan Jia, Baoyan Liu, Xuezhong Zhou |
J. Biomed. Informatics | 4 |
| 2019 | HerGePred: Heterogeneous Network Embedding Representation for Disease Gene PredictionabstractThe discovery of disease-causing genes is a critical step towards understanding the nature of a disease and determining a possible cure for it. In recent years, many computational methods to identify disease genes have been proposed. However, making full use of disease-related (e.g., symptoms) and gene-related (e.g., gene ontology and protein-protein interactions) information to improve the performance of disease gene prediction is still an issue. Here, we develop a heterogeneous disease-gene-related network (HDGN) embedding representation framework for disease gene prediction (called HerGePred). Based on this framework, a low-dimensional vector representation (LVR) of the nodes in the HDGN can be obtained. Then, we propose two specific algorithms, namely, an LVR-based similarity prediction and a random walk with restart on a reconstructed heterogeneous disease-gene network (RW-RDGN), to predict disease genes with high performance. First, to validate the rationality of the framework, we analyze the similarity-based overlap distribution of disease pairs and design an experiment for disease-gene association recovery, the results of which revealed that the LVR of nodes performs well at preserving the local and global network structure of the HDGN. Then, we apply tenfold cross validation and external validation to compare our methods with other well-known disease gene prediction algorithms. The experimental results show that the RW-RDGN performs better than the state-of-the-art algorithm. The prediction results of disease candidate genes are essential for molecular mechanism investigation and experimental validation. The source codes of HerGePred and experimental data are available at https://github.com/yangkuoone/HerGePred. Kuo Yang 0001, Ruyu Wang, Zixin Shu, Ning Wang 0048, Runshun Zhang, Jian Yu 0001, Xuezhong Zhou |
IEEE J. Biomed. Health Informatics | 4 |