Fuzhong Xue

dblp:50/8541 · DBLP profile ↗
← Back
12ranked-venue papers
0as first author
8since 2021 · last 2026
0000-0003-0378-7956ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 11 · 7 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 MoACNN-XGNet: Interpretable Multi-Omics Convolutional Network for Breast Cancer Subtyping and Prognostic Genes Identification
abstract
Breast cancer, a highly heterogeneous disease at both the phenotypic and molecular levels, presents significant challenges for prognosis and treatment. Accurate subtyping of breast cancer is critical due to its complex biological characteristics, which directly influence disease progression and therapeutic outcomes. In this study, we integrate multi-omics data, including copy number variation, RNA sequencing, and DNA methylation, to generate two-dimensional representations of each sample using Uniform Manifold Approximation and Projection. This transformation enhances data interpretability and supports subsequent learning tasks. Traditional convolutional neural networks have demonstrated potential in medical image analysis but often struggle with high-dimensional omics data. To address this limitation, we propose MoACNN-XGNet, an attention-based convolutional neural network framework that prioritizes key features within image-transformed multi-omics data. Our method significantly improves the precision of subtype classification and effectively overcomes the challenges posed by the high dimensionality and structural complexity of multi-omics data. Furthermore, we employ the Guided Grad-CAM method to enhance model interpretability, enabling the identification of subtype-specific explainable genes. Subsequent enrichment and survival analyses of these genes reveal critical biological pathways and potential therapeutic targets. This study offers a novel approach to refining breast cancer subtyping and highlights the potential for personalized treatment strategies, ultimately aiming to improve patient survival outcomes.
Yaoyao Zhao, Jiayi Teng, Fuzhong Xue, Fan Yang 0068
IEEE J. Biomed. Health Informatics8
2024 MRSL: a causal network pruning algorithm based on GWAS summary data
abstract
Causal discovery is a powerful tool to disclose underlying structures by analyzing purely observational data. Genetic variants can provide useful complementary information for structure learning. Recently, Mendelian randomization (MR) studies have provided abundant marginal causal relationships of traits. Here, we propose a causal network pruning algorithm MRSL (MR-based structure learning algorithm) based on these marginal causal relationships. MRSL combines the graph theory with multivariable MR to learn the conditional causal structure using only genome-wide association analyses (GWAS) summary statistics. Specifically, MRSL utilizes topological sorting to improve the precision of structure learning. It proposes MR-separation instead of d-separation and three candidates of sufficient separating set for MR-separation. The results of simulations revealed that MRSL had up to 2-fold higher F1 score and 100 times faster computing time than other eight competitive methods. Furthermore, we applied MRSL to 26 biomarkers and 44 International Classification of Diseases 10 (ICD10)-defined diseases using GWAS summary data from UK Biobank. The results cover most of the expected causal links that have biological interpretations and several new links supported by clinical case reports or previous observational literatures.
Zhi Geng, Zhongshang Yuan, Fuzhong Xue
Briefings Bioinform.8
2023 MPI-VGAE: protein-metabolite enzymatic reaction link learning by variational graph autoencoders
abstract
Enzymatic reactions are crucial to explore the mechanistic function of metabolites and proteins in cellular processes and to understand the etiology of diseases. The increasing number of interconnected metabolic reactions allows the development of in silico deep learning-based methods to discover new enzymatic reaction links between metabolites and proteins to further expand the landscape of existing metabolite-protein interactome. Computational approaches to predict the enzymatic reaction link by metabolite-protein interaction (MPI) prediction are still very limited. In this study, we developed a Variational Graph Autoencoders (VGAE)-based framework to predict MPI in genome-scale heterogeneous enzymatic reaction networks across ten organisms. By incorporating molecular features of metabolites and proteins as well as neighboring information in the MPI networks, our MPI-VGAE predictor achieved the best predictive performance compared to other machine learning methods. Moreover, when applying the MPI-VGAE framework to reconstruct hundreds of metabolic pathways, functional enzymatic reaction networks and a metabolite-metabolite interaction network, our method showed the most robust performance among all scenarios. To the best of our knowledge, this is the first MPI predictor by VGAE for enzymatic reaction link prediction. Furthermore, we implemented the MPI-VGAE framework to reconstruct the disease-specific MPI network based on the disrupted metabolites and proteins in Alzheimer's disease and colorectal cancer, respectively. A substantial number of novel enzymatic reaction links were identified. We further validated and explored the interactions of these enzymatic reactions using molecular docking. These results highlight the potential of the MPI-VGAE framework for the discovery of novel disease-related enzymatic reactions and facilitate the study of the disrupted metabolisms in diseases.
Chuang Yuan, Ranran Chen, Yuying Shi, Tao Zhang 0127, Fuzhong Xue, Gary J. Patti, Leyi Wei, Qingzhen Hou
Briefings Bioinform.7
2023 Personalized prediction for multiple chronic diseases by developing the multi-task Cox learning model
abstract
Personalized prediction of chronic diseases is crucial for reducing the disease burden. However, previous studies on chronic diseases have not adequately considered the relationship between chronic diseases. To explore the patient-wise risk of multiple chronic diseases, we developed a multitask learning Cox (MTL-Cox) model for personalized prediction of nine typical chronic diseases on the UK Biobank dataset. MTL-Cox employs a multitask learning framework to train semiparametric multivariable Cox models. To comprehensively estimate the performance of the MTL-Cox model, we measured it via five commonly used survival analysis metrics: concordance index, area under the curve (AUC), specificity, sensitivity, and Youden index. In addition, we verified the validity of the MTL-Cox model framework in the Weihai physical examination dataset, from Shandong province, China. The MTL-Cox model achieved a statistically significant (p<0.05) improvement in results compared with competing methods in the evaluation metrics of the concordance index, AUC, sensitivity, and Youden index using the paired-sample Wilcoxon signed-rank test. In particular, the MTL-Cox model improved prediction accuracy by up to 12% compared to other models. We also applied the MTL-Cox model to rank the absolute risk of nine chronic diseases in patients on the UK Biobank dataset. This was the first known study to use the multitask learning-based Cox model to predict the personalized risk of the nine chronic diseases. The study can contribute to early screening, personalized risk ranking, and diagnosing of chronic diseases.
Shuaijie Zhang, Fan Yang 0068, Shucheng Si, Jianmei Zhang, Fuzhong Xue
PLoS Comput. Biol.6
2023 Kernelized Multitask Learning Method for Personalized Signaling Adverse Drug Reactions
abstract
The signaling of the associations between drugs and adverse drug reactions (ADRs) is a challenging task in pharmacovigilance, especially when an association is infrequent or has never previously been reported. Most existing methods for ADR signaling are based on analyzing the frequency with which drugs tend to co-occur with ADRs. In this article, we propose a kernelized multitask learning model, KEMULA, in which information is learned and transferred from the clinical data of other patients as collaborative information to rank distinct lists of ADRs for different patients. We comprehensively compare the performance of KEMULA against three baseline methods, two state-of-the-art ADR signaling methods, and two KEMULA variants. The method is tested on adverse drug event reports retrieved from the FDA Adverse Event Reporting System (FAERS), which includes 4,106,633 unique adverse drug event reports, 7,824 unique ADRs, 114 unique biotech drugs, 1,151 unique small molecule drugs, and 3,363 unique medical conditions. The experimental results demonstrate the advantages of our method and show that it not only can signal frequent ADRs but also has the power to signal infrequent ADRs that cannot be signaled by most existing methods.
Fan Yang 0068, Fuzhong Xue, Yanchun Zhang, George Karypis
IEEE Trans. Knowl. Data Eng.2
2022 Signaling repurposable drug combinations against COVID-19 by developing the heterogeneous deep herb-graph method
abstract
BACKGROUND: Coronavirus disease 2019 (COVID-19) has spurred a boom in uncovering repurposable existing drugs. Drug repurposing is a strategy for identifying new uses for approved or investigational drugs that are outside the scope of the original medical indication. MOTIVATION: Current works of drug repurposing for severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) are mostly limited to only focusing on chemical medicines, analysis of single drug targeting single SARS-CoV-2 protein, one-size-fits-all strategy using the same treatment (same drug) for different infected stages of SARS-CoV-2. To dilute these issues, we initially set the research focusing on herbal medicines. We then proposed a heterogeneous graph embedding method to signaled candidate repurposing herbs for each SARS-CoV-2 protein, and employed the variational graph convolutional network approach to recommend the precision herb combinations as the potential candidate treatments against the specific infected stage. METHOD: We initially employed the virtual screening method to construct the 'Herb-Compound' and 'Compound-Protein' docking graph based on 480 herbal medicines, 12,735 associated chemical compounds and 24 SARS-CoV-2 proteins. Sequentially, the 'Herb-Compound-Protein' heterogeneous network was constructed by means of the metapath-based embedding approach. We then proposed the heterogeneous-information-network-based graph embedding method to generate the candidate ranking lists of herbs that target structural, nonstructural and accessory SARS-CoV-2 proteins, individually. To obtain precision synthetic effective treatments forvarious COVID-19 infected stages, we employed the variational graph convolutional network method to generate candidate herb combinations as the recommended therapeutic therapies. RESULTS: There were 24 ranking lists, each containing top-10 herbs, targeting 24 SARS-CoV-2 proteins correspondingly, and 20 herb combinations were generated as the candidate-specific treatment to target the four infected stages. The code and supplementary materials are freely available at https://github.com/fanyang-AI/TCM-COVID19.
Fan Yang 0068, Shuaijie Zhang, Ruiyuan Yao, Yanchun Zhang, Guoyin Wang 0001, Qianghua Zhang, Yunlong Cheng, Jihua Dong, Chunyang Ruan, Li-Zhen Cui 0001, Hao Wu 0062, Fuzhong Xue
Briefings Bioinform.14
2021 Mendelian randomization under the omnigenic architecture
abstract
Mendelian randomization (MR) is a common analytic tool for exploring the causal relationship among complex traits. Existing MR methods require selecting a small set of single nucleotide polymorphisms (SNPs) to serve as instrument variables. However, selecting a small set of SNPs may not be ideal, as most complex traits have a polygenic or omnigenic architecture and are each influenced by thousands of SNPs. Here, motivated by the recent omnigenic hypothesis, we present an MR method that uses all genome-wide SNPs for causal inference. Our method uses summary statistics from genome-wide association studies as input, accommodates the commonly encountered horizontal pleiotropy effects and relies on a composite likelihood framework for scalable computation. We refer to our method as the omnigenic Mendelian randomization, or OMR. We examine the power and robustness of OMR through extensive simulations including those under various modeling misspecifications. We apply OMR to several real data applications, where we identify multiple complex traits that potentially causally influence coronary artery disease (CAD) and asthma. The identified new associations reveal important roles of blood lipids, blood pressure and immunity underlying CAD as well as important roles of immunity and obesity underlying asthma.
Boran Gao, Fuzhong Xue
Briefings Bioinform.4
2021 SeRenDIP-CE: sequence-based interface prediction for conformational epitopes
abstract
MOTIVATION: Antibodies play an important role in clinical research and biotechnology, with their specificity determined by the interaction with the antigen's epitope region, as a special type of protein-protein interaction (PPI) interface. The ubiquitous availability of sequence data, allows us to predict epitopes from sequence in order to focus time-consuming wet-lab experiments toward the most promising epitope regions. Here, we extend our previously developed sequence-based predictors for homodimer and heterodimer PPI interfaces to predict epitope residues that have the potential to bind an antibody. RESULTS: We collected and curated a high quality epitope dataset from the SAbDab database. Our generic PPI heterodimer predictor obtained an AUC-ROC of 0.666 when evaluated on the epitope test set. We then trained a random forest model specifically on the epitope dataset, reaching AUC 0.694. Further training on the combined heterodimer and epitope datasets, improves our final predictor to AUC 0.703 on the epitope test set. This is better than the best state-of-the-art sequence-based epitope predictor BepiPred-2.0. On one solved antibody-antigen structure of the COVID19 virus spike receptor binding domain, our predictor reaches AUC 0.778. We added the SeRenDIP-CE Conformational Epitope predictors to our webserver, which is simple to use and only requires a single antigen sequence as input, which will help make the method immediately applicable in a wide range of biomedical and biomolecular research. AVAILABILITY AND IMPLEMENTATION: Webserver, source code and datasets at www.ibi.vu.nl/programs/serendipwww/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Qingzhen Hou, Bas Stringer, Katharina Waury, Henriette Capel, Reza Haydarlou, Fuzhong Xue, Sanne Abeln, Jaap Heringa, K. Anton Feenstra
Bioinform.6
2020 Similarity Computation based on Formal Concept Analysis for Colorectal Cancer Patients
abstract
Colorectal cancer is a heterogeneous disease. Its response to targeted therapies is associated with various factors, and the treatment effect differ significantly between individuals. Personalize medical treatment (PMT), which takes into consideration of individual patient characteristics, is the most effective way to deal with this issue. Patient similarity and clustering analysis is an important part in PMT. Earlier works mainly focused on similarity computation among the patients but overlook to preserve relationships. This paper presents a formal concept analysis-based approach for computing the similarity between colorectal cancer patients. The approach not only does the clustering of patients based on their similarity but also can preserve the relations between clusters in hierarchical structural form. This would allow us to build a knowledge base which is helpful for clinicians to take fast and effective decision for treatment and care of colorectal cancer patient.
Jing Xiang, Hanbing Xu, Suresh Pokharel, Jiqing Li, Fuzhong Xue, Ping Zhang 0008
BIBM5
2018 A new insight into underlying disease mechanism through semi-parametric latent differential network model
abstract
BACKGROUND: In genomic studies, to investigate how the structure of a genetic network differs between two experiment conditions is a very interesting but challenging problem, especially in high-dimensional setting. Existing literatures mostly focus on differential network modelling for continuous data. However, in real application, we may encounter discrete data or mixed data, which urges us to propose a unified differential network modelling for various data types. RESULTS: We propose a unified latent Gaussian copula differential network model which provides deeper understanding of the unknown mechanism than that among the observed variables. Adaptive rank-based estimation approaches are proposed with the assumption that the true differential network is sparse. The adaptive estimation approaches do not require precision matrices to be sparse, and thus can allow the individual networks to contain hub nodes. Theoretical analysis shows that the proposed methods achieve the same parametric convergence rate for both the difference of the precision matrices estimation and differential structure recovery, which means that the extra modeling flexibility comes at almost no cost of statistical efficiency. Besides theoretical analysis, thorough numerical simulations are conducted to compare the empirical performance of the proposed methods with some other state-of-the-art methods. The result shows that the proposed methods work quite well for various data types. The proposed method is then applied on gene expression data associated with lung cancer to illustrate its empirical usefulness. CONCLUSIONS: The proposed latent variable differential network models allows for various data-types and thus are more flexible, which also provide deeper understanding of the unknown mechanism than that among the observed variables. Theoretical analysis, numerical simulation and real application all demonstrate the great advantages of the latent differential network modelling and thus are highly recommended.
Jiadong Ji, Fuzhong Xue
BMC Bioinform.5
2017 JDINAC: joint density-based non-parametric differential interaction network analysis and classification using high-dimensional sparse omics data
abstract
MOTIVATION: A complex disease is usually driven by a number of genes interwoven into networks, rather than a single gene product. Network comparison or differential network analysis has become an important means of revealing the underlying mechanism of pathogenesis and identifying clinical biomarkers for disease classification. Most studies, however, are limited to network correlations that mainly capture the linear relationship among genes, or rely on the assumption of a parametric probability distribution of gene measurements. They are restrictive in real application. RESULTS: We propose a new Joint density based non-parametric Differential Interaction Network Analysis and Classification (JDINAC) method to identify differential interaction patterns of network activation between two groups. At the same time, JDINAC uses the network biomarkers to build a classification model. The novelty of JDINAC lies in its potential to capture non-linear relations between molecular interactions using high-dimensional sparse data as well as to adjust confounding factors, without the need of the assumption of a parametric probability distribution of gene measurements. Simulation studies demonstrate that JDINAC provides more accurate differential network estimation and lower classification error than that achieved by other state-of-the-art methods. We apply JDINAC to a Breast Invasive Carcinoma dataset, which includes 114 patients who have both tumor and matched normal samples. The hub genes and differential interaction patterns identified were consistent with existing experimental studies. Furthermore, JDINAC discriminated the tumor and normal sample with high accuracy by virtue of the identified biomarkers. JDINAC provides a general framework for feature selection and classification using high-dimensional sparse omics data. AVAILABILITY AND IMPLEMENTATION: R scripts available at https://github.com/jijiadong/JDINAC. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Jiadong Ji, Di He 0003, Yang Feng 0002, Fuzhong Xue, Lei Xie 0006
Bioinform.5
2016 A powerful score-based statistical test for group difference in weighted biological networks
abstract
BACKGROUND: Complex disease is largely determined by a number of biomolecules interwoven into networks, rather than a single biomolecule. A key but inadequately addressed issue is how to test possible differences of the networks between two groups. Group-level comparison of network properties may shed light on underlying disease mechanisms and benefit the design of drug targets for complex diseases. We therefore proposed a powerful score-based statistic to detect group difference in weighted networks, which simultaneously capture the vertex changes and edge changes. RESULTS: Simulation studies indicated that the proposed network difference measure (NetDifM) was stable and outperformed other methods existed, under various sample sizes and network topology structure. One application to real data about GWAS of leprosy successfully identified the specific gene interaction network contributing to leprosy. For additional gene expression data of ovarian cancer, two candidate subnetworks, PI3K-AKT and Notch signaling pathways, were considered and identified respectively. CONCLUSIONS: The proposed method, accounting for the vertex changes and edge changes simultaneously, is valid and powerful to capture the group difference of biological networks.
Jiadong Ji, Zhongshang Yuan, Xiaoshuai Zhang, Fuzhong Xue
BMC Bioinform.4