Fan Yang 0068

dblp:29/3081-68 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
6since 2021 · last 2026
0000-0003-0717-5420ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 MoACNN-XGNet: Interpretable Multi-Omics Convolutional Network for Breast Cancer Subtyping and Prognostic Genes Identification
abstract
Breast cancer, a highly heterogeneous disease at both the phenotypic and molecular levels, presents significant challenges for prognosis and treatment. Accurate subtyping of breast cancer is critical due to its complex biological characteristics, which directly influence disease progression and therapeutic outcomes. In this study, we integrate multi-omics data, including copy number variation, RNA sequencing, and DNA methylation, to generate two-dimensional representations of each sample using Uniform Manifold Approximation and Projection. This transformation enhances data interpretability and supports subsequent learning tasks. Traditional convolutional neural networks have demonstrated potential in medical image analysis but often struggle with high-dimensional omics data. To address this limitation, we propose MoACNN-XGNet, an attention-based convolutional neural network framework that prioritizes key features within image-transformed multi-omics data. Our method significantly improves the precision of subtype classification and effectively overcomes the challenges posed by the high dimensionality and structural complexity of multi-omics data. Furthermore, we employ the Guided Grad-CAM method to enhance model interpretability, enabling the identification of subtype-specific explainable genes. Subsequent enrichment and survival analyses of these genes reveal critical biological pathways and potential therapeutic targets. This study offers a novel approach to refining breast cancer subtyping and highlights the potential for personalized treatment strategies, ultimately aiming to improve patient survival outcomes.
Yaoyao Zhao, Jiayi Teng, Fuzhong Xue, Fan Yang 0068
IEEE J. Biomed. Health Informatics9
2023 Personalized prediction for multiple chronic diseases by developing the multi-task Cox learning model
abstract
Personalized prediction of chronic diseases is crucial for reducing the disease burden. However, previous studies on chronic diseases have not adequately considered the relationship between chronic diseases. To explore the patient-wise risk of multiple chronic diseases, we developed a multitask learning Cox (MTL-Cox) model for personalized prediction of nine typical chronic diseases on the UK Biobank dataset. MTL-Cox employs a multitask learning framework to train semiparametric multivariable Cox models. To comprehensively estimate the performance of the MTL-Cox model, we measured it via five commonly used survival analysis metrics: concordance index, area under the curve (AUC), specificity, sensitivity, and Youden index. In addition, we verified the validity of the MTL-Cox model framework in the Weihai physical examination dataset, from Shandong province, China. The MTL-Cox model achieved a statistically significant (p<0.05) improvement in results compared with competing methods in the evaluation metrics of the concordance index, AUC, sensitivity, and Youden index using the paired-sample Wilcoxon signed-rank test. In particular, the MTL-Cox model improved prediction accuracy by up to 12% compared to other models. We also applied the MTL-Cox model to rank the absolute risk of nine chronic diseases in patients on the UK Biobank dataset. This was the first known study to use the multitask learning-based Cox model to predict the personalized risk of the nine chronic diseases. The study can contribute to early screening, personalized risk ranking, and diagnosing of chronic diseases.
Shuaijie Zhang, Fan Yang 0068, Shucheng Si, Jianmei Zhang, Fuzhong Xue
PLoS Comput. Biol.2
2023 Kernelized Multitask Learning Method for Personalized Signaling Adverse Drug Reactions
abstract
The signaling of the associations between drugs and adverse drug reactions (ADRs) is a challenging task in pharmacovigilance, especially when an association is infrequent or has never previously been reported. Most existing methods for ADR signaling are based on analyzing the frequency with which drugs tend to co-occur with ADRs. In this article, we propose a kernelized multitask learning model, KEMULA, in which information is learned and transferred from the clinical data of other patients as collaborative information to rank distinct lists of ADRs for different patients. We comprehensively compare the performance of KEMULA against three baseline methods, two state-of-the-art ADR signaling methods, and two KEMULA variants. The method is tested on adverse drug event reports retrieved from the FDA Adverse Event Reporting System (FAERS), which includes 4,106,633 unique adverse drug event reports, 7,824 unique ADRs, 114 unique biotech drugs, 1,151 unique small molecule drugs, and 3,363 unique medical conditions. The experimental results demonstrate the advantages of our method and show that it not only can signal frequent ADRs but also has the power to signal infrequent ADRs that cannot be signaled by most existing methods.
Fan Yang 0068, Fuzhong Xue, Yanchun Zhang, George Karypis
IEEE Trans. Knowl. Data Eng.1
2022 StackTADB: a stacking-based ensemble learning model for predicting the boundaries of topologically associating domains (TADs) accurately in fruit flies
abstract
Chromosome is composed of many distinct chromatin domains, referred to variably as topological domains or topologically associating domains (TADs). The domains are stable across different cell types and highly conserved across species, thus these chromatin domains have been considered as the basic units of chromosome folding and regarded as an important secondary structure in chromosome organization. However, the identification of TAD boundaries is still a great challenge due to the high cost and low resolution of Hi-C data or experiments. In this study, we propose a novel ensemble learning framework, termed as StackTADB, for predicting the boundaries of TADs. StackTADB integrates four base classifiers including Random Forest, Logistic Regression, K-NearestNeighbor and Support Vector Machine. From the analysis of a series of examinations on the data set in the previous study, it is concluded that StackTADB has optimal performance in six metrics, AUC, Accuracy, MCC, Precision, Recall and F1 score, and it is superior to the existing methods. In addition, the comparison of the performance of multiple features shows that Kmers-based features play an essential role in predicting TADs boundaries of fruit flies, and we also apply the SHapley Additive exPlanations (SHAP) framework to interpret the predictions of StackTADB to identify the reason why Kmers-based features are vital. The experimental results show that the subsequences matching the BEAF-32 motif play a crucial role in predicting the boundaries of TADs. The source code is freely available at https://github.com/HaoWuLab-Bioinformatics/StackTADB and the webserver of StackTADB is freely available at http://hwtad.sdu.edu.cn:8002/StackTADB.
Hao Wu 0062, Zhaoheng Ai, Leyi Wei, Hongming Zhang 0002, Fan Yang 0068, Li-Zhen Cui 0001
Briefings Bioinform.6
2022 Signaling repurposable drug combinations against COVID-19 by developing the heterogeneous deep herb-graph method
abstract
BACKGROUND: Coronavirus disease 2019 (COVID-19) has spurred a boom in uncovering repurposable existing drugs. Drug repurposing is a strategy for identifying new uses for approved or investigational drugs that are outside the scope of the original medical indication. MOTIVATION: Current works of drug repurposing for severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) are mostly limited to only focusing on chemical medicines, analysis of single drug targeting single SARS-CoV-2 protein, one-size-fits-all strategy using the same treatment (same drug) for different infected stages of SARS-CoV-2. To dilute these issues, we initially set the research focusing on herbal medicines. We then proposed a heterogeneous graph embedding method to signaled candidate repurposing herbs for each SARS-CoV-2 protein, and employed the variational graph convolutional network approach to recommend the precision herb combinations as the potential candidate treatments against the specific infected stage. METHOD: We initially employed the virtual screening method to construct the 'Herb-Compound' and 'Compound-Protein' docking graph based on 480 herbal medicines, 12,735 associated chemical compounds and 24 SARS-CoV-2 proteins. Sequentially, the 'Herb-Compound-Protein' heterogeneous network was constructed by means of the metapath-based embedding approach. We then proposed the heterogeneous-information-network-based graph embedding method to generate the candidate ranking lists of herbs that target structural, nonstructural and accessory SARS-CoV-2 proteins, individually. To obtain precision synthetic effective treatments forvarious COVID-19 infected stages, we employed the variational graph convolutional network method to generate candidate herb combinations as the recommended therapeutic therapies. RESULTS: There were 24 ranking lists, each containing top-10 herbs, targeting 24 SARS-CoV-2 proteins correspondingly, and 20 herb combinations were generated as the candidate-specific treatment to target the four infected stages. The code and supplementary materials are freely available at https://github.com/fanyang-AI/TCM-COVID19.
Fan Yang 0068, Shuaijie Zhang, Ruiyuan Yao, Yanchun Zhang, Guoyin Wang 0001, Qianghua Zhang, Yunlong Cheng, Jihua Dong, Chunyang Ruan, Li-Zhen Cui 0001, Hao Wu 0062, Fuzhong Xue
Briefings Bioinform.1
2021 Disease Prediction via Graph Neural Networks
abstract
With the increasingly available electronic medical records (EMRs), disease prediction has recently gained immense research attention, where an accurate classifier needs to be trained to map the input prediction signals (e.g., symptoms, patient demographics, etc.) to the estimated diseases for each patient. However, existing machine learning-based solutions heavily rely on abundant manually labeled EMR training data to ensure satisfactory prediction results, impeding their performance in the existence of rare diseases that are subject to severe data scarcity. For each rare disease, the limited EMR data can hardly offer sufficient information for a model to correctly distinguish its identity from other diseases with similar clinical symptoms. Furthermore, most existing disease prediction approaches are based on the sequential EMRs collected for every patient and are unable to handle new patients without historical EMRs, reducing their real-life practicality. In this paper, we introduce an innovative model based on Graph Neural Networks (GNNs) for disease prediction, which utilizes external knowledge bases to augment the insufficient EMR data, and learns highly representative node embeddings for patients, diseases and symptoms from the medical concept graph and patient record graph respectively constructed from the medical knowledge base and EMRs. By aggregating information from directly connected neighbor nodes, the proposed neural graph encoder can effectively generate embeddings that capture knowledge from both data sources, and is able to inductively infer the embeddings for a new patient based on the symptoms reported in her/his EMRs to allow for accurate prediction on both general diseases and rare diseases. Extensive experiments on a real-world EMR dataset have demonstrated the state-of-the-art performance of our proposed model.
Zhenchao Sun, Hongzhi Yin, Hongxu Chen 0002, Tong Chen 0005, Li-Zhen Cui 0001, Fan Yang 0068
IEEE J. Biomed. Health Informatics6
2014 Signaling adverse drug reactions with novel feature-based similarity model
abstract
Adverse drug reactions (ADRs) are a main cause of hospitalization and deaths worldwide. These unanticipated episodes are generally infrequent, but almost all existing ADR signaling techniques are designed to use dataset extracted from spontaneous reporting systems or employed a predefined type of information (e.g., drugs), which suffer from failures to detect unexpected and latent ADRs. In this paper, we propose a novel Feature-based Similarity model (FS) to detect the potential ADRs for medical cases using the electronic patient dataset. FS is tested on the real patient data retrieved from the US Food Drug Administration that includes 54,070 patients detail information and 9,567 ADRs records. Our model ranked all ADRs for the given medical case that combined the information of drugs, medical conditions, and patient profiles and can be applied in therapy decision support systems and unexpected ADR warning systems. The experimental results show that FS outperforms comparing methods. This paper clearly illustrates the great potential along the new direction of ADR signal generate from health care administrative database.
Fan Yang 0068, Xiaohui Yu 0001, George Karypis
BIBM1