Yang Hu 0008

dblp:43/4685-8 · DBLP profile ↗
← Back
27ranked-venue papers
4as first author
11since 2021 · last 2025
0000-0002-4508-5365ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 27 · 4 first-author · 11 since 2021
YearPublicationVenuePosition
2025 MetImputBERT: a pretrained BERT framework for missing value imputation in NMR metabolomics data
abstract
Missing values in nuclear magnetic resonance metabolomics data compromise downstream clinical interpretation. Here, we present MetImputBERT, an imputation method based on a pretrained BERT framework. MetImputBERT uses the masks in the masked language model to simulate missing values and leverages predictions and reconstructions to these positions to simulate the imputation process. The learning of MetImputBERT is driven by minimizing the reconstruction error. MetImputBERT was pretrained on the largest metabolomics dataset to date, comprising data from over 230 000 individuals in the UK Biobank. When new datasets with missing values were encountered, MetImputBERT loaded the pretrained parameters and directly imputed the missing values by inferring their reconstructed estimates. MetImputBERT outperformed commonly used methods-K-nearest neighbors, multiple imputation by chained equations, and singular value decomposition-in imputation performance on two independent test sets. We provide an open-source Python tool that allows users to quickly impute missing values in their own NMR metabolomics data without any additional training.
Shizheng Qiu, Yang Hu 0008, Guiyou Liu, Yadong Wang 0001
Briefings Bioinform.2
2025 Proformer: a multimodal proteomics transformer model for multidisease early risk assessment
abstract
Early identification of individuals at high risk for chronic diseases is crucial for prevention and intervention, yet current risk assessment tools are disease-specific, require extensive clinical data collection, and cannot provide multidisease risk profiles from a single measurement. Several protein large language models have been developed for tasks such as protein structure prediction, function prediction, and sequence design. However, none of these models can be directly applied in clinical settings to predict an individual's future disease risk. Here, we present a multimodal proteomics Transformer (Proformer) model that integrates protein expression, sequence, and function information for multidisease risk assessment. We trained Proformer using real proteomics data from 47 124 individuals from the UK Biobank to evaluate its performance in discriminating the risk of 20 common chronic diseases. Proformer achieved state-of-the-art (SOTA) performance in all 20 diseases compared with five common machine learning and deep learning models. Compared to three common clinical predictors, Proformer's 10-year discriminative performance outperforms Age + Sex model for 19 diseases, outperforms the ASCVD risk score for 16 diseases, and outperforms the panel composed of 35 clinical variables for 11 diseases. These results were replicated in the Scotland and Wales cohort from UK Biobank. In conclusion, Proformer enabled users to directly obtain a 10-year risk report for common chronic diseases by inputting their individual proteomics data.
Shizheng Qiu, Yang Hu 0008, Yadong Wang 0001
Briefings Bioinform.2
2024 Interpretation of 10 years of Alzheimer's disease genetic findings in the perspective of statistical heterogeneity
abstract
Common genetic variants and susceptibility loci associated with Alzheimer's disease (AD) have been discovered through large-scale genome-wide association studies (GWAS), GWAS by proxy (GWAX) and meta-analysis of GWAS and GWAX (GWAS+GWAX). However, due to the very low repeatability of AD susceptibility loci and the low heritability of AD, these AD genetic findings have been questioned. We summarize AD genetic findings from the past 10 years and provide a new interpretation of these findings in the context of statistical heterogeneity. We discovered that only 17% of AD risk loci demonstrated reproducibility with a genome-wide significance of P < 5.00E-08 across all AD GWAS and GWAS+GWAX datasets. We highlighted that the AD GWAS+GWAX with the largest sample size failed to identify the most significant signals, the maximum number of genome-wide significant genetic variants or maximum heritability. Additionally, we identified widespread statistical heterogeneity in AD GWAS+GWAX datasets, but not in AD GWAS datasets. We consider that statistical heterogeneity may have attenuated the statistical power in AD GWAS+GWAX and may contribute to explaining the low repeatability (17%) of genome-wide significant AD susceptibility loci and the decreased AD heritability (40-2%) as the sample size increased. Importantly, evidence supports the idea that a decrease in statistical heterogeneity facilitates the identification of genome-wide significant genetic loci and contributes to an increase in AD heritability. Collectively, current AD GWAX and GWAS+GWAX findings should be meticulously assessed and warrant additional investigation, and AD GWAS+GWAX should employ multiple meta-analysis methods, such as random-effects inverse variance-weighted meta-analysis, which is designed specifically for statistical heterogeneity.
Zhifa Han, Yang Hu 0008, Yanli Xue, Guiyou Liu
Briefings Bioinform.4
2023 Identification of risk genes and biological pathways influencing myopia via transcriptome association study and biomedical ontology methods
abstract
Refractive error remains one of the most common eye diseases in the world, which causes great burden on public health and economy every year. Although genome-wide association studies (GWAS) have identified susceptibility loci for refractive error, the credible causal genes associated with gene expression in brain are still unclear. In addition, the influence of RNA splicing events on refractive error in brain remains unknown. Here, we identified candidate refractive error gene targets and biological pathways. In stage 1, we carried out gene-based association analysis to identify refractive error risk genes. In stage 2, by using transcriptome-wide association studies (TWAS) and colocalization analysis, we identified cis-regulated genes in brain tissues that were positively related to refractive error from the whole gene region and the exon region after RNA splicing, respectively. In stage 3, we reveal biological pathways influencing refractive error using biological ontology and knowledge base methods. We identified 367 risk genes passed gene-based association test. A total of 43 genes in brain could be considered as causal genes using TWAS and colocalization analysis, and 27 were novel. After RNA splicing, RCBTB1 significantly affected refractive error. Risk genes for refractive error significantly enriched in biological processes related to synaptic transmission and acetylcholine. In conclusion, our analysis provides a better understanding of potential therapeutic targets and biological pathways for refractive error.
Shizheng Qiu, Yadong Wang 0001, Yang Hu 0008
BIBM4
2023 HNetGO: protein function prediction via heterogeneous network transformer
abstract
Protein function annotation is one of the most important research topics for revealing the essence of life at molecular level in the post-genome era. Current research shows that integrating multisource data can effectively improve the performance of protein function prediction models. However, the heavy reliance on complex feature engineering and model integration methods limits the development of existing methods. Besides, models based on deep learning only use labeled data in a certain dataset to extract sequence features, thus ignoring a large amount of existing unlabeled sequence data. Here, we propose an end-to-end protein function annotation model named HNetGO, which innovatively uses heterogeneous network to integrate protein sequence similarity and protein-protein interaction network information and combines the pretraining model to extract the semantic features of the protein sequence. In addition, we design an attention-based graph neural network model, which can effectively extract node-level features from heterogeneous networks and predict protein function by measuring the similarity between protein nodes and gene ontology term nodes. Comparative experiments on the human dataset show that HNetGO achieves state-of-the-art performance on cellular component and molecular function branches.
Xiaoshuai Zhang, Huannan Guo, Xuan Wang 0002, Kaitao Wu, Shizheng Qiu, Bo Liu 0023, Yadong Wang 0001, Yang Hu 0008, Junyi Li 0004
Briefings Bioinform.9
2022 MTOR hypermethylation may associate with the susceptibility and survival of SARS-CoV-2 infections to lung adenocarcinoma patients based on multi-omics data and machine learning
abstract
Recent studies have shown that lung adenocarcinoma (LUAD) patients have a higher risk and worse prognosis of COVID-19 caused by SARS-CoV-2 compared to normal samples. Whereas, in addition to the receptor for SARS-CoV-2, other genes also deserve attention. In our study, we identified 19 differentially methylated genes (DMGs) that were co-upregulated in LUAD and COVID-19 samples. These 19 DMGs mainly regulated the immune-related and multiple viral infection signaling pathways. Gene Ontology and pathway enrichment analysis were applied with these genes. Then, 6 key DMGs (MTOR, ACE, IGF1, PTPRC, C3, and PTGS2) were identified by constructing and analyzing the protein-protein interaction (PPI) network. Besides, MTOR was significantly associated with 5 prognostic markers (CDO1, NEURL4, SMAP1, NPEPPS, IQCK) identified by survival analysis based on machine learning. In total, MTOR hypermethylation may be related to the susceptibility of LUAD patients to SARS-CoV-2 and the prognosis of LUAD patients suffering from COVID-19.
Yang Hu 0008, Tianyi Zang
BIBM3
2022 Identification of the expression, prognostic value and cancer immunity of Gasdermin E based on multi-omics data, machine learning and gene ontology
abstract
Gasdermin E (GSDME)-mediated pyroptosis participates in the recruitment of tumor-infiltrating lymphocytes into the tumor microenvironment in primary breast cancers, melanoma and colorectal cancer. However, it is still unknown whether GSDME suppresses tumors in kidney cancer. Here, we comprehensively analyzed the mRNA expression, genetic alteration, prognosis value and cancer immunity of GSDME in three types of kidney cancer, including kidney renal clear cell carcinoma (KIRC), kidney renal papillary cell carcinoma (KIRP) and kidney chromophobe (KICH), based on multi-omics data, machine learning and gene ontology (GO). Compared with normal samples, GSDME expression was significantly down-regulated in KICH and up regulated in KIRP. The genetic alteration of GSDME changed significantly in the process of tumor development, and the results of the univariate cox regression model and log-rank test suggested GSDME as a potential prognostic marker for all the three kidney cancers. Immune infiltration and correlation analysis supported that GSDME-mediated pyroptosis played a critical role in kidney cancer and was highly associated with killer cells and macrophages in the tumor microenvironment. Most GSDME-related genes were enriched in membrane proteins and lysosomal, and enzyme activity-related pathways by GO enrichment, which was consistent with the pyroptosis process. Together, our work provided new insights and theoretical basis for GSDME as a drug target for the treatment of kidney cancer.
Shizheng Qiu, Yang Hu 0008
BIBM3
2022 MHCRoBERTa: pan-specific peptide-MHC class I binding prediction through transfer learning with label-agnostic protein sequences
abstract
Predicting the binding of peptide and major histocompatibility complex (MHC) plays a vital role in immunotherapy for cancer. The success of Alphafold of applying natural language processing (NLP) algorithms in protein secondary struction prediction has inspired us to explore the possibility of NLP methods in predicting peptide-MHC class I binding. Based on the above motivations, we propose the MHCRoBERTa method, RoBERTa pre-training approach, for predicting the binding affinity between type I MHC and peptides. Analysis of the results on benchmark dataset demonstrates that MHCRoBERTa can outperform other state-of-art prediction methods with an increase of the Spearman rank correlation coefficient (SRCC) value. Notably, our model gave a significant improvement on IC50 value. Our method has achieved SRCC value and AUC value as 0.785 and 0.817, respectively. Our SRCC value is 14.3% higher than NetMHCpan3.0 (the second highest SRCC value on pan-specific) and is 3% higher than MHCflurry (the second highest SRCC value on all methods). The AUC value is also better than any other pan-specific methods. Moreover, we visualize the multi-head self-attention for the token representation across the layers and heads by this method. Through the analysis of the representation of each layer and head, we can show whether the model has learned the syntax and semantics necessary to perform the prediction task well. All these results demonstrate that our model can accurately predict the peptide-MHC class I binding affinity and that MHCRoBERTa is a powerful tool for screening potential neoantigens for cancer immunotherapy. MHCRoBERTa is available as an open source software at github (https://github.com/FuxuWang/MHCRoBERTa).
Fuxu Wang, Haoyan Wang, Lizhuang Wang, Haoyu Lu, Shizheng Qiu, Tianyi Zang, Xinjun Zhang, Yang Hu 0008
Briefings Bioinform.8
2021 Differentially Expressed Mutant Genes Reveal Potential Prognostic Markers For Lung Adenocarcinoma
abstract
Lung adenocarcinoma is a serious lung cancer, belonging to the category of non-small cell lung cancer, accounting for 30 to 35 percent of the total, and lung adenocarcinoma is more common in women and non-smoking patients. In this study, LUAD samples of mutation genes and RNA-seq expression dataset were retrieved from TCGA database. WGCNA, Cox regression analysis were used to classify melanoma prognosis. It revealed that eleven mutant and differentially expressed prognosis biomarkers were significantly associated with the prognosis of patients. Therefore, detecting these gene mutations and exploring their corresponding expression could be valuable in predicting the prognosis of patients. The results of the high-throughput data mining provide important fundamental bioinformatics information and a relevant theoretical basis for further exploring the molecular pathogenesis of LUAD and assessing the prognosis of patients.
Yue Liu 0034, Shizheng Qiu, Yang Hu 0008, Yadong Wang 0001
BIBM3
2021 Deep-DRM: a computational method for identifying disease-related metabolites based on graph deep learning approaches
abstract
MOTIVATION: The functional changes of the genes, RNAs and proteins will eventually be reflected in the metabolic level. Increasing number of researchers have researched mechanism, biomarkers and targeted drugs by metabolites. However, compared with our knowledge about genes, RNAs, and proteins, we still know few about diseases-related metabolites. All the few existed methods for identifying diseases-related metabolites ignore the chemical structure of metabolites, fail to recognize the association pattern between metabolites and diseases, and fail to apply to isolated diseases and metabolites. RESULTS: In this study, we present a graph deep learning based method, named Deep-DRM, for identifying diseases-related metabolites. First, chemical structures of metabolites were used to calculate similarities of metabolites. The similarities of diseases were obtained based on their functional gene network and semantic associations. Therefore, both metabolites and diseases network could be built. Next, Graph Convolutional Network (GCN) was applied to encode the features of metabolites and diseases, respectively. Then, the dimension of these features was reduced by Principal components analysis (PCA) with retainment 99% information. Finally, Deep neural network was built for identifying true metabolite-disease pairs (MDPs) based on these features. The 10-cross validations on three testing setups showed outstanding AUC (0.952) and AUPR (0.939) of Deep-DRM compared with previous methods and similar approaches. Ten of top 15 predicted associations between diseases and metabolites got support by other studies, which suggests that Deep-DRM is an efficient method to identify MDPs. CONTACT: [email protected]. AVAILABILITY AND IMPLEMENTATION: https://github.com/zty2009/GPDNN-for-Identify-ing-Disease-related-Metabolites.
Tianyi Zhao 0001, Yang Hu 0008, Liang Cheng 0006
Briefings Bioinform.2
2021 Identifying drug-target interactions based on graph convolutional network and deep neural network
abstract
Identification of new drug-target interactions (DTIs) is an important but a time-consuming and costly step in drug discovery. In recent years, to mitigate these drawbacks, researchers have sought to identify DTIs using computational approaches. However, most existing methods construct drug networks and target networks separately, and then predict novel DTIs based on known associations between the drugs and targets without accounting for associations between drug-protein pairs (DPPs). To incorporate the associations between DPPs into DTI modeling, we built a DPP network based on multiple drugs and proteins in which DPPs are the nodes and the associations between DPPs are the edges of the network. We then propose a novel learning-based framework, 'graph convolutional network (GCN)-DTI', for DTI identification. The model first uses a graph convolutional network to learn the features for each DPP. Second, using the feature representation as an input, it uses a deep neural network to predict the final label. The results of our analysis show that the proposed framework outperforms some state-of-the-art approaches by a large margin.
Tianyi Zhao 0001, Yang Hu 0008, Linda R. Valsdottir, Tianyi Zang, Jiajie Peng
Briefings Bioinform.2
2020 DeepLGP: a novel deep learning method for prioritizing lncRNA target genes
abstract
MOTIVATION: Although long non-coding RNAs (lncRNAs) have limited capacity for encoding proteins, they have been verified as biomarkers in the occurrence and development of complex diseases. Recent wet-lab experiments have shown that lncRNAs function by regulating the expression of protein-coding genes (PCGs), which could also be the mechanism responsible for causing diseases. Currently, lncRNA-related biological data are increasing rapidly. Whereas, no computational methods have been designed for predicting the novel target genes of lncRNA. RESULTS: In this study, we present a graph convolutional network (GCN) based method, named DeepLGP, for prioritizing target PCGs of lncRNA. First, gene and lncRNA features were selected, these included their location in the genome, expression in 13 tissues and miRNA-mediated lncRNA-gene pairs. Next, GCN was applied to convolve a gene interaction network for encoding the features of genes and lncRNAs. Then, these features were used by the convolutional neural network for prioritizing target genes of lncRNAs. In 10-cross validations on two independent datasets, DeepLGP obtained high area under curves (0.90-0.98) and area under precision-recall curves (0.91-0.98). We found that lncRNA pairs with high similarity had more overlapped target genes. Further experiments showed that genes targeted by the same lncRNA sets had a strong likelihood of causing the same diseases, which could help in identifying disease-causing PCGs. AVAILABILITY AND IMPLEMENTATION: https://github.com/zty2009/LncRNA-target-gene. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Tianyi Zhao 0001, Yang Hu 0008, Jiajie Peng, Liang Cheng 0006, Pier Luigi Martelli
Bioinform.2
2020 DRACP: a novel method for identification of anticancer peptides
abstract
BACKGROUND: Millions of people are suffering from cancers, but accurate early diagnosis and effective treatment are still tough for all doctors. Common ways against cancer include surgical operation, radiotherapy and chemotherapy. However, they are all very harmful for patients. Recently, the anticancer peptides (ACPs) have been discovered to be a potential way to treat cancer. Since ACPs are natural biologics, they are safer than other methods. However, the experimental technology is an expensive way to find ACPs so we purpose a new machine learning method to identify the ACPs. RESULTS: Firstly, we extracted the feature of ACPs in two aspects: sequence and chemical characteristics of amino acids. For sequence, average 20 amino acids composition was extracted. For chemical characteristics, we classified amino acids into six groups based on the patterns of hydrophobic and hydrophilic residues. Then, deep belief network has been used to encode the features of ACPs. Finally, we purposed Random Relevance Vector Machines to identify the true ACPs. We call this method 'DRACP' and tested the performance of it on two independent datasets. Its AUC and AUPR are higher than 0.9 in both datasets. CONCLUSION: We developed a novel method named 'DRACP' and compared it with some traditional methods. The cross-validation results showed its effectiveness in identifying ACPs.
Tianyi Zhao 0001, Yang Hu 0008, Tianyi Zang
BMC Bioinform.2
2019 PAGWAS: a manually curated web-based knowledge database of GWAS pathway analysis
abstract
Genome-wide association studies (GWAS) have been widely used to investigate the pathogenesis of human complex diseases, and have yielded important new insights into the genetic mechanisms. However, the newly identified susceptibility loci exert very small risk effects, and cannot fully explain the underlying genetic risk. Fortunately, the existing large-scale GWAS datasets provide strong support for the investigation of human complex disease mechanisms using pathway analysis methods. Here, we developed a web-based knowledge database named PAGWAS to provide a comprehensive catalog of published pathway analysis of GWAS. It provides a convenient way to understand the associations between pathways or genes in these pathways and the given disease or trait of interest. PAGWAS incorporates knowledge from 164 published pathway analysis of GWAS papers and manually curated 5769 pathways and 89 diseases. Also, PAGWAS provides two types of network visualization function to illustrate the relationships between different diseases or pathways. We hope PAGWAS will be a useful tool with great potential for researches on pathogenesis of human complex diseases. Database URL: PAGWAS can be accessed at http://www.bio-annotation.cn:18080/GWASDisease/.
Ningyi Zhang, Yang Hu 0008
BIBM2
2019 Identification of anticancer peptides based on Random Relevance Vector Machines
abstract
Cancer is the most threat to human's health and life. At present, people have developed several ways to against cancer, such as surgical operation, radiotherapy and chemotherapy. However, cancers still cause highly mortality rate. A main part of the reason is that the traditional methods bring treatment effect as well as the negative effect. Recently, the anticancer peptides (ACPs) have been discovered which can be a new way to treat cancer. Since ACPs are natural biologics, they are safer than other methods. However, the experimental technology is an expensive way to find ACPs so we purpose a new machine learning method to identify the ACPs which named Random Relevance Vector Machines (RRVMs). The cross validations experiments show the high accuracy and stability of this new method.
Tianyi Zhao 0001, Tianyi Zang, Yang Hu 0008
BIBM4
2019 BIN1 rs744373 variant shows different association with Alzheimer's disease in Caucasian and Asian populations
abstract
Abstract Background The association between BIN1 rs744373 variant and Alzheimer’s disease (AD) had been identified by genome-wide association studies (GWASs) as well as candidate gene studies in Caucasian populations. But in East Asian populations, both positive and negative results had been identified by association studies. Considering the smaller sample sizes of the studies in East Asian, we believe that the results did not have enough statistical power. Results We conducted a meta-analysis with 71,168 samples (22,395 AD cases and 48,773 controls, from 37 studies of 19 articles). Based on the additive model, we observed significant genetic heterogeneities in pooled populations as well as Caucasians and East Asians. We identified a significant association between rs744373 polymorphism with AD in pooled populations (P = 5 × 10− 07, odds ratio (OR) = 1.12, and 95% confidence interval (CI) 1.07–1.17) and in Caucasian populations (P = 3.38 × 10− 08, OR = 1.16, 95% CI 1.10–1.22). But in the East Asian populations, the association was not identified (P = 0.393, OR = 1.057, and 95% CI 0.95–1.15). Besides, the regression analysis suggested no significant publication bias. The results for sensitivity analysis as well as meta-analysis under the dominant model and recessive model remained consistent, which demonstrated the reliability of our finding. Conclusions The large-scale meta-analysis highlighted the significant association between rs744373 polymorphism and AD risk in Caucasian populations but not in the East Asian populations.
Zhifa Han, Wenyang Zhou, Jian Zong, Yang Hu 0008, Shuilin Jin, Qinghua Jiang
BMC Bioinform.8
2019 Identifying Alzheimer's disease-related proteins by LRRGD
abstract
BACKGROUND: Alzheimer's disease (AD) imposes a heavy burden on society and every family. Therefore, diagnosing AD in advance and discovering new drug targets are crucial, while these could be achieved by identifying AD-related proteins. The time-consuming and money-costing biological experiment makes researchers turn to develop more advanced algorithms to identify AD-related proteins. RESULTS: Firstly, we proposed a hypothesis "similar diseases share similar related proteins". Therefore, five similarity calculation methods are introduced to find out others diseases which are similar to AD. Then, these diseases' related proteins could be obtained by public data set. Finally, these proteins are features of each disease and could be used to map their similarity to AD. We developed a novel method 'LRRGD' which combines Logistic Regression (LR) and Gradient Descent (GD) and borrows the idea of Random Forest (RF). LR is introduced to regress features to similarities. Borrowing the idea of RF, hundreds of LR models have been built by randomly selecting 40 features (proteins) each time. Here, GD is introduced to find out the optimal result. To avoid the drawback of local optimal solution, a good initial value is selected by some known AD-related proteins. Finally, 376 proteins are found to be related to AD. CONCLUSION: Three hundred eight of three hundred seventy-six proteins are the novel proteins. Three case studies are done to prove our method's effectiveness. These 308 proteins could give researchers a basis to do biological experiments to help treatment and diagnostic AD.
Tianyi Zhao 0001, Yang Hu 0008, Tianyi Zang, Liang Cheng 0006
BMC Bioinform.2
2018 A Novel Method for Identifying Alzheimer's Disease-related Proteins
Yang Hu 0008, Jun Zhang 0041, Tianyi Zhao 0001, Liang Cheng 0006, Tianyi Zang
BIBM1
2018 BIN1 rs744373 Variant Is Significantly Associated with Alzheimer's Disease in Caucasian but Not East Asian Populations
Zhifa Han, Wenyang Zhou, Jian Zong, Yang Hu 0008, Shuilin Jin, Qinghua Jiang
ICIC (1)8
2018 DincRNA: a comprehensive web-based bioinformatics toolkit for exploring disease associations and ncRNA function
abstract
Summary: DincRNA aims to provide a comprehensive web-based bioinformatics toolkit to elucidate the entangled relationships among diseases and non-coding RNAs (ncRNAs) from the perspective of disease similarity. The quantitative way to illustrate relationships of pair-wise diseases always depends on their molecular mechanisms, and structures of the directed acyclic graph of Disease Ontology (DO). Corresponding methods for calculating similarity of pair-wise diseases involve Resnik's, Lin's, Wang's, PSB and SemFunSim methods. Recently, disease similarity was validated suitable for calculating functional similarities of ncRNAs and prioritizing ncRNA-disease pairs, and it has been widely applied for predicting the ncRNA function due to the limited biological knowledge from wet lab experiments of these RNAs. For this purpose, a large number of algorithms and priori knowledge need to be integrated. e.g. 'pair-wise best, pairs-average' (PBPA) and 'pair-wise all, pairs-maximum' (PAPM) methods for calculating functional similarities of ncRNAs, and random walk with restart (RWR) method for prioritizing ncRNA-disease pairs. To facilitate the exploration of disease associations and ncRNA function, DincRNA implemented all of the above eight algorithms based on DO and disease-related genes. Currently, it provides the function to query disease similarity scores, miRNA and lncRNA functional similarity scores, and the prioritization scores of lncRNA-disease and miRNA-disease pairs. Availability and implementation: http://bio-annotation.cn:18080/DincRNAClient/. Contact: [email protected] or [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Liang Cheng 0006, Yang Hu 0008, Jie Sun 0021, Meng Zhou 0003, Qinghua Jiang
Bioinform.2
2018 Identifying diseases-related metabolites using random walk
abstract
BACKGROUND: Metabolites disrupted by abnormal state of human body are deemed as the effect of diseases. In comparison with the cause of diseases like genes, these markers are easier to be captured for the prevention and diagnosis of metabolic diseases. Currently, a large number of metabolic markers of diseases need to be explored, which drive us to do this work. METHODS: The existing metabolite-disease associations were extracted from Human Metabolome Database (HMDB) using a text mining tool NCBO annotator as priori knowledge. Next we calculated the similarity of a pair-wise metabolites based on the similarity of disease sets of them. Then, all the similarities of metabolite pairs were utilized for constructing a weighted metabolite association network (WMAN). Subsequently, the network was utilized for predicting novel metabolic markers of diseases using random walk. RESULTS: Totally, 604 metabolites and 228 diseases were extracted from HMDB. From 604 metabolites, 453 metabolites are selected to construct the WMAN, where each metabolite is deemed as a node, and the similarity of two metabolites as the weight of the edge linking them. The performance of the network is validated using the leave one out method. As a result, the high area under the receiver operating characteristic curve (AUC) (0.7048) is achieved. The further case studies for identifying novel metabolites of diabetes mellitus were validated in the recent studies. CONCLUSION: In this paper, we presented a novel method for prioritizing metabolite-disease pairs. The superior performance validates its reliability for exploring novel metabolic markers of diseases.
Yang Hu 0008, Tianyi Zhao 0001, Ningyi Zhang, Tianyi Zang, Jun Zhang 0041, Liang Cheng 0006
BMC Bioinform.1
2017 Identifying diseases-related metabolites based on network
abstract
The collaborations of the diseases might be the key to understand the mechanism of the diseases since it is difficult to detect the role of complex genes and micro RNA in diseases. With the rapid development of technology, several metabolites of many kinds of diseases could be obtained by the advanced machines. Some diseases are related to several metabolites, and some metabolites have strong relationship with several diseases. Since there is certain relationship between different diseases, firstly we should find the similarity of different diseases. Then the similarity of different metabolites could be calculated by the relationship of diseases. Then a network of metabolites' similarity could be built. After building up the network, the diseases might not only relate to originally several metabolites, but more related metabolites to the disease could be found by the lines of the network. It can be used to explain the mechanism of the diseases more precisely. These metabolites could also be candidates to map the diseases in terms of the similarities. We can also sort these potentially relevant metabolites by similarity. It offers researchers a novel way to find out metabolites which is related to diseases.
Lingling Zhao, Tianyi Zhao 0001, Yang Hu 0008
BIBM3
2017 DTWscore: differential expression and cell clustering analysis for time-series single-cell RNA-seq data
abstract
BACKGROUND: The development of single-cell RNA sequencing has enabled profound discoveries in biology, ranging from the dissection of the composition of complex tissues to the identification of novel cell types and dynamics in some specialized cellular environments. However, the large-scale generation of single-cell RNA-seq (scRNA-seq) data collected at multiple time points remains a challenge to effective measurement gene expression patterns in transcriptome analysis. RESULTS: We present an algorithm based on the Dynamic Time Warping score (DTWscore) combined with time-series data, that enables the detection of gene expression changes across scRNA-seq samples and recovery of potential cell types from complex mixtures of multiple cell types. CONCLUSIONS: The DTWscore successfully classify cells of different types with the most highly variable genes from time-series scRNA-seq data. The study was confined to methods that are implemented and available within the R framework. Sample datasets and R packages are available at https://github.com/xiaoxiaoxier/DTWscore .
Shuilin Jin, Guiyou Liu, Xiurui Zhang, Deliang Wu, Yang Hu 0008, Chiping Zhang, Qinghua Jiang, Yadong Wang 0001
BMC Bioinform.7
2016 DisSetSim: An online system for calculating similarity between disease sets
abstract
Functional similarity between molecules results in similar phenotypes, such as diseases. Therefore, it is an effective way to reveal the function of molecules based on their induced diseases. However, the lack of a tool for obtaining the similarity score of pair-wise disease sets (SSDS) limits this type of application. Here, we introduce DisSetSim, an online system to solve this problem in this article. Five state-of-the-art methods involving Resnik's, Lin's, Wang's, PSB, and SemFunSim methods were implemented to measure the similarity score of pair-wise diseases (SSD) first. And then “pair-wise-best pairs-average” (PWBPA) method was implemented to calculated the SSDS by the SSD. The system was applied for calculating the functional similarity of miRNAs based on their induced disease sets. The results were further used to predict potential disease-miRNA relationships. The high area under the receiver operating characteristic curve AUC (0.9296) based on leave-one-out cross validation shows that the PWBPA method achieves a high true positive rate and a low false positive rate. The system can be accessed from http://bio-annotation.cn/DisSetSim.
Yang Hu 0008, Lingling Zhao, Zhiyan Liu, Hong Ju, Peigang Xu, Yadong Wang 0001, Liang Cheng 0006
BIBM1
2016 InfDisSim: A novel method for measuring disease similarity based on information flow
abstract
Similar diseases are often caused by their similar molecular origins, such as disease-related protein-coding genes (PCGs). And nowadays, the function of PCGs has been widely studied on a gene function network, where each node represents a gene and each edge indicates an interaction between pair-wise genes. Therefore, functional interaction between disease-related PCGs should be exploited to measure disease similarity. Actually, functional interaction of pair-wise PCGs has been introduced to calculate disease similarity recently. However, existing method ignores that genes could also be associated based on intermediate nodes in the gene functional network. Here, in this article, we proposed a novel method, InfDisSim, to infer disease similarity. InfDisSim models the information flow to the network based on random walk with damping, in which the entire network could be fully utilized. The performance of InfDisSim was evaluated by a benchmark set of similar disease pairs. The area under the receiver operating characteristic curve (AUC) was calculated to evaluate the performance. As a result, InfDisSim achieves a very high AUC (0.9786), which shows it performs well. Furthermore, based on the disease similarity computed by the infDisSim, we re-validated that similar diseases tend to have common therapeutic drugs (Pearson correlation γ2=0.1315, p=2.2e-16). Finally, InfDisSim disease similarity was exploited to construct a lncRNA similarity network (LSN), which was further applied to predict potential associations between diseases and lncRNAs. High AUC (0.9893) based on leave-one-out cross validation shows the LSN is very suitable for identifying novel disease-related lncRNAs.
Yang Hu 0008, Meng Zhou 0003, Hong Ju, Qinghua Jiang, Liang Cheng 0006
BIBM1
2016 A novel method to identify pre-microRNA in various species knowledge base
abstract
More than 1/3 of human genes are regulated by microRNAs. The identification of microRNA (miRNA) is the precondition of discovering the regulatory mechanism of miRNA and developing the cure for genetic diseases. The traditional identification method is biological experiment, but it has the defects of long period, high cost, and missing the miRNAs that only exist in a specific period or low expression level. Therefore, to overcome these defects, machine learning method is applied to identify miRNAs. In this study, for identifying real and pseudo miRNAs and classifying different species, we extracted 98 dimensional features based on the primary and secondary structure, then we proposed the BP-Adaboost method to figure out the overfitting phenomenon of BP neural network by constructing multiple BP neural network classifiers and distributed weights to these classifiers. The novel method we proposed raised the accuracy and the stability. In this study, we verified the effectiveness and superiority over other methods by experiments.
Tianyi Zhao 0001, Ningyi Zhang, Peigang Xu, Zhiyan Liu, Liang Cheng 0006, Yang Hu 0008
BIBM7
2015 Using Semantic Association to Extend and Infer Literature-Oriented Relativity Between Terms
abstract
Relative terms often appear together in the literature. Methods have been presented for weighting relativity of pairwise terms by their co-occurring literature and inferring new relationship. Terms in the literature are also in the directed acyclic graph of ontologies, such as Gene Ontology and Disease Ontology. Therefore, semantic association between terms may help for establishing relativities between terms in literature. However, current methods do not use these associations. In this paper, an adjusted R-scaled score (ARSS) based on information content (ARSSIC) method is introduced to infer new relationship between terms. First, set inclusion relationship between terms of ontology was exploited to extend relationships between these terms and literature. Next, the ARSS method was presented to measure relativity between terms across ontologies according to these extensional relationships. Then, the ARSSIC method using ratios of information shared of term's ancestors was designed to infer new relationship between terms across ontologies. The result of the experiment shows that ARSS identified more pairs of statistically significant terms based on corresponding gene sets than other methods. And the high average area under the receiver operating characteristic curve (0.9293) shows that ARSSIC achieved a high true positive rate and a low false positive rate. Data is available at http://mlg.hit.edu.cn/ARSSIC/.
Liang Cheng 0006, Jie Li 0055, Yang Hu 0008, Yongzhuang Liu, Yan-Shuo Chu, Yadong Wang 0001
IEEE ACM Trans. Comput. Biol. Bioinform.3