Liang Cheng 0006

dblp:72/5666-6 · DBLP profile ↗
← Back
22ranked-venue papers
6as first author
9since 2021 · last 2024
0000-0002-6665-6710ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 22 · 6 first-author · 9 since 2021
YearPublicationVenuePosition
2024 Attention mechanism models for precision medicine
abstract
The development of deep learning models plays a crucial role in advancing precision medicine. These models enable personalized medical treatments and interventions based on the unique genetic, environmental and lifestyle factors of individual patients, and the promotion of precision medicine is achieved mainly through genomic data analysis, variant annotation and interpretation, pharmacogenomics research, biomarker discovery, disease typing, clinical decision support and disease mechanism interpretation. Extensive research has been conducted to address precision medicine challenges using attention mechanism models such as SAN, GAT and transformers. Especially, the recent popularity of ChatGPT has significantly propelled the application of this model type to a new height. Therefore, I propose a Special Issue for Briefings in Bioinformatics about the topic 'Attention Mechanism Models for Precision Medicine'. This Special Issue aims to provide a comprehensive overview and presentation of innovative researches on the application of graph attention mechanism models in precision medicine.
Liang Cheng 0006
Briefings Bioinform.1
2023 iATMEcell: identification of abnormal tumor microenvironment cells to predict the clinical outcomes in cancer based on cell-cell crosstalk network
abstract
Interactions between Tumor microenvironment (TME) cells shape the unique growth environment, sustaining tumor growth and causing the immune escape of tumor cells. Nonetheless, no studies have reported a systematic analysis of cellular interactions in the identification of cancer-related TME cells. Here, we proposed a novel network-based computational method, named as iATMEcell, to identify the abnormal TME cells associated with the biological outcome of interest based on a cell-cell crosstalk network. In the method, iATMEcell first manually collected TME cell types from multiple published studies and obtained their corresponding gene signatures. Then, a weighted cell-cell crosstalk network was constructed in the context of a specific cancer bulk tissue transcriptome data, where the weight between cells reflects both their biological function similarity and the transcriptional dysregulated activities of gene signatures shared by them. Finally, it used a network propagation algorithm to identify significantly dysregulated TME cells. Using the cancer genome atlas (TCGA) Bladder Urothelial Carcinoma training set and two independent validation sets, we illustrated that iATMEcell could identify significant abnormal cells associated with patient survival and immunotherapy response. iATMEcell was further applied to a pan-cancer analysis, which revealed that four common abnormal immune cells play important roles in the patient prognosis across multiple cancer types. Collectively, we demonstrated that iATMEcell could identify potentially abnormal TME cells based on a cell-cell crosstalk network, which provided a new insight into understanding the effect of TME cells in cancer. iATMEcell is developed as an R package, which is freely available on GitHub (https://github.com/hanjunwei-lab/iATMEcell).
Yuqi Sheng, Jiashuo Wu, Xiangmei Li, Jiayue Qiu, Qinyu Ge, Liang Cheng 0006, Junwei Han 0003
Briefings Bioinform.7
2022 Omics data analysis and integration for COVID-19 patients - editorial
abstract
In the latest year, the SARS-CoV-2 virus has spread around the world leading to the explosion of the coronavirus disease 2019 (COVID-19) pandemic. Up to now over 514 000 000 patients with > 6 200 000 deaths have been reported. Since its rapid spread and high case fatality ratio, researchers have concentrated in exposing the origin, mutational tendency, pathogenesis and vaccine of the virus. Sequencing is producing large amounts of Omics data for SARS-CoV-2 virus and COVID-19 patients. For example, over 300 000 SARS-CoV-2 genomes were reported in GISAID since the first publication of the SARS-CoV-2 genome on 24 January 2020. Thousands of COVID-19 patients were sequenced for screening susceptible SNPs. Currently one of the major challenge is to mine casual molecules and phenotypes by integrating multi-level Omics data using system biology methods, which may expand our knowledge of curing COVID-19. This special issue aims to provide a comprehensive overview and innovative penetration, including but not limit to integrative methods and tools for analyzing multi-level Omics data, design of novel methods for exposing casual molecules and phenotypes of COVID-19 patients, design of novel methods for identifying novel drug targets, identification of molecular signatures of COVID-19 patients and identification of mutational tendency of SARS-CoV-2.
Liang Cheng 0006
Briefings Bioinform.1
2022 A comprehensive review of the analysis and integration of omics data for SARS-CoV-2 and COVID-19
abstract
Since the first report of severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) in December 2019, over 100 million people have been infected by COVID-19, millions of whom have died. In the latest year, a large number of omics data have sprung up and helped researchers broadly study the sequence, chemical structure and function of SARS-CoV-2, as well as molecular abnormal mechanisms of COVID-19 patients. Though some successes have been achieved in these areas, it is necessary to analyze and mine omics data for comprehensively understanding SARS-CoV-2 and COVID-19. Hence, we reviewed the current advantages and limitations of the integration of omics data herein. Firstly, we sorted out the sequence resources and database resources of SARS-CoV-2, including protein chemical structure, potential drug information and research literature resources. Next, we collected omics data of the COVID-19 hosts, including genomics, transcriptomics, microbiology and potential drug information data. And subsequently, based on the integration of omics data, we summarized the existing data analysis methods and the related research results of COVID-19 multi-omics data in recent years. Finally, we put forward SARS-CoV-2 (COVID-19) multi-omics data integration research direction and gave a case study to mine deeper for the disease mechanisms of COVID-19.
Zijun Zhu, Ping Wang 0033, Jianxing Bi, Liang Cheng 0006
Briefings Bioinform.6
2022 MGPLI: exploring multigranular representations for protein-ligand interaction prediction
abstract
MOTIVATION: The capability to predict the potential drug binding affinity against a protein target has always been a fundamental challenge in silico drug discovery. The traditional experiments in vitro and in vivo are costly and time-consuming which need to search over large compound space. Recent years have witnessed significant success on deep learning-based models for drug-target binding affinity prediction task. RESULTS: Following the recent success of the Transformer model, we propose a multigranularity protein-ligand interaction (MGPLI) model, which adopts the Transformer encoders to represent the character-level features and fragment-level features, modeling the possible interaction between residues and atoms or their segments. In addition, we use the convolutional neural network to extract higher-level features based on transformer encoder outputs and a highway layer to fuse the protein and drug features. We evaluate MGPLI on different protein-ligand interaction datasets and show the improvement of prediction performance compared to state-of-the-art baselines. AVAILABILITY AND IMPLEMENTATION: The model scripts are available at https://github.com/IILab-Resource/MGDTA.git.
Junjie Wang 0005, Huiting Sun, Mengdie Xu, Liang Cheng 0006
Bioinform.7
2021 Functional alterations caused by mutations reflect evolutionary trends of SARS-CoV-2
abstract
Since the first report of severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) in December 2019, the COVID-19 pandemic has spread rapidly worldwide. Due to the limited virus strains, few key mutations that would be very important with the evolutionary trends of virus genome were observed in early studies. Here, we downloaded 1809 sequence data of SARS-CoV-2 strains from GISAID before April 2020 to identify mutations and functional alterations caused by these mutations. Totally, we identified 1017 nonsynonymous and 512 synonymous mutations with alignment to reference genome NC_045512, none of which were observed in the receptor-binding domain (RBD) of the spike protein. On average, each of the strains could have about 1.75 new mutations each month. The current mutations may have few impacts on antibodies. Although it shows the purifying selection in whole-genome, ORF3a, ORF8 and ORF10 were under positive selection. Only 36 mutations occurred in 1% and more virus strains were further analyzed to reveal linkage disequilibrium (LD) variants and dominant mutations. As a result, we observed five dominant mutations involving three nonsynonymous mutations C28144T, C14408T and A23403G and two synonymous mutations T8782C, and C3037T. These five mutations occurred in almost all strains in April 2020. Besides, we also observed two potential dominant nonsynonymous mutations C1059T and G25563T, which occurred in most of the strains in April 2020. Further functional analysis shows that these mutations decreased protein stability largely, which could lead to a significant reduction of virus virulence. In addition, the A23403G mutation increases the spike-ACE2 interaction and finally leads to the enhancement of its infectivity. All of these proved that the evolution of SARS-CoV-2 is toward the enhancement of infectivity and reduction of virulence.
Liang Cheng 0006, Zijun Zhu, Changlu Qi, Ping Wang 0033
Briefings Bioinform.1
2021 CNA2Subpathway: identification of dysregulated subpathway driven by copy number alterations in cancer
abstract
Biological pathways reflect the key cellular mechanisms that dictate disease states, drug response and altered cellular function. The local areas of pathways are defined as subpathways (SPs), whose dysfunction has been reported to be associated with the occurrence and development of cancer. With the development of high-throughput sequencing technology, identifying dysfunctional SPs by using multi-omics data has become possible. Moreover, the SPs are not isolated in the biological system but interact with each other. Here, we propose a network-based calculated method, CNA2Subpathway, to identify dysfunctional SPs is driven by somatic copy number alterations (CNAs) in cancer through integrating pathway topology information, multi-omics data and SP crosstalk. This provides a novel way of SP analysis by using the SP interactions in the system biological level. Using data sets from breast cancer and head and neck cancer, we validate the effectiveness of CNA2Subpathway in identifying cancer-relevant SPs driven by the somatic CNAs, which are also shown to be associated with cancer immune and prognosis of patients. We further compare our results with five pathway or SP analysis methods based on CNA and gene expression data without considering SP crosstalk. With these analyses, we show that CNA2Subpathway could help to uncover dysfunctional SPs underlying cancer via the use of SP crosstalk. CNA2Subpathway is developed as an R-based tool, which is freely available on GitHub (https://github.com/hanjunwei-lab/CNA2Subpathway).
Yuqi Sheng, Yang Yang 0009, Xiangmei Li, Jiayue Qiu, Jiashuo Wu, Liang Cheng 0006, Junwei Han 0003
Briefings Bioinform.7
2021 Deep-DRM: a computational method for identifying disease-related metabolites based on graph deep learning approaches
abstract
MOTIVATION: The functional changes of the genes, RNAs and proteins will eventually be reflected in the metabolic level. Increasing number of researchers have researched mechanism, biomarkers and targeted drugs by metabolites. However, compared with our knowledge about genes, RNAs, and proteins, we still know few about diseases-related metabolites. All the few existed methods for identifying diseases-related metabolites ignore the chemical structure of metabolites, fail to recognize the association pattern between metabolites and diseases, and fail to apply to isolated diseases and metabolites. RESULTS: In this study, we present a graph deep learning based method, named Deep-DRM, for identifying diseases-related metabolites. First, chemical structures of metabolites were used to calculate similarities of metabolites. The similarities of diseases were obtained based on their functional gene network and semantic associations. Therefore, both metabolites and diseases network could be built. Next, Graph Convolutional Network (GCN) was applied to encode the features of metabolites and diseases, respectively. Then, the dimension of these features was reduced by Principal components analysis (PCA) with retainment 99% information. Finally, Deep neural network was built for identifying true metabolite-disease pairs (MDPs) based on these features. The 10-cross validations on three testing setups showed outstanding AUC (0.952) and AUPR (0.939) of Deep-DRM compared with previous methods and similar approaches. Ten of top 15 predicted associations between diseases and metabolites got support by other studies, which suggests that Deep-DRM is an efficient method to identify MDPs. CONTACT: [email protected]. AVAILABILITY AND IMPLEMENTATION: https://github.com/zty2009/GPDNN-for-Identify-ing-Disease-related-Metabolites.
Tianyi Zhao 0001, Yang Hu 0008, Liang Cheng 0006
Briefings Bioinform.3
2021 SubtypeDrug: a software package for prioritization of candidate cancer subtype-specific drugs
abstract
SUMMARY: Cancer can be classified into various subtypes by its molecular, histological or clinical characteristics. Discovering cancer-subtype-specific drugs is a crucial step in personalized medicine. SubtypeDrug is a system biology R-based software package that enables the prioritization of subtype-specific drugs based on cancer expression data from samples of many subtypes. This provides a novel approach to identify the subtype-specific drug by considering biological functions regulated by drugs at the subpathway level. The operation modes include extraction of subpathways from biological pathways, identification of dysregulated subpathways induced by each drug, inference of sample-specific subpathway activity profiles, evaluation of drug-disease reverse association at the subpathways level, identification of cancer-subtype-specific drugs through subtype sample set enrichment analysis, and visualization of the results. Its capabilities enable SubtypeDrug to find subtype-specific drugs, which will fill the gaps in the recent tools which only identify the drugs for a particular cancer type. SubtypeDrug may help to facilitate the development of tailored treatment for patients with cancer. AVAILABILITY AND IMPLEMENTATION: The package is implemented in R and available under GPL-2 license from the CRAN website (https://CRAN.R-project.org/package=SubtypeDrug). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Qingfei Kong, Chonghui Liu, Liang Cheng 0006, Junwei Han 0003
Bioinform.4
2020 OntoSem: an Ontology Semantic Representation Methodology for Biomedical Domain
abstract
Ontologies are essential description tools for biomedical concepts and entities, supporting biomedical fundamental research such as semantic similarity analysis, protein-protein interaction prediction and so on. An increasing amount of ontology-like domain knowledge is published in scientific publications, meanwhile, advanced natural language processing (NLP) techniques have been widespread to extract information from text resources automatically, both of which facilitate the exploration of the semantic representation of biomedical ontologies. We propose a novel distributional semantic representation methodology based on the combination of two pre-trained and domain-specific word embedding tools, the non-contextualized Word2Vec and the context-dependent NCBI-blueBERT, to enhance the encoding ability for biomedical ontologies. Furthermore, we utilize a randomly initialized bidirectional LSTM to project the obtained word vector sequence to a fixed-length sentence vector, facilitating a flexible and uniform way for the computation of downstream tasks. We evaluate our method in two categories of tasks: the similarity access of ontology terms, and the ontology annotation-based protein-protein interaction classification. Experimental results demonstrate that our method provides encouraging results compared to the baselines in all tests. Our approach offers promising opportunities for representing ontologies semantics and in turn characterizing entities including proteins in biomedical research.
Lingling Zhao, Junjie Wang 0005, Liang Cheng 0006, Chunyu Wang 0002
BIBM3
2020 Identification of gene signature associated with type 2 diabetes mellitus by integrating mutation and expression data
abstract
Type 2 diabetes mellitus (T2DM) is a frequency occurred chronic disease.The early diagnosis could be very helpful for the treatment of T2DM patients .With the development of sequencing technology, a large number of differentially expressed genes were identified from expression data.However, the method of machine learning can only identify the local optimal solution as the signature.The mutation information obtained by inheritance can better reflect the relationship between genes and diseases.Therefore, we need to integrate mutation information to more accurately identify the signature .To this end, we integrated genome-wide association study (GWAS) data and expression data, combined with expression quantitative trait loci (eQTL) technology to get T2DM predictive signature (T2DMSig-10) .Firstly, we used GWAS data to obtain a list of T2DM susceptible loci.Then, we used eQTL technology to locate risk single nucleotide polymorphisms (SNPs) to genes, and combined with the pancreatic ß-cells gene expression data to obtain 10 protein-coding genes .Next, we combined these genes with equal weights .After receiving receiver operating characteristic (ROC), single gene removal method, gene ontology function enrichment and protein-protein interaction network were used to verify, the results showed that T2DMSig-10 had an excellent predictive effect on T2DM (AUC=0 .99), and was highly robust. In short, we obtained the predictive signature of T2DM, and further analyzed and verified it.
Zijun Zhu, Liang Cheng 0006
BIBM3
2020 psSubpathway: a software package for flexible identification of phenotype-specific subpathways in cancer progression
abstract
SUMMARY: Subpathways, which are defined as local gene subregions within a biological pathway, have been reported to be associated with the occurrence and development of cancer. The recent subpathway identification tools generally identify differentially expressed subpathways between normal and cancer samples. psSubpathway is a novel systems biology R-based software package that enables flexible identification of phenotype-specific subpathways in a cancer dataset with multiple categories (such as multiple subtypes and developmental stages of cancer). The operation modes include extraction of subpathways from pathway networks, inference with subpathway activities in the context of gene expression data, identification of subtype-specific subpathways, identification of dynamic-changed subpathways associated with the cancer developmental stage and visualization of subpathway activities of samples in different phenotypes. Its capabilities enable psSubpathway to find specific abnormal subpathways in the datasets with multi-phenotype categories and to fill the gaps in the recent tools. psSubpathway may identify more specific biomarkers to facilitate the development of tailored treatment for patients with cancer. AVAILABILITY AND IMPLEMENTATION: The package is implemented in R and available under GPL-2 license from the CRAN website (https://cran.r-project.org/web/packages/psSubpathway/). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Junwei Han 0003, Qingfei Kong, Liang Cheng 0006
Bioinform.4
2020 DeepLGP: a novel deep learning method for prioritizing lncRNA target genes
abstract
MOTIVATION: Although long non-coding RNAs (lncRNAs) have limited capacity for encoding proteins, they have been verified as biomarkers in the occurrence and development of complex diseases. Recent wet-lab experiments have shown that lncRNAs function by regulating the expression of protein-coding genes (PCGs), which could also be the mechanism responsible for causing diseases. Currently, lncRNA-related biological data are increasing rapidly. Whereas, no computational methods have been designed for predicting the novel target genes of lncRNA. RESULTS: In this study, we present a graph convolutional network (GCN) based method, named DeepLGP, for prioritizing target PCGs of lncRNA. First, gene and lncRNA features were selected, these included their location in the genome, expression in 13 tissues and miRNA-mediated lncRNA-gene pairs. Next, GCN was applied to convolve a gene interaction network for encoding the features of genes and lncRNAs. Then, these features were used by the convolutional neural network for prioritizing target genes of lncRNAs. In 10-cross validations on two independent datasets, DeepLGP obtained high area under curves (0.90-0.98) and area under precision-recall curves (0.91-0.98). We found that lncRNA pairs with high similarity had more overlapped target genes. Further experiments showed that genes targeted by the same lncRNA sets had a strong likelihood of causing the same diseases, which could help in identifying disease-causing PCGs. AVAILABILITY AND IMPLEMENTATION: https://github.com/zty2009/LncRNA-target-gene. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Tianyi Zhao 0001, Yang Hu 0008, Jiajie Peng, Liang Cheng 0006, Pier Luigi Martelli
Bioinform.4
2019 MetSigDis: a manually curated resource for the metabolic signatures of diseases
abstract
Complex diseases cannot be understood only on the basis of single gene, single mRNA transcript or single protein but the effect of their collaborations. The combination consequence in molecular level can be captured by the alterations of metabolites. With the rapidly developing of biomedical instruments and analytical platforms, a large number of metabolite signatures of complex diseases were identified and documented in the literature. Biologists' hardship in the face of this large amount of papers recorded metabolic signatures of experiments' results calls for an automated data repository. Therefore, we developed MetSigDis aiming to provide a comprehensive resource of metabolite alterations in various diseases. MetSigDis is freely available at http://www.bio-annotation.cn/MetSigDis/. By reviewing hundreds of publications, we collected 6849 curated relationships between 2420 metabolites and 129 diseases across eight species involving Homo sapiens and model organisms. All of these relationships were used in constructing a metabolite disease network (MDN). This network displayed scale-free characteristics according to the degree distribution (power-law distribution with R2 = 0.909), and the subnetwork of MDN for interesting diseases and their related metabolites can be visualized in the Web. The common alterations of metabolites reflect the metabolic similarity of diseases, which is measured using Jaccard index. We observed that metabolite-based similar diseases are inclined to share semantic associations of Disease Ontology. A human disease network was then built, where a node represents a disease, and an edge indicates similarity of pair-wise diseases. The network validated the observation that linked diseases based on metabolites should have more overlapped genes.
Liang Cheng 0006, Haixiu Yang, Hengqiang Zhao, Xiaoya Pei, Jie Sun 0021, Meng Zhou 0003
Briefings Bioinform.1
2019 Identifying Alzheimer's disease-related proteins by LRRGD
abstract
BACKGROUND: Alzheimer's disease (AD) imposes a heavy burden on society and every family. Therefore, diagnosing AD in advance and discovering new drug targets are crucial, while these could be achieved by identifying AD-related proteins. The time-consuming and money-costing biological experiment makes researchers turn to develop more advanced algorithms to identify AD-related proteins. RESULTS: Firstly, we proposed a hypothesis "similar diseases share similar related proteins". Therefore, five similarity calculation methods are introduced to find out others diseases which are similar to AD. Then, these diseases' related proteins could be obtained by public data set. Finally, these proteins are features of each disease and could be used to map their similarity to AD. We developed a novel method 'LRRGD' which combines Logistic Regression (LR) and Gradient Descent (GD) and borrows the idea of Random Forest (RF). LR is introduced to regress features to similarities. Borrowing the idea of RF, hundreds of LR models have been built by randomly selecting 40 features (proteins) each time. Here, GD is introduced to find out the optimal result. To avoid the drawback of local optimal solution, a good initial value is selected by some known AD-related proteins. Finally, 376 proteins are found to be related to AD. CONCLUSION: Three hundred eight of three hundred seventy-six proteins are the novel proteins. Three case studies are done to prove our method's effectiveness. These 308 proteins could give researchers a basis to do biological experiments to help treatment and diagnostic AD.
Tianyi Zhao 0001, Yang Hu 0008, Tianyi Zang, Liang Cheng 0006
BMC Bioinform.4
2018 A Novel Method for Identifying Alzheimer's Disease-related Proteins
Yang Hu 0008, Jun Zhang 0041, Tianyi Zhao 0001, Liang Cheng 0006, Tianyi Zang
BIBM4
2018 DincRNA: a comprehensive web-based bioinformatics toolkit for exploring disease associations and ncRNA function
abstract
Summary: DincRNA aims to provide a comprehensive web-based bioinformatics toolkit to elucidate the entangled relationships among diseases and non-coding RNAs (ncRNAs) from the perspective of disease similarity. The quantitative way to illustrate relationships of pair-wise diseases always depends on their molecular mechanisms, and structures of the directed acyclic graph of Disease Ontology (DO). Corresponding methods for calculating similarity of pair-wise diseases involve Resnik's, Lin's, Wang's, PSB and SemFunSim methods. Recently, disease similarity was validated suitable for calculating functional similarities of ncRNAs and prioritizing ncRNA-disease pairs, and it has been widely applied for predicting the ncRNA function due to the limited biological knowledge from wet lab experiments of these RNAs. For this purpose, a large number of algorithms and priori knowledge need to be integrated. e.g. 'pair-wise best, pairs-average' (PBPA) and 'pair-wise all, pairs-maximum' (PAPM) methods for calculating functional similarities of ncRNAs, and random walk with restart (RWR) method for prioritizing ncRNA-disease pairs. To facilitate the exploration of disease associations and ncRNA function, DincRNA implemented all of the above eight algorithms based on DO and disease-related genes. Currently, it provides the function to query disease similarity scores, miRNA and lncRNA functional similarity scores, and the prioritization scores of lncRNA-disease and miRNA-disease pairs. Availability and implementation: http://bio-annotation.cn:18080/DincRNAClient/. Contact: [email protected] or [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Liang Cheng 0006, Yang Hu 0008, Jie Sun 0021, Meng Zhou 0003, Qinghua Jiang
Bioinform.1
2018 Identifying diseases-related metabolites using random walk
abstract
BACKGROUND: Metabolites disrupted by abnormal state of human body are deemed as the effect of diseases. In comparison with the cause of diseases like genes, these markers are easier to be captured for the prevention and diagnosis of metabolic diseases. Currently, a large number of metabolic markers of diseases need to be explored, which drive us to do this work. METHODS: The existing metabolite-disease associations were extracted from Human Metabolome Database (HMDB) using a text mining tool NCBO annotator as priori knowledge. Next we calculated the similarity of a pair-wise metabolites based on the similarity of disease sets of them. Then, all the similarities of metabolite pairs were utilized for constructing a weighted metabolite association network (WMAN). Subsequently, the network was utilized for predicting novel metabolic markers of diseases using random walk. RESULTS: Totally, 604 metabolites and 228 diseases were extracted from HMDB. From 604 metabolites, 453 metabolites are selected to construct the WMAN, where each metabolite is deemed as a node, and the similarity of two metabolites as the weight of the edge linking them. The performance of the network is validated using the leave one out method. As a result, the high area under the receiver operating characteristic curve (AUC) (0.7048) is achieved. The further case studies for identifying novel metabolites of diabetes mellitus were validated in the recent studies. CONCLUSION: In this paper, we presented a novel method for prioritizing metabolite-disease pairs. The superior performance validates its reliability for exploring novel metabolic markers of diseases.
Yang Hu 0008, Tianyi Zhao 0001, Ningyi Zhang, Tianyi Zang, Jun Zhang 0041, Liang Cheng 0006
BMC Bioinform.6
2016 DisSetSim: An online system for calculating similarity between disease sets
abstract
Functional similarity between molecules results in similar phenotypes, such as diseases. Therefore, it is an effective way to reveal the function of molecules based on their induced diseases. However, the lack of a tool for obtaining the similarity score of pair-wise disease sets (SSDS) limits this type of application. Here, we introduce DisSetSim, an online system to solve this problem in this article. Five state-of-the-art methods involving Resnik's, Lin's, Wang's, PSB, and SemFunSim methods were implemented to measure the similarity score of pair-wise diseases (SSD) first. And then “pair-wise-best pairs-average” (PWBPA) method was implemented to calculated the SSDS by the SSD. The system was applied for calculating the functional similarity of miRNAs based on their induced disease sets. The results were further used to predict potential disease-miRNA relationships. The high area under the receiver operating characteristic curve AUC (0.9296) based on leave-one-out cross validation shows that the PWBPA method achieves a high true positive rate and a low false positive rate. The system can be accessed from http://bio-annotation.cn/DisSetSim.
Yang Hu 0008, Lingling Zhao, Zhiyan Liu, Hong Ju, Peigang Xu, Yadong Wang 0001, Liang Cheng 0006
BIBM8
2016 InfDisSim: A novel method for measuring disease similarity based on information flow
abstract
Similar diseases are often caused by their similar molecular origins, such as disease-related protein-coding genes (PCGs). And nowadays, the function of PCGs has been widely studied on a gene function network, where each node represents a gene and each edge indicates an interaction between pair-wise genes. Therefore, functional interaction between disease-related PCGs should be exploited to measure disease similarity. Actually, functional interaction of pair-wise PCGs has been introduced to calculate disease similarity recently. However, existing method ignores that genes could also be associated based on intermediate nodes in the gene functional network. Here, in this article, we proposed a novel method, InfDisSim, to infer disease similarity. InfDisSim models the information flow to the network based on random walk with damping, in which the entire network could be fully utilized. The performance of InfDisSim was evaluated by a benchmark set of similar disease pairs. The area under the receiver operating characteristic curve (AUC) was calculated to evaluate the performance. As a result, InfDisSim achieves a very high AUC (0.9786), which shows it performs well. Furthermore, based on the disease similarity computed by the infDisSim, we re-validated that similar diseases tend to have common therapeutic drugs (Pearson correlation γ2=0.1315, p=2.2e-16). Finally, InfDisSim disease similarity was exploited to construct a lncRNA similarity network (LSN), which was further applied to predict potential associations between diseases and lncRNAs. High AUC (0.9893) based on leave-one-out cross validation shows the LSN is very suitable for identifying novel disease-related lncRNAs.
Yang Hu 0008, Meng Zhou 0003, Hong Ju, Qinghua Jiang, Liang Cheng 0006
BIBM6
2016 A novel method to identify pre-microRNA in various species knowledge base
abstract
More than 1/3 of human genes are regulated by microRNAs. The identification of microRNA (miRNA) is the precondition of discovering the regulatory mechanism of miRNA and developing the cure for genetic diseases. The traditional identification method is biological experiment, but it has the defects of long period, high cost, and missing the miRNAs that only exist in a specific period or low expression level. Therefore, to overcome these defects, machine learning method is applied to identify miRNAs. In this study, for identifying real and pseudo miRNAs and classifying different species, we extracted 98 dimensional features based on the primary and secondary structure, then we proposed the BP-Adaboost method to figure out the overfitting phenomenon of BP neural network by constructing multiple BP neural network classifiers and distributed weights to these classifiers. The novel method we proposed raised the accuracy and the stability. In this study, we verified the effectiveness and superiority over other methods by experiments.
Tianyi Zhao 0001, Ningyi Zhang, Peigang Xu, Zhiyan Liu, Liang Cheng 0006, Yang Hu 0008
BIBM6
2015 Using Semantic Association to Extend and Infer Literature-Oriented Relativity Between Terms
abstract
Relative terms often appear together in the literature. Methods have been presented for weighting relativity of pairwise terms by their co-occurring literature and inferring new relationship. Terms in the literature are also in the directed acyclic graph of ontologies, such as Gene Ontology and Disease Ontology. Therefore, semantic association between terms may help for establishing relativities between terms in literature. However, current methods do not use these associations. In this paper, an adjusted R-scaled score (ARSS) based on information content (ARSSIC) method is introduced to infer new relationship between terms. First, set inclusion relationship between terms of ontology was exploited to extend relationships between these terms and literature. Next, the ARSS method was presented to measure relativity between terms across ontologies according to these extensional relationships. Then, the ARSSIC method using ratios of information shared of term's ancestors was designed to infer new relationship between terms across ontologies. The result of the experiment shows that ARSS identified more pairs of statistically significant terms based on corresponding gene sets than other methods. And the high average area under the receiver operating characteristic curve (0.9293) shows that ARSSIC achieved a high true positive rate and a low false positive rate. Data is available at http://mlg.hit.edu.cn/ARSSIC/.
Liang Cheng 0006, Jie Li 0055, Yang Hu 0008, Yongzhuang Liu, Yan-Shuo Chu, Yadong Wang 0001
IEEE ACM Trans. Comput. Biol. Bioinform.1