EDBT 2026 Demo / reviewers in the wild / expert
Lin Gao 0006
dblp:92/2834-6
· DBLP profile ↗
50ranked-venue papers
0as first author
29since 2021 · last 2026
0000-0001-6396-0787ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 45 · 28 since 2021Artificial intelligence and machine learning · 3 · 1 since 2021Databases, data management, data science and information retrieval · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Semantic-enhanced heterogeneous graph learning for identifying ncRNAs associated with drug resistanceabstractMOTIVATION: Identifying non-coding RNAs (ncRNAs) associated with drug resistance is critical for elucidating molecular mechanisms underlying drug response, facilitating drug screening, and discovering novel therapeutic targets. While several graph neural network-based methods have been proposed to infer ncRNA-drug resistance associations, they remain fundamentally constrained by semantic distortion induced by a sparse bipartite network and neglect of relational semantics among molecular entities, ultimately compromising both predictive reliability and biological interpretability. RESULTS: In this study, we propose iNcRD-HG, a novel framework for identifying ncRNA-drug resistance associations. The framework addresses three critical aspects: constructing a context-enriched heterogeneous network that integrates six distinct molecular interaction types with bio-entity-specific attributes, developing a semantic-enhanced graph learning architecture that implements relation-type-aware message passing to capture complex contextual dependencies, and introducing an interpretability mechanism to reveal potential synergistic pathways underlying drug response. Experimental results demonstrate that iNcRD-HG achieves superior predictive performance across diverse benchmark datasets while deriving association features with strong discriminative capability. By identifying molecular synergistic contexts, iNcRD-HG provides mechanistically interpretable insights into ncRNA-mediated drug resistance. AVAILABILITY AND IMPLEMENTATION: Datasets and source codes are available at https://github.com/Biohang/iNcRD-HG. Hang Wei 0005, Yuran Xie, Wenxiang Zhang, Linyang Li, Shuai Wu 0001, Lin Gao 0006 |
Bioinform. | 6 |
| 2026 | Bridging cancer cell-intrinsic driver genes and -extrinsic cell-cell communication with Driver2CommabstractTumor development and progression are affected not only by cancer cell-intrinsic factors comprising complex genetic variations, but also by -extrinsic factors such as cell-cell communication (CCC)-mediated immunosuppression. However, whether and how these two types of factors influence each other remains an open question. We present Driver2Comm, a general computational framework designed to systematically identify intrinsic-extrinsic (IE) pathways that functionally connect cancer cell driver genes with their associated CCC signatures in the tumor microenvironment (TME). By applying Driver2Comm to single-cell and spatial transcriptomic datasets of multiple cancer types, we find that driver gene-associated CCC signatures play critical roles in immune regulation, metastasis, and therapy response. These signatures not only illuminate mechanisms of TME remodeling but also demonstrate clinical value in predicting patient survival and response to immune checkpoint blockade. Furthermore, Driver2Comm captures higher-order, cell-type-pair-specific CCC functional modules and spatially coherent CCC patterns in tissue contexts. As a generalizable tool, Driver2Comm bridges cancer genomics and cellular ecosystems, offering insights into biomarker discovery and combination therapy strategies. Runzhi Xie, Junping Li, Yuxuan Hu 0004, Lin Gao 0006 |
PLoS Comput. Biol. | 4 |
| 2025 | Functional Module Identification in Spatial Cellular Communication Networks via Variational Graph AutoencoderabstractCell-cell communication plays a pivotal role in tissue development, homeostasis, and disease progression. However, existing approaches often lack systematic modeling of networklevel structure and often overlook spatial modular organization. To address these limitations, we propose CCNVGAE, an unsupervised framework based on variational graph autoencoder (VGAE) that integrates spatial coordinates, ligand-receptor gene expression and graph structural information for functional module identification. Experiments on spatial transcriptomic datasets from mouse brain and human lung cancer demonstrate that CCNVGAE achieves competitive performance in spatial segmentation and biological interpretability. Moreover, it effectively uncovers module specific ligand-receptor interactions and their potential regulatory mechanisms, providing a powerful tool for elucidating spatial communication structures in complex tissues. Xingli Guo, Yuxuan Hu 0004, Bingbo Wang, Lin Gao 0006 |
BIBM | 6 |
| 2025 | DCAM: A Deep Context-Aware Model for Predicting the Function of Genomic Regulatory RegionsabstractAccurately annotating non-coding regulatory DNA, critical for gene regulation and implicated in numerous diseases, remains a major challenge. Current deep learning models struggle to integrate local motifs with long-range dependencies due to limited architectures and simplistic feature fusion. We propose DCAM, a dual-branch neural network where a global branch captures long-range context via a CNN-BiLSTM, while a local branch extracts fine-grained motifs using dna2vec and CNNs. These features are dynamically integrated via a sequential channel-spatial attention mechanism. Evaluated on diverse benchmark datasets, DCAM significantly outperforms state-of-the-art models. DCAM also accurately predicts the functional impact of single nucleotide variants (SNVs), establishing it as a powerful tool for prioritizing pathogenic non-coding variants. Bingbo Wang, Xingli Guo, Lin Gao 0006 |
BIBM | 6 |
| 2025 | RNA language model and graph attention network for RNA and small molecule binding sites predictionabstractMOTIVATION: The structural complexities enable RNA to serve as a versatile molecular scaffold capable of binding small molecules with high specificity. Understanding these interactions is essential for elucidating RNA's role in disease mechanisms and developing RNA-targeted therapeutics. However, predicting RNA-small molecule binding sites remains a significant challenge due to their conformational flexibility, structural diversity, and the limited availability of high-resolution structural data. RESULTS: In this study, we propose RLsite, a novel computational framework integrating pre-trained RNA language models with graph attention networks (GAT) to predict small-molecule binding sites on RNA. Our method effectively captures both sequential and structural features of RNA by leveraging large-scale RNA sequence data to learn intrinsic patterns and processing graph-based RNA structures to highlight key topological and spatial features. Compared to existing methods, RLsite demonstrates superior accuracy, generalizability, and biological relevance, achieving a Precision of 0.749, a Recall of 0.654, an MCC of 0.474, and an AUC of 0.828 on the public test set, which significantly outperforms the previous models, such as CapBind (an AUC of 0.770), MultiModRLBP (an AUC of 0.780), and RNABind (an AUC of 0.471). Notably, a case study of the PreQ1 riboswitch has achieved strong predictive performance (AUC = 0.97, Recall = 0.9), and its predicted binding sites have been confirmed experimentally. These results underscore our method as a potentially powerful tool for RNA-targeted drug discovery and advancing our understanding of RNA-ligand interactions. AVAILABILITY AND IMPLEMENTATION: The resource codes and data can be accessed at https://github.com/SaisaiSun/RLsite. Saisai Sun, Jianyi Yang 0002, Lin Gao 0006, Pengyong Li |
Bioinform. | 3 |
| 2024 | CellFeature: Cell and Feature Co-Embedding from Single-Cell Multi-Omics with Heterogeneous Graph ModelabstractMost current single-cell multi-omics analysis methods are limited to the co-embedding of different omics cells and lack the ability to directly analyze the relationship between different omics cells and different types of features. Here, we describe a multi-omics cells and heterogeneous features co-embedding algorithm, named CellFeature, by eliminating the heterogeneity between different types of nodes based on heterogeneous graph representation learning framework. Using different single-cell multimodal datasets, we compared multi-omics cell embeddings with current state-of-the-art methods and achieved comparable results. By leveraging co-embedding of cells and features, we demonstrate that CellFeature can identify cell-type-specific features with better performance than existing methods. Based on the co-embeddings, CellFeature can also perform cell subtypes discovery and enable trajectory-specific genes identification. Enling Li, Lin Gao 0006, Yusen Ye |
BIBM | 2 |
| 2024 | Exploring Hierarchical Structures of Cell Types in scRNA-seq Data
Haojie Zhai, Yusen Ye, Yuxuan Hu 0004, Lin Gao 0006 |
ISBRA (2) | 5 |
| 2024 | DeepGRNCS: deep learning-based framework for jointly inferring gene regulatory networks across cell subpopulationsabstractInferring gene regulatory networks (GRNs) allows us to obtain a deeper understanding of cellular function and disease pathogenesis. Recent advances in single-cell RNA sequencing (scRNA-seq) technology have improved the accuracy of GRN inference. However, many methods for inferring individual GRNs from scRNA-seq data are limited because they overlook intercellular heterogeneity and similarities between different cell subpopulations, which are often present in the data. Here, we propose a deep learning-based framework, DeepGRNCS, for jointly inferring GRNs across cell subpopulations. We follow the commonly accepted hypothesis that the expression of a target gene can be predicted based on the expression of transcription factors (TFs) due to underlying regulatory relationships. We initially processed scRNA-seq data by discretizing data scattering using the equal-width method. Then, we trained deep learning models to predict target gene expression from TFs. By individually removing each TF from the expression matrix, we used pre-trained deep model predictions to infer regulatory relationships between TFs and genes, thereby constructing the GRN. Our method outperforms existing GRN inference methods for various simulated and real scRNA-seq datasets. Finally, we applied DeepGRNCS to non-small cell lung cancer scRNA-seq data to identify key genes in each cell subpopulation and analyzed their biological relevance. In conclusion, DeepGRNCS effectively predicts cell subpopulation-specific GRNs. The source code is available at https://github.com/Nastume777/DeepGRNCS. Yahui Lei, Xingli Guo, Kei Hang Katie Chan, Lin Gao 0006 |
Briefings Bioinform. | 5 |
| 2024 | Statistical modeling and significance estimation of multi-way chromatin contacts with HyperloopFinderabstractRecent advances in chromatin conformation capture technologies, such as SPRITE and Pore-C, have enabled the detection of simultaneous contacts among multiple chromatin loci. This has made it possible to investigate the cooperative transcriptional regulation involving multiple genes and regulatory elements at the resolution of a single molecule. However, these technologies are unavoidably subject to the random polymer looping effect and technical biases, making it challenging to distinguish genuine regulatory relationships directly from random polymer interactions. Here, we present HyperloopFinder, a method for identifying regulatory multi-way chromatin contacts (hyperloops) by jointly modeling the random polymer looping effect and technical biases to estimate the statistical significance of multi-way contacts. The results show that our model can accurately estimate the expected interaction frequency of multi-way contacts based on the distance distribution of pairwise contacts, revealing that most multi-way contacts can be formed by randomly linking the pairwise contacts adjacent to each other. Moreover, we observed the spatial colocalization of the interaction sites of hyperloops from image-based data. Our results also revealed that hyperloops can function as scaffolds for the cooperation among multiple genes and regulatory elements. In summary, our work contributes novel insights into higher-order chromatin structures and functions and has the potential to enhance our understanding of transcriptional regulation and other cellular processes. Weibing Wang, Yusen Ye, Lin Gao 0006 |
Briefings Bioinform. | 3 |
| 2024 | DiSMVC: a multi-view graph collaborative learning framework for measuring disease similarityabstractMOTIVATION: Exploring potential associations between diseases can help in understanding pathological mechanisms of diseases and facilitating the discovery of candidate biomarkers and drug targets, thereby promoting disease diagnosis and treatment. Some computational methods have been proposed for measuring disease similarity. However, these methods describe diseases without considering their latent multi-molecule regulation and valuable supervision signal, resulting in limited biological interpretability and efficiency to capture association patterns. RESULTS: In this study, we propose a new computational method named DiSMVC. Different from existing predictors, DiSMVC designs a supervised graph collaborative framework to measure disease similarity. Multiple bio-entity associations related to genes and miRNAs are integrated via cross-view graph contrastive learning to extract informative disease representation, and then association pattern joint learning is implemented to compute disease similarity by incorporating phenotype-annotated disease associations. The experimental results show that DiSMVC can draw discriminative characteristics for disease pairs, and outperform other state-of-the-art methods. As a result, DiSMVC is a promising method for predicting disease associations with molecular interpretability. AVAILABILITY AND IMPLEMENTATION: Datasets and source codes are available at https://github.com/Biohang/DiSMVC. Hang Wei 0005, Lin Gao 0006, Shuai Wu 0001, Yina Jiang, Bin Liu 0014 |
Bioinform. | 2 |
| 2024 | Contrastive pre-training and 3D convolution neural network for RNA and small molecule binding affinity predictionabstractMOTIVATION: The diverse structures and functions inherent in RNAs present a wealth of potential drug targets. Some small molecules are anticipated to serve as leading compounds, providing guidance for the development of novel RNA-targeted therapeutics. Consequently, the determination of RNA-small molecule binding affinity is a critical undertaking in the landscape of RNA-targeted drug discovery and development. Nevertheless, to date, only one computational method for RNA-small molecule binding affinity prediction has been proposed. The prediction of RNA-small molecule binding affinity remains a significant challenge. The development of a computational model is deemed essential to effectively extract relevant features and predict RNA-small molecule binding affinity accurately. RESULTS: In this study, we introduced RLaffinity, a novel deep learning model designed for the prediction of RNA-small molecule binding affinity based on 3D structures. RLaffinity integrated information from RNA pockets and small molecules, utilizing a 3D convolutional neural network (3D-CNN) coupled with a contrastive learning-based self-supervised pre-training model. To the best of our knowledge, RLaffinity was the first deep learning based method for the prediction of RNA-small molecule binding affinity. Our experimental results exhibited RLaffinity's superior performance compared to baseline methods, revealed by all metrics. The efficacy of RLaffinity underscores the capability of 3D-CNN to accurately extract both global pocket information and local neighbor nucleotide information within RNAs. Notably, the integration of a self-supervised pre-training model significantly enhanced predictive performance. Ultimately, RLaffinity was also proved as a potential tool for RNA-targeted drugs virtual screening. AVAILABILITY AND IMPLEMENTATION: https://github.com/SaisaiSun/RLaffinity. Saisai Sun, Lin Gao 0006 |
Bioinform. | 2 |
| 2023 | Mining disease-associated genes based on heterogeneous graph transformerabstractIdentification of disease-associated genes is a crucial step in unraveling the underlying molecular mechanisms of complex diseases. As a potent computational tool for recognizing potential associations between genes and diseases, heterogeneous graphs allow for more comprehensive and accurate modeling and analysis of intricate molecular interactions within unified networks. In this study, a method based on Heterogeneous Graph Transformer (HGT) is proposed for mining associations between genes and diseases from heterogeneous graphs. Six types of association data are collected among four biological components: Protein-coding (PC) genes, long non-coding RNAs (lncRNAs), microRNAs, and diseases. A series of processing and integration steps are performed on these raw data to construct an extensive heterogeneous graph. For the vertexes in the graph, their description texts are collected and specialized natural language processing model, BioBERT, is employed to generate semantic vectors as the initial feature vectors for the vertexes. The topological structure of this heterogeneous graph along with the nodes' initial features collectively serve as input to the HGT, enabling information aggregation. Finally, the gene and disease vector representations produced by HGT are used to compute the association scores, which are further applied in the association prediction. Experimental results demonstrate that the proposed method outperforms existing state-of-the-art methods in terms of accuracy and exhibits a remarkably high degree of reliability. Moreover, the method has the capability to implicitly capture meaningful biological mechanisms through learned implicit meta-paths, which can offer substantial guidance for subsequent biomedical validations. Xingli Guo, Yao Yun, Lin Gao 0006 |
BIBM | 4 |
| 2023 | Application of Multi-Dimensional SNV Features in Supervised LearningabstractSingle nucleotide variation (SNV) is closely related to the occurrence of cancer, and effective extraction of SNV sample information will be beneficial for accurate early diagnosis of cancer. The main method of traditional research to extract SNV features is to combine SNV with two adjacent nucleotides to form a trinucleotide, and mutation features are extracted from the pattern of trinucleotides. However, single-dimensional feature extraction may lead to partial information loss and poor model performance. Therefore, we propose a method to extract single-nucleotide variation (SNV) features under multiple feature dimensions. We treat SNV as a one-dimensional feature, change the feature dimension by adding adjacent nucleotides, and achieve resampling. Simultaneously, we store multiple sets of features to ensure the integrity of the information carried by the SNV. We extend the method to the cancer marker identification application scenario to verify the improvement of the prediction performance by the extracted features. Using a dataset obtained from The Cancer Genome Atlas, based on six supervised learning algorithms, including KNN(K-Nearest Neighbors), SVM(Support Vector Machine), and random forest, we verified the feasibility of multidimensional SNV features. Compared with the original SNV feature extraction method, the feature extraction method proposed in this paper has significantly improved the prediction performance of the model. Additionally, this paper compares multi-dimensional features with the K-mer algorithm under the same dataset, and multi-dimensional features show obvious advantages in terms of acquisition time, storage space, and sample distinction Liang Yu 0002, Lin Gao 0006, Hongxia Hao |
BIBM | 2 |
| 2023 | Gene Expression Profile Prediction under Drug Action Based on Generative Adversarial NetworksabstractGene expression profiles play a significant role in drug research. If the gene expression profile under the action of drugs can be obtained quickly, such as through computational methods, the analysis of the relationship between the drug and the disease will become more comprehensive. The efficiency can be improved and costs can be reduced while exploring the effect of the drug. We developed an algorithm (ppc-GAN, predict-profile-conditional Generative Adversarial Networks) for predicting gene expression profiles for drug effects, which can efficiently and accurately obtain the gene expression profiles after drug administration. Compared with traditional algorithms, ppc-GAN does not require more prior knowledge. Therefore, the final prediction result will not be affected by the preference of prior knowledge. Our ppc-GAN mainly includes two parts—an autoencoder and a generative adversarial network (GAN). We trained the autoencoder through all gene expression profile data in the LINCS database and then merged the trained autoencoder into the GAN for data compression and decompression. Besides, we chose bortezomib as the case drug. Our results show that our model is flexible and has high representative power. Furthermore, the state of the gene expression profile after using the drug can be estimated by the deep learning models. Liang Yu 0002, Huan Zhu, Da Dong, Lin Gao 0006 |
BIBM | 4 |
| 2023 | Potent antibiotic design via guided search from antibacterial activity evaluationsabstractMOTIVATION: The emergence of drug-resistant bacteria makes the discovery of new antibiotics an urgent issue, but finding new molecules with the desired antibacterial activity is an extremely difficult task. To address this challenge, we established a framework, MDAGS (Molecular Design via Attribute-Guided Search), to optimize and generate potent antibiotic molecules. RESULTS: By designing the antibacterial activity latent space and guiding the optimization of functional compounds based on this space, the model MDAGS can generate novel compounds with desirable antibacterial activity without the need for extensive expensive and time-consuming evaluations. Compared with existing antibiotics, candidate antibacterial compounds generated by MDAGS always possessed significantly better antibacterial activity and ensured high similarity. Furthermore, although without explicit constraints on similarity to known antibiotics, these candidate antibacterial compounds all exhibited the highest structural similarity to antibiotics of expected function in the DrugBank database query. Overall, our approach provides a viable solution to the problem of bacterial drug resistance. AVAILABILITY AND IMPLEMENTATION: Code of the model and datasets can be downloaded from GitHub (https://github.com/LiangYu-Xidian/MDAGS). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Liang Yu 0002, Lin Gao 0006 |
Bioinform. | 3 |
| 2023 | HiSV: A control-free method for structural variation detection from Hi-C dataabstractStructural variations (SVs) play an essential role in the evolution of human genomes and are associated with cancer genetics and rare disease. High-throughput chromosome capture (Hi-C) technology probed all genome-wide crosslinked chromatin to study the spatial architecture of chromosomes. Hi-C read pairs can span megabases, making the technology useful for detecting large-scale SVs. So far, the identification of SVs from Hi-C data is still in the early stages with only a few methods available. Therefore, we developed HiSV (Hi-C for Structural Variation), a control-free method for identifying large-scale SVs from a Hi-C sample. Inspired by the single image saliency detection model, HiSV constructed a saliency map of interaction frequencies and extracted saliency segments as large-scale SVs. By evaluating both simulated and real data, HiSV not only detected all variant types, but also achieved a higher level of accuracy and sensitivity than most existing methods. Moreover, our results on cancer cell lines showed that HiSV effectively detected eight complex SV events and identified two novel SVs of key factors associated with cancer development. Finally, we found that integrating the result of HiSV helped the WGS method to identify a total number of 94 novel SVs in two cancer cell lines. Junping Li, Lin Gao 0006, Yusen Ye |
PLoS Comput. Biol. | 2 |
| 2022 | Predicting the functional effects of human non-coding variants based on stacking ensemble learningabstractPredicting the functional impact of genetic variants in non-coding regions of the human genome can aid in the elucidation of the etiology of diseases or traits. In recent years, an increasing number of methods to predict the impact of sequence variation in non-coding regions of the human genome have been developed. However, most of current studies are limited to predict specific types of non-coding variants. To address this problem, here we propose a non-coding SNVs prediction method based on stacking integration strategy. The method consists of three stacking models built using the same strategy based on different causality assumptions to predict functional, pathogenic, and cancer driver non-coding SNVs, respectively. We demonstrate that our method outperforms the other seven methods. In addition, a comparison of our proposed model with other methods for non-coding de novo mutations in autism spectrum disease reveals that our model has the highest discriminative ability, indicating that its performance is stable and superior in different scenarios. Kei Hang Katie Chan, Lin Gao 0006 |
BIBM | 4 |
| 2022 | IMRDriver: coding and non-coding cancer driver genes identification based on network propagationabstractIn cancer genomics, the identification of Cancer Driver Genes (CDGs) is a major scientific interest. CDGs can be identified by numerous methods, however the false positive rate still remains high. In addition, non-coding genes, such as miRNAs, can also operate as CDGs due to their regulatory functions in the development of cancer. In this paper, we present IMRDriver, a novel method for identifying both protein-coding and non-coding CDGs based on network propagation. The method first employs gene expression data, copy number variation data, single nucleotide variation data, and gene interaction data to construct a node-weighted gene network. Then, the network topology is combined with the reverse network propagation to rank all genes, with the top ranked genes predicted to be CDG candidates. We compared the prediction results of IMRDriver to twelve other methods and found that IMRDriver outperforms in terms of accuracy, recall, and F1 score. In addition, IMRDriver identified a number of miRNAs as non-coding CDGs, the majority of which have been verified in the scientific literature. In summary, IMRDriver is an effective approach for predicting CDGs. Source code of our paper is available at https://github.com/cczxsong/IMRDriver. Kei Hang Katie Chan, Lin Gao 0006 |
BIBM | 4 |
| 2022 | Reconstruction of human protein-coding gene functional association network based on machine learningabstractNetworks consisting of molecular interactions are intrinsically dynamical systems of an organism. These interactions curated in molecular interaction databases are still not complete and contain false positives introduced by high-throughput screening experiments. In this study, we propose a framework to integrate interactions of functional associated protein-coding genes from 31 data sources to reconstruct a network with high coverage and quality. For each interaction, 369 features were constructed including properties of both the interaction and the involved genes. The training and validation sets were built on the pathway interactions as positives and the potential negative instances resulting from our proposed semi-supervised strategy. Random forest classification method was then applied to train and predict multiple times to give a score for each interaction. After setting a threshold estimated by a Binomial distribution, a Human protein-coding Gene Functional Association Network (HuGFAN) was reconstructed with 20 383 genes and 1185 429 high confidence interactions. Then, HuGFAN was compared with other networks from data sources with respect to network properties, suggesting that HuGFAN is more function and pathway related. Finally, HuGFAN was applied to identify cancer driver through two famous network-based methods (DriverNet and HotNet2) to show its outstanding performance compared with other networks. HuGFAN and other supplementary files are freely available at https://github.com/xthuang226/HuGFAN. Songwei Jia, Lin Gao 0006 |
Briefings Bioinform. | 3 |
| 2022 | MiRNA-disease association prediction based on meta-pathsabstractSince miRNAs can participate in the posttranscriptional regulation of gene expression, they may provide ideas for the development of new drugs or become new biomarkers for drug targets or disease diagnosis. In this work, we propose an miRNA-disease association prediction method based on meta-paths (MDPBMP). First, an miRNA-disease-gene heterogeneous information network was constructed, and seven symmetrical meta-paths were defined according to different semantics. After constructing the initial feature vector for the node, the vector information carried by all nodes on the meta-path instance is extracted and aggregated to update the feature vector of the starting node. Then, the vector information obtained by the nodes on different meta-paths is aggregated. Finally, miRNA and disease embedding feature vectors are used to calculate their associated scores. Compared with the other methods, MDPBMP obtained the highest AUC value of 0.9214. Among the top 50 predicted miRNAs for lung neoplasms, esophageal neoplasms, colon neoplasms and breast neoplasms, 49, 48, 49 and 50 have been verified. Furthermore, for breast neoplasms, we deleted all the known associations between breast neoplasms and miRNAs from the training set. These results also show that for new diseases without known related miRNA information, our model can predict their potential miRNAs. Code and data are available at https://github.com/LiangYu-Xidian/MDPBMP. Liang Yu 0002, Lin Gao 0006 |
Briefings Bioinform. | 3 |
| 2022 | Research progress of miRNA-disease association prediction and comparison of related algorithmsabstractWith an in-depth understanding of noncoding ribonucleic acid (RNA), many studies have shown that microRNA (miRNA) plays an important role in human diseases. Because traditional biological experiments are time-consuming and laborious, new calculation methods have recently been developed to predict associations between miRNA and diseases. In this review, we collected various miRNA-disease association prediction models proposed in recent years and used two common data sets to evaluate the performance of the prediction models. First, we systematically summarized the commonly used databases and similarity data for predicting miRNA-disease associations, and then divided the various calculation models into four categories for summary and detailed introduction. In this study, two independent datasets (D5430 and D6088) were compiled to systematically evaluate 11 publicly available prediction tools for miRNA-disease associations. The experimental results indicate that the methods based on information dissemination and the method based on scoring function require shorter running time. The method based on matrix transformation often requires a longer running time, but the overall prediction result is better than the previous two methods. We hope that the summary of work related to miRNA and disease will provide comprehensive knowledge for predicting the relationship between miRNA and disease and contribute to advanced computation tools in the future. Liang Yu 0002, Bingyi Ju, Chunyan Ao, Lin Gao 0006 |
Briefings Bioinform. | 5 |
| 2022 | Multidrug representation learning based on pretraining model and molecular graph for drug interaction and combination predictionabstractMOTIVATION: Approaches for the diagnosis and treatment of diseases often adopt the multidrug therapy method because it can increase the efficacy or reduce the toxic side effects of drugs. Using different drugs simultaneously may trigger unexpected pharmacological effects. Therefore, efficient identification of drug interactions is essential for the treatment of complex diseases. Currently proposed calculation methods are often limited by the collection of redundant drug features, a small amount of labeled data and low model generalization capabilities. Meanwhile, there is also a lack of unique methods for multidrug representation learning, which makes it more difficult to take full advantage of the originally scarce data. RESULTS: Inspired by graph models and pretraining models, we integrated a large amount of unlabeled drug molecular graph information and target information, then designed a pretraining framework, MGP-DR (Molecular Graph Pretraining for Drug Representation), specifically for drug pair representation learning. The model uses self-supervised learning strategies to mine the contextual information within and between drug molecules to predict drug-drug interactions and drug combinations. The results achieved promising performance across multiple metrics compared with other state-of-the-art methods. Our MGP-DR model can be used to provide a reliable candidate set for the combined use of multiple drugs. AVAILABILITY AND IMPLEMENTATION: Code of the model, datasets and results can be downloaded from GitHub (https://github.com/LiangYu-Xidian/MGP-DR). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Shujie Ren, Liang Yu 0002, Lin Gao 0006 |
Bioinform. | 3 |
| 2021 | The peripheral and core regions of virus-host network of COVID-19abstractTwo thousand nineteen novel coronavirus SARS-CoV-2, the pathogen of COVID-19, has caused a catastrophic pandemic, which has a profound and widespread impact on human lives and social economy globally. However, the molecular perturbations induced by the SARS-CoV-2 infection remain unknown. In this paper, from the perspective of omnigenic, we analyze the properties of the neighborhood perturbed by SARS-CoV-2 in the human interactome and disclose the peripheral and core regions of virus-host network (VHN). We find that the virus-host proteins (VHPs) form a significantly connected VHN, among which highly perturbed proteins aggregate into an observable core region. The non-core region of VHN forms a large scale but relatively low perturbed periphery. We further validate that the periphery is non-negligible and conducive to identifying comorbidities and detecting drug repurposing candidates for COVID-19. We particularly put forward a flower model for COVID-19, SARS and H1N1 based on their peripheral regions, and the flower model shows more correlations between COVID-19 and other two similar diseases in common functional pathways and candidate drugs. Overall, our periphery-core pattern can not only offer insights into interconnectivity of SARS-CoV-2 VHPs but also facilitate the research on therapeutic drugs. Bingbo Wang, Xianan Dong, Lin Gao 0006 |
Briefings Bioinform. | 7 |
| 2021 | Improving Single-Cell RNA-seq Clustering by Integrating PathwaysabstractSingle-cell clustering is an important part of analyzing single-cell RNA-sequencing data. However, the accuracy and robustness of existing methods are disturbed by noise. One promising approach for addressing this challenge is integrating pathway information, which can alleviate noise and improve performance. In this work, we studied the impact on accuracy and robustness of existing single-cell clustering methods by integrating pathways. We collected 10 state-of-the-art single-cell clustering methods, 26 scRNA-seq datasets and four pathway databases, combined the AUCell method and the similarity network fusion to integrate pathway data and scRNA-seq data, and introduced three accuracy indicators, three noise generation strategies and robustness indicators. Experiments on this framework showed that integrating pathways can significantly improve the accuracy and robustness of most single-cell clustering methods. Chenxing Zhang, Lin Gao 0006, Bingbo Wang, Yong Gao 0001 |
Briefings Bioinform. | 2 |
| 2021 | CCIP: predicting CTCF-mediated chromatin loops with transitivityabstractMOTIVATION: CTCF-mediated chromatin loops underlie the formation of topological associating domains and serve as the structural basis for transcriptional regulation. However, the formation mechanism of these loops remains unclear, and the genome-wide mapping of these loops is costly and difficult. Motivated by the recent studies on the formation mechanism of CTCF-mediated loops, we studied the possibility of making use of transitivity-related information of interacting CTCF anchors to predict CTCF loops computationally. In this context, transitivity arises when two CTCF anchors interact with the same third anchor by the loop extrusion mechanism and bring themselves close to each other spatially to form an indirect loop. RESULTS: To determine whether transitivity is informative for predicting CTCF loops and to obtain an accurate and low-cost predicting method, we proposed a two-stage random-forest-based machine learning method, CTCF-mediated Chromatin Interaction Prediction (CCIP), to predict CTCF-mediated chromatin loops. Our two-stage learning approach makes it possible for us to train a prediction model by taking advantage of transitivity-related information as well as functional genome data and genomic data. Experimental studies showed that our method predicts CTCF-mediated loops more accurately than other methods and that transitivity, when used as a properly defined attribute, is informative for predicting CTCF loops. Furthermore, we found that transitivity explains the formation of tandem CTCF loops and facilitates enhancer-promoter interactions. Our work contributes to the understanding of the formation mechanism and function of CTCF-mediated chromatin loops. AVAILABILITY AND IMPLEMENTATION: The source code of CCIP can be accessed at: https://github.com/GaoLabXDU/CCIP. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Weibing Wang, Lin Gao 0006, Yusen Ye, Yong Gao 0001 |
Bioinform. | 2 |
| 2021 | CeNet Omnibus: an R/Shiny application to the construction and analysis of competing endogenous RNA networkabstractBACKGROUND: The competing endogenous RNA (ceRNA) regulation is a newly discovered post-transcriptional regulation mechanism and plays significant roles in physiological and pathological progress. CeRNA networks provide global views to help understand the regulation of ceRNAs. CeRNA networks have been widely used to detect survival biomarkers, select candidate regulators of disease genes, and predict long noncoding RNA functions. However, there is no software platform to provide overall functions from the construction to analysis of ceRNA networks. RESULTS: To fill this gap, we introduce CeNet Omnibus, an R/Shiny application, which provides a unified framework for the construction and analysis of ceRNA network. CeNet Omnibus enables users to select multiple measurements, such as Pearson correlation coefficient (PCC), mutual information (MI), and liquid association (LA), to identify ceRNA pairs and construct ceRNA networks. Furthermore, CeNet Omnibus provides a one-stop solution to analyze the topological properties of ceRNA networks, detect modules, and perform gene enrichment analysis and survival analysis. CeNet Omnibus intends to cover comprehensiveness, high efficiency, high expandability, and user customizability, and it also offers a web-based user-friendly interface to users to obtain the output intuitionally. CONCLUSION: CeNet Omnibus is a comprehensive platform for the construction and analysis of ceRNA networks. It is highly customizable and outputs the results in intuitive and interactive. We expect that CeNet Omnibus will assist researchers to understand the property of ceRNA networks and associated biological phenomena. CeNet Omnibus is an R/Shiny application based on the Shiny framework developed by RStudio. The R package and detailed tutorial are available on our GitHub page with the URL https://github.com/GaoLabXDU/CeNetOmnibus . Lin Gao 0006, Tuo Song, Chaoqun Jiang |
BMC Bioinform. | 2 |
| 2021 | Prediction of drug-target interactions based on multi-layer network representation learning
Yifan Shang, Lin Gao 0006, Quan Zou 0001, Liang Yu 0002 |
Neurocomputing | 2 |
| 2021 | Evaluation and comparison of multi-omics data integration methods for cancer subtypingabstractComputational integrative analysis has become a significant approach in the data-driven exploration of biological problems. Many integration methods for cancer subtyping have been proposed, but evaluating these methods has become a complicated problem due to the lack of gold standards. Moreover, questions of practical importance remain to be addressed regarding the impact of selecting appropriate data types and combinations on the performance of integrative studies. Here, we constructed three classes of benchmarking datasets of nine cancers in TCGA by considering all the eleven combinations of four multi-omics data types. Using these datasets, we conducted a comprehensive evaluation of ten representative integration methods for cancer subtyping in terms of accuracy measured by combining both clustering accuracy and clinical significance, robustness, and computational efficiency. We subsequently investigated the influence of different omics data on cancer subtyping and the effectiveness of their combinations. Refuting the widely held intuition that incorporating more types of omics data always produces better results, our analyses showed that there are situations where integrating more omics data negatively impacts the performance of integration methods. Our analyses also suggested several effective combinations for most cancers under our studies, which may be of particular interest to researchers in omics data analysis. Ran Duan 0001, Lin Gao 0006, Yong Gao 0001, Yuxuan Hu 0004, Mingfeng Huang, Kuo Song, Hongda Wang, Yongqiang Dong, Chaoqun Jiang, Chenxing Zhang, Songwei Jia |
PLoS Comput. Biol. | 2 |
| 2021 | Predicting therapeutic drugs for hepatocellular carcinoma based on tissue-specific pathwaysabstractHepatocellular carcinoma (HCC) is a significant health problem worldwide with poor prognosis. Drug repositioning represents a profitable strategy to accelerate drug discovery in the treatment of HCC. In this study, we developed a new approach for predicting therapeutic drugs for HCC based on tissue-specific pathways and identified three newly predicted drugs that are likely to be therapeutic drugs for the treatment of HCC. We validated these predicted drugs by analyzing their overlapping drug indications reported in PubMed literature. By using the cancer cell line data in the database, we constructed a Connectivity Map (CMap) profile similarity analysis and KEGG enrichment analysis on their related genes. By experimental validation, we found securinine and ajmaline significantly inhibited cell viability of HCC cells and induced apoptosis. Among them, securinine has lower toxicity to normal liver cell line, which is worthy of further research. Our results suggested that the proposed approach was effective and accurate for discovering novel therapeutic options for HCC. This method also could be used to indicate unmarked drug-disease associations in the Comparative Toxicogenomics Database. Meanwhile, our method could also be applied to predict the potential drugs for other types of tumors by changing the database. Liang Yu 0002, Fengdan Xu, Lin Gao 0006, Xiangzhi Li |
PLoS Comput. Biol. | 7 |
| 2020 | C3: connect separate connected components to form a succinct disease moduleabstractBACKGROUND: Precise disease module is conducive to understanding the molecular mechanism of disease causation and identifying drug targets. However, due to the fragmentization of disease module in incomplete human interactome, how to determine connectivity pattern and detect a complete neighbourhood of disease based on this is still an open question. RESULTS: In this paper, we perform exploratory analysis leading to an important observation that through a few intermediate nodes, most separate connected components formed by disease-associated proteins can be effectively connected and eventually form a complete disease module. And based on the topological properties of these intermediate nodes, we propose a connect separate connected components (C3) method to detect a succinct disease module by introducing a relatively small number of intermediate nodes, which allows us to obtain more pure disease module than other methods. Then we apply C3 across a large corpus of diseases to validate this connectivity pattern of disease module. Furthermore, the connectivity of the perturbed genes in multi-omics data such as The Cancer Genome Atlas also fits this pattern. CONCLUSIONS: C3 tool is not only useful in detecting a clearly-defined connected disease neighbourhood of 299 diseases and cancer with multi-omics data, but also helpful in better understanding the interconnection of phenotypically related genes in different omics data and studying complex pathological processes. Bingbo Wang, Chenxing Zhang, Yuanjun Zhou, Liang Yu 0002, Xingli Guo, Lin Gao 0006, Yunru Chen |
BMC Bioinform. | 8 |
| 2020 | Cluster correlation based method for lncRNA-disease association predictionabstractBACKGROUND: In recent years, increasing evidences have indicated that long non-coding RNAs (lncRNAs) are deeply involved in a wide range of human biological pathways. The mutations and disorders of lncRNAs are closely associated with many human diseases. Therefore, it is of great importance to predict potential associations between lncRNAs and complex diseases for the diagnosis and cure of complex diseases. However, the functional mechanisms of the majority of lncRNAs are still remain unclear. As a result, it remains a great challenge to predict potential associations between lncRNAs and diseases. RESULTS: Here, we proposed a new method to predict potential lncRNA-disease associations. First, we constructed a bipartite network based on known associations between diseases and lncRNAs/protein coding genes. Then the cluster association scores were calculated to evaluate the strength of the inner relationships between disease clusters and gene clusters. Finally, the gene-disease association scores are defined based on disease-gene cluster association scores and used to measure the strength for potential gene-disease associations. CONCLUSIONS: Leave-One Out Cross Validation (LOOCV) and 5-fold cross validation tests were implemented to evaluate the performance of our method. As a result, our method achieved reliable performance in the LOOCV (AUCs of 0.8169 and 0.8410 based on Yang's dataset and Lnc2cancer 2.0 database, respectively), and 5-fold cross validation (AUCs of 0.7573 and 0.8198 based on Yang's dataset and Lnc2cancer 2.0 database, respectively), which were significantly higher than the other three comparative methods. Furthermore, our method is simple and efficient. Only the known gene-disease associations are exploited in a graph manner and further new gene-disease associations can be easily incorporated in our model. The results for melanoma and ovarian cancer have been verified by other researches. The case studies indicated that our method can provide informative clues for further investigation. Qianqian Yuan, Xingli Guo, Lin Gao 0006 |
BMC Bioinform. | 5 |
| 2020 | Detection of Driver Modules with Rarely Mutated Genes in CancersabstractIdentifying driver modules or pathways is a key challenge to interpret the molecular mechanisms and pathogenesis underlying cancer. An increasing number of studies suggest that rarely mutated genes are important for the development of cancer. However, the driver modules consisting of mutated genes with low-frequency driver mutations are not well characterized. To identify driver modules with rarely mutated genes, we propose a functional similarity index to quantify the functional relationship between rarely mutated genes and other ones in the same module. Then, we develop a method to detect Driver Modules with Rarely mutated Genes (DMRG) by incorporating the functional similarity, coverage and mutual exclusivity. By applying DMRG on TCGA cancer dataset on three networks: HINT+HI2012, iRefIndex and MultiNet, we detect driver modules intersecting with the well-known signalling pathways and protein complexes, such as the cell cycle pathway and the mediator complex. DMRG can also detect driver modules effectively with 20, 40, 60 and 80 percent of samples by random selection. When compared with HotNet2, DMRG detects more rarely mutated cancer genes and has higher pathway enrichment. Overall, DMRG provides an effective method for the identification of driver modules with rarely mutated genes. Feng Li 0033, Lin Gao 0006, Bingbo Wang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2019 | miES: predicting the essentiality of miRNAs with machine learning and sequence featuresabstractMOTIVATION: MicroRNAs (miRNAs) are one class of small noncoding RNA molecules, which regulate gene expression at the post-transcriptional level and play important roles in health and disease. To dissect the critical miRNAs in miRNAome, it is needed to predict the essentiality of miRNAs, however, bioinformatics methods for this purpose are limited. RESULTS: Here we propose miES, a novel algorithm, for the prioritization of miRNA essentiality. miES implements a machine learning strategy based on learning from positive and unlabeled samples. miES uses sequence features of known essential miRNAs and performs miRNAome-wide searching for new essential miRNAs. miES achieves an AUC of 0.9 for 5-fold cross validation. Moreover, experiments further show that the miES score is significantly correlated with some established biological metrics for miRNA importance, such as miRNA conservation, miRNA disease spectrum width (DSW) and expression level. AVAILABILITY AND IMPLEMENTATION: The R source code is available at the download page of the web server, http://www.cuilab.cn/mies. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Chunmei Cui, Lin Gao 0006, Qinghua Cui |
Bioinform. | 3 |
| 2019 | Human Pathway-Based Disease NetworkabstractConstructing disease-disease similarity network is important in elucidating the associations between the origin and molecular mechanism of diseases, and in researching disease function and medical research. In this paper, we use a high-quality protein interaction network and a collection of pathway databases to construct a Human Pathway-based Disease Network (HPDN) to explore the relationship between diseases and their intrinsic interactions. We find that the similarity of two diseases has a strong correlation with the number of their shared functional pathways and the interaction between their related gene sets. Comparing HPDN with disease networks based on genes and symptoms respectively, we find the three networks have high overlap rates. Additionally, HPDN can predict new disease-disease correlations, which are supported by Comparative Toxicogenomics Database (CTD) benchmark and large-scale biomedical literature database. The comprehensive, high-quality relations between diseases based on pathways can further be applied to study important matters in systems medicine, for instance, drug repurposing. Based on a dense subgraph in our network, we find two drugs, prednisone and folic acid, may have new indications, which will provide potential directions for the treatments of complex diseases. Liang Yu 0002, Lin Gao 0006 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2018 | Inferring Dysregulated Pathways of Driving Cancer Subtypes Through Multi-omics Integration
Kai Shi 0004, Lin Gao 0006, Bingbo Wang |
ISBRA | 2 |
| 2018 | Feature related multi-view nonnegative matrix factorization for identifying conserved functional modules in multiple biological networksabstractBACKGROUND: Comprehensive analyzing multi-omics biological data in different conditions is important for understanding biological mechanism in system level. Multiple or multi-layer network model gives us a new insight into simultaneously analyzing these data, for instance, to identify conserved functional modules in multiple biological networks. However, because of the larger scale and more complicated structure of multiple networks than single network, how to accurate and efficient detect conserved functional biological modules remains a significant challenge. RESULTS: Here, we propose an efficient method, named ConMod, to discover conserved functional modules in multiple biological networks. We introduce two features to characterize multiple networks, thus all networks are compressed into two feature matrices. The module detection is only performed in the feature matrices by using multi-view non-negative matrix factorization (NMF), which is independent of the number of input networks. Experimental results on both synthetic and real biological networks demonstrate that our method is promising in identifying conserved modules in multiple networks since it improves the accuracy and efficiency comparing with state-of-the-art methods. Furthermore, applying ConMod to co-expression networks of different cancers, we find cancer shared gene modules, the majority of which have significantly functional implications, such as ribosome biogenesis and immune response. In addition, analyzing on brain tissue-specific protein interaction networks, we detect conserved modules related to nervous system development, mRNA processing, etc. CONCLUSIONS: ConMod facilitates finding conserved modules in any number of networks with a low time and space complexity, thereby serve as a valuable tool for inference shared traits and biological functions of multiple biological system. Peizhuo Wang, Lin Gao 0006, Yuxuan Hu 0004, Feng Li 0033 |
BMC Bioinform. | 2 |
| 2018 | Extracting Stage-Specific and Dynamic Modules Through Analyzing Multiple Networks Associated with Cancer ProgressionabstractDetermining the dynamics of pathways associated with cancer progression is critical for understanding the etiology of diseases. Advances in biological technology have facilitated the simultaneous genomic profiling of multiple patients at different clinical stages, thus generating the dynamic genomic data for cancers. Such data provide enable investigation of the dynamics of related pathways. However, methods for integrative analysis of dynamic genomic data are inadequate. In this study, we develop a novel nonnegative matrix factorization algorithm for dynamic modules ( NMF-DM), which simultaneously analyzes multiple networks for the identification of stage-specific and dynamic modules. NMF-DM applies the temporal smoothness framework by balancing the networks at the current stage and the previous stage. Experimental results indicate that the NMF-DM algorithm is more accurate than the state-of-the-art methods in artificial dynamic networks. In breast cancer networks, NMF-DM reveals the dynamic modules that are important for cancer stage transitions. Furthermore, the stage-specific and dynamic modules have distinct topological and biochemical properties. Finally, we demonstrate that the stage-specific modules significantly improve the accuracy of cancer stage prediction. The proposed algorithm provides an effective way to explore the time-dependent cancer genomic data. Xiaoke Ma 0001, Wanxin Tang, Peizhuo Wang, Xingli Guo, Lin Gao 0006 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2017 | Drug repositioning based on triangularly balanced structure for tissue-specific diseases in incomplete interactome
Liang Yu 0002, Lin Gao 0006 |
Artif. Intell. Medicine | 3 |
| 2017 | Prediction of Novel Drugs for Hepatocellular Carcinoma Based on Multi-Source Random WalkabstractComputational approaches for predicting drug-disease associations by integrating gene expression and biological network provide great insights to the complex relationships among drugs, targets, disease genes, and diseases at a system level. Hepatocellular carcinoma (HCC) is one of the most common malignant tumors with a high rate of morbidity and mortality. We provide an integrative framework to predict novel d rugs for HCC based on multi-source random walk (PD-MRW). Firstly, based on gene expression and protein interaction network, we construct a gene-gene weighted i nteraction network (GWIN). Then, based on multi-source random walk in GWIN, we build a drug-drug similarity network. Finally, based on the known drugs for HCC, we score all drugs in the drug-drug similarity network. The robustness of our predictions, their overlap with those reported in Comparative Toxicogenomics Database (CTD) and literatures, and their enriched KEGG pathway demonstrate our approach can effectively identify new drug indications. Specifically, regorafenib (Rank = 9 in top-20 list) is proven to be effective in Phase I and II clinical trials of HCC, and the Phase III trial is ongoing. And, it has 11 overlapping pathways with HCC with lower p-values. Focusing on a particular disease, we believe our approach is more accurate and possesses better scalability. Liang Yu 0002, Ruidan Su, Bingbo Wang, Yapeng Zou, Lin Gao 0006 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 7 |
| 2016 | Integrating phenotypic features and tissue-specific information to prioritize disease genes
Yue Deng 0006, Lin Gao 0006, Xingli Guo, Bingbo Wang |
Sci. China Inf. Sci. | 2 |
| 2016 | Systematic tracking of coordinated differential network motifs identifies novel disease-related genes by integrating multiple data
Kai Shi 0004, Lin Gao 0006, Bingbo Wang |
Neurocomputing | 2 |
| 2016 | Network-Based Method for Inferring Cancer Progression at the Pathway Level from Cross-Sectional Mutation DataabstractLarge-scale cancer genomics projects are providing a wealth of somatic mutation data from a large number of cancer patients. However, it is difficult to obtain several samples with a temporal order from one patient in evaluating the cancer progression. Therefore, one of the most challenging problems arising from the data is to infer the temporal order of mutations across many patients. To solve the problem efficiently, we present a Network-based method (NetInf) to Infer cancer progression at the pathway level from cross-sectional data across many patients, leveraging on the exclusive property of driver mutations within a pathway and the property of linear progression between pathways. To assess the robustness of NetInf, we apply it on simulated data with the addition of different levels of noise. To verify the performance of NetInf, we apply it to analyze somatic mutation data from three real cancer studies with large number of samples. Experimental results reveal that the pathways detected by NetInf show significant enrichment. Our method reduces computational complexity by constructing gene networks without assigning the number of pathways, which also provides new insights on the temporal order of somatic mutations at the pathway level rather than at the gene level. Hao Wu 0062, Lin Gao 0006, Nikola K. Kasabov |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2015 | Identifying overlapping mutated driver pathways by constructing gene networks in cancerabstractBACKGROUND: Large-scale cancer genomic projects are providing lots of data on genomic, epigenomic and gene expression aberrations in many cancer types. One key challenge is to detect functional driver pathways and to filter out nonfunctional passenger genes in cancer genomics. Vandin et al. introduced the Maximum Weight Sub-matrix Problem to find driver pathways and showed that it is an NP-hard problem. METHODS: To find a better solution and solve the problem more efficiently, we present a network-based method (NBM) to detect overlapping driver pathways automatically. This algorithm can directly find driver pathways or gene sets de novo from somatic mutation data utilizing two combinatorial properties, high coverage and high exclusivity, without any prior information. We firstly construct gene networks based on the approximate exclusivity between each pair of genes using somatic mutation data from many cancer patients. Secondly, we present a new greedy strategy to add or remove genes for obtaining overlapping gene sets with driver mutations according to the properties of high exclusivity and high coverage. RESULTS: To assess the efficiency of the proposed NBM, we apply the method on simulated data and compare results obtained from the NBM, RME, Dendrix and Multi-Dendrix. NBM obtains optimal results in less than nine seconds on a conventional computer and the time complexity is much less than the three other methods. To further verify the performance of NBM, we apply the method to analyze somatic mutation data from five real biological data sets such as the mutation profiles of 90 glioblastoma tumor samples and 163 lung carcinoma samples. NBM detects groups of genes which overlap with known pathways, including P53, RB and RTK/RAS/PI(3)K signaling pathways. New gene sets with p-value less than 1e-3 are found from the somatic mutation data. CONCLUSIONS: NBM can detect more biologically relevant gene sets. Results show that NBM outperforms other algorithms for detecting driver pathways or gene sets. Further research will be conducted with the use of novel machine learning techniques. Hao Wu 0062, Lin Gao 0006, Feng Li 0033, Xiaofei Yang 0003, Nikola K. Kasabov |
BMC Bioinform. | 2 |
| 2015 | Identifying module biomarker in type 2 diabetes mellitus by discriminative area of functional activityabstractBACKGROUND: Identifying diagnosis and prognosis biomarkers from expression profiling data is of great significance for achieving personalized medicine and designing therapeutic strategy in complex diseases. However, the reproducibility of identified biomarkers across tissues and experiments is still a challenge for this issue. RESULTS: We propose a strategy based on discriminative area of module activities to identify gene biomarkers which interconnect as a subnetwork or module by integrating gene expression data and protein-protein interactions. Then, we implement the procedure in T2DM as a case study and identify a module biomarker with 32 genes from mRNA expression data in skeletal muscle for T2DM. This module biomarker is enriched with known causal genes and related functions of T2DM. Further analysis shows that the module biomarker is of superior performance in classification, and has consistently high accuracies across tissues and experiments. CONCLUSION: The proposed approach can efficiently identify robust and functionally meaningful module biomarkers in T2DM, and could be employed in biomarker discovery of other complex diseases characterized by expression profiles. Lin Gao 0006, Zhi-Ping Liu, Luonan Chen |
BMC Bioinform. | 2 |
| 2013 | Maximizing modularity intensity for community partition and evolution
Peng-Gang Sun, Lin Gao 0006 |
Inf. Sci. | 2 |
| 2012 | Predicting protein complexes in protein interaction networks using a core-attachment algorithm based on graph communicability
Xiaoke Ma 0001, Lin Gao 0006 |
Inf. Sci. | 2 |
| 2011 | Global Network Alignment Based on Multiple Hub SeedsabstractWe present a heuristic global network alignment algorithm, called Alignment based Multiple Hubs(AMH). AMH is efficient in resolving the fatal problem of most conventional algorithms that the initialization selected seeds have a direct influence on the alignment result. This algorithm outperforms state-of-the-art algorithms at detecting conserved functional modules and retrieves in particular 86% more conserved interactions than IsoRank. Bingbo Wang, Lin Gao 0006 |
BIBM | 2 |
| 2010 | A centrality measure based on spectral optimization of modularity density
Lidong Fu, Lin Gao 0006, Xiaoke Ma 0001 |
Sci. China Inf. Sci. | 2 |
| 2009 | Evaluation of subgraph searching algorithms detecting network motif in biological networks
Jialu Hu, Lin Gao 0006, Guimin Qin |
Frontiers Comput. Sci. China | 2 |
| 2005 | A DNA Based Evolutionary Algorithm for the Minimal Set Cover Problem
Xiangou Zhu, Guandong Xu, Qiang Zhang 0008, Lin Gao 0006 |
ICIC (2) | 5 |