Lei Wang 0121

dblp:181/2817-121 · DBLP profile ↗
← Back
66ranked-venue papers
16as first author
49since 2021 · last 2026
0000-0003-0184-307XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 56 · 13 first-author · 40 since 2021Artificial intelligence and machine learning · 9 · 3 first-author · 8 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Dual-Channel Learning Framework for Zero-Shot CircRNA-miRNA Interaction Prediction via State Space Modeling
abstract
CircRNA-miRNA interaction (CMI) plays a pivotal role in disease therapeutics and drug discovery. However, existing methods face several challenges in modeling complex biological networks and zero-shot learning scenarios. Biological networks encapsulate rich biological information, yet current approaches often fail to fully exploit this depth. Moreover, zero-shot prediction requires models to identify new interactions without relying on previously observed samples, imposing stringent requirements on generalization capabilities. To address these limitations, we propose a dual-channel learning framework leveraging State space modeling for Zero-shot CMI prediction (ZeroStem). ZeroStem first enhances the biological relevance of node using prior knowledge, and employs a graph Transformer to extract macro-topological representations. Subsequently, it generates semantic subgraphs based on meta-paths to focus on specific biological relationships, utilizing the Mamba to extract micro-semantic representations via state space modeling. Finally, macro-topological and micro-semantic representations are seamlessly integrated through linear transformation and residual connections, enabling high-precision zero-shot CMI prediction. Extensive experiments on multiple benchmark datasets demonstrate that ZeroStem significantly outperforms existing methods, validating its efficiency and robust generalization in CMI prediction. Case studies further illustrate that ZeroStem offers novel insights into the molecular mechanisms underlying intricate disease-associated networks.
Mengmeng Wei, Lei Wang 0121, Zhu-Hong You, Pengwei Hu 0001, Bo-Wei Zhao, Zhi-an Huang
AAAI2
2026 PEGNet-CDA: A Propagation-Enhanced Graph Network for CircRNA-Disease Association Prediction
Yue-Chao Li, Yao-Lu Li, Chen-Yv Yang, Mengmeng Wei, Xinfei Wang 0001, Lei Wang 0121, Zhi-an Huang, Zhu-Hong You
ICIC (27)6
2026 A Hybrid Transformer-GCN Framework for CircRNA-Disease Association Prediction
Mengmeng Wei, Ziyuan Shen, Ziqi Xia, Lei Wang 0121, Yong Zhou 0003
ICIC (27)5
2026 A Representation Learning Framework for CircRNA-Disease Association Prediction Based on Symmetric Convolutional Networks
Meineng Wang, Hanjing Jiang, Zhi-Xiao Wang, Lei Wang 0121
ICIC (30)4
2026 Spatial-spectral fusion enables drug repositioning by capturing indirect and long-range associations in biological networks
Lei Wang 0121, Runzhou Tang, Zhi-an Huang, Feng Tan 0002, Lun Hu, Zhu-Hong You, Pengwei Hu 0001
Bioinform.2
2026 Multi-hop graph structural modeling for cancer-related circRNA-miRNA interaction prediction
Mengmeng Wei, Lei Wang 0121, Xiao-Rui Su 0001, Bo-Wei Zhao, Zhu-Hong You
Pattern Recognit.2
2026 MuGNet-CMI: Multi-Head Hybrid Graph Neural Network for Predicting circRNA-miRNA Interactions With Global High-Order and Local Low-Order Information
abstract
Circular RNAs (circRNAs) are non-coding RNA molecules that play a crucial role in regulating genes and contributing to disease progression. CircRNAs can function as sponges for microRNAs (miRNAs), thereby regulating gene expression and influencing disease outcomes. Identifying associations between circRNAs and miRNAs through computational methods enhances the understanding of complex disease mechanisms and offers a reliable tool for pre-selecting candidates for experimental validation. Existing models, however, are limited in their ability to capture either global or local node information, the prediction of circRNA and miRNA interactions is still challenging. In order to effectively deal with this problem, we propose a novel framework for predicting circRNA-miRNA interactions (CMIs), known as MuGNet-CMI, which leverages multi-head hybrid graph neural network and global high-order and local low-order information. The model employs the MetaPath2Vec algorithm to generate high-quality node embeddings within the circRNA-miRNA heterogeneous matrix. The multi-head dynamic attention mechanism, combined with GraphSAGE, is incorporated to efficiently capture both global high-order and local low-order node information. Additionally, we integrate neural aggregators into the multi-head dynamic attention mechanism to aggregate feature information from the captured nodes. Validation using three real datasets demonstrates that MuGNet-CMI delivers good performance in predicting CMIs, offering valuable insights to guide experimental research in gene regulation.
Lei Wang 0121, Zhu-Hong You, Xinfei Wang 0001, Mengmeng Wei, Mianshuo Lu
IEEE Trans. Big Data2
2025 A Novel Sparse-Aware Topology Reconstruction and Global Dependency Enhanced Method for Predicting Human Microbe-Disease Associations
abstract
The study of human microbe-disease associations (MDAs) contributes to early diagnosis, personalized treatment, and novel drug and biomarker discovery. However, experimental verification is time-consuming and labor-intensive, underscoring the need for efficient computational prediction methods. However, existing methods have some limitations in dealing with data sparsity and effectively modeling global dependencies. To address these issues, we propose a sparseaware topology reconstruction and global dependencyenhanced method (STAGE) for MDA prediction. STAGE firstly employs an encoder-decoder structure combining Graph Attention Networks (GAT) and Graph Convolutional Networks (GCN). The GAT encoder captures key features from sparse networks, while the GCN decoder reconstructs potential associations to supplement missing information. An adaptive gating mechanism dynamically fuses original and reconstructed information to strengthen representation learning. Furthermore, an improved domain transformer, DAFormer, integrates relative position encoding, biased multi-head attention, and soft masking to enhance global dependency modeling while preserving graph topology. Finally, a multilayer perceptron (MLP) produces the final prediction scores. Experimental results demonstrate that STAGE outperforms existing methods, and case studies further validate its effectiveness and generalization capability.
Yuehu Wu, Lei Wang 0121, Zhengwei Li 0001, Mengmeng Wei, Changchun Liu 0003
BIBM2
2025 Multi-view fusion based on graph convolutional network with attention mechanism for predicting miRNA related to drugs
abstract
MicroRNAs (miRNAs) play crucial roles in cancer progression, invasion, and response to treatment, particularly in regulating anticancer drug resistance and sensitivity. Identifying potential human miRNA-drug associations (MDAs) that manifest as resistance or sensitivity relationships offers valuable insights for cancer treatment and drug development. With the growing availability of biological data, computational methods have emerged as powerful tools to complement experimental approaches. However, limited attention has been paid to computational prediction of MDAs. Furthermore, existing approaches typically rely on known MDA information, overlooking the valuable insights available from multi-source data related to miRNAs and drugs. In this study, we present a multi-view fusion-based graph convolutional network with attention mechanism (MGCNA) to predict miRNA-associated drug resistance/sensitivity. Specifically, MGCNA integrates macro- and micro- level information of miRNAs and drugs to construct multi-view node features from different perspectives. The proposed multi-view graph convolutional network (GCN) encoder obtains miRNA and disease features from different views and learns adaptive importance weights of the embedding using an attention mechanism. Extensive experiments on manually curated benchmark datasets demonstrate that MGCNA outperforms existing baseline methods. Case studies of two common drugs further establish MGCNA's effectiveness in discovering novel MDAs.
Nan Sheng, Yun-Zhi Liu, Lei Wang 0121, Lan Huang 0002, Yan Wang 0028
PLoS Comput. Biol.4
2025 Hypergraph representation learning for identifying circRNA-disease associations
Yang Li 0111, Xuegang Hu, Pei-Pei Li 0001, Lei Wang 0121, Zhu-Hong You
Pattern Recognit.4
2025 Collaborative Framework for circRNA-Disease Associations Prediction Using Dual Variational Graph
abstract
Many experiments have shown that circular RNA (circRNA) can act as biomarkers for complex diseases and play significant regulatory roles in multiple pathological processes. However, most circRNA-disease associations remain unknown, and discovering these associations through biological experimental approach is expensive and time-consuming. Taking into account the shortcomings of current methods, we introduce a new collaborative framework that utilizes multi-heterogeneous graphs, along with variational graph auto-encoders (VGAE) to predict associations between circRNA and diseases. First, we build multi-similarity networks using various biological attributes of circRNA and diseases, and integrate these similarity networks. Two subnetworks are constructed from association matrix and combined similarity network, which included a circRNA-based network and a disease-based network. We then employ random walk with restart and Singular Value Decomposition, for feature extraction from the similarity matrix. Finally, we use collaborative framework to predict the circRNAdisease association scores based on the two subnetworks. We integrate the two score matrices to obtain a final prediction scoring matrix. Using 5-fold cross-validation on the CircR2Disease dataset, our model achieved an AUC score of 0.9828 and an AUPR score of 0.9820. Additionally, among the top 30 highest-scoring circRNA-disease association pairs, 26 associations have already been validated. Our model shows strong performance and can accurately predict associations between circRNA and diseases, according to experimental results.
Changchun Liu 0003, Lei Wang 0121, Bo-Wei Zhao, Mengmeng Wei, Yang Li 0111, Mianshuo Lu, Si-Zhe Liang
IEEE Trans. Big Data2
2025 Self-Supervised Contrastive Learning on Attribute and Topology Graphs for Predicting Relationships Among lncRNAs, miRNAs and Diseases
abstract
Exploring associations between long non-coding RNAs (lncRNAs), microRNAs (miRNAs) and diseases is crucial for disease prevention, diagnosis and treatment. While determining these relationships experimentally is resource-intensive and time-consuming, computational methods have emerged as an attractive way. However, existing computational methods tend to focus on single tasks, neglecting the benefits of leveraging multiple biomolecular interactions and domain-specific knowledge for multi-task prediction. Furthermore, the scarcity of labeled data for lncRNA-disease associations (LDAs), miRNA-disease associations (MDAs) and lncRNA-miRNA interactions (LMIs) poses challenges for comprehensive node embedding learning. This paper proposes a multi-task prediction model (called SSCLMD) that employs self-supervised contrastive learning on attribute and topology graphs to identify potential LDAs, MDAs and LMIs. Firstly, domain knowledge of lncRNAs, miRNAs and diseases as well as their interactions are exploited to construct attribute graph and topology graph, respectively. Then, the nodes are encoded in the attribute and topology spaces to extract the specific and common feature. Meanwhile, the attention mechanism is performed to adaptively fuse the embedding from different views. SSCLMD incorporates contrastive self-supervised learning as a regularize to guide node embedding learning in both attribute and topology space without relying on labels. Severing as a regularize in multi-task learning paradigm, it to improves the model.s generalization capabilities. Extensive experiments on 2 manually curated datasets demonstrate that SSCLMD significantly outperforms baseline methods in LDA, MDA and LMI prediction tasks. Case studies on both old and new datasets further supported SSCLMD's ability to uncover novel disease-related lncRNAs and miRNAs.
Lan Huang 0002, Nan Sheng, Lei Wang 0121, Wenju Hou, Yan Wang 0028
IEEE J. Biomed. Health Informatics4
2025 Integrating Transformer and Graph Attention Network for circRNA-miRNA Interaction Prediction
abstract
CircRNA-miRNA interaction (CMI) plays a crucial role in the gene regulatory network of the cell. Numerous experiments have shown that abnormalities in CMI can impact molecular functions and physiological processes, leading to the occurrence of specific diseases. Current computational models for predicting CMI typically focus on local molecular entity relationships, thereby neglecting inherent molecular attributes and global structural information. To address these limitations, we propose a multi-feature fusion prediction model based on the transformer and graph attention network, named EGATCMI. Specifically, EGATCMI combines the transformer architecture with Word2vec to pre-train the sequence of circRNA and miRNA, capturing their sequence feature representation and sequence similarity. By leveraging the self-attention mechanism, EGATCMI extracts global structural feature from the CMI network. EGATCMI effectively integrates the obtained multi-feature for prediction, achieving AUC values of 0.9106 and 0.9470 on the CMI-9905 and CircBank datasets, respectively, outperforming existing methods. In case studies that the prediction of interactions between three miRNAs that are closely related to diseases and circRNAs, 8 out of 10 pairs were accurately predicted and validated. Extensive experimental results demonstrate the potential of EGATCMI as a reliable tool for candidate screening in biological investigations.
Mengmeng Wei, Lei Wang 0121, Bo-Wei Zhao, Xiao-Rui Su 0001, Zhu-Hong You, De-Shuang Huang
IEEE J. Biomed. Health Informatics2
2025 Local-Global Structure-Aware Geometric Equivariant Graph Representation Learning for Predicting Protein-Ligand Binding Affinity
abstract
Predicting protein-ligand binding affinities is a critical problem in drug discovery and design. A majority of existing methods fail to accurately characterize and exploit the geometrically invariant structures of protein-ligand complexes for predicting binding affinities. In this study, we propose Geo-protein-ligand binding affinity (PLA), a geometric equivariant graph representation learning framework with local-global structure awareness, to predict binding affinity by capturing the geometric information of protein-ligand complexes. Specifically, the local structural information of 3-D protein-ligand complexes is extracted by using an equivariant graph neural network (EGNN), which iteratively updates node representations while preserving the equivariance of coordinate transformations. Meanwhile, a graph transformer is utilized to capture long-range interactions among atoms, offering a global view that adaptively focuses on complex regions with a significant impact on binding affinities. Furthermore, the multiscale information from the two channels is integrated to enhance the predictive capability of the model. Extensive experimental studies on two benchmark datasets confirm the superior performance of Geo-PLA. Moreover, the visual interpretation of the learned protein-ligand complexes further indicates that our model offers valuable biological insights for virtual screening and drug repositioning.
Zhu-Hong You, Xuequn Shang 0001, Lei Wang 0121, Zhen Wang 0020
IEEE Trans. Neural Networks Learn. Syst.6
2024 Predicting CircRNA-Disease Associations Through Non-negative Matrix Factorization and Adversarially Regularized Variational Graph Autoencoder
abstract
Circular RNA (circRNA) is an RNA molecule that plays an important role in both pathology and physiology. Accurate identification of associations between circRNAs and diseases is crucial for further physiology research. However, verifying the circRNA-disease associations (CDA) through biological experimental methods is time-consuming. Here, we propose a novel method combines Non-negative Matrix Factorization (NMF) and Adversarially Regularized Variational Graph Autoencoder (ARVGA) to accurately predict CDA. Our model first fuses multi-source information in order to build circRNA similarity, disease similarity and circRNA-disease association matrices. Thereby our model constructs graphs for circRNA and disease respectively and optimize them using a K-means clustering algorithm. We obtain linear features using NMF and non-linear features using ARVGA. Finally, an Extremely Randomized Trees classifier is employed to predict CDA. On the gold standard dataset CircR2Disease, our model achieved a prediction accuracy of 94.8% and an AUC of 0.984 under 5-fold cross-validation. In case study, 19 of top 20 predicted circRNAs associated with Hepatocellular Carcinoma were confirmed in relevant literature. Furthermore, ablation experiment, classifier experiment and independent datasets test fully demonstrate the effectiveness and robustness of our model.
Mianshuo Lu, Lei Wang 0121, Jinzhu Sun, Yang Li 0111, Mengmeng Wei, Changchun Liu 0003, Zhengwei Li 0001
BIBM2
2024 EELMCDA: Combining evolutionary ensemble learning with matrix feature decomposition for predicting circRNA-disease associations
abstract
Recent studies have indicated that circular RNAs (circRNAs) play a significant role in the diagnosis and treatment of disease. However, the prediction of associations between circRNAs and diseases using conventional biological methods is constrained by numerous factors. In this study, we proposed a novel computational model called EELMCDA that combines evolutionary ensemble learning (EEL) approach and matrix feature decomposition method to predict potential circRNA-disease associations. The model firstly integrates circRNA function information, disease semantic information, and circRNA and disease gaussian interaction profile kernel (GIPK) information into an integrated matrix and constructed the corresponding feature matrix, then uses the matrix feature decomposition algorithm to obtain its important feature, and finally adopted evolutionary ensemble learning module to predict circRNA-disease associations. The average accuracy of the EELMCDA model by 5-fold cross-validation on CircR2Disease, CircAtlasv2.0, Circ2Disease, and CircRNADisease datasets were 92.40%, 92.90%, 88.91%, and 90.74%, respectively. Moreover, in case studies, the 21 of the top 30 circRNA-disease pairs with the highest EELMCDA scores were validated in recent literatures. These results further demonstrate the effectiveness of EELMCDA in predicting circRNA-disease associations.
Zheng Wang 0065, Lei Wang 0030, Zhu-Hong You, Lei Wang 0121, Yang Li 0111
BIBM4
2024 Likelihood-based feature representation learning combined with neighborhood information for predicting circRNA-miRNA associations
abstract
Connections between circular RNAs (circRNAs) and microRNAs (miRNAs) assume a pivotal position in the onset, evolution, diagnosis and treatment of diseases and tumors. Selecting the most potential circRNA-related miRNAs and taking advantage of them as the biological markers or drug targets could be conducive to dealing with complex human diseases through preventive strategies, diagnostic procedures and therapeutic approaches. Compared to traditional biological experiments, leveraging computational models to integrate diverse biological data in order to infer potential associations proves to be a more efficient and cost-effective approach. This paper developed a model of Convolutional Autoencoder for CircRNA-MiRNA Associations (CA-CMA) prediction. Initially, this model merged the natural language characteristics of the circRNA and miRNA sequence with the features of circRNA-miRNA interactions. Subsequently, it utilized all circRNA-miRNA pairs to construct a molecular association network, which was then fine-tuned by labeled samples to optimize the network parameters. Finally, the prediction outcome is obtained by utilizing the deep neural networks classifier. This model innovatively combines the likelihood objective that preserves the neighborhood through optimization, to learn the continuous feature representation of words and preserve the spatial information of two-dimensional signals. During the process of 5-fold cross-validation, CA-CMA exhibited exceptional performance compared to numerous prior computational approaches, as evidenced by its mean area under the receiver operating characteristic curve of 0.9138 and a minimal SD of 0.0024. Furthermore, recent literature has confirmed the accuracy of 25 out of the top 30 circRNA-miRNA pairs identified with the highest CA-CMA scores during case studies. The results of these experiments highlight the robustness and versatility of our model.
Lu-Xiang Guo, Lei Wang 0121, Zhu-Hong You, Meng-Lei Hu, Bo-Wei Zhao, Yang Li 0111
Briefings Bioinform.2
2024 Biolinguistic graph fusion model for circRNA-miRNA association prediction
abstract
Emerging clinical evidence suggests that sophisticated associations with circular ribonucleic acids (RNAs) (circRNAs) and microRNAs (miRNAs) are a critical regulatory factor of various pathological processes and play a critical role in most intricate human diseases. Nonetheless, the above correlations via wet experiments are error-prone and labor-intensive, and the underlying novel circRNA-miRNA association (CMA) has been validated by numerous existing computational methods that rely only on single correlation data. Considering the inadequacy of existing machine learning models, we propose a new model named BGF-CMAP, which combines the gradient boosting decision tree with natural language processing and graph embedding methods to infer associations between circRNAs and miRNAs. Specifically, BGF-CMAP extracts sequence attribute features and interaction behavior features by Word2vec and two homogeneous graph embedding algorithms, large-scale information network embedding and graph factorization, respectively. Multitudinous comprehensive experimental analysis revealed that BGF-CMAP successfully predicted the complex relationship between circRNAs and miRNAs with an accuracy of 82.90% and an area under receiver operating characteristic of 0.9075. Furthermore, 23 of the top 30 miRNA-associated circRNAs of the studies on data were confirmed in relevant experiences, showing that the BGF-CMAP model is superior to others. BGF-CMAP can serve as a helpful model to provide a scientific theoretical basis for the study of CMA prediction.
Lu-Xiang Guo, Lei Wang 0121, Zhu-Hong You, Meng-Lei Hu, Bo-Wei Zhao, Yang Li 0111
Briefings Bioinform.2
2024 A multi-task prediction method based on neighborhood structure embedding and signed graph representation learning to infer the relationship between circRNA, miRNA, and cancer
abstract
MOTIVATION: Research shows that competing endogenous RNA is widely involved in gene regulation in cells, and identifying the association between circular RNA (circRNA), microRNA (miRNA), and cancer can provide new hope for disease diagnosis, treatment, and prognosis. However, affected by reductionism, previous studies regarded the prediction of circRNA-miRNA interaction, circRNA-cancer association, and miRNA-cancer association as separate studies. Currently, few models are capable of simultaneously predicting these three associations. RESULTS: Inspired by holism, we propose a multi-task prediction method based on neighborhood structure embedding and signed graph representation learning, CMCSG, to infer the relationship between circRNA, miRNA, and cancer. Our method aims to extract feature descriptors of all molecules from the circRNA-miRNA-cancer regulatory network using known types of association information to predict unknown types of molecular associations. Specifically, we first constructed the circRNA-miRNA-cancer association network (CMCN), which is constructed based on the experimentally verified biomedical entity regulatory network; next, we combine topological structure embedding methods to extract feature representations in CMCN from local and global perspectives, and use denoising autoencoder for enhancement; then, combined with balance theory and state theory, molecular features are extracted from the point of social relations through the propagation and aggregation of signed graph attention network; finally, the GBDT classifier is used to predict the association of molecules. The results show that CMCSG can effectively predict the relationship between circRNA, miRNA, and cancer. Additionally, the case studies also demonstrate that CMCSG is capable of accurately identifying biomarkers across various types of cancer. The data and source code can be found at https://github.com/1axin/CMCSG.
Lan Huang 0002, Xinfei Wang 0001, Yan Wang 0028, Renchu Guan, Nan Sheng, Xuping Xie, Lei Wang 0121
Briefings Bioinform.7
2024 HHOMR: a hybrid high-order moment residual model for miRNA-disease association prediction
abstract
Numerous studies have demonstrated that microRNAs (miRNAs) are critically important for the prediction, diagnosis, and characterization of diseases. However, identifying miRNA-disease associations through traditional biological experiments is both costly and time-consuming. To further explore these associations, we proposed a model based on hybrid high-order moments combined with element-level attention mechanisms (HHOMR). This model innovatively fused hybrid higher-order statistical information along with structural and community information. Specifically, we first constructed a heterogeneous graph based on existing associations between miRNAs and diseases. HHOMR employs a structural fusion layer to capture structure-level embeddings and leverages a hybrid high-order moments encoder layer to enhance features. Element-level attention mechanisms are then used to adaptively integrate the features of these hybrid moments. Finally, a multi-layer perceptron is utilized to calculate the association scores between miRNAs and diseases. Through five-fold cross-validation on HMDD v2.0, we achieved a mean AUC of 93.28%. Compared with four state-of-the-art models, HHOMR exhibited superior performance. Additionally, case studies on three diseases-esophageal neoplasms, lymphoma, and prostate neoplasms-were conducted. Among the top 50 miRNAs with high disease association scores, 46, 47, and 45 associated with these diseases were confirmed by the dbDEMC and miR2Disease databases, respectively. Our results demonstrate that HHOMR not only outperforms existing models but also shows significant potential in predicting miRNA-disease associations.
Zhengwei Li 0001, Lei Wang 0121, Ru Nie
Briefings Bioinform.3
2024 BEROLECMI: a novel prediction method to infer circRNA-miRNA interaction from the role definition of molecular attributes and biological networks
abstract
Circular RNA (CircRNA)-microRNA (miRNA) interaction (CMI) is an important model for the regulation of biological processes by non-coding RNA (ncRNA), which provides a new perspective for the study of human complex diseases. However, the existing CMI prediction models mainly rely on the nearest neighbor structure in the biological network, ignoring the molecular network topology, so it is difficult to improve the prediction performance. In this paper, we proposed a new CMI prediction method, BEROLECMI, which uses molecular sequence attributes, molecular self-similarity, and biological network topology to define the specific role feature representation for molecules to infer the new CMI. BEROLECMI effectively makes up for the lack of network topology in the CMI prediction model and achieves the highest prediction performance in three commonly used data sets. In the case study, 14 of the 15 pairs of unknown CMIs were correctly predicted.
Xinfei Wang 0001, Zhu-Hong You, Yan Wang 0028, Lan Huang 0002, Yan Qiao 0002, Lei Wang 0121, Zhengwei Li 0001
BMC Bioinform.7
2024 BioKG-CMI: a multi-source feature fusion model based on biological knowledge graph for predicting circRNA-miRNA interactions
Mengmeng Wei, Lei Wang 0121, Bo-Wei Zhao, Xiao-Rui Su 0001, Zhu-Hong You
Sci. China Inf. Sci.2
2024 AMDECDA: Attention Mechanism Combined With Data Ensemble Strategy for Predicting CircRNA-Disease Association
abstract
Accumulating evidence from recent research reveals that circRNA is tightly bound to human complex disease and plays an important regulatory role in disease progression. Identifying disease-associated circRNA occupies a key role in the research of disease pathogenesis. In this study, we propose a new model AMDECDA for predicting circRNA-disease association (CDA) by combining attention mechanism and data ensemble strategy. Firstly, we fuse the heterogeneous information including circRNA Gaussian interaction profile (GIP), disease semantics and disease GIP, and then use the attention mechanism of Graph Attention Network (GAT) to focus on the critical information of data, reasonably allocate resources and extract their essential features. Finally, the ensemble deep RVFL network (edRVFL) is utilized to quickly and accurately predict CDA in the non-iterative manner of closed-form solutions. In the five-fold cross-validation experiment on the benchmark data set, AMDECDA achieves an accuracy of 93.10% with a sensitivity of 97.56% in 0.9235 AUC. In comparison with previous models, AMDECDA exhibits highly competitiveness. Furthermore, 26 of the top 30 unknown CDAs of AMDECDA predicted scores are proved by the related literature. These results indicate that AMDECDA can effectively anticipate latent CDA and provide help for further biological wet experiments.
Lei Wang 0121, Leon Wong, Zhu-Hong You, De-Shuang Huang
IEEE Trans. Big Data1
2024 LMGATCDA: Graph Neural Network With Labeling Trick for Predicting circRNA-Disease Associations
abstract
Previous studies have proven that circular RNAs (circRNAs) are inextricably connected to the etiology and pathophysiology of complicated diseases. Since conventional biological research are frequently small-scale, expensive, and time-consuming, it is essential to establish an efficient and reasonable computation-based method to identify disease-related circRNAs. In this article, we proposed a novel ensemble model for predicting probable circRNA-disease associations based on multi-source similarity information(LMGATCDA). In particular, LMGATCDA first incorporates information on circRNA functional similarity, disease semantic similarity, and the Gaussian interaction profile (GIP) kernel similarity as explicit features, along with node-labeling of the three-hop subgraphs extracted from each linked target node as graph structural features. After that, the fused features are used as input, and further implied features are extracted by graph sampling aggregation (GraphSAGE) and multi-hop attention graph neural network (MAGNA). Finally, the prediction scores are obtained through a fully connected layer. With five-fold cross-validation, LMGATCDA demonstrated excellent competitiveness against gold standard data, reaching 95.37% accuracy and 91.31% recall with an AUC of 94.25% on the circR2Disease benchmark dataset. Collectively, the noteworthy findings from these case studies support our conclusion that the LMGATCDA model can provide reliable circRNA-disease associations for clinical research while helping to mitigate experimental uncertainties in wet-lab investigations.
Pengyong Han, Zhengwei Li 0001, Ru Nie, Kangwei Wang, Lei Wang 0121, Hongmei Liao
IEEE ACM Trans. Comput. Biol. Bioinform.6
2024 GSLCDA: An Unsupervised Deep Graph Structure Learning Method for Predicting CircRNA-Disease Association
abstract
Growing studies reveal that Circular RNAs (circRNAs) are broadly engaged in physiological processes of cell proliferation, differentiation, aging, apoptosis, and are closely associated with the pathogenesis of numerous diseases. Clarification of the correlation among diseases and circRNAs is of great clinical importance to provide new therapeutic strategies for complex diseases. However, previous circRNA-disease association prediction methods rely excessively on the graph network, and the model performance is dramatically reduced when noisy connections occur in the graph structure. To address this problem, this paper proposes an unsupervised deep graph structure learning method GSLCDA to predict potential CDAs. Concretely, we first integrate circRNA and disease multi-source data to constitute the CDA heterogeneous network. Then the network topology is learned using the graph structure, and the original graph is enhanced in an unsupervised manner by maximize the inter information of the learned and original graphs to uncover their essential features. Finally, graph space sensitive k-nearest neighbor (KNN) algorithm is employed to search for latent CDAs. In the benchmark dataset, GSLCDA obtained 92.67% accuracy with 0.9279 AUC. GSLCDA also exhibits exceptional performance on independent datasets. Furthermore, 14, 12 and 14 of the top 16 circRNAs with the most points GSLCDA prediction scores were confirmed in the relevant literature in the breast cancer, colorectal cancer and lung cancer case studies, respectively. Such results demonstrated that GSLCDA can validly reveal underlying CDA and offer new perspectives for the diagnosis and therapy of complex human diseases.
Lei Wang 0121, Zhengwei Li 0001, Zhu-Hong You, De-Shuang Huang, Leon Wong
IEEE J. Biomed. Health Informatics1
2024 MAGCDA: A Multi-Hop Attention Graph Neural Networks Method for CircRNA-Disease Association Prediction
abstract
With a growing body of evidence establishing circular RNAs (circRNAs) are widely exploited in eukaryotic cells and have a significant contribution in the occurrence and development of many complex human diseases. Disease-associated circRNAs can serve as clinical diagnostic biomarkers and therapeutic targets, providing novel ideas for biopharmaceutical research. However, available computation methods for predicting circRNA-disease associations (CDAs) do not sufficiently consider the contextual information of biological network nodes, making their performance limited. In this work, we propose a multi-hop attention graph neural network-based approach MAGCDA to infer potential CDAs. Specifically, we first construct a multi-source attribute heterogeneous network of circRNAs and diseases, then use a multi-hop strategy of graph nodes to deeply aggregate node context information through attention diffusion, thus enhancing topological structure information and mining data hidden features, and finally use random forest to accurately infer potential CDAs. In the four gold standard data sets, MAGCDA achieved prediction accuracy of 92.58%, 91.42%, 83.46% and 91.12%, respectively. MAGCDA has also presented prominent achievements in ablation experiments and in comparisons with other models. Additionally, 18 and 17 potential circRNAs in top 20 predicted scores for MAGCDA prediction scores were confirmed in case studies of the complex diseases breast cancer and Almozheimer's disease, respectively. These results suggest that MAGCDA can be a practical tool to explore potential disease-associated circRNAs and provide a theoretical basis for disease diagnosis and treatment.
Lei Wang 0121, Zhengwei Li 0001, Zhu-Hong You, De-Shuang Huang, Leon Wong
IEEE J. Biomed. Health Informatics1
2023 SPRDA: a link prediction approach based on the structural perturbation to infer disease-associated Piwi-interacting RNAs
abstract
piRNA and PIWI proteins have been confirmed for disease diagnosis and treatment as novel biomarkers due to its abnormal expression in various cancers. However, the current research is not strong enough to further clarify the functions of piRNA in cancer and its underlying mechanism. Therefore, how to provide large-scale and serious piRNA candidates for biological research has grown up to be a pressing issue. In this study, a novel computational model based on the structural perturbation method is proposed to predict potential disease-associated piRNAs, called SPRDA. Notably, SPRDA belongs to positive-unlabeled learning, which is unaffected by negative examples in contrast to previous approaches. In the 5-fold cross-validation, SPRDA shows high performance on the benchmark dataset piRDisease, with an AUC of 0.9529. Furthermore, the predictive performance of SPRDA for 10 diseases shows the robustness of the proposed method. Overall, the proposed approach can provide unique insights into the pathogenesis of the disease and will advance the field of oncology diagnosis and treatment.
Kai Zheng 0020, Xin-Lu Zhang, Lei Wang 0121, Zhu-Hong You, Zhengwei Li 0001
Briefings Bioinform.3
2023 GKLOMLI: a link prediction model for inferring miRNA-lncRNA interactions by using Gaussian kernel-based method on network profile and linear optimization algorithm
abstract
BACKGROUND: The limited knowledge of miRNA-lncRNA interactions is considered as an obstruction of revealing the regulatory mechanism. Accumulating evidence on Human diseases indicates that the modulation of gene expression has a great relationship with the interactions between miRNAs and lncRNAs. However, such interaction validation via crosslinking-immunoprecipitation and high-throughput sequencing (CLIP-seq) experiments that inevitably costs too much money and time but with unsatisfactory results. Therefore, more and more computational prediction tools have been developed to offer many reliable candidates for a better design of further bio-experiments. METHODS: In this work, we proposed a novel link prediction model based on Gaussian kernel-based method and linear optimization algorithm for inferring miRNA-lncRNA interactions (GKLOMLI). Given an observed miRNA-lncRNA interaction network, the Gaussian kernel-based method was employed to output two similarity matrixes of miRNAs and lncRNAs. Based on the integrated matrix combined with similarity matrixes and the observed interaction network, a linear optimization-based link prediction model was trained for inferring miRNA-lncRNA interactions. RESULTS: To evaluate the performance of our proposed method, k-fold cross-validation (CV) and leave-one-out CV were implemented, in which each CV experiment was carried out 100 times on a training set generated randomly. The high area under the curves (AUCs) at 0.8623 ± 0.0027 (2-fold CV), 0.9053 ± 0.0017 (5-fold CV), 0.9151 ± 0.0013 (10-fold CV), and 0.9236 (LOO-CV), illustrated the precision and reliability of our proposed method. CONCLUSION: GKLOMLI with high performance is anticipated to be used to reveal underlying interactions between miRNA and their target lncRNAs, and deciphers the potential mechanisms of the complex diseases.
Leon Wong, Lei Wang 0121, Zhu-Hong You, Chang-an Yuan 0001, Mei-Yuan Cao
BMC Bioinform.2
2023 Relation-propagation meta-learning on an explicit preference graph for cold-start recommendation
Huiting Liu 0001, Lei Wang 0121, Pei-Pei Li 0001, Peng Zhao 0010, Xindong Wu 0001
Knowl. Based Syst.2
2023 Predicting MiRNA-Disease Associations by Graph Representation Learning Based on Jumping Knowledge Networks
abstract
Growing studies have shown that miRNAs are inextricably linked with many human diseases, and a great deal of effort has been spent on identifying their potential associations. Compared with traditional experimental methods, computational approaches have achieved promising results. In this article, we propose a graph representation learning method to predict miRNA-disease associations. Specifically, we first integrate the verified miRNA-disease associations with the similarity information of miRNA and disease to construct a miRNA-disease heterogeneous graph. Then, we apply a graph attention network to aggregate the neighbor information of nodes in each layer, and then feed the representation of the hidden layer into the structure-aware jumping knowledge network to obtain the global features of nodes. The output features of miRNAs and diseases are then concatenated and fed into a fully connected layer to score the potential associations. Through five-fold cross-validation, the average AUC, accuracy and precision values of our model are 93.30%, 85.18% and 88.90%, respectively. In addition, for three case studies of the esophageal tumor, lymphoma and prostate tumor, 46, 45 and 45 of the top 50 miRNAs predicted by our model were confirmed by relevant databases. Overall, our method could provide a reliable alternative for miRNA-disease association prediction.
Zhengwei Li 0001, Chang-an Yuan 0001, Pengyong Han, Zhu-Hong You, Lei Wang 0121
IEEE ACM Trans. Comput. Biol. Bioinform.6
2023 MGRCDA: Metagraph Recommendation Method for Predicting CircRNA-Disease Association
abstract
Clinical evidence began to accumulate, suggesting that circRNAs can be novel therapeutic targets for various diseases and play a critical role in human health. However, limited by the complex mechanism of circRNA, it is difficult to quickly and large-scale explore the relationship between disease and circRNA in the wet-lab experiment. In this work, we design a new computational model MGRCDA on account of the metagraph recommendation theory to predict the potential circRNA-disease associations. Specifically, we first regard the circRNA-disease association prediction problem as the system recommendation problem, and design a series of metagraphs according to the heterogeneous biological networks; then extract the semantic information of the disease and the Gaussian interaction profile kernel (GIPK) similarity of circRNA and disease as network attributes; finally, the iterative search of the metagraph recommendation algorithm is used to calculate the scores of the circRNA-disease pair. On the gold standard dataset circR2Disease, MGRCDA achieved a prediction accuracy of 92.49% with an area under the ROC curve of 0.9298, which is significantly higher than other state-of-the-art models. Furthermore, among the top 30 disease-related circRNAs recommended by the model, 25 have been verified by the latest published literature. The experimental results prove that MGRCDA is feasible and efficient, and it can recommend reliable candidates to further wet-lab experiment and reduce the scope of the experiment.
Lei Wang 0121, Zhu-Hong You, De-Shuang Huang, Jianqiang Li 0001
IEEE Trans. Cybern.1
2023 PPAEDTI: Personalized Propagation Auto-Encoder Model for Predicting Drug-Target Interactions
abstract
Identifying protein targets for drugs establishes an indispensable knowledge foundation for drug repurposing and drug development. Though expensive and time-consuming, vitro trials are widely employed to discover drug targets, and the existing relevant computational algorithms still cannot satisfy the demand for real application in drug R&D with regards to the prediction accuracy and performance efficiency, which are urgently needed to be improved. To this end, we propose here the PPAEDTI model, which uses the graph personalized propagation technique to predict drug-target interactions from the known interaction network. To evaluate the prediction performance, six benchmark datasets were used for testing with some state-of-the-art methods compared. As a result, using the 5-fold cross-validation, the proposed PPAEDTI model achieves average AUCs>90% on 5 collected datasets. We also manually checked the top-20 prediction list for 2 proteins (hsa:775 and hsa:779) and a kind of drug (D00618), and successfully confirmed 18, 17, and 20 items from the public datasets, respectively. The experimental results indicate that, given known drug-target interactions, the PPAEDTI model can provide accurate predictions for the new ones, which is anticipated to serve as a useful tool for pharmacology research. Using the proposed model that was trained with the collected datasets, we have built a computational platform that is accessible at http://120.77.11.78/PPAEDTI/ and corresponding codes and datasets are also released.
Yue-Chao Li, Zhu-Hong You, Lei Wang 0121, Leon Wong, Lun Hu, Pengwei Hu 0001
IEEE J. Biomed. Health Informatics4
2023 Biomedical Knowledge Graph Embedding With Capsule Network for Multi-Label Drug-Drug Interaction Prediction
abstract
Drug-drug interaction (DDI) plays an important role in drug development and administration. Most of existing network-based computation models regard the DDI prediction as a binary classification problem and generate negative DDI samples randomly, but the binary classification is not in line with the real problem since there are dozens of types of DDI and randomly generating negative samples may introduce false-negative samples since the non-observed facts can be either false or just missing. To address the above limitations, we propose a new framework called KG2ECapsule that explicitly models the multi-relational DDI data based on biomedical knowledge graphs in an end-to-end fashion. It first generates high-quality negative samples based on the average number of tail entities and head entities for each relation to reduce false-negative samples to some extent. KG2ECapsule then refines the representations of entities by recursively propagating the embeddings from the attention-based receptive fields of entities. Empirical results on three biomedical knowledge graphs of different scales show that KG2ECapsule outperforms the state-of-the-art methods consistently in multi-label DDI prediction task and further studies verify the efficacy of both probability-based sampling strategy and non-linear transformation for modeling multi-relational data.
Xiao-Rui Su 0001, Zhu-Hong You, De-Shuang Huang, Lei Wang 0121, Leon Wong, Bo-Wei Zhao
IEEE Trans. Knowl. Data Eng.4
2022 Predicting circRNA-disease associations using similarity assessing graph convolution from multi-source information networks
abstract
Circular RNA (circRNA), a novel endogenous noncoding RNA molecule with a closed-loop structure, can be used as a biomarker for many complex human diseases. Determining the relationship between circRNAs and diseases helps us to understand the diagnosis, treatment, and pathogenesis of complex diseases, which plays a critical role in clinical research. Nevertheless, the discovery of new circRNA-disease associations by wet-lab methods is not only time-consuming and costly but also randomized and blinded, which is also limited to small-scale studies. Thus, there is an urgent need to establish efficient and reliable computational methods to infer potential circRNA-disease associations on a large scale to effectively reduce costs and save time, and avoid high false-positive rates. In this paper, we propose a novel computational method for predicting circRNA-disease association based on the Similarity Assessing Graph Convolution Network (SAGCN) algorithm, which combines the multi-source similarity network constructed by circRNA and disease. Firstly, we fuse the multi-source similarity information of circRNAs and diseases and construct the multi-source similarity network respectively. Then we use the SAGCN algorithm to extract the hidden feature representations of circRNAs and diseases efficiently and objectively in the way of measuring the similarity between different nodes in the network. Finally, the obtained high-level features of circRNAs and diseases are fed to the multilayer perceptron (MLP) classifier for accurate prediction. Using the 5-fold cross-validation method, the AUC scores of the four SAGCN algorithms, on the benchmark circR2Disease dataset are 93.30%, 92.98%, 92.22% and 91.94%, respectively. Furthermore, case studies further validated that the proposed model was supported by biological experiments, and 25 of the top 30 circRNA-disease associations with the highest scores were confirmed by recent literature. Based on these reliable results, it can be anticipated that the proposed model can be used as an effective computational tool to predict circRNA-disease associations and can provide the most promising candidates for biological experiments.
Yang Li 0111, Xue-Gang Hu, Pei-Pei Li 0001, Lei Wang 0121, Zhu-Hong You
BIBM4
2022 A novel circRNA-miRNA association prediction model based on structural deep neural network embedding
abstract
A large amount of clinical evidence began to mount, showing that circular ribonucleic acids (RNAs; circRNAs) perform a very important function in complex diseases by participating in transcription and translation regulation of microRNA (miRNA) target genes. However, with strict high-throughput techniques based on traditional biological experiments and the conditions and environment, the association between circRNA and miRNA can be discovered to be labor-intensive, expensive, time-consuming, and inefficient. In this paper, we proposed a novel computational model based on Word2vec, Structural Deep Network Embedding (SDNE), Convolutional Neural Network and Deep Neural Network, which predicts the potential circRNA-miRNA associations, called Word2vec, SDNE, Convolutional Neural Network and Deep Neural Network (WSCD). Specifically, the WSCD model extracts attribute feature and behaviour feature by word embedding and graph embedding algorithm, respectively, and ultimately feed them into a feature fusion model constructed by combining Convolutional Neural Network and Deep Neural Network to deduce potential circRNA-miRNA interactions. The proposed method is proved on dataset and obtained a prediction accuracy and an area under the receiver operating characteristic curve of 81.61% and 0.8898, respectively, which is shown to have much higher accuracy than the state-of-the-art models and classifier models in prediction. In addition, 23 miRNA-related circular RNAs (circRNAs) from the top 30 were confirmed in relevant experiences. In these works, all results represent that WSCD would be a helpful supplementary reliable method for predicting potential miRNA-circRNA associations compared to wet laboratory experiments.
Lu-Xiang Guo, Zhu-Hong You, Lei Wang 0121, Bo-Wei Zhao, Zhong-Hao Ren, Jie Pan 0007
Briefings Bioinform.3
2022 MNMDCDA: prediction of circRNA-disease associations by learning mixed neighborhood information from multiple distances
abstract
Emerging evidence suggests that circular RNA (circRNA) is an important regulator of a variety of pathological processes and serves as a promising biomarker for many complex human diseases. Nevertheless, there are relatively few known circRNA-disease associations, and uncovering new circRNA-disease associations by wet-lab methods is time consuming and costly. Considering the limitations of existing computational methods, we propose a novel approach named MNMDCDA, which combines high-order graph convolutional networks (high-order GCNs) and deep neural networks to infer associations between circRNAs and diseases. Firstly, we computed different biological attribute information of circRNA and disease separately and used them to construct multiple multi-source similarity networks. Then, we used the high-order GCN algorithm to learn feature embedding representations with high-order mixed neighborhood information of circRNA and disease from the constructed multi-source similarity networks, respectively. Finally, the deep neural network classifier was implemented to predict associations of circRNAs with diseases. The MNMDCDA model obtained AUC scores of 95.16%, 94.53%, 89.80% and 91.83% on four benchmark datasets, i.e., CircR2Disease, CircAtlas v2.0, Circ2Disease and CircRNADisease, respectively, using the 5-fold cross-validation approach. Furthermore, 25 of the top 30 circRNA-disease pairs with the best scores of MNMDCDA in the case study were validated by recent literature. Numerous experimental results indicate that MNMDCDA can be used as an effective computational tool to predict circRNA-disease associations and can provide the most promising candidates for biological experiments.
Yang Li 0111, Xue-Gang Hu, Lei Wang 0121, Pei-Pei Li 0001, Zhu-Hong You
Briefings Bioinform.3
2022 A deep learning method for repurposing antiviral drugs against new viruses via multi-view nonnegative matrix factorization and its application to SARS-CoV-2
abstract
The outbreak of COVID-19 caused by SARS-coronavirus (CoV)-2 has made millions of deaths since 2019. Although a variety of computational methods have been proposed to repurpose drugs for treating SARS-CoV-2 infections, it is still a challenging task for new viruses, as there are no verified virus-drug associations (VDAs) between them and existing drugs. To efficiently solve the cold-start problem posed by new viruses, a novel constrained multi-view nonnegative matrix factorization (CMNMF) model is designed by jointly utilizing multiple sources of biological information. With the CMNMF model, the similarities of drugs and viruses can be preserved from their own perspectives when they are projected onto a unified latent feature space. Based on the CMNMF model, we propose a deep learning method, namely VDA-DLCMNMF, for repurposing drugs against new viruses. VDA-DLCMNMF first initializes the node representations of drugs and viruses with their corresponding latent feature vectors to avoid a random initialization and then applies graph convolutional network to optimize their representations. Given an arbitrary drug, its probability of being associated with a new virus is computed according to their representations. To evaluate the performance of VDA-DLCMNMF, we have conducted a series of experiments on three VDA datasets created for SARS-CoV-2. Experimental results demonstrate that the promising prediction accuracy of VDA-DLCMNMF. Moreover, incorporating the CMNMF model into deep learning gains new insight into the drug repurposing for SARS-CoV-2, as the results of molecular docking experiments reveal that four antiviral drugs identified by VDA-DLCMNMF have the potential ability to treat SARS-CoV-2 infections.
Xiao-Rui Su 0001, Lun Hu, Zhu-Hong You, Pengwei Hu 0001, Lei Wang 0121, Bo-Wei Zhao
Briefings Bioinform.5
2022 A machine learning framework based on multi-source feature fusion for circRNA-disease association prediction
abstract
Circular RNAs (circRNAs) are involved in the regulatory mechanisms of multiple complex diseases, and the identification of their associations is critical to the diagnosis and treatment of diseases. In recent years, many computational methods have been designed to predict circRNA-disease associations. However, most of the existing methods rely on single correlation data. Here, we propose a machine learning framework for circRNA-disease association prediction, called MLCDA, which effectively fuses multiple sources of heterogeneous information including circRNA sequences and disease ontology. Comprehensive evaluation in the gold standard dataset showed that MLCDA can successfully capture the complex relationships between circRNAs and diseases and accurately predict their potential associations. In addition, the results of case studies on real data show that MLCDA significantly outperforms other existing methods. MLCDA can serve as a useful tool for circRNA-disease association prediction, providing mechanistic insights for disease research and thus facilitating the progress of disease treatment.
Lei Wang 0121, Leon Wong, Zhengwei Li 0001, Xiao-Rui Su 0001, Bo-Wei Zhao, Zhu-Hong You
Briefings Bioinform.1
2022 iGRLCDA: identifying circRNA-disease association based on graph representation learning
abstract
While the technologies of ribonucleic acid-sequence (RNA-seq) and transcript assembly analysis have continued to improve, a novel topology of RNA transcript was uncovered in the last decade and is called circular RNA (circRNA). Recently, researchers have revealed that they compete with messenger RNA (mRNA) and long noncoding for combining with microRNA in gene regulation. Therefore, circRNA was assumed to be associated with complex disease and discovering the relationship between them would contribute to medical research. However, the work of identifying the association between circRNA and disease in vitro takes a long time and usually without direction. During these years, more and more associations were verified by experiments. Hence, we proposed a computational method named identifying circRNA-disease association based on graph representation learning (iGRLCDA) for the prediction of the potential association of circRNA and disease, which utilized a deep learning model of graph convolution network (GCN) and graph factorization (GF). In detail, iGRLCDA first derived the hidden feature of known associations between circRNA and disease using the Gaussian interaction profile (GIP) kernel combined with disease semantic information to form a numeric descriptor. After that, it further used the deep learning model of GCN and GF to extract hidden features from the descriptor. Finally, the random forest classifier is introduced to identify the potential circRNA-disease association. The five-fold cross-validation of iGRLCDA shows strong competitiveness in comparison with other excellent prediction models at the gold standard data and achieved an average area under the receiver operating characteristic curve of 0.9289 and an area under the precision-recall curve of 0.9377. On reviewing the prediction results from the relevant literature, 22 of the top 30 predicted circRNA-disease associations were noted in recent published papers. These exceptional results make us believe that iGRLCDA can provide reliable circRNA-disease associations for medical research and reduce the blindness of wet-lab experiments.
Lei Wang 0121, Zhu-Hong You, Lun Hu, Bo-Wei Zhao, Zhengwei Li 0001, Yang-Ming Li
Briefings Bioinform.2
2022 HINGRL: predicting drug-disease associations with graph representation learning on heterogeneous information networks
abstract
Identifying new indications for drugs plays an essential role at many phases of drug research and development. Computational methods are regarded as an effective way to associate drugs with new indications. However, most of them complete their tasks by constructing a variety of heterogeneous networks without considering the biological knowledge of drugs and diseases, which are believed to be useful for improving the accuracy of drug repositioning. To this end, a novel heterogeneous information network (HIN) based model, namely HINGRL, is proposed to precisely identify new indications for drugs based on graph representation learning techniques. More specifically, HINGRL first constructs a HIN by integrating drug-disease, drug-protein and protein-disease biological networks with the biological knowledge of drugs and diseases. Then, different representation strategies are applied to learn the features of nodes in the HIN from the topological and biological perspectives. Finally, HINGRL adopts a Random Forest classifier to predict unknown drug-disease associations based on the integrated features of drugs and diseases obtained in the previous step. Experimental results demonstrate that HINGRL achieves the best performance on two real datasets when compared with state-of-the-art models. Besides, our case studies indicate that the simultaneous consideration of network topology and biological knowledge of drugs and diseases allows HINGRL to precisely predict drug-disease associations from a more comprehensive perspective. The promising performance of HINGRL also reveals that the utilization of rich heterogeneous information provides an alternative view for HINGRL to identify novel drug-disease associations especially for new diseases.
Bo-Wei Zhao, Lun Hu, Zhu-Hong You, Lei Wang 0121, Xiao-Rui Su 0001
Briefings Bioinform.4
2022 Line graph attention networks for predicting disease-associated Piwi-interacting RNAs
abstract
PIWI proteins and Piwi-Interacting RNAs (piRNAs) are commonly detected in human cancers, especially in germline and somatic tissues, and correlate with poorer clinical outcomes, suggesting that they play a functional role in cancer. As the problem of combinatorial explosions between ncRNA and disease exposes gradually, new bioinformatics methods for large-scale identification and prioritization of potential associations are therefore of interest. However, in the real world, the network of interactions between molecules is enormously intricate and noisy, which poses a problem for efficient graph mining. Line graphs can extend many heterogeneous networks to replace dichotomous networks. In this study, we present a new graph neural network framework, line graph attention networks (LGAT). And we apply it to predict PiRNA disease association (GAPDA). In the experiment, GAPDA performs excellently in 5-fold cross-validation with an AUC of 0.9038. Not only that, it still has superior performance compared with methods based on collaborative filtering and attribute features. The experimental results show that GAPDA ensures the prospect of the graph neural network on such problems and can be an excellent supplement for future biomedical research.
Kai Zheng 0020, Xin-Lu Zhang, Lei Wang 0121, Zhu-Hong You, Zhaohui Zhan
Briefings Bioinform.3
2022 NSECDA: Natural Semantic Enhancement for CircRNA-Disease Association Prediction
abstract
Increasing evidence suggest that circRNA, as one of the most promising emerging biomarkers, has a very close relationship with diseases. Exploring the relationship between circRNA and diseases can provide novel perspective for diseases diagnosis and pathogenesis. The existing circRNA-disease association (CDA) prediction models, however, generally treat the data attributes equally, do not pay special attention to the attributes with more significant influence, and do not make full use of the correlation and symbiosis between attributes to dig into the latent semantic information of the data. Therefore, in response to the above problems, this paper proposes a natural semantic enhancement method NSECDA to predict CDA. In practical terms, we first recognize the circRNA sequence as a biological language, and analyze its natural semantic properties through the natural language understanding theory; then integrate it with disease attributes, circRNA and disease Gaussian Interaction Profile (GIP) kernel attributes, and use Graph Attention Network (GAT) to focus on the influential attributes, so as to mine the deeply hidden features; finally, the Rotation Forest (RoF) classifier was used to accurately determine CDA. In the gold standard data set CircR2Disease, NSECDA achieved 92.49% accuracy with 0.9225 AUC score. In comparison with the non-natural semantic enhancement model and other classifier models, NSECDA also shows competitive performance. Additionally, 25 of the CDA pairs with unknown associations in the top 30 prediction scores of NSECDA have been proven by newly reported studies. These achievements suggest that NSECDA is an effective model to predict CDA, which can provide credible candidate for subsequent wet experiments, thus significantly reducing the scope of investigations.
Lei Wang 0121, Leon Wong, Zhu-Hong You, De-Shuang Huang, Xiao-Rui Su 0001, Bo-Wei Zhao
IEEE J. Biomed. Health Informatics1
2021 Predicting miRNA-Disease Associations via a New MeSH Headings Representation of Diseases and eXtreme Gradient Boosting
Zhu-Hong You, Lei Wang 0121, Leon Wong, Xiao-Rui Su 0001, Bo-Wei Zhao
ICIC (3)3
2021 CNNEMS: Using Convolutional Neural Networks to Predict Drug-Target Interactions by Combining Protein Evolution and Molecular Structures Information
Zhu-Hong You, Lei Wang 0121, Peng-Peng Chen
ICIC (3)3
2021 Predicting microRNA-disease associations from lncRNA-microRNA interactions via Multiview Multitask Learning
abstract
MOTIVATION: Identifying microRNAs that are associated with different diseases as biomarkers is a problem of great medical significance. Existing computational methods for uncovering such microRNA-diseases associations (MDAs) are mostly developed under the assumption that similar microRNAs tend to associate with similar diseases. Since such an assumption is not always valid, these methods may not always be applicable to all kinds of MDAs. Considering that the relationship between long noncoding RNA (lncRNA) and different diseases and the co-regulation relationships between the biological functions of lncRNA and microRNA have been established, we propose here a multiview multitask method to make use of the known lncRNA-microRNA interaction to predict MDAs on a large scale. The investigation is performed in the absence of complete information of microRNAs and any similarity measurement for it and to the best knowledge, the work represents the first ever attempt to discover MDAs based on lncRNA-microRNA interactions. RESULTS: In this paper, we propose to develop a deep learning model called MVMTMDA that can create a multiview representation of microRNAs. The model is trained based on an end-to-end multitasking approach to machine learning so that, based on it, missing data in the side information can be determined automatically. Experimental results show that the proposed model yields an average area under ROC curve of 0.8410+/-0.018, 0.8512+/-0.012 and 0.8521+/-0.008 when k is set to 2, 5 and 10, respectively. In addition, we also propose here a statistical approach to predicting lncRNA-disease associations based on these associations and the MDA discovered using MVMTMDA. AVAILABILITY: Python code and the datasets used in our studies are made available at https://github.com/yahuang1991polyu/MVMTMDA/.
Keith C. C. Chan, Zhu-Hong You, Pengwei Hu 0001, Lei Wang 0121, Zhi-an Huang
Briefings Bioinform.5
2021 SGANRDA: semi-supervised generative adversarial networks for predicting circRNA-disease associations
abstract
Emerging research shows that circular RNA (circRNA) plays a crucial role in the diagnosis, occurrence and prognosis of complex human diseases. Compared with traditional biological experiments, the computational method of fusing multi-source biological data to identify the association between circRNA and disease can effectively reduce cost and save time. Considering the limitations of existing computational models, we propose a semi-supervised generative adversarial network (GAN) model SGANRDA for predicting circRNA-disease association. This model first fused the natural language features of the circRNA sequence and the features of disease semantics, circRNA and disease Gaussian interaction profile kernel, and then used all circRNA-disease pairs to pre-train the GAN network, and fine-tune the network parameters through labeled samples. Finally, the extreme learning machine classifier is employed to obtain the prediction result. Compared with the previous supervision model, SGANRDA innovatively introduced circRNA sequences and utilized all the information of circRNA-disease pairs during the pre-training process. This step can increase the information content of the feature to some extent and reduce the impact of too few known associations on the model performance. SGANRDA obtained AUC scores of 0.9411 and 0.9223 in leave-one-out cross-validation and 5-fold cross-validation, respectively. Prediction results on the benchmark dataset show that SGANRDA outperforms other existing models. In addition, 25 of the top 30 circRNA-disease pairs with the highest scores of SGANRDA in case studies were verified by recent literature. These experimental results demonstrate that SGANRDA is a useful model to predict the circRNA-disease association and can provide reliable candidates for biological experiments.
Lei Wang 0121, Zhu-Hong You, Xi Zhou 0007
Briefings Bioinform.1
2021 LDGRNMF: LncRNA-disease associations prediction based on graph regularized non-negative matrix factorization
Meineng Wang, Zhu-Hong You, Lei Wang 0121, Liping Li 0003, Kai Zheng 0020
Neurocomputing3
2021 MISSIM: An Incremental Learning-Based Model With Applications to the Prediction of miRNA-Disease Association
abstract
In the past few years, the prediction models have shown remarkable performance in most biological correlation prediction tasks. These tasks traditionally use a fixed dataset, and the model, once trained, is deployed as is. These models often encounter training issues such as sensitivity to hyperparameter tuning and "catastrophic forgetting" when adding new data. However, with the development of biomedicine and the accumulation of biological data, new predictive models are required to face the challenge of adapting to change. To this end, we propose a computational approach based on Broad learning system (BLS) to predict potential disease-associated miRNAs that retain the ability to distinguish prior training associations when new data need to be adapted. In particular, we are introducing incremental learning to the field of biological association prediction for the first time and proposed a new method for quantifying sequence similarity. In the performance evaluation, the AUC in the 5-fold cross-validation was 0.9400 +/- 0.0041. To better assess the effectiveness of MISSIM, we compared it with various classifiers and former prediction models. Its performance is superior to the previous method. Besides, the case study on identifying miRNAs associated with breast neoplasms, lung neoplasms and esophageal neoplasms show that 34, 36 and 35 out of the top 40 associations predicted by MISSIM are confirmed by recent biomedical resources. These results provide ample convincing evidence of this approach have potential value and prospect in promoting biomedical research productivity.
Kai Zheng 0020, Zhu-Hong You, Lei Wang 0121, Ji-Ren Zhou, Haitao Zeng
IEEE ACM Trans. Comput. Biol. Bioinform.3
2021 IMS-CDA: Prediction of CircRNA-Disease Associations From the Integration of Multisource Similarity Information With Deep Stacked Autoencoder Model
abstract
Emerging evidence indicates that circular RNA (circRNA) has been an indispensable role in the pathogenesis of human complex diseases and many critical biological processes. Using circRNA as a molecular marker or therapeutic target opens up a new avenue for our treatment and detection of human complex diseases. The traditional biological experiments, however, are usually limited to small scale and are time consuming, so the development of an effective and feasible computational-based approach for predicting circRNA-disease associations is increasingly favored. In this study, we propose a new computational-based method, called IMS-CDA, to predict potential circRNA-disease associations based on multisource biological information. More specifically, IMS-CDA combines the information from the disease semantic similarity, the Jaccard and Gaussian interaction profile kernel similarity of disease and circRNA, and extracts the hidden features using the stacked autoencoder (SAE) algorithm of deep learning. After training in the rotation forest (RF) classifier, IMS-CDA achieves 88.08% area under the ROC curve with 88.36% accuracy at the sensitivity of 91.38% on the CIRCR2Disease dataset. Compared with the state-of-the-art support vector machine and K -nearest neighbor models and different descriptor models, IMS-CDA achieves the best overall performance. In the case studies, eight of the top 15 circRNA-disease associations with the highest prediction score were confirmed by recent literature. These results indicated that IMS-CDA has an outstanding ability to predict new circRNA-disease associations and can provide reliable candidates for biological experiments.
Lei Wang 0121, Zhu-Hong You, Jianqiang Li 0001
IEEE Trans. Cybern.1
2020 GCNSP: A Novel Prediction Method of Self-Interacting Proteins Based on Graph Convolutional Networks
Lei Wang 0121, Zhu-Hong You, Kai Zheng 0020, Zhengwei Li 0001
ICIC (2)1
2020 DTIFS: A Novel Computational Approach for Predicting Drug-Target Interactions from Drug Structure and Protein Sequence
Zhu-Hong You, Lei Wang 0121, Liping Li 0003, Kai Zheng 0020, Meineng Wang
ICIC (2)3
2020 Image Classification Based on Deep Belief Network and YELM
ChengYong Zhang, Zhengwei Li 0001, Ru Nie, Lei Wang 0121
ICIC (1)4
2020 Predicting Human Disease-Associated piRNAs Based on Multi-source Information and Random Forest
Kai Zheng 0020, Zhu-Hong You, Lei Wang 0121
ICIC (2)3
2020 Inferring Disease-Associated Piwi-Interacting RNAs via Graph Attention Networks
Kai Zheng 0020, Zhu-Hong You, Lei Wang 0121, Leon Wong
ICIC (2)3
2020 An efficient approach based on multi-sources information to predict circRNA-disease associations using deep convolutional neural network
abstract
MOTIVATION: Emerging evidence indicates that circular RNA (circRNA) plays a crucial role in human disease. Using circRNA as biomarker gives rise to a new perspective regarding our diagnosing of diseases and understanding of disease pathogenesis. However, detection of circRNA-disease associations by biological experiments alone is often blind, limited to small scale, high cost and time consuming. Therefore, there is an urgent need for reliable computational methods to rapidly infer the potential circRNA-disease associations on a large scale and to provide the most promising candidates for biological experiments. RESULTS: In this article, we propose an efficient computational method based on multi-source information combined with deep convolutional neural network (CNN) to predict circRNA-disease associations. The method first fuses multi-source information including disease semantic similarity, disease Gaussian interaction profile kernel similarity and circRNA Gaussian interaction profile kernel similarity, and then extracts its hidden deep feature through the CNN and finally sends them to the extreme learning machine classifier for prediction. The 5-fold cross-validation results show that the proposed method achieves 87.21% prediction accuracy with 88.50% sensitivity at the area under the curve of 86.67% on the CIRCR2Disease dataset. In comparison with the state-of-the-art SVM classifier and other feature extraction methods on the same dataset, the proposed model achieves the best results. In addition, we also obtained experimental support for prediction results by searching published literature. As a result, 7 of the top 15 circRNA-disease pairs with the highest scores were confirmed by literature. These results demonstrate that the proposed model is a suitable method for predicting circRNA-disease associations and can provide reliable candidates for biological experiments. AVAILABILITY AND IMPLEMENTATION: The source code and datasets explored in this work are available at https://github.com/look0012/circRNA-Disease-association. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Lei Wang 0121, Zhu-Hong You, De-Shuang Huang, Keith C. C. Chan
Bioinform.1
2020 GCNCDA: A new method for predicting circRNA-disease associations based on Graph Convolutional Network Algorithm
abstract
Numerous evidences indicate that Circular RNAs (circRNAs) are widely involved in the occurrence and development of diseases. Identifying the association between circRNAs and diseases plays a crucial role in exploring the pathogenesis of complex diseases and improving the diagnosis and treatment of diseases. However, due to the complex mechanisms between circRNAs and diseases, it is expensive and time-consuming to discover the new circRNA-disease associations by biological experiment. Therefore, there is increasingly urgent need for utilizing the computational methods to predict novel circRNA-disease associations. In this study, we propose a computational method called GCNCDA based on the deep learning Fast learning with Graph Convolutional Networks (FastGCN) algorithm to predict the potential disease-associated circRNAs. Specifically, the method first forms the unified descriptor by fusing disease semantic similarity information, disease and circRNA Gaussian Interaction Profile (GIP) kernel similarity information based on known circRNA-disease associations. The FastGCN algorithm is then used to objectively extract the high-level features contained in the fusion descriptor. Finally, the new circRNA-disease associations are accurately predicted by the Forest by Penalizing Attributes (Forest PA) classifier. The 5-fold cross-validation experiment of GCNCDA achieved 91.2% accuracy with 92.78% sensitivity at the AUC of 90.90% on circR2Disease benchmark dataset. In comparison with different classifier models, feature extraction models and other state-of-the-art methods, GCNCDA shows strong competitiveness. Furthermore, we conducted case study experiments on diseases including breast cancer, glioma and colorectal cancer. The results showed that 16, 15 and 17 of the top 20 candidate circRNAs with the highest prediction scores were respectively confirmed by relevant literature and databases. These results suggest that GCNCDA can effectively predict potential circRNA-disease associations and provide highly credible candidates for biological experiments.
Lei Wang 0121, Zhu-Hong You, Yang-Ming Li, Kai Zheng 0020
PLoS Comput. Biol.1
2020 iCDA-CGR: Identification of circRNA-disease associations based on Chaos Game Representation
abstract
Found in recent research, tumor cell invasion, proliferation, or other biological processes are controlled by circular RNA. Understanding the association between circRNAs and diseases is an important way to explore the pathogenesis of complex diseases and promote disease-targeted therapy. Most methods, such as k-mer and PSSM, based on the analysis of high-throughput expression data have the tendency to think functionally similar nucleic acid lack direct linear homology regardless of positional information and only quantify nonlinear sequence relationships. However, in many complex diseases, the sequence nonlinear relationship between the pathogenic nucleic acid and ordinary nucleic acid is not much different. Therefore, the analysis of positional information expression can help to predict the complex associations between circRNA and disease. To fill up this gap, we propose a new method, named iCDA-CGR, to predict the circRNA-disease associations. In particular, we introduce circRNA sequence information and quantifies the sequence nonlinear relationship of circRNA by Chaos Game Representation (CGR) technology based on the biological sequence position information for the first time in the circRNA-disease prediction model. In the cross-validation experiment, our method achieved 0.8533 AUC, which was significantly higher than other existing methods. In the validation of independent data sets including circ2Disease, circRNADisease and CRDD, the prediction accuracy of iCDA-CGR reached 95.18%, 90.64% and 95.89%. Moreover, in the case studies, 19 of the top 30 circRNA-disease associations predicted by iCDA-CGR on circRDisease dataset were confirmed by newly published literature. These results demonstrated that iCDA-CGR has outstanding robustness and stability, and can provide highly credible candidates for biological experiments.
Kai Zheng 0020, Zhu-Hong You, Jianqiang Li 0001, Lei Wang 0121, Zhen-Hao Guo
PLoS Comput. Biol.4
2020 Combining High Speed ELM Learning with a Deep Convolutional Neural Network Feature Encoding for Predicting Protein-RNA Interactions
abstract
Emerging evidence has shown that RNA plays a crucial role in many cellular processes, and their biological functions are primarily achieved by binding with a variety of proteins. High-throughput biological experiments provide a lot of valuable information for the initial identification of RNA-protein interactions (RPIs), but with the increasing complexity of RPIs networks, this method gradually falls into expensive and time-consuming situations. Therefore, there is an urgent need for high speed and reliable methods to predict RNA-protein interactions. In this study, we propose a computational method for predicting the RNA-protein interactions using sequence information. The deep learning convolution neural network (CNN) algorithm is utilized to mine the hidden high-level discriminative features from the RNA and protein sequences and feed it into the extreme learning machine (ELM) classifier. The experimental results with 5-fold cross-validation indicate that the proposed method achieves superior performance on benchmark datasets (RPI1807, RPI2241, and RPI369) with the accuracy of 98.83, 90.83, and 85.63 percent, respectively. We further evaluate the performance of the proposed model by comparing it with the state-of-the-art SVM classifier and other existing methods on the same benchmark data set. In addition, we predicted the independent NPInter v2.0 data set using the model trained on RPI369. The experimental results show that our model can serve as a useful tool for predicting RNA-protein interactions.
Lei Wang 0121, Zhu-Hong You, De-Shuang Huang, Fengfeng Zhou
IEEE ACM Trans. Comput. Biol. Bioinform.1
2019 Predicting circRNA-disease associations using deep generative adversarial network based on multi-source fusion information
abstract
Circular RNA (circRNA) is a kind of novel discovered non-coding RNA molecule with a closed loop structure, which plays a critical regulatory role in human diseases. Identifying the association between circRNAs and diseases has important potential value for the diagnosis and treatment of complex human diseases. Although biological experiments can more accurately identify the association between circRNAs and diseases, they are usually blind and limited by small scale and high cost. Therefore, there is an urgent need for efficient and feasible computational methods to predict the potential circRNA-disease associations on a large scale, so as to provide the most promising candidate for biological experiments. In this paper, we propose a novel computational method based on the deep Generative Adversarial Network (GAN) algorithm combined with the multi-source similarity information to predict the circRNA-disease associations. Firstly, we fuse the multi-source information of disease semantic similarity, disease and circRNA Gaussian interaction profile kernel similarity, and then use GAN to extract the hidden features of fusion information objectively and effectively in the way of confrontation learning, and finally send them to Logistic Model Tree (LMT) classifier for accurate prediction. The 5-fold cross-validation experiment of the proposed model achieved 89.2% accuracy with 89.4% precision at the AUC of 90.6% on the CIRCR2Disease dataset. Compared with the state-of-the-art SVM classifier and other feature extraction methods, the proposed model shows strong competitiveness. In addition, the predicted results of this model are supported by the biological experiments, and 9 of the top 15 circRNA-disease associations with the highest scores were confirmed by recently published literature. These promising results indicate that the proposed model is an effective tool for predicting circRNA-disease associations and can provide reliable candidates for biological experiments.
Lei Wang 0121, Zhu-Hong You, Liping Li 0003, Kai Zheng 0020
BIBM1
2019 Precise Prediction of Pathogenic Microorganisms Using 16S rRNA Gene Sequences
Zhi-an Huang, Zhu-Hong You, Pengwei Hu 0001, Liping Li 0003, Zhengwei Li 0001, Lei Wang 0121
ICIC (2)7
2019 MISSIM: Improved miRNA-Disease Association Prediction Model Based on Chaos Game Representation and Broad Learning System
Kai Zheng 0020, Zhu-Hong You, Lei Wang 0121, Hanjing Jiang
ICIC (3)3
2019 LMTRDA: Using logistic model tree to predict MiRNA-disease associations by fusing multi-source information of sequences and similarities
abstract
Emerging evidence has shown microRNAs (miRNAs) play an important role in human disease research. Identifying potential association among them is significant for the development of pathology, diagnose and therapy. However, only a tiny portion of all miRNA-disease pairs in the current datasets are experimentally validated. This prompts the development of high-precision computational methods to predict real interaction pairs. In this paper, we propose a new model of Logistic Model Tree for predicting miRNA-Disease Association (LMTRDA) by fusing multi-source information including miRNA sequences, miRNA functional similarity, disease semantic similarity, and known miRNA-disease associations. In particular, we introduce miRNA sequence information and extract its features using natural language processing technique for the first time in the miRNA-disease prediction model. In the cross-validation experiment, LMTRDA obtained 90.51% prediction accuracy with 92.55% sensitivity at the AUC of 90.54% on the HMDD V3.0 dataset. To further evaluate the performance of LMTRDA, we compared it with different classifier and feature descriptor models. In addition, we also validate the predictive ability of LMTRDA in human diseases including Breast Neoplasms, Breast Neoplasms and Lymphoma. As a result, 28, 27 and 26 out of the top 30 miRNAs associated with these diseases were verified by experiments in different kinds of case studies. These experimental results demonstrate that LMTRDA is a reliable model for predicting the association among miRNAs and diseases.
Lei Wang 0121, Zhu-Hong You, Xing Chen 0001, Yang-Ming Li, Ya-Nan Dong, Liping Li 0003, Kai Zheng 0020
PLoS Comput. Biol.1
2018 Predicting miRNA-disease association based on inductive matrix completion
abstract
Motivation: It has been shown that microRNAs (miRNAs) play key roles in variety of biological processes associated with human diseases. In Consideration of the cost and complexity of biological experiments, computational methods for predicting potential associations between miRNAs and diseases would be an effective complement. Results: This paper presents a novel model of Inductive Matrix Completion for MiRNA-Disease Association prediction (IMCMDA). The integrated miRNA similarity and disease similarity are calculated based on miRNA functional similarity, disease semantic similarity and Gaussian interaction profile kernel similarity. The main idea is to complete the missing miRNA-disease association based on the known associations and the integrated miRNA similarity and disease similarity. IMCMDA achieves AUC of 0.8034 based on leave-one-out-cross-validation and improved previous models. In addition, IMCMDA was applied to five common human diseases in three types of case studies. In the first type, respectively, 42, 44, 45 out of top 50 predicted miRNAs of Colon Neoplasms, Kidney Neoplasms, Lymphoma were confirmed by experimental reports. In the second type of case study for new diseases without any known miRNAs, we chose Breast Neoplasms as the test example by hiding the association information between the miRNAs and Breast Neoplasms. As a result, 50 out of top 50 predicted Breast Neoplasms-related miRNAs are verified. In the third type of case study, IMCMDA was tested on HMDD V1.0 to assess the robustness of IMCMDA, 49 out of top 50 predicted Esophageal Neoplasms-related miRNAs are verified. Availability and implementation: The code and dataset of IMCMDA are freely available at https://github.com/IMCMDAsourcecode/IMCMDA. Supplementary information: Supplementary data are available at Bioinformatics online.
Xing Chen 0001, Lei Wang 0121, Na-Na Guan, Jianqiang Li 0001
Bioinform.2
2018 BNPMDA: Bipartite Network Projection for MiRNA-Disease Association prediction
abstract
Motivation: A large number of resources have been devoted to exploring the associations between microRNAs (miRNAs) and diseases in the recent years. However, the experimental methods are expensive and time-consuming. Therefore, the computational methods to predict potential miRNA-disease associations have been paid increasing attention. Results: In this paper, we proposed a novel computational model of Bipartite Network Projection for MiRNA-Disease Association prediction (BNPMDA) based on the known miRNA-disease associations, integrated miRNA similarity and integrated disease similarity. We firstly described the preference degree of a miRNA for its related disease and the preference degree of a disease for its related miRNA with the bias ratings. We constructed bias ratings for miRNAs and diseases by using agglomerative hierarchical clustering according to the three types of networks. Then, we implemented the bipartite network recommendation algorithm to predict the potential miRNA-disease associations by assigning transfer weights to resource allocation links between miRNAs and diseases based on the bias ratings. BNPMDA had been shown to improve the prediction accuracy in comparison with previous models according to the area under the receiver operating characteristics (ROC) curve (AUC) results of three typical cross validations. As a result, the AUCs of Global LOOCV, Local LOOCV and 5-fold cross validation obtained by implementing BNPMDA were 0.9028, 0.8380 and 0.8980 ± 0.0013, respectively. We further implemented two types of case studies on several important human complex diseases to confirm the effectiveness of BNPMDA. In conclusion, BNPMDA could effectively predict the potential miRNA-disease associations at a high accuracy level. Availability and implementation: BNPMDA is available via http://www.escience.cn/system/file?fileId=99559. Supplementary information: Supplementary data are available at Bioinformatics online.
Xing Chen 0001, Di Xie, Lei Wang 0121, Qi Zhao 0010, Zhu-Hong You, Hongsheng Liu 0001
Bioinform.3
2018 An improved efficient rotation forest algorithm to predict the interactions among proteins
Lei Wang 0121, Zhu-Hong You, Shixiong Xia, Xing Chen 0001, Yong Zhou 0003, Feng Liu 0039
Soft Comput.1
2017 Computational Methods for the Prediction of Drug-Target Interactions from Drug Fingerprints and Protein Sequences by Stacked Auto-Encoder Deep Neural Network
Lei Wang 0121, Zhu-Hong You, Xing Chen 0001, Shixiong Xia, Feng Liu 0039, Yong Zhou 0003
ISBRA1