EDBT 2026 Demo / reviewers in the wild / expert
Bo-Wei Zhao
dblp:276/7969 · also Bowei Zhao
· DBLP profile ↗
44ranked-venue papers
9as first author
42since 2021 · last 2026
0000-0001-8200-6016ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 38 · 8 first-author · 36 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Dual-Channel Learning Framework for Zero-Shot CircRNA-miRNA Interaction Prediction via State Space ModelingabstractCircRNA-miRNA interaction (CMI) plays a pivotal role in disease therapeutics and drug discovery. However, existing methods face several challenges in modeling complex biological networks and zero-shot learning scenarios. Biological networks encapsulate rich biological information, yet current approaches often fail to fully exploit this depth. Moreover, zero-shot prediction requires models to identify new interactions without relying on previously observed samples, imposing stringent requirements on generalization capabilities. To address these limitations, we propose a dual-channel learning framework leveraging State space modeling for Zero-shot CMI prediction (ZeroStem). ZeroStem first enhances the biological relevance of node using prior knowledge, and employs a graph Transformer to extract macro-topological representations. Subsequently, it generates semantic subgraphs based on meta-paths to focus on specific biological relationships, utilizing the Mamba to extract micro-semantic representations via state space modeling. Finally, macro-topological and micro-semantic representations are seamlessly integrated through linear transformation and residual connections, enabling high-precision zero-shot CMI prediction. Extensive experiments on multiple benchmark datasets demonstrate that ZeroStem significantly outperforms existing methods, validating its efficiency and robust generalization in CMI prediction. Case studies further illustrate that ZeroStem offers novel insights into the molecular mechanisms underlying intricate disease-associated networks. Mengmeng Wei, Lei Wang 0121, Zhu-Hong You, Pengwei Hu 0001, Bo-Wei Zhao, Zhi-an Huang |
AAAI | 5 |
| 2026 | Multi-hop graph structural modeling for cancer-related circRNA-miRNA interaction prediction
Mengmeng Wei, Lei Wang 0121, Xiao-Rui Su 0001, Bo-Wei Zhao, Zhu-Hong You |
Pattern Recognit. | 4 |
| 2025 | FlexSecure: Enhancing Flexibility and Security of Shared Terminals in Real-Time Collaborative Programming EnvironmentsabstractReal-time collaborative programming is an emerging technology that supports a team of programmers to concurrently view and edit source code documents, with the benefits of enhancing team productivity and reducing project cost. Shared terminal is one crucial component of real-time collaborative programming environments, which facilitates interactive and instant peer support in debugging scenarios. In this study, we propose a novel approach named FlexSecure to address two major challenges in existing shared terminals. FlexSecure supports unconstrained and flexible shared terminal sessions that allow any collaborator to initiate, and meanwhile, preserves the local security of the initiator by incorporating fine-grained permission control to prevent risky command execution. Prototype implementation has validated the feasibility of FlexSecure, and user evaluation has demonstrated its effectiveness and satisfactory performance. Bicheng Fang, Chengbin Lu, Jinfeng Jiang, Liyou Wang, Bo-Wei Zhao, Hongfei Fan |
SMC | 6 |
| 2025 | Prediction of Budd-Chiari syndrome based on attention mechanisms of high-risk factors in multi-hop graph learning
Mengmeng Wei, Bo-Wei Zhao, Maoheng Zu, Qingqiao Zhang, Zhu-Hong You |
Sci. China Inf. Sci. | 6 |
| 2025 | Regulation-aware graph learning for drug repositioning over heterogeneous biological network
Bo-Wei Zhao, Xiao-Rui Su 0001, Yue Yang 0035, Dongxu Li 0002, Pengwei Hu 0001, Zhu-Hong You, Xin Luo 0001, Lun Hu |
Inf. Sci. | 1 |
| 2025 | Collaborative Framework for circRNA-Disease Associations Prediction Using Dual Variational GraphabstractMany experiments have shown that circular RNA (circRNA) can act as biomarkers for complex diseases and play significant regulatory roles in multiple pathological processes. However, most circRNA-disease associations remain unknown, and discovering these associations through biological experimental approach is expensive and time-consuming. Taking into account the shortcomings of current methods, we introduce a new collaborative framework that utilizes multi-heterogeneous graphs, along with variational graph auto-encoders (VGAE) to predict associations between circRNA and diseases. First, we build multi-similarity networks using various biological attributes of circRNA and diseases, and integrate these similarity networks. Two subnetworks are constructed from association matrix and combined similarity network, which included a circRNA-based network and a disease-based network. We then employ random walk with restart and Singular Value Decomposition, for feature extraction from the similarity matrix. Finally, we use collaborative framework to predict the circRNAdisease association scores based on the two subnetworks. We integrate the two score matrices to obtain a final prediction scoring matrix. Using 5-fold cross-validation on the CircR2Disease dataset, our model achieved an AUC score of 0.9828 and an AUPR score of 0.9820. Additionally, among the top 30 highest-scoring circRNA-disease association pairs, 26 associations have already been validated. Our model shows strong performance and can accurately predict associations between circRNA and diseases, according to experimental results. Changchun Liu 0003, Lei Wang 0121, Bo-Wei Zhao, Mengmeng Wei, Yang Li 0111, Mianshuo Lu, Si-Zhe Liang |
IEEE Trans. Big Data | 3 |
| 2025 | DeepHIV: A Sequence-Based Deep Learning Model for Predicting HIV-1 Protease Cleavage SitesabstractHuman immunodeficiency virus type 1 (HIV-1) is one of the main causative agents of acquired immunodeficiency syndrome (AIDS), and effectively identifying HIV-1 protease cleavage sites (PCSs) is of great importance for the design of new anti-AIDS inhibitors. Computational prediction of HIV-1 PCSs can be used to discover new cleavable substrates, and further facilitates the understanding of substrate specificity. A novel deep learning model, namely DeepHIV, is designed to predict HIV-1 PCSs from substrate sequence information alone. In particular, DeepHIV first applies a convolutional neural network combined with an attention mechanism to capture the rich contextual information of position-specific amino acids in the substrate sequences, thus improving the quality of features learned for substrates. Considering the imbalance observed between cleavable and uncleavable substrates, a biased support vector machine is adopted as the classifier of DeepHIV to complete the prediction task. Experimental results demonstrate that DeepHIV outperforms several state-of-the-art prediction methods across all benchmark datasets and evaluation metrics. Hence, DeepHIV is an accurate and robust tool to predict HIV-1 PCSs. Moreover, the promising predictive performance of DeepHIV also reveals that our deep learning model is capable of fully leveraging the sequence information to effectively learn the latent features of substrates. Dongxu Li 0002, Zhenfeng Li, Bo-Wei Zhao, Xiao-Rui Su 0001, Lun Hu |
IEEE Trans. Comput. Biol. Bioinform. | 3 |
| 2025 | Knowledge Graph Neural Network With Spatial-Aware Capsule for Drug-Drug Interaction PredictionabstractUncovering novel drug-drug interactions (DDIs) plays a pivotal role in advancing drug development and improving clinical treatment. The outstanding effectiveness of graph neural networks (GNNs) has garnered significant interest in the field of DDI prediction. Consequently, there has been a notable surge in the development of network-based computational approaches for predicting DDIs. However, current approaches face limitations in capturing the spatial relationships between neighboring nodes and their higher-level features during the aggregation of neighbor representations. To address this issue, this study introduces a novel model, KGCNN, designed to comprehensively tackle DDI prediction tasks by considering spatial relationships between molecules within the biomedical knowledge graph (BKG). KGCNN is built upon a message-passing GNN framework, consisting of propagation and aggregation. In the context of the BKG, KGCNN governs the propagation of information based on semantic relationships, which determine the flow and exchange of information between different molecules. In contrast to traditional linear aggregators, KGCNN introduces a spatial-aware capsule aggregator, which effectively captures the spatial relationships among neighboring molecules and their higher-level features within the graph structure. The ultimate goal is to leverage these learned drug representations to predict potential DDIs. To evaluate the effectiveness of KGCNN, it undergoes testing on two datasets. Extensive experimental results demonstrate its superiority in DDI predictions and quantified performance. Xiao-Rui Su 0001, Bo-Wei Zhao, Jun Zhang 0003, Pengwei Hu 0001, Zhu-Hong You, Lun Hu |
IEEE J. Biomed. Health Informatics | 2 |
| 2025 | Integrating Transformer and Graph Attention Network for circRNA-miRNA Interaction PredictionabstractCircRNA-miRNA interaction (CMI) plays a crucial role in the gene regulatory network of the cell. Numerous experiments have shown that abnormalities in CMI can impact molecular functions and physiological processes, leading to the occurrence of specific diseases. Current computational models for predicting CMI typically focus on local molecular entity relationships, thereby neglecting inherent molecular attributes and global structural information. To address these limitations, we propose a multi-feature fusion prediction model based on the transformer and graph attention network, named EGATCMI. Specifically, EGATCMI combines the transformer architecture with Word2vec to pre-train the sequence of circRNA and miRNA, capturing their sequence feature representation and sequence similarity. By leveraging the self-attention mechanism, EGATCMI extracts global structural feature from the CMI network. EGATCMI effectively integrates the obtained multi-feature for prediction, achieving AUC values of 0.9106 and 0.9470 on the CMI-9905 and CircBank datasets, respectively, outperforming existing methods. In case studies that the prediction of interactions between three miRNAs that are closely related to diseases and circRNAs, 8 out of 10 pairs were accurately predicted and validated. Extensive experimental results demonstrate the potential of EGATCMI as a reliable tool for candidate screening in biological investigations. Mengmeng Wei, Lei Wang 0121, Bo-Wei Zhao, Xiao-Rui Su 0001, Zhu-Hong You, De-Shuang Huang |
IEEE J. Biomed. Health Informatics | 3 |
| 2024 | Likelihood-based feature representation learning combined with neighborhood information for predicting circRNA-miRNA associationsabstractConnections between circular RNAs (circRNAs) and microRNAs (miRNAs) assume a pivotal position in the onset, evolution, diagnosis and treatment of diseases and tumors. Selecting the most potential circRNA-related miRNAs and taking advantage of them as the biological markers or drug targets could be conducive to dealing with complex human diseases through preventive strategies, diagnostic procedures and therapeutic approaches. Compared to traditional biological experiments, leveraging computational models to integrate diverse biological data in order to infer potential associations proves to be a more efficient and cost-effective approach. This paper developed a model of Convolutional Autoencoder for CircRNA-MiRNA Associations (CA-CMA) prediction. Initially, this model merged the natural language characteristics of the circRNA and miRNA sequence with the features of circRNA-miRNA interactions. Subsequently, it utilized all circRNA-miRNA pairs to construct a molecular association network, which was then fine-tuned by labeled samples to optimize the network parameters. Finally, the prediction outcome is obtained by utilizing the deep neural networks classifier. This model innovatively combines the likelihood objective that preserves the neighborhood through optimization, to learn the continuous feature representation of words and preserve the spatial information of two-dimensional signals. During the process of 5-fold cross-validation, CA-CMA exhibited exceptional performance compared to numerous prior computational approaches, as evidenced by its mean area under the receiver operating characteristic curve of 0.9138 and a minimal SD of 0.0024. Furthermore, recent literature has confirmed the accuracy of 25 out of the top 30 circRNA-miRNA pairs identified with the highest CA-CMA scores during case studies. The results of these experiments highlight the robustness and versatility of our model. Lu-Xiang Guo, Lei Wang 0121, Zhu-Hong You, Meng-Lei Hu, Bo-Wei Zhao, Yang Li 0111 |
Briefings Bioinform. | 6 |
| 2024 | Biolinguistic graph fusion model for circRNA-miRNA association predictionabstractEmerging clinical evidence suggests that sophisticated associations with circular ribonucleic acids (RNAs) (circRNAs) and microRNAs (miRNAs) are a critical regulatory factor of various pathological processes and play a critical role in most intricate human diseases. Nonetheless, the above correlations via wet experiments are error-prone and labor-intensive, and the underlying novel circRNA-miRNA association (CMA) has been validated by numerous existing computational methods that rely only on single correlation data. Considering the inadequacy of existing machine learning models, we propose a new model named BGF-CMAP, which combines the gradient boosting decision tree with natural language processing and graph embedding methods to infer associations between circRNAs and miRNAs. Specifically, BGF-CMAP extracts sequence attribute features and interaction behavior features by Word2vec and two homogeneous graph embedding algorithms, large-scale information network embedding and graph factorization, respectively. Multitudinous comprehensive experimental analysis revealed that BGF-CMAP successfully predicted the complex relationship between circRNAs and miRNAs with an accuracy of 82.90% and an area under receiver operating characteristic of 0.9075. Furthermore, 23 of the top 30 miRNA-associated circRNAs of the studies on data were confirmed in relevant experiences, showing that the BGF-CMAP model is superior to others. BGF-CMAP can serve as a helpful model to provide a scientific theoretical basis for the study of CMA prediction. Lu-Xiang Guo, Lei Wang 0121, Zhu-Hong You, Meng-Lei Hu, Bo-Wei Zhao, Yang Li 0111 |
Briefings Bioinform. | 6 |
| 2024 | BioKG-CMI: a multi-source feature fusion model based on biological knowledge graph for predicting circRNA-miRNA interactions
Mengmeng Wei, Lei Wang 0121, Bo-Wei Zhao, Xiao-Rui Su 0001, Zhu-Hong You |
Sci. China Inf. Sci. | 5 |
| 2024 | Fuzzy-Based Deep Attributed Graph ClusteringabstractAttributed graph (AG) clustering is a fundamental, yet challenging, task for studying underlying network structures. Recently, a variety of graph representation learning models has been proposed to effectively infer the node embeddings, which are then incorporated into conventional clustering techniques to identify meaningful clusters. While these models tend to preserve node proximities, which reflect the similarity between nodes in both structural and attribute dimensions, for representation learning, they generally overlook the crucial dependencies between node embeddings and the resulting clusters. To overcome this problem, we propose a novel fuzzy-based deep AG clustering model, namely FDAGC, which is capable of achieving the task in a purely unsupervised and end-to-end manner without additionally incorporating conventional clustering techniques. In particular, FDAGC first encodes network structures and node attributes into a compact representation with graph convolution. A reconstruction error is then estimated to minimize the information loss during network message-passing. Besides, we utilize a self-monitoring training strategy to optimize node embeddings, thus improving the cluster cohesion by guiding them toward cluster centers. In the training phase, our expectations about resulting clusters are explicitly incorporated into the optimization of FDAGC via the concept of fuzzy clustering, thus leading to more accurate clustering by coupling the dependency between graph representation learning and AG clustering. Extensive experiments have demonstrated the superior performance of FDAGC in terms of several evaluation metrics, such as accuracy, normalized mutual information, F1-score and adjusted rand index, on six real-world AGs with different scales. Yue Yang 0035, Xiao-Rui Su 0001, Bo-Wei Zhao, Pengwei Hu 0001, Jun Zhang 0003, Lun Hu |
IEEE Trans. Fuzzy Syst. | 3 |
| 2024 | Discovering Consensus Regions for Interpretable Identification of RNA N6-Methyladenosine Modification Sites via Graph Contrastive ClusteringabstractAs a pivotal post-transcriptional modification of RNA, N6-methyladenosine (m6A) has a substantial influence on gene expression modulation and cellular fate determination. Although a variety of computational models have been developed to accurately identify potential m6A modification sites, few of them are capable of interpreting the identification process with insights gained from consensus knowledge. To overcome this problem, we propose a deep learning model, namely M6A-DCR, by discovering consensus regions for interpretable identification of m6A modification sites. In particular, M6A-DCR first constructs an instance graph for each RNA sequence by integrating specific positions and types of nucleotides. The discovery of consensus regions is then formulated as a graph clustering problem in light of aggregating all instance graphs. After that, M6A-DCR adopts a motif-aware graph reconstruction optimization process to learn high-quality embeddings of input RNA sequences, thus achieving the identification of m6A modification sites in an end-to-end manner. Experimental results demonstrate the superior performance of M6A-DCR by comparing it with several state-of-the-art identification models. The consideration of consensus regions empowers our model to make interpretable predictions at the motif level. The analysis of cross validation through different species and tissues further verifies the consistency between the identification results of M6A-DCR and the evolutionary relationships among species. Bo-Wei Zhao, Xiao-Rui Su 0001, Yue Yang 0035, Pengwei Hu 0001, Xi Zhou 0007, Lun Hu |
IEEE J. Biomed. Health Informatics | 2 |
| 2024 | Motif-Aware miRNA-Disease Association Prediction via Hierarchical Attention NetworkabstractAs post-transcriptional regulators of gene expression, micro-ribonucleic acids (miRNAs) are regarded as potential biomarkers for a variety of diseases. Hence, the prediction of miRNA-disease associations (MDAs) is of great significance for an in-depth understanding of disease pathogenesis and progression. Existing prediction models are mainly concentrated on incorporating different sources of biological information to perform the MDA prediction task while failing to consider the fully potential utility of MDA network information at the motif-level. To overcome this problem, we propose a novel motif-aware MDA prediction model, namely MotifMDA, by fusing a variety of high- and low-order structural information. In particular, we first design several motifs of interest considering their ability to characterize how miRNAs are associated with diseases through different network structural patterns. Then, MotifMDA adopts a two-layer hierarchical attention to identify novel MDAs. Specifically, the first attention layer learns high-order motif preferences based on their occurrences in the given MDA network, while the second one learns the final embeddings of miRNAs and diseases through coupling high- and low-order preferences. Experimental results on two benchmark datasets have demonstrated the superior performance of MotifMDA over several state-of-the-art prediction models. This strongly indicates that accurate MDA prediction can be achieved by relying solely on MDA network information. Furthermore, our case studies indicate that the incorporation of motif-level structure information allows MotifMDA to discover novel MDAs from different perspectives. Bo-Wei Zhao, Xiao-Rui Su 0001, Yue Yang 0035, Pengwei Hu 0001, Zhu-Hong You, Lun Hu |
IEEE J. Biomed. Health Informatics | 1 |
| 2023 | Learning RNA sequence patterns to interpretably identify m6A modification sitesabstractN6-methyladenosine (m6A) regulates RNA post-transcriptional modification and translation processes, thereby regulating gene expression and cell fate. Hence, accurate identification of potential m6A modification sites is a key step to further reveal their biological functions and understand multiple biological processes such as gene regulation and epigenetic variation. Many computational methods have been developed to address this challenge. However, fewer studies have focused on an interpretable process of m6A modification site identification. Here, we propose an interpretable end-to-end predictor, called M6AInter, which learns the RNA sequence patterns related to modification sites through contrastive learning frameworks to achieve accurate identification of m6A modification sites. Specifically, M6AInter first utilizes chaos game representation theory and one-hot encoding to initialize the position and type information of nucleotides, respectively. On this basis, M6AInter extracts the position and type correlations shared by RNA sequences, and predicts the common sequence patterns by utilizing a graph contrastive clustering framework. These motifs and patterns are involved in describing the associations between RNA sequences and obtaining their low-dimensional representations. Finally, through a designed bias fusion block, these representations are combined with the frequency information of nucleotides to realize the identification of m6A modification sites. Extensive experimental results show that our model can accurately identify modified RNA sequences and can adaptively locate sequential regions associated with m6A modification sites on RNA sequences. Importantly, by exploring the role of these patterns in the identification tasks, M6AInter provides interpretable predictions and analysis at the sequence level. Bo-Wei Zhao, Xiao-Rui Su 0001, Yue Yang 0035, Pengwei Hu 0001, Lun Hu |
BIBM | 2 |
| 2023 | A Novel Graph Representation Learning Model for Drug Repositioning Using Graph Transition Probability Matrix Over Heterogenous Information Networks
Dongxu Li 0002, Bo-Wei Zhao, Xiao-Rui Su 0001, Zhu-Hong You, Pengwei Hu 0001, Lun Hu |
ICIC (3) | 3 |
| 2023 | A Deep Learning Approach Incorporating Data Missing Mechanism in Predicting Acute Kidney Injury in ICU
Zhengbo Zhang, Lei Zha, Fengcong, Xiao-Rui Su 0001, Bo-Wei Zhao, Lun Hu, Pengwei Hu 0001 |
ICIC (3) | 7 |
| 2023 | Multi-level Subgraph Representation Learning for Drug-Disease Association Prediction Over Heterogeneous Biological Information Network
Bo-Wei Zhao, Xiao-Rui Su 0001, Yue Yang 0035, Dongxu Li 0002, Pengwei Hu 0001, Zhu-Hong You, Lun Hu |
ICIC (3) | 1 |
| 2023 | Incorporating higher order network structures to improve miRNA-disease association prediction based on functional modularityabstractAs microRNAs (miRNAs) are involved in many essential biological processes, their abnormal expressions can serve as biomarkers and prognostic indicators to prevent the development of complex diseases, thus providing accurate early detection and prognostic evaluation. Although a number of computational methods have been proposed to predict miRNA-disease associations (MDAs) for further experimental verification, their performance is limited primarily by the inadequacy of exploiting lower order patterns characterizing known MDAs to identify missing ones from MDA networks. Hence, in this work, we present a novel prediction model, namely HiSCMDA, by incorporating higher order network structures for improved performance of MDA prediction. To this end, HiSCMDA first integrates miRNA similarity network, disease similarity network and MDA network to preserve the advantages of all these networks. After that, it identifies overlapping functional modules from the integrated network by predefining several higher order connectivity patterns of interest. Last, a path-based scoring function is designed to infer potential MDAs based on network paths across related functional modules. HiSCMDA yields the best performance across all datasets and evaluation metrics in the cross-validation and independent validation experiments. Furthermore, in the case studies, 49 and 50 out of the top 50 miRNAs, respectively, predicted for colon neoplasms and lung neoplasms have been validated by well-established databases. Experimental results show that rich higher order organizational structures exposed in the MDA network gain new insight into the MDA prediction based on higher order connectivity patterns. Yue Yang 0035, Xiao-Rui Su 0001, Bo-Wei Zhao, Shengwu Xiong 0001, Lun Hu |
Briefings Bioinform. | 4 |
| 2023 | iGRLDTI: an improved graph representation learning method for predicting drug-target interactions over heterogeneous biological information networkabstractMOTIVATION: The task of predicting drug-target interactions (DTIs) plays a significant role in facilitating the development of novel drug discovery. Compared with laboratory-based approaches, computational methods proposed for DTI prediction are preferred due to their high-efficiency and low-cost advantages. Recently, much attention has been attracted to apply different graph neural network (GNN) models to discover underlying DTIs from heterogeneous biological information network (HBIN). Although GNN-based prediction methods achieve better performance, they are prone to encounter the over-smoothing simulation when learning the latent representations of drugs and targets with their rich neighborhood information in HBIN, and thereby reduce the discriminative ability in DTI prediction. RESULTS: In this work, an improved graph representation learning method, namely iGRLDTI, is proposed to address the above issue by better capturing more discriminative representations of drugs and targets in a latent feature space. Specifically, iGRLDTI first constructs an HBIN by integrating the biological knowledge of drugs and targets with their interactions. After that, it adopts a node-dependent local smoothing strategy to adaptively decide the propagation depth of each biomolecule in HBIN, thus significantly alleviating over-smoothing by enhancing the discriminative ability of feature representations of drugs and targets. Finally, a Gradient Boosting Decision Tree classifier is used by iGRLDTI to predict novel DTIs. Experimental results demonstrate that iGRLDTI yields better performance that several state-of-the-art computational methods on the benchmark dataset. Besides, our case study indicates that iGRLDTI can successfully identify novel DTIs with more distinguishable features of drugs and targets. AVAILABILITY AND IMPLEMENTATION: Python codes and dataset are available at https://github.com/stevejobws/iGRLDTI/. Bo-Wei Zhao, Xiao-Rui Su 0001, Pengwei Hu 0001, Zhu-Hong You, Lun Hu |
Bioinform. | 1 |
| 2023 | Biocaiv: an integrative webserver for motif-based clustering analysis and interactive visualization of biological networksabstractBACKGROUND: As an important task in bioinformatics, clustering analysis plays a critical role in understanding the functional mechanisms of many complex biological systems, which can be modeled as biological networks. The purpose of clustering analysis in biological networks is to identify functional modules of interest, but there is a lack of online clustering tools that visualize biological networks and provide in-depth biological analysis for discovered clusters. RESULTS: Here we present BioCAIV, a novel webserver dedicated to maximize its accessibility and applicability on the clustering analysis of biological networks. This, together with its user-friendly interface, assists biological researchers to perform an accurate clustering analysis for biological networks and identify functionally significant modules for further assessment. CONCLUSIONS: BioCAIV is an efficient clustering analysis webserver designed for a variety of biological networks. BioCAIV is freely available without registration requirements at http://bioinformatics.tianshanzw.cn:8888/BioCAIV/ . Dongxu Li 0002, Bo-Wei Zhao, Xiao-Rui Su 0001, Jun Zhang 0003, Pengwei Hu 0001, Lun Hu |
BMC Bioinform. | 3 |
| 2023 | PDA-PRGCN: identification of Piwi-interacting RNA-disease associations through subgraph projection and residual scaling-based feature augmentationabstractBACKGROUND: Emerging evidences show that Piwi-interacting RNAs (piRNAs) play a pivotal role in numerous complex human diseases. Identifying potential piRNA-disease associations (PDAs) is crucial for understanding disease pathogenesis at molecular level. Compared to the biological wet experiments, the computational methods provide a cost-effective strategy. However, few computational methods have been developed so far. RESULTS: Here, we proposed an end-to-end model, referred to as PDA-PRGCN (PDA prediction using subgraph Projection and Residual scaling-based feature augmentation through Graph Convolutional Network). Specifically, starting with the known piRNA-disease associations represented as a graph, we applied subgraph projection to construct piRNA-piRNA and disease-disease subgraphs for the first time, followed by a residual scaling-based feature augmentation algorithm for node initial representation. Then, we adopted graph convolutional network (GCN) to learn and identify potential PDAs as a link prediction task on the constructed heterogeneous graph. Comprehensive experiments, including the performance comparison of individual components in PDA-PRGCN, indicated the significant improvement of integrating subgraph projection, node feature augmentation and dual-loss mechanism into GCN for PDA prediction. Compared with state-of-the-art approaches, PDA-PRGCN gave more accurate and robust predictions. Finally, the case studies further corroborated that PDA-PRGCN can reliably detect PDAs. CONCLUSION: PDA-PRGCN provides a powerful method for PDA prediction, which can also serve as a screening tool for studies of complex diseases. Ping Zhang 0027, Weicheng Sun, Dengguo Wei, Jinsheng Xu, Zhu-Hong You, Bo-Wei Zhao, Li Li 0057 |
BMC Bioinform. | 7 |
| 2023 | Biomedical Knowledge Graph Embedding With Capsule Network for Multi-Label Drug-Drug Interaction PredictionabstractDrug-drug interaction (DDI) plays an important role in drug development and administration. Most of existing network-based computation models regard the DDI prediction as a binary classification problem and generate negative DDI samples randomly, but the binary classification is not in line with the real problem since there are dozens of types of DDI and randomly generating negative samples may introduce false-negative samples since the non-observed facts can be either false or just missing. To address the above limitations, we propose a new framework called KG2ECapsule that explicitly models the multi-relational DDI data based on biomedical knowledge graphs in an end-to-end fashion. It first generates high-quality negative samples based on the average number of tail entities and head entities for each relation to reduce false-negative samples to some extent. KG2ECapsule then refines the representations of entities by recursively propagating the embeddings from the attention-based receptive fields of entities. Empirical results on three biomedical knowledge graphs of different scales show that KG2ECapsule outperforms the state-of-the-art methods consistently in multi-label DDI prediction task and further studies verify the efficacy of both probability-based sampling strategy and non-linear transformation for modeling multi-relational data. Xiao-Rui Su 0001, Zhu-Hong You, De-Shuang Huang, Lei Wang 0121, Leon Wong, Bo-Wei Zhao |
IEEE Trans. Knowl. Data Eng. | 7 |
| 2022 | Cost and Care Insight: An Interactive and Scalable Hierarchical Learning System for Identifying Cost Saving Opportunities
David Koepke, Bibo Hao, Jing Mei, Xu Min, Rachna Gupta, Rajashree Joshi, Fiona McNaughton, Bo-Wei Zhao, Lun Hu, Pengwei Hu 0001 |
ICIC (1) | 10 |
| 2022 | Predicting Drug-Disease Associations via Meta-path Representation Learning based on Heterogeneous Information Net works
Menglong Zhang, Bo-Wei Zhao, Lun Hu, Zhu-Hong You |
ICIC (2) | 2 |
| 2022 | MRLDTI: A Meta-path-Based Representation Learning Model for Drug-Target Interaction Prediction
Bo-Wei Zhao, Lun Hu, Pengwei Hu 0001, Zhu-Hong You, Xiao-Rui Su 0001, Dongxu Li 0002, Ping Zhang 0027 |
ICIC (2) | 1 |
| 2022 | A Novel Fuzzy-Based MOPSO Algorithm for Identifying Clusters From Complex NetworksabstractMany complicated systems can be modeled as complex networks, and a variety of graph clustering algorithms have been proposed to perform accurate clustering analysis for better understanding system behaviors. However, most of them suffer the disadvantage of slow convergence. In this paper, we incorporate multi-objective particle swarm optimization (MOPSO) into a well-established fuzzy clustering algorithm, i.e., FCAN, and propose an improved Fuzzy-based Graph Clustering Algorithm, namely IMFCAN, which retains all the benefits gained with FCAN while achieving significantly fast convergence rate. Specially, IMFCAN enhances the ability of handling the imbalance observed in the distribution of fuzzy membership of nodes by introducing an instance-frequency-weighted regularization (IR) scheme. After that, IMFCAN develops an effective solution to reach a consensus optimization among them by balancing global exploration and local exploitation abilities of particles. Experimental results on four practical datasets demonstrate that IMFCAN performs better than several state-of-the-art clustering algorithm in terms of accuracy and convergence. Hence, IMFCAN is a promising algorithm for addressing the clustering analysis of complex networks. Yue Yang 0035, Xiao-Rui Su 0001, Bo-Wei Zhao, Lun Hu |
ICTAI | 3 |
| 2022 | A novel circRNA-miRNA association prediction model based on structural deep neural network embeddingabstractA large amount of clinical evidence began to mount, showing that circular ribonucleic acids (RNAs; circRNAs) perform a very important function in complex diseases by participating in transcription and translation regulation of microRNA (miRNA) target genes. However, with strict high-throughput techniques based on traditional biological experiments and the conditions and environment, the association between circRNA and miRNA can be discovered to be labor-intensive, expensive, time-consuming, and inefficient. In this paper, we proposed a novel computational model based on Word2vec, Structural Deep Network Embedding (SDNE), Convolutional Neural Network and Deep Neural Network, which predicts the potential circRNA-miRNA associations, called Word2vec, SDNE, Convolutional Neural Network and Deep Neural Network (WSCD). Specifically, the WSCD model extracts attribute feature and behaviour feature by word embedding and graph embedding algorithm, respectively, and ultimately feed them into a feature fusion model constructed by combining Convolutional Neural Network and Deep Neural Network to deduce potential circRNA-miRNA interactions. The proposed method is proved on dataset and obtained a prediction accuracy and an area under the receiver operating characteristic curve of 81.61% and 0.8898, respectively, which is shown to have much higher accuracy than the state-of-the-art models and classifier models in prediction. In addition, 23 miRNA-related circular RNAs (circRNAs) from the top 30 were confirmed in relevant experiences. In these works, all results represent that WSCD would be a helpful supplementary reliable method for predicting potential miRNA-circRNA associations compared to wet laboratory experiments. Lu-Xiang Guo, Zhu-Hong You, Lei Wang 0121, Bo-Wei Zhao, Zhong-Hao Ren, Jie Pan 0007 |
Briefings Bioinform. | 5 |
| 2022 | A deep learning method for repurposing antiviral drugs against new viruses via multi-view nonnegative matrix factorization and its application to SARS-CoV-2abstractThe outbreak of COVID-19 caused by SARS-coronavirus (CoV)-2 has made millions of deaths since 2019. Although a variety of computational methods have been proposed to repurpose drugs for treating SARS-CoV-2 infections, it is still a challenging task for new viruses, as there are no verified virus-drug associations (VDAs) between them and existing drugs. To efficiently solve the cold-start problem posed by new viruses, a novel constrained multi-view nonnegative matrix factorization (CMNMF) model is designed by jointly utilizing multiple sources of biological information. With the CMNMF model, the similarities of drugs and viruses can be preserved from their own perspectives when they are projected onto a unified latent feature space. Based on the CMNMF model, we propose a deep learning method, namely VDA-DLCMNMF, for repurposing drugs against new viruses. VDA-DLCMNMF first initializes the node representations of drugs and viruses with their corresponding latent feature vectors to avoid a random initialization and then applies graph convolutional network to optimize their representations. Given an arbitrary drug, its probability of being associated with a new virus is computed according to their representations. To evaluate the performance of VDA-DLCMNMF, we have conducted a series of experiments on three VDA datasets created for SARS-CoV-2. Experimental results demonstrate that the promising prediction accuracy of VDA-DLCMNMF. Moreover, incorporating the CMNMF model into deep learning gains new insight into the drug repurposing for SARS-CoV-2, as the results of molecular docking experiments reveal that four antiviral drugs identified by VDA-DLCMNMF have the potential ability to treat SARS-CoV-2 infections. Xiao-Rui Su 0001, Lun Hu, Zhu-Hong You, Pengwei Hu 0001, Lei Wang 0121, Bo-Wei Zhao |
Briefings Bioinform. | 6 |
| 2022 | Attention-based Knowledge Graph Representation Learning for Predicting Drug-drug InteractionsabstractDrug-drug interactions (DDIs) are known as the main cause of life-threatening adverse events, and their identification is a key task in drug development. Existing computational algorithms mainly solve this problem by using advanced representation learning techniques. Though effective, few of them are capable of performing their tasks on biomedical knowledge graphs (KGs) that provide more detailed information about drug attributes and drug-related triple facts. In this work, an attention-based KG representation learning framework, namely DDKG, is proposed to fully utilize the information of KGs for improved performance of DDI prediction. In particular, DDKG first initializes the representations of drugs with their embeddings derived from drug attributes with an encoder-decoder layer, and then learns the representations of drugs by recursively propagating and aggregating first-order neighboring information along top-ranked network paths determined by neighboring node embeddings and triple facts. Last, DDKG estimates the probability of being interacting for pairwise drugs with their representations in an end-to-end manner. To evaluate the effectiveness of DDKG, extensive experiments have been conducted on two practical datasets with different sizes, and the results demonstrate that DDKG is superior to state-of-the-art algorithms on the DDI prediction task in terms of different evaluation metrics across all datasets. Xiao-Rui Su 0001, Lun Hu, Zhu-Hong You, Pengwei Hu 0001, Bo-Wei Zhao |
Briefings Bioinform. | 5 |
| 2022 | A machine learning framework based on multi-source feature fusion for circRNA-disease association predictionabstractCircular RNAs (circRNAs) are involved in the regulatory mechanisms of multiple complex diseases, and the identification of their associations is critical to the diagnosis and treatment of diseases. In recent years, many computational methods have been designed to predict circRNA-disease associations. However, most of the existing methods rely on single correlation data. Here, we propose a machine learning framework for circRNA-disease association prediction, called MLCDA, which effectively fuses multiple sources of heterogeneous information including circRNA sequences and disease ontology. Comprehensive evaluation in the gold standard dataset showed that MLCDA can successfully capture the complex relationships between circRNAs and diseases and accurately predict their potential associations. In addition, the results of case studies on real data show that MLCDA significantly outperforms other existing methods. MLCDA can serve as a useful tool for circRNA-disease association prediction, providing mechanistic insights for disease research and thus facilitating the progress of disease treatment. Lei Wang 0121, Leon Wong, Zhengwei Li 0001, Xiao-Rui Su 0001, Bo-Wei Zhao, Zhu-Hong You |
Briefings Bioinform. | 6 |
| 2022 | iGRLCDA: identifying circRNA-disease association based on graph representation learningabstractWhile the technologies of ribonucleic acid-sequence (RNA-seq) and transcript assembly analysis have continued to improve, a novel topology of RNA transcript was uncovered in the last decade and is called circular RNA (circRNA). Recently, researchers have revealed that they compete with messenger RNA (mRNA) and long noncoding for combining with microRNA in gene regulation. Therefore, circRNA was assumed to be associated with complex disease and discovering the relationship between them would contribute to medical research. However, the work of identifying the association between circRNA and disease in vitro takes a long time and usually without direction. During these years, more and more associations were verified by experiments. Hence, we proposed a computational method named identifying circRNA-disease association based on graph representation learning (iGRLCDA) for the prediction of the potential association of circRNA and disease, which utilized a deep learning model of graph convolution network (GCN) and graph factorization (GF). In detail, iGRLCDA first derived the hidden feature of known associations between circRNA and disease using the Gaussian interaction profile (GIP) kernel combined with disease semantic information to form a numeric descriptor. After that, it further used the deep learning model of GCN and GF to extract hidden features from the descriptor. Finally, the random forest classifier is introduced to identify the potential circRNA-disease association. The five-fold cross-validation of iGRLCDA shows strong competitiveness in comparison with other excellent prediction models at the gold standard data and achieved an average area under the receiver operating characteristic curve of 0.9289 and an area under the precision-recall curve of 0.9377. On reviewing the prediction results from the relevant literature, 22 of the top 30 predicted circRNA-disease associations were noted in recent published papers. These exceptional results make us believe that iGRLCDA can provide reliable circRNA-disease associations for medical research and reduce the blindness of wet-lab experiments. Lei Wang 0121, Zhu-Hong You, Lun Hu, Bo-Wei Zhao, Zhengwei Li 0001, Yang-Ming Li |
Briefings Bioinform. | 5 |
| 2022 | HINGRL: predicting drug-disease associations with graph representation learning on heterogeneous information networksabstractIdentifying new indications for drugs plays an essential role at many phases of drug research and development. Computational methods are regarded as an effective way to associate drugs with new indications. However, most of them complete their tasks by constructing a variety of heterogeneous networks without considering the biological knowledge of drugs and diseases, which are believed to be useful for improving the accuracy of drug repositioning. To this end, a novel heterogeneous information network (HIN) based model, namely HINGRL, is proposed to precisely identify new indications for drugs based on graph representation learning techniques. More specifically, HINGRL first constructs a HIN by integrating drug-disease, drug-protein and protein-disease biological networks with the biological knowledge of drugs and diseases. Then, different representation strategies are applied to learn the features of nodes in the HIN from the topological and biological perspectives. Finally, HINGRL adopts a Random Forest classifier to predict unknown drug-disease associations based on the integrated features of drugs and diseases obtained in the previous step. Experimental results demonstrate that HINGRL achieves the best performance on two real datasets when compared with state-of-the-art models. Besides, our case studies indicate that the simultaneous consideration of network topology and biological knowledge of drugs and diseases allows HINGRL to precisely predict drug-disease associations from a more comprehensive perspective. The promising performance of HINGRL also reveals that the utilization of rich heterogeneous information provides an alternative view for HINGRL to identify novel drug-disease associations especially for new diseases. Bo-Wei Zhao, Lun Hu, Zhu-Hong You, Lei Wang 0121, Xiao-Rui Su 0001 |
Briefings Bioinform. | 1 |
| 2022 | A geometric deep learning framework for drug repositioning over heterogeneous information networksabstractDrug repositioning (DR) is a promising strategy to discover new indicators of approved drugs with artificial intelligence techniques, thus improving traditional drug discovery and development. However, most of DR computational methods fall short of taking into account the non-Euclidean nature of biomedical network data. To overcome this problem, a deep learning framework, namely DDAGDL, is proposed to predict drug-drug associations (DDAs) by using geometric deep learning (GDL) over heterogeneous information network (HIN). Incorporating complex biological information into the topological structure of HIN, DDAGDL effectively learns the smoothed representations of drugs and diseases with an attention mechanism. Experiment results demonstrate the superior performance of DDAGDL on three real-world datasets under 10-fold cross-validation when compared with state-of-the-art DR methods in terms of several evaluation metrics. Our case studies and molecular docking experiments indicate that DDAGDL is a promising DR tool that gains new insights into exploiting the geometric prior knowledge for improved efficacy. Bo-Wei Zhao, Xiao-Rui Su 0001, Pengwei Hu 0001, Yu-Peng Ma, Xi Zhou 0007, Lun Hu |
Briefings Bioinform. | 1 |
| 2022 | Multi-view heterogeneous molecular network representation learning for protein-protein interaction predictionabstractBACKGROUND: Protein-protein interaction (PPI) plays an important role in regulating cells and signals. Despite the ongoing efforts of the bioassay group, continued incomplete data limits our ability to understand the molecular roots of human disease. Therefore, it is urgent to develop a computational method to predict PPIs from the perspective of molecular system. METHODS: In this paper, a highly efficient computational model, MTV-PPI, is proposed for PPI prediction based on a heterogeneous molecular network by learning inter-view protein sequences and intra-view interactions between molecules simultaneously. On the one hand, the inter-view feature is extracted from the protein sequence by k-mer method. On the other hand, we use a popular embedding method LINE to encode the heterogeneous molecular network to obtain the intra-view feature. Thus, the protein representation used in MTV-PPI is constructed by the aggregation of its inter-view feature and intra-view feature. Finally, random forest is integrated to predict potential PPIs. RESULTS: To prove the effectiveness of MTV-PPI, we conduct extensive experiments on a collected heterogeneous molecular network with the accuracy of 86.55%, sensitivity of 82.49%, precision of 89.79%, AUC of 0.9301 and AUPR of 0.9308. Further comparison experiments are performed with various protein representations and classifiers to indicate the effectiveness of MTV-PPI in predicting PPIs based on a complex network. CONCLUSION: The achieved experimental results illustrate that MTV-PPI is a promising tool for PPI prediction, which may provide a new perspective for the future interactions prediction researches based on heterogeneous molecular network. Xiao-Rui Su 0001, Lun Hu, Zhu-Hong You, Pengwei Hu 0001, Bo-Wei Zhao |
BMC Bioinform. | 5 |
| 2022 | RLFDDA: a meta-path based graph representation learning model for drug-disease association predictionabstractBACKGROUND: Drug repositioning is a very important task that provides critical information for exploring the potential efficacy of drugs. Yet developing computational models that can effectively predict drug-disease associations (DDAs) is still a challenging task. Previous studies suggest that the accuracy of DDA prediction can be improved by integrating different types of biological features. But how to conduct an effective integration remains a challenging problem for accurately discovering new indications for approved drugs. METHODS: In this paper, we propose a novel meta-path based graph representation learning model, namely RLFDDA, to predict potential DDAs on heterogeneous biological networks. RLFDDA first calculates drug-drug similarities and disease-disease similarities as the intrinsic biological features of drugs and diseases. A heterogeneous network is then constructed by integrating DDAs, disease-protein associations and drug-protein associations. With such a network, RLFDDA adopts a meta-path random walk model to learn the latent representations of drugs and diseases, which are concatenated to construct joint representations of drug-disease associations. As the last step, we employ the random forest classifier to predict potential DDAs with their joint representations. RESULTS: To demonstrate the effectiveness of RLFDDA, we have conducted a series of experiments on two benchmark datasets by following a ten-fold cross-validation scheme. The results show that RLFDDA yields the best performance in terms of AUC and F1-score when compared with several state-of-the-art DDAs prediction models. We have also conducted a case study on two common diseases, i.e., paclitaxel and lung tumors, and found that 7 out of top-10 diseases and 8 out of top-10 drugs have already been validated for paclitaxel and lung tumors respectively with literature evidence. Hence, the promising performance of RLFDDA may provide a new perspective for novel DDAs discovery over heterogeneous networks. Menglong Zhang, Bo-Wei Zhao, Xiao-Rui Su 0001, Yue Yang 0035, Lun Hu |
BMC Bioinform. | 2 |
| 2022 | NSECDA: Natural Semantic Enhancement for CircRNA-Disease Association PredictionabstractIncreasing evidence suggest that circRNA, as one of the most promising emerging biomarkers, has a very close relationship with diseases. Exploring the relationship between circRNA and diseases can provide novel perspective for diseases diagnosis and pathogenesis. The existing circRNA-disease association (CDA) prediction models, however, generally treat the data attributes equally, do not pay special attention to the attributes with more significant influence, and do not make full use of the correlation and symbiosis between attributes to dig into the latent semantic information of the data. Therefore, in response to the above problems, this paper proposes a natural semantic enhancement method NSECDA to predict CDA. In practical terms, we first recognize the circRNA sequence as a biological language, and analyze its natural semantic properties through the natural language understanding theory; then integrate it with disease attributes, circRNA and disease Gaussian Interaction Profile (GIP) kernel attributes, and use Graph Attention Network (GAT) to focus on the influential attributes, so as to mine the deeply hidden features; finally, the Rotation Forest (RoF) classifier was used to accurately determine CDA. In the gold standard data set CircR2Disease, NSECDA achieved 92.49% accuracy with 0.9225 AUC score. In comparison with the non-natural semantic enhancement model and other classifier models, NSECDA also shows competitive performance. Additionally, 25 of the CDA pairs with unknown associations in the top 30 prediction scores of NSECDA have been proven by newly reported studies. These achievements suggest that NSECDA is an effective model to predict CDA, which can provide credible candidate for subsequent wet experiments, thus significantly reducing the scope of investigations. Lei Wang 0121, Leon Wong, Zhu-Hong You, De-Shuang Huang, Xiao-Rui Su 0001, Bo-Wei Zhao |
IEEE J. Biomed. Health Informatics | 6 |
| 2021 | Predicting miRNA-Disease Associations via a New MeSH Headings Representation of Diseases and eXtreme Gradient Boosting
Zhu-Hong You, Lei Wang 0121, Leon Wong, Xiao-Rui Su 0001, Bo-Wei Zhao |
ICIC (3) | 6 |
| 2021 | Detection of Drug-Drug Interactions Through Knowledge Graph Integrating Multi-attention with Capsule Network
Xiao-Rui Su 0001, Zhu-Hong You, Bo-Wei Zhao |
ICIC (3) | 4 |
| 2021 | A Multi-graph Deep Learning Model for Predicting Drug-Disease Associations
Bo-Wei Zhao, Zhu-Hong You, Lun Hu, Leon Wong, Ping Zhang 0027 |
ICIC (3) | 1 |
| 2021 | Predicting Large-scale Protein-protein Interactions by Extracting Coevolutionary Patterns with MapReduce ParadigmabstractProtein-protein interactions are of great significance for us to understand the functional mechanisms of proteins. With the rapid development of high-throughput genomic technology, the amount of protein-protein interaction data has become so big that most of existing prediction algorithms are no longer applicable. To address this problem, we develop a distributed framework by reimplementing one of state-of-the-art algorithms, i.e., CoFex, by using MapReduce. In particular, we adopt a novel tree-based data structure to reduce the heavy memory consumption cased by the huge sequence information of proteins. After that, the procedure of CoFex is modified by following the paradigm of MapReduce such that the prediction task can be completed in a distributed manner, thus fulfilling the demanding requirements of large-scale protein-protein interaction prediction. A series of experiments have been conducted to evaluate the performance of the proposed distributed framework in terms of both efficiency and effectiveness. Experimental results demonstrate that the proposed framework can considerably improve the efficiency of CoFex by achieving more than two-orders-of-magnitude improvement in computational efficiency while retaining a comparable level of accuracy. Lun Hu, Bo-Wei Zhao, Shicheng Yang, Xin Luo 0001, MengChu Zhou |
SMC | 2 |
| 2020 | A Novel Computational Method for Predicting LncRNA-Disease Associations from Heterogeneous Information Network with SDNE Embedding Model
Ping Zhang 0027, Bo-Wei Zhao, Leon Wong, Zhu-Hong You, Zhen-Hao Guo |
ICIC (2) | 2 |
| 2020 | Predicting LncRNA-miRNA Interactions via Network Embedding with Integrated Structure and Attribute Information
Bo-Wei Zhao, Ping Zhang 0027, Zhu-Hong You, Ji-Ren Zhou, Xiao Li 0007 |
ICIC (2) | 1 |