EDBT 2026 Demo / reviewers in the wild / expert
Yue Yang 0035
dblp:54/6179-35
· DBLP profile ↗
20ranked-venue papers
5as first author
20since 2021 · last 2026
0000-0001-7729-595XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 13 · 1 first-author · 13 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LLM-DDI: Leveraging Large Language Models for Drug-Drug Interaction Prediction on Biomedical Knowledge GraphabstractDrug-drug interaction (DDI) refers to the interaction relationships between drugs. Discovering new DDIs is crucial for advancing drug development and enhancing clinical treatments. Given the significant progress achieved through graph neural networks (GNNs), network-based models have become a prevalent approach for tackling this challenge. However, current network-based approaches are incapable of seamlessly integrating a wide range of information. Motivated by this discovery, we propose a novel model, namely LLM-DDI, which aims to comprehensively tackle DDI prediction tasks by integrating various information of molecules in the BKG. LLM-DDI initially incorporates the generative pre-trained transformer (GPT) model to generate embeddings for each molecule within the biomedical knowledge graph (BKG). These embeddings encompass diverse types of information pertaining to each molecule. Subsequently, LLM-DDI utilizes a message-passing GNN framework to enhance the learning of molecular representations with the embeddings derived from GPT as input. LLM-DDI governs the propagation of information within the BKG by semantic relationships. These semantic relationships determine how information flows and is exchanged between different entities in the BKG. Finally, LLM-DDI leverages the learned drug representations to predict potential DDIs. Experiments show the effectiveness of LLM-DDI, as it achieves the best performance on two real-world datasets, providing valuable guidance for drug development and clinical treatment. Dongxu Li 0002, Yue Yang 0035, Ziwen Cui, Hengchuang Yin, Pengwei Hu 0001, Lun Hu |
IEEE J. Biomed. Health Informatics | 2 |
| 2026 | Multi-View Contrastive Learning for Drug-Drug Interaction Event PredictionabstractDrug-drug interactions (DDIs) represent a critical challenge in pharmacology, often leading to adverse effects and compromised therapeutic efficacy. Accurate prediction of DDI events, which involve not only identifying interacting drug pairs but also characterizing the specific nature and context of their interactions, is essential for drug safety and personalized medicine. In this study, we propose a novel Multi-view Contrastive Learning framework, namely MCL-DDI, for DDI Event Prediction by leveraging multi-view representations of drugs to enhance predictive performance. MCL-DDI integrates molecular structures and network features, capturing complementary information about drug properties and interactions. By employing contrastive learning, we align and unify drug representations across these diverse views, enabling the framework to distinguish complex interaction patterns. Extensive experiments on benchmark datasets demonstrate that MCL-DDI outperforms state-of-the-art methods in terms of predictive accuracy. Furthermore, case studies highlight the model's ability to identify clinically relevant DDIs, offering practical insights for drug development and risk assessment. Our work establishes a robust and accurate paradigm for DDI event prediction, paving the way for safer and more effective pharmacological interventions. Dongxu Li 0002, Feifan Zhao, Yue Yang 0035, Ziwen Cui, Pengwei Hu 0001, Lun Hu |
IEEE J. Biomed. Health Informatics | 3 |
| 2025 | Improving Cancer Gene Identification via Mixture-of-Experts-Based Graph Representation LearningabstractAccurately identifying cancer driver genes is crucial for understanding tumorigenesis and advancing precision oncology. However, integrating multi-omics data within complex biological networks remains challenging, particularly when it comes to capturing diverse structural information and leveraging the distinct signals from different omics modalities. While graphbased methods have demonstrated high accuracy in cancer gene identification, they might overlook the heterogeneity between omics features. To address this limitation, we propose CGI-MoE, a Mixture-of-Experts-inspired graph representation learning framework that incorporates omics-feature-specific expert modules based on graph transformers, together with adaptive gating and subgraph aggregation mechanisms. CGI-MoE extracts both local and global structural encodings for each node by sampling multiple subgraphs, enabling the model to capture comprehensive and robust network features. Each expert module focuses on a specific omics modality, and their outputs are fused by a lightweight gating network that dynamically weighs their contributions. When evaluated on both homogeneous and heterogeneous benchmark datasets, CGI-MoE achieves state-of-the-art performance in terms of accuracy, AUC, and AUPR, consistently surpassing existing methods. Ablation studies further highlight the critical roles of each component in achieving robust performance. Using the trained models, CGI-MoE predicted 46 novel cancer gene candidates from all unlabeled genes, demonstrating its potential for novel discovery and for deepening our understanding of cancer development. The code is available at https://github.com/moomight/CGI-MoE. Ying Chang, Yue Yang 0035, Dongxu Li 0002, Ziwen Cui, Hengchuang Yin, Pengwei Hu 0001, Lun Hu |
BIBM | 2 |
| 2025 | Missed Abortion Prediction Using Dual Network on Large-Scale Clinical Data with Multiple FeaturesabstractMissed abortion is a common obstetric complication with serious consequences for maternal health and fertility. Accurately predicting the risk of missed abortion occurrence is important for timely intervention and prevention. Given the large amount and complexity of clinical data, as well as the nonlinear inter-correlations that exist among multiple features, it is difficult for general machine learning methods to effectively mine the potential crucial patterns that dominate the occurrence of missed abortions. To address this problem, this paper proposes a dual network framework, namely DNMAP, which integrates a local feature extraction module to capture correlations among clinical indicators with a feature integration module to aggregate these features for missed abortion prediction. In addition, DNMAP extends the prediction problem to a multi-classification task to better explore the association between missed abortion and recurrent miscarriage where two or more miscarriages have occurred. Different from previous studies, this paper is the first to investigate the problem of missed abortion in this way, with the largest practical clinical miscarriage dataset to our knowledge. A series of experiments have been conducted and the results demonstrate that the proposed DNMAP significantly outperforms various types of models in the task of missed abortion prediction. DNMAP can not only predict the risk of missed abortion a simple and effective manner, but also provide more refined risk stratification, which suggests novel insights and methods to further improve maternal health and reproductive outcomes. Jun Zhang 0003, Xiaoli Bo, Yue Yang 0035, Xiwen Yang, Yating Yang |
BIBM | 4 |
| 2025 | Regulation-aware graph learning for drug repositioning over heterogeneous biological network
Bo-Wei Zhao, Xiao-Rui Su 0001, Yue Yang 0035, Dongxu Li 0002, Pengwei Hu 0001, Zhu-Hong You, Xin Luo 0001, Lun Hu |
Inf. Sci. | 3 |
| 2025 | A bijective inference network for interpretable identification of RNA N6-methyladenosine modification sites
Yue Yang 0035, Dongxu Li 0002, Xiao-Rui Su 0001, Zhi Zeng 0001, Pengwei Hu 0001, Lun Hu |
Pattern Recognit. | 2 |
| 2025 | Link-Based Attributed Graph Clustering via Approximate Generative Bayesian LearningabstractTo understand the mechanisms of complex systems, attributed graphs (AGs) are recognized as a valuable model by their capability of describing nontrivial topological structures and rich node contents, and their emergence raises new challenges on the task of graph clustering. Although a variety of computational algorithms have been proposed to perform accurate clustering analysis on AGs, most of them are incapable of inferring the cluster labels of nodes through links, thus falling short of explaining node behaviors on how to formulate overlapping clusters. Moreover, the vast amount of links considerably decreases the computation efficiency if they are explicitly taken into account for AG clustering. To overcome this problem, we present a novel variational Bayesian learning model, which avoids generating a complete AG by only simulating the generative process of its skeleton with the prior knowledge on the cluster labels of links. When addressing the inference problem, we develop an efficient algorithm, namely, LCAAG, for determining the optimal cluster labels of nodes by estimating local community structures of links. The convergence of LCAAG has been proved theoretically. Compared with several state-of-the-art algorithms, LCAAG has demonstrated its promising performance in terms of both accuracy and scalability on five different scaled benchmark datasets. The source code and datasets are available at https://github.com/shallowdreamoon/LCAAG.git. Yue Yang 0035, Lun Hu, Dongxu Li 0002, Pengwei Hu 0001, Xin Luo 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2025 | FMvPCI: A Multiview Fusion Neural Network for Identifying Protein Complex via Fuzzy ClusteringabstractProtein complexes play a crucial role in regulating various biological processes that govern cell activities. Numerous computational algorithms have been proposed to identify protein complexes from protein-protein interaction (PPI) networks. However, many of these algorithms face limitations in effectively leveraging multiview biological information of proteins, restricting their ability to capture the intricate characteristics of protein complexes in PPI networks. While deep learning-based algorithms have significantly advanced the identification of protein complexes, they often integrate graph representation learning techniques into traditional clustering algorithms without explicitly capturing the dependency between protein embeddings and resulting complexes. To address these issues, we present a multiview fusion neural network, named FMvPCI, for protein complex identification via fuzzy clustering. In FMvPCI, we introduce a novel multiview graph convolution encoder to effectively manipulate and fuse the biological information of proteins from different perspectives. Subsequently, the optimization of FMvPCI incorporates our expectations about protein complexes through the concept of fuzzy clustering. This approach unifies the embeddings of proteins and their cluster memberships within a coherent framework. Leveraging a heuristic search strategy, FMvPCI can discover overlapping protein complexes based on the cluster memberships of proteins. A series of experiments on five different PPI networks collected from two species have been conducted to evaluate the performance of FMvPCI by comparing it with state-of-the-art identification algorithms, and the results demonstrate the superior performance of FMvPCI by significantly improving the identification accuracy for protein complexes. Yue Yang 0035, Lun Hu, Dongxu Li 0002, Pengwei Hu 0001, Xin Luo 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2024 | A Multi-view Nested Contrastive Learning Framework for Predicting Drug-Drug Interaction EventsabstractExploring drug-drug interactions (DDIs) is crucial for avoiding unknown physicochemical incompatibilities between coadministered drugs. While most studies concentrate on detecting the presence or absence of DDIs, they often overlook the diversity of DDI event types that can significantly enhance drug research and guide scientific drug use. To address this limitation, we propose MNCLDDI, a multi-view nested contrastive learning model designed for the precise prediction of DDI events. MN-CLDDI begins by employing a relational graph convolutional network to capture the various explicit relationships between drugs within a multi-relational DDI graph. This is followed by a transformer framework combined with a convolutional neural network (CNN) to learn the biological features of drugs from their Smiles information. The model then integrates these two feature types into a novel multi-view nested contrastive learning framework, thereby improving the expressiveness of drug embeddings from multiple biological perspectives. Experimental results on two real-world datasets demonstrate that MNCLDDI outperforms state-of-the-art models in predicting DDI events. Moreover, our case studies reveal that considering the multi-view features of drugs simultaneously enables MNCLDDI to predict DDI events with greater accuracy and from a more comprehensive perspective, offering valuable insights into the study of DDI events. Dongxu Li 0002, Yue Yang 0035, Pengwei Hu 0001, Lun Hu |
BIBM | 2 |
| 2024 | Knowledge-guided Protein Complex Identification with Fuzzy-based Graph Representation LearningabstractProtein complexes are essential in regulating various cellular processes. A number of computational algorithms have been developed to identify protein complexes from protein-protein interaction (PPI) networks, but they are limited in their ability to effectively leverage diverse biological knowledge of proteins. Additionally, while deep learning-based algorithms perform well in identifying protein complexes, they fail to explicitly capture the dependency between protein embeddings and resulting complexes. To address these challenges, this paper proposes a knowledge-guided protein complex identification algorithm with fuzzy-based graph representation learning, named KPCI-FGRL. In particular, a fuzzy-based graph representation learning framework is developed by KPCI-FGRL to manipulate and fuse network structure with multi-view biological knowledge of proteins. During the training phase of KPCI-FGRL, besides employing self-supervised loss to improve the cohesion of the complexes, we also specifically incorporate the expectation about protein complexes based on fuzzy clustering concept, and thus the dependency between protein embeddings and complexes can be coupled. Furthermore, KPCI-FGRL is capable of achieving the identification of overlapping protein complexes through a heuristic search strategy upon fuzzy memberships of proteins. Extensive experimental results on four different PPI networks collected from two species demonstrate that KPCI-FGRL significantly outperforms several state-of-the-art protein complex identification algorithms. Yue Yang 0035, Dongxu Li 0002, Pengwei Hu 0001, Lun Hu |
BIBM | 1 |
| 2024 | Fuzzy-Based Deep Attributed Graph ClusteringabstractAttributed graph (AG) clustering is a fundamental, yet challenging, task for studying underlying network structures. Recently, a variety of graph representation learning models has been proposed to effectively infer the node embeddings, which are then incorporated into conventional clustering techniques to identify meaningful clusters. While these models tend to preserve node proximities, which reflect the similarity between nodes in both structural and attribute dimensions, for representation learning, they generally overlook the crucial dependencies between node embeddings and the resulting clusters. To overcome this problem, we propose a novel fuzzy-based deep AG clustering model, namely FDAGC, which is capable of achieving the task in a purely unsupervised and end-to-end manner without additionally incorporating conventional clustering techniques. In particular, FDAGC first encodes network structures and node attributes into a compact representation with graph convolution. A reconstruction error is then estimated to minimize the information loss during network message-passing. Besides, we utilize a self-monitoring training strategy to optimize node embeddings, thus improving the cluster cohesion by guiding them toward cluster centers. In the training phase, our expectations about resulting clusters are explicitly incorporated into the optimization of FDAGC via the concept of fuzzy clustering, thus leading to more accurate clustering by coupling the dependency between graph representation learning and AG clustering. Extensive experiments have demonstrated the superior performance of FDAGC in terms of several evaluation metrics, such as accuracy, normalized mutual information, F1-score and adjusted rand index, on six real-world AGs with different scales. Yue Yang 0035, Xiao-Rui Su 0001, Bo-Wei Zhao, Pengwei Hu 0001, Jun Zhang 0003, Lun Hu |
IEEE Trans. Fuzzy Syst. | 1 |
| 2024 | Discovering Consensus Regions for Interpretable Identification of RNA N6-Methyladenosine Modification Sites via Graph Contrastive ClusteringabstractAs a pivotal post-transcriptional modification of RNA, N6-methyladenosine (m6A) has a substantial influence on gene expression modulation and cellular fate determination. Although a variety of computational models have been developed to accurately identify potential m6A modification sites, few of them are capable of interpreting the identification process with insights gained from consensus knowledge. To overcome this problem, we propose a deep learning model, namely M6A-DCR, by discovering consensus regions for interpretable identification of m6A modification sites. In particular, M6A-DCR first constructs an instance graph for each RNA sequence by integrating specific positions and types of nucleotides. The discovery of consensus regions is then formulated as a graph clustering problem in light of aggregating all instance graphs. After that, M6A-DCR adopts a motif-aware graph reconstruction optimization process to learn high-quality embeddings of input RNA sequences, thus achieving the identification of m6A modification sites in an end-to-end manner. Experimental results demonstrate the superior performance of M6A-DCR by comparing it with several state-of-the-art identification models. The consideration of consensus regions empowers our model to make interpretable predictions at the motif level. The analysis of cross validation through different species and tissues further verifies the consistency between the identification results of M6A-DCR and the evolutionary relationships among species. Bo-Wei Zhao, Xiao-Rui Su 0001, Yue Yang 0035, Pengwei Hu 0001, Xi Zhou 0007, Lun Hu |
IEEE J. Biomed. Health Informatics | 4 |
| 2024 | Motif-Aware miRNA-Disease Association Prediction via Hierarchical Attention NetworkabstractAs post-transcriptional regulators of gene expression, micro-ribonucleic acids (miRNAs) are regarded as potential biomarkers for a variety of diseases. Hence, the prediction of miRNA-disease associations (MDAs) is of great significance for an in-depth understanding of disease pathogenesis and progression. Existing prediction models are mainly concentrated on incorporating different sources of biological information to perform the MDA prediction task while failing to consider the fully potential utility of MDA network information at the motif-level. To overcome this problem, we propose a novel motif-aware MDA prediction model, namely MotifMDA, by fusing a variety of high- and low-order structural information. In particular, we first design several motifs of interest considering their ability to characterize how miRNAs are associated with diseases through different network structural patterns. Then, MotifMDA adopts a two-layer hierarchical attention to identify novel MDAs. Specifically, the first attention layer learns high-order motif preferences based on their occurrences in the given MDA network, while the second one learns the final embeddings of miRNAs and diseases through coupling high- and low-order preferences. Experimental results on two benchmark datasets have demonstrated the superior performance of MotifMDA over several state-of-the-art prediction models. This strongly indicates that accurate MDA prediction can be achieved by relying solely on MDA network information. Furthermore, our case studies indicate that the incorporation of motif-level structure information allows MotifMDA to discover novel MDAs from different perspectives. Bo-Wei Zhao, Xiao-Rui Su 0001, Yue Yang 0035, Pengwei Hu 0001, Zhu-Hong You, Lun Hu |
IEEE J. Biomed. Health Informatics | 4 |
| 2023 | Learning RNA sequence patterns to interpretably identify m6A modification sitesabstractN6-methyladenosine (m6A) regulates RNA post-transcriptional modification and translation processes, thereby regulating gene expression and cell fate. Hence, accurate identification of potential m6A modification sites is a key step to further reveal their biological functions and understand multiple biological processes such as gene regulation and epigenetic variation. Many computational methods have been developed to address this challenge. However, fewer studies have focused on an interpretable process of m6A modification site identification. Here, we propose an interpretable end-to-end predictor, called M6AInter, which learns the RNA sequence patterns related to modification sites through contrastive learning frameworks to achieve accurate identification of m6A modification sites. Specifically, M6AInter first utilizes chaos game representation theory and one-hot encoding to initialize the position and type information of nucleotides, respectively. On this basis, M6AInter extracts the position and type correlations shared by RNA sequences, and predicts the common sequence patterns by utilizing a graph contrastive clustering framework. These motifs and patterns are involved in describing the associations between RNA sequences and obtaining their low-dimensional representations. Finally, through a designed bias fusion block, these representations are combined with the frequency information of nucleotides to realize the identification of m6A modification sites. Extensive experimental results show that our model can accurately identify modified RNA sequences and can adaptively locate sequential regions associated with m6A modification sites on RNA sequences. Importantly, by exploring the role of these patterns in the identification tasks, M6AInter provides interpretable predictions and analysis at the sequence level. Bo-Wei Zhao, Xiao-Rui Su 0001, Yue Yang 0035, Pengwei Hu 0001, Lun Hu |
BIBM | 4 |
| 2023 | Multi-level Subgraph Representation Learning for Drug-Disease Association Prediction Over Heterogeneous Biological Information Network
Bo-Wei Zhao, Xiao-Rui Su 0001, Yue Yang 0035, Dongxu Li 0002, Pengwei Hu 0001, Zhu-Hong You, Lun Hu |
ICIC (3) | 3 |
| 2023 | Incorporating higher order network structures to improve miRNA-disease association prediction based on functional modularityabstractAs microRNAs (miRNAs) are involved in many essential biological processes, their abnormal expressions can serve as biomarkers and prognostic indicators to prevent the development of complex diseases, thus providing accurate early detection and prognostic evaluation. Although a number of computational methods have been proposed to predict miRNA-disease associations (MDAs) for further experimental verification, their performance is limited primarily by the inadequacy of exploiting lower order patterns characterizing known MDAs to identify missing ones from MDA networks. Hence, in this work, we present a novel prediction model, namely HiSCMDA, by incorporating higher order network structures for improved performance of MDA prediction. To this end, HiSCMDA first integrates miRNA similarity network, disease similarity network and MDA network to preserve the advantages of all these networks. After that, it identifies overlapping functional modules from the integrated network by predefining several higher order connectivity patterns of interest. Last, a path-based scoring function is designed to infer potential MDAs based on network paths across related functional modules. HiSCMDA yields the best performance across all datasets and evaluation metrics in the cross-validation and independent validation experiments. Furthermore, in the case studies, 49 and 50 out of the top 50 miRNAs, respectively, predicted for colon neoplasms and lung neoplasms have been validated by well-established databases. Experimental results show that rich higher order organizational structures exposed in the MDA network gain new insight into the MDA prediction based on higher order connectivity patterns. Yue Yang 0035, Xiao-Rui Su 0001, Bo-Wei Zhao, Shengwu Xiong 0001, Lun Hu |
Briefings Bioinform. | 2 |
| 2023 | PPISB: A Novel Network-Based Algorithm of Predicting Protein-Protein Interactions With Mixed Membership Stochastic BlockmodelabstractProtein-protein interactions (PPIs) play an essential role for most of biological processes in cells. Many computational algorithms have thus been proposed to predict PPIs. However, most of them heavily rest on the biological information of proteins while ignoring the latent structural features of proteins presented in a PPI network. In this paper, we propose an efficient network-based prediction algorithm, namely PPISB, based on a mixed membership stochastic blockmodel. By simulating the generative process of a PPI network, PPISB is able to capture the latent community structures. The inference procedure adopted by PPISB further optimizes the membership distributions of proteins over different complexes. After that, a distance measure is designed to compute the similarity between two proteins in terms of their likelihoods of being in the same complex, thus verifying whether they interact with each other or not. To evaluate the performance of PPISB, a series of extensive experiments have been conducted with five PPI networks collected from different species and the results demonstrate that PPISB has a promising performance when applied to predict PPIs in terms of several evaluation metrics. Hence, we reason that PPISB is preferred over state-of-the-art network-based prediction algorithms especially for predicting potential PPIs. Wen Yang 0019, Yue Yang 0035, Jun Zhang 0003, Lusheng Wang 0001, Lun Hu |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2023 | FCAN-MOPSO: An Improved Fuzzy-Based Graph Clustering Algorithm for Complex Networks With Multiobjective Particle Swarm OptimizationabstractPerforming an accurate clustering analysis is of great significance for us to understand the behavior of complex networks, and a variety of graph clustering algorithms have, thus, been proposed to do so by taking into account network topology and node attributes. Among them, fuzzy clustering algorithm for complex networks (FCAN) is an established fuzzy clustering algorithm that optimizes the memberships of nodes based on dense structures and content relevance. This article proposes an improved fuzzy-based graph clustering algorithm, namely FCAN-multi objective particle swarm optimization (MOPSO) that retains all the benefits associated with FCAN while achieving significantly increased convergence rate using multiobjective particle swarm optimization (MOPSO). To do so, FCAN-MOPSO first modifies the original optimization model of FCAN by adopting an instance-frequency-weighted regularization, which enhances the ability of FCAN-MOPSO to handle the imbalance observed in the distribution of fuzzy memberships of nodes. After that, FCAN-MOPSO decomposes its optimization problem into a set of suboptimization problems. Following the MOPSO framework, FCAN-MOPSO develops an effective solution to reach a consensus optimization among them by balancing the global exploration and local exploitation abilities of particles. A theoretical analysis is provided to prove the global convergence of FCAN-MOPSO. Extensive experiments have been conducted to evaluate the performance of FCAN-MOPSO on five real-world complex networks with different scale, and experimental results demonstrate that when compared with state-of-the-art clustering algorithms, FCAN-MOPSO achieves a better accuracy performance with improved convergence. Hence, FCAN-MOPSO is a promising graph clustering algorithm to precisely and efficiently discover clusters in complex networks. Lun Hu, Yue Yang 0035, Zehai Tang, Xin Luo 0001 |
IEEE Trans. Fuzzy Syst. | 2 |
| 2022 | A Novel Fuzzy-Based MOPSO Algorithm for Identifying Clusters From Complex NetworksabstractMany complicated systems can be modeled as complex networks, and a variety of graph clustering algorithms have been proposed to perform accurate clustering analysis for better understanding system behaviors. However, most of them suffer the disadvantage of slow convergence. In this paper, we incorporate multi-objective particle swarm optimization (MOPSO) into a well-established fuzzy clustering algorithm, i.e., FCAN, and propose an improved Fuzzy-based Graph Clustering Algorithm, namely IMFCAN, which retains all the benefits gained with FCAN while achieving significantly fast convergence rate. Specially, IMFCAN enhances the ability of handling the imbalance observed in the distribution of fuzzy membership of nodes by introducing an instance-frequency-weighted regularization (IR) scheme. After that, IMFCAN develops an effective solution to reach a consensus optimization among them by balancing global exploration and local exploitation abilities of particles. Experimental results on four practical datasets demonstrate that IMFCAN performs better than several state-of-the-art clustering algorithm in terms of accuracy and convergence. Hence, IMFCAN is a promising algorithm for addressing the clustering analysis of complex networks. Yue Yang 0035, Xiao-Rui Su 0001, Bo-Wei Zhao, Lun Hu |
ICTAI | 1 |
| 2022 | RLFDDA: a meta-path based graph representation learning model for drug-disease association predictionabstractBACKGROUND: Drug repositioning is a very important task that provides critical information for exploring the potential efficacy of drugs. Yet developing computational models that can effectively predict drug-disease associations (DDAs) is still a challenging task. Previous studies suggest that the accuracy of DDA prediction can be improved by integrating different types of biological features. But how to conduct an effective integration remains a challenging problem for accurately discovering new indications for approved drugs. METHODS: In this paper, we propose a novel meta-path based graph representation learning model, namely RLFDDA, to predict potential DDAs on heterogeneous biological networks. RLFDDA first calculates drug-drug similarities and disease-disease similarities as the intrinsic biological features of drugs and diseases. A heterogeneous network is then constructed by integrating DDAs, disease-protein associations and drug-protein associations. With such a network, RLFDDA adopts a meta-path random walk model to learn the latent representations of drugs and diseases, which are concatenated to construct joint representations of drug-disease associations. As the last step, we employ the random forest classifier to predict potential DDAs with their joint representations. RESULTS: To demonstrate the effectiveness of RLFDDA, we have conducted a series of experiments on two benchmark datasets by following a ten-fold cross-validation scheme. The results show that RLFDDA yields the best performance in terms of AUC and F1-score when compared with several state-of-the-art DDAs prediction models. We have also conducted a case study on two common diseases, i.e., paclitaxel and lung tumors, and found that 7 out of top-10 diseases and 8 out of top-10 drugs have already been validated for paclitaxel and lung tumors respectively with literature evidence. Hence, the promising performance of RLFDDA may provide a new perspective for novel DDAs discovery over heterogeneous networks. Menglong Zhang, Bo-Wei Zhao, Xiao-Rui Su 0001, Yue Yang 0035, Lun Hu |
BMC Bioinform. | 5 |