VLDB 2026 Research / reviewers in the wild / expert
Xiao-Rui Su 0001
dblp:276/8163-1 · also Xiaorui Su 0001
· DBLP profile ↗
39ranked-venue papers
13as first author
36since 2021 · last 2026
0000-0001-5468-6085ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 30 · 9 first-author · 27 since 2021Artificial intelligence and machine learning · 7 · 3 first-author · 7 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-hop graph structural modeling for cancer-related circRNA-miRNA interaction prediction
Mengmeng Wei, Lei Wang 0121, Xiao-Rui Su 0001, Bo-Wei Zhao, Zhu-Hong You |
Pattern Recognit. | 3 |
| 2025 | KGARevion: An AI Agent for Knowledge-Intensive Biomedical QAabstractBiomedical reasoning integrates structured, codified knowledge with tacit, experience-driven insights. Depending on the context, quantity, and nature of available evidence, researchers and clinicians use diverse strategies, including rule-based, prototype-based, and case-based reasoning. Effective medical AI models must handle this complexity while ensuring reliability and adaptability. We introduce KGARevion, a knowledge graph-based agent that answers knowledge-intensive questions. Upon receiving a query, KGARevion generates relevant triplets by leveraging the latent knowledge embedded in a large language model. It then verifies these triplets against a grounded knowledge graph, filtering out errors and retaining only accurate, contextually relevant information for the final answer. This multi-step process strengthens reasoning, adapts to different models of medical inference, and outperforms retrieval-augmented generation-based approaches that lack effective verification mechanisms. Evaluations on medical QA benchmarks show that KGARevion improves accuracy by over 5.2% over 15 models in handling complex medical queries. To further assess its effectiveness, we curated three new medical QA datasets with varying levels of semantic complexity, where KGARevion improved accuracy by 10.4%. The agent integrates with different LLMs and biomedical knowledge graphs for broad applicability across knowledge-intensive tasks. We evaluated KGARevion on AfriMed-QA, a newly introduced dataset focused on African healthcare, demonstrating its strong zero-shot generalization to underrepresented medical contexts. Xiao-Rui Su 0001, Yibo Wang 0001, Shanghua Gao, Xiaolong Liu 0012, Valentina Giunchiglia, Djork-Arné Clevert, Marinka Zitnik |
ICLR | 1 |
| 2025 | Multimodal Medical Code TokenizerabstractFoundation models trained on patient electronic health records (EHRs) require tokenizing medical data into sequences of discrete vocabulary items. Existing tokenizers treat medical codes from EHRs as isolated textual tokens. However, each medical code is defined by its textual description, its position in ontological hierarchies, and its relationships to other codes, such as disease co-occurrences and drug-treatment associations. Medical vocabularies contain more than 600,000 codes with critical information for clinical reasoning. We introduce MedTok, a multimodal medical code tokenizer that uses the text descriptions and relational context of codes. MedTok processes text using a language model encoder and encodes the relational structure with a graph encoder. It then quantizes both modalities into a unified token space, preserving modality-specific and cross-modality information. We integrate MedTok into five EHR models and evaluate it on operational and clinical tasks across in-patient and out-patient datasets, including outcome prediction, diagnosis classification, drug recommendation, and risk stratification. Swapping standard EHR tokenizers with MedTok improves AUPRC across all EHR models, by 4.10% on MIMIC-III, 4.78% on MIMIC-IV, and 11.32% on EHRShot, with the largest gains in drug recommendation. Beyond EHR modeling, we demonstrate using MedTok tokenizer with medical QA systems. Our results demonstrate the potential of MedTok as a unified tokenizer for medical codes, improving tokenization for medical foundation models. Xiao-Rui Su 0001, Shvat Messica, Yepeng Huang, Ruth Johnson, Lukas Fesser, Shanghua Gao, Faryad Sahneh, Marinka Zitnik |
ICML | 1 |
| 2025 | Regulation-aware graph learning for drug repositioning over heterogeneous biological network
Bo-Wei Zhao, Xiao-Rui Su 0001, Yue Yang 0035, Dongxu Li 0002, Pengwei Hu 0001, Zhu-Hong You, Xin Luo 0001, Lun Hu |
Inf. Sci. | 2 |
| 2025 | A bijective inference network for interpretable identification of RNA N6-methyladenosine modification sites
Yue Yang 0035, Dongxu Li 0002, Xiao-Rui Su 0001, Zhi Zeng 0001, Pengwei Hu 0001, Lun Hu |
Pattern Recognit. | 4 |
| 2025 | DeepHIV: A Sequence-Based Deep Learning Model for Predicting HIV-1 Protease Cleavage SitesabstractHuman immunodeficiency virus type 1 (HIV-1) is one of the main causative agents of acquired immunodeficiency syndrome (AIDS), and effectively identifying HIV-1 protease cleavage sites (PCSs) is of great importance for the design of new anti-AIDS inhibitors. Computational prediction of HIV-1 PCSs can be used to discover new cleavable substrates, and further facilitates the understanding of substrate specificity. A novel deep learning model, namely DeepHIV, is designed to predict HIV-1 PCSs from substrate sequence information alone. In particular, DeepHIV first applies a convolutional neural network combined with an attention mechanism to capture the rich contextual information of position-specific amino acids in the substrate sequences, thus improving the quality of features learned for substrates. Considering the imbalance observed between cleavable and uncleavable substrates, a biased support vector machine is adopted as the classifier of DeepHIV to complete the prediction task. Experimental results demonstrate that DeepHIV outperforms several state-of-the-art prediction methods across all benchmark datasets and evaluation metrics. Hence, DeepHIV is an accurate and robust tool to predict HIV-1 PCSs. Moreover, the promising predictive performance of DeepHIV also reveals that our deep learning model is capable of fully leveraging the sequence information to effectively learn the latent features of substrates. Dongxu Li 0002, Zhenfeng Li, Bo-Wei Zhao, Xiao-Rui Su 0001, Lun Hu |
IEEE Trans. Comput. Biol. Bioinform. | 4 |
| 2025 | Knowledge Graph Neural Network With Spatial-Aware Capsule for Drug-Drug Interaction PredictionabstractUncovering novel drug-drug interactions (DDIs) plays a pivotal role in advancing drug development and improving clinical treatment. The outstanding effectiveness of graph neural networks (GNNs) has garnered significant interest in the field of DDI prediction. Consequently, there has been a notable surge in the development of network-based computational approaches for predicting DDIs. However, current approaches face limitations in capturing the spatial relationships between neighboring nodes and their higher-level features during the aggregation of neighbor representations. To address this issue, this study introduces a novel model, KGCNN, designed to comprehensively tackle DDI prediction tasks by considering spatial relationships between molecules within the biomedical knowledge graph (BKG). KGCNN is built upon a message-passing GNN framework, consisting of propagation and aggregation. In the context of the BKG, KGCNN governs the propagation of information based on semantic relationships, which determine the flow and exchange of information between different molecules. In contrast to traditional linear aggregators, KGCNN introduces a spatial-aware capsule aggregator, which effectively captures the spatial relationships among neighboring molecules and their higher-level features within the graph structure. The ultimate goal is to leverage these learned drug representations to predict potential DDIs. To evaluate the effectiveness of KGCNN, it undergoes testing on two datasets. Extensive experimental results demonstrate its superiority in DDI predictions and quantified performance. Xiao-Rui Su 0001, Bo-Wei Zhao, Jun Zhang 0003, Pengwei Hu 0001, Zhu-Hong You, Lun Hu |
IEEE J. Biomed. Health Informatics | 1 |
| 2025 | Integrating Transformer and Graph Attention Network for circRNA-miRNA Interaction PredictionabstractCircRNA-miRNA interaction (CMI) plays a crucial role in the gene regulatory network of the cell. Numerous experiments have shown that abnormalities in CMI can impact molecular functions and physiological processes, leading to the occurrence of specific diseases. Current computational models for predicting CMI typically focus on local molecular entity relationships, thereby neglecting inherent molecular attributes and global structural information. To address these limitations, we propose a multi-feature fusion prediction model based on the transformer and graph attention network, named EGATCMI. Specifically, EGATCMI combines the transformer architecture with Word2vec to pre-train the sequence of circRNA and miRNA, capturing their sequence feature representation and sequence similarity. By leveraging the self-attention mechanism, EGATCMI extracts global structural feature from the CMI network. EGATCMI effectively integrates the obtained multi-feature for prediction, achieving AUC values of 0.9106 and 0.9470 on the CMI-9905 and CircBank datasets, respectively, outperforming existing methods. In case studies that the prediction of interactions between three miRNAs that are closely related to diseases and circRNAs, 8 out of 10 pairs were accurately predicted and validated. Extensive experimental results demonstrate the potential of EGATCMI as a reliable tool for candidate screening in biological investigations. Mengmeng Wei, Lei Wang 0121, Bo-Wei Zhao, Xiao-Rui Su 0001, Zhu-Hong You, De-Shuang Huang |
IEEE J. Biomed. Health Informatics | 4 |
| 2024 | Dual-Channel Learning Framework for Drug-Drug Interaction Prediction via Relation-Aware Heterogeneous Graph TransformerabstractIdentifying novel drug-drug interactions (DDIs) is a crucial task in pharmacology, as the interference between pharmacological substances can pose serious medical risks. In recent years, several network-based techniques have emerged for predicting DDIs. However, they primarily focus on local structures within DDI-related networks, often overlooking the significance of indirect connections between pairwise drug nodes from a global perspective. Additionally, effectively handling heterogeneous information present in both biomedical knowledge graphs and drug molecular graphs remains a challenge for improved performance of DDI prediction. To address these limitations, we propose a Transformer-based relatIon-aware Graph rEpresentation leaRning framework (TIGER) for DDI prediction. TIGER leverages the Transformer architecture to effectively exploit the structure of heterogeneous graph, which allows it direct learning of long dependencies and high-order structures. Furthermore, TIGER incorporates a relation-aware self-attention mechanism, capturing a diverse range of semantic relations that exist between pairs of nodes in heterogeneous graph. In addition to these advancements, TIGER enhances predictive accuracy by modeling DDI prediction task using a dual-channel network, where drug molecular graph and biomedical knowledge graph are fed into two respective channels. By incorporating embeddings obtained at graph and node levels, TIGER can benefit from structural properties of drugs as well as rich contextual information provided by biomedical knowledge graph. Extensive experiments conducted on three real-world datasets demonstrate the effectiveness of TIGER in DDI prediction. Furthermore, case studies highlight its ability to provide a deeper understanding of underlying mechanisms of DDIs. Xiao-Rui Su 0001, Pengwei Hu 0001, Zhu-Hong You, Philip S. Yu, Lun Hu |
AAAI | 1 |
| 2024 | BioKG-CMI: a multi-source feature fusion model based on biological knowledge graph for predicting circRNA-miRNA interactions
Mengmeng Wei, Lei Wang 0121, Bo-Wei Zhao, Xiao-Rui Su 0001, Zhu-Hong You |
Sci. China Inf. Sci. | 6 |
| 2024 | Fuzzy-Based Deep Attributed Graph ClusteringabstractAttributed graph (AG) clustering is a fundamental, yet challenging, task for studying underlying network structures. Recently, a variety of graph representation learning models has been proposed to effectively infer the node embeddings, which are then incorporated into conventional clustering techniques to identify meaningful clusters. While these models tend to preserve node proximities, which reflect the similarity between nodes in both structural and attribute dimensions, for representation learning, they generally overlook the crucial dependencies between node embeddings and the resulting clusters. To overcome this problem, we propose a novel fuzzy-based deep AG clustering model, namely FDAGC, which is capable of achieving the task in a purely unsupervised and end-to-end manner without additionally incorporating conventional clustering techniques. In particular, FDAGC first encodes network structures and node attributes into a compact representation with graph convolution. A reconstruction error is then estimated to minimize the information loss during network message-passing. Besides, we utilize a self-monitoring training strategy to optimize node embeddings, thus improving the cluster cohesion by guiding them toward cluster centers. In the training phase, our expectations about resulting clusters are explicitly incorporated into the optimization of FDAGC via the concept of fuzzy clustering, thus leading to more accurate clustering by coupling the dependency between graph representation learning and AG clustering. Extensive experiments have demonstrated the superior performance of FDAGC in terms of several evaluation metrics, such as accuracy, normalized mutual information, F1-score and adjusted rand index, on six real-world AGs with different scales. Yue Yang 0035, Xiao-Rui Su 0001, Bo-Wei Zhao, Pengwei Hu 0001, Jun Zhang 0003, Lun Hu |
IEEE Trans. Fuzzy Syst. | 2 |
| 2024 | Discovering Consensus Regions for Interpretable Identification of RNA N6-Methyladenosine Modification Sites via Graph Contrastive ClusteringabstractAs a pivotal post-transcriptional modification of RNA, N6-methyladenosine (m6A) has a substantial influence on gene expression modulation and cellular fate determination. Although a variety of computational models have been developed to accurately identify potential m6A modification sites, few of them are capable of interpreting the identification process with insights gained from consensus knowledge. To overcome this problem, we propose a deep learning model, namely M6A-DCR, by discovering consensus regions for interpretable identification of m6A modification sites. In particular, M6A-DCR first constructs an instance graph for each RNA sequence by integrating specific positions and types of nucleotides. The discovery of consensus regions is then formulated as a graph clustering problem in light of aggregating all instance graphs. After that, M6A-DCR adopts a motif-aware graph reconstruction optimization process to learn high-quality embeddings of input RNA sequences, thus achieving the identification of m6A modification sites in an end-to-end manner. Experimental results demonstrate the superior performance of M6A-DCR by comparing it with several state-of-the-art identification models. The consideration of consensus regions empowers our model to make interpretable predictions at the motif level. The analysis of cross validation through different species and tissues further verifies the consistency between the identification results of M6A-DCR and the evolutionary relationships among species. Bo-Wei Zhao, Xiao-Rui Su 0001, Yue Yang 0035, Pengwei Hu 0001, Xi Zhou 0007, Lun Hu |
IEEE J. Biomed. Health Informatics | 3 |
| 2024 | Motif-Aware miRNA-Disease Association Prediction via Hierarchical Attention NetworkabstractAs post-transcriptional regulators of gene expression, micro-ribonucleic acids (miRNAs) are regarded as potential biomarkers for a variety of diseases. Hence, the prediction of miRNA-disease associations (MDAs) is of great significance for an in-depth understanding of disease pathogenesis and progression. Existing prediction models are mainly concentrated on incorporating different sources of biological information to perform the MDA prediction task while failing to consider the fully potential utility of MDA network information at the motif-level. To overcome this problem, we propose a novel motif-aware MDA prediction model, namely MotifMDA, by fusing a variety of high- and low-order structural information. In particular, we first design several motifs of interest considering their ability to characterize how miRNAs are associated with diseases through different network structural patterns. Then, MotifMDA adopts a two-layer hierarchical attention to identify novel MDAs. Specifically, the first attention layer learns high-order motif preferences based on their occurrences in the given MDA network, while the second one learns the final embeddings of miRNAs and diseases through coupling high- and low-order preferences. Experimental results on two benchmark datasets have demonstrated the superior performance of MotifMDA over several state-of-the-art prediction models. This strongly indicates that accurate MDA prediction can be achieved by relying solely on MDA network information. Furthermore, our case studies indicate that the incorporation of motif-level structure information allows MotifMDA to discover novel MDAs from different perspectives. Bo-Wei Zhao, Xiao-Rui Su 0001, Yue Yang 0035, Pengwei Hu 0001, Zhu-Hong You, Lun Hu |
IEEE J. Biomed. Health Informatics | 3 |
| 2023 | Learning RNA sequence patterns to interpretably identify m6A modification sitesabstractN6-methyladenosine (m6A) regulates RNA post-transcriptional modification and translation processes, thereby regulating gene expression and cell fate. Hence, accurate identification of potential m6A modification sites is a key step to further reveal their biological functions and understand multiple biological processes such as gene regulation and epigenetic variation. Many computational methods have been developed to address this challenge. However, fewer studies have focused on an interpretable process of m6A modification site identification. Here, we propose an interpretable end-to-end predictor, called M6AInter, which learns the RNA sequence patterns related to modification sites through contrastive learning frameworks to achieve accurate identification of m6A modification sites. Specifically, M6AInter first utilizes chaos game representation theory and one-hot encoding to initialize the position and type information of nucleotides, respectively. On this basis, M6AInter extracts the position and type correlations shared by RNA sequences, and predicts the common sequence patterns by utilizing a graph contrastive clustering framework. These motifs and patterns are involved in describing the associations between RNA sequences and obtaining their low-dimensional representations. Finally, through a designed bias fusion block, these representations are combined with the frequency information of nucleotides to realize the identification of m6A modification sites. Extensive experimental results show that our model can accurately identify modified RNA sequences and can adaptively locate sequential regions associated with m6A modification sites on RNA sequences. Importantly, by exploring the role of these patterns in the identification tasks, M6AInter provides interpretable predictions and analysis at the sequence level. Bo-Wei Zhao, Xiao-Rui Su 0001, Yue Yang 0035, Pengwei Hu 0001, Lun Hu |
BIBM | 3 |
| 2023 | A Novel Graph Representation Learning Model for Drug Repositioning Using Graph Transition Probability Matrix Over Heterogenous Information Networks
Dongxu Li 0002, Bo-Wei Zhao, Xiao-Rui Su 0001, Zhu-Hong You, Pengwei Hu 0001, Lun Hu |
ICIC (3) | 4 |
| 2023 | A Deep Learning Approach Incorporating Data Missing Mechanism in Predicting Acute Kidney Injury in ICU
Zhengbo Zhang, Lei Zha, Fengcong, Xiao-Rui Su 0001, Bo-Wei Zhao, Lun Hu, Pengwei Hu 0001 |
ICIC (3) | 6 |
| 2023 | Multi-level Subgraph Representation Learning for Drug-Disease Association Prediction Over Heterogeneous Biological Information Network
Bo-Wei Zhao, Xiao-Rui Su 0001, Yue Yang 0035, Dongxu Li 0002, Pengwei Hu 0001, Zhu-Hong You, Lun Hu |
ICIC (3) | 2 |
| 2023 | Incorporating higher order network structures to improve miRNA-disease association prediction based on functional modularityabstractAs microRNAs (miRNAs) are involved in many essential biological processes, their abnormal expressions can serve as biomarkers and prognostic indicators to prevent the development of complex diseases, thus providing accurate early detection and prognostic evaluation. Although a number of computational methods have been proposed to predict miRNA-disease associations (MDAs) for further experimental verification, their performance is limited primarily by the inadequacy of exploiting lower order patterns characterizing known MDAs to identify missing ones from MDA networks. Hence, in this work, we present a novel prediction model, namely HiSCMDA, by incorporating higher order network structures for improved performance of MDA prediction. To this end, HiSCMDA first integrates miRNA similarity network, disease similarity network and MDA network to preserve the advantages of all these networks. After that, it identifies overlapping functional modules from the integrated network by predefining several higher order connectivity patterns of interest. Last, a path-based scoring function is designed to infer potential MDAs based on network paths across related functional modules. HiSCMDA yields the best performance across all datasets and evaluation metrics in the cross-validation and independent validation experiments. Furthermore, in the case studies, 49 and 50 out of the top 50 miRNAs, respectively, predicted for colon neoplasms and lung neoplasms have been validated by well-established databases. Experimental results show that rich higher order organizational structures exposed in the MDA network gain new insight into the MDA prediction based on higher order connectivity patterns. Yue Yang 0035, Xiao-Rui Su 0001, Bo-Wei Zhao, Shengwu Xiong 0001, Lun Hu |
Briefings Bioinform. | 3 |
| 2023 | iGRLDTI: an improved graph representation learning method for predicting drug-target interactions over heterogeneous biological information networkabstractMOTIVATION: The task of predicting drug-target interactions (DTIs) plays a significant role in facilitating the development of novel drug discovery. Compared with laboratory-based approaches, computational methods proposed for DTI prediction are preferred due to their high-efficiency and low-cost advantages. Recently, much attention has been attracted to apply different graph neural network (GNN) models to discover underlying DTIs from heterogeneous biological information network (HBIN). Although GNN-based prediction methods achieve better performance, they are prone to encounter the over-smoothing simulation when learning the latent representations of drugs and targets with their rich neighborhood information in HBIN, and thereby reduce the discriminative ability in DTI prediction. RESULTS: In this work, an improved graph representation learning method, namely iGRLDTI, is proposed to address the above issue by better capturing more discriminative representations of drugs and targets in a latent feature space. Specifically, iGRLDTI first constructs an HBIN by integrating the biological knowledge of drugs and targets with their interactions. After that, it adopts a node-dependent local smoothing strategy to adaptively decide the propagation depth of each biomolecule in HBIN, thus significantly alleviating over-smoothing by enhancing the discriminative ability of feature representations of drugs and targets. Finally, a Gradient Boosting Decision Tree classifier is used by iGRLDTI to predict novel DTIs. Experimental results demonstrate that iGRLDTI yields better performance that several state-of-the-art computational methods on the benchmark dataset. Besides, our case study indicates that iGRLDTI can successfully identify novel DTIs with more distinguishable features of drugs and targets. AVAILABILITY AND IMPLEMENTATION: Python codes and dataset are available at https://github.com/stevejobws/iGRLDTI/. Bo-Wei Zhao, Xiao-Rui Su 0001, Pengwei Hu 0001, Zhu-Hong You, Lun Hu |
Bioinform. | 2 |
| 2023 | Biocaiv: an integrative webserver for motif-based clustering analysis and interactive visualization of biological networksabstractBACKGROUND: As an important task in bioinformatics, clustering analysis plays a critical role in understanding the functional mechanisms of many complex biological systems, which can be modeled as biological networks. The purpose of clustering analysis in biological networks is to identify functional modules of interest, but there is a lack of online clustering tools that visualize biological networks and provide in-depth biological analysis for discovered clusters. RESULTS: Here we present BioCAIV, a novel webserver dedicated to maximize its accessibility and applicability on the clustering analysis of biological networks. This, together with its user-friendly interface, assists biological researchers to perform an accurate clustering analysis for biological networks and identify functionally significant modules for further assessment. CONCLUSIONS: BioCAIV is an efficient clustering analysis webserver designed for a variety of biological networks. BioCAIV is freely available without registration requirements at http://bioinformatics.tianshanzw.cn:8888/BioCAIV/ . Dongxu Li 0002, Bo-Wei Zhao, Xiao-Rui Su 0001, Jun Zhang 0003, Pengwei Hu 0001, Lun Hu |
BMC Bioinform. | 4 |
| 2023 | Predicting Drug-Target Interactions Over Heterogeneous Information NetworkabstractIdentifying Drug-Target Interactions (DTIs) is a critical step in studying pathogenesis and drug development. Due to the fact that conventional experimental methods usually suffer from high costs and low efficiency, various computational methods have been proposed to detect potential DTIs by extracting features from the biological information of drugs and their target proteins. Though effective, most of them fall short of considering the topological structure of the DTI network, which provides a global view to discover novel DTIs. In this paper, a network-based computational method, namely LG-DTI, is proposed to accurately predict DTIs over a heterogeneous information network. For drugs and target proteins, LG-DTI first learns not only their local representations from drug molecular structures and protein sequences, but also their global representations by using a semi-supervised heterogeneous network embedding method. These two kinds of representations consist of the final representations of drugs and target proteins, which are then incorporated into a Random Forest classifier to complete the task of DTI prediction. The performance of LG-DTI has been evaluated on two independent datasets and also compared with several state-of-the-art methods. Experimental results show the superior performance of LG-DTI. Moreover, our case study indicates that LG-DTI can be a valuable tool for identifying novel DTIs. Xiao-Rui Su 0001, Pengwei Hu 0001, Zhu-Hong You, Lun Hu |
IEEE J. Biomed. Health Informatics | 1 |
| 2023 | Biomedical Knowledge Graph Embedding With Capsule Network for Multi-Label Drug-Drug Interaction PredictionabstractDrug-drug interaction (DDI) plays an important role in drug development and administration. Most of existing network-based computation models regard the DDI prediction as a binary classification problem and generate negative DDI samples randomly, but the binary classification is not in line with the real problem since there are dozens of types of DDI and randomly generating negative samples may introduce false-negative samples since the non-observed facts can be either false or just missing. To address the above limitations, we propose a new framework called KG2ECapsule that explicitly models the multi-relational DDI data based on biomedical knowledge graphs in an end-to-end fashion. It first generates high-quality negative samples based on the average number of tail entities and head entities for each relation to reduce false-negative samples to some extent. KG2ECapsule then refines the representations of entities by recursively propagating the embeddings from the attention-based receptive fields of entities. Empirical results on three biomedical knowledge graphs of different scales show that KG2ECapsule outperforms the state-of-the-art methods consistently in multi-label DDI prediction task and further studies verify the efficacy of both probability-based sampling strategy and non-linear transformation for modeling multi-relational data. Xiao-Rui Su 0001, Zhu-Hong You, De-Shuang Huang, Lei Wang 0121, Leon Wong, Bo-Wei Zhao |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2022 | MRLDTI: A Meta-path-Based Representation Learning Model for Drug-Target Interaction Prediction
Bo-Wei Zhao, Lun Hu, Pengwei Hu 0001, Zhu-Hong You, Xiao-Rui Su 0001, Dongxu Li 0002, Ping Zhang 0027 |
ICIC (2) | 5 |
| 2022 | A Novel Fuzzy-Based MOPSO Algorithm for Identifying Clusters From Complex NetworksabstractMany complicated systems can be modeled as complex networks, and a variety of graph clustering algorithms have been proposed to perform accurate clustering analysis for better understanding system behaviors. However, most of them suffer the disadvantage of slow convergence. In this paper, we incorporate multi-objective particle swarm optimization (MOPSO) into a well-established fuzzy clustering algorithm, i.e., FCAN, and propose an improved Fuzzy-based Graph Clustering Algorithm, namely IMFCAN, which retains all the benefits gained with FCAN while achieving significantly fast convergence rate. Specially, IMFCAN enhances the ability of handling the imbalance observed in the distribution of fuzzy membership of nodes by introducing an instance-frequency-weighted regularization (IR) scheme. After that, IMFCAN develops an effective solution to reach a consensus optimization among them by balancing global exploration and local exploitation abilities of particles. Experimental results on four practical datasets demonstrate that IMFCAN performs better than several state-of-the-art clustering algorithm in terms of accuracy and convergence. Hence, IMFCAN is a promising algorithm for addressing the clustering analysis of complex networks. Yue Yang 0035, Xiao-Rui Su 0001, Bo-Wei Zhao, Lun Hu |
ICTAI | 2 |
| 2022 | A deep learning method for repurposing antiviral drugs against new viruses via multi-view nonnegative matrix factorization and its application to SARS-CoV-2abstractThe outbreak of COVID-19 caused by SARS-coronavirus (CoV)-2 has made millions of deaths since 2019. Although a variety of computational methods have been proposed to repurpose drugs for treating SARS-CoV-2 infections, it is still a challenging task for new viruses, as there are no verified virus-drug associations (VDAs) between them and existing drugs. To efficiently solve the cold-start problem posed by new viruses, a novel constrained multi-view nonnegative matrix factorization (CMNMF) model is designed by jointly utilizing multiple sources of biological information. With the CMNMF model, the similarities of drugs and viruses can be preserved from their own perspectives when they are projected onto a unified latent feature space. Based on the CMNMF model, we propose a deep learning method, namely VDA-DLCMNMF, for repurposing drugs against new viruses. VDA-DLCMNMF first initializes the node representations of drugs and viruses with their corresponding latent feature vectors to avoid a random initialization and then applies graph convolutional network to optimize their representations. Given an arbitrary drug, its probability of being associated with a new virus is computed according to their representations. To evaluate the performance of VDA-DLCMNMF, we have conducted a series of experiments on three VDA datasets created for SARS-CoV-2. Experimental results demonstrate that the promising prediction accuracy of VDA-DLCMNMF. Moreover, incorporating the CMNMF model into deep learning gains new insight into the drug repurposing for SARS-CoV-2, as the results of molecular docking experiments reveal that four antiviral drugs identified by VDA-DLCMNMF have the potential ability to treat SARS-CoV-2 infections. Xiao-Rui Su 0001, Lun Hu, Zhu-Hong You, Pengwei Hu 0001, Lei Wang 0121, Bo-Wei Zhao |
Briefings Bioinform. | 1 |
| 2022 | Attention-based Knowledge Graph Representation Learning for Predicting Drug-drug InteractionsabstractDrug-drug interactions (DDIs) are known as the main cause of life-threatening adverse events, and their identification is a key task in drug development. Existing computational algorithms mainly solve this problem by using advanced representation learning techniques. Though effective, few of them are capable of performing their tasks on biomedical knowledge graphs (KGs) that provide more detailed information about drug attributes and drug-related triple facts. In this work, an attention-based KG representation learning framework, namely DDKG, is proposed to fully utilize the information of KGs for improved performance of DDI prediction. In particular, DDKG first initializes the representations of drugs with their embeddings derived from drug attributes with an encoder-decoder layer, and then learns the representations of drugs by recursively propagating and aggregating first-order neighboring information along top-ranked network paths determined by neighboring node embeddings and triple facts. Last, DDKG estimates the probability of being interacting for pairwise drugs with their representations in an end-to-end manner. To evaluate the effectiveness of DDKG, extensive experiments have been conducted on two practical datasets with different sizes, and the results demonstrate that DDKG is superior to state-of-the-art algorithms on the DDI prediction task in terms of different evaluation metrics across all datasets. Xiao-Rui Su 0001, Lun Hu, Zhu-Hong You, Pengwei Hu 0001, Bo-Wei Zhao |
Briefings Bioinform. | 1 |
| 2022 | A machine learning framework based on multi-source feature fusion for circRNA-disease association predictionabstractCircular RNAs (circRNAs) are involved in the regulatory mechanisms of multiple complex diseases, and the identification of their associations is critical to the diagnosis and treatment of diseases. In recent years, many computational methods have been designed to predict circRNA-disease associations. However, most of the existing methods rely on single correlation data. Here, we propose a machine learning framework for circRNA-disease association prediction, called MLCDA, which effectively fuses multiple sources of heterogeneous information including circRNA sequences and disease ontology. Comprehensive evaluation in the gold standard dataset showed that MLCDA can successfully capture the complex relationships between circRNAs and diseases and accurately predict their potential associations. In addition, the results of case studies on real data show that MLCDA significantly outperforms other existing methods. MLCDA can serve as a useful tool for circRNA-disease association prediction, providing mechanistic insights for disease research and thus facilitating the progress of disease treatment. Lei Wang 0121, Leon Wong, Zhengwei Li 0001, Xiao-Rui Su 0001, Bo-Wei Zhao, Zhu-Hong You |
Briefings Bioinform. | 5 |
| 2022 | HINGRL: predicting drug-disease associations with graph representation learning on heterogeneous information networksabstractIdentifying new indications for drugs plays an essential role at many phases of drug research and development. Computational methods are regarded as an effective way to associate drugs with new indications. However, most of them complete their tasks by constructing a variety of heterogeneous networks without considering the biological knowledge of drugs and diseases, which are believed to be useful for improving the accuracy of drug repositioning. To this end, a novel heterogeneous information network (HIN) based model, namely HINGRL, is proposed to precisely identify new indications for drugs based on graph representation learning techniques. More specifically, HINGRL first constructs a HIN by integrating drug-disease, drug-protein and protein-disease biological networks with the biological knowledge of drugs and diseases. Then, different representation strategies are applied to learn the features of nodes in the HIN from the topological and biological perspectives. Finally, HINGRL adopts a Random Forest classifier to predict unknown drug-disease associations based on the integrated features of drugs and diseases obtained in the previous step. Experimental results demonstrate that HINGRL achieves the best performance on two real datasets when compared with state-of-the-art models. Besides, our case studies indicate that the simultaneous consideration of network topology and biological knowledge of drugs and diseases allows HINGRL to precisely predict drug-disease associations from a more comprehensive perspective. The promising performance of HINGRL also reveals that the utilization of rich heterogeneous information provides an alternative view for HINGRL to identify novel drug-disease associations especially for new diseases. Bo-Wei Zhao, Lun Hu, Zhu-Hong You, Lei Wang 0121, Xiao-Rui Su 0001 |
Briefings Bioinform. | 5 |
| 2022 | A geometric deep learning framework for drug repositioning over heterogeneous information networksabstractDrug repositioning (DR) is a promising strategy to discover new indicators of approved drugs with artificial intelligence techniques, thus improving traditional drug discovery and development. However, most of DR computational methods fall short of taking into account the non-Euclidean nature of biomedical network data. To overcome this problem, a deep learning framework, namely DDAGDL, is proposed to predict drug-drug associations (DDAs) by using geometric deep learning (GDL) over heterogeneous information network (HIN). Incorporating complex biological information into the topological structure of HIN, DDAGDL effectively learns the smoothed representations of drugs and diseases with an attention mechanism. Experiment results demonstrate the superior performance of DDAGDL on three real-world datasets under 10-fold cross-validation when compared with state-of-the-art DR methods in terms of several evaluation metrics. Our case studies and molecular docking experiments indicate that DDAGDL is a promising DR tool that gains new insights into exploiting the geometric prior knowledge for improved efficacy. Bo-Wei Zhao, Xiao-Rui Su 0001, Pengwei Hu 0001, Yu-Peng Ma, Xi Zhou 0007, Lun Hu |
Briefings Bioinform. | 2 |
| 2022 | Multi-view heterogeneous molecular network representation learning for protein-protein interaction predictionabstractBACKGROUND: Protein-protein interaction (PPI) plays an important role in regulating cells and signals. Despite the ongoing efforts of the bioassay group, continued incomplete data limits our ability to understand the molecular roots of human disease. Therefore, it is urgent to develop a computational method to predict PPIs from the perspective of molecular system. METHODS: In this paper, a highly efficient computational model, MTV-PPI, is proposed for PPI prediction based on a heterogeneous molecular network by learning inter-view protein sequences and intra-view interactions between molecules simultaneously. On the one hand, the inter-view feature is extracted from the protein sequence by k-mer method. On the other hand, we use a popular embedding method LINE to encode the heterogeneous molecular network to obtain the intra-view feature. Thus, the protein representation used in MTV-PPI is constructed by the aggregation of its inter-view feature and intra-view feature. Finally, random forest is integrated to predict potential PPIs. RESULTS: To prove the effectiveness of MTV-PPI, we conduct extensive experiments on a collected heterogeneous molecular network with the accuracy of 86.55%, sensitivity of 82.49%, precision of 89.79%, AUC of 0.9301 and AUPR of 0.9308. Further comparison experiments are performed with various protein representations and classifiers to indicate the effectiveness of MTV-PPI in predicting PPIs based on a complex network. CONCLUSION: The achieved experimental results illustrate that MTV-PPI is a promising tool for PPI prediction, which may provide a new perspective for the future interactions prediction researches based on heterogeneous molecular network. Xiao-Rui Su 0001, Lun Hu, Zhu-Hong You, Pengwei Hu 0001, Bo-Wei Zhao |
BMC Bioinform. | 1 |
| 2022 | RLFDDA: a meta-path based graph representation learning model for drug-disease association predictionabstractBACKGROUND: Drug repositioning is a very important task that provides critical information for exploring the potential efficacy of drugs. Yet developing computational models that can effectively predict drug-disease associations (DDAs) is still a challenging task. Previous studies suggest that the accuracy of DDA prediction can be improved by integrating different types of biological features. But how to conduct an effective integration remains a challenging problem for accurately discovering new indications for approved drugs. METHODS: In this paper, we propose a novel meta-path based graph representation learning model, namely RLFDDA, to predict potential DDAs on heterogeneous biological networks. RLFDDA first calculates drug-drug similarities and disease-disease similarities as the intrinsic biological features of drugs and diseases. A heterogeneous network is then constructed by integrating DDAs, disease-protein associations and drug-protein associations. With such a network, RLFDDA adopts a meta-path random walk model to learn the latent representations of drugs and diseases, which are concatenated to construct joint representations of drug-disease associations. As the last step, we employ the random forest classifier to predict potential DDAs with their joint representations. RESULTS: To demonstrate the effectiveness of RLFDDA, we have conducted a series of experiments on two benchmark datasets by following a ten-fold cross-validation scheme. The results show that RLFDDA yields the best performance in terms of AUC and F1-score when compared with several state-of-the-art DDAs prediction models. We have also conducted a case study on two common diseases, i.e., paclitaxel and lung tumors, and found that 7 out of top-10 diseases and 8 out of top-10 drugs have already been validated for paclitaxel and lung tumors respectively with literature evidence. Hence, the promising performance of RLFDDA may provide a new perspective for novel DDAs discovery over heterogeneous networks. Menglong Zhang, Bo-Wei Zhao, Xiao-Rui Su 0001, Yue Yang 0035, Lun Hu |
BMC Bioinform. | 3 |
| 2022 | NSECDA: Natural Semantic Enhancement for CircRNA-Disease Association PredictionabstractIncreasing evidence suggest that circRNA, as one of the most promising emerging biomarkers, has a very close relationship with diseases. Exploring the relationship between circRNA and diseases can provide novel perspective for diseases diagnosis and pathogenesis. The existing circRNA-disease association (CDA) prediction models, however, generally treat the data attributes equally, do not pay special attention to the attributes with more significant influence, and do not make full use of the correlation and symbiosis between attributes to dig into the latent semantic information of the data. Therefore, in response to the above problems, this paper proposes a natural semantic enhancement method NSECDA to predict CDA. In practical terms, we first recognize the circRNA sequence as a biological language, and analyze its natural semantic properties through the natural language understanding theory; then integrate it with disease attributes, circRNA and disease Gaussian Interaction Profile (GIP) kernel attributes, and use Graph Attention Network (GAT) to focus on the influential attributes, so as to mine the deeply hidden features; finally, the Rotation Forest (RoF) classifier was used to accurately determine CDA. In the gold standard data set CircR2Disease, NSECDA achieved 92.49% accuracy with 0.9225 AUC score. In comparison with the non-natural semantic enhancement model and other classifier models, NSECDA also shows competitive performance. Additionally, 25 of the CDA pairs with unknown associations in the top 30 prediction scores of NSECDA have been proven by newly reported studies. These achievements suggest that NSECDA is an effective model to predict CDA, which can provide credible candidate for subsequent wet experiments, thus significantly reducing the scope of investigations. Lei Wang 0121, Leon Wong, Zhu-Hong You, De-Shuang Huang, Xiao-Rui Su 0001, Bo-Wei Zhao |
IEEE J. Biomed. Health Informatics | 5 |
| 2021 | Predicting miRNA-Disease Associations via a New MeSH Headings Representation of Diseases and eXtreme Gradient Boosting
Zhu-Hong You, Lei Wang 0121, Leon Wong, Xiao-Rui Su 0001, Bo-Wei Zhao |
ICIC (3) | 5 |
| 2021 | Protein-Protein Interaction Prediction by Integrating Sequence Information and Heterogeneous Network Representation
Xiao-Rui Su 0001, Zhu-Hong You, Zhen-Hao Guo |
ICIC (3) | 1 |
| 2021 | Detection of Drug-Drug Interactions Through Knowledge Graph Integrating Multi-attention with Capsule Network
Xiao-Rui Su 0001, Zhu-Hong You, Bo-Wei Zhao |
ICIC (3) | 1 |
| 2021 | In silico drug repositioning using deep learning and comprehensive similarity measuresabstractBACKGROUND: Drug repositioning, meanings finding new uses for existing drugs, which can accelerate the processing of new drugs research and development. Various computational methods have been presented to predict novel drug-disease associations for drug repositioning based on similarity measures among drugs and diseases. However, there are some known associations between drugs and diseases that previous studies not utilized. METHODS: In this work, we develop a deep gated recurrent units model to predict potential drug-disease interactions using comprehensive similarity measures and Gaussian interaction profile kernel. More specifically, the similarity measure is used to exploit discriminative feature for drugs based on their chemical fingerprints. Meanwhile, the Gaussian interactions profile kernel is employed to obtain efficient feature of diseases based on known disease-disease associations. Then, a deep gated recurrent units model is developed to predict potential drug-disease interactions. RESULTS: The performance of the proposed model is evaluated on two benchmark datasets under tenfold cross-validation. And to further verify the predictive ability, case studies for predicting new potential indications of drugs were carried out. CONCLUSION: The experimental results proved the proposed model is a useful tool for predicting new indications for drugs or new treatments for diseases, and can accelerate drug repositioning and related drug research and discovery. Zhu-Hong You, Lei Wang 0065, Xiao-Rui Su 0001, Xi Zhou 0007, Tonghai Jiang |
BMC Bioinform. | 4 |
| 2020 | Prediction of LncRNA-Disease Associations Based on Network Representation LearningabstractMassive observations have indicated that long noncoding RNAs (lncRNAs) are crucial in a number of biological processes and associated with various human diseases. Developing an efficient calculation model to predict the associations between lncRNA and diseases is not only beneficial to disease diagnosis, treatment, prognosis and potential drug targets in drug discovery, but also avoid the waste of human and material resources brought by biological experiments. In this paper, we proposed a novel prediction of lncRNA-disease associations based on complex and comprehensive molecular associations network (MAN), which integrated nine kinds of interactions among five molecules, including lncRNA, miRNA, disease, drug and protein. Network embedding Node2vec method was applied to extract behavior feature from MAN to generate a low-dimension vector containing nodes and edges information. After implementing 5-fold cross validation, the proposed method yielded good prediction performance with an average Accuracy of 91.91%, Sensitivity of 94.05%, Specificity of 89.76%, Precision of 90.21%, MCC value of 83.91%, AUC value of 0.9746 and AUPR of 0.9693. Comparative experiment indicates the behavior feature extracted by Node2vec is more representative than attribute features of lncRNA adopted 3-mer and diseases extracted by semantic similarity. Moreover, breast cancer, colon cancer and lung cancer are explored in case study. As a results, more than half of top 5 interactions are successfully confirmed for each disease by other datasets. Based on these reliable results, it is anticipated that proposed model is feasible and effective to predict lncRNA-disease associations at a global molecules level, which is a new respective for future biomedical researches. Xiao-Rui Su 0001, Zhu-Hong You |
BIBM | 1 |
| 2020 | A Novel Computational Approach for Predicting Drug-Target Interactions via Network Representation Learning
Xiao-Rui Su 0001, Zhu-Hong You, Ji-Ren Zhou, Xiao Li 0007 |
ICIC (2) | 1 |
| 2020 | A Unified Deep Biological Sequence Representation Learning with Pretrained Encoder-Decoder Model
Zhu-Hong You, Xiao-Rui Su 0001, De-Shuang Huang, Zhen-Hao Guo |
ICIC (2) | 3 |