EDBT 2026 Demo / reviewers in the wild / expert
Mengyun Yang
dblp:139/1941
· DBLP profile ↗
23ranked-venue papers
5as first author
15since 2021 · last 2026
0000-0002-2703-533XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 22 · 4 first-author · 14 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TriCloud: Drug-Target-Disease Ternary Network for Drug Repositioning Research Based on Point Cloud ModelingabstractIn recent years, the "drug-target-disease" association has become increasingly complex and data-scarce. Existing methods are limited by information loss caused by explicit graph construction when modeling triplet relationships, making it difficult to effectively capture geometric structures and long-range dependencies, thereby affecting prediction performance. To address this, this paper proposes a new method based on point cloud modeling-TriCloud. This method represents each triad as a spatial point cloud, encoding semantic and topological relationships through geometric coordinates, and performs feature learning directly on an unordered point set, avoiding reliance on predefined graph structures. Based on the PointNet architecture, it introduces a multi-view feature extraction and fusion mechanism to enhance the modeling capability of global structures and complex spatial patterns. Experimental results show that TriCloud significantly outperforms existing methods on multiple benchmark datasets, achieving an AUC of 0.9995 and an AUPR of 0.9996, with all metrics ranking first. External validation demonstrates its excellent generalization ability. Feature analysis reveals that the geometric-semantic joint features of positive samples play a dominant role in classification. This study provides an efficient and reliable computational framework for drug repurposing, contributing to the development of precision medicine. Xiwei Tang, Wanjun Ma, Anzheng Gao, Mengyun Yang, Weijun Liang, Wenjun Li 0001 |
IEEE Trans. Comput. Biol. Bioinform. | 4 |
| 2026 | DLP: Duplex Link Prediction via Subspace Segmentation for Predicting Drug-MiRNA AssociationsabstractThe arduous and costly journey of drug discovery is increasingly intersecting with computational approaches, which promise to accelerate the analysis of bioassays and biomedical literature. The critical role of microRNAs (miRNAs) in disease progression has been underscored in recent studies, elevating them as potential therapeutic targets. This emphasizes the need for the development of sophisticated computational models that can effectively identify promising drug targets such as miRNAs. Herein, we present a novel method, termed Duplex Link Prediction (DLP), rooted in subspace segmentation, to pinpoint potential miRNA targets. Our approach initiates with the application of the Network Enhancement (NE) algorithm to refine the similarity metric between miRNAs. Thereafter, we construct two matrices by pre-loading the association matrix from both the drug and miRNA perspectives, employing the K Nearest Neighbors (KNN) technique. The DLSR algorithm is then applied to predict potential associations. The final predicted association scores are ascertained through the weighted mean of the two matrices. Our empirical findings suggest that the DLP algorithm outperforms current methodologies in the realm of identifying potential miRNA drug targets. Case study validations further reinforce the real-world applicability and effectiveness of our proposed method. Kai Zheng 0020, Guihua Duan, Qichang Zhao, Mengyun Yang, Jianxin Wang 0001 |
IEEE Trans. Comput. Biol. Bioinform. | 4 |
| 2026 | BiBLDR: Bidirectional Behavior Learning for Drug RepositioningabstractMany deep learning methods represented by graph-based approaches achieve significant progress in drug repositioning. However, these graph-based methods face a critical limitation: they often fail in cold-start scenarios because the graph structure relies heavily on known association information from both the drug and disease sides. To address this challenge, we propose a bidirectional behavior learning strategy for drug repositioning, BiBLDR, an innovative framework that reformulates drug repositioning as a behavior sequence learning task. First, we construct bidirectional behavioral sequences based on drug and disease sides. Bidirectional behavior sequences ensure sufficient information for model learning in both drug and disease cold-start scenarios, while providing more precise feature representations for association prediction tasks. Subsequently, we propose a two-stage strategy for drug repositioning. In the first stage, we construct prototype spaces to characterise the representational attributes of drugs and diseases. In the second stage, these refined prototypes and bidirectional behavior sequence data are leveraged to predict potential drug-disease associations. This design allows BiBLDR to more robustly capture hidden pharmacological relationships from bidirectional behavioral sequences, delivering significant benefits in cold-start scenarios. Extensive experiments demonstrate that our method achieves state-of-the-art performance on benchmark datasets. Meanwhile, BiBLDR demonstrates significantly superior performance compared to previous methods in cold-start scenarios. Renye Zhang, Mengyun Yang, Qichang Zhao, Jianxin Wang 0001 |
IEEE J. Biomed. Health Informatics | 2 |
| 2025 | A Novel Sample Selection for Deep Learning Model in Computational Drug Repositioning
Renye Zhang, Mengyun Yang |
ISBRA (1) | 4 |
| 2025 | LRTM: Left-Right Transition Matrices for Molecular Association PredictionabstractMolecular associations are central to most biological processes. The discovery and identification of potential associations between molecules can provide insights into biological exploration, diagnostic and therapeutic interventions, and drug development. So far many relevant computational methods have been proposed, but most of them are usually limited to specific domains and rely on complex preprocessing procedures, which restricts the models' ability to be applied to other tasks. Therefore, it remains a challenge to explore a generalized approach to accurately predicting potential associations. In this study, We propose Left-Right Transition Matrices (LRTM) for molecular association prediction. From the perspective on the diffusion model, we construct two transition matrices to model undirected graph information propagation. This allows modeling the transition probabilities of links, which facilitates link prediction in molecular bipartite networks. The extensive experimental results show that the proposed LRTM algorithm performs better than the compared methods. Also, the proposed algorithm has the potential for cross-task prediction. Furthermore, case studies show that LRTM is a powerful tool that can be effectively applied to practical applications. Kai Zheng 0020, Guihua Duan, Mengyun Yang, Wei Wu 0011, Yaohang Li, Jianxin Wang 0001 |
IEEE Trans. Comput. Biol. Bioinform. | 3 |
| 2024 | MFCM-DTI model of multimodal feature fusion: prediction of drug-target interactionabstractDrug repositioning is a vital area of biomedicine, where confirming interactions between drugs and specific targets is essential for establishing the efficacy of pharmaceutical agents. Traditional in vitro screening methods have limitations, prompting the use of computer simulations as an effective alternative for predicting drug-target interactions (DTI). This approach has gained significant attention in the scientific community. In this study, we introduce MFCM-DTI, a DTI prediction model that employs multimodal features to accurately capture the intricate interactions between drug molecular structures and key amino acids of target proteins. Our results demonstrate that MFCM-DTI outperforms existing models in prediction accuracy and robustness. Furthermore, MFCM-DTI has been successfully used to predict interactions between key SARS-CoV-2 proteins and existing drugs, providing a solid foundation for developing therapeutic agents against SARS-CoV-2 infection. This study underscores the broad applicability and strong predictive capabilities of MFCM-DTI in drug-protein interaction prediction, opening new avenues for research. The predicted drug targets and interaction data offer valuable insights for future experimental validation and clinical trials, potentially driving innovation in the biomedical field. Wenjun Li 0001, Wanjun Ma, Mengyun Yang, Xiwei Tang |
BIBM | 3 |
| 2023 | AttentionDTA: Drug-Target Binding Affinity Prediction by Sequence-Based Deep Learning With Attention MechanismabstractThe identification of drug-target relations (DTRs) is substantial in drug development. A large number of methods treat DTRs as drug-target interactions (DTIs), a binary classification problem. The main drawback of these methods are the lack of reliable negative samples and the absence of many important aspects of DTR, including their dose dependence and quantitative affinities. With increasing number of publications of drug-protein binding affinity data recently, DTRs prediction can be viewed as a regression problem of drug-target affinities (DTAs) which reflects how tightly the drug binds to the target and can present more detailed and specific information than DTIs. The growth of affinity data enables the use of deep learning architectures, which have been shown to be among the state-of-the-art methods in binding affinity prediction. Although relatively effective, due to the black-box nature of deep learning, these models are less biologically interpretable. In this study, we proposed a deep learning-based model, named AttentionDTA, which uses attention mechanism to predict DTAs. Different from the models using 3D structures of drug-target complexes or graph representation of drugs and proteins, the novelty of our work is to use attention mechanism to focus on key subsequences which are important in drug and protein sequences when predicting its affinity. We use two separate one-dimensional Convolution Neural Networks (1D-CNNs) to extract the semantic information of drug's SMILES string and protein's amino acid sequence. Furthermore, a two-side multi-head attention mechanism is developed and embedded to our model to explore the relationship between drug features and protein features. We evaluate our model on three established DTA benchmark datasets, Davis, Metz, and KIBA. AttentionDTA outperforms the state-of-the-art deep learning methods under different evaluation metrics. The results show that the attention-based model can effectively extract protein features related to drug information and drug features related to protein information to better predict drug target affinities. It is worth mentioning that we test our model on IC50 dataset, which provides the binding sites between drugs and proteins, to evaluate the ability of our model to locate binding sites. Finally, we visualize the attention weight to demonstrate the biological significance of the model. The source code of AttentionDTA can be downloaded from https://github.com/zhaoqichang/AttentionDTA_TCBB. Qichang Zhao, Guihua Duan, Mengyun Yang, Zhongjian Cheng, Yaohang Li, Jianxin Wang 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2022 | Drug repositioning based on multi-view learning with matrix completionabstractDetermining drug indications is a critical part of the drug development process. However, traditional drug discovery is expensive and time-consuming. Drug repositioning aims to find potential indications for existing drugs, which is considered as an important alternative to the traditional drug discovery. In this article, we propose a multi-view learning with matrix completion (MLMC) method to predict the potential associations between drugs and diseases. Specifically, MLMC first learns the comprehensive similarity matrices from five drug similarity matrices and two disease similarity matrices based on the multi-view learning (ML) with Laplacian graph regularization, and updates the drug-disease association matrix simultaneously. Then, we introduce matrix completion (MC) to add some positive entries in original association matrix based on low-rank structure, and re-execute the multi-view learning algorithm for association prediction. At last, the prediction results of the above two operations are integrated as the final output. Evaluated by 10-fold cross-validation and de novo tests, MLMC achieves higher prediction accuracy than the current state-of-the-art methods. Moreover, case studies confirm the ability of our method in novel drug-disease association discovery. The codes of MLMC are available at https://github.com/BioinformaticsCSU/MLMC. Contact: [email protected]. Yixin Yan, Mengyun Yang, Guihua Duan, Xiaoqing Peng, Jianxin Wang 0001 |
Briefings Bioinform. | 2 |
| 2022 | Biomedical Data and Deep Learning Computational Models for Predicting Compound-Protein RelationsabstractThe identification of compound-protein relations (CPRs), which includes compound-protein interactions (CPIs) and compound-protein affinities (CPAs), is critical to drug development. A common method for compound-protein relation identification is the use of in vitro screening experiments. However, the number of compounds and proteins is massive, and in vitro screening experiments are labor-intensive, expensive, and time-consuming with high failure rates. Researchers have developed a computational field called virtual screening (VS) to aid experimental drug development. These methods utilize experimentally validated biological interaction information to generate datasets and use the physicochemical and structural properties of compounds and target proteins as input information to train computational prediction models. At present, deep learning has been widely used in computer vision and natural language processing and has experienced epoch-making progress. At the same time, deep learning has also been used in the field of biomedicine widely, and the prediction of CPRs based on deep learning has developed rapidly and has achieved good results. The purpose of this study is to investigate and discuss the latest applications of deep learning techniques in CPR prediction. First, we describe the datasets and feature engineering (i.e., compound and protein representations and descriptors) commonly used in CPR prediction methods. Then, we review and classify recent deep learning approaches in CPR prediction. Next, a comprehensive comparison is performed to demonstrate the prediction performance of representative methods on classical datasets. Finally, we discuss the current state of the field, including the existing challenges and our proposed future directions. We believe that this investigation will provide sufficient references and insight for researchers to understand and develop new deep learning methods to enhance CPR predictions. Qichang Zhao, Mengyun Yang, Zhongjian Cheng, Yaohang Li, Jianxin Wang 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2022 | Feature and Nuclear Norm Minimization for Matrix CompletionabstractMatrix completion, whose goal is to recover a matrix from a few entries observed, is a fundamental model behind many applications. Our study shows that, in many applications, the to-be-complete matrix can be represented as the sum of a low-rank matrix and a sparse matrix associating with side information matrices. The low-rank matrix depicts the global patterns while the sparse matrix characterizes the local patterns, which are often described by the side information. Accordingly, to achieve high-quality matrix completion, we propose a Feature and Nuclear Norm Minimization (FNNM) model. The rationale of FNNM is to employ transductive completion to generalize the global pattern and inductive completion to recover the local pattern. Alternative minimization algorithm based on fixed-point iteration is developed to numerically solve the FNNM model. FNNM has demonstrated promising results on a variety of applications, including movie recommendation, drug-target interaction prediction, and multi-label learning, consistently outperforming the state-of-the-art matrix completion algorithms. Mengyun Yang, Yaohang Li, Jianxin Wang 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2021 | Biomedical data and computational models for drug repositioning: a comprehensive reviewabstractDrug repositioning can drastically decrease the cost and duration taken by traditional drug research and development while avoiding the occurrence of unforeseen adverse events. With the rapid advancement of high-throughput technologies and the explosion of various biological data and medical data, computational drug repositioning methods have been appealing and powerful techniques to systematically identify potential drug-target interactions and drug-disease interactions. In this review, we first summarize the available biomedical data and public databases related to drugs, diseases and targets. Then, we discuss existing drug repositioning approaches and group them based on their underlying computational models consisting of classical machine learning, network propagation, matrix factorization and completion, and deep learning based models. We also comprehensively analyze common standard data sets and evaluation metrics used in drug repositioning, and give a brief comparison of various prediction methods on the gold standard data sets. Finally, we conclude our review with a brief discussion on challenges in computational drug repositioning, which includes the problem of reducing the noise and incompleteness of biomedical data, the ensemble of various computation drug repositioning methods, the importance of designing reliable negative samples selection methods, new techniques dealing with the data sparseness problem, the construction of large-scale and comprehensive benchmark data sets and the analysis and explanation of the underlying mechanisms of predicted interactions. Huimin Luo, Min Li 0007, Mengyun Yang, Fang-Xiang Wu, Yaohang Li, Jianxin Wang 0001 |
Briefings Bioinform. | 3 |
| 2021 | Computational drug repositioning based on multi-similarities bilinear matrix factorizationabstractWith the development of high-throughput technology and the accumulation of biomedical data, the prior information of biological entity can be calculated from different aspects. Specifically, drug-drug similarities can be measured from target profiles, drug-drug interaction and side effects. Similarly, different methods and data sources to calculate disease ontology can result in multiple measures of pairwise disease similarities. Therefore, in computational drug repositioning, developing a dynamic method to optimize the fusion process of multiple similarities is a crucial and challenging task. In this study, we propose a multi-similarities bilinear matrix factorization (MSBMF) method to predict promising drug-associated indications for existing and novel drugs. Instead of fusing multiple similarities into a single similarity matrix, we concatenate these similarity matrices of drug and disease, respectively. Applying matrix factorization methods, we decompose the drug-disease association matrix into a drug-feature matrix and a disease-feature matrix. At the same time, using these feature matrices as basis, we extract effective latent features representing the drug and disease similarity matrices to infer missing drug-disease associations. Moreover, these two factored matrices are constrained by non-negative factorization to ensure that the completed drug-disease association matrix is biologically interpretable. In addition, we numerically solve the MSBMF model by an efficient alternating direction method of multipliers algorithm. The computational experiment results show that MSBMF obtains higher prediction accuracy than the state-of-the-art drug repositioning methods in cross-validation experiments. Case studies also demonstrate the effectiveness of our proposed method in practical applications. Availability: The data and code of MSBMF are freely available at https://github.com/BioinformaticsCSU/MSBMF. Corresponding author: Jianxin Wang, School of Computer Science and Engineering, Central South University, Changsha, Hunan 410083, P. R. China. E-mail: [email protected] Supplementary Data: Supplementary data are available online at https://academic.oup.com/bib. Mengyun Yang, Gaoyan Wu, Qichang Zhao, Yaohang Li, Jianxin Wang 0001 |
Briefings Bioinform. | 1 |
| 2021 | IsoResolve: predicting splice isoform functions by integrating gene and isoform-level features with domain adaptationabstractMOTIVATION: High resolution annotation of gene functions is a central goal in functional genomics. A single gene may produce multiple isoforms with different functions through alternative splicing. Conventional approaches, however, consider a gene as a single entity without differentiating these functionally different isoforms. Towards understanding gene functions at higher resolution, recent efforts have focused on predicting the functions of isoforms. However, the performance of existing methods is far from satisfactory mainly because of the lack of isoform-level functional annotation. RESULTS: We present IsoResolve, a novel approach for isoform function prediction, which leverages the information from gene function prediction models with domain adaptation (DA). IsoResolve treats gene-level and isoform-level features as source and target domains, respectively. It uses DA to project the two domains into a latent variable space in such a way that the latent variables from the two domains have similar distribution, which enables the gene domain information to be leveraged for isoform function prediction. We systematically evaluated the performance of IsoResolve in predicting functions. Compared with five state-of-the-art methods, IsoResolve achieved significantly better performance. IsoResolve was further validated by case studies of genes with isoform-level functional annotation. AVAILABILITY AND IMPLEMENTATION: IsoResolve is freely available at https://github.com/genemine/IsoResolve. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Hong-Dong Li, Changhuo Yang, Mengyun Yang, Fang-Xiang Wu, Gilbert S. Omenn, Jianxin Wang 0001 |
Bioinform. | 4 |
| 2021 | Heterogeneous graph inference with matrix completion for computational drug repositioningabstractMOTIVATION: Emerging evidence presents that traditional drug discovery experiment is time-consuming and high costs. Computational drug repositioning plays a critical role in saving time and resources for drug research and discovery. Therefore, developing more accurate and efficient approaches is imperative. Heterogeneous graph inference is a classical method in computational drug repositioning, which not only has high convergence precision, but also has fast convergence speed. However, the method has not fully considered the sparsity of heterogeneous association network. In addition, rough similarity measure can reduce the performance in identifying drug-associated indications. RESULTS: In this article, we propose a heterogeneous graph inference with matrix completion (HGIMC) method to predict potential indications for approved and novel drugs. First, we use a bounded matrix completion (BMC) model to prefill a part of the missing entries in original drug-disease association matrix. This step can add more positive and formative drug-disease edges between drug network and disease network. Second, Gaussian radial basis function (GRB) is employed to improve the drug and disease similarities since the performance of heterogeneous graph inference more relies on similarity measures. Next, based on the updated drug-disease associations and new similarity measures of drug and disease, we construct a novel heterogeneous drug-disease network. Finally, HGIMC utilizes the heterogeneous network to infer the scores of unknown association pairs, and then recommend the promising indications for drugs. To evaluate the performance of our method, HGIMC is compared with five state-of-the-art approaches of drug repositioning in the 10-fold cross-validation and de novo tests. As the numerical results shown, HGIMC not only achieves a better prediction performance but also has an excellent computation efficiency. In addition, cases studies also confirm the effectiveness of our method in practical application. AVAILABILITYAND IMPLEMENTATION: The HGIMC software and data are freely available at https://github.com/BioinformaticsCSU/HGIMC, https://hub.docker.com/repository/docker/yangmy84/hgimc and http://doi.org/10.5281/zenodo.4285640. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Mengyun Yang, Yunpei Xu, Chengqian Lu, Jianxin Wang 0001 |
Bioinform. | 1 |
| 2021 | Collaborative Matrix Factorization with Soft Regularization for Drug-Target Interaction Prediction
Li-Gang Gao, Mengyun Yang, Jianxin Wang 0001 |
J. Comput. Sci. Technol. | 2 |
| 2020 | De novo Prediction of Drug-Target Interaction via Laplacian Regularized Schatten-p Norm Minimization
Gaoyan Wu, Mengyun Yang, Yaohang Li, Jianxin Wang 0001 |
ISBRA | 2 |
| 2020 | miRTMC: A miRNA Target Prediction Method Based on Matrix Completion AlgorithmabstractmicroRNAs (miRNAs) are small non-coding RNAs which modulate the stability of gene targets and their rates of translation into proteins at transcriptional level and post-transcriptional level. miRNA dysfunctions can lead to human diseases because of dysregulation of their targets. Correct miRNA target prediction will lead to better understanding of the mechanisms of human diseases and provide hints on curing them. In recent years, computational miRNA target prediction methods have been proposed according to the interaction rules between miRNAs and targets. However, these methods suffer from high false positive rates due to the complicated relationship between miRNAs and their targets. The rapidly growing number of experimentally validated miRNA targets enables predicting miRNA targets with high precision via accurate data analysis. Taking advantage of these known miRNA targets, a novel recommendation system model (miRTMC) for miRNA target prediction is established using a new matrix completion algorithm. In miRTMC, a heterogeneous network is constructed by integrating the miRNA similarity network, the gene similarity network, and the miRNA-gene interaction network. Our assumption is that the latent factors determining whether a gene is the target of miRNA or not are highly correlated, i.e., the adjacency matrix of the heterogeneous network is low-rank, which is then completed by using a nuclear norm regularized linear least squares model under non-negative constraints. Alternating direction method of multipliers (ADMM) is adopted to numerically solve the matrix completion problem. Our results show that miRTMC outperforms the competing methods in terms of various evaluation metrics. Our software package is available at https://github.com/hjiangcsu/miRTMC. Hui Jiang 0008, Mengyun Yang, Xiang Chen 0029, Min Li 0007, Yaohang Li, Jianxin Wang 0001 |
IEEE J. Biomed. Health Informatics | 2 |
| 2020 | Predicting Human lncRNA-Disease Associations Based on Geometric Matrix CompletionabstractRecently, increasing evidences reveal that dysregulations of long non-coding RNAs (lncRNAs) are relevant to diverse diseases. However, the number of experimentally verified lncRNA-disease associations is limited. Prioritizing potential associations is beneficial not only for disease diagnosis, but also disease treatment, more important apprehending disease mechanisms at lncRNA level. Various computational methods have been proposed, but precise prediction and full use of data's intrinsic structure are still challenging. In this work, we design a new method, denominated GMCLDA (Geometric Matrix Completion lncRNA-Disease Association), to infer underlying associations based on geometric matrix completion. Utilizing association patterns among functionally similar lncRNAs and phenotypically similar diseases, GMCLCA makes use of the intrinsic structure embedded in the association matrix. Besides, limiting the scope of the predicted values gives rise to a certain sparsity in computation and enhances the robustness of GMCLDA. GMCLDA computes disease semantic similarity according to the Disease Ontology (DO) hierarchy and lncRNA Gaussian interaction profile kernel similarity according to known interaction profiles. Then, GMCLDA measures lncRNA sequence similarity using Needleman-Wunsch algorithm. For a new lncRNA, GMCLDA prefills interaction profile on account of its K-nearest neighbors defined by sequence similarity. Finally, GMCLDA estimates the missing entries of the association matrix based on geometric matrix completion model. Compared with state-of-the-art methods, GMCLDA can provide more accurate lncRNA-disease prediction. Further case studies prove that GMCLDA is able to correctly infer possible lncRNAs for renal cancer. Chengqian Lu, Mengyun Yang, Min Li 0007, Yaohang Li, Fang-Xiang Wu, Jianxin Wang 0001 |
IEEE J. Biomed. Health Informatics | 2 |
| 2019 | Drug and disease similarity calculation platform for drug repositioningabstractDrug repositioning, aiming to infer potential indications for drugs efficiently, has achieved remarkable results in reducing the cycle, cost and risk of drug Research and Development (R&D), and mining new uses of known drugs. Currently, many computational drug repositioning strategies have been proposed. The similarity calculation, as one of the key steps of drug repositioning, has an important impact on the accuracy of computational drug repositioning. However, the biological data used for the similarity calculation come from a wide range of sources with different formats, and similarity calculation methods are developed in different programming languages, thus the similarity calculation methods are varying. To facilitate the similarity calculation for drug repositioning, we developed a computational platform consisting of various datasets and similarity measures for drugs and diseases, and four programming languages (Java, R, Python and MATLAB) are supported by our platform. Users can use relevant data and methods directly according to their needs and customize similarity calculation methods. The platform is available at: http://bioinformatics.csu.edu.cn/artemis/. Huimin Luo, Mengyun Yang, Fang-Xiang Wu, Jianxin Wang 0001 |
BIBM | 3 |
| 2019 | AttentionDTA: prediction of drug-target binding affinity using attention modelabstractIn bioinformatics, machine learning-based prediction of drug-target interaction (DTI) plays an important role in virtual screening of drug discovery. DTI prediction, which have been treated as a binary classification problem, depends on the concentration of two molecules, the interaction between two molecules, and other factors. The degree of affinity between a drug molecule (such as a drug compound) and a target molecule (such as a receptor or protein kinase) reflects how tightly the drug binds to a particular target and is quantified by the measurement which can reflect more detailed and specific information than binary relationship. In this study, we proposed an end-to-end model, named AttentionDTA, based on deep learning, which associates attention mechanism to predict the binding affinity of DTI. The novelty in this work is to use attentional mechanisms to consider which subsequences in a protein are more important for a drug and which subsequences in a drug are more important for a protein when predicting its affinity. So that the representational ability of the model is stronger. The model uses one-dimensional Convolution Neural Networks (1D-CNNs) to extract the abstract information of drug and protein, and makes the drug and protein representations mutually adapt through the attention mechanisms. We evaluate our model on two established drug-target affinity benchmark datasets, Davis and KIBA. The model outperforms DeepDTA, a state-of-the-art deep learning method for drug-target binding affinity prediction, with better Mean Squared Error (MSE), Concordance Index (CI), rm2, and Area Under Precision Recall Curve (AUPR). Our results show that the attention-based model can effectively extract effective representations by calculating the weight of the representation between the drug and the protein. Finally, we visualize the attention weight. It proves our model can obtain the information of binding sites. Qichang Zhao, Fen Xiao, Mengyun Yang, Yaohang Li, Jianxin Wang 0001 |
BIBM | 3 |
| 2019 | Drug repositioning based on bounded nuclear norm regularizationabstractMOTIVATION: Computational drug repositioning is a cost-effective strategy to identify novel indications for existing drugs. Drug repositioning is often modeled as a recommendation system problem. Taking advantage of the known drug-disease associations, the objective of the recommendation system is to identify new treatments by filling out the unknown entries in the drug-disease association matrix, which is known as matrix completion. Underpinned by the fact that common molecular pathways contribute to many different diseases, the recommendation system assumes that the underlying latent factors determining drug-disease associations are highly correlated. In other words, the drug-disease matrix to be completed is low-rank. Accordingly, matrix completion algorithms efficiently constructing low-rank drug-disease matrix approximations consistent with known associations can be of immense help in discovering the novel drug-disease associations. RESULTS: In this article, we propose to use a bounded nuclear norm regularization (BNNR) method to complete the drug-disease matrix under the low-rank assumption. Instead of strictly fitting the known elements, BNNR is designed to tolerate the noisy drug-drug and disease-disease similarities by incorporating a regularization term to balance the approximation error and the rank properties. Moreover, additional constraints are incorporated into BNNR to ensure that all predicted matrix entry values are within the specific interval. BNNR is carried out on an adjacency matrix of a heterogeneous drug-disease network, which integrates the drug-drug, drug-disease and disease-disease networks. It not only makes full use of available drugs, diseases and their association information, but also is capable of dealing with cold start naturally. Our computational results show that BNNR yields higher drug-disease association prediction accuracy than the current state-of-the-art methods. The most significant gain is in prediction precision measured as the fraction of the positive predictions that are truly positive, which is particularly useful in drug design practice. Cases studies also confirm the accuracy and reliability of BNNR. AVAILABILITY AND IMPLEMENTATION: The code of BNNR is freely available at https://github.com/BioinformaticsCSU/BNNR. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Mengyun Yang, Huimin Luo, Yaohang Li, Jianxin Wang 0001 |
Bioinform. | 1 |
| 2019 | Overlap matrix completion for predicting drug-associated indicationsabstractIdentification of potential drug-associated indications is critical for either approved or novel drugs in drug repositioning. Current computational methods based on drug similarity and disease similarity have been developed to predict drug-disease associations. When more reliable drug- or disease-related information becomes available and is integrated, the prediction precision can be continuously improved. However, it is a challenging problem to effectively incorporate multiple types of prior information, representing different characteristics of drugs and diseases, to identify promising drug-disease associations. In this study, we propose an overlap matrix completion (OMC) for bilayer networks (OMC2) and tri-layer networks (OMC3) to predict potential drug-associated indications, respectively. OMC is able to efficiently exploit the underlying low-rank structures of the drug-disease association matrices. In OMC2, first of all, we construct one bilayer network from drug-side aspect and one from disease-side aspect, and then obtain their corresponding block adjacency matrices. We then propose the OMC2 algorithm to fill out the values of the missing entries in these two adjacency matrices, and predict the scores of unknown drug-disease pairs. Moreover, we further extend OMC2 to OMC3 to handle tri-layer networks. Computational experiments on various datasets indicate that our OMC methods can effectively predict the potential drug-disease associations. Compared with the other state-of-the-art approaches, our methods yield higher prediction accuracy in 10-fold cross-validation and de novo experiments. In addition, case studies also confirm the effectiveness of our methods in identifying promising indications for existing drugs in practical applications. Mengyun Yang, Huimin Luo, Yaohang Li, Fang-Xiang Wu, Jianxin Wang 0001 |
PLoS Comput. Biol. | 1 |
| 2018 | Prediction of lncRNA-disease associations based on inductive matrix completionabstractMotivation: Accumulating evidences indicate that long non-coding RNAs (lncRNAs) play pivotal roles in various biological processes. Mutations and dysregulations of lncRNAs are implicated in miscellaneous human diseases. Predicting lncRNA-disease associations is beneficial to disease diagnosis as well as treatment. Although many computational methods have been developed, precisely identifying lncRNA-disease associations, especially for novel lncRNAs, remains challenging. Results: In this study, we propose a method (named SIMCLDA) for predicting potential lncRNA-disease associations based on inductive matrix completion. We compute Gaussian interaction profile kernel of lncRNAs from known lncRNA-disease interactions and functional similarity of diseases based on disease-gene and gene-gene onotology associations. Then, we extract primary feature vectors from Gaussian interaction profile kernel of lncRNAs and functional similarity of diseases by principal component analysis, respectively. For a new lncRNA, we calculate the interaction profile according to the interaction profiles of its neighbors. At last, we complete the association matrix based on the inductive matrix completion framework using the primary feature vectors from the constructed feature matrices. Computational results show that SIMCLDA can effectively predict lncRNA-disease associations with higher accuracy compared with previous methods. Furthermore, case studies show that SIMCLDA can effectively predict candidate lncRNAs for renal cancer, gastric cancer and prostate cancer. Availability and implementation: https://github.com//bioinfomaticsCSU/SIMCLDA. Supplementary information: Supplementary data are available at Bioinformatics online. Chengqian Lu, Mengyun Yang, Feng Luo 0001, Fang-Xiang Wu, Min Li 0007, Yi Pan 0001, Yaohang Li, Jianxin Wang 0001 |
Bioinform. | 2 |