Chengqian Lu

dblp:193/8023 · DBLP profile ↗
← Back
19ranked-venue papers
6as first author
15since 2021 · last 2026
0000-0001-9201-6912ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 17 · 5 first-author · 13 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Predicting Drug Synergy via Cross-modal Contrastive Learning and Masked Multi-omics Hypergraphs
Xinqiang Wen, Chengqian Lu, Xixi Yang, Xuan Lin, Xiangmao Meng
ISBRA (1)3
2026 Multi-scale cross-attention integrates dynamic and static features for protein-RNA prediction
Chengqian Lu, Xiangmao Meng, Min Zeng 0004, Yi Pan 0001, Jianxin Wang 0001
Pattern Recognit.1
2026 MotifGT-DTI: Pivotal Motif-Based Graph Transformer Model Improves Drug-Target Interaction Prediction
abstract
Predicting drug-target interactions (DTIs) plays an essential role in drug discovery and drug repurposing. Although significant performance improvements have been achieved in DTI prediction, existing methods have not fully explored the properties of protein and drug molecular structure to make results interpretable. In this study, we propose MotifGT-DTI, a novel motif-based model with a graph transformer (GT) for DTI prediction. Specifically, MotifGT-DTI captures complex molecular patterns of drug molecular graph motifs and protein 3-D pocket subgraphs with GT. To attain protein characteristics more comprehensively, MotifGT-DTI fuses 1-D sequence and 3-D structure features with cross-attention from two views. Then, the structural-level association patterns of drug molecules and proteins are connected via a bilinear attention network. Experimental results show that MotifGT-DTI achieves the best accuracy compared to state-of-the-art baselines on four public datasets. In the three cold-start scenarios, the prediction results provided by our method are competitive in accuracy, generalization ability, and stability, highlighting its promising potential for practical applications. Furthermore, the visualization study demonstrates that MotifGT-DTI finds functional molecular motifs and provides interpretability for predicted results. The datasets and codes are publicly available at https://github.com/Dimpleney/MotifGT-DTI.
Min Zeng 0004, Jianxin Wang 0001, Chengqian Lu
IEEE Trans. Neural Networks Learn. Syst.4
2025 RNALoc-LM: RNA subcellular localization prediction using pre-trained RNA language model
abstract
MOTIVATION: Accurately predicting RNA subcellular localization is crucial for understanding the cellular functions and regulatory mechanisms of RNAs. Although many computational methods have been developed to predict the subcellular localization of lncRNAs, miRNAs, and circRNAs, very few of them are designed to simultaneously predict the subcellular localization of multiple types of RNAs. In addition, the emergence of pre-trained RNA language model has shown remarkable performance in various bioinformatics tasks, such as structure prediction and functional annotation. Despite these advancements, there remains a significant gap in applying pre-trained RNA language models specifically for predicting RNA subcellular localization. RESULTS: In this study, we proposed RNALoc-LM, the first interpretable deep-learning framework that leverages a pre-trained RNA language model for predicting RNA subcellular localization. RNALoc-LM uses a pre-trained RNA language model to encode RNA sequences, then captures local patterns and long-range dependencies through TextCNN and BiLSTM modules. A multi-head attention mechanism is used to focus on important regions within the RNA sequences. The results demonstrate that RNALoc-LM significantly outperforms both deep-learning baselines and existing state-of-the-art predictors. Additionally, motif analysis highlights RNALoc-LM's potential for discovering important motifs, while an ablation study confirms the effectiveness of the RNA sequence embeddings generated by the pre-trained RNA language model. AVAILABILITY AND IMPLEMENTATION: The RNALoc-LM web server is available at http://csuligroup.com:8000/RNALoc-LM. The source code can be obtained from https://github.com/CSUBioGroup/RNALoc-LM.
Min Zeng 0004, Chengqian Lu, Rui Yin 0002, Fei Guo 0001, Min Li 0007
Bioinform.4
2025 2OMe-LM: predicting 2′-O-methylation sites in human RNA using a pre-trained RNA language model
abstract
MOTIVATION: 2'-O-methylation (2OMe) is a common post-transcriptional modification in RNA that plays a crucial role in regulating gene expression and is implicated in various biological processes and diseases. Computational methods offer an efficient alternative to the time-consuming and costly experimental identification of 2OMe sites. Recent advancements in RNA pre-trained language models have revolutionized RNA bioinformatics. However, there remains a gap in their application specifically for predicting 2OMe sites. RESULTS: In the study, we propose a novel deep learning framework, 2OMe-LM, for predicting 2OMe sites in RNA. 2OMe-LM integrates RNA sequence features derived from RNA pre-trained language models with those obtained from the word2vec technique. Then, 2OMe-LM employs fully connected layers and a bidirectional long short-term memory network to process the two types of features separately, followed by a feature fusion module for the final prediction. Additionally, an attention block is incorporated to provide the interpretability of the prediction results. The results demonstrate that 2OMe-LM significantly outperforms existing state-of-the-art predictors, with features from RNA pre-trained language models proving to be critical. Motif analysis further demonstrates 2OMe-LM's potential for discovering 2OMe-related motifs. AVAILABILITY AND IMPLEMENTATION: The 2OMe-LM web server is available at https://csuligroup.com:9200/2OMe-LM. The source code can be obtained from https://github.com/CSUBioGroup/2OMe-LM.
Qianpei Liu, Min Zeng 0004, Chengqian Lu, Shichao Kan, Fei Guo 0001, Min Li 0007
Bioinform.4
2025 CellCircLoc: Deep Neural Network for Predicting and Explaining Cell Line-Specific CircRNA Subcellular Localization
abstract
The subcellular localization of circular RNAs (circRNAs) is crucial for understanding their functional relevance and regulatory mechanisms. CircRNA subcellular localization exhibits variations across different cell lines, demonstrating the diversity and complexity of circRNA regulation within distinct cellular contexts. However, existing computational methods for predicting circRNA subcellular localization often ignore the importance of cell line specificity and instead train a general model on aggregated data from all cell lines. Considering the diversity and context-dependent behavior of circRNAs across different cell lines, it is imperative to develop cell line-specific models to accurately predict circRNA subcellular localization. In the study, we proposed CellCircLoc, a sequence-based deep learning model for circRNA subcellular localization prediction, which is trained for different cell lines. CellCircLoc utilizes a combination of convolutional neural networks, Transformer blocks, and bidirectional long short-term memory to capture both sequence local features and long-range dependencies within the sequences. In the Transformer blocks, CellCircLoc uses an attentive convolution mechanism to capture the importance of individual nucleotides. Extensive experiments demonstrate the effectiveness of CellCircLoc in accurately predicting circRNA subcellular localization across different cell lines, outperforming other computational models that do not consider cell line specificity. Moreover, the interpretability of CellCircLoc facilitates the discovery of important motifs associated with circRNA subcellular localization.
Min Zeng 0004, Jingwei Lu, Chengqian Lu, Shichao Kan, Fei Guo 0001, Min Li 0007
IEEE J. Biomed. Health Informatics4
2023 GraphLncLoc: long non-coding RNA subcellular localization prediction using graph convolutional networks based on sequence to graph transformation
abstract
The subcellular localization of long non-coding RNAs (lncRNAs) is crucial for understanding lncRNA functions. Most of existing lncRNA subcellular localization prediction methods use k-mer frequency features to encode lncRNA sequences. However, k-mer frequency features lose sequence order information and fail to capture sequence patterns and motifs of different lengths. In this paper, we proposed GraphLncLoc, a graph convolutional network-based deep learning model, for predicting lncRNA subcellular localization. Unlike previous studies encoding lncRNA sequences by using k-mer frequency features, GraphLncLoc transforms lncRNA sequences into de Bruijn graphs, which transforms the sequence classification problem into a graph classification problem. To extract the high-level features from the de Bruijn graph, GraphLncLoc employs graph convolutional networks to learn latent representations. Then, the high-level feature vectors derived from de Bruijn graph are fed into a fully connected layer to perform the prediction task. Extensive experiments show that GraphLncLoc achieves better performance than traditional machine learning models and existing predictors. In addition, our analyses show that transforming sequences into graphs has more distinguishable features and is more robust than k-mer frequency features. The case study shows that GraphLncLoc can uncover important motifs for nucleus subcellular localization. GraphLncLoc web server is available at http://csuligroup.com:8000/GraphLncLoc/.
Min Li 0007, Baoying Zhao, Rui Yin 0002, Chengqian Lu, Fei Guo 0001, Min Zeng 0004
Briefings Bioinform.4
2023 Inferring disease-associated circRNAs by multi-source aggregation based on heterogeneous graph neural network
abstract
Emerging evidence has proved that circular RNAs (circRNAs) are implicated in pathogenic processes. They are regarded as promising biomarkers for diagnosis due to covalently closed loop structures. As opposed to traditional experiments, computational approaches can identify circRNA-disease associations at a lower cost. Aggregating multi-source pathogenesis data helps to alleviate data sparsity and infer potential associations at the system level. The majority of computational approaches construct a homologous network using multi-source data, but they lose the heterogeneity of the data. Effective methods that use the features of multi-source data are considered as a matter of urgency. In this paper, we propose a model (CDHGNN) based on edge-weighted graph attention and heterogeneous graph neural networks for potential circRNA-disease association prediction. The circRNA network, micro RNA network, disease network and heterogeneous network are constructed based on multi-source data. To reflect association probabilities between nodes, an edge-weighted graph attention network model is designed for node features. To assign attention weights to different types of edges and learn contextual meta-path, CDHGNN infers potential circRNA-disease association based on heterogeneous neural networks. CDHGNN outperforms state-of-the-art algorithms in terms of accuracy. Edge-weighted graph attention networks and heterogeneous graph networks have both improved performance significantly. Furthermore, case studies suggest that CDHGNN is capable of identifying specific molecular associations and investigating biomolecular regulatory relationships in pathogenesis. The code of CDHGNN is freely available at https://github.com/BioinformaticsCSU/CDHGNN.
Chengqian Lu, Lishen Zhang, Min Zeng 0004, Wei Lan 0001, Guihua Duan, Jianxin Wang 0001
Briefings Bioinform.1
2023 CRMSS: predicting circRNA-RBP binding sites based on multi-scale characterizing sequence and structure features
abstract
Circular RNAs (circRNAs) are reverse-spliced and covalently closed RNAs. Their interactions with RNA-binding proteins (RBPs) have multiple effects on the progress of many diseases. Some computational methods are proposed to identify RBP binding sites on circRNAs but suffer from insufficient accuracy, robustness and explanation. In this study, we first take the characteristics of both RNA and RBP into consideration. We propose a method for discriminating circRNA-RBP binding sites based on multi-scale characterizing sequence and structure features, called CRMSS. For circRNAs, we use sequence ${k}\hbox{-}{mer}$ embedding and the forming probabilities of local secondary structures as features. For RBPs, we combine sequence and structure frequencies of RNA-binding domain regions to generate features. We capture binding patterns with multi-scale residual blocks. With BiLSTM and attention mechanism, we obtain the contextual information of high-level representation for circRNA-RBP binding. To validate the effectiveness of CRMSS, we compare its predictive performance with other methods on 37 RBPs. Taking the properties of both circRNAs and RBPs into account, CRMSS achieves superior performance over state-of-the-art methods. In the case study, our model provides reliable predictions and correctly identifies experimentally verified circRNA-RBP pairs. The code of CRMSS is freely available at https://github.com/BioinformaticsCSU/CRMSS.
Lishen Zhang, Chengqian Lu, Min Zeng 0004, Yaohang Li, Jianxin Wang 0001
Briefings Bioinform.2
2023 LncLocFormer: a Transformer-based deep learning model for multi-label lncRNA subcellular localization prediction by using localization-specific attention mechanism
abstract
MOTIVATION: There is mounting evidence that the subcellular localization of lncRNAs can provide valuable insights into their biological functions. In the real world of transcriptomes, lncRNAs are usually localized in multiple subcellular localizations. Furthermore, lncRNAs have specific localization patterns for different subcellular localizations. Although several computational methods have been developed to predict the subcellular localization of lncRNAs, few of them are designed for lncRNAs that have multiple subcellular localizations, and none of them take motif specificity into consideration. RESULTS: In this study, we proposed a novel deep learning model, called LncLocFormer, which uses only lncRNA sequences to predict multi-label lncRNA subcellular localization. LncLocFormer utilizes eight Transformer blocks to model long-range dependencies within the lncRNA sequence and shares information across the lncRNA sequence. To exploit the relationship between different subcellular localizations and find distinct localization patterns for different subcellular localizations, LncLocFormer employs a localization-specific attention mechanism. The results demonstrate that LncLocFormer outperforms existing state-of-the-art predictors on the hold-out test set. Furthermore, we conducted a motif analysis and found LncLocFormer can capture known motifs. Ablation studies confirmed the contribution of the localization-specific attention mechanism in improving the prediction performance. AVAILABILITY AND IMPLEMENTATION: The LncLocFormer web server is available at http://csuligroup.com:9000/LncLocFormer. The source code can be obtained from https://github.com/CSUBioGroup/LncLocFormer.
Min Zeng 0004, Yifan Wu 0008, Rui Yin 0002, Chengqian Lu, Junwen Duan, Min Li 0007
Bioinform.5
2022 DeepLncLoc: a deep learning framework for long non-coding RNA subcellular localization prediction based on subsequence embedding
abstract
Long non-coding RNAs (lncRNAs) are a class of RNA molecules with more than 200 nucleotides. A growing amount of evidence reveals that subcellular localization of lncRNAs can provide valuable insights into their biological functions. Existing computational methods for predicting lncRNA subcellular localization use k-mer features to encode lncRNA sequences. However, the sequence order information is lost by using only k-mer features. We proposed a deep learning framework, DeepLncLoc, to predict lncRNA subcellular localization. In DeepLncLoc, we introduced a new subsequence embedding method that keeps the order information of lncRNA sequences. The subsequence embedding method first divides a sequence into some consecutive subsequences and then extracts the patterns of each subsequence, last combines these patterns to obtain a complete representation of the lncRNA sequence. After that, a text convolutional neural network is employed to learn high-level features and perform the prediction task. Compared with traditional machine learning models, popular representation methods and existing predictors, DeepLncLoc achieved better performance, which shows that DeepLncLoc could effectively predict lncRNA subcellular localization. Our study not only presented a novel computational model for predicting lncRNA subcellular localization but also introduced a new subsequence embedding method which is expected to be applied in other sequence-based prediction tasks. The DeepLncLoc web server is freely accessible at http://bioinformatics.csu.edu.cn/DeepLncLoc/, and source code and datasets can be downloaded from https://github.com/CSUBioGroup/DeepLncLoc.
Min Zeng 0004, Yifan Wu 0008, Chengqian Lu, Fuhao Zhang, Fang-Xiang Wu, Min Li 0007
Briefings Bioinform.3
2021 Improving circRNA-disease association prediction by sequence and ontology representations with convolutional and recurrent neural networks
abstract
MOTIVATION: Emerging studies indicate that circular RNAs (circRNAs) are widely involved in the progression of human diseases. Due to its special structure which is stable, circRNAs are promising diagnostic and prognostic biomarkers for diseases. However, the experimental verification of circRNA-disease associations is expensive and limited to small-scale. Effective computational methods for predicting potential circRNA-disease associations are regarded as a matter of urgency. Although several models have been proposed, over-reliance on known associations and the absence of characteristics of biological functions make precise predictions are still challenging. RESULTS: In this study, we propose a method for predicting CircRNA-disease associations based on sequence and ontology representations, named CDASOR, with convolutional and recurrent neural networks. For sequences of circRNAs, we encode them with continuous k-mers, get low-dimensional vectors of k-mers, extract their local feature vectors with 1D CNN and learn their long-term dependencies with bi-directional long short-term memory. For diseases, we serialize disease ontology into sentences containing the hierarchy of ontology, obtain low-dimensional vectors for disease ontology terms and get terms' dependencies. Furthermore, we get association patterns of circRNAs and diseases from known circRNA-disease associations with neural networks. After the above steps, we get circRNAs' and diseases' high-level representations, which are informative to improve the prediction. The experimental results show that CDASOR provides an accurate prediction. Importing the characteristics of biological functions, CDASOR achieves impressive predictions in the de novo test. In addition, 6 of the top-10 predicted results are verified by the published literature in the case studies. AVAILABILITY AND IMPLEMENTATION: The code and data of CDASOR are freely available at https://github.com/BioinformaticsCSU/CDASOR.
Chengqian Lu, Min Zeng 0004, Fang-Xiang Wu, Min Li 0007, Jianxin Wang 0001
Bioinform.1
2021 Heterogeneous graph inference with matrix completion for computational drug repositioning
abstract
MOTIVATION: Emerging evidence presents that traditional drug discovery experiment is time-consuming and high costs. Computational drug repositioning plays a critical role in saving time and resources for drug research and discovery. Therefore, developing more accurate and efficient approaches is imperative. Heterogeneous graph inference is a classical method in computational drug repositioning, which not only has high convergence precision, but also has fast convergence speed. However, the method has not fully considered the sparsity of heterogeneous association network. In addition, rough similarity measure can reduce the performance in identifying drug-associated indications. RESULTS: In this article, we propose a heterogeneous graph inference with matrix completion (HGIMC) method to predict potential indications for approved and novel drugs. First, we use a bounded matrix completion (BMC) model to prefill a part of the missing entries in original drug-disease association matrix. This step can add more positive and formative drug-disease edges between drug network and disease network. Second, Gaussian radial basis function (GRB) is employed to improve the drug and disease similarities since the performance of heterogeneous graph inference more relies on similarity measures. Next, based on the updated drug-disease associations and new similarity measures of drug and disease, we construct a novel heterogeneous drug-disease network. Finally, HGIMC utilizes the heterogeneous network to infer the scores of unknown association pairs, and then recommend the promising indications for drugs. To evaluate the performance of our method, HGIMC is compared with five state-of-the-art approaches of drug repositioning in the 10-fold cross-validation and de novo tests. As the numerical results shown, HGIMC not only achieves a better prediction performance but also has an excellent computation efficiency. In addition, cases studies also confirm the effectiveness of our method in practical application. AVAILABILITYAND IMPLEMENTATION: The HGIMC software and data are freely available at https://github.com/BioinformaticsCSU/HGIMC, https://hub.docker.com/repository/docker/yangmy84/hgimc and http://doi.org/10.5281/zenodo.4285640. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Mengyun Yang, Yunpei Xu, Chengqian Lu, Jianxin Wang 0001
Bioinform.4
2021 DMFLDA: A Deep Learning Framework for Predicting lncRNA-Disease Associations
abstract
A growing amount of evidence suggests that long non-coding RNAs (lncRNAs) play important roles in the regulation of biological processes in many human diseases. However, the number of experimentally verified lncRNA-disease associations is very limited. Thus, various computational approaches are proposed to predict lncRNA-disease associations. Current matrix factorization-based methods cannot capture the complex non-linear relationship between lncRNAs and diseases, and traditional machine learning-based methods are not sufficiently powerful to learn the representation of lncRNAs and diseases. Considering these limitations in existing computational methods, we propose a deep matrix factorization model to predict lncRNA-disease associations (DMFLDA in short). DMFLDA uses a cascade of non-linear hidden layers to learn latent representation to represent lncRNAs and diseases. By using non-linear hidden layers, DMFLDA captures the more complex non-linear relationship between lncRNAs and diseases than traditional matrix factorization-based methods. In addition, DMFLDA learns features directly from the lncRNA-disease interaction matrix and thus can obtain more accurate representation learning for lncRNAs and diseases than traditional machine learning methods. The low dimensional representations of the lncRNAs and diseases are fused to estimate the new interaction value. To evaluate the performance of DMFLDA, we perform leave-one-out cross-validation and 5-fold cross-validation on known experimentally verified lncRNA-disease associations. The experimental results show that DMFLDA performs better than the existing methods. The case studies show that many predicted interactions of colorectal cancer, prostate cancer, and renal cancer have been verified by recent biomedical literature. The source code and datasets can be obtained from https://github.com/CSUBioGroup/DMFLDA.
Min Zeng 0004, Chengqian Lu, Zhihui Fei, Fang-Xiang Wu, Yaohang Li, Jianxin Wang 0001, Min Li 0007
IEEE ACM Trans. Comput. Biol. Bioinform.2
2021 Deep Matrix Factorization Improves Prediction of Human CircRNA-Disease Associations
abstract
In recent years, more and more evidence indicates that circular RNAs (circRNAs) with covalently closed loop play various roles in biological processes. Dysregulation and mutation of circRNAs may be implicated in diseases. Due to its stable structure and resistance to degradation, circRNAs provide great potential to be diagnostic biomarkers. Therefore, predicting circRNA-disease associations is helpful in disease diagnosis. However, there are few experimentally validated associations between circRNAs and diseases. Although several computational methods have been proposed, precisely representing underlying features and grasping the complex structures of data are still challenging. In this paper, we design a new method, called DMFCDA (Deep Matrix Factorization CircRNA-Disease Association), to infer potential circRNA-disease associations. DMFCDA takes both explicit and implicit feedback into account. Then, it uses a projection layer to automatically learn latent representations of circRNAs and diseases. With multi-layer neural networks, DMFCDA can model the non-linear associations to grasp the complex structure of data. We assess the performance of DMFCDA using leave-one cross-validation and 5-fold cross-validation on two datasets. Computational results show that DMFCDA efficiently infers circRNA-disease associations according to AUC values, the percentage of precisely retrieved associations in various top ranks, and statistical comparison. We also conduct case studies to evaluate DMFCDA. All results show that DMFCDA provides accurate predictions.
Chengqian Lu, Min Zeng 0004, Fuhao Zhang, Fang-Xiang Wu, Min Li 0007, Jianxin Wang 0001
IEEE J. Biomed. Health Informatics1
2020 Predicting Human lncRNA-Disease Associations Based on Geometric Matrix Completion
abstract
Recently, increasing evidences reveal that dysregulations of long non-coding RNAs (lncRNAs) are relevant to diverse diseases. However, the number of experimentally verified lncRNA-disease associations is limited. Prioritizing potential associations is beneficial not only for disease diagnosis, but also disease treatment, more important apprehending disease mechanisms at lncRNA level. Various computational methods have been proposed, but precise prediction and full use of data's intrinsic structure are still challenging. In this work, we design a new method, denominated GMCLDA (Geometric Matrix Completion lncRNA-Disease Association), to infer underlying associations based on geometric matrix completion. Utilizing association patterns among functionally similar lncRNAs and phenotypically similar diseases, GMCLCA makes use of the intrinsic structure embedded in the association matrix. Besides, limiting the scope of the predicted values gives rise to a certain sparsity in computation and enhances the robustness of GMCLDA. GMCLDA computes disease semantic similarity according to the Disease Ontology (DO) hierarchy and lncRNA Gaussian interaction profile kernel similarity according to known interaction profiles. Then, GMCLDA measures lncRNA sequence similarity using Needleman-Wunsch algorithm. For a new lncRNA, GMCLDA prefills interaction profile on account of its K-nearest neighbors defined by sequence similarity. Finally, GMCLDA estimates the missing entries of the association matrix based on geometric matrix completion model. Compared with state-of-the-art methods, GMCLDA can provide more accurate lncRNA-disease prediction. Further case studies prove that GMCLDA is able to correctly infer possible lncRNAs for renal cancer.
Chengqian Lu, Mengyun Yang, Min Li 0007, Yaohang Li, Fang-Xiang Wu, Jianxin Wang 0001
IEEE J. Biomed. Health Informatics1
2019 LncRNA-disease association prediction through combining linear and non-linear features with matrix factorization and deep learning techniques
abstract
Long non-coding RNAs (lncRNAs) are the foundation for understanding mechanisms of many human diseases. Considering the limited number of known experimentally verified associations between lncRNAs and diseases, it is appealing to develop accurate and effective computational methods to identify lncRNA-disease associations. Conventional matrix factorization-based methods cannot model complicated associations between lncRNAs and diseases. In this study, we propose a novel computational framework, through combining linear and non-linear features, which is used for lncRNA-disease association prediction. In our model, a conventional matrix factorization method is applied to extract linear features between lncRNAs and diseases. Deep learning techniques (fully connected layers) are applied to extract nonlinear features between lncRNAs and diseases. Finally, linear and non-linear features are fused to improve predictive performance. Compared to previous studies, our model can take advantages of the combination of linear and non-linear features between lncRNAs and diseases, and thus can effectively identify potential lncRNA-disease associations. The results show that our method achieves state-of-the-art performance in the leave-one-out cross-validation. The source codes of our method can be found at https://github.com/CSUBioGroup/DMFLDA2.
Min Zeng 0004, Chengqian Lu, Fuhao Zhang, Zhangli Lu, Fang-Xiang Wu, Yaohang Li, Min Li 0007
BIBM2
2018 Prediction of lncRNA-disease associations based on inductive matrix completion
abstract
Motivation: Accumulating evidences indicate that long non-coding RNAs (lncRNAs) play pivotal roles in various biological processes. Mutations and dysregulations of lncRNAs are implicated in miscellaneous human diseases. Predicting lncRNA-disease associations is beneficial to disease diagnosis as well as treatment. Although many computational methods have been developed, precisely identifying lncRNA-disease associations, especially for novel lncRNAs, remains challenging. Results: In this study, we propose a method (named SIMCLDA) for predicting potential lncRNA-disease associations based on inductive matrix completion. We compute Gaussian interaction profile kernel of lncRNAs from known lncRNA-disease interactions and functional similarity of diseases based on disease-gene and gene-gene onotology associations. Then, we extract primary feature vectors from Gaussian interaction profile kernel of lncRNAs and functional similarity of diseases by principal component analysis, respectively. For a new lncRNA, we calculate the interaction profile according to the interaction profiles of its neighbors. At last, we complete the association matrix based on the inductive matrix completion framework using the primary feature vectors from the constructed feature matrices. Computational results show that SIMCLDA can effectively predict lncRNA-disease associations with higher accuracy compared with previous methods. Furthermore, case studies show that SIMCLDA can effectively predict candidate lncRNAs for renal cancer, gastric cancer and prostate cancer. Availability and implementation: https://github.com//bioinfomaticsCSU/SIMCLDA. Supplementary information: Supplementary data are available at Bioinformatics online.
Chengqian Lu, Mengyun Yang, Feng Luo 0001, Fang-Xiang Wu, Min Li 0007, Yi Pan 0001, Yaohang Li, Jianxin Wang 0001
Bioinform.1
2016 Predicting microRNA-environmental factor interactions based on bi-random walk and multi-label learning
abstract
Increasing evidences have shown that microRNAs (miRNAs) play important roles in many diseases. The environmental factors (EFs) can regulate the expression level of miRNAs in human tissues. Therefore, identifying potential miRNA-environmental factor interactions is helpful not only for understanding the pathogenesis of diseases, but also for disease diagnosis, prognosis and treatment. In this paper, we propose a computational framework, MEI-BRWMLL (MiRNA-EF Interaction prediction based on Bi-Random walk and Multi-Label Learning), to identify interactions between miRNAs and environmental factors. The sequence and topology information of miRNA and structure, anatomical therapeutic chemical and topology information of environmental factor are employed to measure similarity of miRNAs and environmental factors, respectively. In addition, we use similarity network fusion method to integrate biological information of miRNAs and environmental factors, respectively. In the last, the bi-random walk and multi-label learning method are utilized to identify potential miRNA-environmental factor interactions. In order to evaluate the performance of MEI-BRWMLL, we implement the ten-fold cross validation in the experiment. The MEI-BRWMLL achieves an AUC of 0.8208. It has been shown that MEI-BRWMLL is able to identify known miRNA-environmental factor interactions.
Wei Lan 0001, Jianxin Wang 0001, Min Li 0007, Chengqian Lu, Fang-Xiang Wu, Yi Pan 0001
BIBM4