VLDB 2026 Research / reviewers in the wild / expert
Xing Chen 0001
dblp:89/120-1
· DBLP profile ↗
53ranked-venue papers
21as first author
25since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 48 · 21 first-author · 25 since 2021Artificial intelligence and machine learning · 5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Guest Editorial:Application of Computational Techniques in Drug Discovery and Disease Treatment
Xing Chen 0001, Qi Zhao 0010 |
IEEE J. Biomed. Health Informatics | 1 |
| 2025 | Exploring Microbe-Drug Association Prediction via Multi-Attribute Dual-Decoder Graph AutoencoderabstractPredicting potential microbe-drug associations (MDA) can help study pathogenesis, expedite pharmaceutical innovation, and enhance targeted therapeutics. Given the time and labor intensity of traditional biological experiments, an increasing number of computational approaches are being employed to predict MDA. The method based on graph embedding is one of the most widely used. However, most of these methods only consider node embedding or graph structure information in isolation, which leads to restricted predictive accuracy. In this work, we propose a method called exploring microbe-drug association prediction via multi-attribute dual-decoder graph autoencoder (MDGAEMDA). Specifically, a heterogeneous network containing microbe similarity, drug similarity, and known associations is constructed. Second, to enrich the node information, the multi-attribute features are obtained by importing the topological information of microbe and drug. Then, two heterogeneous networks constructed by the graph masking strategy are input into dual-decoder graph autoencoder that contains one encoder and two decoders (node decoder and structure decoder) to learn both node embedding and graph structure information. Finally, two low-dimensional features are spliced into the features of MDA pairs and predicted by random forest. The model was compared with multiple advanced methods using public datasets. The experimental outcomes showed that our model significantly outperformed other methods. The case study of widely used drugs demonstrated the reliability of the proposed method to predict MDA. Wei Liu 0150, Xiangcheng Deng, Xingen Sun, Xu Lu 0002, Xing Chen 0001 |
IEEE J. Biomed. Health Informatics | 5 |
| 2025 | DTI-MvSCA: An Anti-Over-Smoothing Multi-View Framework With Negative Sample Selection for Predicting Drug-Target InteractionsabstractPredicting potential drug-target interactions (DTIs) facilitates to accelerate drug discovery and reduce development cost. Current deep learning-based methods exhibit high-performance predictions, but three challenges remain: first, the absence of negative DTIs severely limits the model performance. Moreover, existing graph neural networks are beset with the scalability due to the model complexity and graph size. More importantly, most methods focus on learning the topological features while ignoring node features during DTI representation learning. To solve the limitations, here, we develop a multi-view neural network framework called DTI-MvSCA for DTI identification. This framework begins with constructing a drug-protein pair (DPP) network with matrix operation-based negative DTI selection, and then learns the DPP representations through aMulti-view neural network, finally classifies each DPP based on multilayer perceptron. Particularly, the multi-view neural network integrates graph topological feature learning based on the self-attention mechanism andSHADOW graph attention network, node feature learning based on 1DConvolutional neural network, and theAttention mechanism. An in-depth experiment on DrugBank V3.0 and V5.0 showed that DTI-MvSCA obtained precise and robust predictions against five state-of-the-art baseline methods. Furthermore, visualizing the feature distributions of the selected negative DTIs exhibits a more distinguishable and clearer boundary. In summary, DTI-MvSCA provides a useful deep learning tool to investigate potential DTIs. Lihong Peng, Zongzheng Bai, Longlong Liu, Xin Liu 0116, Min Chen 0028, Xing Chen 0001 |
IEEE J. Biomed. Health Informatics | 7 |
| 2024 | Computational model for drug researchabstractThis special issue focuses on computational model for drug research regarding drug bioactivity prediction, drug-related interaction prediction, modelling for immunotherapy and modelling for treatment of a specific disease, as conveyed by the following six research and four review articles. Notably, these 10 papers described a wide variety of in-depth drug research from the computational perspective and may represent a snapshot of the wide research landscape. Xing Chen 0001 |
Briefings Bioinform. | 1 |
| 2024 | Drug-drug interaction prediction: databases, web servers and computational modelsabstractIn clinical treatment, two or more drugs (i.e. drug combination) are simultaneously or successively used for therapy with the purpose of primarily enhancing the therapeutic efficacy or reducing drug side effects. However, inappropriate drug combination may not only fail to improve efficacy, but even lead to adverse reactions. Therefore, according to the basic principle of improving the efficacy and/or reducing adverse reactions, we should study drug-drug interactions (DDIs) comprehensively and thoroughly so as to reasonably use drug combination. In this review, we first introduced the basic conception and classification of DDIs. Further, some important publicly available databases and web servers about experimentally verified or predicted DDIs were briefly described. As an effective auxiliary tool, computational models for predicting DDIs can not only save the cost of biological experiments, but also provide relevant guidance for combination therapy to some extent. Therefore, we summarized three types of prediction models (including traditional machine learning-based models, deep learning-based models and score function-based models) proposed during recent years and discussed the advantages as well as limitations of them. Besides, we pointed out the problems that need to be solved in the future research of DDIs prediction and provided corresponding suggestions. Xing Chen 0001 |
Briefings Bioinform. | 5 |
| 2024 | CellDialog: A Computational Framework for Ligand-Receptor-Mediated Cell-Cell Communication AnalysisabstractIntercellular communication significantly influences tumor progression, metastasis, and therapy resistance. An intercellular communication inference method includes two main procedures: ligand-receptor interaction (LRI) curation and LRI-mediated intercellular communication strength measurement. The construction of a comprehensive, high-confident and well-organized LRI database contributes to intercellular communication inference. Here, we developed a computational framework named CellDialog to reconstruct an intercellular connectivity network based on the combined expression of ligands and receptors involved in sender and receiver cells. CellDialog first captures high-confident LRIs through LRI feature extraction, feature selection, and classification. Furthermore, CellDialog uses a three-point estimation approach to measure the LRI-mediated intercellular communication strength by combining LRI filtering and single-cell RNA sequencing data. A comparison analysis of CellDialog and the other tools was conducted, and it was found that CellDialog can efficiently decode intercellular communications. Additionally, CellDialog offers a heatmap view and network view for intercellular communication visualization. In summary, CellDialog provides a tool that allows researchers to analyze intercellular signal transduction. It is freely available at https://github.com/plhhnu/CellDialog. Lihong Peng, Chendi Han, Xing Chen 0001 |
IEEE J. Biomed. Health Informatics | 5 |
| 2023 | Computational model for disease researchabstractComputational analysis of vast public and private omics data [1–5] generated by high-throughput technologies [6, 7] aids in deciphering complex mechanisms [8–11] and relevant gene functions [12–15] of various diseases ranging from viral infections to cancers. Focusing on computational model for disease research, this special issue include the following nine research and three review articles: three of them related to understanding viral infections, two aimed to benefit cancer diagnosis, four seeking to uncover disease development mechanisms from bulk gene expression data, two with a similar objective achieved by modelling single-cell expression data, and the remaining one based on either bulk or single-cell data to identify critical states of diseases. Yin et al. [16] developed a machine learning framework for predicting variable-length linear B-cell epitopes of human-adapted viruses, which could potentially accelerate vaccine design. The capability of handling variable-length sequences was enabled by deploying QR decomposition and composition, transition, and distribution (CTD) descriptors to respectively encode Protvec representation of peptides and physicochemical properties of amino acids. The proposed framework was applicable to both traditional classifiers and deep learning models. Xing Chen 0001 |
Briefings Bioinform. | 1 |
| 2023 | MCFF-MTDDI: multi-channel feature fusion for multi-typed drug-drug interaction predictionabstractAdverse drug-drug interactions (DDIs) have become an increasingly serious problem in the medical and health system. Recently, the effective application of deep learning and biomedical knowledge graphs (KGs) have improved the DDI prediction performance of computational models. However, the problems of feature redundancy and KG noise also arise, bringing new challenges for researchers. To overcome these challenges, we proposed a Multi-Channel Feature Fusion model for multi-typed DDI prediction (MCFF-MTDDI). Specifically, we first extracted drug chemical structure features, drug pairs' extra label features, and KG features of drugs. Then, these different features were effectively fused by a multi-channel feature fusion module. Finally, multi-typed DDIs were predicted through the fully connected neural network. To our knowledge, we are the first to integrate the extra label information into KG-based multi-typed DDI prediction; besides, we innovatively proposed a novel KG feature learning method and a State Encoder to obtain target drug pairs' KG-based features which contained more abundant and more key drug-related KG information with less noise; furthermore, a Gated Recurrent Unit-based multi-channel feature fusion module was proposed in an innovative way to yield more comprehensive feature information about drug pairs, effectively alleviating the problem of feature redundancy. We experimented with four datasets in the multi-class and the multi-label prediction tasks to comprehensively evaluate the performance of MCFF-MTDDI for predicting interactions of known-known drugs, known-new drugs and new-new drugs. In addition, we further conducted ablation studies and case studies. All the results fully demonstrated the effectiveness of MCFF-MTDDI. Chen-Di Han, Chun-Chun Wang, Xing Chen 0001 |
Briefings Bioinform. | 4 |
| 2023 | SNRMPACDC: computational model focused on Siamese network and random matrix projection for anticancer synergistic drug combination predictionabstractSynergistic drug combinations can improve the therapeutic effect and reduce the drug dosage to avoid toxicity. In previous years, an in vitro approach was utilized to screen synergistic drug combinations. However, the in vitro method is time-consuming and expensive. With the rapid growth of high-throughput data, computational methods are becoming efficient tools to predict potential synergistic drug combinations. Considering the limitations of the previous computational methods, we developed a new model named Siamese Network and Random Matrix Projection for AntiCancer Drug Combination prediction (SNRMPACDC). Firstly, the Siamese convolutional network and random matrix projection were used to process the features of the two drugs into drug combination features. Then, the features of the cancer cell line were processed through the convolutional network. Finally, the processed features were integrated and input into the multi-layer perceptron network to get the predicted score. Compared with the traditional method of splicing drug features into drug combination features, SNRMPACDC improved the interpretability of drug combination features to a certain extent. In addition, the introduction of convolutional networks can better extract the potential information in the features. SNRMPACDC achieved the root mean-squared error of 15.01 and the Pearson correlation coefficient of 0.75 in 5-fold cross-validation of regression prediction for response data. In addition, SNRMPACDC achieved the AUC of 0.91 ± 0.03 and the AUPR of 0.62 ± 0.05 in 5-fold cross-validation of classification prediction of synergistic or not. These results are almost better than all the previous models. SNRMPACDC would be an effective approach to infer potential anticancer synergistic drug combinations. Tian-Hao Li, Chun-Chun Wang, Xing Chen 0001 |
Briefings Bioinform. | 4 |
| 2022 | A deep learning-based unsupervised learning method for spatially resolved transcriptomic data analysistabstractSpatially resolved transcriptomic data provide a large quantity of high-throughput gene expression and spatial structure information of tissues. Spatial clusters obtained by spatial transcriptome helps us to identify co-expressed regions and gene modules corresponding to cell types. In this study, we developed a Deep learning-based spatial clustering algorithm (RkDeep) by combining Ratio-cut and k-means. We first preprocessed spatial transcriptome data using graph neural network, and conducted dimensional reduction on the preprocessed data with denoising autoencoder. Finally, we clustered spatial transcriptome data by combining ratio cut and k-means. We compared our proposed RkDeep method with the other two spatial clustering methods, Seurat and Panoview. The results show that RkDeep computed the smallest Davide-Bouldin index and the largest Caliniski Harabaz index, adjusted rand index and normalized mutual information. Moreover, RkDeep was applied to analyze spatial transcriptome data of adult mouse brain, adult mouse kidney, and breast cancer. The results show that RkDeep can more accurately identify cell types from spatial transcriptome data. Lihong Peng, Xianzhi He, Xinhuai Peng, Yuankang Lu, Xing Chen 0001 |
BIBM | 7 |
| 2022 | Analyses of cell-to-cell communication combining a heterogeneous deep ensemble framework and scoring approaches from single-cell RNA sequencing dataabstractCell-to-cell communication (CCC) plays essential roles in multicellular organisms. the identification of CCC between cancer cells themselves and one between cancer cells and normal cells in tumor microenvironment contributes to the understanding of carcinogenesis, cancer development and metastasis. CCC is usually mediated by Ligand-Receptor Interactions (LRIs). In this manuscript, we developed an LRI-mediated CCC estimation framework (LRI-EnABCLG) by incorporating LRI collection, prediction and filtering, CCC inference and visualization. First, four LRI datasets were collected. Second, LRIs were predicted by a heterogeneous deep ensemble model. Third, LRIs were filtered by combining single-cell sequencing (scRNA-seq) data. Fourth, CCC was inferred by combining the filtered LRIs and scRNA-seq data. Finally, the proposed CCC prediction framework was applied to CCC analysis in colorectal tumor tissues. Our proposed LRI-EnABCLG model obtained better LRI prediction performance. Case study demonstrated that fibroblasts was more likely to communicate with colorectal cancer cells, which was in accord with the results from iTALK (a classical CCC analysis pipeline). We anticipate that this work can contribute to diagnosis and treatment of cancers. Lihong Peng, Ruya Yuan, Chendi Han, Jingwei Tan, Min Chen 0003, Xing Chen 0001 |
BIBM | 7 |
| 2022 | Computational model for ncRNA researchabstractThe explosion of research on non-coding RNAs (ncRNAs) in the past few decades has transformed the original notion of regarding such RNAs as ‘transcriptional noise’ [1, 2] before 1980s to gene expression regulators at transcriptional, RNA processing and translational levels [3–6]. It is evident from PubMed that each week new studies are published to reveal altered ncRNA expressions in diseases or discover novel non-coding transcripts [7]. Abundant in eukaryotes ranging from Homo sapiens to Caenorhabditis elegans [8, 9] and persistent in prokaryotes such as bacteria [10, 11], ncRNAs fall into two main classes according to transcript length [12, 13]: long ncRNAs (lncRNAs, with length > 200 bp) and small ncRNAs (sncRNAs, with length < 200 bp). Based on conformation and cellular function [14–17], the former can be further divided into linear RNAs and circular RNAs (circRNAs), whereas the latter can be grouped into more than a dozen subclasses, among which microRNAs (miRNAs) and piwi-interacting RNAs (piRNAs) have recently attracted considerable research attention. Xing Chen 0001 |
Briefings Bioinform. | 1 |
| 2022 | Updated review of advances in microRNAs and complex diseases: taxonomy, trends and challenges of computational modelsabstractSince the problem proposed in late 2000s, microRNA-disease association (MDA) predictions have been implemented based on the data fusion paradigm. Integrating diverse data sources gains a more comprehensive research perspective, and brings a challenge to algorithm design for generating accurate, concise and consistent representations of the fused data. After more than a decade of research progress, a relatively simple algorithm like the score function or a single computation layer may no longer be sufficient for further improving predictive performance. Advanced model design has become more frequent in recent years, particularly in the form of reasonably combing multiple algorithms, a process known as model fusion. In the current review, we present 29 state-of-the-art models and introduce the taxonomy of computational models for MDA prediction based on model fusion and non-fusion. The new taxonomy exhibits notable changes in the algorithmic architecture of models, compared with that of earlier ones in the 2017 review by Chen et al. Moreover, we discuss the progresses that have been made towards overcoming the obstacles to effective MDA prediction since 2017 and elaborated on how future models can be designed according to a set of new schemas. Lastly, we analysed the strengths and weaknesses of each model category in the proposed taxonomy and proposed future research directions from diverse perspectives for enhancing model performance. Xing Chen 0001 |
Briefings Bioinform. | 3 |
| 2022 | Updated review of advances in microRNAs and complex diseases: experimental results, databases, webservers and data fusionabstractMicroRNAs (miRNAs) are gene regulators involved in the pathogenesis of complex diseases such as cancers, and thus serve as potential diagnostic markers and therapeutic targets. The prerequisite for designing effective miRNA therapies is accurate discovery of miRNA-disease associations (MDAs), which has attracted substantial research interests during the last 15 years, as reflected by more than 55 000 related entries available on PubMed. Abundant experimental data gathered from the wealth of literature could effectively support the development of computational models for predicting novel associations. In 2017, Chen et al. published the first-ever comprehensive review on MDA prediction, presenting various relevant databases, 20 representative computational models, and suggestions for building more powerful ones. In the current review, as the continuation of the previous study, we revisit miRNA biogenesis, detection techniques and functions; summarize recent experimental findings related to common miRNA-associated diseases; introduce recent updates of miRNA-relevant databases and novel database releases since 2017, present mainstream webservers and new webserver releases since 2017 and finally elaborate on how fusion of diverse data sources has contributed to accurate MDA prediction. Xing Chen 0001 |
Briefings Bioinform. | 3 |
| 2022 | Updated review of advances in microRNAs and complex diseases: towards systematic evaluation of computational modelsabstractCurrently, there exist no generally accepted strategies of evaluating computational models for microRNA-disease associations (MDAs). Though K-fold cross validations and case studies seem to be must-have procedures, the value of K, the evaluation metrics, and the choice of query diseases as well as the inclusion of other procedures (such as parameter sensitivity tests, ablation studies and computational cost reports) are all determined on a case-by-case basis and depending on the researchers' choices. In the current review, we include a comprehensive analysis on how 29 state-of-the-art models for predicting MDAs were evaluated. Based on the analytical results, we recommend a feasible evaluation workflow that would suit any future model to facilitate fair and systematic assessment of predictive performance. Xing Chen 0001 |
Briefings Bioinform. | 3 |
| 2022 | Prediction of potential miRNA-disease associations based on stacked autoencoderabstractIn recent years, increasing biological experiments and scientific studies have demonstrated that microRNA (miRNA) plays an important role in the development of human complex diseases. Therefore, discovering miRNA-disease associations can contribute to accurate diagnosis and effective treatment of diseases. Identifying miRNA-disease associations through computational methods based on biological data has been proven to be low-cost and high-efficiency. In this study, we proposed a computational model named Stacked Autoencoder for potential MiRNA-Disease Association prediction (SAEMDA). In SAEMDA, all the miRNA-disease samples were used to pretrain a Stacked Autoencoder (SAE) in an unsupervised manner. Then, the positive samples and the same number of selected negative samples were utilized to fine-tune SAE in a supervised manner after adding an output layer with softmax classifier to the SAE. SAEMDA can make full use of the feature information of all unlabeled miRNA-disease pairs. Therefore, SAEMDA is suitable for our dataset containing small labeled samples and large unlabeled samples. As a result, SAEMDA achieved AUCs of 0.9210 and 0.8343 in global and local leave-one-out cross validation. Besides, SAEMDA obtained an average AUC and standard deviation of 0.9102 ± /-0.0029 in 100 times of 5-fold cross validation. These results were better than those of previous models. Moreover, we carried out three case studies to further demonstrate the predictive accuracy of SAEMDA. As a result, 82% (breast neoplasms), 100% (lung neoplasms) and 90% (esophageal neoplasms) of the top 50 predicted miRNAs were verified by databases. Thus, SAEMDA could be a useful and reliable model to predict potential miRNA-disease associations. Chun-Chun Wang, Tian-Hao Li, Xing Chen 0001 |
Briefings Bioinform. | 4 |
| 2022 | Dual-Network Collaborative Matrix Factorization for predicting small molecule-miRNA associationsabstractMicroRNAs (miRNAs) play crucial roles in multiple biological processes and human diseases and can be considered as therapeutic targets of small molecules (SMs). Because biological experiments used to verify SM-miRNA associations are time-consuming and expensive, it is urgent to propose new computational models to predict new SM-miRNA associations. Here, we proposed a novel method called Dual-network Collaborative Matrix Factorization (DCMF) for predicting the potential SM-miRNA associations. Firstly, we utilized the Weighted K Nearest Known Neighbors (WKNKN) method to preprocess SM-miRNA association matrix. Then, we constructed matrix factorization model to obtain two feature matrices containing latent features of SM and miRNA, respectively. Finally, the predicted SM-miRNA association score matrix was obtained by calculating the inner product of two feature matrices. The main innovations of this method were that the use of WKNKN method can preprocess the missing values of association matrix and the introduction of dual network can integrate more diverse similarity information into DCMF. For evaluating the validity of DCMF, we implemented four different cross validations (CVs) based on two distinct datasets and two different case studies. Finally, based on dataset 1 (dataset 2), DCMF achieved Area Under receiver operating characteristic Curves (AUC) of 0.9868 (0.8770), 0.9833 (0.8836), 0.8377 (0.7591) and 0.9836 ± 0.0030 (0.8632 ± 0.0042) in global Leave-One-Out Cross Validation (LOOCV), miRNA-fixed local LOOCV, SM-fixed local LOOCV and 5-fold CV, respectively. For case studies, plenty of predicted associations have been confirmed by published experimental literature. Therefore, DCMF is an effective tool to predict potential SM-miRNA associations. Shu-Hao Wang, Chun-Chun Wang, Lian-Ying Miao, Xing Chen 0001 |
Briefings Bioinform. | 5 |
| 2022 | Ensemble of kernel ridge regression-based small molecule-miRNA association prediction in human diseaseabstractMicroRNAs (miRNAs) play crucial roles in human disease and can be targeted by small molecule (SM) drugs according to numerous studies, which shows that identifying SM-miRNA associations in human disease is important for drug development and disease treatment. We proposed the method of Ensemble of Kernel Ridge Regression-based Small Molecule-MiRNA Association prediction (EKRRSMMA) to uncover potential SM-miRNA associations by combing feature dimensionality reduction and ensemble learning. First, we constructed different feature subsets for both SMs and miRNAs. Then, we trained homogeneous base learners based on distinct feature subsets and took the average of scores obtained from these base learners as SM-miRNA association score. In EKRRSMMA, feature dimensionality reduction technology was employed in the process of construction of feature subsets to reduce the influence of noisy data. Besides, the base learner, namely KRR_avg, was the combination of two classifiers constructed under SM space and miRNA space, which could make full use of the information of SM and miRNA. To assess the prediction performance of EKRRSMMA, we conducted Leave-One-Out Cross-Validation (LOOCV), SM-fixed local LOOCV, miRNA-fixed local LOOCV and 5-fold CV based on two datasets. For Dataset 1 (Dataset 2), EKRRSMMA got the Area Under receiver operating characteristic Curves (AUCs) of 0.9793 (0.8871), 0.8071 (0.7705), 0.9732 (0.8586) and 0.9767 ± 0.0014 (0.8560 ± 0.0027). Besides, we conducted four case studies. As a result, 32 (5-Fluorouracil), 19 (17β-Estradiol), 26 (5-Aza-2'-deoxycytidine) and 11 (cyclophosphamide) out of top 50 predicted potentially associated miRNAs were confirmed by database or experimental literature. Above evaluation results demonstrated that EKRRSMMA is reliable for predicting SM-miRNA associations. Chun-Chun Wang, Chi-Chi Zhu, Xing Chen 0001 |
Briefings Bioinform. | 3 |
| 2022 | Predicting drug-target binding affinity through molecule representation block based on multi-head attention and skip connectionabstractExiting computational models for drug-target binding affinity prediction have much room for improvement in prediction accuracy, robustness and generalization ability. Most deep learning models lack interpretability analysis and few studies provide application examples. Based on these observations, we presented a novel model named Molecule Representation Block-based Drug-Target binding Affinity prediction (MRBDTA). MRBDTA is composed of embedding and positional encoding, molecule representation block and interaction learning module. The advantages of MRBDTA are reflected in three aspects: (i) developing Trans block to extract molecule features through improving the encoder of transformer, (ii) introducing skip connection at encoder level in Trans block and (iii) enhancing the ability to capture interaction sites between proteins and drugs. The test results on two benchmark datasets manifest that MRBDTA achieves the best performance compared with 11 state-of-the-art models. Besides, through replacing Trans block with single Trans encoder and removing skip connection in Trans block, we verified that Trans block and skip connection could effectively improve the prediction accuracy and reliability of MRBDTA. Then, relying on multi-head attention mechanism, we performed interpretability analysis to illustrate that MRBDTA can correctly capture part of interaction sites between proteins and drugs. In case studies, we firstly employed MRBDTA to predict binding affinities between Food and Drug Administration-approved drugs and severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) replication-related proteins. Secondly, we compared true binding affinities between 3C-like proteinase and 185 drugs with those predicted by MRBDTA. The final results of case studies reveal reliable performance of MRBDTA in drug design for SARS-CoV-2. Chun-Chun Wang, Xing Chen 0001 |
Briefings Bioinform. | 3 |
| 2021 | Deep-belief network for predicting potential miRNA-disease associationsabstractMicroRNA (miRNA) plays an important role in the occurrence, development, diagnosis and treatment of diseases. More and more researchers begin to pay attention to the relationship between miRNA and disease. Compared with traditional biological experiments, computational method of integrating heterogeneous biological data to predict potential associations can effectively save time and cost. Considering the limitations of the previous computational models, we developed the model of deep-belief network for miRNA-disease association prediction (DBNMDA). We constructed feature vectors to pre-train restricted Boltzmann machines for all miRNA-disease pairs and applied positive samples and the same number of selected negative samples to fine-tune DBN to obtain the final predicted scores. Compared with the previous supervised models that only use pairs with known label for training, DBNMDA innovatively utilizes the information of all miRNA-disease pairs during the pre-training process. This step could reduce the impact of too few known associations on prediction accuracy to some extent. DBNMDA achieves the AUC of 0.9104 based on global leave-one-out cross validation (LOOCV), the AUC of 0.8232 based on local LOOCV and the average AUC of 0.9048 ± 0.0026 based on 5-fold cross validation. These AUCs are better than other previous models. In addition, three different types of case studies for three diseases were implemented to demonstrate the accuracy of DBNMDA. As a result, 84% (breast neoplasms), 100% (lung neoplasms) and 88% (esophageal neoplasms) of the top 50 predicted miRNAs were verified by recent literature. Therefore, we could conclude that DBNMDA is an effective method to predict potential miRNA-disease associations. Xing Chen 0001, Tian-Hao Li, Chun-Chun Wang, Chi-Chi Zhu |
Briefings Bioinform. | 1 |
| 2021 | Predicting potential small molecule-miRNA associations based on bounded nuclear norm regularizationabstractMounting evidence has demonstrated the significance of taking microRNAs (miRNAs) as the target of small molecule (SM) drugs for disease treatment. Given the fact that exploring new SM-miRNA associations through biological experiments is extremely expensive, several computing models have been constructed to reveal the possible SM-miRNA associations. Here, we built a computing model of Bounded Nuclear Norm Regularization for SM-miRNA Associations prediction (BNNRSMMA). Specifically, we first constructed a heterogeneous SM-miRNA network utilizing miRNA similarity, SM similarity, confirmed SM-miRNA associations and defined a matrix to represent the heterogeneous network. Then, we constructed a model to complete this matrix by minimizing its nuclear norm. The Alternating Direction Method of Multipliers was adopted to minimize the nuclear norm and obtain predicted scores. The main innovation lies in two aspects. During completion, we limited all elements of the matrix within the interval of (0,1) to make sure they have practical significance. Besides, instead of strictly fitting all known elements, a regularization term was incorporated to tolerate the noise in integrated similarities. Furthermore, four kinds of cross-validations on two datasets and two types of case studies were performed to evaluate the predictive performance of BNNRSMMA. Finally, BNNRSMMA attained areas under the curve of 0.9822 (0.8433), 0.9793 (0.8852), 0.8253 (0.7350) and 0.9758 ± 0.0029 (0.8759 ± 0.0041) under global leave-one-out cross-validation (LOOCV), miRNA-fixed LOOCV, SM-fixed LOOCV and 5-fold cross-validation based on Dataset 1(Dataset 2), respectively. With regard to case studies, plenty of predicted associations have been verified by experimental literatures. All these results confirmed that BNNRSMMA is a reliable tool for inferring associations. Xing Chen 0001, Chun-Chun Wang |
Briefings Bioinform. | 1 |
| 2021 | Circular RNAs and complex diseases: from experimental results to computational modelsabstractAbstract Circular RNAs (circRNAs) are a class of single-stranded, covalently closed RNA molecules with a variety of biological functions. Studies have shown that circRNAs are involved in a variety of biological processes and play an important role in the development of various complex diseases, so the identification of circRNA-disease associations would contribute to the diagnosis and treatment of diseases. In this review, we summarize the discovery, classifications and functions of circRNAs and introduce four important diseases associated with circRNAs. Then, we list some significant and publicly accessible databases containing comprehensive annotation resources of circRNAs and experimentally validated circRNA-disease associations. Next, we introduce some state-of-the-art computational models for predicting novel circRNA-disease associations and divide them into two categories, namely network algorithm-based and machine learning-based models. Subsequently, several evaluation methods of prediction performance of these computational models are summarized. Finally, we analyze the advantages and disadvantages of different types of computational models and provide some suggestions to promote the development of circRNA-disease association identification from the perspective of the construction of new computational models and the accumulation of circRNA-related data. Chun-Chun Wang, Chendi Han, Qi Zhao 0010, Xing Chen 0001 |
Briefings Bioinform. | 4 |
| 2021 | Drug-pathway association prediction: from experimental results to computational modelsabstractEffective drugs are urgently needed to overcome human complex diseases. However, the research and development of novel drug would take long time and cost much money. Traditional drug discovery follows the rule of one drug-one target, while some studies have demonstrated that drugs generally perform their task by affecting related pathway rather than targeting single target. Thus, the new strategy of drug discovery, namely pathway-based drug discovery, have been proposed. Obviously, identifying associations between drugs and pathways plays a key role in the development of pathway-based drug discovery. Revealing the drug-pathway associations by experiment methods would take much time and cost. Therefore, some computational models were established to predict potential drug-pathway associations. In this review, we first introduced the background of drug and the concept of drug-pathway associations. Then, some publicly accessible databases and web servers about drug-pathway associations were listed. Next, we summarized some state-of-the-art computational methods in the past years for inferring drug-pathway associations and divided these methods into three classes, namely Bayesian spare factor-based, matrix decomposition-based and other machine learning methods. In addition, we introduced several evaluation strategies to estimate the predictive performance of various computational models. In the end, we discussed the advantages and limitations of existing computational methods and provided some suggestions about the future directions of the data collection and the calculation models development. Chun-Chun Wang, Xing Chen 0001 |
Briefings Bioinform. | 3 |
| 2021 | Microbes and complex diseases: from experimental results to computational modelsabstractStudies have shown that the number of microbes in humans is almost 10 times that of cells. These microbes have been proven to play an important role in a variety of physiological processes, such as enhancing immunity, improving the digestion of gastrointestinal tract and strengthening metabolic function. In addition, in recent years, more and more research results have indicated that there are close relationships between the emergence of the human noncommunicable diseases and microbes, which provides a novel insight for us to further understand the pathogenesis of the diseases. An in-depth study about the relationships between diseases and microbes will not only contribute to exploring new strategies for the diagnosis and treatment of diseases but also significantly heighten the efficiency of new drugs development. However, applying the methods of biological experimentation to reveal the microbe-disease associations is costly and inefficient. In recent years, more and more researchers have constructed multiple computational models to predict microbes that are potentially associated with diseases. Here, we start with a brief introduction of microbes and databases as well as web servers related to them. Then, we mainly introduce four kinds of computational models, including score function-based models, network algorithm-based models, machine learning-based models and experimental analysis-based models. Finally, we summarize the advantages as well as disadvantages of them and set the direction for the future work of revealing microbe-disease associations based on computational models. We firmly believe that computational models are expected to be important tools in large-scale predictions of disease-related microbes. Chun-Chun Wang, Xing Chen 0001 |
Briefings Bioinform. | 3 |
| 2021 | Identification of miRNA-disease associations via multiple information integration with Bayesian rankingabstractIn recent years, increasing microRNA (miRNA)-disease associations were identified through traditionally biological experiments. These associations contribute to revealing molecular mechanism of diseases and preventing and curing diseases. To improve the efficiency of miRNA-disease association discovery, some calculation methods were developed as auxiliary tools for researchers. In the current study, we raised a novel model named Bayesian Ranking for MiRNA-Disease Association prediction (BRMDA) by improving Bayesian Personalized Ranking from three aspects: (i) taking advantage of similarity of diseases and miRNAs; (ii) incorporating miRNA bias for miRNAs associated with different number of diseases; and (iii) implementing neighborhood-based approach for new miRNAs and diseases. For each investigated disease, BRMDA used the set of triples (i.e. disease, labeled miRNA, unlabeled miRNA) that reflected association preference of the disease to miRNAs as training set, which made full use of unknown samples rather than simply considering them as negative samples. To investigate the predictive performance of BRMDA, we employed leave-one-out cross-validation and obtained Area Under the Curve of 0.8697, which outperformed many classical methods. Besides, we further implemented three distinct classes of case studies for three common Neoplasms. As a result, there are 44 (Colon Neoplasms), 49 (Esophageal Neoplasms) and 49 (Lung Neoplasms) among the top 50 predicted miRNAs validated through experiments. In short, BRMDA would be a trustable tool for inferring valuable associations. Chi-Chi Zhu, Chun-Chun Wang, Mingcheng Zuo, Xing Chen 0001 |
Briefings Bioinform. | 5 |
| 2020 | MicroRNA-small molecule association identification: from experimental results to computational modelsabstractSmall molecule is a kind of low molecular weight organic compound with variety of biological functions. Studies have indicated that small molecules can inhibit a specific function of a multifunctional protein or disrupt protein-protein interactions and may have beneficial or detrimental effect against diseases. MicroRNAs (miRNAs) play crucial roles in cellular biology, which makes it possible to develop miRNA as diagnostics and therapeutic targets. Several drug-like compound libraries were screened successfully against different miRNAs in cellular assays further demonstrating the possibility of targeting miRNAs with small molecules. In this review, we summarized the concept and functions of small molecule and miRNAs. Especially, five aspects of miRNA functions were exhibited in detail with individual examples. In addition, four disease states that have been linked to miRNA alterations were summed up. Then, small molecules related to four important miRNAs miR-21, 122, 4644 and 27 were selected for introduction. Some important publicly accessible databases and web servers of the experimentally validated or potential small molecule-miRNA associations were discussed. Identifying small molecule targeting miRNAs has become an important goal of biomedical research. Thus, several experimental and computational models have been developed and implemented to identify novel small molecule-miRNA associations. Here, we reviewed four experimental techniques used in the past few years to search for small-molecule inhibitors of miRNAs, as well as three types of models of predicting small molecule-miRNA associations from different perspectives. Finally, we summarized the limitations of existing methods and discussed the future directions for further development of computational models. Xing Chen 0001, Na-Na Guan, Ya-Zhou Sun, Jianqiang Li 0001 |
Briefings Bioinform. | 1 |
| 2020 | Adaptive boosting-based computational model for predicting potential miRNA-disease associationsabstractBioinformatics (2019) doi: 10.1093/bioinformatics/btz297 In the above mentioned article, several mathematical variables were inadvertently duplicated. These duplicates have now been removed. The publisher apologises for this error. Xing Chen 0001 |
Bioinform. | 2 |
| 2019 | RNA methylation and diseases: experimental results, databases, Web servers and computational modelsabstractRibonucleic acid (RNA) methylation is a type of posttranscriptional modifications occurring in all kingdoms of life. It is strongly related to important biological process, thus making it linked to a number of human diseases. Owing to the development of high-throughput sequencing technology, plenty of achievement had been obtained in RNA methylation research recently. Meanwhile, various computational models have been developed to analyze and mining increasing RNA methylation data. In this review, we first made a brief introduction about eight types of most popular RNA methylation, the biological functions of RNA methylation, the relationship between RNA methylation and disease and five important RNA methylation-related diseases. The research of RNA methylation is based on sequencing data processing, and effective bioinformatics techniques can benefit better understanding of RNA methylation. We further introduced seven publicly available RNA methylation-related databases, and some important publicly available RNA-methylation-related Web servers and software for RNA methylation site identification, differential analysis and so on. Furthermore, we provided detailed analysis of the state-of-the-art computational models used in these Web servers and software. We also analyzed the limitations of these models and discussed the future directions of developing computational models for RNA methylation research. Xing Chen 0001, Ya-Zhou Sun, Hui Liu 0024, Lin Zhang 0015, Jianqiang Li 0001, Jia Meng 0001 |
Briefings Bioinform. | 1 |
| 2019 | MicroRNAs and complex diseases: from experimental results to computational modelsabstractCircular RNAs (circRNAs) are a class of single-stranded, covalently closed RNA molecules with a variety of biological functions. Studies have shown that circRNAs are involved in a variety of biological processes and play an important role in the development of various complex diseases, so the identification of circRNA-disease associations would contribute to the diagnosis and treatment of diseases. In this review, we summarize the discovery, classifications and functions of circRNAs and introduce four important diseases associated with circRNAs. Then, we list some significant and publicly accessible databases containing comprehensive annotation resources of circRNAs and experimentally validated circRNA-disease associations. Next, we introduce some state-of-the-art computational models for predicting novel circRNA-disease associations and divide them into two categories, namely network algorithm-based and machine learning-based models. Subsequently, several evaluation methods of prediction performance of these computational models are summarized. Finally, we analyze the advantages and disadvantages of different types of computational models and provide some suggestions to promote the development of circRNA-disease association identification from the perspective of the construction of new computational models and the accumulation of circRNA-related data. Xing Chen 0001, Di Xie, Qi Zhao 0010, Zhu-Hong You |
Briefings Bioinform. | 1 |
| 2019 | Adaptive boosting-based computational model for predicting potential miRNA-disease associationsabstractMOTIVATION: Recent studies have shown that microRNAs (miRNAs) play a critical part in several biological processes and dysregulation of miRNAs is related with numerous complex human diseases. Thus, in-depth research of miRNAs and their association with human diseases can help us to solve many problems. RESULTS: Due to the high cost of traditional experimental methods, revealing disease-related miRNAs through computational models is a more economical and efficient way. Considering the disadvantages of previous models, in this paper, we developed adaptive boosting for miRNA-disease association prediction (ABMDA) to predict potential associations between diseases and miRNAs. We balanced the positive and negative samples by performing random sampling based on k-means clustering on negative samples, whose process was quick and easy, and our model had higher efficiency and scalability for large datasets than previous methods. As a boosting technology, ABMDA was able to improve the accuracy of given learning algorithm by integrating weak classifiers that could score samples to form a strong classifier based on corresponding weights. Here, we used decision tree as our weak classifier. As a result, the area under the curve (AUC) of global and local leave-one-out cross validation reached 0.9170 and 0.8220, respectively. What is more, the mean and the standard deviation of AUCs achieved 0.9023 and 0.0016, respectively in 5-fold cross validation. Besides, in the case studies of three important human cancers, 49, 50 and 50 out of the top 50 predicted miRNAs for colon neoplasms, hepatocellular carcinoma and breast neoplasms were confirmed by the databases and experimental literatures. AVAILABILITY AND IMPLEMENTATION: The code and dataset of ABMDA are freely available at https://github.com/githubcode007/ABMDA. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Xing Chen 0001 |
Bioinform. | 2 |
| 2019 | Integrating random walk and binary regression to identify novel miRNA-disease associationabstractBACKGROUND: In the last few decades, cumulative experimental researches have witnessed and verified the important roles of microRNAs (miRNAs) in the development of human complex diseases. Benefitting from the rapid growth both in the availability of miRNA-related data and the development of various analysis methodologies, up until recently, some computational models have been developed to predict human disease related miRNAs, efficiently and quickly. RESULTS: In this work, we proposed a computational model of Random Walk and Binary Regression-based MiRNA-Disease Association prediction (RWBRMDA). RWBRMDA extracted features for each miRNA from random walk with restart on the integrated miRNA similarity network for binary logistic regression to predict potential miRNA-disease associations. RWBRMDA obtained AUC of 0.8076 in the leave-one-out cross validation. Additionally, we carried out three different patterns of case studies on four human complex diseases. Specifically, Esophageal cancer and Prostate cancer were conducted as one kind of case study based on known miRNA-disease associations in HMDD v2.0 database. Out of the top 50 predicted miRNAs, 94 and 90% were respectively confirmed by recent experimental reports. To simulate new disease without known related miRNAs, the information of known Breast cancer related miRNAs was removed. As a result, 98% of the top 50 predicted miRNAs for Breast cancer were confirmed. Lymphoma, the verified ratio of which was 88%, was used to assess the prediction robustness of RWBRMDA based on the association records in HMDD v1.0 database. CONCLUSIONS: We anticipated that RWBRMDA could benefit the future experimental investigations about the relation between human disease and miRNAs by generating promising and testable top-ranked miRNAs, and significantly reducing the effort and cost of identification works. Ya-Wei Niu, Guanghui Wang 0002, Guiying Yan, Xing Chen 0001 |
BMC Bioinform. | 4 |
| 2019 | Prediction of potential miRNA-disease associations using matrix decomposition and label propagation
Xing Chen 0001, Zhengwei Li 0001 |
Knowl. Based Syst. | 2 |
| 2019 | Ensemble of decision tree reveals potential miRNA-disease associationsabstractIn recent years, increasing associations between microRNAs (miRNAs) and human diseases have been identified. Based on accumulating biological data, many computational models for potential miRNA-disease associations inference have been developed, which saves time and expenditure on experimental studies, making great contributions to researching molecular mechanism of human diseases and developing new drugs for disease treatment. In this paper, we proposed a novel computational method named Ensemble of Decision Tree based MiRNA-Disease Association prediction (EDTMDA), which innovatively built a computational framework integrating ensemble learning and dimensionality reduction. For each miRNA-disease pair, the feature vector was extracted by calculating the statistical measures, graph theoretical measures, and matrix factorization results for the miRNA and disease, respectively. Then multiple base learnings were built to yield many decision trees (DTs) based on random selection of negative samples and miRNA/disease features. Particularly, Principal Components Analysis was applied to each base learning to reduce feature dimensionality and hence remove the noise or redundancy. Average strategy was adopted for these DTs to get final association scores between miRNAs and diseases. In model performance evaluation, EDTMDA showed AUC of 0.9309 in global leave-one-out cross validation (LOOCV) and AUC of 0.8524 in local LOOCV. Additionally, AUC of 0.9192+/-0.0009 in 5-fold cross validation proved the model's reliability and stability. Furthermore, three types of case studies for four human diseases were implemented. As a result, 94% (Esophageal Neoplasms), 86% (Kidney Neoplasms), 96% (Breast Neoplasms) and 88% (Carcinoma Hepatocellular) of top 50 predicted miRNAs were confirmed by experimental evidences in literature. Xing Chen 0001, Chi-Chi Zhu |
PLoS Comput. Biol. | 1 |
| 2019 | LMTRDA: Using logistic model tree to predict MiRNA-disease associations by fusing multi-source information of sequences and similaritiesabstractEmerging evidence has shown microRNAs (miRNAs) play an important role in human disease research. Identifying potential association among them is significant for the development of pathology, diagnose and therapy. However, only a tiny portion of all miRNA-disease pairs in the current datasets are experimentally validated. This prompts the development of high-precision computational methods to predict real interaction pairs. In this paper, we propose a new model of Logistic Model Tree for predicting miRNA-Disease Association (LMTRDA) by fusing multi-source information including miRNA sequences, miRNA functional similarity, disease semantic similarity, and known miRNA-disease associations. In particular, we introduce miRNA sequence information and extract its features using natural language processing technique for the first time in the miRNA-disease prediction model. In the cross-validation experiment, LMTRDA obtained 90.51% prediction accuracy with 92.55% sensitivity at the AUC of 90.54% on the HMDD V3.0 dataset. To further evaluate the performance of LMTRDA, we compared it with different classifier and feature descriptor models. In addition, we also validate the predictive ability of LMTRDA in human diseases including Breast Neoplasms, Breast Neoplasms and Lymphoma. As a result, 28, 27 and 26 out of the top 30 miRNAs associated with these diseases were verified by experiments in different kinds of case studies. These experimental results demonstrate that LMTRDA is a reliable model for predicting the association among miRNAs and diseases. Lei Wang 0121, Zhu-Hong You, Xing Chen 0001, Yang-Ming Li, Ya-Nan Dong, Liping Li 0003, Kai Zheng 0020 |
PLoS Comput. Biol. | 3 |
| 2018 | A novel approach based on KATZ measure to predict associations of human microbiota with non-infectious diseasesabstractBioinformatics (2017) 33 (5): 733–739. DOI: https://doi.org/10.1093/bioinformatics/btw715 The publisher wishes to inform readers that the footnote for the † symbol was erroneously removed from the paper as first published. The paper has now been corrected online to include the footnote: ‘†The authors wish it to be known that, in their opinion, the first two authors should be regarded as Joint First Authors’. Xing Chen 0001, Zhu-Hong You, Guiying Yan, Xuesong Wang 0001 |
Bioinform. | 1 |
| 2018 | Predicting miRNA-disease association based on inductive matrix completionabstractMotivation: It has been shown that microRNAs (miRNAs) play key roles in variety of biological processes associated with human diseases. In Consideration of the cost and complexity of biological experiments, computational methods for predicting potential associations between miRNAs and diseases would be an effective complement. Results: This paper presents a novel model of Inductive Matrix Completion for MiRNA-Disease Association prediction (IMCMDA). The integrated miRNA similarity and disease similarity are calculated based on miRNA functional similarity, disease semantic similarity and Gaussian interaction profile kernel similarity. The main idea is to complete the missing miRNA-disease association based on the known associations and the integrated miRNA similarity and disease similarity. IMCMDA achieves AUC of 0.8034 based on leave-one-out-cross-validation and improved previous models. In addition, IMCMDA was applied to five common human diseases in three types of case studies. In the first type, respectively, 42, 44, 45 out of top 50 predicted miRNAs of Colon Neoplasms, Kidney Neoplasms, Lymphoma were confirmed by experimental reports. In the second type of case study for new diseases without any known miRNAs, we chose Breast Neoplasms as the test example by hiding the association information between the miRNAs and Breast Neoplasms. As a result, 50 out of top 50 predicted Breast Neoplasms-related miRNAs are verified. In the third type of case study, IMCMDA was tested on HMDD V1.0 to assess the robustness of IMCMDA, 49 out of top 50 predicted Esophageal Neoplasms-related miRNAs are verified. Availability and implementation: The code and dataset of IMCMDA are freely available at https://github.com/IMCMDAsourcecode/IMCMDA. Supplementary information: Supplementary data are available at Bioinformatics online. Xing Chen 0001, Lei Wang 0121, Na-Na Guan, Jianqiang Li 0001 |
Bioinform. | 1 |
| 2018 | BNPMDA: Bipartite Network Projection for MiRNA-Disease Association predictionabstractMotivation: A large number of resources have been devoted to exploring the associations between microRNAs (miRNAs) and diseases in the recent years. However, the experimental methods are expensive and time-consuming. Therefore, the computational methods to predict potential miRNA-disease associations have been paid increasing attention. Results: In this paper, we proposed a novel computational model of Bipartite Network Projection for MiRNA-Disease Association prediction (BNPMDA) based on the known miRNA-disease associations, integrated miRNA similarity and integrated disease similarity. We firstly described the preference degree of a miRNA for its related disease and the preference degree of a disease for its related miRNA with the bias ratings. We constructed bias ratings for miRNAs and diseases by using agglomerative hierarchical clustering according to the three types of networks. Then, we implemented the bipartite network recommendation algorithm to predict the potential miRNA-disease associations by assigning transfer weights to resource allocation links between miRNAs and diseases based on the bias ratings. BNPMDA had been shown to improve the prediction accuracy in comparison with previous models according to the area under the receiver operating characteristics (ROC) curve (AUC) results of three typical cross validations. As a result, the AUCs of Global LOOCV, Local LOOCV and 5-fold cross validation obtained by implementing BNPMDA were 0.9028, 0.8380 and 0.8980 ± 0.0013, respectively. We further implemented two types of case studies on several important human complex diseases to confirm the effectiveness of BNPMDA. In conclusion, BNPMDA could effectively predict the potential miRNA-disease associations at a high accuracy level. Availability and implementation: BNPMDA is available via http://www.escience.cn/system/file?fileId=99559. Supplementary information: Supplementary data are available at Bioinformatics online. Xing Chen 0001, Di Xie, Lei Wang 0121, Qi Zhao 0010, Zhu-Hong You, Hongsheng Liu 0001 |
Bioinform. | 1 |
| 2018 | DroidDet: Effective and robust detection of android malware using static analysis along with rotation forest model
Zhu-Hong You, Zexuan Zhu 0001, Wei-Lei Shi, Xing Chen 0001 |
Neurocomputing | 5 |
| 2018 | MDHGI: Matrix Decomposition and Heterogeneous Graph Inference for miRNA-disease association predictionabstractRecently, a growing number of biological research and scientific experiments have demonstrated that microRNA (miRNA) affects the development of human complex diseases. Discovering miRNA-disease associations plays an increasingly vital role in devising diagnostic and therapeutic tools for diseases. However, since uncovering associations via experimental methods is expensive and time-consuming, novel and effective computational methods for association prediction are in demand. In this study, we developed a computational model of Matrix Decomposition and Heterogeneous Graph Inference for miRNA-disease association prediction (MDHGI) to discover new miRNA-disease associations by integrating the predicted association probability obtained from matrix decomposition through sparse learning method, the miRNA functional similarity, the disease semantic similarity, and the Gaussian interaction profile kernel similarity for diseases and miRNAs into a heterogeneous network. Compared with previous computational models based on heterogeneous networks, our model took full advantage of matrix decomposition before the construction of heterogeneous network, thereby improving the prediction accuracy. MDHGI obtained AUCs of 0.8945 and 0.8240 in the global and the local leave-one-out cross validation, respectively. Moreover, the AUC of 0.8794+/-0.0021 in 5-fold cross validation confirmed its stability of predictive performance. In addition, to further evaluate the model's accuracy, we applied MDHGI to four important human cancers in three different kinds of case studies. In the first type, 98% (Esophageal Neoplasms) and 98% (Lymphoma) of top 50 predicted miRNAs have been confirmed by at least one of the two databases (dbDEMC and miR2Disease) or at least one experimental literature in PubMed. In the second type of case study, what made a difference was that we removed all known associations between the miRNAs and Lung Neoplasms before implementing MDHGI on Lung Neoplasms. As a result, 100% (Lung Neoplasms) of top 50 related miRNAs have been indexed by at least one of the three databases (dbDEMC, miR2Disease and HMDD V2.0) or at least one experimental literature in PubMed. Furthermore, we also tested our prediction method on the HMDD V1.0 database to prove the applicability of MDHGI to different datasets. The results showed that 50 out of top 50 miRNAs related with the breast neoplasms were validated by at least one of the three databases (HMDD V2.0, dbDEMC, and miR2Disease) or at least one experimental literature. Xing Chen 0001 |
PLoS Comput. Biol. | 1 |
| 2018 | An improved efficient rotation forest algorithm to predict the interactions among proteins
Lei Wang 0121, Zhu-Hong You, Shixiong Xia, Xing Chen 0001, Yong Zhou 0003, Feng Liu 0039 |
Soft Comput. | 4 |
| 2017 | Computational Methods for the Prediction of Drug-Target Interactions from Drug Fingerprints and Protein Sequences by Stacked Auto-Encoder Deep Neural Network
Lei Wang 0121, Zhu-Hong You, Xing Chen 0001, Shixiong Xia, Feng Liu 0039, Yong Zhou 0003 |
ISBRA | 3 |
| 2017 | Long non-coding RNAs and complex diseases: from experimental results to computational modelsabstractLncRNAs have attracted lots of attentions from researchers worldwide in recent decades. With the rapid advances in both experimental technology and computational prediction algorithm, thousands of lncRNA have been identified in eukaryotic organisms ranging from nematodes to humans in the past few years. More and more research evidences have indicated that lncRNAs are involved in almost the whole life cycle of cells through different mechanisms and play important roles in many critical biological processes. Therefore, it is not surprising that the mutations and dysregulations of lncRNAs would contribute to the development of various human complex diseases. In this review, we first made a brief introduction about the functions of lncRNAs, five important lncRNA-related diseases, five critical disease-related lncRNAs and some important publicly available lncRNA-related databases about sequence, expression, function, etc. Nowadays, only a limited number of lncRNAs have been experimentally reported to be related to human diseases. Therefore, analyzing available lncRNA-disease associations and predicting potential human lncRNA-disease associations have become important tasks of bioinformatics, which would benefit human complex diseases mechanism understanding at lncRNA level, disease biomarker detection and disease diagnosis, treatment, prognosis and prevention. Furthermore, we introduced some state-of-the-art computational models, which could be effectively used to identify disease-related lncRNAs on a large scale and select the most promising disease-related lncRNAs for experimental validation. We also analyzed the limitations of these models and discussed the future directions of developing computational models for lncRNA research. Xing Chen 0001, Chenggang Yan 0001, Xu Zhang 0028, Zhu-Hong You |
Briefings Bioinform. | 1 |
| 2017 | A novel approach based on KATZ measure to predict associations of human microbiota with non-infectious diseasesabstractMotivation: Accumulating clinical observations have indicated that microbes living in the human body are closely associated with a wide range of human noninfectious diseases, which provides promising insights into the complex disease mechanism understanding. Predicting microbe-disease associations could not only boost human disease diagnostic and prognostic, but also improve the new drug development. However, little efforts have been attempted to understand and predict human microbe-disease associations on a large scale until now. Results: In this work, we constructed a microbe-human disease association network and further developed a novel computational model of KATZ measure for Human Microbe-Disease Association prediction (KATZHMDA) based on the assumption that functionally similar microbes tend to have similar interaction and non-interaction patterns with noninfectious diseases, and vice versa. To our knowledge, KATZHMDA is the first tool for microbe-disease association prediction. The reliable prediction performance could be attributed to the use of KATZ measurement, and the introduction of Gaussian interaction profile kernel similarity for microbes and diseases. LOOCV and k-fold cross validation were implemented to evaluate the effectiveness of this novel computational model based on known microbe-disease associations obtained from HMDAD database. As a result, KATZHMDA achieved reliable performance with average AUCs of 0.8130 ± 0.0054, 0.8301 ± 0.0033 and 0.8382 in 2-fold and 5-fold cross validation and LOOCV framework, respectively. It is anticipated that KATZHMDA could be used to obtain more novel microbes associated with important noninfectious human diseases and therefore benefit drug discovery and human medical improvement. Availability and Implementation: Matlab codes and dataset explored in this work are available at http://dwz.cn/4oX5mS . Contacts: [email protected] or [email protected] or [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Xing Chen 0001, Zhu-Hong You, Guiying Yan, Xuesong Wang 0001 |
Bioinform. | 1 |
| 2017 | HAMDA: Hybrid Approach for MiRNA-Disease Association prediction
Xing Chen 0001, Ya-Wei Niu, Guanghui Wang 0002, Guiying Yan |
J. Biomed. Informatics | 1 |
| 2017 | LRSSLMDA: Laplacian Regularized Sparse Subspace Learning for MiRNA-Disease Association predictionabstractPredicting novel microRNA (miRNA)-disease associations is clinically significant due to miRNAs' potential roles of diagnostic biomarkers and therapeutic targets for various human diseases. Previous studies have demonstrated the viability of utilizing different types of biological data to computationally infer new disease-related miRNAs. Yet researchers face the challenge of how to effectively integrate diverse datasets and make reliable predictions. In this study, we presented a computational model named Laplacian Regularized Sparse Subspace Learning for MiRNA-Disease Association prediction (LRSSLMDA), which projected miRNAs/diseases' statistical feature profile and graph theoretical feature profile to a common subspace. It used Laplacian regularization to preserve the local structures of the training data and a L1-norm constraint to select important miRNA/disease features for prediction. The strength of dimensionality reduction enabled the model to be easily extended to much higher dimensional datasets than those exploited in this study. Experimental results showed that LRSSLMDA outperformed ten previous models: the AUC of 0.9178 in global leave-one-out cross validation (LOOCV) and the AUC of 0.8418 in local LOOCV indicated the model's superior prediction accuracy; and the average AUC of 0.9181+/-0.0004 in 5-fold cross validation justified its accuracy and stability. In addition, three types of case studies further demonstrated its predictive power. Potential miRNAs related to Colon Neoplasms, Lymphoma, Kidney Neoplasms, Esophageal Neoplasms and Breast Neoplasms were predicted by LRSSLMDA. Respectively, 98%, 88%, 96%, 98% and 98% out of the top 50 predictions were validated by experimental evidences. Therefore, we conclude that LRSSLMDA would be a valuable computational tool for miRNA-disease association prediction. Xing Chen 0001 |
PLoS Comput. Biol. | 1 |
| 2017 | PBMDA: A novel and effective path-based computational model for miRNA-disease association predictionabstractIn the recent few years, an increasing number of studies have shown that microRNAs (miRNAs) play critical roles in many fundamental and important biological processes. As one of pathogenetic factors, the molecular mechanisms underlying human complex diseases still have not been completely understood from the perspective of miRNA. Predicting potential miRNA-disease associations makes important contributions to understanding the pathogenesis of diseases, developing new drugs, and formulating individualized diagnosis and treatment for diverse human complex diseases. Instead of only depending on expensive and time-consuming biological experiments, computational prediction models are effective by predicting potential miRNA-disease associations, prioritizing candidate miRNAs for the investigated diseases, and selecting those miRNAs with higher association probabilities for further experimental validation. In this study, Path-Based MiRNA-Disease Association (PBMDA) prediction model was proposed by integrating known human miRNA-disease associations, miRNA functional similarity, disease semantic similarity, and Gaussian interaction profile kernel similarity for miRNAs and diseases. This model constructed a heterogeneous graph consisting of three interlinked sub-graphs and further adopted depth-first search algorithm to infer potential miRNA-disease associations. As a result, PBMDA achieved reliable performance in the frameworks of both local and global LOOCV (AUCs of 0.8341 and 0.9169, respectively) and 5-fold cross validation (average AUC of 0.9172). In the cases studies of three important human diseases, 88% (Esophageal Neoplasms), 88% (Kidney Neoplasms) and 90% (Colon Neoplasms) of top-50 predicted miRNAs have been manually confirmed by previous experimental reports from literatures. Through the comparison performance between PBMDA and other previous models in case studies, the reliable performance also demonstrates that PBMDA could serve as a powerful computational tool to accelerate the identification of disease-miRNA associations. Zhu-Hong You, Zhi-an Huang, Zexuan Zhu 0001, Guiying Yan, Zhengwei Li 0001, Zhenkun Wen, Xing Chen 0001 |
PLoS Comput. Biol. | 7 |
| 2017 | PSPEL: In Silico Prediction of Self-Interacting Proteins from Amino Acids Sequences Using Ensemble LearningabstractSelf interacting proteins (SIPs) play an important role in various aspects of the structural and functional organization of the cell. Detecting SIPs is one of the most important issues in current molecular biology. Although a large number of SIPs data has been generated by experimental methods, wet laboratory approaches are both time-consuming and costly. In addition, they yield high false negative and positive rates. Thus, there is a great need for in silico methods to predict SIPs accurately and efficiently. In this study, a new sequence-based method is proposed to predict SIPs. The evolutionary information contained in Position-Specific Scoring Matrix (PSSM) is extracted from of protein with known sequence. Then, features are fed to an ensemble classifier to distinguish the self-interacting and non-self-interacting proteins. When performed on Saccharomyces cerevisiae and Human SIPs data sets, the proposed method can achieve high accuracies of 86.86 and 91.30 percent, respectively. Our method also shows a good performance when compared with the SVM classifier and previous methods. Consequently, the proposed method can be considered to be a novel promising tool to predict SIPs. Jianqiang Li 0001, Zhu-Hong You, Xiao Li 0007, Zhong Ming 0001, Xing Chen 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2016 | Drug-target interaction prediction: databases, web servers and computational modelsabstractIdentification of drug-target interactions is an important process in drug discovery. Although high-throughput screening and other biological assays are becoming available, experimental methods for drug-target interaction identification remain to be extremely costly, time-consuming and challenging even nowadays. Therefore, various computational models have been developed to predict potential drug-target associations on a large scale. In this review, databases and web servers involved in drug-target identification and drug discovery are summarized. In addition, we mainly introduced some state-of-the-art computational models for drug-target interactions prediction, including network-based method, machine learning-based method and so on. Specially, for the machine learning-based method, much attention was paid to supervised and semi-supervised models, which have essential difference in the adoption of negative samples. Although significant improvements for drug-target interaction prediction have been obtained by many effective computational models, both network-based and machine learning-based methods have their disadvantages, respectively. Furthermore, we discuss the future directions of the network-based drug discovery and network approach for personalized drug discovery based on personalized medicine, genome sequencing, tumor clone-based network and cancer hallmark-based network. Finally, we discussed the new evaluation validation framework and the formulation of drug-target interactions prediction problem by more realistic regression formulation based on quantitative bioactivity data. Xing Chen 0001, Chenggang Yan 0001, Xu Zhang 0028, Jian Yin 0003, Yongdong Zhang 0001 |
Briefings Bioinform. | 1 |
| 2016 | Sequence-based prediction of protein-protein interactions using weighted sparse representation model combined with global encodingabstractBACKGROUND: Proteins are the important molecules which participate in virtually every aspect of cellular function within an organism in pairs. Although high-throughput technologies have generated considerable protein-protein interactions (PPIs) data for various species, the processes of experimental methods are both time-consuming and expensive. In addition, they are usually associated with high rates of both false positive and false negative results. Accordingly, a number of computational approaches have been developed to effectively and accurately predict protein interactions. However, most of these methods typically perform worse when other biological data sources (e.g., protein structure information, protein domains, or gene neighborhoods information) are not available. Therefore, it is very urgent to develop effective computational methods for prediction of PPIs solely using protein sequence information. RESULTS: In this study, we present a novel computational model combining weighted sparse representation based classifier (WSRC) and global encoding (GE) of amino acid sequence. Two kinds of protein descriptors, composition and transition, are extracted for representing each protein sequence. On the basis of such a feature representation, novel weighted sparse representation based classifier is introduced to predict protein interaction class. When the proposed method was evaluated with the PPIs data of S. cerevisiae, Human and H. pylori, it achieved high prediction accuracies of 96.82, 97.66 and 92.83 % respectively. Extensive experiments were performed for cross-species PPIs prediction and the prediction accuracies were also very promising. CONCLUSIONS: To further evaluate the performance of the proposed method, we then compared its performance with the method based on support vector machine (SVM). The results show that the proposed method achieved a significant improvement. Thus, the proposed method is a very efficient method to predict PPIs and may be a useful supplementary tool for future proteomics studies. Zhu-Hong You, Xing Chen 0001, Keith C. C. Chan, Xin Luo 0001 |
BMC Bioinform. | 3 |
| 2016 | Construction of reliable protein-protein interaction networks using weighted sparse representation based classifier with pseudo substitution matrix representation features
Zhu-Hong You, Xiao Li 0007, Xing Chen 0001, Pengwei Hu 0001, Shuai Li 0002, Xin Luo 0001 |
Neurocomputing | 4 |
| 2016 | Distributed image understanding with semantic dictionary and semantic expansion
Liang Li 0003, Chenggang Yan 0001, Xing Chen 0001, Chunjie Zhang 0001, Jian Yin 0003, Baochen Jiang, Qingming Huang |
Neurocomputing | 3 |
| 2016 | NLLSS: Predicting Synergistic Drug Combinations Based on Semi-supervised LearningabstractFungal infection has become one of the leading causes of hospital-acquired infections with high mortality rates. Furthermore, drug resistance is common for fungus-causing diseases. Synergistic drug combinations could provide an effective strategy to overcome drug resistance. Meanwhile, synergistic drug combinations can increase treatment efficacy and decrease drug dosage to avoid toxicity. Therefore, computational prediction of synergistic drug combinations for fungus-causing diseases becomes attractive. In this study, we proposed similar nature of drug combinations: principal drugs which obtain synergistic effect with similar adjuvant drugs are often similar and vice versa. Furthermore, we developed a novel algorithm termed Network-based Laplacian regularized Least Square Synergistic drug combination prediction (NLLSS) to predict potential synergistic drug combinations by integrating different kinds of information such as known synergistic drug combinations, drug-target interactions, and drug chemical structures. We applied NLLSS to predict antifungal synergistic drug combinations and showed that it achieved excellent performance both in terms of cross validation and independent prediction. Finally, we performed biological experiments for fungal pathogen Candida albicans to confirm 7 out of 13 predicted antifungal synergistic drug combinations. NLLSS provides an efficient strategy to identify potential synergistic antifungal combinations. Xing Chen 0001, Biao Ren, Quanxin Wang, Lixin Zhang 0006, Guiying Yan |
PLoS Comput. Biol. | 1 |
| 2013 | Novel human lncRNA-disease association inference based on lncRNA expression profilesabstractMOTIVATION: More and more evidences have indicated that long-non-coding RNAs (lncRNAs) play critical roles in many important biological processes. Therefore, mutations and dysregulations of these lncRNAs would contribute to the development of various complex diseases. Developing powerful computational models for potential disease-related lncRNAs identification would benefit biomarker identification and drug discovery for human disease diagnosis, treatment, prognosis and prevention. RESULTS: In this article, we proposed the assumption that similar diseases tend to be associated with functionally similar lncRNAs. Then, we further developed the method of Laplacian Regularized Least Squares for LncRNA-Disease Association (LRLSLDA) in the semisupervised learning framework. Although known disease-lncRNA associations in the database are rare, LRLSLDA still obtained an AUC of 0.7760 in the leave-one-out cross validation, significantly improving the performance of previous methods. We also illustrated the performance of LRLSLDA is not sensitive (even robust) to the parameters selection and it can obtain a reliable performance in all the test classes. Plenty of potential disease-lncRNA associations were publicly released and some of them have been confirmed by recent results in biological experiments. It is anticipated that LRLSLDA could be an effective and important biological tool for biomedical research. AVAILABILITY: The code of LRLSLDA is freely available at http://asdcd.amss.ac.cn/Software/Details/2. Xing Chen 0001, Guiying Yan |
Bioinform. | 1 |