EDBT 2026 Demo / reviewers in the wild / expert
Jianyi Yang 0002
dblp:124/1315
· DBLP profile ↗
26ranked-venue papers
4as first author
10since 2021 · last 2025
0000-0003-2912-7737ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 26 · 4 first-author · 10 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | RNA language model and graph attention network for RNA and small molecule binding sites predictionabstractMOTIVATION: The structural complexities enable RNA to serve as a versatile molecular scaffold capable of binding small molecules with high specificity. Understanding these interactions is essential for elucidating RNA's role in disease mechanisms and developing RNA-targeted therapeutics. However, predicting RNA-small molecule binding sites remains a significant challenge due to their conformational flexibility, structural diversity, and the limited availability of high-resolution structural data. RESULTS: In this study, we propose RLsite, a novel computational framework integrating pre-trained RNA language models with graph attention networks (GAT) to predict small-molecule binding sites on RNA. Our method effectively captures both sequential and structural features of RNA by leveraging large-scale RNA sequence data to learn intrinsic patterns and processing graph-based RNA structures to highlight key topological and spatial features. Compared to existing methods, RLsite demonstrates superior accuracy, generalizability, and biological relevance, achieving a Precision of 0.749, a Recall of 0.654, an MCC of 0.474, and an AUC of 0.828 on the public test set, which significantly outperforms the previous models, such as CapBind (an AUC of 0.770), MultiModRLBP (an AUC of 0.780), and RNABind (an AUC of 0.471). Notably, a case study of the PreQ1 riboswitch has achieved strong predictive performance (AUC = 0.97, Recall = 0.9), and its predicted binding sites have been confirmed experimentally. These results underscore our method as a potentially powerful tool for RNA-targeted drug discovery and advancing our understanding of RNA-ligand interactions. AVAILABILITY AND IMPLEMENTATION: The resource codes and data can be accessed at https://github.com/SaisaiSun/RLsite. Saisai Sun, Jianyi Yang 0002, Lin Gao 0006, Pengyong Li |
Bioinform. | 2 |
| 2024 | RNA threading with secondary structure and sequence profileabstractMOTIVATION: RNA threading aims to identify remote homologies for template-based modeling of RNA 3D structure. Existing RNA alignment methods primarily rely on secondary structure alignment. They are often time- and memory-consuming, limiting large-scale applications. In addition, the accuracy is far from satisfactory. RESULTS: Using RNA secondary structure and sequence profile, we developed a novel RNA threading algorithm, named RNAthreader. To enhance the alignment process and minimize memory usage, a novel approach has been introduced to simplify RNA secondary structures into compact diagrams. RNAthreader employs a two-step methodology. Initially, integer programming and dynamic programming are combined to create an initial alignment for the simplified diagram. Subsequently, the final alignment is obtained using dynamic programming, taking into account the initial alignment derived from the previous step. The benchmark test on 80 RNAs illustrates that RNAthreader generates more accurate alignments than other methods, especially for RNAs with pseudoknots. Another benchmark, involving 30 RNAs from the RNA-Puzzles experiments, exhibits that the models constructed using RNAthreader templates have a lower average RMSD than those created by alternative methods. Remarkably, RNAthreader takes less than two hours to complete alignments with ∼5000 RNAs, which is 3-40 times faster than other methods. These compelling results suggest that RNAthreader is a promising algorithm for RNA template detection. AVAILABILITY AND IMPLEMENTATION: https://yanglab.qd.sdu.edu.cn/RNAthreader. Zongyang Du, Zhen-Ling Peng, Jianyi Yang 0002 |
Bioinform. | 3 |
| 2023 | A unified approach to protein domain parsing with inter-residue distance matrixabstractMOTIVATION: It is fundamental to cut multi-domain proteins into individual domains, for precise domain-based structural and functional studies. In the past, sequence-based and structure-based domain parsing was carried out independently with different methodologies. The recent progress in deep learning-based protein structure prediction provides the opportunity to unify sequence-based and structure-based domain parsing. RESULTS: Based on the inter-residue distance matrix, which can be either derived from the input structure or predicted by trRosettaX, we can decode the domain boundaries under a unified framework. We name the proposed method UniDoc. The principle of UniDoc is based on the well-accepted physical concept of maximizing intra-domain interaction while minimizing inter-domain interaction. Comprehensive tests on five benchmark datasets indicate that UniDoc outperforms other state-of-the-art methods in terms of both accuracy and speed, for both sequence-based and structure-based domain parsing. The major contribution of UniDoc is providing a unified framework for structure-based and sequence-based domain parsing. We hope that UniDoc would be a convenient tool for protein domain analysis. AVAILABILITY AND IMPLEMENTATION: https://yanglab.nankai.edu.cn/UniDoc/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Kun Zhu 0023, Hong Su, Zhen-Ling Peng, Jianyi Yang 0002 |
Bioinform. | 4 |
| 2022 | Toward the assessment of predicted inter-residue distanceabstractMOTIVATION: Significant progress has been achieved in distance-based protein folding, due to improved prediction of inter-residue distance by deep learning. Many efforts are thus made to improve distance prediction in recent years. However, it remains unknown what is the best way of objectively assessing the accuracy of predicted distance. RESULTS: A total of 19 metrics were proposed to measure the accuracy of predicted distance. These metrics were discussed and compared quantitatively on three benchmark datasets, with distance and structure models predicted by the trRosetta pipeline. The experiments show that a few metrics, such as distance precision, have a high correlation with the model accuracy measure TM-score (Pearson's correlation coefficient >0.7). In addition, the metrics are applied to rank the distance prediction groups in CASP14. The ranking by our metrics coincides largely with the official version. These data suggest that the proposed metrics are effective for measuring distance prediction. We anticipate that this study paves the way for objectively monitoring the progress of inter-residue distance prediction. A web server and a standalone package are provided to implement the proposed metrics. AVAILABILITY AND IMPLEMENTATION: http://yanglab.nankai.edu.cn/APD. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Zongyang Du, Zhen-Ling Peng, Jianyi Yang 0002 |
Bioinform. | 3 |
| 2022 | On Monomeric and Multimeric Structures-Based Protein-Ligand InteractionsabstractMany ligands simultaneously interact with multiple protein chains in quaternary structure (QS). However, a significant number of previous studies on template-based modeling of protein-ligand interactions were based on monomeric structure (MS), which may suffer from incomplete binding information. The defects of using MS rather than QS have not been systematically studied before. In this work, based on molecular docking experiments and binding free energy estimations, we performed a large-scale comparison of the protein-ligand interactions in both forms of structures. We found that 1) about 18.6 percent biologically relevant ligands bind multiple chains in QS simultaneously. 2) For more than 95 percent complexes with multiple chains involved in the interactions, the binding free energy is lower for the QS form than the MS form. 3) For over 70 percent complexes with multi-chain binding pockets, docking with QS yields more accurate ligand conformations than with MS. While for about 1.82 percent complexes, accurate docking conformations were obtained by MS. Based on this work, it is encouraged to make use of QS rather than MS in future studies on protein-ligand interactions. Yajun Dai, Yang Li 0042, Zhen-Ling Peng, Jianyi Yang 0002 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2021 | Human host status inference from temporal microbiome changes via recurrent neural networksabstractWith the rapid increase in sequencing data, human host status inference (e.g. healthy or sick) from microbiome data has become an important issue. Existing studies are mostly based on single-point microbiome composition, while it is rare that the host status is predicted from longitudinal microbiome data. However, single-point-based methods cannot capture the dynamic patterns between the temporal changes and host status. Therefore, it remains challenging to build good predictive models as well as scaling to different microbiome contexts. On the other hand, existing methods are mainly targeted for disease prediction and seldom investigate other host statuses. To fill the gap, we propose a comprehensive deep learning-based framework that utilizes longitudinal microbiome data as input to infer the human host status. Specifically, the framework is composed of specific data preparation strategies and a recurrent neural network tailored for longitudinal microbiome data. In experiments, we evaluated the proposed method on both semi-synthetic and real datasets based on different sequencing technologies and metagenomic contexts. The results indicate that our method achieves robust performance compared to other baseline and state-of-the-art classifiers and provides a significant reduction in prediction time. Xingjian Chen, Lingjing Liu, Jianyi Yang 0002, Ka-Chun Wong |
Briefings Bioinform. | 4 |
| 2021 | Recognition of small molecule-RNA binding sites using RNA sequence and structureabstractMOTIVATION: RNA molecules become attractive small molecule drug targets to treat disease in recent years. Computer-aided drug design can be facilitated by detecting the RNA sites that bind small molecules. However, very limited progress has been reported for the prediction of small molecule-RNA binding sites. RESULTS: We developed a novel method RNAsite to predict small molecule-RNA binding sites using sequence profile- and structure-based descriptors. RNAsite was shown to be competitive with the state-of-the-art methods on the experimental structures of two independent test sets. When predicted structure models were used, RNAsite outperforms other methods by a large margin. The possibility of improving RNAsite by geometry-based binding pocket detection was investigated. The influence of RNA structure's flexibility and the conformational changes caused by ligand binding on RNAsite were also discussed. RNAsite is anticipated to be a useful tool for the design of RNA-targeting small molecule drugs. AVAILABILITY AND IMPLEMENTATION: http://yanglab.nankai.edu.cn/RNAsite. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Hong Su, Zhen-Ling Peng, Jianyi Yang 0002 |
Bioinform. | 3 |
| 2021 | RNA inter-nucleotide 3D closeness prediction by deep residual neural networksabstractMOTIVATION: Recent years have witnessed that the inter-residue contact/distance in proteins could be accurately predicted by deep neural networks, which significantly improve the accuracy of predicted protein structure models. In contrast, fewer studies have been done for the prediction of RNA inter-nucleotide 3D closeness. RESULTS: We proposed a new algorithm named RNAcontact for the prediction of RNA inter-nucleotide 3D closeness. RNAcontact was built based on the deep residual neural networks. The covariance information from multiple sequence alignments and the predicted secondary structure were used as the input features of the networks. Experiments show that RNAcontact achieves the respective precisions of 0.8 and 0.6 for the top L/10 and L (where L is the length of an RNA) predictions on an independent test set, significantly higher than other evolutionary coupling methods. Analysis shows that about 1/3 of the correctly predicted 3D closenesses are not base pairings of secondary structure, which are critical to the determination of RNA structure. In addition, we demonstrated that the predicted 3D closeness could be used as distance restraints to guide RNA structure folding by the 3dRNA package. More accurate models could be built by using the predicted 3D closeness than the models without using 3D closeness. AVAILABILITY AND IMPLEMENTATION: The webserver and a standalone package are available at: http://yanglab.nankai.edu.cn/RNAcontact/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Saisai Sun, Zhen-Ling Peng, Jianyi Yang 0002 |
Bioinform. | 4 |
| 2021 | Improved estimation of model quality using predicted inter-residue distanceabstractMOTIVATION: Protein model quality assessment (QA) is an essential component in protein structure prediction, which aims to estimate the quality of a structure model and/or select the most accurate model out from a pool of structure models, without knowing the native structure. QA remains a challenging task in protein structure prediction. RESULTS: Based on the inter-residue distance predicted by the recent deep learning-based structure prediction algorithm trRosetta, we developed QDistance, a new approach to the estimation of both global and local qualities. QDistance works for both single- and multi-models inputs. We designed several distance-based features to assess the agreement between the predicted and model-derived inter-residue distances. Together with a few widely used features, they are fed into a simple yet powerful linear regression model to infer the global QA scores. The local QA scores for each structure model are predicted based on a comparative analysis with a set of selected reference models. For multi-models input, the reference models are selected from the input based on the predicted global QA scores. For single-model input, the reference models are predicted by trRosetta. With the informative distance-based features, QDistance can predict the global quality with satisfactory accuracy. Benchmark tests on the CASP13 and the CAMEO structure models suggested that QDistance was competitive with other methods. Blind tests in the CASP14 experiments showed that QDistance was robust and ranked among the top predictors. Especially, QDistance was the top 3 local QA method and made the most accurate local QA prediction for unreliable local region. Analysis showed that this superior performance can be attributed to the inclusion of the predicted inter-residue distance. AVAILABILITY AND IMPLEMENTATION: http://yanglab.nankai.edu.cn/QDistance. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Lisha Ye, Peikun Wu, Zhen-Ling Peng, Jianzhao Gao, Jian Liu 0040, Jianyi Yang 0002 |
Bioinform. | 6 |
| 2021 | RNA Flexibility Prediction With Sequence Profile and Predicted Solvent AccessibilityabstractStructural flexibility plays an essential role in many biological processes. B-factor is an important indicator to measure the flexibility of protein or RNA structures. Many methods were developed to predict protein B-factors, but few studies have been done for RNA B-factor prediction. In this paper, we proposed a new method RNAbval to predict RNA B-factors using random forest. The method was developed using a comprehensive set of features, including the sequence profile and predicted solvent accessibility. RNAbval achieved an improvement of 9.2-20.5 percent over the state-of-the-art method on two benchmark test datasets. The proposed method is available at http://yanglab.nankai.edu.cn/RNAbval/. Boling Wang, Jianyi Yang 0002, Jianzhao Gao |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2020 | CATHER: a novel threading algorithm with predicted contactsabstractMOTIVATION: Threading is one of the most effective methods for protein structure prediction. In recent years, the increasing accuracy in protein contact map prediction opens a new avenue to improve the performance of threading algorithms. Several preliminary studies suggest that with predicted contacts, the performance of threading algorithms can be improved greatly. There is still much room to explore to make better use of predicted contacts. RESULTS: We have developed a new contact-assisted threading algorithm named CATHER using both conventional sequential profiles and contact map predicted by a deep learning-based algorithm. Benchmark tests on an independent test set and the CASP12 targets demonstrated that CATHER made significant improvement over other methods which only use either sequential profile or predicted contact map. Our method was ranked at the Top 10 among all 39 participated server groups on the 32 free modeling targets in the blind tests of the CASP13 experiment. These data suggest that it is promising to push forward the threading algorithms by using predicted contacts. AVAILABILITY AND IMPLEMENTATION: http://yanglab.nankai.edu.cn/CATHER/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Zongyang Du, Shuo Pan, Qi Wu 0016, Zhen-Ling Peng, Jianyi Yang 0002 |
Bioinform. | 5 |
| 2020 | Protein contact prediction using metagenome sequence data and residual neural networksabstractMOTIVATION: Almost all protein residue contact prediction methods rely on the availability of deep multiple sequence alignments (MSAs). However, many proteins from the poorly populated families do not have sufficient number of homologs in the conventional UniProt database. Here we aim to solve this issue by exploring the rich sequence data from the metagenome sequencing projects. RESULTS: Based on the improved MSA constructed from the metagenome sequence data, we developed MapPred, a new deep learning-based contact prediction method. MapPred consists of two component methods, DeepMSA and DeepMeta, both trained with the residual neural networks. DeepMSA was inspired by the recent method DeepCov, which was trained on 441 matrices of covariance features. By considering the symmetry of contact map, we reduced the number of matrices to 231, which makes the training more efficient in DeepMSA. Experiments show that DeepMSA outperforms DeepCov by 10-13% in precision. DeepMeta works by combining predicted contacts and other sequence profile features. Experiments on three benchmark datasets suggest that the contribution from the metagenome sequence data is significant with P-values less than 4.04E-17. MapPred is shown to be complementary and comparable the state-of-the-art methods. The success of MapPred is attributed to three factors: the deeper MSA from the metagenome sequence data, improved feature design in DeepMSA and optimized training by the residual neural networks. AVAILABILITY AND IMPLEMENTATION: http://yanglab.nankai.edu.cn/mappred/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Qi Wu 0016, Zhen-Ling Peng, Ivan Anishchenko, Qian Cong, David Baker 0001, Jianyi Yang 0002 |
Bioinform. | 6 |
| 2019 | Improving the prediction of protein-nucleic acids binding residues via multiple sequence profiles and the consensus of complementary methodsabstractMOTIVATION: The interactions between protein and nucleic acids play a key role in various biological processes. Accurate recognition of the residues that bind nucleic acids can facilitate the study of uncharacterized protein-nucleic acids interactions. The accuracy of existing nucleic acids-binding residues prediction methods is relatively low. RESULTS: In this work, we introduce NucBind, a novel method for the prediction of nucleic acids-binding residues. NucBind combines the predictions from a support vector machine-based ab-initio method SVMnuc and a template-based method COACH-D. SVMnuc was trained with features from three complementary sequence profiles. COACH-D predicts the binding residues based on homologous templates identified from a nucleic acids-binding library. The proposed methods were assessed and compared with other peering methods on three benchmark datasets. Experimental results show that NucBind consistently outperforms other state-of-the-art methods. Though with higher accuracy, similar to many other ab-initio methods, cross prediction between DNA and RNA-binding residues was also observed in SVMnuc and NucBind. We attribute the success of NucBind to two folds. The first is the utilization of improved features extracted from three complementary sequence profiles in SVMnuc. The second is the combination of two complementary methods: the ab-initio method SVMnuc and the template-based method COACH-D. AVAILABILITY AND IMPLEMENTATION: http://yanglab.nankai.edu.cn/NucBind. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Hong Su, Mengchen Liu, Saisai Sun, Zhen-Ling Peng, Jianyi Yang 0002 |
Bioinform. | 5 |
| 2019 | Enhanced prediction of RNA solvent accessibility with long short-term memory neural networks and improved sequence profilesabstractMOTIVATION: The de novo prediction of RNA tertiary structure remains a grand challenge. Predicted RNA solvent accessibility provides an opportunity to address this challenge. To the best of our knowledge, there is only one method (RNAsnap) available for RNA solvent accessibility prediction. However, its performance is unsatisfactory for protein-free RNAs. RESULTS: We developed RNAsol, a new algorithm to predict RNA solvent accessibility. RNAsol was built based on improved sequence profiles from the covariance models and trained with the long short-term memory (LSTM) neural networks. Independent tests on the same datasets from RNAsnap show that RNAsol achieves the mean Pearson's correlation coefficient (PCC) of 0.43/0.26 for the protein-bound/protein-free RNA molecules, which is 26.5%/136.4% higher than that of RNAsnap. When the training set is enlarged to include both types of RNAs, the PCCs increase to 0.49 and 0.46 for protein-bound and protein-free RNAs, respectively. The success of RNAsol is attributed to two aspects, including the improved sequence profiles constructed by the sequence-profile alignment and the enhanced training by the LSTM neural networks. AVAILABILITY AND IMPLEMENTATION: http://yanglab.nankai.edu.cn/RNAsol/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Saisai Sun, Qi Wu 0016, Zhen-Ling Peng, Jianyi Yang 0002 |
Bioinform. | 4 |
| 2018 | A large-scale comparative assessment of methods for residue-residue contact predictionabstractSequence-based prediction of residue-residue contact in proteins becomes increasingly more important for improving protein structure prediction in the big data era. In this study, we performed a large-scale comparative assessment of 15 locally installed contact predictors. To assess these methods, we collected a big data set consisting of 680 nonredundant proteins covering different structural classes and target difficulties. We investigated a wide range of factors that may influence the precision of contact prediction, including target difficulty, structural class, the alignment depth and distribution of contact pairs in a protein structure. We found that: (1) the machine learning-based methods outperform the direct-coupling-based methods for short-range contact prediction, while the latter are significantly better for long-range contact prediction. The consensus-based methods, which combine machine learning and direct-coupling methods, perform the best. (2) The target difficulty does not have clear influence on the machine learning-based methods, while it does affect the direct-coupling and consensus-based methods significantly. (3) The alignment depth has relatively weak effect on the machine learning-based methods. However, for the direct-coupling-based methods and consensus-based methods, the predicted contacts for targets with deeper alignment tend to be more accurate. (4) All methods perform relatively better on β and α + β proteins than on α proteins. (5) Residues buried in the core of protein structure are more prone to be in contact than residues on the surface (22 versus 6%). We believe these are useful results for guiding future development of new approach to contact prediction. Qiqige Wuyun, Wei Zheng 0013, Zhen-Ling Peng, Jianyi Yang 0002 |
Briefings Bioinform. | 4 |
| 2018 | mTM-align: an algorithm for fast and accurate multiple protein structure alignmentabstractMotivation: As protein structure is more conserved than sequence during evolution, multiple structure alignment can be more informative than multiple sequence alignment, especially for distantly related proteins. With the rapid increase of the number of protein structures in the Protein Data Bank, it becomes urgent to develop efficient algorithms for multiple structure alignment. Results: A new multiple structure alignment algorithm (mTM-align) was proposed, which is an extension of the highly efficient pairwise structure alignment program TM-align. The algorithm was benchmarked on four widely used datasets, HOMSTRAD, SABmark_sup, SABmark_twi and SISY-multiple, showing that mTM-align consistently outperforms other algorithms. In addition, the comparison with the manually curated alignments in the HOMSTRAD database shows that the automated alignments built by mTM-align are in general more accurate. Therefore, mTM-align may be used as a reliable complement to construct multiple structure alignments for real-world applications. Availability and implementation: http://yanglab.nankai.edu.cn/mTM-align. Contact: [email protected] or [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Runze Dong, Zhen-Ling Peng, Yang Zhang 0040, Jianyi Yang 0002 |
Bioinform. | 4 |
| 2018 | CoABind: a novel algorithm for Coenzyme A (CoA)- and CoA derivatives-binding residues predictionabstractMotivation: Coenzyme A (CoA)-protein binding plays an important role in various cellular functions and metabolic pathways. However, no computational methods can be employed for CoA-binding residues prediction. Results: We developed three methods for the prediction of CoA- and CoA derivatives-binding residues, including an ab initio method SVMpred, a template-based method TemPred and a consensus-based method CoABind. In SVMpred, a comprehensive set of features are designed from two complementary sequence profiles and the predicted secondary structure and solvent accessibility. The engine for classification in SVMpred is selected as the support vector machine. For TemPred, the prediction is transferred from homologous templates in the training set, which are detected by the program HHsearch. The assessment on an independent test set consisting of 73 proteins shows that SVMpred and TemPred achieve Matthews correlation coefficient (MCC) of 0.438 and 0.481, respectively. Analysis on the predictions by SVMpred and TemPred shows that these two methods are complementary to each other. Therefore, we combined them together, forming the third method CoABind, which further improves the MCC to 0.489 on the same set. Experiments demonstrate that the proposed methods significantly outperform the state-of-the-art general-purpose ligand-binding residues prediction algorithm COACH. As the first-of-its-kind method, we anticipate CoABind to be helpful for studying CoA-protein interaction. Availability and implementation: http://yanglab.nankai.edu.cn/CoABind. Supplementary information: Supplementary data are available at Bioinformatics online. Qiaozhen Meng, Zhen-Ling Peng, Jianyi Yang 0002 |
Bioinform. | 3 |
| 2017 | DLTree: efficient and accurate phylogeny reconstruction using the dynamical language methodabstractSUMMARY: A number of alignment-free methods have been proposed for phylogeny reconstruction over the past two decades. But there are some long-standing challenges in these methods, including requirement of huge computer memory and CPU time, and existence of duplicate computations. In this article, we address these challenges with the idea of compressed vector, fingerprint and scalable memory management. With these ideas we developed the DLTree algorithm for efficient implementation of the dynamical language model and whole genome-based phylogenetic analysis. The DLTree algorithm was compared with other alignment-free tools, demonstrating that it is more efficient and accurate for phylogeny reconstruction. AVAILABILITY AND IMPLEMENTATION: The DLTree algorithm is freely available at http://dltree.xtu.edu.cn. CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Qi Wu 0016, Jianyi Yang 0002 |
Bioinform. | 3 |
| 2017 | An ensemble approach to protein fold classification by integration of template-based assignment and support vector machine classifierabstractMotivation: Protein fold classification is a critical step in protein structure prediction. There are two possible ways to classify protein folds. One is through template-based fold assignment and the other is ab-initio prediction using machine learning algorithms. Combination of both solutions to improve the prediction accuracy was never explored before. Results: We developed two algorithms, HH-fold and SVM-fold for protein fold classification. HH-fold is a template-based fold assignment algorithm using the HHsearch program. SVM-fold is a support vector machine-based ab-initio classification algorithm, in which a comprehensive set of features are extracted from three complementary sequence profiles. These two algorithms are then combined, resulting to the ensemble approach TA-fold. We performed a comprehensive assessment for the proposed methods by comparing with ab-initio methods and template-based threading methods on six benchmark datasets. An accuracy of 0.799 was achieved by TA-fold on the DD dataset that consists of proteins from 27 folds. This represents improvement of 5.4-11.7% over ab-initio methods. After updating this dataset to include more proteins in the same folds, the accuracy increased to 0.971. In addition, TA-fold achieved >0.9 accuracy on a large dataset consisting of 6451 proteins from 184 folds. Experiments on the LE dataset show that TA-fold consistently outperforms other threading methods at the family, superfamily and fold levels. The success of TA-fold is attributed to the combination of template-based fold assignment and ab-initio classification using features from complementary sequence profiles that contain rich evolution information. Availability and Implementation: http://yanglab.nankai.edu.cn/TA-fold/. Contact: [email protected] or [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Jiaqi Xia, Zhen-Ling Peng, Dawei Qi, Hongbo Mu, Jianyi Yang 0002 |
Bioinform. | 5 |
| 2016 | Recognizing metal and acid radical ion-binding sites by integrating ab initio modeling with template-based transferalsabstractMOTIVATION: More than half of proteins require binding of metal and acid radical ions for their structure and function. Identification of the ion-binding locations is important for understanding the biological functions of proteins. Due to the small size and high versatility of the metal and acid radical ions, however, computational prediction of their binding sites remains difficult. RESULTS: ) that are most frequently seen in protein databases. A sequence-based ab initio model is first trained on sequence profiles, where a modified AdaBoost algorithm is extended to balance binding and non-binding residue samples. A composite method IonCom is then developed to combine the ab initio model with multiple threading alignments for further improving the robustness of the binding site predictions. The pipeline was tested using 5-fold cross validations on a comprehensive set of 2,100 non-redundant proteins bound with 3,075 small ion ligands. Significant advantage was demonstrated compared with the state of the art ligand-binding methods including COACH and TargetS for high-accuracy ion-binding site identification. Detailed data analyses show that the major advantage of IonCom lies at the integration of complementary ab initio and template-based components. Ion-specific feature design and binding library selection also contribute to the improvement of small ion ligand binding predictions. AVAILABILITY AND IMPLEMENTATION: http://zhanglab.ccmb.med.umich.edu/IonCom CONTACT: [email protected] or [email protected] information: Supplementary data are available at Bioinformatics online. Xiuzhen Hu, Qiwen Dong, Jianyi Yang 0002, Yang Zhang 0040 |
Bioinform. | 3 |
| 2015 | GLASS: a comprehensive database for experimentally validated GPCR-ligand associationsabstractMOTIVATION: G protein-coupled receptors (GPCRs) are probably the most attractive drug target membrane proteins, which constitute nearly half of drug targets in the contemporary drug discovery industry. While the majority of drug discovery studies employ existing GPCR and ligand interactions to identify new compounds, there remains a shortage of specific databases with precisely annotated GPCR-ligand associations. RESULTS: We have developed a new database, GLASS, which aims to provide a comprehensive, manually curated resource for experimentally validated GPCR-ligand associations. A new text-mining algorithm was proposed to collect GPCR-ligand interactions from the biomedical literature, which is then crosschecked with five primary pharmacological datasets, to enhance the coverage and accuracy of GPCR-ligand association data identifications. A special architecture has been designed to allow users for making homologous ligand search with flexible bioactivity parameters. The current database contains ∼500 000 unique entries, of which the vast majority stems from ligand associations with rhodopsin- and secretin-like receptors. The GLASS database should find its most useful application in various in silico GPCR screening and functional annotation studies. AVAILABILITY AND IMPLEMENTATION: The website of GLASS database is freely available at http://zhanglab.ccmb.med.umich.edu/GLASS/. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Wallace K. B. Chan, Hongjiu Zhang, Jianyi Yang 0002, Jeffrey R. Brender, Junguk Hur, Arzucan Özgür, Yang Zhang 0040 |
Bioinform. | 3 |
| 2013 | Protein-ligand binding site recognition using complementary binding-specific substructure comparison and sequence profile alignmentabstractMOTIVATION: Identification of protein-ligand binding sites is critical to protein function annotation and drug discovery. However, there is no method that could generate optimal binding site prediction for different protein types. Combination of complementary predictions is probably the most reliable solution to the problem. RESULTS: We develop two new methods, one based on binding-specific substructure comparison (TM-SITE) and another on sequence profile alignment (S-SITE), for complementary binding site predictions. The methods are tested on a set of 500 non-redundant proteins harboring 814 natural, drug-like and metal ion molecules. Starting from low-resolution protein structure predictions, the methods successfully recognize >51% of binding residues with average Matthews correlation coefficient (MCC) significantly higher (with P-value <10(-9) in student t-test) than other state-of-the-art methods, including COFACTOR, FINDSITE and ConCavity. When combining TM-SITE and S-SITE with other structure-based programs, a consensus approach (COACH) can increase MCC by 15% over the best individual predictions. COACH was examined in the recent community-wide COMEO experiment and consistently ranked as the best method in last 22 individual datasets with the Area Under the Curve score 22.5% higher than the second best method. These data demonstrate a new robust approach to protein-ligand binding site recognition, which is ready for genome-wide structure-based function annotations. AVAILABILITY: http://zhanglab.ccmb.med.umich.edu/COACH/ Jianyi Yang 0002, Ambrish Roy, Yang Zhang 0040 |
Bioinform. | 1 |
| 2011 | A Consensus Approach to Predicting Protein Contact Map via Logistic Regression
Jianyi Yang 0002, Xin Chen 0037 |
ISBRA | 1 |
| 2010 | An improved classification of G-protein-coupled receptors using sequence-derived featuresabstractBACKGROUND: G-protein-coupled receptors (GPCRs) play a key role in diverse physiological processes and are the targets of almost two-thirds of the marketed drugs. The 3 D structures of GPCRs are largely unavailable; however, a large number of GPCR primary sequences are known. To facilitate the identification and characterization of novel receptors, it is therefore very valuable to develop a computational method to accurately predict GPCRs from the protein primary sequences. RESULTS: We propose a new method called PCA-GPCR, to predict GPCRs using a comprehensive set of 1497 sequence-derived features. The principal component analysis is first employed to reduce the dimension of the feature space to 32. Then, the resulting 32-dimensional feature vectors are fed into a simple yet powerful classification algorithm, called intimate sorting, to predict GPCRs at five levels. The prediction at the first level determines whether a protein is a GPCR or a non-GPCR. If it is predicted to be a GPCR, then it will be further predicted into certain family, subfamily, sub-subfamily and subtype by the classifiers at the second, third, fourth, and fifth levels, respectively. To train the classifiers applied at five levels, a non-redundant dataset is carefully constructed, which contains 3178, 1589, 4772, 4924, and 2741 protein sequences at the respective levels. Jackknife tests on this training dataset show that the overall accuracies of PCA-GPCR at five levels (from the first to the fifth) can achieve up to 99.5%, 88.8%, 80.47%, 80.3%, and 92.34%, respectively. We further perform predictions on a dataset of 1238 GPCRs at the second level, and on another two datasets of 167 and 566 GPCRs respectively at the fourth level. The overall prediction accuracies of our method are consistently higher than those of the existing methods to be compared. CONCLUSIONS: The comprehensive set of 1497 features is believed to be capable of capturing information about amino acid composition, sequence order as well as various physicochemical properties of proteins. Therefore, high accuracies are achieved when predicting GPCRs at all the five levels with our proposed method. Zhen-Ling Peng, Jianyi Yang 0002, Xin Chen 0037 |
BMC Bioinform. | 2 |
| 2010 | Prediction of protein structural classes for low-homology sequences based on predicted secondary structureabstractBACKGROUND: Prediction of protein structural classes (alpha, beta, alpha + beta and alpha/beta) from amino acid sequences is of great importance, as it is beneficial to study protein function, regulation and interactions. Many methods have been developed for high-homology protein sequences, and the prediction accuracies can achieve up to 90%. However, for low-homology sequences whose average pairwise sequence identity lies between 20% and 40%, they perform relatively poorly, yielding the prediction accuracy often below 60%. RESULTS: We propose a new method to predict protein structural classes on the basis of features extracted from the predicted secondary structures of proteins rather than directly from their amino acid sequences. It first uses PSIPRED to predict the secondary structure for each protein sequence. Then, the chaos game representation is employed to represent the predicted secondary structure as two time series, from which we generate a comprehensive set of 24 features using recurrence quantification analysis, K-string based information entropy and segment-based analysis. The resulting feature vectors are finally fed into a simple yet powerful Fisher's discriminant algorithm for the prediction of protein structural classes. We tested the proposed method on three benchmark datasets in low homology and achieved the overall prediction accuracies of 82.9%, 83.1% and 81.3%, respectively. Comparisons with ten existing methods showed that our method consistently performs better for all the tested datasets and the overall accuracy improvements range from 2.3% to 27.5%. A web server that implements the proposed method is freely available at http://www1.spms.ntu.edu.sg/~chenxin/RKS_PPSC/. CONCLUSION: The high prediction accuracy achieved by our proposed method is attributed to the design of a comprehensive feature set on the predicted secondary structure sequences, which is capable of characterizing the sequence order information, local interactions of the secondary structural elements, and spacial arrangements of alpha helices and beta strands. Thus, it is a valuable method to predict protein structural classes particularly for low-homology amino acid sequences. Jianyi Yang 0002, Zhen-Ling Peng, Xin Chen 0037 |
BMC Bioinform. | 1 |
| 2008 | Human Pol II promoter recognition based on primary sequences and free energy of dinucleotidesabstractBACKGROUND: Promoter region plays an important role in determining where the transcription of a particular gene should be initiated. Computational prediction of eukaryotic Pol II promoter sequences is one of the most significant problems in sequence analysis. Existing promoter prediction methods are still far from being satisfactory. RESULTS: We attempt to recognize the human Pol II promoter sequences from the non-promoter sequences which are made up of exon and intron sequences. Four methods are used: two kinds of multifractal analysis performed on the numeric sequences obtained from the dinucleotide free energy, Z curve analysis and global descriptor of the promoter/non-promoter primary sequences. A total of 141 parameters are extracted from these methods and categorized into seven groups (methods). They are used to generate certain spaces and then each promoter/non-promoter sequence is represented by a point in the corresponding space. All the 120 possible combinations of the seven methods are tested. Based on Fisher's linear discriminant algorithm, with a relatively smaller number of parameters (96 and 117), we get satisfactory discriminant accuracies. Particularly, in the case of 117 parameters, the accuracies for the training and test sets reach 90.43% and 89.79%, respectively. A comparison with five other existing methods indicates that our methods have a better performance. Using the global descriptor method (36 parameters), 17 of the 18 experimentally verified promoter sequences of human chromosome 22 are correctly identified. CONCLUSION: The high accuracies achieved suggest that the methods of this paper are useful for understanding the difficult problem of promoter prediction. Jianyi Yang 0002, Vo V. Anh, Li-Qian Zhou |
BMC Bioinform. | 1 |