EDBT 2026 Demo / reviewers in the wild / expert
Kuo-Chen Chou
dblp:64/6584
· DBLP profile ↗
40ranked-venue papers
2as first author
2since 2021 · last 2021
0000-0001-8857-7063ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 40 · 2 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
30 papers |
Bioinformatics and computational biology · 100% |
Topics — the 30 heaviest of 39, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology
protein function prediction |
1.4 | 6 | 2019 | Bastion3: a two-layer ensemble predictor of type III secreted effectors · Bioinform. 2019 PROSPERous: high-throughput prediction of substrate cleavage sites for 90 proteases with improved accuracy · Bioinform. 2018 pLoc-mHum: predict subcellular localization of multi-location human proteins via general PseAAC to winnow out the crucial GO information · Bioinform. 2018 |
Bioinformatics and computational biology
genomics |
1.2 | 4 | 2018 | iPromoter-2L: a two-layer predictor for identifying promoters and their types by multi-window-based PseKNC · Bioinform. 2018 iRO-3wPseKNC: identify DNA replication origins by three-window-based PseKNC · Bioinform. 2018 iRSpot-EL: identify recombination spots with an ensemble learning approach · Bioinform. 2017 |
Bioinformatics and computational biology › protein function prediction
protein subcellular localization prediction |
1.1 | 5 | 2019 | pLoc_bal-mAnimal: predict subcellular localization of animal proteins by balancing training dataset and PseAAC · Bioinform. 2019 pLoc-mHum: predict subcellular localization of multi-location human proteins via general PseAAC to winnow out the crucial GO information · Bioinform. 2018 pLoc-mAnimal: predict subcellular localization of animal proteins with both single and multiple sites · Bioinform. 2017 |
Bioinformatics and computational biology › gene regulation › promoter analysis
promoter prediction |
0.7 | 2 | 2019 | MULTiPly: a novel multi-layer predictor for discovering general and specific types of promoters · Bioinform. 2019 iPromoter-2L: a two-layer predictor for identifying promoters and their types by multi-window-based PseKNC · Bioinform. 2018 |
Bioinformatics and computational biology
protein sequence analysis |
0.6 | 2 | 2018 | iFeature: a Python package and web server for features extraction and selection from protein and peptide sequences · Bioinform. 2018 POSSUM: a bioinformatics toolkit for generating numerical sequence feature descriptors based on PSSM profiles · Bioinform. 2017 |
Bioinformatics and computational biology › gene regulation › regulatory element discovery
enhancer prediction |
0.6 | 2 | 2018 | iEnhancer-EL: identifying enhancers and their strength with ensemble learning approach · Bioinform. 2018 iEnhancer-2L: a two-layer predictor for identifying enhancers and their strength by pseudo k-tuple nucleotide composition · Bioinform. 2016 |
Bioinformatics and computational biology › sequence analysis
sequence-based prediction |
0.5 | 2 | 2018 | iLoc-lncRNA: predict the subcellular location of lncRNAs by incorporating octamer composition into general PseKNC · Bioinform. 2018 iNuc-PseKNC: a sequence-based predictor for predicting nucleosome positioning in genomes with pseudo k-tuple nucleotide composition · Bioinform. 2014 |
Bioinformatics and computational biology › proteomics
post-translational modification site prediction |
0.5 | 2 | 2016 | iPTM-mLys: identifying multiple lysine PTM sites and their different types · Bioinform. 2016 pSumo-CD: predicting sumoylation sites in proteins with covariance discriminant algorithm by incorporating sequence-coupled effects into general PseAAC · Bioinform. 2016 |
Bioinformatics and computational biology › gene regulation
regulatory element discovery |
0.5 | 2 | 2016 | iDHS-EL: identifying DNase I hypersensitive sites by fusing three different modes of pseudo nucleotide composition into an ensemble learning framework · Bioinform. 2016 iEnhancer-2L: a two-layer predictor for identifying enhancers and their strength by pseudo k-tuple nucleotide composition · Bioinform. 2016 |
Bioinformatics and computational biology › sequence analysis
DNA sequence analysis |
0.4 | 2 | 2015 | repDNA: a Python package to generate various modes of feature vectors for DNA sequences by incorporating user-defined physicochemical properties and sequence-order effects · Bioinform. 2015 PseKNC-General: a cross-platform package for generating various modes of pseudo nucleotide compositions · Bioinform. 2015 |
Bioinformatics and computational biology
feature selection |
0.4 | 1 | 2019 | MULTiPly: a novel multi-layer predictor for discovering general and specific types of promoters · Bioinform. 2019 |
Bioinformatics and computational biology › protein function prediction
type III secreted effector prediction |
0.4 | 1 | 2019 | Bastion3: a two-layer ensemble predictor of type III secreted effectors · Bioinform. 2019 |
Bioinformatics and computational biology › proteomics › post-translational modification site prediction
protease cleavage site prediction |
0.4 | 2 | 2018 | PROSPERous: high-throughput prediction of substrate cleavage sites for 90 proteases with improved accuracy · Bioinform. 2018 Bio-support vector machines for computational proteomics · Bioinform. 2004 |
Bioinformatics and computational biology › gene regulation
gene regulation analysis |
0.3 | 1 | 2018 | iEnhancer-EL: identifying enhancers and their strength with ensemble learning approach · Bioinform. 2018 |
Bioinformatics and computational biology › RNA biology › RNA analysis › RNA bioinformatics
lncRNA subcellular localization prediction |
0.3 | 1 | 2018 | iLoc-lncRNA: predict the subcellular location of lncRNAs by incorporating octamer composition into general PseKNC · Bioinform. 2018 |
Bioinformatics and computational biology › proteomics › post-translational modification prediction
phosphorylation site prediction |
0.3 | 1 | 2018 | Quokka: a comprehensive tool for rapid and accurate prediction of kinase family-specific phosphorylation sites in the human proteome · Bioinform. 2018 |
Bioinformatics and computational biology › proteomics
post-translational modification prediction |
0.3 | 1 | 2018 | Quokka: a comprehensive tool for rapid and accurate prediction of kinase family-specific phosphorylation sites in the human proteome · Bioinform. 2018 |
Bioinformatics and computational biology › genome annotation
replication origin prediction |
0.3 | 1 | 2018 | iRO-3wPseKNC: identify DNA replication origins by three-window-based PseKNC · Bioinform. 2018 |
Bioinformatics and computational biology › RNA biology › RNA analysis
RNA bioinformatics |
0.3 | 1 | 2018 | iLoc-lncRNA: predict the subcellular location of lncRNAs by incorporating octamer composition into general PseKNC · Bioinform. 2018 |
Bioinformatics and computational biology
sequence analysis |
0.3 | 1 | 2018 | iPromoter-2L: a two-layer predictor for identifying promoters and their types by multi-window-based PseKNC · Bioinform. 2018 |
Bioinformatics and computational biology › drug discovery
anatomical therapeutic chemical classification |
0.3 | 1 | 2017 | iATC-mISF: a multi-label classifier for predicting the classes of anatomical therapeutic chemicals · Bioinform. 2017 |
Bioinformatics and computational biology › molecular informatics › cheminformatics
chemogenomics |
0.3 | 1 | 2017 | iATC-mISF: a multi-label classifier for predicting the classes of anatomical therapeutic chemicals · Bioinform. 2017 |
Bioinformatics and computational biology
drug discovery |
0.3 | 1 | 2017 | iATC-mISF: a multi-label classifier for predicting the classes of anatomical therapeutic chemicals · Bioinform. 2017 |
Bioinformatics and computational biology › genomics
genome analysis |
0.2 | 1 | 2016 | iEnhancer-2L: a two-layer predictor for identifying enhancers and their strength by pseudo k-tuple nucleotide composition · Bioinform. 2016 |
Bioinformatics and computational biology › protein analysis
protein bioinformatics |
0.2 | 1 | 2016 | pSumo-CD: predicting sumoylation sites in proteins with covariance discriminant algorithm by incorporating sequence-coupled effects into general PseAAC · Bioinform. 2016 |
Bioinformatics and computational biology › sequence analysis
sequence feature extraction |
0.2 | 1 | 2015 | PseKNC-General: a cross-platform package for generating various modes of pseudo nucleotide compositions · Bioinform. 2015 |
Bioinformatics and computational biology › epigenomics › chromatin analysis
nucleosome positioning |
0.2 | 1 | 2014 | iNuc-PseKNC: a sequence-based predictor for predicting nucleosome positioning in genomes with pseudo k-tuple nucleotide composition · Bioinform. 2014 |
Bioinformatics and computational biology › protein sequence analysis
protein homology detection |
0.2 | 1 | 2014 | Combining evolutionary information extracted from frequency profiles with sequence-based kernels for protein remote homology detection · Bioinform. 2014 |
Bioinformatics and computational biology › protein analysis › protein bioinformatics
protein representation |
0.2 | 1 | 2014 | Combining evolutionary information extracted from frequency profiles with sequence-based kernels for protein remote homology detection · Bioinform. 2014 |
Bioinformatics and computational biology › sequence analysis › homology detection
remote homology detection |
0.2 | 1 | 2014 | Combining evolutionary information extracted from frequency profiles with sequence-based kernels for protein remote homology detection · Bioinform. 2014 |
Methods — techniques the papers use, named apart from their topics
ensemble learning · 2.2support vector machine · 1.7pseudo k-tuple nucleotide composition · 1.4cross-validation · 1.2PseAAC · 1.0multi-label classification · 0.9logistic regression · 0.7kmer · 0.6random forest · 0.5pseudo nucleotide composition · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | Using CHOU'S 5-Steps Rule to Predict O-Linked Serine Glycosylation Sites by Blending Position Relative Features and Statistical MomentabstractGlycosylation of proteins in eukaryote cells is an important and complicated post-translation modification due to its pivotal role and association with crucial physiological functions within most of the proteins. Identification of glycosylation sites in a polypeptide chain is not an easy task due to multiple impediments. Analytical identification of these sites is expensive and laborious. There is a dire need to develop a reliable computational method for precise determination of such sites which can help researchers to save time and effort. Herein, we propose a novel predictor namely iGlycoS-PseAAC by integrating the Chou's Pseudo Amino Acid Composition (PseAAC) and relative/absolute position-based features. The self-consistency results show that the accuracy revealed by the model using the benchmark dataset for prediction of O-linked glycosylation having serine sites is 98.8 percent. The overall accuracy of predictor achieved through 10-fold cross validation by combining the positive and negative results is 97.2 percent. The overall accuracy achieved through Jackknife test is 96.195 percent by aggregating of all the prediction results. Thus the proposed predictor can help in predicting the O-linked glycosylated serine sites in an efficient and accurate way. The overall results show that the accuracy of the iGlycoS-PseAAC is higher than the existing tools. Muhammad Aizaz Akmal, Nouman Rasool, Yaser Daanial Khan, Sher Afzal Khan, Kuo-Chen Chou |
IEEE ACM Trans. Comput. Biol. Bioinform. | 6 |
| 2021 | iPhosH-PseAAC: Identify Phosphohistidine Sites in Proteins by Blending Statistical Moments and Position Relative Features According to the Chou's 5-Step Rule and General Pseudo Amino Acid CompositionabstractProtein phosphorylation is one of the key mechanism in prokaryotes and eukaryotes and is responsible for various biological functions such as protein degradation, intracellular localization, the multitude of cellular processes, molecular association, cytoskeletal dynamics, and enzymatic inhibition/activation. Phosphohistidine (PhosH) has a key role in a number of biological processes, including central metabolism to signalling in eukaryotes and bacteria. Thus, identification of phosphohistidine sites in a protein sequence is crucial, and experimental identification can be expensive, time-taking, and laborious. To address this problem, here, we propose a novel computational model namely iPhosH-PseAAC for prediction of phosphohistidine sites in a given protein sequence using pseudo amino acid composition (PseAAC), statistical moments, and position relative features. The results of the proposed predictor are validated through self-consistency testing, 10-fold cross-validation, and jackknife testing. The self-consistency validation gave the 100 percent accuracy, whereas, for cross-validation, the accuracy achieved is 94.26 percent. Moreover, jackknife testing gave 97.07 percent accuracy for the proposed model. Thus, the proposed model iPhosH-PseAAC for prediction of iPhosH site has the great ability to predict the PhosH sites in given proteins. Muhammad Awais 0011, Yaser Daanial Khan, Nouman Rasool, Sher Afzal Khan, Kuo-Chen Chou |
IEEE ACM Trans. Comput. Biol. Bioinform. | 6 |
| 2020 | iLearn : an integrated platform and meta-learner for feature engineering, machine-learning analysis and modeling of DNA, RNA and protein sequence dataabstractWith the explosive growth of biological sequences generated in the post-genomic era, one of the most challenging problems in bioinformatics and computational biology is to computationally characterize sequences, structures and functions in an efficient, accurate and high-throughput manner. A number of online web servers and stand-alone tools have been developed to address this to date; however, all these tools have their limitations and drawbacks in terms of their effectiveness, user-friendliness and capacity. Here, we present iLearn, a comprehensive and versatile Python-based toolkit, integrating the functionality of feature extraction, clustering, normalization, selection, dimensionality reduction, predictor construction, best descriptor/model selection, ensemble learning and results visualization for DNA, RNA and protein sequences. iLearn was designed for users that only want to upload their data set and select the functions they need calculated from it, while all necessary procedures and optimal settings are completed automatically by the software. iLearn includes a variety of descriptors for DNA, RNA and proteins, and four feature output formats are supported so as to facilitate direct output usage or communication with other computational tools. In total, iLearn encompasses 16 different types of feature clustering, selection, normalization and dimensionality reduction algorithms, and five commonly used machine-learning algorithms, thereby greatly facilitating feature analysis and predictor construction. iLearn is made freely available via an online web server and a stand-alone toolkit. Zhen Chen 0009, Fuyi Li, Tatiana T. Marquez-Lago, André Leier, Jerico Revote, Yan Zhu 0006, David R. Powell, Tatsuya Akutsu, Geoffrey I. Webb, Kuo-Chen Chou, Alexander Ian Smith, Roger J. Daly, Jian Li 0052, Jiangning Song |
Briefings Bioinform. | 11 |
| 2019 | Large-scale comparative assessment of computational predictors for lysine post-translational modification sitesabstractLysine post-translational modifications (PTMs) play a crucial role in regulating diverse functions and biological processes of proteins. However, because of the large volumes of sequencing data generated from genome-sequencing projects, systematic identification of different types of lysine PTM substrates and PTM sites in the entire proteome remains a major challenge. In recent years, a number of computational methods for lysine PTM identification have been developed. These methods show high diversity in their core algorithms, features extracted and feature selection techniques and evaluation strategies. There is therefore an urgent need to revisit these methods and summarize their methodologies, to improve and further develop computational techniques to identify and characterize lysine PTMs from the large amounts of sequence data. With this goal in mind, we first provide a comprehensive survey on a large collection of 49 state-of-the-art approaches for lysine PTM prediction. We cover a variety of important aspects that are crucial for the development of successful predictors, including operating algorithms, sequence and structural features, feature selection, model performance evaluation and software utility. We further provide our thoughts on potential strategies to improve the model performance. Second, in order to examine the feasibility of using deep learning for lysine PTM prediction, we propose a novel computational framework, termed MUscADEL (Multiple Scalable Accurate Deep Learner for lysine PTMs), using deep, bidirectional, long short-term memory recurrent neural networks for accurate and systematic mapping of eight major types of lysine PTMs in the human and mouse proteomes. Extensive benchmarking tests show that MUscADEL outperforms current methods for lysine PTM characterization, demonstrating the potential and power of deep learning techniques in protein PTM prediction. The web server of MUscADEL, together with all the data sets assembled in this study, is freely available at http://muscadel.erc.monash.edu/. We anticipate this comprehensive review and the application of deep learning will provide practical guide and useful insights into PTM prediction and inspire future bioinformatics studies in the related fields. Zhen Chen 0009, Xuhan Liu, Fuyi Li, Chen Li 0021, Tatiana T. Marquez-Lago, André Leier, Tatsuya Akutsu, Geoffrey I. Webb, Dakang Xu, Alexander Ian Smith, Lei Li 0013, Kuo-Chen Chou, Jiangning Song |
Briefings Bioinform. | 12 |
| 2019 | Twenty years of bioinformatics research for protease-specific substrate and cleavage site prediction: a comprehensive revisit and benchmarking of existing methodsabstractThe roles of proteolytic cleavage have been intensively investigated and discussed during the past two decades. This irreversible chemical process has been frequently reported to influence a number of crucial biological processes (BPs), such as cell cycle, protein regulation and inflammation. A number of advanced studies have been published aiming at deciphering the mechanisms of proteolytic cleavage. Given its significance and the large number of functionally enriched substrates targeted by specific proteases, many computational approaches have been established for accurate prediction of protease-specific substrates and their cleavage sites. Consequently, there is an urgent need to systematically assess the state-of-the-art computational approaches for protease-specific cleavage site prediction to further advance the existing methodologies and to improve the prediction performance. With this goal in mind, in this article, we carefully evaluated a total of 19 computational methods (including 8 scoring function-based methods and 11 machine learning-based methods) in terms of their underlying algorithm, calculated features, performance evaluation and software usability. Then, extensive independent tests were performed to assess the robustness and scalability of the reviewed methods using our carefully prepared independent test data sets with 3641 cleavage sites (specific to 10 proteases). The comparative experimental results demonstrate that PROSPERous is the most accurate generic method for predicting eight protease-specific cleavage sites, while GPS-CCD and LabCaS outperformed other predictors for calpain-specific cleavage sites. Based on our review, we then outlined some potential ways to improve the prediction performance and ease the computational burden by applying ensemble learning, deep learning, positive unlabeled learning and parallel and distributed computing techniques. We anticipate that our study will serve as a practical and useful guide for interested readers to further advance next-generation bioinformatics tools for protease-specific cleavage site prediction. Fuyi Li, Yanan Wang 0003, Chen Li 0021, Tatiana T. Marquez-Lago, André Leier, Neil D. Rawlings, Gholamreza Haffari, Jerico Revote, Tatsuya Akutsu, Kuo-Chen Chou, Anthony W. Purcell, Robert N. Pike, Geoffrey I. Webb, Alexander Ian Smith, Trevor Lithgow, Roger J. Daly, James C. Whisstock, Jiangning Song |
Briefings Bioinform. | 10 |
| 2019 | iProt-Sub: a comprehensive package for accurately mapping and predicting protease-specific substrates and cleavage sitesabstractRegulation of proteolysis plays a critical role in a myriad of important cellular processes. The key to better understanding the mechanisms that control this process is to identify the specific substrates that each protease targets. To address this, we have developed iProt-Sub, a powerful bioinformatics tool for the accurate prediction of protease-specific substrates and their cleavage sites. Importantly, iProt-Sub represents a significantly advanced version of its successful predecessor, PROSPER. It provides optimized cleavage site prediction models with better prediction performance and coverage for more species-specific proteases (4 major protease families and 38 different proteases). iProt-Sub integrates heterogeneous sequence and structural features and uses a two-step feature selection procedure to further remove redundant and irrelevant features in an effort to improve the cleavage site prediction accuracy. Features used by iProt-Sub are encoded by 11 different sequence encoding schemes, including local amino acid sequence profile, secondary structure, solvent accessibility and native disorder, which will allow a more accurate representation of the protease specificity of approximately 38 proteases and training of the prediction models. Benchmarking experiments using cross-validation and independent tests showed that iProt-Sub is able to achieve a better performance than several existing generic tools. We anticipate that iProt-Sub will be a powerful tool for proteome-wide prediction of protease-specific substrates and their cleavage sites, and will facilitate hypothesis-driven functional interrogation of protease-specific substrate cleavage and proteolytic events. Jiangning Song, Yanan Wang 0003, Fuyi Li, Tatsuya Akutsu, Neil D. Rawlings, Geoffrey I. Webb, Kuo-Chen Chou |
Briefings Bioinform. | 7 |
| 2019 | Computational analysis and prediction of lysine malonylation sites by exploiting informative features in an integrative machine-learning frameworkabstractAs a newly discovered post-translational modification (PTM), lysine malonylation (Kmal) regulates a myriad of cellular processes from prokaryotes to eukaryotes and has important implications in human diseases. Despite its functional significance, computational methods to accurately identify malonylation sites are still lacking and urgently needed. In particular, there is currently no comprehensive analysis and assessment of different features and machine learning (ML) methods that are required for constructing the necessary prediction models. Here, we review, analyze and compare 11 different feature encoding methods, with the goal of extracting key patterns and characteristics from residue sequences of Kmal sites. We identify optimized feature sets, with which four commonly used ML methods (random forest, support vector machines, K-nearest neighbor and logistic regression) and one recently proposed [Light Gradient Boosting Machine (LightGBM)] are trained on data from three species, namely, Escherichia coli, Mus musculus and Homo sapiens, and compared using randomized 10-fold cross-validation tests. We show that integration of the single method-based models through ensemble learning further improves the prediction performance and model robustness on the independent test. When compared to the existing state-of-the-art predictor, MaloPred, the optimal ensemble models were more accurate for all three species (AUC: 0.930, 0.923 and 0.944 for E. coli, M. musculus and H. sapiens, respectively). Using the ensemble models, we developed an accessible online predictor, kmal-sp, available at http://kmalsp.erc.monash.edu/. We hope that this comprehensive survey and the proposed strategy for building more accurate models can serve as a useful guide for inspiring future developments of computational methods for PTM site prediction, expedite the discovery of new malonylation and other PTM types and facilitate hypothesis-driven experimental validation of novel malonylated substrates and malonylation sites. Yanju Zhang, Ruopeng Xie, Jiawei Wang 0002, André Leier, Tatiana T. Marquez-Lago, Tatsuya Akutsu, Geoffrey I. Webb, Kuo-Chen Chou, Jiangning Song |
Briefings Bioinform. | 8 |
| 2019 | pLoc_bal-mAnimal: predict subcellular localization of animal proteins by balancing training dataset and PseAACabstractMotivation: A cell contains numerous protein molecules. One of the fundamental goals in cell biology is to determine their subcellular locations, which can provide useful clues about their functions. Knowledge of protein subcellular localization is also indispensable for prioritizing and selecting the right targets for drug development. With the avalanche of protein sequences emerging in the post-genomic age, it is highly desired to develop computational tools for timely and effectively identifying their subcellular localization based on the sequence information alone. Recently, a predictor called 'pLoc-mAnimal' was developed for identifying the subcellular localization of animal proteins. Its performance is overwhelmingly better than that of the other predictors for the same purpose, particularly in dealing with the multi-label systems in which some proteins, called 'multiplex proteins', may simultaneously occur in two or more subcellular locations. Although it is indeed a very powerful predictor, more efforts are definitely needed to further improve it. This is because pLoc-mAnimal was trained by an extremely skewed dataset in which some subset (subcellular location) was about 128 times the size of the other subsets. Accordingly, such an uneven training dataset will inevitably cause a biased consequence. Results: To alleviate such biased consequence, we have developed a new and bias-reducing predictor called pLoc_bal-mAnimal by quasi-balancing the training dataset. Cross-validation tests on exactly the same experiment-confirmed dataset have indicated that the proposed new predictor is remarkably superior to pLoc-mAnimal, the existing state-of-the-art predictor, in identifying the subcellular localization of animal proteins. Availability and implementation: To maximize the convenience for the vast majority of experimental scientists, a user-friendly web-server for the new predictor has been established at http://www.jci-bioinfo.cn/pLoc_bal-mAnimal/, by which users can easily get their desired results without the need to go through the complicated mathematics. Supplementary information: Supplementary data are available at Bioinformatics online. Wei-Zhong Lin, Kuo-Chen Chou |
Bioinform. | 4 |
| 2019 | Bastion3: a two-layer ensemble predictor of type III secreted effectorsabstractMOTIVATION: Type III secreted effectors (T3SEs) can be injected into host cell cytoplasm via type III secretion systems (T3SSs) to modulate interactions between Gram-negative bacterial pathogens and their hosts. Due to their relevance in pathogen-host interactions, significant computational efforts have been put toward identification of T3SEs and these in turn have stimulated new T3SE discoveries. However, as T3SEs with new characteristics are discovered, these existing computational tools reveal important limitations: (i) most of the trained machine learning models are based on the N-terminus (or incorporating also the C-terminus) instead of the proteins' complete sequences, and (ii) the underlying models (trained with classic algorithms) employed only few features, most of which were extracted based on sequence-information alone. To achieve better T3SE prediction, we must identify more powerful, informative features and investigate how to effectively integrate these into a comprehensive model. RESULTS: In this work, we present Bastion3, a two-layer ensemble predictor developed to accurately identify type III secreted effectors from protein sequence data. In contrast with existing methods that employ single models with few features, Bastion3 explores a wide range of features, from various types, trains single models based on these features and finally integrates these models through ensemble learning. We trained the models using a new gradient boosting machine, LightGBM and further boosted the models' performances through a novel genetic algorithm (GA) based two-step parameter optimization strategy. Our benchmark test demonstrates that Bastion3 achieves a much better performance compared to commonly used methods, with an ACC value of 0.959, F-value of 0.958, MCC value of 0.917 and AUC value of 0.956, which comprehensively outperformed all other toolkits by more than 5.6% in ACC value, 5.7% in F-value, 12.4% in MCC value and 5.8% in AUC value. Based on our proposed two-layer ensemble model, we further developed a user-friendly online toolkit, maximizing convenience for experimental scientists toward T3SE prediction. With its design to ease future discoveries of novel T3SEs and improved performance, Bastion3 is poised to become a widely used, state-of-the-art toolkit for T3SE prediction. AVAILABILITY AND IMPLEMENTATION: http://bastion3.erc.monash.edu/. CONTACT: [email protected] or [email protected] or or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jiawei Wang 0002, Jiahui Li 0007, Bingjiao Yang, Ruopeng Xie, Tatiana T. Marquez-Lago, André Leier, Morihiro Hayashida, Tatsuya Akutsu, Yanju Zhang, Kuo-Chen Chou, Joel Selkrig, Tieli Zhou, Jiangning Song, Trevor Lithgow |
Bioinform. | 10 |
| 2019 | MULTiPly: a novel multi-layer predictor for discovering general and specific types of promotersabstractMOTIVATION: Promoters are short DNA consensus sequences that are localized proximal to the transcription start sites of genes, allowing transcription initiation of particular genes. However, the precise prediction of promoters remains a challenging task because individual promoters often differ from the consensus at one or more positions. RESULTS: In this study, we present a new multi-layer computational approach, called MULTiPly, for recognizing promoters and their specific types. MULTiPly took into account the sequences themselves, including both local information such as k-tuple nucleotide composition, dinucleotide-based auto covariance and global information of the entire samples based on bi-profile Bayes and k-nearest neighbour feature encodings. Specifically, the F-score feature selection method was applied to identify the best unique type of feature prediction results, in combination with other types of features that were subsequently added to further improve the prediction performance of MULTiPly. Benchmarking experiments on the benchmark dataset and comparisons with five state-of-the-art tools show that MULTiPly can achieve a better prediction performance on 5-fold cross-validation and jackknife tests. Moreover, the superiority of MULTiPly was also validated on a newly constructed independent test dataset. MULTiPly is expected to be used as a useful tool that will facilitate the discovery of both general and specific types of promoters in the post-genomic era. AVAILABILITY AND IMPLEMENTATION: The MULTiPly webserver and curated datasets are freely available at http://flagshipnt.erc.monash.edu/MULTiPly/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Meng Zhang 0046, Fuyi Li, Tatiana T. Marquez-Lago, André Leier, Cunshuo Fan, Chee Keong Kwoh 0001, Kuo-Chen Chou, Jiangning Song, Cangzhi Jia |
Bioinform. | 7 |
| 2019 | Positive-unlabelled learning of glycosylation sites in the human proteomeabstractBACKGROUND: As an important type of post-translational modification (PTM), protein glycosylation plays a crucial role in protein stability and protein function. The abundance and ubiquity of protein glycosylation across three domains of life involving Eukarya, Bacteria and Archaea demonstrate its roles in regulating a variety of signalling and metabolic pathways. Mutations on and in the proximity of glycosylation sites are highly associated with human diseases. Accordingly, accurate prediction of glycosylation can complement laboratory-based methods and greatly benefit experimental efforts for characterization and understanding of functional roles of glycosylation. For this purpose, a number of supervised-learning approaches have been proposed to identify glycosylation sites, demonstrating a promising predictive performance. To train a conventional supervised-learning model, both reliable positive and negative samples are required. However, in practice, a large portion of negative samples (i.e. non-glycosylation sites) are mislabelled due to the limitation of current experimental technologies. Moreover, supervised algorithms often fail to take advantage of large volumes of unlabelled data, which can aid in model learning in conjunction with positive samples (i.e. experimentally verified glycosylation sites). RESULTS: In this study, we propose a positive unlabelled (PU) learning-based method, PA2DE (V2.0), based on the AlphaMax algorithm for protein glycosylation site prediction. The predictive performance of this proposed method was evaluated by a range of glycosylation data collected over a ten-year period based on an interval of three years. Experiments using both benchmarking and independent tests show that our method outperformed the representative supervised-learning algorithms (including support vector machines and random forests) and one-class learners, as well as currently available prediction methods in terms of F1 score, accuracy and AUC measures. In addition, we developed an online web server as an implementation of the optimized model (available at http://glycomine.erc.monash.edu/Lab/GlycoMine_PU/ ) to facilitate community-wide efforts for accurate prediction of protein glycosylation sites. CONCLUSION: The proposed PU learning approach achieved a competitive predictive performance compared with currently available methods. This PU learning schema may also be effectively employed and applied to address the prediction problems of other important types of protein PTM site and functional sites. Fuyi Li, Yang Zhang 0010, Anthony W. Purcell, Geoffrey I. Webb, Kuo-Chen Chou, Trevor Lithgow, Chen Li 0021, Jiangning Song |
BMC Bioinform. | 5 |
| 2018 | iFeature: a Python package and web server for features extraction and selection from protein and peptide sequencesabstractSummary: Structural and physiochemical descriptors extracted from sequence data have been widely used to represent sequences and predict structural, functional, expression and interaction profiles of proteins and peptides as well as DNAs/RNAs. Here, we present iFeature, a versatile Python-based toolkit for generating various numerical feature representation schemes for both protein and peptide sequences. iFeature is capable of calculating and extracting a comprehensive spectrum of 18 major sequence encoding schemes that encompass 53 different types of feature descriptors. It also allows users to extract specific amino acid properties from the AAindex database. Furthermore, iFeature integrates 12 different types of commonly used feature clustering, selection and dimensionality reduction algorithms, greatly facilitating training, analysis and benchmarking of machine-learning models. The functionality of iFeature is made freely available via an online web server and a stand-alone toolkit. Availability and implementation: http://iFeature.erc.monash.edu/; https://github.com/Superzchen/iFeature/. Supplementary information: Supplementary data are available at Bioinformatics online. Zhen Chen 0009, Fuyi Li, André Leier, Tatiana T. Marquez-Lago, Yanan Wang 0003, Geoffrey I. Webb, Alexander Ian Smith, Roger J. Daly, Kuo-Chen Chou, Jiangning Song |
Bioinform. | 10 |
| 2018 | pLoc-mHum: predict subcellular localization of multi-location human proteins via general PseAAC to winnow out the crucial GO informationabstractMotivation: For in-depth understanding the functions of proteins in a cell, the knowledge of their subcellular localization is indispensable. The current study is focused on human protein subcellular location prediction based on the sequence information alone. Although considerable efforts have been made in this regard, the problem is far from being solved yet. Most existing methods can be used to deal with single-location proteins only. Actually, proteins with multi-locations may have some special biological functions that are particularly important for both basic research and drug design. Results: Using the multi-label theory, we present a new predictor called 'pLoc-mHum' by extracting the crucial GO (Gene Ontology) information into the general PseAAC (Pseudo Amino Acid Composition). Rigorous cross-validations on a same stringent benchmark dataset have indicated that the proposed pLoc-mHum predictor is remarkably superior to iLoc-Hum, the state-of-the-art method in predicting the human protein subcellular localization. Availability and implementation: To maximize the convenience of most experimental scientists, a user-friendly web-server for the new predictor has been established at http://www.jci-bioinfo.cn/pLoc-mHum/, by which users can easily get their desired results without the need to go through the complicated mathematics involved. Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Kuo-Chen Chou |
Bioinform. | 3 |
| 2018 | Quokka: a comprehensive tool for rapid and accurate prediction of kinase family-specific phosphorylation sites in the human proteomeabstractMotivation: Kinase-regulated phosphorylation is a ubiquitous type of post-translational modification (PTM) in both eukaryotic and prokaryotic cells. Phosphorylation plays fundamental roles in many signalling pathways and biological processes, such as protein degradation and protein-protein interactions. Experimental studies have revealed that signalling defects caused by aberrant phosphorylation are highly associated with a variety of human diseases, especially cancers. In light of this, a number of computational methods aiming to accurately predict protein kinase family-specific or kinase-specific phosphorylation sites have been established, thereby facilitating phosphoproteomic data analysis. Results: In this work, we present Quokka, a novel bioinformatics tool that allows users to rapidly and accurately identify human kinase family-regulated phosphorylation sites. Quokka was developed by using a variety of sequence scoring functions combined with an optimized logistic regression algorithm. We evaluated Quokka based on well-prepared up-to-date benchmark and independent test datasets, curated from the Phospho.ELM and UniProt databases, respectively. The independent test demonstrates that Quokka improves the prediction performance compared with state-of-the-art computational tools for phosphorylation prediction. In summary, our tool provides users with high-quality predicted human phosphorylation sites for hypothesis generation and biological validation. Availability and implementation: The Quokka webserver and datasets are freely available at http://quokka.erc.monash.edu/. Supplementary information: Supplementary data are available at Bioinformatics online. Fuyi Li, Chen Li 0021, Tatiana T. Marquez-Lago, André Leier, Tatsuya Akutsu, Anthony W. Purcell, Alexander Ian Smith, Trevor Lithgow, Roger J. Daly, Jiangning Song, Kuo-Chen Chou |
Bioinform. | 11 |
| 2018 | iEnhancer-EL: identifying enhancers and their strength with ensemble learning approachabstractMotivation: Identification of enhancers and their strength is important because they play a critical role in controlling gene expression. Although some bioinformatics tools were developed, they are limited in discriminating enhancers from non-enhancers only. Recently, a two-layer predictor called 'iEnhancer-2L' was developed that can be used to predict the enhancer's strength as well. However, its prediction quality needs further improvement to enhance the practical application value. Results: A new predictor called 'iEnhancer-EL' was proposed that contains two layer predictors: the first one (for identifying enhancers) is formed by fusing an array of six key individual classifiers, and the second one (for their strength) formed by fusing an array of ten key individual classifiers. All these key classifiers were selected from 171 elementary classifiers formed by SVM (Support Vector Machine) based on kmer, subsequence profile and PseKNC (Pseudo K-tuple Nucleotide Composition), respectively. Rigorous cross-validations have indicated that the proposed predictor is remarkably superior to the existing state-of-the-art one in this area. Availability and implementation: A web server for the iEnhancer-EL has been established at http://bioinformatics.hitsz.edu.cn/iEnhancer-EL/, by which users can easily get their desired results without the need to go through the mathematical details. Supplementary information: Supplementary data are available at Bioinformatics online. Bin Liu 0014, De-Shuang Huang, Kuo-Chen Chou |
Bioinform. | 4 |
| 2018 | iRO-3wPseKNC: identify DNA replication origins by three-window-based PseKNCabstractMotivation: DNA replication is the key of the genetic information transmission, and it is initiated from the replication origins. Identifying the replication origins is crucial for understanding the mechanism of DNA replication. Although several discriminative computational predictors were proposed to identify DNA replication origins of yeast species, they could only be used to identify very tiny parts (250 or 300 bp) of the replication origins. Besides, none of the existing predictors could successfully capture the 'GC asymmetry bias' of yeast species reported by experimental observations. Hence it would not be surprising why their power is so limited. To grasp the CG asymmetry feature and make the prediction able to cover the entire replication regions of yeast species, we develop a new predictor called 'iRO-3wPseKNC'. Results: Rigorous cross validations on the benchmark datasets from four yeast species (Saccharomyces cerevisiae, Schizosaccharomyces pombe, Kluyveromyces lactis and Pichia pastoris) have indicated that the proposed predictor is really very powerful for predicting the entire DNA duplication origins. Availability and implementation: The web-server for the iRO-3wPseKNC predictor is available at http://bioinformatics.hitsz.edu.cn/iRO-3wPseKNC/, by which users can easily get their desired results without the need to go through the mathematical details. Supplementary information: Supplementary data are available at Bioinformatics online. Bin Liu 0014, Fan Weng, De-Shuang Huang, Kuo-Chen Chou |
Bioinform. | 4 |
| 2018 | iPromoter-2L: a two-layer predictor for identifying promoters and their types by multi-window-based PseKNCabstractMotivation: Being responsible for initiating transaction of a particular gene in genome, promoter is a short region of DNA. Promoters have various types with different functions. Owing to their importance in biological process, it is highly desired to develop computational tools for timely identifying promoters and their types. Such a challenge has become particularly critical and urgent in facing the avalanche of DNA sequences discovered in the postgenomic age. Although some prediction methods were developed, they can only be used to discriminate a specific type of promoters from non-promoters. None of them has the ability to identify the types of promoters. This is due to the facts that different types of promoters may share quite similar consensus sequence pattern, and that the promoters of same type may have considerably different consensus sequences. Results: To overcome such difficulty, using the multi-window-based PseKNC (pseudo K-tuple nucleotide composition) approach to incorporate the short-, middle-, and long-range sequence information, we have developed a two-layer seamless predictor named as 'iPromoter-2 L'. The first layer serves to identify a query DNA sequence as a promoter or non-promoter, and the second layer to predict which of the following six types the identified promoter belongs to: σ24, σ28, σ32, σ38, σ54 and σ70. Availability and implementation: For the convenience of most experimental scientists, a user-friendly and publicly accessible web-server for the powerful new predictor has been established at http://bioinformatics.hitsz.edu.cn/iPromoter-2L/. It is anticipated that iPromoter-2 L will become a very useful high throughput tool for genome analysis. Contact: [email protected] or [email protected] or [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Bin Liu 0014, De-Shuang Huang, Kuo-Chen Chou |
Bioinform. | 4 |
| 2018 | PROSPERous: high-throughput prediction of substrate cleavage sites for 90 proteases with improved accuracyabstractSummary: Proteases are enzymes that specifically cleave the peptide backbone of their target proteins. As an important type of irreversible post-translational modification, protein cleavage underlies many key physiological processes. When dysregulated, proteases' actions are associated with numerous diseases. Many proteases are highly specific, cleaving only those target substrates that present certain particular amino acid sequence patterns. Therefore, tools that successfully identify potential target substrates for proteases may also identify previously unknown, physiologically relevant cleavage sites, thus providing insights into biological processes and guiding hypothesis-driven experiments aimed at verifying protease-substrate interaction. In this work, we present PROSPERous, a tool for rapid in silico prediction of protease-specific cleavage sites in substrate sequences. Our tool is based on logistic regression models and uses different scoring functions and their pairwise combinations to subsequently predict potential cleavage sites. PROSPERous represents a state-of-the-art tool that enables fast, accurate and high-throughput prediction of substrate cleavage sites for 90 proteases. Availability and implementation: http://prosperous.erc.monash.edu/. Contact: [email protected] or [email protected] or [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Jiangning Song, Fuyi Li, André Leier, Tatiana T. Marquez-Lago, Tatsuya Akutsu, Gholamreza Haffari, Kuo-Chen Chou, Geoffrey I. Webb, Robert N. Pike |
Bioinform. | 7 |
| 2018 | iLoc-lncRNA: predict the subcellular location of lncRNAs by incorporating octamer composition into general PseKNCabstractMotivation: Long non-coding RNAs (lncRNAs) are a class of RNA molecules with more than 200 nucleotides. They have important functions in cell development and metabolism, such as genetic markers, genome rearrangements, chromatin modifications, cell cycle regulation, transcription and translation. Their functions are generally closely related to their localization in the cell. Therefore, knowledge about their subcellular locations can provide very useful clues or preliminary insight into their biological functions. Although biochemical experiments could determine the localization of lncRNAs in a cell, they are both time-consuming and expensive. Therefore, it is highly desirable to develop bioinformatics tools for fast and effective identification of their subcellular locations. Results: We developed a sequence-based bioinformatics tool called 'iLoc-lncRNA' to predict the subcellular locations of LncRNAs by incorporating the 8-tuple nucleotide features into the general PseKNC (Pseudo K-tuple Nucleotide Composition) via the binomial distribution approach. Rigorous jackknife tests have shown that the overall accuracy achieved by the new predictor on a stringent benchmark dataset is 86.72%, which is over 20% higher than that by the existing state-of-the-art predictor evaluated on the same tests. Availability and implementation: A user-friendly webserver has been established at http://lin-group.cn/server/iLoc-LncRNA, by which users can easily obtain their desired results. Supplementary information: Supplementary data are available at Bioinformatics online. Zhen-Dong Su 0001, Zhao-Yue Zhang 0002, Ya-Wei Zhao, Dong Wang 0011, Wei Chen 0064, Kuo-Chen Chou, Hao Lin 0001 |
Bioinform. | 7 |
| 2018 | Bastion6: a bioinformatics approach for accurate prediction of type VI secreted effectorsabstractMotivation: Many Gram-negative bacteria use type VI secretion systems (T6SS) to export effector proteins into adjacent target cells. These secreted effectors (T6SEs) play vital roles in the competitive survival in bacterial populations, as well as pathogenesis of bacteria. Although various computational analyses have been previously applied to identify effectors secreted by certain bacterial species, there is no universal method available to accurately predict T6SS effector proteins from the growing tide of bacterial genome sequence data. Results: We extracted a wide range of features from T6SE protein sequences and comprehensively analyzed the prediction performance of these features through unsupervised and supervised learning. By integrating these features, we subsequently developed a two-layer SVM-based ensemble model with fine-grain optimized parameters, to identify potential T6SEs. We further validated the predictive model using an independent dataset, which showed that the proposed model achieved an impressive performance in terms of ACC (0.943), F-value (0.946), MCC (0.892) and AUC (0.976). To demonstrate applicability, we employed this method to correctly identify two very recently validated T6SE proteins, which represent challenging prediction targets because they significantly differed from previously known T6SEs in terms of their sequence similarity and cellular function. Furthermore, a genome-wide prediction across 12 bacterial species, involving in total 54 212 protein sequences, was carried out to distinguish 94 putative T6SE candidates. We envisage both this information and our publicly accessible web server will facilitate future discoveries of novel T6SEs. Availability and implementation: http://bastion6.erc.monash.edu/. Supplementary information: Supplementary data are available at Bioinformatics online. Jiawei Wang 0002, Bingjiao Yang, André Leier, Tatiana T. Marquez-Lago, Morihiro Hayashida, Andrea Rocker, Yanju Zhang, Tatsuya Akutsu, Kuo-Chen Chou, Richard A. Strugnell, Jiangning Song, Trevor Lithgow |
Bioinform. | 9 |
| 2017 | pLoc-mAnimal: predict subcellular localization of animal proteins with both single and multiple sitesabstractMOTIVATION: Cells are deemed the basic unit of life. However, many important functions of cells as well as their growth and reproduction are performed via the protein molecules located at their different organelles or locations. Facing explosive growth of protein sequences, we are challenged to develop fast and effective method to annotate their subcellular localization. However, this is by no means an easy task. Particularly, mounting evidences have indicated proteins have multi-label feature meaning that they may simultaneously exist at, or move between, two or more different subcellular location sites. Unfortunately, most of the existing computational methods can only be used to deal with the single-label proteins. Although the 'iLoc-Animal' predictor developed recently is quite powerful that can be used to deal with the animal proteins with multiple locations as well, its prediction quality needs to be improved, particularly in enhancing the absolute true rate and reducing the absolute false rate. RESULTS: Here we propose a new predictor called 'pLoc-mAnimal', which is superior to iLoc-Animal as shown by the compelling facts. When tested by the most rigorous cross-validation on the same high-quality benchmark dataset, the absolute true success rate achieved by the new predictor is 37% higher and the absolute false rate is four times lower in comparison with the state-of-the-art predictor. AVAILABILITY AND IMPLEMENTATION: To maximize the convenience of most experimental scientists, a user-friendly web-server for the new predictor has been established at http://www.jci-bioinfo.cn/pLoc-mAnimal/, by which users can easily get their desired results without the need to go through the complicated mathematics involved. CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Shu-Guang Zhao, Wei-Zhong Lin, Kuo-Chen Chou |
Bioinform. | 5 |
| 2017 | iATC-mISF: a multi-label classifier for predicting the classes of anatomical therapeutic chemicalsabstractMotivation: Given a compound, can we predict which anatomical therapeutic chemical (ATC) class/classes it belongs to? It is a challenging problem since the information thus obtained can be used to deduce its possible active ingredients, as well as its therapeutic, pharmacological and chemical properties. And hence the pace of drug development could be substantially expedited. But this problem is by no means an easy one. Particularly, some drugs or compounds may belong to two or more ATC classes. Results: To address it, a multi-label classifier, called iATC-mISF, was developed by incorporating the information of chemical–chemical interaction, the information of the structural similarity, and the information of the fingerprintal similarity. Rigorous cross-validations showed that the proposed predictor achieved remarkably higher prediction quality than its cohorts for the same purpose, particularly in the absolute true rate, the most important and harsh metrics for the multi-label systems. Availability and Implementation: The web-server for iATC-mISF is accessible at http://www.jci-bioinfo.cn/iATC-mISF. Furthermore, to maximize the convenience for most experimental scientists, a step-by-step guide was provided, by which users can easily get their desired results without needing to go through the complicated mathematical equations. Their inclusion in this article is just for the integrity of the new method and stimulating more powerful methods to deal with various multi-label systems in biology. Contact: [email protected] Supplementary Information: Supplementary data are available at Bioinformatics online. Shu-Guang Zhao, Kuo-Chen Chou |
Bioinform. | 4 |
| 2017 | iATC-mISF: a multi-label classifier for predicting the classes of anatomical therapeutic chemicalsabstractBioinformatics (2017) 33(3), 341–346 The horizontal axis of Figure 1 was incorrectly marked as ‘Parameter θ’, which should be changed to ‘Parameter 1/θ’. Accordingly, the sentence ‘As shown in Figure 1, when θ = 16, the absolute true rate reached its highest score’ should be changed to ‘As shown in Figure 1, when θ = 1/16, the absolute true rate reached its highest score’. The article has now been corrected online. Shu-Guang Zhao, Kuo-Chen Chou |
Bioinform. | 4 |
| 2017 | iRSpot-EL: identify recombination spots with an ensemble learning approachabstractMOTIVATION: Coexisting in a DNA system, meiosis and recombination are two indispensible aspects for cell reproduction and growth. With the avalanche of genome sequences emerging in the post-genomic age, it is an urgent challenge to acquire the information of DNA recombination spots because it can timely provide very useful insights into the mechanism of meiotic recombination and the process of genome evolution. RESULTS: To address such a challenge, we have developed a predictor, called IRSPOT-EL: , by fusing different modes of pseudo K-tuple nucleotide composition and mode of dinucleotide-based auto-cross covariance into an ensemble classifier of clustering approach. Five-fold cross tests on a widely used benchmark dataset have indicated that the new predictor remarkably outperforms its existing counterparts. Particularly, far beyond their reach, the new predictor can be easily used to conduct the genome-wide analysis and the results obtained are quite consistent with the experimental map. AVAILABILITY AND IMPLEMENTATION: For the convenience of most experimental scientists, a user-friendly web-server for iRSpot-EL has been established at http://bioinformatics.hitsz.edu.cn/iRSpot-EL/, by which users can easily obtain their desired results without the need to go through the complicated mathematical equations involved. CONTACT: [email protected] or [email protected] information: Supplementary data are available at Bioinformatics online. Bin Liu 0014, Shanyi Wang, Ren Long, Kuo-Chen Chou |
Bioinform. | 4 |
| 2017 | POSSUM: a bioinformatics toolkit for generating numerical sequence feature descriptors based on PSSM profilesabstractSUMMARY: Evolutionary information in the form of a Position-Specific Scoring Matrix (PSSM) is a widely used and highly informative representation of protein sequences. Accordingly, PSSM-based feature descriptors have been successfully applied to improve the performance of various predictors of protein attributes. Even though a number of algorithms have been proposed in previous studies, there is currently no universal web server or toolkit available for generating this wide variety of descriptors. Here, we present POSSUM ( Po sition- S pecific S coring matrix-based feat u re generator for m achine learning), a versatile toolkit with an online web server that can generate 21 types of PSSM-based feature descriptors, thereby addressing a crucial need for bioinformaticians and computational biologists. We envisage that this comprehensive toolkit will be widely used as a powerful tool to facilitate feature extraction, selection, and benchmarking of machine learning-based models, thereby contributing to a more effective analysis and modeling pipeline for bioinformatics research. AVAILABILITY AND IMPLEMENTATION: http://possum.erc.monash.edu/ . CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jiawei Wang 0002, Bingjiao Yang, Jerico Revote, André Leier, Tatiana T. Marquez-Lago, Geoffrey I. Webb, Jiangning Song, Kuo-Chen Chou, Trevor Lithgow |
Bioinform. | 8 |
| 2016 | iEnhancer-2L: a two-layer predictor for identifying enhancers and their strength by pseudo k-tuple nucleotide compositionabstractMOTIVATION: Enhancers are of short regulatory DNA elements. They can be bound with proteins (activators) to activate transcription of a gene, and hence play a critical role in promoting gene transcription in eukaryotes. With the avalanche of DNA sequences generated in the post-genomic age, it is a challenging task to develop computational methods for timely identifying enhancers from extremely complicated DNA sequences. Although some efforts have been made in this regard, they were limited at only identifying whether a query DNA element being of an enhancer or not. According to the distinct levels of biological activities and regulatory effects on target genes, however, enhancers should be further classified into strong and weak ones in strength. RESULTS: In view of this, a two-layer predictor called ' IENHANCER-2L: ' was proposed by formulating DNA elements with the 'pseudo k-tuple nucleotide composition', into which the six DNA local parameters were incorporated. To the best of our knowledge, it is the first computational predictor ever established for identifying not only enhancers, but also their strength. Rigorous cross-validation tests have indicated that IENHANCER-2L: holds very high potential to become a useful tool for genome analysis. AVAILABILITY AND IMPLEMENTATION: For the convenience of most experimental scientists, a web server for the two-layer predictor was established at http://bioinformatics.hitsz.edu.cn/iEnhancer-2L/, by which users can easily get their desired results without the need to go through the mathematical details. CONTACT: [email protected], [email protected], [email protected], [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Bin Liu 0014, Longyun Fang, Ren Long, Xun Lan, Kuo-Chen Chou |
Bioinform. | 5 |
| 2016 | pSumo-CD: predicting sumoylation sites in proteins with covariance discriminant algorithm by incorporating sequence-coupled effects into general PseAACabstractMOTIVATION: Sumoylation is a post-translational modification (PTM) process, in which small ubiquitin-related modifier (SUMO) is attaching by covalent bonds to substrate protein. It is critical to many different biological processes such as replicating genome, expressing gene, localizing and stabilizing proteins; unfortunately, it is also involved with many major disorders including Alzheimer's and Parkinson's diseases. Therefore, for both basic research and drug development, it is important to identify the sumoylation sites in proteins. RESULTS: To address such a problem, we developed a predictor called pSumo-CD by incorporating the sequence-coupled information into the general pseudo-amino acid composition (PseAAC) and introducing the covariance discriminant (CD) algorithm, in which a bias-adjustment term, which has the function to automatically adjust the errors caused by the bias due to the imbalance of training data, had been incorporated. Rigorous cross-validations indicated that the new predictor remarkably outperformed the existing state-of-the-art prediction method for the same purpose. AVAILABILITY AND IMPLEMENTATION: For the convenience of most experimental scientists, a user-friendly web-server for pSumo-CD has been established at http://www.jci-bioinfo.cn/pSumo-CD, by which users can easily obtain their desired results without the need to go through the complicated mathematical equations involved. CONTACT: [email protected], [email protected] or [email protected] information: Supplementary data are available at Bioinformatics online. Jianhua Jia, Liuxia Zhang, Zi Liu, Kuo-Chen Chou |
Bioinform. | 5 |
| 2016 | iDHS-EL: identifying DNase I hypersensitive sites by fusing three different modes of pseudo nucleotide composition into an ensemble learning frameworkabstractMOTIVATION: Regulatory DNA elements are associated with DNase I hypersensitive sites (DHSs). Accordingly, identification of DHSs will provide useful insights for in-depth investigation into the function of noncoding genomic regions. RESULTS: In this study, using the strategy of ensemble learning framework, we proposed a new predictor called iDHS-EL for identifying the location of DHS in human genome. It was formed by fusing three individual Random Forest (RF) classifiers into an ensemble predictor. The three RF operators were respectively based on the three special modes of the general pseudo nucleotide composition (PseKNC): (i) kmer, (ii) reverse complement kmer and (iii) pseudo dinucleotide composition. It has been demonstrated that the new predictor remarkably outperforms the relevant state-of-the-art methods in both accuracy and stability. AVAILABILITY AND IMPLEMENTATION: For the convenience of most experimental scientists, a web server for iDHS-EL is established at http://bioinformatics.hitsz.edu.cn/iDHS-EL, which is the first web-server predictor ever established for identifying DHSs, and by which users can easily get their desired results without the need to go through the mathematical details. We anticipate that IDHS-EL: will become a very useful high throughput tool for genome analysis. CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Bin Liu 0014, Ren Long, Kuo-Chen Chou |
Bioinform. | 3 |
| 2016 | iPTM-mLys: identifying multiple lysine PTM sites and their different typesabstractMOTIVATION: Post-translational modification, abbreviated as PTM, refers to the change of the amino acid side chains of a protein after its biosynthesis. Owing to its significance for in-depth understanding various biological processes and developing effective drugs, prediction of PTM sites in proteins have currently become a hot topic in bioinformatics. Although many computational methods were established to identify various single-label PTM types and their occurrence sites in proteins, no method has ever been developed for multi-label PTM types. As one of the most frequently observed PTMs, the K-PTM, namely, the modification occurring at lysine (K), can be usually accommodated with many different types, such as 'acetylation', 'crotonylation', 'methylation' and 'succinylation'. Now we are facing an interesting challenge: given an uncharacterized protein sequence containing many K residues, which ones can accommodate two or more types of PTM, which ones only one, and which ones none? RESULTS: To address this problem, a multi-label predictor called IPTM-MLYS: has been developed. It represents the first multi-label PTM predictor ever established. The novel predictor is featured by incorporating the sequence-coupled effects into the general PseAAC, and by fusing an array of basic random forest classifiers into an ensemble system. Rigorous cross-validations via a set of multi-label metrics indicate that the first multi-label PTM predictor is very promising and encouraging. AVAILABILITY AND IMPLEMENTATION: For the convenience of most experimental scientists, a user-friendly web-server for iPTM-mLys has been established at http://www.jci-bioinfo.cn/iPTM-mLys, by which users can easily obtain their desired results without the need to go through the complicated mathematical equations involved. CONTACT: [email protected], [email protected], [email protected] information: Supplementary data are available at Bioinformatics online. Wangren Qiu, Bi-Qian Sun, Kuo-Chen Chou |
Bioinform. | 5 |
| 2015 | PseKNC-General: a cross-platform package for generating various modes of pseudo nucleotide compositionsabstractSUMMARY: The avalanche of genomic sequences generated in the post-genomic age requires efficient computational methods for rapidly and accurately identifying biological features from sequence information. Towards this goal, we developed a freely available and open-source package, called PseKNC-General (the general form of pseudo k-tuple nucleotide composition), that allows for fast and accurate computation of all the widely used nucleotide structural and physicochemical properties of both DNA and RNA sequences. PseKNC-General can generate several modes of pseudo nucleotide compositions, including conventional k-tuple nucleotide compositions, Moreau-Broto autocorrelation coefficient, Moran autocorrelation coefficient, Geary autocorrelation coefficient, Type I PseKNC and Type II PseKNC. In every mode, >100 physicochemical properties are available for choosing. Moreover, it is flexible enough to allow the users to calculate PseKNC with user-defined properties. The package can be run on Linux, Mac and Windows systems and also provides a graphical user interface. AVAILABILITY AND IMPLEMENTATION: The package is freely available at: http://lin.uestc.edu.cn/server/pseknc. Wei Chen 0064, Xitong Zhang, Jordan Brooker, Hao Lin 0001, Liqing Zhang 0002, Kuo-Chen Chou |
Bioinform. | 6 |
| 2015 | repDNA: a Python package to generate various modes of feature vectors for DNA sequences by incorporating user-defined physicochemical properties and sequence-order effectsabstractUNLABELLED: In order to develop powerful computational predictors for identifying the biological features or attributes of DNAs, one of the most challenging problems is to find a suitable approach to effectively represent the DNA sequences. To facilitate the studies of DNAs and nucleotides, we developed a Python package called representations of DNAs (repDNA) for generating the widely used features reflecting the physicochemical properties and sequence-order effects of DNAs and nucleotides. There are three feature groups composed of 15 features. The first group calculates three nucleic acid composition features describing the local sequence information by means of kmers; the second group calculates six autocorrelation features describing the level of correlation between two oligonucleotides along a DNA sequence in terms of their specific physicochemical properties; the third group calculates six pseudo nucleotide composition features, which can be used to represent a DNA sequence with a discrete model or vector yet still keep considerable sequence-order information via the physicochemical properties of its constituent oligonucleotides. In addition, these features can be easily calculated based on both the built-in and user-defined properties via using repDNA. AVAILABILITY AND IMPLEMENTATION: The repDNA Python package is freely accessible to the public at http://bioinformatics.hitsz.edu.cn/repDNA/. CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Bin Liu 0014, Fule Liu, Longyun Fang, Xiaolong Wang 0001, Kuo-Chen Chou |
Bioinform. | 5 |
| 2014 | iNuc-PseKNC: a sequence-based predictor for predicting nucleosome positioning in genomes with pseudo k-tuple nucleotide compositionabstractMOTIVATION: Nucleosome positioning participates in many cellular activities and plays significant roles in regulating cellular processes. With the avalanche of genome sequences generated in the post-genomic age, it is highly desired to develop automated methods for rapidly and effectively identifying nucleosome positioning. Although some computational methods were proposed, most of them were species specific and neglected the intrinsic local structural properties that might play important roles in determining the nucleosome positioning on a DNA sequence. RESULTS: Here a predictor called 'iNuc-PseKNC' was developed for predicting nucleosome positioning in Homo sapiens, Caenorhabditis elegans and Drosophila melanogaster genomes, respectively. In the new predictor, the samples of DNA sequences were formulated by a novel feature-vector called 'pseudo k-tuple nucleotide composition', into which six DNA local structural properties were incorporated. It was observed by the rigorous cross-validation tests on the three stringent benchmark datasets that the overall success rates achieved by iNuc-PseKNC in predicting the nucleosome positioning of the aforementioned three genomes were 86.27%, 86.90% and 79.97%, respectively. Meanwhile, the results obtained by iNuc-PseKNC on various benchmark datasets used by the previous investigators for different genomes also indicated that the current predictor remarkably outperformed its counterparts. AVAILABILITY: A user-friendly web-server, iNuc-PseKNC is freely accessible at http://lin.uestc.edu.cn/server/iNuc-PseKNC. Shou-Hui Guo, En-Ze Deng, Liqin Xu, Hui Ding 0005, Hao Lin 0001, Wei Chen 0064, Kuo-Chen Chou |
Bioinform. | 7 |
| 2014 | Combining evolutionary information extracted from frequency profiles with sequence-based kernels for protein remote homology detectionabstractMOTIVATION: Owing to its importance in both basic research (such as molecular evolution and protein attribute prediction) and practical application (such as timely modeling the 3D structures of proteins targeted for drug development), protein remote homology detection has attracted a great deal of interest. It is intriguing to note that the profile-based approach is promising and holds high potential in this regard. To further improve protein remote homology detection, a key step is how to find an optimal means to extract the evolutionary information into the profiles. RESULTS: Here, we propose a novel approach, the so-called profile-based protein representation, to extract the evolutionary information via the frequency profiles. The latter can be calculated from the multiple sequence alignments generated by PSI-BLAST. Three top performing sequence-based kernels (SVM-Ngram, SVM-pairwise and SVM-LA) were combined with the profile-based protein representation. Various tests were conducted on a SCOP benchmark dataset that contains 54 families and 23 superfamilies. The results showed that the new approach is promising, and can obviously improve the performance of the three kernels. Furthermore, our approach can also provide useful insights for studying the features of proteins in various families. It has not escaped our notice that the current approach can be easily combined with the existing sequence-based methods so as to improve their performance as well. AVAILABILITY AND IMPLEMENTATION: For users' convenience, the source code of generating the profile-based proteins and the multiple kernel learning was also provided at http://bioinformatics.hitsz.edu.cn/main/~binliu/remote/ Bin Liu 0014, Deyuan Zhang, Ruifeng Xu 0001, Jinghao Xu, Xiaolong Wang 0001, Qingcai Chen, Qiwen Dong, Kuo-Chen Chou |
Bioinform. | 8 |
| 2010 | Predicting the network of substrate-enzyme-product triads by combining compound similarity and functional domain compositionabstractBACKGROUND: Metabolic pathway is a highly regulated network consisting of many metabolic reactions involving substrates, enzymes, and products, where substrates can be transformed into products with particular catalytic enzymes. Since experimental determination of the network of substrate-enzyme-product triad (whether the substrate can be transformed into the product with a given enzyme) is both time-consuming and expensive, it would be very useful to develop a computational approach for predicting the network of substrate-enzyme-product triads. RESULTS: A mathematical model for predicting the network of substrate-enzyme-product triads was developed. Meanwhile, a benchmark dataset was constructed that contains 744,192 substrate-enzyme-product triads, of which 14,592 are networking triads, and 729,600 are non-networking triads; i.e., the number of the negative triads was about 50 times the number of the positive triads. The molecular graph was introduced to calculate the similarity between the substrate compounds and between the product compounds, while the functional domain composition was introduced to calculate the similarity between enzyme molecules. The nearest neighbour algorithm was utilized as a prediction engine, in which a novel metric was introduced to measure the "nearness" between triads. To train and test the prediction engine, one tenth of the positive triads and one tenth of the negative triads were randomly picked from the benchmark dataset as the testing samples, while the remaining were used to train the prediction model. It was observed that the overall success rate in predicting the network for the testing samples was 98.71%, with 95.41% success rate for the 1,460 testing networking triads and 98.77% for the 72,960 testing non-networking triads. CONCLUSIONS: It is quite promising and encouraged to use the molecular graph to calculate the similarity between compounds and use the functional domain composition to calculate the similarity between enzymes for studying the substrate-enzyme-product network system. The software is available upon request. Lei Chen 0007, Kai-Yan Feng, Yu-Dong Cai 0001, Kuo-Chen Chou |
BMC Bioinform. | 4 |
| 2006 | Ensemble classifier for protein fold pattern recognitionabstractMOTIVATION: Prediction of protein folding patterns is one level deeper than that of protein structural classes, and hence is much more complicated and difficult. To deal with such a challenging problem, the ensemble classifier was introduced. It was formed by a set of basic classifiers, with each trained in different parameter systems, such as predicted secondary structure, hydrophobicity, van der Waals volume, polarity, polarizability, as well as different dimensions of pseudo-amino acid composition, which were extracted from a training dataset. The operation engine for the constituent individual classifiers was OET-KNN (optimized evidence-theoretic k-nearest neighbors) rule. Their outcomes were combined through a weighted voting to give a final determination for classifying a query protein. The recognition was to find the true fold among the 27 possible patterns. RESULTS: The overall success rate thus obtained was 62% for a testing dataset where most of the proteins have <25% sequence identity with the proteins used in training the classifier. Such a rate is 6-21% higher than the corresponding rates obtained by various existing NN (neural networks) and SVM (support vector machines) approaches, implying that the ensemble classifier is very promising and might become a useful vehicle in protein science, as well as proteomics and bioinformatics. AVAILABILITY: The ensemble classifier, called PFP-Pred, is available as a web-server at http://202.120.37.186/bioinf/fold/PFP-Pred.htm for public usage. Hong-Bin Shen, Kuo-Chen Chou |
Bioinform. | 2 |
| 2005 | Using amphiphilic pseudo amino acid composition to predict enzyme subfamily classesabstractMOTIVATION: With protein sequences entering into databanks at an explosive pace, the early determination of the family or subfamily class for a newly found enzyme molecule becomes important because this is directly related to the detailed information about which specific target it acts on, as well as to its catalytic process and biological function. Unfortunately, it is both time-consuming and costly to do so by experiments alone. In a previous study, the covariant-discriminant algorithm was introduced to identify the 16 subfamily classes of oxidoreductases. Although the results were quite encouraging, the entire prediction process was based on the amino acid composition alone without including any sequence-order information. Therefore, it is worthy of further investigation. RESULTS: To incorporate the sequence-order effects into the predictor, the 'amphiphilic pseudo amino acid composition' is introduced to represent the statistical sample of a protein. The novel representation contains 20 + 2lambda discrete numbers: the first 20 numbers are the components of the conventional amino acid composition; the next 2lambda numbers are a set of correlation factors that reflect different hydrophobicity and hydrophilicity distribution patterns along a protein chain. Based on such a concept and formulation scheme, a new predictor is developed. It is shown by the self-consistency test, jackknife test and independent dataset tests that the success rates obtained by the new predictor are all significantly higher than those by the previous predictors. The significant enhancement in success rates also implies that the distribution of hydrophobicity and hydrophilicity of the amino acid residues along a protein chain plays a very important role to its structure and function. Kuo-Chen Chou |
Bioinform. | 1 |
| 2005 | Predicting protein localization in budding YeastabstractMOTIVATION: Most of the existing methods in predicting protein subcellular location were used to deal with the cases limited within the scope from two to five localizations, and only a few of them can be effectively extended to cover the cases of 12-14 localizations. This is because the more the locations involved are, the poorer the success rate would be. Besides, some proteins may occur in several different subcellular locations, i.e. bear the feature of 'multiplex locations'. So far there is no method that can be used to effectively treat the difficult multiplex location problem. The present study was initiated in an attempt to address (1) how to efficiently identify the localization of a query protein among many possible subcellular locations, and (2) how to deal with the case of multiplex locations. RESULTS: By hybridizing gene ontology, functional domain and pseudo amino acid composition approaches, a new method has been developed that can be used to predict subcellular localization of proteins with multiplex location feature. A global analysis of the proteins in budding yeast classified into 22 locations was performed by jack-knife cross-validation with the new method. The overall success identification rate thus obtained is 70%. In contrast to this, the corresponding rates obtained by some other existing methods were only 13-14%, indicating that the new method is very powerful and promising. Furthermore, predictions were made for the four proteins whose localizations could not be determined by experiments, as well as for the 236 proteins whose localizations in budding yeast were ambiguous according to experimental observations. However, according to our predicted results, many of these 'ambiguous proteins' were found to have the same score and ranking for several different subcellular locations, implying that they may simultaneously exist, or move around, in these locations. This finding is intriguing because it reflects the dynamic feature of these proteins in a cell that may be associated with some special biological functions. Kuo-Chen Chou, Yu-Dong Cai 0001 |
Bioinform. | 1 |
| 2004 | Predicting subcellular localization of proteins in a hybridization spaceabstractMOTIVATION: The localization of a protein in a cell is closely correlated with its biological function. With the number of sequences entering into databanks rapidly increasing, the importance of developing a powerful high-throughput tool to determine protein subcellular location has become self-evident. In view of this, the Nearest Neighbour Algorithm was developed for predicting the protein subcellular location using the strategy of hybridizing the information derived from the recent development in gene ontology with that from the functional domain composition as well as the pseudo amino acid composition. RESULTS: As a showcase, the same plant and non-plant protein datasets as investigated by the previous investigators were used for demonstration. The overall success rate of the jackknife test for the plant protein dataset was 86%, and that for the non-plant protein dataset 91.2%. These are the highest success rates achieved so far for the two datasets by following a rigorous cross-validation test procedure, suggesting that such a hybrid approach (particularly by incorporating the knowledge of gene ontology) may become a very useful high-throughput tool in the area of bioinformatics, proteomics, as well as molecular cell biology. AVAILABILITY: The software would be made available on sending a request to the authors. Yu-Dong Cai 0001, Kuo-Chen Chou |
Bioinform. | 2 |
| 2004 | Bio-support vector machines for computational proteomicsabstractMOTIVATION: One of the most important issues in computational proteomics is to produce a prediction model for the classification or annotation of biological function of novel protein sequences. In order to improve the prediction accuracy, much attention has been paid to the improvement of the performance of the algorithms used, few is for solving the fundamental issue, namely, amino acid encoding as most existing pattern recognition algorithms are unable to recognize amino acids in protein sequences. Importantly, the most commonly used amino acid encoding method has the flaw that leads to large computational cost and recognition bias. RESULTS: By replacing kernel functions of support vector machines (SVMs) with amino acid similarity measurement matrices, we have modified SVMs, a new type of pattern recognition algorithm for analysing protein sequences, particularly for proteolytic cleavage site prediction. We refer to the modified SVMs as bio-support vector machine. When applied to the prediction of HIV protease cleavage sites, the new method has shown a remarkable advantage in reducing the model complexity and enhancing the model robustness. Zheng Rong Yang, Kuo-Chen Chou |
Bioinform. | 2 |
| 2004 | Predicting the linkage sites in glycoproteins using bio-basis function neural networkabstractMOTIVATION: Although, it is known that O-glycosidically linked oligosaccharides are commonly conjugated to a serine, threonine or hydroxylysine residue of the polypeptide, the chemical nature of the anchoring monosaccharide and the size of the oligosaccharide unit varies. Among different types, O-linked or mucin-type oligosaccharides are intimately involved in the secretion of proteins, be they enzymes, hormones or structural glycoproteins. Knowledge of the linkage sites in glycoproteins is critical to the design of specific and efficient inhibitors against the enzyme to catalyse the formation of the carbohydrate-peptide linkage. RESULTS: We present a method for predicting the linkage sites in O-linked glycoproteins using bio-basis function neural networks. The mean prediction accuracy of this method is 91.15 +/- 2.75% while it is 82.28 +/- 6.45% using back-propagation neural networks. Importantly, this method has significantly reduced the CPU time for modelling. Zheng Rong Yang, Kuo-Chen Chou |
Bioinform. | 2 |