Tzong-Yi Lee

dblp:01/1288 · DBLP profile ↗
← Back
46ranked-venue papers
4as first author
23since 2021 · last 2025
0000-0001-8475-7868ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 46 · 4 first-author · 23 since 2021
YearPublicationVenuePosition
2025 Self-Attention Enhanced Deep Learning Models for Immune Cell Deconvolution from Bulk RNA-Seq
abstract
Accurate immune cell composition profiling is crucial for understanding immunological dynamics and disease mechanisms. Bulk RNA sequencing (bulk RNA-seq) is widely employed due to its cost-effectiveness and scalability; however, it lacks the resolution to identify cell-specific gene expression. To address this limitation, we propose a self-attention enhanced deep learning model designed for precise immune cell deconvolution from bulk RNA-seq data. We systematically annotated immune cell types from four single-cell RNA-seq (scRNA-seq) peripheral blood mononuclear cell (PBMC) datasets and validated these annotations against established automated identification tools (SingleR, Seurat, scPred, ScType). Leveraging these annotations, we generated realistic pseudo-bulk RNA-seq training samples using Dirichlet-distribution-based composition sampling, significantly enhancing the model’s performance, particularly for rare cell populations. Comparative evaluations demonstrated that our self-attention enhanced deep learning model consistently outperformed existing approaches, including CIBERSORTx and Scaden, achieving lower prediction errors and higher correlations on benchmark PBMC datasets. Integrating multi-head self-attention allowed the model to dynamically capture intricate dependencies among gene expression features, substantially improving deconvolution accuracy for specific cell subsets. While demonstrating robust performance on PBMC datasets, we acknowledge that broader validation is essential due to potential limitations in generalizability across different tissue types and conditions. Our study highlights the potential of self-attention mechanisms and realistic training data generation strategies to enhance computational deconvolution techniques, providing valuable tools for clinical diagnostics and translational immunology research.
Chia-Ru Chung, Yen-Lin Chen, Justin Bo-Kai Hsu, Li-Ching Wu, Tzong-Yi Lee, Jorng-Tzong Horng
CIBCB6
2025 Explainable AI-Enhanced Kinase Activity Profiling Through Phosphoproteomics
abstract
Kinases play a critical role in regulating fundamental cellular processes, including metabolism, signal transduction, and cell growth, primarily through phosphorylation. The dysregulation of kinase activity is implicated in various diseases, highlighting the urgent need for robust and interpretable methodologies to profile this activity. Current approaches frequently depend on overly complex or limited datasets, lack generalizability, or fail to provide meaningful biological insights into the mechanisms governing kinase activity. To address these challenges, we developed an explainable deep learning framework that leverages mass spectrometry-based phosphoproteomics data to profile kinase activity effectively. Our study systematically evaluated deep neural networks (DNNs) and convolutional neural networks (CNNs), incorporating a diverse set of feature inputs, including phosphorylation sites and kinase-substrate relationships. A notable finding was that a three-layer CNN, optimized through rigorous feature selection techniques, demonstrated superior performance, achieving substantial improvements in prediction accuracy and stability when compared to established methods such as kinase-substrate enrichment analysis (KSEA) and the kinase activity ranking pipeline (KARP). We integrated Shapley additive explanations (SHAP) values to enhance interpretability, illuminating biologically significant phosphorylation sites. For example, PAK2-related phosphorylation sites associated with the progression of colon adenocarcinoma and CAMK2D sites integral to adrenergic signaling were identified, thereby effectively linking computational predictions to established molecular pathways. This research illustrates the potential of explainable artificial intelligence in advancing kinase activity profiling by providing accurate and interpretable predictions. Our framework is valuable for elucidating disease mechanisms and identifying therapeutic targets, facilitating broader applications in precision medicine.
Chia-Ru Chung, Ming-Feng Ho, Li-Ching Wu, Justin Bo-Kai Hsu, Tzong-Yi Lee, Jorng-Tzong Horng
CIBCB6
2025 miSAM: Robust MicroRNA Expression Estimation Using Bi-Objective Evolutionary Learning Algorithm in Cancer Transcriptomics
abstract
MicroRNAs (miRNAs) play crucial regulatory roles in cancer biology, but accurately quantifying their expression remains a significant unmet challenge in single-cell and spatial transcriptomics. Current sequencing technologies predominantly capture polyadenylated messenger RNAs (mRNAs), rendering them incapable of directly profiling miRNAs, which lack poly(A) tails. To address this gap, we thus propose miSAM, a novel computational framework for estimation of miRNA expressions based on a bi-objective combinatorial genetic algorithm in conjunction with support vector regression. The miSAM jointly optimizes the selection of a minimal subset of mRNAs, called signatures, while maximizing the Spearman correlation coefficient (SCC) between inferred and actual miRNA expression levels. Evaluation on The Cancer Genome Atlas (TCGA)-BRCA dataset, comprising 1,095 breast cancer samples with expression profiles of 1,881 miRNAs and 19,937 mRNAs, demonstrates the effectiveness of miSAM. Using the top five prognostic miRNA biomarkers for breast cancer, miSAM achieved a mean SCC of 0.633 on the test set while utilizing only 32.6 mRNAs in average, significantly outperforming baseline approaches including: (1) differential expression filtering-based SVR using 887 mRNAs (SCC = 0.562), and (2) LASSO-based SVR using 147.4 mRNAs (SCC = 0.470). miSAM also outperformed XGBoost, which yielded a SCC ofapproximately 0.45–0.50 across various cancer types. Furthermore, the identified mRNAs signatures offer explainable insights into the regulatory associations between miRNAs and their corresponding mRNA targets. These results underscore miSAM’s potential as a robust, interpretable, and scalable tool for miRNA inference in spatial transcriptomics, single-cell sequencing analyses, and precision oncology.
Yann-Lin Ho, Yann-Jen Ho, Shinn-Ying Ho, Tzong-Yi Lee
CIBCB4
2025 EL-DRP: An Evolutionary Learning-Based Multi-Omics Framework for Drug Response Prediction with Interpretable Biomarker Selection
abstract
Cancer is a complex disease driven by diverse genetic, epigenetic, and microenvironmental alterations that result in dysregulated cell proliferation and therapeutic resistance. Both inter- and intratumoral heterogeneity further complicate treatment outcomes, highlighting the critical need for effective and efficient drug response prediction in precision oncology. While deep learning models have yielded robust predictive performance, their limited interpretability remains a critical barrier to clinical translation. To overcome this, we propose EL-DRP, an evolutionary learning framework for drug response prediction that integrates multi-omics data, including mRNA expression, copy number variation, and single nucleotide polymorphisms. Central to EL-DRP is the use of an Inheritable Bi-objective Combinatorial Genetic Algorithm (IBCGA) for optimized feature selection within each omics modality. IBCGA identifies minimal yet informative subsets of features, which are subsequently integrated to construct a unified multi-omics predictive model. EL-DRP achieved an average AUROC of 0.791 ± 0.072 for 21 drugs using an average of 106 biomarkers per drug, most of which align with known drug mechanisms of action. For instance, among the camptothecin response–associated biomarkers, KDM4A and SCML2 are known to participate in camptothecin-induced DNA damage response, whereas SALL2, RBM10, BAHCC1, and PBRM1 regulate chromatin structure and genome maintenance, corroborating their mechanistic relevance. Importantly, this minimal biomarker panel alone retains robust predictive power with explainable biomarkers. These results provide a new perspective for discovering potential biomarkers associated with drug response prediction and lay the groundwork for future studies integrating clinical data with translational potential.
Yann-Jen Ho, Paik You Sheng, Yu-Ruo Chen, Shinn-Ying Ho, Tzong-Yi Lee
CIBCB6
2025 Integrative Multi-Omics Prognostic Modeling of Glioma Recurrence Using Variational Autoencoder and Similarity Network Fusion
abstract
Gliomas represent the most prevalent form of malignant brain tumors, with glioblastoma (GBM) classified as the most aggressive subtype and lower-grade gliomas (LGGs) showing a high recurrence rate despite better survival outcomes. Traditional prognostic methods rely on clinical and molecular markers, yet they often fail to leverage the predictive potential of integrative multiomics data. Hence, early and accurate identification of glioma recurrence is critical for optimizing treatment strategies and improving patient outcomes. To address this challenge, we developed a deep learning-based approach that integrates multiomics data with clinical features to enhance recurrence risk stratification. Multiomics data, including mRNA, miRNA, DNA methylation, and CNV profiles from 512 LGG and 350 GBM patients in The Cancer Genome Atlas (TCGA), were processed using the variational autoencoder for nonlinear feature extraction, followed by similarity network fusion to capture cross-omics relationships. Recurrence-guided feature selection identified a robust biomarker panel that effectively categorized patients into high-and low-risk recurrence groups, with Kaplan-Meier survival analysis demonstrating a notable significant difference with p-value less than 0.001. Integration of clinical factors (age, tumor grade, IDH mutation status) further improved predictive performance, yielding a C-index of 0.695 (p < 2e-16) for LGG and 0.619 (p = 3.346e-11) for GBM. Together, these findings establish a multiomics-based predictive model that enables refined glioma recurrence risk assessment, offering personalized treatment strategies for LGG and GBM patients.
Phuong Lam Tran, Justin Bo-Kai Hsu, Tzong-Yi Lee
CIBCB4
2025 Towards Accurate Identification of Anti-Hepatitis C Peptides Using Stack-AHCP
abstract
Hepatitis C virus (HCV) infection remains a significant global health burden, contributing to progressive hepatic pathologies including chronic hepatitis, cirrhosis, and hepatocellular carcinoma. While anti-hepatitis C peptides (AHCPs) have emerged as promising therapeutic candidates with distinct antiviral mechanisms, conventional wet-lab approaches for AHCP discovery face critical limitations in throughput and scalability. To overcome these constraints, we present Stack-Ahcp, an innovative stacked ensemble learning framework that synergistically integrates multiple machine learning algorithms through a meta-classification strategy. Our model achieves unprecedented predictive performance with 93.1% accuracy and an MCC of 0.863, substantially outperforming existing computational methods. Through comprehensive Shapley additive explanations (SHAP) analysis, we further delineate critical key intrinsic features and determinants governing AHCP bioactivity, enhancing the mechanistic interpretability of the prediction system. To facilitate translational applications, we have implemented an intuitive web interface (accessible at https://awi.cuhk.edu.cn/~biosequence/StackAHCP/index.php) that enables rapid screening and prioritization of candidate peptides. This resource is anticipated to streamline the identification of next-generation peptide therapeutics against HCV while reducing experimental validation costs. Beyond virology applications, our methodological framework establishes a paradigm for interpretable machine learning in biological sequence analysis, with potential adaptability to diverse multi-omics investigation scenarios.
Lantian Yao, Yen-Peng Chiu, Jiahui Guan, Peilin Xie, Yulan Liu, Yunlu Peng, Ying-Chih Chiang, Tzong-Yi Lee
CIBCB12
2025 Graph-RPI: predicting RNA-protein interactions via graph autoencoder and self-supervised learning strategies
abstract
RNA-protein interactions (RPIs) are essential for many biological functions and are associated with various diseases. Traditional methods for detecting RPIs are labor-intensive and costly, necessitating efficient computational methods. In this study, we proposed a novel sequence-based RPI prediction framework based on graph neural networks (GNNs) that addressed key limitations of existing methods, such as inadequate feature integration and negative sample construction. Our method represented RNAs and proteins as nodes in a unified interaction graph, enhancing the representation of RPI pairs through multi-feature fusion and employing self-supervised learning strategies for model training. The model's performance was validated through five-fold cross-validation, achieving accuracy of 0.880, 0.811, 0.950, 0.979, 0.910, and 0.924 on the RPI488, RPI369, RPI2241, RPI1807, RPI1446, and RPImerged datasets, respectively. Additionally, in cross-species generalization tests, our method outperformed existing methods, achieving an overall accuracy of 0.989 across 10 093 RPI pairs. Compared with other state-of-the-art RPI prediction methods, our approach demonstrates greater robustness and stability in RPI prediction, highlighting its potential for broad biological applications and large-scale RPI analysis.
Jiahui Guan, Lantian Yao, Peilin Xie, Dian Meng, Tzong-Yi Lee, Junwen Wang, Ying-Chih Chiang
Briefings Bioinform.6
2025 STForte: tissue context-specific encoding and consistency-aware spatial imputation for spatially resolved transcriptomics
abstract
Encoding spatially resolved transcriptomics (SRT) data serves to identify the biological semantics of RNA expression within the tissue while preserving spatial characteristics. Depending on the analytical scenario, one may focus on different contextual structures of tissues. For instance, anatomical regions reveal consistent patterns by focusing on spatial homogeneity, while elucidating complex tumor micro-environments requires more expression heterogeneity. However, current spatial encoding methods lack consideration of the tissue context. Meanwhile, most developed SRT technologies are still limited in providing exact patterns of intact tissues due to limitations such as low resolution or missed measurements. Here, we propose STForte, a novel pairwise graph autoencoder-based approach with cross-reconstruction and adversarial distribution matching, to model the spatial homogeneity and expression heterogeneity of SRT data. STForte extracts interpretable latent encodings, enabling downstream analysis by accurately portraying various tissue contexts. Moreover, STForte allows spatial imputation using only spatial consistency to restore the biological patterns of unobserved locations or low-quality cells, thereby providing fine-grained views to enhance the SRT analysis. Extensive evaluations of datasets under different scenarios and SRT platforms demonstrate that STForte is a scalable and versatile tool for providing enhanced insights into spatial data analysis.
Yuxuan Pang, Chunxuan Wang, Seiya Imoto, Tzong-Yi Lee
Briefings Bioinform.6
2025 Toward high-efficiency, low-resource, and explainable neuropeptide prediction with MSKDNP
abstract
Neuropeptides are essential signaling molecules produced in the nervous system that regulate diverse physiological processes and are closely implicated in the pathogenesis of neurodegenerative and neuropsychiatric disorders. Investigating neuropeptides contributes to a better understanding of their regulatory mechanisms and offers new insights into therapeutic strategies for related diseases. Therefore, accurate identification of neuropeptides is crucial for advancing biomedical research and drug development. Due to the high cost of experimental validation, various artificial intelligence methods have been developed for rapid neuropeptide identification. However, existing approaches often suffer from high computational resource consumption, slow processing speed, and poor deploy ability. Moreover, a user-friendly web server for practical application is still lacking. To this end, we propose MSKDNP, a neuropeptide prediction model based on a multi-stage knowledge distillation framework. With only 1.2% of the parameters, MSKDNP attains performance comparable to a fully fine-tuned protein language model while achieving state-of-the-art results in neuropeptide recognition. Moreover, MSKDNP provides favorable interpretability, facilitating biological understanding. A freely accessible web server is available at https://awi.cuhk.edu.cn/∼biosequence/MSKDNP/index.php.
Peilin Xie, Jiahui Guan, Yulan Liu, Zhang Cheng, Xuxin He, Zhenglong Sun 0001, Tzong-Yi Lee, Lantian Yao, Ying-Chih Chiang
Briefings Bioinform.10
2024 A two-stage computational framework for identifying antiviral peptides and their functional types based on contrastive learning and multi-feature fusion strategy
abstract
Antiviral peptides (AVPs) have shown potential in inhibiting viral attachment, preventing viral fusion with host cells and disrupting viral replication due to their unique action mechanisms. They have now become a broad-spectrum, promising antiviral therapy. However, identifying effective AVPs is traditionally slow and costly. This study proposed a new two-stage computational framework for AVP identification. The first stage identifies AVPs from a wide range of peptides, and the second stage recognizes AVPs targeting specific families or viruses. This method integrates contrastive learning and multi-feature fusion strategy, focusing on sequence information and peptide characteristics, significantly enhancing predictive ability and interpretability. The evaluation results of the model show excellent performance, with accuracy of 0.9240 and Matthews correlation coefficient (MCC) score of 0.8482 on the non-AVP independent dataset, and accuracy of 0.9934 and MCC score of 0.9869 on the non-AMP independent dataset. Furthermore, our model can predict antiviral activities of AVPs against six key viral families (Coronaviridae, Retroviridae, Herpesviridae, Paramyxoviridae, Orthomyxoviridae, Flaviviridae) and eight viruses (FIV, HCV, HIV, HPIV3, HSV1, INFVA, RSV, SARS-CoV). Finally, to facilitate user accessibility, we built a user-friendly web interface deployed at https://awi.cuhk.edu.cn/∼dbAMP/AVP/.
Jiahui Guan, Lantian Yao, Peilin Xie, Chia-Ru Chung, Yixian Huang, Ying-Chih Chiang, Tzong-Yi Lee
Briefings Bioinform.7
2024 Guided diffusion for molecular generation with interaction prompt
abstract
Molecular generative models have exhibited promising capabilities in designing molecules from scratch with high binding affinities in a predetermined protein pocket, offering potential synergies with traditional structural-based drug design strategy. However, the generative processes of such models are random and the atomic interaction information between ligand and protein are ignored. On the other hand, the ligand has high propensity to bind with residues called hotspots. Hotspot residues contribute to the majority of the binding free energies and have been recognized as appealing targets for designed molecules. In this work, we develop an interaction prompt guided diffusion model, InterDiff to deal with the challenges. Four kinds of atomic interactions are involved in our model and represented as learnable vector embeddings. These embeddings serve as conditions for individual residue to guide the molecular generative process. Comprehensive in silico experiments evince that our model could generate molecules with desired ligand-protein interactions in a guidable way. Furthermore, we validate InterDiff on two realistic protein-based therapeutic agents. Results show that InterDiff could generate molecules with better or similar binding mode compared to known targeted drugs.
Huabin Du, Yingchao Yan, Tzong-Yi Lee
Briefings Bioinform.4
2024 ACP-CapsPred: an explainable computational framework for identification and functional prediction of anticancer peptides based on capsule network
abstract
Cancer is a severe illness that significantly threatens human life and health. Anticancer peptides (ACPs) represent a promising therapeutic strategy for combating cancer. In silico methods enable rapid and accurate identification of ACPs without extensive human and material resources. This study proposes a two-stage computational framework called ACP-CapsPred, which can accurately identify ACPs and characterize their functional activities across different cancer types. ACP-CapsPred integrates a protein language model with evolutionary information and physicochemical properties of peptides, constructing a comprehensive profile of peptides. ACP-CapsPred employs a next-generation neural network, specifically capsule networks, to construct predictive models. Experimental results demonstrate that ACP-CapsPred exhibits satisfactory predictive capabilities in both stages, reaching state-of-the-art performance. In the first stage, ACP-CapsPred achieves accuracies of 80.25% and 95.71%, as well as F1-scores of 79.86% and 95.90%, on benchmark datasets Set 1 and Set 2, respectively. In the second stage, tasked with characterizing the functional activities of ACPs across five selected cancer types, ACP-CapsPred attains an average accuracy of 90.75% and an F1-score of 91.38%. Furthermore, ACP-CapsPred demonstrates excellent interpretability, revealing regions and residues associated with anticancer activity. Consequently, ACP-CapsPred presents a promising solution to expedite the development of ACPs and offers a novel perspective for other biological sequence analyses.
Lantian Yao, Peilin Xie, Jiahui Guan, Chia-Ru Chung, Wenyang Zhang, Junyang Deng, Yixian Huang, Ying-Chih Chiang, Tzong-Yi Lee
Briefings Bioinform.9
2023 Extraction of microRNA-target interaction sentences from biomedical literature by deep learning approach
abstract
MicroRNA (miRNA)-target interaction (MTI) plays a substantial role in various cell activities, molecular regulations and physiological processes. Published biomedical literature is the carrier of high-confidence MTI knowledge. However, digging out this knowledge in an efficient manner from large-scale published articles remains challenging. To address this issue, we were motivated to construct a deep learning-based model. We applied the pre-trained language models to biomedical text to obtain the representation, and subsequently fed them into a deep neural network with gate mechanism layers and a fully connected layer for the extraction of MTI information sentences. Performances of the proposed models were evaluated using two datasets constructed on the basis of text data obtained from miRTarBase. The validation and test results revealed that incorporating both PubMedBERT and SciBERT for sentence level encoding with the long short-term memory (LSTM)-based deep neural network can yield an outstanding performance, with both F1 and accuracy being higher than 80% on validation data and test data. Additionally, the proposed deep learning method outperformed the following machine learning methods: random forest, support vector machine, logistic regression and bidirectional LSTM. This work would greatly facilitate studies on MTI analysis and regulations. It is anticipated that this work can assist in large-scale screening of miRNAs, thereby revealing their functional roles in various diseases, which is important for the development of highly specific drugs with fewer side effects. Source code and corpus are publicly available at https://github.com/qi29.
Mengqi Luo, Shangfu Li, Yuxuan Pang, Lantian Yao, Renfei Ma, Hsi-Yuan Huang, Hsien-Da Huang, Tzong-Yi Lee
Briefings Bioinform.8
2023 Holistic similarity-based prediction of phosphorylation sites for understudied kinases
abstract
Phosphorylation is an essential mechanism for regulating protein activities. Determining kinase-specific phosphorylation sites by experiments involves time-consuming and expensive analyzes. Although several studies proposed computational methods to model kinase-specific phosphorylation sites, they typically required abundant experimentally verified phosphorylation sites to yield reliable predictions. Nevertheless, the number of experimentally verified phosphorylation sites for most kinases is relatively small, and the targeting phosphorylation sites are still unidentified for some kinases. In fact, there is little research related to these understudied kinases in the literature. Thus, this study aims to create predictive models for these understudied kinases. A kinase-kinase similarity network was generated by merging the sequence-, functional-, protein-domain- and 'STRING'-related similarities. Thus, besides sequence data, protein-protein interactions and functional pathways were also considered to aid predictive modelling. This similarity network was then integrated with a classification of kinase groups to yield highly similar kinases to a specific understudied type of kinase. Their experimentally verified phosphorylation sites were leveraged as positive sites to train predictive models. The experimentally verified phosphorylation sites of the understudied kinase were used for validation. Results demonstrate that 82 out of 116 understudied kinases were predicted with adequate performance via the proposed modelling strategy, achieving a balanced accuracy of 0.81, 0.78, 0.84, 0.84, 0.85, 0.82, 0.90, 0.82 and 0.85, for the 'TK', 'Other', 'STE', 'CAMK', 'TKL', 'CMGC', 'AGC', 'CK1' and 'Atypical' groups, respectively. Therefore, this study demonstrates that web-like predictive networks can reliably capture the underlying patterns in such understudied kinases by harnessing relevant sources of similarities to predict their specific phosphorylation sites.
Renfei Ma, Shangfu Li, Luca Parisi, Hsien-Da Huang, Tzong-Yi Lee
Briefings Bioinform.6
2023 Identification of species-specific RNA N6-methyladinosine modification sites from RNA sequences
abstract
N6-methyladinosine (m6A) modification is the most abundant co-transcriptional modification in eukaryotic RNA and plays important roles in cellular regulation. Traditional high-throughput sequencing experiments used to explore functional mechanisms are time-consuming and labor-intensive, and most of the proposed methods focused on limited species types. To further understand the relevant biological mechanisms among different species with the same RNA modification, it is necessary to develop a computational scheme that can be applied to different species. To achieve this, we proposed an attention-based deep learning method, adaptive-m6A, which consists of convolutional neural network, bi-directional long short-term memory and an attention mechanism, to identify m6A sites in multiple species. In addition, three conventional machine learning (ML) methods, including support vector machine, random forest and logistic regression classifiers, were considered in this work. In addition to the performance of ML methods for multi-species prediction, the optimal performance of adaptive-m6A yielded an accuracy of 0.9832 and the area under the receiver operating characteristic curve of 0.98. Moreover, the motif analysis and cross-validation among different species were conducted to test the robustness of one model towards multiple species, which helped improve our understanding about the sequence characteristics and biological functions of RNA modifications in different species.
Rulan Wang, Chia-Ru Chung, Hsien-Da Huang, Tzong-Yi Lee
Briefings Bioinform.4
2023 A risk assessment framework for multidrug-resistant Staphylococcus aureus using machine learning and mass spectrometry technology
abstract
The emergence of multidrug-resistant bacteria is a critical global crisis that poses a serious threat to public health, particularly with the rise of multidrug-resistant Staphylococcus aureus. Accurate assessment of drug resistance is essential for appropriate treatment and prevention of transmission of these deadly pathogens. Early detection of drug resistance in patients is critical for providing timely treatment and reducing the spread of multidrug-resistant bacteria. This study aims to develop a novel risk assessment framework for S. aureus that can accurately determine the resistance to multiple antibiotics. The comprehensive 7-year study involved ˃20 000 isolates with susceptibility testing profiles of six antibiotics. By incorporating mass spectrometry and machine learning, the study was able to predict the susceptibility to four different antibiotics with high accuracy. To validate the accuracy of our models, we externally tested on an independent cohort and achieved impressive results with an area under the receiver operating characteristic curve of 0. 94, 0.90, 0.86 and 0.91, and an area under the precision-recall curve of 0.93, 0.87, 0.87 and 0.81, respectively, for oxacillin, clindamycin, erythromycin and trimethoprim-sulfamethoxazole. In addition, the framework evaluated the level of multidrug resistance of the isolates by using the predicted drug resistance probabilities, interpreting them in the context of a multidrug resistance risk score and analyzing the performance contribution of different sample groups. The results of this study provide an efficient method for early antibiotic decision-making and a better understanding of the multidrug resistance risk of S. aureus.
Yuxuan Pang, Chia-Ru Chung, Hsin-Yao Wang, Haiyan Cui, Ying-Chih Chiang, Jorng-Tzong Horng, Jang-Jih Lu, Tzong-Yi Lee
Briefings Bioinform.9
2022 Integrating transformer and imbalanced multi-label learning to identify antimicrobial peptides and their functional activities
abstract
MOTIVATION: Antimicrobial peptides (AMPs) have the potential to inhibit multiple types of pathogens and to heal infections. Computational strategies can assist in characterizing novel AMPs from proteome or collections of synthetic sequences and discovering their functional abilities toward different microbial targets without intensive labor. RESULTS: Here, we present a deep learning-based method for computer-aided novel AMP discovery that utilizes the transformer neural network architecture with knowledge from natural language processing to extract peptide sequence information. We implemented the method for two AMP-related tasks: the first is to discriminate AMPs from other peptides, and the second task is identifying AMPs functional activities related to seven different targets (gram-negative bacteria, gram-positive bacteria, fungi, viruses, cancer cells, parasites and mammalian cell inhibition), which is a multi-label problem. In addition, asymmetric loss was adopted to resolve the intrinsic imbalance of dataset, particularly for the multi-label scenarios. The evaluation showed that our proposed scheme achieves the best performance for the first task (96.85% balanced accuracy) and has a more unbiased prediction for the second task (79.83% balanced accuracy averaged across all functional activities) when compared with that of strategies without imbalanced learning or deep learning. AVAILABILITY AND IMPLEMENTATION: The source code and data of this study are available at https://github.com/BiOmicsLab/TransImbAMP. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Yuxuan Pang, Lantian Yao, Tzong-Yi Lee
Bioinform.5
2021 Identifying anti-coronavirus peptides by incorporating different negative datasets and imbalanced learning strategies
abstract
As the current worldwide outbreaks of the SARS-CoV-2, it is urgently needed to develop effective therapeutic agents for inhibiting the pathogens or treating the related diseases. Antimicrobial peptides (AMP) with functional activity against coronavirus could be a considerable solution, yet there is no research for identifying anti-coronavirus (anti-CoV) peptides with the computational approach. In this study, we first investigated the physiochemical and compositional properties of the collected anti-CoV peptides by comparing against three other negative sets: antivirus peptides without anti-CoV function (antivirus), regular AMP without antivirus functions (non-AVP) and peptides without antimicrobial functions (non-AMP). Then, we established classifiers for identifying anti-CoV peptides between different negative sets based on random forest. Imbalanced learning strategies were adopted due to the severe class-imbalance within the datasets. The geometric mean of the sensitivity and specificity (GMean) under the identification from antivirus, non-AVP and non-AMP reaches 83.07%, 85.51% and 98.82%, respectively. Then, to pursue identifying anti-CoV peptides from broad-spectrum peptides, we designed a double-stages classifier based on the collected datasets. In the first stage, the classifier characterizes AMPs from regular peptides. It achieves an area under the receiver operating curve (AUCROC) value of 97.31%. The second stage is to identify the anti-CoV peptides between the combined negatives of other AMPs. Here, the GMean of evaluation on the independent test set is 79.42%. The proposed approach is considered as an applicable scheme for assisting the development of novel anti-CoV peptides. The datasets and source codes used in this study are available at https://github.com/poncey/PreAntiCoV.
Yuxuan Pang, Jhih-Hua Jhong, Tzong-Yi Lee
Briefings Bioinform.4
2021 AVPIden: a new scheme for identification and functional prediction of antiviral peptides based on machine learning approaches
abstract
Antiviral peptide (AVP) is a kind of antimicrobial peptide (AMP) that has the potential ability to fight against virus infection. Machine learning-based prediction with a computational biology approach can facilitate the development of the novel therapeutic agents. In this study, we proposed a double-stage classification scheme, named AVPIden, for predicting the AVPs and their functional activities against different viruses. The first stage is to distinguish the AVP from a broad-spectrum peptide collection, including not only the regular peptides (non-AMP) but also the AMPs without antiviral functions (non-AVP). The second stage is responsible for characterizing one or more virus families or species that the AVP targets. Imbalanced learning is utilized to improve the performance of prediction. The AVPIden uses multiple descriptors to precisely demonstrate the peptide properties and adopts explainable machine learning strategies based on Shapley value to exploit how the descriptors impact the antiviral activities. Finally, the evaluation performance of the proposed model suggests its ability to predict the antivirus activities and their potential functions against six virus families (Coronaviridae, Retroviridae, Herpesviridae, Paramyxoviridae, Orthomyxoviridae, Flaviviridae) and eight kinds of virus (FIV, HCV, HIV, HPIV3, HSV1, INFVA, RSV, SARS-CoV). The AVPIden gives an option for reinforcing the development of AVPs with the computer-aided method and has been deployed at http://awi.cuhk.edu.cn/AVPIden/.
Yuxuan Pang, Lantian Yao, Jhih-Hua Jhong, Tzong-Yi Lee
Briefings Bioinform.5
2021 A large-scale investigation and identification of methicillin-resistant Staphylococcus aureus based on peaks binning of matrix-assisted laser desorption ionization-time of flight MS spectra
abstract
Recent studies have demonstrated that the matrix-assisted laser desorption ionization-time of flight mass spectrometry (MALDI-TOF MS) could be used to detect superbugs, such as methicillin-resistant Staphylococcus aureus (MRSA). Due to an increasingly clinical need to classify between MRSA and methicillin-sensitive Staphylococcus aureus (MSSA) efficiently and effectively, we were motivated to develop a systematic pipeline based on a large-scale dataset of MS spectra. However, the shifting problem of peaks in MS spectra induced a low effectiveness in the classification between MRSA and MSSA isolates. Unlike previous works emphasizing on specific peaks, this study employs a binning method to cluster MS shifting ions into several representative peaks. A variety of bin sizes were evaluated to coalesce drifted or shifted MS peaks to a well-defined structured data. Then, various machine learning methods were performed to carry out the classification between MRSA and MSSA samples. Totally 4858 MS spectra of unique S. aureus isolates, including 2500 MRSA and 2358 MSSA instances, were collected by Chang Gung Memorial Hospitals, at Linkou and Kaohsiung branches, Taiwan. Based on the evaluation of Pearson correlation coefficients and the strategy of forward feature selection, a total of 200 peaks (with the bin size of 10 Da) were identified as the marker attributes for the construction of predictive models. These selected peaks, such as bins 2410-2419, 2450-2459 and 6590-6599 Da, have indicated remarkable differences between MRSA and MSSA, which were effective in the prediction of MRSA. The independent testing has revealed that the random forest model can provide a promising prediction with the area under the receiver operating characteristic curve (AUC) at 0.8450. When comparing to previous works conducted with hundreds of MS spectra, the proposed scheme demonstrates that incorporating machine learning method with a large-scale dataset of clinical MS spectra may be a feasible means for clinical physicians on the administration of correct antibiotics in shorter turn-around-time, which could reduce mortality, avoid drug resistance and shorten length of stay in hospital in the future.
Hsin-Yao Wang, Chia-Ru Chung, Shangfu Li, Bo-Yu Chu, Jorng-Tzong Horng, Jang-Jih Lu, Tzong-Yi Lee
Briefings Bioinform.8
2021 Large-scale mass spectrometry data combined with demographics analysis rapidly predicts methicillin resistance in Staphylococcus aureus
abstract
BACKGROUND: A mass spectrometry-based assessment of methicillin resistance in Staphylococcus aureus would have huge potential in addressing fast and effective prediction of antibiotic resistance. Since delays in the traditional antibiotic susceptibility testing, methicillin-resistant S. aureus remains a serious threat to human health. RESULTS: Here, linking a 7 years of longitudinal study from two cohorts in the Taiwan area of over 20 000 individually resolved methicillin susceptibility testing results, we identify associations of methicillin resistance with the demographics and mass spectrometry data. When combined together, these connections allow for machine-learning-based predictions of methicillin resistance, with an area under the receiver operating characteristic curve of >0.85 in both the discovery [95% confidence interval (CI) 0.88-0.90] and replication (95% CI 0.84-0.86) populations. CONCLUSIONS: Our predictive model facilitates early detection for methicillin resistance of patients with S. aureus infection. The large-scale antibiotic resistance study has unbiasedly highlighted putative candidates that could improve trials of treatment efficiency and inform on prescriptions.
Hsin-Yao Wang, Chia-Ru Chung, Jorng-Tzong Horng, Jang-Jih Lu, Tzong-Yi Lee
Briefings Bioinform.6
2021 A representation and deep learning model for annotating ubiquitylation sentences stating E3 ligase - substrate interaction
abstract
BACKGROUND: Ubiquitylation is an important post-translational modification of proteins that not only plays a central role in cellular coding, but is also closely associated with the development of a variety of diseases. The specific selection of substrate by ligase E3 is the key in ubiquitylation. As various high-throughput analytical techniques continue to be applied to the study of ubiquitylation, a large amount of ubiquitylation site data, and records of E3-substrate interactions continue to be generated. Biomedical literature is an important vehicle for information on E3-substrate interactions in ubiquitylation and related new discoveries, as well as an important channel for researchers to obtain such up to date data. The continuous explosion of ubiquitylation related literature poses a great challenge to researchers in acquiring and analyzing the information. Therefore, automatic annotation of these E3-substrate interaction sentences from the available literature is urgently needed. RESULTS: In this research, we proposed a model based on representation and attention mechanism based deep learning methods, to automatic annotate E3-substrate interaction sentences in biomedical literature. Focusing on the sentences with E3 protein inside, we applied several natural language processing methods and a Long Short-Term Memory (LSTM)-based deep learning classifier to train the model. Experimental results had proved the effectiveness of our proposed model. And also, the proposed attention mechanism deep learning method outperforms other statistical machine learning methods. We also created a manual corpus of E3-substrate interaction sentences, in which the E3 proteins and substrate proteins are also labeled, in order to construct our model. The corpus and model proposed by our research are definitely able to be very useful and valuable resource for advancement of ubiquitylation-related research. CONCLUSION: Having the entire manual corpus of E3-substrate interaction sentences readily available in electronic form will greatly facilitate subsequent text mining and machine learning analyses. Automatic annotating ubiquitylation sentences stating E3 ligase-substrate interaction is significantly benefited from semantic representation and deep learning. The model enables rapid information accessing and can assist in further screening of key ubiquitylation ligase substrates for in-depth studies.
Mengqi Luo, Zhongyan Li, Shangfu Li, Tzong-Yi Lee
BMC Bioinform.4
2021 Incorporating support vector machine with sequential minimal optimization to identify anticancer peptides
abstract
BACKGROUND: Cancer is one of the major causes of death worldwide. To treat cancer, the use of anticancer peptides (ACPs) has attracted increased attention in recent years. ACPs are a unique group of small molecules that can target and kill cancer cells fast and directly. However, identifying ACPs by wet-lab experiments is time-consuming and labor-intensive. Therefore, it is significant to develop computational tools for ACPs prediction. Though some ACP prediction tools have been developed recently, their performances are not well enough and most of them do not offer a function to distinguish ACPs from antimicrobial peptides (AMPs). Considering the fact that a growing number of studies have shown that some AMPs exhibit anticancer function, this work tries to build a model for distinguishing AMPs from ACPs in addition to a model that predicts ACPs from whole peptides. RESULTS: This study chooses amino acid composition, N5C5, k-space, position-specific scoring matrix (PSSM) as features, and analyzes them by machine learning methods, including support vector machine (SVM) and sequential minimal optimization (SMO) to build a model (model 2) for distinguishing ACPs from whole peptides. Another model (model 1) that distinguishes ACPs from AMPs is also developed. Comparing to previous models, models developed in this research show better performance (accuracy: 85.5% for model 1 and 95.2% for model 2). CONCLUSIONS: This work utilizes a new feature, PSSM, which contributes to better performance than other features. In addition to SVM, SMO is used in this research for optimizing SVM and the SMO-optimized models show better performance than non-optimized models. Last but not least, this work provides two different functions, including distinguishing ACPs from AMPs and distinguishing ACPs from all peptides. The second SMO-optimized model, which utilizes PSSM as a feature, performs better than all other existing tools.
Tzong-Yi Lee
BMC Bioinform.3
2020 Characterization and identification of antimicrobial peptides with different functional activities
abstract
In recent years, antimicrobial peptides (AMPs) have become an emerging area of focus when developing therapeutics hot spot residues of proteins are dominant against infections. Importantly, AMPs are produced by virtually all known living organisms and are able to target a wide range of pathogenic microorganisms, including viruses, parasites, bacteria and fungi. Although several studies have proposed different machine learning methods to predict peptides as being AMPs, most do not consider the diversity of AMP activities. On this basis, we specifically investigated the sequence features of AMPs with a range of functional activities, including anti-parasitic, anti-viral, anti-cancer and anti-fungal activities and those that target mammals, Gram-positive and Gram-negative bacteria. A new scheme is proposed to systematically characterize and identify AMPs and their functional activities. The 1st stage of the proposed approach is to identify the AMPs, while the 2nd involves further characterization of their functional activities. Sequential forward selection was employed to extract potentially informative features that are possibly associated with the functional activities of the AMPs. These features include hydrophobicity, the normalized van der Waals volume, polarity, charge and solvent accessibility-all of which are essential attributes in classifying between AMPs and non-AMPs. The results revealed the 1st stage AMP classifier was able to achieve an area under the receiver operating characteristic curve (AUC) value of 0.9894. During the 2nd stage, we found pseudo amino acid composition to be an informative attribute when differentiating between AMPs in terms of their functional activities. The independent testing results demonstrated that the AUCs of the multi-class models were 0.7773, 0.9404, 0.8231, 0.8578, 0.8648, 0.8745 and 0.8672 for anti-parasitic, anti-viral, anti-cancer, anti-fungal AMPs and those that target mammals, Gram-positive and Gram-negative bacteria, respectively. The proposed scheme helps facilitate biological experiments related to the functional analysis of AMPs. Additionally, it was implemented as a user-friendly web server (AMPfun, http://fdblab.csie.ncu.edu.tw/AMPfun/index.html) that allows individuals to explore the antimicrobial functions of peptides of interest.
Chia-Ru Chung, Ting-Rung Kuo, Li-Ching Wu, Tzong-Yi Lee, Jorng-Tzong Horng
Briefings Bioinform.4
2019 Characterization and identification of lysine glutarylation based on intrinsic interdependence between positions in the substrate sites
abstract
BACKGROUND: Glutarylation, the addition of a glutaryl group (five carbons) to a lysine residue of a protein molecule, is an important post-translational modification and plays a regulatory role in a variety of physiological and biological processes. As the number of experimentally identified glutarylated peptides increases, it becomes imperative to investigate substrate motifs to enhance the study of protein glutarylation. We carried out a bioinformatics investigation of glutarylation sites based on amino acid composition using a public database containing information on 430 non-homologous glutarylation sites. RESULTS: The TwoSampleLogo analysis indicates that positively charged and polar amino acids surrounding glutarylated sites may be associated with the specificity in substrate site of protein glutarylation. Additionally, the chi-squared test was utilized to explore the intrinsic interdependence between two positions around glutarylation sites. Further, maximal dependence decomposition (MDD), which consists of partitioning a large-scale dataset into subgroups with statistically significant amino acid conservation, was used to capture motif signatures of glutarylation sites. We considered single features, such as amino acid composition (AAC), amino acid pair composition (AAPC), and composition of k-spaced amino acid pairs (CKSAAP), as well as the effectiveness of incorporating MDD-identified substrate motifs into an integrated prediction model. Evaluation by five-fold cross-validation showed that AAC was most effective in discriminating between glutarylation and non-glutarylation sites, according to support vector machine (SVM). CONCLUSIONS: The SVM model integrating MDD-identified substrate motifs performed well, with a sensitivity of 0.677, a specificity of 0.619, an accuracy of 0.638, and a Matthews Correlation Coefficient (MCC) value of 0.28. Using an independent testing dataset (46 glutarylated and 92 non-glutarylated sites) obtained from the literature, we demonstrated that the integrated SVM model could improve the predictive performance effectively, yielding a balanced sensitivity and specificity of 0.652 and 0.739, respectively. This integrated SVM model has been implemented as a web-based system (MDDGlutar), which is now freely available at http://csb.cse.yzu.edu.tw/MDDGlutar/ .
Kai-Yao Huang, Hui-Ju Kao, Justin Bo-Kai Hsu, Shun-Long Weng, Tzong-Yi Lee
BMC Bioinform.5
2019 Rapid classification of group B Streptococcus serotypes based on matrix-assisted laser desorption ionization-time of flight mass spectrometry and machine learning techniques
abstract
BACKGROUND: Group B streptococcus (GBS) is an important pathogen that is responsible for invasive infections, including sepsis and meningitis. GBS serotyping is an essential means for the investigation of possible infection outbreaks and can identify possible sources of infection. Although it is possible to determine GBS serotypes by either immuno-serotyping or geno-serotyping, both traditional methods are time-consuming and labor-intensive. In recent years, the matrix-assisted laser desorption ionization-time of flight mass spectrometry (MALDI-TOF MS) has been reported as an effective tool for the determination of GBS serotypes in a more rapid and accurate manner. Thus, this work aims to investigate GBS serotypes by incorporating machine learning techniques with MALDI-TOF MS to carry out the identification. RESULTS: In this study, a total of 787 GBS isolates, obtained from three research and teaching hospitals, were analyzed by MALDI-TOF MS, and the serotype of the GBS was determined by a geno-serotyping experiment. The peaks of mass-to-charge ratios were regarded as the attributes to characterize the various serotypes of GBS. Machine learning algorithms, such as support vector machine (SVM) and random forest (RF), were then used to construct predictive models for the five different serotypes (Types Ia, Ib, III, V, and VI). After optimization of feature selection and model generation based on training datasets, the accuracies of the selected models attained 54.9-87.1% for various serotypes based on independent testing data. Specifically, for the major serotypes, namely type III and type VI, the accuracies were 73.9 and 70.4%, respectively. CONCLUSION: The proposed models have been adopted to implement a web-based tool (GBSTyper), which is now freely accessible at http://csb.cse.yzu.edu.tw/GBSTyper/, for providing efficient and effective detection of GBS serotypes based on a MALDI-TOF MS spectrum. Overall, this work has demonstrated that the combination of MALDI-TOF MS and machine intelligence could provide a practical means of clinical pathogen testing.
Hsin-Yao Wang, Wen-Chi Li, Kai-Yao Huang, Chia-Ru Chung, Jorng-Tzong Horng, Jen-Fu Hsu, Jang-Jih Lu, Tzong-Yi Lee
BMC Bioinform.8
2018 Genome-wide discovery of viral microRNAs based on phylogenetic analysis and structural evolution of various human papillomavirus subtypes
abstract
In mammals, microRNAs (miRNAs) play key roles in controlling posttranscriptional regulation through binding to the mRNAs of target genes. Recently, it was discovered that viral miRNAs may be involved in human cancers and diseases. It is likely that viral miRNAs help viruses enter the latent phase of their life cycle and become undetected by the host's immune system, while increasing the host's risk for cancer development. Cervical cancer is typically related to the infection of human papillomavirus (HPV) through sexual transmission. To further understand the molecular mechanisms underlying the associations of HPV infection with genital diseases, we developed a systematic method for viral miRNA identification and viral miRNA-mediated regulatory network construction based on genome-wide sequence analysis. The complete genomes of certain high-risk HPV subtypes were used to predict putative viral pre-miRNAs by bioinformatics approaches. In addition, small RNA libraries in human cervical lesions from existing publications were collected to validate the predicted HPV pre-miRNAs. For the construction of virally encoded miRNA-mediated regulatory network of HPV infection, cervical squamous epithelial carcinoma gene expression data were extracted from the RNA sequencing platform in The Cancer Genome Atlas; the differentially expressed genes were used to identify the putative targets of viral miRNAs. Predicted cellular target genes of HPV-encoded miRNAs provide an overview of these viral miRNA's putative functions. Finally, a large-scale genome analysis was carried out to examine the phylogenetic relationship and structural evolution among genital HPV types that have the potential to cause genital cancer. In this study, we discovered putative HPV-encoded miRNAs, which were validated against the small RNA libraries in human cervical lesions. Furthermore, as indicated by their biological functions, host genes targeted by HPV-encoded miRNAs may play significant roles in virus infection and carcinogenesis. These viral miRNAs pose as promising candidates for the development of antiviral drugs. More importantly, the identified subtype-specific miRNAs have the potential to be used as biomarkers for HPV subtype determination.
Shun-Long Weng, Kai-Yao Huang, Julia Tzu-Ya Weng, Fang-Yu Hung, Tzu-Hao Chang, Tzong-Yi Lee
Briefings Bioinform.6
2017 Investigation and identification of protein carbonylation sites based on position-specific amino acid composition and physicochemical features
abstract
BACKGROUND: Protein carbonylation, an irreversible and non-enzymatic post-translational modification (PTM), is often used as a marker of oxidative stress. When reactive oxygen species (ROS) oxidized the amino acid side chains, carbonyl (CO) groups are produced especially on Lysine (K), Arginine (R), Threonine (T), and Proline (P). Nevertheless, due to the lack of information about the carbonylated substrate specificity, we were encouraged to develop a systematic method for a comprehensive investigation of protein carbonylation sites. RESULTS: After the removal of redundant data from multipe carbonylation-related articles, totally 226 carbonylated proteins in human are regarded as training dataset, which consisted of 307, 126, 128, and 129 carbonylation sites for K, R, T and P residues, respectively. To identify the useful features in predicting carbonylation sites, the linear amino acid sequence was adopted not only to build up the predictive model from training dataset, but also to compare the effectiveness of prediction with other types of features including amino acid composition (AAC), amino acid pair composition (AAPC), position-specific scoring matrix (PSSM), positional weighted matrix (PWM), solvent-accessible surface area (ASA), and physicochemical properties. The investigation of position-specific amino acid composition revealed that the positively charged amino acids (K and R) are remarkably enriched surrounding the carbonylated sites, which may play a functional role in discriminating between carbonylation and non-carbonylation sites. A variety of predictive models were built using various features and three different machine learning methods. Based on the evaluation by five-fold cross-validation, the models trained with PWM feature could provide better sensitivity in the positive training dataset, while the models trained with AAindex feature achieved higher specificity in the negative training dataset. Additionally, the model trained using hybrid features, including PWM, AAC and AAindex, obtained best MCC values of 0.432, 0.472, 0.443 and 0.467 on K, R, T and P residues, respectively. CONCLUSION: When comparing to an existing prediction tool, the selected models trained with hybrid features provided a promising accuracy on an independent testing dataset. In short, this work not only characterized the carbonylated substrate preference, but also demonstrated that the proposed method could provide a feasible means for accelerating preliminary discovery of protein carbonylation.
Shun-Long Weng, Kai-Yao Huang, Fergie Joanda Kaunang, Chien-Hsun Huang, Hui-Ju Kao, Tzu-Hao Chang, Hsin-Yao Wang, Jang-Jih Lu, Tzong-Yi Lee
BMC Bioinform.9
2017 A New Scheme to Characterize and Identify Protein Ubiquitination Sites
abstract
Protein ubiquitination, involving the conjugation of ubiquitin on lysine residue, serves as an important modulator of many cellular functions in eukaryotes. Recent advancements in proteomic technology have stimulated increasing interest in identifying ubiquitination sites. However, most computational tools for predicting ubiquitination sites are focused on small-scale data. With an increasing number of experimentally verified ubiquitination sites, we were motivated to design a predictive model for identifying lysine ubiquitination sites for large-scale proteome dataset. This work assessed not only single features, such as amino acid composition (AAC), amino acid pair composition (AAPC) and evolutionary information, but also the effectiveness of incorporating two or more features into a hybrid approach to model construction. The support vector machine (SVM) was applied to generate the prediction models for ubiquitination site identification. Evaluation by five-fold cross-validation showed that the SVM models learned from the combination of hybrid features delivered a better prediction performance. Additionally, a motif discovery tool, MDDLogo, was adopted to characterize the potential substrate motifs of ubiquitination sites. The SVM models integrating the MDDLogo-identified substrate motifs could yield an average accuracy of 68.70 percent. Furthermore, the independent testing result showed that the MDDLogo-clustered SVM models could provide a promising accuracy (78.50 percent) and perform better than other prediction tools. Two cases have demonstrated the effective prediction of ubiquitination sites with corresponding substrate motifs.
Van-Nui Nguyen, Kai-Yao Huang, Chien-Hsun Huang, K. Robert Lai, Tzong-Yi Lee
IEEE ACM Trans. Comput. Biol. Bioinform.5
2016 MDD-SOH: exploiting maximal dependence decomposition to identify S-sulfenylation sites with substrate motifs
abstract
UNLABELLED: S-sulfenylation (S-sulphenylation, or sulfenic acid), the covalent attachment of S-hydroxyl (-SOH) to cysteine thiol, plays a significant role in redox regulation of protein functions. Although sulfenic acid is transient and labile, most of its physiological activities occur under control of S-hydroxylation. Therefore, discriminating the substrate site of S-sulfenylated proteins is an essential task in computational biology for the furtherance of protein structures and functions. Research into S-sulfenylated protein is currently very limited, and no dedicated tools are available for the computational identification of SOH sites. Given a total of 1096 experimentally verified S-sulfenylated proteins from humans, this study carries out a bioinformatics investigation on SOH sites based on amino acid composition and solvent-accessible surface area. A TwoSampleLogo indicates that the positively and negatively charged amino acids flanking the SOH sites may impact the formulation of S-sulfenylation in closed three-dimensional environments. In addition, the substrate motifs of SOH sites are studied using the maximal dependence decomposition (MDD). Based on the concept of binary classification between SOH and non-SOH sites, Support vector machine (SVM) is applied to learn the predictive model from MDD-identified substrate motifs. According to the evaluation results of 5-fold cross-validation, the integrated SVM model learned from substrate motifs yields an average accuracy of 0.87, significantly improving the prediction of SOH sites. Furthermore, the integrated SVM model also effectively improves the predictive performance in an independent testing set. Finally, the integrated SVM model is applied to implement an effective web resource, named MDD-SOH, to identify SOH sites with their corresponding substrate motifs. AVAILABILITY AND IMPLEMENTATION: The MDD-SOH is now freely available to all interested users at http://csb.cse.yzu.edu.tw/MDDSOH/. All of the data set used in this work is also available for download in the website. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. CONTACT: [email protected].
Van-Minh Bui, Cheng-Tsung Lu, Trang-Thi Ho, Tzong-Yi Lee
Bioinform.4
2016 Gene expression profiling identifies candidate biomarkers for active and latent tuberculosis
abstract
BACKGROUND: Tuberculosis (TB) is a serious infectious disease in that 90% of those latently infected with Mycobacterium tuberculosis present no symptoms, but possess a 10% lifetime chance of developing active TB. To prevent the spread of the disease, early diagnosis is crucial. However, current methods of detection require improvement in sensitivity, efficiency or specificity. In the present study, we conducted a microarray experiment, comparing the gene expression profiles in the peripheral blood mononuclear cells among individuals with active TB, latent infection, and healthy conditions in a Taiwanese population. RESULTS: Bioinformatics analysis revealed that most of the differentially expressed genes belonged to immune responses, inflammation pathways, and cell cycle control. Subsequent RT-PCR validation identified four differentially expressed genes, NEMF, ASUN, DHX29, and PTPRC, as potential biomarkers for the detection of active and latent TB infections. Receiver operating characteristic analysis showed that the expression level of PTPRC may discriminate active TB patients from healthy individuals, while ASUN could differentiate between the latent state of TB infection and healthy condidtion. In contrast, DHX29 may be used to identify latently infected individuals among active TB patients or healthy individuals. To test the concept of using these biomarkers as diagnostic support, we constructed classification models using these candidate biomarkers and found the Naïve Bayes-based model built with ASUN, DHX29, and PTPRC to yield the best performance. CONCLUSIONS: Our study demonstrated that gene expression profiles in the blood can be used to identify not only active TB patients, but also to differentiate latently infected patients from their healthy counterparts. Validation of the constructed computational model in a larger sample size would confirm the reliability of the biomarkers and facilitate the development of a cost-effective and sensitive molecular diagnostic platform for TB.
Shih-Wei Lee, Lawrence Shih-Hsin Wu, Guan-Mau Huang, Kai-Yao Huang, Tzong-Yi Lee, Julia Tzu-Ya Weng
BMC Bioinform.5
2015 An interpretable rule-based diagnostic classification of diabetic nephropathy among type 2 diabetes patients
abstract
BACKGROUND: The prevalence of type 2 diabetes is increasing at an alarming rate. Various complications are associated with type 2 diabetes, with diabetic nephropathy being the leading cause of renal failure among diabetics. Often, when patients are diagnosed with diabetic nephropathy, their renal functions have already been significantly damaged. Therefore, a risk prediction tool may be beneficial for the implementation of early treatment and prevention. RESULTS: In the present study, we developed a decision tree-based model integrating genetic and clinical features in a gender-specific classification for the identification of diabetic nephropathy among type 2 diabetic patients. Clinical and genotyping data were obtained from a previous genetic association study involving 345 type 2 diabetic patients (185 with diabetic nephropathy and 160 without diabetic nephropathy). Using a five-fold cross-validation approach, the performance of using clinical or genetic features alone in various classifiers (decision tree, random forest, Naïve Bayes, and support vector machine) was compared with that of utilizing a combination of attributes. The inclusion of genetic features and the implementation of an additional gender-based rule yielded better classification results. CONCLUSIONS: The current model supports the notion that genes and gender are contributing factors of diabetic nephropathy. Further refinement of the proposed approach has the potential to facilitate the early identification of diabetic nephropathy and the development of more efficient treatment in a clinical setting.
Guan-Mau Huang, Kai-Yao Huang, Tzong-Yi Lee, Julia Tzu-Ya Weng
BMC Bioinform.3
2015 ViralmiR: a support-vector-machine-based method for predicting viral microRNA precursors
abstract
BACKGROUND: microRNAs (miRNAs) play a vital role in development, oncogenesis, and apoptosis by binding to mRNAs to regulate the posttranscriptional level of coding genes in mammals, plants, and insects. Recent studies have demonstrated that the expression of viral miRNAs is associated with the ability of the virus to infect a host. Identifying potential viral miRNAs from experimental sequence data is valuable for deciphering virus-host interactions. Thus far, a specific predictive model for viral miRNA identification has yet to be developed. METHODS AND RESULTS: Here, we present ViralmiR for identifying viral miRNA precursors on the basis of sequencing and structural information. We collected 263 experimentally validated miRNA precursors (pre-miRNAs) from 26 virus species and generated sequencing fragments from virus and human genomes as the negative dataset. Support vector machine and random forest models were established using 54 features from RNA sequences and secondary structural information. The results show that ViralmiR achieved a balanced accuracy higher than 83%, which is superior to that of previously developed tools for identifying pre-miRNAs. CONCLUSIONS: The easy-to-use ViralmiR web interface has been provided as a helpful resource for researchers to use in analyzing and deciphering virus-host interactions. The web interface of ViralmiR can be accessed at http://csb.cse.yzu.edu.tw/viralmir/.
Kai-Yao Huang, Tzong-Yi Lee, Yu-Chuan Teng, Tzu-Hao Chang
BMC Bioinform.2
2015 A two-layered machine learning method to identify protein O-GlcNAcylation sites with O-GlcNAc transferase substrate motifs
abstract
Protein O-GlcNAcylation, involving the β-attachment of single N-acetylglucosamine (GlcNAc) to the hydroxyl group of serine or threonine residues, is an O-linked glycosylation catalyzed by O-GlcNAc transferase (OGT). Molecular level investigation of the basis for OGT's substrate specificity should aid understanding how O-GlcNAc contributes to diverse cellular processes. Due to an increasing number of O-GlcNAcylated peptides with site-specific information identified by mass spectrometry (MS)-based proteomics, we were motivated to characterize substrate site motifs of O-GlcNAc transferases. In this investigation, a non-redundant dataset of 410 experimentally verified O-GlcNAcylation sites were manually extracted from dbOGAP, OGlycBase and UniProtKB. After detection of conserved motifs by using maximal dependence decomposition, profile hidden Markov model (profile HMM) was adopted to learn a first-layered model for each identified OGT substrate motif. Support Vector Machine (SVM) was then used to generate a second-layered model learned from the output values of profile HMMs in first layer. The two-layered predictive model was evaluated using a five-fold cross validation which yielded a sensitivity of 85.4%, a specificity of 84.1%, and an accuracy of 84.7%. Additionally, an independent testing set from PhosphoSitePlus, which was really non-homologous to the training data of predictive model, was used to demonstrate that the proposed method could provide a promising accuracy (84.05%) and outperform other O-GlcNAcylation site prediction tools. A case study indicated that the proposed method could be a feasible means of conducting preliminary analyses of protein O-GlcNAcylation and has been implemented as a web-based system, OGTSite, which is now freely available at http://csb.cse.yzu.edu.tw/OGTSite/.
Hui-Ju Kao, Chien-Hsun Huang, Neil Arvin Bretaña, Cheng-Tsung Lu, Kai-Yao Huang, Shun-Long Weng, Tzong-Yi Lee
BMC Bioinform.7
2015 Characterization and identification of ubiquitin conjugation sites with E3 ligase recognition specificities
abstract
BACKGROUND: In eukaryotes, ubiquitin-conjugation is an important mechanism underlying proteasome-mediated degradation of proteins, and as such, plays an essential role in the regulation of many cellular processes. In the ubiquitin-proteasome pathway, E3 ligases play important roles by recognizing a specific protein substrate and catalyzing the attachment of ubiquitin to a lysine (K) residue. As more and more experimental data on ubiquitin conjugation sites become available, it becomes possible to develop prediction models that can be scaled to big data. However, no development that focuses on the investigation of ubiquitinated substrate specificities has existed. Herein, we present an approach that exploits an iteratively statistical method to identify ubiquitin conjugation sites with substrate site specificities. RESULTS: In this investigation, totally 6259 experimentally validated ubiquitinated proteins were obtained from dbPTM. After having filtered out homologous fragments with 40% sequence identity, the training data set contained 2658 ubiquitination sites (positive data) and 5532 non-ubiquitinated sites (negative data). Due to the difficulty in characterizing the substrate site specificities of E3 ligases by conventional sequence logo analysis, a recursively statistical method has been applied to obtain significant conserved motifs. The profile hidden Markov model (profile HMM) was adopted to construct the predictive models learned from the identified substrate motifs. A five-fold cross validation was then used to evaluate the predictive model, achieving sensitivity, specificity, and accuracy of 73.07%, 65.46%, and 67.93%, respectively. Additionally, an independent testing set, completely blind to the training data of the predictive model, was used to demonstrate that the proposed method could provide a promising accuracy (76.13%) and outperform other ubiquitination site prediction tool. CONCLUSION: A case study demonstrated the effectiveness of the characterized substrate motifs for identifying ubiquitination sites. The proposed method presents a practical means of preliminary analysis and greatly diminishes the total number of potential targets required for further experimental confirmation. This method may help unravel their mechanisms and roles in E3 recognition and ubiquitin-mediated protein degradation.
Van-Nui Nguyen, Kai-Yao Huang, Chien-Hsun Huang, Tzu-Hao Chang, Neil Arvin Bretaña, K. Robert Lai, Julia Tzu-Ya Weng, Tzong-Yi Lee
BMC Bioinform.8
2014 dbGSH: a database of S-glutathionylation
abstract
UNLABELLED: S-glutathionylation, the reversible protein posttranslational modification (PTM) that generates a mixed disulfide bond between glutathione and cysteine residue, critically regulates protein activity, stability and redox regulation. Due to its importance in regulating oxidative/nitrosative stress and balance in cellular response, a number of methods have been rapidly developed to study S-glutathionylation, thus expanding the dataset of experimentally determined glutathionylation sites. However, there is currently no database dedicated to the integration of all experimentally verified S-glutathionylation sites along with their characteristics or structural or functional information. Thus, the dbGSH database has been created to integrate all available datasets and to provide the relevant structural analysis. As of January 31, 2014, dbGSH has manually collected >2200 experimentally verified S-glutathionylated peptides from 169 research articles using a text-mining approach. To solve the problem of heterogeneity of the data collected from different sources, the sequence identity of the reported S-glutathionylated peptides is mapped to UniProtKB protein entries. To delineate the structural correlations and consensus motifs of these S-glutathionylation sites, the dbGSH database also provides structural and functional analyses, including the motifs of substrate sites, solvent accessibility, protein secondary and tertiary structures, protein domains and gene ontology. AVAILABILITY AND IMPLEMENTATION: dbGSH is now freely accessible at http://csb.cse.yzu.edu.tw/dbGSH/. The database content is regularly updated with new data collected by the continuous survey of research articles.
Yi-Ju Chen, Cheng-Tsung Lu, Tzong-Yi Lee, Yu-Ju Chen
Bioinform.3
2014 Characterization and identification of protein O-GlcNAcylation sites with substrate specificity
abstract
BACKGROUND: Protein O-GlcNAcylation, involving the attachment of single N-acetylglucosamine (GlcNAc) to the hydroxyl group of serine or threonine residues. Elucidation of O-GlcNAcylation sites on proteins is required in order to decipher its crucial roles in regulating cellular processes and aid in drug design. With an increasing number of O-GlcNAcylation sites identified by mass spectrometry (MS)-based proteomics, several methods have been proposed for the computational identification of O-GlcNAcylation sites. However, no development that focuses on the investigation of O-GlcNAcylated substrate motifs has existed. Thus, we were motivated to design a new method for the identification of protein O-GlcNAcylation sites with the consideration of substrate site specificity. RESULTS: In this study, 375 experimentally verified O-GlcNAcylation sites were collected from dbOGAP, which is an integrated resource for protein O-GlcNAcylation. Due to the difficulty in characterizing the substrate motifs by conventional sequence logo analysis, a recursively statistical method has been applied to obtain significant conserved motifs. To construct the predictive models learned from the identified substrate motifs, we adopted Support Vector Machines (SVMs). A five-fold cross validation was used to evaluate the predictive model, achieving sensitivity, specificity, and accuracy of 0.76, 0.80, and 0.78, respectively. Additionally, an independent testing set, which was really blind to the training data of predictive model, was used to demonstrate that the proposed method could provide a promising accuracy (0.94) and outperform three other O-GlcNAcylation site prediction tools. CONCLUSION: This work proposed a computational method to identify informative substrate motifs for O-GlcNAcylation sites. The evaluation of cross validation and independent testing indicated that the identified motifs were effective in the identification of O-GlcNAcylation sites. A case study demonstrated that the proposed method could be a feasible means of conducting preliminary analyses of protein O-GlcNAcylation. We also anticipated that the revealed substrate motif may facilitate the study of extensive crosstalk between O-GlcNAcylation and phosphorylation. This method may help unravel their mechanisms and roles in signaling, transcription, chronic disease, and cancer.
Hsin-Yi Wu, Cheng-Tsung Lu, Hui-Ju Kao, Yi-Ju Chen, Yu-Ju Chen, Tzong-Yi Lee
BMC Bioinform.6
2013 ViralPhos: incorporating a recursively statistical method to predict phosphorylation sites on virus proteins
abstract
BACKGROUND: The phosphorylation of virus proteins by host kinases is linked to viral replication. This leads to an inhibition of normal host-cell functions. Further elucidation of phosphorylation in virus proteins is required in order to aid in drug design and treatment. However, only a few studies have investigated substrate motifs in identifying virus phosphorylation sites. Additionally, existing bioinformatics tool do not consider potential host kinases that may initiate the phosphorylation of a virus protein. RESULTS: 329 experimentally verified phosphorylation fragments on 111 virus proteins were collected from virPTM. These were clustered into subgroups of significantly conserved motifs using a recursively statistical method. Two-layered Support Vector Machines (SVMs) were then applied to train a predictive model for the identified substrate motifs. The SVM models were evaluated using a five-fold cross validation which yields an average accuracy of 0.86 for serine, and 0.81 for threonine. Furthermore, the proposed method is shown to perform at par with three other phosphorylation site prediction tools: PPSP, KinasePhos 2.0 and GPS 2.1. CONCLUSION: In this study, we propose a computational method, ViralPhos, which aims to investigate virus substrate site motifs and identify potential phosphorylation sites on virus proteins. We identified informative substrate motifs that matched with several well-studied kinase groups as potential catalytic kinases for virus protein substrates. The identified substrate motifs were further exploited to identify potential virus phosphorylation sites. The proposed method is shown to be capable of predicting virus phosphorylation sites and has been implemented as a web server http://csb.cse.yzu.edu.tw/ViralPhos/.
Kai-Yao Huang, Cheng-Tsung Lu, Neil Arvin Bretaña, Tzong-Yi Lee, Tzu-Hao Chang
BMC Bioinform.4
2013 Incorporating substrate sequence motifs and spatial amino acid composition to identify kinase-specific phosphorylation sites on protein three-dimensional structures
abstract
BACKGROUND: Protein phosphorylation catalyzed by kinases plays crucial regulatory roles in cellular processes. Given the high-throughput mass spectrometry-based experiments, the desire to annotate the catalytic kinases for in vivo phosphorylation sites has motivated. Thus, a variety of computational methods have been developed for performing a large-scale prediction of kinase-specific phosphorylation sites. However, most of the proposed methods solely rely on the local amino acid sequences surrounding the phosphorylation sites. An increasing number of three-dimensional structures make it possible to physically investigate the structural environment of phosphorylation sites. RESULTS: In this work, all of the experimental phosphorylation sites are mapped to the protein entries of Protein Data Bank by sequence identity. It resulted in a total of 4508 phosphorylation sites containing the protein three-dimensional (3D) structures. To identify phosphorylation sites on protein 3D structures, this work incorporates support vector machines (SVMs) with the information of linear motifs and spatial amino acid composition, which is determined for each kinase group by calculating the relative frequencies of 20 amino acid types within a specific radial distance from central phosphorylated amino acid residue. After the cross-validation evaluation, most of the kinase-specific models trained with the consideration of structural information outperform the models considering only the sequence information. Furthermore, the independent testing set which is not included in training set has demonstrated that the proposed method could provide a comparable performance to other popular tools. CONCLUSION: The proposed method is shown to be capable of predicting kinase-specific phosphorylation sites on 3D structures and has been implemented as a web server which is freely accessible at http://csb.cse.yzu.edu.tw/PhosK3D/. Due to the difficulty of identifying the kinase-specific phosphorylation sites with similar sequenced motifs, this work also integrates the 3D structural information to improve the cross classifying specificity.
Min-Gang Su, Tzong-Yi Lee
BMC Bioinform.2
2012 dbSNO: a database of cysteine S-nitrosylation
abstract
UNLABELLED: S-nitrosylation (SNO), a selective and reversible protein post-translational modification that involves the covalent attachment of nitric oxide (NO) to the sulfur atom of cysteine, critically regulates protein activity, localization and stability. Due to its importance in regulating protein functions and cell signaling, a mass spectrometry-based proteomics method rapidly evolved to increase the dataset of experimentally determined SNO sites. However, there is currently no database dedicated to the integration of all experimentally verified S-nitrosylation sites with their structural or functional information. Thus, the dbSNO database is created to integrate all available datasets and to provide their structural analysis. Up to April 15, 2012, the dbSNO has manually accumulated >3000 experimentally verified S-nitrosylated peptides from 219 research articles using a text mining approach. To solve the heterogeneity among the data collected from different sources, the sequence identity of these reported S-nitrosylated peptides are mapped to the UniProtKB protein entries. To delineate the structural correlation and consensus motif of these SNO sites, the dbSNO database also provides structural and functional analyses, including the motifs of substrate sites, solvent accessibility, protein secondary and tertiary structures, protein domains and gene ontology. AVAILABILITY: The dbSNO is now freely accessible via http://dbSNO.mbc.nctu.edu.tw. The database content is regularly updated upon collecting new data obtained from continuously surveying research articles.
Tzong-Yi Lee, Yi-Ju Chen, Cheng-Tsung Lu, Wei-Chieh Ching, Yu-Chuan Teng, Hsien-Da Huang, Yu-Ju Chen
Bioinform.1
2011 Prediction of transporter targets using efficient RBF networks with PSSM profiles and biochemical properties
abstract
SUMMARY: Transporters are proteins that are involved in the movement of ions or molecules across biological membranes. Currently, our knowledge about the functions of transporters is limited due to the paucity of their 3D structures. Hence, computational techniques are necessary to annotate the functions of transporters. In this work, we focused on an important functional aspect of transporters, namely annotation of targets for transport proteins. We have systematically analyzed four major classes of transporters with different transporter targets: (i) electron, (ii) protein/mRNA, (iii) ion and (iv) others, using amino acid properties. We have developed a radial basis function network-based method for predicting transport targets with amino acid properties and position specific scoring matrix profiles. Our method showed a 10-fold cross-validation accuracy of 90.1, 80.1, 70.3 and 82.3% for electron transporters, protein/mRNA transporters, ion transporters and others, respectively, in a dataset of 543 transporters. We have also evaluated the performance of the method with an independent dataset of 108 proteins and we obtained similar accuracy. We suggest that our method could be an effective tool for functional annotation of transport proteins. AVAILABILITY: http://rbf.bioinfo.tw/~sachen/ttrbf.html
Shu-An Chen, Yu-Yen Ou, Tzong-Yi Lee, M. Michael Gromiha
Bioinform.3
2011 Exploiting maximal dependence decomposition to identify conserved motifs from a group of aligned signal sequences
abstract
UNLABELLED: Bioinformatics research often requires conservative analyses of a group of sequences associated with a specific biological function (e.g. transcription factor binding sites, micro RNA target sites or protein post-translational modification sites). Due to the difficulty in exploring conserved motifs on a large-scale sequence data involved with various signals, a new method, MDDLogo, is developed. MDDLogo applies maximal dependence decomposition (MDD) to cluster a group of aligned signal sequences into subgroups containing statistically significant motifs. In order to extract motifs that contain a conserved biochemical property of amino acids in protein sequences, the set of 20 amino acids is further categorized according to their physicochemical properties, e.g. hydrophobicity, charge or molecular size. MDDLogo has been demonstrated to accurately identify the kinase-specific substrate motifs in 1221 human phosphorylation sites associated with seven well-known kinase families from Phospho.ELM. Moreover, in a set of plant phosphorylation data-lacking kinase information, MDDLogo has been applied to help in the investigation of substrate motifs of potential kinases and in the improvement of the identification of plant phosphorylation sites with various substrate specificities. In this study, MDDLogo is comparable with another well-known motif discover tool, Motif-X. CONTACT: [email protected]
Tzong-Yi Lee, Zong-Qing Lin, Neil Arvin Bretaña, Cheng-Tsung Lu
Bioinform.1
2011 miRTar: an integrated system for identifying miRNA-target interactions in Human
abstract
BACKGROUND: MicroRNAs (miRNAs) are small non-coding RNA molecules that are ~22-nt-long sequences capable of suppressing protein synthesis. Previous research has suggested that miRNAs regulate 30% or more of the human protein-coding genes. The aim of this work is to consider various analyzing scenarios in the identification of miRNA-target interactions, as well as to provide an integrated system that will aid in facilitating investigation on the influence of miRNA targets by alternative splicing and the biological function of miRNAs in biological pathways. RESULTS: This work presents an integrated system, miRTar, which adopts various analyzing scenarios to identify putative miRNA target sites of the gene transcripts and elucidates the biological functions of miRNAs toward their targets in biological pathways. The system has three major features. First, the prediction system is able to consider various analyzing scenarios (1 miRNA:1 gene, 1:N, N:1, N:M, all miRNAs:N genes, and N miRNAs: genes involved in a pathway) to easily identify the regulatory relationships between interesting miRNAs and their targets, in 3'UTR, 5'UTR and coding regions. Second, miRTar can analyze and highlight a group of miRNA-regulated genes that participate in particular KEGG pathways to elucidate the biological roles of miRNAs in biological pathways. Third, miRTar can provide further information for elucidating the miRNA regulation, i.e., miRNA-target interactions, affected by alternative splicing. CONCLUSIONS: In this work, we developed an integrated resource, miRTar, to enable biologists to easily identify the biological functions and regulatory relationships between a group of known/putative miRNAs and protein coding genes. miRTar is now available at http://miRTar.mbc.nctu.edu.tw/.
Justin Bo-Kai Hsu, Chih-Min Chiu, Sheng-Da Hsu, Wei-Yun Huang, Chia-Hung Chien, Tzong-Yi Lee, Hsien-Da Huang
BMC Bioinform.6
2011 PlantPhos: using Maximal Dependence Decomposition to Identify Plant Phosphorylation Sites with Substrate Site Specificity
abstract
BACKGROUND: Protein phosphorylation catalyzed by kinases plays crucial regulatory roles in intracellular signal transduction. Due to the difficulty in performing high-throughput mass spectrometry-based experiment, there is a desire to predict phosphorylation sites using computational methods. However, previous studies regarding in silico prediction of plant phosphorylation sites lack the consideration of kinase-specific phosphorylation data. Thus, we are motivated to propose a new method that investigates different substrate specificities in plant phosphorylation sites. RESULTS: Experimentally verified phosphorylation data were extracted from TAIR9-a protein database containing 3006 phosphorylation data from the plant species Arabidopsis thaliana. In an attempt to investigate the various substrate motifs in plant phosphorylation, maximal dependence decomposition (MDD) is employed to cluster a large set of phosphorylation data into subgroups containing significantly conserved motifs. Profile hidden Markov model (HMM) is then applied to learn a predictive model for each subgroup. Cross-validation evaluation on the MDD-clustered HMMs yields an average accuracy of 82.4% for serine, 78.6% for threonine, and 89.0% for tyrosine models. Moreover, independent test results using Arabidopsis thaliana phosphorylation data from UniProtKB/Swiss-Prot show that the proposed models are able to correctly predict 81.4% phosphoserine, 77.1% phosphothreonine, and 83.7% phosphotyrosine sites. Interestingly, several MDD-clustered subgroups are observed to have similar amino acid conservation with the substrate motifs of well-known kinases from Phospho.ELM-a database containing kinase-specific phosphorylation data from multiple organisms. CONCLUSIONS: This work presents a novel method for identifying plant phosphorylation sites with various substrate motifs. Based on cross-validation and independent testing, results show that the MDD-clustered models outperform models trained without using MDD. The proposed method has been implemented as a web-based plant phosphorylation prediction tool, PlantPhos http://csb.cse.yzu.edu.tw/PlantPhos/. Additionally, two case studies have been demonstrated to further evaluate the effectiveness of PlantPhos.
Tzong-Yi Lee, Neil Arvin Bretaña, Cheng-Tsung Lu
BMC Bioinform.1
2011 Investigation and identification of protein γ-glutamyl carboxylation sites
abstract
BACKGROUND: Carboxylation is a modification of glutamate (Glu) residues which occurs post-translation that is catalyzed by γ-glutamyl carboxylase in the lumen of the endoplasmic reticulum. Vitamin K is a critical co-factor in the post-translational conversion of Glu residues to γ-carboxyglutamate (Gla) residues. It has been shown that the process of carboxylation is involved in the blood clotting cascade, bone growth, and extraosseous calcification. However, studies in this field have been limited by the difficulty of experimentally studying substrate site specificity in γ-glutamyl carboxylation. In silico investigations have the potential for characterizing carboxylated sites before experiments are carried out. RESULTS: Because of the importance of γ-glutamyl carboxylation in biological mechanisms, this study investigates the substrate site specificity in carboxylation sites. It considers not only the composition of amino acids that surround carboxylation sites, but also the structural characteristics of these sites, including secondary structure and solvent-accessible surface area (ASA). The explored features are used to establish a predictive model for differentiating between carboxylation sites and non-carboxylation sites. A support vector machine (SVM) is employed to establish a predictive model with various features. A five-fold cross-validation evaluation reveals that the SVM model, trained with the combined features of positional weighted matrix (PWM), amino acid composition (AAC), and ASA, yields the highest accuracy (0.892). Furthermore, an independent testing set is constructed to evaluate whether the predictive model is over-fitted to the training set. CONCLUSIONS: Independent testing data that did not undergo the cross-validation process shows that the proposed model can differentiate between carboxylation sites and non-carboxylation sites. This investigation is the first to study carboxylation sites and to develop a system for identifying them. The proposed method is a practical means of preliminary analysis and greatly diminishes the total number of potential carboxylation sites requiring further experimental confirmation.
Tzong-Yi Lee, Cheng-Tsung Lu, Shu-An Chen, Neil Arvin Bretaña, Tzu-Hsiu Cheng, Min-Gang Su, Kai-Yao Huang
BMC Bioinform.1
2010 Incorporating significant amino acid pairs to identify O-linked glycosylation sites on transmembrane proteins and non-transmembrane proteins
abstract
BACKGROUND: While occurring enzymatically in biological systems, O-linked glycosylation affects protein folding, localization and trafficking, protein solubility, antigenicity, biological activity, as well as cell-cell interactions on membrane proteins. Catalytic enzymes involve glycotransferases, sugar-transferring enzymes and glycosidases which trim specific monosaccharides from precursors to form intermediate structures. Due to the difficulty of experimental identification, several works have used computational methods to identify glycosylation sites. RESULTS: By investigating glycosylated sites that contain various motifs between Transmembrane (TM) and non-Transmembrane (non-TM) proteins, this work presents a novel method, GlycoRBF, that implements radial basis function (RBF) networks with significant amino acid pairs (SAAPs) for identifying O-linked glycosylated serine and threonine on TM proteins and non-TM proteins. Additionally, a membrane topology is considered for reducing the false positives on glycosylated TM proteins. Based on an evaluation using five-fold cross-validation, the consideration of a membrane topology can reduce 31.4% of the false positives when identifying O-linked glycosylation sites on TM proteins. Via an independent test, GlycoRBF outperforms previous O-linked glycosylation site prediction schemes. CONCLUSION: A case study of Cyclic AMP-dependent transcription factor ATF-6 alpha was presented to demonstrate the effectiveness of GlycoRBF. Web-based GlycoRBF, which can be accessed at http://GlycoRBF.bioinfo.tw, can identify O-linked glycosylated serine and threonine effectively and efficiently. Moreover, the structural topology of Transmembrane (TM) proteins with glycosylation sites is provided to users. The stand-alone version of GlycoRBF is also available for high throughput data analysis.
Shu-An Chen, Tzong-Yi Lee, Yu-Yen Ou
BMC Bioinform.2