Yu Xue 0001

dblp:05/6904-1 · DBLP profile ↗
← Back
14ranked-venue papers
3as first author
4since 2021 · last 2025
0000-0002-9403-6869ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 14 · 3 first-author · 4 since 2021
YearPublicationVenuePosition
2025 GPSD: a hybrid learning framework for the prediction of phosphatase-specific dephosphorylation sites
abstract
Protein phosphorylation is dynamically and reversibly regulated by protein kinases and protein phosphatases, and plays an essential role in orchestrating a wide range of biological processes. Although a number of tools have been developed for predicting kinase-specific phosphorylation sites (p-sites), computational prediction of phosphatase-specific dephosphorylation sites remains to be a great challenge. In this study, we manually curated 4393 experimentally identified site-specific phosphatase-substrate relationships for 3463 dephosphorylation sites occurring on phosphoserine, phosphothreonine, and/or phosphotyrosine residues, from the literature and public databases. Then, we developed a hybrid learning framework, the group-based prediction system for the prediction of phosphatase-specific dephosphorylation sites (GPSD). For model training, we integrated 10 types of sequence features and utilized three types of machine learning methods, including penalized logistic regression, deep neural networks, and transformer neural networks. First, a pretrained model was constructed using 561 416 nonredundant p-sites and then fine-tuned to generate computational models for predicting general dephosphorylation sites. In addition, 103 individual phosphatase-specific predictors were constructed via transfer learning and meta-learning. For site prediction, one or multiple protein sequences in FASTA format could be inputted, and the prediction results will be shown together with additional annotations, such as protein-protein interactions, structural information, and disorder propensity. The online service of GPSD is freely available at https://gpsd.biocuckoo.cn/. We believe that GPSD can serve as a valuable tool for further analysis of dephosphorylation.
Shanshan Fu, Yujie Gou, Chi Zhang 0080, Xinhe Huang, Leming Xiao, Miaoying Zhao, Yu Xue 0001
Briefings Bioinform.13
2022 GPS-Uber: a hybrid-learning framework for prediction of general and E3-specific lysine ubiquitination sites
abstract
As an important post-translational modification, lysine ubiquitination participates in numerous biological processes and is involved in human diseases, whereas the site specificity of ubiquitination is mainly decided by ubiquitin-protein ligases (E3s). Although numerous ubiquitination predictors have been developed, computational prediction of E3-specific ubiquitination sites is still a great challenge. Here, we carefully reviewed the existing tools for the prediction of general ubiquitination sites. Also, we developed a tool named GPS-Uber for the prediction of general and E3-specific ubiquitination sites. From the literature, we manually collected 1311 experimentally identified site-specific E3-substrate relations, which were classified into different clusters based on corresponding E3s at different levels. To predict general ubiquitination sites, we integrated 10 types of sequence and structure features, as well as three types of algorithms including penalized logistic regression, deep neural network and convolutional neural network. Compared with other existing tools, the general model in GPS-Uber exhibited a highly competitive accuracy, with an area under curve values of 0.7649. Then, transfer learning was adopted for each E3 cluster to construct E3-specific models, and in total 112 individual E3-specific predictors were implemented. Using GPS-Uber, we conducted a systematic prediction of human cancer-associated ubiquitination events, which could be helpful for further experimental consideration. GPS-Uber will be regularly updated, and its online service is free for academic research at http://gpsuber.biocuckoo.cn/.
Xiaodan Tan, Dachao Tang, Yujie Gou, Wanshan Ning, Shaofeng Lin, Weizhi Zhang 0002, Yu Xue 0001
Briefings Bioinform.11
2021 EPSD: a well-annotated data resource of protein phosphorylation sites in eukaryotes
abstract
As an important post-translational modification (PTM), protein phosphorylation is involved in the regulation of almost all of biological processes in eukaryotes. Due to the rapid progress in mass spectrometry-based phosphoproteomics, a large number of phosphorylation sites (p-sites) have been characterized but remain to be curated. Here, we briefly summarized the current progresses in the development of data resources for the collection, curation, integration and annotation of p-sites in eukaryotic proteins. Also, we designed the eukaryotic phosphorylation site database (EPSD), which contained 1 616 804 experimentally identified p-sites in 209 326 phosphoproteins from 68 eukaryotic species. In EPSD, we not only collected 1 451 629 newly identified p-sites from high-throughput (HTP) phosphoproteomic studies, but also integrated known p-sites from 13 additional databases. Moreover, we carefully annotated the phosphoproteins and p-sites of eight model organisms by integrating the knowledge from 100 additional resources that covered 15 aspects, including phosphorylation regulator, genetic variation and mutation, functional annotation, structural annotation, physicochemical property, functional domain, disease-associated information, protein-protein interaction, drug-target relation, orthologous information, biological pathway, transcriptional regulator, mRNA expression, protein expression/proteomics and subcellular localization. We anticipate that the EPSD can serve as a useful resource for further analysis of eukaryotic phosphorylation. With a data volume of 14.1 GB, EPSD is free for all users at http://epsd.biocuckoo.cn/.
Shaofeng Lin, Jiaqi Zhou 0003, Chen Ruan, Yiran Tu, Lan Yao, Yu Xue 0001
Briefings Bioinform.9
2021 GPS-Palm: a deep learning-based graphic presentation system for the prediction of S-palmitoylation sites in proteins
abstract
As an important reversible lipid modification, S-palmitoylation mainly occurs at specific cysteine residues in proteins, participates in regulating various biological processes and is associated with human diseases. Besides experimental assays, computational prediction of S-palmitoylation sites can efficiently generate helpful candidates for further experimental consideration. Here, we reviewed the current progress in the development of S-palmitoylation site predictors, as well as training data sets, informative features and algorithms used in these tools. Then, we compiled a benchmark data set containing 3098 known S-palmitoylation sites identified from small- or large-scale experiments, and developed a new method named data quality discrimination (DQD) to distinguish data quality weights (DQWs) between the two types of the sites. Besides DQD and our previous methods, we encoded sequence similarity values into images, constructed a deep learning framework of convolutional neural networks (CNNs) and developed a novel algorithm of graphic presentation system (GPS) 6.0. We further integrated nine additional types of sequence-based and structural features, implemented parallel CNNs (pCNNs) and designed a new predictor called GPS-Palm. Compared with other existing tools, GPS-Palm showed a >31.3% improvement of the area under the curve (AUC) value (0.855 versus 0.651) for general prediction of S-palmitoylation sites. We also produced two species-specific predictors, with corresponding AUC values of 0.900 and 0.897 for predicting human- and mouse-specific sites, respectively. GPS-Palm is free for academic research at http://gpspalm.biocuckoo.cn/.
Wanshan Ning, Peiran Jiang, Yaping Guo, Xiaodan Tan, Weizhi Zhang 0002, Yu Xue 0001
Briefings Bioinform.8
2017 Computational prediction of methylation types of covalently modified lysine and arginine residues in proteins
abstract
Protein methylation is an essential posttranslational modification (PTM) mostly occurs at lysine and arginine residues, and regulates a variety of cellular processes. Owing to the rapid progresses in the large-scale identification of methylation sites, the available data set was dramatically expanded, and more attention has been paid on the identification of specific methylation types of modification residues. Here, we briefly summarized the current progresses in computational prediction of methylation sites, which provided an accurate, rapid and efficient approach in contrast with labor-intensive experiments. We collected 5421 methyllysines and methylarginines in 2592 proteins from the literature, and classified most of the sites into different types. Data analyses demonstrated that different types of methylated proteins were preferentially involved in different biological processes and pathways, whereas a unique sequence preference was observed for each type of methylation sites. Thus, we developed a predictor of GPS-MSP, which can predict mono-, di- and tri-methylation types for specific lysines, and mono-, symmetric di- and asymmetrical di-methylation types for specific arginines. We critically evaluated the performance of GPS-MSP, and compared it with other existing tools. The satisfying results exhibited that the classification of methylation sites into different types for training can considerably improve the prediction accuracy. Taken together, we anticipate that our study provides a new lead for future computational analysis of protein methylation, and the prediction of methylation types of covalently modified lysine and arginine residues can generate more useful information for further experimental manipulation.
Wankun Deng, Ying Zhang 0052, Yu Xue 0001
Briefings Bioinform.6
2015 IBS: an illustrator for the presentation and visualization of biological sequences
abstract
UNLABELLED: Biological sequence diagrams are fundamental for visualizing various functional elements in protein or nucleotide sequences that enable a summarization and presentation of existing information as well as means of intuitive new discoveries. Here, we present a software package called illustrator of biological sequences (IBS) that can be used for representing the organization of either protein or nucleotide sequences in a convenient, efficient and precise manner. Multiple options are provided in IBS, and biological sequences can be manipulated, recolored or rescaled in a user-defined mode. Also, the final representational artwork can be directly exported into a publication-quality figure. AVAILABILITY AND IMPLEMENTATION: The standalone package of IBS was implemented in JAVA, while the online service was implemented in HTML5 and JavaScript. Both the standalone package and online service are freely available at http://ibs.biocuckoo.org. CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Wenzhong Liu, Yubin Xie, Jiyong Ma, Xiaotong Luo, Zhixiang Zuo, Urs Lahrmann, Qi Zhao 0009, Yueyuan Zheng, Yong Zhao 0013, Yu Xue 0001, Jian Ren 0002
Bioinform.11
2013 Systematic analysis of the Plk-mediated phosphoregulation in eukaryotes
abstract
Substantial evidence has confirmed that Polo-like kinases (Plks) play a crucial role in a variety of cellular processes via phosphorylation-mediated signaling transduction. Identification of Plk phospho-binding proteins and phosphorylation substrates is fundamental for elucidating the molecular mechanisms of Plks. Here, we present an integrative approach for the analysis of Plk-specific phospho-binding and phosphorylation sites (p-sites) in proteins. From the currently available phosphoproteomic data, we predicted tens of thousands of potential Plk phospho-binding and phosphorylation sites in eukaryotes, respectively. Furthermore, statistical analysis suggested that Plk phospho-binding proteins are more closely implicated in mitosis than their phosphorylation substrates. Additional computational analysis together with in vitro and in vivo experimental assays demonstrated that human Mis18B is a novel interacting partner of Plk1, while pT14 and pS48 of Mis18B were identified as phospho-binding sites. Taken together, this systematic analysis provides a global landscape of the complexity and diversity of potential Plk-mediated phosphoregulation, and the prediction results can be helpful for further experimental investigation.
Zexian Liu, Jian Ren 0002, Xuebiao Yao, Changjiang Jin, Yu Xue 0001
Briefings Bioinform.7
2012 CPSS: a computational platform for the analysis of small RNA deep sequencing data
abstract
UNLABELLED: Next generation sequencing (NGS) techniques have been widely used to document the small ribonucleic acids (RNAs) implicated in a variety of biological, physiological and pathological processes. An integrated computational tool is needed for handling and analysing the enormous datasets from small RNA deep sequencing approach. Herein, we present a novel web server, CPSS (a computational platform for the analysis of small RNA deep sequencing data), designed to completely annotate and functionally analyse microRNAs (miRNAs) from NGS data on one platform with a single data submission. Small RNA NGS data can be submitted to this server with analysis results being returned in two parts: (i) annotation analysis, which provides the most comprehensive analysis for small RNA transcriptome, including length distribution and genome mapping of sequencing reads, small RNA quantification, prediction of novel miRNAs, identification of differentially expressed miRNAs, piwi-interacting RNAs and other non-coding small RNAs between paired samples and detection of miRNA editing and modifications and (ii) functional analysis, including prediction of miRNA targeted genes by multiple tools, enrichment of gene ontology terms, signalling pathway involvement and protein-protein interaction analysis for the predicted genes. CPSS, a ready-to-use web server that integrates most functions of currently available bioinformatics tools, provides all the information wanted by the majority of users from small RNA deep sequencing datasets. AVAILABILITY: CPSS is implemented in PHP/PERL+MySQL+R and can be freely accessed at http://mcg.ustc.edu.cn/db/cpss/index.html or http://mcg.ustc.edu.cn/sdap1/cpss/index.html.
Yuanwei Zhang, Yifan Yang 0001, Rongjun Ban, Xiaohua Jiang, Howard J. Cooke, Yu Xue 0001, Qinghua Shi
Bioinform.8
2011 Prediction of novel pre-microRNAs with high accuracy through boosting and SVM
abstract
UNLABELLED: High-throughput deep-sequencing technology has generated an unprecedented number of expressed short sequence reads, presenting not only an opportunity but also a challenge for prediction of novel microRNAs. To verify the existence of candidate microRNAs, we have to show that these short sequences can be processed from candidate pre-microRNAs. However, it is laborious and time consuming to verify these using existing experimental techniques. Therefore, here, we describe a new method, miRD, which is constructed using two feature selection strategies based on support vector machines (SVMs) and boosting method. It is a high-efficiency tool for novel pre-microRNA prediction with accuracy up to 94.0% among different species. AVAILABILITY: miRD is implemented in PHP/PERL+MySQL+R and can be freely accessed at http://mcg.ustc.edu.cn/rpg/mird/mird.php.
Yuanwei Zhang, Yifan Yang 0001, Xiaohua Jiang, Yu Xue 0001, Yunxia Cao, Qian Zhai, Yong Zhai, Mingqing Xu, Howard J. Cooke, Qinghua Shi
Bioinform.6
2008 Proteome-Wide Analysis of Amino Acid Absence in Composition and Plasticity
YuZhong Zhao, Changjiang Jin, Xinjiao Gao, Yu Xue 0001, Xuebiao Yao
ICIC (1)6
2006 HSPPIP: An Online Tool for Prediction of Protein-Protein Interactions in Humans
Yu Xue 0001, Changjiang Jin, Xuebiao Yao
ICIC (3)1
2006 CSS-Palm: palmitoylation site prediction with a clustering and scoring strategy (CSS)
abstract
UNLABELLED: Palmitoylation is an important post-translational lipid modification of proteins. Unlike prenylation and myristoylation, palmitoylation is a reversible covalent modification, allowing for dynamic regulation of multiple complex cellular systems. However, in vivo or in vitro identification of palmitoylation sites is usually time-consuming and labor-intensive. So in silico predictions could help to narrow down the possible palmitoylation sites, which can be used to guide further experimental design. Previous studies suggested that there is no unique canonical motif for palmitoylation sites, so we hypothesize that the bona fide pattern might be compromised by heterogeneity of multiple structural determinants with different features. Based on this hypothesis, we partition the known palmitoylation sites into three clusters and score the similarity between the query peptide and the training ones based on BLOSUM62 matrix. We have implemented a computer program for palmitoylation site prediction, Clustering and Scoring Strategy for Palmitoylation Sites Prediction (CSS-Palm) system, and found that the program's prediction performance is encouraging with highly positive Jack-Knife validation results (sensitivity 82.16% and specificity 83.17% for cut-off score 2.6). Our analyses indicate that CSS-Palm could provide a powerful and effective tool to studies of palmitoylation sites. AVAILABILITY: CSS-Palm is implemented in PHP/PERL+MySQL and can be freely accessed at http://bioinformatics.lcd-ustc.org/css_palm/ CONTACT: [email protected]; [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bionformatics online.
Fengfeng Zhou, Yu Xue 0001, Xuebiao Yao, Ying Xu 0001
Bioinform.2
2006 NBA-Palm: prediction of palmitoylation site implemented in Naïve Bayes algorithm
abstract
BACKGROUND: Protein palmitoylation, an essential and reversible post-translational modification (PTM), has been implicated in cellular dynamics and plasticity. Although numerous experimental studies have been performed to explore the molecular mechanisms underlying palmitoylation processes, the intrinsic feature of substrate specificity has remained elusive. Thus, computational approaches for palmitoylation prediction are much desirable for further experimental design. RESULTS: In this work, we present NBA-Palm, a novel computational method based on Naïve Bayes algorithm for prediction of palmitoylation site. The training data is curated from scientific literature (PubMed) and includes 245 palmitoylated sites from 105 distinct proteins after redundancy elimination. The proper window length for a potential palmitoylated peptide is optimized as six. To evaluate the prediction performance of NBA-Palm, 3-fold cross-validation, 8-fold cross-validation and Jack-Knife validation have been carried out. Prediction accuracies reach 85.79% for 3-fold cross-validation, 86.72% for 8-fold cross-validation and 86.74% for Jack-Knife validation. Two more algorithms, RBF network and support vector machine (SVM), also have been employed and compared with NBA-Palm. CONCLUSION: Taken together, our analyses demonstrate that NBA-Palm is a useful computational program that provides insights for further experimentation. The accuracy of NBA-Palm is comparable with our previously described tool CSS-Palm. The NBA-Palm is freely accessible from: http://www.bioinfo.tsinghua.edu.cn/NBA-Palm.
Yu Xue 0001, Changjiang Jin, Zhirong Sun, Xuebiao Yao
BMC Bioinform.1
2006 PPSP: prediction of PK-specific phosphorylation site with Bayesian decision theory
abstract
BACKGROUND: As a reversible and dynamic post-translational modification (PTM) of proteins, phosphorylation plays essential regulatory roles in a broad spectrum of the biological processes. Although many studies have been contributed on the molecular mechanism of phosphorylation dynamics, the intrinsic feature of substrates specificity is still elusive and remains to be delineated. RESULTS: In this work, we present a novel, versatile and comprehensive program, PPSP (Prediction of PK-specific Phosphorylation site), deployed with approach of Bayesian decision theory (BDT). PPSP could predict the potential phosphorylation sites accurately for approximately 70 PK (Protein Kinase) groups. Compared with four existing tools Scansite, NetPhosK, KinasePhos and GPS, PPSP is more accurate and powerful than these tools. Moreover, PPSP also provides the prediction for many novel PKs, say, TRK, mTOR, SyK and MET/RON, etc. The accuracy of these novel PKs are also satisfying. CONCLUSION: Taken together, we propose that PPSP could be a potentially powerful tool for the experimentalists who are focusing on phosphorylation substrates with their PK-specific sites identification. Moreover, the BDT strategy could also be a ubiquitous approach for PTMs, such as sumoylation and ubiquitination, etc.
Yu Xue 0001, Ao Li 0001, Lirong Wang, Huanqing Feng, Xuebiao Yao
BMC Bioinform.1