VLDB 2026 Research / reviewers in the wild / expert
Vladimir N. Uversky
dblp:38/312
· DBLP profile ↗
17ranked-venue papers
1as first author
5since 2021 · last 2023
0000-0002-4037-5857ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 17 · 1 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | TFBSnet: A deep learning-based tool for predicting transcription factor binding site from DNA sequencesabstractTranscription factors (TFs) are crucial proteins that regulate gene transcription by binding to specific sites on DNA, known as transcription factor binding sites (TFBSs). Identifying TFBSs enables the design of drugs to modulate gene expression, making it important for drug design and gene therapy. While deep learning-based methods have been proposed for predicting TFBSs, there is room for improvement. This study introduces TFBSnet, a novel deep learning-based technique that accurately predicts TFBSs by extracting diverse feature data from DNA sequences and utilizing a convolutional neural network (CNN) combined with SKNet. Experimental results show that TFBSnet outperforms existing methods. It also demonstrates accurate prediction of TF binding sites in human cells without label data and exceptional performance in predicting TFBSs in plant cells using 265 TFs in Arabidopsis. Ablation analysis highlights the integration of different features and advanced feature extraction by SKNet as contributors to TFBSnet's superior predictive capability. Zhihua Du, Tianyou Huang, Jianqiang Li 0001, Vladimir N. Uversky |
BIBM | 4 |
| 2023 | Predicting TF Proteins by Incorporating Evolution Information Through PSSMabstractTranscription factors (TFs) are DNA binding proteins involved in the regulation of gene expression. They exist in all organisms and activate or repress transcription by binding to specific DNA sequences. Traditionally, TFs have been identified by experimental methods that are time-consuming and costly. In recent years, various computational methods have been developed to identify TF to overcome these limitations. However, there is a room for further improvement in the predictive performance of these tools in terms of accuracy. We report here a novel computational tool, TFnet, that provides accurate and comprehensive TF predictions from protein sequences. The accuracy of these predictions is substantially better than the results of the existing TF predictors and methods. Especially, it outperforms comparable methods significantly when sequence similarity to other known sequences in the database drops below 40%. Ablation tests reveal that the high predictive performance stems from innovative ways used in TFnet to derive sequence Position-Specific Scoring Matrix (PSSM) and encode inputs. Zhihua Du, Tianyou Huang, Vladimir N. Uversky, Jianqiang Li 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2022 | DeepCLD: An Efficient Sequence-Based Predictor of Intrinsically Disordered ProteinsabstractIntrinsic disorder is common in proteins, plays important roles in protein functionality, and is commonly associated with various human diseases. To have an accurate tool for the annotation of intrinsic disorder in proteins, this paper proposes a novel algorithm, DeepCLD, for sequence-based prediction of intrinsically disordered proteins. This algorithm uses amino acid position specific scoring matrix (PSSM) to capture the intrinsic variability characteristic of sequence patterns, ResNet to preserve feature space structure, and bidirectional CudnnLSTM as recurrent layer to further improve the efficiency. Futhermore, DeepCLD also utilized the attention mechanism to solve the problem of gradient disappearing in deep network. Comparative analyses show that DeepCLD has faster training speed and higher prediction accuracy than comparable methods. Yufeng He, Zhihua Du, Vladimir N. Uversky |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2021 | Understanding structural malleability of the SARS-CoV-2 proteins and relation to the comorbiditiesabstractSevere acute respiratory syndrome coronavirus 2 (SARS-CoV-2), a causative agent of the coronavirus disease (COVID-19), is a part of the $\beta $-Coronaviridae family. The virus contains five major protein classes viz., four structural proteins [nucleocapsid (N), membrane (M), envelop (E) and spike glycoprotein (S)] and replicase polyproteins (R), synthesized as two polyproteins (ORF1a and ORF1ab). Due to the severity of the pandemic, most of the SARS-CoV-2-related research are focused on finding therapeutic solutions. However, studies on the sequences and structure space throughout the evolutionary time frame of viral proteins are limited. Besides, the structural malleability of viral proteins can be directly or indirectly associated with the dysfunctionality of the host cell proteins. This dysfunctionality may lead to comorbidities during the infection and may continue at the post-infection stage. In this regard, we conduct the evolutionary sequence-structure analysis of the viral proteins to evaluate their malleability. Subsequently, intrinsic disorder propensities of these viral proteins have been studied to confirm that the short intrinsically disordered regions play an important role in enhancing the likelihood of the host proteins interacting with the viral proteins. These interactions may result in molecular dysfunctionality, finally leading to different diseases. Based on the host cell proteins, the diseases are divided in two distinct classes: (i) proteins, directly associated with the set of diseases while showing similar activities, and (ii) cytokine storm-mediated pro-inflammation (e.g. acute respiratory distress syndrome, malignancies) and neuroinflammation (e.g. neurodegenerative and neuropsychiatric diseases). Finally, the study unveils that males and postmenopausal females can be more vulnerable to SARS-CoV-2 infection due to the androgen-mediated protein transmembrane serine protease 2. Sagnik Sen 0002, Ashmita Dey, Sanghamitra Bandhyopadhyay, Vladimir N. Uversky, Ujjwal Maulik |
Briefings Bioinform. | 4 |
| 2021 | Insights into the evolutionary forces that shape the codon usage in the viral genome segments encoding intrinsically disordered protein regionsabstractIntrinsically disordered regions/proteins (IDRs) are abundant across all the domains of life, where they perform important regulatory roles and supplement the biological functions of structured proteins/regions (SRs). Despite the multifunctionality features of IDRs, several interrogations on the evolution of viral genomic regions encoding IDRs in diverse viral proteins remain unreciprocated. To fill this gap, we benchmarked the findings of two most widely used and reliable intrinsic disorder prediction algorithms (IUPred2A and ESpritz) to a dataset of 6108 reference viral proteomes to unravel the multifaceted evolutionary forces that shape the codon usage in the viral genomic regions encoding for IDRs and SRs. We found persuasive evidence that the natural selection predominantly governs the evolution of codon usage in regions encoding IDRs by most of the viruses. In addition, we confirm not only that codon usage in regions encoding IDRs is less optimized for the protein synthesis machinery (transfer RNAs pool) of their host than for those encoding SRs, but also that the selective constraints imposed by codon bias sustain this reduced optimization in IDRs. Our analysis also establishes that IDRs in viruses are likely to tolerate more translational errors than SRs. All these findings hold true, irrespective of the disorder prediction algorithms used to classify IDRs. In conclusion, our study offers a novel perspective on the evolution of viral IDRs and the evolutionary adaptability to multiple taxonomically divergent hosts. Rahul Kaushik, Chandana Tennakoon, Vladimir N. Uversky, Sonia Longhi, Kam Y. J. Zhang, Sandeep Bhatia |
Briefings Bioinform. | 4 |
| 2015 | Erratum to: Improving protein order-disorder classification using charge-hydropathy plotsabstractDuring the production of our manuscript [1], an incorrect Figure Figure11 was put in place of the one we submitted, and we erred in not appropriately informing the production staff of this mistake. The corrected figure and the text describing this figure, as well as the figure legend, are present herein.
Figure 1.
Charge-Hydropathy plots. In (A) the IDP-Hydropathy scale was used, in (B) the Guy (1985) Hydropathy scale was used, and in (C) the Kyte-Doolittle (1981) hydropathy scale was used. Red circles indicate disordered proteins, blue circles indicate structured ...
As stated in our manuscript [1], “The C-H plots generated using scale SVM parameters scale, Kyte-Doolittle hydropathy scale, and Guy hydropathy scale for whole protein prediction are shown in Figure Figure1.1. Figure Figure1A,1A, which is derived by SVM parameters scale, shows many fewer misclassified disordered proteins on the ordered side, compared to Figure Figure1B1B and and1C1C.”
As further stated in reference [1]: “Figure 1. Charge-Hydropathy plots. In (A) the IDP-Hydropathy scale was used, in (B) the Guy (1985) Hydropathy scale was used, and in (C) the Kyte-Doolittle (1981) hydropathy scale was used. Red circles indicate disordered proteins, blue circles indicate structured proteins. For these plots, each scale was normalized to be in the interval of 0 to 1. The Guy’s scale is multiplied by -1 prior to normalization to conform to the energy rule set by Kyte-Doolittle scale. In (A) the function describing the boundary is: = 3.31 -0.97. In (B) the function describing the boundary is: = 2.32 -0.93. In (C), the function describing the boundary is = 1.35 -0.49.”
For more details, the reader is referred to the published manuscript [1]. Christopher J. Oldfield, Wei-Lun Hsu, Jingwei Meng, Li Shen 0001, Pedro Romero, Vladimir N. Uversky, A. Keith Dunker |
BMC Bioinform. | 9 |
| 2014 | Improving protein order-disorder classification using charge-hydropathy plotsabstractBACKGROUND: The earliest whole protein order/disorder predictor (Uversky et al., Proteins, 41: 415-427 (2000)), herein called the charge-hydropathy (C-H) plot, was originally developed using the Kyte-Doolittle (1982) hydropathy scale (Kyte & Doolittle., J. Mol. Biol, 157: 105-132(1982)). Here the goal is to determine whether the performance of the C-H plot in separating structured and disordered proteins can be improved by using an alternative hydropathy scale. RESULTS: Using the performance of the CH-plot as the metric, we compared 19 alternative hydropathy scales, with the finding that the Guy (1985) hydropathy scale (Guy, Biophys. J, 47:61-70(1985)) was the best of the tested hydropathy scales for separating large collections structured proteins and intrinsically disordered proteins (IDPs) on the C-H plot. Next, we developed a new scale, named IDP-Hydropathy, which further improves the discrimination between structured proteins and IDPs. Applying the C-H plot to a dataset containing 109 IDPs and 563 non-homologous fully structured proteins, the Kyte-Doolittle (1982) hydropathy scale, the Guy (1985) hydropathy scale, and the IDP-Hydropathy scale gave balanced two-state classification accuracies of 79%, 84%, and 90%, respectively, indicating a very substantial overall improvement is obtained by using different hydropathy scales. A correlation study shows that IDP-Hydropathy is strongly correlated with other hydropathy scales, thus suggesting that IDP-Hydropathy probably has only minor contributions from amino acid properties other than hydropathy. CONCLUSION: We suggest that IDP-Hydropathy would likely be the best scale to use for any type of algorithm developed to predict protein disorder. Christopher J. Oldfield, Wei-Lun Hsu, Jingwei Meng, Li Shen 0001, Pedro Romero, Vladimir N. Uversky, A. Keith Dunker |
BMC Bioinform. | 9 |
| 2012 | MoRFpred, a computational tool for sequence-based prediction and characterization of short disorder-to-order transitioning binding regions in proteinsabstractMOTIVATION: Molecular recognition features (MoRFs) are short binding regions located within longer intrinsically disordered regions that bind to protein partners via disorder-to-order transitions. MoRFs are implicated in important processes including signaling and regulation. However, only a limited number of experimentally validated MoRFs is known, which motivates development of computational methods that predict MoRFs from protein chains. RESULTS: We introduce a new MoRF predictor, MoRFpred, which identifies all MoRF types (α, β, coil and complex). We develop a comprehensive dataset of annotated MoRFs to build and empirically compare our method. MoRFpred utilizes a novel design in which annotations generated by sequence alignment are fused with predictions generated by a Support Vector Machine (SVM), which uses a custom designed set of sequence-derived features. The features provide information about evolutionary profiles, selected physiochemical properties of amino acids, and predicted disorder, solvent accessibility and B-factors. Empirical evaluation on several datasets shows that MoRFpred outperforms related methods: α-MoRF-Pred that predicts α-MoRFs and ANCHOR which finds disordered regions that become ordered when bound to a globular partner. We show that our predicted (new) MoRF regions have non-random sequence similarity with native MoRFs. We use this observation along with the fact that predictions with higher probability are more accurate to identify putative MoRF regions. We also identify a few sequence-derived hallmarks of MoRFs. They are characterized by dips in the disorder predictions and higher hydrophobicity and stability when compared to adjacent (in the chain) residues. AVAILABILITY: http://biomine.ece.ualberta.ca/MoRFpred/; http://biomine.ece.ualberta.ca/MoRFpred/Supplement.pdf. Fatemeh Miri Disfani, Wei-Lun Hsu, Marcin J. Mizianty, Christopher J. Oldfield, A. Keith Dunker, Vladimir N. Uversky, Lukasz A. Kurgan |
Bioinform. | 7 |
| 2012 | Disease-Associated Mutations Disrupt Functionally Important Regions of Intrinsic Protein DisorderabstractThe effects of disease mutations on protein structure and function have been extensively investigated, and many predictors of the functional impact of single amino acid substitutions are publicly available. The majority of these predictors are based on protein structure and evolutionary conservation, following the assumption that disease mutations predominantly affect folded and conserved protein regions. However, the prevalence of the intrinsically disordered proteins (IDPs) and regions (IDRs) in the human proteome together with their lack of fixed structure and low sequence conservation raise a question about the impact of disease mutations in IDRs. Here, we investigate annotated missense disease mutations and show that 21.7% of them are located within such intrinsically disordered regions. We further demonstrate that 20% of disease mutations in IDRs cause local disorder-to-order transitions, which represents a 1.7-2.7 fold increase compared to annotated polymorphisms and neutral evolutionary substitutions, respectively. Secondary structure predictions show elevated rates of transition from helices and strands into loops and vice versa in the disease mutations dataset. Disease disorder-to-order mutations also influence predicted molecular recognition features (MoRFs) more often than the control mutations. The repertoire of disorder-to-order transition mutations is limited, with five most frequent mutations (R→W, R→C, E→K, R→H, R→Q) collectively accounting for 44% of all deleterious disorder-to-order transitions. As a proof of concept, we performed accelerated molecular dynamics simulations on a deleterious disorder-to-order transition mutation of tumor protein p63 and, in agreement with our predictions, observed an increased α-helical propensity of the region harboring the mutation. Our findings highlight the importance of mutations in IDRs and refine the traditional structure-centric view of disease mutations. The results of this study offer a new perspective on the role of mutations in disease, with implications for improving predictors of the functional impact of missense mutations. Vladimir Vacic, Phineus R. L. Markwick, Christopher J. Oldfield, Xiaoyue Zhao, Chad Haynes, Vladimir N. Uversky, Lilia M. Iakoucheva |
PLoS Comput. Biol. | 6 |
| 2011 | In-silico prediction of disorder content using hybrid sequence representationabstractBACKGROUND: Intrinsically disordered proteins play important roles in various cellular activities and their prevalence was implicated in a number of human diseases. The knowledge of the content of the intrinsic disorder in proteins is useful for a variety of studies including estimation of the abundance of disorder in protein families, classes, and complete proteomes, and for the analysis of disorder-related protein functions. The above investigations currently utilize the disorder content derived from the per-residue disorder predictions. We show that these predictions may over-or under-predict the overall amount of disorder, which motivates development of novel tools for direct and accurate sequence-based prediction of the disorder content. RESULTS: We hypothesize that sequence-level aggregation of input information may provide more accurate content prediction when compared with the content extracted from the local window-based residue-level disorder predictors. We propose a novel predictor, DisCon, that takes advantage of a small set of 29 custom-designed descriptors that aggregate and hybridize information concerning sequence, evolutionary profiles, and predicted secondary structure, solvent accessibility, flexibility, and annotation of globular domains. Using these descriptors and a ridge regression model, DisCon predicts the content with low, 0.05, mean squared error and high, 0.68, Pearson correlation. This is a statistically significant improvement over the content computed from outputs of ten modern disorder predictors on a test dataset with proteins that share low sequence identity with the training sequences. The proposed predictive model is analyzed to discuss factors related to the prediction of the disorder content. CONCLUSIONS: DisCon is a high-quality alternative for high-throughput annotation of the disorder content. We also empirically demonstrate that the DisCon's predictions can be used to improve binary annotations of the disordered residues from the real-value disorder propensities generated by current residue-level disorder predictors. The web server that implements the DisCon is available at http://biomine.ece.ualberta.ca/DisCon/. Marcin J. Mizianty, Yaoqi Zhou, A. Keith Dunker, Vladimir N. Uversky, Lukasz A. Kurgan |
BMC Bioinform. | 6 |
| 2009 | Influence of Sequence Changes and Environment on Intrinsically Disordered ProteinsabstractMany large-scale studies on intrinsically disordered proteins are implicitly based on the structural models deposited in the Protein Data Bank. Yet, the static nature of deposited models supplies little insight into variation of protein structure and function under diverse cellular and environmental conditions. While the computational predictability of disordered regions provides practical evidence that disorder is an intrinsic property of proteins, the robustness of disordered regions to changes in sequence or environmental conditions has not been systematically studied. We analyzed intrinsically disordered regions in the same or similar proteins crystallized independently and studied their sensitivity to changes in protein sequence and parameters of crystallographic experiments. The observed changes in the existence, position, and length of disordered regions indicate that their appearance in X-ray structures dramatically depends on changes in amino acid sequence and peculiarities of the crystallographic experiment. Our study also raises general questions regarding protein evolution and the regulation of protein structure, dynamics, and function via variations in cellular and environmental conditions. Amrita Mohan, Vladimir N. Uversky, Predrag Radivojac |
PLoS Comput. Biol. | 2 |
| 2008 | Malleable Machines in Transcription Regulation: The Mediator ComplexabstractThe Mediator complex provides an interface between gene-specific regulatory proteins and the general transcription machinery including RNA polymerase II (RNAP II). The complex has a modular architecture (Head, Middle, and Tail) and cryoelectron microscopy analysis suggested that it undergoes dramatic conformational changes upon interactions with activators and RNAP II. These rearrangements have been proposed to play a role in the assembly of the preinitiation complex and also to contribute to the regulatory mechanism of Mediator. In analogy to many regulatory and transcriptional proteins, we reasoned that Mediator might also utilize intrinsically disordered regions (IDRs) to facilitate structural transitions and transmit transcriptional signals. Indeed, a high prevalence of IDRs was found in various subunits of Mediator from both Saccharomyces cerevisiae and Homo sapiens, especially in the Tail and the Middle modules. The level of disorder increases from yeast to man, although in both organisms it significantly exceeds that of multiprotein complexes of a similar size. IDRs can contribute to Mediator's function in three different ways: they can individually serve as target sites for multiple partners having distinctive structures; they can act as malleable linkers connecting globular domains that impart modular functionality on the complex; and they can also facilitate assembly and disassembly of complexes in response to regulatory signals. Short segments of IDRs, termed molecular recognition features (MoRFs) distinguished by a high protein-protein interaction propensity, were identified in 16 and 19 subunits of the yeast and human Mediator, respectively. In Saccharomyces cerevisiae, the functional roles of 11 MoRFs have been experimentally verified, and those in the Med8/Med18/Med20 and Med7/Med21 complexes were structurally confirmed. Although the Saccharomyces cerevisiae and Homo sapiens Mediator sequences are only weakly conserved, the arrangements of the disordered regions and their embedded interaction sites are quite similar in the two organisms. All of these data suggest an integral role for intrinsic disorder in Mediator's function. Ágnes Tóth-Petróczy, Christopher J. Oldfield, István Simon, Yuichiro Takagi, A. Keith Dunker, Vladimir N. Uversky, Mónika Fuxreiter |
PLoS Comput. Biol. | 6 |
| 2007 | Intrinsically Disordered Proteins: Predictions and ApplicationsabstractAbout 10 years ago we published our first predictor of intrinsically disordered protein residues in another IEEE journal, theProceedingsoftheIEEEInternationalConferenceonNeural Networks. Others call such proteins "natively unfolded" and "intrinsically unstructured." Since then, we and others have substantially improved the prediction of intrinsically disordered residues. The prediction of protein intrinsic disorder is similar to the prediction of secondary structure in terms of methodology, but, at the structural level, secondary structure (especially random coil) and intrinsic disorder differ completely in their dynamic motion. First, we will briefly describe the prediction of protein disorder, show the progress from ~ 70 % to ~ 85 % per residue prediction accuracy, and show that intrinsically disordered proteins are common over the three domains of life, but are especially common among the eukaryotes. Next we will discuss our methods for deducing functions that are associated with disordered rather than structured proteins. In brief, structured proteins have advantages for catalysis while disordered proteins and regions have advantages for the reversible, weak binding often observed in signaling, control, and regulation. After that we will discuss how disorder facilitates binding diversity in protein-protein interaction networks, both for single disordered regions binding to many partners and for many disordered regions with different sequences binding to a common site on the surface of one structured protein. Part three presents data indicating that alternative splicing is more prevalent in regions of RNA that code for disorder than those that code for structure, thus providing a means for evolving tissue-specific signaling networks. Finally, we will present a novel approach to drug discovery based on disordered protein. A. Keith Dunker, Christopher J. Oldfield, Jingwei Meng, Pedro Romero, Jack Y. Yang, Zoran Obradovic, Vladimir N. Uversky |
BIBE | 7 |
| 2007 | Intrinsically Disordered Proteins: An UpdateabstractJust over 10 years ago, in June, 1997, in the Proceedings of the IEEE International Conference on Neural Networks, we published our first predictor of intrinsically disordered protein [1]. Since then, we have substantially improved our predictors, and more than 20 other laboratory groups have joined in efforts to improve the prediction of protein disorder. At the algorithmic level, prediction of protein intrinsic disorder is similar to the prediction of secondary structure, but, at the structural level, secondary structure and intrinsic disorder are entirely different. The secondary structure class called random coil or irregular differs from intrinsic disorder due to very different dynamic properties, with the secondary structure class being much less mobile than the region of disorder. At the biological level, unlike the prediction of secondary structure, the prediction of intrinsic disorder has been revolutionary. That is, for many years, experimentalists have provided evidence that some proteins lack fixed structure or are disordered (or unfolded) under physiological conditions. Experimentalists further are showing that, for some proteins, functions depended on the unstructured rather than structured state. However, these examples have been mostly ignored. To our knowledge, not one disordered protein or disorder-associated function is discussed in any biochemistry textbook, even though such examples began to be discovered more than 50 years ago. Disorder prediction has been important for showing that the few experimentally characterized examples represent a very large cohort that is found all across all three domains of life. A. Keith Dunker, Christopher J. Oldfield, Jingwei Meng, Pedro Romero, Jack Y. Yang, Zoran Obradovic, Vladimir N. Uversky |
BIBE | 7 |
| 2007 | Intrinsically Disordered Proteins in Human DiseasesabstractIntrinsically disordered proteins lack stable tertiary and/or secondary structure under physiological conditions in vitro. They are highly abundant in nature, with ~25-30% of eukaryotic proteins being mostly disordered, and with >50% of eukaryotic proteins and > 70% of signaling proteins having long disordered regions. Functional repertoire of intrinsically disordered proteins is very broad and complements functions of ordered proteins. Often, intrinsically disordered proteins are involved in regulation, signaling and control pathways, where binding to multiple partners and high-speciflcity/low-afflnity interactions play a crucial role. We have found that out of the 711 Swiss-Prot functional keywords associated with at least 20 proteins, 262 were strongly positively correlated with long intrinsically disordered regions, and 302 were strongly negatively correlated. It is suggested that functions of intrinsically disordered proteins may arise from the specific disorder form, from inter-conversion of disordered forms, or from transitions between disordered and ordered conformations. The choice between these conformations is determined by the peculiarities of the protein environment, and many intrinsically disordered proteins possess an exceptional ability to fold in a template-dependent manner. Intrinsically disordered proteins are key players in protein-protein interaction networks being highly abundant among hubs. Furthermore, regions of mRNA which undergo alternative splicing code for disordered proteins much more often than they code for structured proteins. This association of alternative splicing and intrinsic disorder helps proteins to avoid folding difficulties and provides a novel mechanism for developing tissue-specific protein interaction networks. Numerous intrinsically disordered proteins are associated with such human diseases as cancer, cardiovascular disease, amyloidoses, neurodegenerative diseases, diabetes and others. Our bioinformatics analysis revealed that many human diseases are strongly correlated with proteins predicted to be disordered. Contrary to this, we did not find disease associated proteins to be strongly correlated with absence of disorder. Overall, there is an intriguing interconnection between intrinsic disorder, cell signaling and human diseases, which suggests that protein conformational diseases may result not only from protein misfolding, but also from misidentification and missignaling. Intrinsically disordered proteins, such as alpha-synuclein, tau protein, p53, BRCA1 and many other disease-associated hub proteins represent attractive targets for drugs modulating protein-protein interactions. Therefore, novel strategies for drug discovery are based on intrinsically disordered proteins. Vladimir N. Uversky, Christopher J. Oldfield, A. Keith Dunker |
BIBE | 1 |
| 2007 | Composition Profiler: a tool for discovery and visualization of amino acid composition differencesabstractBACKGROUND: Composition Profiler is a web-based tool for semi-automatic discovery of enrichment or depletion of amino acids, either individually or grouped by their physico-chemical or structural properties. RESULTS: The program takes two samples of amino acids as input: a query sample and a reference sample. The latter provides a suitable background amino acid distribution, and should be chosen according to the nature of the query sample, for example, a standard protein database (e.g. SwissProt, PDB), a representative sample of proteins from the organism under study, or a group of proteins with a contrasting functional annotation. The results of the analysis of amino acid composition differences are summarized in textual and graphical form. CONCLUSION: As an exploratory data mining tool, our software can be used to guide feature selection for protein function or structure predictors. For classes of proteins with significant differences in frequencies of amino acids having particular physico-chemical (e.g. hydrophobicity or charge) or structural (e.g. alpha helix propensity) properties, Composition Profiler can be used as a rough, light-weight visual classifier. Vladimir Vacic, Vladimir N. Uversky, A. Keith Dunker, Stefano Lonardi |
BMC Bioinform. | 2 |
| 2006 | Intrinsic Disorder Is a Common Feature of Hub Proteins from Four Eukaryotic InteractomesabstractRecent proteome-wide screening approaches have provided a wealth of information about interacting proteins in various organisms. To test for a potential association between protein connectivity and the amount of predicted structural disorder, the disorder propensities of proteins with various numbers of interacting partners from four eukaryotic organisms (Caenorhabditis elegans, Saccharomyces cerevisiae, Drosophila melanogaster, and Homo sapiens) were investigated. The results of PONDR VL-XT disorder analysis show that for all four studied organisms, hub proteins, defined here as those that interact with > or = 10 partners, are significantly more disordered than end proteins, defined here as those that interact with just one partner. The proportion of predicted disordered residues, the average disorder score, and the number of predicted disordered regions of various lengths were higher overall in hubs than in ends. A binary classification of hubs and ends into ordered and disordered subclasses using the consensus prediction method showed a significant enrichment of wholly disordered proteins and a significant depletion of wholly ordered proteins in hubs relative to ends in worm, fly, and human. The functional annotation of yeast hubs and ends using GO categories and the correlation of these annotations with disorder predictions demonstrate that proteins with regulation, transcription, and development annotations are enriched in disorder, whereas proteins with catalytic activity, transport, and membrane localization annotations are depleted in disorder. The results of this study demonstrate that intrinsic structural disorder is a distinctive and common characteristic of eukaryotic hub proteins, and that disorder may serve as a determinant of protein interactivity. Chad Haynes, Christopher J. Oldfield, Niels Klitgord, Michael E. Cusick, Predrag Radivojac, Vladimir N. Uversky, Marc Vidal, Lilia M. Iakoucheva |
PLoS Comput. Biol. | 7 |