Silvio C. E. Tosatto

dblp:88/5583 · DBLP profile ↗
← Back
34ranked-venue papers
1as first author
8since 2021 · last 2025
0000-0003-4525-7793ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 34 · 1 first-author · 8 since 2021
YearPublicationVenuePosition
2025 MobiDB-lite 4.0: faster prediction of intrinsic protein disorder and structural compactness
abstract
MOTIVATION: In recent years, many disorder predictors have been developed to identify intrinsically disordered regions (IDRs) in proteins, achieving high accuracy. However, it may be difficult to interpret differences in predictions across methods. Consensus methods offer a simple solution, highlighting reliable predictions while filtering out uncertain positions. Here, we present a new version of MobiDB-lite, a consensus method designed to predict long IDRs and classify them based on compositional biases and conformational properties. RESULTS: MobiDB-lite 4.0 pipeline was optimized to be ten times faster than the previous version. It now provides compactness annotations based on predicted apparent scaling exponent. The newly added features and disorder subclassifications allow the users to get a comprehensive insight into the protein's function and characteristics. MobiDB-lite 4.0 is integrated into the MobiDB and DisProt databases. A version without the compactness predictor is integrated into InterProScan, propagating MobiDB-lite annotations to UniProtKB. AVAILABILITY AND IMPLEMENTATION: The MobiDB-lite 4.0 source code and a Docker container are available from the GitHub repository: https://github.com/BioComputingUP/MobiDB-lite.
Mahta Mehdiabadi, Matthias Blum, Giulio Tesei, Sören von Bülow, Kresten Lindorff-Larsen, Silvio C. E. Tosatto, Damiano Piovesan
Bioinform.6
2025 GeomeTRe: accurate calculation of geometrical descriptors of tandem repeat proteins
abstract
MOTIVATION: Structured tandem repeat proteins (STRPs) are characterized by preserved structural motifs arranged in a modular way. The structural and functional diversity of STRPs makes them particularly important for studying evolution and novel structure-function relationships, and ultimately for designing new synthetic proteins with specific functions. One crucial aspect of their classification is the estimation of geometrical parameters, which can provide better insight into their properties and the relationship between the spatial arrangement of repeated units and protein function. Calculating geometric descriptors for STRPs is challenging because naturally occurring repeats are not "perfect" and often contain insertions and deletions. Existing tools for predicting structural symmetry work well on simple cases but often fail for most natural proteins. RESULTS: Here, we present GeomeTRe, an algorithm that calculates geometrical descriptors such as curvature (yaw), twist (roll), and pitch for a protein structure with known repeat unit positions. The algorithm simulates the movement of consecutive units, identifies rotational axes, and calculates the corresponding Tait-Bryan angles. GeomeTRe's parameters can enhance STRP annotation and classification by identifying variations in geometric arrangements among different functional groups. The package is fast and suitable for processing large protein structure datasets when repeat region information (e.g. from RepeatsDB) is available. AVAILABILITY AND IMPLEMENTATION: GeomeTRe is available as a Python package; source code and documentation can be found at https://github.com/BioComputingUP/GeomeTRe.
Zarifa Osmanli, Elisa Ferrero, Alexander Miguel Monzon, Silvio C. E. Tosatto, Damiano Piovesan
Bioinform.4
2025 Mathematical modeling and simulation of tumor-induced angiogenesis in retinal hemangioblastoma
abstract
Retinal Hemangioblastoma (RH) is the most frequent manifestation of the von Hippel-Lindau syndrome (VHL), a rare disease associated with the germline mutation of the von Hippel-Lindau protein (pVHL). An emblematic feature of RH is the high vascularity, which is explained by the overexpression of angiogenic factors (AFs) arising from the pVHL impairment. The introduction of Optical Coherence Tomography Angiography (OCTA) allowed observing this feature with exceptional detail. Here, we combine OCTA images and a mechanistic model to investigate tumor growth and vascular development in a patient-specific way. We derived our model from the agreed pathology for RH and focused on the earliest stages of tumor-induced angiogenesis. Our simulations closely resemble the medical images, supporting the capability of our model to simulate vascular patterning in actual patients. Our results also suggest that angiogenesis in RH occurs upon reaching a critical dimension (around 200 μm), followed by the rapid formation of stable vascular networks. These findings open a new perspective on the crucial role of time in antiangiogenic therapy in RH, which has resulted in ineffective control. Indeed, it might be that when RH is diagnosed, angiogenesis is already too advanced to be effectively targeted with any effective means. Moreover, our simulations suggest that vascularization in RH is not a continuous process but an inconstant development with long, stable phases and rapid episodes of vascular sprouting.
Franco Pradelli, Giovanni Minervini, Pradeep Venkatesh, Shorya Azad, Hector Gomez, Silvio C. E. Tosatto
PLoS Comput. Biol.6
2024 STRPsearch: fast detection of structured tandem repeat proteins
abstract
MOTIVATION: Structured Tandem Repeats Proteins (STRPs) constitute a subclass of tandem repeats characterized by repetitive structural motifs. These proteins exhibit distinct secondary structures that form repetitive tertiary arrangements, often resulting in large molecular assemblies. Despite highly variable sequences, STRPs can perform important and diverse biological functions, maintaining a consistent structure with a variable number of repeat units. With the advent of protein structure prediction methods, millions of 3D models of proteins are now publicly available. However, automatic detection of STRPs remains challenging with current state-of-the-art tools due to their lack of accuracy and long execution times, hindering their application on large datasets. In most cases, manual curation remains the most accurate method for detecting and classifying STRPs, making it impracticable to annotate millions of structures. RESULTS: We introduce STRPsearch, a novel tool for the rapid identification, classification, and mapping of STRPs. Leveraging manually curated entries from RepeatsDB as the known conformational space of STRPs, STRPsearch uses the latest advances in structural alignment for a fast and accurate detection of repeated structural motifs in proteins, followed by an innovative approach to map units and insertions through the generation of TM-score profiles. STRPsearch is highly scalable, efficiently processing large datasets, and can be applied to both experimental structures and predicted models. In addition, it demonstrates superior performance compared to existing tools, offering researchers a reliable and comprehensive solution for STRP analysis across diverse proteomes. AVAILABILITY AND IMPLEMENTATION: STRPsearch is coded in Python. All scripts and associated documentation are available from: https://github.com/BioComputingUP/STRPsearch.
Soroush Mozaffari, Paula Nazarena Arrías, Damiano Clementel, Damiano Piovesan, Carlo Ferrari, Silvio C. E. Tosatto, Alexander Miguel Monzon
Bioinform.6
2023 RING-PyMOL: residue interaction networks of structural ensembles and molecular dynamics
abstract
RING-PyMOL is a plugin for PyMOL providing a set of analysis tools for structural ensembles and molecular dynamic simulations. RING-PyMOL combines residue interaction networks, as provided by the RING software, with structural clustering to enhance the analysis and visualization of the conformational complexity. It combines precise calculation of non-covalent interactions with the power of PyMOL to manipulate and visualize protein structures. The plugin identifies and highlights correlating contacts and interaction patterns that can explain structural allostery, active sites, and structural heterogeneity connected with molecular function. It is easy to use and extremely fast, processing and rendering hundreds of models and long trajectories in seconds. RING-PyMOL generates a number of interactive plots and output files for use with external tools. The underlying RING software has been improved extensively. It is 10 times faster, can process mmCIF files and it identifies typed interactions also for nucleic acids. AVAILABILITY AND IMPLEMENTATION: https://github.com/BioComputingUP/ring-pymol.
Alessio Del Conte, Alexander Miguel Monzon, Damiano Clementel, Giorgia F. Camagni, Giovanni Minervini, Silvio C. E. Tosatto, Damiano Piovesan
Bioinform.6
2022 ProSeqViewer: an interactive, responsive and efficient TypeScript library for visualization of sequences and alignments in web applications
abstract
SUMMARY: Biological data is ever-increasing in amount and complexity. The mapping of this data to biological entities such as nucleotide and amino acid sequences supports biological data analysis, classification and prediction. Sequence alignments and comparison allow the transfer of knowledge to evolutionary-related entities, the mapping of functional domains, the identification of binding and modification sites. To support these types of studies, we developed ProSeqViewer, a tool to visualize annotation on single sequences and multiple sequence alignments. This state-of-the-art multifunctional library was developed as a modular component to be integrated into static or dynamic web resources and support intuitive visualization of sequence features. ProseSeqViewer is extremely lightweight, fast, interactive, dynamic, responsive and works at any screen size. It generates pure HTML which is compatible with any browser and operating system. ProSeqViewer can exchange events with other visualization components and is already used by multiple biological databases. AVAILABILITY AND IMPLEMENTATION: ProSeqViewer is an open-source TypeScript library compatible with state-of-the-art website environments. The source code and an extensive documentation including use cases are available from the URL: https://github.com/BioComputingUP/ProSeqViewer.
Martina Bevilacqua, Lisanna Paladin, Silvio C. E. Tosatto, Damiano Piovesan
Bioinform.3
2022 Mocafe: a comprehensive Python library for simulating cancer development with Phase Field Models
abstract
SUMMARY: Mathematical models are effective in studying cancer development at different scales from metabolism to tissue. Phase Field Models (PFMs) have been shown to reproduce accurately cancer growth and other related phenomena, including expression of relevant molecules, extracellular matrix remodeling and angiogenesis. However, implementations of such models are rarely published, reducing access to these techniques. To reduce this gap, we developed Mocafe, a modular open-source Python package that implements some of the most important PFMs reported in the literature. Mocafe is designed to handle both PFMs purely based on differential equations and hybrid agent-based PFMs. Moreover, Mocafe is meant to be extensible, allowing the inclusion of new models in future releases. AVAILABILITY AND IMPLEMENTATION: Mocafe is a Python package based on FEniCS, a popular computing platform for solving partial differential equations. The source code, extensive documentation and demos are provided on GitHub at URL: https://github.com/BioComputingUP/mocafe. Moreover, we uploaded on Zenodo an archive of the package, which is available at https://doi.org/10.5281/zenodo.6366052.
Franco Pradelli, Giovanni Minervini, Silvio C. E. Tosatto
Bioinform.3
2021 MobiDB-lite 3.0: fast consensus annotation of intrinsic disorder flavors in proteins
abstract
MOTIVATION: The earlier version of MobiDB-lite is currently used in large-scale proteome annotation platforms to detect intrinsic disorder. However, new theoretical models allow for the classification of intrinsically disordered regions into subtypes from sequence features associated with specific polymeric properties or compositional bias. RESULTS: MobiDB-lite 3.0 maintains its previous speed and performance but also provides a finer classification of disorder by identifying regions with characteristics of polyolyampholytes, positive or negative polyelectrolytes, low-complexity regions or enriched in cysteine, proline or glycine or polar residues. Subregions are abundantly detected in IDRs of the human proteome. The new version of MobiDB-lite represents a new step for the proteome level analysis of protein disorder. AVAILABILITY AND IMPLEMENTATION: Both the MobiDB-lite 3.0 source code and a docker container are available from the GitHub repository: https://github.com/BioComputingUP/MobiDB-lite.
Marco Necci, Damiano Piovesan, Damiano Clementel, Zsuzsanna Dosztányi, Silvio C. E. Tosatto
Bioinform.5
2020 Disentangling the complexity of low complexity proteins
abstract
There are multiple definitions for low complexity regions (LCRs) in protein sequences, with all of them broadly considering LCRs as regions with fewer amino acid types compared to an average composition. Following this view, LCRs can also be defined as regions showing composition bias. In this critical review, we focus on the definition of sequence complexity of LCRs and their connection with structure. We present statistics and methodological approaches that measure low complexity (LC) and related sequence properties. Composition bias is often associated with LC and disorder, but repeats, while compositionally biased, might also induce ordered structures. We illustrate this dichotomy, and more generally the overlaps between different properties related to LCRs, using examples. We argue that statistical measures alone cannot capture all structural aspects of LCRs and recommend the combined usage of a variety of predictive tools and measurements. While the methodologies available to study LCRs are already very advanced, we foresee that a more comprehensive annotation of sequences in the databases will enable the improvement of predictions and a better understanding of the evolution and the connection between structure and function of LCRs. This will require the use of standards for the generation and exchange of data describing all aspects of LCRs. SHORT ABSTRACT: There are multiple definitions for low complexity regions (LCRs) in protein sequences. In this critical review, we focus on the definition of sequence complexity of LCRs and their connection with structure. We present statistics and methodological approaches that measure low complexity (LC) and related sequence properties. Composition bias is often associated with LC and disorder, but repeats, while compositionally biased, might also induce ordered structures. We illustrate this dichotomy, plus overlaps between different properties related to LCRs, using examples.
Pablo Mier, Lisanna Paladin, Stella Tamana, Sophia Petrosian, Borbála Hajdu-Soltész, Annika Urbanek, Aleksandra Gruca, Dariusz Plewczynski, Marcin Grynberg, Pau Bernadó, Zoltán Gáspári, Christos A. Ouzounis, Vasilis J. Promponas, Andrey V. Kajava, John M. Hancock, Silvio C. E. Tosatto, Zsuzsanna Dosztányi, Miguel A. Andrade-Navarro
Briefings Bioinform.16
2020 The Feature-Viewer: a visualization tool for positional annotations on a sequence
abstract
SUMMARY: The Feature-Viewer is a lightweight library for the visualization of biological data mapped to a protein or nucleotide sequence. It is designed for ease of use while allowing for a full customization. The library is already used by several biological data resources and allows intuitive visual mapping of a full spectra of sequence features for different usages. AVAILABILITY AND IMPLEMENTATION: The Feature-Viewer is open source, compatible with state-of-the-art development technologies and responsive, also for mobile viewing. Documentation and usage examples are available online.
Lisanna Paladin, Mathieu Schaeffer, Pascale Gaudet, Monique Zahn-Zabal, Pierre-André Michel, Damiano Piovesan, Silvio C. E. Tosatto, Amos Bairoch
Bioinform.7
2020 Assessing predictors for new post translational modification sites: A case study on hydroxylation
abstract
Post-translational modification (PTM) sites have become popular for predictor development. However, with the exception of phosphorylation and a handful of other examples, PTMs suffer from a limited number of available training examples and sparsity in protein sequences. Here, proline hydroxylation is taken as an example to compare different methods and evaluate their performance on new experimentally determined sites. As a guide for effective experimental design, predictors require both high specificity and sensitivity. However, the self-reported performance may often not be indicative of prediction quality and detection of new sites is not guaranteed. We have benchmarked seven published hydroxylation site predictors on two newly constructed independent datasets. The self-reported performance is found to widely overestimate the real accuracy measured on independent datasets. No predictor performs better than random on new examples, indicating the refined models do not sufficiently generalize to detect new sites. The number of false positives is high and precision low, in particular for non-collagen proteins whose motifs are not conserved. As hydroxylation site predictors do not generalize for new data, caution is advised when using PTM predictors in the absence of independent evaluations, in particular for highly specific sites involved in signalling.
Damiano Piovesan, András Hatos, Giovanni Minervini, Federica Quaglia, Alexander Miguel Monzon, Silvio C. E. Tosatto
PLoS Comput. Biol.6
2019 Genotype-phenotype relations of the von Hippel-Lindau tumor suppressor inferred from a large-scale analysis of disease mutations and interactors
abstract
Familiar cancers represent a privileged point of view for studying the complex cellular events inducing tumor transformation. Von Hippel-Lindau syndrome, a familiar predisposition to develop cancer is a clear example. Here, we present our efforts to decipher the role of von Hippel-Lindau tumor suppressor protein (pVHL) in cancer insurgence. We collected high quality information about both pVHL mutations and interactors to investigate the association between patient phenotypes, mutated protein surface and impaired interactions. Our data suggest that different phenotypes correlate with localized perturbations of the pVHL structure, with specific cell functions associated to different protein surfaces. We propose five different pVHL interfaces to be selectively involved in modulating proteins regulating gene expression, protein homeostasis as well as to address extracellular matrix (ECM) and ciliogenesis associated functions. These data were used to drive molecular docking of pVHL with its interactors and guide Petri net simulations of the most promising alterations. We predict that disruption of pVHL association with certain interactors can trigger tumor transformation, inducing metabolism imbalance and ECM remodeling. Collectively taken, our findings provide novel insights into VHL-associated tumorigenesis. This highly integrated in silico approach may help elucidate novel treatment paradigms for VHL disease.
Giovanni Minervini, Federica Quaglia, Francesco Tabaro, Silvio C. E. Tosatto
PLoS Comput. Biol.4
2018 A comprehensive assessment of long intrinsic protein disorder from the DisProt database
abstract
Motivation: Intrinsic disorder (ID), i.e. the lack of a unique folded conformation at physiological conditions, is a common feature for many proteins, which requires specialized biochemical experiments that are not high-throughput. Missing X-ray residues from the PDB have been widely used as a proxy for ID when developing computational methods. This may lead to a systematic bias, where predictors deviate from biologically relevant ID. Large benchmarking sets on experimentally validated ID are scarce. Recently, the DisProt database has been renewed and expanded to include manually curated ID annotations for several hundred new proteins. This provides a large benchmark set which has not yet been used for training ID predictors. Results: Here, we describe the first systematic benchmarking of ID predictors on the new DisProt dataset. In contrast to previous assessments based on missing X-ray data, this dataset contains mostly long ID regions and a significant amount of fully ID proteins. The benchmarking shows that ID predictors work quite well on the new dataset, especially for long ID segments. However, a large fraction of ID still goes virtually undetected and the ranking of methods is different than for PDB data. In particular, many predictors appear to confound ID and regions outside X-ray structures. This suggests that the ID prediction methods capture different flavors of disorder and can benefit from highly accurate curated examples. Availability and implementation: The raw data used for the evaluation are available from URL: http://www.disprot.org/assessment/. Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Marco Necci, Damiano Piovesan, Zsuzsanna Dosztányi, Peter Tompa, Silvio C. E. Tosatto
Bioinform.5
2018 Mobi 2.0: an improved method to define intrinsic disorder, mobility and linear binding regions in protein structures
abstract
Motivation: The structures contained in the Protein Data Bank (PDB) database are of paramount importance to define our knowledge of folded proteins. While providing mainly circumstantial evidence, PDB data is also increasingly used to define the lack of unique structure, represented by mobile regions and even intrinsic disorder (ID). However, alternative definitions are used by different authors and potentially limit the generality of the analyses being carried out. Results: Here we present Mobi 2.0, a completely re-written version of the Mobi software for the determination of mobile and potentially disordered regions from PDB structures. Mobi 2.0 provides robust definitions of mobility based on four main sources of information: (i) missing residues, (ii) residues with high temperature factors, (iii) mobility between different models of the same structure and (iv) binding to another protein or nucleotide chain. Mobi 2.0 is well suited to aggregate information across different PDB structures for the same UniProt protein sequence, providing consensus annotations. The software is expected to standardize the treatment of mobility, allowing an easier comparison across different studies related to ID. Availability: Mobi 2.0 provides the structure-based annotation for the MobiDB database. The software is available from URL http://protein.bio.unipd.it/mobi2/. Contact: [email protected].
Damiano Piovesan, Silvio C. E. Tosatto
Bioinform.2
2017 MobiDB-lite: fast and highly specific consensus prediction of intrinsic disorder in proteins
abstract
Motivation: Intrinsic disorder (ID) is established as an important feature of protein sequences. Its use in proteome annotation is however hampered by the availability of many methods with similar performance at the single residue level, which have mostly not been optimized to predict long ID regions of size comparable to domains. Results: Here, we have focused on providing a single consensus-based prediction, MobiDB-lite, optimized for highly specific (i.e. few false positive) predictions of long disorder. The method uses eight different predictors to derive a consensus which is then filtered for spurious short predictions. Consensus prediction is shown to outperform the single methods when annotating long ID regions. MobiDB-lite can be useful in large-scale annotation scenarios and has indeed already been integrated in the MobiDB, DisProt and InterPro databases. Availability and Implementation: MobiDB-lite is available as part of the MobiDB database from URL: http://mobidb.bio.unipd.it/. An executable can be downloaded from URL: http://protein.bio.unipd.it/mobidblite/. Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Marco Necci, Damiano Piovesan, Zsuzsanna Dosztányi, Silvio C. E. Tosatto
Bioinform.4
2017 FELLS: fast estimator of latent local structure
abstract
MOTIVATION: The behavior of a protein is encoded in its sequence, which can be used to predict distinct features such as secondary structure, intrinsic disorder or amphipathicity. Integrating these and other features can help explain the context-dependent behavior of proteins. However, most tools focus on a single aspect, hampering a holistic understanding of protein structure. Here, we present Fast Estimator of Latent Local Structure (FELLS) to visualize structural features from the protein sequence. FELLS provides disorder, aggregation and low complexity predictions as well as estimated local propensities including amphipathicity. A novel fast estimator of secondary structure (FESS) is also trained to provide a fast response. The calculations required for FELLS are extremely fast and suited for large-scale analysis while providing a detailed analysis of difficult cases. AVAILABILITY AND IMPLEMENTATION: The FELLS web server is available from URL: http://protein.bio.unipd.it/fells/ . The server also exposes RESTful functionality allowing programmatic prediction requests. An executable version of FESS for Linux can be downloaded from URL: protein.bio.unipd.it/download/. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Damiano Piovesan, Ian Walsh, Giovanni Minervini, Silvio C. E. Tosatto
Bioinform.4
2017 Conformational diversity analysis reveals three functional mechanisms in proteins
abstract
Protein motions are a key feature to understand biological function. Recently, a large-scale analysis of protein conformational diversity showed a positively skewed distribution with a peak at 0.5 Å C-alpha root-mean-square-deviation (RMSD). To understand this distribution in terms of structure-function relationships, we studied a well curated and large dataset of ~5,000 proteins with experimentally determined conformational diversity. We searched for global behaviour patterns studying how structure-based features change among the available conformer population for each protein. This procedure allowed us to describe the RMSD distribution in terms of three main protein classes sharing given properties. The largest of these protein subsets (~60%), which we call "rigid" (average RMSD = 0.83 Å), has no disordered regions, shows low conformational diversity, the largest tunnels and smaller and buried cavities. The two additional subsets contain disordered regions, but with differential sequence composition and behaviour. Partially disordered proteins have on average 67% of their conformers with disordered regions, average RMSD = 1.1 Å, the highest number of hinges and the longest disordered regions. In contrast, malleable proteins have on average only 25% of disordered conformers and average RMSD = 1.3 Å, flexible cavities affected in size by the presence of disordered regions and show the highest diversity of cognate ligands. Proteins in each set are mostly non-homologous to each other, share no given fold class, nor functional similarity but do share features derived from their conformer population. These shared features could represent conformational mechanisms related with biological functions.
Alexander Miguel Monzon, Diego Javier Zea, María Silvina Fornasari, Tadeo E. Saldaño, Sebastian Fernandez Alberti, Silvio C. E. Tosatto, Gustavo D. Parisi
PLoS Comput. Biol.6
2016 Correct machine learning on protein sequences: a peer-reviewing perspective
abstract
Machine learning methods are becoming increasingly popular to predict protein features from sequences. Machine learning in bioinformatics can be powerful but carries also the risk of introducing unexpected biases, which may lead to an overestimation of the performance. This article espouses a set of guidelines to allow both peer reviewers and authors to avoid common machine learning pitfalls. Understanding biology is necessary to produce useful data sets, which have to be large and diverse. Separating the training and test process is imperative to avoid over-selling method performance, which is also dependent on several hidden parameters. A novel predictor has always to be compared with several existing methods, including simple baseline strategies. Using the presented guidelines will help nonspecialists to appreciate the critical issues in machine learning.
Ian Walsh, Gianluca Pollastri, Silvio C. E. Tosatto
Briefings Bioinform.3
2015 The Victor C++ library for protein representation and advanced manipulation
abstract
MOTIVATION: Protein sequence and structure representation and manipulation require dedicated software libraries to support methods of increasing complexity. Here, we describe the VIrtual Constrution TOol for pRoteins (Victor) C++ library, an open source platform dedicated to enabling inexperienced users to develop advanced tools and gathering contributions from the community. The provided application examples cover statistical energy potentials, profile-profile sequence alignments and ab initio loop modeling. Victor was used over the last 15 years in several publications and optimized for efficiency. It is provided as a GitHub repository with source files and unit tests, plus extensive online documentation, including a Wiki with help files and tutorials, examples and Doxygen documentation. AVAILABILITY AND IMPLEMENTATION: The C++ library and online documentation, distributed under a GPL license are available from URL: http://protein.bio.unipd.it/victor/.
Layla Hirsh, Damiano Piovesan, Manuel Giollo, Carlo Ferrari, Silvio C. E. Tosatto
Bioinform.5
2015 Comprehensive large-scale assessment of intrinsic protein disorder
abstract
MOTIVATION: Intrinsically disordered regions are key for the function of numerous proteins. Due to the difficulties in experimental disorder characterization, many computational predictors have been developed with various disorder flavors. Their performance is generally measured on small sets mainly from experimentally solved structures, e.g. Protein Data Bank (PDB) chains. MobiDB has only recently started to collect disorder annotations from multiple experimental structures. RESULTS: MobiDB annotates disorder for UniProt sequences, allowing us to conduct the first large-scale assessment of fast disorder predictors on 25 833 different sequences with X-ray crystallographic structures. In addition to a comprehensive ranking of predictors, this analysis produced the following interesting observations. (i) The predictors cluster according to their disorder definition, with a consensus giving more confidence. (ii) Previous assessments appear over-reliant on data annotated at the PDB chain level and performance is lower on entire UniProt sequences. (iii) Long disordered regions are harder to predict. (iv) Depending on the structural and functional types of the proteins, differences in prediction performance of up to 10% are observed. AVAILABILITY: The datasets are available from Web site at URL: http://mobidb.bio.unipd.it/lsd. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Ian Walsh, Manuel Giollo, Tomás Di Domenico, Carlo Ferrari, Olav Zimmermann, Silvio C. E. Tosatto
Bioinform.6
2013 Analysis and consensus of currently available intrinsic protein disorder annotation sources in the MobiDB database
abstract
BACKGROUND: Intrinsic protein disorder is becoming an increasingly important topic in protein science. During the last few years, intrinsically disordered proteins (IDPs) have been shown to play a role in many important biological processes, e.g. protein signalling and regulation. This has sparked a need to better understand and characterize different types of IDPs, their functions and roles. Our recently published database, MobiDB, provides a centralized resource for accessing and analysing intrinsic protein disorder annotations. RESULTS: Here, we present a thorough description and analysis of the data made available by MobiDB, providing descriptive statistics on the various available annotation sources. Version 1.2.1 of the database contains annotations for ca. 4,500,000 UniProt sequences, covering all eukaryotic proteomes. In addition, we describe a novel consensus annotation calculation and its related weighting scheme. The comparison between disorder information sources highlights how the MobiDB consensus captures the main features of intrinsic disorder and correlates well with manually curated datasets. Finally, we demonstrate the annotation of 13 eukaryotic model organisms through MobiDB's datasets, and of an example protein through the interactive user interface. CONCLUSIONS: MobiDB is a central resource for intrinsic disorder research, containing both experimental data and predictions. In the future it will be expanded to include additional information for all known proteins.
Tomás Di Domenico, Ian Walsh, Silvio C. E. Tosatto
BMC Bioinform.3
2012 MobiDB: a comprehensive database of intrinsic protein disorder annotations
abstract
MOTIVATION: Disordered protein regions are key to the function of numerous processes within an organism and to the determination of a protein's biological role. The most common source for protein disorder annotations, DisProt, covers only a fraction of the available sequences. Alternatively, the Protein Data Bank (PDB) has been mined for missing residues in X-ray crystallographic structures. Herein, we provide a centralized source for data on different flavours of disorder in protein structures, MobiDB, building on and expanding the content provided by already existing sources. In addition to the DisProt and PDB X-ray structures, we have added experimental information from NMR structures and five different flavours of two disorder predictors (ESpritz and IUpred). These are combined into a weighted consensus disorder used to classify disordered regions into flexible and constrained disorder. Users are encouraged to submit manual annotations through a submission form. MobiDB features experimental annotations for 17 285 proteins, covering the entire PDB and predictions for the SwissProt database, with 565 200 annotated sequences. Depending on the disorder flavour, 6-20% of the residues are predicted as disordered. AVAILABILITY: The database is freely available at http://mobidb.bio.unipd.it/. CONTACT: [email protected].
Tomás Di Domenico, Ian Walsh, Alberto J. M. Martin, Silvio C. E. Tosatto
Bioinform.4
2012 Bluues server: electrostatic properties of wild-type and mutated protein structures
abstract
MOTIVATION: Electrostatic calculations are an important tool for deciphering many functional mechanisms in proteins. Generalized Born (GB) models offer a fast and convenient computational approximation over other implicit solvent-based electrostatic models. Here we present a novel GB-based web server, using the program Bluues, to calculate numerous electrostatic features including pKa-values and surface potentials. The output is organized allowing both experts and beginners to rapidly sift the data. A novel feature of the Bluues server is that it explicitly allows to find electrostatic differences between wild-type and mutant structures. AVAILABILITY: The Bluues server, examples and extensive help files are available for non-commercial use at URL: http://protein.bio.unipd.it/bluues/.
Ian Walsh, Giovanni Minervini, Alessandra Corazza, Gennaro Esposito, Silvio C. E. Tosatto, Federico Fogolari
Bioinform.5
2012 ESpritz: accurate and fast prediction of protein disorder
abstract
MOTIVATION: Intrinsically disordered regions are key for the function of numerous proteins, and the scant available experimental annotations suggest the existence of different disorder flavors. While efficient predictions are required to annotate entire genomes, most existing methods require sequence profiles for disorder prediction, making them cumbersome for high-throughput applications. RESULTS: In this work, we present an ensemble of protein disorder predictors called ESpritz. These are based on bidirectional recursive neural networks and trained on three different flavors of disorder, including a novel NMR flexibility predictor. ESpritz can produce fast and accurate sequence-only predictions, annotating entire genomes in the order of hours on a single processor core. Alternatively, a slower but slightly more accurate ESpritz variant using sequence profiles can be used for applications requiring maximum performance. Two levels of prediction confidence allow either to maximize reasonable disorder detection or to limit expected false positives to 5%. ESpritz performs consistently well on the recent CASP9 data, reaching a S(w) measure of 54.82 and area under the receiver operator curve of 0.856. The fast predictor is four orders of magnitude faster and remains better than most publicly available CASP9 methods, making it ideal for genomic scale predictions. CONCLUSIONS: ESpritz predicts three flavors of disorder at two distinct false positive rates, either with a fast or slower and slightly more accurate approach. Given its state-of-the-art performance, it can be especially useful for high-throughput applications. AVAILABILITY: Both a web server for high-throughput analysis and a Linux executable version of ESpritz are available from: http://protein.bio.unipd.it/espritz/.
Ian Walsh, Alberto J. M. Martin, Tomás Di Domenico, Silvio C. E. Tosatto
Bioinform.4
2012 RAPHAEL: recognition, periodicity and insertion assignment of solenoid protein structures
abstract
MOTIVATION: Repeat proteins form a distinct class of structures where folding is greatly simplified. Several classes have been defined, with solenoid repeats of periodicity between ca. 5 and 40 being the most challenging to detect. Such proteins evolve quickly and their periodicity may be rapidly hidden at sequence level. From a structural point of view, finding solenoids may be complicated by the presence of insertions or multiple domains. To the best of our knowledge, no automated methods are available to characterize solenoid repeats from structure. RESULTS: Here we introduce RAPHAEL, a novel method for the detection of solenoids in protein structures. It reliably solves three problems of increasing difficulty: (1) recognition of solenoid domains, (2) determination of their periodicity and (3) assignment of insertions. RAPHAEL uses a geometric approach mimicking manual classification, producing several numeric parameters that are optimized for maximum performance. The resulting method is very accurate, with 89.5% of solenoid proteins and 97.2% of non-solenoid proteins correctly classified. RAPHAEL periodicities have a Spearman correlation coefficient of 0.877 against the manually established ones. A baseline algorithm for insertion detection in identified solenoids has a Q(2) value of 79.8%, suggesting room for further improvement. RAPHAEL finds 1931 highly confident repeat structures not previously annotated as solenoids in the Protein Data Bank records.
Ian Walsh, Francesco Sirocco, Giovanni Minervini, Tomás Di Domenico, Carlo Ferrari, Silvio C. E. Tosatto
Bioinform.6
2011 RING: networking interacting residues, evolutionary information and energetics in protein structures
abstract
MOTIVATION: Residue interaction networks (RINs) have been used in the literature to describe the protein 3D structure as a graph where nodes represent residues and edges physico-chemical interactions, e.g. hydrogen bonds or van-der-Waals contacts. Topological network parameters can be calculated over RINs and have been correlated with various aspects of protein structure and function. Here we present a novel web server, RING, to construct physico-chemically valid RINs interactively from PDB files for subsequent visualization in the Cytoscape platform. The additional structure-based parameters secondary structure, solvent accessibility and experimental uncertainty can be combined with information regarding residue conservation, mutual information and residue-based energy scoring functions. Different visualization styles are provided to facilitate visualization and standard plugins can be used to calculate topological parameters in Cytoscape. A sample use case analyzing the active site of glutathione peroxidase is presented. AVAILABILITY: The RING server, supplementary methods, examples and tutorials are available for non-commercial use at URL: http://protein.bio.unipd.it/ring/.
Alberto J. M. Martin, Michele Vidotto, Filippo Boscariol, Tomás Di Domenico, Ian Walsh, Silvio C. E. Tosatto
Bioinform.6
2010 MOBI: a web server to define and visualize structural mobility in NMR protein ensembles
abstract
MOTIVATION: MOBI is a web server for the identification of structurally mobile regions in NMR protein ensembles. It provides a binary mobility definition that is analogous to the commonly used definition of intrinsic disorder in X-ray crystallographic structures. At least three different use cases can be envisaged: (i) visualization of NMR mobility for structural analysis; (ii) definition of regions for reliable comparative modelling in protein structure prediction and (iii) definition of mobility in analogy to intrinsic disorder. MOBI uses structural superposition and local conformational differences to derive a robust binary mobility definition that is in excellent agreement with the manually curated definition used in the CASP8 experiment for intrinsic disorder in NMR structure. The output includes mobility-coloured PDB files, mobility plots and a FASTA formatted sequence file summarizing the mobility results. AVAILABILITY: The MOBI server and supplementary methods are available for non-commercial use at URL: http://protein.bio.unipd.it/mobi/.
Alberto J. M. Martin, Ian Walsh, Silvio C. E. Tosatto
Bioinform.3
2010 FRASS: the web-server for RNA structural comparison
abstract
BACKGROUND: The impressive increase of novel RNA structures, during the past few years, demands automated methods for structure comparison. While many algorithms handle only small motifs, few techniques, developed in recent years, (ARTS, DIAL, SARA, SARSA, and LaJolla) are available for the structural comparison of large and intact RNA molecules. RESULTS: The FRASS web-server represents a RNA chain with its Gauss integrals and allows one to compare structures of RNA chains and to find similar entries in a database derived from the Protein Data Bank. We observed that FRASS scores correlate well with the ARTS and LaJolla similarity scores. Moreover, the-web server can also reproduce satisfactorily the DARTS classification of RNA 3D structures and the classification of the SCOR functions that was obtained by the SARA method. CONCLUSIONS: The FRASS web-server can be easily used to detect relationships among RNA molecules and to scan efficiently the rapidly enlarging structural databases.
Svetlana Kirillova, Silvio C. E. Tosatto, Oliviero Carugo
BMC Bioinform.2
2009 REPETITA: detection and discrimination of the periodicity of protein solenoid repeats by discrete Fourier transform
abstract
MOTIVATION: Proteins with solenoid repeats evolve more quickly than non-repetitive ones and their periodicity may be rapidly hidden at sequence level, while still evident in structure. In order to identify these repeats, we propose here a novel method based on a metric characterizing amino-acid properties (polarity, secondary structure, molecular volume, codon diversity, electric charge) using five previously derived numerical functions. RESULTS: The five spectra of the candidate sequences coding for structural repeats, obtained by Discrete Fourier Transform (DFT), show common features allowing determination of repeat periodicity with excellent results. Moreover it is possible to introduce a phase space parameterized by two quantities related to the Fourier spectra which allow for a clear distinction between a non-homologous set of globular proteins and proteins with solenoid repeats. The DFT method is shown to be competitive with other state of the art methods in the detection of solenoid structures, while improving its performance especially in the identification of periodicities, since it is able to recognize the actual repeat length in most cases. Moreover it highlights the relevance of local structural propensities in determining solenoid repeats. AVAILABILITY: A web tool implementing the algorithm presented in the article (REPETITA) is available with additional details on the data sets at the URL: http://protein.bio.unipd.it/repetita/.
Luca Marsella, Francesco Sirocco, Antonio Trovato, Flavio Seno, Silvio C. E. Tosatto
Bioinform.5
2008 TESE: generating specific protein structure test set ensembles
abstract
UNLABELLED: TESE is a web server for the generation of test sets of protein sequences and structures fulfilling a number of different criteria. At least three different use cases can be envisaged: (i) benchmarking of novel methods; (ii) test sets tailored for special needs and (iii) extending available datasets. The CATH structure classification is used to control structural/sequence redundancy and a variety of structural quality parameters can be used to interactively select protein subsets with specific characteristics, e.g. all X-ray structures of alpha-helical repeat proteins with more than 120 residues and resolution <2.0 A. The output includes FASTA-formatted sequences, PDB files and a clickable HTML index file containing images of the selected proteins. Multiple subsets for cross-validation are also supported. AVAILABILITY: The TESE server is available for non-commercial use at URL: http://protein.bio.unipd.it/tese/.
Francesco Sirocco, Silvio C. E. Tosatto
Bioinform.2
2007 TAP score: torsion angle propensity normalization applied to local protein structure evaluation
abstract
BACKGROUND: Experimentally determined protein structures may contain errors and require validation. Conformational criteria based on the Ramachandran plot are mainly used to distinguish between distorted and adequately refined models. While the readily available criteria are sufficient to detect totally wrong structures, establishing the more subtle differences between plausible structures remains more challenging. RESULTS: A new criterion, called TAP score, measuring local sequence to structure fitness based on torsion angle propensities normalized against the global minimum and maximum is introduced. It is shown to be more accurate than previous methods at estimating the validity of a protein model in terms of commonly used experimental quality parameters on two test sets representing the full PDB database and a subset of obsolete PDB structures. Highly selective TAP thresholds are derived to recognize over 90% of the top experimental structures in the absence of experimental information. Both a web server and an executable version of the TAP score are available at http://protein.cribi.unipd.it/tap/. CONCLUSION: A novel procedure for energy normalization (TAP) has significantly improved the possibility to recognize the best experimental structures. It will allow the user to more reliably isolate problematic structures in the context of automated experimental structure determination.
Silvio C. E. Tosatto, Roberto Battistutta
BMC Bioinform.1
2006 Improving the quality of protein structure models by selecting from alignment alternatives
abstract
BACKGROUND: In the area of protein structure prediction, recently a lot of effort has gone into the development of Model Quality Assessment Programs (MQAPs). MQAPs distinguish high quality protein structure models from inferior models. Here, we propose a new method to use an MQAP to improve the quality of models. With a given target sequence and template structure, we construct a number of different alignments and corresponding models for the sequence. The quality of these models is scored with an MQAP and used to choose the most promising model. An SVM-based selection scheme is suggested for combining MQAP partial potentials, in order to optimize for improved model selection. RESULTS: The approach has been tested on a representative set of proteins. The ability of the method to improve models was validated by comparing the MQAP-selected structures to the native structures with the model quality evaluation program TM-score. Using the SVM-based model selection, a significant increase in model quality is obtained (as shown with a Wilcoxon signed rank test yielding p-values below 10(-15)). The average increase in TMscore is 0.016, the maximum observed increase in TM-score is 0.29. CONCLUSION: In template-based protein structure prediction alignment is known to be a bottleneck limiting the overall model quality. Here we show that a combination of systematic alignment variation and modern model scoring functions can significantly improve the quality of alignment-based models.
Ingolf Sommer, Stefano Toppo, Oliver Sander, Thomas Lengauer, Silvio C. E. Tosatto
BMC Bioinform.5
2005 The SSEA server for protein secondary structure alignment
abstract
SUMMARY: We present a web server that computes alignments of protein secondary structures. The server supports both performing pairwise alignments and searching a secondary structure against a library of domain folds. It can calculate global and local secondary structure element alignments. A combination of local and global alignment steps can be used to search for domains inside the query sequence or help in the discrimination of novel folds. Both the SCOP and PDB fold libraries, clustered at 95 and 40% sequence identity, are available for alignment. AVAILABILITY: The web server interface is freely accessible to academic users at http://protein.cribi.unipd.it/ssea/. The executable version and benchmarking data are available from the same web page.
Paolo Fontana, Eckart Bindewald, Stefano Toppo, Riccardo Velasco, Giorgio Valle, Silvio C. E. Tosatto
Bioinform.6
2005 A decoy set for the thermostable subdomain from chicken villin headpiece, comparison of different free energy estimators
abstract
BACKGROUND: Estimators of free energies are routinely used to judge the quality of protein structural models. As these estimators still present inaccuracies, they are frequently evaluated by discriminating native or native-like conformations from large ensembles of so-called decoy structures. RESULTS: A decoy set is obtained from snapshots taken from 5 long (100 ns) molecular dynamics (MD) simulations of the thermostable subdomain from chicken villin headpiece. An evaluation of the energy of the decoys is given using: i) a residue based contact potential supplemented by a term for the quality of dihedral angles; ii) a recently introduced combination of four statistical scoring functions for model quality estimation (FRST); iii) molecular mechanics with solvation energy estimated either according to the generalized Born surface area (GBSA) or iv) the Poisson-Boltzmann surface area (PBSA) method. CONCLUSION: The decoy set presented here has the following features which make it attractive for testing energy scoring functions:1) it covers a broad range of RMSD values (from less than 2.0 A to more than 12 A);2) it has been obtained from molecular dynamics trajectories, starting from different non-native-like conformations which have diverse behaviour, with secondary structure elements correctly or incorrectly formed, and in one case folding to a native-like structure. This allows not only for scoring of static structures, but also for studying, using free energy estimators, the kinetics of folding;3) all structures have been obtained from accurate MD simulations in explicit solvent and after molecular mechanics (MM) energy minimization using an implicit solvent method. The quality of the covalent structure therefore does not suffer from steric or covalent problems. The statistical and physical effective energy functions tested on the set behave differently when native simulation snapshots are included or not in the set and when averaging over the trajectory is performed.
Federico Fogolari, Silvio C. E. Tosatto, Giorgio Colombo 0001
BMC Bioinform.2