Damiano Piovesan

dblp:05/11248 · DBLP profile ↗
← Back
16ranked-venue papers
6as first author
6since 2021 · last 2025
0000-0001-8210-2390ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 15 · 5 first-author · 6 since 2021Systems, architecture and hardware · 1 · 1 first-author
YearPublicationVenuePosition
2025 MobiDB-lite 4.0: faster prediction of intrinsic protein disorder and structural compactness
abstract
MOTIVATION: In recent years, many disorder predictors have been developed to identify intrinsically disordered regions (IDRs) in proteins, achieving high accuracy. However, it may be difficult to interpret differences in predictions across methods. Consensus methods offer a simple solution, highlighting reliable predictions while filtering out uncertain positions. Here, we present a new version of MobiDB-lite, a consensus method designed to predict long IDRs and classify them based on compositional biases and conformational properties. RESULTS: MobiDB-lite 4.0 pipeline was optimized to be ten times faster than the previous version. It now provides compactness annotations based on predicted apparent scaling exponent. The newly added features and disorder subclassifications allow the users to get a comprehensive insight into the protein's function and characteristics. MobiDB-lite 4.0 is integrated into the MobiDB and DisProt databases. A version without the compactness predictor is integrated into InterProScan, propagating MobiDB-lite annotations to UniProtKB. AVAILABILITY AND IMPLEMENTATION: The MobiDB-lite 4.0 source code and a Docker container are available from the GitHub repository: https://github.com/BioComputingUP/MobiDB-lite.
Mahta Mehdiabadi, Matthias Blum, Giulio Tesei, Sören von Bülow, Kresten Lindorff-Larsen, Silvio C. E. Tosatto, Damiano Piovesan
Bioinform.7
2025 GeomeTRe: accurate calculation of geometrical descriptors of tandem repeat proteins
abstract
MOTIVATION: Structured tandem repeat proteins (STRPs) are characterized by preserved structural motifs arranged in a modular way. The structural and functional diversity of STRPs makes them particularly important for studying evolution and novel structure-function relationships, and ultimately for designing new synthetic proteins with specific functions. One crucial aspect of their classification is the estimation of geometrical parameters, which can provide better insight into their properties and the relationship between the spatial arrangement of repeated units and protein function. Calculating geometric descriptors for STRPs is challenging because naturally occurring repeats are not "perfect" and often contain insertions and deletions. Existing tools for predicting structural symmetry work well on simple cases but often fail for most natural proteins. RESULTS: Here, we present GeomeTRe, an algorithm that calculates geometrical descriptors such as curvature (yaw), twist (roll), and pitch for a protein structure with known repeat unit positions. The algorithm simulates the movement of consecutive units, identifies rotational axes, and calculates the corresponding Tait-Bryan angles. GeomeTRe's parameters can enhance STRP annotation and classification by identifying variations in geometric arrangements among different functional groups. The package is fast and suitable for processing large protein structure datasets when repeat region information (e.g. from RepeatsDB) is available. AVAILABILITY AND IMPLEMENTATION: GeomeTRe is available as a Python package; source code and documentation can be found at https://github.com/BioComputingUP/GeomeTRe.
Zarifa Osmanli, Elisa Ferrero, Alexander Miguel Monzon, Silvio C. E. Tosatto, Damiano Piovesan
Bioinform.5
2024 STRPsearch: fast detection of structured tandem repeat proteins
abstract
MOTIVATION: Structured Tandem Repeats Proteins (STRPs) constitute a subclass of tandem repeats characterized by repetitive structural motifs. These proteins exhibit distinct secondary structures that form repetitive tertiary arrangements, often resulting in large molecular assemblies. Despite highly variable sequences, STRPs can perform important and diverse biological functions, maintaining a consistent structure with a variable number of repeat units. With the advent of protein structure prediction methods, millions of 3D models of proteins are now publicly available. However, automatic detection of STRPs remains challenging with current state-of-the-art tools due to their lack of accuracy and long execution times, hindering their application on large datasets. In most cases, manual curation remains the most accurate method for detecting and classifying STRPs, making it impracticable to annotate millions of structures. RESULTS: We introduce STRPsearch, a novel tool for the rapid identification, classification, and mapping of STRPs. Leveraging manually curated entries from RepeatsDB as the known conformational space of STRPs, STRPsearch uses the latest advances in structural alignment for a fast and accurate detection of repeated structural motifs in proteins, followed by an innovative approach to map units and insertions through the generation of TM-score profiles. STRPsearch is highly scalable, efficiently processing large datasets, and can be applied to both experimental structures and predicted models. In addition, it demonstrates superior performance compared to existing tools, offering researchers a reliable and comprehensive solution for STRP analysis across diverse proteomes. AVAILABILITY AND IMPLEMENTATION: STRPsearch is coded in Python. All scripts and associated documentation are available from: https://github.com/BioComputingUP/STRPsearch.
Soroush Mozaffari, Paula Nazarena Arrías, Damiano Clementel, Damiano Piovesan, Carlo Ferrari, Silvio C. E. Tosatto, Alexander Miguel Monzon
Bioinform.4
2023 RING-PyMOL: residue interaction networks of structural ensembles and molecular dynamics
abstract
RING-PyMOL is a plugin for PyMOL providing a set of analysis tools for structural ensembles and molecular dynamic simulations. RING-PyMOL combines residue interaction networks, as provided by the RING software, with structural clustering to enhance the analysis and visualization of the conformational complexity. It combines precise calculation of non-covalent interactions with the power of PyMOL to manipulate and visualize protein structures. The plugin identifies and highlights correlating contacts and interaction patterns that can explain structural allostery, active sites, and structural heterogeneity connected with molecular function. It is easy to use and extremely fast, processing and rendering hundreds of models and long trajectories in seconds. RING-PyMOL generates a number of interactive plots and output files for use with external tools. The underlying RING software has been improved extensively. It is 10 times faster, can process mmCIF files and it identifies typed interactions also for nucleic acids. AVAILABILITY AND IMPLEMENTATION: https://github.com/BioComputingUP/ring-pymol.
Alessio Del Conte, Alexander Miguel Monzon, Damiano Clementel, Giorgia F. Camagni, Giovanni Minervini, Silvio C. E. Tosatto, Damiano Piovesan
Bioinform.7
2022 ProSeqViewer: an interactive, responsive and efficient TypeScript library for visualization of sequences and alignments in web applications
abstract
SUMMARY: Biological data is ever-increasing in amount and complexity. The mapping of this data to biological entities such as nucleotide and amino acid sequences supports biological data analysis, classification and prediction. Sequence alignments and comparison allow the transfer of knowledge to evolutionary-related entities, the mapping of functional domains, the identification of binding and modification sites. To support these types of studies, we developed ProSeqViewer, a tool to visualize annotation on single sequences and multiple sequence alignments. This state-of-the-art multifunctional library was developed as a modular component to be integrated into static or dynamic web resources and support intuitive visualization of sequence features. ProseSeqViewer is extremely lightweight, fast, interactive, dynamic, responsive and works at any screen size. It generates pure HTML which is compatible with any browser and operating system. ProSeqViewer can exchange events with other visualization components and is already used by multiple biological databases. AVAILABILITY AND IMPLEMENTATION: ProSeqViewer is an open-source TypeScript library compatible with state-of-the-art website environments. The source code and an extensive documentation including use cases are available from the URL: https://github.com/BioComputingUP/ProSeqViewer.
Martina Bevilacqua, Lisanna Paladin, Silvio C. E. Tosatto, Damiano Piovesan
Bioinform.4
2021 MobiDB-lite 3.0: fast consensus annotation of intrinsic disorder flavors in proteins
abstract
MOTIVATION: The earlier version of MobiDB-lite is currently used in large-scale proteome annotation platforms to detect intrinsic disorder. However, new theoretical models allow for the classification of intrinsically disordered regions into subtypes from sequence features associated with specific polymeric properties or compositional bias. RESULTS: MobiDB-lite 3.0 maintains its previous speed and performance but also provides a finer classification of disorder by identifying regions with characteristics of polyolyampholytes, positive or negative polyelectrolytes, low-complexity regions or enriched in cysteine, proline or glycine or polar residues. Subregions are abundantly detected in IDRs of the human proteome. The new version of MobiDB-lite represents a new step for the proteome level analysis of protein disorder. AVAILABILITY AND IMPLEMENTATION: Both the MobiDB-lite 3.0 source code and a docker container are available from the GitHub repository: https://github.com/BioComputingUP/MobiDB-lite.
Marco Necci, Damiano Piovesan, Damiano Clementel, Zsuzsanna Dosztányi, Silvio C. E. Tosatto
Bioinform.2
2020 The Feature-Viewer: a visualization tool for positional annotations on a sequence
abstract
SUMMARY: The Feature-Viewer is a lightweight library for the visualization of biological data mapped to a protein or nucleotide sequence. It is designed for ease of use while allowing for a full customization. The library is already used by several biological data resources and allows intuitive visual mapping of a full spectra of sequence features for different usages. AVAILABILITY AND IMPLEMENTATION: The Feature-Viewer is open source, compatible with state-of-the-art development technologies and responsive, also for mobile viewing. Documentation and usage examples are available online.
Lisanna Paladin, Mathieu Schaeffer, Pascale Gaudet, Monique Zahn-Zabal, Pierre-André Michel, Damiano Piovesan, Silvio C. E. Tosatto, Amos Bairoch
Bioinform.6
2020 Assessing predictors for new post translational modification sites: A case study on hydroxylation
abstract
Post-translational modification (PTM) sites have become popular for predictor development. However, with the exception of phosphorylation and a handful of other examples, PTMs suffer from a limited number of available training examples and sparsity in protein sequences. Here, proline hydroxylation is taken as an example to compare different methods and evaluate their performance on new experimentally determined sites. As a guide for effective experimental design, predictors require both high specificity and sensitivity. However, the self-reported performance may often not be indicative of prediction quality and detection of new sites is not guaranteed. We have benchmarked seven published hydroxylation site predictors on two newly constructed independent datasets. The self-reported performance is found to widely overestimate the real accuracy measured on independent datasets. No predictor performs better than random on new examples, indicating the refined models do not sufficiently generalize to detect new sites. The number of false positives is high and precision low, in particular for non-collagen proteins whose motifs are not conserved. As hydroxylation site predictors do not generalize for new data, caution is advised when using PTM predictors in the absence of independent evaluations, in particular for highly specific sites involved in signalling.
Damiano Piovesan, András Hatos, Giovanni Minervini, Federica Quaglia, Alexander Miguel Monzon, Silvio C. E. Tosatto
PLoS Comput. Biol.1
2018 A comprehensive assessment of long intrinsic protein disorder from the DisProt database
abstract
Motivation: Intrinsic disorder (ID), i.e. the lack of a unique folded conformation at physiological conditions, is a common feature for many proteins, which requires specialized biochemical experiments that are not high-throughput. Missing X-ray residues from the PDB have been widely used as a proxy for ID when developing computational methods. This may lead to a systematic bias, where predictors deviate from biologically relevant ID. Large benchmarking sets on experimentally validated ID are scarce. Recently, the DisProt database has been renewed and expanded to include manually curated ID annotations for several hundred new proteins. This provides a large benchmark set which has not yet been used for training ID predictors. Results: Here, we describe the first systematic benchmarking of ID predictors on the new DisProt dataset. In contrast to previous assessments based on missing X-ray data, this dataset contains mostly long ID regions and a significant amount of fully ID proteins. The benchmarking shows that ID predictors work quite well on the new dataset, especially for long ID segments. However, a large fraction of ID still goes virtually undetected and the ranking of methods is different than for PDB data. In particular, many predictors appear to confound ID and regions outside X-ray structures. This suggests that the ID prediction methods capture different flavors of disorder and can benefit from highly accurate curated examples. Availability and implementation: The raw data used for the evaluation are available from URL: http://www.disprot.org/assessment/. Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Marco Necci, Damiano Piovesan, Zsuzsanna Dosztányi, Peter Tompa, Silvio C. E. Tosatto
Bioinform.2
2018 Mobi 2.0: an improved method to define intrinsic disorder, mobility and linear binding regions in protein structures
abstract
Motivation: The structures contained in the Protein Data Bank (PDB) database are of paramount importance to define our knowledge of folded proteins. While providing mainly circumstantial evidence, PDB data is also increasingly used to define the lack of unique structure, represented by mobile regions and even intrinsic disorder (ID). However, alternative definitions are used by different authors and potentially limit the generality of the analyses being carried out. Results: Here we present Mobi 2.0, a completely re-written version of the Mobi software for the determination of mobile and potentially disordered regions from PDB structures. Mobi 2.0 provides robust definitions of mobility based on four main sources of information: (i) missing residues, (ii) residues with high temperature factors, (iii) mobility between different models of the same structure and (iv) binding to another protein or nucleotide chain. Mobi 2.0 is well suited to aggregate information across different PDB structures for the same UniProt protein sequence, providing consensus annotations. The software is expected to standardize the treatment of mobility, allowing an easier comparison across different studies related to ID. Availability: Mobi 2.0 provides the structure-based annotation for the MobiDB database. The software is available from URL http://protein.bio.unipd.it/mobi2/. Contact: [email protected].
Damiano Piovesan, Silvio C. E. Tosatto
Bioinform.1
2017 MobiDB-lite: fast and highly specific consensus prediction of intrinsic disorder in proteins
abstract
Motivation: Intrinsic disorder (ID) is established as an important feature of protein sequences. Its use in proteome annotation is however hampered by the availability of many methods with similar performance at the single residue level, which have mostly not been optimized to predict long ID regions of size comparable to domains. Results: Here, we have focused on providing a single consensus-based prediction, MobiDB-lite, optimized for highly specific (i.e. few false positive) predictions of long disorder. The method uses eight different predictors to derive a consensus which is then filtered for spurious short predictions. Consensus prediction is shown to outperform the single methods when annotating long ID regions. MobiDB-lite can be useful in large-scale annotation scenarios and has indeed already been integrated in the MobiDB, DisProt and InterPro databases. Availability and Implementation: MobiDB-lite is available as part of the MobiDB database from URL: http://mobidb.bio.unipd.it/. An executable can be downloaded from URL: http://protein.bio.unipd.it/mobidblite/. Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Marco Necci, Damiano Piovesan, Zsuzsanna Dosztányi, Silvio C. E. Tosatto
Bioinform.2
2017 FELLS: fast estimator of latent local structure
abstract
MOTIVATION: The behavior of a protein is encoded in its sequence, which can be used to predict distinct features such as secondary structure, intrinsic disorder or amphipathicity. Integrating these and other features can help explain the context-dependent behavior of proteins. However, most tools focus on a single aspect, hampering a holistic understanding of protein structure. Here, we present Fast Estimator of Latent Local Structure (FELLS) to visualize structural features from the protein sequence. FELLS provides disorder, aggregation and low complexity predictions as well as estimated local propensities including amphipathicity. A novel fast estimator of secondary structure (FESS) is also trained to provide a fast response. The calculations required for FELLS are extremely fast and suited for large-scale analysis while providing a detailed analysis of difficult cases. AVAILABILITY AND IMPLEMENTATION: The FELLS web server is available from URL: http://protein.bio.unipd.it/fells/ . The server also exposes RESTful functionality allowing programmatic prediction requests. An executable version of FESS for Linux can be downloaded from URL: protein.bio.unipd.it/download/. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Damiano Piovesan, Ian Walsh, Giovanni Minervini, Silvio C. E. Tosatto
Bioinform.1
2015 The Victor C++ library for protein representation and advanced manipulation
abstract
MOTIVATION: Protein sequence and structure representation and manipulation require dedicated software libraries to support methods of increasing complexity. Here, we describe the VIrtual Constrution TOol for pRoteins (Victor) C++ library, an open source platform dedicated to enabling inexperienced users to develop advanced tools and gathering contributions from the community. The provided application examples cover statistical energy potentials, profile-profile sequence alignments and ab initio loop modeling. Victor was used over the last 15 years in several publications and optimized for efficiency. It is provided as a GitHub repository with source files and unit tests, plus extensive online documentation, including a Wiki with help files and tutorials, examples and Doxygen documentation. AVAILABILITY AND IMPLEMENTATION: The C++ library and online documentation, distributed under a GPL license are available from URL: http://protein.bio.unipd.it/victor/.
Layla Hirsh, Damiano Piovesan, Manuel Giollo, Carlo Ferrari, Silvio C. E. Tosatto
Bioinform.2
2013 How to inherit statistically validated annotation within BAR+ protein clusters
abstract
BACKGROUND: In the genomic era a key issue is protein annotation, namely how to endow protein sequences, upon translation from the corresponding genes, with structural and functional features. Routinely this operation is electronically done by deriving and integrating information from previous knowledge. The reference database for protein sequences is UniProtKB divided into two sections, UniProtKB/TrEMBL which is automatically annotated and not reviewed and UniProtKB/Swiss-Prot which is manually annotated and reviewed. The annotation process is essentially based on sequence similarity search. The question therefore arises as to which extent annotation based on transfer by inheritance is valuable and specifically if it is possible to statistically validate inherited features when little homology exists among the target sequence and its template(s). RESULTS: In this paper we address the problem of annotating protein sequences in a statistically validated manner considering as a reference annotation resource UniProtKB. The test case is the set of 48,298 proteins recently released by the Critical Assessment of Function Annotations (CAFA) organization. We show that we can transfer after validation, Gene Ontology (GO) terms of the three main categories and Pfam domains to about 68% and 72% of the sequences, respectively. This is possible after alignment of the CAFA sequences towards BAR+, our annotation resource that allows discriminating among statistically validated and not statistically validated annotation. By comparing with a direct UniProtKB annotation, we find that besides validating annotation of some 78% of the CAFA set, we assign new and statistically validated annotation to 14.8% of the sequences and find new structural templates for about 25% of the chains, half of which share less than 30% sequence identity to the corresponding template/s. CONCLUSION: Inheritance of annotation by transfer generally requires a careful selection of the identity value among the target and the template in order to transfer structural and/or functional features. Here we prove that even distantly remote homologs can be safely endowed with structural templates and GO and/or Pfam terms provided that annotation is done within clusters collecting cluster-related protein sequences and where a statistical validation of the shared structural and functional features is possible.
Damiano Piovesan, Pier Luigi Martelli, Piero Fariselli, Giuseppe Profiti, Andrea Zauli, Ivan Rossi, Rita Casadio
BMC Bioinform.1
2013 Extended and Robust Protein Sequence Annotation over Conservative Nonhierarchical Clusters: The Case Study of the ABC Transporters
abstract
Genome annotation is one of the most important issues in the genomic era. The exponential growth rate of newly sequenced genomes and proteomes urges the development of fast and reliable annotation methods, suited to exploit all the information available in curated databases of protein sequences and structures. To this aim we developed BAR+, the Bologna Annotation Resource. 1 The basic notion is that sequences with high identity value to a counterpart can inherit the same function/s and structure, if available. As a case study we describe how the ATP-binding domain of the ABC transporters can be found and modeled in over 30,000 new sequences not annotated before. We also mapped into BAR+ all the ABC transporters listed in the Transporter Classification DataBase 2 and found that within our environment annotation could be extended to another 256,866 sequences.
Damiano Piovesan, Giuseppe Profiti, Pier Luigi Martelli, Piero Fariselli, Rita Casadio
ACM J. Emerg. Technol. Comput. Syst.1
2012 The human "magnesome": detecting magnesium binding sites on human proteins
abstract
BACKGROUND: Magnesium research is increasing in molecular medicine due to the relevance of this ion in several important biological processes and associated molecular pathogeneses. It is still difficult to predict from the protein covalent structure whether a human chain is or not involved in magnesium binding. This is mainly due to little information on the structural characteristics of magnesium binding sites in proteins and protein complexes. Magnesium binding features, differently from those of other divalent cations such as calcium and zinc, are elusive. Here we address a question that is relevant in protein annotation: how many human proteins can bind Mg2+? Our analysis is performed taking advantage of the recently implemented Bologna Annotation Resource (BAR-PLUS), a non hierarchical clustering method that relies on the pair wise sequence comparison of about 14 millions proteins from over 300.000 species and their grouping into clusters where annotation can safely be inherited after statistical validation. RESULTS: After cluster assignment of the latest version of the human proteome, the total number of human proteins for which we can assign putative Mg binding sites is 3,751. Among these proteins, 2,688 inherit annotation directly from human templates and 1,063 inherit annotation from templates of other organisms. Protein structures are highly conserved inside a given cluster. Transfer of structural properties is possible after alignment of a given sequence with the protein structures that characterise a given cluster as obtained with a Hidden Markov Model (HMM) based procedure. Interestingly a set of 370 human sequences inherit Mg2+ binding sites from templates sharing less than 30% sequence identity with the template. CONCLUSION: We describe and deliver the "human magnesome", a set of proteins of the human proteome that inherit putative binding of magnesium ions. With our BAR-hMG, 251 clusters including 1,341 magnesium binding protein structures corresponding to 387 sequences are sufficient to annotate some 13,689 residues in 3,751 human sequences as "magnesium binding". Protein structures act therefore as three dimensional seeds for structural and functional annotation of human sequences. The data base collects specifically all the human proteins that can be annotated according to our procedure as "magnesium binding", the corresponding structures and BAR+ clusters from where they derive the annotation (http://bar.biocomp.unibo.it/mg).
Damiano Piovesan, Giuseppe Profiti, Pier Luigi Martelli, Rita Casadio
BMC Bioinform.1