Castrense Savojardo

dblp:09/9970 · DBLP profile ↗
← Back
20ranked-venue papers
15as first author
4since 2021 · last 2025
0000-0002-7359-0633ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 20 · 15 first-author · 4 since 2021
YearPublicationVenuePosition
2025 DDGemb: predicting protein stability change upon single- and multi-point variations with embeddings and deep learning
abstract
MOTIVATION: The knowledge of protein stability upon residue variation is an important step for functional protein design and for understanding how protein variants can promote disease onset. Computational methods are important to complement experimental approaches and allow a fast screening of large datasets of variations. RESULTS: In this work, we present DDGemb, a novel method combining protein language model embeddings and transformer architectures to predict protein ΔΔG upon both single- and multi-point variations. DDGemb has been trained on a high-quality dataset derived from literature and tested on available benchmark datasets of single- and multi-point variations. DDGemb performs at the state of the art in both single- and multi-point variations. AVAILABILITY AND IMPLEMENTATION: DDGemb is available as web server at https://ddgemb.biocomp.unibo.it. Datasets used in this study are available at https://ddgemb.biocomp.unibo.it/datasets.
Castrense Savojardo, Matteo Manfredi, Pier Luigi Martelli, Rita Casadio
Bioinform.1
2023 CoCoNat: a novel method based on deep learning for coiled-coil prediction
abstract
MOTIVATION: Coiled-coil domains (CCD) are widespread in all organisms and perform several crucial functions. Given their relevance, the computational detection of CCD is very important for protein functional annotation. State-of-the-art prediction methods include the precise identification of CCD boundaries, the annotation of the typical heptad repeat pattern along the coiled-coil helices as well as the prediction of the oligomerization state. RESULTS: In this article, we describe CoCoNat, a novel method for predicting coiled-coil helix boundaries, residue-level register annotation, and oligomerization state. Our method encodes sequences with the combination of two state-of-the-art protein language models and implements a three-step deep learning procedure concatenated with a Grammatical-Restrained Hidden Conditional Random Field for CCD identification and refinement. A final neural network predicts the oligomerization state. When tested on a blind test set routinely adopted, CoCoNat obtains a performance superior to the current state-of-the-art both for residue-level and segment-level CCD. CoCoNat significantly outperforms the most recent state-of-the-art methods on register annotation and prediction of oligomerization states. AVAILABILITY AND IMPLEMENTATION: CoCoNat web server is available at https://coconat.biocomp.unibo.it. Standalone version is available on GitHub at https://github.com/BolognaBiocomp/coconat.
Giovanni Madeo, Castrense Savojardo, Matteo Manfredi, Pier Luigi Martelli, Rita Casadio
Bioinform.2
2022 E-SNPs&GO: embedding of protein sequence and function improves the annotation of human pathogenic variants
abstract
MOTIVATION: The advent of massive DNA sequencing technologies is producing a huge number of human single-nucleotide polymorphisms occurring in protein-coding regions and possibly changing their sequences. Discriminating harmful protein variations from neutral ones is one of the crucial challenges in precision medicine. Computational tools based on artificial intelligence provide models for protein sequence encoding, bypassing database searches for evolutionary information. We leverage the new encoding schemes for an efficient annotation of protein variants. RESULTS: E-SNPs&GO is a novel method that, given an input protein sequence and a single amino acid variation, can predict whether the variation is related to diseases or not. The proposed method adopts an input encoding completely based on protein language models and embedding techniques, specifically devised to encode protein sequences and GO functional annotations. We trained our model on a newly generated dataset of 101 146 human protein single amino acid variants in 13 661 proteins, derived from public resources. When tested on a blind set comprising 10 266 variants, our method well compares to recent approaches released in literature for the same task, reaching a Matthews Correlation Coefficient score of 0.72. We propose E-SNPs&GO as a suitable, efficient and accurate large-scale annotator of protein variant datasets. AVAILABILITY AND IMPLEMENTATION: The method is available as a webserver at https://esnpsandgo.biocomp.unibo.it. Datasets and predictions are available at https://esnpsandgo.biocomp.unibo.it/datasets. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Matteo Manfredi, Castrense Savojardo, Pier Luigi Martelli, Rita Casadio
Bioinform.2
2021 On the critical review of five machine learning-based algorithms for predicting protein stability changes upon mutation
abstract
A review, recently published in this journal by Fang (2019), showed that methods trained for the prediction of protein stability changes upon mutation have a very critical bias: they neglect that a protein variation (A- > B) and its reverse (B- > A) must have the opposite value of the free energy difference (ΔΔGAB = - ΔΔGBA). In this letter, we complement the Fang's paper presenting a more general view of the problem. In particular, a machine learning-based method, published in 2015 (INPS), addressed the bias issue directly. We include the analysis of the missing method, showing that INPS is nearly insensitive to the addressed problem.
Castrense Savojardo, Pier Luigi Martelli, Rita Casadio, Piero Fariselli
Briefings Bioinform.1
2020 DeepMito: accurate prediction of protein sub-mitochondrial localization using convolutional neural networks
abstract
MOTIVATION: The correct localization of proteins in cell compartments is a key issue for their function. Particularly, mitochondrial proteins are physiologically active in different compartments and their aberrant localization contributes to the pathogenesis of human mitochondrial pathologies. Many computational methods exist to assign protein sequences to subcellular compartments such as nucleus, cytoplasm and organelles. However, a substantial lack of experimental evidence in public sequence databases hampered so far a finer grain discrimination, including also intra-organelle compartments. RESULTS: We describe DeepMito, a novel method for predicting protein sub-mitochondrial cellular localization. Taking advantage of powerful deep-learning approaches, such as convolutional neural networks, our method is able to achieve very high prediction performances when discriminating among four different mitochondrial compartments (matrix, outer, inner and intermembrane regions). The method is trained and tested in cross-validation on a newly generated, high-quality dataset comprising 424 mitochondrial proteins with experimental evidence for sub-organelle localizations. We benchmark DeepMito towards the only one recent approach developed for the same task. Results indicate that DeepMito performances are superior. Finally, genomic-scale prediction on a highly-curated dataset of human mitochondrial proteins further confirms the effectiveness of our approach and suggests that DeepMito is a good candidate for genome-scale annotation of mitochondrial protein subcellular localization. AVAILABILITY AND IMPLEMENTATION: The DeepMito web server as well as all datasets used in this study are available at http://busca.biocomp.unibo.it/deepmito. A standalone version of DeepMito is available on DockerHub at https://hub.docker.com/r/bolognabiocomp/deepmito. DeepMito source code is available on GitHub at https://github.com/BolognaBiocomp/deepmito. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Castrense Savojardo, Niccolò Bruciaferri, Giacomo Tartari, Pier Luigi Martelli, Rita Casadio
Bioinform.1
2020 Large-scale prediction and analysis of protein sub-mitochondrial localization with DeepMito
abstract
BACKGROUND: The prediction of protein subcellular localization is a key step of the big effort towards protein functional annotation. Many computational methods exist to identify high-level protein subcellular compartments such as nucleus, cytoplasm or organelles. However, many organelles, like mitochondria, have their own internal compartmentalization. Knowing the precise location of a protein inside mitochondria is crucial for its accurate functional characterization. We recently developed DeepMito, a new method based on a 1-Dimensional Convolutional Neural Network (1D-CNN) architecture outperforming other similar approaches available in literature. RESULTS: Here, we explore the adoption of DeepMito for the large-scale annotation of four sub-mitochondrial localizations on mitochondrial proteomes of five different species, including human, mouse, fly, yeast and Arabidopsis thaliana. A significant fraction of the proteins from these organisms lacked experimental information about sub-mitochondrial localization. We adopted DeepMito to fill the gap, providing complete characterization of protein localization at sub-mitochondrial level for each protein of the five proteomes. Moreover, we identified novel mitochondrial proteins fishing on the set of proteins lacking any subcellular localization annotation using available state-of-the-art subcellular localization predictors. We finally performed additional functional characterization of proteins predicted by DeepMito as localized into the four different sub-mitochondrial compartments using both available experimental and predicted GO terms. All data generated in this study were collected into a database called DeepMitoDB (available at http://busca.biocomp.unibo.it/deepmitodb ), providing complete functional characterization of 4307 mitochondrial proteins from the five species. CONCLUSIONS: DeepMitoDB offers a comprehensive view of mitochondrial proteins, including experimental and predicted fine-grain sub-cellular localization and annotated and predicted functional annotations. The database complements other similar resources providing characterization of new proteins. Furthermore, it is also unique in including localization information at the sub-mitochondrial level. For this reason, we believe that DeepMitoDB can be a valuable resource for mitochondrial research.
Castrense Savojardo, Pier Luigi Martelli, Giacomo Tartari, Rita Casadio
BMC Bioinform.1
2019 On the biases in predictions of protein stability changes upon variations: the INPS test case
abstract
INPS (Fariselli et al., 2015) is a simple method to predict the protein stability changes upon single point variation. It is based on Support Vector Machine and it was trained using 7 input features that are: (i) Blosum62 substitution, (ii) residue mutability, (iii) molecular weight wild type, (iv) molecular weight variant, (v) hydrophobicity wild type, (vi) hydrophobicity variant, (vii) difference between Viterbi scores of the wild type and variant (Fariselli et al., 2015). The first two input features are not anti-symmetric, the last is anti-symmetric by construction, while the other features can be learned to be anti-symmetric. To enforce anti-symmetry INPS was trained using both direct and inverse variations (Fariselli et al., 2015). When the alignment is poor we can expect a deviation from the anti-symmetry. INPS3D (Savojardo et al., 2016), adds to INPS information about contact potential and solvent accessibility. The solvent accessibility is intrinsically non anti-symmetric, and for this reason INPS3D was trained only with direct variations (Savojardo et al., 2016). The Ssym dataset is a manually curated selection of variations from the ProTherm database (Pucci et al., 2018). It contains variations with experimental ΔΔG values for which the 3D structures of both the wild-type and variant proteins were solved by X-ray crystallography. Ssym consists of 684 variations, half of which are direct (reported in the literature) and half are obtained by anti-symmetry (Pucci et al., 2018). To overcome the bias problem, a new method, PopMuSiCSYM, was explicitly developed to furnish completely anti-symmetric predictions of the variations (Pucci et al., 2015, 2018). The authors of the Ssym dataset estimated the performance of fifteen selected predictors using the average bias parameter (<δ> = ∑(ΔΔGdir+ΔΔGinv⁠)/N) and the linear correlation coefficient (rdir-inv) between the predicted ΔΔG values of the direct and the corresponding inverse variations. A perfect predictor should have bias close to zero (<δ> = 0) and a correlation coefficient close to -1 (rdir-inv = −1). In the same paper, they also tested the performances of these fifteen predictors using the experimental ΔΔG values on both direct and inverse variations using the root mean square deviation (σdir, σinv) and the linear correlation coefficient between the predicted and experimental values (rdir, rinv). In Table 1 we report the predictor performances and biases taken from the original paper (Pucci et al., 2018) and we add those computed using INPS and INPS3D. From Table 1, it is clear that the partial anti-symmetry of the INPS input and its training performed on both direct and inverse variations paid off. INPS shows a very low bias (<δ>, just second best), the highest correlation between direct and inverse variations (rdir-inv) and the highest performance on the inverse variation sets (σinv and rinv). The graph of the correlation between direct and inverse variation predicted by INPS is presented in Figure 1. INPS and INPS3D prediction on Ssym. The INPS (black) and INPS3D (grey) predictions for the 342 pairs (direct and inverse) of single-point variations in Ssym (Pucci et al., 2018) are shown. ΔΔG for the direct (x-axis) and inverse (y-axis) variations are reported as dots in the graph. The ‘ideal’ relationship ΔΔGAB + ΔΔGBA = 0 is shown as a solid line Bias analysis on the Ssym dataset (Pucci et al. 2018) Notes: Standard deviations σ and the average bias <δ> are in kcal/mol. All the values, except those of INPS and INPS3D, are taken from (Pucci et al., 2018). Bold character highlights INPS method. Bias analysis on the Ssym dataset (Pucci et al. 2018) Notes: Standard deviations σ and the average bias <δ> are in kcal/mol. All the values, except those of INPS and INPS3D, are taken from (Pucci et al., 2018). Bold character highlights INPS method. The INPS training set includes some of the proteins in the Ssym dataset (which is also true for all the other tested methods reported in Table 1). For sake of comparison, we also computed the performances of a ‘blind’ INPS. In this case, each ΔΔG prediction is obtained by using a svm model that does not include the protein to predict or a similar one (sequence identity <25%) in the training set. The standard deviations and Pearson correlation coefficients computed for the blind-INPS are: 1.44 kcal/mol and 0.48 for direct variations and 1.45 kcal/mol and 0.47 for inverse variations. The correlation between the predictions of direct and inverse variations is -0.99 while the average bias (measured through the δ) is -0.06 kcal/mol. The structure-based INPS3D has the second best correlation between direct and inverse correlation but a considerably higher bias. This is expected given that INPS3D has not been trained on inverse variations and conains as input feature the solvent accessibility, which is not anti-symmetric. Finally, it must be stressed that the correlations reported in Table 1 are only intended to assess the anti-symmetricity of the methods, and not as estimators of their overall performances, which is a very difficult task that also requires an evaluation of stability of the dataset (Yang et al., 2018). A second dataset was built by Usmanova et al. (2018), by extracting high-resolution pairs of proteins from the Protein Data Bank (PDB) differing by one to ten amino acids. Here we test INPS on the subset of all the single-site variations extracted from the Usmanova et al. (2018) dataset, that is, all the pdb pairs whose sequences differ by exactly one residue. This subset comprises 1000 pairs of single-point protein variations. Although for this set the experimental ΔΔG values are not known, given its important size, it is very informative for assessing predictor biases. In Usmanova et al. (2018), bias for an individual variation is estimated as: =(ΔΔGAB + ΔΔGBA)/2. In Table 2 we add the INPS and INPS3D anti-symmetry indexes to those previously computed on other predictors (Usmanova et al., 2018). The bias of INPS is the smallest and the correlation between direct and inverse variations is the highest. INPS3D, while showing less anti-symmetricity than INPS, has the second best bias and the second best correlation. Bias analysis on the Usmanova et al. (2018) dataset Notes: First column: mean and standard error of the mean for the biases. Second column: r, Pearson correlation coefficient with the associated P-value, between direct and inverse variations. All the values except INPS and INPS3D, are taken from Usmanova et al. (2018). 0.0*: P-value set to 0, since it is undetectable by the machine precision. Bold character highlights INPS method. Bias analysis on the Usmanova et al. (2018) dataset Notes: First column: mean and standard error of the mean for the biases. Second column: r, Pearson correlation coefficient with the associated P-value, between direct and inverse variations. All the values except INPS and INPS3D, are taken from Usmanova et al. (2018). 0.0*: P-value set to 0, since it is undetectable by the machine precision. Bold character highlights INPS method. The very high correlation of -0.95 between direct and inverse variations obtained by INPS on this set is shown in the graph of Figure 2, to be contrasted with those reported for the other methods in the figures of the paper by Usmanova et al. (2018). In Figure 2, the points very far from the diagonal correspond to variations whose protein sequences are poorly aligned (number of aligned sequences <20). INPS and INPS3D prediction for the 1000 pairs (direct and inverse) of single-point variations (Usmanova et al., 2018). ΔΔG for the direct (x-axis) and inverse (y-axis) variations are reported as dots in the graph. The ‘ideal’ relationship ΔΔGAB + ΔΔGBA = 0 is shown as a solid line Recently, two groups (Pucci et al., 2018; Usmanova et al., 2018) compiled two datasets to test a very important bias that affects most of the computational methods designed to predict free energy changes upon protein variation. This bias is the lack of anti-symmetry in the ΔΔG predictions between direct and inverse variations. These studies verified this bias on several relevant methods. Here we add the test on INPS (Fariselli et al., 2015) that was specifically designed to take into account anti-symmetry problem using an input that was partially anti-symmetric (the HMM Viterbi score in particular). Here we show that INPS performs very well on both datasets proposed in the two papers. The fact that INPS uses only sequence information makes it less sensitive to structural rearrangement upon residue substitution and in turn more robust and suitable for testing any kind of protein variations. This work has been supported by EBA-PRISM an Israel-Italy collaborative project, the Israel Ministry of Science and Technology and Italian Ministry of Foreign Affair and International Cooperation. This research received funding specifically appointed to Department of Medical Sciences from the Italian Ministry for Education, University and Research (MIUR) under the programme “Dipartimenti di Eccellenza 2018 – 2022” Project code D15D18000410001. Conflict of Interest: none declared.
Ludovica Montanucci, Castrense Savojardo, Pier Luigi Martelli, Rita Casadio, Piero Fariselli
Bioinform.2
2018 DeepSig: deep learning improves signal peptide detection in proteins
abstract
Motivation: The identification of signal peptides in protein sequences is an important step toward protein localization and function characterization. Results: Here, we present DeepSig, an improved approach for signal peptide detection and cleavage-site prediction based on deep learning methods. Comparative benchmarks performed on an updated independent dataset of proteins show that DeepSig is the current best performing method, scoring better than other available state-of-the-art approaches on both signal peptide detection and precise cleavage-site identification. Availability and implementation: DeepSig is available as both standalone program and web server at https://deepsig.biocomp.unibo.it. All datasets used in this study can be obtained from the same website. Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Castrense Savojardo, Pier Luigi Martelli, Piero Fariselli, Rita Casadio
Bioinform.1
2017 ISPRED4: interaction sites PREDiction in protein structures with a refining grammar model
abstract
MOTIVATION: The identification of protein-protein interaction (PPI) sites is an important step towards the characterization of protein functional integration in the cell complexity. Experimental methods are costly and time-consuming and computational tools for predicting PPI sites can fill the gaps of PPI present knowledge. RESULTS: We present ISPRED4, an improved structure-based predictor of PPI sites on unbound monomer surfaces. ISPRED4 relies on machine-learning methods and it incorporates features extracted from protein sequence and structure. Cross-validation experiments are carried out on a new dataset that includes 151 high-resolution protein complexes and indicate that ISPRED4 achieves a per-residue Matthew Correlation Coefficient of 0.48 and an overall accuracy of 0.85. Benchmarking results show that ISPRED4 is one of the top-performing PPI site predictors developed so far. CONTACT: [email protected]. AVAILABILITY AND IMPLEMENTATION: ISPRED4 and datasets used in this study are available at http://ispred4.biocomp.unibo.it .
Castrense Savojardo, Piero Fariselli, Pier Luigi Martelli, Rita Casadio
Bioinform.1
2017 SChloro: directing Viridiplantae proteins to six chloroplastic sub-compartments
abstract
Motivation: Chloroplasts are organelles found in plants and involved in several important cell processes. Similarly to other compartments in the cell, chloroplasts have an internal structure comprising several sub-compartments, where different proteins are targeted to perform their functions. Given the relation between protein function and localization, the availability of effective computational tools to predict protein sub-organelle localizations is crucial for large-scale functional studies. Results: In this paper we present SChloro, a novel machine-learning approach to predict protein sub-chloroplastic localization, based on targeting signal detection and membrane protein information. The proposed approach performs multi-label predictions discriminating six chloroplastic sub-compartments that include inner membrane, outer membrane, stroma, thylakoid lumen, plastoglobule and thylakoid membrane. In comparative benchmarks, the proposed method outperforms current state-of-the-art methods in both single- and multi-compartment predictions, with an overall multi-label accuracy of 74%. The results demonstrate the relevance of the approach that is eligible as a good candidate for integration into more general large-scale annotation pipelines of protein subcellular localization. Availability and Implementation: The method is available as web server at http://schloro.biocomp.unibo.it Contact: [email protected].
Castrense Savojardo, Pier Luigi Martelli, Piero Fariselli, Rita Casadio
Bioinform.1
2016 INPS-MD: a web server to predict stability of protein variants from sequence and structure
abstract
MOTIVATION: Protein function depends on its structural stability. The effects of single point variations on protein stability can elucidate the molecular mechanisms of human diseases and help in developing new drugs. Recently, we introduced INPS, a method suited to predict the effect of variations on protein stability from protein sequence and whose performance is competitive with the available state-of-the-art tools. RESULTS: In this article, we describe INPS-MD (Impact of Non synonymous variations on Protein Stability-Multi-Dimension), a web server for the prediction of protein stability changes upon single point variation from protein sequence and/or structure. Here, we complement INPS with a new predictor (INPS3D) that exploits features derived from protein 3D structure. INPS3D scores with Pearson's correlation to experimental ΔΔG values of 0.58 in cross validation and of 0.72 on a blind test set. The sequence-based INPS scores slightly lower than the structure-based INPS3D and both on the same blind test sets well compare with the state-of-the-art methods. AVAILABILITY AND IMPLEMENTATION: INPS and INPS3D are available at the same web server: http://inpsmd.biocomp.unibo.it SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. CONTACT: [email protected].
Castrense Savojardo, Piero Fariselli, Pier Luigi Martelli, Rita Casadio
Bioinform.1
2015 INPS: predicting the impact of non-synonymous variations on protein stability from sequence
abstract
MOTIVATION: A tool for reliably predicting the impact of variations on protein stability is extremely important for both protein engineering and for understanding the effects of Mendelian and somatic mutations in the genome. Next Generation Sequencing studies are constantly increasing the number of protein sequences. Given the huge disproportion between protein sequences and structures, there is a need for tools suited to annotate the effect of mutations starting from protein sequence without relying on the structure. Here, we describe INPS, a novel approach for annotating the effect of non-synonymous mutations on the protein stability from its sequence. INPS is based on SVM regression and it is trained to predict the thermodynamic free energy change upon single-point variations in protein sequences. RESULTS: We show that INPS performs similarly to the state-of-the-art methods based on protein structure when tested in cross-validation on a non-redundant dataset. INPS performs very well also on a newly generated dataset consisting of a number of variations occurring in the tumor suppressor protein p53. Our results suggest that INPS is a tool suited for computing the effect of non-synonymous polymorphisms on protein stability when the protein structure is not available. We also show that INPS predictions are complementary to those of the state-of-the-art, structure-based method mCSM. When the two methods are combined, the overall prediction on the p53 set scores significantly higher than those of the single methods. AVAILABILITY AND IMPLEMENTATION: The presented method is available as web server at http://inps.biocomp.unibo.it. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary Materials are available at Bioinformatics online.
Piero Fariselli, Pier Luigi Martelli, Castrense Savojardo, Rita Casadio
Bioinform.3
2015 TPpred3 detects and discriminates mitochondrial and chloroplastic targeting peptides in eukaryotic proteins
abstract
MOTIVATION: Molecular recognition of N-terminal targeting peptides is the most common mechanism controlling the import of nuclear-encoded proteins into mitochondria and chloroplasts. When experimental information is lacking, computational methods can annotate targeting peptides, and determine their cleavage sites for characterizing protein localization, function, and mature protein sequences. The problem of discriminating mitochondrial from chloroplastic propeptides is particularly relevant when annotating proteomes of photosynthetic Eukaryotes, endowed with both types of sequences. RESULTS: Here, we introduce TPpred3, a computational method that given any Eukaryotic protein sequence performs three different tasks: (i) the detection of targeting peptides; (ii) their classification as mitochondrial or chloroplastic and (iii) the precise localization of the cleavage sites in an organelle-specific framework. Our implementation is based on our TPpred previously introduced. Here, we integrate a new N-to-1 Extreme Learning Machine specifically designed for the classification task (ii). For the last task, we introduce an organelle-specific Support Vector Machine that exploits sequence motifs retrieved with an extensive motif-discovery analysis of a large set of mitochondrial and chloroplastic proteins. We show that TPpred3 outperforms the state-of-the-art methods in all the three tasks. AVAILABILITY AND IMPLEMENTATION: The method server and datasets are available at http://tppred3.biocomp.unibo.it. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Castrense Savojardo, Pier Luigi Martelli, Piero Fariselli, Rita Casadio
Bioinform.1
2014 TPpred2: improving the prediction of mitochondrial targeting peptide cleavage sites by exploiting sequence motifs
abstract
SUMMARY: Targeting peptides are N-terminal sorting signals in proteins that promote their translocation to mitochondria through the interaction with different protein machineries. We recently developed TPpred, a machine learning-based method scoring among the best ones available to predict the presence of a targeting peptide into a protein sequence and its cleavage site. Here we introduce TPpred2 that improves TPpred performances in the task of identifying the cleavage site of the targeting peptides. TPpred2 is now available as a web interface and as a stand-alone version for users who can freely download and adopt it for processing large volumes of sequences. Availability and implementaion: TPpred2 is available both as web server and stand-alone version at http://tppred2.biocomp.unibo.it. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Castrense Savojardo, Pier Luigi Martelli, Piero Fariselli, Rita Casadio
Bioinform.1
2013 The prediction of organelle-targeting peptides in eukaryotic proteins with Grammatical-Restrained Hidden Conditional Random Fields
abstract
MOTIVATION: Targeting peptides are the most important signal controlling the import of nuclear encoded proteins into mitochondria and plastids. In the lack of experimental information, their prediction is an essential step when proteomes are annotated for inferring both the localization and the sequence of mature proteins. RESULTS: We developed TPpred a new predictor of organelle-targeting peptides based on Grammatical-Restrained Hidden Conditional Random Fields. TPpred is trained on a non-redundant dataset of proteins where the presence of a target peptide was experimentally validated, comprising 297 sequences. When tested on the 297 positive and some other 8010 negative examples, TPpred outperformed available methods in both accuracy and Matthews correlation index (96% and 0.58, respectively). Given its very low-false-positive rate (3.0%), TPpred is, therefore, well suited for large-scale analyses at the proteome level. We predicted that from ∼4 to 9% of the sequences of human, Arabidopsis thaliana and yeast proteomes contain targeting peptides and are, therefore, likely to be localized in mitochondria and plastids. TPpred predictions correlate to a good extent with the experimental annotation of the subcellular localization, when available. TPpred was also trained and tested to predict the cleavage site of the organelle-targeting peptide: on this task, the average error of TPpred on mitochondrial and plastidic proteins is 7 and 15 residues, respectively. This value is lower than the error reported by other methods currently available. AVAILABILITY: The TPpred datasets are available at http://biocomp.unibo.it/valentina/TPpred/. TPpred is available on request from the authors. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Valentina Indio, Pier Luigi Martelli, Castrense Savojardo, Piero Fariselli, Rita Casadio
Bioinform.3
2013 BETAWARE: a machine-learning tool to detect and predict transmembrane beta-barrel proteins in prokaryotes
abstract
SUMMARY: The annotation of membrane proteins in proteomes is an important problem of Computational Biology, especially after the development of high-throughput techniques that allow fast and efficient genome sequencing. Among membrane proteins, transmembrane β-barrels (TMBBs) are poorly represented in the database of protein structures (PDB) and difficult to identify with experimental approaches. They are, however, extremely important, playing key roles in several cell functions and bacterial pathogenicity. TMBBs are included in the lipid bilayer with a β-barrel structure and are presently found in the outer membranes of Gram-negative bacteria, mitochondria and chloroplasts. Recently, we developed two top-performing methods based on machine-learning approaches to tackle both the detection of TMBBs in sets of proteins and the prediction of their topology. Here, we present our BETAWARE program that includes both approaches and can run as a standalone program on a linux-based computer to easily address in-home massive protein annotation or filtering. AVAILABILITY AND IMPLEMENTATION: http://www.biocomp.unibo.it/∼savojard/betawarecl .
Castrense Savojardo, Piero Fariselli, Rita Casadio
Bioinform.1
2013 BCov: a method for predicting β-sheet topology using sparse inverse covariance estimation and integer programming
abstract
MOTIVATION: Prediction of protein residue contacts, even at the coarse-grain level, can help in finding solutions to the protein structure prediction problem. Unlike α-helices that are locally stabilized, β-sheets result from pairwise hydrogen bonding of two or more disjoint regions of the protein backbone. The problem of predicting contacts among β-strands in proteins has been addressed by several supervised computational approaches. Recently, prediction of residue contacts based on correlated mutations has been greatly improved and finally allows the prediction of 3D structures of the proteins. RESULTS: In this article, we describe BCov, which is the first unsupervised method to predict the β-sheet topology starting from the protein sequence and its secondary structure. BCov takes advantage of the sparse inverse covariance estimation to define β-strand partner scores. Then an optimization based on integer programming is carried out to predict the β-sheet connectivity. When tested on the prediction of β-strand pairing, BCov scores with average values of Matthews Correlation Coefficient (MCC) and F1 equal to 0.56 and 0.61, respectively, on a non-redundant dataset of 916 protein chains known with atomic resolution. Our approach well compares with the state-of-the-art methods trained so far for this specific task. AVAILABILITY AND IMPLEMENTATION: The method is freely available under General Public License at http://biocomp.unibo.it/savojard/bcov/bcov-1.0.tar.gz. The new dataset BetaSheet1452 can be downloaded at http://biocomp.unibo.it/savojard/bcov/BetaSheet1452.dat.
Castrense Savojardo, Piero Fariselli, Pier Luigi Martelli, Rita Casadio
Bioinform.1
2013 Prediction of disulfide connectivity in proteins with machine-learning methods and correlated mutations
abstract
BACKGROUND: Recently, information derived by correlated mutations in proteins has regained relevance for predicting protein contacts. This is due to new forms of mutual information analysis that have been proven to be more suitable to highlight direct coupling between pairs of residues in protein structures and to the large number of protein chains that are currently available for statistical validation. It was previously discussed that disulfide bond topology in proteins is also constrained by correlated mutations. RESULTS: In this paper we exploit information derived from a corrected mutual information analysis and from the inverse of the covariance matrix to address the problem of the prediction of the topology of disulfide bonds in Eukaryotes. Recently, we have shown that Support Vector Regression (SVR) can improve the prediction for the disulfide connectivity patterns. Here we show that the inclusion of the correlated mutation information increases of 5 percentage points the SVR performance (from 54% to 59%). When this approach is used in combination with a method previously developed by us and scoring at the state of art in predicting both location and topology of disulfide bonds in Eukaryotes (DisLocate), the per-protein accuracy is 38%, 2 percentage points higher than that previously obtained. CONCLUSIONS: In this paper we show that the inclusion of information derived from correlated mutations can improve the performance of the state of the art methods for predicting disulfide connectivity patterns in Eukaryotic proteins. Our analysis also provides support to the notion that improving methods to extract evolutionary information from multiple sequence alignments greatly contributes to the scoring performance of predictors suited to detect relevant features from protein chains.
Castrense Savojardo, Piero Fariselli, Pier Luigi Martelli, Rita Casadio
BMC Bioinform.1
2011 Improving the prediction of disulfide bonds in Eukaryotes with machine learning methods and protein subcellular localization
abstract
MOTIVATION: Disulfide bonds stabilize protein structures and play relevant roles in their functions. Their formation requires an oxidizing environment and their stability is consequently depending on the redox ambient potential, which may differ according to the subcellular compartment. Several methods are available to predict cysteine-bonding state and connectivity patterns. However, none of them takes into consideration the relevance of protein subcellular localization. RESULTS: Here we develop DISLOCATE, a two-step method based on machine learning models for predicting both the bonding state and the connectivity patterns of cysteine residues in a protein chain. We find that the inclusion of protein subcellular localization improves the performance of these predictive steps by 3 and 2 percentage points, respectively. When compared with previously developed methods for predicting disulfide bonds from sequence, DISLOCATE improves the overall performance by more than 10 percentage points. AVAILABILITY: The method and the dataset are available at the Web page http://www.biocomp.unibo.it/savojard/Dislocate.html. GRHCRF code is available at http://www.biocomp.unibo.it/savojard/biocrf.html. CONTACT: [email protected].
Castrense Savojardo, Piero Fariselli, Monther Alhamdoosh, Pier Luigi Martelli, Andrea Pierleoni, Rita Casadio
Bioinform.1
2011 Improving the detection of transmembrane β-barrel chains with N-to-1 extreme learning machines
abstract
MOTIVATION: Transmembrane β-barrels (TMBBs) are extremely important proteins that play key roles in several cell functions. They cross the lipid bilayer with β-barrel structures. TMBBs are presently found in the outer membranes of Gram-negative bacteria and of mitochondria and chloroplasts. Loop exposure outside the bacterial cell membranes makes TMBBs important targets for vaccine or drug therapies. In genomes, they are not highly represented and are difficult to identify with experimental approaches. Several computational methods have been developed to discriminate TMBBs from other types of proteins. However, the best performing approaches have a high fraction of false positive predictions. RESULTS: In this article, we introduce a new machine learning approach for TMBB detection based on N-to-1 Extreme Learning Machines that significantly outperforms previous methods achieving a Matthews correlation coefficient of 0.82, a probability of correct prediction of 0.92 and a sensitivity of 0.73.
Castrense Savojardo, Piero Fariselli, Rita Casadio
Bioinform.1