Pier Luigi Martelli

dblp:30/3511 · DBLP profile ↗
← Back
36ranked-venue papers
1as first author
5since 2021 · last 2025
0000-0002-0274-5669ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 35 · 1 first-author · 5 since 2021Systems, architecture and hardware · 1
YearPublicationVenuePosition
2025 DDGemb: predicting protein stability change upon single- and multi-point variations with embeddings and deep learning
abstract
MOTIVATION: The knowledge of protein stability upon residue variation is an important step for functional protein design and for understanding how protein variants can promote disease onset. Computational methods are important to complement experimental approaches and allow a fast screening of large datasets of variations. RESULTS: In this work, we present DDGemb, a novel method combining protein language model embeddings and transformer architectures to predict protein ΔΔG upon both single- and multi-point variations. DDGemb has been trained on a high-quality dataset derived from literature and tested on available benchmark datasets of single- and multi-point variations. DDGemb performs at the state of the art in both single- and multi-point variations. AVAILABILITY AND IMPLEMENTATION: DDGemb is available as web server at https://ddgemb.biocomp.unibo.it. Datasets used in this study are available at https://ddgemb.biocomp.unibo.it/datasets.
Castrense Savojardo, Matteo Manfredi, Pier Luigi Martelli, Rita Casadio
Bioinform.3
2023 CoCoNat: a novel method based on deep learning for coiled-coil prediction
abstract
MOTIVATION: Coiled-coil domains (CCD) are widespread in all organisms and perform several crucial functions. Given their relevance, the computational detection of CCD is very important for protein functional annotation. State-of-the-art prediction methods include the precise identification of CCD boundaries, the annotation of the typical heptad repeat pattern along the coiled-coil helices as well as the prediction of the oligomerization state. RESULTS: In this article, we describe CoCoNat, a novel method for predicting coiled-coil helix boundaries, residue-level register annotation, and oligomerization state. Our method encodes sequences with the combination of two state-of-the-art protein language models and implements a three-step deep learning procedure concatenated with a Grammatical-Restrained Hidden Conditional Random Field for CCD identification and refinement. A final neural network predicts the oligomerization state. When tested on a blind test set routinely adopted, CoCoNat obtains a performance superior to the current state-of-the-art both for residue-level and segment-level CCD. CoCoNat significantly outperforms the most recent state-of-the-art methods on register annotation and prediction of oligomerization states. AVAILABILITY AND IMPLEMENTATION: CoCoNat web server is available at https://coconat.biocomp.unibo.it. Standalone version is available on GitHub at https://github.com/BolognaBiocomp/coconat.
Giovanni Madeo, Castrense Savojardo, Matteo Manfredi, Pier Luigi Martelli, Rita Casadio
Bioinform.4
2022 Quantum computing algorithms: getting closer to critical problems in computational biology
abstract
The recent biotechnological progress has allowed life scientists and physicians to access an unprecedented, massive amount of data at all levels (molecular, supramolecular, cellular and so on) of biological complexity. So far, mostly classical computational efforts have been dedicated to the simulation, prediction or de novo design of biomolecules, in order to improve the understanding of their function or to develop novel therapeutics. At a higher level of complexity, the progress of omics disciplines (genomics, transcriptomics, proteomics and metabolomics) has prompted researchers to develop informatics means to describe and annotate new biomolecules identified with a resolution down to the single cell, but also with a high-throughput speed. Machine learning approaches have been implemented to both the modelling studies and the handling of biomedical data. Quantum computing (QC) approaches hold the promise to resolve, speed up or refine the analysis of a wide range of these computational problems. Here, we review and comment on recently developed QC algorithms for biocomputing, with a particular focus on multi-scale modelling and genomic analyses. Indeed, differently from other computational approaches such as protein structure prediction, these problems have been shown to be adequately mapped onto quantum architectures, the main limit for their immediate use being the number of qubits and decoherence effects in the available quantum machines. Possible advantages over the classical counterparts are highlighted, along with a description of some hybrid classical/quantum approaches, which could be the closest to be realistically applied in biocomputation.
Laura Marchetti, Riccardo Nifosì, Pier Luigi Martelli, Eleonora Da Pozzo, Valentina Cappello, Francesco Banterle, Maria Letizia Trincavelli, Claudia Martini, Massimo D'elia
Briefings Bioinform.3
2022 E-SNPs&GO: embedding of protein sequence and function improves the annotation of human pathogenic variants
abstract
MOTIVATION: The advent of massive DNA sequencing technologies is producing a huge number of human single-nucleotide polymorphisms occurring in protein-coding regions and possibly changing their sequences. Discriminating harmful protein variations from neutral ones is one of the crucial challenges in precision medicine. Computational tools based on artificial intelligence provide models for protein sequence encoding, bypassing database searches for evolutionary information. We leverage the new encoding schemes for an efficient annotation of protein variants. RESULTS: E-SNPs&GO is a novel method that, given an input protein sequence and a single amino acid variation, can predict whether the variation is related to diseases or not. The proposed method adopts an input encoding completely based on protein language models and embedding techniques, specifically devised to encode protein sequences and GO functional annotations. We trained our model on a newly generated dataset of 101 146 human protein single amino acid variants in 13 661 proteins, derived from public resources. When tested on a blind set comprising 10 266 variants, our method well compares to recent approaches released in literature for the same task, reaching a Matthews Correlation Coefficient score of 0.72. We propose E-SNPs&GO as a suitable, efficient and accurate large-scale annotator of protein variant datasets. AVAILABILITY AND IMPLEMENTATION: The method is available as a webserver at https://esnpsandgo.biocomp.unibo.it. Datasets and predictions are available at https://esnpsandgo.biocomp.unibo.it/datasets. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Matteo Manfredi, Castrense Savojardo, Pier Luigi Martelli, Rita Casadio
Bioinform.3
2021 On the critical review of five machine learning-based algorithms for predicting protein stability changes upon mutation
abstract
A review, recently published in this journal by Fang (2019), showed that methods trained for the prediction of protein stability changes upon mutation have a very critical bias: they neglect that a protein variation (A- > B) and its reverse (B- > A) must have the opposite value of the free energy difference (ΔΔGAB = - ΔΔGBA). In this letter, we complement the Fang's paper presenting a more general view of the problem. In particular, a machine learning-based method, published in 2015 (INPS), addressed the bias issue directly. We include the analysis of the missing method, showing that INPS is nearly insensitive to the addressed problem.
Castrense Savojardo, Pier Luigi Martelli, Rita Casadio, Piero Fariselli
Briefings Bioinform.2
2020 DeepMito: accurate prediction of protein sub-mitochondrial localization using convolutional neural networks
abstract
MOTIVATION: The correct localization of proteins in cell compartments is a key issue for their function. Particularly, mitochondrial proteins are physiologically active in different compartments and their aberrant localization contributes to the pathogenesis of human mitochondrial pathologies. Many computational methods exist to assign protein sequences to subcellular compartments such as nucleus, cytoplasm and organelles. However, a substantial lack of experimental evidence in public sequence databases hampered so far a finer grain discrimination, including also intra-organelle compartments. RESULTS: We describe DeepMito, a novel method for predicting protein sub-mitochondrial cellular localization. Taking advantage of powerful deep-learning approaches, such as convolutional neural networks, our method is able to achieve very high prediction performances when discriminating among four different mitochondrial compartments (matrix, outer, inner and intermembrane regions). The method is trained and tested in cross-validation on a newly generated, high-quality dataset comprising 424 mitochondrial proteins with experimental evidence for sub-organelle localizations. We benchmark DeepMito towards the only one recent approach developed for the same task. Results indicate that DeepMito performances are superior. Finally, genomic-scale prediction on a highly-curated dataset of human mitochondrial proteins further confirms the effectiveness of our approach and suggests that DeepMito is a good candidate for genome-scale annotation of mitochondrial protein subcellular localization. AVAILABILITY AND IMPLEMENTATION: The DeepMito web server as well as all datasets used in this study are available at http://busca.biocomp.unibo.it/deepmito. A standalone version of DeepMito is available on DockerHub at https://hub.docker.com/r/bolognabiocomp/deepmito. DeepMito source code is available on GitHub at https://github.com/BolognaBiocomp/deepmito. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Castrense Savojardo, Niccolò Bruciaferri, Giacomo Tartari, Pier Luigi Martelli, Rita Casadio
Bioinform.4
2020 LabxDB: versatile databases for genomic sequencing and lab management
abstract
SUMMARY: Experimental laboratory management and data-driven science require centralized software for sharing information, such as lab collections or genomic sequencing datasets. Although database servers such as PostgreSQL can store such information with multiple-user access, they lack user-friendly graphical and programmatic interfaces for easy data access and inputting. We developed LabxDB, a versatile open-source solution for organizing and sharing structured data. We provide several out-of-the-box databases for deployment in the cloud including simple mutant or plasmid collections and purchase-tracking databases. We also developed a high-throughput sequencing (HTS) database, LabxDB seq, dedicated to storage of hierarchical sample annotations. Scientists can import their own or publicly available HTS data into LabxDB seq to manage them from production to publication. Using LabxDB's programmatic access (REST API), annotations can be easily integrated into bioinformatics pipelines. LabxDB is modular, offering a flexible framework that scientists can leverage to build new database interfaces adapted to their needs. AVAILABILITY AND IMPLEMENTATION: LabxDB is available at https://gitlab.com/vejnar/labxdb and https://labxdb.vejnar.org for documentation. LabxDB is licensed under the terms of the Mozilla Public License 2.0. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Charles E. Vejnar, Antonio J. Giraldez, Pier Luigi Martelli
Bioinform.3
2020 Feature selection and classification of noisy proteomics mass spectrometry data based on one-bit perturbed compressed sensing
abstract
MOTIVATION: The classification of high-throughput protein data based on mass spectrometry (MS) is of great practical significance in medical diagnosis. Generally, MS data are characterized by high dimension, which inevitably leads to prohibitive cost of computation. To solve this problem, one-bit compressed sensing (CS), which is an extreme case of quantized CS, has been employed on MS data to select important features with low dimension. Though enjoying remarkably reduction of computation complexity, the current one-bit CS method does not consider the unavoidable noise contained in MS dataset, and does not exploit the inherent structure of the underlying MS data. RESULTS: We propose two feature selection (FS) methods based on one-bit CS to deal with the noise and the underlying block-sparsity features, respectively. In the first method, the FS problem is modeled as a perturbed one-bit CS problem, where the perturbation represents the noise in MS data. By iterating between perturbation refinement and FS, this method selects the significant features from noisy data. The second method formulates the problem as a perturbed one-bit block CS problem and selects the features block by block. Such block extraction is due to the fact that the significant features in the first method usually cluster in groups. Experiments show that, the two proposed methods have better classification performance for real MS data when compared with the existing method, and the second one outperforms the first one. AVAILABILITY AND IMPLEMENTATION: The source code of our methods is available at: https://github.com/tianyan8023/OBCS. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Wenbo Xu 0003, Siye Wang, Yupeng Cui, Pier Luigi Martelli
Bioinform.5
2020 DeepLGP: a novel deep learning method for prioritizing lncRNA target genes
abstract
MOTIVATION: Although long non-coding RNAs (lncRNAs) have limited capacity for encoding proteins, they have been verified as biomarkers in the occurrence and development of complex diseases. Recent wet-lab experiments have shown that lncRNAs function by regulating the expression of protein-coding genes (PCGs), which could also be the mechanism responsible for causing diseases. Currently, lncRNA-related biological data are increasing rapidly. Whereas, no computational methods have been designed for predicting the novel target genes of lncRNA. RESULTS: In this study, we present a graph convolutional network (GCN) based method, named DeepLGP, for prioritizing target PCGs of lncRNA. First, gene and lncRNA features were selected, these included their location in the genome, expression in 13 tissues and miRNA-mediated lncRNA-gene pairs. Next, GCN was applied to convolve a gene interaction network for encoding the features of genes and lncRNAs. Then, these features were used by the convolutional neural network for prioritizing target genes of lncRNAs. In 10-cross validations on two independent datasets, DeepLGP obtained high area under curves (0.90-0.98) and area under precision-recall curves (0.91-0.98). We found that lncRNA pairs with high similarity had more overlapped target genes. Further experiments showed that genes targeted by the same lncRNA sets had a strong likelihood of causing the same diseases, which could help in identifying disease-causing PCGs. AVAILABILITY AND IMPLEMENTATION: https://github.com/zty2009/LncRNA-target-gene. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Tianyi Zhao 0001, Yang Hu 0008, Jiajie Peng, Liang Cheng 0006, Pier Luigi Martelli
Bioinform.5
2020 Large-scale prediction and analysis of protein sub-mitochondrial localization with DeepMito
abstract
BACKGROUND: The prediction of protein subcellular localization is a key step of the big effort towards protein functional annotation. Many computational methods exist to identify high-level protein subcellular compartments such as nucleus, cytoplasm or organelles. However, many organelles, like mitochondria, have their own internal compartmentalization. Knowing the precise location of a protein inside mitochondria is crucial for its accurate functional characterization. We recently developed DeepMito, a new method based on a 1-Dimensional Convolutional Neural Network (1D-CNN) architecture outperforming other similar approaches available in literature. RESULTS: Here, we explore the adoption of DeepMito for the large-scale annotation of four sub-mitochondrial localizations on mitochondrial proteomes of five different species, including human, mouse, fly, yeast and Arabidopsis thaliana. A significant fraction of the proteins from these organisms lacked experimental information about sub-mitochondrial localization. We adopted DeepMito to fill the gap, providing complete characterization of protein localization at sub-mitochondrial level for each protein of the five proteomes. Moreover, we identified novel mitochondrial proteins fishing on the set of proteins lacking any subcellular localization annotation using available state-of-the-art subcellular localization predictors. We finally performed additional functional characterization of proteins predicted by DeepMito as localized into the four different sub-mitochondrial compartments using both available experimental and predicted GO terms. All data generated in this study were collected into a database called DeepMitoDB (available at http://busca.biocomp.unibo.it/deepmitodb ), providing complete functional characterization of 4307 mitochondrial proteins from the five species. CONCLUSIONS: DeepMitoDB offers a comprehensive view of mitochondrial proteins, including experimental and predicted fine-grain sub-cellular localization and annotated and predicted functional annotations. The database complements other similar resources providing characterization of new proteins. Furthermore, it is also unique in including localization information at the sub-mitochondrial level. For this reason, we believe that DeepMitoDB can be a valuable resource for mitochondrial research.
Castrense Savojardo, Pier Luigi Martelli, Giacomo Tartari, Rita Casadio
BMC Bioinform.2
2019 A natural upper bound to the accuracy of predicting protein stability changes upon mutations
abstract
MOTIVATION: Accurate prediction of protein stability changes upon single-site variations (ΔΔG) is important for protein design, as well as for our understanding of the mechanisms of genetic diseases. The performance of high-throughput computational methods to this end is evaluated mostly based on the Pearson correlation coefficient between predicted and observed data, assuming that the upper bound would be 1 (perfect correlation). However, the performance of these predictors can be limited by the distribution and noise of the experimental data. Here we estimate, for the first time, a theoretical upper-bound to the ΔΔG prediction performances imposed by the intrinsic structure of currently available ΔΔG data. RESULTS: Given a set of measured ΔΔG protein variations, the theoretically "best predictor" is estimated based on its similarity to another set of experimentally determined ΔΔG values. We investigate the correlation between pairs of measured ΔΔG variations, where one is used as a predictor for the other. We analytically derive an upper bound to the Pearson correlation as a function of the noise and distribution of the ΔΔG data. We also evaluate the available datasets to highlight the effect of the noise in conjunction with ΔΔG distribution. We conclude that the upper bound is a function of both uncertainty and spread of the ΔΔG values, and that with current data the best performance should be between 0.7 and 0.8, depending on the dataset used; higher Pearson correlations might be indicative of overtraining. It also follows that comparisons of predictors using different datasets are inherently misleading. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Ludovica Montanucci, Pier Luigi Martelli, Nir Ben-Tal, Piero Fariselli
Bioinform.2
2019 On the biases in predictions of protein stability changes upon variations: the INPS test case
abstract
INPS (Fariselli et al., 2015) is a simple method to predict the protein stability changes upon single point variation. It is based on Support Vector Machine and it was trained using 7 input features that are: (i) Blosum62 substitution, (ii) residue mutability, (iii) molecular weight wild type, (iv) molecular weight variant, (v) hydrophobicity wild type, (vi) hydrophobicity variant, (vii) difference between Viterbi scores of the wild type and variant (Fariselli et al., 2015). The first two input features are not anti-symmetric, the last is anti-symmetric by construction, while the other features can be learned to be anti-symmetric. To enforce anti-symmetry INPS was trained using both direct and inverse variations (Fariselli et al., 2015). When the alignment is poor we can expect a deviation from the anti-symmetry. INPS3D (Savojardo et al., 2016), adds to INPS information about contact potential and solvent accessibility. The solvent accessibility is intrinsically non anti-symmetric, and for this reason INPS3D was trained only with direct variations (Savojardo et al., 2016). The Ssym dataset is a manually curated selection of variations from the ProTherm database (Pucci et al., 2018). It contains variations with experimental ΔΔG values for which the 3D structures of both the wild-type and variant proteins were solved by X-ray crystallography. Ssym consists of 684 variations, half of which are direct (reported in the literature) and half are obtained by anti-symmetry (Pucci et al., 2018). To overcome the bias problem, a new method, PopMuSiCSYM, was explicitly developed to furnish completely anti-symmetric predictions of the variations (Pucci et al., 2015, 2018). The authors of the Ssym dataset estimated the performance of fifteen selected predictors using the average bias parameter (<δ> = ∑(ΔΔGdir+ΔΔGinv⁠)/N) and the linear correlation coefficient (rdir-inv) between the predicted ΔΔG values of the direct and the corresponding inverse variations. A perfect predictor should have bias close to zero (<δ> = 0) and a correlation coefficient close to -1 (rdir-inv = −1). In the same paper, they also tested the performances of these fifteen predictors using the experimental ΔΔG values on both direct and inverse variations using the root mean square deviation (σdir, σinv) and the linear correlation coefficient between the predicted and experimental values (rdir, rinv). In Table 1 we report the predictor performances and biases taken from the original paper (Pucci et al., 2018) and we add those computed using INPS and INPS3D. From Table 1, it is clear that the partial anti-symmetry of the INPS input and its training performed on both direct and inverse variations paid off. INPS shows a very low bias (<δ>, just second best), the highest correlation between direct and inverse variations (rdir-inv) and the highest performance on the inverse variation sets (σinv and rinv). The graph of the correlation between direct and inverse variation predicted by INPS is presented in Figure 1. INPS and INPS3D prediction on Ssym. The INPS (black) and INPS3D (grey) predictions for the 342 pairs (direct and inverse) of single-point variations in Ssym (Pucci et al., 2018) are shown. ΔΔG for the direct (x-axis) and inverse (y-axis) variations are reported as dots in the graph. The ‘ideal’ relationship ΔΔGAB + ΔΔGBA = 0 is shown as a solid line Bias analysis on the Ssym dataset (Pucci et al. 2018) Notes: Standard deviations σ and the average bias <δ> are in kcal/mol. All the values, except those of INPS and INPS3D, are taken from (Pucci et al., 2018). Bold character highlights INPS method. Bias analysis on the Ssym dataset (Pucci et al. 2018) Notes: Standard deviations σ and the average bias <δ> are in kcal/mol. All the values, except those of INPS and INPS3D, are taken from (Pucci et al., 2018). Bold character highlights INPS method. The INPS training set includes some of the proteins in the Ssym dataset (which is also true for all the other tested methods reported in Table 1). For sake of comparison, we also computed the performances of a ‘blind’ INPS. In this case, each ΔΔG prediction is obtained by using a svm model that does not include the protein to predict or a similar one (sequence identity <25%) in the training set. The standard deviations and Pearson correlation coefficients computed for the blind-INPS are: 1.44 kcal/mol and 0.48 for direct variations and 1.45 kcal/mol and 0.47 for inverse variations. The correlation between the predictions of direct and inverse variations is -0.99 while the average bias (measured through the δ) is -0.06 kcal/mol. The structure-based INPS3D has the second best correlation between direct and inverse correlation but a considerably higher bias. This is expected given that INPS3D has not been trained on inverse variations and conains as input feature the solvent accessibility, which is not anti-symmetric. Finally, it must be stressed that the correlations reported in Table 1 are only intended to assess the anti-symmetricity of the methods, and not as estimators of their overall performances, which is a very difficult task that also requires an evaluation of stability of the dataset (Yang et al., 2018). A second dataset was built by Usmanova et al. (2018), by extracting high-resolution pairs of proteins from the Protein Data Bank (PDB) differing by one to ten amino acids. Here we test INPS on the subset of all the single-site variations extracted from the Usmanova et al. (2018) dataset, that is, all the pdb pairs whose sequences differ by exactly one residue. This subset comprises 1000 pairs of single-point protein variations. Although for this set the experimental ΔΔG values are not known, given its important size, it is very informative for assessing predictor biases. In Usmanova et al. (2018), bias for an individual variation is estimated as: =(ΔΔGAB + ΔΔGBA)/2. In Table 2 we add the INPS and INPS3D anti-symmetry indexes to those previously computed on other predictors (Usmanova et al., 2018). The bias of INPS is the smallest and the correlation between direct and inverse variations is the highest. INPS3D, while showing less anti-symmetricity than INPS, has the second best bias and the second best correlation. Bias analysis on the Usmanova et al. (2018) dataset Notes: First column: mean and standard error of the mean for the biases. Second column: r, Pearson correlation coefficient with the associated P-value, between direct and inverse variations. All the values except INPS and INPS3D, are taken from Usmanova et al. (2018). 0.0*: P-value set to 0, since it is undetectable by the machine precision. Bold character highlights INPS method. Bias analysis on the Usmanova et al. (2018) dataset Notes: First column: mean and standard error of the mean for the biases. Second column: r, Pearson correlation coefficient with the associated P-value, between direct and inverse variations. All the values except INPS and INPS3D, are taken from Usmanova et al. (2018). 0.0*: P-value set to 0, since it is undetectable by the machine precision. Bold character highlights INPS method. The very high correlation of -0.95 between direct and inverse variations obtained by INPS on this set is shown in the graph of Figure 2, to be contrasted with those reported for the other methods in the figures of the paper by Usmanova et al. (2018). In Figure 2, the points very far from the diagonal correspond to variations whose protein sequences are poorly aligned (number of aligned sequences <20). INPS and INPS3D prediction for the 1000 pairs (direct and inverse) of single-point variations (Usmanova et al., 2018). ΔΔG for the direct (x-axis) and inverse (y-axis) variations are reported as dots in the graph. The ‘ideal’ relationship ΔΔGAB + ΔΔGBA = 0 is shown as a solid line Recently, two groups (Pucci et al., 2018; Usmanova et al., 2018) compiled two datasets to test a very important bias that affects most of the computational methods designed to predict free energy changes upon protein variation. This bias is the lack of anti-symmetry in the ΔΔG predictions between direct and inverse variations. These studies verified this bias on several relevant methods. Here we add the test on INPS (Fariselli et al., 2015) that was specifically designed to take into account anti-symmetry problem using an input that was partially anti-symmetric (the HMM Viterbi score in particular). Here we show that INPS performs very well on both datasets proposed in the two papers. The fact that INPS uses only sequence information makes it less sensitive to structural rearrangement upon residue substitution and in turn more robust and suitable for testing any kind of protein variations. This work has been supported by EBA-PRISM an Israel-Italy collaborative project, the Israel Ministry of Science and Technology and Italian Ministry of Foreign Affair and International Cooperation. This research received funding specifically appointed to Department of Medical Sciences from the Italian Ministry for Education, University and Research (MIUR) under the programme “Dipartimenti di Eccellenza 2018 – 2022” Project code D15D18000410001. Conflict of Interest: none declared.
Ludovica Montanucci, Castrense Savojardo, Pier Luigi Martelli, Rita Casadio, Piero Fariselli
Bioinform.3
2018 DeepSig: deep learning improves signal peptide detection in proteins
abstract
Motivation: The identification of signal peptides in protein sequences is an important step toward protein localization and function characterization. Results: Here, we present DeepSig, an improved approach for signal peptide detection and cleavage-site prediction based on deep learning methods. Comparative benchmarks performed on an updated independent dataset of proteins show that DeepSig is the current best performing method, scoring better than other available state-of-the-art approaches on both signal peptide detection and precise cleavage-site identification. Availability and implementation: DeepSig is available as both standalone program and web server at https://deepsig.biocomp.unibo.it. All datasets used in this study can be obtained from the same website. Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Castrense Savojardo, Pier Luigi Martelli, Piero Fariselli, Rita Casadio
Bioinform.2
2017 ISPRED4: interaction sites PREDiction in protein structures with a refining grammar model
abstract
MOTIVATION: The identification of protein-protein interaction (PPI) sites is an important step towards the characterization of protein functional integration in the cell complexity. Experimental methods are costly and time-consuming and computational tools for predicting PPI sites can fill the gaps of PPI present knowledge. RESULTS: We present ISPRED4, an improved structure-based predictor of PPI sites on unbound monomer surfaces. ISPRED4 relies on machine-learning methods and it incorporates features extracted from protein sequence and structure. Cross-validation experiments are carried out on a new dataset that includes 151 high-resolution protein complexes and indicate that ISPRED4 achieves a per-residue Matthew Correlation Coefficient of 0.48 and an overall accuracy of 0.85. Benchmarking results show that ISPRED4 is one of the top-performing PPI site predictors developed so far. CONTACT: [email protected]. AVAILABILITY AND IMPLEMENTATION: ISPRED4 and datasets used in this study are available at http://ispred4.biocomp.unibo.it .
Castrense Savojardo, Piero Fariselli, Pier Luigi Martelli, Rita Casadio
Bioinform.3
2017 SChloro: directing Viridiplantae proteins to six chloroplastic sub-compartments
abstract
Motivation: Chloroplasts are organelles found in plants and involved in several important cell processes. Similarly to other compartments in the cell, chloroplasts have an internal structure comprising several sub-compartments, where different proteins are targeted to perform their functions. Given the relation between protein function and localization, the availability of effective computational tools to predict protein sub-organelle localizations is crucial for large-scale functional studies. Results: In this paper we present SChloro, a novel machine-learning approach to predict protein sub-chloroplastic localization, based on targeting signal detection and membrane protein information. The proposed approach performs multi-label predictions discriminating six chloroplastic sub-compartments that include inner membrane, outer membrane, stroma, thylakoid lumen, plastoglobule and thylakoid membrane. In comparative benchmarks, the proposed method outperforms current state-of-the-art methods in both single- and multi-compartment predictions, with an overall multi-label accuracy of 74%. The results demonstrate the relevance of the approach that is eligible as a good candidate for integration into more general large-scale annotation pipelines of protein subcellular localization. Availability and Implementation: The method is available as web server at http://schloro.biocomp.unibo.it Contact: [email protected].
Castrense Savojardo, Pier Luigi Martelli, Piero Fariselli, Rita Casadio
Bioinform.2
2016 NET-GE: a web-server for NETwork-based human gene enrichment
abstract
MOTIVATION: Gene enrichment is a requisite for the interpretation of biological complexity related to specific molecular pathways and biological processes. Furthermore, when interpreting NGS data and human variations, including those related to pathologies, gene enrichment allows the inclusion of other genes that in the human interactome space may also play important key roles in the emergency of the phenotype. Here, we describe NET-GE, a web server for associating biological processes and pathways to sets of human proteins involved in the same phenotype RESULTS: NET-GE is based on protein-protein interaction networks, following the notion that for a set of proteins, the context of their specific interactions can better define their function and the processes they can be related to in the biological complexity of the cell. Our method is suited to extract statistically validated enriched terms from Gene Ontology, KEGG and REACTOME annotation databases. Furthermore, NET-GE is effective even when the number of input proteins is small. AVAILABILITY AND IMPLEMENTATION: NET-GE web server is publicly available and accessible at http://net-ge.biocomp.unibo.it/enrich CONTACT: [email protected] information: Supplementary data are available at Bioinformatics online.
Samuele Bovo, Pietro Di Lena, Pier Luigi Martelli, Piero Fariselli, Rita Casadio
Bioinform.3
2016 INPS-MD: a web server to predict stability of protein variants from sequence and structure
abstract
MOTIVATION: Protein function depends on its structural stability. The effects of single point variations on protein stability can elucidate the molecular mechanisms of human diseases and help in developing new drugs. Recently, we introduced INPS, a method suited to predict the effect of variations on protein stability from protein sequence and whose performance is competitive with the available state-of-the-art tools. RESULTS: In this article, we describe INPS-MD (Impact of Non synonymous variations on Protein Stability-Multi-Dimension), a web server for the prediction of protein stability changes upon single point variation from protein sequence and/or structure. Here, we complement INPS with a new predictor (INPS3D) that exploits features derived from protein 3D structure. INPS3D scores with Pearson's correlation to experimental ΔΔG values of 0.58 in cross validation and of 0.72 on a blind test set. The sequence-based INPS scores slightly lower than the structure-based INPS3D and both on the same blind test sets well compare with the state-of-the-art methods. AVAILABILITY AND IMPLEMENTATION: INPS and INPS3D are available at the same web server: http://inpsmd.biocomp.unibo.it SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. CONTACT: [email protected].
Castrense Savojardo, Piero Fariselli, Pier Luigi Martelli, Rita Casadio
Bioinform.3
2015 INPS: predicting the impact of non-synonymous variations on protein stability from sequence
abstract
MOTIVATION: A tool for reliably predicting the impact of variations on protein stability is extremely important for both protein engineering and for understanding the effects of Mendelian and somatic mutations in the genome. Next Generation Sequencing studies are constantly increasing the number of protein sequences. Given the huge disproportion between protein sequences and structures, there is a need for tools suited to annotate the effect of mutations starting from protein sequence without relying on the structure. Here, we describe INPS, a novel approach for annotating the effect of non-synonymous mutations on the protein stability from its sequence. INPS is based on SVM regression and it is trained to predict the thermodynamic free energy change upon single-point variations in protein sequences. RESULTS: We show that INPS performs similarly to the state-of-the-art methods based on protein structure when tested in cross-validation on a non-redundant dataset. INPS performs very well also on a newly generated dataset consisting of a number of variations occurring in the tumor suppressor protein p53. Our results suggest that INPS is a tool suited for computing the effect of non-synonymous polymorphisms on protein stability when the protein structure is not available. We also show that INPS predictions are complementary to those of the state-of-the-art, structure-based method mCSM. When the two methods are combined, the overall prediction on the p53 set scores significantly higher than those of the single methods. AVAILABILITY AND IMPLEMENTATION: The presented method is available as web server at http://inps.biocomp.unibo.it. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary Materials are available at Bioinformatics online.
Piero Fariselli, Pier Luigi Martelli, Castrense Savojardo, Rita Casadio
Bioinform.2
2015 TPpred3 detects and discriminates mitochondrial and chloroplastic targeting peptides in eukaryotic proteins
abstract
MOTIVATION: Molecular recognition of N-terminal targeting peptides is the most common mechanism controlling the import of nuclear-encoded proteins into mitochondria and chloroplasts. When experimental information is lacking, computational methods can annotate targeting peptides, and determine their cleavage sites for characterizing protein localization, function, and mature protein sequences. The problem of discriminating mitochondrial from chloroplastic propeptides is particularly relevant when annotating proteomes of photosynthetic Eukaryotes, endowed with both types of sequences. RESULTS: Here, we introduce TPpred3, a computational method that given any Eukaryotic protein sequence performs three different tasks: (i) the detection of targeting peptides; (ii) their classification as mitochondrial or chloroplastic and (iii) the precise localization of the cleavage sites in an organelle-specific framework. Our implementation is based on our TPpred previously introduced. Here, we integrate a new N-to-1 Extreme Learning Machine specifically designed for the classification task (ii). For the last task, we introduce an organelle-specific Support Vector Machine that exploits sequence motifs retrieved with an extensive motif-discovery analysis of a large set of mitochondrial and chloroplastic proteins. We show that TPpred3 outperforms the state-of-the-art methods in all the three tasks. AVAILABILITY AND IMPLEMENTATION: The method server and datasets are available at http://tppred3.biocomp.unibo.it. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Castrense Savojardo, Pier Luigi Martelli, Piero Fariselli, Rita Casadio
Bioinform.2
2014 TPpred2: improving the prediction of mitochondrial targeting peptide cleavage sites by exploiting sequence motifs
abstract
SUMMARY: Targeting peptides are N-terminal sorting signals in proteins that promote their translocation to mitochondria through the interaction with different protein machineries. We recently developed TPpred, a machine learning-based method scoring among the best ones available to predict the presence of a targeting peptide into a protein sequence and its cleavage site. Here we introduce TPpred2 that improves TPpred performances in the task of identifying the cleavage site of the targeting peptides. TPpred2 is now available as a web interface and as a stand-alone version for users who can freely download and adopt it for processing large volumes of sequences. Availability and implementaion: TPpred2 is available both as web server and stand-alone version at http://tppred2.biocomp.unibo.it. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Castrense Savojardo, Pier Luigi Martelli, Piero Fariselli, Rita Casadio
Bioinform.2
2013 The prediction of organelle-targeting peptides in eukaryotic proteins with Grammatical-Restrained Hidden Conditional Random Fields
abstract
MOTIVATION: Targeting peptides are the most important signal controlling the import of nuclear encoded proteins into mitochondria and plastids. In the lack of experimental information, their prediction is an essential step when proteomes are annotated for inferring both the localization and the sequence of mature proteins. RESULTS: We developed TPpred a new predictor of organelle-targeting peptides based on Grammatical-Restrained Hidden Conditional Random Fields. TPpred is trained on a non-redundant dataset of proteins where the presence of a target peptide was experimentally validated, comprising 297 sequences. When tested on the 297 positive and some other 8010 negative examples, TPpred outperformed available methods in both accuracy and Matthews correlation index (96% and 0.58, respectively). Given its very low-false-positive rate (3.0%), TPpred is, therefore, well suited for large-scale analyses at the proteome level. We predicted that from ∼4 to 9% of the sequences of human, Arabidopsis thaliana and yeast proteomes contain targeting peptides and are, therefore, likely to be localized in mitochondria and plastids. TPpred predictions correlate to a good extent with the experimental annotation of the subcellular localization, when available. TPpred was also trained and tested to predict the cleavage site of the organelle-targeting peptide: on this task, the average error of TPpred on mitochondrial and plastidic proteins is 7 and 15 residues, respectively. This value is lower than the error reported by other methods currently available. AVAILABILITY: The TPpred datasets are available at http://biocomp.unibo.it/valentina/TPpred/. TPpred is available on request from the authors. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Valentina Indio, Pier Luigi Martelli, Castrense Savojardo, Piero Fariselli, Rita Casadio
Bioinform.2
2013 BCov: a method for predicting β-sheet topology using sparse inverse covariance estimation and integer programming
abstract
MOTIVATION: Prediction of protein residue contacts, even at the coarse-grain level, can help in finding solutions to the protein structure prediction problem. Unlike α-helices that are locally stabilized, β-sheets result from pairwise hydrogen bonding of two or more disjoint regions of the protein backbone. The problem of predicting contacts among β-strands in proteins has been addressed by several supervised computational approaches. Recently, prediction of residue contacts based on correlated mutations has been greatly improved and finally allows the prediction of 3D structures of the proteins. RESULTS: In this article, we describe BCov, which is the first unsupervised method to predict the β-sheet topology starting from the protein sequence and its secondary structure. BCov takes advantage of the sparse inverse covariance estimation to define β-strand partner scores. Then an optimization based on integer programming is carried out to predict the β-sheet connectivity. When tested on the prediction of β-strand pairing, BCov scores with average values of Matthews Correlation Coefficient (MCC) and F1 equal to 0.56 and 0.61, respectively, on a non-redundant dataset of 916 protein chains known with atomic resolution. Our approach well compares with the state-of-the-art methods trained so far for this specific task. AVAILABILITY AND IMPLEMENTATION: The method is freely available under General Public License at http://biocomp.unibo.it/savojard/bcov/bcov-1.0.tar.gz. The new dataset BetaSheet1452 can be downloaded at http://biocomp.unibo.it/savojard/bcov/BetaSheet1452.dat.
Castrense Savojardo, Piero Fariselli, Pier Luigi Martelli, Rita Casadio
Bioinform.3
2013 MIMO: an efficient tool for molecular interaction maps overlap
abstract
BACKGROUND: Molecular pathways represent an ensemble of interactions occurring among molecules within the cell and between cells. The identification of similarities between molecular pathways across organisms and functions has a critical role in understanding complex biological processes. For the inference of such novel information, the comparison of molecular pathways requires to account for imperfect matches (flexibility) and to efficiently handle complex network topologies. To date, these characteristics are only partially available in tools designed to compare molecular interaction maps. RESULTS: Our approach MIMO (Molecular Interaction Maps Overlap) addresses the first problem by allowing the introduction of gaps and mismatches between query and template pathways and permits -when necessary- supervised queries incorporating a priori biological information. It then addresses the second issue by relying directly on the rich graph topology described in the Systems Biology Markup Language (SBML) standard, and uses multidigraphs to efficiently handle multiple queries on biological graph databases. The algorithm has been here successfully used to highlight the contact point between various human pathways in the Reactome database. CONCLUSIONS: MIMO offers a flexible and efficient graph-matching tool for comparing complex biological pathways.
Pietro Di Lena, Pier Luigi Martelli, Rita Casadio, Christine Nardini
BMC Bioinform.3
2013 How to inherit statistically validated annotation within BAR+ protein clusters
abstract
BACKGROUND: In the genomic era a key issue is protein annotation, namely how to endow protein sequences, upon translation from the corresponding genes, with structural and functional features. Routinely this operation is electronically done by deriving and integrating information from previous knowledge. The reference database for protein sequences is UniProtKB divided into two sections, UniProtKB/TrEMBL which is automatically annotated and not reviewed and UniProtKB/Swiss-Prot which is manually annotated and reviewed. The annotation process is essentially based on sequence similarity search. The question therefore arises as to which extent annotation based on transfer by inheritance is valuable and specifically if it is possible to statistically validate inherited features when little homology exists among the target sequence and its template(s). RESULTS: In this paper we address the problem of annotating protein sequences in a statistically validated manner considering as a reference annotation resource UniProtKB. The test case is the set of 48,298 proteins recently released by the Critical Assessment of Function Annotations (CAFA) organization. We show that we can transfer after validation, Gene Ontology (GO) terms of the three main categories and Pfam domains to about 68% and 72% of the sequences, respectively. This is possible after alignment of the CAFA sequences towards BAR+, our annotation resource that allows discriminating among statistically validated and not statistically validated annotation. By comparing with a direct UniProtKB annotation, we find that besides validating annotation of some 78% of the CAFA set, we assign new and statistically validated annotation to 14.8% of the sequences and find new structural templates for about 25% of the chains, half of which share less than 30% sequence identity to the corresponding template/s. CONCLUSION: Inheritance of annotation by transfer generally requires a careful selection of the identity value among the target and the template in order to transfer structural and/or functional features. Here we prove that even distantly remote homologs can be safely endowed with structural templates and GO and/or Pfam terms provided that annotation is done within clusters collecting cluster-related protein sequences and where a statistical validation of the shared structural and functional features is possible.
Damiano Piovesan, Pier Luigi Martelli, Piero Fariselli, Giuseppe Profiti, Andrea Zauli, Ivan Rossi, Rita Casadio
BMC Bioinform.2
2013 Prediction of disulfide connectivity in proteins with machine-learning methods and correlated mutations
abstract
BACKGROUND: Recently, information derived by correlated mutations in proteins has regained relevance for predicting protein contacts. This is due to new forms of mutual information analysis that have been proven to be more suitable to highlight direct coupling between pairs of residues in protein structures and to the large number of protein chains that are currently available for statistical validation. It was previously discussed that disulfide bond topology in proteins is also constrained by correlated mutations. RESULTS: In this paper we exploit information derived from a corrected mutual information analysis and from the inverse of the covariance matrix to address the problem of the prediction of the topology of disulfide bonds in Eukaryotes. Recently, we have shown that Support Vector Regression (SVR) can improve the prediction for the disulfide connectivity patterns. Here we show that the inclusion of the correlated mutation information increases of 5 percentage points the SVR performance (from 54% to 59%). When this approach is used in combination with a method previously developed by us and scoring at the state of art in predicting both location and topology of disulfide bonds in Eukaryotes (DisLocate), the per-protein accuracy is 38%, 2 percentage points higher than that previously obtained. CONCLUSIONS: In this paper we show that the inclusion of information derived from correlated mutations can improve the performance of the state of the art methods for predicting disulfide connectivity patterns in Eukaryotic proteins. Our analysis also provides support to the notion that improving methods to extract evolutionary information from multiple sequence alignments greatly contributes to the scoring performance of predictors suited to detect relevant features from protein chains.
Castrense Savojardo, Piero Fariselli, Pier Luigi Martelli, Rita Casadio
BMC Bioinform.3
2013 Extended and Robust Protein Sequence Annotation over Conservative Nonhierarchical Clusters: The Case Study of the ABC Transporters
abstract
Genome annotation is one of the most important issues in the genomic era. The exponential growth rate of newly sequenced genomes and proteomes urges the development of fast and reliable annotation methods, suited to exploit all the information available in curated databases of protein sequences and structures. To this aim we developed BAR+, the Bologna Annotation Resource. 1 The basic notion is that sequences with high identity value to a counterpart can inherit the same function/s and structure, if available. As a case study we describe how the ATP-binding domain of the ABC transporters can be found and modeled in over 30,000 new sequences not annotated before. We also mapped into BAR+ all the ABC transporters listed in the Transporter Classification DataBase 2 and found that within our environment annotation could be extended to another 256,866 sequences.
Damiano Piovesan, Giuseppe Profiti, Pier Luigi Martelli, Piero Fariselli, Rita Casadio
ACM J. Emerg. Technol. Comput. Syst.3
2012 The human "magnesome": detecting magnesium binding sites on human proteins
abstract
BACKGROUND: Magnesium research is increasing in molecular medicine due to the relevance of this ion in several important biological processes and associated molecular pathogeneses. It is still difficult to predict from the protein covalent structure whether a human chain is or not involved in magnesium binding. This is mainly due to little information on the structural characteristics of magnesium binding sites in proteins and protein complexes. Magnesium binding features, differently from those of other divalent cations such as calcium and zinc, are elusive. Here we address a question that is relevant in protein annotation: how many human proteins can bind Mg2+? Our analysis is performed taking advantage of the recently implemented Bologna Annotation Resource (BAR-PLUS), a non hierarchical clustering method that relies on the pair wise sequence comparison of about 14 millions proteins from over 300.000 species and their grouping into clusters where annotation can safely be inherited after statistical validation. RESULTS: After cluster assignment of the latest version of the human proteome, the total number of human proteins for which we can assign putative Mg binding sites is 3,751. Among these proteins, 2,688 inherit annotation directly from human templates and 1,063 inherit annotation from templates of other organisms. Protein structures are highly conserved inside a given cluster. Transfer of structural properties is possible after alignment of a given sequence with the protein structures that characterise a given cluster as obtained with a Hidden Markov Model (HMM) based procedure. Interestingly a set of 370 human sequences inherit Mg2+ binding sites from templates sharing less than 30% sequence identity with the template. CONCLUSION: We describe and deliver the "human magnesome", a set of proteins of the human proteome that inherit putative binding of magnesium ions. With our BAR-hMG, 251 clusters including 1,341 magnesium binding protein structures corresponding to 387 sequences are sufficient to annotate some 13,689 residues in 3,751 human sequences as "magnesium binding". Protein structures act therefore as three dimensional seeds for structural and functional annotation of human sequences. The data base collects specifically all the human proteins that can be annotated according to our procedure as "magnesium binding", the corresponding structures and BAR+ clusters from where they derive the annotation (http://bar.biocomp.unibo.it/mg).
Damiano Piovesan, Giuseppe Profiti, Pier Luigi Martelli, Rita Casadio
BMC Bioinform.3
2011 MemLoci: predicting subcellular localization of membrane proteins in eukaryotes
abstract
MOTIVATION: Subcellular localization is a key feature in the process of functional annotation of both globular and membrane proteins. In the absence of experimental data, protein localization is inferred on the basis of annotation transfer upon sequence similarity search. However, predictive tools are necessary when the localization of homologs is not known. This is so particularly for membrane proteins. Furthermore, most of the available predictors of subcellular localization are specifically trained on globular proteins and poorly perform on membrane proteins. RESULTS: Here we develop MemLoci, a new support vector machine-based tool that discriminates three membrane protein localizations: plasma, internal and organelle membrane. When tested on an independent set, MemLoci outperforms existing methods, reaching an overall accuracy of 70% on predicting the location in the three membrane types, with a generalized correlation coefficient as high as 0.50. AVAILABILITY: The MemLoci server is freely available on the web at: http://mu2py.biocomp.unibo.it/memloci. Datasets described in the article can be downloaded at the same site.
Andrea Pierleoni, Pier Luigi Martelli, Rita Casadio
Bioinform.2
2011 Improving the prediction of disulfide bonds in Eukaryotes with machine learning methods and protein subcellular localization
abstract
MOTIVATION: Disulfide bonds stabilize protein structures and play relevant roles in their functions. Their formation requires an oxidizing environment and their stability is consequently depending on the redox ambient potential, which may differ according to the subcellular compartment. Several methods are available to predict cysteine-bonding state and connectivity patterns. However, none of them takes into consideration the relevance of protein subcellular localization. RESULTS: Here we develop DISLOCATE, a two-step method based on machine learning models for predicting both the bonding state and the connectivity patterns of cysteine residues in a protein chain. We find that the inclusion of protein subcellular localization improves the performance of these predictive steps by 3 and 2 percentage points, respectively. When compared with previously developed methods for predicting disulfide bonds from sequence, DISLOCATE improves the overall performance by more than 10 percentage points. AVAILABILITY: The method and the dataset are available at the Web page http://www.biocomp.unibo.it/savojard/Dislocate.html. GRHCRF code is available at http://www.biocomp.unibo.it/savojard/biocrf.html. CONTACT: [email protected].
Castrense Savojardo, Piero Fariselli, Monther Alhamdoosh, Pier Luigi Martelli, Andrea Pierleoni, Rita Casadio
Bioinform.4
2008 Predicting protein thermostability changes from sequence upon multiple mutations
abstract
MOTIVATION: A basic question in protein science is to which extent mutations affect protein thermostability. This knowledge would be particularly relevant for engineering thermostable enzymes. In several experimental approaches, this issue has been serendipitously addressed. It would be therefore convenient providing a computational method that predicts when a given protein mutant is more thermostable than its corresponding wild-type. RESULTS: We present a new method based on support vector machines that is able to predict whether a set of mutations (including insertion and deletions) can enhance the thermostability of a given protein sequence. When trained and tested on a redundancy-reduced dataset, our predictor achieves 88% accuracy and a correlation coefficient equal to 0.75. Our predictor also correctly classifies 12 out of 14 experimentally characterized protein mutants with enhanced thermostability. Finally, it correctly detects all the 11 mutated proteins whose increase in stability temperature is >10 degrees C. AVAILABILITY: The dataset and the list of protein clusters adopted for the SVM cross-validation are available at the web site http://lipid.biocomp.unibo.it/~ludovica/thermo-meso-MUT.
Ludovica Montanucci, Piero Fariselli, Pier Luigi Martelli, Rita Casadio
ISMB3
2008 PredGPI: a GPI-anchor predictor
abstract
BACKGROUND: Several eukaryotic proteins associated to the extracellular leaflet of the plasma membrane carry a Glycosylphosphatidylinositol (GPI) anchor, which is linked to the C-terminal residue after a proteolytic cleavage occurring at the so called omega-site. Computational methods were developed to discriminate proteins that undergo this post-translational modification starting from their aminoacidic sequences. However more accurate methods are needed for a reliable annotation of whole proteomes. RESULTS: Here we present PredGPI, a prediction method that, by coupling a Hidden Markov Model (HMM) and a Support Vector Machine (SVM), is able to efficiently predict both the presence of the GPI-anchor and the position of the omega-site. PredGPI is trained on a non-redundant dataset of experimentally characterized GPI-anchored proteins whose annotation was carefully checked in the literature. CONCLUSION: PredGPI outperforms all the other previously described methods and is able to correctly replicate the results of previously published high-throughput experiments. PredGPI reaches a lower rate of false positive predictions with respect to other available methods and it is therefore a costless, rapid and accurate method for screening whole proteomes.
Andrea Pierleoni, Pier Luigi Martelli, Rita Casadio
BMC Bioinform.2
2005 A new decoding algorithm for hidden Markov models improves the prediction of the topology of all-beta membrane proteins
abstract
BACKGROUND: Structure prediction of membrane proteins is still a challenging computational problem. Hidden Markov models (HMM) have been successfully applied to the problem of predicting membrane protein topology. In a predictive task, the HMM is endowed with a decoding algorithm in order to assign the most probable state path, and in turn the labels, to an unknown sequence. The Viterbi and the posterior decoding algorithms are the most common. The former is very efficient when one path dominates, while the latter, even though does not guarantee to preserve the HMM grammar, is more effective when several concurring paths have similar probabilities. A third good alternative is 1-best, which was shown to perform equal or better than Viterbi. RESULTS: In this paper we introduce the posterior-Viterbi (PV) a new decoding which combines the posterior and Viterbi algorithms. PV is a two step process: first the posterior probability of each state is computed and then the best posterior allowed path through the model is evaluated by a Viterbi algorithm. CONCLUSION: We show that PV decoding performs better than other algorithms when tested on the problem of the prediction of the topology of beta-barrel membrane proteins.
Piero Fariselli, Pier Luigi Martelli, Rita Casadio
BMC Bioinform.2
2003 In silico prediction of the structure of membrane proteins: Is it feasible?
abstract
In the 'omic' era, hundreds of genomes are available for protein sequence analysis, and some 30 per cent of all sequences are of membrane proteins. Unlike globular proteins, a 3D model for membrane proteins can hardly be computed starting from the sequence. Why is this so? What can we really compute and with what reliability? These and other matters are outlined.
Rita Casadio, Piero Fariselli, Pier Luigi Martelli
Briefings Bioinform.3
2003 MaxSubSeq: an algorithm for segment-length optimization. The case study of the transmembrane spanning segments
abstract
MOTIVATION: A problem in predicting the topography of transmembrane proteins is the optimal localization of the transmembrane segments along the protein sequences, provided that each residue is associated with a propensity of being or not being included in the transmembrane protein region. From previous work it is known that post-processing of propensity signals with suited algorithms can greatly improve the quality and the accuracy of the predictions. In this paper we describe a general dynamic programming-like algorithm (MaxSubSeq, Maximal SubSequence) specifically designed to optimize the number and length of segments with constrained length in a given protein sequence. Previous application of our algorithm, has proved its effectiveness in the optimization task of both neural network and hidden Markov models output, and in this paper we present the detailed description of MaxSubSeq. RESULTS: We describe the application of MaxSubSeq to the location of both helical and beta strand transmembrane segments, optimizing the outputs derived with different predictive algorithms. For all-alpha transmembrane proteins we use both the standard Kyte-Doolittle (KD) hydropathy scale and the TMHMM predictor (http://www.cbs.dtu.dk/). Using a set of 188 well characterized membrane proteins, MaxSubSeq nearly doubles the correct location of transmembrane segments as compared to the standard KD hydrophobicity plot, reaching 51% accuracy. If MaxSubSeq is used to optimize the TMHMM method the accuracy increases from 68 to 72%. When used to regularize the prediction of beta transmembrane strands, obtained using both a neural network and a HMM based predictors, MaxSubSeq increases the accuracy per protein up to 72 and 73% respectively. AVAILABILITY: The program is available upon request to the authors, or it is accessible through our web server (http://gpcr.biocomp.unibo.it/predictors/)
Piero Fariselli, Michele Finelli, Davide Marchignoli, Pier Luigi Martelli, Ivan Rossi, Rita Casadio
Bioinform.4
2002 A sequence-profile-based HMM for predicting and discriminating beta barrel membrane proteins
abstract
MOTIVATION: Membrane proteins are an abundant and functionally relevant subset of proteins that putatively include from about 15 up to 30% of the proteome of organisms fully sequenced. These estimates are mainly computed on the basis of sequence comparison and membrane protein prediction. It is therefore urgent to develop methods capable of selecting membrane proteins especially in the case of outer membrane proteins, barely taken into consideration when proteome wide analysis is performed. This will also help protein annotation when no homologous sequence is found in the database. Outer membrane proteins solved so far at atomic resolution interact with the external membrane of bacteria with a characteristic beta barrel structure comprising different even numbers of beta strands (beta barrel membrane proteins). In this they differ from the membrane proteins of the cytoplasmic membrane endowed with alpha helix bundles (all alpha membrane proteins) and need specialised predictors. RESULTS: We develop a HMM model, which can predict the topology of beta barrel membrane proteins using, as input, evolutionary information. The model is cyclic with 6 types of states: two for the beta strand transmembrane core, one for the beta strand cap on either side of the membrane, one for the inner loop, one for the outer loop and one for the globular domain state in the middle of each loop. The development of a specific input for HMM based on multiple sequence alignment is novel. The accuracy per residue of the model is 83% when a jack knife procedure is adopted. With a model optimisation method using a dynamic programming algorithm seven topological models out of the twelve proteins included in the testing set are also correctly predicted. When used as a discriminator, the model is rather selective. At a fixed probability value, it retains 84% of a non-redundant set comprising 145 sequences of well-annotated outer membrane proteins. Concomitantly, it correctly rejects 90% of a set of globular proteins including about 1200 chains with low sequence identity (<30%) and 90% of a set of all alpha membrane proteins, including 188 chains.
Pier Luigi Martelli, Piero Fariselli, Anders Krogh, Rita Casadio
ISMB1
1999 A Data Base of Minimally Frustrated Alpha-Helical Segments Extracted from Proteins According to an Entropy Criterion
Rita Casadio, Mario Compiani, Piero Fariselli, Pier Luigi Martelli
ISMB4