EDBT 2026 Demo / reviewers in the wild / expert
Piero Fariselli
dblp:39/3328
· DBLP profile ↗
65ranked-venue papers
10as first author
13since 2021 · last 2026
0000-0003-1811-4762ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 57 · 10 first-author · 8 since 2021Artificial intelligence and machine learning · 5 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Systems, architecture and hardware · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PatientFlow: Learning to generate mixed-type longitudinal clinical data with flow matchingabstractSynthetic longitudinal clinical data, with static and temporal mixed-type components, can help unlock large-scale deep learning models to tackle complex diseases. However, learning to generate realistic patients faces dual challenges: modeling the inherently complex structure of longitudinal data and protecting patient privacy. We introduce PatientFlow, a generative modeling method combining Variational Autoencoders for data representation with Flow Matching for patient generation. We extensively evaluated the generative model on a longitudinal cohort of patients with Amyotrophic Lateral Sclerosis (N = 1560) using both qualitative and quantitative methods. The ability of the method to generate realistic patient data, further validated by expert clinicians, shows its potential application to other diseases. Prognostic models trained on synthetic data across five clinically relevant endpoints matched and sometimes outperformed the models trained on real data. Our results demonstrate that PatientFlow can effectively model longitudinal clinical data with high fidelity, opening promising avenues for sharing and augmenting datasets for deep learning applications in healthcare without compromising privacy. Ruben Branco, Marta Gromicho, Mamede de Carvalho, Piero Fariselli, Sara C. Madeira |
Artif. Intell. Medicine | 4 |
| 2024 | Survival Model Optimization via Federated Learning: A Study Combining Simulations and ExperimentsabstractFederated Learning is an emerging, powerful approach that allows training an artificial intelligence model in distributed setting. Two survival models, Cox and DeepSurv, have been trained in a federated setting, exploiting both code simulations and real experiments on the new platform, developed by the GenoMed4All consortium. Different scenarios have been tested by splitting a Myelodysplastic Syndrome dataset into three nodes and performing feature removal. A significant gain in model performance has been observed due to federated aggregation. Francesco Casadei, Luciana Carota, Gianluca Asti, Saverio D'Amico, Davide Piscia, Santiago Zazo, Patricia A. Apellániz, Juan Parras, Claudia Sala, Cesare Rollo, Nono S. C. Merleau, Piero Fariselli, Matteo Giovanni Della Porta, Tiziana Sanavia, Federico Álvarez García, Gastone C. Castellani, Enrico Giampieri |
IEEE Big Data | 12 |
| 2024 | iDPP@CLEF 2024: The Intelligent Disease Progression Prediction Challenge
Helena Aidos, Roberto Bergamaschi, Paola Cavalla, Adriano Chiò, Arianna Dagliati, Barbara Di Camillo, Mamede de Carvalho, Nicola Ferro 0001, Piero Fariselli, Jose Manuel García Dominguez, Sara C. Madeira, Eleonora Tavazzi |
ECIR (6) | 9 |
| 2024 | MUSE-XAE: MUtational Signature Extraction with eXplainable AutoEncoder enhances tumour types classificationabstractMOTIVATION: Mutational signatures are a critical component in deciphering the genetic alterations that underlie cancer development and have become a valuable resource to understand the genomic changes during tumorigenesis. Therefore, it is essential to employ precise and accurate methods for their extraction to ensure that the underlying patterns are reliably identified and can be effectively utilized in new strategies for diagnosis, prognosis, and treatment of cancer patients. RESULTS: We present MUSE-XAE, a novel method for mutational signature extraction from cancer genomes using an explainable autoencoder. Our approach employs a hybrid architecture consisting of a nonlinear encoder that can capture nonlinear interactions among features, and a linear decoder which ensures the interpretability of the active signatures. We evaluated and compared MUSE-XAE with other available tools on both synthetic and real cancer datasets and demonstrated that it achieves superior performance in terms of precision and sensitivity in recovering mutational signature profiles. MUSE-XAE extracts highly discriminative mutational signature profiles by enhancing the classification of primary tumour types and subtypes in real world settings. This approach could facilitate further research in this area, with neural networks playing a critical role in advancing our understanding of cancer genomics. AVAILABILITY AND IMPLEMENTATION: MUSE-XAE software is freely available at https://github.com/compbiomed-unito/MUSE-XAE. Corrado Pancotti, Cesare Rollo, Francesco Codicè, Giovanni Birolo, Piero Fariselli, Tiziana Sanavia |
Bioinform. | 5 |
| 2024 | Parallel intersection counting on shared-memory multiprocessors and GPUsabstractComputing intersections among sets of one-dimensional intervals is an ubiquitous problem in computational geometry with important applications in bioinformatics, where the size of typical inputs is large and it is therefore important to use efficient algorithms. In this paper we propose a parallel algorithm for the 1D intersection-counting problem, that is, the problem of counting the number of intersections between each interval in a given set A and every interval in a set B. Our algorithm is suitable for shared-memory architectures (e.g., multicore CPUs) and GPUs. The algorithm is work-efficient because it performs the same amount of work as the best serial algorithm for this kind of problem. Our algorithm has been implemented in C++ using the Thrust parallel algorithms library, enabling the generation of optimized programs for multicore CPUs and GPUs from the same source code. The performance of our algorithm is evaluated on synthetic and real datasets, showing good scalability on different generations of hardware. Moreno Marzolla, Giovanni Birolo, Gabriele D'Angelo, Piero Fariselli |
Future Gener. Comput. Syst. | 4 |
| 2023 | iDPP@CLEF 2023: The Intelligent Disease Progression Prediction Challenge
Helena Aidos, Roberto Bergamaschi, Paola Cavalla, Adriano Chiò, Arianna Dagliati, Barbara Di Camillo, Mamede de Carvalho, Nicola Ferro 0001, Piero Fariselli, Jose Manuel García Dominguez, Sara C. Madeira, Eleonora Tavazzi |
ECIR (3) | 9 |
| 2023 | Artificial intelligence and statistical methods for stratification and prediction of progression in amyotrophic lateral sclerosis: A systematic reviewabstractBACKGROUND: Amyotrophic Lateral Sclerosis (ALS) is a fatal neurodegenerative disorder characterised by the progressive loss of motor neurons in the brain and spinal cord. The fact that ALS's disease course is highly heterogeneous, and its determinants not fully known, combined with ALS's relatively low prevalence, renders the successful application of artificial intelligence (AI) techniques particularly arduous. OBJECTIVE: This systematic review aims at identifying areas of agreement and unanswered questions regarding two notable applications of AI in ALS, namely the automatic, data-driven stratification of patients according to their phenotype, and the prediction of ALS progression. Differently from previous works, this review is focused on the methodological landscape of AI in ALS. METHODS: We conducted a systematic search of the Scopus and PubMed databases, looking for studies on data-driven stratification methods based on unsupervised techniques resulting in (A) automatic group discovery or (B) a transformation of the feature space allowing patient subgroups to be identified; and for studies on internally or externally validated methods for the prediction of ALS progression. We described the selected studies according to the following characteristics, when applicable: variables used, methodology, splitting criteria and number of groups, prediction outcomes, validation schemes, and metrics. RESULTS: Of the starting 1604 unique reports (2837 combined hits between Scopus and PubMed), 239 were selected for thorough screening, leading to the inclusion of 15 studies on patient stratification, 28 on prediction of ALS progression, and 6 on both stratification and prediction. In terms of variables used, most stratification and prediction studies included demographics and features derived from the ALSFRS or ALSFRS-R scores, which were also the main prediction targets. The most represented stratification methods were K-means, and hierarchical and expectation-maximisation clustering; while random forests, logistic regression, the Cox proportional hazard model, and various flavours of deep learning were the most widely used prediction methods. Predictive model validation was, albeit unexpectedly, quite rarely performed in absolute terms (leading to the exclusion of 78 eligible studies), with the overwhelming majority of included studies resorting to internal validation only. CONCLUSION: This systematic review highlighted a general agreement in terms of input variable selection for both stratification and prediction of ALS progression, and in terms of prediction targets. A striking lack of validated models emerged, as well as a general difficulty in reproducing many published studies, mainly due to the absence of the corresponding parameter lists. While deep learning seems promising for prediction applications, its superiority with respect to traditional methods has not been established; there is, instead, ample room for its application in the subfield of patient stratification. Finally, an open question remains on the role of new environmental and behavioural variables collected via novel, real-time sensors. Erica Tavazzi, Enrico Longato, Martina Vettoretti, Helena Aidos, Isotta Trescato, Chiara Roversi, Andreia S. Martins, Eduardo N. Castanho, Ruben Branco, Diogo F. Soares, Alessandro Guazzo, Giovanni Birolo, Daniele Pala, Pietro Bosoni, Adriano Chiò, Umberto Manera, Mamede de Carvalho, Bruno Miranda, Marta Gromicho, Inês Alves, Riccardo Bellazzi, Arianna Dagliati, Piero Fariselli, Sara C. Madeira, Barbara Di Camillo |
Artif. Intell. Medicine | 23 |
| 2023 | Nonlinear data fusion over Entity-Relation graphs for Drug-Target Interaction predictionabstractMOTIVATION: The prediction of reliable Drug-Target Interactions (DTIs) is a key task in computer-aided drug design and repurposing. Here, we present a new approach based on data fusion for DTI prediction built on top of the NXTfusion library, which generalizes the Matrix Factorization paradigm by extending it to the nonlinear inference over Entity-Relation graphs. RESULTS: We benchmarked our approach on five datasets and we compared our models against state-of-the-art methods. Our models outperform most of the existing methods and, simultaneously, retain the flexibility to predict both DTIs as binary classification and regression of the real-valued drug-target affinity, competing with models built explicitly for each task. Moreover, our findings suggest that the validation of DTI methods should be stricter than what has been proposed in some previous studies, focusing more on mimicking real-life DTI settings where predictions for previously unseen drugs, proteins, and drug-protein pairs are needed. These settings are exactly the context in which the benefit of integrating heterogeneous information with our Entity-Relation data fusion approach is the most evident. AVAILABILITY AND IMPLEMENTATION: All software and data are available at https://github.com/eugeniomazzone/CPI-NXTFusion and https://pypi.org/project/NXTfusion/. Eugenio Mazzone, Yves Moreau, Piero Fariselli, Daniele Raimondi |
Bioinform. | 3 |
| 2023 | Recombulator-X: A fast and user-friendly tool for estimating X chromosome recombination rates in forensic geneticsabstractGenetic markers (especially short tandem repeats or STRs) located on the X chromosome are a valuable resource to solve complex kinship cases in forensic genetics in addition or alternatively to autosomal STRs. Groups of tightly linked markers are combined into haplotypes, thus increasing the discriminating power of tests. However, this approach requires precise knowledge of the recombination rates between adjacent markers. The International Society of Forensic Genetics recommends that recombination rate estimation on the X chromosome is performed from pedigree genetic data while taking into account the confounding effect of mutations. However, implementations that satisfy these requirements have several drawbacks: they were never publicly released, they are very slow and/or need cluster-level hardware and strong computational expertise to use. In order to address these key concerns we developed Recombulator-X, a new open-source Python tool. The most challenging issue, namely the running time, was addressed with dynamic programming techniques to greatly reduce the computational complexity of the algorithm. Compared to the previous methods, Recombulator-X reduces the estimation times from weeks or months to less than one hour for typical datasets. Moreover, the estimation process, including preprocessing, has been streamlined and packaged into a simple command-line tool that can be run on a normal PC. Where previous approaches were limited to small panels of STR markers (up to 15), our tool can handle greater numbers (up to 100) of mixed STR and non-STR markers. In conclusion, Recombulator-X makes the estimation process much simpler, faster and accessible to researchers without a computational background, hopefully spurring increased adoption of best practices. Serena Aneli, Piero Fariselli, Elena Chierto, Carla Bini, Carlo Robino, Giovanni Birolo |
PLoS Comput. Biol. | 2 |
| 2022 | Predicting protein stability changes upon single-point mutation: a thorough comparison of the available tools on a new datasetabstractPredicting the difference in thermodynamic stability between protein variants is crucial for protein design and understanding the genotype-phenotype relationships. So far, several computational tools have been created to address this task. Nevertheless, most of them have been trained or optimized on the same and 'all' available data, making a fair comparison unfeasible. Here, we introduce a novel dataset, collected and manually cleaned from the latest version of the ThermoMutDB database, consisting of 669 variants not included in the most widely used training datasets. The prediction performance and the ability to satisfy the antisymmetry property by considering both direct and reverse variants were evaluated across 21 different tools. The Pearson correlations of the tested tools were in the ranges of 0.21-0.5 and 0-0.45 for the direct and reverse variants, respectively. When both direct and reverse variants are considered, the antisymmetric methods perform better achieving a Pearson correlation in the range of 0.51-0.62. The tested methods seem relatively insensitive to the physiological conditions, performing well also on the variants measured with more extreme pH and temperature values. A common issue with all the tested methods is the compression of the $\Delta \Delta G$ predictions toward zero. Furthermore, the thermodynamic stability of the most significantly stabilizing variants was found to be more challenging to predict. This study is the most extensive comparisons of prediction methods using an entirely novel set of variants never tested before. Corrado Pancotti, Silvia Benevenuta, Giovanni Birolo, Virginia Alberini, Valeria Repetto, Tiziana Sanavia, Emidio Capriotti, Piero Fariselli |
Briefings Bioinform. | 8 |
| 2021 | DNA sequence symmetries from randomness: the origin of the Chargaff's second parity ruleabstractMost living organisms rely on double-stranded DNA (dsDNA) to store their genetic information and perpetuate themselves. This biological information has been considered as the main target of evolution. However, here we show that symmetries and patterns in the dsDNA sequence can emerge from the physical peculiarities of the dsDNA molecule itself and the maximum entropy principle alone, rather than from biological or environmental evolutionary pressure. The randomness justifies the human codon biases and context-dependent mutation patterns in human populations. Thus, the DNA 'exceptional symmetries,' emerged from the randomness, have to be taken into account when looking for the DNA encoded information. Our results suggest that the double helix energy constraints and, more generally, the physical properties of the dsDNA are the hard drivers of the overall DNA sequence architecture, whereas the selective biological processes act as soft drivers, which only under extraordinary circumstances overtake the overall entropy content of the genome. Piero Fariselli, Cristian Taccioli, Luca Pagani 0002, Amos Maritan |
Briefings Bioinform. | 1 |
| 2021 | On the critical review of five machine learning-based algorithms for predicting protein stability changes upon mutationabstractA review, recently published in this journal by Fang (2019), showed that methods trained for the prediction of protein stability changes upon mutation have a very critical bias: they neglect that a protein variation (A- > B) and its reverse (B- > A) must have the opposite value of the free energy difference (ΔΔGAB = - ΔΔGBA). In this letter, we complement the Fang's paper presenting a more general view of the problem. In particular, a machine learning-based method, published in 2015 (INPS), addressed the bias issue directly. We include the analysis of the missing method, showing that INPS is nearly insensitive to the addressed problem. Castrense Savojardo, Pier Luigi Martelli, Rita Casadio, Piero Fariselli |
Briefings Bioinform. | 4 |
| 2021 | Calibrating variant-scoring methods for clinical decision makingabstractSUMMARY: Identifying pathogenic variants and annotating them is a major challenge in human genetics, especially for the non-coding ones. Several tools have been developed and used to predict the functional effect of genetic variants. However, the calibration assessment of the predictions has received little attention. Calibration refers to the idea that if a model predicts a group of variants to be pathogenic with a probability P, it is expected that the same fraction P of true positive is found in the observed set. For instance, a well-calibrated classifier should label the variants such that among the ones to which it gave a probability value close to 0.7, approximately 70% actually belong to the pathogenic class. Poorly calibrated algorithms can be misleading and potentially harmful for clinical decision making. AVALIABILITY AND IMPLEMENTATION: The dataset used for testing the methods is available through the DOI:10.5281/zenodo.4448197. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Silvia Benevenuta, Emidio Capriotti, Piero Fariselli |
Bioinform. | 3 |
| 2020 | Insight into the protein solubility driving forces with neural attentionabstractProtein solubility is a key aspect for many biotechnological, biomedical and industrial processes, such as the production of active proteins and antibodies. In addition, understanding the molecular determinants of the solubility of proteins may be crucial to shed light on the molecular mechanisms of diseases caused by aggregation processes such as amyloidosis. Here we present SKADE, a novel Neural Network protein solubility predictor and we show how it can provide novel insight into the protein solubility mechanisms, thanks to its neural attention architecture. First, we show that SKADE positively compares with state of the art tools while using just the protein sequence as input. Then, thanks to the neural attention mechanism, we use SKADE to investigate the patterns learned during training and we analyse its decision process. We use this peculiarity to show that, while the attention profiles do not correlate with obvious sequence aspects such as biophysical properties of the aminoacids, they suggest that N- and C-termini are the most relevant regions for solubility prediction and are predictive for complex emergent properties such as aggregation-prone regions involved in beta-amyloidosis and contact density. Moreover, SKADE is able to identify mutations that increase or decrease the overall solubility of the protein, allowing it to be used to perform large scale in-silico mutagenesis of proteins in order to maximize their solubility. Daniele Raimondi, Gabriele Orlando, Piero Fariselli, Yves Moreau |
PLoS Comput. Biol. | 3 |
| 2019 | Improving the prediction of cardiovascular risk with machine-learning and DNA methylation dataabstractClassically, the cardiovascular risk of individual is evaluated using phenomenological variables (PV)such as blood pressure, body mass, smoker status, gender, age etc. Here we show that, on prospective study (after 10-15 years)these PV display a poor agreement with case-control samples. We were able to obtain more accurate predictions using both DNA methylation data and PV as input features of a Random Forest model, achieving a ROC-AUC of 0.74. Furthermore, the Random Forest output correlates with the reliability of the predictions producing a ROC-AUC of 0.90 when only the most reliable predictions are taken into consideration. Giovanni Cugliari, Silvia Benevenuta, Simonetta Guarrera, Carlotta Sacerdote, Salvatore Panico, Vittorio Krogh, Rosario Tumino, Paolo Vineis, Piero Fariselli, Giuseppe Matullo |
CIBCB | 9 |
| 2019 | A natural upper bound to the accuracy of predicting protein stability changes upon mutationsabstractMOTIVATION: Accurate prediction of protein stability changes upon single-site variations (ΔΔG) is important for protein design, as well as for our understanding of the mechanisms of genetic diseases. The performance of high-throughput computational methods to this end is evaluated mostly based on the Pearson correlation coefficient between predicted and observed data, assuming that the upper bound would be 1 (perfect correlation). However, the performance of these predictors can be limited by the distribution and noise of the experimental data. Here we estimate, for the first time, a theoretical upper-bound to the ΔΔG prediction performances imposed by the intrinsic structure of currently available ΔΔG data. RESULTS: Given a set of measured ΔΔG protein variations, the theoretically "best predictor" is estimated based on its similarity to another set of experimentally determined ΔΔG values. We investigate the correlation between pairs of measured ΔΔG variations, where one is used as a predictor for the other. We analytically derive an upper bound to the Pearson correlation as a function of the noise and distribution of the ΔΔG data. We also evaluate the available datasets to highlight the effect of the noise in conjunction with ΔΔG distribution. We conclude that the upper bound is a function of both uncertainty and spread of the ΔΔG values, and that with current data the best performance should be between 0.7 and 0.8, depending on the dataset used; higher Pearson correlations might be indicative of overtraining. It also follows that comparisons of predictors using different datasets are inherently misleading. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Ludovica Montanucci, Pier Luigi Martelli, Nir Ben-Tal, Piero Fariselli |
Bioinform. | 4 |
| 2019 | On the biases in predictions of protein stability changes upon variations: the INPS test caseabstractINPS (Fariselli et al., 2015) is a simple method to predict the protein stability changes upon single point variation. It is based on Support Vector Machine and it was trained using 7 input features that are: (i) Blosum62 substitution, (ii) residue mutability, (iii) molecular weight wild type, (iv) molecular weight variant, (v) hydrophobicity wild type, (vi) hydrophobicity variant, (vii) difference between Viterbi scores of the wild type and variant (Fariselli et al., 2015). The first two input features are not anti-symmetric, the last is anti-symmetric by construction, while the other features can be learned to be anti-symmetric. To enforce anti-symmetry INPS was trained using both direct and inverse variations (Fariselli et al., 2015). When the alignment is poor we can expect a deviation from the anti-symmetry. INPS3D (Savojardo et al., 2016), adds to INPS information about contact potential and solvent accessibility. The solvent accessibility is intrinsically non anti-symmetric, and for this reason INPS3D was trained only with direct variations (Savojardo et al., 2016). The Ssym dataset is a manually curated selection of variations from the ProTherm database (Pucci et al., 2018). It contains variations with experimental ΔΔG values for which the 3D structures of both the wild-type and variant proteins were solved by X-ray crystallography. Ssym consists of 684 variations, half of which are direct (reported in the literature) and half are obtained by anti-symmetry (Pucci et al., 2018). To overcome the bias problem, a new method, PopMuSiCSYM, was explicitly developed to furnish completely anti-symmetric predictions of the variations (Pucci et al., 2015, 2018). The authors of the Ssym dataset estimated the performance of fifteen selected predictors using the average bias parameter (<δ> = ∑(ΔΔGdir+ΔΔGinv)/N) and the linear correlation coefficient (rdir-inv) between the predicted ΔΔG values of the direct and the corresponding inverse variations. A perfect predictor should have bias close to zero (<δ> = 0) and a correlation coefficient close to -1 (rdir-inv = −1). In the same paper, they also tested the performances of these fifteen predictors using the experimental ΔΔG values on both direct and inverse variations using the root mean square deviation (σdir, σinv) and the linear correlation coefficient between the predicted and experimental values (rdir, rinv). In Table 1 we report the predictor performances and biases taken from the original paper (Pucci et al., 2018) and we add those computed using INPS and INPS3D. From Table 1, it is clear that the partial anti-symmetry of the INPS input and its training performed on both direct and inverse variations paid off. INPS shows a very low bias (<δ>, just second best), the highest correlation between direct and inverse variations (rdir-inv) and the highest performance on the inverse variation sets (σinv and rinv). The graph of the correlation between direct and inverse variation predicted by INPS is presented in Figure 1. INPS and INPS3D prediction on Ssym. The INPS (black) and INPS3D (grey) predictions for the 342 pairs (direct and inverse) of single-point variations in Ssym (Pucci et al., 2018) are shown. ΔΔG for the direct (x-axis) and inverse (y-axis) variations are reported as dots in the graph. The ‘ideal’ relationship ΔΔGAB + ΔΔGBA = 0 is shown as a solid line Bias analysis on the Ssym dataset (Pucci et al. 2018) Notes: Standard deviations σ and the average bias <δ> are in kcal/mol. All the values, except those of INPS and INPS3D, are taken from (Pucci et al., 2018). Bold character highlights INPS method. Bias analysis on the Ssym dataset (Pucci et al. 2018) Notes: Standard deviations σ and the average bias <δ> are in kcal/mol. All the values, except those of INPS and INPS3D, are taken from (Pucci et al., 2018). Bold character highlights INPS method. The INPS training set includes some of the proteins in the Ssym dataset (which is also true for all the other tested methods reported in Table 1). For sake of comparison, we also computed the performances of a ‘blind’ INPS. In this case, each ΔΔG prediction is obtained by using a svm model that does not include the protein to predict or a similar one (sequence identity <25%) in the training set. The standard deviations and Pearson correlation coefficients computed for the blind-INPS are: 1.44 kcal/mol and 0.48 for direct variations and 1.45 kcal/mol and 0.47 for inverse variations. The correlation between the predictions of direct and inverse variations is -0.99 while the average bias (measured through the δ) is -0.06 kcal/mol. The structure-based INPS3D has the second best correlation between direct and inverse correlation but a considerably higher bias. This is expected given that INPS3D has not been trained on inverse variations and conains as input feature the solvent accessibility, which is not anti-symmetric. Finally, it must be stressed that the correlations reported in Table 1 are only intended to assess the anti-symmetricity of the methods, and not as estimators of their overall performances, which is a very difficult task that also requires an evaluation of stability of the dataset (Yang et al., 2018). A second dataset was built by Usmanova et al. (2018), by extracting high-resolution pairs of proteins from the Protein Data Bank (PDB) differing by one to ten amino acids. Here we test INPS on the subset of all the single-site variations extracted from the Usmanova et al. (2018) dataset, that is, all the pdb pairs whose sequences differ by exactly one residue. This subset comprises 1000 pairs of single-point protein variations. Although for this set the experimental ΔΔG values are not known, given its important size, it is very informative for assessing predictor biases. In Usmanova et al. (2018), bias for an individual variation is estimated as: =(ΔΔGAB + ΔΔGBA)/2. In Table 2 we add the INPS and INPS3D anti-symmetry indexes to those previously computed on other predictors (Usmanova et al., 2018). The bias of INPS is the smallest and the correlation between direct and inverse variations is the highest. INPS3D, while showing less anti-symmetricity than INPS, has the second best bias and the second best correlation. Bias analysis on the Usmanova et al. (2018) dataset Notes: First column: mean and standard error of the mean for the biases. Second column: r, Pearson correlation coefficient with the associated P-value, between direct and inverse variations. All the values except INPS and INPS3D, are taken from Usmanova et al. (2018). 0.0*: P-value set to 0, since it is undetectable by the machine precision. Bold character highlights INPS method. Bias analysis on the Usmanova et al. (2018) dataset Notes: First column: mean and standard error of the mean for the biases. Second column: r, Pearson correlation coefficient with the associated P-value, between direct and inverse variations. All the values except INPS and INPS3D, are taken from Usmanova et al. (2018). 0.0*: P-value set to 0, since it is undetectable by the machine precision. Bold character highlights INPS method. The very high correlation of -0.95 between direct and inverse variations obtained by INPS on this set is shown in the graph of Figure 2, to be contrasted with those reported for the other methods in the figures of the paper by Usmanova et al. (2018). In Figure 2, the points very far from the diagonal correspond to variations whose protein sequences are poorly aligned (number of aligned sequences <20). INPS and INPS3D prediction for the 1000 pairs (direct and inverse) of single-point variations (Usmanova et al., 2018). ΔΔG for the direct (x-axis) and inverse (y-axis) variations are reported as dots in the graph. The ‘ideal’ relationship ΔΔGAB + ΔΔGBA = 0 is shown as a solid line Recently, two groups (Pucci et al., 2018; Usmanova et al., 2018) compiled two datasets to test a very important bias that affects most of the computational methods designed to predict free energy changes upon protein variation. This bias is the lack of anti-symmetry in the ΔΔG predictions between direct and inverse variations. These studies verified this bias on several relevant methods. Here we add the test on INPS (Fariselli et al., 2015) that was specifically designed to take into account anti-symmetry problem using an input that was partially anti-symmetric (the HMM Viterbi score in particular). Here we show that INPS performs very well on both datasets proposed in the two papers. The fact that INPS uses only sequence information makes it less sensitive to structural rearrangement upon residue substitution and in turn more robust and suitable for testing any kind of protein variations. This work has been supported by EBA-PRISM an Israel-Italy collaborative project, the Israel Ministry of Science and Technology and Italian Ministry of Foreign Affair and International Cooperation. This research received funding specifically appointed to Department of Medical Sciences from the Italian Ministry for Education, University and Research (MIUR) under the programme “Dipartimenti di Eccellenza 2018 – 2022” Project code D15D18000410001. Conflict of Interest: none declared. Ludovica Montanucci, Castrense Savojardo, Pier Luigi Martelli, Rita Casadio, Piero Fariselli |
Bioinform. | 5 |
| 2019 | DDGun: an untrained method for the prediction of protein stability changes upon single and multiple point variationsabstractBACKGROUND: Predicting the effect of single point variations on protein stability constitutes a crucial step toward understanding the relationship between protein structure and function. To this end, several methods have been developed to predict changes in the Gibbs free energy of unfolding (∆∆G) between wild type and variant proteins, using sequence and structure information. Most of the available methods however do not exhibit the anti-symmetric prediction property, which guarantees that the predicted ∆∆G value for a variation is the exact opposite of that predicted for the reverse variation, i.e., ∆∆G(A → B) = -∆∆G(B → A), where A and B are amino acids. RESULTS: Here we introduce simple anti-symmetric features, based on evolutionary information, which are combined to define an untrained method, DDGun (DDG untrained). DDGun is a simple approach based on evolutionary information that predicts the ∆∆G for single and multiple variations from sequence and structure information (DDGun3D). Our method achieves remarkable performance without any training on the experimental datasets, reaching Pearson correlation coefficients between predicted and measured ∆∆G values of ~ 0.5 and ~ 0.4 for single and multiple site variations, respectively. Surprisingly, DDGun performances are comparable with those of state of the art methods. DDGun also naturally predicts multiple site variations, thereby defining a benchmark method for both single site and multiple site predictors. DDGun is anti-symmetric by construction predicting the value of the ∆∆G of a reciprocal variation as almost equal (depending on the sequence profile) to -∆∆G of the direct variation. This is a valuable property that is missing in the majority of the methods. CONCLUSIONS: Evolutionary information alone combined in an untrained method can achieve remarkably high performances in the prediction of ∆∆G upon protein mutation. Non-trained approaches like DDGun represent a valid benchmark both for scoring the predictive power of the individual features and for assessing the learning capability of supervised methods. Ludovica Montanucci, Emidio Capriotti, Yotam Frank, Nir Ben-Tal, Piero Fariselli |
BMC Bioinform. | 5 |
| 2018 | DeepSig: deep learning improves signal peptide detection in proteinsabstractMotivation: The identification of signal peptides in protein sequences is an important step toward protein localization and function characterization. Results: Here, we present DeepSig, an improved approach for signal peptide detection and cleavage-site prediction based on deep learning methods. Comparative benchmarks performed on an updated independent dataset of proteins show that DeepSig is the current best performing method, scoring better than other available state-of-the-art approaches on both signal peptide detection and precise cleavage-site identification. Availability and implementation: DeepSig is available as both standalone program and web server at https://deepsig.biocomp.unibo.it. All datasets used in this study can be obtained from the same website. Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Castrense Savojardo, Pier Luigi Martelli, Piero Fariselli, Rita Casadio |
Bioinform. | 3 |
| 2017 | ISPRED4: interaction sites PREDiction in protein structures with a refining grammar modelabstractMOTIVATION: The identification of protein-protein interaction (PPI) sites is an important step towards the characterization of protein functional integration in the cell complexity. Experimental methods are costly and time-consuming and computational tools for predicting PPI sites can fill the gaps of PPI present knowledge. RESULTS: We present ISPRED4, an improved structure-based predictor of PPI sites on unbound monomer surfaces. ISPRED4 relies on machine-learning methods and it incorporates features extracted from protein sequence and structure. Cross-validation experiments are carried out on a new dataset that includes 151 high-resolution protein complexes and indicate that ISPRED4 achieves a per-residue Matthew Correlation Coefficient of 0.48 and an overall accuracy of 0.85. Benchmarking results show that ISPRED4 is one of the top-performing PPI site predictors developed so far. CONTACT: [email protected]. AVAILABILITY AND IMPLEMENTATION: ISPRED4 and datasets used in this study are available at http://ispred4.biocomp.unibo.it . Castrense Savojardo, Piero Fariselli, Pier Luigi Martelli, Rita Casadio |
Bioinform. | 2 |
| 2017 | SChloro: directing Viridiplantae proteins to six chloroplastic sub-compartmentsabstractMotivation: Chloroplasts are organelles found in plants and involved in several important cell processes. Similarly to other compartments in the cell, chloroplasts have an internal structure comprising several sub-compartments, where different proteins are targeted to perform their functions. Given the relation between protein function and localization, the availability of effective computational tools to predict protein sub-organelle localizations is crucial for large-scale functional studies. Results: In this paper we present SChloro, a novel machine-learning approach to predict protein sub-chloroplastic localization, based on targeting signal detection and membrane protein information. The proposed approach performs multi-label predictions discriminating six chloroplastic sub-compartments that include inner membrane, outer membrane, stroma, thylakoid lumen, plastoglobule and thylakoid membrane. In comparative benchmarks, the proposed method outperforms current state-of-the-art methods in both single- and multi-compartment predictions, with an overall multi-label accuracy of 74%. The results demonstrate the relevance of the approach that is eligible as a good candidate for integration into more general large-scale annotation pipelines of protein subcellular localization. Availability and Implementation: The method is available as web server at http://schloro.biocomp.unibo.it Contact: [email protected]. Castrense Savojardo, Pier Luigi Martelli, Piero Fariselli, Rita Casadio |
Bioinform. | 3 |
| 2016 | NET-GE: a web-server for NETwork-based human gene enrichmentabstractMOTIVATION: Gene enrichment is a requisite for the interpretation of biological complexity related to specific molecular pathways and biological processes. Furthermore, when interpreting NGS data and human variations, including those related to pathologies, gene enrichment allows the inclusion of other genes that in the human interactome space may also play important key roles in the emergency of the phenotype. Here, we describe NET-GE, a web server for associating biological processes and pathways to sets of human proteins involved in the same phenotype RESULTS: NET-GE is based on protein-protein interaction networks, following the notion that for a set of proteins, the context of their specific interactions can better define their function and the processes they can be related to in the biological complexity of the cell. Our method is suited to extract statistically validated enriched terms from Gene Ontology, KEGG and REACTOME annotation databases. Furthermore, NET-GE is effective even when the number of input proteins is small. AVAILABILITY AND IMPLEMENTATION: NET-GE web server is publicly available and accessible at http://net-ge.biocomp.unibo.it/enrich CONTACT: [email protected] information: Supplementary data are available at Bioinformatics online. Samuele Bovo, Pietro Di Lena, Pier Luigi Martelli, Piero Fariselli, Rita Casadio |
Bioinform. | 4 |
| 2016 | INPS-MD: a web server to predict stability of protein variants from sequence and structureabstractMOTIVATION: Protein function depends on its structural stability. The effects of single point variations on protein stability can elucidate the molecular mechanisms of human diseases and help in developing new drugs. Recently, we introduced INPS, a method suited to predict the effect of variations on protein stability from protein sequence and whose performance is competitive with the available state-of-the-art tools. RESULTS: In this article, we describe INPS-MD (Impact of Non synonymous variations on Protein Stability-Multi-Dimension), a web server for the prediction of protein stability changes upon single point variation from protein sequence and/or structure. Here, we complement INPS with a new predictor (INPS3D) that exploits features derived from protein 3D structure. INPS3D scores with Pearson's correlation to experimental ΔΔG values of 0.58 in cross validation and of 0.72 on a blind test set. The sequence-based INPS scores slightly lower than the structure-based INPS3D and both on the same blind test sets well compare with the state-of-the-art methods. AVAILABILITY AND IMPLEMENTATION: INPS and INPS3D are available at the same web server: http://inpsmd.biocomp.unibo.it SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. CONTACT: [email protected]. Castrense Savojardo, Piero Fariselli, Pier Luigi Martelli, Rita Casadio |
Bioinform. | 2 |
| 2015 | INPS: predicting the impact of non-synonymous variations on protein stability from sequenceabstractMOTIVATION: A tool for reliably predicting the impact of variations on protein stability is extremely important for both protein engineering and for understanding the effects of Mendelian and somatic mutations in the genome. Next Generation Sequencing studies are constantly increasing the number of protein sequences. Given the huge disproportion between protein sequences and structures, there is a need for tools suited to annotate the effect of mutations starting from protein sequence without relying on the structure. Here, we describe INPS, a novel approach for annotating the effect of non-synonymous mutations on the protein stability from its sequence. INPS is based on SVM regression and it is trained to predict the thermodynamic free energy change upon single-point variations in protein sequences. RESULTS: We show that INPS performs similarly to the state-of-the-art methods based on protein structure when tested in cross-validation on a non-redundant dataset. INPS performs very well also on a newly generated dataset consisting of a number of variations occurring in the tumor suppressor protein p53. Our results suggest that INPS is a tool suited for computing the effect of non-synonymous polymorphisms on protein stability when the protein structure is not available. We also show that INPS predictions are complementary to those of the state-of-the-art, structure-based method mCSM. When the two methods are combined, the overall prediction on the p53 set scores significantly higher than those of the single methods. AVAILABILITY AND IMPLEMENTATION: The presented method is available as web server at http://inps.biocomp.unibo.it. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary Materials are available at Bioinformatics online. Piero Fariselli, Pier Luigi Martelli, Castrense Savojardo, Rita Casadio |
Bioinform. | 1 |
| 2015 | AlignBucket: a tool to speed up 'all-against-all' protein sequence alignments optimizing length constraintsabstractMOTIVATION: The next-generation sequencing era requires reliable, fast and efficient approaches for the accurate annotation of the ever-increasing number of biological sequences and their variations. Transfer of annotation upon similarity search is a standard approach. The procedure of all-against-all protein comparison is a preliminary step of different available methods that annotate sequences based on information already present in databases. Given the actual volume of sequences, methods are necessary to pre-process data to reduce the time of sequence comparison. RESULTS: We present an algorithm that optimizes the partition of a large volume of sequences (the whole database) into sets where sequence length values (in residues) are constrained depending on a bounded minimal and expected alignment coverage. The idea is to optimally group protein sequences according to their length, and then computing the all-against-all sequence alignments among sequences that fall in a selected length range. We describe a mathematically optimal solution and we show that our method leads to a 5-fold speed-up in real world cases. AVAILABILITY AND IMPLEMENTATION: The software is available for downloading at http://www.biocomp.unibo.it/∼giuseppe/partitioning.html. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Giuseppe Profiti, Piero Fariselli, Rita Casadio |
Bioinform. | 2 |
| 2015 | TPpred3 detects and discriminates mitochondrial and chloroplastic targeting peptides in eukaryotic proteinsabstractMOTIVATION: Molecular recognition of N-terminal targeting peptides is the most common mechanism controlling the import of nuclear-encoded proteins into mitochondria and chloroplasts. When experimental information is lacking, computational methods can annotate targeting peptides, and determine their cleavage sites for characterizing protein localization, function, and mature protein sequences. The problem of discriminating mitochondrial from chloroplastic propeptides is particularly relevant when annotating proteomes of photosynthetic Eukaryotes, endowed with both types of sequences. RESULTS: Here, we introduce TPpred3, a computational method that given any Eukaryotic protein sequence performs three different tasks: (i) the detection of targeting peptides; (ii) their classification as mitochondrial or chloroplastic and (iii) the precise localization of the cleavage sites in an organelle-specific framework. Our implementation is based on our TPpred previously introduced. Here, we integrate a new N-to-1 Extreme Learning Machine specifically designed for the classification task (ii). For the last task, we introduce an organelle-specific Support Vector Machine that exploits sequence motifs retrieved with an extensive motif-discovery analysis of a large set of mitochondrial and chloroplastic proteins. We show that TPpred3 outperforms the state-of-the-art methods in all the three tasks. AVAILABILITY AND IMPLEMENTATION: The method server and datasets are available at http://tppred3.biocomp.unibo.it. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Castrense Savojardo, Pier Luigi Martelli, Piero Fariselli, Rita Casadio |
Bioinform. | 3 |
| 2014 | TPpred2: improving the prediction of mitochondrial targeting peptide cleavage sites by exploiting sequence motifsabstractSUMMARY: Targeting peptides are N-terminal sorting signals in proteins that promote their translocation to mitochondria through the interaction with different protein machineries. We recently developed TPpred, a machine learning-based method scoring among the best ones available to predict the presence of a targeting peptide into a protein sequence and its cleavage site. Here we introduce TPpred2 that improves TPpred performances in the task of identifying the cleavage site of the targeting peptides. TPpred2 is now available as a web interface and as a stand-alone version for users who can freely download and adopt it for processing large volumes of sequences. Availability and implementaion: TPpred2 is available both as web server and stand-alone version at http://tppred2.biocomp.unibo.it. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Castrense Savojardo, Pier Luigi Martelli, Piero Fariselli, Rita Casadio |
Bioinform. | 3 |
| 2013 | The prediction of organelle-targeting peptides in eukaryotic proteins with Grammatical-Restrained Hidden Conditional Random FieldsabstractMOTIVATION: Targeting peptides are the most important signal controlling the import of nuclear encoded proteins into mitochondria and plastids. In the lack of experimental information, their prediction is an essential step when proteomes are annotated for inferring both the localization and the sequence of mature proteins. RESULTS: We developed TPpred a new predictor of organelle-targeting peptides based on Grammatical-Restrained Hidden Conditional Random Fields. TPpred is trained on a non-redundant dataset of proteins where the presence of a target peptide was experimentally validated, comprising 297 sequences. When tested on the 297 positive and some other 8010 negative examples, TPpred outperformed available methods in both accuracy and Matthews correlation index (96% and 0.58, respectively). Given its very low-false-positive rate (3.0%), TPpred is, therefore, well suited for large-scale analyses at the proteome level. We predicted that from ∼4 to 9% of the sequences of human, Arabidopsis thaliana and yeast proteomes contain targeting peptides and are, therefore, likely to be localized in mitochondria and plastids. TPpred predictions correlate to a good extent with the experimental annotation of the subcellular localization, when available. TPpred was also trained and tested to predict the cleavage site of the organelle-targeting peptide: on this task, the average error of TPpred on mitochondrial and plastidic proteins is 7 and 15 residues, respectively. This value is lower than the error reported by other methods currently available. AVAILABILITY: The TPpred datasets are available at http://biocomp.unibo.it/valentina/TPpred/. TPpred is available on request from the authors. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Valentina Indio, Pier Luigi Martelli, Castrense Savojardo, Piero Fariselli, Rita Casadio |
Bioinform. | 4 |
| 2013 | BETAWARE: a machine-learning tool to detect and predict transmembrane beta-barrel proteins in prokaryotesabstractSUMMARY: The annotation of membrane proteins in proteomes is an important problem of Computational Biology, especially after the development of high-throughput techniques that allow fast and efficient genome sequencing. Among membrane proteins, transmembrane β-barrels (TMBBs) are poorly represented in the database of protein structures (PDB) and difficult to identify with experimental approaches. They are, however, extremely important, playing key roles in several cell functions and bacterial pathogenicity. TMBBs are included in the lipid bilayer with a β-barrel structure and are presently found in the outer membranes of Gram-negative bacteria, mitochondria and chloroplasts. Recently, we developed two top-performing methods based on machine-learning approaches to tackle both the detection of TMBBs in sets of proteins and the prediction of their topology. Here, we present our BETAWARE program that includes both approaches and can run as a standalone program on a linux-based computer to easily address in-home massive protein annotation or filtering. AVAILABILITY AND IMPLEMENTATION: http://www.biocomp.unibo.it/∼savojard/betawarecl . Castrense Savojardo, Piero Fariselli, Rita Casadio |
Bioinform. | 2 |
| 2013 | BCov: a method for predicting β-sheet topology using sparse inverse covariance estimation and integer programmingabstractMOTIVATION: Prediction of protein residue contacts, even at the coarse-grain level, can help in finding solutions to the protein structure prediction problem. Unlike α-helices that are locally stabilized, β-sheets result from pairwise hydrogen bonding of two or more disjoint regions of the protein backbone. The problem of predicting contacts among β-strands in proteins has been addressed by several supervised computational approaches. Recently, prediction of residue contacts based on correlated mutations has been greatly improved and finally allows the prediction of 3D structures of the proteins. RESULTS: In this article, we describe BCov, which is the first unsupervised method to predict the β-sheet topology starting from the protein sequence and its secondary structure. BCov takes advantage of the sparse inverse covariance estimation to define β-strand partner scores. Then an optimization based on integer programming is carried out to predict the β-sheet connectivity. When tested on the prediction of β-strand pairing, BCov scores with average values of Matthews Correlation Coefficient (MCC) and F1 equal to 0.56 and 0.61, respectively, on a non-redundant dataset of 916 protein chains known with atomic resolution. Our approach well compares with the state-of-the-art methods trained so far for this specific task. AVAILABILITY AND IMPLEMENTATION: The method is freely available under General Public License at http://biocomp.unibo.it/savojard/bcov/bcov-1.0.tar.gz. The new dataset BetaSheet1452 can be downloaded at http://biocomp.unibo.it/savojard/bcov/BetaSheet1452.dat. Castrense Savojardo, Piero Fariselli, Pier Luigi Martelli, Rita Casadio |
Bioinform. | 2 |
| 2013 | How to inherit statistically validated annotation within BAR+ protein clustersabstractBACKGROUND: In the genomic era a key issue is protein annotation, namely how to endow protein sequences, upon translation from the corresponding genes, with structural and functional features. Routinely this operation is electronically done by deriving and integrating information from previous knowledge. The reference database for protein sequences is UniProtKB divided into two sections, UniProtKB/TrEMBL which is automatically annotated and not reviewed and UniProtKB/Swiss-Prot which is manually annotated and reviewed. The annotation process is essentially based on sequence similarity search. The question therefore arises as to which extent annotation based on transfer by inheritance is valuable and specifically if it is possible to statistically validate inherited features when little homology exists among the target sequence and its template(s). RESULTS: In this paper we address the problem of annotating protein sequences in a statistically validated manner considering as a reference annotation resource UniProtKB. The test case is the set of 48,298 proteins recently released by the Critical Assessment of Function Annotations (CAFA) organization. We show that we can transfer after validation, Gene Ontology (GO) terms of the three main categories and Pfam domains to about 68% and 72% of the sequences, respectively. This is possible after alignment of the CAFA sequences towards BAR+, our annotation resource that allows discriminating among statistically validated and not statistically validated annotation. By comparing with a direct UniProtKB annotation, we find that besides validating annotation of some 78% of the CAFA set, we assign new and statistically validated annotation to 14.8% of the sequences and find new structural templates for about 25% of the chains, half of which share less than 30% sequence identity to the corresponding template/s. CONCLUSION: Inheritance of annotation by transfer generally requires a careful selection of the identity value among the target and the template in order to transfer structural and/or functional features. Here we prove that even distantly remote homologs can be safely endowed with structural templates and GO and/or Pfam terms provided that annotation is done within clusters collecting cluster-related protein sequences and where a statistical validation of the shared structural and functional features is possible. Damiano Piovesan, Pier Luigi Martelli, Piero Fariselli, Giuseppe Profiti, Andrea Zauli, Ivan Rossi, Rita Casadio |
BMC Bioinform. | 3 |
| 2013 | Prediction of disulfide connectivity in proteins with machine-learning methods and correlated mutationsabstractBACKGROUND: Recently, information derived by correlated mutations in proteins has regained relevance for predicting protein contacts. This is due to new forms of mutual information analysis that have been proven to be more suitable to highlight direct coupling between pairs of residues in protein structures and to the large number of protein chains that are currently available for statistical validation. It was previously discussed that disulfide bond topology in proteins is also constrained by correlated mutations. RESULTS: In this paper we exploit information derived from a corrected mutual information analysis and from the inverse of the covariance matrix to address the problem of the prediction of the topology of disulfide bonds in Eukaryotes. Recently, we have shown that Support Vector Regression (SVR) can improve the prediction for the disulfide connectivity patterns. Here we show that the inclusion of the correlated mutation information increases of 5 percentage points the SVR performance (from 54% to 59%). When this approach is used in combination with a method previously developed by us and scoring at the state of art in predicting both location and topology of disulfide bonds in Eukaryotes (DisLocate), the per-protein accuracy is 38%, 2 percentage points higher than that previously obtained. CONCLUSIONS: In this paper we show that the inclusion of information derived from correlated mutations can improve the performance of the state of the art methods for predicting disulfide connectivity patterns in Eukaryotic proteins. Our analysis also provides support to the notion that improving methods to extract evolutionary information from multiple sequence alignments greatly contributes to the scoring performance of predictors suited to detect relevant features from protein chains. Castrense Savojardo, Piero Fariselli, Pier Luigi Martelli, Rita Casadio |
BMC Bioinform. | 2 |
| 2013 | Extended and Robust Protein Sequence Annotation over Conservative Nonhierarchical Clusters: The Case Study of the ABC TransportersabstractGenome annotation is one of the most important issues in the genomic era. The exponential growth rate of newly sequenced genomes and proteomes urges the development of fast and reliable annotation methods, suited to exploit all the information available in curated databases of protein sequences and structures. To this aim we developed BAR+, the Bologna Annotation Resource. 1 The basic notion is that sequences with high identity value to a counterpart can inherit the same function/s and structure, if available. As a case study we describe how the ATP-binding domain of the ABC transporters can be found and modeled in over 30,000 new sequences not annotated before. We also mapped into BAR+ all the ABC transporters listed in the Transporter Classification DataBase 2 and found that within our environment annotation could be extended to another 256,866 sequences. Damiano Piovesan, Giuseppe Profiti, Pier Luigi Martelli, Piero Fariselli, Rita Casadio |
ACM J. Emerg. Technol. Comput. Syst. | 4 |
| 2011 | Improving the prediction of disulfide bonds in Eukaryotes with machine learning methods and protein subcellular localizationabstractMOTIVATION: Disulfide bonds stabilize protein structures and play relevant roles in their functions. Their formation requires an oxidizing environment and their stability is consequently depending on the redox ambient potential, which may differ according to the subcellular compartment. Several methods are available to predict cysteine-bonding state and connectivity patterns. However, none of them takes into consideration the relevance of protein subcellular localization. RESULTS: Here we develop DISLOCATE, a two-step method based on machine learning models for predicting both the bonding state and the connectivity patterns of cysteine residues in a protein chain. We find that the inclusion of protein subcellular localization improves the performance of these predictive steps by 3 and 2 percentage points, respectively. When compared with previously developed methods for predicting disulfide bonds from sequence, DISLOCATE improves the overall performance by more than 10 percentage points. AVAILABILITY: The method and the dataset are available at the Web page http://www.biocomp.unibo.it/savojard/Dislocate.html. GRHCRF code is available at http://www.biocomp.unibo.it/savojard/biocrf.html. CONTACT: [email protected]. Castrense Savojardo, Piero Fariselli, Monther Alhamdoosh, Pier Luigi Martelli, Andrea Pierleoni, Rita Casadio |
Bioinform. | 2 |
| 2011 | Improving the detection of transmembrane β-barrel chains with N-to-1 extreme learning machinesabstractMOTIVATION: Transmembrane β-barrels (TMBBs) are extremely important proteins that play key roles in several cell functions. They cross the lipid bilayer with β-barrel structures. TMBBs are presently found in the outer membranes of Gram-negative bacteria and of mitochondria and chloroplasts. Loop exposure outside the bacterial cell membranes makes TMBBs important targets for vaccine or drug therapies. In genomes, they are not highly represented and are difficult to identify with experimental approaches. Several computational methods have been developed to discriminate TMBBs from other types of proteins. However, the best performing approaches have a high fraction of false positive predictions. RESULTS: In this article, we introduce a new machine learning approach for TMBB detection based on N-to-1 Extreme Learning Machines that significantly outperforms previous methods achieving a Matthews correlation coefficient of 0.82, a probability of correct prediction of 0.92 and a sensitivity of 0.73. Castrense Savojardo, Piero Fariselli, Rita Casadio |
Bioinform. | 2 |
| 2011 | Is There an Optimal Substitution Matrix for Contact Prediction with Correlated Mutations?abstractCorrelated mutations in proteins are believed to occur in order to preserve the protein functional folding through evolution. Their values can be deduced from sequence and/or structural alignments and are indicative of residue contacts in the protein three-dimensional structure. A correlation among pairs of residues is routinely evaluated with the Pearson correlation coefficient and the MCLACHLAN similarity matrix. In literature, there is no justification for the adoption of the MCLACHLAN instead of other substitution matrices. In this paper, we approach the problem of computing the optimal similarity matrix for contact prediction with correlated mutations, i.e., the similarity matrix that maximizes the accuracy of contact prediction with correlated mutations. We describe an optimization procedure, based on the gradient descent method, for computing the optimal similarity matrix and perform an extensive number of experimental tests. Our tests show that there is a large number of optimal matrices that perform similarly to MCLACHLAN. We also obtain that the upper limit to the accuracy achievable in protein contact prediction is independent of the optimized similarity matrix. This suggests that the poor scoring of the correlated mutations approach may be due to the choice of the linear correlation function in evaluating correlated mutations. Pietro Di Lena, Piero Fariselli, Luciano Margara, Marco Vassura, Rita Casadio |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2010 | Fast overlapping of protein contact maps by alignment of eigenvectorsabstractMOTIVATION: Searching for structural similarity is a key issue of protein functional annotation. The maximum contact map overlap (CMO) is one of the possible measures of protein structure similarity. Exact and approximate methods known to optimize the CMO are computationally expensive and this hampers their applicability to large-scale comparison of protein structures. RESULTS: In this article, we describe a heuristic algorithm (Al-Eigen) for finding a solution to the CMO problem. Our approach relies on the approximation of contact maps by eigendecomposition. We obtain good overlaps of two contact maps by computing the optimal global alignment of few principal eigenvectors. Our algorithm is simple, fast and its running time is independent of the amount of contacts in the map. Experimental testing indicates that the algorithm is comparable to exact CMO methods in terms of the overlap quality, to structural alignment methods in terms of structure similarity detection and it is fast enough to be suited for large-scale comparison of protein structures. Furthermore, our preliminary tests indicates that it is quite robust to noise, which makes it suitable for structural similarity detection also for noisy and incomplete contact maps. AVAILABILITY: Available at http://bioinformatics.cs.unibo.it/Al-Eigen. Pietro Di Lena, Piero Fariselli, Luciano Margara, Marco Vassura, Rita Casadio |
Bioinform. | 2 |
| 2009 | On the Upper Bound of the Prediction Accuracy of Residue Contacts in Proteins with Correlated Mutations: The Case Study of the Similarity Matrices
Pietro Di Lena, Piero Fariselli, Luciano Margara, Marco Vassura, Rita Casadio |
WABI | 2 |
| 2009 | A graph theoretic approach to protein structure selection
Marco Vassura, Luciano Margara, Piero Fariselli, Rita Casadio |
Artif. Intell. Medicine | 3 |
| 2009 | Progress and challenges in predicting protein-protein interaction sitesabstractThe identification of protein-protein interaction sites is an essential intermediate step for mutant design and the prediction of protein networks. In recent years a significant number of methods have been developed to predict these interface residues and here we review the current status of the field. Progress in this area requires a clear view of the methodology applied, the data sets used for training and testing the systems, and the evaluation procedures. We have analysed the impact of a representative set of features and algorithms and highlighted the problems inherent in generating reliable protein data sets and in the posterior analysis of the results. Although it is clear that there have been some improvements in methods for predicting interacting sites, several major bottlenecks remain. Proteins in complexes are still under-represented in the structural databases and in particular many proteins involved in transient complexes are still to be crystallized. We provide suggestions for effective feature selection, and make it clear that community standards for testing, training and performance measures are necessary for progress in the field. Iakes Ezkurdia, Lisa Bartoli, Piero Fariselli, Rita Casadio, Alfonso Valencia, Michael L. Tress |
Briefings Bioinform. | 3 |
| 2009 | CCHMM_PROF: a HMM-based coiled-coil predictor with evolutionary informationabstractMOTIVATION: The widespread coiled-coil structural motif in proteins is known to mediate a variety of biological interactions. Recognizing a coiled-coil containing sequence and locating its coiled-coil domains are key steps towards the determination of the protein structure and function. Different tools are available for predicting coiled-coil domains in protein sequences, including those based on position-specific score matrices and machine learning methods. RESULTS: In this article, we introduce a hidden Markov model (CCHMM_PROF) that exploits the information contained in multiple sequence alignments (profiles) to predict coiled-coil regions. The new method discriminates coiled-coil sequences with an accuracy of 97% and achieves a true positive rate of 79% with only 1% of false positives. Furthermore, when predicting the location of coiled-coil segments in protein sequences, the method reaches an accuracy of 80% at the residue level and a best per-segment and per-protein efficiency of 81% and 80%, respectively. The results indicate that CCHMM_PROF outperforms all the existing tools and can be adopted for large-scale genome annotation. AVAILABILITY: The dataset is available at http://www.biocomp.unibo.it/ approximately lisa/coiled-coils. The predictor is freely available at http://gpcr.biocomp.unibo.it/cgi/predictors/cchmmprof/pred_cchmmprof.cgi. CONTACT: [email protected]. Lisa Bartoli, Piero Fariselli, Anders Krogh, Rita Casadio |
Bioinform. | 2 |
| 2008 | Predicting protein thermostability changes from sequence upon multiple mutationsabstractMOTIVATION: A basic question in protein science is to which extent mutations affect protein thermostability. This knowledge would be particularly relevant for engineering thermostable enzymes. In several experimental approaches, this issue has been serendipitously addressed. It would be therefore convenient providing a computational method that predicts when a given protein mutant is more thermostable than its corresponding wild-type. RESULTS: We present a new method based on support vector machines that is able to predict whether a set of mutations (including insertion and deletions) can enhance the thermostability of a given protein sequence. When trained and tested on a redundancy-reduced dataset, our predictor achieves 88% accuracy and a correlation coefficient equal to 0.75. Our predictor also correctly classifies 12 out of 14 experimentally characterized protein mutants with enhanced thermostability. Finally, it correctly detects all the 11 mutated proteins whose increase in stability temperature is >10 degrees C. AVAILABILITY: The dataset and the list of protein clusters adopted for the SVM cross-validation are available at the web site http://lipid.biocomp.unibo.it/~ludovica/thermo-meso-MUT. Ludovica Montanucci, Piero Fariselli, Pier Luigi Martelli, Rita Casadio |
ISMB | 2 |
| 2008 | FT-COMAR: fault tolerant three-dimensional structure reconstruction from protein contact mapsabstractAbstract Summary: Fault Tolerant Contact Map Reconstruction (FT-COMAR) is a heuristic algorithm for the reconstruction of the protein three-dimensional structure from (possibly) incomplete (i.e. containing unknown entries) and noisy contact maps. FT-COMAR runs within minutes, allowing its application to a large-scale number of predictions. Availability: http://bioinformatics.cs.unibo.it/FT-COMAR Contact: [email protected] Supplementary information: Supplementary data are available on Bioinformatics online. Marco Vassura, Luciano Margara, Pietro Di Lena, Filippo Medri, Piero Fariselli, Rita Casadio |
Bioinform. | 5 |
| 2008 | A three-state prediction of single point mutations on protein stability changesabstractBACKGROUND: A basic question of protein structural studies is to which extent mutations affect the stability. This question may be addressed starting from sequence and/or from structure. In proteomics and genomics studies prediction of protein stability free energy change (DeltaDeltaG) upon single point mutation may also help the annotation process. The experimental DeltaDeltaG values are affected by uncertainty as measured by standard deviations. Most of the DeltaDeltaG values are nearly zero (about 32% of the DeltaDeltaG data set ranges from -0.5 to 0.5 kcal/mole) and both the value and sign of DeltaDeltaG may be either positive or negative for the same mutation blurring the relationship among mutations and expected DeltaDeltaG value. In order to overcome this problem we describe a new predictor that discriminates between 3 mutation classes: destabilizing mutations (DeltaDeltaG<-1.0 kcal/mol), stabilizing mutations (DeltaDeltaG>1.0 kcal/mole) and neutral mutations (-1.0</=DeltaDeltaG</=1.0 kcal/mole). RESULTS: In this paper a support vector machine starting from the protein sequence or structure discriminates between stabilizing, destabilizing and neutral mutations. We rank all the possible substitutions according to a three state classification system and show that the overall accuracy of our predictor is as high as 56% when performed starting from sequence information and 61% when the protein structure is available, with a mean value correlation coefficient of 0.27 and 0.35, respectively. These values are about 20 points per cent higher than those of a random predictor. CONCLUSIONS: Our method improves the quality of the prediction of the free energy change due to single point protein mutations by adopting a hypothesis of thermodynamic reversibility of the existing experimental data. By this we both recast the thermodynamic symmetry of the problem and balance the distribution of the available experimental measurements of free energy changes. This eliminates possible overestimations of the previously described methods trained on an unbalanced data set comprising a number of destabilizing mutations higher than stabilizing ones. Emidio Capriotti, Piero Fariselli, Ivan Rossi, Rita Casadio |
BMC Bioinform. | 2 |
| 2008 | Reconstruction of 3D Structures From Protein Contact MapsabstractThe prediction of the protein tertiary structure from solely its residue sequence (the so called Protein Folding Problem) is one of the most challenging problems in Structural Bioinformatics. We focus on the protein residue contact map. When this map is assigned it is possible to reconstruct the 3D structure of the protein backbone. The general problem of recovering a set of 3D coordinates consistent with some given contact map is known as a unit-disk-graph realization problem and it has been recently proven to be NP-Hard. In this paper we describe a heuristic method (COMAR) that is able to reconstruct with an unprecedented rate (3-15 seconds) a 3D model that exactly matches the target contact map of a protein. Working with a non-redundant set of 1760 proteins, we find that the scoring efficiency of finding a 3D model very close to the protein native structure depends on the threshold value adopted to compute the protein residue contact map. Contact maps whose threshold values range from 10 to 18 Angstroms allow reconstructing 3D models that are very similar to the proteins native structure. Marco Vassura, Luciano Margara, Pietro Di Lena, Filippo Medri, Piero Fariselli, Rita Casadio |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2007 | Reconstruction of 3D Structures from Protein Contact Maps
Marco Vassura, Luciano Margara, Filippo Medri, Pietro Di Lena, Piero Fariselli, Rita Casadio |
ISBRA | 5 |
| 2007 | Fault Tolerance for Large Scale Protein 3D Reconstruction from Contact Maps
Marco Vassura, Luciano Margara, Pietro Di Lena, Filippo Medri, Piero Fariselli, Rita Casadio |
WABI | 5 |
| 2007 | The WWWH of remote homolog detection: The state of the artabstractThe detection of remote homolog pairs of proteins using computational methods is a pivotal problem in structural bioinformatics, aiming to compute protein folds on the basis of information in the database of known structures. In the last 25 years, several methods have been developed to tackle this problem, based on different approaches including sequence-sequence alignments and/or structure comparison. In this article, we will briefly discuss When, Why, Where and How (WWWH) to perform remote homology search, reviewing some of the most widely adopted computational approaches. The specific aim is highlighting the basic criteria implemented by different research groups and commenting on the status of the art as well as on still-open questions. Piero Fariselli, Ivan Rossi, Emidio Capriotti, Rita Casadio |
Briefings Bioinform. | 1 |
| 2007 | A computational approach for detecting peptidases and their specific inhibitors at the genome levelabstractBACKGROUND: Peptidases are proteolytic enzymes responsible for fundamental cellular activities in all organisms. Apparently about 2-5% of the genes encode for peptidases, irrespectively of the organism source. The basic peptidase function is "protein digestion" and this can be potentially dangerous in living organisms when it is not strictly controlled by specific inhibitors. In genome annotation a basic question is to predict gene function. Here we describe a computational approach that can filter peptidases and their inhibitors out of a given proteome. Furthermore and as an added value to MEROPS, a specific database for peptidases already available in the public domain, our method can predict whether a pair of peptidase/inhibitor can interact, eventually listing all possible predicted ligands (peptidases and/or inhibitors). RESULTS: We show that by adopting a decision-tree approach the accuracy of PROSITE and HMMER in detecting separately the four major peptidase types (Serine, Aspartic, Cysteine and Metallo- Peptidase) and their inhibitors among a non redundant set of globular proteins can be improved by some percentage points with respect to that obtained with each method separately. More importantly, our method can then predict pairs of peptidases and interacting inhibitors, scoring a joint global accuracy of 99% with coverage for the positive cases (peptidase/inhibitor) close to 100% and a correlation coefficient of 0.91%. In this task the decision-tree approach outperforms the single methods. CONCLUSION: The decision-tree can reliably classify protein sequences as peptidases or inhibitors, belonging to a certain class, and can provide a comprehensive list of possible interacting pairs of peptidase/inhibitor. This information can help the design of experiments to detect interacting peptidase/inhibitor complexes and can speed up the selection of possible interacting candidates, without searching for them separately and manually combining the obtained results. A web server specifically developed for annotating peptidases and their inhibitors (HIPPIE) is available at http://gpcr.biocomp.unibo.it/cgi/predictors/hippie/pred_hippie.cgi. Lisa Bartoli, Remo Calabrese, Piero Fariselli, Damiano G. Mita, Rita Casadio |
BMC Bioinform. | 3 |
| 2005 | A new decoding algorithm for hidden Markov models improves the prediction of the topology of all-beta membrane proteinsabstractBACKGROUND: Structure prediction of membrane proteins is still a challenging computational problem. Hidden Markov models (HMM) have been successfully applied to the problem of predicting membrane protein topology. In a predictive task, the HMM is endowed with a decoding algorithm in order to assign the most probable state path, and in turn the labels, to an unknown sequence. The Viterbi and the posterior decoding algorithms are the most common. The former is very efficient when one path dominates, while the latter, even though does not guarantee to preserve the HMM grammar, is more effective when several concurring paths have similar probabilities. A third good alternative is 1-best, which was shown to perform equal or better than Viterbi. RESULTS: In this paper we introduce the posterior-Viterbi (PV) a new decoding which combines the posterior and Viterbi algorithms. PV is a two step process: first the posterior probability of each state is computed and then the best posterior allowed path through the model is evaluated by a Viterbi algorithm. CONCLUSION: We show that PV decoding performs better than other algorithms when tested on the problem of the prediction of the topology of beta-barrel membrane proteins. Piero Fariselli, Pier Luigi Martelli, Rita Casadio |
BMC Bioinform. | 1 |
| 2004 | ConSeq: the identification of functionally and structurally important residues in protein sequencesabstractMOTIVATION: ConSeq is a web server for the identification of biologically important residues in protein sequences. Functionally important residues that take part, e.g. in ligand binding and protein-protein interactions, are often evolutionarily conserved and are most likely to be solvent-accessible, whereas conserved residues within the protein core most probably have an important structural role in maintaining the protein's fold. Thus, estimated evolutionary rates, as well as relative solvent accessibility predictions, are assigned to each amino acid in the sequence; both are subsequently used to indicate residues that have potential structural or functional importance. AVAILABILITY: The ConSeq web server is available at http://conseq.bioinfo.tau.ac.il/ SUPPLEMENTARY INFORMATION: The ConSeq methodology, a description of its performance in a set of five well-documented proteins, a comparison to other methods, and the outcome of its application to a set of 111 proteins of unknown function, are presented at http://conseq.bioinfo.tau.ac.il/ under 'OVERVIEW', 'VALIDATION', 'COMPARISON' and 'PREDICTIONS', respectively. Carine Berezin, Fabian Glaser, Josef Rosenberg, Inbal Paz, Tal Pupko, Piero Fariselli, Rita Casadio, Nir Ben-Tal |
Bioinform. | 6 |
| 2003 | In silico prediction of the structure of membrane proteins: Is it feasible?abstractIn the 'omic' era, hundreds of genomes are available for protein sequence analysis, and some 30 per cent of all sequences are of membrane proteins. Unlike globular proteins, a 3D model for membrane proteins can hardly be computed starting from the sequence. Why is this so? What can we really compute and with what reliability? These and other matters are outlined. Rita Casadio, Piero Fariselli, Pier Luigi Martelli |
Briefings Bioinform. | 2 |
| 2003 | SPEPlip: the detection of signal peptide and lipoprotein cleavage sitesabstractSUMMARY: SPEPlip is a neural network-based method, trained and tested on a set of experimentally derived signal peptides from eukaryotes and prokaryotes. SPEPlip identifies the presence of sorting signals and predicts their cleavage sites. The accuracy in cross-validation is similar to that of other available programs: the rate of false positives is 4 and 6%, for prokaryotes and eukaryotes respectively and that of false negatives is 3% in both cases. When a set of 409 prokaryotic lipoproteins is predicted, SPEPlip predicts 97% of the chains in the signal peptide class. However, by integrating SPEPlip with a regular expression search utility based on the PROSITE pattern, we can successfully discriminate signal peptide-containing chains from lipoproteins. We propose the method for detecting and discriminating signal peptides containing chains and lipoproteins. AVAILABILITY: It can be accessed through the web page at http://gpcr.biocomp.unibo.it/predictors/ Piero Fariselli, Giacomo Finocchiaro, Rita Casadio |
Bioinform. | 1 |
| 2003 | MaxSubSeq: an algorithm for segment-length optimization. The case study of the transmembrane spanning segmentsabstractMOTIVATION: A problem in predicting the topography of transmembrane proteins is the optimal localization of the transmembrane segments along the protein sequences, provided that each residue is associated with a propensity of being or not being included in the transmembrane protein region. From previous work it is known that post-processing of propensity signals with suited algorithms can greatly improve the quality and the accuracy of the predictions. In this paper we describe a general dynamic programming-like algorithm (MaxSubSeq, Maximal SubSequence) specifically designed to optimize the number and length of segments with constrained length in a given protein sequence. Previous application of our algorithm, has proved its effectiveness in the optimization task of both neural network and hidden Markov models output, and in this paper we present the detailed description of MaxSubSeq. RESULTS: We describe the application of MaxSubSeq to the location of both helical and beta strand transmembrane segments, optimizing the outputs derived with different predictive algorithms. For all-alpha transmembrane proteins we use both the standard Kyte-Doolittle (KD) hydropathy scale and the TMHMM predictor (http://www.cbs.dtu.dk/). Using a set of 188 well characterized membrane proteins, MaxSubSeq nearly doubles the correct location of transmembrane segments as compared to the standard KD hydrophobicity plot, reaching 51% accuracy. If MaxSubSeq is used to optimize the TMHMM method the accuracy increases from 68 to 72%. When used to regularize the prediction of beta transmembrane strands, obtained using both a neural network and a HMM based predictors, MaxSubSeq increases the accuracy per protein up to 72 and 73% respectively. AVAILABILITY: The program is available upon request to the authors, or it is accessible through our web server (http://gpcr.biocomp.unibo.it/predictors/) Piero Fariselli, Michele Finelli, Davide Marchignoli, Pier Luigi Martelli, Ivan Rossi, Rita Casadio |
Bioinform. | 1 |
| 2002 | A sequence-profile-based HMM for predicting and discriminating beta barrel membrane proteinsabstractMOTIVATION: Membrane proteins are an abundant and functionally relevant subset of proteins that putatively include from about 15 up to 30% of the proteome of organisms fully sequenced. These estimates are mainly computed on the basis of sequence comparison and membrane protein prediction. It is therefore urgent to develop methods capable of selecting membrane proteins especially in the case of outer membrane proteins, barely taken into consideration when proteome wide analysis is performed. This will also help protein annotation when no homologous sequence is found in the database. Outer membrane proteins solved so far at atomic resolution interact with the external membrane of bacteria with a characteristic beta barrel structure comprising different even numbers of beta strands (beta barrel membrane proteins). In this they differ from the membrane proteins of the cytoplasmic membrane endowed with alpha helix bundles (all alpha membrane proteins) and need specialised predictors. RESULTS: We develop a HMM model, which can predict the topology of beta barrel membrane proteins using, as input, evolutionary information. The model is cyclic with 6 types of states: two for the beta strand transmembrane core, one for the beta strand cap on either side of the membrane, one for the inner loop, one for the outer loop and one for the globular domain state in the middle of each loop. The development of a specific input for HMM based on multiple sequence alignment is novel. The accuracy per residue of the model is 83% when a jack knife procedure is adopted. With a model optimisation method using a dynamic programming algorithm seven topological models out of the twelve proteins included in the testing set are also correctly predicted. When used as a discriminator, the model is rather selective. At a fixed probability value, it retains 84% of a non-redundant set comprising 145 sequences of well-annotated outer membrane proteins. Concomitantly, it correctly rejects 90% of a set of globular proteins including about 1200 chains with low sequence identity (<30%) and 90% of a set of all alpha membrane proteins, including 188 chains. Pier Luigi Martelli, Piero Fariselli, Anders Krogh, Rita Casadio |
ISMB | 2 |
| 2001 | Prediction of disulfide connectivity in proteinsabstractMOTIVATION: A major problem in protein structure prediction is the correct location of disulfide bridges in cysteine-rich proteins. In protein-folding prediction, the location of disulfide bridges can strongly reduce the search in the conformational space. Therefore the correct prediction of the disulfide connectivity starting from the protein residue sequence may also help in predicting its 3D structure. RESULTS: In this paper we equate the problem of predicting the disulfide connectivity in proteins to a problem of finding the graph matching with the maximum weight. The graph vertices are the residues of cysteine-forming disulfide bridges, and the weight edges are contact potentials. In order to solve this problem we develop and test different residue contact potentials. The best performing one, based on the Edmonds-Gabow algorithm and Monte-Carlo simulated annealing reaches an accuracy significantly higher than that obtained with a general mean force contact potential. Significantly, in the case of proteins with four disulfide bonds in the structure, the accuracy is 17 times higher than that of a random predictor. The method presented here can be used to locate putative disulfide bridges in protein-folding. AVAILABILITY: The program is available upon request from the authors. CONTACT: [email protected]; [email protected]. Piero Fariselli, Rita Casadio |
Bioinform. | 1 |
| 2001 | RCNPRED: prediction of the residue co-ordination numbers in proteinsabstractUNLABELLED: The RCNPRED server implements a neural network-based method to predict the co-ordination numbers of residues starting from the protein sequence. Using evolutionary information as input, RCNPRED predicts the residue states of the proteins in the database with 69% accuracy and scores 12 percentage points higher than a simple statistical method. Moreover the server implements a neural network to predict the relative solvent accessibility of each residue. A protein sequence can be directly submitted to RCNPRED: residue co-ordination numbers and solvent accessibility for each chain are returned via e-mail. AVAILABILITY: Freely available to non-commercial users at http://prion.biocomp.unibo.it/rcnpred.html. Piero Fariselli, Rita Casadio |
Bioinform. | 1 |
| 2000 | Prediction of the Number of Residue Contacts in Proteins
Piero Fariselli, Rita Casadio |
ISMB | 1 |
| 1999 | A Data Base of Minimally Frustrated Alpha-Helical Segments Extracted from Proteins According to an Entropy Criterion
Rita Casadio, Mario Compiani, Piero Fariselli, Pier Luigi Martelli |
ISMB | 3 |
| 1997 | Self-Organizing Neural Maps of the Coding Sequences of G-protein-coupled Receptors Reveal Local Domains Associated with Potentially Functional Determinants in the Proteins
P. Arrigo, Piero Fariselli, Rita Casadio |
ISMB | 2 |
| 1997 | The Prediction of Protein Secondary Structure with a Cascade Correlation Learning Architecture of Neural Networks
Francesco Vivarelli, Piero Fariselli, Rita Casadio |
Neural Comput. Appl. | 2 |
| 1996 | Refining Neural Network Predictions for Helical Transmembrane Proteins by Dynamic Programming
Burkhard Rost, Rita Casadio, Piero Fariselli |
ISMB | 3 |
| 1996 | HTP: a neural network-based method for predicting the topology of helical transmembrane domains in proteinsabstractIn this paper we describe a microcomputer program (HTP) for predicting the location and orientation of alpha-helical transmembrane segments in integral membrane proteins. HTP is a neural network-based tool which gives as output the protein membrane topology based on the statistical propensity of residues to be located in external and internal loops. This method, which uses single protein sequences as input to the network system, correctly predicts the topology of 71 out of 92 membrane proteins of putative membrane orientation, independently of the protein source. Piero Fariselli, Rita Casadio |
Comput. Appl. Biosci. | 1 |
| 1995 | Predicting Free Energy Contribution to the Conformational Stability of Folded Proteins From the Residue Sequence with Radial Basis Function Networks
Rita Casadio, Mario Compiani, Piero Fariselli, Francesco Vivarelli |
ISMB | 3 |
| 1995 | LGANN: a parallel system combining a local genetic algorithm and neural networks for the prediction of secondary structure of proteinsabstractIn this work we describe a parallel system consisting of feed-forward neural networks supervised by a local genetic algorithm. The system is implemented in a transputer architecture and is used to predict the secondary structures of globular proteins. This method allows a wide search in the parameter space of the neural networks and the determination of their optimal topology for the predictive task. Different neural network topologies are selected by the genetic algorithm on the basis of minimal values of mean square errors on the testing set. When the alpha-helix, beta-strand and random coil motifs of secondary structures are discriminated, the maximal efficiency obtained is 0.62, with correlation coefficients of 0.35, 0.31 and 0.37 respectively. This level of accuracy is similar to that previously attained by means of neural networks without hidden layers and using single protein sequences as input. The results validate the neural network topologies used for the prediction of protein secondary structures and highlight the relevance of the input information in determining the limit of their performance. Francesco Vivarelli, G. Giusti, Marco Villani 0001, Renato Campanini, Piero Fariselli, Mario Compiani, Rita Casadio |
Comput. Appl. Biosci. | 5 |