EDBT 2026 Demo / reviewers in the wild / expert
Fabrizio Pucci
dblp:155/4009
· DBLP profile ↗
12ranked-venue papers
4as first author
7since 2021 · last 2025
0000-0003-2916-022XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 12 · 4 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Residue conservation and solvent accessibility are (almost) all you need for predicting mutational effects in proteinsabstractMOTIVATION: Predicting how mutations impact protein biophysical properties remains a significant challenge in computational biology. In recent years, numerous predictors, primarily deep learning models, have been developed to address this problem; however, issues such as their lack of interpretability and limited accuracy persist. RESULTS: We showed that a simple evolutionary score, based on the log-odd ratio of wild-type and mutated residue frequencies in evolutionary related proteins, when scaled by the residue's relative solvent accessibility, performs on par with or slightly outperforms most of the benchmarked predictors, many of which are considerably more complex. The evaluation is performed on mutations from the ProteinGym deep mutational scanning dataset collection, which measures various properties such as stability, activity or fitness. This raises further questions about what these complex models actually learn and highlights their limitations in addressing prediction of mutational landscape. AVAILABILITY AND IMPLEMENTATION: The RSALOR model is available as a user-friendly Python package that can be installed from the PyPI repository. The code is freely available at https://github.com/3BioCompBio/RSALOR. Matsvei Tsishyn, Pauline Hermans, Marianne Rooman, Fabrizio Pucci |
Bioinform. | 4 |
| 2024 | Quantification of biases in predictions of protein-protein binding affinity changes upon mutationsabstractUnderstanding the impact of mutations on protein-protein binding affinity is a key objective for a wide range of biotechnological applications and for shedding light on disease-causing mutations, which are often located at protein-protein interfaces. Over the past decade, many computational methods using physics-based and/or machine learning approaches have been developed to predict how protein binding affinity changes upon mutations. They all claim to achieve astonishing accuracy on both training and test sets, with performances on standard benchmarks such as SKEMPI 2.0 that seem overly optimistic. Here we benchmarked eight well-known and well-used predictors and identified their biases and dataset dependencies, using not only SKEMPI 2.0 as a test set but also deep mutagenesis data on the severe acute respiratory syndrome coronavirus 2 spike protein in complex with the human angiotensin-converting enzyme 2. We showed that, even though most of the tested methods reach a significant degree of robustness and accuracy, they suffer from limited generalizability properties and struggle to predict unseen mutations. Interestingly, the generalizability problems are more severe for pure machine learning approaches, while physics-based methods are less affected by this issue. Moreover, undesirable prediction biases toward specific mutation properties, the most marked being toward destabilizing mutations, are also observed and should be carefully considered by method developers. We conclude from our analyses that there is room for improvement in the prediction models and suggest ways to check, assess and improve their generalizability and robustness. Matsvei Tsishyn, Fabrizio Pucci, Marianne Rooman |
Briefings Bioinform. | 2 |
| 2024 | pycofitness - Evaluating the fitness landscape of RNA and protein sequencesabstractMOTIVATION: The accurate prediction of how mutations change biophysical properties of proteins or RNA is a major goal in computational biology with tremendous impacts on protein design and genetic variant interpretation. Evolutionary approaches such as coevolution can help solving this issue. RESULTS: We present pycofitness, a standalone Python-based software package for the in silico mutagenesis of protein and RNA sequences. It is based on coevolution and, more specifically, on a popular inverse statistical approach, namely direct coupling analysis by pseudo-likelihood maximization. Its efficient implementation and user-friendly command line interface make it an easy-to-use tool even for researchers with no bioinformatics background. To illustrate its strengths, we present three applications in which pycofitness efficiently predicts the deleteriousness of genetic variants and the effect of mutations on protein fitness and thermodynamic stability. AVAILABILITY AND IMPLEMENTATION: https://github.com/KIT-MBS/pycofitness. Fabrizio Pucci, Mehari B. Zerihun, Marianne Rooman, Alexander Schug |
Bioinform. | 1 |
| 2023 | Critical review of conformational B-cell epitope prediction methodsabstractAccurate in silico prediction of conformational B-cell epitopes would lead to major improvements in disease diagnostics, drug design and vaccine development. A variety of computational methods, mainly based on machine learning approaches, have been developed in the last decades to tackle this challenging problem. Here, we rigorously benchmarked nine state-of-the-art conformational B-cell epitope prediction webservers, including generic and antibody-specific methods, on a dataset of over 250 antibody-antigen structures. The results of our assessment and statistical analyses show that all the methods achieve very low performances, and some do not perform better than randomly generated patches of surface residues. In addition, we also found that commonly used consensus strategies that combine the results from multiple webservers are at best only marginally better than random. Finally, we applied all the predictors to the SARS-CoV-2 spike protein as an independent case study, and showed that they perform poorly in general, which largely recapitulates our benchmarking conclusions. We hope that these results will lead to greater caution when using these tools until the biases and issues that limit current methods have been addressed, promote the use of state-of-the-art evaluation methodologies in future publications and suggest new strategies to improve the performance of conformational B-cell epitope prediction methods. Gabriel Cia, Fabrizio Pucci, Marianne Rooman |
Briefings Bioinform. | 2 |
| 2022 | SpikePro: a webserver to predict the fitness of SARS-CoV-2 variantsabstractMOTIVATION: The SARS-CoV-2 virus has shown a remarkable ability to evolve and spread across the globe through successive waves of variants since the original Wuhan lineage. Despite all the efforts of the last 2 years, the early and accurate prediction of variant severity is still a challenging issue which needs to be addressed to help, for example, the decision of activating COVID-19 plans long before the peak of new waves. Upstream preparation would indeed make it possible to avoid the overflow of health systems and limit the most severe cases. RESULTS: We recently developed SpikePro, a structure-based computational model capable of quickly and accurately predicting the viral fitness of a variant from its spike protein sequence. It is based on the impact of mutations on the stability of the spike protein as well as on its binding affinity for the angiotensin-converting enzyme 2 (ACE2) and for a set of neutralizing antibodies. It yields a precise indication of the virus transmissibility, infectivity, immune escape and basic reproduction rate. We present here an updated version of the model that is now available on an easy-to-use webserver, and illustrate its power in a retrospective study of fitness evolution and reproduction rate of the main viral lineages. SpikePro is thus expected to be great help to assess the fitness of newly emerging SARS-CoV-2 variants in genomic surveillance and viral evolution programs. AVAILABILITY AND IMPLEMENTATION: SpikePro webserver http://babylone.ulb.ac.be/SpikePro/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Gabriel Cia, Jean Marc Kwasigroch, Marianne Rooman, Fabrizio Pucci |
Bioinform. | 4 |
| 2021 | MutaFrame - an interpretative visualization framework for deleteriousness prediction of missense variants in the human exomeabstractMOTIVATION: High-throughput experiments are generating ever increasing amounts of various -omics data, so shedding new light on the link between human disorders, their genetic causes and the related impact on protein behavior and structure. While numerous bioinformatics tools now exist that predict which variants in the human exome cause diseases, few tools predict the reasons why they might do so. Yet, understanding the impact of variants at the molecular level is a prerequisite for the rational development of targeted drugs or personalized therapies. RESULTS: We present the updated MutaFrame webserver, which aims to meet this need. It offers two deleteriousness prediction softwares, DEOGEN2 and SNPMuSiC, and is designed for bioinformaticians and medical researchers who want to gain insights into the origins of monogenic diseases. It contains information at two levels for each human protein: its amino acid sequence and its three-dimensional structure; we used the experimental structures whenever available, and modeled structures otherwise. MutaFrame also includes higher-level information, such as protein essentiality and protein-protein interactions. It has a user-friendly interface for the interpretation of results and a convenient visualization system for protein structures, in which the variant positions introduced by the user and other structural information are shown. In this way, MutaFrame aids our understanding of the pathogenic processes caused by single-site mutations and their molecular and contextual interpretation. AVAILABILITY AND IMPLEMENTATION: Mutaframe webserver at http://mutaframe.com/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. François Ancien, Fabrizio Pucci, Wim F. Vranken, Marianne Rooman |
Bioinform. | 2 |
| 2021 | SWOTein: a structure-based approach to predict stability Strengths and Weaknesses of prOTEINsabstractMOTIVATION: Although structured proteins adopt their lowest free energy conformation in physiological conditions, the individual residues are generally not in their lowest free energy conformation. Residues that are stability weaknesses are often involved in functional regions, whereas stability strengths ensure local structural stability. The detection of strengths and weaknesses provides key information to guide protein engineering experiments aiming to modulate folding and various functional processes. RESULTS: We developed the SWOTein predictor which identifies strong and weak residues in proteins on the basis of three types of statistical energy functions describing local interactions along the chain, hydrophobic forces and tertiary interactions. The large-scale analysis of the different types of strengths and weaknesses demonstrated their complementarity and the enhancement of the information they provide. Moreover, a good average correlation was observed between predicted and experimental strengths and weaknesses obtained from native hydrogen exchange data. SWOTein application to three test cases further showed its suitability to predict and interpret strong and weak residues in the context of folding, conformational changes and protein-protein binding. In summary, SWOTein is both fast and accurate and can be applied at small and large scale to analyze and modulate folding and molecular recognition processes. AVAILABILITY: The SWOTein webserver provides the list of predicted strengths and weaknesses and a protein structure visualization tool that facilitates the interpretation of the predictions. It is freely available for academic use at http://babylone.ulb.ac.be/SWOTein/. Qingzhen Hou, Fabrizio Pucci, François Ancien, Jean Marc Kwasigroch, Raphaël Bourgeas, Marianne Rooman |
Bioinform. | 2 |
| 2020 | SOLart: a structure-based method to predict protein solubility and aggregationabstractMOTIVATION: The solubility of a protein is often decisive for its proper functioning. Lack of solubility is a major bottleneck in high-throughput structural genomic studies and in high-concentration protein production, and the formation of protein aggregates causes a wide variety of diseases. Since solubility measurements are time-consuming and expensive, there is a strong need for solubility prediction tools. RESULTS: We have recently introduced solubility-dependent distance potentials that are able to unravel the role of residue-residue interactions in promoting or decreasing protein solubility. Here, we extended their construction by defining solubility-dependent potentials based on backbone torsion angles and solvent accessibility, and integrated them, together with other structure- and sequence-based features, into a random forest model trained on a set of Escherichia coli proteins with experimental structures and solubility values. We thus obtained the SOLart protein solubility predictor, whose most informative features turned out to be folding free energy differences computed from our solubility-dependent statistical potentials. SOLart performances are very good, with a Pearson correlation coefficient between experimental and predicted solubility values of almost 0.7 both in cross-validation on the training dataset and in an independent set of Saccharomyces cerevisiae proteins. On test sets of modeled structures, only a limited drop in performance is observed. SOLart can thus be used with both high-resolution and low-resolution structures, and clearly outperforms state-of-art solubility predictors. It is available through a user-friendly webserver, which is easy to use by non-expert scientists. AVAILABILITY AND IMPLEMENTATION: The SOLart webserver is freely available at http://babylone.ulb.ac.be/SOLART/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Qingzhen Hou, Jean Marc Kwasigroch, Marianne Rooman, Fabrizio Pucci |
Bioinform. | 4 |
| 2020 | pydca v1.0: a comprehensive software for direct coupling analysis of RNA and protein sequencesabstractMOTIVATION: The ongoing advances in sequencing technologies have provided a massive increase in the availability of sequence data. This made it possible to study the patterns of correlated substitution between residues in families of homologous proteins or RNAs and to retrieve structural and stability information. Direct coupling analysis (DCA) infers coevolutionary couplings between pairs of residues indicating their spatial proximity, making such information a valuable input for subsequent structure prediction. RESULTS: Here, we present pydca, a standalone Python-based software package for the DCA of protein- and RNA-homologous families. It is based on two popular inverse statistical approaches, namely, the mean-field and the pseudo-likelihood maximization and is equipped with a series of functionalities that range from multiple sequence alignment trimming to contact map visualization. Thanks to its efficient implementation, features and user-friendly command line interface, pydca is a modular and easy-to-use tool that can be used by researchers with a wide range of backgrounds. AVAILABILITY AND IMPLEMENTATION: pydca can be obtained from https://github.com/KIT-MBS/pydca or from the Python Package Index under the MIT License. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Mehari B. Zerihun, Fabrizio Pucci, Emanuel K. Peter, Alexander Schug |
Bioinform. | 2 |
| 2018 | Quantification of biases in predictions of protein stability changes upon mutationsabstractMotivation: Bioinformatics tools that predict protein stability changes upon point mutations have made a lot of progress in the last decades and have become accurate and fast enough to make computational mutagenesis experiments feasible, even on a proteome scale. Despite these achievements, they still suffer from important issues that must be solved to allow further improving their performances and utilizing them to deepen our insights into protein folding and stability mechanisms. One of these problems is their bias toward the learning datasets which, being dominated by destabilizing mutations, causes predictions to be better for destabilizing than for stabilizing mutations. Results: We thoroughly analyzed the biases in the prediction of folding free energy changes upon point mutations (ΔΔG0) and proposed some unbiased solutions. We started by constructing a dataset Ssym of experimentally measured ΔΔG0s with an equal number of stabilizing and destabilizing mutations, by collecting mutations for which the structure of both the wild-type and mutant protein is available. On this balanced dataset, we assessed the performances of 15 widely used ΔΔG0 predictors. After the astonishing observation that almost all these methods are strongly biased toward destabilizing mutations, especially those that use black-box machine learning, we proposed an elegant way to solve the bias issue by imposing physical symmetries under inverse mutations on the model structure, which we implemented in PoPMuSiCsym. This new predictor constitutes an efficient trade-off between accuracy and absence of biases. Some final considerations and suggestions for further improvement of the predictors are discussed. Supplementary information: Supplementary data are available at Bioinformatics online. Note: The article 10.1093/bioinformatics/bty340/, published alongside this paper, also addresses the problem of biases in protein stability change predictions. Fabrizio Pucci, Katrien V. Bernaerts, Jean Marc Kwasigroch, Marianne Rooman |
Bioinform. | 1 |
| 2017 | SCooP: an accurate and fast predictor of protein stability curves as a function of temperatureabstractMOTIVATION: The molecular bases of protein stability remain far from elucidated even though substantial progress has been made through both computational and experimental investigations. One of the most challenging goals is the development of accurate prediction tools of the temperature dependence of the standard folding free energy ΔG(T). Such predictors have an enormous series of potential applications, which range from drug design in the biopharmaceutical sector to the optimization of enzyme activity for biofuel production. There is thus an important demand for novel, reliable and fast predictors. RESULTS: We present the SCooP algorithm, which is a significant step towards accurate temperature-dependent stability prediction. This automated tool uses the protein structure and the host organism as sole entries and predicts the full T-dependent stability curve of monomeric proteins assumed to follow a two-state folding transition. Equivalently, it predicts all the thermodynamic quantities associated to the folding transition, namely the melting temperature Tm, the standard folding enthalpy ΔHm measured at Tm, and the standard folding heat capacity ΔCp. The cross-validated performances are good, with correlation coefficients between predicted and experimental values equal to [0.80, 0.83, 0.72] for ΔHm, ΔCp and Tm, respectively, which increase up to [0.88, 0.90, 0.78] upon the removal of 10% outliers. Moreover, the stability curve prediction of a target protein is very fast: it takes less than a minute. SCooP can thus potentially be applied on a structurome scale. This opens new perspectives of large-scale analyses of protein stability, which is of considerable interest for protein engineering. AVAILABILITY AND IMPLEMENTATION: The SCooP webserver is freely available at http://babylone.ulb.ac.be/SCooP. CONTACT: [email protected], [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Fabrizio Pucci, Jean Marc Kwasigroch, Marianne Rooman |
Bioinform. | 1 |
| 2014 | Stability Curve Prediction of Homologous Proteins Using Temperature-Dependent Statistical PotentialsabstractThe unraveling and control of protein stability at different temperatures is a fundamental problem in biophysics that is substantially far from being quantitatively and accurately solved, as it requires a precise knowledge of the temperature dependence of amino acid interactions. In this paper we attempt to gain insight into the thermal stability of proteins by designing a tool to predict the full stability curve as a function of the temperature for a set of 45 proteins belonging to 11 homologous families, given their sequence and structure, as well as the melting temperature (Tm) and the change in heat capacity (ΔCP) of proteins belonging to the same family. Stability curves constitute a fundamental instrument to analyze in detail the thermal stability and its relation to the thermodynamic stability, and to estimate the enthalpic and entropic contributions to the folding free energy. In summary, our approach for predicting the protein stability curves relies on temperature-dependent statistical potentials derived from three datasets of protein structures with targeted thermal stability properties. Using these potentials, the folding free energies (ΔG) at three different temperatures were computed for each protein. The Gibbs-Helmholtz equation was then used to predict the protein's stability curve as the curve that best fits these three points. The results are quite encouraging: the standard deviations between the experimental and predicted Tm's, ΔCP's and folding free energies at room temperature (ΔG25) are equal to 13° C, 1.3 kcal/(mol° C) and 4.1 kcal/mol, respectively, in cross-validation. The main sources of error and some further improvements and perspectives are briefly discussed. Fabrizio Pucci, Marianne Rooman |
PLoS Comput. Biol. | 1 |