EDBT 2026 Demo / reviewers in the wild / expert
Frederic Rousseau 0001
dblp:45/3248
· DBLP profile ↗
23ranked-venue papers
0as first author
8since 2021 · last 2025
0000-0002-9189-7399ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 23 · 8 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | FoldX force field revisited, an improved versionabstractMOTIVATION: The FoldX force field was originally validated with a database of 1000 mutants at a time when there were few high-resolution structures. Here, we have manually curated a database of 5556 mutants affecting protein stability, resulting in 2484 highly confident mutations denominated FoldX stability dataset (FSD), represented in non-redundant X-ray structures with <2.5 Å resolution, not involving duplicates, metals, or prosthetic groups. Using this database, we have created a new version of the FoldX force field by introducing pi stacking, pH dependency for all charged residues, improving aromatic-aromatic interactions, modifying the Ncap contribution and α-helix dipole, recalibrating the side-chain entropy of methionine, adjusting the H-bond parameters, and modifying the solvation contribution of tryptophan and others. RESULTS: These changes have led to significant improvements for the prediction of specific mutants involving the above residues/interactions and a statistically significant increase of FoldX predictions, as well as for the majority of the 20 aa. Removing all training sets data from FSD [Validation FoldX Stability Dataset (VFSD) dataset] resulted in improved predictions from R = 0.693 (RMSE = 1.277 kcal/mol) to R = 0.706 (RMSE = 1.252 kcal/mol) when compared with the previously released version. FoldX achieves 95% accuracy considering an error of ±0.85 kcal/mol in prediction and an area under the curve = 0.78 for the VFSD, predicting the sign of the energy change upon mutation. AVAILABILITY AND IMPLEMENTATION: FoldX versions 4.1 and 5.1 are freely available for academics at https://foldxsuite.crg.eu/. Javier Delgado, Raul Reche, Damiano Cianferoni, Gabriele Orlando, Rob van der Kant, Frederic Rousseau 0001, Joost Schymkowitz, Luis Serrano |
Bioinform. | 6 |
| 2025 | AXZ viewer: a web application to visualize unprocessed AFM-IR dataabstractMOTIVATION: Atomic Force Microscopy-based Infrared spectroscopy (AFM-IR) is a novel and innovative method for label-free high-resolution structural biology. However, the nature of the data files generated by AFM-IR instruments precludes investigation by conventional open-source scientific image analysis software suites. As a result, reporting of AFM-IR datasets is not standardized and the data itself is difficult to audit. RESULTS: We have developed a web application that allows anyone to open, review, and audit raw AFM-IR data files easily and without deep knowledge of the method. It also exposes all metadata recorded by the microscope at the time of measurement. The web application is based on a Python package that supports custom data analyses within the scientific Python ecosystem. This tool provides an accessible, transparent solution for AFM-IR data review, with the potential to support reproducibility and standardization in AFM-IR research and encourage wider adoption of this innovative spectroscopy method. AVAILABILITY AND IMPLEMENTATION: The web app is hosted at https://anasys-python-tools-gui.streamlit.app. Its source code is listed at https://github.com/wduverger/anasys-python-tools-gui. The underlying Python package is available at https://github.com/GeorgRamer/anasys-python-tools and can be installed using pip. Wouter Duverger, Georg Ramer, Nikolaos N. Louros, Joost Schymkowitz, Frederic Rousseau 0001 |
Bioinform. | 5 |
| 2025 | Charting the structure-sequence landscape of light chain amyloidsabstractMOTIVATION: Light chain amyloidosis is a disease where misfolded antibody light chains (LCs) form toxic amyloid fibrils, leading to organ damage. Although LC overproduction occurs in all cases, only certain individuals develop the disease, suggesting that specific LC sequences and properties drive amyloid formation. This process is complex, involving both protein sequence and environmental factors, but mutations that destabilize the LC fold are linked to amyloid aggregation. Despite the significance of the disease, our understanding of LC fibril formation remains limited due to the lack of extensive data and technical challenges in studying amyloid structures. To address this, a tool is needed to compare unknown LC sequences with known structures and predict which amyloids are likely to adopt new conformations, guiding experimental investigations. RESULTS: HMMSTUFF addresses this by using a Hidden Markov Model to generate similarity scores between LC sequences and existing PDB templates, eventually modeling the LC amyloid structures similar enough to known templates. HMMSTUFF on one side expands our understanding of LC amyloid fibril conformations, and on the other highlights the gaps in our current knowledge of LC structural space. AVAILABILITY AND IMPLEMENTATION: HMMSTUFF is available as pypi package and as source code at https://github.com/grogdrinker/hmmstuff. Gabriele Orlando, Rodrigo Gallardo, Alicia Colla, Joost Schymkowitz, Frederic Rousseau 0001 |
Bioinform. | 5 |
| 2024 | CORDAX web server: an online platform for the prediction and 3D visualization of aggregation motifs in protein sequencesabstractMOTIVATION: Proteins, the molecular workhorses of biological systems, execute a multitude of critical functions dictated by their precise three-dimensional structures. In a complex and dynamic cellular environment, proteins can undergo misfolding, leading to the formation of aggregates that take up various forms, including amorphous and ordered aggregation in the shape of amyloid fibrils. This phenomenon is closely linked to a spectrum of widespread debilitating pathologies, such as Alzheimer's disease, Parkinson's disease, type-II diabetes, and several other proteinopathies, but also hampers the engineering of soluble agents, as in the case of antibody development. As such, the accurate prediction of aggregation propensity within protein sequences has become pivotal due to profound implications in understanding disease mechanisms, as well as in improving biotechnological and therapeutic applications. RESULTS: We previously developed Cordax, a structure-based predictor that utilizes logistic regression to detect aggregation motifs in protein sequences based on their structural complementarity to the amyloid cross-beta architecture. Here, we present a dedicated web server interface for Cordax. This online platform combines several features including detailed scoring of sequence aggregation propensity, as well as 3D visualization with several customization options for topology models of the structural cores formed by predicted aggregation motifs. In addition, information is provided on experimentally determined aggregation-prone regions that exhibit sequence similarity to predicted motifs, scores, and links to other predictor outputs, as well as simultaneous predictions of relevant sequence propensities, such as solubility, hydrophobicity, and secondary structure propensity. AVAILABILITY AND IMPLEMENTATION: The Cordax webserver is freely accessible at https://cordax.switchlab.org/. Nikolaos N. Louros, Frederic Rousseau 0001, Joost Schymkowitz |
Bioinform. | 2 |
| 2024 | Integrating physics in deep learning algorithms: a force field as a PyTorch moduleabstractMOTIVATION: Deep learning algorithms applied to structural biology often struggle to converge to meaningful solutions when limited data is available, since they are required to learn complex physical rules from examples. State-of-the-art force-fields, however, cannot interface with deep learning algorithms due to their implementation. RESULTS: We present MadraX, a forcefield implemented as a differentiable PyTorch module, able to interact with deep learning algorithms in an end-to-end fashion. AVAILABILITY AND IMPLEMENTATION: MadraX documentation, together with tutorials and installation guide, is available at madrax.readthedocs.io. Gabriele Orlando, Luis Serrano, Joost Schymkowitz, Frederic Rousseau 0001 |
Bioinform. | 4 |
| 2023 | SNPeffect 5.0: large-scale structural phenotyping of protein coding variants extracted from next-generation sequencing data using AlphaFold modelsabstractBACKGROUND: Next-generation sequencing technologies yield large numbers of genetic alterations, of which a subset are missense variants that alter an amino acid in the protein product. These variants can have a potentially destabilizing effect leading to an increased risk of misfolding and aggregation. Multiple software tools exist to predict the effect of single-nucleotide variants on proteins, however, a pipeline integrating these tools while starting from an NGS data output list of variants is lacking. RESULTS: The previous version SNPeffect 4.0 (De Baets in Nucleic Acids Res 40(D1):D935-D939, 2011) provided an online database containing pre-calculated variant effects and low-throughput custom variant analysis. Here, we built an automated and parallelized pipeline that analyzes the impact of missense variants on the aggregation propensity and structural stability of proteins starting from the Variant Call Format as input. The pipeline incorporates the AlphaFold Protein Structure Database to achieve high coverage for structural stability analyses using the FoldX force field. The effect on aggregation-propensity is analyzed using the established predictors TANGO and WALTZ. The pipeline focuses solely on the human proteome and can be used to analyze proteome stability/damage in a given sample based on sequencing results. CONCLUSION: We provide a bioinformatics pipeline that allows structural phenotyping from sequencing data using established stability and aggregation predictors including FoldX, TANGO, and WALTZ; and structural proteome coverage provided by the AlphaFold database. The pipeline and installation guide are freely available for academic users on https://github.com/vibbits/snpeffect and requires a computer cluster. Kobe Janssen, Ramon Duran-Romaña, Guy Bottu, Mainak Guharoy, Alexander Botzki, Frederic Rousseau 0001, Joost Schymkowitz |
BMC Bioinform. | 6 |
| 2022 | StAmP-DB: a platform for structures of polymorphic amyloid fibril coresabstractSUMMARY: Amyloid polymorphism is emerging as a key property that is differentially linked to various conformational diseases, including major neurodegenerative disorders, but also as a feature that potentially relates to complex structural mechanisms mediating transmissibility barriers and selective vulnerability of amyloids. In response to the rapidly expanding number of amyloid fibril structures formed by full-length proteins, we here have developed StAmP-DB, a public database that supports the curation and cross-comparison of experimentally determined three-dimensional amyloid polymorph structures. AVAILABILITY AND IMPLEMENTATION: StAmP-DB is freely accessible for queries and downloads at https://stamp.switchlab.org. Nikolaos N. Louros, Rob van der Kant, Joost Schymkowitz, Frederic Rousseau 0001 |
Bioinform. | 4 |
| 2021 | In silico prediction of in vitro protein liquid-liquid phase separation experiments outcomes with multi-head neural attentionabstractMOTIVATION: Proteins able to undergo liquid-liquid phase separation (LLPS) in vivo and in vitro are drawing a lot of interest, due to their functional relevance for cell life. Nevertheless, the proteome-scale experimental screening of these proteins seems unfeasible, because besides being expensive and time-consuming, LLPS is heavily influenced by multiple environmental conditions such as concentration, pH and temperature, thus requiring a combinatorial number of experiments for each protein. RESULTS: To overcome this problem, we propose a neural network model able to predict the LLPS behavior of proteins given specified experimental conditions, effectively predicting the outcome of in vitro experiments. Our model can be used to rapidly screen proteins and experimental conditions searching for LLPS, thus reducing the search space that needs to be covered experimentally. We experimentally validate Droppler's prediction on the TAR DNA-binding protein in different experimental conditions, showing the consistency of its predictions. AVAILABILITY AND IMPLEMENTATION: A python implementation of Droppler is available at https://bitbucket.org/grogdrinker/droppler. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Daniele Raimondi, Gabriele Orlando, Emiel Michiels, Donya Pakravan, Anna Bratek-Skicki, Ludo Van Den Bosch, Yves Moreau, Frederic Rousseau 0001, Joost Schymkowitz |
Bioinform. | 8 |
| 2020 | Protein Homeostasis Database: protein quality control in E.coliabstractMOTIVATION: In vivo protein folding is governed by molecular chaperones, that escort proteins from their translational birth to their proteolytic degradation. In E.coli the main classes of chaperones that interact with the nascent chain are trigger factor, DnaK/J and GroEL/ES and several authors have performed whole-genome experiments to construct exhaustive client lists for each of these. RESULTS: We constructed a database collecting all publicly available data of experimental chaperone-interaction and -dependency data for the E.coli proteome, and enriched it with an extensive set of protein-specific as well as cell context-dependent proteostatic parameters. We made this publicly accessible via a web interface that allows to search for proteins or chaperone client lists, but also to profile user-specified datasets against all the collected parameters. We hope this will accelerate research in this field by quickly identifying differentiating features in datasets. AVAILABILITY AND IMPLEMENTATION: The Protein Homeostasis Database is freely available without any registration requirement at http://PHDB.switchlab.org/. Reshmi Ramakrishnan, Bert Houben, Lukasz Kreft, Alexander Botzki, Joost Schymkowitz, Frederic Rousseau 0001 |
Bioinform. | 6 |
| 2019 | Differential proteostatic regulation of insoluble and abundant proteinsabstractMOTIVATION: Despite intense effort, it has been difficult to explain chaperone dependencies of proteins from sequence or structural properties. RESULTS: We constructed a database collecting all publicly available data of experimental chaperone interaction and dependency data for the Escherichia coli proteome, and enriched it with an extensive set of protein-specific as well as cell-context-dependent proteostatic parameters. Employing this new resource, we performed a comprehensive meta-analysis of the key determinants of chaperone interaction. Our study confirms that GroEL client proteins are biased toward insoluble proteins of low abundance, but for client proteins of the Trigger Factor/DnaK axis, we instead find that cellular parameters such as high protein abundance, translational efficiency and mRNA turnover are key determinants. We experimentally confirmed the finding that chaperone dependence is a function of translation rate and not protein-intrinsic parameters by tuning chaperone dependence of Green Fluorescent Protein (GFP) in E.coli by synonymous mutations only. The juxtaposition of both protein-intrinsic and cell-contextual chaperone triage mechanisms explains how the E.coli proteome achieves combining reliable production of abundant and conserved proteins, while also enabling the evolution of diverging metabolic functions. AVAILABILITY AND IMPLEMENTATION: The database will be made available via http://phdb.switchlab.org. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Reshmi Ramakrishnan, Bert Houben, Frederic Rousseau 0001, Joost Schymkowitz |
Bioinform. | 3 |
| 2016 | From Binding-Induced Dynamic Effects in SH3 Structures to Evolutionary Conserved SectorsabstractSrc Homology 3 domains are ubiquitous small interaction modules known to act as docking sites and regulatory elements in a wide range of proteins. Prior experimental NMR work on the SH3 domain of Src showed that ligand binding induces long-range dynamic changes consistent with an induced fit mechanism. The identification of the residues that participate in this mechanism produces a chart that allows for the exploration of the regulatory role of such domains in the activity of the encompassing protein. Here we show that a computational approach focusing on the changes in side chain dynamics through ligand binding identifies equivalent long-range effects in the Src SH3 domain. Mutation of a subset of the predicted residues elicits long-range effects on the binding energetics, emphasizing the relevance of these positions in the definition of intramolecular cooperative networks of signal transduction in this domain. We find further support for this mechanism through the analysis of seven other publically available SH3 domain structures of which the sequences represent diverse SH3 classes. By comparing the eight predictions, we find that, in addition to a dynamic pathway that is relatively conserved throughout all SH3 domains, there are dynamic aspects specific to each domain and homologous subgroups. Our work shows for the first time from a structural perspective, which transduction mechanisms are common between a subset of closely related and distal SH3 domains, while at the same time highlighting the differences in signal transduction that make each family member unique. These results resolve the missing link between structural predictions of dynamic changes and the domain sectors recently identified for SH3 domains through sequence analysis. Ana Zafra Ruano, Elisa Cilia, José R. Couceiro, Javier Ruiz Sanz, Joost Schymkowitz, Frederic Rousseau 0001, Irene Luque, Tom Lenaerts |
PLoS Comput. Biol. | 6 |
| 2015 | Solubis: optimize your proteinabstractAbstract Motivation: Protein aggregation is associated with a number of protein misfolding diseases and is a major concern for therapeutic proteins. Aggregation is caused by the presence of aggregation-prone regions (APRs) in the amino acid sequence of the protein. The lower the aggregation propensity of APRs and the better they are protected by native interactions within the folded structure of the protein, the more aggregation is prevented. Therefore, both the local thermodynamic stability of APRs in the native structure and their intrinsic aggregation propensity are a key parameter that needs to be optimized to prevent protein aggregation. Results: The Solubis method presented here automates the process of carefully selecting point mutations that minimize the intrinsic aggregation propensity while improving local protein stability. Availability and implementation: All information about the Solubis plugin is available at http://solubisyasara.switchlab.org/. Contact: [email protected] or [email protected] Supplementary information: Supplementary data are available at Bioinformatics online. Greet De Baets, Joost J. J. van Durme, Rob van der Kant, Joost Schymkowitz, Frederic Rousseau 0001 |
Bioinform. | 5 |
| 2015 | WALTZ-DB: a benchmark database of amyloidogenic hexapeptidesabstractAbstract Summary: Accurate prediction of amyloid-forming amino acid sequences remains an important challenge. We here present an online database that provides open access to the largest set of experimentally characterized amyloid forming hexapeptides. To this end, we expanded our previous set of 280 hexapeptides used to develop the Waltz algorithm with 89 peptides from literature review and by systematic experimental characterisation of the aggregation of 720 hexapeptides by transmission electron microscopy, dye binding and Fourier transform infrared spectroscopy. This brings the total number of experimentally characterized hexapeptides in the WALTZ-DB database to 1089, of which 244 are annotated as positive for amyloid formation. Availability and implementation: The WALTZ-DB database is freely available without any registration requirement at http://waltzdb.switchlab.org. Contact: [email protected] or [email protected] Jacinte Beerten, Joost J. J. van Durme, Rodrigo Gallardo, Emidio Capriotti, Louise C. Serpell, Frederic Rousseau 0001, Joost Schymkowitz |
Bioinform. | 6 |
| 2015 | Increased Aggregation Is More Frequently Associated to Human Disease-Associated Mutations Than to Neutral PolymorphismsabstractProtein aggregation is a hallmark of over 30 human pathologies. In these diseases, the aggregation of one or a few specific proteins is often toxic, leading to cellular degeneration and/or organ disruption in addition to the loss-of-function resulting from protein misfolding. Although the pathophysiological consequences of these diseases are overt, the molecular dysregulations leading to aggregate toxicity are still unclear and appear to be diverse and multifactorial. The molecular mechanisms of protein aggregation and therefore the biophysical parameters favoring protein aggregation are better understood. Here we perform an in silico survey of the impact of human sequence variation on the aggregation propensity of human proteins. We find that disease-associated variations are statistically significantly enriched in mutations that increase the aggregation potential of human proteins when compared to neutral sequence variations. These findings suggest that protein aggregation might have a broader impact on human disease than generally assumed and that beyond loss-of-function, the aggregation of mutant proteins involved in cancer, immune disorders or inflammation could potentially further contribute to disease by additional burden on cellular protein homeostasis. Greet De Baets, Loic Van Doorn, Frederic Rousseau 0001, Joost Schymkowitz |
PLoS Comput. Biol. | 3 |
| 2015 | What Makes a Protein Sequence a Prion?abstractTypical amyloid diseases such as Alzheimer's and Parkinson's were thought to exclusively result from de novo aggregation, but recently it was shown that amyloids formed in one cell can cross-seed aggregation in other cells, following a prion-like mechanism. Despite the large experimental effort devoted to understanding the phenomenon of prion transmissibility, it is still poorly understood how this property is encoded in the primary sequence. In many cases, prion structural conversion is driven by the presence of relatively large glutamine/asparagine (Q/N) enriched segments. Several studies suggest that it is the amino acid composition of these regions rather than their specific sequence that accounts for their priogenicity. However, our analysis indicates that it is instead the presence and potency of specific short amyloid-prone sequences that occur within intrinsically disordered Q/N-rich regions that determine their prion behaviour, modulated by the structural and compositional context. This provides a basis for the accurate identification and evaluation of prion candidate sequences in proteomes in the context of a unified framework for amyloid formation and prion propagation. Raimon Sabate, Frederic Rousseau 0001, Joost Schymkowitz, Salvador Ventura |
PLoS Comput. Biol. | 2 |
| 2011 | A graphical interface for the FoldX forcefieldabstractAbstract Summary: A graphical user interface for the FoldX protein design program has been developed as a plugin for the YASARA molecular graphics suite. The most prominent FoldX commands such as free energy difference upon mutagenesis and interaction energy calculations can now be run entirely via a windowed menu system and the results are immediately shown on screen. Availability and Implementation: The plugin is written in Python and is freely available for download at http://foldxyasara.switchlab.org/ and supported on Linux, MacOSX and MS Windows. Contact: [email protected]; [email protected]; [email protected] Joost J. J. van Durme, Javier Delgado Blanco, François Stricher, Luis Serrano, Joost Schymkowitz, Frederic Rousseau 0001 |
Bioinform. | 6 |
| 2011 | An Evolutionary Trade-Off between Protein Turnover Rate and Protein Aggregation Favors a Higher Aggregation Propensity in Fast Degrading ProteinsabstractWe previously showed the existence of selective pressure against protein aggregation by the enrichment of aggregation-opposing 'gatekeeper' residues at strategic places along the sequence of proteins. Here we analyzed the relationship between protein lifetime and protein aggregation by combining experimentally determined turnover rates, expression data, structural data and chaperone interaction data on a set of more than 500 proteins. We find that selective pressure on protein sequences against aggregation is not homogeneous but that short-living proteins on average have a higher aggregation propensity and fewer chaperone interactions than long-living proteins. We also find that short-living proteins are more often associated to deposition diseases. These findings suggest that the efficient degradation of high-turnover proteins is sufficient to preclude aggregation, but also that factors that inhibit proteasomal activity, such as physiological ageing, will primarily affect the aggregation of short-living proteins. Greet De Baets, Joke Reumers, Javier Delgado Blanco, Joaquín Dopazo, Joost Schymkowitz, Frederic Rousseau 0001 |
PLoS Comput. Biol. | 6 |
| 2010 | Modeling protein-peptide interactions using protein fragments: fitting the pieces?abstractAn estimated 15-40% of all interactions in the cell are mediated through protein-peptide interactions [1,2] meaning that, at the most extreme, nearly every protein is affected either directly or indirectly by peptide-binding events. Peter Vanhee, François Stricher, Lies Baeten, Erik Verschueren, Luis Serrano, Frederic Rousseau 0001, Joost Schymkowitz |
BMC Bioinform. | 6 |
| 2009 | Using structural bioinformatics to investigate the impact of non synonymous SNPs and disease mutations: scope and limitationsabstractBACKGROUND: Linking structural effects of mutations to functional outcomes is a major issue in structural bioinformatics, and many tools and studies have shown that specific structural properties such as stability and residue burial can be used to distinguish neutral variations and disease associated mutations. RESULTS: We have investigated 39 structural properties on a set of SNPs and disease mutations from the Uniprot Knowledge Base that could be mapped on high quality crystal structures and show that none of these properties can be used as a sole classification criterion to separate the two data sets. Furthermore, we have reviewed the annotation process from mutation to result and identified the liabilities in each step. CONCLUSION: Although excellent annotation results of various research groups underline the great potential of using structural bioinformatics to investigate the mechanisms underlying disease, the interpretation of such annotations cannot always be extrapolated to proteome wide variation studies. Difficulties for large-scale studies can be found both on the technical level, i.e. the scarcity of data and the incompleteness of the structural tool suites, and on the conceptual level, i.e. the correct interpretation of the results in a cellular context. Joke Reumers, Joost Schymkowitz, Frederic Rousseau 0001 |
BMC Bioinform. | 3 |
| 2009 | Accurate Prediction of DnaK-Peptide Binding via Homology Modelling and Experimental DataabstractMolecular chaperones are essential elements of the protein quality control machinery that governs translocation and folding of nascent polypeptides, refolding and degradation of misfolded proteins, and activation of a wide range of client proteins. The prokaryotic heat-shock protein DnaK is the E. coli representative of the ubiquitous Hsp70 family, which specializes in the binding of exposed hydrophobic regions in unfolded polypeptides. Accurate prediction of DnaK binding sites in E. coli proteins is an essential prerequisite to understand the precise function of this chaperone and the properties of its substrate proteins. In order to map DnaK binding sites in protein sequences, we have developed an algorithm that combines sequence information from peptide binding experiments and structural parameters from homology modelling. We show that this combination significantly outperforms either single approach. The final predictor had a Matthews correlation coefficient (MCC) of 0.819 when assessed over the 144 tested peptide sequences to detect true positives and true negatives. To test the robustness of the learning set, we have conducted a simulated cross-validation, where we omit sequences from the learning sets and calculate the rate of repredicting them. This resulted in a surprisingly good MCC of 0.703. The algorithm was also able to perform equally well on a blind test set of binders and non-binders, of which there was no prior knowledge in the learning sets. The algorithm is freely available at http://limbo.vib.be. Joost J. J. van Durme, Sebastian Maurer-Stroh, Rodrigo Gallardo, Hannah Wilkinson, Frederic Rousseau 0001, Joost Schymkowitz |
PLoS Comput. Biol. | 5 |
| 2008 | Reconstruction of Protein Backbones from the BriX Collection of Canonical Protein FragmentsabstractAs modeling of changes in backbone conformation still lacks a computationally efficient solution, we developed a discretisation of the conformational states accessible to the protein backbone similar to the successful rotamer approach in side chains. The BriX fragment database, consisting of fragments from 4 to 14 residues long, was realized through identification of recurrent backbone fragments from a non-redundant set of high-resolution protein structures. BriX contains an alphabet of more than 1,000 frequently observed conformations per peptide length for 6 different variation levels. Analysis of the performance of BriX revealed an average structural coverage of protein structures of more than 99% within a root mean square distance (RMSD) of 1 Angstrom. Globally, we are able to reconstruct protein structures with an average accuracy of 0.48 Angstrom RMSD. As expected, regular structures are well covered, but, interestingly, many loop regions that appear irregular at first glance are also found to form a recurrent structural motif, albeit with lower frequency of occurrence than regular secondary structures. Larger loop regions could be completely reconstructed from smaller recurrent elements, between 4 and 8 residues long. Finally, we observed that a significant amount of short sequences tend to display strong structural ambiguity between alpha helix and extended conformations. When the sequence length increases, this so-called sequence plasticity is no longer observed, illustrating the context dependency of polypeptide structures. Lies Baeten, Joke Reumers, Vicente Tur, François Stricher, Tom Lenaerts, Luis Serrano, Frederic Rousseau 0001, Joost Schymkowitz |
PLoS Comput. Biol. | 7 |
| 2008 | Genome-Wide Prediction of SH2 Domain Targets Using Structural Information and the FoldX AlgorithmabstractCurrent experiments likely cover only a fraction of all protein-protein interactions. Here, we developed a method to predict SH2-mediated protein-protein interactions using the structure of SH2-phosphopeptide complexes and the FoldX algorithm. We show that our approach performs similarly to experimentally derived consensus sequences and substitution matrices at predicting known in vitro and in vivo targets of SH2 domains. We use our method to provide a set of high-confidence interactions for human SH2 domains with known structure filtered on secondary structure and phosphorylation state. We validated the predictions using literature-derived SH2 interactions and a probabilistic score obtained from a naive Bayes integration of information on coexpression, conservation of the interaction in other species, shared interaction partners, and functions. We show how our predictions lead to a new hypothesis for the role of SH2 domains in signaling. Ignacio E. Sánchez, Pedro Beltrão, François Stricher, Joost Schymkowitz, Jesper Ferkinghoff-Borg, Frederic Rousseau 0001, Luis Serrano |
PLoS Comput. Biol. | 6 |
| 2006 | SNPeffect v2.0: a new step in investigating the molecular phenotypic effects of human non-synonymous SNPsabstractUNLABELLED: Single nucleotide polymorphisms (SNPs) constitute the most fundamental type of genetic variation in human populations. About 75 000 of these reported variations cause an amino acid change in the translated protein. An important goal in genomic research is to understand how this variability affects protein function, and whether or not particular SNPs are associated to disease susceptibility. Accordingly, the SNPeffect database uses sequence- and structure-based bioinformatics tools to predict the effect of non-synonymous SNPs on the molecular phenotype of proteins. SNPeffect analyses the effect of SNPs on three categories of functional properties: (1) structural and thermodynamic properties affecting protein dynamics and stability (2) the integrity of functional and binding sites and (3) changes in posttranslational processing and cellular localization of proteins. The search interface of the database can be used to search specifically for polymorphisms that are predicted to cause a change in one of these properties. Now based on the Ensembl human databases, the SNPeffect database has been remodeled to better fit an automatically updatable structure. The current edition holds the molecular phenotype of 74 567 nsSNPs in 23 426 proteins. AVAILABILITY: SNPeffect can be accessed through http://snpeffect.vib.be. Joke Reumers, Sebastian Maurer-Stroh, Joost Schymkowitz, Frederic Rousseau 0001 |
Bioinform. | 4 |