EDBT 2026 Demo / reviewers in the wild / expert
Tom L. Blundell
dblp:43/583
· DBLP profile ↗
29ranked-venue papers
0as first author
10since 2021 · last 2024
0000-0002-2708-8992ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 27 · 8 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Evaluating Representation Learning on the Protein Structure UniverseabstractWe introduce ProteinWorkshop, a comprehensive benchmark suite for representation learning on protein structures with Geometric Graph Neural Networks. We consider large-scale pre-training and downstream tasks on both experimental and predicted structures to enable the systematic evaluation of the quality of the learned structural representation and their usefulness in capturing functional relationships for downstream tasks. We find that: (1) large-scale pretraining on AlphaFold structures and auxiliary tasks consistently improve the performance of both rotation-invariant and equivariant GNNs, and (2) more expressive equivariant GNNs benefit from pretraining to a greater extent compared to invariant models.
We aim to establish a common ground for the machine learning and computational biology communities to rigorously compare and advance protein structure representation learning. Our open-source codebase reduces the barrier to entry for working with large protein structure datasets by providing: (1) storage-efficient dataloaders for large-scale structural databases including AlphaFoldDB and ESM Atlas, as well as (2) utilities for constructing new tasks from the entire PDB. ProteinWorkshop is available at: github.com/a-r-j/ProteinWorkshop. Arian Rokkum Jamasb, Alex Morehead, Chaitanya K. Joshi, Zuobai Zhang, Kieran Didi, Simon V. Mathis, Charles Harris, Jian Tang 0005, Jianlin Cheng, Pietro Liò, Tom L. Blundell |
ICLR | 11 |
| 2024 | Melodia: a Python library for protein structure analysisabstractSUMMARY: Analysing protein structure similarities is an important step in protein engineering and drug discovery. Methodologies that are more advanced than simple RMSD are available but often require extensive mathematical or computational knowledge for implementation. Grouping and optimizing such tools in an efficient open-source library increases accessibility and encourages the adoption of more advanced metrics. Melodia is a Python library with a complete set of components devised for describing, comparing and analysing the shape of protein structures using differential geometry of 3D curves and knot theory. It can generate robust geometric descriptors for thousands of shapes in just a few minutes. Those descriptors are more sensitive to structural feature variation than RMSD deviation. Melodia also incorporates sequence structural annotation and 3D visualizations. AVAILABILITY AND IMPLEMENTATION: Melodia is an open-source Python library freely available on https://github.com/rwmontalvao/Melodia_py, along with interactive Jupyter Notebook tutorials. Rinaldo W. Montalvão, William R. Pitt, Vitor Bernardes Pinheiro, Tom L. Blundell |
Bioinform. | 4 |
| 2022 | Graphein - a Python Library for Geometric Deep Learning and Network Analysis on Biomolecular Structures and Interaction NetworksabstractGeometric deep learning has broad applications in biology, a domain where relational structure in data is often intrinsic to modelling the underlying phenomena. Currently, efforts in both geometric deep learning and, more broadly, deep learning applied to biomolecular tasks have been hampered by a scarcity of appropriate datasets accessible to domain specialists and machine learning researchers alike. To address this, we introduce Graphein as a turn-key tool for transforming raw data from widely-used bioinformatics databases into machine learning-ready datasets in a high-throughput and flexible manner. Graphein is a Python library for constructing graph and surface-mesh representations of biomolecular structures, such as proteins, nucleic acids and small molecules, and biological interaction networks for computational analysis and machine learning. Graphein provides utilities for data retrieval from widely-used bioinformatics databases for structural data, including the Protein Data Bank, the AlphaFold Structure Database, chemical data from ZINC and ChEMBL, and for biomolecular interaction networks from STRINGdb, BioGrid, TRRUST and RegNetwork. The library interfaces with popular geometric deep learning libraries: DGL, Jraph, PyTorch Geometric and PyTorch3D though remains framework agnostic as it is built on top of the PyData ecosystem to enable inter-operability with scientific computing tools and libraries. Graphein is designed to be highly flexible, allowing the user to specify each step of the data preparation, scalable to facilitate working with large protein complexes and interaction graphs, and contains useful pre-processing tools for preparing experimental files. Graphein facilitates network-based, graph-theoretic and topological analyses of structural and interaction datasets in a high-throughput manner. We envision that Graphein will facilitate developments in computational biology, graph representation learning and drug discovery. Availability and implementation: Graphein is written in Python. Source code, example usage and tutorials, datasets, and documentation are made freely available under the MIT License at the following URL: https://anonymous.4open.science/r/graphein-3472/README.md Arian Rokkum Jamasb, Ramón Viñas 0001, Eric Ma, Yuanqi Du, Charles Harris, Dominic Hall, Pietro Liò, Tom L. Blundell |
NeurIPS | 9 |
| 2022 | Unheeded SARS-CoV-2 proteins? A deep look into negative-sense RNAabstractSARS-CoV-2 is a novel positive-sense single-stranded RNA virus from the Coronaviridae family (genus Betacoronavirus), which has been established as causing the COVID-19 pandemic. The genome of SARS-CoV-2 is one of the largest among known RNA viruses, comprising of at least 26 known protein-coding loci. Studies thus far have outlined the coding capacity of the positive-sense strand of the SARS-CoV-2 genome, which can be used directly for protein translation. However, it has been recently shown that transcribed negative-sense viral RNA intermediates that arise during viral genome replication from positive-sense viruses can also code for proteins. No studies have yet explored the potential for negative-sense SARS-CoV-2 RNA intermediates to contain protein-coding loci. Thus, using sequence and structure-based bioinformatics methodologies, we have investigated the presence and validity of putative negative-sense ORFs (nsORFs) in the SARS-CoV-2 genome. Nine nsORFs were discovered to contain strong eukaryotic translation initiation signals and high codon adaptability scores, and several of the nsORFs were predicted to interact with RNA-binding proteins. Evolutionary conservation analyses indicated that some of the nsORFs are deeply conserved among related coronaviruses. Three-dimensional protein modeling revealed the presence of higher order folding among all putative SARS-CoV-2 nsORFs, and subsequent structural mimicry analyses suggest similarity of the nsORFs to DNA/RNA-binding proteins and proteins involved in immune signaling pathways. Altogether, these results suggest the potential existence of still undescribed SARS-CoV-2 proteins, which may play an important role in the viral lifecycle and COVID-19 pathogenesis. Martin Bartas, Adriana Volná, Christopher A. Beaudoin, Ebbe Toftgaard Poulsen, Jirí Cerven, Václav Brázda, Vladimír Spunda, Tom L. Blundell, Petr Pecinka |
Briefings Bioinform. | 8 |
| 2022 | Structural landscapes of PPI interfacesabstractProteins are capable of highly specific interactions and are responsible for a wide range of functions, making them attractive in the pursuit of new therapeutic options. Previous studies focusing on overall geometry of protein-protein interfaces, however, concluded that PPI interfaces were generally flat. More recently, this idea has been challenged by their structural and thermodynamic characterisation, suggesting the existence of concave binding sites that are closer in character to traditional small-molecule binding sites, rather than exhibiting complete flatness. Here, we present a large-scale analysis of binding geometry and physicochemical properties of all protein-protein interfaces available in the Protein Data Bank. In this review, we provide a comprehensive overview of the protein-protein interface landscape, including evidence that even for overall larger, more flat interfaces that utilize discontinuous interacting regions, small and potentially druggable pockets are utilized at binding sites. Carlos H. M. Rodrigues, Douglas E. V. Pires, Tom L. Blundell, David B. Ascher |
Briefings Bioinform. | 3 |
| 2021 | SARS-CoV-2 3D database: understanding the coronavirus proteome and evaluating possible drug targetsabstractThe severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) is a rapidly growing infectious disease, widely spread with high mortality rates. Since the release of the SARS-CoV-2 genome sequence in March 2020, there has been an international focus on developing target-based drug discovery, which also requires knowledge of the 3D structure of the proteome. Where there are no experimentally solved structures, our group has created 3D models with coverage of 97.5% and characterized them using state-of-the-art computational approaches. Models of protomers and oligomers, together with predictions of substrate and allosteric binding sites, protein-ligand docking, SARS-CoV-2 protein interactions with human proteins, impacts of mutations, and mapped solved experimental structures are freely available for download. These are implemented in SARS CoV-2 3D, a comprehensive and user-friendly database, available at https://sars3d.com/. This provides essential information for drug discovery, both to evaluate targets and design new potential therapeutics. Ali F. Alsulami, Sherine E. Thomas, Arian Rokkum Jamasb, Christopher A. Beaudoin, Ismail Moghul, Bridget Bannerman, Liviu Copoiu, Sundeep Chaitanya Vedithi, Pedro H. M. Torres, Tom L. Blundell |
Briefings Bioinform. | 10 |
| 2021 | COSMIC Cancer Gene Census 3D database: understanding the impacts of mutations on cancer targetsabstractMutations in hallmark genes are believed to be the main drivers of cancer progression. These mutations are reported in the Catalogue of Somatic Mutations in Cancer (COSMIC). Structural appreciation of where these mutations appear, in protein-protein interfaces, active sites or deoxyribonucleic acid (DNA) interfaces, and predicting the impacts of these mutations using a variety of computational tools are crucial for successful drug discovery and development. Currently, there are 723 genes presented in the COSMIC Cancer Gene Census. Due to the complexity of the gene products, structures of only 87 genes have been solved experimentally with structural coverage between 90% and 100%. Here, we present a comprehensive, user-friendly, web interface (https://cancer-3d.com/) of 714 modelled cancer-related genes, including homo-oligomers, hetero-oligomers, transmembrane proteins and complexes with DNA, ribonucleic acid, ligands and co-factors. Using SDM and mCSM software, we have predicted the impacts of reported mutations on protein stability, protein-protein interfaces affinity and protein-nucleic acid complexes affinity. Furthermore, we also predicted intrinsically disordered regions using DISOPRED3. Ali F. Alsulami, Pedro H. M. Torres, Ismail Moghul, Sheikh Mohammed Arif, Amanda K. Chaplin, Sundeep Chaitanya Vedithi, Tom L. Blundell |
Briefings Bioinform. | 7 |
| 2021 | Utilizing graph machine learning within drug discovery and developmentabstractGraph machine learning (GML) is receiving growing interest within the pharmaceutical and biotechnology industries for its ability to model biomolecular structures, the functional relationships between them, and integrate multi-omic datasets - amongst other data types. Herein, we present a multidisciplinary academic-industrial review of the topic within the context of drug discovery and development. After introducing key terms and modelling approaches, we move chronologically through the drug development pipeline to identify and summarize work incorporating: target identification, design of small molecules and biologics, and drug repurposing. Whilst the field is still emerging, key milestones including repurposed drugs entering in vivo studies, suggest GML will become a modelling framework of choice within biomedical machine learning. Thomas Gaudelet, Ben Day, Arian Rokkum Jamasb, Jyothish Soman, Cristian Regep, Gertrude Liu, Jeremy B. R. Hayter, Richard Vickers, Charles Roberts, Jian Tang 0005, David Roblin, Tom L. Blundell, Michael M. Bronstein, Jake P. Taylor-King |
Briefings Bioinform. | 12 |
| 2021 | ProtCHOIR: a tool for proteome-scale generation of homo-oligomersabstractThe rapid developments in gene sequencing technologies achieved in the recent decades, along with the expansion of knowledge on the three-dimensional structures of proteins, have enabled the construction of proteome-scale databases of protein models such as the Genome3D and ModBase. Nevertheless, although gene products are usually expressed as individual polypeptide chains, most biological processes are associated with either transient or stable oligomerisation. In the PDB databank, for example, ~40% of the deposited structures contain at least one homo-oligomeric interface. Unfortunately, databases of protein models are generally devoid of multimeric structures. To tackle this particular issue, we have developed ProtCHOIR, a tool that is able to generate homo-oligomeric structures in an automated fashion, providing detailed information for the input protein and output complex. ProtCHOIR requires input of either a sequence or a protomeric structure that is queried against a pre-constructed local database of homo-oligomeric structures, then extensively analyzed using well-established tools such as PSI-Blast, MAFFT, PISA and Molprobity. Finally, MODELLER is employed to achieve the construction of the homo-oligomers. The output complex is thoroughly analyzed taking into account its stereochemical quality, interfacial stabilities, hydrophobicity and conservation profile. All these data are then summarized in a user-friendly HTML report that can be saved or printed as a PDF file. The software is easily parallelizable and also outputs a comma-separated file with summary statistics that can straightforwardly be concatenated as a spreadsheet-like document for large-scale data analyses. As a proof-of-concept, we built oligomeric models for the Mabellini Mycobacterium abscessus structural proteome database. ProtCHOIR can be run as a web-service and the code can be obtained free-of-charge at http://lmdm.biof.ufrj.br/protchoir. Pedro H. M. Torres, Artur D. Rossi, Tom L. Blundell |
Briefings Bioinform. | 3 |
| 2021 | A base measure of precision for protein stability predictors: structural sensitivityabstractBACKGROUND: Prediction of the change in fold stability (ΔΔG) of a protein upon mutation is of major importance to protein engineering and screening of disease-causing variants. Many prediction methods can use 3D structural information to predict ΔΔG. While the performance of these methods has been extensively studied, a new problem has arisen due to the abundance of crystal structures: How precise are these methods in terms of structure input used, which structure should be used, and how much does it matter? Thus, there is a need to quantify the structural sensitivity of protein stability prediction methods. RESULTS: We computed the structural sensitivity of six widely-used prediction methods by use of saturated computational mutagenesis on a diverse set of 87 structures of 25 proteins. Our results show that structural sensitivity varies massively and surprisingly falls into two very distinct groups, with methods that take detailed account of the local environment showing a sensitivity of ~ 0.6 to 0.8 kcal/mol, whereas machine-learning methods display much lower sensitivity (~ 0.1 kcal/mol). We also observe that the precision correlates with the accuracy for mutation-type-balanced data sets but not generally reported accuracy of the methods, indicating the importance of mutation-type balance in both contexts. CONCLUSIONS: The structural sensitivity of stability prediction methods varies greatly and is caused mainly by the models and less by the actual protein structural differences. As a new recommended standard, we therefore suggest that ΔΔG values are evaluated on three protein structures when available and the associated standard deviation reported, to emphasize not just the accuracy but also the precision of the method in a specific study. Our observation that machine-learning methods deemphasize structure may indicate that folded wild-type structures alone, without the folded mutant and unfolded structures, only add modest value for assessing protein stability effects, and that side-chain-sensitive methods overstate the significance of the folded wild-type structure. Octav Caldararu, Tom L. Blundell, Kasper P. Kepp |
BMC Bioinform. | 2 |
| 2014 | mCSM: predicting the effects of mutations in proteins using graph-based signaturesabstractMOTIVATION: Mutations play fundamental roles in evolution by introducing diversity into genomes. Missense mutations in structural genes may become either selectively advantageous or disadvantageous to the organism by affecting protein stability and/or interfering with interactions between partners. Thus, the ability to predict the impact of mutations on protein stability and interactions is of significant value, particularly in understanding the effects of Mendelian and somatic mutations on the progression of disease. Here, we propose a novel approach to the study of missense mutations, called mCSM, which relies on graph-based signatures. These encode distance patterns between atoms and are used to represent the protein residue environment and to train predictive models. To understand the roles of mutations in disease, we have evaluated their impacts not only on protein stability but also on protein-protein and protein-nucleic acid interactions. RESULTS: We show that mCSM performs as well as or better than other methods that are used widely. The mCSM signatures were successfully used in different tasks demonstrating that the impact of a mutation can be correlated with the atomic-distance patterns surrounding an amino acid residue. We showed that mCSM can predict stability changes of a wide range of mutations occurring in the tumour suppressor protein p53, demonstrating the applicability of the proposed method in a challenging disease scenario. AVAILABILITY AND IMPLEMENTATION: A web server is available at http://structure.bioc.cam.ac.uk/mcsm. Douglas E. V. Pires, David B. Ascher, Tom L. Blundell |
Bioinform. | 3 |
| 2014 | Polyphony: superposition independent methods for ensemble-based drug discoveryabstractBACKGROUND: Structure-based drug design is an iterative process, following cycles of structural biology, computer-aided design, synthetic chemistry and bioassay. In favorable circumstances, this process can lead to the structures of hundreds of protein-ligand crystal structures. In addition, molecular dynamics simulations are increasingly being used to further explore the conformational landscape of these complexes. Currently, methods capable of the analysis of ensembles of crystal structures and MD trajectories are limited and usually rely upon least squares superposition of coordinates. RESULTS: Novel methodologies are described for the analysis of multiple structures of a protein. Statistical approaches that rely upon residue equivalence, but not superposition, are developed. Tasks that can be performed include the identification of hinge regions, allosteric conformational changes and transient binding sites. The approaches are tested on crystal structures of CDK2 and other CMGC protein kinases and a simulation of p38α. Known interaction - conformational change relationships are highlighted but also new ones are revealed. A transient but druggable allosteric pocket in CDK2 is predicted to occur under the CMGC insert. Furthermore, an evolutionarily-conserved conformational link from the location of this pocket, via the αEF-αF loop, to phosphorylation sites on the activation loop is discovered. CONCLUSIONS: New methodologies are described and validated for the superimposition independent conformational analysis of large collections of structures or simulation snapshots of the same protein. The methodologies are encoded in a Python package called Polyphony, which is released as open source to accompany this paper [http://wrpitt.bitbucket.org/polyphony/]. William R. Pitt, Rinaldo W. Montalvão, Tom L. Blundell |
BMC Bioinform. | 3 |
| 2011 | Comprehensive, atomic-level characterization of structurally characterized protein-protein interactions: the PICCOLO databaseabstractBACKGROUND: Structural studies are increasingly providing huge amounts of information on multi-protein assemblies. Although a complete understanding of cellular processes will be dependent on an explicit characterization of the intermolecular interactions that underlie these assemblies and mediate molecular recognition, these are not well described by standard representations. RESULTS: Here we present PICCOLO, a comprehensive relational database capturing the details of structurally characterized protein-protein interactions. Interactions are described at the level of interacting pairs of atoms, residues and polypeptide chains, with the physico-chemical nature of the interactions being characterized. Distance and angle terms are used to distinguish 12 different interaction types, including van der Waals contacts, hydrogen bonds and hydrophobic contacts. The explicit aim of PICCOLO is to underpin large-scale analyses of the properties of protein-protein interfaces. This is exemplified by an analysis of residue propensity and interface contact preferences derived from a much larger data set than previously reported. However, PICCOLO also supports detailed inspection of particular systems of interest. CONCLUSIONS: The current PICCOLO database comprises more than 260 million interacting atom pairs from 38,202 protein complexes. A web interface for the database is available at http://www-cryst.bioc.cam.ac.uk/piccolo. George R. Bickerton, Alicia P. Higueruelo, Tom L. Blundell |
BMC Bioinform. | 3 |
| 2009 | BIPA: a database for protein-nucleic acid interaction in 3D structuresabstractUNLABELLED: BIPA is a database for protein-nucleic acid interactions in 3D structures. The database provides various physicochemical features of protein-nucleic acid interface such as size, shape, residue propensity, secondary structure composition and intermolecular interactions. The database also contains multiple structural alignments of nucleic acid-binding protein families with annotations of local environments in order to allow definition of features that influence acceptability of mutations at a particular position in a protein family. A web interface has been designed to present the results of these analyses and facilitate navigation of protein-nucleic acid interfaces. AVAILABILITY: http://www-cryst.bioc.cam.ac.uk/bipa SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Semin Lee, Tom L. Blundell |
Bioinform. | 2 |
| 2009 | Ulla: a program for calculating environment-specific amino acid substitution tablesabstractSUMMARY: Amino acid residues are under various kinds of local environmental restraints, which influence substitution patterns. Ulla,(1) a program for calculating environment-specific substitution tables, reads protein sequence alignments and local environment annotations. The program produces a substitution table for every possible combination of environment features. Sparse data is handled using an entropy-based smoothing procedure to estimate robust substitution probabilities. AVAILABILITY: The Ruby source code is available under a Creative Commons Attribution-Noncommercial License along with additional documentation from http://www-cryst.bioc.cam.ac.uk/ulla. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Semin Lee, Tom L. Blundell |
Bioinform. | 2 |
| 2008 | Structural assembly of two-domain proteins by rigid-body dockingabstractBACKGROUND: Modelling proteins with multiple domains is one of the central challenges in Structural Biology. Although homology modelling has successfully been applied for prediction of protein structures, very often domain-domain interactions cannot be inferred from the structures of homologues and their prediction requires ab initio methods. Here we present a new structural prediction approach for modelling two-domain proteins based on rigid-body domain-domain docking. RESULTS: Here we focus on interacting domain pairs that are part of the same peptide chain and thus have an inter-domain peptide region (so called linker). We have developed a method called pyDockTET (tethered-docking), which uses rigid-body docking to generate domain-domain poses that are further scored by binding energy and a pseudo-energy term based on restraints derived from linker end-to-end distances. The method has been benchmarked on a set of 77 non-redundant pairs of domains with available X-ray structure. We have evaluated the docking method ZDOCK, which is able to generate acceptable domain-domain orientations in 51 out of the 77 cases. Among them, our method pyDockTET finds the correct assembly within the top 10 solutions in over 60% of the cases. As a further test, on a subset of 20 pairs where domains were built by homology modelling, ZDOCK generates acceptable orientations in 13 out of the 20 cases, among which the correct assembly is ranked lower than 10 in around 70% of the cases by our pyDockTET method. CONCLUSION: Our results show that rigid-body docking approach plus energy scoring and linker-based restraints are useful for modelling domain-domain interactions. These positive results will encourage development of new methods for structural prediction of macromolecules with multiple (more than two) domains. Tammy M. K. Cheng, Tom L. Blundell, Juan Fernández-Recio |
BMC Bioinform. | 2 |
| 2008 | Prediction by Graph Theoretic Measures of Structural Effects in Proteins Arising from Non-Synonymous Single Nucleotide PolymorphismsabstractRecent analyses of human genome sequences have given rise to impressive advances in identifying non-synonymous single nucleotide polymorphisms (nsSNPs). By contrast, the annotation of nsSNPs and their links to diseases are progressing at a much slower pace. Many of the current approaches to analysing disease-associated nsSNPs use primarily sequence and evolutionary information, while structural information is relatively less exploited. In order to explore the potential of such information, we developed a structure-based approach, Bongo (Bonds ON Graph), to predict structural effects of nsSNPs. Bongo considers protein structures as residue-residue interaction networks and applies graph theoretical measures to identify the residues that are critical for maintaining structural stability by assessing the consequences on the interaction network of single point mutations. Our results show that Bongo is able to identify mutations that cause both local and global structural effects, with a remarkably low false positive rate. Application of the Bongo method to the prediction of 506 disease-associated nsSNPs resulted in a performance (positive predictive value, PPV, 78.5%) similar to that of PolyPhen (PPV, 77.2%) and PANTHER (PPV, 72.2%). As the Bongo method is solely structure-based, our results indicate that the structural changes resulting from nsSNPs are closely associated to their pathological consequences. Tammy M. K. Cheng, Yu-En Lu, Michele Vendruscolo, Pietro Liò, Tom L. Blundell |
PLoS Comput. Biol. | 5 |
| 2008 | Discarding Functional Residues from the Substitution Table Improves Predictions of Active Sites within Three-Dimensional StructuresabstractSubstitutions of individual amino acids in proteins may be under very different evolutionary restraints depending on their structural and functional roles. The Environment Specific Substitution Table (ESST) describes the pattern of substitutions in terms of amino acid location within elements of secondary structure, solvent accessibility, and the existence of hydrogen bonds between side chains and neighbouring amino acid residues. Clearly amino acids that have very different local environments in their functional state compared to those in the protein analysed will give rise to inconsistencies in the calculation of amino acid substitution tables. Here, we describe how the calculation of ESSTs can be improved by discarding the functional residues from the calculation of substitution tables. Four categories of functions are examined in this study: protein-protein interactions, protein-nucleic acid interactions, protein-ligand interactions, and catalytic activity of enzymes. Their contributions to residue conservation are measured and investigated. We test our new ESSTs using the program CRESCENDO, designed to predict functional residues by exploiting knowledge of amino acid substitutions, and compare the benchmark results with proteins whose functions have been defined experimentally. The new methodology increases the Z-score by 98% at the active site residues and finds 16% more active sites compared with the old ESST. We also find that discarding amino acids responsible for protein-protein interactions helps in the prediction of those residues although they are not as conserved as the residues of active sites. Our methodology can make the substitution tables better reflect and describe the substitution patterns of amino acids that are under structural restraints only. Sungsam Gong, Tom L. Blundell |
PLoS Comput. Biol. | 2 |
| 2007 | Andante: reducing side-chain rotamer search space during comparative modeling using environment-specific substitution probabilitiesabstractMOTIVATION: The accurate placement of side chains in computational protein modeling and design involves the searching of vast numbers of rotamer combinations. RESULTS: We have applied the information contained within structurally aligned homologous families, in the form of conserved chi angle conservation rules, to the problem of the comparative modeling. This allows the accurate borrowing of entire side-chain conformations and/or the restriction to high probability rotamer bins. The application of these rules consistently reduces the number of rotamer combinations that need to be searched to trivial values and also reduces the overall side-chain root mean square deviation (rmsd) of the final model. The approach is complementary to current side-chain placement algorithms that use the decomposition of interacting clusters to increase the speed of the placement process. Richard E. Smith, Simon C. Lovell, David F. Burke, Rinaldo W. Montalvão, Tom L. Blundell |
Bioinform. | 5 |
| 2007 | Genome bioinformatic analysis of nonsynonymous SNPsabstractBACKGROUND: Genome-wide association studies of common diseases for common, low penetrance causal variants are underway. A proportion of these will alter protein sequences, the most common of which is the non-synonymous single nucleotide polymorphism (nsSNP). It would be an advantage if the functional effects of an nsSNP on protein structure and function could be predicted, both for the final identification process of a causal variant in a disease-associated chromosome region, and in further functional analyses of the nsSNP and its disease-associated protein. RESULTS: In the present report we have compared and contrasted structure- and sequence-based methods of prediction to over 5500 genes carrying nearly 24,000 nsSNPs, by employing an automatic comparative modelling procedure to build models for the genes. The nsSNP information came from two sources, the OMIM database which are rare (minor allele frequency, MAF, < 0.01) and are known to cause penetrant, monogenic diseases. Secondly, nsSNP information came from dbSNP125, for which the vast majority of nsSNPs, mostly MAF > 0.05, have no known link to a disease. For over 40% of the nsSNPs, structure-based methods predicted which of these sequence changes are likely to either disrupt the structure of the protein or interfere with the function or interactions of the protein. For the remaining 60%, we generated sequence-based predictions. CONCLUSION: We show that, in general, the prediction tools are able distinguish disease causing mutations from those mutations which are thought to have a neutral affect. We give examples of mutations in genes that are predicted to be deleterious and may have a role in disease. Contrary to previous reports, we also show that rare mutations are consistently predicted to be deleterious as often as commonly occurring nsSNPs. David F. Burke, Catherine L. Worth, Eva-Maria Priego, Tammy M. K. Cheng, Luc J. Smink, John A. Todd, Tom L. Blundell |
BMC Bioinform. | 7 |
| 2005 | PROVAT: a tool for Voronoi tessellation analysis of protein structures and complexesabstractSUMMARY: Voronoi tessellation has proved to be a useful tool in protein structure analysis. We have developed PROVAT, a versatile public domain software that enables computation and visualization of Voronoi tessellations of proteins and protein complexes. It is a set of Python scripts that integrate freely available specialized software (Qhull, Pymol etc.) into a pipeline. The calculation component of the tool computes Voronoi tessellation of a given protein system in a way described by a user-supplied XML recipe and stores resulting neighbourhood information as text files with various styles. The Python pickle file generated in the process is used by the visualization component, a Pymol plug-in, that offers a GUI to explore the tessellation visually. AVAILABILITY: PROVAT source code can be downloaded from http://raven.bioc.cam.ac.uk/~swanand/Provat1, which also provides a webserver for its calculation component, documentation and examples. Swanand P. Gore, David F. Burke, Tom L. Blundell |
Bioinform. | 3 |
| 2005 | CHORAL: a differential geometry approach to the prediction of the cores of protein structuresabstractMOTIVATION: Although the cores of homologous proteins are relatively well conserved, amino acid substitutions lead to significant differences in the structures of divergent superfamilies. Thus, the classification of amino acid sequence patterns and the selection of appropriate fragments of the protein cores of homologues of known structure are important for accurate comparative modelling. RESULTS: CHORAL utilizes a knowledge-based method comprising an amalgam of differential geometry and pattern recognition algorithms to identify conserved structural patterns in homologous protein families. Propensity tables are used to classify and to select patterns that most likely represent the structure of the core for a target protein. In our benchmark, CHORAL demonstrates a performance equivalent to that of MODELLER. Rinaldo W. Montalvão, Richard E. Smith, Simon C. Lovell, Tom L. Blundell |
Bioinform. | 4 |
| 2005 | PROVAT - a versatile tool for Voronoi tessellation analysis of protein structures and complexes
Swanand P. Gore, David F. Burke, Tom L. Blundell |
BMC Bioinform. | 3 |
| 2003 | DDBASE2.0: updated domain database with improved identification of structural domainsabstractMOTIVATION: Although many methods are available for the identification of structural domains from protein three-dimensional structures, accurate definition of protein domains and the curation of such data for a large number of proteins are often possible only after manual intervention. The availability of domain definitions for protein structural entries is useful for the sequence analysis of aligned domains, structure comparison, fold recognition procedures and understanding protein folding, domain stability and flexibility. RESULTS: We have improved our method of domain identification starting from the concept of clustering secondary structural elements, but with an intention of reducing the number of discontinuous segments in identified domains. The results of our modified and automatic approach have been compared with the domain definitions from other databases. On a test data set of 55 proteins, this method acquires high agreement (88%) in the number of domains with the crystallographers' definition and resources such as SCOP, CATH, DALI, 3Dee and PDP databases. This method also obtains 98% overlap score with the other resources in the definition of domain boundaries of the 55 proteins. We have examined the domain arrangements of 4592 non-redundant protein chains using the improved method to include 5409 domains leading to an update of the structural domain database. AVAILABILITY: The latest version of the domain database and online domain identification methods are available from http://www.ncbs.res.in/~faculty/mini/ddbase/ddbase.html SUPPLEMENTARY INFORMATION: http://www.ncbs.res.in/~faculty/mini/ddbase/supplementary/supplementary.html A. Vinayagam, Jiye Shi, Ganesan Pugalenthi, B. Meenakshi 0004, Tom L. Blundell, Ramanathan Sowdhamini |
Bioinform. | 5 |
| 2001 | HOMSTRAD: adding sequence information to structure-based alignments of homologous protein familiesabstractsummary: We describe an extension to the Homologous Structure Alignment Database (HOMSTRAD; Mizuguchi et al., Protein Sci., 7, 2469-2471, 1998a) to include homologous sequences derived from the protein families database Pfam (Bateman et al., Nucleic Acids Res., 28, 263-266, 2000). HOMSTRAD is integrated with the server FUGUE (Shi et al., submitted, 2001) for recognition and alignment of homologues, benefitting from the combination of abundant sequence information and accurate structure-based alignments. AVAILABILITY The HOMSTRAD database is available at: http://www-cryst.bioc.cam.ac.uk/homstrad/. Query sequences can be submitted to the homology recognition/alignment server FUGUE at: http://www-cryst.bioc.cam.ac.uk/fugue/. Paul I. W. de Bakker, Alex Bateman, David F. Burke, Ricardo Núñez Miguel, Kenji Mizuguchi, Jiye Shi, Hiroki Shirai, Tom L. Blundell |
Bioinform. | 8 |
| 2001 | SCORE: predicting the core of protein modelsabstractAbstract Motivation: The prediction of the regions of homology models that can be ‘restrained by’ or ‘copied from’ the basis structures is a vital step in correct model generation, because these regions are the models most accurate part. However, there is no ideal method for the identification of their limits. In most algorithms their length depends on the number of family members and definitions of secondary structure. Results: The algorithm SCORE steps away from the conventional definitions of the core to identify from large numbers of basis structures those regions that can be considered structurally related to a target sequence. The use of \batchmode \documentclass[fleqn,10pt,legalpaper]{article} \usepackage{amssymb} \usepackage{amsfonts} \usepackage{amsmath} \pagestyle{empty} \begin{document} \({\phi},\ {\psi}\) \end{document}constraints to accurately pinpoint the regions that are conserved across a family and environmentally constrained substitution tables to extend these regions allows SCORE to rapidly (generally in under 1 s, an order of magnitude faster than methods such as MODELLER) identify and build the core of homology models from the alignments of the target sequence to the basis structures. The SCORE algorithm was used to build 114 model cores. In only two cases was the core size less than 50% of the structure and all the cores built had an RMSD of 3.7 Âor less to the target structure. Availability: The algorithm is available upon request. Contact: [email protected] * To whom correspondence should be addressed. Charlotte M. Deane, Quentin Kaas, Tom L. Blundell |
Bioinform. | 3 |
| 2000 | Browsing the SLoop database of structurally classified loops connecting elements of protein secondary structureabstractWe describe a web server, which provides easy access to the SLoop database of loop conformations connecting elements of protein secondary structure. The loops are classified according to their length, the type of bounding secondary structures and the conformation of the mainchain. The current release of the database consists of over 8000 loops of up to 20 residues in length. A loop prediction method, which selects conformers on the basis of the sequence and the positions of the elements of secondary structure, is also implemented. These web pages are freely accessible over the internet at http://www-cryst.bioc.cam.ac.uk/ approximately sloop. David F. Burke, Charlotte M. Deane, Tom L. Blundell |
Bioinform. | 3 |
| 2000 | Analysis of conservation and substitutions of secondary structure elements within protein superfamiliesabstractAbstract Motivation: Structural alignments of superfamily members often exhibit insertions and deletions of secondary structure elements (SSEs), yet conserved subsets of SSEs appear to be important for maintaining the fold and facilitating common functionalities. Results: A database of aligned SSEs was constructed from the structure-based alignments of protein superfamily members in the CAMPASS database. SSEs were classified into several types on the basis of their length and solvent accessibility and counts were made for the replacements of SSEs in different types at structurally aligned positions. The results, summarized as log-odds substitution matrices, can be used for two types of comparisons: (1) structure against structure, both with secondary structure assignments; and (2) structure against sequence with predicted secondary structures. The conservation of SSEs at each alignment position was defined as the deviation of observed SSE frequencies from the uniform distribution. This offers a useful resource to define and examine the core of superfamily folds. Even when the structure of only a single member of a superfamily is known, the extended method can be used to predict the conservation of SSEs. Such information will be useful when modelling the structure of other members of a superfamily or identifying structurally and functionally important positions in the fold. Availability: The database is available on the world wide web at http://www-cryst.bioc.cam.ac.uk/~kenji/ssdb/. The conservation of SSEs was translated into colour values and can be visualized through the web interface. Contact: [email protected] * To whom correspondence should be addressed. Kenji Mizuguchi, Tom L. Blundell |
Bioinform. | 2 |
| 1998 | JOY: protein sequence-structure representation and analysisabstractMOTIVATION: JOY is a program to annotate protein sequence alignments with three-dimensional (3D) structural features. It was developed to display 3D structural information in a sequence alignment and to help understand the conservation of amino acids in their specific local environments. RESULTS: : The JOY representation now constitutes an essential part of the two databases of protein structure alignments: HOMSTRAD (http://www-cryst.bioc.cam.ac.uk/homstrad ) and CAMPASS (http://www-cryst.bioc.cam.ac. uk/campass). It has also been successfully used for identifying distant evolutionary relationships. AVAILABILITY: The program can be obtained via anonymous ftp from torsa.bioc.cam.ac.uk from the directory /pub/joy/. The address for the JOY server is http://www-cryst.bioc.cam.ac.uk/cgi-bin/joy.cgi. CONTACT: [email protected] Kenji Mizuguchi, Charlotte M. Deane, Tom L. Blundell, Mark S. Johnson, John P. Overington |
Bioinform. | 3 |