Charlotte M. Deane

dblp:82/1876 · DBLP profile ↗
← Back
71ranked-venue papers
1as first author
24since 2021 · last 2025
0000-0003-1388-2252ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 68 · 1 first-author · 21 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021
YearPublicationVenuePosition
2025 Robustly interrogating machine learning-based scoring functions: what are they learning?
abstract
MOTIVATION: Machine learning-based scoring functions (MLBSFs) have been found to exhibit inconsistent performance on different benchmarks and be prone to learning dataset bias. For the field to develop MLBSFs that learn a generalizable understanding of physics, a more rigorous understanding of how they perform is required. RESULTS: In this work, we compared the performance of a diverse set of popular MLBSFs (RFScore, SIGN, OnionNet-2, Pafnucy, and PointVS) to our proposed baseline models that can only learn dataset biases on a range of benchmarks. We found that these baseline models were competitive in accuracy to these MLBSFs in almost all proposed benchmarks, indicating these models only learn dataset biases. Our tests and provided platform, ToolBoxSF, will enable researchers to robustly interrogate MLBSF performance and determine the effect of dataset biases on their predictions. AVAILABILITY AND IMPLEMENTATION: https://github.com/guydurant/toolboxsf.
Guy Durant, Fergus Boyles, Kristian Birchall, Brian D. Marsden, Charlotte M. Deane
Bioinform.5
2025 STCRpy: a software suite for T-cell receptor structure parsing, interaction profiling, and machine learning dataset preparation
abstract
SUMMARY: Computational methods to guide early-stage TCR drug discovery and TCR repertoire informatics currently under-utilize solved and predicted structure data. Here, we streamline use of these data through an open-source python package for high-throughput TCR structure handling and analysis (STCRpy), facilitating analyses such as TCR:peptide-MHC complex orientation calculation/scoring, root-mean-square-distance evaluation, interaction profiling, and machine learning dataset curation. AVAILABILITY AND IMPLEMENTATION: Freely available as a Python package at https://github.com/oxpig/STCRpy.
Nele Quast, Charlotte M. Deane, Matthew I. J. Raybould
Bioinform.2
2024 Kernel-Based Evaluation of Conditional Biological Sequence Models
abstract
We propose a set of kernel-based tools to evaluate the designs and tune the hyperparameters of conditional sequence models, with a focus on problems in computational biology. The backbone of our tools is a new measure of discrepancy between the true conditional distribution and the model's estimate, called the Augmented Conditional Maximum Mean Discrepancy (ACMMD). Provided that the model can be sampled from, the ACMMD can be estimated unbiasedly from data to quantify absolute model fit, integrated within hypothesis tests, and used to evaluate model reliability. We demonstrate the utility of our approach by analyzing a popular protein design model, ProteinMPNN. We are able to reject the hypothesis that ProteinMPNN fits its data for various protein families, and tune the model's temperature hyperparameter to achieve a better fit.
Pierre Glaser, Steffanie Paul, Alissa M. Hummer, Charlotte M. Deane, Debora S. Marks, Alan Nawzad Amin
ICML4
2024 Context-Guided Diffusion for Out-of-Distribution Molecular and Protein Design
abstract
Generative models have the potential to accelerate key steps in the discovery of novel molecular therapeutics and materials. Diffusion models have recently emerged as a powerful approach, excelling at unconditional sample generation and, with data-driven guidance, conditional generation within their training domain. Reliably sampling from high-value regions beyond the training data, however, remains an open challenge---with current methods predominantly focusing on modifying the diffusion process itself. In this paper, we develop context-guided diffusion (CGD), a simple plug-and-play method that leverages unlabeled data and smoothness constraints to improve the out-of-distribution generalization of guided diffusion models. We demonstrate that this approach leads to substantial performance gains across various settings, including continuous, discrete, and graph-structured diffusion processes with applications across drug discovery, materials science, and protein design.
Leo Klarner, Tim G. J. Rudner, Garrett M. Morris, Charlotte M. Deane, Yee Whye Teh
ICML4
2024 Towards the accurate modelling of antibody-antigen complexes from sequence using machine learning and information-driven docking
abstract
MOTIVATION: Antibody-antigen complex modelling is an important step in computational workflows for therapeutic antibody design. While experimentally determined structures of both antibody and the cognate antigen are often not available, recent advances in machine learning-driven protein modelling have enabled accurate prediction of both antibody and antigen structures. Here, we analyse the ability of protein-protein docking tools to use machine learning generated input structures for information-driven docking. RESULTS: In an information-driven scenario, we find that HADDOCK can generate accurate models of antibody-antigen complexes using an ensemble of antibody structures generated by machine learning tools and AlphaFold2 predicted antigen structures. Targeted docking using knowledge of the complementary determining regions on the antibody and some information about the targeted epitope allows the generation of high-quality models of the complex with reduced sampling, resulting in a computationally cheap protocol that outperforms the ZDOCK baseline. AVAILABILITY AND IMPLEMENTATION: The source code of HADDOCK3 is freely available at github.com/haddocking/haddock3. The code to generate and analyse the data is available at github.com/haddocking/ai-antibodies. The full runs, including docking models from all modules of a workflow have been deposited in our lab collection (data.sbgrid.org/labs/32/1139) at the SBGRID data repository.
Marco Giulini, Constantin Schneider, Daniel Cutting, Nikita Desai, Charlotte M. Deane, Alexandre M. J. J. Bonvin
Bioinform.5
2024 ABodyBuilder3: improved and scalable antibody structure predictions
abstract
SUMMARY: In this article, we introduce ABodyBuilder3, an improved and scalable antibody structure prediction model based on ABodyBuilder2. We achieve a new state-of-the-art accuracy in the modelling of CDR loops by leveraging language model embeddings, and show how predicted structures can be further improved through careful relaxation strategies. Finally, we incorporate a predicted Local Distance Difference Test into the model output to allow for a more accurate estimation of uncertainties. AVAILABILITY AND IMPLEMENTATION: The software package is available at https://github.com/Exscientia/ABodyBuilder3 with model weights and data at https://zenodo.org/records/11354577.
Henry Kenlay, Frédéric A. Dreyer, Daniel Cutting, Daniel A. Nissley, Charlotte M. Deane
Bioinform.5
2024 Addressing the antibody germline bias and its effect on language models for improved antibody design
abstract
MOTIVATION: The versatile binding properties of antibodies have made them an extremely important class of biotherapeutics. However, therapeutic antibody development is a complex, expensive, and time-consuming task, with the final antibody needing to not only have strong and specific binding but also be minimally impacted by developability issues. The success of transformer-based language models in protein sequence space and the availability of vast amounts of antibody sequences, has led to the development of many antibody-specific language models to help guide antibody design. Antibody diversity primarily arises from V(D)J recombination, mutations within the CDRs, and/or from a few nongermline mutations outside the CDRs. Consequently, a significant portion of the variable domain of all natural antibody sequences remains germline. This affects the pre-training of antibody-specific language models, where this facet of the sequence data introduces a prevailing bias toward germline residues. This poses a challenge, as mutations away from the germline are often vital for generating specific and potent binding to a target, meaning that language models need be able to suggest key mutations away from germline. RESULTS: In this study, we explore the implications of the germline bias, examining its impact on both general-protein and antibody-specific language models. We develop and train a series of new antibody-specific language models optimized for predicting nongermline residues. We then compare our final model, AbLang-2, with current models and show how it suggests a diverse set of valid mutations with high cumulative probability. AVAILABILITY AND IMPLEMENTATION: AbLang-2 is trained on both unpaired and paired data, and is freely available at https://github.com/oxpig/AbLang2.git.
Tobias Hegelund Olsen, Iain H. Moal, Charlotte M. Deane
Bioinform.3
2024 p-IgGen: a paired antibody generative language model
abstract
SUMMARY: A key challenge in antibody drug discovery is designing novel sequences that are free from developability issues-such as aggregation, polyspecificity, poor expression, or low solubility. Here, we present p-IgGen, a protein language model for paired heavy-light chain antibody generation. The model generates diverse, antibody-like sequences with pairing properties found in natural antibodies. We also create a finetuned version of p-IgGen that biases the model to generate antibodies with 3D biophysical properties that fall within distributions seen in clinical-stage therapeutic antibodies. AVAILABILITY AND IMPLEMENTATION: The model and inference code are freely available at www.github.com/oxpig/p-IgGen. Cleaned training data are deposited at doi.org/10.5281/zenodo.13880874.
Oliver M. Turnbull, Dino Oglic, Rebecca Croasdale-Wood, Charlotte M. Deane
Bioinform.4
2024 It is theoretically possible to avoid misfolding into non-covalent lasso entanglements using small molecule drugs
abstract
A novel class of protein misfolding characterized by either the formation of non-native noncovalent lasso entanglements in the misfolded structure or loss of native entanglements has been predicted to exist and found circumstantial support through biochemical assays and limited-proteolysis mass spectrometry data. Here, we examine whether it is possible to design small molecule compounds that can bind to specific folding intermediates and thereby avoid these misfolded states in computer simulations under idealized conditions (perfect drug-binding specificity, zero promiscuity, and a smooth energy landscape). Studying two proteins, type III chloramphenicol acetyltransferase (CAT-III) and D-alanyl-D-alanine ligase B (DDLB), that were previously suggested to form soluble misfolded states through a mechanism involving a failure-to-form of native entanglements, we explore two different drug design strategies using coarse-grained structure-based models. The first strategy, in which the native entanglement is stabilized by drug binding, failed to decrease misfolding because it formed an alternative entanglement at a nearby region. The second strategy, in which a small molecule was designed to bind to a non-native tertiary structure and thereby destabilize the native entanglement, succeeded in decreasing misfolding and increasing the native state population. This strategy worked because destabilizing the entanglement loop provided more time for the threading segment to position itself correctly to be wrapped by the loop to form the native entanglement. Further, we computationally identified several FDA-approved drugs with the potential to bind these intermediate states and rescue misfolding in these proteins. This study suggests it is possible for small molecule drugs to prevent protein misfolding of this type.
Yang Jiang 0004, Charlotte M. Deane, Garrett M. Morris, Edward P. O'Brien
PLoS Comput. Biol.2
2024 Large scale paired antibody language models
abstract
Antibodies are proteins produced by the immune system that can identify and neutralise a wide variety of antigens with high specificity and affinity, and constitute the most successful class of biotherapeutics. With the advent of next-generation sequencing, billions of antibody sequences have been collected in recent years, though their application in the design of better therapeutics has been constrained by the sheer volume and complexity of the data. To address this challenge, we present IgBert and IgT5, the best performing antibody-specific language models developed to date which can consistently handle both paired and unpaired variable region sequences as input. These models are trained comprehensively using the more than two billion unpaired sequences and two million paired sequences of light and heavy chains present in the Observed Antibody Space dataset. We show that our models outperform existing antibody and protein language models on a diverse range of design and regression tasks relevant to antibody engineering. This advancement marks a significant leap forward in leveraging machine learning, large scale data sets and high-performance computing for enhancing antibody design for therapeutic development.
Henry Kenlay, Frédéric A. Dreyer, Aleksandr Kovaltsuk, Dom Miketa, Douglas E. V. Pires, Charlotte M. Deane
PLoS Comput. Biol.6
2023 Drug Discovery under Covariate Shift with Domain-Informed Prior Distributions over Functions
abstract
Accelerating the discovery of novel and more effective therapeutics is an important pharmaceutical problem in which deep learning is playing an increasingly significant role. However, real-world drug discovery tasks are often characterized by a scarcity of labeled data and significant covariate shift---a setting that poses a challenge to standard deep learning methods. In this paper, we present Q-SAVI, a probabilistic model able to address these challenges by encoding explicit prior knowledge of the data-generating process into a prior distribution over functions, presenting researchers with a transparent and probabilistically principled way to encode data-driven modeling preferences. Building on a novel, gold-standard bioactivity dataset that facilitates a meaningful comparison of models in an extrapolative regime, we explore different approaches to induce data shift and construct a challenging evaluation setup. We then demonstrate that using Q-SAVI to integrate contextualized prior knowledge of drug-like chemical space into the modeling process affords substantial gains in predictive accuracy and calibration, outperforming a broad range of state-of-the-art self-supervised pre-training and domain adaptation techniques.
Leo Klarner, Tim G. J. Rudner, Michael Reutlinger, Torsten Schindler, Garrett M. Morris, Charlotte M. Deane, Yee Whye Teh
ICML6
2023 Paragraph - antibody paratope prediction using graph neural networks with minimal feature vectors
abstract
SUMMARY: The development of new vaccines and antibody therapeutics typically takes several years and requires over $1bn in investment. Accurate knowledge of the paratope (antibody binding site) can speed up and reduce the cost of this process by improving our understanding of antibody-antigen binding. We present Paragraph, a structure-based paratope prediction tool that outperforms current state-of-the-art tools using simpler feature vectors and no antigen information. AVAILABILITY AND IMPLEMENTATION: Source code is freely available at www.github.com/oxpig/Paragraph. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Lewis Chinery, Newton Wahome, Iain H. Moal, Charlotte M. Deane
Bioinform.4
2022 ABlooper: fast accurate antibody CDR loop structure prediction with accuracy estimation
abstract
MOTIVATION: Antibodies are a key component of the immune system and have been extensively used as biotherapeutics. Accurate knowledge of their structure is central to understanding their antigen-binding function. The key area for antigen binding and the main area of structural variation in antibodies are concentrated in the six complementarity determining regions (CDRs), with the most important for binding and most variable being the CDR-H3 loop. The sequence and structural variability of CDR-H3 make it particularly challenging to model. Recently deep learning methods have offered a step change in our ability to predict protein structures. RESULTS: In this work, we present ABlooper, an end-to-end equivariant deep learning-based CDR loop structure prediction tool. ABlooper rapidly predicts the structure of CDR loops with high accuracy and provides a confidence estimate for each of its predictions. On the models of the Rosetta Antibody Benchmark, ABlooper makes predictions with an average CDR-H3 RMSD of 2.49 Å, which drops to 2.05 Å when considering only its 75% most confident predictions. AVAILABILITY AND IMPLEMENTATION: https://github.com/oxpig/ABlooper. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Brennan Abanades, Guy Georges, Alexander Bujotzek, Charlotte M. Deane
Bioinform.4
2022 Current structure predictors are not learning the physics of protein folding
abstract
SUMMARY: Motivation. Predicting the native state of a protein has long been considered a gateway problem for understanding protein folding. Recent advances in structural modeling driven by deep learning have achieved unprecedented success at predicting a protein's crystal structure, but it is not clear if these models are learning the physics of how proteins dynamically fold into their equilibrium structure or are just accurate knowledge-based predictors of the final state. Results. In this work, we compare the pathways generated by state-of-the-art protein structure prediction methods to experimental data about protein folding pathways. The methods considered were AlphaFold 2, RoseTTAFold, trRosetta, RaptorX, DMPfold, EVfold, SAINT2 and Rosetta. We find evidence that their simulated dynamics capture some information about the folding pathway, but their predictive ability is worse than a trivial classifier using sequence-agnostic features like chain length. The folding trajectories produced are also uncorrelated with experimental observables such as intermediate structures and the folding rate constant. These results suggest that recent advances in structure prediction do not yet provide an enhanced understanding of protein folding. Availability. The data underlying this article are available in GitHub at https://github.com/oxpig/structure-vs-folding/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Carlos Outeiral, Daniel A. Nissley, Charlotte M. Deane
Bioinform.3
2022 DLAB: deep learning methods for structure-based virtual screening of antibodies
abstract
MOTIVATION: Antibodies are one of the most important classes of pharmaceuticals, with over 80 approved molecules currently in use against a wide variety of diseases. The drug discovery process for antibody therapeutic candidates however is time- and cost-intensive and heavily reliant on in vivo and in vitro high throughput screens. Here, we introduce a framework for structure-based deep learning for antibodies (DLAB) which can virtually screen putative binding antibodies against antigen targets of interest. DLAB is built to be able to predict antibody-antigen binding for antigens with no known antibody binders. RESULTS: We demonstrate that DLAB can be used both to improve antibody-antigen docking and structure-based virtual screening of antibody drug candidates. DLAB enables improved pose ranking for antibody docking experiments as well as selection of antibody-antigen pairings for which accurate poses are generated and correctly ranked. We also show that DLAB can identify binding antibodies against specific antigens in a case study. Our results demonstrate the promise of deep learning methods for structure-based virtual screening of antibodies. AVAILABILITY AND IMPLEMENTATION: The DLAB source code and pre-trained models are available at https://github.com/oxpig/dlab-public. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Constantin Schneider, Andrew Buchanan, Bruck Taddese, Charlotte M. Deane
Bioinform.4
2021 COGENT: evaluating the consistency of gene co-expression networks
abstract
SUMMARY: Gene co-expression networks can be constructed in multiple different ways, both in the use of different measures of co-expression, and in the thresholds applied to the calculated co-expression values, from any given dataset. It is often not clear which co-expression network construction method should be preferred. COGENT provides a set of tools designed to aid the choice of network construction method without the need for any external validation data. AVAILABILITY AND IMPLEMENTATION: https://github.com/lbozhilova/COGENT. SUPPLEMENTARY INFORMATION: Supplementary information is available at Bioinformatics online.
Lyuba V. Bozhilova, Javier Pardo-Diaz, Gesine Reinert, Charlotte M. Deane
Bioinform.4
2021 Generating property-matched decoy molecules using deep learning
abstract
MOTIVATION: An essential step in the development of virtual screening methods is the use of established sets of actives and decoys for benchmarking and training. However, the decoy molecules in commonly used sets are biased meaning that methods often exploit these biases to separate actives and decoys, and do not necessarily learn to perform molecular recognition. This fundamental issue prevents generalization and hinders virtual screening method development. RESULTS: We have developed a deep learning method (DeepCoy) that generates decoys to a user's preferred specification in order to remove such biases or construct sets with a defined bias. We validated DeepCoy using two established benchmarks, DUD-E and DEKOIS 2.0. For all 102 DUD-E targets and 80 of the 81 DEKOIS 2.0 targets, our generated decoy molecules more closely matched the active molecules' physicochemical properties while introducing no discernible additional risk of false negatives. The DeepCoy decoys improved the Deviation from Optimal Embedding (DOE) score by an average of 81% and 66%, respectively, decreasing from 0.166 to 0.032 for DUD-E and from 0.109 to 0.038 for DEKOIS 2.0. Further, the generated decoys are harder to distinguish than the original decoy molecules via docking with Autodock Vina, with virtual screening performance falling from an AUC ROC of 0.70 to 0.63. AVAILABILITY AND IMPLEMENTATION: The code is available at https://github.com/oxpig/DeepCoy. Generated molecules can be downloaded from http://opig.stats.ox.ac.uk/resources. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Fergus Imrie, Anthony R. Bradley, Charlotte M. Deane
Bioinform.3
2021 Humanization of antibodies using a machine learning approach on large-scale repertoire data
abstract
MOTIVATION: Monoclonal antibody (mAb) therapeutics are often produced from non-human sources (typically murine), and can therefore generate immunogenic responses in humans. Humanization procedures aim to produce antibody therapeutics that do not elicit an immune response and are safe for human use, without impacting efficacy. Humanization is normally carried out in a largely trial-and-error experimental process. We have built machine learning classifiers that can discriminate between human and non-human antibody variable domain sequences using the large amount of repertoire data now available. RESULTS: Our classifiers consistently outperform the current best-in-class model for distinguishing human from murine sequences, and our output scores exhibit a negative relationship with the experimental immunogenicity of existing antibody therapeutics. We used our classifiers to develop a novel, computational humanization tool, Hu-mAb, that suggests mutations to an input sequence to reduce its immunogenicity. For a set of therapeutic antibodies with known precursor sequences, the mutations suggested by Hu-mAb show substantial overlap with those deduced experimentally. Hu-mAb is therefore an effective replacement for trial-and-error humanization experiments, producing similar results in a fraction of the time. AVAILABILITY AND IMPLEMENTATION: Hu-mAb (humanness scoring and humanization) is freely available to use at opig.stats.ox.ac.uk/webapps/humab. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Claire Marks, Alissa M. Hummer, Mark Chin, Charlotte M. Deane
Bioinform.4
2021 Ribosome occupancy profiles are conserved between structurally and evolutionarily related yeast domains
abstract
MOTIVATION: Protein synthesis is a non-equilibrium process, meaning that the speed of translation can influence the ability of proteins to fold and function. Assuming that structurally similar proteins fold by similar pathways, the profile of translation speed along an mRNA should be evolutionarily conserved between related proteins to direct correct folding and downstream function. The only evidence to date for such conservation of translation speed between homologous proteins has used codon rarity as a proxy for translation speed. There are, however, many other factors including mRNA structure and the chemistry of the amino acids in the A- and P-sites of the ribosome that influence the speed of amino acid addition. RESULTS: Ribosome profiling experiments provide a signal directly proportional to the underlying translation times at the level of individual codons. We compared ribosome occupancy profiles (extracted from five different large-scale yeast ribosome profiling studies) between related protein domains to more directly test if their translation schedule was conserved. Our analysis reveals that the ribosome occupancy profiles of paralogous domains tend to be significantly more similar to one another than to profiles of non-paralogous domains. This trend does not depend on domain length, structural classes, amino acid composition or sequence similarity. Our results indicate that entire ribosome occupancy profiles and not just rare codon locations are conserved between even distantly related domains in yeast, providing support for the hypothesis that translation schedule is conserved between structurally related domains to retain folding pathways and facilitate efficient folding. AVAILABILITY AND IMPLEMENTATION: Python3 code is available on GitHub at https://github.com/DanNissley/Compare-ribosome-occupancy. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Daniel A. Nissley, Anna Carbery, Mark Chonofsky, Charlotte M. Deane
Bioinform.4
2021 Robust gene coexpression networks using signed distance correlation
abstract
MOTIVATION: Even within well studied organisms, many genes lack useful functional annotations. One way to generate such functional information is to infer biological relationships between genes/proteins, using a network of gene coexpression data that includes functional annotations. However, the lack of trustworthy functional annotations can impede the validation of such networks. Hence, there is a need for a principled method to construct gene coexpression networks that capture biological information and are structurally stable even in the absence of functional information. RESULTS: We introduce the concept of signed distance correlation as a measure of dependency between two variables, and apply it to generate gene coexpression networks. Distance correlation offers a more intuitive approach to network construction than commonly used methods such as Pearson correlation and mutual information. We propose a framework to generate self-consistent networks using signed distance correlation purely from gene expression data, with no additional information. We analyse data from three different organisms to illustrate how networks generated with our method are more stable and capture more biological information compared to networks obtained from Pearson correlation or mutual information. SUPPLEMENTARY INFORMATION: Supplementary Information and code are available at Bioinformatics and https://github.com/javier-pardodiaz/sdcorGCN online.
Javier Pardo-Diaz, Lyuba V. Bozhilova, Mariano Beguerisse-Díaz, Philip S. Poole, Charlotte M. Deane, Gesine Reinert
Bioinform.5
2021 CoV-AbDab: the coronavirus antibody database
abstract
MOTIVATION: The emergence of a novel strain of betacoronavirus, SARS-CoV-2, has led to a pandemic that has been associated with over 700 000 deaths as of August 5, 2020. Research is ongoing around the world to create vaccines and therapies to minimize rates of disease spread and mortality. Crucial to these efforts are molecular characterizations of neutralizing antibodies to SARS-CoV-2. Such antibodies would be valuable for measuring vaccine efficacy, diagnosing exposure and developing effective biotherapeutics. Here, we describe our new database, CoV-AbDab, which already contains data on over 1400 published/patented antibodies and nanobodies known to bind to at least one betacoronavirus. This database is the first consolidation of antibodies known to bind SARS-CoV-2 as well as other betacoronaviruses such as SARS-CoV-1 and MERS-CoV. It contains relevant metadata including evidence of cross-neutralization, antibody/nanobody origin, full variable domain sequence (where available) and germline assignments, epitope region, links to relevant PDB entries, homology models and source literature. RESULTS: On August 5, 2020, CoV-AbDab referenced sequence information on 1402 anti-coronavirus antibodies and nanobodies, spanning 66 papers and 21 patents. Of these, 1131 bind to SARS-CoV-2. AVAILABILITYAND IMPLEMENTATION: CoV-AbDab is free to access and download without registration at http://opig.stats.ox.ac.uk/webapps/coronavirus. Community submissions are encouraged. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Matthew I. J. Raybould, Aleksandr Kovaltsuk, Claire Marks, Charlotte M. Deane
Bioinform.4
2021 Co-evolutionary distance predictions contain flexibility information
abstract
MOTIVATION: Co-evolution analysis can be used to accurately predict residue-residue contacts from multiple sequence alignments. The introduction of machine-learning techniques has enabled substantial improvements in precision and a shift from predicting binary contacts to predict distances between pairs of residues. These developments have significantly improved the accuracy of de novo prediction of static protein structures. With AlphaFold2 lifting the accuracy of some predicted protein models close to experimental levels, structure prediction research will move on to other challenges. One of those areas is the prediction of more than one conformation of a protein. Here, we examine the potential of residue-residue distance predictions to be informative of protein flexibility rather than simply static structure. RESULTS: We used DMPfold to predict distance distributions for every residue pair in a set of proteins that showed both rigid and flexible behaviour. Residue pairs that were in contact in at least one reference structure were classified as rigid, flexible or neither. The predicted distance distribution of each residue pair was analysed for local maxima of probability indicating the most likely distance or distances between a pair of residues. We found that rigid residue pairs tended to have only a single local maximum in their predicted distance distributions while flexible residue pairs more often had multiple local maxima. These results suggest that the shape of predicted distance distributions contains information on the rigidity or flexibility of a protein and its constituent residues. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Dominik Schwarz, Guy Georges, Sebastian Kelm, Jiye Shi, Anna Vangone, Charlotte M. Deane
Bioinform.6
2021 Public Baseline and shared response structures support the theory of antibody repertoire functional commonality
abstract
The naïve antibody/B-cell receptor (BCR) repertoires of different individuals ought to exhibit significant functional commonality, given that most pathogens trigger an effective antibody response to immunodominant epitopes. Sequence-based repertoire analysis has so far offered little evidence for this phenomenon. For example, a recent study estimated the number of shared ('public') antibody clonotypes in circulating baseline repertoires to be around 0.02% across ten unrelated individuals. However, to engage the same epitope, antibodies only require a similar binding site structure and the presence of key paratope interactions, which can occur even when their sequences are dissimilar. Here, we search for evidence of geometric similarity/convergence across human antibody repertoires. We first structurally profile naïve ('baseline') antibody diversity using snapshots from 41 unrelated individuals, predicting all modellable distinct structures within each repertoire. This analysis uncovers a high (much greater than random) degree of structural commonality. For instance, around 3% of distinct structures are common to the ten most diverse individual samples ('Public Baseline' structures). Our approach is the first computational method to find levels of BCR commonality commensurate with epitope immunodominance and could therefore be harnessed to find more genetically distant antibodies with same-epitope complementarity. We then apply the same structural profiling approach to repertoire snapshots from three individuals before and after flu vaccination, detecting a convergent structural drift indicative of recognising similar epitopes ('Public Response' structures). We show that Antibody Model Libraries derived from Public Baseline and Public Response structures represent a powerful geometric basis set of low-immunogenicity candidates exploitable for general or target-focused therapeutic antibody screening.
Matthew I. J. Raybould, Claire Marks, Aleksandr Kovaltsuk, Alan P. Lewis, Jiye Shi, Charlotte M. Deane
PLoS Comput. Biol.6
2021 Epitope profiling using computational structural modelling demonstrated on coronavirus-binding antibodies
abstract
Identifying the epitope of an antibody is a key step in understanding its function and its potential as a therapeutic. Sequence-based clonal clustering can identify antibodies with similar epitope complementarity, however, antibodies from markedly different lineages but with similar structures can engage the same epitope. We describe a novel computational method for epitope profiling based on structural modelling and clustering. Using the method, we demonstrate that sequence dissimilar but functionally similar antibodies can be found across the Coronavirus Antibody Database, with high accuracy (92% of antibodies in multiple-occupancy structural clusters bind to consistent domains). Our approach functionally links antibodies with distinct genetic lineages, species origins, and coronavirus specificities. This indicates greater convergence exists in the immune responses to coronaviruses than is suggested by sequence-based approaches. Our results show that applying structural analytics to large class-specific antibody databases will enable high confidence structure-function relationships to be drawn, yielding new opportunities to identify functional convergence hitherto missed by sequence-only analysis.
Sarah A. Robinson, Matthew I. J. Raybould, Constantin Schneider, Wing Ki Wong, Claire Marks, Charlotte M. Deane
PLoS Comput. Biol.6
2020 Learning from the ligand: using ligand-based features to improve binding affinity prediction
abstract
MOTIVATION: Machine learning scoring functions for protein-ligand binding affinity prediction have been found to consistently outperform classical scoring functions. Structure-based scoring functions for universal affinity prediction typically use features describing interactions derived from the protein-ligand complex, with limited information about the chemical or topological properties of the ligand itself. RESULTS: We demonstrate that the performance of machine learning scoring functions are consistently improved by the inclusion of diverse ligand-based features. For example, a Random Forest (RF) combining the features of RF-Score v3 with RDKit molecular descriptors achieved Pearson correlation coefficients of up to 0.836, 0.780 and 0.821 on the PDBbind 2007, 2013 and 2016 core sets, respectively, compared to 0.790, 0.746 and 0.814 when using the features of RF-Score v3 alone. Excluding proteins and/or ligands that are similar to those in the test sets from the training set has a significant effect on scoring function performance, but does not remove the predictive power of ligand-based features. Furthermore a RF using only ligand-based features is predictive at a level similar to classical scoring functions and it appears to be predicting the mean binding affinity of a ligand for its protein targets. AVAILABILITY AND IMPLEMENTATION: Data and code to reproduce all the results are freely available at http://opig.stats.ox.ac.uk/resources. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Fergus Boyles, Charlotte M. Deane, Garrett M. Morris
Bioinform.2
2020 The evolution of contact prediction: evidence that contact selection in statistical contact prediction is changing
abstract
MOTIVATION: Over the last few years, the field of protein structure prediction has been transformed by increasingly accurate contact prediction software. These methods are based on the detection of coevolutionary relationships between residues from multiple sequence alignments (MSAs). However, despite speculation, there is little evidence of a link between contact prediction and the physico-chemical interactions which drive amino-acid coevolution. Furthermore, existing protocols predict only a fraction of all protein contacts and it is not clear why some contacts are favoured over others. Using a dataset of 863 protein domains, we assessed the physico-chemical interactions of contacts predicted by CCMpred, MetaPSICOV and DNCON2, as examples of direct coupling analysis, meta-prediction and deep learning. RESULTS: We considered correctly predicted contacts and compared their properties against the protein contacts that were not predicted. Predicted contacts tend to form more bonds than non-predicted contacts, which suggests these contacts may be more important than contacts that were not predicted. Comparing the contacts predicted by each method, we found that metaPSICOV and DNCON2 favour accuracy, whereas CCMPred detects contacts with more bonds. This suggests that the push for higher accuracy may lead to a loss of physico-chemically important contacts. These results underscore the connection between protein physico-chemistry and the coevolutionary couplings that can be derived from MSAs. This relationship is likely to be relevant to protein structure prediction and functional analysis of protein structure and may be key to understanding their utility for different problems in structural biology. AVAILABILITY AND IMPLEMENTATION: We use publicly available databases. Our code is available for download at https://opig.stats.ox.ac.uk/. SUPPLEMENTARY INFORMATION: Supplementary information is available at Bioinformatics online.
Mark Chonofsky, Saulo Henrique Pires de Oliveira, Konrad Krawczyk, Charlotte M. Deane
Bioinform.4
2020 TCRBuilder: multi-state T-cell receptor structure prediction
abstract
MOTIVATION: T-cell receptors (TCRs) are immune proteins that primarily target peptide antigens presented by the major histocompatibility complex. They tend to have lower specificity and affinity than their antibody counterparts, and their binding sites have been shown to adopt multiple conformations, which is potentially an important factor for their polyspecificity. None of the current TCR-modelling tools predict this variability which limits our ability to accurately predict TCR binding. RESULTS: We present TCRBuilder, a multi-state TCR structure prediction tool. Given a paired αβTCR sequence, TCRBuilder returns a model or an ensemble of models covering the potential conformations of the binding site. This enables the analysis of structurally driven polyspecificity in TCRs, which is not possible with existing tools. AVAILABILITY AND IMPLEMENTATION: http://opig.stats.ox.ac.uk/resources. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Wing Ki Wong, Claire Marks, Jinwoo Leem, Alan P. Lewis, Jiye Shi, Charlotte M. Deane
Bioinform.6
2020 Structural diversity of B-cell receptor repertoires along the B-cell differentiation axis in humans and mice
abstract
Most current analysis tools for antibody next-generation sequencing data work with primary sequence descriptors, leaving accompanying structural information unharnessed. We have used novel rapid methods to structurally characterize the complementary-determining regions (CDRs) of more than 180 million human and mouse B-cell receptor (BCR) repertoire sequences. These structurally annotated CDRs provide unprecedented insights into both the structural predetermination and dynamics of the adaptive immune response. We show that B-cell types can be distinguished based solely on these structural properties. Antigen-unexperienced BCR repertoires use the highest number and diversity of CDR structures and these patterns of naïve repertoire paratope usage are highly conserved across subjects. In contrast, more differentiated B-cells are more personalized in terms of CDR structure usage. Our results establish the CDR structure differences in BCR repertoires and have applications for many fields including immunodiagnostics, phage display library generation, and "humanness" assessment of BCR repertoires from transgenic animals. The software tool for structural annotation of BCR repertoires, SAAB+, is available at https://github.com/oxpig/saab_plus.
Aleksandr Kovaltsuk, Matthew I. J. Raybould, Wing Ki Wong, Claire Marks, Sebastian Kelm, James Snowden, Johannes Trück, Charlotte M. Deane
PLoS Comput. Biol.8
2019 Increasing the accuracy of protein loop structure prediction with evolutionary constraints
abstract
MOTIVATION: Accurate prediction of loop structures remains challenging. This is especially true for long loops where the large conformational space and limited coverage of experimentally determined structures often leads to low accuracy. Co-evolutionary contact predictors, which provide information about the proximity of pairs of residues, have been used to improve whole-protein models generated through de novo techniques. Here we investigate whether these evolutionary constraints can enhance the prediction of long loop structures. RESULTS: As a first stage, we assess the accuracy of predicted contacts that involve loop regions. We find that these are less accurate than contacts in general. We also observe that some incorrectly predicted contacts can be identified as they are never satisfied in any of our generated loop conformations. We examined two different strategies for incorporating contacts, and on a test set of long loops (10 residues or more), both approaches improve the accuracy of prediction. For a set of 135 loops, contacts were predicted and hence our methods were applicable in 97 cases. Both strategies result in an increase in the proportion of near-native decoys in the ensemble, leading to more accurate predictions and in some cases improving the root-mean-square deviation of the final model by more than 3 Å. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Claire Marks, Charlotte M. Deane
Bioinform.2
2019 SCALOP: sequence-based antibody canonical loop structure annotation
abstract
MOTIVATION: Canonical forms of the antibody complementarity-determining regions (CDRs) were first described in 1987 and have been redefined on multiple occasions since. The canonical forms are often used to approximate the antibody binding site shape as they can be predicted from sequence. A rapid predictor would facilitate the annotation of CDR structures in the large amounts of repertoire data now becoming available from next generation sequencing experiments. RESULTS: SCALOP annotates CDR canonical forms for antibody sequences, supported by an auto-updating database to capture the latest cluster information. Its accuracy is comparable to that of a standard structural predictor but it is 800 times faster. The auto-updating nature of SCALOP ensures that it always attains the best possible coverage. AVAILABILITY AND IMPLEMENTATION: SCALOP is available as a web application and for download under a GPLv3 license at opig.stats.ox.ac.uk/webapps/scalop. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Wing Ki Wong, Guy Georges, Francesca Ros, Sebastian Kelm, Alan P. Lewis, Bruck Taddese, Jinwoo Leem, Charlotte M. Deane
Bioinform.8
2019 Measuring rank robustness in scored protein interaction networks
abstract
BACKGROUND: Protein interaction databases often provide confidence scores for each recorded interaction based on the available experimental evidence. Protein interaction networks (PINs) are then built by thresholding on these scores, so that only interactions of sufficiently high quality are included. These networks are used to identify biologically relevant motifs or nodes using metrics such as degree or betweenness centrality. This type of analysis can be sensitive to the choice of threshold. If a node metric is to be useful for extracting biological signal, it should induce similar node rankings across PINs obtained at different reasonable confidence score thresholds. RESULTS: We propose three measures-rank continuity, identifiability, and instability-to evaluate how robust a node metric is to changes in the score threshold. We apply our measures to twenty-five metrics and identify four as the most robust: the number of edges in the step-1 ego network, as well as the leave-one-out differences in average redundancy, average number of edges in the step-1 ego network, and natural connectivity. Our measures show good agreement across PINs from different species and data sources. Analysis of synthetically generated scored networks shows that robustness results are context-specific, and depend both on network topology and on how scores are placed across network edges. CONCLUSION: Due to the uncertainty associated with protein interaction detection, and therefore network structure, for PIN analysis to be reproducible, it should yield similar results across different confidence score thresholds. We demonstrate that while certain node metrics are robust with respect to threshold choice, this is not always the case. Promisingly, our results suggest that there are some metrics that are robust across networks constructed from different databases, and different scoring procedures.
Lyuba V. Bozhilova, Alan Whitmore, Jonny Wray, Gesine Reinert, Charlotte M. Deane
BMC Bioinform.5
2019 MHC binding affects the dynamics of different T-cell receptors in different ways
abstract
T cells use their T-cell receptors (TCRs) to scan other cells for antigenic peptides presented by MHC molecules (pMHC). If a TCR encounters a pMHC, it can trigger a signalling pathway that could lead to the activation of the T cell and the initiation of an immune response. It is currently not clear how the binding of pMHC to the TCR initiates signalling within the T cell. One hypothesis is that conformational changes in the TCR lead to further downstream signalling. Here we investigate four different TCRs in their free state as well as in their pMHC bound state using large scale molecular simulations totalling 26 000 ns. We find that the dynamical features within TCRs differ significantly between unbound TCR and TCR/pMHC simulations. However, apart from expected results such as reduced solvent accessibility and flexibility of the interface residues, these features are not conserved among different TCR types. The presence of a pMHC alone is not sufficient to cause cross-TCR-conserved dynamical features within a TCR. Our results argue against models of TCR triggering involving conserved allosteric conformational changes.
Bernhard Knapp, P. Anton van der Merwe, Omer Dushek, Charlotte M. Deane
PLoS Comput. Biol.4
2018 pyHVis3D: visualising molecular simulation deduced H-bond networks in 3D: application to T-cell receptor interactions
abstract
Motivation: Hydrogen bonds (H-bonds) play an essential role for many molecular interactions but are also often transient, making visualising them in a flexible system challenging. Results: We provide pyHVis3D which allows for an easy to interpret 3D visualisation of H-bonds resulting from molecular simulations. We demonstrate the power of pyHVis3D by using it to explain the changes in experimentally measured binding affinities for three T-cell receptor/peptide/MHC complexes and mutants of each of these complexes. Availability and implementation: pyHVis3D can be downloaded for free from http://opig.stats.ox.ac.uk/resources. Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Bernhard Knapp, Marta Alcalá, Clare E. West, P. Anton van der Merwe, Charlotte M. Deane
Bioinform.6
2018 In silico structural modeling of multiple epigenetic marks on DNA
abstract
Abstract There are four known epigenetic cytosine modifications in mammals: methylation (5mC), hydroxymethylation (5hmC), formylation (5fC) and carboxylation (5caC). The biological effects of 5mC are well understood but the roles of the remaining modifications remain elusive. Experimental and computational studies suggest that a single epigenetic mark has little structural effect but six of them can radically change the structure of DNA to a new form, F-DNA. Investigating the collective effect of multiple epigenetic marks requires the ability to interrogate all possible combinations of epigenetic states (e.g. methylated/non-methylated) along a stretch of DNA. Experiments on such complex systems are only feasible on small, isolated examples and there currently exist no systematic computational solutions to this problem. We address this issue by extending the use of Natural Move Monte Carlo to simulate the conformations of epigenetic marks. We validate our protocol by reproducing in silico experimental observations from two recently published high-resolution crystal structures that contain epigenetic marks 5hmC and 5fC. We further demonstrate that our protocol correctly finds either the F-DNA or the B-DNA states more energetically favorable depending on the configuration of the epigenetic marks. We hope that the computational efficiency and ease of use of this novel simulation framework would form the basis for future protocols and facilitate our ability to rapidly interrogate diverse epigenetic systems. Availability and implementation The code together with examples and tutorials are available from http://www.cs.ox.ac.uk/mosaics Supplementary information Supplementary data are available at Bioinformatics online.
Konrad Krawczyk, Samuel Demharter, Bernhard Knapp, Charlotte M. Deane, Peter Minary
Bioinform.4
2018 CommWalker: correctly evaluating modules in molecular networks in light of annotation bias
abstract
Motivation: Detecting novel functional modules in molecular networks is an important step in biological research. In the absence of gold standard functional modules, functional annotations are often used to verify whether detected modules/communities have biological meaning. However, as we show, the uneven distribution of functional annotations means that such evaluation methods favor communities of well-studied proteins. Results: We propose a novel framework for the evaluation of communities as functional modules. Our proposed framework, CommWalker, takes communities as inputs and evaluates them in their local network environment by performing short random walks. We test CommWalker's ability to overcome annotation bias using input communities from four community detection methods on two protein interaction networks. We find that modules accepted by CommWalker are similarly co-expressed as those accepted by current methods. Crucially, CommWalker performs well not only in well-annotated regions, but also in regions otherwise obscured by poor annotation. CommWalker community prioritization both faithfully captures well-validated communities and identifies functional modules that may correspond to more novel biology. Availability and implementation: The CommWalker algorithm is freely available at opig.stats.ox.ac.uk/resources or as a docker image on the Docker Hub at hub.docker.com/r/lueckenmd/commwalker/. Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Malte Lücken, M. J. T. Page, A. J. Crosby, S. Mason, Gesine Reinert, Charlotte M. Deane
Bioinform.6
2018 Predicting loop conformational ensembles
abstract
Motivation: Protein function is often facilitated by the existence of multiple stable conformations. Structure prediction algorithms need to be able to model these different conformations accurately and produce an ensemble of structures that represent a target's conformational diversity rather than just a single state. Here, we investigate whether current loop prediction algorithms are capable of this. We use the algorithms to predict the structures of loops with multiple experimentally determined conformations, and the structures of loops with only one conformation, and assess their ability to generate and select decoys that are close to any, or all, of the observed structures. Results: We find that while loops with only one known conformation are predicted well, conformationally diverse loops are modelled poorly, and in most cases the predictions returned by the methods do not resemble any of the known conformers. Our results contradict the often-held assumption that multiple native conformations will be present in the decoy set, making the production of accurate conformational ensembles impossible, and hence indicating that current methodologies are not well suited to prediction of conformationally diverse, often functionally important protein regions. Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Claire Marks, Jiye Shi, Charlotte M. Deane
Bioinform.3
2018 Combining co-evolution and secondary structure prediction to improve fragment library generation
abstract
Motivation: Recent advances in co-evolution techniques have made possible the accurate prediction of protein structures in the absence of a template. Here, we provide a general approach that further utilizes co-evolution constraints to generate better fragment libraries for fragment-based protein structure prediction. Results: We have compared five different fragment library generation programmes on three different datasets encompassing over 400 unique protein folds. We show that considering the secondary structure of the fragments when assembling these libraries provides a critical way to assess their usefulness to structure prediction. We then use co-evolution constraints to improve the fragment libraries by enriching them with fragments that satisfy constraints and discarding those that do not. These improved libraries have better precision and lead to consistently better modelling results. Availability and implementation: Data is available for download from: http://opig.stats.ox.ac.uk/resources. Flib-Coevo is available for download from: https://github.com/sauloho/Flib-Coevo. Supplementary information: Supplementary data are available at Bioinformatics online.
Saulo Henrique Pires de Oliveira, Charlotte M. Deane
Bioinform.2
2018 Sequential search leads to faster, more efficient fragment-based de novo protein structure prediction
abstract
Motivation: Most current de novo structure prediction methods randomly sample protein conformations and thus require large amounts of computational resource. Here, we consider a sequential sampling strategy, building on ideas from recent experimental work which shows that many proteins fold cotranslationally. Results: We have investigated whether a pseudo-greedy search approach, which begins sequentially from one of the termini, can improve the performance and accuracy of de novo protein structure prediction. We observed that our sequential approach converges when fewer than 20 000 decoys have been produced, fewer than commonly expected. Using our software, SAINT2, we also compared the run time and quality of models produced in a sequential fashion against a standard, non-sequential approach. Sequential prediction produces an individual decoy 1.5-2.5 times faster than non-sequential prediction. When considering the quality of the best model, sequential prediction led to a better model being produced for 31 out of 41 soluble protein validation cases and for 18 out of 24 transmembrane protein cases. Correct models (TM-Score > 0.5) were produced for 29 of these cases by the sequential mode and for only 22 by the non-sequential mode. Our comparison reveals that a sequential search strategy can be used to drastically reduce computational time of de novo protein structure prediction and improve accuracy. Availability and implementation: Data are available for download from: http://opig.stats.ox.ac.uk/resources. SAINT2 is available for download from: https://github.com/sauloho/SAINT2. Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Saulo Henrique Pires de Oliveira, Eleanor C. Law, Jiye Shi, Charlotte M. Deane
Bioinform.4
2017 Sphinx: merging knowledge-based and ab initio approaches to improve protein loop prediction
abstract
Motivation: Loops are often vital for protein function, however, their irregular structures make them difficult to model accurately. Current loop modelling algorithms can mostly be divided into two categories: knowledge-based, where databases of fragments are searched to find suitable conformations and ab initio, where conformations are generated computationally. Existing knowledge-based methods only use fragments that are the same length as the target, even though loops of slightly different lengths may adopt similar conformations. Here, we present a novel method, Sphinx, which combines ab initio techniques with the potential extra structural information contained within loops of a different length to improve structure prediction. Results: We show that Sphinx is able to generate high-accuracy predictions and decoy sets enriched with near-native loop conformations, performing better than the ab initio algorithm on which it is based. In addition, it is able to provide predictions for every target, unlike some knowledge-based methods. Sphinx can be used successfully for the difficult problem of antibody H3 prediction, outperforming RosettaAntibody, one of the leading H3-specific ab initio methods, both in accuracy and speed. Availability and Implementation: Sphinx is available at http://opig.stats.ox.ac.uk/webapps/sphinx. Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Claire Marks, Jaroslaw Nowak, Stefan Klostermann, Guy Georges, James Dunbar, Jiye Shi, Sebastian Kelm, Charlotte M. Deane
Bioinform.8
2017 Comparing co-evolution methods and their application to template-free protein structure prediction
abstract
Motivation: Co-evolution methods have been used as contact predictors to identify pairs of residues that share spatial proximity. Such contact predictors have been compared in terms of the precision of their predictions, but there is no study that compares their usefulness to model generation. Results: We compared eight different co-evolution methods for a set of ∼3500 proteins and found that metaPSICOV stage 2 produces, on average, the most precise predictions. Precision of all the methods is dependent on SCOP class, with most methods predicting contacts in all α and membrane proteins poorly. The contact predictions were then used to assist in de novo model generation. We found that it was not the method with the highest average precision, but rather metaPSICOV stage 1 predictions that consistently led to the best models being produced. Our modelling results show a correlation between the proportion of predicted long range contacts that are satisfied on a model and its quality. We used this proportion to effectively classify models as correct/incorrect; discarding decoys classified as incorrect led to an enrichment in the proportion of good decoys in our final ensemble by a factor of seven. For 17 out of the 18 cases where correct answers were generated, the best models were not discarded by this approach. We were also able to identify eight cases where no correct decoy had been generated. Availability and Implementation: Data is available for download from: http://opig.stats.ox.ac.uk/resources. Contact: [email protected] Supplimentary Information: Supplementary data are available at Bioinformatics online.
Saulo Henrique Pires de Oliveira, Jiye Shi, Charlotte M. Deane
Bioinform.3
2017 Ten simple rules for surviving an interdisciplinary PhD
abstract
DOAJ is a unique and extensive index of diverse open access journals from around the world, driven by a growing community, committed to ensuring quality content is freely available online for everyone.
Samuel Demharter, Nicholas Pearce, Kylie Beattie, Isabel Frost, Jinwoo Leem, Alistair Martin, Robert Oppenheimer, Cristian Regep, Tammo Rukat, Alexander Skates, Nicola Trendel, David Gavaghan, Charlotte M. Deane, Bernhard Knapp
PLoS Comput. Biol.13
2016 Progress and challenges in predicting protein interfaces
abstract
The majority of biological processes are mediated via protein-protein interactions. Determination of residues participating in such interactions improves our understanding of molecular mechanisms and facilitates the development of therapeutics. Experimental approaches to identifying interacting residues, such as mutagenesis, are costly and time-consuming and thus, computational methods for this purpose could streamline conventional pipelines. Here we review the field of computational protein interface prediction. We make a distinction between methods which address proteins in general and those targeted at antibodies, owing to the radically different binding mechanism of antibodies. We organize the multitude of currently available methods hierarchically based on required input and prediction principles to provide an overview of the field.
Reyhaneh Esmaielbeiki, Konrad Krawczyk, Bernhard Knapp, Jean-Christophe Nebel, Charlotte M. Deane
Briefings Bioinform.5
2016 ANARCI: antigen receptor numbering and receptor classification
abstract
MOTIVATION: Antibody amino-acid sequences can be numbered to identify equivalent positions. Such annotations are valuable for antibody sequence comparison, protein structure modelling and engineering. Multiple different numbering schemes exist, they vary in the nomenclature they use to annotate residue positions, their definitions of position equivalence and their popularity within different scientific disciplines. However, currently no publicly available software exists that can apply all the most widely used schemes or for which an executable can be obtained under an open license. RESULTS: ANARCI is a tool to classify and number antibody and T-cell receptor amino-acid variable domain sequences. It can annotate sequences with the five most popular numbering schemes: Kabat, Chothia, Enhanced Chothia, IMGT and AHo. AVAILABILITY AND IMPLEMENTATION: ANARCI is available for download under GPLv3 license at opig.stats.ox.ac.uk/webapps/anarci. A web-interface to the program is available at the same address. CONTACT: [email protected].
James Dunbar, Charlotte M. Deane
Bioinform.2
2016 Exploring peptide/MHC detachment processes using hierarchical natural move Monte Carlo
abstract
MOTIVATION: The binding between a peptide and a major histocompatibility complex (MHC) is one of the most important processes for the induction of an adaptive immune response. Many algorithms have been developed to predict peptide/MHC (pMHC) binding. However, no approach has yet been able to give structural insight into how peptides detach from the MHC. RESULTS: In this study, we used a combination of coarse graining, hierarchical natural move Monte Carlo and stochastic conformational optimization to explore the detachment processes of 32 different peptides from HLA-A*02:01. We performed 100 independent repeats of each stochastic simulation and found that the presence of experimentally known anchor amino acids affects the detachment trajectories of our peptides. Comparison with experimental binding affinity data indicates the reliability of our approach (area under the receiver operating characteristic curve 0.85). We also compared to a 1000 ns molecular dynamics simulation of a non-binding peptide (AAAKTPVIV) and HLA-A*02:01. Even in this simulation, the longest published for pMHC, the peptide does not fully detach. Our approach is orders of magnitude faster and as such allows us to explore pMHC detachment processes in a way not possible with all-atom molecular dynamics simulations. AVAILABILITY AND IMPLEMENTATION: The source code is freely available for download at http://www.cs.ox.ac.uk/mosaics/. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Bernhard Knapp, Samuel Demharter, Charlotte M. Deane, Peter Minary
Bioinform.3
2015 Current status and future challenges in T-cell receptor/peptide/MHC molecular dynamics simulations
abstract
The interaction between T-cell receptors (TCRs) and major histocompatibility complex (MHC)-bound epitopes is one of the most important processes in the adaptive human immune response. Several hypotheses on TCR triggering have been proposed. Many of them involve structural and dynamical adjustments in the TCR/peptide/MHC interface. Molecular Dynamics (MD) simulations are a computational technique that is used to investigate structural dynamics at atomic resolution. Such simulations are used to improve understanding of signalling on a structural level. Here we review how MD simulations of the TCR/peptide/MHC complex have given insight into immune system reactions not achievable with current experimental methods. Firstly, we summarize methods of TCR/peptide/MHC complex modelling and TCR/peptide/MHC MD trajectory analysis methods. Then we classify recently published simulations into categories and give an overview of approaches and results. We show that current studies do not come to the same conclusions about TCR/peptide/MHC interactions. This discrepancy might be caused by too small sample sizes or intrinsic differences between each interaction process. As computational power increases future studies will be able to and should have larger sample sizes, longer runtimes and additional parts of the immunological synapse included.
Bernhard Knapp, Samuel Demharter, Reyhaneh Esmaielbeiki, Charlotte M. Deane
Briefings Bioinform.4
2015 Structural Bridges through Fold Space
abstract
Several protein structure classification schemes exist that partition the protein universe into structural units called folds. Yet these schemes do not discuss how these units sit relative to each other in a global structure space. In this paper we construct networks that describe such global relationships between folds in the form of structural bridges. We generate these networks using four different structural alignment methods across multiple score thresholds. The networks constructed using the different methods remain a similar distance apart regardless of the probability threshold defining a structural bridge. This suggests that at least some structural bridges are method specific and that any attempt to build a picture of structural space should not be reliant on a single structural superposition method. Despite these differences all representations agree on an organisation of fold space into five principal community structures: all-α, all-β sandwiches, all-β barrels, α/β and α + β. We project estimated fold ages onto the networks and find that not only are the pairings of unconnected folds associated with higher age differences than bridged folds, but this difference increases with the number of networks displaying an edge. We also examine different centrality measures for folds within the networks and how these relate to fold age. While these measures interpret the central core of fold space in varied ways they all identify the disposition of ancestral folds to fall within this core and that of the more recently evolved structures to provide the peripheral landscape. These findings suggest that evolutionary information is encoded along these structural bridges. Finally, we identify four highly central pivotal folds representing dominant topological features which act as key attractors within our landscapes.
Hannah Edwards, Charlotte M. Deane
PLoS Comput. Biol.2
2015 Ten Simple Rules for a Successful Cross-Disciplinary Collaboration
abstract
Cross-disciplinary collaborations have become an increasingly important part of science. They are seen as a key factor for finding solutions to pressing societal challenges on a global scale including green technologies, sustainable food production and drug development. This has also been realized by regulators and policy-makers, as it is reflected in the 80 billion Euro "Horizon 2020" EU Framework Programme for Research and Innovation. This programme puts special emphasis at breaking down barriers between fields to create a path breaking environment for knowledge, research and innovation. However, igniting and successfully maintaining cross-disciplinary collaborations can be a delicate task. In this article we focus on the specific challenges associated with cross-disciplinary research in particular from the perspective of the theoretician. As research fellows of the 2020 Science project (http://www.2020science.net) and collaboration partners, we bring broad experience of developing interdisciplinary collaborations [2–12]. We intend this guide for early career computational researchers as well as more senior scientists who are entering a cross disciplinary setting for the first time. We describe the key benefits, as well as some possible pitfalls, arising from collaborations between scientists with backgrounds in very different fields. This paper has inter alia been cited by Times Higher education: http://www.timeshighereducation.co.uk/news/people/the-secrets-to-successful-interdisciplinary-work/2020267.article .
Bernhard Knapp, Rémi Bardenet, Miguel O. Bernabeu, Rafel Bordas, Maria Bruna, Ben Calderhead, Jonathan Cooper, Alexander G. Fletcher, Derek Groen, Bram Kuijper, Joanna Lewis, Greg J. McInerny, Timo Minssen, James M. Osborne, Verena Paulitschke, Joe Pitt-Francis, Jelena Todoric, Christian A. Yates, David Gavaghan, Charlotte M. Deane
PLoS Comput. Biol.20
2014 Alignment-free protein interaction network comparison
abstract
MOTIVATION: Biological network comparison software largely relies on the concept of alignment where close matches between the nodes of two or more networks are sought. These node matches are based on sequence similarity and/or interaction patterns. However, because of the incomplete and error-prone datasets currently available, such methods have had limited success. Moreover, the results of network alignment are in general not amenable for distance-based evolutionary analysis of sets of networks. In this article, we describe Netdis, a topology-based distance measure between networks, which offers the possibility of network phylogeny reconstruction. RESULTS: We first demonstrate that Netdis is able to correctly separate different random graph model types independent of network size and density. The biological applicability of the method is then shown by its ability to build the correct phylogenetic tree of species based solely on the topology of current protein interaction networks. Our results provide new evidence that the topology of protein interaction networks contains information about evolutionary processes, despite the lack of conservation of individual interactions. As Netdis is applicable to all networks because of its speed and simplicity, we apply it to a large collection of biological and non-biological networks where it clusters diverse networks by type. AVAILABILITY AND IMPLEMENTATION: The source code of the program is freely available at http://www.stats.ox.ac.uk/research/proteins/resources. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Tiago Rito, Gesine Reinert, Fengzhu Sun, Charlotte M. Deane
Bioinform.5
2014 Improving B-cell epitope prediction and its application to global antibody-antigen docking
abstract
MOTIVATION: Antibodies are currently the most important class of biopharmaceuticals. Development of such antibody-based drugs depends on costly and time-consuming screening campaigns. Computational techniques such as antibody-antigen docking hold the potential to facilitate the screening process by rapidly providing a list of initial poses that approximate the native complex. RESULTS: We have developed a new method to identify the epitope region on the antigen, given the structures of the antibody and the antigen-EpiPred. The method combines conformational matching of the antibody-antigen structures and a specific antibody-antigen score. We have tested the method on both a large non-redundant set of antibody-antigen complexes and on homology models of the antibodies and/or the unbound antigen structure. On a non-redundant test set, our epitope prediction method achieves 44% recall at 14% precision against 23% recall at 14% precision for a background random distribution. We use our epitope predictions to rescore the global docking results of two rigid-body docking algorithms: ZDOCK and ClusPro. In both cases including our epitope, prediction increases the number of near-native poses found among the top decoys. AVAILABILITY AND IMPLEMENTATION: Our software is available from http://www.stats.ox.ac.uk/research/proteins/resources.
Konrad Krawczyk, Terry Baker, Jiye Shi, Charlotte M. Deane
Bioinform.5
2014 Examining Variable Domain Orientations in Antigen Receptors Gives Insight into TCR-Like Antibody Design
abstract
The variable domains of antibodies and T-Cell receptors (TCRs) share similar structures. Both molecules act as sensors for the immune system but recognise their respective antigens in different ways. Antibodies bind to a diverse set of antigenic shapes whilst TCRs only recognise linear peptides presented by a major histocompatibility complex (MHC). The antigen specificity and affinity of both receptors is determined primarily by the sequence and structure of their complementarity determining regions (CDRs). In antibodies the binding site is also known to be affected by the relative orientation of the variable domains, VH and VL. Here, the corresponding property for TCRs, the Vβ-Vα orientation, is investigated and compared with that of antibodies. We find that TCR and antibody orientations are distinct. General antibody orientations are found to be incompatible with binding to the MHC in a canonical TCR-like mode. Finally, factors that cause the orientation of TCRs and antibodies to be different are investigated. Packing of the long Vα CDR3 in the domain-domain interface is found to be influential. In antibodies, a similar packing affect can be achieved using a bulky residue at IMGT position 50 on the VH domain. Along with IMGT VH 50, other positions are identified that may help to promote a TCR-like orientation in antibodies. These positions should provide useful considerations in the engineering of therapeutic TCR-like antibodies.
James Dunbar, Bernhard Knapp, Angelika Fuchs, Jiye Shi, Charlotte M. Deane
PLoS Comput. Biol.5
2014 Large Scale Characterization of the LC13 TCR and HLA-B8 Structural Landscape in Reaction to 172 Altered Peptide Ligands: A Molecular Dynamics Simulation Study
abstract
The interplay between T cell receptors (TCRs) and peptides bound by major histocompatibility complexes (MHCs) is one of the most important interactions in the adaptive immune system. Several previous studies have computationally investigated their structural dynamics. On the basis of these simulations several structural and dynamical properties have been proposed as effectors of the immunogenicity. Here we present the results of a large scale Molecular Dynamics simulation study consisting of 100 ns simulations of 172 different complexes. These complexes consisted of all possible point mutations of the Epstein Barr Virus peptide FLRGRAYGL bound by HLA-B*08:01 and presented to the LC13 TCR. We compare the results of these 172 structural simulations with experimental immunogenicity data. We found that simulations with more immunogenic peptides and those with less immunogenic peptides are in fact highly similar and on average only minor differences in the hydrogen binding footprints, interface distances, and the relative orientation between the TCR chains are present. Thus our large scale data analysis shows that many previously suggested dynamical and structural properties of the TCR/peptide/MHC interface are unlikely to be conserved causal factors for peptide immunogenicity.
Bernhard Knapp, James Dunbar, Charlotte M. Deane
PLoS Comput. Biol.3
2014 Ten Simple Rules for Effective Computational Research
abstract
In order to attempt to understand the complexity inherent in nature, mathematical, statistical and computational techniques are increasingly being employed in the life sciences. In particular, the use and development of software tools is becoming vital for investigating scientific hypotheses, and a wide range of scientists are finding software development playing a more central role in their day-to-day research. In fields such as biology and ecology, there has been a noticeable trend towards the use of quantitative methods for both making sense of ever-increasing amounts of data [1] and building or selecting models [2]. As Research Fellows of the “2020 Science” project (http://www.2020science.net), funded jointly by the EPSRC (Engineering and Physical Sciences Research Council) and Microsoft Research, we have firsthand experience of the challenges associated with carrying out multidisciplinary computation-based science [3]–[5]. In this paper we offer a jargon-free guide to best practice when developing and using software for scientific research. While many guides to software development exist, they are often aimed at computer scientists [6] or concentrate on large open-source projects [7]; the present guide is aimed specifically at the vast majority of scientific researchers: those without formal training in computer science. We present our ten simple rules with the aim of enabling scientists to be more effective in undertaking research and therefore maximise the impact of this research within the scientific community. While these rules are described individually, collectively they form a single vision for how to approach the practical side of computational science. Our rules are presented in roughly the chronological order in which they should be undertaken, beginning with things that, as a computational scientist, you should do before you even think about writing any code. For each rule, guides on getting started, links to relevant tutorials, and further reading are provided in the supplementary material (Text S1).
James M. Osborne, Miguel O. Bernabeu, Maria Bruna, Ben Calderhead, Jonathan Cooper, Neil Dalchau, Sara-Jane Dunn, Alexander G. Fletcher, Robin Freeman, Derek Groen, Bernhard Knapp, Greg J. McInerny, Gary R. Mirams, Joe Pitt-Francis, Biswa Sengupta, David W. Wright 0001, Christian A. Yates, David Gavaghan, Stephen Emmott, Charlotte M. Deane
PLoS Comput. Biol.20
2013 MP-T: improving membrane protein alignment for structure prediction
abstract
MOTIVATION: Membrane proteins are clinically relevant, yet their crystal structures are rare. Models of membrane proteins are typically built from template structures with low sequence identity to the target sequence, using a sequence-structure alignment as a blueprint. This alignment is usually made with programs designed for use on soluble proteins. Biological membranes have layers of varying hydrophobicity, and membrane proteins have different amino-acid substitution preferences from their soluble counterparts. Here we include these factors into an alignment method to improve alignments and consequently improve membrane protein models. RESULTS: We developed Membrane Protein Threader (MP-T), a sequence-structure alignment tool for membrane proteins based on multiple sequence alignment. Alignment accuracy is tested against seven other alignment methods over 165 non-redundant alignments of membrane proteins. MP-T produces more accurate alignments than all other methods tested (δF(M) from +0.9 to +5.5%). Alignments generated by MP-T also lead to significantly better models than those of the best alternative alignment tool (one-fourth of models see an increase in GDT_TS of ≥4%). AVAILABILITY: All source code, alignments and models are available at http://www.stats.ox.ac.uk/proteins/resources
Jamie R. Hill, Charlotte M. Deane
Bioinform.2
2013 Exploring Fold Space Preferences of New-born and Ancient Protein Superfamilies
abstract
The evolution of proteins is one of the fundamental processes that has delivered the diversity and complexity of life we see around ourselves today. While we tend to define protein evolution in terms of sequence level mutations, insertions and deletions, it is hard to translate these processes to a more complete picture incorporating a polypeptide's structure and function. By considering how protein structures change over time we can gain an entirely new appreciation of their long-term evolutionary dynamics. In this work we seek to identify how populations of proteins at different stages of evolution explore their possible structure space. We use an annotation of superfamily age to this space and explore the relationship between these ages and a diverse set of properties pertaining to a superfamily's sequence, structure and function. We note several marked differences between the populations of newly evolved and ancient structures, such as in their length distributions, secondary structure content and tertiary packing arrangements. In particular, many of these differences suggest a less elaborate structure for newly evolved superfamilies when compared with their ancient counterparts. We show that the structural preferences we report are not a residual effect of a more fundamental relationship with function. Furthermore, we demonstrate the robustness of our results, using significant variation in the algorithm used to estimate the ages. We present these age estimates as a useful tool to analyse protein populations. In particularly, we apply this in a comparison of domains containing greek key or jelly roll motifs.
Hannah Edwards, Sanne Abeln, Charlotte M. Deane
PLoS Comput. Biol.3
2012 What Evidence Is There for the Homology of Protein-Protein Interactions?
abstract
The notion that sequence homology implies functional similarity underlies much of computational biology. In the case of protein-protein interactions, an interaction can be inferred between two proteins on the basis that sequence-similar proteins have been observed to interact. The use of transferred interactions is common, but the legitimacy of such inferred interactions is not clear. Here we investigate transferred interactions and whether data incompleteness explains the lack of evidence found for them. Using definitions of homology associated with functional annotation transfer, we estimate that conservation rates of interactions are low even after taking interactome incompleteness into account. For example, at a blastp E-value threshold of 10(-70), we estimate the conservation rate to be about 11 % between S. cerevisiae and H. sapiens. Our method also produces estimates of interactome sizes (which are similar to those previously proposed). Using our estimates of interaction conservation we estimate the rate at which protein-protein interactions are lost across species. To our knowledge, this is the first such study based on large-scale data. Previous work has suggested that interactions transferred within species are more reliable than interactions transferred across species. By controlling for factors that are specific to within-species interaction prediction, we propose that the transfer of interactions within species might be less reliable than transfers between species. Protein-protein interactions appear to be very rarely conserved unless very high sequence similarity is observed. Consequently, inferred interactions should be used with care.
Anna C. F. Lewis, Nick S. Jones, Mason A. Porter, Charlotte M. Deane
PLoS Comput. Biol.4
2011 Environment specific substitution tables improve membrane protein alignment
abstract
MOTIVATION: Membrane proteins are both abundant and important in cells, but the small number of solved structures restricts our understanding of them. Here we consider whether membrane proteins undergo different substitutions from their soluble counterparts and whether these can be used to improve membrane protein alignments, and therefore improve prediction of their structure. RESULTS: We construct substitution tables for different environments within membrane proteins. As data is scarce, we develop a general metric to assess the quality of these asymmetric tables. Membrane proteins show markedly different substitution preferences from soluble proteins. For example, substitution preferences in lipid tail-contacting parts of membrane proteins are found to be distinct from all environments in soluble proteins, including buried residues. A principal component analysis of the tables identifies the greatest variation in substitution preferences to be due to changes in hydrophobicity; the second largest variation relates to secondary structure. We demonstrate the use of our tables in pairwise sequence-to-structure alignments (also known as 'threading') of membrane proteins using the FUGUE alignment program. On average, in the 10-25% sequence identity range, alignments are improved by 28 correctly aligned residues compared with alignments made using FUGUE's default substitution tables. Our alignments also lead to improved structural models. AVAILABILITY: Substitution tables are available at: http://www.stats.ox.ac.uk/proteins/resources.
Jamie R. Hill, Sebastian Kelm, Jiye Shi, Charlotte M. Deane
Bioinform.4
2010 MEDELLER: homology-based coordinate generation for membrane proteins
abstract
MOTIVATION: Membrane proteins (MPs) are important drug targets but knowledge of their exact structure is limited to relatively few examples. Existing homology-based structure prediction methods are designed for globular, water-soluble proteins. However, we are now beginning to have enough MP structures to justify the development of a homology-based approach specifically for them. RESULTS: We present a MP-specific homology-based coordinate generation method, MEDELLER, which is optimized to build highly reliable core models. The method outperforms the popular structure prediction programme Modeller on MPs. The comparison of the two methods was performed on 616 target-template pairs of MPs, which were classified into four test sets by their sequence identity. Across all targets, MEDELLER gave an average backbone root mean square deviation (RMSD) of 2.62 Å versus 3.16 Å for Modeller. On our 'easy' test set, MEDELLER achieves an average accuracy of 0.93 Å backbone RMSD versus 1.56 Å for Modeller. AVAILABILITY AND IMPLEMENTATION: http://medeller.info; Implemented in Python, Bash and Perl CGI for use on Linux systems; Supplementary data are available at http://www.stats.ox.ac.uk/proteins/resources.
Sebastian Kelm, Jiye Shi, Charlotte M. Deane
Bioinform.3
2010 Exploring the potential of template-based modelling
abstract
MOTIVATION: Template-based modelling can approximate the unknown structure of a target protein using an homologous template structure. The core of the resulting prediction then comprises the structural regions conserved between template and target. Target prediction could be improved by rigidly repositioning such single template, structurally conserved fragment regions. The purpose of this article is to quantify the extent to which such improvements are possible and to relate this extent to properties of the target, the template and their alignment. RESULTS: The improvement in accuracy achievable when rigid fragments from a single template are optimally positioned was calculated using structure pairs from the HOMSTRAD database, as well as CASP7 and CASP8 target/best template pairs. Over the union of the structurally conserved regions, improvements of 0.7 A in root mean squared deviation (RMSD) and 6% in GDT_HA were commonly observed. A generalized linear model revealed that the extent to which a template can be improved can be predicted using four variables. Templates with the greatest scope for improvement tend to have relatively more fragments, shorter fragments, higher percentage of helical secondary structure and lower sequence identity. Optimal positioning of the template fragments offers the potential for improving loop modelling. These results demonstrate that substantial improvement could be made on many templates if the conserved fragments were to be optimally positioned. They also provide a basis for identifying templates for which modification of fragment positions may yield such improvements.
Braddon K. Lance, Charlotte M. Deane, Graham R. Wood
Bioinform.2
2010 How threshold behaviour affects the use of subgraphs for network comparison
abstract
MOTIVATION: A wealth of protein-protein interaction (PPI) data has recently become available. These data are organized as PPI networks and an efficient and biologically meaningful method to compare such PPI networks is needed. As a first step, we would like to compare observed networks to established network models, under the aspect of small subgraph counts, as these are conjectured to relate to functional modules in the PPI network. We employ the software tool GraphCrunch with the Graphlet Degree Distribution Agreement (GDDA) score to examine the use of such counts for network comparison. RESULTS: Our results show that the GDDA score has a pronounced dependency on the number of edges and vertices of the networks being considered. This should be taken into account when testing the fit of models. We provide a method for assessing the statistical significance of the fit between random graph models and biological networks based on non-parametric tests. Using this method we examine the fit of Erdös-Rényi (ER), ER with fixed degree distribution and geometric (3D) models to PPI networks. Under these rigorous tests none of these models fit to the PPI networks. The GDDA score is not stable in the region of graph density relevant to current PPI networks. We hypothesize that this score instability is due to the networks under consideration having a graph density in the threshold region for the appearance of small subgraphs. This is true for both geometric (3D) and ER random graph models. Such threshold behaviour may be linked to the robustness and efficiency properties of the PPI networks.
Tiago Rito, Charlotte M. Deane, Gesine Reinert
Bioinform.3
2010 Directionality in protein fold prediction
abstract
BACKGROUND: Ever since the ground-breaking work of Anfinsen et al. in which a denatured protein was found to refold to its native state, it has been frequently stated by the protein fold prediction community that all the information required for protein folding lies in the amino acid sequence. Recent in vitro experiments and in silico computational studies, however, have shown that cotranslation may affect the folding pathway of some proteins, especially those of ancient folds. In this paper aspects of cotranslational folding have been incorporated into a protein structure prediction algorithm by adapting the Rosetta program to fold proteins as the nascent chain elongates. This makes it possible to conduct a pairwise comparison of folding accuracy, by comparing folds created sequentially from each end of the protein. RESULTS: A single main result emerged: in 94% of proteins analyzed, following the sense of translation, from N-terminus to C-terminus, produced better predictions than following the reverse sense of translation, from the C-terminus to N-terminus. Two secondary results emerged. First, this superiority of N-terminus to C-terminus folding was more marked for proteins showing stronger evidence of cotranslation and second, an algorithm following the sense of translation produced predictions comparable to, and occasionally better than, Rosetta. CONCLUSIONS: There is a directionality effect in protein fold prediction. At present, prediction methods appear to be too noisy to take advantage of this effect; as techniques refine, it may be possible to draw benefit from a sequential approach to protein fold prediction.
Jonathan J. Ellis, Fabien P. E. Huard, Charlotte M. Deane, Sheenal Srivastava, Graham R. Wood
BMC Bioinform.3
2010 Revisiting Date and Party Hubs: Novel Approaches to Role Assignment in Protein Interaction Networks
abstract
The idea of "date" and "party" hubs has been influential in the study of protein-protein interaction networks. Date hubs display low co-expression with their partners, whilst party hubs have high co-expression. It was proposed that party hubs are local coordinators whereas date hubs are global connectors. Here, we show that the reported importance of date hubs to network connectivity can in fact be attributed to a tiny subset of them. Crucially, these few, extremely central, hubs do not display particularly low expression correlation, undermining the idea of a link between this quantity and hub function. The date/party distinction was originally motivated by an approximately bimodal distribution of hub co-expression; we show that this feature is not always robust to methodological changes. Additionally, topological properties of hubs do not in general correlate with co-expression. However, we find significant correlations between interaction centrality and the functional similarity of the interacting proteins. We suggest that thinking in terms of a date/party dichotomy for hubs in protein interaction networks is not meaningful, and it might be more useful to conceive of roles for protein-protein interactions rather than for individual proteins.
Sumeet Agarwal, Charlotte M. Deane, Mason A. Porter, Nick S. Jones
PLoS Comput. Biol.2
2009 Functionally guided alignment of protein interaction networks for module detection
abstract
MOTIVATION: Functional module detection within protein interaction networks is a challenging problem due to the sparsity of data and presence of errors. Computational techniques for this task range from purely graph theoretical approaches involving single networks to alignment of multiple networks from several species. Current network alignment methods all rely on protein sequence similarity to map proteins across species. RESULTS: Here we carry out network alignment using a protein functional similarity measure. We show that using functional similarity to map proteins across species improves network alignment in terms of functional coherence and overlap with experimentally verified protein complexes. Moreover, the results from functional similarity-based network alignment display little overlap (<15%) with sequence similarity-based alignment. Our combined approach integrating sequence and function-based network alignment alongside graph clustering properties offers a 200% increase in coverage of experimental datasets and comparable accuracy to current network alignment methods. AVAILABILITY: Program binaries and source code is freely available at http://www.stats.ox.ac.uk/research/bioinfo/resources. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Charlotte M. Deane
Bioinform.2
2009 iMembrane: homology-based membrane-insertion of proteins
abstract
Abstract Summary: iMembrane is a homology-based method, which predicts a membrane protein's position within a lipid bilayer. It projects the results of coarse-grained molecular dynamics simulations onto any membrane protein structure or sequence provided by the user. iMembrane is simple to use and is currently the only computational method allowing the rapid prediction of a membrane protein's lipid bilayer insertion. Bilayer insertion data are essential in the accurate structural modelling of membrane proteins or the design of drugs that target them. Availability: http://imembrane.info. iMembrane is available under a non-commercial open-source licence, upon request. Contact: [email protected] Supplementary information: Supplementary data are available at Bioinformatics online and at http://www.stats.ox.ac.uk/proteins/resources.
Sebastian Kelm, Jiye Shi, Charlotte M. Deane
Bioinform.3
2008 An assessment of the uses of homologous interactions
abstract
MOTIVATION: Protein-protein interactions have proved to be a valuable starting point for understanding the inner workings of the cell. Computational methodologies have been built which both predict interactions and use interaction datasets in order to predict other protein features. Such methods require gold standard positive (GSP) and negative (GSN) interaction sets. Here we examine and demonstrate the usefulness of homologous interactions in predicting good quality positive and negative interaction datasets. RESULTS: We generate GSP interaction sets as subsets from experimental data using only interaction and sequence information. We can therefore produce sets for several species (many of which at present have no identified GSPs). Comprehensive error rate testing demonstrates the power of the method. We also show how the use of our datasets significantly improves the predictive power of algorithms for interaction prediction and function prediction. Furthermore, we generate GSN interaction sets for yeast and examine the use of homology along with other protein properties such as localization, expression and function. Using a novel method to assess the accuracy of a negative interaction set, we find that the best single selector for negative interactions is a lack of co-function. However, an integrated method using all the characteristics shows significant improvement over any current method for identifying GSN interactions. The nature of homologous interactions is also examined and we demonstrate that interologs are found more commonly within species than across species. CONCLUSION: GSP sets built using our homologous verification method are demonstrably better than standard sets in terms of predictive ability. We can build such GSP sets for several species. When generating GSNs we show a combination of protein features and lack of homologous interactions gives the highest quality interaction sets. AVAILABILITY: GSP and GSN datasets for all the studied species can be downloaded from http://www.stats.ox.ac.uk/~deane/HPIV.
Ramazan Saeed, Charlotte M. Deane
Bioinform.2
2008 Predicting and Validating Protein Interactions Using Network Structure
abstract
Protein interactions play a vital part in the function of a cell. As experimental techniques for detection and validation of protein interactions are time consuming, there is a need for computational methods for this task. Protein interactions appear to form a network with a relatively high degree of local clustering. In this paper we exploit this clustering by suggesting a score based on triplets of observed protein interactions. The score utilises both protein characteristics and network properties. Our score based on triplets is shown to complement existing techniques for predicting protein interactions, outperforming them on data sets which display a high degree of clustering. The predicted interactions score highly against test measures for accuracy. Compared to a similar score derived from pairwise interactions only, the triplet score displays higher sensitivity and specificity. By looking at specific examples, we show how an experimental set of interactions can be enriched and validated. As part of this work we also examine the effect of different prior databases upon the accuracy of prediction and find that the interactions from the same kingdom give better results than from across kingdoms, suggesting that there may be fundamental differences between the networks. These results all emphasize that network structure is important and helps in the accurate prediction of protein interactions. The protein interaction data set and the program used in our analysis, and a list of predictions and validations, are available at http://www.stats.ox.ac.uk/bioinfo/resources/PredictingInteractions.
Pao-Yang Chen, Charlotte M. Deane, Gesine Reinert
PLoS Comput. Biol.2
2007 A statistical approach using network structure in the prediction of protein characteristics
abstract
MOTIVATION: The Majority Vote approach has demonstrated that protein-protein interactions can be used to predict the structure or function of a protein. In this article we propose a novel method for the prediction of such protein characteristics based on frequencies of pairwise interactions. In addition, we study a second new approach using the pattern frequencies of triplets of proteins, thus for the first time taking network structure explicitly into account. Both these methods are extended to jointly consider multiple organisms and multiple characteristics. RESULTS: Compared to the standard non-network-based method, namely the Majority Vote method, in large networks our predictions tend to be more accurate. For structure prediction, the Frequency-based method reaches up to 71% accuracy, and the Triplet-based method reaches up to 72% accuracy, whereas for function prediction, both the Triplet-based method and the Frequency-based method reach up to 90% accuracy. Function prediction on proteins without homologues showed slightly less but comparable accuracies. Including partially annotated proteins substantially increases the number of proteins for which our methods predict their characteristics with reasonable accuracy. We find that the enhanced Triplet-based method does not currently yield significantly better results than the enhanced Frequency-based method, suggesting that triplets of interactions do not contain substantially more information about protein characteristics than interaction pairs. Our methods offer two main improvements over current approaches--first, multiple protein characteristics are considered simultaneously, and second, data is integrated from multiple species. In addition, the Triplet-based method includes network structure more explicitly than the Majority Vote and the Frequency-based method. AVAILABILITY: The program is available upon request. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Pao-Yang Chen, Charlotte M. Deane, Gesine Reinert
Bioinform.2
2007 Using Phylogeny to Improve Genome-Wide Distant Homology Recognition
abstract
The gap between the number of known protein sequences and structures continues to widen, particularly as a result of sequencing projects for entire genomes. Recently there have been many attempts to generate structural assignments to all genes on sets of completed genomes using fold-recognition methods. We developed a method that detects false positives made by these genome-wide structural assignment experiments by identifying isolated occurrences. The method was tested using two sets of assignments, generated by SUPERFAMILY and PSI-BLAST, on 150 completed genomes. A phylogeny of these genomes was built and a parsimony algorithm was used to identify isolated occurrences by detecting occurrences that cause a gain at leaf level. Isolated occurrences tend to have high e-values, and in both sets of assignments, a sudden increase in isolated occurrences is observed for e-values >10(-8) for SUPERFAMILY and >10(-4) for PSI-BLAST. Conditions to predict false positives are based on these results. Independent tests confirm that the predicted false positives are indeed more likely to be incorrectly assigned. Evaluation of the predicted false positives also showed that the accuracy of profile-based fold-recognition methods might depend on secondary structure content and sequence length. We show that false positives generated by fold-recognition methods can be identified by considering structural occurrence patterns on completed genomes; occurrences that are isolated within the phylogeny tend to be less reliable. The method provides a new independent way to examine the quality of fold assignments and may be used to improve the output of any genome-wide fold assignment method.
Sanne Abeln, Carlo Teubner, Charlotte M. Deane
PLoS Comput. Biol.3
2006 Protein protein interactions, evolutionary rate, abundance and age
abstract
BACKGROUND: Does a relationship exist between a protein's evolutionary rate and its number of interactions? This relationship has been put forward many times, based on a biological premise that a highly interacting protein will be more restricted in its sequence changes. However, to date several studies have voiced conflicting views on the presence or absence of such a relationship. RESULTS: Here we perform a large scale study over multiple data sets in order to demonstrate that the major reason for conflict between previous studies is the use of different but overlapping datasets. We show that lack of correlation, between evolutionary rate and number of interactions in a data set is related to the error rate. We also demonstrate that the correlation is not an artifact of the underlying distributions of evolutionary distance and interactions and is therefore likely to be biologically relevant. Further to this, we consider the claim that the dependence is due to gene expression levels and find some supporting evidence. A strong and positive correlation between the number of interactions and the age of a protein is also observed and we show this relationship is independent of expression levels. CONCLUSION: A correlation between number of interactions and evolutionary rate is observed but is dependent on the accuracy of the dataset being used. However it appears that the number of interactions a protein participates in depends more on the age of the protein than the rate at which it changes.
Ramazan Saeed, Charlotte M. Deane
BMC Bioinform.2
2001 SCORE: predicting the core of protein models
abstract
Abstract Motivation: The prediction of the regions of homology models that can be ‘restrained by’ or ‘copied from’ the basis structures is a vital step in correct model generation, because these regions are the models most accurate part. However, there is no ideal method for the identification of their limits. In most algorithms their length depends on the number of family members and definitions of secondary structure. Results: The algorithm SCORE steps away from the conventional definitions of the core to identify from large numbers of basis structures those regions that can be considered structurally related to a target sequence. The use of \batchmode \documentclass[fleqn,10pt,legalpaper]{article} \usepackage{amssymb} \usepackage{amsfonts} \usepackage{amsmath} \pagestyle{empty} \begin{document} \({\phi},\ {\psi}\) \end{document}constraints to accurately pinpoint the regions that are conserved across a family and environmentally constrained substitution tables to extend these regions allows SCORE to rapidly (generally in under 1 s, an order of magnitude faster than methods such as MODELLER) identify and build the core of homology models from the alignments of the target sequence to the basis structures. The SCORE algorithm was used to build 114 model cores. In only two cases was the core size less than 50% of the structure and all the cores built had an RMSD of 3.7 Âor less to the target structure. Availability: The algorithm is available upon request. Contact: [email protected] * To whom correspondence should be addressed.
Charlotte M. Deane, Quentin Kaas, Tom L. Blundell
Bioinform.1
2000 Browsing the SLoop database of structurally classified loops connecting elements of protein secondary structure
abstract
We describe a web server, which provides easy access to the SLoop database of loop conformations connecting elements of protein secondary structure. The loops are classified according to their length, the type of bounding secondary structures and the conformation of the mainchain. The current release of the database consists of over 8000 loops of up to 20 residues in length. A loop prediction method, which selects conformers on the basis of the sequence and the positions of the elements of secondary structure, is also implemented. These web pages are freely accessible over the internet at http://www-cryst.bioc.cam.ac.uk/ approximately sloop.
David F. Burke, Charlotte M. Deane, Tom L. Blundell
Bioinform.2
1998 JOY: protein sequence-structure representation and analysis
abstract
MOTIVATION: JOY is a program to annotate protein sequence alignments with three-dimensional (3D) structural features. It was developed to display 3D structural information in a sequence alignment and to help understand the conservation of amino acids in their specific local environments. RESULTS: : The JOY representation now constitutes an essential part of the two databases of protein structure alignments: HOMSTRAD (http://www-cryst.bioc.cam.ac.uk/homstrad ) and CAMPASS (http://www-cryst.bioc.cam.ac. uk/campass). It has also been successfully used for identifying distant evolutionary relationships. AVAILABILITY: The program can be obtained via anonymous ftp from torsa.bioc.cam.ac.uk from the directory /pub/joy/. The address for the JOY server is http://www-cryst.bioc.cam.ac.uk/cgi-bin/joy.cgi. CONTACT: [email protected]
Kenji Mizuguchi, Charlotte M. Deane, Tom L. Blundell, Mark S. Johnson, John P. Overington
Bioinform.2