Christine A. Orengo

dblp:71/4733 · DBLP profile ↗
← Back
40ranked-venue papers
0as first author
12since 2021 · last 2026
0000-0002-7141-8936ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 38 · 12 since 2021Artificial intelligence and machine learning · 2Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 TEDLH: domain HMMs for sensitive detection of remote homologues
abstract
MOTIVATION: The Encyclopedia of Domains (TED) provides domain annotations for proteins in the AlphaFold Protein Structure Database (AFDB) using a consensus of three state-of-the-art structure-based methods. We used these annotations to construct profile Hidden Markov models (HMMs), collectively forming the TED Library of HMMs (TEDLH). TEDLH enables sensitive sequence and profile searches, supporting systematic exploration of protein domain families and their evolutionary relationships. RESULTS: TEDLH links 934,186 domain HMMs to experimentally determined CATH-PDB structures through direct (primary) and transitive (secondary and tertiary) relationships. Fewer than half of TEDLH HMMs are directly linked to a CATH-PDB domain; the remaining models are connected through transitive relationships. These transitive links extend coverage into more divergent regions of sequence space and better represent CATH superfamily diversity. HMM-HMM comparisons within CATH superfamily 3.30.70.100 illustrate how transitive relationships expand sequence coverage. In this superfamily, 5640 TEDLH HMMs are connected to 173 CATH-PDB representatives. Primary, secondary, and tertiary relationships progressively capture more divergent sequences (pairwise sequence identity <20%) that retain structural similarity (TM-score ≥0.6) and a conserved two-layer α/β sandwich core fold. All-against-all HMM-HMM comparisons across TEDLH also reveal sequence similarities across the CATH hierarchy (cross-hits). At low query coverage (<50%), cross-hits are more frequent between CATH classes, architectures and topologies, whereas at higher coverage thresholds (≥70%) they predominantly occur between superfamilies. These cross-hits are not driven by superfamily size or sequence diversity and can provide guidance for CATH curation. As an example, analysis of cross-hits between superfamilies 2.170.130.30 and 3.10.20.30 reveals evolutionary relationships between these groups. AVAILABILITY: TEDLH is compatible with HH-suite3 and is available from FigShare https://doi.org/10.6084/m9.figshare.28531754 for local use.
Claudia Alvarez-Carreño, Anton S. Petrov, Vaishali P. Waman, Ian Sillitoe, Christine A. Orengo
Bioinform.5
2024 Chainsaw: protein domain segmentation with fully convolutional neural networks
abstract
MOTIVATION: Protein domains are fundamental units of protein structure and play a pivotal role in understanding folding, function, evolution, and design. The advent of accurate structure prediction techniques has resulted in an influx of new structural data, making the partitioning of these structures into domains essential for inferring evolutionary relationships and functional classification. RESULTS: This article presents Chainsaw, a supervised learning approach to domain parsing that achieves accuracy that surpasses current state-of-the-art methods. Chainsaw uses a fully convolutional neural network which is trained to predict the probability that each pair of residues is in the same domain. Domain predictions are then derived from these pairwise predictions using an algorithm that searches for the most likely assignment of residues to domains given the set of pairwise co-membership probabilities. Chainsaw matches CATH domain annotations in 78% of protein domains versus 72% for the next closest method. When predicting on AlphaFold models, expert human evaluators were twice as likely to prefer Chainsaw's predictions versus the next best method. AVAILABILITY AND IMPLEMENTATION: github.com/JudeWells/Chainsaw.
Jude Wells, Alex Hawkins-Hooker, Nicola Bordin, Ian Sillitoe, Brooks Paige, Christine A. Orengo
Bioinform.6
2023 CATHe: detection of remote homologues for CATH superfamilies using embeddings from protein language models
abstract
MOTIVATION: CATH is a protein domain classification resource that exploits an automated workflow of structure and sequence comparison alongside expert manual curation to construct a hierarchical classification of evolutionary and structural relationships. The aim of this study was to develop algorithms for detecting remote homologues missed by state-of-the-art hidden Markov model (HMM)-based approaches. The method developed (CATHe) combines a neural network with sequence representations obtained from protein language models. It was assessed using a dataset of remote homologues having less than 20% sequence identity to any domain in the training set. RESULTS: The CATHe models trained on 1773 largest and 50 largest CATH superfamilies had an accuracy of 85.6 ± 0.4% and 98.2 ± 0.3%, respectively. As a further test of the power of CATHe to detect more remote homologues missed by HMMs derived from CATH domains, we used a dataset consisting of protein domains that had annotations in Pfam, but not in CATH. By using highly reliable CATHe predictions (expected error rate <0.5%), we were able to provide CATH annotations for 4.62 million Pfam domains. For a subset of these domains from Homo sapiens, we structurally validated 90.86% of the predictions by comparing their corresponding AlphaFold2 structures with structures from the CATH superfamilies to which they were assigned. AVAILABILITY AND IMPLEMENTATION: The code for the developed models is available on https://github.com/vam-sin/CATHe, and the datasets developed in this study can be accessed on https://zenodo.org/record/6327572. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Vamsi Nallapareddy, Nicola Bordin, Ian Sillitoe, Michael Heinzinger, Maria Littmann, Vaishali P. Waman, Neeladri Sen, Burkhard Rost, Christine A. Orengo
Bioinform.9
2022 Characterizing and explaining the impact of disease-associated mutations in proteins without known structures or structural homologs
abstract
Mutations in human proteins lead to diseases. The structure of these proteins can help understand the mechanism of such diseases and develop therapeutics against them. With improved deep learning techniques, such as RoseTTAFold and AlphaFold, we can predict the structure of proteins even in the absence of structural homologs. We modeled and extracted the domains from 553 disease-associated human proteins without known protein structures or close homologs in the Protein Databank. We noticed that the model quality was higher and the Root mean square deviation (RMSD) lower between AlphaFold and RoseTTAFold models for domains that could be assigned to CATH families as compared to those which could only be assigned to Pfam families of unknown structure or could not be assigned to either. We predicted ligand-binding sites, protein-protein interfaces and conserved residues in these predicted structures. We then explored whether the disease-associated missense mutations were in the proximity of these predicted functional sites, whether they destabilized the protein structure based on ddG calculations or whether they were predicted to be pathogenic. We could explain 80% of these disease-associated mutations based on proximity to functional sites, structural destabilization or pathogenicity. When compared to polymorphisms, a larger percentage of disease-associated missense mutations were buried, closer to predicted functional sites, predicted as destabilizing and pathogenic. Usage of models from the two state-of-the-art techniques provide better confidence in our predictions, and we explain 93 additional mutations based on RoseTTAFold models which could not be explained based solely on AlphaFold models.
Neeladri Sen, Ivan Anishchenko, Nicola Bordin, Ian Sillitoe, Sameer Velankar, David Baker 0001, Christine A. Orengo
Briefings Bioinform.7
2022 Srinivasan (1962-2021) in Bioinformatics and beyond
abstract
Dear Editor, Last year the Bioinformatics community lost one of its pioneers, a scientist renowned for his talent, creativity and rigour but also for his commitment to supporting his research community and particularly the young scientists he trained. He was an inspiring role model for his field and a scientist who will be remembered very fondly by his many friends in the community for his warmth, humour and kindness. For more than three decades Srinivasan developed timely and novel computational strategies for analyzing proteins and was regarded in high esteem internationally for the insights he provided and the resources he established based on the underpinning concepts. His discoveries cover many areas fundamental to structural biology and pathogen research. Although he was a computational scientist, he worked closely with experimental groups to maximize the impact of his research. He published more than 300 papers, with nearly 10 000 citations altogether. Srinivasan joined the faculty of the Molecular Biophysics Unit, Indian Institute of Science, Bangalore in 1998, after leaving the Madras Biophysics Group (he did his Masters studies from 1982-84). He acquired his PhD degree in the Molecular Biophysics Department (the same Department where he later worked as a faculty member) within the GN Ramachandran school of peptide and peptide stereochemistry. His postdoctoral tenure was in Prof. Sir Tom Blundell’s laboratory (1991–1998), Birkbeck College, UK, with a brief stint in Prof. Mike Waterfield’s laboratory at the Ludwig Institute for Cancer Research, UK. He arrived in London as a seemingly shy young man, but it soon became clear that he was a real expert in protein structures and thought very deeply about their evolution. During these times, his research was largely focused on homology modelling (Johnson et al., 1994) and the study of proteins involved in signal transduction (e.g. Srinivasan et al., 1994, 1996). After coming back to MBU, he headed the ‘Proteins: structure, function and evolutionary’ group. He made major contributions to the understanding the structure and functions of proteins, particularly on protein kinases in a wide range of model organisms (e.g. Krupa and Srinivasan, 2002; Krupa et al., 2004a, b). Specifically, his lab was focused on computational genomics, bioinformatics and structural biology, particularly involving the relationships between protein structure, function and interactions, including protein–protein interactions, cellular signal transduction and biological pathways. Srinivasan’s interest in protein families and protein evolution drove research into strategies for improving multiple alignments of relatives and for better characterizing phylogenetic relationships (which resulted in many useful resources like SUPFAM, MulPSSM, PALI and DoSA). It also drove the design of methods to detect extremely remote homologues, which have been valuable for extending structural and functional annotation of genomes. Srinivasan’s group also showed that sequence-based connections of distantly related proteins can be enabled through the design of artificial sequences (Mudgal et al., 2014). This work was highly innovative and can help to bring valuable annotations for pathogen proteins, which are typically difficult to characterize by more conventional, less-sensitive strategies. His strategies allowed a much deeper characterization of fold space to inform protein engineering. Srinivasan is also highly renowned for his analyses of how changes in the protein structure and sequence impact function. Some protein families, like the kinases, were a major focus of his research and gave him much international acclaim. He studied kinases for >20 years and contributed numerous insights important for understanding their mechanisms and for enabling drug design. For example, structural fluctuations, classifications based on key functional site properties (Kalaivani et al., 2018), mechanisms of stabilization of their key functional sites through specific residue interactions. He also characterized the ways in which domain partnerships modify structure (Vishwanath et al., 2018), kinase functionality and characterized how splicing extends the kinase functional repertoire. This large family is implicated in many human diseases, including cancer, and these discoveries have informed drug design. However, the biological role of proteins is determined by their interactions and Srinivasan applied his precise analytical skills in this arena, too, revealing key insights into the properties of the interfaces involved in assembling protein complexes. He produced a substantial body of very rigorous studies, including analyses of the characteristics of transient complexes and the effect of protein associations on global structural dynamics. He robustly captured this knowledge in the PIC protein interactions calculator (Tina et al., 2007), a valuable tool that is freely available to biologists and very popular among researchers to obtain structural data on various non-covalent interactions within a protein or between proteins in a complex. An important application of these methods was the characterization of interactions between viral proteins and their host proteins, which provided key data for understanding pathogenicity and enabling drug design. For example, Srinivasan performed various studies characterizing toxin–antitoxin systems (Tandon et al., 2019), protein interactions between human erythrocytes and Plasmodium falciparum and Helicobacter pylori and human. Photo taken at the fifth IIT Madras-Tokyo Tech joint symposium on ‘Current Trends in Bioinformatics: Big Data Analysis, Machine Learning and Drug Design’ with leading Bioinformatics scientists in India (March 2020). From left to right D. Velmurugan, N. Manoj, S. Selvaraj, P.K. Ponnuswamy, G.P.S. Raghava, K. Veluraja, Shandar Ahmad, M. Michael Gromiha, N. Srinivasan, R. Sowdhamini Photo taken at the fifth IIT Madras-Tokyo Tech joint symposium on ‘Current Trends in Bioinformatics: Big Data Analysis, Machine Learning and Drug Design’ with leading Bioinformatics scientists in India (March 2020). From left to right D. Velmurugan, N. Manoj, S. Selvaraj, P.K. Ponnuswamy, G.P.S. Raghava, K. Veluraja, Shandar Ahmad, M. Michael Gromiha, N. Srinivasan, R. Sowdhamini In collaboration with multiple laboratories, Srini’s group studied fascinating biological systems, including protein assemblies of ribosomes and spliceosomes (Bhat et al., 2015; Pudi et al., 2003; Yazhini et al., 2022) and developed powerful computational tools for studying structures of large assemblies derived from cryo-electron microscopy (Joseph et al., 2016; Rakesh et al., 2016). As is true of several structural bioinformaticians, his group relied on publicly available structural data and were concerned with the quality of protein–ligand data, deposited through X-ray or cryo-EM studies (Chakraborti et al., 2021). He was always excited to discuss the Ramachandran map. Last year, he attended the fifth IIT Madras—Tokyo Tech joint symposium on Bioinformatics and his lecture on the Ramachandran map was fascinating. Using modern computational tools and along with late Prof. C. Ramakrishnan (his PhD mentor and the student behind the original Ramachandran map) and one of his students Ashraya Ravikumar, he re-examined the classical and renowned Ramachandran map. They clearly demonstrated that it is possible to consider deviations in the ‘allowed’ regions within this map by considering slight deviations in internal parameters from ideal values of the peptide bond (Ravikumar et al., 2019). Srinivasan’s group also participated in several consortia such as the Open Source Drug Discovery program [with the groups of Prof. Tom Blundell (University of Cambridge, UK), Nagasuma Chandra (Indian Institute of Science, India) and Sowdhamini (National Centre for Biological Sciences, India)], the UKIERI study of protein assemblies [with the groups of Profs. Jim Warwicker (University of Manchester, UK), Pinak Chakrabarti (Bose Institute, India), Nagasuma Chandra and Sowdhamini)] and collaborations such as the Indo-French CEFIPRA project on protein alphabets [with Dr. Alexandre de Brevern (INSERM Paris, France) and Dr. Bernard Offmann (University of Nantes, France)], and the Centre for Excellence on protein-protein interactions [with Profs. Sowdhamini and Satyajit Mayor (National Centre for Biological Sciences, India) and Nagasuma Chandra (Indian Institute of Science, India)] and toxin-antitoxin systems [with Prof. Raghavan Varadarajan (Indian Institute of Science, India)]. Throughout his career Srinivasan applied his knowledge, data and computational tools to characterize the protein structures, functions and virus–host interactions of multiple pathogenic bacteria affecting human health, including mycobacterial pathogens (e.g. Mtb), malaria, H. Pylori, Dengue and several gut pathogens. Understanding the critical residues in the protein interface is essential for drug design to reduce infection and pathogenicity. For many years he collaborated with the group of Professor Tom Blundell in Cambridge, UK. As well as detailed analyses, he established the SInCRe structural interactome resource for Mtb in 2015 (Metri et al., 2015), which contributed to studies on the repurposing of drugs for this pathogen. His tools have been applied in a number of medical contexts with promising clinical results. His group also applied docking tools to FDA-approved drugs to SARS-CoV2 (Chakraborti et al., 2020) and his most recent work on inhibitors for the main protease of SARS-CoV2 led to compounds already in clinical trials. Prof. Srinivasan made significant contributions to Bioinformatics and his whole-hearted involvement in scientific activities will not be forgotten. As a scientist, Srinivasan was very highly focused and meticulous. He always set high standards—whether in creating high-quality datasets or in his interpretations of data or in responding to reviewers’ comments. As well as being a multitalented and well-known researcher, Srinivasan actively participated in many university and external committees, commented on PhD theses, and delivered popular and invited lectures in most of the leading conferences in India. He was an elected fellow in all the three major academies in India (Indian National Science Academy, New Delhi, National Academy of Sciences, Allahabad and Indian Academy of Sciences, Bangalore). He also received the most prestigious awards in India including Shanti Swarup Bhatnagar Prize for Science and Technology from Council of Scientific and Industrial Research, National Bioscience Award from the Department of Biotechnology and J.C. Bose National Fellowship from the Department of Science and Technology, Government of India. Srinivasan had the special ability to cordially relate with others he respected and had a very positive attitude towards the work of his fellow researchers. He was a faithful chairperson of the Department (serving between 2018 and 2020) and always supported and wished his younger colleagues to do well. He remained active even when he was critically ill. During this time he still managed to publish around 10 papers, enable six of his lab colleagues to reach higher positions and also attended to multiple student-thesis-related matters. He was very enthusiastic about his research on protein structures and his ability to explain major concepts in a simple accessible manner was extremely impressive. He had a passion for naming his students with ‘amino acids’, each with a background story and spent considerable time with his students in the midst of his busy schedule. He always encouraged young researchers and provided valuable advice for their research. He also had an uncanny enthusiasm and ability to make sure that the people around him felt included and important—he would not hesitate to talk to prospective students, spend a long time on discussions with visitors and provide his undivided attention on work discussions with his students. Many of his students travelled widely and benefitted laboratories and science throughout the world. His students made a large impact wherever they went because their knowledge was always deep and impressive and they showed great enthusiasm for their work, mirroring their mentor. They often liked to discuss their work in detail and place it in the wider context of global knowledge. Although based in India for most of his career, Srinivasan travelled widely and has had an impact on many scientists involved in protein structure analysis. He was also a great host for visitors, ensuring their well-being and spending time discussing their work and ideas. Srinivasan had been a very special and unusual personality—with unlimited affection and love for the people around him. He was able to sense people in trouble and would often go out of his way to help them. His smiling face is not forgettable at any time and evidenced a deeply contented person, very proud of his family, his students and his science and always happy to discuss anything to do with proteins. His positivity and passion for science were infectious. His too-early passing is a great loss to science and to everyone who knew him. Financial Support: none declared. Conflict of Interest: The authors declare that there are no conflicts of interest.
M. Michael Gromiha, Christine A. Orengo, Ramanathan Sowdhamini, Janet M. Thornton
Bioinform.2
2022 Characterizing domain-specific open educational resources by linking ISCB Communities of Special Interest to Wikipedia
abstract
MOTIVATION: Wikipedia is one of the most important channels for the public communication of science and is frequently accessed as an educational resource in computational biology. Joint efforts between the International Society for Computational Biology (ISCB) and the Computational Biology taskforce of WikiProject Molecular Biology (a group of expert Wikipedia editors) have considerably improved computational biology representation on Wikipedia in recent years. However, there is still an urgent need for further improvement in quality, especially when compared to related scientific fields such as genetics and medicine. Facilitating involvement of members from ISCB Communities of Special Interest (COSIs) would improve a vital open education resource in computational biology, additionally allowing COSIs to provide a quality educational resource highly specific to their subfield. RESULTS: We generate a list of around 1500 English Wikipedia articles relating to computational biology and describe the development of a binary COSI-Article matrix, linking COSIs to relevant articles and thereby defining domain-specific open educational resources. Our analysis of the COSI-Article matrix data provides a quantitative assessment of computational biology representation on Wikipedia against other fields and at a COSI-specific level. Furthermore, we conducted similarity analysis and subsequent clustering of COSI-Article data to provide insight into potential relationships between COSIs. Finally, based on our analysis, we suggest courses of action to improve the quality of computational biology representation on Wikipedia.
Alastair M. Kilpatrick, Farzana Rahman, Audra Anjum, Sayane Shome, K. M. Salim Andalib, Shrabonti Banik, Sanjana F. Chowdhury, Peter Coombe, Yesid Cuesta Astroz, J. Maxwell Douglas, Pradeep Eranti, Aleyna D. Kiran, Sachendra Kumar, Hyeri Lim, Valentina Lorenzi, Tiago Lubiana, Sakib Mahmud, Rafael Puche, Agnieszka Rybarczyk, Syed Muktadir Al Sium, David Twesigomwe, Tomasz Zok, Christine A. Orengo, Iddo Friedberg, Janet Kelso, Lonnie R. Welch
Bioinform.23
2022 Assigning protein function from domain-function associations using DomFun
abstract
BACKGROUND: Protein function prediction remains a key challenge. Domain composition affects protein function. Here we present DomFun, a Ruby gem that uses associations between protein domains and functions, calculated using multiple indices based on tripartite network analysis. These domain-function associations are combined at the protein level, to generate protein-function predictions. RESULTS: We analysed 16 tripartite networks connecting homologous superfamily and FunFam domains from CATH-Gene3D with functional annotations from the three Gene Ontology (GO) sub-ontologies, KEGG, and Reactome. We validated the results using the CAFA 3 benchmark platform for GO annotation, finding that out of the multiple association metrics and domain datasets tested, Simpson index for FunFam domain-function associations combined with Stouffer's method leads to the best performance in almost all scenarios. We also found that using FunFams led to better performance than superfamilies, and better results were found for GO molecular function compared to GO biological process terms. DomFun performed as well as the highest-performing method in certain CAFA 3 evaluation procedures in terms of [Formula: see text] and [Formula: see text] We also implemented our own benchmark procedure, Pathway Prediction Performance (PPP), which can be used to validate function prediction for additional annotations sources, such as KEGG and Reactome. Using PPP, we found similar results to those found with CAFA 3 for GO, moreover we found good performance for the other annotation sources. As with CAFA 3, Simpson index with Stouffer's method led to the top performance in almost all scenarios. CONCLUSIONS: DomFun shows competitive performance with other methods evaluated in CAFA 3 when predicting proteins function with GO, although results vary depending on the evaluation procedure. Through our own benchmark procedure, PPP, we have shown it can also make accurate predictions for KEGG and Reactome. It performs best when using FunFams, combining Simpson index derived domain-function associations using Stouffer's method. The tool has been implemented so that it can be easily adapted to incorporate other protein features, such as domain data from other sources, amino acid k-mers and motifs. The DomFun Ruby gem is available from https://rubygems.org/gems/DomFun . Code maintained at https://github.com/ElenaRojano/DomFun . Validation procedure scripts can be found at https://github.com/ElenaRojano/DomFun_project .
Elena Rojano, Fernando Moreno Jabato, James Richard Perkins, José Córdoba-Caballero, Federico García-Criado, Ian Sillitoe, Christine A. Orengo, Juan Garcia Ranea, Pedro Seoane
BMC Bioinform.7
2022 Dissecting peripheral protein-membrane interfaces
abstract
Peripheral membrane proteins (PMPs) include a wide variety of proteins that have in common to bind transiently to the chemically complex interfacial region of membranes through their interfacial binding site (IBS). In contrast to protein-protein or protein-DNA/RNA interfaces, peripheral protein-membrane interfaces are poorly characterized. We collected a dataset of PMP domains representative of the variety of PMP functions: membrane-targeting domains (Annexin, C1, C2, discoidin C2, PH, PX), enzymes (PLA, PLC/D) and lipid-transfer proteins (START). The dataset contains 1328 experimental structures and 1194 AphaFold models. We mapped the amino acid composition and structural patterns of the IBS of each protein in this dataset, and evaluated which were more likely to be found at the IBS compared to the rest of the domains' accessible surface. In agreement with earlier work we find that about two thirds of the PMPs in the dataset have protruding hydrophobes (Leu, Ile, Phe, Tyr, Trp and Met) at their IBS. The three aromatic amino acids Trp, Tyr and Phe are a hallmark of PMPs IBS regardless of whether they protrude on loops or not. This is also the case for lysines but not arginines suggesting that, unlike for Arg-rich membrane-active peptides, the less membrane-disruptive lysine is preferred in PMPs. Another striking observation was the over-representation of glycines at the IBS of PMPs compared to the rest of their surface, possibly procuring IBS loops a much-needed flexibility to insert in-between membrane lipids. The analysis of the 9 superfamilies revealed amino acid distribution patterns in agreement with their known functions and membrane-binding mechanisms. Besides revealing novel amino acids patterns at protein-membrane interfaces, our work contributes a new PMP dataset and an analysis pipeline that can be further built upon for future studies of PMPs properties, or for developing PMPs prediction tools using for example, machine learning approaches.
Thibault Tubiana, Ian Sillitoe, Christine A. Orengo, Nathalie Reuter
PLoS Comput. Biol.3
2021 The impact of structural bioinformatics tools and resources on SARS-CoV-2 research and therapeutic strategies
abstract
SARS-CoV-2 is the causative agent of COVID-19, the ongoing global pandemic. It has posed a worldwide challenge to human health as no effective treatment is currently available to combat the disease. Its severity has led to unprecedented collaborative initiatives for therapeutic solutions against COVID-19. Studies resorting to structure-based drug design for COVID-19 are plethoric and show good promise. Structural biology provides key insights into 3D structures, critical residues/mutations in SARS-CoV-2 proteins, implicated in infectivity, molecular recognition and susceptibility to a broad range of host species. The detailed understanding of viral proteins and their complexes with host receptors and candidate epitope/lead compounds is the key to developing a structure-guided therapeutic design. Since the discovery of SARS-CoV-2, several structures of its proteins have been determined experimentally at an unprecedented speed and deposited in the Protein Data Bank. Further, specialized structural bioinformatics tools and resources have been developed for theoretical models, data on protein dynamics from computer simulations, impact of variants/mutations and molecular therapeutics. Here, we provide an overview of ongoing efforts on developing structural bioinformatics tools and resources for COVID-19 research. We also discuss the impact of these resources and structure-based studies, to understand various aspects of SARS-CoV-2 infection and therapeutic development. These include (i) understanding differences between SARS-CoV-2 and SARS-CoV, leading to increased infectivity of SARS-CoV-2, (ii) deciphering key residues in the SARS-CoV-2 involved in receptor-antibody recognition, (iii) analysis of variants in host proteins that affect host susceptibility to infection and (iv) analyses facilitating structure-based drug and vaccine design against SARS-CoV-2.
Vaishali P. Waman, Neeladri Sen, Mihaly Varadi, Antoine Daina, Shoshana J. Wodak, Vincent Zoete, Sameer Velankar, Christine A. Orengo
Briefings Bioinform.8
2021 CATH functional families predict functional sites in proteins
abstract
MOTIVATION: Identification of functional sites in proteins is essential for functional characterization, variant interpretation and drug design. Several methods are available for predicting either a generic functional site, or specific types of functional site. Here, we present FunSite, a machine learning predictor that identifies catalytic, ligand-binding and protein-protein interaction functional sites using features derived from protein sequence and structure, and evolutionary data from CATH functional families (FunFams). RESULTS: FunSite's prediction performance was rigorously benchmarked using cross-validation and a holdout dataset. FunSite outperformed other publicly available functional site prediction methods. We show that conserved residues in FunFams are enriched in functional sites. We found FunSite's performance depends greatly on the quality of functional site annotations and the information content of FunFams in the training data. Finally, we analyze which structural and evolutionary features are most predictive for functional sites. AVAILABILITYAND IMPLEMENTATION: https://github.com/UCL/cath-funsite-predictor. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Sayoni Das, Harry M. Scholes, Neeladri Sen, Christine A. Orengo
Bioinform.4
2021 Clustering FunFams using sequence embeddings improves EC purity
abstract
MOTIVATION: Classifying proteins into functional families can improve our understanding of protein function and can allow transferring annotations within one family. For this, functional families need to be 'pure', i.e., contain only proteins with identical function. Functional Families (FunFams) cluster proteins within CATH superfamilies into such groups of proteins sharing function. 11% of all FunFams (22 830 of 203 639) contain EC annotations and of those, 7% (1526 of 22 830) have inconsistent functional annotations. RESULTS: We propose an approach to further cluster FunFams into functionally more consistent sub-families by encoding their sequences through embeddings. These embeddings originate from language models transferring knowledge gained from predicting missing amino acids in a sequence (ProtBERT) and have been further optimized to distinguish between proteins belonging to the same or a different CATH superfamily (PB-Tucker). Using distances between embeddings and DBSCAN to cluster FunFams and identify outliers, doubled the number of pure clusters per FunFam compared to random clustering. Our approach was not limited to FunFams but also succeeded on families created using sequence similarity alone. Complementing EC annotations, we observed similar results for binding annotations. Thus, we expect an increased purity also for other aspects of function. Our results can help generating FunFams; the resulting clusters with improved functional consistency allow more reliable inference of annotations. We expect this approach to succeed equally for any other grouping of proteins by their phenotypes. AVAILABILITY AND IMPLEMENTATION: Code and embeddings are available via GitHub: https://github.com/Rostlab/FunFamsClustering. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Maria Littmann, Nicola Bordin, Michael Heinzinger, Konstantin Schütze, Christian Dallago, Christine A. Orengo, Burkhard Rost
Bioinform.6
2021 Biological impact of mutually exclusive exon switching
abstract
Alternative splicing can expand the diversity of proteomes. Homologous mutually exclusive exons (MXEs) originate from the same ancestral exon and result in polypeptides with similar structural properties but altered sequence. Why would some genes switch homologous exons and what are their biological impact? Here, we analyse the extent of sequence, structural and functional variability in MXEs and report the first large scale, structure-based analysis of the biological impact of MXE events from different genomes. MXE-specific residues tend to map to single domains, are highly enriched in surface exposed residues and cluster at or near protein functional sites. Thus, MXE events are likely to maintain the protein fold, but alter specificity and selectivity of protein function. This comprehensive resource of MXE events and their annotations is available at: http://gene3d.biochem.ucl.ac.uk/mxemod/. These findings highlight how small, but significant changes at critical positions on a protein surface are exploited in evolution to alter function.
Su Datt Lam, M. Madan Babu, Jonathan G. Lees, Christine A. Orengo
PLoS Comput. Biol.4
2020 The ELIXIR Core Data Resources: fundamental infrastructure for the life sciences
abstract
SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Rachel Drysdale, Charles E. Cook, Robert Petryszak, Vivienne Baillie Gerritsen, Mary Barlow, Elisabeth Gasteiger, Franziska Gruhl, Jerry Lanfear, Rodrigo Lopez, Nicole Redaschi, Heinz Stockinger, Daniel Teixeira, Aravind Venkatesan, Alex Bateman, Alan J. Bridge, Guy Cochrane, Robert D. Finn, Frank Oliver Glöckner, Marc Hanauer, Thomas M. Keane, Luana Licata, Per Oksvold, Sandra E. Orchard, Christine A. Orengo, Helen E. Parkinson, Bengt Persson, Pablo Porras, Jordi Rambla De Argila, Ana Rath, Charlotte Rodwell, Ugis Sarkans, Dietmar Schomburg, Ian Sillitoe, J. Dylan Spalding, Mathias Uhlen, Sameer Velankar, Juan Antonio Vizcaíno, Kalle von Feilitzen, Christian von Mering, Andy Yates, Niklas Blomberg, Christine Durinx, Johanna R. McEntyre
Bioinform.26
2019 FunFam protein families improve residue level molecular function prediction
abstract
BACKGROUND: The CATH database provides a hierarchical classification of protein domain structures including a sub-classification of superfamilies into functional families (FunFams). We analyzed the similarity of binding site annotations in these FunFams and incorporated FunFams into the prediction of protein binding residues. RESULTS: FunFam members agreed, on average, in 36.9 ± 0.6% of their binding residue annotations. This constituted a 6.7-fold increase over randomly grouped proteins and a 1.2-fold increase (1.1-fold on the same dataset) over proteins with the same enzymatic function (identical Enzyme Commission, EC, number). Mapping de novo binding residue prediction methods (BindPredict-CCS, BindPredict-CC) onto FunFam resulted in consensus predictions for those residues that were aligned and predicted alike (binding/non-binding) within a FunFam. This simple consensus increased the F1-score (for binding) 1.5-fold over the original prediction method. Variation of the threshold for how many proteins in the consensus prediction had to agree provided a convenient control of accuracy/precision and coverage/recall, e.g. reaching a precision as high as 60.8 ± 0.4% for a stringent threshold. CONCLUSIONS: The FunFams outperformed even the carefully curated EC numbers in terms of agreement of binding site residues. Additionally, we assume that our proof-of-principle through the prediction of protein binding residues will be relevant for many other solutions profiting from FunFams to infer functional information at the residue level.
Linus Scheibenreif, Maria Littmann, Christine A. Orengo, Burkhard Rost
BMC Bioinform.3
2017 ISCB's initial reaction to New England Journal of Medicine editorial on data sharing
abstract
The recent editorial by Dr Longo and Dr Drazen in the New England Journal of Medicine (Longo and Drazen, 2016) has stirred up quite a bit of controversy. As Executive Officers of the International Society of Computational Biology, Inc. (ISCB), we express our deep concern about the restrictive and potentially damaging opinions voiced in this editorial, and while ISCB works to write a detailed response, we felt it necessary to promptly address the editorial with this reaction. Although some of the concerns voiced by the authors of the editorial are worth considering, large parts of the statement purport an obsolete view of hegemony over data that is neither in line with today’s spirit of open access nor furthering an atmosphere where the potential of data can be fully realized. ISCB acknowledges that the additional comment on the editorial (Drazen, 2016) eases some of the polemics unfortunately without addressing some of the core issues. We still feel, however, that we need to contrast the opinion voiced in the editorial with what we consider the axioms of our scientific society, statements that lead into a fruitful future of data-driven science: Data produced with public money should be public in benefit of the science and society Restrictions on the use of public data hamper science and slow progress Open data is the best way to combat fraud and misinterpretations Current large data collections proceed from many sources, are continually accumulated, and require a variety of analytical approaches. Data generation and data analysis overlap in time and are continually updated with new data sets produced by new techniques and new analysis methodologies. Furthermore, in many cases current science functions in consortia in which scientists collaborate toward common goals while preserving their own scientific objectives. Dividing scientists into data providers and data analysts is simplistic and gives a misleading impression of the actual state of biological and biomedical science. ISCB very much supports collaboration between disciplines, including experimental and clinical as well as bioinformatics, as the best way forward to address complex biological problems. But this collaboration cannot be based on imposed restrictions to data access and cannot be contained in professional silos. (The use of expressions such as ‘research parasites’ clearly does not help.) Many bio-communities have made significant progress by endorsing open data policies and, gratefully, public funding agencies have connected to the spirit that they are distributing taxpayers’ money to science and that, therefore, the data that are generated in the course belong to the public. It is, perhaps, natural that some areas of biomedical research are slow in adopting these policies. History and the confidential nature of the relevant data are surely among the reasons. However, in our opinion data hegemony is another, a reason that has to be overcome. The sooner these barriers to progress are removed the sooner the patients will benefit from the current flourishing of biomedical research. Conflict of Interest: none declared.
Bonnie Berger, Terry Gaasterland, Thomas Lengauer, Christine A. Orengo, Bruno Gaëta, Scott Markel, Alfonso Valencia
Bioinform.4
2017 Analysis of temporal transcription expression profiles reveal links between protein function and developmental stages of Drosophila melanogaster
abstract
Accurate gene or protein function prediction is a key challenge in the post-genome era. Most current methods perform well on molecular function prediction, but struggle to provide useful annotations relating to biological process functions due to the limited power of sequence-based features in that functional domain. In this work, we systematically evaluate the predictive power of temporal transcription expression profiles for protein function prediction in Drosophila melanogaster. Our results show significantly better performance on predicting protein function when transcription expression profile-based features are integrated with sequence-derived features, compared with the sequence-derived features alone. We also observe that the combination of expression-based and sequence-based features leads to further improvement of accuracy on predicting all three domains of gene function. Based on the optimal feature combinations, we then propose a novel multi-classifier-based function prediction method for Drosophila melanogaster proteins, FFPred-fly+. Interpreting our machine learning models also allows us to identify some of the underlying links between biological processes and developmental stages of Drosophila melanogaster.
Cen Wan, Jonathan G. Lees, Federico Minneci, Christine A. Orengo, David T. Jones
PLoS Comput. Biol.4
2016 Functional classification of CATH superfamilies: a domain-based approach for protein function annotation
abstract
Bioinformatics (2015) 31(21):3460–3467. Author, Sayoni Das, would like to report a typographical error in Equation (2) in the above article.
Sayoni Das, David A. Lee, Ian Sillitoe, Natalie L. Willhoft, Jonathan G. Lees, Christine A. Orengo
Bioinform.6
2016 ISCB's Initial Reaction to The New England Journal of Medicine Editorial on Data Sharing
abstract
This message is a response from the ISCB in light of the recent the New England Journal of Medicine (NEJM) editorial around data sharing.
Bonnie Berger, Terry Gaasterland, Thomas Lengauer, Christine A. Orengo, Bruno Gaëta, Scott Markel, Alfonso Valencia
PLoS Comput. Biol.4
2016 Novel Computational Protocols for Functionally Classifying and Characterising Serine Beta-Lactamases
abstract
Beta-lactamases represent the main bacterial mechanism of resistance to beta-lactam antibiotics and are a significant challenge to modern medicine. We have developed an automated classification and analysis protocol that exploits structure- and sequence-based approaches and which allows us to propose a grouping of serine beta-lactamases that more consistently captures and rationalizes the existing three classification schemes: Classes, (A, C and D, which vary in their implementation of the mechanism of action); Types (that largely reflect evolutionary distance measured by sequence similarity); and Variant groups (which largely correspond with the Bush-Jacoby clinical groups). Our analysis platform exploits a suite of in-house and public tools to identify Functional Determinants (FDs), i.e. residue sites, responsible for conferring different phenotypes between different classes, different types and different variants. We focused on Class A beta-lactamases, the most highly populated and clinically relevant class, to identify FDs implicated in the distinct phenotypes associated with different Class A Types and Variants. We show that our FunFHMMer method can separate the known beta-lactamase classes and identify those positions likely to be responsible for the different implementations of the mechanism of action in these enzymes. Two novel algorithms, ASSP and SSPA, allow detection of FD sites likely to contribute to the broadening of the substrate profiles. Using our approaches, we recognise 151 Class A types in UniProt. Finally, we used our beta-lactamase FunFams and ASSP profiles to detect 4 novel Class A types in microbiome samples. Our platforms have been validated by literature studies, in silico analysis and some targeted experimental verification. Although developed for the serine beta-lactamases they could be used to classify and analyse any diverse protein superfamily where sub-families have diverged over both long and short evolutionary timescales.
David A. Lee, Sayoni Das, Natalie L. Willhoft, Dragana Dobrijevic, John Ward, Christine A. Orengo
PLoS Comput. Biol.6
2015 Functional classification of CATH superfamilies: a domain-based approach for protein function annotation
abstract
MOTIVATION: Computational approaches that can predict protein functions are essential to bridge the widening function annotation gap especially since <1.0% of all proteins in UniProtKB have been experimentally characterized. We present a domain-based method for protein function classification and prediction of functional sites that exploits functional sub-classification of CATH superfamilies. The superfamilies are sub-classified into functional families (FunFams) using a hierarchical clustering algorithm supervised by a new classification method, FunFHMMer. RESULTS: FunFHMMer generates more functionally coherent groupings of protein sequences than other domain-based protein classifications. This has been validated using known functional information. The conserved positions predicted by the FunFams are also found to be enriched in known functional residues. Moreover, the functional annotations provided by the FunFams are found to be more precise than other domain-based resources. FunFHMMer currently identifies 110,439 FunFams in 2735 superfamilies which can be used to functionally annotate>16 million domain sequences. AVAILABILITY AND IMPLEMENTATION: All FunFam annotation data are made available through the CATH webpages (http://www.cathdb.info). The FunFHMMer webserver (http://www.cathdb.info/search/by_funfhmmer) allows users to submit query sequences for assignment to a CATH FunFam. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Sayoni Das, David A. Lee, Ian Sillitoe, Natalie L. Willhoft, Jonathan G. Lees, Christine A. Orengo
Bioinform.6
2015 FUN-L: gene prioritization for RNAi screens
abstract
MOTIVATION: Most biological processes remain only partially characterized with many components still to be identified. Given that a whole genome can usually not be tested in a functional assay, identifying the genes most likely to be of interest is of critical importance to avoid wasting resources. RESULTS: Given a set of known functionally related genes and using a state-of-the-art approach to data integration and mining, our Functional Lists (FUN-L) method provides a ranked list of candidate genes for testing. Validation of predictions from FUN-L with independent RNAi screens confirms that FUN-L-produced lists are enriched in genes with the expected phenotypes. In this article, we describe a website front end to FUN-L. AVAILABILITY AND IMPLEMENTATION: The website is freely available to use at http://funl.org
Jonathan G. Lees, Jean-Karim Hériché, Ian Morilla, José María Fernández 0001, Priit Adler, Martin Krallinger, Jaak Vilo, Alfonso Valencia, Jan Ellenberg, Juan Garcia Ranea, Christine A. Orengo
Bioinform.11
2013 Protein function prediction using domain families
abstract
Here we assessed the use of domain families for predicting the functions of whole proteins. These 'functional families' (FunFams) were derived using a protocol that combines sequence clustering with supervised cluster evaluation, relying on available high-quality Gene Ontology (GO) annotation data in the latter step. In essence, the protocol groups domain sequences belonging to the same superfamily into families based on the GO annotations of their parent proteins. An initial test based on enzyme sequences confirmed that the FunFams resemble enzyme (domain) families much better than do families produced by sequence clustering alone. For the CAFA 2011 experiment, we further associated the FunFams with GO terms probabilistically. All target proteins were first submitted to domain superfamily assignment, followed by FunFam assignment and, eventually, function assignment. The latter included an integration step for multi-domain target proteins. The CAFA results put our domain-based approach among the top ten of 31 competing groups and 56 prediction methods, confirming that it outperforms simple pairwise whole-protein sequence comparisons.
Robert Rentzsch, Christine A. Orengo
BMC Bioinform.2
2012 Exploring the Evolution of Novel Enzyme Functions within Structurally Defined Protein Superfamilies
abstract
In order to understand the evolution of enzyme reactions and to gain an overview of biological catalysis we have combined sequence and structural data to generate phylogenetic trees in an analysis of 276 structurally defined enzyme superfamilies, and used these to study how enzyme functions have evolved. We describe in detail the analysis of two superfamilies to illustrate different paradigms of enzyme evolution. Gathering together data from all the superfamilies supports and develops the observation that they have all evolved to act on a diverse set of substrates, whilst the evolution of new chemistry is much less common. Despite that, by bringing together so much data, we can provide a comprehensive overview of the most common and rare types of changes in function. Our analysis demonstrates on a larger scale than previously studied, that modifications in overall chemistry still occur, with all possible changes at the primary level of the Enzyme Commission (E.C.) classification observed to a greater or lesser extent. The phylogenetic trees map out the evolutionary route taken within a superfamily, as well as all the possible changes within a superfamily. This has been used to generate a matrix of observed exchanges from one enzyme function to another, revealing the scale and nature of enzyme evolution and that some types of exchanges between and within E.C. classes are more prevalent than others. Surprisingly a large proportion (71%) of all known enzyme functions are performed by this relatively small set of 276 superfamilies. This reinforces the hypothesis that relatively few ancient enzymatic domain superfamilies were progenitors for most of the chemistry required for life.
Nicholas Furnham, Ian Sillitoe, Gemma L. Holliday, Alison L. Cuff, Roman A. Laskowski, Christine A. Orengo, Janet M. Thornton
PLoS Comput. Biol.6
2011 Characterization of pathogenic germline mutations in human Protein Kinases
abstract
BACKGROUND: Protein Kinases are a superfamily of proteins involved in crucial cellular processes such as cell cycle regulation and signal transduction. Accordingly, they play an important role in cancer biology. To contribute to the study of the relation between kinases and disease we compared pathogenic mutations to neutral mutations as an extension to our previous analysis of cancer somatic mutations. First, we analyzed native and mutant proteins in terms of amino acid composition. Secondly, mutations were characterized according to their potential structural effects and finally, we assessed the location of the different classes of polymorphisms with respect to kinase-relevant positions in terms of subfamily specificity, conservation, accessibility and functional sites. RESULTS: Pathogenic Protein Kinase mutations perturb essential aspects of protein function, including disruption of substrate binding and/or effector recognition at family-specific positions. Interestingly these mutations in Protein Kinases display a tendency to avoid structurally relevant positions, what represents a significant difference with respect to the average distribution of pathogenic mutations in other protein families. CONCLUSIONS: Disease-associated mutations display sound differences with respect to neutral mutations: several amino acids are specific of each mutation type, different structural properties characterize each class and the distribution of pathogenic mutations within the consensus structure of the Protein Kinase domain is substantially different to that for non-pathogenic mutations. This preferential distribution confirms previous observations about the functional and structural distribution of the controversial cancer driver and passenger somatic mutations and their use as a proxy for the study of the involvement of somatic mutations in cancer development.
José M. G. Izarzugaza, Lisa E. M. Hopcroft, Anja Baresic, Christine A. Orengo, Andrew C. R. Martin, Alfonso Valencia
BMC Bioinform.4
2010 A fast and automated solution for accurately resolving protein domain architectures
abstract
MOTIVATION: Accurate prediction of the domain content and arrangement in multi-domain proteins (which make up >65% of the large-scale protein databases) provides a valuable tool for function prediction, comparative genomics and studies of molecular evolution. However, scanning a multi-domain protein against a database of domain sequence profiles can often produce conflicting and overlapping matches. We have developed a novel method that employs heaviest weighted clique-finding (HCF), which we show significantly outperforms standard published approaches based on successively assigning the best non-overlapping match (Best Match Cascade, BMC). RESULTS: We created benchmark data set of structural domain assignments in the CATH database and a corresponding set of Hidden Markov Model-based domain predictions. Using these, we demonstrate that by considering all possible combinations of matches using the HCF approach, we achieve much higher prediction accuracy than the standard BMC method. We also show that it is essential to allow overlapping domain matches to a query in order to identify correct domain assignments. Furthermore, we introduce a straightforward and effective protocol for resolving any overlapping assignments, and producing a single set of non-overlapping predicted domains. AVAILABILITY AND IMPLEMENTATION: The new approach will be used to determine MDAs for UniProt and Ensembl, and made available via the Gene3D website: http://gene3d.biochem.ucl.ac.uk/Gene3D/. The software has been implemented in C++ and compiled for Linux: source code and binaries can be found at: ftp://ftp.biochem.ucl.ac.uk/pub/gene3d_data/DomainFinder3/ CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Corin Yeats, Oliver Redfern, Christine A. Orengo
Bioinform.3
2010 Finding the "Dark Matter" in Human and Yeast Protein Network Prediction and Modelling
abstract
Accurate modelling of biological systems requires a deeper and more complete knowledge about the molecular components and their functional associations than we currently have. Traditionally, new knowledge on protein associations generated by experiments has played a central role in systems modelling, in contrast to generally less trusted bio-computational predictions. However, we will not achieve realistic modelling of complex molecular systems if the current experimental designs lead to biased screenings of real protein networks and leave large, functionally important areas poorly characterised. To assess the likelihood of this, we have built comprehensive network models of the yeast and human proteomes by using a meta-statistical integration of diverse computationally predicted protein association datasets. We have compared these predicted networks against combined experimental datasets from seven biological resources at different level of statistical significance. These eukaryotic predicted networks resemble all the topological and noise features of the experimentally inferred networks in both species, and we also show that this observation is not due to random behaviour. In addition, the topology of the predicted networks contains information on true protein associations, beyond the constitutive first order binary predictions. We also observe that most of the reliable predicted protein associations are experimentally uncharacterised in our models, constituting the hidden or "dark matter" of networks by analogy to astronomical systems. Some of this dark matter shows enrichment of particular functions and contains key functional elements of protein networks, such as hubs associated with important functional areas like the regulation of Ras protein signal transduction in human cells. Thus, characterising this large and functionally important dark matter, elusive to established experimental designs, may be crucial for modelling biological systems. In any case, these predictions provide a valuable guide to these experimentally elusive regions.
Juan Garcia Ranea, Ian Morilla, Jonathan G. Lees, Adam James Reid, Corin Yeats, Andrew B. Clegg, Francisca Sánchez-Jiménez, Christine A. Orengo
PLoS Comput. Biol.8
2009 An integrated approach to the interpretation of Single Amino Acid Polymorphisms within the framework of CATH and Gene3D
abstract
BACKGROUND: The phenotypic effects of sequence variations in protein-coding regions come about primarily via their effects on the resulting structures, for example by disrupting active sites or affecting structural stability. In order better to understand the mechanisms behind known mutant phenotypes, and predict the effects of novel variations, biologists need tools to gauge the impacts of DNA mutations in terms of their structural manifestation. Although many mutations occur within domains whose structure has been solved, many more occur within genes whose protein products have not been structurally characterized. RESULTS: Here we present 3DSim (3D Structural Implication of Mutations), a database and web application facilitating the localization and visualization of single amino acid polymorphisms (SAAPs) mapped to protein structures even where the structure of the protein of interest is unknown. The server displays information on 6514 point mutations, 4865 of them known to be associated with disease. These polymorphisms are drawn from SAAPdb, which aggregates data from various sources including dbSNP and several pathogenic mutation databases. While the SAAPdb interface displays mutations on known structures, 3DSim projects mutations onto known sequence domains in Gene3D. This resource contains sequences annotated with domains predicted to belong to structural families in the CATH database. Mappings between domain sequences in Gene3D and known structures in CATH are obtained using a MUSCLE alignment. 1210 three-dimensional structures corresponding to CATH structural domains are currently included in 3DSim; these domains are distributed across 396 CATH superfamilies, and provide a comprehensive overview of the distribution of mutations in structural space. CONCLUSION: The server is publicly available at http://3DSim.bioinfo.cnio.es/. In addition, the database containing the mapping between SAAPdb, Gene3D and CATH is available on request and most of the functionality is available through programmatic web service access.
José M. G. Izarzugaza, Anja Baresic, Lisa E. M. McMillan, Corin Yeats, Andrew B. Clegg, Christine A. Orengo, Andrew C. R. Martin, Alfonso Valencia
BMC Bioinform.6
2009 FLORA: A Novel Method to Predict Protein Function from Structure in Diverse Superfamilies
abstract
Predicting protein function from structure remains an active area of interest, particularly for the structural genomics initiatives where a substantial number of structures are initially solved with little or no functional characterisation. Although global structure comparison methods can be used to transfer functional annotations, the relationship between fold and function is complex, particularly in functionally diverse superfamilies that have evolved through different secondary structure embellishments to a common structural core. The majority of prediction algorithms employ local templates built on known or predicted functional residues. Here, we present a novel method (FLORA) that automatically generates structural motifs associated with different functional sub-families (FSGs) within functionally diverse domain superfamilies. Templates are created purely on the basis of their specificity for a given FSG, and the method makes no prior prediction of functional sites, nor assumes specific physico-chemical properties of residues. FLORA is able to accurately discriminate between homologous domains with different functions and substantially outperforms (a 2-3 fold increase in coverage at low error rates) popular structure comparison methods and a leading function prediction method. We benchmark FLORA on a large data set of enzyme superfamilies from all three major protein classes (alpha, beta, alphabeta) and demonstrate the functional relevance of the motifs it identifies. We also provide novel predictions of enzymatic activity for a large number of structures solved by the Protein Structure Initiative. Overall, we show that FLORA is able to effectively detect functionally similar protein domain structures by purely using patterns of structural conservation of all residues.
Oliver Redfern, Benoit H. Dessailly, Timothy Dallman, Ian Sillitoe, Christine A. Orengo
PLoS Comput. Biol.5
2007 Methods of remote homology detection can be combined to increase coverage by 10% in the midnight zone
abstract
MOTIVATION: A recent development in sequence-based remote homologue detection is the introduction of profile-profile comparison methods. These are more powerful than previous technologies and can detect potentially homologous relationships missed by structural classifications such as CATH and SCOP. As structural classifications traditionally act as the gold standard of homology this poses a challenge in benchmarking them. RESULTS: We present a novel approach which allows an accurate benchmark of these methods against the CATH structural classification. We then apply this approach to assess the accuracy of a range of publicly available methods for remote homology detection including several profile-profile methods (COMPASS, HHSearch, PRC) from two perspectives. First, in distinguishing homologous domains from non-homologues and second, in annotating proteomes with structural domain families. PRC is shown to be the best method for distinguishing homologues. We show that SAM is the best practical method for annotating genomes, whilst using COMPASS for the most remote homologues would increase coverage. Finally, we introduce a simple approach to increase the sensitivity of remote homologue detection by up to 10%. This is achieved by combining multiple methods with a jury vote. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Adam James Reid, Corin Yeats, Christine A. Orengo
Bioinform.3
2007 Establishing a major cause of discrepancy in the calibration of Affymetrix GeneChips
abstract
BACKGROUND: Affymetrix GeneChips are a popular platform for performing whole-genome experiments on the transcriptome. There are a range of different calibration steps, and users are presented with choices of different background subtractions, normalisations and expression measures. We wished to establish which of the calibration steps resulted in the biggest uncertainty in the sets of genes reported to be differentially expressed. RESULTS: Our results indicate that the sets of genes identified as being most significantly differentially expressed, as estimated by the z-score of fold change, is relatively insensitive to the choice of background subtraction and normalisation. However, the contents of the gene list are most sensitive to the choice of expression measure. This is irrespective of whether the experiment uses a rat, mouse or human chip and whether the chip definition is made using probe mappings from Unigene, RefSeq, Entrez Gene or the original Affymetrix definitions. It is also irrespective of whether both Present and Absent, or just Present, Calls from the MAS5 algorithm are used to filter genelists, and this conclusion holds for genes of differing intensities. We also reach the same conclusion after assigning genes to be differentially expressed using t-statistics, although this approach results in a large amount of false positives in the sets of genes identified due to the small numbers of replicates typically used in microarray experiments. CONCLUSION: The major calibration uncertainty that biologists need to consider when analysing Affymetrix data is how their multiple probe values are condensed into one expression measure.
Andrew P. Harrison, Caroline E. Johnston, Christine A. Orengo
BMC Bioinform.3
2007 Towards a comprehensive structural coverage of completed genomes: a structural genomics viewpoint
abstract
BACKGROUND: Structural genomics initiatives were established with the aim of solving protein structures on a large-scale. For many initiatives, such as the Protein Structure Initiative (PSI), the primary aim of target selection is focussed towards structurally characterising protein families which, so far, lack a structural representative. It is therefore of considerable interest to gain insights into the number and distribution of these families, and what efforts may be required to achieve a comprehensive structural coverage across all protein families. RESULTS: In this analysis we have derived a comprehensive domain annotation of the genomes using CATH, Pfam-A and Newfam domain families. We consider what proportions of structurally uncharacterized families are accessible to high-throughput structural genomics pipelines, specifically those targeting families containing multiple prokaryotic orthologues. In measuring the domain coverage of the genomes, we show the benefits of selecting targets from both structurally uncharacterized domain families, whilst in addition, pursuing additional targets from large structurally characterised protein superfamilies. CONCLUSION: This work suggests that such a combined approach to target selection is essential if structural genomics is to achieve a comprehensive structural coverage of the genomes, leading to greater insights into structure and the mechanisms that underlie protein evolution.
Russell L. Marsden, Tony A. Lewis, Christine A. Orengo
BMC Bioinform.3
2007 Inferring Function Using Patterns of Native Disorder in Proteins
abstract
Natively unstructured regions are a common feature of eukaryotic proteomes. Between 30% and 60% of proteins are predicted to contain long stretches of disordered residues, and not only have many of these regions been confirmed experimentally, but they have also been found to be essential for protein function. In this study, we directly address the potential contribution of protein disorder in predicting protein function using standard Gene Ontology (GO) categories. Initially we analyse the occurrence of protein disorder in the human proteome and report ontology categories that are enriched in disordered proteins. Pattern analysis of the distributions of disordered regions in human sequences demonstrated that the functions of intrinsically disordered proteins are both length- and position-dependent. These dependencies were then encoded in feature vectors to quantify the contribution of disorder in human protein function prediction using Support Vector Machine classifiers. The prediction accuracies of 26 GO categories relating to signalling and molecular recognition are improved using the disorder features. The most significant improvements were observed for kinase, phosphorylation, growth factor, and helicase categories. Furthermore, we provide predicted GO term assignments using these classifiers for a set of unannotated and orphan human proteins. In this study, the importance of capturing protein disorder information and its value in function prediction is demonstrated. The GO category classifiers generated can be used to provide more reliable predictions and further insights into the behaviour of orphan and unannotated proteins.
Anna E. Lobley, Mark B. Swindells, Christine A. Orengo, David T. Jones
PLoS Comput. Biol.3
2007 Predicting Protein Function with Hierarchical Phylogenetic Profiles: The Gene3D Phylo-Tuner Method Applied to Eukaryotic Genomes
abstract
"Phylogenetic profiling" is based on the hypothesis that during evolution functionally or physically interacting genes are likely to be inherited or eliminated in a codependent manner. Creating presence-absence profiles of orthologous genes is now a common and powerful way of identifying functionally associated genes. In this approach, correctly determining orthology, as a means of identifying functional equivalence between two genes, is a critical and nontrivial step and largely explains why previous work in this area has mainly focused on using presence-absence profiles in prokaryotic species. Here, we demonstrate that eukaryotic genomes have a high proportion of multigene families whose phylogenetic profile distributions are poor in presence-absence information content. This feature makes them prone to orthology mis-assignment and unsuited to standard profile-based prediction methods. Using CATH structural domain assignments from the Gene3D database for 13 complete eukaryotic genomes, we have developed a novel modification of the phylogenetic profiling method that uses genome copy number of each domain superfamily to predict functional relationships. In our approach, superfamilies are subclustered at ten levels of sequence identity-from 30% to 100%-and phylogenetic profiles built at each level. All the profiles are compared using normalised Euclidean distances to identify those with correlated changes in their domain copy number. We demonstrate that two protein families will "auto-tune" with strong co-evolutionary signals when their profiles are compared at the similarity levels that capture their functional relationship. Our method finds functional relationships that are not detectable by the conventional presence-absence profile comparisons, and it does not require a priori any fixed criteria to define orthologous genes.
Juan Garcia Ranea, Corin Yeats, Alastair Grant, Christine A. Orengo
PLoS Comput. Biol.4
2007 CATHEDRAL: A Fast and Effective Algorithm to Predict Folds and Domain Boundaries from Multidomain Protein Structures
abstract
We present CATHEDRAL, an iterative protocol for determining the location of previously observed protein folds in novel multidomain protein structures. CATHEDRAL builds on the features of a fast secondary-structure-based method (using graph theory) to locate known folds within a multidomain context and a residue-based, double-dynamic programming algorithm, which is used to align members of the target fold groups against the query protein structure to identify the closest relative and assign domain boundaries. To increase the fidelity of the assignments, a support vector machine is used to provide an optimal scoring scheme. Once a domain is verified, it is excised, and the search protocol is repeated in an iterative fashion until all recognisable domains have been identified. We have performed an initial benchmark of CATHEDRAL against other publicly available structure comparison methods using a consensus dataset of domains derived from the CATH and SCOP domain classifications. CATHEDRAL shows superior performance in fold recognition and alignment accuracy when compared with many equivalent methods. If a novel multidomain structure contains a known fold, CATHEDRAL will locate it in 90% of cases, with <1% false positives. For nearly 80% of assigned domains in a manually validated test set, the boundaries were correctly delineated within a tolerance of ten residues. For the remaining cases, previously classified domains were very remotely related to the query chain so that embellishments to the core of the fold caused significant differences in domain sizes and manual refinement of the boundaries was necessary. To put this performance in context, a well-established sequence method based on hidden Markov models was only able to detect 65% of domains, with 33% of the subsequent boundaries assigned within ten residues. Since, on average, 50% of newly determined protein structures contain more than one domain unit, and typically 90% or more of these domains are already classified in CATH, CATHEDRAL will considerably facilitate the automation of protein structure classification.
Oliver Redfern, Andrew P. Harrison, Timothy Dallman, Frances M. G. Pearl, Christine A. Orengo
PLoS Comput. Biol.5
2004 A structural perspective on genome evolution
abstract
At UCL we have developed several automated protocols for generating protein family resources (CATH; Gene3D). These resources can be used to perform comparative genome analyses in order to understand the evolution of protein families. Also to identify biologically and/or medically interesting families for which no structural data currently exists and which may therefore be important targets for structure genomics initiatives.The CATH domain structure database, established by Orengo and Thornton in 1993, now contains a significant proportion of protein structures from the PDB clustered into 1400 evolutionary families. Relationships have been identified using robust structure comparison methods (SSAP, CATHEDRAL). We have also benchmarked and optimised various 1D-profiles and HMM based protocols for assigning genome sequences to families within the resource (e.g. SAM-T99, SAMOSA, CATH-ISL).In this way we can assign structural data to a large proportion (up to 60%) of whole or partial sequences in completed genomes and >80% of genes coding for enzymes and other proteins in biochemical pathways. However, in order to include all families regardless of whether their structure is known or not, a new protein family resource has been developed (Gene3D). In Gene3D, complete genes have been clustered according to sequence similarity alone, using a robust clustering method (Pfscape). 120 completed genomes from all kingdoms have been clustered into 220,000 gene families, 70,000 of which contain 2 or more sequences. Subsequently, we have labelled those gene families for which CATH structural or Pfam functional domain annotations can be provided for all or part of the gene.Preliminary analysis of the genome annotations reveals that a significant proportion (up to 70%) of CATH annotated genes or gene regions in genomes are assigned to domain families that are common to all three kingdoms of life. However, only 20% of the genome sequences are assigned to gene families common to all kingdoms. Since a large proportion of these genes are multidomain proteins this supports the view that a great deal of functional diversity within the genomes has been achieved by combining domain modules in different ways.In collaboration with Professor Janet Thornton, we have analysed a subset of 56 bacterial genomes to determine the recurrence of specific domain structure families within the genomes. This revealed a small but essential group of universal, and in some cases, highly recurring domain families. For some size-dependent families, domain recurrence is highly correlated with increase in genome size, whilst in other size-independent families no correlation is observed. Statistical analysis allowed us to distinguish three groups. Within the size-dependent families we differentiated two groups: linearly-distributed and non-linearly-distributed. Functional annotation using the COGs revealed that these domains were predominantly involved in metabolism and regulation, respectively. Whilst a third group of Evenly-distributed size independent domains are primarily involved in protein translation and biosynthesis.By mapping CATH and Pfam domains families onto all the genome sequences in Gene3D we observe that a few hundred highly recurrent families are dominating at least 50% of whole or partial genome sequences. Many of these families are common to both prokaryotes and eukaryotes and are performing essential generic functions. In many of the largest families, significant divergence in sequence has been accompanied by modifications in structure and function. Targetting representatives in these families for structure determination will allow the structure genomics initiatives to map both fold and function space and reveal the mechanisms by which divergence in protein families promotes evolution of new functions.
David A. Lee, Alastair Grant, Ian Sillitoe, Mark Dibley, Juan Garcia Ranea, Christine A. Orengo
RECOMB6
2004 A practical and robust sequence search strategy for structural genomics target selection
abstract
MOTIVATION: Target selection strategies for structural genomic projects must be able to prioritize gene regions on the basis of significant sequence similarity with proteins that have already been structurally determined. With the rapid development of protein comparison software a robust prioritization scheme should be independent of the choice of algorithm and be able to incorporate different sequence similarity thresholds. RESULTS: A robust target selection strategy has been developed that can assign a priority level to all genes in any genome. Structural assignments to genome sequences are calculated at two thresholds and six levels (1-6) describe the prioritization of all whole genes and partial gene regions. This simple two-threshold approach can be implemented with any fold recognition or homology detection algorithms. The results for 10 genomes are presented using the SSEARCH and PSI-BLAST programs. AVAILABILITY: Programs are available on request from the authors.
James E. Bray, Russell L. Marsden, Stuart C. G. Rison, Alexei Savchenko, Aled M. Edwards, Janet M. Thornton, Christine A. Orengo
Bioinform.7
2003 Recognizing the fold of a protein structure
abstract
This paper reports a graph-theoretic program, GRATH, that rapidly, and accurately, matches a novel structure against a library of domain structures to find the most similar ones. GRATH generates distributions of scores by comparing the novel domain against the different types of folds that have been classified previously in the CATH database of structural domains. GRATH uses a measure of similarity that details the geometric information, number of secondary structures and number of residues within secondary structures, that any two protein structures share. Although GRATH builds on well established approaches for secondary structure comparison, a novel scoring scheme has been introduced to allow ranking of any matches identified by the algorithm. More importantly, we have benchmarked the algorithm using a large dataset of 1702 non-redundant structures from the CATH database which have already been classified into fold groups, with manual validation. This has facilitated introduction of further constraints, optimization of parameters and identification of reliable thresholds for fold identification. Following these benchmarking trials, the correct fold can be identified with the top score with a frequency of 90%. It is identified within the ten most likely assignments with a frequency of 98%. GRATH has been implemented to use via a server (http://www.biochem.ucl.ac.uk/cgi-bin/cath/Grath.pl). GRATH's speed and accuracy means that it can be used as a reliable front-end filter for the more accurate, but computationally expensive, residue based structure comparison algorithm SSAP, currently used to classify domain structures in the CATH database. With an increasing number of structures being solved by the structural genomics initiatives, the GRATH server also provides an essential resource for determining whether newly determined structures are related to any known structures from which functional properties may be inferred.
Andrew P. Harrison, Frances M. G. Pearl, Ian Sillitoe, Tim Slidel, Richard Mott, Janet M. Thornton, Christine A. Orengo
Bioinform.7
2002 PFDB: a generic protein family database integrating the CATH domain structure database with sequence based protein family resources
abstract
MOTIVATION: The PFDB (Protein Family Database) is a new database designed to integrate protein family-related data with relevant functional and genomic data. It currently manages biological data for three projects-the CATH protein domain database (Orengo et al., 1997; Pearl et al., 2001), the VIDA virus domains database (Albà et al., 2001) and the Gene3D database (Buchan et al., 2001). The PFDB has been designed to accommodate protein families identified by a variety of sequence based or structure based protocols and provides a generic resource for biological research by enabling mapping between different protein families and diverse biochemical and genetic data, including complete genomes. RESULTS: A characteristic feature of the PFDB is that it has a number of meta-level entities (for example aggregation, collection and inclusion) represented as base tables in the final design. The explicit representation of relationships at the meta-level has a number of advantages, including flexibility-both in terms of the range of queries that can be formulated and the ability to integrate new biological entities within the existing design. A potential drawback with this approach-poor performance caused by the number of joins across meta-level tables-is avoided by implementing the PFDB with materialized views using the mature relational database technology of Oracle 8i. The resultant database is both fast and flexible. This paper presents the principles on which the database has been designed and implemented, and describes the current status of the database and query facilities supported.
Adrian J. Shepherd, Nigel J. Martin 0001, Roger G. Johnson, Paul Kellam, Christine A. Orengo
Bioinform.5
2002 A framework for modelling virus gene expression data
Paul Kellam, Xiaohui Liu 0001, Nigel J. Martin 0001, Christine A. Orengo, Stephen Swift, Allan Tucker
Intell. Data Anal.4
2001 A Framework for Modelling Short, High-Dimensional Multivariate Time Series: Preliminary Results in Virus Gene Expression Data Analysis
Paul Kellam, Xiaohui Liu 0001, Nigel J. Martin 0001, Christine A. Orengo, Stephen Swift, Allan Tucker
IDA4