EDBT 2026 Demo / reviewers in the wild / expert
Janet M. Thornton
dblp:82/3324
· DBLP profile ↗
35ranked-venue papers
0as first author
1since 2021 · last 2022
0000-0003-0824-4096ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 34 · 1 since 2021Software engineering, systems software and programming languages · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
23 papers |
Bioinformatics and computational biology · 100% Computational science and engineering · 0% |
Topics — the 30 heaviest of 36, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology
enzymatic reaction analysis |
0.6 | 2 | 2018 | Transform-MinER: transforming molecules in enzyme reactions · Bioinform. 2018 Reaction Decoder Tool (RDT): extracting features from chemical reactions · Bioinform. 2016 |
Bioinformatics and computational biology › gene regulation › binding site prediction
protein-ligand binding site prediction |
0.4 | 1 | 2020 | GRaSP: a graph-based residue neighborhood strategy to predict binding sites · Bioinform. 2020 |
Bioinformatics and computational biology
genomics |
0.4 | 1 | 2019 | VarMap: a web tool for mapping genomic coordinates to protein sequence and structure and retrieving protein structural annotations · Bioinform. 2019 |
Bioinformatics and computational biology › genome annotation
genomic variant annotation |
0.4 | 1 | 2019 | VarMap: a web tool for mapping genomic coordinates to protein sequence and structure and retrieving protein structural annotations · Bioinform. 2019 |
Bioinformatics and computational biology › biological database
protein data bank |
0.4 | 1 | 2019 | Finding enzyme cofactors in Protein Data Bank · Bioinform. 2019 |
Bioinformatics and computational biology › protein design
enzyme design |
0.3 | 1 | 2018 | Transform-MinER: transforming molecules in enzyme reactions · Bioinform. 2018 |
Bioinformatics and computational biology
structural bioinformatics |
0.3 | 5 | 2009 | WSsas: a web service for the annotation of functional residues through structural homologues · Bioinform. 2009 Analysis of binding site similarity, small-molecule similarity and experimental binding profiles in the human cytosolic sulfotransferase family · Bioinform. 2007 Real spherical harmonic expansion coefficients as 3D shape descriptors for protein binding pocket and ligand comparisons · Bioinform. 2005 |
Bioinformatics and computational biology
protein structure analysis |
0.3 | 5 | 2009 | PoreLogo: a new tool to analyse, visualize and compare channels in transmembrane proteins · Bioinform. 2009 HTHquery: a method for detecting DNA-binding proteins with a helix-turn-helix structural motif · Bioinform. 2005 An examination of the conservation of surface patch polarity for proteins · Bioinform. 2004 |
Bioinformatics and computational biology › molecular informatics › cheminformatics
atom mapping |
0.2 | 1 | 2016 | Reaction Decoder Tool (RDT): extracting features from chemical reactions · Bioinform. 2016 |
Bioinformatics and computational biology › systems bioinformatics › pathway analysis
metabolic pathway analysis |
0.2 | 1 | 2016 | Reaction Decoder Tool (RDT): extracting features from chemical reactions · Bioinform. 2016 |
Bioinformatics and computational biology
survival analysis |
0.2 | 1 | 2015 | SurvCurv database and online survival analysis platform update · Bioinform. 2015 |
Bioinformatics and computational biology › systems bioinformatics
pathway analysis |
0.2 | 1 | 2014 | Comparison of the mammalian insulin signalling pathway to invertebrates in the context of FOXO-mediated ageing · Bioinform. 2014 |
Bioinformatics and computational biology
biological database |
0.2 | 2 | 2015 | The CoFactor database: organic cofactors in enzyme catalysis · Bioinform. 2010 SurvCurv database and online survival analysis platform update · Bioinform. 2015 |
Bioinformatics and computational biology › protein function prediction › sequence-based protein function prediction
homology-based function transfer |
0.1 | 1 | 2009 | WSsas: a web service for the annotation of functional residues through structural homologues · Bioinform. 2009 |
Bioinformatics and computational biology › protein structure analysis
membrane protein analysis |
0.1 | 1 | 2009 | PoreLogo: a new tool to analyse, visualize and compare channels in transmembrane proteins · Bioinform. 2009 |
Bioinformatics and computational biology
protein function prediction |
0.1 | 1 | 2009 | WSsas: a web service for the annotation of functional residues through structural homologues · Bioinform. 2009 |
Bioinformatics and computational biology › protein analysis › protein bioinformatics
protein annotation |
0.1 | 1 | 2008 | The Protein Feature Ontology: a tool for the unification of protein feature annotations · Bioinform. 2008 |
Bioinformatics and computational biology › drug discovery
drug design |
0.1 | 1 | 2007 | Analysis of binding site similarity, small-molecule similarity and experimental binding profiles in the human cytosolic sulfotransferase family · Bioinform. 2007 |
Bioinformatics and computational biology › protein structure analysis › protein binding site analysis
protein binding site comparison |
0.1 | 1 | 2007 | Analysis of binding site similarity, small-molecule similarity and experimental binding profiles in the human cytosolic sulfotransferase family · Bioinform. 2007 |
Bioinformatics and computational biology › genomics
structural genomics |
0.1 | 2 | 2004 | A practical and robust sequence search strategy for structural genomics target selection · Bioinform. 2004 Recognizing the fold of a protein structure · Bioinform. 2003 |
Bioinformatics and computational biology › protein structure prediction › template-based modeling
fold recognition |
0.1 | 2 | 2004 | Recognizing the fold of a protein structure · Bioinform. 2003 A practical and robust sequence search strategy for structural genomics target selection · Bioinform. 2004 |
Bioinformatics and computational biology › protein function prediction › protein classification
DNA-binding protein prediction |
0.1 | 1 | 2005 | HTHquery: a method for detecting DNA-binding proteins with a helix-turn-helix structural motif · Bioinform. 2005 |
Bioinformatics and computational biology › protein structure analysis › protein binding site analysis
protein pocket comparison |
0.1 | 1 | 2005 | Real spherical harmonic expansion coefficients as 3D shape descriptors for protein binding pocket and ligand comparisons · Bioinform. 2005 |
Bioinformatics and computational biology
shape descriptor |
0.1 | 1 | 2005 | Real spherical harmonic expansion coefficients as 3D shape descriptors for protein binding pocket and ligand comparisons · Bioinform. 2005 |
Bioinformatics and computational biology › bioinformatics infrastructure
biological data management |
0.0 | 1 | 2004 | Software Engineering Challenges in Bioinformatics · ICSE 2004 |
Bioinformatics and computational biology
data integration |
0.0 | 1 | 2004 | Software Engineering Challenges in Bioinformatics · ICSE 2004 |
Bioinformatics and computational biology › genomics › structural genomics
target selection |
0.0 | 1 | 2004 | A practical and robust sequence search strategy for structural genomics target selection · Bioinform. 2004 |
Bioinformatics and computational biology › data integration
biological data integration |
0.0 | 1 | 2008 | The Protein Feature Ontology: a tool for the unification of protein feature annotations · Bioinform. 2008 |
Computational science and engineering
controlled vocabulary |
0.0 | 1 | 2008 | The Protein Feature Ontology: a tool for the unification of protein feature annotations · Bioinform. 2008 |
Bioinformatics and computational biology › biological database
enzyme database |
0.0 | 1 | 2005 | MACiE: a database of enzyme reaction mechanisms · Bioinform. 2005 |
Methods — techniques the papers use, named apart from their topics
supervised learning · 0.4graph modeling · 0.4web tool · 0.4data pipeline · 0.4molecular similarity · 0.3dynamic programming · 0.2statistical survival analysis · 0.2literature curation · 0.2in silico gene perturbation · 0.2sequence conservation analysis · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Srinivasan (1962-2021) in Bioinformatics and beyondabstractDear Editor, Last year the Bioinformatics community lost one of its pioneers, a scientist renowned for his talent, creativity and rigour but also for his commitment to supporting his research community and particularly the young scientists he trained. He was an inspiring role model for his field and a scientist who will be remembered very fondly by his many friends in the community for his warmth, humour and kindness. For more than three decades Srinivasan developed timely and novel computational strategies for analyzing proteins and was regarded in high esteem internationally for the insights he provided and the resources he established based on the underpinning concepts. His discoveries cover many areas fundamental to structural biology and pathogen research. Although he was a computational scientist, he worked closely with experimental groups to maximize the impact of his research. He published more than 300 papers, with nearly 10 000 citations altogether. Srinivasan joined the faculty of the Molecular Biophysics Unit, Indian Institute of Science, Bangalore in 1998, after leaving the Madras Biophysics Group (he did his Masters studies from 1982-84). He acquired his PhD degree in the Molecular Biophysics Department (the same Department where he later worked as a faculty member) within the GN Ramachandran school of peptide and peptide stereochemistry. His postdoctoral tenure was in Prof. Sir Tom Blundell’s laboratory (1991–1998), Birkbeck College, UK, with a brief stint in Prof. Mike Waterfield’s laboratory at the Ludwig Institute for Cancer Research, UK. He arrived in London as a seemingly shy young man, but it soon became clear that he was a real expert in protein structures and thought very deeply about their evolution. During these times, his research was largely focused on homology modelling (Johnson et al., 1994) and the study of proteins involved in signal transduction (e.g. Srinivasan et al., 1994, 1996). After coming back to MBU, he headed the ‘Proteins: structure, function and evolutionary’ group. He made major contributions to the understanding the structure and functions of proteins, particularly on protein kinases in a wide range of model organisms (e.g. Krupa and Srinivasan, 2002; Krupa et al., 2004a, b). Specifically, his lab was focused on computational genomics, bioinformatics and structural biology, particularly involving the relationships between protein structure, function and interactions, including protein–protein interactions, cellular signal transduction and biological pathways. Srinivasan’s interest in protein families and protein evolution drove research into strategies for improving multiple alignments of relatives and for better characterizing phylogenetic relationships (which resulted in many useful resources like SUPFAM, MulPSSM, PALI and DoSA). It also drove the design of methods to detect extremely remote homologues, which have been valuable for extending structural and functional annotation of genomes. Srinivasan’s group also showed that sequence-based connections of distantly related proteins can be enabled through the design of artificial sequences (Mudgal et al., 2014). This work was highly innovative and can help to bring valuable annotations for pathogen proteins, which are typically difficult to characterize by more conventional, less-sensitive strategies. His strategies allowed a much deeper characterization of fold space to inform protein engineering. Srinivasan is also highly renowned for his analyses of how changes in the protein structure and sequence impact function. Some protein families, like the kinases, were a major focus of his research and gave him much international acclaim. He studied kinases for >20 years and contributed numerous insights important for understanding their mechanisms and for enabling drug design. For example, structural fluctuations, classifications based on key functional site properties (Kalaivani et al., 2018), mechanisms of stabilization of their key functional sites through specific residue interactions. He also characterized the ways in which domain partnerships modify structure (Vishwanath et al., 2018), kinase functionality and characterized how splicing extends the kinase functional repertoire. This large family is implicated in many human diseases, including cancer, and these discoveries have informed drug design. However, the biological role of proteins is determined by their interactions and Srinivasan applied his precise analytical skills in this arena, too, revealing key insights into the properties of the interfaces involved in assembling protein complexes. He produced a substantial body of very rigorous studies, including analyses of the characteristics of transient complexes and the effect of protein associations on global structural dynamics. He robustly captured this knowledge in the PIC protein interactions calculator (Tina et al., 2007), a valuable tool that is freely available to biologists and very popular among researchers to obtain structural data on various non-covalent interactions within a protein or between proteins in a complex. An important application of these methods was the characterization of interactions between viral proteins and their host proteins, which provided key data for understanding pathogenicity and enabling drug design. For example, Srinivasan performed various studies characterizing toxin–antitoxin systems (Tandon et al., 2019), protein interactions between human erythrocytes and Plasmodium falciparum and Helicobacter pylori and human. Photo taken at the fifth IIT Madras-Tokyo Tech joint symposium on ‘Current Trends in Bioinformatics: Big Data Analysis, Machine Learning and Drug Design’ with leading Bioinformatics scientists in India (March 2020). From left to right D. Velmurugan, N. Manoj, S. Selvaraj, P.K. Ponnuswamy, G.P.S. Raghava, K. Veluraja, Shandar Ahmad, M. Michael Gromiha, N. Srinivasan, R. Sowdhamini Photo taken at the fifth IIT Madras-Tokyo Tech joint symposium on ‘Current Trends in Bioinformatics: Big Data Analysis, Machine Learning and Drug Design’ with leading Bioinformatics scientists in India (March 2020). From left to right D. Velmurugan, N. Manoj, S. Selvaraj, P.K. Ponnuswamy, G.P.S. Raghava, K. Veluraja, Shandar Ahmad, M. Michael Gromiha, N. Srinivasan, R. Sowdhamini In collaboration with multiple laboratories, Srini’s group studied fascinating biological systems, including protein assemblies of ribosomes and spliceosomes (Bhat et al., 2015; Pudi et al., 2003; Yazhini et al., 2022) and developed powerful computational tools for studying structures of large assemblies derived from cryo-electron microscopy (Joseph et al., 2016; Rakesh et al., 2016). As is true of several structural bioinformaticians, his group relied on publicly available structural data and were concerned with the quality of protein–ligand data, deposited through X-ray or cryo-EM studies (Chakraborti et al., 2021). He was always excited to discuss the Ramachandran map. Last year, he attended the fifth IIT Madras—Tokyo Tech joint symposium on Bioinformatics and his lecture on the Ramachandran map was fascinating. Using modern computational tools and along with late Prof. C. Ramakrishnan (his PhD mentor and the student behind the original Ramachandran map) and one of his students Ashraya Ravikumar, he re-examined the classical and renowned Ramachandran map. They clearly demonstrated that it is possible to consider deviations in the ‘allowed’ regions within this map by considering slight deviations in internal parameters from ideal values of the peptide bond (Ravikumar et al., 2019). Srinivasan’s group also participated in several consortia such as the Open Source Drug Discovery program [with the groups of Prof. Tom Blundell (University of Cambridge, UK), Nagasuma Chandra (Indian Institute of Science, India) and Sowdhamini (National Centre for Biological Sciences, India)], the UKIERI study of protein assemblies [with the groups of Profs. Jim Warwicker (University of Manchester, UK), Pinak Chakrabarti (Bose Institute, India), Nagasuma Chandra and Sowdhamini)] and collaborations such as the Indo-French CEFIPRA project on protein alphabets [with Dr. Alexandre de Brevern (INSERM Paris, France) and Dr. Bernard Offmann (University of Nantes, France)], and the Centre for Excellence on protein-protein interactions [with Profs. Sowdhamini and Satyajit Mayor (National Centre for Biological Sciences, India) and Nagasuma Chandra (Indian Institute of Science, India)] and toxin-antitoxin systems [with Prof. Raghavan Varadarajan (Indian Institute of Science, India)]. Throughout his career Srinivasan applied his knowledge, data and computational tools to characterize the protein structures, functions and virus–host interactions of multiple pathogenic bacteria affecting human health, including mycobacterial pathogens (e.g. Mtb), malaria, H. Pylori, Dengue and several gut pathogens. Understanding the critical residues in the protein interface is essential for drug design to reduce infection and pathogenicity. For many years he collaborated with the group of Professor Tom Blundell in Cambridge, UK. As well as detailed analyses, he established the SInCRe structural interactome resource for Mtb in 2015 (Metri et al., 2015), which contributed to studies on the repurposing of drugs for this pathogen. His tools have been applied in a number of medical contexts with promising clinical results. His group also applied docking tools to FDA-approved drugs to SARS-CoV2 (Chakraborti et al., 2020) and his most recent work on inhibitors for the main protease of SARS-CoV2 led to compounds already in clinical trials. Prof. Srinivasan made significant contributions to Bioinformatics and his whole-hearted involvement in scientific activities will not be forgotten. As a scientist, Srinivasan was very highly focused and meticulous. He always set high standards—whether in creating high-quality datasets or in his interpretations of data or in responding to reviewers’ comments. As well as being a multitalented and well-known researcher, Srinivasan actively participated in many university and external committees, commented on PhD theses, and delivered popular and invited lectures in most of the leading conferences in India. He was an elected fellow in all the three major academies in India (Indian National Science Academy, New Delhi, National Academy of Sciences, Allahabad and Indian Academy of Sciences, Bangalore). He also received the most prestigious awards in India including Shanti Swarup Bhatnagar Prize for Science and Technology from Council of Scientific and Industrial Research, National Bioscience Award from the Department of Biotechnology and J.C. Bose National Fellowship from the Department of Science and Technology, Government of India. Srinivasan had the special ability to cordially relate with others he respected and had a very positive attitude towards the work of his fellow researchers. He was a faithful chairperson of the Department (serving between 2018 and 2020) and always supported and wished his younger colleagues to do well. He remained active even when he was critically ill. During this time he still managed to publish around 10 papers, enable six of his lab colleagues to reach higher positions and also attended to multiple student-thesis-related matters. He was very enthusiastic about his research on protein structures and his ability to explain major concepts in a simple accessible manner was extremely impressive. He had a passion for naming his students with ‘amino acids’, each with a background story and spent considerable time with his students in the midst of his busy schedule. He always encouraged young researchers and provided valuable advice for their research. He also had an uncanny enthusiasm and ability to make sure that the people around him felt included and important—he would not hesitate to talk to prospective students, spend a long time on discussions with visitors and provide his undivided attention on work discussions with his students. Many of his students travelled widely and benefitted laboratories and science throughout the world. His students made a large impact wherever they went because their knowledge was always deep and impressive and they showed great enthusiasm for their work, mirroring their mentor. They often liked to discuss their work in detail and place it in the wider context of global knowledge. Although based in India for most of his career, Srinivasan travelled widely and has had an impact on many scientists involved in protein structure analysis. He was also a great host for visitors, ensuring their well-being and spending time discussing their work and ideas. Srinivasan had been a very special and unusual personality—with unlimited affection and love for the people around him. He was able to sense people in trouble and would often go out of his way to help them. His smiling face is not forgettable at any time and evidenced a deeply contented person, very proud of his family, his students and his science and always happy to discuss anything to do with proteins. His positivity and passion for science were infectious. His too-early passing is a great loss to science and to everyone who knew him. Financial Support: none declared. Conflict of Interest: The authors declare that there are no conflicts of interest. M. Michael Gromiha, Christine A. Orengo, Ramanathan Sowdhamini, Janet M. Thornton |
Bioinform. | 4 |
| 2020 | GRaSP: a graph-based residue neighborhood strategy to predict binding sitesabstractMOTIVATION: The discovery of protein-ligand-binding sites is a major step for elucidating protein function and for investigating new functional roles. Detecting protein-ligand-binding sites experimentally is time-consuming and expensive. Thus, a variety of in silico methods to detect and predict binding sites was proposed as they can be scalable, fast and present low cost. RESULTS: We proposed Graph-based Residue neighborhood Strategy to Predict binding sites (GRaSP), a novel residue centric and scalable method to predict ligand-binding site residues. It is based on a supervised learning strategy that models the residue environment as a graph at the atomic level. Results show that GRaSP made compatible or superior predictions when compared with methods described in the literature. GRaSP outperformed six other residue-centric methods, including the one considered as state-of-the-art. Also, our method achieved better results than the method from CAMEO independent assessment. GRaSP ranked second when compared with five state-of-the-art pocket-centric methods, which we consider a significant result, as it was not devised to predict pockets. Finally, our method proved scalable as it took 10-20 s on average to predict the binding site for a protein complex whereas the state-of-the-art residue-centric method takes 2-5 h on average. AVAILABILITY AND IMPLEMENTATION: The source code and datasets are available at https://github.com/charles-abreu/GRaSP. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Charles Abreu Santana, Sabrina de Azevedo Silveira, João P. A. Moraes, Sandro C. Izidoro, Raquel Cardoso de Melo Minardi, António J. M. Ribeiro, Jonathan D. Tyzack, Neera Borkakoti, Janet M. Thornton |
Bioinform. | 9 |
| 2020 | An automated protocol for modelling peptide substrates to proteasesabstractBACKGROUND: Proteases are key drivers in many biological processes, in part due to their specificity towards their substrates. However, depending on the family and molecular function, they can also display substrate promiscuity which can also be essential. Databases compiling specificity matrices derived from experimental assays have provided valuable insights into protease substrate recognition. Despite this, there are still gaps in our knowledge of the structural determinants. Here, we compile a set of protease crystal structures with bound peptide-like ligands to create a protocol for modelling substrates bound to protease structures, and for studying observables associated to the binding recognition. RESULTS: As an application, we modelled a subset of protease-peptide complexes for which experimental cleavage data are available to compare with informational entropies obtained from protease-specificity matrices. The modelled complexes were subjected to conformational sampling using the Backrub method in Rosetta, and multiple observables from the simulations were calculated and compared per peptide position. We found that some of the calculated structural observables, such as the relative accessible surface area and the interaction energy, can help characterize a protease's substrate recognition, giving insights for the potential prediction of novel substrates by combining additional approaches. CONCLUSION: Overall, our approach provides a repository of protease structures with annotated data, and an open source computational protocol to reproduce the modelling and dynamic analysis of the protease-peptide complexes. Rodrigo Ochoa, Mikhail D. Magnitov, Roman A. Laskowski, Pilar Cossio, Janet M. Thornton |
BMC Bioinform. | 5 |
| 2019 | Finding enzyme cofactors in Protein Data BankabstractMOTIVATION: Cofactors are essential for many enzyme reactions. The Protein Data Bank (PDB) contains >67 000 entries containing enzyme structures, many with bound cofactor or cofactor-like molecules. This work aims to identify and categorize these small molecules in the PDB and make it easier to find them. RESULTS: The Protein Data Bank in Europe (PDBe; pdbe.org) has implemented a pipeline to identify enzyme cofactor and cofactor-like molecules, which are now part of the PDBe weekly release process. AVAILABILITY AND IMPLEMENTATION: Information is made available on the individual PDBe entry pages at pdbe.org and programmatically through the PDBe REST API (pdbe.org/api). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Abhik Mukhopadhyay, Neera Borkakoti, Lukás Pravda, Jonathan D. Tyzack, Janet M. Thornton, Sameer Velankar |
Bioinform. | 5 |
| 2019 | VarMap: a web tool for mapping genomic coordinates to protein sequence and structure and retrieving protein structural annotationsabstractMOTIVATION: Understanding the protein structural context and patterning on proteins of genomic variants can help to separate benign from pathogenic variants and reveal molecular consequences. However, mapping genomic coordinates to protein structures is non-trivial, complicated by alternative splicing and transcript evidence. RESULTS: Here we present VarMap, a web tool for mapping a list of chromosome coordinates to canonical UniProt sequences and associated protein 3D structures, including validation checks, and annotating them with structural information. AVAILABILITY AND IMPLEMENTATION: https://www.ebi.ac.uk/thornton-srv/databases/VarMap. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. James D. Stephenson, Roman A. Laskowski, Andrew Nightingale, Matthew E. Hurles, Janet M. Thornton |
Bioinform. | 5 |
| 2019 | Using the drug-protein interactome to identify anti-ageing compounds for humansabstractAdvancing age is the dominant risk factor for most of the major killer diseases in developed countries. Hence, ameliorating the effects of ageing may prevent multiple diseases simultaneously. Drugs licensed for human use against specific diseases have proved to be effective in extending lifespan and healthspan in animal models, suggesting that there is scope for drug repurposing in humans. New bioinformatic methods to identify and prioritise potential anti-ageing compounds for humans are therefore of interest. In this study, we first used drug-protein interaction information, to rank 1,147 drugs by their likelihood of targeting ageing-related gene products in humans. Among 19 statistically significant drugs, 6 have already been shown to have pro-longevity properties in animal models (p < 0.001). Using the targets of each drug, we established their association with ageing at multiple levels of biological action including pathways, functions and protein interactions. Finally, combining all the data, we calculated a ranked list of drugs that identified tanespimycin, an inhibitor of HSP-90, as the top-ranked novel anti-ageing candidate. We experimentally validated the pro-longevity effect of tanespimycin through its HSP-90 target in Caenorhabditis elegans. Matías Fuentealba, Handan Melike Dönertas, Rhianna Williams, Johnathan Labbadia, Janet M. Thornton, Linda Partridge |
PLoS Comput. Biol. | 5 |
| 2018 | Transform-MinER: transforming molecules in enzyme reactionsabstractMotivation: One goal of synthetic biology is to make new enzymes to generate new products, but identifying the starting enzymes for further investigation is often elusive and relies on expert knowledge, intensive literature searching and trial and error. Results: We present Transform Molecules in Enzyme Reactions, an online computational tool that transforms query substrate molecules into products using enzyme reactions. The most similar native enzyme reactions for each transformation are found, highlighting those that may be of most interest for enzyme design and directed evolution approaches. Availability and implementation: https://www.ebi.ac.uk/thornton-srv/transform-miner. Jonathan D. Tyzack, António J. M. Ribeiro, Neera Borkakoti, Janet M. Thornton |
Bioinform. | 4 |
| 2016 | Reaction Decoder Tool (RDT): extracting features from chemical reactionsabstractUNLABELLED: Extracting chemical features like Atom-Atom Mapping (AAM), Bond Changes (BCs) and Reaction Centres from biochemical reactions helps us understand the chemical composition of enzymatic reactions. Reaction Decoder is a robust command line tool, which performs this task with high accuracy. It supports standard chemical input/output exchange formats i.e. RXN/SMILES, computes AAM, highlights BCs and creates images of the mapped reaction. This aids in the analysis of metabolic pathways and the ability to perform comparative studies of chemical reactions based on these features. AVAILABILITY AND IMPLEMENTATION: This software is implemented in Java, supported on Windows, Linux and Mac OSX, and freely available at https://github.com/asad/ReactionDecoder CONTACT: : [email protected] or [email protected]. Syed Asad Rahman, Gilliean Torrance, Lorenzo Baldacci, Sergio Martínez Cuesta, Franz Fenninger, Nimish Gopal, Saket Choudhary, John W. May, Gemma L. Holliday, Christoph Steinbeck, Janet M. Thornton |
Bioinform. | 11 |
| 2015 | SurvCurv database and online survival analysis platform updateabstractUNLABELLED: Understanding the biology of ageing is an important and complex challenge. Survival experiments are one of the primary approaches for measuring changes in ageing. Here, we present a major update to SurvCurv, a database and online resource for survival data in animals. As well as a substantial increase in data and additions to existing graphical and statistical survival analysis features, SurvCurv now includes extended mathematical mortality modelling functions and survival density plots for more advanced representation of groups of survival cohorts. AVAILABILITY AND IMPLEMENTATION: The database is freely available at https://www.ebi.ac.uk/thornton-srv/databases/SurvCurv/. All data are published under the Creative Commons Attribution License. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Matthias Ziehm, Dobril K. Ivanov, Aditi Bhat, Linda Partridge, Janet M. Thornton |
Bioinform. | 5 |
| 2015 | Comparisons of Allergenic and Metazoan Parasite Proteins: Allergy the Price of ImmunityabstractAllergic reactions can be considered as maladaptive IgE immune responses towards environmental antigens. Intriguingly, these mechanisms are observed to be very similar to those implicated in the acquisition of an important degree of immunity against metazoan parasites (helminths and arthropods) in mammalian hosts. Based on the hypothesis that IgE-mediated immune responses evolved in mammals to provide extra protection against metazoan parasites rather than to cause allergy, we predict that the environmental allergens will share key properties with the metazoan parasite antigens that are specifically targeted by IgE in infected human populations. We seek to test this prediction by examining if significant similarity exists between molecular features of allergens and helminth proteins that induce an IgE response in the human host. By employing various computational approaches, 2712 unique protein molecules that are known IgE antigens were searched against a dataset of proteins from helminths and parasitic arthropods, resulting in a comprehensive list of 2445 parasite proteins that show significant similarity through sequence and structure with allergenic proteins. Nearly half of these parasite proteins from 31 species fall within the 10 most abundant allergenic protein domain families (EF-hand, Tropomyosin, CAP, Profilin, Lipocalin, Trypsin-like serine protease, Cupin, BetV1, Expansin and Prolamin). We identified epitopic-like regions in 206 parasite proteins and present the first example of a plant protein (BetV1) that is the commonest allergen in pollen in a worm, and confirming it as the target of IgE in schistosomiasis infected humans. The identification of significant similarity, inclusive of the epitopic regions, between allergens and helminth proteins against which IgE is an observed marker of protective immunity explains the 'off-target' effects of the IgE-mediated immune system in allergy. All these findings can impact the discovery and design of molecules used in immunotherapy of allergic conditions. Nidhi Tyagi, Edward J. Farnell, Colin M. Fitzsimmons, Stephanie Ryan, Edridah M. Tukahebwa, Rick M. Maizels, David W. Dunne, Janet M. Thornton, Nicholas Furnham |
PLoS Comput. Biol. | 8 |
| 2014 | Comparison of the mammalian insulin signalling pathway to invertebrates in the context of FOXO-mediated ageingabstractMOTIVATION: A large number of experimental studies on ageing focus on the effects of genetic perturbations of the insulin/insulin-like growth factor signalling pathway (IIS) on lifespan. Short-lived invertebrate laboratory model organisms are extensively used to quickly identify ageing-related genes and pathways. It is important to extrapolate this knowledge to longer lived mammalian organisms, such as mouse and eventually human, where such analyses are difficult or impossible to perform. Computational tools are needed to integrate and manipulate pathway knowledge in different species. RESULTS: We performed a literature review and curation of the IIS and target of rapamycin signalling pathways in Mus Musculus. We compare this pathway model to the equivalent models in Drosophila melanogaster and Caenorhabtitis elegans. Although generally well-conserved, they exhibit important differences. In general, the worm and mouse pathways include a larger number of feedback loops and interactions than the fly. We identify 'functional orthologues' that share similar molecular interactions, but have moderate sequence similarity. Finally, we incorporate the mouse model into the web-service NetEffects and perform in silico gene perturbations of IIS components and analyses of experimental results. We identify sub-paths that, given a mutation in an IIS component, could potentially antagonize the primary effects on ageing via FOXO in mouse and via SKN-1 in worm. Finally, we explore the effects of FOXO knockouts in three different mouse tissues. AVAILABILITY AND IMPLEMENTATION: http://www.ebi.ac.uk/thornton-srv/software/NetEffects. Irene Papatheodorou, Rudolfs Petrovs, Janet M. Thornton |
Bioinform. | 3 |
| 2013 | Amino Acid Changes in Disease-Associated Variants Differ Radically from Variants Observed in the 1000 Genomes Project DatasetabstractThe 1000 Genomes Project data provides a natural background dataset for amino acid germline mutations in humans. Since the direction of mutation is known, the amino acid exchange matrix generated from the observed nucleotide variants is asymmetric and the mutabilities of the different amino acids are very different. These differences predominantly reflect preferences for nucleotide mutations in the DNA (especially the high mutation rate of the CpG dinucleotide, which makes arginine mutability very much higher than other amino acids) rather than selection imposed by protein structure constraints, although there is evidence for the latter as well. The variants occur predominantly on the surface of proteins (82%), with a slight preference for sites which are more exposed and less well conserved than random. Mutations to functional residues occur about half as often as expected by chance. The disease-associated amino acid variant distributions in OMIM are radically different from those expected on the basis of the 1000 Genomes dataset. The disease-associated variants preferentially occur in more conserved sites, compared to 1000 Genomes mutations. Many of the amino acid exchange profiles appear to exhibit an anti-correlation, with common exchanges in one dataset being rare in the other. Disease-associated variants exhibit more extreme differences in amino acid size and hydrophobicity. More modelling of the mutational processes at the nucleotide level is needed, but these observations should contribute to an improved prediction of the effects of specific variants in humans. Tjaart A. P. de Beer, Roman A. Laskowski, Sarah L. Parks, Botond Sipos, Nick Goldman, Janet M. Thornton |
PLoS Comput. Biol. | 6 |
| 2012 | Exploring the Evolution of Novel Enzyme Functions within Structurally Defined Protein SuperfamiliesabstractIn order to understand the evolution of enzyme reactions and to gain an overview of biological catalysis we have combined sequence and structural data to generate phylogenetic trees in an analysis of 276 structurally defined enzyme superfamilies, and used these to study how enzyme functions have evolved. We describe in detail the analysis of two superfamilies to illustrate different paradigms of enzyme evolution. Gathering together data from all the superfamilies supports and develops the observation that they have all evolved to act on a diverse set of substrates, whilst the evolution of new chemistry is much less common. Despite that, by bringing together so much data, we can provide a comprehensive overview of the most common and rare types of changes in function. Our analysis demonstrates on a larger scale than previously studied, that modifications in overall chemistry still occur, with all possible changes at the primary level of the Enzyme Commission (E.C.) classification observed to a greater or lesser extent. The phylogenetic trees map out the evolutionary route taken within a superfamily, as well as all the possible changes within a superfamily. This has been used to generate a matrix of observed exchanges from one enzyme function to another, revealing the scale and nature of enzyme evolution and that some types of exchanges between and within E.C. classes are more prevalent than others. Surprisingly a large proportion (71%) of all known enzyme functions are performed by this relatively small set of 276 superfamilies. This reinforces the hypothesis that relatively few ancient enzymatic domain superfamilies were progenitors for most of the chemistry required for life. Nicholas Furnham, Ian Sillitoe, Gemma L. Holliday, Alison L. Cuff, Roman A. Laskowski, Christine A. Orengo, Janet M. Thornton |
PLoS Comput. Biol. | 7 |
| 2010 | The CoFactor database: organic cofactors in enzyme catalysisabstractMOTIVATION: Organic enzyme cofactors are involved in many enzyme reactions. Therefore, the analysis of cofactors is crucial to gain a better understanding of enzyme catalysis. To aid this, we have created the CoFactor database. RESULTS: CoFactor provides a web interface to access hand-curated data extracted from the literature on organic enzyme cofactors in biocatalysis, as well as automatically collected information. CoFactor includes information on the conformational and solvent accessibility variation of the enzyme-bound cofactors, as well as mechanistic and structural information about the hosting enzymes. AVAILABILITY: The database is publicly available and can be accessed at http://www.ebi.ac.uk/thornton-srv/databases/CoFactor. Julia D. Fischer, Gemma L. Holliday, Janet M. Thornton |
Bioinform. | 3 |
| 2009 | Metal-MACiE: a database of metals involved in biological catalysisabstractSUMMARY: Metal-MACiE is a new publicly available web-based database, held in MySQL, which aims to organize the available information on the properties and the roles of metals in the context of the catalytic mechanisms of metalloenzymes. Metal-MACiE, which currently covers 75% of metal-dependent enzyme commission (EC) sub-sub-classes and is continuously growing, exploits the existing MACiE database for the annotation of the reaction mechanisms. The two databases constitute complementary sources of information for enzymology, biochemistry and molecular pharmacology studies. AVAILABILITY: http://www.ebi.ac.uk/thornton-srv/databases/Metal_MACiE/home.html. Claudia Andreini, Ivano Bertini, Gabriele Cavallaro, Gemma L. Holliday, Janet M. Thornton |
Bioinform. | 5 |
| 2009 | PoreLogo: a new tool to analyse, visualize and compare channels in transmembrane proteinsabstractUNLABELLED: The increasing number of available atomic 3D structures of transmembrane channel proteins represents a valuable resource for better understanding their structure-function relationships and to eventually predict their selectivity. Herein, we present PoreLogo, an automatic tool for analysing, visualizing and comparing the amino acid composition of transmembrane channels and its conservation across the corresponding protein family. AVAILABILITY: PoreLogo is accessible as a public web server at http://www.ebi.ac.uk/thornton-srv/software/PoreLogo/. Romina Oliva, Janet M. Thornton, Marialuisa Pellegrini-Calace |
Bioinform. | 2 |
| 2009 | WSsas: a web service for the annotation of functional residues through structural homologuesabstractMOTIVATION: Annotation tools help scientists to traverse the gap between characterized and uncharacterized proteins. Tools for the prediction of protein function include those which predict the function of entire proteins or complexes, those annotating functional domains and those which predict specific residues within the domain. We have developed WSsas, a web service focused on the annotation of essential functional residues. WSsas uses similarity searches and pairwise alignments to transfer functional information about binding, catalytic and protein-protein interaction residues from solved structures to query sequences. In addition, WSsas can supply information about the relevant functional atoms. The web service definition (WSDL) file and a Perl client are freely available at http://www.ebi.ac.uk/thornton-srv/databases/WSsas/. David Talavera, Roman A. Laskowski, Janet M. Thornton |
Bioinform. | 3 |
| 2009 | Predicting Protein Ligand Binding Sites by Combining Evolutionary Sequence Conservation and 3D StructureabstractIdentifying a protein's functional sites is an important step towards characterizing its molecular function. Numerous structure- and sequence-based methods have been developed for this problem. Here we introduce ConCavity, a small molecule binding site prediction algorithm that integrates evolutionary sequence conservation estimates with structure-based methods for identifying protein surface cavities. In large-scale testing on a diverse set of single- and multi-chain protein structures, we show that ConCavity substantially outperforms existing methods for identifying both 3D ligand binding pockets and individual ligand binding residues. As part of our testing, we perform one of the first direct comparisons of conservation-based and structure-based methods. We find that the two approaches provide largely complementary information, which can be combined to improve upon either approach alone. We also demonstrate that ConCavity has state-of-the-art performance in predicting catalytic sites and drug binding pockets. Overall, the algorithms and analysis presented here significantly improve our ability to identify ligand binding sites and further advance our understanding of the relationship between evolutionary sequence conservation and structural and functional attributes of proteins. Data, source code, and prediction visualizations are available on the ConCavity web site (http://compbio.cs.princeton.edu/concavity/). John A. Capra, Roman A. Laskowski, Janet M. Thornton, Mona Singh 0001, Thomas A. Funkhouser |
PLoS Comput. Biol. | 3 |
| 2009 | PoreWalker: A Novel Tool for the Identification and Characterization of Channels in Transmembrane Proteins from Their Three-Dimensional StructureabstractTransmembrane channel proteins play pivotal roles in maintaining the homeostasis and responsiveness of cells and the cross-membrane electrochemical gradient by mediating the transport of ions and molecules through biological membranes. Therefore, computational methods which, given a set of 3D coordinates, can automatically identify and describe channels in transmembrane proteins are key tools to provide insights into how they function.Herein we present PoreWalker, a fully automated method, which detects and fully characterises channels in transmembrane proteins from their 3D structures. A stepwise procedure is followed in which the pore centre and pore axis are first identified and optimised using geometric criteria, and then the biggest and longest cavity through the channel is detected. Finally, pore features, including diameter profiles, pore-lining residues, size, shape and regularity of the pore are calculated, providing a quantitative and visual characterization of the channel. To illustrate the use of this tool, the method was applied to several structures of transmembrane channel proteins and was able to identify shape/size/residue features representative of specific channel families. The software is available as a web-based resource at http://www.ebi.ac.uk/thornton-srv/software/PoreWalker/. Marialuisa Pellegrini-Calace, Tim Maiwald, Janet M. Thornton |
PLoS Comput. Biol. | 3 |
| 2008 | The Protein Feature Ontology: a tool for the unification of protein feature annotationsabstractMOTIVATION: The advent of sequencing and structural genomics projects has provided a dramatic boost in the number of uncharacterized protein structures and sequences. Consequently, many computational tools have been developed to help elucidate protein function. However, such services are spread throughout the world, often with standalone web pages. Integration of these methods is needed and so far this has not been possible as there was no common vocabulary available that could be used as a standard language. RESULTS: The Protein Feature Ontology has been developed to provide a structured controlled vocabulary for features on a protein sequence or structure and comprises approximately 100 positional terms, now integrated into the Sequence Ontology (SO) and 40 non-positional terms which describe features relating to the whole-protein sequence. In addition, post-translational modifications are described by using a pre-existing ontology, the Protein Modification Ontology (MOD). This ontology is being used to integrate over 150 distinct annotations provided by the BioSapiens Network of Excellence, a consortium comprising 19 partner sites in Europe. AVAILABILITY: The Protein Feature Ontology can be browsed by accessing the ontology lookup service at the European Bioinformatics Institute (http://www.ebi.ac.uk/ontology-lookup/browse.do?ontName=BS). Gabrielle A. Reeves, Karen Eilbeck, Michele Magrane, Claire O'Donovan, Luisa Montecchi-Palazzi, Midori A. Harris, Sandra E. Orchard, Rafael C. Jiménez, Andreas Prlic, Tim J. P. Hubbard, Henning Hermjakob, Janet M. Thornton |
Bioinform. | 12 |
| 2007 | Analysis of binding site similarity, small-molecule similarity and experimental binding profiles in the human cytosolic sulfotransferase familyabstractMOTIVATION: In the present work we combine computational analysis and experimental data to explore the extent to which binding site similarities between members of the human cytosolic sulfotransferase family correlate with small-molecule binding profiles. Conversely, from a small-molecule point of view, we explore the extent to which structural similarities between small molecules correlate to protein binding profiles. RESULTS: The comparison of binding site structural similarities and small-molecule binding profiles shows that proteins with similar small-molecule binding profiles tend to have a higher degree of binding site similarity but the latter is not sufficient to predict small-molecule binding patterns, highlighting the difficulty of predicting small-molecule binding patterns from sequence or structure. Likewise, from a small-molecule perspective, small molecules with similar protein binding profiles tend to be topologically similar but topological similarity is not sufficient to predict their protein binding patterns. These observations have important consequences for function prediction and drug design. Rafael Najmanovich, Abdellah Allali-Hassani, Richard J. Morris, Ludmila Dombrovsky, Patricia W. Pan, Masoud Vedadi, Alexander N. Plotnikov, Aled M. Edwards, Cheryl H. Arrowsmith, Janet M. Thornton |
Bioinform. | 10 |
| 2007 | Variation of geometrical and physicochemical properties in protein binding pockets and their ligandsabstractPhysicochemical complementarity is commonly believed to be the driving force for molecular binding. The complementarity for example of electrostatic potentials is regarded as the force that draws the ligand from the solvent into the binding site [ 1 ]. If this hypothesis is true, the same ligand should encounter complementarity environmental properties in all proteins to which it binds. We have used our recently published ligand and binding pocket matching algorithm [ 2 ] to test this common assumption by searching for property distributions that are similar for the same ligand bound to different proteins. The algorithm bases on real spherical harmonic functions, which are applicable to approximate any property function on a unit sphere. These property functions can either be of geometrical or physicochemical nature. For our current analysis we used the shape of binding pockets to test their geometrical similarity and mapped electrostatic, van der Waals and hydrophobicity potentials of the protein on the ligand surface to simulate the physicochemical forces that a ligand may feel in its binding site. It was discovered that, of these properties the two that vary least for a given ligand are the binding conformation of the ligand followed by the shape of the binding pocket. Conversely, the same ligand encountered very different electrostatic and van der Waals potential environments in the different proteins to which it is bound. These properties were often found not to be complementary to the ligand's properties, which is in conflict with the general assumption stated above. However, the hydrophobicity of the binding pocket did seem to correlate with the properties of the ligand bound to the protein. Hydrophobic parts of the ligand are often confronted with hydrophobic parts of the protein, giving rise to similar hydrophobicity distributions within different binding pockets binding the same ligand (see Figure 1 ). A set of Adenosine-mono-phosphate (AMP) ligands bound to non-homologous binding sites is shown . Each row displays different geometrical and physicochemical properties of the binding site and ligand respectively. From top to bottom are shown the variation of the ligand shape, the binding pocket shape, the hydrophobicity of the protein mapped on the ligand shape, the van der Waals potential and the electrostatic potential both again mapped on the ligand shape. The properties were ordered according to their average degree of similarity among the different binding pockets from highest to lowest from top to bottom. In addition the AMP binding pockets were ranked according to the similarity of their bound ligand to the AMP ligand of the Protein Data Bank [3] structure 1amu [4]. These results demonstrate that binding sites that bind the same ligand can exhibit a large variation of properties by facing different physicochemical forces within different binding sites. The results urge a re-evaluation of the total contribution of some physicochemical properties to molecular recognition and the factors that drive molecular binding. Abdullah Kahraman, Richard J. Morris, Roman A. Laskowski, Janet M. Thornton |
BMC Bioinform. | 4 |
| 2007 | Construction, Visualisation, and Clustering of Transcription Networks from Microarray Expression DataabstractNetwork analysis transcends conventional pairwise approaches to data analysis as the context of components in a network graph can be taken into account. Such approaches are increasingly being applied to genomics data, where functional linkages are used to connect genes or proteins. However, while microarray gene expression datasets are now abundant and of high quality, few approaches have been developed for analysis of such data in a network context. We present a novel approach for 3-D visualisation and analysis of transcriptional networks generated from microarray data. These networks consist of nodes representing transcripts connected by virtue of their expression profile similarity across multiple conditions. Analysing genome-wide gene transcription across 61 mouse tissues, we describe the unusual topography of the large and highly structured networks produced, and demonstrate how they can be used to visualise, cluster, and mine large datasets. This approach is fast, intuitive, and versatile, and allows the identification of biological relationships that may be missed by conventional analysis techniques. This work has been implemented in a freely available open-source application named BioLayout Express(3D). Tom C. Freeman, Leon Goldovsky, Markus Brosch, Stijn van Dongen, Pierre Mazière, Russell J. Grocock, Shiri Freilich, Janet M. Thornton, Anton J. Enright |
PLoS Comput. Biol. | 8 |
| 2007 | Evolutionary Models for Formation of Network Motifs and Modularity in the Saccharomyces Transcription Factor NetworkabstractMany natural and artificial networks contain overrepresented subgraphs, which have been termed network motifs. In this article, we investigate the processes that led to the formation of the two most common network motifs in eukaryote transcription factor networks: the bi-fan motif and the feed-forward loop. Around 100 million y ago, the common ancestor of the Saccharomyces clade underwent a whole-genome duplication event. The simultaneous duplication of the genes created by this event enabled the origin of many network motifs to be established. The data suggest that there are two primary mechanisms that are involved in motif formation. The first mechanism, enabled by the substantial plasticity in promoter regions, is rewiring of connections as a result of positive environmental selection. The second is duplication of transcription factors, which is also shown to be involved in the formation of intermediate-scale network modularity. These two evolutionary processes are complementary, with the pre-existence of network motifs enabling duplicated transcription factors to bind different targets despite structural constraints on their DNA-binding specificities. This process may facilitate the creation of novel expression states and the increases in regulatory complexity associated with higher eukaryotes. Jonathan J. Ward, Janet M. Thornton |
PLoS Comput. Biol. | 2 |
| 2005 | HTHquery: a method for detecting DNA-binding proteins with a helix-turn-helix structural motifabstractSUMMARY: HTHquery is a web-based service to determine if a protein structure has a helix-turn-helix structural motif which could bind to DNA. It is based on a similarity with a set of structural templates, the accessibility of a putative structural motif and a positive electrostatic potential in the neighbourhood of the putative motif. A set of scores are computed, based on each template, using a linear predictor. From the training set used, the predictor has a true positive rate of 83.5% and a false positive rate of 0.8%. The emphasis for the website is on providing a straightforward interface which can be easily used by a bench-based scientist. AVAILABILITY: HTHquery is implemented using a set of Perl scripts and C program and can be accessed freely on the website http://www.ebi.ac.uk/thornton-srv/databases/HTHquery. Carles Ferrer-Costa, Hugh P. Shanahan, Susan Jones, Janet M. Thornton |
Bioinform. | 4 |
| 2005 | MACiE: a database of enzyme reaction mechanismsabstractSUMMARY: MACiE (mechanism, annotation and classification in enzymes) is a publicly available web-based database, held in CMLReact (an XML application), that aims to help our understanding of the evolution of enzyme catalytic mechanisms and also to create a classification system which reflects the actual chemical mechanism (catalytic steps) of an enzyme reaction, not only the overall reaction. AVAILABILITY: http://www-mitchell.ch.cam.ac.uk/macie/. Gemma L. Holliday, Gail J. Bartlett, Daniel E. Almonacid, Noel M. O'Boyle, Peter Murray-Rust, Janet M. Thornton, John B. O. Mitchell |
Bioinform. | 6 |
| 2005 | Real spherical harmonic expansion coefficients as 3D shape descriptors for protein binding pocket and ligand comparisonsabstractMOTIVATION: An increasing number of protein structures are being determined for which no biochemical characterization is available. The analysis of protein structure and function assignment is becoming an unexpected challenge and a major bottleneck towards the goal of well-annotated genomes. As shape plays a crucial role in biomolecular recognition and function, the examination and development of shape description and comparison techniques is likely to be of prime importance for understanding protein structure-function relationships. RESULTS: A novel technique is presented for the comparison of protein binding pockets. The method uses the coefficients of a real spherical harmonics expansion to describe the shape of a protein's binding pocket. Shape similarity is computed as the L2 distance in coefficient space. Such comparisons in several thousands per second can be carried out on a standard linux PC. Other properties such as the electrostatic potential fit seamlessly into the same framework. The method can also be used directly for describing the shape of proteins and other molecules. AVAILABILITY: A limited version of the software for the real spherical harmonics expansion of a set of points in PDB format is freely available upon request from the authors. Binding pocket comparisons and ligand prediction will be made available through the protein structure annotation pipeline Profunc (written by Roman Laskowski) which will be accessible from the EBI website shortly. Richard J. Morris, Rafael Najmanovich, Abdullah Kahraman, Janet M. Thornton |
Bioinform. | 4 |
| 2005 | Living longer by dieting: analysis of transcriptional response after caloric restrictionabstractData generation for systems biology: functional genomics C loric restriction extends mean and maximum lifespan in a range of eukaryotic species, including yeast, flies and mice, and retards age-associated pathologies such as cancer in mice. However, the molecular mechanisms for this are not well understood. In rodents, it has been suggested that there is little similarity between the transcriptional responses of different tissues to caloric restriction, but this has not been examined simultaneously in the same sample population. We used gene expression arrays to determine the transcriptional profiles of liver, skeletal muscle, hypothalamus and colon in mice subjected to caloric restriction for 48 hours. All of these tissues are known to be affected by caloric restriction and all appear to be important to the metabolic and physiological adaptations that occur in response to caloric restriction. We compared the transcription changes between our tissues in two ways: based on overlap between lists of differentially expressed genes and overlap in functional annotation categories overrepresented in lists of differentially expressed genes (calculated using EASE). from BioSysBio: Bioinformatics and Systems Biology Conference Edinburgh, UK, 14–15 July 2005 Nicola D. Kerrison, Colin Selman, Steven J. Lingard, Janet M. Thornton, Linda Partridge, Dominic J. Withers |
BMC Bioinform. | 4 |
| 2004 | Software Engineering Challenges in BioinformaticsabstractData from biological research is proliferating rapidly and advanced data storage and analysis methods are required to manage it. We introduce the main sources of biological data available and outline some of the domain specific problems associated with automated analysis. We discuss two major areas in which we are likely experience software engineering challenges over the next ten years: data integration and presentation. Jonathan A. Barker, Janet M. Thornton |
ICSE | 2 |
| 2004 | A practical and robust sequence search strategy for structural genomics target selectionabstractMOTIVATION: Target selection strategies for structural genomic projects must be able to prioritize gene regions on the basis of significant sequence similarity with proteins that have already been structurally determined. With the rapid development of protein comparison software a robust prioritization scheme should be independent of the choice of algorithm and be able to incorporate different sequence similarity thresholds. RESULTS: A robust target selection strategy has been developed that can assign a priority level to all genes in any genome. Structural assignments to genome sequences are calculated at two thresholds and six levels (1-6) describe the prioritization of all whole genes and partial gene regions. This simple two-threshold approach can be implemented with any fold recognition or homology detection algorithms. The results for 10 genomes are presented using the SSEARCH and PSI-BLAST programs. AVAILABILITY: Programs are available on request from the authors. James E. Bray, Russell L. Marsden, Stuart C. G. Rison, Alexei Savchenko, Aled M. Edwards, Janet M. Thornton, Christine A. Orengo |
Bioinform. | 6 |
| 2004 | An examination of the conservation of surface patch polarity for proteinsabstractMOTIVATION: The solubility of a protein is crucial for its function and is therefore an evolutionary constraint. As the solubility of a protein is related to the distribution of polar and hydrophobic residues on its solvent accessible surface, such a constraint should provide a valuable insight into the evolution of protein surfaces. We examine how the surfaces of proteins have evolved by considering how the average hydrophobicities of patches of surface residues vary across homologous proteins. We derive distributions for the average hydrophobicity/philicity of surface patches at a residue-based level-which we refer to as the residue hydrophobic density. This is computed for a set of 28 monomeric proteins and their homologues. The resulting distributions are compared with a set of randomized sequences, with the same residue content. RESULTS: We find that the patches, involving typically more than 10 residues, maintain a more hydrophilic surface than one would expect from a random substitution model, indicating a cooperative behaviour for these surfaces residues in terms of this single variable. SUPPLEMENTARY INFORMATION: Additional plots for all of the proteins examined in this paper can be found at: http://www.ebi.ac.uk/~shanahan/PCon/index.html Hugh P. Shanahan, Janet M. Thornton |
Bioinform. | 2 |
| 2003 | An algorithm for constraint-based structural template matching: application to 3D templates with statistical analysisabstractMOTIVATION: Structural templates consisting of a few atoms in a specific geometric conformation provide a powerful tool for studying the relationship between protein structure and function. Current methods for template searching constrain template syntax and semantics by their design. Hence there is a need for a more flexible core algorithm upon which to build more sophisticated tools. Statistical analysis of structural similarity is still in its infancy when compared with its analogue in sequence alignment. In the context of template matching, there is an urgent need for normalization of scores so that results from templates with differing sensitivity may be compared directly. RESULTS: We introduce Jess, a fast and flexible algorithm for searching protein structures for small groups of atoms under arbitrary constraints on geometry and chemistry. We apply the algorithm to a set of manually derived enzyme active site templates, and derive an empirical measure for estimating the relative significance of hits encountered using differing templates. Jonathan A. Barker, Janet M. Thornton |
Bioinform. | 2 |
| 2003 | Recognizing the fold of a protein structureabstractThis paper reports a graph-theoretic program, GRATH, that rapidly, and accurately, matches a novel structure against a library of domain structures to find the most similar ones. GRATH generates distributions of scores by comparing the novel domain against the different types of folds that have been classified previously in the CATH database of structural domains. GRATH uses a measure of similarity that details the geometric information, number of secondary structures and number of residues within secondary structures, that any two protein structures share. Although GRATH builds on well established approaches for secondary structure comparison, a novel scoring scheme has been introduced to allow ranking of any matches identified by the algorithm. More importantly, we have benchmarked the algorithm using a large dataset of 1702 non-redundant structures from the CATH database which have already been classified into fold groups, with manual validation. This has facilitated introduction of further constraints, optimization of parameters and identification of reliable thresholds for fold identification. Following these benchmarking trials, the correct fold can be identified with the top score with a frequency of 90%. It is identified within the ten most likely assignments with a frequency of 98%. GRATH has been implemented to use via a server (http://www.biochem.ucl.ac.uk/cgi-bin/cath/Grath.pl). GRATH's speed and accuracy means that it can be used as a reliable front-end filter for the more accurate, but computationally expensive, residue based structure comparison algorithm SSAP, currently used to classify domain structures in the CATH database. With an increasing number of structures being solved by the structural genomics initiatives, the GRATH server also provides an essential resource for determining whether newly determined structures are related to any known structures from which functional properties may be inferred. Andrew P. Harrison, Frances M. G. Pearl, Ian Sillitoe, Tim Slidel, Richard Mott, Janet M. Thornton, Christine A. Orengo |
Bioinform. | 6 |
| 1999 | Motif-based searching in TOPS protein topology databasesabstractMOTIVATION: TOPS cartoons are a schematic ion of protein three-dimensional structures in two dimensions, and are used for understanding and manual comparison of protein folds. Recently, an algorithm that produces the cartoons automatically from protein structures has been devised and cartoons have been generated to represent all the structures in the structural databank. There is now a need to be able to define target topological patterns and to search the database for matching domains. RESULTS: We have devised a formal language for describing TOPS diagrams and patterns, and have designed an efficient algorithm to match a pattern to a set of diagrams. A pattern-matching system has been implemented, and tested on a database derived from all the current entries in the Protein Data Bank (15,000 domains). Users can search on patterns selected from a library of motifs or, alternatively, they can define their own search patterns. AVAILABILITY: The system is accessible over the Web at http://tops.ebi.ac.uk/tops David R. Gilbert, David R. Westhead, Nozomi Nagano, Janet M. Thornton |
Bioinform. | 4 |
| 1992 | The rapid generation of mutation data matrices from protein sequencesabstractAn efficient means for generating mutation data matrices from large numbers of protein sequences is presented here. By means of an approximate peptide-based sequence comparison algorithm, the set sequences are clustered at the 85% identity level. The closest relating pairs of sequences are aligned, and observed amino acid exchanges tallied in a matrix. The raw mutation frequency matrix is processed in a similar way to that described by Dayhoff et al. (1978), and so the resulting matrices may be easily used in current sequence analysis applications, in place of the standard mutation data matrices, which have not been updated for 13 years. The method is fast enough to process the entire SWISS-PROT databank in 20 h on a Sun SPARCstation 1, and is fast enough to generate a matrix from a specific family or class of proteins in minutes. Differences observed between our 250 PAM mutation data matrix and the matrix calculated by Dayhoff et al. are briefly discussed. David T. Jones, William R. Taylor, Janet M. Thornton |
Comput. Appl. Biosci. | 3 |