EDBT 2026 Demo / reviewers in the wild / expert
Alex Bateman
dblp:77/2690
· DBLP profile ↗
48ranked-venue papers
15as first author
5since 2021 · last 2026
0000-0002-6982-4660ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 48 · 15 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GOFlowLLM - curating miRNA literature with large language models and flowchartsabstractMOTIVATION: The exponential growth of non-coding RNA research-with over 230 000 papers published since 2000-has created an urgent knowledge management crisis in molecular biology. Despite their crucial regulatory roles, microRNAs (miRNAs) face a significant curation bottleneck, with only 1400 articles manually curated to the Gene Ontology (GO) knowledgebase over a decade. This highlights the critical need for automated systems that can accelerate biocuration while maintaining high-quality standards. RESULTS: We present GOFlowLLM, an automated curation pipeline powered by reasoning-enabled Large Language Models (LLMs) that follows established GO curation flowcharts to extract and structure miRNA-mediated gene silencing data at scale. When evaluated on existing curation, GOFlowLLM selects the correct GO term in 90% of cases, with curators agreeing with 95% of the system's reasoning steps and 90% of the evidence selected. Applied to 6996 previously uncurated articles using the Qwen QwQ-32B model, our system identified 2538 new candidate GO annotations on 1785 articles in just 58 hours-potentially doubling the available miRNA GO curation. Manual review shows curators agreed with the selected term in 87% of cases, the model's reasoning in 92% of cases, and the extracted evidence in 93%. The integration of reasoning traces provides transparent justification for annotations that can be reviewed by human curators, addressing a key challenge in adopting AI for scientific curation. AVAILABILITY AND IMPLEMENTATION: GOFlowLLM is implemented as an automated pipeline that follows expert-designed reasoning frameworks to maintain curation quality. The system is available on GitHub: https://github.com/RNAcentral/GO_Flow_LLM. Andrew Green 0001, Nancy Ontiveros-Palacios, Isaac Jandalala, Simona Panni, Valerie Wood, Giulia Antonazzo, Helen Attrill, Alex Bateman, Blake A. Sweeney |
Bioinform. | 8 |
| 2023 | Annotation of biologically relevant ligands in UniProtKB using ChEBIabstractMOTIVATION: To provide high quality, computationally tractable annotation of binding sites for biologically relevant (cognate) ligands in UniProtKB using the chemical ontology ChEBI (Chemical Entities of Biological Interest), to better support efforts to study and predict functionally relevant interactions between protein sequences and structures and small molecule ligands. RESULTS: We structured the data model for cognate ligand binding site annotations in UniProtKB and performed a complete reannotation of all cognate ligand binding sites using stable unique identifiers from ChEBI, which we now use as the reference vocabulary for all such annotations. We developed improved search and query facilities for cognate ligands in the UniProt website, REST API and SPARQL endpoint that leverage the chemical structure data, nomenclature and classification that ChEBI provides. AVAILABILITY AND IMPLEMENTATION: Binding site annotations for cognate ligands described using ChEBI are available for UniProtKB protein sequence records in several formats (text, XML and RDF) and are freely available to query and download through the UniProt website (www.uniprot.org), REST API (www.uniprot.org/help/api), SPARQL endpoint (sparql.uniprot.org/) and FTP site (https://ftp.uniprot.org/pub/databases/uniprot/). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Elisabeth Coudert, Sebastien Gehant, Edouard De Castro, Monica Pozzato, Delphine Baratin, Teresa Batista Neto, Christian J. A. Sigrist, Nicole Redaschi, Alan J. Bridge, Lucila Aimo, Ghislaine Argoud-Puy, Andrea H. Auchincloss, Kristian B. Axelsen, Parit Bansal, Marie-Claude Blatter, Jerven T. Bolleman, Emmanuel Boutet, Lionel Breuza, Blanca Cabrera Gil, Cristina Casals-Casas, Kamal Chikh Echioukh, Béatrice A. Cuche, Anne Estreicher, Maria Livia Famiglietti, Marc Feuermann, Elisabeth Gasteiger, Pascale Gaudet, Vivienne Baillie Gerritsen, Arnaud Gos, Nadine Gruaz-Gumowski, Chantal Hulo, Nevila Hyka-Nouspikel, Florence Jungo, Arnaud Kerhornou, Philippe Le Mercier, Damien Lieberherr, Patrick Masson, Anne Morgat, Venkatesh Muthukrishnan, Salvo Paesano, Ivo Pedruzzi, Sandrine Pilbout, Lucille Pourcel, Sylvain Poux, Manuela Pruess, Catherine Rivoire, Karin Sonesson, Shyamala Sundaram, Alex Bateman, Maria Jesus Martin, Sandra E. Orchard, Michele Magrane, Shadab Ahmad, Emanuele Alpi, Emily H. Bowler-Barnett, Ramona Britto, Hema Bye-A-Jee, Austra Cukura, Paul Denny 0002, Tunca Dogan, Thankgod Ebenezer, Penelope Garmiri, Leonardo Jose da Costa Gonzales, Emma Hatton-Ellis, Abdulrahman Hussein, Alexandr Ignatchenko, Giuseppe Insana, Rizwan Ishtiaq, Vishal Joshi, Dushyanth Jyothi, Swaathi Kandasamy, Antonia Lock, Aurelien Luciani, Marija Lugaric, Yvonne Lussi, Alistair MacDougall, Fábio Madeira, Mahdi Mahmoudy, Alok Mishra 0004, Katie Moulang, Andrew Nightingale, Sangya Pundir, Guoying Qi, Shriya Raj, Pedro Raposo, Daniel Rice, Rabie Saidi, Elena Speretta, James D. Stephenson, Prabhat Totoo, Edward Turner, Nidhi Tyagi, Preethi Vasudev, Kate Warner, Xavier Watkins, Rossana Zaru, Hermann Zellner, Cathy H. Wu, Cecilia N. Arighi, Leslie Arminski, Chuming Chen, Yongxing Chen, Hongzhan Huang, Kati Laiho, Peter B. McGarvey, Darren A. Natale, Karen E. Ross, C. R. Vinayaka, Qinghua Wang 0003 |
Bioinform. | 49 |
| 2022 | DPCfam: Unsupervised protein family classification by Density Peak Clustering of large sequence datasetsabstractProteins that are known only at a sequence level outnumber those with an experimental characterization by orders of magnitude. Classifying protein regions (domains) into homologous families can generate testable functional hypotheses for yet unannotated sequences. Existing domain family resources typically use at least some degree of manual curation: they grow slowly over time and leave a large fraction of the protein sequence space unclassified. We here describe automatic clustering by Density Peak Clustering of UniRef50 v. 2017_07, a protein sequence database including approximately 23M sequences. We performed a radical re-implementation of a pipeline we previously developed in order to allow handling millions of sequences and data volumes of the order of 3 TeraBytes. The modified pipeline, which we call DPCfam, finds ∼ 45,000 protein clusters in UniRef50. Our automatic classification is in close correspondence to the ones of the Pfam and ECOD resources: in particular, about 81% of medium-large Pfam families and 72% of ECOD families can be mapped to clusters generated by DPCfam. In addition, our protocol finds more than 14,000 clusters constituted of protein regions with no Pfam annotation, which are therefore candidates for representing novel protein families. These results are made available to the scientific community through a dedicated repository. Elena Tea Russo, Federico Barone, Alex Bateman, Stefano Cozzini, Marco Punta, Alessandro Laio |
PLoS Comput. Biol. | 3 |
| 2021 | Computational strategies to combat COVID-19: useful tools to accelerate SARS-CoV-2 and coronavirus researchabstractSARS-CoV-2 (severe acute respiratory syndrome coronavirus 2) is a novel virus of the family Coronaviridae. The virus causes the infectious disease COVID-19. The biology of coronaviruses has been studied for many years. However, bioinformatics tools designed explicitly for SARS-CoV-2 have only recently been developed as a rapid reaction to the need for fast detection, understanding and treatment of COVID-19. To control the ongoing COVID-19 pandemic, it is of utmost importance to get insight into the evolution and pathogenesis of the virus. In this review, we cover bioinformatics workflows and tools for the routine detection of SARS-CoV-2 infection, the reliable analysis of sequencing data, the tracking of the COVID-19 pandemic and evaluation of containment measures, the study of coronavirus evolution, the discovery of potential drug targets and development of therapeutic strategies. For each tool, we briefly describe its use case and how it advances research specifically for SARS-CoV-2. All tools are free to use and available online, either through web applications or public code repositories. Contact:[email protected]. Franziska Hufsky, Kevin Lamkiewicz, Alexandre Almeida, Abdel Aouacheria, Cecilia N. Arighi, Alex Bateman, Jan Baumbach, Niko Beerenwinkel, Christian Brandt, Marco Cacciabue, Sara Chuguransky, Oliver Drechsel, Robert D. Finn, Adrian Fritz, Stephan Fuchs, Georges Hattab, Anne-Christin Hauschild, Dominik Heider, Marie Hoffmann, Martin Hölzer, Stefan Hoops, Lars Kaderali, Ioanna Kalvari, Max von Kleist, Renó Kmiecinski, Denise Kühnert, Gorka Lasso, Pieter Libin, Markus List, Hannah F. Löchel, Maria Jesus Martin, Roman Martin, Julian O. Matschinske, Alice C. McHardy, Pedro Mendes 0001, Jaina Mistry, Vincent Navratil, Eric P. Nawrocki, Áine Niamh O'toole, Nancy Ontiveros-Palacios, Anton I. Petrov, Guillermo Rangel-Pineros, Nicole Redaschi, Susanne Reimering, Knut Reinert, Lorna J. Richardson, David L. Robertson, Sepideh Sadegh, Joshua B. Singer, Kristof Theys, Chris Upton, Marius Welzel, Lowri Williams, Manja Marz |
Briefings Bioinform. | 6 |
| 2021 | Ten simple rules to make your computing more environmentally sustainableabstractRule 1: Calculate the carbon footprint of your workWe live in a world ruled by data, where a problem doesn't exist until it has been measured.There is still very limited information available about the carbon footprint of computational Loïc Lannelongue, Jason Grealey, Alex Bateman, Michael Inouye |
PLoS Comput. Biol. | 3 |
| 2020 | The ELIXIR Core Data Resources: fundamental infrastructure for the life sciencesabstractSUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Rachel Drysdale, Charles E. Cook, Robert Petryszak, Vivienne Baillie Gerritsen, Mary Barlow, Elisabeth Gasteiger, Franziska Gruhl, Jerry Lanfear, Rodrigo Lopez, Nicole Redaschi, Heinz Stockinger, Daniel Teixeira, Aravind Venkatesan, Alex Bateman, Alan J. Bridge, Guy Cochrane, Robert D. Finn, Frank Oliver Glöckner, Marc Hanauer, Thomas M. Keane, Luana Licata, Per Oksvold, Sandra E. Orchard, Christine A. Orengo, Helen E. Parkinson, Bengt Persson, Pablo Porras, Jordi Rambla De Argila, Ana Rath, Charlotte Rodwell, Ugis Sarkans, Dietmar Schomburg, Ian Sillitoe, J. Dylan Spalding, Mathias Uhlen, Sameer Velankar, Juan Antonio Vizcaíno, Kalle von Feilitzen, Christian von Mering, Andy Yates, Niklas Blomberg, Christine Durinx, Johanna R. McEntyre |
Bioinform. | 15 |
| 2019 | TADOSS: computational estimation of tandem domain swap stabilityabstractSUMMARY: Proteins with highly similar tandem domains have shown an increased propensity for misfolding and aggregation. Several molecular explanations have been put forward, such as swapping of adjacent domains, but there is a lack of computational tools to systematically analyze them. We present the TAndem DOmain Swap Stability predictor (TADOSS), a method to computationally estimate the stability of tandem domain-swapped conformations from the structures of single domains, based on previous coarse-grained simulation studies. The tool is able to discriminate domains susceptible to domain swapping and to identify structural regions with high propensity to form hinge loops. TADOSS is a scalable method and suitable for large scale analyses. AVAILABILITY AND IMPLEMENTATION: Source code and documentation are freely available under an MIT license on GitHub at https://github.com/lafita/tadoss. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Aleix Lafita, Robert B. Best, Alex Bateman |
Bioinform. | 4 |
| 2017 | On expert curation and scalability: UniProtKB/Swiss-Prot as a case studyabstractMOTIVATION: Biological knowledgebases, such as UniProtKB/Swiss-Prot, constitute an essential component of daily scientific research by offering distilled, summarized and computable knowledge extracted from the literature by expert curators. While knowledgebases play an increasingly important role in the scientific community, their ability to keep up with the growth of biomedical literature is under scrutiny. Using UniProtKB/Swiss-Prot as a case study, we address this concern via multiple literature triage approaches. RESULTS: With the assistance of the PubTator text-mining tool, we tagged more than 10 000 articles to assess the ratio of papers relevant for curation. We first show that curators read and evaluate many more papers than they curate, and that measuring the number of curated publications is insufficient to provide a complete picture as demonstrated by the fact that 8000-10 000 papers are curated in UniProt each year while curators evaluate 50 000-70 000 papers per year. We show that 90% of the papers in PubMed are out of the scope of UniProt, that a maximum of 2-3% of the papers indexed in PubMed each year are relevant for UniProt curation, and that, despite appearances, expert curation in UniProt is scalable. AVAILABILITY AND IMPLEMENTATION: UniProt is freely available at http://www.uniprot.org/. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Sylvain Poux, Cecilia N. Arighi, Michele Magrane, Alex Bateman, Chih-Hsuan Wei, Zhiyong Lu, Emmanuel Boutet, Hema Bye-A-Jee, Maria Livia Famiglietti, Bernd Roechert |
Bioinform. | 4 |
| 2016 | UniProt-DAAC: domain architecture alignment and classification, a new method for automatic functional annotation in UniProtKBabstractMOTIVATION: Similarity-based methods have been widely used in order to infer the properties of genes and gene products containing little or no experimental annotation. New approaches that overcome the limitations of methods that rely solely upon sequence similarity are attracting increased attention. One of these novel approaches is to use the organization of the structural domains in proteins. RESULTS: We propose a method for the automatic annotation of protein sequences in the UniProt Knowledgebase (UniProtKB) by comparing their domain architectures, classifying proteins based on the similarities and propagating functional annotation. The performance of this method was measured through a cross-validation analysis using the Gene Ontology (GO) annotation of a sub-set of UniProtKB/Swiss-Prot. The results demonstrate the effectiveness of this approach in detecting functional similarity with an average F-score: 0.85. We applied the method on nearly 55.3 million uncharacterized proteins in UniProtKB/TrEMBL resulted in 44 818 178 GO term predictions for 12 172 114 proteins. 22% of these predictions were for 2 812 016 previously non-annotated protein entries indicating the significance of the value added by this approach. AVAILABILITY AND IMPLEMENTATION: The results of the method are available at: ftp://ftp.ebi.ac.uk/pub/contrib/martin/DAAC/ CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Tunca Dogan, Alistair MacDougall, Rabie Saidi, Diego Poggioli, Alex Bateman, Claire O'Donovan, Maria Jesus Martin |
Bioinform. | 5 |
| 2014 | Structure and computational analysis of a novel protein with metallopeptidase-like and circularly permuted winged-helix-turn-helix domains reveals a possible role in modified polysaccharide biosynthesisabstractBACKGROUND: CA_C2195 from Clostridium acetobutylicum is a protein of unknown function. Sequence analysis predicted that part of the protein contained a metallopeptidase-related domain. There are over 200 homologs of similar size in large sequence databases such as UniProt, with pairwise sequence identities in the range of ~40-60%. CA_C2195 was chosen for crystal structure determination for structure-based function annotation of novel protein sequence space. RESULTS: The structure confirmed that CA_C2195 contained an N-terminal metallopeptidase-like domain. The structure revealed two extra domains: an α+β domain inserted in the metallopeptidase-like domain and a C-terminal circularly permuted winged-helix-turn-helix domain. CONCLUSIONS: Based on our sequence and structural analyses using the crystal structure of CA_C2195 we provide a view into the possible functions of the protein. From contextual information from gene-neighborhood analysis, we propose that rather than being a peptidase, CA_C2195 and its homologs might play a role in biosynthesis of a modified cell-surface carbohydrate in conjunction with several sugar-modification enzymes. These results provide the groundwork for the experimental verification of the function. Debanu Das, Alexey G. Murzin, Neil D. Rawlings, Robert D. Finn, Penny C. Coggill, Alex Bateman, Adam Godzik, L. Aravind |
BMC Bioinform. | 6 |
| 2013 | Two Pfam protein families characterized by a crystal structure of protein lpg2210 from Legionella pneumophilaabstractBACKGROUND: Every genome contains a large number of uncharacterized proteins that may encode entirely novel biological systems. Many of these uncharacterized proteins fall into related sequence families. By applying sequence and structural analysis we hope to provide insight into novel biology. RESULTS: We analyze a previously uncharacterized Pfam protein family called DUF4424 [Pfam:PF14415]. The recently solved three-dimensional structure of the protein lpg2210 from Legionella pneumophila provides the first structural information pertaining to this family. This protein additionally includes the first representative structure of another Pfam family called the YARHG domain [Pfam:PF13308]. The Pfam family DUF4424 adopts a 19-stranded beta-sandwich fold that shows similarity to the N-terminal domain of leukotriene A-4 hydrolase. The YARHG domain forms an all-helical domain at the C-terminus. Structure analysis allows us to recognize distant similarities between the DUF4424 domain and individual domains of M1 aminopeptidases and tricorn proteases, which form massive proteasome-like capsids in both archaea and bacteria. CONCLUSIONS: Based on our analyses we hypothesize that the DUF4424 domain may have a role in forming large, multi-component enzyme complexes. We suggest that the YARGH domain may play a role in binding a moiety in proximity with peptidoglycan, such as a hydrophobic outer membrane lipid or lipopolysaccharide. Penny C. Coggill, Ruth Y. Eberhardt, Robert D. Finn, Yuanyuan Chang, Lukasz Jaroszewski, Adam Godzik, Debanu Das, Qingping Xu, Herbert L. Axelrod, L. Aravind, Alexey G. Murzin, Alex Bateman |
BMC Bioinform. | 12 |
| 2013 | Filling out the structural map of the NTF2-like superfamilyabstractBACKGROUND: The NTF2-like superfamily is a versatile group of protein domains sharing a common fold. The sequences of these domains are very diverse and they share no common sequence motif. These domains serve a range of different functions within the proteins in which they are found, including both catalytic and non-catalytic versions. Clues to the function of protein domains belonging to such a diverse superfamily can be gleaned from analysis of the proteins and organisms in which they are found. RESULTS: Here we describe three protein domains of unknown function found mainly in bacteria: DUF3828, DUF3887 and DUF4878. Structures of representatives of each of these domains: BT_3511 from Bacteroides thetaiotaomicron (strain VPI-5482) [PDB:3KZT], Cj0202c from Campylobacter jejuni subsp. jejuni serotype O:2 (strain NCTC 11168) [PDB:3K7C], rumgna_01855) and RUMGNA_01855 from Ruminococcus gnavus (strain ATCC 29149) [PDB:4HYZ] have been solved by X-ray crystallography. All three domains are similar in structure and all belong to the NTF2-like superfamily. Although the function of these domains remains unknown at present, our analysis enables us to present a hypothesis concerning their role. CONCLUSIONS: Our analysis of these three protein domains suggests a potential non-catalytic ligand-binding role. This may regulate the activities of domains with which they are combined in the same polypeptide or via operonic linkages, such as signaling domains (e.g. serine/threonine protein kinase), peptidoglycan-processing hydrolases (e.g. NlpC/P60 peptidases) or nucleic acid binding domains (e.g. Zn-ribbons). Ruth Y. Eberhardt, Yuanyuan Chang, Alex Bateman, Alexey G. Murzin, Herbert L. Axelrod, William C. Hwang, L. Aravind |
BMC Bioinform. | 3 |
| 2013 | LUD, a new protein domain associated with lactate utilizationabstractBACKGROUND: A novel highly conserved protein domain, DUF162 [Pfam: PF02589], can be mapped to two proteins: LutB and LutC. Both proteins are encoded by a highly conserved LutABC operon, which has been implicated in lactate utilization in bacteria. Based on our analysis of its sequence, structure, and recent experimental evidence reported by other groups, we hereby redefine DUF162 as the LUD domain family. RESULTS: JCSG solved the first crystal structure [PDB:2G40] from the LUD domain family: LutC protein, encoded by ORF DR_1909, of Deinococcus radiodurans. LutC shares features with domains in the functionally diverse ISOCOT superfamily. We have observed that the LUD domain has an increased abundance in the human gut microbiome. CONCLUSIONS: We propose a model for the substrate and cofactor binding and regulation in LUD domain. The significance of LUD-containing proteins in the human gut microbiome, and the implication of lactate metabolism in the radiation-resistance of Deinococcus radiodurans are discussed. William C. Hwang, Constantina Bakolitsa, Marco Punta, Penny C. Coggill, Alex Bateman, Herbert L. Axelrod, Neil D. Rawlings, Mayya Sedova, Scott N. Peterson, Ruth Y. Eberhardt, L. Aravind, Jaime Pascual, Adam Godzik |
BMC Bioinform. | 5 |
| 2013 | ISCB Computational Biology Wikipedia CompetitionabstractThe International Society for Computational Biology is pleased to announce the 2013 ISCB Computational Biology Wikipedia competition. The competition, in which entrants create or improve the content of any Wikipedia article in the field of computational biology, is open to all students and trainees. Further information about the competition can be found here: http://en.wikipedia.org/wiki/Wikipedia:WikiProject_Computational_Biology/ISCB_competition_announcement_2013
The mission of the ISCB is to promote the use of computational biology and to help educate the next generation of computational biologists. The society has numerous activities that help to address these aims, including conferences, training and mentoring initiatives, and an active student council.
As the world's largest online encyclopedia, Wikipedia has become an indispensable resource for those seeking information on all scientific and technical topics. The English language version of Wikipedia contains over 4.2 million articles, and Wikipedia is now available in 286 languages. The global rise in smartphone use, which allows access to Wikipedia, means that a large fraction of the world's population can now gain access to the world's knowledge. Wikipedia is the most successful example of crowd-sourcing with about 80,000 active editors updating its content each month.
But is Wikipedia a good source of information for computational biology? Certainly, many people are reading the articles. For example, the Bioinformatics article has been visited 1,600 times per day over the last 3 months. Wikipedia contains articles on algorithms, biological databases, software packages, and biographies of eminent computational biologists. The computational biology content ranges from incomplete, a mere “stub” of an article in Wikipedia parlance, to highly detailed Featured Articles. A group of Wikipedia editors have formed the Computational Biology Wikiproject (http://en.wikipedia.org/wiki/Wikipedia:WikiProject_Computational_Biology). This group oversees the computational biology articles and rates them for their importance and their quality. Figure 1 shows the current state of the articles (see also Figure S1). In total, there are over 1,140 articles that have been considered as falling under Computational Biology. There are a small number of articles that have been brought up to the highest levels of quality (Featured Article and Good Article) such as Multiple Sequence Alignment, Genome Wide Association Study, and Folding@home.
Figure 1
The computational biology articles rated by quality and importance by the Wikipedia Computational Biology Wikiproject.
The 2012 competition began 9th September 2012 (coinciding with the start of the European Conference on Computational Biology) and finished four months later on the 10th January 2013. Each article entered in the competition was reviewed for a difference in article quality between these two dates. In 2012, there were 13 substantive entries into the competition. Six of these articles were shortlisted by members of the ISCB Student Council and then considered by the judging panel. The judging panel considered articles based on the criteria of clarity of the writing, depth of knowledge of the subject, and quality of figures and images used. In one case, it was clear that the article was largely derived from a published review, and was not considered further. For the other entries, the quantity and quality of the contributions were very good, and it was a challenge to rank the articles. After much deliberation, the judging panel selected the following as the winners of the 2012 ISCB Wikipedia competition:
1st prize: James Estevez for improvements to the Genomics Article.
2nd prize: Benjamin Moore for improvements to the European Nucleotide Archive article.
3rd prize: Luis Pedro Coelho for improvements to the Bioimage Analysis article.
We are keen to grow the depth and quality of computational biology articles and wish to encourage the widest possible range of students and trainees to take part. We envisage that teachers, tutors, and lecturers could use the competition as an opportunity to train students in literature research on topics of computational biology. This approach to literature review provides the students with a thorough grounding in the subject area of the article. In addition, the collaborative writing environment of Wikipedia encourages critical thinking and improves literature research skills. Furthermore, compared to traditional literature reviews carried out by students, which typically end up unread in a filing cabinet, contributing to Wikipedia means that the students' scholarly contributions will be publicly visible.
We hope that the ISCB Wikipedia competition will continue to grow and help improve the quality of Computational Biology information freely available on the Internet. We are interested in improving not just the articles in Wikipedia, but also the associated media, such as images and figures on Wikimedia Commons, and data through Wikidata. We encourage you to get involved by either entering the competition if you are a student or trainee, or getting your own students to participate. Alex Bateman, Janet Kelso, Daniel Mietchen, Geoff MacIntyre, Tomás Di Domenico, Thomas Abeel, Darren W. Logan, Predrag Radivojac, Burkhard Rost |
PLoS Comput. Biol. | 1 |
| 2012 | Bioimage informatics: a new category in BioinformaticsabstractThe last two decades have witnessed great advances in biological tissue labeling and automated microscopic imaging that, in turn, have revolutionized how biologists visualize molecular, sub-cellular, cellular, and super-cellular structures and study their respective functions. Tremendous volumes of multi-dimensional bioimaging data are now being generated in almost every branch of biology. How to interpret such image datasets in a quantitative, objective, automatic and efficient way has become a major challenge in current computational biology. Bioimage informatics methods have begun to turn image data into useful biological knowledge (Peng, 2008; Swedlow, et al., 2009; Shamir, et al., 2010; Danuser, 2011). The essential methods of bioimage informatics involve large-scale bioimage generation, visualization, analysis and management. Bioimage informatics also encompasses both hypothesis- and data-driven exploratory approaches, with an emphasis on how to generate biological knowledge and/or gain new insights that would otherwise be hard to achieve.
Early work in bioimage informatics began in the late 1990s. Increasingly, computer vision, image analysis, data mining, machine learning and pattern recognition methods have been applied to microscopic images to extract biological information and to generate ontology databases. The growing amount of bioimage data are quickly imposing additional demands on how to store, manage and retrieve such image datasets as well as the associated secondary meta-data. Data analysis, fusion and reconstruction techniques have also been developed to facilitate better image acquisition and formation. Joint analysis of image data in combination with other biological datasets, such as genomes and gene expression profiles, is also becoming more and more commonplace.
To meet the need of this growing field, the first international workshop on Bioimage Informatics was organized at Stanford University in 2005. It grew to be an annual event in this field. Other meetings on similar topics and related applications have also emerged since then. In 2010, the annual conference on Intelligent Systems for Molecular Biology (ISMB) established a paper-submission track on bioimaging data analysis and visualization.
While there is a noticeable need to publish high quality papers on bioimage informatics, so far no high-impact journal explicitly accepts this category of papers. We believe it is an appropriate time to create this new category in Bioinformatics. As of February 2012, Bioinformatics now includes a new paper submission category in the scope described by the journal at its website as follows:
‘Informatics methods for the acquisition, analysis, mining and visualization of images produced by modern microscopy, with an emphasis on the application of novel computing techniques to solve challenging and significant biological and medical problems at the molecular, sub-cellular, cellular, and super-cellular (organ, organism, and population) levels. This category also encourages large-scale image informatics methods/applications/software, various enabling techniques (e.g. cyber-infrastructures, quantitative validation experiments, pattern recognition, etc.) for such large-scale studies, and joint analysis of multiple heterogeneous datasets that include images as a component. Bioimage related ontology and databases studies, image-oriented large-scale machine learning, data mining, and other analytics techniques are also encouraged.
We will not consider image analysis and pattern recognition methods that are solely based on tuning parameters or swapping computational sub-steps, without an in-depth description or demonstration of why such changes are significantly superior for one or more biological problems.’
We would like to thank a number of colleagues and practitioners who have contributed to the creation of this new category. Especially, we thank Eugene Myers, B.S.Manjunath, Badri Roysam, Manfred Auer, Michael Hawrylycz, Jean-Christophe Olivo-Marin, Anne Carpenter and Vebjorn Ljosa in helping to define this new category of paper submissions. We hope the journal Bioinformatics becomes a valuable venue for bioimage informatics researchers to publish their most important work.
Contact: gro.imhh.ailenaj@hgnep Hanchuan Peng, Alex Bateman, Alfonso Valencia, Jonathan D. Wren |
Bioinform. | 2 |
| 2011 | The rise and fall of supervised machine learning techniquesabstractMachine learning is of immense importance in bioinformatics and biomedical science more generally (Larrañaga et al., 2006; Tarca et al., 2007). In particular, supervised machine learning has been used to great effect in numerous bioinformatics prediction methods. Through many years of editing and reviewing manuscripts, we noticed that some supervised machine learning techniques seem to be gaining in popularity while others seemed, at least to our eyes, to be looking ‘unfashionable’. We were motivated to create a league table of machine learning techniques to learn what is hot and what is not in the machine learning field. In this editorial, we only include those that we considered major league and leave analysis of the minor league methods as an exercise for the interested reader. To create our league table, we created a list of supervised machine learning techniques commonly used in bioinformatics and their common synonyms, plural forms and abbreviations. We then searched this list against the PubMed titles and abstracts to identify the number of papers published per year for each machine learning technique. To match as many papers as possible, searches were case insensitive and allowed for variation in hyphenation. To our surprise, the artificial neural network (ANN) is not only the dominant league leader in 2011 but has been in this position since at least the 1970s (see Fig. 1). However, in recent years the usage of support vector machines (SVMs) grew tremendously, and we predict that SVMs will challenge ANNs for the dominant position in the coming decade. Since 2007 the number of publications using ANNs has decreased by 21%, which we hypothesize may be directly attributed to researchers increasingly using SVMs in place of ANNs. SVMs caught up with and overtook Markov models in 2004 to gain second spot in our machine learning league. The growth of supervised machine learning methods in PubMed. As for the question of ‘what is hot?’, one can see that Random forests are a rapidly growing method with not a single mention of them before 2003 and now a total of 407 papers published to date. We were hoping to find techniques that were not so hot and perhaps going out of fashion. The results show that none of the major league methods has gone out of fashion, but we do see moderate decreases in the use of both ANNs and Markov models in the literature. We were also curious to find out if certain machine learning techniques were used in combination with each other. To investigate this, we looked at what machine learning methods are co-mentioned in articles (See Fig. 2). For all pairs of methods from the Supervised Machine Learning Top-5, we counted the number of abstracts that mention both methods and normalized the counts with the number of co-occurrences that would be expected by chance (based on the frequencies with which the methods are mentioned over the years). The strongest correlation (185 times higher than random expectation) is seen between decision trees and random forests, which is to be expected as random forests are ensembles of decision trees. Apart from this, the next strongest correlation (88 times higher than random expectation) is found between the two newest methods on the list, namely SVMs and random forests. We hypothesize that this is due to many researchers using these algorithms through machine learning frameworks such as Weka (Frank et al., 2004), which allows many different algorithms to easily be applied to the same dataset. Heatmap showing the co-occurrence of machine learning techniques within articles. Applications of supervised machine learning methodology continue to grow in the biomedical literature. Despite new methods growing in usage, for example support vector machines and random forests, we see little evidence that any widely adopted methods are falling out of use. Conflict of Interest: none declared. Lars Juhl Jensen, Alex Bateman |
Bioinform. | 2 |
| 2010 | Curators of the world unite: the International Society of BiocurationabstractWe often take the wealth of biological data that is available for granted. As computational biologists, we can often search the Internet and find a dataset that will fulfill our needs. These are usually made available through one of the hundreds of biological databases that have been created over the last 20 years. The people who marshal this information and put it into the formats that allow us to easily work with it are the unsung heroes of molecular biology. The recently formed International Society of Biocuration (ISB) gives these people a voice. Biocuration can be summed up as the transformation of biological data into an organized form. It is only achieved through the combined efforts of the experimental community who generate the data, the biocurators who organize the data and the software and database developers who make the data available for all to use. The ISB grew out of the International Biocuration Conferences, which provided a forum for biocurators to discuss the scientific obstacles as well as new developments in the field. The next meeting is the Fourth International Biocurator Meeting that will be held in Chiba, Japan, in October 2010. Further details can be found on the ISB's web site at http://www.biocurator.org/. I would like to encourage all biocurators and developers with an interest in curation to register as a member of the ISB and play a part in the growth of this exciting new body. The Society was founded in early 2009. The first election of the executive board was held in September 2009, which is now composed of Pascale Gaudet (President), Lorna Richardson (Secretary), Lydie Bougueleret (Treasurer), Terri Attwood, Tanya Berardini, Tadashi Imanishi, Owen White, Ioannis Xenarios and myself. The mission of the ISB is to define the work of biocurators; propose discussion and job forums; organize conferences and workshops; build relationships with journals and publishers to improve links between journals and databases; provide documentation and gold standards; and foster connections with user communities to ensure that databases meet their needs. A further important purpose is working to raise awareness of biocuration and biological databases with funding agencies to help secure long-term funding for biocuration. Despite the billions spent each year on generating biological data, there is still a reluctance to invest in the relatively small fraction of funding needed to maximize the use of this data through curation. Next time you download a dataset for your work, spare a thought for the hardworking biocurator that has made your life so much easier. If you are a biocurator or database developer, then please join the ISB to support their important work. I would like to wish the International Society of Biocuration a warm welcome and every future success. Conflict of Interest: A.B. is a member of the Executive Board of the ISB. Alex Bateman |
Bioinform. | 1 |
| 2010 | Ten Simple Rules for Editing WikipediaabstractWikipedia is the world's most successful online encyclopedia, now containing over 3.3 million English language articles. It is probably the largest collection of knowledge ever assembled, and is certainly the most widely accessible. Wikipedia can be edited by anyone with Internet access that chooses to, but does it provide reliable information? A 2005 study by Nature found that a selection of Wikipedia articles on scientific subjects were comparable to a professionally edited encyclopedia [1], suggesting a community of volunteers can generate and sustain surprisingly accurate content.
For better or worse, people are guided to Wikipedia when searching the Web for biomedical information [2]. So there is an increasing need for the scientific community to engage with Wikipedia to ensure that the information it contains is accurate and current. For scientists, contributing to Wikipedia is an excellent way of fulfilling public engagement responsibilities and sharing expertise. For example, some Wikipedian scientists have successfully integrated biological data with Wikipedia to promote community annotation [3], [4]. This, in turn, encourages wider access to the linked data via Wikipedia. Others have used the wiki model to develop their own specialist, collaborative databases [5]–[8]. Taking your first steps into Wikipedia can be daunting, but here we provide some tips that should make the editing process go smoothly. Darren W. Logan, Massimo Sandal, Paul P. Gardner, Magnus Manske, Alex Bateman |
PLoS Comput. Biol. | 5 |
| 2009 | Phospholipid scramblases and Tubby-like proteins belong to a new superfamily of membrane tethered transcription factorsabstractMOTIVATION: Phospholipid scramblases (PLSCRs) constitute a family of cytoplasmic membrane-associated proteins that were identified based upon their capacity to mediate a Ca(2+)-dependent bidirectional movement of phospholipids across membrane bilayers, thereby collapsing the normally asymmetric distribution of such lipids in cell membranes. The exact function and mechanism(s) of these proteins nevertheless remains obscure: data from several laboratories now suggest that in addition to their putative role in mediating transbilayer flip/flop of membrane lipids, the PLSCRs may also function to regulate diverse processes including signaling, apoptosis, cell proliferation and transcription. A major impediment to deducing the molecular details underlying the seemingly disparate biology of these proteins is the current absence of any representative molecular structures to provide guidance to the experimental investigation of their function. RESULTS: Here, we show that the enigmatic PLSCR family of proteins is directly related to another family of cellular proteins with a known structure. The Arabidopsis protein At5g01750 from the DUF567 family was solved by X-ray crystallography and provides the first structural model for this family. This model identifies that the presumed C-terminal transmembrane helix is buried within the core of the PLSCR structure, suggesting that palmitoylation may represent the principal membrane anchorage for these proteins. The fold of the PLSCR family is also shared by Tubby-like proteins. A search of the PDB with the HHpred server suggests a common evolutionary ancestry. Common functional features also suggest that tubby and PLSCR share a functional origin as membrane tethered transcription factors with capacity to modulate phosphoinositide-based signaling. Alex Bateman, Robert D. Finn, Peter J. Sims, Therese Wiedmer, Andreas Biegert, Johannes Söding |
Bioinform. | 1 |
| 2009 | EditorialabstractThe transforming aspect of the Human Genome Project was not the completion of the genome sequence itself, but rather the technologies that enabled, and were enabled by, the sequencing of that first reference genome. The evolution of ‘omic science through microarray transcriptomics, metabolomics, proteomics, and whole-genome SNP-omics has in many ways come full circle with a new focus on genomics and genome sequencing. Next-generation sequencing technologies have begun to revolutionise genomics and their effects are becoming increasingly widespread. The 1000 genomes project (http://www.1000genomes.org/) will create a new map of genetic variation for our genome going far beyond the detail captured in the HapMap. Other projects are helping to catalogue genes involved in cancer, alternative splicing in different tissues and transcription factor binding, for example. The growing number of robust applications and the steadily falling cost for generating sequence-based data suggest that these next-generation technologies will continue to rapidly open new applications in the biological sciences and generate new opportunities for software and algorithm development. Given the vast amount of data produced (currently greater than a gigabase per run, with this constantly increasing as well), developing a sound data storage and management solution and creating informatics tools to effectively analyze the data are essential to successful application of the technology. During the past year, a large number of new software applications and algorithms have been developed to deal with this new data. A recent advert in Nature from the Illumina, one of the providers of next-generation sequencing technology, highlighted significant papers in the area of bioinformatics; our journal published 7 of the 16 listed papers. In addition to those cited in the advertisement, there have been many other tools and algorithms published in Bioinformatics that are relevant to next-generation sequencing applications. To celebrate this contribution we have gathered these together in a ‘Bioinformatics for Next Generation Sequencing’ virtual issue (http://www.oxfordjournals.org/our_journals/bioinformatics/nextgenerationsequencing.html). This will be a living resource that we will continually update to include the very latest papers in this area to help researchers keep abreast of the latest developments. To date, the majority of the papers have described methods to take the short sequences produced by the Illumina Genome Analyzer and Applied Biosystems SOLiD machines and align them to a reference genome. This is a crucial and basic requirement for many applications and a variety of techniques have been applied to make the tools sufficiently fast to deal with millions of sequences. We have also included papers that address the issue of assembly of these short reads. Now that there are many of these tools available the Bioinformatics community has begun to make applications that are useful for specific applications such as identifying likely sites of interaction in CHIP-seq. A summary of the inaugural collection is included in the Table 1. We sincerely hope that you find this resource useful and that the collected references lead to additional development in an area that we view as critical to the continued development of genomics and bioinformatics. Tools recently published in the journal Tools recently published in the journal Alex Bateman, John Quackenbush |
Bioinform. | 1 |
| 2009 | Cloud computingabstractContact: [email protected] We are used to having huge datasets pouring out of high-throughput genome centres, but with the advent of ultra high-throughput sequencing, genotyping and other functional genomics in every laboratory we are facing a scary new era in petabyte scale data. For example, the 1000 genomes' projects will probably produce about 1 Tb of finished data. To process data, this project required about 100 Tbs of scratch disk. Working at this level, real technical limitations start to hamper progress. One has to consider storage, but not just having enough, but making sure its available to your compute (network), that you have sufficient I/O to do anything in real time. Software language and implementation become critical factors when dealing with terabytes of data. With such high-intensity computing, power (getting enough), cooling, etc. become real issues. How do you let anyone else access the data? Is the data backed up and even if it is how many years would it take to restore from tape? So how will we solve all these technical hurdles? Each of these can be solved with technical knowledge. But you do not want to have to worry about working within these constraints. When working with large datasets, these constraints can continually hamper progress on getting real research done. Whilst one can choose to solve each of these individual problems, the impact of these constrains on the scientific workflow can be considerable. It would be wiser to optimize for productivity. In software development, similar constrains are addressed with abstraction layers. Database access is mediated through relational mapping tools, visualization is aided with powerful graphical packages preventing individual research groups from having to reinvent the wheel. Rails, Eclipse, Processing, Hibernate, Catalyst. Cloud computing offers a similar level of abstraction for many of the constraints encountered when dealing with extremely large (?) datasets. You might have encountered similar ideas when using hosted services such as Google Mail, ManyEyes (http://manyeyes.alphaworks.ibm.com), others. These tools provide an example of what we would ideally like in the perfect world of Bioinformatics. We do not have to worry about how the data is stored, keeping the software up to date. Its all taken care of for you. First steps have been taken along these lines by companies such as Amazon, Google and Microsoft. Amazon has started to provide Bioinformatics datasets in their publicly hosted datasets (http://aws.amazon.com/publicdatasets/) such as Ensembl and Genbank. A recent requirement to assemble a full human genome from 454 short read data provided a good real life example of these approaches. With 140 million individual reads requiring alignment using SSAHA exceeding the available compute capacity in our own data centre, a build was performed on Amazon's elastic compute service, EC2. In an afternoon, a scalable, ad hoc cluster with queue management and replicated data storage was constructed with nothing more than a few web service calls and a valid credit card. No service contracts. No consultation with the vendor, just 100 nodes performing SSAHA alignments. Implications for large scale data centres, currently engineered to provide peak capacity, which often goes unused in idle periods. The elastic, pay-as-you-go nature of cloud services such as AWS means lower infrastructure overheads, as only in-use compute and storage is billable. Cloud computing has green credentials too, so long as the off-site compute is located where renewable sources of energy are used preferentially. Additionally, whilst unused compute may still require cooling and power in a local data centre, it can be reused by others in the cloud. The transfer of large datasets can also be simplified with cloud approaches. As an alternative to shipping the data for others to analyse, cloud approaches allow the compute to remain close to the data. Allowing others to access your compute infrastructure may be preferable to distributing large datasets. Conflict of Interest: none declared. Alex Bateman, Matt Wood |
Bioinform. | 1 |
| 2009 | Ten Simple Rules for Chairing a Scientific SessionabstractChairing a session at a scientific conference is a thankless task.If you get it right, no one is likely to notice.But there are many ways to get it wrong and a little preparation goes a long way to making the session a success.Here are a few pointers that we have picked up over the years. Alex Bateman, Philip E. Bourne |
PLoS Comput. Biol. | 1 |
| 2008 | Pfam 10 years on: 10 000 families and still growingabstractClassifications of proteins into groups of related sequences are in some respects like a periodic table for biology, allowing us to understand the underlying molecular biology of any organism. Pfam is a large collection of protein domains and families. Its scientific goal is to provide a complete and accurate classification of protein families and domains. The next release of the database will contain over 10,000 entries, which leads us to reflect on how far we are from completing this work. Currently Pfam matches 72% of known protein sequences, but for proteins with known structure Pfam matches 95%, which we believe represents the likely upper bound. Based on our analysis a further 28,000 families would be required to achieve this level of coverage for the current sequence database. We also show that as more sequences are added to the sequence databases the fraction of sequences that Pfam matches is reduced, suggesting that continued addition of new families is essential to maintain its relevance. Stephen John Sammut, Robert D. Finn, Alex Bateman |
Briefings Bioinform. | 3 |
| 2008 | Databases, data tombs and dust in the windabstractUNLABELLED: As biomedical data accumulates, the need to store, share and organize it grows. Consequently, the number of Internet-accessible databases has been rapidly growing on an annual basis. Bioinformatics regularly publishes descriptions of biomedically relevant databases, Nucleic Acids Research has published an annual database issue since 1996 and now a new open-access journal, DATABASE: The Journal of Biological DATABASEs and Curation, will soon be launched by Oxford University Press in 2009 (http://www.oxfordjournals.org/our_journals/databa/). Since databases can be made publicly available on the Internet without publication, it is worth considering what factors prioritize publication of database descriptions in a peer-reviewed journal. In general, publication of a database description in a journal advertises it as a valuable resource for scientific research. Implicitly, it is assumed that this resource is publicly available (most likely for free) and will be maintained. However, therein lies the problem: DATABASE papers are simply not of the same nature as regular research articles. Over time, some databases simply become inaccessible, some are created but not maintained or updated, and some databases are never used (Galperin, 2006). Thus, for database creators, reviewers and journal editors, there are several additional considerations to judge, prior to publication, how potentially valuable these new databases may be. Jonathan D. Wren, Alex Bateman |
Bioinform. | 2 |
| 2007 | SCOOP: a simple method for identification of novel protein superfamily relationshipsabstractMOTIVATION: Profile searches of sequence databases are a sensitive way to detect sequence relationships. Sophisticated profile-profile comparison algorithms that have been recently introduced increase search sensitivity even further. RESULTS: In this article, a simpler approach than profile-profile comparison is presented that has a comparable performance to state-of-the-art tools such as COMPASS, HHsearch and PRC. This approach is called SCOOP (Simple Comparison Of Outputs Program), and is shown to find known relationships between families in the Pfam database as well as detect novel distant relationships between families. Several novel discoveries are presented including the discovery that a domain of unknown function (DUF283) found in Dicer proteins is related to double-stranded RNA-binding domains. AVAILABILITY: SCOOP is freely available under a GNU GPL license from http://www.sanger.ac.uk/Users/agb/SCOOP/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Alex Bateman, Robert D. Finn |
Bioinform. | 1 |
| 2007 | EditorialabstractDuring 2006 we have received over 2000 manuscript submissions. Our acceptance rate over the past year has been close to 30%. We are pleased to see that more authors are submitting novel biological insights through our Discovery Notes articles, including a growing number of novel protein domain discoveries. We are proud to have published a number of special issues including full volumes dedicated to the two main conferences in the field: ISMB and ECCB 2006. We are also pleased to note that the latest impact factor for Bioinformatics from the ISI (Institute for Scientific Information) has continued to increase, from 5.7 to 6.0. Papers are published rapidly online ahead of print, within one week of acceptance on average, and usually appear in the print version of the journal within 12 weeks. The journal has continued to publish Open Access papers as part of the Oxford Open initiative (Author Webpage). During 2006, over 20% of Bioinformatics authors have chosen to publish their paper under the Open Access model. This is the highest uptake seen by any journal in the optional Oxford Open initiative. In addition, all content is made freely available online 12 months after publication. In the New Year we will increase the number of review articles in key topics for our community and we are grateful to Jonathan Wren the Associate Editor who is coordinating this. Over the last year we have continued to bring on board new Associate Editors to replace those who are stepping down. Over the past year five editors have stepped down. We would like to thank Charlie Hodgman, Nikolaus Rajewsky, Alvis Brazma, Satoru Miyano and Christos Ouzounis for their great contribution to the journal. We have extended the terms of Martin Bishop, John Quackenbush and Thomas Lengauer for a further three years. We would like to welcome Trey Ideker, Olga Troyanskaya, Limsoon Wong and Burkhard Rost, our new Asso\ciate Editors who have joined through the year. In particular we would like to mention the immense contribution of Christos Ouzounis in shaping Bioinformatics into what it is now. Along with Martin Bishop, Christos was a very active handling editor of the journal for the many years in which the office was at the EBI in Hinxton with Chris Sander as Executive Editor. Among many other contributions Christos was the one behind the creation of the Discovery Notes section, the association with ISMB for the publication of their annual special issue, and for the creation of the Bioinformatics Associate Editor role. He will now continue working with the journal as member of the Editorial Board from his new position in Tessaloniki. We would also like to thank our Editorial Board for helpful advice and suggestions throughout the year. Finally we would like to give our thanks to the many thousands of referees who have freely given up their time to review and improve the work submitted to the journal. Alfonso Valencia, Alex Bateman |
Bioinform. | 2 |
| 2007 | Predicting active site residue annotations in the Pfam databaseabstractBACKGROUND: Approximately 5% of Pfam families are enzymatic, but only a small fraction of the sequences within these families (<0.5%) have had the residues responsible for catalysis determined. To increase the active site annotations in the Pfam database, we have developed a strict set of rules, chosen to reduce the rate of false positives, which enable the transfer of experimentally determined active site residue data to other sequences within the same Pfam family. DESCRIPTION: We have created a large database of predicted active site residues. On comparing our active site predictions to those found in UniProtKB, Catalytic Site Atlas, PROSITE and MEROPS we find that we make many novel predictions. On investigating the small subset of predictions made by these databases that are not predicted by us, we found these sequences did not meet our strict criteria for prediction. We assessed the sensitivity and specificity of our methodology and estimate that only 3% of our predicted sequences are false positives. CONCLUSION: We have predicted 606110 active site residues, of which 94% are not found in UniProtKB, and have increased the active site annotations in Pfam by more than 200 fold. Although implemented for Pfam, the tool we have developed for transferring the data can be applied to any alignment with associated experimental active site data and is available for download. Our active site predictions are re-calculated at each Pfam release to ensure they are comprehensive and up to date. They provide one of the largest available databases of active site annotation. Jaina Mistry, Alex Bateman, Robert D. Finn |
BMC Bioinform. | 2 |
| 2007 | Reuse of structural domain-domain interactions in protein networksabstractBACKGROUND: Protein interactions are thought to be largely mediated by interactions between structural domains. Databases such as iPfam relate interactions in protein structures to known domain families. Here, we investigate how the domain interactions from the iPfam database are distributed in protein interactions taken from the HPRD, MPact, BioGRID, DIP and IntAct databases. RESULTS: We find that known structural domain interactions can only explain a subset of 4-19% of the available protein interactions, nevertheless this fraction is still significantly bigger than expected by chance. There is a correlation between the frequency of a domain interaction and the connectivity of the proteins it occurs in. Furthermore, a large proportion of protein interactions can be attributed to a small number of domain interactions. We conclude that many, but not all, domain interactions constitute reusable modules of molecular recognition. A substantial proportion of domain interactions are conserved between E. coli, S. cerevisiae and H. sapiens. These domains are related to essential cellular functions, suggesting that many domain interactions were already present in the last universal common ancestor. CONCLUSION: Our results support the concept of domain interactions as reusable, conserved building blocks of protein interactions, but also highlight the limitations currently imposed by the small number of available protein structures. Benjamin Schuster-Böckler, Alex Bateman |
BMC Bioinform. | 2 |
| 2006 | Bioinformatics - The new home for protein sequence motifsabstractProtein domains and sequence motifs have been very influential in the field of molecular biology. These units are the common currency of protein structure and function. The Protein Sequence Motifs series in Trends in Biochemical Sciences was an extremely popular forum to publish these results, but sadly the journal published the final one in 2004 (McEntyre and Gibson 2004). Bioinformatics is well placed to publish reports of novel protein domains and over the past year we have published reports of the MEDs and PocR domains (Anantharaman and Aravind, 2005), the G5 domain (Bateman et al., 2005), the Why domain (Ciccarelli and Bork 2005) and the OCRE domain (Callebaut and Mornon 2005). In this issue we publish the PilZ domain (pronounced ‘pills’) by Amikam and Galperin (2005). This manuscript strongly suggests that the PilZ domain mediates binding to c-di-GMP a universal bacterial second messenger. Because the PilZ is found in hundreds of bacterial proteins this work will have a large impact in understanding bacteria. I would like to invite you to submit your next novel domain finding as a Discovery Note to Bioinformatics. We will also consider papers that describe the unification of existing protein families and domains into a larger superfamily. Discovery notes are papers intended for the reporting of biologically interesting discoveries using computational techniques. Discovery notes can be up to 4 journal pages length. The instructions to authors can be found at the following URL: Author Webpage If you would like a quick assessment of your domain discovery please send me a pre-submission enquiry outlining the domain and the biological significance of the finding. Alex Bateman |
Bioinform. | 1 |
| 2006 | Structural genomics meets computational biologyabstractA meeting recently organized by the NIH NIGMS Protein Structure Initiative (PSI, Author Webpage) has made crystal clear the urgency and importance of the development of computational methods for the analysis of protein families, definition of protein domains and regions for expression, and annotation of protein function. No really new problems, but problems made now even more important for the development of the Structural Genomics projects. PSI is now in the first year of the production phase (after a 5 year pilot project) with a projected annual cost of 66 million dollars. PSI is now composed of four large production centres: Northeast Structural Genomics Consortium (Author Webpage), Midwest Center for Structural Genomics (Author Webpage), Joint Center for Structural Genomics (Author Webpage) and New York Structural GenomiX Research Consortium (Author Webpage), and six technology-based centers, modeling centers and a database (KnowledgeBase) and material repository associated project to be funded in September 2006. The project has the core activities around a number of committees, including the target selection Steering Subcommittee that includes the Bioinformatics group comprising our colleagues A. Godzik, A. Fiser, C. Orengo and B. Rost. This subcommittee was the one that organized the meeting to discuss the best strategies for selecting proteins to enter into the protein structure resolution pipeline of the PSI centers. It was also this committee that was lucky enough to choose the four most rainy days in Bethesda in the last 50 years. Target selection is technically and scientifically an important issue that was discussed under the growing impression that the biological community at large is not well informed of the progress and are apparently not aware of the developments and achievements during the pilot phase of this project. We were convinced during the meeting that the appropriate selection of sensible and clearly define targets will certainly contribute positively to increase the interest of biologist in the developments in structural genomics, and that the interest that can be generated by a clear description of the targets to be solved will have to be reinforced/rewarded/maintained by providing access to the right metrics of progress and success. In this scenario computational biology is essential not only to set the objectives and select the protein targets in the most effective possible way, but also, and perhaps more importantly, to make accessible to the community all the information generated by the SG projects. A number of specific problems related to the selection of targets and the definition of milestones were discussed during the meeting. These problems include the definition of protein families at different levels of granularity, the range of sequences that can be modeled with a given structure, the available strategies for evaluating the quality of the models and the strategies for selecting the more interesting targets in large protein families. The issues related with the biological/biomedical interest of the targets and the possibilities in the difficult issue of measuring the coverage of the function space were also considered to be of great interest for the definition of the target selection strategy. It was rewarding for us to see how all these problems of fundamental importance for the development of SG are directly part of the realm of bioinformatics/computational biology. It is also to be said that it is a great community responsibility to redouble our efforts to make available additional methods and resources to address these problems. Specific proposals such as the organization of an annual conference to stimulate the work on the subject related with target selection and analysis with the scientific community outside the SG projects will certainly be steps in the right direction. Fostering the undergoing efforts in the SG projects to make openly available, and easily accessible to computational biologist, the results of the on-going experiments will also create a very positive flow of research and development. For example, it is urgent to make publicly accessible the biophysical results obtained for the many constructions that the SG centers have tested for expression, solubility and crystallization. This resource can be invaluable for the development of domain boundary prediction methods, which in turn will contribute to speed-up the experimental work with complex eukaryotic proteins. Finally, given the importance of SG for the development of biology and biomedicine, and the fundamental importance of bioinformatics for the organization, analysis and exploitation of the results, it is unavoidable to think that the effort dedicated in the different Structural Genomics international projects to the computational analysis should be increased urgently. Alex Bateman, Alfonso Valencia |
Bioinform. | 1 |
| 2006 | Software patents in BioinformaticsabstractBioinformatics has published papers describing new software for over 20 years (Nilsson and Klein 1985). During this time the world of software has changed considerably particularly with the irresistible rise of initiatives to build freely accessible software as has opening access to data resources. The Internet and the Web have also changed the way we use and distribute software. This social and technical revolution is also changing the structure of the relations between commercial and academic software-based activities, for which patents and software protection are key elements. In this and the following issue we publish two editorials addressing the general topics of software accessibility, patents and intellectual property. In this issue, Steven L. Salzberg and John Quackenbush (past and present Associate Editors, respectively) present one perspective on the issue. In the next issue another of our Associate Editors, Jonathan D. Wren will put forward a different perspective. We welcome additional contributions to this discussion from our readership, which will help the journal in the process of adapting our publication guidelines to better serve the development of Bioinformatics. Alfonso Valencia, Alex Bateman |
Bioinform. | 2 |
| 2005 | The G5 domain: a potential N-acetylglucosamine recognition domain involved in biofilm formationabstractSUMMARY: Biofilms are complex microbial communities found at surfaces that are often associated with extracellular polysaccharides. Biofilm formation is a complex process that is being understood at the molecular level only recently. We have identified a novel domain that we call the G5 domain (named after its conserved glycine residues), which is found in a variety of enzymes such as Streptococcal IgA peptidases and various glycosyl hydrolases in bacteria. The G5 domain is found in the Accumulation Associated Protein (AAP), which is an important component in biofilm formation in Staphylococcus aureus. A common feature of the proteins containing G5 domains is N-acetylglucosamine binding, and we attribute this function to the G5 domain. CONTACT: [email protected]. Alex Bateman, Matthew T. G. Holden, Corin Yeats |
Bioinform. | 1 |
| 2005 | An update from the Bioinformatics EditorsabstractIn this editorial we take the opportunity to highlight changes in the journal Bioinformatics, during 2005. During the past couple of years the journal had accepted more papers than it could publish and this resulted in a backlog of manuscripts waiting some months to appear in print. We are very pleased to say that this has been relieved by publishing four large issues in April and May of this year. It now takes only 10 weeks from acceptance until publication in the print issue; manuscripts are also rapidly published online ahead of print (on the ‘Advance Access’ page) within 4 days of acceptance on average. We have now taken steps to carefully monitor the acceptance levels in the journal to improve quality even further and to ensure that we do not increase the acceptance to print times in the future. The current acceptance rate is 25%. We have also improved our review process so that 80% of submissions receive a final editorial decision in 40 days. During 2005 we have published special supplement issues of the journal containing the proceedings of both ISMB and ECCB conferences. The journal also occasionally publishes special sections where a small number of papers from a conference are included in a regular issue. Currently these are submitted on an ad hoc basis by conference organisers. We wish now to regularise this process and hereby request proposals for conference proceeding papers for publication in 2006. The deadline for these proposals is 30th January 2006 and we also welcome preliminary proposals for conferences in 2007. Details about the type of information we will require about conference proposals can be found in the journal's Instructions to Authors. It is expected that the conference papers put forward for publication in the journal will be peer-reviewed by the organisers, liaising with a Bioinformatics Associate Editor, to ensure a high standard. Since July 2005, and following the successful experience of our sister journal Nucleic Acids Research, we have launched a new Open Access option. Authors can now choose whether or not to publish their work ‘open access’. For more information about the Oxford Open initiative visit Author Webpage. If an author does not choose the Open Access option their paper will be made freely available online twelve months following publication. We are convinced that this important decision will best satisfy the desires of our authors, as expressed in our author survey (read about the results of this survey in Bioinformatics, 22, 4071–4072). During 2005, the following new Associate Editors have joined Bioinformatics: Keith Crandall, Joaquin Dopazo, Dmitrij Frishman, Chris Stoeckert and Anna Tramontano. More recently we welcome on board the following new Associate Editors: David Rocke, an expert in statistics with particular dedication to DNA array analysis, Jonathan Wren, a bioinformatician now working in areas related with information extraction and text mining, and Golan Yona, a computer scientist particularly interested in the analysis of networks. Additionally during the year the following scientists have joined the journal's Editorial Board: former Associate Editors Russ Altman, Carlos D Bustamante, Gary Williams and Michael Zhang, and Mikhail Gelfand, Adam Godzik, Jaap Heringa, Ina Koch, Wentian Li, Isidore Rigoutsos, Burkhard Rost, N Srinivasan, Olga Troyanskaya and Limsoon Wong. The Associate Editors are responsible for arranging the peer review of submissions, making editorial decisions and working together to decide the future direction and policies of the journal. To increase the transparency of the editorial process, and following suggestions from our readers and authors, we have recently decided to make known the name of the Associate Editor responsible for each manuscript, and to include their names on appropriate papers in the published version. The Editorial Board members also play an active role in Bioinformatics, acting as the main consulting body for journal policies, aims and scope, and helping with difficult editorial decisions. We would like to thank those who are stepping down in 2005: Associate Editors Phil Bourne, Frank Dudbridge, Steen Knudsen, and Steve Salzberg, and Editorial Board members Terry Gaasterland, Mike Gribskov, Steven Henikoff, Webb Miller, and Eugene Myers. Without their dedication and hard work it would not be possible to produce this journal. Alex Bateman, Alfonso Valencia |
Bioinform. | 1 |
| 2005 | iPfam: visualization of protein?Cprotein interactions in PDB at domain and amino acid resolutionsabstractSUMMARY: There are many resources that contain information about binary interactions between proteins. However, protein interactions are defined by only a subset of residues in any protein. We have implemented a web resource that allows the investigation of protein interactions in the Protein Data Bank structures at the level of Pfam domains and amino acid residues. This detailed knowledge relies on the fact that there are a large number of multidomain proteins and protein complexes being deposited in the structure databases. The resource called iPfam is hosted within the Pfam UK website. Most resources focus on the interactions between proteins; iPfam includes these as well as interactions between domains in a single protein. AVAILABILITY: iPfam is available on the Web for browsing at http://www.sanger.ac.uk/Software/Pfam/iPfam/; the source-data for iPfam is freely available in relational tables via the ftp site ftp://ftp.sanger.ac.uk/pub/databases/Pfam/database_files/. Robert D. Finn, Mhairi Marshall, Alex Bateman |
Bioinform. | 3 |
| 2005 | Visualizing profile-profile alignment: pairwise HMM logosabstractUNLABELLED: The availability of advanced profile-profile comparison tools, such as PRC or HHsearch demands sophisticated visualization tools not presently available. We introduce an approach built upon the concept of HMM logos. The method illustrates the similarities of pairs of protein family profiles in an intuitive way. Two HMM logos, one for each profile, are drawn one upon the other. The aligned states are then highlighted and connected. AVAILABILITY: A web interface offering online creation of pairwise HMM logos is available at http://www.sanger.ac.uk/Software/analysis/logomat-p. Furthermore, software developers may download a Perl package that includes methods for creation of pairwise HMM logos locally. CONTACT: [email protected]. Benjamin Schuster-Böckler, Alex Bateman |
Bioinform. | 2 |
| 2005 | Increasing the Impact of BioinformaticsabstractThe year 2004 has been very successful for Bioinformatics. The journal's latest impact factor from the Institute for Scientific Information has increased from 4.615 to 6.701. This is quite an exceptional increase reflecting the increasing standard of work in the journal as well as the increasing stature of the field. Early in 2004, we implemented a new system for Advance online access which allows scientists to access research in Bioinformatics as rapidly as possible. The number of manuscripts submitted to the journal continues to grow. In 2003 we received 1300 submissions while in 2004 we received over 1800 submissions. That we have coped with this large increase is testament to the dedication and hard work of our team of Associate Editors, referees and the Editorial Office. To cope with future growth we are adapting our editorial structure and processes. From 2005 the journal will appear 24 times per year rather than the 18 issues per year previously. This will allow us to publish more of the high quality research and applications that are being submitted to us each day. However, the increase in submissions has outstripped the increase in journal pages. This means that our acceptance rate is decreasing to ∼20%. We are in the process of appointing new Associate Editors to replace those who have stepped down and to increase our coverage of new areas. We would like to thank Gert Vriend, Debbie Marks and Fritz Roth for their invaluable contribution to the journal. We are also in the process of expanding the membership of our Editorial Board. In 2005, we will be introducing a new scheme of categories for papers. During the submission process authors will be asked to choose which category their paper belongs to. This will improve the assignment of manuscripts to editors as well as helping us to formulate a clear definition of the scope of the journal within each category, thus improving organization, in terms of layout and editorial process. Open Access is a topic that is very important to many of our authors and readers. Along this line, we as Editors, and Oxford University Press as a not-for-profit academic publisher, are very much in favour of the principle of making scientific publications freely available. For a well-established journal, as Bioinformatics is now, it is also important to preserve the reputation and financial security of the journal to which authors, referees, editors and readers have contributed over the many years. During 2005 we will learn a great deal from the experience of our sister journal, Nucleic Acids Research, which is introducing a full Open Access model. We are also seeking the opinions of our authors and readers through a survey exploring publication models. We will be sure to base our decision on whether to move forward with an Open Access initiative on the response from our readers, authors and their institutions. So please let us know your views. If we are encouraged by the journal's community to experiment with Open Access, and with a carefully studied new business model in place, we see Open Access as a real opportunity for the journal in the near future. This significant collection of changes will make 2005 an important year for the journal. Our new cover design reflects this spirit of change which will make our journal, and the bioinformatics field, more open to science. Alfonso Valencia, Alex Bateman |
Bioinform. | 2 |
| 2004 | New Leadership for Bioinformatics
Alex Bateman, Alfonso Valencia |
Bioinform. | 1 |
| 2004 | Enhanced protein domain discovery using taxonomyabstractBACKGROUND: It is well known that different species have different protein domain repertoires, and indeed that some protein domains are kingdom specific. This information has not yet been incorporated into statistical methods for finding domains in sequences of amino acids. RESULTS: We show that by incorporating our understanding of the taxonomic distribution of specific protein domains, we can enhance domain recognition in protein sequences. We identify 4447 new instances of Pfam domains in the SP-TREMBL database using this technique, equivalent to the coverage increase given by the last 8.3% of Pfam families and to a 0.7% increase in the number of domain predictions. We use PSI-BLAST to cross-validate our new predictions. We also benchmark our approach using a SCOP test set of proteins of known structure, and demonstrate improvements relative to standard Hidden Markov model techniques. CONCLUSIONS: Explicitly including knowledge about the taxonomic distribution of protein domains can enhance protein domain recognition. Our method can also incorporate other context-specific domain distributions - such as domain co-occurrence and protein localisation. Lachlan James M. Coin, Alex Bateman, Richard Durbin |
BMC Bioinform. | 2 |
| 2004 | The Hotdog fold: wrapping up a superfamily of thioesterases and dehydratasesabstractBACKGROUND: The Hotdog fold was initially identified in the structure of Escherichia coli FabA and subsequently in 4-hydroxybenzoyl-CoA thioesterase from Pseudomonas sp. strain CBS. Since that time structural determinations have shown a number of other apparently unrelated proteins also share the Hotdog fold. RESULTS: Using sequence analysis we unify a large superfamily of HotDog domains. Membership includes numerous prokaryotic, archaeal and eukaryotic proteins involved in several related, but distinct, catalytic activities, from metabolic roles such as thioester hydrolysis in fatty acid metabolism, to degradation of phenylacetic acid and the environmental pollutant 4-chlorobenzoate. The superfamily also includes FapR, a non-catalytic bacterial homologue that is involved in transcriptional regulation of fatty acid biosynthesis. We have defined 17 subfamilies, with some characterisation. Operon analysis has revealed numerous HotDog domain-containing proteins to be fusion proteins, where two genes, once separate but adjacent open-reading frames, have been fused into one open-reading frame to give a protein with two functional domains. Finally we have generated a Hidden Markov Model library from our analysis, which can be used as a tool for predicting the occurrence of HotDog domains in any protein sequence. CONCLUSIONS: The HotDog domain is both an ancient and ubiquitous motif, with members found in the three branches of life. Shane C. Dillon, Alex Bateman |
BMC Bioinform. | 2 |
| 2003 | The TROVE module: A common element in Telomerase, Ro and Vault ribonucleoproteinsabstractBACKGROUND: Ribonucleoproteins carry out a variety of important tasks in the cell. In this study we show that a number of these contain a novel module, that we speculate mediates RNA-binding. RESULTS: The TROVE module--Telomerase, Ro and Vault module--is found in TEP1 and Ro60 the protein components of three ribonucleoprotein particles. This novel module, consisting of one or more domains, may be involved in binding the RNA components of the three RNPs, which are telomerase RNA, Y RNA and vault RNA. A second conserved region in these proteins is shown to be a member of the vWA domain family. The vWA domain in TEP1 is closely related to the previously recognised vWA domain in VPARP a second component of the vault particle. This vWA domain may mediate interactions between these vault components or bind as yet unidentified components of the RNPs. CONCLUSIONS: This work suggests that a number of ribonucleoprotein components use a common RNA-binding module. The TROVE module is also found in bacterial ribonucleoproteins suggesting an ancient origin for these ribonucleoproteins. Alex Bateman, Valerie A. Kickhoefer |
BMC Bioinform. | 1 |
| 2003 | A comparison of Pfam and MEROPS: Two databases, one comprehensive, and one specialisedabstractBACKGROUND: We wished to compare two databases based on sequence similarity: one that aims to be comprehensive in its coverage of known sequences, and one that specialises in a relatively small subset of known sequences. One of the motivations behind this study was quality control. Pfam is a comprehensive collection of alignments and hidden Markov models representing families of proteins and domains. MEROPS is a catalogue and classification of enzymes with proteolytic activity (peptidases or proteases). These secondary databases are used by researchers worldwide, yet their contents are not peer reviewed. Therefore, we hoped that a systematic comparison of the contents of Pfam and MEROPS would highlight missing members and false-positives leading to improvements in quality of both databases. An additional reason for carrying out this study was to explore the extent of consensus in the definition of a protein family. RESULTS: About half (89 out of 174) of the peptidase families in MEROPS overlapped single Pfam families. A further 32 MEROPS families overlapped multiple Pfam families. Where possible, new Pfam families were built to represent most of the MEROPS families that did not overlap Pfam. When comparing the numbers of sequences found in the overlap between a MEROPS family and its corresponding Pfam family, in most cases the overlap was substantial (52 pairs of MEROPS and Pfam families had an intersection size of greater than 75% of the union) but there were some differences in the sets of sequences included in the MEROPS families versus the overlapping Pfam families. CONCLUSIONS: A number of the discrepancies between MEROPS families and their corresponding Pfam families arose from differences in the aims and philosophies of the two databases. Examination of some of the discrepancies highlighted additional members of families, which have subsequently been added in both Pfam and MEROPS. This has led to improvements in the quality of both databases. Overall there was a great deal of consensus between the databases in definitions of a protein family. David J. Studholme, Neil D. Rawlings, Alan J. Barrett, Alex Bateman |
BMC Bioinform. | 4 |
| 2002 | HMM-based databases in InterProabstractProtein family databases are an important resource for protein annotation and understanding protein evolution and function. In recent years hidden Markov models (HMMs) have become one of the key technologies used for detection of members of these families. This paper reviews the Pfam, TIGRFAMs and SMART databases that use the profile-HMMs provided by the HMMER package. Alex Bateman, Daniel H. Haft |
Briefings Bioinform. | 1 |
| 2002 | InterPro: An Integrated Documentation Resource for Protein Families, Domains and Functional SitesabstractThe exponential increase in the submission of nucleotide sequences to the nucleotide sequence database by genome sequencing centres has resulted in a need for rapid, automatic methods for classification of the resulting protein sequences. There are several signature and sequence cluster-based methods for protein classification, each resource having distinct areas of optimum application owing to the differences in the underlying analysis methods. In recognition of this, InterPro was developed as an integrated documentation resource for protein families, domains and functional sites, to rationalise the complementary efforts of the individual protein signature database projects. The member databases - PRINTS, PROSITE, Pfam, ProDom, SMART and TIGRFAMs - form the InterPro core. Related signatures from each member database are unified into single InterPro entries. Each InterPro entry includes a unique accession number, functional descriptions and literature references, and links are made back to the relevant member database(s). Release 4.0 of InterPro (November 2001) contains 4,691 entries, representing 3,532 families, 1,068 domains, 74 repeats and 15 sites of post-translational modification (PTMs) encoded by different regular expressions, profiles, fingerprints and hidden Markov models (HMMs). Each InterPro entry lists all the matches against SWISS-PROT and TrEMBL (2,141,621 InterPro hits from 586,124 SWISS-PROT and TrEMBL protein sequences). The database is freely accessible for text- and sequence-based searches. Nicola J. Mulder, Rolf Apweiler, Terri K. Attwood, Amos Bairoch, Alex Bateman, David Binns, Margaret Biswas, Paul Bradley, Peer Bork, Philipp Bucher, Richard R. Copley, Emmanuel Courcelle, Richard Durbin, Laurent Falquet, Wolfgang Fleischmann, Jérôme Gouzy, Sam Griffiths-Jones, Daniel H. Haft, Henning Hermjakob, Nicolas Hulo, Daniel Kahn, Alexander Kanapin, Maria Krestyaninova, Rodrigo Lopez, Ivica Letunic, Sue Orchard, Marco Pagni, David Peyruc, Chris P. Ponting, Florence Servant, Christian J. A. Sigrist |
Briefings Bioinform. | 5 |
| 2002 | The use of structure information to increase alignment accuracy does not aid homologue detection with profile HMMsabstractMOTIVATION: The best quality multiple sequence alignments are generally considered to derive from structural superposition. However, no previous work has studied the relative performance of profile hidden Markov models (HMMs) derived from such alignments. Therefore several alignment methods have been used to generate multiple sequence alignments from 348 structurally aligned families in the HOMSTRAD database. The performance of profile HMMs derived from the structural and sequence-based alignments has been assessed for homologue detection. RESULTS: The best alignment methods studied here correctly align nearly 80% of residues with respect to structure alignments. Alignment quality and model sensitivity are found to be dependent on average number, length, and identity of sequences in the alignment. The striking conclusion is that, although structural data may improve the quality of multiple sequence alignments, this does not add to the ability of the derived profile HMMs to find sequence homologues. SUPPLEMENTARY INFORMATION: A list of HOMSTRAD families used in this study and the corresponding Pfam families is available at http://www.sanger.ac.uk/Users/sgj/alignments/map.html CONTACT: [email protected] Sam Griffiths-Jones, Alex Bateman |
Bioinform. | 2 |
| 2002 | QuickTree: building huge Neighbour-Joining trees of protein sequencesabstractAbstract Summary: We have written a fast implementation of the popular Neighbor-Joining tree building algorithm. QuickTree allows the reconstruction of phylogenies for very large protein families (including the largest Pfam alignment containing 27000 HIV GP120 glycoprotein sequences) that would be infeasible using other popular methods. Availability: The source-code for QuickTree, written in ANSI~C, is freely available via the world wide web at http://www.sanger.ac.uk/Software/analysis/quicktree Contact: [email protected] * To whom correspondence should be addressed. Kevin L. Howe, Alex Bateman, Richard Durbin |
Bioinform. | 2 |
| 2002 | The SGS3 protein involved in PTGS finds a familyabstractBACKGROUND: Post transcriptional gene silencing (PTGS) is a recently discovered phenomenon that is an area of intense research interest. Components of the PTGS machinery are being discovered by genetic and bioinformatics approaches, but the picture is not yet complete. RESULTS: The gene for the PTGS impaired Arabidopsis mutant sgs3 was recently cloned and was not found to have similarity to any other known protein. By a detailed analysis of the sequence of SGS3 we have defined three new protein domains: the XH domain, the XS domain and the zf-XS domain, that are shared with a large family of uncharacterised plant proteins. This work implicates these plant proteins in PTGS. CONCLUSION: The enigmatic SGS3 protein has been found to contain two predicted domains in common with a family of plant proteins. The other members of this family have been predicted to be transcription factors, however this function seems unlikely based on this analysis. A bioinformatics approach has implicated a new family of plant proteins related to SGS3 as potential candidates for PTGS related functions. Alex Bateman |
BMC Bioinform. | 1 |
| 2001 | HOMSTRAD: adding sequence information to structure-based alignments of homologous protein familiesabstractsummary: We describe an extension to the Homologous Structure Alignment Database (HOMSTRAD; Mizuguchi et al., Protein Sci., 7, 2469-2471, 1998a) to include homologous sequences derived from the protein families database Pfam (Bateman et al., Nucleic Acids Res., 28, 263-266, 2000). HOMSTRAD is integrated with the server FUGUE (Shi et al., submitted, 2001) for recognition and alignment of homologues, benefitting from the combination of abundant sequence information and accurate structure-based alignments. AVAILABILITY The HOMSTRAD database is available at: http://www-cryst.bioc.cam.ac.uk/homstrad/. Query sequences can be submitted to the homology recognition/alignment server FUGUE at: http://www-cryst.bioc.cam.ac.uk/fugue/. Paul I. W. de Bakker, Alex Bateman, David F. Burke, Ricardo Núñez Miguel, Kenji Mizuguchi, Jiye Shi, Hiroki Shirai, Tom L. Blundell |
Bioinform. | 2 |
| 2000 | InterPro-an integrated documentation resource for protein families, domains and functional sitesabstractMOTIVATION: InterPro is a new integrated documentation resource for protein families, domains and functional sites, developed initially as a means of rationalising the complementary efforts of the PROSITE, PRINTS, Pfam and ProDom database projects. RESULTS: Merged annotations from PRINTS, PROSITE and Pfam form the InterPro core. Each combined InterPro entry includes functional descriptions and literature references, and links are made back to the relevant parent database(s), allowing users to see at a glance whether a particular family or domain has associated patterns, profiles, fingerprints, etc. Merged and individual entries (i.e. those that have no counterpart in the companion resources) are assigned unique accession numbers. Release 1.2 of InterPro (June 2000) contains over 3000 entries, representing families, domains, repeats and sites of post-translational modification (PTMs) encoded by 6581 different regular expressions, profiles, fingerprints and Hidden Markov Models (HMMs). Each InterPro entry lists all the matches against SWISS-PROT and TrEMBL (more than 1000000 hits from 264333 different proteins out of 384572 in SWISS-PROT and TrEMBL). Rolf Apweiler, Terri K. Attwood, Amos Bairoch, Alex Bateman, Ewan Birney, Margaret Biswas, Philipp Bucher, Lorenzo Cerutti, Florence Corpet, Michael D. R. Croning, Richard Durbin, Laurent Falquet, Wolfgang Fleischmann, Jérôme Gouzy, Henning Hermjakob, Nicolas Hulo, Inge Jonassen, Daniel Kahn, Alexander Kanapin, Youla Karavidopoulou, Rodrigo Lopez, Beate Marx, Nicola J. Mulder, Thomas M. Oinn, Marco Pagni, Florence Servant, Christian J. A. Sigrist, Evgeny M. Zdobnov |
Bioinform. | 4 |