Alfonso Valencia

dblp:49/1173 · DBLP profile ↗
← Back
92ranked-venue papers
7as first author
6since 2021 · last 2026
0000-0002-8937-6789ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 90 · 7 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 myAURA : a personalized health library for epilepsy management via knowledge graph sparsification and visualization
abstract
OBJECTIVES: Report the development of the patient-centered myAURA application and suite of methods designed to aid epilepsy patients, caregivers, and clinicians in making decisions about self-management and care. MATERIALS AND METHODS: myAURA rests on an unprecedented collection of epilepsy-relevant heterogeneous data resources, such as biomedical databases, social media, and electronic health records (EHRs). We use a patient-centered biomedical dictionary to link the collected data in a multilayer knowledge graph (KG) computed with a generalizable, open-source methodology. RESULTS: Our approach is based on a novel network sparsification method that uses the metric backbone of weighted graphs to discover important edges for inference, recommendation, and visualization. We demonstrate by studying drug-drug interaction from EHRs, extracting epilepsy-focused digital cohorts from social media, and generating a multilayer KG visualization. We also present our patient-centered design and pilot-testing of myAURA, including its user interface. DISCUSSION: The ability to search and explore myAURA's heterogeneous data sources in a single, sparsified, multilayer KG is highly useful for a range of epilepsy studies and stakeholder support. CONCLUSION: Our stakeholder-driven, scalable approach to integrating traditional and nontraditional data sources enables both clinical discovery and data-powered patient self-management in epilepsy and can be generalized to other chronic conditions.
Rion Brattig Correia, Jordan C. Rozum, Leonard E. Cross, Jack Felag, Michael Gallant, Bruce W. Herr, Aehong Min, Jon Sanchez-Valle, Deborah Stungis Rocha, Alfonso Valencia, Katy Börner, Wendy Miller, Luis M. Rocha
J. Am. Medical Informatics Assoc.11
2024 3Dmapper: a command line tool for BioBank-scale mapping of variants to protein structures
abstract
MOTIVATION: The interpretation of genomic data is crucial to understand the molecular mechanisms of biological processes. Protein structures play a vital role in facilitating this interpretation by providing functional context to genetic coding variants. However, mapping genes to proteins is a tedious and error-prone task due to inconsistencies in data formats. Over the past two decades, numerous tools and databases have been developed to automatically map annotated positions and variants to protein structures. However, most of these tools are web-based and not well-suited for large-scale genomic data analysis. RESULTS: To address this issue, we introduce 3Dmapper, a stand-alone command-line tool developed in Python and R. It systematically maps annotated protein positions and variants to protein structures, providing a solution that is both efficient and reliable. AVAILABILITY AND IMPLEMENTATION: https://github.com/vicruiser/3Dmapper.
Victoria Ruiz-Serra, Samuel Valentini, Sergi Madroñero, Alfonso Valencia, Eduard Porta-Pardo
Bioinform.4
2022 ECCB2022: the 21st European Conference on Computational Biology
abstract
This volume of Bioinformatics includes the proceedings papers of the 21st European Conference in Computational Biology (ECCB), an annual international conference for research in computational biology and bioinformatics. The conference is being held jointly with the Intelligent Systems for Molecular Biology (ISMB) Conference in the odd-numbered years and independently in the even-numbered years. This year, the 21st ECCB (ECCB2022) took place in a hybrid format in Sitges (Barcelona, Spain) from September 12 to September 21, 2022, under the motto of Planetary Health and Biodiversity. The positive experience of running ECCB2020 in a virtual format due to the COVID-19 pandemic and the still existing health concerns and travel restrictions have made ECCB2022 the first edition of the conference held in a hybrid format. Information on ECCB2022 and the conference’s satellite meetings can be found at http://eccb2022.org. With more than 850 in-person participants from academia and industry, ECCB is the leading European Conference on Computational Biology and Bioinformatics and the second largest internationally, next to ISMB. The work presented in ECCB is related to all domains of the field of computational biology and bioinformatics, ranging from molecular to systems biology level. The proceedings papers present new computational methodologies and tools for addressing challenging problems in the field, with a strong presence of predictive modelling approaches and multi-omics data integration (Fig. 1). Recovering the traditional in-person format allowed us to open the calls for Highlight Papers, Applications and Posters that were not included in the ECCB2020 programme. Word cloud representing the most frequent keywords accompanying the submissions received. The keywords ‘machine learning’, ‘deep learning’, ‘cancer’, ‘gene expression’ and ‘data integration’ were the most frequently used keywords ECCB2022 leveraged experiences from previous editions, including the last two ones, held online despite being initially programmed to be hosted in-person in Lyon, France (Dessimoz and Przytycka, 2021, ISMB/ECCB 2021) and Sitges, Barcelona, Spain (Capella-Gutierrez et al. 2020, ECCB2020). This was the second time that ECCB was hosted in Spain after the successful experience of the fourth edition back in 2005 in Madrid, Spain (Guigo et al., 2005, ECCB 2005). ECCB2022 was held under the auspices of the Spanish National Bioinformatics Institute (INB; http://inb-elixir.es), founded in 2003, which has been the node of ELIXIR in Spain since 2015. The Life Sciences Department of the Barcelona Supercomputing Center (BSC) was the main organiser of 2022’s conference. BSC is the national supercomputing facility in Spain with more than 800 R&D experts and professionals distributed across four domains: Engineering, Computer, Earth and Life Sciences. The Life Sciences Department includes more than 160 professionals, ranging from MSc and PhD students to postdoctoral researchers and research engineers. The department’s mission is to understand living organisms through theoretical and computational methods, from genomics, structural bioinformatics, molecular and cellular modelling and biomedical research infrastructures, all of them in the framework of European projects and infrastructures (ELIXIR in particular). ECCB2022 received manuscripts for Proceedings talks from institutions across 39 countries. The submissions were organised in five themes: (1) Data (organisation, management, categorization, integration, analysis of data, knowledge discovery), (2) Genes (expression, function, regulation, transcription, translation, geno-phenotype), (3) Genomes (sequence analysis, alignment, evolution, phylogeny, genetics, epigenetics, 3D conformation) (4) Proteins (structure, function, alterations, assemblies, interactions, design, proteomics) and (5) Systems (systems biology, pathways, molecular networks, dynamics, signalling, multi-scale modelling). All proceedings submissions were subjected to a rigorous peer-review process (two to four reviews per paper), organised by the members of the ECCB2022 Programme Committee of the corresponding theme. The Programme Committee, chaired by Ana Conesa and the respective area chair: Josep Lluís Gelpí, Artemis Hatzigeorgiou, Toni Gabaldón, Mark Wass and Patrick Aloy and one to three co-chairs per theme (Heidi Peterson, Rory Johnson, Irene Papatheodorou, Sonia Tarazona Campos, Stephane Rombauts, Shilpa Garg, Franca Fraternali, Jolanda van Leeuwen and Anaïs Baudot) and a Programme Committee with 185 reviewers (see the Supplementary File ‘ECCB2022 Committees’). The review process was handled by the EasyChair system (www.easychair.org), focusing on the impact and reproducibility of the submitted research, as well as its suitability for the ECCB2022 audience. Once the review was completed, the committee chairs and co-chairs selected 25 papers to be included in the ECCB2022 proceedings with an acceptance rate of 17.4% across the different areas. All Proceedings papers and supplementary files are freely available in the electronic form of Oxford University Press journal Bioinformatics as a special issue of September 2022 (https://academic.oup.com/bioinformatics/issue/38/Supplement_2). The ECCB2022 call for Highlight Papers invited in-person presentations of recently published works, i.e. up to 1 year since its publication. The track received a total of 79 submissions that were ranked according to relevance and impact on biology and/or medicine, and their suitability to be presented in front of a large and heterogeneous audience. The existence of complementary work additional to the published paper was also considered positively. Finally, 30% of the submissions were presented at the conference. The conference programme included an Applications track in which participants were invited to present their works on software, services and infrastructures developed in a broader context of academia and industry. The proposals were evaluated by a committee of six members chaired by Javier De Las Rivas and co-chaired by Sara Aibar. A total of 15 out of 29 submissions were selected to be presented at the conference after ranking them based on their impact and potential in bridging different researcher and professional profiles together. ECCB2022 hosted two poster sessions where researchers introduced and discussed their work in person. The Posters track received 567 contributions from authors affiliated with 54 different countries. The submissions were reviewed by a committee of 28 members, chaired by R. Gonzalo Parra and co-chaired by Handan Melike Dönertaş and Alexander Miguel Monzon. Overall the ECCB2022 tracks received the participation of authors from 66 countries. The countries with the higher number of accepted submissions in ECCB2022 tracks (Fig. 2) were Germany, Spain, the UK and France, jointly contributing to half of the overall accepted submissions. Authors from Belgium, Switzerland, Italy and Poland followed the ranking adding up a comparable number of accepted proposals (∼25–30 submissions) and the USA and India were the non-European countries with the higher number of authors with accepted submissions. Proportion of countries’ affiliation from the authors with accepted submissions at ECCB2022. The Programme Committee accepted a total number of 626 submissions, most of which were authored by researchers affiliated with institutions from Germany, Spain, the UK and France The programme incorporated a new track aligned with the motto of the conference: Climate Crisis and Health. The track consisted of two sessions with talks from invited speakers on air quality, heat stress, infectious diseases and biodiversity. Taking advantage of the hybrid format, ECCB2022 hosted a dedicated track to foster the interactions between different geographically distributed communities. During this edition, two sessions were dedicated to fostering interactions between ELIXIR, the pan-European infrastructure for data in Life Sciences and Latin America Bioinformatics societies. The main topic revolved around research data management and how to facilitate the knowledge exchange between communities. These joint sessions just represented the starting point for further collaborations. The programme also included four ELIXIR sessions in which the speakers showcased the latest outputs and services on data integration and data management strategies developed in this European research infrastructure context. Additionally, the Quest for Orthologs consortium organised a satellite meeting on the scope of the conference, which included a day and a half meeting and an open workshop to discuss new approaches for improving orthology predictions. Six distinguished keynote speakers presented their work: Prof. Cesar Hidalgo from the Universities of Toulouse, Manchester and Harvard, Prof. Raul Rabadan from Columbia University, Dr Ana Freitas from the Institute for Systems and Computer Engineering, Technology and Science—Technical University of Lisbon, Dr Maria Rodriguez-Martínez from IBM Research Europe, Dr Graciela Gonzalez-Hernandez from the University of Pennsylvania and Prof. Mar Albà from the Catalan Institution for Research and Advanced Studies and Hospital del Mar Medical Research Institute. The keynotes covered broad and diverse research areas, including evolutionary genomics, personalized medicine and new methods in artificial intelligence and natural language processing for computational modelling of biological data. Following the pilot experience from ECCB2020, this edition also had a programme of virtual workshops and tutorials under the umbrella of New Trends in Bioinformatics by ECCB. This format allowed the spread of 3-h virtual sessions during the week before ECCB2022, enabling the participation of attendees in multiple sessions and making it easier for the global community to join them. For the 2022 edition, the New Trends in Bioinformatics by ECCB combined a total of 9 virtual sessions and 10 in-person sessions. The virtual events were programmed in two blocks, the first from 13.30 to 16.30 and the second from 17.00 to 20.00 following the Central European Summer Time. All sessions were recorded so participants could visit them again through the ECCB2022 virtual platform and, starting on January 1, 2023, through the ISCB.tv channel. Among the 19 scheduled events were 11 workshops, 7 tutorials and 1 session led by ELIXIR. The workshops aimed to provide participants with the opportunity to discuss different perspectives on the cutting edge of a selected research area through presenting technical issues, exchanging research ideas and sharing practical experiences. The 11 workshops were selected out of a total of 17 applications. In contrast, the purpose of the tutorials program is to provide participants with specialized lectures and hands-on training on the most important and emerging topics in bioinformatics and computational biology research. The tutorials offered topics ranging from early and basic steps of computational analysis, e.g. machine learning, cellular processes modelling or data management, to advanced computational skills in important established topics, e.g. functional analyses of single-cell transcriptomics data. A total of 7 tutorials were selected as part of the New Trends in Bioinformatics by ECCB out of 10 applications. The programme of workshops and tutorials was round out by a workshop organised by ELIXIR on practical approaches for applying FAIR guiding principles in data reuse. The specific sessions for the New Trends in Bioinformatics by ECCB were as follows. Tutorials are denoted by a T, Workshops are denoted by a W and ELIXIR workshop contains an E, to form the events ID. NTB-T01 Computational challenges in phospho-proteomics and systems biology of cellular signalling, organised by Filipa Blasco Tavares Pereira-Lopes (Case Western Reserve University, USA), Marzieh Ayati (University of Texas Rio Grande Valley, USA), Serhan Yilmaz (Case Western Reserve University, USA), Daniela Schlatzer (Case Western Reserve University, USA), Mehmet Koyutürk (Case Western Reserve University, USA) and Mark Chance (Case Western Reserve University, USA). NTB-T02 To rarefy or not to rarefy microbiome data? What are the alpha diversity metrics?, organised by Violeta Larios-Serrato (Winter Genomics, México), Maira Nayeli Luis-Vargas (Winter Genomics, México), Karla Ruiz (Winter Genomics, México) and Kenya Contreras (Winter Genomics, México). NTB-T03 Deep learning for biological sequence data: from convolutional neural networks to transformers, organised by Panagiotis Alexiou (Masaryk University, Czech Republic), Petr Simecek (Masaryk University, Czech Republic), David Cechak (Masaryk University, Czech Republic) and Vlastimil Martinek (Masaryk University, Czech Republic). NTB-T04 Functional analysis of single-cell transcriptomics, organised by Pau Badia i Mompel (Heidelberg University, Germany), Robin Browaeys (VIB Center for Inflammation Research, Belgium) and Daniel Dimitrov (Heidelberg University, Germany). NTB-T05 Guidelines for the assessment and analysis of lrRNA-seq data for transcript identification and quantification (LRGASP challenge), organised by Ana Conesa (Institute for Integrative Systems Biology, Spain), Fairlie Reese (University of California at Irvine, USA), Dennis Mulligan (University of California at Santa Cruz, USA), Ying Chen (Genome Institute of Singapore, Singapore), Ralf Herwig (Max Planck Institute for Molecular Genetics, Germany) and Sílvia Carbonell-Sala (Center for Genomic Regulation, Spain). NTB-T06 Boost your data management planning, organised by Helena Schnitzer (ELIXIR Germany and Forschungszentrum Jülich, Germany) and Daniel Wibberg (ELIXIR Germany and Forschungszentrum Jülich, Germany). NTB-T07 Software containerization in bioinformatics: how to make reproducible, portable and reusable bioinformatics software and pipelines, organised by Giacomo Baruzzo (University of Padova, Italy), Barbara Di Camillo (University of Padova, Italy), Marco Cappellato (University of Padova, Italy), Giulia Cesaro (University of Padova, Italy), Mikele Milia (University of Padova, Italy). NTB-W01 Machine learning good practices—DOME recommendations for better machine learning in computational biology, organised by Jennifer Harrow (ELIXIR, UK), Fotis E. Psomopoulos (CERTH, Greece), Silvio Tosatto (University of Padova, Italy) and Leyla Jael García-Castro (ZB MED Information Centre for Life Sciences, Germany). NTB-W02 FAIRification of multi-omics metadata, organised by Gary Saunders (EATRIS-ERIC Data Director, The Netherlands), Emanuela Oldoni (EATRIS-ERIC Scientific Programme Manager, The Netherlands) and Anna Niehues (Bioinformatician at Radboud University Medical Center, The Netherlands). NTB-W03 Simulating cellular behaviours: advancing HPC-enabled computational biology, organised by Arnau Montagud (BSC, Spain), Marta Lloret-Llinares (EMBL-EBI, UK), Renata Giménez (BSC, Spain) and Mariola Tàrrega-Moltó (BSC, Spain). NTB-W04 Spatial transcriptomics and cell–cell communication modelling: new opportunities to study the cellular dynamics of biological systems, organised by Giacomo Baruzzo (University of Padova, Padova, Italy), Enrica Calura (University of Padova, Italy), Davide Risso (University of Padova, Italy), Chiara Romualdi (University of Padova, Italy), Gabriele Sales (University of Padova, Italy), Luz García-Alonso (Wellcome Sanger Institute, UK), Roser Vento-Tormo (Wellcome Sanger Institute, UK), Julio Saez-Rodriguez (Heidelberg University, Germany) and Yvan Saeys (VIB, Belgium). NTB-W05 Building high-quality reference genome assemblies of eukaryotes, organised by Nadège Guiglielmoni (University of Cologne, Germany), Joanna Malukiewicz (Deutsches Primatenzentrum, Germany; University of Sao Paulo, Brazil), Lino Ometto (University of Pavia, Italy) and Robert (University of Lausanne and Swiss Institute of Bioinformatics, Switzerland). NTB-W06 Tools and techniques to make sensitive data discoverable (use-cases, hands-on session of Beacon implementation), organised by Babita Singh (European Genome-phenome Archive, Spain) and Lauren Fromont (European Genome-phenome Archive, Spain). NTB-W07 Sex and gender dimension in biomedical research, organised by Àtia Cortés (BSC, Spain), Davide Cirillo (BSC, Spain) and Fatemeh Baghdadi (BSC, Spain). NTB-W08 Integration of large-scale data for reference genome development in biodiversity, organised by Shilpa Garg (University of Copenhagen, Denmark) and Josiah Kuja (University of Copenhagen, Denmark). NTB-W09 Annual European Bioinformatics Core Community (AEBC2) workshop 2022, organised by Camille Stephan-Otto Attolini (IRB Barcelona, Spain), Dieter Beule (Berlin Institute of Health, Germany), Sven Rahmann (Saarland University, Germany), Sven Nahnsen (University of Tübingen, Germany) and Daniel J. Stekhoven (ETH Zurich, Germany). NTB-W10 Computational modelling of immunological mechanisms: from statistical approaches to interpretable machine learning, organised by María Rodríguez-Martínez (IBM—Zurich Research Laboratory, Switzerland), Anna Niarakis (University of Évry Val d'Essonne and University of Paris-Saclay) and Matteo Barberis (University of Surrey, UK). NTB-W11 Novel challenges in the quest for orthologs, organised by Ingo Ebersberger (Goethe University Frankfurt, Germany), Michael Hiller (Senckenberg Society for Nature Research, Germany), Thomas Rattei (University of Vienna, Austria), Paul D. Thomas (University of Southern California, USA) and Sofia Kirke Forslund (Max-Delbrück-Centre for Molecular Medicine, Germany). NTB-EW01 ELIXIR | FAIR applied: a practical FAIRification guide for life science data from FAIRplus, organised by Tony Burdett (EMBL-EBI, UK), Ibrahim Emam (Imperial College London, UK), Nick Juty (University of Manchester, UK), (University of UK), (University of and (EMBL-EBI, UK). The New Trends in Bioinformatics by ECCB in format in with the European The was organised by the European of the for and by researchers to a to exchange research ideas and The was co-chaired by from Spain, and Maria from are The work and of the and participants of the have led to a and successful the last 10 making the an part of ECCB. Following previous of a number of by were to students and postdoctoral with to members affiliated in or countries to their at the conference and the New Trends in Bioinformatics by ECCB. A total of were received from from countries. a were by the ECCB2022 Committee and The ECCB2022 Committee is to gender for better and in this was considered in all the steps of the the of review out of 10 chairs and 9 out of 15 co-chairs were the and the made an important to the review by as as and a for the keynotes represented four out of the six distinguished The programme included jointly organised with the to further to in the field of bioinformatics. the of the were after in the field of bioinformatics and a was to to further to their contributions in the ECCB2020 established a of which was leveraged and by ECCB2022. The ECCB2022 of a and to and had to to in had been a of the of The ECCB2022 Committee to through their work to the conference’s The co-chairs and reviewers have been the of the conference and workshops, manuscripts and talks for and making this conference a are to them. ECCB2022 are to the ECCB2022 Committee and the ECCB2022 Committee for their and during the of the of ECCB2022. are also to the to for their and for contributing to the of ECCB2022 at the international and for This conference not be the of and ELIXIR the conference as a was the and the poster also further Sciences and The and Spanish Supercomputing were of ECCB2022. to all for their and make ECCB2022 an and conference. A number of to the of ECCB2022. their call of to make this conference and as part of Conference were for the members of the and the Life Sciences Department at BSC this conference: and the the Finally, the of a conference is its the conference for its ECCB2022 all of for us at the first hybrid edition of ECCB. the and the of previous of ECCB with the of ECCB2022 to the next conference in Supplementary Supplementary data are available at Bioinformatics This paper was published as part of a special issue by ECCB2022. of
Salvador Capella-Gutiérrez, Eva Alloza, Laura Rubinat-Ripoll, Ana Conesa, Alfonso Valencia
Bioinform.5
2022 Detection of oncogenic and clinically actionable mutations in cancer genomes critically depends on variant calling tools
abstract
MOTIVATION: The analysis of cancer genomes provides fundamental information about its etiology, the processes driving cell transformation or potential treatments. While researchers and clinicians are often only interested in the identification of oncogenic mutations, actionable variants or mutational signatures, the first crucial step in the analysis of any tumor genome is the identification of somatic variants in cancer cells (i.e. those that have been acquired during their evolution). For that purpose, a wide range of computational tools have been developed in recent years to detect somatic mutations in sequencing data from tumor samples. While there have been some efforts to benchmark somatic variant calling tools and strategies, the extent to which variant calling decisions impact the results of downstream analyses of tumor genomes remains unknown. RESULTS: Here, we quantify the impact of variant calling decisions by comparing the results obtained in three important analyses of cancer genomics data (identification of cancer driver genes, quantification of mutational signatures and detection of clinically actionable variants) when changing the somatic variant caller (MuSE, MuTect2, SomaticSniper and VarScan2) or the strategy to combine them (Consensus of two, Consensus of three and Union) across all 33 cancer types from The Cancer Genome Atlas. Our results show that variant calling decisions have a significant impact on these analyses, creating important differences that could even impact treatment decisions for some patients. Moreover, the Consensus of three calling strategy to combine the output of multiple variant calling tools, a very widely used strategy by the research community, can lead to the loss of some cancer driver genes and actionable mutations. Overall, our results highlight the limitations of widespread practices within the cancer genomics community and point to important differences in critical analyses of tumor sequencing data depending on variant calling, affecting even the identification of clinically actionable variants. AVAILABILITY AND IMPLEMENTATION: Code is available at https://github.com/carlosgarciaprieto/VariantCallingClinicalBenchmark. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Carlos A. Garcia-Prieto, Francisco Martínez-Jiménez, Alfonso Valencia, Eduard Porta-Pardo
Bioinform.3
2022 Parallel model exploration for tumor treatment simulations
abstract
Abstract Computational systems and methods are often being used in biological research, including the understanding of cancer and the development of treatments. Simulations of tumor growth and its response to different drugs are of particular importance, but also challenging complexity. The main challenges are first to calibrate the simulators so as to reproduce real‐world cases, and second, to search for specific values of the parameter space concerning effective drug treatments. In this work, we combine a multi‐scale simulator for tumor cell growth and a genetic algorithm (GA) as a heuristic search method for finding good parameter configurations in reasonable time. The two modules are integrated into a single workflow that can be executed in parallel on high performance computing infrastructures. In effect, the GA is used to calibrate the simulator, and then to explore different drug delivery schemes. Among these schemes, we aim to find those that minimize tumor cell size and the probability of emergence of drug resistant cells in the future. Experimental results illustrate the effectiveness and computational efficiency of the approach.
Charilaos Akasiadis, Miguel Ponce de Leon, Arnau Montagud, Evangelos Michelioudakis, Alexia Atsidakou, Elias Alevizos, Alexander Artikis, Alfonso Valencia, Georgios Paliouras
Comput. Intell.8
2022 The structural coverage of the human proteome before and after AlphaFold
abstract
The protein structure field is experiencing a revolution. From the increased throughput of techniques to determine experimental structures, to developments such as cryo-EM that allow us to find the structures of large protein complexes or, more recently, the development of artificial intelligence tools, such as AlphaFold, that can predict with high accuracy the folding of proteins for which the availability of homology templates is limited. Here we quantify the effect of the recently released AlphaFold database of protein structural models in our knowledge on human proteins. Our results indicate that our current baseline for structural coverage of 48%, considering experimentally-derived or template-based homology models, elevates up to 76% when including AlphaFold predictions. At the same time the fraction of dark proteome is reduced from 26% to just 10% when AlphaFold models are considered. Furthermore, although the coverage of disease-associated genes and mutations was near complete before AlphaFold release (69% of Clinvar pathogenic mutations and 88% of oncogenic mutations), AlphaFold models still provide an additional coverage of 3% to 13% of these critically important sets of biomedical genes and mutations. Finally, we show how the contribution of AlphaFold models to the structural coverage of non-human organisms, including important pathogenic bacteria, is significantly larger than that of the human proteome. Overall, our results show that the sequence-structure gap of human proteins has almost disappeared, an outstanding success of direct consequences for the knowledge on the human genome and the derived medical applications.
Eduard Porta-Pardo, Victoria Ruiz-Serra, Samuel Valentini, Alfonso Valencia
PLoS Comput. Biol.4
2020 TiFoSi: an efficient tool for mechanobiology simulations of epithelia
abstract
MOTIVATION: Emerging phenomena in developmental biology and tissue engineering are the result of feedbacks between gene expression and cell biomechanics. In that context, in silico experiments are a powerful tool to understand fundamental mechanisms and to formulate and test hypotheses. RESULTS: Here, we present TiFoSi, a computational tool to simulate the cellular dynamics of planar epithelia. TiFoSi allows to model feedbacks between cellular mechanics and gene expression (either in a deterministic or a stochastic way), the interaction between different cell populations, the custom design of the cell cycle and cleavage properties, the protein number partitioning upon cell division, and the modeling of cell communication (juxtacrine and paracrine signaling). AVAILABILITY AND IMPLEMENTATION: http://tifosi.thesimbiosys.com. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Oriol Canela-Xandri, Samira Anbari, Javier Buceta, Alfonso Valencia
Bioinform.4
2020 ECCB2020: the 19th European Conference on Computational Biology
abstract
This volume of Bioinformatics includes the proceedings papers of the 19th European Conference in Computational Biology (ECCB), an annual international conference for research in computational biology and bioinformatics. The conference is being held jointly with the Intelligent Systems for Molecular Biology (ISMB) Conference in the odd-numbered years and independently in the even-numbered years. This year, the 19th ECCB (ECCB2020) took place virtually from 31 August to 8 September, 2020. Health concerns and travel restrictions due to the COVID-19 pandemics have made ECCB2020 the first edition of the conference held virtually. Indeed, ECCB2022 is planned to be held in-person in the same localization (Sitges, Barcelona, Spain) as ECCB2020. Information on the ECCB2020 main conference as well as satellite meetings can be found at eccb20.org and will be later archived at the EBI website. With more than 850 participants from academia and industry in this virtual format, ECCB is the leading European Conference on Computational Biology and Bioinformatics, and the second largest internationally, next to ISMB. The work presented in ECCB is related to all domains of the field of computational biology and bioinformatics, ranging from molecular to systems biology level. The proceedings papers present new computational methodologies and tools for addressing challenging problems in the field, including research in single cell dynamics, 3D organization of the genome, new cancer therapies using big-data approaches and others. Adjusting the conference format to a virtual one led to producing a compact but yet attractive program. Due to the reduced format, the traditional Highlights and Applications tracks were not carried out. Previous ECCB editions were organized in Basel, Switzerland (Bromberg et al., 2019); Athens, Greece (Hatzigeorgiou et al., 2018); Prague, Czech Republic (Beerenwinkel and Bromberg, 2017); The Hague, The Netherlands (Heringa and Reinders, 2016); Dublin, Ireland (Moreau and Beerenwinkel, 2015); Strasbourg, France (Devignes and Moreau, 2014); Berlin, Germany (Ben-Tal, 2013); Basel, Switzerland (Schwede and Iber, 2012); Vienna, Austria (Gaasterland and Vingron, 2011); Ghent, Belgium (Moreau and Heringa, 2010); Stockholm, Sweden (Gusfield and Tramontano, 2009); Cagliari, Italy (Tramontano, 2008); Vienna, Austria (Lengauer et al., 2007); Eilat, Israel (Wolfson and Safer, 2007); Madrid, Spain (Guigo et al., 2005); Glasgow, UK (Thornton et al., 2004); Paris, France (Lenhof and Sagot, 2003); and Saarbrücken, Germany (Lengauer, 2002). ECCB2020 continues the 19-year tradition and was held under the auspices of the Spanish National Bioinformatics Institute (http://inb-elixir.es), founded in 2003, which is the node of ELIXIR in Spain since 2015, and the bioinformatics technology platform of the Carlos III Health Institute (‘Instituto de Salud Carlos III’) since 2018. The Life Sciences department at the Barcelona Supercomputing Center (BSC) has been the main organizer of this year’s conference. BSC is the national supercomputing center in Spain with more than 700 R&D experts and professionals distributed across four fields: Computer Sciences, Life Sciences, Earth Sciences and Computer Applications in Science and Engineering. The Life Sciences department comprises more than 120 professionals including MSc and PhD students, post-doctoral researchers and research engineers. The department’s mission is to understand living organisms by means of theoretical and computational methods, e.g. molecular modeling, genomics, proteomics, as well as contributing to the development of research infrastructures in the frame of different European projects, organizations and international efforts. ECCB2020 received a record number of 203 applications for proceedings talks from institutions across 33 countries. The submissions were organized in five themes, according to their topic: (i) data (organization, integration, knowledge discovery and multi-scale modeling), (ii) genes (expression, function, editing and geno/phenotype), (iii) Genome (sequence analysis, evolution, phylogeny and microbiome), (iv) proteins (structure, function, alterations and drug design) and (v) systems (molecular pathways, signaling and metabolomics). All proceedings submissions were subjected to a rigorous peer-review process (2–4 reviews per paper), organized by the members of the ECCB2020 Programme Committees of the corresponding theme. The Programme Committees consisted of a Theme Chair (Josep Lluís Gelpí, Ana Conesa, Toni Gabaldón, Modesto Orozco, Patrick Aloy, respectively) and up to three co-chairs per theme (Artemis Hatzigeorgiou, Jacques van Helden, Mark Robinson, Stephane Rombauts, Mark Wass, Amelie Stein, Pedro Beltrao). For the review, the EasyChair system (www.easychair.org) was used by 208 reviewers, recruited by the Programme Committees. Reviewers evaluated the impact and reproducibility of the presented research, as well as its suitability for the ECCB2020 audience. Once the review was completed, the committee chairs and co-chairs selected 41 papers to be included in the ECCB2020 proceedings with an acceptance rate of 20.2%, ensuring proportion to the initial theme submissions. The accepted works required, in some cases, minor revisions and the authors had 2 weeks to modify them accordingly. All Proceedings Track papers and any supplementary files accompanying them are freely available in the electronic form of Oxford University Press journal Bioinformatics, as a special issue of December 2020. Due to the virtual format, ECCB2020 did not have a dedicated session to posters where researchers could introduce their works to the whole community in-person. In contrast, the conference proposed a newly established format to get to know research activities from far and from near under the umbrella of the Glimpse into Global Bioinformatics Communities. ECCB2020 held two of these sessions, one dedicated to Latin America and another one to Europe. The aim of the Iberoamerican Society for Bioinformatics session served to bridge the gap with the Latin American communities by highlighting emerging researchers from different countries working on a broad range of topics, including national initiatives on precision medicine, city-scale genomic cartographies for infectious diseases response, molecular modeling and novel transcriptional regulatory interactions. The aim of the ELIXIR session as European representatives was to showcase the latest outputs and services developed in the context of this research infrastructure. Presentations had a strong focus on the response of ELIXIR to the challenges posed by the spread of COVID-19 from different angles including the development of specific ontologies, data management strategies to the release of dedicated structural and functional data for the SARS-CoV-2 proteome. Six distinguished keynote speakers presented their work: Prof. Modesto Orozco from Institute from Research in Biomedicine (IRB Barcelona), Prof. Geneviève Almouzni from Institut Curie, Science Academy in France and LifeTime Initiative, Assistant Prof. Debora Marks from the Harvard Medical School, Prof. Fabian Theis from Institute of Computational Biology at Helmholtz Zentrum München, Prof. Bissan Al-Lazikani from The Institute of Cancer Research and Prof. Londa Schiebinger from Stanford University. Keynotes covered broad and diverse areas of research including single cell genomics modeling, probabilistic generative modeling of genetic variation and gender in research, policy and practice. During ECCB2020, 11 exhibitor and career virtual booths were set up, namely: BSC, INB/ELIXIR-ES, ELIXIR, EMBL-EBI, GOBLET, the International Society for Computational Biology (ISCB), Genes-MDPI, Saint Jude Children’s Research Hospital. Exhibitors showcased the latest trends in computational biology and bioinformatics technology, in training and scientific literature, as well as in high-performance computing. The organizing committee invited participants to visit the exhibition area and the career fair to support these community-minded organizations, who deliver a strong message in supporting the Computational Biology and Bioinformatics scientific field. For the 2020 edition, the organizing committee designed a tailored format for virtual workshops and tutorials under the umbrella of the brand-new New Trends in Bioinformatics by ECCB. New Trends in Bioinformatics by ECCB aims to create virtual spaces for closer interactions among participants in specific areas of Bioinformatics and Computational Biology for this and future editions. Rather than having many parallel 1 or 2 days long face-to-face sessions prior to the main conference, this new format aims to facilitate 3-h long virtual workshops and tutorials during the week prior to ECCB. By programing those sessions considering the different time-zones of participants, it is possible to have a truly global audience taking part in these activities. For the 2020 edition, the New Trends in Bioinformatics by ECCB programed a total of 19 events in 2 parallel tracks and 2 blocks, the first from 13.30 to 16.30 and the second from 17.00 to 20.00, all following the Central European Summer Time. All sessions were recorded so participants can visit them at any time through the ECCB2020 virtual platform and, starting on 1 January, 2021, through the ISCBtv channel. Among the 19 events, there were 5 Workshops including 2 from Special Interest Groups—SIGs, 10 Tutorials and 4 sessions led by ELIXIR. The aim of workshops was to provide participants the opportunity to discuss different perspectives on the cutting edge of a selected research area through presenting technical issues, exchanging research ideas and sharing practical experiences. The five workshops of the 2020 edition were selected out of a total of 14 applications. In contrast, the purpose of the tutorial program is to provide participants with specialized lectures and hands-on training to the most important and emerging topics in bioinformatics and computational biology research. The tutorials offered skills ranging from early and basic steps of computational analysis, e.g. machine learning, cellular processes modeling or data visualization, to advanced computational skills in important established topics, e.g. epigenomics and RNA-seq analysis using long-reads. A total of 10 tutorials were selected as part of the New Trends in Bioinformatics by ECCB2020 out of 19 applications. This initial selection of workshops and tutorials was complemented by four events organized by ELIXIR, the strategic partner for the ECCB series of conferences. Specifically, ELIXIR organized three short workshops to share with the broad community advances in the development of standards and reference implementations for the use and re-use of research data even when data are under access control. Those workshops were complemented by a tutorial on integrating structural and functional data to support in silico predictions for drug design. The specific sessions for the New Trends in Bioinformatics by ECCB were as follows. Tutorials are denoted by a T, Workshops are denoted by a W and ELIXIR events contain an E to form the events ID. NTB-T01 machine learning and omics data: opportunities for advancing biomedical data analysis in Galaxy, organized by Anup Kumar, Alireza Khanteymoori and Björn Grüning (all at de. NBI, ELIXIR-DE, Germany) and Fotis Psomopoulos (INAB | CERTH, ELIXIR-GR, Greece). NTB-T02 powerful presentations tutorial—how to prepare, design and deliver high-impact presentations, organized by the Bioinfo4Women programme (BSC, Spain) and Grifols as strategic partner to the programme. NTB-T03 using deep learning for image and sequence analysis, organized by Petr Simecek and Panagiotis Alexiou (Central European Institute of Technology, Masaryk University, Czech Republic). NTB-T04 keeping up with epigenomic analysis: theory and practice, organized by Marcel Schulz (Institute of Cardiovascular Regeneration, Uniklinikum and Goethe University Frankfurt, Germany) and Ivan G. Costa (Institute for Computational Genomics, RWTH Aachen, Germany). NTB-T05 computational modeling of cellular processes: regulatory versus metabolic systems, organized by Anna Niarakis (Université d'Évry, Université de Paris-Saclay, France), Dagmar Waltemath (University Medicine, Greifswald, Germany),Pedro Monteiro (Universidade de Lisboa, Portugal), Vincent Noel (Institut Curie, France), Marta Cascante (Universitat de Barcelona, Spain) and Miguel Ponce de Leon (BSC, Spain). NTB-T06 reconstruction, analysis and visualization of phylogenomic data with the ETE Toolkit, organized by Jaime Huerta Cepas (Centre for Plant Biotechnology and Genomics, Spain) andFrançois Serra (BSC, Spain). NTB-T07 deep dive into metagenomic data using metagenome-atlas and MMseqs2, organized by Silas Kieser (PHYME, University of Geneva, Switzerland), Milot Mirdita and Johannes Söding (both at the MPI for Biophysical Chemistry, Germany), Martin Steinegger (Seoul National University, Korea) and Sofie Thijs (Hasselt University, Belgium). NTB-T08 full-length RNA-Seq analysis using PacBio long-reads: from reads to functional interpretation, organized by Ana Conesa (University of Florida, United State) and Elizabeth Tseng (Pacific Biosciences, United States). NTB-T09 introduction to structural bioinformatics for evolutionary analysis, organized by Claudia Alvarez Carreño (Georgia Institute of Technology, United States) and Alma Carolina Sanchez Rocha and Vyacheslav Tretyachenko (both at Charles University, Czech Republic). NTB-T10 biomedical data and text processing using shell scripting, organized by Francisco M. Couto (LASIGE, Universidade de Lisboa, Portugal). NTB-W01 CRISPR informatics for functional genomics, cancer targeting, and beyond, organized by Traver Hart (MD Anderson Cancer Center, United States) and Leopold Parts (Wellcome Sanger Institute, UK). NTB-W02 Annual European Bioinformatics Core Community (AEBC2) Workshop 2020, organized by Dieter Beule (Berlin Institute of Health, Germany), Sven Rahmann (University of Duisburg-Essen, Germany), Sven Nahnsen (University of Tübingen, Germany), Daniel J. Stekhoven (ETH Zurich, Switzerland), Camille Stephan Otto Attolini (IRB Barcelona, Spain) and Chris Evelo (Maastricht University, Netherlands). NTB-W03 BioNetVisA: biological network reconstruction, data visualization and analysis in biology and medicine, organized by Emmanuel Barillot, Inna Kuperstein, Cristóbal Monraz Gómez and Andrei Zinovyev (all at Institut Curie, France), Hioraki Kitano (RIKEN Center for Integrative Medical Sciences, Japan), Alfonso Valencia (Spanish National Bioinformatics Institute, Spain), Samik Ghosh (Systems Biology Institute, Japan) andRobin Haw (Ontario Institute for Cancer Research, Canada). NTB-W04 advances in computational modeling of cellular processes and high-performance computing, organized by Anna Niarakis (University Evry, University of Paris-Saclay, France), Arnau Montagud and Miguel Ponce de León (both at the BSC, Spain). NTB-W05 computational pangenomics: algorithms & applications, organized by Solon P. Pissis (Centrum Wiskunde & Informatica, The Netherlands) and Yuri Pirola (Università degli Studi di Milano-Bicocca, Italy). NTB-EW01 ELIXIR | Workshop on FAIR Computational Workflows, organized byIgnacio Eguinoa and Frederik Coppens (both at VIB & ELIXIR-BE, Belgium), Björn Grüning (University of Freiburg & ELIXIR-DE, Germany), Carole Goble and Stian Soiland-Reyes (both at The University of Manchester & ELIXIR-UK, UK) and Salvador Capella-Gutierrez (BSC & ELIXIR-ES, Spain). NTB-EW02 ELIXIR::GA4GH: advancing genomics through expedited data access enabled by standards and ontologies, organized by Gary Saunders (ELIXIR Hub, UK), Melanie Courtot (EMBL-EBI, UK), Michael Baudis (University of Zurich, ELIXIR-CH, Switzerland), Tommi Nyronen (CSC, ELIXIR-FI, UK) and Jordi Rambla (CRG, ELIXIR-ES, Spain). NTB-EW03 ELIXIR | Biological data analysis using InterMine, organized by Rachel Lyne and Daniela Butano (both at the University of Cambridge, UK). NTB-ET01 ELIXIR | 3D-Bioinfo: integrating structural and functional data to support in silico predictions in drug design, organized by Christine Orengo (University College London, UK), Preeti Choudhary (Protein Data Bank in Europe, UK) and Vincent Zoete (University of Lausanne, Switzerland). The New Trends in Bioinformatics by ECCB was followed by the Sixth European Student Council Symposium (ESCS 2020), which preceded the ECCB2020 main conference. This event was organized by the Student Council of the ISCB for and by early-stage researchers aiming to create a space to exchange research ideas and establish future networks. The event was chaired by Gabriel Olguín Orellana, from ISCB Student Council (ISCBSC)-Regional Student Groups (RSG) Chile and by the co-chair Sofia Papadimitriou, from ISCBSC-RSG Belgium. RSG are student groups of the ISCBSC. ESCS 2020 highlights will be published in F1000Research via the ISCBSC channel. The enthusiasm, hard work and genuine scientific interest by the organizers and participants of the ESCS series have led to a continuous and successful collaboration over the last 10 years, which in turn has made the ESCS an inseparable part of ECCB. Following previous editions of ECCB, a number of attendance waiver fellowships, sponsored by ISCB, were given to students and post-doctoral fellows with priority going to members from low or middle-income countries to attend the conference and the New Trends in Bioinformatics by ECCB. A total of 38 applications were received from scientists from 19 countries. After a careful review process, 31 registration fellowships were awarded by the ECCB2020 Organizing Committee together with ISCB. The conference ran under the ECCB2020 Code of to a and to at all and to have to turn to in there had been a to the Code of The ECCB2020 Organizing Committee to who through their work to this virtual Conference a the The and as well all the reviewers, who have been the of the Conference and workshops and this Conference a are truly to ECCB2020 organizers are to the ECCB2020 Committee and the ECCB for their and continuous support during the of the organization of ECCB2020 including the of for the of ECCB2020 into a virtual to ECCB for the support and knowledge on the organization of a conference of are to the ISCB, to for their and support in the of a newly format, in contributing to the of ECCB2020 at the international and for student This Conference not be possible the support of and was the of ECCB2020. ELIXIR the conference as ISCB, Spanish Supercomputing Jude Children’s Research and Bioinfo4Women programme were Genes-MDPI, University Press and Grifols were A to all for their and for to ECCB2020 an and A number of to the organization of ECCB2020. their of for this conference and as part of Conference were for the registration virtual technical and members of the and the Life Sciences at BSC this the of a conference is its presentations, the conference for its ECCB2020 organizers all of for at the first virtual edition of ECCB. the and ECCB2022 will the of previous editions of ECCB together with the of ECCB2020 to a face-to-face conference in the same localization (Sitges, Barcelona, Spain) planned for ECCB2020. The Spanish National Bioinformatics Institute is the Bioinformatics is by of the Salud of the de a de by the de Salud Carlos III and European of
Salvador Capella-Gutiérrez, Eva Alloza, Edurne Gallastegui, Ioannis Kavakiotis, Jennifer L. Harrow, Alberto Langtry, Alfonso Valencia
Bioinform.7
2020 On the inconsistent treatment of gene-protein-reaction rules in context-specific metabolic models
abstract
SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Miguel Ponce de Leon, Iñigo Apaolaza, Alfonso Valencia, Francisco J. Planes
Bioinform.3
2020 Parasitologist-level classification of apicomplexan parasites and host cell with deep cycle transfer learning (DCTL)
abstract
MOTIVATION: Apicomplexan parasites, including Toxoplasma, Plasmodium and Babesia, are important pathogens that affect billions of humans and animals worldwide. Usually a microscope is used to detect these parasites, but it is difficult to use microscopes and clinician requires to be trained. Finding a cost-effective solution to detect these parasites is of particular interest in developing countries, in which infection is more common. RESULTS: Here, we propose an alternative method, deep cycle transfer learning (DCTL), to detect apicomplexan parasites, by utilizing deep learning-based microscopic image analysis. DCTL is based on observations of parasitologists that Toxoplasma is banana-shaped, Plasmodium is generally ring-shaped, and Babesia is typically pear-shaped. Our approach aims to connect those microscopic objects (Toxoplasma, Plasmodium, Babesia and erythrocyte) with their morphological similar macro ones (banana, ring, pear and apple) through a cycle transfer of knowledge. In the experiments, we conduct DCTL on 24 358 microscopic images of parasites. Results demonstrate high accuracy and effectiveness of DCTL, with an average accuracy of 95.7% and an area under the curve of 0.995 for all parasites types. This article is the first work to apply knowledge from parasitologists to apicomplexan parasite recognition, and it opens new ground for developing AI-powered microscopy image diagnostic systems. AVAILABILITY AND IMPLEMENTATION: Code and dataset available at https://github.com/senli2018/DCTL. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Hao Jiang 0028, Jesús A. Cortés-Vecino, Yang Zhang 0057, Alfonso Valencia
Bioinform.6
2020 Chromatin network markers of leukemia
abstract
MOTIVATION: The structure of chromatin impacts gene expression. Its alteration has been shown to coincide with the occurrence of cancer. A key challenge is in understanding the role of chromatin structure (CS) in cellular processes and its implications in diseases. RESULTS: We propose a comparative pipeline to analyze CSs and apply it to study chronic lymphocytic leukemia (CLL). We model the chromatin of the affected and control cells as networks and analyze the network topology by state-of-the-art methods. Our results show that CSs are a rich source of new biological and functional information about DNA elements and cells that can complement protein-protein and co-expression data. Importantly, we show the existence of structural markers of cancer-related DNA elements in the chromatin. Surprisingly, CLL driver genes are characterized by specific local wiring patterns not only in the CS network of CLL cells, but also of healthy cells. This allows us to successfully predict new CLL-related DNA elements. Importantly, this shows that we can identify cancer-related DNA elements in other cancer types by investigating the CS network of the healthy cell of origin, a key new insight paving the road to new therapeutic strategies. This gives us an opportunity to exploit chromosome conformation data in healthy cells to predict new drivers. AVAILABILITY AND IMPLEMENTATION: Our predicted CLL genes and RNAs are provided as a free resource to the community at https://life.bsc.es/iconbi/chromatin/index.html. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Noël Malod-Dognin, Vera Pancaldi, Alfonso Valencia, Natasa Przulj
Bioinform.3
2019 Precision medicine needs pioneering clinical bioinformaticians
abstract
Success in precision medicine depends on accessing high-quality genetic and molecular data from large, well-annotated patient cohorts that couple biological samples to comprehensive clinical data, which in conjunction can lead to effective therapies. From such a scenario emerges the need for a new professional profile, an expert bioinformatician with training in clinical areas who can make sense of multi-omics data to improve therapeutic interventions in patients, and the design of optimized basket trials. In this review, we first describe the main policies and international initiatives that focus on precision medicine. Secondly, we review the currently ongoing clinical trials in precision medicine, introducing the concept of 'precision bioinformatics', and we describe current pioneering bioinformatics efforts aimed at implementing tools and computational infrastructures for precision medicine in health institutions around the world. Thirdly, we discuss the challenges related to the clinical training of bioinformaticians, and the urgent need for computational specialists capable of assimilating medical terminologies and protocols to address real clinical questions. We also propose some skills required to carry out common tasks in clinical bioinformatics and some tips for emergent groups. Finally, we explore the future perspectives and the challenges faced by precision medicine bioinformatics.
Gonzalo Gómez-López, Joaquín Dopazo, Juan C. Cigudosa, Alfonso Valencia, Fátima Al-Shahrour
Briefings Bioinform.4
2019 BIOLITMAP: a web-based geolocated, temporal and thematic visualization of the evolution of bioinformatics publications
abstract
MOTIVATION: The fast growth of bioinformatics adds a significant difficulty to assess the contribution, geographical and thematic distribution of the research publications. RESULTS: To help researchers, grant agencies and general public to assess the progress in bioinformatics, we have developed BIOLITMAP, a web-based geolocation system that allows an easy and sensible exploration of the publications by institution, year and topic. AVAILABILITY AND IMPLEMENTATION: BIOLITMAP is available at http://socialanalytics.bsc.es/biolitmap and the sources have been deposited at https://github.com/inab/BIOLITMAP. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Adrian Bazaga, Alfonso Valencia, María-José Rementeria
Bioinform.2
2019 vulcanSpot: a tool to prioritize therapeutic vulnerabilities in cancer
abstract
MOTIVATION: Genetic alterations lead to tumor progression and cell survival but also uncover cancer-specific vulnerabilities on gene dependencies that can be therapeutically exploited. RESULTS: vulcanSpot is a novel computational approach implemented to expand the therapeutic options in cancer beyond known-driver genes unlocking alternative ways to target undruggable genes. The method integrates genome-wide information provided by massive screening experiments to detect genetic vulnerabilities associated to tumors. Then, vulcanSpot prioritizes drugs to target cancer-specific gene dependencies using a weighted scoring system based on well known drug-gene relationships and drug repositioning strategies. AVAILABILITY AND IMPLEMENTATION: http://www.vulcanspot.org. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Javier Perales-Patón, Tomás Di Domenico, Coral Fustero Torre, Elena Piñeiro-Yáñez, Carlos Carretero-Puche, Héctor Tejero, Alfonso Valencia, Gonzalo Gómez-López, Fátima Al-Shahrour
Bioinform.7
2019 Reviewer-coerced citation: case report, update on journal policy and suggestions for future prevention
abstract
A case was recently brought to the journal’s attention regarding a reviewer who had requested a large number of citations to their own papers as part of their review. After investigation of their most recent reviews, we found that in every review this reviewer requested an average of 35 citations be added, ∼90% of which were to their own papers and the remainder to papers that both cited them extensively and mentioned them by name in the title. The reviewer’s phrasing strongly suggested that inclusion of these citations would influence their recommendation to the editor to accept or reject the paper. The reviewer was unable to provide a satisfactory justification for these requests and Bioinformatics has therefore banned them as a reviewer. Our investigation also suggests that the reviewer has behaved similarly in reviewing for other journals. This case has alerted us to how the peer-review system is vulnerable to unethical behavior, and prompted us to clarify the journal’s policy on when it is appropriate for reviewers to request citations to their own work, and to suggest how some of the current weak points in the peer-review system can be mitigated, so that this behavior can be detected more quickly and efficiently. Peer-reviewers are typically selected based on their expertise in the areas of research associated with newly submitted manuscripts. They are among the most likely to be familiar with prior publications pertinent to the submission, and reviewer feedback on the completeness and accuracy of a manuscript’s reference list is desirable and welcome. It is therefore not unusual that one or a few requested citations may be to the reviewer’s own research. However, reviewers should be aware that a rigorous scientific justification for the inclusion of a new citation must be provided. Since it is easy to provide a tenuous justification for inclusion, as this reviewer often did (e.g. that his papers also involved analysis of sequence data), it should instead be stated why the authors would be remiss or the paper weaker if the citation were not included. Likewise, editors and authors should be aware of the imbalance of power that exists in the review process, and should ensure that any citations added in response to a reviewer comment are relevant and important. Citations have been called the ‘currency’ of science, meaning that they could be considered a quantifiable and objective metric of the impact of a scientist’s research. Specific metrics, such as the H-index, are intended to reflect this and scientists therefore have an incentive to try to improve their H-index. Indeed, the reviewer we caught requesting extensive citation of their work has a webpage that includes prominent mentions of both their high H-index and past awards they received from Thomson Reuters for being a highly cited researcher. When a reviewer agrees to review a paper with the intention of inflating the number of citations to their work, this is a conflict of interest and an unethical manipulation of the peer-review system. One might ask how this reviewer got away with submitting multiple reviews containing coercive requests for citation before being banned. The shortest explanation is that excessive self-citation demands are generally not seen as an ethical problem until a pattern is established, and a decentralized peer-review system is not amenable to detecting patterns. Editors may overlook requests, authors seem reluctant to bring it up explicitly to the editor, reviewer comments are anonymous and scattered across different editors and different journals, and even editors that spot such patterns may not even be aware what options they have, and therefore take the least-energy option of no longer inviting that reviewer. Because accusing someone of unethical conduct is a serious matter, the editors and authors involved are hesitant to do so, particularly if all they have is one instance to base their actions upon. Combined, this creates a system whereby such behavior can persist for a very long time. Even in the rare event that such misbehavior is detected, there is no global solution. Following the guidelines of the Committee on Publication Ethics (COPE), the first step is to contact the reviewer for an explanation. If it is unsatisfactory, then bring the matter to the attention of their immediate supervisor. This second step is only effective in a well-organized academic system, which was not the case here. Despite the actions taken by Bioinformatics, it is likely that this reviewer will continue to review for other journals after this editorial is published. Readers have no doubt noticed that this reviewer is not mentioned by name. This is because we have concluded that we have an obligation to maintain peer-reviewer anonymity, and thus, we can only alert others to the general problem. We debated this last point extensively, and have raised the issue with the COPE as a case report to be discussed. Specific suggestions for reviewers: Motivation. Reviewers should properly motivate their requests for citations and specify how strongly they feel about the addition of references. For example, saying ‘the authors’ review of the field is incomplete, they should add the following references’ is vague, whereas ‘similar studies on the use of X for the purpose of Y were published prior to this one and are needed to alert readers to prior art’ is specific. The less specific a reviewer is regarding motivation, the less weight their request should be given. Moderation. Reviewers should refrain from requesting substantial numbers of references. What is ‘substantial’ will vary by the type of article, with review articles expected to be better in their coverage and short two-page application notes expected to include only the most relevant references. We propose a general rule of thumb to define ‘substantial’ as requesting addition of more than one reference per printed journal page of the paper. In the event the authors’ citations of pertinent prior research is highly incomplete, a reviewer should simply say so and then point them in the right direction with a few citations and let them do the rest. Communication. If a reviewer notices another reviewer has requested excessive or unmerited citations and this has not been commented on by the editor, they should feel free to share their observations and opinions directly with the editor. Reviewers should be cognizant they are also in a position to recognize patterns of abuse from their fellow reviewers. Specific suggestions for journals: Document patterns. Manuscript handling systems should include a checkbox for each reviewer that asks ‘did this reviewer request citation to their own research?’ Editors with a concern about reviewer’s citation requests could then see what percentage of reviews returned contained self-citation requests. Brief but clear guidelines. Instructions should be kept simple and clear. As a result of this case we have updated our reviewer guidelines to state that requests for citations should include ‘a brief, yet specific, rationale as to why their inclusion is merited. This rationale is particularly important if the reviewer requests citation of their own papers.’ Specific suggestions for authors: Voice concerns. Although adding multiple references in response to a reviewer request might seem like an ‘easy’ way to satisfy at least one of the reviewers, each unmerited citation clutters your paper and rewards unethical behavior. Don’t be hesitant to include in your response that you have considered the suggestion and feel they are not necessary. In the event the reviewer responds negatively, you should contact the editor for guidance. This is in accordance with the Ethical Guidelines for Reviewers published by the Committee on Publication Ethics (COPE Council. Ethical guidelines for peer reviewers. September 2017. www.publicationethics.org) Specific suggestions for editors handling papers: Vigilance. The ultimate responsibility in preventing this behavior lies with the editors handling the papers. Careful consideration of the referee reports to detect unethical behavior, including unjustified requests for citations, particularly citations of the reviewer’s own work, is important and all efforts should be made to prevent such requests being made to the authors. Importantly, when a reviewer requests substantial self-citation, this should be reported to the journal so they can investigate whether or not this is part of a pattern, as in the specific case discussed here. Bioinformatics acknowledges that their editorial controls have failed for some time in this particular case, and sincerely apologizes to our authors, referees and readers for not detecting this sooner. This phenomenon of reviewer-coerced citations is not new (Huggett, 2013; Ioannidis, 2015; Resnik et al., 2008; Thombs and Razykov, 2012; Thombs et al., 2015; Wilhite and Fong, 2012), but also not very well explored in terms of how extensive it may be or how it should be dealt with. We hope this editorial will prompt some discussion on the appropriate balance between the need for peer-reviewer anonymity and the need to alert others to potentially unethical behavior once a pattern is established, particularly when it is difficult to detect such patterns. It is possible that eliminating some of the anonymity, either by open peer-review or publishing anonymized peer reviews alongside accepted papers may disincentivize this behavior. Similarly, because highly centralized research resources, such as Publons and ORCID, have been developed, we hope that some ideas or discussion could take place regarding how these or similar centralized resources could be used, responsibly, to help document patterns of ethical concern that are otherwise difficult to detect. Conflict of Interest: none declared.
Jonathan D. Wren, Alfonso Valencia, Janet Kelso
Bioinform.2
2019 Patient Dossier: Healthcare queries over distributed resources
abstract
As with many other aspects of the modern world, in healthcare, the explosion of data and resources opens new opportunities for the development of added-value services. Still, a number of specific conditions on this domain greatly hinders these developments, including ethical and legal issues, fragmentation of the relevant data in different locations, and a level of (meta)data complexity that requires great expertise across technical, clinical, and biological domains. We propose the Patient Dossier paradigm as a way to organize new innovative healthcare services that sorts the current limitations. The Patient Dossier conceptual framework identifies the different issues and suggests how they can be tackled in a safe, efficient, and responsible way while opening options for independent development for different players in the healthcare sector. An initial implementation of the Patient Dossier concepts in the Rbbt framework is available as open-source at https://github.com/mikisvaz and https://github.com/Rbbt-Workflows.
Miguel Vázquez, Alfonso Valencia
PLoS Comput. Biol.2
2017 ISCB's initial reaction to New England Journal of Medicine editorial on data sharing
abstract
The recent editorial by Dr Longo and Dr Drazen in the New England Journal of Medicine (Longo and Drazen, 2016) has stirred up quite a bit of controversy. As Executive Officers of the International Society of Computational Biology, Inc. (ISCB), we express our deep concern about the restrictive and potentially damaging opinions voiced in this editorial, and while ISCB works to write a detailed response, we felt it necessary to promptly address the editorial with this reaction. Although some of the concerns voiced by the authors of the editorial are worth considering, large parts of the statement purport an obsolete view of hegemony over data that is neither in line with today’s spirit of open access nor furthering an atmosphere where the potential of data can be fully realized. ISCB acknowledges that the additional comment on the editorial (Drazen, 2016) eases some of the polemics unfortunately without addressing some of the core issues. We still feel, however, that we need to contrast the opinion voiced in the editorial with what we consider the axioms of our scientific society, statements that lead into a fruitful future of data-driven science: Data produced with public money should be public in benefit of the science and society Restrictions on the use of public data hamper science and slow progress Open data is the best way to combat fraud and misinterpretations Current large data collections proceed from many sources, are continually accumulated, and require a variety of analytical approaches. Data generation and data analysis overlap in time and are continually updated with new data sets produced by new techniques and new analysis methodologies. Furthermore, in many cases current science functions in consortia in which scientists collaborate toward common goals while preserving their own scientific objectives. Dividing scientists into data providers and data analysts is simplistic and gives a misleading impression of the actual state of biological and biomedical science. ISCB very much supports collaboration between disciplines, including experimental and clinical as well as bioinformatics, as the best way forward to address complex biological problems. But this collaboration cannot be based on imposed restrictions to data access and cannot be contained in professional silos. (The use of expressions such as ‘research parasites’ clearly does not help.) Many bio-communities have made significant progress by endorsing open data policies and, gratefully, public funding agencies have connected to the spirit that they are distributing taxpayers’ money to science and that, therefore, the data that are generated in the course belong to the public. It is, perhaps, natural that some areas of biomedical research are slow in adopting these policies. History and the confidential nature of the relevant data are surely among the reasons. However, in our opinion data hegemony is another, a reason that has to be overcome. The sooner these barriers to progress are removed the sooner the patients will benefit from the current flourishing of biomedical research. Conflict of Interest: none declared.
Bonnie Berger, Terry Gaasterland, Thomas Lengauer, Christine A. Orengo, Bruno Gaëta, Scott Markel, Alfonso Valencia
Bioinform.7
2017 CImbinator: a web-based tool for drug synergy analysis in small- and large-scale datasets
abstract
MOTIVATION: Drug synergies are sought to identify combinations of drugs particularly beneficial. User-friendly software solutions that can assist analysis of large-scale datasets are required. RESULTS: CImbinator is a web-service that can aid in batch-wise and in-depth analyzes of data from small-scale and large-scale drug combination screens. CImbinator offers to quantify drug combination effects, using both the commonly employed median effect equation, as well as advanced experimental mathematical models describing dose response relationships. AVAILABILITY AND IMPLEMENTATION: CImbinator is written in Ruby and R. It uses the R package drc for advanced drug response modeling. CImbinator is available at http://cimbinator.bioinfo.cnio.es , the source-code is open and available at https://github.com/Rbbt-Workflows/combination_index . A Docker image is also available at https://hub.docker.com/r/mikisvaz/rbbt-ci_mbinator/ . CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Åsmund Flobak, Miguel Vázquez, Astrid Lægreid, Alfonso Valencia
Bioinform.4
2016 A computational approach inspired by simulated annealing to study the stability of protein interaction networks in cancer and neurological disorders
abstract
Molecular networks provide a powerful tool for the study of biomedical systems, in particular several studies have detected alterations of the network structure associated to disease states. Here we propose that diseases cannot only alter the structure of the network but also its stability. To evaluate network stability we have developed a new methodological framework. Our approach is an adaptation of the classical Deterministic Simulated Annealing algorithm to work with discrete states. Adjusted energy values are used to compare the network stability in disease and control states. The results show that cancer networks are less stable than the Alzheimer’s disease (AD) ones. These results can be interpreted in terms of our previous observations on cancer and AD inverse comorbidity, i.e. AD patients have lower than expected risk to suffer cancer.
Kristina Ibáñez, María Guijarro 0001, Gonzalo Pajares, Alfonso Valencia
Data Min. Knowl. Discov.4
2016 ISCB's Initial Reaction to The New England Journal of Medicine Editorial on Data Sharing
abstract
This message is a response from the ISCB in light of the recent the New England Journal of Medicine (NEJM) editorial around data sharing.
Bonnie Berger, Terry Gaasterland, Thomas Lengauer, Christine A. Orengo, Bruno Gaëta, Scott Markel, Alfonso Valencia
PLoS Comput. Biol.7
2015 Alternative splicing and co-option of transposable elements: the case of TMPO/LAP2α and ZNF451 in mammals
abstract
Transposable elements constitute a large fraction of vertebrate genomes and, during evolution, may be co-opted for new functions. Exonization of transposable elements inserted within or close to host genes is one possible way to generate new genes, and alternative splicing of the new exons may represent an intermediate step in this process. The genes TMPO and ZNF451 are present in all vertebrate lineages. Although they are not evolutionarily related, mammalian TMPO and ZNF451 do have something in common-they both code for splice isoforms that contain LAP2alpha domains. We found that these LAP2alpha domains have sequence similarity to repetitive sequences in non-mammalian genomes, which are in turn related to the first ORF from a DIRS1-like retrotransposon. This retrotransposon domestication happened separately and resulted in proteins that combine retrotransposon and host protein domains. The alternative splicing of the retrotransposed sequence allowed the production of both the new and the untouched original isoforms, which may have contributed to the success of the colonization process. The LAP2alpha-specific isoform of TMPO (LAP2α) has been co-opted for important roles in the cell, whereas the ZNF451 LAP2alpha isoform is evolving under strong purifying selection but remains uncharacterized.
Federico Abascal, Michael L. Tress, Alfonso Valencia
Bioinform.3
2015 FUN-L: gene prioritization for RNAi screens
abstract
MOTIVATION: Most biological processes remain only partially characterized with many components still to be identified. Given that a whole genome can usually not be tested in a functional assay, identifying the genes most likely to be of interest is of critical importance to avoid wasting resources. RESULTS: Given a set of known functionally related genes and using a state-of-the-art approach to data integration and mining, our Functional Lists (FUN-L) method provides a ranked list of candidate genes for testing. Validation of predictions from FUN-L with independent RNAi screens confirms that FUN-L-produced lists are enriched in genes with the expected phenotypes. In this article, we describe a website front end to FUN-L. AVAILABILITY AND IMPLEMENTATION: The website is freely available to use at http://funl.org
Jonathan G. Lees, Jean-Karim Hériché, Ian Morilla, José María Fernández 0001, Priit Adler, Martin Krallinger, Jaak Vilo, Alfonso Valencia, Jan Ellenberg, Juan Garcia Ranea, Christine A. Orengo
Bioinform.8
2015 Detection of significant protein coevolution
abstract
MOTIVATION: The evolution of proteins cannot be fully understood without taking into account the coevolutionary linkages entangling them. From a practical point of view, coevolution between protein families has been used as a way of detecting protein interactions and functional relationships from genomic information. The most common approach to inferring protein coevolution involves the quantification of phylogenetic tree similarity using a family of methodologies termed mirrortree. In spite of their success, a fundamental problem of these approaches is the lack of an adequate statistical framework to assess the significance of a given coevolutionary score (tree similarity). As a consequence, a number of ad hoc filters and arbitrary thresholds are required in an attempt to obtain a final set of confident coevolutionary signals. RESULTS: In this work, we developed a method for associating confidence estimators (P values) to the tree-similarity scores, using a null model specifically designed for the tree comparison problem. We show how this approach largely improves the quality and coverage (number of pairs that can be evaluated) of the detected coevolution in all the stages of the mirrortree workflow, independently of the starting genomic information. This not only leads to a better understanding of protein coevolution and its biological implications, but also to obtain a highly reliable and comprehensive network of predicted interactions, as well as information on the substructure of macromolecular complexes using only genomic information. AVAILABILITY AND IMPLEMENTATION: The software and datasets used in this work are freely available at: http://csbg.cnb.csic.es/pMT/. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
David Ochoa, David de Juan, Alfonso Valencia, Florencio Pazos
Bioinform.3
2015 Structure-PPi: a module for the annotation of cancer-related single-nucleotide variants at protein-protein interfaces
abstract
MOTIVATION: The interpretation of cancer-related single-nucleotide variants (SNVs) considering the protein features they affect, such as known functional sites, protein-protein interfaces, or relation with already annotated mutations, might complement the annotation of genetic variants in the analysis of NGS data. Current tools that annotate mutations fall short on several aspects, including the ability to use protein structure information or the interpretation of mutations in protein complexes. RESULTS: We present the Structure-PPi system for the comprehensive analysis of coding SNVs based on 3D protein structures of protein complexes. The 3D repository used, Interactome3D, includes experimental and modeled structures for proteins and protein-protein complexes. Structure-PPi annotates SNVs with features extracted from UniProt, InterPro, APPRIS, dbNSFP and COSMIC databases. We illustrate the usefulness of Structure-PPi with the interpretation of 1 027 122 non-synonymous SNVs from COSMIC and the 1000G Project that provides a collection of ∼172 700 SNVs mapped onto the protein 3D structure of 8726 human proteins (43.2% of the 20 214 SwissProt-curated proteins in UniProtKB release 2014_06) and protein-protein interfaces with potential functional implications. AVAILABILITY AND IMPLEMENTATION: Structure-PPi, along with a user manual and examples, isavailable at http://structureppi.bioinfo.cnio.es/Structure, the code for local installations at https://github.com/Rbbt-Workflows
Miguel Vázquez, Alfonso Valencia, Tirso Pons
Bioinform.2
2015 Summary of the BioLINK SIG 2013 meeting at ISMB/ECCB 2013
abstract
UNLABELLED: The ISMB Special Interest Group on Linking Literature, Information and Knowledge for Biology (BioLINK) organized a one-day workshop at ISMB/ECCB 2013 in Berlin, Germany. The theme of the workshop was 'Roles for text mining in biomedical knowledge discovery and translational medicine'. This summary reviews the outcomes of the workshop. Meeting themes included concept annotation methods and applications, extraction of biological relationships and the use of text-mined data for biological data analysis. AVAILABILITY AND IMPLEMENTATION: All articles are available at http://biolinksig.org/proceedings-online/.
Karin Verspoor, Hagit Shatkay, Lynette Hirschman, Christian Blaschke, Alfonso Valencia
Bioinform.5
2015 Alternatively Spliced Homologous Exons Have Ancient Origins and Are Highly Expressed at the Protein Level
abstract
Alternative splicing of messenger RNA can generate a wide variety of mature RNA transcripts, and these transcripts may produce protein isoforms with diverse cellular functions. While there is much supporting evidence for the expression of alternative transcripts, the same is not true for the alternatively spliced protein products. Large-scale mass spectroscopy experiments have identified evidence of alternative splicing at the protein level, but with conflicting results. Here we carried out a rigorous analysis of the peptide evidence from eight large-scale proteomics experiments to assess the scale of alternative splicing that is detectable by high-resolution mass spectroscopy. We find fewer splice events than would be expected: we identified peptides for almost 64% of human protein coding genes, but detected just 282 splice events. This data suggests that most genes have a single dominant isoform at the protein level. Many of the alternative isoforms that we could identify were only subtly different from the main splice isoform. Very few of the splice events identified at the protein level disrupted functional domains, in stark contrast to the two thirds of splice events annotated in the human genome that would lead to the loss or damage of functional domains. The most striking result was that more than 20% of the splice isoforms we identified were generated by substituting one homologous exon for another. This is significantly more than would be expected from the frequency of these events in the genome. These homologous exon substitution events were remarkably conserved--all the homologous exons we identified evolved over 460 million years ago--and eight of the fourteen tissue-specific splice isoforms we identified were generated from homologous exons. The combination of proteomics evidence, ancient origin and tissue-specific splicing indicates that isoforms generated from homologous exons may have important cellular roles.
Federico Abascal, Iakes Ezkurdia, Juan Rodriguez-Rivas, Angela del Pozo, Jesús Vázquez, Alfonso Valencia, Michael L. Tress
PLoS Comput. Biol.7
2014 CheNER: chemical named entity recognizer
abstract
MOTIVATION: Chemical named entity recognition is used to automatically identify mentions to chemical compounds in text and is the basis for more elaborate information extraction. However, only a small number of applications are freely available to identify such mentions. Particularly challenging and useful is the identification of International Union of Pure and Applied Chemistry (IUPAC) chemical compounds, which due to the complex morphology of IUPAC names requires more advanced techniques than that of brand names. RESULTS: We present CheNER, a tool for automated identification of systematic IUPAC chemical mentions. We evaluated different systems using an established literature corpus to show that CheNER has a superior performance in identifying IUPAC names specifically, and that it makes better use of computational resources. AVAILABILITY AND IMPLEMENTATION: http://metres.udl.cat/index.php/9-download/4-chener, http://chener.bioinfo.cnio.es/
Anabel Usie, Rui Alves, Francesc Solsona Tehàs, Miguel Vázquez, Alfonso Valencia
Bioinform.5
2014 Predicting Protein Relationshipsto Human Pathways througha Relational Learning ApproachBased on Simple Sequence Features
abstract
UNLABELLED: Biological pathways are important elements of systems biology and in the past decade, an increasing number of pathway databases have been set up to document the growing understanding of complex cellular processes. Although more genome-sequence data are becoming available, a large fraction of it remains functionally uncharacterized. Thus, it is important to be able to predict the mapping of poorly annotated proteins to original pathway models. RESULTS: We have developed a Relational Learning-based Extension (RLE) system to investigate pathway membership through a function prediction approach that mainly relies on combinations of simple properties attributed to each protein. RLE searches for proteins with molecular similarities to specific pathway components. Using RLE, we associated 383 uncharacterized proteins to 28 pre-defined human Reactome pathways, demonstrating relative confidence after proper evaluation. Indeed, in specific cases manual inspection of the database annotations and the related literature supported the proposed classifications. Examples of possible additional components of the Electron transport system, Telomere maintenance and Integrin cell surface interactions pathways are discussed in detail. AVAILABILITY: All the human predicted proteins in the 2009 and 2012 releases 30 and 40 of Reactome are available at http://rle.bioinfo.cnio.es.
Beatriz García Jiménez, Tirso Pons, Araceli Sanchis, Alfonso Valencia
IEEE ACM Trans. Comput. Biol. Bioinform.4
2013 RUbioSeq: a suite of parallelized pipelines to automate exome variation and bisulfite-seq analyses
abstract
MOTIVATION: RUbioSeq has been developed to facilitate the primary and secondary analysis of re-sequencing projects by providing an integrated software suite of parallelized pipelines to detect exome variants (single-nucleotide variants and copy number variations) and to perform bisulfite-seq analyses automatically. RUbioSeq's variant analysis results have been already validated and published. AVAILABILITY: http://rubioseq.sourceforge.net/.
Miriam Rubio-Camarillo, Gonzalo Gómez-López, José María Fernández 0001, Alfonso Valencia, David G. Pisano
Bioinform.4
2013 wKinMut: An integrated tool for the analysis and interpretation of mutations in human protein kinases
abstract
BACKGROUND: Protein kinases are involved in relevant physiological functions and a broad number of mutations in this superfamily have been reported in the literature to affect protein function and stability. Unfortunately, the exploration of the consequences on the phenotypes of each individual mutation remains a considerable challenge. RESULTS: The wKinMut web-server offers direct prediction of the potential pathogenicity of the mutations from a number of methods, including our recently developed prediction method based on the combination of information from a range of diverse sources, including physicochemical properties and functional annotations from FireDB and Swissprot and kinase-specific characteristics such as the membership to specific kinase groups, the annotation with disease-associated GO terms or the occurrence of the mutation in PFAM domains, and the relevance of the residues in determining kinase subfamily specificity from S3Det. This predictor yields interesting results that compare favourably with other methods in the field when applied to protein kinases.Together with the predictions, wKinMut offers a number of integrated services for the analysis of mutations. These include: the classification of the kinase, information about associations of the kinase with other proteins extracted from iHop, the mapping of the mutations onto PDB structures, pathogenicity records from a number of databases and the classification of mutations in large-scale cancer studies. Importantly, wKinMut is connected with the SNP2L system that extracts mentions of mutations directly from the literature, and therefore increases the possibilities of finding interesting functional information associated to the studied mutations. CONCLUSIONS: wKinMut facilitates the exploration of the information available about individual mutations by integrating prediction approaches with the automatic extraction of information from the literature (text mining) and several state-of-the-art databases.wKinMut has been used during the last year for the analysis of the consequences of mutations in the context of a number of cancer genome projects, including the recent analysis of Chronic Lymphocytic Leukemia cases and is publicly available at http://wkinmut.bioinfo.cnio.es.
José M. G. Izarzugaza, Miguel Vázquez, Angela del Pozo, Alfonso Valencia
BMC Bioinform.4
2012 Novel domain combinations in proteins encoded by chimeric transcripts
abstract
MOTIVATION: Chimeric RNA transcripts are generated by different mechanisms including pre-mRNA trans-splicing, chromosomal translocations and/or gene fusions. It was shown recently that at least some of chimeric transcripts can be translated into functional chimeric proteins. RESULTS: To gain a better understanding of the design principles underlying chimeric proteins, we have analyzed 7,424 chimeric RNAs from humans. We focused on the specific domains present in these proteins, comparing their permutations with those of known human proteins. Our method uses genomic alignments of the chimeras, identification of the gene-gene junction sites and prediction of the protein domains. We found that chimeras contain complete protein domains significantly more often than in random data sets. Specifically, we show that eight different types of domains are over-represented among all chimeras as well as in those chimeras confirmed by RNA-seq experiments. Moreover, we discovered that some chimeras potentially encode proteins with novel and unique domain combinations. Given the observed prevalence of entire protein domains in chimeras, we predict that certain putative chimeras that lack activation domains may actively compete with their parental proteins, thereby exerting dominant negative effects. More generally, the production of chimeric transcripts enables a combinatorial increase in the number of protein products available, which may disturb the function of parental genes and influence their protein-protein interaction network. AVAILABILITY: our scripts are available upon request.
Milana Frenkel-Morgenstern, Alfonso Valencia
Bioinform.2
2012 EnrichNet: network-based gene set enrichment analysis
abstract
MOTIVATION: Assessing functional associations between an experimentally derived gene or protein set of interest and a database of known gene/protein sets is a common task in the analysis of large-scale functional genomics data. For this purpose, a frequently used approach is to apply an over-representation-based enrichment analysis. However, this approach has four drawbacks: (i) it can only score functional associations of overlapping gene/proteins sets; (ii) it disregards genes with missing annotations; (iii) it does not take into account the network structure of physical interactions between the gene/protein sets of interest and (iv) tissue-specific gene/protein set associations cannot be recognized. RESULTS: To address these limitations, we introduce an integrative analysis approach and web-application called EnrichNet. It combines a novel graph-based statistic with an interactive sub-network visualization to accomplish two complementary goals: improving the prioritization of putative functional gene/protein set associations by exploiting information from molecular interaction networks and tissue-specific gene expression data and enabling a direct biological interpretation of the results. By using the approach to analyse sets of genes with known involvement in human diseases, new pathway associations are identified, reflecting a dense sub-network of interactions between their corresponding proteins. AVAILABILITY: EnrichNet is freely available at http://www.enrichnet.org. CONTACT: [email protected], [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics Online.
Enrico Glaab, Anaïs Baudot, Natalio Krasnogor, Reinhard Schneider 0002, Alfonso Valencia
Bioinform.5
2012 Mirroring co-evolving trees in the light of their topologies
abstract
MOTIVATION: Determining the interaction partners among protein/domain families poses hard computational problems, in particular in the presence of paralogous proteins. Available approaches aim to identify interaction partners among protein/domain families through maximizing the similarity between trimmed versions of their phylogenetic trees. Since maximization of any natural similarity score is computationally difficult, many approaches employ heuristics to evaluate the distance matrices corresponding to the tree topologies in question. In this article, we devise an efficient deterministic algorithm which directly maximizes the similarity between two leaf labeled trees with edge lengths, obtaining a score-optimal alignment of the two trees in question. RESULTS: Our algorithm is significantly faster than those methods based on distance matrix comparison: 1 min on a single processor versus 730 h on a supercomputer. Furthermore, we outperform the current state-of-the-art exhaustive search approach in terms of precision, while incurring acceptable losses in recall. AVAILABILITY: A C implementation of the method demonstrated in this article is available at http://compbio.cs.sfu.ca/mirrort.htm
Iman Hajirasouliha, Alexander Schönhuth, David de Juan, Alfonso Valencia, Süleyman Cenk Sahinalp
Bioinform.4
2012 JDet: interactive calculation and visualization of function-related conservation patterns in multiple sequence alignments and structures
abstract
UNLABELLED: We have implemented in a single package all the features required for extracting, visualizing and manipulating fully conserved positions as well as those with a family-dependent conservation pattern in multiple sequence alignments. The program allows, among other things, to run different methods for extracting these positions, combine the results and visualize them in protein 3D structures and sequence spaces. AVAILABILITY AND IMPLEMENTATION: JDet is a multiplatform application written in Java. It is freely available, including the source code, at http://csbg.cnb.csic.es/JDet. The package includes two of our recently developed programs for detecting functional positions in protein alignments (Xdet and S3Det), and support for other methods can be added as plug-ins. A help file and a guided tutorial for JDet are also available.
Thilo Muth, Juan A. García-Martín, Antonio Rausell, David de Juan, Alfonso Valencia, Florencio Pazos
Bioinform.5
2012 Bioimage informatics: a new category in Bioinformatics
abstract
The last two decades have witnessed great advances in biological tissue labeling and automated microscopic imaging that, in turn, have revolutionized how biologists visualize molecular, sub-cellular, cellular, and super-cellular structures and study their respective functions. Tremendous volumes of multi-dimensional bioimaging data are now being generated in almost every branch of biology. How to interpret such image datasets in a quantitative, objective, automatic and efficient way has become a major challenge in current computational biology. Bioimage informatics methods have begun to turn image data into useful biological knowledge (Peng, 2008; Swedlow, et al., 2009; Shamir, et al., 2010; Danuser, 2011). The essential methods of bioimage informatics involve large-scale bioimage generation, visualization, analysis and management. Bioimage informatics also encompasses both hypothesis- and data-driven exploratory approaches, with an emphasis on how to generate biological knowledge and/or gain new insights that would otherwise be hard to achieve. Early work in bioimage informatics began in the late 1990s. Increasingly, computer vision, image analysis, data mining, machine learning and pattern recognition methods have been applied to microscopic images to extract biological information and to generate ontology databases. The growing amount of bioimage data are quickly imposing additional demands on how to store, manage and retrieve such image datasets as well as the associated secondary meta-data. Data analysis, fusion and reconstruction techniques have also been developed to facilitate better image acquisition and formation. Joint analysis of image data in combination with other biological datasets, such as genomes and gene expression profiles, is also becoming more and more commonplace. To meet the need of this growing field, the first international workshop on Bioimage Informatics was organized at Stanford University in 2005. It grew to be an annual event in this field. Other meetings on similar topics and related applications have also emerged since then. In 2010, the annual conference on Intelligent Systems for Molecular Biology (ISMB) established a paper-submission track on bioimaging data analysis and visualization. While there is a noticeable need to publish high quality papers on bioimage informatics, so far no high-impact journal explicitly accepts this category of papers. We believe it is an appropriate time to create this new category in Bioinformatics. As of February 2012, Bioinformatics now includes a new paper submission category in the scope described by the journal at its website as follows: ‘Informatics methods for the acquisition, analysis, mining and visualization of images produced by modern microscopy, with an emphasis on the application of novel computing techniques to solve challenging and significant biological and medical problems at the molecular, sub-cellular, cellular, and super-cellular (organ, organism, and population) levels. This category also encourages large-scale image informatics methods/applications/software, various enabling techniques (e.g. cyber-infrastructures, quantitative validation experiments, pattern recognition, etc.) for such large-scale studies, and joint analysis of multiple heterogeneous datasets that include images as a component. Bioimage related ontology and databases studies, image-oriented large-scale machine learning, data mining, and other analytics techniques are also encouraged. We will not consider image analysis and pattern recognition methods that are solely based on tuning parameters or swapping computational sub-steps, without an in-depth description or demonstration of why such changes are significantly superior for one or more biological problems.’ We would like to thank a number of colleagues and practitioners who have contributed to the creation of this new category. Especially, we thank Eugene Myers, B.S.Manjunath, Badri Roysam, Manfred Auer, Michael Hawrylycz, Jean-Christophe Olivo-Marin, Anne Carpenter and Vebjorn Ljosa in helping to define this new category of paper submissions. We hope the journal Bioinformatics becomes a valuable venue for bioimage informatics researchers to publish their most important work. Contact: gro.imhh.ailenaj@hgnep
Hanchuan Peng, Alex Bateman, Alfonso Valencia, Jonathan D. Wren
Bioinform.3
2012 MyMiner: a web application for computer-assisted biocuration and text annotation
abstract
MOTIVATION: The exponential growth of scientific literature has resulted in a massive amount of unstructured natural language data that cannot be directly handled by means of bioinformatics tools. Such tools generally require structured data, often generated through a cumbersome process of manual literature curation. Herein, we present MyMiner, a free and user-friendly text annotation tool aimed to assist in carrying out the main biocuration tasks and to provide labelled data for the development of text mining systems. MyMiner allows easy classification and labelling of textual data according to user-specified classes as well as predefined biological entities. The usefulness and efficiency of this application have been tested for a range of real-life annotation scenarios of various research topics. AVAILABILITY: http://myminer.armi.monash.edu.au.
David Salgado, Martin Krallinger, Marc Depaule, Elodie Drula, Ashish V. Tendulkar, Florian Leitner, Alfonso Valencia, Christophe Marcelle
Bioinform.7
2012 Chapter 14: Cancer Genome Analysis
abstract
Although there is great promise in the benefits to be obtained by analyzing cancer genomes, numerous challenges hinder different stages of the process, from the problem of sample preparation and the validation of the experimental techniques, to the interpretation of the results. This chapter specifically focuses on the technical issues associated with the bioinformatics analysis of cancer genome data. The main issues addressed are the use of database and software resources, the use of analysis workflows and the presentation of clinically relevant action items. We attempt to aid new developers in the field by describing the different stages of analysis and discussing current approaches, as well as by providing practical advice on how to access and use resources, and how to implement recommendations. Real cases from cancer genome projects are used as examples.
Miguel Vázquez, Victor de la Torre, Alfonso Valencia
PLoS Comput. Biol.3
2011 On the organization of bioinformatics core services in biology-based research institutes
abstract
With the growth of genomics, research institutes are increasingly confronted with the task of providing computational support to biologists and the provision of computational facilities, both for laboratory-based bioinformaticians and bioinformatics research groups. The optimal organization of bioinformatics support units to meet these needs is a common topic of discussion and consideration. During the recent evaluation of a bioinformatics core facility, we ended up discussing—in general terms—what we consider basic guidelines for the organization of bioinformatics core units in large research institutes, based on the experience and developments in our own institutes. We thought that part of our discussion could be useful for others, keeping in mind that every organization has specific needs and requirements, and that bioinformatics support is a fast moving area that requires continuous adaptation. In the following, we summarized what we consider some key recommendations: Finally, there is always the very difficult issue regarding the optimal size of the bioinformatics support units. This depends on many factors such as the diversity of the topics being covered in the institute and the level of the computational skills of the biologists. We will, therefore, not give an exact figure, but rather provide figures from our own institutions which are very comparable in this regard. In our own institutions, NKI-AVL, CNIO and FIMM, we have approximately 1 full-time, institutionally funded bioinformatician to support 100 scientists. In addition, the research groups themselves have up to 5–10 times this number of bioinformaticians working in their own research groups. This number varies, of course, depending on the specific topics researched by a particular group. These so-called ‘embedded’ bioinfomatics researchers are involved in more research-oriented work and are largely funded by outside grants. In our opinion it is very important to develop structures that allow the ‘embedded’ bioinformaticians to meet and communicate. Bioinformatics departments have to clearly separate their service units from their research laboratories. This contributes to the transparency of the organization, especially regarding the allocation of funds. The organization of bioinformatics service unit should be clear with defined missions for each component. It is particularly important to separate tasks by topic. For example: support of database development and maintenance, statistical analysis of high-throughput data, automatic image acquisition and analysis of e.g. light microscopy data and analysis of next-generation sequencing data. The unit should install a ‘users committee’ that includes users and specialists of the unit. This committee would be responsible for establishing clear priorities on what technical platforms are supported or aborted, and defining the rules applied to prioritize projects for bioinformatics support. It is important that the services are provided in a highly transparent manner, including publicly accessible information on statistics of users, project progress, tools and use of resources. Particularly important is the interface between the bioinformatics unit and the medical activities in an institute. This requires the integration of biobanking, medical information (records) and the corresponding genomic information. This activity is key for translational medicine and it requires efficient, multi-disciplinary interaction between clinical and computational experts. This process can be greatly enhanced by a seamless integration of the relevant data streams. One of the key missions of the bioinformatics is to provide training to biologists at a basic level. It is important to also incorporate separate advanced training in new technical developments for ‘hybrid users’ or for expert bioinformaticians. This should preferably be done in collaboration with the bioinformatics research laboratories in the institute. It is useful to nominate a bioinformatics support person, whose task is to guide users in the use of public bioinformatics tools and databases as well as tools and methods developed within the institute. The goal is to facilitate the use of bioinformatics by the biologists using state of the art tools. In conclusion, we can state that we favor a model where a small core group of bioinformaticians provide transparently organized support on more general, institute-wide research problems, while the majority of the bioinformaticians are embedded in the research groups. Embedding ensures more direct interaction between the biologist and the bioinformatician, providing both researchers with a sense of ownership of the project. Not only does this elevate the skills level of both parties, but it also greatly enhances productivity.
Olli-P. Kallioniemi, Lodewyk F. A. Wessels, Alfonso Valencia
Bioinform.3
2011 Overview of the BioCreative III Workshop
abstract
BACKGROUND: The overall goal of the BioCreative Workshops is to promote the development of text mining and text processing tools which are useful to the communities of researchers and database curators in the biological sciences. To this end BioCreative I was held in 2004, BioCreative II in 2007, and BioCreative II.5 in 2009. Each of these workshops involved humanly annotated test data for several basic tasks in text mining applied to the biomedical literature. Participants in the workshops were invited to compete in the tasks by constructing software systems to perform the tasks automatically and were given scores based on their performance. The results of these workshops have benefited the community in several ways. They have 1) provided evidence for the most effective methods currently available to solve specific problems; 2) revealed the current state of the art for performance on those problems; 3) and provided gold standard data and results on that data by which future advances can be gauged. This special issue contains overview papers for the three tasks of BioCreative III. RESULTS: The BioCreative III Workshop was held in September of 2010 and continued the tradition of a challenge evaluation on several tasks judged basic to effective text mining in biology, including a gene normalization (GN) task and two protein-protein interaction (PPI) tasks. In total the Workshop involved the work of twenty-three teams. Thirteen teams participated in the GN task which required the assignment of EntrezGene IDs to all named genes in full text papers without any species information being provided to a system. Ten teams participated in the PPI article classification task (ACT) requiring a system to classify and rank a PubMed® record as belonging to an article either having or not having "PPI relevant" information. Eight teams participated in the PPI interaction method task (IMT) where systems were given full text documents and were required to extract the experimental methods used to establish PPIs and a text segment supporting each such method. Gold standard data was compiled for each of these tasks and participants competed in developing systems to perform the tasks automatically.BioCreative III also introduced a new interactive task (IAT), run as a demonstration task. The goal was to develop an interactive system to facilitate a user's annotation of the unique database identifiers for all the genes appearing in an article. This task included ranking genes by importance (based preferably on the amount of described experimental information regarding genes). There was also an optional task to assist the user in finding the most relevant articles about a given gene. For BioCreative III, a user advisory group (UAG) was assembled and played an important role 1) in producing some of the gold standard annotations for the GN task, 2) in critiquing IAT systems, and 3) in providing guidance for a future more rigorous evaluation of IAT systems. Six teams participated in the IAT demonstration task and received feedback on their systems from the UAG group. Besides innovations in the GN and PPI tasks making them more realistic and practical and the introduction of the IAT task, discussions were begun on community data standards to promote interoperability and on user requirements and evaluation metrics to address utility and usability of systems. CONCLUSIONS: In this paper we give a brief history of the BioCreative Workshops and how they relate to other text mining competitions in biology. This is followed by a synopsis of the three tasks GN, PPI, and IAT in BioCreative III with figures for best participant performance on the GN and PPI tasks. These results are discussed and compared with results from previous BioCreative Workshops and we conclude that the best performing systems for GN, PPI-ACT and PPI-IMT in realistic settings are not sufficient for fully automatic use. This provides evidence for the importance of interactive systems and we present our vision of how best to construct an interactive system for a GN or PPI like task in the remainder of the paper.
Cecilia N. Arighi, Zhiyong Lu, Martin Krallinger, Kevin Cohen 0001, W. John Wilbur, Alfonso Valencia, Lynette Hirschman, Cathy H. Wu
BMC Bioinform.6
2011 Selection of organisms for the co-evolution-based study of protein interactions
abstract
BACKGROUND: The prediction and study of protein interactions and functional relationships based on similarity of phylogenetic trees, exemplified by the mirrortree and related methodologies, is being widely used. Although dependence between the performance of these methods and the set of organisms used to build the trees was suspected, so far nobody assessed it in an exhaustive way, and, in general, previous works used as many organisms as possible. In this work we asses the effect of using different sets of organism (chosen according with various phylogenetic criteria) on the performance of this methodology in detecting protein interactions of different nature. RESULTS: We show that the performance of three mirrortree-related methodologies depends on the set of organisms used for building the trees, and it is not always directly related to the number of organisms in a simple way. Certain subsets of organisms seem to be more suitable for the predictions of certain types of interactions. This relationship between type of interaction and optimal set of organism for detecting them makes sense in the light of the phylogenetic distribution of the organisms and the nature of the interactions. CONCLUSIONS: In order to obtain an optimal performance when predicting protein interactions, it is recommended to use different sets of organisms depending on the available computational resources and data, as well as the type of interactions of interest.
Dorota Herman, David Ochoa, David de Juan, Daniel Lopez, Alfonso Valencia, Florencio Pazos
BMC Bioinform.5
2011 Characterization of pathogenic germline mutations in human Protein Kinases
abstract
BACKGROUND: Protein Kinases are a superfamily of proteins involved in crucial cellular processes such as cell cycle regulation and signal transduction. Accordingly, they play an important role in cancer biology. To contribute to the study of the relation between kinases and disease we compared pathogenic mutations to neutral mutations as an extension to our previous analysis of cancer somatic mutations. First, we analyzed native and mutant proteins in terms of amino acid composition. Secondly, mutations were characterized according to their potential structural effects and finally, we assessed the location of the different classes of polymorphisms with respect to kinase-relevant positions in terms of subfamily specificity, conservation, accessibility and functional sites. RESULTS: Pathogenic Protein Kinase mutations perturb essential aspects of protein function, including disruption of substrate binding and/or effector recognition at family-specific positions. Interestingly these mutations in Protein Kinases display a tendency to avoid structurally relevant positions, what represents a significant difference with respect to the average distribution of pathogenic mutations in other protein families. CONCLUSIONS: Disease-associated mutations display sound differences with respect to neutral mutations: several amino acids are specific of each mutation type, different structural properties characterize each class and the distribution of pathogenic mutations within the consensus structure of the Protein Kinase domain is substantially different to that for non-pathogenic mutations. This preferential distribution confirms previous observations about the functional and structural distribution of the controversial cancer driver and passenger somatic mutations and their use as a proxy for the study of the involvement of somatic mutations in cancer development.
José M. G. Izarzugaza, Lisa E. M. Hopcroft, Anja Baresic, Christine A. Orengo, Andrew C. R. Martin, Alfonso Valencia
BMC Bioinform.6
2010 TopoGSA: network topological gene set analysis
abstract
UNLABELLED: TopoGSA (Topology-based Gene Set Analysis) is a web-application dedicated to the computation and visualization of network topological properties for gene and protein sets in molecular interaction networks. Different topological characteristics, such as the centrality of nodes in the network or their tendency to form clusters, can be computed and compared with those of known cellular pathways and processes. AVAILABILITY: Freely available at http://www.infobiotics.net/topogsa.
Enrico Glaab, Anaïs Baudot, Natalio Krasnogor, Alfonso Valencia
Bioinform.4
2010 Extending pathways and processes using molecular interaction networks to analyse cancer genome data
abstract
BACKGROUND: Cellular processes and pathways, whose deregulation may contribute to the development of cancers, are often represented as cascades of proteins transmitting a signal from the cell surface to the nucleus. However, recent functional genomic experiments have identified thousands of interactions for the signalling canonical proteins, challenging the traditional view of pathways as independent functional entities. Combining information from pathway databases and interaction networks obtained from functional genomic experiments is therefore a promising strategy to obtain more robust pathway and process representations, facilitating the study of cancer-related pathways. RESULTS: We present a methodology for extending pre-defined protein sets representing cellular pathways and processes by mapping them onto a protein-protein interaction network, and extending them to include densely interconnected interaction partners. The added proteins display distinctive network topological features and molecular function annotations, and can be proposed as putative new components, and/or as regulators of the communication between the different cellular processes. Finally, these extended pathways and processes are used to analyse their enrichment in pancreatic mutated genes. Significant associations between mutated genes and certain processes are identified, enabling an analysis of the influence of previously non-annotated cancer mutated genes. CONCLUSIONS: The proposed method for extending cellular pathways helps to explain the functions of cancer mutated genes by exploiting the synergies of canonical knowledge and large-scale interaction data.
Enrico Glaab, Anaïs Baudot, Natalio Krasnogor, Alfonso Valencia
BMC Bioinform.4
2010 The PPI affix dictionary (PPIAD) and BioMethod Lexicon: importance of affixes and tags for recognition of entity mentions and experimental protein interactions
abstract
Substantial text mining efforts are being devoted to detect protein mentions and protein-protein interaction (PPI) relations from scientific articles [1,2]. In this context, the BioCreative challenge showed that the correct identification of the individual interactor proteins is still a challenging task, especially when using full text articles [2]. A systematic analysis of particularities of protein mentions in the context of interaction descriptions was nonetheless missing. Experimental biologists often use specific fusion proteins or protein-tags such as -GST, -His, -Myc, FLAG-, antibodies or fluorescent protein (GFP, YFP, CFP and RFP) tags to detect and visualize interactions. These tags are often mentioned as affixes of the target proteins in the literature. The importance of affixes in biomedical text mining had been addressed in case of affixal negation expressions [3], to consider general posttranslational modifications of proteins [4] and can be observed in trigger verbs used for interaction extraction [5]. We carried out a detailed study on the presence of common affixes belonging to interactor protein mentions in full text sentences considered by database curators as evidential support for experimentally characterized physical protein interactions. Furthermore, we tried to determine whether specific affixes might be useful to detect PPI relevant articles and to correlate affix mentions with particular interaction detection methods. Based on examination of over 3,000 of the previously referred interaction evidence passages we have compiled a collection of 277 interaction relevant affixes (89 suffixes, 176 prefixes and 12 that could be both), which were structured into 36 affix tag classes (26 super-affix and 10 combined or sub-affix classes). Figure 1A shows the frequency of mention of each of the affix tag classes. In the resulting PPI affix dictionary (PPIAD), each affix tag class has been manually linked to experimental qualifiers represented by associated PSI-MI ontology [5] concepts by considering their concept definitions. Additionally, statistical associations of affix tag classes to PSI-MI interaction detection method concepts have been derived through curator-based annotations of the evidence passages. To overcome the limited scope and lexical coverage of terms contained in the PSI-MI ontology we build the BioMethod Lexicon, a collection of experimental method terms important for protein interaction and gene regulation relations, and characterized method term co-mentions with affix tag classes. Within a total set of 6,300 interaction evidence sentences, 1,946 (31 %) mentioned at least one interaction relevant affix, which shows that it is a relatively common feature of interaction descriptions. Using statistical analysis of associations between affix classes and interaction detection method annotations (Chi-square test) we discovered that some of the affix classes showed strong associations to interaction methods, such as between: MI:0096 AF_21 (MI: pull down and PPIAD: gst_pull_down_tag), MI:0676 AF_6 (tandem affinity purification and Tandem_Affinity_Purification_tag), MI:0018 AF_10 (two hybrid and Gal4_tag), MI:0006 AF_4 (anti bait coimmunoprecipitation and Antibody_tag), MI:0055 AF_15 (fluorescent * Correspondence: [email protected] Structural Biology and BioComputing Programme, Spanish National Cancer Research Centre, Madrid, Spain Full list of author information is available at the end of the article Krallinger et al. BMC Bioinformatics 2010, 11(Suppl 5):O1 http://www.biomedcentral.com/1471-2105/11/S5/O1
Martin Krallinger, Ashish V. Tendulkar, Florian Leitner, Andrew Chatr-aryamontri, Alfonso Valencia
BMC Bioinform.5
2010 An Overview of BioCreative II.5
abstract
We present the results of the BioCreative II.5 evaluation in association with the FEBS Letters experiment, where authors created Structured Digital Abstracts to capture information about protein-protein interactions. The BioCreative II.5 challenge evaluated automatic annotations from 15 text mining teams based on a gold standard created by reconciling annotations from curators, authors, and automated systems. The tasks were to rank articles for curation based on curatable protein-protein interactions; to identify the interacting proteins (using UniProt identifiers) in the positive articles (61); and to identify interacting protein pairs. There were 595 full-text articles in the evaluation test set, including those both with and without curatable protein interactions. The principal evaluation metrics were the interpolated area under the precision/recall curve (AUC iP/R), and (balanced) F-measure. For article classification, the best AUC iP/R was 0.70; for interacting proteins, the best system achieved good macroaveraged recall (0.73) and interpolated area under the precision/recall curve (0.58), after filtering incorrect species and mapping homonymous orthologs; for interacting protein pairs, the top (filtered, mapped) recall was 0.42 and AUC iP/R was 0.29. Ensemble systems improved performance for the interacting protein task.
Florian Leitner, Scott A. Mardis, Martin Krallinger, Gianni Cesareni, Lynette Hirschman, Alfonso Valencia
IEEE ACM Trans. Comput. Biol. Bioinform.6
2009 Progress and challenges in predicting protein-protein interaction sites
abstract
The identification of protein-protein interaction sites is an essential intermediate step for mutant design and the prediction of protein networks. In recent years a significant number of methods have been developed to predict these interface residues and here we review the current status of the field. Progress in this area requires a clear view of the methodology applied, the data sets used for training and testing the systems, and the evaluation procedures. We have analysed the impact of a representative set of features and algorithms and highlighted the problems inherent in generating reliable protein data sets and in the posterior analysis of the results. Although it is clear that there have been some improvements in methods for predicting interacting sites, several major bottlenecks remain. Proteins in complexes are still under-represented in the structural databases and in particular many proteins involved in transient complexes are still to be crystallized. We provide suggestions for effective feature selection, and make it clear that community standards for testing, training and performance measures are necessary for progress in the field.
Iakes Ezkurdia, Lisa Bartoli, Piero Fariselli, Rita Casadio, Alfonso Valencia, Michael L. Tress
Briefings Bioinform.5
2009 Automated Alphabet Reduction for Protein Datasets
abstract
BACKGROUND: We investigate automated and generic alphabet reduction techniques for protein structure prediction datasets. Reducing alphabet cardinality without losing key biochemical information opens the door to potentially faster machine learning, data mining and optimization applications in structural bioinformatics. Furthermore, reduced but informative alphabets often result in, e.g., more compact and human-friendly classification/clustering rules. In this paper we propose a robust and sophisticated alphabet reduction protocol based on mutual information and state-of-the-art optimization techniques. RESULTS: We applied this protocol to the prediction of two protein structural features: contact number and relative solvent accessibility. For both features we generated alphabets of two, three, four and five letters. The five-letter alphabets gave prediction accuracies statistically similar to that obtained using the full amino acid alphabet. Moreover, the automatically designed alphabets were compared against other reduced alphabets taken from the literature or human-designed, outperforming them. The differences between our alphabets and the alphabets taken from the literature were quantitatively analyzed. All the above process had been performed using a primary sequence representation of proteins. As a final experiment, we extrapolated the obtained five-letter alphabet to reduce a, much richer, protein representation based on evolutionary information for the prediction of the same two features. Again, the performance gap between the full representation and the reduced representation was small, showing that the results of our automated alphabet reduction protocol, even if they were obtained using a simple representation, are also able to capture the crucial information needed for state-of-the-art protein representations. CONCLUSION: Our automated alphabet reduction protocol generates competent reduced alphabets tailored specifically for a variety of protein datasets. This process is done without any domain knowledge, using information theory metrics instead. The reduced alphabets contain some unexpected (but sound) groups of amino acids, thus suggesting new ways of interpreting the data.
Jaume Bacardit, Michael Stout, Jonathan D. Hirst, Alfonso Valencia, Robert E. Smith 0001, Natalio Krasnogor
BMC Bioinform.4
2009 An integrated approach to the interpretation of Single Amino Acid Polymorphisms within the framework of CATH and Gene3D
abstract
BACKGROUND: The phenotypic effects of sequence variations in protein-coding regions come about primarily via their effects on the resulting structures, for example by disrupting active sites or affecting structural stability. In order better to understand the mechanisms behind known mutant phenotypes, and predict the effects of novel variations, biologists need tools to gauge the impacts of DNA mutations in terms of their structural manifestation. Although many mutations occur within domains whose structure has been solved, many more occur within genes whose protein products have not been structurally characterized. RESULTS: Here we present 3DSim (3D Structural Implication of Mutations), a database and web application facilitating the localization and visualization of single amino acid polymorphisms (SAAPs) mapped to protein structures even where the structure of the protein of interest is unknown. The server displays information on 6514 point mutations, 4865 of them known to be associated with disease. These polymorphisms are drawn from SAAPdb, which aggregates data from various sources including dbSNP and several pathogenic mutation databases. While the SAAPdb interface displays mutations on known structures, 3DSim projects mutations onto known sequence domains in Gene3D. This resource contains sequences annotated with domains predicted to belong to structural families in the CATH database. Mappings between domain sequences in Gene3D and known structures in CATH are obtained using a MUSCLE alignment. 1210 three-dimensional structures corresponding to CATH structural domains are currently included in 3DSim; these domains are distributed across 396 CATH superfamilies, and provide a comprehensive overview of the distribution of mutations in structural space. CONCLUSION: The server is publicly available at http://3DSim.bioinfo.cnio.es/. In addition, the database containing the mapping between SAAPdb, Gene3D and CATH is available on request and most of the functionality is available through programmatic web service access.
José M. G. Izarzugaza, Anja Baresic, Lisa E. M. McMillan, Corin Yeats, Andrew B. Clegg, Christine A. Orengo, Andrew C. R. Martin, Alfonso Valencia
BMC Bioinform.8
2009 Extraction of human kinase mutations from literature, databases and genotyping studies
abstract
BACKGROUND: There is a considerable interest in characterizing the biological role of specific protein residue substitutions through mutagenesis experiments. Additionally, recent efforts related to the detection of disease-associated SNPs motivated both the manual annotation, as well as the automatic extraction, of naturally occurring sequence variations from the literature, especially for protein families that play a significant role in signaling processes such as kinases. Systematic integration and comparison of kinase mutation information from multiple sources, covering literature, manual annotation databases and large-scale experiments can result in a more comprehensive view of functional, structural and disease associated aspects of protein sequence variants. Previously published mutation extraction approaches did not sufficiently distinguish between two fundamentally different variation origin categories, namely natural occurring and induced mutations generated through in vitro experiments. RESULTS: We present a literature mining pipeline for the automatic extraction and disambiguation of single-point mutation mentions from both abstracts as well as full text articles, followed by a sequence validation check to link mutations to their corresponding kinase protein sequences. Each mutation is scored according to whether it corresponds to an induced mutation or a natural sequence variant. We were able to provide direct literature links for a considerable fraction of previously annotated kinase mutations, enabling thus more efficient interpretation of their biological characterization and experimental context. In order to test the capabilities of the presented pipeline, the mutations in the protein kinase domain of the kinase family were analyzed. Using our literature extraction system, we were able to recover a total of 643 mutations-protein associations from PubMed abstracts and 6,970 from a large collection of full text articles. When compared to state-of-the-art annotation databases and high throughput genotyping studies, the mutation mentions extracted from the literature overlap to a good extent with the existing knowledgebases, whereas the remaining mentions suggest new mutation records that were not previously annotated in the databases. CONCLUSION: Using the proposed residue disambiguation and classification approach, we were able to differentiate between natural variant and mutagenesis types of mutations with an accuracy of 93.88. The resulting system is useful for constructing a Gold Standard set of mutations extracted from the literature by human experts with minimal manual curation effort, providing direct pointers to relevant evidence sentences. Our system is able to recover mutations from the literature that are not present in state-of-the-art databases. Human expert manual validation of a subset of the literature extracted mutations conducted on 100 mutations from PubMed abstracts highlights that almost three quarters (72%) of the extracted mutations turned out to be correct, and more than half of these had not been previously annotated in databases.
Martin Krallinger, José M. G. Izarzugaza, Carlos Rodríguez Penagos, Alfonso Valencia
BMC Bioinform.4
2008 Protein Interactions Extracted from Genomes and Papers
Alfonso Valencia
APBC1
2008 iProLINK: A Framework for Linking Text Mining with Ontology and Systems Biology
abstract
The ever-increasing scientific literature and the exponential growth of large-scale molecular data have prompted active research in biological text mining to facilitate literature-based curation of molecular databases. Meanwhile, systems biology and bio-ontologies are emerging as critical tools in biological research where complex data in disparate resources are generated, integrated and analyzed. Both rely on literature for data annotation and analysis. The challenges facing us are to develop broadly utilized text mining tools and systems, and to bring together developer and user communities for system development and evaluation. We describe a framework for linking text mining tools with ontology and systems biology, extending from a previously developed text mining resource, iProLINK. We focus on molecular and ontological resources, including genes/proteins, protein-protein interaction (PPI), and Protein Ontology. The framework consists of two major components: a user interface for text mining of PPI from an integrated tool server and software modules to allow text mining outputs to be created, ranked, and used by the community. Use cases are presented for assessing the gaps and making recommendations for future development.
Zhang-Zhi Hu, Kevin Cohen 0001, Lynette Hirschman, Alfonso Valencia, Michelle G. Giglio, Cathy H. Wu
BIBM4
2008 Editorial
abstract
This special issue contains the papers accepted for presentation at the Sixteenth ISMB Conference (http://www.iscb.org/ismb2008/). The ‘Intelligent Systems for Molecular Biology’ conference took place in Toronto, Canada from July 19–July 23, 2008. The papers contained in this volume are a selection of the 287 papers submitted to the conference. The papers were organized in 10 areas edited by two scientists. Half of Area Chairs were new as ISMB editors (see Table 1). In contrast with the previous years we decided to join the ‘Protein Structure’ and ‘Protein Function’ areas. aNew Area Chair this year. The Area Chairs were responsible for selecting the member of the Program Committee that comprised of more than 300 reviewers who were assisted by a sizeable number of sub-reviewers. Most papers received three reviews and many of them four or more reports. There was an intense discussion between referees fostered by the area chairs. In a first round of evaluation, 224 papers were selected for rejection, 32 for acceptance and 31 as undecided. The papers classified for potential acceptance and those classified as undecided were revised by the Conference Chairs and by the Area Chairs, and most of them were referred again to the referees, with additional Area Editors’ comments. The phone conference confirmed the selection of 34 papers and reserved another 15 for additional discussion, in most cases by obtaining new reports or additional editorial advice in light of previous comments. For all the papers in which the referee comments might have confused the authors, editorial comments explaining the basic reasons for the rejection of the papers were introduced. In total, 45 papers were finally accepted, and four more accepted after compulsory revision. Unfortunately, one of these 49 accepted papers was withdrawn by the authors during the production phase. We are happy to say that we have received a single rebuttal letter from the 224 submissions. We truly believe that Area Editors and Referees made a consistent effort to increase the quality of the referee process by selecting only papers of scientific excellence in their area. The selected papers have a balance between the novelty of the methods and the significance of their contribution to the corresponding areas of biology, as well as being of general interest for the conference. After such a large evaluation effort it is difficult to imagine that the quality of the conference can be further improved by additional evaluation efforts. In our opinion the only way to produce this desired quality transition, would be to introduce a second round of evaluation, a topic that is being discussed for ISMB09. The final acceptance rate was 17.1%, a ratio similar to previous year's (15.8% in 2007). The corresponding authors of submitted papers came from 26 countries, 244 authors were from North America, 104 from Europe and Israel, 51 from Asia, 2 from South America and 5 from Australia. We expect the conference to be a scientific success and a solid contribution to the role of ISCB and its contribution to Molecular Biology and Biomedicine. We would like to thank the Area Chairs and reviewers for their effort, dedication and quality of their professional work; to Thomas Lengauer and Mario Albretch for sharing their invaluable experience in ISMB-ECCB07, Andrei Voronkov for technical support with the EasyChair system, the team at Oxford Journals for proof-setting the papers; the Conference Chairs, Burkhard Rost, Michal Linial , Jill Mesirov for their continuous support, and Steven Leard for helping us to deal with the all kind of issues. Alfonso Valencia (ISMB Paper Selection Chair) Ildefonso Cases (Co-Chair)
Alfonso Valencia, Ildefonso Cases
ISMB1
2008 Determination and validation of principal gene products
abstract
MOTIVATION: Alternative splicing has the potential to generate a wide range of protein isoforms. For many computational applications and for experimental research, it is important to be able to concentrate on the isoform that retains the core biological function. For many genes this is far from clear. RESULTS: We have combined five methods into a pipeline that allows us to detect the principal variant for a gene. Most of the methods were based on conservation between species, at the level of both gene and protein. The five methods used were the conservation of exonic structure, the detection of non-neutral evolution, the conservation of functional residues, the existence of a known protein structure and the abundance of vertebrate orthologues. The pipeline was able to determine a principal isoform for 83% of a set of well-annotated genes with multiple variants.
Michael L. Tress, Jan-Jaap Wesselink, Adam Frankish, Gonzalo López, Nick Goldman, Ari Löytynoja, Tim Massingham, Fabio Pardi, Simon Whelan, Jennifer L. Harrow, Alfonso Valencia
Bioinform.11
2008 Enhancing the prediction of protein pairings between interacting families using orthology information
abstract
BACKGROUND: It has repeatedly been shown that interacting protein families tend to have similar phylogenetic trees. These similarities can be used to predicting the mapping between two families of interacting proteins (i.e. which proteins from one family interact with which members of the other). The correct mapping will be that which maximizes the similarity between the trees. The two families may eventually comprise orthologs and paralogs, if members of the two families are present in more than one organism. This fact can be exploited to restrict the possible mappings, simply by impeding links between proteins of different organisms. We present here an algorithm to predict the mapping between families of interacting proteins which is able to incorporate information regarding orthologues, or any other assignment of proteins to "classes" that may restrict possible mappings. RESULTS: For the first time in methods for predicting mappings, we have tested this new approach on a large number of interacting protein domains in order to statistically assess its performance. The method accurately predicts around 80% in the most favourable cases. We also analysed in detail the results of the method for a well defined case of interacting families, the sensor and kinase components of the Ntr-type two-component system, for which up to 98% of the pairings predicted by the method were correct. CONCLUSION: Based on the well established relationship between tree similarity and interactions we developed a method for predicting the mapping between two interacting families using genomic information alone. The program is available through a web interface.
José M. G. Izarzugaza, David de Juan, Carles Pons, Florencio Pazos, Alfonso Valencia
BMC Bioinform.5
2008 Defining functional distances over Gene Ontology
abstract
BACKGROUND: A fundamental problem when trying to define the functional relationships between proteins is the difficulty in quantifying functional similarities, even when well-structured ontologies exist regarding the activity of proteins (i.e. 'gene ontology' -GO-). However, functional metrics can overcome the problems in the comparing and evaluating functional assignments and predictions. As a reference of proximity, previous approaches to compare GO terms considered linkage in terms of ontology weighted by a probability distribution that balances the non-uniform 'richness' of different parts of the Direct Acyclic Graph. Here, we have followed a different approach to quantify functional similarities between GO terms. RESULTS: We propose a new method to derive 'functional distances' between GO terms that is based on the simultaneous occurrence of terms in the same set of Interpro entries, instead of relying on the structure of the GO. The coincidence of GO terms reveals natural biological links between the GO functions and defines a distance model Df which fulfils the properties of a Metric Space. The distances obtained in this way can be represented as a hierarchical 'Functional Tree'. CONCLUSION: The method proposed provides a new definition of distance that enables the similarity between GO terms to be quantified. Additionally, the 'Functional Tree' defines groups with biological meaning enhancing its utility for protein function comparison and prediction. Finally, this approach could be for function-based protein searches in databases, and for analysing the gene clusters produced by DNA array experiments.
Angela del Pozo, Florencio Pazos, Alfonso Valencia
BMC Bioinform.3
2007 Editorial
abstract
During 2006 we have received over 2000 manuscript submissions. Our acceptance rate over the past year has been close to 30%. We are pleased to see that more authors are submitting novel biological insights through our Discovery Notes articles, including a growing number of novel protein domain discoveries. We are proud to have published a number of special issues including full volumes dedicated to the two main conferences in the field: ISMB and ECCB 2006. We are also pleased to note that the latest impact factor for Bioinformatics from the ISI (Institute for Scientific Information) has continued to increase, from 5.7 to 6.0. Papers are published rapidly online ahead of print, within one week of acceptance on average, and usually appear in the print version of the journal within 12 weeks. The journal has continued to publish Open Access papers as part of the Oxford Open initiative (Author Webpage). During 2006, over 20% of Bioinformatics authors have chosen to publish their paper under the Open Access model. This is the highest uptake seen by any journal in the optional Oxford Open initiative. In addition, all content is made freely available online 12 months after publication. In the New Year we will increase the number of review articles in key topics for our community and we are grateful to Jonathan Wren the Associate Editor who is coordinating this. Over the last year we have continued to bring on board new Associate Editors to replace those who are stepping down. Over the past year five editors have stepped down. We would like to thank Charlie Hodgman, Nikolaus Rajewsky, Alvis Brazma, Satoru Miyano and Christos Ouzounis for their great contribution to the journal. We have extended the terms of Martin Bishop, John Quackenbush and Thomas Lengauer for a further three years. We would like to welcome Trey Ideker, Olga Troyanskaya, Limsoon Wong and Burkhard Rost, our new Asso\ciate Editors who have joined through the year. In particular we would like to mention the immense contribution of Christos Ouzounis in shaping Bioinformatics into what it is now. Along with Martin Bishop, Christos was a very active handling editor of the journal for the many years in which the office was at the EBI in Hinxton with Chris Sander as Executive Editor. Among many other contributions Christos was the one behind the creation of the Discovery Notes section, the association with ISMB for the publication of their annual special issue, and for the creation of the Bioinformatics Associate Editor role. He will now continue working with the journal as member of the Editorial Board from his new position in Tessaloniki. We would also like to thank our Editorial Board for helpful advice and suggestions throughout the year. Finally we would like to give our thanks to the many thousands of referees who have freely given up their time to review and improve the work submitted to the journal.
Alfonso Valencia, Alex Bateman
Bioinform.1
2006 Structural genomics meets computational biology
abstract
A meeting recently organized by the NIH NIGMS Protein Structure Initiative (PSI, Author Webpage) has made crystal clear the urgency and importance of the development of computational methods for the analysis of protein families, definition of protein domains and regions for expression, and annotation of protein function. No really new problems, but problems made now even more important for the development of the Structural Genomics projects. PSI is now in the first year of the production phase (after a 5 year pilot project) with a projected annual cost of 66 million dollars. PSI is now composed of four large production centres: Northeast Structural Genomics Consortium (Author Webpage), Midwest Center for Structural Genomics (Author Webpage), Joint Center for Structural Genomics (Author Webpage) and New York Structural GenomiX Research Consortium (Author Webpage), and six technology-based centers, modeling centers and a database (KnowledgeBase) and material repository associated project to be funded in September 2006. The project has the core activities around a number of committees, including the target selection Steering Subcommittee that includes the Bioinformatics group comprising our colleagues A. Godzik, A. Fiser, C. Orengo and B. Rost. This subcommittee was the one that organized the meeting to discuss the best strategies for selecting proteins to enter into the protein structure resolution pipeline of the PSI centers. It was also this committee that was lucky enough to choose the four most rainy days in Bethesda in the last 50 years. Target selection is technically and scientifically an important issue that was discussed under the growing impression that the biological community at large is not well informed of the progress and are apparently not aware of the developments and achievements during the pilot phase of this project. We were convinced during the meeting that the appropriate selection of sensible and clearly define targets will certainly contribute positively to increase the interest of biologist in the developments in structural genomics, and that the interest that can be generated by a clear description of the targets to be solved will have to be reinforced/rewarded/maintained by providing access to the right metrics of progress and success. In this scenario computational biology is essential not only to set the objectives and select the protein targets in the most effective possible way, but also, and perhaps more importantly, to make accessible to the community all the information generated by the SG projects. A number of specific problems related to the selection of targets and the definition of milestones were discussed during the meeting. These problems include the definition of protein families at different levels of granularity, the range of sequences that can be modeled with a given structure, the available strategies for evaluating the quality of the models and the strategies for selecting the more interesting targets in large protein families. The issues related with the biological/biomedical interest of the targets and the possibilities in the difficult issue of measuring the coverage of the function space were also considered to be of great interest for the definition of the target selection strategy. It was rewarding for us to see how all these problems of fundamental importance for the development of SG are directly part of the realm of bioinformatics/computational biology. It is also to be said that it is a great community responsibility to redouble our efforts to make available additional methods and resources to address these problems. Specific proposals such as the organization of an annual conference to stimulate the work on the subject related with target selection and analysis with the scientific community outside the SG projects will certainly be steps in the right direction. Fostering the undergoing efforts in the SG projects to make openly available, and easily accessible to computational biologist, the results of the on-going experiments will also create a very positive flow of research and development. For example, it is urgent to make publicly accessible the biophysical results obtained for the many constructions that the SG centers have tested for expression, solubility and crystallization. This resource can be invaluable for the development of domain boundary prediction methods, which in turn will contribute to speed-up the experimental work with complex eukaryotic proteins. Finally, given the importance of SG for the development of biology and biomedicine, and the fundamental importance of bioinformatics for the organization, analysis and exploitation of the results, it is unavoidable to think that the effort dedicated in the different Structural Genomics international projects to the computational analysis should be increased urgently.
Alex Bateman, Alfonso Valencia
Bioinform.2
2006 Semantic Mining in Biomedicine (Introduction to the papers selected from the SMBM 2005 Symposium, Hinxton, U.K., April 2005)
abstract
Researchers working in the life sciences domain in the past years have witnessed an enormous growth of literature—for the whole field as well as for their highly specialized areas of expertise. Only small portions of the biomedical knowledge are accessible in a structured way, i.e. through formatted databases. These few pieces of textually encoded knowledge that have gone into databases are, by default, manually extracted from documents and manually inserted into databases after careful curation efforts by highly skilled domain experts. Still, the vast majority of biomedical knowledge captured in texts is not at disposal when biomedical databases are queried. Life scientists have realized this loss of possibly highly relevant information and devised various forms of support. The weakest one is provided by information retrieval (IR) systems [for a life-science-centred survey; cf. Hersh (2002)]. Given a user-formulated query the terms from this query are appropriately matched with the terms occurring in documents from a large collection (e.g. the, currently, 14 million abstracts from Medline). Documents matching the query (up to a specified degree) are returned to the user for closer inspection and, possibly, ranked by some relevance-based sorting criterion (e.g. closeness of match). Information extraction (IE) provides a more powerful alternative that has mainly been developed in other areas different from molecular biology by the natural language processing community. IE aims at directly extracting relevant information from natural language documents [usually original text snippets, sentences, relevant phrases or even quasi-logical propositions, such as predicate–argument structures—for a general survey, cf. Gaizauskas and Wilks (1998) and for a life-science-centred view, cf. Blaschke et al. (2002) and Hoffmann et al. (2005)]. Unlike the output of IR systems, which only list relevant documents, IE systems provide immediate access to relevant information pieces via pre-specified information templates. This is achieved, however, at the price of supplying rather sophisticated language processing methodologies [e.g. taggers, chunkers, light semantic interpreters and information extraction rules; cf. for a survey, Hahn and Wermter (2006)], domain-specific developments and resources (e.g. databases and ontologies) and machine learning methodologies usually lack in IR systems. The evaluation of the degree of achievements from a biomedical perspective is an issue of active research. The IR stream is currently mainly investigated in the TREC (Text Retrieval Conference) Genomics track (Author Webpage), whereas there are several challenge evaluation platforms for IE that deal with often complementary problems from a biological perspective, the most important, currently, being the BioCreAtIvE (Critical Assessment of Information Extraction systems in Biology) contest (Author Webpage) [surveyed in Hirschman et al. (2005); see also Blaschke et al. (2005)]. Both forms of activities, IR as well as IE, are often labelled as text mining but miss a major extra requirement, namely the knowledge discovery perspective usually attributed to text mining procedures as well (Hearst, 1999). In particular, this relates to the identification and elimination of redundant knowledge as well as the recognition of (user-new?, expert-new? and community-new?) novel information. This value-adding, summarizing and selective aspect of text mining could be particularly helpful in taming the flood of literature for biomedical researchers, and will certainly be the focus of new developments in the years to come. The challenge evaluations, however, have already revealed some of the most pressing research problems for text analysis in the biomedical domain. In particular, biomedical terminology is extremely hard to deal with, in part because of the poor introduction of standards. It starts from identifying biological terms in a document (terms have a complex internal structure and are often composed of multiple, up to four or five, words), and leads to determining their conceptual type (e.g. genes, proteins and cell lines) and the way they are relationally linked (e.g. in terms of taxonomies or partonomies that in biology are often related with the organization of protein families). Further on, concrete factual biomedical knowledge (sometimes called relation mining) is also hard to extract from documents (e.g. ‘protein X inhibits protein Y’). This step is crucial for any sort of automated functional annotation in biological databases. Once this kind of knowledge has been successfully captured on a large scale making thousands of these propositions available, another severe follow-up problem arises, namely how to communicate this mass of information in a concise, comprehensible and, finally, useful way to the researcher in the laboratory. For this purpose, text mining systems have simply borrowed visualization techniques that were originally developed for numerical data mining. However, symbolic abstraction mechanisms leading, e.g. to the automatic generation of pathway diagrams from this huge dataset are still an area that requires further developments. The above-mentioned research problems have motivated the creation of a Network of Excellence—‘Semantic Interoperability and Data Mining in Biomedicine’ (Semantic Mining, Author Webpage)—which has been funded by the European Community since 2004 under the FP6 Programme ‘Integrating and Strengthening the European Research Area’. The NoE has initiated a series of conferences dedicated to these particular challenges of data mining and text mining in the life sciences [for a life-science-centred survey, cf. Ananiadou and McNaught (2006)]. The first of these symposia was held under the title ‘Semantic Mining in Biomedicine’ (SMBM) in Hinxton (Cambridgeshire, UK) from April 10–13, 2005 organized by Stefan Schulz, Freiburg University Hospital, and Dietrich Rebholz-Schuhmann, EBI-EMBL, Hinxton (see Author Webpage). A specific feature of SMBM meetings is their focus on content-oriented methodologies and semantic resources—either controlled vocabularies, terminologies and formal domain ontologies, or conceptually as well as propositionally annotated corpora—in order to improve text-based biomedical knowledge management, e.g. through document classification, text or fact retrieval, information extraction, or (real) text mining. Also methodologies being discussed should look at applications to real-world problems in molecular biology and biomedicine [for a review of systems currently operational in this domain, see Krallinger and Valencia (2005)]. We had the honour of chairing the programme committee that comprised 21 scientists who evaluated the 28 submissions and selected 12 papers for their presentation in the conference. Four outstanding papers were selected for publication in Bioinformatics, after additional extensive reviews and revisions. Seven full papers plus the abstracts of these selected presentations appeared in the proceedings of the conference [Hahn and Valencia (2005)]. The selected papers cover research performed under the following headings: (1) entity identification—identification of gene names, (2) text classification classification—assignment of sentences to known Gene Ontology (GO) and Medical Subject Headings (MeSH) classes, (3) identification of relations in text—extracting phosphorylation and gene control networks; and (4) identification of new concepts—proposing new GO categories and their corresponding associated genes. Automatic Term List Generation for Entity Tagging by Ted Sandler, Andrew I. Schein and Lyle H. Ungar from the University of Pennsylvania. The basic problem of term characterization is tackled here with an unsupervised approach based on clustering terms (gene names) using additional context information. The clustering approach is related to the distributional clustering technique published previously and the context information provided include neighbouring and syntactic relations. The basic sources of information were sentences from the Biocreative gene tagging challenge and a set of two million Medline abstracts. The results are significantly better than those obtained with standard taggers based on dictionaries of genes. Interestingly enough, the results are still far from matching those obtained in other domains such as newswire information, most probably owing to the additional complexity of biological nomenclature. Automatic Assignment of Biomedical Categories: Toward a Generic Approach by Patrick Ruch from the University Hospitals of Geneva. Describes new results on the automatical assignment of biomedical categories with a system that is designed to be largely data-independent. The system includes a pattern-based identification and vector space retrieval engine, and uses both stems and linguistically motivated information, and it is applied to the classification of sentences in MeSH and GO classes. The results are compared with those obtained in the related BioCreative task. Extraction of Regulatory Gene/Protein Networks from Medline by Jasmin Saric, Lars Juhl Jensen, Rossitza Ouzounova, Isabel Rojas and Peer Bork, from EML Research and EMBL both in Heidelberg. The authors address the problem of extracting two key types of biological relations, which are the regulators of protein function by phosphorylation and the control of gene expression. Their rule-based String-IE system uses organism-specific lexicons that are incorporated in the training of a part-of-speech tagger that uses the GENIA corpus as background information. In practice, the system is able to extract 3319 phosphorylations or gene expression relations, with a sustained level of accuracy across different organisms. Automatic Extension of GO with Flexible Identification of Candidate Terms by Jin-Bok Lee, Jung-jae Kim and Jong C. Park from KAIST in Daejeon, Korea. The authors tackle the problem of identifying new GO concepts using existing GO concepts and their relations in text. The proposed new terms are compared with those created by human experts in subsequent releases of GO. This type of approaches can be useful for speeding up the process of annotation, and for increasing the number of categories in which GO concepts can be divided when they have a large number of genes assigned.
Udo Hahn, Alfonso Valencia
Bioinform.2
2006 Bioinformatics in the human interactome project
abstract
‘In the early days of the Human Interactome Project, a meeting was organized…’. Perhaps, a few years from now, newspapers will describe in those terms how straightforward it was to plan the large-scale mapping of protein interactions in human and other model organisms. Scientists attending the second Cold Spring Harbor Laboratory/Wellcome Trust symposium on ‘Interactome Networks’1 know well that things are not that easy. Important scientific, technical and sociological issues remain before the ‘Human Interactome Project’ can be considered on its way. But things are definitely moving. At the meeting, Marc Vidal proposed some concrete goals for such a project: ‘To produce hundreds of different sets of cloned ORFs (ORFeomes) and 100 million interactions with a 1–5% false positive rate. To add directionality and signs to the interactions (i.e. activation or inhibition), and to study the variation of the interactions associated with diseases’. Nonetheless, the community still has much to decide. It will still have to agree on these goals, to subdivide the project into recognizable milestones, to set a time line for achieving these milestones, to associate cost to each of the operations, and perhaps most importantly, to obtain funding for such an ambitious endeavor. But there is little doubt that having clear goals will help strengthen the ties between researchers in this already very active community, as well as to engage new partners and grant agencies. As with the Human Genome Project, bioinformatics and computational biology will be of profound importance to any protein interaction mapping effort. Following the presentations during the meeting (for a recent review see Sharan and Ideker, 2006) an early and essential bioinformatics task will be the creation of database standards (i.e. the IMEX interaction database standard and the emerging Biopax bio-pathways standard) and analysis/visualization platforms, such as the one provided by the Cytoscape project. It was also interesting to realize the progress that has already been made in the analysis of the structure, function and properties of protein interaction networks (as well as gene control and metabolic networks), even while the number of reliable datasets is still small. It is also rewarding to see also how the first large-scale simulations based on protein interaction data are becoming a reality. These efforts in analysis and simulation are proceeding in parallel with those dedicated to the prediction of new interactions, modules, motifs and functional properties (phenotypes, diseases and others), in most cases by integrating complementary sources of information on functional and structural interactions. All this domain of emerging ‘Network Biology’ offers a direct connection between computational and the experimental analysis. In this respect, two questions emerge from the meeting as critical for the future: (1) the development of methods able to distinguish physical from functional interactions (and/or different types of physical interaction) and (2) the mapping of the details of the physical interactions (i.e. interacting residues) and other information that is necessary for the interpretation of the variation data (SNPs) and for experimental manipulation of interaction networks. Finally, an interesting controversy arose during the meeting that might have consequences for our bioinformatics community. Some argue that it will be more effective to concentrate all efforts into scale-up of the experimental proteomics technology, postponing the bioinformatics analysis to a second phase once the underlying data are fully (or at least mostly) complete. On the contrary, we think that, it is essential to continuously support the development of the methods that will be required for the interpretation of the Human Interactome Project, including network alignments, annotation, analysis and others. The analogy with the Human Genome Project can be useful here. In that case even if the basic alignment techniques were ready since the 70's when the genome sequencing emerged basic bioinformatic technologies were not available (c.f. just remember the challenging analysis of the first bacterial genome in 1995 (Casari et al., 1995), or the struggle to assemble and represent the first draft of the human genome (Istrail et al., 2004). We are convinced that by pushing in parallel experimental and computational developments we can prepare in a more effective way the future of this area of research. Indeed, much of the current interest in large-scale proteomics is related with the impact that the early computational analysis of the first (and imperfect) datasets have had (i.e. the first ‘scale free’ and ‘motif discovery’ papers of Barabasi (Jeong et al., 2000) and Alon (Shen-Orr et al., 2002) teams have captured the imagination of biologist, physicists and theoreticians like few other problems in molecular biology have). Moreover, integrative and computational approaches have already been indispensable for assessing data quality and scoring confidence in specific interactions as well as whole interaction datasets. Finally, at a practical level, what biologists see as a result of large-scale proteomics are computational representations based on the data provided by databases. Therefore, a successful Human Proteome Project depends intimately on ongoing developments in bioinformatics, as they proceed in parallel with the large-scale experiments.
Trey Ideker, Alfonso Valencia
Bioinform.2
2006 Synthetic Biology: challenges ahead
abstract
This expanding scientific discipline is proving extremely popular and is attracting engineering and system design experts to the field of Biology. As Bioinformatics and Computational Biology will be essential components of new technical and scientific developments, it is vital to follow the discussion generated by the recent ESF Exploratory Workshop (October 13–16, 2005, Constructing and de-constructing Life, Magalia, Spain) and the 2005 report of the NEST High-Level Expert Group on Synthetic Biology: Applying Engineering to Biology Author Webpage) Synthetic Biology stands at the meeting-point of two cultures. The first, represented by those interested in ‘deconstructing life’, dissects biological systems in the search for simplified and minimal forms that will help us understand the adaptation and evolution of natural processes. This approach includes experiments to obtain information on isolated parts of biological systems, the simulation of these systems and then the prediction of associated properties followed by further experimental verification. Well-known examples include work on metabolic pathways such as glycolysis (Hans Westerhoff, Free University of Amsterdam) and the simulation of cell systems using stochastic approaches (Luis Serrano, EMBL, Heidelberg). Simplified systems, based on phospholipids (Doron Lancet, Weizmann Institute, Rehovot) or polymers (Steen Rasmussen, Los Alamos National Laboratory), are used to explore possible prebiotic systems. Directly related to this activity is research into minimal forms of life and minimal genomes Tom Knight (MIT Artificial Intelligence Laboratory). An essential part of this approach is the definition of biological systems that make modeling and simulation feasible: biodegradation networks (Alfonso Valencia, National Center of Biotechnology, Madrid), minimal genomes (Andres Moya, Instituto Cavanilles, Valencia), etc.). The development of computer viruses (Chris Adami, Santa Fe Institute) to study properties of biological evolution can also be included in this kind of research. In short, this approach focuses on the definition of material or virtual systems that help investigate the properties of complex biological problems. The second, complementary and symmetrical culture is the ‘Construction of life’ approach. In this case, the goal is to build systems that inspired by general biological principles, use biological or chemical components to reproduce the behavior of live systems. Highlights of this kind of approach are the designs of Drew Endy (MIT Department of Biology) and Ron Weiss (Department Electrical Engineering Princeton University) who have started to address biological phenomena with the conceptual weaponry of electrical engineering. The general underlying notion is to combine autonomous, modular, robust and reusable components. Characteristic specimens of this sort are the input components that sense a given environment, such as interfacing with biological signals, internal components, processing of biological input information inside a synthetic system to minimize side-effects, and output components, which send the signal processed by the synthetic setup back to the endogenous biological system. A very attractive and on-going by-product of this research is the development of a registry of Standard Biological Parts (Author Webpage), which includes lists of formated components, which, as they comply with international standards, can be easily distributed and shared. The second step will involve the combination of these components into working devices: associated research is seeking to define containers for these machines, which could range from simple lipid vesicles (Peter Walde, ETH; Albert Libchaber, the Rockefeller University) to minimal genomes (Hamilton Smith, Venter Institute). There is a clear difference between the intellectual goals of these two fields. The ‘deconstruction’ community seeks to understand biological systems and their evolution, whereas the ‘construction’ community searches for general design principles regardless of their relationship to actual Biology. Both communities, however, hope that the exploration and/or construction of these (biological) systems will expand our understanding of the organizational principles of living molecular systems, and both are linked by their dependence on very similar theoretical, experimental and computational techniques. This new drive towards ‘construction’ and ‘deconstruction’ of biological phenomena gives a new dimension and adds a new value to traditional research into the origin of life on Earth. Such research, now under the umbrella of Synthetic Biology, was for decades restricted to the field of the Chemistry of prebiotic systems. The body of theories and simulations on primitive, pre-cellular metabolism (Eric Smith, Santa Fe Institute) and on the transitions between living and non-living systems (Steen Rasmussen, Los Alamos National Laboratory; Norman Packard, Protolife Srl) provides a wealth of conceptual and material assets that the new field will be able to build on. Both the ‘deconstruction’ and the ‘construction’ approach generate extremely interesting scientific and technical challenges, among which there are at least two that are likely to influence the future of Computational Biology and Bioinformatics. The first issue is how to synthesize chromosomes containing well-defined functional regions (genes) under clear controllable replication, transcription and translation conditions. This fascinating prospect will require all our skills in analyzing, predicting and designing genome elements, with many implications for genome analysis, comparative genomics and transcription regulation. The second key issue is the modulation of functional specificity. Both newly designed components and those extracted from biological systems require a clear understanding of how they adapt to specific working conditions and how the interactions that determine the properties of stability and adaptation of molecular systems could be engineered. A clear example is the design of Transcription factors able to trigger the activation of specific genes, whether these are designed rationally (Homme Hellinga, Duke University) or obtained by directed evolution (Víctor de Lorenzo, National Center of Biotechnology). Computational analysis of protein families and of protein/DNA structures is, of course, essential for the development of these components. Apart of the instrumental role of Bioinformatics in Synthetic Biology in the design of synthetic chromosomes and components of a desired specificity, there are interesting additional opportunities as well for the development of computational models to analyze, simulate and predict the behavior of artificial and synthetic systems, in what is a new growing field.
Víctor de Lorenzo, Luis Serrano, Alfonso Valencia
Bioinform.3
2006 Phylogeny-independent detection of functional residues
abstract
MOTIVATION: Current projects for the massive characterization of proteomes are generating protein sequences and structures with unknown function. The difficulty of experimentally determining functionally important sites calls for the development of computational methods. The first techniques, based on the search for fully conserved positions in multiple sequence alignments (MSAs), were followed by methods for locating family-dependent conserved positions. These rely on the functional classification implicit in the alignment for locating these positions related with functional specificity. The next obvious step, still scarcely explored, is to detect these positions using a functional classification different from the one implicit in the sequence relationships between the proteins. Here, we present two new methods for locating functional positions which can incorporate an arbitrary external functional classification which may or may not coincide with the one implicit in the MSA. The Xdet method is able to use a functional classification with an associated hierarchy or similarity between functions to locate positions related to that classification. The MCdet method uses multivariate statistical analysis to locate positions responsible for each one of the functions within a multifunctional family. RESULTS: We applied the methods to different cases, illustrating scenarios where there is a disagreement between the functional and the phylogenetic relationships, and demonstrated their usefulness for the phylogeny-independent prediction of functional positions.
Florencio Pazos, Antonio Rausell, Alfonso Valencia
Bioinform.3
2006 Software patents in Bioinformatics
abstract
Bioinformatics has published papers describing new software for over 20 years (Nilsson and Klein 1985). During this time the world of software has changed considerably particularly with the irresistible rise of initiatives to build freely accessible software as has opening access to data resources. The Internet and the Web have also changed the way we use and distribute software. This social and technical revolution is also changing the structure of the relations between commercial and academic software-based activities, for which patents and software protection are key elements. In this and the following issue we publish two editorials addressing the general topics of software accessibility, patents and intellectual property. In this issue, Steven L. Salzberg and John Quackenbush (past and present Associate Editors, respectively) present one perspective on the issue. In the next issue another of our Associate Editors, Jonathan D. Wren will put forward a different perspective. We welcome additional contributions to this discussion from our readership, which will help the journal in the process of adapting our publication guidelines to better serve the development of Bioinformatics.
Alfonso Valencia, Alex Bateman
Bioinform.1
2006 An analysis of the Sargasso Sea resource and the consequences for database composition
abstract
BACKGROUND: The environmental sequencing of the Sargasso Sea has introduced a huge new resource of genomic information. Unlike the protein sequences held in the current searchable databases, the Sargasso Sea sequences originate from a single marine environment and have been sequenced from species that are not easily obtainable by laboratory cultivation. The resource also contains very many fragments of whole protein sequences, a side effect of the shotgun sequencing method.These sequences form a significant addendum to the current searchable databases but also present us with some intrinsic difficulties. While it is important to know whether it is possible to assign function to these sequences with the current methods and whether they will increase our capacity to explore sequence space, it is also interesting to know how current bioinformatics techniques will deal with the new sequences in the resource. RESULTS: The Sargasso Sea sequences seem to introduce a bias that decreases the potential of current methods to propose structure and function for new proteins. In particular the high proportion of sequence fragments in the resource seems to result in poor quality multiple alignments. CONCLUSION: These observations suggest that the new sequences should be used with care, especially if the information is to be used in large scale analyses. On a positive note, the results may just spark improvements in computational and experimental methods to take into account the fragments generated by environmental sequencing techniques.
Michael L. Tress, Domenico Cozzetto, Anna Tramontano, Alfonso Valencia
BMC Bioinform.4
2005 An update from the Bioinformatics Editors
abstract
In this editorial we take the opportunity to highlight changes in the journal Bioinformatics, during 2005. During the past couple of years the journal had accepted more papers than it could publish and this resulted in a backlog of manuscripts waiting some months to appear in print. We are very pleased to say that this has been relieved by publishing four large issues in April and May of this year. It now takes only 10 weeks from acceptance until publication in the print issue; manuscripts are also rapidly published online ahead of print (on the ‘Advance Access’ page) within 4 days of acceptance on average. We have now taken steps to carefully monitor the acceptance levels in the journal to improve quality even further and to ensure that we do not increase the acceptance to print times in the future. The current acceptance rate is 25%. We have also improved our review process so that 80% of submissions receive a final editorial decision in 40 days. During 2005 we have published special supplement issues of the journal containing the proceedings of both ISMB and ECCB conferences. The journal also occasionally publishes special sections where a small number of papers from a conference are included in a regular issue. Currently these are submitted on an ad hoc basis by conference organisers. We wish now to regularise this process and hereby request proposals for conference proceeding papers for publication in 2006. The deadline for these proposals is 30th January 2006 and we also welcome preliminary proposals for conferences in 2007. Details about the type of information we will require about conference proposals can be found in the journal's Instructions to Authors. It is expected that the conference papers put forward for publication in the journal will be peer-reviewed by the organisers, liaising with a Bioinformatics Associate Editor, to ensure a high standard. Since July 2005, and following the successful experience of our sister journal Nucleic Acids Research, we have launched a new Open Access option. Authors can now choose whether or not to publish their work ‘open access’. For more information about the Oxford Open initiative visit Author Webpage. If an author does not choose the Open Access option their paper will be made freely available online twelve months following publication. We are convinced that this important decision will best satisfy the desires of our authors, as expressed in our author survey (read about the results of this survey in Bioinformatics, 22, 4071–4072). During 2005, the following new Associate Editors have joined Bioinformatics: Keith Crandall, Joaquin Dopazo, Dmitrij Frishman, Chris Stoeckert and Anna Tramontano. More recently we welcome on board the following new Associate Editors: David Rocke, an expert in statistics with particular dedication to DNA array analysis, Jonathan Wren, a bioinformatician now working in areas related with information extraction and text mining, and Golan Yona, a computer scientist particularly interested in the analysis of networks. Additionally during the year the following scientists have joined the journal's Editorial Board: former Associate Editors Russ Altman, Carlos D Bustamante, Gary Williams and Michael Zhang, and Mikhail Gelfand, Adam Godzik, Jaap Heringa, Ina Koch, Wentian Li, Isidore Rigoutsos, Burkhard Rost, N Srinivasan, Olga Troyanskaya and Limsoon Wong. The Associate Editors are responsible for arranging the peer review of submissions, making editorial decisions and working together to decide the future direction and policies of the journal. To increase the transparency of the editorial process, and following suggestions from our readers and authors, we have recently decided to make known the name of the Associate Editor responsible for each manuscript, and to include their names on appropriate papers in the published version. The Editorial Board members also play an active role in Bioinformatics, acting as the main consulting body for journal policies, aims and scope, and helping with difficult editorial decisions. We would like to thank those who are stepping down in 2005: Associate Editors Phil Bourne, Frank Dudbridge, Steen Knudsen, and Steve Salzberg, and Editorial Board members Terry Gaasterland, Mike Gribskov, Steven Henikoff, Webb Miller, and Eugene Myers. Without their dedication and hard work it would not be possible to produce this journal.
Alex Bateman, Alfonso Valencia
Bioinform.2
2005 Do you do text?
abstract
Retrieving information from text has become an important area in bioinformatics, and not too surprisingly, this journal has published more than 30 papers on this topic since the first article published by the journal in 1998. In addition, ISCB (International Society for Computational Biology) (Author Webpage) has organized special sessions in the ISMB conferences (Intelligent Systems for Molecular Biology, see Author Webpage) and a specialized interest group (Author Webpage) for the last five years. In parallel, major computer science conferences in related areas have begun to include sessions on biology, e.g. the TREC (text retrieval conference) Genomics track (medir.ohsu.edu/∼genomics), ICML (International Conference on Machine Learning) and a series of workshops organized in association with the ACL (Association for Computation Linguistics) and Human Language Technology meetings. The requirements of the text mining community are similar to those in other areas of bioinformatics: Availability of high quality input information A set of objective metrics for the comparison of different methods The need to involve the biologist to keep the focus on developing applications that are suitable for end users and biological databases. Exactly the same issues have been discussed, and partially solved, in the field of protein structure prediction, in part thanks to the organization of CASP (Critical Assessment of techniques for protein Structure Prediction) (Author Webpage) during the last 10 years. With similar evaluation goals in mind, we organized BioCreAtIvE (critical assessment of information extraction in biology), focusing on two tasks. The first dealt with extraction and normalization of gene or protein names from text for three model organism databases (fly, mouse, yeast). The second task addressed issues of extracting functional annotations from text. Overall, 27 groups participated in the assessment (see Author Webpage and Author Webpage). The results and the assessment were discussed in a meeting sponsored by EMBO (European Molecular Biology Organization), in Granada, Spain (Author Webpage). The results for gene/protein name extraction showed that at least four groups provided systems that were able to extract gene names from sentences of MEDLINE abstracts at over 80% balanced precision and recall. For the subtask of recognizing the relation between names and normalized database identifiers, the results ranged from a maximum of 92% balanced precision and recall for yeast to 79% for mouse. These results indicate that this technology may now be mature enough to be used in production environments (e.g. document retrieval). However, the results for gene/protein names lag behind those obtained for identifying persons and locations for online news (90–95%). The identification of the many other entities of interest in biology (chemical compounds, tissues, diseases, species and others) will involve additional challenges. For the functional annotation task, systems were asked to identify a segment of text as evidence for a GO (Gene Ontology) annotation for a given protein in full text articles. In this case, participants were not given training examples of identified text segments. Annotations and text evidence were reviewed by expert annotators from the GO annotation team (Author Webpage) for validity. When both the protein name and the GO annotation were given, several systems provided correct evidence for the GO predictions 25–30% of the time. The average performances were lower in a subtask in which the GO codes were not given. Interestingly, two systems provided a higher rate of correct predictions by focusing on high confidence cases. These results indicate that the retrieval of functional information is a challenging problem. We believe, however, that this first BioCreAtIvE assessment has laid the foundation for rapid progress in this area by providing an infrastructure, particularly training and test datasets, which will encourage researchers to test their systems against these datasets. A technical description of the results can be found at Author Webpage, and the full collection of methods and evaluation papers has been recently published (BMC Bioinformatics, 2005; 6 Suppl. 1). Indeed, BioCreAtIvE also delivered the associated collection of annotated data provided by the organizers and the corresponding evaluations of results from the participating groups (Author Webpage). This collection will complement other datasets such as the GENIA corpus (Author Webpage) as a valuable resource for training and testing methods. Our main conclusions from the meeting are that: A number of groups achieved similar levels of performance, using a variety of technologies. These ranged from those based on natural language processing and linguistic analysis to machine learning and computational biology approaches used in protein structure prediction, gene finding and the like. An unbiased assessment, based on clear standards and objective evaluation, has provided a more realistic view of the state of the art than has been available to date from more limited evaluations reported in the literature. Setting up the assessment required a considerable effort in the preparation of the datasets, and the evaluation of the results required a considerable effort by human experts. In spite of this, the size and the quality of the available datasets are the main limitation of BioCreAtIvE and other assessment initiatives. Indeed, a number of groups representing what could be called traditional bioinformatics are experiencing considerable success in the field and we encourage you to consider exploring this new field: have you tried your best in text mining? Data and additional information are available at Author Webpage The contribution of Alex Morgan, John Wilbur, Lorrie Tanabe and Vivian Lee was essential for the organization and evaluation of BioCreAtIvE. Essential, as well, was the participation of the 27 groups, and their input to the organization, evaluation and discussion of the results. The datasets for the first task were provided by NCBI (National Center for Biotechnology Information) and MITRE and the datasets for the second by EBI–EMBL (European Bioinformatics Institute–European Molecular Biology Laboratory). The MITRE contributions to BioCreAtIvE were supported in part by NSF (grant EIA-0326404), and those to CNB-CSIC and EBI were supported by the European Commission (grants TEMBLOR QLRT-2001-00015 and ORIEL IST-2001-32688).
Christian Blaschke, Alexander S. Yeh, Evelyn Camon, Marc E. Colosimo, Rolf Apweiler, Lynette Hirschman, Alfonso Valencia
Bioinform.7
2005 Increasing the Impact of Bioinformatics
abstract
The year 2004 has been very successful for Bioinformatics. The journal's latest impact factor from the Institute for Scientific Information has increased from 4.615 to 6.701. This is quite an exceptional increase reflecting the increasing standard of work in the journal as well as the increasing stature of the field. Early in 2004, we implemented a new system for Advance online access which allows scientists to access research in Bioinformatics as rapidly as possible. The number of manuscripts submitted to the journal continues to grow. In 2003 we received 1300 submissions while in 2004 we received over 1800 submissions. That we have coped with this large increase is testament to the dedication and hard work of our team of Associate Editors, referees and the Editorial Office. To cope with future growth we are adapting our editorial structure and processes. From 2005 the journal will appear 24 times per year rather than the 18 issues per year previously. This will allow us to publish more of the high quality research and applications that are being submitted to us each day. However, the increase in submissions has outstripped the increase in journal pages. This means that our acceptance rate is decreasing to ∼20%. We are in the process of appointing new Associate Editors to replace those who have stepped down and to increase our coverage of new areas. We would like to thank Gert Vriend, Debbie Marks and Fritz Roth for their invaluable contribution to the journal. We are also in the process of expanding the membership of our Editorial Board. In 2005, we will be introducing a new scheme of categories for papers. During the submission process authors will be asked to choose which category their paper belongs to. This will improve the assignment of manuscripts to editors as well as helping us to formulate a clear definition of the scope of the journal within each category, thus improving organization, in terms of layout and editorial process. Open Access is a topic that is very important to many of our authors and readers. Along this line, we as Editors, and Oxford University Press as a not-for-profit academic publisher, are very much in favour of the principle of making scientific publications freely available. For a well-established journal, as Bioinformatics is now, it is also important to preserve the reputation and financial security of the journal to which authors, referees, editors and readers have contributed over the many years. During 2005 we will learn a great deal from the experience of our sister journal, Nucleic Acids Research, which is introducing a full Open Access model. We are also seeking the opinions of our authors and readers through a survey exploring publication models. We will be sure to base our decision on whether to move forward with an Open Access initiative on the response from our readers, authors and their institutions. So please let us know your views. If we are encouraged by the journal's community to experiment with Open Access, and with a carefully studied new business model in place, we see Open Access as a real opportunity for the journal in the near future. This significant collection of changes will make 2005 an important year for the journal. Our new cover design reflects this spirit of change which will make our journal, and the bioinformatics field, more open to science.
Alfonso Valencia, Alex Bateman
Bioinform.1
2005 Evaluation of BioCreAtIvE assessment of task 2
abstract
BACKGROUND: Molecular Biology accumulated substantial amounts of data concerning functions of genes and proteins. Information relating to functional descriptions is generally extracted manually from textual data and stored in biological databases to build up annotations for large collections of gene products. Those annotation databases are crucial for the interpretation of large scale analysis approaches using bioinformatics or experimental techniques. Due to the growing accumulation of functional descriptions in biomedical literature the need for text mining tools to facilitate the extraction of such annotations is urgent. In order to make text mining tools useable in real world scenarios, for instance to assist database curators during annotation of protein function, comparisons and evaluations of different approaches on full text articles are needed. RESULTS: The Critical Assessment for Information Extraction in Biology (BioCreAtIvE) contest consists of a community wide competition aiming to evaluate different strategies for text mining tools, as applied to biomedical literature. We report on task two which addressed the automatic extraction and assignment of Gene Ontology (GO) annotations of human proteins, using full text articles. The predictions of task 2 are based on triplets of protein--GO term--article passage. The annotation-relevant text passages were returned by the participants and evaluated by expert curators of the GO annotation (GOA) team at the European Institute of Bioinformatics (EBI). Each participant could submit up to three results for each sub-task comprising task 2. In total more than 15,000 individual results were provided by the participants. The curators evaluated in addition to the annotation itself, whether the protein and the GO term were correctly predicted and traceable through the submitted text fragment. CONCLUSION: Concepts provided by GO are currently the most extended set of terms used for annotating gene products, thus they were explored to assess how effectively text mining tools are able to extract those annotations automatically. Although the obtained results are promising, they are still far from reaching the required performance demanded by real world applications. Among the principal difficulties encountered to address the proposed task, were the complex nature of the GO terms and protein names (the large range of variants which are used to express proteins and especially GO terms in free text), and the lack of a standard training set. A range of very different strategies were used to tackle this task. The dataset generated in line with the BioCreative challenge is publicly available and will allow new possibilities for training information extraction methods in the domain of molecular biology.
Christian Blaschke, Eduardo Andrés León, Martin Krallinger, Alfonso Valencia
BMC Bioinform.4
2005 Overview of BioCreAtIvE: critical assessment of information extraction for biology
abstract
BACKGROUND: The goal of the first BioCreAtIvE challenge (Critical Assessment of Information Extraction in Biology) was to provide a set of common evaluation tasks to assess the state of the art for text mining applied to biological problems. The results were presented in a workshop held in Granada, Spain March 28-31, 2004. The articles collected in this BMC Bioinformatics supplement entitled "A critical assessment of text mining methods in molecular biology" describe the BioCreAtIvE tasks, systems, results and their independent evaluation. RESULTS: BioCreAtIvE focused on two tasks. The first dealt with extraction of gene or protein names from text, and their mapping into standardized gene identifiers for three model organism databases (fly, mouse, yeast). The second task addressed issues of functional annotation, requiring systems to identify specific text passages that supported Gene Ontology annotations for specific proteins, given full text articles. CONCLUSION: The first BioCreAtIvE assessment achieved a high level of international participation (27 groups from 10 countries). The assessment provided state-of-the-art performance results for a basic task (gene name finding and normalization), where the best systems achieved a balanced 80% precision / recall or better, which potentially makes them suitable for real applications in biology. The results for the advanced task (functional annotation from free text) were significantly lower, demonstrating the current limitations of text-mining approaches where knowledge extrapolation and interpretation are required. In addition, an important contribution of BioCreAtIvE has been the creation and release of training and test data sets for both tasks. There are 22 articles in this special issue, including six that provide analyses of results or data quality for the data sets, including a novel inter-annotator consistency assessment for the test set used in task 2.
Lynette Hirschman, Alexander S. Yeh, Christian Blaschke, Alfonso Valencia
BMC Bioinform.4
2005 A sentence sliding window approach to extract protein annotations from biomedical articles
abstract
BACKGROUND: Within the emerging field of text mining and statistical natural language processing (NLP) applied to biomedical articles, a broad variety of techniques have been developed during the past years. Nevertheless, there is still a great ned of comparative assessment of the performance of the proposed methods and the development of common evaluation criteria. This issue was addressed by the Critical Assessment of Text Mining Methods in Molecular Biology (BioCreative) contest. The aim of this contest was to assess the performance of text mining systems applied to biomedical texts including tools which recognize named entities such as genes and proteins, and tools which automatically extract protein annotations. RESULTS: The "sentence sliding window" approach proposed here was found to efficiently extract text fragments from full text articles containing annotations on proteins, providing the highest number of correctly predicted annotations. Moreover, the number of correct extractions of individual entities (i.e. proteins and GO terms) involved in the relationships used for the annotations was significantly higher than the correct extractions of the complete annotations (protein-function relations). CONCLUSION: We explored the use of averaging sentence sliding windows for information extraction, especially in a context where conventional training data is unavailable. The combination of our approach with more refined statistical estimators and machine learning techniques might be a way to improve annotation extraction for future biomedical text mining applications.
Martin Krallinger, Maria Padron, Alfonso Valencia
BMC Bioinform.3
2004 New Leadership for Bioinformatics
Alex Bateman, Alfonso Valencia
Bioinform.2
2004 YAdumper: extracting and translating large information volumes from relational databases to structured flat files
abstract
Downloading the information stored in relational databases into XML and other flat formats is a common task in bioinformatics. This periodical dumping of information requires considerable CPU time, disk and memory resources. YAdumper has been developed as a purpose-specific tool to deal with the integral structured information download of relational databases. YAdumper is a Java application that organizes database extraction following an XML template based on an external Document Type Declaration. Compared with other non-native alternatives, YAdumper substantially reduces memory requirements and considerably improves writing performance.
José María Fernández 0001, Alfonso Valencia
Bioinform.2
2004 SQUARE-determining reliable regions in sequence alignments
abstract
The Server for Quick Alignment Reliability Evaluation (SQUARE) is a Web-based version of the method we developed to predict regions of reliably aligned residues in sequence alignments. Given an alignment between a query sequence and a sequence of known structure, SQUARE is able to predict which residues are reliably aligned. The server accesses a database of profiles of sequences of known three-dimensional structures in order to calculate the scores for each residue in the alignment. SQUARE produces a graphical output of the residue profile-derived alignment scores along with an indication of the reliability of the alignment. In addition, the scores can be compared against template secondary structure, conserved residues and important sites.
Michael L. Tress, Osvaldo Graña, Alfonso Valencia
Bioinform.3
2004 SPOC: A widely distributed domain associated with cancer, apoptosis and transcription
abstract
BACKGROUND: The Split ends (Spen) family are large proteins characterised by N-terminal RNA recognition motifs (RRMs) and a conserved SPOC (Spen paralog and ortholog C-terminal) domain. The aim of this study is to characterize the family at the sequence level. RESULTS: We describe undetected members of the Spen family in other lineages (Plasmodium and Plants) and localise SPOC in a new domain context, in a family that is common to all eukaryotes using profile-based sequence searches and structural prediction methods. CONCLUSIONS: The widely distributed DIO (Death inducer-obliterator) family is related to cancer and apoptosis and offers new clues about SPOC domain functionality.
Luis Sánchez-Pulido, Ana María Rojas, Karel H. M. van Wely, Carlos Martínez-A, Alfonso Valencia
BMC Bioinform.5
2003 Evaluation of annotation strategies using an entire genome sequence
abstract
Abstract Motivation: Genome-wide functional annotation either by manual or automatic means has raised considerable concerns regarding the accuracy of assignments and the reproducibility of methodologies. In addition, a performance evaluation of automated systems that attempt to tackle sequence analyses rapidly and reproducibly is generally missing. In order to quantify the accuracy and reproducibility of function assignments on a genome-wide scale, we have re-annotated the entire genome sequence of Chlamydia trachomatis (serovar D), in a collaborative manner. Results: We have encoded all annotations in a structured format to allow further comparison and data exchange and have used a scale that records the different levels of potential annotation errors according to their propensity to propagate in the database due to transitive function assignments. We conclude that genome annotation may entail a considerable amount of errors, ranging from simple typographical errors to complex sequence analysis problems. The most surprising result of this comparative study is that automatic systems might perform as well as the teams of experts annotating genome sequences. Availability and supplementary information: http://www.ebi.ac.uk/research/cgg/annotation/cteval/ Contact: [email protected] * To whom correspondence should be addressed. † INA-EKETA, GR-57001 Thessaloniki, Greece ‡ Computational Biology Center, Memorial Sloan-Kettering Cancer Center, New York, NY 10021, USA § Aetion Technologies LLC, Worthington, OH 43085, USA ¶ Institut Curie, F-75248 Paris, France ∥ CNRS, UMR6543, F-06108 Nice, France ** Alma Bioinformatics, E-28760 Madrid, Spain †† Cap Gemini Ernst & Young, London SW1X 7LX, UK ‡‡ Univ. of Rome ‘La Sapienza’, I-00185 Rome, Italy §§ MWG-Biotech AG, Ebersberg, D-85560 Berlin, Germany ¶¶ Wellcome Trust Biocentre, Univ. of Dundee, Dundee DD1 5HN, UK
Sophia Tsoka, Miguel A. Andrade-Navarro, Anton J. Enright, Mark Carroll, Patrick Poullet, Vasilis J. Promponas, Theodore Liakopoulos, Giorgos Palaios, Claude Pasquier, Stavros J. Hamodrakas, Javier Tamames, Asutosh T. Yagnik, Anna Tramontano, Damien Devos, Christian Blaschke, Alfonso Valencia, David Brett, David M. A. Martin, Christophe Leroy, Isidore Rigoutsos, Chris Sander, Christos A. Ouzounis
Bioinform.17
2003 Early bioinformatics: the birth of a discipline - a personal view
abstract
MOTIVATION: The field of bioinformatics has experienced an explosive growth in the last decade, yet this 'new' field has a long history. Some historical perspectives have been previously provided by the founders of this field. Here, we take the opportunity to review the early stages and follow developments of this discipline from a personal perspective. RESULTS: We review the early days of algorithmic questions and answers in biology, the theoretical foundations of bioinformatics, the development of algorithms and database resources and finally provide a realistic picture of what the field looked like from a resources and finally provide a realistic picture of what the field looked like from a practitioner's viewpoint 10 years ago, with a perspective for future developments.
Christos A. Ouzounis, Alfonso Valencia
Bioinform.2
2003 Meta, Metan and Cyber Servers
Alfonso Valencia
Bioinform.1
2002 Information extraction in molecular biology
abstract
Information extraction has become a very active field in bioinformatics recently and a number of interesting papers have been published. Most of the efforts have been concentrated on a few specific problems, such as the detection of protein-protein interactions and the analysis of DNA expression arrays, although it is obvious that there are many other interesting areas of potential application (document retrieval, protein functional description, and detection of disease-related genes to name a few). Paradoxically, these exciting developments have not yet crystallised into general agreement on a set of standard evaluation criteria, such as the ones developed in fields such as protein structure prediction, which makes it very difficult to compare performance across these different systems. In this review we introduce the general field of information extraction, we outline the status of the applications in molecular biology, and we then discuss some ideas about possible standards for evaluation that are needed for the future development of the field.
Christian Blaschke, Lynette Hirschman, Alfonso Valencia
Briefings Bioinform.3
2002 Clustering of proximal sequence space for the identification of protein families
abstract
MOTIVATION: The study of sequence space, and the deciphering of the structure of protein families and subfamilies, has up to now been required for work in comparative genomics and for the prediction of protein function. With the emergence of structural proteomics projects, it is becoming increasingly important to be able to select protein targets for structural studies that will appropriately cover the space of protein sequences, functions and genomic distribution. These problems are the motivation for the development of methods for clustering protein sequences and building families of potentially orthologous sequences, such as those proposed here. RESULTS: First we developed a clustering strategy (Ncut algorithm) capable of forming groups of related sequences by assessing their pairwise relationships. The results presented for the ras super-family of proteins are similar to those produced by other clustering methods, but without the need for clustering the full sequence space. The Ncut clusters are then used as the input to a process of reconstruction of groups with equilibrated genomic composition formed by closely-related sequences. The results of applying this technique to the data set used in the construction of the COG database are very similar to those derived by the human experts responsible for this database. AVAILABILITY: The analysis of different systems, including the COG equivalent 21 genomes are available at http://www.pdg.cnb.uam.es/GenoClustering.html.
Federico Abascal, Alfonso Valencia
Bioinform.2
2002 Bioinformatics in structural genomics - Editorial
abstract
Burkhard Rost, Barry Honig, Alfonso Valencia; Bioinformatics in structural genomics, Bioinformatics, Volume 18, Issue 7, 1 July 2002, Pages 897, https://doi.org
Burkhard Rost, Barry Honig, Alfonso Valencia
Bioinform.3
2002 Bioinformatics: Biology by other means
Alfonso Valencia
Bioinform.1
2001 EVA: continuous automatic evaluation of protein structure prediction servers
abstract
UNLABELLED: Evaluation of protein structure prediction methods is difficult and time-consuming. Here, we describe EVA, a web server for assessing protein structure prediction methods, in an automated, continuous and large-scale fashion. Currently, EVA evaluates the performance of a variety of prediction methods available through the internet. Every week, the sequences of the latest experimentally determined protein structures are sent to prediction servers, results are collected, performance is evaluated, and a summary is published on the web. EVA has so far collected data for more than 3000 protein chains. These results may provide valuable insight to both developers and users of prediction methods. AVAILABILITY: http://cubic.bioc.columbia.edu/eva. CONTACT: [email protected]
Volker A. Eyrich, Marc A. Martí-Renom, Dariusz Przybylski, M. S. Madhusudhan 0001, András Fiser, Florencio Pazos, Alfonso Valencia, Andrej Sali, Burkhard Rost
Bioinform.7
2001 A hierarchical unsupervised growing neural network for clustering gene expression patterns
abstract
MOTIVATION: We describe a new approach to the analysis of gene expression data coming from DNA array experiments, using an unsupervised neural network. DNA array technologies allow monitoring thousands of genes rapidly and efficiently. One of the interests of these studies is the search for correlated gene expression patterns, and this is usually achieved by clustering them. The Self-Organising Tree Algorithm, (SOTA) (Dopazo,J. and Carazo,J.M. (1997) J. Mol. Evol., 44, 226-233), is a neural network that grows adopting the topology of a binary tree. The result of the algorithm is a hierarchical cluster obtained with the accuracy and robustness of a neural network. RESULTS: SOTA clustering confers several advantages over classical hierarchical clustering methods. SOTA is a divisive method: the clustering process is performed from top to bottom, i.e. the highest hierarchical levels are resolved before going to the details of the lowest levels. The growing can be stopped at the desired hierarchical level. Moreover, a criterion to stop the growing of the tree, based on the approximate distribution of probability obtained by randomisation of the original data set, is provided. By means of this criterion, a statistical support for the definition of clusters is proposed. In addition, obtaining average gene expression patterns is a built-in feature of the algorithm. Different neurons defining the different hierarchical levels represent the averages of the gene expression patterns contained in the clusters. Since SOTA runtimes are approximately linear with the number of items to be classified, it is especially suitable for dealing with huge amounts of data. The method proposed is very general and applies to any data providing that they can be coded as a series of numbers and that a computable measure of similarity between data items can be used. AVAILABILITY: A server running the program can be found at: http://bioinfo.cnio.es/sotarray.
Javier Herrero, Alfonso Valencia, Joaquín Dopazo
Bioinform.2
1999 Automatic Extraction of Biological Information from Scientific Text: Protein-Protein Interactions
Christian Blaschke, Miguel A. Andrade-Navarro, Christos A. Ouzounis, Alfonso Valencia
ISMB4
1999 Automated genome sequence analysis and annotation
abstract
MOTIVATION: Large-scale genome projects generate a rapidly increasing number of sequences, most of them biochemically uncharacterized. Research in bioinformatics contributes to the development of methods for the computational characterization of these sequences. However, the installation and application of these methods require experience and are time consuming. RESULTS: We present here an automatic system for preliminary functional annotation of protein sequences that has been applied to the analysis of sets of sequences from complete genomes, both to refine overall performance and to make new discoveries comparable to those made by human experts. The GeneQuiz system includes a Web-based browser that allows examination of the evidence leading to an automatic annotation and offers additional information, views of the results, and links to biological databases that complement the automatic analysis. System structure and operating principles concerning the use of multiple sequence databases, underlying sequence analysis tools, lexical analyses of database annotations and decision criteria for functional assignments are detailed. The system makes automatic quality assessments of results based on prior experience with the underlying sequence analysis tools; overall error rates in functional assignment are estimated at 2.5-5% for cases annotated with highest reliability ('clear' cases). Sources of over-interpretation of results are discussed with proposals for improvement. A conservative definition for reporting 'new findings' that takes account of database maturity is presented along with examples of possible kinds of discoveries (new function, family and superfamily) made by the system. System performance in relation to sequence database coverage, database dynamics and database search methods is analysed, demonstrating the inherent advantages of an integrated automatic approach using multiple databases and search methods applied in an objective and repeatable manner. AVAILABILITY: The GeneQuiz system is publicly available for analysis of protein sequences through a Web server at http://www.sander.ebi.ac. uk/gqsrv/submit
Miguel A. Andrade-Navarro, Nigel P. Brown, Christophe Leroy, S. Hörsch, Antoine de Daruvar, C. Reich, Angelo Franchini, Javier Tamames, Alfonso Valencia, Christos A. Ouzounis, Chris Sander
Bioinform.9
1999 A platform for integrating threading results with protein family analyses
abstract
Abstract Summary: We have developed a package for the interactive visualization of results from different threading programs. Additionally, we have integrated relevant information about protein sequence, function, evolution, and structure into the interface. Availability: A detailed documentation of THREADLIZE, and the binaries for IRIX, SunOS and Linux are available at http://www.cnb.uam.es/~pazos/threadlize. The package is free for academic users. Contact: [email protected] Supplementary information: http://www.cnb.uam.es/~pazos/threadlize
Florencio Pazos, Burkhard Rost, Alfonso Valencia
Bioinform.3
1998 Automatic extraction of keywords from scientific text: application to the knowledge domain of protein families
abstract
MOTIVATION: Annotation of the biological function of different protein sequences is a time-consuming process currently performed by human experts. Genome analysis tools encounter great difficulty in performing this task. Database curators, developers of genome analysis tools and biologists in general could benefit from access to tools able to suggest functional annotations and facilitate access to functional information. APPROACH: We present here the first prototype of a system for the automatic annotation of protein function. The system is triggered by collections of s related to a given protein, and it is able to extract biological information directly from scientific literature, i.e. MEDLINE abstracts. Relevant keywords are selected by their relative accumulation in comparison with a domain-specific background distribution. Simultaneously, the most representative sentences and MEDLINE abstracts are selected and presented to the end-user. Evolutionary information is considered as a predominant characteristic in the domain of protein function. Our system consequently extracts domain-specific information from the analysis of a set of protein families. RESULTS: The system has been tested with different protein families, of which three examples are discussed in detail here: 'ataxia-telangiectasia associated protein', 'ran GTPase' and 'carbonic anhydrase'. We found generally good correlation between the amount of information provided to the system and the quality of the annotations. Finally, the current limitations and future developments of the system are discussed. AVAILABILITY: The current system can be considered as a prototype system. As such, it can be accessed as a server at http://columba.ebi.ac. uk:8765/andrade/abx. The system accepts text related to the protein or proteins to be evaluated (optimally, the result of a MEDLINE search by keyword) and the results are returned in the form of Web pages for keywords, sentences and s. SUPPLEMENTARY INFORMATION: Web pages containing full information on the examples mentioned in the text are available at: http://www.cnb.uam.es/ approximately cnbprot/keywords/ CONTACT: [email protected]
Miguel A. Andrade-Navarro, Alfonso Valencia
Bioinform.2
1998 EUCLID: automatic classification of proteins in functional classes by their database annotations
abstract
UNLABELLED: A tool is described for the automatic classification of sequences in functional classes using their database annotations. The Euclid system is based on a simple learning procedure from examples provided by human experts. AVAILABILITY: Euclid is freely available for academics at http://www.gredos.cnb.uam.es/EUCLID, with the corresponding dictionaries for the generation of three, eight and 14 functional classes. CONTACT: E-mail: [email protected] SUPPLEMENTARY INFORMATION: The results of the EUCLID classification of different genomes are available at http://www.sander.ebi.ac. uk/genequiz/. A detailed description of the different applications mentioned in the text is available at http://www.gredos.cnb.uam. es/EUCLID/Full_Paper
Javier Tamames, Christos A. Ouzounis, Georg Casari, Chris Sander, Alfonso Valencia
Bioinform.5
1998 Computational space reduction and parallelization of a new clustering approach for large groups of sequences
abstract
MOTIVATION: The explosive growth of the biological sequences databases stimulated by genome projects has modified the framework of several applications in the biological sequence analysis area. In most cases, this new scenario is characterized by studies on large sets of sequences, suggesting the need for effective and automatic methods for their clustering. A more effective clustering of the database could be followed by the application of common family analysis schemes to the groups so formed. RESULTS: In this work, we present a new strategy to reduce the computational cost associated with the clustering of large sets of sequences which are expected to contain several families. The strategy is based on the grouping of the sequences into families by using a dynamic threshold on a pairwise sequence similarity criterion. Routine clustering of large data sets can now be done very efficiently. The method developed here achieves a computational space reduction of about an order of magnitude over more traditional ones of all-versus-all comparisons. The outcome of this approach produces family groupings that reproduce closely already accepted biological results. Our work includes a parallel implementation for distributed memory multiprocessors with a dynamic scheduling strategy for performance optimization. AVAILABILITY: By anonymous ftp at ftp.ac.uma.es (/pub/ots/pCluster directory), or from our Web site http://www.cnb. uam.es/www/software/software_index.html CONTACT: [email protected]
Oswaldo Trelles, Miguel A. Andrade-Navarro, Alfonso Valencia, Emilio L. Zapata, José María Carazo
Bioinform.3
1997 Automatic Annotation for Biological Sequences by Etraction of Keywords from MEDLINE Abstracts: Development of a Prototype System
Miguel A. Andrade-Navarro, Alfonso Valencia
ISMB2
1997 Sequence analysis of the Methanococcus jannaschii genome and the prediction of protein function
abstract
Miguel Andrade, Georg Casari, Antoine de Daruvar, Chris Sander, Reinhard Schneider, Javier Tamames, Alfonso Valencia, Christos Ouzounis; Sequence analysis
Miguel A. Andrade-Navarro, Georg Casari, Antoine de Daruvar, Chris Sander, Reinhard Schneider 0002, Javier Tamames, Alfonso Valencia, Christos A. Ouzounis
Comput. Appl. Biosci.7
1997 A graphical interface for correlated mutations and other protein structure prediction methods
Florencio Pazos, O. Olmea, Alfonso Valencia
Comput. Appl. Biosci.3
1994 GeneQuiz: A Workbench for Sequence Analysis
Michael Scharf, Reinhard Schneider 0002, Georg Casari, Peer Bork, Alfonso Valencia, Christos A. Ouzounis, Chris Sander
ISMB5