Henning Hermjakob

dblp:71/4115 · DBLP profile ↗
← Back
45ranked-venue papers
2as first author
10since 2021 · last 2025
0000-0001-8479-0262ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 45 · 2 first-author · 10 since 2021
YearPublicationVenuePosition
2025 SBMLtoOdin and Menelmacar: interactive visualisation of systems biology models for expert and non-expert audiences
abstract
SUMMARY: Computational models in biology can increase our understanding of biological systems, be used to answer research questions, and make predictions. Accessibility and reusability of computational models is limited and often restricted to experts in programming and mathematics. This is due to the need to implement entire models and solvers from the mathematical notation models are normally presented as. Here, we present SBMLtoOdin, an R package that translates differential equation models in SBML format from the BioModels database into executable R code using the R package odin, allowing researchers to easily reuse models. We also present Menelmacar, a web-based application that provides interactive visualisations of these models by solving their differential equations in the browser. This platform allows non-experts to simulate and investigate models using an easy-to-use interface. AVAILABILITY AND IMPLEMENTATION: SBMLtoOdin is published under the open source Apache 2.0 licence at https://github.com/bacpop/SBMLtoOdin and can be installed as an R package. The code for the Menelmacar website is published under the MIT License at https://github.com/bacpop/odinviewer, and the website can be found at https://biomodels.bacpop.org/.
Leonie J. Lorenz, Antoine Andréoletti, Tung V. N. Nguyen, Henning Hermjakob, Richard G. FitzJohn, Rahuman S. Malik-Sheriff, John A. Lees
Bioinform.4
2025 Talk2Biomodels: AI agent-based open-source LLM initiative for kinetic biological models
abstract
BACKGROUND: Quantitative kinetic models of biological regulatory processes play an important role in understanding disease mechanisms. However, their simulation and analysis require specialized domain expertise. RESULTS: In this study, we present Talk2Biomodels (T2B), an open-source, user-friendly, large language model-based agentic AI platform designed to facilitate access to computational models of biological systems and promote the FAIRification (Findability, Accessibility, Interoperability, and Reusability) principles in systems biology. T2B allows users to interact with and analyse mathematical models of biological systems through conversations in natural language, thereby lowering the barrier to entry for model interpretation and hypothesis-driven exploration. The platform natively supports models encoded in the Systems Biology Markup Language, a widely adopted standard in the computational biology community. T2B is integrated with the BioModels database ( https://www.ebi.ac.uk/biomodels/ ), enabling retrieval, simulation, and analysis of curated systems biology models. We illustrate the platform's capabilities through use cases in precision medicine, infectious disease epidemiology, and the study of emergent network-level properties in cellular systems - demonstrating how both computational experts and domain scientists without formal modelling training can derive actionable insights from complex biological models. Talk2Biomodels is available at https://github.com/VirtualPatientEngine/AIAgents4Pharma . Detailed documentation and use cases are available at https://virtualpatientengine.github.io/AIAgents4Pharma/talk2biomodels/intro/ . CONCLUSIONS: In summary, T2B lowers the barrier for non-experts to engage with and extract insights from computational models of biological systems, while simultaneously providing experts with a streamlined interface for analysing models and overall contributes to the FAIRification of models.
Lilija Wehling, Ahmad Wisnu Mulyadi, Rakesh Hadne Sreenath, Henning Hermjakob, Tung V. N. Nguyen, Thomas Rückle, Mohammed H. Mosa, Henrik Cordes, Tommaso Andreani, Thomas Klabunde, Rahuman S. Malik-Sheriff, Douglas McCloskey
BMC Bioinform.5
2025 Verification and reproducible curation of the BioModels repository
abstract
The BioModels Repository contains over 1000 manually curated mechanistic models from published literature, most often encoded in the Systems Biology Markup Language (SBML). This community-based standard formally specifies each model, but does not describe the computational experimental conditions to run a simulation and collect data. Therefore, it can be challenging to reproduce any figure or result from a publication with an SBML model alone. The Simulation Experiment Description Markup Language (SED-ML) provides a solution: a standard way to specify exactly how to run an experiment corresponding to a specific figure or result. BioModels was established years before SED-ML, and both systems evolved over time, both in content and acceptance. Hence, only about half of the entries in BioModels contained SED-ML files, and these files reflected the version of SED-ML that was available at the time. Additionally, almost all of these SED-ML files had at least one minor mistake that made them impossible to run. To make these models and their results more reproducible, we report here on our work updating, correcting and generating new SED-ML files for 1055 curated mechanistic models in BioModels. In addition, because SED-ML is implementation-independent, it can be used for verification, demonstrating that results hold across multiple simulation engines. We tested, corrected, and improved over 450 existing SED-ML files in the BioModels database, and created basic files for the rest of the entries. Then, we used a wrapper architecture for interpreting SED-ML, and report verification results across five different ODE-based biosimulation engines, after further improving the models, the wrappers, and the engines themselves. Our work with SED-ML and the BioModels collection aims to improve the utility of these models by making them more reproducible and credible. Improved reproducibility means these models are now even more fit for re-use, such as in new investigations and as components of multiscale models.
Lucian P. Smith, Rahuman S. Malik-Sheriff, Tung V. N. Nguyen, Henning Hermjakob, Jonathan R. Karr, Bilal Shaikh, Logan Drescher, Ion I. Moraru, James C. Schaff, Eran Agmon, Alexander A. Patrie, Michael L. Blinov, Joseph L. Hellerstein, Elebeoba E. May, David P. Nickerson, John H. Gennari, Herbert M. Sauro
PLoS Comput. Biol.4
2024 ReactomeGSA: new features to simplify public data reuse
abstract
MOTIVATION: ReactomeGSA is part of the Reactome knowledgebase and one of the leading multi-omics pathway analysis platforms. ReactomeGSA provides access to quantitative pathway analysis methods supporting different 'omics data types. Additionally, ReactomeGSA can process different datasets simultaneously, leading to a comparative pathway analysis that can also be performed across different species. RESULTS: We present a major update to the ReactomeGSA analysis platforms that greatly simplifies the reuse and direct integration of public data. In order to increase the number of available datasets, we developed the new grein_loader Python application that can directly fetch experiments from the GREIN resource. This enabled us to support both EMBL-EBI's Expression Atlas and GEO RNA-seq Experiments Interactive Navigator within ReactomeGSA. To further increase the visibility and simplify the reuse of public datasets, we integrated a novel search function into ReactomeGSA that enables users to search for public datasets across both supported resources. Finally, we completely re-developed ReactomeGSA's web-frontend and R/Bioconductor package to support the new search and loading features, and greatly simplify the use of ReactomeGSA. AVAILABILITY AND IMPLEMENTATION: The new ReactomeGSA web frontend is available at https://www.reactome.org/gsa with an built-in, interactive tutorial. The ReactomeGSA R package (https://bioconductor.org/packages/release/bioc/html/ReactomeGSA.html) is available through Bioconductor and shipped with detailed documentation and vignettes. The grein_loader Python application is available through the Python Package Index (pypi). The complete source code for all applications is available on GitHub at https://github.com/grisslab/grein_loader and https://github.com/reactome.
Alexander Grentner, Eliot Ragueneau, Chuqiao Gong, Adrian Prinz, Sabina Gansberger, Inigo Oyarzun, Henning Hermjakob, Johannes Griss
Bioinform.7
2024 Poincaré and SimBio: a versatile and extensible Python ecosystem for modeling systems
abstract
MOTIVATION: Chemical reaction networks (CRNs) play a pivotal role in diverse fields such as systems biology, biochemistry, chemical engineering, and epidemiology. High-level definitions of CRNs enables to use various simulation approaches, including deterministic and stochastic methods, from the same model. However, existing Python tools for simulation of CRN typically wrap external C/C++ libraries for model definition, translation into equations and/or numerically solving them, limiting their extensibility and integration with the broader Python ecosystem. RESULTS: In response, we developed Poincaré and SimBio, two novel Python packages for simulation of dynamical systems and CRNs. Poincaré serves as a foundation for dynamical systems modeling, while SimBio extends this functionality to CRNs, including support for the Systems Biology Markup Language (SBML). Poincaré and SimBio are developed as pure Python packages enabling users to easily extend their simulation capabilities by writing new or leveraging other Python packages. Moreover, this does not compromise the performance, as code can be just-in-time compiled with Numba. Our benchmark tests using curated models from the BioModels repository demonstrate that these tools may provide a potentially superior performance advantage compared to other existing tools. In addition, to ensure a user-friendly experience, our packages use standard typed modern Python syntax that provides a seamless integration with integrated development environments. Our Python-centric approach significantly enhances code analysis, error detection, and refactoring capabilities, positioning Poincaré and SimBio as valuable tools for the modeling community. AVAILABILITY AND IMPLEMENTATION: Poincaré and SimBio are released under the MIT license. Their source code is available on GitHub (https://github.com/maurosilber/poincare and https://github.com/hgrecco/simbio) and can be installed from PyPI or conda-forge.
Mauro Silberberg, Henning Hermjakob, Rahuman S. Malik-Sheriff, Hernán E. Grecco
Bioinform.2
2022 Addressing barriers in comprehensiveness, accessibility, reusability, interoperability and reproducibility of computational models in systems biology
abstract
Computational models are often employed in systems biology to study the dynamic behaviours of complex systems. With the rise in the number of computational models, finding ways to improve the reusability of these models and their ability to reproduce virtual experiments becomes critical. Correct and effective model annotation in community-supported and standardised formats is necessary for this improvement. Here, we present recent efforts toward a common framework for annotated, accessible, reproducible and interoperable computational models in biology, and discuss key challenges of the field.
Anna Niarakis, Dagmar Waltemath, James A. Glazier, Falk Schreiber, Sarah M. Keating, David P. Nickerson, Claudine Chaouiya, Anne Siegel, Vincent Noel, Henning Hermjakob, Tomás Helikar, Sylvain Soliman, Laurence Calzone
Briefings Bioinform.10
2021 A decoupled, modular and scriptable architecture for tools to curate data platforms
Juniper Tyree, Henning Hermjakob, Manuel Bernal Llinares
Bioinform.2
2021 Identifiers.org: Compact Identifier services in the cloud
abstract
MOTIVATION: Since its launch in 2010, Identifiers.org has become an important tool for the annotation and cross-referencing of Life Science data. In 2016, we established the Compact Identifier (CID) scheme (prefix: accession) to generate globally unique identifiers for data resources using their locally assigned accession identifiers. Since then, we have developed and improved services to support the growing need to create, reference and resolve CIDs, in systems ranging from human readable text to cloud-based e-infrastructures, by providing high availability and low-latency cloud-based services, backed by a high-quality, manually curated resource. RESULTS: We describe a set of services that can be used to construct and resolve CIDs in Life Sciences and beyond. We have developed a new front end for accessing the Identifiers.org registry data and APIs to simplify integration of Identifiers.org CID services with third-party applications. We have also deployed the new Identifiers.org infrastructure in a commercial cloud environment, bringing our services closer to the data. AVAILABILITYAND IMPLEMENTATION: https://identifiers.org.
Manuel Bernal Llinares, Javier Ferrer-Gómez, Nick S. Juty, Carole A. Goble, Sarala M. Wimalaratne, Henning Hermjakob
Bioinform.6
2021 IntAct App: a Cytoscape application for molecular interaction network visualization and analysis
abstract
SUMMARY: IntAct App is a Cytoscape 3 application that grants in-depth access to IntAct's molecular interaction data. It build networks where nodes are interacting molecules (mainly proteins, but also genes, RNA, chemicals…) and edges represent evidence of interaction. Users can query a network by providing its molecules, identified by different fields and optionally include all their interacting partners in the resulting network. The app offers three visualizations: one only displaying interactions, another representing every evidence and the last one emphasizing evidence where mutated versions of proteins were used. Users can also filter networks and click on nodes and edges to access all their related details. Finally, the application supports automation of its main features via Cytoscape commands. AVAILABILITY AND IMPLEMENTATION: Implementation available at https://apps.cytoscape.org/apps/intactapp, while the source code is available at https://github.com/EBI-IntAct/IntactApp.
Eliot Ragueneau, Anjali Shrivastava, John Scotter Morris, Noemi del-Toro, Henning Hermjakob, Pablo Porras
Bioinform.5
2021 The Minimum Information about a Molecular Interaction CAusal STatement (MI2CAST)
abstract
MOTIVATION: A large variety of molecular interactions occurs between biomolecular components in cells. When a molecular interaction results in a regulatory effect, exerted by one component onto a downstream component, a so-called 'causal interaction' takes place. Causal interactions constitute the building blocks in our understanding of larger regulatory networks in cells. These causal interactions and the biological processes they enable (e.g. gene regulation) need to be described with a careful appreciation of the underlying molecular reactions. A proper description of this information enables archiving, sharing and reuse by humans and for automated computational processing. Various representations of causal relationships between biological components are currently used in a variety of resources. RESULTS: Here, we propose a checklist that accommodates current representations, called the Minimum Information about a Molecular Interaction CAusal STatement (MI2CAST). This checklist defines both the required core information, as well as a comprehensive set of other contextual details valuable to the end user and relevant for reusing and reproducing causal molecular interaction information. The MI2CAST checklist can be used as reporting guidelines when annotating and curating causal statements, while fostering uniformity and interoperability of the data across resources. AVAILABILITY AND IMPLEMENTATION: The checklist together with examples is accessible at https://github.com/MI2CAST/MI2CAST. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Vasundra Touré, Steven Vercruysse, Marcio Luis Acencio, Ruth C. Lovering, Sandra E. Orchard, Glyn Bradley, Cristina Casals-Casas, Claudine Chaouiya, Noemi del-Toro, Åsmund Flobak, Pascale Gaudet, Henning Hermjakob, Charles Tapley Hoyt, Luana Licata, Astrid Lægreid, Chris Mungall, Anne Niknejad, Simona Panni, Livia Perfetto, Pablo Porras, Dexter Pratt, Julio Saez-Rodriguez, Denis Thieffry, Paul D. Thomas, Dénes Türei, Martin Kuiper
Bioinform.12
2020 BioModels Parameters: a treasure trove of parameter values from published systems biology models
abstract
MOTIVATION: One of the major bottlenecks in building systems biology models is identification and estimation of model parameters for model calibration. Searching for model parameters from published literature and models is an essential, yet laborious task. RESULTS: We have developed a new service, BioModels Parameters, to facilitate search and retrieval of parameter values from the Systems Biology Markup Language models stored in BioModels. Modellers can now directly search for a model entity (e.g. a protein or drug) to retrieve the rate equations describing it; the associated parameter values (e.g. degradation rate, production rate, Kcat, Michaelis-Menten constant, etc.) and the initial concentrations. Currently, BioModels Parameters contains entries from over 84,000 reactions and 60 different taxa with cross-references. The retrieved rate equations and parameters can be used for scanning parameter ranges, model fitting and model extension. Thus, BioModels Parameters will be a valuable service for systems biology modellers. AVAILABILITY AND IMPLEMENTATION: The data are accessible via web interface and API. BioModels Parameters is free to use and is publicly available at https://www.ebi.ac.uk/biomodels/parameterSearch. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Mihai Glont, Chinmay Arankalle, Krishan K. Tiwari, Tung V. N. Nguyen, Henning Hermjakob, Rahuman S. Malik-Sheriff
Bioinform.5
2019 CausalTAB: the PSI-MITAB 2.8 updated format for signalling data representation and dissemination
abstract
MOTIVATION: Combining multiple layers of information underlying biological complexity into a structured framework represent a challenge in systems biology. A key task is the formalization of such information in models describing how biological entities interact to mediate the response to external and internal signals. Several databases with signalling information, focus on capturing, organizing and displaying signalling interactions by representing them as binary, causal relationships between biological entities. The curation efforts that build these individual databases demand a concerted effort to ensure interoperability among resources. RESULTS: Aware of the enormous benefits of standardization efforts in the molecular interaction research field, representatives of the signalling network community agreed to extend the PSI-MI controlled vocabulary to include additional terms representing aspects of causal interactions. Here, we present a common standard for the representation and dissemination of signalling information: the PSI Causal Interaction tabular format (CausalTAB) which is an extension of the existing PSI-MI tab-delimited format, now designated PSI-MITAB 2.8. We define the new term 'causal interaction', and related child terms, which are children of the PSI-MI 'molecular interaction' term. The new vocabulary terms in this extended PSI-MI format will enable systems biologists to model large-scale signalling networks more precisely and with higher coverage than before. AVAILABILITY AND IMPLEMENTATION: PSI-MITAB 2.8 format and the new reference implementation of PSICQUIC are available online (https://psicquic.github.io/ and https://psicquic.github.io/MITAB28Format.html). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Livia Perfetto, Marcio Luis Acencio, Glyn Bradley, Gianni Cesareni, Noemi del-Toro, Dávid Fazekas, Henning Hermjakob, Tamás Korcsmáros, Martin Kuiper, Astrid Lægreid, Prisca Lo Surdo, Ruth C. Lovering, Sandra E. Orchard, Pablo Porras, Paul D. Thomas, Vasundra Touré, John Zobolas, Luana Licata
Bioinform.7
2018 Reactome diagram viewer: data structures and strategies to boost performance
abstract
Motivation: Reactome is a free, open-source, open-data, curated and peer-reviewed knowledgebase of biomolecular pathways. For web-based pathway visualization, Reactome uses a custom pathway diagram viewer that has been evolved over the past years. Here, we present comprehensive enhancements in usability and performance based on extensive usability testing sessions and technology developments, aiming to optimize the viewer towards the needs of the community. Results: The pathway diagram viewer version 3 achieves consistently better performance, loading and rendering of 97% of the diagrams in Reactome in less than 1 s. Combining the multi-layer html5 canvas strategy with a space partitioning data structure minimizes CPU workload, enabling the introduction of new features that further enhance user experience. Through the use of highly optimized data structures and algorithms, Reactome has boosted the performance and usability of the new pathway diagram viewer, providing a robust, scalable and easy-to-integrate solution to pathway visualization. As graph-based visualization of complex data is a frequent challenge in bioinformatics, many of the individual strategies presented here are applicable to a wide range of web-based bioinformatics resources. Availability and implementation: Reactome is available online at: https://reactome.org. The diagram viewer is part of the Reactome pathway browser (https://reactome.org/PathwayBrowser/) and also available as a stand-alone widget at: https://reactome.org/dev/diagram/. The source code is freely available at: https://github.com/reactome-pwp/diagram. Contact: [email protected] or [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Antonio Fabregat, Konstantinos Sidiropoulos, Guilherme Viteri, Pablo Marín-García, Peipei Ping, Lincoln Stein, Peter D'Eustachio, Henning Hermjakob
Bioinform.8
2018 Encompassing new use cases - level 3.0 of the HUPO-PSI format for molecular interactions
abstract
BACKGROUND: Systems biologists study interaction data to understand the behaviour of whole cell systems, and their environment, at a molecular level. In order to effectively achieve this goal, it is critical that researchers have high quality interaction datasets available to them, in a standard data format, and also a suite of tools with which to analyse such data and form experimentally testable hypotheses from them. The PSI-MI XML standard interchange format was initially published in 2004, and expanded in 2007 to enable the download and interchange of molecular interaction data. PSI-XML2.5 was designed to describe experimental data and to date has fulfilled this basic requirement. However, new use cases have arisen that the format cannot properly accommodate. These include data abstracted from more than one publication such as allosteric/cooperative interactions and protein complexes, dynamic interactions and the need to link kinetic and affinity data to specific mutational changes. RESULTS: The Molecular Interaction workgroup of the HUPO-PSI has extended the existing, well-used XML interchange format for molecular interaction data to meet new use cases and enable the capture of new data types, following extensive community consultation. PSI-MI XML3.0 expands the capabilities of the format beyond simple experimental data, with a concomitant update of the tool suite which serves this format. The format has been implemented by key data producers such as the International Molecular Exchange (IMEx) Consortium of protein interaction databases and the Complex Portal. CONCLUSIONS: PSI-MI XML3.0 has been developed by the data producers, data users, tool developers and database providers who constitute the PSI-MI workgroup. This group now actively supports PSI-MI XML2.5 as the main interchange format for experimental data, PSI-MI XML3.0 which additionally handles more complex data types, and the simpler, tab-delimited MITAB2.5, 2.6 and 2.7 for rapid parsing and download.
M. Sivade Dumousseau, Diego Alonso-López, Mais G. Ammari, Glyn Bradley, Nancy H. Campbell, Arnaud Céol, Gianni Cesareni, Colin W. Combe, Javier De Las Rivas, Noemi del-Toro, Joshua Heimbach, Henning Hermjakob, Igor Jurisica, Luana Licata, Ruth C. Lovering, David J. Lynn, Birgit Meldal, Gos Micklem, Simona Panni, Pablo Porras, Sylvie Ricard-Blum, Bernd Roechert, Lukasz Salwínski, Anjali Shrivastava, Julie M. Sullivan, Nicolas Thierry-Mieg, Yo Yehudi, Kim Van Roey, Sandra E. Orchard
BMC Bioinform.12
2018 Reactome graph database: Efficient access to complex pathway data
abstract
Reactome is a free, open-source, open-data, curated and peer-reviewed knowledgebase of biomolecular pathways. One of its main priorities is to provide easy and efficient access to its high quality curated data. At present, biological pathway databases typically store their contents in relational databases. This limits access efficiency because there are performance issues associated with queries traversing highly interconnected data. The same data in a graph database can be queried more efficiently. Here we present the rationale behind the adoption of a graph database (Neo4j) as well as the new ContentService (REST API) that provides access to these data. The Neo4j graph database and its query language, Cypher, provide efficient access to the complex Reactome data model, facilitating easy traversal and knowledge discovery. The adoption of this technology greatly improved query efficiency, reducing the average query time by 93%. The web service built on top of the graph database provides programmatic access to Reactome data by object oriented queries, but also supports more complex queries that take advantage of the new underlying graph-based data storage. By adopting graph database technology we are providing a high performance pathway data resource to the community. The Reactome graph database use case shows the power of NoSQL database engines for complex biological data types.
Antonio Fabregat, Florian Korninger, Guilherme Viteri, Konstantinos Sidiropoulos, Pablo Marín-García, Peipei Ping, Guanming Wu, Lincoln Stein, Peter D'Eustachio, Henning Hermjakob
PLoS Comput. Biol.10
2017 ComplexViewer: visualization of curated macromolecular complexes
abstract
SUMMARY: Proteins frequently function as parts of complexes, assemblages of multiple proteins and other biomolecules, yet network visualizations usually only show proteins as parts of binary interactions. ComplexViewer visualizes interactions with more than two participants and thereby avoids the need to first expand these into multiple binary interactions. Furthermore, if binding regions between molecules are known then these can be displayed in the context of the larger complex. AVAILABILITY AND IMPLEMENTATION: freely available under Apache version 2 license; EMBL-EBI Complex Portal: http://www.ebi.ac.uk/complexportal; Source code: https://github.com/MICommunity/ComplexViewer; Package: https://www.npmjs.com/package/complexviewer; http://biojs.io/d/complexviewer. Language: JavaScript; Web technology: Scalable Vector Graphics; Libraries: D3.js. CONTACT: [email protected] or [email protected].
Colin W. Combe, Marine Sivade, Henning Hermjakob, Joshua Heimbach, Birgit Meldal, Gos Micklem, Sandra E. Orchard, Juri Rappsilber
Bioinform.3
2017 Reactome enhanced pathway visualization
abstract
MOTIVATION: Reactome is a free, open-source, open-data, curated and peer-reviewed knowledge base of biomolecular pathways. Pathways are arranged in a hierarchical structure that largely corresponds to the GO biological process hierarchy, allowing the user to navigate from high level concepts like immune system to detailed pathway diagrams showing biomolecular events like membrane transport or phosphorylation. Here, we present new developments in the Reactome visualization system that facilitate navigation through the pathway hierarchy and enable efficient reuse of Reactome visualizations for users' own research presentations and publications. RESULTS: For the higher levels of the hierarchy, Reactome now provides scalable, interactive textbook-style diagrams in SVG format, which are also freely downloadable and editable. Repeated diagram elements like 'mitochondrion' or 'receptor' are available as a library of graphic elements. Detailed lower-level diagrams are now downloadable in editable PPTX format as sets of interconnected objects. AVAILABILITY AND IMPLEMENTATION: http://reactome.org. CONTACT: [email protected] or [email protected].
Konstantinos Sidiropoulos, Guilherme Viteri, Cristoffer Sevilla, Steven Jupe, Marissa Webber, Marija Orlic-Milacic, Bijay Jassal, Bruce May, Veronica Shamovsky, Corina Duenas, Karen Rothfels, Lisa Matthews, Heeyeon Song, Lincoln Stein, Robin Haw, Peter D'Eustachio, Peipei Ping, Henning Hermjakob, Antonio Fabregat
Bioinform.18
2017 Reactome pathway analysis: a high-performance in-memory approach
abstract
BACKGROUND: Reactome aims to provide bioinformatics tools for visualisation, interpretation and analysis of pathway knowledge to support basic research, genome analysis, modelling, systems biology and education. Pathway analysis methods have a broad range of applications in physiological and biomedical research; one of the main problems, from the analysis methods performance point of view, is the constantly increasing size of the data samples. RESULTS: Here, we present a new high-performance in-memory implementation of the well-established over-representation analysis method. To achieve the target, the over-representation analysis method is divided in four different steps and, for each of them, specific data structures are used to improve performance and minimise the memory footprint. The first step, finding out whether an identifier in the user's sample corresponds to an entity in Reactome, is addressed using a radix tree as a lookup table. The second step, modelling the proteins, chemicals, their orthologous in other species and their composition in complexes and sets, is addressed with a graph. The third and fourth steps, that aggregate the results and calculate the statistics, are solved with a double-linked tree. CONCLUSION: Through the use of highly optimised, in-memory data structures and algorithms, Reactome has achieved a stable, high performance pathway analysis service, enabling the analysis of genome-wide datasets within seconds, allowing interactive exploration and analysis of high throughput data. The proposed pathway analysis approach is available in the Reactome production web site either via the AnalysisService for programmatic access or the user submission interface integrated into the PathwayBrowser. Reactome is an open data and open source project and all of its source code, including the one described here, is available in the AnalysisTools repository in the Reactome GitHub ( https://github.com/reactome/ ).
Antonio Fabregat, Konstantinos Sidiropoulos, Guilherme Viteri, Oscar Forner-Martinez, Pablo Marín-García, Vicente Arnau, Peter D'Eustachio, Lincoln Stein, Henning Hermjakob
BMC Bioinform.9
2016 Accurate estimation of isoelectric point of protein and peptide based on amino acid sequences
abstract
MOTIVATION: In any macromolecular polyprotic system-for example protein, DNA or RNA-the isoelectric point-commonly referred to as the pI-can be defined as the point of singularity in a titration curve, corresponding to the solution pH value at which the net overall surface charge-and thus the electrophoretic mobility-of the ampholyte sums to zero. Different modern analytical biochemistry and proteomics methods depend on the isoelectric point as a principal feature for protein and peptide characterization. Protein separation by isoelectric point is a critical part of 2-D gel electrophoresis, a key precursor of proteomics, where discrete spots can be digested in-gel, and proteins subsequently identified by analytical mass spectrometry. Peptide fractionation according to their pI is also widely used in current proteomics sample preparation procedures previous to the LC-MS/MS analysis. Therefore accurate theoretical prediction of pI would expedite such analysis. While such pI calculation is widely used, it remains largely untested, motivating our efforts to benchmark pI prediction methods. RESULTS: Using data from the database PIP-DB and one publically available dataset as our reference gold standard, we have undertaken the benchmarking of pI calculation methods. We find that methods vary in their accuracy and are highly sensitive to the choice of basis set. The machine-learning algorithms, especially the SVM-based algorithm, showed a superior performance when studying peptide mixtures. In general, learning-based pI prediction methods (such as Cofactor, SVM and Branca) require a large training dataset and their resulting performance will strongly depend of the quality of that data. In contrast with Iterative methods, machine-learning algorithms have the advantage of being able to add new features to improve the accuracy of prediction. CONTACT: [email protected] AVAILABILITY AND IMPLEMENTATION: The software and data are freely available at https://github.com/ypriverol/pIRSupplementary information: Supplementary data are available at Bioinformatics online.
Enrique Audain, Yassel Ramos, Henning Hermjakob, Darren R. Flower, Yasset Pérez-Riverol
Bioinform.3
2015 ms-data-core-api: an open-source, metadata-oriented library for computational proteomics
abstract
UNLABELLED: The ms-data-core-api is a free, open-source library for developing computational proteomics tools and pipelines. The Application Programming Interface, written in Java, enables rapid tool creation by providing a robust, pluggable programming interface and common data model. The data model is based on controlled vocabularies/ontologies and captures the whole range of data types included in common proteomics experimental workflows, going from spectra to peptide/protein identifications to quantitative results. The library contains readers for three of the most used Proteomics Standards Initiative standard file formats: mzML, mzIdentML, and mzTab. In addition to mzML, it also supports other common mass spectra data formats: dta, ms2, mgf, pkl, apl (text-based), mzXML and mzData (XML-based). Also, it can be used to read PRIDE XML, the original format used by the PRIDE database, one of the world-leading proteomics resources. Finally, we present a set of algorithms and tools whose implementation illustrates the simplicity of developing applications using the library. AVAILABILITY AND IMPLEMENTATION: The software is freely available at https://github.com/PRIDE-Utilities/ms-data-core-api. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online CONTACT: [email protected].
Yasset Pérez-Riverol, Julian Uszkoreit, Aniel Sánchez, Tobias Ternent, Noemi del-Toro, Henning Hermjakob, Juan Antonio Vizcaíno
Bioinform.6
2015 SPARQL-enabled identifier conversion with Identifiers.org
abstract
MOTIVATION: On the semantic web, in life sciences in particular, data is often distributed via multiple resources. Each of these sources is likely to use their own International Resource Identifier for conceptually the same resource or database record. The lack of correspondence between identifiers introduces a barrier when executing federated SPARQL queries across life science data. RESULTS: We introduce a novel SPARQL-based service to enable on-the-fly integration of life science data. This service uses the identifier patterns defined in the Identifiers.org Registry to generate a plurality of identifier variants, which can then be used to match source identifiers with target identifiers. We demonstrate the utility of this identifier integration approach by answering queries across major producers of life science Linked Data. AVAILABILITY AND IMPLEMENTATION: The SPARQL-based identifier conversion service is available without restriction at http://identifiers.org/services/sparql.
Sarala M. Wimalaratne, Jerven T. Bolleman, Nick S. Juty, Toshiaki Katayama, Michel Dumontier, Nicole Redaschi, Nicolas Le Novère, Henning Hermjakob, Camille Laibe
Bioinform.8
2015 Development of data representation standards by the human proteome organization proteomics standards initiative
abstract
OBJECTIVE: To describe the goals of the Proteomics Standards Initiative (PSI) of the Human Proteome Organization, the methods that the PSI has employed to create data standards, the resulting output of the PSI, lessons learned from the PSI's evolution, and future directions and synergies for the group. MATERIALS AND METHODS: The PSI has 5 categories of deliverables that have guided the group. These are minimum information guidelines, data formats, controlled vocabularies, resources and software tools, and dissemination activities. These deliverables are produced via the leadership and working group organization of the initiative, driven by frequent workshops and ongoing communication within the working groups. Official standards are subjected to a rigorous document process that includes several levels of peer review prior to release. RESULTS: We have produced and published minimum information guidelines describing what information should be provided when making data public, either via public repositories or other means. The PSI has produced a series of standard formats covering mass spectrometer input, mass spectrometer output, results of informatics analysis (both qualitative and quantitative analyses), reports of molecular interaction data, and gel electrophoresis analyses. We have produced controlled vocabularies that ensure that concepts are uniformly annotated in the formats and engaged in extensive software development and dissemination efforts so that the standards can efficiently be used by the community.Conclusion In its first dozen years of operation, the PSI has produced many standards that have accelerated the field of proteomics by facilitating data exchange and deposition to data repositories. We look to the future to continue developing standards for new proteomics technologies and workflows and mechanisms for integration with other omics data types. Our products facilitate the translation of genomics and proteomics findings to clinical and biological phenotypes. The PSI website can be accessed at http://www.psidev.info.
Eric W. Deutsch, Juan P. Albar, Pierre-Alain Binz, Martin Eisenacher, Andrew R. Jones, Gerhard Mayer, Gilbert S. Omenn, Sandra E. Orchard, Juan Antonio Vizcaíno, Henning Hermjakob
J. Am. Medical Informatics Assoc.10
2013 BioJS: an open source JavaScript framework for biological data visualization
abstract
SUMMARY: BioJS is an open-source project whose main objective is the visualization of biological data in JavaScript. BioJS provides an easy-to-use consistent framework for bioinformatics application programmers. It follows a community-driven standard specification that includes a collection of components purposely designed to require a very simple configuration and installation. In addition to the programming framework, BioJS provides a centralized repository of components available for reutilization by the bioinformatics community. AVAILABILITY AND IMPLEMENTATION: http://code.google.com/p/biojs/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
John Gómez, Leyla Jael Castro, Gustavo A. Salazar, Jose M. Villaveces, Swanand P. Gore, Alexander García Castro, Maria Jesus Martin, Guillaume Launay, Rafael Alcántara, Noemi del-Toro, Marine Sivade, Sandra E. Orchard, Sameer Velankar, Henning Hermjakob, Chenggong Zong, Peipei Ping, Manuel Corpas, Rafael C. Jiménez
Bioinform.14
2013 iAnn: an event sharing platform for the life sciences
abstract
SUMMARY: We present iAnn, an open source community-driven platform for dissemination of life science events, such as courses, conferences and workshops. iAnn allows automatic visualisation and integration of customised event reports. A central repository lies at the core of the platform: curators add submitted events, and these are subsequently accessed via web services. Thus, once an iAnn widget is incorporated into a website, it permanently shows timely relevant information as if it were native to the remote site. At the same time, announcements submitted to the repository are automatically disseminated to all portals that query the system. To facilitate the visualization of announcements, iAnn provides powerful filtering options and views, integrated in Google Maps and Google Calendar. All iAnn widgets are freely available. AVAILABILITY: http://iann.pro/iannviewer CONTACT: [email protected].
Rafael C. Jiménez, Juan P. Albar, Jong Bhak, Marie-Claude Blatter, Thomas Blicher, Michelle D. Brazas, Catherine Brooksbank, Aidan Budd, Javier De Las Rivas, Jacqueline Dreyer, Marc A. van Driel, Michael J. Dunn, Pedro L. Fernandes, Celia W. G. van Gelder, Henning Hermjakob, Vassilios Ioannidis, David Phillip Judge, Pascal Kahlem, Eija Korpelainen, Hans-Joachim Kraus, Jane E. Loveland, Christine Mayer, Jennifer McDowall, Federico Morán, Nicola J. Mulder, Tommi H. Nyrönen, Kristian Rother, Gustavo A. Salazar, Reinhard Schneider 0002, Allegra Via, Jose M. Villaveces, Maria Victoria Schneider, Terri K. Attwood, Manuel Corpas
Bioinform.15
2012 Hydra: a scalable proteomic search engine which utilizes the Hadoop distributed computing framework
abstract
BACKGROUND: For shotgun mass spectrometry based proteomics the most computationally expensive step is in matching the spectra against an increasingly large database of sequences and their post-translational modifications with known masses. Each mass spectrometer can generate data at an astonishingly high rate, and the scope of what is searched for is continually increasing. Therefore solutions for improving our ability to perform these searches are needed. RESULTS: We present a sequence database search engine that is specifically designed to run efficiently on the Hadoop MapReduce distributed computing framework. The search engine implements the K-score algorithm, generating comparable output for the same input files as the original implementation. The scalability of the system is shown, and the architecture required for the development of such distributed processing is discussed. CONCLUSION: The software is scalable in its ability to handle a large peptide database, numerous modifications and large numbers of spectra. Performance scales with the number of processors in the cluster, allowing throughput to expand with the available resources.
Steven Lewis, Attila Csordas, Sarah A. Killcoyne, Henning Hermjakob, Michael R. Hoopmann, Robert L. Moritz, Eric W. Deutsch, John Boyle
BMC Bioinform.4
2011 Dasty3, a WEB framework for DAS
abstract
MOTIVATION: Dasty3 is a highly interactive and extensible Web-based framework. It provides a rich Application Programming Interface upon which it is possible to develop specialized clients capable of retrieving information from DAS sources as well as from data providers not using the DAS protocol. Dasty3 provides significant improvements on previous Web-based frameworks and is implemented using the 1.6 DAS specification. AVAILABILITY: Dasty3 is an open-source tool freely available at http://www.ebi.ac.uk/dasty/ under the terms of the GNU General public license. Source and documentation can be found at http://code.google.com/p/dasty/. CONTACT: [email protected].
Jose M. Villaveces, Rafael C. Jiménez, Leyla Jael Castro, Gustavo A. Salazar, Bernat Gel, Nicola J. Mulder, Maria Jesus Martin, Alexander García Castro, Henning Hermjakob
Bioinform.9
2011 easyDAS: Automatic creation of DAS servers
abstract
BACKGROUND: The Distributed Annotation System (DAS) has proven to be a successful way to publish and share biological data. Although there are more than 750 active registered servers from around 50 organizations, setting up a DAS server comprises a fair amount of work, making it difficult for many research groups to share their biological annotations. Given the clear advantage that the generalized sharing of relevant biological data is for the research community it would be desirable to facilitate the sharing process. RESULTS: Here we present easyDAS, a web-based system enabling anyone to publish biological annotations with just some clicks. The system, available at http://www.ebi.ac.uk/panda-srv/easydas is capable of reading different standard data file formats, process the data and create a new publicly available DAS source in a completely automated way. The created sources are hosted on the EBI systems and can take advantage of its high storage capacity and network connection, freeing the data provider from any network management work. easyDAS is an open source project under the GNU LGPL license. CONCLUSIONS: easyDAS is an automated DAS source creation system which can help many researchers in sharing their biological data, potentially increasing the amount of relevant biological data available to the scientific community.
Bernat Gel, Andrew M. Jenkinson, Rafael C. Jiménez, Xavier Messeguer, Henning Hermjakob
BMC Bioinform.5
2011 DAS Writeback: A Collaborative Annotation System
abstract
BACKGROUND: Centralised resources such as GenBank and UniProt are perfect examples of the major international efforts that have been made to integrate and share biological information. However, additional data that adds value to these resources needs a simple and rapid route to public access. The Distributed Annotation System (DAS) provides an adequate environment to integrate genomic and proteomic information from multiple sources, making this information accessible to the community. DAS offers a way to distribute and access information but it does not provide domain experts with the mechanisms to participate in the curation process of the available biological entities and their annotations. RESULTS: We designed and developed a Collaborative Annotation System for proteins called DAS Writeback. DAS writeback is a protocol extension of DAS to provide the functionalities of adding, editing and deleting annotations. We implemented this new specification as extensions of both a DAS server and a DAS client. The architecture was designed with the involvement of the DAS community and it was improved after performing usability experiments emulating a real annotation task. CONCLUSIONS: We demonstrate that DAS Writeback is effective, usable and will provide the appropriate environment for the creation and evolution of community protein annotation.
Gustavo A. Salazar, Rafael C. Jiménez, Alexander García Castro, Henning Hermjakob, Nicola J. Mulder, Edwin H. Blake
BMC Bioinform.4
2008 The HUPO proteomics standards initiative - easing communication and minimizing data loss in a changing world
abstract
The amount of data currently being generated by proteomics laboratories around the world is increasing exponentially, making it ever more critical that scientists are able to exchange, compare and retrieve datasets when re-evaluation of their original conclusions becomes important. Only a fraction of this data is published in the literature and important information is being lost every day as data formats become obsolete. The Human Proteome Organisation Proteomics Standards Initiative (HUPO-PSI) was tasked with the creation of data standards and interchange formats to allow both the exchange and storage of such data irrespective of the hardware and software from which it was generated. This article will provide an update on the work of this group, the creation and implementation of these standards and the standards-compliant data repositories being established as result of their efforts.
Sandra E. Orchard, Henning Hermjakob
Briefings Bioinform.2
2008 Rintact: enabling computational analysis of molecular interaction data from the IntAct repository
abstract
MOTIVATION: The IntAct repository is one of the largest and most widely used databases for the curation and storage of molecular interaction data. These datasets need to be analyzed by computational methods. Software packages in the statistical environment R provide powerful tools for conducting such analyses. RESULTS: We introduce Rintact, a Bioconductor package that allows users to transform PSI-MI XML2.5 interaction data files from IntAct into R graph objects. On these, they can use methods from R and Bioconductor for a variety of tasks: determining cohesive subgraphs, computing summary statistics, fitting mathematical models to the data or rendering graphical layouts. Rintact provides a programmatic interface to the IntAct repository and allows the use of the analytic methods provided by R and Bioconductor. AVAILABILITY: Rintact is freely available at http://bioconductor.org
Tony Chiang, Nianhua Li, Sandra E. Orchard, Samuel Kerrien, Henning Hermjakob, Robert Gentleman, Wolfgang Huber
Bioinform.5
2008 Dasty2, an Ajax protein DAS client
abstract
SUMMARY: Dasty2 is a highly interactive web client integrating protein sequence annotations from currently more than 40 sources, using the distributed annotation system (DAS). AVAILABILITY: Dasty2 is an open source tool freely available under the terms of the Apache License 2.0, publicly available at http://www.ebi.ac.uk/dasty/.
Rafael C. Jiménez, Antony F. Quinn, Alexander García Castro, Alberto Labarga, Kieran O'Neill, Gustavo A. Salazar, Henning Hermjakob
Bioinform.8
2008 InteroPORC: automated inference of highly conserved protein interaction networks
abstract
MOTIVATION: Protein-protein interaction networks provide insights into the relationships between the proteins of an organism thereby contributing to a better understanding of cellular processes. Nevertheless, large-scale interaction networks are available for only a few model organisms. Thus, interologs are useful for a systematic transfer of protein interaction networks between organisms. However, no standard tool is available so far for that purpose. RESULTS: In this study, we present an automated prediction tool developed for all sequenced genomes available in Integr8. We also have developed a second method to predict protein-protein interactions in the widely used cyanobacterium Synechocystis. Using these methods, we have constructed a new network of 8783 inferred interactions for Synechocystis. AVAILABILITY: InteroPORC is open-source, downloadable and usable through a web interface at http://biodev.extra.cea.fr/interoporc/.
Magali Michaut, Samuel Kerrien, Luisa Montecchi-Palazzi, Franck Chauvat, Corinne Cassier-Chauvat, Jean-Christophe Aude, Pierre Legrain, Henning Hermjakob
Bioinform.8
2008 The Protein Feature Ontology: a tool for the unification of protein feature annotations
abstract
MOTIVATION: The advent of sequencing and structural genomics projects has provided a dramatic boost in the number of uncharacterized protein structures and sequences. Consequently, many computational tools have been developed to help elucidate protein function. However, such services are spread throughout the world, often with standalone web pages. Integration of these methods is needed and so far this has not been possible as there was no common vocabulary available that could be used as a standard language. RESULTS: The Protein Feature Ontology has been developed to provide a structured controlled vocabulary for features on a protein sequence or structure and comprises approximately 100 positional terms, now integrated into the Sequence Ontology (SO) and 40 non-positional terms which describe features relating to the whole-protein sequence. In addition, post-translational modifications are described by using a pre-existing ontology, the Protein Modification Ontology (MOD). This ontology is being used to integrate over 150 distinct annotations provided by the BioSapiens Network of Excellence, a consortium comprising 19 partner sites in Europe. AVAILABILITY: The Protein Feature Ontology can be browsed by accessing the ontology lookup service at the European Bioinformatics Institute (http://www.ebi.ac.uk/ontology-lookup/browse.do?ontName=BS).
Gabrielle A. Reeves, Karen Eilbeck, Michele Magrane, Claire O'Donovan, Luisa Montecchi-Palazzi, Midori A. Harris, Sandra E. Orchard, Rafael C. Jiménez, Andreas Prlic, Tim J. P. Hubbard, Henning Hermjakob, Janet M. Thornton
Bioinform.11
2008 Integrating biological data - the Distributed Annotation System
abstract
BACKGROUND: The Distributed Annotation System (DAS) is a widely adopted protocol for dynamically integrating a wide range of biological data from geographically diverse sources. DAS continues to expand its applicability and evolve in response to new challenges facing integrative bioinformatics. RESULTS: Here we describe the various infrastructure components of DAS and present a new extended version of the DAS specification. Version 1.53E incorporates several recent developments, including its extension to serve new data types and an ontology for protein features. CONCLUSION: Our extensions to the DAS protocol have facilitated the integration of new data types, and our improvements to the existing DAS infrastructure have addressed recent challenges. The steadily increasing numbers of available data sources demonstrates further adoption of the DAS protocol.
Andrew M. Jenkinson, Mario Albrecht, Ewan Birney, Hagen Blankenburg, Thomas A. Down, Robert D. Finn, Henning Hermjakob, Tim J. P. Hubbard, Rafael C. Jiménez, Philip Jones, Andreas Kähäri, Eugene Kulesha, José R. Macías, Gabrielle A. Reeves, Andreas Prlic
BMC Bioinform.7
2008 InteroPORC: an automated tool to predict highly conserved protein interaction networks
Magali Michaut, Samuel Kerrien, Luisa Montecchi-Palazzi, Corinne Cassier-Chauvat, Franck Chauvat, Jean-Christophe Aude, Pierre Legrain, Henning Hermjakob
BMC Bioinform.8
2008 OntoDas - a tool for facilitating the construction of complex queries to the Gene Ontology
abstract
BACKGROUND: Ontologies such as the Gene Ontology can enable the construction of complex queries over biological information in a conceptual way, however existing systems to do this are too technical. Within the biological domain there is an increasing need for software that facilitates the flexible retrieval of information. OntoDas aims to fulfil this need by allowing the definition of queries by selecting valid ontology terms. RESULTS: OntoDas is a web-based tool that uses information visualisation techniques to provide an intuitive, interactive environment for constructing ontology-based queries against the Gene Ontology Database. Both a comprehensive use case and the interface itself were designed in a participatory manner by working with biologists to ensure that the interface matches the way biologists work. OntoDas was further tested with a separate group of biologists and refined based on their suggestions. CONCLUSION: OntoDas provides a visual and intuitive means for constructing complex queries against the Gene Ontology. It was designed with the participation of biologists and compares favourably with similar tools. It is available at http://ontodas.nbn.ac.za.
Kieran O'Neill, Alexander García Castro, Anita Schwegmann, Rafael C. Jiménez, Dan Jacobson, Henning Hermjakob
BMC Bioinform.6
2007 The Protein Identifier Cross-Referencing (PICR) service: reconciling protein identifiers across multiple source databases
abstract
BACKGROUND: Each major protein database uses its own conventions when assigning protein identifiers. Resolving the various, potentially unstable, identifiers that refer to identical proteins is a major challenge. This is a common problem when attempting to unify datasets that have been annotated with proteins from multiple data sources or querying data providers with one flavour of protein identifiers when the source database uses another. Partial solutions for protein identifier mapping exist but they are limited to specific species or techniques and to a very small number of databases. As a result, we have not found a solution that is generic enough and broad enough in mapping scope to suit our needs. RESULTS: We have created the Protein Identifier Cross-Reference (PICR) service, a web application that provides interactive and programmatic (SOAP and REST) access to a mapping algorithm that uses the UniProt Archive (UniParc) as a data warehouse to offer protein cross-references based on 100% sequence identity to proteins from over 70 distinct source databases loaded into UniParc. Mappings can be limited by source database, taxonomic ID and activity status in the source database. Users can copy/paste or upload files containing protein identifiers or sequences in FASTA format to obtain mappings using the interactive interface. Search results can be viewed in simple or detailed HTML tables or downloaded as comma-separated values (CSV) or Microsoft Excel (XLS) files suitable for use in a local database or a spreadsheet. Alternatively, a SOAP interface is available to integrate PICR functionality in other applications, as is a lightweight REST interface. CONCLUSION: We offer a publicly available service that can interactively map protein identifiers and protein sequences to the majority of commonly used protein databases. Programmatic access is available through a standards-compliant SOAP interface or a lightweight REST interface. The PICR interface, documentation and code examples are available at http://www.ebi.ac.uk/Tools/picr.
Richard G. Côté, Philip Jones, Lennart Martens, Samuel Kerrien, Florian Reisinger, Quan Lin, Rasko Leinonen, Rolf Apweiler, Henning Hermjakob
BMC Bioinform.9
2006 The Ontology Lookup Service, a lightweight cross-platform tool for controlled vocabulary queries
abstract
BACKGROUND: With the vast amounts of biomedical data being generated by high-throughput analysis methods, controlled vocabularies and ontologies are becoming increasingly important to annotate units of information for ease of search and retrieval. Each scientific community tends to create its own locally available ontology. The interfaces to query these ontologies tend to vary from group to group. We saw the need for a centralized location to perform controlled vocabulary queries that would offer both a lightweight web-accessible user interface as well as a consistent, unified SOAP interface for automated queries. RESULTS: The Ontology Lookup Service (OLS) was created to integrate publicly available biomedical ontologies into a single database. All modified ontologies are updated daily. A list of currently loaded ontologies is available online. The database can be queried to obtain information on a single term or to browse a complete ontology using AJAX. Auto-completion provides a user-friendly search mechanism. An AJAX-based ontology viewer is available to browse a complete ontology or subsets of it. A programmatic interface is available to query the webservice using SOAP. The service is described by a WSDL descriptor file available online. A sample Java client to connect to the webservice using SOAP is available for download from SourceForge. All OLS source code is publicly available under the open source Apache Licence. CONCLUSION: The OLS provides a user-friendly single entry point for publicly available ontologies in the Open Biomedical Ontology (OBO) format. It can be accessed interactively or programmatically at http://www.ebi.ac.uk/ontology-lookup/.
Richard G. Côté, Philip Jones, Rolf Apweiler, Henning Hermjakob
BMC Bioinform.4
2005 Dasty and UniProt DAS: a perfect pair for protein feature visualization
abstract
In this study, we present two freely available and complementary Distributed Annotation System (DAS) resources: a DAS reference server that provides up-to-date sequence and annotation from UniProt, with additional feature links and database cross-references from InterPro and a DAS client implemented using Java and Macromedia Flash that is optimized for the display of protein features.
Philip Jones, Nisha Vinod, Thomas A. Down, Andre Hackmann, Andreas Kähäri, Ernst Kretschmann, Antony F. Quinn, Daniela Wieser, Henning Hermjakob, Rolf Apweiler
Bioinform.9
2002 InterPro: An Integrated Documentation Resource for Protein Families, Domains and Functional Sites
abstract
The exponential increase in the submission of nucleotide sequences to the nucleotide sequence database by genome sequencing centres has resulted in a need for rapid, automatic methods for classification of the resulting protein sequences. There are several signature and sequence cluster-based methods for protein classification, each resource having distinct areas of optimum application owing to the differences in the underlying analysis methods. In recognition of this, InterPro was developed as an integrated documentation resource for protein families, domains and functional sites, to rationalise the complementary efforts of the individual protein signature database projects. The member databases - PRINTS, PROSITE, Pfam, ProDom, SMART and TIGRFAMs - form the InterPro core. Related signatures from each member database are unified into single InterPro entries. Each InterPro entry includes a unique accession number, functional descriptions and literature references, and links are made back to the relevant member database(s). Release 4.0 of InterPro (November 2001) contains 4,691 entries, representing 3,532 families, 1,068 domains, 74 repeats and 15 sites of post-translational modification (PTMs) encoded by different regular expressions, profiles, fingerprints and hidden Markov models (HMMs). Each InterPro entry lists all the matches against SWISS-PROT and TrEMBL (2,141,621 InterPro hits from 586,124 SWISS-PROT and TrEMBL protein sequences). The database is freely accessible for text- and sequence-based searches.
Nicola J. Mulder, Rolf Apweiler, Terri K. Attwood, Amos Bairoch, Alex Bateman, David Binns, Margaret Biswas, Paul Bradley, Peer Bork, Philipp Bucher, Richard R. Copley, Emmanuel Courcelle, Richard Durbin, Laurent Falquet, Wolfgang Fleischmann, Jérôme Gouzy, Sam Griffiths-Jones, Daniel H. Haft, Henning Hermjakob, Nicolas Hulo, Daniel Kahn, Alexander Kanapin, Maria Krestyaninova, Rodrigo Lopez, Ivica Letunic, Sue Orchard, Marco Pagni, David Peyruc, Chris P. Ponting, Florence Servant, Christian J. A. Sigrist
Briefings Bioinform.19
2000 InterPro-an integrated documentation resource for protein families, domains and functional sites
abstract
MOTIVATION: InterPro is a new integrated documentation resource for protein families, domains and functional sites, developed initially as a means of rationalising the complementary efforts of the PROSITE, PRINTS, Pfam and ProDom database projects. RESULTS: Merged annotations from PRINTS, PROSITE and Pfam form the InterPro core. Each combined InterPro entry includes functional descriptions and literature references, and links are made back to the relevant parent database(s), allowing users to see at a glance whether a particular family or domain has associated patterns, profiles, fingerprints, etc. Merged and individual entries (i.e. those that have no counterpart in the companion resources) are assigned unique accession numbers. Release 1.2 of InterPro (June 2000) contains over 3000 entries, representing families, domains, repeats and sites of post-translational modification (PTMs) encoded by 6581 different regular expressions, profiles, fingerprints and Hidden Markov Models (HMMs). Each InterPro entry lists all the matches against SWISS-PROT and TrEMBL (more than 1000000 hits from 264333 different proteins out of 384572 in SWISS-PROT and TrEMBL).
Rolf Apweiler, Terri K. Attwood, Amos Bairoch, Alex Bateman, Ewan Birney, Margaret Biswas, Philipp Bucher, Lorenzo Cerutti, Florence Corpet, Michael D. R. Croning, Richard Durbin, Laurent Falquet, Wolfgang Fleischmann, Jérôme Gouzy, Henning Hermjakob, Nicolas Hulo, Inge Jonassen, Daniel Kahn, Alexander Kanapin, Youla Karavidopoulou, Rodrigo Lopez, Beate Marx, Nicola J. Mulder, Thomas M. Oinn, Marco Pagni, Florence Servant, Christian J. A. Sigrist, Evgeny M. Zdobnov
Bioinform.15
2000 VARSPLIC: alternatively-spliced protein sequences derived from SWISS-PROT and TrEMBL
abstract
UNLABELLED: The program varsplic.pl uses information present in the SWISS-PROT and TrEMBL databases to create new records for alternatively spliced isoforms. These new records can be used in similarity searches. AVAILABILITY: The program is available at ftp://ftp.ebi.ac.uk/pub/software/swissprot/, together with regularly updated output files. CONTACT: [email protected]
Paul J. Kersey, Henning Hermjakob, Rolf Apweiler
Bioinform.2
2000 A comparison of signal sequence prediction methods using a test set of signal peptides
abstract
Abstract Summary: We describe the creation of a test set containing secretory and non-secretory proteins. Five existing prediction programs for signal sequences and their cleavage sites are compared on the basis of this test set: SPScan, SigCleave, SignalP V1.1, SignalP V2.0.b2-HMM and SignalP V2.0.b2-NN. Availability: Detailed results, the test set and the script to create the test set are available at ftp://ftp.ebi.ac.uk/pub/contrib/swissprot/testsets/signal/ Contact: [email protected]
Kerstin M. L. Menne, Henning Hermjakob, Rolf Apweiler
Bioinform.2
1999 Swissknife - 'lazy parsing' of SWISS-PROT entries
abstract
UNLABELLED: We present Swissknife, a set of Perl modules which provides a fast and reliable object-oriented interface to parsing and modifying files in SWISS-PROT format. AVAILABILITY: The Swissknife modules are available at ftp://ftp.ebi.ac. uk/pub/software/swissprot/. CONTACT: [email protected]
Henning Hermjakob, Wolfgang Fleischmann, Rolf Apweiler
Bioinform.1
1997 RIFLE: Rapid Identification of Microorganisms by Fragment Length Evaluation
Henning Hermjakob, Robert Giegerich, Walter Arnold
ISMB1