EDBT 2026 Demo / reviewers in the wild / expert
Christoph Steinbeck
dblp:37/2657
· DBLP profile ↗
21ranked-venue papers
0as first author
2since 2021 · last 2022
0000-0001-6966-0814ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 19 · 2 since 2021Artificial intelligence and machine learning · 2Theory of computation · 2Software engineering, systems software and programming languages · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
5 papers |
Bioinformatics and computational biology · 100% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Cloud and datacenter computing · 100% |
Topics — the 7 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology
metabolomics |
0.7 | 2 | 2019 | Interoperable and scalable data analysis with microservices: applications in metabolomics · Bioinform. 2019 mzML2ISA & nmrML2ISA: generating enriched ISA-Tab metadata files from metabolomics XML data · Bioinform. 2017 |
Cloud and datacenter computing
container orchestration |
0.4 | 1 | 2019 | Interoperable and scalable data analysis with microservices: applications in metabolomics · Bioinform. 2019 |
Bioinformatics and computational biology › molecular informatics › cheminformatics
atom mapping |
0.2 | 1 | 2016 | Reaction Decoder Tool (RDT): extracting features from chemical reactions · Bioinform. 2016 |
Bioinformatics and computational biology
enzymatic reaction analysis |
0.2 | 1 | 2016 | Reaction Decoder Tool (RDT): extracting features from chemical reactions · Bioinform. 2016 |
Bioinformatics and computational biology › systems bioinformatics › pathway analysis
metabolic pathway analysis |
0.2 | 1 | 2016 | Reaction Decoder Tool (RDT): extracting features from chemical reactions · Bioinform. 2016 |
Bioinformatics and computational biology › knowledge representation in biology
biomedical ontology |
0.2 | 1 | 2013 | OntoQuery: easy-to-use web-based OWL querying · Bioinform. 2013 |
Bioinformatics and computational biology › systems biology
metabolic modeling |
0.2 | 1 | 2013 | Metingear: a development environment for annotating genome-scale metabolic models · Bioinform. 2013 |
Methods — techniques the papers use, named apart from their topics
microservice architecture · 0.8kubernetes · 0.8docker · 0.8XML parsing · 0.3ISA-Tab format conversion · 0.3dynamic programming · 0.2description logic querying · 0.2database cross-referencing · 0.2chemical structure assignment · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Toward a Framework for Integrative, FAIR, and Reproducible Management of Data on the Dynamic Balance of Microbial CommunitiesabstractThe increasing volumes of data produced by high-throughput instruments coupled with advanced computational infrastructures for scientific computing have enabled what is often called a Fourth Paradigm for scientific research based on the exploration of large datasets. Current scientific research is often interdisciplinary, making data integration a critical technique for combining data from different scientific domains. Research data management is a critical part of this paradigm, through the proposition and development of methods, techniques, and practices for managing scientific data through their life cycle. Research on microbial communities follows the same pattern of production of large amounts of data obtained, for instance, from sequencing organisms present in environmental samples. Data on microbial communities can come from a multitude of sources and can be stored in different formats. For example, data from metagenomics, metatranscriptomics, metabolomics, and biological imaging are often combined in studies. In this article, we describe the design and current state of implementation of an integrative research data management framework for the Cluster of Excellence Balance of the Microverse aiming to allow for data on microbial communities to be more easily discovered, accessed, combined, and reused. This framework is based on research data repositories and best practices for managing workflows used in the analysis of microbial communities, which includes recording provenance information for tracking data derivation. Luiz M. R. Gadelha Jr., Martin Hohmuth, Mahnoor Zulfiqar, David Schöne, Sheeba Samuel, Maria Sorokina, Christoph Steinbeck, Birgitta König-Ries |
e-Science | 7 |
| 2021 | Chemical graph generatorsabstractChemical graph generators are software packages to generate computer representations of chemical structures adhering to certain boundary conditions. Their development is a research topic of cheminformatics. Chemical graph generators are used in areas such as virtual library generation in drug design, in molecular design with specified properties, called inverse QSAR/QSPR, as well as in organic synthesis design, retrosynthesis or in systems for computer-assisted structure elucidation (CASE). CASE systems again have regained interest for the structure elucidation of unknowns in computational metabolomics, a current area of computational biology. Mehmet Aziz Yirik, Christoph Steinbeck |
PLoS Comput. Biol. | 2 |
| 2019 | Interoperable and scalable data analysis with microservices: applications in metabolomicsabstractMOTIVATION: Developing a robust and performant data analysis workflow that integrates all necessary components whilst still being able to scale over multiple compute nodes is a challenging task. We introduce a generic method based on the microservice architecture, where software tools are encapsulated as Docker containers that can be connected into scientific workflows and executed using the Kubernetes container orchestrator. RESULTS: We developed a Virtual Research Environment (VRE) which facilitates rapid integration of new tools and developing scalable and interoperable workflows for performing metabolomics data analysis. The environment can be launched on-demand on cloud resources and desktop computers. IT-expertise requirements on the user side are kept to a minimum, and workflows can be re-used effortlessly by any novice user. We validate our method in the field of metabolomics on two mass spectrometry, one nuclear magnetic resonance spectroscopy and one fluxomics study. We showed that the method scales dynamically with increasing availability of computational resources. We demonstrated that the method facilitates interoperability using integration of the major software suites resulting in a turn-key workflow encompassing all steps for mass-spectrometry-based metabolomics including preprocessing, statistics and identification. Microservices is a generic methodology that can serve any scientific discipline and opens up for new types of large-scale integrative science. AVAILABILITY AND IMPLEMENTATION: The PhenoMeNal consortium maintains a web portal (https://portal.phenomenal-h2020.eu) providing a GUI for launching the Virtual Research Environment. The GitHub repository https://github.com/phnmnl/ hosts the source code of all projects. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Payam Emami Khoonsari, Pablo A. Moreno, Sven Bergmann, Joachim Burman, Marco Capuccini, Matteo Carone, Marta Cascante, Pedro de Atauri, Carles Foguet, Alejandra N. González-Beltrán, Thomas Hankemeier, Kenneth Haug, Sijin He, Stephanie Herman, David Johnson 0006, Namrata Kale, Anders Larsson, Steffen Neumann, Kristian Peters, Luca Pireddu, Philippe Rocca-Serra, Pierrick Roger, Rico Rueedi, Christoph Ruttkies, Noureddin Sadawi, Reza M. Salek, Susanna-Assunta Sansone, Daniel Schober, Vitaly A. Selivanov, Etienne A. Thévenot, Michael van Vliet, Gianluigi Zanetti, Christoph Steinbeck, Kim Kultima, Ola Spjuth |
Bioinform. | 33 |
| 2017 | mzML2ISA & nmrML2ISA: generating enriched ISA-Tab metadata files from metabolomics XML dataabstractSUMMARY: Submission to the MetaboLights repository for metabolomics data currently places the burden of reporting instrument and acquisition parameters in ISA-Tab format on users, who have to do it manually, a process that is time consuming and prone to user input error. Since the large majority of these parameters are embedded in instrument raw data files, an opportunity exists to capture this metadata more accurately. Here we report a set of Python packages that can automatically generate ISA-Tab metadata file stubs from raw XML metabolomics data files. The parsing packages are separated into mzML2ISA (encompassing mzML and imzML formats) and nmrML2ISA (nmrML format only). Overall, the use of mzML2ISA & nmrML2ISA reduces the time needed to capture metadata substantially (capturing 90% of metadata on assay and sample levels), is much less prone to user input errors, improves compliance with minimum information reporting guidelines and facilitates more finely grained data exploration and querying of datasets. AVAILABILITY AND IMPLEMENTATION: mzML2ISA & nmrML2ISA are available under version 3 of the GNU General Public Licence at https://github.com/ISA-tools. Documentation is available from http://2isa.readthedocs.io/en/latest/. CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Martin Larralde, Thomas N. Lawson, Ralf J. M. Weber, Pablo A. Moreno, Kenneth Haug, Philippe Rocca-Serra, Mark R. Viant, Christoph Steinbeck, Reza M. Salek |
Bioinform. | 8 |
| 2016 | Reaction Decoder Tool (RDT): extracting features from chemical reactionsabstractUNLABELLED: Extracting chemical features like Atom-Atom Mapping (AAM), Bond Changes (BCs) and Reaction Centres from biochemical reactions helps us understand the chemical composition of enzymatic reactions. Reaction Decoder is a robust command line tool, which performs this task with high accuracy. It supports standard chemical input/output exchange formats i.e. RXN/SMILES, computes AAM, highlights BCs and creates images of the mapped reaction. This aids in the analysis of metabolic pathways and the ability to perform comparative studies of chemical reactions based on these features. AVAILABILITY AND IMPLEMENTATION: This software is implemented in Java, supported on Windows, Linux and Mac OSX, and freely available at https://github.com/asad/ReactionDecoder CONTACT: : [email protected] or [email protected]. Syed Asad Rahman, Gilliean Torrance, Lorenzo Baldacci, Sergio Martínez Cuesta, Franz Fenninger, Nimish Gopal, Saket Choudhary, John W. May, Gemma L. Holliday, Christoph Steinbeck, Janet M. Thornton |
Bioinform. | 10 |
| 2015 | BiNChE: A web tool and library for chemical enrichment analysis based on the ChEBI ontologyabstractBACKGROUND: Ontology-based enrichment analysis aids in the interpretation and understanding of large-scale biological data. Ontologies are hierarchies of biologically relevant groupings. Using ontology annotations, which link ontology classes to biological entities, enrichment analysis methods assess whether there is a significant over or under representation of entities for ontology classes. While many tools exist that run enrichment analysis for protein sets annotated with the Gene Ontology, there are only a few that can be used for small molecules enrichment analysis. RESULTS: We describe BiNChE, an enrichment analysis tool for small molecules based on the ChEBI Ontology. BiNChE displays an interactive graph that can be exported as a high-resolution image or in network formats. The tool provides plain, weighted and fragment analysis based on either the ChEBI Role Ontology or the ChEBI Structural Ontology. CONCLUSIONS: BiNChE aids in the exploration of large sets of small molecules produced within Metabolomics or other Systems Biology research contexts. The open-source tool provides easy and highly interactive web access to enrichment analysis with the ChEBI ontology tool and is additionally available as a standalone library. Pablo A. Moreno, Stephan Beisken, Bhavana Harsha, Venkatesh Muthukrishnan, Ilinca Tudose, Adriano Dekker, Stefanie Dornfeldt, Franziska Taruttis, Ivo Grosse, Janna Hastings, Steffen Neumann, Christoph Steinbeck |
BMC Bioinform. | 12 |
| 2014 | Building blocks for automated elucidation of metabolites: natural product-likeness for candidate rankingabstractBACKGROUND: In metabolomics experiments, spectral fingerprints of metabolites with no known structural identity are detected routinely. Computer-assisted structure elucidation (CASE) has been used to determine the structural identities of unknown compounds. It is generally accepted that a single 1D NMR spectrum or mass spectrum is usually not sufficient to establish the identity of a hitherto unknown compound. When a suite of spectra from 1D and 2D NMR experiments supplemented with a molecular formula are available, the successful elucidation of the chemical structure for candidates with up to 30 heavy atoms has been reported previously by one of the authors. In high-throughput metabolomics, usually 1D NMR or mass spectrometry experiments alone are conducted for rapid analysis of samples. This method subsequently requires that the spectral patterns are analyzed automatically to quickly identify known and unknown structures. In this study, we investigated whether additional existing knowledge, such as the fact that the unknown compound is a natural product, can be used to improve the ranking of the correct structure in the result list after the structure elucidation process. RESULTS: To identify unknowns using as little spectroscopic information as possible, we implemented an evolutionary algorithm-based CASE mechanism to elucidate candidates in a fully automated fashion, with input of the molecular formula and 13C NMR spectrum of the isolated compound. We also tested how filters like natural product-likeness, a measure that calculates the similarity of the compounds to known natural product space, might enhance the performance and quality of the structure elucidation. The evolutionary algorithm is implemented within the SENECA package for CASE reported previously, and is available for free download under artistic license at http://sourceforge.net/projects/seneca/. The natural product-likeness calculator is incorporated as a plugin within SENECA and is available as a GUI client and command-line executable. Significant improvements in candidate ranking were demonstrated for 41 small test molecules when the CASE system was supplemented by a natural product-likeness filter. CONCLUSIONS: In spectroscopically underdetermined structure elucidation problems, natural product-likeness can contribute to a better ranking of the correct structure in the results list. Kalai Vanii Jayaseelan, Christoph Steinbeck |
BMC Bioinform. | 2 |
| 2013 | Metingear: a development environment for annotating genome-scale metabolic modelsabstractUNLABELLED: Genome-scale metabolic models often lack annotations that would allow them to be used for further analysis. Previous efforts have focused on associating metabolites in the model with a cross reference, but this can be problematic if the reference is not freely available, multiple resources are used or the metabolite is added from a literature review. Associating each metabolite with chemical structure provides unambiguous identification of the components and a more detailed view of the metabolism. We have developed an open-source desktop application that simplifies the process of adding database cross references and chemical structures to genome-scale metabolic models. Annotated models can be exported to the Systems Biology Markup Language open interchange format. AVAILABILITY: Source code, binaries, documentation and tutorials are freely available at http://johnmay.github.com/metingear. The application is implemented in Java with bundles available for MS Windows and Macintosh OS X. John W. May, A. Gordon James, Christoph Steinbeck |
Bioinform. | 3 |
| 2013 | OntoQuery: easy-to-use web-based OWL queryingabstractSUMMARY: The Web Ontology Language (OWL) provides a sophisticated language for building complex domain ontologies and is widely used in bio-ontologies such as the Gene Ontology. The Protégé-OWL ontology editing tool provides a query facility that allows composition and execution of queries with the human-readable Manchester OWL syntax, with syntax checking and entity label lookup. No equivalent query facility such as the Protégé Description Logics (DL) query yet exists in web form. However, many users interact with bio-ontologies such as chemical entities of biological interest and the Gene Ontology using their online Web sites, within which DL-based querying functionality is not available. To address this gap, we introduce the OntoQuery web-based query utility. AVAILABILITY AND IMPLEMENTATION: The source code for this implementation together with instructions for installation is available at http://github.com/IlincaTudose/OntoQuery. OntoQuery software is fully compatible with all OWL-based ontologies and is available for download (CC-0 license). The ChEBI installation, ChEBI OntoQuery, is available at http://www.ebi.ac.uk/chebi/tools/ontoquery. CONTACT: [email protected]. Ilinca Tudose, Janna Hastings, Venkatesh Muthukrishnan, Gareth I. Owen, Steve Turner, Adriano Dekker, Namrata Kale, Marcus Ennis, Christoph Steinbeck |
Bioinform. | 9 |
| 2013 | KNIME-CDK: Workflow-driven CheminformaticsabstractBACKGROUND: Cheminformaticians have to routinely process and analyse libraries of small molecules. Among other things, that includes the standardization of molecules, calculation of various descriptors, visualisation of molecular structures, and downstream analysis. For this purpose, scientific workflow platforms such as the Konstanz Information Miner can be used if provided with the right plug-in. A workflow-based cheminformatics tool provides the advantage of ease-of-use and interoperability between complementary cheminformatics packages within the same framework, hence facilitating the analysis process. RESULTS: KNIME-CDK comprises functions for molecule conversion to/from common formats, generation of signatures, fingerprints, and molecular properties. It is based on the Chemistry Development Toolkit and uses the Chemical Markup Language for persistence. A comparison with the cheminformatics plug-in RDKit shows that KNIME-CDK supports a similar range of chemical classes and adds new functionality to the framework. We describe the design and integration of the plug-in, and demonstrate the usage of the nodes on ChEBI, a library of small molecules of biological interest. CONCLUSIONS: KNIME-CDK is an open-source plug-in for the Konstanz Information Miner, a free workflow platform. KNIME-CDK is build on top of the open-source Chemistry Development Toolkit and allows for efficient cross-vendor structural cheminformatics. Its ease-of-use and modularity enables researchers to automate routine tasks and data analysis, bringing complimentary cheminformatics functionality to the workflow environment. Stephan Beisken, Thorsten Meinl, Bernd Wiswedel, Luis F. de Figueiredo, Michael R. Berthold, Christoph Steinbeck |
BMC Bioinform. | 6 |
| 2013 | The Enzyme Portal: A case study in applying user-centred design methods in bioinformaticsabstractUser-centred design (UCD) is a type of user interface design in which the needs and desires of users are taken into account at each stage of the design process for a service or product; often for software applications and websites. Its goal is to facilitate the design of software that is both useful and easy to use. To achieve this, you must characterise users' requirements, design suitable interactions to meet their needs, and test your designs using prototypes and real life scenarios.For bioinformatics, there is little practical information available regarding how to carry out UCD in practice. To address this we describe a complete, multi-stage UCD process used for creating a new bioinformatics resource for integrating enzyme information, called the Enzyme Portal (http://www.ebi.ac.uk/enzymeportal). This freely-available service mines and displays data about proteins with enzymatic activity from public repositories via a single search, and includes biochemical reactions, biological pathways, small molecule chemistry, disease information, 3D protein structures and relevant scientific literature.We employed several UCD techniques, including: persona development, interviews, 'canvas sort' card sorting, user workflows, usability testing and others. Our hope is that this case study will motivate the reader to apply similar UCD approaches to their own software design for bioinformatics. Indeed, we found the benefits included more effective decision-making for design ideas and technologies; enhanced team-working and communication; cost effectiveness; and ultimately a service that more closely meets the needs of our target audience. Paula de Matos, Jennifer A. Cham, Rafael Alcántara, Francis Rowland, Rodrigo Lopez, Christoph Steinbeck |
BMC Bioinform. | 7 |
| 2012 | Self-organizing ontology of biochemically relevant small moleculesabstractBACKGROUND: The advent of high-throughput experimentation in biochemistry has led to the generation of vast amounts of chemical data, necessitating the development of novel analysis, characterization, and cataloguing techniques and tools. Recently, a movement to publically release such data has advanced biochemical structure-activity relationship research, while providing new challenges, the biggest being the curation, annotation, and classification of this information to facilitate useful biochemical pattern analysis. Unfortunately, the human resources currently employed by the organizations supporting these efforts (e.g. ChEBI) are expanding linearly, while new useful scientific information is being released in a seemingly exponential fashion. Compounding this, currently existing chemical classification and annotation systems are not amenable to automated classification, formal and transparent chemical class definition axiomatization, facile class redefinition, or novel class integration, thus further limiting chemical ontology growth by necessitating human involvement in curation. Clearly, there is a need for the automation of this process, especially for novel chemical entities of biological interest. RESULTS: To address this, we present a formal framework based on Semantic Web technologies for the automatic design of chemical ontology which can be used for automated classification of novel entities. We demonstrate the automatic self-assembly of a structure-based chemical ontology based on 60 MeSH and 40 ChEBI chemical classes. This ontology is then used to classify 200 compounds with an accuracy of 92.7%. We extend these structure-based classes with molecular feature information and demonstrate the utility of our framework for classification of functionally relevant chemicals. Finally, we discuss an iterative approach that we envision for future biochemical ontology development. CONCLUSIONS: We conclude that the proposed methodology can ease the burden of chemical data annotators and dramatically increase their productivity. We anticipate that the use of formal logic in our proposed framework will make chemical classification criteria more transparent to humans and machines alike and will thus facilitate predictive and integrative bioactivity model development. Leonid L. Chepelev, Janna Hastings, Marcus Ennis, Christoph Steinbeck, Michel Dumontier |
BMC Bioinform. | 4 |
| 2012 | Natural product-likeness score revisited: an open-source, open-data implementationabstractBACKGROUND: Natural product-likeness of a molecule, i.e. similarity of this molecule to the structure space covered by natural products, is a useful criterion in screening compound libraries and in designing new lead compounds. A closed source implementation of a natural product-likeness score, that finds its application in virtual screening, library design and compound selection, has been previously reported by one of us. In this note, we report an open-source and open-data re-implementation of this scoring system, illustrate its efficiency in ranking small molecules for natural product likeness and discuss its potential applications. RESULTS: The Natural-Product-Likeness scoring system is implemented as Taverna 2.2 workflows, and is available under Creative Commons Attribution-Share Alike 3.0 Unported License at http://www.myexperiment.org/packs/183.html. It is also available for download as executable standalone java package from http://sourceforge.net/projects/np-likeness/under Academic Free License. CONCLUSIONS: Our open-source, open-data Natural-Product-Likeness scoring system can be used as a filter for metabolites in Computer Assisted Structure Elucidation or to select natural-product-like molecules from molecular libraries for the use as leads in drug discovery. Kalai Vanii Jayaseelan, Pablo A. Moreno, Andreas Truszkowski, Peter Ertl, Christoph Steinbeck |
BMC Bioinform. | 5 |
| 2012 | Bioinformatics Meets User-Centred Design: A PerspectiveabstractDesigners have a saying that "the joy of an early release lasts but a short time. The bitterness of an unusable system lasts for years." It is indeed disappointing to discover that your data resources are not being used to their full potential. Not only have you invested your time, effort, and research grant on the project, but you may face costly redesigns if you want to improve the system later. This scenario would be less likely if the product was designed to provide users with exactly what they need, so that it is fit for purpose before its launch. We work at EMBL-European Bioinformatics Institute (EMBL-EBI), and we consult extensively with life science researchers to find out what they need from biological data resources. We have found that although users believe that the bioinformatics community is providing accurate and valuable data, they often find the interfaces to these resources tricky to use and navigate. We believe that if you can find out what your users want even before you create the first mock-up of a system, the final product will provide a better user experience. This would encourage more people to use the resource and they would have greater access to the data, which could ultimately lead to more scientific discoveries. In this paper, we explore the need for a user-centred design (UCD) strategy when designing bioinformatics resources and illustrate this with examples from our work at EMBL-EBI. Our aim is to introduce the reader to how selected UCD techniques may be successfully applied to software design for bioinformatics. Katrina Pavelin, Jennifer A. Cham, Paula de Matos, Catherine Brooksbank, Graham Cameron, Christoph Steinbeck |
PLoS Comput. Biol. | 6 |
| 2010 | Ontological dependence, dispositions and institutional reality in chemistryabstractBiochemical ‘small molecules’ are involved in all living processes across all biological domains. Chemical ontologies provide structured chemical data and thereby support cross-disciplinary and integrative research across systems biology, chemogenomics and metabolomics. Efforts are underway to align ChEBI with upper level ontologies such as BFO, but confusion persists as to the ontological status of the ChEBI ‘role’ entities, which refer to continuants which inhere in chemical entities by virtue of the activity of the chemical entities. We provide a formal classification of these ‘role’ entities according to the continuants and endurants on which they ontologically depend, discuss the nature of chemical dispositions and the relevance of institutional reality, and address granularity issues in modelling chemical activity. Colin R. Batchelor, Janna Hastings, Christoph Steinbeck |
FOIS | 3 |
| 2010 | What are chemical structures and their relations?abstractIn chemistry, advances in computational technologies have allowed research into molecules that have not been synthesized yet, and may never be, to become widespread. These are described in terms of their structures, which are expressed as chemical graphs. Chemical graphs are a representational artifact, and as such are of a different ontological nature than the molecular entities which they describe. Janna Hastings, Colin R. Batchelor, Christoph Steinbeck, Stefan Schulz 0001 |
FOIS | 3 |
| 2010 | CDK-Taverna: an open workflow environment for cheminformaticsabstractBACKGROUND: Small molecules are of increasing interest for bioinformatics in areas such as metabolomics and drug discovery. The recent release of large open access chemistry databases generates a demand for flexible tools to process them and discover new knowledge. To freely support open science based on these data resources, it is desirable for the processing tools to be open source and available for everyone. RESULTS: Here we describe a novel combination of the workflow engine Taverna and the cheminformatics library Chemistry Development Kit (CDK) resulting in a open source workflow solution for cheminformatics. We have implemented more than 160 different workers to handle specific cheminformatics tasks. We describe the applications of CDK-Taverna in various usage scenarios. CONCLUSIONS: The combination of the workflow engine Taverna and the Chemistry Development Kit provides the first open source cheminformatics workflow solution for the biosciences. With the Taverna-community working towards a more powerful workflow engine and a more user-friendly user interface, CDK-Taverna has the potential to become a free alternative to existing proprietary workflow tools. Egon L. Willighagen, Achim Zielesny, Christoph Steinbeck |
BMC Bioinform. | 4 |
| 2009 | Bioclipse 2: A scriptable integration platform for the life sciencesabstractBACKGROUND: Contemporary biological research integrates neighboring scientific domains to answer complex questions in fields such as systems biology and drug discovery. This calls for tools that are intuitive to use, yet flexible to adapt to new tasks. RESULTS: Bioclipse is a free, open source workbench with advanced features for the life sciences. Version 2.0 constitutes a complete rewrite of Bioclipse, and delivers a stable, scalable integration platform for developers and an intuitive workbench for end users. All functionality is available both from the graphical user interface and from a built-in novel domain-specific language, supporting the scientist in interdisciplinary research and reproducible analyses through advanced visualization of the inputs and the results. New components for Bioclipse 2 include a rewritten editor for chemical structures, a table for multiple molecules that supports gigabyte-sized files, as well as a graphical editor for sequences and alignments. CONCLUSION: Bioclipse 2 is equipped with advanced tools required to carry out complex analysis in the fields of bio- and cheminformatics. Developed as a Rich Client based on Eclipse, Bioclipse 2 leverages on today's powerful desktop computers for providing a responsive user interface, but also takes full advantage of the Web and networked (Web/Cloud) services for more demanding calculations or retrieval of data. The fact that Bioclipse 2 is based on an advanced and widely used service platform ensures wide extensibility, making it easy to add new algorithms, visualizations, as well as scripting commands. The intuitive tools for end users and the extensible architecture make Bioclipse 2 ideal for interdisciplinary and integrative research.Bioclipse 2 is released under the Eclipse Public License (EPL), a flexible open source license that allows additional plugins to be of any license. Bioclipse 2 is implemented in Java and supported on all major platforms; Source code and binaries are freely available at http://www.bioclipse.net. Ola Spjuth, Jonathan Alvarsson, Arvid Berg, Martin Eklund, Stefan Kuhn 0001, Carl Mäsak, Gilleain M. Torrance, Johannes Wagener, Egon L. Willighagen, Christoph Steinbeck, Jarl E. S. Wikberg |
BMC Bioinform. | 10 |
| 2008 | Building blocks for automated elucidation of metabolites: Machine learning methods for NMR predictionabstractBACKGROUND: Current efforts in Metabolomics, such as the Human Metabolome Project, collect structures of biological metabolites as well as data for their characterisation, such as spectra for identification of substances and measurements of their concentration. Still, only a fraction of existing metabolites and their spectral fingerprints are known. Computer-Assisted Structure Elucidation (CASE) of biological metabolites will be an important tool to leverage this lack of knowledge. Indispensable for CASE are modules to predict spectra for hypothetical structures. This paper evaluates different statistical and machine learning methods to perform predictions of proton NMR spectra based on data from our open database NMRShiftDB. RESULTS: A mean absolute error of 0.18 ppm was achieved for the prediction of proton NMR shifts ranging from 0 to 11 ppm. Random forest, J48 decision tree and support vector machines achieved similar overall errors. HOSE codes being a notably simple method achieved a comparatively good result of 0.17 ppm mean absolute error. CONCLUSION: NMR prediction methods applied in the course of this work delivered precise predictions which can serve as a building block for Computer-Assisted Structure Elucidation for biological metabolites. Stefan Kuhn 0001, Björn Egert, Steffen Neumann, Christoph Steinbeck |
BMC Bioinform. | 4 |
| 2007 | Bioclipse: an open source workbench for chemo- and bioinformaticsabstractBACKGROUND: There is a need for software applications that provide users with a complete and extensible toolkit for chemo- and bioinformatics accessible from a single workbench. Commercial packages are expensive and closed source, hence they do not allow end users to modify algorithms and add custom functionality. Existing open source projects are more focused on providing a framework for integrating existing, separately installed bioinformatics packages, rather than providing user-friendly interfaces. No open source chemoinformatics workbench has previously been published, and no successful attempts have been made to integrate chemo- and bioinformatics into a single framework. RESULTS: Bioclipse is an advanced workbench for resources in chemo- and bioinformatics, such as molecules, proteins, sequences, spectra, and scripts. It provides 2D-editing, 3D-visualization, file format conversion, calculation of chemical properties, and much more; all fully integrated into a user-friendly desktop application. Editing supports standard functions such as cut and paste, drag and drop, and undo/redo. Bioclipse is written in Java and based on the Eclipse Rich Client Platform with a state-of-the-art plugin architecture. This gives Bioclipse an advantage over other systems as it can easily be extended with functionality in any desired direction. CONCLUSION: Bioclipse is a powerful workbench for bio- and chemoinformatics as well as an advanced integration platform. The rich functionality, intuitive user interface, and powerful plugin architecture make Bioclipse the most advanced and user-friendly open source workbench for chemo- and bioinformatics. Bioclipse is released under Eclipse Public License (EPL), an open source license which sets no constraints on external plugin licensing; it is totally open for both open source plugins as well as commercial ones. Bioclipse is freely available at http://www.bioclipse.net. Ola Spjuth, Tobias Helmus, Egon L. Willighagen, Stefan Kuhn 0001, Martin Eklund, Johannes Wagener, Peter Murray-Rust, Christoph Steinbeck, Jarl E. S. Wikberg |
BMC Bioinform. | 8 |
| 2007 | Userscripts for the Life SciencesabstractBACKGROUND: The web has seen an explosion of chemistry and biology related resources in the last 15 years: thousands of scientific journals, databases, wikis, blogs and resources are available with a wide variety of types of information. There is a huge need to aggregate and organise this information. However, the sheer number of resources makes it unrealistic to link them all in a centralised manner. Instead, search engines to find information in those resources flourish, and formal languages like Resource Description Framework and Web Ontology Language are increasingly used to allow linking of resources. A recent development is the use of userscripts to change the appearance of web pages, by on-the-fly modification of the web content. This opens possibilities to aggregate information and computational results from different web resources into the web page of one of those resources. RESULTS: Several userscripts are presented that enrich biology and chemistry related web resources by incorporating or linking to other computational or data sources on the web. The scripts make use of Greasemonkey-like plugins for web browsers and are written in JavaScript. Information from third-party resources are extracted using open Application Programming Interfaces, while common Universal Resource Locator schemes are used to make deep links to related information in that external resource. The userscripts presented here use a variety of techniques and resources, and show the potential of such scripts. CONCLUSION: This paper discusses a number of userscripts that aggregate information from two or more web resources. Examples are shown that enrich web pages with information from other resources, and show how information from web pages can be used to link to, search, and process information in other resources. Due to the nature of userscripts, scientists are able to select those scripts they find useful on a daily basis, as the scripts run directly in their own web browser rather than on the web server. This flexibility allows the scientists to tune the features of web resources to optimise their productivity. Egon L. Willighagen, Noel M. O'Boyle, Harini Gopalakrishnan, Dazhi Jiao, Rajarshi Guha, Christoph Steinbeck, David J. Wild 0001 |
BMC Bioinform. | 6 |