EDBT 2026 Demo / reviewers in the wild / expert
Oliver Kohlbacher
dblp:63/2396
· DBLP profile ↗
55ranked-venue papers
3as first author
6since 2021 · last 2026
0000-0003-1739-4598ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 48 · 3 first-author · 6 since 2021Systems, architecture and hardware · 5Graphics, computer vision, multimedia, augmented reality and games · 1Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Interacting Species Database (ISDB): comprehensive resource for interspecies interactions at the molecular levelabstractMOTIVATION: Organisms within ecological systems often engage in molecular interactions that mediate key biological processes, such as protein-protein interactions involved in host-pathogen recognition and symbiosis. Characterization of these interactions at a molecular level is essential for understanding the mechanistic, evolutionary, and functional basis of interspecies interactions, as well as for informing potential therapeutic interventions. However, progress in this field is significantly impeded by the lack of a comprehensive database of interacting species at molecular resolution and the limited availability of experimental data. RESULTS: We introduce the Interacting Species Database (ISDB), a comprehensive resource that catalogs interspecies interactions, annotated with NCBI taxonomic identifiers, interaction types and known molecular interactions. The ISDB encompasses 858 229 interacting species pairs and 171 713 interspecies protein-protein interactions within 261 287 organisms. ISDB is designed to support researchers in searching for, downloading, and depositing interspecies interaction data, which facilitates the study of ecological dynamics across diverse research domains. AVAILABILITY AND IMPLEMENTATION: The ISDB is available via a web interface (https://www.elhabashylab.org/isdb), open-source code on GitHub (https://github.com/ElhabashyLab/ISDB) under the MIT license and is archived on Zenodo (Version v1.0.1, DOI: 10.5281/zenodo.20162385). Michael Mederer, Anupam Gautam, Oliver Kohlbacher, Andrei N. Lupas, Hadeer Elhabashy |
Bioinform. | 3 |
| 2024 | CLAUDIO: automated structural analysis of cross-linking dataabstractMOTIVATION: Cross-linking mass spectrometry has made remarkable advancements in the high-throughput characterization of protein structures and interactions. The resulting pairs of cross-linked peptides typically require geometric assessment and validation, given the availability of their corresponding structures. RESULTS: CLAUDIO (Cross-linking Analysis Using Distances and Overlaps) is an open-source software tool designed for the automated analysis and validation of different varieties of large-scale cross-linking experiments. Many of the otherwise manual processes for structural validation (i.e. structure retrieval and mapping) are performed fully automatically to simplify and accelerate the data interpretation process. In addition, CLAUDIO has the ability to remap intra-protein links as inter-protein links and discover evidence for homo-multimers. AVAILABILITY AND IMPLEMENTATION: CLAUDIO is available as open-source software under the MIT license at https://github.com/KohlbacherLab/CLAUDIO. Alexander Röhl, Eugen Netz, Oliver Kohlbacher, Hadeer Elhabashy |
Bioinform. | 3 |
| 2023 | The personalized cancer network explorer (PeCaX) as a visual analytics tool to support molecular tumor boardsabstractBACKGROUND: Personalized oncology represents a shift in cancer treatment from conventional methods to target specific therapies where the decisions are made based on the patient specific tumor profile. Selection of the optimal therapy relies on a complex interdisciplinary analysis and interpretation of these variants by experts in molecular tumor boards. With up to hundreds of somatic variants identified in a tumor, this process requires visual analytics tools to guide and accelerate the annotation process. RESULTS: The Personal Cancer Network Explorer (PeCaX) is a visual analytics tool supporting the efficient annotation, navigation, and interpretation of somatic genomic variants through functional annotation, drug target annotation, and visual interpretation within the context of biological networks. Starting with somatic variants in a VCF file, PeCaX enables users to explore these variants through a web-based graphical user interface. The most protruding feature of PeCaX is the combination of clinical variant annotation and gene-drug networks with an interactive visualization. This reduces the time and effort the user needs to invest to get to a treatment suggestion and helps to generate new hypotheses. PeCaX is being provided as a platform-independent containerized software package for local or institution-wide deployment. PeCaX is available for download at https://github.com/KohlbacherLab/PeCaX-docker . Mirjam Figaschewski, Bilge Sürün, Thorsten Tiede, Oliver Kohlbacher |
BMC Bioinform. | 4 |
| 2022 | Efficient privacy-preserving whole-genome variant queriesabstractMOTIVATION: Diagnosis and treatment decisions on genomic data have become widespread as the cost of genome sequencing decreases gradually. In this context, disease-gene association studies are of great importance. However, genomic data are very sensitive when compared to other data types and contains information about individuals and their relatives. Many studies have shown that this information can be obtained from the query-response pairs on genomic databases. In this work, we propose a method that uses secure multi-party computation to query genomic databases in a privacy-protected manner. The proposed solution privately outsources genomic data from arbitrarily many sources to the two non-colluding proxies and allows genomic databases to be safely stored in semi-honest cloud environments. It provides data privacy, query privacy and output privacy by using XOR-based sharing and unlike previous solutions, it allows queries to run efficiently on hundreds of thousands of genomic data. RESULTS: We measure the performance of our solution with parameters similar to real-world applications. It is possible to query a genomic database with 3 000 000 variants with five genomic query predicates under 400 ms. Querying 1 048 576 genomes, each containing 1 000 000 variants, for the presence of five different query variants can be achieved approximately in 6 min with a small amount of dedicated hardware and connectivity. These execution times are in the right range to enable real-world applications in medical research and healthcare. Unlike previous studies, it is possible to query multiple databases with response times fast enough for practical application. To the best of our knowledge, this is the first solution that provides this performance for querying large-scale genomic data. AVAILABILITY AND IMPLEMENTATION: https://gitlab.com/DIFUTURE/privacy-preserving-variant-queries. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Mete Akgün, Nico Pfeifer, Oliver Kohlbacher |
Bioinform. | 3 |
| 2022 | De novo identification of maximally deregulated subnetworks based on multi-omics data with DeRegNetabstractBACKGROUND: With a growing amount of (multi-)omics data being available, the extraction of knowledge from these datasets is still a difficult problem. Classical enrichment-style analyses require predefined pathways or gene sets that are tested for significant deregulation to assess whether the pathway is functionally involved in the biological process under study. De novo identification of these pathways can reduce the bias inherent in predefined pathways or gene sets. At the same time, the definition and efficient identification of these pathways de novo from large biological networks is a challenging problem. RESULTS: We present a novel algorithm, DeRegNet, for the identification of maximally deregulated subnetworks on directed graphs based on deregulation scores derived from (multi-)omics data. DeRegNet can be interpreted as maximum likelihood estimation given a certain probabilistic model for de-novo subgraph identification. We use fractional integer programming to solve the resulting combinatorial optimization problem. We can show that the approach outperforms related algorithms on simulated data with known ground truths. On a publicly available liver cancer dataset we can show that DeRegNet can identify biologically meaningful subgraphs suitable for patient stratification. DeRegNet can also be used to find explicitly multi-omics subgraphs which we demonstrate by presenting subgraphs with consistent methylation-transcription patterns. DeRegNet is freely available as open-source software. CONCLUSION: The proposed algorithmic framework and its available implementation can serve as a valuable heuristic hypothesis generation tool contextualizing omics data within biomolecular networks. Sebastian Winkler, Ivana Winkler, Mirjam Figaschewski, Thorsten Tiede, Alfred Nordheim, Oliver Kohlbacher |
BMC Bioinform. | 6 |
| 2021 | Identifying disease-causing mutations with privacy protectionabstractMOTIVATION: The use of genome data for diagnosis and treatment is becoming increasingly common. Researchers need access to as many genomes as possible to interpret the patient genome, to obtain some statistical patterns and to reveal disease-gene relationships. The sensitive information contained in the genome data and the high risk of re-identification increase the privacy and security concerns associated with sharing such data. In this article, we present an approach to identify disease-associated variants and genes while ensuring patient privacy. The proposed method uses secure multi-party computation to find disease-causing mutations under specific inheritance models without sacrificing the privacy of individuals. It discloses only variants or genes obtained as a result of the analysis. Thus, the vast majority of patient data can be kept private. RESULTS: Our prototype implementation performs analyses on thousands of genomic data in milliseconds, and the runtime scales logarithmically with the number of patients. We present the first inheritance model (recessive, dominant and compound heterozygous) based privacy-preserving analyses of genomic data to find disease-causing mutations. Furthermore, we re-implement the privacy-preserving methods (MAX, SETDIFF and INTERSECTION) proposed in a previous study. Our MAX, SETDIFF and INTERSECTION implementations are 2.5, 1122 and 341 times faster than the corresponding operations of the state-of-the-art protocol, respectively. AVAILABILITY AND IMPLEMENTATION: https://gitlab.com/DIFUTURE/privacy-preserving-genomic-diagnosis. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Mete Akgün, Ali Burak Ünal, Bekir Ergüner, Nico Pfeifer, Oliver Kohlbacher |
Bioinform. | 5 |
| 2020 | ClinVAP: a reporting strategy from variants to therapeutic optionsabstractMOTIVATION: Next-generation sequencing has become routine in oncology and opens up new avenues of therapies, particularly in personalized oncology setting. An increasing number of cases also implies a need for a more robust, automated and reproducible processing of long lists of variants for cancer diagnosis and therapy. While solutions for the large-scale analysis of somatic variants have been implemented, existing solutions often have issues with reproducibility, scalability and interoperability. RESULTS: Clinical Variant Annotation Pipeline (ClinVAP) is an automated pipeline which annotates, filters and prioritizes somatic single nucleotide variants provided in variant call format. It augments the variant information with documented or predicted clinical effect. These annotated variants are prioritized based on driver gene status and druggability. ClinVAP is available as a fully containerized, self-contained pipeline maximizing reproducibility and scalability allowing the analysis of larger scale data. The resulting JSON-based report is suited for automated downstream processing, but ClinVAP can also automatically render the information into a user-defined template to yield a human-readable report. AVAILABILITY AND IMPLEMENTATION: ClinVAP is available at https://github.com/PersonalizedOncology/ClinVAP. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Bilge Sürün, Charlotta Schärfe, Mathew R. Divine, Julian Heinrich, Nora C. Toussaint, Lukas Zimmermann, Janina Beha, Oliver Kohlbacher |
Bioinform. | 8 |
| 2019 | Biological Network Knowledge in Molecular Tumor Boards
Thorsten Tiede, Oliver Kohlbacher |
AMIA | 2 |
| 2019 | Reproducible Scientific Workflows for High Performance and Cloud ComputingabstractMany complex data analysis tasks are performed by scientific workflows and pipelines deployed on high performance computing (HPC) or cloud computing resources. The complex software stack required by a workflow and unnoticed dependencies can make the deployment of a pipeline a demanding task. Once deployed, workflows tend to be black boxes, especially for users that did not create the pipeline themselves. At the end of a project a researcher should archive the pipeline in order to ensure reproducibility of published results. This paper illustrates a possible solution for each of the three tasks: reproducible deployment via software containers, automated generation of provenance information to break black boxes, and using the CiTAR service for archiving software containers. Felix Bartusch, Maximilian Hanussek, Jens Krüger 0002, Oliver Kohlbacher |
CCGRID | 4 |
| 2019 | BOOTABLE: Bioinformatics Benchmark Tool SuiteabstractThe interest in analyzing biological data on a large scale has grown over the last years. Bioinformatic applications play an important role when it comes to the analysis of huge amounts of data. Due to the large amount of biological data and/or large problem spaces a considerable amount of computing resources is required to answer the raised research questions. In order to estimate which underlying hardware might be the most suitable for the bioinformatic tools applied, a well-defined benchmark suite is required. Such a benchmark suite can get useful in the case of purchasing hardware and even further for larger projects with the goal to establish a bioinformatics compute infrastructure. With this paper we present BOOTABLE, our bioinformatic benchmark suite. BOOTABLE currently contains six popular and widely used bioinformatic applications representing a broad spectrum of usage characteristics. It further includes an automated installation procedure and all required datasets. BOOTABLE is available from our Github repository (https://github.com/MaximilianHanussek/BOOTABLE) in various formats. Maximilian Hanussek, Felix Bartusch, Jens Krüger 0002, Oliver Kohlbacher |
CCGRID | 4 |
| 2019 | ClinOmicsTrailbc: a visual analytics tool for breast cancer treatment stratificationabstractMOTIVATION: Breast cancer is the second leading cause of cancer death among women. Tumors, even of the same histopathological subtype, exhibit a high genotypic diversity that impedes therapy stratification and that hence must be accounted for in the treatment decision-making process. RESULTS: Here, we present ClinOmicsTrailbc, a comprehensive visual analytics tool for breast cancer decision support that provides a holistic assessment of standard-of-care targeted drugs, candidates for drug repositioning and immunotherapeutic approaches. To this end, our tool analyzes and visualizes clinical markers and (epi-)genomics and transcriptomics datasets to identify and evaluate the tumor's main driver mutations, the tumor mutational burden, activity patterns of core cancer-relevant pathways, drug-specific biomarkers, the status of molecular drug targets and pharmacogenomic influences. In order to demonstrate ClinOmicsTrailbc's rich functionality, we present three case studies highlighting various ways in which ClinOmicsTrailbc can support breast cancer precision medicine. ClinOmicsTrailbc is a powerful integrated visual analytics tool for breast cancer research in general and for therapy stratification in particular, assisting oncologists to find the best possible treatment options for their breast cancer patients based on actionable, evidence-based results. AVAILABILITY AND IMPLEMENTATION: ClinOmicsTrailbc can be freely accessed at https://clinomicstrail.bioinf.uni-sb.de. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Lara Schneider, Tim Kehl, Kristina Thedinga, Nadja Liddy Grammes, Christina Backes, Christopher Mohr, Benjamin Schubert, Kerstin Lenhof, Nico Gerstner, Andreas Daniel Hartkopf, Markus Wallwiener, Oliver Kohlbacher, Andreas Keller, Eckart Meese, Norbert Graf 0001, Hans-Peter Lenhof |
Bioinform. | 12 |
| 2018 | Population-specific design of de-immunized protein biotherapeuticsabstractImmunogenicity is a major problem during the development of biotherapeutics since it can lead to rapid clearance of the drug and adverse reactions.The challenge for biotherapeutic design is therefore to identify mutants of the protein sequence that minimize immunogenicity in a target population whilst retaining pharmaceutical activity and protein function.Current approaches are moderately successful in designing sequences with reduced immunogenicity, but do not account for the varying frequencies of different human leucocyte antigen alleles in a specific population and in addition, since many designs are non-functional, require costly experimental post-screening.Here, we report a new method for de-immunization design using multi-objective combinatorial optimization.The method simultaneously optimizes the likelihood of a functional protein sequence at the same time as minimizing its immunogenicity tailored to a target population.We bypass the need for three-dimensional protein structure or molecular simulations to identify functional designs by automatically generating sequences using probabilistic models that have been used previously for mutation effect prediction and structure prediction.As proof-of-principle we designed sequences of the C2 domain of Factor VIII and tested them experimentally, resulting in a good correlation with the predicted immunogenicity of our model. Author summaryTherapeutic proteins have become an important area of pharmaceutical research and have been successfully applied to treat many diseases in the last decades.However, biotherapeutics suffer from the formation of anti-drug antibodies, which can reduce the efficacy of the drug or even result in severe adverse effects.A main contributor to the antibody formation is a T-cell mediated immune reaction caused by presentation of small immunogenic peptides derived from the biotherapeutic.Targeting these peptides via Benjamin Schubert, Charlotta Schärfe, Pierre Dönnes, Thomas A. Hopf, Debora S. Marks, Oliver Kohlbacher |
PLoS Comput. Biol. | 6 |
| 2017 | SANDPUMA: ensemble predictions of nonribosomal peptide chemistry reveal biosynthetic diversity across ActinobacteriaabstractSUMMARY: Nonribosomally synthesized peptides (NRPs) are natural products with widespread applications in medicine and biotechnology. Many algorithms have been developed to predict the substrate specificities of nonribosomal peptide synthetase adenylation (A) domains from DNA sequences, which enables prioritization and dereplication, and integration with other data types in discovery efforts. However, insufficient training data and a lack of clarity regarding prediction quality have impeded optimal use. Here, we introduce prediCAT, a new phylogenetics-inspired algorithm, which quantitatively estimates the degree of predictability of each A-domain. We then systematically benchmarked all algorithms on a newly gathered, independent test set of 434 A-domain sequences, showing that active-site-motif-based algorithms outperform whole-domain-based methods. Subsequently, we developed SANDPUMA, a powerful ensemble algorithm, based on newly trained versions of all high-performing algorithms, which significantly outperforms individual methods. Finally, we deployed SANDPUMA in a systematic investigation of 7635 Actinobacteria genomes, suggesting that NRP chemical diversity is much higher than previously estimated. SANDPUMA has been integrated into the widely used antiSMASH biosynthetic gene cluster analysis pipeline and is also available as an open-source, standalone tool. AVAILABILITY AND IMPLEMENTATION: SANDPUMA is freely available at https://bitbucket.org/chevrm/sandpuma and as a docker image at https://hub.docker.com/r/chevrm/sandpuma/ under the GNU Public License 3 (GPL3). CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Marc G. Chevrette, Fabian Aicheler, Oliver Kohlbacher, Cameron R. Currie, Marnix H. Medema |
Bioinform. | 3 |
| 2017 | ImmunoNodes - graphical development of complex immunoinformatics workflowsabstractBACKGROUND: Immunoinformatics has become a crucial part in biomedical research. Yet many immunoinformatics tools have command line interfaces only and can be difficult to install. Web-based immunoinformatics tools, on the other hand, are difficult to integrate with other tools, which is typically required for the complex analysis and prediction pipelines required for advanced applications. RESULT: We present ImmunoNodes, an immunoinformatics toolbox that is fully integrated into the visual workflow environment KNIME. By dragging and dropping tools and connecting them to indicate the data flow through the pipeline, it is possible to construct very complex workflows without the need for coding. CONCLUSION: ImmunoNodes allows users to build complex workflows with an easy to use and intuitive interface with a few clicks on any desktop computer. Benjamin Schubert, Luis de la Garza, Christopher Mohr, Mathias Walzer, Oliver Kohlbacher |
BMC Bioinform. | 5 |
| 2016 | BALL-SNPgp - from genetic variants toward computational diagnosticsabstractUNLABELLED: In medical research, it is crucial to understand the functional consequences of genetic alterations, for example, non-synonymous single nucleotide variants (nsSNVs). NsSNVs are known to be causative for several human diseases. However, the genetic basis of complex disorders such as diabetes or cancer comprises multiple factors. Methods to analyze putative synergetic effects of multiple such factors, however, are limited. Here, we concentrate on nsSNVs and present BALL-SNPgp, a tool for structural and functional characterization of nsSNVs, which is aimed to improve pathogenicity assessment in computational diagnostics. Based on annotated SNV data, BALL-SNPgp creates a three-dimensional visualization of the encoded protein, collects available information from different resources concerning disease relevance and other functional annotations, performs cluster analysis, predicts putative binding pockets and provides data on known interaction sites. AVAILABILITY AND IMPLEMENTATION: BALL-SNPgp is based on the comprehensive C ++ framework Biochemical Algorithms Library (BALL) and its visualization front-end BALLView. Our tool is available at www.ccb.uni-saarland.de/BALL-SNPgp CONTACT: [email protected]. Sabine C. Mueller, Christina Backes, Alexander Greß, Nina Baumgarten, Olga V. Kalinina, Andreas Moll, Oliver Kohlbacher, Eckart Meese, Andreas Keller |
Bioinform. | 7 |
| 2016 | FRED 2: an immunoinformatics framework for PythonabstractUNLABELLED: Immunoinformatics approaches are widely used in a variety of applications from basic immunological to applied biomedical research. Complex data integration is inevitable in immunological research and usually requires comprehensive pipelines including multiple tools and data sources. Non-standard input and output formats of immunoinformatics tools make the development of such applications difficult. Here we present FRED 2, an open-source immunoinformatics framework offering easy and unified access to methods for epitope prediction and other immunoinformatics applications. FRED 2 is implemented in Python and designed to be extendable and flexible to allow rapid prototyping of complex applications. AVAILABILITY AND IMPLEMENTATION: FRED 2 is available at http://fred-2.github.io CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Benjamin Schubert, Mathias Walzer, Hans-Philipp Brachvogel, András Szolek, Christopher Mohr, Oliver Kohlbacher |
Bioinform. | 6 |
| 2016 | From the desktop to the grid: scalable bioinformatics via workflow conversionabstractBACKGROUND: Reproducibility is one of the tenets of the scientific method. Scientific experiments often comprise complex data flows, selection of adequate parameters, and analysis and visualization of intermediate and end results. Breaking down the complexity of such experiments into the joint collaboration of small, repeatable, well defined tasks, each with well defined inputs, parameters, and outputs, offers the immediate benefit of identifying bottlenecks, pinpoint sections which could benefit from parallelization, among others. Workflows rest upon the notion of splitting complex work into the joint effort of several manageable tasks. There are several engines that give users the ability to design and execute workflows. Each engine was created to address certain problems of a specific community, therefore each one has its advantages and shortcomings. Furthermore, not all features of all workflow engines are royalty-free -an aspect that could potentially drive away members of the scientific community. RESULTS: We have developed a set of tools that enables the scientific community to benefit from workflow interoperability. We developed a platform-free structured representation of parameters, inputs, outputs of command-line tools in so-called Common Tool Descriptor documents. We have also overcome the shortcomings and combined the features of two royalty-free workflow engines with a substantial user community: the Konstanz Information Miner, an engine which we see as a formidable workflow editor, and the Grid and User Support Environment, a web-based framework able to interact with several high-performance computing resources. We have thus created a free and highly accessible way to design workflows on a desktop computer and execute them on high-performance computing resources. CONCLUSIONS: Our work will not only reduce time spent on designing scientific workflows, but also make executing workflows on remote high-performance computing resources more accessible to technically inexperienced users. We strongly believe that our efforts not only decrease the turnaround time to obtain scientific results but also have a positive impact on reproducibility, thus elevating the quality of obtained scientific results. Luis de la Garza, Johannes Veit, András Szolek, Marc Röttig, Stephan Aiche, Sandra Gesing, Knut Reinert, Oliver Kohlbacher |
BMC Bioinform. | 8 |
| 2016 | Learning from Heterogeneous Data Sources: An Application in Spatial ProteomicsabstractSub-cellular localisation of proteins is an essential post-translational regulatory mechanism that can be assayed using high-throughput mass spectrometry (MS). These MS-based spatial proteomics experiments enable us to pinpoint the sub-cellular distribution of thousands of proteins in a specific system under controlled conditions. Recent advances in high-throughput MS methods have yielded a plethora of experimental spatial proteomics data for the cell biology community. Yet, there are many third-party data sources, such as immunofluorescence microscopy or protein annotations and sequences, which represent a rich and vast source of complementary information. We present a unique transfer learning classification framework that utilises a nearest-neighbour or support vector machine system, to integrate heterogeneous data sources to considerably improve on the quantity and quality of sub-cellular protein assignment. We demonstrate the utility of our algorithms through evaluation of five experimental datasets, from four different species in conjunction with four different auxiliary data sources to classify proteins to tens of sub-cellular compartments with high generalisation accuracy. We further apply the method to an experiment on pluripotent mouse embryonic stem cells to classify a set of previously unknown proteins, and validate our findings against a recent high resolution map of the mouse stem cell proteome. The methodology is distributed as part of the open-source Bioconductor pRoloc suite for spatial proteomics data analysis. Lisa M. Breckels, Sean B. Holden, David Wojnar, Claire M. Mulvey, Andy Christoforou, Arnoud Groen, Matthew W. B. Trotter, Oliver Kohlbacher, Kathryn S. Lilley, Laurent Gatto |
PLoS Comput. Biol. | 8 |
| 2015 | ballaxy: web services for structural bioinformaticsabstractMOTIVATION: Web-based workflow systems have gained considerable momentum in sequence-oriented bioinformatics. In structural bioinformatics, however, such systems are still relatively rare; while commercial stand-alone workflow applications are common in the pharmaceutical industry, academic researchers often still rely on command-line scripting to glue individual tools together. RESULTS: In this work, we address the problem of building a web-based system for workflows in structural bioinformatics. For the underlying molecular modelling engine, we opted for the BALL framework because of its extensive and well-tested functionality in the field of structural bioinformatics. The large number of molecular data structures and algorithms implemented in BALL allows for elegant and sophisticated development of new approaches in the field. We hence connected the versatile BALL library and its visualization and editing front end BALLView with the Galaxy workflow framework. The result, which we call ballaxy, enables the user to simply and intuitively create sophisticated pipelines for applications in structure-based computational biology, integrated into a standard tool for molecular modelling. AVAILABILITY AND IMPLEMENTATION: ballaxy consists of three parts: some minor modifications to the Galaxy system, a collection of tools and an integration into the BALL framework and the BALLView application for molecular modelling. Modifications to Galaxy will be submitted to the Galaxy project, and the BALL and BALLView integrations will be integrated in the next major BALL release. After acceptance of the modifications into the Galaxy project, we will publish all ballaxy tools via the Galaxy toolshed. In the meantime, all three components are available from http://www.ball-project.org/ballaxy. Also, docker images for ballaxy are available at https://registry.hub.docker.com/u/anhi/ballaxy/dockerfile/. ballaxy is licensed under the terms of the GPL. Anna Katharina Hildebrandt, Daniel Stöckel, Nina M. Fischer, Luis de la Garza, Jens Krüger 0002, Stefan Nickels, Marc Röttig, Charlotta Schärfe, Marcel Schumann, Philipp Thiel, Hans-Peter Lenhof, Oliver Kohlbacher, Andreas Hildebrandt 0001 |
Bioinform. | 12 |
| 2015 | EpiToolKit - a web-based workbench for vaccine designabstractUNLABELLED: EpiToolKit is a virtual workbench for immunological questions with a focus on vaccine design. It offers an array of immunoinformatics tools covering MHC genotyping, epitope and neo-epitope prediction, epitope selection for vaccine design, and epitope assembly. In its recently re-implemented version 2.0, EpiToolKit provides a range of new functionality and for the first time allows combining tools into complex workflows. For inexperienced users it offers simplified interfaces to guide the users through the analysis of complex immunological data sets. AVAILABILITY AND IMPLEMENTATION: http://www.epitoolkit.de CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Benjamin Schubert, Hans-Philipp Brachvogel, Christopher Jürges, Oliver Kohlbacher |
Bioinform. | 4 |
| 2015 | Protein (multi-)location prediction: utilizing interdependencies via a generative modelabstractMOTIVATION: Proteins are responsible for a multitude of vital tasks in all living organisms. Given that a protein's function and role are strongly related to its subcellular location, protein location prediction is an important research area. While proteins move from one location to another and can localize to multiple locations, most existing location prediction systems assign only a single location per protein. A few recent systems attempt to predict multiple locations for proteins, however, their performance leaves much room for improvement. Moreover, such systems do not capture dependencies among locations and usually consider locations as independent. We hypothesize that a multi-location predictor that captures location inter-dependencies can improve location predictions for proteins. RESULTS: We introduce a probabilistic generative model for protein localization, and develop a system based on it-which we call MDLoc-that utilizes inter-dependencies among locations to predict multiple locations for proteins. The model captures location inter-dependencies using Bayesian networks and represents dependency between features and locations using a mixture model. We use iterative processes for learning model parameters and for estimating protein locations. We evaluate our classifier MDLoc, on a dataset of single- and multi-localized proteins derived from the DBMLoc dataset, which is the most comprehensive protein multi-localization dataset currently available. Our results, obtained by using MDLoc, significantly improve upon results obtained by an initial simpler classifier, as well as on results reported by other top systems. AVAILABILITY AND IMPLEMENTATION: MDLoc is available at: http://www.eecis.udel.edu/∼compbio/mdloc. Ramanuja Simha, Sebastian Briesemeister, Oliver Kohlbacher, Hagit Shatkay |
Bioinform. | 3 |
| 2014 | Rebuilding KEGG Maps: Algorithms and BenefitsabstractStatic drawings of biological pathways are still an important research tool for biologists. Gerhard Michal created his seminal drawings of metabolic networks in the 1960s and thus defined canonical representations of some key pathways. The Kyoto Encyclopedia of Genes and Genomes (KEGG) provides the most popular static drawings of biological networks of different types, used in a huge number of publications. These drawings are so widely known that they are immediately recognizable to most biologists. This enables collaborative work and simplifies the communication of analysis results. Automatic layout of these pathway maps is complicated by the fact that the information available from KEGG does not contain the entire layout information of the reference maps. Here we present a fully automated algorithm for interactive KEGG layout construction. The algorithm conserves the original KEGG layout to the extent possible while improving readability by removing unnecessary elements (in organism-specific maps). Multiple pathway maps can be laid out simultaneously to facilitate the navigation of larger networks. The algorithm supports the hierarchical layout of sub networks and thus supports interactive exploration of large datasets. Andreas Gerasch, Michael Kaufmann 0001, Oliver Kohlbacher |
PacificVis | 3 |
| 2014 | OptiType: precision HLA typing from next-generation sequencing dataabstractMOTIVATION: The human leukocyte antigen (HLA) gene cluster plays a crucial role in adaptive immunity and is thus relevant in many biomedical applications. While next-generation sequencing data are often available for a patient, deducing the HLA genotype is difficult because of substantial sequence similarity within the cluster and exceptionally high variability of the loci. Established approaches, therefore, rely on specific HLA enrichment and sequencing techniques, coming at an additional cost and extra turnaround time. RESULT: We present OptiType, a novel HLA genotyping algorithm based on integer linear programming, capable of producing accurate predictions from NGS data not specifically enriched for the HLA cluster. We also present a comprehensive benchmark dataset consisting of RNA, exome and whole-genome sequencing data. OptiType significantly outperformed previously published in silico approaches with an overall accuracy of 97% enabling its use in a broad range of applications. András Szolek, Benjamin Schubert, Christopher Mohr, Marc Sturm, Magdalena Feldhahn, Oliver Kohlbacher |
Bioinform. | 6 |
| 2012 | In silico design of targeted SRM-based experimentsabstractSelected reaction monitoring (SRM)-based proteomics approaches enable highly sensitive and reproducible assays for profiling of thousands of peptides in one experiment. The development of such assays involves the determination of retention time, detectability and fragmentation properties of peptides, followed by an optimal selection of transitions. If those properties have to be identified experimentally, the assay development becomes a time-consuming task. We introduce a computational framework for the optimal selection of transitions for a given set of proteins based on their sequence information alone or in conjunction with already existing transition databases. The presented method enables the rapid and fully automated initial development of assays for targeted proteomics. We introduce the relevant methods, report and discuss a step-wise and generic protocol and we also show that we can reach an ad hoc coverage of 80 % of the targeted proteins. The presented algorithmic procedure is implemented in the open-source software package OpenMS/TOPP. Sven Nahnsen, Oliver Kohlbacher |
BMC Bioinform. | 2 |
| 2012 | A Single Sign-On Infrastructure for Science Gateways on a Use Case for Structural Bioinformatics
Sandra Gesing, Richard Grunzke, Jens Krüger 0002, Georg Birkenheuer, Martin Wewior, Patrick Schäfer 0001, Bernd Schuller, Johannes Schuster, Sonja Herres-Pawlis, Sebastian Breuers, Ákos Balaskó, Miklós Kozlovszky, Anna Szikszay Fabri, Lars Packschies, Péter Kacsuk, Dirk Blunk, Thomas Steinke 0001, André Brinkmann, Gregor Fels, Ralph Müller-Pfefferkorn, René Jäkel, Oliver Kohlbacher |
J. Grid Comput. | 22 |
| 2011 | POPISK: T-cell reactivity prediction using support vector machines and string kernelsabstractBACKGROUND: Accurate prediction of peptide immunogenicity and characterization of relation between peptide sequences and peptide immunogenicity will be greatly helpful for vaccine designs and understanding of the immune system. In contrast to the prediction of antigen processing and presentation pathway, the prediction of subsequent T-cell reactivity is a much harder topic. Previous studies of identifying T-cell receptor (TCR) recognition positions were based on small-scale analyses using only a few peptides and concluded different recognition positions such as positions 4, 6 and 8 of peptides with length 9. Large-scale analyses are necessary to better characterize the effect of peptide sequence variations on T-cell reactivity and design predictors of a peptide's T-cell reactivity (and thus immunogenicity). The identification and characterization of important positions influencing T-cell reactivity will provide insights into the underlying mechanism of immunogenicity. RESULTS: This work establishes a large dataset by collecting immunogenicity data from three major immunology databases. In order to consider the effect of MHC restriction, peptides are classified by their associated MHC alleles. Subsequently, a computational method (named POPISK) using support vector machine with a weighted degree string kernel is proposed to predict T-cell reactivity and identify important recognition positions. POPISK yields a mean 10-fold cross-validation accuracy of 68% in predicting T-cell reactivity of HLA-A2-binding peptides. POPISK is capable of predicting immunogenicity with scores that can also correctly predict the change in T-cell reactivity related to point mutations in epitopes reported in previous studies using crystal structures. Thorough analyses of the prediction results identify the important positions 4, 6, 8 and 9, and yield insights into the molecular basis for TCR recognition. Finally, we relate this finding to physicochemical properties and structural features of the MHC-peptide-TCR interaction. CONCLUSIONS: A computational method POPISK is proposed to predict immunogenicity with scores which are useful for predicting immunogenicity changes made by single-residue modifications. The web server of POPISK is freely available at http://iclab.life.nctu.edu.tw/POPISK. Chun-Wei Tung, Matthias Ziehm, Andreas Kämper, Oliver Kohlbacher, Shinn-Ying Ho |
BMC Bioinform. | 4 |
| 2011 | Special Issue: Portals for life sciences - Providing intuitive access to bioinformatic toolsabstractAbstract The topic ‘Portals for life sciences’ includes various research fields, on the one hand many different topics out of life sciences, e.g. mass spectrometry, on the other hand portal technologies and different aspects of computer science, such as usability of user interfaces and security of systems. The main aspect about portals is to simplify the user's interaction with computational resources that are concerted to a supported application domain. Copyright © 2010 John Wiley & Sons, Ltd. Sandra Gesing, Jano I. van Hemert, Péter Kacsuk, Oliver Kohlbacher |
Concurr. Comput. Pract. Exp. | 4 |
| 2010 | TOPP goes Rapid The OpenMS Proteomics Pipeline in a Grid-Enabled Web PortalabstractProteomics, the study of all the proteins contained in a particular sample, e.g., a cell, is a key technology in current biomedical research. The complexity and volume of proteomics data sets produced by mass spectrometric methods clearly suggests the use of grid-based high-performance computing for analysis. TOPP and OpenMS are open-source packages for proteomics data analysis, however, they do not provide support for Grid computing. In this work we present a portal interface for high-throughput data analysis with TOPP. The portal is based on Rapid, a tool for efficiently generating standardized port lets for a wide range of applications. The web-based interface allows the creation and editing of user-defined pipelines and their execution and monitoring on a Grid infrastructure. The portal also supports several file transfer protocols for data staging. It thus provides a simple and complete solution to high-throughput proteomics data analysis for inexperienced users through a convenient portal interface. Sandra Gesing, Jano I. van Hemert, Jos Koetsier, Andreas Bertsch, Oliver Kohlbacher |
CCGRID | 5 |
| 2010 | Going from where to why - interpretable prediction of protein subcellular localizationabstractMOTIVATION: Protein subcellular localization is pivotal in understanding a protein's function. Computational prediction of subcellular localization has become a viable alternative to experimental approaches. While current machine learning-based methods yield good prediction accuracy, most of them suffer from two key problems: lack of interpretability and dealing with multiple locations. RESULTS: We present YLoc, a novel method for predicting protein subcellular localization that addresses these issues. Due to its simple architecture, YLoc can identify the relevant features of a protein sequence contributing to its subcellular localization, e.g. localization signals or motifs relevant to protein sorting. We present several example applications where YLoc identifies the sequence features responsible for protein localization, and thus reveals not only to which location a protein is transported to, but also why it is transported there. YLoc also provides a confidence estimate for the prediction. Thus, the user can decide what level of error is acceptable for a prediction. Due to a probabilistic approach and the use of several thousands of dual-targeted proteins, YLoc is able to predict multiple locations per protein. YLoc was benchmarked using several independent datasets for protein subcellular localization and performs on par with other state-of-the-art predictors. Disregarding low-confidence predictions, YLoc can achieve prediction accuracies of over 90%. Moreover, we show that YLoc is able to reliably predict multiple locations and outperforms the best predictors in this area. AVAILABILITY: www.multiloc.org/YLoc. Sebastian Briesemeister, Jörg Rahnenführer, Oliver Kohlbacher |
Bioinform. | 3 |
| 2010 | BALL - biochemical algorithms library 1.3abstractBACKGROUND: The Biochemical Algorithms Library (BALL) is a comprehensive rapid application development framework for structural bioinformatics. It provides an extensive C++ class library of data structures and algorithms for molecular modeling and structural bioinformatics. Using BALL as a programming toolbox does not only allow to greatly reduce application development times but also helps in ensuring stability and correctness by avoiding the error-prone reimplementation of complex algorithms and replacing them with calls into the library that has been well-tested by a large number of developers. In the ten years since its original publication, BALL has seen a substantial increase in functionality and numerous other improvements. RESULTS: Here, we discuss BALL's current functionality and highlight the key additions and improvements: support for additional file formats, molecular edit-functionality, new molecular mechanics force fields, novel energy minimization techniques, docking algorithms, and support for cheminformatics. CONCLUSIONS: BALL is available for all major operating systems, including Linux, Windows, and MacOS X. It is available free of charge under the Lesser GNU Public License (LPGL). Parts of the code are distributed under the GNU Public License (GPL). BALL is available as source code and binary packages from the project web site at http://www.ball-project.org. Recently, it has been accepted into the debian project; integration into further distributions is currently pursued. Andreas Hildebrandt 0001, Anna Katharina Hildebrandt, Alexander Rurainski, Andreas Bertsch, Marcel Schumann, Nora C. Toussaint, Andreas Moll, Daniel Stöckel, Stefan Nickels, Sabine C. Mueller, Hans-Peter Lenhof, Oliver Kohlbacher |
BMC Bioinform. | 12 |
| 2010 | Exploiting physico-chemical properties in string kernelsabstractBACKGROUND: String kernels are commonly used for the classification of biological sequences, nucleotide as well as amino acid sequences. Although string kernels are already very powerful, when it comes to amino acids they have a major short coming. They ignore an important piece of information when comparing amino acids: the physico-chemical properties such as size, hydrophobicity, or charge. This information is very valuable, especially when training data is less abundant. There have been only very few approaches so far that aim at combining these two ideas. RESULTS: We propose new string kernels that combine the benefits of physico-chemical descriptors for amino acids with the ones of string kernels. The benefits of the proposed kernels are assessed on two problems: MHC-peptide binding classification using position specific kernels and protein classification based on the substring spectrum of the sequences. Our experiments demonstrate that the incorporation of amino acid properties in string kernels yields improved performances compared to standard string kernels and to previously proposed non-substring kernels. CONCLUSIONS: In summary, the proposed modifications, in particular the combination with the RBF substring kernel, consistently yield improvements without affecting the computational complexity. The proposed kernels therefore appear to be the kernels of choice for any protein sequence-based inference. AVAILABILITY: Data sets, code and additional information are available from http://www.fml.tuebingen.mpg.de/raetsch/suppl/aask. Implementations of the developed kernels are available as part of the Shogun toolbox. Nora C. Toussaint, Christian Widmer, Oliver Kohlbacher, Gunnar Rätsch |
BMC Bioinform. | 3 |
| 2010 | Combining Structure and Sequence Information Allows Automated Prediction of Substrate Specificities within Enzyme FamiliesabstractAn important aspect of the functional annotation of enzymes is not only the type of reaction catalysed by an enzyme, but also the substrate specificity, which can vary widely within the same family. In many cases, prediction of family membership and even substrate specificity is possible from enzyme sequence alone, using a nearest neighbour classification rule. However, the combination of structural information and sequence information can improve the interpretability and accuracy of predictive models. The method presented here, Active Site Classification (ASC), automatically extracts the residues lining the active site from one representative three-dimensional structure and the corresponding residues from sequences of other members of the family. From a set of representatives with known substrate specificity, a Support Vector Machine (SVM) can then learn a model of substrate specificity. Applied to a sequence of unknown specificity, the SVM can then predict the most likely substrate. The models can also be analysed to reveal the underlying structural reasons determining substrate specificities and thus yield valuable insights into mechanisms of enzyme specificity. We illustrate the high prediction accuracy achieved on two benchmark data sets and the structural insights gained from ASC by a detailed analysis of the family of decarboxylating dehydrogenases. The ASC web service is available at http://asc.informatik.uni-tuebingen.de/. Marc Röttig, Christian Rausch, Oliver Kohlbacher |
PLoS Comput. Biol. | 3 |
| 2009 | On Open Problems in Biological Network Visualization
Mario Albrecht, Andreas Kerren, Karsten Klein 0001, Oliver Kohlbacher, Petra Mutzel, Wolfgang Paul 0001, Falk Schreiber, Michael Wybrow |
GD | 4 |
| 2009 | FRED - a framework for T-cell epitope detectionabstractUNLABELLED: Over the last decade, immunoinformatics has made significant progress. Computational approaches, in particular the prediction of T-cell epitopes using machine learning methods, are at the core of modern vaccine design. Large-scale analyses and the integration or comparison of different methods become increasingly important. We have developed FRED, an extendable, open source software framework for key tasks in immunoinformatics. In this, its first version, FRED offers easily accessible prediction methods for MHC binding and antigen processing as well as general infrastructure for the handling of antigen sequence data and epitopes. FRED is implemented in Python in a modular way and allows the integration of external methods. AVAILABILITY: FRED is freely available for download at http://www-bs.informatik.uni-tuebingen.de/Software/FRED. Magdalena Feldhahn, Pierre Dönnes, Philipp Thiel, Oliver Kohlbacher |
Bioinform. | 4 |
| 2009 | A novel algorithm for detecting differentially regulated paths based on gene set enrichment analysisabstractMOTIVATION: Deregulated signaling cascades are known to play a crucial role in many pathogenic processes, among them are tumor initiation and progression. In the recent past, modern experimental techniques that allow for measuring the amount of mRNA transcripts of almost all known human genes in a tissue or even in a single cell have opened new avenues for studying the activity of the signaling cascades and for understanding the information flow in the networks. RESULTS: We present a novel dynamic programming algorithm for detecting deregulated signaling cascades. The so-called FiDePa (Finding Deregulated Paths) algorithm interprets differences in the expression profiles of tumor and normal tissues. It relies on the well-known gene set enrichment analysis (GSEA) and efficiently detects all paths in a given regulatory or signaling network that are significantly enriched with differentially expressed genes or proteins. Since our algorithm allows for comparing a single tumor expression profile with the control group, it facilitates the detection of specific regulatory features of a tumor that may help to optimize tumor therapy. To demonstrate the capabilities of our algorithm, we analyzed a glioma expression dataset with respect to a directed graph that combined the regulatory networks of the KEGG and TRANSPATH database. The resulting glioma consensus network that encompasses all detected deregulated paths contained many genes and pathways that are known to be key players in glioma or cancer-related pathogenic processes. Moreover, we were able to correlate clinically relevant features like necrosis or metastasis with the detected paths. AVAILABILITY: C++ source code is freely available, BiNA can be downloaded from http://www.bnplusplus.org/. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Andreas Keller, Christina Backes, Andreas Gerasch, Michael Kaufmann 0001, Oliver Kohlbacher, Eckart Meese, Hans-Peter Lenhof |
Bioinform. | 5 |
| 2009 | KIRMES: kernel-based identification of regulatory modules in euchromatic sequencesabstractMOTIVATION: Understanding transcriptional regulation is one of the main challenges in computational biology. An important problem is the identification of transcription factor (TF) binding sites in promoter regions of potential TF target genes. It is typically approached by position weight matrix-based motif identification algorithms using Gibbs sampling, or heuristics to extend seed oligos. Such algorithms succeed in identifying single, relatively well-conserved binding sites, but tend to fail when it comes to the identification of combinations of several degenerate binding sites, as those often found in cis-regulatory modules. RESULTS: We propose a new algorithm that combines the benefits of existing motif finding with the ones of support vector machines (SVMs) to find degenerate motifs in order to improve the modeling of regulatory modules. In experiments on microarray data from Arabidopsis thaliana, we were able to show that the newly developed strategy significantly improves the recognition of TF targets. AVAILABILITY: The python source code (open source-licensed under GPL), the data for the experiments and a Galaxy-based web service are available at http://www.fml.mpg.de/raetsch/suppl/kirmes/. Sebastian J. Schultheiß, Wolfgang Busch, Jan U. Lohmann, Oliver Kohlbacher, Gunnar Rätsch |
Bioinform. | 4 |
| 2009 | MultiLoc2: integrating phylogeny and Gene Ontology terms improves subcellular protein localization predictionabstractBACKGROUND: Knowledge of subcellular localization of proteins is crucial to proteomics, drug target discovery and systems biology since localization and biological function are highly correlated. In recent years, numerous computational prediction methods have been developed. Nevertheless, there is still a need for prediction methods that show more robustness and higher accuracy. RESULTS: We extended our previous MultiLoc predictor by incorporating phylogenetic profiles and Gene Ontology terms. Two different datasets were used for training the system, resulting in two versions of this high-accuracy prediction method. One version is specialized for globular proteins and predicts up to five localizations, whereas a second version covers all eleven main eukaryotic subcellular localizations. In a benchmark study with five localizations, MultiLoc2 performs considerably better than other methods for animal and plant proteins and comparably for fungal proteins. Furthermore, MultiLoc2 performs clearly better when using a second dataset that extends the benchmark study to all eleven main eukaryotic subcellular localizations. CONCLUSION: MultiLoc2 is an extensive high-performance subcellular protein localization prediction system. By incorporating phylogenetic profiles and Gene Ontology terms MultiLoc2 yields higher accuracies compared to its previous version. Moreover, it outperforms other prediction systems in two benchmarks studies. MultiLoc2 is available as user-friendly and free web-service, available at: http://www-bs.informatik.uni-tuebingen.de/Services/MultiLoc2. Torsten Blum, Sebastian Briesemeister, Oliver Kohlbacher |
BMC Bioinform. | 3 |
| 2009 | KIRMES: kernel-based identification of regulatory modules in euchromatic sequences
Sebastian J. Schultheiß, Wolfgang Busch, Jan U. Lohmann, Oliver Kohlbacher, Gunnar Rätsch |
BMC Bioinform. | 4 |
| 2008 | Multiple Instance Learning Allows MHC Class II Epitope Predictions Across Alleles
Nico Pfeifer, Oliver Kohlbacher |
WABI | 2 |
| 2008 | MetaRoute: fast search for relevant metabolic routes for interactive network navigation and visualizationabstractUNLABELLED: We present MetaRoute, an efficient search algorithm based on atom mapping rules and path weighting schemes that returns relevant or textbook-like routes between a source and a product metabolite within seconds for genome-scale networks. Its speed allows the algorithm to be used interactively through a web interface to visualize relevant routes and local networks for one or multiple organisms based on data from KEGG. AVAILABILITY: http://www-bs.informatik.uni-tuebingen.de/Services/MetaRoute. SUPPLEMENTARY INFORMATION: Supplementary details are available at http://www-bs.informatik.uni-tuebingen.de/Services/MetaRoute. Torsten Blum, Oliver Kohlbacher |
Bioinform. | 2 |
| 2008 | GeneTrailExpress: a web-based pipeline for the statistical evaluation of microarray experimentsabstractBACKGROUND: High-throughput methods that allow for measuring the expression of thousands of genes or proteins simultaneously have opened new avenues for studying biochemical processes. While the noisiness of the data necessitates an extensive pre-processing of the raw data, the high dimensionality requires effective statistical analysis methods that facilitate the identification of crucial biological features and relations. For these reasons, the evaluation and interpretation of expression data is a complex, labor-intensive multi-step process. While a variety of tools for normalizing, analysing, or visualizing expression profiles has been developed in the last years, most of these tools offer only functionality for accomplishing certain steps of the evaluation pipeline. RESULTS: Here, we present a web-based toolbox that provides rich functionality for all steps of the evaluation pipeline. Our tool GeneTrailExpress offers besides standard normalization procedures powerful statistical analysis methods for studying a large variety of biological categories and pathways. Furthermore, an integrated graph visualization tool, BiNA, enables the user to draw the relevant biological pathways applying cutting-edge graph-layout algorithms. CONCLUSION: Our gene expression toolbox with its interactive visualization of the pathways and the expression values projected onto the nodes will simplify the analysis and interpretation of biochemical pathways considerably. Andreas Keller, Christina Backes, Maher Al-Awadhi, Andreas Gerasch, Jan Küntzer, Oliver Kohlbacher, Michael Kaufmann 0001, Hans-Peter Lenhof |
BMC Bioinform. | 6 |
| 2008 | LC-MSsim - a simulation software for liquid chromatography mass spectrometry dataabstractBACKGROUND: Mass Spectrometry coupled to Liquid Chromatography (LC-MS) is commonly used to analyze the protein content of biological samples in large scale studies. The data resulting from an LC-MS experiment is huge, highly complex and noisy. Accordingly, it has sparked new developments in Bioinformatics, especially in the fields of algorithm development, statistics and software engineering. In a quantitative label-free mass spectrometry experiment, crucial steps are the detection of peptide features in the mass spectra and the alignment of samples by correcting for shifts in retention time. At the moment, it is difficult to compare the plethora of algorithms for these tasks. So far, curated benchmark data exists only for peptide identification algorithms but no data that represents a ground truth for the evaluation of feature detection, alignment and filtering algorithms. RESULTS: We present LC-MSsim, a simulation software for LC-ESI-MS experiments. It simulates ESI spectra on the MS level. It reads a list of proteins from a FASTA file and digests the protein mixture using a user-defined enzyme. The software creates an LC-MS data set using a predictor for the retention time of the peptides and a model for peak shapes and elution profiles of the mass spectral peaks. Our software also offers the possibility to add contaminants, to change the background noise level and includes a model for the detectability of peptides in mass spectra. After the simulation, LC-MSsim writes the simulated data to mzData, a public XML format. The software also stores the positions (monoisotopic m/z and retention time) and ion counts of the simulated ions in separate files. CONCLUSION: LC-MSsim generates simulated LC-MS data sets and incorporates models for peak shapes and contaminations. Algorithm developers can match the results of feature detection and alignment algorithms against the simulated ion lists and meaningful error rates can be computed. We anticipate that LC-MSsim will be useful to the wider community to perform benchmark studies and comparisons between computational tools. Ole Schulz-Trieglaff, Nico Pfeifer, Clemens Gröpl, Oliver Kohlbacher, Knut Reinert |
BMC Bioinform. | 4 |
| 2008 | OpenMS - An open-source software framework for mass spectrometryabstractBACKGROUND: Mass spectrometry is an essential analytical technique for high-throughput analysis in proteomics and metabolomics. The development of new separation techniques, precise mass analyzers and experimental protocols is a very active field of research. This leads to more complex experimental setups yielding ever increasing amounts of data. Consequently, analysis of the data is currently often the bottleneck for experimental studies. Although software tools for many data analysis tasks are available today, they are often hard to combine with each other or not flexible enough to allow for rapid prototyping of a new analysis workflow. RESULTS: We present OpenMS, a software framework for rapid application development in mass spectrometry. OpenMS has been designed to be portable, easy-to-use and robust while offering a rich functionality ranging from basic data structures to sophisticated algorithms for data analysis. This has already been demonstrated in several studies. CONCLUSION: OpenMS is available under the Lesser GNU Public License (LGPL) from the project website at http://www.openms.de. Marc Sturm, Andreas Bertsch, Clemens Gröpl, Andreas Hildebrandt 0001, Rene Hussong, Eva Lange, Nico Pfeifer, Ole Schulz-Trieglaff, Alexandra Zerck, Knut Reinert, Oliver Kohlbacher |
BMC Bioinform. | 11 |
| 2008 | Peak intensity prediction in MALDI-TOF mass spectrometry: A machine learning study to support quantitative proteomicsabstractBACKGROUND: Mass spectrometry is a key technique in proteomics and can be used to analyze complex samples quickly. One key problem with the mass spectrometric analysis of peptides and proteins, however, is the fact that absolute quantification is severely hampered by the unclear relationship between the observed peak intensity and the peptide concentration in the sample. While there are numerous approaches to circumvent this problem experimentally (e.g. labeling techniques), reliable prediction of the peak intensities from peptide sequences could provide a peptide-specific correction factor. Thus, it would be a valuable tool towards label-free absolute quantification. RESULTS: In this work we present machine learning techniques for peak intensity prediction for MALDI mass spectra. Features encoding the peptides' physico-chemical properties as well as string-based features were extracted. A feature subset was obtained from multiple forward feature selections on the extracted features. Based on these features, two advanced machine learning methods (support vector regression and local linear maps) are shown to yield good results for this problem (Pearson correlation of 0.68 in a ten-fold cross validation). CONCLUSION: The techniques presented here are a useful first step going beyond the binary prediction of proteotypic peptides towards a more quantitative prediction of peak intensities. These predictions in turn will turn out to be beneficial for mass spectrometry-based quantitative proteomics. Wiebke Timm, Alexandra Scherbart, Sebastian Böcker, Oliver Kohlbacher, Tim W. Nattkemper |
BMC Bioinform. | 4 |
| 2008 | A Mathematical Framework for the Selection of an Optimal Set of Peptides for Epitope-Based VaccinesabstractEpitope-based vaccines (EVs) have a wide range of applications: from therapeutic to prophylactic approaches, from infectious diseases to cancer. The development of an EV is based on the knowledge of target-specific antigens from which immunogenic peptides, so-called epitopes, are derived. Such epitopes form the key components of the EV. Due to regulatory, economic, and practical concerns the number of epitopes that can be included in an EV is limited. Furthermore, as the major histocompatibility complex (MHC) binding these epitopes is highly polymorphic, every patient possesses a set of MHC class I and class II molecules of differing specificities. A peptide combination effective for one person can thus be completely ineffective for another. This renders the optimal selection of these epitopes an important and interesting optimization problem. In this work we present a mathematical framework based on integer linear programming (ILP) that allows the formulation of various flavors of the vaccine design problem and the efficient identification of optimal sets of epitopes. Out of a user-defined set of predicted or experimentally determined epitopes, the framework selects the set with the maximum likelihood of eliciting a broad and potent immune response. Our ILP approach allows an elegant and flexible formulation of numerous variants of the EV design problem. In order to demonstrate this, we show how common immunological requirements for a good EV (e.g., coverage of epitopes from each antigen, coverage of all MHC alleles in a set, or avoidance of epitopes with high mutation rates) can be translated into constraints or modifications of the objective function within the ILP framework. An implementation of the algorithm outperforms a simple greedy strategy as well as a previously suggested evolutionary algorithm and has runtimes on the order of seconds for typical problem sizes. Nora C. Toussaint, Pierre Dönnes, Oliver Kohlbacher |
PLoS Comput. Biol. | 3 |
| 2007 | Electrostatic potentials of proteins in water: a structured continuum approachabstractElectrostatic interactions play a crucial role in many biomolecular processes, including molecular recognition and binding. Biomolecular electrostatics is modulated to a large extent by the water surrounding the molecules. Here, we present a novel approach to the computation of electrostatic potentials which allows the inclusion of water structure into the classical theory of continuum electrostatics. Based on our recent purely differential formulation of nonlocal electrostatics [Hildebrandt, et al. (2004) Phys. Rev. Lett., 93, 108104] we have developed a new algorithm for its efficient numerical solution. The key component of this algorithm is a boundary element solver, having the same computational complexity as established boundary element methods for local continuum electrostatics. This allows, for the first time, the computation of electrostatic potentials and interactions of large biomolecular systems immersed in water including effects of the solvent's structure in a continuum description. We illustrate the applicability of our approach with two examples, the enzymes trypsin and acetylcholinesterase. The approach is applicable to all problems requiring precise prediction of electrostatic interactions in water, such as protein-ligand and protein-protein docking, folding and chromatin regulation. Initial results indicate that this approach may shed new light on biomolecular electrostatics and on aspects of molecular recognition that classical local electrostatics cannot reveal. Andreas Hildebrandt 0001, Ralf Blossey, Sergej Rjasanow, Oliver Kohlbacher, Hans-Peter Lenhof |
Bioinform. | 4 |
| 2007 | TOPP - the OpenMS proteomics pipelineabstractMOTIVATION: Experimental techniques in proteomics have seen rapid development over the last few years. Volume and complexity of the data have both been growing at a similar rate. Accordingly, data management and analysis are one of the major challenges in proteomics. Flexible algorithms are required to handle changing experimental setups and to assist in developing and validating new methods. In order to facilitate these studies, it would be desirable to have a flexible 'toolbox' of versatile and user-friendly applications allowing for rapid construction of computational workflows in proteomics. RESULTS: We describe a set of tools for proteomics data analysis-TOPP, The OpenMS Proteomics Pipeline. TOPP provides a set of computational tools which can be easily combined into analysis pipelines even by non-experts and can be used in proteomics workflows. These applications range from useful utilities (file format conversion, peak picking) over wrapper applications for known applications (e.g. Mascot) to completely new algorithmic techniques for data reduction and data analysis. We anticipate that TOPP will greatly facilitate rapid prototyping of proteomics data evaluation pipelines. As such, we describe the basic concepts and the current abilities of TOPP and illustrate these concepts in the context of two example applications: the identification of peptides from a raw dataset through database search and the complex analysis of a standard addition experiment for the absolute quantitation of biomarkers. The latter example demonstrates TOPP's ability to construct flexible analysis pipelines in support of complex experimental setups. AVAILABILITY: The TOPP components are available as open-source software under the lesser GNU public license (LGPL). Source code is available from the project website at www.OpenMS.de Oliver Kohlbacher, Knut Reinert, Clemens Gröpl, Eva Lange, Nico Pfeifer, Ole Schulz-Trieglaff, Marc Sturm |
Bioinform. | 1 |
| 2007 | SherLoc: high-accuracy prediction of protein subcellular localization by integrating text and protein sequence dataabstractMOTIVATION: Knowing the localization of a protein within the cell helps elucidate its role in biological processes, its function and its potential as a drug target. Thus, subcellular localization prediction is an active research area. Numerous localization prediction systems are described in the literature; some focus on specific localizations or organisms, while others attempt to cover a wide range of localizations. RESULTS: We introduce SherLoc, a new comprehensive system for predicting the localization of eukaryotic proteins. It integrates several types of sequence and text-based features. While applying the widely used support vector machines (SVMs), SherLoc's main novelty lies in the way in which it selects its text sources and features, and integrates those with sequence-based features. We test SherLoc on previously used datasets, as well as on a new set devised specifically to test its predictive power, and show that SherLoc consistently improves on previous reported results. We also report the results of applying SherLoc to a large set of yet-unlocalized proteins. AVAILABILITY: SherLoc, along with Supplementary Information, is available at: http://www-bs.informatik.uni-tuebingen.de/Services/SherLoc/ Hagit Shatkay, Annette Höglund, Scott Brady, Torsten Blum, Pierre Dönnes, Oliver Kohlbacher |
Bioinform. | 6 |
| 2007 | BNDB - The Biochemical Network DatabaseabstractBACKGROUND: Technological advances in high-throughput techniques and efficient data acquisition methods have resulted in a massive amount of life science data. The data is stored in numerous databases that have been established over the last decades and are essential resources for scientists nowadays. However, the diversity of the databases and the underlying data models make it difficult to combine this information for solving complex problems in systems biology. Currently, researchers typically have to browse several, often highly focused, databases to obtain the required information. Hence, there is a pressing need for more efficient systems for integrating, analyzing, and interpreting these data. The standardization and virtual consolidation of the databases is a major challenge resulting in a unified access to a variety of data sources. DESCRIPTION: We present the Biochemical Network Database (BNDB), a powerful relational database platform, allowing a complete semantic integration of an extensive collection of external databases. BNDB is built upon a comprehensive and extensible object model called BioCore, which is powerful enough to model most known biochemical processes and at the same time easily extensible to be adapted to new biological concepts. Besides a web interface for the search and curation of the data, a Java-based viewer (BiNA) provides a powerful platform-independent visualization and navigation of the data. BiNA uses sophisticated graph layout algorithms for an interactive visualization and navigation of BNDB. CONCLUSION: BNDB allows a simple, unified access to a variety of external data sources. Its tight integration with the biochemical network library BN++ offers the possibility for import, integration, analysis, and visualization of the data. BNDB is freely accessible at http://www.bndb.org. Jan Küntzer, Christina Backes, Torsten Blum, Andreas Gerasch, Michael Kaufmann 0001, Oliver Kohlbacher, Hans-Peter Lenhof |
BMC Bioinform. | 6 |
| 2007 | Statistical learning of peptide retention behavior in chromatographic separations: a new kernel-based approach for computational proteomicsabstractBACKGROUND: High-throughput peptide and protein identification technologies have benefited tremendously from strategies based on tandem mass spectrometry (MS/MS) in combination with database searching algorithms. A major problem with existing methods lies within the significant number of false positive and false negative annotations. So far, standard algorithms for protein identification do not use the information gained from separation processes usually involved in peptide analysis, such as retention time information, which are readily available from chromatographic separation of the sample. Identification can thus be improved by comparing measured retention times to predicted retention times. Current prediction models are derived from a set of measured test analytes but they usually require large amounts of training data. RESULTS: We introduce a new kernel function which can be applied in combination with support vector machines to a wide range of computational proteomics problems. We show the performance of this new approach by applying it to the prediction of peptide adsorption/elution behavior in strong anion-exchange solid-phase extraction (SAX-SPE) and ion-pair reversed-phase high-performance liquid chromatography (IP-RP-HPLC). Furthermore, the predicted retention times are used to improve spectrum identifications by a p-value-based filtering approach. The approach was tested on a number of different datasets and shows excellent performance while requiring only very small training sets (about 40 peptides instead of thousands). Using the retention time predictor in our retention time filter improves the fraction of correctly identified peptide mass spectra significantly. CONCLUSION: The proposed kernel function is well-suited for the prediction of chromatographic separation in computational proteomics and requires only a limited amount of training data. The performance of this new method is demonstrated by applying it to peptide retention time prediction in IP-RP-HPLC and prediction of peptide sample fractionation in SAX-SPE. Finally, we incorporate the predicted chromatographic behavior in a p-value based filter to improve peptide identifications based on liquid chromatography-tandem mass spectrometry. Nico Pfeifer, Andreas Leinenbach, Christian G. Huber, Oliver Kohlbacher |
BMC Bioinform. | 4 |
| 2006 | MultiLoc: prediction of protein subcellular localization using N-terminal targeting sequences, sequence motifs and amino acid compositionabstractMOTIVATION: Functional annotation of unknown proteins is a major goal in proteomics. A key annotation is the prediction of a protein's subcellular localization. Numerous prediction techniques have been developed, typically focusing on a single underlying biological aspect or predicting a subset of all possible localizations. An important step is taken towards emulating the protein sorting process by capturing and bringing together biologically relevant information, and addressing the clear need to improve prediction accuracy and localization coverage. RESULTS: Here we present a novel SVM-based approach for predicting subcellular localization, which integrates N-terminal targeting sequences, amino acid composition and protein sequence motifs. We show how this approach improves the prediction based on N-terminal targeting sequences, by comparing our method TargetLoc against existing methods. Furthermore, MultiLoc performs considerably better than comparable methods predicting all major eukaryotic subcellular localizations, and shows better or comparable results to methods that are specialized on fewer localizations or for one organism. AVAILABILITY: http://www-bs.informatik.uni-tuebingen.de/Services/MultiLoc/ Annette Höglund, Pierre Dönnes, Torsten Blum, Hans-Werner Adolph, Oliver Kohlbacher |
Bioinform. | 5 |
| 2006 | BALLView: a tool for research and education in molecular modelingabstractAbstract Summary: We present BALLView, a molecular viewer and modeling tool. It combines state-of-the-art visualization capabilities with powerful modeling functionality including implementations of force field methods and continuum electrostatics models. BALLView is a versatile and extensible tool for research in structural bioinformatics and molecular modeling. Furthermore, the convenient and intuitive graphical user interface offers novice users direct access to the full functionality, rendering it ideal for teaching. Through an interface to the object-oriented scripting language Python it is easily extensible. Availability: BALLView is an open source software and runs on all major platforms (Windows, MacOS X, Linux and most Unix flavors). It is available free of charge under the GNU Public License at Contact: [email protected] Andreas Moll, Andreas Hildebrandt 0001, Hans-Peter Lenhof, Oliver Kohlbacher |
Bioinform. | 4 |
| 2001 | A NMR-spectra-based scoring function for protein dockingabstractA well studied problem in the area of Computational Molecular Biology is the so-called Protein-Protein Docking problem (PPD) that can be formulated as follows: Given two proteins A and B that form a protein complex, compute the 3D-structure of the protein complex AB. Protein docking algorithms can be used to study the driving forces and reaction mechanisms of docking processes. They are also able to speed up the lenghty process of experimental structure elucidation of protein complexes by proposing potential structures. In this paper, we are discussing a variant of the PPD-problem where the input consists of the tertiary structures of A and B plus an unassigned 1H-NMR spectrum of the complex AB. We present a new scoring function for evaluating and ranking potential complex structures produced by a docking algorithm. The scoring function computes a “theoretical” 1H-NMR spectrum for each tentative complex structure and subtracts the calculated spectrum from the experimental spectrum. The absolute areas of the difference spectra are then used to rank the potential complex structures. In contrast to formerly published approaches (e.g. Morelli et. al. [38]) we do not use distance constraints (intermolecular NOE constraints). We have tested the approach with the bound conformations of four protein complexes whose three-dimensional structures are stored in the PDB data bank [5] and whose 1H-NMR shift assignments are available from the BMRB database (BioMagResBank [47]). Oliver Kohlbacher, Andreas Burchardt, Andreas Moll, Andreas Hildebrandt 0001, Peter Bayer, Hans-Peter Lenhof |
RECOMB | 1 |
| 2000 | A combinatorial approach to protein docking with flexible side-chains
Ernst Althaus, Oliver Kohlbacher, Hans-Peter Lenhof, Peter Müller 0008 |
RECOMB | 2 |
| 2000 | BALL-rapid software prototyping in computational molecular biologyabstractAbstract Motivation: Rapid software prototyping can significantly reduce development times in the field of computational molecular biology and molecular modeling. Biochemical Algorithms Library (BALL) is an application framework in C++that has been specifically designed for this purpose. Results: BALL provides an extensive set of data structures as well as classes for molecular mechanics, advanced solvation methods, comparison and analysis of protein structures, file import/export, and visualization. BALL has been carefully designed to be robust, easy to use, and open to extensions. Especially its extensibility which results from an object-oriented and generic programming approach distinguishes it from other software packages. BALL is well suited to serve as a public repository for reliable data structures and algorithms. We show in an example that the implementation of complex methods is greatly simplified when using the data structures and functionality provided by BALL. Availability: BALL is available via internet from http://www.mpi-sb.mpg.de/BALL/. It may be used free of charge for research and teaching. Commercial licenses are available upon request. Contact: O.Kohlbacher, [email protected] To whom correspondence should be addressed. Oliver Kohlbacher, Hans-Peter Lenhof |
Bioinform. | 1 |