EDBT 2026 Demo / reviewers in the wild / expert
Dominik Heider
dblp:29/2345
· DBLP profile ↗
31ranked-venue papers
4as first author
12since 2021 · last 2026
0000-0002-3108-8311ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 23 · 4 first-author · 8 since 2021Artificial intelligence and machine learning · 5 · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Security and privacy · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Expert-aided causal discovery of ancestral graphsabstractPublisher Copyright: © 2026 The Author(s) | openaire: EC/H2020/951847/EU//ELISE Tiago da Silva, Bruna Bazaluk, Eliezer S. Silva, António Góis, Salem Lahlou, Dominik Heider, Samuel Kaski, Diego Mesquita, Adèle Helena Ribeiro |
Inf. Sci. | 6 |
| 2025 | LegionProfiler: a computational tool for the identification of virulence factors and classification of Legionella pneumophila serogroup 1 isolatesabstractSUMMARY: Legionella pneumophila has significantly contributed to multiple cases of pneumonia with a high rate of mortality globally. Its ability to exploit host mechanisms through several expressed virulence factors poses challenges for diagnosis, treatment, and outbreak control. To address this, we developed LegionProfiler, a computational tool that swiftly identifies virulence factor protein domains within genome assemblies of Legionella pneumophila serogroup 1 isolates and classifies them into high- or low-virulence groups. LegionProfiler automates the probing of genome assemblies for virulence-associated protein domains and determines the isolate's potential to cause severe pneumonia infection. The LegionProfiler workflow is made available through a user-friendly interface to enhance technical control of infectious sources and adds important insights to the general epidemiology of clinical isolates. It could also support the development of targeted therapeutic strategies that will improve patient treatment. AVAILABILITY AND IMPLEMENTATION: LegionProfiler is freely accessible as a web service at https://legionprofiler.uni-muenster.de, and can also be run locally in a Docker container. The source code can be found at https://imigitlab.uni-muenster.de/heiderlab/legionprofiler or at Zenodo (DOI:10.5281/zenodo.15592325). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Oluwafemi A. Sarumi, Wilhelm Bertrams, Oliver Schwengers, Jan-Paul Herrmann, Torsten Hain, Laurine Kieper, Markus Petzold, Alexander Goesmann, Bernd Schmeck, Dominik Heider |
Bioinform. | 10 |
| 2025 | Enhancing deep neural network training through learnable adaptive normalizationabstractNormalization is a fundamental preprocessing technique in data science, commonly used to standardize data distributions prior to model training. Its primary role is to maintain consistent statistical properties across features, which facilitates efficient learning and enhances training stability. In deep learning, normalization methods are particularly beneficial, as they regulate input distributions within neural networks, promoting more stable training and faster convergence. This study introduces and evaluates a novel approach: learnable adaptive normalization layers integrated into neural networks. Experiments were conducted across nine datasets encompassing feature, image, and time-series data, utilizing various deep learning architectures, including feed-forward, convolutional, and transformer-based neural networks. The results demonstrate that adaptive normalization consistently outperforms traditional methods, such as mean subtraction, standard deviation scaling, and layer normalization. Moreover, in many scenarios, adaptive normalization achieved performance that was either superior to or comparable with batch normalization, with image classification being the notable exception. These findings indicate that adaptive normalization not only accelerates training convergence but also enhances final model performance in most cases, underscoring its effectiveness. Given that the choice of network architecture is inherently a hyperparameter tuning challenge, we recommend considering adaptive normalization in the preprocessing step for future networks. Furthermore, replacing batch or layer normalization with adaptive normalization may lead to improved training efficiency and final performance, depending on the specific problem. Benedikt Ruhland, Iraj Masoudian, Dominik Heider |
Knowl. Based Syst. | 3 |
| 2024 | Fair swarm learning: Improving incentives for collaboration by a fair reward mechanismabstractSwarm learning is an emerging technique for collaborative machine learning in which several participants train machine learning models without sharing private data. In a standard swarm network, all the nodes in the network receive identical final models regardless of their individual contributions. This mechanism may be deemed unfair from an economic perspective, discouraging organizations with more resources from participating in any collaboration. Here, we present a framework for swarm learning in which nodes receive personalized models based on their contributions. The results of this study demonstrate the efficacy of this approach by showing that all participants experience performance enhancements compared to their local models. However, participants with higher contributions receive better models than those with lower contributions. This fair mechanism results in the highest possible accuracy for the most contributive participant, comparable to the standard swarm learning model. Such incentive structure can motivate resource-rich organizations to engage in collaboration, leading to the development of machine learning models that incorporate data from more resources, which is ultimately beneficial for every party. Mohammad Tajabadi, Dominik Heider |
Knowl. Based Syst. | 2 |
| 2023 | Human-in-the-Loop Integration with Domain-Knowledge Graphs for Explainable Federated Deep LearningabstractAbstract We explore the integration of domain knowledge graphs into Deep Learning for improved interpretability and explainability using Graph Neural Networks (GNNs). Specifically, a protein-protein interaction (PPI) network is masked over a deep neural network for classification, with patient-specific multi-modal genomic features enriched into the PPI graph’s nodes. Subnetworks that are relevant to the classification (referred to as “disease subnetworks”) are detected using explainable AI. Federated learning is enabled by dividing the knowledge graph into relevant subnetworks, constructing an ensemble classifier, and allowing domain experts to analyze and manipulate detected subnetworks using a developed user interface. Furthermore, the human-in-the-loop principle can be applied with the incorporation of experts, interacting through a sophisticated User Interface (UI) driven by Explainable Artificial Intelligence (xAI) methods, changing the datasets to create counterfactual explanations. The adapted datasets could influence the local model’s characteristics and thereby create a federated version that distils their diverse knowledge in a centralized scenario. This work demonstrates the feasibility of the presented strategies, which were originally envisaged in 2021 and most of it has now been materialized into actionable items. In this paper, we report on some lessons learned during this project. Andreas Holzinger, Anna Saranti, Anne-Christin Hauschild, Jacqueline Michelle Metsch, Dominik Heider, Richard Röttger, Heimo Müller, Jan Baumbach, Bastian Pfeifer |
CD-MAKE | 5 |
| 2023 | ODNA: identification of organellar DNA by machine learningabstractMOTIVATION: Identifying organellar DNA, such as mitochondrial or plastid sequences, inside a whole genome assembly, remains challenging and requires biological background knowledge. To address this, we developed ODNA based on genome annotation and machine learning to fulfill. RESULTS: ODNA is a software that classifies organellar DNA sequences within a genome assembly by machine learning based on a predefined genome annotation workflow. We trained our model with 829 769 DNA sequences from 405 genome assemblies and achieved high predictive performance (e.g. matthew's correlation coefficient of 0.61 for mitochondria and 0.73 for chloroplasts) on independent validation data, thus outperforming existing approaches significantly. AVAILABILITY AND IMPLEMENTATION: Our software ODNA is freely accessible as a web service at https://odna.mathematik.uni-marburg.de and can also be run in a docker container. The source code can be found at https://gitlab.com/mosga/odna and the processed data at Zenodo (DOI: 10.5281/zenodo.7506483). Roman Martin, Minh Kien Nguyen, Nick Lowack, Dominik Heider |
Bioinform. | 4 |
| 2023 | Ensemble-GNN: federated ensemble learning with graph neural networks for disease module discovery and classificationabstractSUMMARY: Federated learning enables collaboration in medicine, where data is scattered across multiple centers without the need to aggregate the data in a central cloud. While, in general, machine learning models can be applied to a wide range of data types, graph neural networks (GNNs) are particularly developed for graphs, which are very common in the biomedical domain. For instance, a patient can be represented by a protein-protein interaction (PPI) network where the nodes contain the patient-specific omics features. Here, we present our Ensemble-GNN software package, which can be used to deploy federated, ensemble-based GNNs in Python. Ensemble-GNN allows to quickly build predictive models utilizing PPI networks consisting of various node features such as gene expression and/or DNA methylation. We exemplary show the results from a public dataset of 981 patients and 8469 genes from the Cancer Genome Atlas (TCGA). AVAILABILITY AND IMPLEMENTATION: The source code is available at https://github.com/pievos101/Ensemble-GNN, and the data at Zenodo (DOI: 10.5281/zenodo.8305122). Bastian Pfeifer, Hryhorii Chereda, Roman Martin, Anna Saranti, Sandra Clemens, Anne-Christin Hauschild, Tim Beißbarth, Andreas Holzinger, Dominik Heider |
Bioinform. | 9 |
| 2022 | Federated Random Forests can improve local performance of predictive models for various healthcare applicationsabstractMOTIVATION: Limited data access has hindered the field of precision medicine from exploring its full potential, e.g. concerning machine learning and privacy and data protection rules.Our study evaluates the efficacy of federated Random Forests (FRF) models, focusing particularly on the heterogeneity within and between datasets. We addressed three common challenges: (i) number of parties, (ii) sizes of datasets and (iii) imbalanced phenotypes, evaluated on five biomedical datasets. RESULTS: The FRF outperformed the average local models and performed comparably to the data-centralized models trained on the entire data. With an increasing number of models and decreasing dataset size, the performance of local models decreases drastically. The FRF, however, do not decrease significantly. When combining datasets of different sizes, the FRF vastly improve compared to the average local models. We demonstrate that the FRF remain more robust and outperform the local models by analyzing different class-imbalances.Our results support that FRF overcome boundaries of clinical research and enables collaborations across institutes without violating privacy or legal regulations. Clinicians benefit from a vast collection of unbiased data aggregated from different geographic locations, demographics and other varying factors. They can build more generalizable models to make better clinical decisions, which will have relevance, especially for patients in rural areas and rare or geographically uncommon diseases, enabling personalized treatment. In combination with secure multi-party computation, federated learning has the power to revolutionize clinical practice by increasing the accuracy and robustness of healthcare AI and thus paving the way for precision medicine. AVAILABILITY AND IMPLEMENTATION: The implementation of the federated random forests can be found at https://featurecloud.ai/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Anne-Christin Hauschild, Marta Lemanczyk, Julian O. Matschinske, Tobias Frisch, Olga I. Zolotareva, Andreas Holzinger, Jan Baumbach, Dominik Heider |
Bioinform. | 8 |
| 2022 | Prediction of antimicrobial resistance based on whole-genome sequencing and machine learningabstractMOTIVATION: Antimicrobial resistance (AMR) is one of the biggest global problems threatening human and animal health. Rapid and accurate AMR diagnostic methods are thus very urgently needed. However, traditional antimicrobial susceptibility testing (AST) is time-consuming, low throughput and viable only for cultivable bacteria. Machine learning methods may pave the way for automated AMR prediction based on genomic data of the bacteria. However, comparing different machine learning methods for the prediction of AMR based on different encodings and whole-genome sequencing data without previously known knowledge remains to be done. RESULTS: In this study, we evaluated logistic regression (LR), support vector machine (SVM), random forest (RF) and convolutional neural network (CNN) for the prediction of AMR for the antibiotics ciprofloxacin, cefotaxime, ceftazidime and gentamicin. We could demonstrate that these models can effectively predict AMR with label encoding, one-hot encoding and frequency matrix chaos game representation (FCGR encoding) on whole-genome sequencing data. We trained these models on a large AMR dataset and evaluated them on an independent public dataset. Generally, RFs and CNNs perform better than LR and SVM with AUCs up to 0.96. Furthermore, we were able to identify mutations that are associated with AMR for each antibiotic. AVAILABILITY AND IMPLEMENTATION: Source code in data preparation and model training are provided at GitHub website (https://github.com/YunxiaoRen/ML-iAMR). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yunxiao Ren, Trinad Chakraborty, Swapnil Doijad, Linda Falgenhauer, Jane Falgenhauer, Alexander Goesmann, Anne-Christin Hauschild, Oliver Schwengers, Dominik Heider |
Bioinform. | 9 |
| 2021 | Computational strategies to combat COVID-19: useful tools to accelerate SARS-CoV-2 and coronavirus researchabstractSARS-CoV-2 (severe acute respiratory syndrome coronavirus 2) is a novel virus of the family Coronaviridae. The virus causes the infectious disease COVID-19. The biology of coronaviruses has been studied for many years. However, bioinformatics tools designed explicitly for SARS-CoV-2 have only recently been developed as a rapid reaction to the need for fast detection, understanding and treatment of COVID-19. To control the ongoing COVID-19 pandemic, it is of utmost importance to get insight into the evolution and pathogenesis of the virus. In this review, we cover bioinformatics workflows and tools for the routine detection of SARS-CoV-2 infection, the reliable analysis of sequencing data, the tracking of the COVID-19 pandemic and evaluation of containment measures, the study of coronavirus evolution, the discovery of potential drug targets and development of therapeutic strategies. For each tool, we briefly describe its use case and how it advances research specifically for SARS-CoV-2. All tools are free to use and available online, either through web applications or public code repositories. Contact:[email protected]. Franziska Hufsky, Kevin Lamkiewicz, Alexandre Almeida, Abdel Aouacheria, Cecilia N. Arighi, Alex Bateman, Jan Baumbach, Niko Beerenwinkel, Christian Brandt, Marco Cacciabue, Sara Chuguransky, Oliver Drechsel, Robert D. Finn, Adrian Fritz, Stephan Fuchs, Georges Hattab, Anne-Christin Hauschild, Dominik Heider, Marie Hoffmann, Martin Hölzer, Stefan Hoops, Lars Kaderali, Ioanna Kalvari, Max von Kleist, Renó Kmiecinski, Denise Kühnert, Gorka Lasso, Pieter Libin, Markus List, Hannah F. Löchel, Maria Jesus Martin, Roman Martin, Julian O. Matschinske, Alice C. McHardy, Pedro Mendes 0001, Jaina Mistry, Vincent Navratil, Eric P. Nawrocki, Áine Niamh O'toole, Nancy Ontiveros-Palacios, Anton I. Petrov, Guillermo Rangel-Pineros, Nicole Redaschi, Susanne Reimering, Knut Reinert, Lorna J. Richardson, David L. Robertson, Sepideh Sadegh, Joshua B. Singer, Kristof Theys, Chris Upton, Marius Welzel, Lowri Williams, Manja Marz |
Briefings Bioinform. | 18 |
| 2021 | MOSGA: Modular Open-Source Genome AnnotatorabstractMOTIVATION: The generation of high-quality assemblies, even for large eukaryotic genomes, has become a routine task for many biologists thanks to recent advances in sequencing technologies. However, the annotation of these assemblies-a crucial step toward unlocking the biology of the organism of interest-has remained a complex challenge that often requires advanced bioinformatics expertise. RESULTS: Here, we present MOSGA (Modular Open-Source Genome Annotator), a genome annotation framework for eukaryotic genomes with a user-friendly web-interface that generates and integrates annotations from various tools. The aggregated results can be analyzed with a fully integrated genome browser and are provided in a format ready for submission to NCBI. MOSGA is built on a portable, customizable and easily extendible Snakemake backend, and thus, can be tailored to a wide range of users and projects. AVAILABILITY AND IMPLEMENTATION: We provide MOSGA as a web service at https://mosga.mathematik.uni-marburg.de and as a docker container at registry.gitlab.com/mosga/mosga: latest. Source code can be found at https://gitlab.com/mosga/mosga. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Roman Martin, Thomas Hackl, Georges Hattab, Matthias G. Fischer, Dominik Heider |
Bioinform. | 5 |
| 2021 | Correction: Ten simple rules to colorize biological data visualizationabstract[This corrects the article DOI: 10.1371/journal.pcbi.1008259.]. Georges Hattab, Theresa-Marie Rhyne, Dominik Heider |
PLoS Comput. Biol. | 3 |
| 2020 | Deep learning on chaos game representation for proteinsabstractMOTIVATION: Classification of protein sequences is one big task in bioinformatics and has many applications. Different machine learning methods exist and are applied on these problems, such as support vector machines (SVM), random forests (RF) and neural networks (NN). All of these methods have in common that protein sequences have to be made machine-readable and comparable in the first step, for which different encodings exist. These encodings are typically based on physical or chemical properties of the sequence. However, due to the outstanding performance of deep neural networks (DNN) on image recognition, we used frequency matrix chaos game representation (FCGR) for encoding of protein sequences into images. In this study, we compare the performance of SVMs, RFs and DNNs, trained on FCGR encoded protein sequences. While the original chaos game representation (CGR) has been used mainly for genome sequence encoding and classification, we modified it to work also for protein sequences, resulting in n-flakes representation, an image with several icosagons. RESULTS: We could show that all applied machine learning techniques (RF, SVM and DNN) show promising results compared to the state-of-the-art methods on our benchmark datasets, with DNNs outperforming the other methods and that FCGR is a promising new encoding method for protein sequences. AVAILABILITY AND IMPLEMENTATION: https://cran.r-project.org/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Hannah F. Löchel, Dominic Eger, Theodor Sperlea, Dominik Heider |
Bioinform. | 4 |
| 2020 | MESA: automated assessment of synthetic DNA fragments and simulation of DNA synthesis, storage, sequencing and PCR errorsabstractSUMMARY: The development of de novo DNA synthesis, polymerase chain reaction (PCR), DNA sequencing and molecular cloning gave researchers unprecedented control over DNA and DNA-mediated processes. To reduce the error probabilities of these techniques, DNA composition has to adhere to method-dependent restrictions. To comply with such restrictions, a synthetic DNA fragment is often adjusted manually or by using custom-made scripts. In this article, we present MESA (Mosla Error Simulator), a web application for the assessment of DNA fragments based on limitations of DNA synthesis, amplification, cloning, sequencing methods and biological restrictions of host organisms. Furthermore, MESA can be used to simulate errors during synthesis, PCR, storage and sequencing processes. AVAILABILITY AND IMPLEMENTATION: MESA is available at mesa.mosla.de, with the source code available at github.com/umr-ds/mesa_dna_sim. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Michael Schwarz 0009, Marius Welzel, Tolganay Kabdullayeva, Anke Becker, Bernd Freisleben, Dominik Heider |
Bioinform. | 6 |
| 2020 | Natrix: a Snakemake-based workflow for processing, clustering, and taxonomically assigning amplicon sequencing readsabstractBACKGROUND: Sequencing of marker genes amplified from environmental samples, known as amplicon sequencing, allows us to resolve some of the hidden diversity and elucidate evolutionary relationships and ecological processes among complex microbial communities. The analysis of large numbers of samples at high sequencing depths generated by high throughput sequencing technologies requires efficient, flexible, and reproducible bioinformatics pipelines. Only a few existing workflows can be run in a user-friendly, scalable, and reproducible manner on different computing devices using an efficient workflow management system. RESULTS: We present Natrix, an open-source bioinformatics workflow for preprocessing raw amplicon sequencing data. The workflow contains all analysis steps from quality assessment, read assembly, dereplication, chimera detection, split-sample merging, sequence representative assignment (OTUs or ASVs) to the taxonomic assignment of sequence representatives. The workflow is written using Snakemake, a workflow management engine for developing data analysis workflows. In addition, Conda is used for version control. Thus, Snakemake ensures reproducibility and Conda offers version control of the utilized programs. The encapsulation of rules and their dependencies support hassle-free sharing of rules between workflows and easy adaptation and extension of existing workflows. Natrix is freely available on GitHub ( https://github.com/MW55/Natrix ) or as a Docker container on DockerHub ( https://hub.docker.com/r/mw55/natrix ). CONCLUSION: Natrix is a user-friendly and highly extensible workflow for processing Illumina amplicon data. Marius Welzel, Anja Lange, Dominik Heider, Michael Schwarz 0009, Bernd Freisleben, Manfred Jensen, Jens Boenigk, Daniela Beisser |
BMC Bioinform. | 3 |
| 2020 | Ten simple rules to colorize biological data visualizationabstractMethods for visualization of biological data continue to improve, but there is still a fundamental challenge in colorization of these visualizations (vis).Visual representation of biological data should not overwhelm, obscure, or bias the findings, but rather make them more understandable.This is often due to the challenge of how to use color effectively in creating visualizations.The recent global adoption of data vis has helped address this challenge in some fields, but it remains open in the biological domain.The visualization of biological data deals with the application of computer graphics, scientific visualization, and information visualization in various areas of the life sciences.This paper describes 10 simple rules to colorize biological data visualization.Rule 1: Identify the nature of your data Rule 2: Select a color space Rule 3: Create a color palette based on the selected color space Rule 4: Apply the color palette to your data set for visualization Rule 5: Check for color context in your data vis after the color palette is applied Rule 6: Evaluate interactions of colors in your data visualization Rule 7: Be aware of color conventions and definitions in your particular discipline Rule 8: Assess color deficiencies Rule 9: Consider web content accessibility and print realities Rule 10: Get it right in black and white Georges Hattab, Theresa-Marie Rhyne, Dominik Heider |
PLoS Comput. Biol. | 3 |
| 2019 | FRI-Feature Relevance Intervals for Interpretable and Interactive Data ExplorationabstractMost existing feature selection methods are insufficient for analytic purposes as soon as high dimensional data or redundant sensor signals are dealt with since features can be selected due to spurious effects or correlations rather than causal effects. To support the finding of causal features in biomedical experiments, we hereby present FRI, an open source Python library that can be used to identify all-relevant variables in linear classification and (ordinal) regression problems. Using the recently proposed feature relevance interval method, FRI is able to provide the base for further general experimentation or in specific can facilitate the search for alternative biomarkers. It can be used in an interactive context, by providing model manipulation and visualization methods, or in a batch process as a filter method. Lukas Pfannschmidt, Christina Göpfert, Ursula Neumann, Dominik Heider, Barbara Hammer |
CIBCB | 4 |
| 2019 | The virtual doctor: An interactive clinical-decision-support system based on deep learning for non-invasive prediction of diabetes
Sebastian Spänig, Agnes Emberger-Klein, Jan-Peter Sowa, Ali Canbay, Klaus Menrad, Dominik Heider |
Artif. Intell. Medicine | 6 |
| 2019 | GUESS: projecting machine learning scores to well-calibrated probability estimates for clinical decision-makingabstractMOTIVATION: Clinical decision support systems have been applied in numerous fields, ranging from cancer survival toward drug resistance prediction. Nevertheless, clinical decision support systems typically have a caveat: many of them are perceived as black-boxes by non-experts and, unfortunately, the obtained scores cannot usually be interpreted as class probability estimates. In probability-focused medical applications, it is not sufficient to perform well with regards to discrimination and, consequently, various calibration methods have been developed to enable probabilistic interpretation. The aims of this study were (i) to develop a tool for fast and comparative analysis of different calibration methods, (ii) to demonstrate their limitations for the use on clinical data and (iii) to introduce our novel method GUESS. RESULTS: We compared the performances of two different state-of-the-art calibration methods, namely histogram binning and Bayesian Binning in Quantiles, as well as our novel method GUESS on both, simulated and real-world datasets. GUESS demonstrated calibration performance comparable to the state-of-the-art methods and always retained accurate class discrimination. GUESS showed superior calibration performance in small datasets and therefore may be an optimal calibration method for typical clinical datasets. Moreover, we provide a framework (CalibratR) for R, which can be used to identify the most suitable calibration method for novel datasets in a timely and efficient manner. Using calibrated probability estimates instead of original classifier scores will contribute to the acceptance and dissemination of machine learning based classification models in cost-sensitive applications, such as clinical research. AVAILABILITY AND IMPLEMENTATION: GUESS as part of CalibratR can be downloaded at CRAN. Johanna Schwarz, Dominik Heider |
Bioinform. | 2 |
| 2018 | SCOTCH: subtype A coreceptor tropism classification in HIV-1abstractMotivation: The V3 loop of the gp120 glycoprotein of the Human Immunodeficiency Virus 1 (HIV-1) is considered to be responsible for viral coreceptor tropism. gp120 interacts with the CD4 receptor of the host cell and subsequently V3 binds either CCR5 or CXCR4. Due to the fact that the CCR5 coreceptor is targeted by entry inhibitors, a reliable prediction of the coreceptor usage of HIV-1 is of great interest for antiretroviral therapy. Although several methods for the prediction of coreceptor tropism are available, almost all of them have been developed based on only subtype B sequences, and it has been shown in several studies that the prediction of non-B sequences, in particular subtype A sequences, are less reliable. Thus, the aim of the current study was to develop a reliable prediction model for subtype A viruses. Results: Our new model SCOTCH is based on a stacking approach of classifier ensembles and shows a significantly better performance for subtype A sequences compared to other available models. In particular for low false positive rates (between 0.05 and 0.2, i.e. recommendation in the German and European Guidelines for tropism prediction), SCOTCH shows significantly better prediction performances in terms of partial area under the curves and diagnostic odds ratios compared to existing tools, and thus can be used to reliably predict coreceptor tropism for subtype A sequences. Availability and implementation: SCOTCH can be downloaded/accessed at http://www.heiderlab.de. Hannah F. Löchel, Mona Riemenschneider, Dmitrij Frishman, Dominik Heider |
Bioinform. | 4 |
| 2017 | eccCL: parallelized GPU implementation of Ensemble Classifier ChainsabstractBACKGROUND: Multi-label classification has recently gained great attention in diverse fields of research, e.g., in biomedical application such as protein function prediction or drug resistance testing in HIV. In this context, the concept of Classifier Chains has been shown to improve prediction accuracy, especially when applied as Ensemble Classifier Chains. However, these techniques lack computational efficiency when applied on large amounts of data, e.g., derived from next-generation sequencing experiments. By adapting algorithms for the use of graphics processing units, computational efficiency can be greatly improved due to parallelization of computations. RESULTS: Here, we provide a parallelized and optimized graphics processing unit implementation (eccCL) of Classifier Chains and Ensemble Classifier Chains. Additionally to the OpenCL implementation, we provide an R-Package with an easy to use R-interface for parallelized graphics processing unit usage. CONCLUSION: eccCL is a handy implementation of Classifier Chains on GPUs, which is able to process up to over 25,000 instances per second, and thus can be used efficiently in high-throughput experiments. The software is available at http://www.heiderlab.de . Mona Riemenschneider, Alexander Herbst, Ari Rasch, Sergei Gorlatch, Dominik Heider |
BMC Bioinform. | 5 |
| 2016 | SHIVA - a web application for drug resistance and tropism testing in HIVabstractBACKGROUND: Drug resistance testing is mandatory in antiretroviral therapy in human immunodeficiency virus (HIV) infected patients for successful treatment. The emergence of resistances against antiretroviral agents remains the major obstacle in inhibition of viral replication and thus to control infection. Due to the high mutation rate the virus is able to adapt rapidly under drug pressure leading to the evolution of resistant variants and finally to therapy failure. RESULTS: We developed a web service for drug resistance prediction of commonly used drugs in antiretroviral therapy, i.e., protease inhibitors (PIs), reverse transcriptase inhibitors (NRTIs and NNRTIs), and integrase inhibitors (INIs), but also for the novel drug class of maturation inhibitors. Furthermore, co-receptor tropism (CCR5 or CXCR4) can be predicted as well, which is essential for treatment with entry inhibitors, such as Maraviroc. Currently, SHIVA provides 24 prediction models for several drug classes. SHIVA can be used with single RNA/DNA or amino acid sequences, but also with large amounts of next-generation sequencing data and allows prediction of a user specified selection of drugs simultaneously. Prediction results are provided as clinical reports which are sent via email to the user. CONCLUSIONS: SHIVA represents a novel high performing alternative for hitherto developed drug resistance testing approaches able to process data derived from next-generation sequencing technologies. SHIVA is publicly available via a user-friendly web interface. Mona Riemenschneider, Thomas Hummel 0001, Dominik Heider |
BMC Bioinform. | 3 |
| 2014 | gCUP: rapid GPU-based HIV-1 co-receptor usage prediction for next-generation sequencingabstractSUMMARY: Next-generation sequencing (NGS) has a large potential in HIV diagnostics, and genotypic prediction models have been developed and successfully tested in the recent years. However, albeit being highly accurate, these computational models lack computational efficiency to reach their full potential. In this study, we demonstrate the use of graphics processing units (GPUs) in combination with a computational prediction model for HIV tropism. Our new model named gCUP, parallelized and optimized for GPU, is highly accurate and can classify >175 000 sequences per second on an NVIDIA GeForce GTX 460. The computational efficiency of our new model is the next step to enable NGS technologies to reach clinical significance in HIV diagnostics. Moreover, our approach is not limited to HIV tropism prediction, but can also be easily adapted to other settings, e.g. drug resistance prediction. AVAILABILITY AND IMPLEMENTATION: The source code can be downloaded at http://www.heiderlab.de CONTACT: [email protected]. Michael Olejnik, Michel Steuwer, Sergei Gorlatch, Dominik Heider |
Bioinform. | 4 |
| 2013 | Multilabel classification for exploiting cross-resistance information in HIV-1 drug resistance predictionabstractMOTIVATION: Antiretroviral treatment regimens can sufficiently suppress viral replication in human immunodeficiency virus (HIV)-infected patients and prevent the progression of the disease. However, one of the factors contributing to the progression of the disease despite ongoing antiretroviral treatment is the emergence of drug resistance. The high mutation rate of HIV can lead to a fast adaptation of the virus under drug pressure, thus to failure of antiretroviral treatment due to the evolution of drug-resistant variants. Moreover, cross-resistance phenomena have been frequently found in HIV-1, leading to resistance not only against a drug from the current treatment, but also to other not yet applied drugs. Automatic classification and prediction of drug resistance is increasingly important in HIV research as well as in clinical settings, and to this end, machine learning techniques have been widely applied. Nevertheless, cross-resistance information was not taken explicitly into account, yet. RESULTS: In our study, we demonstrated the use of cross-resistance information to predict drug resistance in HIV-1. We tested a set of more than 600 reverse transcriptase sequences and corresponding resistance information for six nucleoside analogues. Based on multilabel classification models and cross-resistance information, we were able to significantly improve overall prediction accuracy for all drugs, compared with single binary classifiers without any additional information. Moreover, we identified drug-specific patterns within the reverse transcriptase sequences that can be used to determine an optimal order of the classifiers within the classifier chains. These patterns are in good agreement with known resistance mutations and support the use of cross-resistance information in such prediction models. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Dominik Heider, Robin Senge, Weiwei Cheng, Eyke Hüllermeier |
Bioinform. | 1 |
| 2012 | The Brain in a Box - An Encoding Scheme for Natural Neural Networks
Martin Pyka, Tilo Kircher, Sascha Hauke, Dominik Heider |
IJCCI | 4 |
| 2010 | Harnessing Recommendations from Weakly Linked Neighbors in Reputation-Based Trust FormationabstractInteractions between individuals are inherently dependent upon trust, no matter if they occur in the real world or in cyber communities. Over the past years, proposals have been made to model trust relations computationally, either to assist users or for modeling purposes in multi-agent systems. These models rely implicitly on the social networks established by participating entities (be they autonomous agents or internet users). However, state-of-the-art trust frameworks often neglect the structure of those complex networks. In this paper, we present a new approach allowing agent-based trust frameworks to leverage information from so-called weak ties that would otherwise be neglected. An effective and robust voting scheme based on an agreement metric is presented and its benefit is shown through simulations. Sascha Hauke, Martin Pyka, Markus Borschbach, Dominik Heider |
CW | 4 |
| 2010 | Augmenting Reputation-Based Trust Metrics with Rumor-Like Dissemination of Reputation Information
Sascha Hauke, Martin Pyka, Markus Borschbach, Dominik Heider |
SEC | 4 |
| 2010 | Predicting Bevirimat resistance of HIV-1 from genotypeabstractBACKGROUND: Maturation inhibitors are a new class of antiretroviral drugs. Bevirimat (BVM) was the first substance in this class of inhibitors entering clinical trials. While the inhibitory function of BVM is well established, the molecular mechanisms of action and resistance are not well understood. It is known that mutations in the regions CS p24/p2 and p2 can cause phenotypic resistance to BVM. We have investigated a set of p24/p2 sequences of HIV-1 of known phenotypic resistance to BVM to test whether BVM resistance can be predicted from sequence, and to identify possible molecular mechanisms of BVM resistance in HIV-1. RESULTS: We used artificial neural networks and random forests with different descriptors for the prediction of BVM resistance. Random forests with hydrophobicity as descriptor performed best and classified the sequences with an area under the Receiver Operating Characteristics (ROC) curve of 0.93 +/- 0.001. For the collected data we find that p2 sequence positions 369 to 376 have the highest impact on resistance, with positions 370 and 372 being particularly important. These findings are in partial agreement with other recent studies. Apart from the complex machine learning models we derived a number of simple rules that predict BVM resistance from sequence with surprising accuracy. According to computational predictions based on the data set used, cleavage sites are usually not shifted by resistance mutations. However, we found that resistance mutations could shorten and weaken the alpha-helix in p2, which hints at a possible resistance mechanism. CONCLUSIONS: We found that BVM resistance of HIV-1 can be predicted well from the sequence of the p2 peptide, which may prove useful for personalized therapy if maturation inhibitors reach clinical practice. Results of secondary structure analysis are compatible with a possible route to BVM resistance in which mutations weaken a six-helix bundle discovered in recent experiments, and thus ease Gag cleavage by the retroviral protease. Dominik Heider, Jens Verheyen, Daniel Hoffmann |
BMC Bioinform. | 1 |
| 2010 | Prediction of Co-Receptor Usage of HIV-1 from GenotypeabstractHuman Immunodeficiency Virus 1 uses for entry into host cells a receptor (CD4) and one of two co-receptors (CCR5 or CXCR4). Recently, a new class of antiretroviral drugs has entered clinical practice that specifically bind to the co-receptor CCR5, and thus inhibit virus entry. Accurate prediction of the co-receptor used by the virus in the patient is important as it allows for personalized selection of effective drugs and prognosis of disease progression. We have investigated whether it is possible to predict co-receptor usage accurately by analyzing the amino acid sequence of the main determinant of co-receptor usage, i.e., the third variable loop V3 of the gp120 protein. We developed a two-level machine learning approach that in the first level considers two different properties important for protein-protein binding derived from structural models of V3 and V3 sequences. The second level combines the two predictions of the first level. The two-level method predicts usage of CXCR4 co-receptor for new V3 sequences within seconds, with an area under the ROC curve of 0.937+/-0.004. Moreover, it is relatively robust against insertions and deletions, which frequently occur in V3. The approach could help clinicians to find optimal personalized treatments, and it offers new insights into the molecular basis of co-receptor usage. For instance, it quantifies the importance for co-receptor usage of a pocket that probably is responsible for binding sulfated tyrosine. Jan Nikolaj Dybowski, Dominik Heider, Daniel Hoffmann |
PLoS Comput. Biol. | 2 |
| 2008 | Watermarking sexually reproducing diploid organismsabstractUNLABELLED: DNA watermarks are used for hiding messages or for authenticating genetically modified organisms. Recently, we presented an algorithm called DNA-Crypt for generating DNA-based watermarks that can be integrated into the genome by using the characteristics of the degenerative genetic code. DNA-Crypt generates the watermark by replacing single bases and thus creating synonymous codons that encrypt the hidden information. Mutations within the integrated DNA sequence can be corrected using several mutation correction codes, to keep the hidden information intact. This method has successfully been tested in asexually replicating organisms like bacteria or yeast, where the watermark is duplicated with every cell division. It has been shown that DNA watermarks produced by DNA-Crypt do not influence the transcription or translation of a protein. In sexually reproducing diploid organisms, additional problems can occur, e.g. recombination events can destroy hidden information. Using population predictions as well as statistical analyses we identified a coupled Y-chromosomal/mitochondrial DNA watermarking procedure as the most appropriate for diploid organisms. We developed a mitochondria adapted version of DNA-Crypt, which is called Project Mito that can be used in combination with the original program. AVAILABILITY: http://www.uni-muenster.de/Biologie.NeuroVer/Tumorbiologie/DNA-Crypt/index.html Dominik Heider, Daniel Kessler, Angelika Barnekow |
Bioinform. | 1 |
| 2007 | DNA-based watermarks using the DNA-Crypt algorithmabstractBACKGROUND: The aim of this paper is to demonstrate the application of watermarks based on DNA sequences to identify the unauthorized use of genetically modified organisms (GMOs) protected by patents. Predicted mutations in the genome can be corrected by the DNA-Crypt program leaving the encrypted information intact. Existing DNA cryptographic and steganographic algorithms use synthetic DNA sequences to store binary information however, although these sequences can be used for authentication, they may change the target DNA sequence when introduced into living organisms. RESULTS: The DNA-Crypt algorithm and image steganography are based on the same watermark-hiding principle, namely using the least significant base in case of DNA-Crypt and the least significant bit in case of the image steganography. It can be combined with binary encryption algorithms like AES, RSA or Blowfish. DNA-Crypt is able to correct mutations in the target DNA with several mutation correction codes such as the Hamming-code or the WDH-code. Mutations which can occur infrequently may destroy the encrypted information, however an integrated fuzzy controller decides on a set of heuristics based on three input dimensions, and recommends whether or not to use a correction code. These three input dimensions are the length of the sequence, the individual mutation rate and the stability over time, which is represented by the number of generations. In silico experiments using the Ypt7 in Saccharomyces cerevisiae shows that the DNA watermarks produced by DNA-Crypt do not alter the translation of mRNA into protein. CONCLUSION: The program is able to store watermarks in living organisms and can maintain the original information by correcting mutations itself. Pairwise or multiple sequence alignments show that DNA-Crypt produces few mismatches between the sequences similar to all steganographic algorithms. Dominik Heider, Angelika Barnekow |
BMC Bioinform. | 1 |