Jan Baumbach

dblp:32/4397 · DBLP profile ↗
← Back
55ranked-venue papers
7as first author
28since 2021 · last 2026
0000-0002-0282-0462ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 42 · 3 first-author · 24 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 3 since 2021Theory of computation · 4 · 2 first-authorDatabases, data management, data science and information retrieval · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Detection of alternative splicing: deep sequencing or deep learning?
abstract
Alternative splicing is a crucial mechanism of gene regulation that enables condition- and tissue-specific expression of gene isoforms. Its dysregulation plays a role in various diseases such as cancer, neurological disorders, and metabolic conditions. Despite its importance, accurate detection of alternative splicing events remains challenging. Comprehensive alternative splicing event detection typically requires deep sequencing with over 100 million reads; however, much of the publicly accessible RNA sequencing data is of lower sequencing depth. Recent advances, particularly deep learning models working with genomic sequences, offer new avenues for predicting alternative splicing without reliance on high sequencing depth data. Our study addresses the question: Can we utilize the vast repository of publicly available RNA sequencing data for comprehensive alternative splicing detection, despite the low sequencing depth? Our results demonstrate the potential of sequence-based deep learning tools such as AlphaGenome, SpliceAI and DeepSplice for initial hypothesis development and as additional filters in standard RNA sequencing pipelines, especially when sequencing depth is limited. Nonetheless, validation with higher sequencing depths remains essential for confirmation of splice events. Overall, our findings underscore the need for integrative methods combining genomic sequence data and RNA sequencing data for the prediction of tissue- and condition-specific alternative splicing in resource-limited settings.
Lena Maria Hackl, Fabian Neuhaus, Sabine Ameling, Uwe Völker, Jan Baumbach, Olga Tsoy
Briefings Bioinform.5
2026 ViMOP: a user-friendly and field-applicable pipeline for untargeted viral genome nanopore sequencing
abstract
MOTIVATION: Untargeted, also known as metagenomic, nanopore sequencing is a powerful tool for virus genomic surveillance, particularly in resource-limited settings and when paired with the portability of the MinION device (Oxford Nanopore Technologies, ONT). However, a major bottleneck for global access is the absence of a user-friendly software capable of efficiently analyzing untargeted nanopore sequencing data to generate high-quality consensus genomes. RESULTS: We share ViMOP, a pipeline built on our long-term experience in nanopore field sequencing. The pipeline emphasizes field user-friendliness, flexibility and versatility to analyze reads generated directly from human clinical samples. The software assembles de novo contigs, matches contigs to known viral references and uses them to assemble consensus genomes. Executed with a single Nextflow command or via the EPI2ME Desktop interface (ONT), results are summarized in an HTML report. ViMOP, through its user-centered design, lowers the barrier to high-quality virus genome reconstruction and advances capacity for genomic surveillance. AVAILABILITY AND IMPLEMENTATION: ViMOP is freely available for non-commercial use (https://github.com/opr-group-bnitm/vimop and https://zenodo.org/records/17913089), along with the associated database (https://zenodo.org/records/17652512), the scripts used to generate it (https://zenodo.org/records/17632662) and benchmarking code (https://zenodo.org/records/17633185).
Nils Peter Petersen, Mia Le, Annick Renevey, Ehizojie Emua, Sarah Ryter, Giuditta Annibaldis, Jacob Camara, Sanaba Boumbaly, Cyril Erameh, Tanja Laske, Jan Baumbach, Philippe Lemey, Stephan Günther 0003, Sophie Duraffour, Liana Eleni Kafetzopoulou
Bioinform.11
2026 A systematic analysis of the impact of data variation on AI-based histopathological grading of prostate cancer
abstract
The histopathological evaluation of biopsies by human experts is a gold standard in clinical disease diagnosis. While recent artificial intelligence-based (AI) approaches have reached human expert-level performance, they often display shortcomings caused by variations in sample preparation, limiting clinical applicability. This study investigates the impact of data variation on AI-based histopathological grading and explores algorithmic approaches that confer prediction robustness. To evaluate the impact of data variation in histopathology, we collected a multicentric, retrospective, observational prostate cancer (PCa) trial consisting of six cohorts in 3 countries with 25,591 patients, 83,864 images. This includes a high-variance dataset of 8,157 patients and 28,236 images with variations in section thickness, staining protocol, and scanner. This unique training dataset enabled the development of an AI-based PCa grading framework by training on patient outcome, not subjective grading. It was made robust through several algorithmic adaptations, including domain adversarial training and credibility-guided color adaptation. We named the final grading framework PCAI. We compare PCAI to a BASE model and human experts on three external test cohorts, comprising 2,255 patients and 9,437 images. Variations in sample processing, particularly section thickness and staining time, significantly reduced the performance of AI-based PCa grading by up to 8.6 percentage points in the event-ordered concordance index (EOC-Index) thus highlighting serious risks for AI-based histopathological grading. Algorithmic improvements for model robustness, credibility, and training on high-variance data as well as outcome-based severity prediction give rise to robust models with grading performance surpassing experienced pathologists. We demonstrate how our algorithmic enhancements for greater robustness lead to significantly better performance, surpassing expert grading on EOC-Index and 5-year AUROC by up to 21.2 percentage points.
Patrick Fuhlert, Fabian Westhaeusser, Esther Dietrich, Maximilian Lennartz, Robin Khatri, Nico Kaiser, Pontus Röbeck, Roman David Bülow, Saskia Von Stillfried, Anja Witte, Sam Ladjevardi, Anders Drotte, Peter Severgårdh, Jan Baumbach, Victor G. Puelles, Michael Häggman, Michael Brehler, Peter Boor, Peter Walhagen, Anca Dragomir, Christer Busch, Markus Graefen, Ewert Bengtsson, Guido Sauter, Marina Zimmermann, Stefan Bonn
Medical Image Anal.14
2025 Systematic evaluation of normalization approaches in tandem mass tag and label-free protein quantification data using PRONE
abstract
Despite the significant progress in accuracy and reliability in mass spectrometry technology, as well as the development of strategies based on isotopic labeling or internal standards in recent decades, systematic biases originating from non-biological factors remain a significant challenge in data analysis. In addition, the wide range of available normalization methods renders the choice of a suitable normalization method challenging. We systematically evaluated 17 normalization and 2 batch effect correction methods, originally developed for preprocessing DNA microarray data but widely applied in proteomics, on 6 publicly available spike-in and 3 label-free and tandem mass tag datasets. Opposed to state-of-the-art normalization practice, we found that a reduction in intragroup variation is not directly related to the effectiveness of the normalization methods. Furthermore, our results demonstrated that the methods RobNorm and Normics, specifically developed for proteomics data, in line with LoessF performed consistently well across the spike-in datasets, while EigenMS exhibited a high false-positive rate. Finally, based on experimental data, we show that normalization substantially impacts downstream analyses, and the impact is highly dataset-specific, emphasizing the importance of use-case-specific evaluations for novel proteomics datasets. For this, we developed the PROteomics Normalization Evaluator (PRONE), a unifying R package enabling comparative evaluation of normalization methods, including their impact on downstream analyses, while offering considerable flexibility, acknowledging the lack of universally accepted standards. PRONE is available on Bioconductor with a web application accessible at https://exbio.wzw.tum.de/prone/.
Lis Arend, Klaudia Adamowicz, Johannes R. Schmidt, Yuliya Burankova, Olga I. Zolotareva, Olga Tsoy, Josch Pauling, Stefan Kalkhof, Jan Baumbach, Markus List, Tanja Laske
Briefings Bioinform.9
2025 Transcription factor prediction using protein 3D secondary structures
abstract
MOTIVATION: Transcription factors (TFs) are DNA-binding proteins that regulate gene expression. Traditional methods predict a protein as a TF if the protein contains any DNA-binding domains (DBDs) of known TFs. However, this approach fails to identify a novel TF that does not contain any known DBDs. Recently proposed TF prediction methods do not rely on DBDs. Such methods use features of protein sequences to train a machine learning model, and then use the trained model to predict whether a protein is a TF or not. Because the 3-dimensional (3D) structure of a protein captures more information than its sequence, using 3D protein structures will likely allow for more accurate prediction of novel TFs. RESULTS: We propose a deep learning-based TF prediction method (StrucTFactor), which is the first method to utilize 3D secondary structural information of proteins. We compare StrucTFactor with recent state-of-the-art TF prediction methods based on ∼525 000 proteins across 12 datasets, capturing different aspects of data bias (including sequence redundancy) possibly influencing a method's performance. We find that StrucTFactor significantly (P-value < 0.001) outperforms the existing TF prediction methods, improving the performance over its closest competitor by up to 17% based on Matthews correlation coefficient. AVAILABILITY AND IMPLEMENTATION: Data and source code are available at https://github.com/lieboldj/StrucTFactor and on our website at https://apps.cosy.bio/StrucTFactor.
Jeanine Liebold, Fabian Neuhaus, Janina Geiser, Stefan Kurtz, Jan Baumbach, Khalique Newaz
Bioinform.5
2025 DRaCOon: a novel algorithm for pathway-level differential co-expression analysis in transcriptomics
abstract
Understanding the molecular mechanisms underlying diseases is crucial for more precise, personalized medicine. Pathway-level differential co-expression analysis, a powerful approach for transcriptomics, identifies condition-specific changes in gene-gene interaction networks, offering targeted insights. However, a key challenge is the lack of robust methods and benchmarks specifically for evaluating algorithms' ability to identify disrupted gene-gene associations across conditions. We introduce DRaCOoN (Differential Regulatory and Co-expression Networks), a Python package and web tool for pathway-level differential co-expression analysis. DRaCOoN uniquely integrates multiple association and differential metrics, with a novel, computationally efficient permutation test for significance assessment. Crucially, DRaCOoN also provides a benchmarking framework for comprehensive method evaluation. Extensive benchmarking on simulated data and three real-world datasets (bone healing, colorectal cancer, and head/neck carcinoma) showed that DRaCOoN, particularly with an entropy-based association measure and the s differential metric, consistently outperforms eight other methods. It remains highly accurate in balanced datasets, robust to varying gene perturbation levels, and identifies biologically relevant regulatory changes. Furthermore, DRaCOoN serves as both a powerful tool and a benchmarking framework for elucidating disease mechanisms from transcriptomics data, advancing precision medicine by uncovering critical gene regulatory alterations.
Fernando M. Delgado-Chaves, Ferdinand Spurny, Tanja Laske, Mhaned Oubounyt, Jan Baumbach
BMC Bioinform.5
2025 Potentials and limitations in the application of Convolutional Neural Networks for mosquito species identification using wing images
abstract
This study addresses the pressing global health burden of mosquito-borne diseases by investigating the application of Convolutional Neural Networks (CNNs) for mosquito species identification using wing images. Conventional identification methods are hampered by the need for significant expertise and resources, while CNNs offer a promising alternative. Our research aimed to develop a reliable and applicable classification system that can be used under real-world conditions, with a focus on improving model adaptability to unencountered devices, mitigating dataset biases, and ensuring usability across different users without standardized protocols. We utilized a large, diverse dataset of mosquito wing images of 21 taxa and three image-capturing devices (N = 14,888) and a preprocessing pipeline to standardize images and remove undesirable image features. The developed CNN models demonstrated high performance, with an average balanced accuracy of 98.3% and a macro F1-score of 97.6%, effectively distinguishing between the 21 mosquito taxa, including morphologically similar pairs. The preprocessing pipeline improved the model's robustness, reducing performance drops on unfamiliar devices effectively. However, the study also highlights the persistence of inherent dataset biases, which the preprocessing steps could only partially mitigate. The classification system's practical usability was demonstrated through a feasibility study, showing high inter-rater reliability. The results underscore the potential of the proposed workflow to enhance vector surveillance, especially in resource-constrained settings, and suggest its applicability to other winged insect species. The classification system developed in this study is available for public use, providing a valuable tool for vector surveillance and research, supporting efforts to mitigate the spread of mosquito-borne diseases.
Kristopher Nolte, Jan Baumbach, Christian Lins, Jens Johann Georg Lohmann, Philip Kollmannsberger, Felix Gregor Sauer, Renke Lühken
PLoS Comput. Biol.2
2024 NeDRex-Web: An Interactive Web Tool for Drug Repurposing by Exploring Heterogeneous Molecular Networks
abstract
Finding new indications for approved drugs is a promising alternative to the often very lengthy and expensive process of de novo drug development. Systems medicine has brought forth several different approaches to tackle this important task. We recently published NeDRex, a network medicine tool for the identification of disease modules and drug repurposing. NeDRex-Web (https://web.nedrex.net) brings existing and new features of the NeDRex platform to a user-friendly and research-oriented web application, enabling online exploration of large heterogeneous molecular networks. Focusing mainly on drug repurposing, NeDRex-Web implements customizable disease module identification and drug prioritization workflows to support users of diverse backgrounds in their research. Users are assisted during every step of their analysis, including the definition of relevant input sets, the selection from various algorithms for module identification or drug prioritization, and the prioritization of the results by their statistical significance. A guided connectivity search provides an easy way to identify links between node sets of interest and can be used to create user-specific induced networks.
Andreas Maier 0009, Mahdie Rafiei, Elisa Anastasi, Olga I. Zolotareva, James Skelton, Maria L. Elkjaer, Ana I. Casas, Cristian Nogales, Harald H. H. W. Schmidt, Tim Kacprowski, David B. Blumenthal, Anil Wipat, Sepideh Sadegh, Jan Baumbach
BIBM14
2024 Partial reinforcement optimizer: An evolutionary optimization algorithm
Ahmad Taheri, Keyvan RahimiZadeh, Amin Beheshti, Jan Baumbach, Ravipudi Venkata Rao, Seyedali Mirjalili, Amir Hossein Gandomi
Expert Syst. Appl.4
2023 Differentially Private Federated Learning: Privacy and Utility Analysis of Output Perturbation and DP-SGD
abstract
Federated Learning (FL) is a method that allows multiple entities to jointly train a machine learning model using data located in various places. Unlike the conventional approach of gathering private data from distributed locations to a central place, federated learning involves solely exchanging and aggregating the machine learning models. Each party shares only a machine learning model trained locally on their private data, ensuring that the sensitive data remains within the respective silos throughout the process. However, these shared models in FL may still leak sensitive information about the training data in the form of e.g. membership disclosure. To mitigate these residual privacy risks in federated learning, one has to use additional defence techniques such as Differential Privacy (DP), which introduces noise into the training data or the model. Differential Privacy provides a mathematical definition of privacy and can be applied in machine learning via different perturbation mechanisms. This work focuses on the analysis of Differential Privacy in federated learning through (i) output perturbation of the trained machine learning models and (ii) a differentially-private form of stochastic gradient descent (DP-SGD). We consider these two approaches in various settings and analyse their performance in terms of model utility and achieved privacy. To evaluate a model’s privacy risk, we empirically measure the success rate of a membership inference attack. We observe that DP-SGD allows for a better trade-off between privacy and utility in most of the considered settings. In some settings, however, output perturbation can provide a better or similar privacy-utility trade-off and at the same time better communication and computational efficiency.
Anastasia Pustozerova, Jan Baumbach, Rudolf Mayer
IEEE Big Data2
2023 Human-in-the-Loop Integration with Domain-Knowledge Graphs for Explainable Federated Deep Learning
abstract
Abstract We explore the integration of domain knowledge graphs into Deep Learning for improved interpretability and explainability using Graph Neural Networks (GNNs). Specifically, a protein-protein interaction (PPI) network is masked over a deep neural network for classification, with patient-specific multi-modal genomic features enriched into the PPI graph’s nodes. Subnetworks that are relevant to the classification (referred to as “disease subnetworks”) are detected using explainable AI. Federated learning is enabled by dividing the knowledge graph into relevant subnetworks, constructing an ensemble classifier, and allowing domain experts to analyze and manipulate detected subnetworks using a developed user interface. Furthermore, the human-in-the-loop principle can be applied with the incorporation of experts, interacting through a sophisticated User Interface (UI) driven by Explainable Artificial Intelligence (xAI) methods, changing the datasets to create counterfactual explanations. The adapted datasets could influence the local model’s characteristics and thereby create a federated version that distils their diverse knowledge in a centralized scenario. This work demonstrates the feasibility of the presented strategies, which were originally envisaged in 2021 and most of it has now been materialized into actionable items. In this paper, we report on some lessons learned during this project.
Andreas Holzinger, Anna Saranti, Anne-Christin Hauschild, Jacqueline Michelle Metsch, Dominik Heider, Richard Röttger, Heimo Müller, Jan Baumbach, Bastian Pfeifer
CD-MAKE8
2023 Analysing Utility Loss in Federated Learning with Differential Privacy
abstract
Federated learning provides the solution when multiple parties want to collaboratively train a machine learning model without directly sharing sensitive data. In Federated Learning, each party trains a machine learning model locally on its private data and sends only the models’ weights or updates (gradients) to an aggregator, which averages locally trained models into a new global model with higher effectiveness. However, the machine learning models, which have to be shared during the federated learning process, can still leak sensitive information about their training data through e.g. membership inference attacks. Differential Privacy (DP) can mitigate privacy risks in federated learning by introducing noise into machine learning models. In this work, we consider two approaches for achieving Differential Privacy in federated learning: (i) output perturbation of the trained machine learning models and (ii) a differentially-private form of stochastic gradient descent (DP-SGD). We perform an extensive analysis of these two approaches in several federated settings and compare their performance in terms of model utility and achieved privacy. We observe that DP-SGD allows for a better trade-off between privacy and utility.
Anastasia Pustozerova, Jan Baumbach, Rudolf Mayer
TrustCom2
2023 spongEffects: ceRNA modules offer patient-specific insights into the miRNA regulatory landscape
abstract
MOTIVATION: Cancer is one of the leading causes of death worldwide. Despite significant improvements in prevention and treatment, mortality remains high for many cancer types. Hence, innovative methods that use molecular data to stratify patients and identify biomarkers are needed. Promising biomarkers can also be inferred from competing endogenous RNA (ceRNA) networks that capture the gene-miRNA gene regulatory landscape. Thus far, the role of these biomarkers could only be studied globally but not in a sample-specific manner. To mitigate this, we introduce spongEffects, a novel method that infers subnetworks (or modules) from ceRNA networks and calculates patient- or sample-specific scores related to their regulatory activity. RESULTS: We show how spongEffects can be used for downstream interpretation and machine learning tasks such as tumor classification and for identifying subtype-specific regulatory interactions. In a concrete example of breast cancer subtype classification, we prioritize modules impacting the biology of the different subtypes. In summary, spongEffects prioritizes ceRNA modules as biomarkers and offers insights into the miRNA regulatory landscape. Notably, these module scores can be inferred from gene expression data alone and can thus be applied to cohorts where miRNA expression information is lacking. AVAILABILITY AND IMPLEMENTATION: https://bioconductor.org/packages/devel/bioc/html/SPONGE.html.
Fabio Boniolo, Markus Hoffmann, Norman Roggendorf, Bahar Tercan, Jan Baumbach, Mauro A. A. Castro, Gordon Robertson, Dieter Saur, Markus List
Bioinform.5
2023 Systematic analysis of alternative splicing in time course data using Spycone
abstract
MOTIVATION: During disease progression or organism development, alternative splicing may lead to isoform switches that demonstrate similar temporal patterns and reflect the alternative splicing co-regulation of such genes. Tools for dynamic process analysis usually neglect alternative splicing. RESULTS: Here, we propose Spycone, a splicing-aware framework for time course data analysis. Spycone exploits a novel IS detection algorithm and offers downstream analysis such as network and gene set enrichment. We demonstrate the performance of Spycone using simulated and real-world data of SARS-CoV-2 infection. AVAILABILITY AND IMPLEMENTATION: The Spycone package is available as a PyPI package. The source code of Spycone is available under the GPLv3 license at https://github.com/yollct/spycone and the documentation at https://spycone.readthedocs.io/en/latest/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Chit Tong Lio, Gordon Grabert, Zakaria Louadi, Amit Fenn, Jan Baumbach, Tim Kacprowski, Markus List, Olga Tsoy
Bioinform.5
2023 Inference of differential key regulatory networks and mechanistic drug repurposing candidates from scRNA-seq data with SCANet
abstract
MOTIVATION: The reconstruction of small key regulatory networks that explain the differences in the development of cell (sub)types from single-cell RNA sequencing is a yet unresolved computational problem. RESULTS: To this end, we have developed SCANet, an all-in-one package for single-cell profiling that covers the whole differential mechanotyping workflow, from inference of trait/cell-type-specific gene co-expression modules, driver gene detection, and transcriptional gene regulatory network reconstruction to mechanistic drug repurposing candidate prediction. To illustrate the power of SCANet, we examined data from two studies. First, we identify the drivers of the mechanotype of a cytokine storm associated with increased mortality in patients with acute respiratory illness. Secondly, we find 20 drugs for eight potential pharmacological targets in cellular driver mechanisms in the intestinal stem cells of obese mice. AVAILABILITY AND IMPLEMENTATION: SCANet is a free, open-source, and user-friendly Python package that can be seamlessly integrated into single-cell-based systems medicine research and mechanistic drug discovery.
Mhaned Oubounyt, Lorenz Adlung, Fábio Malta de Sá Patroni, Nina Kerstin Wenke, Andreas Maier 0009, Michael Hartung, Jan Baumbach, Maria L. Elkjaer
Bioinform.7
2023 Online bias-aware disease module mining with ROBUST-Web
abstract
SUMMARY: We present ROBUST-Web which implements our recently presented ROBUST disease module mining algorithm in a user-friendly web application. ROBUST-Web features seamless downstream disease module exploration via integrated gene set enrichment analysis, tissue expression annotation, and visualization of drug-protein and disease-gene links. Moreover, ROBUST-Web includes bias-aware edge costs for the underlying Steiner tree model as a new algorithmic feature, which allow to correct for study bias in protein-protein interaction networks and further improves the robustness of the computed modules. AVAILABILITY AND IMPLEMENTATION: Web application: https://robust-web.net. Source code of web application and Python package with new bias-aware edge costs: https://github.com/bionetslab/robust-web, https://github.com/bionetslab/robust_bias_aware.
Suryadipto Sarkar, Marta Lucchetta, Andreas Maier 0009, Mohamed M. Abdrabbou, Jan Baumbach, Markus List, Martin H. Schaefer 0001, David B. Blumenthal
Bioinform.5
2022 Online in silico validation of disease and gene sets, clusterings or subnetworks with DIGEST
abstract
As the development of new drugs reaches its physical and financial limits, drug repurposing has become more important than ever. For mechanistically grounded drug repurposing, it is crucial to uncover the disease mechanisms and to detect clusters of mechanistically related diseases. Various methods for computing candidate disease mechanisms and disease clusters exist. However, in the absence of ground truth, in silico validation is challenging. This constitutes a major hurdle toward the adoption of in silico prediction tools by experimentalists who are often hesitant to carry out wet-lab validations for predicted candidate mechanisms without clearly quantified initial plausibility. To address this problem, we present DIGEST (in silico validation of disease and gene sets, clusterings or subnetworks), a Python-based validation tool available as a web interface (https://digest-validation.net), as a stand-alone package or over a REST API. DIGEST greatly facilitates in silico validation of gene and disease sets, clusterings or subnetworks via fully automated pipelines comprising disease and gene ID mapping, enrichment analysis, comparisons of shared genes and variants and background distribution estimation. Moreover, functionality is provided to automatically update the external databases used by the pipelines. DIGEST hence allows the user to assess the statistical significance of candidate mechanisms with regard to functional and genetic coherence and enables the computation of empirical $P$-values with just a few mouse clicks.
Klaudia Adamowicz, Andreas Maier 0009, Jan Baumbach, David B. Blumenthal
Briefings Bioinform.3
2022 A systematic comparison of novel and existing differential analysis methods for CyTOF data
abstract
Cytometry techniques are widely used to discover cellular characteristics at single-cell resolution. Many data analysis methods for cytometry data focus solely on identifying subpopulations via clustering and testing for differential cell abundance. For differential expression analysis of markers between conditions, only few tools exist. These tools either reduce the data distribution to medians, discarding valuable information, or have underlying assumptions that may not hold for all expression patterns. Here, we systematically evaluated existing and novel approaches for differential expression analysis on real and simulated CyTOF data. We found that methods using median marker expressions compute fast and reliable results when the data are not strongly zero-inflated. Methods using all data detect changes in strongly zero-inflated markers, but partially suffer from overprediction or cannot handle big datasets. We present a new method, CyEMD, based on calculating the earth mover's distance between expression distributions that can handle strong zero-inflation without being too sensitive. Additionally, we developed CYANUS - CYtometry ANalysis Using Shiny - a user-friendly R Shiny App allowing the user to analyze cytometry data with state-of-the-art tools, including well-performing methods from our comparison. A public web interface is available at https://exbio.wzw.tum.de/cyanus/.
Lis Arend, Judith Bernett, Quirin Manz, Melissa Klug, Olga Lazareva, Jan Baumbach, Dario Bongiovanni, Markus List
Briefings Bioinform.6
2022 Robust disease module mining via enumeration of diverse prize-collecting Steiner trees
abstract
MOTIVATION: Disease module mining methods (DMMMs) extract subgraphs that constitute candidate disease mechanisms from molecular interaction networks such as protein-protein interaction (PPI) networks. Irrespective of the employed models, DMMMs typically include non-robust steps in their workflows, i.e. the computed subnetworks vary when running the DMMMs multiple times on equivalent input. This lack of robustness has a negative effect on the trustworthiness of the obtained subnetworks and is hence detrimental for the widespread adoption of DMMMs in the biomedical sciences. RESULTS: To overcome this problem, we present a new DMMM called ROBUST (robust disease module mining via enumeration of diverse prize-collecting Steiner trees). In a large-scale empirical evaluation, we show that ROBUST outperforms competing methods in terms of robustness, scalability and, in most settings, functional relevance of the produced modules, measured via KEGG (Kyoto Encyclopedia of Genes and Genomes) gene set enrichment scores and overlap with DisGeNET disease genes. AVAILABILITY AND IMPLEMENTATION: A Python 3 implementation and scripts to reproduce the results reported in this article are available on GitHub: https://github.com/bionetslab/robust, https://github.com/bionetslab/robust-eval. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Judith Bernett, Dominik Krupke, Sepideh Sadegh, Jan Baumbach, Sándor P. Fekete, Tim Kacprowski, Markus List, David B. Blumenthal
Bioinform.4
2022 dsMTL: a computational framework for privacy-preserving, distributed multi-task machine learning
abstract
MOTIVATION: In multi-cohort machine learning studies, it is critical to differentiate between effects that are reproducible across cohorts and those that are cohort-specific. Multi-task learning (MTL) is a machine learning approach that facilitates this differentiation through the simultaneous learning of prediction tasks across cohorts. Since multi-cohort data can often not be combined into a single storage solution, there would be the substantial utility of an MTL application for geographically distributed data sources. RESULTS: Here, we describe the development of 'dsMTL', a computational framework for privacy-preserving, distributed multi-task machine learning that includes three supervised and one unsupervised algorithms. First, we derive the theoretical properties of these methods and the relevant machine learning workflows to ensure the validity of the software implementation. Second, we implement dsMTL as a library for the R programming language, building on the DataSHIELD platform that supports the federated analysis of sensitive individual-level data. Third, we demonstrate the applicability of dsMTL for comorbidity modeling in distributed data. We show that comorbidity modeling using dsMTL outperformed conventional, federated machine learning, as well as the aggregation of multiple models built on the distributed datasets individually. The application of dsMTL was computationally efficient and highly scalable when applied to moderate-size (n < 500), real expression data given the actual network latency. AVAILABILITY AND IMPLEMENTATION: dsMTL is freely available at https://github.com/transbioZI/dsMTLBase (server-side package) and https://github.com/transbioZI/dsMTLClient (client-side package). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Han Cao 0001, Youcheng Zhang, Jan Baumbach, Paul R. Burton, Dominic B. Dwyer, Nikolaos Koutsouleris, Julian O. Matschinske, Yannick Marcon, Sivanesan Rajan, Thilo Rieg, Patricia Ryser-Welch, Julian Späth, Carl Herrmann, Emanuel Schwarz
Bioinform.3
2022 Federated Random Forests can improve local performance of predictive models for various healthcare applications
abstract
MOTIVATION: Limited data access has hindered the field of precision medicine from exploring its full potential, e.g. concerning machine learning and privacy and data protection rules.Our study evaluates the efficacy of federated Random Forests (FRF) models, focusing particularly on the heterogeneity within and between datasets. We addressed three common challenges: (i) number of parties, (ii) sizes of datasets and (iii) imbalanced phenotypes, evaluated on five biomedical datasets. RESULTS: The FRF outperformed the average local models and performed comparably to the data-centralized models trained on the entire data. With an increasing number of models and decreasing dataset size, the performance of local models decreases drastically. The FRF, however, do not decrease significantly. When combining datasets of different sizes, the FRF vastly improve compared to the average local models. We demonstrate that the FRF remain more robust and outperform the local models by analyzing different class-imbalances.Our results support that FRF overcome boundaries of clinical research and enables collaborations across institutes without violating privacy or legal regulations. Clinicians benefit from a vast collection of unbiased data aggregated from different geographic locations, demographics and other varying factors. They can build more generalizable models to make better clinical decisions, which will have relevance, especially for patients in rural areas and rare or geographically uncommon diseases, enabling personalized treatment. In combination with secure multi-party computation, federated learning has the power to revolutionize clinical practice by increasing the accuracy and robustness of healthcare AI and thus paving the way for precision medicine. AVAILABILITY AND IMPLEMENTATION: The implementation of the federated random forests can be found at https://featurecloud.ai/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Anne-Christin Hauschild, Marta Lemanczyk, Julian O. Matschinske, Tobias Frisch, Olga I. Zolotareva, Andreas Holzinger, Jan Baumbach, Dominik Heider
Bioinform.7
2021 On the Privacy of Federated Pipelines
abstract
Federated learning (FL) is becoming an increasingly popular machine learning paradigm in application scenarios where sensitive data available at various local sites cannot be shared due to privacy protection regulations. In FL, the sensitive data never leaves the local sites and only model parameters are shared with a global aggregator. Nonetheless, it has recently been shown that, under some circumstances, the private data can be reconstructed from the model parameters, which implies that data leakage can occur in FL. In this paper, we draw attention to another risk associated with FL: Even if federated algorithms are individually privacy-preserving, combining them into pipelines is not necessarily privacy-preserving. We provide a concrete example from genome-wide association studies, where the combination of federated principal component analysis and federated linear regression allows the aggregator to retrieve sensitive patient data by solving an instance of the multidimensional subset sum problem. This supports the increasing awareness in the field that, for FL to be truly privacy-preserving, measures have to be undertaken to protect against data leakage at the aggregator.
Reza Nasirigerdeh, Reihaneh Torkzadehmahani, Jan Baumbach, David B. Blumenthal
SIGIR3
2021 Computational strategies to combat COVID-19: useful tools to accelerate SARS-CoV-2 and coronavirus research
abstract
SARS-CoV-2 (severe acute respiratory syndrome coronavirus 2) is a novel virus of the family Coronaviridae. The virus causes the infectious disease COVID-19. The biology of coronaviruses has been studied for many years. However, bioinformatics tools designed explicitly for SARS-CoV-2 have only recently been developed as a rapid reaction to the need for fast detection, understanding and treatment of COVID-19. To control the ongoing COVID-19 pandemic, it is of utmost importance to get insight into the evolution and pathogenesis of the virus. In this review, we cover bioinformatics workflows and tools for the routine detection of SARS-CoV-2 infection, the reliable analysis of sequencing data, the tracking of the COVID-19 pandemic and evaluation of containment measures, the study of coronavirus evolution, the discovery of potential drug targets and development of therapeutic strategies. For each tool, we briefly describe its use case and how it advances research specifically for SARS-CoV-2. All tools are free to use and available online, either through web applications or public code repositories. Contact:[email protected].
Franziska Hufsky, Kevin Lamkiewicz, Alexandre Almeida, Abdel Aouacheria, Cecilia N. Arighi, Alex Bateman, Jan Baumbach, Niko Beerenwinkel, Christian Brandt, Marco Cacciabue, Sara Chuguransky, Oliver Drechsel, Robert D. Finn, Adrian Fritz, Stephan Fuchs, Georges Hattab, Anne-Christin Hauschild, Dominik Heider, Marie Hoffmann, Martin Hölzer, Stefan Hoops, Lars Kaderali, Ioanna Kalvari, Max von Kleist, Renó Kmiecinski, Denise Kühnert, Gorka Lasso, Pieter Libin, Markus List, Hannah F. Löchel, Maria Jesus Martin, Roman Martin, Julian O. Matschinske, Alice C. McHardy, Pedro Mendes 0001, Jaina Mistry, Vincent Navratil, Eric P. Nawrocki, Áine Niamh O'toole, Nancy Ontiveros-Palacios, Anton I. Petrov, Guillermo Rangel-Pineros, Nicole Redaschi, Susanne Reimering, Knut Reinert, Lorna J. Richardson, David L. Robertson, Sepideh Sadegh, Joshua B. Singer, Kristof Theys, Chris Upton, Marius Welzel, Lowri Williams, Manja Marz
Briefings Bioinform.7
2021 On the limits of active module identification
abstract
In network and systems medicine, active module identification methods (AMIMs) are widely used for discovering candidate molecular disease mechanisms. To this end, AMIMs combine network analysis algorithms with molecular profiling data, most commonly, by projecting gene expression data onto generic protein-protein interaction (PPI) networks. Although active module identification has led to various novel insights into complex diseases, there is increasing awareness in the field that the combination of gene expression data and PPI network is problematic because up-to-date PPI networks have a very small diameter and are subject to both technical and literature bias. In this paper, we report the results of an extensive study where we analyzed for the first time whether widely used AMIMs really benefit from using PPI networks. Our results clearly show that, except for the recently proposed AMIM DOMINO, the tested AMIMs do not produce biologically more meaningful candidate disease modules on widely used PPI networks than on random networks with the same node degrees. AMIMs hence mainly learn from the node degrees and mostly fail to exploit the biological knowledge encoded in the edges of the PPI networks. This has far-reaching consequences for the field of active module identification. In particular, we suggest that novel algorithms are needed which overcome the degree bias of most existing AMIMs and/or work with customized, context-specific networks instead of generic PPI networks.
Olga Lazareva, Jan Baumbach, Markus List, David B. Blumenthal
Briefings Bioinform.2
2021 A framework for modeling epistatic interaction
abstract
MOTIVATION: Recently, various tools for detecting single nucleotide polymorphisms (SNPs) involved in epistasis have been developed. However, no studies evaluate the employed statistical epistasis models such as the χ2-test or quadratic regression independently of the tools that use them. Such an independent evaluation is crucial for developing improved epistasis detection tools, for it allows to decide if a tool's performance should be attributed to the epistasis model or to the optimization strategy run on top of it. RESULTS: We present a protocol for evaluating epistasis models independently of the tools they are used in and generalize existing models designed for dichotomous phenotypes to the categorical and quantitative case. In addition, we propose a new model which scores candidate SNP sets by computing maximum likelihood distributions for the observed phenotypes in the cells of their penetrance tables. Extensive experiments show that the proposed maximum likelihood model outperforms three widely used epistasis models in most cases. The experiments also provide valuable insights into the properties of existing models, for instance, that quadratic regression perform particularly well on instances with quantitative phenotypes. AVAILABILITY AND IMPLEMENTATION: The evaluation protocol and all compared models are implemented in C++ and are supported under Linux and macOS. They are available at https://github.com/baumbachlab/genepiseeker/, along with test datasets and scripts to reproduce the experiments. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
David B. Blumenthal, Jan Baumbach, Markus Hoffmann, Tim Kacprowski, Markus List
Bioinform.2
2021 BiCoN: network-constrained biclustering of patients and omics data
abstract
MOTIVATION: Unsupervised learning approaches are frequently used to stratify patients into clinically relevant subgroups and to identify biomarkers such as disease-associated genes. However, clustering and biclustering techniques are oblivious to the functional relationship of genes and are thus not ideally suited to pinpoint molecular mechanisms along with patient subgroups. RESULTS: We developed the network-constrained biclustering approach Biclustering Constrained by Networks (BiCoN) which (i) restricts biclusters to functionally related genes connected in molecular interaction networks and (ii) maximizes the difference in gene expression between two subgroups of patients. This allows BiCoN to simultaneously pinpoint molecular mechanisms responsible for the patient grouping. Network-constrained clustering of genes makes BiCoN more robust to noise and batch effects than typical clustering and biclustering methods. BiCoN can faithfully reproduce known disease subtypes as well as novel, clinically relevant patient subgroups, as we could demonstrate using breast and lung cancer datasets. In summary, BiCoN is a novel systems medicine tool that combines several heuristic optimization strategies for robust disease mechanism extraction. BiCoN is well-documented and freely available as a python package or a web interface. AVAILABILITY AND IMPLEMENTATION: PyPI package: https://pypi.org/project/bicon. WEB INTERFACE: https://exbio.wzw.tum.de/bicon. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Olga Lazareva, Stefan Canzar, Kevin Yuan, Jan Baumbach, David B. Blumenthal, Paolo Tieri, Tim Kacprowski, Markus List
Bioinform.4
2021 EWASex: an efficient R-package to predict sex in epigenome-wide association studies
abstract
SUMMARY: Epigenome-Wide Association Study (EWAS) has become a powerful approach to identify epigenetic variations associated with diseases or health traits. Sex is an important variable to include in EWAS to ensure unbiased data processing and statistical analysis. We introduce the R-package EWASex, which allows for fast and highly accurate sex-estimation using DNA methylation data on a small set of CpG sites located on the X-chromosome under stable X-chromosome inactivation in females. RESULTS: We demonstrate that EWASex outperforms the current state of the art tools by using different EWAS datasets. With EWASex, we offer an efficient way to predict and to verify sex that can be easily implemented in any EWAS using blood samples or even other tissue types. It comes with pre-trained weights to work without prior sex labels and without requiring access to RAW data, which is a necessity for all currently available methods. AVAILABILITY AND IMPLEMENTATION: The EWASex R-package along with tutorials, documentation and source code are available at https://github.com/Silver-Hawk/EWASex. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Jesper Lund, Afsaneh Mohammadnejad, Shuxia Li, Jan Baumbach, Qihua Tan
Bioinform.5
2021 ASimulatoR: splice-aware RNA-Seq data simulation
abstract
SUMMARY: A plethora of tools exist for RNA-Seq data analysis with a focus on alternative splicing (AS). However, appropriate data for their comparative evaluation is missing. The R package ASimulatoR simulates gold standard RNA-Seq datasets with fine-grained control over the distribution of AS events, which allow for evaluating alternative splicing tools, e.g. to study the effect of sequencing depth on the performance of AS event detection. AVAILABILITY AND IMPLEMENTATION: ASimulatoR is freely available at https://github.com/biomedbigdata/ASimulatoR as an R package under GPL-3 license.
Quirin Manz, Olga Tsoy, Amit Fenn, Jan Baumbach, Uwe Völker, Markus List, Tim Kacprowski
Bioinform.4
2020 EpiGEN: an epistasis simulation pipeline
abstract
SUMMARY: Simulated data are crucial for evaluating epistasis detection tools in genome-wide association studies. Existing simulators are limited, as they do not account for linkage disequilibrium (LD), support limited interaction models of single nucleotide polymorphisms (SNPs) and only dichotomous phenotypes or depend on proprietary software. In contrast, EpiGEN supports SNP interactions of arbitrary order, produces realistic LD patterns and generates both categorical and quantitative phenotypes. AVAILABILITY AND IMPLEMENTATION: EpiGEN is implemented in Python 3 and is freely available at https://github.com/baumbachlab/epigen. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
David B. Blumenthal, Lorenzo Viola, Markus List, Jan Baumbach, Paolo Tieri, Tim Kacprowski
Bioinform.4
2019 Fuzzy Inference System for Risk Evaluation in Gestational Diabetes Mellitus
abstract
Remote monitoring health data analysis holds the potential to reduce pregnancy complications, improve patients' quality of life, enhance the efficiency of healthcare delivery and reduce healthcare costs. In this paper, we present a method based on fuzzy inference systems to monitor pregnancies complicated by gestational diabetes mellitus (GDM). The system is simple, fast, flexible and exploits domain expertise in assessing risk levels according to capillary glucose levels from women with GDM. We show that this approach generates an interpretable input, which is valuable in medical applications. To prove the capabilities of the system, we present prediction results from 50 real-world patients and show that the system obtains relevant glycaemic-control data comparable to current monitoring methods that rely on periodic face-to-face physician review. Our systems achieves 95% accuracy. Moreover, we show that the difference in predictions account for a more personalized treatment.
Carlos Salort Sánchez, Suzanne Smyth, Elizabeth Tully, Joanna Griffin, Luke Heaphy, Niamh Redmond, Fionnuala Breathnach, Jan Baumbach, Cristian Axenie
BIBE8
2019 An Online Incremental Clustering Framework for Real-Time Stream Analytics
abstract
With the evolution of data acquisition methods, our ability to collect real time data has increased. This requires the development of real-time analytics, using the most recent data to generate valuable insights. One example is customer profiling, where we want to identify groups of similar clients who were active recently, and improve the quality of the suggestions. Traditional clustering algorithms perform well on finite datasets, but their execution is often not compatible with real-time requirements, especially for rapid changing trends. In this context, we propose a novel approach for the definition of incremental clustering algorithms to work within real-time constraints, in an online fashion, while preserving accuracy. We show the general applicability of the framework by employing this method to three different clustering algorithms. We compare the experimental results between traditional and online approaches evaluating accuracy and computational cost. The results show that algorithms executed in our framework are comparable to their offline implementation in terms of accuracy and with a high gain in execution time, up to three orders of magnitude on average.
Carlos Salort Sánchez, Radu Tudoran, Mohamad Al Hajj Hassan, Stefano Bortoli, Goetz Brasche, Jan Baumbach, Cristian Axenie
ICMLA6
2019 Community effort endorsing multiscale modelling, multiscale data science and multiscale computing for systems medicine
abstract
Systems medicine holds many promises, but has so far provided only a limited number of proofs of principle. To address this road block, possible barriers and challenges of translating systems medicine into clinical practice need to be identified and addressed. The members of the European Cooperation in Science and Technology (COST) Action CA15120 Open Multiscale Systems Medicine (OpenMultiMed) wish to engage the scientific community of systems medicine and multiscale modelling, data science and computing, to provide their feedback in a structured manner. This will result in follow-up white papers and open access resources to accelerate the clinical translation of systems medicine.
Massimiliano Zanin, Ivan Chorbev, Blaz Stres, Egils Stalidzans, Julio Vera, Paolo Tieri, Filippo Castiglione, Derek Groen, Huiru Zheng, Jan Baumbach, Johannes A. Schmid, José Basilio, Peter Klimek, Natasa Debeljak, Damjana Rozman, Harald H. H. W. Schmidt
Briefings Bioinform.10
2019 Comprehensive evaluation of transcriptome-based cell-type quantification methods for immuno-oncology
abstract
MOTIVATION: The composition and density of immune cells in the tumor microenvironment (TME) profoundly influence tumor progression and success of anti-cancer therapies. Flow cytometry, immunohistochemistry staining or single-cell sequencing are often unavailable such that we rely on computational methods to estimate the immune-cell composition from bulk RNA-sequencing (RNA-seq) data. Various methods have been proposed recently, yet their capabilities and limitations have not been evaluated systematically. A general guideline leading the research community through cell type deconvolution is missing. RESULTS: We developed a systematic approach for benchmarking such computational methods and assessed the accuracy of tools at estimating nine different immune- and stromal cells from bulk RNA-seq samples. We used a single-cell RNA-seq dataset of ∼11 000 cells from the TME to simulate bulk samples of known cell type proportions, and validated the results using independent, publicly available gold-standard estimates. This allowed us to analyze and condense the results of more than a hundred thousand predictions to provide an exhaustive evaluation across seven computational methods over nine cell types and ∼1800 samples from five simulated and real-world datasets. We demonstrate that computational deconvolution performs at high accuracy for well-defined cell-type signatures and propose how fuzzy cell-type signatures can be improved. We suggest that future efforts should be dedicated to refining cell population definitions and finding reliable signatures. AVAILABILITY AND IMPLEMENTATION: A snakemake pipeline to reproduce the benchmark is available at https://github.com/grst/immune_deconvolution_benchmark. An R package allows the community to perform integrated deconvolution using different methods (https://grst.github.io/immunedeconv). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Gregor Sturm, Francesca Finotello, Florent Petitprez, Jitao David Zhang, Jan Baumbach, Wolf-Herman Fridman, Markus List, Tatsiana Aneichyk
Bioinform.5
2019 PanRV: Pangenome-reverse vaccinology approach for identifications of potential vaccine candidates in microbial pangenome
abstract
BACKGROUND: A revolutionary diversion from classical vaccinology to reverse vaccinology approach has been observed in the last decade. The ever-increasing genomic and proteomic data has greatly facilitated the vaccine designing and development process. Reverse vaccinology is considered as a cost-effective and proficient approach to screen the entire pathogen genome. To look for broad-spectrum immunogenic targets and analysis of closely-related bacterial species, the assimilation of pangenome concept into reverse vaccinology approach is essential. The categories of species pangenome such as core, accessory, and unique genes sets can be analyzed for the identification of vaccine candidates through reverse vaccinology. RESULTS: We have designed an integrative computational pipeline term as "PanRV" that employs both the pangenome and reverse vaccinology approaches. PanRV comprises of four functional modules including i) Pangenome Estimation Module (PGM) ii) Reverse Vaccinology Module (RVM) iii) Functional Annotation Module (FAM) and iv) Antibiotic Resistance Association Module (ARM). The pipeline is tested by using genomic data from 301 genomes of Staphylococcus aureus and the results are verified by experimentally known antigenic data. CONCLUSION: The proposed pipeline has proved to be the first comprehensive automated pipeline that can precisely identify putative vaccine candidates exploiting the microbial pangenome. PanRV is a Linux based package developed in JAVA language. An executable installer is provided for ease of installation along with a user manual at https://sourceforge.net/projects/panrv2/ .
Kanwal Naz, Anam Naz, Shifa Tariq Ashraf, Jamil Ahmad 0002, Jan Baumbach, Amjad Ali 0001
BMC Bioinform.6
2018 From gene panels to systems medicine
Jan Baumbach
BIBM1
2018 Weighted gene co-expression network analysis of microarray mRNA expression profiling in response to electroacupuncture
Afsaneh Mohammadnejad, Shuxia Li, Hongmei Duan, Jesper Lund, Jan Baumbach, Qihua Tan
BIBM6
2018 On the power of epigenome-wide association studies using a disease-discordant twin design
abstract
Motivation: Many studies have investigated the association between DNA methylation alterations and disease occurrences using two design paradigms, traditional case-control and disease-discordant twins. In the disease-discordant twin design, the affected twin serves as the case and the unaffected twin serves as the control. Theoretically the twin design takes advantage of controlling for the shared genetic make-up, but it is still highly debatable if and how much researchers may benefit from such a design over the traditional case-control design. Results: In this study, we investigate and compare the power of both designs with simulations. A liability threshold model was used assuming that identical twins share the same genetic contribution with respect to the liability of complex human diseases. Varying ranges of parameters have been used to ensure that the simulation is close to real-world scenarios. Our results reveal that the disease-discordant twin design implies greater statistical power over the traditional case-control design. For diseases with moderate and high heritability (>0.3), the disease-discordant twin design allows for large sample size reductions compared to the ordinary case-control design. Our simulation results indicate that the discordant twin design is indeed a powerful tool for epigenetic association studies. Availability and implementation: Computer scripts are available at https://github.com/zickyls/EWAS-Twin-Simulation. Supplementary information: Supplementary data are available at Bioinformatics online.
Lene Christiansen, Jacob v. B. Hjelmborg, Jan Baumbach, Qihua Tan
Bioinform.4
2017 Efficient detection of differentially methylated regions using DiMmeR
abstract
Motivation: Epigenome-wide association studies (EWAS) generate big epidemiological datasets. They aim for detecting differentially methylated DNA regions that are likely to influence transcriptional gene activity and, thus, the regulation of metabolic processes. The by far most widely used technology is the Illumina Methylation BeadChip, which measures the methylation levels of 450 (850) thousand cytosines, in the CpG dinucleotide context in a set of patients compared to a control group. Many bioinformatics tools exist for raw data analysis. However, most of them require some knowledge in the programming language R, have no user interface, and do not offer all necessary steps to guide users from raw data all the way down to statistically significant differentially methylated regions (DMRs) and the associated genes. Results: Here, we present DiMmeR (Discovery of Multiple Differentially Methylated Regions), the first free standalone software that interactively guides with a user-friendly graphical user interface (GUI) scientists the whole way through EWAS data analysis. It offers parallelized statistical methods for efficiently identifying DMRs in both Illumina 450K and 850K EPIC chip data. DiMmeR computes empirical P -values through randomization tests, even for big datasets of hundreds of patients and thousands of permutations within a few minutes on a standard desktop PC. It is independent of any third-party libraries, computes regression coefficients, P -values and empirical P -values, and it corrects for multiple testing. Availability and Implementation: DiMmeR is publicly available at http://dimmer.compbio.sdu.dk . Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Diogo Almeida, Ida Skov, Artur Silva, Fabio Vandin, Qihua Tan, Richard Röttger, Jan Baumbach
Bioinform.7
2016 A Simulated Annealing Algorithm for Maximum Common Edge Subgraph Detection in Biological Networks
abstract
Network alignment is a challenging computational problem that identifies node or edge mappings between two or more networks, with the aim to unravel common patterns among them. Pairwise network alignment is already intractable, making multiple network comparison even more difficult. Here, we introduce a heuristic algorithm for the multiple maximum common edge subgraph problem that is able to detect large common substructures shared across multiple, real-world size networks efficiently. Our algorithm uses a combination of iterated local search, simulated annealing and a pheromone-based perturbation strategy. We implemented multiple local search strategies and annealing schedules, that were evaluated on a range of synthetic networks and real protein-protein interaction networks. Our method is parallelized and well-suited to exploit current multi-core CPU architectures. While it is generic, we apply it to unravel a biochemical backbone inherent in different species, modeled as multiple maximum common subgraphs.
Simon J. Larsen, Frederik G. Alkærsig, Henrik J. Ditzel, Igor Jurisica, Nicolas Alcaraz, Jan Baumbach
GECCO6
2016 CytoGEDEVO - global alignment of biological networks with Cytoscape
abstract
MOTIVATION: In the systems biology era, high-throughput omics technologies have enabled the unraveling of the interplay of some biological entities on a large scale (e.g. genes, proteins, metabolites or RNAs). Huge biological networks have emerged, where nodes correspond to these entities and edges between them model their relations. Protein-protein interaction networks, for instance, show the physical interactions of proteins in an organism. The comparison of such networks promises additional insights into protein and cell function as well as knowledge-transfer across species. Several computational approaches have been developed previously to solve the network alignment (NA) problem, but only a few concentrate on the usability of the implemented tools for the evaluation of protein-protein interactions by the end users (biologists and medical researchers). RESULTS: We have created CytoGEDEVO, a Cytoscape app for visual and user-assisted NA. It extends the previous GEDEVO methodology for global pairwise NAs with new graphical and functional features. Our main focus was on the usability, even by non-programmers and the interpretability of the NA results with Cytoscape. AVAILABILITY AND IMPLEMENTATION: CytoGEDEVO is publicly available from the Cytoscape app store at http://apps.cytoscape.org/apps/cytogedevo In addition, we provide stand-alone command line executables, source code, documentation and step-by-step user instructions at http://cytogedevo.compbio.sdu.dk CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Maximilian Malek, Rashid Ibragimov, Mario Albrecht, Jan Baumbach
Bioinform.4
2014 Complexity of Dense Bicluster Editing Problems
Peng Sun 0008, Jiong Guo, Jan Baumbach
COCOON3
2014 Compactness-Preserving Mapping on Trees
Jan Baumbach, Jiong Guo, Rashid Ibragimov
CPM1
2014 Multiple graph edit distance: simultaneous topological alignment of multiple protein-protein interaction networks with an evolutionary algorithm
abstract
Motivation: We address the problem of multiple protein-protein interaction (PPI) network alignment. Given a set of such networks for different species we might ask how much the network topology is conserved throughout evolution. Solving this problem will help to derive a subset of interactions that is conserved over multiple species thus forming a 'core interactome'. Methods: We model the problem as Topological Multiple one-to-one Network Alignment (TMNA), where we aim to minimize the total Graph Edit Distance (GED) between pairs of the input networks. Here, the GED between two graphs is the number of deleted and inserted edges that are required to make one graph isomorphic to another. By minimizing the GED we indirectly maximize the number of edges that are aligned in multiple networks simultaneously. However, computing an optimal GED value is computationally intractable. We thus propose an evolutionary algorithm and developed a software tool, GEDEVO-M, which is able to align multiple PPI networks using topological information only. We demonstrate the power of our approach by computing a maximal common subnetwork for a set of bacterial and eukaryotic PPI networks. GEDEVO-M thus provides great potential for computing the 'core interactome' of different species. Availability: http://gedevo.mpi-inf.mpg.de/multiple-network-alignment/.
Rashid Ibragimov, Maximilian Malek, Jan Baumbach, Jiong Guo
GECCO3
2014 Microarray R-based analysis of complex lysate experiments with MIRACLE
abstract
MOTIVATION: Reverse-phase protein arrays (RPPAs) allow sensitive quantification of relative protein abundance in thousands of samples in parallel. Typical challenges involved in this technology are antibody selection, sample preparation and optimization of staining conditions. The issue of combining effective sample management and data analysis, however, has been widely neglected. RESULTS: This motivated us to develop MIRACLE, a comprehensive and user-friendly web application bridging the gap between spotting and array analysis by conveniently keeping track of sample information. Data processing includes correction of staining bias, estimation of protein concentration from response curves, normalization for total protein amount per sample and statistical evaluation. Established analysis methods have been integrated with MIRACLE, offering experimental scientists an end-to-end solution for sample management and for carrying out data analysis. In addition, experienced users have the possibility to export data to R for more complex analyses. MIRACLE thus has the potential to further spread utilization of RPPAs as an emerging technology for high-throughput protein analysis. AVAILABILITY: Project URL: http://www.nanocan.org/miracle/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Markus List, Ines Block, Marlene Lemvig Pedersen, Helle Christiansen, Steffen Schmidt, Mads Thomassen, Qihua Tan, Jan Baumbach, Jan Mollenhauer
Bioinform.8
2013 Cluster Editing
Sebastian Böcker, Jan Baumbach
CiE2
2013 Covering Tree with Stars
Jan Baumbach, Jiong Guo, Rashid Ibragimov
COCOON1
2013 Neighborhood-Preserving Mapping between Trees
Jan Baumbach, Jiong Guo, Rashid Ibragimov
WADS1
2013 Density parameter estimation for finding clusters of homologous proteins - tracing actinobacterial pathogenicity lifestyles
abstract
MOTIVATION: Homology detection is a long-standing challenge in computational biology. To tackle this problem, typically all-versus-all BLAST results are coupled with data partitioning approaches resulting in clusters of putative homologous proteins. One of the main problems, however, has been widely neglected: all clustering tools need a density parameter that adjusts the number and size of the clusters. This parameter is crucial but hard to estimate without gold standard data at hand. Developing a gold standard, however, is a difficult and time consuming task. Having a reliable method for detecting clusters of homologous proteins between a huge set of species would open opportunities for better understanding the genetic repertoire of bacteria with different lifestyles. RESULTS: Our main contribution is a method for identifying a suitable and robust density parameter for protein homology detection without a given gold standard. Therefore, we study the core genome of 89 actinobacteria. This allows us to incorporate background knowledge, i.e. the assumption that a set of evolutionarily closely related species should share a comparably high number of evolutionarily conserved proteins (emerging from phylum-specific housekeeping genes). We apply our strategy to find genes/proteins that are specific for certain actinobacterial lifestyles, i.e. different types of pathogenicity. The whole study was performed with transitivity clustering, as it only requires a single intuitive density parameter and has been shown to be well applicable for the task of protein sequence clustering. Note, however, that the presented strategy generally does not depend on our clustering method but can easily be adapted to other clustering approaches. AVAILABILITY: All results are publicly available at http://transclust.mmci.uni-saarland.de/actino_core/ or as Supplementary Material of this article. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Richard Röttger, Prabhav Kalaghatgi, Peng Sun 0008, Siomar de Castro Soares, Vasco Ariston de Carvalho Azevedo, Tobias Wittkop, Jan Baumbach
Bioinform.7
2012 Efficient algorithms for extracting biological key pathways with global constraints
abstract
The integrated analysis of data of different types and with various interdependencies is one of the major challenges in computational biology. Recently, we developed KeyPathwayMiner, a method that combines biological networks modeled as graphs with disease-specific genetic expression data gained from a set of cases (patients, cell lines, tissues, etc.). We aimed for finding all maximal connected sub-graphs where all nodes but $K$ are expressed in all cases but at most $L$, i.e. key pathways. Thereby, we combined biological networks with OMICS data, instead of analyzing these data sets in isolation. Here we present an alternative approach that avoids a certain bias towards hub nodes: We now aim for extracting all maximal connected sub-networks where all but at most $K$ nodes are expressed in all cases but in total (!) at most $L$, i.e. accumulated over all cases and all nodes in a solution. We call this strategy GLONE (global node exceptions); the previous problem we call INES (individual node exceptions). Since finding GLONE-components is computationally hard, we developed an Ant Colony Optimization algorithm and implemented it with the KeyPathwayMiner Cytoscape framework as an alternative to the INES algorithms. KeyPathwayMiner 3.0 now offers both the INES and the GLONE algorithms. It is available as plugin from Cytoscape and online at http://keypathwayminer.mpi-inf.mpg.de.
Jan Baumbach, Tobias Friedrich 0001, Timo Kötzing, Anton Krohmer, Josch Pauling
GECCO1
2012 How Little Do We Actually Know? On the Size of Gene Regulatory Networks
abstract
The National Center for Biotechnology Information (NCBI) recently announced the availability of whole genome sequences for more than 1,000 species. And the number of sequenced individual organisms is growing. Ongoing improvement of DNA sequencing technology will further contribute to this, enabling large-scale evolution and population genetics studies. However, the availability of sequence information is only the first step in understanding how cells survive, reproduce, and adjust their behavior. The genetic control behind organized development and adaptation of complex organisms still remains widely undetermined. One major molecular control mechanism is transcriptional gene regulation. The direct juxtaposition of the total number of sequenced species to the handful of model organisms with known regulations is surprising. Here, we investigate how little we even know about these model organisms. We aim to predict the sizes of the whole-organism regulatory networks of seven species. In particular, we provide statistical lower bounds for the expected number of regulations. For Escherichia coli we estimate at most 37 percent of the expected gene regulatory interactions to be already discovered, 24 percent for Bacillus subtilis, and <3% human, respectively. We conclude that even for our best researched model organisms we still lack substantial understanding of fundamental molecular control mechanisms, at least on a large scale.
Richard Röttger, Ulrich Rückert 0002, Jan Taubert, Jan Baumbach
IEEE ACM Trans. Comput. Biol. Bioinform.4
2011 clusterMaker: a multi-algorithm clustering plugin for Cytoscape
abstract
BACKGROUND: In the post-genomic era, the rapid increase in high-throughput data calls for computational tools capable of integrating data of diverse types and facilitating recognition of biologically meaningful patterns within them. For example, protein-protein interaction data sets have been clustered to identify stable complexes, but scientists lack easily accessible tools to facilitate combined analyses of multiple data sets from different types of experiments. Here we present clusterMaker, a Cytoscape plugin that implements several clustering algorithms and provides network, dendrogram, and heat map views of the results. The Cytoscape network is linked to all of the other views, so that a selection in one is immediately reflected in the others. clusterMaker is the first Cytoscape plugin to implement such a wide variety of clustering algorithms and visualizations, including the only implementations of hierarchical clustering, dendrogram plus heat map visualization (tree view), k-means, k-medoid, SCPS, AutoSOME, and native (Java) MCL. RESULTS: Results are presented in the form of three scenarios of use: analysis of protein expression data using a recently published mouse interactome and a mouse microarray data set of nearly one hundred diverse cell/tissue types; the identification of protein complexes in the yeast Saccharomyces cerevisiae; and the cluster analysis of the vicinal oxygen chelate (VOC) enzyme superfamily. For scenario one, we explore functionally enriched mouse interactomes specific to particular cellular phenotypes and apply fuzzy clustering. For scenario two, we explore the prefoldin complex in detail using both physical and genetic interaction clusters. For scenario three, we explore the possible annotation of a protein as a methylmalonyl-CoA epimerase within the VOC superfamily. Cytoscape session files for all three scenarios are provided in the Additional Files section. CONCLUSIONS: The Cytoscape plugin clusterMaker provides a number of clustering algorithms and visualizations that can be used independently or in combination for analysis and visualization of biological data sets, and for confirming or generating hypotheses about biological function. Several of these visualizations and algorithms are only available to Cytoscape users through the clusterMaker plugin. clusterMaker is available via the Cytoscape plugin manager.
John Scotter Morris, Leonard Apeltsin, Aaron M. Newman, Jan Baumbach, Tobias Wittkop, Gang Su, Gary D. Bader, Thomas E. Ferrin
BMC Bioinform.4
2009 Towards the integrated analysis, visualization and reconstruction of microbial gene regulatory networks
abstract
To handle changing environmental surroundings and to manage unfavorable conditions, microbial organisms have evolved complex transcriptional regulatory networks. To comprehensively analyze these gene regulatory networks, several online available databases and analysis platforms have been implemented and established. In this article, we address the typical cycle of scientific knowledge exploration and integration in the area of procaryotic transcriptional gene regulation. We briefly review five popular, publicly available systems that support (i) the integration of existing knowledge, (ii) visualization capabilities and (iii) computer analysis to predict promising wet lab targets. We exemplify the benefits of such integrated data analysis platforms by means of four application cases exemplarily performed with the corynebacterial reference database CoryneRegNet.
Jan Baumbach, Andreas Tauch, Sven Rahmann
Briefings Bioinform.1
2007 CoryneRegNet 4.0 - A reference database for corynebacterial gene regulatory networks
abstract
BACKGROUND: Detailed information on DNA-binding transcription factors (the key players in the regulation of gene expression) and on transcriptional regulatory interactions of microorganisms deduced from literature-derived knowledge, computer predictions and global DNA microarray hybridization experiments, has opened the way for the genome-wide analysis of transcriptional regulatory networks. The large-scale reconstruction of these networks allows the in silico analysis of cell behavior in response to changing environmental conditions. We previously published CoryneRegNet, an ontology-based data warehouse of corynebacterial transcription factors and regulatory networks. Initially, it was designed to provide methods for the analysis and visualization of the gene regulatory network of Corynebacterium glutamicum. RESULTS: Now we introduce CoryneRegNet release 4.0, which integrates data on the gene regulatory networks of 4 corynebacteria, 2 mycobacteria and the model organism Escherichia coli K12. As the previous versions, CoryneRegNet provides a web-based user interface to access the database content, to allow various queries, and to support the reconstruction, analysis and visualization of regulatory networks at different hierarchical levels. In this article, we present the further improved database content of CoryneRegNet along with novel analysis features. The network visualization feature GraphVis now allows the inter-species comparisons of reconstructed gene regulatory networks and the projection of gene expression levels onto that networks. Therefore, we added stimulon data directly into the database, but also provide Web Service access to the DNA microarray analysis platform EMMA. Additionally, CoryneRegNet now provides a SOAP based Web Service server, which can easily be consumed by other bioinformatics software systems. Stimulons (imported from the database, or uploaded by the user) can be analyzed in the context of known transcriptional regulatory networks to predict putative contradictions or further gene regulatory interactions. Furthermore, it integrates protein clusters by means of heuristically solving the weighted graph cluster editing problem. In addition, it provides Web Service based access to up to date gene annotation data from GenDB. CONCLUSION: The release 4.0 of CoryneRegNet is a comprehensive system for the integrated analysis of procaryotic gene regulatory networks. It is a versatile systems biology platform to support the efficient and large-scale analysis of transcriptional regulation of gene expression in microorganisms. It is publicly available at http://www.CoryneRegNet.DE.
Jan Baumbach
BMC Bioinform.1
2007 Large scale clustering of protein sequences with FORCE -A layout based heuristic for weighted cluster editing
abstract
BACKGROUND: Detecting groups of functionally related proteins from their amino acid sequence alone has been a long-standing challenge in computational genome research. Several clustering approaches, following different strategies, have been published to attack this problem. Today, new sequencing technologies provide huge amounts of sequence data that has to be efficiently clustered with constant or increased accuracy, at increased speed. RESULTS: We advocate that the model of weighted cluster editing, also known as transitive graph projection is well-suited to protein clustering. We present the FORCE heuristic that is based on transitive graph projection and clusters arbitrary sets of objects, given pairwise similarity measures. In particular, we apply FORCE to the problem of protein clustering and show that it outperforms the most popular existing clustering tools (Spectral clustering, TribeMCL, GeneRAGE, Hierarchical clustering, and Affinity Propagation). Furthermore, we show that FORCE is able to handle huge datasets by calculating clusters for all 192 187 prokaryotic protein sequences (66 organisms) obtained from the COG database. Finally, FORCE is integrated into the corynebacterial reference database CoryneRegNet. CONCLUSION: FORCE is an applicable alternative to existing clustering algorithms. Its theoretical foundation, weighted cluster editing, can outperform other clustering paradigms on protein homology clustering. FORCE is open source and implemented in Java. The software, including the source code, the clustering results for COG and CoryneRegNet, and all evaluation datasets are available at http://gi.cebitec.uni-bielefeld.de/comet/force/.
Tobias Wittkop, Jan Baumbach, Francisco Pereira Lobo, Sven Rahmann
BMC Bioinform.2
2006 Graph-based analysis and visualization of experimental results with ONDEX
abstract
MOTIVATION: Assembling the relevant information needed to interpret the output from high-throughput, genome scale, experiments such as gene expression microarrays is challenging. Analysis reveals genes that show statistically significant changes in expression levels, but more information is needed to determine their biological relevance. The challenge is to bring these genes together with biological information distributed across hundreds of databases or buried in the scientific literature (millions of articles). Software tools are needed to automate this task which at present is labor-intensive and requires considerable informatics and biological expertise. RESULTS: This article describes ONDEX and how it can be applied to the task of interpreting gene expression results. ONDEX is a database system that combines the features of semantic database integration and text mining with methods for graph-based analysis. An overview of the ONDEX system is presented, concentrating on recently developed features for graph-based analysis and visualization. A case study is used to show how ONDEX can help to identify causal relationships between stress response genes and metabolic pathways from gene expression data. ONDEX also discovered functional annotations for most of the genes that emerged as significant in the microarray experiment, but were previously of unknown function.
Jacob Köhler, Jan Baumbach, Jan Taubert, Michael Specht, Andre Skusa, Alexander Rüegg, Christopher J. Rawlings, Paul Verrier, Stephan Philippi
Bioinform.2