Holger Fröhlich

dblp:86/3713 · DBLP profile ↗
← Back
45ranked-venue papers
16as first author
6since 2021 · last 2023
0000-0002-5328-1243ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 32 · 8 first-author · 6 since 2021Artificial intelligence and machine learning · 12 · 8 first-authorSystems, architecture and hardware · 3Theory of computation · 1
YearPublicationVenuePosition
2023 AI reveals insights into link between CD33 and cognitive impairment in Alzheimer's Disease
abstract
Modeling biological mechanisms is a key for disease understanding and drug-target identification. However, formulating quantitative models in the field of Alzheimer's Disease is challenged by a lack of detailed knowledge of relevant biochemical processes. Additionally, fitting differential equation systems usually requires time resolved data and the possibility to perform intervention experiments, which is difficult in neurological disorders. This work addresses these challenges by employing the recently published Variational Autoencoder Modular Bayesian Networks (VAMBN) method, which we here trained on combined clinical and patient level gene expression data while incorporating a disease focused knowledge graph. Our approach, called iVAMBN, resulted in a quantitative model that allowed us to simulate a down-expression of the putative drug target CD33, including potential impact on cognitive impairment and brain pathophysiology. Experimental validation demonstrated a high overlap of molecular mechanism predicted to be altered by CD33 perturbation with cell line data. Altogether, our modeling approach may help to select promising drug targets.
Tamara Raschka, Meemansa Sood, Bruce Schultz, Aybuge Altay, Christian Ebeling, Holger Fröhlich
PLoS Comput. Biol.6
2023 A Transformer-Based Model Trained on Large Scale Claims Data for Prediction of Severe COVID-19 Disease Progression
abstract
In situations like the COVID-19 pandemic, healthcare systems are under enormous pressure as they can rapidly collapse under the burden of the crisis. Machine learning (ML) based risk models could lift the burden by identifying patients with a high risk of severe disease progression. Electronic Health Records (EHRs) provide crucial sources of information to develop these models because they rely on routinely collected healthcare data. However, EHR data is challenging for training ML models because it contains irregularly timestamped diagnosis, prescription, and procedure codes. For such data, transformer-based models are promising. We extended the previously published Med-BERT model by including age, sex, medications, quantitative clinical measures, and state information. After pre-training on approximately 988 million EHRs from 3.5 million patients, we developed models to predict Acute Respiratory Manifestations (ARM) risk using the medical history of 80,211 COVID-19 patients. Compared to Random Forests, XGBoost, and RETAIN, our transformer-based models more accurately forecast the risk of developing ARM after COVID-19 infection. We used Integrated Gradients and Bayesian networks to understand the link between the essential features of our model. Finally, we evaluated adapting our model to Austrian in-patient data. Our study highlights the promise of predictive transformer-based models for precision medicine.
Manuel Lentzen, Thomas Linden, Sai Veeranki, Sumit Madan, Diether Kramer, Werner Leodolter, Holger Fröhlich
IEEE J. Biomed. Health Informatics7
2022 GenRisk: a tool for comprehensive genetic risk modeling
abstract
SUMMARY: The genetic architecture of complex traits can be influenced by both many common regulatory variants with small effect sizes and rare deleterious variants in coding regions with larger effect sizes. However, the two kinds of genetic contributions are typically analyzed independently. Here, we present GenRisk, a python package for the computation and the integration of gene scores based on the burden of rare deleterious variants and common-variants-based polygenic risk scores. The derived scores can be analyzed within GenRisk to perform association tests or to derive phenotype prediction models by testing multiple classification and regression approaches. GenRisk is compatible with VCF input file formats. AVAILABILITY AND IMPLEMENTATION: GenRisk is an open source publicly available python package that can be downloaded or installed from Github (https://github.com/AldisiRana/GenRisk). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Rana Aldisi, Emadeldin Hassanin, Sugirthan Sivalingam, Andreas Buness, Hannah Klinkhammer, Andreas Mayr 0001, Holger Fröhlich, Peter M. Krawitz, Carlo Maj
Bioinform.7
2022 Ten quick tips for biomarker discovery and validation analyses using machine learning
abstract
IntroductionAU : Pleaseconfirmthatallheadinglevelsarerepresentedcorrectly
Ramón Díaz-Uriarte, Elisa Gómez de Lope, Rosalba Giugno, Holger Fröhlich, Petr V. Nazarov, Isabel A. Nepomuceno-Chamorro, Armin Rauschenberger, Enrico Glaab
PLoS Comput. Biol.4
2022 GuiltyTargets: Prioritization of Novel Therapeutic Targets With Network Representation Learning
abstract
The majority of clinical trials fail due to low efficacy of investigated drugs, often resulting from a poor choice of target protein. Existing computational approaches aim to support target selection either via genetic evidence or by putting potential targets into the context of a disease specific network reconstruction. The purpose of this work was to investigate whether network representation learning techniques could be used to allow for a machine learning based prioritization of putative targets. We propose a novel target prioritization approach, GuiltyTargets, which relies on attributed network representation learning of a genome-wide protein-protein interaction network annotated with disease-specific differential gene expression and uses positive-unlabeled (PU) machine learning for candidate ranking. We evaluated our approach on 12 datasets from six diseases of different type (cancer, metabolic, neurodegenerative) within a 10 times repeated 5-fold stratified cross-validation and achieved AUROC values between 0.92 - 0.97, significantly outperforming previous approaches that relied on manually engineered topological features. Moreover, we showed that GuiltyTargets allows for target repositioning across related disease areas. An application of GuiltyTargets to Alzheimer's disease resulted in a number of highly ranked candidates that are currently discussed as targets in the literature. Interestingly, one (COMT) is also the target of an approved drug (Tolcapone) for Parkinson's disease, highlighting the potential for target repositioning with our method. The GuiltyTargets Python package is available on PyPI and all code used for analysis can be found under the MIT License at https://github.com/GuiltyTargets. Attributed network representation learning techniques provide an interesting approach to effectively leverage the existing knowledge about the molecular mechanisms in different diseases. In this work, the combination with positive-unlabeled learning for target prioritization demonstrated a clear superiority compared to classical feature engineering approaches. Our work highlights the potential of attributed network representation learning for target prioritization. Given the overarching relevance of networks in computational biology we believe that attributed network representation learning techniques could have a broader impact in the future.
Ozlem Muslu, Charles Tapley Hoyt, Mauricio Lacerda, Martin Hofmann-Apitius, Holger Fröhlich
IEEE ACM Trans. Comput. Biol. Bioinform.5
2021 SEEDS: data driven inference of structural model errors and unknown inputs for dynamic systems biology
abstract
SUMMARY: Dynamic models formulated as ordinary differential equations can provide information about the mechanistic and causal interactions in biological systems to guide targeted interventions and to design further experiments. Inaccurate knowledge about the structure, functional form and parameters of interactions is a major obstacle to mechanistic modeling. A further challenge is the open nature of biological systems which receive unknown inputs from their environment. The R-package SEEDS implements two recently developed algorithms to infer structural model errors and unknown inputs from output measurements. This information can facilitate efficient model recalibration as well as experimental design in the case of misfits between the initial model and data. AVAILABILITY AND IMPLEMENTATION: For the R-package seeds, see the CRAN server https://cran.r-project.org/package=seeds.
Tobias Newmiwaka, Benjamin Engelhardt, Philipp Wendland, Dominik Kahl, Holger Fröhlich, Maik Kschischo
Bioinform.5
2020 PathME: pathway based multi-modal sparse autoencoders for clustering of patient-level multi-omics data
abstract
BACKGROUND: Recent years have witnessed an increasing interest in multi-omics data, because these data allow for better understanding complex diseases such as cancer on a molecular system level. In addition, multi-omics data increase the chance to robustly identify molecular patient sub-groups and hence open the door towards a better personalized treatment of diseases. Several methods have been proposed for unsupervised clustering of multi-omics data. However, a number of challenges remain, such as the magnitude of features and the large difference in dimensionality across different omics data sources. RESULTS: We propose a multi-modal sparse denoising autoencoder framework coupled with sparse non-negative matrix factorization to robustly cluster patients based on multi-omics data. The proposed model specifically leverages pathway information to effectively reduce the dimensionality of omics data into a pathway and patient specific score profile. In consequence, our method allows us to understand, which pathway is a feature of which particular patient cluster. Moreover, recently proposed machine learning techniques allow us to disentangle the specific impact of each individual omics feature on a pathway score. We applied our method to cluster patients in several cancer datasets using gene expression, miRNA expression, DNA methylation and CNVs, demonstrating the possibility to obtain biologically plausible disease subtypes characterized by specific molecular features. Comparison against several competing methods showed a competitive clustering performance. In addition, post-hoc analysis of somatic mutations and clinical data provided supporting evidence and interpretation of the identified clusters. CONCLUSIONS: Our suggested multi-modal sparse denoising autoencoder approach allows for an effective and interpretable integration of multi-omics data on pathway level while addressing the high dimensional character of omics data. Patient specific pathway score profiles derived from our model allow for a robust identification of disease subgroups.
Amina Lemsara, Salima Ouadfel, Holger Fröhlich
BMC Bioinform.3
2017 Towards clinically more relevant dissection of patient heterogeneity via survival-based Bayesian clustering
abstract
MOTIVATION: Discovery of clinically relevant disease sub-types is of prime importance in personalized medicine. Disease sub-type identification has in the past often been explored in an unsupervised machine learning paradigm which involves clustering of patients based on available-omics data, such as gene expression. A follow-up analysis involves determining the clinical relevance of the molecular sub-types such as that reflected by comparing their disease progressions. The above methodology, however, fails to guarantee the separability of the sub-types based on their subtype-specific survival curves. RESULTS: We propose a new algorithm, Survival-based Bayesian Clustering (SBC) which simultaneously clusters heterogeneous-omics and clinical end point data (time to event) in order to discover clinically relevant disease subtypes. For this purpose we formulate a novel Hierarchical Bayesian Graphical Model which combines a Dirichlet Process Gaussian Mixture Model with an Accelerated Failure Time model. In this way we make sure that patients are grouped in the same cluster only when they show similar characteristics with respect to molecular features across data types (e.g. gene expression, mi-RNA) as well as survival times. We extensively test our model in simulation studies and apply it to cancer patient data from the Breast Cancer dataset and The Cancer Genome Atlas repository. Notably, our method is not only able to find clinically relevant sub-groups, but is also able to predict cluster membership and survival on test data in a better way than other competing methods. AVAILABILITY AND IMPLEMENTATION: Our R-code can be accessed as https://github.com/ashar799/SBC. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Ashar Ahmad, Holger Fröhlich
Bioinform.2
2017 Linking metabolic network features to phenotypes using sparse group lasso
abstract
MOTIVATION: Integration of metabolic networks with '-omics' data has been a subject of recent research in order to better understand the behaviour of such networks with respect to differences between biological and clinical phenotypes. Under the conditions of steady state of the reaction network and the non-negativity of fluxes, metabolic networks can be algebraically decomposed into a set of sub-pathways often referred to as extreme currents (ECs). Our objective is to find the statistical association of such sub-pathways with given clinical outcomes, resulting in a particular instance of a self-contained gene set analysis method. In this direction, we propose a method based on sparse group lasso (SGL) to identify phenotype associated ECs based on gene expression data. SGL selects a sparse set of feature groups and also introduces sparsity within each group. Features in our model are clusters of ECs, and feature groups are defined based on correlations among these features. RESULTS: We apply our method to metabolic networks from KEGG database and study the association of network features to prostate cancer (where the outcome is tumor and normal, respectively) as well as glioblastoma multiforme (where the outcome is survival time). In addition, simulations show the superior performance of our method compared to global test, which is an existing self-contained gene set analysis method. AVAILABILITY AND IMPLEMENTATION: R code (compatible with version 3.2.5) is available from http://www.abi.bit.uni-bonn.de/index.php?id=17. CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Satya Swarup Samal, Ovidiu Radulescu, Andreas Weber 0004, Holger Fröhlich
Bioinform.4
2017 Inferring modulators of genetic interactions with epistatic nested effects models
abstract
Maps of genetic interactions can dissect functional redundancies in cellular networks. Gene expression profiles as high-dimensional molecular readouts of combinatorial perturbations provide a detailed view of genetic interactions, but can be hard to interpret if different gene sets respond in different ways (called mixed epistasis). Here we test the hypothesis that mixed epistasis between a gene pair can be explained by the action of a third gene that modulates the interaction. We have extended the framework of Nested Effects Models (NEMs), a type of graphical model specifically tailored to analyze high-dimensional gene perturbation data, to incorporate logical functions that describe interactions between regulators on downstream genes and proteins. We benchmark our approach in the controlled setting of a simulation study and show high accuracy in inferring the correct model. In an application to data from deletion mutants of kinases and phosphatases in S. cerevisiae we show that epistatic NEMs can point to modulators of genetic interactions. Our approach is implemented in the R-package 'epiNEM' available from https://github.com/cbg-ethz/epiNEM and https://bioconductor.org/packages/epiNEM/.
Martin Pirkl, Madeline Diekmann, Marlies van der Wees, Niko Beerenwinkel, Holger Fröhlich, Florian Markowetz
PLoS Comput. Biol.5
2015 Analysis of Reaction Network Systems Using Tropical Geometry
Satya Swarup Samal, Dima Grigoriev, Holger Fröhlich, Ovidiu Radulescu
CASC3
2015 biRte: Bayesian inference of context-specific regulator activities and transcriptional networks
abstract
UNLABELLED: In the last years there has been an increasing effort to computationally model and predict the influence of regulators (transcription factors, miRNAs) on gene expression. Here we introduce biRte as a computationally attractive approach combining Bayesian inference of regulator activities with network reverse engineering. biRte integrates target gene predictions with different omics data entities (e.g. miRNA and mRNA data) into a joint probabilistic framework. The utility of our method is tested in extensive simulation studies and demonstrated with applications from prostate cancer and Escherichia coli growth control. The resulting regulatory networks generally show a good agreement with the biological literature. AVAILABILITY AND IMPLEMENTATION: biRte is available on Bioconductor (http://bioconductor.org). CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Holger Fröhlich
Bioinform.1
2015 NEMix: Single-cell Nested Effects Models for Probabilistic Pathway Stimulation
abstract
Nested effects models have been used successfully for learning subcellular networks from high-dimensional perturbation effects that result from RNA interference (RNAi) experiments. Here, we further develop the basic nested effects model using high-content single-cell imaging data from RNAi screens of cultured cells infected with human rhinovirus. RNAi screens with single-cell readouts are becoming increasingly common, and they often reveal high cell-to-cell variation. As a consequence of this cellular heterogeneity, knock-downs result in variable effects among cells and lead to weak average phenotypes on the cell population level. To address this confounding factor in network inference, we explicitly model the stimulation status of a signaling pathway in individual cells. We extend the framework of nested effects models to probabilistic combinatorial knock-downs and propose NEMix, a nested effects mixture model that accounts for unobserved pathway activation. We analyzed the identifiability of NEMix and developed a parameter inference scheme based on the Expectation Maximization algorithm. In an extensive simulation study, we show that NEMix improves learning of pathway structures over classical NEMs significantly in the presence of hidden pathway stimulation. We applied our model to single-cell imaging data from RNAi screens monitoring human rhinovirus infection, where limited infection efficiency of the assay results in uncertain pathway stimulation. Using a subset of genes with known interactions, we show that the inferred NEMix network has high accuracy and outperforms the classical nested effects model without hidden pathway activity. NEMix is implemented as part of the R/Bioconductor package 'nem' and available at www.cbg.ethz.ch/software/NEMix.
Juliane Siebourg-Polster, Daria Mudrak, Mario Emmenlauer, Pauli Rämö, Christoph Dehio, Urs F. Greber, Holger Fröhlich, Niko Beerenwinkel
PLoS Comput. Biol.7
2014 netClass: an R-package for network based, integrative biomarker signature discovery
abstract
In the past years, there has been a growing interest in methods that incorporate network information into classification algorithms for biomarker signature discovery in personalized medicine. The general hope is that this way the typical low reproducibility of signatures, together with the difficulty to link them to biological knowledge, can be addressed. Complementary to these efforts, there is an increasing interest in integrating different data entities (e.g. gene and miRNA expressions) into comprehensive models. To our knowledge, R-package netClass is the first software that addresses both, network and data integration. Besides several published approaches for network integration, it specifically contains our recently published STSVM method, which allows for additional integration of gene and miRNA expression data into one predictive classifier.
Yupeng Cun, Holger Fröhlich
Bioinform.2
2013 Learning gene network structure from time laps cell imaging in RNAi Knock downs
abstract
MOTIVATION: As RNA interference is becoming a standard method for targeted gene perturbation, computational approaches to reverse engineer parts of biological networks based on measurable effects of RNAi become increasingly relevant. The vast majority of these methods use gene expression data, but little attention has been paid so far to other data types. RESULTS: Here we present a method, which can infer gene networks from high-dimensional phenotypic perturbation effects on single cells recorded by time-lapse microscopy. We use data from the Mitocheck project to extract multiple shape, intensity and texture features at each frame. Features from different cells and movies are then aligned along the cell cycle time. Subsequently we use Dynamic Nested Effects Models (dynoNEMs) to estimate parts of the network structure between perturbed genes via a Markov Chain Monte Carlo approach. Our simulation results indicate a high reconstruction quality of this method. A reconstruction based on 22 gene knock downs yielded a network, where all edges could be explained via the biological literature. AVAILABILITY: The implementation of dynoNEMs is part of the Bioconductor R-package nem.
Henrik Failmezger, Paurush Praveen, Achim Tresch, Holger Fröhlich
Bioinform.4
2013 Unsupervised automated high throughput phenotyping of RNAi time-lapse movies
abstract
BACKGROUND: Gene perturbation experiments in combination with fluorescence time-lapse cell imaging are a powerful tool in reverse genetics. High content applications require tools for the automated processing of the large amounts of data. These tools include in general several image processing steps, the extraction of morphological descriptors, and the grouping of cells into phenotype classes according to their descriptors. This phenotyping can be applied in a supervised or an unsupervised manner. Unsupervised methods are suitable for the discovery of formerly unknown phenotypes, which are expected to occur in high-throughput RNAi time-lapse screens. RESULTS: We developed an unsupervised phenotyping approach based on Hidden Markov Models (HMMs) with multivariate Gaussian emissions for the detection of knockdown-specific phenotypes in RNAi time-lapse movies. The automated detection of abnormal cell morphologies allows us to assign a phenotypic fingerprint to each gene knockdown. By applying our method to the Mitocheck database, we show that a phenotypic fingerprint is indicative of a gene's function. CONCLUSION: Our fully unsupervised HMM-based phenotyping is able to automatically identify cell morphologies that are specific for a certain knockdown. Beyond the identification of genes whose knockdown affects cell morphology, phenotypic fingerprints can be used to find modules of functionally related genes.
Henrik Failmezger, Holger Fröhlich, Achim Tresch
BMC Bioinform.2
2013 Steiner tree methods for optimal sub-network identification: an empirical study
abstract
BACKGROUND: Analysis and interpretation of biological networks is one of the primary goals of systems biology. In this context identification of sub-networks connecting sets of seed proteins or seed genes plays a crucial role. Given that no natural node and edge weighting scheme is available retrieval of a minimum size sub-graph leads to the classical Steiner tree problem, which is known to be NP-complete. Many approximate solutions have been published and theoretically analyzed in the computer science literature, but far less is known about their practical performance in the bioinformatics field. RESULTS: Here we conducted a systematic simulation study of four different approximate and one exact algorithms on a large human protein-protein interaction network with ~14,000 nodes and ~400,000 edges. Moreover, we devised an own algorithm to retrieve a sub-graph of merged Steiner trees. The application of our algorithms was demonstrated for two breast cancer signatures and a sub-network playing a role in male pattern baldness. CONCLUSION: We found a modified version of the shortest paths based approximation algorithm by Takahashi and Matsuyama to lead to accurate solutions, while at the same time being several orders of magnitude faster than the exact approach. Our devised algorithm for merged Steiner trees, which is a further development of the Takahashi and Matsuyama algorithm, proved to be useful for small seed lists. All our implemented methods are available in the R-package SteinerNet on CRAN (http://www.r-project.org) and as a supplement to this paper.
Afshin Sadeghi, Holger Fröhlich
BMC Bioinform.2
2012 Joint Bayesian inference of condition-specific miRNA and transcription factor activities from combined gene and microRNA expression data
abstract
MOTIVATION: There have been many successful experimental and bioinformatics efforts to elucidate transcription factor (TF)-target networks in several organisms. For many organisms, these annotations are complemented by miRNA-target networks of good quality. Attempts that use these networks in combination with gene expression data to draw conclusions on TF or miRNA activity are, however, still relatively sparse. RESULTS: In this study, we propose Bayesian inference of regulation of transcriptional activity (BIRTA) as a novel approach to infer both, TF and miRNA activities, from combined miRNA and mRNA expression data in a condition specific way. That means our model explains mRNA and miRNA expression for a specific experimental condition by the activities of certain miRNAs and TFs, hence allowing for differentiating between switches from active to inactive (negative switch) and inactive to active (positive switch) forms. Extensive simulations of our model reveal its good prediction performance in comparison to other approaches. Furthermore, the utility of BIRTA is demonstrated at the example of Escherichia coli data comparing aerobic and anaerobic growth conditions, and by human expression data from pancreas and ovarian cancer. AVAILABILITY AND IMPLEMENTATION: The method is implemented in the R package birta, which is freely available for Bio-conductor (>=2.10) on http://www.bioconductor.org/packages/release/bioc/html/birta.html.
Benedikt Zacher, Khalid Abnaof, Stephan Gade, Erfan Younesi, Achim Tresch, Holger Fröhlich
Bioinform.6
2012 Prognostic gene signatures for patient stratification in breast cancer - accuracy, stability and interpretability of gene selection approaches using prior knowledge on protein-protein interactions
abstract
BACKGROUND: Stratification of patients according to their clinical prognosis is a desirable goal in cancer treatment in order to achieve a better personalized medicine. Reliable predictions on the basis of gene signatures could support medical doctors on selecting the right therapeutic strategy. However, during the last years the low reproducibility of many published gene signatures has been criticized. It has been suggested that incorporation of network or pathway information into prognostic biomarker discovery could improve prediction performance. In the meanwhile a large number of different approaches have been suggested for the same purpose. METHODS: We found that on average incorporation of pathway information or protein interaction data did not significantly enhance prediction performance, but indeed greatly interpretability of gene signatures. Some methods (specifically network-based SVMs) could greatly enhance gene selection stability, but revealed only a comparably low prediction accuracy, whereas Reweighted Recursive Feature Elimination (RRFE) and average pathway expression led to very clearly interpretable signatures. In addition, average pathway expression, together with elastic net SVMs, showed the highest prediction performance here. RESULTS: The results indicated that no single algorithm to perform best with respect to all three categories in our study. Incorporating network of prior knowledge into gene selection methods in general did not significantly improve classification accuracy, but greatly interpretability of gene signatures compared to classical algorithms.
Yupeng Cun, Holger Fröhlich
BMC Bioinform.2
2012 MC EMiNEM Maps the Interaction Landscape of the Mediator
abstract
The Mediator is a highly conserved, large multiprotein complex that is involved essentially in the regulation of eukaryotic mRNA transcription. It acts as a general transcription factor by integrating regulatory signals from gene-specific activators or repressors to the RNA Polymerase II. The internal network of interactions between Mediator subunits that conveys these signals is largely unknown. Here, we introduce MC EMiNEM, a novel method for the retrieval of functional dependencies between proteins that have pleiotropic effects on mRNA transcription. MC EMiNEM is based on Nested Effects Models (NEMs), a class of probabilistic graphical models that extends the idea of hierarchical clustering. It combines mode-hopping Monte Carlo (MC) sampling with an Expectation-Maximization (EM) algorithm for NEMs to increase sensitivity compared to existing methods. A meta-analysis of four Mediator perturbation studies in Saccharomyces cerevisiae, three of which are unpublished, provides new insight into the Mediator signaling network. In addition to the known modular organization of the Mediator subunits, MC EMiNEM reveals a hierarchical ordering of its internal information flow, which is putatively transmitted through structural changes within the complex. We identify the N-terminus of Med7 as a peripheral entity, entailing only local structural changes upon perturbation, while the C-terminus of Med7 and Med19 appear to play a central role. MC EMiNEM associates Mediator subunits to most directly affected genes, which, in conjunction with gene set enrichment analysis, allows us to construct an interaction map of Mediator subunits and transcription factors.
Theresa Niederberger, Stefanie Etzold, Michael Lidschreiber, Kerstin C. Maier, Dietmar E. Martin, Holger Fröhlich, Patrick Cramer, Achim Tresch
PLoS Comput. Biol.6
2011 Fast and efficient dynamic nested effects models
abstract
MOTIVATION: Targeted interventions in combination with the measurement of secondary effects can be used to computationally reverse engineer features of upstream non-transcriptional signaling cascades. Nested effect models (NEMs) have been introduced as a statistical approach to estimate the upstream signal flow from downstream nested subset structure of perturbation effects. The method was substantially extended later on by several authors and successfully applied to various datasets. The connection of NEMs to Bayesian Networks and factor graph models has been highlighted. RESULTS: Here, we introduce a computationally attractive extension of NEMs that enables the analysis of perturbation time series data, hence allowing to discriminate between direct and indirect signaling and to resolve feedback loops. AVAILABILITY: The implementation (R and C) is part of the Supplement to this article.
Holger Fröhlich, Paurush Praveen, Achim Tresch
Bioinform.1
2011 pathClass: an R-package for integration of pathway knowledge into support vector machines for biomarker discovery
abstract
UNLABELLED: Prognostic and diagnostic biomarker discovery is one of the key issues for a successful stratification of patients according to clinical risk factors. For this purpose, statistical classification methods, such as support vector machines (SVM), are frequently used tools. Different groups have recently shown that the usage of prior biological knowledge significantly improves the classification results in terms of accuracy as well as reproducibility and interpretability of gene lists. Here, we introduce pathClass, a collection of different SVM-based classification methods for improved gene selection and classfication performance. The methods contained in pathClass do not merely rely on gene expression data but also exploit the information that is carried in gene network data. AVAILABILITY: pathClass is open source and freely available as an R-Package on the CRAN repository at http://cran.r-project.org.
Marc Johannes, Holger Fröhlich, Holger Sültmann, Tim Beißbarth
Bioinform.2
2010 Dynamic deterministic effects propagation networks: learning signalling pathways from longitudinal protein array data
abstract
MOTIVATION: Network modelling in systems biology has become an important tool to study molecular interactions in cancer research, because understanding the interplay of proteins is necessary for developing novel drugs and therapies. De novo reconstruction of signalling pathways from data allows to unravel interactions between proteins and make qualitative statements on possible aberrations of the cellular regulatory program. We present a new method for reconstructing signalling networks from time course experiments after external perturbation and show an application of the method to data measuring abundance of phosphorylated proteins in a human breast cancer cell line, generated on reverse phase protein arrays. RESULTS: Signalling dynamics is modelled using active and passive states for each protein at each timepoint. A fixed signal propagation scheme generates a set of possible state transitions on a discrete timescale for a given network hypothesis, reducing the number of theoretically reachable states. A likelihood score is proposed, describing the probability of measurements given the states of the proteins over time. The optimal sequence of state transitions is found via a hidden Markov model and network structure search is performed using a genetic algorithm that optimizes the overall likelihood of a population of candidate networks. Our method shows increased performance compared with two different dynamical Bayesian network approaches. For our real data, we were able to find several known signalling cascades from the ERBB signalling pathway. AVAILABILITY: Dynamic deterministic effects propagation networks is implemented in the R programming language and available at http://www.dkfz.de/mga2/ddepn/.
Christian Bender, Frauke Henjes, Holger Fröhlich, Stefan Wiemann, Ulrike Korf, Tim Beißbarth
Bioinform.3
2010 Integration of pathway knowledge into a reweighted recursive feature elimination approach for risk stratification of cancer patients
abstract
MOTIVATION: One of the main goals of high-throughput gene-expression studies in cancer research is to identify prognostic gene signatures, which have the potential to predict the clinical outcome. It is common practice to investigate these questions using classification methods. However, standard methods merely rely on gene-expression data and assume the genes to be independent. Including pathway knowledge a priori into the classification process has recently been indicated as a promising way to increase classification accuracy as well as the interpretability and reproducibility of prognostic gene signatures. RESULTS: We propose a new method called Reweighted Recursive Feature Elimination. It is based on the hypothesis that a gene with a low fold-change should have an increased influence on the classifier if it is connected to differentially expressed genes. We used a modified version of Google's PageRank algorithm to alter the ranking criterion of the SVM-RFE algorithm. Evaluations of our method on an integrated breast cancer dataset comprising 788 samples showed an improvement of the area under the receiver operator characteristic curve as well as in the reproducibility and interpretability of selected genes. AVAILABILITY: The R code of the proposed algorithm is given in Supplementary Material.
Marc Johannes, Jan C. Brase, Holger Fröhlich, Stephan Gade, Mathias C. Gehrmann, Maria Fälth, Holger Sültmann, Tim Beißbarth
Bioinform.3
2009 Deterministic Effects Propagation Networks for reconstructing protein signaling networks from multiple interventions
abstract
BACKGROUND: Modern gene perturbation techniques, like RNA interference (RNAi), enable us to study effects of targeted interventions in cells efficiently. In combination with mRNA or protein expression data this allows to gain insights into the behavior of complex biological systems. RESULTS: In this paper, we propose Deterministic Effects Propagation Networks (DEPNs) as a special Bayesian Network approach to reverse engineer signaling networks from a combination of protein expression and perturbation data. DEPNs allow to reconstruct protein networks based on combinatorial intervention effects, which are monitored via changes of the protein expression or activation over one or a few time points. Our implementation of DEPNs allows for latent network nodes (i.e. proteins without measurements) and has a built in mechanism to impute missing data. The robustness of our approach was tested on simulated data. We applied DEPNs to reconstruct the ERBB signaling network in de novo trastuzumab resistant human breast cancer cells, where protein expression was monitored on Reverse Phase Protein Arrays (RPPAs) after knockdown of network proteins using RNAi. CONCLUSION: DEPNs offer a robust, efficient and simple approach to infer protein signaling networks from multiple interventions. The method as well as the data have been made part of the latest version of the R package "nem" available as a supplement to this paper and via the Bioconductor repository.
Holger Fröhlich, Özgür Sahin, Dorit Arlt, Christian Bender, Tim Beißbarth
BMC Bioinform.1
2008 Analyzing gene perturbation screens with nested effects models in R and bioconductor
abstract
UNLABELLED: Nested effects models (NEMs) are a class of probabilistic models introduced to analyze the effects of gene perturbation screens visible in high-dimensional phenotypes like microarrays or cell morphology. NEMs reverse engineer upstream/downstream relations of cellular signaling cascades. NEMs take as input a set of candidate pathway genes and phenotypic profiles of perturbing these genes. NEMs return a pathway structure explaining the observed perturbation effects. Here, we describe the package nem, an open-source software to efficiently infer NEMs from data. Our software implements several search algorithms for model fitting and is applicable to a wide range of different data types and representations. The methods we present summarize the current state-of-the-art in NEMs. AVAILABILITY: Our software is written in the R language and freely avail-able via the Bioconductor project at http://www.bioconductor.org.
Holger Fröhlich, Tim Beißbarth, Achim Tresch, Dennis Kostka, Juby Jacob, Rainer Spang, Florian Markowetz
Bioinform.1
2008 Predicting pathway membership via domain signatures
abstract
MOTIVATION: Functional characterization of genes is of great importance for the understanding of complex cellular processes. Valuable information for this purpose can be obtained from pathway databases, like KEGG. However, only a small fraction of genes is annotated with pathway information up to now. In contrast, information on contained protein domains can be obtained for a significantly higher number of genes, e.g. from the InterPro database. RESULTS: We present a classification model, which for a specific gene of interest can predict the mapping to a KEGG pathway, based on its domain signature. The classifier makes explicit use of the hierarchical organization of pathways in the KEGG database. Furthermore, we take into account that a specific gene can be mapped to different pathways at the same time. The classification method produces a scoring of all possible mapping positions of the gene in the KEGG hierarchy. Evaluations of our model, which is a combination of a SVM and ranking perceptron approach, show a high prediction performance. Moreover, for signaling pathways we reveal that it is even possible to forecast accurately the membership to individual pathway components. AVAILABILITY: The R package gene2pathway is a supplement to this article.
Holger Fröhlich, Mark Fellmann, Holger Sültmann, Annemarie Poustka, Tim Beißbarth
Bioinform.1
2008 Estimating large-scale signaling networks through nested effect models with intervention effects from microarray data
abstract
MOTIVATION: Targeted interventions using RNA interference in combination with the measurement of secondary effects with DNA microarrays can be used to computationally reverse engineer features of upstream non-transcriptional signaling cascades based on the nested structure of effects. RESULTS: We extend previous work by Markowetz et al., who proposed a statistical framework to score different network hypotheses. Our extensions go in several directions: we show how prior assumptions on the network structure can be incorporated into the scoring scheme by defining appropriate prior distributions on the network structure as well as on hyperparameters. An approach called module networks is introduced to scale up the original approach, which is limited to around 5 genes, to infer large-scale networks of more than 30 genes. Instead of the data discretization step needed in the original framework, we propose the usage of a beta-uniform mixture distribution on the P-value profile, resulting from differential gene expression calculation, to quantify effects. Extensive simulations on artificial data and application of our module network approach to infer the signaling network between 13 genes in the ER-alpha pathway in human MCF-7 breast cancer cells show that our approach gives sensible results. Using a bootstrapping and a jackknife approach, this reconstruction is found to be statistically stable. AVAILABILITY: The proposed method is available within the Bioconductor R-package nem.
Holger Fröhlich, Mark Fellmann, Holger Sültmann, Annemarie Poustka, Tim Beißbarth
Bioinform.1
2008 Automated classification of the behavior of rats in the forced swimming test with support vector machines
Holger Fröhlich, Andreas Hoenselaar, Jonas Eichner, Holger Rosenbrock, Gerald Birk, Andreas Zell
Neural Networks1
2007 Inferring Gene Regulatory Networks by Machine Learning Methods
Jochen Supper, Holger Fröhlich, Christian Spieth, Andreas Dräger, Andreas Zell
APBC2
2007 Gene Regulatory Network Inference via Regression Based Topological Refinement
Jochen Supper, Holger Fröhlich, Andreas Zell
APBC2
2007 Large scale statistical inference of signaling pathways from RNAi and microarray data
abstract
BACKGROUND: The advent of RNA interference techniques enables the selective silencing of biologically interesting genes in an efficient way. In combination with DNA microarray technology this enables researchers to gain insights into signaling pathways by observing downstream effects of individual knock-downs on gene expression. These secondary effects can be used to computationally reverse engineer features of the upstream signaling pathway. RESULTS: In this paper we address this challenging problem by extending previous work by Markowetz et al., who proposed a statistical framework to score networks hypotheses in a Bayesian manner. Our extensions go in three directions: First, we introduce a way to omit the data discretization step needed in the original framework via a calculation based on p-values instead. Second, we show how prior assumptions on the network structure can be incorporated into the scoring scheme using regularization techniques. Third and most important, we propose methods to scale up the original approach, which is limited to around 5 genes, to large scale networks. CONCLUSION: Comparisons of these methods on artificial data are conducted. Our proposed module network is employed to infer the signaling network between 13 genes in the ER-alpha pathway in human MCF-7 breast cancer cells. Using a bootstrapping approach this reconstruction can be found with good statistical stability. The code for the module network inference method is available in the latest version of the R-package nem, which can be obtained from the Bioconductor homepage.
Holger Fröhlich, Mark Fellmann, Holger Sültmann, Annemarie Poustka, Tim Beißbarth
BMC Bioinform.1
2007 GOSim - an R-package for computation of information theoretic GO similarities between terms and gene products
abstract
BACKGROUND: With the increased availability of high throughput data, such as DNA microarray data, researchers are capable of producing large amounts of biological data. During the analysis of such data often there is the need to further explore the similarity of genes not only with respect to their expression, but also with respect to their functional annotation which can be obtained from Gene Ontology (GO). RESULTS: We present the freely available software package GOSim, which allows to calculate the functional similarity of genes based on various information theoretic similarity concepts for GO terms. GOSim extends existing tools by providing additional lately developed functional similarity measures for genes. These can e.g. be used to cluster genes according to their biological function. Vice versa, they can also be used to evaluate the homogeneity of a given grouping of genes with respect to their GO annotation. GOSim hence provides the researcher with a flexible and powerful tool to combine knowledge stored in GO with experimental data. It can be seen as complementary to other tools that, for instance, search for significantly overrepresented GO terms within a given group of genes. CONCLUSION: GOSim is implemented as a package for the statistical computing environment R and is distributed under GPL within the CRAN project.
Holger Fröhlich, Nora Speer, Annemarie Poustka, Tim Beißbarth
BMC Bioinform.1
2006 Kernel Based Functional Gene Grouping
abstract
During the last years, high throughput experiments have become very popular. During the analysis of such data the need for a functional grouping of genes arises. In this paper, we propose grouping genes according to their biological function by means of kernel functions, which are similarity measures having special mathematical properties and play a crucial role e.g. in SVM classification. Thereby our kernel functions rely on functional information on the genes provided by Gene Ontology annotation. We investigate and compare several provably symmetric, positive semidefinite kernel functions in combination with spectral clustering, dual k-means and average linkage and demonstrate that our approach leads to good clustering results.
Holger Fröhlich, Nora Speer, Christian Spieth, Andreas Zell
IJCNN1
2006 Vibration-based Terrain Classification Using Support Vector Machines
abstract
In outdoor environments, there is a variety of different types of ground surfaces. If some of them are slippery or bumpy, for example, the ground surface itself is a possible hazard for an autonomous mobile vehicle traversing the surface. Therefore, it is beneficial if the vehicle is able to estimate, which terrain it is currently traversing. Using this estimation, the vehicle can adapt its driving style to the terrain. In this paper, we present a method for terrain classification based on vibration induced in the vehicle's body. An accelerometer mounted on the vehicle measures the vibration perpendicular to the ground surface. We experimentally compare representations of the data based on the fast Fourier transform (FFT) and on the power spectral density (PSD). Additionally, we suggest a simpler and more compact representation based on features calculated from the raw data vectors and a combination of this representation with the PSD. We train and classify the data with a support vector machine (SVM). Experiments on a large real-world dataset containing seven different terrain types evaluate our approach
Christian Weiss, Holger Fröhlich, Andreas Zell
IROS2
2005 Functional Distances for Genes Based on GO Feature Maps and their Application to Clustering
Nora Speer, Holger Fröhlich, Christian Spieth, Andreas Zell
CIBCB2
2005 Optimal assignment kernels for attributed molecular graphs
abstract
We propose a new kernel function for attributed molecular graphs, which is based on the idea of computing an optimal assignment from the atoms of one molecule to those of another one, including information on neighborhood, membership to a certain structural element and other characteristics for each atom. As a byproduct this leads to a new class of kernel functions. We demonstrate how the necessary computations can be carried out efficiently. Compared to marginalized graph kernels our method in some cases leads to a significant reduction of the prediction error. Further improvement can be gained, if expert knowledge is combined with our method. We also investigate a reduced graph representation of molecules by collapsing certain structural elements, like e.g. rings, into a single node of the molecular graph.
Holger Fröhlich, Jörg K. Wegner, Florian Sieker, Andreas Zell
ICML1
2005 Which features trigger action potentials in cortical neurons in vivo?
abstract
We study the initiation of action potentials (APs) in in vivo recordings of cortical neurons from cat visual cortex. It was shown that cortical neurons are not simple threshold devices, emitting an AP each time a fixed voltage threshold is reached, but that the emission of an AP partly depends on the rate of change of the membrane potential preceding an AP. In this paper we investigate systematically which features of the membrane potential lead to an AP by means of machine learning methods. We use support vector machines (SVMs) to discriminate between trajectories of the membrane potential which lead to an AP within the next ms and trajectories which do not lead to the initiation of an AP. For every point in a trajectory of the membrane potential (MP) we compute a set of 11 features and use a forward selection algorithm to find out the relevant features for the occurrence of an AP. Based on the results we construct a reduced prediction model. This model suggests that AP occurrences can be predicted best by a combination of the 1st temporal derivative of the MP at distance to the AP maximum, the MP itself and the mean MP over a longer range.
Holger Fröhlich, Björn Naundorf, Maxim Volgushev, Fred Wolf 0002
IJCNN1
2005 Assignment kernels for chemical compounds
abstract
During the last years kernel methods like the support vector machine (SVM) have gained a growing interest in machine learning. One of the strengths of this approach is the ability to deal easily with arbitrarily structured data by means of the kernel function. In this paper we propose a kernel for chemical compounds which is based on the idea of computing optimal assignments between atoms of two different molecules including information about their neighborhood. As a byproduct this leads to a new class of kernel functions. We demonstrate how the necessary computations can be carried out efficiently. We compare our method against the marginalized graph kernels by Kashima et al. and show its good performance on classifying toxicological and human intestinal absorption data.
Holger Fröhlich, Jörg K. Wegner, Andreas Zell
IJCNN1
2005 Efficient parameter selection for support vector machines in classification and regression via model-based global optimization
abstract
Support vector machines (SVMs) have become one of the most popular methods in machine learning during the last years. A special strength is the use of a kernel function to introduce nonlinearity and to deal with arbitrarily structured data. Usually the kernel function depends on certain parameters, which, together with other parameters of the SVM, have to be tuned to achieve good results. However, finding good parameters can become a real computational burden as the number of parameters and the size of the dataset increases. In this paper we propose an algorithm to deal with the model selection problem, which is based on the idea of learning an online Gaussian process model of the error surface in parameter space and sampling systematically at points for which the so called expected improvement is highest. Our experiments show that on this way we can find good parameters very efficiently.
Holger Fröhlich, Andreas Zell
IJCNN1
2005 Functional grouping of genes using spectral clustering and Gene Ontology
abstract
With the invention of high throughput methods, researchers are capable of producing large amounts of biological data. During the analysis of such data the need for a functional grouping of genes arises. In this paper, we propose a new method based on spectral clustering for the partitioning of genes according to their biological function. The functional information is based on Gene Ontology annotation, a mechanism to capture functional knowledge in a shareable and computer processable form. Our functional cluster method promises to automates, speed up and therefore improve biological data analysis.
Nora Speer, Holger Fröhlich, Christian Spieth, Andreas Zell
IJCNN2
2004 Gas Source Declaration with a Mobile Robot
abstract
As a sub-task of the general gas source localisation problem, gas source declaration is the process of determining the certainty that a source is in the immediate vicinity. Due to the turbulent character of gas transport in a natural indoor environment, it is not sufficient to search for instantaneous concentration maxima, in order to solve this task. Therefore, this paper introduces a method to classify whether an object is a gas source or not from a series of concentration measurements, recorded while the robot performs a rotation manoeuvre in front of a possible source. For three different gas source positions, a total of 288 declaration experiments were carried out at different robot-to-source distances. Based on these readings, two machine learning techniques (ANN, SVM) were evaluated in terms of their classification performance. With learning parameters that were optimised by grid search, a maximal hit rate of approximately 87.5% could be obtained using a support vector machine.
Achim J. Lilienthal, Andreas Zell, Holger Ulmer, Holger Fröhlich, Andreas Stützle, Felix Werner
ICRA4
2004 Feature subset selection for support vector machines by incremental regularized risk minimization
abstract
In This work we present a novel feature selection algorithm for SVMs which works by decreasing the regularized risk in an iterative manner by using a combination of a backward elimination procedure together with an exchange algorithm. It is applicable to linear as well as to nonlinear problems. We test this new algorithm on toy and real life data sets and show its good performance in comparison to state-of-the-art feature selection methods.
Holger Fröhlich, Andreas Zell
IJCNN1
2004 Learning to detect proximity to a gas source with a mobile robot
abstract
As a sub-task of the general gas source localisation problem, gas source declaration is the process of determining the certainty that a source is in the immediate vicinity. Due to the turbulent character of gas transport in a natural indoor environment, it is not sufficient to search for instantaneous concentration maxima, in order to solve this task. Therefore, this paper introduces a method to classify whether an object is a gas source from a series of concentration measurements, recorded while the robot performs a rotation manoeuvre in front of a possible source. For three different gas source positions, a total of 1056 declaration experiments were carried out at different robot-to-source distances. Based on these readings, support vector machines (SVM) with optimised learning parameters were trained and the cross-validation classification performance was evaluated. The results demonstrate the feasibility of the approach to detect proximity to a gas source using only gas sensors. The paper also presents an analysis of the classification rate depending on the desired declaration accuracy, and a comparison with the classification rate that can be achieved by selecting an optimal threshold value regarding the mean sensor signal.
Achim J. Lilienthal, Holger Ulmer, Holger Fröhlich, Felix Werner, Andreas Zell
IROS3
2003 Feature Selection for Support Vector Machines by Means of Genetic Algorithms
abstract
The problem of feature selection is a difficult combinatorial task in machine learning and of high practical relevance, e.g. in bioinformatics. genetic algorithms (GAs) offer a natural way to solve this problem. In this paper, we present a special genetic algorithm, which especially takes into account the existing bounds on the generalization error for support vector machines (SVMs). This new approach is compared to the traditional method of performing cross-validation and to other existing algorithms for feature selection.
Holger Fröhlich, Olivier Chapelle, Bernhard Schölkopf
ICTAI1