Tim Beißbarth

dblp:13/1537 · DBLP profile ↗
← Back
30ranked-venue papers
1as first author
8since 2021 · last 2025
0000-0001-6509-2143ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 29 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Prediction of gene expression using histone modification patterns extracted by Particle Swarm Optimization
abstract
MOTIVATION: Histone modifications play an important role in transcription regulation. Although the general importance of some histone modifications for transcription regulation has been previously established, the relevance of others and their interaction is subject to ongoing research. By training Machine Learning models to predict a gene's expression and explaining their decision making process, we can get hints on how histone modifications affect transcription. In previous studies, trained models were either hardly explainable or the models were trained solely on the abundance of histone modifications. Based on other studies, which used histone modification patterns, rather than their abundance, to identify potential regulatory elements, we hypothesize the histone modification pattern in a gene's promoter to be more predictive for gene expression. We used an optimization algorithm to extract predictive histone modification profiles. RESULTS: Our algorithm called PatternChrome achieved an average area under curve (AUC) score of 0.9029 over 56 samples for binary classification, outperforming all previous algorithms for the same task. We explained the models decisions to deduce the effect of specific features, certain histone modifications or promoter positions on transcription regulation. Although the predictive histone modification patterns were extracted for each sample separately, they can be used to predict gene expression in other samples, implying that the created patterns are largely generalizable. Interestingly, the impact of histone modifications on gene regulation appears predominantly indifferent to cellular specificity. Through explanation of the classifier's decisions, we substantiate established literature knowledge while concurrently revealing novel insights into the intricate landscape of transcriptional regulation via histone modification. AVAILABILITY AND IMPLEMENTATION: The code for the PatternChrome algorithm, the scripts for the analyses and the required data can be found at (https://gitlab.gwdg.de/MedBioinf/generegulation/patternchrome).
Niels Benjamin Paul, Jonas C. Wolber, Malte Lennart Sahrhage, Tim Beißbarth, Martin Haubrock
Bioinform.4
2025 HIDE: hierarchical cell-type deconvolution
abstract
MOTIVATION: Cell-type deconvolution is a computational approach to infer cellular distributions from bulk transcriptomics data. Several methods have been proposed, each with its own advantages and disadvantages. Reference based approaches make use of archetypic transcriptomic profiles representing individual cell types. Those reference profiles are ideally chosen such that the observed bulks can be reconstructed as a linear combination thereof. This strategy, however, ignores the fact that cellular populations arise through the process of cellular differentiation, which entails the gradual emergence of cell groups with diverse morphological and functional characteristics. RESULTS: Here, we propose Hierarchical cell-type Deconvolution (HIDE), a cell-type deconvolution approach which incorporates a cell hierarchy for improved performance and interpretability. This is achieved by a hierarchical procedure that preserves estimates of major cell populations while inferring their respective subpopulations. We show in simulation studies that this procedure produces more reliable and more consistent results than other state-of-the-art approaches. Finally, we provide an example application of HIDE to explore breast cancer specimens from TCGA. AVAILABILITY AND IMPLEMENTATION: A python implementation of HIDE is available at zenodo (doi: 10.5281/zenodo.14724906).
Dennis Völkl, Malte Mensching-Buhr, Thomas Sterr, Sarah Bolz, Andreas Schäfer 0005, Nicole Seifert, Jana Tauschke, Austin Rayford, Oddbjørn Straume, Helena U. Zacharias, Sushma Nagaraja Grellscheid, Tim Beißbarth, Michael Altenbuchinger, Franziska Görtler
Bioinform.12
2024 SpaCeNet: Spatial Cellular Networks from Omics Data
Stefan Schrod, Niklas Lück, Robert Lohmayer, Stefan Solbrig, Tina Wipfler, Katherine H. Shutta, Marouen Ben Guebila, Andreas Schäfer 0005, Tim Beißbarth, Helena U. Zacharias, Peter J. Oefner, John Quackenbush, Michael Altenbuchinger
RECOMB9
2024 Stable feature selection utilizing Graph Convolutional Neural Network and Layer-wise Relevance Propagation for biomarker discovery in breast cancer
abstract
High-throughput technologies are becoming increasingly important in discovering prognostic biomarkers and in identifying novel drug targets. With Mammaprint, Oncotype DX, and many other prognostic molecular signatures breast cancer is one of the paradigmatic examples of the utility of high-throughput data to deliver prognostic biomarkers, that can be represented in a form of a rather short gene list. Such gene lists can be obtained as a set of features (genes) that are important for the decisions of a Machine Learning (ML) method applied to high-dimensional gene expression data. Several studies have identified predictive gene lists for patient prognosis in breast cancer, but these lists are unstable and have only a few genes in common. Instability of feature selection impedes biological interpretability: genes that are relevant for cancer pathology should be members of any predictive gene list obtained for the same clinical type of patients. Stability and interpretability of selected features can be improved by including information on molecular networks in ML methods. Graph Convolutional Neural Network (GCNN) is a contemporary deep learning approach applicable to gene expression data structured by a prior knowledge molecular network. Layer-wise Relevance Propagation (LRP) and SHapley Additive exPlanations (SHAP) are methods to explain individual decisions of deep learning models. We used both GCNN+LRP and GCNN+SHAP techniques to construct feature sets by aggregating individual explanations. We suggest a methodology to systematically and quantitatively analyze the stability, the impact on the classification performance, and the interpretability of the selected feature sets. We used this methodology to compare GCNN+LRP to GCNN+SHAP and to more classical ML-based feature selection approaches. Utilizing a large breast cancer gene expression dataset we show that, while feature selection with SHAP is useful in applications where selected features have to be impactful for classification performance, among all studied methods GCNN+LRP delivers the most stable (reproducible) and interpretable gene lists.
Hryhorii Chereda, Andreas Leha, Tim Beißbarth
Artif. Intell. Medicine3
2024 Adaptive digital tissue deconvolution
abstract
MOTIVATION: The inference of cellular compositions from bulk and spatial transcriptomics data increasingly complements data analyses. Multiple computational approaches were suggested and recently, machine learning techniques were developed to systematically improve estimates. Such approaches allow to infer additional, less abundant cell types. However, they rely on training data which do not capture the full biological diversity encountered in transcriptomics analyses; data can contain cellular contributions not seen in the training data and as such, analyses can be biased or blurred. Thus, computational approaches have to deal with unknown, hidden contributions. Moreover, most methods are based on cellular archetypes which serve as a reference; e.g. a generic T-cell profile is used to infer the proportion of T-cells. It is well known that cells adapt their molecular phenotype to the environment and that pre-specified cell archetypes can distort the inference of cellular compositions. RESULTS: We propose Adaptive Digital Tissue Deconvolution (ADTD) to estimate cellular proportions of pre-selected cell types together with possibly unknown and hidden background contributions. Moreover, ADTD adapts prototypic reference profiles to the molecular environment of the cells, which further resolves cell-type specific gene regulation from bulk transcriptomics data. We verify this in simulation studies and demonstrate that ADTD improves existing approaches in estimating cellular compositions. In an application to bulk transcriptomics data from breast cancer patients, we demonstrate that ADTD provides insights into cell-type specific molecular differences between breast cancer subtypes. AVAILABILITY AND IMPLEMENTATION: A python implementation of ADTD and a tutorial are available at Gitlab and zenodo (doi:10.5281/zenodo.7548362).
Franziska Görtler, Malte Mensching-Buhr, Ørjan Skaar, Stefan Schrod, Thomas Sterr, Andreas Schäfer 0005, Tim Beißbarth, Anagha Joshi, Helena U. Zacharias, Sushma Nagaraja Grellscheid, Michael Altenbuchinger
Bioinform.7
2024 CODEX: COunterfactual Deep learning for the in silico EXploration of cancer cell line perturbations
abstract
MOTIVATION: High-throughput screens (HTS) provide a powerful tool to decipher the causal effects of chemical and genetic perturbations on cancer cell lines. Their ability to evaluate a wide spectrum of interventions, from single drugs to intricate drug combinations and CRISPR-interference, has established them as an invaluable resource for the development of novel therapeutic approaches. Nevertheless, the combinatorial complexity of potential interventions makes a comprehensive exploration intractable. Hence, prioritizing interventions for further experimental investigation becomes of utmost importance. RESULTS: We propose CODEX (COunterfactual Deep learning for the in silico EXploration of cancer cell line perturbations) as a general framework for the causal modeling of HTS data, linking perturbations to their downstream consequences. CODEX relies on a stringent causal modeling strategy based on counterfactual reasoning. As such, CODEX predicts drug-specific cellular responses, comprising cell survival and molecular alterations, and facilitates the in silico exploration of drug combinations. This is achieved for both bulk and single-cell HTS. We further show that CODEX provides a rationale to explore complex genetic modifications from CRISPR-interference in silico in single cells. AVAILABILITY AND IMPLEMENTATION: Our implementation of CODEX is publicly available at https://github.com/sschrod/CODEX. All data used in this article are publicly available.
Stefan Schrod, Helena U. Zacharias, Tim Beißbarth, Anne-Christin Hauschild, Michael Altenbuchinger
Bioinform.3
2023 Ensemble-GNN: federated ensemble learning with graph neural networks for disease module discovery and classification
abstract
SUMMARY: Federated learning enables collaboration in medicine, where data is scattered across multiple centers without the need to aggregate the data in a central cloud. While, in general, machine learning models can be applied to a wide range of data types, graph neural networks (GNNs) are particularly developed for graphs, which are very common in the biomedical domain. For instance, a patient can be represented by a protein-protein interaction (PPI) network where the nodes contain the patient-specific omics features. Here, we present our Ensemble-GNN software package, which can be used to deploy federated, ensemble-based GNNs in Python. Ensemble-GNN allows to quickly build predictive models utilizing PPI networks consisting of various node features such as gene expression and/or DNA methylation. We exemplary show the results from a public dataset of 981 patients and 8469 genes from the Cancer Genome Atlas (TCGA). AVAILABILITY AND IMPLEMENTATION: The source code is available at https://github.com/pievos101/Ensemble-GNN, and the data at Zenodo (DOI: 10.5281/zenodo.8305122).
Bastian Pfeifer, Hryhorii Chereda, Roman Martin, Anna Saranti, Sandra Clemens, Anne-Christin Hauschild, Tim Beißbarth, Andreas Holzinger, Dominik Heider
Bioinform.7
2022 BITES: balanced individual treatment effect for survival data
abstract
MOTIVATION: Estimating the effects of interventions on patient outcome is one of the key aspects of personalized medicine. Their inference is often challenged by the fact that the training data comprises only the outcome for the administered treatment, and not for alternative treatments (the so-called counterfactual outcomes). Several methods were suggested for this scenario based on observational data, i.e. data where the intervention was not applied randomly, for both continuous and binary outcome variables. However, patient outcome is often recorded in terms of time-to-event data, comprising right-censored event times if an event does not occur within the observation period. Albeit their enormous importance, time-to-event data are rarely used for treatment optimization. We suggest an approach named BITES (Balanced Individual Treatment Effect for Survival data), which combines a treatment-specific semi-parametric Cox loss with a treatment-balanced deep neural network; i.e. we regularize differences between treated and non-treated patients using Integral Probability Metrics (IPM). RESULTS: We show in simulation studies that this approach outperforms the state of the art. Furthermore, we demonstrate in an application to a cohort of breast cancer patients that hormone treatment can be optimized based on six routine parameters. We successfully validated this finding in an independent cohort. AVAILABILITY AND IMPLEMENTATION: We provide BITES as an easy-to-use python implementation including scheduled hyper-parameter optimization (https://github.com/sschrod/BITES). The data underlying this article are available in the CRAN repository at https://rdrr.io/cran/survival/man/gbsg.html and https://rdrr.io/cran/survival/man/rotterdam.html. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Stefan Schrod, Andreas Schäfer 0005, Stefan Solbrig, Robert Lohmayer, Wolfram Gronwald, Peter J. Oefner, Tim Beißbarth, Rainer Spang, Helena U. Zacharias, Michael Altenbuchinger
Bioinform.7
2015 pwOmics: an R package for pathway-based integration of time-series omics data using public database knowledge
abstract
UNLABELLED: Characterization of biological processes is progressively enabled with the increased generation of omics data on different signaling levels. Here we present a straightforward approach for the integrative analysis of data from different high-throughput technologies based on pathway and interaction models from public databases. pwOmics performs pathway-based level-specific data comparison of coupled human proteomic and genomic/transcriptomic datasets based on their log fold changes. Separate downstream and upstream analyses results on the functional levels of pathways, transcription factors and genes/transcripts are performed in the cross-platform consensus analysis. These provide a basis for the combined interpretation of regulatory effects over time. Via network reconstruction and inference methods (Steiner tree, dynamic Bayesian network inference) consensus graphical networks can be generated for further analyses and visualization. AVAILABILITY AND IMPLEMENTATION: The R package pwOmics is freely available on Bioconductor (http://www.bioconductor.org/). CONTACT: [email protected].
Astrid Wachter, Tim Beißbarth
Bioinform.2
2015 Comparative study on gene set and pathway topology-based enrichment methods
abstract
BACKGROUND: Enrichment analysis is a popular approach to identify pathways or sets of genes which are significantly enriched in the context of differentially expressed genes. The traditional gene set enrichment approach considers a pathway as a simple gene list disregarding any knowledge of gene or protein interactions. In contrast, the new group of so called pathway topology-based methods integrates the topological structure of a pathway into the analysis. METHODS: We comparatively investigated gene set and pathway topology-based enrichment approaches, considering three gene set and four topological methods. These methods were compared in two extensive simulation studies and on a benchmark of 36 real datasets, providing the same pathway input data for all methods. RESULTS: In the benchmark data analysis both types of methods showed a comparable ability to detect enriched pathways. The first simulation study was conducted with KEGG pathways, which showed considerable gene overlaps between each other. In this study with original KEGG pathways, none of the topology-based methods outperformed the gene set approach. Therefore, a second simulation study was performed on non-overlapping pathways created by unique gene IDs. Here, methods accounting for pathway topology reached higher accuracy than the gene set methods, however their sensitivity was lower. CONCLUSIONS: We conducted one of the first comprehensive comparative works on evaluating gene set against pathway topology-based enrichment methods. The topological methods showed better performance in the simulation scenarios with non-overlapping pathways, however, they were not conclusively better in the other scenarios. This suggests that simple gene set approach might be sufficient to detect an enriched pathway under realistic circumstances. Nevertheless, more extensive studies and further benchmark data are needed to systematically evaluate these methods and to assess what gain and cost pathway topology information introduces into enrichment analysis. Both types of methods for enrichment analysis require further improvements in order to deal with the problem of pathway overlaps.
Michaela Bayerlová, Klaus Jung, Frank Kramer 0001, Florian Klemm, Annalen Bleckmann, Tim Beißbarth
BMC Bioinform.6
2014 Adaption of the global test idea to proteomics data with missing values
abstract
MOTIVATION: Global test procedures are frequently used in gene expression analysis to study the relationship between a functional subset of RNA transcripts and an experimental group factor. However, these procedures have been rarely used for the analysis of high-throughput data from other sources, such as proteome expression data. The main difficulties in transferring global test procedures from genomics to proteomics data are the more complicated way of obtaining functional annotations and the handling of missing values in some types of proteomics data. RESULTS: We propose a simple mixed linear model in combination with a permutation procedure and missing values imputation to conduct global tests in proteomics experiments. This new approach is motivated by protein expression data obtained by means of 2-D gel electrophoresis within a mouse experiment of our current research. A simulation study yielded that power and testing level of the mixed model alone can be affected by missing values in the dataset. Imputation of missing values was able to correct for a bias in some simulation settings. Our new approach provides the possibility to rank Gene Ontology (GO) terms associated with protein sets. It is also helpful in the case in which a specific protein is represented by multiple spots on a 2-D gel by considering these spots also as a protein set. Analysis of our data points at correlations between the deficiency of the protein 'calreticulin' and protein sets related to biological processes in the heart muscle. AVAILABILITY AND IMPLEMENTATION: Our proposed approach is included in the R-package 'RepeatedHighDim', which already contains a global test procedure for gene expression data. The package can be retrieved from http://cran.r-project.org/. CONTACT: [email protected].
Klaus Jung, Hassan Dihazi, Asima Bibi, Gry H. Dihazi, Tim Beißbarth
Bioinform.5
2013 rBiopaxParser - an R package to parse, modify and visualize BioPAX data
abstract
MOTIVATION: Biological pathway data, stored in structured databases, is a useful source of knowledge for a wide range of bioinformatics algorithms and tools. The Biological Pathway Exchange (BioPAX) language has been established as a standard to store and annotate pathway information. However, use of these data within statistical analyses can be tedious. On the other hand, the statistical computing environment R has become the standard for bioinformatics analysis of large-scale genomics data. With this package, we hope to enable R users to work with BioPAX data and make use of the always increasing amount of biological pathway knowledge within data analysis methods. RESULTS: rBiopaxParser is a software package that provides a comprehensive set of functions for parsing, viewing and modifying BioPAX pathway data within R. These functions enable the user to access and modify specific parts of the BioPAX model. Furthermore, it allows to generate and layout regulatory graphs of controlling interactions and to visualize BioPAX pathways. AVAILABILITY: rBiopaxParser is an open-source R package and has been submitted to Bioconductor.
Frank Kramer 0001, Michaela Bayerlová, Florian Klemm, Annalen Bleckmann, Tim Beißbarth
Bioinform.5
2011 pathClass: an R-package for integration of pathway knowledge into support vector machines for biomarker discovery
abstract
UNLABELLED: Prognostic and diagnostic biomarker discovery is one of the key issues for a successful stratification of patients according to clinical risk factors. For this purpose, statistical classification methods, such as support vector machines (SVM), are frequently used tools. Different groups have recently shown that the usage of prior biological knowledge significantly improves the classification results in terms of accuracy as well as reproducibility and interpretability of gene lists. Here, we introduce pathClass, a collection of different SVM-based classification methods for improved gene selection and classfication performance. The methods contained in pathClass do not merely rely on gene expression data but also exploit the information that is carried in gene network data. AVAILABILITY: pathClass is open source and freely available as an R-Package on the CRAN repository at http://cran.r-project.org.
Marc Johannes, Holger Fröhlich, Holger Sültmann, Tim Beißbarth
Bioinform.4
2011 Comparison of global tests for functional gene sets in two-group designs and selection of potentially effect-causing genes
abstract
MOTIVATION: An important object in the analysis of high-throughput genomic data is to find an association between the expression profile of functional gene sets and the different levels of a group response. Instead of multiple testing procedures which focus on single genes, global tests are usually used to detect a group effect in an entire gene set. In a simulation study, we compare the power and computation times of four different approaches for global testing. The applicability of one of these methods to gene expression data is demonstrated for the first time. In addition, we propose an algorithm for the detection of those genes which might be responsible for a group effect. RESULTS: We could detect that the power of three of the approaches is comparable in many settings but considerable differences were detected in the computation times. Our proposed gene selection algorithm was able to detect potentially effect-causing genes in artificial sets with high power when many genes were altered with a small effect, while classical multiple testing was more powerful when few genes were altered with a large effect. AVAILABILITY: An R-package called 'RepeatedHighDim' which implements our new global test procedures is made available from http://cran.r-project.org/.
Klaus Jung, Benjamin Becker, Edgar Brunner, Tim Beißbarth
Bioinform.4
2011 Inferring signalling networks from longitudinal data using sampling based approaches in the R-package 'ddepn'
abstract
BACKGROUND: Network inference from high-throughput data has become an important means of current analysis of biological systems. For instance, in cancer research, the functional relationships of cancer related proteins, summarised into signalling networks are of central interest for the identification of pathways that influence tumour development. Cancer cell lines can be used as model systems to study the cellular response to drug treatments in a time-resolved way. Based on these kind of data, modelling approaches for the signalling relationships are needed, that allow to generate hypotheses on potential interference points in the networks. RESULTS: We present the R-package 'ddepn' that implements our recent approach on network reconstruction from longitudinal data generated after external perturbation of network components. We extend our approach by two novel methods: a Markov Chain Monte Carlo method for sampling network structures with two edge types (activation and inhibition) and an extension of a prior model that penalises deviances from a given reference network while incorporating these two types of edges. Further, as alternative prior we include a model that learns signalling networks with the scale-free property. CONCLUSIONS: The package 'ddepn' is freely available on R-Forge and CRAN http://ddepn.r-forge.r-project.org, http://cran.r-project.org. It allows to conveniently perform network inference from longitudinal high-throughput data using two different sampling based network structure search algorithms.
Christian Bender, Silvia von der Heyde, Frauke Henjes, Stefan Wiemann, Ulrike Korf, Tim Beißbarth
BMC Bioinform.6
2011 Graph based fusion of miRNA and mRNA expression data improves clinical outcome prediction in prostate cancer
abstract
BACKGROUND: One of the main goals in cancer studies including high-throughput microRNA (miRNA) and mRNA data is to find and assess prognostic signatures capable of predicting clinical outcome. Both mRNA and miRNA expression changes in cancer diseases are described to reflect clinical characteristics like staging and prognosis. Furthermore, miRNA abundance can directly affect target transcripts and translation in tumor cells. Prediction models are trained to identify either mRNA or miRNA signatures for patient stratification. With the increasing number of microarray studies collecting mRNA and miRNA from the same patient cohort there is a need for statistical methods to integrate or fuse both kinds of data into one prediction model in order to find a combined signature that improves the prediction. RESULTS: Here, we propose a new method to fuse miRNA and mRNA data into one prediction model. Since miRNAs are known regulators of mRNAs we used the correlations between them as well as the target prediction information to build a bipartite graph representing the relations between miRNAs and mRNAs. This graph was used to guide the feature selection in order to improve the prediction. The method is illustrated on a prostate cancer data set comprising 98 patient samples with miRNA and mRNA expression data. The biochemical relapse was used as clinical endpoint. It could be shown that the bipartite graph in combination with both data sets could improve prediction performance as well as the stability of the feature selection. CONCLUSIONS: Fusion of mRNA and miRNA expression data into one prediction model improves clinical outcome prediction in terms of prediction error and stable feature selection. The R source code of the proposed method is available in the supplement.
Stephan Gade, Christine Porzelius, Maria Fälth, Jan C. Brase, Daniela Wuttig, Ruprecht Kuner, Harald Binder, Holger Sültmann, Tim Beißbarth
BMC Bioinform.9
2011 Reporting FDR analogous confidence intervals for the log fold change of differentially expressed genes
abstract
BACKGROUND: Gene expression experiments are common in molecular biology, for example in order to identify genes which play a certain role in a specified biological framework. For that purpose expression levels of several thousand genes are measured simultaneously using DNA microarrays. Comparing two distinct groups of tissue samples to detect those genes which are differentially expressed one statistical test per gene is performed, and resulting p-values are adjusted to control the false discovery rate. In addition, the expression change of each gene is quantified by some effect measure, typically the log fold change. In certain cases, however, a gene with a significant p-value can have a rather small fold change while in other cases a non-significant gene can have a rather large fold change. The biological relevance of the change of gene expression can be more intuitively judged by a fold change then merely by a p-value. Therefore, confidence intervals for the log fold change which accompany the adjusted p-values are desirable. RESULTS: In a new approach, we employ an existing algorithm for adjusting confidence intervals in the case of high-dimensional data and apply it to a widely used linear model for microarray data. Furthermore, we adopt a concept of different relevance categories for effects in clinical trials to assess biological relevance of genes in microarray experiments. In a brief simulation study the properties of the adjusting algorithm are maintained when being combined with the linear model for microarray data. In two cancer data sets the adjusted confidence intervals can indicate significance of large fold changes and distinguish them from other large but non-significant fold changes. Adjusting of confidence intervals also corrects the assessment of biological relevance. CONCLUSIONS: Our new combination approach and the categorization of fold changes facilitates the selection of genes in microarray experiments and helps to interpret their biological relevance.
Klaus Jung, Tim Friede, Tim Beißbarth
BMC Bioinform.3
2011 Sequential Interim Analyses of Survival Data in DNA Microarray Experiments
abstract
BACKGROUND: Discovery of biomarkers that are correlated with therapy response and thus with survival is an important goal of medical research on severe diseases, e.g. cancer. Frequently, microarray studies are performed to identify genes of which the expression levels in pretherapeutic tissue samples are correlated to survival times of patients. Typically, such a study can take several years until the full planned sample size is available.Therefore, interim analyses are desirable, offering the possibility of stopping the study earlier, or of performing additional laboratory experiments to validate the role of the detected genes. While many methods correcting the multiple testing bias introduced by interim analyses have been proposed for studies of one single feature, there are still open questions about interim analyses of multiple features, particularly of high-dimensional microarray data, where the number of features clearly exceeds the number of samples. Therefore, we examine false discovery rates and power rates in microarray experiments performed during interim analyses of survival studies. In addition, the early stopping based on interim results of such studies is evaluated. As stop criterion we employ the achieved average power rate, i.e. the proportion of detected true positives, for which a new estimator is derived and compared to existing estimators. RESULTS: In a simulation study, pre-specified levels of the false discovery rate are maintained in each interim analysis, where reduced levels as used in classical group sequential designs of one single feature are not necessary. Average power rates increase with each interim analysis, and many studies can be stopped prior to their planned end when a certain pre-specified power rate is achieved. The new estimator for the power rate slightly deviates from the true power rate but is comparable to other estimators. CONCLUSIONS: Interim analyses of microarray experiments can provide evidence for early stopping of long-term survival studies. The developed simulation framework, which we also offer as a new R package 'SurvGenesInterim' available at http://survgenesinter.R-Forge.R-Project.org, can be used for sample size planning of the evaluated study design.
Andreas Leha, Tim Beißbarth, Klaus Jung
BMC Bioinform.2
2010 Dynamic deterministic effects propagation networks: learning signalling pathways from longitudinal protein array data
abstract
MOTIVATION: Network modelling in systems biology has become an important tool to study molecular interactions in cancer research, because understanding the interplay of proteins is necessary for developing novel drugs and therapies. De novo reconstruction of signalling pathways from data allows to unravel interactions between proteins and make qualitative statements on possible aberrations of the cellular regulatory program. We present a new method for reconstructing signalling networks from time course experiments after external perturbation and show an application of the method to data measuring abundance of phosphorylated proteins in a human breast cancer cell line, generated on reverse phase protein arrays. RESULTS: Signalling dynamics is modelled using active and passive states for each protein at each timepoint. A fixed signal propagation scheme generates a set of possible state transitions on a discrete timescale for a given network hypothesis, reducing the number of theoretically reachable states. A likelihood score is proposed, describing the probability of measurements given the states of the proteins over time. The optimal sequence of state transitions is found via a hidden Markov model and network structure search is performed using a genetic algorithm that optimizes the overall likelihood of a population of candidate networks. Our method shows increased performance compared with two different dynamical Bayesian network approaches. For our real data, we were able to find several known signalling cascades from the ERBB signalling pathway. AVAILABILITY: Dynamic deterministic effects propagation networks is implemented in the R programming language and available at http://www.dkfz.de/mga2/ddepn/.
Christian Bender, Frauke Henjes, Holger Fröhlich, Stefan Wiemann, Ulrike Korf, Tim Beißbarth
Bioinform.6
2010 QuantProReloaded: quantitative analysis of microspot immunoassays
abstract
UNLABELLED: Protein microarrays are well-established as sensitive tools for proteomics. Particularly, the microspot immunoassay (MIA) platform enables a quantitative analysis of (phospho-) proteins in complex solutions (e.g. cell lysates or blood plasma) and with low consumption of samples and reagents. Despite numerous biological and clinical applications of MIAs there is currently no user-friendly open source data analysis software available with versatile options for data analysis and data visualization. Here, we introduce the open source software QuantProReloaded that is specifically designed for the analysis of data from MIA experiments. AVAILABILITY AND IMPLEMENTATION: QuantProReloaded is written in R and Java and is open for download under the BSB license at http://code.google.com/p/quantproreloaded/.
Anika Jöcker, Johanna Sonntag, Frauke Henjes, Frank Götschel, Achim Tresch, Tim Beißbarth, Stefan Wiemann, Ulrike Korf
Bioinform.6
2010 Integration of pathway knowledge into a reweighted recursive feature elimination approach for risk stratification of cancer patients
abstract
MOTIVATION: One of the main goals of high-throughput gene-expression studies in cancer research is to identify prognostic gene signatures, which have the potential to predict the clinical outcome. It is common practice to investigate these questions using classification methods. However, standard methods merely rely on gene-expression data and assume the genes to be independent. Including pathway knowledge a priori into the classification process has recently been indicated as a promising way to increase classification accuracy as well as the interpretability and reproducibility of prognostic gene signatures. RESULTS: We propose a new method called Reweighted Recursive Feature Elimination. It is based on the hypothesis that a gene with a low fold-change should have an increased influence on the classifier if it is connected to differentially expressed genes. We used a modified version of Google's PageRank algorithm to alter the ranking criterion of the SVM-RFE algorithm. Evaluations of our method on an integrated breast cancer dataset comprising 788 samples showed an improvement of the area under the receiver operator characteristic curve as well as in the reproducibility and interpretability of selected genes. AVAILABILITY: The R code of the proposed algorithm is given in Supplementary Material.
Marc Johannes, Jan C. Brase, Holger Fröhlich, Stephan Gade, Mathias C. Gehrmann, Maria Fälth, Holger Sültmann, Tim Beißbarth
Bioinform.8
2010 RPPanalyzer: Analysis of reverse-phase protein array data
abstract
SUMMARY: RPPanalyzer is a statistical tool developed to read reverse-phase protein array data, to perform the basic data analysis and to visualize the resulting biological information. The R-package provides different functions to compare protein expression levels of different samples and to normalize the data. Implemented plotting functions permit a quality control by monitoring data distribution and signal validity. Finally, the data can be visualized in heatmaps, boxplots, time course plots and correlation plots. RPPanalyzer is a flexible tool and tolerates a huge variety of different experimental designs. AVAILABILITY: The RPPAanalyzer is open source and freely available as an R-Package on the CRAN platform http://cran.r-project.org/.
Heiko A. Mannsperger, Stephan Gade, Frauke Henjes, Tim Beißbarth, Ulrike Korf
Bioinform.4
2009 Deterministic Effects Propagation Networks for reconstructing protein signaling networks from multiple interventions
abstract
BACKGROUND: Modern gene perturbation techniques, like RNA interference (RNAi), enable us to study effects of targeted interventions in cells efficiently. In combination with mRNA or protein expression data this allows to gain insights into the behavior of complex biological systems. RESULTS: In this paper, we propose Deterministic Effects Propagation Networks (DEPNs) as a special Bayesian Network approach to reverse engineer signaling networks from a combination of protein expression and perturbation data. DEPNs allow to reconstruct protein networks based on combinatorial intervention effects, which are monitored via changes of the protein expression or activation over one or a few time points. Our implementation of DEPNs allows for latent network nodes (i.e. proteins without measurements) and has a built in mechanism to impute missing data. The robustness of our approach was tested on simulated data. We applied DEPNs to reconstruct the ERBB signaling network in de novo trastuzumab resistant human breast cancer cells, where protein expression was monitored on Reverse Phase Protein Arrays (RPPAs) after knockdown of network proteins using RNAi. CONCLUSION: DEPNs offer a robust, efficient and simple approach to infer protein signaling networks from multiple interventions. The method as well as the data have been made part of the latest version of the R package "nem" available as a supplement to this paper and via the Bioconductor repository.
Holger Fröhlich, Özgür Sahin, Dorit Arlt, Christian Bender, Tim Beißbarth
BMC Bioinform.5
2008 Analyzing gene perturbation screens with nested effects models in R and bioconductor
abstract
UNLABELLED: Nested effects models (NEMs) are a class of probabilistic models introduced to analyze the effects of gene perturbation screens visible in high-dimensional phenotypes like microarrays or cell morphology. NEMs reverse engineer upstream/downstream relations of cellular signaling cascades. NEMs take as input a set of candidate pathway genes and phenotypic profiles of perturbing these genes. NEMs return a pathway structure explaining the observed perturbation effects. Here, we describe the package nem, an open-source software to efficiently infer NEMs from data. Our software implements several search algorithms for model fitting and is applicable to a wide range of different data types and representations. The methods we present summarize the current state-of-the-art in NEMs. AVAILABILITY: Our software is written in the R language and freely avail-able via the Bioconductor project at http://www.bioconductor.org.
Holger Fröhlich, Tim Beißbarth, Achim Tresch, Dennis Kostka, Juby Jacob, Rainer Spang, Florian Markowetz
Bioinform.2
2008 Predicting pathway membership via domain signatures
abstract
MOTIVATION: Functional characterization of genes is of great importance for the understanding of complex cellular processes. Valuable information for this purpose can be obtained from pathway databases, like KEGG. However, only a small fraction of genes is annotated with pathway information up to now. In contrast, information on contained protein domains can be obtained for a significantly higher number of genes, e.g. from the InterPro database. RESULTS: We present a classification model, which for a specific gene of interest can predict the mapping to a KEGG pathway, based on its domain signature. The classifier makes explicit use of the hierarchical organization of pathways in the KEGG database. Furthermore, we take into account that a specific gene can be mapped to different pathways at the same time. The classification method produces a scoring of all possible mapping positions of the gene in the KEGG hierarchy. Evaluations of our model, which is a combination of a SVM and ranking perceptron approach, show a high prediction performance. Moreover, for signaling pathways we reveal that it is even possible to forecast accurately the membership to individual pathway components. AVAILABILITY: The R package gene2pathway is a supplement to this article.
Holger Fröhlich, Mark Fellmann, Holger Sültmann, Annemarie Poustka, Tim Beißbarth
Bioinform.5
2008 Estimating large-scale signaling networks through nested effect models with intervention effects from microarray data
abstract
MOTIVATION: Targeted interventions using RNA interference in combination with the measurement of secondary effects with DNA microarrays can be used to computationally reverse engineer features of upstream non-transcriptional signaling cascades based on the nested structure of effects. RESULTS: We extend previous work by Markowetz et al., who proposed a statistical framework to score different network hypotheses. Our extensions go in several directions: we show how prior assumptions on the network structure can be incorporated into the scoring scheme by defining appropriate prior distributions on the network structure as well as on hyperparameters. An approach called module networks is introduced to scale up the original approach, which is limited to around 5 genes, to infer large-scale networks of more than 30 genes. Instead of the data discretization step needed in the original framework, we propose the usage of a beta-uniform mixture distribution on the P-value profile, resulting from differential gene expression calculation, to quantify effects. Extensive simulations on artificial data and application of our module network approach to infer the signaling network between 13 genes in the ER-alpha pathway in human MCF-7 breast cancer cells show that our approach gives sensible results. Using a bootstrapping and a jackknife approach, this reconstruction is found to be statistically stable. AVAILABILITY: The proposed method is available within the Bioconductor R-package nem.
Holger Fröhlich, Mark Fellmann, Holger Sültmann, Annemarie Poustka, Tim Beißbarth
Bioinform.5
2008 Extending pathways based on gene lists using InterPro domain signatures
abstract
BACKGROUND: High-throughput technologies like functional screens and gene expression analysis produce extended lists of candidate genes. Gene-Set Enrichment Analysis is a commonly used and well established technique to test for the statistically significant over-representation of particular pathways. A shortcoming of this method is however, that most genes that are investigated in the experiments have very sparse functional or pathway annotation and therefore cannot be the target of such an analysis. The approach presented here aims to assign lists of genes with limited annotation to previously described functional gene collections or pathways. This works by comparing InterPro domain signatures of the candidate gene lists with domain signatures of gene sets derived from known classifications, e.g. KEGG pathways. RESULTS: In order to validate our approach, we designed a simulation study. Based on all pathways available in the KEGG database, we create test gene lists by randomly selecting pathway genes, removing these genes from the known pathways and adding variable amounts of noise in the form of genes not annotated to the pathway. We show that we can recover pathway memberships based on the simulated gene lists with high accuracy. We further demonstrate the applicability of our approach on a biological example. CONCLUSION: Results based on simulation and data analysis show that domain based pathway enrichment analysis is a very sensitive method to test for enrichment of pathways in sparsely annotated lists of genes. An R based software package domainsignatures, to routinely perform this analysis on the results of high-throughput screening, is available via Bioconductor.
Florian Hahne, Alexander Mehrle, Dorit Arlt, Annemarie Poustka, Stefan Wiemann, Tim Beißbarth
BMC Bioinform.6
2007 Large scale statistical inference of signaling pathways from RNAi and microarray data
abstract
BACKGROUND: The advent of RNA interference techniques enables the selective silencing of biologically interesting genes in an efficient way. In combination with DNA microarray technology this enables researchers to gain insights into signaling pathways by observing downstream effects of individual knock-downs on gene expression. These secondary effects can be used to computationally reverse engineer features of the upstream signaling pathway. RESULTS: In this paper we address this challenging problem by extending previous work by Markowetz et al., who proposed a statistical framework to score networks hypotheses in a Bayesian manner. Our extensions go in three directions: First, we introduce a way to omit the data discretization step needed in the original framework via a calculation based on p-values instead. Second, we show how prior assumptions on the network structure can be incorporated into the scoring scheme using regularization techniques. Third and most important, we propose methods to scale up the original approach, which is limited to around 5 genes, to large scale networks. CONCLUSION: Comparisons of these methods on artificial data are conducted. Our proposed module network is employed to infer the signaling network between 13 genes in the ER-alpha pathway in human MCF-7 breast cancer cells. Using a bootstrapping approach this reconstruction can be found with good statistical stability. The code for the module network inference method is available in the latest version of the R-package nem, which can be obtained from the Bioconductor homepage.
Holger Fröhlich, Mark Fellmann, Holger Sültmann, Annemarie Poustka, Tim Beißbarth
BMC Bioinform.5
2007 GOSim - an R-package for computation of information theoretic GO similarities between terms and gene products
abstract
BACKGROUND: With the increased availability of high throughput data, such as DNA microarray data, researchers are capable of producing large amounts of biological data. During the analysis of such data often there is the need to further explore the similarity of genes not only with respect to their expression, but also with respect to their functional annotation which can be obtained from Gene Ontology (GO). RESULTS: We present the freely available software package GOSim, which allows to calculate the functional similarity of genes based on various information theoretic similarity concepts for GO terms. GOSim extends existing tools by providing additional lately developed functional similarity measures for genes. These can e.g. be used to cluster genes according to their biological function. Vice versa, they can also be used to evaluate the homogeneity of a given grouping of genes with respect to their GO annotation. GOSim hence provides the researcher with a flexible and powerful tool to combine knowledge stored in GO with experimental data. It can be seen as complementary to other tools that, for instance, search for significantly overrepresented GO terms within a given group of genes. CONCLUSION: GOSim is implemented as a package for the statistical computing environment R and is distributed under GPL within the CRAN project.
Holger Fröhlich, Nora Speer, Annemarie Poustka, Tim Beißbarth
BMC Bioinform.4
2000 Processing and quality control of DNA array hybridization data
abstract
MOTIVATION: The technology of hybridization to DNA arrays is used to obtain the expression levels of many different genes simultaneously. It enables searching for genes that are expressed specifically under certain conditions. However, the technology produces large amounts of data demanding computational methods for their analysis. It is necessary to find ways to compare data from different experiments and to consider the quality and reproducibility of the data. RESULTS: Data analyzed in this paper have been generated by hybridization of radioactively labeled targets to DNA arrays spotted on nylon membranes. We introduce methods to compare the intensity values of several hybridization experiments. This is essential to find differentially expressed genes or to do pattern analysis. We also discuss possibilities for quality control of the acquired data. AVAILABILITY: http://www.dkfz.de/tbi CONTACT: [email protected]
Tim Beißbarth, Kurt Fellenberg, Benedikt Brors, Rosa Arribas-Prat, Judith M. Boer, Nicole C. Hauser, Marcel Scheideler, Jörg D. Hoheisel, Günther Schütz, Annemarie Poustka, Martin Vingron
Bioinform.1