Rainer Spang

dblp:95/4799 · DBLP profile ↗
← Back
37ranked-venue papers
3as first author
7since 2021 · last 2026
0000-0002-1326-4297ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 37 · 3 first-author · 7 since 2021
YearPublicationVenuePosition
2026 Quantifying uncertainty of predictions from cancer progression models
abstract
MOTIVATION: Cancer progresses through the accumulation of genomic events. Cancer progression models such as Mutual Hazard Networks (MHNs) describe this dynamic, enabling prediction of temporal event positions and patient-specific risks of acquiring mutations. However, current MHN analyses rely on single most likely models and do not quantify the uncertainty inherent to parameter estimation. Assessing forecast stability is essential before using them to anticipate treatment-relevant mutations, adapt targeted therapies, or prioritize monitoring of patients at elevated progression risk. RESULTS: We address a key prerequisite for the responsible clinical use of cancer progression models by making MHN-derived predictions uncertainty-aware. We present a Bayesian framework for MHN that uses Markov Chain Monte Carlo to sample from the posterior distributions of model parameters and derived predictions. For practical use we implemented the Random-Walk Metropolis, Metropolis-Adjusted Langevin Algorithm (MALA), and simplified manifold MALA samplers as part of the existing mhn Python package. Only MALA and smMALA were successful in sampling from MHN posteriors, with MALA performing best. While most MHN parameters and predictions showed low posterior variance, a small subset displayed greater variability across the posterior distribution. This differentiation cannot be obtained from a single most likely model, emphasizing the need for uncertainty quantification, especially in clinical contexts. As an illustrative example, posterior sampling identified a subgroup of STK11$-$, KRAS$+$ lung adenocarcinoma patients with a high predicted short-term risk-with low variance across posterior samples-to develop an STK11 mutation. This subgroup exhibited poorer survival under immunotherapy, resembling patterns observed in STK11+ patients. AVAILABILITY AND IMPLEMENTATION: Our implementation is part of version 1.2.0 of the mhn package (https://github.com/spang-lab/LearnMHN). All analyses including the code to produce all figures in this article can be found under https://github.com/huy29433/MCMC-sampling-for-MHN (https://doi.org/10.5281/zenodo.21160219).
Y. Linda Hu, Simon Pfahler, Andreas Lösch, Stefan Vocht, Stefan Hansch, Kevin Rupp, Niko Beerenwinkel, Tilo Wettig, Rudolf Schill, Rainer Spang
Bioinform.10
2025 Harp: data harmonization for computational tissue deconvolution across diverse transcriptomics platforms
abstract
MOTIVATION: The cellular composition of a solid tissue can be assessed either through the physical dissociation of the tissue followed by single-cell analysis techniques or by computational deconvolution of bulk gene expression profiles. However, both approaches are prone to significant biases. Tissue dissociation often results in disproportionate cell loss, while deconvolution is hindered by biological and technological inconsistencies between the datasets it relies on. RESULTS: Using calibration datasets that include both experimentally measured and deconvolution-based cell compositions, we present a new method, Harp, which reconciles these approaches to produce more reliable deconvolution results in applications where only gene expression data is available. Both on simulated and real data, harmonizing cell reference profiles proved advantageous over competing state-of-the-art deconvolution tools, overcoming technological and biological batch effects. AVAILABILITY AND IMPLEMENTATION: R package available at https://github.com/spang-lab/harp (archived as 10.5281/zenodo.16851930). Code and data for reproducing the results of this paper are available at https://github.com/spang-lab/harplication (archived as 10.5281/zenodo.16851705) and https://doi.org/10.5281/zenodo.15650057, respectively.
Zahra Nozari, Paul Hüttl, Jakob Simeth, Marian Schön, James A. Hutchinson, Rainer Spang
Bioinform.6
2024 Overcoming Observation Bias for Cancer Progression Modeling
Rudolf Schill, Maren Klever, Andreas Lösch, Y. Linda Hu, Stefan Vocht, Kevin Rupp, Lars Grasedyck, Rainer Spang, Niko Beerenwinkel
RECOMB8
2024 Modeling metastatic progression from cross-sectional cancer genomics data
abstract
MOTIVATION: Metastasis formation is a hallmark of cancer lethality. Yet, metastases are generally unobservable during their early stages of dissemination and spread to distant organs. Genomic datasets of matched primary tumors and metastases may offer insights into the underpinnings and the dynamics of metastasis formation. RESULTS: We present metMHN, a cancer progression model designed to deduce the joint progression of primary tumors and metastases using cross-sectional cancer genomics data. The model elucidates the statistical dependencies among genomic events, the formation of metastasis, and the clinical emergence of both primary tumors and their metastatic counterparts. metMHN enables the chronological reconstruction of mutational sequences and facilitates estimation of the timing of metastatic seeding. In a study of nearly 5000 lung adenocarcinomas, metMHN pinpointed TP53 and EGFR as mediators of metastasis formation. Furthermore, the study revealed that post-seeding adaptation is predominantly influenced by frequent copy number alterations. AVAILABILITY AND IMPLEMENTATION: All datasets and code are available on GitHub at https://github.com/cbg-ethz/metMHN.
Kevin Rupp, Andreas Lösch, Y. Linda Hu, Chenxi Nie, Rudolf Schill, Maren Klever, Simon Pfahler, Lars Grasedyck, Tilo Wettig, Niko Beerenwinkel, Rainer Spang
Bioinform.11
2024 Virtual tissue expression analysis
abstract
MOTIVATION: Bulk RNA expression data are widely accessible, whereas single-cell data are relatively scarce in comparison. However, single-cell data offer profound insights into the cellular composition of tissues and cell type-specific gene regulation, both of which remain hidden in bulk expression analysis. RESULTS: Here, we present tissueResolver, an algorithm designed to extract single-cell information from bulk data, enabling us to attribute expression changes to individual cell types. When validated on simulated data tissueResolver outperforms competing methods. Additionally, our study demonstrates that tissueResolver reveals cell type-specific regulatory distinctions between the activated B-cell-like (ABC) and germinal center B-cell-like (GCB) subtypes of diffuse large B-cell lymphomas (DLBCL). AVAILABILITY AND IMPLEMENTATION: R package available at https://github.com/spang-lab/tissueResolver (archived as 10.5281/zenodo.14160846).Code for reproducing the results of this article is available at https://github.com/spang-lab/tissueResolver-docs archived as swh:1:dir:faea2d4f0ded30de774b28e028299ddbdd0c4f89).
Jakob Simeth, Paul Hüttl, Marian Schön, Zahra Nozari, Michael Huttner, Michael Altenbuchinger, Rainer Spang
Bioinform.8
2023 Anomaly detection in mixed high-dimensional molecular data
abstract
MOTIVATION: Mixed molecular data combines continuous and categorical features of the same samples, such as OMICS profiles with genotypes, diagnoses, or patient sex. Like all high-dimensional molecular data, it is prone to incorrect values that can stem from various sources for example the technical limitations of the measurement devices, errors in the sample preparation, or contamination. Most anomaly detection algorithms identify complete samples as outliers or anomalies. However, in most cases, not all measurements of those samples are erroneous but only a few one-dimensional features within the samples are incorrect. These one-dimensional data errors are continuous measurements that are either located outside or inside the normal ranges of their features but in both cases show atypical values given all other continuous and categorical features in the sample. Additionally, categorical anomalies can occur for example when the genotype or diagnosis was submitted wrongly. RESULTS: We introduce ADMIRE (Anomaly Detection using MIxed gRaphical modEls), a novel approach for the detection and correction of anomalies in mixed high-dimensional data. Hereby, we focus on the detection of single (one-dimensional) data errors in the categorical and continuous features of a sample. For that the joint distribution of continuous and categorical features is learned by mixed graphical models, anomalies are detected by the difference between measured and model-based estimations and are corrected using imputation. We evaluated ADMIRE in simulation and by screening for anomalies in one of our own metabolic datasets. In simulation experiments, ADMIRE outperformed the state-of-the-art methods of Local Outlier Factor, stray, and Isolation Forest. AVAILABILITY AND IMPLEMENTATION: All data and code is available at https://github.com/spang-lab/adadmire. ADMIRE is implemented in a Python package called adadmire which can be found at https://pypi.org/project/adadmire.
Lena Buck, Maren Feist, Philipp Schwarzfischer, Dieter Kube, Peter J. Oefner, Helena U. Zacharias, Michael Altenbuchinger, Katja Dettmer, Wolfram Gronwald, Rainer Spang
Bioinform.11
2022 BITES: balanced individual treatment effect for survival data
abstract
MOTIVATION: Estimating the effects of interventions on patient outcome is one of the key aspects of personalized medicine. Their inference is often challenged by the fact that the training data comprises only the outcome for the administered treatment, and not for alternative treatments (the so-called counterfactual outcomes). Several methods were suggested for this scenario based on observational data, i.e. data where the intervention was not applied randomly, for both continuous and binary outcome variables. However, patient outcome is often recorded in terms of time-to-event data, comprising right-censored event times if an event does not occur within the observation period. Albeit their enormous importance, time-to-event data are rarely used for treatment optimization. We suggest an approach named BITES (Balanced Individual Treatment Effect for Survival data), which combines a treatment-specific semi-parametric Cox loss with a treatment-balanced deep neural network; i.e. we regularize differences between treated and non-treated patients using Integral Probability Metrics (IPM). RESULTS: We show in simulation studies that this approach outperforms the state of the art. Furthermore, we demonstrate in an application to a cohort of breast cancer patients that hormone treatment can be optimized based on six routine parameters. We successfully validated this finding in an independent cohort. AVAILABILITY AND IMPLEMENTATION: We provide BITES as an easy-to-use python implementation including scheduled hyper-parameter optimization (https://github.com/sschrod/BITES). The data underlying this article are available in the CRAN repository at https://rdrr.io/cran/survival/man/gbsg.html and https://rdrr.io/cran/survival/man/rotterdam.html. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Stefan Schrod, Andreas Schäfer 0005, Stefan Solbrig, Robert Lohmayer, Wolfram Gronwald, Peter J. Oefner, Tim Beißbarth, Rainer Spang, Helena U. Zacharias, Michael Altenbuchinger
Bioinform.8
2020 Modelling cancer progression using Mutual Hazard Networks
abstract
MOTIVATION: Cancer progresses by accumulating genomic events, such as mutations and copy number alterations, whose chronological order is key to understanding the disease but difficult to observe. Instead, cancer progression models use co-occurrence patterns in cross-sectional data to infer epistatic interactions between events and thereby uncover their most likely order of occurrence. State-of-the-art progression models, however, are limited by mathematical tractability and only allow events to interact in directed acyclic graphs, to promote but not inhibit subsequent events, or to be mutually exclusive in distinct groups that cannot overlap. RESULTS: Here we propose Mutual Hazard Networks (MHN), a new Machine Learning algorithm to infer cyclic progression models from cross-sectional data. MHN model events by their spontaneous rate of fixation and by multiplicative effects they exert on the rates of successive events. MHN compared favourably to acyclic models in cross-validated model fit on four datasets tested. In application to the glioblastoma dataset from The Cancer Genome Atlas, MHN proposed a novel interaction in line with consecutive biopsies: IDH1 mutations are early events that promote subsequent fixation of TP53 mutations. AVAILABILITY AND IMPLEMENTATION: Implementation and data are available at https://github.com/RudiSchill/MHN. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Rudolf Schill, Stefan Solbrig, Tilo Wettig, Rainer Spang
Bioinform.4
2018 Loss-Function Learning for Digital Tissue Deconvolution
Franziska Görtler, Stefan Solbrig, Tilo Wettig, Peter J. Oefner, Rainer Spang, Michael Altenbuchinger
RECOMB5
2017 Reference point insensitive molecular data analysis
abstract
MOTIVATION: In biomedicine, every molecular measurement is relative to a reference point, like a fixed aliquot of RNA extracted from a tissue, a defined number of blood cells, or a defined volume of biofluid. Reference points are often chosen for practical reasons. For example, we might want to assess the metabolome of a diseased organ but can only measure metabolites in blood or urine. In this case, the observable data only indirectly reflects the disease state. The statistical implications of these discrepancies in reference points have not yet been discussed. RESULTS: Here, we show that reference point discrepancies compromise the performance of regression models like the LASSO. As an alternative, we suggest zero-sum regression for a reference point insensitive analysis. We show that zero-sum regression is superior to the LASSO in case of a poor choice of reference point both in simulations and in an application that integrates intestinal microbiome analysis with metabolomics. Moreover, we describe a novel coordinate descent based algorithm to fit zero-sum elastic nets. AVAILABILITY AND IMPLEMENTATION: The R-package "zeroSum" can be downloaded at https://github.com/rehbergT/zeroSum Moreover, we provide all R-scripts and data used to produce the results of this manuscript as Supplementary Material CONTACT: [email protected], [email protected] and [email protected] information: Supplementary material is available at Bioinformatics online.
Michael Altenbuchinger, Thorsten Rehberg, Helena U. Zacharias, Frank Stämmler, K. Dettmer, D. Weber, Andreas Hiergeist, A. Gessner, E. Holler, Peter J. Oefner, Rainer Spang
Bioinform.11
2017 Molecular signatures that can be transferred across different omics platforms
abstract
MOTIVATION: Molecular signatures for treatment recommendations are well researched. Still it is challenging to apply them to data generated by different protocols or technical platforms. RESULTS: We analyzed paired data for the same tumors (Burkitt lymphoma, diffuse large B-cell lymphoma) and features that had been generated by different experimental protocols and analytical platforms including the nanoString nCounter and Affymetrix Gene Chip transcriptomics as well as the SWATH and SRM proteomics platforms. A statistical model that assumes independent sample and feature effects accounted for 69-94% of technical variability. We analyzed how variability is propagated through linear signatures possibly affecting predictions and treatment recommendations. Linear signatures with feature weights adding to zero were substantially more robust than unbalanced signatures. They yielded consistent predictions across data from different platforms, both for transcriptomics and proteomics data. Similarly stable were their predictions across data from fresh frozen and matching formalin-fixed paraffin-embedded human tumor tissue. AVAILABILITY AND IMPLEMENTATION: The R-package 'zeroSum' can be downloaded at https://github.com/rehbergT/zeroSum . Complete data and R codes necessary to reproduce all our results can be received from the authors upon request. CONTACT: [email protected].
Michael Altenbuchinger, Philipp Schwarzfischer, Thorsten Rehberg, Jörg Reinders, Christian W. Kohler, Wolfram Gronwald, Julia Richter, Monika Szczepanowski, Neus Masqué-Soler, Wolfram Klapper, Peter J. Oefner, Rainer Spang
Bioinform.12
2017 Molecular signatures that can be transferred across different omics platforms
abstract
Bioinformatics (2017) 33 (14): i333-i340. The publisher wishes to inform readers that figure 5 was incorrect as published due to a production error re-positioning the letters (c), (d), (e) and (f), in the labelling of the figure, making it out of sync with the caption. The figure has now been corrected online.
Michael Altenbuchinger, Philipp Schwarzfischer, Thorsten Rehberg, Jörg Reinders, Christian W. Kohler, Wolfram Gronwald, Julia Richter, Monika Szczepanowski, Neus Masqué-Soler, Wolfram Klapper, Peter J. Oefner, Rainer Spang
Bioinform.12
2016 Analyzing synergistic and non-synergistic interactions in signalling pathways using Boolean Nested Effect Models
abstract
MOTIVATION: Understanding the structure and interplay of cellular signalling pathways is one of the great challenges in molecular biology. Boolean Networks can infer signalling networks from observations of protein activation. In situations where it is difficult to assess protein activation directly, Nested Effect Models are an alternative. They derive the network structure indirectly from downstream effects of pathway perturbations. To date, Nested Effect Models cannot resolve signalling details like the formation of signalling complexes or the activation of proteins by multiple alternative input signals. Here we introduce Boolean Nested Effect Models (B-NEM). B-NEMs combine the use of downstream effects with the higher resolution of signalling pathway structures in Boolean Networks. RESULTS: We show that B-NEMs accurately reconstruct signal flows in simulated data. Using B-NEM we then resolve BCR signalling via PI3K and TAK1 kinases in BL2 lymphoma cell lines. AVAILABILITY AND IMPLEMENTATION: R code is available at https://github.com/MartinFXP/B-NEM (github). The BCR signalling dataset is available at the GEO database (http://www.ncbi.nlm.nih.gov/geo/) through accession number GSE68761. CONTACT: [email protected], [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Martin Pirkl, Elisabeth Hand, Dieter Kube, Rainer Spang
Bioinform.4
2015 A statistical approach to virtual cellular experiments: improved causal discovery using accumulation IDA (aIDA)
abstract
MOTIVATION: We address the following question: Does inhibition of the expression of a gene X in a cellular assay affect the expression of another gene Y? Rather than inhibiting gene X experimentally, we aim at answering this question computationally using as the only input observational gene expression data. Recently, a new statistical algorithm called Intervention calculus when the Directed acyclic graph is Absent (IDA), has been proposed for this problem. For several biological systems, IDA has been shown to outcompete regression-based methods with respect to the number of true positives versus the number of false positives for the top 5000 predicted effects. Further improvements in the performance of IDA have been realized by stability selection, a resampling method wrapped around IDA that enhances the discovery of true causal effects. Nevertheless, the rate of false positive and false negative predictions is still unsatisfactorily high. RESULTS: We introduce a new resampling approach for causal discovery called accumulation IDA (aIDA). We show that aIDA improves the performance of causal discoveries compared to existing variants of IDA on both simulated and real yeast data. The higher reliability of top causal effect predictions achieved by aIDA promises to increase the rate of success of wet lab intervention experiments for functional studies. AVAILABILITY AND IMPLEMENTATION: R code for aIDA is available in the Supplementary material. CONTACT: [email protected], [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Franziska Taruttis, Rainer Spang, Julia C. Engelmann
Bioinform.2
2015 Causal Modeling of Cancer-Stromal Communication Identifies PAPPA as a Novel Stroma-Secreted Factor Activating NFκB Signaling in Hepatocellular Carcinoma
abstract
Inter-cellular communication with stromal cells is vital for cancer cells. Molecules involved in the communication are potential drug targets. To identify them systematically, we applied a systems level analysis that combined reverse network engineering with causal effect estimation. Using only observational transcriptome profiles we searched for paracrine factors sending messages from activated hepatic stellate cells (HSC) to hepatocellular carcinoma (HCC) cells. We condensed these messages to predict ten proteins that, acting in concert, cause the majority of the gene expression changes observed in HCC cells. Among the 10 paracrine factors were both known and unknown cancer promoting stromal factors, the former including Placental Growth Factor (PGF) and Periostin (POSTN), while Pregnancy-Associated Plasma Protein A (PAPPA) was among the latter. Further support for the predicted effect of PAPPA on HCC cells came from both in vitro studies that showed PAPPA to contribute to the activation of NFκB signaling, and clinical data, which linked higher expression levels of PAPPA to advanced stage HCC. In summary, this study demonstrates the potential of causal modeling in combination with a condensation step borrowed from gene set analysis [Model-based Gene Set Analysis (MGSA)] in the identification of stromal signaling molecules influencing the cancer phenotype.
Julia C. Engelmann, Thomas Amann, Birgitta Ott-Rötzer, Margit Nützel, Yvonne Reinders, Jörg Reinders, Wolfgang E. Thasler, Theresa Kristl, Andreas Teufel, Christian G. Huber, Peter J. Oefner, Rainer Spang, Claus Hellerbrand
PLoS Comput. Biol.12
2014 Exact likelihood computation in Boolean networks with probabilistic time delays, and its application in signal network reconstruction
abstract
MOTIVATION: For biological pathways, it is common to measure a gene expression time series after various knockdowns of genes that are putatively involved in the process of interest. These interventional time-resolved data are most suitable for the elucidation of dynamic causal relationships in signaling networks. Even with this kind of data it is still a major and largely unsolved challenge to infer the topology and interaction logic of the underlying regulatory network. RESULTS: In this work, we present a novel model-based approach involving Boolean networks to reconstruct small to medium-sized regulatory networks. In particular, we solve the problem of exact likelihood computation in Boolean networks with probabilistic exponential time delays. Simulations demonstrate the high accuracy of our approach. We apply our method to data of Ivanova et al. (2006), where RNA interference knockdown experiments were used to build a network of the key regulatory genes governing mouse stem cell maintenance and differentiation. In contrast to previous analyses of that data set, our method can identify feedback loops and provides new insights into the interplay of some master regulators in embryonic stem cell development. AVAILABILITY AND IMPLEMENTATION: The algorithm is implemented in the statistical language R. Code and documentation are available at Bioinformatics online. CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary Materials are available at Bioinfomatics online.
Sebastian Dümcke, Johannes Bräuer, Benedict Anchang, Rainer Spang, Niko Beerenwinkel, Achim Tresch
Bioinform.4
2013 Considering Unknown Unknowns - Reconstruction of Non-confoundable Causal Relations in Biological Networks
Mohammad Javad Sadeh, Giusi Moffa, Rainer Spang
RECOMB3
2011 Estimating classification probabilities in high-dimensional diagnostic studies
abstract
MOTIVATION: Classification algorithms for high-dimensional biological data like gene expression profiles or metabolomic fingerprints are typically evaluated by the number of misclassifications across a test dataset. However, to judge the classification of a single case in the context of clinical diagnosis, we need to assess the uncertainties associated with that individual case rather than the average accuracy across many cases. Reliability of individual classifications can be expressed in terms of class probabilities. While classification algorithms are a well-developed area of research, the estimation of class probabilities is considerably less progressed in biology, with only a few classification algorithms that provide estimated class probabilities. RESULTS: We compared several probability estimators in the context of classification of metabolomics profiles. Evaluation criteria included sparseness biases, calibration of the estimator, the variance of the estimator and its performance in identifying highly reliable classifications. We observed that several of them display artifacts that compromise their use in practice. Classification probabilities based on a combination of local cross-validation error rates and monotone regression prove superior in metabolomic profiling. AVAILABILITY: The source code written in R is freely available at http://compdiag.uni-regensburg.de/software/probEstimation.shtml. CONTACT: [email protected].
Inka J. Appel, Wolfram Gronwald, Rainer Spang
Bioinform.3
2011 Genomic data integration using guided clustering
abstract
MOTIVATION: In biomedical research transcriptomic, proteomic or metabolomic profiles of patient samples are often combined with genomic profiles from experiments in cell lines or animal models. Integrating experimental data with patient data is still a challenging task due to the lack of tailored statistical tools. RESULTS: Here we introduce guided clustering, a new data integration strategy that combines experimental and clinical high-throughput data. Guided clustering identifies sets of genes that stand out in experimental data while at the same time display coherent expression in clinical data. We report on two potential applications: The integration of clinical microarray data with (i) genome-wide chromatin immunoprecipitation assays and (ii) with cell perturbation assays. Unlike other analysis strategies, guided clustering does not analyze the two datasets sequentially but instead in a single joint analysis. In a simulation study and in several biological applications, guided clustering performs favorably when compared with sequential analysis approaches. AVAILABILITY: Guided clustering is available as a R-package from http://compdiag.uni-regensburg.de/software/guidedClustering.shtml. Documented R code of all our analysis is included in the Supplementary Materials. All newly generated data are available at the GEO database (GSE29700). CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Matthias Maneck, Alexandra Schrader, Dieter Kube, Rainer Spang
Bioinform.4
2009 ReseqChip: Automated integration of multiple local context probe data from the MitoChip array in mitochondrial DNA sequence assembly
abstract
BACKGROUND: The Affymetrix MitoChip v2.0 is an oligonucleotide tiling array for the resequencing of the human mitochondrial (mt) genome. For each of 16,569 nucleotide positions of the mt genome it holds two sets of four 25-mer probes each that match the heavy and the light strand of a reference mt genome and vary only at their central position to interrogate all four possible alleles. In addition, the MitoChip v2.0 carries alternative local context probes to account for known mtDNA variants. These probes have been neglected in most studies due to the lack of software for their automated analysis. RESULTS: We provide ReseqChip, a free software that automates the process of resequencing mtDNA using multiple local context probes on the MitoChip v2.0. ReseqChip significantly improves base call rate and sequence accuracy. ReseqChip is available at http://code.open-bio.org/svnweb/index.cgi/bioperl/browse/bioperl-live/trunk/Bio/Microarray/Tools/. CONCLUSIONS: ReseqChip allows for the automated consolidation of base calls from alternative local mt genome context probes. It thereby improves the accuracy of resequencing, while reducing the number of non-called bases.
Marian Thieme, Claudio Lottaz, Harald Niederstätter, Walther Parson, Rainer Spang, Peter J. Oefner
BMC Bioinform.5
2008 Analyzing gene perturbation screens with nested effects models in R and bioconductor
abstract
UNLABELLED: Nested effects models (NEMs) are a class of probabilistic models introduced to analyze the effects of gene perturbation screens visible in high-dimensional phenotypes like microarrays or cell morphology. NEMs reverse engineer upstream/downstream relations of cellular signaling cascades. NEMs take as input a set of candidate pathway genes and phenotypic profiles of perturbing these genes. NEMs return a pathway structure explaining the observed perturbation effects. Here, we describe the package nem, an open-source software to efficiently infer NEMs from data. Our software implements several search algorithms for model fitting and is applicable to a wide range of different data types and representations. The methods we present summarize the current state-of-the-art in NEMs. AVAILABILITY: Our software is written in the R language and freely avail-able via the Bioconductor project at http://www.bioconductor.org.
Holger Fröhlich, Tim Beißbarth, Achim Tresch, Dennis Kostka, Juby Jacob, Rainer Spang, Florian Markowetz
Bioinform.6
2008 Detecting hierarchical structure in molecular characteristics of disease using transitive approximations of directed graphs
abstract
MOTIVATION: Molecular diagnostics aims at classifying diseases into clinically relevant sub-entities based on molecular characteristics. Typically, the entities are split into subgroups, which might contain several variants yielding a hierarchical model of the disease. Recent years have introduced a plethora of new molecular screening technologies to molecular diagnostics. As a result molecular profiles of patients became complex and the classification task more difficult. RESULTS: We present a novel tool for detecting hierarchical structure in binary datasets. We aim for identifying molecular characteristics, which are stochastically implying other characteristics. The final hierarchical structure is encoded in a directed transitive graph where nodes represent molecular characteristics and a directed edge from a node A to a node B denotes that almost all cases with characteristic B also display characteristic A. Naturally, these graphs need to be transitive. In the core of our modeling approach lies the problem of calculating good transitive approximations of given directed but not necessarily transitive graphs. By good transitive approximation we understand transitive graphs, which differ from the reference graph in only a small number of edges. It is known that the problem of finding optimal transitive approximation is NP-complete. Here we develop an efficient heuristic for generating good transitive approximations. We evaluate the computational efficiency of the algorithm in simulations, and demonstrate its use in the context of a large genome-wide study on mature aggressive lymphomas. AVAILABILITY: The software used in our analysis is freely available from http://compdiag.uni-regensburg.de/software/transApproxs.shtml.
Juby Jacob, Marcel Jentsch, Dennis Kostka, Stefan Bentink, Rainer Spang
Bioinform.5
2008 Microarray Based Diagnosis Profits from Better Documentation of Gene Expression Signatures
abstract
Microarray gene expression signatures hold great promise to improve diagnosis and prognosis of disease. However, current documentation standards of such signatures do not allow for an unambiguous application to study-external patients. This hinders independent evaluation, effectively delaying the use of signatures in clinical practice. Data from eight publicly available clinical microarray studies were analyzed and the consistency of study-internal with study-external diagnoses was evaluated. Study-external classifications were based on documented information only. Documenting a signature is conceptually different from reporting a list of genes. We show that even the exact quantitative specification of a classification rule alone does not define a signature unambiguously. We found that discrepancy between study-internal and study-external diagnoses can be as frequent as 30% (worst case) and 18% (median). By using the proposed documentation by value strategy, which documents quantitative preprocessing information, the median discrepancy was reduced to 1%. The process of evaluating microarray gene expression diagnostic signatures and bringing them to clinical practice can be substantially improved and made more reliable by better documentation of the signatures.
Dennis Kostka, Rainer Spang
PLoS Comput. Biol.2
2007 Annotation-based distance measures for patient subgroup discovery in clinical microarray studies
abstract
MOTIVATION: Clustering algorithms are widely used in the analysis of microarray data. In clinical studies, they are often applied to find groups of co-regulated genes. Clustering, however, can also stratify patients by similarity of their gene expression profiles, thereby defining novel disease entities based on molecular characteristics. Several distance-based cluster algorithms have been suggested, but little attention has been given to the distance measure between patients. Even with the Euclidean metric, including and excluding genes from the analysis leads to different distances between the same objects, and consequently different clustering results. RESULTS: We describe a new clustering algorithm, in which gene selection is used to derive biologically meaningful clusterings of samples by combining expression profiles and functional annotation data. According to gene annotations, candidate gene sets with specific functional characterizations are generated. Each set defines a different distance measure between patients, leading to different clusterings. These clusterings are filtered using a resampling-based significance measure. Significant clusterings are reported together with the underlying gene sets and their functional definition. CONCLUSIONS: Our method reports clusterings defined by biologically focused sets of genes. In annotation-driven clusterings, we have recovered clinically relevant patient subgroups through biologically plausible sets of genes as well as new subgroupings. We conjecture that our method has the potential to reveal so far unknown, clinically relevant classes of patients in an unsupervised manner. AVAILABILITY: We provide the R package adSplit as part of Bioconductor release 1.9 and on http://compdiag.molgen.mpg.de/software.
Claudio Lottaz, Joern Toedling, Rainer Spang
Bioinform.3
2007 Inferring cellular networks - a review
abstract
In this review we give an overview of computational and statistical methods to reconstruct cellular networks. Although this area of research is vast and fast developing, we show that most currently used methods can be organized by a few key concepts. The first part of the review deals with conditional independence models including Gaussian graphical models and Bayesian networks. The second part discusses probabilistic and graph-based methods for data from experimental interventions and perturbations.
Florian Markowetz, Rainer Spang
BMC Bioinform.2
2006 Permutation Filtering: A Novel Concept for Significance Analysis of Large-Scale Genomic Data
Stefanie Scheid, Rainer Spang
RECOMB2
2006 OrderedList - a bioconductor package for detecting similarity in ordered gene lists
abstract
UNLABELLED: OrderedList is a Bioconductor compliant package for meta-analysis based on ordered gene lists like those resulting from differential gene expression analysis. Our package quantifies the similarity between gene lists. The significance of the similarity score is estimated from random scores computed on perturbed data. OrderedList illustrates list similarity in intuitive plots and determines the score-driving genes for further analysis. AVAILABILITY: http://www.bioconductor.org CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Please visit our webpage on http://compdiag.molgen.mpg.de/software.
Claudio Lottaz, Xinan Yang, Stefanie Scheid, Rainer Spang
Bioinform.4
2006 Selecting normalization genes for small diagnostic microarrays
abstract
BACKGROUND: Normalization of gene expression microarrays carrying thousands of genes is based on assumptions that do not hold for diagnostic microarrays carrying only few genes. Thus, applying standard microarray normalization strategies to diagnostic microarrays causes new normalization problems. RESULTS: In this paper we point out the differences of normalizing large microarrays and small diagnostic microarrays. We suggest to include additional normalization genes on the small diagnostic microarrays and propose two strategies for selecting them from genomewide microarray studies. The first is a data driven univariate selection of normalization genes. The second is multivariate and based on finding a balanced diagnostic signature. Finally, we compare both methods to standard normalization protocols known from large microarrays. CONCLUSION: Not including additional genes for normalization on small microarrays leads to a loss of diagnostic information. Using house keeping genes from the literature for normalization fails to work for certain datasets. While a data driven selection of additional normalization genes works well, the best results were obtained using a balanced signature.
Jochen Jaeger, Rainer Spang
BMC Bioinform.2
2006 Automated in-silico detection of cell populations in flow cytometry readouts and its application to leukemia disease monitoring
abstract
BACKGROUND: Identification of minor cell populations, e.g. leukemic blasts within blood samples, has become increasingly important in therapeutic disease monitoring. Modern flow cytometers enable researchers to reliably measure six and more variables, describing cellular size, granularity and expression of cell-surface and intracellular proteins, for thousands of cells per second. Currently, analysis of cytometry readouts relies on visual inspection and manual gating of one- or two-dimensional projections of the data. This procedure, however, is labor-intensive and misses potential characteristic patterns in higher dimensions. RESULTS: Leukemic samples from patients with acute lymphoblastic leukemia at initial diagnosis and during induction therapy have been investigated by 4-color flow cytometry. We have utilized multivariate classification techniques, Support Vector Machines (SVM), to automate leukemic cell detection in cytometry. Classifiers were built on conventionally diagnosed training data. We assessed the detection accuracy on independent test data and analyzed marker expression of incongruently classified cells. SVM classification can recover manually gated leukemic cells with 99.78% sensitivity and 98.87% specificity. CONCLUSION: Multivariate classification techniques allow for automating cell population detection in cytometry readouts for diagnostic purposes. They potentially reduce time, costs and arbitrariness associated with these procedures. Due to their multivariate classification rules, they also allow for the reliable detection of small cell populations.
Joern Toedling, Peter Rhein, Richard Ratei, Leonid Karawajew, Rainer Spang
BMC Bioinform.5
2005 Molecular decomposition of complex clinical phenotypes using biologically structured analysis of microarray data
abstract
MOTIVATION: Today, the characterization of clinical phenotypes by gene-expression patterns is widely used in clinical research. If the investigated phenotype is complex from the molecular point of view, new challenges arise and these have not been addressed systematically. For instance, the same clinical phenotype can be caused by various molecular disorders, such that one observes different characteristic expression patterns in different patients. RESULTS: In this paper we describe a novel algorithm called Structured Analysis of Microarrays (StAM), which accounts for molecular heterogeneity of complex clinical phenotypes. Our algorithm goes beyond established methodology in several aspects: in addition to the expression data, it exploits functional annotations from the Gene Ontology database to build biologically focussed classifiers. These are used to uncover potential molecular disease subentities and associate them to biological processes without compromising overall prediction accuracy. AVAILABILITY: Bioconductor compliant R package SUPPLEMENTARY INFORMATION: Complete analyses are available at http://compdiag.molgen.mpg.de/supplements/lottaz05.
Claudio Lottaz, Rainer Spang
Bioinform.2
2005 Non-transcriptional pathway features reconstructed from secondary effects of RNA interference
abstract
MOTIVATION: Cellular signaling pathways, which are not modulated on a transcriptional level, cannot be directly deduced from expression profiling experiments. The situation changes, when external interventions such as RNA interference or gene knock-outs come into play. Even if the expression of the signaling genes is not changed, secondary effects in downstream genes shed light on the pathway, and allow partial reconstruction of its topology. RESULTS: We introduce an algorithm to infer non-transcriptional pathway features based on differential gene expression in silencing assays. We demonstrate the power of our algorithm in the controlled setting of simulation studies, and explain its practical use in the context of an RNA interference dataset investigating the response to microbial challenge in Drosophila melanogaster.
Florian Markowetz, Jacques Bloch, Rainer Spang
Bioinform.3
2005 twilight; a Bioconductor package for estimating the local false discovery rate
abstract
UNLABELLED: twilight is a Bioconductor compatible package for analysing the statistical significance of differentially expressed genes. It is based on the concept of the local false discovery rate (FDR), a generalization of the frequently used global FDR. twilight implements the heuristic search algorithm for estimating the local FDR introduced in our earlier work. In addition to the raw significance measures, it produces diagnostic plots, which provide insight into the extent of differential expression across genes. AVAILABILITY: http://www.bioconductor.org CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Please visit our software webpage on http://compdiag.molgen.mpg.de/software.
Stefanie Scheid, Rainer Spang
Bioinform.2
2005 stam - a Bioconductor compliant R package for structured analysis of microarray data
abstract
BACKGROUND: Genome wide microarray studies have the potential to unveil novel disease entities. Clinically homogeneous groups of patients can have diverse gene expression profiles. The definition of novel subclasses based on gene expression is a difficult problem not addressed systematically by currently available software tools. RESULTS: We present a computational tool for semi-supervised molecular disease entity detection. It automatically discovers molecular heterogeneities in phenotypically defined disease entities and suggests alternative molecular sub-entities of clinical phenotypes. This is done using both gene expression data and functional gene annotations. We provide stam, a Bioconductor compliant software package for the statistical programming environment R. We demonstrate that our tool detects gene expression patterns, which are characteristic for only a subset of patients from an established disease entity. We call such expression patterns molecular symptoms. Furthermore, stam finds novel sub-group stratifications of patients according to the absence or presence of molecular symptoms. CONCLUSION: Our software is easy to install and can be applied to a wide range of datasets. It provides the potential to reveal so far indistinguishable patient sub-groups of clinical relevance.
Claudio Lottaz, Rainer Spang
BMC Bioinform.2
2004 A Stochastic Downhill Search Algorithm for Estimating the Local False Discovery Rate
abstract
Screening for differential gene expression in microarray studies leads to difficult large-scale multiple testing problems. The local false discovery rate is a statistical concept for quantifying uncertainty in multiple testing. In this paper, we introduce a novel estimator for the local false discovery rate that is based on an algorithm which splits all genes into two groups, representing induced and noninduced genes, respectively. Starting from the full set of genes, we successively exclude genes until the gene-wise p-values of the remaining genes look like a typical sample from a uniform distribution. In comparison to other methods, our algorithm performs compatibly in detecting the shape of the local false discovery rate and has a smaller bias with respect to estimating the overall percentage of noninduced genes. Our algorithm is implemented in the Bioconductor compatible R package TWILIGHT version 1.0.1, which is available from http://compdiag.molgen.mpg.de/software or from the Bioconductor project at http://www.bioconductor.org.
Stefanie Scheid, Rainer Spang
IEEE ACM Trans. Comput. Biol. Bioinform.2
2001 Limits of homology detection by pairwise sequence comparison
abstract
Abstract Motivation: Noise in database searches resulting from random sequence similarities increases as the databases expand rapidly. The noise problems are not a technical shortcoming of the database search programs, but a logical consequence of the idea of homology searches. The effect can be observed in simulation experiments. Results: We have investigated noise levels in pairwise alignment based database searches. The noise levels of 38 releases of the SwissProt database, display perfect logarithmic growth with the total length of the databases. Clustering of real biological sequences reduces noise levels, but the effect is marginal. Contact: [email protected]; [email protected] 2 To whom correspondence should be addressed. Pressent address: Duke University, Institute of Statistics and Decision Sciences, Box 90251 Duke University, Durham, NC 27708-0251, USA.
Rainer Spang, Martin Vingron
Bioinform.1
2000 Sequence Database Search Using Jumping Alignments
Rainer Spang, Marc Rehmsmeier, Jens Stoye
ISMB1
1998 Statistics of large-scale sequence searching
abstract
MOTIVATION: Database search programs such as FASTA, BLAST or a rigorous Smith-Waterman algorithm produce lists of database entries, which are assumed to be related to the query. The computation of statistical significance of similarity scores is well established for single pairs of sequences and using purely random models. However, the multi-trial context of a database search poses new problems. The credibility of a certain score obtained in a database search decreases with the amount of data that is compared. To improve p-value computation for database search experiments, statistical properties of the databases, such as the distribution of sequence length and effects induced by frequently repeated sequence patterns, need to be taken into account. RESULTS: We investigated the SWISS-PROT protein database Release 31.0 running extensive simulations of database searches. A discrepancy is observed between the theoretical predictions and the empirical distribution. To correct for this, we evaluate the statistical significance of scores in the context of a database search by a contrasting semi-random model. This model enhances purely random models by one additional parameter reflecting individual statistical properties of real databases. We call this parameter the effective size of the database. CONTACT: [email protected];m.vingron@dkfz-hei del berg.de
Rainer Spang, Martin Vingron
Bioinform.1